Compare commits

...
Author SHA1 Message Date
0c16ec91b0 [feat]: connect Dreamverse creation settings to generation
Apply the backend-wiring changes beyond the UI uplift to
ds8/dreamversev2-dev for review and refactoring.

Source PR: hao-ai-lab/FastVideo#1854
Source range: 90d739a91892302edf37c4b23f807f402c83071d..8c5ee9b51cc3b75f4eba9cb3904fe08578c9dc9e
The 28-file patch is identical to that source range.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-15 17:16:15 -07:00
e57543b79d [feat]: Dreamverse creation studio UI uplift (#1853)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-15 16:45:08 -07:00
Raghav K 9b0e57fe4b [ci] Make Dreamverse provider race test deterministic (#1729) 2026-09-15 14:13:56 -07:00
William Lin 0100218594 [feat] Support the FastH3 8-Step V2 checkpoint: checkpoint-defined shifts, explicit DMD schedule, new example (#1852) 2026-09-15 14:13:45 -07:00
li-lizhe 39718cd54d [bugfix] fix(cosmos): make AdaLayerNorm autocast device-agnostic (#1818) 2026-09-15 07:44:51 -07:00
IshanandSolitaryThinker 37d06a832f [feat] Add fastvideo serve configs for Wan CUDA models (#1801)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 18:23:35 -07:00
sudhirpol522 8839ba8d4d [bugfix] Add OpenAI-compatible image generation endpoint (#1840) 2026-09-14 17:55:56 -07:00
sudhirpol522andSolitaryThinker 61b91220c0 [bugfix] Validate image response format before generation (#1841)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 17:01:40 -07:00
Lele 316f3876c2 [bugfix]: allow MiniMax H3 frame padding at the 15-second limit
Accept the causal-VAE-aligned 362-frame bucket (15.083 s) for 15-second H3 requests.

- Hoist MINIMAX_H3_MIN/MAX_ALIGNED_FRAMES next to align_num_frames and reuse them in the CUDA stage, the MLX runtime, and the LoRA example so the bound is consistent across entry points.
- State the accepted frame range in the error messages.
- Cover seconds="15" -> 362 at OpenAI admission and the MLX resolve_geometry bound.
- Update the Spark/openai cookbook frame caps from 345 to 362.

Known gaps: no GPU-lane coverage for the 362 bucket (SSIM runs 124 frames; the golden gate is a fixed-geometry fingerprint), and explicit num_frames=360 still hits the pre-existing grid gate.
2026-09-14 16:54:41 -07:00
Yaegaki1Erika 614b59543c [bugfix]: respect serialized tensor dtype in parquet dataloader (#1843) 2026-09-14 16:34:25 -07:00
William Lin 1c14afd559 [docs]: refresh AGENTS.md maps and add Wan SP/I2V and CI test notes (#1845) 2026-09-14 16:06:53 -07:00
Yaegaki1Erika 9a3c45779c [bugfix]: fix Wan I2V DMD conditioning under sequence parallelism (#1844)
The pipeline pre-sharded only the image conditioning along the temporal axis while the noise input stayed full-length, so concatenation could not match for sp_world_size > 1. WanTransformer3DModel shards the flattened token sequence after patch embedding, so pass full-length mask+latent conditioning and let the transformer shard.

Also reads temporal_compression_ratio from the VAE config, simplifies the conditioning mask, concatenates in the transformer's (bs, c, t, h, w) layout, and adds unit coverage.

Co-authored-by: Yaegaki1Erika <70182590+Yaegaki1Erika@users.noreply.github.com>
2026-09-14 15:58:13 -07:00
Kyle Hu bfc9c01797 [feat]: convert the MiniMax H3 text encoder to NVFP4 (#1838) 2026-09-12 18:48:07 -07:00
Kyle Hu 3a3ad3d209 [feat]: NVFP4 text encoder for MiniMax H3 (#1837) 2026-09-12 18:31:59 -07:00
lpc0220andClaude Fable 5.1 aef4e9b3b1 [kernel] VSA kernel: one sm_100a / sm_103a image per listed arch; un-gate backward on sm_103a (#1833)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:49:36 -07:00
a943220c11 [bugfix]: drop dead h3_sequential_load from Spark FastH3 presets (#1831)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-08 12:38:06 -07:00
William Lin 556ac7088e [refactor] Simplify Wan sampling and tests (#1825) 2026-09-07 15:57:19 -07:00
Junda Su e7456f1b75 Add H3 support into Dreamverse (#1800) 2026-09-07 13:09:44 -07:00
lpc0220 c993d7393e [kernel] sm_100a CUDA backward for VSA block-sparse attention (blk64) (#1819) 2026-09-06 21:00:18 -07:00
William Lin 7f83164233 [refactor] Move Wan VAE into the Wan package (#1824) 2026-09-05 18:43:55 -07:00
Junda Su 4e52f47d1e [feat] add MXFP8 support on H3 (#1796) 2026-09-05 17:08:58 -07:00
William Lin e19913f6e9 [refactor] Group Wan transformer and config (#1823) 2026-09-05 17:05:32 -07:00
William Lin 2413a57651 Disable old SSIM models (#1820) 2026-09-04 21:20:48 -07:00
SYLAR 7bb76b5ec9 [feat] Add native SM103a VSA support (#1812)
Signed-off-by: lishunyang12 <lishunyang12@163.com>
2026-09-03 15:00:28 -07:00
Aryan KumarandAryan Kumar 0bd19a976b [docs]: add one-Spark FastH3 cookbook runtime with a device-count row (#1811)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-02 10:05:59 -07:00
William Lin 40b93784d2 [docs]: add cookbook link to README (#1810) 2026-09-01 12:42:19 -07:00
Aryan Kumar 33d3478bad [docs] Announce local FastH3 support (#1809) 2026-09-01 12:28:04 -07:00
3d8ac9d14b [feat]: collapse cookbook recipe pages into an accordion layout, add … (#1805)
Co-authored-by: Vaish, Ishan <isvaish@UCSD.EDU>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 01:51:51 -07:00
aaef49bfc6 [feat]: run FastH3 across two DGX Sparks with Ray sequence parallel (#1803)
Co-authored-by: Kyle <shh075@ucsd.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-01 01:37:23 -07:00
Aryan KumarandAryan Kumar cf6a00b9be [feat]: add opt-in CUDA TAEH3 preview decode for FastH3 (#1795)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 23:34:56 -07:00
Shahrad ZomorrodiandShahrad Zomorrodi 1ae39562dd [bugfix] Write generated documentation as UTF-8 (#1797)
Co-authored-by: Shahrad Zomorrodi <264690209+shahradzomorrodi@users.noreply.github.com>
2026-08-31 22:45:03 -07:00
Aryan KumarandAryan Kumar 26064193e2 [feat] Add an H3 server cookbook and prompt playground (#1798)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 22:14:33 -07:00
William Lin 8446fc003e [docs]: add FastH3 Preview v1 news links (#1804) 2026-08-31 22:12:02 -07:00
Aryan KumarandAryan Kumar a28f2bab4b [feat] Add an optional MLX TAEH3 preview decoder (#1794)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 05:08:06 -07:00
Aryan KumarandAryan Kumar f82d8be4bf [perf] Sequential MiniMax H3 start with GPU-direct DiT load (#1793)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:20:21 -07:00
Aryan KumarandAryan Kumar 8e1775183e [perf] Speed up exact MiniMax H3 MLX inference (#1792)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:14:55 -07:00
Aryan KumarandAryan Kumar 620bc36dc4 [docs]: Cookbook catalog improvements (#1790)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 17:50:57 -07:00
Suhaan Khurana 29ff16ec96 [feat] Add MiniMax H3 MLX spatial fast mode (#1789) 2026-08-30 17:47:02 -07:00
Aryan KumarandAryan Kumar 8f9d76a80d [perf]: dispatch wide-M affine H3 MLX linears through dequant plus dense GEMM (#1788)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 14:42:01 -07:00
a4d9a75e2c [perf] Add MiniMax H3 MLX VSA and SIMD attention (#1776)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-30 13:44:54 -07:00
KyleNeverGivesUp b2db0c0a13 [ci]: seed stable GB10 grad-norm references (#1756) 2026-08-30 03:09:18 -07:00
Kevin Lin 6aa7d8a278 [misc] FastVideo Studio UI Additions (H3 Ref2V support) (#1783) 2026-08-30 03:06:47 -07:00
Aryan KumarandAryan Kumar ccc9014430 [docs] Add model-family inference cookbook (#1787)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 02:26:10 -07:00
William Lin a159b63c67 [bugfix] Harden OpenAI serving after post-merge review (#1782) 2026-08-28 22:29:24 -07:00
ac48bb3cd1 [feat] Add MiniMax H3 MLX T2VA inference (#1770)
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
2026-08-28 12:54:33 -07:00
William Lin 3987b9ddcd [feat] Align multimodal OpenAI serving APIs (#1781) 2026-08-28 10:09:58 -07:00
William Lin c7da2f5d60 [chore]: release v0.2.1 (#1778) 2026-08-28 02:03:18 -07:00
William Lin 39ae1decc0 [misc] pin fastvideo-kernel to exact 0.3.5 (#1777) 2026-08-28 02:02:27 -07:00
William Lin 1aed667377 [chore] release fastvideo-kernel 0.3.5 (#1775) 2026-08-27 23:18:59 -07:00
William Lin c1612ff397 [bugfix]: pin fastvideo-kernel to Torch 2.12.0 (#1774) 2026-08-27 23:13:54 -07:00
William Linandshaoxiongduan a534ba20a0 [feat] Add MiniMax H3 LoRA inference and preview launchers (#1771)
Co-authored-by: shaoxiongduan <shaoxiongduan@gmail.com>
2026-08-27 14:43:39 -07:00
KyleNeverGivesUp e9bbaca07d [perf] Disable every offload path on unified memory, unblocking MiniMax H3 generation on one GB10 (#1715) 2026-08-26 15:32:56 -07:00
KyleNeverGivesUp 9bfa585448 [perf]: stop holding the whole checkpoint during DiT load, unblocking MiniMax H3 on one GB10 (#1714) 2026-08-26 15:23:51 -07:00
Raghav K b2062556a9 [perf] VSA Triton: widen the autotune num_stages range (the optimum was outside it) (#1706) 2026-08-26 12:30:56 -07:00
KyleNeverGivesUp c9c5585758 [perf]: MiniMax H3 on GB10 - skip text encoder CPU offload on unified memory (5m49s to 30ms) (#1710) 2026-08-26 12:03:14 -07:00
William Lin 9212f4f218 [ci] make GPU validation change-aware (#1747) 2026-08-25 21:26:25 -07:00
Aryan KumarandAryan Kumar 6388db815b [bugfix] FastMetal-QAD MLX support: refuse CUDA QAD trees, use packed mlx_dit config, stream loads (#1736) (#1758)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-25 14:51:40 -07:00
lpc0220 7a4285189f [kernel] Route block-sparse VSA to the sm_100a forward behind FASTVIDEO_VSA_SM100A (opt-in) (#1754) 2026-08-24 15:26:56 -07:00
William Lin a837fe841a [docs] Update FastH3 README (#1749) 2026-08-23 05:05:58 -07:00
William Lin f9e3680f11 [perf] Align FastH3 optimized inference profile (#1748) 2026-08-23 02:01:03 -07:00
William Lin 98f761ec45 [bugfix] validation: inherit the trained denoising ladder (#1738) 2026-08-22 23:09:26 -07:00
Shao Duan c041318f2c [perf] Add fused NVLink all-to-all for Ulysses (#1740) 2026-08-22 23:06:24 -07:00
William Lin 604e0205a4 [perf] Keep odd MiniMax-H3 VSA tiles on sm100a (#1745) 2026-08-22 18:29:30 -07:00
William Lin 13213395b4 [perf] Parallelize MiniMax-H3 VAE over sequence ranks (#1744) 2026-08-22 18:05:31 -07:00
William Lin 46afee5998 [bugfix] Classify MiniMax-H3 inference controls in schema inventory (#1743) 2026-08-22 17:17:02 -07:00
William Lin c488fa1211 [perf] Add opt-in packed-varlen FA4 for MiniMax-H3 (#1742) 2026-08-22 17:16:47 -07:00
William Lin d3cff517cd [perf] Add opt-in regional fullgraph compile for DiT inference (#1741) 2026-08-22 12:14:39 -07:00
Junda Su 2f3d407406 [perf] Optimize MiniMax H3 VAE decoding (#1734) 2026-08-21 14:57:32 -07:00
Kaiqin Kong bcffa4026e [perf] Optimize MiniMax-H3 text encoder memory (#1732) 2026-08-21 14:57:06 -07:00
William Lin 6d6a10be7a [feat] FastVideo-Minimax-FastH3-Preview few-step example + 64-token-tile VSA-H3 inference path (#1731) 2026-08-21 12:40:09 -05:00
Kaiqin Kong 73dd105f3d [perf] Add opt-in MiniMax-H3 Sol-Engine fusions (#1735) 2026-08-21 12:39:28 -05:00
559 changed files with 61331 additions and 8263 deletions
+99
View File
@@ -0,0 +1,99 @@
---
name: ci-runner
description: Work on FastVideo's Slurm-only, change-aware GPU CI lanes, static Buildkite graph, trusted ci-runner policy, lane scripts, and GB200 validation.
---
# Slinky Slurm CI lanes
FastVideo's `ci-runner` Buildkite queue is the control plane for all active
GPU CI. A private host-owned dispatcher leases GPUs from the Slinky Slurm tray
and runs the immutable PR SHA inside an isolated Enroot container. Buildkite
pipeline upload and Slurm submission occur on the login plane; every test
payload executes on Slurm compute.
The files under `fastvideo/tests/modal/` and `.buildkite/scripts/pr_test.sh`
are dormant rollback code. Never add an active Buildkite or slash-command
route to them. `pr_test.sh` must continue to reject Buildkite invocations.
The private operator bundle is deliberately outside this repository because
it contains site paths and credentials. See
`docs/contributing/ci_architecture.md`; this skill covers the repository half
and the coordination contract with that bundle.
## Invariants
- `.buildkite/pipeline.yml` contains exactly one static step for every active
GPU lane. Each step pins a unique key and label, a 90-minute timeout, the
trusted `/opt/fastvideo-ci-runner/run-ci` command (`run-unit` is the one
compatibility wrapper), step-level internal `TEST_TYPE`, and
`queue: "ci-runner"`.
- Active CI contains no `pr_test.sh` command, Modal invocation, default queue,
Buildkite plugin, `soft_fail`, or job-controlled artifact glob.
- The six Fastcheck lanes use `:microscope:` labels. Full-Suite-only lanes use
`:test_tube:` or `:bar_chart:` so direct reruns update the right aggregate.
- SSIM and vanilla training request all four GPUs. Keep both in the
`fastvideo/slinky/whole-tray` Buildkite concurrency group with a limit of one
so the second job does not consume an agent or command timeout while waiting
for the same tray.
- `/test full` schedules all twenty lanes. `/merge`, `ready`, and new pushes to
ready PRs use the trusted base-branch planner in
`.github/scripts/plan_merge_ci.py`: automatic Fastcheck remains the universal
six-lane baseline, and the merge build adds only path-relevant integration
lanes. Unknown source/build paths fail closed to all fourteen additive lanes.
The trusted uploader still normalizes and validates the complete static graph
before Buildkite evaluates its plan conditions.
- Focused merge builds may pass allowlisted golden-gate and SSIM test basenames.
The private host validates the lane plan and basenames before staging them,
and the in-container scripts validate them again. Direct `/test ssim`,
explicit `/test full`, and the weekly main-branch schedule run the complete
SSIM matrix.
- The trusted uploader serves exactly three entry pipelines:
`pr-fastcheck` for automatic PR builds, `ci` for slash-command/ready-label
API builds, and `fastvideo-performance-lane` for the weekly schedule. Keep
incoming GitHub webhook processing disabled on `ci` so it cannot duplicate
`pr-fastcheck` on every PR update.
- Test payloads live in `.buildkite/scripts/unit_test.sh` or executable
`.buildkite/scripts/lanes/<lane>.sh`. Backend policy (GPU count, extras,
secrets, kernel build, artifacts) stays in the agent-owned lane table.
- Tests must preserve an inherited `MASTER_PORT`. Packed containers share the
tray network namespace, so the private runner assigns a distinct port range
per GPU lease and the SSIM scheduler assigns task offsets within its range.
- The ARM64 runner image includes the pinned FA4 CuTe overlay validated on
GB200. Keep SSIM at `FASTVIDEO_FA4=1` because its references were seeded with
FA4; keep lanes with FA2 baselines at `FASTVIDEO_FA4=0`. A runner image change
must revalidate both the FA4 import and an actual GB200 forward kernel.
- `fastvideo/tests/ssim/ci_runner.py` is the active four-GPU SSIM scheduler.
New SSIM files are discovered through `REQUIRED_GPUS` and
`*_MODEL_TO_PARAMS`; do not wire them through the dormant Modal scheduler.
- The host policy fail-closes unknown tuples. A repository-side lane change is
inert until the operator updates the private lane table and uploader policy
in the same rollout.
## Adding or changing a lane
1. Read the closest `AGENTS.md` and the domain-specific testing guide.
2. Add or update the executable lane payload under `.buildkite/scripts/`.
Keep it deterministic and free of host-specific paths or credential fetches.
3. Add the static pipeline step and canonical `/test <name>` mapping. Keep the
`<name>-ci` alias only when compatibility requires it.
4. Add its source/test path ownership to `.github/scripts/plan_merge_ci.py`.
Prefer the narrowest correctness-preserving lane set; leave unknown paths
fail-closed. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
`test_merge_ci_plan.py`, and focused CPU-only scheduler/policy tests.
5. Coordinate the private lane row: GPU count (1-4), wall time, script, scope
pairs, step key, command, HF cache/token, tracking mode, extras, attention
backend policy, kernel policy, and artifact relay. Active training lanes
keep W&B offline and do not stage a W&B credential.
6. Update the trusted pipeline-uploader schema. A mismatch must reject the
pipeline rather than silently skip a lane.
7. Run `pre-commit run --files <changed paths>`, the planner's representative
diff matrix, contract tests, private driver tests, and a real GB200 canary.
Multi-GPU, hardware-reference, training, performance, and SSIM changes need
their own target-hardware evidence.
## Rollback
Rollback the Slurm routing/configuration change or pause the `ci-runner` queue.
Do not silently reactivate Modal. A manual Modal experiment requires the
explicit local opt-in documented in `ci_architecture.md`; returning it to
production CI needs a separate reviewed decision.
@@ -1,6 +1,6 @@
---
name: reseed-ssim-references
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted model subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
@@ -13,7 +13,7 @@ on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
3. Regenerates through the manual legacy Modal L40S maintenance path.
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
@@ -51,8 +51,9 @@ harder to recover from than failing closed.
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Modal GPU: **L40S**. This is a manual reference-maintenance target, not the
active Slurm CI compute path; changing the SKU also changes the historical
`L40S_reference_videos` contract.
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
+4 -2
View File
@@ -35,7 +35,8 @@ The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an artefact without crashing (the missing-reference assertion at
the end is expected). The skill does not re-test locally; it goes straight
to Modal L40S (which is what CI uses).
to the manual legacy Modal L40S reference-maintenance target. Active CI runs
on the Slinky Slurm cluster and only consumes the resulting references.
## When to use
@@ -61,7 +62,8 @@ Prompt the user for it if they didn't supply it.
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Modal maintenance GPU: **L40S** (hardcoded in
`fastvideo/tests/modal/ssim_test.py`; this is not the active CI compute path).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
seeded by this skill.
+461 -528
View File
File diff suppressed because it is too large Load Diff
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the OpenAI-compatible API lane.
set -euo pipefail
exec pytest ./fastvideo/tests/entrypoints/test_openai_api_integration.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the distillation-DMD lane.
set -euo pipefail
exec pytest ./fastvideo/tests/training/distill/test_distill_dmd.py -vs
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# DreamVerse needs a GPU for import-time device resolution, but it does not
# build or exercise fastvideo-kernel. A checksummed Node archive is installed
# in the disposable Slurm container because the shared CI image is
# Python/CUDA focused.
set -euo pipefail
node_version=v22.23.2
case $(uname -m) in
aarch64 | arm64)
node_arch=arm64
node_archive_sha256=013b59cfd2819703a6f4a14ab891fc46fc2a4e3f5bcd92de3fb4929b43e35b30
;;
x86_64 | amd64)
node_arch=x64
node_archive_sha256=b294a556e639d64338823920e5866c21c02741742d2e1529ee1a225c1ec9252a
;;
*)
echo "Unsupported architecture for DreamVerse Node runtime: $(uname -m)" >&2
exit 2
;;
esac
node_archive="node-${node_version}-linux-${node_arch}.tar.gz"
node_runtime_root=$(mktemp -d -t fastvideo-node.XXXXXX)
node_archive_path="${node_runtime_root}/${node_archive}"
node_install_dir="${node_runtime_root}/${node_archive%.tar.gz}"
curl --proto '=https' --tlsv1.2 --retry 5 --retry-all-errors \
--location --fail --silent --show-error \
"https://nodejs.org/dist/${node_version}/${node_archive}" \
--output "$node_archive_path"
printf '%s %s\n' "$node_archive_sha256" "$node_archive_path" | sha256sum --check --status
tar -xzf "$node_archive_path" -C "$node_runtime_root"
export PATH="${node_install_dir}/bin:${PATH}"
node --version
npm --version
export PYTHONPATH="$(pwd)/apps/dreamverse${PYTHONPATH:+:$PYTHONPATH}"
pytest apps/dreamverse/dreamverse/tests -q
cd apps/dreamverse/web
npm ci
npm run typecheck
npm test
machine_arch=$(uname -m)
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
npx playwright install --with-deps chromium firefox
else
npx playwright install --with-deps chromium webkit firefox
fi
master_port=${MASTER_PORT:-7959}
BACKEND_PORT=${BACKEND_PORT:-$((master_port + 50))}
python -m uvicorn dreamverse.mock_server:app --host 127.0.0.1 --port "$BACKEND_PORT" &
mock_server_pid=$!
cleanup() {
kill "$mock_server_pid" 2>/dev/null || true
wait "$mock_server_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
for _ in {1..30}; do
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz" && break
sleep 1
done
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz"
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
# Playwright WebKit traps before opening a page on Linux ARM64, and its
# bundled Chromium lacks the H.264/AAC codecs used by the fMP4 assertions.
# Firefox covers every flow, including streaming. Chromium and its mobile
# profile still cover all codec-independent UI behavior on GB200.
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- --project=firefox
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=mobile-chromium \
--grep-invert='streams, plays, and surfaces a downloadable clip|starts a new project and switches back to the prior session|saved projects persist across a page reload'
else
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=webkit \
--project=firefox \
--project=mobile-safari \
--project=mobile-chromium
fi
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the encoder lane.
set -euo pipefail
exec pytest ./fastvideo/tests/encoders -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the evaluation lane.
set -euo pipefail
exec pytest ./fastvideo/tests/eval -vs
+35
View File
@@ -0,0 +1,35 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the golden-gate lane. Environment (HF_HOME
# and authentication) is the runner's responsibility.
set -euo pipefail
golden_root=./fastvideo/tests/golden_gate
selected=${FASTVIDEO_GOLDEN_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_GOLDEN_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" = all ]; then
exec pytest "$golden_root" -xvs
fi
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_GOLDEN_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a golden_files <<< "$selected"
golden_paths=()
for golden_file in "${golden_files[@]}"; do
golden_path="$golden_root/$golden_file"
[ -f "$golden_path" ] || {
echo "Selected golden test does not exist: $golden_file" >&2
exit 2
}
golden_paths+=("$golden_path")
done
exec pytest "${golden_paths[@]}" -xvs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-inference lane.
set -euo pipefail
exec pytest ./fastvideo/tests/inference/lora/test_lora_inference_similarity.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VMoBA-inference lane.
set -euo pipefail
exec python fastvideo/tests/inference/vmoba/test_vmoba_inference.py
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the custom-kernel lane.
set -euo pipefail
exec pytest fastvideo-kernel/tests/ -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-extraction lane.
set -euo pipefail
exec pytest ./fastvideo/tests/lora_extraction/test_lora_extraction.py -vs
+52
View File
@@ -0,0 +1,52 @@
#!/usr/bin/env bash
# Canonical Slurm performance lane. Reports are written outside the checkout
# so the trusted host driver can upload them after untrusted code exits.
set -uo pipefail
export PERFORMANCE_TRACKING_ROOT=/tmp/perf-tracking
export PERF_REPORTS_DIR=/workspace/artifacts/performance
mkdir -p "$PERF_REPORTS_DIR"
if [[ ${BUILDKITE_PULL_REQUEST:-false} =~ ^[1-9][0-9]*$ ]]; then
export PERF_RUN_SOURCE=pr
export PERF_UPLOAD_POLICY=pass
elif [ "${BUILDKITE_BRANCH:-}" = main ] \
&& { [ "${BUILDKITE_SOURCE:-}" = schedule ] || [ "${TEST_SCOPE:-}" = full ]; }; then
export PERF_RUN_SOURCE=scheduled_main
export PERF_UPLOAD_POLICY=always
elif [ "${TEST_SCOPE:-}" = direct ]; then
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=pass
else
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=never
fi
nvidia-smi \
--query-gpu=index,timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu \
--format=csv -l 10 > "$PERF_REPORTS_DIR/gpu_telemetry.csv" 2>/dev/null &
telemetry_pid=$!
cleanup() {
kill "$telemetry_pid" 2>/dev/null || true
wait "$telemetry_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
pytest ./fastvideo/tests/performance -vs
pytest_rc=$?
compare_rc=0
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
PERF_PYTEST_RC=$pytest_rc python ./fastvideo/tests/performance/compare_baseline.py
compare_rc=$?
fi
python ./fastvideo/tests/performance/dashboard.py || true
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
final_rc=$pytest_rc
if [ "$final_rc" -eq 0 ]; then
final_rc=$compare_rc
fi
exit "$final_rc"
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the self-forcing lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/self-forcing/test_self_forcing.py -vs
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# Canonical four-GPU SSIM lane for the Slinky Slurm worker.
set -euo pipefail
args=()
if [ "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-0}" = 1 ]; then
args+=(--bootstrap-mode)
fi
selected=${FASTVIDEO_SSIM_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_SSIM_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" != all ]; then
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_SSIM_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a ssim_files <<< "$selected"
for ssim_file in "${ssim_files[@]}"; do
args+=(--test-file "$ssim_file")
done
fi
# MoGe's utils3d dependency builds glcontext from source on ARM64. The current
# runner image predates the baked-in X11 headers below, so keep this guarded
# bootstrap until every deployed image digest contains libx11-dev.
if [ ! -f /usr/include/X11/Xlib.h ]; then
apt-get -o Acquire::Retries=5 update
apt-get -o Acquire::Retries=5 install -y --no-install-recommends libx11-dev
rm -rf /var/lib/apt/lists/*
fi
uv pip install git+https://github.com/microsoft/MoGe.git
uv pip install k_diffusion einops_exts alias_free_torch torchsde
exec python fastvideo/tests/ssim/ci_runner.py "${args[@]}"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the modular training-framework lane.
set -euo pipefail
exec pytest ./fastvideo/tests/train/models ./fastvideo/tests/train/methods -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy vanilla-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/Vanilla -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy LoRA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/lora/test_lora_training.py -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy VSA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/VSA -srP
+9
View File
@@ -0,0 +1,9 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the transformer lane.
set -euo pipefail
# The existing block reference records an absent FASTVIDEO_FA4 (FA2). Keep
# that reference identity; the component lane also selects FA2 explicitly.
env -u FASTVIDEO_FA4 pytest ./fastvideo/tests/golden_gate/test_wan_t2v.py -xvs
pytest ./fastvideo/tests/golden_gate/test_wan_causal.py -xvs
exec pytest ./fastvideo/tests/transformers -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VAE lane.
set -euo pipefail
pytest ./fastvideo/tests/golden_gate/test_wan_vae.py -xvs
exec pytest ./fastvideo/tests/vaes -vs
+13
View File
@@ -1,6 +1,19 @@
#!/bin/bash
set -uo pipefail
# DORMANT ROLLBACK ONLY. Active CI is Slurm-only and pipeline.yml never calls
# this launcher. Refuse every Buildkite invocation even if a stale step or
# operator typo reaches this file; local rollback experiments require an
# explicit opt-in.
if [ -n "${BUILDKITE:-}" ]; then
echo "Legacy Modal CI is disabled; use the Slinky Slurm runner." >&2
exit 2
fi
if [ "${FASTVIDEO_ENABLE_LEGACY_MODAL_CI:-0}" != 1 ]; then
echo "Legacy Modal CI is dormant. Set FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1 only for a manual rollback test." >&2
exit 2
fi
log() {
echo "[$(date '+%Y-%m-%d %H:%M:%S')] $1"
}
+25
View File
@@ -0,0 +1,25 @@
#!/usr/bin/env bash
set -euo pipefail
exec pytest \
./fastvideo/tests/api/ \
./fastvideo/tests/contract/ \
./fastvideo/tests/dataset/ \
./fastvideo/tests/workflow/ \
./fastvideo/tests/entrypoints/ \
./fastvideo/tests/loader/ \
./fastvideo/tests/pipelines/ \
./fastvideo/tests/platforms/ \
./fastvideo/tests/train/ \
./fastvideo/tests/stages/ \
./fastvideo/tests/ops/ \
./fastvideo/tests/worker/ \
./fastvideo/tests/training/test_trackers.py \
./fastvideo/tests/attention/test_sdpa_metadata_mask_contract.py \
./fastvideo/tests/modal/test_kernel_build_cache.py \
./fastvideo/tests/modal/test_pr_test.py \
./fastvideo/tests/modal/test_ssim_test.py \
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
--ignore=./fastvideo/tests/train/models \
--ignore=./fastvideo/tests/train/methods \
-vs
+2 -2
View File
@@ -8,10 +8,10 @@ PR TITLE: Must start with a type tag, e.g.:
MERGE WORKFLOW:
1. Ensure pre-commit passes and you have at least 1 approval
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
3. A path-aware merge gate runs only relevant integration tests → auto-merge on success
ON-DEMAND TESTING (write access required):
/test full — Full Test Suite /test ssim — SSIM regression
/test full — Explicit all-lane run /test ssim — Full SSIM regression
/test training — Training pipeline /test encoder — Encoder tests
/test transformer — Transformer tests /test vae — VAE tests
/test kernel — CUDA kernel tests /test unit — Unit tests
+10 -10
View File
@@ -1,14 +1,14 @@
#!/usr/bin/env bash
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
# Gate the path-aware Buildkite merge plan on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
# already doomed.
# 'ready' label cannot burn path-selected GPU lanes on a head that a cheap
# check has already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no full suite; the next push re-arms via the 'synchronize' trigger)
# no merge gate; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
@@ -29,7 +29,7 @@ set -euo pipefail
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the full suite
# Workflow-level `name:` values that must be green before the merge gate
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
@@ -56,7 +56,7 @@ recheck_ready_label() {
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
"NOT triggering the Buildkite merge gate. Re-add the label to re-arm."
exit 1
fi
else
@@ -84,7 +84,7 @@ while true; do
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
"NOT triggering the Buildkite merge gate. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
@@ -97,7 +97,7 @@ while true; do
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — full suite may proceed."
echo "All watched cheap checks are green — merge gate may proceed."
exit 0
fi
case "$missing" in
@@ -119,14 +119,14 @@ while true; do
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the merge gate WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the merge gate anyway."
exit 0
fi
sleep "$POLL_SECS"
+582
View File
@@ -0,0 +1,582 @@
#!/usr/bin/env python3
"""Select the additive GPU integration lanes needed by a PR diff.
Fastcheck is the universal six-lane baseline and is intentionally not repeated
here. This planner selects only the more expensive merge-gate lanes. Unknown
source/build paths fail closed to the complete integration set, while explicit
documentation and repository-metadata paths require no additional GPU work.
"""
from __future__ import annotations
import argparse
import fnmatch
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import TextIO
MERGE_LANES = (
"golden-gate",
"ssim",
"lora-inference",
"lora-extraction",
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
"inference-vmoba",
"performance",
"api-server",
"train-framework",
"eval",
)
LANE_SCRIPT_TO_KEY = {
"api_server.sh": "api-server",
"distillation_dmd.sh": "distillation",
"eval.sh": "eval",
"golden_gate.sh": "golden-gate",
"inference_lora.sh": "lora-inference",
"inference_vmoba.sh": "inference-vmoba",
"lora_extraction.sh": "lora-extraction",
"performance.sh": "performance",
"self_forcing.sh": "self-forcing",
"ssim.sh": "ssim",
"train_framework.sh": "train-framework",
"training.sh": "training",
"training_lora.sh": "lora-training",
"training_vsa.sh": "training-vsa",
}
FASTCHECK_LANE_SCRIPTS = {
"dreamverse.sh",
"encoder.sh",
"kernel_tests.sh",
"transformer.sh",
"vae.sh",
}
LEGACY_TRAINING_LANES = (
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
)
ALL_TRAINING_LANES = (*LEGACY_TRAINING_LANES, "train-framework")
SSIM_SMOKE_TESTS = (
"test_flux_t2i_similarity.py",
"test_wan_t2v_similarity.py",
)
SAFE_PATTERNS = (
"*.md",
"*.rst",
".agents/**",
".claude/**",
".codex/**",
".github/ISSUE_TEMPLATE/**",
".github/PULL_REQUEST_TEMPLATE.md",
".github/dependabot.yml",
".github/mergify.yml",
".github/scripts/**",
".github/workflows/**",
".buildkite/scripts/pre_commit.sh",
".git-blame-ignore-revs",
".gitattributes",
".gitignore",
".pre-commit-config.yaml",
"AGENTS.md",
"CITATION.cff",
"CODE_OF_CONDUCT.md",
"CONTRIBUTING.md",
"LICENSE",
"NOTICE",
"__init__.py",
"collect_env.py",
"SECURITY.md",
"assets/**",
"comfyui/**",
"docs/**",
"examples/**",
"mkdocs.yml",
"requirements-mkdocs.in",
"requirements-mkdocs.txt",
"scripts/**",
"tests/__init__.py",
"tests/local_tests/**",
)
ALL_IMPACT_PATTERNS = (
".buildkite/pipeline.yml",
"docker/**",
"pyproject.toml",
"requirements*.txt",
"setup.cfg",
"setup.py",
"uv.lock",
)
@dataclass(frozen=True)
class FamilyCoverage:
pattern: re.Pattern[str]
golden_tests: tuple[str, ...]
ssim_tests: tuple[str, ...]
FAMILY_COVERAGE = (
FamilyCoverage(
re.compile(r"(^|[/_.-])dreamx(_world)?([/_.-]|$)"),
("test_dreamx.py", ),
("test_dreamx_world_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux[_-]?2([/_.-]|$)"),
("test_flux2_klein.py", ),
("test_flux2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux(?![_-]?2)([/_.-]|$)"),
("test_flux.py", ),
("test_flux_t2i_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])(hunyuan)?gamecraft([/_.-]|$)"),
("test_gamecraft.py", ),
("test_gamecraft_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])gen3c([/_.-]|$)"),
("test_gen3c.py", ),
("test_gen3c_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])glm[_-]?image([/_.-]|$)"),
("test_glm_image.py", ),
("test_glm_image_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
("test_kandinsky5.py", ),
("test_kandinsky5_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])lingbot([a-z0-9_-]*)([/_.-]|$)"),
("test_lingbot.py", ),
("test_lingbot_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])longcat([/_.-]|$)"),
("test_longcat.py", ),
("test_longcat_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])ltx[_-]?2([/_.-]|$)"),
("test_ltx2.py", ),
("test_ltx2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?2([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?3([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])minimax[_-]?h3([/_.-]|$)"),
("test_minimax_h3_t2v.py", ),
("test_minimax_h3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])sd[_-]?3([._-]?5)?([/_.-]|$)"),
("test_sd35.py", ),
("test_sd35_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])stable[_-]?audio([/_.-]|$)"),
("test_stable_audio.py", ),
("test_stable_audio_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])turbo(diffusion)?([/_.-]|$)"),
(),
("test_turbodiffusion_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)"),
("test_wan_t2v.py", "test_wan_vae.py", "test_wan_causal.py", "test_wan_denoising.py"),
(
"test_causal_similarity.py",
"test_wan_i2v_similarity.py",
"test_wan_t2v_similarity.py",
),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])z[_-]?image([/_.-]|$)"),
("test_zimage.py", ),
("test_zimage_similarity.py", ),
),
)
@dataclass
class MergePlan:
lanes: set[str] = field(default_factory=set)
golden_tests: set[str] = field(default_factory=set)
ssim_tests: set[str] = field(default_factory=set)
golden_all: bool = False
ssim_all: bool = False
reasons: list[str] = field(default_factory=list)
def add_lanes(self, *lanes: str, reason: str) -> None:
unknown = set(lanes) - set(MERGE_LANES)
if unknown:
raise ValueError(f"Unknown merge lanes: {sorted(unknown)}")
self.lanes.update(lanes)
self.reasons.append(reason)
def add_golden(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("golden-gate", reason=reason)
self.golden_tests.update(tests)
def add_ssim(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("ssim", reason=reason)
self.ssim_tests.update(tests)
def require_all(self, reason: str) -> None:
self.lanes.update(MERGE_LANES)
self.golden_all = True
self.ssim_all = True
self.reasons.append(reason)
def ordered_lanes(self) -> tuple[str, ...]:
return tuple(lane for lane in MERGE_LANES if lane in self.lanes)
def encoded_lanes(self) -> str:
lanes = self.ordered_lanes()
return "," + ",".join(lanes or ("none", )) + ","
def encoded_golden_tests(self) -> str:
if "golden-gate" not in self.lanes:
return "none"
if self.golden_all or not self.golden_tests:
return "all"
return ",".join(sorted(self.golden_tests))
def encoded_ssim_tests(self) -> str:
if "ssim" not in self.lanes:
return "none"
if self.ssim_all or not self.ssim_tests:
return "all"
return ",".join(sorted(self.ssim_tests))
def _matches_any(path: str, patterns: tuple[str, ...]) -> bool:
return any(fnmatch.fnmatchcase(path, pattern) for pattern in patterns)
def _family_coverage(path: str) -> tuple[set[str], set[str]]:
normalized = path.lower()
golden: set[str] = set()
ssim: set[str] = set()
for family in FAMILY_COVERAGE:
if family.pattern.search(normalized):
golden.update(family.golden_tests)
ssim.update(family.ssim_tests)
# Select the component actually touched, including compatibility paths.
# Family configs/pipeline wiring can affect all four Wan gates.
if re.search(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)", normalized):
if (normalized.endswith(("/wan/vae.py", "/wan/vae_config.py", "/vaes/wanvae.py"))
or normalized.endswith("/wan/stages/conditioning.py")):
golden = {"test_wan_vae.py"}
elif normalized.endswith(("/wan/causal_transformer.py", "/dits/causal_wanvideo.py",
"/wan/stages/causal_denoising.py")):
golden = {"test_wan_causal.py"}
elif (normalized == "fastvideo/models/dits/wanvideo.py"
or normalized.endswith(("/wan/transformer.py", "/wan/stages/denoising.py", "/wan/stages/dmd.py"))):
golden = {"test_wan_t2v.py", "test_wan_denoising.py"}
return golden, ssim
def _select_output_coverage(plan: MergePlan, path: str) -> None:
golden, ssim = _family_coverage(path)
if golden:
plan.add_golden(tuple(sorted(golden)), reason=f"model-family golden coverage: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared output golden coverage: {path}")
if ssim:
plan.add_ssim(tuple(sorted(ssim)), reason=f"model-family SSIM coverage: {path}")
else:
plan.add_ssim(SSIM_SMOKE_TESTS, reason=f"shared output SSIM smoke coverage: {path}")
def classify_paths(paths: list[str]) -> MergePlan:
plan = MergePlan()
normalized_paths: list[str] = []
for raw_path in paths:
path = raw_path.strip()
while path.startswith("./"):
path = path[2:]
if path:
normalized_paths.append(path)
normalized_paths = sorted(set(normalized_paths))
if not normalized_paths:
plan.require_all("changed-file list was empty; failing closed")
return plan
for path in normalized_paths:
if path == "__FASTVIDEO_CI_PLAN_ALL__":
plan.require_all("changed-file API failed; failing closed")
continue
if path in {"requirements-mkdocs.in", "requirements-mkdocs.txt"}:
plan.reasons.append(f"documentation dependencies need no GPU integration: {path}")
continue
if _matches_any(path, ALL_IMPACT_PATTERNS):
plan.require_all(f"cross-cutting build/runtime surface: {path}")
continue
lane_script_prefix = ".buildkite/scripts/lanes/"
if path.startswith(lane_script_prefix):
script_name = Path(path).name
lane = LANE_SCRIPT_TO_KEY.get(script_name)
if lane is None:
if script_name in FASTCHECK_LANE_SCRIPTS:
plan.reasons.append(f"covered by automatic Fastcheck lane: {path}")
else:
plan.require_all(f"unknown lane script: {path}")
elif lane == "golden-gate":
plan.golden_all = True
plan.add_lanes(lane, reason=f"golden lane implementation: {path}")
elif lane == "ssim":
plan.ssim_all = True
plan.add_lanes(lane, reason=f"SSIM lane implementation: {path}")
else:
plan.add_lanes(lane, reason=f"lane implementation: {path}")
continue
if path.startswith("fastvideo/tests/golden_gate/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_golden((name, ), reason=f"changed golden test: {path}")
elif name in {"AGENTS.md", "README.md"}:
plan.reasons.append(f"golden documentation only: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared golden harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/ssim/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_ssim((name, ), reason=f"changed SSIM test: {path}")
elif path.endswith((".py", ".json", ".pt", ".png", ".mp4")):
plan.ssim_all = True
plan.add_lanes("ssim", reason=f"shared SSIM harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/performance/") or path.startswith(".buildkite/performance-benchmarks/"):
plan.add_lanes("performance", reason=f"performance coverage: {path}")
continue
if path.startswith(("fastvideo/performance/", "fastvideo/performance_dashboard/",
"apps/performance_dashboard/")):
plan.add_lanes("performance", reason=f"performance implementation: {path}")
continue
if path.startswith("fastvideo/benchmarks/"):
if "/mlx_" in path or Path(path).name.startswith("mlx_"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
else:
plan.add_lanes("performance", reason=f"benchmark implementation: {path}")
continue
if path.startswith("fastvideo/tests/eval/") or path.startswith("fastvideo/eval/"):
plan.add_lanes("eval", reason=f"evaluation coverage: {path}")
continue
if path.startswith("fastvideo/third_party/eval/"):
plan.add_lanes("eval", reason=f"vendored evaluation implementation: {path}")
continue
if path.startswith("fastvideo/tests/lora_extraction/") or path.startswith("scripts/lora_extraction/"):
plan.add_lanes("lora-extraction", reason=f"LoRA extraction coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/lora/"):
plan.add_lanes("lora-inference", reason=f"LoRA inference coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/vmoba/"):
plan.add_lanes("inference-vmoba", reason=f"VMoBA inference coverage: {path}")
continue
if path.startswith(("fastvideo/dataset/", "fastvideo/workflow/", "fastvideo/pipelines/preprocess/",
"fastvideo/pipelines/training/")):
plan.add_lanes(*ALL_TRAINING_LANES, reason=f"shared data/training input surface: {path}")
continue
if path.startswith("fastvideo/tests/train/") or path.startswith("fastvideo/train/"):
plan.add_lanes("train-framework", reason=f"modular training coverage: {path}")
continue
if path.startswith("fastvideo/tests/training/"):
lowered = path.lower()
if "/vanilla/" in lowered:
plan.add_lanes("training", reason=f"vanilla training coverage: {path}")
elif "/distill/" in lowered:
plan.add_lanes("distillation", reason=f"distillation coverage: {path}")
elif "/self-forcing/" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing coverage: {path}")
elif "/lora/" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training coverage: {path}")
elif "/vsa/" in lowered:
plan.add_lanes("training-vsa", reason=f"VSA training coverage: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training coverage: {path}")
continue
if path.startswith("fastvideo/training/"):
lowered = path.lower()
if "self_forcing" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing implementation: {path}")
elif "distill" in lowered:
plan.add_lanes("distillation", reason=f"distillation implementation: {path}")
elif "lora" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training implementation: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training implementation: {path}")
continue
lowered = path.lower()
if "vmoba" in lowered and path.startswith(("fastvideo/", ".buildkite/")):
plan.add_lanes("inference-vmoba", reason=f"VMoBA implementation: {path}")
plan.add_golden(("test_wan_t2v.py", ), reason=f"VMoBA end-to-end coverage: {path}")
continue
if "lora" in lowered and path.startswith("fastvideo/"):
plan.add_lanes(
"lora-inference",
"lora-extraction",
"lora-training",
reason=f"shared LoRA implementation: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/entrypoints/") or path.startswith("fastvideo/api/"):
plan.add_lanes("api-server", reason=f"API/entrypoint integration: {path}")
if "openai" not in lowered and "/cli/" not in lowered:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/worker/"):
plan.add_lanes("api-server", reason=f"worker/API integration: {path}")
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/distributed/"):
plan.add_lanes(
"training",
"train-framework",
reason=f"distributed runtime integration: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/hooks/", "fastvideo/platforms/", "fastvideo/third_party/")):
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/models/", "fastvideo/pipelines/", "fastvideo/configs/",
"fastvideo/layers/", "fastvideo/attention/")):
_select_output_coverage(plan, path)
continue
if path in {
"fastvideo/fastvideo_args.py",
"fastvideo/forward_context.py",
"fastvideo/image_processor.py",
"fastvideo/registry.py",
"fastvideo/utils.py",
}:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/mlx_runtime/"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
continue
if path.startswith("fastvideo/logging_utils/") or path in {
"fastvideo/__init__.py",
"fastvideo/envs.py",
"fastvideo/logger.py",
"fastvideo/profiler.py",
"fastvideo/version.py",
}:
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path.startswith(("fastvideo-kernel/", "csrc/")):
plan.add_golden(("test_wan_t2v.py", ), reason=f"kernel integration smoke: {path}")
plan.add_ssim(("test_wan_t2v_similarity.py", ), reason=f"kernel numerical smoke: {path}")
continue
if path.startswith("apps/dreamverse/"):
# DreamVerse is already one of the six automatic Fastcheck lanes.
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
continue
if path.startswith("fastvideo/tests/"):
# The automatic unit/component Fastcheck lanes own the remaining
# package tests. Domain-specific expensive test roots were handled
# above.
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path in {".buildkite/scripts/unit_test.sh", ".buildkite/scripts/pr_test.sh"}:
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if _matches_any(path, SAFE_PATTERNS):
plan.reasons.append(f"no additional GPU integration needed: {path}")
continue
plan.require_all(f"unclassified path; failing closed: {path}")
return plan
def _write_github_output(output: TextIO, plan: MergePlan) -> None:
output.write(f"merge_test_plan={plan.encoded_lanes()}\n")
output.write(f"merge_golden_tests={plan.encoded_golden_tests()}\n")
output.write(f"merge_ssim_tests={plan.encoded_ssim_tests()}\n")
output.write(f"merge_plan_label={','.join(plan.ordered_lanes()) or 'none'}\n")
def _write_summary(output: TextIO, plan: MergePlan) -> None:
output.write("## Change-aware merge test plan\n\n")
output.write("| Selection | Value |\n|---|---|\n")
output.write(f"| Additional Slurm lanes | `{','.join(plan.ordered_lanes()) or 'none'}` |\n")
output.write(f"| Golden tests | `{plan.encoded_golden_tests()}` |\n")
output.write(f"| SSIM tests | `{plan.encoded_ssim_tests()}` |\n\n")
output.write("Fastcheck remains the universal six-lane baseline.\n")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--paths-file", type=Path, required=True)
parser.add_argument("--github-output", type=Path)
parser.add_argument("--summary-file", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
paths = args.paths_file.read_text(encoding="utf-8").splitlines()
plan = classify_paths(paths)
print(f"MERGE_TEST_PLAN={plan.encoded_lanes()}")
print(f"MERGE_GOLDEN_TESTS={plan.encoded_golden_tests()}")
print(f"MERGE_SSIM_TESTS={plan.encoded_ssim_tests()}")
for reason in plan.reasons:
print(f"- {reason}")
if args.github_output:
with args.github_output.open("a", encoding="utf-8") as output:
_write_github_output(output, plan)
if args.summary_file:
with args.summary_file.open("a", encoding="utf-8") as output:
_write_summary(output, plan)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -1
View File
@@ -53,7 +53,7 @@ PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
OTHER='{"name": "Trigger Merge Gate", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
@@ -190,6 +190,7 @@ jobs:
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "Digest: ${{ steps.build-push.outputs.digest }}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
+39 -25
View File
@@ -26,29 +26,48 @@ jobs:
per_page: 100,
});
const bkStatuses = data.statuses.filter(
s => s.context.startsWith('buildkite/ci/')
);
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
// Buildkite derives the GitHub context prefix from the label emoji.
// Keep hard Full Suite lanes in test-tube/bar-chart namespaces and
// Fastcheck lanes in microscope so targeted reruns cannot clear the
// wrong aggregate status. Automatic PR jobs use pr-fastcheck while
// slash-command and Full Suite jobs use ci; normalize the suffix
// and keep the newest status for each logical lane.
const FASTCHECK_PREFIXES = [
'buildkite/pr-fastcheck/microscope-',
'buildkite/ci/microscope-',
];
const FULL_SUITE_PREFIXES = [
'buildkite/ci/test-tube-',
'buildkite/ci/bar-chart-',
];
const fastcheck = bkStatuses.filter(
s => s.context.startsWith(FASTCHECK_PREFIX)
);
const fullSuite = bkStatuses.filter(
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
);
function newestByLane(prefixes) {
const statuses = new Map();
for (const status of data.statuses) {
const prefix = prefixes.find(p => status.context.startsWith(p));
if (!prefix) continue;
const lane = status.context.slice(prefix.length);
const previous = statuses.get(lane);
if (!previous || Date.parse(status.updated_at) > Date.parse(previous.updated_at)) {
statuses.set(lane, status);
}
}
return statuses;
}
if (
fastcheck.length > 0
&& fastcheck.every(s => s.state === 'success')
) {
const fastcheck = newestByLane(FASTCHECK_PREFIXES);
const fullSuiteOnly = newestByLane(FULL_SUITE_PREFIXES);
const fastcheckPassed =
fastcheck.size === 6
&& [...fastcheck.values()].every(s => s.state === 'success');
const fullSuitePassed =
fastcheckPassed
&& fullSuiteOnly.size === 14
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
if (fastcheckPassed) {
core.info(
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -56,17 +75,13 @@ jobs:
sha,
state: 'success',
context: 'fastcheck-passed',
description:
`All ${fastcheck.length} fastcheck tests passed`,
description: `All ${fastcheck.size} fastcheck tests passed`,
});
}
if (
fullSuite.length > 0
&& fullSuite.every(s => s.state === 'success')
) {
if (fullSuitePassed) {
core.info(
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
'All 20 full suite tests passed — updating full-suite-passed'
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -74,7 +89,6 @@ jobs:
sha,
state: 'success',
context: 'full-suite-passed',
description:
`All ${fullSuite.length} full suite tests passed`,
description: 'All 20 full suite tests passed',
});
}
+20 -2
View File
@@ -8,6 +8,8 @@ on:
- "fastvideo/mlx_runtime/**"
- "fastvideo/tests/mlx/**"
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
- "fastvideo/tests/platforms/test_cpu_sdpa.py"
- "fastvideo/platforms/cpu.py"
- "fastvideo/platforms/mps.py"
- "fastvideo/platforms/__init__.py"
- "fastvideo/__init__.py"
@@ -78,6 +80,13 @@ jobs:
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
@@ -90,7 +99,8 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
@@ -134,6 +144,13 @@ jobs:
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
@@ -146,4 +163,5 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
+44
View File
@@ -0,0 +1,44 @@
name: Scheduled Full SSIM
on:
schedule:
- cron: "0 5 * * 0"
workflow_dispatch:
permissions:
contents: read
jobs:
trigger:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- name: Trigger weekly full SSIM on Slinky Slurm
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
SOURCE_SHA: ${{ github.sha }}
SOURCE_BRANCH: ${{ github.event.repository.default_branch }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
set -euo pipefail
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$SOURCE_SHA" \
--arg branch "$SOURCE_BRANCH" \
'{
commit: $commit,
branch: $branch,
message: "Weekly full SSIM on Slinky Slurm",
ignore_pipeline_branch_filters: true,
env: {
TEST_SCOPE: "scheduled",
FULL_SUITE: "false",
TEST_TYPE: "ssim",
PR_NUMBER: "false",
PR_TITLE: "Scheduled full SSIM"
}
}')"
+12 -44
View File
@@ -33,7 +33,6 @@ jobs:
core.setOutput('has_write', String(hasWrite));
- name: Add ready label and react
id: label
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
@@ -48,47 +47,6 @@ jobs:
comment_id: context.payload.comment.id,
content: 'rocket',
});
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
core.setOutput('pr_title', pr.title);
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
PR_TITLE: ${{ steps.label.outputs.pr_title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
parse-command:
if: >-
@@ -129,7 +87,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval unit-ci kernel-ci dreamverse-ci ssim-ci golden-gate-ci encoder-ci vae-ci transformer-ci lora-inference-ci lora-training-ci lora-extraction-ci training-ci distillation-ci self-forcing-ci vsa-ci vmoba-ci performance-ci api-ci train-framework-ci eval-ci full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -137,7 +95,17 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[kernel]=kernel_tests [unit]=unit_test [unit-ci]=unit_test_ci
[kernel-ci]=kernel_tests_ci [dreamverse-ci]=dreamverse_app_ci
[ssim-ci]=ssim_ci [vmoba-ci]=inference_vmoba_ci
[golden-gate-ci]=golden_gate_ci [training-ci]=training_ci
[encoder-ci]=encoder_ci [vae-ci]=vae_ci [transformer-ci]=transformer_ci
[lora-inference-ci]=inference_lora_ci [lora-training-ci]=training_lora_ci
[lora-extraction-ci]=lora_extraction_ci [distillation-ci]=distillation_dmd_ci
[self-forcing-ci]=self_forcing_ci [vsa-ci]=training_vsa_ci
[performance-ci]=performance_ci [api-ci]=api_server_ci
[train-framework-ci]=train_framework_ci [eval-ci]=eval_ci
[dreamverse]=dreamverse_app
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
+66 -13
View File
@@ -1,4 +1,4 @@
name: Trigger Full Suite
name: Trigger Merge Gate
on:
pull_request_target:
@@ -10,7 +10,7 @@ permissions:
actions: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
group: merge-gate-${{ github.event.pull_request.number }}
cancel-in-progress: false
jobs:
@@ -34,29 +34,72 @@ jobs:
});
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
core.setOutput('changed_files', String(pr.changed_files));
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
- name: Cancel previous Buildkite builds
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
# Find running builds for this branch with TEST_SCOPE=full and cancel them
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
# Match both branch and PR number: forks can reuse the same branch name.
builds=$(curl -sS --get -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state=running,scheduled" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds" \
| jq -r --arg pr_number "$PR_NUMBER" \
'.[] | select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number)) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
# Checks out the BASE branch (default for pull_request_target), so PR
# authors cannot tamper with the gate script.
- name: Checkout gate script
# Check out the immutable BASE SHA: pull_request_target must never run a
# planner or gate script from the untrusted PR head.
- name: Checkout trusted merge planner
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
ref: ${{ github.event.pull_request.base.sha }}
persist-credentials: false
- name: Collect changed paths
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ github.event.pull_request.number }}
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
run: |
set -euo pipefail
changed_json="$RUNNER_TEMP/merge-changed-files.json"
changed_paths="$RUNNER_TEMP/merge-changed-paths.txt"
if gh api --paginate --slurp \
"repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/files?per_page=100" \
> "$changed_json"; then
observed=$(jq '[.[][] | .filename] | unique | length' "$changed_json")
if [ "$observed" = "$EXPECTED_CHANGED_FILES" ]; then
jq -r '.[][] | .filename, (.previous_filename // empty)' "$changed_json" \
| sort -u > "$changed_paths"
else
echo "::warning::Changed-file API returned $observed of $EXPECTED_CHANGED_FILES paths; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
else
echo "::warning::Changed-file API failed; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
- name: Select minimal merge tests
id: plan
if: steps.check.outputs.has_ready == 'true'
run: |
python3 .github/scripts/plan_merge_ci.py \
--paths-file "$RUNNER_TEMP/merge-changed-paths.txt" \
--github-output "$GITHUB_OUTPUT" \
--summary-file "$GITHUB_STEP_SUMMARY"
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
@@ -66,7 +109,7 @@ jobs:
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite Full Suite
- name: Trigger Buildkite merge gate
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
@@ -76,6 +119,10 @@ jobs:
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
MERGE_GOLDEN_TESTS: ${{ steps.plan.outputs.merge_golden_tests }}
MERGE_SSIM_TESTS: ${{ steps.plan.outputs.merge_ssim_tests }}
MERGE_PLAN_LABEL: ${{ steps.plan.outputs.merge_plan_label }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
@@ -84,8 +131,11 @@ jobs:
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg message "Merge gate [${MERGE_PLAN_LABEL}] for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--arg merge_test_plan "$MERGE_TEST_PLAN" \
--arg merge_golden_tests "$MERGE_GOLDEN_TESTS" \
--arg merge_ssim_tests "$MERGE_SSIM_TESTS" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -95,8 +145,11 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
TEST_SCOPE: "merge",
FULL_SUITE: "true",
MERGE_TEST_PLAN: $merge_test_plan,
MERGE_GOLDEN_TESTS: $merge_golden_tests,
MERGE_SSIM_TESTS: $merge_ssim_tests,
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
+4 -4
View File
@@ -38,17 +38,17 @@ jobs:
**How our CI works:**
PRs run a two-tier CI system:
PRs run a three-tier CI system:
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
2. **Fastcheck** — six core GPU lanes run automatically via Buildkite (~10-15 min).
3. **Merge gate** — a reviewer adds `ready`; changed paths select only the relevant integration, training, golden, or SSIM coverage.
**Before your PR is reviewed:**
- [ ] `pre-commit run --all-files` passes locally
- [ ] You've added or updated tests for your changes
- [ ] The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
**Useful links:**
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
+27
View File
@@ -13,6 +13,11 @@ on:
required: false
default: false
type: boolean
build_ci_runner_image:
description: 'Build the ARM64 CUDA 13 CI runner image (sm_100)'
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
@@ -198,6 +203,28 @@ jobs:
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# The CI runner is ARM64 like DGX Spark, but targets sm_100 rather than sm_121.
# Publish a single-architecture variant so the self-hosted CI runner can reuse
# the exact prebuilt kernel instead of compiling it in every job.
build-ci-runner-image:
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_ci_runner_image == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12-cuda13.0.0-sm100
runner: ubuntu-24.04-arm
architecture: arm64
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=13.0.0
UV_TORCH_BACKEND=cu130
TORCH_CUDA_ARCH_LIST=10.0
CMAKE_BUILD_PARALLEL_LEVEL=1
FLASH_ATTN_WHEEL_TAG=cu130torch2.12
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
+15 -8
View File
@@ -62,8 +62,9 @@ jobs:
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
# x86_64 builds the full cu126 + cu130 set. cu130 ships the
# data-center Blackwell sm_100a/sm_103a VSA and consumer sm_120a FP4
# kernels.
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
@@ -124,7 +125,7 @@ jobs:
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo apt install -y git gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
@@ -168,17 +169,18 @@ jobs:
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
# + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="9.0a;10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
@@ -194,7 +196,11 @@ jobs:
python -m build --wheel --outdir dist
# Fix the wheel to be manylinux compliant
uv pip install --system auditwheel
# Ubuntu 22.04 ships patchelf 0.14.3, while current auditwheel
# requires at least 0.14.5. Use the stable PyPI binary on both
# x86_64 and aarch64 release runners.
uv pip install --system auditwheel patchelf==0.17.2.4
patchelf --version
# Point auditwheel at torch libs, but do not vendor them into the wheel.
TORCH_LIB_DIR=$(python - <<'PY'
import os
@@ -211,7 +217,8 @@ jobs:
--exclude libtorch.so \
--exclude libc10.so \
--exclude libc10_cuda.so \
--exclude libtorch_python.so
--exclude libtorch_python.so \
--exclude libnccl.so.2
# Move fixed wheels back to dist for upload consistency
rm dist/*.whl
mv fixed_dist/*.whl dist/
+5
View File
@@ -55,6 +55,8 @@ eggs/
# MkDocs documentation
site/
docs/assets/cookbook-serving.json
examples/serving/clients/node_modules/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
@@ -133,6 +135,9 @@ fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
.kernel_bench/
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
+1 -1
View File
@@ -9,7 +9,7 @@ exclude: |
tests/.*|
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
fastvideo/models/(?!wan/(config|vae_config|pipeline_config|definition|__init__)\.py$).*|
^apps/dreamverse/web/.*|
examples/.*|
\.agents/.*|
+4
View File
@@ -66,14 +66,18 @@ Local guidance lives next to the code. Read the in-scope file before editing:
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
| `fastvideo/configs/AGENTS.md` | Arch + pipeline config dataclasses, `param_names_mapping` |
| `fastvideo/models/AGENTS.md` | DiT / VAE / encoder / scheduler / loader layout (pre-commit excluded) |
| `fastvideo/models/wan/AGENTS.md` | Wan family-local transformers, VAE, configs, and the SP sharding invariant |
| `fastvideo/layers/AGENTS.md` | Tensor-parallel linear/attention layer rules for ports |
| `fastvideo/attention/AGENTS.md` | Backend registry + env-var override |
| `fastvideo/pipelines/AGENTS.md` | Stage ABC, `basic/<model>/`, `preprocess/`, presets |
| `fastvideo/pipelines/basic/wan/AGENTS.md` | Wan sampling stages, first-frame conditioning, DMD/causal boundaries |
| `fastvideo/pipelines/basic/magi_human/AGENTS.md` | MagiHuman umbrella repo, lazy-loaded components, packing invariants |
| `fastvideo/training/AGENTS.md` | Legacy monolithic pipelines (frozen for existing models) |
| `fastvideo/train/AGENTS.md` | New modular trainer (methods × models × callbacks, YAML) |
| `fastvideo/tests/AGENTS.md` | Test taxonomy, conftest, pre-commit-excluded path |
| `fastvideo/tests/ssim/AGENTS.md` | GPU SSIM regression authoring + reference video sync |
| `scripts/checkpoint_conversion/AGENTS.md` | Adding a converter for a new HF/official checkpoint |
| `apps/dreamverse/AGENTS.md` | DreamVerse app structure and conventions |
## Critical: Two Training Stacks Coexist
+8 -4
View File
@@ -3,12 +3,15 @@
</div>
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://haoailab.com/FastVideo/cookbook/"><b>Cookbook</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
</p>
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/09/15`: Release [FastH3 8-Step V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2), an eight-forward data-free DMD2 checkpoint distilled from MiniMax-H3 with 80% Video Sparse Attention. Run it with `examples/inference/basic/basic_fasth3_8step.py` or the [FastH3 8-Step V2 recipe](https://haoailab.com/FastVideo/cookbook/minimax-h3/).
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac—follow the [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
@@ -63,9 +66,10 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
> **On an Apple Silicon Mac?** FastVideo runs FastWan text-to-video natively
> through an MLX runtime — a 5-second 480p clip generated locally, no cloud,
> no discrete GPU. Install with `uv pip install -e '.[mlx]'` and follow the
> **On an Apple Silicon Mac?** FastVideo runs FastMetal-QAD through an MLX
> runtime. Install with `uv pip install -e '.[mlx]'`, download
> [`FastVideo/FastMetal-1.3B-QAD`](https://huggingface.co/FastVideo/FastMetal-1.3B-QAD),
> and follow the
> [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
+23 -2
View File
@@ -97,13 +97,33 @@ dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
### Run Dreamverse with FastH3
Select the VSA data-free FastH3 Preview profile when you start the backend:
```bash
DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
The `fast-h3` profile uses four visible GPUs by default. It loads the `MiniMaxAI/MiniMax-H3` base checkpoint and the
`vsa-datafree/adapter_model.safetensors` adapter from
`FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA`. Each request generates a 124-frame, 768×1344 video with
synchronized audio and five sigma-grid points. Dreamverse uses the last frame of each segment as first-frame
conditioning for the following segment.
Set `CUDA_VISIBLE_DEVICES` when you need to choose the four physical GPUs:
```bash
CUDA_VISIBLE_DEVICES=0,1,2,3 DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
> (the default), the backend compiles the segment 1 and segment 2 inference
> paths before it reports ready — this can take **tens of minutes on a cold
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
> warmup finishes. For a faster, uncompiled startup while testing, set
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
> warmup finishes. To defer compilation until the first generated request while
> testing, set `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
## Frontend Setup
@@ -219,6 +239,7 @@ selection, and mock-server behavior:
pytest apps/dreamverse/dreamverse/tests/test_config.py \
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
apps/dreamverse/dreamverse/tests/test_minimax_h3_generation.py \
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
```
+44 -4
View File
@@ -139,7 +139,18 @@ session.
- startup warmup
- user join/leave commands
- `USER_STEP` execution for each segment
- continuation state between segments
- generation-command routing and stream-result delivery
Model generation has a separate ownership boundary inside each GPU process:
- `apps/dreamverse/dreamverse/generation_worker.py` selects the backend that the active model profile declares and owns
the backend lifecycle.
- `apps/dreamverse/dreamverse/ltx2_generation.py` owns LTX-2 generator configuration, video and audio continuation, and
runtime LoRA application.
- `apps/dreamverse/dreamverse/minimax_h3_generation.py` owns the VSA data-free FastH3 adapter, FastH3 generator and
request configuration, and last-frame continuation through MiniMax H3 first-frame conditioning.
- `apps/dreamverse/dreamverse/generation_contracts.py` defines the decoded media and stream-trimming result that both
model backends return to `apps/dreamverse/dreamverse/gpu_pool.py`.
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
@@ -288,18 +299,47 @@ There are three related prompt paths in the current system:
## Initial Image And Segment Handling
The frontend currently sends `initial_image` as part of session init or
The frontend sends `initial_image` and, for first/last frame mode,
`last_frame_image` as part of `session_init_v2`, `project_init_v1`, or
`simple_generate`.
The server:
- validates and persists the image
- uses it only for segment 1 when present
- validates and persists the images
- uses `initial_image` only for segment 1 when present
- keeps continuation state for later segments in the GPU worker
This means the runtime, not the frontend, decides how segment 1 image
conditioning and later continuation conditioning are applied.
## Creation Studio Config
The lobby creation studio sends model, mode, aspect ratio, resolution, and
duration with session init. The server parses these fields into a per-session
creation config and echoes the resolved values back on `gpu_assigned` and
`ltx2_stream_start` as `creation_config`.
Incoming fields on `session_init_v2` and `project_init_v1`:
- `generation_mode`: `t2va`, `fl2va`, or `ref2va` (canonical upstream IDs from #1834)
- `model_id`: `fast-ltx2`, `fast-ltx23`, or `fast-h3`
- `aspect_ratio`: one of `21:9`, `16:9`, `4:3`, `1:1`, `3:4`, `9:16`
- `resolution`: one of `480p`, `720p`, `1080p`, `4k`
- `duration_sec`: `5`, `10`, or `15`
- `initial_image`: optional image payload for reference / first-frame modes
- `last_frame_image`: optional image payload for first/last frame mode
Echoed `creation_config` includes the resolved frame size,
`num_frames`, and `generation_segment_cap` derived from `duration_sec`.
Mode validation:
- `ref2va` requires `initial_image`
- `fl2va` requires both `initial_image` and `last_frame_image`
Per-step generation uses the resolved `frame_width`, `frame_height`, and
`num_frames` from the session creation config.
## Websocket Contract
The websocket is the main integration surface between UI and runtime.
@@ -1,6 +1,6 @@
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
Mirrors how ``apps/dreamverse/dreamverse/video_generation.py`` constructs
Mirrors how ``apps/dreamverse/dreamverse/ltx2_generation.py`` constructs
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
+21 -2
View File
@@ -1,5 +1,6 @@
import os
from pathlib import Path
from typing import cast
_REPO_ROOT = Path(__file__).resolve().parents[1]
_SERVER_ROOT = Path(__file__).resolve().parent
@@ -55,16 +56,34 @@ FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
MODEL_REGISTRY = {
"fast-ltx2": {
"name": "FastLTX2",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
},
"fast-ltx23": {
"name": "FastLTX23",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
},
"fast-h3": {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
@@ -171,7 +190,7 @@ def _optional_env(*names: str) -> str | None:
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", cast(int, MODEL_CONFIG["default_sp_size"])))
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
if DREAMVERSE_MODEL_PATH:
@@ -213,7 +232,7 @@ def _resolve_lora_spec(spec: str) -> str | None:
if not spec:
return None
if spec.lower() == "omninft":
return MODEL_CONFIG.get("lora_repo")
return cast(str | None, MODEL_CONFIG.get("lora_repo"))
if spec.lower() in AVAILABLE_LORAS:
return AVAILABLE_LORAS[spec.lower()]["repo"]
return spec
@@ -0,0 +1,137 @@
from __future__ import annotations
from dataclasses import dataclass
from dreamverse.config import MODEL_REGISTRY
# Canonical upstream wire IDs. FL2VA is tracked in #1834 but not wired on Dreamverse
# streaming backends yet.
LTX_LOBBY_GENERATION_MODES = frozenset({"t2va", "ref2va"})
H3_LOBBY_GENERATION_MODES = frozenset({"t2va", "ref2va"})
LTX_LOBBY_ASPECT_RATIOS = frozenset({"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"})
# Realtime FastLTX serving is validated through 1080p-class outputs; 4K is rejected
# until the runtime path is tested on Dreamverse GPUs.
LTX_LOBBY_RESOLUTIONS = frozenset({"480p", "720p", "1080p"})
# FastH3 serves a fixed 768x1344 (16:9-class) output; lobby resolution is nominal.
H3_LOBBY_ASPECT_RATIOS = frozenset({"16:9"})
H3_LOBBY_RESOLUTIONS = frozenset({"720p"})
LOBBY_DURATION_SEC = frozenset({5, 10, 15})
UNSUPPORTED_GENERATION_MODE_MESSAGES = {
"fl2va": "First/last frame mode (FL2VA) is not supported yet.",
}
@dataclass(frozen=True)
class ModelCreationCapabilities:
generation_modes: frozenset[str]
aspect_ratios: frozenset[str]
resolutions: frozenset[str]
duration_sec: frozenset[int]
unsupported_generation_modes: frozenset[str] = frozenset({"fl2va"})
def as_dict(self) -> dict[str, object]:
unsupported = {
mode: UNSUPPORTED_GENERATION_MODE_MESSAGES[mode]
for mode in sorted(self.unsupported_generation_modes)
if mode in UNSUPPORTED_GENERATION_MODE_MESSAGES
}
return {
"generation_modes": sorted(self.generation_modes),
"aspect_ratios": sorted(self.aspect_ratios),
"resolutions": sorted(self.resolutions),
"duration_sec": sorted(self.duration_sec),
"unsupported_generation_modes": unsupported,
"reference_assets": {
"mime_types": ["image/png", "image/jpeg", "image/webp"],
"max_bytes": 15 * 1024 * 1024,
},
}
LTX_MODEL_CREATION_CAPABILITIES = ModelCreationCapabilities(
generation_modes=LTX_LOBBY_GENERATION_MODES,
aspect_ratios=LTX_LOBBY_ASPECT_RATIOS,
resolutions=LTX_LOBBY_RESOLUTIONS,
duration_sec=LOBBY_DURATION_SEC,
)
H3_MODEL_CREATION_CAPABILITIES = ModelCreationCapabilities(
generation_modes=H3_LOBBY_GENERATION_MODES,
aspect_ratios=H3_LOBBY_ASPECT_RATIOS,
resolutions=H3_LOBBY_RESOLUTIONS,
duration_sec=LOBBY_DURATION_SEC,
)
MODEL_CREATION_CAPABILITIES: dict[str, ModelCreationCapabilities] = {
"fast-ltx2": LTX_MODEL_CREATION_CAPABILITIES,
"fast-ltx23": LTX_MODEL_CREATION_CAPABILITIES,
"fast-h3": H3_MODEL_CREATION_CAPABILITIES,
}
def capabilities_for_model(model_id: str) -> ModelCreationCapabilities:
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unknown model_id: {model_id}")
return MODEL_CREATION_CAPABILITIES.get(model_id, LTX_MODEL_CREATION_CAPABILITIES)
def lobby_capabilities_as_dict() -> dict[str, object]:
model_ids = sorted(MODEL_REGISTRY.keys())
models = {model_id: capabilities_for_model(model_id).as_dict() for model_id in model_ids}
union_modes: set[str] = set()
union_aspects: set[str] = set()
union_resolutions: set[str] = set()
union_durations: set[int] = set()
for caps in MODEL_CREATION_CAPABILITIES.values():
union_modes.update(caps.generation_modes)
union_aspects.update(caps.aspect_ratios)
union_resolutions.update(caps.resolutions)
union_durations.update(caps.duration_sec)
return {
"model_ids": model_ids,
"models": models,
"generation_modes": sorted(union_modes),
"aspect_ratios": sorted(union_aspects),
"resolutions": sorted(union_resolutions),
"duration_sec": sorted(union_durations),
"unsupported_generation_modes": dict(UNSUPPORTED_GENERATION_MODE_MESSAGES),
"reference_assets": {
"mime_types": ["image/png", "image/jpeg", "image/webp"],
"max_bytes": 15 * 1024 * 1024,
},
}
# Backward-compatible alias used in tests.
LOBBY_CREATION_CAPABILITIES = lobby_capabilities_as_dict()
def validate_lobby_creation_config(
*,
model_id: str,
generation_mode: str,
aspect_ratio: str,
resolution: str,
duration_sec: int,
) -> None:
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unknown model_id: {model_id}")
caps = capabilities_for_model(model_id)
if generation_mode in caps.unsupported_generation_modes:
raise ValueError(UNSUPPORTED_GENERATION_MODE_MESSAGES[generation_mode])
if generation_mode not in caps.generation_modes:
raise ValueError(f"Unsupported generation_mode: {generation_mode}")
if aspect_ratio not in caps.aspect_ratios:
raise ValueError(f"Unsupported aspect_ratio: {aspect_ratio}")
if resolution not in caps.resolutions:
raise ValueError(f"Unsupported resolution: {resolution}")
if duration_sec not in caps.duration_sec:
raise ValueError("duration_sec must be 5, 10, or 15.")
@@ -0,0 +1,50 @@
"""Shared contract between DreamVerse generation backends and GPU workers."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Protocol
@dataclass
class StepResult:
"""Decoded media and stream-trimming metadata for one DreamVerse segment."""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict[str, float]
head_trim_frames: int
head_trim_audio_frames: int
class GenerationBackend(Protocol):
"""Model-owned generation operations used by one GPU worker process."""
def initialize(self, model_config: dict | None = None) -> None:
...
def shutdown(self) -> None:
...
def clear_conditioning(self) -> None:
...
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
...
def warmup(self, prompt: str) -> dict[str, float]:
...
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
...
@@ -0,0 +1,103 @@
"""Select and own one model-specific generation backend per GPU process."""
from __future__ import annotations
from dreamverse.config import MODEL_CONFIG
from dreamverse.generation_contracts import GenerationBackend, StepResult
def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBackend:
"""Construct the backend that owns the selected model family's behavior."""
if backend_name == "ltx2":
from dreamverse.ltx2_generation import LTX2GenerationBackend
return LTX2GenerationBackend(gpu_id)
if backend_name == "minimax_h3":
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
return MiniMaxH3GenerationBackend(gpu_id)
raise ValueError(f"Unsupported DreamVerse generation backend: {backend_name!r}")
class VideoGenerationWorker:
"""Delegate GPU lifecycle and generation calls to the active model backend."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.model_config: dict = dict(MODEL_CONFIG)
self.backend_name: str | None = None
self.backend: GenerationBackend | None = None
def initialize(self, model_config: dict | None = None) -> None:
"""Load the requested model through its generation backend.
Model selection belongs here so the GPU process and streaming layers
use one stable media contract without importing model-specific code.
"""
requested_model_config = dict(model_config) if model_config is not None else dict(self.model_config)
backend_name = requested_model_config.get("generation_backend")
if not isinstance(backend_name, str) or not backend_name:
raise ValueError("DreamVerse model configuration requires `generation_backend`.")
candidate_backend = self.backend
if candidate_backend is None or self.backend_name != backend_name:
if candidate_backend is not None:
candidate_backend.shutdown()
candidate_backend = _create_generation_backend(backend_name, self.gpu_id)
try:
candidate_backend.initialize(requested_model_config)
except Exception:
try:
candidate_backend.shutdown()
except Exception as shutdown_error:
print(f"[GPU {self.gpu_id}] Backend cleanup after initialization failure: {shutdown_error}")
self.backend = None
self.backend_name = None
raise
self.model_config = requested_model_config
self.backend = candidate_backend
self.backend_name = backend_name
def _require_backend(self) -> GenerationBackend:
"""Return the initialized backend or fail before processing a command."""
if self.backend is None:
raise RuntimeError("Generation backend is not initialized.")
return self.backend
def shutdown(self) -> None:
"""Release model resources owned by the selected backend."""
if self.backend is not None:
self.backend.shutdown()
def clear_conditioning(self) -> None:
self._require_backend().clear_conditioning()
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Generate one segment through the selected model backend."""
return self._require_backend().generate_step(
prompt,
segment_idx,
image_path,
reset_conditioning,
frame_width=frame_width,
frame_height=frame_height,
num_frames=num_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
return self._require_backend().warmup(prompt)
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
return self._require_backend().apply_lora_stack(stack)
+27 -10
View File
@@ -12,7 +12,7 @@ from enum import Enum
from multiprocessing import Process, Queue
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
DREAMVERSE_SP_SIZE,
MODEL_REGISTRY,
STARTUP_WARMUP_ENABLED,
@@ -54,7 +54,7 @@ from dreamverse.worker_ipc import (
def _parse_requested_gpu_limit() -> int | None:
raw_value = os.getenv("FASTVIDEO_GPU_COUNT", "").strip().lower()
if not raw_value:
return 1
return DREAMVERSE_SP_SIZE
if raw_value == "all":
return None
try:
@@ -164,12 +164,12 @@ def gpu_worker_process(
os.environ["CUDA_VISIBLE_DEVICES"] = cuda_device
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
from dreamverse.video_generation import VideoGenerationWorker
from dreamverse.generation_worker import VideoGenerationWorker
worker = VideoGenerationWorker(gpu_id)
def event_loop(first_cmd: Command = None):
"""Blocking event loop for LTX2; dispatches user commands."""
"""Block on generation commands after the model is initialized."""
print(f"[GPU {gpu_id}] Entering event loop")
def handle_command(cmd: Command):
@@ -189,6 +189,9 @@ def gpu_worker_process(
segment_idx,
image_path=payload.image_path,
reset_conditioning=payload.reset_conditioning,
frame_width=payload.frame_width,
frame_height=payload.frame_height,
num_frames=payload.num_frames,
)
head_trim_frames = step_result.head_trim_frames
head_trim_audio_frames = step_result.head_trim_audio_frames
@@ -435,7 +438,7 @@ class GPUSlot:
self._response_reader_task: asyncio.Task | None = None
self._active: bool = False
self._reader_lock: asyncio.Lock | None = None
self.current_model_id: str = DEFAULT_MODEL_ID
self.current_model_id: str | None = ACTIVE_MODEL_ID
self.shared_stream_buffer = None
self.shared_stream_buffer_size = SHARED_STREAM_BUFFER_BYTES
@@ -690,7 +693,7 @@ class GPUSlot:
async def join_user(self, user_id: str, model_id: str = None) -> JoinAck:
"""Add a user to this GPU."""
if model_id is None:
model_id = DEFAULT_MODEL_ID
model_id = ACTIVE_MODEL_ID
# Reload model if a different one is requested
if model_id != self.current_model_id and model_id in MODEL_REGISTRY:
@@ -705,16 +708,23 @@ class GPUSlot:
self.connected_users.clear()
model_config = MODEL_REGISTRY[model_id]
reload_response = await self._send_command(Command(CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
try:
reload_response = await self._send_command(Command(
CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
except Exception:
self.current_model_id = None
raise
match reload_response:
case ReloadAck():
pass
case WorkerError(message=msg):
self.current_model_id = None
raise RuntimeError(f"Model reload failed: {msg}")
case _:
self.current_model_id = None
raise RuntimeError(f"Unexpected reload response: "
f"{type(reload_response).__name__}")
@@ -746,6 +756,10 @@ class GPUSlot:
segment_idx: int = 1,
image_path: str | None = None,
reset_conditioning: bool = False,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> dict[str, float]:
"""Execute a generation step for a specific user.
@@ -759,6 +773,9 @@ class GPUSlot:
segment_idx=segment_idx,
image_path=image_path,
reset_conditioning=bool(reset_conditioning),
frame_width=frame_width,
frame_height=frame_height,
num_frames=num_frames,
)
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
timeout=1800.0)
@@ -1,9 +1,9 @@
"""LTX2 model lifecycle and continuation conditioning.
"""LTX-2 model lifecycle and continuation conditioning.
Runs inside a GPU worker subprocess. Owns the model, the audio
encoder, and the per-session continuation state carried across
segments. Callers must set ``os.environ["CUDA_VISIBLE_DEVICES"]``
before constructing ``VideoGenerationWorker`` — all ``fastvideo.*``
before constructing ``LTX2GenerationBackend`` — all ``fastvideo.*``
imports are deferred to method bodies so nothing touches CUDA at
module import time.
"""
@@ -14,9 +14,6 @@ import gc
import os
import re
import time
from dataclasses import dataclass
from typing import Any
import numpy as np
import torch
@@ -35,6 +32,7 @@ from dreamverse.config import (
DREAMVERSE_LORA_STACK,
_resolve_lora_spec,
)
from dreamverse.generation_contracts import StepResult
# Multi-frame decoded continuation defaults from
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
@@ -80,22 +78,6 @@ def _reset_lora_registry(worker) -> dict:
return {"status": "lora_registry_reset"}
@dataclass
class StepResult:
"""Output of one generation step.
``head_trim_frames`` / ``head_trim_audio_frames`` are derived here
so downstream AV streaming never needs to import conditioning
constants.
"""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict
head_trim_frames: int
head_trim_audio_frames: int
class ContinuationState:
"""Per-session video + audio conditioning carried across segments."""
@@ -202,7 +184,7 @@ class ContinuationState:
self.audio_latents = latents.detach().clone().cpu()
class VideoGenerationWorker:
class LTX2GenerationBackend:
"""Single-GPU LTX2 generator with continuation state.
Caller must set ``os.environ["CUDA_VISIBLE_DEVICES"]`` before
@@ -472,6 +454,10 @@ class VideoGenerationWorker:
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Execute one generation step; snapshot state for the next segment."""
timings: dict = {}
@@ -482,9 +468,9 @@ class VideoGenerationWorker:
prompt=prompt,
negative_prompt="",
save_video=False,
height=FRAME_HEIGHT,
width=FRAME_WIDTH,
num_frames=NUM_FRAMES,
height=frame_height or FRAME_HEIGHT,
width=frame_width or FRAME_WIDTH,
num_frames=num_frames or NUM_FRAMES,
fps=24,
num_inference_steps=NUM_INFERENCE_STEPS,
guidance_scale=1.0,
+2
View File
@@ -33,6 +33,7 @@ from dreamverse.routes.presets import (
prompt_config_router,
curated_presets_router,
)
from dreamverse.routes.creation import creation_router
from dreamverse.session.controller import SessionController
@@ -92,6 +93,7 @@ app.add_middleware(
app.include_router(build_health_router(lambda: runtime.gpu_pool))
app.include_router(internal_monitor_router)
app.include_router(prompt_config_router)
app.include_router(creation_router)
if DEVTOOLS_ENABLED:
app.include_router(curated_presets_router)
@@ -0,0 +1,302 @@
"""FastH3 model lifecycle and first-frame continuation for DreamVerse."""
from __future__ import annotations
import gc
import os
import time
from typing import TYPE_CHECKING, Any
import numpy as np
import torch
from dreamverse.config import DREAMVERSE_SP_SIZE
from dreamverse.generation_contracts import StepResult
if TYPE_CHECKING:
from PIL.Image import Image
def _required_config_str(model_config: dict, field_name: str) -> str:
"""Read one required non-empty string from a DreamVerse model profile."""
value = model_config.get(field_name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"FastH3 model configuration requires `{field_name}`.")
return value.strip()
class MiniMaxH3GenerationBackend:
"""Run the VSA data-free FastH3 adapter and retain one continuation frame."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
reserved_gib = torch.cuda.memory_reserved() / 1024**3
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
@staticmethod
def _configure_environment(attention_backend: str) -> None:
"""Apply the fixed boot-time switches from the FastH3 reference recipe."""
os.environ.update({
"FASTVIDEO_ATTENTION_BACKEND": attention_backend,
"FASTVIDEO_FA4": "1",
"FASTVIDEO_MINIMAX_H3_FUSIONS": "all",
"FASTVIDEO_VSA_SM100A": "0",
})
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
def initialize(self, model_config: dict | None = None) -> None:
"""Download the fixed Preview adapter and load the FastH3 generator.
The model profile owns the base checkpoint, adapter file, attention
backend, and generation geometry. The backend translates that profile
into FastVideo's typed generator configuration.
"""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("FastH3 initialization requires a model configuration.")
if self.generator is not None:
self.generator.shutdown()
self.generator = None
gc.collect()
torch.cuda.empty_cache()
self.clear_conditioning()
model_path = _required_config_str(self.model_config, "model_path")
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
from huggingface_hub import hf_hub_download
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig,
ComponentConfig,
EngineConfig,
GeneratorConfig,
OffloadConfig,
ParallelismConfig,
PipelineSelection,
)
adapter_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
experimental = {
"attention_backend": attention_backend,
"inference_torch_compile": attention_backend == "FLASH_ATTN",
"vae_parallel_decode": True,
"vae_parallel_decode_strategy": "gather",
}
if attention_backend == "VIDEO_SPARSE_ATTN_H3":
experimental.update({
"VSA_sparsity": 0.9,
"VSA_tile_size": 64,
})
generator_config = GeneratorConfig(
model_path=model_path,
pipeline=PipelineSelection(
components=ComponentConfig(lora_path=adapter_path, lora_strength=1.0),
experimental=experimental,
),
engine=EngineConfig(
num_gpus=DREAMVERSE_SP_SIZE,
parallelism=ParallelismConfig(tp_size=1, sp_size=DREAMVERSE_SP_SIZE),
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
text_encoder=True,
image_encoder=True,
vae=True,
pin_cpu_memory=True,
),
compile=CompileConfig(enabled=False, vae_enabled=True),
use_fsdp_inference=False,
),
)
print(f"[GPU {self.gpu_id}] Loading FastH3 model: {model_path}")
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
self.generator = VideoGenerator.from_config(generator_config)
print(f"[GPU {self.gpu_id}] FastH3 loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
"""Release the FastVideo generator and cached continuation image."""
self.clear_conditioning()
if self.generator is not None:
self.generator.shutdown()
self.generator = None
def clear_conditioning(self) -> None:
"""Release the first-frame image retained for the next segment."""
if self.continuation_image is not None:
self.continuation_image.close()
self.continuation_image = None
@staticmethod
def _load_rgb_image(image_path: str) -> Image:
"""Load an image into an independent RGB buffer with no open file handle."""
from PIL import Image
with Image.open(image_path) as image:
return image.convert("RGB").copy()
def _select_conditioning_image(
self,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> tuple[Image | None, bool]:
"""Select the initial upload or retained last frame for one segment."""
if reset_conditioning:
self.clear_conditioning()
if segment_idx > 1 and self.continuation_image is not None:
return self.continuation_image.copy(), True
if segment_idx > 1 and not reset_conditioning:
raise RuntimeError(f"FastH3 segment {segment_idx} requires a retained continuation frame.")
if segment_idx == 1 and image_path:
return self._load_rgb_image(image_path), False
return None, False
def _build_request(self, prompt: str, conditioning_image: Image | None):
"""Build the typed FastVideo request owned by the FastH3 profile."""
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
return GenerationRequest(
prompt=prompt,
negative_prompt="",
inputs=InputConfig(pil_image=conditioning_image),
sampling=SamplingConfig(
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
num_frames=int(self.model_config["num_frames"]),
fps=24,
num_inference_steps=int(self.model_config["num_inference_steps"]),
guidance_scale=1.0,
batch_cfg=False,
seed=int(self.model_config["seed"]),
),
output=OutputConfig(save_video=False, return_frames=True),
)
def _save_continuation_frame(self, frames: list) -> None:
"""Retain the last decoded frame as first-frame conditioning."""
from PIL import Image
self.clear_conditioning()
self.continuation_image = Image.fromarray(np.ascontiguousarray(frames[-1])).convert("RGB")
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Generate one synchronized FastH3 segment and retain its last frame.
Later segments use MiniMax H3's first-frame-to-video path. The first
conditioned frame and its matching audio duration are trimmed before
streaming so adjacent segments do not duplicate media.
"""
del frame_width, frame_height, num_frames
if self.generator is None:
raise RuntimeError("FastH3 generator is not initialized.")
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
request = self._build_request(prompt, conditioning_image)
started = time.perf_counter()
try:
result = self.generator.generate(request)
finally:
if conditioning_image is not None:
conditioning_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
if isinstance(result, list):
raise RuntimeError("FastH3 returned multiple results for one DreamVerse segment.")
frames = result.frames
if not isinstance(frames, list) or not frames:
raise RuntimeError("FastH3 generation did not return decoded frames.")
audio = result.audio
audio_sample_rate = result.audio_sample_rate
if audio is not None and audio_sample_rate is None:
raise RuntimeError("FastH3 returned audio without an audio sample rate.")
save_started = time.perf_counter()
self._save_continuation_frame(frames)
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
timings = {
"generation_ms": generation_ms,
"generation_time_ms": float(result.generation_time or 0.0) * 1000.0,
"save_conditioning_ms": save_conditioning_ms,
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
}
trim_frames = 1 if uses_continuation else 0
print(f"[GPU {self.gpu_id}] FastH3 segment {segment_idx}: "
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
f"save_conditioning={save_conditioning_ms:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
return StepResult(
frames=frames,
audio=audio,
audio_sample_rate=audio_sample_rate,
timings=timings,
head_trim_frames=trim_frames,
head_trim_audio_frames=trim_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
"""Compile the FastH3 text and first-frame paths before readiness."""
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] FastH3 startup warmup starting "
"(synthetic segments: text-to-video, first-frame-to-video)")
started = time.perf_counter()
text_result = self.generate_step(
warmup_prompt,
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
first_frame_result = self.generate_step(
warmup_prompt,
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
total_ms = (time.perf_counter() - started) * 1000.0
self.clear_conditioning()
text_ms = float(text_result.timings.get("e2e_latency_ms", 0.0))
first_frame_ms = float(first_frame_result.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] FastH3 startup warmup complete: "
f"text_to_video={text_ms:.0f}ms, "
f"first_frame_to_video={first_frame_ms:.0f}ms, "
f"total={total_ms:.0f}ms")
return {
"warmup_text_to_video_ms": text_ms,
"warmup_first_frame_to_video_ms": first_frame_ms,
"warmup_total_ms": total_ms,
}
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
"""Reject runtime LoRA mutation because FastH3 uses one startup adapter."""
del stack
raise RuntimeError("FastH3 uses its fixed startup adapter and does not support runtime LoRA changes.")
+65 -5
View File
@@ -31,6 +31,8 @@ from fastapi.staticfiles import StaticFiles
from dreamverse._deps import require_dreamverse_runtime_deps
from dreamverse.config import FRONTEND_STATIC_DIR_CANDIDATES, GENERATION_SEGMENT_CAP
from dreamverse.creation_capabilities import lobby_capabilities_as_dict
from dreamverse.session_creation_config import parse_session_creation_config, validate_generation_mode_assets
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
LATENCY_MS = 200
@@ -225,6 +227,11 @@ async def prompt_system_config():
}
@app.get("/creation-capabilities")
async def creation_capabilities():
return lobby_capabilities_as_dict()
@app.get("/curated-presets")
async def curated_presets():
presets = [
@@ -290,6 +297,8 @@ async def websocket_endpoint(websocket: WebSocket):
send_lock = asyncio.Lock()
stop_event = asyncio.Event()
session_init_image = None
session_last_frame_image = None
session_creation_config = None
async def ws_send_json(payload: dict) -> None:
async with send_lock:
@@ -348,6 +357,7 @@ async def websocket_endpoint(websocket: WebSocket):
try:
session_init_image = persist_session_init_image(init_data.get("initial_image"))
session_last_frame_image = persist_session_init_image(init_data.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
@@ -356,13 +366,31 @@ async def websocket_endpoint(websocket: WebSocket):
await websocket.close(code=1003, reason="Invalid initial image")
return
try:
session_creation_config = parse_session_creation_config(init_data)
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid creation config")
return
timeout_task = asyncio.create_task(session_timeout())
await ws_send_json({
gpu_assigned_payload: dict[str, object] = {
"type": "gpu_assigned",
"gpu_id": 0,
"session_timeout": SESSION_TIMEOUT_SECONDS,
})
}
if session_creation_config is not None:
gpu_assigned_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(gpu_assigned_payload)
raw_prompt_queue: asyncio.Queue[PromptSubmission] = asyncio.Queue()
ready_prompt_queue: asyncio.Queue[ReadyPrompt] = asyncio.Queue()
@@ -391,8 +419,16 @@ async def websocket_endpoint(websocket: WebSocket):
if previous_session_image is not None:
cleanup_session_init_image(previous_session_image)
def replace_last_frame_image(last_frame_payload: object) -> None:
nonlocal session_last_frame_image
next_last_frame_image = persist_session_init_image(last_frame_payload)
previous_last_frame_image = session_last_frame_image
session_last_frame_image = next_last_frame_image
if previous_last_frame_image is not None:
cleanup_session_init_image(previous_last_frame_image)
async def send_stream_start(seed_reason: str) -> None:
await ws_send_json({
stream_start_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
@@ -400,8 +436,15 @@ async def websocket_endpoint(websocket: WebSocket):
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": 0,
})
"generation_segment_cap": (
session_creation_config.generation_segment_cap
if session_creation_config is not None
else GENERATION_SEGMENT_CAP
),
}
if session_creation_config is not None:
stream_start_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(stream_start_payload)
if seed_reason == "init":
await ws_send_json({
"type": "seed_prompts_updated",
@@ -509,6 +552,7 @@ async def websocket_endpoint(websocket: WebSocket):
nonlocal project_active
nonlocal project_stream_started
nonlocal pending_project_end
nonlocal session_creation_config
next_initial_rollout_prompt = str(payload.get("initial_rollout_prompt") or "").strip()
next_preset_id = str(payload.get("preset_id") or "").strip()
@@ -520,6 +564,21 @@ async def websocket_endpoint(websocket: WebSocket):
try:
replace_session_image(payload.get("initial_image"))
replace_last_frame_image(payload.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
return False
try:
session_creation_config = parse_session_creation_config(payload)
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
@@ -1182,6 +1241,7 @@ async def websocket_endpoint(websocket: WebSocket):
finally:
stop_event.set()
cleanup_session_init_image(session_init_image)
cleanup_session_init_image(session_last_frame_image)
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
@@ -0,0 +1,14 @@
"""Creation studio capability routes."""
from __future__ import annotations
from fastapi import APIRouter
from dreamverse.creation_capabilities import lobby_capabilities_as_dict
creation_router = APIRouter(tags=["creation"])
@creation_router.get("/creation-capabilities")
async def creation_capabilities() -> dict[str, object]:
return lobby_capabilities_as_dict()
+106 -45
View File
@@ -27,10 +27,11 @@ from typing import TYPE_CHECKING
from fastapi import WebSocket, WebSocketDisconnect
from dreamverse.gpu_pool import GPUSlot
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
from dreamverse.session_creation_config import parse_session_creation_config, validate_generation_mode_assets
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
GENERATION_SEGMENT_CAP,
PROMPT_AUTO_SLEEP_MS,
PROMPT_AUTO_TIMEOUT_MS,
@@ -156,6 +157,9 @@ class SessionController:
prompt_worker_task: asyncio.Task | None = None
rewrite_seed_prompts_task: asyncio.Task | None = None
session_init_image = None
session_last_frame_image = None
session_creation_config = None
session_generation_segment_cap = GENERATION_SEGMENT_CAP
async def session_timeout():
"""Close the session after timeout."""
@@ -237,6 +241,7 @@ class SessionController:
try:
session_init_image = persist_session_init_image(init_data.get("initial_image"))
session_last_frame_image = persist_session_init_image(init_data.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
@@ -245,6 +250,22 @@ class SessionController:
await websocket.close(code=1003, reason="Invalid initial image")
return
try:
session_creation_config = parse_session_creation_config(init_data)
session_generation_segment_cap = session_creation_config.generation_segment_cap
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid creation config")
return
if preset_id:
print(f"Client {client_id[:8]} selected preset: {preset_id} "
f"label={preset_label or '(unset)'} "
@@ -256,6 +277,16 @@ class SessionController:
if session_init_image is not None:
print(f"Client {client_id[:8]} uploaded initial image: "
f"{session_init_image.display_name}")
if session_last_frame_image is not None:
print(f"Client {client_id[:8]} uploaded last frame image: "
f"{session_last_frame_image.display_name}")
if session_creation_config is not None:
print(f"Client {client_id[:8]} creation config: "
f"model={session_creation_config.model_id}, "
f"mode={session_creation_config.generation_mode}, "
f"size={session_creation_config.frame_width}x{session_creation_config.frame_height}, "
f"duration={session_creation_config.duration_sec}s, "
f"segment_cap={session_creation_config.generation_segment_cap}")
# Acquire a GPU slot.
gpu_id, slot = await self.gpu_pool.acquire(client_id, websocket)
@@ -264,14 +295,20 @@ class SessionController:
timeout_task = asyncio.create_task(session_timeout())
# Join the engine on this GPU.
await slot.join_user(client_id, model_id=DEFAULT_MODEL_ID)
await slot.join_user(
client_id,
model_id=session_creation_config.model_id if session_creation_config is not None else ACTIVE_MODEL_ID,
)
# Notify client they're connected to a GPU.
await ws_send_json({
gpu_assigned_payload: dict[str, object] = {
"type": "gpu_assigned",
"gpu_id": gpu_id,
"session_timeout": SESSION_TIMEOUT_SECONDS,
})
}
if session_creation_config is not None:
gpu_assigned_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(gpu_assigned_payload)
await log_event(
"gpu_assigned",
{
@@ -315,6 +352,14 @@ class SessionController:
if previous_session_init_image is not None:
cleanup_session_init_image(previous_session_init_image)
def replace_last_frame_image(last_frame_payload: object) -> None:
nonlocal session_last_frame_image
next_last_frame_image = persist_session_init_image(last_frame_payload)
previous_last_frame_image = session_last_frame_image
session_last_frame_image = next_last_frame_image
if previous_last_frame_image is not None:
cleanup_session_init_image(previous_last_frame_image)
async def schedule_simple_generate_request(payload: dict[str, object]) -> None:
nonlocal preset_id
nonlocal preset_label
@@ -452,6 +497,8 @@ class SessionController:
nonlocal project_active
nonlocal project_stream_started
nonlocal pending_project_end
nonlocal session_creation_config
nonlocal session_generation_segment_cap
next_initial_rollout_prompt = str(payload.get("initial_rollout_prompt") or "").strip()
next_enhancement_enabled = bool(payload.get("enhancement_enabled", True))
@@ -498,6 +545,22 @@ class SessionController:
try:
replace_session_init_image(payload.get("initial_image"))
replace_last_frame_image(payload.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
return False
try:
session_creation_config = parse_session_creation_config(payload)
session_generation_segment_cap = session_creation_config.generation_segment_cap
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
@@ -941,7 +1004,7 @@ class SessionController:
"segment_cap":
_resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
continue
@@ -1281,27 +1344,22 @@ class SessionController:
))
else:
project_stream_started = True
await ws_send_json({
"type":
"ltx2_stream_start",
"total_segments":
len(curated_prompts),
"preset_id":
preset_id,
"stream_mode":
"av_fmp4",
"live_mode":
True,
"loop_generation_enabled":
loop_generation_enabled,
"loop_iteration":
loop_iteration,
"generation_segment_cap":
_resolve_generation_segment_cap(
stream_start_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
"stream_mode": "av_fmp4",
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": _resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
}
if session_creation_config is not None:
stream_start_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(stream_start_payload)
await ws_send_json({
"type": "seed_prompts_updated",
"prompts": seed_prompt_memory,
@@ -1340,27 +1398,22 @@ class SessionController:
loop_iteration += 1
project_stream_started = True
await ws_send_json({
"type":
"ltx2_stream_start",
"total_segments":
len(curated_prompts),
"preset_id":
preset_id,
"stream_mode":
"av_fmp4",
"live_mode":
True,
"loop_generation_enabled":
loop_generation_enabled,
"loop_iteration":
loop_iteration,
"generation_segment_cap":
_resolve_generation_segment_cap(
restart_stream_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
"stream_mode": "av_fmp4",
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": _resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
}
if session_creation_config is not None:
restart_stream_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(restart_stream_payload)
if nonlocal_reason == "loop_restart":
await ws_send_json({
"type": "loop_restarted",
@@ -1389,13 +1442,14 @@ class SessionController:
pending_simple_prompt_submission = None
if (not single_clip_mode and not generation_cap_blocked and not rollout_waiting_for_rewrite
and GENERATION_SEGMENT_CAP > 0 and generated_segment_count >= GENERATION_SEGMENT_CAP):
and session_generation_segment_cap > 0
and generated_segment_count >= session_generation_segment_cap):
loop_generation_enabled = False
rollout_waiting_for_rewrite = True
_main_print(
"INFO",
f"Segment cap reached for client {client_id[:8]} "
f"(cap_segments={GENERATION_SEGMENT_CAP}, "
f"(cap_segments={session_generation_segment_cap}, "
f"generated_segments={generated_segment_count}); "
"waiting for rollout rewrite",
)
@@ -1620,6 +1674,9 @@ class SessionController:
pending_reset_conditioning = False
step_image_path = (str(session_init_image.file_path)
if segment_idx == 1 and session_init_image is not None else None)
step_frame_width = session_creation_config.frame_width if session_creation_config is not None else None
step_frame_height = session_creation_config.frame_height if session_creation_config is not None else None
step_num_frames = session_creation_config.num_frames if session_creation_config is not None else None
step_task = asyncio.create_task(
slot.user_step(
client_id,
@@ -1627,6 +1684,9 @@ class SessionController:
segment_idx=segment_idx,
image_path=step_image_path,
reset_conditioning=step_reset_conditioning,
frame_width=step_frame_width,
frame_height=step_frame_height,
num_frames=step_num_frames,
))
segment_generation_active = True
try:
@@ -1808,3 +1868,4 @@ class SessionController:
await self.gpu_pool.release(client_id)
finally:
cleanup_session_init_image(session_init_image)
cleanup_session_init_image(session_last_frame_image)
@@ -0,0 +1,140 @@
from __future__ import annotations
from dataclasses import dataclass
from dreamverse.config import FRAME_HEIGHT, FRAME_WIDTH, GENERATION_SEGMENT_CAP, MODEL_REGISTRY, NUM_FRAMES
from dreamverse.creation_capabilities import validate_lobby_creation_config
LTX_LOBBY_MODEL_IDS = frozenset(MODEL_REGISTRY.keys())
SUPPORTED_GENERATION_MODES = frozenset({"t2va", "fl2va", "ref2va"})
SUPPORTED_ASPECT_RATIOS = frozenset({"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"})
SUPPORTED_RESOLUTIONS = frozenset({"480p", "720p", "1080p", "4k"})
SEGMENT_DURATION_SEC = 5
@dataclass(frozen=True)
class SessionCreationConfig:
model_id: str
generation_mode: str
aspect_ratio: str
resolution: str
duration_sec: int
frame_width: int
frame_height: int
num_frames: int
generation_segment_cap: int
def as_dict(self) -> dict[str, object]:
return {
"model_id": self.model_id,
"generation_mode": self.generation_mode,
"aspect_ratio": self.aspect_ratio,
"resolution": self.resolution,
"duration_sec": self.duration_sec,
"frame_width": self.frame_width,
"frame_height": self.frame_height,
"num_frames": self.num_frames,
"generation_segment_cap": self.generation_segment_cap,
}
def _round_to_multiple(value: float, multiple: int = 32) -> int:
rounded = int(round(value / multiple)) * multiple
return max(multiple, rounded)
def _resolution_base(resolution: str) -> int:
return {
"480p": 480,
"720p": 720,
"1080p": 1080,
"4k": 2160,
}.get(resolution, 720)
def resolve_frame_size(aspect_ratio: str, resolution: str) -> tuple[int, int]:
if aspect_ratio == "16:9" and resolution == "1080p":
return FRAME_WIDTH, FRAME_HEIGHT
base = _resolution_base(resolution)
width_ratio, height_ratio = {
"21:9": (21, 9),
"16:9": (16, 9),
"4:3": (4, 3),
"1:1": (1, 1),
"3:4": (3, 4),
"9:16": (9, 16),
}.get(aspect_ratio, (16, 9))
if width_ratio >= height_ratio:
height = _round_to_multiple(base)
width = _round_to_multiple(height * width_ratio / height_ratio)
else:
width = _round_to_multiple(base)
height = _round_to_multiple(width * height_ratio / width_ratio)
return width, height
def duration_sec_to_segment_cap(duration_sec: int, *, global_cap: int = GENERATION_SEGMENT_CAP) -> int:
requested = max(1, int(round(duration_sec / SEGMENT_DURATION_SEC + 0.0001)))
if global_cap <= 0:
return requested
return max(1, min(requested, global_cap))
def parse_session_creation_config(payload: dict[str, object]) -> SessionCreationConfig:
raw_model_id = str(payload.get("model_id") or "").strip()
model_id = raw_model_id if raw_model_id in LTX_LOBBY_MODEL_IDS else "fast-ltx23"
generation_mode = str(payload.get("generation_mode") or "t2va").strip()
if generation_mode not in SUPPORTED_GENERATION_MODES:
raise ValueError(f"Unsupported generation_mode: {generation_mode}")
aspect_ratio = str(payload.get("aspect_ratio") or "16:9").strip()
if aspect_ratio not in SUPPORTED_ASPECT_RATIOS:
raise ValueError(f"Unsupported aspect_ratio: {aspect_ratio}")
resolution = str(payload.get("resolution") or "720p").strip()
if resolution not in SUPPORTED_RESOLUTIONS:
raise ValueError(f"Unsupported resolution: {resolution}")
try:
duration_sec = int(payload.get("duration_sec") or SEGMENT_DURATION_SEC)
except (TypeError, ValueError) as exc:
raise ValueError("duration_sec must be an integer.") from exc
if duration_sec not in {5, 10, 15}:
raise ValueError("duration_sec must be 5, 10, or 15.")
validate_lobby_creation_config(
model_id=model_id,
generation_mode=generation_mode,
aspect_ratio=aspect_ratio,
resolution=resolution,
duration_sec=duration_sec,
)
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unsupported model_id: {model_id}")
frame_width, frame_height = resolve_frame_size(aspect_ratio, resolution)
return SessionCreationConfig(
model_id=model_id,
generation_mode=generation_mode,
aspect_ratio=aspect_ratio,
resolution=resolution,
duration_sec=duration_sec,
frame_width=frame_width,
frame_height=frame_height,
num_frames=NUM_FRAMES,
generation_segment_cap=duration_sec_to_segment_cap(duration_sec),
)
def validate_generation_mode_assets(
generation_mode: str,
*,
has_initial_image: bool,
has_last_frame_image: bool,
) -> None:
if generation_mode == "ref2va" and not has_initial_image:
raise ValueError("Ref2VA mode requires a reference image.")
@@ -2,13 +2,14 @@ from __future__ import annotations
import importlib.util
from pathlib import Path
from types import ModuleType
import pytest
SERVER_DIR = Path(__file__).resolve().parents[1]
def _load_config_module():
def _load_config_module() -> ModuleType:
spec = importlib.util.spec_from_file_location(
"server_config_test_module",
SERVER_DIR / "config.py",
@@ -20,7 +21,7 @@ def _load_config_module():
return module
def _set_required_prompt_keys(monkeypatch):
def _set_required_prompt_keys(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
@@ -150,3 +151,38 @@ def test_config_rejects_invalid_prompt_provider(monkeypatch):
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_PROMPT_PROVIDER"):
_load_config_module()
def test_config_registers_vsa_datafree_fasth3_profile(monkeypatch):
"""The FastH3 registry entry owns the complete fixed Preview recipe."""
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.MODEL_REGISTRY["fast-h3"] == {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
def test_config_uses_fasth3_sequence_parallel_default(monkeypatch):
"""Selecting FastH3 defaults DreamVerse to its four-GPU topology."""
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "fast-h3")
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
module = _load_config_module()
assert module.ACTIVE_MODEL_ID == "fast-h3"
assert module.MODEL_CONFIG["generation_backend"] == "minimax_h3"
assert module.DREAMVERSE_SP_SIZE == 4
@@ -0,0 +1,74 @@
import pytest
from dreamverse.creation_capabilities import (
capabilities_for_model,
lobby_capabilities_as_dict,
validate_lobby_creation_config,
)
def test_lobby_capabilities_include_all_registry_models():
caps = lobby_capabilities_as_dict()
assert set(caps["model_ids"]) == {"fast-ltx2", "fast-ltx23", "fast-h3"}
assert "fl2va" not in caps["generation_modes"]
assert "4k" not in caps["resolutions"]
def test_fast_h3_capabilities_use_fixed_geometry():
h3_caps = capabilities_for_model("fast-h3")
assert h3_caps.generation_modes == frozenset({"t2va", "ref2va"})
assert h3_caps.aspect_ratios == frozenset({"16:9"})
assert h3_caps.resolutions == frozenset({"720p"})
def test_validate_lobby_creation_config_accepts_supported_t2va():
validate_lobby_creation_config(
model_id="fast-ltx23",
generation_mode="t2va",
aspect_ratio="16:9",
resolution="1080p",
duration_sec=5,
)
def test_validate_lobby_creation_config_accepts_fast_h3():
validate_lobby_creation_config(
model_id="fast-h3",
generation_mode="ref2va",
aspect_ratio="16:9",
resolution="720p",
duration_sec=10,
)
def test_validate_lobby_creation_config_rejects_fl2va():
with pytest.raises(ValueError, match="FL2VA"):
validate_lobby_creation_config(
model_id="fast-ltx23",
generation_mode="fl2va",
aspect_ratio="16:9",
resolution="720p",
duration_sec=5,
)
def test_validate_lobby_creation_config_rejects_4k():
with pytest.raises(ValueError, match="Unsupported resolution"):
validate_lobby_creation_config(
model_id="fast-ltx2",
generation_mode="t2va",
aspect_ratio="16:9",
resolution="4k",
duration_sec=10,
)
def test_validate_lobby_creation_config_rejects_invalid_h3_aspect():
with pytest.raises(ValueError, match="Unsupported aspect_ratio"):
validate_lobby_creation_config(
model_id="fast-h3",
generation_mode="t2va",
aspect_ratio="9:16",
resolution="720p",
duration_sec=5,
)
@@ -0,0 +1,9 @@
from dreamverse.generation_worker import _create_generation_backend
from dreamverse.ltx2_generation import LTX2GenerationBackend
def test_create_generation_backend_ltx2_module_import():
backend = _create_generation_backend("ltx2", gpu_id=3)
assert isinstance(backend, LTX2GenerationBackend)
assert backend.gpu_id == 3
@@ -63,6 +63,14 @@ def test_get_available_gpus_defaults_to_first_visible_device(monkeypatch):
assert gpu_pool.get_available_gpus() == [3]
def test_get_available_gpus_defaults_to_active_model_sequence_parallel_size(monkeypatch):
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "0,1,2,3,4")
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
monkeypatch.setattr(gpu_pool, "DREAMVERSE_SP_SIZE", 4)
assert gpu_pool.get_available_gpus() == [0, 1, 2, 3]
def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "zero")
@@ -71,6 +79,23 @@ def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
gpu_pool.get_available_gpus()
def test_join_user_failed_reload_marks_model_uninitialized(monkeypatch):
"""A failed model reload forces the next join to reload a model."""
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.current_model_id = "fast-ltx2"
async def fake_send_command(command, timeout):
del command, timeout
return gpu_pool.WorkerError(user_id="__reload__", message="load failed")
monkeypatch.setattr(slot, "_send_command", fake_send_command)
with pytest.raises(RuntimeError, match="Model reload failed"):
asyncio.run(slot.join_user("client-id", model_id="fast-h3"))
assert slot.current_model_id is None
def test_send_command_raises_on_worker_death():
"""A worker that consumes a command and exits without replying must
surface as RuntimeError via sentinel detection, not after the long
@@ -92,9 +117,9 @@ def test_send_command_raises_on_worker_death():
ready = resp_q.get(timeout=30.0)
assert ready == "READY"
async def runner():
async def runner() -> None:
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.process = proc
slot.process = proc # type: ignore[assignment]
slot.command_queue = cmd_q
slot.response_queue = resp_q
@@ -17,11 +17,11 @@ FORBIDDEN_PREFIXES = (
)
ALLOWED_INTERNAL_IMPORTS = {
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.audio.ltx2_audio_processing",
),
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.loader.component_loader",
),
}
@@ -0,0 +1,247 @@
from __future__ import annotations
import os
from types import SimpleNamespace
from typing import Any
import numpy as np
import pytest
import dreamverse.generation_worker as generation_worker
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
FASTH3_MODEL_CONFIG = {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
class _RecordingGenerator:
"""Record typed requests and return small synchronized media fixtures."""
def __init__(self) -> None:
self.requests: list[Any] = []
self.conditioning_pixels: list[np.ndarray | None] = []
def generate(self, request):
"""Capture the request and return two tiny video frames with audio."""
self.requests.append(request)
conditioning_image = request.inputs.pil_image
self.conditioning_pixels.append(
None if conditioning_image is None else np.asarray(conditioning_image).copy())
frames = [
np.full((2, 3, 3), 10, dtype=np.uint8),
np.full((2, 3, 3), 20, dtype=np.uint8),
]
return SimpleNamespace(
frames=frames,
audio=np.zeros((2, 16), dtype=np.float32),
audio_sample_rate=44100,
generation_time=0.25,
)
def test_initialize_builds_vsa_datafree_fasth3_generator(monkeypatch):
"""Initialization translates the DreamVerse profile into typed FastVideo config."""
from fastvideo import VideoGenerator
captured = {}
fake_generator = SimpleNamespace(shutdown=lambda: None)
def fake_from_config(config):
captured["config"] = config
return fake_generator
def fake_download(**kwargs):
captured["download"] = kwargs
return f"/models/{kwargs['filename']}"
monkeypatch.setattr("huggingface_hub.hf_hub_download", fake_download)
monkeypatch.setattr(VideoGenerator, "from_config", fake_from_config)
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
monkeypatch.setenv("FASTVIDEO_FA4", "0")
monkeypatch.setenv("FASTVIDEO_MINIMAX_H3_FUSIONS", "0")
monkeypatch.setenv("FASTVIDEO_VSA_SM100A", "1")
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
backend = MiniMaxH3GenerationBackend(gpu_id=0)
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
backend.initialize(FASTH3_MODEL_CONFIG)
config = captured["config"]
assert captured["download"] == {
"repo_id": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"filename": "vsa-datafree/adapter_model.safetensors",
}
assert config.model_path == "MiniMaxAI/MiniMax-H3"
assert config.pipeline.components.lora_path.endswith("vsa-datafree/adapter_model.safetensors")
assert config.pipeline.components.lora_strength == 1.0
assert config.pipeline.experimental == {
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"inference_torch_compile": False,
"vae_parallel_decode": True,
"vae_parallel_decode_strategy": "gather",
"VSA_sparsity": 0.9,
"VSA_tile_size": 64,
}
assert config.engine.num_gpus == 4
assert config.engine.parallelism.tp_size == 1
assert config.engine.parallelism.sp_size == 4
assert config.engine.offload.dit is False
assert config.engine.offload.dit_layerwise is False
assert config.engine.offload.text_encoder is True
assert config.engine.offload.vae is True
assert config.engine.compile.vae_enabled is True
assert config.engine.use_fsdp_inference is False
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "VIDEO_SPARSE_ATTN_H3"
assert os.environ["FASTVIDEO_FA4"] == "1"
assert os.environ["FASTVIDEO_MINIMAX_H3_FUSIONS"] == "all"
assert os.environ["FASTVIDEO_VSA_SM100A"] == "0"
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
def test_initialize_selects_declared_generation_backend(monkeypatch):
"""The GPU worker constructs the backend that the active model profile declares."""
from unittest.mock import Mock
selected_backend = Mock()
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: selected_backend,
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend_name == "minimax_h3"
assert worker.backend is selected_backend
selected_backend.initialize.assert_called_once_with(FASTH3_MODEL_CONFIG)
def test_initialize_failure_clears_backend_ownership(monkeypatch):
"""A failed family change leaves the GPU worker explicitly uninitialized."""
ltx_backend = SimpleNamespace(initialize=lambda config: None, shutdown=lambda: None)
def fail_initialize(config):
del config
raise RuntimeError("load failed")
fasth3_backend = SimpleNamespace(
initialize=fail_initialize,
shutdown=lambda: None,
)
backends = {
"ltx2": ltx_backend,
"minimax_h3": fasth3_backend,
}
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: backends[backend_name],
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize({"generation_backend": "ltx2"})
with pytest.raises(RuntimeError, match="load failed"):
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend is None
assert worker.backend_name is None
assert worker.model_config == {"generation_backend": "ltx2"}
def test_generate_step_uses_last_frame_for_continuation(monkeypatch):
"""A later segment receives the prior segment's last decoded frame."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
first_result = backend.generate_step(
"first prompt",
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
second_result = backend.generate_step(
"second prompt",
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
first_request = backend.generator.requests[0]
assert first_request.inputs.pil_image is None
assert first_request.negative_prompt == ""
assert first_request.sampling.height == 768
assert first_request.sampling.width == 1344
assert first_request.sampling.num_frames == 124
assert first_request.sampling.num_inference_steps == 5
assert first_request.sampling.fps == 24
assert first_request.sampling.guidance_scale == 1.0
assert first_request.sampling.batch_cfg is False
assert first_request.sampling.seed == 1000
assert first_request.output.save_video is False
assert first_request.output.return_frames is True
assert backend.generator.conditioning_pixels[1].tolist() == np.full((2, 3, 3), 20).tolist()
assert first_result.head_trim_frames == 0
assert first_result.head_trim_audio_frames == 0
assert second_result.head_trim_frames == 1
assert second_result.head_trim_audio_frames == 1
assert second_result.audio_sample_rate == 44100
def test_generate_step_reset_uses_text_to_video_path(monkeypatch):
"""Resetting continuation produces an unconditioned text-to-video request."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
backend.generate_step("first prompt", 1, None, True)
reset_result = backend.generate_step("reset prompt", 2, None, True)
assert backend.generator.requests[-1].inputs.pil_image is None
assert reset_result.head_trim_frames == 0
assert reset_result.head_trim_audio_frames == 0
def test_generate_step_missing_continuation_frame(monkeypatch):
"""A later segment fails when no reset or retained frame defines its input."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
backend.generate_step("later prompt", 2, None, False)
assert backend.generator.requests == []
def test_warmup_exercises_text_and_first_frame_paths(monkeypatch):
"""Warmup covers both request shapes used by a DreamVerse session."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
timings = backend.warmup("warmup prompt")
assert backend.generator.conditioning_pixels[0] is None
assert backend.generator.conditioning_pixels[1] is not None
assert backend.continuation_image is None
assert "warmup_text_to_video_ms" in timings
assert "warmup_first_frame_to_video_ms" in timings
@@ -331,11 +331,11 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
]
def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
def test_enhance_prompt_uses_groq_when_it_returns_first():
enhancer = _build_staged_enhancer(
cerebras_payload=_chat_payload_with_content('{"prompt":"Cerebras prompt"}'),
groq_payload=_chat_payload_with_content('{"prompt":"Groq prompt"}'),
cerebras_delay_s=0.01,
cerebras_delay_s=0.08,
groq_delay_s=0.01,
)
@@ -346,12 +346,12 @@ def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
assert result.fallback_used is False
assert result.error is None
assert result.provider == "cerebras"
assert result.provider == "groq"
assert result.model == "gpt-test"
assert result.prompt == "Cerebras prompt"
assert result.prompt == "Groq prompt"
assert enhancer.get_provider_success_counts() == {
"cerebras": 1,
"groq": 0,
"cerebras": 0,
"groq": 1,
}
@@ -0,0 +1,89 @@
import pytest
from dreamverse.session_creation_config import (
duration_sec_to_segment_cap,
parse_session_creation_config,
resolve_frame_size,
validate_generation_mode_assets,
)
def test_parse_session_creation_config_defaults():
config = parse_session_creation_config({})
assert config.model_id == "fast-ltx23"
assert config.generation_mode == "t2va"
assert config.aspect_ratio == "16:9"
assert config.resolution == "720p"
assert config.duration_sec == 5
assert config.generation_segment_cap == 1
def test_parse_session_creation_config_maps_duration_to_segment_cap():
config = parse_session_creation_config(
{
"model_id": "fast-ltx2",
"generation_mode": "ref2va",
"aspect_ratio": "9:16",
"resolution": "480p",
"duration_sec": 15,
},
)
assert config.model_id == "fast-ltx2"
assert config.generation_mode == "ref2va"
assert config.generation_segment_cap == 3
assert config.frame_width >= 480
assert config.frame_height >= 480
def test_resolve_frame_size_uses_model_default_for_1080p_landscape():
width, height = resolve_frame_size("16:9", "1080p")
assert (width, height) == (1920, 1088)
def test_duration_sec_to_segment_cap_respects_global_cap():
assert duration_sec_to_segment_cap(15, global_cap=2) == 2
def test_parse_session_creation_config_accepts_fast_h3():
config = parse_session_creation_config(
{
"model_id": "fast-h3",
"generation_mode": "t2va",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration_sec": 10,
},
)
assert config.model_id == "fast-h3"
assert config.generation_mode == "t2va"
assert config.generation_segment_cap == 2
def test_parse_session_creation_config_rejects_fl2va():
with pytest.raises(ValueError, match="FL2VA"):
parse_session_creation_config(
{
"generation_mode": "fl2va",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration_sec": 5,
},
)
def test_parse_session_creation_config_rejects_4k():
with pytest.raises(ValueError, match="Unsupported resolution"):
parse_session_creation_config(
{
"generation_mode": "t2va",
"aspect_ratio": "16:9",
"resolution": "4k",
"duration_sec": 5,
},
)
def test_validate_generation_mode_assets():
validate_generation_mode_assets("t2va", has_initial_image=False, has_last_frame_image=False)
with pytest.raises(ValueError, match="Ref2VA"):
validate_generation_mode_assets("ref2va", has_initial_image=False, has_last_frame_image=False)
+3
View File
@@ -147,6 +147,9 @@ class UserStepPayload:
segment_idx: int
image_path: str | None
reset_conditioning: bool
frame_width: int | None = None
frame_height: int | None = None
num_frames: int | None = None
@dataclass(frozen=True)
+7 -7
View File
@@ -70,7 +70,7 @@
<mxCell id="dispatcher" value="command dispatcher&#xa;&#xa;gpu_worker_process() branches on&#xa;CommandType; asserts payload type&#xa;&#xa;INIT / WARMUP / RELOAD_MODEL&#xa;USER_JOIN / USER_STEP / USER_LEAVE&#xa;SHUTDOWN" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe6cc;strokeColor=#d79b00;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="120" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;video_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;ltx2_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="460" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="stream_av" value="stream_fmp4()&#xa;av_streaming.py:121&#xa;&#xa;trims overlap, pipes to ffmpeg,&#xa;publishes StreamInit / StreamChunk /&#xa;StreamComplete via callback" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#b1d8d7;strokeColor=#23445d;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
@@ -79,13 +79,13 @@
<mxCell id="Ot8BU52QTIb4EhyRSe7I-2" value="" style="edgeStyle=none;html=1;" parent="1" source="generator" target="Ot8BU52QTIb4EhyRSe7I-1" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;video_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;ltx2_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="460" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="ffmpeg" value="ffmpeg subprocess&#xa;&#xa;libx264 / *_nvenc&#xa;fragmented mp4" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="800" y="1300" width="260" height="100" as="geometry"/>
</mxCell>
<mxCell id="caches" value="ContinuationState&#xa;video_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="caches" value="ContinuationState&#xa;ltx2_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="120" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="e_cp" value="acquire" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="client" target="pool" edge="1">
@@ -250,7 +250,7 @@
<mxPoint x="690" y="880"/>
</Array>
</mxCell>
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender video_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender ltx2_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="39" y="-200" width="270" height="380" as="geometry"/>
</mxCell>
<mxCell id="Ot8BU52QTIb4EhyRSe7I-1" value="FastVideo video_generator" style="whiteSpace=wrap;html=1;fontSize=11;fillColor=#e1d5e7;strokeColor=#9673a6;rounded=1;" parent="1" vertex="1">
@@ -389,10 +389,10 @@
<mxCell id="cw2" value="from fastvideo.entrypoints.video_generator import VideoGenerator&#xa;from fastvideo.models.dits.ltx2 import DEFAULT_LTX2_AUDIO_*&#xa;&#xa;** Dreamverse reaches into fastvideo internals here **" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="695" width="550" height="60" as="geometry"/>
</mxCell>
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (video_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (ltx2_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="765" width="550" height="95" as="geometry"/>
</mxCell>
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (video_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (ltx2_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="870" width="550" height="55" as="geometry"/>
</mxCell>
<mxCell id="cw5" value="enter main worker loop → waits for JOIN_USER / USER_STEP / LEAVE" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
@@ -534,7 +534,7 @@
<mxPoint x="1040" y="1610" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (video_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (ltx2_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="955" y="1640" width="180" height="70" as="geometry"/>
</mxCell>
<mxCell id="dm11" value="10b. resp_q.put(MediaInit / MediaChunk / MediaComplete / StepComplete)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 85 KiB

After

Width:  |  Height:  |  Size: 85 KiB

+4 -12
View File
@@ -17,25 +17,17 @@ test.describe('frontend shell', () => {
});
});
test('composer hydrates with curated preset cards', async ({ page }) => {
test('composer hydrates with creation studio controls', async ({ page }) => {
await page.goto('/');
// The Continuation prompt textarea + Generate button render once
// the FE has hydrated against the public-FastVideo-backed
// dreamverse-server. Their presence proves the integration handshake
// (CORS, /curated-presets, /prompt-system-config) completed.
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible({ timeout: 30_000 });
const generate = page.getByRole('button', { name: /^generate$/i });
await expect(generate).toBeVisible({ timeout: 30_000 });
// Curated presets render as buttons; verify at least one is
// available — that's the only way the user can populate the
// Continuation textarea in the default composer.
const presetCard = page.getByRole('button', {
name: /LEGO Stormtroopers|Clay Stop-Motion|Boy & Dog|School Prank|Gamer Gets Banned|Small Town Oil Strike|Grandpa's Wing Costume/i,
}).first();
await expect(presetCard).toBeVisible({ timeout: 30_000 });
await expect(page.getByText('Direct scenes in seconds')).toBeVisible({ timeout: 30_000 });
await expect(page.getByRole('button', { name: /FastLTX/i }).first()).toBeVisible({ timeout: 30_000 });
await expect(continuation).toHaveAttribute('placeholder', /Describe your video or mention elements/i);
});
});
+4
View File
@@ -42,6 +42,10 @@ const nextConfig: NextConfig = {
source: '/prompt-system-config',
destination: `${backendUrl}/prompt-system-config`,
},
{
source: '/creation-capabilities',
destination: `${backendUrl}/creation-capabilities`,
},
{
source: '/curated-presets',
destination: `${backendUrl}/curated-presets`,
+2262 -113
View File
File diff suppressed because it is too large Load Diff
+4
View File
@@ -27,11 +27,15 @@
"@radix-ui/react-accordion": "^1.2.12",
"@radix-ui/react-checkbox": "^1.3.3",
"@radix-ui/react-collapsible": "^1.1.12",
"@radix-ui/react-dropdown-menu": "^2.1.24",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-popover": "^1.1.23",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.4.7",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-tabs": "^1.1.21",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"framer-motion": "^12.36.0",
+26 -1
View File
@@ -2,6 +2,7 @@
@import "tailwindcss";
@custom-variant dark (&:is(.dark *));
@custom-variant hover-capable (@media (hover: hover) and (pointer: fine));
:root {
color-scheme: light;
@@ -201,7 +202,7 @@ summary::-webkit-details-marker {
@apply border-border;
}
body {
@apply bg-background text-foreground;
@apply bg-background text-foreground antialiased;
}
}
@@ -220,6 +221,14 @@ summary::-webkit-details-marker {
.stroke-dash-anim {
animation: stroke-dash-animation 2s linear infinite;
}
.text-pretty {
text-wrap: pretty;
}
.text-balance {
text-wrap: balance;
}
}
@keyframes stroke-dash-animation {
@@ -287,6 +296,22 @@ html.theme-transition *::after {
/* —————————————— CUSTOM TAILWIND —————————————— */
@layer components {
.studio-control {
@apply transition-[border-color,background-color,box-shadow,color,transform] duration-150 focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40;
}
.studio-control-press {
@apply active:scale-[0.96] motion-reduce:active:scale-100;
}
.studio-hover-surface {
@apply hover-capable:hover:border-border hover-capable:hover:bg-accent/50;
}
.studio-media-outline {
@apply outline outline-1 -outline-offset-1 outline-black/10 dark:outline-white/10;
}
.debug {
@apply border border-rose-500;
}
+300 -136
View File
@@ -1,10 +1,20 @@
"use client";
import { Fragment, useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState } from "react";
import { Fragment, useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState, type Dispatch, type SetStateAction } from "react";
import { AnimatePresence, motion } from "framer-motion";
import { Download, Share2 } from "lucide-react";
import DevtoolsShell from "@/components/devtools/DevtoolsShell";
import MonitorPage from "@/components/MonitorPage";
import ChatBar from "@/components/ChatBar";
import CreationStudio from "@/components/creation/CreationStudio";
import {
buildMentionOptions,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type ResolutionId,
} from "@/lib/creationConfig";
import { toGenerationMode } from "@/lib/generationMode";
import type { SessionCreationConfig } from "@/components/creation/SessionCreationConfigPills";
import SessionTimeoutModal from "@/components/SessionTimeoutModal";
import Sidebar from "@/components/Sidebar";
import Header from "@/components/Header";
@@ -22,6 +32,18 @@ import {
buildRewritePromptWindowSnapshotFromPrompts,
normalizePromptWindowSnapshot,
} from "@/lib/prompts/promptWindowSnapshot";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
clampLobbySelectionToCapabilities,
parseLobbyCapabilitiesBundle,
resolveModelCapabilities,
validateLobbyCreationSelection,
type LobbyCapabilitiesBundle,
} from "@/lib/creationCapabilities";
import {
buildCreationInitPayload,
parseEchoedCreationConfig,
} from "@/lib/creationPayload";
import rawPresets from "@/lib/storyPresetsData";
import { cn } from "@/lib/utils";
import { createWebSocketConnection, detachAndCloseWebSocket } from "@/lib/ws/client";
@@ -78,98 +100,6 @@ function yieldToEventLoop(): Promise<void> {
return new Promise((r) => setTimeout(r, 0));
}
const HERO_WAVE_LIGHT = ["#2A4A98", "#4878E5", "#6FA0F2", "#B0BCC8", "#E8D99E", "#D8C844", "#C2A620"];
const HERO_WAVE_DARK = ["#143468", "#1E58B8", "#3892F0", "#80B8E8", "#B8D0EA", "#E2D498", "#DABB50"];
const HERO_TEXT = "Direct scenes in seconds";
function HeroTagline() {
const ref = useRef<HTMLHeadingElement>(null);
useEffect(() => {
const el = ref.current;
if (!el) return;
let rafId = 0;
function play() {
const chars = el!.querySelectorAll<HTMLSpanElement>("[data-char]");
if (!chars.length) return;
cancelAnimationFrame(rafId);
const isDark = document.documentElement.classList.contains("dark");
const colors = isDark ? HERO_WAVE_DARK : HERO_WAVE_LIGHT;
const waveLen = 10;
const total = chars.length + waveLen;
const duration = 1200;
const maxBlur = 3.5;
const start = performance.now();
function tick() {
const t = Math.min((performance.now() - start) / duration, 1);
const pos = t * total;
chars.forEach((ch, i) => {
const rel = pos - i;
if (rel >= 0 && rel < waveLen) {
const norm = rel / waveLen;
const ci = Math.floor(norm * colors.length);
ch.style.color = colors[Math.min(colors.length - 1, ci)];
let blur = 0;
if (norm < 0.25) {
blur = maxBlur * (1 - norm / 0.25);
} else if (norm > 0.75) {
blur = maxBlur * ((norm - 0.75) / 0.25);
}
ch.style.filter = blur > 0.1 ? `blur(${blur.toFixed(1)}px)` : "";
} else {
ch.style.color = "";
ch.style.filter = "";
}
});
if (t < 1) {
rafId = requestAnimationFrame(tick);
} else {
chars.forEach((ch) => {
ch.style.color = "";
ch.style.filter = "";
});
}
}
rafId = requestAnimationFrame(tick);
}
const initialDelay = setTimeout(play, 400);
const interval = setInterval(play, 5000);
return () => {
clearTimeout(initialDelay);
clearInterval(interval);
cancelAnimationFrame(rafId);
};
}, []);
return (
<h1 ref={ref} className="text-center text-3xl font-medium text-[#343537] dark:text-[#FAFAFB] sm:text-4xl">
{HERO_TEXT.split(" ").map((word, wi) => (
<Fragment key={wi}>
{wi > 0 && (
<span data-char className="transition-[color,filter] duration-150">
{" "}
</span>
)}
<span className="inline-flex">
{word.split("").map((char, ci) => (
<span key={ci} data-char className="inline-block transition-[color,filter] duration-150">
{char}
</span>
))}
</span>
</Fragment>
))}
</h1>
);
}
export default function Page() {
const storesRef = useRef<PageStores | null>(null);
if (!storesRef.current) {
@@ -323,8 +253,33 @@ export default function Page() {
const [ttffValueMs, setTtffValueMs] = useState<number | null>(null);
const ttffIntervalRef = useRef<ReturnType<typeof setInterval> | null>(null);
const pendingInitialPromptRef = useRef("");
const referenceFileRef = useRef<File | null>(null);
const firstFrameFileRef = useRef<File | null>(null);
const lastFrameFileRef = useRef<File | null>(null);
const lastArchivedReplayKeyRef = useRef("");
const [sidebarOpen, setSidebarOpen] = useState(false);
const [creationModelId, setCreationModelId] = useState<CreationModelId>("fast-ltx23");
const [creationModeId, setCreationModeId] = useState<CreationModeId>("t2v");
const [creationAspectRatio, setCreationAspectRatio] = useState<AspectRatioId>("16:9");
const [creationResolution, setCreationResolution] = useState<ResolutionId>("720p");
const [creationDurationSec, setCreationDurationSec] = useState(5);
const [lobbyCapabilitiesBundle, setLobbyCapabilitiesBundle] = useState<LobbyCapabilitiesBundle>(
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
);
const activeModelCapabilities = useMemo(
() => resolveModelCapabilities(lobbyCapabilitiesBundle, creationModelId),
[lobbyCapabilitiesBundle, creationModelId],
);
const [sessionCreationConfig, setSessionCreationConfig] = useState<SessionCreationConfig>({
modelId: "fast-ltx23",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
});
const [referencePreviewUrl, setReferencePreviewUrl] = useState<string | null>(null);
const [firstFramePreviewUrl, setFirstFramePreviewUrl] = useState<string | null>(null);
const [lastFramePreviewUrl, setLastFramePreviewUrl] = useState<string | null>(null);
const [currentThumbnail, setCurrentThumbnail] = useState<string | null>(null);
const currentProjectIdRef = useRef("");
const currentProjectCreatedAtRef = useRef(0);
@@ -345,6 +300,57 @@ export default function Page() {
setIsMobileShareCapable(typeof navigator.canShare === "function" && window.matchMedia("(pointer: coarse)").matches);
}, []);
useEffect(() => {
return () => {
if (referencePreviewUrl) {
URL.revokeObjectURL(referencePreviewUrl);
}
if (firstFramePreviewUrl) {
URL.revokeObjectURL(firstFramePreviewUrl);
}
if (lastFramePreviewUrl) {
URL.revokeObjectURL(lastFramePreviewUrl);
}
};
}, [referencePreviewUrl, firstFramePreviewUrl, lastFramePreviewUrl]);
function setPreviewUrl(setter: Dispatch<SetStateAction<string | null>>, file: File | null) {
setter((current) => {
if (current) URL.revokeObjectURL(current);
return file ? URL.createObjectURL(file) : null;
});
}
function handleReferenceSelect(file: File | null) {
referenceFileRef.current = file;
setPreviewUrl(setReferencePreviewUrl, file);
}
function handleFirstFrameSelect(file: File | null) {
firstFrameFileRef.current = file;
setPreviewUrl(setFirstFramePreviewUrl, file);
}
function handleLastFrameSelect(file: File | null) {
lastFrameFileRef.current = file;
setPreviewUrl(setLastFramePreviewUrl, file);
}
const mentionOptions = useMemo(() => buildMentionOptions(storyPresets as Array<{ id?: string; label?: string; description?: string }>), [storyPresets]);
const lobbyStoryPresets = useMemo(
() =>
(storyPresets as Array<{ id?: string; label?: string; description?: string; segment_prompts?: unknown }>)
.filter((preset) => typeof preset.id === "string" && typeof preset.label === "string")
.map((preset) => ({
id: String(preset.id),
label: String(preset.label),
description: typeof preset.description === "string" ? preset.description : undefined,
segmentCount: Array.isArray(preset.segment_prompts) ? preset.segment_prompts.length : undefined,
})),
[storyPresets],
);
const videoElRef = useRef<HTMLVideoElement | null>(null);
const archivedPlaybackElRef = useRef<HTMLVideoElement | null>(null);
const viewingModePlaybackStateRef = useRef<{
@@ -524,6 +530,65 @@ export default function Page() {
setRuntimeReady(true);
}, []);
function applyLobbyCapabilitiesBundle(bundle: LobbyCapabilitiesBundle) {
setLobbyCapabilitiesBundle(bundle);
const clamped = clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(bundle, creationModelId),
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
setCreationModelId(clamped.modelId);
setCreationModeId(clamped.modeId);
setCreationAspectRatio(clamped.aspectRatio);
setCreationResolution(clamped.resolution);
setCreationDurationSec(clamped.durationSec);
}
function handleCreationModelChange(modelId: CreationModelId) {
const clamped = clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(lobbyCapabilitiesBundle, modelId),
modelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
setCreationModelId(clamped.modelId);
setCreationModeId(clamped.modeId);
setCreationAspectRatio(clamped.aspectRatio);
setCreationResolution(clamped.resolution);
setCreationDurationSec(clamped.durationSec);
}
useEffect(() => {
if (!runtimeReady) return;
let cancelled = false;
void fetch("/creation-capabilities", {
headers: { Accept: "application/json" },
cache: "no-store",
})
.then(async (response) => {
if (!response.ok) return DEFAULT_LOBBY_CAPABILITIES_BUNDLE;
return parseLobbyCapabilitiesBundle(await response.json());
})
.then((bundle) => {
if (!cancelled) {
applyLobbyCapabilitiesBundle(bundle);
}
})
.catch(() => {
if (!cancelled) {
applyLobbyCapabilitiesBundle(DEFAULT_LOBBY_CAPABILITIES_BUNDLE);
}
});
return () => {
cancelled = true;
};
}, [runtimeReady]);
useEffect(() => {
if (!runtimeReady || initializedRef.current) return;
initializedRef.current = true;
@@ -1759,33 +1824,44 @@ export default function Page() {
resetPlaybackState();
}
function buildProjectInitPayload(type: "session_init_v2" | "project_init_v1") {
async function buildProjectInitPayload(type: "session_init_v2" | "project_init_v1") {
const segmentPrompts = getSessionInitPrompts();
setSeedPrompts(segmentPrompts);
const creationPayload = await buildCreationInitPayload({
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
referenceFile: referenceFileRef.current,
firstFrameFile: firstFrameFileRef.current,
lastFrameFile: lastFrameFileRef.current,
});
return {
type,
generation_mode: toGenerationMode(creationModeId),
preset_id: getInitialPresetId(),
preset_label: getInitialPresetLabel(),
curated_prompts: segmentPrompts,
initial_rollout_prompt: normalizeInitialPrompt(pendingInitialPromptRef.current),
initial_image: null,
single_clip_mode: false,
enhancement_enabled: sessionStore.get().enhancementEnabled,
auto_extension_enabled: sessionStore.get().autoExtensionEnabled,
loop_generation_enabled: sessionStore.get().loopGenerationEnabled,
...creationPayload,
};
}
function sendSessionInitMessage() {
async function sendSessionInitMessage() {
const ws = wsRef.current;
if (!ws) return;
ws.send(JSON.stringify(buildProjectInitPayload("session_init_v2")));
ws.send(JSON.stringify(await buildProjectInitPayload("session_init_v2")));
}
function sendProjectInitMessage() {
async function sendProjectInitMessage() {
const ws = wsRef.current;
if (!ws || ws.readyState !== WebSocket.OPEN) return;
ws.send(JSON.stringify(buildProjectInitPayload("project_init_v1")));
ws.send(JSON.stringify(await buildProjectInitPayload("project_init_v1")));
}
function sendEndProjectKeepSession() {
@@ -1822,6 +1898,9 @@ export default function Page() {
return;
}
const normalizedEvent = normalizeSocketMessage(decoded.data);
if (decoded.data?.type === "gpu_assigned" || decoded.data?.type === "ltx2_stream_start") {
applyEchoedCreationConfig(decoded.data);
}
await applyNormalizedSocketEvent(normalizedEvent, {
sessionStore,
promptWindowStore,
@@ -1864,7 +1943,12 @@ export default function Page() {
onOpen: () => {
opened = true;
sessionStore.patch({ connected: true, connecting: false });
sendSessionInitMessage();
void sendSessionInitMessage().catch((error) => {
console.error("Failed to send session init payload:", error);
recoverFailedSessionStart(
error instanceof Error ? error.message : "Failed to prepare session settings.",
);
});
},
onMessage: (event: MessageEvent) => {
wsMessageQueueRef.current = wsMessageQueueRef.current
@@ -1937,10 +2021,29 @@ export default function Page() {
}
}
function syncSessionCreationConfigFromLobby() {
setSessionCreationConfig({
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
}
function applyEchoedCreationConfig(data: unknown) {
const echoed = parseEchoedCreationConfig(data);
if (!echoed) {
return;
}
setSessionCreationConfig(echoed);
}
function beginProjectLocally({ force = false } = {}) {
if (!force && !canStartSession) return;
if (sessionStore.get().sessionStarted || sessionStore.get().projectResetPending) return false;
setTimeoutModalOpen(false);
syncSessionCreationConfigFromLobby();
// Unmute during the user gesture so iOS Safari permits audio playback.
setVideoMuted(false);
if (viewingProject) closeViewingProject();
@@ -1999,13 +2102,34 @@ export default function Page() {
}
async function joinSession({ force = false } = {}) {
const validationError = validateLobbyCreationSelection({
capabilities: activeModelCapabilities,
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
referenceFile: referenceFileRef.current,
firstFrameFile: firstFrameFileRef.current,
lastFrameFile: lastFrameFileRef.current,
});
if (validationError) {
showPreSessionNotice(validationError);
return;
}
if (
wsRef.current
&& wsRef.current.readyState === WebSocket.OPEN
&& sessionStore.get().connected
) {
if (!beginProjectLocally({ force })) return;
sendProjectInitMessage();
try {
await sendProjectInitMessage();
} catch (error) {
console.error("Failed to send project init payload:", error);
showPreSessionNotice(error instanceof Error ? error.message : "Failed to prepare session settings.");
}
return;
}
showPreSessionNotice("");
@@ -2027,7 +2151,12 @@ export default function Page() {
&& wsRef.current.readyState === WebSocket.OPEN
&& sessionStore.get().connected
) {
sendProjectInitMessage();
try {
await sendProjectInitMessage();
} catch (error) {
console.error("Failed to send project init payload:", error);
showPreSessionNotice(error instanceof Error ? error.message : "Failed to prepare session settings.");
}
return;
}
connectWebSocket();
@@ -2641,7 +2770,7 @@ export default function Page() {
/>
<Header timeLeft={headerTimeLeft} formatTime={formatTime} onToggleSidebar={() => setSidebarOpen((prev) => !prev)} />
<div className="relative flex flex-1 min-h-0 flex-col justify-center px-4 pb-2 sm:px-6 sm:pb-12">
<div className={cn("relative flex flex-1 min-h-0 flex-col", showActiveProject || isViewingMode ? "justify-center px-4 pb-2 sm:px-6 sm:pb-12" : "overflow-hidden")}>
{isViewingMode && (
<>
{viewingSelectedClip && (
@@ -2758,44 +2887,79 @@ export default function Page() {
/>
</section>
<AnimatePresence>
{!showActiveProject && (
<motion.div
key="hero-tagline"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0, transition: { duration: 0.2, ease: "easeIn" } }}
transition={{ duration: 0.5, ease: "easeOut" }}
className="pointer-events-none absolute inset-x-0 top-0 bottom-1/2 z-10 flex items-center justify-center px-4"
>
<HeroTagline />
</motion.div>
)}
</AnimatePresence>
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
<ChatBar
sessionStarted={sessionStarted as boolean}
rewritingSeedPrompts={rewritingSeedPrompts as boolean}
{!showActiveProject ? (
<CreationStudio
value={livePromptDraft as string}
disabled={projectResetPending as boolean}
isGenerating={loadingAnimation as boolean}
storyPresets={storyPresets as any[]}
continuationDraft={livePromptDraft as string}
canJoinSession={canStartSession}
canSubmitContinuation={canSubmitContinuation}
sessionExpired={sessionExpired as boolean}
sessionNotice={sessionNotice as string}
projectResetPending={projectResetPending as boolean}
canSubmit={canStartSession}
modelId={creationModelId}
modeId={creationModeId}
aspectRatio={creationAspectRatio}
resolution={creationResolution}
durationSec={creationDurationSec}
referencePreviewUrl={referencePreviewUrl}
firstFramePreviewUrl={firstFramePreviewUrl}
lastFramePreviewUrl={lastFramePreviewUrl}
mentionOptions={mentionOptions}
storyPresets={lobbyStoryPresets}
capabilities={activeModelCapabilities}
onValueChange={(value) => sessionStore.patch({ livePromptDraft: value })}
onSubmit={() => void joinSession()}
onKeyDown={handleLivePromptKeydown}
onModelChange={handleCreationModelChange}
onModeChange={setCreationModeId}
onAspectRatioChange={setCreationAspectRatio}
onResolutionChange={setCreationResolution}
onDurationChange={setCreationDurationSec}
onReferenceSelect={handleReferenceSelect}
onFirstFrameSelect={handleFirstFrameSelect}
onLastFrameSelect={handleLastFrameSelect}
onPresetGenerate={handlePresetGenerate}
onContinuationInput={handleLivePromptInput}
onContinuationKeydown={handleLivePromptKeydown}
onGenerate={joinSession}
onSubmitContinuation={submitLivePrompt}
onLeave={leaveSession}
onStartNewProject={handleStartNewProject}
onSpeechTranscript={handleLivePromptSpeechTranscript}
onSpeechInterimChange={handleLivePromptSpeechInterim}
onOpenProjects={() => setSidebarOpen(true)}
/>
</motion.div>
) : (
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
<ChatBar
sessionStarted={sessionStarted as boolean}
rewritingSeedPrompts={rewritingSeedPrompts as boolean}
isGenerating={loadingAnimation as boolean}
storyPresets={storyPresets as any[]}
continuationDraft={livePromptDraft as string}
canJoinSession={canStartSession}
canSubmitContinuation={canSubmitContinuation}
sessionExpired={sessionExpired as boolean}
sessionNotice={sessionNotice as string}
projectResetPending={projectResetPending as boolean}
sessionCreationConfig={sessionCreationConfig}
configPillsReadOnly
onSessionModelChange={(modelId) => setSessionCreationConfig((current) => ({ ...current, modelId }))}
onSessionModeChange={(modeId) => setSessionCreationConfig((current) => ({ ...current, modeId }))}
onSessionAspectRatioChange={(aspectRatio) => setSessionCreationConfig((current) => ({ ...current, aspectRatio }))}
onSessionResolutionChange={(resolution) => setSessionCreationConfig((current) => ({ ...current, resolution }))}
onSessionDurationChange={(durationSec) => setSessionCreationConfig((current) => ({ ...current, durationSec }))}
onPresetGenerate={handlePresetGenerate}
onContinuationInput={handleLivePromptInput}
onContinuationKeydown={handleLivePromptKeydown}
onGenerate={joinSession}
onSubmitContinuation={submitLivePrompt}
onLeave={leaveSession}
onStartNewProject={handleStartNewProject}
onSpeechTranscript={handleLivePromptSpeechTranscript}
onSpeechInterimChange={handleLivePromptSpeechInterim}
/>
</motion.div>
)}
{sessionNotice && !showActiveProject && (
<div className="mx-auto mt-2 w-full max-w-3xl px-4">
<div className="rounded-xl border border-rose-500/20 bg-rose-500/10 px-4 py-2.5 text-center text-xs text-rose-700 dark:text-rose-300">
{sessionNotice}
</div>
</div>
)}
</div>
</div>
</main>
+62 -197
View File
@@ -2,10 +2,13 @@
import React, { useRef, useState, useCallback, useEffect } from "react";
import Image from "next/image";
import { Film, ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
import { ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
import { Button } from "@/components/ui/button";
import LeaveSessionModal, { shouldShowLeaveWarning } from "@/components/LeaveSessionModal";
import PresetQuickLaunchRail from "@/components/creation/PresetQuickLaunchRail";
import SessionCreationConfigPills, { type SessionCreationConfig } from "@/components/creation/SessionCreationConfigPills";
import SpeechToTextButton from "@/components/SpeechToTextButton";
import type { AspectRatioId, CreationModeId, CreationModelId, ResolutionId } from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
const PROMPT_MAX_LENGTH = 500;
@@ -32,6 +35,13 @@ interface Props {
onBackFromViewing?: () => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
sessionCreationConfig?: SessionCreationConfig | null;
configPillsReadOnly?: boolean;
onSessionModelChange?: (modelId: CreationModelId) => void;
onSessionModeChange?: (modeId: CreationModeId) => void;
onSessionAspectRatioChange?: (aspectRatio: AspectRatioId) => void;
onSessionResolutionChange?: (resolution: ResolutionId) => void;
onSessionDurationChange?: (durationSec: number) => void;
}
export default function ChatBar({
@@ -56,6 +66,13 @@ export default function ChatBar({
onBackFromViewing = () => {},
onSpeechTranscript,
onSpeechInterimChange,
sessionCreationConfig = null,
configPillsReadOnly = false,
onSessionModelChange,
onSessionModeChange,
onSessionAspectRatioChange,
onSessionResolutionChange,
onSessionDurationChange,
}: Props) {
const [sttBusy, setSttBusy] = useState(false);
const [leaveModalOpen, setLeaveModalOpen] = useState(false);
@@ -66,134 +83,11 @@ export default function ChatBar({
: isBusy
? "Generating video\u2026"
: !sessionStarted
? "What video are you imagining?"
? "Describe your video"
: "What do you want to edit?";
const actionLabel = !sessionStarted ? "Generate" : "Rewrite rollout";
const inputRef = useRef<HTMLTextAreaElement>(null);
const scrollRef = useRef<HTMLDivElement>(null);
const [canScrollLeft, setCanScrollLeft] = useState(false);
const [canScrollRight, setCanScrollRight] = useState(false);
const [presetRailDragging, setPresetRailDragging] = useState(false);
const presetDragStateRef = useRef({
pointerId: null as number | null,
startX: 0,
startScrollLeft: 0,
moved: false,
});
const suppressPresetClickRef = useRef(false);
const updateScrollState = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setCanScrollLeft(el.scrollLeft > 2);
setCanScrollRight(el.scrollLeft + el.clientWidth < el.scrollWidth - 2);
}, []);
const handlePresetWheel = useCallback(
(event: React.WheelEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
const dominantDelta = Math.abs(event.deltaX) > Math.abs(event.deltaY)
? event.deltaX
: event.deltaY;
if (!dominantDelta) return;
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
const nextScrollLeft = Math.min(
Math.max(el.scrollLeft + dominantDelta, 0),
maxScrollLeft,
);
if (nextScrollLeft === el.scrollLeft) return;
event.preventDefault();
el.scrollLeft = nextScrollLeft;
updateScrollState();
},
[updateScrollState],
);
const finishPresetDrag = useCallback(() => {
presetDragStateRef.current = {
pointerId: null,
startX: 0,
startScrollLeft: 0,
moved: false,
};
setPresetRailDragging(false);
}, []);
const handlePresetPointerDown = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (event.pointerType !== "mouse" || event.button !== 0) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
suppressPresetClickRef.current = false;
presetDragStateRef.current = {
pointerId: event.pointerId,
startX: event.clientX,
startScrollLeft: el.scrollLeft,
moved: false,
};
},
[],
);
const handlePresetPointerMove = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
const dragState = presetDragStateRef.current;
if (!el || dragState.pointerId !== event.pointerId) return;
const deltaX = event.clientX - dragState.startX;
if (!dragState.moved && Math.abs(deltaX) > 4) {
dragState.moved = true;
suppressPresetClickRef.current = true;
setPresetRailDragging(true);
el.setPointerCapture?.(event.pointerId);
}
if (!dragState.moved) return;
event.preventDefault();
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
el.scrollLeft = Math.min(
Math.max(dragState.startScrollLeft - deltaX, 0),
maxScrollLeft,
);
updateScrollState();
},
[updateScrollState],
);
const handlePresetPointerUp = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el || presetDragStateRef.current.pointerId !== event.pointerId) return;
if (el.hasPointerCapture?.(event.pointerId)) {
el.releasePointerCapture(event.pointerId);
}
finishPresetDrag();
},
[finishPresetDrag],
);
const handlePresetClickCapture = useCallback(
(event: React.MouseEvent<HTMLDivElement>) => {
if (!suppressPresetClickRef.current) return;
suppressPresetClickRef.current = false;
event.preventDefault();
event.stopPropagation();
},
[],
);
useEffect(() => {
updateScrollState();
}, [storyPresets, updateScrollState]);
useEffect(() => {
if (!isBusy && !sttBusy && !window.matchMedia("(pointer: coarse)").matches) {
@@ -239,7 +133,7 @@ export default function ChatBar({
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-6 py-4 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">View-only project</p>
<p className="max-w-md text-xs text-muted-foreground">Project sessions are currently limited to 5 minutes. Start a new project to create more videos.</p>
<p className="max-w-md text-xs text-muted-foreground">Sessions are limited to 5 minutes. Start a new project to keep creating.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onBackFromViewing} variant="outline" size="sm" className="gap-1.5 rounded-full px-4">
@@ -261,7 +155,7 @@ export default function ChatBar({
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-8 py-5 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">Session ended</p>
<p className="max-w-xs text-xs text-muted-foreground">Each project currently has a 5-minute session. Start a new project to continue creating videos.</p>
<p className="max-w-xs text-xs text-muted-foreground">Sessions are limited to 5 minutes. Start a new project to continue.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onStartNewProject} size="sm" className="rounded-full px-5">
@@ -280,51 +174,8 @@ export default function ChatBar({
return (
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
{storyPresets.length > 0 && !sessionStarted && (
<div className={cn("relative transition-opacity duration-200", isGenerating && "pointer-events-none opacity-40")}>
<div
ref={scrollRef}
onScroll={updateScrollState}
onWheel={handlePresetWheel}
onPointerDown={handlePresetPointerDown}
onPointerMove={handlePresetPointerMove}
onPointerUp={handlePresetPointerUp}
onPointerCancel={handlePresetPointerUp}
onLostPointerCapture={finishPresetDrag}
onClickCapture={handlePresetClickCapture}
className={cn(
"scrollbar-hidden flex gap-3 overflow-x-auto px-1 select-none",
presetRailDragging ? "cursor-grabbing" : "cursor-grab",
)}
>
{storyPresets.map((preset) => (
<button
key={preset.id}
type="button"
disabled={isGenerating}
onClick={() => onPresetGenerate(preset.id)}
className="flex flex-col sm:flex-row items-start gap-1.5 shrink-0 rounded-xl border p-2.5 text-left backdrop-blur-sm transition-colors max-w-42 sm:max-w-[215px] border-input bg-card/80 text-muted-foreground hover:bg-slate-200/60 hover:border-slate-400 hover:text-slate-700 dark:bg-slate-800/80 dark:text-slate-300 dark:hover:bg-slate-700/50 dark:hover:border-slate-500 dark:hover:text-slate-200"
>
<Film className="mt-0.5 size-4 shrink-0 opacity-60" />
<span className="flex flex-col gap-1 min-w-0">
<span className="text-[14px] font-medium line-clamp-1">{preset.label}</span>
{preset.description && <span className="text-xs leading-tight opacity-70 line-clamp-3 sm:line-clamp-2">{preset.description}</span>}
</span>
</button>
))}
</div>
<div
className={cn("pointer-events-none absolute inset-y-0 left-0 w-8 bg-background transition-opacity duration-150", canScrollLeft ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to right, black, transparent)", WebkitMaskImage: "linear-gradient(to right, black, transparent)" }}
aria-hidden="true"
/>
<div
className={cn("pointer-events-none absolute inset-y-0 right-0 w-8 bg-background transition-opacity duration-150", canScrollRight ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to left, black, transparent)", WebkitMaskImage: "linear-gradient(to left, black, transparent)" }}
aria-hidden="true"
/>
</div>
{!sessionStarted && (
<PresetQuickLaunchRail storyPresets={storyPresets} disabled={isGenerating} onPresetGenerate={onPresetGenerate} />
)}
{sessionNotice && (
@@ -342,17 +193,30 @@ export default function ChatBar({
{projectResetPending && sessionStarted && (
<div className="rounded-xl border border-sky-500/20 bg-sky-500/10 px-4 py-2.5 text-center text-xs text-sky-700 dark:text-sky-300">
Starting a new project after the current shot finishes. Your GPU session stays active.
Starting a new project when this shot finishes. GPU session stays open.
</div>
)}
<div
className={cn(
"flex min-w-0 items-center gap-1.5 rounded-4xl border py-2.5 pl-5 pr-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
"flex min-w-0 flex-col gap-2 rounded-4xl border py-2.5 pl-5 pr-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
isBusy ? "border-input/60 bg-card/40" : "border-input bg-card/65",
)}
>
<textarea
{sessionStarted && sessionCreationConfig && (
<SessionCreationConfigPills
{...sessionCreationConfig}
disabled={isBusy}
readOnly={configPillsReadOnly}
onModelChange={onSessionModelChange}
onModeChange={onSessionModeChange}
onAspectRatioChange={onSessionAspectRatioChange}
onResolutionChange={onSessionResolutionChange}
onDurationChange={onSessionDurationChange}
/>
)}
<div className="flex min-w-0 items-center gap-1.5">
<textarea
ref={inputRef}
id="continuation-prompt"
aria-label="Continuation prompt"
@@ -367,36 +231,37 @@ export default function ChatBar({
"min-w-0 flex-1 resize-none bg-transparent text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
(isBusy || sttBusy) && "cursor-not-allowed opacity-50",
)}
/>
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
{!sessionStarted ? (
<Button
aria-label={actionLabel}
title={actionLabel}
onClick={onGenerate}
disabled={!canJoinSession || isGenerating || !continuationDraft.trim()}
size="icon-sm"
className="shrink-0 rounded-full"
>
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
</Button>
) : (
<>
/>
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
{!sessionStarted ? (
<Button
aria-label={actionLabel}
title={actionLabel}
onClick={onSubmitContinuation}
disabled={!canSubmitContinuation || showSpinner || projectResetPending || !continuationDraft.trim()}
onClick={onGenerate}
disabled={!canJoinSession || isGenerating || !continuationDraft.trim()}
size="icon-sm"
className="shrink-0 rounded-full"
>
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
</Button>
<Button variant="outline" aria-label="Leave" title="Leave" onClick={() => { if (shouldShowLeaveWarning()) setLeaveModalOpen(true); else onLeave(); }} disabled={isGenerating || projectResetPending} size="icon-sm" className="shrink-0 rounded-full">
<X className="size-5" />
</Button>
</>
)}
) : (
<>
<Button
aria-label={actionLabel}
title={actionLabel}
onClick={onSubmitContinuation}
disabled={!canSubmitContinuation || showSpinner || projectResetPending || !continuationDraft.trim()}
size="icon-sm"
className="shrink-0 rounded-full"
>
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
</Button>
<Button variant="outline" aria-label="Leave" title="Leave" onClick={() => { if (shouldShowLeaveWarning()) setLeaveModalOpen(true); else onLeave(); }} disabled={isGenerating || projectResetPending} size="icon-sm" className="shrink-0 rounded-full">
<X className="size-5" />
</Button>
</>
)}
</div>
</div>
<p className="px-2 text-center text-[11px] text-muted-foreground">
LLM powered by{" "}
@@ -0,0 +1,95 @@
"use client";
import { Fragment, useEffect, useRef } from "react";
const HERO_WAVE_LIGHT = ["#2A4A98", "#4878E5", "#6FA0F2", "#B0BCC8", "#E8D99E", "#D8C844", "#C2A620"];
const HERO_WAVE_DARK = ["#143468", "#1E58B8", "#3892F0", "#80B8E8", "#B8D0EA", "#E2D498", "#DABB50"];
const HERO_TEXT = "Direct scenes in seconds";
export default function HeroTagline() {
const ref = useRef<HTMLHeadingElement>(null);
useEffect(() => {
const el = ref.current;
if (!el) return;
let rafId = 0;
function play() {
const chars = el!.querySelectorAll<HTMLSpanElement>("[data-char]");
if (!chars.length) return;
cancelAnimationFrame(rafId);
const isDark = document.documentElement.classList.contains("dark");
const colors = isDark ? HERO_WAVE_DARK : HERO_WAVE_LIGHT;
const waveLen = 10;
const total = chars.length + waveLen;
const duration = 1200;
const maxBlur = 3.5;
const start = performance.now();
function tick() {
const t = Math.min((performance.now() - start) / duration, 1);
const pos = t * total;
chars.forEach((ch, i) => {
const rel = pos - i;
if (rel >= 0 && rel < waveLen) {
const norm = rel / waveLen;
const ci = Math.floor(norm * colors.length);
ch.style.color = colors[Math.min(colors.length - 1, ci)];
let blur = 0;
if (norm < 0.25) {
blur = maxBlur * (1 - norm / 0.25);
} else if (norm > 0.75) {
blur = maxBlur * ((norm - 0.75) / 0.25);
}
ch.style.filter = blur > 0.1 ? `blur(${blur.toFixed(1)}px)` : "";
} else {
ch.style.color = "";
ch.style.filter = "";
}
});
if (t < 1) {
rafId = requestAnimationFrame(tick);
} else {
chars.forEach((ch) => {
ch.style.color = "";
ch.style.filter = "";
});
}
}
rafId = requestAnimationFrame(tick);
}
const initialDelay = setTimeout(play, 400);
const interval = setInterval(play, 5000);
return () => {
clearTimeout(initialDelay);
clearInterval(interval);
cancelAnimationFrame(rafId);
};
}, []);
return (
<h1 ref={ref} className="text-balance text-center text-3xl font-medium text-[#343537] dark:text-[#FAFAFB] sm:text-4xl">
{HERO_TEXT.split(" ").map((word, wi) => (
<Fragment key={wi}>
{wi > 0 && (
<span data-char className="transition-[color,filter] duration-150">
{" "}
</span>
)}
<span className="inline-flex">
{word.split("").map((char, ci) => (
<span key={ci} data-char className="inline-block transition-[color,filter] duration-150">
{char}
</span>
))}
</span>
</Fragment>
))}
</h1>
);
}
@@ -0,0 +1,66 @@
"use client";
import React from "react";
import { FolderOpen, Home, Sparkles } from "lucide-react";
import { cn } from "@/lib/utils";
export type AppNavSection = "explore" | "create" | "assets";
interface AppNavRailProps {
activeSection?: AppNavSection;
onSectionChange?: (section: AppNavSection) => void;
onOpenProjects?: () => void;
className?: string;
}
const NAV_ITEMS: Array<{ id: AppNavSection; label: string; icon: typeof Home }> = [
{ id: "explore", label: "Explore", icon: Home },
{ id: "create", label: "Create", icon: Sparkles },
{ id: "assets", label: "Assets", icon: FolderOpen },
];
export default function AppNavRail({
activeSection = "create",
onSectionChange = () => {},
onOpenProjects,
className,
}: AppNavRailProps) {
return (
<aside
className={cn(
"hidden shrink-0 flex-col items-center gap-2 border-r border-border/40 bg-background/30 px-2.5 py-5 lg:flex",
className,
)}
aria-label="Primary navigation"
>
{NAV_ITEMS.map((item) => {
const Icon = item.icon;
const isActive = item.id === activeSection;
return (
<button
key={item.id}
type="button"
aria-label={item.label}
aria-current={isActive ? "page" : undefined}
onClick={() => {
if (item.id === "assets") {
onOpenProjects?.();
}
onSectionChange(item.id);
}}
className={cn(
"studio-control studio-control-press flex w-[4.5rem] min-h-11 flex-col items-center gap-1 rounded-xl px-2 py-2.5 text-[10px] font-medium tracking-wide",
isActive
? "bg-secondary/90 text-foreground shadow-sm ring-1 ring-border/60"
: "text-muted-foreground hover-capable:hover:bg-secondary/50 hover-capable:hover:text-foreground",
)}
>
<Icon className={cn("size-[18px]", isActive && "text-accent-blue")} />
{item.label}
</button>
);
})}
</aside>
);
}
@@ -0,0 +1,24 @@
"use client";
import React from "react";
import { cn } from "@/lib/utils";
export default function ConfigPill({
children,
className,
...props
}: React.ButtonHTMLAttributes<HTMLButtonElement>) {
return (
<button
type="button"
className={cn(
"studio-control studio-control-press studio-hover-surface inline-flex h-9 min-h-9 shrink-0 items-center gap-1 rounded-full border border-border/50 bg-background/80 px-2.5 text-[11px] font-medium text-foreground/90",
className,
)}
{...props}
>
{children}
</button>
);
}
@@ -0,0 +1,466 @@
"use client";
import React, { useMemo, useRef, useState } from "react";
import { ArrowUp, Box, ChevronDown, Clock, Monitor, Wand2 } from "lucide-react";
import ConfigPill from "@/components/creation/ConfigPill";
import HeroTagline from "@/components/HeroTagline";
import ReferenceUploadSlot from "@/components/creation/ReferenceUploadSlot";
import { Button } from "@/components/ui/button";
import {
DropdownMenu,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuTrigger,
} from "@/components/ui/dropdown-menu";
import { Popover, PopoverContent, PopoverTrigger } from "@/components/ui/popover";
import { Slider } from "@/components/ui/slider";
import SpeechToTextButton from "@/components/SpeechToTextButton";
import {
ASPECT_RATIOS,
CREATION_MODELS,
CREATION_MODES,
RESOLUTIONS,
UNSUPPORTED_CREATION_MODES,
UNSUPPORTED_RESOLUTIONS,
modeRequiresReference,
modeUsesDualFrames,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type MentionOption,
type ResolutionId,
formatDurationLabel,
formatResolutionLabel,
} from "@/lib/creationConfig";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
isSupportedCreationMode,
isSupportedResolution,
resolveModelCapabilities,
unsupportedModeNotice,
type LobbyCreationCapabilities,
} from "@/lib/creationCapabilities";
import { cn } from "@/lib/utils";
const PROMPT_MAX_LENGTH = 500;
interface CreationComposerProps {
value: string;
disabled?: boolean;
isGenerating?: boolean;
canSubmit?: boolean;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referencePreviewUrl?: string | null;
firstFramePreviewUrl?: string | null;
lastFramePreviewUrl?: string | null;
mentionOptions?: MentionOption[];
onValueChange: (value: string) => void;
onSubmit: () => void;
onKeyDown?: (event: React.KeyboardEvent<HTMLTextAreaElement>) => void;
onModelChange: (modelId: CreationModelId) => void;
onModeChange: (modeId: CreationModeId) => void;
onAspectRatioChange: (aspectRatio: AspectRatioId) => void;
onResolutionChange: (resolution: ResolutionId) => void;
onDurationChange: (durationSec: number) => void;
onReferenceSelect?: (file: File | null) => void;
onFirstFrameSelect?: (file: File | null) => void;
onLastFrameSelect?: (file: File | null) => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
capabilities?: LobbyCreationCapabilities;
}
export default function CreationComposer({
value,
disabled = false,
isGenerating = false,
canSubmit = false,
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
referencePreviewUrl = null,
firstFramePreviewUrl = null,
lastFramePreviewUrl = null,
mentionOptions = [],
onValueChange,
onSubmit,
onKeyDown,
onModelChange,
onModeChange,
onAspectRatioChange,
onResolutionChange,
onDurationChange,
onReferenceSelect,
onFirstFrameSelect,
onLastFrameSelect,
onSpeechTranscript,
onSpeechInterimChange,
capabilities = resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, modelId),
}: CreationComposerProps) {
const inputRef = useRef<HTMLTextAreaElement>(null);
const [sttBusy, setSttBusy] = useState(false);
const [mentionQuery, setMentionQuery] = useState("");
const [mentionOpen, setMentionOpen] = useState(false);
const [mentionStart, setMentionStart] = useState<number | null>(null);
const availableModels = useMemo(
() => CREATION_MODELS.filter((model) => capabilities.model_ids.includes(model.id)),
[capabilities.model_ids],
);
const availableModes = useMemo(
() => CREATION_MODES.filter((mode) => isSupportedCreationMode(mode.id, capabilities)),
[capabilities],
);
const unavailableModes = useMemo(
() =>
UNSUPPORTED_CREATION_MODES.filter(
(mode) => unsupportedModeNotice(mode.id, capabilities) !== null,
),
[capabilities],
);
const availableAspectRatios = useMemo(
() => ASPECT_RATIOS.filter((ratio) => capabilities.aspect_ratios.includes(ratio)),
[capabilities.aspect_ratios],
);
const availableResolutions = useMemo(
() => RESOLUTIONS.filter((item) => isSupportedResolution(item, capabilities)),
[capabilities],
);
const unavailableResolutions = useMemo(
() => UNSUPPORTED_RESOLUTIONS.filter((item) => !isSupportedResolution(item, capabilities)),
[capabilities],
);
const durationMin = capabilities.duration_sec[0] ?? 5;
const durationMax = capabilities.duration_sec[capabilities.duration_sec.length - 1] ?? 15;
const selectedModel = availableModels.find((model) => model.id === modelId) ?? availableModels[0];
const selectedMode = availableModes.find((mode) => mode.id === modeId) ?? availableModes[0];
const usesDualFrames = modeUsesDualFrames(modeId);
const requiresReference = modeRequiresReference(modeId);
const referenceMissing = requiresReference && !referencePreviewUrl;
const submitDisabled = !canSubmit || disabled || isGenerating || !value.trim() || referenceMissing;
const filteredMentions = useMemo(() => {
const query = mentionQuery.trim().toLowerCase();
if (!query) return mentionOptions.slice(0, 6);
return mentionOptions
.filter((option) => option.label.toLowerCase().includes(query) || option.description?.toLowerCase().includes(query))
.slice(0, 6);
}, [mentionOptions, mentionQuery]);
function autoResize() {
const el = inputRef.current;
if (!el) return;
el.style.height = "auto";
const lineHeight = parseFloat(getComputedStyle(el).lineHeight) || 22;
const maxHeight = lineHeight * 4;
el.style.height = `${Math.min(el.scrollHeight, maxHeight)}px`;
el.style.overflowY = el.scrollHeight > maxHeight ? "auto" : "hidden";
}
function updateMentionState(nextValue: string, cursorPosition: number) {
const beforeCursor = nextValue.slice(0, cursorPosition);
const atIndex = beforeCursor.lastIndexOf("@");
if (atIndex === -1 || (atIndex > 0 && !/\s/.test(beforeCursor[atIndex - 1] ?? ""))) {
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
return;
}
const query = beforeCursor.slice(atIndex + 1);
if (/\s/.test(query)) {
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
return;
}
setMentionStart(atIndex);
setMentionQuery(query);
setMentionOpen(true);
}
function insertMention(option: MentionOption) {
if (mentionStart === null) return;
const before = value.slice(0, mentionStart);
const after = value.slice(inputRef.current?.selectionStart ?? value.length);
const mentionText = `@${option.label} `;
const nextValue = `${before}${mentionText}${after}`.slice(0, PROMPT_MAX_LENGTH);
onValueChange(nextValue);
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
requestAnimationFrame(() => {
const el = inputRef.current;
if (!el) return;
const cursor = before.length + mentionText.length;
el.focus();
el.setSelectionRange(cursor, cursor);
autoResize();
});
}
function handleInputChange(event: React.ChangeEvent<HTMLTextAreaElement>) {
const nextValue = event.target.value.slice(0, PROMPT_MAX_LENGTH);
onValueChange(nextValue);
updateMentionState(nextValue, event.target.selectionStart ?? nextValue.length);
requestAnimationFrame(autoResize);
}
function handleKeyDown(event: React.KeyboardEvent<HTMLTextAreaElement>) {
if (mentionOpen && filteredMentions.length > 0) {
if (event.key === "Tab" || (event.key === "Enter" && !event.shiftKey)) {
event.preventDefault();
insertMention(filteredMentions[0]);
return;
}
if (event.key === "Escape") {
setMentionOpen(false);
return;
}
}
onKeyDown?.(event);
}
return (
<section className="mx-auto flex w-full max-w-3xl flex-col gap-5">
<HeroTagline />
<div className="rounded-[32px] border border-border/40 bg-secondary/95 p-4 shadow-[0_24px_80px_-32px_rgba(0,0,0,0.72)] backdrop-blur-xl sm:p-5">
<div className="flex gap-3.5">
{usesDualFrames ? (
<div className="flex shrink-0 gap-2">
<ReferenceUploadSlot
label="Asset"
sublabel="First"
previewUrl={firstFramePreviewUrl}
disabled={disabled}
onSelect={onFirstFrameSelect}
/>
<ReferenceUploadSlot
label="Asset"
sublabel="Last"
previewUrl={lastFramePreviewUrl}
disabled={disabled}
onSelect={onLastFrameSelect}
/>
</div>
) : (
<ReferenceUploadSlot
label="Reference"
previewUrl={referencePreviewUrl}
required={requiresReference}
optional={!requiresReference}
disabled={disabled}
onSelect={onReferenceSelect}
/>
)}
<div className="relative min-w-0 flex-1">
<textarea
ref={inputRef}
id="continuation-prompt"
aria-label="Continuation prompt"
value={value}
onChange={handleInputChange}
onKeyDown={handleKeyDown}
onClick={(event) => updateMentionState(value, event.currentTarget.selectionStart ?? value.length)}
placeholder="Describe your video or mention elements"
disabled={disabled || sttBusy}
rows={3}
className={cn(
"min-h-[92px] w-full resize-none bg-transparent px-0.5 text-base leading-6 text-foreground outline-none placeholder:text-muted-foreground/80 sm:text-sm",
(disabled || sttBusy) && "cursor-not-allowed opacity-50",
)}
/>
{mentionOpen && filteredMentions.length > 0 && (
<div className="absolute left-0 right-0 top-full z-20 mt-2 overflow-hidden rounded-2xl border border-border bg-popover/95 p-1 shadow-xl backdrop-blur-md">
<p className="px-2 py-1 text-[11px] font-medium uppercase tracking-wide text-muted-foreground">Mention</p>
{filteredMentions.map((option) => (
<button
key={option.id}
type="button"
onMouseDown={(event) => {
event.preventDefault();
insertMention(option);
}}
className="studio-control studio-hover-surface flex w-full items-start gap-2 rounded-xl px-2.5 py-2 text-left"
>
<span className="mt-0.5 rounded-md bg-accent px-1.5 py-0.5 text-[10px] font-semibold uppercase tracking-wide text-muted-foreground">
{option.kind}
</span>
<span className="min-w-0">
<span className="block truncate text-sm font-medium text-foreground">{option.label}</span>
{option.description && <span className="block truncate text-xs text-muted-foreground">{option.description}</span>}
</span>
</button>
))}
</div>
)}
</div>
</div>
<div className="mt-4 flex flex-wrap items-center gap-1.5 rounded-2xl bg-muted/35 p-1.5 ring-1 ring-border/25">
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled}>
<Box className="size-3.5" />
{selectedModel.label}
<ChevronDown className="size-3 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-72">
<DropdownMenuLabel>Model</DropdownMenuLabel>
<DropdownMenuSeparator />
{availableModels.map((model) => (
<DropdownMenuItem key={model.id} onClick={() => onModelChange(model.id)} className="flex-col items-start gap-1 py-2.5">
<span className="flex items-center gap-2 text-sm font-medium">
{model.label}
{model.badge && <span className="rounded-full bg-accent-blue/15 px-1.5 py-0.5 text-[10px] text-accent-blue">{model.badge}</span>}
</span>
<span className="text-xs text-muted-foreground">{model.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled}>
<Wand2 className="size-3.5" />
{selectedMode.label}
<ChevronDown className="size-3 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-64">
<DropdownMenuLabel>Mode</DropdownMenuLabel>
<DropdownMenuSeparator />
{availableModes.map((mode) => (
<DropdownMenuItem key={mode.id} onClick={() => onModeChange(mode.id)} className="flex-col items-start gap-1 py-2.5">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">{mode.description}</span>
</DropdownMenuItem>
))}
{unavailableModes.length > 0 && <DropdownMenuSeparator />}
{unavailableModes.map((mode) => (
<DropdownMenuItem key={mode.id} disabled className="flex-col items-start gap-1 py-2.5 opacity-60">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">
{unsupportedModeNotice(mode.id, capabilities) ?? mode.description}
</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled}>
<Monitor className="size-3.5" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-80">
<p className="mb-3 text-xs font-medium text-muted-foreground">Aspect ratio</p>
<div className="grid grid-cols-3 gap-2">
{availableAspectRatios.map((ratio) => (
<button
key={ratio}
type="button"
onClick={() => onAspectRatioChange(ratio)}
className={cn(
"studio-control studio-control-press studio-hover-surface flex flex-col items-center gap-2 rounded-xl border px-2 py-3 text-xs",
aspectRatio === ratio ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
<span className={cn("rounded-sm border border-current/40 bg-muted/40", ratio === "9:16" && "h-7 w-4", ratio === "16:9" && "h-4 w-7", ratio === "1:1" && "size-5", ratio === "4:3" && "h-5 w-6", ratio === "3:4" && "h-6 w-5", ratio === "21:9" && "h-3 w-8")} />
{ratio}
</button>
))}
</div>
<p className="mb-2 mt-4 text-xs font-medium text-muted-foreground">Resolution</p>
<div className="flex flex-wrap gap-2">
{availableResolutions.map((item) => (
<button
key={item}
type="button"
onClick={() => onResolutionChange(item)}
className={cn(
"studio-control studio-control-press studio-hover-surface rounded-full border px-3 py-1.5 text-xs font-medium",
resolution === item ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
{formatResolutionLabel(item)}
</button>
))}
{unavailableResolutions.map((item) => (
<button
key={item}
type="button"
disabled
className="studio-control rounded-full border border-border px-3 py-1.5 text-xs font-medium text-muted-foreground opacity-50"
title="Not supported on FastLTX models yet"
>
{formatResolutionLabel(item)}
</button>
))}
</div>
</PopoverContent>
</Popover>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled}>
<Clock className="size-3.5" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-72">
<p className="mb-3 text-xs font-medium text-muted-foreground">Total duration</p>
<Slider min={durationMin} max={durationMax} step={5} value={[durationSec]} onValueChange={(values) => onDurationChange(values[0] ?? durationMin)} />
<div className="mt-3 flex items-center justify-between text-[11px] text-muted-foreground">
<span>{formatDurationLabel(durationMin)}</span>
<span className="rounded-md border border-border px-2 py-1 text-xs font-medium text-foreground">{formatDurationLabel(durationSec)}</span>
<span>{formatDurationLabel(durationMax)}</span>
</div>
</PopoverContent>
</Popover>
<div className="ml-auto flex items-center gap-1.5">
{onSpeechTranscript && (
<SpeechToTextButton
disabled={disabled || isGenerating}
onTranscript={onSpeechTranscript}
onInterimChange={onSpeechInterimChange}
onBusyChange={setSttBusy}
/>
)}
<Button
aria-label="Generate"
onClick={onSubmit}
disabled={submitDisabled}
size="icon"
className="studio-control-press rounded-full bg-accent-blue text-white shadow-sm hover-capable:hover:bg-accent-blue/90 disabled:bg-muted disabled:text-muted-foreground"
>
<ArrowUp className="size-5" />
</Button>
</div>
</div>
{referenceMissing && value.trim() && (
<p className="mt-3 text-center text-xs leading-5 text-amber-700 dark:text-amber-400">
Upload a reference asset to use Omni reference mode.
</p>
)}
</div>
</section>
);
}
@@ -0,0 +1,73 @@
"use client";
import React from "react";
import AppNavRail, { type AppNavSection } from "@/components/creation/AppNavRail";
import CreationComposer from "@/components/creation/CreationComposer";
import PresetQuickLaunchRail, { type StoryPresetLike } from "@/components/creation/PresetQuickLaunchRail";
import {
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type MentionOption,
type ResolutionId,
} from "@/lib/creationConfig";
import type { LobbyCreationCapabilities } from "@/lib/creationCapabilities";
interface CreationStudioProps {
value: string;
disabled?: boolean;
isGenerating?: boolean;
canSubmit?: boolean;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referencePreviewUrl?: string | null;
firstFramePreviewUrl?: string | null;
lastFramePreviewUrl?: string | null;
mentionOptions?: MentionOption[];
storyPresets?: StoryPresetLike[];
activeSection?: AppNavSection;
onValueChange: (value: string) => void;
onSubmit: () => void;
onKeyDown?: (event: React.KeyboardEvent<HTMLTextAreaElement>) => void;
onModelChange: (modelId: CreationModelId) => void;
onModeChange: (modeId: CreationModeId) => void;
onAspectRatioChange: (aspectRatio: AspectRatioId) => void;
onResolutionChange: (resolution: ResolutionId) => void;
onDurationChange: (durationSec: number) => void;
onReferenceSelect?: (file: File | null) => void;
onFirstFrameSelect?: (file: File | null) => void;
onLastFrameSelect?: (file: File | null) => void;
onPresetGenerate?: (presetId: string) => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
onOpenProjects?: () => void;
capabilities?: LobbyCreationCapabilities;
}
export default function CreationStudio({
activeSection = "create",
onOpenProjects,
storyPresets = [],
onPresetGenerate,
isGenerating = false,
capabilities,
...composerProps
}: CreationStudioProps) {
return (
<div className="flex min-h-0 flex-1">
<AppNavRail activeSection={activeSection} onOpenProjects={onOpenProjects} />
<div className="min-w-0 flex-1 overflow-y-auto">
<div className="mx-auto flex w-full max-w-5xl flex-col gap-5 px-4 py-7 sm:px-6 sm:py-8">
<CreationComposer {...composerProps} isGenerating={isGenerating} capabilities={capabilities} />
{storyPresets.length > 0 && onPresetGenerate && (
<PresetQuickLaunchRail storyPresets={storyPresets} disabled={isGenerating} onPresetGenerate={onPresetGenerate} />
)}
</div>
</div>
</div>
);
}
@@ -0,0 +1,240 @@
"use client";
import React, { useCallback, useEffect, useRef, useState } from "react";
import { ChevronLeft, ChevronRight } from "lucide-react";
import { cn } from "@/lib/utils";
export interface StoryPresetLike {
id: string;
label: string;
description?: string;
segmentCount?: number;
styleTag?: string;
}
interface PresetQuickLaunchRailProps {
storyPresets: StoryPresetLike[];
disabled?: boolean;
onPresetGenerate: (presetId: string) => void;
}
export default function PresetQuickLaunchRail({
storyPresets,
disabled = false,
onPresetGenerate,
}: PresetQuickLaunchRailProps) {
const scrollRef = useRef<HTMLDivElement>(null);
const [canScrollLeft, setCanScrollLeft] = useState(false);
const [canScrollRight, setCanScrollRight] = useState(false);
const [presetRailDragging, setPresetRailDragging] = useState(false);
const presetDragStateRef = useRef({
pointerId: null as number | null,
startX: 0,
startScrollLeft: 0,
moved: false,
});
const suppressPresetClickRef = useRef(false);
const updateScrollState = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setCanScrollLeft(el.scrollLeft > 2);
setCanScrollRight(el.scrollLeft + el.clientWidth < el.scrollWidth - 2);
}, []);
const scrollByAmount = useCallback(
(direction: "left" | "right") => {
const el = scrollRef.current;
if (!el) return;
const delta = direction === "left" ? -220 : 220;
el.scrollBy({ left: delta, behavior: "smooth" });
window.setTimeout(updateScrollState, 220);
},
[updateScrollState],
);
const handlePresetWheel = useCallback(
(event: React.WheelEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
const dominantDelta = Math.abs(event.deltaX) > Math.abs(event.deltaY) ? event.deltaX : event.deltaY;
if (!dominantDelta) return;
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
const nextScrollLeft = Math.min(Math.max(el.scrollLeft + dominantDelta, 0), maxScrollLeft);
if (nextScrollLeft === el.scrollLeft) return;
event.preventDefault();
el.scrollLeft = nextScrollLeft;
updateScrollState();
},
[updateScrollState],
);
const finishPresetDrag = useCallback(() => {
presetDragStateRef.current = {
pointerId: null,
startX: 0,
startScrollLeft: 0,
moved: false,
};
setPresetRailDragging(false);
}, []);
const handlePresetPointerDown = useCallback((event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (event.pointerType !== "mouse" || event.button !== 0) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
suppressPresetClickRef.current = false;
presetDragStateRef.current = {
pointerId: event.pointerId,
startX: event.clientX,
startScrollLeft: el.scrollLeft,
moved: false,
};
}, []);
const handlePresetPointerMove = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
const dragState = presetDragStateRef.current;
if (!el || dragState.pointerId !== event.pointerId) return;
const deltaX = event.clientX - dragState.startX;
if (!dragState.moved && Math.abs(deltaX) > 4) {
dragState.moved = true;
suppressPresetClickRef.current = true;
setPresetRailDragging(true);
el.setPointerCapture?.(event.pointerId);
}
if (!dragState.moved) return;
event.preventDefault();
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
el.scrollLeft = Math.min(Math.max(dragState.startScrollLeft - deltaX, 0), maxScrollLeft);
updateScrollState();
},
[updateScrollState],
);
const handlePresetPointerUp = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el || presetDragStateRef.current.pointerId !== event.pointerId) return;
if (el.hasPointerCapture?.(event.pointerId)) {
el.releasePointerCapture(event.pointerId);
}
finishPresetDrag();
},
[finishPresetDrag],
);
const handlePresetClickCapture = useCallback((event: React.MouseEvent<HTMLDivElement>) => {
if (!suppressPresetClickRef.current) return;
suppressPresetClickRef.current = false;
event.preventDefault();
event.stopPropagation();
}, []);
useEffect(() => {
updateScrollState();
}, [storyPresets, updateScrollState]);
useEffect(() => {
const el = scrollRef.current;
if (!el) return;
const observer = new ResizeObserver(() => updateScrollState());
observer.observe(el);
return () => observer.disconnect();
}, [updateScrollState]);
if (storyPresets.length === 0) return null;
const scrollMaskStyle =
canScrollLeft && canScrollRight
? {
maskImage: "linear-gradient(to right, transparent, black 20px, black calc(100% - 20px), transparent)",
WebkitMaskImage: "linear-gradient(to right, transparent, black 20px, black calc(100% - 20px), transparent)",
}
: canScrollLeft
? {
maskImage: "linear-gradient(to right, transparent, black 20px, black)",
WebkitMaskImage: "linear-gradient(to right, transparent, black 20px, black)",
}
: canScrollRight
? {
maskImage: "linear-gradient(to right, black, black calc(100% - 20px), transparent)",
WebkitMaskImage: "linear-gradient(to right, black, black calc(100% - 20px), transparent)",
}
: undefined;
return (
<div className={cn("mx-auto w-full max-w-3xl transition-opacity duration-200", disabled && "pointer-events-none opacity-40")}>
<div className="grid grid-cols-[auto_minmax(0,1fr)_auto] items-center gap-1 sm:gap-2">
<div className="flex w-8 shrink-0 justify-center">
{canScrollLeft ? (
<button
type="button"
aria-label="Scroll suggested prompts left"
onClick={() => scrollByAmount("left")}
className="studio-control studio-control-press inline-flex size-8 items-center justify-center rounded-full text-muted-foreground hover-capable:hover:bg-muted/60 hover-capable:hover:text-foreground"
>
<ChevronLeft className="size-4" />
</button>
) : null}
</div>
<div
ref={scrollRef}
onScroll={updateScrollState}
onWheel={handlePresetWheel}
onPointerDown={handlePresetPointerDown}
onPointerMove={handlePresetPointerMove}
onPointerUp={handlePresetPointerUp}
onPointerCancel={handlePresetPointerUp}
onLostPointerCapture={finishPresetDrag}
onClickCapture={handlePresetClickCapture}
style={scrollMaskStyle}
className={cn(
"scrollbar-hidden flex gap-2 overflow-x-auto overflow-y-visible py-0.5 select-none",
presetRailDragging ? "cursor-grabbing" : "cursor-grab",
)}
>
{storyPresets.map((preset) => (
<button
key={preset.id}
type="button"
disabled={disabled}
onClick={() => onPresetGenerate(preset.id)}
className="studio-control studio-control-press studio-hover-surface flex w-[12.5rem] shrink-0 flex-col gap-1 rounded-xl border border-border/50 bg-card/70 px-3 py-2.5 text-left"
>
<span className="line-clamp-1 text-sm font-medium text-foreground">{preset.label}</span>
{preset.description && (
<span className="text-pretty line-clamp-2 text-xs leading-5 text-muted-foreground">{preset.description}</span>
)}
</button>
))}
</div>
<div className="flex w-8 shrink-0 justify-center">
{canScrollRight ? (
<button
type="button"
aria-label="Scroll suggested prompts right"
onClick={() => scrollByAmount("right")}
className="studio-control studio-control-press inline-flex size-8 items-center justify-center rounded-full text-muted-foreground hover-capable:hover:bg-muted/60 hover-capable:hover:text-foreground"
>
<ChevronRight className="size-4" />
</button>
) : null}
</div>
</div>
</div>
);
}
@@ -0,0 +1,97 @@
"use client";
import React, { useRef, useState } from "react";
import { ImagePlus } from "lucide-react";
import { REFERENCE_ACCEPT, isReferenceMediaFile } from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
interface ReferenceUploadSlotProps {
label: string;
sublabel?: string;
previewUrl?: string | null;
required?: boolean;
optional?: boolean;
disabled?: boolean;
onSelect?: (file: File | null) => void;
}
export default function ReferenceUploadSlot({
label,
sublabel,
previewUrl = null,
required = false,
optional = false,
disabled = false,
onSelect,
}: ReferenceUploadSlotProps) {
const fileInputRef = useRef<HTMLInputElement>(null);
const [dragActive, setDragActive] = useState(false);
function handleFile(file: File | null) {
if (!file || !isReferenceMediaFile(file)) return;
onSelect?.(file);
}
return (
<div className="flex flex-col gap-1">
<button
type="button"
aria-label={[label, sublabel].filter(Boolean).join(" ")}
onClick={() => fileInputRef.current?.click()}
disabled={disabled}
onDragEnter={(event) => {
event.preventDefault();
event.stopPropagation();
if (!disabled) setDragActive(true);
}}
onDragOver={(event) => {
event.preventDefault();
event.stopPropagation();
if (!disabled) setDragActive(true);
}}
onDragLeave={(event) => {
event.preventDefault();
event.stopPropagation();
setDragActive(false);
}}
onDrop={(event) => {
event.preventDefault();
event.stopPropagation();
setDragActive(false);
if (disabled) return;
handleFile(event.dataTransfer.files?.[0] ?? null);
}}
className={cn(
"studio-control studio-control-press studio-hover-surface relative flex size-[76px] shrink-0 flex-col items-center justify-center gap-1 overflow-hidden rounded-2xl border border-dashed bg-muted/50 px-1 text-center text-[11px] font-medium text-muted-foreground",
required && !previewUrl ? "border-amber-500/50" : "border-border/60",
dragActive && "border-accent-blue bg-accent-blue/10 ring-2 ring-accent-blue/30",
disabled && "pointer-events-none opacity-50",
)}
>
{previewUrl ? (
<img src={previewUrl} alt="" className="studio-media-outline absolute inset-0 size-full object-cover" />
) : (
<>
<ImagePlus className="size-4" />
<span>{label}</span>
{sublabel && <span className="text-[10px] font-normal opacity-70">{sublabel}</span>}
</>
)}
</button>
{(required || optional) && (
<span className="text-center text-[10px] text-muted-foreground">{required ? "Required" : "Optional"}</span>
)}
<input
ref={fileInputRef}
type="file"
accept={REFERENCE_ACCEPT}
className="hidden"
onChange={(event) => {
handleFile(event.target.files?.[0] ?? null);
event.target.value = "";
}}
/>
</div>
);
}
@@ -0,0 +1,213 @@
"use client";
import React from "react";
import { Box, ChevronDown, Clock, Monitor, Wand2 } from "lucide-react";
import ConfigPill from "@/components/creation/ConfigPill";
import {
DropdownMenu,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuTrigger,
} from "@/components/ui/dropdown-menu";
import { Popover, PopoverContent, PopoverTrigger } from "@/components/ui/popover";
import { Slider } from "@/components/ui/slider";
import {
ASPECT_RATIOS,
CREATION_MODELS,
CREATION_MODES,
RESOLUTIONS,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type ResolutionId,
formatDurationLabel,
formatResolutionLabel,
} from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
export interface SessionCreationConfig {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}
interface SessionCreationConfigPillsProps extends SessionCreationConfig {
disabled?: boolean;
readOnly?: boolean;
onModelChange?: (modelId: CreationModelId) => void;
onModeChange?: (modeId: CreationModeId) => void;
onAspectRatioChange?: (aspectRatio: AspectRatioId) => void;
onResolutionChange?: (resolution: ResolutionId) => void;
onDurationChange?: (durationSec: number) => void;
}
export default function SessionCreationConfigPills({
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
disabled = false,
readOnly = false,
onModelChange,
onModeChange,
onAspectRatioChange,
onResolutionChange,
onDurationChange,
}: SessionCreationConfigPillsProps) {
const selectedModel = CREATION_MODELS.find((model) => model.id === modelId) ?? CREATION_MODELS[0];
const selectedMode = CREATION_MODES.find((mode) => mode.id === modeId) ?? CREATION_MODES[0];
const isInteractive = !readOnly && !disabled;
const pillClassName = cn(
"h-9 min-h-9 px-2 text-[11px]",
!isInteractive && "pointer-events-none opacity-70",
);
if (readOnly) {
return (
<div className="flex flex-wrap items-center gap-1.5">
<ConfigPill disabled className={pillClassName} aria-label="Model">
<Box className="size-3" />
{selectedModel.label}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Mode">
<Wand2 className="size-3" />
{selectedMode.label}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Aspect ratio and resolution">
<Monitor className="size-3" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Duration">
<Clock className="size-3" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</div>
);
}
return (
<div className="flex flex-wrap items-center gap-1.5">
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Model">
<Box className="size-3" />
{selectedModel.label}
<ChevronDown className="size-2.5 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-72">
<DropdownMenuLabel>Model</DropdownMenuLabel>
<DropdownMenuSeparator />
{CREATION_MODELS.map((model) => (
<DropdownMenuItem key={model.id} onClick={() => onModelChange?.(model.id)} className="flex-col items-start gap-1 py-2.5">
<span className="flex items-center gap-2 text-sm font-medium">
{model.label}
{model.badge && <span className="rounded-full bg-accent-blue/15 px-1.5 py-0.5 text-[10px] text-accent-blue">{model.badge}</span>}
</span>
<span className="text-xs text-muted-foreground">{model.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Mode">
<Wand2 className="size-3" />
{selectedMode.label}
<ChevronDown className="size-2.5 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-64">
<DropdownMenuLabel>Mode</DropdownMenuLabel>
<DropdownMenuSeparator />
{CREATION_MODES.map((mode) => (
<DropdownMenuItem key={mode.id} onClick={() => onModeChange?.(mode.id)} className="flex-col items-start gap-1 py-2.5">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">{mode.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Aspect ratio and resolution">
<Monitor className="size-3" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-80">
<p className="mb-3 text-xs font-medium text-muted-foreground">Aspect ratio</p>
<div className="grid grid-cols-3 gap-2">
{ASPECT_RATIOS.map((ratio) => (
<button
key={ratio}
type="button"
onClick={() => onAspectRatioChange?.(ratio)}
className={cn(
"studio-control studio-control-press studio-hover-surface flex flex-col items-center gap-2 rounded-xl border px-2 py-3 text-xs",
aspectRatio === ratio ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
<span
className={cn(
"rounded-sm border border-current/40 bg-muted/40",
ratio === "9:16" && "h-7 w-4",
ratio === "16:9" && "h-4 w-7",
ratio === "1:1" && "size-5",
ratio === "4:3" && "h-5 w-6",
ratio === "3:4" && "h-6 w-5",
ratio === "21:9" && "h-3 w-8",
)}
/>
{ratio}
</button>
))}
</div>
<p className="mb-2 mt-4 text-xs font-medium text-muted-foreground">Resolution</p>
<div className="flex flex-wrap gap-2">
{RESOLUTIONS.map((item) => (
<button
key={item}
type="button"
onClick={() => onResolutionChange?.(item)}
className={cn(
"studio-control studio-control-press studio-hover-surface rounded-full border px-3 py-1.5 text-xs font-medium",
resolution === item ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
{formatResolutionLabel(item)}
</button>
))}
</div>
</PopoverContent>
</Popover>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Duration">
<Clock className="size-3" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-72">
<p className="mb-3 text-xs font-medium text-muted-foreground">Total duration</p>
<Slider min={5} max={15} step={5} value={[durationSec]} onValueChange={(values) => onDurationChange?.(values[0] ?? 5)} />
<div className="mt-3 flex items-center justify-between text-[11px] text-muted-foreground">
<span>5s</span>
<span className="rounded-md border border-border px-2 py-1 text-xs font-medium text-foreground">{formatDurationLabel(durationSec)}</span>
<span>15s</span>
</div>
</PopoverContent>
</Popover>
</div>
);
}
@@ -0,0 +1,141 @@
"use client";
import * as React from "react";
import * as DropdownMenuPrimitive from "@radix-ui/react-dropdown-menu";
import { Check, ChevronRight } from "lucide-react";
import { cn } from "@/lib/utils";
const DropdownMenu = DropdownMenuPrimitive.Root;
const DropdownMenuTrigger = DropdownMenuPrimitive.Trigger;
const DropdownMenuGroup = DropdownMenuPrimitive.Group;
const DropdownMenuPortal = DropdownMenuPrimitive.Portal;
const DropdownMenuSub = DropdownMenuPrimitive.Sub;
const DropdownMenuRadioGroup = DropdownMenuPrimitive.RadioGroup;
const DropdownMenuSubTrigger = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.SubTrigger>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.SubTrigger> & { inset?: boolean }
>(({ className, inset, children, ...props }, ref) => (
<DropdownMenuPrimitive.SubTrigger
ref={ref}
className={cn(
"flex cursor-default select-none items-center rounded-xl px-2 py-1.5 text-sm outline-none data-[state=open]:bg-accent focus:bg-accent",
inset && "pl-8",
className,
)}
{...props}
>
{children}
<ChevronRight className="ml-auto size-4" />
</DropdownMenuPrimitive.SubTrigger>
));
DropdownMenuSubTrigger.displayName = DropdownMenuPrimitive.SubTrigger.displayName;
const DropdownMenuSubContent = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.SubContent>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.SubContent>
>(({ className, ...props }, ref) => (
<DropdownMenuPrimitive.SubContent
ref={ref}
className={cn(
"z-50 min-w-[8rem] overflow-hidden rounded-2xl border border-border bg-popover/95 p-1 text-popover-foreground shadow-xl backdrop-blur-md",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
));
DropdownMenuSubContent.displayName = DropdownMenuPrimitive.SubContent.displayName;
const DropdownMenuContent = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Content>
>(({ className, sideOffset = 6, ...props }, ref) => (
<DropdownMenuPrimitive.Portal>
<DropdownMenuPrimitive.Content
ref={ref}
sideOffset={sideOffset}
className={cn(
"z-50 min-w-[12rem] overflow-hidden rounded-2xl border border-border bg-popover/95 p-1.5 text-popover-foreground shadow-xl backdrop-blur-md",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
</DropdownMenuPrimitive.Portal>
));
DropdownMenuContent.displayName = DropdownMenuPrimitive.Content.displayName;
const DropdownMenuItem = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Item>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Item> & { inset?: boolean }
>(({ className, inset, ...props }, ref) => (
<DropdownMenuPrimitive.Item
ref={ref}
className={cn(
"relative flex cursor-default select-none items-center gap-2 rounded-xl px-2.5 py-2 text-sm outline-none transition-colors data-[disabled]:pointer-events-none data-[disabled]:opacity-50 focus:bg-accent focus:text-accent-foreground",
inset && "pl-8",
className,
)}
{...props}
/>
));
DropdownMenuItem.displayName = DropdownMenuPrimitive.Item.displayName;
const DropdownMenuCheckboxItem = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.CheckboxItem>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.CheckboxItem>
>(({ className, children, checked, ...props }, ref) => (
<DropdownMenuPrimitive.CheckboxItem
ref={ref}
className={cn(
"relative flex cursor-default select-none items-center rounded-xl py-2 pl-8 pr-2 text-sm outline-none transition-colors data-[disabled]:pointer-events-none data-[disabled]:opacity-50 focus:bg-accent focus:text-accent-foreground",
className,
)}
checked={checked}
{...props}
>
<span className="absolute left-2 flex size-3.5 items-center justify-center">
<DropdownMenuPrimitive.ItemIndicator>
<Check className="size-4 text-accent-blue" />
</DropdownMenuPrimitive.ItemIndicator>
</span>
{children}
</DropdownMenuPrimitive.CheckboxItem>
));
DropdownMenuCheckboxItem.displayName = DropdownMenuPrimitive.CheckboxItem.displayName;
const DropdownMenuLabel = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Label>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Label> & { inset?: boolean }
>(({ className, inset, ...props }, ref) => (
<DropdownMenuPrimitive.Label ref={ref} className={cn("px-2 py-1.5 text-xs font-semibold text-muted-foreground", inset && "pl-8", className)} {...props} />
));
DropdownMenuLabel.displayName = DropdownMenuPrimitive.Label.displayName;
const DropdownMenuSeparator = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Separator>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Separator>
>(({ className, ...props }, ref) => (
<DropdownMenuPrimitive.Separator ref={ref} className={cn("-mx-1 my-1 h-px bg-border", className)} {...props} />
));
DropdownMenuSeparator.displayName = DropdownMenuPrimitive.Separator.displayName;
export {
DropdownMenu,
DropdownMenuTrigger,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuCheckboxItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuGroup,
DropdownMenuPortal,
DropdownMenuSub,
DropdownMenuSubContent,
DropdownMenuSubTrigger,
DropdownMenuRadioGroup,
};
@@ -0,0 +1,33 @@
"use client";
import * as React from "react";
import * as PopoverPrimitive from "@radix-ui/react-popover";
import { cn } from "@/lib/utils";
const Popover = PopoverPrimitive.Root;
const PopoverTrigger = PopoverPrimitive.Trigger;
const PopoverAnchor = PopoverPrimitive.Anchor;
const PopoverContent = React.forwardRef<
React.ElementRef<typeof PopoverPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof PopoverPrimitive.Content>
>(({ className, align = "center", sideOffset = 6, ...props }, ref) => (
<PopoverPrimitive.Portal>
<PopoverPrimitive.Content
ref={ref}
align={align}
sideOffset={sideOffset}
className={cn(
"z-50 w-72 rounded-2xl border border-border bg-popover/95 p-3 text-popover-foreground shadow-xl backdrop-blur-md outline-none",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
</PopoverPrimitive.Portal>
));
PopoverContent.displayName = PopoverPrimitive.Content.displayName;
export { Popover, PopoverTrigger, PopoverContent, PopoverAnchor };
@@ -0,0 +1,25 @@
"use client";
import * as React from "react";
import * as SliderPrimitive from "@radix-ui/react-slider";
import { cn } from "@/lib/utils";
const Slider = React.forwardRef<
React.ElementRef<typeof SliderPrimitive.Root>,
React.ComponentPropsWithoutRef<typeof SliderPrimitive.Root>
>(({ className, ...props }, ref) => (
<SliderPrimitive.Root
ref={ref}
className={cn("relative flex w-full touch-none select-none items-center", className)}
{...props}
>
<SliderPrimitive.Track className="relative h-1.5 w-full grow overflow-hidden rounded-full bg-muted">
<SliderPrimitive.Range className="absolute h-full bg-accent-blue" />
</SliderPrimitive.Track>
<SliderPrimitive.Thumb className="block size-4 rounded-full border border-accent-blue/40 bg-background shadow transition-colors focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40 disabled:pointer-events-none disabled:opacity-50" />
</SliderPrimitive.Root>
));
Slider.displayName = SliderPrimitive.Root.displayName;
export { Slider };
@@ -0,0 +1,48 @@
"use client";
import * as React from "react";
import * as TabsPrimitive from "@radix-ui/react-tabs";
import { cn } from "@/lib/utils";
const Tabs = TabsPrimitive.Root;
const TabsList = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.List>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.List>
>(({ className, ...props }, ref) => (
<TabsPrimitive.List
ref={ref}
className={cn("inline-flex items-center gap-1 rounded-full bg-muted/60 p-1 text-muted-foreground", className)}
{...props}
/>
));
TabsList.displayName = TabsPrimitive.List.displayName;
const TabsTrigger = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.Trigger>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.Trigger>
>(({ className, ...props }, ref) => (
<TabsPrimitive.Trigger
ref={ref}
className={cn(
"inline-flex items-center justify-center rounded-full px-3 py-1.5 text-xs font-medium whitespace-nowrap transition-all",
"focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40",
"disabled:pointer-events-none disabled:opacity-50",
"data-[state=active]:bg-card data-[state=active]:text-foreground data-[state=active]:shadow-sm",
className,
)}
{...props}
/>
));
TabsTrigger.displayName = TabsPrimitive.Trigger.displayName;
const TabsContent = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.Content>
>(({ className, ...props }, ref) => (
<TabsPrimitive.Content ref={ref} className={cn("mt-4 outline-none", className)} {...props} />
));
TabsContent.displayName = TabsPrimitive.Content.displayName;
export { Tabs, TabsList, TabsTrigger, TabsContent };
@@ -0,0 +1,80 @@
import { describe, expect, it } from "vitest";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
clampLobbySelectionToCapabilities,
parseLobbyCapabilitiesBundle,
resolveModelCapabilities,
validateLobbyCreationSelection,
} from "@/lib/creationCapabilities";
describe("creationCapabilities", () => {
it("parses backend capability payloads with per-model caps", () => {
const bundle = parseLobbyCapabilitiesBundle({
model_ids: ["fast-ltx2", "fast-h3"],
models: {
"fast-ltx2": {
generation_modes: ["t2va"],
resolutions: ["480p", "720p"],
duration_sec: [5, 10],
},
"fast-h3": {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["16:9"],
resolutions: ["720p"],
},
},
});
expect(bundle.model_ids).toEqual(["fast-ltx2", "fast-h3"]);
expect(bundle.models["fast-h3"]?.aspect_ratios).toEqual(["16:9"]);
});
it("includes fast-h3 in default lobby models", () => {
expect(DEFAULT_LOBBY_CAPABILITIES_BUNDLE.model_ids).toContain("fast-h3");
});
it("clamps unsupported lobby selections to model-specific defaults", () => {
expect(
clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-h3"),
modelId: "fast-h3",
modeId: "fl2av",
aspectRatio: "9:16",
resolution: "4k",
durationSec: 99,
}),
).toEqual({
modelId: "fast-h3",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
});
});
it("rejects unsupported generation modes with a clear message", () => {
expect(
validateLobbyCreationSelection({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-ltx23"),
modelId: "fast-ltx23",
modeId: "fl2av",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
}),
).toMatch(/FL2VA/i);
});
it("rejects unsupported resolutions for ltx models", () => {
expect(
validateLobbyCreationSelection({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-ltx23"),
modelId: "fast-ltx23",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "4k",
durationSec: 5,
}),
).toMatch(/resolution/i);
});
});
@@ -0,0 +1,263 @@
import type {
AspectRatioId,
CreationModeId,
CreationModelId,
ResolutionId,
} from "@/lib/creationConfig";
import { fromGenerationMode, toGenerationMode, type GenerationMode } from "@/lib/generationMode";
const ALL_MODEL_IDS: CreationModelId[] = ["fast-ltx23", "fast-ltx2", "fast-h3"];
const ALL_GENERATION_MODES: GenerationMode[] = ["t2va", "fl2va", "ref2va"];
const ALL_ASPECT_RATIOS: AspectRatioId[] = ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"];
const ALL_RESOLUTIONS: ResolutionId[] = ["480p", "720p", "1080p", "4k"];
export interface ModelCreationCapabilities {
generation_modes: GenerationMode[];
aspect_ratios: AspectRatioId[];
resolutions: ResolutionId[];
duration_sec: number[];
unsupported_generation_modes: Record<string, string>;
reference_assets: {
mime_types: string[];
max_bytes: number;
};
}
export interface LobbyCreationCapabilities extends ModelCreationCapabilities {
model_ids: CreationModelId[];
}
export interface LobbyCapabilitiesBundle {
model_ids: CreationModelId[];
models: Partial<Record<CreationModelId, ModelCreationCapabilities>>;
generation_modes: GenerationMode[];
aspect_ratios: AspectRatioId[];
resolutions: ResolutionId[];
duration_sec: number[];
unsupported_generation_modes: Record<string, string>;
reference_assets: {
mime_types: string[];
max_bytes: number;
};
}
const DEFAULT_LTX_MODEL_CAPABILITIES: ModelCreationCapabilities = {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"],
resolutions: ["480p", "720p", "1080p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: {
fl2va: "First/last frame mode (FL2VA) is not supported yet.",
},
reference_assets: {
mime_types: ["image/png", "image/jpeg", "image/webp"],
max_bytes: 15 * 1024 * 1024,
},
};
const DEFAULT_H3_MODEL_CAPABILITIES: ModelCreationCapabilities = {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["16:9"],
resolutions: ["720p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: {
fl2va: "First/last frame mode (FL2VA) is not supported yet.",
},
reference_assets: DEFAULT_LTX_MODEL_CAPABILITIES.reference_assets,
};
export const DEFAULT_LOBBY_CAPABILITIES_BUNDLE: LobbyCapabilitiesBundle = {
model_ids: ALL_MODEL_IDS,
models: {
"fast-ltx2": DEFAULT_LTX_MODEL_CAPABILITIES,
"fast-ltx23": DEFAULT_LTX_MODEL_CAPABILITIES,
"fast-h3": DEFAULT_H3_MODEL_CAPABILITIES,
},
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"],
resolutions: ["480p", "720p", "1080p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: DEFAULT_LTX_MODEL_CAPABILITIES.unsupported_generation_modes,
reference_assets: DEFAULT_LTX_MODEL_CAPABILITIES.reference_assets,
};
function pickStrings<T extends string>(value: unknown, allowed: readonly T[], fallback: readonly T[]): T[] {
if (!Array.isArray(value)) return [...fallback];
return value.filter((item): item is T => typeof item === "string" && allowed.includes(item as T));
}
function parseReferenceAssets(
value: unknown,
fallback: ModelCreationCapabilities["reference_assets"],
): ModelCreationCapabilities["reference_assets"] {
if (!value || typeof value !== "object") return fallback;
const data = value as Record<string, unknown>;
return {
mime_types: Array.isArray(data.mime_types)
? (data.mime_types as string[])
: fallback.mime_types,
max_bytes: typeof data.max_bytes === "number" ? data.max_bytes : fallback.max_bytes,
};
}
function parseModelCreationCapabilities(
value: unknown,
fallback: ModelCreationCapabilities,
): ModelCreationCapabilities {
if (!value || typeof value !== "object") return fallback;
const data = value as Record<string, unknown>;
return {
generation_modes: pickStrings(data.generation_modes, ALL_GENERATION_MODES, fallback.generation_modes),
aspect_ratios: pickStrings(data.aspect_ratios, ALL_ASPECT_RATIOS, fallback.aspect_ratios),
resolutions: pickStrings(data.resolutions, ALL_RESOLUTIONS, fallback.resolutions),
duration_sec: Array.isArray(data.duration_sec)
? data.duration_sec.filter((item): item is number => typeof item === "number")
: fallback.duration_sec,
unsupported_generation_modes:
typeof data.unsupported_generation_modes === "object" && data.unsupported_generation_modes
? (data.unsupported_generation_modes as Record<string, string>)
: fallback.unsupported_generation_modes,
reference_assets: parseReferenceAssets(data.reference_assets, fallback.reference_assets),
};
}
export function parseLobbyCapabilitiesBundle(payload: unknown): LobbyCapabilitiesBundle {
if (!payload || typeof payload !== "object") {
return DEFAULT_LOBBY_CAPABILITIES_BUNDLE;
}
const data = payload as Record<string, unknown>;
const modelIds = pickStrings(data.model_ids, ALL_MODEL_IDS, DEFAULT_LOBBY_CAPABILITIES_BUNDLE.model_ids);
const rawModels = typeof data.models === "object" && data.models ? (data.models as Record<string, unknown>) : {};
const models: Partial<Record<CreationModelId, ModelCreationCapabilities>> = {};
for (const modelId of modelIds) {
const fallback =
DEFAULT_LOBBY_CAPABILITIES_BUNDLE.models[modelId] ??
(modelId === "fast-h3" ? DEFAULT_H3_MODEL_CAPABILITIES : DEFAULT_LTX_MODEL_CAPABILITIES);
models[modelId] = parseModelCreationCapabilities(rawModels[modelId], fallback);
}
const unionFallback = parseModelCreationCapabilities(payload, DEFAULT_LTX_MODEL_CAPABILITIES);
return {
model_ids: modelIds,
models,
generation_modes: unionFallback.generation_modes,
aspect_ratios: unionFallback.aspect_ratios,
resolutions: unionFallback.resolutions,
duration_sec: unionFallback.duration_sec,
unsupported_generation_modes: unionFallback.unsupported_generation_modes,
reference_assets: unionFallback.reference_assets,
};
}
export function resolveModelCapabilities(
bundle: LobbyCapabilitiesBundle,
modelId: CreationModelId,
): LobbyCreationCapabilities {
const modelCaps =
bundle.models[modelId] ??
(modelId === "fast-h3" ? DEFAULT_H3_MODEL_CAPABILITIES : DEFAULT_LTX_MODEL_CAPABILITIES);
return {
model_ids: bundle.model_ids,
...modelCaps,
};
}
export function supportedCreationModes(capabilities: LobbyCreationCapabilities) {
return capabilities.generation_modes.map((wireMode) => ({
wireMode,
modeId: fromGenerationMode(wireMode),
}));
}
export function isSupportedCreationMode(modeId: CreationModeId, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.generation_modes.includes(toGenerationMode(modeId));
}
export function isSupportedResolution(resolution: ResolutionId, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.resolutions.includes(resolution);
}
export function isSupportedReferenceImage(file: File, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.reference_assets.mime_types.includes(file.type);
}
export function unsupportedModeNotice(modeId: CreationModeId, capabilities: LobbyCreationCapabilities): string | null {
const wireMode = toGenerationMode(modeId);
return capabilities.unsupported_generation_modes[wireMode] ?? null;
}
export function clampLobbySelectionToCapabilities(input: {
capabilities: LobbyCreationCapabilities;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}): {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
} {
const { capabilities } = input;
const modelId = capabilities.model_ids.includes(input.modelId)
? input.modelId
: (capabilities.model_ids[0] ?? "fast-ltx23");
const supportedModes = supportedCreationModes(capabilities);
const modeId = isSupportedCreationMode(input.modeId, capabilities)
? input.modeId
: (supportedModes[0]?.modeId ?? "t2v");
const aspectRatio = capabilities.aspect_ratios.includes(input.aspectRatio)
? input.aspectRatio
: (capabilities.aspect_ratios[0] ?? "16:9");
const resolution = isSupportedResolution(input.resolution, capabilities)
? input.resolution
: (capabilities.resolutions[0] ?? "720p");
const durationSec = capabilities.duration_sec.includes(input.durationSec)
? input.durationSec
: (capabilities.duration_sec[0] ?? 5);
return { modelId, modeId, aspectRatio, resolution, durationSec };
}
export function validateLobbyCreationSelection(input: {
capabilities: LobbyCreationCapabilities;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): string | null {
const unsupportedMode = unsupportedModeNotice(input.modeId, input.capabilities);
if (unsupportedMode) return unsupportedMode;
if (!input.capabilities.model_ids.includes(input.modelId)) {
return "Selected model is not supported yet.";
}
if (!isSupportedCreationMode(input.modeId, input.capabilities)) {
return "Selected mode is not supported yet.";
}
if (!input.capabilities.aspect_ratios.includes(input.aspectRatio)) {
return "Selected aspect ratio is not supported for this model yet.";
}
if (!isSupportedResolution(input.resolution, input.capabilities)) {
return "Selected resolution is not supported for this model yet.";
}
if (!input.capabilities.duration_sec.includes(input.durationSec)) {
return "Selected duration is not supported yet.";
}
if (input.modeId === "ref2av" && !input.referenceFile) {
return "Upload a reference image to use reference-guided mode.";
}
if (input.referenceFile && !isSupportedReferenceImage(input.referenceFile, input.capabilities)) {
return "Reference assets must be PNG, JPEG, or WebP images.";
}
if (input.firstFrameFile && !isSupportedReferenceImage(input.firstFrameFile, input.capabilities)) {
return "First frame must be a PNG, JPEG, or WebP image.";
}
if (input.lastFrameFile && !isSupportedReferenceImage(input.lastFrameFile, input.capabilities)) {
return "Last frame must be a PNG, JPEG, or WebP image.";
}
return null;
}
@@ -0,0 +1,62 @@
import { describe, expect, it } from "vitest";
import {
CREATION_MODELS,
buildMentionOptions,
formatDurationLabel,
formatResolutionLabel,
isReferenceMediaFile,
modeRequiresReference,
modeUsesDualFrames,
} from "@/lib/creationConfig";
describe("creationConfig", () => {
it("formats resolution labels", () => {
expect(formatResolutionLabel("480p")).toBe("480P");
expect(formatResolutionLabel("720p")).toBe("720P");
expect(formatResolutionLabel("4k")).toBe("4K");
});
it("formats duration labels", () => {
expect(formatDurationLabel(5)).toBe("5s");
});
it("includes all Dreamverse lobby models", () => {
expect(CREATION_MODELS.map((model) => model.id)).toEqual(["fast-ltx23", "fast-ltx2", "fast-h3"]);
});
it("builds mention options from presets", () => {
expect(
buildMentionOptions([
{ id: "preset-a", label: "Preset A", description: "A short preset" },
{ label: "Missing id" },
]),
).toEqual([
{
id: "preset-a",
label: "Preset A",
kind: "preset",
description: "A short preset",
},
{
id: "Missing id",
label: "Missing id",
kind: "preset",
description: undefined,
},
]);
});
it("derives mode-specific reference requirements", () => {
expect(modeRequiresReference("ref2av")).toBe(true);
expect(modeRequiresReference("t2v")).toBe(false);
expect(modeUsesDualFrames("fl2av")).toBe(true);
expect(modeUsesDualFrames("t2v")).toBe(false);
});
it("accepts image reference files only", () => {
expect(isReferenceMediaFile(new File(["x"], "a.png", { type: "image/png" }))).toBe(true);
expect(isReferenceMediaFile(new File(["x"], "a.mp4", { type: "video/mp4" }))).toBe(false);
expect(isReferenceMediaFile(new File(["x"], "a.txt", { type: "text/plain" }))).toBe(false);
});
});
@@ -0,0 +1,96 @@
export type CreationModeId = "t2v" | "fl2av" | "ref2av";
export type CreationModelId = "fast-ltx2" | "fast-ltx23" | "fast-h3";
export type AspectRatioId = "21:9" | "16:9" | "4:3" | "1:1" | "3:4" | "9:16";
export type ResolutionId = "480p" | "720p" | "1080p" | "4k";
export interface CreationModeOption {
id: CreationModeId;
label: string;
description: string;
}
export interface CreationModelOption {
id: CreationModelId;
label: string;
description: string;
badge?: string;
}
export interface MentionOption {
id: string;
label: string;
kind: "preset" | "asset" | "character";
description?: string;
}
export const CREATION_MODES: CreationModeOption[] = [
{ id: "t2v", label: "Text to video", description: "Generate from a text prompt" },
{ id: "ref2av", label: "Image to video", description: "Guide the first segment with a reference image" },
];
export const UNSUPPORTED_CREATION_MODES: CreationModeOption[] = [
{ id: "fl2av", label: "First and last frame", description: "Coming soon on FastLTX models" },
];
export const CREATION_MODELS: CreationModelOption[] = [
{
id: "fast-ltx23",
label: "FastLTX 2.3",
description: "LTX 2.3 with OmniNFT LoRA",
badge: "New",
},
{
id: "fast-ltx2",
label: "FastLTX 2",
description: "FastLTX 2 for streaming",
},
{
id: "fast-h3",
label: "FastH3",
description: "MiniMax H3 with VSA data-free adapter",
},
];
export const ASPECT_RATIOS: AspectRatioId[] = ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"];
export const RESOLUTIONS: ResolutionId[] = ["480p", "720p", "1080p"];
export const UNSUPPORTED_RESOLUTIONS: ResolutionId[] = ["4k"];
export const DURATION_MARKS = [5, 10, 15] as const;
export const REFERENCE_ACCEPT = "image/png,image/jpeg,image/webp";
export function formatResolutionLabel(resolution: ResolutionId): string {
return resolution === "4k" ? "4K" : resolution.toUpperCase();
}
export function formatDurationLabel(seconds: number): string {
return `${seconds}s`;
}
export function modeRequiresReference(modeId: CreationModeId): boolean {
return modeId === "ref2av";
}
export function modeUsesDualFrames(modeId: CreationModeId): boolean {
return modeId === "fl2av";
}
export function isReferenceMediaFile(file: File): boolean {
return file.type === "image/png" || file.type === "image/jpeg" || file.type === "image/webp";
}
export function buildMentionOptions(storyPresets: Array<{ id?: string; label?: string; description?: string }>): MentionOption[] {
return storyPresets
.filter((preset) => typeof preset.label === "string" && preset.label.trim())
.map((preset) => ({
id: String(preset.id || preset.label),
label: String(preset.label),
kind: "preset" as const,
description: typeof preset.description === "string" ? preset.description : undefined,
}));
}
@@ -0,0 +1,66 @@
import { describe, expect, it } from "vitest";
import { parseEchoedCreationConfig, validateCreationInputs } from "@/lib/creationPayload";
describe("creationPayload", () => {
it("requires a reference asset for omni reference mode", () => {
expect(
validateCreationInputs({
modeId: "ref2av",
referenceFile: null,
}),
).toMatch(/reference asset/i);
});
it("requires both frames for first and last frame mode", () => {
expect(
validateCreationInputs({
modeId: "fl2av",
firstFrameFile: new File(["a"], "first.png", { type: "image/png" }),
lastFrameFile: null,
}),
).toMatch(/both first and last/i);
});
it("accepts text to video without references", () => {
expect(
validateCreationInputs({
modeId: "t2v",
}),
).toBeNull();
});
it("parses echoed creation config from server payloads", () => {
expect(
parseEchoedCreationConfig({
type: "gpu_assigned",
creation_config: {
model_id: "fast-ltx2",
generation_mode: "ref2va",
aspect_ratio: "9:16",
resolution: "480p",
duration_sec: 10,
},
}),
).toEqual({
modelId: "fast-ltx2",
modeId: "ref2av",
aspectRatio: "9:16",
resolution: "480p",
durationSec: 10,
});
});
it("ignores invalid echoed creation config", () => {
expect(parseEchoedCreationConfig({ creation_config: { model_id: "unknown" } })).toBeNull();
});
it("rejects unsupported reference mime types", () => {
expect(
validateCreationInputs({
modeId: "t2v",
referenceFile: new File(["a"], "clip.mp4", { type: "video/mp4" }),
}),
).toMatch(/PNG, JPEG, or WebP/i);
});
});
@@ -0,0 +1,172 @@
import type {
AspectRatioId,
CreationModeId,
CreationModelId,
ResolutionId,
} from "@/lib/creationConfig";
import { fromGenerationMode, type GenerationMode } from "@/lib/generationMode";
const LOBBY_MODEL_IDS = new Set<CreationModelId>(["fast-ltx2", "fast-ltx23", "fast-h3"]);
const ASPECT_RATIO_IDS = new Set<AspectRatioId>(["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"]);
const RESOLUTION_IDS = new Set<ResolutionId>(["480p", "720p", "1080p", "4k"]);
const DURATION_SEC_VALUES = new Set([5, 10, 15]);
export interface EchoedSessionCreationConfig {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}
const MAX_IMAGE_BYTES = 15 * 1024 * 1024;
const SUPPORTED_IMAGE_TYPES = new Set(["image/png", "image/jpeg", "image/webp"]);
export interface InitialImagePayload {
name: string;
mime_type: string;
data_url: string;
}
export interface CreationInitPayload {
model_id: string;
aspect_ratio: string;
resolution: string;
duration_sec: number;
initial_image: InitialImagePayload | null;
last_frame_image: InitialImagePayload | null;
}
function readFileAsDataUrl(file: File): Promise<string> {
return new Promise((resolve, reject) => {
const reader = new FileReader();
reader.onload = () => {
if (typeof reader.result === "string") {
resolve(reader.result);
return;
}
reject(new Error("Failed to read reference image."));
};
reader.onerror = () => reject(new Error("Failed to read reference image."));
reader.readAsDataURL(file);
});
}
export async function fileToInitialImagePayload(file: File): Promise<InitialImagePayload> {
if (!SUPPORTED_IMAGE_TYPES.has(file.type)) {
throw new Error("Reference assets must be PNG, JPEG, or WebP images.");
}
if (file.size > MAX_IMAGE_BYTES) {
throw new Error("Reference image must be 15 MB or smaller.");
}
return {
name: file.name,
mime_type: file.type,
data_url: await readFileAsDataUrl(file),
};
}
export async function resolveCreationImages(input: {
modeId: CreationModeId;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): Promise<Pick<CreationInitPayload, "initial_image" | "last_frame_image">> {
if (input.modeId === "fl2av") {
const firstFrame = input.firstFrameFile ? await fileToInitialImagePayload(input.firstFrameFile) : null;
const lastFrame = input.lastFrameFile ? await fileToInitialImagePayload(input.lastFrameFile) : null;
return {
initial_image: firstFrame,
last_frame_image: lastFrame,
};
}
const reference = input.referenceFile ? await fileToInitialImagePayload(input.referenceFile) : null;
return {
initial_image: reference,
last_frame_image: null,
};
}
export function validateCreationInputs(input: {
modeId: CreationModeId;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): string | null {
if (input.modeId === "ref2av" && !input.referenceFile) {
return "Upload a reference asset to use Omni reference mode.";
}
if (input.modeId === "fl2av") {
if (!input.firstFrameFile || !input.lastFrameFile) {
return "Upload both first and last frame assets.";
}
}
if (input.referenceFile && !SUPPORTED_IMAGE_TYPES.has(input.referenceFile.type)) {
return "Reference assets must be PNG, JPEG, or WebP images.";
}
if (input.firstFrameFile && !SUPPORTED_IMAGE_TYPES.has(input.firstFrameFile.type)) {
return "First frame must be a PNG, JPEG, or WebP image.";
}
if (input.lastFrameFile && !SUPPORTED_IMAGE_TYPES.has(input.lastFrameFile.type)) {
return "Last frame must be a PNG, JPEG, or WebP image.";
}
return null;
}
export function parseEchoedCreationConfig(data: unknown): EchoedSessionCreationConfig | null {
if (!data || typeof data !== "object") {
return null;
}
const creationConfig = (data as Record<string, unknown>).creation_config;
if (!creationConfig || typeof creationConfig !== "object") {
return null;
}
const config = creationConfig as Record<string, unknown>;
const modelId = typeof config.model_id === "string" && LOBBY_MODEL_IDS.has(config.model_id as CreationModelId)
? (config.model_id as CreationModelId)
: null;
const generationMode = typeof config.generation_mode === "string" ? config.generation_mode as GenerationMode : null;
const modeId = generationMode === "t2va" || generationMode === "fl2va" || generationMode === "ref2va"
? fromGenerationMode(generationMode)
: null;
const aspectRatio = typeof config.aspect_ratio === "string" && ASPECT_RATIO_IDS.has(config.aspect_ratio as AspectRatioId)
? (config.aspect_ratio as AspectRatioId)
: null;
const resolution = typeof config.resolution === "string" && RESOLUTION_IDS.has(config.resolution as ResolutionId)
? (config.resolution as ResolutionId)
: null;
const durationSec = typeof config.duration_sec === "number" && DURATION_SEC_VALUES.has(config.duration_sec)
? config.duration_sec
: null;
if (modelId === null || modeId === null || aspectRatio === null || resolution === null || durationSec === null) {
return null;
}
return {
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
};
}
export async function buildCreationInitPayload(input: {
modelId: string;
modeId: CreationModeId;
aspectRatio: string;
resolution: string;
durationSec: number;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): Promise<CreationInitPayload> {
const images = await resolveCreationImages(input);
return {
model_id: input.modelId,
aspect_ratio: input.aspectRatio,
resolution: input.resolution,
duration_sec: input.durationSec,
...images,
};
}
@@ -0,0 +1,39 @@
import { describe, expect, it } from "vitest";
import {
DEFAULT_GENERATION_MODE,
GENERATION_MODES,
fromGenerationMode,
getGenerationMode,
isGenerationMode,
toGenerationMode,
} from "./generationMode";
describe("generation modes", () => {
it("exposes stable wire IDs in the expected product order", () => {
expect(GENERATION_MODES.map((mode) => mode.id)).toEqual([
"t2va",
"fl2va",
"ref2va",
]);
expect(DEFAULT_GENERATION_MODE).toBe("t2va");
});
it("validates and resolves generation mode values", () => {
expect(isGenerationMode("ref2va")).toBe(true);
expect(isGenerationMode("unknown")).toBe(false);
expect(getGenerationMode("fl2va").label).toBe("FL2VA");
});
it("maps creation studio mode IDs to upstream wire values", () => {
expect(toGenerationMode("t2v")).toBe("t2va");
expect(toGenerationMode("fl2av")).toBe("fl2va");
expect(toGenerationMode("ref2av")).toBe("ref2va");
});
it("maps upstream wire values back to creation studio mode IDs", () => {
expect(fromGenerationMode("t2va")).toBe("t2v");
expect(fromGenerationMode("fl2va")).toBe("fl2av");
expect(fromGenerationMode("ref2va")).toBe("ref2av");
});
});
@@ -0,0 +1,54 @@
import type { CreationModeId } from "@/lib/creationConfig";
export const GENERATION_MODES = [
{
id: "t2va",
label: "T2VA",
name: "Text to video + audio",
description: "Start with a text prompt; no reference asset is required.",
},
{
id: "fl2va",
label: "FL2VA",
name: "First/last frames to video + audio",
description: "Provide first and last frame images to control the transition.",
},
{
id: "ref2va",
label: "Ref2VA",
name: "References to video + audio",
description: "Guide the result with ordered image, video, or audio references.",
},
] as const;
export type GenerationMode = (typeof GENERATION_MODES)[number]["id"];
export const DEFAULT_GENERATION_MODE: GenerationMode = "t2va";
const CREATION_MODE_TO_GENERATION_MODE: Record<CreationModeId, GenerationMode> = {
t2v: "t2va",
fl2av: "fl2va",
ref2av: "ref2va",
};
export function isGenerationMode(value: unknown): value is GenerationMode {
return GENERATION_MODES.some((mode) => mode.id === value);
}
export function getGenerationMode(value: GenerationMode) {
return GENERATION_MODES.find((mode) => mode.id === value) ?? GENERATION_MODES[0];
}
const GENERATION_MODE_TO_CREATION_MODE: Record<GenerationMode, CreationModeId> = {
t2va: "t2v",
fl2va: "fl2av",
ref2va: "ref2av",
};
export function fromGenerationMode(mode: GenerationMode): CreationModeId {
return GENERATION_MODE_TO_CREATION_MODE[mode];
}
export function toGenerationMode(modeId: CreationModeId): GenerationMode {
return CREATION_MODE_TO_GENERATION_MODE[modeId];
}
+20 -4
View File
@@ -9,6 +9,7 @@ from __future__ import annotations
import contextlib
import logging
import json
import sqlite3
import threading
from pathlib import Path
@@ -46,8 +47,13 @@ DEFAULT_SETTINGS: dict[str, Any] = {
def _sqlite_row_get(row: sqlite3.Row, key: str, default: Any) -> Any:
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10)."""
return row[key] if key in row else default # noqa: SIM401
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10).
NOTE: `key in row` tests Row *values*, not column names, so the membership
check has to go through .keys() -- otherwise every lookup falls back to the
default and jobs restored from the database lose their stored fields.
"""
return row[key] if key in row.keys() else default # noqa: SIM401, SIM118
def _get_db_path(data_dir: Path) -> Path:
@@ -83,6 +89,9 @@ def _migrate_db(conn: sqlite3.Connection) -> None:
_add_column_if_missing(conn, "jobs", "fps", "INTEGER", "24")
_add_column_if_missing(conn, "jobs", "workload_type", "TEXT", "'t2v'")
_add_column_if_missing(conn, "jobs", "image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "name", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "last_image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "references_json", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "job_type", "TEXT", "'inference'")
_add_column_if_missing(conn, "jobs", "data_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "max_train_steps", "INTEGER", "1000")
@@ -242,7 +251,8 @@ class Database:
self._execute(
"""
INSERT INTO jobs (
id, model_id, prompt, workload_type, image_path, job_type, status,
id, model_id, name, prompt, workload_type, image_path,
last_image_path, references_json, job_type, status,
created_at, started_at, finished_at, error, output_path, log_file_path,
num_inference_steps, num_frames, height, width, guidance_scale,
guidance_rescale, fps, seed, num_gpus, dit_cpu_offload,
@@ -254,14 +264,17 @@ class Database:
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
real_score_guidance_scale,
generator_update_interval, real_score_model_path, fake_score_model_path
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
job["id"],
job["model_id"],
job.get("name", ""),
job["prompt"],
job.get("workload_type", "t2v"),
job.get("image_path", ""),
job.get("last_image_path", ""),
json.dumps(job.get("references") or []),
job.get("job_type", "inference"),
job["status"],
job["created_at"],
@@ -540,9 +553,12 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
result = {
"id": row["id"],
"model_id": row["model_id"],
"name": _sqlite_row_get(row, "name", "") or "",
"prompt": row["prompt"],
"workload_type": _sqlite_row_get(row, "workload_type", "t2v"),
"image_path": _sqlite_row_get(row, "image_path", "") or "",
"last_image_path": _sqlite_row_get(row, "last_image_path", "") or "",
"references": _sqlite_row_get(row, "references_json", "") or "",
"job_type": _sqlite_row_get(row, "job_type", "inference"),
"status": row["status"],
"created_at": row["created_at"],
+239 -58
View File
@@ -10,7 +10,9 @@ from __future__ import annotations
import atexit
import collections
import contextlib
import copy
import enum
import json
import logging
import logging.handlers
import multiprocessing as mp
@@ -123,10 +125,13 @@ class LogBufferHandler(logging.Handler):
class Job:
id: str
model_id: str
prompt: str
name: str = ""
prompt: str = ""
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] = field(default_factory=list)
status: JobStatus = JobStatus.PENDING
created_at: float = field(default_factory=time.time)
started_at: float | None = None
@@ -145,6 +150,7 @@ class Job:
negative_prompt: str = ""
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
@@ -180,10 +186,13 @@ class Job:
return {
"id": self.id,
"model_id": self.model_id,
"name": self.name,
"prompt": self.prompt,
"workload_type": self.workload_type,
"job_type": self.job_type,
"image_path": self.image_path,
"last_image_path": self.last_image_path,
"references": self.references,
"status": self.status.value,
"created_at": self.created_at,
"started_at": self.started_at,
@@ -202,6 +211,7 @@ class Job:
"negative_prompt": self.negative_prompt,
"num_gpus": self.num_gpus,
"dit_cpu_offload": self.dit_cpu_offload,
"dit_layerwise_offload": self.dit_layerwise_offload,
"text_encoder_cpu_offload": self.text_encoder_cpu_offload,
"vae_cpu_offload": self.vae_cpu_offload,
"image_encoder_cpu_offload": self.image_encoder_cpu_offload,
@@ -232,6 +242,75 @@ class Job:
}
MINIMAX_H3_REF2VA_PIPELINE = "MiniMaxH3Ref2VAModularPipeline"
def _build_h3_references(raw: list[dict[str, Any]]) -> list[Any]:
"""Turn the API's reference dicts into MiniMaxH3Reference objects.
Imported lazily so the API server starts without pulling in fastvideo.
"""
from fastvideo.pipelines.basic.minimax_h3 import MiniMaxH3Reference
built = []
for i, ref in enumerate(raw):
source = (ref or {}).get("source")
if not source:
raise ValueError(f"reference {i} has no source")
if not os.path.isfile(source):
raise ValueError(f"reference {i} source not found: {source}")
kwargs: dict[str, Any] = {
"source": source,
"media_type": (ref.get("media_type") or "image"),
}
for opt in ("soundtrack", "fps", "sample_rate"):
if ref.get(opt) not in (None, ""):
kwargs[opt] = ref[opt]
built.append(MiniMaxH3Reference(**kwargs))
return built
JOB_LOG_FILENAME = "out.log"
def _job_log_path(output_dir: str, job_id: str) -> str:
"""Each job's log lives beside its outputs: <output_dir>/<job_id>/out.log."""
return os.path.join(output_dir, job_id, JOB_LOG_FILENAME)
def _decode_references(value: Any) -> list[dict[str, Any]]:
"""Reference lists round-trip through the DB as JSON text."""
if not value:
return []
if isinstance(value, list):
return list(value)
try:
decoded = json.loads(value)
except (TypeError, ValueError):
logger.warning("Could not decode stored references: %r", value)
return []
return list(decoded) if isinstance(decoded, list) else []
def _generator_is_alive(generator: Any) -> bool:
"""True if the generator's worker processes are all still running.
A cached VideoGenerator holds a MultiprocExecutor whose workers are separate
processes; nothing notices when they exit. Probing `proc.is_alive()` is what
the executor itself uses during shutdown. Anything unexpected in the object
graph is treated as alive so a probe failure can never wedge the cache.
"""
executor = getattr(generator, "executor", None)
workers = getattr(executor, "workers", None)
if not workers:
return True
try:
return all(w.proc.is_alive() for w in workers)
except Exception:
logger.debug("Worker liveness probe failed", exc_info=True)
return True
class JobRunner:
"""Manages video generation jobs, their execution, and generator caching."""
@@ -276,7 +355,7 @@ class JobRunner:
"""Populate job's log buffer from its log file if it exists."""
path = job.log_file_path
if not path:
path = os.path.join(self.log_dir, f"{job.id}.log")
path = _job_log_path(self.output_dir, job.id)
if not os.path.isfile(path):
return
try:
@@ -314,10 +393,13 @@ class JobRunner:
job = Job(
id=row["id"],
model_id=row["model_id"],
name=row.get("name", "") or "",
prompt=row["prompt"],
workload_type=row.get("workload_type", "t2v"),
job_type=row.get("job_type", "inference"),
image_path=row.get("image_path", "") or "",
last_image_path=row.get("last_image_path", "") or "",
references=_decode_references(row.get("references")),
data_path=row.get("data_path", "") or "",
max_train_steps=row.get("max_train_steps", 1000),
train_batch_size=row.get("train_batch_size", 1),
@@ -350,6 +432,7 @@ class JobRunner:
negative_prompt=row.get("negative_prompt", "") or "",
num_gpus=row.get("num_gpus", 1),
dit_cpu_offload=row.get("dit_cpu_offload", False),
dit_layerwise_offload=row.get("dit_layerwise_offload", False),
text_encoder_cpu_offload=row.get("text_encoder_cpu_offload", False),
vae_cpu_offload=row.get("vae_cpu_offload", False),
image_encoder_cpu_offload=row.get("image_encoder_cpu_offload", False),
@@ -394,9 +477,12 @@ class JobRunner:
job_id: str,
model_id: str,
prompt: str,
name: str = "",
workload_type: str = "t2v",
job_type: str = "inference",
image_path: str = "",
last_image_path: str = "",
references: list[dict[str, Any]] | None = None,
data_path: str = "",
max_train_steps: int = 1000,
train_batch_size: int = 1,
@@ -422,6 +508,7 @@ class JobRunner:
num_gpus: int = 1,
negative_prompt: str = "",
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -435,10 +522,13 @@ class JobRunner:
job = Job(
id=job_id,
model_id=model_id,
name=(name or "").strip(),
prompt=prompt.strip(),
workload_type=workload_type or "t2v",
job_type=job_type or "inference",
image_path=image_path or "",
last_image_path=last_image_path or "",
references=list(references or []),
data_path=data_path or "",
max_train_steps=max_train_steps,
train_batch_size=train_batch_size,
@@ -464,6 +554,7 @@ class JobRunner:
negative_prompt=negative_prompt or "",
num_gpus=num_gpus,
dit_cpu_offload=dit_cpu_offload,
dit_layerwise_offload=dit_layerwise_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
image_encoder_cpu_offload=image_encoder_cpu_offload,
@@ -521,6 +612,84 @@ class JobRunner:
logger.info("Deleted job %s", job.id)
return True
CONFIG_FIELDS: tuple[str, ...] = (
"model_id",
"name",
"prompt",
"workload_type",
"job_type",
"image_path",
"last_image_path",
"references",
"negative_prompt",
"num_inference_steps",
"num_frames",
"height",
"width",
"guidance_scale",
"guidance_rescale",
"fps",
"seed",
"num_gpus",
"dit_cpu_offload",
"dit_layerwise_offload",
"text_encoder_cpu_offload",
"vae_cpu_offload",
"image_encoder_cpu_offload",
"use_fsdp_inference",
"enable_torch_compile",
"vsa_sparsity",
"tp_size",
"sp_size",
"data_path",
"max_train_steps",
"train_batch_size",
"learning_rate",
"num_latent_t",
"validation_dataset_file",
"lora_rank",
"dmd_use_vsa",
"dmd_vsa_sparsity",
"dmd_denoising_steps",
"real_score_guidance_scale",
"generator_update_interval",
"real_score_model_path",
"fake_score_model_path",
)
def duplicate_job(self, job_id: str, new_job_id: str) -> Job:
"""Create a new pending job with an existing job's configuration.
Runtime state (status, timings, logs, outputs) is not carried over.
"""
with self._jobs_lock:
source = self._jobs.get(job_id)
if source is None:
raise ValueError(f"Job {job_id} not found")
config = {f: copy.deepcopy(getattr(source, f)) for f in self.CONFIG_FIELDS}
return self.create_job(job_id=new_job_id, **config)
#: Editable exactly when startable: the same set start_job() accepts.
EDITABLE_STATUSES = (JobStatus.PENDING, JobStatus.FAILED, JobStatus.STOPPED)
def update_job_config(self, job_id: str, updates: dict[str, Any]) -> Job:
"""Edit the configuration of a job that has not produced a result."""
with self._jobs_lock:
job = self._jobs.get(job_id)
if job is None:
raise ValueError(f"Job {job_id} not found")
if job.status not in self.EDITABLE_STATUSES:
allowed = ", ".join(s.value for s in self.EDITABLE_STATUSES)
raise ValueError(f"Job is {job.status.value}; only {allowed} jobs can be edited. "
"Duplicate it instead.")
unknown = set(updates) - set(self.CONFIG_FIELDS)
if unknown:
raise ValueError(f"Not editable: {', '.join(sorted(unknown))}")
for field_name, value in updates.items():
setattr(job, field_name, value)
self._save_job(job)
return job
def start_job(self, job_id: str) -> Job:
"""Start (or restart) a pending / stopped / failed job.
@@ -623,6 +792,8 @@ class JobRunner:
workload_type: str,
num_gpus: int,
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
override_pipeline_cls_name: str | None = None,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -638,6 +809,10 @@ class JobRunner:
workload_type,
num_gpus,
dit_cpu_offload,
dit_layerwise_offload,
# Ref2VA loads different DiT weights (transformer_ref), so the
# override must key the cache or a t2v/i2v generator gets reused.
override_pipeline_cls_name,
text_encoder_cpu_offload,
vae_cpu_offload,
image_encoder_cpu_offload,
@@ -650,8 +825,21 @@ class JobRunner:
# Generators are cached by model_id and configuration parameters
with self._generators_lock:
if cache_key in self._generators:
return self._generators[cache_key]
cached = self._generators.get(cache_key)
if cached is not None:
if _generator_is_alive(cached):
return cached
# Workers can exit while a generator sits idle in the cache;
# reusing it fails every later job with the same config.
logger.warning(
"Cached generator for %s has dead workers; reloading.",
model_id,
)
self._generators.pop(cache_key, None)
try:
cached.shutdown()
except Exception:
logger.debug("Shutdown of the dead generator failed", exc_info=True)
# Import lazily so starting the server is fast even without a GPU.
from fastvideo import VideoGenerator
@@ -677,6 +865,11 @@ class JobRunner:
gen = VideoGenerator.from_pretrained(
model_id,
workload_type=workload_type,
num_gpus=num_gpus,
dit_layerwise_offload=dit_layerwise_offload,
**({
"override_pipeline_cls_name": override_pipeline_cls_name
} if override_pipeline_cls_name else {}),
dit_cpu_offload=dit_cpu_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
@@ -706,10 +899,9 @@ class JobRunner:
def _run_training_job(self, job: Job):
"""Run a finetuning, distillation, or LoRA job via subprocess."""
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
job_output_dir = os.path.join(self.output_dir, job.id)
os.makedirs(job_output_dir, exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
if not job.data_path or not os.path.isdir(job.data_path):
job.status = JobStatus.FAILED
@@ -827,8 +1019,8 @@ class JobRunner:
def _run_inference_job(self, job: Job):
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
os.makedirs(os.path.join(self.output_dir, job.id), exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
# Add file handler to persist logs
file_handler = logging.FileHandler(job.log_file_path, mode='w', encoding='utf-8')
@@ -875,62 +1067,45 @@ class JobRunner:
buf.phase = "loading model"
logger.info("Loading model...")
# Run generator creation in a background thread so we
# can poll _stop_event while the (potentially slow)
# model download / load is in progress.
_gen_result: list[Any] = []
_gen_error: list[BaseException] = []
# The generator MUST be created on this thread: building it spawns
# the executor's worker processes, and they are torn down if the
# creating thread exits. Running it in a helper thread (to poll
# _stop_event during load) made every collective_rpc fail with
# ConnectionResetError.
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning("Job %s stopped before model loading", job.id)
buf.phase = "stopped"
return
def _load_generator() -> None:
try:
gen = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
_gen_result.append(gen)
except BaseException as exc:
_gen_error.append(exc)
loader = threading.Thread(
target=_load_generator,
daemon=True,
generator = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
dit_layerwise_offload=job.dit_layerwise_offload,
override_pipeline_cls_name=(MINIMAX_H3_REF2VA_PIPELINE if job.references else None),
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
loader.start()
while loader.is_alive():
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning(
"Job %s stopped during model loading",
job.id,
)
buf.phase = "stopped"
return
loader.join(timeout=0.5)
if _gen_error:
raise _gen_error[0]
generator = _gen_result[0]
buf.phase = "generating"
logger.info("Starting generation for job %s (model=%s)", job.id, job.model_id)
# Without a name FastVideo derives the filename from the prompt.
safe_name = re.sub(r'[\\/:*?"<>|]+', "", job.name).strip().strip(".")
output_target = (os.path.join(job_output_dir, f"{safe_name[:80]}.mp4") if safe_name else job_output_dir)
gen_kwargs: dict[str, Any] = {
"prompt": job.prompt,
"output_path": job_output_dir,
"output_path": output_target,
"save_video": True,
"num_inference_steps": job.num_inference_steps,
"num_frames": job.num_frames,
@@ -945,6 +1120,12 @@ class JobRunner:
}
if job.image_path:
gen_kwargs["image_path"] = job.image_path
if job.references:
gen_kwargs["references"] = _build_h3_references(job.references)
if job.last_image_path:
# _prepare_fl2va requires a PIL image, not a path.
from PIL import Image as _PILImage
gen_kwargs["last_image"] = _PILImage.open(job.last_image_path)
generator.generate_video(**gen_kwargs)
buf.phase = "saving"
@@ -977,7 +1158,7 @@ class JobRunner:
except Exception as exception:
error_msg = str(exception)
logger.error("Critical error in job thread: %s", error_msg)
logger.exception("Critical error in job thread: %s", error_msg)
job.status = JobStatus.FAILED
job.error = f"Critical error ({type(exception).__name__}): {error_msg}"
job.finished_at = time.time()
@@ -1,15 +1,20 @@
# SPDX-License-Identifier: Apache-2.0
"""Request model for creating a job."""
from typing import Any
from pydantic import BaseModel
class CreateJobRequest(BaseModel):
model_id: str
name: str = ""
prompt: str
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] | None = None
data_path: str = ""
max_train_steps: int = 1000
train_batch_size: int = 1
@@ -28,6 +33,7 @@ class CreateJobRequest(BaseModel):
seed: int = 1024
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
+92 -2
View File
@@ -18,6 +18,7 @@ import argparse
import contextlib
import logging
import os
import re
import shutil
import signal
import time
@@ -116,7 +117,30 @@ def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _available_models
def _safe_upload_name(filename: str | None, ext: str) -> str:
"""A filesystem-safe version of the client's filename, keeping it readable.
Uploads live under a per-file uuid directory, so the basename does not have
to be unique -- only safe. Keeping the original name means the path stays
self-describing wherever it travels: the database, job logs, and payloads
copied back out to the API.
"""
stem = os.path.basename(filename or "").rsplit(".", 1)[0]
stem = re.sub(r"[^A-Za-z0-9._-]+", "_", stem).strip("._-")
return f"{stem[:80] or 'upload'}{ext}"
def _upload_destination(ext: str, filename: str | None) -> str:
"""<upload_dir>/<uuid4>/<safe original name><ext>"""
directory = os.path.join(upload_dir, uuid.uuid4().hex)
os.makedirs(directory, exist_ok=True)
return os.path.join(directory, _safe_upload_name(filename, ext))
ALLOWED_IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".webp", ".bmp"}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".mov", ".mkv", ".webm", ".avi"}
ALLOWED_AUDIO_EXTENSIONS = {".wav", ".mp3", ".flac", ".m4a", ".ogg"}
ALLOWED_MEDIA_EXTENSIONS = (ALLOWED_IMAGE_EXTENSIONS | ALLOWED_VIDEO_EXTENSIONS | ALLOWED_AUDIO_EXTENSIONS)
@app.post("/api/upload-image")
@@ -136,8 +160,7 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
f"{', '.join(ALLOWED_IMAGE_EXTENSIONS)}"),
)
os.makedirs(upload_dir, exist_ok=True)
unique_name = f"{uuid.uuid4().hex}{ext}"
dest_path = os.path.join(upload_dir, unique_name)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
@@ -150,6 +173,47 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
return {"path": os.path.abspath(dest_path)}
@app.post("/api/upload-media")
async def upload_media(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
"""Upload an image, video or audio file for Ref2VA references.
Returns the absolute path plus the media_type MiniMax-H3 expects, so the
caller does not have to re-derive it from the extension.
"""
global upload_dir # noqa: PLW0603
if not upload_dir:
raise HTTPException(
status_code=503,
detail="Upload directory not configured",
)
ext = Path(file.filename or "").suffix.lower()
if ext not in ALLOWED_MEDIA_EXTENSIONS:
raise HTTPException(
status_code=400,
detail=(f"Invalid file type. Allowed: "
f"{', '.join(sorted(ALLOWED_MEDIA_EXTENSIONS))}"),
)
if ext in ALLOWED_VIDEO_EXTENSIONS:
media_type = "video"
elif ext in ALLOWED_AUDIO_EXTENSIONS:
media_type = "audio"
else:
media_type = "image"
os.makedirs(upload_dir, exist_ok=True)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
f.write(contents)
except OSError as e:
raise HTTPException(
status_code=500,
detail=f"Failed to save upload: {e}",
) from e
return {"path": os.path.abspath(dest_path), "media_type": media_type}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
@@ -282,10 +346,13 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
job = job_runner.create_job(
job_id=str(uuid.uuid4()),
model_id=req.model_id,
name=req.name or "",
prompt=req.prompt,
workload_type=req.workload_type or "t2v",
job_type=job_type,
image_path=req.image_path or "",
last_image_path=req.last_image_path or "",
references=req.references or [],
data_path=data_path,
max_train_steps=req.max_train_steps,
train_batch_size=req.train_batch_size,
@@ -304,6 +371,7 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
seed=req.seed,
num_gpus=req.num_gpus,
dit_cpu_offload=req.dit_cpu_offload,
dit_layerwise_offload=req.dit_layerwise_offload,
text_encoder_cpu_offload=req.text_encoder_cpu_offload,
vae_cpu_offload=req.vae_cpu_offload,
image_encoder_cpu_offload=req.image_encoder_cpu_offload,
@@ -336,6 +404,28 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
return job.to_dict()
@app.post("/api/jobs/{job_id}/duplicate", status_code=201)
def duplicate_job(job_id: str) -> dict[str, Any]:
"""Create a new pending job with the same configuration as an existing one."""
try:
job = job_runner.duplicate_job(job_id, str(uuid.uuid4()))
except ValueError as e:
raise HTTPException(status_code=404, detail=str(e)) from e
return job.to_dict()
@app.patch("/api/jobs/{job_id}")
def update_job(job_id: str, updates: dict[str, Any]) -> dict[str, Any]:
"""Edit a pending job's configuration. Started jobs cannot be edited."""
try:
job = job_runner.update_job_config(job_id, updates)
except ValueError as e:
detail = str(e)
status = 404 if "not found" in detail else 400
raise HTTPException(status_code=status, detail=detail) from e
return job.to_dict()
@app.post("/api/jobs/{job_id}/start")
def start_job(job_id: str) -> dict[str, Any]:
"""Start (or restart) a pending / stopped / failed job."""
@@ -22,15 +22,35 @@ import { useStore } from '@/hooks/useStore';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import {
createJob,
updateJob,
getDatasets,
getModels,
uploadImage,
uploadMedia,
type CreateJobRequest,
type Model,
} from '@/lib/api';
import { getDefaultModelForWorkload } from '@/lib/defaultOptions';
import { WORKLOAD_OPTIONS } from '@/lib/jobConfig';
import type { JobType } from '@/lib/types';
import {
H3_MAX_REFERENCES,
labelReferences,
referencePromptSeed,
validateReferences,
type H3Reference,
} from '@/lib/h3References';
import {
EMPTY_H3_PROMPT_FIELDS,
H3_PROMPT_SECTIONS,
H3_SECTION_HINTS,
H3_SECTION_LABELS,
isEmptyPromptFields,
parseH3Prompt,
serializeH3Prompt,
type H3PromptFields,
} from '@/lib/h3Prompt';
import { jobToFormFields, type JobLike } from '@/lib/jobToFields';
export interface CreateJobModalProps {
isOpen: boolean;
@@ -38,6 +58,10 @@ export interface CreateJobModalProps {
onSuccess: () => void;
jobType: JobType;
workloadType: string;
/** When set, the modal edits this pending job instead of creating a new one. */
editingJob?: JobLike | null;
/** Show the configuration without allowing changes (started/finished jobs). */
readOnly?: boolean;
}
export default function CreateJobModal({
@@ -46,6 +70,8 @@ export default function CreateJobModal({
onSuccess,
jobType,
workloadType,
editingJob,
readOnly = false,
}: CreateJobModalProps) {
const { options } = useStore(defaultOptionsStore);
@@ -55,8 +81,22 @@ export default function CreateJobModal({
const [models, setModels] = React.useState<Model[]>([]);
const [modelId, setModelId] = React.useState('');
const [name, setName] = React.useState('');
const [prompt, setPrompt] = React.useState('');
const [imagePath, setImagePath] = React.useState('');
const [lastImagePath, setLastImagePath] = React.useState('');
const [references, setReferences] = React.useState<H3Reference[]>([]);
const [isUploadingReference, setIsUploadingReference] = React.useState(false);
const [referenceError, setReferenceError] = React.useState<string | null>(null);
const [promptFields, setPromptFields] = React.useState<H3PromptFields>(
EMPTY_H3_PROMPT_FIELDS,
);
const [useGuidedPrompt, setUseGuidedPrompt] = React.useState(true);
const [lastImageFileName, setLastImageFileName] = React.useState('');
const [isUploadingLastImage, setIsUploadingLastImage] = React.useState(false);
const [lastImageUploadError, setLastImageUploadError] = React.useState<
string | null
>(null);
const [imageFileName, setImageFileName] = React.useState('');
const [isUploadingImage, setIsUploadingImage] = React.useState(false);
const [negativePrompt, setNegativePrompt] = React.useState('');
@@ -70,12 +110,50 @@ export default function CreateJobModal({
const [seed, setSeed] = React.useState(1024);
const [numGpus, setNumGpus] = React.useState(1);
const [ditCpuOffload, setDitCpuOffload] = React.useState(false);
const [ditLayerwiseOffload, setDitLayerwiseOffload] = React.useState(false);
const [textEncoderCpuOffload, setTextEncoderCpuOffload] =
React.useState(false);
const [vaeCpuOffload, setVaeCpuOffload] = React.useState(false);
const [imageEncoderCpuOffload, setImageEncoderCpuOffload] =
React.useState(false);
const [useFsdpInference, setUseFsdpInference] = React.useState(false);
// H3 is the only registered model with an end frame or references.
const supportsLastImage = modelId.toLowerCase().includes('minimax-h3');
const usingReferences = supportsLastImage && references.length > 0;
// JobCard re-renders on every job-list poll, so `editingJob` is a fresh
// object each time. Effects must depend on these, never on the object.
const editingJobId = editingJob?.id ?? null;
const editingJobModelId = editingJob?.model_id ?? null;
// Layerwise offload and FSDP compete for the DiT weights and FastVideoArgs
// silently picks a winner (fastvideo_args.py:859); resolve it visibly here.
// dit_cpu_offload is deliberately not interlocked -- it is a modifier, not a
// competing strategy.
const handleDitLayerwiseOffloadChange = React.useCallback((next: boolean) => {
setDitLayerwiseOffload(next);
if (next) {
setUseFsdpInference(false);
}
}, []);
const handleUseFsdpInferenceChange = React.useCallback((next: boolean) => {
setUseFsdpInference(next);
if (next) {
setDitLayerwiseOffload(false);
}
}, []);
const handleNumGpusChange = React.useCallback((next: number) => {
setNumGpus(next);
if (next > 1) {
// Dropping back to one GPU leaves FSDP alone: single-GPU FSDP is a
// valid way to reach its CPU offload (docs/inference/offloading.md).
setUseFsdpInference(true);
setDitLayerwiseOffload(false);
}
}, []);
const [enableTorchCompile, setEnableTorchCompile] = React.useState(false);
const [vsaSparsity, setVsaSparsity] = React.useState(0);
const [tpSize, setTpSize] = React.useState(-1);
@@ -126,6 +204,47 @@ export default function CreateJobModal({
const justOpened = isOpen && !justOpenedRef.current;
justOpenedRef.current = isOpen;
if (!justOpened) return;
if (editingJob) {
// Must not fall through to the defaults below: a partially-seeded
// form silently edits values the user never saw.
const f = jobToFormFields(editingJob);
setModelId(f.modelId);
setName(f.name);
setPrompt(f.prompt);
setNegativePrompt(f.negativePrompt);
setImagePath(f.imagePath);
setImageFileName(f.imagePath.split('/').pop() ?? '');
setLastImagePath(f.lastImagePath);
setLastImageFileName(f.lastImagePath.split('/').pop() ?? '');
setReferences(f.references);
setPromptFields(f.promptFields ?? EMPTY_H3_PROMPT_FIELDS);
setUseGuidedPrompt(f.promptFields !== null);
setNumInferenceSteps(f.numInferenceSteps);
setNumFrames(f.numFrames);
setHeight(f.height);
setWidth(f.width);
setGuidanceScale(f.guidanceScale);
setGuidanceRescale(f.guidanceRescale);
setFps(f.fps);
setSeed(f.seed);
setNumGpus(f.numGpus);
setDitCpuOffload(f.ditCpuOffload);
setDitLayerwiseOffload(f.ditLayerwiseOffload);
setTextEncoderCpuOffload(f.textEncoderCpuOffload);
setVaeCpuOffload(f.vaeCpuOffload);
setImageEncoderCpuOffload(f.imageEncoderCpuOffload);
setUseFsdpInference(f.useFsdpInference);
setEnableTorchCompile(f.enableTorchCompile);
setVsaSparsity(f.vsaSparsity);
setTpSize(f.tpSize);
setSpSize(f.spSize);
setReferenceError(null);
setModelLoadError(null);
setImageUploadError(null);
setLastImageUploadError(null);
setSubmitError(null);
return;
}
const opts = options;
setNumInferenceSteps(opts.numInferenceSteps);
setNumFrames(workloadType === 't2i' ? 1 : opts.numFrames);
@@ -137,6 +256,7 @@ export default function CreateJobModal({
setSeed(opts.seed);
setNumGpus(opts.numGpus);
setDitCpuOffload(opts.ditCpuOffload);
setDitLayerwiseOffload(opts.ditLayerwiseOffload ?? false);
setTextEncoderCpuOffload(opts.textEncoderCpuOffload);
setVaeCpuOffload(opts.vaeCpuOffload);
setImageEncoderCpuOffload(opts.imageEncoderCpuOffload);
@@ -151,8 +271,16 @@ export default function CreateJobModal({
inferenceWorkload as 't2v' | 'i2v' | 't2i',
),
);
setName('');
setImagePath('');
setImageFileName('');
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
setReferences([]);
setReferenceError(null);
setPromptFields(EMPTY_H3_PROMPT_FIELDS);
setUseGuidedPrompt(true);
setSelectedDatasetId('');
setSelectedValidationDatasetId('');
setModelLoadError(null);
@@ -168,7 +296,7 @@ export default function CreateJobModal({
setRealScoreModelPath('');
setFakeScoreModelPath('');
}
}, [isOpen, workloadType, inferenceWorkload, options]);
}, [isOpen, workloadType, inferenceWorkload, options, editingJobId]);
// Load the models available for this workload.
React.useEffect(() => {
@@ -188,7 +316,16 @@ export default function CreateJobModal({
opts,
inferenceWorkload as 't2v' | 'i2v' | 't2i',
);
const chosen = ids.includes(defaultId) ? defaultId : (list[0]?.id ?? '');
// When editing, the job's own model wins over the workload default --
// this resolves after the seeding effect, so choosing a default here
// would silently swap the model out from under the user.
const editedId = editingJobModelId;
const chosen =
editedId && ids.includes(editedId)
? editedId
: ids.includes(defaultId)
? defaultId
: (list[0]?.id ?? '');
setModelId(chosen);
if (workloadType === 'dmd_t2v') {
setRealScoreModelPath(chosen);
@@ -210,7 +347,7 @@ export default function CreateJobModal({
return () => {
stale = true;
};
}, [isOpen, inferenceWorkload, workloadType]);
}, [isOpen, inferenceWorkload, workloadType, editingJobModelId]);
// Training jobs need a dataset; load the ready datasets when relevant.
React.useEffect(() => {
@@ -262,6 +399,106 @@ export default function CreateJobModal({
}
}
async function handleLastImageChange(
e: React.ChangeEvent<HTMLInputElement>,
) {
const file = e.target.files?.[0];
if (!file) {
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
return;
}
setIsUploadingLastImage(true);
setLastImageFileName(file.name);
setLastImageUploadError(null);
try {
const { path } = await uploadImage(file);
setLastImagePath(path);
} catch (error) {
console.error('Failed to upload end image:', error);
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(
error instanceof Error
? `${error.message}. Choose the image again to retry.`
: 'The image could not be uploaded. Choose it again to retry.',
);
} finally {
setIsUploadingLastImage(false);
}
}
async function handleAddReference(
e: React.ChangeEvent<HTMLInputElement>,
) {
const file = e.target.files?.[0];
e.target.value = ''; // allow re-picking the same file
if (!file) return;
setIsUploadingReference(true);
setReferenceError(null);
try {
const { path, media_type } = await uploadMedia(file);
const next: H3Reference[] = [
...references,
{
id: `${Date.now()}-${file.name}`,
source: path,
media_type,
fileName: file.name,
},
];
setReferences(next);
setReferenceError(validateReferences(next));
} catch (error) {
console.error('Failed to upload reference:', error);
setReferenceError(
error instanceof Error ? error.message : 'The file could not be uploaded.',
);
} finally {
setIsUploadingReference(false);
}
}
function removeReference(id: string) {
const next = references.filter((r) => r.id !== id);
setReferences(next);
setReferenceError(validateReferences(next));
}
function seedPromptFields() {
setPromptFields({
...EMPTY_H3_PROMPT_FIELDS,
...referencePromptSeed(references),
});
setUseGuidedPrompt(true);
}
function setPromptField(section: string, value: string) {
setPromptFields((prev) => ({ ...prev, [section]: value }));
}
// Switching between the guided fields and the raw editor keeps whatever was
// typed: serialize on the way out, parse back on the way in.
function toggleGuidedPrompt() {
if (useGuidedPrompt) {
if (!isEmptyPromptFields(promptFields)) {
setPrompt(serializeH3Prompt(promptFields));
}
setUseGuidedPrompt(false);
} else {
const parsed = parseH3Prompt(prompt);
if (parsed) setPromptFields(parsed);
setUseGuidedPrompt(true);
}
}
function clearLastImage() {
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
}
function clearImage() {
setImagePath('');
setImageFileName('');
@@ -271,7 +508,16 @@ export default function CreateJobModal({
async function handleSubmit(e: React.FormEvent<HTMLFormElement>) {
e.preventDefault();
if (isInference && workloadType === 'i2v' && !imagePath) return;
if (isInference && workloadType === 'i2v' && !imagePath && !usingReferences)
return;
if (usingReferences && validateReferences(references)) return;
if (
usingReferences &&
useGuidedPrompt &&
isEmptyPromptFields(promptFields) &&
!prompt.trim()
)
return;
// Send the dataset id; the backend resolves it to the on-disk media dir.
const effectiveDataPath = selectedDatasetId ?? '';
if (!isInference && !selectedDatasetId) return;
@@ -285,14 +531,35 @@ export default function CreateJobModal({
try {
const payload: CreateJobRequest = {
model_id: modelId,
prompt,
name: name.trim(),
prompt:
usingReferences && useGuidedPrompt && !isEmptyPromptFields(promptFields)
? serializeH3Prompt(promptFields)
: prompt,
workload_type: workloadType,
job_type: effectiveJobType,
...(isInference
? {
...(workloadType === 'i2v' && imagePath
// Ref2VA and the FL2VA keyframes are mutually exclusive:
// _prepare_ref2va rejects image_path/last_image_path outright
// when references are present.
...(workloadType === 'i2v' && !usingReferences && imagePath
? { image_path: imagePath }
: {}),
...(workloadType === 'i2v' &&
supportsLastImage &&
!usingReferences &&
lastImagePath
? { last_image_path: lastImagePath }
: {}),
...(workloadType === 'i2v' && supportsLastImage && references.length
? {
references: references.map((r) => ({
source: r.source,
media_type: r.media_type,
})),
}
: {}),
negative_prompt: negativePrompt,
num_inference_steps: numInferenceSteps,
num_frames: numFrames,
@@ -304,6 +571,7 @@ export default function CreateJobModal({
seed,
num_gpus: numGpus,
dit_cpu_offload: ditCpuOffload,
dit_layerwise_offload: ditLayerwiseOffload,
text_encoder_cpu_offload: textEncoderCpuOffload,
vae_cpu_offload: vaeCpuOffload,
image_encoder_cpu_offload: imageEncoderCpuOffload,
@@ -334,7 +602,14 @@ export default function CreateJobModal({
: {}),
}),
};
await createJob(payload);
if (editingJob) {
await updateJob(
editingJob.id,
payload as unknown as Record<string, unknown>,
);
} else {
await createJob(payload);
}
onSuccess();
onClose();
} catch (err) {
@@ -356,9 +631,9 @@ export default function CreateJobModal({
const workloadLabel =
WORKLOAD_OPTIONS[jobType]?.find((o) => o.type === workloadType)?.label ?? '';
const title = `New ${jobType.charAt(0).toUpperCase() + jobType.slice(1)} Job${
workloadLabel ? ` (${workloadLabel})` : ''
}`;
const title = `${readOnly ? 'View' : editingJob ? 'Edit' : 'New'} ${
jobType.charAt(0).toUpperCase() + jobType.slice(1)
} Job${workloadLabel ? ` (${workloadLabel})` : ''}`;
return (
<Dialog
@@ -385,6 +660,23 @@ export default function CreateJobModal({
autoComplete="off"
className="flex flex-col gap-3.5"
>
{/* disabled cascades to every control inside; display:contents
keeps the parent's flex layout. */}
<fieldset
disabled={readOnly}
style={{ display: 'contents' }}
className="contents"
>
<FieldRow htmlFor="modal-name" label="Name (optional)">
<Input
id="modal-name"
value={name}
onChange={(e) => setName(e.target.value)}
placeholder="Shown on the job card and used for the output filename"
disabled={isSubmitting}
/>
</FieldRow>
<FieldRow htmlFor="modal-modelId" label="Model">
<NativeSelect
id="modal-modelId"
@@ -429,12 +721,12 @@ export default function CreateJobModal({
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp"
onChange={handleImageChange}
disabled={isSubmitting || isUploadingImage}
disabled={isSubmitting || isUploadingImage || usingReferences}
aria-describedby={
imageUploadError ? 'modal-image-error' : undefined
}
aria-invalid={imageUploadError ? true : undefined}
required
required={!usingReferences}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{imageFileName && (
@@ -462,24 +754,169 @@ export default function CreateJobModal({
</FieldRow>
)}
<FieldRow
htmlFor="modal-prompt"
label={isInference ? 'Prompt' : 'Description'}
>
<Textarea
id="modal-prompt"
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
rows={isInference ? 3 : 2}
placeholder={
isInference
? 'A curious raccoon peers through a vibrant field of yellow sunflowers…'
: 'Brief description of this training job…'
}
required
disabled={isSubmitting}
/>
</FieldRow>
{isInference && workloadType === 'i2v' && supportsLastImage && (
<FieldRow htmlFor="modal-last-image" label="End Frame (optional)">
<Input
id="modal-last-image"
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp"
onChange={handleLastImageChange}
disabled={isSubmitting || isUploadingLastImage}
aria-describedby={
lastImageUploadError ? 'modal-last-image-error' : undefined
}
aria-invalid={lastImageUploadError ? true : undefined}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{lastImageFileName && (
<span className="mt-0.5 text-xs text-muted-foreground">
{isUploadingLastImage ? 'Uploading…' : lastImageFileName} ·{' '}
<button
type="button"
onClick={clearLastImage}
disabled={isSubmitting || isUploadingLastImage}
className="text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Clear
</button>
</span>
)}
{lastImageUploadError && (
<p
id="modal-last-image-error"
role="alert"
className="text-sm text-destructive"
>
{lastImageUploadError}
</p>
)}
</FieldRow>
)}
{isInference && workloadType === 'i2v' && supportsLastImage && (
<FieldRow htmlFor="modal-reference" label="References (Ref2VA)">
<Input
id="modal-reference"
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp,.mp4,.mov,.mkv,.webm,.avi,.wav,.mp3,.flac,.m4a,.ogg"
onChange={handleAddReference}
disabled={
isSubmitting ||
isUploadingReference ||
references.length >= H3_MAX_REFERENCES
}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{isUploadingReference && (
<span className="mt-0.5 text-xs text-muted-foreground">
Uploading…
</span>
)}
{references.length > 0 && (
<ul className="mt-1 flex list-none flex-col gap-1 p-0">
{references.map((reference, index) => (
<li
key={reference.id}
className="flex items-center gap-2 text-xs text-muted-foreground"
>
<code className="font-mono text-accent-blue">
{labelReferences(references)[index]}
</code>
<span className="truncate">{reference.fileName}</span>
<button
type="button"
onClick={() => removeReference(reference.id)}
disabled={isSubmitting}
className="ml-auto text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Remove
</button>
</li>
))}
</ul>
)}
{references.length > 0 && (
<button
type="button"
onClick={seedPromptFields}
disabled={isSubmitting}
className="mt-1 self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Fill prompt sections from references
</button>
)}
{referenceError && (
<p role="alert" className="mt-0.5 text-xs text-destructive">
{referenceError}
</p>
)}
<span className="mt-0.5 text-xs text-muted-foreground">
Ref2VA replaces the keyframes: up to 9 images, 3 videos, 3 audio
(12 total). Audio needs at least one image or video.
</span>
</FieldRow>
)}
{usingReferences && useGuidedPrompt ? (
/* Six-section format from the model's reference prompt guide. */
<>
{H3_PROMPT_SECTIONS.map((section) => (
<FieldRow
key={section}
htmlFor={`modal-prompt-${section}`}
label={H3_SECTION_LABELS[section]}
>
<Textarea
id={`modal-prompt-${section}`}
value={promptFields[section]}
onChange={(e) => setPromptField(section, e.target.value)}
rows={section === 'detailed_description' ? 5 : 2}
placeholder={H3_SECTION_HINTS[section]}
disabled={isSubmitting}
/>
</FieldRow>
))}
<button
type="button"
onClick={toggleGuidedPrompt}
disabled={isSubmitting}
className="self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Edit as raw prompt
</button>
</>
) : (
<>
<FieldRow
htmlFor="modal-prompt"
label={isInference ? 'Prompt' : 'Description'}
>
<Textarea
id="modal-prompt"
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
rows={isInference ? 3 : 2}
placeholder={
isInference
? 'A curious raccoon peers through a vibrant field of yellow sunflowers…'
: 'Brief description of this training job…'
}
required={!(usingReferences && useGuidedPrompt)}
disabled={isSubmitting}
/>
</FieldRow>
{usingReferences && (
<button
type="button"
onClick={toggleGuidedPrompt}
disabled={isSubmitting}
className="self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Edit as prompt sections
</button>
)}
</>
)}
{isInference && (
<FieldRow htmlFor="modal-negative-prompt" label="Negative Prompt">
@@ -849,6 +1286,13 @@ export default function CreateJobModal({
onChange={setDitCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-dit-layerwise-offload"
label="DiT Layerwise Offload"
checked={ditLayerwiseOffload}
onChange={handleDitLayerwiseOffloadChange}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-text-encoder-cpu-offload"
label="Text Encoder CPU Offload"
@@ -860,7 +1304,7 @@ export default function CreateJobModal({
id="modal-use-fsdp-inference"
label="Use FSDP Inference"
checked={useFsdpInference}
onChange={setUseFsdpInference}
onChange={handleUseFsdpInferenceChange}
disabled={isSubmitting}
/>
<ToggleRow
@@ -891,7 +1335,7 @@ export default function CreateJobModal({
max={8}
step={1}
value={numGpus}
onChange={setNumGpus}
onChange={handleNumGpusChange}
disabled={isSubmitting}
/>
<NumberRow
@@ -906,6 +1350,8 @@ export default function CreateJobModal({
</details>
)}
</fieldset>
<div className="flex flex-col items-start gap-2">
{submitError && (
<p role="alert" className="text-sm text-destructive">
@@ -914,14 +1360,22 @@ export default function CreateJobModal({
)}
<Button
type="submit"
hidden={readOnly}
disabled={
readOnly ||
isSubmitting ||
isUploadingImage ||
!!modelLoadError ||
!!datasetLoadError
}
>
{isSubmitting ? 'Creating…' : 'Create Job'}
{isSubmitting
? editingJob
? 'Saving…'
: 'Creating…'
: editingJob
? 'Save Changes'
: 'Create Job'}
</Button>
</div>
</form>

Some files were not shown because too many files have changed in this diff Show More