Compare commits

..
Author SHA1 Message Date
0c16ec91b0 [feat]: connect Dreamverse creation settings to generation
Apply the backend-wiring changes beyond the UI uplift to
ds8/dreamversev2-dev for review and refactoring.

Source PR: hao-ai-lab/FastVideo#1854
Source range: 90d739a91892302edf37c4b23f807f402c83071d..8c5ee9b51cc3b75f4eba9cb3904fe08578c9dc9e
The 28-file patch is identical to that source range.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-15 17:16:15 -07:00
e57543b79d [feat]: Dreamverse creation studio UI uplift (#1853)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-15 16:45:08 -07:00
1170 changed files with 12536 additions and 105789 deletions
@@ -1,66 +0,0 @@
---
date: 2026-09-29
experiment: Wan-VACE multi-GPU inference on shared Slurm nodes
category: infrastructure
severity: important
---
# Spawn workers fail with ENOENT in `SemLock._rebuild`
## What Happened
On Slurm nodes, standard `VideoGenerator` runs with the `mp` executor sometimes
failed during worker startup. A spawn child raised `FileNotFoundError` from
`multiprocessing.synchronize.SemLock._rebuild` while unpickling its arguments,
even though the parent still held its Queue objects. Failures showed up on
several nodes and looked intermittent.
## Root Cause
The cluster's Slurm epilog `80-epilog-cleanup-shm-tmp.sh` runs
`find /dev/shm -maxdepth 2 -user "$SLURM_JOB_USER" -delete` at the end of
**every** job. When one of a user's jobs ends, the epilog deletes that user's
`/dev/shm` files on the node, including those that the user's *other* running
jobs are still using. Python's POSIX named semaphores (`/dev/shm/sem.mp-*`) are
among them. If the deletion lands between queue creation and a spawn child
unpickling it, the child fails with ENOENT.
Evidence:
- **Timing.** Every recorded deletion burst during a probe job matched, to the
second, the end of another job by the same user on the same node
(`sacct -u $USER -N <node>` end times).
- **Controlled reproduction.**
- Job A held a spawn `Lock` semaphore and a marker file in `/dev/shm`.
- Job B, on the same node, started and then ended.
- Job B's start deleted nothing.
- Within 1 s of job B's end, both of job A's objects were gone, and job A's
finalizer then hit the same ENOENT.
- **Still unexplained.** One historical failure has no same-user job
end in its window, so another deleter cannot be ruled out for it. logind
`RemoveIPC` was the earlier hypothesis; it is not needed to explain the other
cases.
## Fix / Workaround
- **FastVideo hardening.** Standard inference no longer creates the two
streaming queues it never uses. `FastVideoArgs.enable_streaming_ipc_queues`
defaults to `False`, and `StreamingVideoGenerator` sets it to `True`. This
removes the exposure for standard inference only.
- **Still exposed.** Streaming queues, NCCL shared-memory segments, and
DataLoader shared memory on the same node can still be deleted by the epilog.
- **Cluster fix (administrators).** Clean `/dev/shm` only when the user has no
other job on the node, or give each job a private `/dev/shm` with
`JobContainerType=job_container/tmpfs`.
- **User workaround.** Don't co-locate your own jobs on one node
(`--exclusive=user`), or avoid ending short jobs next to long-running ones.
## Prevention
- Before blaming an application for `/dev/shm` ENOENT on Slurm, read the
node-local prolog/epilog hooks (`/cm/local/apps/slurm/var/{prologs,epilogs}`).
Then correlate the failure window with `sacct -u $USER -N <node>` job end
times.
- Do not pass IPC primitives to spawn workers unless the worker needs them.
- `fastvideo/tests/worker/test_multiproc_executor.py` checks that standard
workers survive semaphore removal during spawn.
@@ -1,79 +0,0 @@
---
name: env-var-conventions
description: Add, read, rename, or remove an environment variable in FastVideo, or change the environment-variable policy. Use before touching fastvideo/envs.py, os.environ, os.getenv, or monkeypatch.setenv in fastvideo/, and when fastvideo/tests/contract/test_env_policy.py fails.
---
# Environment Variable Conventions
## Purpose
FastVideo registers its environment variables as typed fields in
`fastvideo/envs.py`. The policy that governs them is
`docs/contributing/env_vars.md`, and the contract test
`fastvideo/tests/contract/test_env_policy.py` enforces the policy in the unit
CI lane. This skill routes an environment-variable change through that policy.
The policy doc is the single source of the rules; read it instead of relying
on a summary here.
## Prerequisites
- Read `docs/contributing/env_vars.md` in full.
- Decide whether the setting belongs in an environment variable or an argument
(rule 5 in the policy doc). Settings that users change per deployment are
arguments; add them through `fastvideo/fastvideo_args.py` instead.
## Inputs
| Parameter | Required | Description |
| ---------- | -------- | -------------------------------------------------------------- |
| `change` | Yes | Add, read, rename, or remove a variable, or change the policy. |
| `variable` | Yes | The variable name, with the `FASTVIDEO_` prefix. |
## Steps
1. **Declare or edit the variable in `fastvideo/envs.py`.**
- Pick the field type and category that the policy doc lists.
- Write a description that states what the variable does and its units.
- To rename, keep the old name in `deprecated_names`. To remove, add the
name to `DEPRECATED_VARIABLES`. Update the uses in `examples/`,
`scripts/`, `docs/`, `apps/`, and the tests.
2. **Read the variable with `envs.NAME.get()` inside a function.**
- In tests, change the value with `envs.NAME.override(value)`, and a variable
outside the registry with `envs.override_external(name, value)`; the
`env_overrides` fixture keeps either until the end of the test.
- Name a variable that only tests read `FASTVIDEO_TEST_*`.
- Do not call `os.environ`, `os.getenv`, or `monkeypatch.setenv` for a
FastVideo variable.
- To set a variable that another tool reads, call `envs.set_external`,
`envs.setdefault_external`, or `envs.unset_external`.
3. **Regenerate the table in the policy doc.**
- Run `python fastvideo/tests/contract/test_env_policy.py`.
4. **Run the contract test.**
- Run `pytest fastvideo/tests/contract/test_env_policy.py`.
- When the test reports a fixed known violation, delete or lower its entry
in `KNOWN_VIOLATIONS`. Never add an entry to `KNOWN_VIOLATIONS`.
5. **When the policy itself changes, update the policy doc and the contract
test in the same pull request.**
- The rules in `docs/contributing/env_vars.md`, the checks and allowlist in
`fastvideo/tests/contract/test_env_policy.py`, and this skill must agree.
## Outputs
- A registry entry in `fastvideo/envs.py` and call sites that use
`envs.NAME.get()`.
- A regenerated table in `docs/contributing/env_vars.md`.
- A passing `fastvideo/tests/contract/test_env_policy.py`.
## Example Usage
```
Add a FASTVIDEO_DEBUG_MY_STAGE switch that logs MyStage inputs.
```
## References
- `docs/contributing/env_vars.md`: the policy, the field types, and the
violation kinds that the contract test reports.
- `fastvideo/envs.py`: the registry.
- `fastvideo/tests/contract/test_env_policy.py`: the contract test,
`EXTERNAL_ALLOWLIST`, and `KNOWN_VIOLATIONS`.
@@ -57,7 +57,7 @@ Hardcoded:
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
`FASTVIDEO_TEST_SSIM_REFERENCE_HF_REPO`).
`FASTVIDEO_SSIM_REFERENCE_HF_REPO`).
- Device folder: `L40S_reference_videos`.
## Prerequisites
@@ -1,51 +0,0 @@
{
"benchmark_id": "wan-t2v-1.3b-1gpu-gb10",
"config_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp1",
"benchmark_version": 3,
"description": "Wan2.1 T2V 1.3B single-GPU inference performance on NVIDIA DGX Spark (GB10). Single-GPU variant of wan-t2v-1.3b (same workload_id for dashboard comparability). Gated to the GB10 via run_config.gpu_types so it does not run on the shared H100/L40S lanes.",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"model_short_name": "Wan2.1-T2V-1.3B"
},
"init_kwargs": {
"num_gpus": 1,
"flow_shift": 7.0,
"sp_size": 1,
"tp_size": 1,
"vae_sp": false,
"vae_tiling": true,
"text_encoder_precisions": ["fp32"]
},
"generation_kwargs": {
"height": 480,
"width": 832,
"num_frames": 45,
"num_inference_steps": 4,
"guidance_scale": 3,
"embedded_cfg_scale": 6,
"seed": 1024,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
},
"test_prompts": [
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
],
"run_config": {
"num_warmup_runs": 2,
"num_measurement_runs": 5,
"required_gpus": 1,
"gpu_types": ["GB10"]
},
"thresholds": {
"GB10": {
"max_generation_time_s": 55.0,
"max_peak_memory_mb": 12000.0
},
"default": {
"max_generation_time_s": 120.0,
"max_peak_memory_mb": 40000.0
}
}
}
+90 -93
View File
@@ -23,100 +23,7 @@ notify:
# dispatcher, and every test payload executes inside the Slinky Slurm tray.
# fastvideo/tests/modal remains available only for an explicit manual rollback;
# no active pipeline or slash-command route invokes it.
# Buildkite hands jobs to free agents in the order they appear here. Golden-gate comes first
# because every later merge lane waits for it; the fastcheck lanes follow from longest to
# shortest measured runtime, so the longest lane never starts last and stretches the build.
steps:
- label: ":test_tube: Golden-Gate Tests"
key: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,golden-gate,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "golden_gate" || build.env("TEST_TYPE") == "golden_gate_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "golden_gate_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: Unit Tests"
key: "unit"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,unit,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "unit_test" || build.env("TEST_TYPE") == "unit_test_ci"))
command: "/opt/fastvideo-ci-runner/run-unit"
timeout_in_minutes: 90
env:
TEST_TYPE: "unit_test_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: Kernel Tests"
key: "kernel-tests"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,kernel-tests,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "kernel_tests" || build.env("TEST_TYPE") == "kernel_tests_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "kernel_tests_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: DreamVerse App Tests"
key: "dreamverse"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,dreamverse,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "dreamverse_app" || build.env("TEST_TYPE") == "dreamverse_app_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "dreamverse_app_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: Encoder Tests"
key: "encoder"
if: |
@@ -186,6 +93,96 @@ steps:
agents:
queue: "ci-runner"
- label: ":microscope: Kernel Tests"
key: "kernel-tests"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,kernel-tests,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "kernel_tests" || build.env("TEST_TYPE") == "kernel_tests_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "kernel_tests_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: Unit Tests"
key: "unit"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,unit,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "unit_test" || build.env("TEST_TYPE") == "unit_test_ci"))
command: "/opt/fastvideo-ci-runner/run-unit"
timeout_in_minutes: 90
env:
TEST_TYPE: "unit_test_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: DreamVerse App Tests"
key: "dreamverse"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,dreamverse,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "dreamverse_app" || build.env("TEST_TYPE") == "dreamverse_app_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "dreamverse_app_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Golden-Gate Tests"
key: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,golden-gate,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "golden_gate" || build.env("TEST_TYPE") == "golden_gate_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "golden_gate_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":bar_chart: SSIM Tests"
key: "ssim"
depends_on: "golden-gate"
+1 -1
View File
@@ -2,4 +2,4 @@
# Canonical Slurm CI selection for the LoRA-extraction lane.
set -euo pipefail
exec pytest ./fastvideo/tests/lora_extraction/ -vs
exec pytest ./fastvideo/tests/lora_extraction/test_lora_extraction.py -vs
+1 -13
View File
@@ -32,7 +32,7 @@ cleanup() {
}
trap cleanup EXIT INT TERM
pytest ./fastvideo/tests/performance/test_inference_performance.py -vs
pytest ./fastvideo/tests/performance -vs
pytest_rc=$?
compare_rc=0
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
@@ -41,18 +41,6 @@ if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
fi
python ./fastvideo/tests/performance/dashboard.py || true
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
# The trusted host relays only .md/.html/.json/.csv from PERF_REPORTS_DIR, so
# mirror each captured worker log with an allowlisted extension.
for worker_log in fastvideo/tests/performance/results/worker_logs/*.log; do
[ -f "$worker_log" ] || continue
base=$(basename "${worker_log%.log}")
# WorkerLogCapture keeps a .log.1 backup after rollover, and read_log_tail
# includes it; mirror that retained history too so the artifact is complete.
if [ -f "$worker_log.1" ]; then
cp -f "$worker_log.1" "$PERF_REPORTS_DIR/${base}.1.md" 2>/dev/null || true
fi
cp -f "$worker_log" "$PERF_REPORTS_DIR/${base}.md" 2>/dev/null || true
done
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
+1 -18
View File
@@ -1,16 +1,7 @@
#!/usr/bin/env bash
set -euo pipefail
# Collect the whole attention directory so new files cannot land uncovered.
# Its FA2/FA3 regression files skip when FA4 is selected (the Modal image
# enables FA4 by default), so pin FA4 off for the directory to be real
# coverage on every runner rather than a nominal collection.
export FASTVIDEO_FA4=0
# The livestream app's tests are CPU-only; its single gpu-marked module is
# deselected, and DreamVerse's GPU tests have their own lane.
exec pytest \
./apps/infinite_livestream/infinite_livestream/tests \
./fastvideo/tests/api/ \
./fastvideo/tests/contract/ \
./fastvideo/tests/dataset/ \
@@ -19,24 +10,16 @@ exec pytest \
./fastvideo/tests/loader/ \
./fastvideo/tests/pipelines/ \
./fastvideo/tests/platforms/ \
./fastvideo/tests/schedulers/ \
./fastvideo/tests/train/ \
./fastvideo/tests/stages/ \
./fastvideo/tests/ops/ \
./fastvideo/tests/worker/ \
./fastvideo/tests/training/test_runner.py \
./fastvideo/tests/training/test_trackers.py \
./fastvideo/tests/training/test_ltx2_rope_fps.py \
./fastvideo/tests/inference/test_basic_fasth3_omniref_pdd.py \
./fastvideo/tests/inference/test_inference_regional_compile.py \
./fastvideo/tests/attention/ \
./fastvideo/tests/layers/test_pdd_linear.py \
./fastvideo/tests/layers/test_triton_fused_norm.py \
./fastvideo/tests/attention/test_sdpa_metadata_mask_contract.py \
./fastvideo/tests/modal/test_kernel_build_cache.py \
./fastvideo/tests/modal/test_pr_test.py \
./fastvideo/tests/modal/test_ssim_test.py \
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
--ignore=./fastvideo/tests/train/models \
--ignore=./fastvideo/tests/train/methods \
-m "not gpu" \
-vs
-7
View File
@@ -61,10 +61,3 @@ Fixes #
**For model/pipeline changes, also check:**
- [ ] I verified SSIM regression tests pass
- [ ] I updated the support matrix if adding a new model
**For new models or serving runtimes:**
- [ ] I added or updated the native serving YAML and [cookbook deployment](https://github.com/hao-ai-lab/FastVideo/blob/main/docs/contributing/cookbook_configuration.md),
or explained why this is not applicable.
- [ ] I documented the GPU type/count and minimum VRAM for the serving setup with supporting measurements,
or marked the minimum as unknown and any VRAM estimate as an estimate.
-10
View File
@@ -160,11 +160,6 @@ FAMILY_COVERAGE = (
("test_glm_image.py", ),
("test_glm_image_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])hunyuan(video)?15([a-z0-9_-]*)([/_.-]|$)"),
(),
("test_hunyuan15_i2v_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
("test_kandinsky5.py", ),
@@ -222,7 +217,6 @@ FAMILY_COVERAGE = (
"test_causal_similarity.py",
"test_wan_i2v_similarity.py",
"test_wan_t2v_similarity.py",
"test_wan_ti2v_similarity.py",
),
),
FamilyCoverage(
@@ -524,10 +518,6 @@ def classify_paths(paths: list[str]) -> MergePlan:
# DreamVerse is already one of the six automatic Fastcheck lanes.
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
continue
if path.startswith("apps/infinite_livestream/"):
# The app's CPU-only tests run in the automatic unit Fastcheck lane.
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if path.startswith("fastvideo/tests/"):
# The automatic unit/component Fastcheck lanes own the remaining
# package tests. Domain-specific expensive test roots were handled
+2 -8
View File
@@ -65,13 +65,7 @@ jobs:
&& fullSuiteOnly.size === 14
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
// Direct reruns may repair a failed suite, never create a gate for
// a suite that did not run.
const failedAggregate = context => data.statuses.some(
s => s.context === context && s.state === 'failure'
);
if (failedAggregate('fastcheck-passed') && fastcheckPassed) {
if (fastcheckPassed) {
core.info(
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
);
@@ -85,7 +79,7 @@ jobs:
});
}
if (failedAggregate('full-suite-passed') && fullSuitePassed) {
if (fullSuitePassed) {
core.info(
'All 20 full suite tests passed — updating full-suite-passed'
);
+11 -14
View File
@@ -33,15 +33,15 @@ jobs:
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: "127.0.0.1"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
GLOO_SOCKET_IFNAME: lo0
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
@@ -49,9 +49,9 @@ jobs:
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.12.0 torchvision torchaudio
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru mlx \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
@@ -65,8 +65,8 @@ jobs:
print("machine:", platform.machine())
print("processor:", platform.processor())
print("mlx default device:", mx.default_device())
device_info = mx.metal.device_info() if mx.metal.is_available() else "metal unavailable"
print("mlx device_info:", device_info)
memory_size = mx.metal.device_info().get("memory_size") if mx.metal.is_available() else "metal unavailable"
print("mlx memory_size:", memory_size)
print("torch:", torch.__version__)
print("torch mps available:", torch.backends.mps.is_available())
PY
@@ -74,7 +74,6 @@ jobs:
- name: Run MLX smoke tests
run: |
python -m pytest \
fastvideo/mlx_runtime/tests/ \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
@@ -93,7 +92,6 @@ jobs:
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
@@ -102,7 +100,7 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s --timeout=120 -o faulthandler_timeout=120
-v -s -o faulthandler_timeout=120
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
@@ -123,6 +121,7 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
@@ -130,16 +129,15 @@ jobs:
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.12.0 torchvision torchaudio
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru "mlx[cpu]" \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Run MLX smoke tests (CPU backend)
run: |
python -m pytest \
fastvideo/mlx_runtime/tests/ \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
@@ -158,7 +156,6 @@ jobs:
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
@@ -167,4 +164,4 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s --timeout=120 -o faulthandler_timeout=120
-v -s -o faulthandler_timeout=120
+3 -23
View File
@@ -32,7 +32,7 @@ jobs:
}
core.setOutput('has_write', String(hasWrite));
- name: Add ready label
- name: Add ready label and react
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
@@ -40,34 +40,14 @@ jobs:
const owner = context.repo.owner;
const repo = context.repo.repo;
const prNumber = context.payload.issue.number;
try { await github.rest.issues.removeLabel({ owner, repo, issue_number: prNumber, name: 'ready' }); } catch {}
await github.rest.issues.addLabels({ owner, repo, issue_number: prNumber, labels: ['ready'] });
- name: React to comment
if: steps.perm.outputs.has_write == 'true'
continue-on-error: true
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
await github.rest.reactions.createForIssueComment({
owner: context.repo.owner,
repo: context.repo.repo,
owner, repo,
comment_id: context.payload.comment.id,
content: 'rocket',
});
trigger-merge-gate:
needs: handle-merge
if: needs.handle-merge.result == 'success'
permissions:
actions: read
contents: read
pull-requests: read
uses: ./.github/workflows/ci-trigger-full-suite.yml
with:
pr_number: ${{ github.event.issue.number }}
secrets:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
parse-command:
if: >-
github.event.issue.pull_request != null
+28 -132
View File
@@ -3,178 +3,74 @@ name: Trigger Merge Gate
on:
pull_request_target:
types: [labeled, synchronize]
workflow_call:
inputs:
pr_number:
description: Pull request number to enter into the merge gate
required: true
type: number
secrets:
BUILDKITE_API_TOKEN:
required: true
permissions:
contents: read
pull-requests: read
actions: read
concurrency:
group: merge-gate-${{ github.event.pull_request.number }}
cancel-in-progress: false
jobs:
trigger:
if: >-
inputs.pr_number > 0
|| (github.event.action == 'labeled' && github.event.label.name == 'ready')
(github.event.action == 'labeled' && github.event.label.name == 'ready')
|| github.event.action == 'synchronize'
runs-on: ubuntu-latest
# Job-level concurrency: only this guarded job acquires the group, so an
# unrelated `labeled` event (which skips the job) cannot cancel an in-flight
# gate and then skip its replacement. The newest real trigger (`ready`,
# push, or `/merge`) supersedes the in-flight run, whose Buildkite build the
# cancel step below replaces.
concurrency:
group: merge-gate-${{ inputs.pr_number || github.event.pull_request.number }}
cancel-in-progress: true
# Gate below may wait for cheap checks (up to MAX_WAIT_SECS = 25 min).
timeout-minutes: 35
steps:
- name: Check ready label
id: check
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
env:
CALLED_PR_NUMBER: ${{ inputs.pr_number }}
with:
script: |
const eventPrNumber = context.payload.pull_request?.number;
const calledPrNumber = Number(process.env.CALLED_PR_NUMBER);
const prNumber = eventPrNumber ?? calledPrNumber;
if (!Number.isSafeInteger(prNumber) || prNumber <= 0) {
core.setFailed(`Invalid pull request number: ${process.env.CALLED_PR_NUMBER}`);
return;
}
const { data: pr } = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: prNumber,
pull_number: context.payload.pull_request.number,
});
if (pr.state !== 'open') {
core.setFailed(`PR #${prNumber} is not open.`);
return;
}
if (pr.base.repo.full_name !== context.payload.repository.full_name
|| pr.base.ref !== context.payload.repository.default_branch) {
core.setFailed(`PR #${prNumber} does not target this repository's default branch.`);
return;
}
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
core.setOutput('changed_files', String(pr.changed_files));
core.setOutput('pr_number', String(pr.number));
core.setOutput('head_sha', pr.head.sha);
core.setOutput('head_ref', pr.head.ref);
core.setOutput('base_sha', pr.base.sha);
core.setOutput('title', pr.title);
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
- name: Cancel previous Buildkite builds
# Cancelling stale builds only saves agent time. If it cannot run, the
# merge gate must still be triggered by the steps below, so a failure
# here is reported and stepped over rather than ending the job.
continue-on-error: true
timeout-minutes: 3
if: steps.check.outputs.has_ready == 'true'
env:
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
set -euo pipefail
response_file=$(mktemp)
builds_file=$(mktemp)
trap 'rm -f "$response_file" "$builds_file"' EXIT
if [[ ! "$PR_NUMBER" =~ ^[1-9][0-9]*$ ]]; then
echo "::warning::Invalid pull request number; stale Buildkite builds may continue."
exit 1
fi
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 --get \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state[]=running" \
--data-urlencode "state[]=scheduled" \
--data-urlencode "state[]=failing" \
--data-urlencode "exclude_jobs=true" \
--data-urlencode "exclude_pipeline=true" \
--output "$response_file" \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds"; then
echo "::warning::Could not list Buildkite builds; stale merge-gate builds may continue."
exit 1
fi
if ! jq -e '
if type != "array" then false
else all(.[];
if type != "object" then false
else
(.number | if type == "number" then . > 0 and floor == . else false end)
and (
(.env? | if . == null then {} else . end) as $env
| if ($env | type) != "object" then false
else
($env.TEST_SCOPE? | . == null or type == "string")
and ($env.PR_NUMBER? | . == null or type == "string")
end
)
end
)
end
' "$response_file" >/dev/null 2>&1; then
# Do not print the response body: it is remote data and may contain
# multiline values that would be interpreted as workflow commands.
echo "::warning::Buildkite returned an invalid build list; stale merge-gate builds may continue."
exit 1
fi
# Match both branch and PR number: forks can reuse the same branch name.
if ! jq -r --arg pr_number "$PR_NUMBER" '
.[]
| select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number))
| .number
' "$response_file" > "$builds_file"; then
echo "::warning::Could not select stale Buildkite builds; stale merge-gate builds may continue."
exit 1
fi
cancellation_failed=0
while IFS= read -r build_num; do
builds=$(curl -sS --get -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state=running,scheduled" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds" \
| jq -r --arg pr_number "$PR_NUMBER" \
'.[] | select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number)) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 -o /dev/null -X PUT \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds/${build_num}/cancel"; then
echo "::warning::Could not cancel Buildkite build #$build_num; trying remaining builds."
cancellation_failed=1
fi
done < "$builds_file"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
if (( cancellation_failed != 0 )); then
exit 1
fi
# Check out the immutable BASE SHA: neither pull_request_target nor the
# privileged slash-command call may run code from the untrusted PR head.
# Check out the immutable BASE SHA: pull_request_target must never run a
# planner or gate script from the untrusted PR head.
- name: Checkout trusted merge planner
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
ref: ${{ steps.check.outputs.base_sha }}
ref: ${{ github.event.pull_request.base.sha }}
persist-credentials: false
- name: Collect changed paths
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
PR_NUMBER: ${{ github.event.pull_request.number }}
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
run: |
set -euo pipefail
@@ -209,18 +105,18 @@ jobs:
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_SHA: ${{ steps.check.outputs.head_sha }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite merge gate
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.check.outputs.head_sha }}
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
PR_TITLE: ${{ steps.check.outputs.title }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
-15
View File
@@ -56,21 +56,6 @@ jobs:
- name: Setup Pages
uses: actions/configure-pages@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '22'
- name: Install cookbook dependencies
run: npm ci --prefix docs
- name: Validate and generate cookbook catalogs
run: |
npm run check:cookbook --prefix docs
npm run test:cookbook --prefix docs
python docs/tests/test_cookbook_assets.py
npm run build:catalog --prefix docs
- name: Build documentation
run: mkdocs build
+9 -11
View File
@@ -68,7 +68,7 @@ jobs:
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
# aarch64 is Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), not
# aarch64 is Blackwell (GB200 sm_100a + DGX Spark / consumer sm_120a), not
# Hopper, and Blackwell needs CUDA >= 12.8 — so only the cu130 leg applies.
# Added via include so x86 keeps cu126 + cu130 while aarch64 stays cu130-only.
include:
@@ -164,20 +164,19 @@ jobs:
cd fastvideo-kernel
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
# Release builds run on GPU-less runners, so set kernels + arch explicitly:
# * aarch64 = Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer)
# covers sm_120a+sm_121a; turbodiffusion covers every listed arch. The
# sm_100 FP4 forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# * aarch64 = Blackwell (GB200 sm_100a + DGX Spark/consumer sm_120a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer, SM120)
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
# + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to requested
# sm_120a/sm_121a and builds the main extension for the full arch list.
# CMAKE_BUILD_PARALLEL_LEVEL caps
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a;12.1a"
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
@@ -251,8 +250,7 @@ jobs:
- name: Download PyPI wheels
# Publish the cu130 (CUDA 13) wheels to PyPI for both architectures:
# x86_64 — Hopper sm_90a TK + consumer Blackwell sm_120a FP4
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_103a/sm_120a/sm_121a)
# + C++ FP4 (sm_120a/sm_121a);
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_120a) + C++ FP4 (sm_120a);
# no TK (Hopper). sm_100 FP4 forward is the FA4 CuTe DSL path in the
# fastvideo package (#1221), shipped/JIT separately.
# The x86_64 cu126 wheel stays available as a build artifact / GitHub-release asset.
-7
View File
@@ -56,9 +56,7 @@ eggs/
# MkDocs documentation
site/
docs/assets/cookbook-serving.json
docs/assets/cookbook-config/
examples/serving/clients/node_modules/
docs/node_modules/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
@@ -136,7 +134,6 @@ openspec/
fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
fastvideo/tests/ssim/.reference_videos_download.lock
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
.kernel_bench/
@@ -145,7 +142,3 @@ fastvideo/tests/ssim/.reference_videos_download.lock
*.nvimlog
.nvimlog
.python-version
/LTX-2-Reference/
/DFDReference/
scripts/benchmarks/minimax_h3_pro6000/headline_results/
fastvideo/tests/ssim/.reference_videos_download.lock
-3
View File
@@ -7,6 +7,3 @@
[submodule "fastvideo/third_party/eval/vbench"]
path = fastvideo/third_party/eval/vbench
url = https://github.com/Vchitect/VBench.git
[submodule "fastvideo/third_party/eval/vqeval"]
path = fastvideo/third_party/eval/vqeval
url = https://github.com/JiusiServe/LongVideoSparseAttention.git
+1 -9
View File
@@ -6,7 +6,7 @@ exclude: |
fastvideo/third_party/.*|
fastvideo-kernel/.*|
assets/.*|
(?<!docs/)tests/.*|
tests/.*|
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/(?!wan/(config|vae_config|pipeline_config|definition|__init__)\.py$).*|
@@ -65,14 +65,6 @@ repos:
language: system
always_run: true
pass_filenames: false
- id: cookbook-javascript
name: Cookbook JavaScript lint and format
entry: npm run check:cookbook --prefix docs
language: system
# CI's trusted manual stage must not execute PR-authored npm scripts.
stages: [pre-commit]
files: ^docs/(js/(build-)?cookbook-.*\.mjs|tests/.*\.test\.mjs|eslint\.config\.mjs|prettier\.config\.mjs|package(-lock)?\.json)$
pass_filenames: false
# Keep `suggestion` last
- id: suggestion
name: Suggestion
+8 -11
View File
@@ -9,12 +9,10 @@
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/10/06`: FastH3 V2 now runs on a single consumer machine: NVIDIA RTX 5090, RTX 4090 and RTX PRO 6000 GPUs, DGX Spark and Apple Silicon. We also release [FastH3 Trim](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4), an experimental pruned model that is 4.2× smaller than base H3 and runs in as little as 8 GB of GPU memory. Get the [models](https://huggingface.co/collections/FastVideo/fastvideo-fasth3) and read the [Blog](https://haoailab.com/blogs/fasth3-rtx/).
- `2026/10/06`: FastVideo now supports [Kandinsky 6](https://x.com/kandinskylab_ai/status/2107374635218055345) from Kandinsky Lab: text- and image-to-video with synchronized audio (base and 10-step distilled pi-Flow checkpoints) plus video super-resolution up to 4x. See the [Kandinsky 6 recipes](https://haoailab.com/FastVideo/cookbook/kandinsky6/).
- `2026/09/15`: Release [FastH3 8-Step V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2), an eight-forward data-free DMD2 checkpoint distilled from MiniMax-H3 with 80% Video Sparse Attention. Run it with `examples/inference/basic/basic_fasth3_8step.py` or the [FastH3 8-Step V2 recipe](https://haoailab.com/FastVideo/cookbook/minimax-h3/).
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac. Follow the [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac—follow the [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
@@ -66,12 +64,13 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
```
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
> **On an Apple Silicon Mac?** Install with `uv pip install -e '.[mlx]'` from
> a clone, then pick a recipe in the
> [cookbook](https://haoailab.com/FastVideo/cookbook/). See the
> [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
> **On an Apple Silicon Mac?** FastVideo runs FastMetal-QAD through an MLX
> runtime. Install with `uv pip install -e '.[mlx]'`, download
> [`FastVideo/FastMetal-1.3B-QAD`](https://huggingface.co/FastVideo/FastMetal-1.3B-QAD),
> and follow the
> [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
@@ -89,7 +88,7 @@ Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virt
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
- Apple Silicon, macOS -> docs/getting_started/installation/mlx.md
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
4. Verify the result:
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
@@ -154,8 +153,6 @@ if __name__ == '__main__':
main()
```
`num_gpus=1` runs the worker in-process (weights load once, no extra Python process). On Colab/Kaggle-style machines with ~16GB host RAM, keep `num_gpus=1`; free-tier system memory does not grow with extra T4s, so `num_gpus>1` is likely to OOM.
Run the script with:
```bash
-34
View File
@@ -158,40 +158,6 @@ dreamverse-server --host 0.0.0.0 --port 8009
The Dreamverse backend defaults to `0.0.0.0:8009` and starts one GPU worker on
the first visible GPU by default.
### Cosmos Predict2.5 DFD continuation (experimental)
Dreamverse can combine two converted Cosmos Predict2.5 2B packages: the
distilled Text2World student creates an unconditioned first segment, then the
Data-Forcing Distillation (DFD) Video2World student conditions each later
segment on the prior terminal frame. Point the runtime at both local converted
packages:
```bash
export DREAMVERSE_MODEL_ID=cosmos25-dfd
export DREAMVERSE_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-Distilled-TrigFlow-FastVideo
export DREAMVERSE_COSMOS25_DFD_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-DFD-FastVideo
export ENABLE_TORCH_COMPILE=0
dreamverse-server --host 0.0.0.0 --port 8009
```
The backend loads and warms both model roles before reporting ready. Both use
BF16, Torch SDPA, 704x1280 output, 24 FPS, and four steps. Bootstrap segments
contain 77 frames. DFD segments contain 81 decoded frames, but Dreamverse drops
the repeated conditioning frame before streaming, leaving 80 new frames. An
initial user image selects DFD immediately without treating that first frame as
a cross-segment overlap.
The profile uses a 30-minute session lease because sequential generation on
GB10-class hardware can exceed Dreamverse's five-minute default while the GPU
is still making progress. Deployments can override the lease with
`FASTVIDEO_SESSION_TIMEOUT_SECONDS`.
Cosmos does not produce audio, so the backend supplies duration-matched silent
24 kHz audio for the existing browser streaming contract and trims 1,000 audio
samples with each repeated DFD boundary frame. Runtime LoRA changes are not
supported. Full segments take roughly 145 seconds on GB10, so this profile is a
continuation-quality integration rather than a real-time configuration.
### Check Readiness
In another shell, verify that the backend process is alive:
+32 -3
View File
@@ -299,18 +299,47 @@ There are three related prompt paths in the current system:
## Initial Image And Segment Handling
The frontend currently sends `initial_image` as part of session init or
The frontend sends `initial_image` and, for first/last frame mode,
`last_frame_image` as part of `session_init_v2`, `project_init_v1`, or
`simple_generate`.
The server:
- validates and persists the image
- uses it only for segment 1 when present
- validates and persists the images
- uses `initial_image` only for segment 1 when present
- keeps continuation state for later segments in the GPU worker
This means the runtime, not the frontend, decides how segment 1 image
conditioning and later continuation conditioning are applied.
## Creation Studio Config
The lobby creation studio sends model, mode, aspect ratio, resolution, and
duration with session init. The server parses these fields into a per-session
creation config and echoes the resolved values back on `gpu_assigned` and
`ltx2_stream_start` as `creation_config`.
Incoming fields on `session_init_v2` and `project_init_v1`:
- `generation_mode`: `t2va`, `fl2va`, or `ref2va` (canonical upstream IDs from #1834)
- `model_id`: `fast-ltx2`, `fast-ltx23`, or `fast-h3`
- `aspect_ratio`: one of `21:9`, `16:9`, `4:3`, `1:1`, `3:4`, `9:16`
- `resolution`: one of `480p`, `720p`, `1080p`, `4k`
- `duration_sec`: `5`, `10`, or `15`
- `initial_image`: optional image payload for reference / first-frame modes
- `last_frame_image`: optional image payload for first/last frame mode
Echoed `creation_config` includes the resolved frame size,
`num_frames`, and `generation_segment_cap` derived from `duration_sec`.
Mode validation:
- `ref2va` requires `initial_image`
- `fl2va` requires both `initial_image` and `last_frame_image`
Per-step generation uses the resolved `frame_width`, `frame_height`, and
`num_frames` from the session creation config.
## Websocket Contract
The websocket is the main integration surface between UI and runtime.
-184
View File
@@ -1,184 +0,0 @@
"""Bounded, runtime-local media library shared by the HTTP and generation APIs."""
from __future__ import annotations
import json
import math
import os
import re
import shutil
import subprocess
import tempfile
import threading
import uuid
from dataclasses import dataclass
from pathlib import Path
from PIL import Image, UnidentifiedImageError
IMAGE_LIMIT = 15 * 1024 * 1024
MEDIA_LIMIT = 100 * 1024 * 1024
STORE_LIMIT = 2 * 1024 * 1024 * 1024
ASSET_LIMIT = 100
MAX_MEDIA_SECONDS = 30
MIME_TYPES = {
"image/png": ("image", ".png"),
"image/jpeg": ("image", ".jpg"),
"image/webp": ("image", ".webp"),
"video/mp4": ("video", ".mp4"),
"video/quicktime": ("video", ".mov"),
"video/webm": ("video", ".webm"),
"audio/mpeg": ("audio", ".mp3"),
"audio/mp4": ("audio", ".m4a"),
"audio/x-m4a": ("audio", ".m4a"),
"audio/wav": ("audio", ".wav"),
"audio/x-wav": ("audio", ".wav"),
"audio/flac": ("audio", ".flac"),
"audio/x-flac": ("audio", ".flac"),
"audio/ogg": ("audio", ".ogg"),
"audio/webm": ("audio", ".webm"),
}
@dataclass(frozen=True)
class StoredAsset:
asset_id: str
kind: str
path: str
name: str
mime_type: str
size: int
def public(self) -> dict:
return {
"asset_id": self.asset_id,
"kind": self.kind,
"name": self.name,
"mime_type": self.mime_type,
"size": self.size,
"url": f"/assets/{self.asset_id}",
}
def validate_media(path: Path, mime_type: str) -> None:
"""Inspect content, not filenames; refuse playlists and non-media uploads."""
kind = MIME_TYPES[mime_type][0]
if kind == "image":
try:
with Image.open(path) as img:
expected = {"image/png": "PNG", "image/jpeg": "JPEG", "image/webp": "WEBP"}[mime_type]
if img.format != expected:
raise ValueError("The image content does not match its file type.")
if img.width * img.height > 16_777_216:
raise ValueError("Images must contain at most 16 megapixels.")
if getattr(img, "is_animated", False):
raise ValueError("Use a still image or upload the animation as a video.")
img.verify()
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
raise ValueError("The image could not be decoded. Use PNG, JPEG, or WebP.") from exc
return
probe = shutil.which(os.getenv("FASTVIDEO_FFPROBE_BIN", "ffprobe"))
if not probe:
raise ValueError("This runtime needs ffprobe installed to accept video and audio assets.")
try:
result = subprocess.run(
[
probe, "-v", "error", "-protocol_whitelist", "file,pipe", "-format_whitelist",
"mov,matroska,webm,mp3,wav,flac,ogg", "-show_format", "-show_streams", "-of", "json",
str(path)
],
check=True,
capture_output=True,
timeout=15,
)
info = json.loads(result.stdout)
formats = set(info.get("format", {}).get("format_name", "").split(","))
if not formats.intersection({"mov", "mp4", "matroska", "webm", "mp3", "wav", "flac", "ogg"}):
raise ValueError("Upload a media file, not a playlist or external reference.")
streams = [stream for stream in info.get("streams", []) if stream.get("codec_type") == kind]
if not streams:
raise ValueError(f"The file contains no {kind} stream.")
for stream in info.get("streams", []):
if stream.get("codec_type") == "audio" and int(stream.get("channels", 0)) not in (1, 2):
raise ValueError("H3 references require mono or stereo audio, including video soundtracks.")
duration = float(info.get("format", {}).get("duration", "nan"))
if not math.isfinite(duration) or not 0 < duration <= MAX_MEDIA_SECONDS:
raise ValueError(f"Reference video and audio must be between 0 and {MAX_MEDIA_SECONDS} seconds long.")
for stream in streams:
if kind == "video" and int(stream.get("width", 0)) * int(stream.get("height", 0)) > 8_294_400:
raise ValueError("Reference videos must be 4K or smaller.")
except (subprocess.SubprocessError, json.JSONDecodeError, OSError) as exc:
raise ValueError("The media file could not be decoded. Check its format and try again.") from exc
class AssetStore:
"""Assets live until deletion or runtime exit; pinned generation inputs cannot be deleted."""
def __init__(self) -> None:
self._directory: tempfile.TemporaryDirectory | None = None
self._assets: dict[str, StoredAsset] = {}
self._pins: dict[str, int] = {}
self._lock = threading.RLock()
def staging_path(self, mime_type: str) -> Path:
with self._lock:
if mime_type not in MIME_TYPES:
raise ValueError("Unsupported media type. Use PNG/JPEG/WebP, MP4/WebM/MOV, or WAV/MP3/M4A/FLAC/OGG.")
if len(self._assets) >= ASSET_LIMIT or sum(item.size for item in self._assets.values()) >= STORE_LIMIT:
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
if self._directory is None:
self._directory = tempfile.TemporaryDirectory(prefix="dreamverse-assets-")
return Path(self._directory.name) / f"{uuid.uuid4().hex}{MIME_TYPES[mime_type][1]}"
def add(self, path: Path, name: str, mime_type: str) -> StoredAsset:
validate_media(path, mime_type)
size = path.stat().st_size
if size == 0 or size > (IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT):
raise ValueError("The asset is empty or exceeds its upload size limit.")
with self._lock:
if len(self._assets) >= ASSET_LIMIT or size + sum(item.size
for item in self._assets.values()) > STORE_LIMIT:
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
if self._directory is None or path.parent != Path(self._directory.name):
raise ValueError("The asset must be uploaded to this runtime.")
display_name = re.sub(r"[\x00-\x1f\x7f/\\]", "_", name).strip()[:200] or "Untitled asset"
asset = StoredAsset(path.stem, MIME_TYPES[mime_type][0], str(path), display_name, mime_type, size)
self._assets[asset.asset_id] = asset
return asset
def get(self, asset_id: str) -> StoredAsset:
with self._lock:
if not isinstance(asset_id, str) or not re.fullmatch(r"[a-f0-9]{32}", asset_id):
raise ValueError("Invalid asset ID. Upload or select an asset from the library.")
asset = self._assets.get(asset_id)
if asset is None or not Path(asset.path).is_file():
raise ValueError("An asset is no longer available. Upload it again and reselect it.")
return asset
def pin(self, asset_ids: list[str]) -> None:
with self._lock:
for asset_id in asset_ids:
self.get(asset_id)
for asset_id in asset_ids:
self._pins[asset_id] = self._pins.get(asset_id, 0) + 1
def release(self, asset_ids: list[str]) -> None:
with self._lock:
for asset_id in asset_ids:
count = self._pins.get(asset_id, 0)
if count > 1:
self._pins[asset_id] = count - 1
else:
self._pins.pop(asset_id, None)
def delete(self, asset_id: str) -> None:
with self._lock:
asset = self.get(asset_id)
if self._pins.get(asset_id, 0):
raise ValueError("This asset is in use by a generation session. End the session before deleting it.")
Path(asset.path).unlink(missing_ok=True)
del self._assets[asset_id]
asset_store = AssetStore()
@@ -111,13 +111,7 @@ def _build_generator_config(model_path: str, enable_compile: bool, num_gpus: int
mode="max-autotune-no-cudagraphs",
dynamic=False),
use_fsdp_inference=False,
# The bundled LTX2 model enables a refinement LoRA during the
# first request. NVFP4 otherwise purges the dense weights that
# FastVideo's LoRA merge path requires.
quantization=QuantizationConfig(
transformer_quant="NVFP4",
transformer_retain_original_weights=True,
),
quantization=QuantizationConfig(transformer_quant="NVFP4"),
),
pipeline=PipelineSelection(
components=components,
+16 -69
View File
@@ -84,37 +84,6 @@ MODEL_REGISTRY = {
"num_inference_steps": 5,
"seed": 1000,
},
"full-h3": {
"name": "MiniMax H3 (Full)",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"attention_backend": "FLASH_ATTN",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 50,
"seed": 1000,
"full_checkpoint": True,
},
"cosmos25-dfd": {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
# Six sequential GB10 segments can exceed the legacy five-minute
# DreamVerse lease even though the GPU is making progress.
"session_timeout_seconds": 1800,
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
@@ -126,6 +95,22 @@ if ACTIVE_MODEL_ID not in MODEL_REGISTRY:
# Active model configuration
MODEL_CONFIG = MODEL_REGISTRY[ACTIVE_MODEL_ID]
# Generation limits
SESSION_TIMEOUT_SECONDS = 300
# Frame settings
NUM_FRAMES = 121
FRAME_HEIGHT = 1088
FRAME_WIDTH = 1920
NUM_INFERENCE_STEPS = 5
JPEG_QUALITY = 100
BATCH_SIZE = 3
# Streaming mode:
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
def _env_int(name: str, default: int) -> int:
value = os.getenv(name)
@@ -202,37 +187,6 @@ def _optional_env(*names: str) -> str | None:
return None
# Generation limits
# Slower backends may own a longer default lease. A profile can set
# ``session_timeout_seconds``; Full H3 loads and generates substantially longer
# than the Preview adapter, which also covers a base/ref pipeline reload inside a
# retained session. An explicit environment override remains available for
# deployment policy: DREAMVERSE_SESSION_TIMEOUT_SECONDS, with
# FASTVIDEO_SESSION_TIMEOUT_SECONDS accepted as an alias.
# Values below 60 seconds are floored so a single segment cannot outlast the session.
_DEFAULT_SESSION_TIMEOUT_SECONDS = cast(
int, MODEL_CONFIG.get("session_timeout_seconds", 7200 if ACTIVE_MODEL_ID == "full-h3" else 300))
SESSION_TIMEOUT_SECONDS = max(
60,
_env_int(
"DREAMVERSE_SESSION_TIMEOUT_SECONDS",
_env_int("FASTVIDEO_SESSION_TIMEOUT_SECONDS", _DEFAULT_SESSION_TIMEOUT_SECONDS),
),
)
# Frame settings
NUM_FRAMES = 121
FRAME_HEIGHT = 1088
FRAME_WIDTH = 1920
NUM_INFERENCE_STEPS = 5
JPEG_QUALITY = 100
BATCH_SIZE = 3
# Streaming mode:
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
@@ -246,13 +200,6 @@ if DREAMVERSE_MODEL_PATH:
"config_model_path": DREAMVERSE_MODEL_PATH,
}
DREAMVERSE_COSMOS25_DFD_MODEL_PATH = (os.getenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", "").strip() or None)
if DREAMVERSE_COSMOS25_DFD_MODEL_PATH and MODEL_CONFIG.get("generation_backend") == "cosmos25_dfd":
MODEL_CONFIG = {
**MODEL_CONFIG,
"continuation_model_path": DREAMVERSE_COSMOS25_DFD_MODEL_PATH,
}
AVAILABLE_LORAS = {
"pixar": {
"repo": "vrgamedevgirl84/LTX_2.3_Pixar_Toon_Style_LoRa",
@@ -1,252 +0,0 @@
"""Cosmos Predict2.5 distilled bootstrap and DFD continuation for DreamVerse."""
from __future__ import annotations
import gc
import os
import time
from typing import TYPE_CHECKING, Any
import numpy as np
import torch
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
if TYPE_CHECKING:
from PIL.Image import Image
_SILENT_AUDIO_SAMPLE_RATE = 24_000
def _required_config_str(model_config: dict, field_name: str) -> str:
value = model_config.get(field_name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"Cosmos Predict2.5 DFD model configuration requires `{field_name}`.")
return value.strip()
class Cosmos25DFDGenerationBackend:
"""Own complementary Cosmos T2W and one-frame-conditioned DFD generators."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.bootstrap_generator: Any | None = None
self.continuation_generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
reserved_gib = torch.cuda.memory_reserved() / 1024**3
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
@staticmethod
def _configure_environment(attention_backend: str) -> None:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = attention_backend
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
@staticmethod
def _load_generator(model_path: str):
from fastvideo import VideoGenerator
return VideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
enable_torch_compile=False,
)
def initialize(self, model_config: dict | None = None) -> None:
"""Load both package roles so bootstrap and continuation are ready."""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("Cosmos Predict2.5 DFD initialization requires a model configuration.")
self.shutdown()
bootstrap_path = _required_config_str(self.model_config, "model_path")
continuation_path = _required_config_str(self.model_config, "continuation_model_path")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
print(f"[GPU {self.gpu_id}] Loading Cosmos T2W bootstrap: {bootstrap_path}")
print(f"[GPU {self.gpu_id}] Before bootstrap load: {self._gpu_mem()}")
self.bootstrap_generator = self._load_generator(bootstrap_path)
print(f"[GPU {self.gpu_id}] Loading Cosmos DFD continuation: {continuation_path}")
self.continuation_generator = self._load_generator(continuation_path)
print(f"[GPU {self.gpu_id}] Cosmos T2W + DFD loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
"""Release both FastVideo generators and the retained terminal frame."""
self.clear_conditioning()
for attr_name in ("bootstrap_generator", "continuation_generator"):
generator = getattr(self, attr_name)
if generator is not None:
try:
generator.shutdown()
except Exception as exc:
print(f"[GPU {self.gpu_id}] Cosmos generator shutdown warning: {exc}")
setattr(self, attr_name, None)
gc.collect()
if torch.cuda.is_available():
torch.cuda.empty_cache()
def clear_conditioning(self) -> None:
if self.continuation_image is not None:
self.continuation_image.close()
self.continuation_image = None
@staticmethod
def _load_rgb_image(image_path: str) -> Image:
from PIL import Image
with Image.open(image_path) as image:
return image.convert("RGB").copy()
def _select_conditioning_image(
self,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> tuple[Image | None, bool]:
if reset_conditioning:
self.clear_conditioning()
if segment_idx > 1 and self.continuation_image is not None:
return self.continuation_image.copy(), True
if segment_idx > 1 and not reset_conditioning:
raise RuntimeError(f"Cosmos DFD segment {segment_idx} requires a retained continuation frame.")
if segment_idx == 1 and image_path:
return self._load_rgb_image(image_path), False
return None, False
def _sampling_param(self, *, conditioned: bool):
# ``num_cond_frames`` is not yet exposed by the typed SamplingConfig,
# so this backend uses the compatibility request until that field lands.
from fastvideo.api.sampling_param import SamplingParam
num_frames_key = "continuation_num_frames" if conditioned else "bootstrap_num_frames"
return SamplingParam(
negative_prompt="",
save_video=False,
return_frames=True,
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
num_frames=int(self.model_config[num_frames_key]),
fps=int(self.model_config["fps"]),
num_inference_steps=int(self.model_config["num_inference_steps"]),
guidance_scale=1.0,
seed=int(self.model_config["seed"]),
num_cond_frames=1 if conditioned else 0,
)
def _save_continuation_frame(self, frame: object) -> None:
from PIL import Image
self.clear_conditioning()
if isinstance(frame, Image.Image):
self.continuation_image = frame.convert("RGB").copy()
return
pixels = np.asarray(frame)
self.continuation_image = Image.fromarray(np.ascontiguousarray(pixels)).convert("RGB")
@staticmethod
def _silent_audio(frame_count: int, fps: int) -> torch.Tensor:
sample_count = max(1, int(round((frame_count / float(fps)) * _SILENT_AUDIO_SAMPLE_RATE)))
return torch.zeros(sample_count, dtype=torch.float32)
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
"""Generate a T2W start or DFD continuation and retain its last frame."""
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
raise ValueError("Cosmos supports text generation only through the generation mode API.")
if self.bootstrap_generator is None or self.continuation_generator is None:
raise RuntimeError("Cosmos T2W + DFD generators are not initialized.")
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
conditioned = conditioning_image is not None
generator = self.continuation_generator if conditioned else self.bootstrap_generator
sampling_param = self._sampling_param(conditioned=conditioned)
started = time.perf_counter()
try:
if conditioned:
sampling_param.pil_image = conditioning_image
result = generator.generate_video(prompt, sampling_param=sampling_param)
finally:
if conditioning_image is not None:
conditioning_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
if not isinstance(result, dict):
raise RuntimeError("Cosmos generation did not return one result dictionary.")
frames = result.get("frames")
expected_frames = int(sampling_param.num_frames)
if not isinstance(frames, list) or len(frames) != expected_frames:
actual_frames = len(frames) if isinstance(frames, list) else None
raise RuntimeError(f"Cosmos generation returned {actual_frames} frames; expected {expected_frames}.")
save_started = time.perf_counter()
self._save_continuation_frame(frames[-1])
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
fps = int(sampling_param.fps)
timings = {
"generation_ms": generation_ms,
"generation_time_ms": float(result.get("generation_time") or 0.0) * 1000.0,
"save_conditioning_ms": save_conditioning_ms,
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
}
trim_frames = 1 if uses_continuation else 0
mode = "DFD continuation" if conditioned else "T2W bootstrap"
print(f"[GPU {self.gpu_id}] Cosmos {mode} segment {segment_idx}: "
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
f"save_conditioning={save_conditioning_ms:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
return StepResult(
frames=frames,
audio=self._silent_audio(len(frames), fps),
audio_sample_rate=_SILENT_AUDIO_SAMPLE_RATE,
timings=timings,
head_trim_frames=trim_frames,
head_trim_audio_frames=trim_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
"""Exercise both T2W bootstrap and retained-frame DFD request shapes."""
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] Cosmos startup warmup starting "
"(synthetic segments: T2W bootstrap, DFD continuation)")
started = time.perf_counter()
bootstrap_result = self.generate_step(warmup_prompt, 1, None, True)
continuation_result = self.generate_step(warmup_prompt, 2, None, False)
total_ms = (time.perf_counter() - started) * 1000.0
self.clear_conditioning()
bootstrap_ms = float(bootstrap_result.timings.get("e2e_latency_ms", 0.0))
continuation_ms = float(continuation_result.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] Cosmos startup warmup complete: "
f"bootstrap={bootstrap_ms:.0f}ms, continuation={continuation_ms:.0f}ms, total={total_ms:.0f}ms")
return {
"warmup_bootstrap_ms": bootstrap_ms,
"warmup_continuation_ms": continuation_ms,
"warmup_total_ms": total_ms,
}
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
del stack
raise RuntimeError("Cosmos Predict2.5 DFD does not support DreamVerse runtime LoRA changes.")
@@ -0,0 +1,137 @@
from __future__ import annotations
from dataclasses import dataclass
from dreamverse.config import MODEL_REGISTRY
# Canonical upstream wire IDs. FL2VA is tracked in #1834 but not wired on Dreamverse
# streaming backends yet.
LTX_LOBBY_GENERATION_MODES = frozenset({"t2va", "ref2va"})
H3_LOBBY_GENERATION_MODES = frozenset({"t2va", "ref2va"})
LTX_LOBBY_ASPECT_RATIOS = frozenset({"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"})
# Realtime FastLTX serving is validated through 1080p-class outputs; 4K is rejected
# until the runtime path is tested on Dreamverse GPUs.
LTX_LOBBY_RESOLUTIONS = frozenset({"480p", "720p", "1080p"})
# FastH3 serves a fixed 768x1344 (16:9-class) output; lobby resolution is nominal.
H3_LOBBY_ASPECT_RATIOS = frozenset({"16:9"})
H3_LOBBY_RESOLUTIONS = frozenset({"720p"})
LOBBY_DURATION_SEC = frozenset({5, 10, 15})
UNSUPPORTED_GENERATION_MODE_MESSAGES = {
"fl2va": "First/last frame mode (FL2VA) is not supported yet.",
}
@dataclass(frozen=True)
class ModelCreationCapabilities:
generation_modes: frozenset[str]
aspect_ratios: frozenset[str]
resolutions: frozenset[str]
duration_sec: frozenset[int]
unsupported_generation_modes: frozenset[str] = frozenset({"fl2va"})
def as_dict(self) -> dict[str, object]:
unsupported = {
mode: UNSUPPORTED_GENERATION_MODE_MESSAGES[mode]
for mode in sorted(self.unsupported_generation_modes)
if mode in UNSUPPORTED_GENERATION_MODE_MESSAGES
}
return {
"generation_modes": sorted(self.generation_modes),
"aspect_ratios": sorted(self.aspect_ratios),
"resolutions": sorted(self.resolutions),
"duration_sec": sorted(self.duration_sec),
"unsupported_generation_modes": unsupported,
"reference_assets": {
"mime_types": ["image/png", "image/jpeg", "image/webp"],
"max_bytes": 15 * 1024 * 1024,
},
}
LTX_MODEL_CREATION_CAPABILITIES = ModelCreationCapabilities(
generation_modes=LTX_LOBBY_GENERATION_MODES,
aspect_ratios=LTX_LOBBY_ASPECT_RATIOS,
resolutions=LTX_LOBBY_RESOLUTIONS,
duration_sec=LOBBY_DURATION_SEC,
)
H3_MODEL_CREATION_CAPABILITIES = ModelCreationCapabilities(
generation_modes=H3_LOBBY_GENERATION_MODES,
aspect_ratios=H3_LOBBY_ASPECT_RATIOS,
resolutions=H3_LOBBY_RESOLUTIONS,
duration_sec=LOBBY_DURATION_SEC,
)
MODEL_CREATION_CAPABILITIES: dict[str, ModelCreationCapabilities] = {
"fast-ltx2": LTX_MODEL_CREATION_CAPABILITIES,
"fast-ltx23": LTX_MODEL_CREATION_CAPABILITIES,
"fast-h3": H3_MODEL_CREATION_CAPABILITIES,
}
def capabilities_for_model(model_id: str) -> ModelCreationCapabilities:
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unknown model_id: {model_id}")
return MODEL_CREATION_CAPABILITIES.get(model_id, LTX_MODEL_CREATION_CAPABILITIES)
def lobby_capabilities_as_dict() -> dict[str, object]:
model_ids = sorted(MODEL_REGISTRY.keys())
models = {model_id: capabilities_for_model(model_id).as_dict() for model_id in model_ids}
union_modes: set[str] = set()
union_aspects: set[str] = set()
union_resolutions: set[str] = set()
union_durations: set[int] = set()
for caps in MODEL_CREATION_CAPABILITIES.values():
union_modes.update(caps.generation_modes)
union_aspects.update(caps.aspect_ratios)
union_resolutions.update(caps.resolutions)
union_durations.update(caps.duration_sec)
return {
"model_ids": model_ids,
"models": models,
"generation_modes": sorted(union_modes),
"aspect_ratios": sorted(union_aspects),
"resolutions": sorted(union_resolutions),
"duration_sec": sorted(union_durations),
"unsupported_generation_modes": dict(UNSUPPORTED_GENERATION_MODE_MESSAGES),
"reference_assets": {
"mime_types": ["image/png", "image/jpeg", "image/webp"],
"max_bytes": 15 * 1024 * 1024,
},
}
# Backward-compatible alias used in tests.
LOBBY_CREATION_CAPABILITIES = lobby_capabilities_as_dict()
def validate_lobby_creation_config(
*,
model_id: str,
generation_mode: str,
aspect_ratio: str,
resolution: str,
duration_sec: int,
) -> None:
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unknown model_id: {model_id}")
caps = capabilities_for_model(model_id)
if generation_mode in caps.unsupported_generation_modes:
raise ValueError(UNSUPPORTED_GENERATION_MODE_MESSAGES[generation_mode])
if generation_mode not in caps.generation_modes:
raise ValueError(f"Unsupported generation_mode: {generation_mode}")
if aspect_ratio not in caps.aspect_ratios:
raise ValueError(f"Unsupported aspect_ratio: {aspect_ratio}")
if resolution not in caps.resolutions:
raise ValueError(f"Unsupported resolution: {resolution}")
if duration_sec not in caps.duration_sec:
raise ValueError("duration_sec must be 5, 10, or 15.")
@@ -5,8 +5,6 @@ from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Protocol
from dreamverse.generation_inputs import GenerationInputs
@dataclass
class StepResult:
@@ -38,7 +36,10 @@ class GenerationBackend(Protocol):
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
...
@@ -1,102 +0,0 @@
"""GPU-independent validation for generation modes and ordered asset handles."""
from __future__ import annotations
from dataclasses import dataclass
from PIL import Image, UnidentifiedImageError
from dreamverse.assets import asset_store
GENERATION_MODES = ("t2va", "fl2va", "ref2va")
@dataclass(frozen=True)
class GenerationAsset:
asset_id: str
kind: str
path: str
role: str
@dataclass(frozen=True)
class GenerationInputs:
mode: str | None = None
assets: tuple[GenerationAsset, ...] = ()
@property
def first_frame_path(self) -> str | None:
return next((asset.path for asset in self.assets if asset.role == "first_frame"), None)
@property
def last_frame_path(self) -> str | None:
return next((asset.path for asset in self.assets if asset.role == "last_frame"), None)
@property
def references(self) -> tuple[GenerationAsset, ...]:
return tuple(asset for asset in self.assets if asset.role == "reference")
def supported_generation_modes(model_id: str) -> tuple[str, ...]:
return GENERATION_MODES if model_id in ("full-h3", "mock") else ("t2va", )
def resolve_generation_inputs(payload: dict, model_id: str) -> GenerationInputs:
mode = payload.get("generation_mode")
raw_assets = payload.get("conditioning_assets", [])
if mode is None and "generation_mode" not in payload:
if raw_assets:
raise ValueError("Select a generation mode before attaching conditioning assets.")
return GenerationInputs()
if not isinstance(mode, str) or mode not in GENERATION_MODES:
raise ValueError("Unknown generation mode. Choose T2VA, FL2VA, or Ref2VA.")
if mode not in supported_generation_modes(model_id):
raise ValueError(f"{mode.upper()} requires the Full H3 runtime. This runtime is running {model_id}.")
if payload.get("initial_image") is not None:
raise ValueError("Use asset IDs for generation modes; do not combine them with the legacy initial_image field.")
if not isinstance(raw_assets, list) or len(raw_assets) > 12:
raise ValueError("conditioning_assets must be an ordered list with at most 12 assets.")
if mode == "t2va" and raw_assets:
raise ValueError("T2VA accepts text only. Remove conditioning assets or choose another mode.")
assets: list[GenerationAsset] = []
for item in raw_assets:
if not isinstance(item, dict) or set(item) != {"asset_id", "role"}:
raise ValueError("Each conditioning asset must contain only asset_id and role.")
role = item["role"]
if role not in ("first_frame", "last_frame", "reference"):
raise ValueError("Asset role must be first_frame, last_frame, or reference.")
stored = asset_store.get(item["asset_id"])
assets.append(GenerationAsset(stored.asset_id, stored.kind, stored.path, role))
if mode == "fl2va":
if any(asset.kind != "image" or asset.role == "reference" for asset in assets):
raise ValueError("FL2VA accepts only first-frame and last-frame images.")
if sum(asset.role == "first_frame" for asset in assets) != 1:
raise ValueError("FL2VA requires exactly one first-frame image.")
if sum(asset.role == "last_frame" for asset in assets) > 1:
raise ValueError("FL2VA accepts at most one last-frame image.")
elif mode == "ref2va":
if not assets or any(asset.role != "reference" for asset in assets):
raise ValueError("Ref2VA requires an ordered list of reference assets, without keyframe roles.")
if not any(asset.kind in ("image", "video") for asset in assets):
raise ValueError("Ref2VA requires at least one image or video; audio alone is not supported.")
for kind, limit in (("image", 9), ("video", 3), ("audio", 3)):
if sum(asset.kind == kind for asset in assets) > limit:
raise ValueError(f"Ref2VA accepts at most {limit} {kind} references.")
for asset in assets:
if asset.kind == "image":
try:
with Image.open(asset.path) as image:
if image.width > 4 * image.height or image.height > 4 * image.width:
raise ValueError(
"Ref2VA image aspect ratios must be between 1:4 and 4:1. Crop this image first.")
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
raise ValueError("A selected reference image could not be decoded. Upload it again.") from exc
return GenerationInputs(mode, tuple(assets))
def pin_generation_inputs(inputs: GenerationInputs) -> None:
asset_store.pin([asset.asset_id for asset in inputs.assets])
def release_generation_inputs(inputs: GenerationInputs) -> None:
asset_store.release([asset.asset_id for asset in inputs.assets])
@@ -4,7 +4,6 @@ from __future__ import annotations
from dreamverse.config import MODEL_CONFIG
from dreamverse.generation_contracts import GenerationBackend, StepResult
from dreamverse.generation_inputs import GenerationInputs
def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBackend:
@@ -17,10 +16,6 @@ def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBack
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
return MiniMaxH3GenerationBackend(gpu_id)
if backend_name == "cosmos25_dfd":
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
return Cosmos25DFDGenerationBackend(gpu_id)
raise ValueError(f"Unsupported DreamVerse generation backend: {backend_name!r}")
@@ -85,7 +80,10 @@ class VideoGenerationWorker:
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Generate one segment through the selected model backend."""
return self._require_backend().generate_step(
@@ -93,7 +91,9 @@ class VideoGenerationWorker:
segment_idx,
image_path,
reset_conditioning,
generation_inputs=generation_inputs,
frame_width=frame_width,
frame_height=frame_height,
num_frames=num_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
+11 -29
View File
@@ -29,7 +29,6 @@ from dreamverse.av_streaming import (
generate_stream_id,
stream_fmp4,
)
from dreamverse.generation_inputs import GenerationInputs, pin_generation_inputs, release_generation_inputs
from dreamverse.worker_ipc import (
CommandPayload,
InitAck,
@@ -190,7 +189,9 @@ def gpu_worker_process(
segment_idx,
image_path=payload.image_path,
reset_conditioning=payload.reset_conditioning,
generation_inputs=payload.generation_inputs,
frame_width=payload.frame_width,
frame_height=payload.frame_height,
num_frames=payload.num_frames,
)
head_trim_frames = step_result.head_trim_frames
head_trim_audio_frames = step_result.head_trim_audio_frames
@@ -434,7 +435,6 @@ class GPUSlot:
self.connected_users: set[str] = set()
self._pending_futures: dict[str, asyncio.Future] = {}
self._stream_queues: dict[str, asyncio.Queue] = {}
self._step_asset_inputs: dict[str, GenerationInputs] = {}
self._response_reader_task: asyncio.Task | None = None
self._active: bool = False
self._reader_lock: asyncio.Lock | None = None
@@ -666,9 +666,6 @@ class GPUSlot:
if isinstance(event, (StepComplete, WarmupComplete)):
event.timings["ipc_get_done_ns"] = time.time_ns()
if isinstance(event, (StepComplete, WorkerError)) and event.user_id is not None:
self._release_step_assets(event.user_id)
user_id = event.user_id
if user_id and user_id in self._pending_futures:
future = self._pending_futures.pop(user_id)
@@ -759,7 +756,10 @@ class GPUSlot:
segment_idx: int = 1,
image_path: str | None = None,
reset_conditioning: bool = False,
generation_inputs: GenerationInputs | None = None,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> dict[str, float]:
"""Execute a generation step for a specific user.
@@ -773,19 +773,12 @@ class GPUSlot:
segment_idx=segment_idx,
image_path=image_path,
reset_conditioning=bool(reset_conditioning),
generation_inputs=generation_inputs,
frame_width=frame_width,
frame_height=frame_height,
num_frames=num_frames,
)
if generation_inputs is not None:
if user_id in self._step_asset_inputs:
raise RuntimeError("The previous generation is still using this project's assets.")
pin_generation_inputs(generation_inputs)
self._step_asset_inputs[user_id] = generation_inputs
# Pins intentionally survive a waiter timeout/cancellation: the GPU
# command keeps running. The response reader releases them when the
# worker actually completes (even if that response is now unmatched).
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
timeout=1800.0)
self._release_step_assets(user_id)
match response:
case StepComplete(timings=timings):
return timings
@@ -795,11 +788,6 @@ class GPUSlot:
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
f"{type(response).__name__}")
def _release_step_assets(self, user_id: str) -> None:
inputs = self._step_asset_inputs.pop(user_id, None)
if inputs is not None:
release_generation_inputs(inputs)
async def apply_lora_stack(
self,
stack: list[tuple[str, float]],
@@ -822,9 +810,7 @@ class GPUSlot:
async def leave_user(self, user_id: str) -> None:
"""Remove a user from this GPU."""
try:
response = await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
if isinstance(response, LeaveAck):
self._release_step_assets(user_id)
await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
except Exception as e:
print(f"[GPU {self.gpu_id}] Leave user error: {e}")
finally:
@@ -861,10 +847,6 @@ class GPUSlot:
except Exception:
pass
if self.process is None or not self.process.is_alive():
for user_id in list(self._step_asset_inputs):
self._release_step_assets(user_id)
for q in (self.command_queue, self.response_queue):
if q is not None:
try:
+8 -14
View File
@@ -33,7 +33,6 @@ from dreamverse.config import (
_resolve_lora_spec,
)
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
# Multi-frame decoded continuation defaults from
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
@@ -291,13 +290,7 @@ class LTX2GenerationBackend:
dynamic=False,
),
use_fsdp_inference=False,
# The bundled LTX2 model enables a refinement LoRA during the
# first request. NVFP4 otherwise purges the dense weights that
# FastVideo's LoRA merge path requires.
quantization=QuantizationConfig(
transformer_quant="NVFP4",
transformer_retain_original_weights=True,
),
quantization=QuantizationConfig(transformer_quant="NVFP4"),
),
pipeline=PipelineSelection(
components=components,
@@ -461,11 +454,12 @@ class LTX2GenerationBackend:
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Execute one generation step; snapshot state for the next segment."""
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
raise ValueError("LTX supports text generation only through the generation mode API.")
timings: dict = {}
prompt = self._inject_style_trigger(prompt)
@@ -474,9 +468,9 @@ class LTX2GenerationBackend:
prompt=prompt,
negative_prompt="",
save_video=False,
height=FRAME_HEIGHT,
width=FRAME_WIDTH,
num_frames=NUM_FRAMES,
height=frame_height or FRAME_HEIGHT,
width=frame_width or FRAME_WIDTH,
num_frames=num_frames or NUM_FRAMES,
fps=24,
num_inference_steps=NUM_INFERENCE_STEPS,
guidance_scale=1.0,
+2 -9
View File
@@ -15,7 +15,6 @@ from dreamverse.gpu_pool import GPUPool, get_available_gpus
from dreamverse.session_logger import SessionEventLogger
from dreamverse.config import (
ACTIVE_MODEL_ID,
AVAILABLE_LORAS,
DEVTOOLS_ENABLED,
FRONTEND_STATIC_DIR_CANDIDATES,
@@ -34,9 +33,8 @@ from dreamverse.routes.presets import (
prompt_config_router,
curated_presets_router,
)
from dreamverse.routes.creation import creation_router
from dreamverse.session.controller import SessionController
from dreamverse.generation_inputs import supported_generation_modes
from dreamverse.routes.assets import router as asset_router
class _HeartbeatAccessLogFilter(logging.Filter):
@@ -95,16 +93,11 @@ app.add_middleware(
app.include_router(build_health_router(lambda: runtime.gpu_pool))
app.include_router(internal_monitor_router)
app.include_router(prompt_config_router)
app.include_router(asset_router)
app.include_router(creation_router)
if DEVTOOLS_ENABLED:
app.include_router(curated_presets_router)
@app.get("/generation-capabilities")
async def generation_capabilities() -> dict:
return {"model_id": ACTIVE_MODEL_ID, "modes": supported_generation_modes(ACTIVE_MODEL_ID), "mock": False}
@app.websocket("/ws")
async def websocket_endpoint(websocket: WebSocket):
controller = SessionController(
@@ -1,4 +1,4 @@
"""Full/Preview H3 lifecycle, conditioning and per-project pipeline selection."""
"""FastH3 model lifecycle and first-frame continuation for DreamVerse."""
from __future__ import annotations
@@ -12,7 +12,6 @@ import torch
from dreamverse.config import DREAMVERSE_SP_SIZE
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
if TYPE_CHECKING:
from PIL.Image import Image
@@ -27,14 +26,13 @@ def _required_config_str(model_config: dict, field_name: str) -> str:
class MiniMaxH3GenerationBackend:
"""Own one H3 pipeline at a time and retain base-pipeline continuation."""
"""Run the VSA data-free FastH3 adapter and retain one continuation frame."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
self.pipeline_mode = "base"
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
@@ -53,37 +51,32 @@ class MiniMaxH3GenerationBackend:
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
def initialize(self, model_config: dict | None = None) -> None:
"""Load the profile's base pipeline; Ref2VA is loaded on first use."""
"""Download the fixed Preview adapter and load the FastH3 generator.
The model profile owns the base checkpoint, adapter file, attention
backend, and generation geometry. The backend translates that profile
into FastVideo's typed generator configuration.
"""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("FastH3 initialization requires a model configuration.")
self._load_pipeline("base")
def _load_pipeline(self, pipeline_mode: str) -> None:
"""Unload the old executor before loading a base or reference transformer.
GPU worker commands are serialized, so a project boundary never swaps
weights while another request is using them. Keeping one executor also
avoids simultaneously retaining two large H3 transformers in VRAM. A
failed load leaves no executor behind so the next step retries it.
"""
full_checkpoint = bool(self.model_config.get("full_checkpoint", False))
if pipeline_mode == "ref2va" and not full_checkpoint:
raise ValueError("Ref2VA requires the full-h3 model profile.")
if self.generator is not None:
previous_generator = self.generator
self.generator.shutdown()
self.generator = None
previous_generator.shutdown()
del previous_generator
gc.collect()
torch.cuda.empty_cache()
self.clear_conditioning()
model_path = _required_config_str(self.model_config, "model_path")
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
from huggingface_hub import hf_hub_download
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig,
@@ -95,20 +88,10 @@ class MiniMaxH3GenerationBackend:
PipelineSelection,
)
components = ComponentConfig()
if not full_checkpoint:
from huggingface_hub import hf_hub_download
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
components.lora_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
components.lora_strength = 1.0
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
if pipeline_mode == "ref2va":
components.override_pipeline_cls_name = "MiniMaxH3Ref2VAModularPipeline"
adapter_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
experimental = {
"attention_backend": attention_backend,
"inference_torch_compile": not full_checkpoint and attention_backend == "FLASH_ATTN",
"inference_torch_compile": attention_backend == "FLASH_ATTN",
"vae_parallel_decode": True,
"vae_parallel_decode_strategy": "gather",
}
@@ -120,8 +103,7 @@ class MiniMaxH3GenerationBackend:
generator_config = GeneratorConfig(
model_path=model_path,
pipeline=PipelineSelection(
workload_type="i2v" if pipeline_mode == "ref2va" else None,
components=components,
components=ComponentConfig(lora_path=adapter_path, lora_strength=1.0),
experimental=experimental,
),
engine=EngineConfig(
@@ -133,23 +115,17 @@ class MiniMaxH3GenerationBackend:
text_encoder=True,
image_encoder=True,
vae=True,
pin_cpu_memory=not full_checkpoint,
pin_cpu_memory=True,
),
compile=CompileConfig(enabled=False, vae_enabled=True),
use_fsdp_inference=full_checkpoint and DREAMVERSE_SP_SIZE > 1,
use_fsdp_inference=False,
),
)
print(f"[GPU {self.gpu_id}] Loading H3 model: {model_path} ({pipeline_mode})")
print(f"[GPU {self.gpu_id}] Loading FastH3 model: {model_path}")
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
try:
self.generator = VideoGenerator.from_config(generator_config)
except Exception:
# The old executor is already gone; leaving no executor behind lets
# the next step retry this load instead of stranding the GPU slot.
self.generator = None
raise
self.pipeline_mode = pipeline_mode
self.generator = VideoGenerator.from_config(generator_config)
print(f"[GPU {self.gpu_id}] FastH3 loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
@@ -190,28 +166,14 @@ class MiniMaxH3GenerationBackend:
return self._load_rgb_image(image_path), False
return None, False
def _build_request(
self,
prompt: str,
conditioning_image: Image | None,
last_image: Image | None = None,
generation_inputs: GenerationInputs | None = None,
):
def _build_request(self, prompt: str, conditioning_image: Image | None):
"""Build the typed FastVideo request owned by the FastH3 profile."""
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
references = None
if generation_inputs is not None and generation_inputs.mode == "ref2va":
from fastvideo.api import MiniMaxH3Reference
references = [
MiniMaxH3Reference(source=str(asset.path), media_type=asset.kind)
for asset in generation_inputs.references
]
return GenerationRequest(
prompt=prompt,
negative_prompt="",
inputs=InputConfig(pil_image=conditioning_image, last_image=last_image, references=references),
inputs=InputConfig(pil_image=conditioning_image),
sampling=SamplingConfig(
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
@@ -238,7 +200,10 @@ class MiniMaxH3GenerationBackend:
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
*,
frame_width: int | None = None,
frame_height: int | None = None,
num_frames: int | None = None,
) -> StepResult:
"""Generate one synchronized FastH3 segment and retain its last frame.
@@ -246,46 +211,21 @@ class MiniMaxH3GenerationBackend:
conditioned frame and its matching audio duration are trimmed before
streaming so adjacent segments do not duplicate media.
"""
mode = generation_inputs.mode if generation_inputs is not None else None
if mode not in (None, "t2va", "fl2va", "ref2va"):
raise ValueError(f"Unsupported H3 generation mode: {mode!r}.")
if mode in ("fl2va", "ref2va") and not self.model_config.get("full_checkpoint", False):
raise ValueError(f"{mode.upper()} requires the full-h3 model profile.")
pipeline_mode = "ref2va" if mode == "ref2va" else "base"
if self.generator is None or self.pipeline_mode != pipeline_mode:
# A failed switch leaves no executor behind; reload here so the
# slot recovers on the next step instead of staying broken.
if segment_idx > 1 and not reset_conditioning:
raise ValueError("Generation mode cannot change in the middle of a project.")
self._load_pipeline(pipeline_mode)
conditioning_image = None
last_image = None
uses_continuation = False
if mode == "ref2va":
# The reference pipeline rejects first/last-frame inputs. Preserve
# all original references for every clip and do not trim overlap.
self.clear_conditioning()
else:
if mode == "fl2va" and generation_inputs is not None:
image_path = generation_inputs.first_frame_path
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
del frame_width, frame_height, num_frames
if self.generator is None:
raise RuntimeError("FastH3 generator is not initialized.")
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
request = self._build_request(prompt, conditioning_image)
started = time.perf_counter()
try:
if (mode == "fl2va" and segment_idx == 1 and generation_inputs is not None
and generation_inputs.last_frame_path):
last_image = self._load_rgb_image(generation_inputs.last_frame_path)
request = self._build_request(prompt, conditioning_image, last_image, generation_inputs)
result = self.generator.generate(request)
finally:
if conditioning_image is not None:
conditioning_image.close()
if last_image is not None:
last_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
@@ -300,8 +240,7 @@ class MiniMaxH3GenerationBackend:
raise RuntimeError("FastH3 returned audio without an audio sample rate.")
save_started = time.perf_counter()
if mode != "ref2va":
self._save_continuation_frame(frames)
self._save_continuation_frame(frames)
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
timings = {
"generation_ms": generation_ms,
+66 -46
View File
@@ -31,15 +31,9 @@ from fastapi.staticfiles import StaticFiles
from dreamverse._deps import require_dreamverse_runtime_deps
from dreamverse.config import FRONTEND_STATIC_DIR_CANDIDATES, GENERATION_SEGMENT_CAP
from dreamverse.creation_capabilities import lobby_capabilities_as_dict
from dreamverse.session_creation_config import parse_session_creation_config, validate_generation_mode_assets
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
from dreamverse.generation_inputs import (
GenerationInputs,
pin_generation_inputs,
release_generation_inputs,
resolve_generation_inputs,
supported_generation_modes,
)
from dreamverse.routes.assets import router as asset_router
LATENCY_MS = 200
SESSION_TIMEOUT_SECONDS = 300
@@ -178,12 +172,6 @@ app.add_middleware(
allow_methods=["*"],
allow_headers=["*"],
)
app.include_router(asset_router)
@app.get("/generation-capabilities")
async def generation_capabilities():
return {"model_id": "mock", "modes": supported_generation_modes("mock"), "mock": True}
@app.get("/healthz")
@@ -239,6 +227,11 @@ async def prompt_system_config():
}
@app.get("/creation-capabilities")
async def creation_capabilities():
return lobby_capabilities_as_dict()
@app.get("/curated-presets")
async def curated_presets():
presets = [
@@ -304,7 +297,8 @@ async def websocket_endpoint(websocket: WebSocket):
send_lock = asyncio.Lock()
stop_event = asyncio.Event()
session_init_image = None
generation_inputs = GenerationInputs()
session_last_frame_image = None
session_creation_config = None
async def ws_send_json(payload: dict) -> None:
async with send_lock:
@@ -362,27 +356,41 @@ async def websocket_endpoint(websocket: WebSocket):
generation_paused = bool(initial_rollout_prompt and not single_clip_mode and len(curated_prompts) == 0)
try:
generation_inputs = resolve_generation_inputs(init_data, "mock")
pin_generation_inputs(generation_inputs)
session_init_image = persist_session_init_image(init_data.get("initial_image"))
session_last_frame_image = persist_session_init_image(init_data.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid initial image")
return
try:
session_creation_config = parse_session_creation_config(init_data)
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid creation config")
return
timeout_task = asyncio.create_task(session_timeout())
await ws_send_json({
gpu_assigned_payload: dict[str, object] = {
"type": "gpu_assigned",
"gpu_id": 0,
"session_timeout": SESSION_TIMEOUT_SECONDS,
"generation_mode": generation_inputs.mode,
"mock": True,
})
}
if session_creation_config is not None:
gpu_assigned_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(gpu_assigned_payload)
raw_prompt_queue: asyncio.Queue[PromptSubmission] = asyncio.Queue()
ready_prompt_queue: asyncio.Queue[ReadyPrompt] = asyncio.Queue()
@@ -411,8 +419,16 @@ async def websocket_endpoint(websocket: WebSocket):
if previous_session_image is not None:
cleanup_session_init_image(previous_session_image)
def replace_last_frame_image(last_frame_payload: object) -> None:
nonlocal session_last_frame_image
next_last_frame_image = persist_session_init_image(last_frame_payload)
previous_last_frame_image = session_last_frame_image
session_last_frame_image = next_last_frame_image
if previous_last_frame_image is not None:
cleanup_session_init_image(previous_last_frame_image)
async def send_stream_start(seed_reason: str) -> None:
await ws_send_json({
stream_start_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
@@ -420,8 +436,15 @@ async def websocket_endpoint(websocket: WebSocket):
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": 0,
})
"generation_segment_cap": (
session_creation_config.generation_segment_cap
if session_creation_config is not None
else GENERATION_SEGMENT_CAP
),
}
if session_creation_config is not None:
stream_start_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(stream_start_payload)
if seed_reason == "init":
await ws_send_json({
"type": "seed_prompts_updated",
@@ -506,7 +529,6 @@ async def websocket_endpoint(websocket: WebSocket):
})
async def apply_project_init_payload(payload: dict[str, object], ) -> bool:
nonlocal generation_inputs
nonlocal preset_id
nonlocal preset_label
nonlocal initial_rollout_prompt
@@ -530,6 +552,7 @@ async def websocket_endpoint(websocket: WebSocket):
nonlocal project_active
nonlocal project_stream_started
nonlocal pending_project_end
nonlocal session_creation_config
next_initial_rollout_prompt = str(payload.get("initial_rollout_prompt") or "").strip()
next_preset_id = str(payload.get("preset_id") or "").strip()
@@ -540,19 +563,25 @@ async def websocket_endpoint(websocket: WebSocket):
]
try:
next_inputs = resolve_generation_inputs(payload, "mock")
pin_generation_inputs(next_inputs)
try:
replace_session_image(payload.get("initial_image"))
except ValueError:
release_generation_inputs(next_inputs)
raise
release_generation_inputs(generation_inputs)
generation_inputs = next_inputs
replace_session_image(payload.get("initial_image"))
replace_last_frame_image(payload.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
return False
try:
session_creation_config = parse_session_creation_config(payload)
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return False
@@ -598,7 +627,6 @@ async def websocket_endpoint(websocket: WebSocket):
return drained
async def enter_project_idle() -> None:
nonlocal generation_inputs
nonlocal seed_prompt_memory
nonlocal curated_prompts
nonlocal curated_idx
@@ -618,8 +646,6 @@ async def websocket_endpoint(websocket: WebSocket):
dropped_raw = drain_queue_nowait(raw_prompt_queue)
dropped_ready = drain_queue_nowait(ready_prompt_queue)
release_generation_inputs(generation_inputs)
generation_inputs = GenerationInputs()
seed_prompt_memory = []
curated_prompts = []
curated_idx = 0
@@ -800,16 +826,10 @@ async def websocket_endpoint(websocket: WebSocket):
continue
try:
if generation_inputs.mode is not None and data.get("initial_image") is not None:
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
"cannot replace generation mode inputs.")
if "generation_mode" in data or "conditioning_assets" in data:
raise ValueError("simple_generate cannot change the mode; use project_init_v1.")
replace_session_image(data.get("initial_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
continue
@@ -1221,7 +1241,7 @@ async def websocket_endpoint(websocket: WebSocket):
finally:
stop_event.set()
cleanup_session_init_image(session_init_image)
release_generation_inputs(generation_inputs)
cleanup_session_init_image(session_last_frame_image)
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
@@ -1,72 +0,0 @@
"""Raw, bounded media uploads keep large binary data out of websocket messages."""
from __future__ import annotations
import asyncio
from urllib.parse import unquote
from fastapi import APIRouter, HTTPException, Request, Response
from fastapi.responses import FileResponse
from starlette.concurrency import run_in_threadpool
from dreamverse.assets import IMAGE_LIMIT, MEDIA_LIMIT, MIME_TYPES, asset_store
router = APIRouter()
_upload_lock = asyncio.Lock()
@router.post("/assets", status_code=201)
async def upload_asset(request: Request) -> dict:
mime_type = request.headers.get("content-type", "").split(";", 1)[0].lower()
if mime_type not in MIME_TYPES:
raise HTTPException(415, "Unsupported asset type. Select a supported image, video, or audio file.")
limit = IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT
try:
if int(request.headers.get("content-length", "0")) > limit:
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
except ValueError as exc:
raise HTTPException(400, "Invalid Content-Length.") from exc
async with _upload_lock:
try:
path = asset_store.staging_path(mime_type)
except ValueError as exc:
raise HTTPException(400, str(exc)) from exc
try:
size = 0
with path.open("xb") as handle:
async for chunk in request.stream():
size += len(chunk)
if size > limit:
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
await run_in_threadpool(handle.write, chunk)
asset = await run_in_threadpool(asset_store.add, path,
unquote(request.headers.get("x-asset-name", "Untitled asset")), mime_type)
return asset.public()
except ValueError as exc:
path.unlink(missing_ok=True)
raise HTTPException(400, str(exc)) from exc
except BaseException:
path.unlink(missing_ok=True)
raise
@router.api_route("/assets/{asset_id}", methods=["GET", "HEAD"])
async def get_asset(asset_id: str) -> FileResponse:
try:
asset = asset_store.get(asset_id)
except ValueError as exc:
raise HTTPException(404, str(exc)) from exc
return FileResponse(asset.path, media_type=asset.mime_type, headers={"X-Content-Type-Options": "nosniff"})
@router.delete("/assets/{asset_id}", status_code=204)
async def delete_asset(asset_id: str) -> Response:
try:
asset_store.get(asset_id)
except ValueError as exc:
raise HTTPException(404, str(exc)) from exc
try:
asset_store.delete(asset_id)
except ValueError as exc:
raise HTTPException(409, str(exc)) from exc
return Response(status_code=204)
@@ -0,0 +1,14 @@
"""Creation studio capability routes."""
from __future__ import annotations
from fastapi import APIRouter
from dreamverse.creation_capabilities import lobby_capabilities_as_dict
creation_router = APIRouter(tags=["creation"])
@creation_router.get("/creation-capabilities")
async def creation_capabilities() -> dict[str, object]:
return lobby_capabilities_as_dict()
+108 -92
View File
@@ -26,13 +26,8 @@ from typing import TYPE_CHECKING
from fastapi import WebSocket, WebSocketDisconnect
from dreamverse.gpu_pool import GPUSlot
from dreamverse.generation_inputs import (
GenerationInputs,
pin_generation_inputs,
release_generation_inputs,
resolve_generation_inputs,
)
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
from dreamverse.session_creation_config import parse_session_creation_config, validate_generation_mode_assets
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
from dreamverse.config import (
@@ -162,7 +157,9 @@ class SessionController:
prompt_worker_task: asyncio.Task | None = None
rewrite_seed_prompts_task: asyncio.Task | None = None
session_init_image = None
generation_inputs: GenerationInputs | None = None
session_last_frame_image = None
session_creation_config = None
session_generation_segment_cap = GENERATION_SEGMENT_CAP
async def session_timeout():
"""Close the session after timeout."""
@@ -198,18 +195,6 @@ class SessionController:
init_data = {}
init_type = init_data.get("type")
try:
next_generation_inputs = resolve_generation_inputs(init_data, ACTIVE_MODEL_ID)
pin_generation_inputs(next_generation_inputs)
generation_inputs = next_generation_inputs
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
await websocket.close(code=1008, reason="Invalid generation inputs")
return
preset_id = init_data.get("preset_id")
preset_label = str(init_data.get("preset_label") or "").strip()
initial_rollout_prompt = str(init_data.get("initial_rollout_prompt") or "").strip()
@@ -256,6 +241,7 @@ class SessionController:
try:
session_init_image = persist_session_init_image(init_data.get("initial_image"))
session_last_frame_image = persist_session_init_image(init_data.get("last_frame_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
@@ -264,6 +250,22 @@ class SessionController:
await websocket.close(code=1003, reason="Invalid initial image")
return
try:
session_creation_config = parse_session_creation_config(init_data)
session_generation_segment_cap = session_creation_config.generation_segment_cap
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid creation config")
return
if preset_id:
print(f"Client {client_id[:8]} selected preset: {preset_id} "
f"label={preset_label or '(unset)'} "
@@ -275,6 +277,16 @@ class SessionController:
if session_init_image is not None:
print(f"Client {client_id[:8]} uploaded initial image: "
f"{session_init_image.display_name}")
if session_last_frame_image is not None:
print(f"Client {client_id[:8]} uploaded last frame image: "
f"{session_last_frame_image.display_name}")
if session_creation_config is not None:
print(f"Client {client_id[:8]} creation config: "
f"model={session_creation_config.model_id}, "
f"mode={session_creation_config.generation_mode}, "
f"size={session_creation_config.frame_width}x{session_creation_config.frame_height}, "
f"duration={session_creation_config.duration_sec}s, "
f"segment_cap={session_creation_config.generation_segment_cap}")
# Acquire a GPU slot.
gpu_id, slot = await self.gpu_pool.acquire(client_id, websocket)
@@ -283,15 +295,20 @@ class SessionController:
timeout_task = asyncio.create_task(session_timeout())
# Join the engine on this GPU.
await slot.join_user(client_id, model_id=ACTIVE_MODEL_ID)
await slot.join_user(
client_id,
model_id=session_creation_config.model_id if session_creation_config is not None else ACTIVE_MODEL_ID,
)
# Notify client they're connected to a GPU.
await ws_send_json({
gpu_assigned_payload: dict[str, object] = {
"type": "gpu_assigned",
"gpu_id": gpu_id,
"session_timeout": SESSION_TIMEOUT_SECONDS,
"generation_mode": generation_inputs.mode,
})
}
if session_creation_config is not None:
gpu_assigned_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(gpu_assigned_payload)
await log_event(
"gpu_assigned",
{
@@ -335,6 +352,14 @@ class SessionController:
if previous_session_init_image is not None:
cleanup_session_init_image(previous_session_init_image)
def replace_last_frame_image(last_frame_payload: object) -> None:
nonlocal session_last_frame_image
next_last_frame_image = persist_session_init_image(last_frame_payload)
previous_last_frame_image = session_last_frame_image
session_last_frame_image = next_last_frame_image
if previous_last_frame_image is not None:
cleanup_session_init_image(previous_last_frame_image)
async def schedule_simple_generate_request(payload: dict[str, object]) -> None:
nonlocal preset_id
nonlocal preset_label
@@ -381,16 +406,10 @@ class SessionController:
return
try:
if generation_inputs.mode is not None and payload.get("initial_image") is not None:
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
"cannot replace generation mode inputs.")
if "generation_mode" in payload or "conditioning_assets" in payload:
raise ValueError("simple_generate cannot change the mode; use project_init_v1.")
replace_session_init_image(payload.get("initial_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return
@@ -443,7 +462,6 @@ class SessionController:
})
async def apply_project_init_payload(payload: dict[str, object]) -> bool:
nonlocal generation_inputs
nonlocal preset_id
nonlocal preset_label
nonlocal initial_rollout_prompt
@@ -479,6 +497,8 @@ class SessionController:
nonlocal project_active
nonlocal project_stream_started
nonlocal pending_project_end
nonlocal session_creation_config
nonlocal session_generation_segment_cap
next_initial_rollout_prompt = str(payload.get("initial_rollout_prompt") or "").strip()
next_enhancement_enabled = bool(payload.get("enhancement_enabled", True))
@@ -523,25 +543,30 @@ class SessionController:
})
return False
next_generation_inputs = None
next_inputs_pinned = False
try:
next_generation_inputs = resolve_generation_inputs(payload, ACTIVE_MODEL_ID)
pin_generation_inputs(next_generation_inputs)
next_inputs_pinned = True
replace_session_init_image(payload.get("initial_image"))
replace_last_frame_image(payload.get("last_frame_image"))
except ValueError as exc:
if next_inputs_pinned:
release_generation_inputs(next_generation_inputs)
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return False
release_generation_inputs(generation_inputs)
generation_inputs = next_generation_inputs
try:
session_creation_config = parse_session_creation_config(payload)
session_generation_segment_cap = session_creation_config.generation_segment_cap
validate_generation_mode_assets(
session_creation_config.generation_mode,
has_initial_image=session_init_image is not None,
has_last_frame_image=session_last_frame_image is not None,
)
except ValueError as exc:
await ws_send_json({
"type": "error",
"message": str(exc),
})
return False
initial_rollout_prompt = next_initial_rollout_prompt
enhancement_enabled = next_enhancement_enabled
@@ -979,7 +1004,7 @@ class SessionController:
"segment_cap":
_resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
continue
@@ -1209,7 +1234,6 @@ class SessionController:
return drained
async def enter_project_idle() -> None:
nonlocal generation_inputs
nonlocal curated_prompts
nonlocal seed_prompt_memory
nonlocal curated_idx
@@ -1260,9 +1284,6 @@ class SessionController:
project_active = False
pending_project_end = False
release_generation_inputs(generation_inputs)
generation_inputs = GenerationInputs()
if project_stream_started:
project_stream_started = False
await ws_send_json({"type": "ltx2_stream_complete"})
@@ -1323,27 +1344,22 @@ class SessionController:
))
else:
project_stream_started = True
await ws_send_json({
"type":
"ltx2_stream_start",
"total_segments":
len(curated_prompts),
"preset_id":
preset_id,
"stream_mode":
"av_fmp4",
"live_mode":
True,
"loop_generation_enabled":
loop_generation_enabled,
"loop_iteration":
loop_iteration,
"generation_segment_cap":
_resolve_generation_segment_cap(
stream_start_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
"stream_mode": "av_fmp4",
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": _resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
}
if session_creation_config is not None:
stream_start_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(stream_start_payload)
await ws_send_json({
"type": "seed_prompts_updated",
"prompts": seed_prompt_memory,
@@ -1382,27 +1398,22 @@ class SessionController:
loop_iteration += 1
project_stream_started = True
await ws_send_json({
"type":
"ltx2_stream_start",
"total_segments":
len(curated_prompts),
"preset_id":
preset_id,
"stream_mode":
"av_fmp4",
"live_mode":
True,
"loop_generation_enabled":
loop_generation_enabled,
"loop_iteration":
loop_iteration,
"generation_segment_cap":
_resolve_generation_segment_cap(
restart_stream_payload: dict[str, object] = {
"type": "ltx2_stream_start",
"total_segments": len(curated_prompts),
"preset_id": preset_id,
"stream_mode": "av_fmp4",
"live_mode": True,
"loop_generation_enabled": loop_generation_enabled,
"loop_iteration": loop_iteration,
"generation_segment_cap": _resolve_generation_segment_cap(
single_clip_mode=single_clip_mode,
cap=GENERATION_SEGMENT_CAP,
cap=session_generation_segment_cap,
),
})
}
if session_creation_config is not None:
restart_stream_payload["creation_config"] = session_creation_config.as_dict()
await ws_send_json(restart_stream_payload)
if nonlocal_reason == "loop_restart":
await ws_send_json({
"type": "loop_restarted",
@@ -1431,13 +1442,14 @@ class SessionController:
pending_simple_prompt_submission = None
if (not single_clip_mode and not generation_cap_blocked and not rollout_waiting_for_rewrite
and GENERATION_SEGMENT_CAP > 0 and generated_segment_count >= GENERATION_SEGMENT_CAP):
and session_generation_segment_cap > 0
and generated_segment_count >= session_generation_segment_cap):
loop_generation_enabled = False
rollout_waiting_for_rewrite = True
_main_print(
"INFO",
f"Segment cap reached for client {client_id[:8]} "
f"(cap_segments={GENERATION_SEGMENT_CAP}, "
f"(cap_segments={session_generation_segment_cap}, "
f"generated_segments={generated_segment_count}); "
"waiting for rollout rewrite",
)
@@ -1662,6 +1674,9 @@ class SessionController:
pending_reset_conditioning = False
step_image_path = (str(session_init_image.file_path)
if segment_idx == 1 and session_init_image is not None else None)
step_frame_width = session_creation_config.frame_width if session_creation_config is not None else None
step_frame_height = session_creation_config.frame_height if session_creation_config is not None else None
step_num_frames = session_creation_config.num_frames if session_creation_config is not None else None
step_task = asyncio.create_task(
slot.user_step(
client_id,
@@ -1669,7 +1684,9 @@ class SessionController:
segment_idx=segment_idx,
image_path=step_image_path,
reset_conditioning=step_reset_conditioning,
generation_inputs=generation_inputs,
frame_width=step_frame_width,
frame_height=step_frame_height,
num_frames=step_num_frames,
))
segment_generation_active = True
try:
@@ -1724,10 +1741,10 @@ class SessionController:
print(f"[GPU {gpu_id}] Unknown AV event: "
f"{type(event).__name__}")
# A GPU command cannot be cancelled by cancelling its
# asyncio waiter. Await completion before releasing pinned
# asset files or making this GPU available to a new user.
timings = await step_task
if not step_task.done():
step_task.cancel()
else:
timings = await step_task
finally:
segment_generation_active = False
if not step_task.done():
@@ -1851,5 +1868,4 @@ class SessionController:
await self.gpu_pool.release(client_id)
finally:
cleanup_session_init_image(session_init_image)
if generation_inputs is not None:
release_generation_inputs(generation_inputs)
cleanup_session_init_image(session_last_frame_image)
@@ -0,0 +1,140 @@
from __future__ import annotations
from dataclasses import dataclass
from dreamverse.config import FRAME_HEIGHT, FRAME_WIDTH, GENERATION_SEGMENT_CAP, MODEL_REGISTRY, NUM_FRAMES
from dreamverse.creation_capabilities import validate_lobby_creation_config
LTX_LOBBY_MODEL_IDS = frozenset(MODEL_REGISTRY.keys())
SUPPORTED_GENERATION_MODES = frozenset({"t2va", "fl2va", "ref2va"})
SUPPORTED_ASPECT_RATIOS = frozenset({"21:9", "16:9", "4:3", "1:1", "3:4", "9:16"})
SUPPORTED_RESOLUTIONS = frozenset({"480p", "720p", "1080p", "4k"})
SEGMENT_DURATION_SEC = 5
@dataclass(frozen=True)
class SessionCreationConfig:
model_id: str
generation_mode: str
aspect_ratio: str
resolution: str
duration_sec: int
frame_width: int
frame_height: int
num_frames: int
generation_segment_cap: int
def as_dict(self) -> dict[str, object]:
return {
"model_id": self.model_id,
"generation_mode": self.generation_mode,
"aspect_ratio": self.aspect_ratio,
"resolution": self.resolution,
"duration_sec": self.duration_sec,
"frame_width": self.frame_width,
"frame_height": self.frame_height,
"num_frames": self.num_frames,
"generation_segment_cap": self.generation_segment_cap,
}
def _round_to_multiple(value: float, multiple: int = 32) -> int:
rounded = int(round(value / multiple)) * multiple
return max(multiple, rounded)
def _resolution_base(resolution: str) -> int:
return {
"480p": 480,
"720p": 720,
"1080p": 1080,
"4k": 2160,
}.get(resolution, 720)
def resolve_frame_size(aspect_ratio: str, resolution: str) -> tuple[int, int]:
if aspect_ratio == "16:9" and resolution == "1080p":
return FRAME_WIDTH, FRAME_HEIGHT
base = _resolution_base(resolution)
width_ratio, height_ratio = {
"21:9": (21, 9),
"16:9": (16, 9),
"4:3": (4, 3),
"1:1": (1, 1),
"3:4": (3, 4),
"9:16": (9, 16),
}.get(aspect_ratio, (16, 9))
if width_ratio >= height_ratio:
height = _round_to_multiple(base)
width = _round_to_multiple(height * width_ratio / height_ratio)
else:
width = _round_to_multiple(base)
height = _round_to_multiple(width * height_ratio / width_ratio)
return width, height
def duration_sec_to_segment_cap(duration_sec: int, *, global_cap: int = GENERATION_SEGMENT_CAP) -> int:
requested = max(1, int(round(duration_sec / SEGMENT_DURATION_SEC + 0.0001)))
if global_cap <= 0:
return requested
return max(1, min(requested, global_cap))
def parse_session_creation_config(payload: dict[str, object]) -> SessionCreationConfig:
raw_model_id = str(payload.get("model_id") or "").strip()
model_id = raw_model_id if raw_model_id in LTX_LOBBY_MODEL_IDS else "fast-ltx23"
generation_mode = str(payload.get("generation_mode") or "t2va").strip()
if generation_mode not in SUPPORTED_GENERATION_MODES:
raise ValueError(f"Unsupported generation_mode: {generation_mode}")
aspect_ratio = str(payload.get("aspect_ratio") or "16:9").strip()
if aspect_ratio not in SUPPORTED_ASPECT_RATIOS:
raise ValueError(f"Unsupported aspect_ratio: {aspect_ratio}")
resolution = str(payload.get("resolution") or "720p").strip()
if resolution not in SUPPORTED_RESOLUTIONS:
raise ValueError(f"Unsupported resolution: {resolution}")
try:
duration_sec = int(payload.get("duration_sec") or SEGMENT_DURATION_SEC)
except (TypeError, ValueError) as exc:
raise ValueError("duration_sec must be an integer.") from exc
if duration_sec not in {5, 10, 15}:
raise ValueError("duration_sec must be 5, 10, or 15.")
validate_lobby_creation_config(
model_id=model_id,
generation_mode=generation_mode,
aspect_ratio=aspect_ratio,
resolution=resolution,
duration_sec=duration_sec,
)
if model_id not in MODEL_REGISTRY:
raise ValueError(f"Unsupported model_id: {model_id}")
frame_width, frame_height = resolve_frame_size(aspect_ratio, resolution)
return SessionCreationConfig(
model_id=model_id,
generation_mode=generation_mode,
aspect_ratio=aspect_ratio,
resolution=resolution,
duration_sec=duration_sec,
frame_width=frame_width,
frame_height=frame_height,
num_frames=NUM_FRAMES,
generation_segment_cap=duration_sec_to_segment_cap(duration_sec),
)
def validate_generation_mode_assets(
generation_mode: str,
*,
has_initial_image: bool,
has_last_frame_image: bool,
) -> None:
if generation_mode == "ref2va" and not has_initial_image:
raise ValueError("Ref2VA mode requires a reference image.")
@@ -139,48 +139,12 @@ def test_config_enables_prompt_safety_when_requested(monkeypatch):
def test_config_uses_five_minute_session_timeout(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.delenv("DREAMVERSE_MODEL_ID", raising=False)
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 300
def test_config_uses_thirty_minute_cosmos25_session_timeout(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 1800
def test_config_allows_session_timeout_override(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 900
def test_config_prefers_dreamverse_session_timeout_over_alias(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.setenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", "1200")
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 1200
def test_config_rejects_invalid_prompt_provider(monkeypatch):
monkeypatch.setenv("FASTVIDEO_PROMPT_PROVIDER", "unsupported")
_set_required_prompt_keys(monkeypatch)
@@ -222,57 +186,3 @@ def test_config_uses_fasth3_sequence_parallel_default(monkeypatch):
assert module.ACTIVE_MODEL_ID == "fast-h3"
assert module.MODEL_CONFIG["generation_backend"] == "minimax_h3"
assert module.DREAMVERSE_SP_SIZE == 4
def test_full_h3_profile_has_no_preview_adapter_and_longer_session(monkeypatch):
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "full-h3")
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.MODEL_CONFIG["full_checkpoint"] is True
assert "adapter_repo" not in module.MODEL_CONFIG
assert module.MODEL_CONFIG["num_inference_steps"] == 50
assert module.DREAMVERSE_SP_SIZE == 4
assert module.SESSION_TIMEOUT_SECONDS == 7200
def test_config_registers_cosmos25_dfd_profile(monkeypatch):
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.MODEL_REGISTRY["cosmos25-dfd"] == {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
"session_timeout_seconds": 1800,
}
def test_config_selects_cosmos25_package_roles(monkeypatch, tmp_path):
_set_required_prompt_keys(monkeypatch)
bootstrap_path = tmp_path / "cosmos25-t2w"
continuation_path = tmp_path / "cosmos25-dfd"
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.setenv("DREAMVERSE_MODEL_PATH", str(bootstrap_path))
monkeypatch.setenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", str(continuation_path))
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
module = _load_config_module()
assert module.ACTIVE_MODEL_ID == "cosmos25-dfd"
assert module.MODEL_CONFIG["generation_backend"] == "cosmos25_dfd"
assert module.MODEL_CONFIG["model_path"] == str(bootstrap_path)
assert module.MODEL_CONFIG["continuation_model_path"] == str(continuation_path)
assert module.DREAMVERSE_SP_SIZE == 1
@@ -1,219 +0,0 @@
from __future__ import annotations
import os
from pathlib import Path
from types import SimpleNamespace
import numpy as np
import pytest
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
from dreamverse.generation_inputs import GenerationInputs
COSMOS_CONFIG = {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "/models/cosmos25-t2w",
"continuation_model_path": "/models/cosmos25-dfd",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
}
class _RecordingGenerator:
def __init__(self, pixel_value: int = 20) -> None:
self.pixel_value = pixel_value
self.calls: list[dict] = []
self.shutdown_calls = 0
def generate_video(self, prompt, sampling_param):
condition = sampling_param.pil_image
self.calls.append({
"prompt": prompt,
"sampling": sampling_param,
"conditioning_pixels": None if condition is None else np.asarray(condition).copy(),
})
frames = [
np.full((2, 3, 3), self.pixel_value, dtype=np.uint8)
for _ in range(sampling_param.num_frames)
]
frames[-1] = np.full((2, 3, 3), self.pixel_value + 1, dtype=np.uint8)
return {
"frames": frames,
"generation_time": 0.25,
}
def shutdown(self):
self.shutdown_calls += 1
@pytest.fixture
def backend(monkeypatch) -> Cosmos25DFDGenerationBackend:
instance = Cosmos25DFDGenerationBackend(gpu_id=0)
instance.model_config = dict(COSMOS_CONFIG)
instance.bootstrap_generator = _RecordingGenerator(pixel_value=20)
instance.continuation_generator = _RecordingGenerator(pixel_value=40)
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.synchronize", lambda: None)
def fake_sampling_param(*, conditioned):
return SimpleNamespace(
negative_prompt="",
save_video=False,
return_frames=True,
height=704,
width=1280,
num_frames=81 if conditioned else 77,
fps=24,
num_inference_steps=4,
guidance_scale=1.0,
seed=42,
num_cond_frames=1 if conditioned else 0,
pil_image=None,
)
monkeypatch.setattr(instance, "_sampling_param", fake_sampling_param)
return instance
def test_initialize_loads_both_package_roles(monkeypatch):
loaded_paths = []
generators = [_RecordingGenerator(), _RecordingGenerator()]
backend = Cosmos25DFDGenerationBackend(gpu_id=0)
def fake_load(model_path):
loaded_paths.append(model_path)
return generators[len(loaded_paths) - 1]
monkeypatch.setattr(backend, "_load_generator", fake_load)
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.gc.collect", lambda: 0)
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.is_available", lambda: False)
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
backend.initialize(COSMOS_CONFIG)
assert loaded_paths == [
"/models/cosmos25-t2w",
"/models/cosmos25-dfd",
]
assert backend.bootstrap_generator is generators[0]
assert backend.continuation_generator is generators[1]
assert backend.model_config == COSMOS_CONFIG
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "TORCH_SDPA"
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
def test_unconditioned_start_uses_t2w_and_retains_terminal_frame(backend):
result = backend.generate_step("first prompt", 1, None, True)
assert len(backend.bootstrap_generator.calls) == 1
assert backend.continuation_generator.calls == []
sampling = backend.bootstrap_generator.calls[0]["sampling"]
assert sampling.height == 704
assert sampling.width == 1280
assert sampling.num_frames == 77
assert sampling.fps == 24
assert sampling.num_inference_steps == 4
assert sampling.guidance_scale == 1.0
assert sampling.seed == 42
assert sampling.num_cond_frames == 0
assert sampling.pil_image is None
assert result.head_trim_frames == 0
assert result.head_trim_audio_frames == 0
assert result.audio_sample_rate == 24_000
assert result.audio.shape == (77_000, )
assert result.audio.count_nonzero() == 0
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 21).tolist()
def test_retained_frame_uses_dfd_and_trims_repeated_boundary(backend):
backend.generate_step("first prompt", 1, None, True)
result = backend.generate_step("pivot right", 2, None, False)
assert len(backend.continuation_generator.calls) == 1
call = backend.continuation_generator.calls[0]
sampling = call["sampling"]
assert sampling.num_frames == 81
assert sampling.num_cond_frames == 1
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 21).tolist()
assert result.head_trim_frames == 1
assert result.head_trim_audio_frames == 1
assert result.audio.shape == (81_000, )
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 41).tolist()
def test_initial_image_uses_dfd_without_stream_trim(backend, tmp_path: Path):
from PIL import Image
image_path = tmp_path / "initial.png"
Image.fromarray(np.full((2, 3, 3), 7, dtype=np.uint8)).save(image_path)
result = backend.generate_step("animate", 1, str(image_path), True)
assert backend.bootstrap_generator.calls == []
call = backend.continuation_generator.calls[0]
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 7).tolist()
assert result.head_trim_frames == 0
assert result.head_trim_audio_frames == 0
def test_generation_mode_api_accepts_text_only_and_rejects_conditioning_modes(backend):
result = backend.generate_step("first prompt", 1, None, True, generation_inputs=GenerationInputs(mode="t2va"))
assert len(backend.bootstrap_generator.calls) == 1
assert result.head_trim_frames == 0
with pytest.raises(ValueError, match="text generation only"):
backend.generate_step("pivot right", 2, None, False, generation_inputs=GenerationInputs(mode="fl2va"))
assert backend.continuation_generator.calls == []
def test_missing_later_continuation_fails_before_generation(backend):
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
backend.generate_step("later prompt", 2, None, False)
assert backend.bootstrap_generator.calls == []
assert backend.continuation_generator.calls == []
def test_reset_later_segment_uses_fresh_t2w_bootstrap(backend):
backend.generate_step("first prompt", 1, None, True)
result = backend.generate_step("new scene", 2, None, True)
assert len(backend.bootstrap_generator.calls) == 2
assert backend.continuation_generator.calls == []
assert result.head_trim_frames == 0
def test_warmup_exercises_bootstrap_and_dfd_paths(backend):
timings = backend.warmup("warmup prompt")
assert len(backend.bootstrap_generator.calls) == 1
assert len(backend.continuation_generator.calls) == 1
assert backend.continuation_image is None
assert "warmup_bootstrap_ms" in timings
assert "warmup_continuation_ms" in timings
assert "warmup_total_ms" in timings
def test_shutdown_releases_both_generators_and_conditioning(backend):
bootstrap = backend.bootstrap_generator
continuation = backend.continuation_generator
backend.generate_step("first prompt", 1, None, True)
backend.shutdown()
assert bootstrap.shutdown_calls == 1
assert continuation.shutdown_calls == 1
assert backend.bootstrap_generator is None
assert backend.continuation_generator is None
assert backend.continuation_image is None
@@ -0,0 +1,74 @@
import pytest
from dreamverse.creation_capabilities import (
capabilities_for_model,
lobby_capabilities_as_dict,
validate_lobby_creation_config,
)
def test_lobby_capabilities_include_all_registry_models():
caps = lobby_capabilities_as_dict()
assert set(caps["model_ids"]) == {"fast-ltx2", "fast-ltx23", "fast-h3"}
assert "fl2va" not in caps["generation_modes"]
assert "4k" not in caps["resolutions"]
def test_fast_h3_capabilities_use_fixed_geometry():
h3_caps = capabilities_for_model("fast-h3")
assert h3_caps.generation_modes == frozenset({"t2va", "ref2va"})
assert h3_caps.aspect_ratios == frozenset({"16:9"})
assert h3_caps.resolutions == frozenset({"720p"})
def test_validate_lobby_creation_config_accepts_supported_t2va():
validate_lobby_creation_config(
model_id="fast-ltx23",
generation_mode="t2va",
aspect_ratio="16:9",
resolution="1080p",
duration_sec=5,
)
def test_validate_lobby_creation_config_accepts_fast_h3():
validate_lobby_creation_config(
model_id="fast-h3",
generation_mode="ref2va",
aspect_ratio="16:9",
resolution="720p",
duration_sec=10,
)
def test_validate_lobby_creation_config_rejects_fl2va():
with pytest.raises(ValueError, match="FL2VA"):
validate_lobby_creation_config(
model_id="fast-ltx23",
generation_mode="fl2va",
aspect_ratio="16:9",
resolution="720p",
duration_sec=5,
)
def test_validate_lobby_creation_config_rejects_4k():
with pytest.raises(ValueError, match="Unsupported resolution"):
validate_lobby_creation_config(
model_id="fast-ltx2",
generation_mode="t2va",
aspect_ratio="16:9",
resolution="4k",
duration_sec=10,
)
def test_validate_lobby_creation_config_rejects_invalid_h3_aspect():
with pytest.raises(ValueError, match="Unsupported aspect_ratio"):
validate_lobby_creation_config(
model_id="fast-h3",
generation_mode="t2va",
aspect_ratio="9:16",
resolution="720p",
duration_sec=5,
)
@@ -1,248 +0,0 @@
"""Contract regressions independent of CUDA and model weights."""
import io
import asyncio
import shutil
import subprocess
import pytest
from fastapi import FastAPI
from fastapi.testclient import TestClient
from PIL import Image
from dreamverse import assets, generation_inputs
from dreamverse.generation_inputs import resolve_generation_inputs
from dreamverse.routes import assets as asset_routes
from dreamverse.tests.test_mock_server import _FakeWebSocket
@pytest.fixture
def library(monkeypatch):
store = assets.AssetStore()
monkeypatch.setattr(asset_routes, "asset_store", store)
monkeypatch.setattr(generation_inputs, "asset_store", store)
app = FastAPI()
app.include_router(asset_routes.router)
with TestClient(app) as client:
yield store, client
def upload_image(client, color="red"):
image_bytes = io.BytesIO()
Image.new("RGB", (32, 32), color).save(image_bytes, format="PNG")
response = client.post("/assets", content=image_bytes.getvalue(),
headers={"Content-Type": "image/png", "X-Asset-Name": "frame.png"})
assert response.status_code == 201, response.text
return response.json()
def conditioning(asset, role):
return {"asset_id": asset["asset_id"], "role": role}
def test_assets_validate_content_and_support_head_range_and_delete(library):
store, client = library
asset = upload_image(client)
assert set(asset) == {"asset_id", "kind", "name", "mime_type", "size", "url"}
assert client.head(asset["url"]).status_code == 200
response = client.get(asset["url"], headers={"Range": "bytes=0-7"})
assert response.status_code == 206
assert response.content == b"\x89PNG\r\n\x1a\n"
assert client.post("/assets", content=b"not an image", headers={"Content-Type": "image/png"}).status_code == 400
assert client.post("/assets", content=b"<svg/>", headers={"Content-Type": "image/svg+xml"}).status_code == 415
assert client.post("/assets", content=b"", headers={"Content-Type": "image/png",
"Content-Length": str(assets.IMAGE_LIMIT + 1)}).status_code == 413
with pytest.raises(ValueError, match="Invalid asset ID"):
store.get("../../etc/passwd")
assert client.delete(asset["url"]).status_code == 204
assert client.head(asset["url"]).status_code == 404
def test_generation_pin_prevents_deletion_until_session_releases(library):
_, client = library
asset = upload_image(client)
inputs = resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
conditioning(asset, "first_frame")
]}, "full-h3")
generation_inputs.pin_generation_inputs(inputs)
generation_inputs.pin_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 409
generation_inputs.release_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 409
generation_inputs.release_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 204
def test_legacy_init_remains_compatible_but_explicit_t2va_is_text_only(library):
_, client = library
assert resolve_generation_inputs({"initial_image": {"old": "payload"}}, "fast-ltx2").mode is None
assert resolve_generation_inputs({"generation_mode": "t2va"}, "fast-ltx2").mode == "t2va"
with pytest.raises(ValueError, match="legacy initial_image"):
resolve_generation_inputs({"generation_mode": "t2va", "initial_image": {}}, "full-h3")
asset = upload_image(client)
with pytest.raises(ValueError, match="text only"):
resolve_generation_inputs({"generation_mode": "t2va", "conditioning_assets": [
conditioning(asset, "reference")
]}, "full-h3")
@pytest.mark.parametrize("mode", ["unknown", None, 3, [], {}])
def test_unknown_mode_fails_before_assets_are_resolved(mode):
with pytest.raises(ValueError, match="Unknown generation mode"):
resolve_generation_inputs({"generation_mode": mode}, "full-h3")
@pytest.mark.parametrize("model_id", ["fast-h3", "fast-ltx2", "fast-ltx23"])
def test_preview_and_ltx_cannot_advertise_full_h3_modes(model_id):
with pytest.raises(ValueError, match="Full H3"):
resolve_generation_inputs({"generation_mode": "ref2va"}, model_id)
def test_fl2va_first_required_last_optional_and_roles_unique(library):
_, client = library
first = upload_image(client)
last = upload_image(client, "blue")
payload = {"generation_mode": "fl2va", "conditioning_assets": [conditioning(first, "first_frame")]}
inputs = resolve_generation_inputs(payload, "full-h3")
assert inputs.first_frame_path.endswith(".png")
assert inputs.last_frame_path is None
payload["conditioning_assets"].append(conditioning(last, "last_frame"))
assert resolve_generation_inputs(payload, "full-h3").last_frame_path is not None
payload["conditioning_assets"].append(conditioning(first, "first_frame"))
with pytest.raises(ValueError, match="exactly one first-frame"):
resolve_generation_inputs(payload, "full-h3")
with pytest.raises(ValueError, match="exactly one first-frame"):
resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
conditioning(last, "last_frame")
]}, "full-h3")
def test_ref_order_is_preserved_and_limits_are_enforced(library):
_, client = library
first, second = upload_image(client), upload_image(client, "blue")
refs = [conditioning(second, "reference"), conditioning(first, "reference")]
payload = {"generation_mode": "ref2va", "conditioning_assets": refs}
inputs = resolve_generation_inputs(payload, "full-h3")
assert [asset.asset_id for asset in inputs.references] == [second["asset_id"], first["asset_id"]]
with pytest.raises(ValueError, match="at most 9 image"):
resolve_generation_inputs({**payload, "conditioning_assets": refs * 5}, "full-h3")
with pytest.raises(ValueError, match="without keyframe roles"):
resolve_generation_inputs({**payload, "conditioning_assets": [conditioning(first, "first_frame")]}, "full-h3")
with pytest.raises(ValueError, match="at most 12"):
resolve_generation_inputs({**payload, "conditioning_assets": refs * 7}, "full-h3")
def test_ref_audio_requires_visual_reference(library, monkeypatch):
store, _ = library
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "audio", "/audio.wav", "audio",
"audio/wav", 100))
with pytest.raises(ValueError, match="audio alone"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
{"asset_id": "a" * 32, "role": "reference"}
]}, "full-h3")
def test_ref_rejects_extreme_image_aspect_before_gpu(library):
_, client = library
content = io.BytesIO()
Image.new("RGB", (500, 50), "blue").save(content, format="PNG")
response = client.post("/assets", content=content.getvalue(), headers={"Content-Type": "image/png"})
assert response.status_code == 201
with pytest.raises(ValueError, match="aspect ratios"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
conditioning(response.json(), "reference")
]}, "full-h3")
def test_ref_reports_undecodable_image_as_invalid_input(library, monkeypatch, tmp_path):
store, _ = library
broken = tmp_path / "broken.png"
broken.write_bytes(b"not an image")
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "image", str(broken), "broken.png",
"image/png", 11))
with pytest.raises(ValueError, match="could not be decoded"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
{"asset_id": "a" * 32, "role": "reference"}
]}, "full-h3")
def test_audio_upload_rejects_surround_sound(library, monkeypatch):
import json
_, client = library
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
info = {"format": {"format_name": "wav", "duration": "1"},
"streams": [{"codec_type": "audio", "channels": 6}]}
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
response = client.post("/assets", content=b"surround wav", headers={"Content-Type": "audio/wav"})
assert response.status_code == 400
assert "mono or stereo" in response.json()["detail"]
@pytest.mark.parametrize("mime,format_name", [("audio/x-m4a", "mov,mp4,m4a,3gp,3g2,mj2"), ("audio/x-flac", "flac")])
def test_legacy_audio_mime_aliases_are_accepted(library, monkeypatch, mime, format_name):
"""Browsers report x- variants for the M4A and FLAC formats the docs promise."""
import json
_, client = library
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
info = {"format": {"format_name": format_name, "duration": "1"},
"streams": [{"codec_type": "audio", "channels": 2}]}
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
response = client.post("/assets", content=b"audio bytes", headers={"Content-Type": mime})
assert response.status_code == 201, response.text
assert response.json()["kind"] == "audio"
assert response.json()["mime_type"] == mime
@pytest.mark.parametrize("entries", [None, {}, "x", [{"path": "/etc/passwd", "role": "reference"}],
[{"asset_id": "x", "role": "unknown"}]])
def test_malformed_conditioning_is_rejected(entries):
with pytest.raises(ValueError):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": entries}, "full-h3")
@pytest.mark.parametrize("mode", ["t2va", "fl2va", "ref2va"])
def test_mock_streams_all_valid_modes_and_releases_assets(library, monkeypatch, mode):
from dreamverse import mock_server
_, client = library
monkeypatch.setattr(mock_server, "MOCK_SEGMENT_BYTES", b"mock-fmp4")
monkeypatch.setattr(mock_server, "LATENCY_MS", 1)
image = upload_image(client)
refs = [] if mode == "t2va" else [conditioning(image, "first_frame" if mode == "fl2va" else "reference")]
ws = _FakeWebSocket([
(0, {"type": "session_init_v2", "generation_mode": mode, "conditioning_assets": refs,
"curated_prompts": ["A bird flies over a lake."], "single_clip_mode": True,
"enhancement_enabled": False}),
(0.15, {"type": "leave"}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
assert not [event for event in ws.sent_json if event["type"] == "error"]
assert any(event["type"] == "media_segment_complete" for event in ws.sent_json)
assert ws.sent_bytes
assert client.delete(image["url"]).status_code == 204
def test_mock_rejects_invalid_mode_before_gpu_assignment(library):
from dreamverse import mock_server
ws = _FakeWebSocket([(0, {"type": "session_init_v2", "generation_mode": "fl2va"})])
asyncio.run(mock_server.websocket_endpoint(ws))
assert not any(event["type"] == "gpu_assigned" for event in ws.sent_json)
errors = [event for event in ws.sent_json if event["type"] == "error"]
assert errors[0]["error_code"] == "invalid_generation_input"
assert "first-frame" in errors[0]["message"]
@pytest.mark.skipif(not shutil.which("ffmpeg") or not shutil.which("ffprobe"), reason="ffmpeg + ffprobe required")
@pytest.mark.parametrize("kind,mime,suffix", [("video", "video/mp4", ".mp4"), ("audio", "audio/wav", ".wav")])
def test_actual_video_and_audio_upload_validation(library, tmp_path, kind, mime, suffix):
_, client = library
media_path = tmp_path / f"sample{suffix}"
source = "testsrc2=size=64x64:rate=24" if kind == "video" else "sine=frequency=440:sample_rate=24000"
command = [shutil.which("ffmpeg"), "-v", "error", "-f", "lavfi", "-i", source, "-t", "0.5", str(media_path)]
subprocess.run(command, check=True, capture_output=True, timeout=30)
response = client.post("/assets", content=media_path.read_bytes(), headers={"Content-Type": mime})
assert response.status_code == 201, response.text
assert response.json()["kind"] == kind
response = client.post("/assets", content=b"#EXTM3U\nhttp://example.com/stream", headers={"Content-Type": mime})
assert response.status_code == 400
@@ -1,214 +0,0 @@
"""CPU contract tests; fake executors do not validate generated-media quality."""
from __future__ import annotations
import importlib.util
import pickle
import sys
from pathlib import Path
from types import ModuleType, SimpleNamespace
from unittest.mock import Mock
import numpy as np
import pytest
from PIL import Image
from dreamverse.config import MODEL_REGISTRY
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
from dreamverse.generation_worker import VideoGenerationWorker
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
from dreamverse.worker_ipc import UserStepPayload
@pytest.fixture
def fastvideo_api(monkeypatch):
"""Use the actual lightweight API schema with only GPU execution replaced."""
schema_path = Path(__file__).resolve().parents[4] / "fastvideo/api/schema.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_api_schema", schema_path)
assert spec is not None and spec.loader is not None
schema = importlib.util.module_from_spec(spec)
monkeypatch.setitem(sys.modules, spec.name, schema)
spec.loader.exec_module(schema)
package = ModuleType("fastvideo")
package.__path__ = []
package.VideoGenerator = SimpleNamespace(from_config=Mock())
monkeypatch.setitem(sys.modules, "fastvideo", package)
monkeypatch.setitem(sys.modules, "fastvideo.api", schema)
schema.MiniMaxH3Reference = lambda **kwargs: SimpleNamespace(**kwargs)
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.empty_cache", lambda: None)
return package.VideoGenerator.from_config
class RecordingGenerator:
def __init__(self):
self.requests = []
self.images = []
self.closed = False
def shutdown(self):
self.closed = True
def generate(self, request):
self.requests.append(request)
self.images.append(tuple(None if image is None else np.asarray(image).copy()
for image in (request.inputs.pil_image, request.inputs.last_image)))
return SimpleNamespace(
frames=[np.full((2, 3, 3), 7, dtype=np.uint8), np.full((2, 3, 3), 29, dtype=np.uint8)],
audio=np.zeros((2, 16), dtype=np.float32),
audio_sample_rate=44100,
generation_time=0.1,
)
def prepared_backend(monkeypatch):
backend = MiniMaxH3GenerationBackend(0)
backend.model_config = dict(MODEL_REGISTRY["full-h3"])
backend.generator = RecordingGenerator()
monkeypatch.setattr(backend, "_gpu_mem", lambda: "fake executor")
return backend
def test_ipc_preserves_immutable_ordered_references():
inputs = GenerationInputs("ref2va", (
GenerationAsset("second", "video", "/assets/second.mp4", "reference"),
GenerationAsset("first", "image", "/assets/first.png", "reference"),
))
payload = UserStepPayload("follow the references", 1, None, True, inputs)
restored = pickle.loads(pickle.dumps(payload))
assert restored == payload
assert [asset.asset_id for asset in restored.generation_inputs.references] == ["second", "first"]
def test_worker_passes_conditioning_to_selected_backend():
inputs = GenerationInputs("t2va")
worker = VideoGenerationWorker(0)
worker.backend = Mock()
worker.generate_step("prompt", 1, None, True, inputs)
worker.backend.generate_step.assert_called_once_with("prompt", 1, None, True, generation_inputs=inputs)
def test_full_h3_uses_full_weights_without_preview_lora(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
old_generator = backend.generator
fastvideo_api.return_value = RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
backend.initialize(MODEL_REGISTRY["full-h3"])
config = fastvideo_api.call_args.args[0]
assert old_generator.closed
assert config.pipeline.components.lora_path is None
assert config.pipeline.components.override_pipeline_cls_name is None
assert config.engine.use_fsdp_inference
assert config.engine.num_gpus == 4
assert not config.pipeline.experimental["inference_torch_compile"]
def test_fl2va_maps_endpoints_only_on_initial_segment(monkeypatch, fastvideo_api, tmp_path):
first = tmp_path / "first.png"
last = tmp_path / "last.png"
Image.new("RGB", (3, 2), (10, 20, 30)).save(first)
Image.new("RGB", (3, 2), (40, 50, 60)).save(last)
inputs = GenerationInputs("fl2va", (
GenerationAsset("first", "image", str(first), "first_frame"),
GenerationAsset("last", "image", str(last), "last_frame"),
))
backend = prepared_backend(monkeypatch)
first_result = backend.generate_step("first", 1, None, True, inputs)
later_result = backend.generate_step("later", 2, None, False, inputs)
assert backend.generator.images[0][0][0, 0].tolist() == [10, 20, 30]
assert backend.generator.images[0][1][0, 0].tolist() == [40, 50, 60]
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
assert backend.generator.images[1][1] is None
assert first_result.head_trim_frames == 0
assert later_result.head_trim_frames == 1
assert backend.generator.requests[0].sampling.num_inference_steps == 50
def test_ref2va_switches_pipeline_and_preserves_reference_order(monkeypatch, fastvideo_api):
inputs = GenerationInputs("ref2va", (
GenerationAsset("video", "video", "/assets/reference.mp4", "reference"),
GenerationAsset("audio", "audio", "/assets/reference.wav", "reference"),
GenerationAsset("image", "image", "/assets/reference.png", "reference"),
))
backend = prepared_backend(monkeypatch)
base_generator = backend.generator
reference_generator = RecordingGenerator()
def load(config):
assert base_generator.closed, "Old executor must release memory before loading reference weights"
assert config.pipeline.components.override_pipeline_cls_name == "MiniMaxH3Ref2VAModularPipeline"
assert config.pipeline.workload_type == "i2v"
assert config.pipeline.components.lora_path is None
return reference_generator
fastvideo_api.side_effect = load
backend.generate_step("first", 1, None, True, inputs)
result = backend.generate_step("second", 2, None, False, inputs)
assert fastvideo_api.call_count == 1
for request in reference_generator.requests:
assert [(reference.media_type, reference.source) for reference in request.inputs.references] == [
("video", "/assets/reference.mp4"), ("audio", "/assets/reference.wav"), ("image", "/assets/reference.png")
]
assert request.inputs.pil_image is None
assert request.inputs.last_image is None
assert result.head_trim_frames == result.head_trim_audio_frames == 0
assert backend.continuation_image is None
fastvideo_api.side_effect = None
fastvideo_api.return_value = RecordingGenerator()
backend.generate_step("new project", 1, None, True, GenerationInputs("t2va"))
assert reference_generator.closed
config = fastvideo_api.call_args.args[0]
assert config.pipeline.components.override_pipeline_cls_name is None
assert backend.pipeline_mode == "base"
def test_ref2va_pipeline_switch_failure_drops_unloaded_executor(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
old_generator = backend.generator
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
assert old_generator.closed
assert backend.generator is None
def test_failed_pipeline_switch_reloads_on_the_next_step(monkeypatch, fastvideo_api):
"""A failed base<->ref2va switch must not strand the slot for later steps."""
backend = prepared_backend(monkeypatch)
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
fastvideo_api.side_effect = None
fastvideo_api.return_value = RecordingGenerator()
backend.generate_step("retry", 1, None, True, GenerationInputs("ref2va"))
assert fastvideo_api.call_count == 2
assert backend.pipeline_mode == "ref2va"
assert backend.generator is not None
def test_mode_cannot_switch_mid_project(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
with pytest.raises(ValueError, match="middle of a project"):
backend.generate_step("prompt", 2, None, False, GenerationInputs("ref2va"))
fastvideo_api.assert_not_called()
@pytest.mark.parametrize("mode", ["fl2va", "ref2va"])
def test_preview_rejects_unsupported_generation_modes(monkeypatch, fastvideo_api, mode):
backend = prepared_backend(monkeypatch)
backend.model_config = dict(MODEL_REGISTRY["fast-h3"])
with pytest.raises(ValueError, match="full-h3"):
backend.generate_step("prompt", 1, None, True, GenerationInputs(mode))
assert backend.generator.requests == []
def test_legacy_h3_continuation_is_preserved(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
backend.generate_step("first", 1, None, False)
result = backend.generate_step("second", 2, None, False)
assert backend.generator.requests[0].inputs.pil_image is None
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
assert result.head_trim_frames == 1
@@ -1,278 +0,0 @@
"""Session-mode validation and IPC handoff without a GPU worker process."""
from __future__ import annotations
import asyncio
import importlib.util
import sys
from pathlib import Path
from types import ModuleType, SimpleNamespace
from unittest.mock import Mock
import pytest
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
@pytest.fixture
def controller_module(monkeypatch):
gpu_pool = ModuleType("dreamverse.gpu_pool")
gpu_pool.GPUSlot = object
monkeypatch.setitem(sys.modules, "dreamverse.gpu_pool", gpu_pool)
path = Path(__file__).resolve().parents[1] / "session/controller.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_session_controller", path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
monkeypatch.setattr(module, "ACTIVE_MODEL_ID", "full-h3")
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
monkeypatch.setattr(module, "release_generation_inputs", Mock())
return module
class Socket:
def __init__(self):
self.incoming = asyncio.Queue()
self.outgoing = asyncio.Queue()
self.messages = []
self.closed = False
async def accept(self):
pass
async def receive_json(self):
return await self.incoming.get()
async def send_json(self, payload):
self.messages.append(payload)
await self.outgoing.put(payload)
async def send_bytes(self, payload):
pass
async def close(self, **kwargs):
self.closed = True
async def wait_for(self, kind):
while True:
payload = await asyncio.wait_for(self.outgoing.get(), 3)
if payload["type"] == kind:
return payload
class Slot:
def __init__(self):
self.shared_stream_buffer = None
self.queue = asyncio.Queue()
self.calls = []
async def join_user(self, *args, **kwargs):
pass
def register_stream_queue(self, client_id):
return self.queue
def unregister_stream_queue(self, client_id):
pass
async def user_step(self, client_id, **kwargs):
self.calls.append(kwargs)
segment_idx = kwargs["segment_idx"]
await self.queue.put(MediaInit(client_id, segment_idx, "test", "video/mp4", False))
await self.queue.put(MediaChunk(client_id, segment_idx, "test", chunk=b"test"))
await self.queue.put(MediaComplete(client_id, segment_idx, "test", 1))
return {"e2e_latency_ms": 1.0}
class Pool:
def __init__(self):
self.slot = Slot()
self.acquire_count = 0
def get_status(self):
return {"queue_size": 0, "available_gpus": 1, "total_gpus": 1}
async def acquire(self, *args):
self.acquire_count += 1
return 0, self.slot
async def release(self, *args):
pass
def start_controller(module, socket, pool):
enhancer = SimpleNamespace(
resolve_rewrite_model=lambda value: "test-model",
resolve_rewrite_system_prompt=lambda value: "test-system",
resolve_rewrite_temperature=lambda value: 1.0,
)
controller = module.SessionController(socket, pool, enhancer, None, None)
return asyncio.create_task(controller.run())
def test_invalid_initial_mode_does_not_acquire_gpu(controller_module):
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({"type": "session_init_v2", "generation_mode": "unknown"})
await asyncio.wait_for(start_controller(controller_module, socket, pool), 3)
error = next(message for message in socket.messages if message["type"] == "error")
assert error["error_code"] == "invalid_generation_input"
assert pool.acquire_count == 0
assert socket.closed
controller_module.pin_generation_inputs.assert_not_called()
asyncio.run(scenario())
def test_new_project_replaces_conditioning_and_passes_it_to_gpu(controller_module, monkeypatch):
first = GenerationInputs("t2va")
second = GenerationInputs("fl2va", (GenerationAsset("first", "image", "/assets/first.png", "first_frame"),))
monkeypatch.setattr(controller_module, "resolve_generation_inputs", Mock(side_effect=[first, second]))
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["first prompt"],
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await socket.wait_for("media_segment_complete")
await socket.incoming.put({"type": "end_project_keep_session"})
await socket.wait_for("project_idle")
assert first in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
await socket.incoming.put({
"type": "project_init_v1", "generation_mode": "fl2va", "curated_prompts": ["second prompt"],
"enhancement_enabled": False,
})
await socket.wait_for("media_segment_complete")
assert [call["generation_inputs"] for call in pool.slot.calls] == [first, second]
assert pool.slot.calls[1]["segment_idx"] == 1
assert pool.slot.calls[1]["reset_conditioning"]
await socket.incoming.put({"type": "leave"})
await asyncio.wait_for(task, 3)
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
assert [call.args[0] for call in controller_module.pin_generation_inputs.call_args_list] == [first, second]
assert second in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
asyncio.run(scenario())
@pytest.mark.parametrize("injection", [
{"initial_image": {"data_url": "not allowed"}},
{"generation_mode": "ref2va"},
{"conditioning_assets": []},
])
def test_simple_generate_cannot_replace_locked_inputs(controller_module, injection):
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "single_clip_mode": True,
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await socket.wait_for("gpu_assigned")
await socket.incoming.put({"type": "simple_generate", "prompt": "prompt", **injection})
error = await socket.wait_for("error")
assert error["error_code"] == "invalid_generation_input"
assert "project" in error["message"]
assert pool.slot.calls == []
await socket.incoming.put({"type": "leave"})
await asyncio.wait_for(task, 3)
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
asyncio.run(scenario())
def test_disconnect_waits_for_worker_before_releasing_assets(controller_module, monkeypatch):
async def scenario():
socket, pool = Socket(), Pool()
worker_started = asyncio.Event()
worker_finished = asyncio.Event()
proceed = asyncio.Event()
async def slow_step(client_id, **kwargs):
worker_started.set()
await proceed.wait()
worker_finished.set()
return {"e2e_latency_ms": 1.0}
pool.slot.user_step = slow_step
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["prompt"],
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await asyncio.wait_for(worker_started.wait(), 3)
await socket.incoming.put({"type": "leave"})
await asyncio.sleep(0.07)
assert not task.done()
controller_module.release_generation_inputs.assert_not_called()
proceed.set()
await asyncio.wait_for(task, 3)
assert worker_finished.is_set()
controller_module.release_generation_inputs.assert_called_once()
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
asyncio.run(scenario())
@pytest.fixture
def gpu_pool_module(monkeypatch):
streaming = ModuleType("dreamverse.av_streaming")
for name in ("StreamChunk", "StreamComplete", "StreamEvent", "StreamInit", "generate_stream_id", "stream_fmp4"):
setattr(streaming, name, object)
streaming.SHARED_STREAM_BUFFER_BYTES = 1024
streaming.USE_SHARED_STREAM_BUFFER = False
monkeypatch.setitem(sys.modules, "dreamverse.av_streaming", streaming)
path = Path(__file__).resolve().parents[1] / "gpu_pool.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_gpu_pool", path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
monkeypatch.setitem(sys.modules, spec.name, module)
spec.loader.exec_module(module)
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
monkeypatch.setattr(module, "release_generation_inputs", Mock())
return module
def test_gpu_step_timeout_keeps_assets_pinned_until_late_worker_completion(gpu_pool_module):
from dreamverse.worker_ipc import StepComplete
async def scenario():
slot = gpu_pool_module.GPUSlot(0, "0")
inputs = GenerationInputs("ref2va", (GenerationAsset("ref", "image", "/assets/ref.png", "reference"),))
async def timeout(command, timeout):
assert command.payload.generation_inputs == inputs
raise asyncio.TimeoutError
slot._send_command_tagged = timeout
with pytest.raises(asyncio.TimeoutError):
await slot.user_step("user", "prompt", generation_inputs=inputs)
gpu_pool_module.pin_generation_inputs.assert_called_once_with(inputs)
gpu_pool_module.release_generation_inputs.assert_not_called()
def late_response(timeout):
slot._active = False
return StepComplete("user", 1, {})
slot.response_queue = SimpleNamespace(get=late_response)
slot._active = True
await slot._response_reader()
gpu_pool_module.release_generation_inputs.assert_called_once_with(inputs)
assert slot._step_asset_inputs == {}
asyncio.run(scenario())
@@ -7,12 +7,3 @@ def test_create_generation_backend_ltx2_module_import():
assert isinstance(backend, LTX2GenerationBackend)
assert backend.gpu_id == 3
def test_create_generation_backend_cosmos25_dfd_module_import():
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
backend = _create_generation_backend("cosmos25_dfd", gpu_id=2)
assert isinstance(backend, Cosmos25DFDGenerationBackend)
assert backend.gpu_id == 2
@@ -1,6 +1,4 @@
import ast
import subprocess
import sys
from pathlib import Path
ALLOWED_PREFIXES = (
@@ -50,39 +48,3 @@ def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
assert bad == [], f"Forbidden internal imports: {bad}"
def test_h3_reference_public_export_is_lazy_and_preserves_type_identity() -> None:
"""Only explicit reference usage should load H3's optional GPU dependencies."""
repo_root = Path(__file__).resolve().parents[4]
# Isolate the import graph: keep the real public API implementation/schema,
# substituting only the unrelated legacy sampling module and heavy H3 leaf.
script = r'''
import importlib
import sys
from pathlib import Path
from types import ModuleType
root = Path(sys.argv[1])
fastvideo = ModuleType("fastvideo")
fastvideo.__path__ = [str(root / "fastvideo")]
sys.modules["fastvideo"] = fastvideo
sampling = ModuleType("fastvideo.api.sampling_param")
sampling.SamplingParam = type("SamplingParam", (), {})
sys.modules[sampling.__name__] = sampling
api = importlib.import_module("fastvideo.api")
assert "MiniMaxH3Reference" in api.__all__
assert "MiniMaxH3Reference" not in vars(api)
assert not any(name.startswith("fastvideo.pipelines") for name in sys.modules)
internal = ModuleType("fastvideo.pipelines.basic.minimax_h3.reference")
internal.MiniMaxH3Reference = type("MiniMaxH3Reference", (), {})
sys.modules[internal.__name__] = internal
from fastvideo.api import MiniMaxH3Reference
assert MiniMaxH3Reference is internal.MiniMaxH3Reference
assert api.MiniMaxH3Reference is internal.MiniMaxH3Reference
assert not hasattr(api, "UnknownReference")
'''
result = subprocess.run([sys.executable, "-c", script, str(repo_root)], capture_output=True, text=True, timeout=30)
assert result.returncode == 0, result.stderr
@@ -0,0 +1,89 @@
import pytest
from dreamverse.session_creation_config import (
duration_sec_to_segment_cap,
parse_session_creation_config,
resolve_frame_size,
validate_generation_mode_assets,
)
def test_parse_session_creation_config_defaults():
config = parse_session_creation_config({})
assert config.model_id == "fast-ltx23"
assert config.generation_mode == "t2va"
assert config.aspect_ratio == "16:9"
assert config.resolution == "720p"
assert config.duration_sec == 5
assert config.generation_segment_cap == 1
def test_parse_session_creation_config_maps_duration_to_segment_cap():
config = parse_session_creation_config(
{
"model_id": "fast-ltx2",
"generation_mode": "ref2va",
"aspect_ratio": "9:16",
"resolution": "480p",
"duration_sec": 15,
},
)
assert config.model_id == "fast-ltx2"
assert config.generation_mode == "ref2va"
assert config.generation_segment_cap == 3
assert config.frame_width >= 480
assert config.frame_height >= 480
def test_resolve_frame_size_uses_model_default_for_1080p_landscape():
width, height = resolve_frame_size("16:9", "1080p")
assert (width, height) == (1920, 1088)
def test_duration_sec_to_segment_cap_respects_global_cap():
assert duration_sec_to_segment_cap(15, global_cap=2) == 2
def test_parse_session_creation_config_accepts_fast_h3():
config = parse_session_creation_config(
{
"model_id": "fast-h3",
"generation_mode": "t2va",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration_sec": 10,
},
)
assert config.model_id == "fast-h3"
assert config.generation_mode == "t2va"
assert config.generation_segment_cap == 2
def test_parse_session_creation_config_rejects_fl2va():
with pytest.raises(ValueError, match="FL2VA"):
parse_session_creation_config(
{
"generation_mode": "fl2va",
"aspect_ratio": "16:9",
"resolution": "720p",
"duration_sec": 5,
},
)
def test_parse_session_creation_config_rejects_4k():
with pytest.raises(ValueError, match="Unsupported resolution"):
parse_session_creation_config(
{
"generation_mode": "t2va",
"aspect_ratio": "16:9",
"resolution": "4k",
"duration_sec": 5,
},
)
def test_validate_generation_mode_assets():
validate_generation_mode_assets("t2va", has_initial_image=False, has_last_frame_image=False)
with pytest.raises(ValueError, match="Ref2VA"):
validate_generation_mode_assets("ref2va", has_initial_image=False, has_last_frame_image=False)
@@ -105,7 +105,6 @@ class _FakeSlot:
segment_idx: int,
reset_conditioning: bool,
image_path: str | None = None,
generation_inputs=None,
):
self.calls.append({
"client_id": client_id,
+3 -3
View File
@@ -15,8 +15,6 @@ from __future__ import annotations
from dataclasses import dataclass
from dreamverse.generation_inputs import GenerationInputs
# ---- User-scoped events (carry user_id) ------------------------------------
@@ -149,7 +147,9 @@ class UserStepPayload:
segment_idx: int
image_path: str | None
reset_conditioning: bool
generation_inputs: GenerationInputs | None = None
frame_width: int | None = None
frame_height: int | None = None
num_frames: int | None = None
@dataclass(frozen=True)
-107
View File
@@ -1,107 +0,0 @@
# Dreamverse on Slurm
Run Full H3 inside a one-node, four-GPU allocation. The maintained H3 examples
default to four GPUs; this is a starting configuration, not a measured minimum.
The full checkpoint supports T2VA, FL2VA, and Ref2VA. The FastH3 Preview profile
is a separate T2VA configuration.
`launch_backend.sh` checks that it is inside an `srun` step, preserves
`CUDA_VISIBLE_DEVICES`, and replaces itself with the backend process. It does
not allocate GPUs, kill existing processes, or source a personal credentials
file. The local `dreamverse-deploy` helper is not suitable for a shared Slurm
cluster because it kills processes by physical GPU and port.
## Prepare and allocate
Keep the checkout, weights, outputs, and logs on storage visible to the compute
node. Source installation is documented in the [GPU guide](../../../../docs/getting_started/installation/gpu.md).
On ARM64 GB200 use CUDA 13, a matching PyTorch build, and kernels built for
`sm_100`; the DGX Spark `sm_121` kernel image is not the GB200 image.
The repository's image workflow publishes an ARM64 GB200 variant under
`ghcr.io/hao-ai-lab/fastvideo/fastvideo-dev:py3.12-cuda13.0.0-sm100-latest`.
Resolve that tag to a digest for reproducible runs. If your compute nodes use
Pyxis/Enroot, pass the approved image or a prepared SquashFS file to
`srun --container-image`, with explicit mounts for your checkout and model cache.
The Dreamverse-specific Docker images are currently AMD64-only.
For the Slinky customer partition, a bounded allocation is:
```bash
salloc --account=customer --qos=normal --partition=hpc-rack-1 \
--nodes=1 --ntasks=1 --cpus-per-task=72 --gres=gpu:nvidia_gb200:4 \
--mem=800G --time=02:00:00 --job-name=dreamverse
srun --ntasks=1 --pty bash
```
Wait for Slurm to grant the allocation before entering the compute step. A
successful SSH login does not grant GPU resources. Inspect pending capacity
with `squeue -u "$USER" --start`; do not attach to another user's job.
The checkpoint includes duplicate release layouts. Download the diffusers
components needed by both base and reference pipelines, rather than the whole
repository (about 210 GB versus about 498 GB at revision
`42ed227ee7df40d41602854ae760620d6eb651fe`):
```bash
hf download MiniMaxAI/MiniMax-H3 \
--revision 42ed227ee7df40d41602854ae760620d6eb651fe \
--include model_index.json --include modular_model_index.json \
--include 'audio_scheduler/*' --include 'audio_vae/*' \
--include 'processor/*' --include 'scheduler/*' \
--include 'text_encoder/*' --include 'tokenizer/*' \
--include 'transformer/*' --include 'transformer_ref/*' --include 'vae/*' \
--local-dir /path/to/models/MiniMax-H3
```
The GPU environment needs `fastvideo[dreamverse]`, the Dreamverse workspace
package, and FFmpeg with H.264/AAC encoders. In a prepared FastVideo image,
install the checked-out code and its Dreamverse dependencies in that image's
Python environment. Keep its matching CUDA/PyTorch/kernel stack intact.
## Start and connect
From the checked-out repository inside the allocated step:
```bash
export DREAMVERSE_PYTHON=/path/to/environment/bin/python
export DREAMVERSE_MODEL_PATH=/path/to/models/MiniMax-H3
export FASTVIDEO_DREAMVERSE_HOME=/path/to/persistent/dreamverse-state
bash apps/dreamverse/scripts/slurm/launch_backend.sh
```
The default backend binds port 8009 on the private compute node. Connect through
the login node from your laptop, replacing `COMPUTE_NODE_IP` with the allocated
node's `NodeAddr` from `scontrol show node`:
```bash
ssh -N -L 8009:COMPUTE_NODE_IP:8009 USER@LOGIN_NODE
```
In another laptop terminal, run the frontend from your local checkout:
```bash
cd apps/dreamverse/web
BACKEND_HOST=127.0.0.1 BACKEND_PORT=8009 npm run dev
```
Open `http://localhost:5299`. `/healthz` reports the server process; `/readyz`
reports model readiness. Full H3 loads and generates more slowly than the
Preview adapter. Keep prompt enhancement disabled in the UI unless the
runtime has the selected provider's credentials.
## Verify and stop
Check all three modes with small, valid user-owned assets. Capture the selected
mode and assets, WebSocket errors or completion events, the generated video and
audio, and GPU memory usage. Also verify actionable validation errors and
backward compatibility with clients that omit `generation_mode`.
Use the frontend Playwright instructions in the
[Dreamverse development guide](../../../../docs/contributing/dreamverse-development.md)
against the forwarded backend. A mock-server demo validates UI and protocol
behavior; it is not evidence of GPU generation.
Stop the backend with Ctrl-C, exit the compute step, and release your allocation.
For a detached allocation, use `scancel YOUR_JOB_ID`. Cancel a pending demo job
when it is no longer needed; do not leave an unattended reservation queued.
@@ -1,49 +0,0 @@
#!/usr/bin/env bash
# Run inside an existing Slurm step. Slurm owns the GPU visibility and lifetime.
set -euo pipefail
if [[ -z "${SLURM_JOB_ID:-}" || -z "${SLURM_STEP_ID:-}" ]]; then
echo "Run this launcher inside an allocated Slurm step (srun), not on the login node." >&2
exit 2
fi
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
repo_root="$(cd -- "${script_dir}/../../../.." && pwd)"
python_bin="${DREAMVERSE_PYTHON:-${repo_root}/.venv/bin/python}"
if [[ ! -x "${python_bin}" ]]; then
echo "Set DREAMVERSE_PYTHON to a Python environment with fastvideo[dreamverse] installed." >&2
exit 2
fi
export DREAMVERSE_MODEL_ID="${DREAMVERSE_MODEL_ID:-full-h3}"
export DREAMVERSE_SP_SIZE="${DREAMVERSE_SP_SIZE:-4}"
export FASTVIDEO_GPU_COUNT="${FASTVIDEO_GPU_COUNT:-${DREAMVERSE_SP_SIZE}}"
export FASTVIDEO_ENABLE_STARTUP_WARMUP="${FASTVIDEO_ENABLE_STARTUP_WARMUP:-0}"
export ENABLE_TORCH_COMPILE="${ENABLE_TORCH_COMPILE:-0}"
export STREAM_MODE="${STREAM_MODE:-av_fmp4}"
export PYTHONPATH="${repo_root}/apps/dreamverse:${repo_root}${PYTHONPATH:+:${PYTHONPATH}}"
export PYTHONUNBUFFERED=1
"${python_bin}" - <<'PY'
import os
import shutil
import torch
expected = int(os.environ["DREAMVERSE_SP_SIZE"])
visible = torch.cuda.device_count()
if expected < 1 or visible < expected:
raise SystemExit(f"The Slurm step exposes {visible} GPUs; DREAMVERSE_SP_SIZE requires {expected}.")
ffmpeg = os.environ.get("FASTVIDEO_FFMPEG_BIN", "ffmpeg")
if not shutil.which(ffmpeg):
raise SystemExit("FFmpeg is missing; install it in the compute environment or set FASTVIDEO_FFMPEG_BIN.")
print(f"Slurm job {os.environ['SLURM_JOB_ID']}: {visible} visible GPUs; using {expected} per worker")
for index in range(expected):
properties = torch.cuda.get_device_properties(index)
print(f" GPU {index}: {properties.name}, {properties.total_memory / 2**30:.1f} GiB")
PY
cd "${repo_root}"
exec "${python_bin}" -m dreamverse.server_entry \
--host "${DREAMVERSE_BIND_HOST:-0.0.0.0}" \
--port "${DREAMVERSE_BACKEND_PORT:-8009}" "$@"
+4 -12
View File
@@ -17,25 +17,17 @@ test.describe('frontend shell', () => {
});
});
test('composer hydrates with curated preset cards', async ({ page }) => {
test('composer hydrates with creation studio controls', async ({ page }) => {
await page.goto('/');
// The Continuation prompt textarea + Generate button render once
// the FE has hydrated against the public-FastVideo-backed
// dreamverse-server. Their presence proves the integration handshake
// (CORS, /curated-presets, /prompt-system-config) completed.
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible({ timeout: 30_000 });
const generate = page.getByRole('button', { name: /^generate$/i });
await expect(generate).toBeVisible({ timeout: 30_000 });
// Curated presets render as buttons; verify at least one is
// available — that's the only way the user can populate the
// Continuation textarea in the default composer.
const presetCard = page.getByRole('button', {
name: /LEGO Stormtroopers|Clay Stop-Motion|Boy & Dog|School Prank|Gamer Gets Banned|Small Town Oil Strike|Grandpa's Wing Costume/i,
}).first();
await expect(presetCard).toBeVisible({ timeout: 30_000 });
await expect(page.getByText('Direct scenes in seconds')).toBeVisible({ timeout: 30_000 });
await expect(page.getByRole('button', { name: /FastLTX/i }).first()).toBeVisible({ timeout: 30_000 });
await expect(continuation).toHaveAttribute('placeholder', /Describe your video or mention elements/i);
});
});
@@ -1,126 +0,0 @@
import { execFileSync } from "node:child_process";
import { readFile } from "node:fs/promises";
import path from "node:path";
import { test, expect } from "@playwright/test";
const imagePath = path.resolve("public/k2.png");
const framePrompt = "A paper fox walks through a sunlit forest, gentle birdsong.";
function makeAudio(sampleRate = 8000, seconds = 1): Buffer {
const sampleCount = sampleRate * seconds;
const bytes = Buffer.alloc(44 + sampleCount * 2);
bytes.write("RIFF", 0); bytes.writeUInt32LE(bytes.length - 8, 4); bytes.write("WAVEfmt ", 8);
bytes.writeUInt32LE(16, 16); bytes.writeUInt16LE(1, 20); bytes.writeUInt16LE(1, 22);
bytes.writeUInt32LE(sampleRate, 24); bytes.writeUInt32LE(sampleRate * 2, 28);
bytes.writeUInt16LE(2, 32); bytes.writeUInt16LE(16, 34); bytes.write("data", 36);
bytes.writeUInt32LE(sampleCount * 2, 40);
for (let i = 0; i < sampleCount; i++) bytes.writeInt16LE(Math.round(Math.sin(i * 440 * 2 * Math.PI / sampleRate) * 1000), 44 + i * 2);
return bytes;
}
test.describe("generation modes through the mock runtime", () => {
for (const mode of ["t2va", "fl2va", "ref2va"] as const) {
test(`${mode} sends validated assets and plays a clearly labeled sample`, async ({ page, request }, testInfo) => {
const response = await request.get("/generation-capabilities");
const capabilities = response.ok() ? await response.json() : {};
test.skip(capabilities.mock !== true, "This test uses the CPU mock runtime; it must not silently allocate a real GPU.");
const sent: Record<string, any>[] = [];
const received: Record<string, any>[] = [];
page.on("websocket", (socket) => {
socket.on("framesent", ({ payload }) => { if (typeof payload === "string") { try { sent.push(JSON.parse(payload)); } catch {} } });
socket.on("framereceived", ({ payload }) => { if (typeof payload === "string") { try { received.push(JSON.parse(payload)); } catch {} } });
});
await page.goto("/");
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
const modeSelect = page.getByRole("combobox", { name: "Generation mode" });
const modeLabel = mode === "ref2va" ? "Ref2VA" : mode.toUpperCase();
await modeSelect.click();
await page.getByRole("option", { name: modeLabel, exact: true }).click();
await expect(modeSelect).toHaveText(modeLabel);
await page.getByLabel("Continuation prompt").fill(framePrompt);
const uploadedIds: string[] = [];
page.on("response", async (uploadResponse) => {
if (uploadResponse.request().method() === "POST" && uploadResponse.url().endsWith("/assets") && uploadResponse.ok()) {
const asset = await uploadResponse.json().catch(() => null);
if (asset?.asset_id) uploadedIds.push(asset.asset_id);
}
});
try {
if (mode === "fl2va") {
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
await page.locator('input[type="file"]').setInputFiles([
{ name: "first-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
{ name: "last-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
]);
await expect(page.getByRole("option", { name: "first-frame.png", exact: true }).first()).toBeAttached();
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "first-frame.png" });
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeEnabled();
await page.getByRole("combobox", { name: "Last frame", exact: true }).selectOption({ label: "last-frame.png" });
}
if (mode === "ref2va") {
const video = execFileSync(process.env.FASTVIDEO_FFMPEG_BIN || "ffmpeg", ["-v", "error", "-f", "lavfi", "-i", "color=c=royalblue:s=64x64:r=8", "-t", "1", "-c:v", "libx264", "-pix_fmt", "yuv420p", "-movflags", "frag_keyframe+empty_moov", "-f", "mp4", "pipe:1"]);
await page.locator('input[type="file"]').setInputFiles([
{ name: "subject.png", mimeType: "image/png", buffer: await readFile(imagePath) },
{ name: "motion.mp4", mimeType: "video/mp4", buffer: video },
{ name: "sound.wav", mimeType: "audio/wav", buffer: makeAudio() },
]);
await expect(page.getByRole("button", { name: "Add sound.wav as reference" })).toBeEnabled();
await page.getByRole("button", { name: "Add sound.wav as reference" }).click();
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
await page.getByRole("button", { name: "Add subject.png as reference" }).click();
await page.getByRole("button", { name: "Add motion.mp4 as reference" }).click();
await page.getByRole("button", { name: "Move sound.wav down" }).click();
const names = await page.getByRole("list", { name: "Ordered references" }).locator("li p.font-medium").allTextContents();
expect(names).toEqual(["subject.png", "sound.wav", "motion.mp4"]);
}
await page.screenshot({ path: testInfo.outputPath(`${mode}-inputs.png`), fullPage: true });
await page.getByRole("button", { name: "Generate", exact: true }).click();
await expect.poll(() => sent.find((item) => item.type === "session_init_v2")?.generation_mode).toBe(mode);
const init = sent.find((item) => item.type === "session_init_v2")!;
expect(init.conditioning_assets.map((item: any) => item.role)).toEqual(mode === "t2va" ? [] : mode === "fl2va" ? ["first_frame", "last_frame"] : ["reference", "reference", "reference"]);
if (mode === "ref2va") expect(init.conditioning_assets.map((item: any) => item.asset_id)).toEqual([uploadedIds[0], uploadedIds[2], uploadedIds[1]]);
await expect.poll(() => received.find((item) => item.type === "gpu_assigned")?.generation_mode).toBe(mode);
await expect.poll(() => received.some((item) => item.type === "media_segment_complete")).toBe(true);
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
await expect(modeSelect).toHaveCount(0);
await expect.poll(async () => page.locator("video:visible").first().evaluate((element: HTMLVideoElement) => element.readyState)).toBeGreaterThanOrEqual(2);
await page.screenshot({ path: testInfo.outputPath(`${mode}-playback.png`), fullPage: true });
if (mode === "ref2va") {
await page.getByRole("button", { name: "Toggle sidebar" }).click();
await page.getByRole("button", { name: "New project", exact: true }).click();
await expect(modeSelect).toHaveText("T2VA");
await modeSelect.click();
await page.getByRole("option", { name: "FL2VA", exact: true }).click();
await expect(modeSelect).toHaveText("FL2VA");
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "subject.png" });
await expect(page.getByRole("combobox", { name: "Last frame", exact: true })).toHaveValue("");
await page.getByLabel("Continuation prompt").fill("The paper fox explores a new scene.");
await page.getByRole("button", { name: "Generate", exact: true }).click();
await expect.poll(() => sent.find((item) => item.type === "project_init_v1")?.generation_mode).toBe("fl2va");
const secondProject = sent.find((item) => item.type === "project_init_v1")!;
expect(secondProject.conditioning_assets).toEqual([{ asset_id: uploadedIds[0], role: "first_frame" }]);
expect(sent.filter((item) => item.type === "session_init_v2")).toHaveLength(1);
await expect.poll(() => received.filter((item) => item.type === "media_segment_complete").length).toBeGreaterThan(1);
}
} finally {
await page.close();
for (const id of uploadedIds) await request.delete(`/assets/${id}`);
}
});
}
test("proxies a media upload larger than Next's default 10 MiB body limit", async ({ request }) => {
const response = await request.get("/generation-capabilities");
const capabilities = response.ok() ? await response.json() : {};
test.skip(capabilities.mock !== true, "Requires the local mock runtime.");
const audio = makeAudio(192000, 29);
expect(audio.length).toBeGreaterThan(10 * 1024 * 1024);
const upload = await request.post("/assets", {
headers: { "Content-Type": "audio/wav", "X-Asset-Name": "large-proxy-check.wav" },
data: audio,
});
expect(upload.status()).toBe(201);
const asset = await upload.json();
try { expect(asset.size).toBe(audio.length); } finally { await request.delete(`/assets/${asset.asset_id}`); }
});
});
+4 -14
View File
@@ -9,8 +9,6 @@ const configDir = path.dirname(fileURLToPath(import.meta.url));
const staticExport = process.env.NEXT_OUTPUT_EXPORT === '1';
const nextConfig: NextConfig = {
// Next 15.5 name for the dev rewrite-proxy body limit; Next 16 renames it to `proxyClientMaxBodySize`.
experimental: { middlewareClientMaxBodySize: 100 * 1024 * 1024 },
...(staticExport ? { output: 'export' as const } : {}),
...(staticExport ? { images: { unoptimized: true } } : {}),
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
@@ -40,22 +38,14 @@ const nextConfig: NextConfig = {
source: '/router/:path*',
destination: `${backendUrl}/router/:path*`
},
{
source: '/generation-capabilities',
destination: `${backendUrl}/generation-capabilities`,
},
{
source: '/assets',
destination: `${backendUrl}/assets`,
},
{
source: '/assets/:path*',
destination: `${backendUrl}/assets/:path*`,
},
{
source: '/prompt-system-config',
destination: `${backendUrl}/prompt-system-config`,
},
{
source: '/creation-capabilities',
destination: `${backendUrl}/creation-capabilities`,
},
{
source: '/curated-presets',
destination: `${backendUrl}/curated-presets`,
+2262 -113
View File
File diff suppressed because it is too large Load Diff
+4
View File
@@ -27,11 +27,15 @@
"@radix-ui/react-accordion": "^1.2.12",
"@radix-ui/react-checkbox": "^1.3.3",
"@radix-ui/react-collapsible": "^1.1.12",
"@radix-ui/react-dropdown-menu": "^2.1.24",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-popover": "^1.1.23",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.4.7",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-tabs": "^1.1.21",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"framer-motion": "^12.36.0",
+26 -1
View File
@@ -2,6 +2,7 @@
@import "tailwindcss";
@custom-variant dark (&:is(.dark *));
@custom-variant hover-capable (@media (hover: hover) and (pointer: fine));
:root {
color-scheme: light;
@@ -201,7 +202,7 @@ summary::-webkit-details-marker {
@apply border-border;
}
body {
@apply bg-background text-foreground;
@apply bg-background text-foreground antialiased;
}
}
@@ -220,6 +221,14 @@ summary::-webkit-details-marker {
.stroke-dash-anim {
animation: stroke-dash-animation 2s linear infinite;
}
.text-pretty {
text-wrap: pretty;
}
.text-balance {
text-wrap: balance;
}
}
@keyframes stroke-dash-animation {
@@ -287,6 +296,22 @@ html.theme-transition *::after {
/* —————————————— CUSTOM TAILWIND —————————————— */
@layer components {
.studio-control {
@apply transition-[border-color,background-color,box-shadow,color,transform] duration-150 focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40;
}
.studio-control-press {
@apply active:scale-[0.96] motion-reduce:active:scale-100;
}
.studio-hover-surface {
@apply hover-capable:hover:border-border hover-capable:hover:bg-accent/50;
}
.studio-media-outline {
@apply outline outline-1 -outline-offset-1 outline-black/10 dark:outline-white/10;
}
.debug {
@apply border border-rose-500;
}
@@ -589,7 +589,6 @@ describe.skip('App websocket integration', () => {
});
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
expect(initMessage.generation_mode).toBe('t2va');
expect(initMessage.preset_id).toBe('test_preset');
expect(initMessage.curated_prompts).toEqual(['segment one', 'segment two']);
expect(initMessage.enhancement_enabled).toBe(true);
@@ -598,42 +597,6 @@ describe.skip('App websocket integration', () => {
expect(initMessage.initial_rollout_prompt).toBe('');
});
it('sends the selected generation mode and locks it after session start', async () => {
const outbound: any[] = [];
server.on('connection', (socket) => {
socket.on('message', (rawMessage) => {
outbound.push(JSON.parse(rawMessage as string));
});
});
const user = userEvent.setup();
render(<Page />);
const modeSelect = await screen.findByRole('combobox', { name: 'Generation mode' });
expect(modeSelect).toHaveTextContent('T2VA');
await user.click(modeSelect);
await user.click(await screen.findByRole('option', { name: 'FL2VA' }));
expect(modeSelect).toHaveTextContent('FL2VA');
expect(modeSelect).toHaveAttribute(
'title',
'First/last frames to video + audio. Start from a first frame image. Add an optional last frame to guide the ending.',
);
const generateButton = await screen.findByRole('button', { name: 'Generate' });
await waitFor(() => expect(generateButton).toBeEnabled());
await user.click(generateButton);
await waitFor(() => {
expect(outbound.some((message) => message.type === 'session_init_v2')).toBe(true);
});
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
expect(initMessage.generation_mode).toBe('fl2va');
expect(screen.queryByRole('combobox', { name: 'Generation mode' }))
.not.toBeInTheDocument();
});
it('starts a streaming session from a custom initial prompt without using curated prompts', async () => {
const outbound: any[] = [];
server.on('connection', (socket) => {
+302 -249
View File
@@ -1,11 +1,20 @@
"use client";
import { Fragment, useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState } from "react";
import { Fragment, useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState, type Dispatch, type SetStateAction } from "react";
import { AnimatePresence, motion } from "framer-motion";
import { Download, Share2 } from "lucide-react";
import DevtoolsShell from "@/components/devtools/DevtoolsShell";
import MonitorPage from "@/components/MonitorPage";
import ChatBar from "@/components/ChatBar";
import AssetList from "@/components/AssetList";
import CreationStudio from "@/components/creation/CreationStudio";
import {
buildMentionOptions,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type ResolutionId,
} from "@/lib/creationConfig";
import { toGenerationMode } from "@/lib/generationMode";
import type { SessionCreationConfig } from "@/components/creation/SessionCreationConfigPills";
import SessionTimeoutModal from "@/components/SessionTimeoutModal";
import Sidebar from "@/components/Sidebar";
import Header from "@/components/Header";
@@ -14,18 +23,27 @@ import Workspace from "@/components/Workspace";
import { saveProject, saveProjectMetadata, listProjects, loadProjectClips, deleteProject, pruneOldProjects, type StoredProject, type StoredClip } from "@/lib/projectStorage";
import { isInfrastructureError } from "@/lib/ws/reducer";
import { useStore } from "@/hooks/useStore";
import { useAssetLibrary } from "@/hooks/useAssetLibrary";
import { useGenerationCapabilities } from "@/hooks/useGenerationCapabilities";
import { resolveDevtoolsMode } from "@/lib/devtoolsMode";
import { createAvPipeline, DEFAULT_AV_MIME } from "@/lib/media/avPipeline";
import { remuxArchivedFmp4Segments } from "@/lib/media/fmp4Remux";
import { DEFAULT_CUSTOM_PRESET_ID, parseStoryPresets, sanitizePresetId } from "@/lib/presets";
import { DEFAULT_GENERATION_MODE, buildGenerationInitFields, validateGenerationInputs, type GenerationMode, type GenerationInitFields, type GenerationAsset } from "@/lib/generationMode";
import {
buildRewritePromptWindowSnapshot,
buildRewritePromptWindowSnapshotFromPrompts,
normalizePromptWindowSnapshot,
} from "@/lib/prompts/promptWindowSnapshot";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
clampLobbySelectionToCapabilities,
parseLobbyCapabilitiesBundle,
resolveModelCapabilities,
validateLobbyCreationSelection,
type LobbyCapabilitiesBundle,
} from "@/lib/creationCapabilities";
import {
buildCreationInitPayload,
parseEchoedCreationConfig,
} from "@/lib/creationPayload";
import rawPresets from "@/lib/storyPresetsData";
import { cn } from "@/lib/utils";
import { createWebSocketConnection, detachAndCloseWebSocket } from "@/lib/ws/client";
@@ -82,98 +100,6 @@ function yieldToEventLoop(): Promise<void> {
return new Promise((r) => setTimeout(r, 0));
}
const HERO_WAVE_LIGHT = ["#2A4A98", "#4878E5", "#6FA0F2", "#B0BCC8", "#E8D99E", "#D8C844", "#C2A620"];
const HERO_WAVE_DARK = ["#143468", "#1E58B8", "#3892F0", "#80B8E8", "#B8D0EA", "#E2D498", "#DABB50"];
const HERO_TEXT = "Direct scenes in seconds";
function HeroTagline() {
const ref = useRef<HTMLHeadingElement>(null);
useEffect(() => {
const el = ref.current;
if (!el) return;
let rafId = 0;
function play() {
const chars = el!.querySelectorAll<HTMLSpanElement>("[data-char]");
if (!chars.length) return;
cancelAnimationFrame(rafId);
const isDark = document.documentElement.classList.contains("dark");
const colors = isDark ? HERO_WAVE_DARK : HERO_WAVE_LIGHT;
const waveLen = 10;
const total = chars.length + waveLen;
const duration = 1200;
const maxBlur = 3.5;
const start = performance.now();
function tick() {
const t = Math.min((performance.now() - start) / duration, 1);
const pos = t * total;
chars.forEach((ch, i) => {
const rel = pos - i;
if (rel >= 0 && rel < waveLen) {
const norm = rel / waveLen;
const ci = Math.floor(norm * colors.length);
ch.style.color = colors[Math.min(colors.length - 1, ci)];
let blur = 0;
if (norm < 0.25) {
blur = maxBlur * (1 - norm / 0.25);
} else if (norm > 0.75) {
blur = maxBlur * ((norm - 0.75) / 0.25);
}
ch.style.filter = blur > 0.1 ? `blur(${blur.toFixed(1)}px)` : "";
} else {
ch.style.color = "";
ch.style.filter = "";
}
});
if (t < 1) {
rafId = requestAnimationFrame(tick);
} else {
chars.forEach((ch) => {
ch.style.color = "";
ch.style.filter = "";
});
}
}
rafId = requestAnimationFrame(tick);
}
const initialDelay = setTimeout(play, 400);
const interval = setInterval(play, 5000);
return () => {
clearTimeout(initialDelay);
clearInterval(interval);
cancelAnimationFrame(rafId);
};
}, []);
return (
<h1 ref={ref} className="text-center text-3xl font-medium text-[#343537] dark:text-[#FAFAFB] sm:text-4xl">
{HERO_TEXT.split(" ").map((word, wi) => (
<Fragment key={wi}>
{wi > 0 && (
<span data-char className="transition-[color,filter] duration-150">
{" "}
</span>
)}
<span className="inline-flex">
{word.split("").map((char, ci) => (
<span key={ci} data-char className="inline-block transition-[color,filter] duration-150">
{char}
</span>
))}
</span>
</Fragment>
))}
</h1>
);
}
export default function Page() {
const storesRef = useRef<PageStores | null>(null);
if (!storesRef.current) {
@@ -327,8 +253,33 @@ export default function Page() {
const [ttffValueMs, setTtffValueMs] = useState<number | null>(null);
const ttffIntervalRef = useRef<ReturnType<typeof setInterval> | null>(null);
const pendingInitialPromptRef = useRef("");
const referenceFileRef = useRef<File | null>(null);
const firstFrameFileRef = useRef<File | null>(null);
const lastFrameFileRef = useRef<File | null>(null);
const lastArchivedReplayKeyRef = useRef("");
const [sidebarOpen, setSidebarOpen] = useState(false);
const [creationModelId, setCreationModelId] = useState<CreationModelId>("fast-ltx23");
const [creationModeId, setCreationModeId] = useState<CreationModeId>("t2v");
const [creationAspectRatio, setCreationAspectRatio] = useState<AspectRatioId>("16:9");
const [creationResolution, setCreationResolution] = useState<ResolutionId>("720p");
const [creationDurationSec, setCreationDurationSec] = useState(5);
const [lobbyCapabilitiesBundle, setLobbyCapabilitiesBundle] = useState<LobbyCapabilitiesBundle>(
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
);
const activeModelCapabilities = useMemo(
() => resolveModelCapabilities(lobbyCapabilitiesBundle, creationModelId),
[lobbyCapabilitiesBundle, creationModelId],
);
const [sessionCreationConfig, setSessionCreationConfig] = useState<SessionCreationConfig>({
modelId: "fast-ltx23",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
});
const [referencePreviewUrl, setReferencePreviewUrl] = useState<string | null>(null);
const [firstFramePreviewUrl, setFirstFramePreviewUrl] = useState<string | null>(null);
const [lastFramePreviewUrl, setLastFramePreviewUrl] = useState<string | null>(null);
const [currentThumbnail, setCurrentThumbnail] = useState<string | null>(null);
const currentProjectIdRef = useRef("");
const currentProjectCreatedAtRef = useRef(0);
@@ -345,23 +296,61 @@ export default function Page() {
const [isMobileShareCapable, setIsMobileShareCapable] = useState(false);
const [videoMuted, setVideoMuted] = useState(true);
const [timeoutModalOpen, setTimeoutModalOpen] = useState(false);
const [generationMode, setGenerationMode] = useState<GenerationMode>(DEFAULT_GENERATION_MODE);
const assetLibrary = useAssetLibrary();
const { capabilities, capabilityNotice, refreshCapabilities } = useGenerationCapabilities();
const joiningRef = useRef(false);
const activeGenerationRef = useRef<{ fields: GenerationInitFields; assets: GenerationAsset[]; mock: boolean } | null>(null);
const generationInputError = validateGenerationInputs(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets);
const generationSupported = capabilities.modes.includes(generationMode);
const generationInputsValid = !generationInputError && generationSupported && !assetLibrary.uploading;
function changeGenerationMode(mode: GenerationMode) {
if (sessionStore.get().sessionStarted || joiningRef.current || !capabilities.modes.includes(mode)) return;
setGenerationMode(mode);
assetLibrary.clearConditioning();
}
useEffect(() => {
setIsMobileShareCapable(typeof navigator.canShare === "function" && window.matchMedia("(pointer: coarse)").matches);
}, []);
useEffect(() => {
return () => {
if (referencePreviewUrl) {
URL.revokeObjectURL(referencePreviewUrl);
}
if (firstFramePreviewUrl) {
URL.revokeObjectURL(firstFramePreviewUrl);
}
if (lastFramePreviewUrl) {
URL.revokeObjectURL(lastFramePreviewUrl);
}
};
}, [referencePreviewUrl, firstFramePreviewUrl, lastFramePreviewUrl]);
function setPreviewUrl(setter: Dispatch<SetStateAction<string | null>>, file: File | null) {
setter((current) => {
if (current) URL.revokeObjectURL(current);
return file ? URL.createObjectURL(file) : null;
});
}
function handleReferenceSelect(file: File | null) {
referenceFileRef.current = file;
setPreviewUrl(setReferencePreviewUrl, file);
}
function handleFirstFrameSelect(file: File | null) {
firstFrameFileRef.current = file;
setPreviewUrl(setFirstFramePreviewUrl, file);
}
function handleLastFrameSelect(file: File | null) {
lastFrameFileRef.current = file;
setPreviewUrl(setLastFramePreviewUrl, file);
}
const mentionOptions = useMemo(() => buildMentionOptions(storyPresets as Array<{ id?: string; label?: string; description?: string }>), [storyPresets]);
const lobbyStoryPresets = useMemo(
() =>
(storyPresets as Array<{ id?: string; label?: string; description?: string; segment_prompts?: unknown }>)
.filter((preset) => typeof preset.id === "string" && typeof preset.label === "string")
.map((preset) => ({
id: String(preset.id),
label: String(preset.label),
description: typeof preset.description === "string" ? preset.description : undefined,
segmentCount: Array.isArray(preset.segment_prompts) ? preset.segment_prompts.length : undefined,
})),
[storyPresets],
);
const videoElRef = useRef<HTMLVideoElement | null>(null);
const archivedPlaybackElRef = useRef<HTMLVideoElement | null>(null);
const viewingModePlaybackStateRef = useRef<{
@@ -408,7 +397,7 @@ export default function Page() {
// --- Derived values ---
const canStartSession = generationInputsValid && !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
const canStartSession = !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
const currentClipLabel = useMemo(() => {
if ((activeClip as Record<string, any>)?.label) return (activeClip as Record<string, any>).label;
@@ -541,6 +530,65 @@ export default function Page() {
setRuntimeReady(true);
}, []);
function applyLobbyCapabilitiesBundle(bundle: LobbyCapabilitiesBundle) {
setLobbyCapabilitiesBundle(bundle);
const clamped = clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(bundle, creationModelId),
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
setCreationModelId(clamped.modelId);
setCreationModeId(clamped.modeId);
setCreationAspectRatio(clamped.aspectRatio);
setCreationResolution(clamped.resolution);
setCreationDurationSec(clamped.durationSec);
}
function handleCreationModelChange(modelId: CreationModelId) {
const clamped = clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(lobbyCapabilitiesBundle, modelId),
modelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
setCreationModelId(clamped.modelId);
setCreationModeId(clamped.modeId);
setCreationAspectRatio(clamped.aspectRatio);
setCreationResolution(clamped.resolution);
setCreationDurationSec(clamped.durationSec);
}
useEffect(() => {
if (!runtimeReady) return;
let cancelled = false;
void fetch("/creation-capabilities", {
headers: { Accept: "application/json" },
cache: "no-store",
})
.then(async (response) => {
if (!response.ok) return DEFAULT_LOBBY_CAPABILITIES_BUNDLE;
return parseLobbyCapabilitiesBundle(await response.json());
})
.then((bundle) => {
if (!cancelled) {
applyLobbyCapabilitiesBundle(bundle);
}
})
.catch(() => {
if (!cancelled) {
applyLobbyCapabilitiesBundle(DEFAULT_LOBBY_CAPABILITIES_BUNDLE);
}
});
return () => {
cancelled = true;
};
}, [runtimeReady]);
useEffect(() => {
if (!runtimeReady || initializedRef.current) return;
initializedRef.current = true;
@@ -764,11 +812,7 @@ export default function Page() {
function recoverFailedSessionStart(notice: string) {
const restoredDraft = normalizeInitialPrompt(pendingInitialPromptRef.current);
if (wsRef.current) {
detachAndCloseWebSocket(wsRef.current);
wsRef.current = null;
}
resetToLobbyState({ preserveSessionNotice: true });
resetToLobbyState();
clearPendingProjectPointers();
pendingInitialPromptRef.current = "";
sessionStore.patch({
@@ -1723,10 +1767,6 @@ export default function Page() {
function resetToLobbyState({ preserveSessionNotice = false, preservePlayback = false } = {}) {
setVideoMuted(true);
if (!preserveSessionNotice) {
setGenerationMode(DEFAULT_GENERATION_MODE);
assetLibrary.clearConditioning();
}
clearCountdownInterval();
pendingInitialPromptRef.current = "";
sessionStore.patch({
@@ -1761,8 +1801,6 @@ export default function Page() {
function resetToProjectLobbyState() {
setVideoMuted(true);
setGenerationMode(DEFAULT_GENERATION_MODE);
assetLibrary.clearConditioning();
pendingInitialPromptRef.current = "";
sessionStore.patch({
sessionStarted: false,
@@ -1786,34 +1824,44 @@ export default function Page() {
resetPlaybackState();
}
function buildProjectInitPayload(type: "session_init_v2" | "project_init_v1") {
async function buildProjectInitPayload(type: "session_init_v2" | "project_init_v1") {
const segmentPrompts = getSessionInitPrompts();
setSeedPrompts(segmentPrompts);
const creationPayload = await buildCreationInitPayload({
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
referenceFile: referenceFileRef.current,
firstFrameFile: firstFrameFileRef.current,
lastFrameFile: lastFrameFileRef.current,
});
return {
type,
...(activeGenerationRef.current?.fields || buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets)),
generation_mode: toGenerationMode(creationModeId),
preset_id: getInitialPresetId(),
preset_label: getInitialPresetLabel(),
curated_prompts: segmentPrompts,
initial_rollout_prompt: normalizeInitialPrompt(pendingInitialPromptRef.current),
initial_image: null,
single_clip_mode: false,
enhancement_enabled: sessionStore.get().enhancementEnabled,
auto_extension_enabled: sessionStore.get().autoExtensionEnabled,
loop_generation_enabled: sessionStore.get().loopGenerationEnabled,
...creationPayload,
};
}
function sendSessionInitMessage() {
async function sendSessionInitMessage() {
const ws = wsRef.current;
if (!ws) return;
ws.send(JSON.stringify(buildProjectInitPayload("session_init_v2")));
ws.send(JSON.stringify(await buildProjectInitPayload("session_init_v2")));
}
function sendProjectInitMessage() {
async function sendProjectInitMessage() {
const ws = wsRef.current;
if (!ws || ws.readyState !== WebSocket.OPEN) return;
ws.send(JSON.stringify(buildProjectInitPayload("project_init_v1")));
ws.send(JSON.stringify(await buildProjectInitPayload("project_init_v1")));
}
function sendEndProjectKeepSession() {
@@ -1840,12 +1888,6 @@ export default function Page() {
return;
}
if (decoded.kind !== "json") return;
if (decoded.data?.type === "error" && sessionStore.get().sessionStarted
&& (!sessionStore.get().gpuAssigned || decoded.data.error_code === "invalid_generation_input")) {
const message = typeof decoded.data.message === "string" ? decoded.data.message : "The generation inputs were rejected. Check the mode and selected assets.";
recoverFailedSessionStart(message);
return;
}
if (decoded.data?.type === "error" && isInfrastructureError(decoded.data)) {
const message = typeof decoded.data?.message === "string" && decoded.data.message.trim()
? decoded.data.message.trim()
@@ -1856,6 +1898,9 @@ export default function Page() {
return;
}
const normalizedEvent = normalizeSocketMessage(decoded.data);
if (decoded.data?.type === "gpu_assigned" || decoded.data?.type === "ltx2_stream_start") {
applyEchoedCreationConfig(decoded.data);
}
await applyNormalizedSocketEvent(normalizedEvent, {
sessionStore,
promptWindowStore,
@@ -1898,7 +1943,12 @@ export default function Page() {
onOpen: () => {
opened = true;
sessionStore.patch({ connected: true, connecting: false });
sendSessionInitMessage();
void sendSessionInitMessage().catch((error) => {
console.error("Failed to send session init payload:", error);
recoverFailedSessionStart(
error instanceof Error ? error.message : "Failed to prepare session settings.",
);
});
},
onMessage: (event: MessageEvent) => {
wsMessageQueueRef.current = wsMessageQueueRef.current
@@ -1971,16 +2021,29 @@ export default function Page() {
}
}
function beginProjectLocally({ force = false, mockRuntime = capabilities.mock === true } = {}) {
function syncSessionCreationConfigFromLobby() {
setSessionCreationConfig({
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
});
}
function applyEchoedCreationConfig(data: unknown) {
const echoed = parseEchoedCreationConfig(data);
if (!echoed) {
return;
}
setSessionCreationConfig(echoed);
}
function beginProjectLocally({ force = false } = {}) {
if (!force && !canStartSession) return;
if (!generationInputsValid) return false;
if (sessionStore.get().sessionStarted || sessionStore.get().projectResetPending) return false;
activeGenerationRef.current = {
fields: buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets),
assets: assetLibrary.assets.filter((asset) => assetLibrary.conditioningAssets.some((item) => item.asset_id === asset.asset_id)),
mock: mockRuntime,
};
setTimeoutModalOpen(false);
syncSessionCreationConfigFromLobby();
// Unmute during the user gesture so iOS Safari permits audio playback.
setVideoMuted(false);
if (viewingProject) closeViewingProject();
@@ -2039,44 +2102,34 @@ export default function Page() {
}
async function joinSession({ force = false } = {}) {
if (joiningRef.current || sessionStore.get().sessionStarted) return;
if (generationInputError || assetLibrary.uploading) {
showPreSessionNotice(generationInputError || "Wait for the asset upload to finish.");
const validationError = validateLobbyCreationSelection({
capabilities: activeModelCapabilities,
modelId: creationModelId,
modeId: creationModeId,
aspectRatio: creationAspectRatio,
resolution: creationResolution,
durationSec: creationDurationSec,
referenceFile: referenceFileRef.current,
firstFrameFile: firstFrameFileRef.current,
lastFrameFile: lastFrameFileRef.current,
});
if (validationError) {
showPreSessionNotice(validationError);
return;
}
joiningRef.current = true;
try {
await startGenerationSession({ force });
} finally {
joiningRef.current = false;
}
}
async function startGenerationSession({ force = false } = {}) {
sessionStore.patch({ sessionNotice: "" });
streamStore.patch({ loadingAnimation: true });
const currentCapabilities = await refreshCapabilities();
if (!currentCapabilities.modes.includes(generationMode)) {
streamStore.patch({ loadingAnimation: false });
showPreSessionNotice(`${generationMode.toUpperCase()} is unavailable on this runtime. Connect a full H3 runtime or choose a supported mode.`);
return;
}
const assetProblem = await assetLibrary.verifySelectedAssets();
if (assetProblem) {
streamStore.patch({ loadingAnimation: false });
showPreSessionNotice(assetProblem);
return;
}
if (
wsRef.current
&& wsRef.current.readyState === WebSocket.OPEN
&& sessionStore.get().connected
) {
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
streamStore.patch({ loadingAnimation: false });
return;
if (!beginProjectLocally({ force })) return;
try {
await sendProjectInitMessage();
} catch (error) {
console.error("Failed to send project init payload:", error);
showPreSessionNotice(error instanceof Error ? error.message : "Failed to prepare session settings.");
}
sendProjectInitMessage();
return;
}
showPreSessionNotice("");
@@ -2088,7 +2141,7 @@ export default function Page() {
showPreSessionNotice(probe.notice);
return;
}
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
if (!beginProjectLocally({ force })) {
streamStore.patch({ loadingAnimation: false });
sessionStore.patch({ connecting: false });
return;
@@ -2098,7 +2151,12 @@ export default function Page() {
&& wsRef.current.readyState === WebSocket.OPEN
&& sessionStore.get().connected
) {
sendProjectInitMessage();
try {
await sendProjectInitMessage();
} catch (error) {
console.error("Failed to send project init payload:", error);
showPreSessionNotice(error instanceof Error ? error.message : "Failed to prepare session settings.");
}
return;
}
connectWebSocket();
@@ -2115,10 +2173,6 @@ export default function Page() {
createdAt: currentProjectCreatedAtRef.current || Date.now(),
lastThumbnail: currentThumbnail,
promptEvents: [...(rewriteStore.get().promptEvents as Record<string, unknown>[])],
generationMode: activeGenerationRef.current?.fields.generation_mode || DEFAULT_GENERATION_MODE,
conditioningAssets: activeGenerationRef.current?.fields.conditioning_assets || [],
assets: activeGenerationRef.current?.assets || [],
mock: activeGenerationRef.current?.mock === true,
};
const clips: StoredClip[] = (streamStore.get().completedClips as any[])
.filter((clip: any) => clip?.blob instanceof Blob)
@@ -2553,24 +2607,6 @@ export default function Page() {
// --- Render ---
const conditioningPanel = generationMode !== "t2va" && !sessionStarted && !sessionExpired ? (
<AssetList
mode={generationMode}
assets={assetLibrary.assets}
conditioning={assetLibrary.conditioningAssets}
locked={Boolean(loadingAnimation || projectResetPending)}
uploading={assetLibrary.uploading}
error={assetLibrary.assetError}
validationNotice={generationInputError}
onUpload={assetLibrary.uploadAssets}
onAssign={assetLibrary.assignAsset}
onRemove={assetLibrary.removeAsset}
onUnselect={assetLibrary.removeConditioning}
onMove={assetLibrary.moveConditioning}
onMissing={assetLibrary.checkAssetAvailability}
/>
) : null;
if (!runtimeReady) {
return null;
}
@@ -2594,17 +2630,13 @@ export default function Page() {
enhancementEnabled={enhancementEnabled as boolean}
autoExtensionEnabled={autoExtensionEnabled as boolean}
loopGenerationEnabled={loopGenerationEnabled as boolean}
canJoinSession={canStartSession}
canJoinSession={canJoinSession as boolean}
canSubmitContinuation={canSubmitContinuation}
editableMode={editableMode as boolean}
demoMode={demoMode as boolean}
editableCanJoin={editableCanJoin as boolean}
curatedPromptLimit={curatedPromptLimit as number}
maxCuratedPromptCount={maxCuratedPromptCount as number}
generationMode={generationMode}
supportedGenerationModes={capabilities.modes}
conditioningPanel={conditioningPanel}
onGenerationModeChange={changeGenerationMode}
onPresetChange={handlePresetSelectionChange}
onEnhancementToggle={handleEnhancementToggle}
onCuratedPromptLimitChange={handleCuratedPromptLimitChange}
@@ -2738,12 +2770,7 @@ export default function Page() {
/>
<Header timeLeft={headerTimeLeft} formatTime={formatTime} onToggleSidebar={() => setSidebarOpen((prev) => !prev)} />
<div className={cn(
"relative flex flex-1 min-h-0 flex-col px-4 pb-2 sm:px-6 sm:pb-12",
!isViewingMode && !showActiveProject && generationMode !== "t2va"
? "justify-start overflow-y-auto pt-4"
: "justify-center",
)}>
<div className={cn("relative flex flex-1 min-h-0 flex-col", showActiveProject || isViewingMode ? "justify-center px-4 pb-2 sm:px-6 sm:pb-12" : "overflow-hidden")}>
{isViewingMode && (
<>
{viewingSelectedClip && (
@@ -2782,8 +2809,6 @@ export default function Page() {
/>
</section>
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
{viewingProject?.project.mock && <p className="mb-2 text-center text-xs text-violet-600 dark:text-violet-300">Demo sample · This saved clip was not generated by an AI model.</p>}
{viewingProject?.project.generationMode && <p className="mb-3 text-center text-xs text-muted-foreground">{viewingProject.project.generationMode.toUpperCase()} · {viewingProject.project.conditioningAssets?.length || 0} saved references. Uploaded originals may expire; your saved video remains available.</p>}
<ChatBar sessionStarted={false} viewingReadOnly={true} onStartNewProject={handleStartNewProject} onBackFromViewing={closeViewingProject} />
</motion.div>
</>
@@ -2862,51 +2887,79 @@ export default function Page() {
/>
</section>
<AnimatePresence>
{!showActiveProject && generationMode === "t2va" && (
<motion.div
key="hero-tagline"
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0, transition: { duration: 0.2, ease: "easeIn" } }}
transition={{ duration: 0.5, ease: "easeOut" }}
className="pointer-events-none absolute inset-x-0 top-0 bottom-1/2 z-10 flex items-center justify-center px-4"
>
<HeroTagline />
</motion.div>
)}
</AnimatePresence>
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
<ChatBar
sessionStarted={sessionStarted as boolean}
rewritingSeedPrompts={rewritingSeedPrompts as boolean}
{!showActiveProject ? (
<CreationStudio
value={livePromptDraft as string}
disabled={projectResetPending as boolean}
isGenerating={loadingAnimation as boolean}
storyPresets={storyPresets as any[]}
continuationDraft={livePromptDraft as string}
canJoinSession={canStartSession}
canSubmitContinuation={canSubmitContinuation}
sessionExpired={sessionExpired as boolean}
sessionNotice={sessionNotice as string}
projectResetPending={projectResetPending as boolean}
generationMode={generationMode}
supportedGenerationModes={capabilities.modes}
generationInputsValid={generationInputsValid}
capabilityNotice={!generationSupported ? `${generationMode.toUpperCase()} is unavailable on this runtime.` : capabilityNotice}
mockRuntime={capabilities.mock}
conditioningPanel={conditioningPanel}
canSubmit={canStartSession}
modelId={creationModelId}
modeId={creationModeId}
aspectRatio={creationAspectRatio}
resolution={creationResolution}
durationSec={creationDurationSec}
referencePreviewUrl={referencePreviewUrl}
firstFramePreviewUrl={firstFramePreviewUrl}
lastFramePreviewUrl={lastFramePreviewUrl}
mentionOptions={mentionOptions}
storyPresets={lobbyStoryPresets}
capabilities={activeModelCapabilities}
onValueChange={(value) => sessionStore.patch({ livePromptDraft: value })}
onSubmit={() => void joinSession()}
onKeyDown={handleLivePromptKeydown}
onModelChange={handleCreationModelChange}
onModeChange={setCreationModeId}
onAspectRatioChange={setCreationAspectRatio}
onResolutionChange={setCreationResolution}
onDurationChange={setCreationDurationSec}
onReferenceSelect={handleReferenceSelect}
onFirstFrameSelect={handleFirstFrameSelect}
onLastFrameSelect={handleLastFrameSelect}
onPresetGenerate={handlePresetGenerate}
onContinuationInput={handleLivePromptInput}
onContinuationKeydown={handleLivePromptKeydown}
onGenerate={joinSession}
onSubmitContinuation={submitLivePrompt}
onLeave={leaveSession}
onStartNewProject={handleStartNewProject}
onGenerationModeChange={changeGenerationMode}
onSpeechTranscript={handleLivePromptSpeechTranscript}
onSpeechInterimChange={handleLivePromptSpeechInterim}
onOpenProjects={() => setSidebarOpen(true)}
/>
</motion.div>
) : (
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
<ChatBar
sessionStarted={sessionStarted as boolean}
rewritingSeedPrompts={rewritingSeedPrompts as boolean}
isGenerating={loadingAnimation as boolean}
storyPresets={storyPresets as any[]}
continuationDraft={livePromptDraft as string}
canJoinSession={canStartSession}
canSubmitContinuation={canSubmitContinuation}
sessionExpired={sessionExpired as boolean}
sessionNotice={sessionNotice as string}
projectResetPending={projectResetPending as boolean}
sessionCreationConfig={sessionCreationConfig}
configPillsReadOnly
onSessionModelChange={(modelId) => setSessionCreationConfig((current) => ({ ...current, modelId }))}
onSessionModeChange={(modeId) => setSessionCreationConfig((current) => ({ ...current, modeId }))}
onSessionAspectRatioChange={(aspectRatio) => setSessionCreationConfig((current) => ({ ...current, aspectRatio }))}
onSessionResolutionChange={(resolution) => setSessionCreationConfig((current) => ({ ...current, resolution }))}
onSessionDurationChange={(durationSec) => setSessionCreationConfig((current) => ({ ...current, durationSec }))}
onPresetGenerate={handlePresetGenerate}
onContinuationInput={handleLivePromptInput}
onContinuationKeydown={handleLivePromptKeydown}
onGenerate={joinSession}
onSubmitContinuation={submitLivePrompt}
onLeave={leaveSession}
onStartNewProject={handleStartNewProject}
onSpeechTranscript={handleLivePromptSpeechTranscript}
onSpeechInterimChange={handleLivePromptSpeechInterim}
/>
</motion.div>
)}
{sessionNotice && !showActiveProject && (
<div className="mx-auto mt-2 w-full max-w-3xl px-4">
<div className="rounded-xl border border-rose-500/20 bg-rose-500/10 px-4 py-2.5 text-center text-xs text-rose-700 dark:text-rose-300">
{sessionNotice}
</div>
</div>
)}
</div>
</div>
</main>
@@ -1,41 +0,0 @@
import { render, screen } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { describe, expect, it, vi } from "vitest";
import AssetList from "./AssetList";
import type { GenerationAsset } from "@/lib/generationMode";
const frame: GenerationAsset = { asset_id: "first", kind: "image", name: "frame.png", mime_type: "image/png", size: 2000, url: "/assets/first" };
const video: GenerationAsset = { asset_id: "video", kind: "video", name: "motion.mp4", mime_type: "video/mp4", size: 2000, url: "/assets/video" };
function props() {
return { assets: [frame, video], onUpload: vi.fn(), onAssign: vi.fn(), onRemove: vi.fn(), onUnselect: vi.fn(), onMove: vi.fn(), onMissing: vi.fn() };
}
describe("Asset List", () => {
it("uploads to the library and assigns images to endpoint roles", async () => {
const callbacks = props();
const user = userEvent.setup();
render(<AssetList {...callbacks} mode="fl2va" conditioning={[]} />);
await user.selectOptions(screen.getByRole("combobox", { name: "First frame" }), "first");
expect(callbacks.onAssign).toHaveBeenCalledWith("first", "first_frame");
expect(screen.getByRole("combobox", { name: "Last frame" })).toHaveValue("");
expect(screen.queryByRole("option", { name: "motion.mp4" })).not.toBeInTheDocument();
const file = new File(["image"], "new.png", { type: "image/png" });
await user.upload(screen.getByLabelText("Upload assets", { selector: "input" }), file);
expect(callbacks.onUpload).toHaveBeenCalledWith([file]);
});
it("exposes accessible ordering and removal controls for multimodal references", async () => {
const callbacks = props();
const user = userEvent.setup();
render(<AssetList {...callbacks} mode="ref2va" conditioning={[{ asset_id: "first", role: "reference" }, { asset_id: "video", role: "reference" }]} />);
await user.click(screen.getByRole("button", { name: "Move motion.mp4 up" }));
expect(callbacks.onMove).toHaveBeenCalledWith(1, 0);
await user.click(screen.getByRole("button", { name: "Unselect frame.png" }));
expect(callbacks.onUnselect).toHaveBeenCalledWith(0);
expect(screen.getByRole("button", { name: "Move frame.png up" })).toBeDisabled();
});
it("locks uploads and assignments while starting generation", () => {
render(<AssetList {...props()} mode="fl2va" conditioning={[]} locked />);
expect(screen.getByRole("button", { name: "Upload assets" })).toBeDisabled();
expect(screen.getByRole("combobox", { name: "First frame" })).toBeDisabled();
});
});
@@ -1,154 +0,0 @@
"use client";
import { useRef, useState } from "react";
import { ArrowDown, ArrowUp, AudioLines, Check, GripVertical, ImagePlus, Plus, Trash2, Upload, X } from "lucide-react";
import { Button } from "@/components/ui/button";
import { NativeSelect } from "@/components/ui/native-select";
import { cn } from "@/lib/utils";
import type { ConditioningAsset, ConditioningRole, GenerationAsset, GenerationMode } from "@/lib/generationMode";
interface AssetListProps {
mode: GenerationMode;
assets: GenerationAsset[];
conditioning: ConditioningAsset[];
locked?: boolean;
uploading?: boolean;
error?: string;
validationNotice?: string | null;
onUpload: (files: File[]) => void;
onAssign: (assetId: string, role: ConditioningRole) => void;
onRemove: (assetId: string) => void;
onUnselect: (index: number) => void;
onMove: (from: number, to: number) => void;
onMissing: (assetId: string) => void;
}
function AssetPreview({ asset, onMissing, compact = false }: {
asset: GenerationAsset;
onMissing: (assetId: string) => void;
compact?: boolean;
}) {
const [previewFailed, setPreviewFailed] = useState(false);
function previewError() {
setPreviewFailed(true);
onMissing(asset.asset_id);
}
const className = cn("h-full w-full object-cover", asset.missing && "opacity-25");
if (asset.missing) return <span className="p-2 text-center text-[10px] text-muted-foreground">Upload again</span>;
if (previewFailed) return <span className="p-2 text-center text-[10px] text-muted-foreground">Preview unavailable</span>;
if (asset.kind === "image") {
return <img src={asset.url} alt={asset.name} className={className} onError={previewError} />;
}
if (asset.kind === "video") {
return <video src={asset.url} aria-label={`Preview ${asset.name}`} className={className} muted playsInline controls={!compact} preload="metadata" onError={previewError} />;
}
return (
<div className="flex h-full w-full flex-col items-center justify-center gap-2 bg-violet-500/10 p-2 text-violet-500">
<AudioLines className="size-6" />
{!compact && <audio src={asset.url} aria-label={`Preview ${asset.name}`} controls preload="metadata" className="h-6 w-full min-w-0" onError={previewError} />}
</div>
);
}
/** A reusable library/picker. The parent asset store owns uploads and selection. */
export default function AssetList({
mode, assets, conditioning, locked = false, uploading = false, error = "", validationNotice,
onUpload, onAssign, onRemove, onUnselect, onMove, onMissing,
}: AssetListProps) {
const inputRef = useRef<HTMLInputElement>(null);
const [libraryOpen, setLibraryOpen] = useState(true);
const [dragIndex, setDragIndex] = useState<number | null>(null);
const disabled = locked || uploading;
const imageAssets = assets.filter((asset) => asset.kind === "image");
return (
<section aria-label="Asset List" className="overflow-hidden rounded-2xl border border-input bg-card/70 shadow-sm backdrop-blur-sm">
<div className="flex items-center justify-between gap-3 px-4 py-3">
<div>
<h2 className="text-xs font-semibold tracking-wide">{mode === "fl2va" ? "Frame guidance" : "Reference sequence"}</h2>
<p className="mt-0.5 text-[11px] text-muted-foreground">{locked ? "Inputs are locked for this project." : mode === "fl2va" ? "Choose your opening image and, optionally, the ending." : "Arrange references in the order you want the model to read them."}</p>
</div>
<Button type="button" variant="outline" size="sm" disabled={disabled} onClick={() => inputRef.current?.click()} className="shrink-0 gap-1.5 rounded-full text-xs">
<Upload className="size-3.5" />{uploading ? "Uploading…" : "Upload assets"}
</Button>
<input ref={inputRef} type="file" aria-label="Upload assets" className="sr-only" multiple accept={mode === "fl2va" ? "image/*" : "image/*,video/*,audio/*"} disabled={disabled} onChange={(event) => {
const files = Array.from(event.target.files || []);
if (files.length) onUpload(files);
event.target.value = "";
}} />
</div>
<div className="max-h-[min(42vh,350px)] overflow-y-auto px-4 pb-3">
{mode === "fl2va" ? (
<div className="grid grid-cols-2 gap-3">
{(["first_frame", "last_frame"] as const).map((role) => {
const label = role === "first_frame" ? "First frame" : "Last frame";
const assetId = conditioning.find((item) => item.role === role)?.asset_id || "";
const asset = assets.find((item) => item.asset_id === assetId);
return (
<div key={role} className="overflow-hidden rounded-xl border border-input bg-background/40 p-2">
<div className="flex h-20 items-center justify-center overflow-hidden rounded-lg bg-muted/60 sm:h-24">
{asset ? <AssetPreview key={asset.asset_id} asset={asset} onMissing={onMissing} /> : <ImagePlus className="size-6 text-muted-foreground/45" />}
</div>
<label htmlFor={`asset-${role}`} className="mb-1 mt-2 block text-[11px] font-medium">{label} <span className="font-normal text-muted-foreground">{role === "first_frame" ? "· required" : "· optional"}</span></label>
<NativeSelect id={`asset-${role}`} aria-label={label} value={assetId} disabled={disabled} className="h-8 text-xs" onChange={(event) => onAssign(event.target.value, role)}>
<option value="">{imageAssets.length ? "Choose an image" : "Upload an image first"}</option>
{imageAssets.map((item) => <option key={item.asset_id} value={item.asset_id} disabled={item.missing}>{item.name}{item.missing ? " (upload again)" : ""}</option>)}
</NativeSelect>
</div>
);
})}
</div>
) : (
<>
{conditioning.length ? (
<ol aria-label="Ordered references" className="flex flex-col gap-2">
{conditioning.map((item, index) => {
const asset = assets.find((entry) => entry.asset_id === item.asset_id);
if (!asset) return null;
return (
<li key={`${item.asset_id}-${index}`} draggable={!disabled} onDragStart={() => setDragIndex(index)} onDragEnd={() => setDragIndex(null)} onDragOver={(event) => { if (!disabled && dragIndex !== null) event.preventDefault(); }} onDrop={(event) => { event.preventDefault(); if (!disabled && dragIndex !== null) onMove(dragIndex, index); setDragIndex(null); }} className={cn("flex items-center gap-2 rounded-xl border border-input bg-background/40 p-2", dragIndex === index && "opacity-50")}>
<GripVertical className="hidden size-3.5 shrink-0 text-muted-foreground/50 sm:block" aria-hidden />
<span className="w-4 text-center text-[11px] font-medium text-muted-foreground">{index + 1}</span>
<div className="flex size-10 shrink-0 items-center justify-center overflow-hidden rounded-md bg-muted"><AssetPreview asset={asset} onMissing={onMissing} compact /></div>
<div className="min-w-0 flex-1"><p className="truncate text-xs font-medium">{asset.name}</p><p className="text-[10px] capitalize text-muted-foreground">{asset.kind}{asset.missing ? " · unavailable" : ""}</p></div>
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Move ${asset.name} up`} disabled={disabled || index === 0} onClick={() => onMove(index, index - 1)}><ArrowUp className="size-3.5" /></Button>
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Move ${asset.name} down`} disabled={disabled || index === conditioning.length - 1} onClick={() => onMove(index, index + 1)}><ArrowDown className="size-3.5" /></Button>
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Unselect ${asset.name}`} disabled={disabled} onClick={() => onUnselect(index)}><X className="size-3.5" /></Button>
</li>
);
})}
</ol>
) : (
<div className="flex items-center gap-3 rounded-xl border border-dashed border-input px-4 py-4 text-muted-foreground"><ImagePlus className="size-6 shrink-0 opacity-50" /><p className="text-xs">Add images, video, or audio from your asset library.<br /><span className="text-[11px] opacity-75">At least one image or video is required.</span></p></div>
)}
<p className="mt-2 text-[10px] text-muted-foreground">{conditioning.length}/12 selected · up to 9 images, 3 videos, 3 audio clips</p>
</>
)}
{assets.length > 0 && !locked && (
<div className="mt-3 border-t border-border/60 pt-2">
<button type="button" className="flex w-full items-center justify-between py-1 text-[11px] font-medium text-muted-foreground" aria-expanded={libraryOpen} onClick={() => setLibraryOpen(!libraryOpen)}><span>Asset library · {assets.length}</span><span>{libraryOpen ? "Hide" : "Show"}</span></button>
{libraryOpen && <div className="mt-2 grid grid-cols-2 gap-2 sm:grid-cols-3">
{assets.map((asset) => {
const selected = conditioning.some((item) => item.asset_id === asset.asset_id);
return (
<div key={asset.asset_id} className={cn("overflow-hidden rounded-lg border bg-background/40", selected ? "border-sky-400/70" : "border-input")}>
<div className="flex h-20 items-center justify-center overflow-hidden bg-muted/50"><AssetPreview asset={asset} onMissing={onMissing} /></div>
<div className="flex items-center gap-1 p-1.5">
<div className="min-w-0 flex-1"><p title={asset.name} className="truncate text-[10px] font-medium">{asset.name}</p><p className="text-[9px] capitalize text-muted-foreground">{asset.missing ? "Upload again" : `${asset.kind} · ${(asset.size / 1024 / 1024).toFixed(1)} MB`}</p></div>
{mode === "ref2va" && <Button type="button" variant="ghost" size="icon-sm" className="size-7" aria-label={`Add ${asset.name} as reference`} disabled={disabled || selected || asset.missing || conditioning.length >= 12} onClick={() => onAssign(asset.asset_id, "reference")}>{selected ? <Check className="size-3.5 text-sky-500" /> : <Plus className="size-3.5" />}</Button>}
<Button type="button" variant="ghost" size="icon-sm" className="size-7 text-muted-foreground" aria-label={`Remove asset ${asset.name}`} disabled={disabled} onClick={() => onRemove(asset.asset_id)}><Trash2 className="size-3" /></Button>
</div>
</div>
);
})}
</div>}
</div>
)}
</div>
{!locked && <p className="px-4 pb-2 text-[10px] text-muted-foreground">Images ≤15 MiB / 16 MP{mode === "ref2va" ? " / 1:4–4:1 aspect ratio" : ""} · video/audio ≤100 MiB / 30 sec · video up to 4K · mono/stereo audio</p>}
{(error || validationNotice) && <p role={error ? "alert" : "status"} className={cn("border-t border-border/60 px-4 py-2 text-[11px]", error ? "bg-rose-500/5 text-rose-600 dark:text-rose-300" : "bg-amber-500/5 text-amber-700 dark:text-amber-300")}>{error || validationNotice}</p>}
</section>
);
}
@@ -1,155 +0,0 @@
import { render, screen, waitFor, within } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { afterAll, beforeAll, describe, expect, it, vi } from "vitest";
import ChatBar from "./ChatBar";
// JSDOM does not implement the pointer/scroll APIs used by the Radix popup.
const domPolyfills = {
hasPointerCapture: () => false,
releasePointerCapture: () => {},
scrollIntoView: () => {},
};
const originalDescriptors = new Map<string, PropertyDescriptor | undefined>();
beforeAll(() => {
for (const [name, implementation] of Object.entries(domPolyfills)) {
originalDescriptors.set(name, Object.getOwnPropertyDescriptor(HTMLElement.prototype, name));
Object.defineProperty(HTMLElement.prototype, name, { configurable: true, value: implementation });
}
vi.stubGlobal("PointerEvent", MouseEvent);
});
afterAll(() => {
for (const [name, descriptor] of originalDescriptors) {
if (descriptor) Object.defineProperty(HTMLElement.prototype, name, descriptor);
else Reflect.deleteProperty(HTMLElement.prototype, name);
}
vi.unstubAllGlobals();
});
describe("ChatBar generation mode selection", () => {
it("places Mode and the prompt input inside the same composer", () => {
render(<ChatBar />);
const composer = screen.getByRole("group", { name: "Prompt composer" });
expect(within(composer).getByText("Mode", { exact: true })).toBeVisible();
expect(within(composer).getByRole("combobox", { name: "Generation mode" }))
.toBeVisible();
expect(within(composer).getByRole("textbox", { name: "Continuation prompt" }))
.toBeVisible();
expect(screen.queryByText("Generation mode", { exact: true }))
.not.toBeInTheDocument();
});
it("shows only mode abbreviations and keeps explanations in the tooltip", async () => {
const user = userEvent.setup();
render(<ChatBar />);
expect(screen.getByRole("combobox", { name: "Generation mode" }))
.toHaveAttribute("title", "Text to video + audio. Start with a text prompt; no reference asset is required.");
expect(screen.queryByText("Start with a text prompt; no reference asset is required."))
.not.toBeInTheDocument();
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
await user.click(screen.getByRole("combobox", { name: "Generation mode" }));
const menu = await screen.findByRole("listbox");
expect(within(menu).getAllByRole("option").map((option) => option.textContent))
.toEqual(["T2VA", "FL2VA", "Ref2VA"]);
});
it("defaults to T2VA and reports a selected mode", async () => {
const onGenerationModeChange = vi.fn();
const user = userEvent.setup();
render(
<ChatBar
canJoinSession
continuationDraft="A lighthouse in a storm"
onGenerationModeChange={onGenerationModeChange}
/>,
);
const modeSelect = screen.getByRole("combobox", { name: "Generation mode" });
expect(modeSelect).toHaveTextContent("T2VA");
await user.click(modeSelect);
await user.click(await screen.findByRole("option", { name: "Ref2VA" }));
expect(onGenerationModeChange).toHaveBeenCalledWith("ref2va");
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
});
it("hides mode selection after generation starts", () => {
render(<ChatBar sessionStarted />);
expect(screen.queryByRole("combobox", { name: "Generation mode" }))
.not.toBeInTheDocument();
expect(within(screen.getByRole("group", { name: "Prompt composer" }))
.getByRole("textbox", { name: "Continuation prompt" })).toBeVisible();
});
it("disables mode selection and prompt editing while generation is busy", async () => {
const onGenerationModeChange = vi.fn();
const user = userEvent.setup();
render(<ChatBar isGenerating onGenerationModeChange={onGenerationModeChange} />);
const composer = screen.getByRole("group", { name: "Prompt composer" });
const modeSelect = within(composer).getByRole("combobox", { name: "Generation mode" });
expect(modeSelect).toBeDisabled();
expect(within(composer).getByRole("textbox", { name: "Continuation prompt" }))
.toBeDisabled();
await user.click(modeSelect);
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
expect(onGenerationModeChange).not.toHaveBeenCalled();
});
it("still submits the prompt with Enter from the combined composer", async () => {
const onGenerate = vi.fn();
const user = userEvent.setup();
render(<ChatBar canJoinSession continuationDraft="A lighthouse in a storm" onGenerate={onGenerate} />);
const input = within(screen.getByRole("group", { name: "Prompt composer" }))
.getByRole("textbox", { name: "Continuation prompt" });
await user.click(input);
await user.keyboard("{Enter}");
expect(onGenerate).toHaveBeenCalledTimes(1);
});
it("disables unsupported modes and labels mock playback", async () => {
const user = userEvent.setup();
render(<ChatBar supportedGenerationModes={["t2va"]} mockRuntime />);
expect(screen.getByText(/No AI model is generating/)).toBeInTheDocument();
await user.click(screen.getByRole("combobox", { name: "Generation mode" }));
expect(await screen.findByRole("option", { name: "FL2VA" })).toHaveAttribute("aria-disabled", "true");
expect(screen.getByRole("option", { name: "Ref2VA" })).toHaveAttribute("aria-disabled", "true");
expect(screen.getByRole("option", { name: "Ref2VA" }))
.toHaveAttribute("title", "References to video + audio (unavailable on this runtime)");
});
it("closes the menu with Escape and restores focus to Mode", async () => {
const onGenerationModeChange = vi.fn();
const user = userEvent.setup();
render(<ChatBar onGenerationModeChange={onGenerationModeChange} />);
const modeSelect = screen.getByRole("combobox", { name: "Generation mode" });
await user.click(modeSelect);
await screen.findByRole("listbox");
await user.keyboard("{Escape}");
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
await waitFor(() => expect(modeSelect).toHaveFocus());
expect(onGenerationModeChange).not.toHaveBeenCalled();
});
it("supports choosing a mode with the keyboard", async () => {
const onGenerationModeChange = vi.fn();
const user = userEvent.setup();
render(<ChatBar onGenerationModeChange={onGenerationModeChange} />);
await user.click(screen.getByRole("textbox", { name: "Continuation prompt" }));
await user.tab();
expect(screen.getByRole("combobox", { name: "Generation mode" })).toHaveFocus();
await user.keyboard("{ArrowDown}");
await waitFor(() => expect(screen.getByRole("option", { name: "T2VA" })).toHaveFocus());
await user.keyboard("{ArrowDown}");
await waitFor(() => expect(screen.getByRole("option", { name: "FL2VA" })).toHaveFocus());
await user.keyboard("{Enter}");
expect(onGenerationModeChange).toHaveBeenCalledWith("fl2va");
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
});
});
+41 -255
View File
@@ -2,18 +2,13 @@
import React, { useRef, useState, useCallback, useEffect } from "react";
import Image from "next/image";
import { Film, ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
import { ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
import { Button } from "@/components/ui/button";
import { Select, SelectContent, SelectItem, SelectTrigger, SelectValue } from "@/components/ui/select";
import LeaveSessionModal, { shouldShowLeaveWarning } from "@/components/LeaveSessionModal";
import PresetQuickLaunchRail from "@/components/creation/PresetQuickLaunchRail";
import SessionCreationConfigPills, { type SessionCreationConfig } from "@/components/creation/SessionCreationConfigPills";
import SpeechToTextButton from "@/components/SpeechToTextButton";
import {
DEFAULT_GENERATION_MODE,
GENERATION_MODES,
getGenerationMode,
isGenerationMode,
type GenerationMode,
} from "@/lib/generationMode";
import type { AspectRatioId, CreationModeId, CreationModelId, ResolutionId } from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
const PROMPT_MAX_LENGTH = 500;
@@ -30,12 +25,6 @@ interface Props {
sessionNotice?: string;
projectResetPending?: boolean;
viewingReadOnly?: boolean;
generationMode?: GenerationMode;
supportedGenerationModes?: readonly GenerationMode[];
generationInputsValid?: boolean;
capabilityNotice?: string;
mockRuntime?: boolean;
conditioningPanel?: React.ReactNode;
onPresetGenerate?: (presetId: string) => void;
onContinuationInput?: (e: React.ChangeEvent<HTMLTextAreaElement>) => void;
onContinuationKeydown?: (e: React.KeyboardEvent<HTMLTextAreaElement>) => void;
@@ -44,9 +33,15 @@ interface Props {
onLeave?: () => void;
onStartNewProject?: () => void;
onBackFromViewing?: () => void;
onGenerationModeChange?: (mode: GenerationMode) => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
sessionCreationConfig?: SessionCreationConfig | null;
configPillsReadOnly?: boolean;
onSessionModelChange?: (modelId: CreationModelId) => void;
onSessionModeChange?: (modeId: CreationModeId) => void;
onSessionAspectRatioChange?: (aspectRatio: AspectRatioId) => void;
onSessionResolutionChange?: (resolution: ResolutionId) => void;
onSessionDurationChange?: (durationSec: number) => void;
}
export default function ChatBar({
@@ -61,12 +56,6 @@ export default function ChatBar({
sessionNotice = "",
projectResetPending = false,
viewingReadOnly = false,
generationMode = DEFAULT_GENERATION_MODE,
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
generationInputsValid = true,
capabilityNotice = "",
mockRuntime = false,
conditioningPanel,
onPresetGenerate = () => {},
onContinuationInput = () => {},
onContinuationKeydown = () => {},
@@ -75,9 +64,15 @@ export default function ChatBar({
onLeave = () => {},
onStartNewProject = () => {},
onBackFromViewing = () => {},
onGenerationModeChange = () => {},
onSpeechTranscript,
onSpeechInterimChange,
sessionCreationConfig = null,
configPillsReadOnly = false,
onSessionModelChange,
onSessionModeChange,
onSessionAspectRatioChange,
onSessionResolutionChange,
onSessionDurationChange,
}: Props) {
const [sttBusy, setSttBusy] = useState(false);
const [leaveModalOpen, setLeaveModalOpen] = useState(false);
@@ -88,135 +83,11 @@ export default function ChatBar({
: isBusy
? "Generating video\u2026"
: !sessionStarted
? "What video are you imagining?"
? "Describe your video"
: "What do you want to edit?";
const actionLabel = !sessionStarted ? "Generate" : "Rewrite rollout";
const selectedGenerationMode = getGenerationMode(generationMode);
const inputRef = useRef<HTMLTextAreaElement>(null);
const scrollRef = useRef<HTMLDivElement>(null);
const [canScrollLeft, setCanScrollLeft] = useState(false);
const [canScrollRight, setCanScrollRight] = useState(false);
const [presetRailDragging, setPresetRailDragging] = useState(false);
const presetDragStateRef = useRef({
pointerId: null as number | null,
startX: 0,
startScrollLeft: 0,
moved: false,
});
const suppressPresetClickRef = useRef(false);
const updateScrollState = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setCanScrollLeft(el.scrollLeft > 2);
setCanScrollRight(el.scrollLeft + el.clientWidth < el.scrollWidth - 2);
}, []);
const handlePresetWheel = useCallback(
(event: React.WheelEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
const dominantDelta = Math.abs(event.deltaX) > Math.abs(event.deltaY)
? event.deltaX
: event.deltaY;
if (!dominantDelta) return;
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
const nextScrollLeft = Math.min(
Math.max(el.scrollLeft + dominantDelta, 0),
maxScrollLeft,
);
if (nextScrollLeft === el.scrollLeft) return;
event.preventDefault();
el.scrollLeft = nextScrollLeft;
updateScrollState();
},
[updateScrollState],
);
const finishPresetDrag = useCallback(() => {
presetDragStateRef.current = {
pointerId: null,
startX: 0,
startScrollLeft: 0,
moved: false,
};
setPresetRailDragging(false);
}, []);
const handlePresetPointerDown = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (event.pointerType !== "mouse" || event.button !== 0) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
suppressPresetClickRef.current = false;
presetDragStateRef.current = {
pointerId: event.pointerId,
startX: event.clientX,
startScrollLeft: el.scrollLeft,
moved: false,
};
},
[],
);
const handlePresetPointerMove = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
const dragState = presetDragStateRef.current;
if (!el || dragState.pointerId !== event.pointerId) return;
const deltaX = event.clientX - dragState.startX;
if (!dragState.moved && Math.abs(deltaX) > 4) {
dragState.moved = true;
suppressPresetClickRef.current = true;
setPresetRailDragging(true);
el.setPointerCapture?.(event.pointerId);
}
if (!dragState.moved) return;
event.preventDefault();
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
el.scrollLeft = Math.min(
Math.max(dragState.startScrollLeft - deltaX, 0),
maxScrollLeft,
);
updateScrollState();
},
[updateScrollState],
);
const handlePresetPointerUp = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el || presetDragStateRef.current.pointerId !== event.pointerId) return;
if (el.hasPointerCapture?.(event.pointerId)) {
el.releasePointerCapture(event.pointerId);
}
finishPresetDrag();
},
[finishPresetDrag],
);
const handlePresetClickCapture = useCallback(
(event: React.MouseEvent<HTMLDivElement>) => {
if (!suppressPresetClickRef.current) return;
suppressPresetClickRef.current = false;
event.preventDefault();
event.stopPropagation();
},
[],
);
useEffect(() => {
updateScrollState();
}, [storyPresets, updateScrollState]);
useEffect(() => {
if (!isBusy && !sttBusy && !window.matchMedia("(pointer: coarse)").matches) {
@@ -262,7 +133,7 @@ export default function ChatBar({
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-6 py-4 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">View-only project</p>
<p className="max-w-md text-xs text-muted-foreground">This saved project is available for playback. Start a new project to create more videos.</p>
<p className="max-w-md text-xs text-muted-foreground">Sessions are limited to 5 minutes. Start a new project to keep creating.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onBackFromViewing} variant="outline" size="sm" className="gap-1.5 rounded-full px-4">
@@ -284,7 +155,7 @@ export default function ChatBar({
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-8 py-5 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">Session ended</p>
<p className="max-w-xs text-xs text-muted-foreground">The runtime session has ended. Your saved videos remain available. Start a new project to continue creating.</p>
<p className="max-w-xs text-xs text-muted-foreground">Sessions are limited to 5 minutes. Start a new project to continue.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onStartNewProject} size="sm" className="rounded-full px-5">
@@ -303,57 +174,8 @@ export default function ChatBar({
return (
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
{storyPresets.length > 0 && !sessionStarted && generationMode === "t2va" && (
<div className={cn("relative transition-opacity duration-200", isGenerating && "pointer-events-none opacity-40")}>
<div
ref={scrollRef}
onScroll={updateScrollState}
onWheel={handlePresetWheel}
onPointerDown={handlePresetPointerDown}
onPointerMove={handlePresetPointerMove}
onPointerUp={handlePresetPointerUp}
onPointerCancel={handlePresetPointerUp}
onLostPointerCapture={finishPresetDrag}
onClickCapture={handlePresetClickCapture}
className={cn(
"scrollbar-hidden flex gap-3 overflow-x-auto px-1 select-none",
presetRailDragging ? "cursor-grabbing" : "cursor-grab",
)}
>
{storyPresets.map((preset) => (
<button
key={preset.id}
type="button"
disabled={isBusy || !generationInputsValid}
onClick={() => onPresetGenerate(preset.id)}
className="flex flex-col sm:flex-row items-start gap-1.5 shrink-0 rounded-xl border p-2.5 text-left backdrop-blur-sm transition-colors max-w-42 sm:max-w-[215px] border-input bg-card/80 text-muted-foreground hover:bg-slate-200/60 hover:border-slate-400 hover:text-slate-700 dark:bg-slate-800/80 dark:text-slate-300 dark:hover:bg-slate-700/50 dark:hover:border-slate-500 dark:hover:text-slate-200"
>
<Film className="mt-0.5 size-4 shrink-0 opacity-60" />
<span className="flex flex-col gap-1 min-w-0">
<span className="text-[14px] font-medium line-clamp-1">{preset.label}</span>
{preset.description && <span className="text-xs leading-tight opacity-70 line-clamp-3 sm:line-clamp-2">{preset.description}</span>}
</span>
</button>
))}
</div>
<div
className={cn("pointer-events-none absolute inset-y-0 left-0 w-8 bg-background transition-opacity duration-150", canScrollLeft ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to right, black, transparent)", WebkitMaskImage: "linear-gradient(to right, black, transparent)" }}
aria-hidden="true"
/>
<div
className={cn("pointer-events-none absolute inset-y-0 right-0 w-8 bg-background transition-opacity duration-150", canScrollRight ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to left, black, transparent)", WebkitMaskImage: "linear-gradient(to left, black, transparent)" }}
aria-hidden="true"
/>
</div>
)}
{mockRuntime && (
<p role="status" className="rounded-xl border border-violet-500/25 bg-violet-500/10 px-4 py-2 text-center text-xs text-violet-700 dark:text-violet-300">
Demo runtime · Sample playback only. No AI model is generating this video.
</p>
{!sessionStarted && (
<PresetQuickLaunchRail storyPresets={storyPresets} disabled={isGenerating} onPresetGenerate={onPresetGenerate} />
)}
{sessionNotice && (
@@ -371,22 +193,30 @@ export default function ChatBar({
{projectResetPending && sessionStarted && (
<div className="rounded-xl border border-sky-500/20 bg-sky-500/10 px-4 py-2.5 text-center text-xs text-sky-700 dark:text-sky-300">
Starting a new project after the current shot finishes. Your GPU session stays active.
Starting a new project when this shot finishes. GPU session stays open.
</div>
)}
{sessionStarted && <p className="px-2 text-center text-[11px] text-muted-foreground">{selectedGenerationMode.label} · Mode and reference inputs are locked for this project.</p>}
{conditioningPanel}
<div
role="group"
aria-label="Prompt composer"
className={cn(
"flex min-w-0 flex-col gap-2 rounded-3xl border p-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
"flex min-w-0 flex-col gap-2 rounded-4xl border py-2.5 pl-5 pr-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
isBusy ? "border-input/60 bg-card/40" : "border-input bg-card/65",
)}
>
<textarea
{sessionStarted && sessionCreationConfig && (
<SessionCreationConfigPills
{...sessionCreationConfig}
disabled={isBusy}
readOnly={configPillsReadOnly}
onModelChange={onSessionModelChange}
onModeChange={onSessionModeChange}
onAspectRatioChange={onSessionAspectRatioChange}
onResolutionChange={onSessionResolutionChange}
onDurationChange={onSessionDurationChange}
/>
)}
<div className="flex min-w-0 items-center gap-1.5">
<textarea
ref={inputRef}
id="continuation-prompt"
aria-label="Continuation prompt"
@@ -398,53 +228,10 @@ export default function ChatBar({
disabled={isBusy || sttBusy}
rows={1}
className={cn(
"w-full min-w-0 resize-none bg-transparent px-2 py-1 text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
"min-w-0 flex-1 resize-none bg-transparent text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
(isBusy || sttBusy) && "cursor-not-allowed opacity-50",
)}
/>
<div className="flex min-w-0 items-center gap-1.5">
{!sessionStarted && (
<div className="flex shrink-0 items-center gap-1 pl-2">
<label htmlFor="generation-mode" className="cursor-pointer text-xs font-medium text-muted-foreground">
Mode
</label>
<Select
value={generationMode}
disabled={isBusy || sttBusy}
onValueChange={(value) => {
if (isGenerationMode(value)) onGenerationModeChange(value);
}}
>
<SelectTrigger
id="generation-mode"
aria-label="Generation mode"
title={`${selectedGenerationMode.name}. ${selectedGenerationMode.description}`}
className="h-8 w-24 cursor-pointer rounded-lg border-0 bg-transparent px-2 py-1 text-xs font-medium shadow-none hover:bg-muted/60 data-[state=open]:bg-muted/80 [&>svg]:size-3 [&>svg]:transition-transform [&[data-state=open]>svg]:rotate-180"
>
<SelectValue />
</SelectTrigger>
<SelectContent
side="bottom"
align="start"
sideOffset={6}
className="min-w-36 rounded-2xl border-input/70 bg-card/95 shadow-xl backdrop-blur-xl"
>
{GENERATION_MODES.map((mode) => (
<SelectItem
key={mode.id}
value={mode.id}
disabled={!supportedGenerationModes.includes(mode.id)}
title={supportedGenerationModes.includes(mode.id) ? mode.name : `${mode.name} (unavailable on this runtime)`}
className="cursor-pointer rounded-xl text-xs transition-colors data-[state=checked]:bg-muted/80 data-[state=checked]:font-semibold [&_svg]:size-3.5 [&_svg]:text-foreground"
>
{mode.label}
</SelectItem>
))}
</SelectContent>
</Select>
</div>
)}
<div className="flex-1" />
/>
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
{!sessionStarted ? (
<Button
@@ -476,7 +263,6 @@ export default function ChatBar({
)}
</div>
</div>
{!sessionStarted && capabilityNotice && <p className="px-2 text-[11px] text-amber-700 dark:text-amber-300">{capabilityNotice}</p>}
<p className="px-2 text-center text-[11px] text-muted-foreground">
LLM powered by{" "}
<a
@@ -0,0 +1,95 @@
"use client";
import { Fragment, useEffect, useRef } from "react";
const HERO_WAVE_LIGHT = ["#2A4A98", "#4878E5", "#6FA0F2", "#B0BCC8", "#E8D99E", "#D8C844", "#C2A620"];
const HERO_WAVE_DARK = ["#143468", "#1E58B8", "#3892F0", "#80B8E8", "#B8D0EA", "#E2D498", "#DABB50"];
const HERO_TEXT = "Direct scenes in seconds";
export default function HeroTagline() {
const ref = useRef<HTMLHeadingElement>(null);
useEffect(() => {
const el = ref.current;
if (!el) return;
let rafId = 0;
function play() {
const chars = el!.querySelectorAll<HTMLSpanElement>("[data-char]");
if (!chars.length) return;
cancelAnimationFrame(rafId);
const isDark = document.documentElement.classList.contains("dark");
const colors = isDark ? HERO_WAVE_DARK : HERO_WAVE_LIGHT;
const waveLen = 10;
const total = chars.length + waveLen;
const duration = 1200;
const maxBlur = 3.5;
const start = performance.now();
function tick() {
const t = Math.min((performance.now() - start) / duration, 1);
const pos = t * total;
chars.forEach((ch, i) => {
const rel = pos - i;
if (rel >= 0 && rel < waveLen) {
const norm = rel / waveLen;
const ci = Math.floor(norm * colors.length);
ch.style.color = colors[Math.min(colors.length - 1, ci)];
let blur = 0;
if (norm < 0.25) {
blur = maxBlur * (1 - norm / 0.25);
} else if (norm > 0.75) {
blur = maxBlur * ((norm - 0.75) / 0.25);
}
ch.style.filter = blur > 0.1 ? `blur(${blur.toFixed(1)}px)` : "";
} else {
ch.style.color = "";
ch.style.filter = "";
}
});
if (t < 1) {
rafId = requestAnimationFrame(tick);
} else {
chars.forEach((ch) => {
ch.style.color = "";
ch.style.filter = "";
});
}
}
rafId = requestAnimationFrame(tick);
}
const initialDelay = setTimeout(play, 400);
const interval = setInterval(play, 5000);
return () => {
clearTimeout(initialDelay);
clearInterval(interval);
cancelAnimationFrame(rafId);
};
}, []);
return (
<h1 ref={ref} className="text-balance text-center text-3xl font-medium text-[#343537] dark:text-[#FAFAFB] sm:text-4xl">
{HERO_TEXT.split(" ").map((word, wi) => (
<Fragment key={wi}>
{wi > 0 && (
<span data-char className="transition-[color,filter] duration-150">
{" "}
</span>
)}
<span className="inline-flex">
{word.split("").map((char, ci) => (
<span key={ci} data-char className="inline-block transition-[color,filter] duration-150">
{char}
</span>
))}
</span>
</Fragment>
))}
</h1>
);
}
@@ -33,7 +33,7 @@ export default function SessionTimeoutModal({
Session ended
</h2>
<p className="text-sm text-muted-foreground">
This project reached the runtime session limit. Your latest video stays on screen, and the project is being kept in the archive so you can come back to it.
This project hit the current 5-minute session limit. Your latest video stays on screen, and the project is being kept in the archive so you can come back to it.
</p>
</div>
<p className="text-sm text-muted-foreground">
@@ -0,0 +1,66 @@
"use client";
import React from "react";
import { FolderOpen, Home, Sparkles } from "lucide-react";
import { cn } from "@/lib/utils";
export type AppNavSection = "explore" | "create" | "assets";
interface AppNavRailProps {
activeSection?: AppNavSection;
onSectionChange?: (section: AppNavSection) => void;
onOpenProjects?: () => void;
className?: string;
}
const NAV_ITEMS: Array<{ id: AppNavSection; label: string; icon: typeof Home }> = [
{ id: "explore", label: "Explore", icon: Home },
{ id: "create", label: "Create", icon: Sparkles },
{ id: "assets", label: "Assets", icon: FolderOpen },
];
export default function AppNavRail({
activeSection = "create",
onSectionChange = () => {},
onOpenProjects,
className,
}: AppNavRailProps) {
return (
<aside
className={cn(
"hidden shrink-0 flex-col items-center gap-2 border-r border-border/40 bg-background/30 px-2.5 py-5 lg:flex",
className,
)}
aria-label="Primary navigation"
>
{NAV_ITEMS.map((item) => {
const Icon = item.icon;
const isActive = item.id === activeSection;
return (
<button
key={item.id}
type="button"
aria-label={item.label}
aria-current={isActive ? "page" : undefined}
onClick={() => {
if (item.id === "assets") {
onOpenProjects?.();
}
onSectionChange(item.id);
}}
className={cn(
"studio-control studio-control-press flex w-[4.5rem] min-h-11 flex-col items-center gap-1 rounded-xl px-2 py-2.5 text-[10px] font-medium tracking-wide",
isActive
? "bg-secondary/90 text-foreground shadow-sm ring-1 ring-border/60"
: "text-muted-foreground hover-capable:hover:bg-secondary/50 hover-capable:hover:text-foreground",
)}
>
<Icon className={cn("size-[18px]", isActive && "text-accent-blue")} />
{item.label}
</button>
);
})}
</aside>
);
}
@@ -0,0 +1,24 @@
"use client";
import React from "react";
import { cn } from "@/lib/utils";
export default function ConfigPill({
children,
className,
...props
}: React.ButtonHTMLAttributes<HTMLButtonElement>) {
return (
<button
type="button"
className={cn(
"studio-control studio-control-press studio-hover-surface inline-flex h-9 min-h-9 shrink-0 items-center gap-1 rounded-full border border-border/50 bg-background/80 px-2.5 text-[11px] font-medium text-foreground/90",
className,
)}
{...props}
>
{children}
</button>
);
}
@@ -0,0 +1,466 @@
"use client";
import React, { useMemo, useRef, useState } from "react";
import { ArrowUp, Box, ChevronDown, Clock, Monitor, Wand2 } from "lucide-react";
import ConfigPill from "@/components/creation/ConfigPill";
import HeroTagline from "@/components/HeroTagline";
import ReferenceUploadSlot from "@/components/creation/ReferenceUploadSlot";
import { Button } from "@/components/ui/button";
import {
DropdownMenu,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuTrigger,
} from "@/components/ui/dropdown-menu";
import { Popover, PopoverContent, PopoverTrigger } from "@/components/ui/popover";
import { Slider } from "@/components/ui/slider";
import SpeechToTextButton from "@/components/SpeechToTextButton";
import {
ASPECT_RATIOS,
CREATION_MODELS,
CREATION_MODES,
RESOLUTIONS,
UNSUPPORTED_CREATION_MODES,
UNSUPPORTED_RESOLUTIONS,
modeRequiresReference,
modeUsesDualFrames,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type MentionOption,
type ResolutionId,
formatDurationLabel,
formatResolutionLabel,
} from "@/lib/creationConfig";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
isSupportedCreationMode,
isSupportedResolution,
resolveModelCapabilities,
unsupportedModeNotice,
type LobbyCreationCapabilities,
} from "@/lib/creationCapabilities";
import { cn } from "@/lib/utils";
const PROMPT_MAX_LENGTH = 500;
interface CreationComposerProps {
value: string;
disabled?: boolean;
isGenerating?: boolean;
canSubmit?: boolean;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referencePreviewUrl?: string | null;
firstFramePreviewUrl?: string | null;
lastFramePreviewUrl?: string | null;
mentionOptions?: MentionOption[];
onValueChange: (value: string) => void;
onSubmit: () => void;
onKeyDown?: (event: React.KeyboardEvent<HTMLTextAreaElement>) => void;
onModelChange: (modelId: CreationModelId) => void;
onModeChange: (modeId: CreationModeId) => void;
onAspectRatioChange: (aspectRatio: AspectRatioId) => void;
onResolutionChange: (resolution: ResolutionId) => void;
onDurationChange: (durationSec: number) => void;
onReferenceSelect?: (file: File | null) => void;
onFirstFrameSelect?: (file: File | null) => void;
onLastFrameSelect?: (file: File | null) => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
capabilities?: LobbyCreationCapabilities;
}
export default function CreationComposer({
value,
disabled = false,
isGenerating = false,
canSubmit = false,
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
referencePreviewUrl = null,
firstFramePreviewUrl = null,
lastFramePreviewUrl = null,
mentionOptions = [],
onValueChange,
onSubmit,
onKeyDown,
onModelChange,
onModeChange,
onAspectRatioChange,
onResolutionChange,
onDurationChange,
onReferenceSelect,
onFirstFrameSelect,
onLastFrameSelect,
onSpeechTranscript,
onSpeechInterimChange,
capabilities = resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, modelId),
}: CreationComposerProps) {
const inputRef = useRef<HTMLTextAreaElement>(null);
const [sttBusy, setSttBusy] = useState(false);
const [mentionQuery, setMentionQuery] = useState("");
const [mentionOpen, setMentionOpen] = useState(false);
const [mentionStart, setMentionStart] = useState<number | null>(null);
const availableModels = useMemo(
() => CREATION_MODELS.filter((model) => capabilities.model_ids.includes(model.id)),
[capabilities.model_ids],
);
const availableModes = useMemo(
() => CREATION_MODES.filter((mode) => isSupportedCreationMode(mode.id, capabilities)),
[capabilities],
);
const unavailableModes = useMemo(
() =>
UNSUPPORTED_CREATION_MODES.filter(
(mode) => unsupportedModeNotice(mode.id, capabilities) !== null,
),
[capabilities],
);
const availableAspectRatios = useMemo(
() => ASPECT_RATIOS.filter((ratio) => capabilities.aspect_ratios.includes(ratio)),
[capabilities.aspect_ratios],
);
const availableResolutions = useMemo(
() => RESOLUTIONS.filter((item) => isSupportedResolution(item, capabilities)),
[capabilities],
);
const unavailableResolutions = useMemo(
() => UNSUPPORTED_RESOLUTIONS.filter((item) => !isSupportedResolution(item, capabilities)),
[capabilities],
);
const durationMin = capabilities.duration_sec[0] ?? 5;
const durationMax = capabilities.duration_sec[capabilities.duration_sec.length - 1] ?? 15;
const selectedModel = availableModels.find((model) => model.id === modelId) ?? availableModels[0];
const selectedMode = availableModes.find((mode) => mode.id === modeId) ?? availableModes[0];
const usesDualFrames = modeUsesDualFrames(modeId);
const requiresReference = modeRequiresReference(modeId);
const referenceMissing = requiresReference && !referencePreviewUrl;
const submitDisabled = !canSubmit || disabled || isGenerating || !value.trim() || referenceMissing;
const filteredMentions = useMemo(() => {
const query = mentionQuery.trim().toLowerCase();
if (!query) return mentionOptions.slice(0, 6);
return mentionOptions
.filter((option) => option.label.toLowerCase().includes(query) || option.description?.toLowerCase().includes(query))
.slice(0, 6);
}, [mentionOptions, mentionQuery]);
function autoResize() {
const el = inputRef.current;
if (!el) return;
el.style.height = "auto";
const lineHeight = parseFloat(getComputedStyle(el).lineHeight) || 22;
const maxHeight = lineHeight * 4;
el.style.height = `${Math.min(el.scrollHeight, maxHeight)}px`;
el.style.overflowY = el.scrollHeight > maxHeight ? "auto" : "hidden";
}
function updateMentionState(nextValue: string, cursorPosition: number) {
const beforeCursor = nextValue.slice(0, cursorPosition);
const atIndex = beforeCursor.lastIndexOf("@");
if (atIndex === -1 || (atIndex > 0 && !/\s/.test(beforeCursor[atIndex - 1] ?? ""))) {
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
return;
}
const query = beforeCursor.slice(atIndex + 1);
if (/\s/.test(query)) {
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
return;
}
setMentionStart(atIndex);
setMentionQuery(query);
setMentionOpen(true);
}
function insertMention(option: MentionOption) {
if (mentionStart === null) return;
const before = value.slice(0, mentionStart);
const after = value.slice(inputRef.current?.selectionStart ?? value.length);
const mentionText = `@${option.label} `;
const nextValue = `${before}${mentionText}${after}`.slice(0, PROMPT_MAX_LENGTH);
onValueChange(nextValue);
setMentionOpen(false);
setMentionStart(null);
setMentionQuery("");
requestAnimationFrame(() => {
const el = inputRef.current;
if (!el) return;
const cursor = before.length + mentionText.length;
el.focus();
el.setSelectionRange(cursor, cursor);
autoResize();
});
}
function handleInputChange(event: React.ChangeEvent<HTMLTextAreaElement>) {
const nextValue = event.target.value.slice(0, PROMPT_MAX_LENGTH);
onValueChange(nextValue);
updateMentionState(nextValue, event.target.selectionStart ?? nextValue.length);
requestAnimationFrame(autoResize);
}
function handleKeyDown(event: React.KeyboardEvent<HTMLTextAreaElement>) {
if (mentionOpen && filteredMentions.length > 0) {
if (event.key === "Tab" || (event.key === "Enter" && !event.shiftKey)) {
event.preventDefault();
insertMention(filteredMentions[0]);
return;
}
if (event.key === "Escape") {
setMentionOpen(false);
return;
}
}
onKeyDown?.(event);
}
return (
<section className="mx-auto flex w-full max-w-3xl flex-col gap-5">
<HeroTagline />
<div className="rounded-[32px] border border-border/40 bg-secondary/95 p-4 shadow-[0_24px_80px_-32px_rgba(0,0,0,0.72)] backdrop-blur-xl sm:p-5">
<div className="flex gap-3.5">
{usesDualFrames ? (
<div className="flex shrink-0 gap-2">
<ReferenceUploadSlot
label="Asset"
sublabel="First"
previewUrl={firstFramePreviewUrl}
disabled={disabled}
onSelect={onFirstFrameSelect}
/>
<ReferenceUploadSlot
label="Asset"
sublabel="Last"
previewUrl={lastFramePreviewUrl}
disabled={disabled}
onSelect={onLastFrameSelect}
/>
</div>
) : (
<ReferenceUploadSlot
label="Reference"
previewUrl={referencePreviewUrl}
required={requiresReference}
optional={!requiresReference}
disabled={disabled}
onSelect={onReferenceSelect}
/>
)}
<div className="relative min-w-0 flex-1">
<textarea
ref={inputRef}
id="continuation-prompt"
aria-label="Continuation prompt"
value={value}
onChange={handleInputChange}
onKeyDown={handleKeyDown}
onClick={(event) => updateMentionState(value, event.currentTarget.selectionStart ?? value.length)}
placeholder="Describe your video or mention elements"
disabled={disabled || sttBusy}
rows={3}
className={cn(
"min-h-[92px] w-full resize-none bg-transparent px-0.5 text-base leading-6 text-foreground outline-none placeholder:text-muted-foreground/80 sm:text-sm",
(disabled || sttBusy) && "cursor-not-allowed opacity-50",
)}
/>
{mentionOpen && filteredMentions.length > 0 && (
<div className="absolute left-0 right-0 top-full z-20 mt-2 overflow-hidden rounded-2xl border border-border bg-popover/95 p-1 shadow-xl backdrop-blur-md">
<p className="px-2 py-1 text-[11px] font-medium uppercase tracking-wide text-muted-foreground">Mention</p>
{filteredMentions.map((option) => (
<button
key={option.id}
type="button"
onMouseDown={(event) => {
event.preventDefault();
insertMention(option);
}}
className="studio-control studio-hover-surface flex w-full items-start gap-2 rounded-xl px-2.5 py-2 text-left"
>
<span className="mt-0.5 rounded-md bg-accent px-1.5 py-0.5 text-[10px] font-semibold uppercase tracking-wide text-muted-foreground">
{option.kind}
</span>
<span className="min-w-0">
<span className="block truncate text-sm font-medium text-foreground">{option.label}</span>
{option.description && <span className="block truncate text-xs text-muted-foreground">{option.description}</span>}
</span>
</button>
))}
</div>
)}
</div>
</div>
<div className="mt-4 flex flex-wrap items-center gap-1.5 rounded-2xl bg-muted/35 p-1.5 ring-1 ring-border/25">
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled}>
<Box className="size-3.5" />
{selectedModel.label}
<ChevronDown className="size-3 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-72">
<DropdownMenuLabel>Model</DropdownMenuLabel>
<DropdownMenuSeparator />
{availableModels.map((model) => (
<DropdownMenuItem key={model.id} onClick={() => onModelChange(model.id)} className="flex-col items-start gap-1 py-2.5">
<span className="flex items-center gap-2 text-sm font-medium">
{model.label}
{model.badge && <span className="rounded-full bg-accent-blue/15 px-1.5 py-0.5 text-[10px] text-accent-blue">{model.badge}</span>}
</span>
<span className="text-xs text-muted-foreground">{model.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled}>
<Wand2 className="size-3.5" />
{selectedMode.label}
<ChevronDown className="size-3 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-64">
<DropdownMenuLabel>Mode</DropdownMenuLabel>
<DropdownMenuSeparator />
{availableModes.map((mode) => (
<DropdownMenuItem key={mode.id} onClick={() => onModeChange(mode.id)} className="flex-col items-start gap-1 py-2.5">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">{mode.description}</span>
</DropdownMenuItem>
))}
{unavailableModes.length > 0 && <DropdownMenuSeparator />}
{unavailableModes.map((mode) => (
<DropdownMenuItem key={mode.id} disabled className="flex-col items-start gap-1 py-2.5 opacity-60">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">
{unsupportedModeNotice(mode.id, capabilities) ?? mode.description}
</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled}>
<Monitor className="size-3.5" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-80">
<p className="mb-3 text-xs font-medium text-muted-foreground">Aspect ratio</p>
<div className="grid grid-cols-3 gap-2">
{availableAspectRatios.map((ratio) => (
<button
key={ratio}
type="button"
onClick={() => onAspectRatioChange(ratio)}
className={cn(
"studio-control studio-control-press studio-hover-surface flex flex-col items-center gap-2 rounded-xl border px-2 py-3 text-xs",
aspectRatio === ratio ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
<span className={cn("rounded-sm border border-current/40 bg-muted/40", ratio === "9:16" && "h-7 w-4", ratio === "16:9" && "h-4 w-7", ratio === "1:1" && "size-5", ratio === "4:3" && "h-5 w-6", ratio === "3:4" && "h-6 w-5", ratio === "21:9" && "h-3 w-8")} />
{ratio}
</button>
))}
</div>
<p className="mb-2 mt-4 text-xs font-medium text-muted-foreground">Resolution</p>
<div className="flex flex-wrap gap-2">
{availableResolutions.map((item) => (
<button
key={item}
type="button"
onClick={() => onResolutionChange(item)}
className={cn(
"studio-control studio-control-press studio-hover-surface rounded-full border px-3 py-1.5 text-xs font-medium",
resolution === item ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
{formatResolutionLabel(item)}
</button>
))}
{unavailableResolutions.map((item) => (
<button
key={item}
type="button"
disabled
className="studio-control rounded-full border border-border px-3 py-1.5 text-xs font-medium text-muted-foreground opacity-50"
title="Not supported on FastLTX models yet"
>
{formatResolutionLabel(item)}
</button>
))}
</div>
</PopoverContent>
</Popover>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled}>
<Clock className="size-3.5" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-72">
<p className="mb-3 text-xs font-medium text-muted-foreground">Total duration</p>
<Slider min={durationMin} max={durationMax} step={5} value={[durationSec]} onValueChange={(values) => onDurationChange(values[0] ?? durationMin)} />
<div className="mt-3 flex items-center justify-between text-[11px] text-muted-foreground">
<span>{formatDurationLabel(durationMin)}</span>
<span className="rounded-md border border-border px-2 py-1 text-xs font-medium text-foreground">{formatDurationLabel(durationSec)}</span>
<span>{formatDurationLabel(durationMax)}</span>
</div>
</PopoverContent>
</Popover>
<div className="ml-auto flex items-center gap-1.5">
{onSpeechTranscript && (
<SpeechToTextButton
disabled={disabled || isGenerating}
onTranscript={onSpeechTranscript}
onInterimChange={onSpeechInterimChange}
onBusyChange={setSttBusy}
/>
)}
<Button
aria-label="Generate"
onClick={onSubmit}
disabled={submitDisabled}
size="icon"
className="studio-control-press rounded-full bg-accent-blue text-white shadow-sm hover-capable:hover:bg-accent-blue/90 disabled:bg-muted disabled:text-muted-foreground"
>
<ArrowUp className="size-5" />
</Button>
</div>
</div>
{referenceMissing && value.trim() && (
<p className="mt-3 text-center text-xs leading-5 text-amber-700 dark:text-amber-400">
Upload a reference asset to use Omni reference mode.
</p>
)}
</div>
</section>
);
}
@@ -0,0 +1,73 @@
"use client";
import React from "react";
import AppNavRail, { type AppNavSection } from "@/components/creation/AppNavRail";
import CreationComposer from "@/components/creation/CreationComposer";
import PresetQuickLaunchRail, { type StoryPresetLike } from "@/components/creation/PresetQuickLaunchRail";
import {
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type MentionOption,
type ResolutionId,
} from "@/lib/creationConfig";
import type { LobbyCreationCapabilities } from "@/lib/creationCapabilities";
interface CreationStudioProps {
value: string;
disabled?: boolean;
isGenerating?: boolean;
canSubmit?: boolean;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referencePreviewUrl?: string | null;
firstFramePreviewUrl?: string | null;
lastFramePreviewUrl?: string | null;
mentionOptions?: MentionOption[];
storyPresets?: StoryPresetLike[];
activeSection?: AppNavSection;
onValueChange: (value: string) => void;
onSubmit: () => void;
onKeyDown?: (event: React.KeyboardEvent<HTMLTextAreaElement>) => void;
onModelChange: (modelId: CreationModelId) => void;
onModeChange: (modeId: CreationModeId) => void;
onAspectRatioChange: (aspectRatio: AspectRatioId) => void;
onResolutionChange: (resolution: ResolutionId) => void;
onDurationChange: (durationSec: number) => void;
onReferenceSelect?: (file: File | null) => void;
onFirstFrameSelect?: (file: File | null) => void;
onLastFrameSelect?: (file: File | null) => void;
onPresetGenerate?: (presetId: string) => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
onOpenProjects?: () => void;
capabilities?: LobbyCreationCapabilities;
}
export default function CreationStudio({
activeSection = "create",
onOpenProjects,
storyPresets = [],
onPresetGenerate,
isGenerating = false,
capabilities,
...composerProps
}: CreationStudioProps) {
return (
<div className="flex min-h-0 flex-1">
<AppNavRail activeSection={activeSection} onOpenProjects={onOpenProjects} />
<div className="min-w-0 flex-1 overflow-y-auto">
<div className="mx-auto flex w-full max-w-5xl flex-col gap-5 px-4 py-7 sm:px-6 sm:py-8">
<CreationComposer {...composerProps} isGenerating={isGenerating} capabilities={capabilities} />
{storyPresets.length > 0 && onPresetGenerate && (
<PresetQuickLaunchRail storyPresets={storyPresets} disabled={isGenerating} onPresetGenerate={onPresetGenerate} />
)}
</div>
</div>
</div>
);
}
@@ -0,0 +1,240 @@
"use client";
import React, { useCallback, useEffect, useRef, useState } from "react";
import { ChevronLeft, ChevronRight } from "lucide-react";
import { cn } from "@/lib/utils";
export interface StoryPresetLike {
id: string;
label: string;
description?: string;
segmentCount?: number;
styleTag?: string;
}
interface PresetQuickLaunchRailProps {
storyPresets: StoryPresetLike[];
disabled?: boolean;
onPresetGenerate: (presetId: string) => void;
}
export default function PresetQuickLaunchRail({
storyPresets,
disabled = false,
onPresetGenerate,
}: PresetQuickLaunchRailProps) {
const scrollRef = useRef<HTMLDivElement>(null);
const [canScrollLeft, setCanScrollLeft] = useState(false);
const [canScrollRight, setCanScrollRight] = useState(false);
const [presetRailDragging, setPresetRailDragging] = useState(false);
const presetDragStateRef = useRef({
pointerId: null as number | null,
startX: 0,
startScrollLeft: 0,
moved: false,
});
const suppressPresetClickRef = useRef(false);
const updateScrollState = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setCanScrollLeft(el.scrollLeft > 2);
setCanScrollRight(el.scrollLeft + el.clientWidth < el.scrollWidth - 2);
}, []);
const scrollByAmount = useCallback(
(direction: "left" | "right") => {
const el = scrollRef.current;
if (!el) return;
const delta = direction === "left" ? -220 : 220;
el.scrollBy({ left: delta, behavior: "smooth" });
window.setTimeout(updateScrollState, 220);
},
[updateScrollState],
);
const handlePresetWheel = useCallback(
(event: React.WheelEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
const dominantDelta = Math.abs(event.deltaX) > Math.abs(event.deltaY) ? event.deltaX : event.deltaY;
if (!dominantDelta) return;
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
const nextScrollLeft = Math.min(Math.max(el.scrollLeft + dominantDelta, 0), maxScrollLeft);
if (nextScrollLeft === el.scrollLeft) return;
event.preventDefault();
el.scrollLeft = nextScrollLeft;
updateScrollState();
},
[updateScrollState],
);
const finishPresetDrag = useCallback(() => {
presetDragStateRef.current = {
pointerId: null,
startX: 0,
startScrollLeft: 0,
moved: false,
};
setPresetRailDragging(false);
}, []);
const handlePresetPointerDown = useCallback((event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (event.pointerType !== "mouse" || event.button !== 0) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
suppressPresetClickRef.current = false;
presetDragStateRef.current = {
pointerId: event.pointerId,
startX: event.clientX,
startScrollLeft: el.scrollLeft,
moved: false,
};
}, []);
const handlePresetPointerMove = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
const dragState = presetDragStateRef.current;
if (!el || dragState.pointerId !== event.pointerId) return;
const deltaX = event.clientX - dragState.startX;
if (!dragState.moved && Math.abs(deltaX) > 4) {
dragState.moved = true;
suppressPresetClickRef.current = true;
setPresetRailDragging(true);
el.setPointerCapture?.(event.pointerId);
}
if (!dragState.moved) return;
event.preventDefault();
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
el.scrollLeft = Math.min(Math.max(dragState.startScrollLeft - deltaX, 0), maxScrollLeft);
updateScrollState();
},
[updateScrollState],
);
const handlePresetPointerUp = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el || presetDragStateRef.current.pointerId !== event.pointerId) return;
if (el.hasPointerCapture?.(event.pointerId)) {
el.releasePointerCapture(event.pointerId);
}
finishPresetDrag();
},
[finishPresetDrag],
);
const handlePresetClickCapture = useCallback((event: React.MouseEvent<HTMLDivElement>) => {
if (!suppressPresetClickRef.current) return;
suppressPresetClickRef.current = false;
event.preventDefault();
event.stopPropagation();
}, []);
useEffect(() => {
updateScrollState();
}, [storyPresets, updateScrollState]);
useEffect(() => {
const el = scrollRef.current;
if (!el) return;
const observer = new ResizeObserver(() => updateScrollState());
observer.observe(el);
return () => observer.disconnect();
}, [updateScrollState]);
if (storyPresets.length === 0) return null;
const scrollMaskStyle =
canScrollLeft && canScrollRight
? {
maskImage: "linear-gradient(to right, transparent, black 20px, black calc(100% - 20px), transparent)",
WebkitMaskImage: "linear-gradient(to right, transparent, black 20px, black calc(100% - 20px), transparent)",
}
: canScrollLeft
? {
maskImage: "linear-gradient(to right, transparent, black 20px, black)",
WebkitMaskImage: "linear-gradient(to right, transparent, black 20px, black)",
}
: canScrollRight
? {
maskImage: "linear-gradient(to right, black, black calc(100% - 20px), transparent)",
WebkitMaskImage: "linear-gradient(to right, black, black calc(100% - 20px), transparent)",
}
: undefined;
return (
<div className={cn("mx-auto w-full max-w-3xl transition-opacity duration-200", disabled && "pointer-events-none opacity-40")}>
<div className="grid grid-cols-[auto_minmax(0,1fr)_auto] items-center gap-1 sm:gap-2">
<div className="flex w-8 shrink-0 justify-center">
{canScrollLeft ? (
<button
type="button"
aria-label="Scroll suggested prompts left"
onClick={() => scrollByAmount("left")}
className="studio-control studio-control-press inline-flex size-8 items-center justify-center rounded-full text-muted-foreground hover-capable:hover:bg-muted/60 hover-capable:hover:text-foreground"
>
<ChevronLeft className="size-4" />
</button>
) : null}
</div>
<div
ref={scrollRef}
onScroll={updateScrollState}
onWheel={handlePresetWheel}
onPointerDown={handlePresetPointerDown}
onPointerMove={handlePresetPointerMove}
onPointerUp={handlePresetPointerUp}
onPointerCancel={handlePresetPointerUp}
onLostPointerCapture={finishPresetDrag}
onClickCapture={handlePresetClickCapture}
style={scrollMaskStyle}
className={cn(
"scrollbar-hidden flex gap-2 overflow-x-auto overflow-y-visible py-0.5 select-none",
presetRailDragging ? "cursor-grabbing" : "cursor-grab",
)}
>
{storyPresets.map((preset) => (
<button
key={preset.id}
type="button"
disabled={disabled}
onClick={() => onPresetGenerate(preset.id)}
className="studio-control studio-control-press studio-hover-surface flex w-[12.5rem] shrink-0 flex-col gap-1 rounded-xl border border-border/50 bg-card/70 px-3 py-2.5 text-left"
>
<span className="line-clamp-1 text-sm font-medium text-foreground">{preset.label}</span>
{preset.description && (
<span className="text-pretty line-clamp-2 text-xs leading-5 text-muted-foreground">{preset.description}</span>
)}
</button>
))}
</div>
<div className="flex w-8 shrink-0 justify-center">
{canScrollRight ? (
<button
type="button"
aria-label="Scroll suggested prompts right"
onClick={() => scrollByAmount("right")}
className="studio-control studio-control-press inline-flex size-8 items-center justify-center rounded-full text-muted-foreground hover-capable:hover:bg-muted/60 hover-capable:hover:text-foreground"
>
<ChevronRight className="size-4" />
</button>
) : null}
</div>
</div>
</div>
);
}
@@ -0,0 +1,97 @@
"use client";
import React, { useRef, useState } from "react";
import { ImagePlus } from "lucide-react";
import { REFERENCE_ACCEPT, isReferenceMediaFile } from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
interface ReferenceUploadSlotProps {
label: string;
sublabel?: string;
previewUrl?: string | null;
required?: boolean;
optional?: boolean;
disabled?: boolean;
onSelect?: (file: File | null) => void;
}
export default function ReferenceUploadSlot({
label,
sublabel,
previewUrl = null,
required = false,
optional = false,
disabled = false,
onSelect,
}: ReferenceUploadSlotProps) {
const fileInputRef = useRef<HTMLInputElement>(null);
const [dragActive, setDragActive] = useState(false);
function handleFile(file: File | null) {
if (!file || !isReferenceMediaFile(file)) return;
onSelect?.(file);
}
return (
<div className="flex flex-col gap-1">
<button
type="button"
aria-label={[label, sublabel].filter(Boolean).join(" ")}
onClick={() => fileInputRef.current?.click()}
disabled={disabled}
onDragEnter={(event) => {
event.preventDefault();
event.stopPropagation();
if (!disabled) setDragActive(true);
}}
onDragOver={(event) => {
event.preventDefault();
event.stopPropagation();
if (!disabled) setDragActive(true);
}}
onDragLeave={(event) => {
event.preventDefault();
event.stopPropagation();
setDragActive(false);
}}
onDrop={(event) => {
event.preventDefault();
event.stopPropagation();
setDragActive(false);
if (disabled) return;
handleFile(event.dataTransfer.files?.[0] ?? null);
}}
className={cn(
"studio-control studio-control-press studio-hover-surface relative flex size-[76px] shrink-0 flex-col items-center justify-center gap-1 overflow-hidden rounded-2xl border border-dashed bg-muted/50 px-1 text-center text-[11px] font-medium text-muted-foreground",
required && !previewUrl ? "border-amber-500/50" : "border-border/60",
dragActive && "border-accent-blue bg-accent-blue/10 ring-2 ring-accent-blue/30",
disabled && "pointer-events-none opacity-50",
)}
>
{previewUrl ? (
<img src={previewUrl} alt="" className="studio-media-outline absolute inset-0 size-full object-cover" />
) : (
<>
<ImagePlus className="size-4" />
<span>{label}</span>
{sublabel && <span className="text-[10px] font-normal opacity-70">{sublabel}</span>}
</>
)}
</button>
{(required || optional) && (
<span className="text-center text-[10px] text-muted-foreground">{required ? "Required" : "Optional"}</span>
)}
<input
ref={fileInputRef}
type="file"
accept={REFERENCE_ACCEPT}
className="hidden"
onChange={(event) => {
handleFile(event.target.files?.[0] ?? null);
event.target.value = "";
}}
/>
</div>
);
}
@@ -0,0 +1,213 @@
"use client";
import React from "react";
import { Box, ChevronDown, Clock, Monitor, Wand2 } from "lucide-react";
import ConfigPill from "@/components/creation/ConfigPill";
import {
DropdownMenu,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuTrigger,
} from "@/components/ui/dropdown-menu";
import { Popover, PopoverContent, PopoverTrigger } from "@/components/ui/popover";
import { Slider } from "@/components/ui/slider";
import {
ASPECT_RATIOS,
CREATION_MODELS,
CREATION_MODES,
RESOLUTIONS,
type AspectRatioId,
type CreationModeId,
type CreationModelId,
type ResolutionId,
formatDurationLabel,
formatResolutionLabel,
} from "@/lib/creationConfig";
import { cn } from "@/lib/utils";
export interface SessionCreationConfig {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}
interface SessionCreationConfigPillsProps extends SessionCreationConfig {
disabled?: boolean;
readOnly?: boolean;
onModelChange?: (modelId: CreationModelId) => void;
onModeChange?: (modeId: CreationModeId) => void;
onAspectRatioChange?: (aspectRatio: AspectRatioId) => void;
onResolutionChange?: (resolution: ResolutionId) => void;
onDurationChange?: (durationSec: number) => void;
}
export default function SessionCreationConfigPills({
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
disabled = false,
readOnly = false,
onModelChange,
onModeChange,
onAspectRatioChange,
onResolutionChange,
onDurationChange,
}: SessionCreationConfigPillsProps) {
const selectedModel = CREATION_MODELS.find((model) => model.id === modelId) ?? CREATION_MODELS[0];
const selectedMode = CREATION_MODES.find((mode) => mode.id === modeId) ?? CREATION_MODES[0];
const isInteractive = !readOnly && !disabled;
const pillClassName = cn(
"h-9 min-h-9 px-2 text-[11px]",
!isInteractive && "pointer-events-none opacity-70",
);
if (readOnly) {
return (
<div className="flex flex-wrap items-center gap-1.5">
<ConfigPill disabled className={pillClassName} aria-label="Model">
<Box className="size-3" />
{selectedModel.label}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Mode">
<Wand2 className="size-3" />
{selectedMode.label}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Aspect ratio and resolution">
<Monitor className="size-3" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
<ConfigPill disabled className={pillClassName} aria-label="Duration">
<Clock className="size-3" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</div>
);
}
return (
<div className="flex flex-wrap items-center gap-1.5">
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Model">
<Box className="size-3" />
{selectedModel.label}
<ChevronDown className="size-2.5 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-72">
<DropdownMenuLabel>Model</DropdownMenuLabel>
<DropdownMenuSeparator />
{CREATION_MODELS.map((model) => (
<DropdownMenuItem key={model.id} onClick={() => onModelChange?.(model.id)} className="flex-col items-start gap-1 py-2.5">
<span className="flex items-center gap-2 text-sm font-medium">
{model.label}
{model.badge && <span className="rounded-full bg-accent-blue/15 px-1.5 py-0.5 text-[10px] text-accent-blue">{model.badge}</span>}
</span>
<span className="text-xs text-muted-foreground">{model.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<DropdownMenu>
<DropdownMenuTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Mode">
<Wand2 className="size-3" />
{selectedMode.label}
<ChevronDown className="size-2.5 opacity-60" />
</ConfigPill>
</DropdownMenuTrigger>
<DropdownMenuContent align="start" className="w-64">
<DropdownMenuLabel>Mode</DropdownMenuLabel>
<DropdownMenuSeparator />
{CREATION_MODES.map((mode) => (
<DropdownMenuItem key={mode.id} onClick={() => onModeChange?.(mode.id)} className="flex-col items-start gap-1 py-2.5">
<span className="text-sm font-medium">{mode.label}</span>
<span className="text-xs text-muted-foreground">{mode.description}</span>
</DropdownMenuItem>
))}
</DropdownMenuContent>
</DropdownMenu>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Aspect ratio and resolution">
<Monitor className="size-3" />
{aspectRatio} {formatResolutionLabel(resolution)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-80">
<p className="mb-3 text-xs font-medium text-muted-foreground">Aspect ratio</p>
<div className="grid grid-cols-3 gap-2">
{ASPECT_RATIOS.map((ratio) => (
<button
key={ratio}
type="button"
onClick={() => onAspectRatioChange?.(ratio)}
className={cn(
"studio-control studio-control-press studio-hover-surface flex flex-col items-center gap-2 rounded-xl border px-2 py-3 text-xs",
aspectRatio === ratio ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
<span
className={cn(
"rounded-sm border border-current/40 bg-muted/40",
ratio === "9:16" && "h-7 w-4",
ratio === "16:9" && "h-4 w-7",
ratio === "1:1" && "size-5",
ratio === "4:3" && "h-5 w-6",
ratio === "3:4" && "h-6 w-5",
ratio === "21:9" && "h-3 w-8",
)}
/>
{ratio}
</button>
))}
</div>
<p className="mb-2 mt-4 text-xs font-medium text-muted-foreground">Resolution</p>
<div className="flex flex-wrap gap-2">
{RESOLUTIONS.map((item) => (
<button
key={item}
type="button"
onClick={() => onResolutionChange?.(item)}
className={cn(
"studio-control studio-control-press studio-hover-surface rounded-full border px-3 py-1.5 text-xs font-medium",
resolution === item ? "border-accent-blue bg-accent-blue/10 text-foreground" : "border-border",
)}
>
{formatResolutionLabel(item)}
</button>
))}
</div>
</PopoverContent>
</Popover>
<Popover>
<PopoverTrigger asChild>
<ConfigPill disabled={disabled} className={pillClassName} aria-label="Duration">
<Clock className="size-3" />
{formatDurationLabel(durationSec)}
</ConfigPill>
</PopoverTrigger>
<PopoverContent align="start" className="w-72">
<p className="mb-3 text-xs font-medium text-muted-foreground">Total duration</p>
<Slider min={5} max={15} step={5} value={[durationSec]} onValueChange={(values) => onDurationChange?.(values[0] ?? 5)} />
<div className="mt-3 flex items-center justify-between text-[11px] text-muted-foreground">
<span>5s</span>
<span className="rounded-md border border-border px-2 py-1 text-xs font-medium text-foreground">{formatDurationLabel(durationSec)}</span>
<span>15s</span>
</div>
</PopoverContent>
</Popover>
</div>
);
}
@@ -17,13 +17,6 @@ import {
SelectValue,
} from '@/components/ui/select';
import { Textarea } from '@/components/ui/textarea';
import {
DEFAULT_GENERATION_MODE,
GENERATION_MODES,
getGenerationMode,
isGenerationMode,
type GenerationMode,
} from '@/lib/generationMode';
interface DevtoolsComposerProps {
connected?: boolean;
@@ -44,10 +37,6 @@ interface DevtoolsComposerProps {
loopGenerationEnabled?: boolean;
curatedPromptLimit?: number;
maxCuratedPromptCount?: number;
generationMode?: GenerationMode;
supportedGenerationModes?: readonly GenerationMode[];
conditioningPanel?: React.ReactNode;
onGenerationModeChange?: (mode: GenerationMode) => void;
rewriteWindowMode?: boolean;
rewritingSeedPrompts?: boolean;
autoExtensionTimeoutHint?: string;
@@ -85,10 +74,6 @@ export default function DevtoolsComposer({
loopGenerationEnabled = false,
curatedPromptLimit = 0,
maxCuratedPromptCount = 0,
generationMode = DEFAULT_GENERATION_MODE,
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
conditioningPanel,
onGenerationModeChange = () => {},
rewriteWindowMode = false,
rewritingSeedPrompts = false,
autoExtensionTimeoutHint = '',
@@ -107,7 +92,6 @@ export default function DevtoolsComposer({
onSpeechInterimChange,
}: DevtoolsComposerProps) {
const [sttBusy, setSttBusy] = useState(false);
const selectedGenerationMode = getGenerationMode(generationMode);
const submitButtonLabel = useMemo(
() =>
rewriteWindowMode
@@ -126,8 +110,7 @@ export default function DevtoolsComposer({
);
return (
<section className="space-y-4">
{conditioningPanel}
<section>
<Card>
<CardContent className="space-y-5 p-5">
<div className="grid gap-5 xl:grid-cols-[minmax(0,1fr)_320px]">
@@ -302,48 +285,6 @@ export default function DevtoolsComposer({
</div>
<div className="space-y-4">
<div className="space-y-2">
<Label htmlFor="devtools-generation-mode">
Generation mode
</Label>
<Select
value={generationMode}
disabled={sessionStarted}
onValueChange={(value) => {
if (isGenerationMode(value)) {
onGenerationModeChange(value);
}
}}
>
<SelectTrigger
id="devtools-generation-mode"
aria-label="Generation mode"
title={selectedGenerationMode.name}
>
<SelectValue />
</SelectTrigger>
<SelectContent>
{GENERATION_MODES.map((mode) => (
<SelectItem
key={mode.id}
value={mode.id}
disabled={!supportedGenerationModes.includes(mode.id)}
title={
supportedGenerationModes.includes(mode.id)
? mode.name
: `${mode.name} (unavailable on this runtime)`
}
>
{mode.label}
</SelectItem>
))}
</SelectContent>
</Select>
<p className="text-sm text-muted-foreground">
{selectedGenerationMode.description}
</p>
</div>
<div className="flex items-start gap-3">
<Checkbox
id="devtools-enhance-prompts"
@@ -35,7 +35,6 @@ describe('DevtoolsShell', () => {
expect(screen.getByText('Devtools Mode')).toBeInTheDocument();
expect(screen.getByText('Your video will appear here')).toBeInTheDocument();
expect(screen.getByLabelText('Story preset')).toBeInTheDocument();
expect(screen.getByLabelText('Generation mode')).toBeInTheDocument();
expect(screen.getByLabelText('Continuation prompt')).toBeDisabled();
expect(screen.getByText('Advanced controls')).toBeInTheDocument();
@@ -9,7 +9,6 @@ import VideoPlayer from '../VideoPlayer';
import RewriteInspector from '../rewrite/RewriteInspector';
import DevtoolsComposer from './DevtoolsComposer';
import DevtoolsDrawer from './DevtoolsDrawer';
import { DEFAULT_GENERATION_MODE, GENERATION_MODES, type GenerationMode } from '@/lib/generationMode';
interface DevtoolsShellProps {
connected?: boolean;
@@ -31,11 +30,6 @@ interface DevtoolsShellProps {
curatedPromptLimit?: number;
maxCuratedPromptCount?: number;
generationMode?: GenerationMode;
supportedGenerationModes?: readonly GenerationMode[];
conditioningPanel?: React.ReactNode;
onGenerationModeChange?: (mode: GenerationMode) => void;
onPresetChange?: (e: React.ChangeEvent<HTMLSelectElement>) => void;
onEnhancementToggle?: (e: React.ChangeEvent<HTMLInputElement>) => void;
onCuratedPromptLimitChange?: (e: React.ChangeEvent<HTMLInputElement>) => void;
@@ -143,11 +137,6 @@ export default function DevtoolsShell({
curatedPromptLimit = 0,
maxCuratedPromptCount = 0,
generationMode = DEFAULT_GENERATION_MODE,
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
conditioningPanel,
onGenerationModeChange = () => {},
onPresetChange = () => {},
onEnhancementToggle = () => {},
onCuratedPromptLimitChange = () => {},
@@ -298,10 +287,6 @@ export default function DevtoolsShell({
loopGenerationEnabled={loopGenerationEnabled}
curatedPromptLimit={curatedPromptLimit}
maxCuratedPromptCount={maxCuratedPromptCount}
generationMode={generationMode}
supportedGenerationModes={supportedGenerationModes}
conditioningPanel={conditioningPanel}
onGenerationModeChange={onGenerationModeChange}
rewriteWindowMode={livePromptRewriteMode}
rewritingSeedPrompts={rewritingSeedPrompts}
autoExtensionTimeoutHint={autoExtensionTimeoutHint}
@@ -0,0 +1,141 @@
"use client";
import * as React from "react";
import * as DropdownMenuPrimitive from "@radix-ui/react-dropdown-menu";
import { Check, ChevronRight } from "lucide-react";
import { cn } from "@/lib/utils";
const DropdownMenu = DropdownMenuPrimitive.Root;
const DropdownMenuTrigger = DropdownMenuPrimitive.Trigger;
const DropdownMenuGroup = DropdownMenuPrimitive.Group;
const DropdownMenuPortal = DropdownMenuPrimitive.Portal;
const DropdownMenuSub = DropdownMenuPrimitive.Sub;
const DropdownMenuRadioGroup = DropdownMenuPrimitive.RadioGroup;
const DropdownMenuSubTrigger = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.SubTrigger>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.SubTrigger> & { inset?: boolean }
>(({ className, inset, children, ...props }, ref) => (
<DropdownMenuPrimitive.SubTrigger
ref={ref}
className={cn(
"flex cursor-default select-none items-center rounded-xl px-2 py-1.5 text-sm outline-none data-[state=open]:bg-accent focus:bg-accent",
inset && "pl-8",
className,
)}
{...props}
>
{children}
<ChevronRight className="ml-auto size-4" />
</DropdownMenuPrimitive.SubTrigger>
));
DropdownMenuSubTrigger.displayName = DropdownMenuPrimitive.SubTrigger.displayName;
const DropdownMenuSubContent = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.SubContent>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.SubContent>
>(({ className, ...props }, ref) => (
<DropdownMenuPrimitive.SubContent
ref={ref}
className={cn(
"z-50 min-w-[8rem] overflow-hidden rounded-2xl border border-border bg-popover/95 p-1 text-popover-foreground shadow-xl backdrop-blur-md",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
));
DropdownMenuSubContent.displayName = DropdownMenuPrimitive.SubContent.displayName;
const DropdownMenuContent = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Content>
>(({ className, sideOffset = 6, ...props }, ref) => (
<DropdownMenuPrimitive.Portal>
<DropdownMenuPrimitive.Content
ref={ref}
sideOffset={sideOffset}
className={cn(
"z-50 min-w-[12rem] overflow-hidden rounded-2xl border border-border bg-popover/95 p-1.5 text-popover-foreground shadow-xl backdrop-blur-md",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
</DropdownMenuPrimitive.Portal>
));
DropdownMenuContent.displayName = DropdownMenuPrimitive.Content.displayName;
const DropdownMenuItem = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Item>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Item> & { inset?: boolean }
>(({ className, inset, ...props }, ref) => (
<DropdownMenuPrimitive.Item
ref={ref}
className={cn(
"relative flex cursor-default select-none items-center gap-2 rounded-xl px-2.5 py-2 text-sm outline-none transition-colors data-[disabled]:pointer-events-none data-[disabled]:opacity-50 focus:bg-accent focus:text-accent-foreground",
inset && "pl-8",
className,
)}
{...props}
/>
));
DropdownMenuItem.displayName = DropdownMenuPrimitive.Item.displayName;
const DropdownMenuCheckboxItem = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.CheckboxItem>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.CheckboxItem>
>(({ className, children, checked, ...props }, ref) => (
<DropdownMenuPrimitive.CheckboxItem
ref={ref}
className={cn(
"relative flex cursor-default select-none items-center rounded-xl py-2 pl-8 pr-2 text-sm outline-none transition-colors data-[disabled]:pointer-events-none data-[disabled]:opacity-50 focus:bg-accent focus:text-accent-foreground",
className,
)}
checked={checked}
{...props}
>
<span className="absolute left-2 flex size-3.5 items-center justify-center">
<DropdownMenuPrimitive.ItemIndicator>
<Check className="size-4 text-accent-blue" />
</DropdownMenuPrimitive.ItemIndicator>
</span>
{children}
</DropdownMenuPrimitive.CheckboxItem>
));
DropdownMenuCheckboxItem.displayName = DropdownMenuPrimitive.CheckboxItem.displayName;
const DropdownMenuLabel = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Label>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Label> & { inset?: boolean }
>(({ className, inset, ...props }, ref) => (
<DropdownMenuPrimitive.Label ref={ref} className={cn("px-2 py-1.5 text-xs font-semibold text-muted-foreground", inset && "pl-8", className)} {...props} />
));
DropdownMenuLabel.displayName = DropdownMenuPrimitive.Label.displayName;
const DropdownMenuSeparator = React.forwardRef<
React.ElementRef<typeof DropdownMenuPrimitive.Separator>,
React.ComponentPropsWithoutRef<typeof DropdownMenuPrimitive.Separator>
>(({ className, ...props }, ref) => (
<DropdownMenuPrimitive.Separator ref={ref} className={cn("-mx-1 my-1 h-px bg-border", className)} {...props} />
));
DropdownMenuSeparator.displayName = DropdownMenuPrimitive.Separator.displayName;
export {
DropdownMenu,
DropdownMenuTrigger,
DropdownMenuContent,
DropdownMenuItem,
DropdownMenuCheckboxItem,
DropdownMenuLabel,
DropdownMenuSeparator,
DropdownMenuGroup,
DropdownMenuPortal,
DropdownMenuSub,
DropdownMenuSubContent,
DropdownMenuSubTrigger,
DropdownMenuRadioGroup,
};
@@ -0,0 +1,33 @@
"use client";
import * as React from "react";
import * as PopoverPrimitive from "@radix-ui/react-popover";
import { cn } from "@/lib/utils";
const Popover = PopoverPrimitive.Root;
const PopoverTrigger = PopoverPrimitive.Trigger;
const PopoverAnchor = PopoverPrimitive.Anchor;
const PopoverContent = React.forwardRef<
React.ElementRef<typeof PopoverPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof PopoverPrimitive.Content>
>(({ className, align = "center", sideOffset = 6, ...props }, ref) => (
<PopoverPrimitive.Portal>
<PopoverPrimitive.Content
ref={ref}
align={align}
sideOffset={sideOffset}
className={cn(
"z-50 w-72 rounded-2xl border border-border bg-popover/95 p-3 text-popover-foreground shadow-xl backdrop-blur-md outline-none",
"data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0 data-[state=closed]:zoom-out-95 data-[state=open]:zoom-in-95",
"data-[side=bottom]:slide-in-from-top-2 data-[side=top]:slide-in-from-bottom-2 data-[side=left]:slide-in-from-right-2 data-[side=right]:slide-in-from-left-2",
className,
)}
{...props}
/>
</PopoverPrimitive.Portal>
));
PopoverContent.displayName = PopoverPrimitive.Content.displayName;
export { Popover, PopoverTrigger, PopoverContent, PopoverAnchor };
@@ -0,0 +1,25 @@
"use client";
import * as React from "react";
import * as SliderPrimitive from "@radix-ui/react-slider";
import { cn } from "@/lib/utils";
const Slider = React.forwardRef<
React.ElementRef<typeof SliderPrimitive.Root>,
React.ComponentPropsWithoutRef<typeof SliderPrimitive.Root>
>(({ className, ...props }, ref) => (
<SliderPrimitive.Root
ref={ref}
className={cn("relative flex w-full touch-none select-none items-center", className)}
{...props}
>
<SliderPrimitive.Track className="relative h-1.5 w-full grow overflow-hidden rounded-full bg-muted">
<SliderPrimitive.Range className="absolute h-full bg-accent-blue" />
</SliderPrimitive.Track>
<SliderPrimitive.Thumb className="block size-4 rounded-full border border-accent-blue/40 bg-background shadow transition-colors focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40 disabled:pointer-events-none disabled:opacity-50" />
</SliderPrimitive.Root>
));
Slider.displayName = SliderPrimitive.Root.displayName;
export { Slider };
@@ -0,0 +1,48 @@
"use client";
import * as React from "react";
import * as TabsPrimitive from "@radix-ui/react-tabs";
import { cn } from "@/lib/utils";
const Tabs = TabsPrimitive.Root;
const TabsList = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.List>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.List>
>(({ className, ...props }, ref) => (
<TabsPrimitive.List
ref={ref}
className={cn("inline-flex items-center gap-1 rounded-full bg-muted/60 p-1 text-muted-foreground", className)}
{...props}
/>
));
TabsList.displayName = TabsPrimitive.List.displayName;
const TabsTrigger = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.Trigger>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.Trigger>
>(({ className, ...props }, ref) => (
<TabsPrimitive.Trigger
ref={ref}
className={cn(
"inline-flex items-center justify-center rounded-full px-3 py-1.5 text-xs font-medium whitespace-nowrap transition-all",
"focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-accent-blue/40",
"disabled:pointer-events-none disabled:opacity-50",
"data-[state=active]:bg-card data-[state=active]:text-foreground data-[state=active]:shadow-sm",
className,
)}
{...props}
/>
));
TabsTrigger.displayName = TabsPrimitive.Trigger.displayName;
const TabsContent = React.forwardRef<
React.ElementRef<typeof TabsPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof TabsPrimitive.Content>
>(({ className, ...props }, ref) => (
<TabsPrimitive.Content ref={ref} className={cn("mt-4 outline-none", className)} {...props} />
));
TabsContent.displayName = TabsPrimitive.Content.displayName;
export { Tabs, TabsList, TabsTrigger, TabsContent };
@@ -1,55 +0,0 @@
import { act, renderHook, waitFor } from "@testing-library/react";
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { useAssetLibrary } from "./useAssetLibrary";
const image = { asset_id: "asset-1", kind: "image", name: "frame.png", mime_type: "image/png", size: 5, url: "/assets/asset-1" };
describe("asset library lifecycle", () => {
beforeEach(() => localStorage.clear());
afterEach(() => vi.unstubAllGlobals());
it("uploads the raw file and keeps assignment state separate from uploaded assets", async () => {
const fetchMock = vi.fn(async () => ({ ok: true, json: async () => image }));
vi.stubGlobal("fetch", fetchMock);
const { result } = renderHook(() => useAssetLibrary());
const file = new File(["image"], "first frame.png", { type: "image/png" });
await act(async () => { await result.current.uploadAssets([file]); });
expect(fetchMock).toHaveBeenCalledWith("/assets", expect.objectContaining({ body: file, headers: { "Content-Type": "image/png", "X-Asset-Name": "first%20frame.png" } }));
act(() => result.current.assignAsset("asset-1", "first_frame"));
expect(result.current.conditioningAssets).toEqual([{ asset_id: "asset-1", role: "first_frame" }]);
act(() => result.current.clearConditioning());
expect(result.current.assets).toHaveLength(1);
expect(result.current.conditioningAssets).toEqual([]);
});
it("identifies stale server assets without downloading their contents", async () => {
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
const fetchMock = vi.fn(async () => ({ ok: false, status: 404 }));
vi.stubGlobal("fetch", fetchMock);
const { result } = renderHook(() => useAssetLibrary());
await waitFor(() => expect(result.current.assets).toHaveLength(1));
act(() => result.current.assignAsset("asset-1", "first_frame"));
let message: string | null = null;
await act(async () => { message = await result.current.verifySelectedAssets(); });
expect(message).toMatch(/Upload it again/);
expect(result.current.assets[0].missing).toBe(true);
expect(fetchMock).toHaveBeenCalledWith("/assets/asset-1", expect.objectContaining({ method: "HEAD" }));
});
it("deletes both the library entry and its selected references", async () => {
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true })));
const { result } = renderHook(() => useAssetLibrary());
await waitFor(() => expect(result.current.assets).toHaveLength(1));
act(() => result.current.assignAsset("asset-1", "reference"));
await act(async () => { await result.current.removeAsset("asset-1"); });
expect(result.current.assets).toEqual([]);
expect(result.current.conditioningAssets).toEqual([]);
});
it("does not mark a valid upload missing when the browser cannot preview its codec", async () => {
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true, status: 200 })));
const { result } = renderHook(() => useAssetLibrary());
await waitFor(() => expect(result.current.assets).toHaveLength(1));
await act(async () => { await result.current.checkAssetAvailability("asset-1"); });
expect(result.current.assets[0].missing).not.toBe(true);
});
});
@@ -1,136 +0,0 @@
"use client";
import { useCallback, useEffect, useRef, useState } from "react";
import type { ConditioningAsset, ConditioningRole, GenerationAsset } from "@/lib/generationMode";
const LIBRARY_KEY = "dreamverse-asset-library-v1";
const MAX_UPLOAD_BYTES = 100 * 1024 * 1024;
async function responseError(response: Response, fallback: string): Promise<Error> {
const payload = await response.json().catch(() => ({}));
return new Error(typeof payload.detail === "string" ? payload.detail : fallback);
}
function isAsset(value: unknown): value is GenerationAsset {
if (!value || typeof value !== "object") return false;
const asset = value as Partial<GenerationAsset>;
return typeof asset.asset_id === "string" && /^[a-zA-Z0-9_-]+$/.test(asset.asset_id)
&& ["image", "video", "audio"].includes(asset.kind || "")
&& typeof asset.name === "string" && typeof asset.mime_type === "string" && typeof asset.size === "number";
}
/** Asset ownership lives here so composers and other pickers share the same library. */
export function useAssetLibrary() {
const [assets, setAssets] = useState<GenerationAsset[]>([]);
const [conditioningAssets, setConditioningAssets] = useState<ConditioningAsset[]>([]);
const [uploading, setUploading] = useState(false);
const [assetError, setAssetError] = useState("");
const [hydrated, setHydrated] = useState(false);
const uploadingRef = useRef(false);
useEffect(() => {
try {
const saved = JSON.parse(localStorage.getItem(LIBRARY_KEY) || "[]");
if (Array.isArray(saved)) setAssets(saved.filter(isAsset).map((asset) => ({
...asset, url: `/assets/${asset.asset_id}`,
})));
} catch { /* Storage is optional; uploads still work in private browsing. */ }
setHydrated(true);
}, []);
useEffect(() => {
if (!hydrated) return;
try { localStorage.setItem(LIBRARY_KEY, JSON.stringify(assets)); } catch { /* Optional cache. */ }
}, [assets, hydrated]);
const uploadAssets = useCallback(async (files: File[]) => {
if (uploadingRef.current) return;
uploadingRef.current = true;
setUploading(true);
setAssetError("");
const errors: string[] = [];
for (const file of files) {
try {
if (!/^(image|video|audio)\//.test(file.type)) throw new Error(`${file.name}: choose an image, video, or audio file.`);
const maxBytes = file.type.startsWith("image/") ? 15 * 1024 * 1024 : MAX_UPLOAD_BYTES;
if (!file.size || file.size > maxBytes) throw new Error(`${file.name}: use a non-empty file up to ${maxBytes / 1024 / 1024} MiB.`);
const response = await fetch("/assets", {
method: "POST",
headers: { "Content-Type": file.type, "X-Asset-Name": encodeURIComponent(file.name) },
body: file,
});
if (!response.ok) throw await responseError(response, `Could not upload ${file.name}.`);
const asset: unknown = await response.json();
if (!isAsset(asset)) throw new Error("The server returned an invalid asset. Please retry the upload.");
setAssets((current) => [...current.filter((item) => item.asset_id !== asset.asset_id), {
...asset, url: `/assets/${asset.asset_id}`, missing: false,
}]);
} catch (error) {
errors.push(error instanceof Error ? error.message : `Could not upload ${file.name}.`);
}
}
setAssetError(errors.join(" "));
setUploading(false);
uploadingRef.current = false;
}, []);
const assignAsset = useCallback((assetId: string, role: ConditioningRole) => {
setConditioningAssets((current) => {
const next = role === "reference" ? current : current.filter((item) => item.role !== role);
if (!assetId || next.some((item) => item.asset_id === assetId && item.role === role)) return next;
return [...next, { asset_id: assetId, role }];
});
}, []);
const removeConditioning = useCallback((index: number) => {
setConditioningAssets((current) => current.filter((_, itemIndex) => itemIndex !== index));
}, []);
const moveConditioning = useCallback((from: number, to: number) => {
setConditioningAssets((current) => {
if (from < 0 || to < 0 || from >= current.length || to >= current.length) return current;
const next = [...current];
next.splice(to, 0, next.splice(from, 1)[0]);
return next;
});
}, []);
const clearConditioning = useCallback(() => setConditioningAssets([]), []);
const markAssetMissing = useCallback((assetId: string) => {
setAssets((current) => current.map((asset) => asset.asset_id === assetId ? { ...asset, missing: true } : asset));
}, []);
const checkAssetAvailability = useCallback(async (assetId: string) => {
try {
const response = await fetch(`/assets/${assetId}`, { method: "HEAD", signal: AbortSignal.timeout(4000) });
if (response.status === 404) markAssetMissing(assetId);
} catch { /* A browser preview failure alone does not mean the upload expired. */ }
}, [markAssetMissing]);
const removeAsset = useCallback(async (assetId: string) => {
setAssetError("");
try {
const response = await fetch(`/assets/${assetId}`, { method: "DELETE" });
if (!response.ok && response.status !== 404) throw await responseError(response, "Could not remove the asset. Retry when the backend is available.");
setAssets((current) => current.filter((asset) => asset.asset_id !== assetId));
setConditioningAssets((current) => current.filter((asset) => asset.asset_id !== assetId));
} catch (error) {
setAssetError(error instanceof Error ? error.message : "Could not remove the asset.");
}
}, []);
const verifySelectedAssets = useCallback(async (): Promise<string | null> => {
const selected = [...new Set(conditioningAssets.map((item) => item.asset_id))];
try {
for (const assetId of selected) {
const response = await fetch(`/assets/${assetId}`, { method: "HEAD", signal: AbortSignal.timeout(4000) });
if (response.status === 404) {
markAssetMissing(assetId);
return "A selected asset expired or was removed from the server. Upload it again, then select the new copy.";
}
if (!response.ok) return "Could not verify the selected assets. Check the backend and try again.";
}
return null;
} catch {
return "Could not verify the selected assets. Check the backend and try again.";
}
}, [conditioningAssets, markAssetMissing]);
return {
assets, conditioningAssets, uploading, assetError, uploadAssets, assignAsset,
removeConditioning, moveConditioning, clearConditioning, removeAsset, checkAssetAvailability, verifySelectedAssets,
};
}
@@ -1,21 +0,0 @@
import { renderHook, waitFor } from "@testing-library/react";
import { afterEach, describe, expect, it, vi } from "vitest";
import { useGenerationCapabilities } from "./useGenerationCapabilities";
describe("generation capabilities", () => {
afterEach(() => vi.unstubAllGlobals());
it("enables only the modes the runtime advertises and identifies mock playback", async () => {
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true, json: async () => ({ model_id: "mock", modes: ["t2va", "fl2va", "ref2va"], mock: true }) })));
const { result } = renderHook(() => useGenerationCapabilities());
await waitFor(() => expect(result.current.loadingCapabilities).toBe(false));
expect(result.current.capabilities.modes).toEqual(["t2va", "fl2va", "ref2va"]);
expect(result.current.capabilities.mock).toBe(true);
});
it("keeps old runtimes text-only when the capabilities endpoint is missing", async () => {
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: false, status: 404 })));
const { result } = renderHook(() => useGenerationCapabilities());
await waitFor(() => expect(result.current.loadingCapabilities).toBe(false));
expect(result.current.capabilities.modes).toEqual(["t2va"]);
expect(result.current.capabilityNotice).toMatch(/Text-only compatibility/);
});
});
@@ -1,37 +0,0 @@
"use client";
import { useCallback, useEffect, useState } from "react";
import { isGenerationMode, type GenerationCapabilities } from "@/lib/generationMode";
const LEGACY_CAPABILITIES: GenerationCapabilities = { model_id: "legacy", modes: ["t2va"] };
export function useGenerationCapabilities() {
const [capabilities, setCapabilities] = useState<GenerationCapabilities>(LEGACY_CAPABILITIES);
const [capabilityNotice, setCapabilityNotice] = useState("");
const [loadingCapabilities, setLoadingCapabilities] = useState(true);
const refreshCapabilities = useCallback(async () => {
try {
const response = await fetch("/generation-capabilities", { signal: AbortSignal.timeout(4000) });
if (!response.ok) throw new Error("Capabilities unavailable");
const payload = await response.json();
if (typeof payload.model_id !== "string" || !Array.isArray(payload.modes)
|| !payload.modes.every(isGenerationMode)) throw new Error("Invalid capabilities");
const next: GenerationCapabilities = {
model_id: payload.model_id,
modes: payload.modes,
mock: payload.mock === true,
};
setCapabilities(next);
setCapabilityNotice("");
return next;
} catch {
setCapabilities(LEGACY_CAPABILITIES);
setCapabilityNotice("Runtime capabilities unavailable. Text-only compatibility mode is available; check the backend to enable image and reference modes.");
return LEGACY_CAPABILITIES;
} finally {
setLoadingCapabilities(false);
}
}, []);
useEffect(() => { void refreshCapabilities(); }, [refreshCapabilities]);
return { capabilities, capabilityNotice, loadingCapabilities, refreshCapabilities };
}
@@ -0,0 +1,80 @@
import { describe, expect, it } from "vitest";
import {
DEFAULT_LOBBY_CAPABILITIES_BUNDLE,
clampLobbySelectionToCapabilities,
parseLobbyCapabilitiesBundle,
resolveModelCapabilities,
validateLobbyCreationSelection,
} from "@/lib/creationCapabilities";
describe("creationCapabilities", () => {
it("parses backend capability payloads with per-model caps", () => {
const bundle = parseLobbyCapabilitiesBundle({
model_ids: ["fast-ltx2", "fast-h3"],
models: {
"fast-ltx2": {
generation_modes: ["t2va"],
resolutions: ["480p", "720p"],
duration_sec: [5, 10],
},
"fast-h3": {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["16:9"],
resolutions: ["720p"],
},
},
});
expect(bundle.model_ids).toEqual(["fast-ltx2", "fast-h3"]);
expect(bundle.models["fast-h3"]?.aspect_ratios).toEqual(["16:9"]);
});
it("includes fast-h3 in default lobby models", () => {
expect(DEFAULT_LOBBY_CAPABILITIES_BUNDLE.model_ids).toContain("fast-h3");
});
it("clamps unsupported lobby selections to model-specific defaults", () => {
expect(
clampLobbySelectionToCapabilities({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-h3"),
modelId: "fast-h3",
modeId: "fl2av",
aspectRatio: "9:16",
resolution: "4k",
durationSec: 99,
}),
).toEqual({
modelId: "fast-h3",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
});
});
it("rejects unsupported generation modes with a clear message", () => {
expect(
validateLobbyCreationSelection({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-ltx23"),
modelId: "fast-ltx23",
modeId: "fl2av",
aspectRatio: "16:9",
resolution: "720p",
durationSec: 5,
}),
).toMatch(/FL2VA/i);
});
it("rejects unsupported resolutions for ltx models", () => {
expect(
validateLobbyCreationSelection({
capabilities: resolveModelCapabilities(DEFAULT_LOBBY_CAPABILITIES_BUNDLE, "fast-ltx23"),
modelId: "fast-ltx23",
modeId: "t2v",
aspectRatio: "16:9",
resolution: "4k",
durationSec: 5,
}),
).toMatch(/resolution/i);
});
});
@@ -0,0 +1,263 @@
import type {
AspectRatioId,
CreationModeId,
CreationModelId,
ResolutionId,
} from "@/lib/creationConfig";
import { fromGenerationMode, toGenerationMode, type GenerationMode } from "@/lib/generationMode";
const ALL_MODEL_IDS: CreationModelId[] = ["fast-ltx23", "fast-ltx2", "fast-h3"];
const ALL_GENERATION_MODES: GenerationMode[] = ["t2va", "fl2va", "ref2va"];
const ALL_ASPECT_RATIOS: AspectRatioId[] = ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"];
const ALL_RESOLUTIONS: ResolutionId[] = ["480p", "720p", "1080p", "4k"];
export interface ModelCreationCapabilities {
generation_modes: GenerationMode[];
aspect_ratios: AspectRatioId[];
resolutions: ResolutionId[];
duration_sec: number[];
unsupported_generation_modes: Record<string, string>;
reference_assets: {
mime_types: string[];
max_bytes: number;
};
}
export interface LobbyCreationCapabilities extends ModelCreationCapabilities {
model_ids: CreationModelId[];
}
export interface LobbyCapabilitiesBundle {
model_ids: CreationModelId[];
models: Partial<Record<CreationModelId, ModelCreationCapabilities>>;
generation_modes: GenerationMode[];
aspect_ratios: AspectRatioId[];
resolutions: ResolutionId[];
duration_sec: number[];
unsupported_generation_modes: Record<string, string>;
reference_assets: {
mime_types: string[];
max_bytes: number;
};
}
const DEFAULT_LTX_MODEL_CAPABILITIES: ModelCreationCapabilities = {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"],
resolutions: ["480p", "720p", "1080p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: {
fl2va: "First/last frame mode (FL2VA) is not supported yet.",
},
reference_assets: {
mime_types: ["image/png", "image/jpeg", "image/webp"],
max_bytes: 15 * 1024 * 1024,
},
};
const DEFAULT_H3_MODEL_CAPABILITIES: ModelCreationCapabilities = {
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["16:9"],
resolutions: ["720p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: {
fl2va: "First/last frame mode (FL2VA) is not supported yet.",
},
reference_assets: DEFAULT_LTX_MODEL_CAPABILITIES.reference_assets,
};
export const DEFAULT_LOBBY_CAPABILITIES_BUNDLE: LobbyCapabilitiesBundle = {
model_ids: ALL_MODEL_IDS,
models: {
"fast-ltx2": DEFAULT_LTX_MODEL_CAPABILITIES,
"fast-ltx23": DEFAULT_LTX_MODEL_CAPABILITIES,
"fast-h3": DEFAULT_H3_MODEL_CAPABILITIES,
},
generation_modes: ["t2va", "ref2va"],
aspect_ratios: ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"],
resolutions: ["480p", "720p", "1080p"],
duration_sec: [5, 10, 15],
unsupported_generation_modes: DEFAULT_LTX_MODEL_CAPABILITIES.unsupported_generation_modes,
reference_assets: DEFAULT_LTX_MODEL_CAPABILITIES.reference_assets,
};
function pickStrings<T extends string>(value: unknown, allowed: readonly T[], fallback: readonly T[]): T[] {
if (!Array.isArray(value)) return [...fallback];
return value.filter((item): item is T => typeof item === "string" && allowed.includes(item as T));
}
function parseReferenceAssets(
value: unknown,
fallback: ModelCreationCapabilities["reference_assets"],
): ModelCreationCapabilities["reference_assets"] {
if (!value || typeof value !== "object") return fallback;
const data = value as Record<string, unknown>;
return {
mime_types: Array.isArray(data.mime_types)
? (data.mime_types as string[])
: fallback.mime_types,
max_bytes: typeof data.max_bytes === "number" ? data.max_bytes : fallback.max_bytes,
};
}
function parseModelCreationCapabilities(
value: unknown,
fallback: ModelCreationCapabilities,
): ModelCreationCapabilities {
if (!value || typeof value !== "object") return fallback;
const data = value as Record<string, unknown>;
return {
generation_modes: pickStrings(data.generation_modes, ALL_GENERATION_MODES, fallback.generation_modes),
aspect_ratios: pickStrings(data.aspect_ratios, ALL_ASPECT_RATIOS, fallback.aspect_ratios),
resolutions: pickStrings(data.resolutions, ALL_RESOLUTIONS, fallback.resolutions),
duration_sec: Array.isArray(data.duration_sec)
? data.duration_sec.filter((item): item is number => typeof item === "number")
: fallback.duration_sec,
unsupported_generation_modes:
typeof data.unsupported_generation_modes === "object" && data.unsupported_generation_modes
? (data.unsupported_generation_modes as Record<string, string>)
: fallback.unsupported_generation_modes,
reference_assets: parseReferenceAssets(data.reference_assets, fallback.reference_assets),
};
}
export function parseLobbyCapabilitiesBundle(payload: unknown): LobbyCapabilitiesBundle {
if (!payload || typeof payload !== "object") {
return DEFAULT_LOBBY_CAPABILITIES_BUNDLE;
}
const data = payload as Record<string, unknown>;
const modelIds = pickStrings(data.model_ids, ALL_MODEL_IDS, DEFAULT_LOBBY_CAPABILITIES_BUNDLE.model_ids);
const rawModels = typeof data.models === "object" && data.models ? (data.models as Record<string, unknown>) : {};
const models: Partial<Record<CreationModelId, ModelCreationCapabilities>> = {};
for (const modelId of modelIds) {
const fallback =
DEFAULT_LOBBY_CAPABILITIES_BUNDLE.models[modelId] ??
(modelId === "fast-h3" ? DEFAULT_H3_MODEL_CAPABILITIES : DEFAULT_LTX_MODEL_CAPABILITIES);
models[modelId] = parseModelCreationCapabilities(rawModels[modelId], fallback);
}
const unionFallback = parseModelCreationCapabilities(payload, DEFAULT_LTX_MODEL_CAPABILITIES);
return {
model_ids: modelIds,
models,
generation_modes: unionFallback.generation_modes,
aspect_ratios: unionFallback.aspect_ratios,
resolutions: unionFallback.resolutions,
duration_sec: unionFallback.duration_sec,
unsupported_generation_modes: unionFallback.unsupported_generation_modes,
reference_assets: unionFallback.reference_assets,
};
}
export function resolveModelCapabilities(
bundle: LobbyCapabilitiesBundle,
modelId: CreationModelId,
): LobbyCreationCapabilities {
const modelCaps =
bundle.models[modelId] ??
(modelId === "fast-h3" ? DEFAULT_H3_MODEL_CAPABILITIES : DEFAULT_LTX_MODEL_CAPABILITIES);
return {
model_ids: bundle.model_ids,
...modelCaps,
};
}
export function supportedCreationModes(capabilities: LobbyCreationCapabilities) {
return capabilities.generation_modes.map((wireMode) => ({
wireMode,
modeId: fromGenerationMode(wireMode),
}));
}
export function isSupportedCreationMode(modeId: CreationModeId, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.generation_modes.includes(toGenerationMode(modeId));
}
export function isSupportedResolution(resolution: ResolutionId, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.resolutions.includes(resolution);
}
export function isSupportedReferenceImage(file: File, capabilities: LobbyCreationCapabilities): boolean {
return capabilities.reference_assets.mime_types.includes(file.type);
}
export function unsupportedModeNotice(modeId: CreationModeId, capabilities: LobbyCreationCapabilities): string | null {
const wireMode = toGenerationMode(modeId);
return capabilities.unsupported_generation_modes[wireMode] ?? null;
}
export function clampLobbySelectionToCapabilities(input: {
capabilities: LobbyCreationCapabilities;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}): {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
} {
const { capabilities } = input;
const modelId = capabilities.model_ids.includes(input.modelId)
? input.modelId
: (capabilities.model_ids[0] ?? "fast-ltx23");
const supportedModes = supportedCreationModes(capabilities);
const modeId = isSupportedCreationMode(input.modeId, capabilities)
? input.modeId
: (supportedModes[0]?.modeId ?? "t2v");
const aspectRatio = capabilities.aspect_ratios.includes(input.aspectRatio)
? input.aspectRatio
: (capabilities.aspect_ratios[0] ?? "16:9");
const resolution = isSupportedResolution(input.resolution, capabilities)
? input.resolution
: (capabilities.resolutions[0] ?? "720p");
const durationSec = capabilities.duration_sec.includes(input.durationSec)
? input.durationSec
: (capabilities.duration_sec[0] ?? 5);
return { modelId, modeId, aspectRatio, resolution, durationSec };
}
export function validateLobbyCreationSelection(input: {
capabilities: LobbyCreationCapabilities;
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): string | null {
const unsupportedMode = unsupportedModeNotice(input.modeId, input.capabilities);
if (unsupportedMode) return unsupportedMode;
if (!input.capabilities.model_ids.includes(input.modelId)) {
return "Selected model is not supported yet.";
}
if (!isSupportedCreationMode(input.modeId, input.capabilities)) {
return "Selected mode is not supported yet.";
}
if (!input.capabilities.aspect_ratios.includes(input.aspectRatio)) {
return "Selected aspect ratio is not supported for this model yet.";
}
if (!isSupportedResolution(input.resolution, input.capabilities)) {
return "Selected resolution is not supported for this model yet.";
}
if (!input.capabilities.duration_sec.includes(input.durationSec)) {
return "Selected duration is not supported yet.";
}
if (input.modeId === "ref2av" && !input.referenceFile) {
return "Upload a reference image to use reference-guided mode.";
}
if (input.referenceFile && !isSupportedReferenceImage(input.referenceFile, input.capabilities)) {
return "Reference assets must be PNG, JPEG, or WebP images.";
}
if (input.firstFrameFile && !isSupportedReferenceImage(input.firstFrameFile, input.capabilities)) {
return "First frame must be a PNG, JPEG, or WebP image.";
}
if (input.lastFrameFile && !isSupportedReferenceImage(input.lastFrameFile, input.capabilities)) {
return "Last frame must be a PNG, JPEG, or WebP image.";
}
return null;
}
@@ -0,0 +1,62 @@
import { describe, expect, it } from "vitest";
import {
CREATION_MODELS,
buildMentionOptions,
formatDurationLabel,
formatResolutionLabel,
isReferenceMediaFile,
modeRequiresReference,
modeUsesDualFrames,
} from "@/lib/creationConfig";
describe("creationConfig", () => {
it("formats resolution labels", () => {
expect(formatResolutionLabel("480p")).toBe("480P");
expect(formatResolutionLabel("720p")).toBe("720P");
expect(formatResolutionLabel("4k")).toBe("4K");
});
it("formats duration labels", () => {
expect(formatDurationLabel(5)).toBe("5s");
});
it("includes all Dreamverse lobby models", () => {
expect(CREATION_MODELS.map((model) => model.id)).toEqual(["fast-ltx23", "fast-ltx2", "fast-h3"]);
});
it("builds mention options from presets", () => {
expect(
buildMentionOptions([
{ id: "preset-a", label: "Preset A", description: "A short preset" },
{ label: "Missing id" },
]),
).toEqual([
{
id: "preset-a",
label: "Preset A",
kind: "preset",
description: "A short preset",
},
{
id: "Missing id",
label: "Missing id",
kind: "preset",
description: undefined,
},
]);
});
it("derives mode-specific reference requirements", () => {
expect(modeRequiresReference("ref2av")).toBe(true);
expect(modeRequiresReference("t2v")).toBe(false);
expect(modeUsesDualFrames("fl2av")).toBe(true);
expect(modeUsesDualFrames("t2v")).toBe(false);
});
it("accepts image reference files only", () => {
expect(isReferenceMediaFile(new File(["x"], "a.png", { type: "image/png" }))).toBe(true);
expect(isReferenceMediaFile(new File(["x"], "a.mp4", { type: "video/mp4" }))).toBe(false);
expect(isReferenceMediaFile(new File(["x"], "a.txt", { type: "text/plain" }))).toBe(false);
});
});
@@ -0,0 +1,96 @@
export type CreationModeId = "t2v" | "fl2av" | "ref2av";
export type CreationModelId = "fast-ltx2" | "fast-ltx23" | "fast-h3";
export type AspectRatioId = "21:9" | "16:9" | "4:3" | "1:1" | "3:4" | "9:16";
export type ResolutionId = "480p" | "720p" | "1080p" | "4k";
export interface CreationModeOption {
id: CreationModeId;
label: string;
description: string;
}
export interface CreationModelOption {
id: CreationModelId;
label: string;
description: string;
badge?: string;
}
export interface MentionOption {
id: string;
label: string;
kind: "preset" | "asset" | "character";
description?: string;
}
export const CREATION_MODES: CreationModeOption[] = [
{ id: "t2v", label: "Text to video", description: "Generate from a text prompt" },
{ id: "ref2av", label: "Image to video", description: "Guide the first segment with a reference image" },
];
export const UNSUPPORTED_CREATION_MODES: CreationModeOption[] = [
{ id: "fl2av", label: "First and last frame", description: "Coming soon on FastLTX models" },
];
export const CREATION_MODELS: CreationModelOption[] = [
{
id: "fast-ltx23",
label: "FastLTX 2.3",
description: "LTX 2.3 with OmniNFT LoRA",
badge: "New",
},
{
id: "fast-ltx2",
label: "FastLTX 2",
description: "FastLTX 2 for streaming",
},
{
id: "fast-h3",
label: "FastH3",
description: "MiniMax H3 with VSA data-free adapter",
},
];
export const ASPECT_RATIOS: AspectRatioId[] = ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"];
export const RESOLUTIONS: ResolutionId[] = ["480p", "720p", "1080p"];
export const UNSUPPORTED_RESOLUTIONS: ResolutionId[] = ["4k"];
export const DURATION_MARKS = [5, 10, 15] as const;
export const REFERENCE_ACCEPT = "image/png,image/jpeg,image/webp";
export function formatResolutionLabel(resolution: ResolutionId): string {
return resolution === "4k" ? "4K" : resolution.toUpperCase();
}
export function formatDurationLabel(seconds: number): string {
return `${seconds}s`;
}
export function modeRequiresReference(modeId: CreationModeId): boolean {
return modeId === "ref2av";
}
export function modeUsesDualFrames(modeId: CreationModeId): boolean {
return modeId === "fl2av";
}
export function isReferenceMediaFile(file: File): boolean {
return file.type === "image/png" || file.type === "image/jpeg" || file.type === "image/webp";
}
export function buildMentionOptions(storyPresets: Array<{ id?: string; label?: string; description?: string }>): MentionOption[] {
return storyPresets
.filter((preset) => typeof preset.label === "string" && preset.label.trim())
.map((preset) => ({
id: String(preset.id || preset.label),
label: String(preset.label),
kind: "preset" as const,
description: typeof preset.description === "string" ? preset.description : undefined,
}));
}
@@ -0,0 +1,66 @@
import { describe, expect, it } from "vitest";
import { parseEchoedCreationConfig, validateCreationInputs } from "@/lib/creationPayload";
describe("creationPayload", () => {
it("requires a reference asset for omni reference mode", () => {
expect(
validateCreationInputs({
modeId: "ref2av",
referenceFile: null,
}),
).toMatch(/reference asset/i);
});
it("requires both frames for first and last frame mode", () => {
expect(
validateCreationInputs({
modeId: "fl2av",
firstFrameFile: new File(["a"], "first.png", { type: "image/png" }),
lastFrameFile: null,
}),
).toMatch(/both first and last/i);
});
it("accepts text to video without references", () => {
expect(
validateCreationInputs({
modeId: "t2v",
}),
).toBeNull();
});
it("parses echoed creation config from server payloads", () => {
expect(
parseEchoedCreationConfig({
type: "gpu_assigned",
creation_config: {
model_id: "fast-ltx2",
generation_mode: "ref2va",
aspect_ratio: "9:16",
resolution: "480p",
duration_sec: 10,
},
}),
).toEqual({
modelId: "fast-ltx2",
modeId: "ref2av",
aspectRatio: "9:16",
resolution: "480p",
durationSec: 10,
});
});
it("ignores invalid echoed creation config", () => {
expect(parseEchoedCreationConfig({ creation_config: { model_id: "unknown" } })).toBeNull();
});
it("rejects unsupported reference mime types", () => {
expect(
validateCreationInputs({
modeId: "t2v",
referenceFile: new File(["a"], "clip.mp4", { type: "video/mp4" }),
}),
).toMatch(/PNG, JPEG, or WebP/i);
});
});
@@ -0,0 +1,172 @@
import type {
AspectRatioId,
CreationModeId,
CreationModelId,
ResolutionId,
} from "@/lib/creationConfig";
import { fromGenerationMode, type GenerationMode } from "@/lib/generationMode";
const LOBBY_MODEL_IDS = new Set<CreationModelId>(["fast-ltx2", "fast-ltx23", "fast-h3"]);
const ASPECT_RATIO_IDS = new Set<AspectRatioId>(["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"]);
const RESOLUTION_IDS = new Set<ResolutionId>(["480p", "720p", "1080p", "4k"]);
const DURATION_SEC_VALUES = new Set([5, 10, 15]);
export interface EchoedSessionCreationConfig {
modelId: CreationModelId;
modeId: CreationModeId;
aspectRatio: AspectRatioId;
resolution: ResolutionId;
durationSec: number;
}
const MAX_IMAGE_BYTES = 15 * 1024 * 1024;
const SUPPORTED_IMAGE_TYPES = new Set(["image/png", "image/jpeg", "image/webp"]);
export interface InitialImagePayload {
name: string;
mime_type: string;
data_url: string;
}
export interface CreationInitPayload {
model_id: string;
aspect_ratio: string;
resolution: string;
duration_sec: number;
initial_image: InitialImagePayload | null;
last_frame_image: InitialImagePayload | null;
}
function readFileAsDataUrl(file: File): Promise<string> {
return new Promise((resolve, reject) => {
const reader = new FileReader();
reader.onload = () => {
if (typeof reader.result === "string") {
resolve(reader.result);
return;
}
reject(new Error("Failed to read reference image."));
};
reader.onerror = () => reject(new Error("Failed to read reference image."));
reader.readAsDataURL(file);
});
}
export async function fileToInitialImagePayload(file: File): Promise<InitialImagePayload> {
if (!SUPPORTED_IMAGE_TYPES.has(file.type)) {
throw new Error("Reference assets must be PNG, JPEG, or WebP images.");
}
if (file.size > MAX_IMAGE_BYTES) {
throw new Error("Reference image must be 15 MB or smaller.");
}
return {
name: file.name,
mime_type: file.type,
data_url: await readFileAsDataUrl(file),
};
}
export async function resolveCreationImages(input: {
modeId: CreationModeId;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): Promise<Pick<CreationInitPayload, "initial_image" | "last_frame_image">> {
if (input.modeId === "fl2av") {
const firstFrame = input.firstFrameFile ? await fileToInitialImagePayload(input.firstFrameFile) : null;
const lastFrame = input.lastFrameFile ? await fileToInitialImagePayload(input.lastFrameFile) : null;
return {
initial_image: firstFrame,
last_frame_image: lastFrame,
};
}
const reference = input.referenceFile ? await fileToInitialImagePayload(input.referenceFile) : null;
return {
initial_image: reference,
last_frame_image: null,
};
}
export function validateCreationInputs(input: {
modeId: CreationModeId;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): string | null {
if (input.modeId === "ref2av" && !input.referenceFile) {
return "Upload a reference asset to use Omni reference mode.";
}
if (input.modeId === "fl2av") {
if (!input.firstFrameFile || !input.lastFrameFile) {
return "Upload both first and last frame assets.";
}
}
if (input.referenceFile && !SUPPORTED_IMAGE_TYPES.has(input.referenceFile.type)) {
return "Reference assets must be PNG, JPEG, or WebP images.";
}
if (input.firstFrameFile && !SUPPORTED_IMAGE_TYPES.has(input.firstFrameFile.type)) {
return "First frame must be a PNG, JPEG, or WebP image.";
}
if (input.lastFrameFile && !SUPPORTED_IMAGE_TYPES.has(input.lastFrameFile.type)) {
return "Last frame must be a PNG, JPEG, or WebP image.";
}
return null;
}
export function parseEchoedCreationConfig(data: unknown): EchoedSessionCreationConfig | null {
if (!data || typeof data !== "object") {
return null;
}
const creationConfig = (data as Record<string, unknown>).creation_config;
if (!creationConfig || typeof creationConfig !== "object") {
return null;
}
const config = creationConfig as Record<string, unknown>;
const modelId = typeof config.model_id === "string" && LOBBY_MODEL_IDS.has(config.model_id as CreationModelId)
? (config.model_id as CreationModelId)
: null;
const generationMode = typeof config.generation_mode === "string" ? config.generation_mode as GenerationMode : null;
const modeId = generationMode === "t2va" || generationMode === "fl2va" || generationMode === "ref2va"
? fromGenerationMode(generationMode)
: null;
const aspectRatio = typeof config.aspect_ratio === "string" && ASPECT_RATIO_IDS.has(config.aspect_ratio as AspectRatioId)
? (config.aspect_ratio as AspectRatioId)
: null;
const resolution = typeof config.resolution === "string" && RESOLUTION_IDS.has(config.resolution as ResolutionId)
? (config.resolution as ResolutionId)
: null;
const durationSec = typeof config.duration_sec === "number" && DURATION_SEC_VALUES.has(config.duration_sec)
? config.duration_sec
: null;
if (modelId === null || modeId === null || aspectRatio === null || resolution === null || durationSec === null) {
return null;
}
return {
modelId,
modeId,
aspectRatio,
resolution,
durationSec,
};
}
export async function buildCreationInitPayload(input: {
modelId: string;
modeId: CreationModeId;
aspectRatio: string;
resolution: string;
durationSec: number;
referenceFile?: File | null;
firstFrameFile?: File | null;
lastFrameFile?: File | null;
}): Promise<CreationInitPayload> {
const images = await resolveCreationImages(input);
return {
model_id: input.modelId,
aspect_ratio: input.aspectRatio,
resolution: input.resolution,
duration_sec: input.durationSec,
...images,
};
}
@@ -3,11 +3,10 @@ import { describe, expect, it } from "vitest";
import {
DEFAULT_GENERATION_MODE,
GENERATION_MODES,
fromGenerationMode,
getGenerationMode,
isGenerationMode,
buildGenerationInitFields,
validateGenerationInputs,
type GenerationAsset,
toGenerationMode,
} from "./generationMode";
describe("generation modes", () => {
@@ -25,34 +24,16 @@ describe("generation modes", () => {
expect(isGenerationMode("unknown")).toBe(false);
expect(getGenerationMode("fl2va").label).toBe("FL2VA");
});
});
const image: GenerationAsset = { asset_id: "img", kind: "image", name: "frame.png", mime_type: "image/png", size: 12, url: "/assets/img" };
const audio: GenerationAsset = { asset_id: "sound", kind: "audio", name: "sound.wav", mime_type: "audio/wav", size: 12, url: "/assets/sound" };
it("maps creation studio mode IDs to upstream wire values", () => {
expect(toGenerationMode("t2v")).toBe("t2va");
expect(toGenerationMode("fl2av")).toBe("fl2va");
expect(toGenerationMode("ref2av")).toBe("ref2va");
});
describe("generation input contract", () => {
it("keeps text-only init valid and rejects accidental references", () => {
expect(buildGenerationInitFields("t2va", [], [])).toEqual({ generation_mode: "t2va", conditioning_assets: [] });
expect(() => buildGenerationInitFields("t2va", [{ asset_id: "img", role: "reference" }], [image])).toThrow("text only");
});
it("requires first frame but permits first-only or both endpoints", () => {
expect(validateGenerationInputs("fl2va", [], [])).toMatch(/first frame/);
const first = { asset_id: "img", role: "first_frame" } as const;
expect(validateGenerationInputs("fl2va", [first], [image])).toBeNull();
expect(validateGenerationInputs("fl2va", [first, { asset_id: "img", role: "last_frame" }], [image])).toBeNull();
expect(validateGenerationInputs("fl2va", [first, first], [image])).toMatch(/one first frame/);
expect(validateGenerationInputs("fl2va", [{ asset_id: "sound", role: "first_frame" }], [audio])).toMatch(/images only/);
});
it("requires a visual reference and preserves multimodal ordering without file bodies", () => {
expect(validateGenerationInputs("ref2va", [{ asset_id: "sound", role: "reference" }], [audio])).toMatch(/image or video/);
const items = [{ asset_id: "sound", role: "reference" }, { asset_id: "img", role: "reference" }] as const;
expect(buildGenerationInitFields("ref2va", items, [image, audio])).toEqual({ generation_mode: "ref2va", conditioning_assets: items });
});
it("rejects per-kind limits, total limits, and stale uploads", () => {
const images = Array.from({ length: 10 }, (_, index) => ({ ...image, asset_id: `img-${index}` }));
expect(validateGenerationInputs("ref2va", images.map((item) => ({ asset_id: item.asset_id, role: "reference" })), images)).toMatch(/at most 9 image/);
const refs = Array.from({ length: 13 }, () => ({ asset_id: "img", role: "reference" as const }));
expect(validateGenerationInputs("ref2va", refs, [image])).toMatch(/at most 12/);
expect(validateGenerationInputs("fl2va", [{ asset_id: "img", role: "first_frame" }], [{ ...image, missing: true }])).toMatch(/no longer on the server/);
it("maps upstream wire values back to creation studio mode IDs", () => {
expect(fromGenerationMode("t2va")).toBe("t2v");
expect(fromGenerationMode("fl2va")).toBe("fl2av");
expect(fromGenerationMode("ref2va")).toBe("ref2av");
});
});
+18 -78
View File
@@ -1,3 +1,5 @@
import type { CreationModeId } from "@/lib/creationConfig";
export const GENERATION_MODES = [
{
id: "t2va",
@@ -9,7 +11,7 @@ export const GENERATION_MODES = [
id: "fl2va",
label: "FL2VA",
name: "First/last frames to video + audio",
description: "Start from a first frame image. Add an optional last frame to guide the ending.",
description: "Provide first and last frame images to control the transition.",
},
{
id: "ref2va",
@@ -23,6 +25,12 @@ export type GenerationMode = (typeof GENERATION_MODES)[number]["id"];
export const DEFAULT_GENERATION_MODE: GenerationMode = "t2va";
const CREATION_MODE_TO_GENERATION_MODE: Record<CreationModeId, GenerationMode> = {
t2v: "t2va",
fl2av: "fl2va",
ref2av: "ref2va",
};
export function isGenerationMode(value: unknown): value is GenerationMode {
return GENERATION_MODES.some((mode) => mode.id === value);
}
@@ -31,84 +39,16 @@ export function getGenerationMode(value: GenerationMode) {
return GENERATION_MODES.find((mode) => mode.id === value) ?? GENERATION_MODES[0];
}
export type AssetKind = "image" | "video" | "audio";
export type ConditioningRole = "first_frame" | "last_frame" | "reference";
const GENERATION_MODE_TO_CREATION_MODE: Record<GenerationMode, CreationModeId> = {
t2va: "t2v",
fl2va: "fl2av",
ref2va: "ref2av",
};
/** Runtime-owned uploads; project metadata keeps references, never file contents. */
export interface GenerationAsset {
asset_id: string;
kind: AssetKind;
name: string;
mime_type: string;
size: number;
url: string;
missing?: boolean;
export function fromGenerationMode(mode: GenerationMode): CreationModeId {
return GENERATION_MODE_TO_CREATION_MODE[mode];
}
export interface ConditioningAsset {
asset_id: string;
role: ConditioningRole;
}
export interface GenerationCapabilities {
model_id: string;
modes: GenerationMode[];
mock?: boolean;
}
export interface GenerationInitFields {
generation_mode: GenerationMode;
conditioning_assets: ConditioningAsset[];
}
export const REFERENCE_LIMITS = { image: 9, video: 3, audio: 3, total: 12 } as const;
export function validateGenerationInputs(
mode: GenerationMode,
conditioning: readonly ConditioningAsset[],
assets: readonly GenerationAsset[],
): string | null {
if (mode === "t2va") {
return conditioning.length ? "T2VA uses text only. Remove the selected references." : null;
}
const resolved = conditioning.map((item) => assets.find((asset) => asset.asset_id === item.asset_id));
if (resolved.some((asset) => !asset || asset.missing)) {
return "A selected asset is no longer on the server. Upload it again and select the new copy.";
}
if (mode === "fl2va") {
if (!conditioning.some((item) => item.role === "first_frame")) return "Choose a first frame image to generate.";
if (conditioning.some((item) => item.role === "reference") || resolved.some((asset) => asset?.kind !== "image")) {
return "FL2VA accepts first and last frame images only.";
}
if (conditioning.filter((item) => item.role === "first_frame").length !== 1
|| conditioning.filter((item) => item.role === "last_frame").length > 1) {
return "Choose one first frame and at most one last frame.";
}
return null;
}
if (conditioning.some((item) => item.role !== "reference")) return "Ref2VA accepts ordered reference assets only.";
if (!resolved.some((asset) => asset?.kind === "image" || asset?.kind === "video")) {
return "Add at least one image or video reference. Audio alone is not enough.";
}
if (conditioning.length > REFERENCE_LIMITS.total) return "Use at most 12 reference assets in total.";
for (const kind of ["image", "video", "audio"] as const) {
if (resolved.filter((asset) => asset?.kind === kind).length > REFERENCE_LIMITS[kind]) {
return `Use at most ${REFERENCE_LIMITS[kind]} ${kind} references.`;
}
}
return null;
}
/** Shared by both session_init_v2 and project_init_v1. */
export function buildGenerationInitFields(
mode: GenerationMode,
conditioning: readonly ConditioningAsset[],
assets: readonly GenerationAsset[],
): GenerationInitFields {
const problem = validateGenerationInputs(mode, conditioning, assets);
if (problem) throw new Error(problem);
return {
generation_mode: mode,
conditioning_assets: conditioning.map(({ asset_id, role }) => ({ asset_id, role })),
};
export function toGenerationMode(modeId: CreationModeId): GenerationMode {
return CREATION_MODE_TO_GENERATION_MODE[modeId];
}
@@ -1,5 +1,3 @@
import type { ConditioningAsset, GenerationAsset, GenerationMode } from "./generationMode";
const DB_NAME = "fastvideo-projects";
const DB_VERSION = 1;
const PROJECTS_STORE = "projects";
@@ -37,11 +35,6 @@ export interface StoredProject {
createdAt: number;
lastThumbnail: string | null;
promptEvents: Record<string, unknown>[];
/** Optional for projects created before generation modes were introduced. */
generationMode?: GenerationMode;
conditioningAssets?: ConditioningAsset[];
assets?: GenerationAsset[];
mock?: boolean;
}
export interface StoredClip {
@@ -1,63 +0,0 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
test.describe('create job interactions', () => {
skipWithoutMock();
for (const jobType of ['inference', 'finetuning', 'distillation']) {
test(`${jobType} remains interactive after repeated dialog dismissals`, async ({ page }) => {
await page.goto(`/${jobType}`);
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
const dialog = page.getByRole('dialog');
// Exercise both dismissal paths and reopen without reloading the page.
for (const closeWithEscape of [false, true]) {
await trigger.click();
await page.getByRole('menuitem').first().click();
await expect(dialog).toBeVisible();
if (closeWithEscape) {
await page.keyboard.press('Escape');
} else {
await dialog.getByRole('button', { name: 'Close', exact: true }).click();
}
await expect(dialog).toBeHidden();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
await expect(trigger).toBeFocused();
}
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
await expect(page).toHaveURL(/\/datasets$/);
});
}
test('preserves keyboard menu dismissal and dialog focus trapping', async ({ page }) => {
await page.goto('/inference');
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
await trigger.focus();
await page.keyboard.press('Enter');
const firstItem = page.getByRole('menuitem').first();
await expect(firstItem).toBeFocused();
await page.keyboard.press('Escape');
await expect(page.getByRole('menu')).toBeHidden();
await expect(trigger).toBeFocused();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
await page.keyboard.press('Enter');
await expect(firstItem).toBeFocused();
await page.keyboard.press('Enter');
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
// Shift+Tab from the first field wraps to Close, then Tab wraps back.
await page.keyboard.press('Shift+Tab');
await expect(dialog.getByRole('button', { name: 'Close', exact: true })).toBeFocused();
await page.keyboard.press('Tab');
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
await page.keyboard.press('Escape');
await expect(dialog).toBeHidden();
await expect(trigger).toBeFocused();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
});
});
+2 -18
View File
@@ -1,6 +1,6 @@
import { expect, test } from '@playwright/test';
import { API_BASE, skipWithoutMock } from './helpers';
import { skipWithoutMock } from './helpers';
/**
* Create-job flow: open the Create Job modal on /inference, fill the prompt
@@ -10,8 +10,7 @@ import { API_BASE, skipWithoutMock } from './helpers';
test.describe('create inference job', () => {
skipWithoutMock();
test('creates a T2V job and starts it without refreshing', async ({ page, request }) => {
await request.put(`${API_BASE}/settings`, { data: { autoStartJob: false } });
test('creates a T2V job and shows it in the queue', async ({ page }) => {
await page.goto('/inference');
// The trigger opens a real menu on click, so this path works for touch,
@@ -39,20 +38,5 @@ test.describe('create inference job', () => {
// Modal closes and the queue refreshes with the newly created job.
await expect(dialog).toBeHidden();
await expect(page.getByText(prompt)).toBeVisible();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
const card = page.getByRole('article').filter({ hasText: prompt });
await expect(card.getByText('pending', { exact: true })).toBeVisible();
const started = page.waitForResponse((response) =>
response.url().startsWith(`${API_BASE}/jobs/`) &&
response.url().endsWith('/start') &&
response.request().method() === 'POST',
);
await card.getByRole('button', { name: 'Start', exact: true }).click();
expect((await started).ok()).toBe(true);
await expect(card.getByText('running', { exact: true })).toBeVisible();
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
await expect(page).toHaveURL(/\/datasets$/);
});
});
+829 -975
View File
File diff suppressed because it is too large Load Diff
+10 -1
View File
@@ -17,11 +17,20 @@
"start:all": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\""
},
"dependencies": {
"@radix-ui/react-dialog": "^1.1.0",
"@radix-ui/react-dropdown-menu": "^2.1.24",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.2.0",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-switch": "^1.1.0",
"@radix-ui/react-tabs": "^1.1.0",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"lucide-react": "^0.577.0",
"next": "15.5.18",
"radix-ui": "^1.6.7",
"react": "^19.1.0",
"react-dom": "^19.1.0",
"sonner": "^2.0.7",
@@ -1,26 +1,24 @@
import { render, screen, waitFor, within } from '@testing-library/react';
import { render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { describe, expect, it, vi } from 'vitest';
import CreateJobButton from './CreateJobButton';
import { getDatasets, getModels } from '@/lib/api';
vi.mock('@/lib/api', () => ({
createJob: vi.fn(),
getModels: vi.fn(),
getDatasets: vi.fn(),
uploadImage: vi.fn(),
getSettings: vi.fn(),
updateSettings: vi.fn(),
vi.mock('./CreateJobModal', () => ({
default: ({
isOpen,
workloadType,
}: {
isOpen: boolean;
workloadType: string;
}) =>
isOpen ? (
<div role="dialog" data-workload-type={workloadType}>
Create job form
</div>
) : null,
}));
beforeEach(() => {
vi.mocked(getModels).mockResolvedValue([
{ id: 'wan/t2v-1.3b', label: 'Wan T2V' },
]);
vi.mocked(getDatasets).mockResolvedValue([]);
});
describe('CreateJobButton', () => {
it('opens the workload menu on click and selects an item', async () => {
const user = userEvent.setup();
@@ -29,9 +27,10 @@ describe('CreateJobButton', () => {
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await user.click(screen.getByRole('menuitem', { name: /I2V/i }));
expect(
screen.getByRole('dialog', { name: 'New Inference Job (I2V)' }),
).toBeInTheDocument();
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
'i2v',
);
});
it('opens and operates the workload menu from the keyboard', async () => {
@@ -46,42 +45,9 @@ describe('CreateJobButton', () => {
expect(firstItem).toHaveFocus();
await user.keyboard('{Enter}');
expect(
screen.getByRole('dialog', { name: 'New Inference Job (T2V)' }),
).toBeInTheDocument();
await user.keyboard('{Escape}');
await waitFor(() =>
expect(screen.queryByRole('dialog')).not.toBeInTheDocument(),
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
't2v',
);
await waitFor(() =>
expect(document.body.style.pointerEvents).not.toBe('none'),
);
expect(trigger).toHaveFocus();
});
it.each(['inference', 'finetuning', 'distillation'] as const)(
'restores page interaction after closing the real %s dialog',
async (jobType) => {
const user = userEvent.setup();
render(<CreateJobButton jobType={jobType} />);
const trigger = screen.getByRole('button', { name: 'Create Job' });
// Keep the real Dialog mounted: mocking it hides conflicting Radix layers.
for (let attempt = 0; attempt < 2; attempt++) {
await user.click(trigger);
await user.click(screen.getAllByRole('menuitem')[0]);
const dialog = screen.getByRole('dialog');
await user.click(
within(dialog).getByRole('button', { name: 'Close' }),
);
await waitFor(() =>
expect(screen.queryByRole('dialog')).not.toBeInTheDocument(),
);
await waitFor(() =>
expect(document.body.style.pointerEvents).not.toBe('none'),
);
expect(trigger).toHaveFocus();
}
},
);
});
@@ -2,7 +2,7 @@
import * as React from 'react';
import { ChevronDown } from 'lucide-react';
import { DropdownMenu } from 'radix-ui';
import * as DropdownMenu from '@radix-ui/react-dropdown-menu';
import CreateJobModal from '@/components/jobs/CreateJobModal';
import { Button } from '@/components/ui/button';
@@ -16,7 +16,6 @@ interface CreateJobButtonProps {
export default function CreateJobButton({ jobType }: CreateJobButtonProps) {
const options = WORKLOAD_OPTIONS[jobType] ?? [];
const triggerRef = React.useRef<HTMLButtonElement>(null);
const [modalOpen, setModalOpen] = React.useState(false);
const [workloadType, setWorkloadType] = React.useState(
@@ -37,7 +36,7 @@ export default function CreateJobButton({ jobType }: CreateJobButtonProps) {
<>
<DropdownMenu.Root>
<DropdownMenu.Trigger asChild>
<Button ref={triggerRef} type="button" className="gap-1.5">
<Button type="button" className="gap-1.5">
Create Job
<ChevronDown className="size-3.5 opacity-85" aria-hidden />
</Button>
@@ -67,11 +66,6 @@ export default function CreateJobButton({ jobType }: CreateJobButtonProps) {
<CreateJobModal
isOpen={modalOpen}
onClose={() => setModalOpen(false)}
onCloseAutoFocus={(event) => {
// This dialog opens from a menu item, so it has no DialogTrigger.
event.preventDefault();
triggerRef.current?.focus();
}}
onSuccess={handleSuccess}
jobType={jobType}
workloadType={workloadType}
@@ -55,9 +55,6 @@ import { jobToFormFields, type JobLike } from '@/lib/jobToFields';
export interface CreateJobModalProps {
isOpen: boolean;
onClose: () => void;
onCloseAutoFocus?: React.ComponentProps<
typeof DialogContent
>['onCloseAutoFocus'];
onSuccess: () => void;
jobType: JobType;
workloadType: string;
@@ -70,7 +67,6 @@ export interface CreateJobModalProps {
export default function CreateJobModal({
isOpen,
onClose,
onCloseAutoFocus,
onSuccess,
jobType,
workloadType,
@@ -648,7 +644,6 @@ export default function CreateJobModal({
>
<DialogContent
className="max-h-[90vh] w-[90vw] max-w-[850px] overflow-y-auto"
onCloseAutoFocus={onCloseAutoFocus}
onEscapeKeyDown={(e) => {
if (isSubmitting) e.preventDefault();
}}

Some files were not shown because too many files have changed in this diff Show More