Compare commits
36
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9ba740e125 | ||
|
|
34346779d1 | ||
|
|
ad9cd63122 | ||
|
|
68e6ffca9e | ||
|
|
64cdcf6be4 | ||
|
|
aa0d98a6b8 | ||
|
|
93b03bc14d | ||
|
|
160f0c9ccf | ||
|
|
e5d1110a0f | ||
|
|
622217ff2a | ||
|
|
cbab605eff | ||
|
|
ac98869aa1 | ||
|
|
9713ea1275 | ||
|
|
37aa382cce | ||
|
|
3f00983287 | ||
|
|
2dc57f4070 | ||
|
|
9df19be719 | ||
|
|
089eea3970 | ||
|
|
dd8447ecc5 | ||
|
|
dca423fd31 | ||
|
|
1b43af8e8e | ||
|
|
628591b620 | ||
|
|
0980ca563f | ||
|
|
b158388733 | ||
|
|
ac56806aff | ||
|
|
aa95a4c18e | ||
|
|
942f7db3db | ||
|
|
aadb23f409 | ||
|
|
f56f567042 | ||
|
|
15a164a052 | ||
|
|
aaaa7a14a3 | ||
|
|
74b409d7cf | ||
|
|
528cef02c4 | ||
|
|
3c3da4d057 | ||
|
|
0399713e7b | ||
|
|
0af2e9e8ef |
@@ -1,99 +0,0 @@
|
||||
---
|
||||
name: ci-runner
|
||||
description: Work on FastVideo's Slurm-only, change-aware GPU CI lanes, static Buildkite graph, trusted ci-runner policy, lane scripts, and GB200 validation.
|
||||
---
|
||||
|
||||
# Slinky Slurm CI lanes
|
||||
|
||||
FastVideo's `ci-runner` Buildkite queue is the control plane for all active
|
||||
GPU CI. A private host-owned dispatcher leases GPUs from the Slinky Slurm tray
|
||||
and runs the immutable PR SHA inside an isolated Enroot container. Buildkite
|
||||
pipeline upload and Slurm submission occur on the login plane; every test
|
||||
payload executes on Slurm compute.
|
||||
|
||||
The files under `fastvideo/tests/modal/` and `.buildkite/scripts/pr_test.sh`
|
||||
are dormant rollback code. Never add an active Buildkite or slash-command
|
||||
route to them. `pr_test.sh` must continue to reject Buildkite invocations.
|
||||
|
||||
The private operator bundle is deliberately outside this repository because
|
||||
it contains site paths and credentials. See
|
||||
`docs/contributing/ci_architecture.md`; this skill covers the repository half
|
||||
and the coordination contract with that bundle.
|
||||
|
||||
## Invariants
|
||||
|
||||
- `.buildkite/pipeline.yml` contains exactly one static step for every active
|
||||
GPU lane. Each step pins a unique key and label, a 90-minute timeout, the
|
||||
trusted `/opt/fastvideo-ci-runner/run-ci` command (`run-unit` is the one
|
||||
compatibility wrapper), step-level internal `TEST_TYPE`, and
|
||||
`queue: "ci-runner"`.
|
||||
- Active CI contains no `pr_test.sh` command, Modal invocation, default queue,
|
||||
Buildkite plugin, `soft_fail`, or job-controlled artifact glob.
|
||||
- The six Fastcheck lanes use `:microscope:` labels. Full-Suite-only lanes use
|
||||
`:test_tube:` or `:bar_chart:` so direct reruns update the right aggregate.
|
||||
- SSIM and vanilla training request all four GPUs. Keep both in the
|
||||
`fastvideo/slinky/whole-tray` Buildkite concurrency group with a limit of one
|
||||
so the second job does not consume an agent or command timeout while waiting
|
||||
for the same tray.
|
||||
- `/test full` schedules all twenty lanes. `/merge`, `ready`, and new pushes to
|
||||
ready PRs use the trusted base-branch planner in
|
||||
`.github/scripts/plan_merge_ci.py`: automatic Fastcheck remains the universal
|
||||
six-lane baseline, and the merge build adds only path-relevant integration
|
||||
lanes. Unknown source/build paths fail closed to all fourteen additive lanes.
|
||||
The trusted uploader still normalizes and validates the complete static graph
|
||||
before Buildkite evaluates its plan conditions.
|
||||
- Focused merge builds may pass allowlisted golden-gate and SSIM test basenames.
|
||||
The private host validates the lane plan and basenames before staging them,
|
||||
and the in-container scripts validate them again. Direct `/test ssim`,
|
||||
explicit `/test full`, and the weekly main-branch schedule run the complete
|
||||
SSIM matrix.
|
||||
- The trusted uploader serves exactly three entry pipelines:
|
||||
`pr-fastcheck` for automatic PR builds, `ci` for slash-command/ready-label
|
||||
API builds, and `fastvideo-performance-lane` for the weekly schedule. Keep
|
||||
incoming GitHub webhook processing disabled on `ci` so it cannot duplicate
|
||||
`pr-fastcheck` on every PR update.
|
||||
- Test payloads live in `.buildkite/scripts/unit_test.sh` or executable
|
||||
`.buildkite/scripts/lanes/<lane>.sh`. Backend policy (GPU count, extras,
|
||||
secrets, kernel build, artifacts) stays in the agent-owned lane table.
|
||||
- Tests must preserve an inherited `MASTER_PORT`. Packed containers share the
|
||||
tray network namespace, so the private runner assigns a distinct port range
|
||||
per GPU lease and the SSIM scheduler assigns task offsets within its range.
|
||||
- The ARM64 runner image includes the pinned FA4 CuTe overlay validated on
|
||||
GB200. Keep SSIM at `FASTVIDEO_FA4=1` because its references were seeded with
|
||||
FA4; keep lanes with FA2 baselines at `FASTVIDEO_FA4=0`. A runner image change
|
||||
must revalidate both the FA4 import and an actual GB200 forward kernel.
|
||||
- `fastvideo/tests/ssim/ci_runner.py` is the active four-GPU SSIM scheduler.
|
||||
New SSIM files are discovered through `REQUIRED_GPUS` and
|
||||
`*_MODEL_TO_PARAMS`; do not wire them through the dormant Modal scheduler.
|
||||
- The host policy fail-closes unknown tuples. A repository-side lane change is
|
||||
inert until the operator updates the private lane table and uploader policy
|
||||
in the same rollout.
|
||||
|
||||
## Adding or changing a lane
|
||||
|
||||
1. Read the closest `AGENTS.md` and the domain-specific testing guide.
|
||||
2. Add or update the executable lane payload under `.buildkite/scripts/`.
|
||||
Keep it deterministic and free of host-specific paths or credential fetches.
|
||||
3. Add the static pipeline step and canonical `/test <name>` mapping. Keep the
|
||||
`<name>-ci` alias only when compatibility requires it.
|
||||
4. Add its source/test path ownership to `.github/scripts/plan_merge_ci.py`.
|
||||
Prefer the narrowest correctness-preserving lane set; leave unknown paths
|
||||
fail-closed. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
|
||||
`test_merge_ci_plan.py`, and focused CPU-only scheduler/policy tests.
|
||||
5. Coordinate the private lane row: GPU count (1-4), wall time, script, scope
|
||||
pairs, step key, command, HF cache/token, tracking mode, extras, attention
|
||||
backend policy, kernel policy, and artifact relay. Active training lanes
|
||||
keep W&B offline and do not stage a W&B credential.
|
||||
6. Update the trusted pipeline-uploader schema. A mismatch must reject the
|
||||
pipeline rather than silently skip a lane.
|
||||
7. Run `pre-commit run --files <changed paths>`, the planner's representative
|
||||
diff matrix, contract tests, private driver tests, and a real GB200 canary.
|
||||
Multi-GPU, hardware-reference, training, performance, and SSIM changes need
|
||||
their own target-hardware evidence.
|
||||
|
||||
## Rollback
|
||||
|
||||
Rollback the Slurm routing/configuration change or pause the `ci-runner` queue.
|
||||
Do not silently reactivate Modal. A manual Modal experiment requires the
|
||||
explicit local opt-in documented in `ci_architecture.md`; returning it to
|
||||
production CI needs a separate reviewed decision.
|
||||
@@ -1,79 +0,0 @@
|
||||
---
|
||||
name: env-var-conventions
|
||||
description: Add, read, rename, or remove an environment variable in FastVideo, or change the environment-variable policy. Use before touching fastvideo/envs.py, os.environ, os.getenv, or monkeypatch.setenv in fastvideo/, and when fastvideo/tests/contract/test_env_policy.py fails.
|
||||
---
|
||||
|
||||
# Environment Variable Conventions
|
||||
|
||||
## Purpose
|
||||
|
||||
FastVideo registers its environment variables as typed fields in
|
||||
`fastvideo/envs.py`. The policy that governs them is
|
||||
`docs/contributing/env_vars.md`, and the contract test
|
||||
`fastvideo/tests/contract/test_env_policy.py` enforces the policy in the unit
|
||||
CI lane. This skill routes an environment-variable change through that policy.
|
||||
The policy doc is the single source of the rules; read it instead of relying
|
||||
on a summary here.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Read `docs/contributing/env_vars.md` in full.
|
||||
- Decide whether the setting belongs in an environment variable or an argument
|
||||
(rule 5 in the policy doc). Settings that users change per deployment are
|
||||
arguments; add them through `fastvideo/fastvideo_args.py` instead.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
| ---------- | -------- | -------------------------------------------------------------- |
|
||||
| `change` | Yes | Add, read, rename, or remove a variable, or change the policy. |
|
||||
| `variable` | Yes | The variable name, with the `FASTVIDEO_` prefix. |
|
||||
|
||||
## Steps
|
||||
|
||||
1. **Declare or edit the variable in `fastvideo/envs.py`.**
|
||||
- Pick the field type and category that the policy doc lists.
|
||||
- Write a description that states what the variable does and its units.
|
||||
- To rename, keep the old name in `deprecated_names`. To remove, add the
|
||||
name to `DEPRECATED_VARIABLES`. Update the uses in `examples/`,
|
||||
`scripts/`, `docs/`, `apps/`, and the tests.
|
||||
2. **Read the variable with `envs.NAME.get()` inside a function.**
|
||||
- In tests, change the value with `envs.NAME.override(value)`, and a variable
|
||||
outside the registry with `envs.override_external(name, value)`; the
|
||||
`env_overrides` fixture keeps either until the end of the test.
|
||||
- Name a variable that only tests read `FASTVIDEO_TEST_*`.
|
||||
- Do not call `os.environ`, `os.getenv`, or `monkeypatch.setenv` for a
|
||||
FastVideo variable.
|
||||
- To set a variable that another tool reads, call `envs.set_external`,
|
||||
`envs.setdefault_external`, or `envs.unset_external`.
|
||||
3. **Regenerate the table in the policy doc.**
|
||||
- Run `python fastvideo/tests/contract/test_env_policy.py`.
|
||||
4. **Run the contract test.**
|
||||
- Run `pytest fastvideo/tests/contract/test_env_policy.py`.
|
||||
- When the test reports a fixed known violation, delete or lower its entry
|
||||
in `KNOWN_VIOLATIONS`. Never add an entry to `KNOWN_VIOLATIONS`.
|
||||
5. **When the policy itself changes, update the policy doc and the contract
|
||||
test in the same pull request.**
|
||||
- The rules in `docs/contributing/env_vars.md`, the checks and allowlist in
|
||||
`fastvideo/tests/contract/test_env_policy.py`, and this skill must agree.
|
||||
|
||||
## Outputs
|
||||
|
||||
- A registry entry in `fastvideo/envs.py` and call sites that use
|
||||
`envs.NAME.get()`.
|
||||
- A regenerated table in `docs/contributing/env_vars.md`.
|
||||
- A passing `fastvideo/tests/contract/test_env_policy.py`.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Add a FASTVIDEO_DEBUG_MY_STAGE switch that logs MyStage inputs.
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- `docs/contributing/env_vars.md`: the policy, the field types, and the
|
||||
violation kinds that the contract test reports.
|
||||
- `fastvideo/envs.py`: the registry.
|
||||
- `fastvideo/tests/contract/test_env_policy.py`: the contract test,
|
||||
`EXTERNAL_ALLOWLIST`, and `KNOWN_VIOLATIONS`.
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: reseed-ssim-references
|
||||
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted model subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
|
||||
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
|
||||
---
|
||||
|
||||
# Re-seed SSIM Reference Videos
|
||||
@@ -13,7 +13,7 @@ on HF — the old refs are overwritten — so the skill always:
|
||||
|
||||
1. Confirms intent with a one-liner the user has to type.
|
||||
2. Downloads the existing refs as a local, timestamped backup.
|
||||
3. Regenerates through the manual legacy Modal L40S maintenance path.
|
||||
3. Regenerates on Modal L40S (same code path that CI uses).
|
||||
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
|
||||
5. Uploads with `--force`, scoped to the single `--model-id`.
|
||||
6. Reminds the user to keep the backup until the PR lands.
|
||||
@@ -51,13 +51,12 @@ harder to recover from than failing closed.
|
||||
|
||||
Hardcoded:
|
||||
|
||||
- Modal GPU: **L40S**. This is a manual reference-maintenance target, not the
|
||||
active Slurm CI compute path; changing the SKU also changes the historical
|
||||
`L40S_reference_videos` contract.
|
||||
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
|
||||
that L40S CI cannot match).
|
||||
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
|
||||
operation.
|
||||
- HF repo: `FastVideo/ssim-reference-videos` (override via
|
||||
`FASTVIDEO_TEST_SSIM_REFERENCE_HF_REPO`).
|
||||
`FASTVIDEO_SSIM_REFERENCE_HF_REPO`).
|
||||
- Device folder: `L40S_reference_videos`.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
@@ -35,8 +35,7 @@ The skill is run **manually**, once per new test. Before invoking it, the user
|
||||
has already sanity-tested the new test locally — it launches `VideoGenerator`
|
||||
and writes an artefact without crashing (the missing-reference assertion at
|
||||
the end is expected). The skill does not re-test locally; it goes straight
|
||||
to the manual legacy Modal L40S reference-maintenance target. Active CI runs
|
||||
on the Slinky Slurm cluster and only consumes the resulting references.
|
||||
to Modal L40S (which is what CI uses).
|
||||
|
||||
## When to use
|
||||
|
||||
@@ -62,8 +61,7 @@ Prompt the user for it if they didn't supply it.
|
||||
|
||||
Everything else is fixed:
|
||||
|
||||
- Modal maintenance GPU: **L40S** (hardcoded in
|
||||
`fastvideo/tests/modal/ssim_test.py`; this is not the active CI compute path).
|
||||
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
|
||||
- Device folder: `L40S_reference_videos`.
|
||||
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
|
||||
seeded by this skill.
|
||||
|
||||
@@ -1,51 +0,0 @@
|
||||
{
|
||||
"benchmark_id": "wan-t2v-1.3b-1gpu-gb10",
|
||||
"config_schema_version": 2,
|
||||
"workload_id": "wan-t2v",
|
||||
"variant_id": "1.3b-sp1",
|
||||
"benchmark_version": 3,
|
||||
"description": "Wan2.1 T2V 1.3B single-GPU inference performance on NVIDIA DGX Spark (GB10). Single-GPU variant of wan-t2v-1.3b (same workload_id for dashboard comparability). Gated to the GB10 via run_config.gpu_types so it does not run on the shared H100/L40S lanes.",
|
||||
"model": {
|
||||
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
"model_short_name": "Wan2.1-T2V-1.3B"
|
||||
},
|
||||
"init_kwargs": {
|
||||
"num_gpus": 1,
|
||||
"flow_shift": 7.0,
|
||||
"sp_size": 1,
|
||||
"tp_size": 1,
|
||||
"vae_sp": false,
|
||||
"vae_tiling": true,
|
||||
"text_encoder_precisions": ["fp32"]
|
||||
},
|
||||
"generation_kwargs": {
|
||||
"height": 480,
|
||||
"width": 832,
|
||||
"num_frames": 45,
|
||||
"num_inference_steps": 4,
|
||||
"guidance_scale": 3,
|
||||
"embedded_cfg_scale": 6,
|
||||
"seed": 1024,
|
||||
"fps": 24,
|
||||
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
|
||||
},
|
||||
"test_prompts": [
|
||||
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
|
||||
],
|
||||
"run_config": {
|
||||
"num_warmup_runs": 2,
|
||||
"num_measurement_runs": 5,
|
||||
"required_gpus": 1,
|
||||
"gpu_types": ["GB10"]
|
||||
},
|
||||
"thresholds": {
|
||||
"GB10": {
|
||||
"max_generation_time_s": 55.0,
|
||||
"max_peak_memory_mb": 12000.0
|
||||
},
|
||||
"default": {
|
||||
"max_generation_time_s": 120.0,
|
||||
"max_peak_memory_mb": 40000.0
|
||||
}
|
||||
}
|
||||
}
|
||||
+528
-464
File diff suppressed because it is too large
Load Diff
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the OpenAI-compatible API lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/entrypoints/test_openai_api_integration.py -vs
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the distillation-DMD lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/training/distill/test_distill_dmd.py -vs
|
||||
@@ -1,87 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# DreamVerse needs a GPU for import-time device resolution, but it does not
|
||||
# build or exercise fastvideo-kernel. A checksummed Node archive is installed
|
||||
# in the disposable Slurm container because the shared CI image is
|
||||
# Python/CUDA focused.
|
||||
set -euo pipefail
|
||||
|
||||
node_version=v22.23.2
|
||||
case $(uname -m) in
|
||||
aarch64 | arm64)
|
||||
node_arch=arm64
|
||||
node_archive_sha256=013b59cfd2819703a6f4a14ab891fc46fc2a4e3f5bcd92de3fb4929b43e35b30
|
||||
;;
|
||||
x86_64 | amd64)
|
||||
node_arch=x64
|
||||
node_archive_sha256=b294a556e639d64338823920e5866c21c02741742d2e1529ee1a225c1ec9252a
|
||||
;;
|
||||
*)
|
||||
echo "Unsupported architecture for DreamVerse Node runtime: $(uname -m)" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
node_archive="node-${node_version}-linux-${node_arch}.tar.gz"
|
||||
node_runtime_root=$(mktemp -d -t fastvideo-node.XXXXXX)
|
||||
node_archive_path="${node_runtime_root}/${node_archive}"
|
||||
node_install_dir="${node_runtime_root}/${node_archive%.tar.gz}"
|
||||
curl --proto '=https' --tlsv1.2 --retry 5 --retry-all-errors \
|
||||
--location --fail --silent --show-error \
|
||||
"https://nodejs.org/dist/${node_version}/${node_archive}" \
|
||||
--output "$node_archive_path"
|
||||
printf '%s %s\n' "$node_archive_sha256" "$node_archive_path" | sha256sum --check --status
|
||||
tar -xzf "$node_archive_path" -C "$node_runtime_root"
|
||||
export PATH="${node_install_dir}/bin:${PATH}"
|
||||
node --version
|
||||
npm --version
|
||||
|
||||
export PYTHONPATH="$(pwd)/apps/dreamverse${PYTHONPATH:+:$PYTHONPATH}"
|
||||
pytest apps/dreamverse/dreamverse/tests -q
|
||||
|
||||
cd apps/dreamverse/web
|
||||
npm ci
|
||||
npm run typecheck
|
||||
npm test
|
||||
machine_arch=$(uname -m)
|
||||
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
|
||||
npx playwright install --with-deps chromium firefox
|
||||
else
|
||||
npx playwright install --with-deps chromium webkit firefox
|
||||
fi
|
||||
|
||||
master_port=${MASTER_PORT:-7959}
|
||||
BACKEND_PORT=${BACKEND_PORT:-$((master_port + 50))}
|
||||
python -m uvicorn dreamverse.mock_server:app --host 127.0.0.1 --port "$BACKEND_PORT" &
|
||||
mock_server_pid=$!
|
||||
cleanup() {
|
||||
kill "$mock_server_pid" 2>/dev/null || true
|
||||
wait "$mock_server_pid" 2>/dev/null || true
|
||||
}
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
for _ in {1..30}; do
|
||||
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz" && break
|
||||
sleep 1
|
||||
done
|
||||
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz"
|
||||
|
||||
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
|
||||
# Playwright WebKit traps before opening a page on Linux ARM64, and its
|
||||
# bundled Chromium lacks the H.264/AAC codecs used by the fMP4 assertions.
|
||||
# Firefox covers every flow, including streaming. Chromium and its mobile
|
||||
# profile still cover all codec-independent UI behavior on GB200.
|
||||
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
|
||||
npm run e2e -- --project=firefox
|
||||
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
|
||||
npm run e2e -- \
|
||||
--project=chromium \
|
||||
--project=mobile-chromium \
|
||||
--grep-invert='streams, plays, and surfaces a downloadable clip|starts a new project and switches back to the prior session|saved projects persist across a page reload'
|
||||
else
|
||||
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
|
||||
npm run e2e -- \
|
||||
--project=chromium \
|
||||
--project=webkit \
|
||||
--project=firefox \
|
||||
--project=mobile-safari \
|
||||
--project=mobile-chromium
|
||||
fi
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the encoder lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/encoders -vs
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the evaluation lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/eval -vs
|
||||
@@ -1,35 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the golden-gate lane. Environment (HF_HOME
|
||||
# and authentication) is the runner's responsibility.
|
||||
set -euo pipefail
|
||||
|
||||
golden_root=./fastvideo/tests/golden_gate
|
||||
selected=${FASTVIDEO_GOLDEN_TEST_FILES-}
|
||||
if [ -z "$selected" ]; then
|
||||
if [ "${TEST_SCOPE:-}" = merge ]; then
|
||||
echo "Missing FASTVIDEO_GOLDEN_TEST_FILES for merge scope" >&2
|
||||
exit 2
|
||||
fi
|
||||
selected=all
|
||||
fi
|
||||
if [ "$selected" = all ]; then
|
||||
exec pytest "$golden_root" -xvs
|
||||
fi
|
||||
|
||||
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
|
||||
echo "Invalid FASTVIDEO_GOLDEN_TEST_FILES selection" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
IFS=, read -r -a golden_files <<< "$selected"
|
||||
golden_paths=()
|
||||
for golden_file in "${golden_files[@]}"; do
|
||||
golden_path="$golden_root/$golden_file"
|
||||
[ -f "$golden_path" ] || {
|
||||
echo "Selected golden test does not exist: $golden_file" >&2
|
||||
exit 2
|
||||
}
|
||||
golden_paths+=("$golden_path")
|
||||
done
|
||||
|
||||
exec pytest "${golden_paths[@]}" -xvs
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the LoRA-inference lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/inference/lora/test_lora_inference_similarity.py -vs
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the VMoBA-inference lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec python fastvideo/tests/inference/vmoba/test_vmoba_inference.py
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the custom-kernel lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest fastvideo-kernel/tests/ -vs
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the LoRA-extraction lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/lora_extraction/ -vs
|
||||
@@ -1,64 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm performance lane. Reports are written outside the checkout
|
||||
# so the trusted host driver can upload them after untrusted code exits.
|
||||
set -uo pipefail
|
||||
|
||||
export PERFORMANCE_TRACKING_ROOT=/tmp/perf-tracking
|
||||
export PERF_REPORTS_DIR=/workspace/artifacts/performance
|
||||
mkdir -p "$PERF_REPORTS_DIR"
|
||||
|
||||
if [[ ${BUILDKITE_PULL_REQUEST:-false} =~ ^[1-9][0-9]*$ ]]; then
|
||||
export PERF_RUN_SOURCE=pr
|
||||
export PERF_UPLOAD_POLICY=pass
|
||||
elif [ "${BUILDKITE_BRANCH:-}" = main ] \
|
||||
&& { [ "${BUILDKITE_SOURCE:-}" = schedule ] || [ "${TEST_SCOPE:-}" = full ]; }; then
|
||||
export PERF_RUN_SOURCE=scheduled_main
|
||||
export PERF_UPLOAD_POLICY=always
|
||||
elif [ "${TEST_SCOPE:-}" = direct ]; then
|
||||
export PERF_RUN_SOURCE=unknown
|
||||
export PERF_UPLOAD_POLICY=pass
|
||||
else
|
||||
export PERF_RUN_SOURCE=unknown
|
||||
export PERF_UPLOAD_POLICY=never
|
||||
fi
|
||||
|
||||
nvidia-smi \
|
||||
--query-gpu=index,timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu \
|
||||
--format=csv -l 10 > "$PERF_REPORTS_DIR/gpu_telemetry.csv" 2>/dev/null &
|
||||
telemetry_pid=$!
|
||||
cleanup() {
|
||||
kill "$telemetry_pid" 2>/dev/null || true
|
||||
wait "$telemetry_pid" 2>/dev/null || true
|
||||
}
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
pytest ./fastvideo/tests/performance/test_inference_performance.py -vs
|
||||
pytest_rc=$?
|
||||
compare_rc=0
|
||||
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
|
||||
PERF_PYTEST_RC=$pytest_rc python ./fastvideo/tests/performance/compare_baseline.py
|
||||
compare_rc=$?
|
||||
fi
|
||||
python ./fastvideo/tests/performance/dashboard.py || true
|
||||
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
|
||||
# The trusted host relays only .md/.html/.json/.csv from PERF_REPORTS_DIR, so
|
||||
# mirror each captured worker log with an allowlisted extension.
|
||||
for worker_log in fastvideo/tests/performance/results/worker_logs/*.log; do
|
||||
[ -f "$worker_log" ] || continue
|
||||
base=$(basename "${worker_log%.log}")
|
||||
# WorkerLogCapture keeps a .log.1 backup after rollover, and read_log_tail
|
||||
# includes it; mirror that retained history too so the artifact is complete.
|
||||
if [ -f "$worker_log.1" ]; then
|
||||
cp -f "$worker_log.1" "$PERF_REPORTS_DIR/${base}.1.md" 2>/dev/null || true
|
||||
fi
|
||||
cp -f "$worker_log" "$PERF_REPORTS_DIR/${base}.md" 2>/dev/null || true
|
||||
done
|
||||
|
||||
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
|
||||
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
|
||||
|
||||
final_rc=$pytest_rc
|
||||
if [ "$final_rc" -eq 0 ]; then
|
||||
final_rc=$compare_rc
|
||||
fi
|
||||
exit "$final_rc"
|
||||
@@ -1,6 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the self-forcing lane.
|
||||
set -euo pipefail
|
||||
|
||||
export WANDB_MODE=offline
|
||||
exec pytest ./fastvideo/tests/training/self-forcing/test_self_forcing.py -vs
|
||||
@@ -1,40 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical four-GPU SSIM lane for the Slinky Slurm worker.
|
||||
set -euo pipefail
|
||||
|
||||
args=()
|
||||
if [ "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-0}" = 1 ]; then
|
||||
args+=(--bootstrap-mode)
|
||||
fi
|
||||
selected=${FASTVIDEO_SSIM_TEST_FILES-}
|
||||
if [ -z "$selected" ]; then
|
||||
if [ "${TEST_SCOPE:-}" = merge ]; then
|
||||
echo "Missing FASTVIDEO_SSIM_TEST_FILES for merge scope" >&2
|
||||
exit 2
|
||||
fi
|
||||
selected=all
|
||||
fi
|
||||
if [ "$selected" != all ]; then
|
||||
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
|
||||
echo "Invalid FASTVIDEO_SSIM_TEST_FILES selection" >&2
|
||||
exit 2
|
||||
}
|
||||
IFS=, read -r -a ssim_files <<< "$selected"
|
||||
for ssim_file in "${ssim_files[@]}"; do
|
||||
args+=(--test-file "$ssim_file")
|
||||
done
|
||||
fi
|
||||
|
||||
# MoGe's utils3d dependency builds glcontext from source on ARM64. The current
|
||||
# runner image predates the baked-in X11 headers below, so keep this guarded
|
||||
# bootstrap until every deployed image digest contains libx11-dev.
|
||||
if [ ! -f /usr/include/X11/Xlib.h ]; then
|
||||
apt-get -o Acquire::Retries=5 update
|
||||
apt-get -o Acquire::Retries=5 install -y --no-install-recommends libx11-dev
|
||||
rm -rf /var/lib/apt/lists/*
|
||||
fi
|
||||
|
||||
uv pip install git+https://github.com/microsoft/MoGe.git
|
||||
uv pip install k_diffusion einops_exts alias_free_torch torchsde
|
||||
|
||||
exec python fastvideo/tests/ssim/ci_runner.py "${args[@]}"
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the modular training-framework lane.
|
||||
set -euo pipefail
|
||||
|
||||
exec pytest ./fastvideo/tests/train/models ./fastvideo/tests/train/methods -vs
|
||||
@@ -1,6 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the legacy vanilla-training lane.
|
||||
set -euo pipefail
|
||||
|
||||
export WANDB_MODE=offline
|
||||
exec pytest ./fastvideo/tests/training/Vanilla -srP
|
||||
@@ -1,6 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the legacy LoRA-training lane.
|
||||
set -euo pipefail
|
||||
|
||||
export WANDB_MODE=offline
|
||||
exec pytest ./fastvideo/tests/training/lora/test_lora_training.py -srP
|
||||
@@ -1,6 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the legacy VSA-training lane.
|
||||
set -euo pipefail
|
||||
|
||||
export WANDB_MODE=offline
|
||||
exec pytest ./fastvideo/tests/training/VSA -srP
|
||||
@@ -1,9 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the transformer lane.
|
||||
set -euo pipefail
|
||||
|
||||
# The existing block reference records an absent FASTVIDEO_FA4 (FA2). Keep
|
||||
# that reference identity; the component lane also selects FA2 explicitly.
|
||||
env -u FASTVIDEO_FA4 pytest ./fastvideo/tests/golden_gate/test_wan_t2v.py -xvs
|
||||
pytest ./fastvideo/tests/golden_gate/test_wan_causal.py -xvs
|
||||
exec pytest ./fastvideo/tests/transformers -vs
|
||||
@@ -1,6 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Canonical Slurm CI selection for the VAE lane.
|
||||
set -euo pipefail
|
||||
|
||||
pytest ./fastvideo/tests/golden_gate/test_wan_vae.py -xvs
|
||||
exec pytest ./fastvideo/tests/vaes -vs
|
||||
@@ -1,19 +1,6 @@
|
||||
#!/bin/bash
|
||||
set -uo pipefail
|
||||
|
||||
# DORMANT ROLLBACK ONLY. Active CI is Slurm-only and pipeline.yml never calls
|
||||
# this launcher. Refuse every Buildkite invocation even if a stale step or
|
||||
# operator typo reaches this file; local rollback experiments require an
|
||||
# explicit opt-in.
|
||||
if [ -n "${BUILDKITE:-}" ]; then
|
||||
echo "Legacy Modal CI is disabled; use the Slinky Slurm runner." >&2
|
||||
exit 2
|
||||
fi
|
||||
if [ "${FASTVIDEO_ENABLE_LEGACY_MODAL_CI:-0}" != 1 ]; then
|
||||
echo "Legacy Modal CI is dormant. Set FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1 only for a manual rollback test." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
log() {
|
||||
echo "[$(date '+%Y-%m-%d %H:%M:%S')] $1"
|
||||
}
|
||||
|
||||
@@ -1,41 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
# Collect the whole attention directory so new files cannot land uncovered.
|
||||
# Its FA2/FA3 regression files skip when FA4 is selected (the Modal image
|
||||
# enables FA4 by default), so pin FA4 off for the directory to be real
|
||||
# coverage on every runner rather than a nominal collection.
|
||||
export FASTVIDEO_FA4=0
|
||||
|
||||
# The livestream app's tests are CPU-only; its single gpu-marked module is
|
||||
# deselected, and DreamVerse's GPU tests have their own lane.
|
||||
exec pytest \
|
||||
./apps/infinite_livestream/infinite_livestream/tests \
|
||||
./fastvideo/tests/api/ \
|
||||
./fastvideo/tests/contract/ \
|
||||
./fastvideo/tests/dataset/ \
|
||||
./fastvideo/tests/workflow/ \
|
||||
./fastvideo/tests/entrypoints/ \
|
||||
./fastvideo/tests/loader/ \
|
||||
./fastvideo/tests/pipelines/ \
|
||||
./fastvideo/tests/platforms/ \
|
||||
./fastvideo/tests/schedulers/ \
|
||||
./fastvideo/tests/train/ \
|
||||
./fastvideo/tests/stages/ \
|
||||
./fastvideo/tests/ops/ \
|
||||
./fastvideo/tests/worker/ \
|
||||
./fastvideo/tests/training/test_runner.py \
|
||||
./fastvideo/tests/training/test_trackers.py \
|
||||
./fastvideo/tests/inference/test_basic_fasth3_omniref_pdd.py \
|
||||
./fastvideo/tests/inference/test_inference_regional_compile.py \
|
||||
./fastvideo/tests/attention/ \
|
||||
./fastvideo/tests/layers/test_pdd_linear.py \
|
||||
./fastvideo/tests/layers/test_triton_fused_norm.py \
|
||||
./fastvideo/tests/modal/test_kernel_build_cache.py \
|
||||
./fastvideo/tests/modal/test_pr_test.py \
|
||||
./fastvideo/tests/modal/test_ssim_test.py \
|
||||
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
|
||||
--ignore=./fastvideo/tests/train/models \
|
||||
--ignore=./fastvideo/tests/train/methods \
|
||||
-m "not gpu" \
|
||||
-vs
|
||||
@@ -8,10 +8,10 @@ PR TITLE: Must start with a type tag, e.g.:
|
||||
MERGE WORKFLOW:
|
||||
1. Ensure pre-commit passes and you have at least 1 approval
|
||||
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
|
||||
3. A path-aware merge gate runs only relevant integration tests → auto-merge on success
|
||||
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
|
||||
|
||||
ON-DEMAND TESTING (write access required):
|
||||
/test full — Explicit all-lane run /test ssim — Full SSIM regression
|
||||
/test full — Full Test Suite /test ssim — SSIM regression
|
||||
/test training — Training pipeline /test encoder — Encoder tests
|
||||
/test transformer — Transformer tests /test vae — VAE tests
|
||||
/test kernel — CUDA kernel tests /test unit — Unit tests
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
#!/usr/bin/env bash
|
||||
# Gate the path-aware Buildkite merge plan on the cheap GitHub checks.
|
||||
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
|
||||
#
|
||||
# Polls the workflow runs for the PR head commit and only exits 0 once the
|
||||
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
|
||||
# 'ready' label cannot burn path-selected GPU lanes on a head that a cheap
|
||||
# check has already doomed.
|
||||
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
|
||||
# already doomed.
|
||||
#
|
||||
# Semantics:
|
||||
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
|
||||
# no merge gate; the next push re-arms via the 'synchronize' trigger)
|
||||
# no full suite; the next push re-arms via the 'synchronize' trigger)
|
||||
# - watched run cancelled -> still pending: the docs
|
||||
# workflow's repo-global 'pages' concurrency group cancels runs superseded
|
||||
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
|
||||
@@ -29,7 +29,7 @@ set -euo pipefail
|
||||
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
|
||||
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
|
||||
|
||||
# Workflow-level `name:` values that must be green before the merge gate
|
||||
# Workflow-level `name:` values that must be green before the full suite
|
||||
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
|
||||
# legitimately never exist; pre-commit always runs, so it must appear.
|
||||
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
|
||||
@@ -56,7 +56,7 @@ recheck_ready_label() {
|
||||
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
|
||||
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
|
||||
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
|
||||
"NOT triggering the Buildkite merge gate. Re-add the label to re-arm."
|
||||
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
@@ -84,7 +84,7 @@ while true; do
|
||||
| map(.name) | join(", ")' <<<"$state")
|
||||
if [ -n "$failed" ]; then
|
||||
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
|
||||
"NOT triggering the Buildkite merge gate. Push a fix (the 'ready'" \
|
||||
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
|
||||
"label re-arms on every push), or re-run the failed check and then" \
|
||||
"re-run this workflow."
|
||||
exit 1
|
||||
@@ -97,7 +97,7 @@ while true; do
|
||||
if [ "$pending" -eq 0 ]; then
|
||||
if [ -z "$missing" ]; then
|
||||
recheck_ready_label
|
||||
echo "All watched cheap checks are green — merge gate may proceed."
|
||||
echo "All watched cheap checks are green — full suite may proceed."
|
||||
exit 0
|
||||
fi
|
||||
case "$missing" in
|
||||
@@ -119,14 +119,14 @@ while true; do
|
||||
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
|
||||
if [ "$api_fails" -ge 3 ]; then
|
||||
recheck_ready_label
|
||||
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the merge gate WITHOUT the cheap-check gate."
|
||||
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
|
||||
recheck_ready_label
|
||||
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the merge gate anyway."
|
||||
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
|
||||
exit 0
|
||||
fi
|
||||
sleep "$POLL_SECS"
|
||||
|
||||
@@ -1,592 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Select the additive GPU integration lanes needed by a PR diff.
|
||||
|
||||
Fastcheck is the universal six-lane baseline and is intentionally not repeated
|
||||
here. This planner selects only the more expensive merge-gate lanes. Unknown
|
||||
source/build paths fail closed to the complete integration set, while explicit
|
||||
documentation and repository-metadata paths require no additional GPU work.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import fnmatch
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import TextIO
|
||||
|
||||
MERGE_LANES = (
|
||||
"golden-gate",
|
||||
"ssim",
|
||||
"lora-inference",
|
||||
"lora-extraction",
|
||||
"training",
|
||||
"distillation",
|
||||
"self-forcing",
|
||||
"lora-training",
|
||||
"training-vsa",
|
||||
"inference-vmoba",
|
||||
"performance",
|
||||
"api-server",
|
||||
"train-framework",
|
||||
"eval",
|
||||
)
|
||||
|
||||
LANE_SCRIPT_TO_KEY = {
|
||||
"api_server.sh": "api-server",
|
||||
"distillation_dmd.sh": "distillation",
|
||||
"eval.sh": "eval",
|
||||
"golden_gate.sh": "golden-gate",
|
||||
"inference_lora.sh": "lora-inference",
|
||||
"inference_vmoba.sh": "inference-vmoba",
|
||||
"lora_extraction.sh": "lora-extraction",
|
||||
"performance.sh": "performance",
|
||||
"self_forcing.sh": "self-forcing",
|
||||
"ssim.sh": "ssim",
|
||||
"train_framework.sh": "train-framework",
|
||||
"training.sh": "training",
|
||||
"training_lora.sh": "lora-training",
|
||||
"training_vsa.sh": "training-vsa",
|
||||
}
|
||||
|
||||
FASTCHECK_LANE_SCRIPTS = {
|
||||
"dreamverse.sh",
|
||||
"encoder.sh",
|
||||
"kernel_tests.sh",
|
||||
"transformer.sh",
|
||||
"vae.sh",
|
||||
}
|
||||
|
||||
LEGACY_TRAINING_LANES = (
|
||||
"training",
|
||||
"distillation",
|
||||
"self-forcing",
|
||||
"lora-training",
|
||||
"training-vsa",
|
||||
)
|
||||
|
||||
ALL_TRAINING_LANES = (*LEGACY_TRAINING_LANES, "train-framework")
|
||||
|
||||
SSIM_SMOKE_TESTS = (
|
||||
"test_flux_t2i_similarity.py",
|
||||
"test_wan_t2v_similarity.py",
|
||||
)
|
||||
|
||||
SAFE_PATTERNS = (
|
||||
"*.md",
|
||||
"*.rst",
|
||||
".agents/**",
|
||||
".claude/**",
|
||||
".codex/**",
|
||||
".github/ISSUE_TEMPLATE/**",
|
||||
".github/PULL_REQUEST_TEMPLATE.md",
|
||||
".github/dependabot.yml",
|
||||
".github/mergify.yml",
|
||||
".github/scripts/**",
|
||||
".github/workflows/**",
|
||||
".buildkite/scripts/pre_commit.sh",
|
||||
".git-blame-ignore-revs",
|
||||
".gitattributes",
|
||||
".gitignore",
|
||||
".pre-commit-config.yaml",
|
||||
"AGENTS.md",
|
||||
"CITATION.cff",
|
||||
"CODE_OF_CONDUCT.md",
|
||||
"CONTRIBUTING.md",
|
||||
"LICENSE",
|
||||
"NOTICE",
|
||||
"__init__.py",
|
||||
"collect_env.py",
|
||||
"SECURITY.md",
|
||||
"assets/**",
|
||||
"comfyui/**",
|
||||
"docs/**",
|
||||
"examples/**",
|
||||
"mkdocs.yml",
|
||||
"requirements-mkdocs.in",
|
||||
"requirements-mkdocs.txt",
|
||||
"scripts/**",
|
||||
"tests/__init__.py",
|
||||
"tests/local_tests/**",
|
||||
)
|
||||
|
||||
ALL_IMPACT_PATTERNS = (
|
||||
".buildkite/pipeline.yml",
|
||||
"docker/**",
|
||||
"pyproject.toml",
|
||||
"requirements*.txt",
|
||||
"setup.cfg",
|
||||
"setup.py",
|
||||
"uv.lock",
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class FamilyCoverage:
|
||||
pattern: re.Pattern[str]
|
||||
golden_tests: tuple[str, ...]
|
||||
ssim_tests: tuple[str, ...]
|
||||
|
||||
|
||||
FAMILY_COVERAGE = (
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])dreamx(_world)?([/_.-]|$)"),
|
||||
("test_dreamx.py", ),
|
||||
("test_dreamx_world_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])flux[_-]?2([/_.-]|$)"),
|
||||
("test_flux2_klein.py", ),
|
||||
("test_flux2_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])flux(?![_-]?2)([/_.-]|$)"),
|
||||
("test_flux.py", ),
|
||||
("test_flux_t2i_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])(hunyuan)?gamecraft([/_.-]|$)"),
|
||||
("test_gamecraft.py", ),
|
||||
("test_gamecraft_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])gen3c([/_.-]|$)"),
|
||||
("test_gen3c.py", ),
|
||||
("test_gen3c_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])glm[_-]?image([/_.-]|$)"),
|
||||
("test_glm_image.py", ),
|
||||
("test_glm_image_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])hunyuan(video)?15([a-z0-9_-]*)([/_.-]|$)"),
|
||||
(),
|
||||
("test_hunyuan15_i2v_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
|
||||
("test_kandinsky5.py", ),
|
||||
("test_kandinsky5_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])lingbot([a-z0-9_-]*)([/_.-]|$)"),
|
||||
("test_lingbot.py", ),
|
||||
("test_lingbot_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])longcat([/_.-]|$)"),
|
||||
("test_longcat.py", ),
|
||||
("test_longcat_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])ltx[_-]?2([/_.-]|$)"),
|
||||
("test_ltx2.py", ),
|
||||
("test_ltx2_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])matrixgame[_-]?2([/_.-]|$)"),
|
||||
("test_matrixgame.py", ),
|
||||
("test_matrixgame2_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])matrixgame[_-]?3([/_.-]|$)"),
|
||||
("test_matrixgame.py", ),
|
||||
("test_matrixgame3_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])minimax[_-]?h3([/_.-]|$)"),
|
||||
("test_minimax_h3_t2v.py", ),
|
||||
("test_minimax_h3_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])sd[_-]?3([._-]?5)?([/_.-]|$)"),
|
||||
("test_sd35.py", ),
|
||||
("test_sd35_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])stable[_-]?audio([/_.-]|$)"),
|
||||
("test_stable_audio.py", ),
|
||||
("test_stable_audio_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])turbo(diffusion)?([/_.-]|$)"),
|
||||
(),
|
||||
("test_turbodiffusion_similarity.py", ),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)"),
|
||||
("test_wan_t2v.py", "test_wan_vae.py", "test_wan_causal.py", "test_wan_denoising.py"),
|
||||
(
|
||||
"test_causal_similarity.py",
|
||||
"test_wan_i2v_similarity.py",
|
||||
"test_wan_t2v_similarity.py",
|
||||
"test_wan_ti2v_similarity.py",
|
||||
),
|
||||
),
|
||||
FamilyCoverage(
|
||||
re.compile(r"(^|[/_.-])z[_-]?image([/_.-]|$)"),
|
||||
("test_zimage.py", ),
|
||||
("test_zimage_similarity.py", ),
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
@dataclass
|
||||
class MergePlan:
|
||||
lanes: set[str] = field(default_factory=set)
|
||||
golden_tests: set[str] = field(default_factory=set)
|
||||
ssim_tests: set[str] = field(default_factory=set)
|
||||
golden_all: bool = False
|
||||
ssim_all: bool = False
|
||||
reasons: list[str] = field(default_factory=list)
|
||||
|
||||
def add_lanes(self, *lanes: str, reason: str) -> None:
|
||||
unknown = set(lanes) - set(MERGE_LANES)
|
||||
if unknown:
|
||||
raise ValueError(f"Unknown merge lanes: {sorted(unknown)}")
|
||||
self.lanes.update(lanes)
|
||||
self.reasons.append(reason)
|
||||
|
||||
def add_golden(self, tests: tuple[str, ...], reason: str) -> None:
|
||||
self.add_lanes("golden-gate", reason=reason)
|
||||
self.golden_tests.update(tests)
|
||||
|
||||
def add_ssim(self, tests: tuple[str, ...], reason: str) -> None:
|
||||
self.add_lanes("ssim", reason=reason)
|
||||
self.ssim_tests.update(tests)
|
||||
|
||||
def require_all(self, reason: str) -> None:
|
||||
self.lanes.update(MERGE_LANES)
|
||||
self.golden_all = True
|
||||
self.ssim_all = True
|
||||
self.reasons.append(reason)
|
||||
|
||||
def ordered_lanes(self) -> tuple[str, ...]:
|
||||
return tuple(lane for lane in MERGE_LANES if lane in self.lanes)
|
||||
|
||||
def encoded_lanes(self) -> str:
|
||||
lanes = self.ordered_lanes()
|
||||
return "," + ",".join(lanes or ("none", )) + ","
|
||||
|
||||
def encoded_golden_tests(self) -> str:
|
||||
if "golden-gate" not in self.lanes:
|
||||
return "none"
|
||||
if self.golden_all or not self.golden_tests:
|
||||
return "all"
|
||||
return ",".join(sorted(self.golden_tests))
|
||||
|
||||
def encoded_ssim_tests(self) -> str:
|
||||
if "ssim" not in self.lanes:
|
||||
return "none"
|
||||
if self.ssim_all or not self.ssim_tests:
|
||||
return "all"
|
||||
return ",".join(sorted(self.ssim_tests))
|
||||
|
||||
|
||||
def _matches_any(path: str, patterns: tuple[str, ...]) -> bool:
|
||||
return any(fnmatch.fnmatchcase(path, pattern) for pattern in patterns)
|
||||
|
||||
|
||||
def _family_coverage(path: str) -> tuple[set[str], set[str]]:
|
||||
normalized = path.lower()
|
||||
golden: set[str] = set()
|
||||
ssim: set[str] = set()
|
||||
for family in FAMILY_COVERAGE:
|
||||
if family.pattern.search(normalized):
|
||||
golden.update(family.golden_tests)
|
||||
ssim.update(family.ssim_tests)
|
||||
# Select the component actually touched, including compatibility paths.
|
||||
# Family configs/pipeline wiring can affect all four Wan gates.
|
||||
if re.search(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)", normalized):
|
||||
if (normalized.endswith(("/wan/vae.py", "/wan/vae_config.py", "/vaes/wanvae.py"))
|
||||
or normalized.endswith("/wan/stages/conditioning.py")):
|
||||
golden = {"test_wan_vae.py"}
|
||||
elif normalized.endswith(("/wan/causal_transformer.py", "/dits/causal_wanvideo.py",
|
||||
"/wan/stages/causal_denoising.py")):
|
||||
golden = {"test_wan_causal.py"}
|
||||
elif (normalized == "fastvideo/models/dits/wanvideo.py"
|
||||
or normalized.endswith(("/wan/transformer.py", "/wan/stages/denoising.py", "/wan/stages/dmd.py"))):
|
||||
golden = {"test_wan_t2v.py", "test_wan_denoising.py"}
|
||||
return golden, ssim
|
||||
|
||||
|
||||
def _select_output_coverage(plan: MergePlan, path: str) -> None:
|
||||
golden, ssim = _family_coverage(path)
|
||||
if golden:
|
||||
plan.add_golden(tuple(sorted(golden)), reason=f"model-family golden coverage: {path}")
|
||||
else:
|
||||
plan.golden_all = True
|
||||
plan.add_lanes("golden-gate", reason=f"shared output golden coverage: {path}")
|
||||
if ssim:
|
||||
plan.add_ssim(tuple(sorted(ssim)), reason=f"model-family SSIM coverage: {path}")
|
||||
else:
|
||||
plan.add_ssim(SSIM_SMOKE_TESTS, reason=f"shared output SSIM smoke coverage: {path}")
|
||||
|
||||
|
||||
def classify_paths(paths: list[str]) -> MergePlan:
|
||||
plan = MergePlan()
|
||||
normalized_paths: list[str] = []
|
||||
for raw_path in paths:
|
||||
path = raw_path.strip()
|
||||
while path.startswith("./"):
|
||||
path = path[2:]
|
||||
if path:
|
||||
normalized_paths.append(path)
|
||||
normalized_paths = sorted(set(normalized_paths))
|
||||
if not normalized_paths:
|
||||
plan.require_all("changed-file list was empty; failing closed")
|
||||
return plan
|
||||
|
||||
for path in normalized_paths:
|
||||
if path == "__FASTVIDEO_CI_PLAN_ALL__":
|
||||
plan.require_all("changed-file API failed; failing closed")
|
||||
continue
|
||||
|
||||
if path in {"requirements-mkdocs.in", "requirements-mkdocs.txt"}:
|
||||
plan.reasons.append(f"documentation dependencies need no GPU integration: {path}")
|
||||
continue
|
||||
|
||||
if _matches_any(path, ALL_IMPACT_PATTERNS):
|
||||
plan.require_all(f"cross-cutting build/runtime surface: {path}")
|
||||
continue
|
||||
|
||||
lane_script_prefix = ".buildkite/scripts/lanes/"
|
||||
if path.startswith(lane_script_prefix):
|
||||
script_name = Path(path).name
|
||||
lane = LANE_SCRIPT_TO_KEY.get(script_name)
|
||||
if lane is None:
|
||||
if script_name in FASTCHECK_LANE_SCRIPTS:
|
||||
plan.reasons.append(f"covered by automatic Fastcheck lane: {path}")
|
||||
else:
|
||||
plan.require_all(f"unknown lane script: {path}")
|
||||
elif lane == "golden-gate":
|
||||
plan.golden_all = True
|
||||
plan.add_lanes(lane, reason=f"golden lane implementation: {path}")
|
||||
elif lane == "ssim":
|
||||
plan.ssim_all = True
|
||||
plan.add_lanes(lane, reason=f"SSIM lane implementation: {path}")
|
||||
else:
|
||||
plan.add_lanes(lane, reason=f"lane implementation: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/tests/golden_gate/"):
|
||||
name = Path(path).name
|
||||
if name.startswith("test_") and name.endswith(".py"):
|
||||
plan.add_golden((name, ), reason=f"changed golden test: {path}")
|
||||
elif name in {"AGENTS.md", "README.md"}:
|
||||
plan.reasons.append(f"golden documentation only: {path}")
|
||||
else:
|
||||
plan.golden_all = True
|
||||
plan.add_lanes("golden-gate", reason=f"shared golden harness/reference: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/tests/ssim/"):
|
||||
name = Path(path).name
|
||||
if name.startswith("test_") and name.endswith(".py"):
|
||||
plan.add_ssim((name, ), reason=f"changed SSIM test: {path}")
|
||||
elif path.endswith((".py", ".json", ".pt", ".png", ".mp4")):
|
||||
plan.ssim_all = True
|
||||
plan.add_lanes("ssim", reason=f"shared SSIM harness/reference: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/tests/performance/") or path.startswith(".buildkite/performance-benchmarks/"):
|
||||
plan.add_lanes("performance", reason=f"performance coverage: {path}")
|
||||
continue
|
||||
if path.startswith(("fastvideo/performance/", "fastvideo/performance_dashboard/",
|
||||
"apps/performance_dashboard/")):
|
||||
plan.add_lanes("performance", reason=f"performance implementation: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/benchmarks/"):
|
||||
if "/mlx_" in path or Path(path).name.startswith("mlx_"):
|
||||
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
|
||||
else:
|
||||
plan.add_lanes("performance", reason=f"benchmark implementation: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/eval/") or path.startswith("fastvideo/eval/"):
|
||||
plan.add_lanes("eval", reason=f"evaluation coverage: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/third_party/eval/"):
|
||||
plan.add_lanes("eval", reason=f"vendored evaluation implementation: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/lora_extraction/") or path.startswith("scripts/lora_extraction/"):
|
||||
plan.add_lanes("lora-extraction", reason=f"LoRA extraction coverage: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/inference/lora/"):
|
||||
plan.add_lanes("lora-inference", reason=f"LoRA inference coverage: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/inference/vmoba/"):
|
||||
plan.add_lanes("inference-vmoba", reason=f"VMoBA inference coverage: {path}")
|
||||
continue
|
||||
if path.startswith(("fastvideo/dataset/", "fastvideo/workflow/", "fastvideo/pipelines/preprocess/",
|
||||
"fastvideo/pipelines/training/")):
|
||||
plan.add_lanes(*ALL_TRAINING_LANES, reason=f"shared data/training input surface: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/train/") or path.startswith("fastvideo/train/"):
|
||||
plan.add_lanes("train-framework", reason=f"modular training coverage: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/tests/training/"):
|
||||
lowered = path.lower()
|
||||
if "/vanilla/" in lowered:
|
||||
plan.add_lanes("training", reason=f"vanilla training coverage: {path}")
|
||||
elif "/distill/" in lowered:
|
||||
plan.add_lanes("distillation", reason=f"distillation coverage: {path}")
|
||||
elif "/self-forcing/" in lowered:
|
||||
plan.add_lanes("self-forcing", reason=f"self-forcing coverage: {path}")
|
||||
elif "/lora/" in lowered:
|
||||
plan.add_lanes("lora-training", reason=f"LoRA training coverage: {path}")
|
||||
elif "/vsa/" in lowered:
|
||||
plan.add_lanes("training-vsa", reason=f"VSA training coverage: {path}")
|
||||
else:
|
||||
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training coverage: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/training/"):
|
||||
lowered = path.lower()
|
||||
if "self_forcing" in lowered:
|
||||
plan.add_lanes("self-forcing", reason=f"self-forcing implementation: {path}")
|
||||
elif "distill" in lowered:
|
||||
plan.add_lanes("distillation", reason=f"distillation implementation: {path}")
|
||||
elif "lora" in lowered:
|
||||
plan.add_lanes("lora-training", reason=f"LoRA training implementation: {path}")
|
||||
else:
|
||||
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training implementation: {path}")
|
||||
continue
|
||||
|
||||
lowered = path.lower()
|
||||
if "vmoba" in lowered and path.startswith(("fastvideo/", ".buildkite/")):
|
||||
plan.add_lanes("inference-vmoba", reason=f"VMoBA implementation: {path}")
|
||||
plan.add_golden(("test_wan_t2v.py", ), reason=f"VMoBA end-to-end coverage: {path}")
|
||||
continue
|
||||
if "lora" in lowered and path.startswith("fastvideo/"):
|
||||
plan.add_lanes(
|
||||
"lora-inference",
|
||||
"lora-extraction",
|
||||
"lora-training",
|
||||
reason=f"shared LoRA implementation: {path}",
|
||||
)
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
|
||||
if path.startswith("fastvideo/entrypoints/") or path.startswith("fastvideo/api/"):
|
||||
plan.add_lanes("api-server", reason=f"API/entrypoint integration: {path}")
|
||||
if "openai" not in lowered and "/cli/" not in lowered:
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path.startswith("fastvideo/worker/"):
|
||||
plan.add_lanes("api-server", reason=f"worker/API integration: {path}")
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path.startswith("fastvideo/distributed/"):
|
||||
plan.add_lanes(
|
||||
"training",
|
||||
"train-framework",
|
||||
reason=f"distributed runtime integration: {path}",
|
||||
)
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path.startswith(("fastvideo/hooks/", "fastvideo/platforms/", "fastvideo/third_party/")):
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path.startswith(("fastvideo/models/", "fastvideo/pipelines/", "fastvideo/configs/",
|
||||
"fastvideo/layers/", "fastvideo/attention/")):
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path in {
|
||||
"fastvideo/fastvideo_args.py",
|
||||
"fastvideo/forward_context.py",
|
||||
"fastvideo/image_processor.py",
|
||||
"fastvideo/registry.py",
|
||||
"fastvideo/utils.py",
|
||||
}:
|
||||
_select_output_coverage(plan, path)
|
||||
continue
|
||||
if path.startswith("fastvideo/mlx_runtime/"):
|
||||
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/logging_utils/") or path in {
|
||||
"fastvideo/__init__.py",
|
||||
"fastvideo/envs.py",
|
||||
"fastvideo/logger.py",
|
||||
"fastvideo/profiler.py",
|
||||
"fastvideo/version.py",
|
||||
}:
|
||||
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
|
||||
continue
|
||||
if path.startswith(("fastvideo-kernel/", "csrc/")):
|
||||
plan.add_golden(("test_wan_t2v.py", ), reason=f"kernel integration smoke: {path}")
|
||||
plan.add_ssim(("test_wan_t2v_similarity.py", ), reason=f"kernel numerical smoke: {path}")
|
||||
continue
|
||||
|
||||
if path.startswith("apps/dreamverse/"):
|
||||
# DreamVerse is already one of the six automatic Fastcheck lanes.
|
||||
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
|
||||
continue
|
||||
if path.startswith("apps/infinite_livestream/"):
|
||||
# The app's CPU-only tests run in the automatic unit Fastcheck lane.
|
||||
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
|
||||
continue
|
||||
if path.startswith("fastvideo/tests/"):
|
||||
# The automatic unit/component Fastcheck lanes own the remaining
|
||||
# package tests. Domain-specific expensive test roots were handled
|
||||
# above.
|
||||
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
|
||||
continue
|
||||
if path in {".buildkite/scripts/unit_test.sh", ".buildkite/scripts/pr_test.sh"}:
|
||||
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
|
||||
continue
|
||||
if _matches_any(path, SAFE_PATTERNS):
|
||||
plan.reasons.append(f"no additional GPU integration needed: {path}")
|
||||
continue
|
||||
|
||||
plan.require_all(f"unclassified path; failing closed: {path}")
|
||||
|
||||
return plan
|
||||
|
||||
|
||||
def _write_github_output(output: TextIO, plan: MergePlan) -> None:
|
||||
output.write(f"merge_test_plan={plan.encoded_lanes()}\n")
|
||||
output.write(f"merge_golden_tests={plan.encoded_golden_tests()}\n")
|
||||
output.write(f"merge_ssim_tests={plan.encoded_ssim_tests()}\n")
|
||||
output.write(f"merge_plan_label={','.join(plan.ordered_lanes()) or 'none'}\n")
|
||||
|
||||
|
||||
def _write_summary(output: TextIO, plan: MergePlan) -> None:
|
||||
output.write("## Change-aware merge test plan\n\n")
|
||||
output.write("| Selection | Value |\n|---|---|\n")
|
||||
output.write(f"| Additional Slurm lanes | `{','.join(plan.ordered_lanes()) or 'none'}` |\n")
|
||||
output.write(f"| Golden tests | `{plan.encoded_golden_tests()}` |\n")
|
||||
output.write(f"| SSIM tests | `{plan.encoded_ssim_tests()}` |\n\n")
|
||||
output.write("Fastcheck remains the universal six-lane baseline.\n")
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--paths-file", type=Path, required=True)
|
||||
parser.add_argument("--github-output", type=Path)
|
||||
parser.add_argument("--summary-file", type=Path)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
paths = args.paths_file.read_text(encoding="utf-8").splitlines()
|
||||
plan = classify_paths(paths)
|
||||
print(f"MERGE_TEST_PLAN={plan.encoded_lanes()}")
|
||||
print(f"MERGE_GOLDEN_TESTS={plan.encoded_golden_tests()}")
|
||||
print(f"MERGE_SSIM_TESTS={plan.encoded_ssim_tests()}")
|
||||
for reason in plan.reasons:
|
||||
print(f"- {reason}")
|
||||
if args.github_output:
|
||||
with args.github_output.open("a", encoding="utf-8") as output:
|
||||
_write_github_output(output, plan)
|
||||
if args.summary_file:
|
||||
with args.summary_file.open("a", encoding="utf-8") as output:
|
||||
_write_summary(output, plan)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -53,7 +53,7 @@ PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion
|
||||
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
|
||||
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
|
||||
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
|
||||
OTHER='{"name": "Trigger Merge Gate", "id": 3, "status": "in_progress", "conclusion": null}'
|
||||
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
|
||||
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
|
||||
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
|
||||
|
||||
|
||||
@@ -190,7 +190,6 @@ jobs:
|
||||
if: ${{ !inputs.push_by_digest }}
|
||||
run: |
|
||||
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
|
||||
echo "Digest: ${{ steps.build-push.outputs.digest }}"
|
||||
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
|
||||
|
||||
- name: Digest success message
|
||||
|
||||
@@ -26,54 +26,29 @@ jobs:
|
||||
per_page: 100,
|
||||
});
|
||||
|
||||
// Buildkite derives the GitHub context prefix from the label emoji.
|
||||
// Keep hard Full Suite lanes in test-tube/bar-chart namespaces and
|
||||
// Fastcheck lanes in microscope so targeted reruns cannot clear the
|
||||
// wrong aggregate status. Automatic PR jobs use pr-fastcheck while
|
||||
// slash-command and Full Suite jobs use ci; normalize the suffix
|
||||
// and keep the newest status for each logical lane.
|
||||
const FASTCHECK_PREFIXES = [
|
||||
'buildkite/pr-fastcheck/microscope-',
|
||||
'buildkite/ci/microscope-',
|
||||
];
|
||||
const bkStatuses = data.statuses.filter(
|
||||
s => s.context.startsWith('buildkite/ci/')
|
||||
);
|
||||
|
||||
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
|
||||
const FULL_SUITE_PREFIXES = [
|
||||
'buildkite/ci/test-tube-',
|
||||
'buildkite/ci/bar-chart-',
|
||||
];
|
||||
|
||||
function newestByLane(prefixes) {
|
||||
const statuses = new Map();
|
||||
for (const status of data.statuses) {
|
||||
const prefix = prefixes.find(p => status.context.startsWith(p));
|
||||
if (!prefix) continue;
|
||||
const lane = status.context.slice(prefix.length);
|
||||
const previous = statuses.get(lane);
|
||||
if (!previous || Date.parse(status.updated_at) > Date.parse(previous.updated_at)) {
|
||||
statuses.set(lane, status);
|
||||
}
|
||||
}
|
||||
return statuses;
|
||||
}
|
||||
|
||||
const fastcheck = newestByLane(FASTCHECK_PREFIXES);
|
||||
const fullSuiteOnly = newestByLane(FULL_SUITE_PREFIXES);
|
||||
const fastcheckPassed =
|
||||
fastcheck.size === 6
|
||||
&& [...fastcheck.values()].every(s => s.state === 'success');
|
||||
const fullSuitePassed =
|
||||
fastcheckPassed
|
||||
&& fullSuiteOnly.size === 14
|
||||
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
|
||||
|
||||
// Direct reruns may repair a failed suite, never create a gate for
|
||||
// a suite that did not run.
|
||||
const failedAggregate = context => data.statuses.some(
|
||||
s => s.context === context && s.state === 'failure'
|
||||
const fastcheck = bkStatuses.filter(
|
||||
s => s.context.startsWith(FASTCHECK_PREFIX)
|
||||
);
|
||||
const fullSuite = bkStatuses.filter(
|
||||
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
|
||||
);
|
||||
|
||||
if (failedAggregate('fastcheck-passed') && fastcheckPassed) {
|
||||
if (
|
||||
fastcheck.length > 0
|
||||
&& fastcheck.every(s => s.state === 'success')
|
||||
) {
|
||||
core.info(
|
||||
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
|
||||
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
|
||||
);
|
||||
await github.rest.repos.createCommitStatus({
|
||||
owner: context.repo.owner,
|
||||
@@ -81,13 +56,17 @@ jobs:
|
||||
sha,
|
||||
state: 'success',
|
||||
context: 'fastcheck-passed',
|
||||
description: `All ${fastcheck.size} fastcheck tests passed`,
|
||||
description:
|
||||
`All ${fastcheck.length} fastcheck tests passed`,
|
||||
});
|
||||
}
|
||||
|
||||
if (failedAggregate('full-suite-passed') && fullSuitePassed) {
|
||||
if (
|
||||
fullSuite.length > 0
|
||||
&& fullSuite.every(s => s.state === 'success')
|
||||
) {
|
||||
core.info(
|
||||
'All 20 full suite tests passed — updating full-suite-passed'
|
||||
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
|
||||
);
|
||||
await github.rest.repos.createCommitStatus({
|
||||
owner: context.repo.owner,
|
||||
@@ -95,6 +74,7 @@ jobs:
|
||||
sha,
|
||||
state: 'success',
|
||||
context: 'full-suite-passed',
|
||||
description: 'All 20 full suite tests passed',
|
||||
description:
|
||||
`All ${fullSuite.length} full suite tests passed`,
|
||||
});
|
||||
}
|
||||
|
||||
@@ -8,8 +8,6 @@ on:
|
||||
- "fastvideo/mlx_runtime/**"
|
||||
- "fastvideo/tests/mlx/**"
|
||||
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
|
||||
- "fastvideo/tests/platforms/test_cpu_sdpa.py"
|
||||
- "fastvideo/platforms/cpu.py"
|
||||
- "fastvideo/platforms/mps.py"
|
||||
- "fastvideo/platforms/__init__.py"
|
||||
- "fastvideo/__init__.py"
|
||||
@@ -33,15 +31,15 @@ jobs:
|
||||
env:
|
||||
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
|
||||
TOKENIZERS_PARALLELISM: "false"
|
||||
MASTER_ADDR: "127.0.0.1"
|
||||
MASTER_ADDR: localhost
|
||||
MASTER_PORT: "29513"
|
||||
GLOO_SOCKET_IFNAME: lo0
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: "3.12"
|
||||
cache: pip
|
||||
|
||||
- uses: astral-sh/setup-uv@v3
|
||||
|
||||
@@ -49,9 +47,9 @@ jobs:
|
||||
run: |
|
||||
uv pip install --system \
|
||||
--index-url https://download.pytorch.org/whl/cpu \
|
||||
torch==2.12.0 torchvision torchaudio
|
||||
torch==2.11.0 torchvision torchaudio
|
||||
uv pip install --system \
|
||||
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
|
||||
pytest numpy scipy pillow imageio einops cloudpickle filelock \
|
||||
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru mlx \
|
||||
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
|
||||
|
||||
@@ -65,8 +63,8 @@ jobs:
|
||||
print("machine:", platform.machine())
|
||||
print("processor:", platform.processor())
|
||||
print("mlx default device:", mx.default_device())
|
||||
device_info = mx.metal.device_info() if mx.metal.is_available() else "metal unavailable"
|
||||
print("mlx device_info:", device_info)
|
||||
memory_size = mx.metal.device_info().get("memory_size") if mx.metal.is_available() else "metal unavailable"
|
||||
print("mlx memory_size:", memory_size)
|
||||
print("torch:", torch.__version__)
|
||||
print("torch mps available:", torch.backends.mps.is_available())
|
||||
PY
|
||||
@@ -74,26 +72,17 @@ jobs:
|
||||
- name: Run MLX smoke tests
|
||||
run: |
|
||||
python -m pytest \
|
||||
fastvideo/mlx_runtime/tests/ \
|
||||
fastvideo/tests/mlx/test_dmd_sampling.py \
|
||||
fastvideo/tests/mlx/test_memory_limits.py \
|
||||
fastvideo/tests/mlx/test_quant_capability.py \
|
||||
fastvideo/tests/mlx/test_mlx_dit_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_compile_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_checkpoint.py \
|
||||
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
|
||||
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
|
||||
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
|
||||
fastvideo/tests/mlx/test_taehv_decode.py \
|
||||
fastvideo/tests/mlx/test_frame_upsample.py \
|
||||
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
|
||||
fastvideo/tests/mlx/test_mlx_refine.py \
|
||||
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
|
||||
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
|
||||
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
|
||||
fastvideo/tests/mlx/test_wan22_sample.py \
|
||||
@@ -101,8 +90,7 @@ jobs:
|
||||
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
|
||||
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
|
||||
fastvideo/tests/platforms/test_mps_vsa_error.py \
|
||||
fastvideo/tests/platforms/test_cpu_sdpa.py \
|
||||
-v -s --timeout=120 -o faulthandler_timeout=120
|
||||
-q
|
||||
|
||||
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
|
||||
# slower to schedule; this Linux job gives fast PR signal on the identical
|
||||
@@ -123,6 +111,7 @@ jobs:
|
||||
- uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: "3.12"
|
||||
cache: pip
|
||||
|
||||
- uses: astral-sh/setup-uv@v3
|
||||
|
||||
@@ -130,35 +119,26 @@ jobs:
|
||||
run: |
|
||||
uv pip install --system \
|
||||
--index-url https://download.pytorch.org/whl/cpu \
|
||||
torch==2.12.0 torchvision torchaudio
|
||||
torch==2.11.0 torchvision torchaudio
|
||||
uv pip install --system \
|
||||
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
|
||||
pytest numpy scipy pillow imageio einops cloudpickle filelock \
|
||||
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru "mlx[cpu]" \
|
||||
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
|
||||
|
||||
- name: Run MLX smoke tests (CPU backend)
|
||||
run: |
|
||||
python -m pytest \
|
||||
fastvideo/mlx_runtime/tests/ \
|
||||
fastvideo/tests/mlx/test_dmd_sampling.py \
|
||||
fastvideo/tests/mlx/test_memory_limits.py \
|
||||
fastvideo/tests/mlx/test_quant_capability.py \
|
||||
fastvideo/tests/mlx/test_mlx_dit_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_compile_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_checkpoint.py \
|
||||
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
|
||||
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
|
||||
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
|
||||
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
|
||||
fastvideo/tests/mlx/test_taehv_decode.py \
|
||||
fastvideo/tests/mlx/test_frame_upsample.py \
|
||||
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
|
||||
fastvideo/tests/mlx/test_mlx_refine.py \
|
||||
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
|
||||
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
|
||||
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
|
||||
fastvideo/tests/mlx/test_wan22_sample.py \
|
||||
@@ -166,5 +146,4 @@ jobs:
|
||||
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
|
||||
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
|
||||
fastvideo/tests/platforms/test_mps_vsa_error.py \
|
||||
fastvideo/tests/platforms/test_cpu_sdpa.py \
|
||||
-v -s --timeout=120 -o faulthandler_timeout=120
|
||||
-q
|
||||
|
||||
@@ -1,44 +0,0 @@
|
||||
name: Scheduled Full SSIM
|
||||
|
||||
on:
|
||||
schedule:
|
||||
- cron: "0 5 * * 0"
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
trigger:
|
||||
if: github.repository == 'hao-ai-lab/FastVideo'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Trigger weekly full SSIM on Slinky Slurm
|
||||
env:
|
||||
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
|
||||
SOURCE_SHA: ${{ github.sha }}
|
||||
SOURCE_BRANCH: ${{ github.event.repository.default_branch }}
|
||||
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
|
||||
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
curl -sS --fail-with-body -X POST \
|
||||
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
|
||||
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
-H "Content-Type: application/json" \
|
||||
--data-raw "$(jq -n \
|
||||
--arg commit "$SOURCE_SHA" \
|
||||
--arg branch "$SOURCE_BRANCH" \
|
||||
'{
|
||||
commit: $commit,
|
||||
branch: $branch,
|
||||
message: "Weekly full SSIM on Slinky Slurm",
|
||||
ignore_pipeline_branch_filters: true,
|
||||
env: {
|
||||
TEST_SCOPE: "scheduled",
|
||||
FULL_SUITE: "false",
|
||||
TEST_TYPE: "ssim",
|
||||
PR_NUMBER: "false",
|
||||
PR_TITLE: "Scheduled full SSIM"
|
||||
}
|
||||
}')"
|
||||
@@ -32,7 +32,8 @@ jobs:
|
||||
}
|
||||
core.setOutput('has_write', String(hasWrite));
|
||||
|
||||
- name: Add ready label
|
||||
- name: Add ready label and react
|
||||
id: label
|
||||
if: steps.perm.outputs.has_write == 'true'
|
||||
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
|
||||
with:
|
||||
@@ -40,33 +41,54 @@ jobs:
|
||||
const owner = context.repo.owner;
|
||||
const repo = context.repo.repo;
|
||||
const prNumber = context.payload.issue.number;
|
||||
try { await github.rest.issues.removeLabel({ owner, repo, issue_number: prNumber, name: 'ready' }); } catch {}
|
||||
await github.rest.issues.addLabels({ owner, repo, issue_number: prNumber, labels: ['ready'] });
|
||||
|
||||
- name: React to comment
|
||||
if: steps.perm.outputs.has_write == 'true'
|
||||
continue-on-error: true
|
||||
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
|
||||
with:
|
||||
script: |
|
||||
await github.rest.reactions.createForIssueComment({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
owner, repo,
|
||||
comment_id: context.payload.comment.id,
|
||||
content: 'rocket',
|
||||
});
|
||||
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
|
||||
core.setOutput('pr_sha', pr.head.sha);
|
||||
core.setOutput('pr_branch', pr.head.ref);
|
||||
core.setOutput('pr_number', String(prNumber));
|
||||
core.setOutput('pr_title', pr.title);
|
||||
|
||||
trigger-merge-gate:
|
||||
needs: handle-merge
|
||||
if: needs.handle-merge.result == 'success'
|
||||
permissions:
|
||||
actions: read
|
||||
contents: read
|
||||
pull-requests: read
|
||||
uses: ./.github/workflows/ci-trigger-full-suite.yml
|
||||
with:
|
||||
pr_number: ${{ github.event.issue.number }}
|
||||
secrets:
|
||||
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
|
||||
- name: Trigger Full Suite
|
||||
if: steps.perm.outputs.has_write == 'true'
|
||||
env:
|
||||
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
|
||||
PR_SHA: ${{ steps.label.outputs.pr_sha }}
|
||||
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
|
||||
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
|
||||
PR_TITLE: ${{ steps.label.outputs.pr_title }}
|
||||
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
|
||||
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
|
||||
run: |
|
||||
curl -sS --fail-with-body -X POST \
|
||||
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
|
||||
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
-H "Content-Type: application/json" \
|
||||
--data-raw "$(jq -n \
|
||||
--arg commit "$PR_SHA" \
|
||||
--arg branch "$PR_BRANCH" \
|
||||
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
|
||||
--arg pr_title "$PR_TITLE" \
|
||||
--argjson pr_id "$PR_NUMBER" \
|
||||
'{
|
||||
commit: $commit,
|
||||
branch: $branch,
|
||||
message: $message,
|
||||
ignore_pipeline_branch_filters: true,
|
||||
pull_request_id: $pr_id,
|
||||
pull_request_base_branch: "main",
|
||||
env: {
|
||||
TEST_SCOPE: "full",
|
||||
FULL_SUITE: "true",
|
||||
PR_NUMBER: ($pr_id | tostring),
|
||||
PR_TITLE: $pr_title
|
||||
}
|
||||
}')"
|
||||
|
||||
parse-command:
|
||||
if: >-
|
||||
@@ -107,7 +129,7 @@ jobs:
|
||||
set -euo pipefail
|
||||
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
|
||||
|
||||
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval unit-ci kernel-ci dreamverse-ci ssim-ci golden-gate-ci encoder-ci vae-ci transformer-ci lora-inference-ci lora-training-ci lora-extraction-ci training-ci distillation-ci self-forcing-ci vsa-ci vmoba-ci performance-ci api-ci train-framework-ci eval-ci full fastcheck pre-commit"
|
||||
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
|
||||
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
|
||||
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
|
||||
exit 1
|
||||
@@ -115,17 +137,7 @@ jobs:
|
||||
|
||||
declare -A MAP=(
|
||||
[encoder]=encoder [vae]=vae [transformer]=transformer
|
||||
[kernel]=kernel_tests [unit]=unit_test [unit-ci]=unit_test_ci
|
||||
[kernel-ci]=kernel_tests_ci [dreamverse-ci]=dreamverse_app_ci
|
||||
[ssim-ci]=ssim_ci [vmoba-ci]=inference_vmoba_ci
|
||||
[golden-gate-ci]=golden_gate_ci [training-ci]=training_ci
|
||||
[encoder-ci]=encoder_ci [vae-ci]=vae_ci [transformer-ci]=transformer_ci
|
||||
[lora-inference-ci]=inference_lora_ci [lora-training-ci]=training_lora_ci
|
||||
[lora-extraction-ci]=lora_extraction_ci [distillation-ci]=distillation_dmd_ci
|
||||
[self-forcing-ci]=self_forcing_ci [vsa-ci]=training_vsa_ci
|
||||
[performance-ci]=performance_ci [api-ci]=api_server_ci
|
||||
[train-framework-ci]=train_framework_ci [eval-ci]=eval_ci
|
||||
[dreamverse]=dreamverse_app
|
||||
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
|
||||
[ssim]=ssim [golden-gate]=golden_gate [training]=training
|
||||
[lora-inference]=inference_lora [lora-training]=training_lora
|
||||
[lora-extraction]=lora_extraction
|
||||
|
||||
@@ -1,232 +1,81 @@
|
||||
name: Trigger Merge Gate
|
||||
name: Trigger Full Suite
|
||||
|
||||
on:
|
||||
pull_request_target:
|
||||
types: [labeled, synchronize]
|
||||
workflow_call:
|
||||
inputs:
|
||||
pr_number:
|
||||
description: Pull request number to enter into the merge gate
|
||||
required: true
|
||||
type: number
|
||||
secrets:
|
||||
BUILDKITE_API_TOKEN:
|
||||
required: true
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
pull-requests: read
|
||||
actions: read
|
||||
|
||||
concurrency:
|
||||
group: full-suite-${{ github.event.pull_request.number }}
|
||||
cancel-in-progress: false
|
||||
|
||||
jobs:
|
||||
trigger:
|
||||
if: >-
|
||||
inputs.pr_number > 0
|
||||
|| (github.event.action == 'labeled' && github.event.label.name == 'ready')
|
||||
(github.event.action == 'labeled' && github.event.label.name == 'ready')
|
||||
|| github.event.action == 'synchronize'
|
||||
runs-on: ubuntu-latest
|
||||
# Job-level concurrency: only this guarded job acquires the group, so an
|
||||
# unrelated `labeled` event (which skips the job) cannot cancel an in-flight
|
||||
# gate and then skip its replacement. The newest real trigger (`ready`,
|
||||
# push, or `/merge`) supersedes the in-flight run, whose Buildkite build the
|
||||
# cancel step below replaces.
|
||||
concurrency:
|
||||
group: merge-gate-${{ inputs.pr_number || github.event.pull_request.number }}
|
||||
cancel-in-progress: true
|
||||
# Gate below may wait for cheap checks (up to MAX_WAIT_SECS = 25 min).
|
||||
timeout-minutes: 35
|
||||
steps:
|
||||
- name: Check ready label
|
||||
id: check
|
||||
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
|
||||
env:
|
||||
CALLED_PR_NUMBER: ${{ inputs.pr_number }}
|
||||
with:
|
||||
script: |
|
||||
const eventPrNumber = context.payload.pull_request?.number;
|
||||
const calledPrNumber = Number(process.env.CALLED_PR_NUMBER);
|
||||
const prNumber = eventPrNumber ?? calledPrNumber;
|
||||
if (!Number.isSafeInteger(prNumber) || prNumber <= 0) {
|
||||
core.setFailed(`Invalid pull request number: ${process.env.CALLED_PR_NUMBER}`);
|
||||
return;
|
||||
}
|
||||
const { data: pr } = await github.rest.pulls.get({
|
||||
owner: context.repo.owner,
|
||||
repo: context.repo.repo,
|
||||
pull_number: prNumber,
|
||||
pull_number: context.payload.pull_request.number,
|
||||
});
|
||||
if (pr.state !== 'open') {
|
||||
core.setFailed(`PR #${prNumber} is not open.`);
|
||||
return;
|
||||
}
|
||||
if (pr.base.repo.full_name !== context.payload.repository.full_name
|
||||
|| pr.base.ref !== context.payload.repository.default_branch) {
|
||||
core.setFailed(`PR #${prNumber} does not target this repository's default branch.`);
|
||||
return;
|
||||
}
|
||||
const hasReady = pr.labels.some(l => l.name === 'ready');
|
||||
core.setOutput('has_ready', String(hasReady));
|
||||
core.setOutput('changed_files', String(pr.changed_files));
|
||||
core.setOutput('pr_number', String(pr.number));
|
||||
core.setOutput('head_sha', pr.head.sha);
|
||||
core.setOutput('head_ref', pr.head.ref);
|
||||
core.setOutput('base_sha', pr.base.sha);
|
||||
core.setOutput('title', pr.title);
|
||||
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
|
||||
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
|
||||
|
||||
- name: Cancel previous Buildkite builds
|
||||
# Cancelling stale builds only saves agent time. If it cannot run, the
|
||||
# merge gate must still be triggered by the steps below, so a failure
|
||||
# here is reported and stepped over rather than ending the job.
|
||||
continue-on-error: true
|
||||
timeout-minutes: 3
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
env:
|
||||
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
|
||||
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
|
||||
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
|
||||
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
|
||||
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
|
||||
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
response_file=$(mktemp)
|
||||
builds_file=$(mktemp)
|
||||
trap 'rm -f "$response_file" "$builds_file"' EXIT
|
||||
|
||||
if [[ ! "$PR_NUMBER" =~ ^[1-9][0-9]*$ ]]; then
|
||||
echo "::warning::Invalid pull request number; stale Buildkite builds may continue."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 --get \
|
||||
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
--data-urlencode "branch=$PR_BRANCH" \
|
||||
--data-urlencode "state[]=running" \
|
||||
--data-urlencode "state[]=scheduled" \
|
||||
--data-urlencode "state[]=failing" \
|
||||
--data-urlencode "exclude_jobs=true" \
|
||||
--data-urlencode "exclude_pipeline=true" \
|
||||
--output "$response_file" \
|
||||
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds"; then
|
||||
echo "::warning::Could not list Buildkite builds; stale merge-gate builds may continue."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! jq -e '
|
||||
if type != "array" then false
|
||||
else all(.[];
|
||||
if type != "object" then false
|
||||
else
|
||||
(.number | if type == "number" then . > 0 and floor == . else false end)
|
||||
and (
|
||||
(.env? | if . == null then {} else . end) as $env
|
||||
| if ($env | type) != "object" then false
|
||||
else
|
||||
($env.TEST_SCOPE? | . == null or type == "string")
|
||||
and ($env.PR_NUMBER? | . == null or type == "string")
|
||||
end
|
||||
)
|
||||
end
|
||||
)
|
||||
end
|
||||
' "$response_file" >/dev/null 2>&1; then
|
||||
# Do not print the response body: it is remote data and may contain
|
||||
# multiline values that would be interpreted as workflow commands.
|
||||
echo "::warning::Buildkite returned an invalid build list; stale merge-gate builds may continue."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Match both branch and PR number: forks can reuse the same branch name.
|
||||
if ! jq -r --arg pr_number "$PR_NUMBER" '
|
||||
.[]
|
||||
| select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number))
|
||||
| .number
|
||||
' "$response_file" > "$builds_file"; then
|
||||
echo "::warning::Could not select stale Buildkite builds; stale merge-gate builds may continue."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
cancellation_failed=0
|
||||
while IFS= read -r build_num; do
|
||||
# Find running builds for this branch with TEST_SCOPE=full and cancel them
|
||||
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
|
||||
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
|
||||
for build_num in $builds; do
|
||||
echo "Cancelling Buildkite build #$build_num"
|
||||
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 -o /dev/null -X PUT \
|
||||
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds/${build_num}/cancel"; then
|
||||
echo "::warning::Could not cancel Buildkite build #$build_num; trying remaining builds."
|
||||
cancellation_failed=1
|
||||
fi
|
||||
done < "$builds_file"
|
||||
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
|
||||
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
|
||||
done
|
||||
|
||||
if (( cancellation_failed != 0 )); then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Check out the immutable BASE SHA: neither pull_request_target nor the
|
||||
# privileged slash-command call may run code from the untrusted PR head.
|
||||
- name: Checkout trusted merge planner
|
||||
# Checks out the BASE branch (default for pull_request_target), so PR
|
||||
# authors cannot tamper with the gate script.
|
||||
- name: Checkout gate script
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
|
||||
with:
|
||||
ref: ${{ steps.check.outputs.base_sha }}
|
||||
persist-credentials: false
|
||||
|
||||
- name: Collect changed paths
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
|
||||
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
changed_json="$RUNNER_TEMP/merge-changed-files.json"
|
||||
changed_paths="$RUNNER_TEMP/merge-changed-paths.txt"
|
||||
if gh api --paginate --slurp \
|
||||
"repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/files?per_page=100" \
|
||||
> "$changed_json"; then
|
||||
observed=$(jq '[.[][] | .filename] | unique | length' "$changed_json")
|
||||
if [ "$observed" = "$EXPECTED_CHANGED_FILES" ]; then
|
||||
jq -r '.[][] | .filename, (.previous_filename // empty)' "$changed_json" \
|
||||
| sort -u > "$changed_paths"
|
||||
else
|
||||
echo "::warning::Changed-file API returned $observed of $EXPECTED_CHANGED_FILES paths; selecting all merge lanes."
|
||||
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
|
||||
fi
|
||||
else
|
||||
echo "::warning::Changed-file API failed; selecting all merge lanes."
|
||||
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
|
||||
fi
|
||||
|
||||
- name: Select minimal merge tests
|
||||
id: plan
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
run: |
|
||||
python3 .github/scripts/plan_merge_ci.py \
|
||||
--paths-file "$RUNNER_TEMP/merge-changed-paths.txt" \
|
||||
--github-output "$GITHUB_OUTPUT" \
|
||||
--summary-file "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
- name: Wait for pre-commit and docs build
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
PR_SHA: ${{ steps.check.outputs.head_sha }}
|
||||
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
|
||||
PR_SHA: ${{ github.event.pull_request.head.sha }}
|
||||
PR_NUMBER: ${{ github.event.pull_request.number }}
|
||||
run: bash .github/scripts/gate_full_suite.sh
|
||||
|
||||
- name: Trigger Buildkite merge gate
|
||||
- name: Trigger Buildkite Full Suite
|
||||
if: steps.check.outputs.has_ready == 'true'
|
||||
env:
|
||||
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
|
||||
PR_SHA: ${{ steps.check.outputs.head_sha }}
|
||||
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
|
||||
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
|
||||
PR_TITLE: ${{ steps.check.outputs.title }}
|
||||
PR_SHA: ${{ github.event.pull_request.head.sha }}
|
||||
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
|
||||
PR_NUMBER: ${{ github.event.pull_request.number }}
|
||||
PR_TITLE: ${{ github.event.pull_request.title }}
|
||||
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
|
||||
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
|
||||
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
|
||||
MERGE_GOLDEN_TESTS: ${{ steps.plan.outputs.merge_golden_tests }}
|
||||
MERGE_SSIM_TESTS: ${{ steps.plan.outputs.merge_ssim_tests }}
|
||||
MERGE_PLAN_LABEL: ${{ steps.plan.outputs.merge_plan_label }}
|
||||
run: |
|
||||
curl -sS --fail-with-body -X POST \
|
||||
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
|
||||
@@ -235,11 +84,8 @@ jobs:
|
||||
--data-raw "$(jq -n \
|
||||
--arg commit "$PR_SHA" \
|
||||
--arg branch "$PR_BRANCH" \
|
||||
--arg message "Merge gate [${MERGE_PLAN_LABEL}] for PR #${PR_NUMBER}" \
|
||||
--arg message "Full Suite for PR #${PR_NUMBER}" \
|
||||
--arg pr_title "$PR_TITLE" \
|
||||
--arg merge_test_plan "$MERGE_TEST_PLAN" \
|
||||
--arg merge_golden_tests "$MERGE_GOLDEN_TESTS" \
|
||||
--arg merge_ssim_tests "$MERGE_SSIM_TESTS" \
|
||||
--argjson pr_id "$PR_NUMBER" \
|
||||
'{
|
||||
commit: $commit,
|
||||
@@ -249,11 +95,8 @@ jobs:
|
||||
pull_request_id: $pr_id,
|
||||
pull_request_base_branch: "main",
|
||||
env: {
|
||||
TEST_SCOPE: "merge",
|
||||
TEST_SCOPE: "full",
|
||||
FULL_SUITE: "true",
|
||||
MERGE_TEST_PLAN: $merge_test_plan,
|
||||
MERGE_GOLDEN_TESTS: $merge_golden_tests,
|
||||
MERGE_SSIM_TESTS: $merge_ssim_tests,
|
||||
PR_NUMBER: ($pr_id | tostring),
|
||||
PR_TITLE: $pr_title
|
||||
}
|
||||
|
||||
@@ -38,17 +38,17 @@ jobs:
|
||||
|
||||
**How our CI works:**
|
||||
|
||||
PRs run a three-tier CI system:
|
||||
PRs run a two-tier CI system:
|
||||
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
|
||||
2. **Fastcheck** — six core GPU lanes run automatically via Buildkite (~10-15 min).
|
||||
3. **Merge gate** — a reviewer adds `ready`; changed paths select only the relevant integration, training, golden, or SSIM coverage.
|
||||
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
|
||||
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
|
||||
|
||||
**Before your PR is reviewed:**
|
||||
- [ ] `pre-commit run --all-files` passes locally
|
||||
- [ ] You've added or updated tests for your changes
|
||||
- [ ] The PR description explains what and why
|
||||
|
||||
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
|
||||
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
|
||||
|
||||
**Useful links:**
|
||||
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
|
||||
|
||||
@@ -13,11 +13,6 @@ on:
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
build_ci_runner_image:
|
||||
description: 'Build the ARM64 CUDA 13 CI runner image (sm_100)'
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
# Auto-rebuild the CUDA images when a repository-controlled image input
|
||||
# changes on main. This includes the trusted SM89 kernel artifact's source,
|
||||
# metadata/key helper, ABI dependency metadata, and build orchestration.
|
||||
@@ -203,28 +198,6 @@ jobs:
|
||||
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
|
||||
docker buildx imagetools inspect "${TAGS[0]}"
|
||||
|
||||
# The CI runner is ARM64 like DGX Spark, but targets sm_100 rather than sm_121.
|
||||
# Publish a single-architecture variant so the self-hosted CI runner can reuse
|
||||
# the exact prebuilt kernel instead of compiling it in every job.
|
||||
build-ci-runner-image:
|
||||
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_ci_runner_image == 'true' }}
|
||||
uses: ./.github/workflows/_template-build-image.yml
|
||||
with:
|
||||
python_version: '3.12'
|
||||
dockerfile_path: docker/Dockerfile
|
||||
tag_suffix: py3.12-cuda13.0.0-sm100
|
||||
runner: ubuntu-24.04-arm
|
||||
architecture: arm64
|
||||
build_args: |
|
||||
PYTHON_VERSION=3.12
|
||||
CUDA_VERSION=13.0.0
|
||||
UV_TORCH_BACKEND=cu130
|
||||
TORCH_CUDA_ARCH_LIST=10.0
|
||||
CMAKE_BUILD_PARALLEL_LEVEL=1
|
||||
FLASH_ATTN_WHEEL_TAG=cu130torch2.12
|
||||
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
|
||||
secrets: inherit
|
||||
|
||||
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
|
||||
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
|
||||
# required FA4 dependency stack is available and validated on arm64.
|
||||
|
||||
@@ -62,13 +62,12 @@ jobs:
|
||||
cuda-version: '13.0.0'
|
||||
torch-cuda-short: 'cu130'
|
||||
platform:
|
||||
# x86_64 builds the full cu126 + cu130 set. cu130 ships the
|
||||
# data-center Blackwell sm_100a/sm_103a VSA and consumer sm_120a FP4
|
||||
# kernels.
|
||||
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
|
||||
# Blackwell sm_120a FP4 kernels).
|
||||
- os: ubuntu-22.04
|
||||
arch: x86_64
|
||||
wheel-plat: manylinux_2_35_x86_64
|
||||
# aarch64 is Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), not
|
||||
# aarch64 is Blackwell (GB200 sm_100a + DGX Spark / consumer sm_120a), not
|
||||
# Hopper, and Blackwell needs CUDA >= 12.8 — so only the cu130 leg applies.
|
||||
# Added via include so x86 keeps cu126 + cu130 while aarch64 stays cu130-only.
|
||||
include:
|
||||
@@ -125,7 +124,7 @@ jobs:
|
||||
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
|
||||
run: |
|
||||
sudo apt update
|
||||
sudo apt install -y git gcc-11 g++-11 clang-11
|
||||
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
|
||||
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
|
||||
|
||||
# Allow Git to Access Safe Directory
|
||||
@@ -164,24 +163,22 @@ jobs:
|
||||
cd fastvideo-kernel
|
||||
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
|
||||
# Release builds run on GPU-less runners, so set kernels + arch explicitly:
|
||||
# * aarch64 = Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), NOT
|
||||
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer)
|
||||
# covers sm_120a+sm_121a; turbodiffusion covers every listed arch. The
|
||||
# sm_100 FP4 forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
|
||||
# * aarch64 = Blackwell (GB200 sm_100a + DGX Spark/consumer sm_120a), NOT
|
||||
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer, SM120)
|
||||
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
|
||||
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
|
||||
# JIT-compiled at runtime — not built into this wheel.
|
||||
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
|
||||
# + consumer Blackwell sm_120a FP4.
|
||||
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
|
||||
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
|
||||
# The per-arch split in CMakeLists pins the FP4 targets to requested
|
||||
# sm_120a/sm_121a and builds the main extension for the full arch list.
|
||||
# CMAKE_BUILD_PARALLEL_LEVEL caps
|
||||
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
|
||||
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
|
||||
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
|
||||
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
|
||||
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a;12.1a"
|
||||
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
|
||||
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
|
||||
export CMAKE_BUILD_PARALLEL_LEVEL=1
|
||||
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
|
||||
export TORCH_CUDA_ARCH_LIST="9.0a;10.0a;10.3a;12.0a"
|
||||
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
|
||||
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
|
||||
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
|
||||
export CMAKE_BUILD_PARALLEL_LEVEL=1
|
||||
@@ -197,11 +194,7 @@ jobs:
|
||||
python -m build --wheel --outdir dist
|
||||
|
||||
# Fix the wheel to be manylinux compliant
|
||||
# Ubuntu 22.04 ships patchelf 0.14.3, while current auditwheel
|
||||
# requires at least 0.14.5. Use the stable PyPI binary on both
|
||||
# x86_64 and aarch64 release runners.
|
||||
uv pip install --system auditwheel patchelf==0.17.2.4
|
||||
patchelf --version
|
||||
uv pip install --system auditwheel
|
||||
# Point auditwheel at torch libs, but do not vendor them into the wheel.
|
||||
TORCH_LIB_DIR=$(python - <<'PY'
|
||||
import os
|
||||
@@ -218,8 +211,7 @@ jobs:
|
||||
--exclude libtorch.so \
|
||||
--exclude libc10.so \
|
||||
--exclude libc10_cuda.so \
|
||||
--exclude libtorch_python.so \
|
||||
--exclude libnccl.so.2
|
||||
--exclude libtorch_python.so
|
||||
# Move fixed wheels back to dist for upload consistency
|
||||
rm dist/*.whl
|
||||
mv fixed_dist/*.whl dist/
|
||||
@@ -251,8 +243,7 @@ jobs:
|
||||
- name: Download PyPI wheels
|
||||
# Publish the cu130 (CUDA 13) wheels to PyPI for both architectures:
|
||||
# x86_64 — Hopper sm_90a TK + consumer Blackwell sm_120a FP4
|
||||
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_103a/sm_120a/sm_121a)
|
||||
# + C++ FP4 (sm_120a/sm_121a);
|
||||
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_120a) + C++ FP4 (sm_120a);
|
||||
# no TK (Hopper). sm_100 FP4 forward is the FA4 CuTe DSL path in the
|
||||
# fastvideo package (#1221), shipped/JIT separately.
|
||||
# The x86_64 cu126 wheel stays available as a build artifact / GitHub-release asset.
|
||||
|
||||
-10
@@ -55,8 +55,6 @@ eggs/
|
||||
|
||||
# MkDocs documentation
|
||||
site/
|
||||
docs/assets/cookbook-serving.json
|
||||
examples/serving/clients/node_modules/
|
||||
docs/getting_started/examples/
|
||||
docs/examples/
|
||||
docs/inference/examples/
|
||||
@@ -134,16 +132,8 @@ openspec/
|
||||
fastvideo/tests/ssim/reference_videos/**
|
||||
!fastvideo/tests/ssim/reference_videos/**/*.mp4
|
||||
!fastvideo/tests/ssim/reference_videos/**/*.png
|
||||
fastvideo/tests/ssim/.reference_videos_download.lock
|
||||
|
||||
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
|
||||
.kernel_bench/
|
||||
|
||||
# Editor logs and local Python version pins (accidentally committed)
|
||||
*.nvimlog
|
||||
.nvimlog
|
||||
.python-version
|
||||
/LTX-2-Reference/
|
||||
/DFDReference/
|
||||
scripts/benchmarks/minimax_h3_pro6000/headline_results/
|
||||
fastvideo/tests/ssim/.reference_videos_download.lock
|
||||
|
||||
@@ -7,6 +7,3 @@
|
||||
[submodule "fastvideo/third_party/eval/vbench"]
|
||||
path = fastvideo/third_party/eval/vbench
|
||||
url = https://github.com/Vchitect/VBench.git
|
||||
[submodule "fastvideo/third_party/eval/vqeval"]
|
||||
path = fastvideo/third_party/eval/vqeval
|
||||
url = https://github.com/JiusiServe/LongVideoSparseAttention.git
|
||||
|
||||
@@ -9,7 +9,7 @@ exclude: |
|
||||
tests/.*|
|
||||
scripts/.*|
|
||||
fastvideo/dataset/.*|
|
||||
fastvideo/models/(?!wan/(config|vae_config|pipeline_config|definition|__init__)\.py$).*|
|
||||
fastvideo/models/.*|
|
||||
^apps/dreamverse/web/.*|
|
||||
examples/.*|
|
||||
\.agents/.*|
|
||||
|
||||
@@ -66,18 +66,14 @@ Local guidance lives next to the code. Read the in-scope file before editing:
|
||||
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
|
||||
| `fastvideo/configs/AGENTS.md` | Arch + pipeline config dataclasses, `param_names_mapping` |
|
||||
| `fastvideo/models/AGENTS.md` | DiT / VAE / encoder / scheduler / loader layout (pre-commit excluded) |
|
||||
| `fastvideo/models/wan/AGENTS.md` | Wan family-local transformers, VAE, configs, and the SP sharding invariant |
|
||||
| `fastvideo/layers/AGENTS.md` | Tensor-parallel linear/attention layer rules for ports |
|
||||
| `fastvideo/attention/AGENTS.md` | Backend registry + env-var override |
|
||||
| `fastvideo/pipelines/AGENTS.md` | Stage ABC, `basic/<model>/`, `preprocess/`, presets |
|
||||
| `fastvideo/pipelines/basic/wan/AGENTS.md` | Wan sampling stages, first-frame conditioning, DMD/causal boundaries |
|
||||
| `fastvideo/pipelines/basic/magi_human/AGENTS.md` | MagiHuman umbrella repo, lazy-loaded components, packing invariants |
|
||||
| `fastvideo/training/AGENTS.md` | Legacy monolithic pipelines (frozen for existing models) |
|
||||
| `fastvideo/train/AGENTS.md` | New modular trainer (methods × models × callbacks, YAML) |
|
||||
| `fastvideo/tests/AGENTS.md` | Test taxonomy, conftest, pre-commit-excluded path |
|
||||
| `fastvideo/tests/ssim/AGENTS.md` | GPU SSIM regression authoring + reference video sync |
|
||||
| `scripts/checkpoint_conversion/AGENTS.md` | Adding a converter for a new HF/official checkpoint |
|
||||
| `apps/dreamverse/AGENTS.md` | DreamVerse app structure and conventions |
|
||||
|
||||
## Critical: Two Training Stacks Coexist
|
||||
|
||||
|
||||
@@ -3,18 +3,13 @@
|
||||
</div>
|
||||
|
||||
<p align="center">
|
||||
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://haoailab.com/FastVideo/cookbook/"><b>Cookbook</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
|
||||
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
|
||||
</p>
|
||||
|
||||
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
|
||||
|
||||
## NEWS
|
||||
- `2026/10/06`: FastH3 V2 now runs on a single consumer machine: NVIDIA RTX 5090, RTX 4090 and RTX PRO 6000 GPUs, DGX Spark and Apple Silicon. We also release [FastH3 Trim](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4), an experimental pruned model that is 4.2× smaller than base H3 and runs in as little as 8 GB of GPU memory. Get the [models](https://huggingface.co/collections/FastVideo/fastvideo-fasth3) and read the [Blog](https://haoailab.com/blogs/fasth3-rtx/).
|
||||
- `2026/10/06`: FastVideo now supports [Kandinsky 6](https://x.com/kandinskylab_ai/status/2107374635218055345) from Kandinsky Lab: text- and image-to-video with synchronized audio (base and 10-step distilled pi-Flow checkpoints) plus video super-resolution up to 4x. See the [Kandinsky 6 recipes](https://haoailab.com/FastVideo/cookbook/kandinsky6/).
|
||||
- `2026/09/15`: Release [FastH3 8-Step V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2), an eight-forward data-free DMD2 checkpoint distilled from MiniMax-H3 with 80% Video Sparse Attention. Run it with `examples/inference/basic/basic_fasth3_8step.py` or the [FastH3 8-Step V2 recipe](https://haoailab.com/FastVideo/cookbook/minimax-h3/).
|
||||
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
|
||||
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
|
||||
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac. Follow the [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
|
||||
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac—follow the [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
|
||||
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
|
||||
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
|
||||
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
|
||||
@@ -66,12 +61,12 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
|
||||
```
|
||||
|
||||
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
|
||||
[MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
|
||||
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
|
||||
|
||||
> **On an Apple Silicon Mac?** Install with `uv pip install -e '.[mlx]'` from
|
||||
> a clone, then pick a recipe in the
|
||||
> [cookbook](https://haoailab.com/FastVideo/cookbook/). See the
|
||||
> [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
|
||||
> **On an Apple Silicon Mac?** FastVideo runs FastWan text-to-video natively
|
||||
> through an MLX runtime — a 5-second 480p clip generated locally, no cloud,
|
||||
> no discrete GPU. Install with `uv pip install -e '.[mlx]'` and follow the
|
||||
> [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
|
||||
|
||||
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
|
||||
|
||||
@@ -89,7 +84,7 @@ Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virt
|
||||
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
|
||||
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
|
||||
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
|
||||
- Apple Silicon, macOS -> docs/getting_started/installation/mlx.md
|
||||
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
|
||||
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
|
||||
4. Verify the result:
|
||||
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
|
||||
@@ -154,8 +149,6 @@ if __name__ == '__main__':
|
||||
main()
|
||||
```
|
||||
|
||||
`num_gpus=1` runs the worker in-process (weights load once, no extra Python process). On Colab/Kaggle-style machines with ~16GB host RAM, keep `num_gpus=1`; free-tier system memory does not grow with extra T4s, so `num_gpus>1` is likely to OOM.
|
||||
|
||||
Run the script with:
|
||||
|
||||
```bash
|
||||
|
||||
@@ -97,33 +97,13 @@ dreamverse-server --port 8009
|
||||
dreamverse-mock-server --port 8009
|
||||
```
|
||||
|
||||
### Run Dreamverse with FastH3
|
||||
|
||||
Select the VSA data-free FastH3 Preview profile when you start the backend:
|
||||
|
||||
```bash
|
||||
DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
|
||||
```
|
||||
|
||||
The `fast-h3` profile uses four visible GPUs by default. It loads the `MiniMaxAI/MiniMax-H3` base checkpoint and the
|
||||
`vsa-datafree/adapter_model.safetensors` adapter from
|
||||
`FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA`. Each request generates a 124-frame, 768×1344 video with
|
||||
synchronized audio and five sigma-grid points. Dreamverse uses the last frame of each segment as first-frame
|
||||
conditioning for the following segment.
|
||||
|
||||
Set `CUDA_VISIBLE_DEVICES` when you need to choose the four physical GPUs:
|
||||
|
||||
```bash
|
||||
CUDA_VISIBLE_DEVICES=0,1,2,3 DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
|
||||
```
|
||||
|
||||
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
|
||||
> (the default), the backend compiles the segment 1 and segment 2 inference
|
||||
> paths before it reports ready — this can take **tens of minutes on a cold
|
||||
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
|
||||
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
|
||||
> warmup finishes. To defer compilation until the first generated request while
|
||||
> testing, set `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
|
||||
> warmup finishes. For a faster, uncompiled startup while testing, set
|
||||
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
|
||||
|
||||
## Frontend Setup
|
||||
|
||||
@@ -158,40 +138,6 @@ dreamverse-server --host 0.0.0.0 --port 8009
|
||||
The Dreamverse backend defaults to `0.0.0.0:8009` and starts one GPU worker on
|
||||
the first visible GPU by default.
|
||||
|
||||
### Cosmos Predict2.5 DFD continuation (experimental)
|
||||
|
||||
Dreamverse can combine two converted Cosmos Predict2.5 2B packages: the
|
||||
distilled Text2World student creates an unconditioned first segment, then the
|
||||
Data-Forcing Distillation (DFD) Video2World student conditions each later
|
||||
segment on the prior terminal frame. Point the runtime at both local converted
|
||||
packages:
|
||||
|
||||
```bash
|
||||
export DREAMVERSE_MODEL_ID=cosmos25-dfd
|
||||
export DREAMVERSE_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-Distilled-TrigFlow-FastVideo
|
||||
export DREAMVERSE_COSMOS25_DFD_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-DFD-FastVideo
|
||||
export ENABLE_TORCH_COMPILE=0
|
||||
dreamverse-server --host 0.0.0.0 --port 8009
|
||||
```
|
||||
|
||||
The backend loads and warms both model roles before reporting ready. Both use
|
||||
BF16, Torch SDPA, 704x1280 output, 24 FPS, and four steps. Bootstrap segments
|
||||
contain 77 frames. DFD segments contain 81 decoded frames, but Dreamverse drops
|
||||
the repeated conditioning frame before streaming, leaving 80 new frames. An
|
||||
initial user image selects DFD immediately without treating that first frame as
|
||||
a cross-segment overlap.
|
||||
|
||||
The profile uses a 30-minute session lease because sequential generation on
|
||||
GB10-class hardware can exceed Dreamverse's five-minute default while the GPU
|
||||
is still making progress. Deployments can override the lease with
|
||||
`FASTVIDEO_SESSION_TIMEOUT_SECONDS`.
|
||||
|
||||
Cosmos does not produce audio, so the backend supplies duration-matched silent
|
||||
24 kHz audio for the existing browser streaming contract and trims 1,000 audio
|
||||
samples with each repeated DFD boundary frame. Runtime LoRA changes are not
|
||||
supported. Full segments take roughly 145 seconds on GB10, so this profile is a
|
||||
continuation-quality integration rather than a real-time configuration.
|
||||
|
||||
### Check Readiness
|
||||
|
||||
In another shell, verify that the backend process is alive:
|
||||
@@ -273,7 +219,6 @@ selection, and mock-server behavior:
|
||||
pytest apps/dreamverse/dreamverse/tests/test_config.py \
|
||||
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
|
||||
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
|
||||
apps/dreamverse/dreamverse/tests/test_minimax_h3_generation.py \
|
||||
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
|
||||
```
|
||||
|
||||
|
||||
+1
-12
@@ -139,18 +139,7 @@ session.
|
||||
- startup warmup
|
||||
- user join/leave commands
|
||||
- `USER_STEP` execution for each segment
|
||||
- generation-command routing and stream-result delivery
|
||||
|
||||
Model generation has a separate ownership boundary inside each GPU process:
|
||||
|
||||
- `apps/dreamverse/dreamverse/generation_worker.py` selects the backend that the active model profile declares and owns
|
||||
the backend lifecycle.
|
||||
- `apps/dreamverse/dreamverse/ltx2_generation.py` owns LTX-2 generator configuration, video and audio continuation, and
|
||||
runtime LoRA application.
|
||||
- `apps/dreamverse/dreamverse/minimax_h3_generation.py` owns the VSA data-free FastH3 adapter, FastH3 generator and
|
||||
request configuration, and last-frame continuation through MiniMax H3 first-frame conditioning.
|
||||
- `apps/dreamverse/dreamverse/generation_contracts.py` defines the decoded media and stream-trimming result that both
|
||||
model backends return to `apps/dreamverse/dreamverse/gpu_pool.py`.
|
||||
- continuation state between segments
|
||||
|
||||
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
|
||||
|
||||
|
||||
@@ -1,184 +0,0 @@
|
||||
"""Bounded, runtime-local media library shared by the HTTP and generation APIs."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import tempfile
|
||||
import threading
|
||||
import uuid
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
from PIL import Image, UnidentifiedImageError
|
||||
|
||||
IMAGE_LIMIT = 15 * 1024 * 1024
|
||||
MEDIA_LIMIT = 100 * 1024 * 1024
|
||||
STORE_LIMIT = 2 * 1024 * 1024 * 1024
|
||||
ASSET_LIMIT = 100
|
||||
MAX_MEDIA_SECONDS = 30
|
||||
MIME_TYPES = {
|
||||
"image/png": ("image", ".png"),
|
||||
"image/jpeg": ("image", ".jpg"),
|
||||
"image/webp": ("image", ".webp"),
|
||||
"video/mp4": ("video", ".mp4"),
|
||||
"video/quicktime": ("video", ".mov"),
|
||||
"video/webm": ("video", ".webm"),
|
||||
"audio/mpeg": ("audio", ".mp3"),
|
||||
"audio/mp4": ("audio", ".m4a"),
|
||||
"audio/x-m4a": ("audio", ".m4a"),
|
||||
"audio/wav": ("audio", ".wav"),
|
||||
"audio/x-wav": ("audio", ".wav"),
|
||||
"audio/flac": ("audio", ".flac"),
|
||||
"audio/x-flac": ("audio", ".flac"),
|
||||
"audio/ogg": ("audio", ".ogg"),
|
||||
"audio/webm": ("audio", ".webm"),
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class StoredAsset:
|
||||
asset_id: str
|
||||
kind: str
|
||||
path: str
|
||||
name: str
|
||||
mime_type: str
|
||||
size: int
|
||||
|
||||
def public(self) -> dict:
|
||||
return {
|
||||
"asset_id": self.asset_id,
|
||||
"kind": self.kind,
|
||||
"name": self.name,
|
||||
"mime_type": self.mime_type,
|
||||
"size": self.size,
|
||||
"url": f"/assets/{self.asset_id}",
|
||||
}
|
||||
|
||||
|
||||
def validate_media(path: Path, mime_type: str) -> None:
|
||||
"""Inspect content, not filenames; refuse playlists and non-media uploads."""
|
||||
kind = MIME_TYPES[mime_type][0]
|
||||
if kind == "image":
|
||||
try:
|
||||
with Image.open(path) as img:
|
||||
expected = {"image/png": "PNG", "image/jpeg": "JPEG", "image/webp": "WEBP"}[mime_type]
|
||||
if img.format != expected:
|
||||
raise ValueError("The image content does not match its file type.")
|
||||
if img.width * img.height > 16_777_216:
|
||||
raise ValueError("Images must contain at most 16 megapixels.")
|
||||
if getattr(img, "is_animated", False):
|
||||
raise ValueError("Use a still image or upload the animation as a video.")
|
||||
img.verify()
|
||||
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
|
||||
raise ValueError("The image could not be decoded. Use PNG, JPEG, or WebP.") from exc
|
||||
return
|
||||
|
||||
probe = shutil.which(os.getenv("FASTVIDEO_FFPROBE_BIN", "ffprobe"))
|
||||
if not probe:
|
||||
raise ValueError("This runtime needs ffprobe installed to accept video and audio assets.")
|
||||
try:
|
||||
result = subprocess.run(
|
||||
[
|
||||
probe, "-v", "error", "-protocol_whitelist", "file,pipe", "-format_whitelist",
|
||||
"mov,matroska,webm,mp3,wav,flac,ogg", "-show_format", "-show_streams", "-of", "json",
|
||||
str(path)
|
||||
],
|
||||
check=True,
|
||||
capture_output=True,
|
||||
timeout=15,
|
||||
)
|
||||
info = json.loads(result.stdout)
|
||||
formats = set(info.get("format", {}).get("format_name", "").split(","))
|
||||
if not formats.intersection({"mov", "mp4", "matroska", "webm", "mp3", "wav", "flac", "ogg"}):
|
||||
raise ValueError("Upload a media file, not a playlist or external reference.")
|
||||
streams = [stream for stream in info.get("streams", []) if stream.get("codec_type") == kind]
|
||||
if not streams:
|
||||
raise ValueError(f"The file contains no {kind} stream.")
|
||||
for stream in info.get("streams", []):
|
||||
if stream.get("codec_type") == "audio" and int(stream.get("channels", 0)) not in (1, 2):
|
||||
raise ValueError("H3 references require mono or stereo audio, including video soundtracks.")
|
||||
duration = float(info.get("format", {}).get("duration", "nan"))
|
||||
if not math.isfinite(duration) or not 0 < duration <= MAX_MEDIA_SECONDS:
|
||||
raise ValueError(f"Reference video and audio must be between 0 and {MAX_MEDIA_SECONDS} seconds long.")
|
||||
for stream in streams:
|
||||
if kind == "video" and int(stream.get("width", 0)) * int(stream.get("height", 0)) > 8_294_400:
|
||||
raise ValueError("Reference videos must be 4K or smaller.")
|
||||
except (subprocess.SubprocessError, json.JSONDecodeError, OSError) as exc:
|
||||
raise ValueError("The media file could not be decoded. Check its format and try again.") from exc
|
||||
|
||||
|
||||
class AssetStore:
|
||||
"""Assets live until deletion or runtime exit; pinned generation inputs cannot be deleted."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._directory: tempfile.TemporaryDirectory | None = None
|
||||
self._assets: dict[str, StoredAsset] = {}
|
||||
self._pins: dict[str, int] = {}
|
||||
self._lock = threading.RLock()
|
||||
|
||||
def staging_path(self, mime_type: str) -> Path:
|
||||
with self._lock:
|
||||
if mime_type not in MIME_TYPES:
|
||||
raise ValueError("Unsupported media type. Use PNG/JPEG/WebP, MP4/WebM/MOV, or WAV/MP3/M4A/FLAC/OGG.")
|
||||
if len(self._assets) >= ASSET_LIMIT or sum(item.size for item in self._assets.values()) >= STORE_LIMIT:
|
||||
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
|
||||
if self._directory is None:
|
||||
self._directory = tempfile.TemporaryDirectory(prefix="dreamverse-assets-")
|
||||
return Path(self._directory.name) / f"{uuid.uuid4().hex}{MIME_TYPES[mime_type][1]}"
|
||||
|
||||
def add(self, path: Path, name: str, mime_type: str) -> StoredAsset:
|
||||
validate_media(path, mime_type)
|
||||
size = path.stat().st_size
|
||||
if size == 0 or size > (IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT):
|
||||
raise ValueError("The asset is empty or exceeds its upload size limit.")
|
||||
with self._lock:
|
||||
if len(self._assets) >= ASSET_LIMIT or size + sum(item.size
|
||||
for item in self._assets.values()) > STORE_LIMIT:
|
||||
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
|
||||
if self._directory is None or path.parent != Path(self._directory.name):
|
||||
raise ValueError("The asset must be uploaded to this runtime.")
|
||||
display_name = re.sub(r"[\x00-\x1f\x7f/\\]", "_", name).strip()[:200] or "Untitled asset"
|
||||
asset = StoredAsset(path.stem, MIME_TYPES[mime_type][0], str(path), display_name, mime_type, size)
|
||||
self._assets[asset.asset_id] = asset
|
||||
return asset
|
||||
|
||||
def get(self, asset_id: str) -> StoredAsset:
|
||||
with self._lock:
|
||||
if not isinstance(asset_id, str) or not re.fullmatch(r"[a-f0-9]{32}", asset_id):
|
||||
raise ValueError("Invalid asset ID. Upload or select an asset from the library.")
|
||||
asset = self._assets.get(asset_id)
|
||||
if asset is None or not Path(asset.path).is_file():
|
||||
raise ValueError("An asset is no longer available. Upload it again and reselect it.")
|
||||
return asset
|
||||
|
||||
def pin(self, asset_ids: list[str]) -> None:
|
||||
with self._lock:
|
||||
for asset_id in asset_ids:
|
||||
self.get(asset_id)
|
||||
for asset_id in asset_ids:
|
||||
self._pins[asset_id] = self._pins.get(asset_id, 0) + 1
|
||||
|
||||
def release(self, asset_ids: list[str]) -> None:
|
||||
with self._lock:
|
||||
for asset_id in asset_ids:
|
||||
count = self._pins.get(asset_id, 0)
|
||||
if count > 1:
|
||||
self._pins[asset_id] = count - 1
|
||||
else:
|
||||
self._pins.pop(asset_id, None)
|
||||
|
||||
def delete(self, asset_id: str) -> None:
|
||||
with self._lock:
|
||||
asset = self.get(asset_id)
|
||||
if self._pins.get(asset_id, 0):
|
||||
raise ValueError("This asset is in use by a generation session. End the session before deleting it.")
|
||||
Path(asset.path).unlink(missing_ok=True)
|
||||
del self._assets[asset_id]
|
||||
|
||||
|
||||
asset_store = AssetStore()
|
||||
@@ -1,6 +1,6 @@
|
||||
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
|
||||
|
||||
Mirrors how ``apps/dreamverse/dreamverse/ltx2_generation.py`` constructs
|
||||
Mirrors how ``apps/dreamverse/dreamverse/video_generation.py`` constructs
|
||||
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
|
||||
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
|
||||
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
|
||||
@@ -111,13 +111,7 @@ def _build_generator_config(model_path: str, enable_compile: bool, num_gpus: int
|
||||
mode="max-autotune-no-cudagraphs",
|
||||
dynamic=False),
|
||||
use_fsdp_inference=False,
|
||||
# The bundled LTX2 model enables a refinement LoRA during the
|
||||
# first request. NVFP4 otherwise purges the dense weights that
|
||||
# FastVideo's LoRA merge path requires.
|
||||
quantization=QuantizationConfig(
|
||||
transformer_quant="NVFP4",
|
||||
transformer_retain_original_weights=True,
|
||||
),
|
||||
quantization=QuantizationConfig(transformer_quant="NVFP4"),
|
||||
),
|
||||
pipeline=PipelineSelection(
|
||||
components=components,
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
import os
|
||||
from pathlib import Path
|
||||
from typing import cast
|
||||
|
||||
_REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
_SERVER_ROOT = Path(__file__).resolve().parent
|
||||
@@ -56,65 +55,16 @@ FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
|
||||
MODEL_REGISTRY = {
|
||||
"fast-ltx2": {
|
||||
"name": "FastLTX2",
|
||||
"generation_backend": "ltx2",
|
||||
"default_sp_size": 1,
|
||||
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
|
||||
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
|
||||
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
|
||||
},
|
||||
"fast-ltx23": {
|
||||
"name": "FastLTX23",
|
||||
"generation_backend": "ltx2",
|
||||
"default_sp_size": 1,
|
||||
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
|
||||
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
|
||||
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
|
||||
},
|
||||
"fast-h3": {
|
||||
"name": "FastH3",
|
||||
"generation_backend": "minimax_h3",
|
||||
"default_sp_size": 4,
|
||||
"model_path": "MiniMaxAI/MiniMax-H3",
|
||||
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
|
||||
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
|
||||
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
|
||||
"height": 768,
|
||||
"width": 1344,
|
||||
"num_frames": 124,
|
||||
"num_inference_steps": 5,
|
||||
"seed": 1000,
|
||||
},
|
||||
"full-h3": {
|
||||
"name": "MiniMax H3 (Full)",
|
||||
"generation_backend": "minimax_h3",
|
||||
"default_sp_size": 4,
|
||||
"model_path": "MiniMaxAI/MiniMax-H3",
|
||||
"attention_backend": "FLASH_ATTN",
|
||||
"height": 768,
|
||||
"width": 1344,
|
||||
"num_frames": 124,
|
||||
"num_inference_steps": 50,
|
||||
"seed": 1000,
|
||||
"full_checkpoint": True,
|
||||
},
|
||||
"cosmos25-dfd": {
|
||||
"name": "Cosmos Predict2.5 DFD",
|
||||
"generation_backend": "cosmos25_dfd",
|
||||
"default_sp_size": 1,
|
||||
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
|
||||
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
|
||||
"attention_backend": "TORCH_SDPA",
|
||||
"height": 704,
|
||||
"width": 1280,
|
||||
"bootstrap_num_frames": 77,
|
||||
"continuation_num_frames": 81,
|
||||
"fps": 24,
|
||||
"num_inference_steps": 4,
|
||||
"seed": 42,
|
||||
# Six sequential GB10 segments can exceed the legacy five-minute
|
||||
# DreamVerse lease even though the GPU is making progress.
|
||||
"session_timeout_seconds": 1800,
|
||||
},
|
||||
}
|
||||
|
||||
DEFAULT_MODEL_ID = "fast-ltx2"
|
||||
@@ -126,6 +76,22 @@ if ACTIVE_MODEL_ID not in MODEL_REGISTRY:
|
||||
# Active model configuration
|
||||
MODEL_CONFIG = MODEL_REGISTRY[ACTIVE_MODEL_ID]
|
||||
|
||||
# Generation limits
|
||||
SESSION_TIMEOUT_SECONDS = 300
|
||||
|
||||
# Frame settings
|
||||
NUM_FRAMES = 121
|
||||
FRAME_HEIGHT = 1088
|
||||
FRAME_WIDTH = 1920
|
||||
NUM_INFERENCE_STEPS = 5
|
||||
JPEG_QUALITY = 100
|
||||
BATCH_SIZE = 3
|
||||
|
||||
# Streaming mode:
|
||||
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
|
||||
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
|
||||
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
|
||||
|
||||
|
||||
def _env_int(name: str, default: int) -> int:
|
||||
value = os.getenv(name)
|
||||
@@ -202,42 +168,10 @@ def _optional_env(*names: str) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
# Generation limits
|
||||
# Slower backends may own a longer default lease. A profile can set
|
||||
# ``session_timeout_seconds``; Full H3 loads and generates substantially longer
|
||||
# than the Preview adapter, which also covers a base/ref pipeline reload inside a
|
||||
# retained session. An explicit environment override remains available for
|
||||
# deployment policy: DREAMVERSE_SESSION_TIMEOUT_SECONDS, with
|
||||
# FASTVIDEO_SESSION_TIMEOUT_SECONDS accepted as an alias.
|
||||
# Values below 60 seconds are floored so a single segment cannot outlast the session.
|
||||
_DEFAULT_SESSION_TIMEOUT_SECONDS = cast(
|
||||
int, MODEL_CONFIG.get("session_timeout_seconds", 7200 if ACTIVE_MODEL_ID == "full-h3" else 300))
|
||||
SESSION_TIMEOUT_SECONDS = max(
|
||||
60,
|
||||
_env_int(
|
||||
"DREAMVERSE_SESSION_TIMEOUT_SECONDS",
|
||||
_env_int("FASTVIDEO_SESSION_TIMEOUT_SECONDS", _DEFAULT_SESSION_TIMEOUT_SECONDS),
|
||||
),
|
||||
)
|
||||
|
||||
# Frame settings
|
||||
NUM_FRAMES = 121
|
||||
FRAME_HEIGHT = 1088
|
||||
FRAME_WIDTH = 1920
|
||||
NUM_INFERENCE_STEPS = 5
|
||||
JPEG_QUALITY = 100
|
||||
BATCH_SIZE = 3
|
||||
|
||||
# Streaming mode:
|
||||
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
|
||||
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
|
||||
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
|
||||
|
||||
|
||||
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
|
||||
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
|
||||
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
|
||||
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", cast(int, MODEL_CONFIG["default_sp_size"])))
|
||||
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
|
||||
|
||||
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
|
||||
if DREAMVERSE_MODEL_PATH:
|
||||
@@ -247,13 +181,6 @@ if DREAMVERSE_MODEL_PATH:
|
||||
"config_model_path": DREAMVERSE_MODEL_PATH,
|
||||
}
|
||||
|
||||
DREAMVERSE_COSMOS25_DFD_MODEL_PATH = (os.getenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", "").strip() or None)
|
||||
if DREAMVERSE_COSMOS25_DFD_MODEL_PATH and MODEL_CONFIG.get("generation_backend") == "cosmos25_dfd":
|
||||
MODEL_CONFIG = {
|
||||
**MODEL_CONFIG,
|
||||
"continuation_model_path": DREAMVERSE_COSMOS25_DFD_MODEL_PATH,
|
||||
}
|
||||
|
||||
AVAILABLE_LORAS = {
|
||||
"pixar": {
|
||||
"repo": "vrgamedevgirl84/LTX_2.3_Pixar_Toon_Style_LoRa",
|
||||
@@ -286,7 +213,7 @@ def _resolve_lora_spec(spec: str) -> str | None:
|
||||
if not spec:
|
||||
return None
|
||||
if spec.lower() == "omninft":
|
||||
return cast(str | None, MODEL_CONFIG.get("lora_repo"))
|
||||
return MODEL_CONFIG.get("lora_repo")
|
||||
if spec.lower() in AVAILABLE_LORAS:
|
||||
return AVAILABLE_LORAS[spec.lower()]["repo"]
|
||||
return spec
|
||||
|
||||
@@ -1,252 +0,0 @@
|
||||
"""Cosmos Predict2.5 distilled bootstrap and DFD continuation for DreamVerse."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import gc
|
||||
import os
|
||||
import time
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from dreamverse.generation_contracts import StepResult
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from PIL.Image import Image
|
||||
|
||||
_SILENT_AUDIO_SAMPLE_RATE = 24_000
|
||||
|
||||
|
||||
def _required_config_str(model_config: dict, field_name: str) -> str:
|
||||
value = model_config.get(field_name)
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ValueError(f"Cosmos Predict2.5 DFD model configuration requires `{field_name}`.")
|
||||
return value.strip()
|
||||
|
||||
|
||||
class Cosmos25DFDGenerationBackend:
|
||||
"""Own complementary Cosmos T2W and one-frame-conditioned DFD generators."""
|
||||
|
||||
def __init__(self, gpu_id: int):
|
||||
self.gpu_id = gpu_id
|
||||
self.bootstrap_generator: Any | None = None
|
||||
self.continuation_generator: Any | None = None
|
||||
self.model_config: dict = {}
|
||||
self.continuation_image: Image | None = None
|
||||
|
||||
def _gpu_mem(self) -> str:
|
||||
allocated_gib = torch.cuda.memory_allocated() / 1024**3
|
||||
reserved_gib = torch.cuda.memory_reserved() / 1024**3
|
||||
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
|
||||
|
||||
@staticmethod
|
||||
def _configure_environment(attention_backend: str) -> None:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = attention_backend
|
||||
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
|
||||
|
||||
@staticmethod
|
||||
def _load_generator(model_path: str):
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
return VideoGenerator.from_pretrained(
|
||||
model_path,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
enable_torch_compile=False,
|
||||
)
|
||||
|
||||
def initialize(self, model_config: dict | None = None) -> None:
|
||||
"""Load both package roles so bootstrap and continuation are ready."""
|
||||
if model_config is not None:
|
||||
self.model_config = dict(model_config)
|
||||
if not self.model_config:
|
||||
raise ValueError("Cosmos Predict2.5 DFD initialization requires a model configuration.")
|
||||
|
||||
self.shutdown()
|
||||
bootstrap_path = _required_config_str(self.model_config, "model_path")
|
||||
continuation_path = _required_config_str(self.model_config, "continuation_model_path")
|
||||
attention_backend = _required_config_str(self.model_config, "attention_backend")
|
||||
self._configure_environment(attention_backend)
|
||||
|
||||
print(f"[GPU {self.gpu_id}] Loading Cosmos T2W bootstrap: {bootstrap_path}")
|
||||
print(f"[GPU {self.gpu_id}] Before bootstrap load: {self._gpu_mem()}")
|
||||
self.bootstrap_generator = self._load_generator(bootstrap_path)
|
||||
print(f"[GPU {self.gpu_id}] Loading Cosmos DFD continuation: {continuation_path}")
|
||||
self.continuation_generator = self._load_generator(continuation_path)
|
||||
print(f"[GPU {self.gpu_id}] Cosmos T2W + DFD loaded: {self._gpu_mem()} (warmup pending)")
|
||||
|
||||
def shutdown(self) -> None:
|
||||
"""Release both FastVideo generators and the retained terminal frame."""
|
||||
self.clear_conditioning()
|
||||
for attr_name in ("bootstrap_generator", "continuation_generator"):
|
||||
generator = getattr(self, attr_name)
|
||||
if generator is not None:
|
||||
try:
|
||||
generator.shutdown()
|
||||
except Exception as exc:
|
||||
print(f"[GPU {self.gpu_id}] Cosmos generator shutdown warning: {exc}")
|
||||
setattr(self, attr_name, None)
|
||||
gc.collect()
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.empty_cache()
|
||||
|
||||
def clear_conditioning(self) -> None:
|
||||
if self.continuation_image is not None:
|
||||
self.continuation_image.close()
|
||||
self.continuation_image = None
|
||||
|
||||
@staticmethod
|
||||
def _load_rgb_image(image_path: str) -> Image:
|
||||
from PIL import Image
|
||||
|
||||
with Image.open(image_path) as image:
|
||||
return image.convert("RGB").copy()
|
||||
|
||||
def _select_conditioning_image(
|
||||
self,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
) -> tuple[Image | None, bool]:
|
||||
if reset_conditioning:
|
||||
self.clear_conditioning()
|
||||
if segment_idx > 1 and self.continuation_image is not None:
|
||||
return self.continuation_image.copy(), True
|
||||
if segment_idx > 1 and not reset_conditioning:
|
||||
raise RuntimeError(f"Cosmos DFD segment {segment_idx} requires a retained continuation frame.")
|
||||
if segment_idx == 1 and image_path:
|
||||
return self._load_rgb_image(image_path), False
|
||||
return None, False
|
||||
|
||||
def _sampling_param(self, *, conditioned: bool):
|
||||
# ``num_cond_frames`` is not yet exposed by the typed SamplingConfig,
|
||||
# so this backend uses the compatibility request until that field lands.
|
||||
from fastvideo.api.sampling_param import SamplingParam
|
||||
|
||||
num_frames_key = "continuation_num_frames" if conditioned else "bootstrap_num_frames"
|
||||
return SamplingParam(
|
||||
negative_prompt="",
|
||||
save_video=False,
|
||||
return_frames=True,
|
||||
height=int(self.model_config["height"]),
|
||||
width=int(self.model_config["width"]),
|
||||
num_frames=int(self.model_config[num_frames_key]),
|
||||
fps=int(self.model_config["fps"]),
|
||||
num_inference_steps=int(self.model_config["num_inference_steps"]),
|
||||
guidance_scale=1.0,
|
||||
seed=int(self.model_config["seed"]),
|
||||
num_cond_frames=1 if conditioned else 0,
|
||||
)
|
||||
|
||||
def _save_continuation_frame(self, frame: object) -> None:
|
||||
from PIL import Image
|
||||
|
||||
self.clear_conditioning()
|
||||
if isinstance(frame, Image.Image):
|
||||
self.continuation_image = frame.convert("RGB").copy()
|
||||
return
|
||||
pixels = np.asarray(frame)
|
||||
self.continuation_image = Image.fromarray(np.ascontiguousarray(pixels)).convert("RGB")
|
||||
|
||||
@staticmethod
|
||||
def _silent_audio(frame_count: int, fps: int) -> torch.Tensor:
|
||||
sample_count = max(1, int(round((frame_count / float(fps)) * _SILENT_AUDIO_SAMPLE_RATE)))
|
||||
return torch.zeros(sample_count, dtype=torch.float32)
|
||||
|
||||
def generate_step(
|
||||
self,
|
||||
prompt: str,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> StepResult:
|
||||
"""Generate a T2W start or DFD continuation and retain its last frame."""
|
||||
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
|
||||
raise ValueError("Cosmos supports text generation only through the generation mode API.")
|
||||
if self.bootstrap_generator is None or self.continuation_generator is None:
|
||||
raise RuntimeError("Cosmos T2W + DFD generators are not initialized.")
|
||||
|
||||
conditioning_image, uses_continuation = self._select_conditioning_image(
|
||||
segment_idx,
|
||||
image_path,
|
||||
reset_conditioning,
|
||||
)
|
||||
conditioned = conditioning_image is not None
|
||||
generator = self.continuation_generator if conditioned else self.bootstrap_generator
|
||||
sampling_param = self._sampling_param(conditioned=conditioned)
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
if conditioned:
|
||||
sampling_param.pil_image = conditioning_image
|
||||
result = generator.generate_video(prompt, sampling_param=sampling_param)
|
||||
finally:
|
||||
if conditioning_image is not None:
|
||||
conditioning_image.close()
|
||||
torch.cuda.synchronize()
|
||||
generation_ms = (time.perf_counter() - started) * 1000.0
|
||||
|
||||
if not isinstance(result, dict):
|
||||
raise RuntimeError("Cosmos generation did not return one result dictionary.")
|
||||
frames = result.get("frames")
|
||||
expected_frames = int(sampling_param.num_frames)
|
||||
if not isinstance(frames, list) or len(frames) != expected_frames:
|
||||
actual_frames = len(frames) if isinstance(frames, list) else None
|
||||
raise RuntimeError(f"Cosmos generation returned {actual_frames} frames; expected {expected_frames}.")
|
||||
|
||||
save_started = time.perf_counter()
|
||||
self._save_continuation_frame(frames[-1])
|
||||
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
|
||||
fps = int(sampling_param.fps)
|
||||
timings = {
|
||||
"generation_ms": generation_ms,
|
||||
"generation_time_ms": float(result.get("generation_time") or 0.0) * 1000.0,
|
||||
"save_conditioning_ms": save_conditioning_ms,
|
||||
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
|
||||
}
|
||||
trim_frames = 1 if uses_continuation else 0
|
||||
mode = "DFD continuation" if conditioned else "T2W bootstrap"
|
||||
print(f"[GPU {self.gpu_id}] Cosmos {mode} segment {segment_idx}: "
|
||||
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
|
||||
f"save_conditioning={save_conditioning_ms:.0f}ms, "
|
||||
f"e2e={timings['e2e_latency_ms']:.0f}ms")
|
||||
return StepResult(
|
||||
frames=frames,
|
||||
audio=self._silent_audio(len(frames), fps),
|
||||
audio_sample_rate=_SILENT_AUDIO_SAMPLE_RATE,
|
||||
timings=timings,
|
||||
head_trim_frames=trim_frames,
|
||||
head_trim_audio_frames=trim_frames,
|
||||
)
|
||||
|
||||
def warmup(self, prompt: str) -> dict[str, float]:
|
||||
"""Exercise both T2W bootstrap and retained-frame DFD request shapes."""
|
||||
warmup_prompt = (prompt or "").strip()
|
||||
if not warmup_prompt:
|
||||
raise RuntimeError("Startup warmup prompt must be non-empty.")
|
||||
print(f"[GPU {self.gpu_id}] Cosmos startup warmup starting "
|
||||
"(synthetic segments: T2W bootstrap, DFD continuation)")
|
||||
started = time.perf_counter()
|
||||
bootstrap_result = self.generate_step(warmup_prompt, 1, None, True)
|
||||
continuation_result = self.generate_step(warmup_prompt, 2, None, False)
|
||||
total_ms = (time.perf_counter() - started) * 1000.0
|
||||
self.clear_conditioning()
|
||||
bootstrap_ms = float(bootstrap_result.timings.get("e2e_latency_ms", 0.0))
|
||||
continuation_ms = float(continuation_result.timings.get("e2e_latency_ms", 0.0))
|
||||
print(f"[GPU {self.gpu_id}] Cosmos startup warmup complete: "
|
||||
f"bootstrap={bootstrap_ms:.0f}ms, continuation={continuation_ms:.0f}ms, total={total_ms:.0f}ms")
|
||||
return {
|
||||
"warmup_bootstrap_ms": bootstrap_ms,
|
||||
"warmup_continuation_ms": continuation_ms,
|
||||
"warmup_total_ms": total_ms,
|
||||
}
|
||||
|
||||
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
|
||||
del stack
|
||||
raise RuntimeError("Cosmos Predict2.5 DFD does not support DreamVerse runtime LoRA changes.")
|
||||
@@ -1,49 +0,0 @@
|
||||
"""Shared contract between DreamVerse generation backends and GPU workers."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Protocol
|
||||
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
|
||||
@dataclass
|
||||
class StepResult:
|
||||
"""Decoded media and stream-trimming metadata for one DreamVerse segment."""
|
||||
|
||||
frames: list
|
||||
audio: Any
|
||||
audio_sample_rate: int | None
|
||||
timings: dict[str, float]
|
||||
head_trim_frames: int
|
||||
head_trim_audio_frames: int
|
||||
|
||||
|
||||
class GenerationBackend(Protocol):
|
||||
"""Model-owned generation operations used by one GPU worker process."""
|
||||
|
||||
def initialize(self, model_config: dict | None = None) -> None:
|
||||
...
|
||||
|
||||
def shutdown(self) -> None:
|
||||
...
|
||||
|
||||
def clear_conditioning(self) -> None:
|
||||
...
|
||||
|
||||
def generate_step(
|
||||
self,
|
||||
prompt: str,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> StepResult:
|
||||
...
|
||||
|
||||
def warmup(self, prompt: str) -> dict[str, float]:
|
||||
...
|
||||
|
||||
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
|
||||
...
|
||||
@@ -1,102 +0,0 @@
|
||||
"""GPU-independent validation for generation modes and ordered asset handles."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
from PIL import Image, UnidentifiedImageError
|
||||
|
||||
from dreamverse.assets import asset_store
|
||||
|
||||
GENERATION_MODES = ("t2va", "fl2va", "ref2va")
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class GenerationAsset:
|
||||
asset_id: str
|
||||
kind: str
|
||||
path: str
|
||||
role: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class GenerationInputs:
|
||||
mode: str | None = None
|
||||
assets: tuple[GenerationAsset, ...] = ()
|
||||
|
||||
@property
|
||||
def first_frame_path(self) -> str | None:
|
||||
return next((asset.path for asset in self.assets if asset.role == "first_frame"), None)
|
||||
|
||||
@property
|
||||
def last_frame_path(self) -> str | None:
|
||||
return next((asset.path for asset in self.assets if asset.role == "last_frame"), None)
|
||||
|
||||
@property
|
||||
def references(self) -> tuple[GenerationAsset, ...]:
|
||||
return tuple(asset for asset in self.assets if asset.role == "reference")
|
||||
|
||||
|
||||
def supported_generation_modes(model_id: str) -> tuple[str, ...]:
|
||||
return GENERATION_MODES if model_id in ("full-h3", "mock") else ("t2va", )
|
||||
|
||||
|
||||
def resolve_generation_inputs(payload: dict, model_id: str) -> GenerationInputs:
|
||||
mode = payload.get("generation_mode")
|
||||
raw_assets = payload.get("conditioning_assets", [])
|
||||
if mode is None and "generation_mode" not in payload:
|
||||
if raw_assets:
|
||||
raise ValueError("Select a generation mode before attaching conditioning assets.")
|
||||
return GenerationInputs()
|
||||
if not isinstance(mode, str) or mode not in GENERATION_MODES:
|
||||
raise ValueError("Unknown generation mode. Choose T2VA, FL2VA, or Ref2VA.")
|
||||
if mode not in supported_generation_modes(model_id):
|
||||
raise ValueError(f"{mode.upper()} requires the Full H3 runtime. This runtime is running {model_id}.")
|
||||
if payload.get("initial_image") is not None:
|
||||
raise ValueError("Use asset IDs for generation modes; do not combine them with the legacy initial_image field.")
|
||||
if not isinstance(raw_assets, list) or len(raw_assets) > 12:
|
||||
raise ValueError("conditioning_assets must be an ordered list with at most 12 assets.")
|
||||
if mode == "t2va" and raw_assets:
|
||||
raise ValueError("T2VA accepts text only. Remove conditioning assets or choose another mode.")
|
||||
assets: list[GenerationAsset] = []
|
||||
for item in raw_assets:
|
||||
if not isinstance(item, dict) or set(item) != {"asset_id", "role"}:
|
||||
raise ValueError("Each conditioning asset must contain only asset_id and role.")
|
||||
role = item["role"]
|
||||
if role not in ("first_frame", "last_frame", "reference"):
|
||||
raise ValueError("Asset role must be first_frame, last_frame, or reference.")
|
||||
stored = asset_store.get(item["asset_id"])
|
||||
assets.append(GenerationAsset(stored.asset_id, stored.kind, stored.path, role))
|
||||
if mode == "fl2va":
|
||||
if any(asset.kind != "image" or asset.role == "reference" for asset in assets):
|
||||
raise ValueError("FL2VA accepts only first-frame and last-frame images.")
|
||||
if sum(asset.role == "first_frame" for asset in assets) != 1:
|
||||
raise ValueError("FL2VA requires exactly one first-frame image.")
|
||||
if sum(asset.role == "last_frame" for asset in assets) > 1:
|
||||
raise ValueError("FL2VA accepts at most one last-frame image.")
|
||||
elif mode == "ref2va":
|
||||
if not assets or any(asset.role != "reference" for asset in assets):
|
||||
raise ValueError("Ref2VA requires an ordered list of reference assets, without keyframe roles.")
|
||||
if not any(asset.kind in ("image", "video") for asset in assets):
|
||||
raise ValueError("Ref2VA requires at least one image or video; audio alone is not supported.")
|
||||
for kind, limit in (("image", 9), ("video", 3), ("audio", 3)):
|
||||
if sum(asset.kind == kind for asset in assets) > limit:
|
||||
raise ValueError(f"Ref2VA accepts at most {limit} {kind} references.")
|
||||
for asset in assets:
|
||||
if asset.kind == "image":
|
||||
try:
|
||||
with Image.open(asset.path) as image:
|
||||
if image.width > 4 * image.height or image.height > 4 * image.width:
|
||||
raise ValueError(
|
||||
"Ref2VA image aspect ratios must be between 1:4 and 4:1. Crop this image first.")
|
||||
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
|
||||
raise ValueError("A selected reference image could not be decoded. Upload it again.") from exc
|
||||
return GenerationInputs(mode, tuple(assets))
|
||||
|
||||
|
||||
def pin_generation_inputs(inputs: GenerationInputs) -> None:
|
||||
asset_store.pin([asset.asset_id for asset in inputs.assets])
|
||||
|
||||
|
||||
def release_generation_inputs(inputs: GenerationInputs) -> None:
|
||||
asset_store.release([asset.asset_id for asset in inputs.assets])
|
||||
@@ -1,103 +0,0 @@
|
||||
"""Select and own one model-specific generation backend per GPU process."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dreamverse.config import MODEL_CONFIG
|
||||
from dreamverse.generation_contracts import GenerationBackend, StepResult
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
|
||||
def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBackend:
|
||||
"""Construct the backend that owns the selected model family's behavior."""
|
||||
if backend_name == "ltx2":
|
||||
from dreamverse.ltx2_generation import LTX2GenerationBackend
|
||||
|
||||
return LTX2GenerationBackend(gpu_id)
|
||||
if backend_name == "minimax_h3":
|
||||
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
|
||||
|
||||
return MiniMaxH3GenerationBackend(gpu_id)
|
||||
if backend_name == "cosmos25_dfd":
|
||||
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
|
||||
|
||||
return Cosmos25DFDGenerationBackend(gpu_id)
|
||||
raise ValueError(f"Unsupported DreamVerse generation backend: {backend_name!r}")
|
||||
|
||||
|
||||
class VideoGenerationWorker:
|
||||
"""Delegate GPU lifecycle and generation calls to the active model backend."""
|
||||
|
||||
def __init__(self, gpu_id: int):
|
||||
self.gpu_id = gpu_id
|
||||
self.model_config: dict = dict(MODEL_CONFIG)
|
||||
self.backend_name: str | None = None
|
||||
self.backend: GenerationBackend | None = None
|
||||
|
||||
def initialize(self, model_config: dict | None = None) -> None:
|
||||
"""Load the requested model through its generation backend.
|
||||
|
||||
Model selection belongs here so the GPU process and streaming layers
|
||||
use one stable media contract without importing model-specific code.
|
||||
"""
|
||||
requested_model_config = dict(model_config) if model_config is not None else dict(self.model_config)
|
||||
backend_name = requested_model_config.get("generation_backend")
|
||||
if not isinstance(backend_name, str) or not backend_name:
|
||||
raise ValueError("DreamVerse model configuration requires `generation_backend`.")
|
||||
|
||||
candidate_backend = self.backend
|
||||
if candidate_backend is None or self.backend_name != backend_name:
|
||||
if candidate_backend is not None:
|
||||
candidate_backend.shutdown()
|
||||
candidate_backend = _create_generation_backend(backend_name, self.gpu_id)
|
||||
|
||||
try:
|
||||
candidate_backend.initialize(requested_model_config)
|
||||
except Exception:
|
||||
try:
|
||||
candidate_backend.shutdown()
|
||||
except Exception as shutdown_error:
|
||||
print(f"[GPU {self.gpu_id}] Backend cleanup after initialization failure: {shutdown_error}")
|
||||
self.backend = None
|
||||
self.backend_name = None
|
||||
raise
|
||||
|
||||
self.model_config = requested_model_config
|
||||
self.backend = candidate_backend
|
||||
self.backend_name = backend_name
|
||||
|
||||
def _require_backend(self) -> GenerationBackend:
|
||||
"""Return the initialized backend or fail before processing a command."""
|
||||
if self.backend is None:
|
||||
raise RuntimeError("Generation backend is not initialized.")
|
||||
return self.backend
|
||||
|
||||
def shutdown(self) -> None:
|
||||
"""Release model resources owned by the selected backend."""
|
||||
if self.backend is not None:
|
||||
self.backend.shutdown()
|
||||
|
||||
def clear_conditioning(self) -> None:
|
||||
self._require_backend().clear_conditioning()
|
||||
|
||||
def generate_step(
|
||||
self,
|
||||
prompt: str,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> StepResult:
|
||||
"""Generate one segment through the selected model backend."""
|
||||
return self._require_backend().generate_step(
|
||||
prompt,
|
||||
segment_idx,
|
||||
image_path,
|
||||
reset_conditioning,
|
||||
generation_inputs=generation_inputs,
|
||||
)
|
||||
|
||||
def warmup(self, prompt: str) -> dict[str, float]:
|
||||
return self._require_backend().warmup(prompt)
|
||||
|
||||
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
|
||||
return self._require_backend().apply_lora_stack(stack)
|
||||
@@ -12,7 +12,7 @@ from enum import Enum
|
||||
from multiprocessing import Process, Queue
|
||||
|
||||
from dreamverse.config import (
|
||||
ACTIVE_MODEL_ID,
|
||||
DEFAULT_MODEL_ID,
|
||||
DREAMVERSE_SP_SIZE,
|
||||
MODEL_REGISTRY,
|
||||
STARTUP_WARMUP_ENABLED,
|
||||
@@ -29,7 +29,6 @@ from dreamverse.av_streaming import (
|
||||
generate_stream_id,
|
||||
stream_fmp4,
|
||||
)
|
||||
from dreamverse.generation_inputs import GenerationInputs, pin_generation_inputs, release_generation_inputs
|
||||
from dreamverse.worker_ipc import (
|
||||
CommandPayload,
|
||||
InitAck,
|
||||
@@ -55,7 +54,7 @@ from dreamverse.worker_ipc import (
|
||||
def _parse_requested_gpu_limit() -> int | None:
|
||||
raw_value = os.getenv("FASTVIDEO_GPU_COUNT", "").strip().lower()
|
||||
if not raw_value:
|
||||
return DREAMVERSE_SP_SIZE
|
||||
return 1
|
||||
if raw_value == "all":
|
||||
return None
|
||||
try:
|
||||
@@ -165,12 +164,12 @@ def gpu_worker_process(
|
||||
os.environ["CUDA_VISIBLE_DEVICES"] = cuda_device
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
|
||||
|
||||
from dreamverse.generation_worker import VideoGenerationWorker
|
||||
from dreamverse.video_generation import VideoGenerationWorker
|
||||
|
||||
worker = VideoGenerationWorker(gpu_id)
|
||||
|
||||
def event_loop(first_cmd: Command = None):
|
||||
"""Block on generation commands after the model is initialized."""
|
||||
"""Blocking event loop for LTX2; dispatches user commands."""
|
||||
print(f"[GPU {gpu_id}] Entering event loop")
|
||||
|
||||
def handle_command(cmd: Command):
|
||||
@@ -190,7 +189,6 @@ def gpu_worker_process(
|
||||
segment_idx,
|
||||
image_path=payload.image_path,
|
||||
reset_conditioning=payload.reset_conditioning,
|
||||
generation_inputs=payload.generation_inputs,
|
||||
)
|
||||
head_trim_frames = step_result.head_trim_frames
|
||||
head_trim_audio_frames = step_result.head_trim_audio_frames
|
||||
@@ -434,11 +432,10 @@ class GPUSlot:
|
||||
self.connected_users: set[str] = set()
|
||||
self._pending_futures: dict[str, asyncio.Future] = {}
|
||||
self._stream_queues: dict[str, asyncio.Queue] = {}
|
||||
self._step_asset_inputs: dict[str, GenerationInputs] = {}
|
||||
self._response_reader_task: asyncio.Task | None = None
|
||||
self._active: bool = False
|
||||
self._reader_lock: asyncio.Lock | None = None
|
||||
self.current_model_id: str | None = ACTIVE_MODEL_ID
|
||||
self.current_model_id: str = DEFAULT_MODEL_ID
|
||||
self.shared_stream_buffer = None
|
||||
self.shared_stream_buffer_size = SHARED_STREAM_BUFFER_BYTES
|
||||
|
||||
@@ -666,9 +663,6 @@ class GPUSlot:
|
||||
if isinstance(event, (StepComplete, WarmupComplete)):
|
||||
event.timings["ipc_get_done_ns"] = time.time_ns()
|
||||
|
||||
if isinstance(event, (StepComplete, WorkerError)) and event.user_id is not None:
|
||||
self._release_step_assets(event.user_id)
|
||||
|
||||
user_id = event.user_id
|
||||
if user_id and user_id in self._pending_futures:
|
||||
future = self._pending_futures.pop(user_id)
|
||||
@@ -696,7 +690,7 @@ class GPUSlot:
|
||||
async def join_user(self, user_id: str, model_id: str = None) -> JoinAck:
|
||||
"""Add a user to this GPU."""
|
||||
if model_id is None:
|
||||
model_id = ACTIVE_MODEL_ID
|
||||
model_id = DEFAULT_MODEL_ID
|
||||
|
||||
# Reload model if a different one is requested
|
||||
if model_id != self.current_model_id and model_id in MODEL_REGISTRY:
|
||||
@@ -711,23 +705,16 @@ class GPUSlot:
|
||||
self.connected_users.clear()
|
||||
|
||||
model_config = MODEL_REGISTRY[model_id]
|
||||
try:
|
||||
reload_response = await self._send_command(Command(
|
||||
CommandType.RELOAD_MODEL,
|
||||
payload=ReloadModelPayload(model_config=model_config),
|
||||
user_id="__reload__"),
|
||||
timeout=600.0)
|
||||
except Exception:
|
||||
self.current_model_id = None
|
||||
raise
|
||||
reload_response = await self._send_command(Command(CommandType.RELOAD_MODEL,
|
||||
payload=ReloadModelPayload(model_config=model_config),
|
||||
user_id="__reload__"),
|
||||
timeout=600.0)
|
||||
match reload_response:
|
||||
case ReloadAck():
|
||||
pass
|
||||
case WorkerError(message=msg):
|
||||
self.current_model_id = None
|
||||
raise RuntimeError(f"Model reload failed: {msg}")
|
||||
case _:
|
||||
self.current_model_id = None
|
||||
raise RuntimeError(f"Unexpected reload response: "
|
||||
f"{type(reload_response).__name__}")
|
||||
|
||||
@@ -759,7 +746,6 @@ class GPUSlot:
|
||||
segment_idx: int = 1,
|
||||
image_path: str | None = None,
|
||||
reset_conditioning: bool = False,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> dict[str, float]:
|
||||
"""Execute a generation step for a specific user.
|
||||
|
||||
@@ -773,19 +759,9 @@ class GPUSlot:
|
||||
segment_idx=segment_idx,
|
||||
image_path=image_path,
|
||||
reset_conditioning=bool(reset_conditioning),
|
||||
generation_inputs=generation_inputs,
|
||||
)
|
||||
if generation_inputs is not None:
|
||||
if user_id in self._step_asset_inputs:
|
||||
raise RuntimeError("The previous generation is still using this project's assets.")
|
||||
pin_generation_inputs(generation_inputs)
|
||||
self._step_asset_inputs[user_id] = generation_inputs
|
||||
# Pins intentionally survive a waiter timeout/cancellation: the GPU
|
||||
# command keeps running. The response reader releases them when the
|
||||
# worker actually completes (even if that response is now unmatched).
|
||||
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
|
||||
timeout=1800.0)
|
||||
self._release_step_assets(user_id)
|
||||
match response:
|
||||
case StepComplete(timings=timings):
|
||||
return timings
|
||||
@@ -795,11 +771,6 @@ class GPUSlot:
|
||||
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
|
||||
f"{type(response).__name__}")
|
||||
|
||||
def _release_step_assets(self, user_id: str) -> None:
|
||||
inputs = self._step_asset_inputs.pop(user_id, None)
|
||||
if inputs is not None:
|
||||
release_generation_inputs(inputs)
|
||||
|
||||
async def apply_lora_stack(
|
||||
self,
|
||||
stack: list[tuple[str, float]],
|
||||
@@ -822,9 +793,7 @@ class GPUSlot:
|
||||
async def leave_user(self, user_id: str) -> None:
|
||||
"""Remove a user from this GPU."""
|
||||
try:
|
||||
response = await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
|
||||
if isinstance(response, LeaveAck):
|
||||
self._release_step_assets(user_id)
|
||||
await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
|
||||
except Exception as e:
|
||||
print(f"[GPU {self.gpu_id}] Leave user error: {e}")
|
||||
finally:
|
||||
@@ -861,10 +830,6 @@ class GPUSlot:
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
if self.process is None or not self.process.is_alive():
|
||||
for user_id in list(self._step_asset_inputs):
|
||||
self._release_step_assets(user_id)
|
||||
|
||||
for q in (self.command_queue, self.response_queue):
|
||||
if q is not None:
|
||||
try:
|
||||
|
||||
@@ -15,7 +15,6 @@ from dreamverse.gpu_pool import GPUPool, get_available_gpus
|
||||
from dreamverse.session_logger import SessionEventLogger
|
||||
|
||||
from dreamverse.config import (
|
||||
ACTIVE_MODEL_ID,
|
||||
AVAILABLE_LORAS,
|
||||
DEVTOOLS_ENABLED,
|
||||
FRONTEND_STATIC_DIR_CANDIDATES,
|
||||
@@ -35,8 +34,6 @@ from dreamverse.routes.presets import (
|
||||
curated_presets_router,
|
||||
)
|
||||
from dreamverse.session.controller import SessionController
|
||||
from dreamverse.generation_inputs import supported_generation_modes
|
||||
from dreamverse.routes.assets import router as asset_router
|
||||
|
||||
|
||||
class _HeartbeatAccessLogFilter(logging.Filter):
|
||||
@@ -95,16 +92,10 @@ app.add_middleware(
|
||||
app.include_router(build_health_router(lambda: runtime.gpu_pool))
|
||||
app.include_router(internal_monitor_router)
|
||||
app.include_router(prompt_config_router)
|
||||
app.include_router(asset_router)
|
||||
if DEVTOOLS_ENABLED:
|
||||
app.include_router(curated_presets_router)
|
||||
|
||||
|
||||
@app.get("/generation-capabilities")
|
||||
async def generation_capabilities() -> dict:
|
||||
return {"model_id": ACTIVE_MODEL_ID, "modes": supported_generation_modes(ACTIVE_MODEL_ID), "mock": False}
|
||||
|
||||
|
||||
@app.websocket("/ws")
|
||||
async def websocket_endpoint(websocket: WebSocket):
|
||||
controller = SessionController(
|
||||
|
||||
@@ -1,363 +0,0 @@
|
||||
"""Full/Preview H3 lifecycle, conditioning and per-project pipeline selection."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import gc
|
||||
import os
|
||||
import time
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from dreamverse.config import DREAMVERSE_SP_SIZE
|
||||
from dreamverse.generation_contracts import StepResult
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from PIL.Image import Image
|
||||
|
||||
|
||||
def _required_config_str(model_config: dict, field_name: str) -> str:
|
||||
"""Read one required non-empty string from a DreamVerse model profile."""
|
||||
value = model_config.get(field_name)
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ValueError(f"FastH3 model configuration requires `{field_name}`.")
|
||||
return value.strip()
|
||||
|
||||
|
||||
class MiniMaxH3GenerationBackend:
|
||||
"""Own one H3 pipeline at a time and retain base-pipeline continuation."""
|
||||
|
||||
def __init__(self, gpu_id: int):
|
||||
self.gpu_id = gpu_id
|
||||
self.generator: Any | None = None
|
||||
self.model_config: dict = {}
|
||||
self.continuation_image: Image | None = None
|
||||
self.pipeline_mode = "base"
|
||||
|
||||
def _gpu_mem(self) -> str:
|
||||
allocated_gib = torch.cuda.memory_allocated() / 1024**3
|
||||
reserved_gib = torch.cuda.memory_reserved() / 1024**3
|
||||
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
|
||||
|
||||
@staticmethod
|
||||
def _configure_environment(attention_backend: str) -> None:
|
||||
"""Apply the fixed boot-time switches from the FastH3 reference recipe."""
|
||||
os.environ.update({
|
||||
"FASTVIDEO_ATTENTION_BACKEND": attention_backend,
|
||||
"FASTVIDEO_FA4": "1",
|
||||
"FASTVIDEO_MINIMAX_H3_FUSIONS": "all",
|
||||
"FASTVIDEO_VSA_SM100A": "0",
|
||||
})
|
||||
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
|
||||
|
||||
def initialize(self, model_config: dict | None = None) -> None:
|
||||
"""Load the profile's base pipeline; Ref2VA is loaded on first use."""
|
||||
if model_config is not None:
|
||||
self.model_config = dict(model_config)
|
||||
if not self.model_config:
|
||||
raise ValueError("FastH3 initialization requires a model configuration.")
|
||||
self._load_pipeline("base")
|
||||
|
||||
def _load_pipeline(self, pipeline_mode: str) -> None:
|
||||
"""Unload the old executor before loading a base or reference transformer.
|
||||
|
||||
GPU worker commands are serialized, so a project boundary never swaps
|
||||
weights while another request is using them. Keeping one executor also
|
||||
avoids simultaneously retaining two large H3 transformers in VRAM. A
|
||||
failed load leaves no executor behind so the next step retries it.
|
||||
"""
|
||||
full_checkpoint = bool(self.model_config.get("full_checkpoint", False))
|
||||
if pipeline_mode == "ref2va" and not full_checkpoint:
|
||||
raise ValueError("Ref2VA requires the full-h3 model profile.")
|
||||
if self.generator is not None:
|
||||
previous_generator = self.generator
|
||||
self.generator = None
|
||||
previous_generator.shutdown()
|
||||
del previous_generator
|
||||
gc.collect()
|
||||
torch.cuda.empty_cache()
|
||||
|
||||
self.clear_conditioning()
|
||||
model_path = _required_config_str(self.model_config, "model_path")
|
||||
attention_backend = _required_config_str(self.model_config, "attention_backend")
|
||||
self._configure_environment(attention_backend)
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
CompileConfig,
|
||||
ComponentConfig,
|
||||
EngineConfig,
|
||||
GeneratorConfig,
|
||||
OffloadConfig,
|
||||
ParallelismConfig,
|
||||
PipelineSelection,
|
||||
)
|
||||
|
||||
components = ComponentConfig()
|
||||
if not full_checkpoint:
|
||||
from huggingface_hub import hf_hub_download
|
||||
|
||||
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
|
||||
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
|
||||
components.lora_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
|
||||
components.lora_strength = 1.0
|
||||
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
|
||||
if pipeline_mode == "ref2va":
|
||||
components.override_pipeline_cls_name = "MiniMaxH3Ref2VAModularPipeline"
|
||||
experimental = {
|
||||
"attention_backend": attention_backend,
|
||||
"inference_torch_compile": not full_checkpoint and attention_backend == "FLASH_ATTN",
|
||||
"vae_parallel_decode": True,
|
||||
"vae_parallel_decode_strategy": "gather",
|
||||
}
|
||||
if attention_backend == "VIDEO_SPARSE_ATTN_H3":
|
||||
experimental.update({
|
||||
"VSA_sparsity": 0.9,
|
||||
"VSA_tile_size": 64,
|
||||
})
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=model_path,
|
||||
pipeline=PipelineSelection(
|
||||
workload_type="i2v" if pipeline_mode == "ref2va" else None,
|
||||
components=components,
|
||||
experimental=experimental,
|
||||
),
|
||||
engine=EngineConfig(
|
||||
num_gpus=DREAMVERSE_SP_SIZE,
|
||||
parallelism=ParallelismConfig(tp_size=1, sp_size=DREAMVERSE_SP_SIZE),
|
||||
offload=OffloadConfig(
|
||||
dit=False,
|
||||
dit_layerwise=False,
|
||||
text_encoder=True,
|
||||
image_encoder=True,
|
||||
vae=True,
|
||||
pin_cpu_memory=not full_checkpoint,
|
||||
),
|
||||
compile=CompileConfig(enabled=False, vae_enabled=True),
|
||||
use_fsdp_inference=full_checkpoint and DREAMVERSE_SP_SIZE > 1,
|
||||
),
|
||||
)
|
||||
|
||||
print(f"[GPU {self.gpu_id}] Loading H3 model: {model_path} ({pipeline_mode})")
|
||||
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
|
||||
try:
|
||||
self.generator = VideoGenerator.from_config(generator_config)
|
||||
except Exception:
|
||||
# The old executor is already gone; leaving no executor behind lets
|
||||
# the next step retry this load instead of stranding the GPU slot.
|
||||
self.generator = None
|
||||
raise
|
||||
self.pipeline_mode = pipeline_mode
|
||||
print(f"[GPU {self.gpu_id}] FastH3 loaded: {self._gpu_mem()} (warmup pending)")
|
||||
|
||||
def shutdown(self) -> None:
|
||||
"""Release the FastVideo generator and cached continuation image."""
|
||||
self.clear_conditioning()
|
||||
if self.generator is not None:
|
||||
self.generator.shutdown()
|
||||
self.generator = None
|
||||
|
||||
def clear_conditioning(self) -> None:
|
||||
"""Release the first-frame image retained for the next segment."""
|
||||
if self.continuation_image is not None:
|
||||
self.continuation_image.close()
|
||||
self.continuation_image = None
|
||||
|
||||
@staticmethod
|
||||
def _load_rgb_image(image_path: str) -> Image:
|
||||
"""Load an image into an independent RGB buffer with no open file handle."""
|
||||
from PIL import Image
|
||||
|
||||
with Image.open(image_path) as image:
|
||||
return image.convert("RGB").copy()
|
||||
|
||||
def _select_conditioning_image(
|
||||
self,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
) -> tuple[Image | None, bool]:
|
||||
"""Select the initial upload or retained last frame for one segment."""
|
||||
if reset_conditioning:
|
||||
self.clear_conditioning()
|
||||
if segment_idx > 1 and self.continuation_image is not None:
|
||||
return self.continuation_image.copy(), True
|
||||
if segment_idx > 1 and not reset_conditioning:
|
||||
raise RuntimeError(f"FastH3 segment {segment_idx} requires a retained continuation frame.")
|
||||
if segment_idx == 1 and image_path:
|
||||
return self._load_rgb_image(image_path), False
|
||||
return None, False
|
||||
|
||||
def _build_request(
|
||||
self,
|
||||
prompt: str,
|
||||
conditioning_image: Image | None,
|
||||
last_image: Image | None = None,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
):
|
||||
"""Build the typed FastVideo request owned by the FastH3 profile."""
|
||||
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
|
||||
|
||||
references = None
|
||||
if generation_inputs is not None and generation_inputs.mode == "ref2va":
|
||||
from fastvideo.api import MiniMaxH3Reference
|
||||
|
||||
references = [
|
||||
MiniMaxH3Reference(source=str(asset.path), media_type=asset.kind)
|
||||
for asset in generation_inputs.references
|
||||
]
|
||||
return GenerationRequest(
|
||||
prompt=prompt,
|
||||
negative_prompt="",
|
||||
inputs=InputConfig(pil_image=conditioning_image, last_image=last_image, references=references),
|
||||
sampling=SamplingConfig(
|
||||
height=int(self.model_config["height"]),
|
||||
width=int(self.model_config["width"]),
|
||||
num_frames=int(self.model_config["num_frames"]),
|
||||
fps=24,
|
||||
num_inference_steps=int(self.model_config["num_inference_steps"]),
|
||||
guidance_scale=1.0,
|
||||
batch_cfg=False,
|
||||
seed=int(self.model_config["seed"]),
|
||||
),
|
||||
output=OutputConfig(save_video=False, return_frames=True),
|
||||
)
|
||||
|
||||
def _save_continuation_frame(self, frames: list) -> None:
|
||||
"""Retain the last decoded frame as first-frame conditioning."""
|
||||
from PIL import Image
|
||||
|
||||
self.clear_conditioning()
|
||||
self.continuation_image = Image.fromarray(np.ascontiguousarray(frames[-1])).convert("RGB")
|
||||
|
||||
def generate_step(
|
||||
self,
|
||||
prompt: str,
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> StepResult:
|
||||
"""Generate one synchronized FastH3 segment and retain its last frame.
|
||||
|
||||
Later segments use MiniMax H3's first-frame-to-video path. The first
|
||||
conditioned frame and its matching audio duration are trimmed before
|
||||
streaming so adjacent segments do not duplicate media.
|
||||
"""
|
||||
mode = generation_inputs.mode if generation_inputs is not None else None
|
||||
if mode not in (None, "t2va", "fl2va", "ref2va"):
|
||||
raise ValueError(f"Unsupported H3 generation mode: {mode!r}.")
|
||||
if mode in ("fl2va", "ref2va") and not self.model_config.get("full_checkpoint", False):
|
||||
raise ValueError(f"{mode.upper()} requires the full-h3 model profile.")
|
||||
pipeline_mode = "ref2va" if mode == "ref2va" else "base"
|
||||
if self.generator is None or self.pipeline_mode != pipeline_mode:
|
||||
# A failed switch leaves no executor behind; reload here so the
|
||||
# slot recovers on the next step instead of staying broken.
|
||||
if segment_idx > 1 and not reset_conditioning:
|
||||
raise ValueError("Generation mode cannot change in the middle of a project.")
|
||||
self._load_pipeline(pipeline_mode)
|
||||
|
||||
conditioning_image = None
|
||||
last_image = None
|
||||
uses_continuation = False
|
||||
if mode == "ref2va":
|
||||
# The reference pipeline rejects first/last-frame inputs. Preserve
|
||||
# all original references for every clip and do not trim overlap.
|
||||
self.clear_conditioning()
|
||||
else:
|
||||
if mode == "fl2va" and generation_inputs is not None:
|
||||
image_path = generation_inputs.first_frame_path
|
||||
conditioning_image, uses_continuation = self._select_conditioning_image(
|
||||
segment_idx,
|
||||
image_path,
|
||||
reset_conditioning,
|
||||
)
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
if (mode == "fl2va" and segment_idx == 1 and generation_inputs is not None
|
||||
and generation_inputs.last_frame_path):
|
||||
last_image = self._load_rgb_image(generation_inputs.last_frame_path)
|
||||
request = self._build_request(prompt, conditioning_image, last_image, generation_inputs)
|
||||
result = self.generator.generate(request)
|
||||
finally:
|
||||
if conditioning_image is not None:
|
||||
conditioning_image.close()
|
||||
if last_image is not None:
|
||||
last_image.close()
|
||||
torch.cuda.synchronize()
|
||||
generation_ms = (time.perf_counter() - started) * 1000.0
|
||||
|
||||
if isinstance(result, list):
|
||||
raise RuntimeError("FastH3 returned multiple results for one DreamVerse segment.")
|
||||
frames = result.frames
|
||||
if not isinstance(frames, list) or not frames:
|
||||
raise RuntimeError("FastH3 generation did not return decoded frames.")
|
||||
audio = result.audio
|
||||
audio_sample_rate = result.audio_sample_rate
|
||||
if audio is not None and audio_sample_rate is None:
|
||||
raise RuntimeError("FastH3 returned audio without an audio sample rate.")
|
||||
|
||||
save_started = time.perf_counter()
|
||||
if mode != "ref2va":
|
||||
self._save_continuation_frame(frames)
|
||||
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
|
||||
timings = {
|
||||
"generation_ms": generation_ms,
|
||||
"generation_time_ms": float(result.generation_time or 0.0) * 1000.0,
|
||||
"save_conditioning_ms": save_conditioning_ms,
|
||||
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
|
||||
}
|
||||
trim_frames = 1 if uses_continuation else 0
|
||||
print(f"[GPU {self.gpu_id}] FastH3 segment {segment_idx}: "
|
||||
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
|
||||
f"save_conditioning={save_conditioning_ms:.0f}ms, "
|
||||
f"e2e={timings['e2e_latency_ms']:.0f}ms")
|
||||
return StepResult(
|
||||
frames=frames,
|
||||
audio=audio,
|
||||
audio_sample_rate=audio_sample_rate,
|
||||
timings=timings,
|
||||
head_trim_frames=trim_frames,
|
||||
head_trim_audio_frames=trim_frames,
|
||||
)
|
||||
|
||||
def warmup(self, prompt: str) -> dict[str, float]:
|
||||
"""Compile the FastH3 text and first-frame paths before readiness."""
|
||||
warmup_prompt = (prompt or "").strip()
|
||||
if not warmup_prompt:
|
||||
raise RuntimeError("Startup warmup prompt must be non-empty.")
|
||||
print(f"[GPU {self.gpu_id}] FastH3 startup warmup starting "
|
||||
"(synthetic segments: text-to-video, first-frame-to-video)")
|
||||
started = time.perf_counter()
|
||||
text_result = self.generate_step(
|
||||
warmup_prompt,
|
||||
segment_idx=1,
|
||||
image_path=None,
|
||||
reset_conditioning=True,
|
||||
)
|
||||
first_frame_result = self.generate_step(
|
||||
warmup_prompt,
|
||||
segment_idx=2,
|
||||
image_path=None,
|
||||
reset_conditioning=False,
|
||||
)
|
||||
total_ms = (time.perf_counter() - started) * 1000.0
|
||||
self.clear_conditioning()
|
||||
text_ms = float(text_result.timings.get("e2e_latency_ms", 0.0))
|
||||
first_frame_ms = float(first_frame_result.timings.get("e2e_latency_ms", 0.0))
|
||||
print(f"[GPU {self.gpu_id}] FastH3 startup warmup complete: "
|
||||
f"text_to_video={text_ms:.0f}ms, "
|
||||
f"first_frame_to_video={first_frame_ms:.0f}ms, "
|
||||
f"total={total_ms:.0f}ms")
|
||||
return {
|
||||
"warmup_text_to_video_ms": text_ms,
|
||||
"warmup_first_frame_to_video_ms": first_frame_ms,
|
||||
"warmup_total_ms": total_ms,
|
||||
}
|
||||
|
||||
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
|
||||
"""Reject runtime LoRA mutation because FastH3 uses one startup adapter."""
|
||||
del stack
|
||||
raise RuntimeError("FastH3 uses its fixed startup adapter and does not support runtime LoRA changes.")
|
||||
@@ -32,14 +32,6 @@ from fastapi.staticfiles import StaticFiles
|
||||
from dreamverse._deps import require_dreamverse_runtime_deps
|
||||
from dreamverse.config import FRONTEND_STATIC_DIR_CANDIDATES, GENERATION_SEGMENT_CAP
|
||||
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
|
||||
from dreamverse.generation_inputs import (
|
||||
GenerationInputs,
|
||||
pin_generation_inputs,
|
||||
release_generation_inputs,
|
||||
resolve_generation_inputs,
|
||||
supported_generation_modes,
|
||||
)
|
||||
from dreamverse.routes.assets import router as asset_router
|
||||
|
||||
LATENCY_MS = 200
|
||||
SESSION_TIMEOUT_SECONDS = 300
|
||||
@@ -178,12 +170,6 @@ app.add_middleware(
|
||||
allow_methods=["*"],
|
||||
allow_headers=["*"],
|
||||
)
|
||||
app.include_router(asset_router)
|
||||
|
||||
|
||||
@app.get("/generation-capabilities")
|
||||
async def generation_capabilities():
|
||||
return {"model_id": "mock", "modes": supported_generation_modes("mock"), "mock": True}
|
||||
|
||||
|
||||
@app.get("/healthz")
|
||||
@@ -304,7 +290,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
send_lock = asyncio.Lock()
|
||||
stop_event = asyncio.Event()
|
||||
session_init_image = None
|
||||
generation_inputs = GenerationInputs()
|
||||
|
||||
async def ws_send_json(payload: dict) -> None:
|
||||
async with send_lock:
|
||||
@@ -362,13 +347,10 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
generation_paused = bool(initial_rollout_prompt and not single_clip_mode and len(curated_prompts) == 0)
|
||||
|
||||
try:
|
||||
generation_inputs = resolve_generation_inputs(init_data, "mock")
|
||||
pin_generation_inputs(generation_inputs)
|
||||
session_init_image = persist_session_init_image(init_data.get("initial_image"))
|
||||
except ValueError as exc:
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
await websocket.close(code=1003, reason="Invalid initial image")
|
||||
@@ -380,8 +362,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
"type": "gpu_assigned",
|
||||
"gpu_id": 0,
|
||||
"session_timeout": SESSION_TIMEOUT_SECONDS,
|
||||
"generation_mode": generation_inputs.mode,
|
||||
"mock": True,
|
||||
})
|
||||
|
||||
raw_prompt_queue: asyncio.Queue[PromptSubmission] = asyncio.Queue()
|
||||
@@ -506,7 +486,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
})
|
||||
|
||||
async def apply_project_init_payload(payload: dict[str, object], ) -> bool:
|
||||
nonlocal generation_inputs
|
||||
nonlocal preset_id
|
||||
nonlocal preset_label
|
||||
nonlocal initial_rollout_prompt
|
||||
@@ -540,19 +519,10 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
]
|
||||
|
||||
try:
|
||||
next_inputs = resolve_generation_inputs(payload, "mock")
|
||||
pin_generation_inputs(next_inputs)
|
||||
try:
|
||||
replace_session_image(payload.get("initial_image"))
|
||||
except ValueError:
|
||||
release_generation_inputs(next_inputs)
|
||||
raise
|
||||
release_generation_inputs(generation_inputs)
|
||||
generation_inputs = next_inputs
|
||||
replace_session_image(payload.get("initial_image"))
|
||||
except ValueError as exc:
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
return False
|
||||
@@ -598,7 +568,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
return drained
|
||||
|
||||
async def enter_project_idle() -> None:
|
||||
nonlocal generation_inputs
|
||||
nonlocal seed_prompt_memory
|
||||
nonlocal curated_prompts
|
||||
nonlocal curated_idx
|
||||
@@ -618,8 +587,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
|
||||
dropped_raw = drain_queue_nowait(raw_prompt_queue)
|
||||
dropped_ready = drain_queue_nowait(ready_prompt_queue)
|
||||
release_generation_inputs(generation_inputs)
|
||||
generation_inputs = GenerationInputs()
|
||||
seed_prompt_memory = []
|
||||
curated_prompts = []
|
||||
curated_idx = 0
|
||||
@@ -800,17 +767,10 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
continue
|
||||
|
||||
try:
|
||||
if generation_inputs.mode is not None and data.get("initial_image") is not None:
|
||||
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
|
||||
"cannot replace generation mode inputs.")
|
||||
if "generation_mode" in data or "conditioning_assets" in data:
|
||||
raise ValueError(
|
||||
"simple_generate cannot change the mode; use project_init_v1.")
|
||||
replace_session_image(data.get("initial_image"))
|
||||
except ValueError as exc:
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
continue
|
||||
@@ -1222,7 +1182,6 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
finally:
|
||||
stop_event.set()
|
||||
cleanup_session_init_image(session_init_image)
|
||||
release_generation_inputs(generation_inputs)
|
||||
|
||||
|
||||
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
|
||||
|
||||
@@ -1,72 +0,0 @@
|
||||
"""Raw, bounded media uploads keep large binary data out of websocket messages."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from urllib.parse import unquote
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Request, Response
|
||||
from fastapi.responses import FileResponse
|
||||
from starlette.concurrency import run_in_threadpool
|
||||
|
||||
from dreamverse.assets import IMAGE_LIMIT, MEDIA_LIMIT, MIME_TYPES, asset_store
|
||||
|
||||
router = APIRouter()
|
||||
_upload_lock = asyncio.Lock()
|
||||
|
||||
|
||||
@router.post("/assets", status_code=201)
|
||||
async def upload_asset(request: Request) -> dict:
|
||||
mime_type = request.headers.get("content-type", "").split(";", 1)[0].lower()
|
||||
if mime_type not in MIME_TYPES:
|
||||
raise HTTPException(415, "Unsupported asset type. Select a supported image, video, or audio file.")
|
||||
limit = IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT
|
||||
try:
|
||||
if int(request.headers.get("content-length", "0")) > limit:
|
||||
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
|
||||
except ValueError as exc:
|
||||
raise HTTPException(400, "Invalid Content-Length.") from exc
|
||||
async with _upload_lock:
|
||||
try:
|
||||
path = asset_store.staging_path(mime_type)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(400, str(exc)) from exc
|
||||
try:
|
||||
size = 0
|
||||
with path.open("xb") as handle:
|
||||
async for chunk in request.stream():
|
||||
size += len(chunk)
|
||||
if size > limit:
|
||||
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
|
||||
await run_in_threadpool(handle.write, chunk)
|
||||
asset = await run_in_threadpool(asset_store.add, path,
|
||||
unquote(request.headers.get("x-asset-name", "Untitled asset")), mime_type)
|
||||
return asset.public()
|
||||
except ValueError as exc:
|
||||
path.unlink(missing_ok=True)
|
||||
raise HTTPException(400, str(exc)) from exc
|
||||
except BaseException:
|
||||
path.unlink(missing_ok=True)
|
||||
raise
|
||||
|
||||
|
||||
@router.api_route("/assets/{asset_id}", methods=["GET", "HEAD"])
|
||||
async def get_asset(asset_id: str) -> FileResponse:
|
||||
try:
|
||||
asset = asset_store.get(asset_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(404, str(exc)) from exc
|
||||
return FileResponse(asset.path, media_type=asset.mime_type, headers={"X-Content-Type-Options": "nosniff"})
|
||||
|
||||
|
||||
@router.delete("/assets/{asset_id}", status_code=204)
|
||||
async def delete_asset(asset_id: str) -> Response:
|
||||
try:
|
||||
asset_store.get(asset_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(404, str(exc)) from exc
|
||||
try:
|
||||
asset_store.delete(asset_id)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(409, str(exc)) from exc
|
||||
return Response(status_code=204)
|
||||
@@ -26,17 +26,11 @@ from typing import TYPE_CHECKING
|
||||
|
||||
from fastapi import WebSocket, WebSocketDisconnect
|
||||
from dreamverse.gpu_pool import GPUSlot
|
||||
from dreamverse.generation_inputs import (
|
||||
GenerationInputs,
|
||||
pin_generation_inputs,
|
||||
release_generation_inputs,
|
||||
resolve_generation_inputs,
|
||||
)
|
||||
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
|
||||
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
|
||||
|
||||
from dreamverse.config import (
|
||||
ACTIVE_MODEL_ID,
|
||||
DEFAULT_MODEL_ID,
|
||||
GENERATION_SEGMENT_CAP,
|
||||
PROMPT_AUTO_SLEEP_MS,
|
||||
PROMPT_AUTO_TIMEOUT_MS,
|
||||
@@ -162,7 +156,6 @@ class SessionController:
|
||||
prompt_worker_task: asyncio.Task | None = None
|
||||
rewrite_seed_prompts_task: asyncio.Task | None = None
|
||||
session_init_image = None
|
||||
generation_inputs: GenerationInputs | None = None
|
||||
|
||||
async def session_timeout():
|
||||
"""Close the session after timeout."""
|
||||
@@ -198,18 +191,6 @@ class SessionController:
|
||||
init_data = {}
|
||||
|
||||
init_type = init_data.get("type")
|
||||
try:
|
||||
next_generation_inputs = resolve_generation_inputs(init_data, ACTIVE_MODEL_ID)
|
||||
pin_generation_inputs(next_generation_inputs)
|
||||
generation_inputs = next_generation_inputs
|
||||
except ValueError as exc:
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
await websocket.close(code=1008, reason="Invalid generation inputs")
|
||||
return
|
||||
preset_id = init_data.get("preset_id")
|
||||
preset_label = str(init_data.get("preset_label") or "").strip()
|
||||
initial_rollout_prompt = str(init_data.get("initial_rollout_prompt") or "").strip()
|
||||
@@ -283,14 +264,13 @@ class SessionController:
|
||||
timeout_task = asyncio.create_task(session_timeout())
|
||||
|
||||
# Join the engine on this GPU.
|
||||
await slot.join_user(client_id, model_id=ACTIVE_MODEL_ID)
|
||||
await slot.join_user(client_id, model_id=DEFAULT_MODEL_ID)
|
||||
|
||||
# Notify client they're connected to a GPU.
|
||||
await ws_send_json({
|
||||
"type": "gpu_assigned",
|
||||
"gpu_id": gpu_id,
|
||||
"session_timeout": SESSION_TIMEOUT_SECONDS,
|
||||
"generation_mode": generation_inputs.mode,
|
||||
})
|
||||
await log_event(
|
||||
"gpu_assigned",
|
||||
@@ -381,16 +361,10 @@ class SessionController:
|
||||
return
|
||||
|
||||
try:
|
||||
if generation_inputs.mode is not None and payload.get("initial_image") is not None:
|
||||
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
|
||||
"cannot replace generation mode inputs.")
|
||||
if "generation_mode" in payload or "conditioning_assets" in payload:
|
||||
raise ValueError("simple_generate cannot change the mode; use project_init_v1.")
|
||||
replace_session_init_image(payload.get("initial_image"))
|
||||
except ValueError as exc:
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
return
|
||||
@@ -443,7 +417,6 @@ class SessionController:
|
||||
})
|
||||
|
||||
async def apply_project_init_payload(payload: dict[str, object]) -> bool:
|
||||
nonlocal generation_inputs
|
||||
nonlocal preset_id
|
||||
nonlocal preset_label
|
||||
nonlocal initial_rollout_prompt
|
||||
@@ -523,26 +496,15 @@ class SessionController:
|
||||
})
|
||||
return False
|
||||
|
||||
next_generation_inputs = None
|
||||
next_inputs_pinned = False
|
||||
try:
|
||||
next_generation_inputs = resolve_generation_inputs(payload, ACTIVE_MODEL_ID)
|
||||
pin_generation_inputs(next_generation_inputs)
|
||||
next_inputs_pinned = True
|
||||
replace_session_init_image(payload.get("initial_image"))
|
||||
except ValueError as exc:
|
||||
if next_inputs_pinned:
|
||||
release_generation_inputs(next_generation_inputs)
|
||||
await ws_send_json({
|
||||
"type": "error",
|
||||
"error_code": "invalid_generation_input",
|
||||
"message": str(exc),
|
||||
})
|
||||
return False
|
||||
|
||||
release_generation_inputs(generation_inputs)
|
||||
generation_inputs = next_generation_inputs
|
||||
|
||||
initial_rollout_prompt = next_initial_rollout_prompt
|
||||
enhancement_enabled = next_enhancement_enabled
|
||||
auto_extension_enabled = next_auto_extension_enabled
|
||||
@@ -1209,7 +1171,6 @@ class SessionController:
|
||||
return drained
|
||||
|
||||
async def enter_project_idle() -> None:
|
||||
nonlocal generation_inputs
|
||||
nonlocal curated_prompts
|
||||
nonlocal seed_prompt_memory
|
||||
nonlocal curated_idx
|
||||
@@ -1260,9 +1221,6 @@ class SessionController:
|
||||
project_active = False
|
||||
pending_project_end = False
|
||||
|
||||
release_generation_inputs(generation_inputs)
|
||||
generation_inputs = GenerationInputs()
|
||||
|
||||
if project_stream_started:
|
||||
project_stream_started = False
|
||||
await ws_send_json({"type": "ltx2_stream_complete"})
|
||||
@@ -1669,7 +1627,6 @@ class SessionController:
|
||||
segment_idx=segment_idx,
|
||||
image_path=step_image_path,
|
||||
reset_conditioning=step_reset_conditioning,
|
||||
generation_inputs=generation_inputs,
|
||||
))
|
||||
segment_generation_active = True
|
||||
try:
|
||||
@@ -1724,10 +1681,10 @@ class SessionController:
|
||||
print(f"[GPU {gpu_id}] Unknown AV event: "
|
||||
f"{type(event).__name__}")
|
||||
|
||||
# A GPU command cannot be cancelled by cancelling its
|
||||
# asyncio waiter. Await completion before releasing pinned
|
||||
# asset files or making this GPU available to a new user.
|
||||
timings = await step_task
|
||||
if not step_task.done():
|
||||
step_task.cancel()
|
||||
else:
|
||||
timings = await step_task
|
||||
finally:
|
||||
segment_generation_active = False
|
||||
if not step_task.done():
|
||||
@@ -1851,5 +1808,3 @@ class SessionController:
|
||||
await self.gpu_pool.release(client_id)
|
||||
finally:
|
||||
cleanup_session_init_image(session_init_image)
|
||||
if generation_inputs is not None:
|
||||
release_generation_inputs(generation_inputs)
|
||||
|
||||
@@ -2,14 +2,13 @@ from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
from pathlib import Path
|
||||
from types import ModuleType
|
||||
|
||||
import pytest
|
||||
|
||||
SERVER_DIR = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
def _load_config_module() -> ModuleType:
|
||||
def _load_config_module():
|
||||
spec = importlib.util.spec_from_file_location(
|
||||
"server_config_test_module",
|
||||
SERVER_DIR / "config.py",
|
||||
@@ -21,7 +20,7 @@ def _load_config_module() -> ModuleType:
|
||||
return module
|
||||
|
||||
|
||||
def _set_required_prompt_keys(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
def _set_required_prompt_keys(monkeypatch):
|
||||
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
|
||||
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
|
||||
|
||||
@@ -139,140 +138,15 @@ def test_config_enables_prompt_safety_when_requested(monkeypatch):
|
||||
|
||||
def test_config_uses_five_minute_session_timeout(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.delenv("DREAMVERSE_MODEL_ID", raising=False)
|
||||
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.SESSION_TIMEOUT_SECONDS == 300
|
||||
|
||||
|
||||
def test_config_uses_thirty_minute_cosmos25_session_timeout(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
|
||||
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.SESSION_TIMEOUT_SECONDS == 1800
|
||||
|
||||
|
||||
def test_config_allows_session_timeout_override(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
|
||||
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.SESSION_TIMEOUT_SECONDS == 900
|
||||
|
||||
|
||||
def test_config_prefers_dreamverse_session_timeout_over_alias(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
|
||||
monkeypatch.setenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", "1200")
|
||||
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.SESSION_TIMEOUT_SECONDS == 1200
|
||||
|
||||
|
||||
def test_config_rejects_invalid_prompt_provider(monkeypatch):
|
||||
monkeypatch.setenv("FASTVIDEO_PROMPT_PROVIDER", "unsupported")
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
|
||||
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_PROMPT_PROVIDER"):
|
||||
_load_config_module()
|
||||
|
||||
|
||||
def test_config_registers_vsa_datafree_fasth3_profile(monkeypatch):
|
||||
"""The FastH3 registry entry owns the complete fixed Preview recipe."""
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.MODEL_REGISTRY["fast-h3"] == {
|
||||
"name": "FastH3",
|
||||
"generation_backend": "minimax_h3",
|
||||
"default_sp_size": 4,
|
||||
"model_path": "MiniMaxAI/MiniMax-H3",
|
||||
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
|
||||
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
|
||||
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
|
||||
"height": 768,
|
||||
"width": 1344,
|
||||
"num_frames": 124,
|
||||
"num_inference_steps": 5,
|
||||
"seed": 1000,
|
||||
}
|
||||
|
||||
|
||||
def test_config_uses_fasth3_sequence_parallel_default(monkeypatch):
|
||||
"""Selecting FastH3 defaults DreamVerse to its four-GPU topology."""
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "fast-h3")
|
||||
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.ACTIVE_MODEL_ID == "fast-h3"
|
||||
assert module.MODEL_CONFIG["generation_backend"] == "minimax_h3"
|
||||
assert module.DREAMVERSE_SP_SIZE == 4
|
||||
|
||||
|
||||
def test_full_h3_profile_has_no_preview_adapter_and_longer_session(monkeypatch):
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "full-h3")
|
||||
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
|
||||
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
|
||||
module = _load_config_module()
|
||||
assert module.MODEL_CONFIG["full_checkpoint"] is True
|
||||
assert "adapter_repo" not in module.MODEL_CONFIG
|
||||
assert module.MODEL_CONFIG["num_inference_steps"] == 50
|
||||
assert module.DREAMVERSE_SP_SIZE == 4
|
||||
assert module.SESSION_TIMEOUT_SECONDS == 7200
|
||||
|
||||
|
||||
def test_config_registers_cosmos25_dfd_profile(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.MODEL_REGISTRY["cosmos25-dfd"] == {
|
||||
"name": "Cosmos Predict2.5 DFD",
|
||||
"generation_backend": "cosmos25_dfd",
|
||||
"default_sp_size": 1,
|
||||
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
|
||||
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
|
||||
"attention_backend": "TORCH_SDPA",
|
||||
"height": 704,
|
||||
"width": 1280,
|
||||
"bootstrap_num_frames": 77,
|
||||
"continuation_num_frames": 81,
|
||||
"fps": 24,
|
||||
"num_inference_steps": 4,
|
||||
"seed": 42,
|
||||
"session_timeout_seconds": 1800,
|
||||
}
|
||||
|
||||
|
||||
def test_config_selects_cosmos25_package_roles(monkeypatch, tmp_path):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
bootstrap_path = tmp_path / "cosmos25-t2w"
|
||||
continuation_path = tmp_path / "cosmos25-dfd"
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
|
||||
monkeypatch.setenv("DREAMVERSE_MODEL_PATH", str(bootstrap_path))
|
||||
monkeypatch.setenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", str(continuation_path))
|
||||
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert module.ACTIVE_MODEL_ID == "cosmos25-dfd"
|
||||
assert module.MODEL_CONFIG["generation_backend"] == "cosmos25_dfd"
|
||||
assert module.MODEL_CONFIG["model_path"] == str(bootstrap_path)
|
||||
assert module.MODEL_CONFIG["continuation_model_path"] == str(continuation_path)
|
||||
assert module.DREAMVERSE_SP_SIZE == 1
|
||||
|
||||
@@ -1,219 +0,0 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
COSMOS_CONFIG = {
|
||||
"name": "Cosmos Predict2.5 DFD",
|
||||
"generation_backend": "cosmos25_dfd",
|
||||
"default_sp_size": 1,
|
||||
"model_path": "/models/cosmos25-t2w",
|
||||
"continuation_model_path": "/models/cosmos25-dfd",
|
||||
"attention_backend": "TORCH_SDPA",
|
||||
"height": 704,
|
||||
"width": 1280,
|
||||
"bootstrap_num_frames": 77,
|
||||
"continuation_num_frames": 81,
|
||||
"fps": 24,
|
||||
"num_inference_steps": 4,
|
||||
"seed": 42,
|
||||
}
|
||||
|
||||
|
||||
class _RecordingGenerator:
|
||||
def __init__(self, pixel_value: int = 20) -> None:
|
||||
self.pixel_value = pixel_value
|
||||
self.calls: list[dict] = []
|
||||
self.shutdown_calls = 0
|
||||
|
||||
def generate_video(self, prompt, sampling_param):
|
||||
condition = sampling_param.pil_image
|
||||
self.calls.append({
|
||||
"prompt": prompt,
|
||||
"sampling": sampling_param,
|
||||
"conditioning_pixels": None if condition is None else np.asarray(condition).copy(),
|
||||
})
|
||||
frames = [
|
||||
np.full((2, 3, 3), self.pixel_value, dtype=np.uint8)
|
||||
for _ in range(sampling_param.num_frames)
|
||||
]
|
||||
frames[-1] = np.full((2, 3, 3), self.pixel_value + 1, dtype=np.uint8)
|
||||
return {
|
||||
"frames": frames,
|
||||
"generation_time": 0.25,
|
||||
}
|
||||
|
||||
def shutdown(self):
|
||||
self.shutdown_calls += 1
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def backend(monkeypatch) -> Cosmos25DFDGenerationBackend:
|
||||
instance = Cosmos25DFDGenerationBackend(gpu_id=0)
|
||||
instance.model_config = dict(COSMOS_CONFIG)
|
||||
instance.bootstrap_generator = _RecordingGenerator(pixel_value=20)
|
||||
instance.continuation_generator = _RecordingGenerator(pixel_value=40)
|
||||
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.synchronize", lambda: None)
|
||||
|
||||
def fake_sampling_param(*, conditioned):
|
||||
return SimpleNamespace(
|
||||
negative_prompt="",
|
||||
save_video=False,
|
||||
return_frames=True,
|
||||
height=704,
|
||||
width=1280,
|
||||
num_frames=81 if conditioned else 77,
|
||||
fps=24,
|
||||
num_inference_steps=4,
|
||||
guidance_scale=1.0,
|
||||
seed=42,
|
||||
num_cond_frames=1 if conditioned else 0,
|
||||
pil_image=None,
|
||||
)
|
||||
|
||||
monkeypatch.setattr(instance, "_sampling_param", fake_sampling_param)
|
||||
return instance
|
||||
|
||||
|
||||
def test_initialize_loads_both_package_roles(monkeypatch):
|
||||
loaded_paths = []
|
||||
generators = [_RecordingGenerator(), _RecordingGenerator()]
|
||||
backend = Cosmos25DFDGenerationBackend(gpu_id=0)
|
||||
|
||||
def fake_load(model_path):
|
||||
loaded_paths.append(model_path)
|
||||
return generators[len(loaded_paths) - 1]
|
||||
|
||||
monkeypatch.setattr(backend, "_load_generator", fake_load)
|
||||
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
|
||||
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.gc.collect", lambda: 0)
|
||||
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.is_available", lambda: False)
|
||||
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
|
||||
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
|
||||
|
||||
backend.initialize(COSMOS_CONFIG)
|
||||
|
||||
assert loaded_paths == [
|
||||
"/models/cosmos25-t2w",
|
||||
"/models/cosmos25-dfd",
|
||||
]
|
||||
assert backend.bootstrap_generator is generators[0]
|
||||
assert backend.continuation_generator is generators[1]
|
||||
assert backend.model_config == COSMOS_CONFIG
|
||||
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "TORCH_SDPA"
|
||||
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
|
||||
|
||||
|
||||
def test_unconditioned_start_uses_t2w_and_retains_terminal_frame(backend):
|
||||
result = backend.generate_step("first prompt", 1, None, True)
|
||||
|
||||
assert len(backend.bootstrap_generator.calls) == 1
|
||||
assert backend.continuation_generator.calls == []
|
||||
sampling = backend.bootstrap_generator.calls[0]["sampling"]
|
||||
assert sampling.height == 704
|
||||
assert sampling.width == 1280
|
||||
assert sampling.num_frames == 77
|
||||
assert sampling.fps == 24
|
||||
assert sampling.num_inference_steps == 4
|
||||
assert sampling.guidance_scale == 1.0
|
||||
assert sampling.seed == 42
|
||||
assert sampling.num_cond_frames == 0
|
||||
assert sampling.pil_image is None
|
||||
assert result.head_trim_frames == 0
|
||||
assert result.head_trim_audio_frames == 0
|
||||
assert result.audio_sample_rate == 24_000
|
||||
assert result.audio.shape == (77_000, )
|
||||
assert result.audio.count_nonzero() == 0
|
||||
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 21).tolist()
|
||||
|
||||
|
||||
def test_retained_frame_uses_dfd_and_trims_repeated_boundary(backend):
|
||||
backend.generate_step("first prompt", 1, None, True)
|
||||
result = backend.generate_step("pivot right", 2, None, False)
|
||||
|
||||
assert len(backend.continuation_generator.calls) == 1
|
||||
call = backend.continuation_generator.calls[0]
|
||||
sampling = call["sampling"]
|
||||
assert sampling.num_frames == 81
|
||||
assert sampling.num_cond_frames == 1
|
||||
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 21).tolist()
|
||||
assert result.head_trim_frames == 1
|
||||
assert result.head_trim_audio_frames == 1
|
||||
assert result.audio.shape == (81_000, )
|
||||
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 41).tolist()
|
||||
|
||||
|
||||
def test_initial_image_uses_dfd_without_stream_trim(backend, tmp_path: Path):
|
||||
from PIL import Image
|
||||
|
||||
image_path = tmp_path / "initial.png"
|
||||
Image.fromarray(np.full((2, 3, 3), 7, dtype=np.uint8)).save(image_path)
|
||||
|
||||
result = backend.generate_step("animate", 1, str(image_path), True)
|
||||
|
||||
assert backend.bootstrap_generator.calls == []
|
||||
call = backend.continuation_generator.calls[0]
|
||||
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 7).tolist()
|
||||
assert result.head_trim_frames == 0
|
||||
assert result.head_trim_audio_frames == 0
|
||||
|
||||
|
||||
def test_generation_mode_api_accepts_text_only_and_rejects_conditioning_modes(backend):
|
||||
result = backend.generate_step("first prompt", 1, None, True, generation_inputs=GenerationInputs(mode="t2va"))
|
||||
|
||||
assert len(backend.bootstrap_generator.calls) == 1
|
||||
assert result.head_trim_frames == 0
|
||||
|
||||
with pytest.raises(ValueError, match="text generation only"):
|
||||
backend.generate_step("pivot right", 2, None, False, generation_inputs=GenerationInputs(mode="fl2va"))
|
||||
assert backend.continuation_generator.calls == []
|
||||
|
||||
|
||||
def test_missing_later_continuation_fails_before_generation(backend):
|
||||
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
|
||||
backend.generate_step("later prompt", 2, None, False)
|
||||
|
||||
assert backend.bootstrap_generator.calls == []
|
||||
assert backend.continuation_generator.calls == []
|
||||
|
||||
|
||||
def test_reset_later_segment_uses_fresh_t2w_bootstrap(backend):
|
||||
backend.generate_step("first prompt", 1, None, True)
|
||||
|
||||
result = backend.generate_step("new scene", 2, None, True)
|
||||
|
||||
assert len(backend.bootstrap_generator.calls) == 2
|
||||
assert backend.continuation_generator.calls == []
|
||||
assert result.head_trim_frames == 0
|
||||
|
||||
|
||||
def test_warmup_exercises_bootstrap_and_dfd_paths(backend):
|
||||
timings = backend.warmup("warmup prompt")
|
||||
|
||||
assert len(backend.bootstrap_generator.calls) == 1
|
||||
assert len(backend.continuation_generator.calls) == 1
|
||||
assert backend.continuation_image is None
|
||||
assert "warmup_bootstrap_ms" in timings
|
||||
assert "warmup_continuation_ms" in timings
|
||||
assert "warmup_total_ms" in timings
|
||||
|
||||
|
||||
def test_shutdown_releases_both_generators_and_conditioning(backend):
|
||||
bootstrap = backend.bootstrap_generator
|
||||
continuation = backend.continuation_generator
|
||||
backend.generate_step("first prompt", 1, None, True)
|
||||
|
||||
backend.shutdown()
|
||||
|
||||
assert bootstrap.shutdown_calls == 1
|
||||
assert continuation.shutdown_calls == 1
|
||||
assert backend.bootstrap_generator is None
|
||||
assert backend.continuation_generator is None
|
||||
assert backend.continuation_image is None
|
||||
@@ -1,248 +0,0 @@
|
||||
"""Contract regressions independent of CUDA and model weights."""
|
||||
|
||||
import io
|
||||
import asyncio
|
||||
import shutil
|
||||
import subprocess
|
||||
|
||||
import pytest
|
||||
from fastapi import FastAPI
|
||||
from fastapi.testclient import TestClient
|
||||
from PIL import Image
|
||||
|
||||
from dreamverse import assets, generation_inputs
|
||||
from dreamverse.generation_inputs import resolve_generation_inputs
|
||||
from dreamverse.routes import assets as asset_routes
|
||||
from dreamverse.tests.test_mock_server import _FakeWebSocket
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def library(monkeypatch):
|
||||
store = assets.AssetStore()
|
||||
monkeypatch.setattr(asset_routes, "asset_store", store)
|
||||
monkeypatch.setattr(generation_inputs, "asset_store", store)
|
||||
app = FastAPI()
|
||||
app.include_router(asset_routes.router)
|
||||
with TestClient(app) as client:
|
||||
yield store, client
|
||||
|
||||
|
||||
def upload_image(client, color="red"):
|
||||
image_bytes = io.BytesIO()
|
||||
Image.new("RGB", (32, 32), color).save(image_bytes, format="PNG")
|
||||
response = client.post("/assets", content=image_bytes.getvalue(),
|
||||
headers={"Content-Type": "image/png", "X-Asset-Name": "frame.png"})
|
||||
assert response.status_code == 201, response.text
|
||||
return response.json()
|
||||
|
||||
|
||||
def conditioning(asset, role):
|
||||
return {"asset_id": asset["asset_id"], "role": role}
|
||||
|
||||
|
||||
def test_assets_validate_content_and_support_head_range_and_delete(library):
|
||||
store, client = library
|
||||
asset = upload_image(client)
|
||||
assert set(asset) == {"asset_id", "kind", "name", "mime_type", "size", "url"}
|
||||
assert client.head(asset["url"]).status_code == 200
|
||||
response = client.get(asset["url"], headers={"Range": "bytes=0-7"})
|
||||
assert response.status_code == 206
|
||||
assert response.content == b"\x89PNG\r\n\x1a\n"
|
||||
assert client.post("/assets", content=b"not an image", headers={"Content-Type": "image/png"}).status_code == 400
|
||||
assert client.post("/assets", content=b"<svg/>", headers={"Content-Type": "image/svg+xml"}).status_code == 415
|
||||
assert client.post("/assets", content=b"", headers={"Content-Type": "image/png",
|
||||
"Content-Length": str(assets.IMAGE_LIMIT + 1)}).status_code == 413
|
||||
with pytest.raises(ValueError, match="Invalid asset ID"):
|
||||
store.get("../../etc/passwd")
|
||||
assert client.delete(asset["url"]).status_code == 204
|
||||
assert client.head(asset["url"]).status_code == 404
|
||||
|
||||
|
||||
def test_generation_pin_prevents_deletion_until_session_releases(library):
|
||||
_, client = library
|
||||
asset = upload_image(client)
|
||||
inputs = resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
|
||||
conditioning(asset, "first_frame")
|
||||
]}, "full-h3")
|
||||
generation_inputs.pin_generation_inputs(inputs)
|
||||
generation_inputs.pin_generation_inputs(inputs)
|
||||
assert client.delete(asset["url"]).status_code == 409
|
||||
generation_inputs.release_generation_inputs(inputs)
|
||||
assert client.delete(asset["url"]).status_code == 409
|
||||
generation_inputs.release_generation_inputs(inputs)
|
||||
assert client.delete(asset["url"]).status_code == 204
|
||||
|
||||
|
||||
def test_legacy_init_remains_compatible_but_explicit_t2va_is_text_only(library):
|
||||
_, client = library
|
||||
assert resolve_generation_inputs({"initial_image": {"old": "payload"}}, "fast-ltx2").mode is None
|
||||
assert resolve_generation_inputs({"generation_mode": "t2va"}, "fast-ltx2").mode == "t2va"
|
||||
with pytest.raises(ValueError, match="legacy initial_image"):
|
||||
resolve_generation_inputs({"generation_mode": "t2va", "initial_image": {}}, "full-h3")
|
||||
asset = upload_image(client)
|
||||
with pytest.raises(ValueError, match="text only"):
|
||||
resolve_generation_inputs({"generation_mode": "t2va", "conditioning_assets": [
|
||||
conditioning(asset, "reference")
|
||||
]}, "full-h3")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", ["unknown", None, 3, [], {}])
|
||||
def test_unknown_mode_fails_before_assets_are_resolved(mode):
|
||||
with pytest.raises(ValueError, match="Unknown generation mode"):
|
||||
resolve_generation_inputs({"generation_mode": mode}, "full-h3")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("model_id", ["fast-h3", "fast-ltx2", "fast-ltx23"])
|
||||
def test_preview_and_ltx_cannot_advertise_full_h3_modes(model_id):
|
||||
with pytest.raises(ValueError, match="Full H3"):
|
||||
resolve_generation_inputs({"generation_mode": "ref2va"}, model_id)
|
||||
|
||||
|
||||
def test_fl2va_first_required_last_optional_and_roles_unique(library):
|
||||
_, client = library
|
||||
first = upload_image(client)
|
||||
last = upload_image(client, "blue")
|
||||
payload = {"generation_mode": "fl2va", "conditioning_assets": [conditioning(first, "first_frame")]}
|
||||
inputs = resolve_generation_inputs(payload, "full-h3")
|
||||
assert inputs.first_frame_path.endswith(".png")
|
||||
assert inputs.last_frame_path is None
|
||||
payload["conditioning_assets"].append(conditioning(last, "last_frame"))
|
||||
assert resolve_generation_inputs(payload, "full-h3").last_frame_path is not None
|
||||
payload["conditioning_assets"].append(conditioning(first, "first_frame"))
|
||||
with pytest.raises(ValueError, match="exactly one first-frame"):
|
||||
resolve_generation_inputs(payload, "full-h3")
|
||||
with pytest.raises(ValueError, match="exactly one first-frame"):
|
||||
resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
|
||||
conditioning(last, "last_frame")
|
||||
]}, "full-h3")
|
||||
|
||||
|
||||
def test_ref_order_is_preserved_and_limits_are_enforced(library):
|
||||
_, client = library
|
||||
first, second = upload_image(client), upload_image(client, "blue")
|
||||
refs = [conditioning(second, "reference"), conditioning(first, "reference")]
|
||||
payload = {"generation_mode": "ref2va", "conditioning_assets": refs}
|
||||
inputs = resolve_generation_inputs(payload, "full-h3")
|
||||
assert [asset.asset_id for asset in inputs.references] == [second["asset_id"], first["asset_id"]]
|
||||
with pytest.raises(ValueError, match="at most 9 image"):
|
||||
resolve_generation_inputs({**payload, "conditioning_assets": refs * 5}, "full-h3")
|
||||
with pytest.raises(ValueError, match="without keyframe roles"):
|
||||
resolve_generation_inputs({**payload, "conditioning_assets": [conditioning(first, "first_frame")]}, "full-h3")
|
||||
with pytest.raises(ValueError, match="at most 12"):
|
||||
resolve_generation_inputs({**payload, "conditioning_assets": refs * 7}, "full-h3")
|
||||
|
||||
|
||||
def test_ref_audio_requires_visual_reference(library, monkeypatch):
|
||||
store, _ = library
|
||||
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "audio", "/audio.wav", "audio",
|
||||
"audio/wav", 100))
|
||||
with pytest.raises(ValueError, match="audio alone"):
|
||||
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
|
||||
{"asset_id": "a" * 32, "role": "reference"}
|
||||
]}, "full-h3")
|
||||
|
||||
|
||||
def test_ref_rejects_extreme_image_aspect_before_gpu(library):
|
||||
_, client = library
|
||||
content = io.BytesIO()
|
||||
Image.new("RGB", (500, 50), "blue").save(content, format="PNG")
|
||||
response = client.post("/assets", content=content.getvalue(), headers={"Content-Type": "image/png"})
|
||||
assert response.status_code == 201
|
||||
with pytest.raises(ValueError, match="aspect ratios"):
|
||||
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
|
||||
conditioning(response.json(), "reference")
|
||||
]}, "full-h3")
|
||||
|
||||
|
||||
def test_ref_reports_undecodable_image_as_invalid_input(library, monkeypatch, tmp_path):
|
||||
store, _ = library
|
||||
broken = tmp_path / "broken.png"
|
||||
broken.write_bytes(b"not an image")
|
||||
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "image", str(broken), "broken.png",
|
||||
"image/png", 11))
|
||||
with pytest.raises(ValueError, match="could not be decoded"):
|
||||
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
|
||||
{"asset_id": "a" * 32, "role": "reference"}
|
||||
]}, "full-h3")
|
||||
|
||||
|
||||
def test_audio_upload_rejects_surround_sound(library, monkeypatch):
|
||||
import json
|
||||
_, client = library
|
||||
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
|
||||
info = {"format": {"format_name": "wav", "duration": "1"},
|
||||
"streams": [{"codec_type": "audio", "channels": 6}]}
|
||||
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
|
||||
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
|
||||
response = client.post("/assets", content=b"surround wav", headers={"Content-Type": "audio/wav"})
|
||||
assert response.status_code == 400
|
||||
assert "mono or stereo" in response.json()["detail"]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mime,format_name", [("audio/x-m4a", "mov,mp4,m4a,3gp,3g2,mj2"), ("audio/x-flac", "flac")])
|
||||
def test_legacy_audio_mime_aliases_are_accepted(library, monkeypatch, mime, format_name):
|
||||
"""Browsers report x- variants for the M4A and FLAC formats the docs promise."""
|
||||
import json
|
||||
_, client = library
|
||||
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
|
||||
info = {"format": {"format_name": format_name, "duration": "1"},
|
||||
"streams": [{"codec_type": "audio", "channels": 2}]}
|
||||
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
|
||||
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
|
||||
response = client.post("/assets", content=b"audio bytes", headers={"Content-Type": mime})
|
||||
assert response.status_code == 201, response.text
|
||||
assert response.json()["kind"] == "audio"
|
||||
assert response.json()["mime_type"] == mime
|
||||
|
||||
|
||||
@pytest.mark.parametrize("entries", [None, {}, "x", [{"path": "/etc/passwd", "role": "reference"}],
|
||||
[{"asset_id": "x", "role": "unknown"}]])
|
||||
def test_malformed_conditioning_is_rejected(entries):
|
||||
with pytest.raises(ValueError):
|
||||
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": entries}, "full-h3")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", ["t2va", "fl2va", "ref2va"])
|
||||
def test_mock_streams_all_valid_modes_and_releases_assets(library, monkeypatch, mode):
|
||||
from dreamverse import mock_server
|
||||
_, client = library
|
||||
monkeypatch.setattr(mock_server, "MOCK_SEGMENT_BYTES", b"mock-fmp4")
|
||||
monkeypatch.setattr(mock_server, "LATENCY_MS", 1)
|
||||
image = upload_image(client)
|
||||
refs = [] if mode == "t2va" else [conditioning(image, "first_frame" if mode == "fl2va" else "reference")]
|
||||
ws = _FakeWebSocket([
|
||||
(0, {"type": "session_init_v2", "generation_mode": mode, "conditioning_assets": refs,
|
||||
"curated_prompts": ["A bird flies over a lake."], "single_clip_mode": True,
|
||||
"enhancement_enabled": False}),
|
||||
(0.15, {"type": "leave"}),
|
||||
])
|
||||
asyncio.run(mock_server.websocket_endpoint(ws))
|
||||
assert not [event for event in ws.sent_json if event["type"] == "error"]
|
||||
assert any(event["type"] == "media_segment_complete" for event in ws.sent_json)
|
||||
assert ws.sent_bytes
|
||||
assert client.delete(image["url"]).status_code == 204
|
||||
|
||||
|
||||
def test_mock_rejects_invalid_mode_before_gpu_assignment(library):
|
||||
from dreamverse import mock_server
|
||||
ws = _FakeWebSocket([(0, {"type": "session_init_v2", "generation_mode": "fl2va"})])
|
||||
asyncio.run(mock_server.websocket_endpoint(ws))
|
||||
assert not any(event["type"] == "gpu_assigned" for event in ws.sent_json)
|
||||
errors = [event for event in ws.sent_json if event["type"] == "error"]
|
||||
assert errors[0]["error_code"] == "invalid_generation_input"
|
||||
assert "first-frame" in errors[0]["message"]
|
||||
|
||||
|
||||
@pytest.mark.skipif(not shutil.which("ffmpeg") or not shutil.which("ffprobe"), reason="ffmpeg + ffprobe required")
|
||||
@pytest.mark.parametrize("kind,mime,suffix", [("video", "video/mp4", ".mp4"), ("audio", "audio/wav", ".wav")])
|
||||
def test_actual_video_and_audio_upload_validation(library, tmp_path, kind, mime, suffix):
|
||||
_, client = library
|
||||
media_path = tmp_path / f"sample{suffix}"
|
||||
source = "testsrc2=size=64x64:rate=24" if kind == "video" else "sine=frequency=440:sample_rate=24000"
|
||||
command = [shutil.which("ffmpeg"), "-v", "error", "-f", "lavfi", "-i", source, "-t", "0.5", str(media_path)]
|
||||
subprocess.run(command, check=True, capture_output=True, timeout=30)
|
||||
response = client.post("/assets", content=media_path.read_bytes(), headers={"Content-Type": mime})
|
||||
assert response.status_code == 201, response.text
|
||||
assert response.json()["kind"] == kind
|
||||
response = client.post("/assets", content=b"#EXTM3U\nhttp://example.com/stream", headers={"Content-Type": mime})
|
||||
assert response.status_code == 400
|
||||
@@ -1,214 +0,0 @@
|
||||
"""CPU contract tests; fake executors do not validate generated-media quality."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
import pickle
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from types import ModuleType, SimpleNamespace
|
||||
from unittest.mock import Mock
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
from PIL import Image
|
||||
|
||||
from dreamverse.config import MODEL_REGISTRY
|
||||
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
|
||||
from dreamverse.generation_worker import VideoGenerationWorker
|
||||
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
|
||||
from dreamverse.worker_ipc import UserStepPayload
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fastvideo_api(monkeypatch):
|
||||
"""Use the actual lightweight API schema with only GPU execution replaced."""
|
||||
schema_path = Path(__file__).resolve().parents[4] / "fastvideo/api/schema.py"
|
||||
spec = importlib.util.spec_from_file_location("dreamverse_test_api_schema", schema_path)
|
||||
assert spec is not None and spec.loader is not None
|
||||
schema = importlib.util.module_from_spec(spec)
|
||||
monkeypatch.setitem(sys.modules, spec.name, schema)
|
||||
spec.loader.exec_module(schema)
|
||||
package = ModuleType("fastvideo")
|
||||
package.__path__ = []
|
||||
package.VideoGenerator = SimpleNamespace(from_config=Mock())
|
||||
monkeypatch.setitem(sys.modules, "fastvideo", package)
|
||||
monkeypatch.setitem(sys.modules, "fastvideo.api", schema)
|
||||
schema.MiniMaxH3Reference = lambda **kwargs: SimpleNamespace(**kwargs)
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.empty_cache", lambda: None)
|
||||
return package.VideoGenerator.from_config
|
||||
|
||||
|
||||
class RecordingGenerator:
|
||||
def __init__(self):
|
||||
self.requests = []
|
||||
self.images = []
|
||||
self.closed = False
|
||||
|
||||
def shutdown(self):
|
||||
self.closed = True
|
||||
|
||||
def generate(self, request):
|
||||
self.requests.append(request)
|
||||
self.images.append(tuple(None if image is None else np.asarray(image).copy()
|
||||
for image in (request.inputs.pil_image, request.inputs.last_image)))
|
||||
return SimpleNamespace(
|
||||
frames=[np.full((2, 3, 3), 7, dtype=np.uint8), np.full((2, 3, 3), 29, dtype=np.uint8)],
|
||||
audio=np.zeros((2, 16), dtype=np.float32),
|
||||
audio_sample_rate=44100,
|
||||
generation_time=0.1,
|
||||
)
|
||||
|
||||
|
||||
def prepared_backend(monkeypatch):
|
||||
backend = MiniMaxH3GenerationBackend(0)
|
||||
backend.model_config = dict(MODEL_REGISTRY["full-h3"])
|
||||
backend.generator = RecordingGenerator()
|
||||
monkeypatch.setattr(backend, "_gpu_mem", lambda: "fake executor")
|
||||
return backend
|
||||
|
||||
|
||||
def test_ipc_preserves_immutable_ordered_references():
|
||||
inputs = GenerationInputs("ref2va", (
|
||||
GenerationAsset("second", "video", "/assets/second.mp4", "reference"),
|
||||
GenerationAsset("first", "image", "/assets/first.png", "reference"),
|
||||
))
|
||||
payload = UserStepPayload("follow the references", 1, None, True, inputs)
|
||||
restored = pickle.loads(pickle.dumps(payload))
|
||||
assert restored == payload
|
||||
assert [asset.asset_id for asset in restored.generation_inputs.references] == ["second", "first"]
|
||||
|
||||
|
||||
def test_worker_passes_conditioning_to_selected_backend():
|
||||
inputs = GenerationInputs("t2va")
|
||||
worker = VideoGenerationWorker(0)
|
||||
worker.backend = Mock()
|
||||
worker.generate_step("prompt", 1, None, True, inputs)
|
||||
worker.backend.generate_step.assert_called_once_with("prompt", 1, None, True, generation_inputs=inputs)
|
||||
|
||||
|
||||
def test_full_h3_uses_full_weights_without_preview_lora(monkeypatch, fastvideo_api):
|
||||
backend = prepared_backend(monkeypatch)
|
||||
old_generator = backend.generator
|
||||
fastvideo_api.return_value = RecordingGenerator()
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
|
||||
backend.initialize(MODEL_REGISTRY["full-h3"])
|
||||
config = fastvideo_api.call_args.args[0]
|
||||
assert old_generator.closed
|
||||
assert config.pipeline.components.lora_path is None
|
||||
assert config.pipeline.components.override_pipeline_cls_name is None
|
||||
assert config.engine.use_fsdp_inference
|
||||
assert config.engine.num_gpus == 4
|
||||
assert not config.pipeline.experimental["inference_torch_compile"]
|
||||
|
||||
|
||||
def test_fl2va_maps_endpoints_only_on_initial_segment(monkeypatch, fastvideo_api, tmp_path):
|
||||
first = tmp_path / "first.png"
|
||||
last = tmp_path / "last.png"
|
||||
Image.new("RGB", (3, 2), (10, 20, 30)).save(first)
|
||||
Image.new("RGB", (3, 2), (40, 50, 60)).save(last)
|
||||
inputs = GenerationInputs("fl2va", (
|
||||
GenerationAsset("first", "image", str(first), "first_frame"),
|
||||
GenerationAsset("last", "image", str(last), "last_frame"),
|
||||
))
|
||||
backend = prepared_backend(monkeypatch)
|
||||
first_result = backend.generate_step("first", 1, None, True, inputs)
|
||||
later_result = backend.generate_step("later", 2, None, False, inputs)
|
||||
assert backend.generator.images[0][0][0, 0].tolist() == [10, 20, 30]
|
||||
assert backend.generator.images[0][1][0, 0].tolist() == [40, 50, 60]
|
||||
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
|
||||
assert backend.generator.images[1][1] is None
|
||||
assert first_result.head_trim_frames == 0
|
||||
assert later_result.head_trim_frames == 1
|
||||
assert backend.generator.requests[0].sampling.num_inference_steps == 50
|
||||
|
||||
|
||||
def test_ref2va_switches_pipeline_and_preserves_reference_order(monkeypatch, fastvideo_api):
|
||||
inputs = GenerationInputs("ref2va", (
|
||||
GenerationAsset("video", "video", "/assets/reference.mp4", "reference"),
|
||||
GenerationAsset("audio", "audio", "/assets/reference.wav", "reference"),
|
||||
GenerationAsset("image", "image", "/assets/reference.png", "reference"),
|
||||
))
|
||||
backend = prepared_backend(monkeypatch)
|
||||
base_generator = backend.generator
|
||||
reference_generator = RecordingGenerator()
|
||||
|
||||
def load(config):
|
||||
assert base_generator.closed, "Old executor must release memory before loading reference weights"
|
||||
assert config.pipeline.components.override_pipeline_cls_name == "MiniMaxH3Ref2VAModularPipeline"
|
||||
assert config.pipeline.workload_type == "i2v"
|
||||
assert config.pipeline.components.lora_path is None
|
||||
return reference_generator
|
||||
|
||||
fastvideo_api.side_effect = load
|
||||
backend.generate_step("first", 1, None, True, inputs)
|
||||
result = backend.generate_step("second", 2, None, False, inputs)
|
||||
assert fastvideo_api.call_count == 1
|
||||
for request in reference_generator.requests:
|
||||
assert [(reference.media_type, reference.source) for reference in request.inputs.references] == [
|
||||
("video", "/assets/reference.mp4"), ("audio", "/assets/reference.wav"), ("image", "/assets/reference.png")
|
||||
]
|
||||
assert request.inputs.pil_image is None
|
||||
assert request.inputs.last_image is None
|
||||
assert result.head_trim_frames == result.head_trim_audio_frames == 0
|
||||
assert backend.continuation_image is None
|
||||
|
||||
fastvideo_api.side_effect = None
|
||||
fastvideo_api.return_value = RecordingGenerator()
|
||||
backend.generate_step("new project", 1, None, True, GenerationInputs("t2va"))
|
||||
assert reference_generator.closed
|
||||
config = fastvideo_api.call_args.args[0]
|
||||
assert config.pipeline.components.override_pipeline_cls_name is None
|
||||
assert backend.pipeline_mode == "base"
|
||||
|
||||
|
||||
def test_ref2va_pipeline_switch_failure_drops_unloaded_executor(monkeypatch, fastvideo_api):
|
||||
backend = prepared_backend(monkeypatch)
|
||||
old_generator = backend.generator
|
||||
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
|
||||
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
|
||||
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
|
||||
assert old_generator.closed
|
||||
assert backend.generator is None
|
||||
|
||||
|
||||
def test_failed_pipeline_switch_reloads_on_the_next_step(monkeypatch, fastvideo_api):
|
||||
"""A failed base<->ref2va switch must not strand the slot for later steps."""
|
||||
backend = prepared_backend(monkeypatch)
|
||||
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
|
||||
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
|
||||
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
|
||||
|
||||
fastvideo_api.side_effect = None
|
||||
fastvideo_api.return_value = RecordingGenerator()
|
||||
backend.generate_step("retry", 1, None, True, GenerationInputs("ref2va"))
|
||||
|
||||
assert fastvideo_api.call_count == 2
|
||||
assert backend.pipeline_mode == "ref2va"
|
||||
assert backend.generator is not None
|
||||
|
||||
|
||||
def test_mode_cannot_switch_mid_project(monkeypatch, fastvideo_api):
|
||||
backend = prepared_backend(monkeypatch)
|
||||
with pytest.raises(ValueError, match="middle of a project"):
|
||||
backend.generate_step("prompt", 2, None, False, GenerationInputs("ref2va"))
|
||||
fastvideo_api.assert_not_called()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", ["fl2va", "ref2va"])
|
||||
def test_preview_rejects_unsupported_generation_modes(monkeypatch, fastvideo_api, mode):
|
||||
backend = prepared_backend(monkeypatch)
|
||||
backend.model_config = dict(MODEL_REGISTRY["fast-h3"])
|
||||
with pytest.raises(ValueError, match="full-h3"):
|
||||
backend.generate_step("prompt", 1, None, True, GenerationInputs(mode))
|
||||
assert backend.generator.requests == []
|
||||
|
||||
|
||||
def test_legacy_h3_continuation_is_preserved(monkeypatch, fastvideo_api):
|
||||
backend = prepared_backend(monkeypatch)
|
||||
backend.generate_step("first", 1, None, False)
|
||||
result = backend.generate_step("second", 2, None, False)
|
||||
assert backend.generator.requests[0].inputs.pil_image is None
|
||||
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
|
||||
assert result.head_trim_frames == 1
|
||||
@@ -1,278 +0,0 @@
|
||||
"""Session-mode validation and IPC handoff without a GPU worker process."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import importlib.util
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from types import ModuleType, SimpleNamespace
|
||||
from unittest.mock import Mock
|
||||
|
||||
import pytest
|
||||
|
||||
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
|
||||
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def controller_module(monkeypatch):
|
||||
gpu_pool = ModuleType("dreamverse.gpu_pool")
|
||||
gpu_pool.GPUSlot = object
|
||||
monkeypatch.setitem(sys.modules, "dreamverse.gpu_pool", gpu_pool)
|
||||
path = Path(__file__).resolve().parents[1] / "session/controller.py"
|
||||
spec = importlib.util.spec_from_file_location("dreamverse_test_session_controller", path)
|
||||
assert spec is not None and spec.loader is not None
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
monkeypatch.setattr(module, "ACTIVE_MODEL_ID", "full-h3")
|
||||
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
|
||||
monkeypatch.setattr(module, "release_generation_inputs", Mock())
|
||||
return module
|
||||
|
||||
|
||||
class Socket:
|
||||
def __init__(self):
|
||||
self.incoming = asyncio.Queue()
|
||||
self.outgoing = asyncio.Queue()
|
||||
self.messages = []
|
||||
self.closed = False
|
||||
|
||||
async def accept(self):
|
||||
pass
|
||||
|
||||
async def receive_json(self):
|
||||
return await self.incoming.get()
|
||||
|
||||
async def send_json(self, payload):
|
||||
self.messages.append(payload)
|
||||
await self.outgoing.put(payload)
|
||||
|
||||
async def send_bytes(self, payload):
|
||||
pass
|
||||
|
||||
async def close(self, **kwargs):
|
||||
self.closed = True
|
||||
|
||||
async def wait_for(self, kind):
|
||||
while True:
|
||||
payload = await asyncio.wait_for(self.outgoing.get(), 3)
|
||||
if payload["type"] == kind:
|
||||
return payload
|
||||
|
||||
|
||||
class Slot:
|
||||
def __init__(self):
|
||||
self.shared_stream_buffer = None
|
||||
self.queue = asyncio.Queue()
|
||||
self.calls = []
|
||||
|
||||
async def join_user(self, *args, **kwargs):
|
||||
pass
|
||||
|
||||
def register_stream_queue(self, client_id):
|
||||
return self.queue
|
||||
|
||||
def unregister_stream_queue(self, client_id):
|
||||
pass
|
||||
|
||||
async def user_step(self, client_id, **kwargs):
|
||||
self.calls.append(kwargs)
|
||||
segment_idx = kwargs["segment_idx"]
|
||||
await self.queue.put(MediaInit(client_id, segment_idx, "test", "video/mp4", False))
|
||||
await self.queue.put(MediaChunk(client_id, segment_idx, "test", chunk=b"test"))
|
||||
await self.queue.put(MediaComplete(client_id, segment_idx, "test", 1))
|
||||
return {"e2e_latency_ms": 1.0}
|
||||
|
||||
|
||||
class Pool:
|
||||
def __init__(self):
|
||||
self.slot = Slot()
|
||||
self.acquire_count = 0
|
||||
|
||||
def get_status(self):
|
||||
return {"queue_size": 0, "available_gpus": 1, "total_gpus": 1}
|
||||
|
||||
async def acquire(self, *args):
|
||||
self.acquire_count += 1
|
||||
return 0, self.slot
|
||||
|
||||
async def release(self, *args):
|
||||
pass
|
||||
|
||||
|
||||
def start_controller(module, socket, pool):
|
||||
enhancer = SimpleNamespace(
|
||||
resolve_rewrite_model=lambda value: "test-model",
|
||||
resolve_rewrite_system_prompt=lambda value: "test-system",
|
||||
resolve_rewrite_temperature=lambda value: 1.0,
|
||||
)
|
||||
controller = module.SessionController(socket, pool, enhancer, None, None)
|
||||
return asyncio.create_task(controller.run())
|
||||
|
||||
|
||||
def test_invalid_initial_mode_does_not_acquire_gpu(controller_module):
|
||||
async def scenario():
|
||||
socket, pool = Socket(), Pool()
|
||||
await socket.incoming.put({"type": "session_init_v2", "generation_mode": "unknown"})
|
||||
await asyncio.wait_for(start_controller(controller_module, socket, pool), 3)
|
||||
error = next(message for message in socket.messages if message["type"] == "error")
|
||||
assert error["error_code"] == "invalid_generation_input"
|
||||
assert pool.acquire_count == 0
|
||||
assert socket.closed
|
||||
controller_module.pin_generation_inputs.assert_not_called()
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
|
||||
def test_new_project_replaces_conditioning_and_passes_it_to_gpu(controller_module, monkeypatch):
|
||||
first = GenerationInputs("t2va")
|
||||
second = GenerationInputs("fl2va", (GenerationAsset("first", "image", "/assets/first.png", "first_frame"),))
|
||||
monkeypatch.setattr(controller_module, "resolve_generation_inputs", Mock(side_effect=[first, second]))
|
||||
|
||||
async def scenario():
|
||||
socket, pool = Socket(), Pool()
|
||||
await socket.incoming.put({
|
||||
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["first prompt"],
|
||||
"enhancement_enabled": False,
|
||||
})
|
||||
task = start_controller(controller_module, socket, pool)
|
||||
try:
|
||||
await socket.wait_for("media_segment_complete")
|
||||
await socket.incoming.put({"type": "end_project_keep_session"})
|
||||
await socket.wait_for("project_idle")
|
||||
assert first in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
|
||||
await socket.incoming.put({
|
||||
"type": "project_init_v1", "generation_mode": "fl2va", "curated_prompts": ["second prompt"],
|
||||
"enhancement_enabled": False,
|
||||
})
|
||||
await socket.wait_for("media_segment_complete")
|
||||
assert [call["generation_inputs"] for call in pool.slot.calls] == [first, second]
|
||||
assert pool.slot.calls[1]["segment_idx"] == 1
|
||||
assert pool.slot.calls[1]["reset_conditioning"]
|
||||
await socket.incoming.put({"type": "leave"})
|
||||
await asyncio.wait_for(task, 3)
|
||||
finally:
|
||||
if not task.done():
|
||||
task.cancel()
|
||||
await asyncio.gather(task, return_exceptions=True)
|
||||
assert [call.args[0] for call in controller_module.pin_generation_inputs.call_args_list] == [first, second]
|
||||
assert second in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
|
||||
@pytest.mark.parametrize("injection", [
|
||||
{"initial_image": {"data_url": "not allowed"}},
|
||||
{"generation_mode": "ref2va"},
|
||||
{"conditioning_assets": []},
|
||||
])
|
||||
def test_simple_generate_cannot_replace_locked_inputs(controller_module, injection):
|
||||
async def scenario():
|
||||
socket, pool = Socket(), Pool()
|
||||
await socket.incoming.put({
|
||||
"type": "session_init_v2", "generation_mode": "t2va", "single_clip_mode": True,
|
||||
"enhancement_enabled": False,
|
||||
})
|
||||
task = start_controller(controller_module, socket, pool)
|
||||
try:
|
||||
await socket.wait_for("gpu_assigned")
|
||||
await socket.incoming.put({"type": "simple_generate", "prompt": "prompt", **injection})
|
||||
error = await socket.wait_for("error")
|
||||
assert error["error_code"] == "invalid_generation_input"
|
||||
assert "project" in error["message"]
|
||||
assert pool.slot.calls == []
|
||||
await socket.incoming.put({"type": "leave"})
|
||||
await asyncio.wait_for(task, 3)
|
||||
finally:
|
||||
if not task.done():
|
||||
task.cancel()
|
||||
await asyncio.gather(task, return_exceptions=True)
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
|
||||
def test_disconnect_waits_for_worker_before_releasing_assets(controller_module, monkeypatch):
|
||||
async def scenario():
|
||||
socket, pool = Socket(), Pool()
|
||||
worker_started = asyncio.Event()
|
||||
worker_finished = asyncio.Event()
|
||||
proceed = asyncio.Event()
|
||||
|
||||
async def slow_step(client_id, **kwargs):
|
||||
worker_started.set()
|
||||
await proceed.wait()
|
||||
worker_finished.set()
|
||||
return {"e2e_latency_ms": 1.0}
|
||||
|
||||
pool.slot.user_step = slow_step
|
||||
await socket.incoming.put({
|
||||
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["prompt"],
|
||||
"enhancement_enabled": False,
|
||||
})
|
||||
task = start_controller(controller_module, socket, pool)
|
||||
try:
|
||||
await asyncio.wait_for(worker_started.wait(), 3)
|
||||
await socket.incoming.put({"type": "leave"})
|
||||
await asyncio.sleep(0.07)
|
||||
assert not task.done()
|
||||
controller_module.release_generation_inputs.assert_not_called()
|
||||
proceed.set()
|
||||
await asyncio.wait_for(task, 3)
|
||||
assert worker_finished.is_set()
|
||||
controller_module.release_generation_inputs.assert_called_once()
|
||||
finally:
|
||||
if not task.done():
|
||||
task.cancel()
|
||||
await asyncio.gather(task, return_exceptions=True)
|
||||
|
||||
asyncio.run(scenario())
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def gpu_pool_module(monkeypatch):
|
||||
streaming = ModuleType("dreamverse.av_streaming")
|
||||
for name in ("StreamChunk", "StreamComplete", "StreamEvent", "StreamInit", "generate_stream_id", "stream_fmp4"):
|
||||
setattr(streaming, name, object)
|
||||
streaming.SHARED_STREAM_BUFFER_BYTES = 1024
|
||||
streaming.USE_SHARED_STREAM_BUFFER = False
|
||||
monkeypatch.setitem(sys.modules, "dreamverse.av_streaming", streaming)
|
||||
path = Path(__file__).resolve().parents[1] / "gpu_pool.py"
|
||||
spec = importlib.util.spec_from_file_location("dreamverse_test_gpu_pool", path)
|
||||
assert spec is not None and spec.loader is not None
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
monkeypatch.setitem(sys.modules, spec.name, module)
|
||||
spec.loader.exec_module(module)
|
||||
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
|
||||
monkeypatch.setattr(module, "release_generation_inputs", Mock())
|
||||
return module
|
||||
|
||||
|
||||
def test_gpu_step_timeout_keeps_assets_pinned_until_late_worker_completion(gpu_pool_module):
|
||||
from dreamverse.worker_ipc import StepComplete
|
||||
|
||||
async def scenario():
|
||||
slot = gpu_pool_module.GPUSlot(0, "0")
|
||||
inputs = GenerationInputs("ref2va", (GenerationAsset("ref", "image", "/assets/ref.png", "reference"),))
|
||||
|
||||
async def timeout(command, timeout):
|
||||
assert command.payload.generation_inputs == inputs
|
||||
raise asyncio.TimeoutError
|
||||
|
||||
slot._send_command_tagged = timeout
|
||||
with pytest.raises(asyncio.TimeoutError):
|
||||
await slot.user_step("user", "prompt", generation_inputs=inputs)
|
||||
gpu_pool_module.pin_generation_inputs.assert_called_once_with(inputs)
|
||||
gpu_pool_module.release_generation_inputs.assert_not_called()
|
||||
|
||||
def late_response(timeout):
|
||||
slot._active = False
|
||||
return StepComplete("user", 1, {})
|
||||
|
||||
slot.response_queue = SimpleNamespace(get=late_response)
|
||||
slot._active = True
|
||||
await slot._response_reader()
|
||||
gpu_pool_module.release_generation_inputs.assert_called_once_with(inputs)
|
||||
assert slot._step_asset_inputs == {}
|
||||
|
||||
asyncio.run(scenario())
|
||||
@@ -1,18 +0,0 @@
|
||||
from dreamverse.generation_worker import _create_generation_backend
|
||||
from dreamverse.ltx2_generation import LTX2GenerationBackend
|
||||
|
||||
|
||||
def test_create_generation_backend_ltx2_module_import():
|
||||
backend = _create_generation_backend("ltx2", gpu_id=3)
|
||||
|
||||
assert isinstance(backend, LTX2GenerationBackend)
|
||||
assert backend.gpu_id == 3
|
||||
|
||||
|
||||
def test_create_generation_backend_cosmos25_dfd_module_import():
|
||||
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
|
||||
|
||||
backend = _create_generation_backend("cosmos25_dfd", gpu_id=2)
|
||||
|
||||
assert isinstance(backend, Cosmos25DFDGenerationBackend)
|
||||
assert backend.gpu_id == 2
|
||||
@@ -63,14 +63,6 @@ def test_get_available_gpus_defaults_to_first_visible_device(monkeypatch):
|
||||
assert gpu_pool.get_available_gpus() == [3]
|
||||
|
||||
|
||||
def test_get_available_gpus_defaults_to_active_model_sequence_parallel_size(monkeypatch):
|
||||
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "0,1,2,3,4")
|
||||
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
|
||||
monkeypatch.setattr(gpu_pool, "DREAMVERSE_SP_SIZE", 4)
|
||||
|
||||
assert gpu_pool.get_available_gpus() == [0, 1, 2, 3]
|
||||
|
||||
|
||||
def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
|
||||
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
|
||||
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "zero")
|
||||
@@ -79,23 +71,6 @@ def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
|
||||
gpu_pool.get_available_gpus()
|
||||
|
||||
|
||||
def test_join_user_failed_reload_marks_model_uninitialized(monkeypatch):
|
||||
"""A failed model reload forces the next join to reload a model."""
|
||||
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
|
||||
slot.current_model_id = "fast-ltx2"
|
||||
|
||||
async def fake_send_command(command, timeout):
|
||||
del command, timeout
|
||||
return gpu_pool.WorkerError(user_id="__reload__", message="load failed")
|
||||
|
||||
monkeypatch.setattr(slot, "_send_command", fake_send_command)
|
||||
|
||||
with pytest.raises(RuntimeError, match="Model reload failed"):
|
||||
asyncio.run(slot.join_user("client-id", model_id="fast-h3"))
|
||||
|
||||
assert slot.current_model_id is None
|
||||
|
||||
|
||||
def test_send_command_raises_on_worker_death():
|
||||
"""A worker that consumes a command and exits without replying must
|
||||
surface as RuntimeError via sentinel detection, not after the long
|
||||
@@ -117,9 +92,9 @@ def test_send_command_raises_on_worker_death():
|
||||
ready = resp_q.get(timeout=30.0)
|
||||
assert ready == "READY"
|
||||
|
||||
async def runner() -> None:
|
||||
async def runner():
|
||||
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
|
||||
slot.process = proc # type: ignore[assignment]
|
||||
slot.process = proc
|
||||
slot.command_queue = cmd_q
|
||||
slot.response_queue = resp_q
|
||||
|
||||
|
||||
@@ -1,6 +1,4 @@
|
||||
import ast
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
ALLOWED_PREFIXES = (
|
||||
@@ -19,11 +17,11 @@ FORBIDDEN_PREFIXES = (
|
||||
)
|
||||
ALLOWED_INTERNAL_IMPORTS = {
|
||||
(
|
||||
"ltx2_generation.py",
|
||||
"video_generation.py",
|
||||
"fastvideo.models.audio.ltx2_audio_processing",
|
||||
),
|
||||
(
|
||||
"ltx2_generation.py",
|
||||
"video_generation.py",
|
||||
"fastvideo.models.loader.component_loader",
|
||||
),
|
||||
}
|
||||
@@ -50,39 +48,3 @@ def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
|
||||
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
|
||||
|
||||
assert bad == [], f"Forbidden internal imports: {bad}"
|
||||
|
||||
|
||||
def test_h3_reference_public_export_is_lazy_and_preserves_type_identity() -> None:
|
||||
"""Only explicit reference usage should load H3's optional GPU dependencies."""
|
||||
repo_root = Path(__file__).resolve().parents[4]
|
||||
# Isolate the import graph: keep the real public API implementation/schema,
|
||||
# substituting only the unrelated legacy sampling module and heavy H3 leaf.
|
||||
script = r'''
|
||||
import importlib
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from types import ModuleType
|
||||
|
||||
root = Path(sys.argv[1])
|
||||
fastvideo = ModuleType("fastvideo")
|
||||
fastvideo.__path__ = [str(root / "fastvideo")]
|
||||
sys.modules["fastvideo"] = fastvideo
|
||||
sampling = ModuleType("fastvideo.api.sampling_param")
|
||||
sampling.SamplingParam = type("SamplingParam", (), {})
|
||||
sys.modules[sampling.__name__] = sampling
|
||||
|
||||
api = importlib.import_module("fastvideo.api")
|
||||
assert "MiniMaxH3Reference" in api.__all__
|
||||
assert "MiniMaxH3Reference" not in vars(api)
|
||||
assert not any(name.startswith("fastvideo.pipelines") for name in sys.modules)
|
||||
|
||||
internal = ModuleType("fastvideo.pipelines.basic.minimax_h3.reference")
|
||||
internal.MiniMaxH3Reference = type("MiniMaxH3Reference", (), {})
|
||||
sys.modules[internal.__name__] = internal
|
||||
from fastvideo.api import MiniMaxH3Reference
|
||||
assert MiniMaxH3Reference is internal.MiniMaxH3Reference
|
||||
assert api.MiniMaxH3Reference is internal.MiniMaxH3Reference
|
||||
assert not hasattr(api, "UnknownReference")
|
||||
'''
|
||||
result = subprocess.run([sys.executable, "-c", script, str(repo_root)], capture_output=True, text=True, timeout=30)
|
||||
assert result.returncode == 0, result.stderr
|
||||
|
||||
@@ -1,247 +0,0 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from types import SimpleNamespace
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
import dreamverse.generation_worker as generation_worker
|
||||
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
|
||||
|
||||
|
||||
FASTH3_MODEL_CONFIG = {
|
||||
"name": "FastH3",
|
||||
"generation_backend": "minimax_h3",
|
||||
"default_sp_size": 4,
|
||||
"model_path": "MiniMaxAI/MiniMax-H3",
|
||||
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
|
||||
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
|
||||
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
|
||||
"height": 768,
|
||||
"width": 1344,
|
||||
"num_frames": 124,
|
||||
"num_inference_steps": 5,
|
||||
"seed": 1000,
|
||||
}
|
||||
|
||||
|
||||
class _RecordingGenerator:
|
||||
"""Record typed requests and return small synchronized media fixtures."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.requests: list[Any] = []
|
||||
self.conditioning_pixels: list[np.ndarray | None] = []
|
||||
|
||||
def generate(self, request):
|
||||
"""Capture the request and return two tiny video frames with audio."""
|
||||
self.requests.append(request)
|
||||
conditioning_image = request.inputs.pil_image
|
||||
self.conditioning_pixels.append(
|
||||
None if conditioning_image is None else np.asarray(conditioning_image).copy())
|
||||
frames = [
|
||||
np.full((2, 3, 3), 10, dtype=np.uint8),
|
||||
np.full((2, 3, 3), 20, dtype=np.uint8),
|
||||
]
|
||||
return SimpleNamespace(
|
||||
frames=frames,
|
||||
audio=np.zeros((2, 16), dtype=np.float32),
|
||||
audio_sample_rate=44100,
|
||||
generation_time=0.25,
|
||||
)
|
||||
|
||||
|
||||
def test_initialize_builds_vsa_datafree_fasth3_generator(monkeypatch):
|
||||
"""Initialization translates the DreamVerse profile into typed FastVideo config."""
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
captured = {}
|
||||
fake_generator = SimpleNamespace(shutdown=lambda: None)
|
||||
|
||||
def fake_from_config(config):
|
||||
captured["config"] = config
|
||||
return fake_generator
|
||||
|
||||
def fake_download(**kwargs):
|
||||
captured["download"] = kwargs
|
||||
return f"/models/{kwargs['filename']}"
|
||||
|
||||
monkeypatch.setattr("huggingface_hub.hf_hub_download", fake_download)
|
||||
monkeypatch.setattr(VideoGenerator, "from_config", fake_from_config)
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
|
||||
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
|
||||
monkeypatch.setenv("FASTVIDEO_FA4", "0")
|
||||
monkeypatch.setenv("FASTVIDEO_MINIMAX_H3_FUSIONS", "0")
|
||||
monkeypatch.setenv("FASTVIDEO_VSA_SM100A", "1")
|
||||
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
|
||||
|
||||
backend = MiniMaxH3GenerationBackend(gpu_id=0)
|
||||
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
|
||||
backend.initialize(FASTH3_MODEL_CONFIG)
|
||||
|
||||
config = captured["config"]
|
||||
assert captured["download"] == {
|
||||
"repo_id": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
|
||||
"filename": "vsa-datafree/adapter_model.safetensors",
|
||||
}
|
||||
assert config.model_path == "MiniMaxAI/MiniMax-H3"
|
||||
assert config.pipeline.components.lora_path.endswith("vsa-datafree/adapter_model.safetensors")
|
||||
assert config.pipeline.components.lora_strength == 1.0
|
||||
assert config.pipeline.experimental == {
|
||||
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
|
||||
"inference_torch_compile": False,
|
||||
"vae_parallel_decode": True,
|
||||
"vae_parallel_decode_strategy": "gather",
|
||||
"VSA_sparsity": 0.9,
|
||||
"VSA_tile_size": 64,
|
||||
}
|
||||
assert config.engine.num_gpus == 4
|
||||
assert config.engine.parallelism.tp_size == 1
|
||||
assert config.engine.parallelism.sp_size == 4
|
||||
assert config.engine.offload.dit is False
|
||||
assert config.engine.offload.dit_layerwise is False
|
||||
assert config.engine.offload.text_encoder is True
|
||||
assert config.engine.offload.vae is True
|
||||
assert config.engine.compile.vae_enabled is True
|
||||
assert config.engine.use_fsdp_inference is False
|
||||
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "VIDEO_SPARSE_ATTN_H3"
|
||||
assert os.environ["FASTVIDEO_FA4"] == "1"
|
||||
assert os.environ["FASTVIDEO_MINIMAX_H3_FUSIONS"] == "all"
|
||||
assert os.environ["FASTVIDEO_VSA_SM100A"] == "0"
|
||||
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
|
||||
|
||||
|
||||
def test_initialize_selects_declared_generation_backend(monkeypatch):
|
||||
"""The GPU worker constructs the backend that the active model profile declares."""
|
||||
from unittest.mock import Mock
|
||||
|
||||
selected_backend = Mock()
|
||||
monkeypatch.setattr(
|
||||
generation_worker,
|
||||
"_create_generation_backend",
|
||||
lambda backend_name, gpu_id: selected_backend,
|
||||
)
|
||||
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
|
||||
|
||||
worker.initialize(FASTH3_MODEL_CONFIG)
|
||||
|
||||
assert worker.backend_name == "minimax_h3"
|
||||
assert worker.backend is selected_backend
|
||||
selected_backend.initialize.assert_called_once_with(FASTH3_MODEL_CONFIG)
|
||||
|
||||
|
||||
def test_initialize_failure_clears_backend_ownership(monkeypatch):
|
||||
"""A failed family change leaves the GPU worker explicitly uninitialized."""
|
||||
ltx_backend = SimpleNamespace(initialize=lambda config: None, shutdown=lambda: None)
|
||||
|
||||
def fail_initialize(config):
|
||||
del config
|
||||
raise RuntimeError("load failed")
|
||||
|
||||
fasth3_backend = SimpleNamespace(
|
||||
initialize=fail_initialize,
|
||||
shutdown=lambda: None,
|
||||
)
|
||||
backends = {
|
||||
"ltx2": ltx_backend,
|
||||
"minimax_h3": fasth3_backend,
|
||||
}
|
||||
monkeypatch.setattr(
|
||||
generation_worker,
|
||||
"_create_generation_backend",
|
||||
lambda backend_name, gpu_id: backends[backend_name],
|
||||
)
|
||||
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
|
||||
worker.initialize({"generation_backend": "ltx2"})
|
||||
|
||||
with pytest.raises(RuntimeError, match="load failed"):
|
||||
worker.initialize(FASTH3_MODEL_CONFIG)
|
||||
|
||||
assert worker.backend is None
|
||||
assert worker.backend_name is None
|
||||
assert worker.model_config == {"generation_backend": "ltx2"}
|
||||
|
||||
|
||||
def test_generate_step_uses_last_frame_for_continuation(monkeypatch):
|
||||
"""A later segment receives the prior segment's last decoded frame."""
|
||||
backend = MiniMaxH3GenerationBackend(gpu_id=0)
|
||||
backend.model_config = dict(FASTH3_MODEL_CONFIG)
|
||||
backend.generator = _RecordingGenerator()
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
|
||||
|
||||
first_result = backend.generate_step(
|
||||
"first prompt",
|
||||
segment_idx=1,
|
||||
image_path=None,
|
||||
reset_conditioning=True,
|
||||
)
|
||||
second_result = backend.generate_step(
|
||||
"second prompt",
|
||||
segment_idx=2,
|
||||
image_path=None,
|
||||
reset_conditioning=False,
|
||||
)
|
||||
|
||||
first_request = backend.generator.requests[0]
|
||||
assert first_request.inputs.pil_image is None
|
||||
assert first_request.negative_prompt == ""
|
||||
assert first_request.sampling.height == 768
|
||||
assert first_request.sampling.width == 1344
|
||||
assert first_request.sampling.num_frames == 124
|
||||
assert first_request.sampling.num_inference_steps == 5
|
||||
assert first_request.sampling.fps == 24
|
||||
assert first_request.sampling.guidance_scale == 1.0
|
||||
assert first_request.sampling.batch_cfg is False
|
||||
assert first_request.sampling.seed == 1000
|
||||
assert first_request.output.save_video is False
|
||||
assert first_request.output.return_frames is True
|
||||
assert backend.generator.conditioning_pixels[1].tolist() == np.full((2, 3, 3), 20).tolist()
|
||||
assert first_result.head_trim_frames == 0
|
||||
assert first_result.head_trim_audio_frames == 0
|
||||
assert second_result.head_trim_frames == 1
|
||||
assert second_result.head_trim_audio_frames == 1
|
||||
assert second_result.audio_sample_rate == 44100
|
||||
|
||||
|
||||
def test_generate_step_reset_uses_text_to_video_path(monkeypatch):
|
||||
"""Resetting continuation produces an unconditioned text-to-video request."""
|
||||
backend = MiniMaxH3GenerationBackend(gpu_id=0)
|
||||
backend.model_config = dict(FASTH3_MODEL_CONFIG)
|
||||
backend.generator = _RecordingGenerator()
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
|
||||
|
||||
backend.generate_step("first prompt", 1, None, True)
|
||||
reset_result = backend.generate_step("reset prompt", 2, None, True)
|
||||
|
||||
assert backend.generator.requests[-1].inputs.pil_image is None
|
||||
assert reset_result.head_trim_frames == 0
|
||||
assert reset_result.head_trim_audio_frames == 0
|
||||
|
||||
|
||||
def test_generate_step_missing_continuation_frame(monkeypatch):
|
||||
"""A later segment fails when no reset or retained frame defines its input."""
|
||||
backend = MiniMaxH3GenerationBackend(gpu_id=0)
|
||||
backend.model_config = dict(FASTH3_MODEL_CONFIG)
|
||||
backend.generator = _RecordingGenerator()
|
||||
|
||||
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
|
||||
backend.generate_step("later prompt", 2, None, False)
|
||||
|
||||
assert backend.generator.requests == []
|
||||
|
||||
|
||||
def test_warmup_exercises_text_and_first_frame_paths(monkeypatch):
|
||||
"""Warmup covers both request shapes used by a DreamVerse session."""
|
||||
backend = MiniMaxH3GenerationBackend(gpu_id=0)
|
||||
backend.model_config = dict(FASTH3_MODEL_CONFIG)
|
||||
backend.generator = _RecordingGenerator()
|
||||
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
|
||||
|
||||
timings = backend.warmup("warmup prompt")
|
||||
|
||||
assert backend.generator.conditioning_pixels[0] is None
|
||||
assert backend.generator.conditioning_pixels[1] is not None
|
||||
assert backend.continuation_image is None
|
||||
assert "warmup_text_to_video_ms" in timings
|
||||
assert "warmup_first_frame_to_video_ms" in timings
|
||||
@@ -331,11 +331,11 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
|
||||
]
|
||||
|
||||
|
||||
def test_enhance_prompt_uses_groq_when_it_returns_first():
|
||||
def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
|
||||
enhancer = _build_staged_enhancer(
|
||||
cerebras_payload=_chat_payload_with_content('{"prompt":"Cerebras prompt"}'),
|
||||
groq_payload=_chat_payload_with_content('{"prompt":"Groq prompt"}'),
|
||||
cerebras_delay_s=0.08,
|
||||
cerebras_delay_s=0.01,
|
||||
groq_delay_s=0.01,
|
||||
)
|
||||
|
||||
@@ -346,12 +346,12 @@ def test_enhance_prompt_uses_groq_when_it_returns_first():
|
||||
|
||||
assert result.fallback_used is False
|
||||
assert result.error is None
|
||||
assert result.provider == "groq"
|
||||
assert result.provider == "cerebras"
|
||||
assert result.model == "gpt-test"
|
||||
assert result.prompt == "Groq prompt"
|
||||
assert result.prompt == "Cerebras prompt"
|
||||
assert enhancer.get_provider_success_counts() == {
|
||||
"cerebras": 0,
|
||||
"groq": 1,
|
||||
"cerebras": 1,
|
||||
"groq": 0,
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -105,7 +105,6 @@ class _FakeSlot:
|
||||
segment_idx: int,
|
||||
reset_conditioning: bool,
|
||||
image_path: str | None = None,
|
||||
generation_inputs=None,
|
||||
):
|
||||
self.calls.append({
|
||||
"client_id": client_id,
|
||||
|
||||
+23
-15
@@ -1,9 +1,9 @@
|
||||
"""LTX-2 model lifecycle and continuation conditioning.
|
||||
"""LTX2 model lifecycle and continuation conditioning.
|
||||
|
||||
Runs inside a GPU worker subprocess. Owns the model, the audio
|
||||
encoder, and the per-session continuation state carried across
|
||||
segments. Callers must set ``os.environ["CUDA_VISIBLE_DEVICES"]``
|
||||
before constructing ``LTX2GenerationBackend`` — all ``fastvideo.*``
|
||||
before constructing ``VideoGenerationWorker`` — all ``fastvideo.*``
|
||||
imports are deferred to method bodies so nothing touches CUDA at
|
||||
module import time.
|
||||
"""
|
||||
@@ -14,6 +14,9 @@ import gc
|
||||
import os
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
@@ -32,8 +35,6 @@ from dreamverse.config import (
|
||||
DREAMVERSE_LORA_STACK,
|
||||
_resolve_lora_spec,
|
||||
)
|
||||
from dreamverse.generation_contracts import StepResult
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
# Multi-frame decoded continuation defaults from
|
||||
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
|
||||
@@ -79,6 +80,22 @@ def _reset_lora_registry(worker) -> dict:
|
||||
return {"status": "lora_registry_reset"}
|
||||
|
||||
|
||||
@dataclass
|
||||
class StepResult:
|
||||
"""Output of one generation step.
|
||||
|
||||
``head_trim_frames`` / ``head_trim_audio_frames`` are derived here
|
||||
so downstream AV streaming never needs to import conditioning
|
||||
constants.
|
||||
"""
|
||||
frames: list
|
||||
audio: Any
|
||||
audio_sample_rate: int | None
|
||||
timings: dict
|
||||
head_trim_frames: int
|
||||
head_trim_audio_frames: int
|
||||
|
||||
|
||||
class ContinuationState:
|
||||
"""Per-session video + audio conditioning carried across segments."""
|
||||
|
||||
@@ -185,7 +202,7 @@ class ContinuationState:
|
||||
self.audio_latents = latents.detach().clone().cpu()
|
||||
|
||||
|
||||
class LTX2GenerationBackend:
|
||||
class VideoGenerationWorker:
|
||||
"""Single-GPU LTX2 generator with continuation state.
|
||||
|
||||
Caller must set ``os.environ["CUDA_VISIBLE_DEVICES"]`` before
|
||||
@@ -291,13 +308,7 @@ class LTX2GenerationBackend:
|
||||
dynamic=False,
|
||||
),
|
||||
use_fsdp_inference=False,
|
||||
# The bundled LTX2 model enables a refinement LoRA during the
|
||||
# first request. NVFP4 otherwise purges the dense weights that
|
||||
# FastVideo's LoRA merge path requires.
|
||||
quantization=QuantizationConfig(
|
||||
transformer_quant="NVFP4",
|
||||
transformer_retain_original_weights=True,
|
||||
),
|
||||
quantization=QuantizationConfig(transformer_quant="NVFP4"),
|
||||
),
|
||||
pipeline=PipelineSelection(
|
||||
components=components,
|
||||
@@ -461,11 +472,8 @@ class LTX2GenerationBackend:
|
||||
segment_idx: int,
|
||||
image_path: str | None,
|
||||
reset_conditioning: bool,
|
||||
generation_inputs: GenerationInputs | None = None,
|
||||
) -> StepResult:
|
||||
"""Execute one generation step; snapshot state for the next segment."""
|
||||
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
|
||||
raise ValueError("LTX supports text generation only through the generation mode API.")
|
||||
timings: dict = {}
|
||||
|
||||
prompt = self._inject_style_trigger(prompt)
|
||||
@@ -15,8 +15,6 @@ from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
from dreamverse.generation_inputs import GenerationInputs
|
||||
|
||||
# ---- User-scoped events (carry user_id) ------------------------------------
|
||||
|
||||
|
||||
@@ -149,7 +147,6 @@ class UserStepPayload:
|
||||
segment_idx: int
|
||||
image_path: str | None
|
||||
reset_conditioning: bool
|
||||
generation_inputs: GenerationInputs | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
|
||||
@@ -70,7 +70,7 @@
|
||||
<mxCell id="dispatcher" value="command dispatcher

gpu_worker_process() branches on
CommandType; asserts payload type

INIT / WARMUP / RELOAD_MODEL
USER_JOIN / USER_STEP / USER_LEAVE
SHUTDOWN" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe6cc;strokeColor=#d79b00;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
|
||||
<mxGeometry x="120" y="1120" width="240" height="120" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()
ltx2_generation.py:380

reads + updates ContinuationState,
calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
|
||||
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()
video_generation.py:380

reads + updates ContinuationState,
calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
|
||||
<mxGeometry x="460" y="1120" width="240" height="120" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="stream_av" value="stream_fmp4()
av_streaming.py:121

trims overlap, pipes to ffmpeg,
publishes StreamInit / StreamChunk /
StreamComplete via callback" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#b1d8d7;strokeColor=#23445d;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
|
||||
@@ -79,13 +79,13 @@
|
||||
<mxCell id="Ot8BU52QTIb4EhyRSe7I-2" value="" style="edgeStyle=none;html=1;" parent="1" source="generator" target="Ot8BU52QTIb4EhyRSe7I-1" edge="1">
|
||||
<mxGeometry relative="1" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="generator" value="VideoGenerator (fastvideo)

LTX2 DiT + refine upsampler
FP4 quant, torch.compile

owned by VideoGenerationWorker
ltx2_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
|
||||
<mxCell id="generator" value="VideoGenerator (fastvideo)

LTX2 DiT + refine upsampler
FP4 quant, torch.compile

owned by VideoGenerationWorker
video_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
|
||||
<mxGeometry x="460" y="1300" width="240" height="100" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="ffmpeg" value="ffmpeg subprocess

libx264 / *_nvenc
fragmented mp4" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=11;" parent="1" vertex="1">
|
||||
<mxGeometry x="800" y="1300" width="260" height="100" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="caches" value="ContinuationState
ltx2_generation.py:89

• video_images: list[PIL.Image]
• audio_latents: torch.Tensor (CPU)

carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
|
||||
<mxCell id="caches" value="ContinuationState
video_generation.py:89

• video_images: list[PIL.Image]
• audio_latents: torch.Tensor (CPU)

carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
|
||||
<mxGeometry x="120" y="1300" width="240" height="100" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="e_cp" value="acquire" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="client" target="pool" edge="1">
|
||||
@@ -250,7 +250,7 @@
|
||||
<mxPoint x="690" y="880"/>
|
||||
</Array>
|
||||
</mxCell>
|
||||
<mxCell id="legend" value="Legend

■ blue client / external
■ green main-process pool/slot
 (methods — italic label)
■ yellow containers (routing state)
■ red IPC primitives (mp.Queue, mp.RawArray)

Worker subprocess modules:
■ orange gpu_pool.py (dispatcher)
■ lavender ltx2_generation.py
■ teal av_streaming.py
■ gray worker_ipc.py (shared types)

Flow:
 client → pool → slot
 → _send_command(_tagged) → command_queue
 → dispatcher → generate_step()
 → stream_fmp4() → ffmpeg
 → shared_buf + response_queue
 → _response_reader → futures / stream_queues
 → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
|
||||
<mxCell id="legend" value="Legend

■ blue client / external
■ green main-process pool/slot
 (methods — italic label)
■ yellow containers (routing state)
■ red IPC primitives (mp.Queue, mp.RawArray)

Worker subprocess modules:
■ orange gpu_pool.py (dispatcher)
■ lavender video_generation.py
■ teal av_streaming.py
■ gray worker_ipc.py (shared types)

Flow:
 client → pool → slot
 → _send_command(_tagged) → command_queue
 → dispatcher → generate_step()
 → stream_fmp4() → ffmpeg
 → shared_buf + response_queue
 → _response_reader → futures / stream_queues
 → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
|
||||
<mxGeometry x="39" y="-200" width="270" height="380" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="Ot8BU52QTIb4EhyRSe7I-1" value="FastVideo video_generator" style="whiteSpace=wrap;html=1;fontSize=11;fillColor=#e1d5e7;strokeColor=#9673a6;rounded=1;" parent="1" vertex="1">
|
||||
@@ -389,10 +389,10 @@
|
||||
<mxCell id="cw2" value="from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo.models.dits.ltx2 import DEFAULT_LTX2_AUDIO_*

** Dreamverse reaches into fastvideo internals here **" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxGeometry x="675" y="695" width="550" height="60" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="cw3" value="on Command(INIT):
 VideoGenerationWorker.initialize() (ltx2_generation.py:247)
 maybe_download_model(model_id)
 VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)
 load audio VAE, resolve refine upsampler
 resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxCell id="cw3" value="on Command(INIT):
 VideoGenerationWorker.initialize() (video_generation.py:247)
 maybe_download_model(model_id)
 VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)
 load audio VAE, resolve refine upsampler
 resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxGeometry x="675" y="765" width="550" height="95" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:
 VideoGenerationWorker.warmup(payload.prompt) (ltx2_generation.py:518)
 two synthetic segments prime caches + torch.compile
 resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:
 VideoGenerationWorker.warmup(payload.prompt) (video_generation.py:518)
 two synthetic segments prime caches + torch.compile
 resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxGeometry x="675" y="870" width="550" height="55" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="cw5" value="enter main worker loop → waits for JOIN_USER / USER_STEP / LEAVE" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
|
||||
@@ -534,7 +534,7 @@
|
||||
<mxPoint x="1040" y="1610" as="targetPoint"/>
|
||||
</mxGeometry>
|
||||
</mxCell>
|
||||
<mxCell id="dm11a" value="10a. worker runs:
VideoGenerationWorker.generate_step()
 (ltx2_generation.py:380)
 → generator.generate_video()
 → updates ContinuationState
then stream_fmp4() (av_streaming.py:121)
 → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxCell id="dm11a" value="10a. worker runs:
VideoGenerationWorker.generate_step()
 (video_generation.py:380)
 → generator.generate_video()
 → updates ContinuationState
then stream_fmp4() (av_streaming.py:121)
 → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
|
||||
<mxGeometry x="955" y="1640" width="180" height="70" as="geometry"/>
|
||||
</mxCell>
|
||||
<mxCell id="dm11" value="10b. resp_q.put(MediaInit / MediaChunk / MediaComplete / StepComplete)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
|
||||
|
||||
File diff suppressed because one or more lines are too long
|
Before Width: | Height: | Size: 85 KiB After Width: | Height: | Size: 85 KiB |
@@ -1,107 +0,0 @@
|
||||
# Dreamverse on Slurm
|
||||
|
||||
Run Full H3 inside a one-node, four-GPU allocation. The maintained H3 examples
|
||||
default to four GPUs; this is a starting configuration, not a measured minimum.
|
||||
The full checkpoint supports T2VA, FL2VA, and Ref2VA. The FastH3 Preview profile
|
||||
is a separate T2VA configuration.
|
||||
|
||||
`launch_backend.sh` checks that it is inside an `srun` step, preserves
|
||||
`CUDA_VISIBLE_DEVICES`, and replaces itself with the backend process. It does
|
||||
not allocate GPUs, kill existing processes, or source a personal credentials
|
||||
file. The local `dreamverse-deploy` helper is not suitable for a shared Slurm
|
||||
cluster because it kills processes by physical GPU and port.
|
||||
|
||||
## Prepare and allocate
|
||||
|
||||
Keep the checkout, weights, outputs, and logs on storage visible to the compute
|
||||
node. Source installation is documented in the [GPU guide](../../../../docs/getting_started/installation/gpu.md).
|
||||
On ARM64 GB200 use CUDA 13, a matching PyTorch build, and kernels built for
|
||||
`sm_100`; the DGX Spark `sm_121` kernel image is not the GB200 image.
|
||||
|
||||
The repository's image workflow publishes an ARM64 GB200 variant under
|
||||
`ghcr.io/hao-ai-lab/fastvideo/fastvideo-dev:py3.12-cuda13.0.0-sm100-latest`.
|
||||
Resolve that tag to a digest for reproducible runs. If your compute nodes use
|
||||
Pyxis/Enroot, pass the approved image or a prepared SquashFS file to
|
||||
`srun --container-image`, with explicit mounts for your checkout and model cache.
|
||||
The Dreamverse-specific Docker images are currently AMD64-only.
|
||||
|
||||
For the Slinky customer partition, a bounded allocation is:
|
||||
|
||||
```bash
|
||||
salloc --account=customer --qos=normal --partition=hpc-rack-1 \
|
||||
--nodes=1 --ntasks=1 --cpus-per-task=72 --gres=gpu:nvidia_gb200:4 \
|
||||
--mem=800G --time=02:00:00 --job-name=dreamverse
|
||||
srun --ntasks=1 --pty bash
|
||||
```
|
||||
|
||||
Wait for Slurm to grant the allocation before entering the compute step. A
|
||||
successful SSH login does not grant GPU resources. Inspect pending capacity
|
||||
with `squeue -u "$USER" --start`; do not attach to another user's job.
|
||||
|
||||
The checkpoint includes duplicate release layouts. Download the diffusers
|
||||
components needed by both base and reference pipelines, rather than the whole
|
||||
repository (about 210 GB versus about 498 GB at revision
|
||||
`42ed227ee7df40d41602854ae760620d6eb651fe`):
|
||||
|
||||
```bash
|
||||
hf download MiniMaxAI/MiniMax-H3 \
|
||||
--revision 42ed227ee7df40d41602854ae760620d6eb651fe \
|
||||
--include model_index.json --include modular_model_index.json \
|
||||
--include 'audio_scheduler/*' --include 'audio_vae/*' \
|
||||
--include 'processor/*' --include 'scheduler/*' \
|
||||
--include 'text_encoder/*' --include 'tokenizer/*' \
|
||||
--include 'transformer/*' --include 'transformer_ref/*' --include 'vae/*' \
|
||||
--local-dir /path/to/models/MiniMax-H3
|
||||
```
|
||||
|
||||
The GPU environment needs `fastvideo[dreamverse]`, the Dreamverse workspace
|
||||
package, and FFmpeg with H.264/AAC encoders. In a prepared FastVideo image,
|
||||
install the checked-out code and its Dreamverse dependencies in that image's
|
||||
Python environment. Keep its matching CUDA/PyTorch/kernel stack intact.
|
||||
|
||||
## Start and connect
|
||||
|
||||
From the checked-out repository inside the allocated step:
|
||||
|
||||
```bash
|
||||
export DREAMVERSE_PYTHON=/path/to/environment/bin/python
|
||||
export DREAMVERSE_MODEL_PATH=/path/to/models/MiniMax-H3
|
||||
export FASTVIDEO_DREAMVERSE_HOME=/path/to/persistent/dreamverse-state
|
||||
bash apps/dreamverse/scripts/slurm/launch_backend.sh
|
||||
```
|
||||
|
||||
The default backend binds port 8009 on the private compute node. Connect through
|
||||
the login node from your laptop, replacing `COMPUTE_NODE_IP` with the allocated
|
||||
node's `NodeAddr` from `scontrol show node`:
|
||||
|
||||
```bash
|
||||
ssh -N -L 8009:COMPUTE_NODE_IP:8009 USER@LOGIN_NODE
|
||||
```
|
||||
|
||||
In another laptop terminal, run the frontend from your local checkout:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
BACKEND_HOST=127.0.0.1 BACKEND_PORT=8009 npm run dev
|
||||
```
|
||||
|
||||
Open `http://localhost:5299`. `/healthz` reports the server process; `/readyz`
|
||||
reports model readiness. Full H3 loads and generates more slowly than the
|
||||
Preview adapter. Keep prompt enhancement disabled in the UI unless the
|
||||
runtime has the selected provider's credentials.
|
||||
|
||||
## Verify and stop
|
||||
|
||||
Check all three modes with small, valid user-owned assets. Capture the selected
|
||||
mode and assets, WebSocket errors or completion events, the generated video and
|
||||
audio, and GPU memory usage. Also verify actionable validation errors and
|
||||
backward compatibility with clients that omit `generation_mode`.
|
||||
|
||||
Use the frontend Playwright instructions in the
|
||||
[Dreamverse development guide](../../../../docs/contributing/dreamverse-development.md)
|
||||
against the forwarded backend. A mock-server demo validates UI and protocol
|
||||
behavior; it is not evidence of GPU generation.
|
||||
|
||||
Stop the backend with Ctrl-C, exit the compute step, and release your allocation.
|
||||
For a detached allocation, use `scancel YOUR_JOB_ID`. Cancel a pending demo job
|
||||
when it is no longer needed; do not leave an unattended reservation queued.
|
||||
@@ -1,49 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Run inside an existing Slurm step. Slurm owns the GPU visibility and lifetime.
|
||||
set -euo pipefail
|
||||
|
||||
if [[ -z "${SLURM_JOB_ID:-}" || -z "${SLURM_STEP_ID:-}" ]]; then
|
||||
echo "Run this launcher inside an allocated Slurm step (srun), not on the login node." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
repo_root="$(cd -- "${script_dir}/../../../.." && pwd)"
|
||||
python_bin="${DREAMVERSE_PYTHON:-${repo_root}/.venv/bin/python}"
|
||||
if [[ ! -x "${python_bin}" ]]; then
|
||||
echo "Set DREAMVERSE_PYTHON to a Python environment with fastvideo[dreamverse] installed." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
export DREAMVERSE_MODEL_ID="${DREAMVERSE_MODEL_ID:-full-h3}"
|
||||
export DREAMVERSE_SP_SIZE="${DREAMVERSE_SP_SIZE:-4}"
|
||||
export FASTVIDEO_GPU_COUNT="${FASTVIDEO_GPU_COUNT:-${DREAMVERSE_SP_SIZE}}"
|
||||
export FASTVIDEO_ENABLE_STARTUP_WARMUP="${FASTVIDEO_ENABLE_STARTUP_WARMUP:-0}"
|
||||
export ENABLE_TORCH_COMPILE="${ENABLE_TORCH_COMPILE:-0}"
|
||||
export STREAM_MODE="${STREAM_MODE:-av_fmp4}"
|
||||
export PYTHONPATH="${repo_root}/apps/dreamverse:${repo_root}${PYTHONPATH:+:${PYTHONPATH}}"
|
||||
export PYTHONUNBUFFERED=1
|
||||
|
||||
"${python_bin}" - <<'PY'
|
||||
import os
|
||||
import shutil
|
||||
|
||||
import torch
|
||||
|
||||
expected = int(os.environ["DREAMVERSE_SP_SIZE"])
|
||||
visible = torch.cuda.device_count()
|
||||
if expected < 1 or visible < expected:
|
||||
raise SystemExit(f"The Slurm step exposes {visible} GPUs; DREAMVERSE_SP_SIZE requires {expected}.")
|
||||
ffmpeg = os.environ.get("FASTVIDEO_FFMPEG_BIN", "ffmpeg")
|
||||
if not shutil.which(ffmpeg):
|
||||
raise SystemExit("FFmpeg is missing; install it in the compute environment or set FASTVIDEO_FFMPEG_BIN.")
|
||||
print(f"Slurm job {os.environ['SLURM_JOB_ID']}: {visible} visible GPUs; using {expected} per worker")
|
||||
for index in range(expected):
|
||||
properties = torch.cuda.get_device_properties(index)
|
||||
print(f" GPU {index}: {properties.name}, {properties.total_memory / 2**30:.1f} GiB")
|
||||
PY
|
||||
|
||||
cd "${repo_root}"
|
||||
exec "${python_bin}" -m dreamverse.server_entry \
|
||||
--host "${DREAMVERSE_BIND_HOST:-0.0.0.0}" \
|
||||
--port "${DREAMVERSE_BACKEND_PORT:-8009}" "$@"
|
||||
@@ -1,126 +0,0 @@
|
||||
import { execFileSync } from "node:child_process";
|
||||
import { readFile } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { test, expect } from "@playwright/test";
|
||||
|
||||
const imagePath = path.resolve("public/k2.png");
|
||||
const framePrompt = "A paper fox walks through a sunlit forest, gentle birdsong.";
|
||||
|
||||
function makeAudio(sampleRate = 8000, seconds = 1): Buffer {
|
||||
const sampleCount = sampleRate * seconds;
|
||||
const bytes = Buffer.alloc(44 + sampleCount * 2);
|
||||
bytes.write("RIFF", 0); bytes.writeUInt32LE(bytes.length - 8, 4); bytes.write("WAVEfmt ", 8);
|
||||
bytes.writeUInt32LE(16, 16); bytes.writeUInt16LE(1, 20); bytes.writeUInt16LE(1, 22);
|
||||
bytes.writeUInt32LE(sampleRate, 24); bytes.writeUInt32LE(sampleRate * 2, 28);
|
||||
bytes.writeUInt16LE(2, 32); bytes.writeUInt16LE(16, 34); bytes.write("data", 36);
|
||||
bytes.writeUInt32LE(sampleCount * 2, 40);
|
||||
for (let i = 0; i < sampleCount; i++) bytes.writeInt16LE(Math.round(Math.sin(i * 440 * 2 * Math.PI / sampleRate) * 1000), 44 + i * 2);
|
||||
return bytes;
|
||||
}
|
||||
|
||||
test.describe("generation modes through the mock runtime", () => {
|
||||
for (const mode of ["t2va", "fl2va", "ref2va"] as const) {
|
||||
test(`${mode} sends validated assets and plays a clearly labeled sample`, async ({ page, request }, testInfo) => {
|
||||
const response = await request.get("/generation-capabilities");
|
||||
const capabilities = response.ok() ? await response.json() : {};
|
||||
test.skip(capabilities.mock !== true, "This test uses the CPU mock runtime; it must not silently allocate a real GPU.");
|
||||
const sent: Record<string, any>[] = [];
|
||||
const received: Record<string, any>[] = [];
|
||||
page.on("websocket", (socket) => {
|
||||
socket.on("framesent", ({ payload }) => { if (typeof payload === "string") { try { sent.push(JSON.parse(payload)); } catch {} } });
|
||||
socket.on("framereceived", ({ payload }) => { if (typeof payload === "string") { try { received.push(JSON.parse(payload)); } catch {} } });
|
||||
});
|
||||
await page.goto("/");
|
||||
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
|
||||
const modeSelect = page.getByRole("combobox", { name: "Generation mode" });
|
||||
const modeLabel = mode === "ref2va" ? "Ref2VA" : mode.toUpperCase();
|
||||
await modeSelect.click();
|
||||
await page.getByRole("option", { name: modeLabel, exact: true }).click();
|
||||
await expect(modeSelect).toHaveText(modeLabel);
|
||||
await page.getByLabel("Continuation prompt").fill(framePrompt);
|
||||
const uploadedIds: string[] = [];
|
||||
page.on("response", async (uploadResponse) => {
|
||||
if (uploadResponse.request().method() === "POST" && uploadResponse.url().endsWith("/assets") && uploadResponse.ok()) {
|
||||
const asset = await uploadResponse.json().catch(() => null);
|
||||
if (asset?.asset_id) uploadedIds.push(asset.asset_id);
|
||||
}
|
||||
});
|
||||
try {
|
||||
if (mode === "fl2va") {
|
||||
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
|
||||
await page.locator('input[type="file"]').setInputFiles([
|
||||
{ name: "first-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
|
||||
{ name: "last-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
|
||||
]);
|
||||
await expect(page.getByRole("option", { name: "first-frame.png", exact: true }).first()).toBeAttached();
|
||||
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "first-frame.png" });
|
||||
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeEnabled();
|
||||
await page.getByRole("combobox", { name: "Last frame", exact: true }).selectOption({ label: "last-frame.png" });
|
||||
}
|
||||
if (mode === "ref2va") {
|
||||
const video = execFileSync(process.env.FASTVIDEO_FFMPEG_BIN || "ffmpeg", ["-v", "error", "-f", "lavfi", "-i", "color=c=royalblue:s=64x64:r=8", "-t", "1", "-c:v", "libx264", "-pix_fmt", "yuv420p", "-movflags", "frag_keyframe+empty_moov", "-f", "mp4", "pipe:1"]);
|
||||
await page.locator('input[type="file"]').setInputFiles([
|
||||
{ name: "subject.png", mimeType: "image/png", buffer: await readFile(imagePath) },
|
||||
{ name: "motion.mp4", mimeType: "video/mp4", buffer: video },
|
||||
{ name: "sound.wav", mimeType: "audio/wav", buffer: makeAudio() },
|
||||
]);
|
||||
await expect(page.getByRole("button", { name: "Add sound.wav as reference" })).toBeEnabled();
|
||||
await page.getByRole("button", { name: "Add sound.wav as reference" }).click();
|
||||
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
|
||||
await page.getByRole("button", { name: "Add subject.png as reference" }).click();
|
||||
await page.getByRole("button", { name: "Add motion.mp4 as reference" }).click();
|
||||
await page.getByRole("button", { name: "Move sound.wav down" }).click();
|
||||
const names = await page.getByRole("list", { name: "Ordered references" }).locator("li p.font-medium").allTextContents();
|
||||
expect(names).toEqual(["subject.png", "sound.wav", "motion.mp4"]);
|
||||
}
|
||||
await page.screenshot({ path: testInfo.outputPath(`${mode}-inputs.png`), fullPage: true });
|
||||
await page.getByRole("button", { name: "Generate", exact: true }).click();
|
||||
await expect.poll(() => sent.find((item) => item.type === "session_init_v2")?.generation_mode).toBe(mode);
|
||||
const init = sent.find((item) => item.type === "session_init_v2")!;
|
||||
expect(init.conditioning_assets.map((item: any) => item.role)).toEqual(mode === "t2va" ? [] : mode === "fl2va" ? ["first_frame", "last_frame"] : ["reference", "reference", "reference"]);
|
||||
if (mode === "ref2va") expect(init.conditioning_assets.map((item: any) => item.asset_id)).toEqual([uploadedIds[0], uploadedIds[2], uploadedIds[1]]);
|
||||
await expect.poll(() => received.find((item) => item.type === "gpu_assigned")?.generation_mode).toBe(mode);
|
||||
await expect.poll(() => received.some((item) => item.type === "media_segment_complete")).toBe(true);
|
||||
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
|
||||
await expect(modeSelect).toHaveCount(0);
|
||||
await expect.poll(async () => page.locator("video:visible").first().evaluate((element: HTMLVideoElement) => element.readyState)).toBeGreaterThanOrEqual(2);
|
||||
await page.screenshot({ path: testInfo.outputPath(`${mode}-playback.png`), fullPage: true });
|
||||
if (mode === "ref2va") {
|
||||
await page.getByRole("button", { name: "Toggle sidebar" }).click();
|
||||
await page.getByRole("button", { name: "New project", exact: true }).click();
|
||||
await expect(modeSelect).toHaveText("T2VA");
|
||||
await modeSelect.click();
|
||||
await page.getByRole("option", { name: "FL2VA", exact: true }).click();
|
||||
await expect(modeSelect).toHaveText("FL2VA");
|
||||
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "subject.png" });
|
||||
await expect(page.getByRole("combobox", { name: "Last frame", exact: true })).toHaveValue("");
|
||||
await page.getByLabel("Continuation prompt").fill("The paper fox explores a new scene.");
|
||||
await page.getByRole("button", { name: "Generate", exact: true }).click();
|
||||
await expect.poll(() => sent.find((item) => item.type === "project_init_v1")?.generation_mode).toBe("fl2va");
|
||||
const secondProject = sent.find((item) => item.type === "project_init_v1")!;
|
||||
expect(secondProject.conditioning_assets).toEqual([{ asset_id: uploadedIds[0], role: "first_frame" }]);
|
||||
expect(sent.filter((item) => item.type === "session_init_v2")).toHaveLength(1);
|
||||
await expect.poll(() => received.filter((item) => item.type === "media_segment_complete").length).toBeGreaterThan(1);
|
||||
}
|
||||
} finally {
|
||||
await page.close();
|
||||
for (const id of uploadedIds) await request.delete(`/assets/${id}`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
test("proxies a media upload larger than Next's default 10 MiB body limit", async ({ request }) => {
|
||||
const response = await request.get("/generation-capabilities");
|
||||
const capabilities = response.ok() ? await response.json() : {};
|
||||
test.skip(capabilities.mock !== true, "Requires the local mock runtime.");
|
||||
const audio = makeAudio(192000, 29);
|
||||
expect(audio.length).toBeGreaterThan(10 * 1024 * 1024);
|
||||
const upload = await request.post("/assets", {
|
||||
headers: { "Content-Type": "audio/wav", "X-Asset-Name": "large-proxy-check.wav" },
|
||||
data: audio,
|
||||
});
|
||||
expect(upload.status()).toBe(201);
|
||||
const asset = await upload.json();
|
||||
try { expect(asset.size).toBe(audio.length); } finally { await request.delete(`/assets/${asset.asset_id}`); }
|
||||
});
|
||||
});
|
||||
@@ -9,8 +9,6 @@ const configDir = path.dirname(fileURLToPath(import.meta.url));
|
||||
const staticExport = process.env.NEXT_OUTPUT_EXPORT === '1';
|
||||
|
||||
const nextConfig: NextConfig = {
|
||||
// Next 15.5 name for the dev rewrite-proxy body limit; Next 16 renames it to `proxyClientMaxBodySize`.
|
||||
experimental: { middlewareClientMaxBodySize: 100 * 1024 * 1024 },
|
||||
...(staticExport ? { output: 'export' as const } : {}),
|
||||
...(staticExport ? { images: { unoptimized: true } } : {}),
|
||||
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
|
||||
@@ -40,18 +38,6 @@ const nextConfig: NextConfig = {
|
||||
source: '/router/:path*',
|
||||
destination: `${backendUrl}/router/:path*`
|
||||
},
|
||||
{
|
||||
source: '/generation-capabilities',
|
||||
destination: `${backendUrl}/generation-capabilities`,
|
||||
},
|
||||
{
|
||||
source: '/assets',
|
||||
destination: `${backendUrl}/assets`,
|
||||
},
|
||||
{
|
||||
source: '/assets/:path*',
|
||||
destination: `${backendUrl}/assets/:path*`,
|
||||
},
|
||||
{
|
||||
source: '/prompt-system-config',
|
||||
destination: `${backendUrl}/prompt-system-config`,
|
||||
|
||||
@@ -589,7 +589,6 @@ describe.skip('App websocket integration', () => {
|
||||
});
|
||||
|
||||
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
|
||||
expect(initMessage.generation_mode).toBe('t2va');
|
||||
expect(initMessage.preset_id).toBe('test_preset');
|
||||
expect(initMessage.curated_prompts).toEqual(['segment one', 'segment two']);
|
||||
expect(initMessage.enhancement_enabled).toBe(true);
|
||||
@@ -598,42 +597,6 @@ describe.skip('App websocket integration', () => {
|
||||
expect(initMessage.initial_rollout_prompt).toBe('');
|
||||
});
|
||||
|
||||
it('sends the selected generation mode and locks it after session start', async () => {
|
||||
const outbound: any[] = [];
|
||||
server.on('connection', (socket) => {
|
||||
socket.on('message', (rawMessage) => {
|
||||
outbound.push(JSON.parse(rawMessage as string));
|
||||
});
|
||||
});
|
||||
|
||||
const user = userEvent.setup();
|
||||
render(<Page />);
|
||||
|
||||
const modeSelect = await screen.findByRole('combobox', { name: 'Generation mode' });
|
||||
expect(modeSelect).toHaveTextContent('T2VA');
|
||||
|
||||
await user.click(modeSelect);
|
||||
await user.click(await screen.findByRole('option', { name: 'FL2VA' }));
|
||||
expect(modeSelect).toHaveTextContent('FL2VA');
|
||||
expect(modeSelect).toHaveAttribute(
|
||||
'title',
|
||||
'First/last frames to video + audio. Start from a first frame image. Add an optional last frame to guide the ending.',
|
||||
);
|
||||
|
||||
const generateButton = await screen.findByRole('button', { name: 'Generate' });
|
||||
await waitFor(() => expect(generateButton).toBeEnabled());
|
||||
await user.click(generateButton);
|
||||
|
||||
await waitFor(() => {
|
||||
expect(outbound.some((message) => message.type === 'session_init_v2')).toBe(true);
|
||||
});
|
||||
|
||||
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
|
||||
expect(initMessage.generation_mode).toBe('fl2va');
|
||||
expect(screen.queryByRole('combobox', { name: 'Generation mode' }))
|
||||
.not.toBeInTheDocument();
|
||||
});
|
||||
|
||||
it('starts a streaming session from a custom initial prompt without using curated prompts', async () => {
|
||||
const outbound: any[] = [];
|
||||
server.on('connection', (socket) => {
|
||||
|
||||
@@ -5,7 +5,6 @@ import { Download, Share2 } from "lucide-react";
|
||||
import DevtoolsShell from "@/components/devtools/DevtoolsShell";
|
||||
import MonitorPage from "@/components/MonitorPage";
|
||||
import ChatBar from "@/components/ChatBar";
|
||||
import AssetList from "@/components/AssetList";
|
||||
import SessionTimeoutModal from "@/components/SessionTimeoutModal";
|
||||
import Sidebar from "@/components/Sidebar";
|
||||
import Header from "@/components/Header";
|
||||
@@ -14,13 +13,10 @@ import Workspace from "@/components/Workspace";
|
||||
import { saveProject, saveProjectMetadata, listProjects, loadProjectClips, deleteProject, pruneOldProjects, type StoredProject, type StoredClip } from "@/lib/projectStorage";
|
||||
import { isInfrastructureError } from "@/lib/ws/reducer";
|
||||
import { useStore } from "@/hooks/useStore";
|
||||
import { useAssetLibrary } from "@/hooks/useAssetLibrary";
|
||||
import { useGenerationCapabilities } from "@/hooks/useGenerationCapabilities";
|
||||
import { resolveDevtoolsMode } from "@/lib/devtoolsMode";
|
||||
import { createAvPipeline, DEFAULT_AV_MIME } from "@/lib/media/avPipeline";
|
||||
import { remuxArchivedFmp4Segments } from "@/lib/media/fmp4Remux";
|
||||
import { DEFAULT_CUSTOM_PRESET_ID, parseStoryPresets, sanitizePresetId } from "@/lib/presets";
|
||||
import { DEFAULT_GENERATION_MODE, buildGenerationInitFields, validateGenerationInputs, type GenerationMode, type GenerationInitFields, type GenerationAsset } from "@/lib/generationMode";
|
||||
import {
|
||||
buildRewritePromptWindowSnapshot,
|
||||
buildRewritePromptWindowSnapshotFromPrompts,
|
||||
@@ -345,19 +341,6 @@ export default function Page() {
|
||||
const [isMobileShareCapable, setIsMobileShareCapable] = useState(false);
|
||||
const [videoMuted, setVideoMuted] = useState(true);
|
||||
const [timeoutModalOpen, setTimeoutModalOpen] = useState(false);
|
||||
const [generationMode, setGenerationMode] = useState<GenerationMode>(DEFAULT_GENERATION_MODE);
|
||||
const assetLibrary = useAssetLibrary();
|
||||
const { capabilities, capabilityNotice, refreshCapabilities } = useGenerationCapabilities();
|
||||
const joiningRef = useRef(false);
|
||||
const activeGenerationRef = useRef<{ fields: GenerationInitFields; assets: GenerationAsset[]; mock: boolean } | null>(null);
|
||||
const generationInputError = validateGenerationInputs(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets);
|
||||
const generationSupported = capabilities.modes.includes(generationMode);
|
||||
const generationInputsValid = !generationInputError && generationSupported && !assetLibrary.uploading;
|
||||
function changeGenerationMode(mode: GenerationMode) {
|
||||
if (sessionStore.get().sessionStarted || joiningRef.current || !capabilities.modes.includes(mode)) return;
|
||||
setGenerationMode(mode);
|
||||
assetLibrary.clearConditioning();
|
||||
}
|
||||
useEffect(() => {
|
||||
setIsMobileShareCapable(typeof navigator.canShare === "function" && window.matchMedia("(pointer: coarse)").matches);
|
||||
}, []);
|
||||
@@ -408,7 +391,7 @@ export default function Page() {
|
||||
|
||||
// --- Derived values ---
|
||||
|
||||
const canStartSession = generationInputsValid && !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
|
||||
const canStartSession = !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
|
||||
|
||||
const currentClipLabel = useMemo(() => {
|
||||
if ((activeClip as Record<string, any>)?.label) return (activeClip as Record<string, any>).label;
|
||||
@@ -764,11 +747,7 @@ export default function Page() {
|
||||
|
||||
function recoverFailedSessionStart(notice: string) {
|
||||
const restoredDraft = normalizeInitialPrompt(pendingInitialPromptRef.current);
|
||||
if (wsRef.current) {
|
||||
detachAndCloseWebSocket(wsRef.current);
|
||||
wsRef.current = null;
|
||||
}
|
||||
resetToLobbyState({ preserveSessionNotice: true });
|
||||
resetToLobbyState();
|
||||
clearPendingProjectPointers();
|
||||
pendingInitialPromptRef.current = "";
|
||||
sessionStore.patch({
|
||||
@@ -1723,10 +1702,6 @@ export default function Page() {
|
||||
|
||||
function resetToLobbyState({ preserveSessionNotice = false, preservePlayback = false } = {}) {
|
||||
setVideoMuted(true);
|
||||
if (!preserveSessionNotice) {
|
||||
setGenerationMode(DEFAULT_GENERATION_MODE);
|
||||
assetLibrary.clearConditioning();
|
||||
}
|
||||
clearCountdownInterval();
|
||||
pendingInitialPromptRef.current = "";
|
||||
sessionStore.patch({
|
||||
@@ -1761,8 +1736,6 @@ export default function Page() {
|
||||
|
||||
function resetToProjectLobbyState() {
|
||||
setVideoMuted(true);
|
||||
setGenerationMode(DEFAULT_GENERATION_MODE);
|
||||
assetLibrary.clearConditioning();
|
||||
pendingInitialPromptRef.current = "";
|
||||
sessionStore.patch({
|
||||
sessionStarted: false,
|
||||
@@ -1791,7 +1764,6 @@ export default function Page() {
|
||||
setSeedPrompts(segmentPrompts);
|
||||
return {
|
||||
type,
|
||||
...(activeGenerationRef.current?.fields || buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets)),
|
||||
preset_id: getInitialPresetId(),
|
||||
preset_label: getInitialPresetLabel(),
|
||||
curated_prompts: segmentPrompts,
|
||||
@@ -1840,12 +1812,6 @@ export default function Page() {
|
||||
return;
|
||||
}
|
||||
if (decoded.kind !== "json") return;
|
||||
if (decoded.data?.type === "error" && sessionStore.get().sessionStarted
|
||||
&& (!sessionStore.get().gpuAssigned || decoded.data.error_code === "invalid_generation_input")) {
|
||||
const message = typeof decoded.data.message === "string" ? decoded.data.message : "The generation inputs were rejected. Check the mode and selected assets.";
|
||||
recoverFailedSessionStart(message);
|
||||
return;
|
||||
}
|
||||
if (decoded.data?.type === "error" && isInfrastructureError(decoded.data)) {
|
||||
const message = typeof decoded.data?.message === "string" && decoded.data.message.trim()
|
||||
? decoded.data.message.trim()
|
||||
@@ -1971,15 +1937,9 @@ export default function Page() {
|
||||
}
|
||||
}
|
||||
|
||||
function beginProjectLocally({ force = false, mockRuntime = capabilities.mock === true } = {}) {
|
||||
function beginProjectLocally({ force = false } = {}) {
|
||||
if (!force && !canStartSession) return;
|
||||
if (!generationInputsValid) return false;
|
||||
if (sessionStore.get().sessionStarted || sessionStore.get().projectResetPending) return false;
|
||||
activeGenerationRef.current = {
|
||||
fields: buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets),
|
||||
assets: assetLibrary.assets.filter((asset) => assetLibrary.conditioningAssets.some((item) => item.asset_id === asset.asset_id)),
|
||||
mock: mockRuntime,
|
||||
};
|
||||
setTimeoutModalOpen(false);
|
||||
// Unmute during the user gesture so iOS Safari permits audio playback.
|
||||
setVideoMuted(false);
|
||||
@@ -2039,43 +1999,12 @@ export default function Page() {
|
||||
}
|
||||
|
||||
async function joinSession({ force = false } = {}) {
|
||||
if (joiningRef.current || sessionStore.get().sessionStarted) return;
|
||||
if (generationInputError || assetLibrary.uploading) {
|
||||
showPreSessionNotice(generationInputError || "Wait for the asset upload to finish.");
|
||||
return;
|
||||
}
|
||||
joiningRef.current = true;
|
||||
try {
|
||||
await startGenerationSession({ force });
|
||||
} finally {
|
||||
joiningRef.current = false;
|
||||
}
|
||||
}
|
||||
|
||||
async function startGenerationSession({ force = false } = {}) {
|
||||
sessionStore.patch({ sessionNotice: "" });
|
||||
streamStore.patch({ loadingAnimation: true });
|
||||
const currentCapabilities = await refreshCapabilities();
|
||||
if (!currentCapabilities.modes.includes(generationMode)) {
|
||||
streamStore.patch({ loadingAnimation: false });
|
||||
showPreSessionNotice(`${generationMode.toUpperCase()} is unavailable on this runtime. Connect a full H3 runtime or choose a supported mode.`);
|
||||
return;
|
||||
}
|
||||
const assetProblem = await assetLibrary.verifySelectedAssets();
|
||||
if (assetProblem) {
|
||||
streamStore.patch({ loadingAnimation: false });
|
||||
showPreSessionNotice(assetProblem);
|
||||
return;
|
||||
}
|
||||
if (
|
||||
wsRef.current
|
||||
&& wsRef.current.readyState === WebSocket.OPEN
|
||||
&& sessionStore.get().connected
|
||||
) {
|
||||
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
|
||||
streamStore.patch({ loadingAnimation: false });
|
||||
return;
|
||||
}
|
||||
if (!beginProjectLocally({ force })) return;
|
||||
sendProjectInitMessage();
|
||||
return;
|
||||
}
|
||||
@@ -2088,7 +2017,7 @@ export default function Page() {
|
||||
showPreSessionNotice(probe.notice);
|
||||
return;
|
||||
}
|
||||
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
|
||||
if (!beginProjectLocally({ force })) {
|
||||
streamStore.patch({ loadingAnimation: false });
|
||||
sessionStore.patch({ connecting: false });
|
||||
return;
|
||||
@@ -2115,10 +2044,6 @@ export default function Page() {
|
||||
createdAt: currentProjectCreatedAtRef.current || Date.now(),
|
||||
lastThumbnail: currentThumbnail,
|
||||
promptEvents: [...(rewriteStore.get().promptEvents as Record<string, unknown>[])],
|
||||
generationMode: activeGenerationRef.current?.fields.generation_mode || DEFAULT_GENERATION_MODE,
|
||||
conditioningAssets: activeGenerationRef.current?.fields.conditioning_assets || [],
|
||||
assets: activeGenerationRef.current?.assets || [],
|
||||
mock: activeGenerationRef.current?.mock === true,
|
||||
};
|
||||
const clips: StoredClip[] = (streamStore.get().completedClips as any[])
|
||||
.filter((clip: any) => clip?.blob instanceof Blob)
|
||||
@@ -2553,24 +2478,6 @@ export default function Page() {
|
||||
|
||||
// --- Render ---
|
||||
|
||||
const conditioningPanel = generationMode !== "t2va" && !sessionStarted && !sessionExpired ? (
|
||||
<AssetList
|
||||
mode={generationMode}
|
||||
assets={assetLibrary.assets}
|
||||
conditioning={assetLibrary.conditioningAssets}
|
||||
locked={Boolean(loadingAnimation || projectResetPending)}
|
||||
uploading={assetLibrary.uploading}
|
||||
error={assetLibrary.assetError}
|
||||
validationNotice={generationInputError}
|
||||
onUpload={assetLibrary.uploadAssets}
|
||||
onAssign={assetLibrary.assignAsset}
|
||||
onRemove={assetLibrary.removeAsset}
|
||||
onUnselect={assetLibrary.removeConditioning}
|
||||
onMove={assetLibrary.moveConditioning}
|
||||
onMissing={assetLibrary.checkAssetAvailability}
|
||||
/>
|
||||
) : null;
|
||||
|
||||
if (!runtimeReady) {
|
||||
return null;
|
||||
}
|
||||
@@ -2594,17 +2501,13 @@ export default function Page() {
|
||||
enhancementEnabled={enhancementEnabled as boolean}
|
||||
autoExtensionEnabled={autoExtensionEnabled as boolean}
|
||||
loopGenerationEnabled={loopGenerationEnabled as boolean}
|
||||
canJoinSession={canStartSession}
|
||||
canJoinSession={canJoinSession as boolean}
|
||||
canSubmitContinuation={canSubmitContinuation}
|
||||
editableMode={editableMode as boolean}
|
||||
demoMode={demoMode as boolean}
|
||||
editableCanJoin={editableCanJoin as boolean}
|
||||
curatedPromptLimit={curatedPromptLimit as number}
|
||||
maxCuratedPromptCount={maxCuratedPromptCount as number}
|
||||
generationMode={generationMode}
|
||||
supportedGenerationModes={capabilities.modes}
|
||||
conditioningPanel={conditioningPanel}
|
||||
onGenerationModeChange={changeGenerationMode}
|
||||
onPresetChange={handlePresetSelectionChange}
|
||||
onEnhancementToggle={handleEnhancementToggle}
|
||||
onCuratedPromptLimitChange={handleCuratedPromptLimitChange}
|
||||
@@ -2738,12 +2641,7 @@ export default function Page() {
|
||||
/>
|
||||
<Header timeLeft={headerTimeLeft} formatTime={formatTime} onToggleSidebar={() => setSidebarOpen((prev) => !prev)} />
|
||||
|
||||
<div className={cn(
|
||||
"relative flex flex-1 min-h-0 flex-col px-4 pb-2 sm:px-6 sm:pb-12",
|
||||
!isViewingMode && !showActiveProject && generationMode !== "t2va"
|
||||
? "justify-start overflow-y-auto pt-4"
|
||||
: "justify-center",
|
||||
)}>
|
||||
<div className="relative flex flex-1 min-h-0 flex-col justify-center px-4 pb-2 sm:px-6 sm:pb-12">
|
||||
{isViewingMode && (
|
||||
<>
|
||||
{viewingSelectedClip && (
|
||||
@@ -2782,8 +2680,6 @@ export default function Page() {
|
||||
/>
|
||||
</section>
|
||||
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
|
||||
{viewingProject?.project.mock && <p className="mb-2 text-center text-xs text-violet-600 dark:text-violet-300">Demo sample · This saved clip was not generated by an AI model.</p>}
|
||||
{viewingProject?.project.generationMode && <p className="mb-3 text-center text-xs text-muted-foreground">{viewingProject.project.generationMode.toUpperCase()} · {viewingProject.project.conditioningAssets?.length || 0} saved references. Uploaded originals may expire; your saved video remains available.</p>}
|
||||
<ChatBar sessionStarted={false} viewingReadOnly={true} onStartNewProject={handleStartNewProject} onBackFromViewing={closeViewingProject} />
|
||||
</motion.div>
|
||||
</>
|
||||
@@ -2863,7 +2759,7 @@ export default function Page() {
|
||||
</section>
|
||||
|
||||
<AnimatePresence>
|
||||
{!showActiveProject && generationMode === "t2va" && (
|
||||
{!showActiveProject && (
|
||||
<motion.div
|
||||
key="hero-tagline"
|
||||
initial={{ opacity: 0 }}
|
||||
@@ -2889,12 +2785,6 @@ export default function Page() {
|
||||
sessionExpired={sessionExpired as boolean}
|
||||
sessionNotice={sessionNotice as string}
|
||||
projectResetPending={projectResetPending as boolean}
|
||||
generationMode={generationMode}
|
||||
supportedGenerationModes={capabilities.modes}
|
||||
generationInputsValid={generationInputsValid}
|
||||
capabilityNotice={!generationSupported ? `${generationMode.toUpperCase()} is unavailable on this runtime.` : capabilityNotice}
|
||||
mockRuntime={capabilities.mock}
|
||||
conditioningPanel={conditioningPanel}
|
||||
onPresetGenerate={handlePresetGenerate}
|
||||
onContinuationInput={handleLivePromptInput}
|
||||
onContinuationKeydown={handleLivePromptKeydown}
|
||||
@@ -2902,7 +2792,6 @@ export default function Page() {
|
||||
onSubmitContinuation={submitLivePrompt}
|
||||
onLeave={leaveSession}
|
||||
onStartNewProject={handleStartNewProject}
|
||||
onGenerationModeChange={changeGenerationMode}
|
||||
onSpeechTranscript={handleLivePromptSpeechTranscript}
|
||||
onSpeechInterimChange={handleLivePromptSpeechInterim}
|
||||
/>
|
||||
|
||||
@@ -1,41 +0,0 @@
|
||||
import { render, screen } from "@testing-library/react";
|
||||
import userEvent from "@testing-library/user-event";
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import AssetList from "./AssetList";
|
||||
import type { GenerationAsset } from "@/lib/generationMode";
|
||||
|
||||
const frame: GenerationAsset = { asset_id: "first", kind: "image", name: "frame.png", mime_type: "image/png", size: 2000, url: "/assets/first" };
|
||||
const video: GenerationAsset = { asset_id: "video", kind: "video", name: "motion.mp4", mime_type: "video/mp4", size: 2000, url: "/assets/video" };
|
||||
function props() {
|
||||
return { assets: [frame, video], onUpload: vi.fn(), onAssign: vi.fn(), onRemove: vi.fn(), onUnselect: vi.fn(), onMove: vi.fn(), onMissing: vi.fn() };
|
||||
}
|
||||
|
||||
describe("Asset List", () => {
|
||||
it("uploads to the library and assigns images to endpoint roles", async () => {
|
||||
const callbacks = props();
|
||||
const user = userEvent.setup();
|
||||
render(<AssetList {...callbacks} mode="fl2va" conditioning={[]} />);
|
||||
await user.selectOptions(screen.getByRole("combobox", { name: "First frame" }), "first");
|
||||
expect(callbacks.onAssign).toHaveBeenCalledWith("first", "first_frame");
|
||||
expect(screen.getByRole("combobox", { name: "Last frame" })).toHaveValue("");
|
||||
expect(screen.queryByRole("option", { name: "motion.mp4" })).not.toBeInTheDocument();
|
||||
const file = new File(["image"], "new.png", { type: "image/png" });
|
||||
await user.upload(screen.getByLabelText("Upload assets", { selector: "input" }), file);
|
||||
expect(callbacks.onUpload).toHaveBeenCalledWith([file]);
|
||||
});
|
||||
it("exposes accessible ordering and removal controls for multimodal references", async () => {
|
||||
const callbacks = props();
|
||||
const user = userEvent.setup();
|
||||
render(<AssetList {...callbacks} mode="ref2va" conditioning={[{ asset_id: "first", role: "reference" }, { asset_id: "video", role: "reference" }]} />);
|
||||
await user.click(screen.getByRole("button", { name: "Move motion.mp4 up" }));
|
||||
expect(callbacks.onMove).toHaveBeenCalledWith(1, 0);
|
||||
await user.click(screen.getByRole("button", { name: "Unselect frame.png" }));
|
||||
expect(callbacks.onUnselect).toHaveBeenCalledWith(0);
|
||||
expect(screen.getByRole("button", { name: "Move frame.png up" })).toBeDisabled();
|
||||
});
|
||||
it("locks uploads and assignments while starting generation", () => {
|
||||
render(<AssetList {...props()} mode="fl2va" conditioning={[]} locked />);
|
||||
expect(screen.getByRole("button", { name: "Upload assets" })).toBeDisabled();
|
||||
expect(screen.getByRole("combobox", { name: "First frame" })).toBeDisabled();
|
||||
});
|
||||
});
|
||||
@@ -1,154 +0,0 @@
|
||||
"use client";
|
||||
|
||||
import { useRef, useState } from "react";
|
||||
import { ArrowDown, ArrowUp, AudioLines, Check, GripVertical, ImagePlus, Plus, Trash2, Upload, X } from "lucide-react";
|
||||
import { Button } from "@/components/ui/button";
|
||||
import { NativeSelect } from "@/components/ui/native-select";
|
||||
import { cn } from "@/lib/utils";
|
||||
import type { ConditioningAsset, ConditioningRole, GenerationAsset, GenerationMode } from "@/lib/generationMode";
|
||||
|
||||
interface AssetListProps {
|
||||
mode: GenerationMode;
|
||||
assets: GenerationAsset[];
|
||||
conditioning: ConditioningAsset[];
|
||||
locked?: boolean;
|
||||
uploading?: boolean;
|
||||
error?: string;
|
||||
validationNotice?: string | null;
|
||||
onUpload: (files: File[]) => void;
|
||||
onAssign: (assetId: string, role: ConditioningRole) => void;
|
||||
onRemove: (assetId: string) => void;
|
||||
onUnselect: (index: number) => void;
|
||||
onMove: (from: number, to: number) => void;
|
||||
onMissing: (assetId: string) => void;
|
||||
}
|
||||
|
||||
function AssetPreview({ asset, onMissing, compact = false }: {
|
||||
asset: GenerationAsset;
|
||||
onMissing: (assetId: string) => void;
|
||||
compact?: boolean;
|
||||
}) {
|
||||
const [previewFailed, setPreviewFailed] = useState(false);
|
||||
function previewError() {
|
||||
setPreviewFailed(true);
|
||||
onMissing(asset.asset_id);
|
||||
}
|
||||
const className = cn("h-full w-full object-cover", asset.missing && "opacity-25");
|
||||
if (asset.missing) return <span className="p-2 text-center text-[10px] text-muted-foreground">Upload again</span>;
|
||||
if (previewFailed) return <span className="p-2 text-center text-[10px] text-muted-foreground">Preview unavailable</span>;
|
||||
if (asset.kind === "image") {
|
||||
return <img src={asset.url} alt={asset.name} className={className} onError={previewError} />;
|
||||
}
|
||||
if (asset.kind === "video") {
|
||||
return <video src={asset.url} aria-label={`Preview ${asset.name}`} className={className} muted playsInline controls={!compact} preload="metadata" onError={previewError} />;
|
||||
}
|
||||
return (
|
||||
<div className="flex h-full w-full flex-col items-center justify-center gap-2 bg-violet-500/10 p-2 text-violet-500">
|
||||
<AudioLines className="size-6" />
|
||||
{!compact && <audio src={asset.url} aria-label={`Preview ${asset.name}`} controls preload="metadata" className="h-6 w-full min-w-0" onError={previewError} />}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
/** A reusable library/picker. The parent asset store owns uploads and selection. */
|
||||
export default function AssetList({
|
||||
mode, assets, conditioning, locked = false, uploading = false, error = "", validationNotice,
|
||||
onUpload, onAssign, onRemove, onUnselect, onMove, onMissing,
|
||||
}: AssetListProps) {
|
||||
const inputRef = useRef<HTMLInputElement>(null);
|
||||
const [libraryOpen, setLibraryOpen] = useState(true);
|
||||
const [dragIndex, setDragIndex] = useState<number | null>(null);
|
||||
const disabled = locked || uploading;
|
||||
const imageAssets = assets.filter((asset) => asset.kind === "image");
|
||||
|
||||
return (
|
||||
<section aria-label="Asset List" className="overflow-hidden rounded-2xl border border-input bg-card/70 shadow-sm backdrop-blur-sm">
|
||||
<div className="flex items-center justify-between gap-3 px-4 py-3">
|
||||
<div>
|
||||
<h2 className="text-xs font-semibold tracking-wide">{mode === "fl2va" ? "Frame guidance" : "Reference sequence"}</h2>
|
||||
<p className="mt-0.5 text-[11px] text-muted-foreground">{locked ? "Inputs are locked for this project." : mode === "fl2va" ? "Choose your opening image and, optionally, the ending." : "Arrange references in the order you want the model to read them."}</p>
|
||||
</div>
|
||||
<Button type="button" variant="outline" size="sm" disabled={disabled} onClick={() => inputRef.current?.click()} className="shrink-0 gap-1.5 rounded-full text-xs">
|
||||
<Upload className="size-3.5" />{uploading ? "Uploading…" : "Upload assets"}
|
||||
</Button>
|
||||
<input ref={inputRef} type="file" aria-label="Upload assets" className="sr-only" multiple accept={mode === "fl2va" ? "image/*" : "image/*,video/*,audio/*"} disabled={disabled} onChange={(event) => {
|
||||
const files = Array.from(event.target.files || []);
|
||||
if (files.length) onUpload(files);
|
||||
event.target.value = "";
|
||||
}} />
|
||||
</div>
|
||||
|
||||
<div className="max-h-[min(42vh,350px)] overflow-y-auto px-4 pb-3">
|
||||
{mode === "fl2va" ? (
|
||||
<div className="grid grid-cols-2 gap-3">
|
||||
{(["first_frame", "last_frame"] as const).map((role) => {
|
||||
const label = role === "first_frame" ? "First frame" : "Last frame";
|
||||
const assetId = conditioning.find((item) => item.role === role)?.asset_id || "";
|
||||
const asset = assets.find((item) => item.asset_id === assetId);
|
||||
return (
|
||||
<div key={role} className="overflow-hidden rounded-xl border border-input bg-background/40 p-2">
|
||||
<div className="flex h-20 items-center justify-center overflow-hidden rounded-lg bg-muted/60 sm:h-24">
|
||||
{asset ? <AssetPreview key={asset.asset_id} asset={asset} onMissing={onMissing} /> : <ImagePlus className="size-6 text-muted-foreground/45" />}
|
||||
</div>
|
||||
<label htmlFor={`asset-${role}`} className="mb-1 mt-2 block text-[11px] font-medium">{label} <span className="font-normal text-muted-foreground">{role === "first_frame" ? "· required" : "· optional"}</span></label>
|
||||
<NativeSelect id={`asset-${role}`} aria-label={label} value={assetId} disabled={disabled} className="h-8 text-xs" onChange={(event) => onAssign(event.target.value, role)}>
|
||||
<option value="">{imageAssets.length ? "Choose an image" : "Upload an image first"}</option>
|
||||
{imageAssets.map((item) => <option key={item.asset_id} value={item.asset_id} disabled={item.missing}>{item.name}{item.missing ? " (upload again)" : ""}</option>)}
|
||||
</NativeSelect>
|
||||
</div>
|
||||
);
|
||||
})}
|
||||
</div>
|
||||
) : (
|
||||
<>
|
||||
{conditioning.length ? (
|
||||
<ol aria-label="Ordered references" className="flex flex-col gap-2">
|
||||
{conditioning.map((item, index) => {
|
||||
const asset = assets.find((entry) => entry.asset_id === item.asset_id);
|
||||
if (!asset) return null;
|
||||
return (
|
||||
<li key={`${item.asset_id}-${index}`} draggable={!disabled} onDragStart={() => setDragIndex(index)} onDragEnd={() => setDragIndex(null)} onDragOver={(event) => { if (!disabled && dragIndex !== null) event.preventDefault(); }} onDrop={(event) => { event.preventDefault(); if (!disabled && dragIndex !== null) onMove(dragIndex, index); setDragIndex(null); }} className={cn("flex items-center gap-2 rounded-xl border border-input bg-background/40 p-2", dragIndex === index && "opacity-50")}>
|
||||
<GripVertical className="hidden size-3.5 shrink-0 text-muted-foreground/50 sm:block" aria-hidden />
|
||||
<span className="w-4 text-center text-[11px] font-medium text-muted-foreground">{index + 1}</span>
|
||||
<div className="flex size-10 shrink-0 items-center justify-center overflow-hidden rounded-md bg-muted"><AssetPreview asset={asset} onMissing={onMissing} compact /></div>
|
||||
<div className="min-w-0 flex-1"><p className="truncate text-xs font-medium">{asset.name}</p><p className="text-[10px] capitalize text-muted-foreground">{asset.kind}{asset.missing ? " · unavailable" : ""}</p></div>
|
||||
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Move ${asset.name} up`} disabled={disabled || index === 0} onClick={() => onMove(index, index - 1)}><ArrowUp className="size-3.5" /></Button>
|
||||
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Move ${asset.name} down`} disabled={disabled || index === conditioning.length - 1} onClick={() => onMove(index, index + 1)}><ArrowDown className="size-3.5" /></Button>
|
||||
<Button type="button" variant="ghost" size="icon-sm" aria-label={`Unselect ${asset.name}`} disabled={disabled} onClick={() => onUnselect(index)}><X className="size-3.5" /></Button>
|
||||
</li>
|
||||
);
|
||||
})}
|
||||
</ol>
|
||||
) : (
|
||||
<div className="flex items-center gap-3 rounded-xl border border-dashed border-input px-4 py-4 text-muted-foreground"><ImagePlus className="size-6 shrink-0 opacity-50" /><p className="text-xs">Add images, video, or audio from your asset library.<br /><span className="text-[11px] opacity-75">At least one image or video is required.</span></p></div>
|
||||
)}
|
||||
<p className="mt-2 text-[10px] text-muted-foreground">{conditioning.length}/12 selected · up to 9 images, 3 videos, 3 audio clips</p>
|
||||
</>
|
||||
)}
|
||||
|
||||
{assets.length > 0 && !locked && (
|
||||
<div className="mt-3 border-t border-border/60 pt-2">
|
||||
<button type="button" className="flex w-full items-center justify-between py-1 text-[11px] font-medium text-muted-foreground" aria-expanded={libraryOpen} onClick={() => setLibraryOpen(!libraryOpen)}><span>Asset library · {assets.length}</span><span>{libraryOpen ? "Hide" : "Show"}</span></button>
|
||||
{libraryOpen && <div className="mt-2 grid grid-cols-2 gap-2 sm:grid-cols-3">
|
||||
{assets.map((asset) => {
|
||||
const selected = conditioning.some((item) => item.asset_id === asset.asset_id);
|
||||
return (
|
||||
<div key={asset.asset_id} className={cn("overflow-hidden rounded-lg border bg-background/40", selected ? "border-sky-400/70" : "border-input")}>
|
||||
<div className="flex h-20 items-center justify-center overflow-hidden bg-muted/50"><AssetPreview asset={asset} onMissing={onMissing} /></div>
|
||||
<div className="flex items-center gap-1 p-1.5">
|
||||
<div className="min-w-0 flex-1"><p title={asset.name} className="truncate text-[10px] font-medium">{asset.name}</p><p className="text-[9px] capitalize text-muted-foreground">{asset.missing ? "Upload again" : `${asset.kind} · ${(asset.size / 1024 / 1024).toFixed(1)} MB`}</p></div>
|
||||
{mode === "ref2va" && <Button type="button" variant="ghost" size="icon-sm" className="size-7" aria-label={`Add ${asset.name} as reference`} disabled={disabled || selected || asset.missing || conditioning.length >= 12} onClick={() => onAssign(asset.asset_id, "reference")}>{selected ? <Check className="size-3.5 text-sky-500" /> : <Plus className="size-3.5" />}</Button>}
|
||||
<Button type="button" variant="ghost" size="icon-sm" className="size-7 text-muted-foreground" aria-label={`Remove asset ${asset.name}`} disabled={disabled} onClick={() => onRemove(asset.asset_id)}><Trash2 className="size-3" /></Button>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
})}
|
||||
</div>}
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
{!locked && <p className="px-4 pb-2 text-[10px] text-muted-foreground">Images ≤15 MiB / 16 MP{mode === "ref2va" ? " / 1:4–4:1 aspect ratio" : ""} · video/audio ≤100 MiB / 30 sec · video up to 4K · mono/stereo audio</p>}
|
||||
{(error || validationNotice) && <p role={error ? "alert" : "status"} className={cn("border-t border-border/60 px-4 py-2 text-[11px]", error ? "bg-rose-500/5 text-rose-600 dark:text-rose-300" : "bg-amber-500/5 text-amber-700 dark:text-amber-300")}>{error || validationNotice}</p>}
|
||||
</section>
|
||||
);
|
||||
}
|
||||
@@ -1,155 +0,0 @@
|
||||
import { render, screen, waitFor, within } from "@testing-library/react";
|
||||
import userEvent from "@testing-library/user-event";
|
||||
import { afterAll, beforeAll, describe, expect, it, vi } from "vitest";
|
||||
|
||||
import ChatBar from "./ChatBar";
|
||||
|
||||
// JSDOM does not implement the pointer/scroll APIs used by the Radix popup.
|
||||
const domPolyfills = {
|
||||
hasPointerCapture: () => false,
|
||||
releasePointerCapture: () => {},
|
||||
scrollIntoView: () => {},
|
||||
};
|
||||
const originalDescriptors = new Map<string, PropertyDescriptor | undefined>();
|
||||
beforeAll(() => {
|
||||
for (const [name, implementation] of Object.entries(domPolyfills)) {
|
||||
originalDescriptors.set(name, Object.getOwnPropertyDescriptor(HTMLElement.prototype, name));
|
||||
Object.defineProperty(HTMLElement.prototype, name, { configurable: true, value: implementation });
|
||||
}
|
||||
vi.stubGlobal("PointerEvent", MouseEvent);
|
||||
});
|
||||
afterAll(() => {
|
||||
for (const [name, descriptor] of originalDescriptors) {
|
||||
if (descriptor) Object.defineProperty(HTMLElement.prototype, name, descriptor);
|
||||
else Reflect.deleteProperty(HTMLElement.prototype, name);
|
||||
}
|
||||
vi.unstubAllGlobals();
|
||||
});
|
||||
|
||||
describe("ChatBar generation mode selection", () => {
|
||||
it("places Mode and the prompt input inside the same composer", () => {
|
||||
render(<ChatBar />);
|
||||
|
||||
const composer = screen.getByRole("group", { name: "Prompt composer" });
|
||||
expect(within(composer).getByText("Mode", { exact: true })).toBeVisible();
|
||||
expect(within(composer).getByRole("combobox", { name: "Generation mode" }))
|
||||
.toBeVisible();
|
||||
expect(within(composer).getByRole("textbox", { name: "Continuation prompt" }))
|
||||
.toBeVisible();
|
||||
expect(screen.queryByText("Generation mode", { exact: true }))
|
||||
.not.toBeInTheDocument();
|
||||
});
|
||||
|
||||
it("shows only mode abbreviations and keeps explanations in the tooltip", async () => {
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar />);
|
||||
|
||||
expect(screen.getByRole("combobox", { name: "Generation mode" }))
|
||||
.toHaveAttribute("title", "Text to video + audio. Start with a text prompt; no reference asset is required.");
|
||||
expect(screen.queryByText("Start with a text prompt; no reference asset is required."))
|
||||
.not.toBeInTheDocument();
|
||||
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
|
||||
await user.click(screen.getByRole("combobox", { name: "Generation mode" }));
|
||||
const menu = await screen.findByRole("listbox");
|
||||
expect(within(menu).getAllByRole("option").map((option) => option.textContent))
|
||||
.toEqual(["T2VA", "FL2VA", "Ref2VA"]);
|
||||
});
|
||||
|
||||
it("defaults to T2VA and reports a selected mode", async () => {
|
||||
const onGenerationModeChange = vi.fn();
|
||||
const user = userEvent.setup();
|
||||
|
||||
render(
|
||||
<ChatBar
|
||||
canJoinSession
|
||||
continuationDraft="A lighthouse in a storm"
|
||||
onGenerationModeChange={onGenerationModeChange}
|
||||
/>,
|
||||
);
|
||||
|
||||
const modeSelect = screen.getByRole("combobox", { name: "Generation mode" });
|
||||
expect(modeSelect).toHaveTextContent("T2VA");
|
||||
|
||||
await user.click(modeSelect);
|
||||
await user.click(await screen.findByRole("option", { name: "Ref2VA" }));
|
||||
|
||||
expect(onGenerationModeChange).toHaveBeenCalledWith("ref2va");
|
||||
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
|
||||
});
|
||||
|
||||
it("hides mode selection after generation starts", () => {
|
||||
render(<ChatBar sessionStarted />);
|
||||
|
||||
expect(screen.queryByRole("combobox", { name: "Generation mode" }))
|
||||
.not.toBeInTheDocument();
|
||||
expect(within(screen.getByRole("group", { name: "Prompt composer" }))
|
||||
.getByRole("textbox", { name: "Continuation prompt" })).toBeVisible();
|
||||
});
|
||||
|
||||
it("disables mode selection and prompt editing while generation is busy", async () => {
|
||||
const onGenerationModeChange = vi.fn();
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar isGenerating onGenerationModeChange={onGenerationModeChange} />);
|
||||
|
||||
const composer = screen.getByRole("group", { name: "Prompt composer" });
|
||||
const modeSelect = within(composer).getByRole("combobox", { name: "Generation mode" });
|
||||
expect(modeSelect).toBeDisabled();
|
||||
expect(within(composer).getByRole("textbox", { name: "Continuation prompt" }))
|
||||
.toBeDisabled();
|
||||
await user.click(modeSelect);
|
||||
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
|
||||
expect(onGenerationModeChange).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("still submits the prompt with Enter from the combined composer", async () => {
|
||||
const onGenerate = vi.fn();
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar canJoinSession continuationDraft="A lighthouse in a storm" onGenerate={onGenerate} />);
|
||||
|
||||
const input = within(screen.getByRole("group", { name: "Prompt composer" }))
|
||||
.getByRole("textbox", { name: "Continuation prompt" });
|
||||
await user.click(input);
|
||||
await user.keyboard("{Enter}");
|
||||
expect(onGenerate).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it("disables unsupported modes and labels mock playback", async () => {
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar supportedGenerationModes={["t2va"]} mockRuntime />);
|
||||
expect(screen.getByText(/No AI model is generating/)).toBeInTheDocument();
|
||||
await user.click(screen.getByRole("combobox", { name: "Generation mode" }));
|
||||
expect(await screen.findByRole("option", { name: "FL2VA" })).toHaveAttribute("aria-disabled", "true");
|
||||
expect(screen.getByRole("option", { name: "Ref2VA" })).toHaveAttribute("aria-disabled", "true");
|
||||
expect(screen.getByRole("option", { name: "Ref2VA" }))
|
||||
.toHaveAttribute("title", "References to video + audio (unavailable on this runtime)");
|
||||
});
|
||||
|
||||
it("closes the menu with Escape and restores focus to Mode", async () => {
|
||||
const onGenerationModeChange = vi.fn();
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar onGenerationModeChange={onGenerationModeChange} />);
|
||||
const modeSelect = screen.getByRole("combobox", { name: "Generation mode" });
|
||||
await user.click(modeSelect);
|
||||
await screen.findByRole("listbox");
|
||||
await user.keyboard("{Escape}");
|
||||
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
|
||||
await waitFor(() => expect(modeSelect).toHaveFocus());
|
||||
expect(onGenerationModeChange).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("supports choosing a mode with the keyboard", async () => {
|
||||
const onGenerationModeChange = vi.fn();
|
||||
const user = userEvent.setup();
|
||||
render(<ChatBar onGenerationModeChange={onGenerationModeChange} />);
|
||||
await user.click(screen.getByRole("textbox", { name: "Continuation prompt" }));
|
||||
await user.tab();
|
||||
expect(screen.getByRole("combobox", { name: "Generation mode" })).toHaveFocus();
|
||||
await user.keyboard("{ArrowDown}");
|
||||
await waitFor(() => expect(screen.getByRole("option", { name: "T2VA" })).toHaveFocus());
|
||||
await user.keyboard("{ArrowDown}");
|
||||
await waitFor(() => expect(screen.getByRole("option", { name: "FL2VA" })).toHaveFocus());
|
||||
await user.keyboard("{Enter}");
|
||||
expect(onGenerationModeChange).toHaveBeenCalledWith("fl2va");
|
||||
expect(screen.queryByRole("listbox")).not.toBeInTheDocument();
|
||||
});
|
||||
});
|
||||
@@ -4,16 +4,8 @@ import React, { useRef, useState, useCallback, useEffect } from "react";
|
||||
import Image from "next/image";
|
||||
import { Film, ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
|
||||
import { Button } from "@/components/ui/button";
|
||||
import { Select, SelectContent, SelectItem, SelectTrigger, SelectValue } from "@/components/ui/select";
|
||||
import LeaveSessionModal, { shouldShowLeaveWarning } from "@/components/LeaveSessionModal";
|
||||
import SpeechToTextButton from "@/components/SpeechToTextButton";
|
||||
import {
|
||||
DEFAULT_GENERATION_MODE,
|
||||
GENERATION_MODES,
|
||||
getGenerationMode,
|
||||
isGenerationMode,
|
||||
type GenerationMode,
|
||||
} from "@/lib/generationMode";
|
||||
import { cn } from "@/lib/utils";
|
||||
|
||||
const PROMPT_MAX_LENGTH = 500;
|
||||
@@ -30,12 +22,6 @@ interface Props {
|
||||
sessionNotice?: string;
|
||||
projectResetPending?: boolean;
|
||||
viewingReadOnly?: boolean;
|
||||
generationMode?: GenerationMode;
|
||||
supportedGenerationModes?: readonly GenerationMode[];
|
||||
generationInputsValid?: boolean;
|
||||
capabilityNotice?: string;
|
||||
mockRuntime?: boolean;
|
||||
conditioningPanel?: React.ReactNode;
|
||||
onPresetGenerate?: (presetId: string) => void;
|
||||
onContinuationInput?: (e: React.ChangeEvent<HTMLTextAreaElement>) => void;
|
||||
onContinuationKeydown?: (e: React.KeyboardEvent<HTMLTextAreaElement>) => void;
|
||||
@@ -44,7 +30,6 @@ interface Props {
|
||||
onLeave?: () => void;
|
||||
onStartNewProject?: () => void;
|
||||
onBackFromViewing?: () => void;
|
||||
onGenerationModeChange?: (mode: GenerationMode) => void;
|
||||
onSpeechTranscript?: (text: string) => void;
|
||||
onSpeechInterimChange?: (text: string) => void;
|
||||
}
|
||||
@@ -61,12 +46,6 @@ export default function ChatBar({
|
||||
sessionNotice = "",
|
||||
projectResetPending = false,
|
||||
viewingReadOnly = false,
|
||||
generationMode = DEFAULT_GENERATION_MODE,
|
||||
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
|
||||
generationInputsValid = true,
|
||||
capabilityNotice = "",
|
||||
mockRuntime = false,
|
||||
conditioningPanel,
|
||||
onPresetGenerate = () => {},
|
||||
onContinuationInput = () => {},
|
||||
onContinuationKeydown = () => {},
|
||||
@@ -75,7 +54,6 @@ export default function ChatBar({
|
||||
onLeave = () => {},
|
||||
onStartNewProject = () => {},
|
||||
onBackFromViewing = () => {},
|
||||
onGenerationModeChange = () => {},
|
||||
onSpeechTranscript,
|
||||
onSpeechInterimChange,
|
||||
}: Props) {
|
||||
@@ -91,7 +69,6 @@ export default function ChatBar({
|
||||
? "What video are you imagining?"
|
||||
: "What do you want to edit?";
|
||||
const actionLabel = !sessionStarted ? "Generate" : "Rewrite rollout";
|
||||
const selectedGenerationMode = getGenerationMode(generationMode);
|
||||
|
||||
const inputRef = useRef<HTMLTextAreaElement>(null);
|
||||
const scrollRef = useRef<HTMLDivElement>(null);
|
||||
@@ -262,7 +239,7 @@ export default function ChatBar({
|
||||
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-6 py-4 text-center shadow-md backdrop-blur-sm">
|
||||
<div className="flex flex-col gap-1">
|
||||
<p className="text-sm font-semibold text-foreground">View-only project</p>
|
||||
<p className="max-w-md text-xs text-muted-foreground">This saved project is available for playback. Start a new project to create more videos.</p>
|
||||
<p className="max-w-md text-xs text-muted-foreground">Project sessions are currently limited to 5 minutes. Start a new project to create more videos.</p>
|
||||
</div>
|
||||
<div className="mt-1 flex items-center gap-2">
|
||||
<Button onClick={onBackFromViewing} variant="outline" size="sm" className="gap-1.5 rounded-full px-4">
|
||||
@@ -284,7 +261,7 @@ export default function ChatBar({
|
||||
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-8 py-5 text-center shadow-md backdrop-blur-sm">
|
||||
<div className="flex flex-col gap-1">
|
||||
<p className="text-sm font-semibold text-foreground">Session ended</p>
|
||||
<p className="max-w-xs text-xs text-muted-foreground">The runtime session has ended. Your saved videos remain available. Start a new project to continue creating.</p>
|
||||
<p className="max-w-xs text-xs text-muted-foreground">Each project currently has a 5-minute session. Start a new project to continue creating videos.</p>
|
||||
</div>
|
||||
<div className="mt-1 flex items-center gap-2">
|
||||
<Button onClick={onStartNewProject} size="sm" className="rounded-full px-5">
|
||||
@@ -303,7 +280,7 @@ export default function ChatBar({
|
||||
|
||||
return (
|
||||
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
|
||||
{storyPresets.length > 0 && !sessionStarted && generationMode === "t2va" && (
|
||||
{storyPresets.length > 0 && !sessionStarted && (
|
||||
<div className={cn("relative transition-opacity duration-200", isGenerating && "pointer-events-none opacity-40")}>
|
||||
<div
|
||||
ref={scrollRef}
|
||||
@@ -324,7 +301,7 @@ export default function ChatBar({
|
||||
<button
|
||||
key={preset.id}
|
||||
type="button"
|
||||
disabled={isBusy || !generationInputsValid}
|
||||
disabled={isGenerating}
|
||||
onClick={() => onPresetGenerate(preset.id)}
|
||||
className="flex flex-col sm:flex-row items-start gap-1.5 shrink-0 rounded-xl border p-2.5 text-left backdrop-blur-sm transition-colors max-w-42 sm:max-w-[215px] border-input bg-card/80 text-muted-foreground hover:bg-slate-200/60 hover:border-slate-400 hover:text-slate-700 dark:bg-slate-800/80 dark:text-slate-300 dark:hover:bg-slate-700/50 dark:hover:border-slate-500 dark:hover:text-slate-200"
|
||||
>
|
||||
@@ -350,12 +327,6 @@ export default function ChatBar({
|
||||
</div>
|
||||
)}
|
||||
|
||||
{mockRuntime && (
|
||||
<p role="status" className="rounded-xl border border-violet-500/25 bg-violet-500/10 px-4 py-2 text-center text-xs text-violet-700 dark:text-violet-300">
|
||||
Demo runtime · Sample playback only. No AI model is generating this video.
|
||||
</p>
|
||||
)}
|
||||
|
||||
{sessionNotice && (
|
||||
<div
|
||||
className={cn(
|
||||
@@ -375,14 +346,9 @@ export default function ChatBar({
|
||||
</div>
|
||||
)}
|
||||
|
||||
{sessionStarted && <p className="px-2 text-center text-[11px] text-muted-foreground">{selectedGenerationMode.label} · Mode and reference inputs are locked for this project.</p>}
|
||||
{conditioningPanel}
|
||||
|
||||
<div
|
||||
role="group"
|
||||
aria-label="Prompt composer"
|
||||
className={cn(
|
||||
"flex min-w-0 flex-col gap-2 rounded-3xl border p-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
|
||||
"flex min-w-0 items-center gap-1.5 rounded-4xl border py-2.5 pl-5 pr-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
|
||||
isBusy ? "border-input/60 bg-card/40" : "border-input bg-card/65",
|
||||
)}
|
||||
>
|
||||
@@ -398,85 +364,40 @@ export default function ChatBar({
|
||||
disabled={isBusy || sttBusy}
|
||||
rows={1}
|
||||
className={cn(
|
||||
"w-full min-w-0 resize-none bg-transparent px-2 py-1 text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
|
||||
"min-w-0 flex-1 resize-none bg-transparent text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
|
||||
(isBusy || sttBusy) && "cursor-not-allowed opacity-50",
|
||||
)}
|
||||
/>
|
||||
<div className="flex min-w-0 items-center gap-1.5">
|
||||
{!sessionStarted && (
|
||||
<div className="flex shrink-0 items-center gap-1 pl-2">
|
||||
<label htmlFor="generation-mode" className="cursor-pointer text-xs font-medium text-muted-foreground">
|
||||
Mode
|
||||
</label>
|
||||
<Select
|
||||
value={generationMode}
|
||||
disabled={isBusy || sttBusy}
|
||||
onValueChange={(value) => {
|
||||
if (isGenerationMode(value)) onGenerationModeChange(value);
|
||||
}}
|
||||
>
|
||||
<SelectTrigger
|
||||
id="generation-mode"
|
||||
aria-label="Generation mode"
|
||||
title={`${selectedGenerationMode.name}. ${selectedGenerationMode.description}`}
|
||||
className="h-8 w-24 cursor-pointer rounded-lg border-0 bg-transparent px-2 py-1 text-xs font-medium shadow-none hover:bg-muted/60 data-[state=open]:bg-muted/80 [&>svg]:size-3 [&>svg]:transition-transform [&[data-state=open]>svg]:rotate-180"
|
||||
>
|
||||
<SelectValue />
|
||||
</SelectTrigger>
|
||||
<SelectContent
|
||||
side="bottom"
|
||||
align="start"
|
||||
sideOffset={6}
|
||||
className="min-w-36 rounded-2xl border-input/70 bg-card/95 shadow-xl backdrop-blur-xl"
|
||||
>
|
||||
{GENERATION_MODES.map((mode) => (
|
||||
<SelectItem
|
||||
key={mode.id}
|
||||
value={mode.id}
|
||||
disabled={!supportedGenerationModes.includes(mode.id)}
|
||||
title={supportedGenerationModes.includes(mode.id) ? mode.name : `${mode.name} (unavailable on this runtime)`}
|
||||
className="cursor-pointer rounded-xl text-xs transition-colors data-[state=checked]:bg-muted/80 data-[state=checked]:font-semibold [&_svg]:size-3.5 [&_svg]:text-foreground"
|
||||
>
|
||||
{mode.label}
|
||||
</SelectItem>
|
||||
))}
|
||||
</SelectContent>
|
||||
</Select>
|
||||
</div>
|
||||
)}
|
||||
<div className="flex-1" />
|
||||
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
|
||||
{!sessionStarted ? (
|
||||
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
|
||||
{!sessionStarted ? (
|
||||
<Button
|
||||
aria-label={actionLabel}
|
||||
title={actionLabel}
|
||||
onClick={onGenerate}
|
||||
disabled={!canJoinSession || isGenerating || !continuationDraft.trim()}
|
||||
size="icon-sm"
|
||||
className="shrink-0 rounded-full"
|
||||
>
|
||||
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
|
||||
</Button>
|
||||
) : (
|
||||
<>
|
||||
<Button
|
||||
aria-label={actionLabel}
|
||||
title={actionLabel}
|
||||
onClick={onGenerate}
|
||||
disabled={!canJoinSession || isGenerating || !continuationDraft.trim()}
|
||||
onClick={onSubmitContinuation}
|
||||
disabled={!canSubmitContinuation || showSpinner || projectResetPending || !continuationDraft.trim()}
|
||||
size="icon-sm"
|
||||
className="shrink-0 rounded-full"
|
||||
>
|
||||
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
|
||||
</Button>
|
||||
) : (
|
||||
<>
|
||||
<Button
|
||||
aria-label={actionLabel}
|
||||
title={actionLabel}
|
||||
onClick={onSubmitContinuation}
|
||||
disabled={!canSubmitContinuation || showSpinner || projectResetPending || !continuationDraft.trim()}
|
||||
size="icon-sm"
|
||||
className="shrink-0 rounded-full"
|
||||
>
|
||||
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
|
||||
</Button>
|
||||
<Button variant="outline" aria-label="Leave" title="Leave" onClick={() => { if (shouldShowLeaveWarning()) setLeaveModalOpen(true); else onLeave(); }} disabled={isGenerating || projectResetPending} size="icon-sm" className="shrink-0 rounded-full">
|
||||
<X className="size-5" />
|
||||
</Button>
|
||||
</>
|
||||
)}
|
||||
</div>
|
||||
<Button variant="outline" aria-label="Leave" title="Leave" onClick={() => { if (shouldShowLeaveWarning()) setLeaveModalOpen(true); else onLeave(); }} disabled={isGenerating || projectResetPending} size="icon-sm" className="shrink-0 rounded-full">
|
||||
<X className="size-5" />
|
||||
</Button>
|
||||
</>
|
||||
)}
|
||||
</div>
|
||||
{!sessionStarted && capabilityNotice && <p className="px-2 text-[11px] text-amber-700 dark:text-amber-300">{capabilityNotice}</p>}
|
||||
<p className="px-2 text-center text-[11px] text-muted-foreground">
|
||||
LLM powered by{" "}
|
||||
<a
|
||||
|
||||
@@ -33,7 +33,7 @@ export default function SessionTimeoutModal({
|
||||
Session ended
|
||||
</h2>
|
||||
<p className="text-sm text-muted-foreground">
|
||||
This project reached the runtime session limit. Your latest video stays on screen, and the project is being kept in the archive so you can come back to it.
|
||||
This project hit the current 5-minute session limit. Your latest video stays on screen, and the project is being kept in the archive so you can come back to it.
|
||||
</p>
|
||||
</div>
|
||||
<p className="text-sm text-muted-foreground">
|
||||
|
||||
@@ -17,13 +17,6 @@ import {
|
||||
SelectValue,
|
||||
} from '@/components/ui/select';
|
||||
import { Textarea } from '@/components/ui/textarea';
|
||||
import {
|
||||
DEFAULT_GENERATION_MODE,
|
||||
GENERATION_MODES,
|
||||
getGenerationMode,
|
||||
isGenerationMode,
|
||||
type GenerationMode,
|
||||
} from '@/lib/generationMode';
|
||||
|
||||
interface DevtoolsComposerProps {
|
||||
connected?: boolean;
|
||||
@@ -44,10 +37,6 @@ interface DevtoolsComposerProps {
|
||||
loopGenerationEnabled?: boolean;
|
||||
curatedPromptLimit?: number;
|
||||
maxCuratedPromptCount?: number;
|
||||
generationMode?: GenerationMode;
|
||||
supportedGenerationModes?: readonly GenerationMode[];
|
||||
conditioningPanel?: React.ReactNode;
|
||||
onGenerationModeChange?: (mode: GenerationMode) => void;
|
||||
rewriteWindowMode?: boolean;
|
||||
rewritingSeedPrompts?: boolean;
|
||||
autoExtensionTimeoutHint?: string;
|
||||
@@ -85,10 +74,6 @@ export default function DevtoolsComposer({
|
||||
loopGenerationEnabled = false,
|
||||
curatedPromptLimit = 0,
|
||||
maxCuratedPromptCount = 0,
|
||||
generationMode = DEFAULT_GENERATION_MODE,
|
||||
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
|
||||
conditioningPanel,
|
||||
onGenerationModeChange = () => {},
|
||||
rewriteWindowMode = false,
|
||||
rewritingSeedPrompts = false,
|
||||
autoExtensionTimeoutHint = '',
|
||||
@@ -107,7 +92,6 @@ export default function DevtoolsComposer({
|
||||
onSpeechInterimChange,
|
||||
}: DevtoolsComposerProps) {
|
||||
const [sttBusy, setSttBusy] = useState(false);
|
||||
const selectedGenerationMode = getGenerationMode(generationMode);
|
||||
const submitButtonLabel = useMemo(
|
||||
() =>
|
||||
rewriteWindowMode
|
||||
@@ -126,8 +110,7 @@ export default function DevtoolsComposer({
|
||||
);
|
||||
|
||||
return (
|
||||
<section className="space-y-4">
|
||||
{conditioningPanel}
|
||||
<section>
|
||||
<Card>
|
||||
<CardContent className="space-y-5 p-5">
|
||||
<div className="grid gap-5 xl:grid-cols-[minmax(0,1fr)_320px]">
|
||||
@@ -302,48 +285,6 @@ export default function DevtoolsComposer({
|
||||
</div>
|
||||
|
||||
<div className="space-y-4">
|
||||
<div className="space-y-2">
|
||||
<Label htmlFor="devtools-generation-mode">
|
||||
Generation mode
|
||||
</Label>
|
||||
<Select
|
||||
value={generationMode}
|
||||
disabled={sessionStarted}
|
||||
onValueChange={(value) => {
|
||||
if (isGenerationMode(value)) {
|
||||
onGenerationModeChange(value);
|
||||
}
|
||||
}}
|
||||
>
|
||||
<SelectTrigger
|
||||
id="devtools-generation-mode"
|
||||
aria-label="Generation mode"
|
||||
title={selectedGenerationMode.name}
|
||||
>
|
||||
<SelectValue />
|
||||
</SelectTrigger>
|
||||
<SelectContent>
|
||||
{GENERATION_MODES.map((mode) => (
|
||||
<SelectItem
|
||||
key={mode.id}
|
||||
value={mode.id}
|
||||
disabled={!supportedGenerationModes.includes(mode.id)}
|
||||
title={
|
||||
supportedGenerationModes.includes(mode.id)
|
||||
? mode.name
|
||||
: `${mode.name} (unavailable on this runtime)`
|
||||
}
|
||||
>
|
||||
{mode.label}
|
||||
</SelectItem>
|
||||
))}
|
||||
</SelectContent>
|
||||
</Select>
|
||||
<p className="text-sm text-muted-foreground">
|
||||
{selectedGenerationMode.description}
|
||||
</p>
|
||||
</div>
|
||||
|
||||
<div className="flex items-start gap-3">
|
||||
<Checkbox
|
||||
id="devtools-enhance-prompts"
|
||||
|
||||
@@ -35,7 +35,6 @@ describe('DevtoolsShell', () => {
|
||||
expect(screen.getByText('Devtools Mode')).toBeInTheDocument();
|
||||
expect(screen.getByText('Your video will appear here')).toBeInTheDocument();
|
||||
expect(screen.getByLabelText('Story preset')).toBeInTheDocument();
|
||||
expect(screen.getByLabelText('Generation mode')).toBeInTheDocument();
|
||||
expect(screen.getByLabelText('Continuation prompt')).toBeDisabled();
|
||||
|
||||
expect(screen.getByText('Advanced controls')).toBeInTheDocument();
|
||||
|
||||
@@ -9,7 +9,6 @@ import VideoPlayer from '../VideoPlayer';
|
||||
import RewriteInspector from '../rewrite/RewriteInspector';
|
||||
import DevtoolsComposer from './DevtoolsComposer';
|
||||
import DevtoolsDrawer from './DevtoolsDrawer';
|
||||
import { DEFAULT_GENERATION_MODE, GENERATION_MODES, type GenerationMode } from '@/lib/generationMode';
|
||||
|
||||
interface DevtoolsShellProps {
|
||||
connected?: boolean;
|
||||
@@ -31,11 +30,6 @@ interface DevtoolsShellProps {
|
||||
curatedPromptLimit?: number;
|
||||
maxCuratedPromptCount?: number;
|
||||
|
||||
generationMode?: GenerationMode;
|
||||
supportedGenerationModes?: readonly GenerationMode[];
|
||||
conditioningPanel?: React.ReactNode;
|
||||
onGenerationModeChange?: (mode: GenerationMode) => void;
|
||||
|
||||
onPresetChange?: (e: React.ChangeEvent<HTMLSelectElement>) => void;
|
||||
onEnhancementToggle?: (e: React.ChangeEvent<HTMLInputElement>) => void;
|
||||
onCuratedPromptLimitChange?: (e: React.ChangeEvent<HTMLInputElement>) => void;
|
||||
@@ -143,11 +137,6 @@ export default function DevtoolsShell({
|
||||
curatedPromptLimit = 0,
|
||||
maxCuratedPromptCount = 0,
|
||||
|
||||
generationMode = DEFAULT_GENERATION_MODE,
|
||||
supportedGenerationModes = GENERATION_MODES.map((mode) => mode.id),
|
||||
conditioningPanel,
|
||||
onGenerationModeChange = () => {},
|
||||
|
||||
onPresetChange = () => {},
|
||||
onEnhancementToggle = () => {},
|
||||
onCuratedPromptLimitChange = () => {},
|
||||
@@ -298,10 +287,6 @@ export default function DevtoolsShell({
|
||||
loopGenerationEnabled={loopGenerationEnabled}
|
||||
curatedPromptLimit={curatedPromptLimit}
|
||||
maxCuratedPromptCount={maxCuratedPromptCount}
|
||||
generationMode={generationMode}
|
||||
supportedGenerationModes={supportedGenerationModes}
|
||||
conditioningPanel={conditioningPanel}
|
||||
onGenerationModeChange={onGenerationModeChange}
|
||||
rewriteWindowMode={livePromptRewriteMode}
|
||||
rewritingSeedPrompts={rewritingSeedPrompts}
|
||||
autoExtensionTimeoutHint={autoExtensionTimeoutHint}
|
||||
|
||||
@@ -1,55 +0,0 @@
|
||||
import { act, renderHook, waitFor } from "@testing-library/react";
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { useAssetLibrary } from "./useAssetLibrary";
|
||||
|
||||
const image = { asset_id: "asset-1", kind: "image", name: "frame.png", mime_type: "image/png", size: 5, url: "/assets/asset-1" };
|
||||
|
||||
describe("asset library lifecycle", () => {
|
||||
beforeEach(() => localStorage.clear());
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
it("uploads the raw file and keeps assignment state separate from uploaded assets", async () => {
|
||||
const fetchMock = vi.fn(async () => ({ ok: true, json: async () => image }));
|
||||
vi.stubGlobal("fetch", fetchMock);
|
||||
const { result } = renderHook(() => useAssetLibrary());
|
||||
const file = new File(["image"], "first frame.png", { type: "image/png" });
|
||||
await act(async () => { await result.current.uploadAssets([file]); });
|
||||
expect(fetchMock).toHaveBeenCalledWith("/assets", expect.objectContaining({ body: file, headers: { "Content-Type": "image/png", "X-Asset-Name": "first%20frame.png" } }));
|
||||
act(() => result.current.assignAsset("asset-1", "first_frame"));
|
||||
expect(result.current.conditioningAssets).toEqual([{ asset_id: "asset-1", role: "first_frame" }]);
|
||||
act(() => result.current.clearConditioning());
|
||||
expect(result.current.assets).toHaveLength(1);
|
||||
expect(result.current.conditioningAssets).toEqual([]);
|
||||
});
|
||||
it("identifies stale server assets without downloading their contents", async () => {
|
||||
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
|
||||
const fetchMock = vi.fn(async () => ({ ok: false, status: 404 }));
|
||||
vi.stubGlobal("fetch", fetchMock);
|
||||
const { result } = renderHook(() => useAssetLibrary());
|
||||
await waitFor(() => expect(result.current.assets).toHaveLength(1));
|
||||
act(() => result.current.assignAsset("asset-1", "first_frame"));
|
||||
let message: string | null = null;
|
||||
await act(async () => { message = await result.current.verifySelectedAssets(); });
|
||||
expect(message).toMatch(/Upload it again/);
|
||||
expect(result.current.assets[0].missing).toBe(true);
|
||||
expect(fetchMock).toHaveBeenCalledWith("/assets/asset-1", expect.objectContaining({ method: "HEAD" }));
|
||||
});
|
||||
it("deletes both the library entry and its selected references", async () => {
|
||||
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
|
||||
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true })));
|
||||
const { result } = renderHook(() => useAssetLibrary());
|
||||
await waitFor(() => expect(result.current.assets).toHaveLength(1));
|
||||
act(() => result.current.assignAsset("asset-1", "reference"));
|
||||
await act(async () => { await result.current.removeAsset("asset-1"); });
|
||||
expect(result.current.assets).toEqual([]);
|
||||
expect(result.current.conditioningAssets).toEqual([]);
|
||||
});
|
||||
|
||||
it("does not mark a valid upload missing when the browser cannot preview its codec", async () => {
|
||||
localStorage.setItem("dreamverse-asset-library-v1", JSON.stringify([image]));
|
||||
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true, status: 200 })));
|
||||
const { result } = renderHook(() => useAssetLibrary());
|
||||
await waitFor(() => expect(result.current.assets).toHaveLength(1));
|
||||
await act(async () => { await result.current.checkAssetAvailability("asset-1"); });
|
||||
expect(result.current.assets[0].missing).not.toBe(true);
|
||||
});
|
||||
});
|
||||
@@ -1,136 +0,0 @@
|
||||
"use client";
|
||||
|
||||
import { useCallback, useEffect, useRef, useState } from "react";
|
||||
import type { ConditioningAsset, ConditioningRole, GenerationAsset } from "@/lib/generationMode";
|
||||
|
||||
const LIBRARY_KEY = "dreamverse-asset-library-v1";
|
||||
const MAX_UPLOAD_BYTES = 100 * 1024 * 1024;
|
||||
|
||||
async function responseError(response: Response, fallback: string): Promise<Error> {
|
||||
const payload = await response.json().catch(() => ({}));
|
||||
return new Error(typeof payload.detail === "string" ? payload.detail : fallback);
|
||||
}
|
||||
|
||||
function isAsset(value: unknown): value is GenerationAsset {
|
||||
if (!value || typeof value !== "object") return false;
|
||||
const asset = value as Partial<GenerationAsset>;
|
||||
return typeof asset.asset_id === "string" && /^[a-zA-Z0-9_-]+$/.test(asset.asset_id)
|
||||
&& ["image", "video", "audio"].includes(asset.kind || "")
|
||||
&& typeof asset.name === "string" && typeof asset.mime_type === "string" && typeof asset.size === "number";
|
||||
}
|
||||
|
||||
/** Asset ownership lives here so composers and other pickers share the same library. */
|
||||
export function useAssetLibrary() {
|
||||
const [assets, setAssets] = useState<GenerationAsset[]>([]);
|
||||
const [conditioningAssets, setConditioningAssets] = useState<ConditioningAsset[]>([]);
|
||||
const [uploading, setUploading] = useState(false);
|
||||
const [assetError, setAssetError] = useState("");
|
||||
const [hydrated, setHydrated] = useState(false);
|
||||
const uploadingRef = useRef(false);
|
||||
|
||||
useEffect(() => {
|
||||
try {
|
||||
const saved = JSON.parse(localStorage.getItem(LIBRARY_KEY) || "[]");
|
||||
if (Array.isArray(saved)) setAssets(saved.filter(isAsset).map((asset) => ({
|
||||
...asset, url: `/assets/${asset.asset_id}`,
|
||||
})));
|
||||
} catch { /* Storage is optional; uploads still work in private browsing. */ }
|
||||
setHydrated(true);
|
||||
}, []);
|
||||
useEffect(() => {
|
||||
if (!hydrated) return;
|
||||
try { localStorage.setItem(LIBRARY_KEY, JSON.stringify(assets)); } catch { /* Optional cache. */ }
|
||||
}, [assets, hydrated]);
|
||||
|
||||
const uploadAssets = useCallback(async (files: File[]) => {
|
||||
if (uploadingRef.current) return;
|
||||
uploadingRef.current = true;
|
||||
setUploading(true);
|
||||
setAssetError("");
|
||||
const errors: string[] = [];
|
||||
for (const file of files) {
|
||||
try {
|
||||
if (!/^(image|video|audio)\//.test(file.type)) throw new Error(`${file.name}: choose an image, video, or audio file.`);
|
||||
const maxBytes = file.type.startsWith("image/") ? 15 * 1024 * 1024 : MAX_UPLOAD_BYTES;
|
||||
if (!file.size || file.size > maxBytes) throw new Error(`${file.name}: use a non-empty file up to ${maxBytes / 1024 / 1024} MiB.`);
|
||||
const response = await fetch("/assets", {
|
||||
method: "POST",
|
||||
headers: { "Content-Type": file.type, "X-Asset-Name": encodeURIComponent(file.name) },
|
||||
body: file,
|
||||
});
|
||||
if (!response.ok) throw await responseError(response, `Could not upload ${file.name}.`);
|
||||
const asset: unknown = await response.json();
|
||||
if (!isAsset(asset)) throw new Error("The server returned an invalid asset. Please retry the upload.");
|
||||
setAssets((current) => [...current.filter((item) => item.asset_id !== asset.asset_id), {
|
||||
...asset, url: `/assets/${asset.asset_id}`, missing: false,
|
||||
}]);
|
||||
} catch (error) {
|
||||
errors.push(error instanceof Error ? error.message : `Could not upload ${file.name}.`);
|
||||
}
|
||||
}
|
||||
setAssetError(errors.join(" "));
|
||||
setUploading(false);
|
||||
uploadingRef.current = false;
|
||||
}, []);
|
||||
|
||||
const assignAsset = useCallback((assetId: string, role: ConditioningRole) => {
|
||||
setConditioningAssets((current) => {
|
||||
const next = role === "reference" ? current : current.filter((item) => item.role !== role);
|
||||
if (!assetId || next.some((item) => item.asset_id === assetId && item.role === role)) return next;
|
||||
return [...next, { asset_id: assetId, role }];
|
||||
});
|
||||
}, []);
|
||||
const removeConditioning = useCallback((index: number) => {
|
||||
setConditioningAssets((current) => current.filter((_, itemIndex) => itemIndex !== index));
|
||||
}, []);
|
||||
const moveConditioning = useCallback((from: number, to: number) => {
|
||||
setConditioningAssets((current) => {
|
||||
if (from < 0 || to < 0 || from >= current.length || to >= current.length) return current;
|
||||
const next = [...current];
|
||||
next.splice(to, 0, next.splice(from, 1)[0]);
|
||||
return next;
|
||||
});
|
||||
}, []);
|
||||
const clearConditioning = useCallback(() => setConditioningAssets([]), []);
|
||||
const markAssetMissing = useCallback((assetId: string) => {
|
||||
setAssets((current) => current.map((asset) => asset.asset_id === assetId ? { ...asset, missing: true } : asset));
|
||||
}, []);
|
||||
const checkAssetAvailability = useCallback(async (assetId: string) => {
|
||||
try {
|
||||
const response = await fetch(`/assets/${assetId}`, { method: "HEAD", signal: AbortSignal.timeout(4000) });
|
||||
if (response.status === 404) markAssetMissing(assetId);
|
||||
} catch { /* A browser preview failure alone does not mean the upload expired. */ }
|
||||
}, [markAssetMissing]);
|
||||
const removeAsset = useCallback(async (assetId: string) => {
|
||||
setAssetError("");
|
||||
try {
|
||||
const response = await fetch(`/assets/${assetId}`, { method: "DELETE" });
|
||||
if (!response.ok && response.status !== 404) throw await responseError(response, "Could not remove the asset. Retry when the backend is available.");
|
||||
setAssets((current) => current.filter((asset) => asset.asset_id !== assetId));
|
||||
setConditioningAssets((current) => current.filter((asset) => asset.asset_id !== assetId));
|
||||
} catch (error) {
|
||||
setAssetError(error instanceof Error ? error.message : "Could not remove the asset.");
|
||||
}
|
||||
}, []);
|
||||
const verifySelectedAssets = useCallback(async (): Promise<string | null> => {
|
||||
const selected = [...new Set(conditioningAssets.map((item) => item.asset_id))];
|
||||
try {
|
||||
for (const assetId of selected) {
|
||||
const response = await fetch(`/assets/${assetId}`, { method: "HEAD", signal: AbortSignal.timeout(4000) });
|
||||
if (response.status === 404) {
|
||||
markAssetMissing(assetId);
|
||||
return "A selected asset expired or was removed from the server. Upload it again, then select the new copy.";
|
||||
}
|
||||
if (!response.ok) return "Could not verify the selected assets. Check the backend and try again.";
|
||||
}
|
||||
return null;
|
||||
} catch {
|
||||
return "Could not verify the selected assets. Check the backend and try again.";
|
||||
}
|
||||
}, [conditioningAssets, markAssetMissing]);
|
||||
|
||||
return {
|
||||
assets, conditioningAssets, uploading, assetError, uploadAssets, assignAsset,
|
||||
removeConditioning, moveConditioning, clearConditioning, removeAsset, checkAssetAvailability, verifySelectedAssets,
|
||||
};
|
||||
}
|
||||
@@ -1,21 +0,0 @@
|
||||
import { renderHook, waitFor } from "@testing-library/react";
|
||||
import { afterEach, describe, expect, it, vi } from "vitest";
|
||||
import { useGenerationCapabilities } from "./useGenerationCapabilities";
|
||||
|
||||
describe("generation capabilities", () => {
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
it("enables only the modes the runtime advertises and identifies mock playback", async () => {
|
||||
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: true, json: async () => ({ model_id: "mock", modes: ["t2va", "fl2va", "ref2va"], mock: true }) })));
|
||||
const { result } = renderHook(() => useGenerationCapabilities());
|
||||
await waitFor(() => expect(result.current.loadingCapabilities).toBe(false));
|
||||
expect(result.current.capabilities.modes).toEqual(["t2va", "fl2va", "ref2va"]);
|
||||
expect(result.current.capabilities.mock).toBe(true);
|
||||
});
|
||||
it("keeps old runtimes text-only when the capabilities endpoint is missing", async () => {
|
||||
vi.stubGlobal("fetch", vi.fn(async () => ({ ok: false, status: 404 })));
|
||||
const { result } = renderHook(() => useGenerationCapabilities());
|
||||
await waitFor(() => expect(result.current.loadingCapabilities).toBe(false));
|
||||
expect(result.current.capabilities.modes).toEqual(["t2va"]);
|
||||
expect(result.current.capabilityNotice).toMatch(/Text-only compatibility/);
|
||||
});
|
||||
});
|
||||
@@ -1,37 +0,0 @@
|
||||
"use client";
|
||||
|
||||
import { useCallback, useEffect, useState } from "react";
|
||||
import { isGenerationMode, type GenerationCapabilities } from "@/lib/generationMode";
|
||||
|
||||
const LEGACY_CAPABILITIES: GenerationCapabilities = { model_id: "legacy", modes: ["t2va"] };
|
||||
|
||||
export function useGenerationCapabilities() {
|
||||
const [capabilities, setCapabilities] = useState<GenerationCapabilities>(LEGACY_CAPABILITIES);
|
||||
const [capabilityNotice, setCapabilityNotice] = useState("");
|
||||
const [loadingCapabilities, setLoadingCapabilities] = useState(true);
|
||||
const refreshCapabilities = useCallback(async () => {
|
||||
try {
|
||||
const response = await fetch("/generation-capabilities", { signal: AbortSignal.timeout(4000) });
|
||||
if (!response.ok) throw new Error("Capabilities unavailable");
|
||||
const payload = await response.json();
|
||||
if (typeof payload.model_id !== "string" || !Array.isArray(payload.modes)
|
||||
|| !payload.modes.every(isGenerationMode)) throw new Error("Invalid capabilities");
|
||||
const next: GenerationCapabilities = {
|
||||
model_id: payload.model_id,
|
||||
modes: payload.modes,
|
||||
mock: payload.mock === true,
|
||||
};
|
||||
setCapabilities(next);
|
||||
setCapabilityNotice("");
|
||||
return next;
|
||||
} catch {
|
||||
setCapabilities(LEGACY_CAPABILITIES);
|
||||
setCapabilityNotice("Runtime capabilities unavailable. Text-only compatibility mode is available; check the backend to enable image and reference modes.");
|
||||
return LEGACY_CAPABILITIES;
|
||||
} finally {
|
||||
setLoadingCapabilities(false);
|
||||
}
|
||||
}, []);
|
||||
useEffect(() => { void refreshCapabilities(); }, [refreshCapabilities]);
|
||||
return { capabilities, capabilityNotice, loadingCapabilities, refreshCapabilities };
|
||||
}
|
||||
@@ -1,58 +0,0 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
|
||||
import {
|
||||
DEFAULT_GENERATION_MODE,
|
||||
GENERATION_MODES,
|
||||
getGenerationMode,
|
||||
isGenerationMode,
|
||||
buildGenerationInitFields,
|
||||
validateGenerationInputs,
|
||||
type GenerationAsset,
|
||||
} from "./generationMode";
|
||||
|
||||
describe("generation modes", () => {
|
||||
it("exposes stable wire IDs in the expected product order", () => {
|
||||
expect(GENERATION_MODES.map((mode) => mode.id)).toEqual([
|
||||
"t2va",
|
||||
"fl2va",
|
||||
"ref2va",
|
||||
]);
|
||||
expect(DEFAULT_GENERATION_MODE).toBe("t2va");
|
||||
});
|
||||
|
||||
it("validates and resolves generation mode values", () => {
|
||||
expect(isGenerationMode("ref2va")).toBe(true);
|
||||
expect(isGenerationMode("unknown")).toBe(false);
|
||||
expect(getGenerationMode("fl2va").label).toBe("FL2VA");
|
||||
});
|
||||
});
|
||||
|
||||
const image: GenerationAsset = { asset_id: "img", kind: "image", name: "frame.png", mime_type: "image/png", size: 12, url: "/assets/img" };
|
||||
const audio: GenerationAsset = { asset_id: "sound", kind: "audio", name: "sound.wav", mime_type: "audio/wav", size: 12, url: "/assets/sound" };
|
||||
|
||||
describe("generation input contract", () => {
|
||||
it("keeps text-only init valid and rejects accidental references", () => {
|
||||
expect(buildGenerationInitFields("t2va", [], [])).toEqual({ generation_mode: "t2va", conditioning_assets: [] });
|
||||
expect(() => buildGenerationInitFields("t2va", [{ asset_id: "img", role: "reference" }], [image])).toThrow("text only");
|
||||
});
|
||||
it("requires first frame but permits first-only or both endpoints", () => {
|
||||
expect(validateGenerationInputs("fl2va", [], [])).toMatch(/first frame/);
|
||||
const first = { asset_id: "img", role: "first_frame" } as const;
|
||||
expect(validateGenerationInputs("fl2va", [first], [image])).toBeNull();
|
||||
expect(validateGenerationInputs("fl2va", [first, { asset_id: "img", role: "last_frame" }], [image])).toBeNull();
|
||||
expect(validateGenerationInputs("fl2va", [first, first], [image])).toMatch(/one first frame/);
|
||||
expect(validateGenerationInputs("fl2va", [{ asset_id: "sound", role: "first_frame" }], [audio])).toMatch(/images only/);
|
||||
});
|
||||
it("requires a visual reference and preserves multimodal ordering without file bodies", () => {
|
||||
expect(validateGenerationInputs("ref2va", [{ asset_id: "sound", role: "reference" }], [audio])).toMatch(/image or video/);
|
||||
const items = [{ asset_id: "sound", role: "reference" }, { asset_id: "img", role: "reference" }] as const;
|
||||
expect(buildGenerationInitFields("ref2va", items, [image, audio])).toEqual({ generation_mode: "ref2va", conditioning_assets: items });
|
||||
});
|
||||
it("rejects per-kind limits, total limits, and stale uploads", () => {
|
||||
const images = Array.from({ length: 10 }, (_, index) => ({ ...image, asset_id: `img-${index}` }));
|
||||
expect(validateGenerationInputs("ref2va", images.map((item) => ({ asset_id: item.asset_id, role: "reference" })), images)).toMatch(/at most 9 image/);
|
||||
const refs = Array.from({ length: 13 }, () => ({ asset_id: "img", role: "reference" as const }));
|
||||
expect(validateGenerationInputs("ref2va", refs, [image])).toMatch(/at most 12/);
|
||||
expect(validateGenerationInputs("fl2va", [{ asset_id: "img", role: "first_frame" }], [{ ...image, missing: true }])).toMatch(/no longer on the server/);
|
||||
});
|
||||
});
|
||||
@@ -1,114 +0,0 @@
|
||||
export const GENERATION_MODES = [
|
||||
{
|
||||
id: "t2va",
|
||||
label: "T2VA",
|
||||
name: "Text to video + audio",
|
||||
description: "Start with a text prompt; no reference asset is required.",
|
||||
},
|
||||
{
|
||||
id: "fl2va",
|
||||
label: "FL2VA",
|
||||
name: "First/last frames to video + audio",
|
||||
description: "Start from a first frame image. Add an optional last frame to guide the ending.",
|
||||
},
|
||||
{
|
||||
id: "ref2va",
|
||||
label: "Ref2VA",
|
||||
name: "References to video + audio",
|
||||
description: "Guide the result with ordered image, video, or audio references.",
|
||||
},
|
||||
] as const;
|
||||
|
||||
export type GenerationMode = (typeof GENERATION_MODES)[number]["id"];
|
||||
|
||||
export const DEFAULT_GENERATION_MODE: GenerationMode = "t2va";
|
||||
|
||||
export function isGenerationMode(value: unknown): value is GenerationMode {
|
||||
return GENERATION_MODES.some((mode) => mode.id === value);
|
||||
}
|
||||
|
||||
export function getGenerationMode(value: GenerationMode) {
|
||||
return GENERATION_MODES.find((mode) => mode.id === value) ?? GENERATION_MODES[0];
|
||||
}
|
||||
|
||||
export type AssetKind = "image" | "video" | "audio";
|
||||
export type ConditioningRole = "first_frame" | "last_frame" | "reference";
|
||||
|
||||
/** Runtime-owned uploads; project metadata keeps references, never file contents. */
|
||||
export interface GenerationAsset {
|
||||
asset_id: string;
|
||||
kind: AssetKind;
|
||||
name: string;
|
||||
mime_type: string;
|
||||
size: number;
|
||||
url: string;
|
||||
missing?: boolean;
|
||||
}
|
||||
|
||||
export interface ConditioningAsset {
|
||||
asset_id: string;
|
||||
role: ConditioningRole;
|
||||
}
|
||||
|
||||
export interface GenerationCapabilities {
|
||||
model_id: string;
|
||||
modes: GenerationMode[];
|
||||
mock?: boolean;
|
||||
}
|
||||
|
||||
export interface GenerationInitFields {
|
||||
generation_mode: GenerationMode;
|
||||
conditioning_assets: ConditioningAsset[];
|
||||
}
|
||||
|
||||
export const REFERENCE_LIMITS = { image: 9, video: 3, audio: 3, total: 12 } as const;
|
||||
|
||||
export function validateGenerationInputs(
|
||||
mode: GenerationMode,
|
||||
conditioning: readonly ConditioningAsset[],
|
||||
assets: readonly GenerationAsset[],
|
||||
): string | null {
|
||||
if (mode === "t2va") {
|
||||
return conditioning.length ? "T2VA uses text only. Remove the selected references." : null;
|
||||
}
|
||||
const resolved = conditioning.map((item) => assets.find((asset) => asset.asset_id === item.asset_id));
|
||||
if (resolved.some((asset) => !asset || asset.missing)) {
|
||||
return "A selected asset is no longer on the server. Upload it again and select the new copy.";
|
||||
}
|
||||
if (mode === "fl2va") {
|
||||
if (!conditioning.some((item) => item.role === "first_frame")) return "Choose a first frame image to generate.";
|
||||
if (conditioning.some((item) => item.role === "reference") || resolved.some((asset) => asset?.kind !== "image")) {
|
||||
return "FL2VA accepts first and last frame images only.";
|
||||
}
|
||||
if (conditioning.filter((item) => item.role === "first_frame").length !== 1
|
||||
|| conditioning.filter((item) => item.role === "last_frame").length > 1) {
|
||||
return "Choose one first frame and at most one last frame.";
|
||||
}
|
||||
return null;
|
||||
}
|
||||
if (conditioning.some((item) => item.role !== "reference")) return "Ref2VA accepts ordered reference assets only.";
|
||||
if (!resolved.some((asset) => asset?.kind === "image" || asset?.kind === "video")) {
|
||||
return "Add at least one image or video reference. Audio alone is not enough.";
|
||||
}
|
||||
if (conditioning.length > REFERENCE_LIMITS.total) return "Use at most 12 reference assets in total.";
|
||||
for (const kind of ["image", "video", "audio"] as const) {
|
||||
if (resolved.filter((asset) => asset?.kind === kind).length > REFERENCE_LIMITS[kind]) {
|
||||
return `Use at most ${REFERENCE_LIMITS[kind]} ${kind} references.`;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Shared by both session_init_v2 and project_init_v1. */
|
||||
export function buildGenerationInitFields(
|
||||
mode: GenerationMode,
|
||||
conditioning: readonly ConditioningAsset[],
|
||||
assets: readonly GenerationAsset[],
|
||||
): GenerationInitFields {
|
||||
const problem = validateGenerationInputs(mode, conditioning, assets);
|
||||
if (problem) throw new Error(problem);
|
||||
return {
|
||||
generation_mode: mode,
|
||||
conditioning_assets: conditioning.map(({ asset_id, role }) => ({ asset_id, role })),
|
||||
};
|
||||
}
|
||||
@@ -1,5 +1,3 @@
|
||||
import type { ConditioningAsset, GenerationAsset, GenerationMode } from "./generationMode";
|
||||
|
||||
const DB_NAME = "fastvideo-projects";
|
||||
const DB_VERSION = 1;
|
||||
const PROJECTS_STORE = "projects";
|
||||
@@ -37,11 +35,6 @@ export interface StoredProject {
|
||||
createdAt: number;
|
||||
lastThumbnail: string | null;
|
||||
promptEvents: Record<string, unknown>[];
|
||||
/** Optional for projects created before generation modes were introduced. */
|
||||
generationMode?: GenerationMode;
|
||||
conditioningAssets?: ConditioningAsset[];
|
||||
assets?: GenerationAsset[];
|
||||
mock?: boolean;
|
||||
}
|
||||
|
||||
export interface StoredClip {
|
||||
|
||||
@@ -9,7 +9,6 @@ from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import logging
|
||||
import json
|
||||
import sqlite3
|
||||
import threading
|
||||
from pathlib import Path
|
||||
@@ -47,13 +46,8 @@ DEFAULT_SETTINGS: dict[str, Any] = {
|
||||
|
||||
|
||||
def _sqlite_row_get(row: sqlite3.Row, key: str, default: Any) -> Any:
|
||||
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10).
|
||||
|
||||
NOTE: `key in row` tests Row *values*, not column names, so the membership
|
||||
check has to go through .keys() -- otherwise every lookup falls back to the
|
||||
default and jobs restored from the database lose their stored fields.
|
||||
"""
|
||||
return row[key] if key in row.keys() else default # noqa: SIM401, SIM118
|
||||
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10)."""
|
||||
return row[key] if key in row else default # noqa: SIM401
|
||||
|
||||
|
||||
def _get_db_path(data_dir: Path) -> Path:
|
||||
@@ -89,9 +83,6 @@ def _migrate_db(conn: sqlite3.Connection) -> None:
|
||||
_add_column_if_missing(conn, "jobs", "fps", "INTEGER", "24")
|
||||
_add_column_if_missing(conn, "jobs", "workload_type", "TEXT", "'t2v'")
|
||||
_add_column_if_missing(conn, "jobs", "image_path", "TEXT", "''")
|
||||
_add_column_if_missing(conn, "jobs", "name", "TEXT", "''")
|
||||
_add_column_if_missing(conn, "jobs", "last_image_path", "TEXT", "''")
|
||||
_add_column_if_missing(conn, "jobs", "references_json", "TEXT", "''")
|
||||
_add_column_if_missing(conn, "jobs", "job_type", "TEXT", "'inference'")
|
||||
_add_column_if_missing(conn, "jobs", "data_path", "TEXT", "''")
|
||||
_add_column_if_missing(conn, "jobs", "max_train_steps", "INTEGER", "1000")
|
||||
@@ -251,8 +242,7 @@ class Database:
|
||||
self._execute(
|
||||
"""
|
||||
INSERT INTO jobs (
|
||||
id, model_id, name, prompt, workload_type, image_path,
|
||||
last_image_path, references_json, job_type, status,
|
||||
id, model_id, prompt, workload_type, image_path, job_type, status,
|
||||
created_at, started_at, finished_at, error, output_path, log_file_path,
|
||||
num_inference_steps, num_frames, height, width, guidance_scale,
|
||||
guidance_rescale, fps, seed, num_gpus, dit_cpu_offload,
|
||||
@@ -264,17 +254,14 @@ class Database:
|
||||
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
|
||||
real_score_guidance_scale,
|
||||
generator_update_interval, real_score_model_path, fake_score_model_path
|
||||
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""",
|
||||
(
|
||||
job["id"],
|
||||
job["model_id"],
|
||||
job.get("name", ""),
|
||||
job["prompt"],
|
||||
job.get("workload_type", "t2v"),
|
||||
job.get("image_path", ""),
|
||||
job.get("last_image_path", ""),
|
||||
json.dumps(job.get("references") or []),
|
||||
job.get("job_type", "inference"),
|
||||
job["status"],
|
||||
job["created_at"],
|
||||
@@ -553,12 +540,9 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
|
||||
result = {
|
||||
"id": row["id"],
|
||||
"model_id": row["model_id"],
|
||||
"name": _sqlite_row_get(row, "name", "") or "",
|
||||
"prompt": row["prompt"],
|
||||
"workload_type": _sqlite_row_get(row, "workload_type", "t2v"),
|
||||
"image_path": _sqlite_row_get(row, "image_path", "") or "",
|
||||
"last_image_path": _sqlite_row_get(row, "last_image_path", "") or "",
|
||||
"references": _sqlite_row_get(row, "references_json", "") or "",
|
||||
"job_type": _sqlite_row_get(row, "job_type", "inference"),
|
||||
"status": row["status"],
|
||||
"created_at": row["created_at"],
|
||||
|
||||
@@ -1,63 +0,0 @@
|
||||
import { expect, test } from '@playwright/test';
|
||||
|
||||
import { skipWithoutMock } from './helpers';
|
||||
|
||||
test.describe('create job interactions', () => {
|
||||
skipWithoutMock();
|
||||
|
||||
for (const jobType of ['inference', 'finetuning', 'distillation']) {
|
||||
test(`${jobType} remains interactive after repeated dialog dismissals`, async ({ page }) => {
|
||||
await page.goto(`/${jobType}`);
|
||||
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
|
||||
const dialog = page.getByRole('dialog');
|
||||
|
||||
// Exercise both dismissal paths and reopen without reloading the page.
|
||||
for (const closeWithEscape of [false, true]) {
|
||||
await trigger.click();
|
||||
await page.getByRole('menuitem').first().click();
|
||||
await expect(dialog).toBeVisible();
|
||||
if (closeWithEscape) {
|
||||
await page.keyboard.press('Escape');
|
||||
} else {
|
||||
await dialog.getByRole('button', { name: 'Close', exact: true }).click();
|
||||
}
|
||||
await expect(dialog).toBeHidden();
|
||||
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
|
||||
await expect(trigger).toBeFocused();
|
||||
}
|
||||
|
||||
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
|
||||
await expect(page).toHaveURL(/\/datasets$/);
|
||||
});
|
||||
}
|
||||
|
||||
test('preserves keyboard menu dismissal and dialog focus trapping', async ({ page }) => {
|
||||
await page.goto('/inference');
|
||||
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
|
||||
await trigger.focus();
|
||||
await page.keyboard.press('Enter');
|
||||
const firstItem = page.getByRole('menuitem').first();
|
||||
await expect(firstItem).toBeFocused();
|
||||
await page.keyboard.press('Escape');
|
||||
await expect(page.getByRole('menu')).toBeHidden();
|
||||
await expect(trigger).toBeFocused();
|
||||
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
|
||||
|
||||
await page.keyboard.press('Enter');
|
||||
await expect(firstItem).toBeFocused();
|
||||
await page.keyboard.press('Enter');
|
||||
const dialog = page.getByRole('dialog');
|
||||
await expect(dialog).toBeVisible();
|
||||
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
|
||||
|
||||
// Shift+Tab from the first field wraps to Close, then Tab wraps back.
|
||||
await page.keyboard.press('Shift+Tab');
|
||||
await expect(dialog.getByRole('button', { name: 'Close', exact: true })).toBeFocused();
|
||||
await page.keyboard.press('Tab');
|
||||
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
|
||||
await page.keyboard.press('Escape');
|
||||
await expect(dialog).toBeHidden();
|
||||
await expect(trigger).toBeFocused();
|
||||
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
|
||||
});
|
||||
});
|
||||
@@ -1,6 +1,6 @@
|
||||
import { expect, test } from '@playwright/test';
|
||||
|
||||
import { API_BASE, skipWithoutMock } from './helpers';
|
||||
import { skipWithoutMock } from './helpers';
|
||||
|
||||
/**
|
||||
* Create-job flow: open the Create Job modal on /inference, fill the prompt
|
||||
@@ -10,8 +10,7 @@ import { API_BASE, skipWithoutMock } from './helpers';
|
||||
test.describe('create inference job', () => {
|
||||
skipWithoutMock();
|
||||
|
||||
test('creates a T2V job and starts it without refreshing', async ({ page, request }) => {
|
||||
await request.put(`${API_BASE}/settings`, { data: { autoStartJob: false } });
|
||||
test('creates a T2V job and shows it in the queue', async ({ page }) => {
|
||||
await page.goto('/inference');
|
||||
|
||||
// The trigger opens a real menu on click, so this path works for touch,
|
||||
@@ -39,20 +38,5 @@ test.describe('create inference job', () => {
|
||||
// Modal closes and the queue refreshes with the newly created job.
|
||||
await expect(dialog).toBeHidden();
|
||||
await expect(page.getByText(prompt)).toBeVisible();
|
||||
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
|
||||
|
||||
const card = page.getByRole('article').filter({ hasText: prompt });
|
||||
await expect(card.getByText('pending', { exact: true })).toBeVisible();
|
||||
const started = page.waitForResponse((response) =>
|
||||
response.url().startsWith(`${API_BASE}/jobs/`) &&
|
||||
response.url().endsWith('/start') &&
|
||||
response.request().method() === 'POST',
|
||||
);
|
||||
await card.getByRole('button', { name: 'Start', exact: true }).click();
|
||||
expect((await started).ok()).toBe(true);
|
||||
await expect(card.getByText('running', { exact: true })).toBeVisible();
|
||||
|
||||
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
|
||||
await expect(page).toHaveURL(/\/datasets$/);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -10,9 +10,7 @@ from __future__ import annotations
|
||||
import atexit
|
||||
import collections
|
||||
import contextlib
|
||||
import copy
|
||||
import enum
|
||||
import json
|
||||
import logging
|
||||
import logging.handlers
|
||||
import multiprocessing as mp
|
||||
@@ -125,13 +123,10 @@ class LogBufferHandler(logging.Handler):
|
||||
class Job:
|
||||
id: str
|
||||
model_id: str
|
||||
name: str = ""
|
||||
prompt: str = ""
|
||||
prompt: str
|
||||
workload_type: str = "t2v"
|
||||
job_type: str = "inference"
|
||||
image_path: str = ""
|
||||
last_image_path: str = ""
|
||||
references: list[dict[str, Any]] = field(default_factory=list)
|
||||
status: JobStatus = JobStatus.PENDING
|
||||
created_at: float = field(default_factory=time.time)
|
||||
started_at: float | None = None
|
||||
@@ -150,7 +145,6 @@ class Job:
|
||||
negative_prompt: str = ""
|
||||
num_gpus: int = 1
|
||||
dit_cpu_offload: bool = False
|
||||
dit_layerwise_offload: bool = False
|
||||
text_encoder_cpu_offload: bool = False
|
||||
vae_cpu_offload: bool = False
|
||||
image_encoder_cpu_offload: bool = False
|
||||
@@ -186,13 +180,10 @@ class Job:
|
||||
return {
|
||||
"id": self.id,
|
||||
"model_id": self.model_id,
|
||||
"name": self.name,
|
||||
"prompt": self.prompt,
|
||||
"workload_type": self.workload_type,
|
||||
"job_type": self.job_type,
|
||||
"image_path": self.image_path,
|
||||
"last_image_path": self.last_image_path,
|
||||
"references": self.references,
|
||||
"status": self.status.value,
|
||||
"created_at": self.created_at,
|
||||
"started_at": self.started_at,
|
||||
@@ -211,7 +202,6 @@ class Job:
|
||||
"negative_prompt": self.negative_prompt,
|
||||
"num_gpus": self.num_gpus,
|
||||
"dit_cpu_offload": self.dit_cpu_offload,
|
||||
"dit_layerwise_offload": self.dit_layerwise_offload,
|
||||
"text_encoder_cpu_offload": self.text_encoder_cpu_offload,
|
||||
"vae_cpu_offload": self.vae_cpu_offload,
|
||||
"image_encoder_cpu_offload": self.image_encoder_cpu_offload,
|
||||
@@ -242,75 +232,6 @@ class Job:
|
||||
}
|
||||
|
||||
|
||||
MINIMAX_H3_REF2VA_PIPELINE = "MiniMaxH3Ref2VAModularPipeline"
|
||||
|
||||
|
||||
def _build_h3_references(raw: list[dict[str, Any]]) -> list[Any]:
|
||||
"""Turn the API's reference dicts into MiniMaxH3Reference objects.
|
||||
|
||||
Imported lazily so the API server starts without pulling in fastvideo.
|
||||
"""
|
||||
from fastvideo.pipelines.basic.minimax_h3 import MiniMaxH3Reference
|
||||
|
||||
built = []
|
||||
for i, ref in enumerate(raw):
|
||||
source = (ref or {}).get("source")
|
||||
if not source:
|
||||
raise ValueError(f"reference {i} has no source")
|
||||
if not os.path.isfile(source):
|
||||
raise ValueError(f"reference {i} source not found: {source}")
|
||||
kwargs: dict[str, Any] = {
|
||||
"source": source,
|
||||
"media_type": (ref.get("media_type") or "image"),
|
||||
}
|
||||
for opt in ("soundtrack", "fps", "sample_rate"):
|
||||
if ref.get(opt) not in (None, ""):
|
||||
kwargs[opt] = ref[opt]
|
||||
built.append(MiniMaxH3Reference(**kwargs))
|
||||
return built
|
||||
|
||||
|
||||
JOB_LOG_FILENAME = "out.log"
|
||||
|
||||
|
||||
def _job_log_path(output_dir: str, job_id: str) -> str:
|
||||
"""Each job's log lives beside its outputs: <output_dir>/<job_id>/out.log."""
|
||||
return os.path.join(output_dir, job_id, JOB_LOG_FILENAME)
|
||||
|
||||
|
||||
def _decode_references(value: Any) -> list[dict[str, Any]]:
|
||||
"""Reference lists round-trip through the DB as JSON text."""
|
||||
if not value:
|
||||
return []
|
||||
if isinstance(value, list):
|
||||
return list(value)
|
||||
try:
|
||||
decoded = json.loads(value)
|
||||
except (TypeError, ValueError):
|
||||
logger.warning("Could not decode stored references: %r", value)
|
||||
return []
|
||||
return list(decoded) if isinstance(decoded, list) else []
|
||||
|
||||
|
||||
def _generator_is_alive(generator: Any) -> bool:
|
||||
"""True if the generator's worker processes are all still running.
|
||||
|
||||
A cached VideoGenerator holds a MultiprocExecutor whose workers are separate
|
||||
processes; nothing notices when they exit. Probing `proc.is_alive()` is what
|
||||
the executor itself uses during shutdown. Anything unexpected in the object
|
||||
graph is treated as alive so a probe failure can never wedge the cache.
|
||||
"""
|
||||
executor = getattr(generator, "executor", None)
|
||||
workers = getattr(executor, "workers", None)
|
||||
if not workers:
|
||||
return True
|
||||
try:
|
||||
return all(w.proc.is_alive() for w in workers)
|
||||
except Exception:
|
||||
logger.debug("Worker liveness probe failed", exc_info=True)
|
||||
return True
|
||||
|
||||
|
||||
class JobRunner:
|
||||
"""Manages video generation jobs, their execution, and generator caching."""
|
||||
|
||||
@@ -355,7 +276,7 @@ class JobRunner:
|
||||
"""Populate job's log buffer from its log file if it exists."""
|
||||
path = job.log_file_path
|
||||
if not path:
|
||||
path = _job_log_path(self.output_dir, job.id)
|
||||
path = os.path.join(self.log_dir, f"{job.id}.log")
|
||||
if not os.path.isfile(path):
|
||||
return
|
||||
try:
|
||||
@@ -393,13 +314,10 @@ class JobRunner:
|
||||
job = Job(
|
||||
id=row["id"],
|
||||
model_id=row["model_id"],
|
||||
name=row.get("name", "") or "",
|
||||
prompt=row["prompt"],
|
||||
workload_type=row.get("workload_type", "t2v"),
|
||||
job_type=row.get("job_type", "inference"),
|
||||
image_path=row.get("image_path", "") or "",
|
||||
last_image_path=row.get("last_image_path", "") or "",
|
||||
references=_decode_references(row.get("references")),
|
||||
data_path=row.get("data_path", "") or "",
|
||||
max_train_steps=row.get("max_train_steps", 1000),
|
||||
train_batch_size=row.get("train_batch_size", 1),
|
||||
@@ -432,7 +350,6 @@ class JobRunner:
|
||||
negative_prompt=row.get("negative_prompt", "") or "",
|
||||
num_gpus=row.get("num_gpus", 1),
|
||||
dit_cpu_offload=row.get("dit_cpu_offload", False),
|
||||
dit_layerwise_offload=row.get("dit_layerwise_offload", False),
|
||||
text_encoder_cpu_offload=row.get("text_encoder_cpu_offload", False),
|
||||
vae_cpu_offload=row.get("vae_cpu_offload", False),
|
||||
image_encoder_cpu_offload=row.get("image_encoder_cpu_offload", False),
|
||||
@@ -477,12 +394,9 @@ class JobRunner:
|
||||
job_id: str,
|
||||
model_id: str,
|
||||
prompt: str,
|
||||
name: str = "",
|
||||
workload_type: str = "t2v",
|
||||
job_type: str = "inference",
|
||||
image_path: str = "",
|
||||
last_image_path: str = "",
|
||||
references: list[dict[str, Any]] | None = None,
|
||||
data_path: str = "",
|
||||
max_train_steps: int = 1000,
|
||||
train_batch_size: int = 1,
|
||||
@@ -508,7 +422,6 @@ class JobRunner:
|
||||
num_gpus: int = 1,
|
||||
negative_prompt: str = "",
|
||||
dit_cpu_offload: bool = False,
|
||||
dit_layerwise_offload: bool = False,
|
||||
text_encoder_cpu_offload: bool = False,
|
||||
vae_cpu_offload: bool = False,
|
||||
image_encoder_cpu_offload: bool = False,
|
||||
@@ -522,13 +435,10 @@ class JobRunner:
|
||||
job = Job(
|
||||
id=job_id,
|
||||
model_id=model_id,
|
||||
name=(name or "").strip(),
|
||||
prompt=prompt.strip(),
|
||||
workload_type=workload_type or "t2v",
|
||||
job_type=job_type or "inference",
|
||||
image_path=image_path or "",
|
||||
last_image_path=last_image_path or "",
|
||||
references=list(references or []),
|
||||
data_path=data_path or "",
|
||||
max_train_steps=max_train_steps,
|
||||
train_batch_size=train_batch_size,
|
||||
@@ -554,7 +464,6 @@ class JobRunner:
|
||||
negative_prompt=negative_prompt or "",
|
||||
num_gpus=num_gpus,
|
||||
dit_cpu_offload=dit_cpu_offload,
|
||||
dit_layerwise_offload=dit_layerwise_offload,
|
||||
text_encoder_cpu_offload=text_encoder_cpu_offload,
|
||||
vae_cpu_offload=vae_cpu_offload,
|
||||
image_encoder_cpu_offload=image_encoder_cpu_offload,
|
||||
@@ -612,84 +521,6 @@ class JobRunner:
|
||||
logger.info("Deleted job %s", job.id)
|
||||
return True
|
||||
|
||||
CONFIG_FIELDS: tuple[str, ...] = (
|
||||
"model_id",
|
||||
"name",
|
||||
"prompt",
|
||||
"workload_type",
|
||||
"job_type",
|
||||
"image_path",
|
||||
"last_image_path",
|
||||
"references",
|
||||
"negative_prompt",
|
||||
"num_inference_steps",
|
||||
"num_frames",
|
||||
"height",
|
||||
"width",
|
||||
"guidance_scale",
|
||||
"guidance_rescale",
|
||||
"fps",
|
||||
"seed",
|
||||
"num_gpus",
|
||||
"dit_cpu_offload",
|
||||
"dit_layerwise_offload",
|
||||
"text_encoder_cpu_offload",
|
||||
"vae_cpu_offload",
|
||||
"image_encoder_cpu_offload",
|
||||
"use_fsdp_inference",
|
||||
"enable_torch_compile",
|
||||
"vsa_sparsity",
|
||||
"tp_size",
|
||||
"sp_size",
|
||||
"data_path",
|
||||
"max_train_steps",
|
||||
"train_batch_size",
|
||||
"learning_rate",
|
||||
"num_latent_t",
|
||||
"validation_dataset_file",
|
||||
"lora_rank",
|
||||
"dmd_use_vsa",
|
||||
"dmd_vsa_sparsity",
|
||||
"dmd_denoising_steps",
|
||||
"real_score_guidance_scale",
|
||||
"generator_update_interval",
|
||||
"real_score_model_path",
|
||||
"fake_score_model_path",
|
||||
)
|
||||
|
||||
def duplicate_job(self, job_id: str, new_job_id: str) -> Job:
|
||||
"""Create a new pending job with an existing job's configuration.
|
||||
|
||||
Runtime state (status, timings, logs, outputs) is not carried over.
|
||||
"""
|
||||
with self._jobs_lock:
|
||||
source = self._jobs.get(job_id)
|
||||
if source is None:
|
||||
raise ValueError(f"Job {job_id} not found")
|
||||
config = {f: copy.deepcopy(getattr(source, f)) for f in self.CONFIG_FIELDS}
|
||||
return self.create_job(job_id=new_job_id, **config)
|
||||
|
||||
#: Editable exactly when startable: the same set start_job() accepts.
|
||||
EDITABLE_STATUSES = (JobStatus.PENDING, JobStatus.FAILED, JobStatus.STOPPED)
|
||||
|
||||
def update_job_config(self, job_id: str, updates: dict[str, Any]) -> Job:
|
||||
"""Edit the configuration of a job that has not produced a result."""
|
||||
with self._jobs_lock:
|
||||
job = self._jobs.get(job_id)
|
||||
if job is None:
|
||||
raise ValueError(f"Job {job_id} not found")
|
||||
if job.status not in self.EDITABLE_STATUSES:
|
||||
allowed = ", ".join(s.value for s in self.EDITABLE_STATUSES)
|
||||
raise ValueError(f"Job is {job.status.value}; only {allowed} jobs can be edited. "
|
||||
"Duplicate it instead.")
|
||||
unknown = set(updates) - set(self.CONFIG_FIELDS)
|
||||
if unknown:
|
||||
raise ValueError(f"Not editable: {', '.join(sorted(unknown))}")
|
||||
for field_name, value in updates.items():
|
||||
setattr(job, field_name, value)
|
||||
self._save_job(job)
|
||||
return job
|
||||
|
||||
def start_job(self, job_id: str) -> Job:
|
||||
"""Start (or restart) a pending / stopped / failed job.
|
||||
|
||||
@@ -792,8 +623,6 @@ class JobRunner:
|
||||
workload_type: str,
|
||||
num_gpus: int,
|
||||
dit_cpu_offload: bool = False,
|
||||
dit_layerwise_offload: bool = False,
|
||||
override_pipeline_cls_name: str | None = None,
|
||||
text_encoder_cpu_offload: bool = False,
|
||||
vae_cpu_offload: bool = False,
|
||||
image_encoder_cpu_offload: bool = False,
|
||||
@@ -809,10 +638,6 @@ class JobRunner:
|
||||
workload_type,
|
||||
num_gpus,
|
||||
dit_cpu_offload,
|
||||
dit_layerwise_offload,
|
||||
# Ref2VA loads different DiT weights (transformer_ref), so the
|
||||
# override must key the cache or a t2v/i2v generator gets reused.
|
||||
override_pipeline_cls_name,
|
||||
text_encoder_cpu_offload,
|
||||
vae_cpu_offload,
|
||||
image_encoder_cpu_offload,
|
||||
@@ -825,21 +650,8 @@ class JobRunner:
|
||||
|
||||
# Generators are cached by model_id and configuration parameters
|
||||
with self._generators_lock:
|
||||
cached = self._generators.get(cache_key)
|
||||
if cached is not None:
|
||||
if _generator_is_alive(cached):
|
||||
return cached
|
||||
# Workers can exit while a generator sits idle in the cache;
|
||||
# reusing it fails every later job with the same config.
|
||||
logger.warning(
|
||||
"Cached generator for %s has dead workers; reloading.",
|
||||
model_id,
|
||||
)
|
||||
self._generators.pop(cache_key, None)
|
||||
try:
|
||||
cached.shutdown()
|
||||
except Exception:
|
||||
logger.debug("Shutdown of the dead generator failed", exc_info=True)
|
||||
if cache_key in self._generators:
|
||||
return self._generators[cache_key]
|
||||
|
||||
# Import lazily so starting the server is fast even without a GPU.
|
||||
from fastvideo import VideoGenerator
|
||||
@@ -865,11 +677,6 @@ class JobRunner:
|
||||
gen = VideoGenerator.from_pretrained(
|
||||
model_id,
|
||||
workload_type=workload_type,
|
||||
num_gpus=num_gpus,
|
||||
dit_layerwise_offload=dit_layerwise_offload,
|
||||
**({
|
||||
"override_pipeline_cls_name": override_pipeline_cls_name
|
||||
} if override_pipeline_cls_name else {}),
|
||||
dit_cpu_offload=dit_cpu_offload,
|
||||
text_encoder_cpu_offload=text_encoder_cpu_offload,
|
||||
vae_cpu_offload=vae_cpu_offload,
|
||||
@@ -899,9 +706,10 @@ class JobRunner:
|
||||
def _run_training_job(self, job: Job):
|
||||
"""Run a finetuning, distillation, or LoRA job via subprocess."""
|
||||
buf = job._log_buf
|
||||
os.makedirs(self.log_dir, exist_ok=True)
|
||||
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
|
||||
job_output_dir = os.path.join(self.output_dir, job.id)
|
||||
os.makedirs(job_output_dir, exist_ok=True)
|
||||
job.log_file_path = _job_log_path(self.output_dir, job.id)
|
||||
|
||||
if not job.data_path or not os.path.isdir(job.data_path):
|
||||
job.status = JobStatus.FAILED
|
||||
@@ -1019,8 +827,8 @@ class JobRunner:
|
||||
|
||||
def _run_inference_job(self, job: Job):
|
||||
buf = job._log_buf
|
||||
os.makedirs(os.path.join(self.output_dir, job.id), exist_ok=True)
|
||||
job.log_file_path = _job_log_path(self.output_dir, job.id)
|
||||
os.makedirs(self.log_dir, exist_ok=True)
|
||||
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
|
||||
|
||||
# Add file handler to persist logs
|
||||
file_handler = logging.FileHandler(job.log_file_path, mode='w', encoding='utf-8')
|
||||
@@ -1067,45 +875,62 @@ class JobRunner:
|
||||
buf.phase = "loading model"
|
||||
logger.info("Loading model...")
|
||||
|
||||
# The generator MUST be created on this thread: building it spawns
|
||||
# the executor's worker processes, and they are torn down if the
|
||||
# creating thread exits. Running it in a helper thread (to poll
|
||||
# _stop_event during load) made every collective_rpc fail with
|
||||
# ConnectionResetError.
|
||||
if job._stop_event.is_set():
|
||||
job.status = JobStatus.STOPPED
|
||||
job.finished_at = time.time()
|
||||
self._save_job(job)
|
||||
logger.warning("Job %s stopped before model loading", job.id)
|
||||
buf.phase = "stopped"
|
||||
return
|
||||
# Run generator creation in a background thread so we
|
||||
# can poll _stop_event while the (potentially slow)
|
||||
# model download / load is in progress.
|
||||
_gen_result: list[Any] = []
|
||||
_gen_error: list[BaseException] = []
|
||||
|
||||
generator = self._get_or_create_generator(
|
||||
job.model_id,
|
||||
job.workload_type,
|
||||
job.num_gpus,
|
||||
dit_cpu_offload=job.dit_cpu_offload,
|
||||
dit_layerwise_offload=job.dit_layerwise_offload,
|
||||
override_pipeline_cls_name=(MINIMAX_H3_REF2VA_PIPELINE if job.references else None),
|
||||
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
|
||||
vae_cpu_offload=job.vae_cpu_offload,
|
||||
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
|
||||
use_fsdp_inference=job.use_fsdp_inference,
|
||||
enable_torch_compile=(job.enable_torch_compile),
|
||||
vsa_sparsity=job.vsa_sparsity,
|
||||
tp_size=job.tp_size,
|
||||
sp_size=job.sp_size,
|
||||
log_queue=log_queue,
|
||||
def _load_generator() -> None:
|
||||
try:
|
||||
gen = self._get_or_create_generator(
|
||||
job.model_id,
|
||||
job.workload_type,
|
||||
job.num_gpus,
|
||||
dit_cpu_offload=job.dit_cpu_offload,
|
||||
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
|
||||
vae_cpu_offload=job.vae_cpu_offload,
|
||||
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
|
||||
use_fsdp_inference=job.use_fsdp_inference,
|
||||
enable_torch_compile=(job.enable_torch_compile),
|
||||
vsa_sparsity=job.vsa_sparsity,
|
||||
tp_size=job.tp_size,
|
||||
sp_size=job.sp_size,
|
||||
log_queue=log_queue,
|
||||
)
|
||||
_gen_result.append(gen)
|
||||
except BaseException as exc:
|
||||
_gen_error.append(exc)
|
||||
|
||||
loader = threading.Thread(
|
||||
target=_load_generator,
|
||||
daemon=True,
|
||||
)
|
||||
loader.start()
|
||||
|
||||
while loader.is_alive():
|
||||
if job._stop_event.is_set():
|
||||
job.status = JobStatus.STOPPED
|
||||
job.finished_at = time.time()
|
||||
self._save_job(job)
|
||||
logger.warning(
|
||||
"Job %s stopped during model loading",
|
||||
job.id,
|
||||
)
|
||||
buf.phase = "stopped"
|
||||
return
|
||||
loader.join(timeout=0.5)
|
||||
|
||||
if _gen_error:
|
||||
raise _gen_error[0]
|
||||
|
||||
generator = _gen_result[0]
|
||||
buf.phase = "generating"
|
||||
logger.info("Starting generation for job %s (model=%s)", job.id, job.model_id)
|
||||
|
||||
# Without a name FastVideo derives the filename from the prompt.
|
||||
safe_name = re.sub(r'[\\/:*?"<>|]+', "", job.name).strip().strip(".")
|
||||
output_target = (os.path.join(job_output_dir, f"{safe_name[:80]}.mp4") if safe_name else job_output_dir)
|
||||
gen_kwargs: dict[str, Any] = {
|
||||
"prompt": job.prompt,
|
||||
"output_path": output_target,
|
||||
"output_path": job_output_dir,
|
||||
"save_video": True,
|
||||
"num_inference_steps": job.num_inference_steps,
|
||||
"num_frames": job.num_frames,
|
||||
@@ -1120,12 +945,6 @@ class JobRunner:
|
||||
}
|
||||
if job.image_path:
|
||||
gen_kwargs["image_path"] = job.image_path
|
||||
if job.references:
|
||||
gen_kwargs["references"] = _build_h3_references(job.references)
|
||||
if job.last_image_path:
|
||||
# _prepare_fl2va requires a PIL image, not a path.
|
||||
from PIL import Image as _PILImage
|
||||
gen_kwargs["last_image"] = _PILImage.open(job.last_image_path)
|
||||
generator.generate_video(**gen_kwargs)
|
||||
|
||||
buf.phase = "saving"
|
||||
@@ -1158,7 +977,7 @@ class JobRunner:
|
||||
|
||||
except Exception as exception:
|
||||
error_msg = str(exception)
|
||||
logger.exception("Critical error in job thread: %s", error_msg)
|
||||
logger.error("Critical error in job thread: %s", error_msg)
|
||||
job.status = JobStatus.FAILED
|
||||
job.error = f"Critical error ({type(exception).__name__}): {error_msg}"
|
||||
job.finished_at = time.time()
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user