Compare commits
73
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5a01741ea0 | ||
|
|
de4d430d87 | ||
|
|
633d393568 | ||
|
|
5854aec2ce | ||
|
|
30e45c2411 | ||
|
|
2a4fe697a6 | ||
|
|
921db7479d | ||
|
|
7f539424cb | ||
|
|
19a838f54f | ||
|
|
d922ab2cbc | ||
|
|
9ea77d37f3 | ||
|
|
2e35b0c6bd | ||
|
|
1c627a3f98 | ||
|
|
a931efe33a | ||
|
|
041e5e9029 | ||
|
|
efcc245c2e | ||
|
|
922e7e0813 | ||
|
|
c62a8514b0 | ||
|
|
3eb8081801 | ||
|
|
3505d09564 | ||
|
|
570607945c | ||
|
|
d3a821cdcf | ||
|
|
3f24578139 | ||
|
|
89fcf08378 | ||
|
|
c2930b2aa1 | ||
|
|
019239690b | ||
|
|
d6119c1f82 | ||
|
|
84214c80bb | ||
|
|
c2d7143c72 | ||
|
|
afdb6fbfa5 | ||
|
|
ba4c02d883 | ||
|
|
2c137931f3 | ||
|
|
36682797a0 | ||
|
|
0ef1357a77 | ||
|
|
a75d19786a | ||
|
|
be548a78ea | ||
|
|
6a610e2bc9 | ||
|
|
321d5112b4 | ||
|
|
ba75ad82db | ||
|
|
58caa5109f | ||
|
|
2f3ca8aaad | ||
|
|
3a67319cb6 | ||
|
|
266fa044b3 | ||
|
|
f3398db868 | ||
|
|
fda02036bc | ||
|
|
68179cd752 | ||
|
|
2dd5760291 | ||
|
|
af2ee9c78a | ||
|
|
1c80371b27 | ||
|
|
eef473225d | ||
|
|
44fb84ef6a | ||
|
|
e8597b7448 | ||
|
|
e2252c0a5e | ||
|
|
72cb427cd9 | ||
|
|
63030cf6ec | ||
|
|
773d44b875 | ||
|
|
1df513922f | ||
|
|
30c45620a2 | ||
|
|
6b2c731596 | ||
|
|
e6022c20b2 | ||
|
|
460f6e398e | ||
|
|
d2ffec5cce | ||
|
|
cb12e88713 | ||
|
|
0958c344b8 | ||
|
|
71dc27ea7b | ||
|
|
1263449d2a | ||
|
|
b4de5a9f1b | ||
|
|
8acd8e21f9 | ||
|
|
d45d82334f | ||
|
|
6392bd40a9 | ||
|
|
17f07bc313 | ||
|
|
325861fb99 | ||
|
|
c5088670c8 |
@@ -0,0 +1,41 @@
|
||||
---
|
||||
date: 2026-05-22
|
||||
experiment: PR #1386 DreamVerse app CI backend tests
|
||||
category: infrastructure
|
||||
severity: important
|
||||
---
|
||||
|
||||
# DreamVerse App CI Streaming Imports Need GPU
|
||||
|
||||
## What Happened
|
||||
|
||||
DreamVerse app CI backend pytest collection imports FastVideo streaming surfaces.
|
||||
When those tests run in a CPU-only Modal environment, collection can fail before
|
||||
any app assertions run with Triton reporting:
|
||||
|
||||
```text
|
||||
RuntimeError: 0 active drivers
|
||||
```
|
||||
|
||||
## Root Cause
|
||||
|
||||
Some streaming import paths can import `fastvideo_kernel` at module import time.
|
||||
Triton then probes for an active GPU driver during pytest collection. A CPU-only
|
||||
Modal container has no active driver, so the failure appears as an import-time
|
||||
collection error rather than a DreamVerse app behavior failure.
|
||||
|
||||
## Fix / Workaround
|
||||
|
||||
For PR #1386, use a surgical CI fix: allocate a GPU to
|
||||
`run_dreamverse_app_tests`. Do not refactor core streaming/kernel imports just to
|
||||
unstick this app CI path.
|
||||
|
||||
Keep `build_kernel=False` for this job. The DreamVerse app backend test imports
|
||||
streaming surfaces but does not need to rebuild or exercise custom kernels.
|
||||
|
||||
## Prevention
|
||||
|
||||
When adding or modifying DreamVerse app CI jobs that import FastVideo streaming
|
||||
modules, make the GPU requirement explicit if the import graph may touch
|
||||
`fastvideo_kernel`. Prefer small CI resource fixes for app test collection issues
|
||||
unless the product code genuinely requires lazy import cleanup.
|
||||
@@ -31,7 +31,7 @@ FastVideo-WorldModel/
|
||||
│ │ │ └── distribution_matching/ # DMD2Method, SelfForcingMethod
|
||||
│ │ ├── models/ # Per-role model wrappers (ModelBase, CausalModelBase)
|
||||
│ │ │ ├── wan/ # WanModel, WanCausalModel
|
||||
│ │ │ └── matrixgame/ # MatrixGameModel, MatrixGameCausalModel
|
||||
│ │ │ └── matrixgame2/ # MatrixGame2Model, MatrixGame2CausalModel
|
||||
│ │ ├── callbacks/ # Composable hooks (grad_clip, ema, validation)
|
||||
│ │ └── utils/ # Config, builder, checkpoint, optimizer, tracking
|
||||
│ ├── training/ # Legacy training infrastructure (being phased out)
|
||||
@@ -44,7 +44,7 @@ FastVideo-WorldModel/
|
||||
│ │ ├── wan_distillation_pipeline.py # Wan distillation
|
||||
│ │ ├── self_forcing_distillation_pipeline.py # Self-forcing distill
|
||||
│ │ ├── ltx2_training_pipeline.py # LTX-2 training
|
||||
│ │ └── matrixgame_training_pipeline.py # MatrixGame training
|
||||
│ │ └── matrixgame2_training_pipeline.py # Matrix-Game 2.0 training
|
||||
│ ├── attention/ # Attention backends
|
||||
│ ├── distributed/ # Sequence/tensor parallel utilities
|
||||
│ ├── layers/ # Tensor-parallel layers
|
||||
@@ -95,7 +95,7 @@ FastVideo-WorldModel/
|
||||
| Wan distillation (DMD) | `fastvideo/training/wan_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| Self-forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| MatrixGame | `fastvideo/training/matrixgame_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| Matrix-Game 2.0 | `fastvideo/training/matrixgame2_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
|
||||
## W&B Integration
|
||||
|
||||
|
||||
@@ -8,6 +8,8 @@ follow-up actions see [open-threads.md](open-threads.md).
|
||||
|
||||
**Last updated:** 2026-05-06 (added D-21 — chunk-stutter root cause is software libx264 encoding consuming ~22% of segment wall-time, NOT a migration regression; verified by 3 parallel explore agents that NVFP4 + torch.compile coverage matches FastVideo-internal exactly; landed opt-in NVENC build path in install_native_ffmpeg.sh + `--nvenc`/`--no-nvenc` flag in dreamverse-deploy.sh + `apps/dreamverse/server/benchmarks/benchmark_av_streaming.py` regression test + memory dir update; default codec stays `libx264` for backward compat, opt-in via `--nvenc`. Earlier: added D-12 — GpuPool layer separation, Oracle review post-#1257-merge; added D-13 — prompt enhancer / LLMProvider abstraction shape, Oracle review pre-#1258-merge; added D-14 — streaming auxiliaries cohesion, Oracle review during #1284 review cycle; added D-15 — streaming router placement + sticky/active-active deferral, Oracle review during #1286 review cycle; added D-16 — streaming router polish round 2, second-pass review on top of D-15 covering bridge cancellation hygiene, registry state machine, httpx hard-fail, replica YAML parsing, and `websockets` dep; added D-17 — strategy reversal: abandon 6-PR split in favor of single mega-PR #1288 on `will/ltx2_sr_port`; added D-18 — Option B+ chosen: Dreamverse FE+product-server move into FastVideo as `apps/dreamverse/` subfolder while generic backend stays at `fastvideo.entrypoints.streaming.*`; integration-review.md deprecated, integration-plan.md is the executable migration plan; added D-19 — D-18 executed: 5 commits land on `will/dreamverse-monorepo`, fix-up commits corrected the integration-plan's invalid "delete generic-merged, import public substitutes" assumption — generic-merged files carried product-local instead, e2e passes against migrated code with /proc-verified evidence; added D-20 — segment-2 BrokenPipe root cause was a TWO-direction silent drop of LTX-2 audio kwargs in public `VideoGenerator`).
|
||||
|
||||
**Update 2026-05:** Dreamverse frontend tooling migrated from standalone pnpm to standalone npm. `apps/dreamverse/web/package-lock.json` is authoritative; see PR #1385.
|
||||
|
||||
## Status legend
|
||||
|
||||
- ✅ **Resolved** — decision made and implementation complete (or no implementation needed)
|
||||
@@ -297,14 +299,14 @@ Playwright (8/8 PASS in 5.1s):
|
||||
- **Python ML library** stays at root: `fastvideo/`, `fastvideo-kernel/`.
|
||||
- **Generic backend** stays at `fastvideo.entrypoints.streaming.*` (already there per #1257/#1258/#1284/#1286/#1288).
|
||||
- **Dreamverse product** moves into `apps/dreamverse/{server,web,prompts,serve_configs,scripts}/`.
|
||||
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone pnpm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
|
||||
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone npm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
|
||||
|
||||
**Rationale:**
|
||||
|
||||
- Drops the cross-repo coordination overhead identified in the post-#1286 rebase cycle (D-17 handled by consolidating into mega-PR; D-18 prevents the next round of cross-repo coordination from happening).
|
||||
- Keeps the architectural separation Option D recommended (FastVideo owns reusable runtime; product owns product). The boundary is now `apps/dreamverse/` directory rather than two repos.
|
||||
- Single repo means atomic cross-cutting refactors (e.g. GpuPool API change + Dreamverse adoption) ship as one PR.
|
||||
- OSS precedents support the shape (chainlit uv-workspace + pnpm; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
|
||||
- OSS precedents support the shape (chainlit uv-workspace + frontend package manager; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
|
||||
|
||||
**Why not Option D (separate repos):**
|
||||
|
||||
|
||||
@@ -311,7 +311,7 @@ Notes:
|
||||
development and CI.
|
||||
- Product server release is Docker/deploy workflow, not PyPI.
|
||||
|
||||
### Frontend build: standalone pnpm
|
||||
### Frontend build: standalone npm
|
||||
|
||||
Do not add a root `package.json`.
|
||||
|
||||
@@ -319,14 +319,14 @@ Keep all frontend tooling under:
|
||||
|
||||
```text
|
||||
apps/dreamverse/web/package.json
|
||||
apps/dreamverse/web/pnpm-lock.yaml
|
||||
apps/dreamverse/web/package-lock.json
|
||||
apps/dreamverse/web/playwright.config.*
|
||||
```
|
||||
|
||||
Rationale:
|
||||
|
||||
- FastVideo remains primarily a Python ML library.
|
||||
- Python contributors should not need Node or pnpm for normal work.
|
||||
- Python contributors should not need Node or npm for normal work.
|
||||
- This intentionally diverges from chainlit's root JS workspace pattern and
|
||||
follows the simpler open-webui-style split.
|
||||
|
||||
@@ -455,31 +455,24 @@ jobs:
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
# IMPORTANT: pnpm/action-setup MUST run BEFORE setup-node when using
|
||||
# cache: pnpm — setup-node otherwise can't find pnpm to populate cache.
|
||||
- name: Setup pnpm
|
||||
uses: pnpm/action-setup@v4
|
||||
with:
|
||||
version: 9
|
||||
|
||||
- name: Setup Node
|
||||
uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: '22'
|
||||
cache: pnpm
|
||||
cache-dependency-path: apps/dreamverse/web/pnpm-lock.yaml
|
||||
cache: npm
|
||||
cache-dependency-path: apps/dreamverse/web/package-lock.json
|
||||
|
||||
- name: Install dependencies
|
||||
run: pnpm install --frozen-lockfile
|
||||
run: npm ci
|
||||
|
||||
- name: Typecheck
|
||||
run: pnpm run typecheck --if-present
|
||||
run: npm run typecheck --if-present
|
||||
|
||||
- name: Unit tests
|
||||
run: pnpm run test --if-present
|
||||
run: npm run test --if-present
|
||||
|
||||
- name: Build
|
||||
run: pnpm run build
|
||||
run: npm run build
|
||||
|
||||
# NOTE: Playwright tests require a running backend. Until Phase 4 lands
|
||||
# `/healthz`/`/readyz`/`/status`/`/prompt-system-config`/`/curated-presets`
|
||||
@@ -489,13 +482,13 @@ jobs:
|
||||
# `fastvideo.entrypoints.streaming.build_app` and add product routes).
|
||||
#
|
||||
# - name: Install Playwright browsers
|
||||
# run: pnpm exec playwright install --with-deps chromium
|
||||
# run: npm exec -- playwright install --with-deps chromium
|
||||
#
|
||||
# - name: Playwright (re-enable in Phase 4)
|
||||
# run: pnpm exec playwright test
|
||||
# run: npm exec -- playwright test
|
||||
```
|
||||
|
||||
**Playwright config update needed** when moving FE: `Dreamverse/apps/web/playwright.config.ts` line 39 currently uses `npm run dev`; change to `pnpm run dev` post-move ([source](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39)).
|
||||
**Playwright config update needed** when moving FE: keep `Dreamverse/apps/web/playwright.config.ts` line 39 on `npm run dev` after the move ([source](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39)).
|
||||
|
||||
#### `.github/workflows/ci-dreamverse-backend.yml`
|
||||
|
||||
@@ -604,7 +597,7 @@ monorepo CI path.
|
||||
| uv workspaces | https://docs.astral.sh/uv/concepts/workspaces/ | Authoritative Python workspace model. |
|
||||
| Hatch monorepo | https://hatch.pypa.io/latest/how-to/environment/workspace/ | Alternative workspace model; not selected. |
|
||||
| chainlit | https://github.com/Chainlit/chainlit | uv workspace precedent plus **per-language CI split** (separate `check-frontend.yaml` / `check-backend.yaml` workflows path-filtered by directory). Not a PR-level split — independent of D-17 single-mega-PR decision. |
|
||||
| open-webui | https://github.com/open-webui/open-webui | Frontend path filtering (`paths-ignore` on backend-only changes) and separate release tracks. **Note:** open-webui has a root `package.json`; we are choosing standalone-pnpm despite the precedent, to avoid forcing Python-only contributors to install Node. |
|
||||
| open-webui | https://github.com/open-webui/open-webui | Frontend path filtering (`paths-ignore` on backend-only changes) and separate release tracks. **Note:** open-webui has a root `package.json`; we are choosing standalone npm under `apps/dreamverse/web/` despite the precedent, to avoid forcing Python-only contributors to install Node. |
|
||||
| streamlit | https://github.com/streamlit/streamlit | Split Python and JS testing in one repo. |
|
||||
| gradio | https://github.com/gradio-app/gradio | Python package plus JS workspace precedent. |
|
||||
| full-stack-fastapi-template-nextjs | https://github.com/nemanjam/full-stack-fastapi-template-nextjs | Separate frontend/backend build and deploy workflows. |
|
||||
@@ -623,8 +616,8 @@ class runs where, because not all tests can run on `ubuntu-latest` CI.
|
||||
| **Unit** | (none / `unit`) | `ci-dreamverse-backend.yml` (ubuntu-latest CI) + locally | Pure logic, no GPU, no live service. Mocked FastVideo backends, schema validation, helper functions. | `test_config.py`, `test_rewrite_prompt_payload.py`, `test_session_init_image.py`, the new `test_import_contract.py` |
|
||||
| **Integration (fakes)** | `integration` | `ci-dreamverse-backend.yml` + locally | FastAPI test client + in-process fakes/mocks for GPU pool. Validates routes, request/response shapes, session state machine. | `test_health_endpoints.py`, `test_mock_server.py`, `test_entrypoints.py`, `test_prompt_safety.py`, `test_batching.py` (deleted) |
|
||||
| **Live-service GPU** | `gpu` (skip-by-default in CI) | **Local GPU4 manual QA** + Buildkite-Modal (when added) | Real `fastvideo serve` process + real model weights + real WebSocket round-trips. Validates LTX-2 streaming, NVFP4 wiring, continuation state, frame emission. | `test_realtime_stress.py` (947 LOC), `test_session_logging.py` (1278 LOC) — these spin up real workers per their current shape |
|
||||
| **Frontend unit / build** | (n/a — pnpm) | `ci-dreamverse-frontend.yml` (ubuntu-latest, no GPU) | Vitest + tsc + Next.js build. No backend needed. | `apps/dreamverse/web/src/**/*.test.ts(x)` |
|
||||
| **Frontend Playwright E2E** | (n/a — pnpm) | **Local GPU4 manual QA** until Phase 4 lands public health routes; then `ci-dreamverse-frontend.yml` against a mock backend OR a deployed staging | Real browser → real backend WebSocket flow. Requires `/healthz`, `/readyz`, `/status`, `/prompt-system-config`, `/curated-presets`, `/v1/stream`. | `apps/dreamverse/web/e2e/{backend-health,frontend-shell,preset-prompt-generation}.spec.ts` |
|
||||
| **Frontend unit / build** | (n/a — npm) | `ci-dreamverse-frontend.yml` (ubuntu-latest, no GPU) | Vitest + tsc + Next.js build. No backend needed. | `apps/dreamverse/web/src/**/*.test.ts(x)` |
|
||||
| **Frontend Playwright E2E** | (n/a — npm) | **Local GPU4 manual QA** until Phase 4 lands public health routes; then `ci-dreamverse-frontend.yml` against a mock backend OR a deployed staging | Real browser → real backend WebSocket flow. Requires `/healthz`, `/readyz`, `/status`, `/prompt-system-config`, `/curated-presets`, `/v1/stream`. | `apps/dreamverse/web/e2e/{backend-health,frontend-shell,preset-prompt-generation}.spec.ts` |
|
||||
| **FastVideo public contract** | (none) | Existing FastVideo CI (`ci-precommit` + Buildkite for GPU) | Schema/shape guards that this migration must not break. | `fastvideo/tests/contract/test_dreamverse_shape.py`, `test_dynamo_shape.py`, `test_generate_async.py` |
|
||||
| **FastVideo SSIM regression** | (Buildkite path-filter) | Buildkite-Modal | Inference-quality gates for ported models. | `fastvideo/tests/ssim/test_*.py` |
|
||||
|
||||
@@ -709,7 +702,7 @@ CUDA_VISIBLE_DEVICES=4 uv run --locked --package dreamverse-server \
|
||||
# Smoke-test from another terminal
|
||||
curl -s http://localhost:8009/health | jq .
|
||||
curl -s http://localhost:8009/readyz | jq . # Phase 4+ only
|
||||
# Drive the FE against it: cd apps/dreamverse/web && pnpm run dev
|
||||
# Drive the FE against it: cd apps/dreamverse/web && npm run dev
|
||||
```
|
||||
|
||||
This is the **manual QA gate** for any phase that touches the live-service
|
||||
@@ -907,7 +900,7 @@ mass move. This should be a small PR.
|
||||
6. Add `apps/dreamverse/web/.*` to pre-commit global exclude.
|
||||
7. Add `docs/contributing/dreamverse-development.md` with local dev commands:
|
||||
backend `uv run --locked --package dreamverse-server --extra test pytest ...`;
|
||||
frontend `cd apps/dreamverse/web && pnpm install && pnpm run build`.
|
||||
frontend `cd apps/dreamverse/web && npm ci && npm run build`.
|
||||
8. Run `uv lock` to regenerate `uv.lock` with the new workspace member; commit
|
||||
the lock change in the same PR.
|
||||
9. Land D-12-A docstring caveat: mark `GpuPool` experimental/server-internal
|
||||
@@ -990,7 +983,7 @@ These are **explicit shims** — Phase 4 is responsible for their promotion to p
|
||||
|
||||
- [`config.py:13`](file:///home/william5lin/Dreamverse/server/config.py#L13): `_APP_ROOT / "apps" / "web"` → `_APP_ROOT / "web"` (since `_APP_ROOT` will resolve to `apps/dreamverse/` in the new layout).
|
||||
- [`apps/web/next.config.ts:11`](file:///home/william5lin/Dreamverse/apps/web/next.config.ts#L11): `outputFileTracingRoot: path.resolve(__dirname, "../..")` → `path.resolve(__dirname, "../../..")` (one extra `..` since the FE is one level deeper in the monorepo).
|
||||
- [`playwright.config.ts:39`](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39): `command: "npm run dev"` → `command: "pnpm run dev"` (matches Phase 1 tooling decision).
|
||||
- [`playwright.config.ts:39`](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39): keep `command: "npm run dev"` (matches Phase 1 tooling decision).
|
||||
- Any hardcoded `../FastVideo` paths in scripts/configs → make repo-root-relative since they now share a repo.
|
||||
|
||||
**Steps:**
|
||||
@@ -1084,9 +1077,9 @@ scripts, and product docs after backend tests are green.
|
||||
|
||||
**Verification gate:**
|
||||
|
||||
- `cd apps/dreamverse/web && pnpm install --frozen-lockfile` succeeds.
|
||||
- `cd apps/dreamverse/web && pnpm run build` succeeds.
|
||||
- `cd apps/dreamverse/web && pnpm run test --if-present` succeeds (Vitest + tsc).
|
||||
- `cd apps/dreamverse/web && npm ci` succeeds.
|
||||
- `cd apps/dreamverse/web && npm run build` succeeds.
|
||||
- `cd apps/dreamverse/web && npm run test --if-present` succeeds (Vitest + tsc).
|
||||
- **Frontend CI Playwright is intentionally DEFERRED to Phase 4** — at Phase 3
|
||||
the public `build_app` does not yet expose `/healthz`+`/readyz`+`/status`+
|
||||
`/prompt-system-config`+`/curated-presets`. The `ci-dreamverse-frontend.yml`
|
||||
@@ -1094,7 +1087,7 @@ scripts, and product docs after backend tests are green.
|
||||
reactivates them. No PR note required.
|
||||
- **Manual GPU4 Playwright smoke** (recommended): on this dev node, run
|
||||
`apps/dreamverse/web/e2e/frontend-shell.spec.ts` against the GPU4-deployed
|
||||
backend from Phase 2 manual QA + a `pnpm run dev` frontend at port 5274
|
||||
backend from Phase 2 manual QA + a `npm run dev` frontend at port 5274
|
||||
to confirm shell hydration. `backend-health.spec.ts` and
|
||||
`preset-prompt-generation.spec.ts` will fail until Phase 4 — that is
|
||||
expected; document as deferred.
|
||||
@@ -1249,7 +1242,7 @@ product-specific prompt orchestration.
|
||||
| CI cost increase | Medium | Medium | Add Dreamverse path-specific workflows; add `paths-ignore` to broad workflows; rely on Buildkite monorepo diff watch lists. | CI owner |
|
||||
| FastVideo PyPI release accidentally includes Dreamverse app | High | Low | Add `apps*` to `[tool.setuptools.packages.find]` and `[tool.wheel]` excludes; verify built wheel contents. | Release owner |
|
||||
| Release cadence coupling | Medium | Medium | Keep FastVideo PyPI version release unchanged; Dreamverse uses Docker/Vercel deploy from app paths. | Release owner |
|
||||
| Frontend tooling drift | Medium | Medium | Pin pnpm lockfile in `apps/dreamverse/web/`; no root JS workspace. | Frontend owner |
|
||||
| Frontend tooling drift | Medium | Medium | Pin npm lockfile in `apps/dreamverse/web/`; no root JS workspace. | Frontend owner |
|
||||
| Security surface enlargement | Medium | Medium | Product routes stay in `apps/dreamverse/server`; only generic health/streaming routes go into FastVideo. | Backend owner |
|
||||
| Product-specific API leakage into `fastvideo.*` | High | Medium | Enforce import/module boundary; keep curated presets and prompt UX product-local. | Architecture owner |
|
||||
| Migration regression | High | Medium | Phase gates; backend before frontend; can stop after any phase with Dreamverse repo still usable until Phase 7. | Migration owner |
|
||||
|
||||
@@ -6,7 +6,10 @@ recommended next action.
|
||||
For why each item is open see [decisions-log.md](decisions-log.md). For
|
||||
PR-level context see [pr-roadmap.md](pr-roadmap.md).
|
||||
|
||||
**Last updated:** 2026-05-12 (added DR-4 follow-up for PR #1330's skipped
|
||||
**Last updated:** 2026-05-14 (added DR-7 follow-up to remove LTX2 debug
|
||||
logging env vars and replace them with per-pipeline/model state. Earlier: 2026-05-14 added DR-6 follow-up for PR #1335's Gemma
|
||||
lazy-load/device-placement behavior outside compiled `forward`. Earlier: 2026-05-13 added DR-5 follow-up for PR #1333's LTX2
|
||||
distilled SSIM reference refresh. Earlier: 2026-05-12 added DR-4 follow-up for PR #1330's skipped
|
||||
`App websocket integration` suite. Earlier: 2026-05-05 D-20 broken-pipe root cause + fix landed on
|
||||
`will/dreamverse-monorepo` @ `5eaf0a13`; added new thread D-20-CP for
|
||||
cherry-picking the public-API audio routing fix to `will/ltx2_sr_port`
|
||||
@@ -34,6 +37,9 @@ different vehicle.).
|
||||
| **DR-2** | Med | Decide `cerebras_ifm` provider path: (a) public Literal + `CerebrasIFMProvider` shipped, OR (b) Dreamverse-side custom provider via `enhancer.register_provider(...)` | S (decision) + S-M (impl) | Resolves the cerebras_ifm gap left by PR #1258. Same item as legacy #3 below; DR-2 is the Dreamverse-side framing. |
|
||||
| **DR-3** | Low | Replace Dreamverse `PromptEnhancer._run_blocking_request` manual thread/queue polling with `asyncio.to_thread` after the PR #1327 prompt-enhancer compatibility surface is retired or isolated | S | Review comment #1327 (`prompt_enhancer.py`) is valid, but deferred to avoid patching the local fork in this PR. |
|
||||
| **DR-4** | Low | Investigate unskipping PR #1330's skipped public `App websocket integration` suite | S-M | Public PR #1330 has `describe.skip(...)` around 27 websocket tests while the internal equivalent suite is active with 25 tests. The 2 public-only tests cover backend unreachable / GPU workers not ready. Not blocking while skipped, but stale assertions may need safe refresh before unskip. |
|
||||
| **DR-5** | Low | Regenerate LTX2-Distilled latent SSIM references under the intended neutral/distilled defaults, then remove the PR #1333 historical full-guidance pins | S-M | PR #1333 changed public LTX2 distilled defaults to neutral/distilled values, but existing LTX2 latent SSIM references appear to have been generated with historical full-guidance defaults. The current PR pins the SSIM test to old values to keep CI compatible until references are refreshed. |
|
||||
| **DR-6** | Low | Decide whether `LTX2GemmaTextEncoderModel` needs a non-forward device-placement hook after lazy Gemma load | S | PR #1335 should remove the `model.device` / `model.to(...)` guard from `forward` for Dynamo/fullgraph compatibility. Non-compiled runs probably do not need it because `gemma_model` moves Gemma at first load, but a later wrapper `.to(...)` after lazy load could leave Gemma on the old device unless lifecycle placement handles it. |
|
||||
| **DR-7** | Low | Remove LTX2 debug logging env-var plumbing and replace it with per-pipeline/model debug state | S-M | PR #1335 review flagged that `initialize_pipeline()` mutates process-global LTX2 debug env vars, which can leak/race across pipeline instances. Deleting only the mutation is low-risk but loses config-driven debug logging; deleting all reads without replacement would remove useful SSIM/latent drift diagnostics. |
|
||||
| **3** | Med | Add `cerebras_ifm` to `PromptEnhancerConfig.provider` Literal + provider | S-M | Public-side resolution if DR-2 picks (a) |
|
||||
| **4** | Med | Expose `layer_profile` on typed `engine.quantization` | M | Removes Dreamverse's `experimental["pipeline_config"]` dodge for stage profiles |
|
||||
| **5** | Med | Design typed `dit_config.quant_config` carrier | L design + L impl | Removes broader `experimental["pipeline_config"]` escape hatch |
|
||||
@@ -421,6 +427,65 @@ current PR #1330 review.
|
||||
stack end-to-end and compare the public assertions against the internal active
|
||||
suite.
|
||||
|
||||
### Item DR-6: Gemma lazy-load device placement outside `forward`
|
||||
|
||||
**Why:** PR #1335 review flagged this pattern in
|
||||
`fastvideo/models/encoders/gemma.py::LTX2GemmaTextEncoderModel.forward`:
|
||||
|
||||
```py
|
||||
if model.device != target_device:
|
||||
model.to(device=target_device)
|
||||
```
|
||||
|
||||
The immediate concern is the compiled/Dynamo path: `model.device` and
|
||||
`model.to(...)` inside `forward` can introduce non-tensor/device parsing work
|
||||
that fullgraph tracing should not see. Removing the guard from `forward` is the
|
||||
right PR #1335 review fix.
|
||||
|
||||
For non-compiled execution, the guard is mostly defensive rather than required:
|
||||
`gemma_model` already moves the lazily loaded HF Gemma model to the wrapper's
|
||||
current parameter device on first load. The remaining edge case is a lifecycle
|
||||
sequence where Gemma is loaded, then the parent wrapper is later moved to a
|
||||
different device; in that case Gemma could stay behind unless placement is
|
||||
handled outside `forward`.
|
||||
|
||||
**Action:** After PR #1335 review is unblocked, decide whether FastVideo needs a
|
||||
small lifecycle hook/helper for this class so lazy Gemma is moved whenever the
|
||||
wrapper/device placement changes. If yes, implement it outside `forward`; if no,
|
||||
document that Gemma must be loaded after final device placement.
|
||||
|
||||
**Effort:** Small.
|
||||
|
||||
**Dependencies:** Not blocking PR #1335 if the forward-path guard is removed and
|
||||
the normal load path keeps placing Gemma on the wrapper's current device.
|
||||
|
||||
### Item DR-7: Remove LTX2 debug logging env-var plumbing
|
||||
|
||||
**Why:** PR #1335 review flagged that
|
||||
`fastvideo/pipelines/basic/ltx2/ltx2_pipeline.py::initialize_pipeline()` sets
|
||||
and pops LTX2 debug env vars such as `LTX2_PIPELINE_DEBUG_LOG`,
|
||||
`LTX2_PIPELINE_DEBUG_PATH`, `LTX2_DEBUG_DETAIL`, and
|
||||
`LTX2_PIPELINE_DEBUG_DETAIL_PATH`. Those env vars are process-global, so one
|
||||
pipeline instance can enable, overwrite, or clear debug behavior for another
|
||||
pipeline instance running in the same process.
|
||||
|
||||
Deleting only the `initialize_pipeline()` env mutation is low risk for normal
|
||||
generation, but it would stop config-driven debug logging unless replaced.
|
||||
Deleting all env-var reads without replacement is riskier because these logs are
|
||||
useful for SSIM/latent drift diagnosis and may be used by local debug scripts.
|
||||
|
||||
**Action:** Replace LTX2 debug env-var plumbing with per-pipeline/model debug
|
||||
state. Use pipeline/model config for construction-time hooks and a
|
||||
`ForwardContext`/`ForwardBatch`-style carrier for forward-time logging. Keep
|
||||
external env-var compatibility only if there is a documented operator workflow
|
||||
that still needs it.
|
||||
|
||||
**Effort:** Small-Medium.
|
||||
|
||||
**Dependencies:** Not blocking PR #1335 if the immediate fix is limited to
|
||||
removing process-global mutation from pipeline initialization while preserving
|
||||
existing externally supplied env-var reads.
|
||||
|
||||
### Item D-12-C: Avoid locking `PoolAssignment.gpu_id: int` as public
|
||||
|
||||
**Why:** Today `PoolAssignment` exposes `gpu_id: int`, assuming
|
||||
|
||||
@@ -12,7 +12,7 @@ _Last updated: 2026-03-02_
|
||||
|
||||
| Metric | Category | Status | Location | Trust |
|
||||
|--------|----------|--------|----------|-------|
|
||||
| **FVD** | Distribution | ✅ Implemented | `benchmarks/fvd/` | High |
|
||||
| **FVD** | Distribution | ✅ Implemented | `fastvideo/eval/metrics/common/fvd/` | High |
|
||||
| **SSIM** | Reference | ✅ Implemented | `fastvideo/tests/ssim/` | High |
|
||||
| **LPIPS** | Perceptual | ✅ Implemented | `scripts/lora_extraction/` | Medium |
|
||||
| **Loss trajectory** | Training signal | ✅ Implemented | W&B `train_loss` | Medium |
|
||||
@@ -27,8 +27,8 @@ _Last updated: 2026-03-02_
|
||||
### FVD — Fréchet Video Distance
|
||||
|
||||
**Category**: Distribution-level quality metric
|
||||
**Status**: ✅ Fully implemented in `benchmarks/fvd/`
|
||||
**Trust**: High — standard protocol, I3D feature extractor
|
||||
**Status**: ✅ Registered as the `common.fvd` eval metric in `fastvideo/eval/metrics/common/fvd/`
|
||||
**Trust**: High — standard protocol, I3D feature extractor (CLIP / VideoMAE backbones also available, research-grade)
|
||||
|
||||
#### What It Measures
|
||||
FVD measures the distance between the **distribution** of generated videos and
|
||||
@@ -59,30 +59,39 @@ Lower FVD = generated videos are more statistically similar to real videos.
|
||||
#### How to Use
|
||||
|
||||
```python
|
||||
# Programmatic
|
||||
from benchmarks.fvd import compute_fvd_with_config, FVDConfig
|
||||
# Programmatic — drive the metric directly for custom kwargs
|
||||
from fastvideo.eval import get_metric
|
||||
|
||||
config = FVDConfig.fvd2048_16f() # Standard: 2048 videos, 16 frames
|
||||
results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
|
||||
print(f"FVD: {results['fvd']:.2f}")
|
||||
metric = get_metric("common.fvd", extractor="i3d") # or "clip" / "videomae"
|
||||
metric.to("cuda")
|
||||
metric.setup()
|
||||
metric.reset()
|
||||
|
||||
# First sample carries the reference set; later samples reuse the cache.
|
||||
metric.accumulate({"video": gen_tensors[0], "reference": real_tensors})
|
||||
for gen in gen_tensors[1:]:
|
||||
metric.accumulate({"video": gen})
|
||||
|
||||
result = metric.finalize()
|
||||
print(f"FVD: {result.score:.2f}")
|
||||
```
|
||||
|
||||
```bash
|
||||
# CLI
|
||||
python -m benchmarks.fvd.cli \
|
||||
--real-path data/real/ \
|
||||
--gen-path outputs/gen/ \
|
||||
--protocol fvd2048_16f
|
||||
# CLI — folder of generated mp4s vs a reference folder
|
||||
python examples/inference/eval/eval_fvd.py \
|
||||
--gen-dir outputs/gen/ \
|
||||
--reference-dir data/real/ \
|
||||
--extractor i3d \
|
||||
--output fvd_scores.json
|
||||
```
|
||||
|
||||
**Preset protocols**:
|
||||
| Protocol | Videos | Frames | Use Case |
|
||||
|----------|--------|--------|----------|
|
||||
| `fvd2048_16f` | 2048 | 16 | Standard benchmark (papers) |
|
||||
| `fvd2048_128f` | 2048 | 128 | Long video evaluation |
|
||||
| `quick_test` | 100 | 16 | Fast dev iteration |
|
||||
**Feature extractors**: `i3d` (default, standard FVD spec used in papers),
|
||||
`clip`, `videomae` (research-grade; not directly comparable to published
|
||||
FVD numbers).
|
||||
|
||||
**Feature extractors**: `i3d` (default, standard), `clip`, `videomae`
|
||||
**Protocol**: standard FVD uses 2048 generated + 2048 reference videos at
|
||||
16 frames each. A warning fires below 256 — the score becomes
|
||||
statistically unreliable.
|
||||
|
||||
#### Interpretation
|
||||
| FVD Range | Interpretation |
|
||||
|
||||
@@ -14,7 +14,7 @@ based on **Wan2.1** (SkyReels-V2) DiT models with causal attention for
|
||||
auto-regressive streaming generation.
|
||||
|
||||
**Key techniques you will work with:**
|
||||
- Full finetuning and LoRA on Wan / LTX-2 / MatrixGame models
|
||||
- Full finetuning and LoRA on Wan / LTX-2 / Matrix-Game 2.0 models
|
||||
- DMD-based distillation (few-step generation)
|
||||
- Self-Forcing distillation (causal streaming)
|
||||
- Diffusion-Forcing SFT (DFSFT) for causal models
|
||||
@@ -225,10 +225,10 @@ callbacks:
|
||||
|----------|-----------|----------|
|
||||
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | Standard text-to-video finetune / LoRA |
|
||||
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | Image-to-video (first frame conditioned) |
|
||||
| MatrixGame finetune | `fastvideo/training/matrixgame_training_pipeline.py` | Action-conditioned world model |
|
||||
| MatrixGame AR diffusion | `fastvideo/training/matrixgame_ar_diffusion_pipeline.py` | AR diffusion-forcing training |
|
||||
| MatrixGame ODE-init | `fastvideo/training/matrixgame_ode_causal_pipeline.py` | ODE-trajectory init |
|
||||
| MatrixGame self-forcing distill | `fastvideo/training/matrixgame_self_forcing_distillation_pipeline.py` | Self-forcing distillation |
|
||||
| Matrix-Game 2.0 finetune | `fastvideo/training/matrixgame2_training_pipeline.py` | Action-conditioned world model |
|
||||
| Matrix-Game 2.0 AR diffusion | `fastvideo/training/matrixgame2_ar_diffusion_pipeline.py` | AR diffusion-forcing training |
|
||||
| Matrix-Game 2.0 ODE-init | `fastvideo/training/matrixgame2_ode_causal_pipeline.py` | ODE-trajectory init |
|
||||
| Matrix-Game 2.0 self-forcing distill | `fastvideo/training/matrixgame2_self_forcing_distillation_pipeline.py` | Self-forcing distillation |
|
||||
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | LTX-2 architecture finetuning |
|
||||
| Wan DMD distillation | `fastvideo/training/wan_distillation_pipeline.py` | Few-step distillation via DMD |
|
||||
| Self-Forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | Causal streaming distillation |
|
||||
@@ -258,7 +258,7 @@ Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
|
||||
|--------|-------------|-------|
|
||||
| **Loss trajectory** | Every run, real-time from W&B | Medium |
|
||||
| **SSIM** | When comparing against reference outputs | High |
|
||||
| **FVD** | For benchmarking model quality (`benchmarks/fvd/`) | High |
|
||||
| **FVD** | For benchmarking model quality (`common.fvd` eval metric; example: `examples/inference/eval/eval_fvd.py`) | High |
|
||||
| **LPIPS** | LoRA merge validation | Medium |
|
||||
| **Human preference** | Major checkpoints | Highest |
|
||||
|
||||
@@ -278,8 +278,8 @@ Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
|
||||
|
||||
## World Model–Specific Concepts
|
||||
|
||||
### Action Injection (MatrixGame)
|
||||
The MatrixGame pipeline adds **action modules** to each DiT block, enabling
|
||||
### Action Injection (Matrix-Game 2.0)
|
||||
The Matrix-Game 2.0 pipeline adds **action modules** to each DiT block, enabling
|
||||
frame-level mouse/keyboard input conditioning. The action sequence is injected
|
||||
per-frame alongside the latent video tokens.
|
||||
|
||||
|
||||
@@ -197,6 +197,12 @@ Tolerance guide:
|
||||
Element-wise `assert_close` alone is not enough for deep full-DiT parity. Also
|
||||
log global abs-mean drift and per-modality summaries.
|
||||
|
||||
When a non-skip component parity run is numerically red after weight/input
|
||||
checks, invoke `../add-model-08-trace/SKILL.md` before adding bespoke forward
|
||||
hooks. Use `docs/contributing/activation_trace.md` to keep
|
||||
`FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and `FASTVIDEO_TRACE_STEPS`
|
||||
identical across FastVideo and upstream traces.
|
||||
|
||||
Useful local commands:
|
||||
|
||||
```bash
|
||||
|
||||
@@ -94,8 +94,12 @@ Run the shared parity-debug loop. The component test command is:
|
||||
pytest <parity_test> -v -s
|
||||
```
|
||||
|
||||
For numerical drift, narrow the first divergent block with per-block hooks or
|
||||
intermediate tensor comparisons before changing layers.
|
||||
For numerical drift, use `../add-model-08-trace/SKILL.md` before writing bespoke
|
||||
hooks. Start with FastVideo's activation trace (`fastvideo/hooks/activation_trace.py`;
|
||||
`docs/contributing/activation_trace.md`) and a block-level regex such as
|
||||
`FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$"`. Only fall back to custom
|
||||
per-block hooks if the needed boundary or statistic is not exposed by
|
||||
`FASTVIDEO_TRACE_STATS`.
|
||||
|
||||
## Escape Hatches
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
name: add-model-08-trace
|
||||
description: Use during /add-model Phase 6 when component parity has failed and root cause requires layer-by-layer divergence analysis. Instruments both the official reference and FastVideo port with forward hooks to find the first numerical divergence point.
|
||||
description: Use during /add-model Phase 6 when component parity has failed and root cause requires layer-by-layer divergence analysis. Uses FastVideo activation trace first, falling back to custom hooks only for boundaries or stats the utility cannot observe.
|
||||
---
|
||||
|
||||
# Add-Model Trace
|
||||
@@ -38,104 +38,90 @@ Required inputs before starting:
|
||||
- Shared deterministic test inputs (same tensors on both sides).
|
||||
- The component parity test file path and its current failure output.
|
||||
|
||||
## Hard Rules: Instrumentation Hierarchy
|
||||
## Primary Path: FastVideo Activation Trace
|
||||
|
||||
Apply these in priority order. Use the highest-priority method that works for
|
||||
the target site.
|
||||
Use FastVideo's first-class activation trace before writing custom hooks:
|
||||
`fastvideo/hooks/activation_trace.py`, documented in
|
||||
`docs/contributing/activation_trace.md`.
|
||||
|
||||
### (1) Forward hooks (PREFERRED)
|
||||
Pipeline runs attach trace to the transformer during pipeline initialization.
|
||||
Component-only parity harnesses may call `attach_activation_trace(model)` from
|
||||
local test/debug code; do not add trace calls to production model code.
|
||||
|
||||
Prefix the failing parity command with a tight layer regex:
|
||||
|
||||
```bash
|
||||
FASTVIDEO_TRACE_ACTIVATIONS=1 \
|
||||
FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$" \
|
||||
FASTVIDEO_TRACE_STATS="abs_mean,sum,max,shape" \
|
||||
FASTVIDEO_TRACE_STEPS="0" \
|
||||
FASTVIDEO_TRACE_OUTPUT="/tmp/opencode/fv_trace.jsonl" \
|
||||
pytest tests/local_tests -k "parity" -v -s
|
||||
```
|
||||
|
||||
Match the layer regex to the actual `model.named_modules()` names. Empty or
|
||||
broad regexes are expensive; prefer block-level names first, then narrow to
|
||||
submodules after the first divergent block is known.
|
||||
|
||||
## Trace Compare Contract
|
||||
|
||||
One JSONL file per side. FastVideo output should use `FASTVIDEO_TRACE_OUTPUT`;
|
||||
the upstream harness should emit the same JSONL shape:
|
||||
|
||||
```json
|
||||
{"module":"block.layers.0","tensor":"out","step":0,"abs_mean":0.0123,"sum":1.0,"max":0.5,"shape":[1,16,32]}
|
||||
```
|
||||
|
||||
Compare rows by `(module, step, tensor)`. The first row whose `shape`,
|
||||
`abs_mean`, or `max` diverges beyond the component tolerance is the first broken
|
||||
boundary. Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
|
||||
`FASTVIDEO_TRACE_STEPS` identical between sides; if row order differs, sort or
|
||||
normalize before diffing.
|
||||
|
||||
## Drill-Down Loop
|
||||
|
||||
**Initial run:** trace every top-level block (`^block\.layers\.[0-9]+$` or the
|
||||
family's equivalent). Identify the first block index where `abs_mean` or `max`
|
||||
drifts beyond tolerance while earlier blocks match.
|
||||
|
||||
**Drill run:** tighten `FASTVIDEO_TRACE_LAYERS` to submodules inside the first
|
||||
divergent block: attention output, MLP projections, norm outputs, modality
|
||||
adapters, or other named boundaries exposed by `named_modules()`.
|
||||
|
||||
**Iterate:** if the first divergent operation is a free function or tensor op not
|
||||
visible as an `nn.Module`, use the fallback instrumentation hierarchy below.
|
||||
|
||||
The loop ends when the first divergent submodule or operation is identified with
|
||||
a file:line citation in the official source.
|
||||
|
||||
## Fallback Instrumentation Hierarchy
|
||||
|
||||
Use these only when activation trace cannot observe the needed boundary or
|
||||
statistic.
|
||||
|
||||
### (1) Custom forward hooks
|
||||
|
||||
`module.register_forward_hook(...)` and `register_forward_pre_hook(...)`.
|
||||
Always within `try/finally` with `handle.remove()`. Zero source residue.
|
||||
|
||||
```python
|
||||
handle = module.register_forward_hook(fn)
|
||||
try:
|
||||
output = model(inputs)
|
||||
finally:
|
||||
handle.remove()
|
||||
```
|
||||
|
||||
### (2) Runtime monkey-patch (PREFERRED over source edits)
|
||||
### (2) Runtime monkey-patch
|
||||
|
||||
`module.attr = wrapped_func` or `cls.method = wrapped_method`, restored via
|
||||
`try/finally` (save original first). Use for free functions and non-Module
|
||||
sites such as activation functions (`swiglu`, `apply_rotary_emb`) that cannot
|
||||
be hooked as `nn.Module` submodules.
|
||||
|
||||
```python
|
||||
original = cls.method
|
||||
cls.method = wrapped
|
||||
try:
|
||||
output = model(inputs)
|
||||
finally:
|
||||
cls.method = original
|
||||
```
|
||||
`try/finally` (save original first). Use for free functions and non-Module sites
|
||||
such as activation functions (`swiglu`, `apply_rotary_emb`).
|
||||
|
||||
### (3) Source edits in FastVideo's own code
|
||||
|
||||
Only when (1) and (2) are insufficient. Track all edits within a single named
|
||||
`git stash` boundary OR a temporary branch. Run `git diff` before closing the
|
||||
investigation to confirm the stash or branch is clean. The cleanup gate
|
||||
enforces this.
|
||||
investigation to confirm cleanup.
|
||||
|
||||
### (4) Source edits in official repo source
|
||||
|
||||
Allowed if EITHER:
|
||||
|
||||
- (a) The official repo is a git-tracked clone (e.g. `daVinci-MagiHuman/` at
|
||||
the repo root): use `git diff` in the clone path to verify cleanup.
|
||||
- (b) It's installed editable (`pip install -e .`): use `git diff` in the
|
||||
editable source path to verify cleanup.
|
||||
|
||||
If the official repo is installed non-editable in site-packages: back up the
|
||||
target file (`cp original.py original.py.trace-backup`) before editing, then
|
||||
restore from backup at the end (or `pip install --force-reinstall <pkg>`).
|
||||
The cleanup gate verifies via diff-against-backup or zero-diff-in-clone.
|
||||
|
||||
## Logging Contract
|
||||
|
||||
One log file per side. Paths:
|
||||
|
||||
```
|
||||
/tmp/opencode/<family>_<component>_up_layers.log
|
||||
/tmp/opencode/<family>_<component>_fv_layers.log
|
||||
```
|
||||
|
||||
Format: one line per captured tensor, space-separated:
|
||||
|
||||
```
|
||||
<name> <shape> <abs_mean> <sum> <min> <max>
|
||||
```
|
||||
|
||||
Example:
|
||||
|
||||
```
|
||||
block[00] (1,512,1024) 0.012345 6.3210 -0.4321 0.4321
|
||||
```
|
||||
|
||||
Keep the format diff-friendly. Running `diff /tmp/opencode/x_up.log
|
||||
/tmp/opencode/x_fv.log` should highlight the first divergent line directly.
|
||||
Retain side-by-side stdout output alongside the per-side files for human
|
||||
review.
|
||||
|
||||
## Drill-Down Loop
|
||||
|
||||
**Initial run:** attach hooks to every top-level block (`model.block.layers[i]`
|
||||
or equivalent). Identify the first block index `NN` where abs_mean relative
|
||||
drift exceeds 0.5% compared to the previous block.
|
||||
|
||||
**Drill run:** set `<FAMILY>_DEBUG_DRILL_LAYER=NN` and re-run. The script
|
||||
attaches submodule hooks inside block `NN`: attention output, mlp.pre_norm,
|
||||
mlp.up_gate_proj, mlp.down_proj input (via pre-hook) and output, mlp output,
|
||||
attn_post_norm (if present), mlp_post_norm (if present).
|
||||
|
||||
**Iterate:** if the drill run points to a free function (e.g. an activation
|
||||
not wrapped in an `nn.Module`), switch to a monkey-patch (method 2) to
|
||||
intercept its output via the next module's pre-hook.
|
||||
|
||||
The loop ends when the first divergent submodule is identified with a
|
||||
file:line citation in the official source.
|
||||
Allowed only when hook and monkey-patch approaches cannot capture the site.
|
||||
For git-tracked or editable official clones, use `git diff` in the clone path to
|
||||
verify cleanup. For non-editable site-packages, back up the target file before
|
||||
editing and restore it before handoff.
|
||||
|
||||
## Hypothesis Toggles
|
||||
|
||||
@@ -194,11 +180,13 @@ Escalate to the calling bucket skill when:
|
||||
|
||||
Return to the calling subagent with:
|
||||
|
||||
- File paths to per-side logs (`/tmp/opencode/<family>_<component>_{up,fv}_layers.log`).
|
||||
- The identified first divergent layer or submodule name.
|
||||
- FastVideo trace JSONL path and upstream trace JSONL path.
|
||||
- Trace settings used: `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
|
||||
`FASTVIDEO_TRACE_STEPS`.
|
||||
- The first divergent `(module, step, tensor)` row and observed drift.
|
||||
- The upstream file:line citation where the divergence originates.
|
||||
- Hypothesis verdict if an A/B toggle was used (e.g. "PATCH_LINEAR=1 closes
|
||||
the gap, confirming dtype-cast difference in PackedExpertLinear").
|
||||
- Fallback hook/patch verdict if activation trace could not observe the boundary.
|
||||
- Hypothesis verdict if an A/B toggle was used, for example `PATCH_LINEAR=1`.
|
||||
- Cleanup-gate status: `[cleanup-gate] PASS` or a list of unresolved items.
|
||||
|
||||
The calling agent uses this to scope the production fix in the FastVideo
|
||||
@@ -206,10 +194,14 @@ component file.
|
||||
|
||||
## References
|
||||
|
||||
- `templates/block_trace_debug.py` in this skill directory: the canonical
|
||||
template this skill generalizes.
|
||||
- `docs/contributing/activation_trace.md` for canonical activation-trace env vars,
|
||||
JSONL output, cost model, and troubleshooting.
|
||||
- `fastvideo/hooks/activation_trace.py` for the implementation and
|
||||
`attach_activation_trace(model)` entry point.
|
||||
- `templates/block_trace_debug.py` in this skill directory: fallback custom-hook
|
||||
template when activation trace cannot observe the needed boundary or stat.
|
||||
- `tests/local_tests/transformers/_debug_magi_human_block_parity.py` in the
|
||||
FastVideo3 repo: the worked magi-human example this skill was extracted from.
|
||||
FastVideo3 repo: historical worked example for custom hook/patch debugging.
|
||||
- `add-model/SKILL.md` Phase 6: the calling context for this skill.
|
||||
- `add-model-03-port-dit/SKILL.md`, `add-model-04-port-vae/SKILL.md`,
|
||||
`add-model-05-port-encoder/SKILL.md`, `add-model-06-port-generic/SKILL.md`:
|
||||
|
||||
@@ -157,6 +157,13 @@ Debug pipeline drift in this order:
|
||||
channel order, sample rate, FPS, and final slicing.
|
||||
7. Add targeted stage-level diagnostics to identify the first divergent stage.
|
||||
|
||||
If stage diagnostics show the first bad stage is transformer/denoising or a
|
||||
mid-DiT block, enable activation trace before adding ad hoc pipeline prints; see
|
||||
`docs/contributing/activation_trace.md` and `../add-model-08-trace/SKILL.md`.
|
||||
Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
|
||||
`FASTVIDEO_TRACE_STEPS` identical across reruns so pipeline parity traces diff
|
||||
one-to-one.
|
||||
|
||||
If the first divergence belongs to component implementation, strict loading, or
|
||||
conversion mapping, stop pipeline edits and return `next_step=return_to_phase_6`
|
||||
with the exact failing evidence. Do not patch conversion from this skill.
|
||||
|
||||
@@ -267,6 +267,11 @@ conversion, route it through `../add-model-07-conversion/SKILL.md` with a retry
|
||||
request matching `contracts/conversion_request.md`, then resume the component
|
||||
skill with the updated conversion handoff.
|
||||
|
||||
When a component failure narrows to layer-by-layer numerical drift, load
|
||||
`../add-model-08-trace/SKILL.md` before writing custom hooks. It uses
|
||||
`fastvideo/hooks/activation_trace.py`; canonical env vars and JSONL format are
|
||||
documented in `docs/contributing/activation_trace.md`.
|
||||
|
||||
Phase 6 ends only when every required component handoff reports
|
||||
`parity_status=non_skip_pass`, or when a precise blocker or escape hatch is
|
||||
recorded in `port_state_file`.
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
name: add-reward-model
|
||||
description: Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.
|
||||
---
|
||||
|
||||
# Add Reward Model
|
||||
|
||||
Use for reward models consumed by RL methods.
|
||||
|
||||
## Placement
|
||||
|
||||
- Put reusable reward code under `fastvideo/train/methods/rl/rewards/`.
|
||||
- Expose public builders from `fastvideo/train/methods/rl/rewards/__init__.py`.
|
||||
- Keep method-specific aggregation or advantage logic out of reward classes.
|
||||
|
||||
## Media Inputs
|
||||
|
||||
- Reward callables receive decoded media tensors.
|
||||
- Accept single-frame tensors as `[B, C, H, W]` and multi-frame tensors as `[B, C, T, H, W]` when practical.
|
||||
- Frame selection is reward-specific. Frame scorers such as PickScore and CLIPScore should explicitly select frame `0`; temporal rewards should inspect whichever frames they need.
|
||||
- Return one scalar reward per prompt/sample.
|
||||
|
||||
## Attribution
|
||||
|
||||
- If code is ported or closely adapted from another repo, add a short comment or docstring naming the source file/function.
|
||||
- Preserve SPDX headers used by FastVideo files.
|
||||
|
||||
## Tests
|
||||
|
||||
- Unit-test tensor layout handling without loading large reward checkpoints.
|
||||
- Allow fake scorer injection for multi-reward tests.
|
||||
- Test weighted reward aggregation and metric keys.
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: add-rl-method
|
||||
description: Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.
|
||||
---
|
||||
|
||||
# Add RL Method
|
||||
|
||||
Use for new RL methods in the modular `fastvideo/train` stack.
|
||||
|
||||
## Required Shape
|
||||
|
||||
- Add the method under `fastvideo/train/methods/rl/`.
|
||||
- Subclass `TrainingMethod`.
|
||||
- Keep model-family logic in `ModelBase` wrappers.
|
||||
- Decode generated latents through `ModelBase.decode_latents`; add that hook to the new model wrapper instead of decoding inside the RL method.
|
||||
- Use `fastvideo/train/methods/rl/common/sampling.py` for generation unless the method has a documented reason to avoid sampling.
|
||||
- Use `fastvideo/train/methods/rl/common/prompt_sampling.py` for reusable grouped prompt sampling patterns such as DiffusionNFT K-repeat.
|
||||
- Use `fastvideo/train/methods/rl/rewards/` for reward models.
|
||||
|
||||
## Optimization
|
||||
|
||||
- Return `manages_optimization() == True` only when the method must own a nonstandard outer/inner loop.
|
||||
- If using managed optimization, implement `managed_train_step(data_stream, iteration)`.
|
||||
- Existing trainer callbacks, checkpointing, tracking, and validation should still work.
|
||||
|
||||
## Config
|
||||
|
||||
- Put method knobs under `method`.
|
||||
- Put sampler knobs under `method.sampling`.
|
||||
- Do not put scheduler or trajectory policy into model configs.
|
||||
- Do not split a diffusers-style scheduler from its built-in `step()` solver in YAML; use `trajectory` only for higher-level ODE vs re-noise behavior.
|
||||
- Avoid fixed timestep lists in examples unless reproducing a known baseline; prefer scheduler-generated defaults.
|
||||
|
||||
## Tests
|
||||
|
||||
- Add fake-model tests for sampler/method behavior.
|
||||
- Add config parse tests for the public YAML.
|
||||
- Confirm existing train methods stay on the default Trainer path.
|
||||
@@ -1,3 +1,8 @@
|
||||
---
|
||||
name: dreamverse-deploy
|
||||
description: Use when redeploying the migrated Dreamverse app backend and frontend on a chosen local GPU; tears down existing ports, launches services, and waits for readiness checks.
|
||||
---
|
||||
|
||||
# dreamverse-deploy — redeploy migrated Dreamverse on a chosen GPU
|
||||
|
||||
**Scope:** project (lives in this repo at `.agents/skills/dreamverse-deploy/`)
|
||||
@@ -17,7 +22,7 @@ both `/readyz` and the FE root to return 200.
|
||||
`cerebras-cloud-sdk`, `openai` installed (override the default path with
|
||||
`DREAMVERSE_PYTHON=/path/to/python`)
|
||||
- `~/.env` exporting `CEREBRAS_API_KEY`, `GROQ_API_KEY`, etc.
|
||||
- pnpm installed at `/home/william5lin/.local/share/pnpm/pnpm` (or in `$PATH`)
|
||||
- npm available in `$PATH` (or set `NPM=/path/to/npm`)
|
||||
- `gcc-13` + `g++-13` at `/usr/bin/` (workaround for nvcc gcc-15 rejection)
|
||||
- **Recommended:** native ffmpeg env file at `apps/dreamverse/scripts/ffmpeg-env.sh`
|
||||
(built once via `bash apps/dreamverse/scripts/install_native_ffmpeg.sh`).
|
||||
@@ -105,7 +110,7 @@ Flags can appear in any position relative to the positional args. Explicit flag
|
||||
5. Launches the backend via `apps/dreamverse/scripts/dreamverse-server` in a
|
||||
detached `setsid` session, captures PID.
|
||||
6. Polls `/readyz` until 200 (max 5 min).
|
||||
7. Launches the frontend via `pnpm run dev:devtools` in a detached session,
|
||||
7. Launches the frontend via `npm run dev:devtools` in a detached session,
|
||||
captures PID.
|
||||
8. Polls FE `/` until 200 (max 60s).
|
||||
9. Prints URLs, PIDs, and log paths.
|
||||
@@ -120,7 +125,7 @@ Flags can appear in any position relative to the positional args. Explicit flag
|
||||
PLAYWRIGHT_SKIP_WEBSERVER=1 BACKEND_URL=http://127.0.0.1:8009 \
|
||||
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
|
||||
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
|
||||
pnpm exec playwright test
|
||||
npm exec -- playwright test
|
||||
```
|
||||
The fast suite (8 specs, ~5s) runs by default; the long-running
|
||||
two-segment audio-continuation spec is gated behind
|
||||
@@ -146,7 +151,7 @@ PLAYWRIGHT_SKIP_WEBSERVER=1 \
|
||||
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
|
||||
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
|
||||
PLAYWRIGHT_LONG_RUNNING=1 \
|
||||
pnpm exec playwright test e2e/long-running-segments.spec.ts
|
||||
npm exec -- playwright test e2e/long-running-segments.spec.ts
|
||||
```
|
||||
|
||||
Expected runtime: ~7-9 minutes on a B200 (torch.compile max-autotune
|
||||
|
||||
@@ -185,17 +185,10 @@ CONDA_ENV_PYTHON="${DREAMVERSE_PYTHON:-${HOME}/miniconda3/envs/fv-main/bin/pytho
|
||||
"${CONDA_ENV_PYTHON}" -c 'import flashinfer' 2>/dev/null \
|
||||
|| bail "flashinfer-python not installed in ${CONDA_ENV_PYTHON} (run: ${CONDA_ENV_PYTHON} -m pip install flashinfer-python --no-build-isolation)"
|
||||
|
||||
PNPM="${PNPM:-}"
|
||||
if [[ -n "${PNPM}" ]]; then
|
||||
PNPM_REQUESTED="${PNPM}"
|
||||
PNPM="$(command -v "${PNPM}" 2>/dev/null || true)"
|
||||
[[ -n "${PNPM}" ]] || bail "pnpm not executable or not in PATH: ${PNPM_REQUESTED} (set PNPM to override)"
|
||||
elif [[ -x "${HOME}/.local/share/pnpm/pnpm" ]]; then
|
||||
PNPM="${HOME}/.local/share/pnpm/pnpm"
|
||||
else
|
||||
PNPM="$(command -v pnpm 2>/dev/null || true)"
|
||||
fi
|
||||
[[ -n "${PNPM}" ]] && [[ -x "${PNPM}" ]] || bail "pnpm not found. Set PNPM, install at ${HOME}/.local/share/pnpm/pnpm, or add pnpm to PATH"
|
||||
NPM="${NPM:-npm}"
|
||||
NPM_REQUESTED="${NPM}"
|
||||
NPM="$(command -v "${NPM}" 2>/dev/null || true)"
|
||||
[[ -n "${NPM}" ]] && [[ -x "${NPM}" ]] || bail "npm not executable or not in PATH: ${NPM_REQUESTED} (set NPM to override)"
|
||||
|
||||
GCC13="$(command -v "${GCC13:-gcc-13}" 2>/dev/null || true)"
|
||||
GPP13="$(command -v "${GPP13:-g++-13}" 2>/dev/null || true)"
|
||||
@@ -386,7 +379,7 @@ frontend_log="${LOG_DIR}/frontend-port${FRONTEND_PORT}.log"
|
||||
# requested port differs, run `next dev --port` directly with devtools env.
|
||||
fe_cmd="run dev:devtools"
|
||||
if [[ "${FRONTEND_PORT}" != "5274" ]]; then
|
||||
fe_cmd="exec next dev --port ${FRONTEND_PORT}"
|
||||
fe_cmd="exec -- next dev --port ${FRONTEND_PORT}"
|
||||
fi
|
||||
|
||||
setsid bash -c "
|
||||
@@ -395,7 +388,7 @@ setsid bash -c "
|
||||
export BACKEND_URL=http://127.0.0.1:${BACKEND_PORT}
|
||||
export BACKEND_HOST=127.0.0.1
|
||||
export BACKEND_PORT=${BACKEND_PORT}
|
||||
exec '${PNPM}' ${fe_cmd}
|
||||
exec '${NPM}' ${fe_cmd}
|
||||
" > "${frontend_log}" 2>&1 < /dev/null &
|
||||
disown
|
||||
|
||||
@@ -462,5 +455,5 @@ cat <<SUMMARY
|
||||
BACKEND_URL=http://127.0.0.1:${BACKEND_PORT} \\
|
||||
PLAYWRIGHT_BASE_URL=http://127.0.0.1:${FRONTEND_PORT} \\
|
||||
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \\
|
||||
pnpm exec playwright test
|
||||
npm exec -- playwright test
|
||||
SUMMARY
|
||||
|
||||
@@ -6,4 +6,8 @@
|
||||
{"name": "seed-ssim-references", "description": "Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline", "path": "seed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "reseed-ssim-references", "description": "Re-seed (overwrite) HF reference videos for an existing fastvideo/tests/ssim/ test and a single model id on Modal L40S. Always backs up current refs first, regenerates on Modal, pauses for the user to eyeball before-vs-after, then uploads with --force scoped to --model-id. Sister skill to seed-ssim-references; use when intentional code change has invalidated existing refs", "path": "reseed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "decompose-pipeline-pr", "description": "Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body. Worked example: PR #1280 daVinci-MagiHuman (9.8k LOC) decomposed into 10 stacked PRs.", "path": "decompose-pipeline-pr/SKILL.md", "status": "tested", "trust": "medium"}
|
||||
{"name": "reseed-performance-baseline", "description": "Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline must be advanced by replicating one reviewed shifted source result into three success=true records, or five records when explicitly requested", "path": "reseed-performance-baseline/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-model", "description": "Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.", "path": "add-model/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "rlhf-training-abstractions", "description": "Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.", "path": "rlhf-training-abstractions/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-rl-method", "description": "Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.", "path": "add-rl-method/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-reward-model", "description": "Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.", "path": "add-reward-model/SKILL.md", "status": "draft", "trust": "low"}
|
||||
|
||||
@@ -38,7 +38,7 @@ entrypoint, and applying defaults from the closest example script.
|
||||
| `finetune` (Wan T2V) | `fastvideo/training/wan_training_pipeline.py` |
|
||||
| `finetune` (Wan I2V) | `fastvideo/training/wan_i2v_training_pipeline.py` |
|
||||
| `finetune` (LTX-2) | `fastvideo/training/ltx2_training_pipeline.py` |
|
||||
| `finetune` (MatrixGame) | `fastvideo/training/matrixgame_training_pipeline.py` |
|
||||
| `finetune` (Matrix-Game 2.0) | `fastvideo/training/matrixgame2_training_pipeline.py` |
|
||||
| `distill-dmd` | `fastvideo/training/wan_distillation_pipeline.py` |
|
||||
| `self-forcing` | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` |
|
||||
|
||||
|
||||
@@ -0,0 +1,476 @@
|
||||
---
|
||||
name: reseed-performance-baseline
|
||||
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
|
||||
---
|
||||
|
||||
# Re-seed Performance Baseline
|
||||
|
||||
## Purpose
|
||||
|
||||
Replace or advance the rolling performance baseline for a single
|
||||
`(model_id, gpu_type)` pair in the HF dataset
|
||||
`FastVideo/performance-tracking`.
|
||||
|
||||
Performance comparison uses the median of up to the last 5 successful records
|
||||
for the same model and GPU. Failed records are useful audit history, but they
|
||||
do not move the future baseline because `compare_baseline.py` loads records
|
||||
with `successful_only=True`.
|
||||
|
||||
This skill now reseeds from a reviewed batch of one or more source performance
|
||||
JSONs. It uploads one new `success=true` record per accepted source JSON; it
|
||||
does not blindly replicate one measurement into 3 or 5 records. The effective
|
||||
reseed size is therefore dynamic and equals the number of provided, validated,
|
||||
internally consistent source JSONs.
|
||||
|
||||
If the operator provides fewer than 3 records, call out that the last-5 rolling
|
||||
median may not move immediately. If the operator provides 3 consistent shifted
|
||||
records, the rolling median usually moves immediately. If the operator provides
|
||||
5 consistent shifted records, the last-5 window is effectively reset to the new
|
||||
runtime profile.
|
||||
|
||||
These records are intentional operator-approved baseline resets, not ordinary
|
||||
independent main-branch persistence. Mark them clearly with provenance fields
|
||||
so the HF history remains auditable.
|
||||
|
||||
Use this skill when a performance test fails for an intentional and reviewed
|
||||
reason, such as a torch/runtime/container upgrade that legitimately increases
|
||||
peak memory or changes timings. This is the performance equivalent of
|
||||
`reseed-ssim-references`: backup first, scope tightly, require explicit human
|
||||
approval, then upload reviewed accepted baseline records.
|
||||
|
||||
## When to use
|
||||
|
||||
- A PR or main run failed the rolling performance comparison by more than the
|
||||
allowed regression threshold, and maintainers agree the shift is caused by
|
||||
an intentional runtime, dependency, hardware image, or benchmark environment
|
||||
change rather than a FastVideo logic regression.
|
||||
- One or more shifted source result JSONs have been reviewed and accepted, and
|
||||
the operator wants to use those exact reviewed results to advance the rolling
|
||||
baseline.
|
||||
- The source batch is internally consistent: no provided source JSON regresses
|
||||
against the batch median by more than the configured tolerance.
|
||||
|
||||
## When not to use
|
||||
|
||||
- The benchmark failure might be a real code regression. Fix or investigate
|
||||
the code path first.
|
||||
- The fixed benchmark thresholds in
|
||||
`.buildkite/performance-benchmarks/tests/*.json` are too low. Those are a
|
||||
separate gate from the rolling HF baseline and may need a code review change.
|
||||
- There is no clear source run, commit, and rationale. Baseline history is a
|
||||
production signal; do not edit it without provenance.
|
||||
- The provided source JSONs disagree materially with each other. Rerun or
|
||||
investigate instead of uploading a noisy reseed batch.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
|
||||
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
|
||||
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
|
||||
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `PERF_MAX_REGRESSION` if set, otherwise `0.05` (5%). |
|
||||
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
|
||||
|
||||
Hardcoded defaults:
|
||||
|
||||
- HF repo: `FastVideo/performance-tracking` (`HF_REPO_ID` override is
|
||||
supported by the code, but use the default unless the user explicitly asks).
|
||||
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
|
||||
is supported).
|
||||
- Backup root: `/tmp/performance_reseed_backup`.
|
||||
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
|
||||
- Baseline window: last 5 `success=true` records for the same
|
||||
`(model_id, gpu_type)`.
|
||||
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
|
||||
source JSON.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Validate the target and source results
|
||||
|
||||
Normalize `source_results` to a list. If the user passes a single
|
||||
`source_result`, treat it as a one-element `source_results` list and report
|
||||
that a single record may not move the last-5 median immediately.
|
||||
|
||||
If any source result is a Buildkite artifact URL, download it first into a
|
||||
local scratch directory under `/tmp/performance_reseed_source/` and use the
|
||||
downloaded JSON path for the rest of the workflow. If the agent cannot access
|
||||
the artifact because Buildkite authentication is missing, ask the user to
|
||||
download the artifact manually and provide the local path.
|
||||
|
||||
Prefer the normalized Buildkite artifact emitted by `compare_baseline.py`:
|
||||
|
||||
```text
|
||||
perf_reports/results/normalized_perf_*.json
|
||||
```
|
||||
|
||||
That file is already in the HF tracking schema. Load each normalized JSON
|
||||
directly:
|
||||
|
||||
```python
|
||||
import json
|
||||
|
||||
with open(source_result, encoding="utf-8") as f:
|
||||
record = json.load(f)
|
||||
```
|
||||
|
||||
Stop if any normalized record's `model_id` or `gpu_type` does not match the
|
||||
requested `model_id` and `gpu_type`.
|
||||
|
||||
The source records may have `success: false` when they came from failed
|
||||
rolling baseline comparisons. That is expected; only the reviewed reseed
|
||||
records become new `success: true` baseline records after explicit approval.
|
||||
|
||||
Sort validated source records by their original `timestamp` ascending before
|
||||
preparing the seed records. If a source timestamp is missing or unparsable,
|
||||
preserve input order for those records and print a warning. This makes the
|
||||
fresh reseed timestamps deterministic and makes it clear which records enter
|
||||
the last-5 window when more than 5 source JSONs are provided.
|
||||
|
||||
Check that `HF_API_KEY` is exported. The sync path may be public, but the
|
||||
upload path requires write access.
|
||||
|
||||
### 1a. Check source batch consistency
|
||||
|
||||
Before syncing or preparing uploads, reject source batches that are internally
|
||||
inconsistent. Use the same metric direction as `compare_baseline.py`:
|
||||
|
||||
- Lower is better: `latency`, `memory`, `text_encoder_time_s`, `dit_time_s`,
|
||||
`vae_decode_time_s`.
|
||||
- Higher is better: `throughput`.
|
||||
|
||||
For each metric with at least two non-null source values:
|
||||
|
||||
1. Compute the source batch median.
|
||||
2. For lower-is-better metrics, compute `(source_value - batch_median) / batch_median`.
|
||||
3. For `throughput`, compute `(batch_median - source_value) / batch_median`.
|
||||
4. Stop if any source record regresses against the batch median by more than
|
||||
`max_intra_batch_regression`.
|
||||
|
||||
Default `max_intra_batch_regression` to `PERF_MAX_REGRESSION` when set,
|
||||
otherwise `0.05`. Print a table with per-source values, batch median, and
|
||||
worst intra-batch regression.
|
||||
|
||||
This check prevents uploading a mixed batch where one JSON is materially
|
||||
slower or faster than the others. If the batch fails this check, ask the user
|
||||
to provide a cleaner batch or explicitly investigate the variance. Do not
|
||||
silently drop outliers unless the user gives a concrete reviewed reason and a
|
||||
new source list.
|
||||
|
||||
### 1b. How to obtain source results from CI
|
||||
|
||||
The performance CI exports normalized source results for failed rolling
|
||||
baseline comparisons when `compare_baseline.py` ran. The preferred artifacts
|
||||
come from:
|
||||
|
||||
```text
|
||||
perf_reports/results/normalized_perf_*.json
|
||||
```
|
||||
|
||||
The normal operator flow is:
|
||||
|
||||
1. Open the failed Buildkite performance job or several reruns of the same
|
||||
benchmark after the accepted environment shift.
|
||||
2. Download the `normalized_perf_*.json` artifacts for the target benchmark.
|
||||
3. Pass all reviewed local paths or artifact URLs as `source_results`.
|
||||
|
||||
Do not scrape the Markdown performance summary to reconstruct JSON. The
|
||||
normalized JSON artifacts are the only supported source of truth for reseed
|
||||
metrics and provenance. Raw `fastvideo/tests/performance/results/perf_*.json`
|
||||
artifacts are not accepted by this skill. If no normalized JSON artifact is
|
||||
present, that run is not a valid source for baseline reseeding.
|
||||
|
||||
### 2. Sync and back up existing HF records under /tmp
|
||||
|
||||
Use `fastvideo/tests/performance/hf_store.py` helpers directly. Do **not** use
|
||||
`compare_baseline.py` as a sync shortcut; on full main runs it can persist
|
||||
records, while this step must only fetch and back up existing history.
|
||||
|
||||
The sync command pattern is:
|
||||
|
||||
```bash
|
||||
export PERFORMANCE_TRACKING_ROOT="${PERFORMANCE_TRACKING_ROOT:-/tmp/perf-tracking}"
|
||||
export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
|
||||
PYTHONPATH=fastvideo/tests/performance python -c 'from hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
|
||||
```
|
||||
|
||||
Then back up only the sanitized model directory under `/tmp`:
|
||||
|
||||
```bash
|
||||
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
|
||||
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
|
||||
MODEL_SAFE=$(PYTHONPATH=fastvideo/tests/performance python - <<'PY'
|
||||
from hf_store import sanitize
|
||||
print(sanitize("<model_id>"))
|
||||
PY
|
||||
)
|
||||
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_${MODEL_SAFE}"
|
||||
mkdir -p "$BACKUP_DIR"
|
||||
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
|
||||
```
|
||||
|
||||
Write provenance next to the backup:
|
||||
|
||||
```bash
|
||||
cat > "$BACKUP_DIR/PROVENANCE.txt" <<EOF
|
||||
model_id: <model_id>
|
||||
gpu_type: <gpu_type>
|
||||
source_results:
|
||||
- <source_result_1>
|
||||
- <source_result_2>
|
||||
reseed_record_count: <len(source_results)>
|
||||
max_intra_batch_regression: <threshold>
|
||||
head_commit: $(git rev-parse HEAD)
|
||||
timestamp_utc: $(date -u +%FT%TZ)
|
||||
reason: <intent_rationale>
|
||||
EOF
|
||||
```
|
||||
|
||||
If the backup has no prior records, this is not a destructive reseed; it is a
|
||||
first baseline seed. Continue, but report that baseline history was empty.
|
||||
|
||||
### 3. Compute old baseline and candidate shift
|
||||
|
||||
Load the last 5 successful records for the target:
|
||||
|
||||
```python
|
||||
from hf_store import load_records_for_model
|
||||
|
||||
records = load_records_for_model(
|
||||
"/tmp/perf-tracking",
|
||||
"<model_id>",
|
||||
"<gpu_type>",
|
||||
last_n=5,
|
||||
successful_only=True,
|
||||
)
|
||||
```
|
||||
|
||||
Print a small table showing old medians, source batch medians, candidate
|
||||
medians after appending the proposed seed records, and source batch spread for:
|
||||
|
||||
- `latency`
|
||||
- `throughput`
|
||||
- `memory`
|
||||
- `text_encoder_time_s`
|
||||
- `dit_time_s`
|
||||
- `vae_decode_time_s`
|
||||
|
||||
Also print how many successful old records exist. Make clear:
|
||||
|
||||
- 1 seed record usually does not move a last-5 median by itself.
|
||||
- 3 consistent seed records usually move the last-5 median immediately.
|
||||
- 5 consistent seed records effectively reset the last-5 window.
|
||||
- The records are intentional approved baseline resets and must be labeled
|
||||
that way.
|
||||
|
||||
### 4. Confirm intent
|
||||
|
||||
Require an explicit confirmation phrase before preparing the upload:
|
||||
|
||||
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
|
||||
> This will upload `<N>` new `success=true` records to
|
||||
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
|
||||
> source JSON.
|
||||
>
|
||||
> Reason: `<intent_rationale>`
|
||||
> Source results: `<source_results>`
|
||||
> Reseed record count: `<N>`
|
||||
> Max intra-batch regression: `<threshold>`
|
||||
> Note: these records come from a reviewed source batch and are intended to
|
||||
> move the rolling median to the accepted runtime profile. They are not
|
||||
> ordinary main-branch persistence.
|
||||
> HEAD: `<git rev-parse --short=12 HEAD>`
|
||||
> Backup: `<BACKUP_DIR>`
|
||||
>
|
||||
> Reply `confirm performance reseed` to proceed, anything else to abort.
|
||||
|
||||
Do not continue unless the user types exactly `confirm performance reseed`.
|
||||
|
||||
### 5. Create the accepted seed records
|
||||
|
||||
Create one seed record from each normalized source result. Do not copy the
|
||||
source JSON wholesale.
|
||||
|
||||
Infer the baseline field allowlist from all existing HF records for the target
|
||||
`(model_id, gpu_type)` after syncing, including both `success=true` and
|
||||
`success=false` records. Use the union of non-provenance keys present in those
|
||||
target records, preserving only fields that also exist in the normalized
|
||||
source record or are explicitly set by the reseed workflow. Always include
|
||||
`model_id`, `timestamp`, and `success` because the upload path and baseline
|
||||
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
|
||||
`success` to `true`. Do not include unrelated source-only fields that are
|
||||
absent from existing HF records.
|
||||
|
||||
Exclude existing provenance or operator metadata from the inferred baseline
|
||||
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
|
||||
any fields known to be local-only audit metadata.
|
||||
|
||||
If there are no previous HF records for the target model/GPU, fall back to this
|
||||
default baseline field list:
|
||||
|
||||
- `model_id`
|
||||
- `timestamp`
|
||||
- `commit_sha`
|
||||
- `gpu_type`
|
||||
- `latency`
|
||||
- `throughput`
|
||||
- `memory`
|
||||
- `text_encoder_time_s`
|
||||
- `dit_time_s`
|
||||
- `vae_decode_time_s`
|
||||
- `success`
|
||||
|
||||
Do not upload extra fields from the source artifact.
|
||||
|
||||
Optional provenance fields are allowed and useful:
|
||||
|
||||
- `baseline_reseed: true`
|
||||
- `baseline_reseed_reason`
|
||||
- `baseline_reseed_source_result`
|
||||
- `baseline_reseed_source_timestamp`
|
||||
- `baseline_reseed_source_success`
|
||||
- `baseline_reseed_batch_size`
|
||||
- `baseline_reseed_batch_index`
|
||||
- `baseline_reseed_operator`
|
||||
- `baseline_reseed_max_intra_batch_regression`
|
||||
|
||||
Use a fresh reseed timestamp for each seed record, not the original source
|
||||
result timestamp. This is required because
|
||||
`load_records_for_model(..., last_n=5)` keeps the last records after loading
|
||||
the model directory; stale filenames/timestamps may not enter the last-5
|
||||
window and therefore may not move the median. Preserve the original source
|
||||
timestamp in `baseline_reseed_source_timestamp`.
|
||||
|
||||
Use the existing filename convention from `_write_tracking_record()`:
|
||||
`<sanitize(timestamp)>_<sanitize(commit_sha)>.json` under the sanitized model
|
||||
directory, but include a deterministic suffix such as `_reseed_01`,
|
||||
`_reseed_02`, and so on before `.json` so multiple records from the same
|
||||
batch do not overwrite each other.
|
||||
|
||||
If a source record already exists on HF with `success=false`, do not edit it
|
||||
in place unless the user explicitly asked for an audit-preserving correction.
|
||||
Prefer uploading new accepted seed records so failed history remains visible.
|
||||
|
||||
### 6. Pause before upload
|
||||
|
||||
Print:
|
||||
|
||||
- Backup directory path under `/tmp`.
|
||||
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
|
||||
- HF paths that will receive the new records.
|
||||
- Old rolling medians.
|
||||
- Source batch medians, source batch spread, reseed count, and candidate
|
||||
medians.
|
||||
- Rationale.
|
||||
|
||||
Ask the user to reply exactly `upload`. Anything else aborts and leaves the
|
||||
prepared records plus backup on disk.
|
||||
|
||||
### 7. Upload only the scoped records
|
||||
|
||||
Use the shared storage helper so the path and repo type match CI:
|
||||
|
||||
```python
|
||||
from hf_store import upload_record
|
||||
|
||||
upload_record("<local_record_path>", record, strict=True)
|
||||
```
|
||||
|
||||
Run it once per prepared record. Each upload goes to:
|
||||
|
||||
```text
|
||||
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
|
||||
```
|
||||
|
||||
Never bulk upload the whole tracking root. Never modify another model's
|
||||
directory in the same operation.
|
||||
|
||||
### 8. Report outcome and offer cleanup
|
||||
|
||||
Report:
|
||||
|
||||
- Uploaded HF paths.
|
||||
- Backup directory under `/tmp`.
|
||||
- Local tracking root, usually `/tmp/perf-tracking`.
|
||||
- Old baseline window count and medians.
|
||||
- Source batch medians, source batch spread, reseed count, and candidate
|
||||
medians.
|
||||
- Expected effect based on reseed count.
|
||||
- Any separate threshold changes still needed in
|
||||
`.buildkite/performance-benchmarks/tests/*.json`.
|
||||
|
||||
Include the `intent_rationale` in the PR or follow-up comment so reviewers can
|
||||
distinguish an accepted baseline shift from a hidden regression.
|
||||
|
||||
After the upload is verified, ask whether the user wants to clear temporary
|
||||
local state. Explain what each directory is for:
|
||||
|
||||
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
|
||||
mirror of `FastVideo/performance-tracking` plus the prepared local seed
|
||||
records used for scoped upload.
|
||||
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
|
||||
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
|
||||
audited or corrected.
|
||||
- `/tmp/performance_reseed_source/<...>` when used: downloaded source JSON
|
||||
artifacts from Buildkite URLs.
|
||||
|
||||
Ask:
|
||||
|
||||
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
|
||||
> source downloads, and reseed backup under `/tmp`? These files are local
|
||||
> safety/audit artifacts only; HF already has the uploaded records.
|
||||
>
|
||||
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
|
||||
|
||||
Do not delete anything unless the user replies exactly
|
||||
`cleanup reseed temp`. If cleanup is requested, remove only the specific
|
||||
directories created for this reseed. Never remove unrelated `/tmp` contents.
|
||||
|
||||
## Failure modes and handling
|
||||
|
||||
- **`HF_API_KEY` unset.** Stop before upload. Do not create an untracked
|
||||
process that appears to have reseeded but never reached HF.
|
||||
- **Source result does not match target.** Stop. The wrong benchmark or GPU
|
||||
would poison a separate baseline.
|
||||
- **Source batch is internally inconsistent.** Stop if any source regresses
|
||||
against the source batch median by more than `max_intra_batch_regression`.
|
||||
Ask for cleaner sources or a reviewed explanation before continuing.
|
||||
- **Too few source records to move the median.** Continue only after making
|
||||
clear that one or two records may not immediately move the last-5 median.
|
||||
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
|
||||
those measurements into the baseline, so they must be reviewed first.
|
||||
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
|
||||
make the old baseline look missing.
|
||||
- **Candidate still violates fixed thresholds.** Report that this skill only
|
||||
handles the rolling HF baseline; update benchmark JSON thresholds in code
|
||||
review if maintainers accept the new absolute limit.
|
||||
- **The user aborts at either confirmation.** Leave the backup and prepared
|
||||
records on disk. Nothing should be uploaded.
|
||||
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
|
||||
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
|
||||
place for audit/debugging.
|
||||
- **A bad seed was uploaded.** Use the backup and HF history to identify the
|
||||
uploaded file, then remove or supersede it with an explicitly reviewed
|
||||
corrective record. Do not silently rewrite unrelated history.
|
||||
|
||||
## References
|
||||
|
||||
- `.agents/skills/reseed-ssim-references/SKILL.md` — safety pattern for
|
||||
intentional baseline replacement.
|
||||
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
|
||||
median comparison, and persistence rules.
|
||||
- `fastvideo/tests/performance/hf_store.py` — HF sync, record loading,
|
||||
`sanitize()`, and `upload_record()`.
|
||||
- `fastvideo/tests/performance/test_inference_performance.py` — source result
|
||||
JSON schema.
|
||||
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
|
||||
thresholds, separate from rolling baseline comparisons.
|
||||
|
||||
## Changelog
|
||||
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
|
||||
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
|
||||
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
name: rlhf-training-abstractions
|
||||
description: Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.
|
||||
---
|
||||
|
||||
# RLHF Training Abstractions
|
||||
|
||||
Use this skill before editing RLHF-style training code in `fastvideo/train`.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- RL methods live under `fastvideo/train/methods/rl/` and own algorithm logic: reward collection, advantage computation, policy loss, KL/reference terms, and optimizer cadence.
|
||||
- Rewards live under `fastvideo/train/methods/rl/rewards/` and must be reusable across RL methods.
|
||||
- RL methods pass decoded media to rewards; each reward decides whether to use the first frame, sampled frames, or the full video.
|
||||
- Sampling lives under `fastvideo/train/methods/rl/common/` and must use `ModelBase` primitives plus scheduler math, not model-family inference pipelines.
|
||||
- Model wrappers under `fastvideo/train/models/` own model-specific forward details.
|
||||
- Model wrappers also own model-specific latent decoding via `ModelBase.decode_latents`; RL methods should not reach into VAE normalization internals.
|
||||
- Shared RL helpers such as K-repeat prompt sampling belong under `fastvideo/train/methods/rl/common/` when they are reusable across RL methods.
|
||||
|
||||
## Anti-Patterns
|
||||
|
||||
- Do not bind RL methods to inference pipeline classes such as `WanDMDPipeline`.
|
||||
- Do not hardcode timestep lists in a method when the scheduler can generate them.
|
||||
- Do not put reward-model code inside one RL method.
|
||||
- Do not make existing non-RL methods use method-managed optimization unless explicitly requested.
|
||||
|
||||
## Sampling Policy
|
||||
|
||||
- Prefer YAML-configured `method.sampling` with `scheduler`, `trajectory`, `num_steps`, `timesteps`, and `sigmas`.
|
||||
- Treat diffusers-style scheduler classes as owning both the timestep schedule and their `step()` update rule; avoid a separate `solver` field unless a new sampler truly implements solver math outside the scheduler object.
|
||||
- Missing `timesteps` means “ask the scheduler”; explicit `timesteps` or `sigmas` are overrides.
|
||||
- ODE-style trajectories should not re-noise between denoising steps.
|
||||
- SDE/re-noise behavior must be explicit in config.
|
||||
|
||||
## Validation
|
||||
|
||||
- Run focused local tests for sampler config and Trainer opt-in behavior.
|
||||
- Verify existing train methods still report `manages_optimization() == False`.
|
||||
- Keep fixed-prompt validation helpers in `fastvideo/train/methods/rl/common/validation.py` so new RL methods can reuse sharding and captions.
|
||||
- Test distributed prompt grouping helpers separately from heavyweight model loading.
|
||||
- Run `pre-commit run --files <changed paths>`; respect configured excludes.
|
||||
@@ -89,5 +89,4 @@ The following land in follow-up PRs:
|
||||
|
||||
- **MIND** metrics (depends on a separate `vipe` submodule).
|
||||
- **VBench-2.0** sibling package.
|
||||
- Native conversion of **FVD** under `fastvideo/eval/metrics/fvd/`.
|
||||
- The training-time `EvalCallback`.
|
||||
|
||||
@@ -36,7 +36,10 @@
|
||||
"thresholds": {
|
||||
"L40S": {
|
||||
"max_generation_time_s": 34.0,
|
||||
"max_peak_memory_mb": 11000.0
|
||||
"max_peak_memory_mb": 11000.0,
|
||||
"max_text_encoder_time_s": 5.0,
|
||||
"max_dit_time_s": 10.0,
|
||||
"max_vae_decode_time_s": 10.0
|
||||
},
|
||||
"default": {
|
||||
"max_generation_time_s": 120.0,
|
||||
|
||||
@@ -77,6 +77,17 @@ steps:
|
||||
limit: 2
|
||||
agents:
|
||||
queue: "default"
|
||||
- label: ":microscope: DreamVerse App Tests"
|
||||
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
|
||||
command: "timeout 90m .buildkite/scripts/pr_test.sh"
|
||||
retry:
|
||||
automatic:
|
||||
- exit_status: 128
|
||||
limit: 3
|
||||
- exit_status: -1
|
||||
limit: 2
|
||||
agents:
|
||||
queue: "default"
|
||||
|
||||
# --- Full-suite-scope direct tests ---
|
||||
- label: ":bar_chart: SSIM Tests"
|
||||
@@ -289,6 +300,16 @@ steps:
|
||||
- TEST_TYPE=unit_test
|
||||
agents:
|
||||
queue: "default"
|
||||
- path:
|
||||
- "apps/dreamverse/**"
|
||||
- "pyproject.toml"
|
||||
config:
|
||||
command: "timeout 30m .buildkite/scripts/pr_test.sh"
|
||||
label: ":microscope: DreamVerse App Tests"
|
||||
env:
|
||||
- TEST_TYPE=dreamverse_app
|
||||
agents:
|
||||
queue: "default"
|
||||
|
||||
# ============================================================
|
||||
# Full Suite: Runs when TEST_SCOPE=full
|
||||
|
||||
@@ -121,6 +121,19 @@ upload_performance_artifacts() {
|
||||
fi
|
||||
}
|
||||
|
||||
_upload_normalized_perf_results() {
|
||||
local found=0
|
||||
while IFS= read -r -d '' target; do
|
||||
found=1
|
||||
log "Found normalized performance result: $target. Uploading to Buildkite..."
|
||||
buildkite-agent artifact upload "$target"
|
||||
done < <(find "$LOCAL_DIR" -path "*/results/normalized_perf_*.json" -print0)
|
||||
|
||||
if [ "$found" -eq 0 ]; then
|
||||
log "No normalized performance result artifacts found. This is expected when the rolling performance comparison did not run."
|
||||
fi
|
||||
}
|
||||
|
||||
_cleanup_modal_volume() {
|
||||
log "Cleaning up perf_reports/ from Modal Volume..."
|
||||
if modal volume rm hf-model-weights "perf_reports/" --recursive; then
|
||||
@@ -139,6 +152,7 @@ upload_performance_artifacts() {
|
||||
_download_reports || { _cleanup_local; return 1; }
|
||||
_upload_dashboard
|
||||
_upload_perf_summary
|
||||
_upload_normalized_perf_results
|
||||
_cleanup_modal_volume
|
||||
_cleanup_local
|
||||
}
|
||||
@@ -197,6 +211,10 @@ case "$TEST_TYPE" in
|
||||
log "Running unit tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_unit_test"
|
||||
;;
|
||||
"dreamverse_app")
|
||||
log "Running DreamVerse app tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_dreamverse_app_tests"
|
||||
;;
|
||||
"train_framework")
|
||||
log "Running fastvideo.train framework tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_train_framework_tests"
|
||||
|
||||
@@ -12,6 +12,18 @@ on:
|
||||
tag_suffix:
|
||||
required: true
|
||||
type: string
|
||||
image_name:
|
||||
required: false
|
||||
type: string
|
||||
default: fastvideo-dev
|
||||
build_args:
|
||||
required: false
|
||||
type: string
|
||||
default: ''
|
||||
include_latest_tags:
|
||||
required: false
|
||||
type: boolean
|
||||
default: true
|
||||
|
||||
jobs:
|
||||
build-and-push:
|
||||
@@ -67,11 +79,14 @@ jobs:
|
||||
run: |
|
||||
SHORT_SHA=$(echo ${{ github.sha }} | cut -c1-7)
|
||||
|
||||
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest"
|
||||
TAGS="${TAGS}\ntype=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
|
||||
|
||||
TAGS="type=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
|
||||
|
||||
if [[ "${{ inputs.include_latest_tags }}" == "true" ]]; then
|
||||
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest\n${TAGS}"
|
||||
fi
|
||||
|
||||
# Set Python 3.10 as the default image
|
||||
if [[ "${{ inputs.python_version }}" == "3.10" ]]; then
|
||||
if [[ "${{ inputs.include_latest_tags }}" == "true" && "${{ inputs.python_version }}" == "3.10" ]]; then
|
||||
TAGS="${TAGS}\ntype=raw,value=latest"
|
||||
fi
|
||||
|
||||
@@ -85,7 +100,7 @@ jobs:
|
||||
id: meta
|
||||
uses: docker/metadata-action@v5
|
||||
with:
|
||||
images: ghcr.io/${{ github.repository }}/fastvideo-dev
|
||||
images: ghcr.io/${{ github.repository }}/${{ inputs.image_name }}
|
||||
tags: ${{ steps.prepare-tags.outputs.tags }}
|
||||
|
||||
- name: Build and push Docker image
|
||||
@@ -97,10 +112,11 @@ jobs:
|
||||
push: true
|
||||
tags: ${{ steps.meta.outputs.tags }}
|
||||
labels: ${{ steps.meta.outputs.labels }}
|
||||
build-args: ${{ inputs.build_args }}
|
||||
cache-from: type=gha
|
||||
cache-to: type=gha,mode=max
|
||||
|
||||
- name: Success message
|
||||
run: |
|
||||
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/fastvideo-dev:${{ inputs.tag_suffix }}-latest"
|
||||
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
|
||||
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/${{ inputs.image_name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
|
||||
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
|
||||
|
||||
@@ -125,7 +125,7 @@ jobs:
|
||||
set -euo pipefail
|
||||
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
|
||||
|
||||
VALID="encoder vae transformer kernel unit ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework full fastcheck pre-commit"
|
||||
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework full fastcheck pre-commit"
|
||||
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
|
||||
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
|
||||
exit 1
|
||||
@@ -133,7 +133,7 @@ jobs:
|
||||
|
||||
declare -A MAP=(
|
||||
[encoder]=encoder [vae]=vae [transformer]=transformer
|
||||
[kernel]=kernel_tests [unit]=unit_test
|
||||
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
|
||||
[ssim]=ssim [training]=training
|
||||
[lora-inference]=inference_lora [lora-training]=training_lora
|
||||
[distillation]=distillation_dmd [self-forcing]=self_forcing
|
||||
|
||||
@@ -23,6 +23,11 @@ on:
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
dreamverse_cuda_12_9:
|
||||
description: 'Build Dreamverse CUDA 12.9 backend-only and UI images'
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
|
||||
|
||||
permissions:
|
||||
@@ -65,3 +70,27 @@ jobs:
|
||||
dockerfile_path: docker/Dockerfile.python3.12.cuda12.9.1
|
||||
tag_suffix: py3.12-cuda12.9.1
|
||||
secrets: inherit
|
||||
|
||||
build-dreamverse-backend-cuda-12-9:
|
||||
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
|
||||
uses: ./.github/workflows/_template-build-image.yml
|
||||
with:
|
||||
python_version: '3.12'
|
||||
dockerfile_path: apps/dreamverse/docker/Dockerfile
|
||||
tag_suffix: dreamverse-backend-cuda12.9.1
|
||||
image_name: dreamverse
|
||||
build_args: BUILD_DREAMVERSE_UI=0
|
||||
include_latest_tags: false
|
||||
secrets: inherit
|
||||
|
||||
build-dreamverse-ui-cuda-12-9:
|
||||
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
|
||||
uses: ./.github/workflows/_template-build-image.yml
|
||||
with:
|
||||
python_version: '3.12'
|
||||
dockerfile_path: apps/dreamverse/docker/Dockerfile
|
||||
tag_suffix: dreamverse-ui-cuda12.9.1
|
||||
image_name: dreamverse
|
||||
build_args: BUILD_DREAMVERSE_UI=1
|
||||
include_latest_tags: false
|
||||
secrets: inherit
|
||||
|
||||
@@ -34,6 +34,8 @@ env
|
||||
*.log
|
||||
weights/
|
||||
logs/
|
||||
official_weights/
|
||||
converted_weights/
|
||||
|
||||
# SSIM test outputs
|
||||
fastvideo/tests/ssim/generated_videos/
|
||||
|
||||
@@ -42,6 +42,8 @@ FastVideo has the following features:
|
||||
- Support H100, A100, 4090
|
||||
- Support Linux, Windows, MacOS
|
||||
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
|
||||
- Realtime video generation & editing
|
||||
- [Dreamverse](apps/dreamverse/README.md): stream and "vibe direct" video in realtime ([live demo](https://dreamverse.fastvideo.org/)), deployable on local GPU, a self-hosted B200 server, Docker, or serverless Modal
|
||||
|
||||
## Getting Started
|
||||
|
||||
@@ -69,6 +71,19 @@ See below for recipes and datasets:
|
||||
| [FastWan2.1-T2V-1.3B](https://huggingface.co/FastVideo/FastWan2.1-T2V-1.3B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.1-T2V/Wan-Syn-Data-480P) | [FastVideo Synthetic Wan2.1 480P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x448x832_600k) |
|
||||
| [FastWan2.2-TI2V-5B](https://huggingface.co/FastVideo/FastWan2.2-TI2V-5B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.2-TI2V-5B-Diffusers/Data-free) | [FastVideo Synthetic Wan2.2 720P](https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k) |
|
||||
|
||||
## Dreamverse — Realtime Video Generation & Editing
|
||||
|
||||
[Dreamverse](apps/dreamverse/README.md) is FastVideo's realtime video generation
|
||||
and editing platform — "vibe directing" a video as it streams. It lives in the
|
||||
monorepo under [`apps/dreamverse/`](apps/dreamverse/) and ships its own backend
|
||||
(`dreamverse-server`) plus a web UI.
|
||||
|
||||
Try the [live demo](https://dreamverse.fastvideo.org/), read the
|
||||
[blog](https://haoailab.com/blogs/dreamverse/), or run it yourself. Dreamverse
|
||||
deploys on a local GPU, a self-hosted B200 server over SSH, Docker, or
|
||||
serverless [Modal](apps/dreamverse/scripts/modal/README.md) — see the
|
||||
[Dreamverse README](apps/dreamverse/README.md).
|
||||
|
||||
## Inference
|
||||
|
||||
### Generating Your First Video
|
||||
|
||||
+68
-11
@@ -2,6 +2,8 @@
|
||||
|
||||
Dreamverse is the FastVideo realtime video generation & editing platform. It lives in this monorepo under `apps/dreamverse/`.
|
||||
|
||||
**Deploy on:** [local GPU](#quick-start-local-gpu) · [self-hosted B200 (SSH)](#server-b200-deployment-ssh) · [Docker](docker/README.md) · [Modal](scripts/modal/README.md)
|
||||
|
||||
## Install Dreamverse
|
||||
|
||||
You can install Dreamverse using one of the methods below.
|
||||
@@ -43,13 +45,36 @@ See `apps/dreamverse/docker/README.md` for Docker build and run option details.
|
||||
## Optional: Building FFmpeg For Better Performance
|
||||
|
||||
For full streaming performance in a non-Docker install, build a custom FFmpeg
|
||||
binary:
|
||||
binary from a FastVideo source checkout. The command below is repo-relative,
|
||||
so run it from the repository root:
|
||||
|
||||
```bash
|
||||
bash apps/dreamverse/scripts/install_native_ffmpeg.sh
|
||||
```
|
||||
|
||||
This builds and installs into `~/opt/ffmpeg-native/` and writes
|
||||
The installer supports Linux `x86_64` and `aarch64`. It prefers conda-forge
|
||||
triplet compilers when those commands are on `PATH`, otherwise it falls back to
|
||||
system `gcc`/`g++` (plain venv). On `x86_64`, x264's hand-tuned SIMD also
|
||||
requires `nasm`; install via whichever path fits your host:
|
||||
|
||||
```bash
|
||||
sudo apt install nasm # Debian/Ubuntu
|
||||
conda install -c conda-forge nasm # inside an active conda env
|
||||
```
|
||||
|
||||
No sudo and no conda? Build `nasm` from source (~30s, installs into `$HOME`):
|
||||
|
||||
```bash
|
||||
(
|
||||
mkdir -p "$HOME/src" "$HOME/opt" && cd "$HOME/src"
|
||||
curl -fsSL -O https://www.nasm.us/pub/nasm/releasebuilds/2.16.03/nasm-2.16.03.tar.gz
|
||||
tar -xf nasm-2.16.03.tar.gz && cd nasm-2.16.03
|
||||
./configure --prefix="$HOME/opt/nasm" && make -j"$(nproc)" && make install
|
||||
)
|
||||
export PATH="$HOME/opt/nasm/bin:$PATH" # add to ~/.bashrc to persist
|
||||
```
|
||||
|
||||
The installer writes to `~/opt/ffmpeg-native/` and emits
|
||||
`apps/dreamverse/scripts/ffmpeg-env.sh`. Source it before starting the backend
|
||||
so Dreamverse uses the custom FFmpeg binary:
|
||||
|
||||
@@ -59,8 +84,7 @@ dreamverse-server
|
||||
```
|
||||
|
||||
Docker images already run this FFmpeg build during image creation and source the
|
||||
generated environment file at container startup. The installer supports Linux
|
||||
`x86_64` and `aarch64`.
|
||||
generated environment file at container startup.
|
||||
|
||||
## Launch Dreamverse
|
||||
|
||||
@@ -71,17 +95,24 @@ dreamverse-server --port 8009
|
||||
dreamverse-mock-server --port 8009
|
||||
```
|
||||
|
||||
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
|
||||
> (the default), the backend compiles the segment 1 and segment 2 inference
|
||||
> paths before it reports ready — this can take **tens of minutes on a cold
|
||||
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
|
||||
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
|
||||
> warmup finishes. For a faster, uncompiled startup while testing, set
|
||||
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
|
||||
|
||||
## Frontend Setup
|
||||
|
||||
Install the web dependencies once from the FastVideo checkout:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
pnpm install --frozen-lockfile
|
||||
npm ci
|
||||
```
|
||||
|
||||
The frontend package also has an npm lockfile, but the bundled launch scripts
|
||||
use `pnpm`.
|
||||
The frontend package uses `package-lock.json`; use npm for installs and scripts.
|
||||
|
||||
## Quick Start: Local GPU
|
||||
|
||||
@@ -137,11 +168,37 @@ Start the frontend:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
BACKEND_HOST=localhost BACKEND_PORT=8009 pnpm run dev
|
||||
BACKEND_HOST=localhost BACKEND_PORT=8009 npm run dev
|
||||
```
|
||||
|
||||
Open `http://localhost:5299`.
|
||||
|
||||
## Server B200 deployment (SSH)
|
||||
|
||||
Deploying on a remote GPU host (for example a B200 box) is a local install run
|
||||
over SSH, plus a few server-specific concerns. Two paths:
|
||||
|
||||
### Option A: Native (source install)
|
||||
|
||||
SSH in, then follow [Install → From source](#method-2-from-source) and
|
||||
(recommended) [Building FFmpeg](#optional-building-ffmpeg-for-better-performance),
|
||||
then start the backend as in [Quick Start: Local GPU](#quick-start-local-gpu).
|
||||
For a remote host, a few things differ from localhost:
|
||||
|
||||
- Bind all interfaces: `dreamverse-server --host 0.0.0.0 --port 8009`.
|
||||
- Point the frontend/client at the host: `BACKEND_HOST=<b200-host> BACKEND_PORT=8009 npm run dev`.
|
||||
- Keep the backend alive across SSH sessions (`tmux` / `systemd` / `nohup`).
|
||||
- Expose / firewall port `8009`, or front it with a reverse proxy + auth.
|
||||
|
||||
### Option B: Docker (on the server)
|
||||
|
||||
SSH in, then follow [Install → Using Docker](#method-3-using-docker) and the run
|
||||
steps in [`docker/README.md`](docker/README.md):
|
||||
|
||||
```bash
|
||||
CEREBRAS_API_KEY="<key>" GROQ_API_KEY="<key>" apps/dreamverse/docker/docker_run.sh
|
||||
```
|
||||
|
||||
## Quick Start: Mock Backend (For UI development)
|
||||
|
||||
The mock server emulates the Dreamverse backend protocol and streams a
|
||||
@@ -173,14 +230,14 @@ Run the frontend tests:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
pnpm test
|
||||
npm test
|
||||
```
|
||||
|
||||
Run the frontend e2e tests:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
pnpm run e2e
|
||||
npm run e2e
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
@@ -216,7 +273,7 @@ Only one GPU is used
|
||||
Frontend cannot connect to backend
|
||||
|
||||
- confirm the backend is running on `8009`; if not, point the frontend at it
|
||||
with `BACKEND_HOST=<host> BACKEND_PORT=<port> pnpm run dev`
|
||||
with `BACKEND_HOST=<host> BACKEND_PORT=<port> npm run dev`
|
||||
- confirm `http://localhost:8009/healthz` responds before starting the frontend
|
||||
- confirm `http://localhost:8009/readyz` returns `200` before clicking Generate
|
||||
- use `apps/dreamverse/scripts/smoke_local.sh` for a repeatable local startup
|
||||
|
||||
@@ -1,40 +0,0 @@
|
||||
.git/
|
||||
**/.git/
|
||||
.venv/
|
||||
**/.venv/
|
||||
**/__pycache__/
|
||||
**/*.pyc
|
||||
**/*.pyo
|
||||
**/*.egg-info/
|
||||
.pytest_cache/
|
||||
**/.pytest_cache/
|
||||
.mypy_cache/
|
||||
.ruff_cache/
|
||||
.cache/
|
||||
|
||||
apps/dreamverse/web/node_modules/
|
||||
apps/dreamverse/web/.next/
|
||||
apps/dreamverse/web/out/
|
||||
apps/dreamverse/web/dist/
|
||||
apps/dreamverse/web/test-results/
|
||||
apps/dreamverse/web/playwright-report/
|
||||
|
||||
apps/dreamverse/outputs/
|
||||
apps/dreamverse/dreamverse/outputs/
|
||||
apps/dreamverse/dreamverse/prompts.local/
|
||||
apps/dreamverse/logs/
|
||||
outputs/
|
||||
logs/
|
||||
slurm-logs/
|
||||
wandb/
|
||||
|
||||
.env
|
||||
.env.*
|
||||
**/prompts.local/
|
||||
.codex/
|
||||
.agents/exploration/
|
||||
.vscode/
|
||||
.idea/
|
||||
*.log
|
||||
*.tmp
|
||||
*.pdf
|
||||
@@ -3,6 +3,7 @@ ARG CUDA_TAG=12.9.1-cudnn-devel-ubuntu22.04
|
||||
FROM nvidia/cuda:${CUDA_TAG}
|
||||
|
||||
ARG BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=0
|
||||
ARG BUILD_DREAMVERSE_UI=0
|
||||
|
||||
ENV DEBIAN_FRONTEND=noninteractive \
|
||||
PYTHONUNBUFFERED=1 \
|
||||
@@ -50,6 +51,18 @@ RUN if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
|
||||
echo "Skipping source fastvideo-kernel build; using installed fastvideo-kernel package."; \
|
||||
fi
|
||||
|
||||
RUN if [[ "${BUILD_DREAMVERSE_UI}" == "1" ]]; then \
|
||||
curl -fsSL https://deb.nodesource.com/setup_22.x | bash - \
|
||||
&& apt-get install -y --no-install-recommends nodejs \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& cd /opt/FastVideo/apps/dreamverse/web \
|
||||
&& npm ci --ignore-scripts \
|
||||
&& NEXT_OUTPUT_EXPORT=1 npm run build \
|
||||
&& rm -rf node_modules .next; \
|
||||
else \
|
||||
echo "Skipping Dreamverse UI build; Building backend-only image."; \
|
||||
fi
|
||||
|
||||
# The monorepo ffmpeg installer force-selects conda compiler triplets for
|
||||
# local dev shells. Inside this image we explicitly opt into the system
|
||||
# gcc/g++ toolchain.
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
**/.git/
|
||||
.venv/
|
||||
**/.venv/
|
||||
.*dreamverse*/
|
||||
**/__pycache__/
|
||||
**/*.pyc
|
||||
**/*.pyo
|
||||
@@ -24,6 +25,9 @@ apps/dreamverse/dreamverse/outputs/
|
||||
apps/dreamverse/dreamverse/prompts.local/
|
||||
apps/dreamverse/logs/
|
||||
outputs/
|
||||
outputs_video/
|
||||
quality_check_outputs/
|
||||
data/
|
||||
logs/
|
||||
slurm-logs/
|
||||
wandb/
|
||||
@@ -32,7 +36,9 @@ wandb/
|
||||
.env.*
|
||||
**/prompts.local/
|
||||
.codex/
|
||||
.agents/exploration/
|
||||
.agents/
|
||||
.opencode/
|
||||
.agent_tmp/
|
||||
.vscode/
|
||||
.idea/
|
||||
*.log
|
||||
|
||||
@@ -1,9 +1,9 @@
|
||||
# Dreamverse Docker Image
|
||||
|
||||
This folder contains the backend-only Docker image for Dreamverse inside the
|
||||
FastVideo monorepo. Build commands use the FastVideo repository root as the
|
||||
Docker context, so run the helper scripts from this folder or from any path in
|
||||
the checkout.
|
||||
This folder contains the Docker image for Dreamverse inside the FastVideo
|
||||
monorepo. Build commands use the FastVideo repository root as the Docker
|
||||
context, so run the helper scripts from this folder or from any path in the
|
||||
checkout.
|
||||
|
||||
## Build
|
||||
|
||||
@@ -17,6 +17,16 @@ The image defaults to `dreamverse:dev`. Override it with:
|
||||
DREAMVERSE_IMAGE=dreamverse:local apps/dreamverse/docker/docker_build.sh
|
||||
```
|
||||
|
||||
Backend-only remains the default image. To include the static Dreamverse UI
|
||||
served by the backend, set `BUILD_DREAMVERSE_UI=1` and choose a specific image
|
||||
tag:
|
||||
|
||||
```bash
|
||||
BUILD_DREAMVERSE_UI=1 DREAMVERSE_IMAGE=<image-tag> apps/dreamverse/docker/docker_build.sh
|
||||
```
|
||||
|
||||
Prefer SHA-specific tags for deployable images; avoid `latest`.
|
||||
|
||||
The Dockerfile builds a CUDA 12.9.1 image, installs FastVideo from this
|
||||
checkout with the `dreamverse` extra, installs the FA4
|
||||
flash-attention fork, builds native FFmpeg, and installs FlashInfer for NVFP4
|
||||
|
||||
@@ -4,13 +4,44 @@ set -euo pipefail
|
||||
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO_ROOT="$(cd -- "${SCRIPT_DIR}/../../.." && pwd)"
|
||||
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
|
||||
ROOT_DOCKERIGNORE="${REPO_ROOT}/.dockerignore"
|
||||
DOCKERFILE_DOCKERIGNORE="${SCRIPT_DIR}/Dockerfile.dockerignore"
|
||||
CREATED_ROOT_DOCKERIGNORE_SYMLINK=0
|
||||
|
||||
cleanup_root_dockerignore_symlink() {
|
||||
if [[ "${CREATED_ROOT_DOCKERIGNORE_SYMLINK}" == "1" && -L "${ROOT_DOCKERIGNORE}" ]] && \
|
||||
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
|
||||
rm -- "${ROOT_DOCKERIGNORE}"
|
||||
fi
|
||||
}
|
||||
trap cleanup_root_dockerignore_symlink EXIT INT TERM
|
||||
|
||||
if [[ -z "${DOCKER_BUILDKIT:-}" ]] && docker buildx version >/dev/null 2>&1; then
|
||||
export DOCKER_BUILDKIT=1
|
||||
fi
|
||||
|
||||
build_args=()
|
||||
[[ -n "${CUDA_TAG:-}" ]] && build_args+=(--build-arg "CUDA_TAG=${CUDA_TAG}")
|
||||
[[ -n "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE:-}" ]] && \
|
||||
build_args+=(--build-arg "BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}")
|
||||
build_args+=(--build-arg "BUILD_DREAMVERSE_UI=${BUILD_DREAMVERSE_UI:-0}")
|
||||
|
||||
exec docker build \
|
||||
if [[ "${DOCKER_BUILDKIT:-}" == "1" ]]; then
|
||||
if [[ -L "${ROOT_DOCKERIGNORE}" ]] && \
|
||||
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
|
||||
rm -- "${ROOT_DOCKERIGNORE}"
|
||||
printf 'Removed stale temporary root .dockerignore symlink: %s\n' "${ROOT_DOCKERIGNORE}"
|
||||
fi
|
||||
elif [[ -e "${ROOT_DOCKERIGNORE}" || -L "${ROOT_DOCKERIGNORE}" ]]; then
|
||||
printf 'Using existing root .dockerignore: %s\n' "${ROOT_DOCKERIGNORE}"
|
||||
else
|
||||
ln -s -- "${DOCKERFILE_DOCKERIGNORE}" "${ROOT_DOCKERIGNORE}"
|
||||
CREATED_ROOT_DOCKERIGNORE_SYMLINK=1
|
||||
printf 'Created temporary root .dockerignore symlink for legacy Docker builder: %s -> %s\n' \
|
||||
"${ROOT_DOCKERIGNORE}" "${DOCKERFILE_DOCKERIGNORE}"
|
||||
fi
|
||||
|
||||
docker build \
|
||||
-f "${SCRIPT_DIR}/Dockerfile" \
|
||||
-t "${IMAGE}" \
|
||||
"${build_args[@]}" \
|
||||
|
||||
@@ -4,6 +4,7 @@ from pathlib import Path
|
||||
_REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
_SERVER_ROOT = Path(__file__).resolve().parent
|
||||
_FASTVIDEO_DREAMVERSE_HOME = os.environ.get("FASTVIDEO_DREAMVERSE_HOME")
|
||||
_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT = os.environ.get("FASTVIDEO_DREAMVERSE_FRONTEND_ROOT")
|
||||
_XDG_STATE_HOME = os.environ.get("XDG_STATE_HOME")
|
||||
_DEFAULT_STATE_ROOT = (Path(_FASTVIDEO_DREAMVERSE_HOME) if _FASTVIDEO_DREAMVERSE_HOME else
|
||||
(Path(_XDG_STATE_HOME) if _XDG_STATE_HOME else Path.home() / ".local/state") /
|
||||
@@ -16,18 +17,39 @@ _APP_ROOT = _REPO_ROOT
|
||||
|
||||
def _resolve_frontend_root() -> Path:
|
||||
for candidate in (
|
||||
Path(_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT) if _FASTVIDEO_DREAMVERSE_FRONTEND_ROOT else None,
|
||||
_APP_ROOT / "web",
|
||||
_APP_ROOT / "prod-ui",
|
||||
):
|
||||
if candidate.is_dir():
|
||||
if candidate is not None and candidate.is_dir():
|
||||
return candidate
|
||||
if _FASTVIDEO_DREAMVERSE_FRONTEND_ROOT:
|
||||
return Path(_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT)
|
||||
return _APP_ROOT / "web"
|
||||
|
||||
|
||||
def _resolve_frontend_static_dir_candidates() -> tuple[str, ...]:
|
||||
roots = (
|
||||
FRONTEND_ROOT,
|
||||
Path.cwd() / "apps/dreamverse/web",
|
||||
Path.cwd() / "web",
|
||||
Path("/opt/FastVideo/apps/dreamverse/web"),
|
||||
)
|
||||
candidates: list[str] = []
|
||||
seen: set[Path] = set()
|
||||
for root in roots:
|
||||
resolved_root = root.resolve(strict=False)
|
||||
if resolved_root in seen:
|
||||
continue
|
||||
seen.add(resolved_root)
|
||||
candidates.extend(str(resolved_root / dirname) for dirname in ("out", "dist"))
|
||||
return tuple(candidates)
|
||||
|
||||
|
||||
FRONTEND_ROOT = _resolve_frontend_root()
|
||||
_CLIENT_PROMPTS_ROOT = FRONTEND_ROOT / "prompts"
|
||||
_CLIENT_PROMPTS_LOCAL_ROOT = FRONTEND_ROOT / "prompts.local"
|
||||
FRONTEND_STATIC_DIR_CANDIDATES = tuple(str(FRONTEND_ROOT / dirname) for dirname in ("out", "dist"))
|
||||
FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
|
||||
|
||||
# Model registry
|
||||
MODEL_REGISTRY = {
|
||||
@@ -35,18 +57,24 @@ MODEL_REGISTRY = {
|
||||
"name": "FastLTX2",
|
||||
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
|
||||
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
|
||||
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
|
||||
},
|
||||
"fast-ltx23": {
|
||||
"name": "FastLTX23",
|
||||
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
|
||||
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
|
||||
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
|
||||
},
|
||||
}
|
||||
|
||||
DEFAULT_MODEL_ID = "fast-ltx2"
|
||||
|
||||
ACTIVE_MODEL_ID = (os.getenv("DREAMVERSE_MODEL_ID", "").strip() or DEFAULT_MODEL_ID)
|
||||
if ACTIVE_MODEL_ID not in MODEL_REGISTRY:
|
||||
ACTIVE_MODEL_ID = DEFAULT_MODEL_ID
|
||||
|
||||
# Active model configuration
|
||||
MODEL_CONFIG = MODEL_REGISTRY[DEFAULT_MODEL_ID]
|
||||
MODEL_CONFIG = MODEL_REGISTRY[ACTIVE_MODEL_ID]
|
||||
|
||||
# Generation limits
|
||||
SESSION_TIMEOUT_SECONDS = 300
|
||||
@@ -142,6 +170,76 @@ def _optional_env(*names: str) -> str | None:
|
||||
|
||||
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
|
||||
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
|
||||
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
|
||||
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
|
||||
|
||||
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
|
||||
if DREAMVERSE_MODEL_PATH:
|
||||
MODEL_CONFIG = {
|
||||
**MODEL_CONFIG,
|
||||
"model_path": DREAMVERSE_MODEL_PATH,
|
||||
"config_model_path": DREAMVERSE_MODEL_PATH,
|
||||
}
|
||||
|
||||
AVAILABLE_LORAS = {
|
||||
"pixar": {
|
||||
"repo": "vrgamedevgirl84/LTX_2.3_Pixar_Toon_Style_LoRa",
|
||||
"trigger": "P1x4r",
|
||||
"model": "fast-ltx23",
|
||||
"position": "prepend",
|
||||
"label": "Pixar Toon",
|
||||
},
|
||||
"transition": {
|
||||
"repo": "valiantcat/LTX-2.3-Transition-LORA",
|
||||
"trigger": "zhuanchang",
|
||||
"model": "fast-ltx23",
|
||||
"position": "append",
|
||||
"label": "Transition",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _active_model_key() -> str:
|
||||
return ACTIVE_MODEL_ID
|
||||
|
||||
|
||||
def _available_styles_for_active_model() -> list[str]:
|
||||
active = _active_model_key()
|
||||
return [k for k, v in AVAILABLE_LORAS.items() if v.get("model") == active]
|
||||
|
||||
|
||||
def _resolve_lora_spec(spec: str) -> str | None:
|
||||
spec = (spec or "").strip()
|
||||
if not spec:
|
||||
return None
|
||||
if spec.lower() == "omninft":
|
||||
return MODEL_CONFIG.get("lora_repo")
|
||||
if spec.lower() in AVAILABLE_LORAS:
|
||||
return AVAILABLE_LORAS[spec.lower()]["repo"]
|
||||
return spec
|
||||
|
||||
|
||||
def _parse_lora_stack(raw: str) -> list[tuple[str, float]]:
|
||||
out: list[tuple[str, float]] = []
|
||||
for item in (raw or "").split(","):
|
||||
item = item.strip()
|
||||
if not item:
|
||||
continue
|
||||
spec, _, stren = item.partition("@")
|
||||
resolved = _resolve_lora_spec(spec)
|
||||
if resolved:
|
||||
try:
|
||||
strength = float(stren) if stren.strip() else 1.0
|
||||
except ValueError:
|
||||
strength = 1.0
|
||||
out.append((resolved, strength))
|
||||
return out
|
||||
|
||||
|
||||
DREAMVERSE_LORA_PATH = _resolve_lora_spec(os.getenv("DREAMVERSE_LORA_PATH", ""))
|
||||
DREAMVERSE_LORA_NICKNAME = (os.getenv("DREAMVERSE_LORA_NICKNAME", "omninft").strip() or "omninft")
|
||||
DREAMVERSE_LORA_STRENGTH = _env_float("DREAMVERSE_LORA_STRENGTH", 1.0)
|
||||
DREAMVERSE_LORA_STACK = _parse_lora_stack(os.getenv("DREAMVERSE_LORA_STACK", ""))
|
||||
|
||||
|
||||
def _resolve_devtools_paths(
|
||||
|
||||
@@ -13,6 +13,7 @@ from multiprocessing import Process, Queue
|
||||
|
||||
from dreamverse.config import (
|
||||
DEFAULT_MODEL_ID,
|
||||
DREAMVERSE_SP_SIZE,
|
||||
MODEL_REGISTRY,
|
||||
STARTUP_WARMUP_ENABLED,
|
||||
STARTUP_WARMUP_PROMPT,
|
||||
@@ -33,6 +34,8 @@ from dreamverse.worker_ipc import (
|
||||
InitAck,
|
||||
JoinAck,
|
||||
LeaveAck,
|
||||
LoraAck,
|
||||
LoraStackPayload,
|
||||
MediaChunk,
|
||||
MediaComplete,
|
||||
MediaInit,
|
||||
@@ -80,6 +83,7 @@ class CommandType(Enum):
|
||||
USER_STEP = "user_step"
|
||||
USER_LEAVE = "user_leave"
|
||||
RELOAD_MODEL = "reload_model"
|
||||
APPLY_LORA = "apply_lora"
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -282,6 +286,25 @@ def gpu_worker_process(
|
||||
message=str(e),
|
||||
))
|
||||
|
||||
elif cmd.type == CommandType.APPLY_LORA:
|
||||
try:
|
||||
assert isinstance(cmd.payload, LoraStackPayload), (f"APPLY_LORA requires LoraStackPayload, "
|
||||
f"got {type(cmd.payload).__name__}")
|
||||
trigger, position = worker.apply_lora_stack(cmd.payload.stack)
|
||||
response_queue.put(
|
||||
LoraAck(
|
||||
user_id=cmd.user_id,
|
||||
style_trigger=trigger,
|
||||
style_trigger_position=position,
|
||||
))
|
||||
except Exception as e:
|
||||
print(f"[GPU {gpu_id}] Apply LoRA error: {e}")
|
||||
traceback.print_exc()
|
||||
response_queue.put(WorkerError(
|
||||
user_id=cmd.user_id,
|
||||
message=str(e),
|
||||
))
|
||||
|
||||
return True # Continue loop
|
||||
|
||||
if first_cmd is not None:
|
||||
@@ -359,6 +382,25 @@ def gpu_worker_process(
|
||||
message=str(e),
|
||||
))
|
||||
|
||||
elif cmd.type == CommandType.APPLY_LORA:
|
||||
try:
|
||||
assert isinstance(cmd.payload, LoraStackPayload), (f"APPLY_LORA requires LoraStackPayload, "
|
||||
f"got {type(cmd.payload).__name__}")
|
||||
trigger, position = worker.apply_lora_stack(cmd.payload.stack)
|
||||
response_queue.put(
|
||||
LoraAck(
|
||||
user_id=cmd.user_id,
|
||||
style_trigger=trigger,
|
||||
style_trigger_position=position,
|
||||
))
|
||||
except Exception as e:
|
||||
print(f"[GPU {gpu_id}] Apply LoRA error: {e}")
|
||||
traceback.print_exc()
|
||||
response_queue.put(WorkerError(
|
||||
user_id=cmd.user_id,
|
||||
message=str(e),
|
||||
))
|
||||
|
||||
elif cmd.type in (CommandType.USER_JOIN, CommandType.USER_STEP, CommandType.USER_LEAVE):
|
||||
event_loop(first_cmd=cmd)
|
||||
break
|
||||
@@ -729,6 +771,25 @@ class GPUSlot:
|
||||
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
|
||||
f"{type(response).__name__}")
|
||||
|
||||
async def apply_lora_stack(
|
||||
self,
|
||||
stack: list[tuple[str, float]],
|
||||
) -> LoraAck:
|
||||
"""Re-apply a runtime LoRA stack on this GPU's worker."""
|
||||
self._active = True
|
||||
user_id = "__lora__"
|
||||
payload = LoraStackPayload(stack=stack)
|
||||
response = await self._send_command_tagged(Command(CommandType.APPLY_LORA, payload=payload, user_id=user_id),
|
||||
timeout=120.0)
|
||||
match response:
|
||||
case LoraAck() as ack:
|
||||
return ack
|
||||
case WorkerError(message=msg):
|
||||
raise RuntimeError(f"Apply LoRA failed on GPU {self.gpu_id}: {msg}")
|
||||
case _:
|
||||
raise RuntimeError(f"Unexpected LoRA response on GPU {self.gpu_id}: "
|
||||
f"{type(response).__name__}")
|
||||
|
||||
async def leave_user(self, user_id: str) -> None:
|
||||
"""Remove a user from this GPU."""
|
||||
try:
|
||||
@@ -785,8 +846,16 @@ class GPUPool:
|
||||
"""Manages multiple GPU worker subprocesses."""
|
||||
|
||||
def __init__(self, gpu_ids: list[int]):
|
||||
self.gpu_ids = gpu_ids
|
||||
self.slots: dict[int, GPUSlot] = {gpu_id: GPUSlot(gpu_id, str(gpu_id)) for gpu_id in gpu_ids}
|
||||
sp_size = DREAMVERSE_SP_SIZE
|
||||
groups = [gpu_ids[i:i + sp_size] for i in range(0, len(gpu_ids), sp_size)]
|
||||
groups = [g for g in groups if len(g) == sp_size]
|
||||
if not groups:
|
||||
raise RuntimeError(f"Not enough GPUs for DREAMVERSE_SP_SIZE={sp_size}: available={gpu_ids}")
|
||||
if sp_size > 1:
|
||||
print(f"[INFO] Sequence-parallel slots (sp_size={sp_size}): " + ", ".join("{" + ",".join(map(str, g)) + "}"
|
||||
for g in groups))
|
||||
self.gpu_ids = [g[0] for g in groups]
|
||||
self.slots: dict[int, GPUSlot] = {g[0]: GPUSlot(g[0], ",".join(str(x) for x in g)) for g in groups}
|
||||
self.waiting_list: list[tuple[str, asyncio.Event, WebSocket]] = []
|
||||
self.client_gpu_map: dict[str, int] = {}
|
||||
self._pool_lock = asyncio.Lock()
|
||||
@@ -903,6 +972,17 @@ class GPUPool:
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
async def apply_lora_stack(
|
||||
self,
|
||||
stack: list[tuple[str, float]],
|
||||
) -> dict[int, str | None]:
|
||||
"""Re-apply a runtime LoRA stack across all ready GPU workers."""
|
||||
ready_slots = [slot for slot in self.slots.values() if slot.ready]
|
||||
if not ready_slots:
|
||||
raise RuntimeError("No ready GPU workers to apply LoRA stack.")
|
||||
results = await asyncio.gather(*(slot.apply_lora_stack(stack) for slot in ready_slots))
|
||||
return {slot.gpu_id: ack.style_trigger for slot, ack in zip(ready_slots, results, strict=False)}
|
||||
|
||||
async def shutdown(self):
|
||||
"""Shutdown all GPU workers."""
|
||||
print("Shutting down GPU pool...")
|
||||
|
||||
@@ -6,18 +6,22 @@ import os
|
||||
from contextlib import asynccontextmanager
|
||||
from pathlib import Path
|
||||
|
||||
from fastapi import FastAPI, WebSocket
|
||||
from fastapi import FastAPI, HTTPException, WebSocket
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
from pydantic import BaseModel
|
||||
from fastvideo.entrypoints.streaming import build_health_router
|
||||
from dreamverse.gpu_pool import GPUPool, get_available_gpus
|
||||
from dreamverse.session_logger import SessionEventLogger
|
||||
|
||||
from dreamverse.config import (
|
||||
AVAILABLE_LORAS,
|
||||
DEVTOOLS_ENABLED,
|
||||
FRONTEND_STATIC_DIR_CANDIDATES,
|
||||
PROMPT_SAFETY_ENABLED,
|
||||
SESSION_LOG_ROOT,
|
||||
_available_styles_for_active_model,
|
||||
_resolve_lora_spec,
|
||||
)
|
||||
from dreamverse.prompt_enhancer import PromptEnhancer
|
||||
from dreamverse.prompt_safety import PromptSafetyFilter
|
||||
@@ -104,6 +108,83 @@ async def websocket_endpoint(websocket: WebSocket):
|
||||
await controller.run()
|
||||
|
||||
|
||||
class LoraRequest(BaseModel):
|
||||
strength: float = 1.0
|
||||
styles: dict[str, float] = {}
|
||||
style: str = ""
|
||||
|
||||
|
||||
@app.get("/lora/options")
|
||||
async def lora_options() -> dict:
|
||||
if not DEVTOOLS_ENABLED:
|
||||
raise HTTPException(status_code=404, detail="Not found")
|
||||
styles = _available_styles_for_active_model()
|
||||
has_base_lora = _resolve_lora_spec("omninft") is not None
|
||||
labels = {"none": "None"}
|
||||
labels.update({k: AVAILABLE_LORAS[k].get("label", k) for k in styles})
|
||||
return {
|
||||
"styles": ["none", *styles],
|
||||
"labels": labels,
|
||||
"has_base_lora": has_base_lora,
|
||||
}
|
||||
|
||||
|
||||
@app.post("/lora")
|
||||
async def apply_lora(request: LoraRequest) -> dict:
|
||||
if not DEVTOOLS_ENABLED:
|
||||
raise HTTPException(status_code=404, detail="Not found")
|
||||
if runtime.gpu_pool is None:
|
||||
raise HTTPException(status_code=503, detail="GPU pool not ready")
|
||||
|
||||
strength = max(0.0, min(1.0, float(request.strength)))
|
||||
allowed_styles = _available_styles_for_active_model()
|
||||
|
||||
requested = dict(request.styles) if request.styles else {}
|
||||
if not requested and request.style and request.style.strip().lower() != "none":
|
||||
requested = {request.style: 1.0}
|
||||
|
||||
styles_map: dict[str, float] = {}
|
||||
for name, intensity in requested.items():
|
||||
key = str(name).strip().lower()
|
||||
if key in ("", "none"):
|
||||
continue
|
||||
if key not in allowed_styles:
|
||||
raise HTTPException(status_code=400, detail=f"Unknown style for active model: {key}")
|
||||
value = max(0.0, min(1.0, float(intensity)))
|
||||
if value <= 0.0:
|
||||
continue
|
||||
styles_map[key] = value
|
||||
|
||||
stack: list[tuple[str, float]] = []
|
||||
if _resolve_lora_spec("omninft") is not None:
|
||||
stack.append(("omninft", strength))
|
||||
for key, intensity in styles_map.items():
|
||||
stack.append((key, intensity))
|
||||
|
||||
if not stack:
|
||||
raise HTTPException(status_code=400,
|
||||
detail="Active model has no OmniNFT LoRA and no style selected; nothing to apply.")
|
||||
|
||||
try:
|
||||
triggers = await runtime.gpu_pool.apply_lora_stack(stack)
|
||||
except Exception as exc:
|
||||
raise HTTPException(status_code=500, detail=str(exc)) from exc
|
||||
|
||||
return {
|
||||
"applied": True,
|
||||
"strength": strength,
|
||||
"styles": styles_map,
|
||||
"triggers": {
|
||||
key: AVAILABLE_LORAS[key]["trigger"]
|
||||
for key in styles_map
|
||||
},
|
||||
"gpus": {
|
||||
str(gpu_id): trigger
|
||||
for gpu_id, trigger in triggers.items()
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
# Serve an exported frontend bundle when present.
|
||||
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
|
||||
if os.path.isdir(static_dir):
|
||||
|
||||
@@ -41,7 +41,7 @@ MOCK_FPS = 24
|
||||
MOCK_DURATION_SECONDS = 5.0
|
||||
MOCK_STREAM_CHUNK_SIZE_BYTES = 256 * 1024
|
||||
MOCK_STREAM_CHUNK_DELAY_MS = 15
|
||||
MOCK_AV_MIME = 'video/mp4; codecs="avc1.42E01E"'
|
||||
MOCK_AV_MIME = 'video/mp4; codecs="avc1.42E01E,mp4a.40.2"'
|
||||
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
|
||||
MOCK_SEGMENT_BYTES: bytes | None = None
|
||||
|
||||
@@ -75,9 +75,14 @@ def _build_mock_segment_bytes() -> bytes:
|
||||
"lavfi",
|
||||
"-i",
|
||||
f"testsrc2=size={MOCK_FRAME_WIDTH}x{MOCK_FRAME_HEIGHT}:rate={MOCK_FPS}",
|
||||
# Silent stereo AAC track. We don't need audible content — the FE's
|
||||
# audio-decode coverage only needs a real AAC track present
|
||||
"-f",
|
||||
"lavfi",
|
||||
"-i",
|
||||
"anullsrc=channel_layout=stereo:sample_rate=48000",
|
||||
"-t",
|
||||
f"{MOCK_DURATION_SECONDS}",
|
||||
"-an",
|
||||
"-c:v",
|
||||
"libx264",
|
||||
"-preset",
|
||||
@@ -90,6 +95,12 @@ def _build_mock_segment_bytes() -> bytes:
|
||||
"baseline",
|
||||
"-level",
|
||||
"3.0",
|
||||
"-c:a",
|
||||
"aac",
|
||||
"-b:a",
|
||||
"128k",
|
||||
"-ar",
|
||||
"48000",
|
||||
"-movflags",
|
||||
"+empty_moov+default_base_moof+frag_keyframe",
|
||||
"-frag_duration",
|
||||
@@ -183,6 +194,55 @@ async def status():
|
||||
"queue_size": 0,
|
||||
"available_gpus": 1,
|
||||
"total_gpus": 1,
|
||||
"warmup_enabled": False,
|
||||
"warmup_successful_gpus": 1,
|
||||
"warmup_failed_gpus": 0,
|
||||
"gpu_status": {
|
||||
0: {
|
||||
"ready": True,
|
||||
"available": True,
|
||||
"client_count": 0,
|
||||
"current_model_id": "mock-ltx2",
|
||||
"process_alive": True,
|
||||
"warmup_enabled": False,
|
||||
"warmup_success": True,
|
||||
"warmup_error": None,
|
||||
"warmup_timings": {},
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@app.get("/prompt-system-config")
|
||||
async def prompt_system_config():
|
||||
return {
|
||||
"next_segment_system_prompt": "Mock next segment prompt.",
|
||||
"auto_extension_system_prompt": "Mock auto extension prompt.",
|
||||
"rewrite_window_system_prompt": "Mock rewrite window prompt.",
|
||||
"rewrite_user_system_prompt": "Mock rewrite user prompt.",
|
||||
"rewrite_model": "mock-rewrite-model",
|
||||
"rewrite_temperature": 0.0,
|
||||
}
|
||||
|
||||
|
||||
@app.get("/curated-presets")
|
||||
async def curated_presets():
|
||||
presets = [
|
||||
{
|
||||
"id": "mock_story",
|
||||
"label": "Mock Story",
|
||||
"description": "CI-safe mock story preset.",
|
||||
"segment_prompts": [
|
||||
"A tiny robot opens a glowing door.",
|
||||
"The robot waves at a friendly moon.",
|
||||
],
|
||||
},
|
||||
]
|
||||
return {
|
||||
"presets": presets,
|
||||
"count": len(presets),
|
||||
"file_path": None,
|
||||
"fallback_file_path": None,
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -124,12 +124,21 @@ def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_
|
||||
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
|
||||
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
|
||||
)
|
||||
assert module.FRONTEND_STATIC_DIR_CANDIDATES == (
|
||||
assert module.FRONTEND_STATIC_DIR_CANDIDATES[:2] == (
|
||||
str(module.FRONTEND_ROOT / "out"),
|
||||
str(module.FRONTEND_ROOT / "dist"),
|
||||
)
|
||||
|
||||
|
||||
def test_config_includes_docker_checkout_static_frontend_candidates(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
|
||||
module = _load_config_module()
|
||||
|
||||
assert "/opt/FastVideo/apps/dreamverse/web/out" in module.FRONTEND_STATIC_DIR_CANDIDATES
|
||||
assert "/opt/FastVideo/apps/dreamverse/web/dist" in module.FRONTEND_STATIC_DIR_CANDIDATES
|
||||
|
||||
|
||||
def test_config_enables_prompt_safety_when_requested(monkeypatch):
|
||||
_set_required_prompt_keys(monkeypatch)
|
||||
monkeypatch.setenv("FASTVIDEO_ENABLE_PROMPT_SAFETY", "true")
|
||||
|
||||
@@ -171,3 +171,69 @@ def test_mock_server_cli_updates_latency(monkeypatch):
|
||||
assert mock_server.LATENCY_MS == 321
|
||||
finally:
|
||||
mock_server.LATENCY_MS = old_latency_ms
|
||||
|
||||
|
||||
def _lora_client(monkeypatch, captured):
|
||||
server_main = _import_server_main(monkeypatch)
|
||||
monkeypatch.setattr(server_main, "DEVTOOLS_ENABLED", True)
|
||||
monkeypatch.setattr(server_main, "_available_styles_for_active_model", lambda: ["pixar", "transition"])
|
||||
monkeypatch.setattr(
|
||||
server_main,
|
||||
"_resolve_lora_spec",
|
||||
lambda spec: "FastVideo/OmniNFT" if str(spec).strip().lower() == "omninft" else spec,
|
||||
)
|
||||
|
||||
class _FakePool:
|
||||
|
||||
async def apply_lora_stack(self, stack):
|
||||
captured["stack"] = list(stack)
|
||||
return {0: "trigger"}
|
||||
|
||||
server_main.runtime.gpu_pool = _FakePool()
|
||||
return TestClient(server_main.app)
|
||||
|
||||
|
||||
def test_apply_lora_stacks_multiple_styles_with_intensities(monkeypatch):
|
||||
captured: dict = {}
|
||||
client = _lora_client(monkeypatch, captured)
|
||||
response = client.post("/lora", json={"strength": 0.8, "styles": {"pixar": 0.45, "transition": 0.5}})
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["applied"] is True
|
||||
assert body["styles"] == {"pixar": 0.45, "transition": 0.5}
|
||||
assert captured["stack"] == [("omninft", 0.8), ("pixar", 0.45), ("transition", 0.5)]
|
||||
|
||||
|
||||
def test_apply_lora_rejects_unknown_style(monkeypatch):
|
||||
captured: dict = {}
|
||||
client = _lora_client(monkeypatch, captured)
|
||||
response = client.post("/lora", json={"strength": 0.5, "styles": {"watercolor": 0.5}})
|
||||
assert response.status_code == 400
|
||||
assert "stack" not in captured
|
||||
|
||||
|
||||
def test_apply_lora_accepts_legacy_single_style(monkeypatch):
|
||||
captured: dict = {}
|
||||
client = _lora_client(monkeypatch, captured)
|
||||
response = client.post("/lora", json={"strength": 0.6, "style": "pixar"})
|
||||
assert response.status_code == 200
|
||||
assert captured["stack"] == [("omninft", 0.6), ("pixar", 1.0)]
|
||||
|
||||
|
||||
def test_apply_lora_clamps_and_drops_zero_intensity(monkeypatch):
|
||||
captured: dict = {}
|
||||
client = _lora_client(monkeypatch, captured)
|
||||
response = client.post("/lora", json={"strength": 1.5, "styles": {"pixar": 2.0, "transition": 0.0}})
|
||||
assert response.status_code == 200
|
||||
body = response.json()
|
||||
assert body["strength"] == 1.0
|
||||
assert body["styles"] == {"pixar": 1.0}
|
||||
assert captured["stack"] == [("omninft", 1.0), ("pixar", 1.0)]
|
||||
|
||||
|
||||
def test_lora_endpoints_hidden_without_devtools(monkeypatch):
|
||||
server_main = _import_server_main(monkeypatch)
|
||||
monkeypatch.setattr(server_main, "DEVTOOLS_ENABLED", False)
|
||||
client = TestClient(server_main.app)
|
||||
assert client.get("/lora/options").status_code == 404
|
||||
assert client.post("/lora", json={"strength": 0.5}).status_code == 404
|
||||
|
||||
@@ -20,9 +20,11 @@ os.environ.setdefault("GROQ_API_KEY", "dummy")
|
||||
|
||||
import dreamverse.main as server_main # noqa: E402
|
||||
import dreamverse.runtime as runtime # noqa: E402
|
||||
from dreamverse.config import PROMPT_TIMEOUT_MS # noqa: E402
|
||||
from dreamverse.session_logger import SessionEventLogger # noqa: E402
|
||||
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit # noqa: E402
|
||||
from dreamverse.session import controller as session_controller # noqa: E402
|
||||
from dreamverse.utils import PROMPT_EXTENSION_FAILURE_USER_MESSAGE # noqa: E402
|
||||
|
||||
pytestmark = pytest.mark.gpu
|
||||
|
||||
@@ -338,7 +340,7 @@ def test_session_event_logger_initializes_hostname_folder_and_utc_filename(
|
||||
logger = SessionEventLogger(tmp_path)
|
||||
assert logger.directory == tmp_path / socket.gethostname()
|
||||
assert logger.directory.is_dir()
|
||||
assert re.match(r"^\d{6}_\d{6}\.jsonl$", logger.path.name)
|
||||
assert re.match(r"^\d{6}_\d{6}_\d{6}\.jsonl$", logger.path.name)
|
||||
assert logger.path.is_file()
|
||||
|
||||
|
||||
@@ -833,7 +835,7 @@ def test_initial_custom_rollout_prompt_generates_seed_window_before_streaming():
|
||||
"rewrite_instruction": "A moonbase corridor thriller with flooding",
|
||||
"rewrite_model": "gpt-4.1-mini",
|
||||
"rewrite_temperature": 0.4,
|
||||
"timeout_ms": server_main.PROMPT_TIMEOUT_MS,
|
||||
"timeout_ms": PROMPT_TIMEOUT_MS,
|
||||
"system_prompt_override": "Session-specific rewrite prompt",
|
||||
}
|
||||
assert fake_pool._slot.calls
|
||||
@@ -1256,7 +1258,7 @@ def test_single5s_enhancement_fallback_does_not_start_generation():
|
||||
assert fallback_payloads[0]["prompt_id"] == "simple_custom_prompt"
|
||||
assert (
|
||||
fallback_payloads[0]["error"]
|
||||
== server_main.PROMPT_EXTENSION_FAILURE_USER_MESSAGE
|
||||
== PROMPT_EXTENSION_FAILURE_USER_MESSAGE
|
||||
)
|
||||
assert fallback_payloads[0]["source"] == "user_enhancement_failed"
|
||||
|
||||
|
||||
@@ -12,6 +12,7 @@ module import time.
|
||||
# mypy: ignore-errors
|
||||
import gc
|
||||
import os
|
||||
import re
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
@@ -20,11 +21,19 @@ import numpy as np
|
||||
import torch
|
||||
|
||||
from dreamverse.config import (
|
||||
AVAILABLE_LORAS,
|
||||
FRAME_HEIGHT,
|
||||
FRAME_WIDTH,
|
||||
MODEL_CONFIG,
|
||||
NUM_FRAMES,
|
||||
NUM_INFERENCE_STEPS,
|
||||
DREAMVERSE_MAX_AUTOTUNE,
|
||||
DREAMVERSE_SP_SIZE,
|
||||
DREAMVERSE_LORA_PATH,
|
||||
DREAMVERSE_LORA_NICKNAME,
|
||||
DREAMVERSE_LORA_STRENGTH,
|
||||
DREAMVERSE_LORA_STACK,
|
||||
_resolve_lora_spec,
|
||||
)
|
||||
|
||||
# Multi-frame decoded continuation defaults from
|
||||
@@ -62,6 +71,15 @@ DEFAULT_LTX2_AUDIO_HOP_LENGTH = 160
|
||||
DEFAULT_LTX2_AUDIO_DOWNSAMPLE = 4
|
||||
|
||||
|
||||
def _reset_lora_registry(worker) -> dict:
|
||||
pipeline = getattr(worker, "pipeline", None)
|
||||
if pipeline is None:
|
||||
return {"status": "no_pipeline"}
|
||||
pipeline.cur_adapter_name = ""
|
||||
pipeline.cur_adapter_strength = 1.0
|
||||
return {"status": "lora_registry_reset"}
|
||||
|
||||
|
||||
@dataclass
|
||||
class StepResult:
|
||||
"""Output of one generation step.
|
||||
@@ -199,6 +217,8 @@ class VideoGenerationWorker:
|
||||
self.continuation = ContinuationState()
|
||||
self.audio_encoder_module = None
|
||||
self.audio_processor_module = None
|
||||
self.active_style_prepend: list[str] = []
|
||||
self.active_style_append: list[str] = []
|
||||
|
||||
def _gpu_mem(self) -> str:
|
||||
a = torch.cuda.memory_allocated() / 1024**3
|
||||
@@ -257,6 +277,7 @@ class VideoGenerationWorker:
|
||||
or self.current_model_config["model_path"])
|
||||
|
||||
enable_compile = os.getenv("ENABLE_TORCH_COMPILE", "1") == "1"
|
||||
compile_mode = "max-autotune-no-cudagraphs" if DREAMVERSE_MAX_AUTOTUNE else None
|
||||
|
||||
components = ComponentConfig(
|
||||
config_root=config_model_path,
|
||||
@@ -269,7 +290,7 @@ class VideoGenerationWorker:
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=model_root,
|
||||
engine=EngineConfig(
|
||||
num_gpus=1,
|
||||
num_gpus=DREAMVERSE_SP_SIZE,
|
||||
offload=OffloadConfig(
|
||||
dit=False,
|
||||
dit_layerwise=False,
|
||||
@@ -280,9 +301,10 @@ class VideoGenerationWorker:
|
||||
compile=CompileConfig(
|
||||
enabled=enable_compile,
|
||||
text_encoder_enabled=enable_compile,
|
||||
vae_enabled=enable_compile,
|
||||
backend="inductor",
|
||||
fullgraph=True,
|
||||
mode="max-autotune-no-cudagraphs",
|
||||
mode=compile_mode,
|
||||
dynamic=False,
|
||||
),
|
||||
use_fsdp_inference=False,
|
||||
@@ -304,9 +326,88 @@ class VideoGenerationWorker:
|
||||
|
||||
self.generator = VideoGenerator.from_pretrained(config=generator_config)
|
||||
print(f"[GPU {self.gpu_id}] After model load: {self._gpu_mem()}")
|
||||
|
||||
lora_stack = DREAMVERSE_LORA_STACK or ([(DREAMVERSE_LORA_PATH,
|
||||
DREAMVERSE_LORA_STRENGTH)] if DREAMVERSE_LORA_PATH else [])
|
||||
for i, (lora_path, lora_strength) in enumerate(lora_stack):
|
||||
nickname = DREAMVERSE_LORA_NICKNAME if i == 0 else f"{DREAMVERSE_LORA_NICKNAME}_{i}"
|
||||
print(f"[GPU {self.gpu_id}] Applying LoRA '{nickname}' from {lora_path} "
|
||||
f"@{lora_strength} accumulate={i > 0}")
|
||||
self.generator.set_lora_adapter(
|
||||
lora_nickname=nickname,
|
||||
lora_path=lora_path,
|
||||
strength=lora_strength,
|
||||
accumulate=(i > 0),
|
||||
)
|
||||
if lora_stack:
|
||||
print(f"[GPU {self.gpu_id}] LoRA stack applied ({len(lora_stack)})")
|
||||
|
||||
self._load_audio_encoder(model_root)
|
||||
print(f"[GPU {self.gpu_id}] LTX2 model loaded (warmup pending)")
|
||||
|
||||
def apply_lora_stack(
|
||||
self,
|
||||
stack: list[tuple[str, float]],
|
||||
) -> tuple[str | None, str | None]:
|
||||
"""Re-apply a runtime LoRA stack and update the active style trigger."""
|
||||
if self.generator is None:
|
||||
raise RuntimeError("Generator not initialized; cannot apply LoRA stack.")
|
||||
|
||||
try:
|
||||
self.generator.unmerge_lora_weights()
|
||||
except Exception as exc:
|
||||
print(f"[GPU {self.gpu_id}] unmerge_lora_weights skipped: {exc}")
|
||||
|
||||
try:
|
||||
self.generator.executor.collective_rpc(_reset_lora_registry)
|
||||
except Exception as exc:
|
||||
print(f"[GPU {self.gpu_id}] lora registry reset skipped: {exc}")
|
||||
|
||||
resolved_stack: list[tuple[str, float]] = []
|
||||
for spec, strength in stack:
|
||||
resolved = _resolve_lora_spec(spec)
|
||||
if resolved:
|
||||
resolved_stack.append((resolved, float(strength)))
|
||||
|
||||
for i, (lora_path, lora_strength) in enumerate(resolved_stack):
|
||||
nickname = DREAMVERSE_LORA_NICKNAME if i == 0 else f"{DREAMVERSE_LORA_NICKNAME}_{i}"
|
||||
print(f"[GPU {self.gpu_id}] Re-applying LoRA '{nickname}' from {lora_path} "
|
||||
f"@{lora_strength} accumulate={i > 0}")
|
||||
self.generator.set_lora_adapter(
|
||||
lora_nickname=nickname,
|
||||
lora_path=lora_path,
|
||||
strength=lora_strength,
|
||||
accumulate=(i > 0),
|
||||
)
|
||||
print(f"[GPU {self.gpu_id}] Runtime LoRA stack applied ({len(resolved_stack)})")
|
||||
|
||||
prepend: list[str] = []
|
||||
append: list[str] = []
|
||||
for spec, _ in stack:
|
||||
key = (spec or "").strip().lower()
|
||||
if key in AVAILABLE_LORAS:
|
||||
trigger = AVAILABLE_LORAS[key]["trigger"]
|
||||
if AVAILABLE_LORAS[key].get("position", "append") == "prepend":
|
||||
prepend.append(trigger)
|
||||
else:
|
||||
append.append(trigger)
|
||||
self.active_style_prepend = prepend
|
||||
self.active_style_append = append
|
||||
combined = " ".join(prepend + append) or None
|
||||
return combined, None
|
||||
|
||||
def _inject_style_trigger(self, prompt: str) -> str:
|
||||
result = prompt
|
||||
for trigger in self.active_style_prepend:
|
||||
trigger = (trigger or "").strip()
|
||||
if trigger and not re.search(rf"\b{re.escape(trigger)}\b", result, re.IGNORECASE):
|
||||
result = f"{trigger} {result}".strip()
|
||||
for trigger in self.active_style_append:
|
||||
trigger = (trigger or "").strip()
|
||||
if trigger and not re.search(rf"\b{re.escape(trigger)}\b", result, re.IGNORECASE):
|
||||
result = f"{result} {trigger}".strip()
|
||||
return result
|
||||
|
||||
def _load_audio_encoder(self, model_root: str) -> None:
|
||||
if not ENABLE_AUDIO_RE_ENCODE:
|
||||
return
|
||||
@@ -375,6 +476,8 @@ class VideoGenerationWorker:
|
||||
"""Execute one generation step; snapshot state for the next segment."""
|
||||
timings: dict = {}
|
||||
|
||||
prompt = self._inject_style_trigger(prompt)
|
||||
|
||||
request_kwargs = dict(
|
||||
prompt=prompt,
|
||||
negative_prompt="",
|
||||
|
||||
@@ -48,6 +48,13 @@ class ReloadAck:
|
||||
user_id: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LoraAck:
|
||||
user_id: str | None
|
||||
style_trigger: str | None
|
||||
style_trigger_position: str | None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class WarmupComplete:
|
||||
user_id: str | None
|
||||
@@ -117,6 +124,7 @@ WorkerEvent = (StepComplete
|
||||
| JoinAck
|
||||
| LeaveAck
|
||||
| ReloadAck
|
||||
| LoraAck
|
||||
| WarmupComplete
|
||||
| MediaInit
|
||||
| MediaChunk
|
||||
@@ -154,4 +162,9 @@ class ReloadModelPayload:
|
||||
model_config: dict
|
||||
|
||||
|
||||
CommandPayload = UserStepPayload | WarmupPayload | ReloadModelPayload
|
||||
@dataclass(frozen=True)
|
||||
class LoraStackPayload:
|
||||
stack: list[tuple[str, float]]
|
||||
|
||||
|
||||
CommandPayload = (UserStepPayload | WarmupPayload | ReloadModelPayload | LoraStackPayload)
|
||||
|
||||
@@ -32,14 +32,21 @@
|
||||
# FFMPEG_NATIVE_CC explicit C compiler command for native builds
|
||||
# FFMPEG_NATIVE_CXX explicit C++ compiler command for native builds
|
||||
#
|
||||
# Toolchain selection: CC / CXX / AS are pinned to the conda-forge
|
||||
# triplet matching `uname -m`. Inherited values are intentionally
|
||||
# ignored — conda envs that have BOTH `gcc_linux-64` and
|
||||
# `gcc_linux-aarch64` installed export the cross-compiler triplet on
|
||||
# every `conda activate` (the `aarch64` activation script sorts later
|
||||
# and wins), which silently breaks x264's compiler probe on the
|
||||
# opposite host. If you genuinely need a non-host toolchain, set
|
||||
# FFMPEG_NATIVE_CC and/or FFMPEG_NATIVE_CXX explicitly.
|
||||
# Toolchain selection (precedence, highest first):
|
||||
# 1. FFMPEG_NATIVE_CC / FFMPEG_NATIVE_CXX, if set — explicit override.
|
||||
# 2. The conda-forge triplet matching `uname -m`
|
||||
# (`x86_64-conda-linux-gnu-cc` or `aarch64-conda-linux-gnu-cc`),
|
||||
# if present on PATH. Pinning the triplet sidesteps a bug where
|
||||
# conda envs with BOTH `gcc_linux-64` and `gcc_linux-aarch64`
|
||||
# installed export the cross-compiler triplet on every
|
||||
# `conda activate` (the `aarch64` activation script sorts later
|
||||
# and wins), silently breaking x264's compiler probe on the
|
||||
# opposite host.
|
||||
# 3. System `gcc` / `g++` — fallback so the script also works in a
|
||||
# plain venv or bare shell with no conda toolchain installed.
|
||||
# Inherited bare CC / CXX from the caller environment are ignored in
|
||||
# cases (2) and (3); use FFMPEG_NATIVE_CC/CXX to inject a non-host
|
||||
# toolchain.
|
||||
set -euo pipefail
|
||||
|
||||
# ─── Defaults (override via env) ──────────────────────────────────────────
|
||||
@@ -61,8 +68,8 @@ MAKE_JOBS="${MAKE_JOBS:-$(( NPROC < 16 ? NPROC : 16 ))}"
|
||||
ARCH="$(uname -m)"
|
||||
case "$ARCH" in
|
||||
x86_64)
|
||||
DEFAULT_CC=x86_64-conda-linux-gnu-cc
|
||||
DEFAULT_CXX=x86_64-conda-linux-gnu-c++
|
||||
CONDA_CC=x86_64-conda-linux-gnu-cc
|
||||
CONDA_CXX=x86_64-conda-linux-gnu-c++
|
||||
AS=nasm
|
||||
X264_CFLAGS="-O3 -march=native -mtune=native -fPIC -flto"
|
||||
X264_LDFLAGS="-flto -fuse-linker-plugin"
|
||||
@@ -70,8 +77,8 @@ case "$ARCH" in
|
||||
FFMPEG_LDFLAGS="-flto -Wl,-rpath,$INSTALL_PREFIX/lib"
|
||||
;;
|
||||
aarch64)
|
||||
DEFAULT_CC=aarch64-conda-linux-gnu-cc
|
||||
DEFAULT_CXX=aarch64-conda-linux-gnu-c++
|
||||
CONDA_CC=aarch64-conda-linux-gnu-cc
|
||||
CONDA_CXX=aarch64-conda-linux-gnu-c++
|
||||
unset AS # GNU as on ARM
|
||||
X264_CFLAGS="-O3 -mcpu=native -fPIC -flto"
|
||||
X264_LDFLAGS="-flto -fuse-linker-plugin"
|
||||
@@ -84,8 +91,27 @@ case "$ARCH" in
|
||||
;;
|
||||
esac
|
||||
|
||||
# Prefer the conda-forge triplet when it's actually on PATH; otherwise fall
|
||||
# back to system gcc/g++ so a plain venv works too. FFMPEG_NATIVE_CC/CXX
|
||||
# overrides both.
|
||||
if command -v -- "$CONDA_CC" >/dev/null 2>&1 \
|
||||
&& command -v -- "$CONDA_CXX" >/dev/null 2>&1; then
|
||||
DEFAULT_CC="$CONDA_CC"
|
||||
DEFAULT_CXX="$CONDA_CXX"
|
||||
default_source="conda-forge ($ARCH triplet)"
|
||||
else
|
||||
DEFAULT_CC=gcc
|
||||
DEFAULT_CXX=g++
|
||||
default_source="system gcc/g++"
|
||||
fi
|
||||
|
||||
CC="${FFMPEG_NATIVE_CC:-$DEFAULT_CC}"
|
||||
CXX="${FFMPEG_NATIVE_CXX:-$DEFAULT_CXX}"
|
||||
if [[ -n "${FFMPEG_NATIVE_CC:-}" || -n "${FFMPEG_NATIVE_CXX:-}" ]]; then
|
||||
toolchain_source="FFMPEG_NATIVE_CC/CXX override"
|
||||
else
|
||||
toolchain_source="$default_source"
|
||||
fi
|
||||
|
||||
require_compiler() {
|
||||
local name="$1" compiler="$2"
|
||||
@@ -95,6 +121,8 @@ require_compiler() {
|
||||
fi
|
||||
if ! command -v -- "$compiler" >/dev/null 2>&1; then
|
||||
echo "[install_native_ffmpeg] $name is unavailable: $compiler" >&2
|
||||
echo "[install_native_ffmpeg] install a compiler, or set FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX" \
|
||||
"to compiler commands on PATH." >&2
|
||||
exit 1
|
||||
fi
|
||||
if ! "$compiler" --version >/dev/null 2>&1; then
|
||||
@@ -107,7 +135,7 @@ require_compiler CC "$CC"
|
||||
require_compiler CXX "$CXX"
|
||||
export CC CXX
|
||||
[[ -n "${AS:-}" ]] && export AS
|
||||
echo "[install_native_ffmpeg] toolchain: CC=$CC CXX=$CXX AS=${AS:-<gnu-as>} (uname -m=$ARCH)"
|
||||
echo "[install_native_ffmpeg] toolchain: CC=$CC CXX=$CXX AS=${AS:-<gnu-as>} (uname -m=$ARCH, source: $toolchain_source)"
|
||||
|
||||
# ─── Step 0: probe required tools ─────────────────────────────────────────
|
||||
required=("$CC" "$CXX" make pkg-config git)
|
||||
@@ -118,7 +146,12 @@ for cmd in "${required[@]}"; do
|
||||
done
|
||||
if (( ${#missing[@]} > 0 )); then
|
||||
echo "[install_native_ffmpeg] missing required tools: ${missing[*]}" >&2
|
||||
echo "[install_native_ffmpeg] install them, or set FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX explicitly." >&2
|
||||
echo "[install_native_ffmpeg] install the missing tools. FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX" \
|
||||
"only override compiler selection; they do not provide make, pkg-config, git, or nasm." >&2
|
||||
if [[ " ${missing[*]} " == *" nasm "* ]]; then
|
||||
echo "[install_native_ffmpeg] nasm is required on x86_64 for x264 SIMD; install nasm" \
|
||||
"(for example via apt or conda-forge) and retry." >&2
|
||||
fi
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ NO_BROWSER=1 apps/dreamverse/scripts/launch/launch_demo.sh
|
||||
`launch_backend_dreamverse.sh` starts the full Dreamverse backend path used by
|
||||
the web app.
|
||||
|
||||
`launch_frontend.sh` starts the Next.js frontend and installs `pnpm`
|
||||
`launch_frontend.sh` starts the Next.js frontend and installs npm
|
||||
dependencies when `node_modules/` is missing.
|
||||
|
||||
`launch_backend_fastvideo.sh` starts the typed `fastvideo serve --config` path
|
||||
|
||||
@@ -7,8 +7,8 @@
|
||||
# FRONTEND_MODE=dev bash launch_frontend.sh # plain dev (5299, no devtools)
|
||||
# FRONTEND_MODE=single5s bash launch_frontend.sh # single-5s product mode
|
||||
#
|
||||
# The script ``cd``'s into ``web`` and shells out to pnpm. It runs
|
||||
# ``pnpm install --frozen-lockfile`` only when ``node_modules/`` is missing so repeat
|
||||
# The script ``cd``'s into ``web`` and shells out to npm. It runs
|
||||
# ``npm ci`` only when ``node_modules/`` is missing so repeat
|
||||
# launches are fast.
|
||||
|
||||
set -euo pipefail
|
||||
@@ -34,21 +34,21 @@ fi
|
||||
cd "${WEB_ROOT}"
|
||||
|
||||
if [[ ! -d node_modules ]]; then
|
||||
echo "[launch-demo] node_modules missing — running pnpm install --frozen-lockfile"
|
||||
pnpm install --frozen-lockfile
|
||||
echo "[launch-demo] node_modules missing — running npm ci"
|
||||
npm ci
|
||||
fi
|
||||
|
||||
case "${FRONTEND_MODE}" in
|
||||
devtools)
|
||||
echo "[launch-demo] starting Next.js dev:devtools (port 5274)"
|
||||
exec pnpm run dev:devtools -- "$@"
|
||||
exec npm run dev:devtools -- "$@"
|
||||
;;
|
||||
dev)
|
||||
echo "[launch-demo] starting Next.js dev (port 5299)"
|
||||
exec pnpm run dev -- "$@"
|
||||
exec npm run dev -- "$@"
|
||||
;;
|
||||
single5s)
|
||||
echo "[launch-demo] starting Next.js dev:single5s (port 5274)"
|
||||
exec pnpm run dev:single5s -- "$@"
|
||||
exec npm run dev:single5s -- "$@"
|
||||
;;
|
||||
esac
|
||||
|
||||
@@ -0,0 +1,127 @@
|
||||
# Dreamverse Modal deployment
|
||||
|
||||
This directory contains the Modal wrapper for running the registry-built
|
||||
Dreamverse Docker image on a Modal B200 container. The wrapper pulls a prebuilt Docker image and exposes the dreamverse application running inside.
|
||||
|
||||
|
||||
## 1. Install Modal CLI
|
||||
|
||||
Install and configure the Modal CLI for the target workspace/profile:
|
||||
|
||||
```bash
|
||||
pip install modal
|
||||
modal token set --token-id <token-id> --token-secret <token-secret> --profile=<profile-name>
|
||||
modal profile activate <profile-name>
|
||||
```
|
||||
|
||||
## 2. Create the API key secret
|
||||
|
||||
Dreamverse needs several API keys for prompt-rewriter LLMs and model access in the Modal secret named
|
||||
`dreamverse-api-keys`. Create or replace it with placeholders like this:
|
||||
|
||||
```bash
|
||||
modal secret create dreamverse-api-keys \
|
||||
CEREBRAS_API_KEY=<your-cerebras-key> \
|
||||
GROQ_API_KEY=<your-groq-key> \
|
||||
HF_TOKEN=<your-huggingface-token> \
|
||||
--force
|
||||
```
|
||||
|
||||
## 3. Deploy with Docker image
|
||||
|
||||
`DREAMVERSE_IMAGE` is required at deploy time:
|
||||
|
||||
```bash
|
||||
DREAMVERSE_IMAGE=ghcr.io/<org>/<repo>/dreamverse:<tag> \
|
||||
modal deploy apps/dreamverse/scripts/modal/modal_app.py
|
||||
```
|
||||
|
||||
Use a SHA-specific tag, not `latest`. Use a `dreamverse-backend-cuda12.9.1-sha-*`
|
||||
tag for backend-only deploys, or a `dreamverse-ui-cuda12.9.1-sha-*` tag for an
|
||||
image that includes the static UI served by the backend.
|
||||
|
||||
For local image build details, see
|
||||
`apps/dreamverse/docker/README.md` and `apps/dreamverse/docker/docker_build.sh`.
|
||||
|
||||
|
||||
## 4. Validate the deployment
|
||||
|
||||
Set the URL returned by `modal deploy`, then probe the backend:
|
||||
|
||||
```bash
|
||||
URL=https://<workspace>--dreamverse-b200-serve.modal.run
|
||||
curl -fsS --max-time 300 "$URL/healthz"
|
||||
curl -fsS --max-time 300 "$URL/status"
|
||||
curl -fsS --max-time 300 "$URL/readyz"
|
||||
```
|
||||
|
||||
Check logs with function and container IDs so you can confirm only one container
|
||||
is active:
|
||||
|
||||
```bash
|
||||
modal app logs dreamverse-b200 \
|
||||
--tail 200 \
|
||||
--timestamps \
|
||||
--show-function-id \
|
||||
--show-container-id
|
||||
```
|
||||
|
||||
Useful Modal commands:
|
||||
|
||||
```bash
|
||||
modal app list
|
||||
modal app list --json
|
||||
modal billing report --for today --resolution h
|
||||
modal app stop dreamverse-b200
|
||||
```
|
||||
|
||||
## Runtime details
|
||||
|
||||
### Volumes and assumed locations
|
||||
|
||||
`modal_app.py` creates the named volumes if they do not exist and mounts them at
|
||||
the locations expected by the image:
|
||||
|
||||
| Modal volume | Container path | Environment variable | Purpose |
|
||||
| --- | --- | --- | --- |
|
||||
| `dreamverse-hf-cache` | `/root/.cache/huggingface` | `HF_HOME` | Hugging Face model/cache data |
|
||||
| `dreamverse-state` | `/var/lib/dreamverse` | `FASTVIDEO_DREAMVERSE_HOME` | Dreamverse outputs, session logs, and runtime state |
|
||||
|
||||
### Useful dev env vars
|
||||
|
||||
Most deploys only need the Modal secret and required `DREAMVERSE_IMAGE`.
|
||||
`DREAMVERSE_IMAGE` is read from your local shell at deploy time; the runtime
|
||||
knobs below live in the image or `modal_app.py` unless you intentionally change them:
|
||||
|
||||
- `DREAMVERSE_IMAGE`: deploy a prebuilt SHA-specific image tag without editing `modal_app.py`.
|
||||
- `FASTVIDEO_ENABLE_DEVTOOLS`: enable Dreamverse devtools behavior.
|
||||
- `FASTVIDEO_PROMPT_PROVIDER` / `FASTVIDEO_PROMPT_*_MODEL`: try prompt-rewriter provider or model choices.
|
||||
- `FASTVIDEO_ENABLE_STARTUP_WARMUP`: trade slower startup for a warmer first request.
|
||||
- `DREAMVERSE_MAX_AUTOTUNE`: enable or disable PyTorch Inductor max-autotune for the compiled Dreamverse runtime.
|
||||
|
||||
The Modal wrapper defaults to torch compile with Inductor max-autotune enabled
|
||||
for the fastest generation after startup warmup. If you need shorter
|
||||
compile/warmup time, disable max-autotune via:
|
||||
|
||||
```bash
|
||||
DREAMVERSE_IMAGE=ghcr.io/<org>/<repo>/dreamverse:<tag> \
|
||||
DREAMVERSE_MAX_AUTOTUNE=0 \
|
||||
modal deploy apps/dreamverse/scripts/modal/modal_app.py
|
||||
```
|
||||
|
||||
### Autoscaling and cost safety
|
||||
|
||||
The current deployed script keeps exactly one B200 container warm by setting
|
||||
`min_containers=1` and `max_containers=1`. `min_containers=1` prevents Modal
|
||||
from scaling the deployment down to zero after idle periods, which avoids paying
|
||||
the expensive torch-compile/startup-warmup cost again on the next request.
|
||||
`max_containers=1` caps concurrency so concurrent requests queue onto the warm
|
||||
container instead of spawning additional B200 containers and multiplying cost.
|
||||
|
||||
### Common gotchas
|
||||
|
||||
- Modal profile/workspace matters: create the secret and deploy with the same active profile.
|
||||
- `modal secret create ... --force` replaces the existing `dreamverse-api-keys` secret.
|
||||
- Use the exact URL printed by `modal deploy`; the placeholder URL above is only an example shape.
|
||||
- B200 cold starts and first model downloads can make `/readyz` slow. Check logs before redeploying.
|
||||
- `modal app stop dreamverse-b200` stops the app; it is not a read-only inspection command.
|
||||
@@ -0,0 +1,77 @@
|
||||
# pyright: reportAttributeAccessIssue=false
|
||||
"""Modal deployment entrypoint for the Dreamverse B200 backend."""
|
||||
|
||||
import os
|
||||
import subprocess
|
||||
|
||||
import modal
|
||||
|
||||
IMAGE = os.environ.get("DREAMVERSE_IMAGE")
|
||||
if not IMAGE:
|
||||
raise RuntimeError(
|
||||
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
|
||||
"for example a dreamverse-backend-cuda12.9.1-sha-* tag or a "
|
||||
"dreamverse-ui-cuda12.9.1-sha-* tag if serving the static UI."
|
||||
)
|
||||
|
||||
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
|
||||
# ENTRYPOINT (``docker/docker_entrypoint.sh``). That entrypoint normally
|
||||
# ``:?``-validates these secret entries and ``source``-s ``ffmpeg-env.sh``;
|
||||
# neither runs on Modal, so we replicate both here.
|
||||
_REQUIRED_SECRET_KEYS: tuple[str, ...] = ("CEREBRAS_API_KEY", "GROQ_API_KEY")
|
||||
|
||||
image = modal.Image.from_registry(IMAGE)
|
||||
image = image.env({
|
||||
"DREAMVERSE_IMAGE": IMAGE,
|
||||
"HF_HOME": "/root/.cache/huggingface",
|
||||
"FASTVIDEO_DREAMVERSE_HOME": "/var/lib/dreamverse",
|
||||
"FASTVIDEO_ENABLE_STARTUP_WARMUP": "1",
|
||||
"FASTVIDEO_GPU_COUNT": "1",
|
||||
"ENABLE_TORCH_COMPILE": "1",
|
||||
"DREAMVERSE_MAX_AUTOTUNE": os.environ.get("DREAMVERSE_MAX_AUTOTUNE", "1"),
|
||||
"STREAM_MODE": "av_fmp4",
|
||||
# Native ffmpeg is built into ``/opt/ffmpeg-native`` by the Dockerfile but
|
||||
# is not on the image's ``PATH``. Without these, ``av_streaming``'s
|
||||
# ``shutil.which("ffmpeg")`` returns ``None`` and ``av_fmp4`` muxing has
|
||||
# no encoder. ``ffmpeg-env.sh`` would normally set them.
|
||||
"FASTVIDEO_FFMPEG_BIN": "/opt/ffmpeg-native/bin/ffmpeg",
|
||||
"FASTVIDEO_VIDEO_CODEC": "libx264",
|
||||
})
|
||||
|
||||
app = modal.App("dreamverse-b200")
|
||||
|
||||
hf_cache = modal.Volume.from_name("dreamverse-hf-cache", create_if_missing=True)
|
||||
dreamverse_state = modal.Volume.from_name("dreamverse-state", create_if_missing=True)
|
||||
|
||||
|
||||
@app.function(
|
||||
image=image,
|
||||
gpu="B200",
|
||||
cpu=16,
|
||||
memory=65536,
|
||||
timeout=7200,
|
||||
startup_timeout=4800,
|
||||
min_containers=1,
|
||||
max_containers=1,
|
||||
secrets=[modal.Secret.from_name("dreamverse-api-keys")],
|
||||
volumes={
|
||||
"/root/.cache/huggingface": hf_cache,
|
||||
"/var/lib/dreamverse": dreamverse_state,
|
||||
},
|
||||
)
|
||||
@modal.web_server(8009, startup_timeout=4800)
|
||||
def serve():
|
||||
# ``or ""`` collapses ``None`` (unset) into an empty string, ``.strip()``
|
||||
# collapses whitespace-only values (e.g. ``" "``) — both should be
|
||||
# treated as missing.
|
||||
missing = [
|
||||
k for k in _REQUIRED_SECRET_KEYS
|
||||
if not (os.environ.get(k) or "").strip()
|
||||
]
|
||||
if missing:
|
||||
raise RuntimeError(
|
||||
"dreamverse-api-keys secret is missing required entries: "
|
||||
f"{', '.join(missing)}. Add them with `modal secret create "
|
||||
"dreamverse-api-keys ... --force` and redeploy "
|
||||
"(see apps/dreamverse/scripts/modal/README.md).")
|
||||
subprocess.Popen(["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"])
|
||||
@@ -87,9 +87,9 @@ if [[ "${started_backend}" == "1" ]]; then
|
||||
echo "Backend log: ${BACKEND_LOG_PATH}"
|
||||
echo "Backend is still running so you can launch the frontend:"
|
||||
echo " cd ${ROOT_DIR}/web"
|
||||
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
|
||||
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} npm run dev"
|
||||
else
|
||||
echo "You can now launch the frontend:"
|
||||
echo " cd ${ROOT_DIR}/web"
|
||||
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
|
||||
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} npm run dev"
|
||||
fi
|
||||
|
||||
@@ -11,7 +11,7 @@ test.describe('backend health', () => {
|
||||
expect(response.ok()).toBeTruthy();
|
||||
const body = await response.json();
|
||||
expect(body.status).toBe('ok');
|
||||
expect(body.service).toBe('ltx2-streaming-backend');
|
||||
expect(['ltx2-streaming-backend', 'ltx2-streaming-mock-server']).toContain(body.service);
|
||||
});
|
||||
|
||||
test('readyz reports gpu pool state', async ({ request }) => {
|
||||
|
||||
@@ -18,7 +18,7 @@ import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
|
||||
* PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
|
||||
* NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
|
||||
* PLAYWRIGHT_LONG_RUNNING=1 \
|
||||
* pnpm exec playwright test e2e/long-running-segments.spec.ts
|
||||
* npm exec -- playwright test e2e/long-running-segments.spec.ts
|
||||
*
|
||||
* The full run takes ~7-9 minutes on a B200: ~3-4min for torch.compile
|
||||
* max-autotune to warm both DiT + text-encoder graphs, then ~30s for
|
||||
|
||||
@@ -0,0 +1,311 @@
|
||||
import { stat } from 'node:fs/promises';
|
||||
|
||||
import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
|
||||
|
||||
test.describe('mock-backed generation smoke', () => {
|
||||
test('accepts a custom prompt and enters the streaming state', async ({ page, request }) => {
|
||||
const health = await request.get('/healthz');
|
||||
const body = health.ok() ? await health.json() : {};
|
||||
test.skip(
|
||||
body.service !== 'ltx2-streaming-mock-server',
|
||||
'Mock-backed smoke requires dreamverse.mock_server on BACKEND_PORT (default 8009).',
|
||||
);
|
||||
|
||||
const wsEventTypes: string[] = [];
|
||||
|
||||
page.on('websocket', (ws: PWWebSocket) => {
|
||||
ws.on('framereceived', ({ payload }) => {
|
||||
if (typeof payload !== 'string') return;
|
||||
try {
|
||||
const parsed = JSON.parse(payload);
|
||||
if (typeof parsed?.type === 'string') wsEventTypes.push(parsed.type);
|
||||
} catch {
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
await page.goto('/');
|
||||
|
||||
await expect(page.getByRole('img', { name: 'FastVideo' })).toBeVisible();
|
||||
|
||||
const continuation = page.getByLabel('Continuation prompt');
|
||||
await expect(continuation).toBeVisible();
|
||||
await continuation.fill('A glass whale swims above a neon forest');
|
||||
|
||||
await page.getByRole('button', { name: /^generate$/i }).click();
|
||||
|
||||
await expect(continuation).toBeDisabled({ timeout: 30_000 });
|
||||
await expect(continuation).toHaveAttribute('placeholder', /generating video/i);
|
||||
await expect(page.getByRole('button', { name: /^leave$/i })).toBeVisible({ timeout: 30_000 });
|
||||
await expect(page.locator('video').first()).toHaveCount(1);
|
||||
|
||||
await expect.poll(() => wsEventTypes).toEqual(
|
||||
expect.arrayContaining(['ltx2_stream_start', 'ltx2_segment_start', 'media_init']),
|
||||
);
|
||||
});
|
||||
|
||||
test('streams, plays, and surfaces a downloadable clip', async ({ page, request }) => {
|
||||
const health = await request.get('/healthz');
|
||||
const body = health.ok() ? await health.json() : {};
|
||||
test.skip(
|
||||
body.service !== 'ltx2-streaming-mock-server',
|
||||
'Mock-backed playback assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
|
||||
);
|
||||
|
||||
await page.addInitScript(() => {
|
||||
(window as unknown as { __sharedFiles: unknown }).__sharedFiles = null;
|
||||
const stub = async (data: { files?: File[] }) => {
|
||||
const files = Array.isArray(data?.files) ? data.files : [];
|
||||
(window as unknown as { __sharedFiles: unknown }).__sharedFiles = files.map((f) => ({
|
||||
name: f.name,
|
||||
type: f.type,
|
||||
size: f.size,
|
||||
}));
|
||||
};
|
||||
Object.defineProperty(Navigator.prototype, 'share', {
|
||||
value: stub,
|
||||
configurable: true,
|
||||
writable: true,
|
||||
});
|
||||
Object.defineProperty(Navigator.prototype, 'canShare', {
|
||||
value: (data: { files?: File[] }) => Array.isArray(data?.files),
|
||||
configurable: true,
|
||||
writable: true,
|
||||
});
|
||||
});
|
||||
|
||||
await page.goto('/');
|
||||
|
||||
const continuation = page.getByLabel('Continuation prompt');
|
||||
await expect(continuation).toBeVisible();
|
||||
await continuation.fill('A glass whale swims above a neon forest');
|
||||
await page.getByRole('button', { name: /^generate$/i }).click();
|
||||
|
||||
const liveVideo = page.locator('video:not(.hidden)');
|
||||
await expect(liveVideo).toHaveCount(1);
|
||||
|
||||
await test.step('MSE pipeline attaches the live <video>', async () => {
|
||||
// Chromium/Firefox attach via blob: URL; Safari uses srcObject with ManagedMediaSource.
|
||||
await expect
|
||||
.poll(
|
||||
async () =>
|
||||
liveVideo.evaluate(
|
||||
(v: HTMLVideoElement) =>
|
||||
v.src.startsWith('blob:') || v.srcObject !== null,
|
||||
),
|
||||
{ timeout: 30_000 },
|
||||
)
|
||||
.toBe(true);
|
||||
});
|
||||
|
||||
await test.step('SourceBuffer accepts fMP4 and decoder produces frames', async () => {
|
||||
await expect
|
||||
.poll(
|
||||
async () =>
|
||||
liveVideo.evaluate((v: HTMLVideoElement) =>
|
||||
v.buffered.length > 0 ? v.buffered.end(0) : 0,
|
||||
),
|
||||
{ timeout: 60_000 },
|
||||
)
|
||||
.toBeGreaterThan(0);
|
||||
|
||||
// readyState >= 3 = HAVE_FUTURE_DATA.
|
||||
await expect
|
||||
.poll(
|
||||
async () => liveVideo.evaluate((v: HTMLVideoElement) => v.readyState),
|
||||
{ timeout: 60_000 },
|
||||
)
|
||||
.toBeGreaterThanOrEqual(3);
|
||||
});
|
||||
|
||||
await test.step('<video> playback advances past the first second', async () => {
|
||||
await liveVideo.evaluate(async (v: HTMLVideoElement) => {
|
||||
if (v.paused) {
|
||||
try {
|
||||
await v.play();
|
||||
} catch {
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
await expect
|
||||
.poll(
|
||||
async () => liveVideo.evaluate((v: HTMLVideoElement) => v.currentTime),
|
||||
{ timeout: 30_000 },
|
||||
)
|
||||
.toBeGreaterThan(0);
|
||||
|
||||
await expect
|
||||
.poll(
|
||||
async () =>
|
||||
liveVideo.evaluate((v: HTMLVideoElement) =>
|
||||
v.ended ? v.duration : v.currentTime,
|
||||
),
|
||||
{ timeout: 10_000 },
|
||||
)
|
||||
.toBeGreaterThanOrEqual(1.0);
|
||||
});
|
||||
|
||||
await test.step('the AAC audio track is demuxed and decoded', async () => {
|
||||
// The mock emits AAC (mp4a.40.2), matching the real backend
|
||||
// (av_streaming.py). Proves the FE decoded the audio track.
|
||||
await expect
|
||||
.poll(
|
||||
async () =>
|
||||
liveVideo.evaluate((el: HTMLVideoElement) => {
|
||||
const v = el as HTMLVideoElement & {
|
||||
audioTracks?: { length: number };
|
||||
mozHasAudio?: boolean;
|
||||
webkitAudioDecodedByteCount?: number;
|
||||
};
|
||||
return (
|
||||
(v.audioTracks?.length ?? 0) > 0 ||
|
||||
v.mozHasAudio === true ||
|
||||
(v.webkitAudioDecodedByteCount ?? 0) > 0
|
||||
);
|
||||
}),
|
||||
{ timeout: 10_000 },
|
||||
)
|
||||
.toBe(true);
|
||||
});
|
||||
|
||||
await test.step('completed clip surfaces a working download', async () => {
|
||||
const downloadButton = page.getByRole('button', { name: /download video|share video/i });
|
||||
await expect(downloadButton).toBeVisible({ timeout: 30_000 });
|
||||
|
||||
const buttonLabel = (await downloadButton.getAttribute('aria-label')) ?? '';
|
||||
const isShareFlow = /share video/i.test(buttonLabel);
|
||||
|
||||
if (isShareFlow) {
|
||||
await downloadButton.click();
|
||||
await expect
|
||||
.poll(
|
||||
async () => page.evaluate(() => (window as { __sharedFiles?: unknown }).__sharedFiles),
|
||||
{ timeout: 10_000 },
|
||||
)
|
||||
.not.toBeNull();
|
||||
const shared = (await page.evaluate(
|
||||
() => (window as { __sharedFiles?: Array<{ name: string; type: string; size: number }> }).__sharedFiles,
|
||||
)) ?? [];
|
||||
expect(shared).toHaveLength(1);
|
||||
expect(shared[0].name).toMatch(/\.(mp4|webm)$/);
|
||||
expect(shared[0].size).toBeGreaterThan(0);
|
||||
} else {
|
||||
const downloadPromise = page.waitForEvent('download');
|
||||
await downloadButton.click();
|
||||
const download = await downloadPromise;
|
||||
expect(download.suggestedFilename()).toMatch(/\.(mp4|webm)$/);
|
||||
|
||||
const savedPath = await download.path();
|
||||
expect(savedPath).not.toBeNull();
|
||||
const { size } = await stat(savedPath!);
|
||||
expect(size).toBeGreaterThan(0);
|
||||
}
|
||||
});
|
||||
|
||||
await test.step('project history sidebar lists the current session', async () => {
|
||||
const sidebar = page.getByRole('complementary', { name: 'Project history' });
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
|
||||
await expect(sidebar).toBeInViewport();
|
||||
|
||||
await expect(sidebar.getByText('Current', { exact: true })).toBeVisible();
|
||||
await expect(sidebar.getByText('Active', { exact: true })).toBeVisible();
|
||||
});
|
||||
|
||||
const lastError = await liveVideo.evaluate((v: HTMLVideoElement) =>
|
||||
v.error ? `${v.error.code}: ${v.error.message ?? ''}` : null,
|
||||
);
|
||||
expect(lastError).toBeNull();
|
||||
});
|
||||
|
||||
test('starts a new project and switches back to the prior session', async ({ page, request }) => {
|
||||
const health = await request.get('/healthz');
|
||||
const body = health.ok() ? await health.json() : {};
|
||||
test.skip(
|
||||
body.service !== 'ltx2-streaming-mock-server',
|
||||
'Mock-backed project lifecycle assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
|
||||
);
|
||||
|
||||
await page.goto('/');
|
||||
|
||||
await test.step('complete a generation so there is a project to save', async () => {
|
||||
const continuation = page.getByLabel('Continuation prompt');
|
||||
await expect(continuation).toBeVisible();
|
||||
await continuation.fill('Aurora over a frozen lake');
|
||||
await page.getByRole('button', { name: /^generate$/i }).click();
|
||||
await expect(
|
||||
page.getByRole('button', { name: /download video|share video/i }),
|
||||
).toBeVisible({ timeout: 60_000 });
|
||||
});
|
||||
|
||||
const sidebar = page.getByRole('complementary', { name: 'Project history' });
|
||||
|
||||
await test.step('"New project" closes the sidebar and resets the composer', async () => {
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
await expect(sidebar).toBeInViewport();
|
||||
await sidebar.getByRole('button', { name: /^new project$/i }).click();
|
||||
await expect(sidebar).not.toBeInViewport();
|
||||
await expect(page.getByRole('button', { name: /^generate$/i })).toBeVisible({ timeout: 30_000 });
|
||||
const continuation = page.getByLabel('Continuation prompt');
|
||||
await expect(continuation).toBeEnabled();
|
||||
await expect(continuation).toHaveValue('');
|
||||
});
|
||||
|
||||
await test.step('the prior session appears under timeline', async () => {
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
await expect(sidebar).toBeInViewport();
|
||||
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
|
||||
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
|
||||
});
|
||||
|
||||
await test.step('clicking the prior session enters viewing mode', async () => {
|
||||
const priorRow = sidebar.locator('div[role="button"]').filter({ hasText: /just now|\d+m ago/ }).first();
|
||||
await priorRow.click();
|
||||
await expect(sidebar).not.toBeInViewport();
|
||||
await expect(page.locator('video[autoplay][loop]')).toBeVisible({ timeout: 30_000 });
|
||||
});
|
||||
});
|
||||
|
||||
test('saved projects persist across a page reload', async ({ page, request }) => {
|
||||
const health = await request.get('/healthz');
|
||||
const body = health.ok() ? await health.json() : {};
|
||||
test.skip(
|
||||
body.service !== 'ltx2-streaming-mock-server',
|
||||
'Mock-backed persistence assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
|
||||
);
|
||||
|
||||
await page.goto('/');
|
||||
|
||||
await test.step('complete a generation', async () => {
|
||||
const continuation = page.getByLabel('Continuation prompt');
|
||||
await expect(continuation).toBeVisible();
|
||||
await continuation.fill('Aurora over a frozen lake');
|
||||
await page.getByRole('button', { name: /^generate$/i }).click();
|
||||
await expect(
|
||||
page.getByRole('button', { name: /download video|share video/i }),
|
||||
).toBeVisible({ timeout: 60_000 });
|
||||
});
|
||||
|
||||
const sidebar = page.getByRole('complementary', { name: 'Project history' });
|
||||
|
||||
await test.step('persist via "New project" and confirm it lands under Previous', async () => {
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
await expect(sidebar).toBeInViewport();
|
||||
await sidebar.getByRole('button', { name: /^new project$/i }).click();
|
||||
await expect(page.getByRole('button', { name: /^generate$/i })).toBeVisible({ timeout: 30_000 });
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
await expect(sidebar).toBeInViewport();
|
||||
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
|
||||
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
|
||||
});
|
||||
|
||||
await test.step('after page reload, the prior project is still in Previous', async () => {
|
||||
await page.reload();
|
||||
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
|
||||
await expect(sidebar).toBeInViewport();
|
||||
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
|
||||
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -6,10 +6,13 @@ const backendHost = process.env.BACKEND_HOST || '127.0.0.1';
|
||||
const backendPort = Number(process.env.BACKEND_PORT) || 8009;
|
||||
const backendUrl = `http://${backendHost}:${backendPort}`;
|
||||
const configDir = path.dirname(fileURLToPath(import.meta.url));
|
||||
const staticExport = process.env.NEXT_OUTPUT_EXPORT === '1';
|
||||
|
||||
const nextConfig: NextConfig = {
|
||||
...(staticExport ? { output: 'export' as const } : {}),
|
||||
...(staticExport ? { images: { unoptimized: true } } : {}),
|
||||
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
|
||||
async rewrites() {
|
||||
...(staticExport ? {} : { async rewrites() {
|
||||
return [
|
||||
{
|
||||
source: '/ws',
|
||||
@@ -47,8 +50,16 @@ const nextConfig: NextConfig = {
|
||||
source: '/curated-presets/:path*',
|
||||
destination: `${backendUrl}/curated-presets/:path*`,
|
||||
},
|
||||
{
|
||||
source: '/lora',
|
||||
destination: `${backendUrl}/lora`,
|
||||
},
|
||||
{
|
||||
source: '/lora/:path*',
|
||||
destination: `${backendUrl}/lora/:path*`,
|
||||
},
|
||||
];
|
||||
},
|
||||
} }),
|
||||
webpack: (config) => {
|
||||
config.module.rules.push({
|
||||
test: /\.jsonl$/,
|
||||
|
||||
Generated
+11669
-11443
File diff suppressed because it is too large
Load Diff
@@ -38,7 +38,7 @@
|
||||
"geist": "^1.7.0",
|
||||
"lucide-react": "^0.577.0",
|
||||
"mp4box": "^2.3.0",
|
||||
"next": "^15.3.3",
|
||||
"next": "15.5.18",
|
||||
"radix-ui": "^1.4.3",
|
||||
"react": "^19.1.0",
|
||||
"react-dom": "^19.1.0",
|
||||
@@ -59,9 +59,16 @@
|
||||
"@vitest/coverage-v8": "^3.2.4",
|
||||
"jsdom": "^26.1.0",
|
||||
"mock-socket": "^9.3.1",
|
||||
"postcss": "^8.5.8",
|
||||
"postcss": "8.5.10",
|
||||
"tailwindcss": "^4.2.1",
|
||||
"typescript": "^5.8.3",
|
||||
"vitest": "^3.2.4"
|
||||
},
|
||||
"overrides": {
|
||||
"esbuild": "0.25.12",
|
||||
"picomatch": "4.0.4",
|
||||
"postcss": "8.5.10",
|
||||
"rollup": "4.59.0",
|
||||
"vite": "7.3.2"
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,12 +1,16 @@
|
||||
import { defineConfig, devices } from '@playwright/test';
|
||||
|
||||
const backendHost = process.env.BACKEND_HOST || '127.0.0.1';
|
||||
const backendPort = process.env.BACKEND_PORT || '8009';
|
||||
|
||||
/**
|
||||
* Playwright config for Dreamverse end-to-end tests.
|
||||
*
|
||||
* Tests assume the Dreamverse Python server is reachable at
|
||||
* BACKEND_HOST:BACKEND_PORT (default 127.0.0.1:8009) and the Next.js frontend
|
||||
* runs on port 5299. The webServer block boots `pnpm run dev` if no
|
||||
* server is already listening so tests work both locally and in CI.
|
||||
* runs on port 5299. The webServer block boots the mock backend and
|
||||
* `npm run dev` if no server is already listening so tests work both
|
||||
* locally and in CI.
|
||||
*/
|
||||
export default defineConfig({
|
||||
testDir: './e2e',
|
||||
@@ -26,23 +30,34 @@ export default defineConfig({
|
||||
viewport: { width: 1280, height: 720 },
|
||||
screenshot: 'only-on-failure',
|
||||
trace: 'retain-on-failure',
|
||||
video: process.env.CI ? 'retain-on-failure' : 'on',
|
||||
},
|
||||
projects: [
|
||||
{
|
||||
name: 'chromium',
|
||||
use: { ...devices['Desktop Chrome'] },
|
||||
},
|
||||
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
|
||||
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
|
||||
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
|
||||
{ name: 'msedge', use: { ...devices['Desktop Edge'], channel: 'msedge' } },
|
||||
{ name: 'mobile-safari', use: { ...devices['iPhone 14'] } },
|
||||
{ name: 'mobile-chromium', use: { ...devices['Pixel 7'] } },
|
||||
],
|
||||
webServer: process.env.PLAYWRIGHT_SKIP_WEBSERVER
|
||||
? undefined
|
||||
: {
|
||||
command: 'pnpm run dev',
|
||||
url: 'http://127.0.0.1:5299',
|
||||
reuseExistingServer: true,
|
||||
timeout: 120_000,
|
||||
env: {
|
||||
BACKEND_HOST: process.env.BACKEND_HOST || '127.0.0.1',
|
||||
BACKEND_PORT: process.env.BACKEND_PORT || '8009',
|
||||
: [
|
||||
{
|
||||
command: `PYTHONPATH=.. python3 -m uvicorn dreamverse.mock_server:app --host ${backendHost} --port ${backendPort}`,
|
||||
url: `http://${backendHost}:${backendPort}/healthz`,
|
||||
reuseExistingServer: true,
|
||||
timeout: 120_000,
|
||||
},
|
||||
},
|
||||
{
|
||||
command: 'npm run dev',
|
||||
url: 'http://127.0.0.1:5299',
|
||||
reuseExistingServer: true,
|
||||
timeout: 120_000,
|
||||
env: {
|
||||
BACKEND_HOST: backendHost,
|
||||
BACKEND_PORT: backendPort,
|
||||
},
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
Generated
-5199
File diff suppressed because it is too large
Load Diff
@@ -2489,6 +2489,8 @@ export default function Page() {
|
||||
if (devtoolsMode) {
|
||||
return (
|
||||
<DevtoolsShell
|
||||
videoRef={videoRefCallback}
|
||||
archivedPlaybackRef={archivedPlaybackRefCallback}
|
||||
connected={connected as boolean}
|
||||
gpuAssigned={gpuAssigned as boolean}
|
||||
sessionStarted={sessionStarted as boolean}
|
||||
|
||||
@@ -161,6 +161,7 @@ export default function VideoPlayer({
|
||||
}}
|
||||
size="icon"
|
||||
variant="outline"
|
||||
aria-label={canShare ? "Share video" : "Download video"}
|
||||
className="absolute top-3 left-3 z-10 cursor-pointer bg-slate-800/50 text-white/90 shadow-md backdrop-blur-sm transition-all border-white/30 hover:bg-slate-800/85 hover:border-white/50 hover:text-white hover:scale-105"
|
||||
>
|
||||
{canShare ? <Share className="size-5" /> : <Download className="size-5" />}
|
||||
|
||||
@@ -9,6 +9,7 @@ import { Label } from '@/components/ui/label';
|
||||
import { ScrollArea } from '@/components/ui/scroll-area';
|
||||
import { Textarea } from '@/components/ui/textarea';
|
||||
import { cn } from '@/lib/utils';
|
||||
import LoraControls from '@/components/devtools/LoraControls';
|
||||
|
||||
interface DevtoolsDrawerProps {
|
||||
editableMode?: boolean;
|
||||
@@ -126,6 +127,8 @@ export default function DevtoolsDrawer({
|
||||
</div>
|
||||
|
||||
<div className="grid gap-4 xl:grid-cols-2">
|
||||
<LoraControls />
|
||||
|
||||
{/* Prompt window memory */}
|
||||
<details className="group overflow-hidden rounded-2xl border border-border bg-card shadow-sm" open>
|
||||
<summary className="flex cursor-pointer list-none items-start justify-between gap-4 px-5 py-4">
|
||||
|
||||
@@ -113,6 +113,8 @@ interface DevtoolsShellProps {
|
||||
formatTime?: (seconds: number) => string;
|
||||
formatDurationMs?: (durationMs: number) => string;
|
||||
onPlaying?: () => void;
|
||||
videoRef?: React.RefCallback<HTMLVideoElement>;
|
||||
archivedPlaybackRef?: React.RefCallback<HTMLVideoElement>;
|
||||
}
|
||||
|
||||
export default function DevtoolsShell({
|
||||
@@ -209,6 +211,8 @@ export default function DevtoolsShell({
|
||||
formatTime = (seconds) => `${seconds}`,
|
||||
formatDurationMs = (durationMs) => `${durationMs}`,
|
||||
onPlaying = () => {},
|
||||
videoRef,
|
||||
archivedPlaybackRef,
|
||||
}: DevtoolsShellProps) {
|
||||
return (
|
||||
<WorkspaceShell
|
||||
@@ -235,6 +239,8 @@ export default function DevtoolsShell({
|
||||
workspace={
|
||||
<div className="flex flex-col gap-4">
|
||||
<VideoPlayer
|
||||
videoRef={videoRef}
|
||||
archivedPlaybackRef={archivedPlaybackRef}
|
||||
activeClip={activeClip}
|
||||
sessionStarted={sessionStarted}
|
||||
avPlaybackStarted={avPlaybackStarted}
|
||||
|
||||
@@ -0,0 +1,181 @@
|
||||
'use client';
|
||||
|
||||
import React, { useEffect, useRef, useState } from 'react';
|
||||
|
||||
import { Badge } from '@/components/ui/badge';
|
||||
import { Label } from '@/components/ui/label';
|
||||
|
||||
const STYLE_LABELS: Record<string, string> = {
|
||||
none: 'None',
|
||||
pixar: 'Pixar Toon',
|
||||
transition: 'Transition',
|
||||
};
|
||||
|
||||
export default function LoraControls() {
|
||||
const [styleKeys, setStyleKeys] = useState<string[]>([]);
|
||||
const [labels, setLabels] = useState<Record<string, string>>(STYLE_LABELS);
|
||||
const [strength, setStrength] = useState(0.8);
|
||||
const [enabled, setEnabled] = useState<Record<string, boolean>>({});
|
||||
const [intensity, setIntensity] = useState<Record<string, number>>({});
|
||||
const [status, setStatus] = useState('idle');
|
||||
|
||||
const strengthRef = useRef(0.8);
|
||||
const enabledRef = useRef<Record<string, boolean>>({});
|
||||
const intensityRef = useRef<Record<string, number>>({});
|
||||
const debounceRef = useRef<ReturnType<typeof setTimeout> | null>(null);
|
||||
const requestIdRef = useRef(0);
|
||||
|
||||
useEffect(() => {
|
||||
fetch('/lora/options')
|
||||
.then((response) => (response.ok ? response.json() : null))
|
||||
.then((data) => {
|
||||
if (data && Array.isArray(data.styles)) {
|
||||
const keys = data.styles.filter((s: string) => s !== 'none');
|
||||
setStyleKeys(keys);
|
||||
const initIntensity: Record<string, number> = {};
|
||||
keys.forEach((k: string) => {
|
||||
initIntensity[k] = 1.0;
|
||||
});
|
||||
setIntensity(initIntensity);
|
||||
intensityRef.current = initIntensity;
|
||||
}
|
||||
if (data && data.labels && typeof data.labels === 'object') {
|
||||
setLabels((prev) => ({ ...prev, ...data.labels }));
|
||||
}
|
||||
})
|
||||
.catch(() => setStatus('options unavailable'));
|
||||
|
||||
return () => {
|
||||
if (debounceRef.current) {
|
||||
clearTimeout(debounceRef.current);
|
||||
}
|
||||
};
|
||||
}, []);
|
||||
|
||||
const applyNow = () => {
|
||||
const reqId = ++requestIdRef.current;
|
||||
const stylesPayload: Record<string, number> = {};
|
||||
for (const key of Object.keys(enabledRef.current)) {
|
||||
if (enabledRef.current[key]) {
|
||||
stylesPayload[key] = intensityRef.current[key] ?? 1.0;
|
||||
}
|
||||
}
|
||||
setStatus('applying…');
|
||||
fetch('/lora', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ strength: strengthRef.current, styles: stylesPayload }),
|
||||
})
|
||||
.then(async (response) => {
|
||||
if (!response.ok) {
|
||||
throw new Error(await response.text());
|
||||
}
|
||||
return response.json();
|
||||
})
|
||||
.then((data) => {
|
||||
if (reqId !== requestIdRef.current) return;
|
||||
const active = Object.keys(data?.styles ?? {});
|
||||
const desc = active.length
|
||||
? active.map((k) => `${labels[k] ?? k}@${data.styles[k]}`).join(' + ')
|
||||
: 'no style';
|
||||
setStatus(`applied OmniNFT@${data.strength} + ${desc}`);
|
||||
})
|
||||
.catch((error) => {
|
||||
if (reqId !== requestIdRef.current) return;
|
||||
setStatus(`error: ${String(error).slice(0, 120)}`);
|
||||
});
|
||||
};
|
||||
|
||||
const scheduleApply = () => {
|
||||
if (debounceRef.current) {
|
||||
clearTimeout(debounceRef.current);
|
||||
}
|
||||
debounceRef.current = setTimeout(applyNow, 250);
|
||||
};
|
||||
|
||||
return (
|
||||
<details
|
||||
className="group overflow-hidden rounded-2xl border border-border bg-card shadow-sm xl:col-span-2"
|
||||
open
|
||||
>
|
||||
<summary className="flex cursor-pointer list-none items-start justify-between gap-4 px-5 py-4">
|
||||
<div className="space-y-1">
|
||||
<span className="block text-lg font-semibold text-foreground">
|
||||
LoRA stack
|
||||
</span>
|
||||
<span className="block text-sm leading-6 text-muted-foreground">
|
||||
OmniNFT strength + stackable style adapters, each with its own intensity (live).
|
||||
</span>
|
||||
</div>
|
||||
<Badge variant="secondary">{status}</Badge>
|
||||
</summary>
|
||||
<div className="space-y-5 border-t border-border px-5 py-4">
|
||||
<div className="space-y-2">
|
||||
<Label htmlFor="lora-strength">
|
||||
OmniNFT strength · {strength.toFixed(2)}
|
||||
</Label>
|
||||
<input
|
||||
id="lora-strength"
|
||||
type="range"
|
||||
min={0}
|
||||
max={1}
|
||||
step={0.05}
|
||||
value={strength}
|
||||
onChange={(event) => {
|
||||
const next = Number(event.target.value);
|
||||
setStrength(next);
|
||||
strengthRef.current = next;
|
||||
scheduleApply();
|
||||
}}
|
||||
className="w-full accent-sky-400"
|
||||
/>
|
||||
</div>
|
||||
|
||||
{styleKeys.length === 0 ? (
|
||||
<p className="text-sm text-muted-foreground">
|
||||
No style adapters for the active model.
|
||||
</p>
|
||||
) : (
|
||||
styleKeys.map((key) => (
|
||||
<div key={key} className="space-y-2">
|
||||
<div className="flex items-center justify-between">
|
||||
<label className="flex cursor-pointer items-center gap-2 text-sm font-medium text-foreground">
|
||||
<input
|
||||
type="checkbox"
|
||||
checked={!!enabled[key]}
|
||||
onChange={(event) => {
|
||||
const next = { ...enabledRef.current, [key]: event.target.checked };
|
||||
enabledRef.current = next;
|
||||
setEnabled(next);
|
||||
scheduleApply();
|
||||
}}
|
||||
className="accent-sky-400"
|
||||
/>
|
||||
{labels[key] ?? STYLE_LABELS[key] ?? key}
|
||||
</label>
|
||||
<span className="text-xs text-muted-foreground">
|
||||
{(intensity[key] ?? 1).toFixed(2)}
|
||||
</span>
|
||||
</div>
|
||||
<input
|
||||
type="range"
|
||||
min={0}
|
||||
max={1}
|
||||
step={0.05}
|
||||
value={intensity[key] ?? 1}
|
||||
disabled={!enabled[key]}
|
||||
onChange={(event) => {
|
||||
const next = { ...intensityRef.current, [key]: Number(event.target.value) };
|
||||
intensityRef.current = next;
|
||||
setIntensity(next);
|
||||
scheduleApply();
|
||||
}}
|
||||
className="w-full accent-sky-400 disabled:opacity-40"
|
||||
/>
|
||||
</div>
|
||||
))
|
||||
)}
|
||||
</div>
|
||||
</details>
|
||||
);
|
||||
}
|
||||
@@ -1,106 +0,0 @@
|
||||
# FVD (Fréchet Video Distance) Benchmark
|
||||
|
||||
Evaluate generated video quality using FVD with the I3D feature extractor.
|
||||
|
||||
## Quick Start
|
||||
|
||||
**Run the benchmark:**
|
||||
|
||||
```bash
|
||||
bash benchmarks/scripts/run.sh
|
||||
```
|
||||
|
||||
That's it! The script auto-installs dependencies and runs the benchmark.
|
||||
|
||||
**To customize:** Edit `benchmarks/fvd/run_fvd.py` to change:
|
||||
- Video paths (`real_dir`, `gen_dir`)
|
||||
- Number of videos, frames, sampling strategy
|
||||
- Device, batch size, caching, etc.
|
||||
|
||||
## Advanced Usage (CLI)
|
||||
|
||||
For more control without editing Python files, use the CLI.
|
||||
|
||||
**First-time setup** (one-time per pod/environment):
|
||||
|
||||
```bash
|
||||
bash benchmarks/scripts/setup_fvd.sh
|
||||
```
|
||||
|
||||
Then run any configuration you want:
|
||||
|
||||
```bash
|
||||
# Custom configuration
|
||||
python -m benchmarks.fvd.cli \
|
||||
--real-path data/real/ \
|
||||
--gen-path outputs/gen/ \
|
||||
--num-videos 1024 \
|
||||
--num-frames 32 \
|
||||
--clip-strategy random \
|
||||
--batch-size 32 \
|
||||
--seed 42 \
|
||||
--extractor clip
|
||||
```
|
||||
|
||||
**Standard protocols:**
|
||||
|
||||
```bash
|
||||
# Use predefined protocols
|
||||
python -m benchmarks.fvd.cli \
|
||||
--real-path data/real/ \
|
||||
--gen-path outputs/gen/ \
|
||||
--protocol fvd2048_16f # or fvd2048_128f, quick_test, etc.
|
||||
```
|
||||
|
||||
This would use i3d model by default as the feature extractor
|
||||
|
||||
**Feature caching** (speed up repeated evaluations):
|
||||
|
||||
```bash
|
||||
python -m benchmarks.fvd.cli \
|
||||
--real-path data/real/ \
|
||||
--gen-path outputs/gen/ \
|
||||
--protocol fvd2048_16f \
|
||||
--cache-real-features fvd-cache/extractor_name # Directory path (will save/load fvd-cache/extractor_name/extractor-name_real_features.pkl)
|
||||
```
|
||||
|
||||
Run `python -m benchmarks.fvd.cli --help` for all options.
|
||||
|
||||
## Available Protocols
|
||||
|
||||
- `fvd2048_16f` - Standard (2048 videos, 16 frames)
|
||||
- `fvd2048_128f` - Long videos (128 frames)
|
||||
- `fvd2048_128f_subsample8` - Subsampled long videos
|
||||
- `quick_test` - Fast testing (10 videos)
|
||||
|
||||
## Configuration Options
|
||||
|
||||
Key options in `FVDConfig`:
|
||||
|
||||
```python
|
||||
num_videos=2048, # Videos to evaluate
|
||||
num_frames_per_clip=16, # Frames per clip
|
||||
clip_strategy='beginning', # beginning|random|uniform|middle|sliding
|
||||
frame_stride=1, # Frame subsampling
|
||||
batch_size=32, # GPU batch size
|
||||
device='cuda', # cuda|cpu
|
||||
cache_real_features=None, # Cache path for speed
|
||||
seed=42, # Reproducibility
|
||||
extractor='i3d', # i3d|clip|videomae
|
||||
```
|
||||
|
||||
## Programmatic Usage
|
||||
|
||||
```python
|
||||
from benchmarks.fvd import compute_fvd_with_config, FVDConfig
|
||||
|
||||
config = FVDConfig.fvd2048_16f() # or custom config
|
||||
results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
|
||||
print(f"FVD: {results['fvd']:.2f}")
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- Requires minimum 10 frames per clip
|
||||
- Supports both video files (.mp4, .avi, etc.) and frame directories
|
||||
- `--cache-real-features` expects a **directory path** (e.g., `cache/real`), it will automatically create/load `real_features.pkl` inside that directory
|
||||
@@ -1,37 +0,0 @@
|
||||
"""
|
||||
FastVideo Frechet Video Distance (FVD) Benchmark Module.
|
||||
>>> from fastvideo.benchmarks.fvd import compute_fvd_with_config, FVDConfig
|
||||
>>> config = FVDConfig.fvd2048_16f() # Standard protocol
|
||||
>>> results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
|
||||
>>> print(f"FVD: {results['fvd']:.2f}")
|
||||
"""
|
||||
|
||||
from .fvd import (
|
||||
compute_fvd,
|
||||
compute_fvd_with_config,
|
||||
compute_frechet_distance,
|
||||
compute_statistics,
|
||||
FVDConfig,
|
||||
)
|
||||
from .feature_extractors import (BaseFeatureExtractor, I3DFeatureExtractor, load_extractor)
|
||||
from .video_utils import (
|
||||
load_video_auto,
|
||||
sample_clips_from_video,
|
||||
load_video_clips_streaming,
|
||||
ClipSamplingStrategy,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
'compute_fvd',
|
||||
'compute_fvd_with_config',
|
||||
'compute_frechet_distance',
|
||||
'compute_statistics',
|
||||
'FVDConfig',
|
||||
'BaseFeatureExtractor',
|
||||
'I3DFeatureExtractor',
|
||||
'load_extractor',
|
||||
'load_video_auto',
|
||||
'sample_clips_from_video',
|
||||
'load_video_clips_streaming',
|
||||
'ClipSamplingStrategy',
|
||||
]
|
||||
@@ -1,77 +0,0 @@
|
||||
import argparse
|
||||
import sys
|
||||
import traceback
|
||||
from .fvd import compute_fvd_with_config, FVDConfig
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description='Compute Fréchet Video Distance (FVD)')
|
||||
|
||||
# Required arguments
|
||||
parser.add_argument('--real-path', type=str, required=True, help='Path to real videos')
|
||||
parser.add_argument('--gen-path', type=str, required=True, help='Path to generated videos')
|
||||
|
||||
# Extractor selection
|
||||
parser.add_argument('--extractor',
|
||||
type=str,
|
||||
default='i3d',
|
||||
choices=['i3d', 'clip', 'videomae'],
|
||||
help='Feature extractor model to use (default: i3d)')
|
||||
|
||||
# Standard args
|
||||
parser.add_argument('--seed', type=int, default=None, help='Random seed for reproducibility')
|
||||
parser.add_argument('--protocol',
|
||||
type=str,
|
||||
default=None,
|
||||
choices=['fvd2048_16f', 'fvd2048_128f', 'quick_test'],
|
||||
help='Use standard protocol (overrides other settings)')
|
||||
parser.add_argument('--num-videos', type=int, default=2048, help='Number of videos to use')
|
||||
parser.add_argument('--num-frames', type=int, default=16, help='Number of frames per clip')
|
||||
parser.add_argument('--clip-strategy', type=str, default='beginning', help='Clip sampling strategy')
|
||||
parser.add_argument('--batch-size', type=int, default=32, help='Batch size for feature extraction')
|
||||
parser.add_argument('--device', type=str, default='cuda', help='Device to use (cuda or cpu)')
|
||||
parser.add_argument('--cache-real-features', type=str, default=None, help='Path to cache real video features')
|
||||
parser.add_argument('--quiet', action='store_true', help='Suppress progress output')
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
# Create config
|
||||
if args.protocol:
|
||||
protocol_map = {
|
||||
'fvd2048_16f': FVDConfig.fvd2048_16f,
|
||||
'fvd2048_128f': FVDConfig.fvd2048_128f,
|
||||
'quick_test': FVDConfig.quick_test,
|
||||
}
|
||||
config = protocol_map[args.protocol]()
|
||||
# Apply overrides
|
||||
config.device = args.device
|
||||
config.cache_real_features = args.cache_real_features
|
||||
config.extractor_model = args.extractor # Apply extractor arg
|
||||
else:
|
||||
config = FVDConfig(
|
||||
num_videos=args.num_videos,
|
||||
num_frames_per_clip=args.num_frames,
|
||||
extractor_model=args.extractor, # Apply extractor arg
|
||||
clip_strategy=args.clip_strategy,
|
||||
batch_size=args.batch_size,
|
||||
device=args.device,
|
||||
cache_real_features=args.cache_real_features,
|
||||
seed=args.seed)
|
||||
|
||||
try:
|
||||
_ = compute_fvd_with_config(
|
||||
args.real_path, # noqa: F841
|
||||
args.gen_path,
|
||||
config,
|
||||
verbose=not args.quiet)
|
||||
|
||||
return 0
|
||||
|
||||
except Exception as e:
|
||||
print(f"Error: {e}", file=sys.stderr)
|
||||
traceback.print_exc(file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
sys.exit(main())
|
||||
@@ -1,232 +0,0 @@
|
||||
"""
|
||||
Pluggable Feature Extractors for FVD Computation.
|
||||
Supports I3D (standard), CLIP, and VideoMAE via a common interface.
|
||||
"""
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
from abc import ABC, abstractmethod
|
||||
from huggingface_hub import hf_hub_download
|
||||
from tqdm import tqdm
|
||||
|
||||
try:
|
||||
from transformers import CLIPModel, CLIPProcessor, VideoMAEModel
|
||||
TRANSFORMERS_AVAILABLE = True
|
||||
except ImportError:
|
||||
TRANSFORMERS_AVAILABLE = False
|
||||
|
||||
|
||||
class BaseFeatureExtractor(ABC, nn.Module):
|
||||
"""Abstract base class for all video feature extractors."""
|
||||
|
||||
def __init__(self, device: str = 'cuda'):
|
||||
super().__init__()
|
||||
self.device = torch.device(device if torch.cuda.is_available() else 'cpu')
|
||||
|
||||
@property
|
||||
@abstractmethod
|
||||
def feature_dim(self) -> int:
|
||||
"""Dimension of the output feature vector."""
|
||||
pass
|
||||
|
||||
@abstractmethod
|
||||
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
"""
|
||||
Args:
|
||||
videos: [B, T, C, H, W] in [0, 255] range.
|
||||
Returns:
|
||||
Preprocessed tensor ready for the model.
|
||||
"""
|
||||
pass
|
||||
|
||||
@abstractmethod
|
||||
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
"""
|
||||
Extract features for a single batch.
|
||||
Args:
|
||||
videos: [B, T, C, H, W] (raw input)
|
||||
Returns:
|
||||
Features: [B, feature_dim]
|
||||
"""
|
||||
pass
|
||||
|
||||
@torch.no_grad()
|
||||
def extract_features(self, videos: torch.Tensor, batch_size: int = 32, verbose: bool = True) -> torch.Tensor:
|
||||
"""
|
||||
Extract features for a large tensor of videos by batching.
|
||||
"""
|
||||
N = len(videos)
|
||||
all_features = []
|
||||
|
||||
iterator = range(0, N, batch_size)
|
||||
if verbose:
|
||||
iterator = tqdm(iterator, desc=f"Extracting features ({self.__class__.__name__})")
|
||||
|
||||
for i in iterator:
|
||||
batch = videos[i:i + batch_size].to(self.device)
|
||||
features = self.extract_features_batch(batch)
|
||||
all_features.append(features.cpu())
|
||||
|
||||
return torch.cat(all_features, dim=0)
|
||||
|
||||
|
||||
# 1. I3D Extractor (The Standard FVD Metric)
|
||||
class I3DFeatureExtractor(BaseFeatureExtractor):
|
||||
REPO_ID = 'flateon/FVD-I3D-torchscript'
|
||||
MODEL_FILENAME = 'i3d_torchscript.pt'
|
||||
|
||||
def __init__(self, device: str = 'cuda', cache_dir: str | None = None):
|
||||
super().__init__(device)
|
||||
self.cache_dir = cache_dir
|
||||
self.model = self._load_model()
|
||||
self.model.eval()
|
||||
self.model.to(self.device)
|
||||
|
||||
@property
|
||||
def feature_dim(self) -> int:
|
||||
return 400
|
||||
|
||||
def _load_model(self) -> torch.nn.Module:
|
||||
try:
|
||||
model_path = hf_hub_download(repo_id=self.REPO_ID, filename=self.MODEL_FILENAME, cache_dir=self.cache_dir)
|
||||
return torch.jit.load(model_path, map_location=self.device)
|
||||
except Exception as e:
|
||||
raise RuntimeError(f"Failed to load I3D model: {e}") from e
|
||||
|
||||
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
"""Standard I3D preprocessing: Resize to 224, Norm to [-1, 1]."""
|
||||
B, T, C, H, W = videos.shape
|
||||
|
||||
if T < 10:
|
||||
raise ValueError(f"I3D requires at least 10 frames, got {T}")
|
||||
|
||||
# Normalize to [0, 1]
|
||||
if videos.max() > 1.0:
|
||||
videos = videos / 255.0
|
||||
|
||||
# Scale to [-1, 1]
|
||||
videos = videos * 2.0 - 1.0
|
||||
|
||||
# Resize to 224x224
|
||||
if H != 224 or W != 224:
|
||||
videos = videos.reshape(B * T, C, H, W)
|
||||
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
|
||||
videos = videos.reshape(B, T, C, 224, 224)
|
||||
|
||||
# [B, T, C, H, W] -> [B, C, T, H, W]
|
||||
return videos.permute(0, 2, 1, 3, 4).contiguous()
|
||||
|
||||
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
batch = self.preprocess(videos)
|
||||
# TorchScript I3D returns raw logits when return_features=True
|
||||
return self.model(batch, rescale=False, resize=False, return_features=True)
|
||||
|
||||
|
||||
# 2. CLIP Extractor (Semantic/Content Quality)
|
||||
class CLIPFeatureExtractor(BaseFeatureExtractor):
|
||||
|
||||
def __init__(self, device: str = 'cuda', model_name: str = "openai/clip-vit-base-patch32"):
|
||||
if not TRANSFORMERS_AVAILABLE:
|
||||
raise ImportError("Please install transformers: uv pip install transformers")
|
||||
super().__init__(device)
|
||||
self.processor = CLIPProcessor.from_pretrained(model_name)
|
||||
self.model = CLIPModel.from_pretrained(model_name).to(self.device)
|
||||
self.model.eval()
|
||||
self._feature_dim = self.model.config.projection_dim
|
||||
|
||||
@property
|
||||
def feature_dim(self) -> int:
|
||||
return self._feature_dim
|
||||
|
||||
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
# Ensure values are [0, 255]
|
||||
if videos.max() <= 1.0:
|
||||
videos = videos * 255.0
|
||||
|
||||
return videos.to(torch.uint8)
|
||||
|
||||
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
# Input: [B, T, C, H, W]
|
||||
B, T, C, H, W = videos.shape
|
||||
videos = self.preprocess(videos)
|
||||
|
||||
# Flatten B*T to treat frames as images
|
||||
images = videos.view(B * T, C, H, W)
|
||||
|
||||
# HF Processor
|
||||
inputs = self.processor(images=images, return_tensors="pt", padding=True)
|
||||
inputs = {k: v.to(self.device) for k, v in inputs.items()}
|
||||
|
||||
# Extract features [B*T, Dim]
|
||||
outputs = self.model.get_image_features(**inputs)
|
||||
|
||||
# Reshape [B, T, Dim] and Average Pooling over time
|
||||
outputs = outputs.view(B, T, -1)
|
||||
return outputs.mean(dim=1)
|
||||
|
||||
|
||||
# 3. VideoMAE Extractor (Structure/Motion Quality)
|
||||
class VideoMAEFeatureExtractor(BaseFeatureExtractor):
|
||||
|
||||
def __init__(self, device: str = 'cuda', model_name: str = "MCG-NJU/videomae-base"):
|
||||
if not TRANSFORMERS_AVAILABLE:
|
||||
raise ImportError("Please install transformers: uv pip install transformers")
|
||||
super().__init__(device)
|
||||
self.model = VideoMAEModel.from_pretrained(model_name).to(self.device)
|
||||
self.model.eval()
|
||||
|
||||
self.register_buffer('mean', torch.tensor([0.485, 0.456, 0.406], device=self.device).view(1, 1, 3, 1, 1))
|
||||
self.register_buffer('std', torch.tensor([0.229, 0.224, 0.225], device=self.device).view(1, 1, 3, 1, 1))
|
||||
|
||||
@property
|
||||
def feature_dim(self) -> int:
|
||||
return self.model.config.hidden_size
|
||||
|
||||
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
"""
|
||||
Efficient GPU-based preprocessing.
|
||||
Input: [B, T, C, H, W] in range [0, 255]
|
||||
"""
|
||||
B, T, C, H, W = videos.shape
|
||||
|
||||
# 1. Resize to 224x224
|
||||
if H != 224 or W != 224:
|
||||
videos = videos.view(B * T, C, H, W)
|
||||
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
|
||||
videos = videos.view(B, T, C, 224, 224)
|
||||
|
||||
# 2. Normalize to [0, 1]
|
||||
if videos.dtype != torch.float32:
|
||||
videos = videos.float()
|
||||
|
||||
if videos.max() > 1.0:
|
||||
videos = videos / 255.0
|
||||
|
||||
# 3. Apply ImageNet Mean/Std
|
||||
return (videos - self.mean) / self.std
|
||||
|
||||
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
# Input: [B, T, C, H, W]
|
||||
|
||||
# Fast GPU Preprocessing
|
||||
pixel_values = self.preprocess(videos)
|
||||
|
||||
# Forward pass
|
||||
outputs = self.model(pixel_values)
|
||||
|
||||
# Global Average Pooling of last hidden state [B, T_patches, 768] -> [B, 768]
|
||||
return outputs.last_hidden_state.mean(dim=1)
|
||||
|
||||
|
||||
# Factory
|
||||
def load_extractor(name: str, device: str = 'cuda') -> BaseFeatureExtractor:
|
||||
name = name.lower()
|
||||
if name == 'i3d':
|
||||
return I3DFeatureExtractor(device)
|
||||
elif name == 'clip':
|
||||
return CLIPFeatureExtractor(device)
|
||||
elif name == 'videomae':
|
||||
return VideoMAEFeatureExtractor(device)
|
||||
else:
|
||||
raise ValueError(f"Unknown extractor: {name}. Options: i3d, clip, videomae")
|
||||
@@ -1,384 +0,0 @@
|
||||
import numpy as np
|
||||
import scipy.linalg
|
||||
import torch
|
||||
from pathlib import Path
|
||||
from collections.abc import Iterator
|
||||
import pickle
|
||||
from dataclasses import dataclass, field
|
||||
from .feature_extractors import BaseFeatureExtractor, load_extractor
|
||||
from .video_utils import ClipSamplingStrategy, load_video_clips_streaming
|
||||
|
||||
|
||||
def compute_statistics(features: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
|
||||
"""Compute mean and covariance."""
|
||||
mu = np.mean(features, axis=0)
|
||||
sigma = np.cov(features, rowvar=False)
|
||||
return mu, sigma
|
||||
|
||||
|
||||
def compute_frechet_distance(mu1: np.ndarray,
|
||||
sigma1: np.ndarray,
|
||||
mu2: np.ndarray,
|
||||
sigma2: np.ndarray,
|
||||
eps: float = 1e-6) -> float:
|
||||
"""
|
||||
Compute Fréchet distance between two Gaussians.
|
||||
"""
|
||||
sigma1 = sigma1 + eps * np.eye(sigma1.shape[0])
|
||||
sigma2 = sigma2 + eps * np.eye(sigma2.shape[0])
|
||||
|
||||
diff = mu1 - mu2
|
||||
mean_distance = np.sum(diff**2)
|
||||
|
||||
trace_sum = np.trace(sigma1 + sigma2)
|
||||
|
||||
covmean = scipy.linalg.sqrtm(sigma1 @ sigma2)
|
||||
|
||||
if np.iscomplexobj(covmean):
|
||||
if not np.allclose(np.diagonal(covmean).imag, 0, atol=1e-3):
|
||||
print(f"Warning: Imaginary component: {np.max(np.abs(covmean.imag))}")
|
||||
covmean = covmean.real
|
||||
|
||||
trace_product = np.trace(covmean)
|
||||
|
||||
fvd = mean_distance + trace_sum - 2 * trace_product
|
||||
|
||||
return float(fvd)
|
||||
|
||||
|
||||
@dataclass
|
||||
class FVDConfig:
|
||||
# default configuration for FVD computation:
|
||||
|
||||
# Video selection
|
||||
num_videos: int = 2048
|
||||
|
||||
# Feature Extractor Selection
|
||||
extractor_model: str = 'i3d' # Options: 'i3d', 'clip', 'videomae'
|
||||
|
||||
# Clip sampling
|
||||
num_frames_per_clip: int = 16
|
||||
num_clips_per_video: int = 1
|
||||
clip_strategy: str | ClipSamplingStrategy = 'beginning'
|
||||
|
||||
# Temporal subsampling
|
||||
frame_stride: int = 1 # 1=no subsampling, 2=every 2nd, 8=every 8th
|
||||
temporal_stride: int = 1 # For sliding window clips
|
||||
|
||||
# Data processing
|
||||
video_extensions: list[str] = field(default_factory=lambda: ['.mp4', '.avi', '.mov', '.mkv'])
|
||||
support_frame_dirs: bool = True
|
||||
|
||||
# Computation
|
||||
batch_size: int = 32
|
||||
device: str = 'cuda'
|
||||
|
||||
use_streaming: bool = True
|
||||
resize_before_extraction: bool = True
|
||||
|
||||
# Caching
|
||||
cache_real_features: str | None = None
|
||||
i3d_model_path: str | None = None
|
||||
|
||||
# Reproducibility
|
||||
seed: int | None = None
|
||||
|
||||
@classmethod
|
||||
def fvd2048_16f(cls) -> 'FVDConfig':
|
||||
"""Standard FVD protocol: 2048 videos, 16 frames, beginning clip."""
|
||||
return cls(num_videos=2048, num_frames_per_clip=16, clip_strategy='beginning', use_streaming=True)
|
||||
|
||||
@classmethod
|
||||
def fvd2048_128f(cls) -> 'FVDConfig':
|
||||
"""Long video protocol: 2048 videos, 128 frames."""
|
||||
return cls(num_videos=2048, num_frames_per_clip=128, clip_strategy='beginning', use_streaming=True)
|
||||
|
||||
@classmethod
|
||||
def quick_test(cls) -> 'FVDConfig':
|
||||
"""Quick test config: 100 videos, 16 frames."""
|
||||
return cls(num_videos=100, num_frames_per_clip=16, clip_strategy='beginning')
|
||||
|
||||
def to_dict(self) -> dict:
|
||||
"""Export config to dict for logging"""
|
||||
d = self.__dict__.copy()
|
||||
d['clip_strategy'] = str(self.clip_strategy)
|
||||
return d
|
||||
|
||||
def __str__(self) -> str:
|
||||
"""Human-readable protocol name"""
|
||||
desc = f"FVD_{self.extractor_model.upper()}_{self.num_videos}_{self.num_frames_per_clip}f"
|
||||
if self.frame_stride > 1:
|
||||
desc += f"_subsample{self.frame_stride}"
|
||||
if self.num_clips_per_video > 1:
|
||||
desc += f"_{self.num_clips_per_video}clips"
|
||||
if self.clip_strategy != 'beginning':
|
||||
desc += f"_{self.clip_strategy}"
|
||||
return desc
|
||||
|
||||
|
||||
def extract_features_streaming(video_generator: Iterator[torch.Tensor],
|
||||
extractor: BaseFeatureExtractor,
|
||||
batch_size: int = 32,
|
||||
max_clips: int | None = None,
|
||||
verbose: bool = True) -> np.ndarray:
|
||||
"""
|
||||
Extract features from a video clip generator using streaming.
|
||||
"""
|
||||
all_features = []
|
||||
batch = []
|
||||
|
||||
if verbose:
|
||||
print(f"Extracting features with batch_size={batch_size}...")
|
||||
|
||||
with torch.no_grad():
|
||||
for clip_count, clip in enumerate(video_generator):
|
||||
batch.append(clip)
|
||||
|
||||
# Process batch when full
|
||||
if len(batch) == batch_size:
|
||||
batch_tensor = torch.stack(batch).to(extractor.device)
|
||||
features = extractor.extract_features_batch(batch_tensor)
|
||||
|
||||
all_features.append(features.detach().cpu().numpy())
|
||||
batch = []
|
||||
|
||||
if verbose and clip_count % (batch_size * 10) == 0:
|
||||
print(f"Processed {clip_count} clips...")
|
||||
|
||||
if max_clips is not None and clip_count >= max_clips:
|
||||
break
|
||||
|
||||
# Process remaining clips
|
||||
if len(batch) > 0:
|
||||
batch_tensor = torch.stack(batch).to(extractor.device)
|
||||
features = extractor.extract_features_batch(batch_tensor)
|
||||
all_features.append(features.detach().cpu().numpy())
|
||||
|
||||
if len(all_features) == 0:
|
||||
raise RuntimeError("No features extracted - check video loading")
|
||||
|
||||
features = np.concatenate(all_features, axis=0)
|
||||
|
||||
if verbose:
|
||||
print(f"Extracted {len(features)} feature vectors")
|
||||
|
||||
return features
|
||||
|
||||
|
||||
def load_or_compute_features(videos: str | Path | torch.Tensor,
|
||||
extractor: BaseFeatureExtractor,
|
||||
config: FVDConfig,
|
||||
cache_path: str | None = None,
|
||||
cache_name: str = "real_features") -> np.ndarray:
|
||||
"""Load features from cache or compute (with streaming support)"""
|
||||
|
||||
if cache_path is not None:
|
||||
script_dir = Path(__file__).parent
|
||||
cache_dir = script_dir / cache_path
|
||||
cache_file = cache_dir / f"{config.extractor_model}_{cache_name}.pkl"
|
||||
|
||||
if cache_file.exists():
|
||||
print(f"Loading cached features from {cache_file}")
|
||||
with open(cache_file, 'rb') as f:
|
||||
features = pickle.load(f)
|
||||
|
||||
# Validate and limit based on config
|
||||
max_features = config.num_videos * config.num_clips_per_video
|
||||
|
||||
if len(features) < max_features:
|
||||
print(f"WARNING: Cache has {len(features)} features but need {max_features}")
|
||||
print("Cached features insufficient - will recompute...")
|
||||
elif len(features) > max_features:
|
||||
print(f"Using {max_features} features from cache (truncated from {len(features)})")
|
||||
features = features[:max_features]
|
||||
return features
|
||||
else:
|
||||
print(f"Using all {len(features)} cached features")
|
||||
return features
|
||||
|
||||
print("Computing features from scratch...")
|
||||
|
||||
if isinstance(videos, (str | Path)):
|
||||
target_size = (224, 224) if config.resize_before_extraction else None
|
||||
|
||||
video_generator = load_video_clips_streaming(videos,
|
||||
num_frames=config.num_frames_per_clip,
|
||||
max_videos=config.num_videos,
|
||||
clip_strategy=config.clip_strategy,
|
||||
frame_stride=config.frame_stride,
|
||||
num_clips_per_video=config.num_clips_per_video,
|
||||
video_extensions=config.video_extensions,
|
||||
support_frame_dirs=config.support_frame_dirs,
|
||||
target_size=target_size,
|
||||
verbose=True)
|
||||
|
||||
max_clips = config.num_videos * config.num_clips_per_video
|
||||
features = extract_features_streaming(video_generator,
|
||||
extractor,
|
||||
batch_size=config.batch_size,
|
||||
max_clips=max_clips,
|
||||
verbose=True)
|
||||
else:
|
||||
print(f"Extracting features from {len(videos)} video tensors...")
|
||||
features = extractor.extract_features(videos, batch_size=config.batch_size, verbose=True)
|
||||
features = features.numpy()
|
||||
|
||||
# Validate feature count
|
||||
expected_count = config.num_videos * config.num_clips_per_video
|
||||
if len(features) < expected_count:
|
||||
raise ValueError(f"ERROR: Only extracted {len(features)} features, but need {expected_count}!\n"
|
||||
f"Found fewer videos than expected. Check your video directory.")
|
||||
elif len(features) > expected_count:
|
||||
print(f"Truncating {len(features)} features to {expected_count}")
|
||||
features = features[:expected_count]
|
||||
|
||||
# Cache features if requested
|
||||
if cache_path is not None:
|
||||
script_dir = Path(__file__).parent
|
||||
cache_dir = script_dir / cache_path
|
||||
cache_dir.mkdir(parents=True, exist_ok=True)
|
||||
cache_file = cache_dir / f"{config.extractor_model}_{cache_name}.pkl"
|
||||
print(f"Caching features to {cache_file}")
|
||||
with open(cache_file, 'wb') as f:
|
||||
pickle.dump(features, f)
|
||||
|
||||
return features
|
||||
|
||||
|
||||
def compute_fvd_with_config(real_videos: str | Path | torch.Tensor,
|
||||
gen_videos: str | Path | torch.Tensor,
|
||||
config: FVDConfig,
|
||||
verbose: bool = True) -> dict:
|
||||
"""
|
||||
Compute FVD using a standardized configuration.
|
||||
|
||||
This is the recommended way to compute FVD for reproducibility.
|
||||
|
||||
Args:
|
||||
real_videos: Path or tensors
|
||||
gen_videos: Path or tensors
|
||||
config: FVDConfig specifying protocol
|
||||
verbose: Print progress
|
||||
|
||||
Returns:
|
||||
results: Dictionary with:
|
||||
- 'fvd': FVD score (float)
|
||||
- 'protocol': Protocol name (str)
|
||||
- 'model': Feature extractor model name (str)
|
||||
- 'config': Configuration dict
|
||||
|
||||
Example:
|
||||
>>> config = FVDConfig.fvd2048_16f()
|
||||
>>> results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
|
||||
>>> print(f"FVD: {results['fvd']:.2f}")
|
||||
"""
|
||||
|
||||
# Seed for reproducibility
|
||||
if config.seed is not None:
|
||||
import random as _rnd
|
||||
_rnd.seed(config.seed)
|
||||
np.random.seed(config.seed)
|
||||
torch.manual_seed(config.seed)
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.manual_seed_all(config.seed)
|
||||
|
||||
if verbose:
|
||||
print("=" * 70)
|
||||
print(f"Computing FVD with protocol: {config}")
|
||||
print(f"Model: {config.extractor_model.upper()}")
|
||||
print("=" * 70)
|
||||
print("\nConfiguration:")
|
||||
for key, value in config.to_dict().items():
|
||||
print(f" {key}: {value}")
|
||||
print()
|
||||
|
||||
# Initialize Extractor using Factory
|
||||
if verbose:
|
||||
print(f"\nInitializing {config.extractor_model.upper()} model on {config.device}...")
|
||||
|
||||
extractor = load_extractor(config.extractor_model, device=config.device)
|
||||
|
||||
# Extract features
|
||||
if verbose:
|
||||
print(f"\n{'='*70}")
|
||||
print("Extracting REAL video features...")
|
||||
print(f"{'='*70}")
|
||||
|
||||
real_features = load_or_compute_features(videos=real_videos,
|
||||
extractor=extractor,
|
||||
config=config,
|
||||
cache_path=config.cache_real_features,
|
||||
cache_name="real_features")
|
||||
|
||||
if verbose:
|
||||
print(f"\n{'='*70}")
|
||||
print("Extracting GENERATED video features...")
|
||||
print(f"{'='*70}")
|
||||
|
||||
gen_features = load_or_compute_features(videos=gen_videos,
|
||||
extractor=extractor,
|
||||
config=config,
|
||||
cache_path=None,
|
||||
cache_name="gen_features")
|
||||
|
||||
if verbose:
|
||||
print(f"\nReal videos/clips: {len(real_features)}")
|
||||
print(f"Generated videos/clips: {len(gen_features)}")
|
||||
print(f"\n{'='*70}")
|
||||
print("Computing statistics...")
|
||||
print(f"{'='*70}")
|
||||
|
||||
mu_real, sigma_real = compute_statistics(real_features)
|
||||
mu_gen, sigma_gen = compute_statistics(gen_features)
|
||||
|
||||
if verbose:
|
||||
print(f"\n{'='*70}")
|
||||
print("Computing Fréchet distance...")
|
||||
print(f"{'='*70}")
|
||||
|
||||
fvd = compute_frechet_distance(mu_real, sigma_real, mu_gen, sigma_gen)
|
||||
|
||||
if verbose:
|
||||
print(f"\n{'='*70}")
|
||||
print(f"FVD Score ({config.extractor_model.upper()}): {fvd:.4f}")
|
||||
print(f"Protocol: {config}")
|
||||
print(f"{'='*70}\n")
|
||||
|
||||
results = {
|
||||
'fvd': fvd,
|
||||
'protocol': str(config),
|
||||
'model': config.extractor_model,
|
||||
'config': config.to_dict(),
|
||||
}
|
||||
|
||||
return results
|
||||
|
||||
|
||||
def compute_fvd(real_videos: str | Path | torch.Tensor,
|
||||
gen_videos: str | Path | torch.Tensor,
|
||||
num_frames: int = 16,
|
||||
batch_size: int = 32,
|
||||
device: str = 'cuda',
|
||||
num_videos: int | None = 2048,
|
||||
cache_real_features: str | None = None,
|
||||
i3d_model_path: str | None = None,
|
||||
seed: int | None = None,
|
||||
verbose: bool = True) -> float:
|
||||
"""
|
||||
Backward compatibility wrapper for computing FVD (defaults to I3D).
|
||||
"""
|
||||
num_videos = num_videos if num_videos is not None else 2048
|
||||
|
||||
config = FVDConfig(
|
||||
num_videos=num_videos,
|
||||
num_frames_per_clip=num_frames,
|
||||
extractor_model='i3d', # Default to I3D
|
||||
batch_size=batch_size,
|
||||
device=device,
|
||||
cache_real_features=cache_real_features,
|
||||
i3d_model_path=i3d_model_path,
|
||||
seed=seed,
|
||||
)
|
||||
|
||||
result = compute_fvd_with_config(real_videos, gen_videos, config, verbose)
|
||||
return result['fvd']
|
||||
@@ -1,124 +0,0 @@
|
||||
"""I3D Feature Extractor for FVD Computation"""
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
from pathlib import Path
|
||||
from huggingface_hub import hf_hub_download
|
||||
from tqdm import tqdm
|
||||
from contextlib import suppress
|
||||
|
||||
|
||||
class I3DFeatureExtractor(nn.Module):
|
||||
"""
|
||||
I3D feature extractor for FVD computation.
|
||||
Extracts 400-dimensional features from videos using I3D model
|
||||
trained on Kinetics-400.
|
||||
"""
|
||||
|
||||
REPO_ID = 'flateon/FVD-I3D-torchscript'
|
||||
MODEL_FILENAME = 'i3d_torchscript.pt'
|
||||
|
||||
def __init__(self, device: str = 'cuda', cache_dir: str | Path | None = None):
|
||||
super().__init__()
|
||||
|
||||
self.device_str = device
|
||||
if device == 'cuda' and not torch.cuda.is_available():
|
||||
print("Warning: CUDA requested but not available – falling back to CPU")
|
||||
self.device = torch.device('cpu')
|
||||
else:
|
||||
self.device = torch.device(device)
|
||||
|
||||
self.cache_dir: str | None
|
||||
if cache_dir is not None:
|
||||
self.cache_dir = str(Path(cache_dir).resolve())
|
||||
else:
|
||||
self.cache_dir = None # Use HF default cache
|
||||
|
||||
self.model = self._load_model()
|
||||
self.model.eval()
|
||||
|
||||
with suppress(Exception):
|
||||
self.model.to(self.device)
|
||||
|
||||
def _load_model(self) -> torch.nn.Module:
|
||||
"""Download and load I3D TorchScript model from Hugging Face Hub."""
|
||||
print(f"Loading I3D model from Hugging Face Hub ({self.REPO_ID})...")
|
||||
|
||||
try:
|
||||
# Download model from Hugging Face Hub
|
||||
model_path = hf_hub_download(repo_id=self.REPO_ID, filename=self.MODEL_FILENAME, cache_dir=self.cache_dir)
|
||||
|
||||
# Load directly to chosen device
|
||||
model = torch.jit.load(model_path, map_location=self.device)
|
||||
print("I3D model loaded successfully")
|
||||
return model
|
||||
|
||||
except Exception as e:
|
||||
raise RuntimeError(f"Failed to load I3D model from Hugging Face Hub. Error: {e}\n"
|
||||
f"Ensure you have internet connection and huggingface_hub installed:\n"
|
||||
f"uv pip install huggingface_hub") from e
|
||||
|
||||
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
|
||||
"""
|
||||
Preprocess videos for I3D.
|
||||
|
||||
Args:
|
||||
videos: [B, T, C, H, W], values in [0, 255]
|
||||
|
||||
Returns:
|
||||
Preprocessed videos [B, C, T, 224, 224] (normalized and resized)
|
||||
"""
|
||||
B, T, C, H, W = videos.shape
|
||||
|
||||
if T < 10:
|
||||
raise ValueError(f"I3D requires at least 10 frames, got {T}")
|
||||
|
||||
# Normalize to [0, 1] if needed
|
||||
if videos.max() > 1.0:
|
||||
videos = videos / 255.0
|
||||
|
||||
# Resize to 224x224 if needed
|
||||
if H != 224 or W != 224:
|
||||
videos = videos.reshape(B * T, C, H, W)
|
||||
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
|
||||
videos = videos.reshape(B, T, C, 224, 224)
|
||||
|
||||
# Convert to [B, C, T, H, W] format
|
||||
videos = videos.permute(0, 2, 1, 3, 4).contiguous()
|
||||
|
||||
return videos
|
||||
|
||||
@torch.no_grad()
|
||||
def extract_features(self, videos: torch.Tensor, batch_size: int = 32, verbose: bool = True) -> torch.Tensor:
|
||||
"""
|
||||
Extract I3D features
|
||||
|
||||
Args:
|
||||
videos: [N, T, C, H, W], values in [0, 255]
|
||||
batch_size: Batch size for processing
|
||||
verbose: Show progress bar
|
||||
|
||||
Returns:
|
||||
Features [N, 400]
|
||||
"""
|
||||
N = len(videos)
|
||||
all_features = []
|
||||
|
||||
iterator = range(0, N, batch_size)
|
||||
if verbose:
|
||||
iterator = tqdm(iterator, desc="Extracting I3D features")
|
||||
|
||||
for i in iterator:
|
||||
batch = videos[i:i + batch_size].to(self.device)
|
||||
batch = self.preprocess(batch) # Now returns [B, C, T, H, W]
|
||||
|
||||
# Use the HF model without rescale/resize (we handle it in preprocess)
|
||||
features = self.model(batch, rescale=False, resize=False, return_features=True)
|
||||
|
||||
all_features.append(features.cpu())
|
||||
|
||||
return torch.cat(all_features, dim=0)
|
||||
|
||||
def __call__(self, videos: torch.Tensor, batch_size: int = 32) -> torch.Tensor:
|
||||
return self.extract_features(videos, batch_size=batch_size)
|
||||
@@ -1,51 +0,0 @@
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
root_dir = Path(__file__).parent.parent.parent
|
||||
sys.path.insert(0, str(root_dir))
|
||||
|
||||
from benchmarks.fvd.fvd import FVDConfig, compute_fvd_with_config # noqa: E402
|
||||
|
||||
|
||||
def main() -> None:
|
||||
script_dir = Path(__file__).parent.resolve()
|
||||
|
||||
# Define directories
|
||||
real_dir = "benchmarks/data/real_videos"
|
||||
gen_dir = "benchmarks/data/generated_videos"
|
||||
|
||||
# Compare all 3 models
|
||||
models_to_test = ['i3d', 'clip', 'videomae']
|
||||
|
||||
print(f"\n{'='*60}")
|
||||
print("STARTING COMPARISON BENCHMARK")
|
||||
print(f"{'='*60}")
|
||||
|
||||
for model_name in models_to_test:
|
||||
print(f"\n>>> Running evaluation with {model_name.upper()}...")
|
||||
|
||||
try:
|
||||
cfg = FVDConfig(
|
||||
num_videos=650,
|
||||
num_frames_per_clip=16,
|
||||
extractor_model=model_name,
|
||||
clip_strategy='beginning',
|
||||
device='cuda',
|
||||
seed=42,
|
||||
# Use separate cache folders for each model to avoid conflicts
|
||||
cache_real_features=str(script_dir / f'fvd-cache/{model_name}'),
|
||||
)
|
||||
|
||||
results = compute_fvd_with_config(real_dir, gen_dir, cfg, verbose=False)
|
||||
print(f"FVD: {results['fvd']}\nModel: {results['model']}")
|
||||
|
||||
except Exception as e:
|
||||
print(f"{model_name.upper()} Failed: {e}")
|
||||
|
||||
print(f"\n{'='*60}")
|
||||
print("BENCHMARK COMPLETE")
|
||||
print(f"{'='*60}")
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -1,89 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
import shutil
|
||||
import random
|
||||
from fvd import compute_fvd_with_config, FVDConfig
|
||||
|
||||
script_path = Path(__file__).resolve()
|
||||
fastvideo_root = script_path.parent.parent.parent
|
||||
sys.path.insert(0, str(fastvideo_root))
|
||||
|
||||
|
||||
def split_videos(video_dir: Path, n_per_subset: int = 128, seed: int = 42):
|
||||
subset_a = video_dir.parent / 'bair_full_subset_A'
|
||||
subset_b = video_dir.parent / 'bair_full_subset_B'
|
||||
|
||||
if subset_a.exists():
|
||||
shutil.rmtree(subset_a)
|
||||
if subset_b.exists():
|
||||
shutil.rmtree(subset_b)
|
||||
|
||||
subset_a.mkdir(parents=True)
|
||||
subset_b.mkdir(parents=True)
|
||||
|
||||
videos = sorted(video_dir.glob('*.mp4'))
|
||||
|
||||
random.seed(seed)
|
||||
shuffled = list(videos)
|
||||
random.shuffle(shuffled)
|
||||
|
||||
needed = n_per_subset * 2
|
||||
if len(shuffled) > needed:
|
||||
shuffled = shuffled[:needed]
|
||||
|
||||
mid = len(shuffled) // 2
|
||||
|
||||
print(f"\nSplitting {len(shuffled)} BAIR FULL videos:")
|
||||
print(f" Subset A: {mid} videos")
|
||||
print(f" Subset B: {len(shuffled) - mid} videos")
|
||||
|
||||
for v in shuffled[:mid]:
|
||||
shutil.copy2(v, subset_a / v.name)
|
||||
|
||||
for v in shuffled[mid:]:
|
||||
shutil.copy2(v, subset_b / v.name)
|
||||
|
||||
return subset_a, subset_b, mid
|
||||
|
||||
|
||||
def validate_fvd(subset_a: Path, subset_b: Path, num_videos: int):
|
||||
config = FVDConfig(num_videos=num_videos,
|
||||
num_frames_per_clip=16,
|
||||
clip_strategy='beginning',
|
||||
batch_size=8,
|
||||
device='cuda',
|
||||
seed=42)
|
||||
|
||||
print("\n" + "=" * 70)
|
||||
print("TEST 1: Identity Test")
|
||||
print("=" * 70)
|
||||
|
||||
result1 = compute_fvd_with_config(real_videos=str(subset_a), gen_videos=str(subset_a), config=config, verbose=False)
|
||||
fvd_identity = result1['fvd']
|
||||
print(f"\nIdentity FVD: {fvd_identity:.2f}")
|
||||
|
||||
print("\n" + "=" * 70)
|
||||
print("TEST 2: Real vs Real")
|
||||
print("=" * 70)
|
||||
|
||||
result2 = compute_fvd_with_config(real_videos=str(subset_a), gen_videos=str(subset_b), config=config, verbose=False)
|
||||
fvd_real = result2['fvd']
|
||||
print(f"\nReal vs Real FVD: {fvd_real:.2f}")
|
||||
|
||||
print("\n" + "=" * 70)
|
||||
print("RESULTS")
|
||||
print("=" * 70)
|
||||
print(f"Identity: {fvd_identity:.2f}")
|
||||
print(f"Real vs Real: {fvd_real:.2f}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
bair_dir = Path('benchmarks/data/bair_full_videos')
|
||||
|
||||
subset_a, subset_b, count = split_videos(bair_dir, n_per_subset=128, seed=42)
|
||||
validate_fvd(subset_a, subset_b, count)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -1,460 +0,0 @@
|
||||
import torch
|
||||
import cv2
|
||||
import numpy as np
|
||||
from pathlib import Path
|
||||
from collections.abc import Iterator
|
||||
from tqdm import tqdm
|
||||
from enum import Enum
|
||||
|
||||
|
||||
class ClipSamplingStrategy(Enum):
|
||||
"""Clip sampling strategies for FVD evaluation."""
|
||||
BEGINNING = 'beginning' # Take first N frames (most common)
|
||||
RANDOM = 'random' # Random N consecutive frames
|
||||
UNIFORM = 'uniform' # Uniformly spaced frames across video
|
||||
MIDDLE = 'middle' # Middle N frames
|
||||
SLIDING = 'sliding' # Multiple sliding windows
|
||||
ALL = 'all' # All possible clips
|
||||
|
||||
|
||||
def _load_video_cv2(video_path: str | Path,
|
||||
num_frames: int | None = 16,
|
||||
sample_strategy: str = 'uniform') -> torch.Tensor:
|
||||
"""
|
||||
Load video from video file using OpenCV.
|
||||
|
||||
Args:
|
||||
video_path: Path to video file (MP4, AVI, MOV, MKV)
|
||||
num_frames: Number of frames to extract
|
||||
sample_strategy: 'uniform' or 'random'
|
||||
|
||||
Returns:
|
||||
video: [T, C, H, W]
|
||||
"""
|
||||
video_path = str(video_path)
|
||||
cap = cv2.VideoCapture(video_path)
|
||||
|
||||
if not cap.isOpened():
|
||||
raise RuntimeError(f"Cannot open video: {video_path}")
|
||||
|
||||
frames = []
|
||||
total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
|
||||
|
||||
if num_frames is None:
|
||||
# Read all available frames
|
||||
while True:
|
||||
ret, frame = cap.read()
|
||||
if not ret:
|
||||
break
|
||||
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
|
||||
frames.append(frame)
|
||||
|
||||
cap.release()
|
||||
if len(frames) == 0:
|
||||
raise RuntimeError(f"Video has 0 frames: {video_path}")
|
||||
|
||||
frames = np.stack(frames) # [T, H, W, C]
|
||||
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
|
||||
return frames
|
||||
|
||||
if total_frames == 0:
|
||||
raise RuntimeError(f"Video has 0 frames: {video_path}")
|
||||
|
||||
# Determine frame indices for sampling
|
||||
if total_frames < num_frames:
|
||||
frame_indices = list(range(total_frames)) + [total_frames - 1] * (num_frames - total_frames)
|
||||
elif sample_strategy == 'uniform':
|
||||
frame_indices = np.linspace(0, total_frames - 1, num_frames, dtype=int).tolist()
|
||||
elif sample_strategy == 'random':
|
||||
frame_indices = sorted(np.random.choice(total_frames, num_frames, replace=False))
|
||||
else:
|
||||
raise ValueError(f"Unknown sample_strategy: {sample_strategy}")
|
||||
|
||||
# Extract frames
|
||||
for idx in frame_indices:
|
||||
cap.set(cv2.CAP_PROP_POS_FRAMES, idx)
|
||||
ret, frame = cap.read()
|
||||
|
||||
if not ret:
|
||||
if len(frames) > 0:
|
||||
frames.append(frames[-1].copy())
|
||||
else:
|
||||
h = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
|
||||
w = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
|
||||
frames.append(np.zeros((h, w, 3), dtype=np.uint8))
|
||||
continue
|
||||
|
||||
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
|
||||
frames.append(frame)
|
||||
|
||||
cap.release()
|
||||
|
||||
frames = np.stack(frames) # [T, H, W, C]
|
||||
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
|
||||
|
||||
return frames
|
||||
|
||||
|
||||
def _load_video_from_frames(frame_dir: str | Path,
|
||||
num_frames: int | None = 16,
|
||||
sample_strategy: str = 'uniform',
|
||||
frame_extensions: list[str] | None = None) -> torch.Tensor:
|
||||
"""
|
||||
Load video from directory of frame images.
|
||||
|
||||
Args:
|
||||
frame_dir: Directory containing frames
|
||||
num_frames: Number of frames to sample
|
||||
sample_strategy: 'uniform' or 'random'
|
||||
frame_extensions: Image file extensions to look for
|
||||
|
||||
Returns:
|
||||
video: [T, C, H, W]
|
||||
"""
|
||||
if frame_extensions is None:
|
||||
frame_extensions = ['.jpg', '.png', '.jpeg', '.bmp']
|
||||
|
||||
frame_dir = Path(frame_dir)
|
||||
|
||||
if not frame_dir.exists():
|
||||
raise FileNotFoundError(f"Frame directory not found: {frame_dir}")
|
||||
|
||||
# Find all frames
|
||||
frame_files: list[Path] = []
|
||||
for ext in frame_extensions:
|
||||
frame_files.extend(frame_dir.glob(f"*{ext}"))
|
||||
|
||||
if len(frame_files) == 0:
|
||||
raise ValueError(f"No frames found in {frame_dir} with extensions {frame_extensions}")
|
||||
|
||||
frame_files = sorted(frame_files, key=lambda x: x.name)
|
||||
total_frames = len(frame_files)
|
||||
|
||||
# Determine frame indices
|
||||
if num_frames is None:
|
||||
frame_indices = list(range(total_frames))
|
||||
else:
|
||||
if total_frames < num_frames:
|
||||
frame_indices = list(range(total_frames)) + [total_frames - 1] * (num_frames - total_frames)
|
||||
elif sample_strategy == 'uniform':
|
||||
frame_indices = np.linspace(0, total_frames - 1, num_frames, dtype=int).tolist()
|
||||
elif sample_strategy == 'random':
|
||||
frame_indices = sorted(np.random.choice(total_frames, num_frames, replace=False))
|
||||
else:
|
||||
raise ValueError(f"Unknown sample_strategy: {sample_strategy}")
|
||||
|
||||
# Load frames
|
||||
frames = []
|
||||
for idx in frame_indices:
|
||||
frame_path = frame_files[idx]
|
||||
frame = cv2.imread(str(frame_path))
|
||||
|
||||
if frame is None:
|
||||
raise RuntimeError(f"Failed to load frame: {frame_path}")
|
||||
|
||||
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
|
||||
frames.append(frame)
|
||||
|
||||
# Stack and convert to tensor
|
||||
frames = np.stack(frames) # [T, H, W, C]
|
||||
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
|
||||
|
||||
return frames
|
||||
|
||||
|
||||
def _detect_video_format(path: str | Path) -> str:
|
||||
"""
|
||||
Detect if path is a video file or frame directory.
|
||||
|
||||
Returns:
|
||||
'video_file', 'frame_directory', or 'unknown'
|
||||
"""
|
||||
path = Path(path)
|
||||
|
||||
if path.is_file():
|
||||
return 'video_file'
|
||||
elif path.is_dir():
|
||||
# Check if contains image files
|
||||
image_extensions = ['.jpg', '.jpeg', '.png', '.bmp']
|
||||
for ext in image_extensions:
|
||||
if list(path.glob(f"*{ext}")):
|
||||
return 'frame_directory'
|
||||
return 'unknown'
|
||||
else:
|
||||
raise ValueError(f"Path does not exist: {path}")
|
||||
|
||||
|
||||
def load_video_auto(video_path: str | Path,
|
||||
num_frames: int | None = 16,
|
||||
sample_strategy: str = 'uniform') -> torch.Tensor:
|
||||
"""
|
||||
Automatically detect format and load video.
|
||||
|
||||
Supports:
|
||||
- Video files (MP4, AVI, MOV, MKV)
|
||||
- Frame directories (JPG, PNG)
|
||||
|
||||
Args:
|
||||
video_path: Path to video file or frame directory
|
||||
num_frames: Number of frames to extract
|
||||
sample_strategy: 'uniform' or 'random'
|
||||
|
||||
Returns:
|
||||
video: [T, C, H, W]
|
||||
"""
|
||||
format_type = _detect_video_format(video_path)
|
||||
|
||||
if format_type == 'video_file':
|
||||
return _load_video_cv2(video_path, num_frames, sample_strategy)
|
||||
elif format_type == 'frame_directory':
|
||||
return _load_video_from_frames(video_path, num_frames, sample_strategy)
|
||||
else:
|
||||
raise ValueError(f"Unknown video format at {video_path}")
|
||||
|
||||
|
||||
def sample_clips_from_video(video: torch.Tensor,
|
||||
num_frames_per_clip: int = 16,
|
||||
num_clips: int = 1,
|
||||
strategy: str | ClipSamplingStrategy = ClipSamplingStrategy.BEGINNING,
|
||||
frame_stride: int = 1,
|
||||
temporal_stride: int = 1) -> list[torch.Tensor]:
|
||||
"""
|
||||
Sample clips from a video with various strategies.
|
||||
|
||||
Args:
|
||||
video: [T, C, H, W] full video
|
||||
num_frames_per_clip: Frames per clip
|
||||
num_clips: Number of clips to extract
|
||||
strategy: ClipSamplingStrategy or string ('beginning', 'random', etc.)
|
||||
frame_stride: Skip frames (FPS control: 1=all, 2=every 2nd, 8=every 8th)
|
||||
temporal_stride: Stride between clips for sliding window
|
||||
|
||||
Returns:
|
||||
List of clips, each [num_frames_per_clip, C, H, W]
|
||||
|
||||
Examples:
|
||||
>>> # Beginning clip (most common for FVD)
|
||||
>>> clips = sample_clips_from_video(video, 16, strategy='beginning')
|
||||
|
||||
>>> # Multiple random clips
|
||||
>>> clips = sample_clips_from_video(video, 16, num_clips=4, strategy='random')
|
||||
|
||||
>>> # Subsample FPS by 2x (every 2nd frame)
|
||||
>>> clips = sample_clips_from_video(video, 16, frame_stride=2)
|
||||
|
||||
>>> # Sliding window with overlap
|
||||
>>> clips = sample_clips_from_video(video, 16, strategy='sliding', temporal_stride=8)
|
||||
"""
|
||||
# Convert string to enum if needed
|
||||
if isinstance(strategy, str):
|
||||
strategy = ClipSamplingStrategy(strategy)
|
||||
|
||||
T, C, H, W = video.shape
|
||||
|
||||
# Apply frame stride (FPS subsampling)
|
||||
if frame_stride > 1:
|
||||
video = video[::frame_stride]
|
||||
T = len(video)
|
||||
|
||||
effective_clip_length = num_frames_per_clip
|
||||
|
||||
# Handle videos shorter than clip length
|
||||
if effective_clip_length > T:
|
||||
pad_length = effective_clip_length - T
|
||||
last_frame = video[-1:].repeat(pad_length, 1, 1, 1)
|
||||
video = torch.cat([video, last_frame], dim=0)
|
||||
T = len(video)
|
||||
|
||||
clips = []
|
||||
|
||||
if strategy == ClipSamplingStrategy.BEGINNING:
|
||||
# Take first clip (most common for FVD evaluation)
|
||||
clip = video[:effective_clip_length]
|
||||
clips.append(clip)
|
||||
|
||||
elif strategy == ClipSamplingStrategy.MIDDLE:
|
||||
# Take middle clip
|
||||
start = (T - effective_clip_length) // 2
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
|
||||
elif strategy == ClipSamplingStrategy.RANDOM:
|
||||
# Sample N random clips
|
||||
for _ in range(num_clips):
|
||||
start = 0 if effective_clip_length == T else np.random.randint(0, T - effective_clip_length + 1)
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
|
||||
elif strategy == ClipSamplingStrategy.UNIFORM:
|
||||
# Uniformly spaced clips
|
||||
if num_clips == 1:
|
||||
# Single clip from middle
|
||||
start = (T - effective_clip_length) // 2
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
else:
|
||||
# Multiple uniformly spaced clips
|
||||
step = (T - effective_clip_length) / (num_clips - 1) if num_clips > 1 else 0
|
||||
for i in range(num_clips):
|
||||
start = int(i * step)
|
||||
start = min(start, T - effective_clip_length)
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
|
||||
elif strategy == ClipSamplingStrategy.SLIDING:
|
||||
# Sliding window with stride
|
||||
for start in range(0, T - effective_clip_length + 1, temporal_stride):
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
if len(clips) >= num_clips:
|
||||
break
|
||||
|
||||
elif strategy == ClipSamplingStrategy.ALL:
|
||||
# All possible clips (overlapping)
|
||||
for start in range(T - effective_clip_length + 1):
|
||||
clip = video[start:start + effective_clip_length]
|
||||
clips.append(clip)
|
||||
|
||||
else:
|
||||
raise ValueError(f"Unknown strategy: {strategy}")
|
||||
|
||||
return clips
|
||||
|
||||
|
||||
def load_video_clips_streaming(directory: str | Path,
|
||||
num_frames: int = 16,
|
||||
max_videos: int | None = None,
|
||||
clip_strategy: str
|
||||
| ClipSamplingStrategy = 'beginning',
|
||||
frame_stride: int = 1,
|
||||
num_clips_per_video: int = 1,
|
||||
video_extensions: list[str] | None = None,
|
||||
support_frame_dirs: bool = True,
|
||||
target_size: tuple[int, int] | None = (224, 224),
|
||||
verbose: bool = True) -> Iterator[torch.Tensor]:
|
||||
"""
|
||||
This generator yields clips one-by-one instead of loading all videos into RAM.
|
||||
Perfect for large datasets where memory is limited.
|
||||
|
||||
Args:
|
||||
directory: Path to directory with videos
|
||||
num_frames: Frames per clip
|
||||
max_videos: Max videos to load
|
||||
clip_strategy: 'beginning', 'random', 'uniform', etc.
|
||||
frame_stride: Frame skip (1=all, 2=every 2nd, 8=every 8th)
|
||||
num_clips_per_video: Number of clips per video
|
||||
video_extensions: Video file extensions
|
||||
support_frame_dirs: Also load frame directories
|
||||
target_size: Resize clips to (H, W). If None, keep original size.
|
||||
verbose: Show progress
|
||||
|
||||
Yields:
|
||||
clip: [T, C, H, W] individual clips
|
||||
|
||||
Example:
|
||||
>>> for clip in load_video_clips_streaming('data/videos/', num_frames=16):
|
||||
>>> features = model.extract_features(clip.unsqueeze(0))
|
||||
>>> # Process one clip at a time - low memory usage!
|
||||
"""
|
||||
if video_extensions is None:
|
||||
video_extensions = ['.mp4', '.avi', '.mov', '.mkv']
|
||||
|
||||
directory = Path(directory)
|
||||
|
||||
if not directory.exists():
|
||||
raise FileNotFoundError(f"Directory not found: {directory}")
|
||||
|
||||
# Find video paths
|
||||
video_paths: list[Path] = []
|
||||
|
||||
# Find video files
|
||||
for ext in video_extensions:
|
||||
video_paths.extend(directory.glob(f"**/*{ext}"))
|
||||
|
||||
# Find frame directories if enabled
|
||||
if support_frame_dirs:
|
||||
for subdir in directory.iterdir():
|
||||
if subdir.is_dir():
|
||||
# Check if it contains frames
|
||||
image_extensions = ['.jpg', '.jpeg', '.png', '.bmp']
|
||||
for ext in image_extensions:
|
||||
if list(subdir.glob(f"*{ext}")):
|
||||
video_paths.append(subdir)
|
||||
break
|
||||
|
||||
if len(video_paths) == 0:
|
||||
raise ValueError(f"No videos found in {directory}")
|
||||
|
||||
video_paths = sorted(video_paths)
|
||||
|
||||
if max_videos is not None:
|
||||
video_paths = video_paths[:max_videos]
|
||||
|
||||
if verbose:
|
||||
print(f"Found {len(video_paths)} videos in {directory}")
|
||||
if num_clips_per_video > 1:
|
||||
print(f"Extracting {num_clips_per_video} clips per video...")
|
||||
if frame_stride > 1:
|
||||
print(f"Subsampling frames with stride {frame_stride}...")
|
||||
if target_size:
|
||||
print(f"Resizing clips to {target_size}...")
|
||||
|
||||
# Track statistics
|
||||
failed_count = 0
|
||||
total_clips = 0
|
||||
|
||||
iterator = tqdm(video_paths, desc="Loading videos") if verbose else video_paths
|
||||
|
||||
for video_path in iterator:
|
||||
try:
|
||||
# Load full video
|
||||
video = load_video_auto(video_path, num_frames=None, sample_strategy='uniform')
|
||||
|
||||
# Sample clips from video
|
||||
clips = sample_clips_from_video(video,
|
||||
num_frames_per_clip=num_frames,
|
||||
num_clips=num_clips_per_video,
|
||||
strategy=clip_strategy,
|
||||
frame_stride=frame_stride)
|
||||
|
||||
if target_size is not None:
|
||||
resized_clips = []
|
||||
for clip in clips:
|
||||
T, C, H, W = clip.shape
|
||||
if target_size != (H, W):
|
||||
# Resize to target size
|
||||
clip = clip.contiguous() # Fix non-contiguous tensors first
|
||||
clip_flat = clip.view(T * C, H, W).unsqueeze(0) # [1, T*C, H, W]
|
||||
clip_resized = torch.nn.functional.interpolate(clip_flat,
|
||||
size=target_size,
|
||||
mode='bilinear',
|
||||
align_corners=False)
|
||||
clip = clip_resized.squeeze(0).view(T, C, target_size[0],
|
||||
target_size[1]) # Back to [T, C, H, W]
|
||||
resized_clips.append(clip)
|
||||
clips = resized_clips
|
||||
|
||||
# Yield clips one by one
|
||||
for clip in clips:
|
||||
yield clip
|
||||
total_clips += 1
|
||||
|
||||
# Free memory
|
||||
del video, clips
|
||||
|
||||
except Exception as e:
|
||||
failed_count += 1
|
||||
if verbose:
|
||||
print(f"\nWarning: Failed to load {video_path}: {e}")
|
||||
continue
|
||||
|
||||
# Validate
|
||||
if total_clips == 0:
|
||||
raise RuntimeError(f"Failed to load any videos from {directory}")
|
||||
|
||||
failure_rate = failed_count / len(video_paths)
|
||||
if failure_rate > 0.1: # More than 10% failed
|
||||
print(f"\nWARNING: {failure_rate:.1%} of videos failed to load ({failed_count}/{len(video_paths)})")
|
||||
|
||||
if verbose:
|
||||
print(f"\nSuccessfully loaded {total_clips} clips from {len(video_paths) - failed_count} videos")
|
||||
@@ -1,7 +0,0 @@
|
||||
#!/bin/bash
|
||||
|
||||
# 1. Install missing dependency
|
||||
uv pip install -q opencv-python-headless transformers huggingface_hub
|
||||
|
||||
# 2. Run FVD script
|
||||
python benchmarks/fvd/run_fvd.py
|
||||
@@ -1,4 +0,0 @@
|
||||
#!/bin/bash
|
||||
|
||||
# 1. Install missing dependency
|
||||
uv pip install -q opencv-python-headless
|
||||
@@ -0,0 +1,71 @@
|
||||
# FastVideo Next-Gen Runtime — Two-Page Summary
|
||||
|
||||
**Companion to** `design.md` (v19, 2026-06-12) · **Status:** draft for discussion · **Ask:** read this, then dive into the sections you own.
|
||||
|
||||
---
|
||||
|
||||
## The problem
|
||||
|
||||
FastVideo's pipeline abstraction has been outgrown by its own model zoo. Four facts, all on `main` today:
|
||||
|
||||
- **The denoise/sampling loop exists in four copies** — inference stages (`pipelines/stages/denoising.py`, a 1,381-line file), `train/` distillation methods, legacy `training/` monoliths, and the just-landed RL work. The fourth copy documents the cause in its own docstring: `DiffusionSampler` (PR #1450) *"intentionally does not call FastVideo's full inference pipelines"* because the only consumable units are family-bound pipeline classes. Every new post-training method must pick between a wrong dependency and a private loop.
|
||||
- **There is no serving runtime.** One request at a time, no queue, no cross-request batching. Dreamverse — our shipping product — hand-rolls its own GPU pool, queue, warmup, and streaming relay at a cost of **one B200 per user session**.
|
||||
- **Cosmos3 outgrows both the stage abstraction and the alternatives.** Its AR text reasoner and multimodal diffusion denoiser are *the same resident weights* driven by two loop types within one request — our full port (incl. action modality, on `feat/cosmos3-reasoning`) runs it only by bypassing stages with one monolithic block. Multi-engine DAG stacks (vllm-omni, sglang-omni) compose separable stages with disjoint weights; none can express a mixture-of-transformers model. This is the forcing function.
|
||||
- **RL has arrived and pays the tax already** — likelihood-free DiffusionNFT for Wan (#1450): in-process rollouts with zero serving-grade optimizations (no CFG, dense attention, full 25-step ODE), a vendored sampler, and a parallel validation path built *because* inference pipelines aren't consumable as a library.
|
||||
|
||||
## The design
|
||||
|
||||
```
|
||||
Request plane OpenAI-compatible server (videos/images/audio/chat) · AsyncEngine
|
||||
OmniRequest ─► admission ─► queue ─► OmniOutput (typed modality parts)
|
||||
Pipeline plane PipelineSpec: declarative graph per family
|
||||
nodes: Stage | LoopStage (DenoiseLoop, ARDecodeLoop) · typed Artifacts
|
||||
policies: CFG, ExpertRouting, AttnMetadata, Precision, FlowShift
|
||||
Execution plane StepScheduler (multiplexes denoise + AR steps) · worker pools (TP/SP + CFG-parallel)
|
||||
CacheManager (paged text-KV ▪ chunked causal-video KV ▪ feature caches) · connectors
|
||||
```
|
||||
|
||||
Default deployment is exactly today's: one SPMD pool, co-located nodes, synchronous call. Serving is additive configuration, not a different code path. Five load-bearing decisions:
|
||||
|
||||
1. **Loop inversion.** Loops become `LoopStage` nodes exposing `init / step / finalize`; the runtime owns iteration, families own step bodies (with a custom-step escape hatch — the runtime never dictates step factoring). This is the enabler for step-level scheduling, streaming, MoT interleaving, and one shared loop across inference, distillation, and RL. Stated honestly: **no surveyed system does this at scheduler granularity** — the risk is retired by Phase-1 bit-identical parity gates and a measured falsifier, not borrowed validation.
|
||||
2. **Cost-model scheduling.** Denoise steps and AR tokens are incommensurable (bidirectional attention is O(L²) per step with zero KV amortization; steps differ ~1000×) — the budget currency is **predicted GPU-time** from a per-(model, phase, shape) cost model, calibrated by the profiler and published to Dynamo's router/Planner as the same artifact.
|
||||
3. **One substrate for inference, training, and RL** — models, loaders, configs, schedulers, parallel state, loop step bodies — under a strict `engine never imports train` rule. Trainers keep their internals; their embedded sampling paths migrate onto the shared loops.
|
||||
4. **Consistency is a declared, measured contract**, not a hope: **C0** corrected (profiles differ, TIS/MIS fixes it) / **C1** kernel-pinned (RL default; drift gated in CI) / **C2** bitwise (batch-invariant kernels + Behavior Record, for goldens and MoE parity). One repo ≠ automatic parity — the ladder is what makes the single-runtime bet honest.
|
||||
5. **Extensions, never monkeypatching**: read-only observers (ParityAligner, ActivationTrace, Profiler, NaNWatch) and compute-altering interceptors at declared points — **cache-dit** is the first interceptor. **Dynamo is the fleet layer** (first-class partner, not a dependency we rebuild): registration/health/cost contract in-engine, seven concrete upstream asks (affinity key spaces, cost interface, media streaming, role-graph disagg, RL weight plane, KVBM generalization, sessions) — each with a fallback.
|
||||
|
||||
## Why this is the moat
|
||||
|
||||
Unlike LLMs — where inference optimization is post-hoc on frozen weights — **a usable video model is itself a post-training artifact**. Every inference capability we ship is a *(recipe, runtime)* pair: step distillation ↔ few-step samplers; self-forcing ↔ causal KV streaming; QAT-NVFP4 ↔ FP4 kernels; VSA ↔ sparse attention backend; RL ↔ samplers + capture. The training loop *embeds* the inference loop, so whoever owns both sides of the pair owns the optimization frontier. The industry's RL pain proves the converse: verl-omni re-implements Wan2.2 inside vLLM-Omni and corrects the numerics afterward; miles' headline features are all mismatch patches for two runtimes with different kernels. We answer with one model definition, one kernel set, one measured ladder — at FastVideo's 1–30B FSDP2 scale, where the bet is viable.
|
||||
|
||||
## What's pulling on it
|
||||
|
||||
| Customer | Pull | Proof point |
|
||||
|---|---|---|
|
||||
| **Cosmos3 / omni** | MoT loops, packed sequences, reasoner KV, world-model rollout | 150-test parity suite on `feat/cosmos3-reasoning` |
|
||||
| **Dreamverse** | Engine-client replaces hand-rolled pool; capacity = duty cycle + cost-model admission + distillation | today 1 B200/session; Phase-3 gate: ≥2 sessions/GPU on a recorded duty-cycle trace, p95 within SLO |
|
||||
| **RL (landed)** | #1450 migrates onto shared loops (Phase 1), engine-client rollouts (Phase 2+); GRPO-class next | the vendored-sampler docstring; C1-by-construction discipline |
|
||||
| **ComfyUI funnel** | embed (nodes) → **compile** (workflow→PipelineSpec, accelerated cloud) → productize (Studio) | tier-1 ~20-node static sublanguage maps onto PipelineSpec |
|
||||
|
||||
## Migration — seven phases, each independently shippable
|
||||
|
||||
| Phase | Ships | Gate |
|
||||
|---|---|---|
|
||||
| **−1** | Merge cosmos3 chain; seed SSIM for uncovered families | baselines exist |
|
||||
| **0** | Typed omni I/O; config freeze (`compat.py` shrinks monotonically to zero) | all SSIM suites unchanged |
|
||||
| **1** | **Loop inversion** + policies + extension core (cache-dit, ParityAligner); RL migrates off its vendored sampler | old vs new loop **bit-identical**; per-method grad-norm refs (#1396) extended |
|
||||
| **2** | AsyncEngine + StepScheduler; LTX-2 linear graph; Dynamo worker (stock); colocated weight sync | ≤2% batch-1 latency regression; Dreamverse single-session parity; RL engine-client parity |
|
||||
| **3** | PipelineSpec graphs, role pools, declarative parallelism, ComfyUI compiler MVP, general WeightSyncPlan | ≥2 Dreamverse sessions/GPU on recorded duty-cycle trace |
|
||||
| **4** | Cosmos3 native; AR continuous batching + paged KV (arriving *with* their workload, per N5); RL hardening (C1/C2, Behavior Record) | Cosmos3 parity suite on new runtime; drift ≈ 0 on a Wan RL run |
|
||||
| **5** | Deletion: legacy `training/` retires, then `ComposedPipelineBase`, legacy loop, `forward_context.py`, `compat.py`, `RayDistributedExecutor` | the deletion diff — **4 loop copies → 1** |
|
||||
|
||||
Enforcement the last freeze lacked (it was broken 19×): `compat.py` frozen from Phase 0; CI path gates + CODEOWNERS once Phase 1 lands; new families land on new abstractions from Phase-1 completion.
|
||||
|
||||
## What we are deliberately not doing
|
||||
|
||||
Datacenter orchestration (Dynamo's job) · trainer internals (frozen `training/`; `train/` is a consumer) · replacing the bit-exact porting methodology · migrating 20+ families at once · **standalone LLM-serving excellence** — AR machinery arrives only at the sophistication omni workloads pull (N5).
|
||||
|
||||
## Decisions we need from this review
|
||||
|
||||
1. **sglang `multimodal_gen` relationship** — upstream, friendly fork, or shared core (decide by Phase 2; drift is a strategic cost either way).
|
||||
2. **Dynamo asks** — green-light proposing the Phase-2 asks (A2 cost interface, A3 media streaming) to the team first, with A5 (RL weight plane) queued behind them?
|
||||
3. **Phase −1 start** — merge the cosmos3 chain and seed SSIM baselines now; it blocks everything else.
|
||||
+830
@@ -0,0 +1,830 @@
|
||||
# FastVideo v3 — A Model-Native Runtime for the (Recipe, Runtime) Era
|
||||
|
||||
**Status:** unconstrained north-star. This document assumes we are free to build a brand-new architecture with no
|
||||
backward-compatibility, no migration tax, and no obligation to the current code. It exists to define the *ceiling*:
|
||||
the system FastVideo should be if nothing held it back. Migration is a separate, later question — deliberately out of
|
||||
scope here.
|
||||
|
||||
**Lineage:** this is the synthesis of `design.md` (the strategic thesis, product pull, and hard-won serving realism)
|
||||
and `designv2.md` (the model-native center and typed contracts), with the open tensions of both resolved rather than
|
||||
hedged.
|
||||
|
||||
---
|
||||
|
||||
## Table of contents
|
||||
|
||||
1. [The thesis](#1-the-thesis)
|
||||
2. [The two signature ideas](#2-the-two-signature-ideas)
|
||||
3. [Planes and their dependency order](#3-planes-and-their-dependency-order)
|
||||
4. [Model Plane — the center](#4-model-plane--the-center)
|
||||
5. [The loop contract — driven loops](#5-the-loop-contract--driven-loops)
|
||||
6. [Runtime and scheduler — one currency, one WorkUnit](#6-runtime-and-scheduler)
|
||||
7. [Memory, cache, transport, compile](#7-memory-cache-transport-compile)
|
||||
8. [Parallelism as a model contract](#8-parallelism)
|
||||
9. [Correctness — parity as a typed gate](#9-correctness)
|
||||
10. [Training and RL on the same loops](#10-training-and-rl)
|
||||
11. [Extensions — observers and interceptors](#11-extensions)
|
||||
12. [Request, session, artifact, stream](#12-request-session-artifact-stream)
|
||||
13. [Programs and workflows](#13-programs-and-workflows)
|
||||
14. [Deployment and fleet](#14-deployment-and-fleet)
|
||||
15. [Worked examples](#15-worked-examples)
|
||||
16. [What this unlocks](#16-what-this-unlocks)
|
||||
17. [Honest unknowns and falsifiers](#17-honest-unknowns-and-falsifiers)
|
||||
18. [Package layout](#18-package-layout)
|
||||
19. [Reference synthesis](#19-reference-synthesis)
|
||||
|
||||
---
|
||||
|
||||
## 1. The thesis
|
||||
|
||||
Three facts about video generation, taken together, dictate the architecture.
|
||||
|
||||
**A deployable video model is a post-training artifact.** Unlike an LLM — where inference optimizes frozen weights
|
||||
post-hoc — a *usable* video model is *created* by training: step distillation is mandatory for usable latency, low
|
||||
precision needs QAT, and causal/world models are made by distillation plus self-forcing. Every inference capability is
|
||||
therefore a **(recipe, runtime) pair**: the weights and the loop that produced-and-assumes them are inseparable. A
|
||||
"4-step NVFP4 FastWan" is not weights plus a flag; it is a distillation recipe, a sampler, a precision path, and a
|
||||
parity contract that are one object.
|
||||
|
||||
**Video systems are loop systems.** Denoise timesteps, AR decode, chunked world-model rollout, VAE tiles, encoder
|
||||
chunks, audio tokens, reward batches, optimizer steps, media chunks — the work is iteration, not a single `forward`. A
|
||||
runtime that reduces everything to `forward()` cannot schedule, batch, cancel, stream, reserve memory for, or capture
|
||||
the behavior of the thing that actually runs.
|
||||
|
||||
**Omni models share weights across loop types within one request.** Cosmos3's text reasoner and multimodal denoiser
|
||||
are the *same resident weights*, driven by an AR loop and then a diffusion loop in one request. This cannot be a
|
||||
DAG of separate engines (that doubles 30B+ of weights and severs the shared KV/denoise state); it must be one resident
|
||||
model instance running many loop types. This is achievable — vllm-omni's `bagel_single_stage`/`lance` already run one
|
||||
resident MoT instance doing AR `generate_text` and diffusion `generate_image` on co-resident experts in a single
|
||||
request — *but* they bury that interleaving inside one opaque `DIFFUSION` stage their scheduler never sees inside,
|
||||
request-scheduled with `max_num_running_reqs` forced to 1. The hard part, and the differentiation, is not *expressing*
|
||||
the shared-weight loops — it is making them **runtime-visible, step-scheduled, batchable, and cost-priced.**
|
||||
|
||||
The architecture that falls out:
|
||||
|
||||
> **The atomic unit is the (recipe, runtime) pair, owned by a typed `ModelCard`.** Everything — serving, training, RL,
|
||||
> products, deployment — is a *view* over that card. The **runtime owns loop *lifecycle*** (admission, scheduling,
|
||||
> batching, caching, cancellation, streaming, behavior capture); the **model owns loop *semantics*** (typed state
|
||||
> transitions and kernel execution). One resident model instance runs many loops; one scheduler schedules the
|
||||
> *steps* of all of them in a single currency; one parity contract binds the train-forward to the serve-forward so
|
||||
> the recipe and the runtime never silently drift apart.
|
||||
|
||||
The single invariant, stated once:
|
||||
|
||||
```text
|
||||
Model cards own components, loops, recipes, and parity.
|
||||
Programs compose loops into tasks.
|
||||
The scheduler executes the steps of loops as WorkUnits under one budget.
|
||||
Caches are correct by key, not by hope.
|
||||
Training records behavior on the same loops it serves.
|
||||
Deployment places and routes; it does not define semantics.
|
||||
Products stream artifacts; they do not reach into the model.
|
||||
```
|
||||
|
||||
Everything below is the elaboration of that invariant.
|
||||
|
||||
---
|
||||
|
||||
## 2. The two signature ideas
|
||||
|
||||
Two ideas do most of the work and are what an unconstrained design can reach that an incremental one cannot.
|
||||
|
||||
### 2.1 The (recipe, runtime) pair is a first-class, versioned, typed object
|
||||
|
||||
A model in v3 is not a checkpoint. It is a `ModelCard` that owns, as one versioned unit:
|
||||
|
||||
- the **components** (weights, loaders, layouts),
|
||||
- the **loops** it can run (the runtime semantics),
|
||||
- the **recipe** that produced the weights (distillation/QAT/RL config, teacher, data contract, the sampler the
|
||||
recipe assumes), and
|
||||
- the **parity contract** asserting that the train-forward and the serve-forward agree to a declared level.
|
||||
|
||||
You cannot ship the weights without the loop they assume, and you cannot change the loop without re-proving parity.
|
||||
This makes design.md's "(recipe, runtime) pair" *literal*: the deployable artifact carries its own provenance and its
|
||||
own correctness obligation. It is the thing that turns "we do training and inference in one repo" from an org chart
|
||||
into a guarantee — and it is essentially un-retrofittable, which is exactly why it belongs in a clean design.
|
||||
|
||||
### 2.2 Driven loops: the model owns control flow, the runtime owns execution
|
||||
|
||||
Loop inversion, done right, is not a heavy `plan_step`/`run_step` contract and not a hidden `for t in timesteps`. It
|
||||
is a **driven loop**: the model describes the *next step it needs*, the runtime *decides when and with whom that step
|
||||
runs*, and the model folds the result back into its own state and decides what to do next. The model keeps its control
|
||||
flow (so content-adaptive decisions — cache-dit skips, EOS, VSA tile selection — are natural); the runtime keeps the
|
||||
`await` (so admission, batching, cancellation, streaming, and behavior capture are universal). Per-request state lives
|
||||
in the loop's own typed `LoopState`, never in module globals, so interleaving requests through one model instance
|
||||
cannot smear state — the failure mode that makes naive loop-inversion dangerous is *structurally* excluded.
|
||||
|
||||
These two ideas are developed in §4–§5. The rest of the system is their consequence.
|
||||
|
||||
---
|
||||
|
||||
## 3. Planes and their dependency order
|
||||
|
||||
```text
|
||||
Products: Python · CLI · OpenAI API · ComfyUI · Dreamverse · RTC · Trainer
|
||||
│ (thin: validate intent, make requests/sessions, subscribe)
|
||||
Request / Session / Artifact / Stream
|
||||
│ (typed runtime objects, cancellation, streaming)
|
||||
Program Plane ← typed loop programs + compiled workflows
|
||||
│
|
||||
┌──────────────── Model Plane (CENTER) ────────────────┐
|
||||
│ ModelCard: components · loops · recipe · parity │
|
||||
│ capabilities · caches · parallelism · precision │
|
||||
└───────────────────────┬──────────────────────────────┘
|
||||
│
|
||||
┌─────────────────────────────┼─────────────────────────────┐
|
||||
│ Runtime / Scheduler │ Training / RL │ (same loops, different capture)
|
||||
│ WorkUnits · GPU-time budget │ rollout · reward · weight-sync│
|
||||
└─────────────────────────────┼─────────────────────────────┘
|
||||
│
|
||||
Memory · Cache · Transport · Compile (typed CacheKey, per-class pools, CuMem sleep/wake)
|
||||
│
|
||||
Parallelism (named axes → DeviceMesh, validated, part of the cache key)
|
||||
│
|
||||
Deployment / Fleet (DeploymentCard → Dynamo; never the core)
|
||||
```
|
||||
|
||||
Dependency rules (enforced at the package boundary, §18):
|
||||
|
||||
- Products do not define model semantics. Workflows do not define model semantics. Deployment does not define model
|
||||
semantics. Training does not redefine model semantics. **All of them reference the Model Plane.**
|
||||
- The runtime *executes* model loops but does not *own* their math. Training *captures* behavior on serving loops but
|
||||
does not *fork* them. Cross-cutting concerns — extensions (§11), parallelism (§8), and parity (§9) — are contracts
|
||||
declared on the card, not features bolted onto the runtime.
|
||||
|
||||
---
|
||||
|
||||
## 4. Model Plane — the center
|
||||
|
||||
### 4.1 ModelCard
|
||||
|
||||
```python
|
||||
class ModelCard:
|
||||
model_id: str # "fastwan-1.3b-nvfp4-4step"
|
||||
family: str # "wan"
|
||||
components: dict[str, ComponentSpec]
|
||||
loops: dict[str, LoopSpec]
|
||||
capabilities: CapabilityMatrix # text_to_video, image_to_video, reasoning_text, vae_decode, ...
|
||||
recipe: RecipeSpec # ← what produced these weights (signature idea §2.1)
|
||||
parity: ParitySpec # ← train-forward ≡ serve-forward, to a declared level (§9)
|
||||
caches: dict[str, CacheContract]
|
||||
parallelism: ParallelismContract
|
||||
precision: PrecisionContract
|
||||
checkpoint: CheckpointManifest # explicit components, layouts, key maps — no name-detector guessing
|
||||
```
|
||||
|
||||
The card is both a **declarative contract** (strict enough to validate before any GPU touches it) and a **runtime
|
||||
factory** (it knows how to instantiate components, bind loops, and resolve caches). It is hub-interchange compatible
|
||||
with diffusers' `modular_model_index.json` / `ComponentSpec` so models published either way load both ways.
|
||||
|
||||
`CheckpointManifest` replaces today's implicit `model_index.json` + name-detector resolution with explicit declared
|
||||
components and `required_for` / `optional_for` task sets (the Cosmos3 lazy-sound-VAE problem becomes a declaration, not
|
||||
an `if env_var` inside `forward`).
|
||||
|
||||
### 4.2 RecipeSpec — the provenance half of the pair
|
||||
|
||||
```python
|
||||
class RecipeSpec:
|
||||
method: str # "dmd2" | "self_forcing" | "attn_qat_nvfp4" | "diffusion_nft" | "base"
|
||||
parents: list[str] # teacher / base model_ids this was distilled or RL'd from
|
||||
data_contract: DataRef # what the recipe trained on (for governance and reproduction)
|
||||
assumes_loop: str # the loop_id this recipe's weights require at serve time
|
||||
assumes_precision: str # the precision the QAT recipe baked in
|
||||
consistency_required: str # the minimum parity level this recipe's outputs are valid under (§9)
|
||||
```
|
||||
|
||||
`assumes_loop` and `assumes_precision` are the teeth: a 4-step distilled model whose `assumes_loop = "ddim_4step"`
|
||||
cannot be served under a 50-step sampler without a typed mismatch error. The recipe and the runtime are bound.
|
||||
|
||||
### 4.3 ComponentSpec and LoopSpec
|
||||
|
||||
```python
|
||||
class ComponentSpec:
|
||||
component_id: str
|
||||
kind: str # dit | vae | text_encoder | reasoner_tower | reward_head | ...
|
||||
load_id: str
|
||||
config_schema: type
|
||||
io_schema: tuple[type, type]
|
||||
precision_policy: PrecisionPolicy
|
||||
placement_policy: PlacementPolicy
|
||||
parallel_constraints: ParallelConstraint
|
||||
parity_tests: list[ParityTestSpec]
|
||||
|
||||
class LoopSpec:
|
||||
loop_id: str # diffusion_denoise | ar_decode | chunk_rollout | vae_tile | ...
|
||||
state_schema: type # the typed LoopState (no dicts)
|
||||
step_schema: type # the typed WorkPlan a step emits
|
||||
result_schema: type # the typed StepResult a step returns
|
||||
behavior_schema: type | None # what to capture for RL (None if not training-relevant)
|
||||
step_cost_model: CostModel # predicted GPU-time per step at (shape, precision, policy) — §6
|
||||
valid_parallel_plans: list[ParallelPlanPattern]
|
||||
graph_capture: GraphCapturePolicy
|
||||
cache_policy: CachePolicy
|
||||
```
|
||||
|
||||
A `ModelInstance` is a resident, loaded card: component instances, model state, caches, compiled graphs, and a
|
||||
parallel plan. **A request may run several of the card's loops against one `ModelInstance`.** That single sentence is
|
||||
the difference between this design and a stage-only design, and it is what makes omni native:
|
||||
|
||||
```text
|
||||
one Cosmos3 ModelInstance, one request:
|
||||
ar_decode(reasoner) → pack → diffusion_denoise(vision[+action][+sound]) → vae_tile_decode → audio_decode
|
||||
└────────────── same resident weights, shared packed state, scheduled as steps ──────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. The loop contract — driven loops
|
||||
|
||||
### 5.1 The contract
|
||||
|
||||
A loop is a **serializable state machine** the runtime drives:
|
||||
|
||||
```python
|
||||
class Loop(Protocol):
|
||||
def init(self, req: Request, model: ModelState, ctx: LoopContext) -> LoopState: ...
|
||||
def next(self, state: LoopState) -> WorkPlan | Done: ... # describe the next step; NO GPU kernels here
|
||||
def advance(self, state: LoopState, result: StepResult) -> LoopState: ... # fold result in; decide what's next
|
||||
def finalize(self, state: LoopState) -> LoopResult: ...
|
||||
```
|
||||
|
||||
The runtime's driver — the only place iteration lives:
|
||||
|
||||
```python
|
||||
state = loop.init(req, model_state, ctx)
|
||||
while True:
|
||||
plan = loop.next(state) # typed WorkPlan: resources, cache reads/writes, shape, sinks, cancel-scope
|
||||
if isinstance(plan, Done):
|
||||
break
|
||||
result = await ctx.execute(plan) # ← THE INVERSION POINT: runtime admits, batches, places, runs, returns
|
||||
state = loop.advance(state, result) # content-adaptive: next() can branch on everything in state, incl. result
|
||||
for chunk in plan.emits:
|
||||
ctx.emit(chunk) # streaming falls out
|
||||
return loop.finalize(state)
|
||||
```
|
||||
|
||||
Why this is the right contract, point by point against the failure modes:
|
||||
|
||||
- **Content-adaptive steps are natural.** `next()` reads `state`, and `advance()` has already folded in the last
|
||||
`StepResult` — so cache-dit's skip decision (a residual comparison from the prior step), AR's EOS, and VSA's
|
||||
content-dependent tile selection are ordinary control flow in the model. This is the tension `designv2.md`'s
|
||||
"pure plan_step that pre-declares shape" could not resolve; here it dissolves, because planning the *next* step is
|
||||
allowed to depend on the *previous* result. `next()` is still kernel-free (it *describes* work; it does not run it),
|
||||
which is all the scheduler needs.
|
||||
- **Cross-request state safety is structural.** All per-request state is in the loop's typed `LoopState`. There are no
|
||||
module-level residual/KV globals (the bug that silently corrupts cache-dit and TeaCache forks under concurrency).
|
||||
Interleaving requests through one `ModelInstance` cannot smear state because there is no shared mutable state to
|
||||
smear. This is the safety property naive loop inversion lacks, made impossible-to-get-wrong by construction.
|
||||
- **The runtime owns everything it needs and nothing it doesn't.** `execute(plan)` is the single seam for admission,
|
||||
memory reservation, cross-request batching, placement, graph dispatch, cancellation, and behavior capture. The model
|
||||
never sees the scheduler; the scheduler never sees the model's math.
|
||||
- **Serializable, therefore migratable and resumable.** `LoopState` is typed and serializable, so a half-finished
|
||||
1000-step job is a resume point: preempt by stopping the driver, migrate by shipping `LoopState` to another worker,
|
||||
recover from a crash by replaying from the last serialized state. (A coroutine that keeps state in a suspended Python
|
||||
frame — the tempting sugar — cannot do this; the explicit state machine is the price of resumability, and it is
|
||||
worth paying.)
|
||||
|
||||
Custom step bodies are first-class, not an escape hatch with an asterisk: a family whose math is genuinely braided
|
||||
(Cosmos's EDM coefficients consumed inside the CFG branch with an x0-space combine; LTX-2's 1–4 runtime-decided
|
||||
guidance passes) writes `next()`/`advance()` by hand using samplers and CFG utilities as a *library*. The runtime
|
||||
requires only the four methods; *how* a step body is factored is the model's business. Policies (CFG, expert routing,
|
||||
precision, flow-shift, conditioning) are the *default* decomposition that deletes duplication for the families that
|
||||
fit — never an admission requirement.
|
||||
|
||||
### 5.2 Loop granularity
|
||||
|
||||
Chosen by runtime value, not purity:
|
||||
|
||||
- too coarse → cannot cancel/batch/reserve/stream/record at useful points;
|
||||
- too fine → scheduler overhead dominates, graph capture fragments;
|
||||
- good default → one denoise step (or window), one AR decode batch, one encoder chunk, one VAE-tile batch, one
|
||||
reward/logprob batch.
|
||||
|
||||
The runtime may **fuse** adjacent compatible WorkPlans after planning (an optimization); the unfused boundary remains
|
||||
the semantic model, so parity and behavior capture are defined on the unfused loop.
|
||||
|
||||
### 5.3 CFG is a policy over *one* shared denoise body (verified)
|
||||
|
||||
A natural worry: CFG changes the *shape* of the step (one forward vs two vs a batched pair vs a data-parallel split),
|
||||
so can one shared denoise loop really host all of it by swapping a policy? **Yes — and it is proven by existing code,
|
||||
not aspiration.** vllm-omni's `CFGParallelMixin.predict_noise_maybe_with_cfg` + `combine_cfg_noise`
|
||||
(`diffusion/.../cfg_parallel.py:76-212`) already runs sequential-2-forward, batched-1-forward, *and* cfg-parallel
|
||||
through **one** pair, with the loop body unaware of which. The clean cut is three layers:
|
||||
|
||||
- **In-loop `CFGPolicy`** — branch vocabulary (`[cond]`, `[cond, uncond]`, per-modality, STG-perturbed),
|
||||
the combine formula (standard `uncond + s·(cond−uncond)`, CFG-zero `st_star`, `cfg_normalize`/`guidance_rescale`),
|
||||
and **per-request mutable state** (the adaptive-gate cached delta with model-id self-invalidation is the canonical
|
||||
state case — and exactly why state lives in `LoopState`, §5.1). **Batched-vs-two-forward is a *dispatch detail
|
||||
inside one policy*, not a separate mechanism.** This covers classic / batched / adaptive-gate / per-modality.
|
||||
- **`cfg`-parallel is a *parallelism axis*, not a policy** — it shards the policy's branches across ranks and runs the
|
||||
*same rank-invariant `combine`* on every rank. It composes *under* any `CFGPolicy`; you own a `BatchedCFG` policy
|
||||
**or** a `cfg` group, never both (the §9 build-guard).
|
||||
- **Companions are an *orchestrator pattern*, not in the loop** — splitting a request into companion sub-requests
|
||||
upstream of diffusion (the conditioning is precomputed and bundled in; the loop is unchanged).
|
||||
|
||||
Two caveats keep the first pass honest: the `combine` runs in *the step body's* numeric space (Cosmos combines in
|
||||
x0-space after EDM preconditioning, not noise-space — the body fixes the space, the policy fixes the algebra), and
|
||||
embedded-guidance (Flux) is a **degenerate single-branch identity-combine policy** (guidance rides inside the forward
|
||||
kwarg), kept *inside* the same abstraction rather than special-cased as "no CFG." This is the same shared denoise body
|
||||
that RL rollout reuses (§10) — one CFG taxonomy serves both serving and rollout.
|
||||
|
||||
---
|
||||
|
||||
## 6. Runtime and scheduler
|
||||
|
||||
### 6.1 One WorkUnit, one currency
|
||||
|
||||
Every `await ctx.execute(plan)` produces a **WorkUnit**: the smallest schedulable action with a resource reservation
|
||||
and a loop boundary. Kinds: `ar_prefill`, `ar_token`, `diffusion_step`, `diffusion_window`, `chunk_step`,
|
||||
`encoder_chunk`, `vae_tile`, `audio_chunk`, `reward_batch`, `logprob_batch`, `transfer`, `cache_io`, `graph_capture`.
|
||||
Tokens are *one kind*, not the scheduler — this is the generalization of vLLM's token scheduler that diffusion forces.
|
||||
|
||||
**The budget currency is predicted GPU-time, not counts.** A bidirectional denoise step re-attends the full latent at
|
||||
O(L²) with zero KV amortization (every step pays full price); an AR decode step is ~O(context) against a cache; a
|
||||
chunked-causal step sits between. Counting "steps" or "tokens" puts items three orders of magnitude apart in one
|
||||
bucket. So each WorkUnit converts to GPU-seconds via its `LoopSpec.step_cost_model`, calibrated online by the Profiler
|
||||
(§11). **The same cost model is the interface published to the fleet** (§14): the scheduler's internal budget and
|
||||
Dynamo's routing/autoscaling input are one object, built once.
|
||||
|
||||
Two honesty caveats, kept from design.md's contact with reality:
|
||||
|
||||
- **Admission uses the conservative baseline.** The design's own flagship features make realized cost unknowable in
|
||||
advance — cache-dit skips are residual comparisons, VSA tiles are content-dependent, AR length is unbounded
|
||||
(budgeted at the `max_tokens` cap, refunded on early EOS). Telemetry refines calibration; it never licenses
|
||||
admission optimism.
|
||||
- **A denoise step is indivisible.** A 30s-1080p step bounds iteration latency no matter the budget. Mitigations are
|
||||
first-class, not afterthoughts: cost-class pools (jumbo steps don't co-schedule with latency-class work), SP within
|
||||
a node to shrink jumbo wall-time, and admission-time SLO classes so the fleet planner scales pools per class. This
|
||||
is why the scheduler is *cost-aware*, not just *count-aware*.
|
||||
|
||||
### 6.2 WorkPlan and admission
|
||||
|
||||
```python
|
||||
class WorkPlan:
|
||||
loop_id: str
|
||||
instance_id: str
|
||||
kind: str
|
||||
shape_sig: ShapeSignature # for batch compatibility + graph capture key
|
||||
resources: ResourceRequest # compute (GPU-s), resident bytes, peak-activation bytes, cache blocks, xfer bw, sinks
|
||||
cache: CachePlan # typed reads/writes (§7)
|
||||
placement: PlacementHint
|
||||
cancel_scope: CancelScope
|
||||
emits: list[StreamChunk]
|
||||
class Done: result: LoopResult
|
||||
```
|
||||
|
||||
**Admission rule (the soundness condition of multiplexing):** *do not admit a waiting WorkUnit unless every resource
|
||||
it requests can be reserved* — compute budget **and** memory (resident + worst-case peak) **and** cache blocks **and**
|
||||
transfer bandwidth **and** graph-capture shape **and** output sinks. Two requests that fit individually but jointly OOM
|
||||
are rejected at admission, not discovered at step 37. This is vLLM's "token budget is half the story, `allocate_slots`
|
||||
is the other half," generalized.
|
||||
|
||||
### 6.3 The scheduler, in layers (each testable on a fake pool, no GPU)
|
||||
|
||||
1. **RequestScheduler** — accepts requests/sessions, selects programs, starts loop drivers.
|
||||
2. **LoopScheduler** — drives `next()`, collects pending WorkPlans.
|
||||
3. **BatchScheduler** — groups compatible WorkPlans by `(instance, loop_kind, shape_sig, precision, parallel_plan,
|
||||
graph_key)`; image diffusion and AR decode batch across requests, jumbo video stays batch-of-1, Cosmos-style
|
||||
token-budget packing is an opt-in.
|
||||
4. **PlacementScheduler** — worker, role pool, instance, device mesh.
|
||||
5. **TransferScheduler** — tensor / cache / artifact movement as scheduled WorkUnits.
|
||||
6. **AdmissionController** — the reservation gate of §6.2.
|
||||
|
||||
Policies: running loops first (vLLM); preempt only at loop step boundaries; cancel only at declared scopes; prefer
|
||||
cache hits when latency/fairness allow; never starve a long denoise loop behind short AR requests.
|
||||
|
||||
### 6.4 SPMD consistency and failure isolation
|
||||
|
||||
All ranks of a pool must make identical scheduling decisions or NCCL deadlocks. **Rank-0 decides, broadcasts** — the
|
||||
existing discipline, now also the channel for the **abort broadcast** (failure isolation and scheduling share one
|
||||
consistency mechanism). Failure classes: *request-fatal* (NaN flagged by NaNWatch, one request's step error) →
|
||||
SPMD-consistent abort of that request, deliver partial artifacts with a structured error; *pool-fatal* (illegal
|
||||
access, NCCL desync) → pool re-init, invalidate pool caches, resume requests from serialized `LoopState` where one
|
||||
exists. **Cancellation is common-path, not exceptional** — vibe-directing makes abandoning in-flight work the *normal*
|
||||
user action; cancel takes effect at the next step boundary, drops queued WorkUnits, releases `LoopState` and cache
|
||||
handles, reports `cancelled`.
|
||||
|
||||
---
|
||||
|
||||
## 7. Memory, cache, transport, compile
|
||||
|
||||
Video and omni inference are memory systems as much as compute systems. This plane is explicit and typed.
|
||||
|
||||
### 7.1 Cache correctness is a contract
|
||||
|
||||
```python
|
||||
class CacheKey:
|
||||
model_id: str; component_id: str; loop_id: str | None
|
||||
weights_version: str; adapter_versions: dict[str, str]
|
||||
precision: str; parallel_plan_hash: str
|
||||
shape_sig: str; layout_sig: str
|
||||
scheduler_sig: str | None; guidance_sig: str | None; seed: int | None
|
||||
input_hashes: dict[str, str]; step_index: int | None
|
||||
contract_version: str
|
||||
```
|
||||
|
||||
**If a field can change output semantics, it is in the key.** Incorrect reuse is worse than no reuse. The serving
|
||||
hazard this kills: a workflow-cloud request that shares a prompt but differs in te-LoRA stack must not serve stale
|
||||
embeddings — so the key is *partitioned* by `adapter_versions`, not flushed. An RL `update_weights` bumps
|
||||
`weights_version` and invalidates wholesale.
|
||||
|
||||
### 7.2 Per-class pools (the granularity reality)
|
||||
|
||||
There is **no single unified block pool**, because a unified pool requires uniform bytes-per-block and our cache
|
||||
classes differ by 150–500× in natural granularity (a text-KV page ≈ 64 KB/layer; a causal-video latent-chunk slab is
|
||||
9.6–32 MB/layer) and their demand is workload-decoupled. Each class gets a statically budgeted pool behind one
|
||||
`CacheHandle`: paged text-KV (`ar_decode`), slab chunk-KV (`chunk_rollout`, with a declared training mode that
|
||||
disables mid-rollout recycling and keeps grad-aware index snapshots), feature caches (text/vision-encoder, content-hash
|
||||
keyed, reference-counted FIFO), residual caches (cache-dit, scoped per `LoopState`), weight/adapter cache
|
||||
(disk→CPU→GPU LRU for the workflow cloud). MoT falls out: the und pathway draws paged KV, the gen pathway draws slab or
|
||||
nothing — independent budgets, no interference. **KV is the minority case** — a pure bidirectional deployment allocates
|
||||
none of it; the machinery materializes only when a card declares KV-bearing loops.
|
||||
|
||||
### 7.3 Memory, transport, compile
|
||||
|
||||
- **Memory** — tagged pools, sleep/wake by tag (CuMem-style; tags are component names), reservation before admission,
|
||||
per-role budgets, host-pinned staging. Sleep/wake is component-granular for RL (drop DiT + caches, keep
|
||||
VAE/text-encoder resident).
|
||||
- **Transport** — manifest-based and pluggable: in-proc reference → SHM → CUDA IPC → NCCL/UCXX/NIXL/RDMA →
|
||||
object-store. KV/cache-bearing edges speak a `KVConnector`-shaped protocol (scheduler-side query/alloc/finish +
|
||||
worker-side async load/save) so NIXL/LMCache/Mooncake/KVBM implement it directly. Transfers are scheduled WorkUnits,
|
||||
not side effects.
|
||||
- **Compile** — CUDA graphs and `torch.compile` managed by a `CompileCache` keyed on `(model, component, loop,
|
||||
work_kind, shape_sig, precision, parallel_plan, backend)`. **Never full-graph across the engine** (per vLLM's own
|
||||
reversal): per-block compile where it pays, manual fused ops permitted in model code, breakable CUDA graphs as an
|
||||
*optimization tier* over an always-correct eager baseline. Graph capture is planned by the scheduler (padding,
|
||||
bucketing, capture sizes affect admission and batching).
|
||||
|
||||
---
|
||||
|
||||
## 8. Parallelism as a model contract
|
||||
|
||||
Parallelism is not a launch flag; it affects cache keys, scheduling, transport, capture, and parity, so it lives on
|
||||
the card.
|
||||
|
||||
```python
|
||||
class ParallelPlan:
|
||||
axes: dict[str, int] # dp, tp, sp(=ulysses×ring), cp, cfgp(≤2), pp_patch, vae, ep, fsdp, role, replica
|
||||
mesh_order: list[str]
|
||||
placement: PlacementSpec
|
||||
communication: CommunicationSpec
|
||||
```
|
||||
|
||||
Declarative, validated, compiled to a PyTorch `DeviceMesh` via a `ParallelDims`-style builder
|
||||
(product-of-degrees validation, cached submeshes). **Pre-flight or it fails at load, never halfway.** Ownership
|
||||
conflicts are build errors (CFG owned by a `BatchedCFG` *policy* or a `cfgp` *group*, never both). Applicability
|
||||
conditions travel with axes: `pp_patch` (PipeFusion displaced-patch pipelining) is **invalid for causal/AR** (stale KV
|
||||
breaks causality) and the validator enforces it per card. Degree-one axes exist as trivial groups so component code
|
||||
needs no special cases. **Pools are single-node**; multi-node scale is *multiple pools* fronted by the fleet (§14) —
|
||||
the engine never owns a cross-node NCCL mesh inside one pool.
|
||||
|
||||
---
|
||||
|
||||
## 9. Correctness — parity as a typed gate
|
||||
|
||||
This is the section both prior documents needed and neither fully had.
|
||||
|
||||
### 9.1 The parity contract
|
||||
|
||||
Every card carries a `ParitySpec`. Parity is **measured, never assumed**, by a `ParityAligner` observer (§11): record
|
||||
named taps per step/block from a reference (the official framework, or a pre-change build); compare-mode replays with
|
||||
fixed seeds and reports the first divergence beyond per-tap tolerance. This is the engine behind the "old loop vs new
|
||||
loop, bit-identical" gate and the standing instrument for every port, precision change, and kernel swap.
|
||||
|
||||
### 9.2 The consistency ladder (with the rung both prior docs missed)
|
||||
|
||||
```text
|
||||
C0 component parity — VAE, encoder, transformer block, scheduler step in isolation
|
||||
C1 loop parity — full denoise trajectory / AR logits, fixed seed
|
||||
C2 behavioral identity — the train-forward and serve-forward agree on the quantity the RL objective uses:
|
||||
· likelihood-based methods (GRPO-class): per-step log-prob identity
|
||||
· likelihood-free methods (DiffusionNFT-class): seeded final-sample +
|
||||
prediction-space identity (old_deviate / ref-MSE) — there are NO log-probs to match
|
||||
C3 distribution parity — rollout distribution under allowed nondeterminism
|
||||
C4 artifact quality — SSIM-class, reward agreement, human-preference (gates product claims; needs the eval system)
|
||||
```
|
||||
|
||||
The C2 split is load-bearing and is the lesson of the landed RL stack: the shipped Wan DiffusionNFT is
|
||||
**likelihood-free** — it captures only final clean latents and contrasts the student against an implicit negative
|
||||
policy in prediction space, so "log-prob identity" is *undefined* for it. A ladder that assumes log-probs (as both
|
||||
`design.md`'s and `designv2.md`'s early framings did) cannot describe the only RL method actually in the tree. RL
|
||||
methods declare their required level on the `RecipeSpec`.
|
||||
|
||||
### 9.3 The gate that catches what batch-of-1 cannot
|
||||
|
||||
Loop inversion's real hazard is **cross-request state smearing under interleaving** — and a batch-of-1 parity gate is
|
||||
*structurally blind* to it, because the corruption only manifests when two requests share a loop. §5.1 excludes the
|
||||
hazard by construction (state in `LoopState`, never globals), but construction-arguments need a test. So v3 makes a
|
||||
**batch-of-N interleave parity test** a *required* gate: two (or more) concurrent requests, interleaved at step
|
||||
granularity, must be bit-identical to the same requests run serially. This is the test the whole loop-inversion bet
|
||||
lives or dies on, and it is named here as a first-class obligation, not left implicit.
|
||||
|
||||
### 9.4 Three execution profiles, one definition
|
||||
|
||||
Even in one runtime there are three forwards: the **serve** forward (no-grad, graphed, cached, possibly quantized), the
|
||||
**rollout** forward (serve profile + behavior capture), and the **train** forward (grad, checkpointed, FSDP-gathered).
|
||||
They share *one* loop definition; they differ only in grad mode and capture. The ladder measures the gap; the recipe
|
||||
declares the level it needs. "Train BF16, serve FP8" is legal only at C2-corrected with importance-sampling, and the
|
||||
card says so. This is how the (recipe, runtime) pair stays honest: the contract is typed and tested, not trusted.
|
||||
|
||||
---
|
||||
|
||||
## 10. Training and RL on the same loops
|
||||
|
||||
```text
|
||||
serve : request → program → loop → WorkUnits → artifacts
|
||||
rollout : prompt batch → program → loop → WorkUnits → BehaviorRecords → rewards → update
|
||||
```
|
||||
|
||||
The loop kernel is shared; the only difference is output capture and training policy — not a second interpretation of
|
||||
the model. This is design.md's §8 thesis and v2's training plane, with the dependency rule kept absolute: **`training`
|
||||
may require behavior records but must not fork serving loop logic; the engine never imports `training`.** The engine
|
||||
*is* the rollout engine (it already runs the loops); the trainer is a client.
|
||||
|
||||
**This is the moat — and it is the one place a serving-only runtime structurally cannot follow.** vllm-omni proves
|
||||
omni serving can be production-grade, but it has *no* training/RL plane at all; verl-omni and miles prove the
|
||||
alternative — a standalone trainer-side sampler on a *different* runtime than serving — costs the two-runtime tax
|
||||
forever. The whole point of collocation is that **the rollout forward *is* the serve forward plus capture**: same loop,
|
||||
same caches, same batcher, same numerics. Three consequences nothing else gets:
|
||||
|
||||
- **Every serving optimization is automatically a rollout optimization.** Distilled few-step samplers, cache-dit
|
||||
skips, CFG-parallel, paged/feature caches, step batching — the recipe team builds them once for serving and the RL
|
||||
rollout inherits them for free. FastVideo's *own* landed DiffusionNFT is the negative example that proves the
|
||||
point: it vendors a bare-model `for`-loop (`rl/common/sampling.py`, whose docstring says it "intentionally does not
|
||||
call FastVideo's full inference pipelines"), and DMD2 vendors a *second* one (`dmd2.py::_student_rollout`) — so
|
||||
today's rollout runs with **zero serving-grade optimizations** (no CFG, dense attention, full 25-step ODE,
|
||||
one-sample-at-a-time). Collocation deletes both private loops.
|
||||
- **RL rollout is a *better* batching case than open-world serving — not a worse one.** A GRPO/NFT group is K
|
||||
*identical-config* samples of one prompt: same shape, same schedule, same CFG branch. The landed config is K=24
|
||||
(`num_video_per_prompt: 24`), 6 prompts/batch × 48 batches = **288 prompt-slots/GPU/epoch, each a 24-wide homogeneous
|
||||
denoise batch** — zero bucketing required (serving must bucket heterogeneous resolutions/steps/CFG across users; a
|
||||
GRPO group is homogeneous *by construction*). And all K samples share one prompt embedding, so the content-hash
|
||||
feature cache computes the text encoder **once per group instead of 24×**. The vendored sampler captures none of
|
||||
this; it carries the embedding per sample and runs one shape at a time.
|
||||
- **One numerics surface.** Serve-forward and rollout-forward differ only in grad mode and capture (§9.4), so there is
|
||||
no rollout-vs-train kernel gap to patch — the consistency ladder *measures* the gap rather than a correction layer
|
||||
*papering over* it. For the landed likelihood-free NFT, "reuse holds" means it holds at the **C2 behavioral rung**
|
||||
(seeded sample + prediction-space identity), under a `CFGPolicy` that is conditional-only and a `WeightSyncPlan`
|
||||
whose role is the decay-blended old policy — all of which the card already declares.
|
||||
|
||||
- **BehaviorRecord** — captured at generation time (reconstructing later is fragile): seeds, scheduler trajectory,
|
||||
timesteps, latents-or-refs, logprobs *where applicable*, sampled/action tokens, guidance, reward in/out, cache
|
||||
assumptions, precision, parallel plan, attention backend, deterministic flags, `weights_version`. Sized honestly:
|
||||
full MoE-routing capture is GB/sample for Cosmos3-class requests, so it is an **opt-in instrument** for goldens and
|
||||
debugging, not always-on.
|
||||
- **Weight-sync lifecycle** — freeze admission for the affected role/version → drain or boundary-stop in-flight loops →
|
||||
transfer weights/deltas → bump `weights_version` → invalidate incompatible caches and graphs → publish version →
|
||||
resume. A `WeightSyncPlan` is three inputs (mesh specs + per-model layout adapters + transport), validated
|
||||
pre-flight, CPU-testable on fake pools. RL ships a *role*, not "the weights": student / EMA / decay-blended old
|
||||
policy is declared (the landed NFT behavior policy is the *old* copy, not the student — the plan must carry that).
|
||||
- **Roles** (policy, rollout, reference, reward, critic, evaluator, data, coordinator) reference the same cards and
|
||||
loops; they are deployment concerns, scaled by the fleet.
|
||||
- **The industry tax we delete:** verl-omni re-implements Wan inside vLLM-Omni and corrects numerics afterward; miles'
|
||||
headline features (TIS/MIS, bitwise logprobs, R3 routing replay, unified FP8) are all mismatch patches for *two
|
||||
runtimes with different kernels*. One model definition, one kernel set, one measured ladder is the answer — viable
|
||||
at FastVideo's 1–30B FSDP2 scale (the boundary condition: a Megatron-class trainer at 100B+ re-enters the
|
||||
two-runtime world, and the ladder is the fallback there).
|
||||
|
||||
---
|
||||
|
||||
## 11. Extensions — observers and interceptors
|
||||
|
||||
The optimization, debugging, and parity surface, as versioned hook points assembled at loop build (an unused hook is
|
||||
*literally absent* from the hot path). It composes with §5 cleanly: the hooks wrap `ctx.execute(plan)`.
|
||||
|
||||
- **Observers (read-only):** `ParityAligner` (§9), `Profiler` (per-step wall+CUDA, calibrates the cost model),
|
||||
`NaNWatch` (first-NaN localization), `ActivationTrace`. They cannot mutate state.
|
||||
- **Interceptors (compute-altering):** `StepInterceptor` (step-skip / cached-prediction) and `BlockInterceptor`
|
||||
(cache-dit's DBCache/FBCache/TaylorSeer). State lives in `LoopState.plugin_state[id]`, keyed **per request and per
|
||||
CFG branch** — the structural fix for the module-global residual state that silently corrupts cache-dit/TeaCache
|
||||
forks under concurrency. cache-dit is the reference integration (the library sglang's serving already uses);
|
||||
conflicting interceptors are rejected pre-flight; a 4-step distilled card *rejects* step-skip caches rather than
|
||||
producing garbage.
|
||||
|
||||
**Trust boundary:** plugins are enabled at deploy scope only (never a per-request `plugins=[...]` field that would wire
|
||||
third-party code selection into the public API); requests only *parameterize* pre-enabled plugins through validated
|
||||
schemas, and exact-mode requests reject `distribution_altering` parameterization outright.
|
||||
|
||||
---
|
||||
|
||||
## 12. Request, session, artifact, stream
|
||||
|
||||
Typed runtime objects, not IDs in a batch (the Dreamverse/LiveKit lesson):
|
||||
|
||||
- `Request` — one generation, scoring, encoding, training-sample, or conversion job.
|
||||
- `Session` — a long-lived interactive context: prompt memory, media streams, cancellation, partial updates,
|
||||
cross-request chunk-KV that persists for a game/scene session.
|
||||
- `Artifact` — a *named, typed* output with provenance (which node produced it): `VideoArtifact`, `AudioArtifact(
|
||||
sample_rate)`, `TextArtifact(token_ids, text)`, `TensorArtifact`, `LatentArtifact`. This kills the `extra["audio"]`
|
||||
pattern — audio carries its sample rate as a first-class artifact, not a dict passenger.
|
||||
- `Stream` — one ordered event channel for previews, media chunks, progress, logs, finals.
|
||||
- `CancelScope` — structured cancellation target (request / loop / stream / session).
|
||||
|
||||
Typed event taxonomy (`request.*`, `session.*`, `artifact.*`, `media.{init,chunk,complete}`, `trace.*`). A
|
||||
`media.chunk` must know its stream, byte-range or shared-buffer ref, codec/container, timestamp range, and
|
||||
preview-vs-final — invalid combinations are unrepresentable.
|
||||
|
||||
The **request is the only currency crossing the product boundary.** A typed `Request` carries `task: TaskType`
|
||||
(declared, never inferred), `inputs: list[ModalPart]` (Text/Image/Video/Audio/Action/Latent), AR `sampling` vs
|
||||
`diffusion` params, an `OutputSpec` (requested modalities + streaming + capture flags), and per-node overrides. Task is
|
||||
declared; heuristics may only *suggest* a default at the boundary.
|
||||
|
||||
---
|
||||
|
||||
## 13. Programs and workflows
|
||||
|
||||
A **Program** composes a card's loops into a task; the card says what loops *exist*, the program says how to *run* them
|
||||
for this request. Kinds: `InlineProgram` (many loops, one resident instance — the omni default), `DisaggregatedProgram`
|
||||
(encoder→denoiser→decoder role pools), `WorkflowProgram` (compiled from ComfyUI), `TrainingProgram`,
|
||||
`RealtimeProgram`. Nodes: `ModelLoopNode`, `ComponentNode`, `ExternalNode`, `ArtifactNode`, `ControlNode`,
|
||||
`StreamNode`, `TransferNode`. Edges are typed (`TensorEdge`, `ArtifactEdge`, `StreamEdge`, `ControlEdge`, `CacheEdge`,
|
||||
`BehaviorEdge`). Linear pipelines are the degenerate case; branches/fan-out/fan-in are real (video and audio decode in
|
||||
parallel after a joint denoise). A separate deploy config maps nodes → pools/devices/parallelism, defaulting to "one
|
||||
pool, everything colocated."
|
||||
|
||||
**Workflows compile, they are not the runtime.** A ComfyUI workflow's tier-1/tier-2 static sublanguage maps onto a
|
||||
`Program` (`CheckpointLoaderSimple→card`, `KSampler→diffusion_denoise` with sampler/CFG policies,
|
||||
`LoraLoader→adapter hot-swap`, `ControlNetApply→ConditioningInjector`); unknown nodes become `ExternalNode`s or a
|
||||
coverage rejection — never silent wrongness. The moat: an orchestrator can run stock workflows on rented silicon;
|
||||
substituting a *credibly faster* model requires owning the recipe (§2.1) — which an orchestrator structurally cannot
|
||||
do. Equivalence is a quality-metric vs a reference render (C4), never a bit-parity claim.
|
||||
|
||||
---
|
||||
|
||||
## 14. Deployment and fleet
|
||||
|
||||
The engine exports a `DeploymentCard` and lets a fleet orchestrator (Dynamo) route — Dynamo orchestrates engines, it
|
||||
is never the engine core.
|
||||
|
||||
```python
|
||||
class DeploymentCard:
|
||||
engine_id: str; model_cards: list[str]
|
||||
capabilities: CapabilityMatrix; role_pools: list[RolePoolSpec]
|
||||
supported_programs: list[str]; supported_parallel_plans: list[ParallelPlan]
|
||||
cache_events: list[CacheEventSpec]; transfer_endpoints: list[TransferEndpoint]
|
||||
cost_model: CostModel # the SAME §6 cost model — one object, two consumers
|
||||
health: HealthSchema; slo: SLOSchema
|
||||
```
|
||||
|
||||
Clean line: the **fleet** owns global routing, tenant policy, cold start, role-pool scaling, cross-node transfer,
|
||||
placement-by-SLO, health/failover, multi-engine upgrades, global cache routing. The **engine** owns model load, loop
|
||||
execution, local scheduling, local memory/cache, model-specific behavior, parity, WorkUnit batching. The asks of the
|
||||
fleet are concrete and each has a fallback: generic affinity key-spaces (checkpoint/session/lora/weight_version beyond
|
||||
token prefixes), a heterogeneous request-cost interface (the §6 cost model), chunked media streaming through the
|
||||
frontend, role-graph disagg (N roles, not two), an RL weight plane (versioned broadcast + staleness-aware routing),
|
||||
cache-object tiering (KVBM generalized to latent/session caches), and session lifecycle as a routing primitive.
|
||||
|
||||
---
|
||||
|
||||
## 15. Worked examples
|
||||
|
||||
**(a) Text → video, one instance.** `Request(T2V)` → `InlineProgram` → `diffusion_denoise` loop. Driver: `init`
|
||||
builds sigmas/latents; `next` emits a `diffusion_step` WorkPlan (batch-of-1, SP+CFG-parallel); `advance` folds the
|
||||
model output and the CFG combine; cache-dit's `BlockInterceptor` may skip blocks based on the prior residual; at
|
||||
`Done`, a `vae_tile_decode` loop runs; output is a named `VideoArtifact`. Compiles/captures exactly like today's inner
|
||||
loop — nothing tensor-level changes for batch-of-1.
|
||||
|
||||
**(b) Cosmos3 omni, one request, shared weights.** `InlineProgram` over one `ModelInstance`: `ar_decode(reasoner)`
|
||||
yields `ar_token` WorkUnits that join the AR continuous-batching group → `pack` → `diffusion_denoise(vision+action+
|
||||
sound)` yields `diffusion_step` WorkUnits → fan-out `vae_tile_decode` + `audio_decode`. The reasoner's tokens and the
|
||||
denoiser's steps hit the *same resident weights*; the scheduler is the mode multiplexer. AR decode runs data-parallel
|
||||
across the cfg×sp weight-replica axes (decode is sequence-length-1; SP has nothing to shard). This is the workload no
|
||||
DAG-of-engines can express.
|
||||
|
||||
**(c) Image serving at scale.** Many `Request(T2I)` → the `BatchScheduler` groups `diffusion_step` WorkUnits by
|
||||
resolution bucket and batches across requests every step — the case where cross-request batching pays most. The
|
||||
*same* scheduler that runs (a) and (b).
|
||||
|
||||
**(d) RL rollout.** A `TrainingProgram` drives the *same* `diffusion_denoise` loop with `OutputSpec(capture=behavior)`;
|
||||
each step emits a `BehaviorRecord` slice; rollouts run C2 by construction (in-process, trainer kernels, pinned
|
||||
attention). For likelihood-free NFT the behavior is seeded final latents + prediction-space deviations; for a future
|
||||
GRPO-class method it is per-step log-probs — the loop is identical, the capture differs, the ladder rung is declared.
|
||||
|
||||
**(e) Dreamverse session.** A `Session(realtime_video_continue)` holds chunk-KV across 5s segments; `push_text`
|
||||
updates prompt memory; `stream` yields `media.chunk` previews from loop `emit`s; a direction change throws
|
||||
`Cancelled` at the next step boundary and starts a new segment. Capacity comes from duty cycle + cost-model admission +
|
||||
distillation — interleaving is fairness, not throughput.
|
||||
|
||||
**(f) ComfyUI compile.** `workflow.compile(json)` → a `WorkflowProgram` of `ModelLoopNode`/`ComponentNode` over a
|
||||
weight-fleet-cached card, with stacked-LoRA patch/unpatch priced by the §6 weight-transition cost — same runtime, new
|
||||
frontend.
|
||||
|
||||
---
|
||||
|
||||
## 16. What this unlocks (the unconstrained payoff)
|
||||
|
||||
Things no incremental design — and neither prior document — could actually claim:
|
||||
|
||||
- **True omni/MoT serving.** One resident model, many loop types, scheduled at step granularity, in one request. Not a
|
||||
monolith bypassing the abstraction (the Cosmos3 port's necessary hack), not a DAG doubling weights — native.
|
||||
- **Train ≡ serve by construction.** Because rollout and serve are the *same loop*, the (recipe, runtime) flywheel is
|
||||
real and measured, not aspirational: Dreamverse's directing sessions emit preference data → the RL plane → faster
|
||||
distilled cards → a better product, with the ladder guaranteeing the preferences collected under the serving profile
|
||||
transfer into training.
|
||||
- **Real-time interactive omni.** The driven-loop contract + sessions + WebRTC frame/PTS streaming + step-boundary
|
||||
cancellation make the <100ms motion-to-photon interactive world-model loop expressible in the same runtime that
|
||||
serves batch T2V.
|
||||
- **One substrate, three personas.** A research vehicle (new ports land as cards), a product engine (Dreamverse, the
|
||||
workflow cloud), and an RL rollout engine — without three codebases. The dependency rules keep them from fusing into
|
||||
mud.
|
||||
- **Correctness you can sign.** A deployable card is a *(recipe, runtime)* pair with a typed parity obligation; "this
|
||||
fast model is equivalent" is a claim with a test behind it, which is the one thing an orchestrator-without-recipes
|
||||
can never say.
|
||||
|
||||
---
|
||||
|
||||
## 17. Honest unknowns and falsifiers
|
||||
|
||||
An unconstrained design is not an unfalsifiable one. The bets, stated with the experiment that kills each:
|
||||
|
||||
- **The novelty is concentrated and real.** Runtime-owned diffusion iteration has a **narrow, opt-in precedent** —
|
||||
vllm-omni's `SupportsStepExecution` (`prepare_encode/denoise_step/step_scheduler/post_decode`,
|
||||
`diffusion/models/interface.py:44-67`) is exactly runtime-owned diffusion iteration at step granularity, and maps
|
||||
almost 1:1 onto our `init/next/advance/finalize` — but it is Qwen-Image-only and off in every shipped deploy. What
|
||||
is unprecedented is making it the **always-on universal contract** *and* a fully general WorkUnit scheduler over
|
||||
heterogeneous units. The risk is not the loop contract (a state machine is well-understood, and now demonstrably
|
||||
shippable); it is whether step-level cross-request scheduling *pays* for video. **Falsifier:** publish a load profile and targets from a real duty-cycle trace; if step-level scheduling does
|
||||
not beat a request-level baseline (≥2 concurrent sessions/GPU, p95 within SLO), the scheduler degrades to
|
||||
request-level dispatch and the loop contract keeps only its streaming/cancellation/behavior seams — which still
|
||||
justify it. The contract is safe even if the scheduling bet loses; that is the design's insurance.
|
||||
- **The general WorkUnit scheduler may be over-general.** Scheduling VAE tiles, transfers, and graph-captures through
|
||||
the *same* admission machinery as denoise steps is elegant and unproven. **Falsifier:** if, after Phase 2, the
|
||||
non-diffusion/non-AR WorkUnit kinds (tile, transfer, cache_io) gain nothing from unified scheduling over a simple
|
||||
in-loop call, collapse them back to in-loop operations and keep WorkUnits for the step-bearing kinds only.
|
||||
- **Cost-model admission is a modeling bet.** It converges toward cost-class pool routing once the indivisible-step
|
||||
reality is respected — which is close to what request-level pooling + a fleet planner already do. The fine-grained
|
||||
interleave win has a *narrow* window (many small concurrent jobs); it should be argued on that window, measured.
|
||||
- **The clean-slate premise is the elephant.** This document deliberately ignores migration. The org that would build
|
||||
it broke its own freeze 19 times and ships 20+ families, a live product, and a landed RL stack. A clean-slate
|
||||
rebuild is the highest-risk path that exists for *this* org; the responsible realization is to build v3 as a
|
||||
*parallel* engine around one forcing-function card (Cosmos3), prove it on the parity ladder, then migrate families
|
||||
onto it behind an adapter while everything keeps shipping — i.e., reach this architecture incrementally. That plan is
|
||||
out of scope here by request; it is non-optional in reality.
|
||||
- **Quality is unmeasured.** C4 (artifact quality / human preference) and the eval system it needs do not exist yet,
|
||||
and they gate every product claim ("fast mode is equivalent", RL reward validity, distillation comparisons). Named
|
||||
as a required, currently-absent subsystem, not assumed.
|
||||
|
||||
---
|
||||
|
||||
## 18. Package layout
|
||||
|
||||
```text
|
||||
fastvideo/
|
||||
card/ specs, components, loops, recipes, parity, checkpoints, capabilities # the Model Plane
|
||||
loop/ driver, loopstate, workplan, policies (cfg, expert, precision, flowshift, conditioning)
|
||||
runtime/ engine, scheduler/{request,loop,batch,placement,transfer,admission}, workers, events
|
||||
cache/ keys, classes/{paged_kv, slab_kv, feature, residual, weight_fleet}, policies
|
||||
memory/ allocator, sleep_wake, reservations
|
||||
transport/ manifests, backends/{shm, cuda_ipc, nccl, nixl, kvbm}, relay
|
||||
parallel/ plans, mesh, process_groups, validation
|
||||
parity/ aligner, ladder, interleave_gate # §9 is its own home
|
||||
extend/ observers, interceptors, cache_dit, registry, trust
|
||||
program/ specs, compiler, workflows
|
||||
request/ requests, sessions, artifacts, streams, cancel
|
||||
training/ rollout, behavior, rewards, weight_sync, methods # imports card/loop/runtime; never imported by them
|
||||
deploy/ cards, role_pools, dynamo_adapter
|
||||
integrations/ comfyui, dreamverse, livekit, diffusers
|
||||
```
|
||||
|
||||
Enforced boundaries: `card/` imports no product/runtime; `runtime/` executes `card/` loops but defines no semantics;
|
||||
`training/` may require behavior records but forks no loop; `integrations/` adapt external systems into core specs and
|
||||
events, never bypass them. **`parity/` is a first-class package**, not a test folder — it is how the (recipe, runtime)
|
||||
pair is kept honest.
|
||||
|
||||
---
|
||||
|
||||
## 19. Reference synthesis
|
||||
|
||||
| Source | Take | Constrain / reject |
|
||||
|---|---|---|
|
||||
| Cosmos3 (official + port) | Shared model instance across reasoning/diffusion/action/sound; packed multimodal sequences; component+scheduler parity matrices | A strong `ModelCard`, not the framework; no Cosmos-specific branching in global runtime |
|
||||
| vLLM core | Running-first scheduling, reservation-before-admission, model-owned state, encoder/KV cache managers, CuMem sleep/wake, CUDA-graph dispatch, KV-connector split | Token scheduling is one WorkUnit kind; never full-graph compile |
|
||||
| sglang `multimodal_gen` | Role pools, request lifecycle, capacity dispatch, transfer manifests, disagg state machine, cache-dit integration | No large mutable `Req`/`ForwardBatch` as the stable API; not single-item diffusion scheduling |
|
||||
| vLLM-Omni | Frozen pipeline spec separate from deploy YAML (verified, adopt); `OmniConnectorBase` + `chunk_ready` readiness; **`SupportsStepExecution` as loop-inversion prior art** (opt-in, Qwen-Image-only — we generalize to always-on); TP-rank- and CFG-branch-aware KV-copy transfer; **`CFGParallelMixin` proves CFG-as-policy over one shared denoise body** (§5.3); 3 separate cache subsystems confirm per-class pools | Expresses shared-weight MoT (`bagel`/`lance`) only as **one opaque request-scheduled stage** the scheduler never sees inside — no step visibility, no cross-request batching by default; cross-stage KV is a *copy*, not a shared live cache; **no cost model** (per-stage count budgets); readiness-parking, not credit flow; RDMA = Mooncake/Mori/Yuanrong, not NIXL/NCCL |
|
||||
| sglang-omni | The `next/wait_for/merge_fn/stream_to` edge vocabulary; Relay transport + **credit-based flow control** (this is sglang-omni's, not vllm-omni's) | Stages own disjoint weights; hybrid AR+diffusion only as AR-stage → DiT-stage; per-model bootstrap duplication |
|
||||
| Dynamo | Fleet routing, disagg role pools, KV-aware routing, KVBM, SLA planner, ModelExpress cold-start/weight streaming | Orchestrates engines; never the engine core. Export a `DeploymentCard` + cost model to it |
|
||||
| diffusers Modular | `ComponentSpec`/`modular_model_index.json` interchange; Guiders ≈ CFG policies | A Python pipeline interpreter is not the performance boundary; import is lossy |
|
||||
| xDiT | DiT parallelism catalog (USP, ring/ulysses, PipeFusion, CFG-parallel, DistVAE) + world-size validation | Parallelism lives in the runtime + card, not a wrapper-per-model library; `pp_patch` invalid for causal |
|
||||
| TorchTitan | Named mesh axes, `ParallelDims` validation, ModelSpec discipline, TorchStore weight-sync, batch-invariance utils | Adopt the discipline, not the stack; DCP/TorchStore don't reshard — `WeightSyncPlan` owns layout |
|
||||
| verl-omni / miles / cosmos-rl | Rollout adapters, per-step capture, async rewards, group-relative advantage, TIS/MIS, deterministic/batch-invariant modes, per-payload `weight_version`, AIPO/off-policy masking | The two-runtime tax is the thing to delete; capture behavior *in* the serving loop, not after the fact |
|
||||
| ComfyUI | Workflow graph, node-signature cache, model memory management, App-Mode (workflows-as-products) | Compile to `Program`; dynamic node execution is not the serving/training core; GPL hygiene |
|
||||
| Dreamverse | Sessions, prompt memory, typed media IPC, cancellation, the duty-cycle capacity reality, the preference-data flywheel | Product/session behavior is first-class in the request plane, never merged into the model core |
|
||||
| LiveKit | Realtime sessions, push audio/video, interruptions, turn/activity state, frame+PTS streaming | Realtime triggers only when they fire (<100ms interactive); don't force RTC onto offline jobs |
|
||||
| Thinking Machines (batch-invariance) | The C2/C3 mechanism: batch-invariant kernels for bitwise rollout↔train identity | Scoped to goldens; the conservative baseline governs admission |
|
||||
|
||||
---
|
||||
|
||||
## Final position
|
||||
|
||||
```text
|
||||
A model card is a (recipe, runtime) pair with a parity obligation.
|
||||
The model owns loop semantics; the runtime owns loop lifecycle.
|
||||
One resident instance runs many loops; one scheduler runs their steps in one currency.
|
||||
Caches are correct by key; parity is correct by test; the interleave gate is non-negotiable.
|
||||
Training records behavior on the same loops it serves.
|
||||
Deployment places and routes; products stream artifacts; neither defines the model.
|
||||
```
|
||||
|
||||
This is the ceiling: a model-native runtime where omni is native, train and serve are the same loops by construction,
|
||||
correctness is a typed contract you can sign, and the (recipe, runtime) flywheel is real. The constraint we removed to
|
||||
see it was migration. Putting that constraint back is the next document, not this one.
|
||||
+1895
File diff suppressed because it is too large
Load Diff
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,11 +55,13 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
// Add a one-click copy button for the rendered MkDocs article.
|
||||
(function () {
|
||||
const BUTTON_ID = "copy-page-button";
|
||||
|
||||
function text(value) {
|
||||
return (value || "").replace(/\s+/g, " ").trim();
|
||||
}
|
||||
|
||||
function codeLanguage(code) {
|
||||
const classes = Array.from(code.classList || []);
|
||||
const language = classes.find((name) => name.startsWith("language-"));
|
||||
return language ? language.replace("language-", "") : "";
|
||||
}
|
||||
|
||||
function serializeInline(node) {
|
||||
if (node.nodeType === Node.TEXT_NODE) {
|
||||
return node.textContent || "";
|
||||
}
|
||||
|
||||
if (node.nodeType !== Node.ELEMENT_NODE) {
|
||||
return "";
|
||||
}
|
||||
|
||||
const tagName = node.tagName.toLowerCase();
|
||||
|
||||
if (tagName === "code" && node.parentElement && node.parentElement.tagName.toLowerCase() !== "pre") {
|
||||
return "`" + (node.textContent || "").trim() + "`";
|
||||
}
|
||||
|
||||
if (tagName === "a") {
|
||||
if (node.classList.contains("headerlink")) {
|
||||
return "";
|
||||
}
|
||||
|
||||
const label = text(Array.from(node.childNodes).map(serializeInline).join(""));
|
||||
const href = node.href;
|
||||
return href && label ? `${label} (${href})` : label;
|
||||
}
|
||||
|
||||
if (tagName === "img") {
|
||||
const alt = node.getAttribute("alt") || "image";
|
||||
const src = node.src || "";
|
||||
return src ? `[${alt}](${src})` : `[${alt}]`;
|
||||
}
|
||||
|
||||
if (tagName === "br") {
|
||||
return "\n";
|
||||
}
|
||||
|
||||
return Array.from(node.childNodes).map(serializeInline).join("");
|
||||
}
|
||||
|
||||
function serializeTable(table) {
|
||||
const rows = Array.from(table.rows);
|
||||
if (rows.length === 0) return "";
|
||||
|
||||
const mdRows = rows.map(
|
||||
(row) => "| " + Array.from(row.children).map((cell) => text(serializeInline(cell))).join(" | ") + " |"
|
||||
);
|
||||
const separator = "| " + Array.from(rows[0].children)
|
||||
.map(() => "---")
|
||||
.join(" | ") + " |";
|
||||
mdRows.splice(1, 0, separator);
|
||||
return mdRows.join("\n");
|
||||
}
|
||||
|
||||
function serializeBlock(node, listDepth = 0) {
|
||||
if (node.nodeType === Node.TEXT_NODE) {
|
||||
return text(node.textContent);
|
||||
}
|
||||
|
||||
if (node.nodeType !== Node.ELEMENT_NODE) {
|
||||
return "";
|
||||
}
|
||||
|
||||
const tagName = node.tagName.toLowerCase();
|
||||
|
||||
if (["script", "style", "nav", "button"].includes(tagName) || node.id === BUTTON_ID) {
|
||||
return "";
|
||||
}
|
||||
|
||||
if (/^h[1-6]$/.test(tagName)) {
|
||||
const level = Number(tagName.slice(1));
|
||||
return `${"#".repeat(level)} ${text(serializeInline(node))}`;
|
||||
}
|
||||
|
||||
if (tagName === "pre") {
|
||||
const code = node.querySelector("code");
|
||||
const content = code ? code.textContent || "" : node.textContent || "";
|
||||
return `\`\`\`${code ? codeLanguage(code) : ""}\n${content.replace(/\n$/, "")}\n\`\`\``;
|
||||
}
|
||||
|
||||
if (["p", "figcaption"].includes(tagName)) {
|
||||
return text(serializeInline(node));
|
||||
}
|
||||
|
||||
if (tagName === "blockquote") {
|
||||
return serializeChildren(node, listDepth)
|
||||
.split("\n")
|
||||
.map((line) => (line ? `> ${line}` : ">"))
|
||||
.join("\n");
|
||||
}
|
||||
|
||||
if (tagName === "ul" || tagName === "ol") {
|
||||
return Array.from(node.children)
|
||||
.filter((child) => child.tagName && child.tagName.toLowerCase() === "li")
|
||||
.map((item, index) => serializeListItem(item, tagName === "ol", index, listDepth))
|
||||
.join("\n");
|
||||
}
|
||||
|
||||
if (tagName === "table") {
|
||||
return serializeTable(node);
|
||||
}
|
||||
|
||||
if (["hr"].includes(tagName)) {
|
||||
return "---";
|
||||
}
|
||||
|
||||
return serializeChildren(node, listDepth);
|
||||
}
|
||||
|
||||
function serializeListItem(item, ordered, index, listDepth) {
|
||||
const marker = ordered ? `${index + 1}. ` : "- ";
|
||||
const indent = " ".repeat(listDepth);
|
||||
const childBlocks = [];
|
||||
const inlineParts = [];
|
||||
|
||||
Array.from(item.childNodes).forEach((child) => {
|
||||
if (child.nodeType === Node.ELEMENT_NODE && ["ul", "ol"].includes(child.tagName.toLowerCase())) {
|
||||
childBlocks.push(serializeBlock(child, listDepth + 1));
|
||||
} else {
|
||||
const content = serializeInline(child);
|
||||
if (content) inlineParts.push(content);
|
||||
}
|
||||
});
|
||||
|
||||
const firstLine = `${indent}${marker}${text(inlineParts.join(" "))}`.trimEnd();
|
||||
return [firstLine, ...childBlocks.filter(Boolean)].join("\n");
|
||||
}
|
||||
|
||||
function serializeChildren(node, listDepth = 0) {
|
||||
return Array.from(node.childNodes)
|
||||
.map((child) => serializeBlock(child, listDepth))
|
||||
.map((value) => value.trim())
|
||||
.filter(Boolean)
|
||||
.join("\n\n");
|
||||
}
|
||||
|
||||
function articleText(article) {
|
||||
const clone = article.cloneNode(true);
|
||||
clone.querySelectorAll("script, style, .headerlink, .md-clipboard, #copy-page-button").forEach((node) => node.remove());
|
||||
|
||||
const content = serializeChildren(clone).trim();
|
||||
const title = document.querySelector("h1") || document.querySelector("title");
|
||||
const pageTitle = title ? text(title.textContent) : "";
|
||||
|
||||
if (pageTitle && !content.startsWith("# ")) {
|
||||
return `# ${pageTitle}\n\n${content}`.trim();
|
||||
}
|
||||
|
||||
return content;
|
||||
}
|
||||
|
||||
async function copyArticle(button, article) {
|
||||
const originalLabel = button.textContent;
|
||||
try {
|
||||
await navigator.clipboard.writeText(articleText(article));
|
||||
button.textContent = "Copied!";
|
||||
button.classList.add("copy-page-button--copied");
|
||||
} catch (error) {
|
||||
button.textContent = "Copy failed";
|
||||
button.classList.add("copy-page-button--error");
|
||||
console.error("Failed to copy page", error);
|
||||
}
|
||||
|
||||
window.setTimeout(() => {
|
||||
button.textContent = originalLabel;
|
||||
button.classList.remove("copy-page-button--copied", "copy-page-button--error");
|
||||
}, 2000);
|
||||
}
|
||||
|
||||
function addCopyButton() {
|
||||
const article = document.querySelector("article.md-content__inner");
|
||||
if (!article || document.getElementById(BUTTON_ID)) {
|
||||
return;
|
||||
}
|
||||
|
||||
const button = document.createElement("button");
|
||||
button.id = BUTTON_ID;
|
||||
button.type = "button";
|
||||
button.className = "copy-page-button md-button md-button--primary";
|
||||
button.textContent = "Copy page";
|
||||
button.setAttribute("aria-label", "Copy this page as plain text");
|
||||
button.addEventListener("click", () => copyArticle(button, article));
|
||||
|
||||
article.insertBefore(button, article.firstChild);
|
||||
}
|
||||
|
||||
if (typeof document$ !== "undefined") {
|
||||
document$.subscribe(addCopyButton);
|
||||
}
|
||||
|
||||
document.addEventListener("DOMContentLoaded", addCopyButton);
|
||||
window.addEventListener("load", addCopyButton);
|
||||
})();
|
||||
@@ -41,3 +41,20 @@ img {
|
||||
display: block;
|
||||
margin: 0 auto;
|
||||
}
|
||||
|
||||
.md-typeset .copy-page-button.md-button {
|
||||
float: right;
|
||||
margin: 0 0 1rem 1rem;
|
||||
padding: 0.2em 0.6em;
|
||||
font-size: 1em;
|
||||
font-weight: 400;
|
||||
line-height: 1.2;
|
||||
}
|
||||
|
||||
.md-typeset .copy-page-button.md-button.copy-page-button--copied {
|
||||
background-color: var(--md-accent-fg-color);
|
||||
}
|
||||
|
||||
.md-typeset .copy-page-button.md-button.copy-page-button--error {
|
||||
background-color: var(--md-code-hl-number-color);
|
||||
}
|
||||
|
||||
@@ -85,6 +85,124 @@ Each line in `/tmp/fv_trace.jsonl` is a JSON record:
|
||||
5. The first divergent line identifies the first layer where FastVideo and the
|
||||
upstream produce different outputs. Start debugging there.
|
||||
|
||||
## Performance impact
|
||||
|
||||
When the master toggle is off, overhead is nil. When on, the cost is
|
||||
proportional to how broadly the layer regex matches.
|
||||
|
||||
### What runs when tracing is on
|
||||
|
||||
For every match against `model.named_modules()`, FastVideo registers an
|
||||
`ActivationStatHook` that runs after the module's forward returns:
|
||||
|
||||
1. Walks the (possibly nested) output `tuple` / `list` / `dict` and extracts
|
||||
every `torch.Tensor` leaf.
|
||||
2. Computes each enabled stat on the tensor's `.detach().float()` view — the
|
||||
cast is required for bf16 inputs because some reductions are not stable
|
||||
in bf16.
|
||||
3. Writes one JSON record per tensor to the line-buffered JSONL sink.
|
||||
|
||||
Cost is roughly `O(num_matched_modules × num_output_tensors × num_stats × tensor_numel)`
|
||||
per forward, dominated by `tensor_numel` for value-stats (`abs_mean`, `sum`,
|
||||
`min`, `max`, `mean`, `std`). The `shape` and `dtype` stats are O(1).
|
||||
|
||||
### Cost-shaping knobs
|
||||
|
||||
The default config (empty `FASTVIDEO_TRACE_LAYERS` is treated as `.*`, default
|
||||
two stats, all steps) is intentionally blunt — useful only for a one-shot
|
||||
smoke run. For real debugging, scope down:
|
||||
|
||||
| Knob | Effect |
|
||||
|---|---|
|
||||
| Tighten `FASTVIDEO_TRACE_LAYERS` to a regex matching <50 modules | Linear reduction in hook count |
|
||||
| Drop unused stats from `FASTVIDEO_TRACE_STATS` (`mean`, `std`, `min`, `max`) if you only need divergence detection | One full reduction per stat saved per matched module per forward |
|
||||
| Set `FASTVIDEO_TRACE_STEPS="0,15,31"` for a 32-step run | ~10x reduction vs all-steps (the hook still fires but exits early when the step doesn't match) |
|
||||
| Use `shape` + `dtype` only on layers where you only care about layout | Skips tensor reductions entirely on those layers |
|
||||
|
||||
### Disk
|
||||
|
||||
Output is line-buffered (`open(..., buffering=1)`), so every record flushes
|
||||
on write. A typical 32-step run that traces 40 DiT blocks with 4 stats
|
||||
writes roughly 5K records — about 1 MB of JSONL. Point
|
||||
`FASTVIDEO_TRACE_OUTPUT` at a fast local disk for parity runs; slow network
|
||||
mounts will dominate runtime once tracing is on.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### "I set the env var but no JSONL file appears"
|
||||
|
||||
Three things to check, in order:
|
||||
|
||||
1. **Toggle semantics.** `FASTVIDEO_TRACE_ACTIVATIONS` uses a strict
|
||||
not-equal-to-`"0"` test. Setting it to `1`, `true`, or even `""` all
|
||||
enable tracing. Only an unset variable or `FASTVIDEO_TRACE_ACTIVATIONS=0`
|
||||
disables it.
|
||||
2. **Module exposure.** The hook attaches to `pipeline.modules.get("transformer")`
|
||||
at the end of `post_init`. If your pipeline does not expose a module
|
||||
under that key (e.g. a non-standard custom pipeline whose DiT is
|
||||
reachable only via `pipeline.modules["sr_transformer"]`), the trace is
|
||||
silently a no-op. Check the
|
||||
`Activation trace attached to N modules` log line at startup; if
|
||||
`N=0`, either the regex didn't match anything or the expected module
|
||||
isn't exposed.
|
||||
3. **Output path failure.** The parent of `FASTVIDEO_TRACE_OUTPUT` is
|
||||
auto-created. If creation fails (permissions, read-only mount),
|
||||
`JsonlSink.__init__` raises at startup — look for an `OSError` early
|
||||
in the log.
|
||||
|
||||
### "My regex isn't filtering the way I expect"
|
||||
|
||||
`FASTVIDEO_TRACE_LAYERS` is compiled with `re.compile(spec)` and matched
|
||||
with `pattern.search(name)`. Two consequences:
|
||||
|
||||
- `search`, not `fullmatch`. `block.layers` matches
|
||||
`transformer.block.layers.0.attn`. Anchor with `^...$` if you want
|
||||
exact matches.
|
||||
- Module names use Python dot notation (`transformer.block.layers.0`),
|
||||
not slashes. The `.` in your regex is a metacharacter — escape it as
|
||||
`\.` if you want a literal dot.
|
||||
|
||||
Run once with `FASTVIDEO_LOGGING_LEVEL=DEBUG` to see the
|
||||
`Activation trace attached to N modules (pattern=...)` line. `N` is the
|
||||
ground truth for how many modules survived your regex.
|
||||
|
||||
### "Stats are NaN or `<error: ...>` for some layers"
|
||||
|
||||
Some outputs hit numerical edge cases:
|
||||
|
||||
- `mean` / `std` on a 0-dim tensor returns NaN.
|
||||
- `abs_mean` on an empty tensor or one full of `inf` returns NaN.
|
||||
- Non-tensor outputs (e.g. a Python `bool` from a verification gate) are
|
||||
silently skipped — the hook only walks `torch.Tensor` leaves.
|
||||
|
||||
Dropping `mean`/`std` and using `abs_mean`/`max` is more robust. For
|
||||
layout-only debugging, `shape`+`dtype` never fail.
|
||||
|
||||
### "Tracing slows the run by 5x"
|
||||
|
||||
You're probably matching too broadly. An empty `FASTVIDEO_TRACE_LAYERS` is
|
||||
treated as `.*` and matches every named module — for a 15B-param DiT
|
||||
that's hundreds of submodules, each running stat reductions on every
|
||||
forward. Tighten to a single block depth: `^transformer\.blocks\.\d+$`
|
||||
typically matches a few dozen modules, which is a manageable trace.
|
||||
|
||||
### "Trace records are missing tensors I expect"
|
||||
|
||||
The hook walks `tuple` / `list` / `dict` outputs recursively but does
|
||||
**not** unpack custom dataclasses or named tuples — those are silently
|
||||
skipped. If your module returns
|
||||
`BlockOutput(hidden_states=..., attn_logits=...)`, no records are
|
||||
emitted. Workaround: either return a plain `dict` (`{"hidden_states": ..., "attn_logits": ...}`)
|
||||
or attach the hook to a deeper module that already returns a raw tensor.
|
||||
|
||||
### "I want tracing inside a `torch.compile`'d region"
|
||||
|
||||
Module forward hooks run on the eager wrapper. If the entire module is
|
||||
compiled, the hook sees only the wrapped op's output, not internal FX
|
||||
nodes. This is by design — Extension 1 (FX backend rewrite) in
|
||||
[Future extensions](#future-extensions-design-only-not-yet-implemented)
|
||||
is the planned path when inside-graph granularity is required.
|
||||
|
||||
## Architecture (Extension 0: module forward hooks)
|
||||
|
||||
At pipeline initialization, `attach_activation_trace()` reads the env vars once.
|
||||
|
||||
@@ -73,6 +73,7 @@ replicate CI results before pushing.
|
||||
| Transformer Tests | `transformer` | `fastvideo/models/dits/**`, `fastvideo/models/loader/**`, `fastvideo/tests/transformers/**`, `fastvideo/layers/**`, `fastvideo/attention/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
|
||||
| Kernel Tests | `kernel_tests` | `fastvideo-kernel/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
|
||||
| Unit Tests | `unit_test` | `fastvideo/**`, `.buildkite/**`, `.github/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
|
||||
| DreamVerse App Tests | `dreamverse_app` | `apps/dreamverse/**`, `pyproject.toml` |
|
||||
|
||||
A Fastcheck failure means a component-level regression. Check the Buildkite build log for the
|
||||
failing test's output.
|
||||
@@ -99,7 +100,7 @@ failing test's output.
|
||||
| LoRA Training Tests | `training_lora` | 15 min |
|
||||
| Training Tests VSA | `training_vsa` | 15 min |
|
||||
| Inference Tests VMoBA | `inference_vmoba` | 15 min |
|
||||
| Performance Tests | `performance` | 30 min |
|
||||
| [Performance Tests](performance_benchmarks.md) | `performance` | 30 min |
|
||||
| API Server Tests | `api_server` | 30 min |
|
||||
| Train Framework Tests | `train_framework` | 30 min |
|
||||
|
||||
@@ -268,6 +269,7 @@ Triggers a specific Buildkite test or suite on the current PR branch.
|
||||
| `/test transformer` | Transformer Tests (Fastcheck) | `transformer` |
|
||||
| `/test kernel` | Kernel Tests (Fastcheck) | `kernel_tests` |
|
||||
| `/test unit` | Unit Tests (Fastcheck) | `unit_test` |
|
||||
| `/test dreamverse` | DreamVerse App Tests (Fastcheck) | `dreamverse_app` |
|
||||
| `/test ssim` | SSIM regression tests | `ssim` |
|
||||
| `/test training` | Training pipeline tests | `training` |
|
||||
| `/test lora-inference` | LoRA inference tests | `inference_lora` |
|
||||
|
||||
@@ -0,0 +1,100 @@
|
||||
# Dreamverse Development
|
||||
|
||||
Dreamverse lives under `apps/dreamverse/` as a product app inside the
|
||||
FastVideo monorepo. Backend code uses the local FastVideo workspace package;
|
||||
frontend tooling remains standalone under `apps/dreamverse/web/`.
|
||||
|
||||
## Backend tests
|
||||
|
||||
Run CPU-safe backend tests from the FastVideo repository root:
|
||||
|
||||
```bash
|
||||
uv run --locked --package dreamverse --extra test pytest apps/dreamverse/server/tests/ -m 'not gpu' -q
|
||||
```
|
||||
|
||||
## Backend launch
|
||||
|
||||
Launch the migrated backend through the installed console commands:
|
||||
|
||||
```bash
|
||||
dreamverse-server --port 8009
|
||||
dreamverse-mock-server --port 8009
|
||||
```
|
||||
|
||||
If `dreamverse-server` is missing, install FastVideo with the `dreamverse`
|
||||
extra from the checkout:
|
||||
|
||||
```bash
|
||||
uv pip install -e ".[dreamverse]"
|
||||
```
|
||||
|
||||
## Frontend build and tests
|
||||
|
||||
Run frontend commands from the standalone web app:
|
||||
|
||||
```bash
|
||||
cd apps/dreamverse/web
|
||||
npm ci
|
||||
npm run build
|
||||
npm test
|
||||
```
|
||||
|
||||
Playwright is intentionally run against a live backend as part of the GPU4
|
||||
manual verification flow, not in the Phase 3 migration gate.
|
||||
|
||||
## Local GPU4 verification hook
|
||||
|
||||
Use physical GPU 4 for migration smoke tests. `CUDA_VISIBLE_DEVICES=4` makes
|
||||
that GPU appear as logical GPU 0 inside the process, preserving the previous
|
||||
Dreamverse deployment behavior.
|
||||
|
||||
```bash
|
||||
CUDA_VISIBLE_DEVICES=4 dreamverse-server --host 0.0.0.0 --port 8009
|
||||
```
|
||||
|
||||
In another shell, verify the service:
|
||||
|
||||
```bash
|
||||
curl -s http://localhost:8009/healthz
|
||||
```
|
||||
|
||||
Phase 4 adds the public `/healthz`, `/readyz`, `/status`,
|
||||
`/prompt-system-config`, and `/curated-presets` route coverage needed for the
|
||||
full Playwright suite.
|
||||
|
||||
## Phase 0 production-equivalent prerequisites
|
||||
|
||||
For the production-equivalent NVFP4 path, install these dependencies
|
||||
in the FastVideo `.venv` before GPU smoke tests:
|
||||
|
||||
```bash
|
||||
uv pip install --python .venv/bin/python \
|
||||
flashinfer-python flash-attn cerebras-cloud-sdk openai \
|
||||
--no-build-isolation
|
||||
```
|
||||
|
||||
| Package | Why |
|
||||
|---|---|
|
||||
| `flashinfer-python` | Required for NVFP4 quantization. Without it, model load fails with `ImportError: NVFP4 quantization requires flashinfer`. |
|
||||
| `flash-attn` | Optional but recommended; without it attention falls back to Torch SDPA (functional but slower). |
|
||||
| `cerebras-cloud-sdk` | Required by the migrated prompt enhancer for the default `cerebras` provider. |
|
||||
| `openai` | Required by the prompt enhancer's OpenAI-compatible providers + downstream rewrites. |
|
||||
|
||||
### B200 / sm_100a + gcc-15 conda toolchain (flashinfer JIT workaround)
|
||||
|
||||
On hosts where the conda toolchain ships gcc-15 (which nvcc rejects with
|
||||
`#error -- unsupported GNU version! gcc versions later than 14 are not
|
||||
supported!`), set these env vars before launching anything that triggers
|
||||
flashinfer's JIT kernel build:
|
||||
|
||||
```bash
|
||||
export CC=/usr/bin/gcc-13
|
||||
export CXX=/usr/bin/g++-13
|
||||
export CUDAHOSTCXX=/usr/bin/g++-13
|
||||
export NVCC_PREPEND_FLAGS="-ccbin /usr/bin/gcc-13 -allow-unsupported-compiler"
|
||||
```
|
||||
|
||||
`dreamverse-server` does NOT set these — they need to come from the launching
|
||||
shell. The `dreamverse-deploy` skill
|
||||
([`.agents/skills/dreamverse-deploy/`](../../.agents/skills/dreamverse-deploy/SKILL.md))
|
||||
sets them for you and is the recommended local-deploy path.
|
||||
@@ -11,14 +11,14 @@ Use this guide when you are:
|
||||
- Adding a new metric (native or wrapping a third-party library).
|
||||
- Porting a benchmark (e.g. VBench, MIND, EvalCrafter) whose Python
|
||||
code needs to be importable from a pinned upstream.
|
||||
- Adding a new metric group (audio, vlm, etc.).
|
||||
- Adding a new metric group (audio, videoscore2, etc.).
|
||||
|
||||
## TL;DR
|
||||
|
||||
Metrics are auto-discovered from
|
||||
`fastvideo/eval/metrics/<group>/<name>/metric.py`. Each declares itself
|
||||
with `@register("<group>.<name>")` and subclasses `BaseMetric`. Three
|
||||
recipes:
|
||||
with `@register("<group>.<name>")` and subclasses `BaseMetric`. Five
|
||||
recipes, depending on how the metric ships and what its licence allows:
|
||||
|
||||
1. **Native metric** (pure-PyTorch, no submodule). Drop a file,
|
||||
declare deps, implement `compute(sample)`.
|
||||
@@ -26,11 +26,18 @@ recipes:
|
||||
Same as above, plus route the library's cache through
|
||||
`get_cache_dir()` if it has a `download_root=` / `cache_dir=`
|
||||
kwarg.
|
||||
3. **Upstream-submodule-wrapped metric** (vbench-style). Pin upstream
|
||||
as a git submodule under `fastvideo/third_party/eval/<bench>/`. The
|
||||
adapter `__init__.py` does the `sys.path` insert and any runtime
|
||||
compat shims for modern dep versions. Patches live as Python in
|
||||
that file rather than as on-disk patches to the submodule.
|
||||
3. **Submodule-wrapped metric** (vbench-style). Pin upstream as a git
|
||||
submodule under `fastvideo/third_party/eval/<bench>/`. The adapter
|
||||
`__init__.py` does the `sys.path` insert and any runtime compat
|
||||
shims. Best for large research packages with stable layouts.
|
||||
4. **Vendored upstream** (synchformer / glmasr-style). Copy a small,
|
||||
surgical piece of upstream into `fastvideo/third_party/eval/<name>/`
|
||||
with its `LICENSE` alongside. Best for permissive-licensed
|
||||
(MIT / Apache-2.0) source you need a few files from.
|
||||
5. **Git-source dep via `[tool.uv.sources]`** (ImageBind-style). For
|
||||
license-restricted upstream (e.g. CC BY-NC-SA) that cannot be
|
||||
redistributed in the FastVideo tree. uv pulls the source at install
|
||||
time pinned to a SHA in `pyproject.toml`.
|
||||
|
||||
The full recipes are below.
|
||||
|
||||
@@ -43,7 +50,8 @@ fastvideo/eval/metrics/
|
||||
├── base.py # BaseMetric + lifecycle contract
|
||||
├── common/ # group: SSIM, PSNR, LPIPS
|
||||
├── optical_flow/ # group: gt_optical_flow, synthetic_optical_flow
|
||||
├── vlm/ # group: VideoScore-2
|
||||
├── audio/ # group: CLAP, AudioBox, KL, FAD, WER, DeSync, ImageBind
|
||||
├── videoscore2/ # VideoScore-2 (single metric at group level)
|
||||
├── physics_iq/ # group + sub-metrics
|
||||
└── vbench/ # group: 16 sub-metrics
|
||||
├── __init__.py # sys.path bootstrap + runtime compat shims
|
||||
@@ -546,10 +554,6 @@ scores ± tolerance, and add a calibration test under
|
||||
|
||||
## 10) When not to add a metric
|
||||
|
||||
- **Set-vs-set distribution metrics** (FVD, FID-style) do not fit
|
||||
`BaseMetric.compute(sample)` cleanly; they need a population.
|
||||
Adding them requires a stateful accumulator interface that does
|
||||
not exist yet. Open an issue first.
|
||||
- **Metrics requiring a single-GPU model larger than available
|
||||
memory.** Eval is not the place for tensor-parallel sharding;
|
||||
metrics are expected to fit on one GPU.
|
||||
|
||||
@@ -0,0 +1,312 @@
|
||||
# Performance Benchmarks
|
||||
|
||||
FastVideo's performance benchmark suite measures end-to-end inference latency,
|
||||
throughput, peak GPU memory, and component-level pipeline timings for
|
||||
representative pipeline configurations. It tracks those metrics over time
|
||||
against a rolling baseline stored on the Hugging Face Hub.
|
||||
|
||||
It serves three audiences:
|
||||
|
||||
* **CI** — gates pull requests against a per-GPU static threshold and a
|
||||
rolling-median regression check.
|
||||
* **Maintainers** — surfaces regressions in a Markdown summary on every
|
||||
performance build and a long-form Plotly dashboard.
|
||||
* **Local developers** — lets you run the same benchmark on your own machine,
|
||||
then compare against the historical baseline for the same model and GPU.
|
||||
|
||||
## Quick start (local)
|
||||
|
||||
```bash
|
||||
# Run all benchmarks; writes raw perf_*.json under
|
||||
# fastvideo/tests/performance/results/
|
||||
pytest fastvideo/tests/performance/ -vs
|
||||
|
||||
# Optional: compare against the rolling HF baseline (read-only outside CI).
|
||||
# PERF_REPORTS_DIR defaults to /root/data/perf_reports for Modal/CI, so
|
||||
# override it when running outside the container.
|
||||
PERF_REPORTS_DIR=/tmp/fastvideo_perf_reports \
|
||||
python fastvideo/tests/performance/compare_baseline.py
|
||||
|
||||
# Optional: build the Plotly dashboard locally.
|
||||
PERF_REPORTS_DIR=/tmp/fastvideo_perf_reports \
|
||||
python fastvideo/tests/performance/dashboard.py
|
||||
```
|
||||
|
||||
The pytest run never uploads anything. `compare_baseline.py` only writes to
|
||||
the HF dataset when `TEST_SCOPE=full` *and* `BUILDKITE_BRANCH=main`, so local
|
||||
runs are always read-only. The report directory default is container-oriented;
|
||||
set `PERF_REPORTS_DIR` to a writable local path when generating dashboards or
|
||||
when you want local Markdown/normalized-result artifacts from the comparator.
|
||||
`compare_baseline.py` reads every `perf_*.json` currently present in
|
||||
`fastvideo/tests/performance/results/`; remove stale result files if you only
|
||||
want to compare the latest local run.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
.buildkite/performance-benchmarks/tests/*.json
|
||||
└── per-benchmark configs: model, gen kwargs, per-GPU thresholds
|
||||
|
||||
fastvideo/tests/performance/
|
||||
├── test_inference_performance.py
|
||||
│ └── pytest test that runs each config, writes perf_*.json
|
||||
│ with latency, memory, throughput, and component timings
|
||||
├── compare_baseline.py
|
||||
│ └── normalizes raw results, compares against HF rolling baseline,
|
||||
│ writes Markdown summary + (optionally) uploads new records
|
||||
├── dashboard.py
|
||||
│ └── builds time-series Plotly HTML from HF history
|
||||
└── hf_store.py # shared HF I/O + DataFrame helpers
|
||||
```
|
||||
|
||||
The HF dataset (`FastVideo/performance-tracking` by default) holds one
|
||||
normalized JSON per `(model_id, gpu_type, run)` tuple. The rolling baseline is
|
||||
the median of the last 5 successful records for that model+GPU.
|
||||
|
||||
## Planned Coverage
|
||||
|
||||
The current rollout tracks a small set of representative inference workloads.
|
||||
Broader coverage is planned for additional models, GPU types, attention
|
||||
backends, workload shapes, and inference recipes. As that coverage lands, the
|
||||
performance tracking system will also add environment-specific considerations
|
||||
so comparisons remain meaningful across hardware, runtime, attention backend,
|
||||
and recipe changes instead of treating all records for a model as equivalent.
|
||||
|
||||
## Metrics
|
||||
|
||||
Each benchmark records six metrics:
|
||||
|
||||
| Metric | Raw key | Normalized key | Direction |
|
||||
|---|---|---|---|
|
||||
| End-to-end generation latency | `avg_generation_time_s` | `latency` | Lower is better |
|
||||
| Video throughput | `throughput_fps` | `throughput` | Higher is better |
|
||||
| Peak GPU memory | `max_peak_memory_mb` | `memory` | Lower is better |
|
||||
| Text encoder time | `text_encoder_time_s` | `text_encoder_time_s` | Lower is better |
|
||||
| DiT denoising time | `dit_time_s` | `dit_time_s` | Lower is better |
|
||||
| VAE decode time | `vae_decode_time_s` | `vae_decode_time_s` | Lower is better |
|
||||
|
||||
`test_inference_performance.py` temporarily sets `FASTVIDEO_STAGE_LOGGING=1`
|
||||
while it runs so pipeline stage execution times are available in
|
||||
`generate_video(...).logging_info`. It maps `TextEncodingStage` to
|
||||
`text_encoder_time_s`, `DenoisingStage` and `DmdDenoisingStage` to
|
||||
`dit_time_s`, and `DecodingStage` to `vae_decode_time_s`. If a pipeline does
|
||||
not report one of those stages, that component metric is stored as `null` and
|
||||
is skipped by the static threshold and rolling baseline checks.
|
||||
|
||||
## The two gates
|
||||
|
||||
There are **two independent regression gates** — they protect against
|
||||
different failure modes and are not redundant.
|
||||
|
||||
### Static thresholds (per-GPU)
|
||||
|
||||
Defined in `.buildkite/performance-benchmarks/tests/<benchmark>.json` under
|
||||
`thresholds`. Example:
|
||||
|
||||
```json
|
||||
"thresholds": {
|
||||
"L40S": {
|
||||
"max_generation_time_s": 34.0,
|
||||
"max_peak_memory_mb": 11000.0,
|
||||
"max_text_encoder_time_s": 5.0,
|
||||
"max_dit_time_s": 10.0,
|
||||
"max_vae_decode_time_s": 10.0
|
||||
},
|
||||
"default": { "max_generation_time_s": 120.0, "max_peak_memory_mb": 30000.0 }
|
||||
}
|
||||
```
|
||||
|
||||
Selection: `_get_thresholds(cfg)` matches the current GPU name (substring
|
||||
match) against keys; falls back to `default` if no GPU matches.
|
||||
|
||||
`max_generation_time_s` and `max_peak_memory_mb` are required for every
|
||||
selected threshold block. Component limits are optional: if
|
||||
`max_text_encoder_time_s`, `max_dit_time_s`, or `max_vae_decode_time_s` is
|
||||
absent, the pytest static-threshold gate skips that component.
|
||||
|
||||
These are **fail-safes** — they catch order-of-magnitude regressions,
|
||||
unrealistic memory growth, and optionally large component-specific slowdowns
|
||||
even when the rolling baseline is empty. They are hand-set with generous
|
||||
headroom and almost never need touching.
|
||||
|
||||
### Rolling baseline (per `(model_id, gpu_type)`)
|
||||
|
||||
`compare_baseline.py` loads the last 5 successful records for the same
|
||||
`(model_id, gpu_type)` from the HF dataset, computes the median for each
|
||||
available metric, and fails if the current run regresses by more than
|
||||
`PERF_MAX_REGRESSION` (default 5%). For latency, memory, and component times,
|
||||
higher values are regressions. For throughput, lower values are regressions.
|
||||
|
||||
This is the **drift detector** — it catches sub-threshold regressions that
|
||||
slowly add up. It only persists new records when running the full suite on
|
||||
`main`. Local and pull-request runs can compare against the HF baseline, but
|
||||
they do not update it.
|
||||
|
||||
When the baseline shifts for a legitimate reason (torch upgrade, kernel
|
||||
change, etc.) and CI starts failing, use the
|
||||
[`reseed-performance-baseline`](https://github.com/hao-ai-lab/FastVideo/blob/main/.agents/skills/reseed-performance-baseline/SKILL.md)
|
||||
agent skill to advance the rolling median.
|
||||
|
||||
## Schemas
|
||||
|
||||
### Raw record (`results/perf_*.json`)
|
||||
|
||||
Written by `test_inference_performance.py`. One file per benchmark run.
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"benchmark_id": "wan-t2v-1.3b-2gpu",
|
||||
"model_short_name": "Wan2.1-T2V-1.3B-Diffusers",
|
||||
"device": "NVIDIA L40S",
|
||||
"num_gpus": 2,
|
||||
"num_warmup_runs": 1,
|
||||
"num_measurement_runs": 3,
|
||||
"avg_generation_time_s": 28.4,
|
||||
"individual_times_s": [28.5, 28.3, 28.4],
|
||||
"throughput_fps": 1.58,
|
||||
"max_peak_memory_mb": 10840.0,
|
||||
"individual_peak_memories_mb": [10840.0, 10822.0, 10833.0],
|
||||
"thresholds": {
|
||||
"max_generation_time_s": 34.0,
|
||||
"max_peak_memory_mb": 11000.0,
|
||||
"max_text_encoder_time_s": 5.0,
|
||||
"max_dit_time_s": 10.0,
|
||||
"max_vae_decode_time_s": 10.0
|
||||
},
|
||||
"commit": "<full sha>",
|
||||
"pr_number": "1234",
|
||||
"timestamp": "2026-05-08T22:00:00+00:00",
|
||||
"text_encoder_time_s": 2.141,
|
||||
"dit_time_s": 8.437,
|
||||
"vae_decode_time_s": 3.208
|
||||
}
|
||||
```
|
||||
|
||||
### Normalized record (HF dataset, also dumped as `normalized_perf_*.json`)
|
||||
|
||||
Written by `compare_baseline.py:_normalize_record`. One file per benchmark
|
||||
result, used as the rolling-baseline source of truth.
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"model_id": "wan-t2v-1.3b-2gpu",
|
||||
"timestamp": "2026-05-08T22:00:00+00:00",
|
||||
"commit_sha": "<full sha>",
|
||||
"gpu_type": "NVIDIA L40S",
|
||||
"latency": 28.4,
|
||||
"throughput": 1.58,
|
||||
"memory": 10840.0,
|
||||
"text_encoder_time_s": 2.141,
|
||||
"dit_time_s": 8.437,
|
||||
"vae_decode_time_s": 3.208,
|
||||
"success": true
|
||||
}
|
||||
```
|
||||
|
||||
### Compatibility with legacy records
|
||||
|
||||
Older records in the HF dataset may not have component timing fields. The
|
||||
comparator ignores missing or `null` metrics when computing a median, and the
|
||||
dashboard lists skipped plots for metric series that have no non-null values.
|
||||
|
||||
## Environment variable reference
|
||||
|
||||
| Variable | Default | Used by | Purpose |
|
||||
|---|---|---|---|
|
||||
| `PERF_MAX_REGRESSION` | `0.05` | `compare_baseline.py` | Per-metric regression fraction that fails the build. |
|
||||
| `PERFORMANCE_TRACKING_ROOT` | `/tmp/perf-tracking` | `compare_baseline.py`, `dashboard.py` | Local directory the HF dataset is synced to. |
|
||||
| `PERF_REPORTS_DIR` | `/root/data/perf_reports` | `compare_baseline.py`, `dashboard.py` | Where the Markdown summary and Plotly HTML get written for Buildkite to pick up. |
|
||||
| `HF_REPO_ID` | `FastVideo/performance-tracking` | `hf_store.py` | HF dataset repo holding rolling-baseline records. |
|
||||
| `HF_API_KEY` | unset | `hf_store.py` | Required for upload (main-branch full-suite only); reads work without it. |
|
||||
| `TEST_SCOPE` | unset | `compare_baseline.py` | Set to `full` together with `BUILDKITE_BRANCH=main` to enable HF persistence. |
|
||||
| `BUILDKITE_BRANCH`, `BUILDKITE_COMMIT`, `BUILDKITE_PULL_REQUEST` | unset | `compare_baseline.py`, `test_inference_performance.py` | CI metadata stamped into records. |
|
||||
| `DASHBOARD_DAYS` | `30` | `dashboard.py` | Lookback window for the Plotly trend pages. |
|
||||
| `PERFORMANCE_TRACKING_SYNC_REUSE_TTL_SECONDS` | `3600` | `hf_store.py` | Freshness window for reusing an existing HF sync when requested by dashboard consumers. |
|
||||
| `FASTVIDEO_STAGE_LOGGING` | set by the pytest test | `test_inference_performance.py` | Enables pipeline stage timing capture for component metrics during benchmark runs. |
|
||||
|
||||
## CI integration
|
||||
|
||||
The performance step can run on demand with `/test performance` and as part of
|
||||
the Full Suite (see [CI Architecture](ci_architecture.md)). The Modal entry
|
||||
point is `fastvideo/tests/modal/pr_test.py:run_performance_tests` and the
|
||||
Buildkite artifact upload is in
|
||||
`.buildkite/scripts/pr_test.sh:upload_performance_artifacts`.
|
||||
|
||||
Each performance build runs pytest first. If that fixed-threshold phase fails,
|
||||
`compare_baseline.py` is skipped, so Markdown summaries and normalized JSON
|
||||
artifacts are not emitted. The dashboard still runs best-effort for
|
||||
observability. When pytest passes, the rolling-baseline phase emits:
|
||||
|
||||
* **Markdown summary** — appended to `$GITHUB_STEP_SUMMARY` when that variable
|
||||
is set, and written as `perf_<sha>_<ts>.md` for Buildkite upload. Contains a
|
||||
per-benchmark row with current vs. baseline values for latency, throughput,
|
||||
memory, text encoder time, DiT time, and VAE decode time.
|
||||
* **Plotly dashboard** — `dashboard_<sha>_<ts>.html` showing time-series for
|
||||
each metric grouped by `(model_id, gpu_type)`.
|
||||
* **Normalized records** — `normalized_perf_*.json`, one per benchmark.
|
||||
Useful as input to the
|
||||
[`reseed-performance-baseline`](https://github.com/hao-ai-lab/FastVideo/blob/main/.agents/skills/reseed-performance-baseline/SKILL.md)
|
||||
skill.
|
||||
|
||||
## Adding a new benchmark
|
||||
|
||||
1. Drop a new JSON config into
|
||||
`.buildkite/performance-benchmarks/tests/<name>.json`. Required keys:
|
||||
|
||||
```json
|
||||
{
|
||||
"benchmark_id": "<unique-id>",
|
||||
"model": { "model_path": "...", "model_short_name": "..." },
|
||||
"init_kwargs": { "num_gpus": 1, ... },
|
||||
"generation_kwargs": { "num_frames": 45, ... },
|
||||
"test_prompts": ["..."],
|
||||
"run_config": { "required_gpus": 1,
|
||||
"num_warmup_runs": 1, "num_measurement_runs": 3 },
|
||||
"thresholds": {
|
||||
"L40S": {
|
||||
"max_generation_time_s": 34.0,
|
||||
"max_peak_memory_mb": 11000.0,
|
||||
"max_text_encoder_time_s": 5.0,
|
||||
"max_dit_time_s": 10.0,
|
||||
"max_vae_decode_time_s": 10.0
|
||||
},
|
||||
"default": { "max_generation_time_s": 120.0, "max_peak_memory_mb": 30000.0 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
2. The pytest test auto-discovers all configs — no test code needed. CI
|
||||
picks it up on the next `/test performance` run.
|
||||
|
||||
3. The first persisted main-branch run with no HF history initializes the
|
||||
baseline (passes automatically). Subsequent runs compare against it. Local
|
||||
and pull-request runs with no HF history also pass, but they do not seed the
|
||||
shared baseline.
|
||||
|
||||
4. If the benchmark targets a GPU not currently in `thresholds`, either add
|
||||
that GPU as a key or rely on the `default` block. Note that `default` is
|
||||
intended for slower fallback GPUs, so its values should be relaxed
|
||||
relative to the fastest entry.
|
||||
|
||||
5. Add component thresholds only when the stage timing is stable enough to be
|
||||
a useful fixed gate. The rolling baseline will still track component times
|
||||
when static component thresholds are omitted.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**"No baseline for ... Initializing"** — first run for this `(model_id,
|
||||
gpu_type)`. Run will pass and (if persisting) seed the first record.
|
||||
|
||||
**Persistent failure right after a torch / kernel / image upgrade** —
|
||||
genuine regression *or* baseline drift. Compare the failing normalized record
|
||||
with recent successful records in the HF dataset. If the shift is expected and
|
||||
reviewed, use the `reseed-performance-baseline` skill.
|
||||
|
||||
**Dashboard reports skipped metric plots** — the loaded records do not have
|
||||
non-null values for that metric. This is expected for older records or for
|
||||
pipelines that did not report a mapped component stage.
|
||||
|
||||
**Component timing is `null`** — the generated result did not include a mapped
|
||||
stage in `logging_info.stages`. Check that the pipeline emits stage logging
|
||||
and that the stage name is listed in `STAGE_METRIC_MAP` in
|
||||
`test_inference_performance.py`.
|
||||
@@ -51,3 +51,11 @@ Traces can be visualized using <https://ui.perfetto.dev/>.
|
||||
- Keep the profiled step count small; traces can be large and slow down job shutdown while the profiler flushes data.
|
||||
- After profiling, clean up trace directories to avoid filling disk storage.
|
||||
- When adding new regions, register them in `fastvideo.profiler` and wrap the corresponding code block with `with self.profiler_controller.region("your_region"):` or the `@profile_region` decorator.
|
||||
|
||||
## Related: Activation Trace Mode
|
||||
|
||||
For per-layer **numerical-divergence** debugging (parity bring-up against an
|
||||
upstream reference), use the env-gated [Activation Trace Mode](activation_trace.md)
|
||||
instead of the torch profiler. Activation trace dumps per-tensor stats to
|
||||
JSONL for offline `diff`-ing; the torch profiler captures kernel timing.
|
||||
They solve different problems and can run together if needed.
|
||||
|
||||
@@ -6,10 +6,12 @@ This guide explains how to add and run tests in FastVideo. The testing suite is
|
||||
|
||||
* **Unit Tests**: Located in `fastvideo/tests/api`, `fastvideo/tests/dataset`, `fastvideo/tests/entrypoints`, `fastvideo/tests/workflow`, and the CPU-only subset of `fastvideo/tests/train` (callbacks, utils). These test individual functions and classes.
|
||||
* **Component Tests**: Located in `fastvideo/tests/encoders`, `fastvideo/tests/transformers`, and `fastvideo/tests/vaes`. These verify the loading and basic functionality of model components.
|
||||
* **Train Framework Tests** (GPU): Located in `fastvideo/tests/train/models`. Cover model loading + forward smoke for the new `fastvideo/train/` framework. Triggered via `/test train-framework` or as part of the Full Suite.
|
||||
* **Train Framework Tests** (GPU): Located in `fastvideo/tests/train/models` (model loading + forward smoke) and `fastvideo/tests/train/methods` (per-method single training step). Cover the new `fastvideo/train/` framework end-to-end on real checkpoints with tiny synthetic batches. Triggered via `/test train-framework` or as part of the Full Suite.
|
||||
* **SSIM Tests**: Located in `fastvideo/tests/ssim`. These are regression tests that compare generated videos against reference videos using the Structural Similarity Index Measure (SSIM) to detect quality degradation.
|
||||
* **Training Tests**: Located in `fastvideo/tests/training`. These validate training loops, loss calculations, and specific training techniques like LoRA, Distillation, and VSA.
|
||||
* **Inference Tests**: Located in `fastvideo/tests/inference`. These test specialized inference pipelines and optimizations (e.g., VSA, V-MoBA).
|
||||
* **Performance Tests**: Located in `fastvideo/tests/performance`. End-to-end latency, throughput, and peak-memory benchmarks gated against per-GPU static thresholds and a rolling Hugging Face baseline. See [Performance Benchmarks](performance_benchmarks.md) for the full workflow, schema and future environment-specific considerations.
|
||||
* **DreamVerse App Tests**: Located in `apps/dreamverse`. These validate the DreamVerse app frontend and related app code.
|
||||
|
||||
For now, we will focus on **SSIM Tests**.
|
||||
|
||||
@@ -212,6 +214,7 @@ from a PR comment. The workflow reacts with a 🚀 emoji to confirm the command
|
||||
/test transformer # Transformer / DiT tests (Fastcheck)
|
||||
/test kernel # CUDA kernel tests (Fastcheck)
|
||||
/test unit # Unit tests (Fastcheck)
|
||||
/test dreamverse # DreamVerse app tests (Fastcheck)
|
||||
/test full # Entire Full Suite
|
||||
/test fastcheck # Entire Fastcheck suite
|
||||
```
|
||||
|
||||
@@ -29,24 +29,46 @@ surfaces:
|
||||
vae_cpu_offload: generator.engine.offload.vae
|
||||
pin_cpu_memory: generator.engine.offload.pin_cpu_memory
|
||||
enable_torch_compile: generator.engine.compile.enabled
|
||||
enable_torch_compile_text_encoder: generator.engine.compile.text_encoder_enabled
|
||||
enable_torch_compile_vae: generator.engine.compile.vae_enabled
|
||||
enable_torch_compile_audio_vae: generator.engine.compile.audio_vae_enabled
|
||||
torch_compile_kwargs: generator.engine.compile.backend,fullgraph,mode,dynamic,extras
|
||||
torch_compile_kwargs_dit: generator.engine.compile.dit_kwargs
|
||||
torch_compile_kwargs_text_encoder: generator.engine.compile.text_encoder_kwargs
|
||||
torch_compile_kwargs_vae: generator.engine.compile.vae_kwargs
|
||||
torch_compile_kwargs_audio_vae: generator.engine.compile.audio_vae_kwargs
|
||||
transformer_quant: generator.engine.quantization.transformer_quant
|
||||
disable_autocast: generator.engine.disable_autocast
|
||||
enable_stage_verification: generator.engine.enable_stage_verification
|
||||
prompt_txt: request.inputs.prompt_path
|
||||
override_text_encoder_safetensors: generator.pipeline.components.text_encoder_weights
|
||||
override_text_encoder_quant: generator.engine.quantization.text_encoder_quant
|
||||
transformer_quant: generator.engine.quantization.transformer_quant
|
||||
override_transformer_cls_name: generator.pipeline.components.override_transformer_cls_name
|
||||
init_weights_from_safetensors: generator.pipeline.components.transformer_weights
|
||||
init_weights_from_safetensors_2: generator.pipeline.components.transformer_2_weights
|
||||
override_pipeline_cls_name: generator.pipeline.components.override_pipeline_cls_name
|
||||
boundary_ratio: request.sampling.boundary_ratio
|
||||
ltx2_vae_tiling: generator.pipeline.vae_tiling
|
||||
refine_enabled: generator.pipeline.preset_overrides.refine.enabled
|
||||
refine_upsampler_path: generator.pipeline.components.upsampler_weights
|
||||
refine_lora_path: generator.pipeline.components.lora_path
|
||||
refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
|
||||
refine_guidance_scale: request.stage_overrides.refine.guidance_scale
|
||||
refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
|
||||
ltx2_refine_enabled: generator.pipeline.preset_overrides.refine.enabled
|
||||
ltx2_refine_upsampler_path: generator.pipeline.components.upsampler_weights
|
||||
ltx2_refine_lora_path: generator.pipeline.components.lora_path
|
||||
ltx2_refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
|
||||
ltx2_refine_guidance_scale: request.stage_overrides.refine.guidance_scale
|
||||
ltx2_refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
|
||||
preset_owned:
|
||||
ltx2_vae_spatial_tile_size_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_size_in_pixels
|
||||
ltx2_vae_spatial_tile_overlap_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_overlap_in_pixels
|
||||
ltx2_vae_temporal_tile_size_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_size_in_frames
|
||||
ltx2_vae_temporal_tile_overlap_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_overlap_in_frames
|
||||
ltx2_initial_latent_path: request.extensions.ltx2.initial_latent_path
|
||||
ltx2_audio_latent_path: request.extensions.ltx2.audio_latent_path
|
||||
compatibility_only:
|
||||
mode: "Legacy multi-mode FastVideoArgs switch; typed inference config should not expose execution mode."
|
||||
inference_mode: "Legacy boolean mirror of mode; kept only through adapters while FastVideoArgs remains."
|
||||
@@ -56,6 +78,14 @@ surfaces:
|
||||
VSA_sparsity: "Model-specific inference optimization not yet represented in the typed public schema."
|
||||
moba_config_path: "Model-specific MoBA optimization surface not yet represented in the typed public schema."
|
||||
master_port: "Executor/bootstrap compatibility field; not part of the canonical inference schema."
|
||||
refine_transformer_path: "Generic stage-2 refine transformer override; no typed equivalent yet."
|
||||
refine_noise_path: "Generic stage-2 refine noise override; no typed equivalent yet."
|
||||
refine_audio_noise_path: "Generic stage-2 refine audio noise override; no typed equivalent yet."
|
||||
ltx2_refine_transformer_path: "LTX-2 refine transformer carrier; no typed equivalent yet."
|
||||
ltx2_refine_noise_path: "LTX-2 refine noise carrier; no typed equivalent yet."
|
||||
ltx2_refine_audio_noise_path: "LTX-2 refine audio noise carrier; no typed equivalent yet."
|
||||
ltx2_legacy_native_noise_order: "LTX-2 SSIM compatibility knob preserving legacy native latent noise ordering."
|
||||
ltx2_use_distilled_sigmas: "LTX-2 compatibility knob gating use of distilled sigma schedule."
|
||||
private_only:
|
||||
ray_placement_group: "Ray deployment-only field."
|
||||
ray_runtime_env: "Ray deployment-only field."
|
||||
@@ -78,6 +108,7 @@ surfaces:
|
||||
vae_sp: generator.pipeline.preset_overrides.vae_sp
|
||||
dmd_denoising_steps: generator.pipeline.preset_overrides.dmd_denoising_steps
|
||||
ti2v_task: generator.pipeline.preset_overrides.ti2v_task
|
||||
lucy_edit_task: generator.pipeline.preset_overrides.lucy_edit_task
|
||||
boundary_ratio: generator.pipeline.preset_overrides.boundary_ratio
|
||||
compatibility_only:
|
||||
model_path: "Redundant with generator.model_path."
|
||||
@@ -95,9 +126,18 @@ surfaces:
|
||||
text_encoder_configs: "Legacy internal component config object."
|
||||
preprocess_text_funcs: "Internal text preprocessing hooks."
|
||||
postprocess_text_funcs: "Internal text postprocessing hooks."
|
||||
scheduler_step_in_fp32: "Runtime scheduler precision toggle; not part of the public typed inference API."
|
||||
|
||||
pipeline_config_extensions:
|
||||
preset_owned:
|
||||
flux2_text_encoder_type:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
|
||||
text_encoder_out_layers:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
|
||||
conditioning_strategy:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.cosmos.CosmosConfig
|
||||
@@ -231,8 +271,8 @@ surfaces:
|
||||
- fastvideo.configs.pipelines.turbodiffusion.TurboDiffusionT2V_1_3B_Config
|
||||
- fastvideo.configs.pipelines.wan.FastWan2_1_T2V_480P_Config
|
||||
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
|
||||
- fastvideo.configs.pipelines.wan.MatrixGameBaseI2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig
|
||||
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2BaseI2V480PConfig
|
||||
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.SelfForcingWan2_2_T2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.SelfForcingWanT2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.WANV2VConfig
|
||||
@@ -254,8 +294,8 @@ surfaces:
|
||||
- fastvideo.configs.pipelines.turbodiffusion.TurboDiffusionT2V_1_3B_Config
|
||||
- fastvideo.configs.pipelines.wan.FastWan2_1_T2V_480P_Config
|
||||
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
|
||||
- fastvideo.configs.pipelines.wan.MatrixGameBaseI2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig
|
||||
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2BaseI2V480PConfig
|
||||
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.SelfForcingWan2_2_T2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.SelfForcingWanT2V480PConfig
|
||||
- fastvideo.configs.pipelines.wan.WANV2VConfig
|
||||
@@ -303,9 +343,9 @@ surfaces:
|
||||
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
|
||||
- fastvideo.configs.pipelines.wan.Wan2_2_TI2V_5B_Config
|
||||
context_noise:
|
||||
sources: [fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig]
|
||||
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
|
||||
num_frames_per_block:
|
||||
sources: [fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig]
|
||||
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
|
||||
audio_channels:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
|
||||
@@ -330,6 +370,40 @@ surfaces:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
|
||||
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
|
||||
audio_txt_guidance_scale:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
cfg_number:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
cfg_trick_start_frame:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
cfg_trick_value:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
noise_value:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
sr_audio_noise_scale:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
sr_height:
|
||||
sources:
|
||||
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig
|
||||
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR1080pConfig
|
||||
sr_num_inference_steps:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
sr_video_txt_guidance_scale:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
sr_width:
|
||||
sources:
|
||||
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig
|
||||
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR1080pConfig
|
||||
t5_gemma_target_length:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
use_cfg_trick:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
|
||||
video_guidance_high_t_threshold:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
video_guidance_low_t_value:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
video_txt_guidance_scale:
|
||||
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
|
||||
compatibility_only:
|
||||
batch_size: "Gen3C inference-only tuning field pending typed batching design."
|
||||
gradient_checkpointing: "Gen3C inference-only compatibility field pending typed batching design."
|
||||
@@ -338,6 +412,15 @@ surfaces:
|
||||
internal_only:
|
||||
audio_decoder_config: "Legacy internal component config object."
|
||||
audio_decoder_precision: "Precision override pending dedicated component precision design."
|
||||
audio_vae_config: "MagiHuman internal audio VAE component config object."
|
||||
coords_style: "MagiHuman internal data-proxy coordinate convention."
|
||||
frame_receptive_field: "MagiHuman internal data-proxy receptive-field setting."
|
||||
image_conditioning: "MagiHuman preset variant marker for reference-image conditioning."
|
||||
ref_audio_offset: "MagiHuman internal data-proxy audio alignment offset."
|
||||
sr_local_attn_layers: "MagiHuman SR internal sparse-attention layer selection."
|
||||
text_offset: "MagiHuman internal data-proxy text alignment offset."
|
||||
vae_stride: "MagiHuman internal VAE/data-proxy stride setting."
|
||||
z_dim: "MagiHuman internal VAE latent channel setting."
|
||||
vocoder_config: "Legacy internal component config object."
|
||||
vocoder_precision: "Precision override pending dedicated component precision design."
|
||||
|
||||
@@ -404,6 +487,11 @@ surfaces:
|
||||
ltx2_stg_scale_audio: request.extensions.ltx2.stg_scale_audio
|
||||
ltx2_stg_blocks_video: request.extensions.ltx2.stg_blocks_video
|
||||
ltx2_stg_blocks_audio: request.extensions.ltx2.stg_blocks_audio
|
||||
ltx2_images: request.extensions.ltx2.images
|
||||
ltx2_image_crf: request.stage_overrides.refine.image_crf
|
||||
ltx2_conditioning_latent_stage1: request.extensions.ltx2.conditioning_latent_stage1
|
||||
ltx2_conditioning_latent_stage2: request.extensions.ltx2.conditioning_latent_stage2
|
||||
ltx2_video_conditions: request.extensions.ltx2.video_conditions
|
||||
audio_start_in_s: request.extensions.stable_audio.audio_start_in_s
|
||||
audio_end_in_s: request.extensions.stable_audio.audio_end_in_s
|
||||
init_audio: request.extensions.stable_audio.init_audio
|
||||
@@ -413,6 +501,8 @@ surfaces:
|
||||
inpaint_mask: request.extensions.stable_audio.inpaint_mask
|
||||
internal_only:
|
||||
data_type: "Derived from the request shape and not a public input."
|
||||
latents: "Pre-generated diffusion latents supplied by parity/debug harnesses; not a public input."
|
||||
max_sequence_length: "Model-specific text-encoder sequence cap; not part of the public typed inference API."
|
||||
|
||||
sampling_param_extensions: {}
|
||||
|
||||
|
||||
@@ -168,6 +168,11 @@ How this maps to FastVideo:
|
||||
|
||||
- Attention backends live in `fastvideo/attention/` and can be selected via
|
||||
`FASTVIDEO_ATTENTION_BACKEND`.
|
||||
- SageAttention3 is split into two selectable backends:
|
||||
`SAGE_ATTN_THREE` for the regular upstream package and
|
||||
`ATTN_QAT_INFER` for the FastVideoKernel-backed inference variant.
|
||||
- `ATTN_QAT_TRAIN` is a separate FastVideoKernel Triton backend for the QAT attention
|
||||
path.
|
||||
- `LocalAttention` is used for cross-attention and most attention layers.
|
||||
- `DistributedAttention` is used for full-sequence self-attention in the DiT.
|
||||
- Tensor-parallel layers live in `fastvideo/layers/`.
|
||||
|
||||
@@ -0,0 +1,354 @@
|
||||
# Dynamo Native Backend Integration
|
||||
|
||||
FastVideo exposes a stable Python API that the
|
||||
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo) project consumes
|
||||
as a pure-Python import, same tier as `vllm`, `sglang`, `trtllm`.
|
||||
|
||||
**FastVideo hosts no Dynamo code.** The backend subpackage
|
||||
(`components/src/dynamo/fastvideo/`) lives in the Dynamo repo. This doc
|
||||
is the reference integrators copy when standing up that package — it
|
||||
mirrors the structure used by `dynamo/components/src/dynamo/sglang/`
|
||||
and is known to satisfy the (closed) draft
|
||||
[ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
|
||||
pattern.
|
||||
|
||||
## What FastVideo provides
|
||||
|
||||
The public surface Dynamo imports is intentionally small:
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
ContinuationState,
|
||||
GenerationRequest,
|
||||
InputConfig,
|
||||
OutputConfig,
|
||||
SamplingConfig,
|
||||
# Post-PR 7.10:
|
||||
VideoEvent, VideoProgressEvent, VideoPartialEvent, VideoFinalEvent,
|
||||
VideoResult,
|
||||
)
|
||||
```
|
||||
|
||||
| Surface | Availability | Notes |
|
||||
| --- | --- | --- |
|
||||
| `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` | Today | `typed_kwargs` is a stable subset from `GeneratorConfig` — no flat legacy LTX-2 kwargs (guaranteed after PR 6) |
|
||||
| `VideoGenerator.generate(request: GenerationRequest) -> GenerationResult` | Today | Aggregated; Dynamo wraps in `asyncio.to_thread` under `asyncio.Lock` |
|
||||
| `VideoGenerator.generate_async(request) -> AsyncGenerator[VideoEvent, None]` | **PR 7.10** | Canonical execution substrate; sync wrapper reroutes through this |
|
||||
| `VideoGenerator.default_health_check_request() -> GenerationRequest` | **PR 7.10** | 256x256 / 8 frames / 1 step; lets Dynamo build its health payload without knowing any FastVideo internals |
|
||||
| `fastvideo.api.GenerationRequest` / `SamplingConfig` / `InputConfig` | Today | Stable public dataclasses |
|
||||
| `fastvideo.api.ContinuationState` | Today (PR 7) | JSON-safe envelope; kind-versioned payloads |
|
||||
| `fastvideo.api.VideoResult` | Today | `frames`, `video_path`, `state`, `metadata` |
|
||||
| `config_to_dict(cfg)` | Today | Used by Dynamo's `dump_config(path, config)` |
|
||||
|
||||
## Backend package layout
|
||||
|
||||
Modeled on `components/src/dynamo/sglang/`:
|
||||
|
||||
```
|
||||
components/src/dynamo/fastvideo/
|
||||
├── __init__.py
|
||||
├── __main__.py # Entry: python -m dynamo.fastvideo
|
||||
├── main.py # worker() dispatch — mirrors sglang/main.py
|
||||
├── args.py # FastVideoArgGroup — CLI → GeneratorConfig
|
||||
├── backend_args.py # Dynamo runtime flags (namespace, fs_url, ...)
|
||||
├── init_video_generation.py # init_video_generation(runtime, config)
|
||||
├── register.py # register_video_generation_model() for Dynamo
|
||||
├── backend.py # VideoGenerationWorkerHandler
|
||||
├── health_check.py # FastVideoHealthCheckPayload
|
||||
├── protocol.py # NvCreateVideoRequest ↔ GenerationRequest adapter
|
||||
├── request_handlers/
|
||||
│ └── video_generation/
|
||||
│ └── video_generation_handler.py # async generate(req, ctx)
|
||||
├── README.md
|
||||
└── CLAUDE.md # per-backend guidance
|
||||
```
|
||||
|
||||
None of these files live in FastVideo.
|
||||
|
||||
## Request/response mapping
|
||||
|
||||
Dynamo's `NvCreateVideoRequest` / `VideoNvExt` / `NvVideosResponse` map
|
||||
one-to-one onto FastVideo's typed schema:
|
||||
|
||||
```
|
||||
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
|
||||
prompt -> request.prompt
|
||||
size="WxH" -> request.sampling.width, height
|
||||
seconds -> seconds * nvext.fps -> request.sampling.num_frames
|
||||
input_reference -> request.inputs.image_path / video_path
|
||||
nvext.fps -> request.sampling.fps
|
||||
nvext.num_frames -> request.sampling.num_frames (overrides seconds*fps)
|
||||
nvext.num_inference_steps -> request.sampling.num_inference_steps
|
||||
nvext.guidance_scale -> request.sampling.guidance_scale
|
||||
nvext.seed -> request.sampling.seed
|
||||
nvext.negative_prompt -> request.sampling.negative_prompt
|
||||
nvext.continuation_state -> request.state (opaque ContinuationState)
|
||||
response_format -> (handled by adapter at output)
|
||||
|
||||
VideoFinalEvent -> NvVideosResponse
|
||||
video_bytes -> data[0].b64_json (if response_format=b64_json)
|
||||
uploaded URL -> data[0].url (if response_format=url)
|
||||
metadata.inference_time_s -> inference_time_s
|
||||
continuation_state -> nvext.continuation_state (reserved for disagg)
|
||||
```
|
||||
|
||||
## Example: aggregated handler (sync wrap)
|
||||
|
||||
Satisfies the PR #7544 shape; works today against
|
||||
`VideoGenerator.generate`, upgrades cleanly to `generate_async` after
|
||||
PR 7.10.
|
||||
|
||||
```python
|
||||
# components/src/dynamo/fastvideo/request_handlers/video_generation/
|
||||
# video_generation_handler.py
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import base64
|
||||
import time
|
||||
from typing import Any, AsyncGenerator
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
|
||||
|
||||
|
||||
class VideoGenerationWorkerHandler:
|
||||
def __init__(self, generator: VideoGenerator, config, fs=None):
|
||||
self.generator = generator
|
||||
self.config = config
|
||||
self.fs = fs
|
||||
self._lock = asyncio.Lock() # aggregated = one-in-flight
|
||||
|
||||
async def generate(
|
||||
self,
|
||||
request: dict[str, Any],
|
||||
context,
|
||||
) -> AsyncGenerator[dict[str, Any], None]:
|
||||
req = _to_fastvideo_request(request)
|
||||
t0 = time.perf_counter()
|
||||
async with self._lock:
|
||||
result = await asyncio.to_thread(self.generator.generate, req)
|
||||
elapsed = time.perf_counter() - t0
|
||||
|
||||
video_bytes = _materialize(result, self.fs, request.get("response_format"))
|
||||
yield {
|
||||
"data": [video_bytes],
|
||||
"inference_time_s": elapsed,
|
||||
"model": request.get("model"),
|
||||
}
|
||||
|
||||
|
||||
def _to_fastvideo_request(request: dict[str, Any]) -> GenerationRequest:
|
||||
nvext = request.get("nvext") or {}
|
||||
fps = nvext.get("fps", 24)
|
||||
num_frames = nvext.get("num_frames") or (request.get("seconds") or 4) * fps
|
||||
width, height = _parse_size(request.get("size"))
|
||||
|
||||
return GenerationRequest(
|
||||
prompt=request["prompt"],
|
||||
negative_prompt=nvext.get("negative_prompt"),
|
||||
inputs=InputConfig(
|
||||
image_path=request.get("input_reference"),
|
||||
),
|
||||
sampling=SamplingConfig(
|
||||
width=width, height=height,
|
||||
num_frames=num_frames, fps=fps,
|
||||
num_inference_steps=nvext.get("num_inference_steps", 50),
|
||||
guidance_scale=nvext.get("guidance_scale", 1.0),
|
||||
seed=nvext.get("seed", 1024),
|
||||
),
|
||||
output=OutputConfig(save_video=False, return_frames=False),
|
||||
state=nvext.get("continuation_state"), # public ContinuationState
|
||||
)
|
||||
```
|
||||
|
||||
`_parse_size` and `_materialize` are small adapter helpers owned by the
|
||||
Dynamo backend package; they never appear in FastVideo.
|
||||
|
||||
## Example: streaming handler (post-PR 7.10)
|
||||
|
||||
```python
|
||||
async def generate(self, request, context):
|
||||
req = _to_fastvideo_request(request)
|
||||
async for event in self.generator.generate_async(req):
|
||||
if event.__class__.__name__ == "VideoProgressEvent":
|
||||
yield {"status": "generating", "progress": event.step / event.total_steps}
|
||||
elif event.__class__.__name__ == "VideoFinalEvent":
|
||||
yield {
|
||||
"data": [{"b64_json": base64.b64encode(event.video_bytes).decode()}],
|
||||
"inference_time_s": event.metadata.get("inference_time_s"),
|
||||
"nvext": {"continuation_state": _serialize_state(event.continuation_state)},
|
||||
}
|
||||
```
|
||||
|
||||
Aggregated and streaming differ only in which events the handler
|
||||
forwards; both share one `generate_async` substrate.
|
||||
|
||||
## Example: health check
|
||||
|
||||
```python
|
||||
# components/src/dynamo/fastvideo/health_check.py
|
||||
from dynamo.health_check import HealthCheckPayload
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
|
||||
class FastVideoHealthCheckPayload(HealthCheckPayload):
|
||||
def __init__(self, generator: VideoGenerator) -> None:
|
||||
# Post-PR 7.10: generator.default_health_check_request() returns a
|
||||
# typed GenerationRequest; dump it into the same dict shape that
|
||||
# FastVideo's adapter accepts.
|
||||
req = generator.default_health_check_request()
|
||||
self.default_payload = {
|
||||
"prompt": req.prompt or "test",
|
||||
"size": f"{req.sampling.width}x{req.sampling.height}",
|
||||
"response_format": "b64_json",
|
||||
"nvext": {
|
||||
"fps": req.sampling.fps,
|
||||
"num_frames": req.sampling.num_frames,
|
||||
"num_inference_steps": req.sampling.num_inference_steps,
|
||||
"guidance_scale": req.sampling.guidance_scale,
|
||||
},
|
||||
}
|
||||
super().__init__()
|
||||
```
|
||||
|
||||
Fallback (pre-PR 7.10) — hardcoded 256×256 / 8 frames / 1 step, matching
|
||||
[`VideoGenerationHealthCheckPayload`](https://github.com/ai-dynamo/dynamo/blob/main/components/src/dynamo/sglang/health_check.py#L198-L226).
|
||||
|
||||
## Init function sketch
|
||||
|
||||
```python
|
||||
# components/src/dynamo/fastvideo/init_video_generation.py
|
||||
async def init_video_generation(runtime, config, shutdown_endpoints):
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import config_to_dict
|
||||
|
||||
server_args, dynamo_args = config.server_args, config.dynamo_args
|
||||
generator = VideoGenerator.from_pretrained(**config.fastvideo_kwargs())
|
||||
|
||||
dump_config(dynamo_args.dump_config_to, config)
|
||||
|
||||
endpoint = runtime.endpoint(
|
||||
f"{dynamo_args.namespace}.{dynamo_args.component}.{dynamo_args.endpoint}"
|
||||
)
|
||||
shutdown_endpoints[:] = [endpoint]
|
||||
|
||||
handler = VideoGenerationWorkerHandler(
|
||||
generator, config, fs=get_fs(dynamo_args.media_output_fs_url)
|
||||
)
|
||||
payload = FastVideoHealthCheckPayload(generator).to_dict()
|
||||
|
||||
await asyncio.gather(
|
||||
endpoint.serve_endpoint(
|
||||
handler.generate,
|
||||
graceful_shutdown=True,
|
||||
health_check_payload=payload,
|
||||
),
|
||||
register_video_generation_model(
|
||||
generator, endpoint, server_args,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Args adapter
|
||||
|
||||
`FastVideoArgGroup` (Dynamo-side) converts CLI flags into a typed
|
||||
`GeneratorConfig` — **never** into legacy flat kwargs. Because PR 6
|
||||
added typed homes for every kwarg the internal `gpu_pool.py` used,
|
||||
this adapter can build the config purely from the public typed schema:
|
||||
|
||||
```python
|
||||
def build_generator_config(args) -> "GeneratorConfig":
|
||||
from fastvideo.api import (
|
||||
CompileConfig, ComponentConfig, EngineConfig, GeneratorConfig,
|
||||
OffloadConfig, ParallelismConfig, PipelineSelection,
|
||||
)
|
||||
return GeneratorConfig(
|
||||
model_path=args.model_path,
|
||||
engine=EngineConfig(
|
||||
num_gpus=args.num_gpus,
|
||||
parallelism=ParallelismConfig(tp_size=args.tp_size, sp_size=args.sp_size),
|
||||
offload=OffloadConfig(dit=args.dit_offload, text_encoder=args.te_offload),
|
||||
compile=CompileConfig(enabled=args.compile, mode=args.compile_mode),
|
||||
),
|
||||
pipeline=PipelineSelection(
|
||||
workload_type=args.workload or "t2v",
|
||||
preset=args.preset, # e.g. "ltx2_two_stage"
|
||||
components=ComponentConfig(
|
||||
upsampler_weights=args.refine_upsampler,
|
||||
lora_path=args.refine_lora,
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Registration
|
||||
|
||||
Dynamo's Rust side skips HuggingFace `config.json` downloads for
|
||||
`ModelType::Videos`, same fast path used by image diffusion. The
|
||||
Python-side registration:
|
||||
|
||||
```python
|
||||
# components/src/dynamo/fastvideo/register.py
|
||||
from dynamo.llm import ModelDeploymentCard, ModelType, register_model
|
||||
|
||||
|
||||
async def register_video_generation_model(generator, endpoint, server_args):
|
||||
mdc = ModelDeploymentCard.with_name_only(server_args.model_name or server_args.model_path)
|
||||
await register_model(endpoint, mdc, ModelType.Videos, readiness_gate=asyncio.Event())
|
||||
```
|
||||
|
||||
## Contract guarantees
|
||||
|
||||
These guardrails let the Dynamo backend be written once and not
|
||||
re-chase FastVideo drift:
|
||||
|
||||
1. `GenerationRequest` field paths are stable across PR 6 onward. Any
|
||||
breaking rename triggers a major bump and appears in
|
||||
[`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml).
|
||||
2. `ContinuationState.payload` is JSON-serializable or references
|
||||
opaque blob ids. Dynamo can round-trip it through RPC without
|
||||
special-casing torch tensors.
|
||||
3. `VideoGenerator.from_pretrained` accepts a typed `GeneratorConfig`;
|
||||
legacy flat kwargs are compatibility-only and deprecate in PR 13.
|
||||
4. `generate_async` (PR 7.10+) emits events in order
|
||||
`Progress* → Partial* → Final`; the final event always has exactly
|
||||
one occurrence per request.
|
||||
5. `default_health_check_request()` (PR 7.10+) returns a request that
|
||||
passes `parse_config` and produces a non-zero-latency but bounded
|
||||
workload (256×256 / 8 frames / 1 step).
|
||||
|
||||
FastVideo's contract tests (`fastvideo/tests/contract/`) assert these
|
||||
with mocked Dynamo-style handlers that import only the public surface.
|
||||
If a change to FastVideo breaks the adapter pattern, those tests fail
|
||||
at FastVideo's CI — before the Dynamo-side integration even knows.
|
||||
|
||||
## What the Dynamo adapter MUST NOT import
|
||||
|
||||
* Anything under `fastvideo.pipelines.*` directly (pipelines are
|
||||
internal; presets identify them by name on
|
||||
`PipelineSelection.preset`).
|
||||
* `fastvideo.fastvideo_args.FastVideoArgs` (legacy compat type).
|
||||
* `fastvideo.api.compat.*` private helpers
|
||||
(`_validate_continuation_state` etc.) — the public boundary is
|
||||
`VideoGenerator` + `fastvideo.api`.
|
||||
* Any flat legacy LTX-2 kwarg (`ltx2_refine_upsampler_path`,
|
||||
`torch_compile_kwargs`, etc.) — all have typed homes in
|
||||
`GeneratorConfig`.
|
||||
|
||||
## Future: disaggregated prefill/decode
|
||||
|
||||
PR 7's continuation state was designed to survive RPC transport, so a
|
||||
future Dynamo split where prefill yields state and decode hydrates it
|
||||
is expressible without changing the contract. The streaming server
|
||||
(PR 7.6) already uses the `SessionStore` pattern; Dynamo's disagg
|
||||
could wire a distributed `SessionStore` backend by the same interface.
|
||||
|
||||
## See also
|
||||
|
||||
* [OpenAI HTTP contract](openai.md)
|
||||
* [Streaming WebSocket protocol](streaming.md)
|
||||
* Draft PR reference: [ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
|
||||
* Dynamo SGLang backend (template this doc is modeled on):
|
||||
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo/tree/main/components/src/dynamo/sglang)
|
||||
@@ -0,0 +1,41 @@
|
||||
# Server Contracts
|
||||
|
||||
FastVideo's typed public API (`fastvideo.api`) is consumed by three
|
||||
server-class integrations that must share one execution substrate so we
|
||||
don't grow three near-duplicate progress loops:
|
||||
|
||||
| Consumer | Transport | Request shape | State model |
|
||||
| --- | --- | --- | --- |
|
||||
| [Stateless OpenAI](openai.md) | HTTP POST `/v1/videos` | `VideoGenerationsRequest` → `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; optional `ContinuationState` round-trip |
|
||||
| [Streaming WebSocket](streaming.md) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment | Server-held `SessionStore`, snapshot-on-demand |
|
||||
| [Dynamo native backend](dynamo.md) | Dynamo RPC | `NvCreateVideoRequest` → adapter → `GenerationRequest` | Aggregated today; disaggregated via `ContinuationState` later |
|
||||
|
||||
All three consume the same underlying surface:
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
ContinuationState,
|
||||
GenerationRequest,
|
||||
InputConfig,
|
||||
OutputConfig,
|
||||
SamplingConfig,
|
||||
ServeConfig,
|
||||
)
|
||||
|
||||
# Sync today, async after PR 7.10 lands VideoGenerator.generate_async.
|
||||
result = generator.generate(request)
|
||||
```
|
||||
|
||||
These docs lock down the request/response shapes so drift between
|
||||
FastVideo, the internal UI, and Dynamo can be caught at review time.
|
||||
PR 8 does not ship runtime code; it ships the contract reference and
|
||||
the contract tests that guard it.
|
||||
|
||||
## Related
|
||||
|
||||
- [API refactor design](../overview.md)
|
||||
- Parity inventory: [`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml)
|
||||
- [Streaming server upstream plan](../../../.agents/memory/dreamverse-integration/source-archive/streaming-server-upstream-plan.md)
|
||||
- Draft PR (closed) that establishes the Dynamo shape:
|
||||
https://github.com/ai-dynamo/dynamo/pull/7544
|
||||
@@ -0,0 +1,116 @@
|
||||
# OpenAI-compatible HTTP Contract
|
||||
|
||||
The stateless FastVideo HTTP server lives at
|
||||
[`fastvideo/entrypoints/openai/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/entrypoints/openai).
|
||||
Launch: `fastvideo serve --config serve.yaml`.
|
||||
|
||||
## Endpoints
|
||||
|
||||
| Method | Path | Description |
|
||||
| --- | --- | --- |
|
||||
| `POST` | `/v1/videos/generations` | Synchronous video generation |
|
||||
| `GET` | `/v1/videos` | List prior jobs held in the in-memory store |
|
||||
| `GET` | `/v1/videos/{id}` | Job status / result |
|
||||
| `GET` | `/v1/videos/{id}/content` | Download the MP4 once ready |
|
||||
| `POST` | `/v1/images/generations` | Synchronous image generation |
|
||||
| `GET` | `/v1/models` | Enumerate registered models |
|
||||
| `GET` | `/health` | Liveness probe |
|
||||
|
||||
## `VideoGenerationsRequest` shape
|
||||
|
||||
Mirrors the OpenAI `POST /v1/videos/generations` shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"prompt": "a fox running through snow",
|
||||
"size": "1024x1536",
|
||||
"seconds": 5,
|
||||
"fps": 24,
|
||||
"num_frames": 121,
|
||||
"seed": 42,
|
||||
"num_inference_steps": 8,
|
||||
"guidance_scale": 1.0,
|
||||
"negative_prompt": "blurry, low quality",
|
||||
"input_reference": "/path/to/init.png"
|
||||
}
|
||||
```
|
||||
|
||||
SGLang-compatible extensions carried today:
|
||||
`num_inference_steps`, `guidance_scale`, `guidance_scale_2`,
|
||||
`true_cfg_scale`, `negative_prompt`, `enable_teacache`, `output_path`.
|
||||
|
||||
## Merge precedence
|
||||
|
||||
The server builds a `GenerationRequest` each call using three layers,
|
||||
highest first:
|
||||
|
||||
1. **Request body (client-explicit)** — only fields carried in
|
||||
`request.model_fields_set` (Pydantic v2). Unset fields do not count,
|
||||
even if the Pydantic model has a schema default for them.
|
||||
2. **`ServeConfig.default_request` (operator-explicit)** — projected via
|
||||
[`explicit_request_updates()`](../../../fastvideo/api/compat.py);
|
||||
only fields the operator actually wrote into the YAML count as
|
||||
defaults. Every other field inherits the schema default rather than
|
||||
being pinned.
|
||||
3. **Hardcoded fallback** — e.g. `fps = 24`.
|
||||
|
||||
The gate matters: both surfaces carry schema defaults. Without
|
||||
`model_fields_set` / explicit-path tracking, schema defaults would
|
||||
masquerade as intent and silently shadow the other side.
|
||||
|
||||
See [`video_api.py::_build_generation_kwargs`](../../../fastvideo/entrypoints/openai/video_api.py)
|
||||
for the canonical implementation; the per-request assembly lives there,
|
||||
not in pipeline code.
|
||||
|
||||
## Continuation state
|
||||
|
||||
The stateless surface accepts an opaque `ContinuationState` round-trip.
|
||||
Clients that want continuation pass the prior `state` blob back on the
|
||||
next request, and receive a new one on the response when
|
||||
`request.output.return_state = true`.
|
||||
|
||||
Shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"state": {
|
||||
"kind": "ltx2.v1",
|
||||
"payload": { "schema_version": 1, "segment_index": 3, ... }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Payload is always JSON-serializable. Large tensors may live in an
|
||||
opaque blob-store reference the client simply round-trips; see
|
||||
[`LTX2ContinuationState`](../../../fastvideo/pipelines/basic/ltx2/continuation.py).
|
||||
|
||||
Continuation is not yet wired all the way through to
|
||||
`generator.generate_video(...)` — PR 7.6 (GPU pool upstream) is the
|
||||
pipeline-level consumer. PR 7 locked the envelope so this surface is
|
||||
stable ahead of that plumbing.
|
||||
|
||||
## Error codes
|
||||
|
||||
| HTTP | Condition |
|
||||
| --- | --- |
|
||||
| `400 Bad Request` | Parse/validation failure (unknown field, type mismatch, incompatible preset/state) |
|
||||
| `404 Not Found` | `GET /v1/videos/{id}` for an unknown job |
|
||||
| `409 Conflict` | Job id already exists |
|
||||
| `500 Internal Server Error` | Pipeline raised; body mirrors upstream OpenAI error envelope |
|
||||
| `503 Service Unavailable` | No generator loaded, or shutdown in progress |
|
||||
|
||||
Errors include a JSON body with
|
||||
`{"error": {"type": "...", "message": "..."}}` matching the OpenAI
|
||||
Python SDK's expectation.
|
||||
|
||||
## What does not cross this boundary
|
||||
|
||||
* Flat legacy kwargs (`ltx2_refine_enabled`, `torch_compile_kwargs`,
|
||||
etc.) — these are init-time, configured via `ServeConfig.generator`,
|
||||
never per-request.
|
||||
* Private Dreamverse-only fields — those live in a private adapter on
|
||||
the Dreamverse side; the public FastVideo surface never promises
|
||||
backward compatibility for them.
|
||||
* Raw tensor payloads (`ltx2_audio_clean_latent` et al.) — these are
|
||||
derived by the pipeline from `ContinuationState`, never shipped as
|
||||
request fields.
|
||||
@@ -226,7 +226,7 @@ Dataclass carrying all pipeline state between stages. Key field groups:
|
||||
`width_latents`.
|
||||
- **Scheduler**: `timesteps`, `num_inference_steps`, `guidance_scale`,
|
||||
`sigmas`.
|
||||
- **Task-specific**: `mouse_cond`/`keyboard_cond` (MatrixGame), `pose`
|
||||
- **Task-specific**: `mouse_cond`/`keyboard_cond` (Matrix-Game 2.0), `pose`
|
||||
(HYWorld), `camera_states` (GameCraft), `c2ws_plucker_emb`
|
||||
(LingBotWorld).
|
||||
- **Output**: `output: Tensor | None`.
|
||||
@@ -252,7 +252,7 @@ Standard stages (typical execution order):
|
||||
|
||||
Specialized variants: `CausalDenoisingStage`, `LTX2DenoisingStage`,
|
||||
`LongCatDenoisingStage`, `GameCraftDenoisingStage`,
|
||||
`HYWorldDenoisingStage`, `MatrixGameDenoisingStage`,
|
||||
`HYWorldDenoisingStage`, `MatrixGame2CausalDenoisingStage`,
|
||||
`SRDenoisingStage`, `LTX2AudioDecodingStage`, `SD35ConditioningStage`,
|
||||
`LTX2TextEncodingStage`, `LTX2LatentPreparationStage`.
|
||||
|
||||
|
||||
@@ -107,6 +107,8 @@ If you encounter CUDA out of memory errors:
|
||||
(single GPU) or `use_fsdp_inference=True` (multi-GPU)
|
||||
- Try a smaller model or use distilled versions
|
||||
- Use `num_gpus` > 1 if multiple GPUs are available
|
||||
- Try enabling FSDP inference with `use_fsdp_inference=True` (may slow down generation)
|
||||
- Try enabling DiT layerwise offload with `dit_layerwise_offload=True` (now only a few models support this, but may introduce less overhead than FSDP)
|
||||
|
||||
### Slow Generation
|
||||
|
||||
|
||||
+147
-12
@@ -11,6 +11,9 @@ This page describes the various options for speeding up generation times in Fast
|
||||
- [Sliding Tile Attention (Archived)](#sliding-tile-attention-archived)
|
||||
- [Sage Attention](#sage-attention)
|
||||
- [Sage Attention 3](#sage-attention-3)
|
||||
- [Adaptive Guidance (CFG gating)](#adaptive-guidance-cfg-gating)
|
||||
|
||||
- [torch.compile](#torch-compile)
|
||||
|
||||
## Attention Backends
|
||||
|
||||
@@ -69,9 +72,9 @@ python setup.py install
|
||||
|
||||
### FP4 Flash Attention 4 (Blackwell only)
|
||||
|
||||
**`FLASH_ATTN`** with **`FASTVIDEO_NVFP4_FA4=1`**
|
||||
**`FLASH_ATTN`** with **`--nvfp4_fa4`**
|
||||
|
||||
On Blackwell GPUs (B200/B300), you can enable FP4 quantized Q/K attention for up to **1.39x kernel speedup** over BF16 FA4, peaking at **1801 TFLOPS**. This quantizes Q and K to NVFP4 E2M1 with per-block E4M3 scale factors while keeping V in BF16.
|
||||
On Blackwell GPUs (B200/B300), you can enable FP4 quantized Q/K attention for up to **1.31x kernel speedup** over BF16 FA4, peaking at **2018 TFLOPS**. This quantizes Q and K to NVFP4 E2M1 with per-block E4M3 scale factors while keeping V in BF16 or FP8.
|
||||
|
||||
See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040) and [flash-attention-fp4 benchmark results](https://github.com/hao-ai-lab/flash-attention-fp4/blob/fp4/flash_attn/cute/README.md) for details.
|
||||
|
||||
@@ -83,32 +86,30 @@ See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040) and [flash-attention-
|
||||
|
||||
#### Installation
|
||||
|
||||
Install the FP4 flash attention kernel and its dependencies:
|
||||
Install the FP4 flash attention kernel (without upgrading your existing torch):
|
||||
|
||||
```bash
|
||||
pip install "git+ssh://git@github.com/hao-ai-lab/flash-attention-fp4.git@fp4#subdirectory=flash_attn/cute"
|
||||
pip install --no-deps "git+ssh://git@github.com/hao-ai-lab/flash-attention-fp4.git@fp4#subdirectory=flash_attn/cute"
|
||||
pip install "nvidia-cutlass-dsl>=4.4.2" apache-tvm-ffi flashinfer-python
|
||||
```
|
||||
|
||||
This installs the FP4 kernel and all dependencies (nvidia-cutlass-dsl, flashinfer-python, apache-tvm-ffi).
|
||||
The `--no-deps` flag prevents upgrading torch/torchvision. The kernel requires torch >= 2.4 with CUDA 12.8+ support (already present in FastVideo's environment).
|
||||
|
||||
#### Usage
|
||||
|
||||
Enable FP4 attention via environment variables:
|
||||
Enable FP4 attention via the `--nvfp4_fa4` flag:
|
||||
|
||||
```bash
|
||||
FASTVIDEO_NVFP4_FA4=1 CUTE_DSL_ENABLE_TVM_FFI=1 python examples/inference/optimizations/fp4_attn_wan2_1_1_3b.py --nvfp4_fa4
|
||||
python examples/inference/optimizations/fp4_attn_wan2_1_1_3b.py --nvfp4_fa4
|
||||
```
|
||||
|
||||
Or in Python:
|
||||
Or in Python via the `nvfp4_fa4` kwarg (sets env vars automatically):
|
||||
|
||||
```python
|
||||
import os
|
||||
os.environ["FASTVIDEO_NVFP4_FA4"] = "1"
|
||||
os.environ["CUTE_DSL_ENABLE_TVM_FFI"] = "1"
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
gen = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
nvfp4_fa4=True,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # FSDP is incompatible with FP4 pointer path
|
||||
)
|
||||
@@ -177,6 +178,106 @@ These backends are model-specific and require the corresponding kernels and
|
||||
dependencies. Use the support matrix and model examples to confirm compatibility
|
||||
before enabling them.
|
||||
|
||||
<a id="torch-compile"></a>
|
||||
|
||||
## torch.compile
|
||||
|
||||
FastVideo can `torch.compile` the DiT (transformer) for a substantial
|
||||
end-to-end speedup. It is **off by default** and enabled per-run.
|
||||
|
||||
### Enabling
|
||||
|
||||
```python
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
enable_torch_compile=True,
|
||||
)
|
||||
```
|
||||
|
||||
A complete A/B example (eager vs compiled, warmup excluded) is in
|
||||
[`examples/inference/optimizations/torch_compile_example.py`](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/optimizations/torch_compile_example.py).
|
||||
|
||||
`fastvideo generate` is config-file driven; to enable `torch.compile`
|
||||
from the CLI, set the relevant field in your run config and pass it via
|
||||
`fastvideo generate --config run.yaml`. There is no top-level
|
||||
`--enable-torch-compile` flag on the subcommand.
|
||||
|
||||
Only DiT submodules that declare `_compile_conditions` are compiled
|
||||
(most shipped models). The text encoder and VAE are not compiled by this
|
||||
flag.
|
||||
|
||||
### What to expect
|
||||
|
||||
| Config | Effect |
|
||||
|---|---|
|
||||
| Wan2.1-T2V-1.3B, A100-80GB, 480×832×81f, 50 steps | end-to-end **259.7s → 198.1s (−23.7%)**; per-step **4.91 → 3.78 s/it** |
|
||||
|
||||
The speedup is **configuration-dependent** — it varies with model,
|
||||
resolution, step count, and GPU. Treat the number above as one measured
|
||||
data point, not a guarantee; benchmark your own config (recipe below).
|
||||
|
||||
There is a **one-time graph-build cost** on the first generation (tens of
|
||||
seconds to minutes, model-dependent). It amortizes over subsequent
|
||||
generations with the same input shapes. Always exclude the first
|
||||
(warmup) generation when measuring steady-state latency — measuring the
|
||||
warmup is the most common way to wrongly conclude "compile is slower".
|
||||
|
||||
**Numerics.** Inductor's lowering is designed to preserve eager
|
||||
semantics within floating-point tolerance, but per-model equivalence is
|
||||
not asserted by any standing SSIM regression here — the SSIM tests in
|
||||
[`fastvideo/tests/ssim/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/tests/ssim)
|
||||
run with `enable_torch_compile` disabled. If you depend on compile
|
||||
output staying close to eager (or your previous compiled run), run an
|
||||
MS-SSIM gate on *your* config, especially when combining
|
||||
`enable_torch_compile=True` with other numerics-affecting flags
|
||||
(quantized attention backends, FP4, layerwise offload edge cases).
|
||||
|
||||
### Known interactions
|
||||
|
||||
- **Layerwise CPU offload** (`dit_layerwise_offload=True`, the default):
|
||||
the offload hook previously caused an implicit graph break once per
|
||||
transformer layer, fragmenting the compiled region. Addressed in
|
||||
hao-ai-lab/FastVideo#1365 — keep that fix to get a clean compiled
|
||||
region under the default offload path.
|
||||
- **`mode="reduce-overhead"` / CUDA graphs**: not yet supported
|
||||
end-to-end. The attention dispatch is an untraceable custom op and
|
||||
still breaks the graph, which CUDA-graph trees cannot span. Use the
|
||||
default inductor mode (shown above) until that is resolved.
|
||||
|
||||
Extra `torch.compile` options are passed through `torch_compile_kwargs`
|
||||
(a dict), accepted by `VideoGenerator.from_pretrained(...)` and by the
|
||||
CLI as a JSON string via `--torch-compile-kwargs`. Example (currently
|
||||
**not** recommended — see the CUDA-graphs caveat above):
|
||||
|
||||
```python
|
||||
VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
enable_torch_compile=True,
|
||||
torch_compile_kwargs={"mode": "reduce-overhead"}, # may error today
|
||||
)
|
||||
```
|
||||
|
||||
### Benchmarking torch.compile
|
||||
|
||||
Same discipline as attention backends — same prompt, same seed, same
|
||||
config; **discard the first generation** (graph build):
|
||||
|
||||
```python
|
||||
import time
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
gen = VideoGenerator.from_pretrained("your-model-id", enable_torch_compile=True)
|
||||
req = {"prompt": "Your prompt", "sampling": {"seed": 1024},
|
||||
"output": {"save_video": False}}
|
||||
gen.generate(req) # warmup: graph build, discard
|
||||
t0 = time.perf_counter()
|
||||
gen.generate(req) # measured: shapes reused
|
||||
print(f"compiled steady-state: {time.perf_counter() - t0:.2f}s")
|
||||
```
|
||||
|
||||
See `examples/inference/optimizations/torch_compile_example.py` for a
|
||||
baseline-vs-compile A/B with the warmup correctly excluded.
|
||||
|
||||
## Benchmarking different optimizations
|
||||
|
||||
To benchmark backend performance, generate the same prompt with the same seed and compare end-to-end generation times:
|
||||
@@ -198,3 +299,37 @@ for backend in ["TORCH_SDPA", "FLASH_ATTN", "SAGE_ATTN"]:
|
||||
```
|
||||
|
||||
Note: reinstantiate `VideoGenerator` after changing `FASTVIDEO_ATTENTION_BACKEND`.
|
||||
|
||||
## Adaptive Guidance (CFG gating)
|
||||
|
||||
CFG gating accelerates classifier-free guidance by reusing the cached
|
||||
`noise_pred_cond - noise_pred_uncond` delta after a configurable fraction of
|
||||
the denoising schedule, skipping the unconditional model forward for the
|
||||
remaining steps. The technique is the LinearAG variant of Adaptive Guidance
|
||||
(Castillo et al. 2023, [arXiv:2312.12487](https://arxiv.org/abs/2312.12487)).
|
||||
|
||||
### Enabling
|
||||
|
||||
Set the `FASTVIDEO_CFG_GATE_STEP` environment variable to a float in `[0, 1]`:
|
||||
|
||||
| Value | Behavior |
|
||||
|-------|----------|
|
||||
| `1.0` (default) | Disabled — legacy two-pass CFG every step. |
|
||||
| `0.5` | Cache the delta after `len(timesteps) * 0.5` steps; reuse for the rest. |
|
||||
| `0.0` | Cache from the very first step (most aggressive). |
|
||||
|
||||
```bash
|
||||
export FASTVIDEO_CFG_GATE_STEP=0.5
|
||||
```
|
||||
|
||||
### Trade-offs
|
||||
|
||||
- **Memory**: one extra model-output-sized tensor per rank held during the
|
||||
gating window.
|
||||
- **Quality**: VBench-measured quality is preserved within noise on 4 of 5
|
||||
dimensions at `FASTVIDEO_CFG_GATE_STEP=0.5` for Wan T2V 1.3B per the PR's
|
||||
reported numbers (see [#1372](https://github.com/hao-ai-lab/FastVideo/pull/1372)).
|
||||
- **Speed**: ~22% e2e on 4xL40S and ~24% on 1xH100 at the same settings.
|
||||
|
||||
Default behavior is byte-for-byte equivalent to the legacy two-pass CFG path;
|
||||
the feature is fully opt-in.
|
||||
|
||||
@@ -58,6 +58,7 @@ pipeline initialization and sampling.
|
||||
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
| Wan2.2 TI2V 5B | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | 720P | ⭕ | ⭕ | ✅ | ⭕ | ⭕ |
|
||||
| Lucy Edit Dev 5B*** | `decart-ai/Lucy-Edit-Dev` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Wan2.2 T2V A14B | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
|
||||
| Wan2.2 I2V A14B | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
|
||||
| HunyuanVideo | `hunyuanvideo-community/HunyuanVideo` | 720px1280p<br>544px960p | ❌ | ✅ | ✅ | ⭕ | ⭕ |
|
||||
@@ -70,13 +71,17 @@ pipeline initialization and sampling.
|
||||
| TurboWan2.1 T2V 14B | `loayrashid/TurboWan2.1-T2V-14B-Diffusers` | 480P, 720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| TurboWan2.2 I2V A14B | `loayrashid/TurboWan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| LongCat T2V 13.6B | See note** | 480P<br>720P | ❌ | ❌ | ❌ | ⭕ | ✅ |
|
||||
| Matrix Game 2.0 Base | `FastVideo/Matrix-Game-2.0-Base-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 GTA | `FastVideo/Matrix-Game-2.0-GTA-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 TempleRun | `FastVideo/Matrix-Game-2.0-TempleRun-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 Base Distilled | `FastVideo/Matrix-Game-2.0-Base-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 GTA Distilled | `FastVideo/Matrix-Game-2.0-GTA-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 TempleRun Distilled | `FastVideo/Matrix-Game-2.0-TempleRun-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 3.0 Base Distilled | `FastVideo/Matrix-Game-3.0-Base-Distilled-Diffusers` | 720x1280 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| GEN3C Cosmos 7B | `FastVideo/GEN3C-Cosmos-7B-Diffusers` | 704px1280p | ❌ | ❌ | ❌ | ⭕ | ⭕ |
|
||||
|
||||
**Note**: Wan2.2 TI2V 5B has some quality issues when performing I2V generation. We are working on fixing this issue.
|
||||
|
||||
***Lucy Edit Dev uses a non-commercial model license. FastVideo support is
|
||||
focused on inference integration for video editing workflows.
|
||||
|
||||
`Sliding Tile Attn (Legacy Branch)` entries refer to the archived
|
||||
`sta_do_not_delete` branch workflow, not active `main` inference wiring.
|
||||
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user