Compare commits

...
Author SHA1 Message Date
fbd823df89 [docs] dreamverse-integration: GPU4 smoke validated; flashinfer prereq
Live smoke test on GPU4 confirms public fastvideo serve from will/ltx2_sr_port HEAD d23e71c2 boots cleanly with bf16 fallback (~24s) and empirically validates drift item #4 (only /health exists; /healthz, /readyz, /status, /prompt-system-config, /curated-presets all 404 — Phase 4 promotion target).

Plan doc updates:

* CUDA_VISIBLE_DEVICES=4 pin convention documented (original Dreamverse-side server uses this; smoke commands now consistently include it).

* flashinfer-python + flash_attn marked as Phase 0 prerequisites for production-equivalent NVFP4 smoke; bf16 fallback path documented for hosts without these deps.

* Cleanup commands updated with the verified setsid/disown launch pattern that survived the smoke session.

* Smoke-test evidence (boot time, /health JSON, endpoint matrix) recorded inline as a 2026-05-05 verification anchor.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>

Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>

Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>

Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 10:55:55 -07:00
d23e71c2f3 [docs] dreamverse-integration: align test strategy with B+ + GPU4 hook
Audit found the test strategy in integration-plan.md was partially stale
from the prior Option D / split-PR drafting and missing an explicit
local-GPU verification path. This commit reconciles testing with the
single-mega-PR (D-17) + Option B+ monorepo (D-18) reality and makes the
GPU4-on-this-node hook explicit.

* New "Test strategy" subsection (added before "File-by-file migration
  map"): test taxonomy table (unit / integration-with-fakes / live-GPU
  / FE build / FE Playwright / contract / SSIM), pytest marker
  scheme (`@pytest.mark.gpu`), pyproject.toml addopts to default-skip
  GPU tests in ubuntu-latest CI, and a per-phase responsibility matrix
  (which class runs in which gate per Phase 0..7).

* GPU4 local verification hook documented: this dev node has 8x B200;
  GPU4 currently held by Dreamverse-side server PID 2453227 / port
  8009; concrete kill / redeploy / smoke-test / cleanup commands so
  any phase touching the live-service path (Phase 2/3/4/5) can be
  validated locally without waiting for Buildkite-Modal.

* Phase 2 verification gate: `uv run` command upgraded to
  `--locked --extra test pytest ... -m 'not gpu'` (matches Phase 1
  `--locked` enforcement). Added explicit MANUAL GPU4 QA gate
  required, not optional: kill PID 2453227, deploy `fastvideo serve`
  on GPU4 from this branch, hit `/health`, run `pytest -m gpu`
  against the live deploy, capture output in PR. Rollback expanded
  to handle GPU4-discovered bugs.

* Phase 3 verification gate: Playwright in CI is now explicitly
  DEFERRED to Phase 4 (matches `ci-dreamverse-frontend.yml` scaffold
  which already comments out Playwright steps until health routes
  land). Removed the contradictory "or PR note explains" clause.
  Added recommended manual GPU4 Playwright `frontend-shell.spec.ts`
  against the Phase 2 GPU4 deploy. Same `--locked --extra test`
  upgrade for the backend regression check.

* Phase 4 verification gate: Frontend CI Playwright RE-ENABLED at
  end of this phase (the natural re-baseline once health routes
  land). Manual GPU4 full E2E for all 3 Playwright specs
  (backend-health + frontend-shell + preset-prompt-generation) is
  required to merge.

* OSS precedent table: clarified that chainlit's "split CI" is a
  per-language CI split (frontend vs backend workflows), not a
  PR-level split — independent of D-17's single-mega-PR decision.
  Removes any reading-confusion that the testing strategy carries
  forward from the abandoned split.

No claims, references, or commands tied to the abandoned split-PR
branches (`will/api_7.10`, `will/api_8`, `will/ltx2_sr_runtime`,
`will/ltx2_nvfp4`, `will/ltx2_post_fixes`, `will/agents_cleanup`)
remain in the plan doc. All test commands route through
`apps/dreamverse/server/` and `apps/dreamverse/web/` per the B+
layout.

PR #1288: https://github.com/hao-ai-lab/FastVideo/pull/1288

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 10:43:21 -07:00
1e66c74aec [docs] dreamverse-integration: D-18 — Option B+ monorepo plan
User decision after reviewing integration-review.md's Option D
recommendation. Combines Option D's "generic backend stays generic"
principle with Option B's monorepo subfolder layout for the
Dreamverse product. Result: single repo (FastVideo), Dreamverse
becomes apps/dreamverse/, generic backend stays at
fastvideo.entrypoints.streaming.*, Dreamverse repo gets archived
after migration.

Methodology: 3 parallel exploration agents (FastVideo build/CI
surface, Dreamverse product-vs-generic file split, Python+Next.js
monorepo precedents). Synthesis authored by ultrabrain category.
Oracle review applied as PASS-WITH-FIXES; 10 critical issues
addressed (import contract too tight, Phase 2 missing shim list,
prompt_enhancer misclassified, frontend CI Playwright/pnpm setup
broken, .gitignore product-asset issue, uv.lock missing, backend CI
path filters too narrow, cross-repo history claim wrong, Phase 6
too coarse, missing security/CORS risks).

* integration-plan.md (new, 1205 lines): executable 7-phase
  migration plan with file-by-file map, tooling stack (uv workspace
  + standalone pnpm), pyproject.toml/pre-commit/.gitignore/CI YAML
  diffs, risk register, verification gates, rollback. Phase 0 lands
  #1288; Phases 1-7 add skeleton, move backend, move FE, promote
  generic-pending, retire prompt enhancer fork (DR-1), CI/release
  cutover (split into 6a-6f), archive Dreamverse repo.

* integration-review.md: DEPRECATED banner added at top pointing to
  integration-plan.md. Body kept for drift-audit and OSS precedent
  reference (still authoritative). Reading-guide entry updated.

* decisions-log.md: D-18 entry capturing strategy decision,
  rationale (drops cross-repo coordination overhead from D-17 cycle;
  preserves architectural separation; OSS precedents support shape),
  why not A/B/C/D (concrete reasons for each), implications (apps/
  exclude in setuptools, no cross-repo history preservation, drift
  items fold into phases), and 4 deferred decisions (DR-2, VPO,
  history-import method, CORS/write-endpoint security).

* README.md: Last-reconciled bumped with D-18 reference. Reading-
  guide table updated: integration-plan.md is CURRENT;
  integration-review.md marked DEPRECATED.

Verified: pre-commit clean (memory dir excluded from yapf/ruff/mypy;
spaces-check passes). Oracle PASS-WITH-FIXES, all 10 critical fixes
applied (import contract relaxed to allow fastvideo.{api,configs,
entrypoints.streaming,entrypoints.video_generator}; Phase 2 lists 6
explicit shims for generic-pending; Phase 5 explicitly retires the
prompt fork; ci-dreamverse-frontend.yml uses pnpm/action-setup BEFORE
setup-node; Playwright deferred to Phase 4 until health routes land;
.gitignore unignores product assets; Phase 1 commits uv.lock with
--locked enforcement; backend CI watches fastvideo/api + streaming +
video_generator + configs; cross-repo history disposition documented;
Phase 6 split into 6a-6f substeps with independent verification
gates; CORS/write-endpoint risks added to register).

PR #1288: https://github.com/hao-ai-lab/FastVideo/pull/1288

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 10:37:11 -07:00
907620ff22 [docs] dreamverse-integration: drift audit + integration tradeoff doc
Comprehensive review against ../Dreamverse and ../FastVideo-internal
plus 4-option integration tradeoff analysis with recommendation.

Driven by user request to (1) verify will/ltx2_sr_port has not drifted
from the integration goal, and (2) analyze whether Dreamverse should
remain a separate repo, become a subfolder, or merge into FastVideo's
namespace entirely.

Methodology: 4 parallel exploration agents mapped Dreamverse repo
structure, FastVideo-internal residual, FastVideo public cross-repo
surfaces, and OSS service-in-monorepo precedents (vLLM, BentoML, Ray
Serve, TGI+ChatUI, Transformers.js, ComfyUI, AUTOMATIC1111). Synthesis
authored by ultrabrain category. Oracle review applied as
PASS-WITH-FIXES; critical issues (broad zero-drift wording, stale Dynamo
mapping citation, broken Ray Serve URL, missing D-8/VPO/D-12-A/D-12-B/
SBS coverage) addressed in this commit.

* integration-review.md (new, 863 lines): Part 1 drift audit, Part 2
  integration tradeoffs, Part 3 action items.

  Part 1 — Drift audit
    - Methodology: typed-public-boundary criterion, intentional refactor
      vs drift, evidence sourced from worktree files + memory dir.
    - Zero core typed API drift on the integration surface — but the
      realtime-runtime contract surface (health routes) IS real drift.
    - 8 numbered drift items (Dreamverse README/bootstrap, prompt
      enhancer fork, cerebras_ifm gap, health routes, layerwise offload,
      upsampler CLI, missing example config, LTX-2 stage equivalence
      verification gap).
    - 6 deferred/accepted residual items.
    - 17-row drift summary table with priority / effort / status /
      tracked-where / next-action columns. Includes D-8 (ltx2_image_crf
      flow), VPO (video_position_offset_sec semantics — overdue),
      D-12-A (GpuPool docstring), D-12-B (run_async migration), SBS
      (session/blob lifecycle).

  Part 2 — Integration path tradeoffs
    - Option A: status quo (Dreamverse separate, depends on fastvideo).
    - Option B: Dreamverse as subfolder under FastVideo.
    - Option C: full merge into fastvideo.entrypoints.dreamverse.*.
    - Option D: hybrid — backend merges, frontend stays separate.
    - Comparison matrix on 13 axes.
    - 7 OSS precedent rows with citations (vLLM, BentoML, Ray Serve,
      TGI, Transformers.js, ComfyUI, AUTOMATIC1111).
    - Recommendation: Option D (constrained) — backend merges as
      generic FastVideo streaming, frontend stays separate. Mirrors
      ComfyUI's late-stage frontend split. Strong precedent: TGI +
      ChatUI. Critical librarian finding: NO 1:1 precedent exists for
      "Python ML library + Next.js product merged into library
      namespace" — argues against Option C.
    - Conditions that would change the recommendation enumerated.

  Part 3 — Action items: phased migration sketch (phases 0-5) with
  concrete file paths, effort estimates, and cross-references to
  open-threads.md / decisions-log.md.

* README.md: register integration-review.md in deep-dive reading guide.
  Bump Last reconciled to b36bdbc9 (current PR #1288 head, post
  STACK.md removal + integration-review.md addition); 36 commits ahead.

* state.md: branch tip refresh — b36bdbc9 / 36 commits / now includes
  integration-review.md.

* pr-roadmap.md: PR #1288 head bumped to b36bdbc9.

* open-threads.md: Last updated header bumped to b36bdbc9. Item D
  reference SHA bumped to b36bdbc9.

* authors.md: PR #1288 row bumped to b36bdbc9 / 36 commits / 70 files
  (post STACK.md removal).

Verified:
  - Pre-commit clean (yapf/ruff/mypy/codespell skipped per memory dir
    excludes; spaces-check passes).
  - Oracle review: PASS-WITH-FIXES, all 4 critical issues addressed.
  - Cross-references resolve: design.md, cross-repo-surfaces.md,
    decisions-log.md, open-threads.md, pr-roadmap.md, state.md,
    runbook.md, authors.md, streaming-server.md, quantization.md.

PR #1288: https://github.com/hao-ai-lab/FastVideo/pull/1288

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 04:20:56 -07:00
b36bdbc96d [chore] remove STACK.md — split-PR plan abandoned per D-17
The 10-PR / 6-remaining-slice split tracked by STACK.md is no longer the
landing strategy. Per [decisions-log.md D-17](.agents/memory/dreamverse-integration/decisions-log.md#d-17),
the remaining `will/ltx2_sr_port` content ships as a single mega-PR
(#1288) instead of 6 stacked PRs. The bulk-rebase formula and split
bookmark table in STACK.md no longer reflect reality, so keeping the
file in the repo would mislead future agents/contributors.

Already deleted:
* PR #1287 (was first slice of split) — closed
* origin/will/api_7.10 — remote branch deleted
* Local split bookmarks (will/api_7.10, will/api_8, will/ltx2_sr_runtime,
  will/ltx2_nvfp4, will/ltx2_post_fixes, will/agents_cleanup) — deleted

Kept:
* CO-AUTHORS.md — still the canonical roster reference (mirrored in
  .agents/memory/dreamverse-integration/authors.md but the top-level
  file is the source of truth for the broader stack history).
* will/ltx2_sr_port-pre-1286-rebase — local safety backup of the
  pre-rebase chain, scheduled for deletion after #1288 merges.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:39:17 -07:00
aa64458db1 [docs] dreamverse-integration: D-17 — abandon split, single mega-PR #1288
User decision after the post-#1286 rebase + re-slice cycle. Closing
#1287 (slice 1-3); opening #1288 on will/ltx2_sr_port covering the full
34-commit chain (71 files, +13,074/-583 LOC) at once. STACK.md model
deprecated; split bookmarks no longer maintained.

* decisions-log.md: add D-17 with rationale (review-coordination
  overhead vs structural benefit, layers not actually independent,
  CI/merge-queue simplicity), implications (STACK.md deprecated,
  bookmarks stale, runbook protocol replaced), and watch-outs (PR
  size, fallback to re-split if main moves significantly). Bumps
  Last updated header.

* README.md: switch narrative — single mega-PR #1288 on
  will/ltx2_sr_port @ 39dfa009 replaces the 6-PR split plan. PR #1287
  CLOSED. Cross-link to D-17.

* state.md: header reflects strategy reversal. Branch tips collapse
  the deprecated split bookmarks into one row noting D-17 deprecation.

* pr-roadmap.md: "In flight" is now the mega-PR #1288. New "Closed
  PRs in this scope" section captures #1287's history. New
  "Deprecated split bookmarks (D-17)" section names the abandoned
  branches. "Planned" section trimmed to post-#1288 work.

* open-threads.md: header updated. Item D resolution gate flipped
  from #1287 merge to #1288 merge — same content, different vehicle.

* runbook.md: branch-topology section reflects single-PR model.
  "After a PR merges (re-slice protocol)" replaced by "After PR
  #1288 merges" with simpler 6-step cleanup. The deprecated
  re-slice protocol is preserved in git history at b34d9704 and
  referenced from state.md.

* authors.md: PR #1287 row marked CLOSED; new PR #1288 row added
  with full 34-commit / 71-file / +13,074 LOC scope. All 34 commits
  carry the 4 co-author trailers (verified per the per-commit
  trailer block in every commit message).

PR #1288: https://github.com/hao-ai-lab/FastVideo/pull/1288

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:36:51 -07:00
39dfa0090f [docs] dreamverse-integration: reconcile post-#1286 merge + 7.10 split
PR #1286 merged at 2aaeee2a (squash). will/ltx2_sr_port rebased onto
new origin/main, dropping 4 commits whose content is now in main:
cd76cf51 + 1ac1e732 + b0b7f59c + 40e265b8. Backup preserved on
will/ltx2_sr_port-pre-1286-rebase. PR #1287 opened on will/api_7.10
@ 6ae7a99f (slice 1-3, generate_async + VideoEvent + tests).

* README.md: bump Last reconciled to post-#1286 merge + rebase. Note
  new working tip b34d9704 on will/ltx2_sr_port (33 commits ahead),
  PR #1287 OPEN MERGEABLE, backup branch preserved.

* state.md: rename header to "post-#1286 merge + rebase". Branch-tips
  table refreshed with all 8 relevant branches and the safety backup.
  Add "Post-#1286 rebase summary" describing the 4 dropped commits +
  zero conflicts. Add "New linearized chain" table mapping every
  surviving slice (7.10/8/LTX-2 SR/NVFP4/post-fixes/agents_cleanup)
  to its tip SHA. Mark the historical layered-chain section as
  pre-rebase narrative (SHAs only valid on the backup branch).

* pr-roadmap.md: promote PR #1284 (7.8) and PR #1286 (7.9) to Landed
  with merge SHAs eb3a3942 and 2aaeee2a respectively. Move PR #1287
  (7.10) from Planned to In flight with full scope description.
  Refresh remaining Planned rows with new slice indices and new tip
  SHAs from the rebased chain.

* open-threads.md: bump Last updated header to capture the merge +
  rebase + #1287 opening. Item D (Implement generate_async) flipped
  from "High pri" to "🟢 in flight" with the live PR #1287 reference;
  resolution gate is the merge of #1287 alongside Q-5/Q-9/PR-7.5/
  D-12-B chain.

* runbook.md: bump Last updated. Branch-topology section reflects
  PR #1286 -> merged, PR #1287 -> active. Add a new "After a PR
  merges (re-slice protocol)" section with the 10-step recipe used
  in the post-#1286 rebase, derived from the canonical STACK.md
  bulk-rebase formula adapted for squash-merge drops.

* authors.md: bump Last updated. PR #1286 row promoted to merged with
  squash-merge note about how the trailerless cherry-pick a152cb77
  was absorbed cleanly. PR #1287 row added with all 4 trailers
  verified on every commit. Known-gaps section rewritten — the two
  trailerless commits are no longer reachable from any active branch
  (a152cb77 absorbed by squash, 40e265b8 dropped by rebase). Gap
  permanently resolved; backup branch preserves them archeologically.

Verified: 17/17 router tests pass on rebased branch (router code now
loaded from origin/main). 221 passed across api/contract/NVFP4 suites.
Pre-commit clean (memory dir is yapf/ruff/mypy excluded; only
spaces-check runs). PR #1287 created at
https://github.com/hao-ai-lab/FastVideo/pull/1287, MERGEABLE.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:13:49 -07:00
b34d970442 [docs] dreamverse-integration: add runbook + fresh-context onboarding
Make the memory dir self-sufficient as an onboarding package — a fresh
agent should be able to resume work end-to-end (verify, commit, push,
propagate to PR #1286) using only files in this directory.

* runbook.md (new): operational how-to. Covers worktree contract,
  branch topology, verification (pre-commit binary path, pytest,
  lsp_diagnostics, gh CLI), commit workflow (subject/body conventions,
  required co-author trailers, the multi-`-m` trailer-parsing trap and
  its `-F` workaround), push + PR propagation (cherry-pick from
  ltx2_sr_port to api_7.9 to avoid force-push), memory-dir maintenance
  table, common pitfalls (pre-commit binary location, stash@{0} not
  ours, AbsMaxFP8 pre-existing failure, untracked nested clones, live
  ports, branch-switch by other agents, force-push policy, two known
  trailerless commits in PR #1286), 8-question self-test, and a "first
  60 seconds" copy-paste orientation block.

* README.md: replace single "Reading guide" with a two-tier structure.
  New "Fresh-context onboarding (read in order)" section gives 5
  ordered steps for an agent picking up the work for the first time —
  worktree confirmation via runbook, then state.md, pr-roadmap.md,
  open-threads.md, runbook.md end-to-end, finishing with the runbook
  self-test as a context-loaded check. Original table preserved as
  "Deep-dive reading guide" with a new entry for runbook.md. Bumps
  "Last reconciled" header to 2026-05-05 / `09647a30` and updates the
  PR status line (PR #1284 merged; PR #1286 open at `a152cb77`,
  MERGEABLE).

* state.md: bump "Current State" header from 2026-05-03 to 2026-05-05.
  Branch-tips table refreshed — `will/ltx2_sr_port` now at `09647a30`
  (was `156103b9`), 32 commits ahead of i2v base; new row for
  `will/api_7.9` (PR #1286 head `a152cb77`) noting it's an ancestor of
  the working branch. Added cross-links to runbook.md and authors.md
  in the front matter, plus a note about other agents sharing the
  worktree (covers the branch-switch scenario without inventing
  reconciliation logic).

Verified: pre-commit clean (memory dir is yapf/ruff/mypy excluded; only
spaces-check runs). All cross-links resolve to siblings in this
directory or to top-level CO-AUTHORS.md / STACK.md / AGENTS.md as
appropriate.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
e2dce4c478 [docs] dreamverse-integration: add authors.md + track D-16 router polish
* authors.md (new): dreamverse-integration-scoped mirror of the top-level
  CO-AUTHORS.md. Documents the 4 human co-authors credited on every commit
  in the integration scope, verified against PRs #1257/#1258/#1284/#1286.
  Self-contained so the memory dir is discoverable without traversing to
  the repo root. Cross-references CO-AUTHORS.md and STACK.md. Includes a
  'Known gaps' section flagging that commits a152cb77 / 40e265b8 (the
  [fix] streaming: router polish pair from earlier in this session) are
  missing trailers and need an amend + force-push to fix.

* README.md: register authors.md in the reading guide table.

* decisions-log.md: add D-16 — Streaming router polish round 2 —
  capturing the 5 second-pass fixes applied on top of D-15's pre-merge
  polishes (bridge cancellation hygiene, registry UNKNOWN->HEALTHY
  immediate, httpx hard-fail, URL path/query/fragment/duplicate
  validation, per-index YAML parser errors, explicit websockets
  [streaming] dep, +7 test cases). Bumps Last updated header.

* open-threads.md: bump Last updated header to mention the second-pass
  router commit and link to D-16.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
SolitaryThinker 27d6e378dd [docs] dreamverse-integration: track D-15 + #13/#14/#15 — PR #1286 arch review
Adds D-15 to decisions-log.md: Oracle review of PR #1286 streaming
router. Verdict — keep current shape (Alt A): in-repo at
fastvideo/entrypoints/streaming/router/, FastAPI-based, single-primary
failover, lazy httpx/websockets imports. Reject Alt B (separate
package), Alt C (fold into server), Alt D (delegate to
nginx/envoy/HAProxy as the sole answer), Alt E (sticky now), Alt F
(round-robin now).

Captures the 4 gemini review fixes:
* High: select() docstring claimed round-robin but always returns [0]
  — rewrote with explicit MVP semantics
* Medium: per-probe httpx.AsyncClient prevented connection reuse —
  refactored run_health_check_loop to share one client via
  _build_default_probe() async context manager
* Medium: sequential probes could fall behind interval — now use
  asyncio.gather for parallel per-cycle probes
* Medium: parser duplication — kept manual mapping intentionally
  because YAML has nested health_check: block while RouterConfig is
  flat; restructuring is beyond this PR's scope

Plus 1 Oracle pre-merge polish:
* RouterConfig.__post_init__ validation — empty replicas, non-positive
  intervals/timeouts, thresholds < 1, non-http(s) URLs, and >1
  primary all raise ValueError so misconfigurations surface at
  config-load instead of confusing runtime failures

Plus 1 prep-rebase item carried in:
* Migrated FastAPI @app.on_event(startup/shutdown) to a single
  @contextlib.asynccontextmanager-based _lifespan() handler
  (deprecated API in newer FastAPI; was the 7.9 caveat in
  pr-roadmap.md)

Adds open-threads.md items #13/#14/#15:
* #13: sticky session routing extensibility (forward-compat from D-15)
* #14: bridge backpressure note for high-scale deployments
* #15: multi-primary semantics if active-active becomes required
2026-05-05 03:05:53 -07:00
SolitaryThinker 95e5120cc0 [docs] dreamverse-integration: track D-14 + #12 — PR #1284 arch review
Adds D-14 to decisions-log.md: Oracle review of PR #1284 streaming
auxiliaries. Verdict — keep current shape (Alt A): single PR, 4 modules
under streaming/, mock_server in production module path, concrete
PromptSafetyFilter. Reject Alt B (split into 4 PRs), Alt C (move mock
to tests/), Alt D (premature observability extraction), Alts E/F
(premature Protocol-ization).

Captures the 4 gemini review fixes (high: session_logger race vs close;
medium: rewrite regex in hot path, safety _ensure_loaded race,
streaming extra missing prompt-safety) and 2 pre-merge polishes from
Oracle (remove inert RewriteOptions.user_system_prompt_override field,
sanitize session_id for filename safety as defense-in-depth).

Adds open-threads.md item #12 (Low): when streaming server starts using
PromptSafetyFilter, ensure operator-visible logging on
SafetyDecision.UNAVAILABLE results — surfaces degraded-safety state per
D-14's Watch-Out note.
2026-05-05 03:05:53 -07:00
SolitaryThinker a88920f071 [docs] dreamverse-integration: reflect 7.5/7.6/7.7 merged + 7.8 opened
Updates following the PR landings during this session:

* pr-roadmap.md: move PR 7.5 (#1251 merged 2026-04-26 as 95fd29e0),
  PR 7.6 (#1257 merged 2026-05-04 as eb0a4152), and PR 7.7 (#1258
  merged 2026-05-04 as f673423b) into 'Landed PRs' section. Add merge
  commit refs and link to D-12 / D-13 architecture reviews. Add PR 7.8
  (#1284 OPEN) as the new 'In flight' entry. Add LTX-2 SR / NVFP4 /
  agents-cleanup branches to the Planned section as out-of-band streams
  ready to open.
* open-threads.md: mark items #9 + #10 RESOLVED (commit-message
  cleanups bundled into the will/api_7.8 prep rebase). DR-1 (Dreamverse
  compat shim) marked actionable now that #1258 has merged. Recommended
  pull order updated.
* README.md: bump last-reconciled date and current branch tip
  (89a6484d post-rebase). Note PRs 1257 + 1258 merged; PR 1284 open.
2026-05-05 03:05:53 -07:00
SolitaryThinker 4288572d76 [docs] dreamverse-integration: track D-13 + 9 new follow-ups from this thread
Adds D-13 to decisions-log.md mirroring D-12's format: PR #1258 prompt
enhancer / LLMProvider abstraction shape review by Oracle. Verdict —
keep current shape (Alt A); promote to top-level fastvideo.prompt.* only
when a second non-streaming consumer exists. Don't convert Protocol →
ABC. Three deferred polishes captured.

Adds 9 new items to open-threads.md priority table:

* DR-1 (High): Dreamverse must create prompting/_internal_compat.py
  shim and delete most of the 1933-LOC local prompt_enhancer.py once
  PR #1258 merges.
* DR-2 (Med): Decide cerebras_ifm provider path — public Literal vs.
  Dreamverse-side custom provider via register_provider().
* D-12-A (Med): GpuPool ABC docstring — mark experimental /
  server-internal until PR 7.10 cycle.
* D-12-B (Med): Replace GpuPool.run() -> Any with run_async() ->
  AsyncIterator[VideoEvent] in PR 7.10 cycle.
* D-13-A (Med): Document streaming/prompt/* as streaming-scoped in
  user-facing docs; avoid framework-level framing.
* D-12-C (Low): Avoid locking PoolAssignment.gpu_id: int as public;
  prefer worker_id (already exists) or device_ids: list[int] for
  future multi-GPU-per-worker.
* D-13-B (Low): Optional client_factory parameter for httpx pooling
  if metrics justify.
* #9 (Low): PR 8's three commits still have [8/n] Improve API:
  prefix — regex didn't match single-digit version. Cleanup when
  PR 8 is opened.
* #10 (Low): PR 7.8/7.9 commits have streaming: streaming X
  duplication. Cleanup when 7.8/7.9 are opened.
* #11 (Low): Promote LTX-2 prompt orchestration (locked segments,
  segment_prompts JSON shape) to public when a second consumer
  appears.

Updates recommended pull order to include the new items.
2026-05-05 03:05:53 -07:00
SolitaryThinker 3756c60ca0 [docs] dreamverse-integration: track GpuPool architecture decision (D-12)
Adds D-12 to decisions-log.md, documenting the Oracle review of whether
GpuPool (PR #1257, just merged) should be folded into VideoGenerator.

Decision summary: keep GpuPool separate from VideoGenerator (Alt A) as
interim, evolve to Alt C (thin async executor over generate_async) once
PR 7.10 lands. Do not pursue Alt B (folding into VideoGenerator).

Captures:
* Three alternatives evaluated (A: status quo, B: VideoGenerator absorbs
  pool role, C: async executor over generate_async)
* Key finding: MultiprocExecutor and SubprocessGpuPool are orthogonal,
  not redundant — both spawn subprocesses because CUDA contexts demand
  process boundaries, but they solve different problems (one-call-fast
  vs. many-sessions-throughput)
* Sticky binding stays in pool, NOT in VideoGenerator (different
  consumers want different policies — sticky for streaming, lease for
  HTTP, queue for per-frame real-time)
* Risks flagged: GpuPool.run() sync shape, PoolAssignment.gpu_id: int
  freeze, and the 'is this canonical serving API?' framing question

Action items deferred to PR 7.10 cycle: replace run() with run_async()
returning AsyncIterator[VideoEvent], clarify worker_id vs gpu_id field
naming, mark GpuPool docstring as experimental until 7.10.
2026-05-05 03:05:53 -07:00
4c5144163c [docs] add STACK.md + CO-AUTHORS.md trackers for the will/ltx2_sr_port stack
* STACK.md tracks the 10-PR split of will/ltx2_sr_port: per-PR slice
  ranges, tip SHAs (live, not hardcoded), independence map, landing
  strategy, fast-track candidates, and re-slice commands to run after
  any rebase or trailer injection.

* CO-AUTHORS.md documents the 4 GitHub users credited as co-authors
  via Co-authored-by trailers on every commit in the stack:
  - @Davids048 (Junda Su)
  - @RandNMR73 (Matthew Noto)
  - @XOR-op
  - @jzhang38 (Zhang Peiyuan)

  All 4 verified active in FastVideo-internal git history. Trailer
  emails use GitHub's <id>+<username>@users.noreply.github.com form
  for reliable account linkage.

Both files are temporary trackers (delete after the full stack merges
or as documented at the bottom of each file).

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
67fdb4f438 [docs] .agents/memory: add dreamverse-integration knowledge base
Consolidates the public-API refactor + LTX-2 streaming-server upstream
+ Dreamverse migration + NVFP4 quantization landing into a single agent
memory module. Future agents loading the dreamverse-integration story
can pull a single targeted slice (state, design, PR roadmap, streaming,
cross-repo, quantization, decisions, open threads) rather than wading
through ~200 KB of source documents.

Structure:
  .agents/memory/dreamverse-integration/
  ├── README.md              ← index + glossary + reading guide
  ├── state.md               ← branches, commits, live services
  ├── design.md              ← typed schema design
  ├── pr-roadmap.md          ← PRs 0-17 status
  ├── streaming-server.md    ← PRs 7.5-7.10 + Dynamo + build_app
  ├── cross-repo-surfaces.md ← Dreamverse 3-surface + Dynamo contract
  ├── quantization.md        ← NVFP4 + LinearBase fallback + AbsMaxFP8
  ├── decisions-log.md       ← D-1..D-11 + Q-1..Q-9 status
  ├── open-threads.md        ← 12 active items prioritized
  └── source-archive/        ← 7 archived source docs (pre-synthesis)

Source-archive contents (all previously untracked at repo root or in
.agents/exploration/): apirefactor.md (838 lines), PR-plan.md (1145
lines, was 'PR plan.md'), dreamverse_review.md (390 lines),
handoff-nvfp4-launch-demo.md, streaming-server-upstream-plan.md,
dreamverse_integration.md, video-generator-config-api-design.md.

Registered in .agents/memory/index.jsonl as
{name: dreamverse-integration, status: ready, trust: high}.

Also adds CLEANUP.md (temporary tracker) for the multi-phase .agents/
cleanup. Phase 1 (deletes) shipped in the prior commit; Phase 2+
(rewrites, additions, registry consolidation) are tracked there for a
follow-up.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
bd77fc155b [chore] .agents: drop stale STATUS.md + sync-dashboard SOP + vapor skills
Phase 1 of .agents/ cleanup. All deletions are stale, redundant, or
unused; .agents/ memory + skills + workflows remain functional.

Removed:
* STATUS.md — hand-maintained dashboard, last synced 2026-03-02 with
  wrong counts (claimed 8 skills/4 workflows/4 memory; actual 9/5/5)
  and references to non-existent snake_case filenames. Strictly
  redundant with .agents/{memory,skills}/index.jsonl.
* workflows/sync-dashboard.md — SOP that maintained the deleted
  STATUS.md, with obsolete pre-PR-4 file paths.
* skills/index-related-work/, skills/search-related-work/ — vapor
  skills operating on the empty .agents/memory/related-work/ registry.
  The search skill literally requires 'index has entries' as a
  prerequisite but none have ever been added. Re-introduce when the
  registry is actually populated.

Updated skills/index.jsonl to drop the two removed entries (was 9
entries; now 7).

.agents/scripts/sync-skills.sh ran post-deletion and pruned 2 stale
.claude/skills/ symlinks pointing at the removed skill directories.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
25897b677e [fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep
InputValidationStage normalizes generator to a one-element list when
num_videos_per_prompt == 1, but torch.randn only accepts a single
torch.Generator. Unwrap the list here to keep the single-sample path
working. Batched sampling (>1) currently collapses to the first
generator — flagged in a comment for the future batched-inference path.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
91bd76f2b7 [fix]: avoid model.to() round-trip in Gemma encoder forward
Wrap the device move in an equality guard so Dynamo can DCE it under
fullgraph=True (model.to() goes through _parse_to, which returns a
non-Tensor torch.device that Dynamo can't trace). The model is already
placed on the target device by prepare_for_compile, so the guard is a
runtime no-op. Also drop the orig_device save/restore — the encoder
should stay resident on the compute device between calls.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
6793166bdd test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow
Two new contract test files plus a few stale-name doc fixups so the
LTX-2 NVFP4 path can't silently regress.

* ``test_nvfp4_ltx2_wiring.py`` (6 tests): asserts that
  - ``LTXSelfAttention``'s ``to_q/to_k/to_v/to_out`` are
    ``ReplicatedLinear`` (not plain ``nn.Linear``);
  - ``NVFP4Config()`` attaches ``NVFP4QuantizeMethod`` to the
    quantized subset and the ``layer_prefix`` matches the
    ``ltx2.blocks.<i>.<sub>.<proj>`` paths in
    ``NVFP4Config.fp4_layers``;
  - non-tagged projections (cross-attn K/V, audio attn, audio FFN)
    fall back to ``UnquantizedLinearMethod`` instead of crashing on
    the quant_method assert;
  - ``BasicAVTransformerBlock`` propagates ``quant_config`` and
    ``prefix`` correctly to all 4 attention modules + FFN at once.
  These are CPU-only and don't need flashinfer.

* ``test_typed_quant_flow.py`` (4 tests): asserts that
  - typed ``engine.quantization.transformer_quant: "NVFP4"`` resolves
    through ``compat.py`` to a concrete ``NVFP4Config()`` instance;
  - omitting the typed surface leaves the ``transformer_quant``
    carrier ``None`` so legacy callers that mutate
    ``pipeline_config.dit_config.quant_config`` directly keep working;
  - ``__post_init__._apply_transformer_quant`` pins the carrier onto
    ``dit_config.quant_config``;
  - explicit ``dit_config.quant_config = …`` is preserved (the
    explicit setter wins over the typed carrier).

Plus stale ``FP4Config`` → ``NVFP4Config`` doc updates in
``api/compat.py``, ``fastvideo_args.py``, and ``layers/linear.py``
that the previous rename commit missed (comments and docstrings only,
no behavior change).

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
1e854e42cc refactor(quant): rename FP4 → NVFP4 to disambiguate from other FP4 variants
The FP4 implementation is specifically NVIDIA's block-scaled FP4
(e2m1 mantissa, fp32 alpha, ``layout_128x4`` scale layout, group
size 16) backed by FlashInfer's ``nvfp4_quantize`` and ``mm_fp4``
kernels. Naming the public surface ``FP4`` would collide with other
FP4 variants we may want to support later (OCP-FP4 / MX-FP4 / e3m0).

Mechanical rename, no behavior change:

* ``fp4_config.py`` → ``nvfp4_config.py``
* ``FP4Config`` → ``NVFP4Config``; ``get_name()`` returns ``"nvfp4"``
* ``FP4QuantizeMethod`` → ``NVFP4QuantizeMethod``
* ``convert_model_to_fp4`` → ``convert_model_to_nvfp4``
* ``QuantizationMethods`` literal: ``"FP4"`` → ``"NVFP4"``
* registered buffer names: ``_fp4_weight``/``_fp4_alpha`` →
  ``_nvfp4_weight``/``_nvfp4_alpha``
* loader helper ``_maybe_convert_model_to_fp4`` →
  ``_maybe_convert_model_to_nvfp4``
* test file rename + symbol updates
* Dreamverse worker updates its single import + call site

Internal-scope torch op namespace ``fastvideo_fp4::*`` and
``_get_ltx2_fp4_stage_profile`` left as-is (purely internal naming
that mirrors FastVideo-internal).

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
1dc991809b feat(ltx2): wire FP4 inference through fastvideo.layers.quantization
The public LTX-2 path was previously running full bf16 even when
callers set ``pipeline_config.dit_config.quant_config = FP4Config()``,
because:

* the DiT used plain ``torch.nn.Linear`` for FP4-eligible projections
  (``attn1/attn2.to_q/k/v/out``, ``audio_to_video_attn``,
  ``video_to_audio_attn``, FFN ``fc_in``/``fc_out``), so
  ``LinearBase.quant_method.apply`` was never reached;
* the loader did not call ``convert_model_to_fp4`` after weights
  loaded, so even matched layers had no ``_fp4_weight*`` buffers;
* ``QuantizationMethods`` did not register ``"FP4"``, so the typed
  ``engine.quantization.transformer_quant`` surface had no way to
  select it.

This commit completes the inference wire-up through the existing
``fastvideo.layers.quantization`` registry — no parallel pathway:

* swap ``nn.Linear`` → ``ReplicatedLinear`` for the LTX-2 attention
  ``to_q``/``to_k``/``to_v``/``to_out``/``to_gate_compress`` and the
  FFN ``fc_in``/``fc_out`` (``GELUApprox.proj`` and ``FeedForward``
  ``project_out``). Other linears (timestep MLP, caption proj,
  patchify, AdaLN scale/shift) stay ``nn.Linear`` to match internal.
* port ``_supports_prequantized_input`` and
  ``_linear_project_with_optional_prequant`` so the attention forward
  quantizes input once and reuses the ``(x_fp4, x_scale, x_global_sf)``
  tuple across q/k/v projections — matches internal exactly.
* plumb ``quant_config`` and ``prefix`` through
  ``BasicAVTransformerBlock`` → ``_init_transformer_blocks`` →
  ``LTXModel`` → ``LTX2Transformer3DModel`` so each
  ``ReplicatedLinear`` gets the correct
  ``ltx2.blocks.<i>.<sub>.<proj>`` prefix that
  ``FP4Config.get_quant_method`` matches against.
* register ``"FP4"`` in ``QuantizationMethods`` literal +
  ``get_quantization_config`` registry so typed
  ``engine.quantization.transformer_quant: "FP4"`` resolves to a
  concrete ``FP4Config()`` instance.
* add ``transformer_quant`` carrier on ``FastVideoArgs`` and
  ``__post_init__._apply_transformer_quant`` to pin the resolved
  quant config onto ``dit_config.quant_config`` (without
  overwriting an explicit setter on the dit_config).
* add ``_maybe_convert_model_to_fp4`` in ``fsdp_load.py`` that
  walks the loaded model once and registers
  ``_fp4_weight*``/``_fp4_alpha``/``_weight_global_sf`` buffers on
  layers whose ``quant_method`` is ``FP4QuantizeMethod``. flashinfer
  is imported lazily inside ``convert_model_to_fp4`` so this is a
  no-op on hosts without the FP4 kernels.
* restore ``LinearBase`` fallback to ``UnquantizedLinearMethod`` when
  ``quant_config.get_quant_method`` returns ``None`` — ``FP4Config``
  only tags a curated subset of layers, and the previous
  ``assert quant_method is not None`` would crash any non-tagged
  layer that received a quant_config.

Non-FP4 callers are unaffected: ``UnquantizedLinearMethod.apply``
runs the standard ``F.linear`` path, and parameter names
(``to_q.weight``, ``to_out.0.weight`` from ``nn.ModuleList``) are
identical to the previous ``nn.Sequential`` layout.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
13f3baae97 feat(compile): per-component compile + transformer_refine + prepare hook
Bring the public ``ComposedPipelineBase.post_init`` compile dispatch
to feature parity with FastVideo-internal:

* Compile ``transformer_refine`` alongside ``transformer`` and
  ``transformer_2`` whenever the DiT compile flag is on. Without this
  the LTX-2 stage-2 refine pass silently runs eager while the main
  transformer is compiled — different inductor fusion choices vs.
  eager execution can produce last-bit divergences in bf16.
* Drive the new per-component compile flags from ``FastVideoArgs``:
  text encoder (``enable_torch_compile_text_encoder``), VAE
  (``enable_torch_compile_vae``), audio VAE
  (``enable_torch_compile_audio_vae``). Each picks its own kwargs
  dict (``torch_compile_kwargs_*``) and falls back to the master
  ``torch_compile_kwargs`` when empty.
* Call ``module.prepare_for_compile()`` on each compiled submodule
  before invoking ``torch.compile`` so model-specific external state
  (e.g. lazy HF loads) lands outside Dynamo's tracer. ``Gemma3``
  implements this hook to materialize the Gemma weights up front.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
f041cc30e5 feat(api): typed per-component CompileConfig + FastVideoArgs carriers
Extend the typed inference API so callers can drive torch.compile per
component without falling back to the legacy flat-kwargs pathway.

* ``CompileConfig`` now exposes ``vae_enabled``, ``audio_vae_enabled``,
  ``dit_kwargs``, ``text_encoder_kwargs``, ``vae_kwargs``,
  ``audio_vae_kwargs`` alongside the existing ``enabled`` /
  ``text_encoder_enabled`` switches. The master ``backend``/
  ``fullgraph``/``mode``/``dynamic``/``extras`` continue to apply to
  every compiled submodule unless a per-component kwargs dict is
  non-empty (in which case it overrides entirely — matches the
  FastVideo-internal precedent).
* Add the matching runtime carrier fields on ``FastVideoArgs``:
  ``enable_torch_compile_text_encoder/vae/audio_vae`` plus
  ``torch_compile_kwargs_dit/text_encoder/vae/audio_vae``. These let
  the legacy flat-kwargs path round-trip through the typed config.
* ``api/compat.legacy_from_pretrained_to_config`` now lifts the new
  flat kwargs into ``CompileConfig``, and
  ``generator_config_to_fastvideo_args`` writes them back out so
  ``VideoGenerator.from_pretrained(model_path, ...)`` callers keep
  working unchanged.

No behavior change yet — these surfaces are ports of carriers only.
The composed pipeline base will start consuming them in the next
commit.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
5f05e3e8f6 fix(api): propagate generic refine_* args + match internal randn
Three small parity fixes uncovered while comparing the LTX-2 distilled
path against FastVideo-internal:

* `FastVideoArgs.__post_init__` now calls `_resolve_refine_args()` to
  copy the public-facing generic `refine_*` knobs onto their
  `ltx2_refine_*` runtime carriers (mirrors internal lines 303-322).
  Without this, callers that set `refine_lora_path=...` via the typed
  CLI/config surface would silently drop the value, surfacing later as
  the "applied to 0 layers" warning.
* Revert the patch-noise sampler in `_randn_ltx2_video_latents` from
  `randn_tensor` back to `torch.randn` to bit-match internal under
  single-generator inference. `randn_tensor` is identical for a single
  `torch.Generator` but diverges for `list[Generator]` (per-sample
  seeds), which is the only place the two paths could disagree.
* Classify the 19 previously-unclassified `refine_*` / `ltx2_refine_*`
  / `ltx2_audio_latent_path` / `ltx2_images` / `ltx2_image_crf` /
  `ltx2_conditioning_latent_*` / `ltx2_video_conditions` fields in the
  schema-parity inventory yaml so the parity test passes.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
e729751986 feat(ltx2): full i2v conditioning + continuation latent port
Replaces the T2V-only stub with the full internal port:

* resolve_ltx2_images / _resize_and_center_crop / CRF re-encode
* load_ltx2_conditioning_image + load_ltx2_conditioning_video_clip
* _extract_video_latent / _insert_conditioning_latent
* build_ltx2_image_conditioning composes (clean_latent, denoise_mask)
  from images + video clips + continuation latents (stage1 vs stage2)
  — returns None for plain T2V

ForwardBatch + SamplingParam gain the missing fields the builder
reads (ltx2_images, ltx2_image_crf, ltx2_conditioning_latent_stage1/2,
ltx2_video_conditions). generate_video kwargs flow through
sampling_param.update -> ForwardBatch(**shallow_asdict(...)) so
Dreamverse's ltx2_image_crf=0.0 and continuation latents land on the
batch correctly.

68/68 LTX-2 tests still green.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
d879fbd013 fix(registry): order LTX-2 detectors so distilled wins for distilled paths
The model-name detector loop in get_model_name_for_path tests both
``model_path.lower()`` AND ``pipeline_name`` against each detector,
OR-ing the result. The pipeline_name for the distilled checkpoint is
``ltx2pipeline`` — a string that contains no "distilled" marker — so
the *base* detector's ``and "distilled" not in path`` predicate
returns True against it, even when the absolute model path clearly
contains "distilled".

Result: the resolver matches both LTX-2 base and LTX-2 distilled,
warns "Multiple models matched … Using the first matched: '0'", and
falls back to whichever was registered first. That used to be base
(ltx2_base preset → cfg=3.0 / mod=3.0 / rescale=0.7 / stg=1.0), so
SamplingParam.from_pretrained on a distilled-checkpoint path silently
loaded the *full* LTX-2 sampling defaults — exactly the divergence
the public-vs-internal Dreamverse alignment surfaced.

Reorder so the distilled entry registers first. The registry's
"first matched wins" tiebreak now lands on the more specific preset
when both detectors fire, and SamplingParam.from_pretrained reaches
the internal-aligned distilled defaults (mod=1.0 / rescale=0.0 /
stg=0.0 / cfg_video=1.0 / cfg_audio=1.0).

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
ee4e21544d fix(api): align public SamplingParam ltx2 defaults with distilled
Internal FastVideo declares two LTX-2 SamplingParam variants:
LTX2SamplingParam (mod=3.0/rescale=0.7/stg=1.0 — full LTX-2) and
LTX2DistilledSamplingParam (mod=1.0/rescale=0.0/stg=0.0). Dreamverse
production runs against LTX2-Distilled-Diffusers, so the distilled
defaults are what the runtime should land on.

The public package only has the single SamplingParam class, and its
class-level defaults for these knobs were the *full* LTX-2 values.
Result: when Dreamverse (or any other consumer) called
``SamplingParam.from_pretrained("FastVideo/LTX2-Distilled-Diffusers")``,
the LTX2_DISTILLED preset's defaults dict didn't override these
fields (it didn't list them) so they fell through to the class
defaults — silently picking up *full*-model CFG / modality / rescale /
STG instead of the distilled ones. Streamed video output diverged
from the internal-ui reference.

Switch the class defaults to the distilled values (1.0 / 0.0 / 0.0).
The LTX2_BASE preset already overrides them explicitly to 3.0/0.7/1.0
in its ``defaults`` dict, so users selecting the full-LTX-2 preset
keep the full-model behavior.

The numerical alignment harness still reproduces the same 42.72 dB
PSNR / 52.87% exact pixel match because that test pinned the values
explicitly via legacy kwargs; this fix moves the *default* path on
the public side onto the same numerical track.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
07fd06b763 test(ltx2-sr): pin ltx2 sampling knobs in harness for parity diff
Initial alignment showed PSNR ~12 dB because public's LTX2_BASE preset
sets ``ltx2_modality_scale_video=3.0 / rescale=0.7 / stg_video=1.0``
while internal's default ForwardBatch lands at ``1.0 / 0.0 / 0.0``.
Pinning these explicitly on both runs aligns the denoising mechanics
so the diff measures pipeline parity, not preset divergence.

Result with both sides pinned to mod=1.0/rescale=0.0/stg=0.0:

  max_abs_diff   = 203
  mean_abs_diff  = 0.788
  rms_diff       = 1.864
  psnr_db        = 42.72 dB
  exact pixels   = 52.87%
  within 1 lvl   = 86.94%
  within 4 lvls  = 97.93%

The ~47% near-but-not-exact bucket is consistent with bf16 reduction
non-determinism on B200 plus the fact that the public LoRA matcher
emitted "applied to 0 layers" for the LTX2-Distilled refine LoRA
(internal side did not). Both runs effectively skip the refine LoRA,
so they share the same effective stage-2 model — the residual diff is
the right ballpark for "same algorithm, different bf16 reduction
schedule".

The typed-API path also sets the same overrides via setattr on the
GenerationRequest so the typed path stays comparable. Once the typed
SamplingConfig grows ltx2-specific fields the setattr workaround can
go.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
1eebbe4da6 fix(ltx2-sr): close port gaps surfaced by alignment harness retries
LTX-2 SR pipeline now passes module load + stage graph build on the
public side. Each retry of scripts/ltx2_sr_alignment.py surfaced one
more gap; this commit closes them so the pipeline progresses to the
upsample → stage-2 → decode tail.

* harness: PipelineSelection (real class name; PipelineConfig was a typo)
* FastVideoArgs: generic refine_* fields (preferred user-facing API)
* LTX2Pipeline.initialize_pipeline: getattr-with-default for
  debug_model_sums / debug_model_detail (public args lack the debug
  fields; LTX-2 still flips the env-var toggles when set)
* VAEConfig: use_temporal_scaling_frames default True (LTX-2 latent
  prep reads this in the schedule-shift gate)
* LTX2LatentPreparationStage._randn_ltx2_video_latents: switch to
  randn_tensor so list-of-generators batches work (matches internal
  sampling order)
* ModelRegistry: register LTX2LatentUpsampler in _UPSAMPLERS
* component_loader: register spatial_upsampler / temporal_upsampler
  with UpsamplerLoader (was hitting the generic stub loader);
  UpsamplerLoader gains LTX-2 fallback when pipeline_config.upsampler_config
  isn't a multi-config tuple, plus the 'model.'-prefix-stripping
  weight load path the internal version uses

7 retries -> public side now reaches the upsample stage.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
5622863d75 test(ltx2-sr): add numerical alignment harness — public vs internal
Single script with three modes:

* --label internal --output /tmp/ltx2_sr_internal.pt  : runs the legacy
  from_pretrained kwargs path and saves frames + metadata. Used against
  PYTHONPATH=../FastVideo-internal as the reference run.
* --label public --use-typed-api --output /tmp/ltx2_sr_public.pt : runs
  through the new GeneratorConfig + GenerationRequest typed API with
  pipeline.preset_overrides["refine"] driving the SR path. Used against
  PYTHONPATH=../FastVideo (the public package) for the candidate run.
* --diff --reference R --candidate C : loads both .pt files and reports
  shape parity + max_abs / mean_abs / rms / psnr / pixel-bucket
  histograms over the uint8 frames.

Pinned alignment fixture: PROMPT + seed + 1088x1920x121x24fps + 8 base
steps + 3 refine steps mirrors basic_ltx2_upscale.py upstream so we
exercise the same code path the user has been generating with.
Deliberately skips FP4 (no quant config) and Dreamverse runtime knobs
to keep the bf16 path clean per request.

The two runs are intentionally invoked in separate processes via
PYTHONPATH overrides because both repos register as the `fastvideo`
package and can't be imported together. The diff script then runs in
either env to load both .pt files and compute the alignment metrics.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
241783f955 feat(ltx2): wire SR pipeline graph + port denoising/latent-prep stages
Brings the public LTX2Pipeline into parity with the internal SR-capable
version end-to-end:

* LTX2DenoisingStage gains the stage-2 control surface
  (sigmas_override, num_inference_steps_override, force_guidance_scale,
  initial_audio_latents_key) plus the i2v conditioning round-trip
  (clean_latent + denoise_mask via LTX2_VIDEO_*_KEY) and the audio
  refine path. The file is now ~1:1 with the internal stage; only the
  i2v_conditioning import path differs (rerouted to the public
  ltx2_image_conditioning module).
* LTX2LatentPreparationStage takes vae=, exposes the official LTX-2
  noise sampling order via _randn_ltx2_video_latents (patchify -> noise
  in token order -> unpatchify), and applies image conditioning when
  build_ltx2_image_conditioning returns a state. Matches internal.
* LTX2Pipeline switches its base from ComposedPipelineBase to
  LoRAPipeline (refine-LoRA support); create_pipeline_stages adds the
  refine_init / upsample / refine_lora / refine_denoising chain when
  ltx2_refine_enabled is True; load_modules pulls
  spatial_upsampler (and optionally transformer_refine) from the
  resolved upsampler path, with model_index.json defaults
  (fastvideo_refine_*) feeding the FastVideoArgs.ltx2_refine_*
  carriers.

stages/__init__.py re-exports the new refine stages so importers
outside the basic/ltx2 package keep working. 68/68 LTX-2-related tests
still pass after the rewrite (api translation, gpu_pool, contract,
streaming).

Reachability: setting pipeline.preset_overrides["refine"] = {"enabled":
True, ...} and components.upsampler_weights flows through compat.py
into the new ltx2_refine_* args, the pipeline branches, and load_modules
brings in the upsampler. T2V SR is now end-to-end runnable; numerical
alignment harness comes next.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
2ee0a2b217 feat(ltx2): port LTX-2 SR runtime — upsampler, refine stages, refine args
Lays the runtime-side pieces the LTX2_TWO_STAGE preset will exercise.
Behaviour matches FastVideo-internal 1:1 for the text-to-video SR path;
i2v / continuation conditioning is stubbed with a NotImplementedError
that fires only on those code paths so the public T2V SR run is clean.

* fastvideo/models/upsamplers/ltx2_upsampler.py — verbatim copy of the
  internal LTX-2 latent upsampler (PixelShuffleND, BlurDownsample,
  SpatialRationalResampler, ResBlock, LatentUpsampler, LTX2LatentUpsampler,
  upsample_video). Re-exported via models/upsamplers/__init__.py.
* fastvideo/pipelines/basic/ltx2/stages/ltx2_image_conditioning.py —
  minimal i2v-conditioning helpers: the four ForwardBatch.extra keys,
  apply_ltx2_gaussian_noiser, post_process_ltx2_denoised, and a T2V-only
  build_ltx2_image_conditioning that returns None when the request has
  no image inputs / no continuation latents (the SR T2V path).
* fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py — the three
  refine stages adapted to public imports: LTX2RefineInitStage halves
  the request resolution and stashes the original target;
  LTX2UpsampleStage runs the latent upsampler, optionally re-applies
  conditioning, and mixes refine-noise scaled by the stage-2 sigma;
  LTX2RefineLoRAStage swaps in a refine-specific LoRA (no-op when path
  is unset).
* fastvideo/fastvideo_args.py — add the runtime carrier fields
  ltx2_refine_{enabled,upsampler_path,transformer_path,lora_path,
  num_inference_steps,guidance_scale,add_noise,noise_path,
  audio_noise_path}. The typed public surface (ComponentConfig
  upsampler_weights/lora_path + pipeline.preset_overrides.refine) keeps
  flowing through compat.py:273-279, which expands the dict into these
  ltx2_refine_* kwargs at construction time.

Not yet wired: the LTX2 pipeline's create_pipeline_stages still goes
denoise -> audio_decode -> decode without the refine_init / upsample
/ refine_denoise injection. The remaining piece is teaching
LTX2DenoisingStage to accept sigmas_override / num_inference_steps_override
/ force_guidance_scale / initial_audio_latents_key (the internal
version is 725 lines vs public 395; ~330 lines of stage-2 logic still
need bringing across) and to flip the pipeline graph based on
ltx2_refine_enabled. That follow-up unlocks the numerical alignment run.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
0ee288b587 feat(quantization): upstream LTX-2 FP4Config with lazy flashinfer
Ports the FP4 quantization config + ops from FastVideo-internal
(fastvideo.layers.quantization.fp4_config). flashinfer is imported
lazily inside _require_flashinfer() rather than at module top so
``import fastvideo`` stays cheap on hosts that don't ship flashinfer
— the FP4 kernels themselves still raise a clear ImportError pointing
at ``pip install flashinfer-python`` when invoked.

Behavior is identical to the internal version:

* FP4Config: LTX-2 layer set (48 transformer blocks across attn1/attn2
  /audio_to_video_attn/video_to_audio_attn/ffn plus adaln_single.linear),
  layer_profile selects between "base" and "refine" stage subsets.
* FP4QuantizeMethod: per-layer quantize_input + apply with stage-aware
  routing for refine-only layers (audio_to_video_attn.to_q,
  video_to_audio_attn.to_k, video_to_audio_attn.to_v fall through to
  dense F.linear during the base stage).
* convert_model_to_fp4(): pre-quantize weight buffers + alpha onto
  module attrs at load time so the forward path doesn't pay the
  per-step weight-side quantize cost.

Required for Dreamverse integration: dreamverse-server's GPU worker
runs ``pipeline_config.dit_config.quant_config = FP4Config()`` during
LTX-2 model load, which is the dominant inference path.

test_fp4_config covers the lazy-flashinfer contract: import succeeds
without flashinfer; layer_profile round-trips via from_config; calling
into _require_flashinfer raises ImportError with an actionable
message when flashinfer is missing.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
f32e31eca1 [test] streaming: contract tests for Dreamverse + Dynamo shapes
fastvideo/tests/contract/ guards against drift in the two public-facing
adapter shapes that PR 8 locks down:

* test_dynamo_shape.py: mock Dynamo-style handler implementing the
  adapter template from docs/design/server_contracts/dynamo.md. It
  imports only fastvideo + fastvideo.api, translates a
  Dynamo-shape NvCreateVideoRequest dict into a typed
  GenerationRequest, runs it through an async generator that matches
  endpoint.serve_endpoint(handler.generate, ...), and round-trips the
  continuation-state envelope. A source scan asserts the adapter
  function body touches no private FastVideo module.
* test_dreamverse_shape.py: the flat init-time kwarg bag the internal
  gpu_pool.py passes today lands entirely on typed GeneratorConfig
  fields (no leakage into pipeline.experimental) after PR 6's typed
  replacement work. Per-request Dreamverse shape round-trips through
  legacy_generate_call_to_request -> normalize_generation_request;
  output.return_state (PR 7) stays wired through.

Failures in this suite appear at FastVideo CI — before the Dynamo-side
integration or private Dreamverse adapter see the drift.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
a1eed9a4c7 [docs] streaming: Dynamo native-backend integration reference
docs/design/server_contracts/dynamo.md documents the integration shape
the ai-dynamo/dynamo native backend package consumes. Per the refactor
decision, no Dynamo code lives in FastVideo — the backend package
ships in the Dynamo repo at components/src/dynamo/fastvideo/ with the
layout mirroring components/src/dynamo/sglang/.

Captured here:

* the 8-symbol public surface Dynamo imports (VideoGenerator plus
  ContinuationState, GenerationRequest, SamplingConfig, InputConfig,
  OutputConfig, VideoEvent, VideoResult) and what is available today
  vs. what lands in PR 7.10
* target backend package layout matching the sglang template
* full NvCreateVideoRequest <-> GenerationRequest mapping table and
  VideoFinalEvent <-> NvVideosResponse reverse mapping
* worked examples: aggregated handler (works today with generate_video
  + asyncio.to_thread), streaming handler (post-PR 7.10 via
  generate_async), health check payload (swapping to
  default_health_check_request() after PR 7.10)
* init function sketch, registration (ModelType::Videos fast path),
  args adapter
* contract guarantees that let the adapter be written once and not
  chase FastVideo drift
* explicit no-import list for the adapter

Also wires the server_contracts/ section into mkdocs.yml under Design.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
5cff4e184f [docs] streaming: OpenAI HTTP contract reference
docs/design/server_contracts/ lands the contract reference set that PR
8 locks down ahead of the streaming + Dynamo upstreams. This commit
adds the index page plus the OpenAI HTTP reference:

* endpoint catalogue (/v1/videos/generations, list, status, content,
  images, models, health)
* VideoGenerationsRequest shape with SGLang-compatible extensions
* merge precedence (body > default_request > hardcoded fallback)
  with a pointer to the explicit-path tracking mechanism that makes
  it work
* continuation-state roundtrip on the stateless surface
* HTTP error code table
* explicit non-goals: flat legacy kwargs, private Dreamverse fields,
  raw tensor payloads — all excluded from this boundary by design

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
6ae7a99f18 [test] streaming: generate_async coverage + refreshed streaming test
* tests/contract/test_generate_async.py: event ordering (Progress ->
  Final); exactly-one-final per request; batch expansion emits one
  Final per sub-result; VideoFinalEvent carries frames / full
  GenerationResult / continuation_state as expected; health-check
  request is minimal and round-trips through normalize_generation_
  request; public fastvideo.api exports the new symbols; a mock
  Dynamo-style async handler wraps generate_async using only the
  public import surface.
* tests/api/test_cli_translation.py: the streaming-dispatch test no
  longer expects NotImplementedError (PR 7.5 landed the live server);
  it now injects a fake run_server and asserts the CLI hands the
  typed streaming ServeConfig through unchanged.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
00aa56498e [feat] streaming: VideoGenerator.generate_async + health helper
* VideoGenerator.generate_async(request) yields VideoEvent objects:
  one VideoProgressEvent at start, one VideoFinalEvent per resulting
  GenerationResult (prompt-batch expansion emits one Final per
  sub-result). The aggregated code path runs the existing sync
  generate() inside asyncio.to_thread; future work threads per-step
  progress events from inside the pipeline's denoise loop without
  changing the public contract.
* VideoGenerator.default_health_check_request() returns a 256x256 /
  8-frame / 1-step typed GenerationRequest. Dynamo uses it to derive
  its health_check_payload without knowing FastVideo internals.
* _final_event_from_result packages GenerationResult into a
  VideoFinalEvent, reading the encoded MP4 off disk when the pipeline
  wrote one so streaming consumers don't re-encode.

Also: streaming_server.run_server validates streaming config before
loading the generator so the existing ValueError fires before the
expensive model-load path.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
4411266162 [feat] streaming: VideoEvent hierarchy + VideoResult alias
Add the typed event dataclasses VideoGenerator.generate_async will
emit, plus a VideoResult alias for the public docs:

* VideoProgressEvent / VideoPartialEvent / VideoFinalEvent
* VideoEvent = union of the three
* VideoResult = GenerationResult

Re-exported from fastvideo.api so the Dynamo backend package, the
streaming server, and the stateless OpenAI server share one set of
names.

Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:05:53 -07:00
2aaeee2ab8 [feat] Improve API: streaming router (multi-replica load balancer + ws proxy) (#1286)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:00:08 -07:00
eb3a394224 [feat] Improve API: streaming auxiliaries (safety, rewrite, logger, mock) (#1284)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 00:14:34 -07:00
f673423b51 [feat] Improve API: streaming prompt enhancer with LLMProvider abstraction (#1258)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 13:44:40 -07:00
eb0a41528a [feat] Improve API: streaming server GpuPool + worker subprocess (#1257)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 12:56:31 -07:00
100 changed files with 19332 additions and 601 deletions
-94
View File
@@ -1,94 +0,0 @@
# Agent Infrastructure — Status Dashboard
Developer-maintained overview of all agent components and their maturity.
Use this to understand what exists, how complete it is, and how much to trust it.
_Last synced: 2026-03-02_
> To resync this dashboard, use the workflow: `.agents/workflows/sync-dashboard.md`
---
## Summary
| Category | Total | ✅ Ready | 🟡 Draft | 🔴 Stub | Trust |
|----------|-------|---------|---------|---------|-------|
| Skills | 8 | 0 | 8 | 0 | Low — newly created, untested |
| Workflows (SOPs) | 4 | 0 | 4 | 0 | Low — newly created, untested |
| Memory files | 4 | 1 | 3 | 0 | Medium — codebase_map is solid |
| Lessons | 0 | — | — | — | N/A — empty |
| Exploration logs | 0 | — | — | — | N/A — empty |
---
## Skills (`.agents/skills/`)
| Skill | File | Status | Trust | Tested | Notes |
|-------|------|--------|-------|--------|-------|
| Launch Experiment | `launch-experiment.md` | 🟡 Draft | Low | ❌ | Needs dry-run validation |
| Monitor Experiment | `monitor-experiment.md` | 🟡 Draft | Low | ❌ | Requires W&B API access to test |
| Summarize Run | `summarize-run.md` | 🟡 Draft | Low | ❌ | Pattern from existing test infra |
| Log Experiment | `log-experiment.md` | 🟡 Draft | Low | ❌ | Journal formatting only |
| Evaluate Video Quality | `evaluate-video-quality.md` | 🟡 Draft | Low | ❌ | SSIM section most mature |
| Index Related Work | `index-related-work.md` | 🟡 Draft | Low | ❌ | Schema defined, no entries yet |
| Search Related Work | `search-related-work.md` | 🟡 Draft | Low | ❌ | Depends on indexed entries |
| Skill Template | `SKILL_TEMPLATE.md` | ✅ Ready | High | ✅ | Meta-template, stable |
### Trust Level Definitions
- **High**: Tested in production, validated against real experiments
- **Medium**: Logic is sound, partially tested or based on existing patterns
- **Low**: Newly written, not yet validated
- **None**: Placeholder only
---
## Workflows / SOPs (`.agents/workflows/`)
| Workflow | File | Status | Trust | Tested | Notes |
|----------|------|--------|-------|--------|-------|
| Experiment Lifecycle | `experiment-lifecycle.md` | 🟡 Draft | Low | ❌ | End-to-end flow, untested |
| Evaluation Development | `evaluation-development.md` | 🟡 Draft | Low | ❌ | Metric dev process |
| Experiment Journaling | `experiment-journaling.md` | 🟡 Draft | Low | ❌ | Journaling cadence |
| Lesson Capture | `lesson-capture.md` | 🟡 Draft | Low | ❌ | Post-experiment reflection |
| Sync Dashboard | `sync-dashboard.md` | 🟡 Draft | Low | ❌ | This dashboard's updater |
---
## Memory (`.agents/memory/`)
| File | Status | Trust | Notes |
|------|--------|-------|-------|
| `codebase_map.md` | ✅ Ready | High | Synthesized from full repo research |
| `experiment_journal.md` | 🟡 Draft | Medium | Schema defined, no entries yet |
| `evaluation_registry.md` | 🟡 Draft | Medium | SSIM/loss metrics documented |
| `related_work/README.md` | 🟡 Draft | Medium | Schema defined, no entries yet |
---
## Lessons (`.agents/lessons/`)
| File | Category | Severity | Notes |
|------|----------|----------|-------|
_No lessons captured yet._
---
## Exploration Logs (`.agents/exploration/`)
| File | Status | Topic | Notes |
|------|--------|-------|-------|
_No exploration logs yet._
---
## What to Do Next
1. **Validate skills**: Run a minimal training experiment using the
`experiment-lifecycle` SOP to test `launch-experiment` → `monitor-experiment`
→ `summarize-run` end-to-end.
2. **Index first related work**: Use `index-related-work` to add at least one
paper (e.g., the Self-Forcing paper used in the codebase).
3. **Capture first lesson**: After the validation run, capture any findings.
4. **Promote to Ready**: As each skill/SOP is tested, update its status here.
@@ -0,0 +1,163 @@
# Dreamverse Integration — Memory Index
Living knowledge base for the FastVideo ↔ Dreamverse ↔ Dynamo integration.
Tracks the public API refactor (PRs 0-17), the LTX-2 streaming server
upstream, the Dreamverse switch from `FastVideo-internal` to public
`FastVideo`, and the NVFP4 quantization landing.
**Last reconciled:** 2026-05-05 (**D-18**: Option B+ chosen — Dreamverse
becomes `apps/dreamverse/` subfolder under FastVideo; generic backend stays
at `fastvideo.entrypoints.streaming.*`. See [integration-plan.md](integration-plan.md)
for the executable 7-phase migration plan;
[integration-review.md](integration-review.md) is **deprecated** but kept
for the drift audit and OSS precedent citations.).
FastVideo `will/ltx2_sr_port` @ HEAD (post-D-17 STACK.md removal +
integration-review.md addition + integration-plan.md addition + D-18
reconciliation). Dreamverse `will/integrate-public-fastvideo` @ `ec8ef92`.
PRs #1257 / #1258 / #1284 / #1286 MERGED to main. **PR #1287 CLOSED
(in favor of consolidation); PR #1288 OPEN as the single mega-PR
landing the entire `will/ltx2_sr_port` chain at once** (LTX-2 SR
runtime + NVFP4 + `generate_async`/Dynamo contract + agents memory dir).
Split branches kept as historical bookmarks; STACK.md model **abandoned** —
see [decisions-log.md D-17](decisions-log.md#d-17). Local backup
`will/ltx2_sr_port-pre-1286-rebase` @ `1baa60bb` preserves the
pre-rebase chain.
## Fresh-context onboarding (read in order)
If you're an agent picking up this work for the first time, do these
**5 things in this order**. Once done, you have full context to continue
any open thread, commit correctly, push, and propagate to the open PR.
1. **Confirm worktree state** — run the "First 60 seconds" block in
[runbook.md](runbook.md). Tells you the branch is right, services
are up, and PR #1286's head matches what this dir claims.
2. **Read [state.md](state.md)** — single-page snapshot of branch tips,
live services, test status, pre-existing failures, "do not pop"
stashes.
3. **Read [pr-roadmap.md](pr-roadmap.md)** — what PRs landed, what's in
flight, what's planned. Identifies the active open PR (currently
#1286) and where it sits in the dependency chain.
4. **Read [open-threads.md](open-threads.md)** — prioritized work items
with effort estimates and dependencies. The "Recommended pull order"
section is a ready-made TODO list if you need one.
5. **Skim [runbook.md](runbook.md) end-to-end** — operational how-to:
verify, commit (with co-author trailers), push, propagate to PR
#1286, maintain the memory dir, and the "Common pitfalls" section
that catches the recurring traps.
Skip the deep-context docs (design / streaming-server / cross-repo /
quantization / decisions-log) until you need them — they're indexed in
the "Deep-dive reading guide" below.
Final check: run the "Self-test" block at the bottom of
[runbook.md](runbook.md). If you can answer all 8 questions from this
dir alone, you're ready. If you can't, the gap is a memory-dir bug —
file it in [open-threads.md](open-threads.md) before continuing.
## Deep-dive reading guide
| Question / task | File |
|---|---|
| "What's running right now? What just landed?" | [state.md](state.md) |
| "How do I commit / push / propagate to PR #1286?" | [runbook.md](runbook.md) |
| "Why is the schema typed this way? What's the philosophy?" | [design.md](design.md) |
| "What PRs landed? In flight? Planned?" | [pr-roadmap.md](pr-roadmap.md) |
| "Streaming server, `generate_async`, `build_app` routes?" | [streaming-server.md](streaming-server.md) |
| "How does Dreamverse use FastVideo? What about Dynamo?" | [cross-repo-surfaces.md](cross-repo-surfaces.md) |
| "NVFP4? Layer profiles? `LinearBase` fallback? AbsMaxFP8?" | [quantization.md](quantization.md) |
| "Why was decision X made? What's resolved vs. open?" | [decisions-log.md](decisions-log.md) |
| "What should I work on next? Priority order?" | [open-threads.md](open-threads.md) |
| "Who should be co-authored on commits in this scope?" | [authors.md](authors.md) |
| "How do we execute the Dreamverse → FastVideo monorepo merge?" | [integration-plan.md](integration-plan.md) ← **CURRENT** |
| "Historical drift audit + Option-D evaluation (deprecated by D-18)" | [integration-review.md](integration-review.md) (DEPRECATED) |
## Repo + worktree paths
| Repo | Path | Active branch |
|---|---|---|
| FastVideo (public) | `/home/william5lin/FastVideo` | `will/ltx2_sr_port` |
| Dreamverse | `/home/william5lin/Dreamverse` | `will/integrate-public-fastvideo` |
| FastVideo-internal (read-only ref) | `/home/william5lin/FastVideo-internal` | their `main` |
| Dynamo (read-only ref) | `/home/william5lin/dynamo` | upstream |
## Glossary
- **NVFP4**: NVIDIA's specific block-scaled FP4 (e2m1 mantissa, fp32 alpha,
`layout_128x4` scale layout, group size 16). Distinct from MX-FP4 / OCP-FP4.
- **`GeneratorConfig`**: typed init-time public config (model_path, engine,
pipeline). Replaces flat `from_pretrained(**kwargs)`.
- **`GenerationRequest`**: typed per-call request (prompt, inputs, sampling,
runtime, output, stage_overrides, state, plan, extensions). Replaces flat
`generate_video(**kwargs)`.
- **`ServeConfig`** / **`RunConfig`**: top-level YAML envelopes. ServeConfig
for `fastvideo serve`; RunConfig for offline `fastvideo generate`.
- **`InferencePreset`**: model-owned named preset (e.g. `ltx2_two_stage`)
defining stage topology + per-stage defaults + valid override types.
- **`ContinuationState`**: opaque round-trip state envelope `{kind, payload}`.
Hybrid: server-held for streaming WS, client-round-trip for stateless HTTP.
- **`generate_async`**: future canonical async exec API (PR 7.10) yielding
`VideoProgressEvent` / `VideoPartialEvent` / `VideoFinalEvent`. Substrate
for streaming server, OpenAI server, AND Dynamo backend.
- **`build_app`**: FastAPI app factory in
`fastvideo.entrypoints.streaming.server`. Currently exposes only
`/health` + `/v1/stream`. FE-required `/healthz`+`/readyz`+`/status`
migration is open follow-up #1.
- **`LLMProvider`**: protocol abstraction for prompt enhancer providers
(cerebras, cerebras_ifm, groq). Public schema currently restricts to
`Literal["cerebras", "groq"]`; `cerebras_ifm` is internal-only.
- **`compat.py`**: legacy kwargs translation layer (~370 lines). Scheduled
for death across PRs 14-17.
- **`prepare_for_compile`**: duck-type protocol method called via
`getattr(module, "prepare_for_compile", None)` before `torch.compile`.
Currently only Gemma3 implements it.
- **`SubprocessGpuPool`**: PR 7.6 public replacement for the internal
`realtime/local_runtime.GPUPool`. Per-GPU subprocess workers, typed
`GeneratorConfig` boundary.
- **PR 5.5**: streaming server subpackage skeleton — adds
`fastvideo/entrypoints/streaming/` parallel to `openai/`.
- **PR 7.10**: the unlock PR. Closes Q-5 (audio re-encode), Q-9 (Dynamo
progress), and PR 7.5's mid-segment cancellation TODO simultaneously.
## Live process map (as of 2026-05-03)
| Port | Service | Source |
|---|---|---|
| 8009 | `dreamverse-server` | running, `/readyz` 200, 1 warmed GPU worker |
| 5274 | `next-server` (dev) | running |
| 8000 | unknown FastAPI | not in handoff — verify before launching new BE |
## How this directory is maintained
- Source of truth for the integration story. Update when state changes.
- Each file has a "Last updated" header; bump when you edit.
- Cross-reference siblings via relative links; do NOT duplicate content.
- New entries: register in `../index.jsonl`.
- These files supersede the untracked source docs in the repo root and
`.agents/exploration/` — see [state.md](state.md) "Untracked but
present" section for disposition.
## Source documents (archived 2026-05-03)
The 7 source docs that this directory consolidates have been moved into
[`source-archive/`](source-archive/). They remain available for agents
who want the full unsynthesized rationale, but the synthesized memory
files in this dir are the canonical source of truth.
| Source doc | Lines | Synthesized into |
|---|---|---|
| [`source-archive/apirefactor.md`](source-archive/apirefactor.md) | 838 | [design.md](design.md) |
| [`source-archive/PR-plan.md`](source-archive/PR-plan.md) | 1145 | [pr-roadmap.md](pr-roadmap.md) |
| [`source-archive/dreamverse_review.md`](source-archive/dreamverse_review.md) | 390 | [state.md](state.md) + [decisions-log.md](decisions-log.md) |
| [`source-archive/handoff-nvfp4-launch-demo.md`](source-archive/handoff-nvfp4-launch-demo.md) | 518 | [state.md](state.md) + [quantization.md](quantization.md) + [open-threads.md](open-threads.md) |
| [`source-archive/streaming-server-upstream-plan.md`](source-archive/streaming-server-upstream-plan.md) | 539 | [streaming-server.md](streaming-server.md) + [decisions-log.md](decisions-log.md) |
| [`source-archive/dreamverse_integration.md`](source-archive/dreamverse_integration.md) | 285 | [cross-repo-surfaces.md](cross-repo-surfaces.md) |
| [`source-archive/video-generator-config-api-design.md`](source-archive/video-generator-config-api-design.md) | 93 | [design.md](design.md) (early-draft material) |
| `.agents/exploration/pr-link-review.md` | 29 | already promoted to `.agents/skills/review-pr-link/` (kept in exploration dir) |
See [`source-archive/README.md`](source-archive/README.md) for the
archive policy.
@@ -0,0 +1,150 @@
# Authors — Dreamverse Integration
**Status:** PERMANENT — keep around as the source of truth for who collaborated
on the dreamverse-integration work, even after every PR in the integration
scope has merged.
**Last updated:** 2026-05-05 (strategy reversal — single mega-PR #1288 on `will/ltx2_sr_port` replaces planned 6-PR split; #1287 closed; per [decisions-log.md D-17](decisions-log.md#d-17))
This file documents the human co-authors credited on every commit in the
dreamverse-integration scope (FastVideo public-API refactor, streaming server
upstream, GPU pool, prompt enhancer, NVFP4 wire-up, LTX-2 SR port). The
4 collaborators below worked on the FastVideo-internal precursor of this code
and are credited as co-authors on every public-side upstream commit via Git's
standard
[`Co-authored-by`](https://docs.github.com/en/pull-requests/committing-changes-to-your-project/creating-and-editing-commits/creating-a-commit-with-multiple-authors)
trailer convention.
Scope-wise this is the dreamverse-integration-flavored mirror of the
top-level [`CO-AUTHORS.md`](../../../CO-AUTHORS.md), which is scoped to the
broader `will/ltx2_sr_port` 10-PR stack. The roster is identical; this file
exists so the dreamverse-integration memory dir is self-contained and
discoverable without traversing to the repo root.
## Co-author roster
| GitHub user | Real name | GitHub ID | Trailer email |
|---|---|---|---|
| [`@Davids048`](https://github.com/Davids048) | Junda (David) Su | 90978028 | `90978028+Davids048@users.noreply.github.com` |
| [`@RandNMR73`](https://github.com/RandNMR73) | Matthew Noto | 99706358 | `99706358+RandNMR73@users.noreply.github.com` |
| [`@XOR-op`](https://github.com/XOR-op) | (unset) | 17672363 | `17672363+XOR-op@users.noreply.github.com` |
| [`@jzhang38`](https://github.com/jzhang38) | Zhang Peiyuan | 42993249 | `42993249+jzhang38@users.noreply.github.com` |
## Verification — where these trailers appear
Verified via `gh pr view <PR> --json commits --jq '.commits[].messageBody'`
across every PR in the integration scope:
| PR | Branch | Status | Trailers present on every commit |
|---|---|---|---|
| #1257 | `will/api_7.6` (GPU pool upstream) | ✅ merged 2026-05-04 | yes (4/4) |
| #1258 | `will/api_7.7` (prompt enhancer + LLMProvider) | ✅ merged 2026-05-04 | yes (3/3) |
| #1284 | `will/api_7.8` (streaming auxiliaries) | ✅ merged 2026-05-04 | yes (2/2) |
| #1286 | `will/api_7.9` (streaming router) | ✅ merged 2026-05-05 at `2aaeee2a` (squash) | yes on commits 1-3; commit `a152cb77` (`[fix] streaming: router polish`) was missing trailers but got squashed into the merge commit, so the merge commit on main inherits the trailers from the other 3. The trailerless cherry-pick partner (`40e265b8` on `will/ltx2_sr_port`) was dropped by the post-#1286 rebase — gap permanently resolved. |
| #1287 | `will/api_7.10` (`generate_async` + `VideoEvent`) | ❌ CLOSED 2026-05-05 — superseded by #1288 per [D-17](decisions-log.md#d-17) | yes on all 3 commits (now part of #1288's chain) |
| **#1288** | **`will/ltx2_sr_port`** (mega-PR — full stack: SR runtime + NVFP4 + generate_async + Dynamo contract + agents memory + integration-review) | 🟢 OPEN, MERGEABLE at `b36bdbc9`, 36 commits / 70 files / ~+13.0k LOC (post STACK.md removal) | yes on all 36 commits |
Aggregate count across `will/ltx2_sr_port` (top of stack) at the time of
writing: 32-33 commits per co-author, matching the 32 commits in the stack
on top of base `cfccd292`. Numbers stay consistent because the rebase
command (see "How the trailers were applied" below) walks every commit.
## Trailer block (copy-paste ready)
The trailers added to every commit on `will/ltx2_sr_port` and every
dreamverse-integration PR:
```
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
```
For one-off `git commit -m` invocations, use `--trailer` flags:
```bash
git commit -m "..." \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
```
`--trailer` is idempotent (dedupes by full `key: value`) so re-running is safe.
## Why no-reply emails
GitHub's `<id>+<username>@users.noreply.github.com` form is the most reliable
way to link a `Co-authored-by` trailer to a GitHub account. It:
- Always works regardless of whether the user has a public verified email
- Survives the user changing their primary email
- Doesn't expose anyone's personal email to git history
- Is the format GitHub itself produces when you click "Add co-author" in the
web UI
(All 4 collaborators have this email already used in `FastVideo-internal`
git history, verified via `git log --all` on that repo.)
## How the trailers were applied (bulk rebase)
```bash
git rebase --exec '
git commit --amend --no-edit \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
' origin/main will/ltx2_sr_port
```
After running, re-slice all 10 split branches per [`STACK.md`](../../../STACK.md)
and force-push the published branches (`will/api_7.9`, `will/ltx2_sr_port`).
## How to add a new co-author later
1. Add the user to the roster table above (and the top-level
[`CO-AUTHORS.md`](../../../CO-AUTHORS.md) — keep them in sync).
2. Append their `Co-authored-by` line to the trailer block above.
3. Re-run the bulk rebase command on `will/ltx2_sr_port` — git's trailer
dedupe handles the existing 4; the new one gets appended.
4. Re-slice all split branches per [`STACK.md`](../../../STACK.md).
5. Force-push the published branches.
## What we do NOT add
Per the repo's top-level [`AGENTS.md`](../../../AGENTS.md):
> Never add any coding agent or models such as Claude (or Claude Code), GPT,
> Codex or others as a co-author in commits or PRs. Do not include
> `Co-Authored-By: Claude ...` trailers or "Generated with Claude Code" and
> other such lines.
So no `Co-authored-by: Claude <noreply@anthropic.com>`, no
`Generated with Claude Code` footer, no `Cursor <cursoragent@cursor.com>`
trailer (one such commit exists on `will/ltx2_sr_port` from a pre-policy
external contribution and stays grandfathered; new commits MUST NOT introduce
the pattern). Only human collaborators.
## Known gaps
**Resolved 2026-05-05 by the post-#1286 rebase.** The two trailerless
commits (`a152cb77` on `will/api_7.9` and `40e265b8` on
`will/ltx2_sr_port`) are no longer reachable from any active branch:
- `a152cb77` was absorbed into squash merge `2aaeee2a` on main, which
inherits the trailers from the other 3 commits in the squash.
- `40e265b8` was dropped by the post-#1286 rebase of
`will/ltx2_sr_port`.
Both still exist on the local backup `will/ltx2_sr_port-pre-1286-rebase`
for archeological reference. No further action needed.
## See also
- [`../../../CO-AUTHORS.md`](../../../CO-AUTHORS.md) — top-level stack-scoped
co-authors file (same roster, broader scope)
- [`../../../STACK.md`](../../../STACK.md) — 10-PR split layout for
`will/ltx2_sr_port` (re-slice commands live here)
- [`pr-roadmap.md`](pr-roadmap.md) — per-PR status within the
dreamverse-integration scope
@@ -0,0 +1,277 @@
# Cross-Repo Surfaces — Dreamverse + Dynamo
How Dreamverse consumes FastVideo today, what's already shared, what's
ad hoc, and what migrations land alongside each PR. Plus the Dynamo
backend contract.
For the streaming-server side see [streaming-server.md](streaming-server.md).
For the API design see [design.md](design.md). For PR sequence see
[pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-03.
## The three surfaces
Dreamverse depends on FastVideo across three surfaces (in order of
stability):
1. **Pipeline construction** (stable)
2. **Realtime runtime** (in flight: PRs 7.5/7.6)
3. **Continuation state** (PR 7 typed; PR 7.6 wires server-held)
## Surface 1: Pipeline construction (stable)
`Dreamverse/server/video_generation.py:VideoGenerationWorker` calls
`VideoGenerator.from_pretrained(...)`.
After PR 6 the typed `GeneratorConfig` path exists; **as of `d80c2a8`
(May 2)** Dreamverse migrated to the typed path:
| Dreamverse usage | FastVideo public surface (post-PR 6) |
|---|---|
| `VideoGenerator.from_pretrained(model_path, ltx2_refine_enabled=…, …)` | `VideoGenerator.from_pretrained(config=GeneratorConfig(...))` |
| Flat `torch_compile_kwargs={…}` dict | `engine.compile.{backend,fullgraph,mode,dynamic,extras}` |
| `ltx2_vae_tiling=True` | `pipeline.vae_tiling=True` |
| `ltx2_refine_*` family | `pipeline.preset_overrides.refine.*` + `pipeline.components.upsampler_weights` |
| `enable_torch_compile_text_encoder` | `engine.compile.text_encoder_enabled` |
Refine knobs moved from `ltx2_refine_*` flat kwargs into
`preset_overrides["refine"]`. **The in-memory `pipeline_config` pin**
(`dit_config.quant_config = NVFP4Config()`) keeps using the legacy
`experimental["pipeline_config"]` carrier because typed
`transformer_quant: "NVFP4"` doesn't yet support setting
`layer_profile` (see [open-threads.md](open-threads.md) follow-up #4 +
[quantization.md](quantization.md)).
Legacy flat-kwarg path stays supported via `compat.py`; migration is
opt-in. PR 13's deprecation warnings are the eventual nudge.
## Surface 2: Realtime runtime (in flight: PRs 7.5–7.6)
`Dreamverse/server/runtime/factory.py` selects a runtime backend at
process start:
```python
def create_runtime_pool() -> RuntimePool:
if os.getenv("FASTVIDEO_REALTIME_BASE_URL"):
return FastVideoRealtimePool(base_url=..., ws_url=..., default_model_id=...)
return GPUPool(get_available_gpus()) # in-process, wraps
# fastvideo.entrypoints.realtime.local_runtime
```
Both backends speak the same `RuntimePool` / `RuntimeSlot` Protocol
(`server/runtime/interfaces.py`):
- `acquire(client_id, websocket=None) -> (gpu_id, RuntimeSlot)`
- `release(client_id)`
- `RuntimeSlot.{join_user, user_step, leave_user, register_stream_queue, ...}`
Today both impls reach into FastVideo-internal's
`fastvideo.entrypoints.realtime.local_runtime` (which exposes
`RealtimeRuntimeConfig`, `GPUPool`, `GPUSlot`). The remote backend talks
HTTP+WS to a separately-deployed runtime of the same shape.
**Contract that PR 7.5/7.6 must preserve:**
- `RealtimeRuntimeConfig` accepts `model_registry`, `default_model_id`,
`default_height/width/num_frames/fps/num_inference_steps/guidance_scale/seed/negative_prompt`,
`default_ltx2_image_crf`, `startup_warmup_{enabled,prompt,timeout_seconds}`.
- `GPUPool(gpu_ids: list[int], config: RealtimeRuntimeConfig)` constructor.
- `pool.initialize() / shutdown() / acquire() / release() / get_status()`.
- HTTP endpoints on the remote variant: `GET /healthz`, `GET /readyz`,
`GET /status`, `WS /ws`. (Already match what
`Dreamverse/server/routes/health.py` consumes.)
**These three health routes still need to migrate into FastVideo's
`build_app` to make `BE_FLAVOR=fastvideo` FE-compatible** — see
[streaming-server.md](streaming-server.md) "build_app route contract" +
[open-threads.md](open-threads.md) follow-up #1.
When PR 7.6 lands the upstream of `fastvideo/entrypoints/realtime/`,
Dreamverse should not need any code change unless the import path
renames. Decided: keep `streaming/` (post-PR-5.5 public name); ship
`realtime/__init__.py` as a re-export with `DeprecationWarning` for one
release cycle.
### Note on `default_ltx2_image_crf`
Dreamverse's `RealtimeRuntimeConfig` includes `default_ltx2_image_crf`.
The April 26 Dreamverse review (D-8) showed this getting passed to
`SamplingParam(...)` and **silently dropped** by the public schema. Post
`d80c2a8` (May 2 typed-config refactor), the migration target is
`request.stage_overrides.refine.image_crf` (per
[design.md](design.md) compatibility mapping table).
**Whether `d80c2a8` actually wired this through, or it's still latent,
is unverified.** See [open-threads.md](open-threads.md) item D-8.
## Surface 3: Continuation state (PR 7)
`Dreamverse/server/video_generation.py:89 ContinuationState` is
Dreamverse's hand-rolled per-session state holder. PR 7 introduced the
typed equivalent at
[`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py).
### Field mapping
| Dreamverse | PR 7 `LTX2ContinuationState` | Notes |
|---|---|---|
| `video_images: list[PIL.Image]` | `video_frames: list[np.ndarray]` (uint8 H×W×3) | numpy is leaner; Dreamverse already round-trips PIL→numpy→PIL just to add noise |
| `audio_latents: torch.Tensor` `[B, C, T, mel]` | `audio_latents: torch.Tensor` (safetensors-serialized; bf16-safe) | unchanged shape; safetensors preserves bf16 |
| `LTX2_VIDEO_CONDITIONING_FRAME_IDX` (env) | `video_conditioning_frame_idx: int` | env constant → per-state field |
| `LTX2_VIDEO_CONDITIONING_STRENGTH` (env) | `video_conditioning_strength: float` | env constant → per-state field |
| `AUDIO_CONDITIONING_NUM_FRAMES` (env) | `audio_conditioning_num_frames: int` | env constant → per-state field |
| `AUDIO_CONDITIONING_STRENGTH` (env) | `audio_conditioning_strength: float` | env constant → per-state field |
| `audio_lps` (passed into `apply_audio`) | `audio_sample_rate: int \| None` | analogous; rename worth confirming with audio team |
| Computed `prefix_sec` per segment | `video_position_offset_sec: float` | **see open question below** |
| `segment_idx` (param to `apply_*`) | `segment_index: int` | per-state field |
| `VIDEO_CONTEXT_NOISE`, `AUDIO_CONTEXT_NOISE`, `ENABLE_AUDIO_COND` | not on state | runtime policy / regularization knobs, not portable session data |
| `apply_video / apply_audio / save_video / save_audio_latents / clear` | not on PR-7 state class | state is a pure data carrier; runtime owns lifecycle policy |
PR 7 is a strict superset of Dreamverse's data model **plus** lifts
several env globals into per-session typed fields.
### Lifecycle mapping
| Dreamverse pattern | `SessionStore` API |
|---|---|
| `self.continuation = ContinuationState()` per session | `state = session_store.snapshot(sid) or LTX2ContinuationState()` |
| `apply_video(req_kwargs, segment_idx)` + `apply_audio(req_kwargs, segment_idx, audio_lps)` | `state = session_store.snapshot(sid)`; runtime builds request from `state.video_frames` / `state.audio_latents` |
| `save_video(frames)` + `save_audio_latents(latents)` | runtime constructs new `LTX2ContinuationState`, `session_store.store(sid, ...)` |
| `clear()` at end of session | `session_store.drop(sid)` |
`SessionStore` and `BlobStore` ABCs ship with thread-safe in-memory
defaults (`InMemorySessionStore`, `InMemoryBlobStore`). Dreamverse can
adopt them as-is for the local runtime; remote runtimes can plug in
redis-backed implementations later.
### Wire format (HTTP/WS round-trip)
Dreamverse's `FastVideoRealtimePool` already speaks the realtime
runtime's HTTP+WS protocol. When PR 7.5/7.6 land state emission on the
server side, the on-the-wire payload is the public envelope:
```json
{
"kind": "ltx2.v1",
"payload": {
"schema_version": 1,
"segment_index": 3,
"video_conditioning_frame_idx": 9,
"video_conditioning_strength": 0.75,
"audio_sample_rate": 24000,
"audio_conditioning_num_frames": 5,
"audio_conditioning_strength": 0.5,
"video_position_offset_sec": 0.2,
"video": {"frames_b64": ["..."]},
"audio": {"safetensors_b64": "..."},
"metadata": {}
}
}
```
JSON-serializable end-to-end; safetensors blob preserves audio dtype
(incl. bf16). For payloads above the inline threshold a `BlobStore`
indirection replaces the b64-encoded body with `{"blob_id": "..."}`;
the blob itself stays inside the runtime that produced it.
## Migration plan per PR
| PR | Dreamverse action |
|---|---|
| PR 6 (landed) | Typed `GeneratorConfig` available; flat-kwarg path still works via compat. Optional migration. |
| PR 7 (landed) | Typed `LTX2ContinuationState` available. ~50-line Dreamverse PR: replace `server/video_generation.py:89` import; move `apply_*`/`save_*`/`clear` off the state class onto `VideoGenerationWorker`; read knobs from typed state instead of env globals; swap `list[PIL.Image]` → `list[np.ndarray]`. |
| PR 7.5 (open) | Streaming server skeleton — Dreamverse's `runtime/factory.py` either keeps building `GPUPool` from `RealtimeRuntimeConfig` (current path), or migrates to `ServeConfig.streaming` shape and invokes `fastvideo serve --config realtime.yaml`. Dreamverse's `RuntimePool`/`RuntimeSlot` Protocol can stay in place. |
| PR 7.6 (branch ready) | GPU pool upstream — `local_runtime.py` import becomes a public import with same symbols (`RealtimeRuntimeConfig`, `GPUPool`, `get_available_gpus`). Per-GPU continuation state inside the worker becomes a `SessionStore` reference (Dreamverse doesn't see this). `request.state` / `result.state` round-trip starts working end-to-end on the local runtime. |
| PR 7.10 (planned) | `generate_async` is canonical. Dreamverse's per-segment `user_step` flow can migrate from sync `generate_video(..., **kwargs)` to consuming the typed event stream. Optional; sync wrapper stays. |
## Dynamo backend contract
**FastVideo does not host any Dynamo code.** The backend package
(`args.py`, `main.py`, `backend.py`, `register.py`, `health_check.py`,
adapter, Dockerfile) lives entirely in the Dynamo repo at
`components/src/dynamo/fastvideo/`, modeled on
`components/src/dynamo/sglang/`.
FastVideo's only obligation is to expose a stable, typed Python API
that Dynamo's backend package imports.
### Contract surface
| Surface | Exposed as |
|---|---|
| Construction | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` (typed_kwargs = a stable subset from `GeneratorConfig`; no flat LTX2 legacy) |
| Sync execution | `generator.generate_video(request: GenerationRequest) -> VideoResult` |
| Async execution | `generator.generate_async(request: GenerationRequest) -> AsyncGenerator[VideoEvent, None]` (PR 7.10) |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` |
| Typed result | `VideoResult` with `video_bytes` or tensor frames + optional `ContinuationState` |
| Continuation | `ContinuationState(kind, payload)` — schema-versioned payloads |
| Health-check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` (256x256 / 8 frames / 1 step) |
| Config dump | `GeneratorConfig.to_dict()` / `ServeConfig.to_dict()` |
### Request/response mapping (Dynamo ↔ FastVideo)
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> (seconds * nvext.fps) -> sampling.num_frames
input_reference -> input.image_path / input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
video_url (after upload) -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
```
All fields exist on FastVideo's typed schema after PR 6 expansion (typed
LTX2 kwargs) + PR 7.10 (`generate_async` + health check).
### Reference: PR ai-dynamo/dynamo#7544
Closed draft establishing the Dynamo backend shape. Two frictions
identified:
1. Flat legacy LTX2 kwargs — solved by PR 6.
2. Sync-only generation — solved by PR 7.10's `generate_async`.
Next iteration of this PR (or its successor) will reopen against PR 8's
docs reference and land cleanly.
## Open questions across surfaces
| # | Question | Source | Status |
|---|---|---|---|
| Q-1 | Multi-model GPU pool | dreamverse_review D-1 | Deferred (production single-model) |
| Q-2 | LTX-2 prompt orchestration promotion to public | dreamverse_review D-2 | Open; consumer-side until 2nd consumer |
| Q-3 | Race-based provider fallback | dreamverse_review D-3 | Open; sequential is current public |
| Q-4 | Router upstream skip on Dreamverse | dreamverse_review D-4 | Resolved (PR 7.9 lands publicly, Dreamverse doesn't consume) |
| Q-5 / D-5 | `generate_async` cutover (audio re-encode) | dreamverse_review | **Blocked on PR 7.10** |
| D-6 | Don't upstream `realtime/local_runtime.py` | dreamverse_review | Resolved (Dreamverse switches to `streaming.gpu_pool.SubprocessGpuPool`) |
| D-7 / Q-6 | FP4Config public colocation | dreamverse_review | **Resolved May 2** — public NVFP4 landed with lazy flashinfer |
| D-8 | `ltx2_image_crf` silently dropped | dreamverse_review | **Unverified post-`d80c2a8`** — see [open-threads.md](open-threads.md) |
| D-9 | `aarch64-conda-linux-gnu-cc` triton compile failure | dreamverse_review | Operational; `ENABLE_TORCH_COMPILE=0` workaround |
| D-10 | Warmup OOM on shared GPU | dreamverse_review | Operational; idle-GPU pre-warm probe |
| D-11 | ffmpeg `Broken pipe` on disconnect | dreamverse_review | Cosmetic logging cleanup |
| — | `video_position_offset_sec` semantics (persistent vs per-segment) | dreamverse_integration | **Open — needs decision before PR 7.6 emits state** |
| — | `SessionStore` / `BlobStore` lifecycle (TTL/eviction/blob-drop) | dreamverse_integration | Open — defer to PR 7.5 design pass |
See [decisions-log.md](decisions-log.md) for full rationale per
decision.
## Don't / Cautions
- **Don't pop the Dreamverse stash on this branch.** It's 3867 lines of
orphan modular refactor with broken absolute imports.
- **Don't change `RealtimeRuntimeConfig` shape without coordinating
with Dreamverse `runtime/factory.py`.**
- **Don't promise public compatibility for private Dreamverse-only
field aliases.** Those belong in the private adapter layer per design
spec.
@@ -0,0 +1,710 @@
# Decisions Log — D + Q Resolutions
Cross-doc consolidated decision log. Each entry: ID, source doc,
question/decision, rationale, current status.
For implementation status see [pr-roadmap.md](pr-roadmap.md). For
follow-up actions see [open-threads.md](open-threads.md).
**Last updated:** 2026-05-05 (added D-12 — GpuPool layer separation, Oracle review post-#1257-merge; added D-13 — prompt enhancer / LLMProvider abstraction shape, Oracle review pre-#1258-merge; added D-14 — streaming auxiliaries cohesion, Oracle review during #1284 review cycle; added D-15 — streaming router placement + sticky/active-active deferral, Oracle review during #1286 review cycle; added D-16 — streaming router polish round 2, second-pass review on top of D-15 covering bridge cancellation hygiene, registry state machine, httpx hard-fail, replica YAML parsing, and `websockets` dep; added D-17 — strategy reversal: abandon 6-PR split in favor of single mega-PR #1288 on `will/ltx2_sr_port`; added D-18 — Option B+ chosen: Dreamverse FE+product-server move into FastVideo as `apps/dreamverse/` subfolder while generic backend stays at `fastvideo.entrypoints.streaming.*`; integration-review.md deprecated, integration-plan.md is the executable migration plan).
## Status legend
- ✅ **Resolved** — decision made and implementation complete (or no implementation needed)
- 🟡 **Deferred** — decision made, implementation deferred to a known PR
- 🔴 **Open** — needs decision
## Post-merge architecture decisions
### D-18: Option B+ — Dreamverse becomes `apps/dreamverse/` subfolder under FastVideo
**Status:** ✅ Resolved 2026-05-05. [integration-plan.md](integration-plan.md) is the executable migration plan; [integration-review.md](integration-review.md) is deprecated but kept for drift audit + OSS precedents.
**Source:** User decision after reviewing [integration-review.md](integration-review.md)'s Option D recommendation.
**Question:** [integration-review.md](integration-review.md) recommended **Option D** — Dreamverse stays a separate repo, generic backend (streaming runtime, GPU pool, prompt enhancer, router) merges into `fastvideo.entrypoints.streaming.*`. The user reviewed this and chose a different shape: keep the generic-backend principle from Option D but ALSO move the Dreamverse FE + product server into FastVideo as a subfolder (`apps/dreamverse/`). Combination is "Option B+" (Option B layout with Option D's backend principle).
**Decision:** Option B+. Concrete shape:
- **One repo**: `hao-ai-lab/FastVideo`. Dreamverse repo gets archived after migration completes.
- **Python ML library** stays at root: `fastvideo/`, `fastvideo-kernel/`.
- **Generic backend** stays at `fastvideo.entrypoints.streaming.*` (already there per #1257/#1258/#1284/#1286/#1288).
- **Dreamverse product** moves into `apps/dreamverse/{server,web,prompts,serve_configs,scripts}/`.
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone pnpm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
**Rationale:**
- Drops the cross-repo coordination overhead identified in the post-#1286 rebase cycle (D-17 handled by consolidating into mega-PR; D-18 prevents the next round of cross-repo coordination from happening).
- Keeps the architectural separation Option D recommended (FastVideo owns reusable runtime; product owns product). The boundary is now `apps/dreamverse/` directory rather than two repos.
- Single repo means atomic cross-cutting refactors (e.g. GpuPool API change + Dreamverse adoption) ship as one PR.
- OSS precedents support the shape (chainlit uv-workspace + pnpm; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
**Why not Option D (separate repos):**
- Each upstream merge into FastVideo invalidates Dreamverse's lockfile/imports; the post-#1286 rebase showed this requires coordination overhead that scales with feature velocity.
- Cross-repo contract tests catch shape drift but not behavior drift.
- Two repos means two `AGENTS.md`, two CI configs, two release stories, two Dependabot dashboards.
**Why not Option C (full merge into `fastvideo.entrypoints.dreamverse.*`):**
- Forces FastVideo to ship Tailwind config + curated preset JSON + Next.js build artifacts.
- Locks Dreamverse product cadence to FastVideo PyPI releases.
- Librarian: "no 1:1 precedent for Python ML library + Next.js product merged into library namespace" — argues against this.
**Why not Option B (subfolder, but generic backend folded into `apps/dreamverse/server/`):**
- Other consumers (Dynamo, future streaming clients) need the backend without the Dreamverse product. Folding the backend under `apps/dreamverse/server/` would force Dynamo to either depend on `apps/` paths (ugly) or carry a fork.
**Implications:**
- [integration-review.md](integration-review.md) is **deprecated** (banner header + reading-guide demotion). Kept in tree for drift audit + OSS precedent reference.
- [integration-plan.md](integration-plan.md) is the **canonical executable plan** with 7 phases (Phase 0: land #1288; Phase 1: skeleton + tooling; Phase 2: backend move; Phase 3: FE move; Phase 4: promote generic-pending; Phase 5: prompt enhancer fork retirement; Phase 6: CI/release cutover; Phase 7: archive Dreamverse repo).
- Dreamverse repo will be **archived** at end of Phase 7 — not before.
- Dreamverse history does NOT migrate cross-repo via `git mv` (technical limitation); original history stays in archived Dreamverse repo, and Phase 2 PR body records the source SHA(s).
- New top-level `apps/` directory created — must be excluded from FastVideo PyPI wheel via `[tool.setuptools.packages.find] exclude = ["apps*", ...]`.
- Drift items from [integration-review.md](integration-review.md) get folded into specific phases of [integration-plan.md](integration-plan.md) (e.g. health routes → Phase 4, DR-1 → Phase 5).
**Open questions deferred to phase planning:**
- DR-2 (`cerebras_ifm`): public Literal vs Dreamverse-side custom provider — decide before Phase 5.
- VPO (`video_position_offset_sec` semantics): persistent vs per-segment — decide in Phase 4.
- Cross-repo history: fresh import vs `git subtree` import — decide before Phase 2.
- CORS / write-endpoint security policy: dev-only vs auth vs firewall — decide before Phase 6.
### D-17: Abandon 6-PR split — land everything as single mega-PR #1288
**Status:** ✅ Resolved 2026-05-05. PR #1287 closed; PR #1288 opened on `will/ltx2_sr_port` covering the full chain.
**Source:** User decision after observing the post-#1286 rebase + re-slice cycle.
**Question:** The original plan ([STACK.md](../../../STACK.md), [pr-roadmap.md](pr-roadmap.md)) called for the remaining `will/ltx2_sr_port` content (after PRs 7.5/7.6/7.7/7.8/7.9 landed) to ship as 6 stacked PRs: 7.10 (#1287, generate_async), 8 (server contract docs), LTX-2 SR runtime, NVFP4, post-fixes, agents-cleanup. PR #1287 was opened on 2026-05-05 as the first slice. Should the remaining 5 slices be opened sequentially as planned, or should everything be consolidated into one PR?
**Decision:** Consolidate. Close #1287; open one mega-PR (#1288) on `will/ltx2_sr_port` covering all 34 commits / 71 files / +13,074 LOC at once.
**Rationale:**
- The post-#1286 rebase + re-slice cycle exposed real overhead: backup branch, interactive rebase with manual `drop` directives, force-push, re-slice 6 bookmarks, push next slice as new remote, open new PR, update memory dir. Repeating that 6 more times for the remaining slices accumulates substantial review-coordination overhead with diminishing structural benefit.
- The 6 layers are not independent in the way that landed PRs 7.5-7.9 were. PR 7.10 (`generate_async`) is the only API-shape change; PR 8 is docs+tests on top; LTX-2 SR / NVFP4 / post-fixes / agents-cleanup are feature/fix/docs work that doesn't shape the public API. Reviewing them as one ordered diff is at least as easy as reviewing 6 stacked PRs whose dependencies must be tracked manually.
- Single PR keeps CI / merge queue simpler and avoids the 6-PR cascade where every upstream merge invalidates the chain below it.
**Implications:**
- [STACK.md](../../../STACK.md) (top-level, 10-PR split tracker) is **deprecated**. Kept in tree as a historical artifact with the merged half (PRs 1-4 of the 10) accurate. Safe to delete in a follow-up.
- [authors.md](authors.md), [`CO-AUTHORS.md`](../../../CO-AUTHORS.md) — co-author roster is unchanged; trailers still apply per-commit on every commit in the consolidated PR.
- [runbook.md](runbook.md) — "After a PR merges (re-slice protocol)" section replaced by a simpler "After PR #1288 merges" section.
- Local split bookmarks (`will/api_7.10`, `will/api_8`, `will/ltx2_sr_runtime`, `will/ltx2_nvfp4`, `will/ltx2_post_fixes`, `will/agents_cleanup`) are no longer maintained; safe to delete locally.
- `origin/will/api_7.10` — pushed during the #1287 cycle; can be deleted on origin once #1287 close-cleanup completes.
**Watch outs:**
- The PR is large (71 files, +13,074 LOC). Reviewers will need commit-by-commit review; the PR body structures the layers in commit order to make this tractable.
- If #1288 becomes too large to merge cleanly later (e.g. main moves significantly underneath it), the fallback is to re-split — but the current expectation is to land it as-is.
### D-12: `GpuPool` layer separation — keep distinct from `VideoGenerator`
**Status:** ✅ Resolved (interim) + 🟡 Deferred long-term shape to PR 7.10.
**Source:** Oracle review on 2026-05-04, post-PR-#1257 merge.
**Question:** Should `fastvideo.entrypoints.streaming.GpuPool` (PR #1257) be
folded into `fastvideo.entrypoints.video_generator.VideoGenerator`, or kept
separate? Three alternatives were evaluated:
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `VideoGenerator` (single inference call) and `GpuPool` (multi-session orchestration) stay separate | ✅ Correct as **interim** |
| B | `VideoGenerator` absorbs the pool's role (`from_pretrained_pool`, `acquire/release/run`) | ❌ **Wrong layer.** Conflates execution with serving scheduler. |
| C | `GpuPool` becomes a thin **session-aware async executor** over PR 7.10's `generate_async` | ✅ Correct **long-term destination** |
**Decision:** Alt A as interim; evolve toward Alt C once PR 7.10 lands
`generate_async`. Do NOT pursue Alt B.
**Rationale:**
- `VideoGenerator` is a library handle — "execute one request, possibly
across ranks via `MultiprocExecutor`/`RayDistributedExecutor`."
- `GpuPool` is serving infrastructure — "schedule N concurrent sessions
across N independent replicas, with sticky session-to-GPU affinity for
cache locality."
- These are different layers driven by different consumers (a Python
script doing `gen.generate(req)` vs. a WebSocket server with sticky
sessions). Folding them muddies both surfaces.
**Key finding — `MultiprocExecutor` and `SubprocessGpuPool` are orthogonal,
not redundant:**
| Layer | Job | Granularity |
|---|---|---|
| `MultiprocExecutor` (`fastvideo/worker/`) | TP/SP shard ONE inference call across N GPU ranks | per-call |
| `streaming_generator.py` (existing real-time path) | Per-frame streaming via `MultiprocExecutor.submit_step`/`get_result` | per-step within one generator |
| `SubprocessGpuPool` (`entrypoints/streaming/`, PR #1257) | Serve N concurrent sessions on N replicas, sticky-bound | per-session |
Both spawn subprocesses because **CUDA contexts demand process boundaries**,
not because they solve the same problem. Sharing low-level lifecycle
utilities (process spawn, queue plumbing, shutdown) is a future refactor;
unifying the abstractions is wrong.
**Sticky binding stays in the pool, NOT in `VideoGenerator`:** sticky
session-to-GPU affinity is a serving policy driven by LTX-2's per-GPU
continuation cache (last-9-decoded-frames + audio-latents). Different
consumers want different policies — stateless OpenAI HTTP wants
per-request leasing; LTX-2 streaming wants sticky affinity; per-frame
real-time streaming wants a continuous queue. Keeping policy in the pool
keeps `VideoGenerator` policy-free.
**Specific risks flagged in PR #1257 (already merged):**
| Risk | Mitigation (when relevant) |
|---|---|
| `GpuPool.run() -> Any` is sync — fine for whole-segment dispatch, blocks on cancellation | Replace with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle (`generate_async` makes this trivial) |
| `PoolAssignment.gpu_id: int` assumes one-GPU-per-worker | Don't lock as public API. Future may need `device_ids: list[int]` for topology-aware pooling (one worker = group of GPUs running internal `MultiprocExecutor`) |
| `GpuPool` could be documented as the canonical FastVideo serving API | Mark as **experimental / server-internal** in docstring until PR 7.10 lands. Don't include in user-facing API docs yet |
| Memory: N processes = N model replicas (~10-50 GB each) | Expected for concurrent serving with crash isolation. CUDA IPC weight sharing loses isolation; CPU-shared-memory loading helps host RAM not device. Real scalable path is topology-aware pooling later. |
**Action items (carried into post-7.10 cycle):**
- [ ] Update `GpuPool` ABC docstring to note "API may change post-PR-7.10"
- [ ] Plan to replace `run()` with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle
- [ ] Don't promote `gpu_id: int` to public API; revisit shape post-7.10
- [ ] Consider clarifying field naming (e.g. `worker_id` is the stable identifier; `gpu_id` is current-impl detail)
- [ ] When opening 7.10's PR, have it consume `generate_async` from `GpuPool.run_async` end-to-end
**Open thread it touches:** PR 7.10 (`open-threads.md` item D — generate_async)
unblocks Alt C and is the natural place to land the API shape change.
### D-15: Streaming router (PR #1286) — keep in-repo, defer sticky / active-active
**Status:** ✅ Resolved (interim). Pre-merge polishes applied. Three follow-up
items tracked.
**Source:** Oracle review on 2026-05-05, during PR #1286 review cycle.
**Question:** Where should the multi-replica WebSocket router live? Should it
ship at all (vs. delegating to nginx/envoy)? Should sticky session routing
or weighted/round-robin balancing be in the initial PR?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `fastvideo/entrypoints/streaming/router/`, FastAPI-based, single-primary failover, lazy `httpx`/`websockets` imports | ✅ **Keep** |
| B | Move to separate package `fastvideo-router/` | ❌ **Premature** — adds packaging/release/compat overhead before evidence of independent adoption |
| C | Fold router into the streaming server itself (one app, mode flag) | ❌ Conflates router/generator lifecycles, mode-dependent config, drags inference deps into routing deployments |
| D | Replace with reverse proxy (nginx/envoy/HAProxy) recipes | ❌ Not as the SOLE answer — mature proxies don't naturally emit FastVideo typed `gpu_unavailable` frames or evolve with FastVideo session semantics. Recommend external proxies as a complement at high scale. |
| E | Add sticky session routing now | ❌ **Defer** — implementing correctly depends on where `session_id` is available (URL/header is easy, first JSON frame is invasive). Reconnects are rare today. |
| F | Add weighted / round-robin now | ❌ **Defer** — active-active without sticky routing is worse for LTX-2 continuation locality than active-passive failover |
**Decision:** Alt A — keep current shape. Apply pre-merge polishes; preserve
forward-compat for sticky routing.
**Rationale:**
- Python router is justified as a FastVideo-aware control-plane component,
not a replacement for Envoy/HAProxy. It can emit typed
`gpu_unavailable` frames, evolve with FastVideo session semantics,
and ship local/dev deployment without ceremony.
- The current abstraction is small + testable: `RouterConfig`,
`ReplicaRegistry`, `ReplicaStatus`, `HttpProbe` (Protocol/structural alias).
Adding strategy registries / telemetry interfaces / active-active policies
now would be over-engineering.
- Active-passive (single primary) is the right MVP for LTX-2 streaming —
preserves continuation cache locality (D-12 sticky binding rationale)
better than naive active-active.
- The biggest architectural risk isn't placement; it's accidentally baking
in unstated semantics. Define single-primary behavior + config validation
now so future active-active or sticky routing becomes additive.
**Pre-merge polishes applied (per gemini + Oracle review):**
| # | What | Why |
|---|---|---|
| 1 | `ReplicaRegistry.select()` docstring rewrite | gemini flagged "round-robin via insertion order" claim was misleading — implementation always returns `[0]`. Replaced with explicit "first healthy primary, else first healthy non-primary; this MVP picks first match within tier; round-robin/weighted deferred". |
| 2 | Refactored `run_health_check_loop` to share single `httpx.AsyncClient` across the loop's lifetime via `_build_default_probe()` async context manager | gemini flagged per-probe client instantiation as inefficient. With ~1 probe/second default polling, TCP/TLS handshake overhead is non-trivial; now reuses connection. Tests inject probes directly so the path stays bypassable. |
| 3 | Probe all replicas concurrently per cycle via `asyncio.gather(..., return_exceptions=True)` | gemini flagged sequential probes risk falling behind `health_check_interval_seconds` if replicas time out. Now per-cycle wall time = max(probe latencies), not sum. |
| 4 | `RouterConfig.__post_init__` validation | Oracle recommended: empty replicas, non-positive intervals/timeouts, thresholds < 1, non-`http(s)://` URLs, and >1 primary all `raise ValueError`. Surfaces misconfiguration at config-load instead of confusing runtime failures. |
| 5 | Migrated `@app.on_event("startup"/"shutdown")` to `@contextlib.asynccontextmanager`-based `_lifespan()` | Pre-merge — FastAPI deprecated the old API. Was tracked as the 7.9 caveat in pr-roadmap.md. |
**One review comment intentionally not implemented:**
| Comment | Decision |
|---|---|
| gemini medium: `_load_router_config` duplicates `fastvideo.api.parser.parse_config` logic | Kept manual flat-from-nested mapping. The YAML schema has nested `health_check:` block but `RouterConfig` is flat; using `parse_config` directly would require either restructuring `RouterConfig` to have a nested `HealthCheckConfig` (schema change beyond this PR's scope) or accepting incomplete parsing. Manual mapping is intentional and well-typed. |
All 4 review threads marked resolved on the GitHub PR.
**Action items (deferred):**
- [ ] Track sticky session routing extensibility — when needed, add
`ReplicaRegistry.select(routing_key: str | None = None)` so registry
evolution is additive; document upfront where `session_id` should
appear (URL/header preferred over first JSON frame to avoid
buffering/peeking)
- [ ] Track `_bridge_session()` backpressure note — fine for MVP because
`websockets` library provides basic transport backpressure, but at
high scale add max_size/timeouts or recommend Envoy/HAProxy in front
- [ ] If active-active multi-primary becomes a requirement, define
behavior (round-robin within healthy primaries, weighted, sticky-by-key)
rather than letting `select()` silently pick `[0]`
**Watch outs:**
- `session_id` in WebSocket URL/headers is the cleanest sticky-routing
hook. If it ends up only in the first JSON message, sticky routing
later will require buffering/peeking before backend selection.
- Multi-primary configs are now explicitly rejected by validation;
documented + enforced.
- `_bridge_session()` is fine for MVP (the libraries provide basic
backpressure), but not production-grade for edge load. Document the
limit.
**Open thread it touches:** open-threads.md items #13 (sticky routing),
#14 (bridge backpressure), #15 (multi-primary semantics).
### D-16: Streaming router polish round 2 — second-pass fixes on top of D-15
**Status:** ✅ Resolved. Applied as `[fix] streaming: router polish — bridge
cancel + state machine + deps` (`a152cb77` on `will/api_7.9`, `40e265b8` on
`will/ltx2_sr_port`).
**Source:** Second-pass review on PR #1286, 2026-05-05, after D-15's pre-merge
polishes landed.
**Question:** D-15 closed the structural review (placement, sticky/active-active
deferral, basic `__post_init__` validation). On a second pass through the same
files, five latent issues surfaced that weren't covered by gemini's first pass
or Oracle's structural review. Apply them on top of the merged D-15 polishes,
or queue for a follow-up PR?
**Decision:** Apply on top of `will/api_7.9` directly. All five are bug-class
or DX-class — none are scope-expanding architecture changes — so folding them
into PR #1286 keeps the router landing in one reviewable unit instead of
shipping a router PR plus an immediate follow-up fix PR.
**Fixes applied:**
| # | File | What | Why |
|---|---|---|---|
| 1 | `router/main.py::_bridge_session` | Replaced `asyncio.gather()` with `wait(FIRST_COMPLETED)` + explicit `cancel()`/drain + `_is_normal_disconnect()` classifier | `gather` waited for both directions; on client disconnect, the backend-reader task leaked and stayed pending. Backend `ConnectionClosed` also surfaced as an unhandled exception in server logs. New shape: first task to finish triggers explicit cancel of the other, both are drained, and only non-routine exceptions re-raise. |
| 2 | `router/registry.py::record_success` | Split state transitions: `UNKNOWN -> HEALTHY` is now immediate on first successful probe; only `UNHEALTHY -> HEALTHY` remains gated by `recovery_threshold` | Previously a fresh registry needed `recovery_threshold` consecutive successes before any replica was selectable. With default `recovery_threshold=2` and `health_check_interval=1s`, that meant 2-3s of `gpu_unavailable` rejections at startup. Now the first probe promotes immediately; recovery gating still protects against flapping replicas. |
| 3 | `router/registry.py::_build_default_probe` | Missing `httpx` now raises `RuntimeError` with install hint instead of yielding a "disabled" probe stub | Previous behavior: silently returned `(0.0, "httpx not installed; ...")` for every probe, which `record_failure` then folded into `UNHEALTHY` after `failure_threshold` cycles. Operators saw replicas drop UNHEALTHY with a confusing reason and no clear remediation. Hard-fail at startup is the right surface. |
| 4 | `router/config.py::__post_init__` | Extended D-15 polish #4 with: rejects `urlparse(url).path not in ("", "/")`, rejects `query`/`fragment`, rejects duplicate URLs across replicas | D-15's validation rejected non-`http(s)://` URLs and >1 primary; it didn't catch `http://host/api` (the router appends `/health` and `/v1/stream` itself, so a base-URL with path yields malformed routes) or `[{url: x}, {url: x}]` (replica registry keys by URL — duplicates would silently collapse to one entry, masking the misconfiguration). |
| 5 | `cli/router_serve.py::_load_router_config` | Replaced silent list-comprehension filter (`for r in replicas_raw if isinstance(r, dict) and r.get("url")`) with per-index `raise ValueError` | Original parser silently dropped malformed YAML entries. A single typo in `replicas[2].url` would yield 2 replicas instead of 3 with no log line. New shape: explicit per-index error message ("missing required key 'url'", "must be a mapping"). |
| 6 | `pyproject.toml::[streaming]` extra | Added `websockets` as explicit dep | `router/main.py::_bridge_session` does `import websockets` lazily and raises `RuntimeError` if missing. The `[streaming]` extra was an implicit transitive — anyone installing only `[streaming]` (and not the broader requirements) hit the runtime error. Now explicit. |
**Tests added (7 cases in `fastvideo/tests/entrypoints/streaming/test_router.py`):**
- `TestUnknownToHealthyImmediate.test_first_success_promotes_unknown` — first probe success transitions `UNKNOWN -> HEALTHY` regardless of `recovery_threshold`
- `TestUnknownToHealthyImmediate.test_unhealthy_recovery_still_gated_by_threshold` — `UNHEALTHY -> HEALTHY` still requires `recovery_threshold` successes
- `TestConfigValidation.test_rejects_path_in_url` / `test_rejects_query_in_url` / `test_rejects_fragment_in_url` / `test_rejects_duplicate_urls` / `test_accepts_trailing_slash` — `__post_init__` URL validation matrix
**Verification:** 17/17 router tests pass on both branches. `pre-commit run`
clean (yapf / ruff / codespell / mypy). `lsp_diagnostics` clean on changed
regions; the one pre-existing `Task` generic-type warning at `main.py:37` is
unrelated and predates this commit.
**In-flight pre-commit corrections (not part of the 6 fixes themselves):**
- yapf auto-reformatted 4 files (kept verbatim).
- ruff `UP038`: rewrote `isinstance(exc, (CancelledError, WebSocketDisconnect))`
to `isinstance(exc, CancelledError | WebSocketDisconnect)`.
- mypy `[misc]`: renamed loop var `exc` (inside `for task in done`) to
`task_exc` to avoid name collision with the outer
`except ImportError as exc` binding.
**Open thread it touches:** None new. Item #14 (bridge backpressure) and
item #13 (sticky routing) from D-15 remain deferred — this round addressed
**cancellation/disconnect** semantics on the bridge, which is distinct from
**throughput backpressure**. Item #14 still applies: at higher load, add
`_bridge_session()` max-size + timeout limits or recommend Envoy/HAProxy
in front.
### D-14: Streaming auxiliaries (PR #1284) — cohesion + concrete-vs-Protocol scoping
**Status:** ✅ Resolved (interim). Two polish items applied during review; one
operational caveat tracked.
**Source:** Oracle review on 2026-05-04, during PR #1284 review cycle.
**Question:** Is PR #1284's bundle of 4 streaming-server auxiliary modules
(`prompt/safety.py`, `prompt/rewrite.py`, `session_logger.py`,
`mock_server.py`) correctly scoped? Should `mock_server` live in production
module path? Should `PromptSafetyFilter` be a Protocol? Should the bundle
have been split into 4 PRs?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — single PR, 4 modules under `streaming/`, mock_server in production path, concrete safety filter | ✅ **Keep** |
| B | Split into 4 separate PRs | ❌ Process overhead, not architectural improvement |
| C | Move `mock_server.py` into `tests/` | ❌ Would reduce discoverability + install-time usability of `python -m fastvideo.entrypoints.streaming.mock_server` |
| D | Move `session_logger.py` to `streaming/observability/` (or top-level `fastvideo/observability/`) | ❌ Premature — currently session-shaped + streaming-specific; promote when a non-streaming consumer appears |
| E | Convert `PromptSafetyFilter` to Protocol (like `LLMProvider`) | ❌ Premature abstraction — only one classifier exists; small duck-typed surface preserves future Protocol introduction without breaking the concrete |
| F | Convert `MockGenerator` to Protocol | ❌ Same — small duck-typed surface; no second mock generator exists |
**Decision:** Alt A — keep current shape. Apply two polish items from
Oracle's review before merge.
**Rationale:**
- "Streaming-server auxiliaries" is cohesive enough at 730 LOC with
isolated modules + tests. Each module has independent code path but
shared deployment context (the streaming server boots them all).
- `mock_server.py` in production path is a strength: reuses
`build_app()` for protocol parity. Hiding it under `tests/` would lose
`python -m fastvideo.entrypoints.streaming.mock_server` CLI access for
FE devs.
- Concrete `PromptSafetyFilter` matches "ship what we have, abstract
later" pattern. Internal had multi-classifier composition; public
ships single + leaves chaining as a Dreamverse-side concern (per D-2).
- Same pattern for `MockGenerator`: small duck-typed `_GeneratorLike`
surface lets a second mock implementation drop in without inheritance.
- `threading.Lock` (not `asyncio.Lock`) in `session_logger.py` is
correct — writes come from real encoder/control threads via
`run_in_executor`, not from coroutines directly. `asyncio.Lock` would
be the wrong primitive for cross-thread concurrency.
**Pre-merge polishes applied (per Oracle):**
| Polish | What | Why |
|---|---|---|
| 1 | Removed `RewriteOptions.user_system_prompt_override` | Inert public field — was declared but never threaded through to `enhancer.rewrite()`. Shipping unused public options is more likely to bite than any structural choice. Re-add when actually wired through. |
| 2 | Sanitized `session_id` filename in `session_logger.SessionLogger._get_file()` | Defense-in-depth: today session_id is server-generated UUID, but a future code path that accepts client-supplied ids would otherwise allow path traversal via `../`. Added `_FILENAME_SANITIZE_RE = re.compile(r"[^A-Za-z0-9._-]")` + sub before `os.path.join`. |
**Operational caveat tracked (not a code change):**
- `SafetyDecision.UNAVAILABLE` is treated as `ALLOW` by callers — a
policy choice that's correct for an opt-in safety filter, but
callers should log loudly so operators know the filter is degraded.
Tracked as open-threads.md item #12.
**Pre-merge review feedback (4 of 4 resolved on the GitHub PR):**
| # | File:Line | Severity | Issue | Fix applied |
|---|---|---|---|---|
| 1 | `session_logger.py:57` | High | `log()` race vs `close()` — `KeyError` on `_locks[session_id]` | Atomic capture in `_get_file()`; master `_registry_lock`; `with lock, contextlib.suppress(ValueError):` |
| 2 | `rewrite.py:71` | Medium | `re.compile()` in hot path | Module-level `_LEADING_MARKER_RE`, top-level `import re` |
| 3 | `safety.py:105` | Medium | `_ensure_loaded()` race on concurrent fastText load | `_load_lock = threading.Lock()` + double-check pattern |
| 4 | `pyproject.toml:145` | Medium | `streaming` extra missing `prompt-safety` | Added to aggregator |
All 4 review threads marked resolved via GraphQL `resolveReviewThread`.
**Action items (deferred):**
- [ ] Track `SafetyDecision.UNAVAILABLE` log loudness in
open-threads.md item #12 — when streaming server starts using the
safety filter, ensure operator-visible logging on `UNAVAILABLE`
results
- [ ] If a second safety classifier appears (Perspective API, Detoxify,
custom rules), promote `PromptSafetyFilter` to a Protocol — same
pattern as `LLMProvider` per D-13
- [ ] If a second mock generator appears (different frame patterns,
different latency models), promote `MockGenerator` to a Protocol
**Open thread it touches:** PR #1284 itself; future safety-classifier
Protocol promotion; future observability module extraction.
### D-13: Prompt enhancer / `LLMProvider` abstraction shape — keep streaming-scoped
**Status:** ✅ Resolved (interim) + 🟡 Three deferred polishes after metrics or 2nd consumer.
**Source:** Oracle review on 2026-05-04, pre-PR-#1258-merge.
**Question:** Is PR #1258's `fastvideo.entrypoints.streaming.prompt.*` module
correctly designed? Should it be (a) Protocol-based vs ABC, (b) under
`streaming/` vs top-level `fastvideo.prompt.*`, (c) closed 3-op enum vs
open `complete()` API?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `streaming/prompt/*`, Protocol provider, fixed 3 ops, lazy `httpx`, per-call `AsyncClient` | ✅ **Keep** |
| B | Move to top-level `fastvideo.prompt.*` (decouple from streaming) | ❌ **Premature.** No second consumer exists yet. |
| C | Convert `LLMProvider` Protocol → ABC with default impls + retry classification | ❌ **Wrong direction.** Biases extension toward OpenAI shape; `_openai_compat.py` already factors that as helper not inheritance. |
**Decision:** Alt A as interim. Promote to Alt B only when a second
non-streaming consumer (OpenAI server, batch generation, tooling) actually
needs the prompt enhancer. Don't pursue Alt C.
**Rationale:**
- Public contract is tiny — `name: str` + `async complete(LLMRequest) -> LLMResponse`. ABC adds zero value.
- `_openai_compat.py` is the right place for shared logic — helper, not base class. Anthropic / local / custom providers stay first-class.
- The 3 ops (enhance / auto_extend / rewrite) are LTX-2 streaming concepts. `auto_extend` (continue prompt sequence) and `rewrite` (multi-line alternatives) come directly from session UX. Calling this "the FastVideo prompt API" misrepresents that.
**Specific risks flagged in PR #1258 (already merged-pending review):**
| Risk | Mitigation (when relevant) |
|---|---|
| API publicity — calling this "the FastVideo prompt API" before a second consumer exists | Document module as "streaming-server prompt enhancement" in user-facing docs; keep it nested under `entrypoints/streaming/` |
| `httpx.AsyncClient` per-call (no connection pooling) | Acceptable for ~6-10 calls per LTX-2 session; LLM latency dominates. Add optional `client_factory` parameter LATER if metrics show connect overhead is meaningful. |
| 3 fixed operations could constrain future generic use | Closed enum is right for application-level orchestration. Future generic consumers should either call `provider.complete()` directly, or get a thin separate enhancer that shares the provider/fallback machinery. |
| `register_provider(priority=-1)` semantics rely on Python's negative-index `list.insert` | Cosmetic concern; docstring is clear. Could be tightened to explicit branch later. |
| `runtime_checkable` Protocol with `name: str` instance attribute — static type checkers may miss missing `name` | Acceptable; runtime check via `isinstance(p, LLMProvider)` works for plugin discovery. |
**Action items (deferred):**
- [ ] Document `fastvideo.entrypoints.streaming.prompt.*` as streaming-scoped in user-facing docs (PR 12 docs migration); avoid promoting as framework-level
- [ ] Add optional `client_factory` parameter to providers when metrics justify pooling
- [ ] Plan future move to `fastvideo.prompt.*` (with import shim) when second non-streaming consumer materializes
- [ ] Track Q-2 reactivation: promote LTX-2 prompt orchestration (locked segments, segment-prompts JSON parsing) to public `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` when a second LTX-2-style consumer appears
**Open thread it touches:** Dreamverse migration (open-threads.md DR-1)
will be the first real test of the public surface. Lessons learned there
inform whether Alt B becomes feasible.
## D-decisions (from `dreamverse_review.md`, Apr 26)
### D-1: Realtime runtime → streaming GpuPool migration shape
**Status:** ✅ Resolved.
Internal `RealtimeRuntimeConfig` had a multi-model registry +
flattened sampling defaults. Public `SubprocessGpuPool` is single-model
+ uses per-request `SamplingConfig`.
**Decision:** Drop multi-model registry on integration branch (not used
in production). Construct `GeneratorConfig` for chosen model and pass to
`SubprocessGpuPool`. Move sampling defaults to a server-side
`default_request: GenerationRequest` template.
**Risk:** Migration branch surfaces missing-model errors if a flow
silently relied on registry to swap models per-session. Integration
tests exercise at least one segment per supported model id before
merging.
### D-2: PR 7.7 prompt enhancer API surface narrower than internal
**Status:** ✅ Resolved.
Public `PromptEnhancer.enhance/auto_extend/rewrite` returns
`LLMResponse(content, provider, model, latency_ms, fallback_used)`.
Internal returns `EnhanceResult(prompt, fallback_used, error, ...)` /
`RewriteResult(prompts, ..., rollout_id, rollout_label, ...)`.
**Decision:** Adapt at the call site via
`Dreamverse/server/prompting/_internal_compat.py` shim. Locked-segment /
next-segment-index plumbing stays Dreamverse-side. Public stays minimal
and provider-agnostic.
**Open question (Q-2):** Promote LTX-2-specific orchestration into
`fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` once a
second consumer appears. Logged for future review.
### D-3: Multi-stage provider race vs. sequential fallback
**Status:** ✅ Resolved (public stays sequential).
Internal enhancer runs all providers in a stage in parallel
(`_run_provider_race`). Public enhancer runs sequentially with
retryable-error fallback.
**Decision:** Public stays sequential for PR 7.7. Race is a
Dreamverse-specific tail-latency optimization that depends on parallel
API budgets.
**Risk / Q-3:** First-segment latency on Dreamverse may regress
slightly when Cerebras has a bad minute (sequential waits 20s before
trying Groq). If real production concern, add public
`concurrency: int = 1` knob behind a race path — but only after measuring.
### D-4: Skip PR 7.9 router for the integration branch
**Status:** ✅ Resolved.
Internal stack ships `router/main.py` for multi-replica load balancing.
Dreamverse deployment uses single replica per region.
**Decision:** Land PR 7.9 publicly (upstream the surface). Skip wiring
into Dreamverse integration branch. Dreamverse's `server/main.py` does
not import from `router/`.
### D-5: Audio re-encode (PR 7.10) needed for streaming, deferred
**Status:** 🟡 Deferred to PR 7.10.
Internal streaming server's per-step path runs `_re_encode_audio` inside
`_stream_av_fmp4_events` so each fMP4 segment ships with
continuation-conditioning audio. Whole-segment `pool.run()` path doesn't
need this.
**Decision:** Land PR 7.10's `generate_async` publicly. Dreamverse
integration branch initially keeps using `pool.run()` (whole segment, no
re-encode). Follow-up branch swaps to `generate_async` + audio re-encode.
**Open question (Q-5):** Acceptable for first switch, or does
Dreamverse audio quality regress vs. internal until 7.10 wires in?
### D-6: `realtime/local_runtime.py` is NOT upstreamed
**Status:** ✅ Resolved.
It was the FastVideo-internal precursor to `streaming.gpu_pool`.
Upstreaming both would create two GPU pool implementations in public.
**Decision:** Don't upstream `realtime/local_runtime.py`. Dreamverse
switches to `streaming.gpu_pool.SubprocessGpuPool` on integration
branch. Internal module can be deleted at follow-up.
### D-7 / Q-6: `FP4Config` is private-only
**Status:** ✅ **Resolved May 2.**
April 26: `Dreamverse/server/video_generation.py:271` imported
`fastvideo.layers.quantization.fp4_config.FP4Config` from
FastVideo-internal only. The 411-line module hard-imported `flashinfer`.
**Two options at the time:**
1. Colocate publicly with `flashinfer` as optional extra
`pip install fastvideo[fp4]`; refactor `FP4QuantizeMethod` to take
layer-prefix list from a pipeline-config field instead of hardcoding
ltx2 paths.
2. Keep private — Dreamverse imports from internal via thin shim.
**Recommendation at the time:** option 1 once API refactor settles.
**Resolution:** May 2 work chose option 1.
- `365a66c7` upstreamed FP4Config with lazy `flashinfer` import in
loader helper (no public hard-dep)
- `94c983a2` renamed FP4 → NVFP4 to disambiguate from MX-FP4 / OCP-FP4
- `42b30bf9` wired through `fastvideo.layers.quantization`
See [quantization.md](quantization.md) for full details.
### D-8: `ltx2_image_crf` silently dropped by public schema
**Status:** 🔴 **Unverified post-`d80c2a8`.**
April 26: Dreamverse's `server/video_generation.py:406` passed
`ltx2_image_crf=0.0` to `SamplingParam(...)`. Public
`fastvideo.api.sampling_param.SamplingParam` did NOT have this field;
the BE logged ERROR and silently dropped the kwarg.
**Migration target** (per [design.md](design.md) compatibility map):
`request.stage_overrides.refine.image_crf`.
**Resolution status:** `d80c2a8` (May 2) refactored
`server/video_generation.py` to use typed `GeneratorConfig` +
`preset_overrides["refine"]`. Whether this PR routed `image_crf`
through the typed `stage_overrides` path or left it silently dropped is
unverified. See [open-threads.md](open-threads.md).
### D-9: `aarch64-conda-linux-gnu-cc` triton compile failure
**Status:** ✅ Resolved (operational).
Conda env injected an ARM cross-compiler ahead of `gcc` on `$PATH`, so
`torch._inductor`'s triton launcher failed compilation. Setting
`ENABLE_TORCH_COMPILE=0` bypasses it.
**Long-term fix:** clean conda env's compiler shadowing or add
`CC=gcc` override in Dreamverse's worker bootstrap.
### D-10: Warmup OOM on shared GPU
**Status:** ✅ Resolved (operational).
When `CUDA_VISIBLE_DEVICES` lands on a GPU another tenant uses, LTX-2
warmup fails with OOM. Picking an idle GPU (4-7 in test setup) is a
manual step.
**Improvement:** pre-warm probe that checks free memory before booting
the pool would prevent this.
### D-11: ffmpeg fragment write `Broken pipe`
**Status:** ✅ Resolved (cosmetic).
When WS client closes before backend finishes streaming first segment,
ffmpeg hits `[Errno 32] Broken pipe`. Currently propagates to
"User step failed". Cosmetic — swallowing pipe-broken on intentional
disconnect would clean up logs.
## Q-questions (from `streaming-server-upstream-plan.md`, Apr 17)
### Q-1: Router placement (in-repo or separate package)
**Status:** ✅ Resolved (in-tree).
**Recommendation at the time:** separate package `fastvideo-router/` or
`fastvideo/contrib/router/`; defer final call to PR 7.9.
**Resolution:** PR 7.9 implementation places router in-tree at
`fastvideo/entrypoints/streaming/router/`.
### Q-2: Session ID authority
**Status:** ✅ Resolved (server-generated).
**Recommendation:** server-generated UUID; accept externally provided
session ID only for resume flows.
### Q-3: Torch compile kwargs typing (opaque vs full vs hybrid)
**Status:** ✅ Resolved (hybrid).
**Recommendation:** hybrid — type the common four (`backend`,
`fullgraph`, `mode`, `dynamic`) + allow `extras: dict[str, Any]`.
**Resolution:** PR 6 + NVFP4 `221cb20a` shipped exactly this hybrid.
### Q-4: Prompt safety / fasttext dependency
**Status:** ✅ Resolved (optional extra).
**Recommendation:** ship as optional extra `pip install fastvideo[prompt-safety]`.
**Resolution:** PR 7.8 implements as optional extra.
### Q-5: Audio-specific tensor payloads in continuation
**Status:** ✅ Resolved (typed `LTX2ContinuationState`).
`ltx2_audio_clean_latent`, `ltx2_audio_denoise_mask`,
`ltx2_audio_latents` not in pre-refactor public schema.
**Recommendation:** classify as opaque fields inside
`LTX2ContinuationState.payload`, not top-level sampling fields.
**Resolution:** PR 7's typed `LTX2ContinuationState` lifts these into
typed fields (see [cross-repo-surfaces.md](cross-repo-surfaces.md)
field mapping table).
### Q-6: Dynamo subpackage home
**Status:** ✅ Resolved (lives in Dynamo repo).
**Resolution:** No Dynamo code in FastVideo. Full backend package
(handler, adapter, registration, health check) owned by Dynamo repo at
`components/src/dynamo/fastvideo/`, same pattern as vllm/sglang.
FastVideo only guarantees the public API contract.
### Q-7 (was Q-6 in dreamverse_review): How to land FP4Config publicly
**Status:** ✅ Resolved May 2 — option 1 (colocate publicly).
See D-7 above.
### Q-8: Disaggregation readiness contract test
**Status:** 🟡 Recommended; not yet shipped.
PR ai-dynamo/dynamo#7544 is aggregated-only. `ContinuationState` hybrid
already supports future prefill/decode split.
**Recommendation:** PR 7.10 explicitly validate `ContinuationState`
survives round-trip through Dynamo-style RPC (pickle or JSON), even
though Dynamo isn't using it today. Cheap regression guard.
### Q-9: Dynamo progress/status passthrough
**Status:** 🟡 Deferred until Dynamo clarifies.
`NvVideosResponse` has `status` and `progress` fields.
**Recommendation:** PR 7.10 stays aggregated-final-only to match PR
#7544 shape; revisit after Dynamo clarifies their streaming/progress
semantics.
## Cross-doc questions still 🔴 OPEN
These need decisions; tracked also in [open-threads.md](open-threads.md):
| ID | Question | Source | Why it matters |
|---|---|---|---|
| **D-8** | Did `d80c2a8` route `ltx2_image_crf` correctly, or is it still silently dropped? | dreamverse_review | Latent silent-drop bug; FP4-disabled paths may degrade |
| **VPO** | `video_position_offset_sec` — persistent accumulation (a) vs per-segment hint (b) | dreamverse_integration | Needs decision before PR 7.6 emits state |
| **SBS** | `SessionStore` / `BlobStore` lifecycle (TTL/eviction/blob-drop on state replacement) | dreamverse_integration | Needs decision in PR 7.5 design pass |
| **#1** | Migrate `/healthz`+`/readyz`+`/status` into FastVideo `build_app` | streaming-upstream-plan + handoff | Closes BE_FLAVOR=fastvideo FE-compatibility |
| **#3** | Add `cerebras_ifm` to public `PromptEnhancerConfig.provider` Literal | handoff | Internal supports it; public schema doesn't |
| **#4** | Expose `layer_profile` on typed `engine.quantization` | handoff | Removes Dreamverse's `experimental["pipeline_config"]` dodge |
| **#5** | Typed `dit_config.quant_config` carrier (design TBD) | handoff | Eliminates the `experimental["pipeline_config"]` escape hatch entirely |
@@ -0,0 +1,332 @@
# Design — Typed Public Inference API
Synthesis of the FastVideo public inference API refactor design philosophy.
For PR-by-PR execution see [pr-roadmap.md](pr-roadmap.md). For the streaming
extension see [streaming-server.md](streaming-server.md).
**Last updated:** 2026-05-03.
## Why the refactor
The pre-refactor public boundary mixed three concerns through `**kwargs`:
- `VideoGenerator.from_pretrained(..., **kwargs)` mixed engine/runtime,
pipeline init, and component overrides.
- `VideoGenerator.generate_video(..., **kwargs)` mixed prompt+inputs,
sampling, output, and model-specific workflow knobs.
- Unknown keys silently filtered or merely logged → API drift hard to detect.
- Multi-stage models (LTX-2 two-stage, Hunyuan15 SR, LongCat distill+refine)
exposed via ad hoc top-level flags.
This was already painful for LTX2/Dreamverse and would worsen as more
multi-stage pipelines came in.
## Core decision
FastVideo has:
1. **Typed nested public schema** — `RunConfig`, `ServeConfig`,
`GeneratorConfig`, `GenerationRequest`, `ContinuationState`.
2. **Model-owned named pipeline presets** — `ltx2_two_stage`,
`longcat_distill_refine`, `hunyuan15_sr_1080p`, etc. All 13 model families
landed presets in PR 4.
3. **Semantic stage overrides by stage name** —
`request.stage_overrides["refine"] = LTX2RefineStageOverride(...)`.
4. **Optional advanced explicit plans** for power users — `GenerationPlan`
(escape hatch only; not the canonical surface).
5. **YAML-first CLI** with dotted overrides —
`fastvideo generate --config run.yaml --request.sampling.seed 42`.
The canonical user experience: choose a model → choose a preset → override
a few typed fields → generate. Dicts/YAML/JSON are supported as
serialization, but parse immediately into typed objects with strict
unknown-key validation.
## Schema surface
Implemented in [`fastvideo/api/`](file:///home/william5lin/FastVideo/fastvideo/api/):
| Type | Role |
|---|---|
| `RunConfig` | Offline envelope: `generator` + `request` |
| `ServeConfig` | Serving envelope: `generator` + `server` + `default_request` + optional `streaming` |
| `GeneratorConfig` | `model_path`, `revision`, `trust_remote_code`, `engine`, `pipeline` |
| `EngineConfig` | parallelism / offload / compile / quantization / flags |
| `PipelineSelection` | `workload_type`, `preset`, `preset_version`, `components`, `preset_overrides`, `experimental` |
| `GenerationRequest` | `prompt`, `negative_prompt`, `inputs`, `sampling`, `runtime`, `output`, `stage_overrides`, `state`, `plan`, `extensions` |
| `ContinuationState` | Opaque envelope `{kind: str, payload: dict[str, Any]}` |
| `GenerationPlan` | Advanced/escape-hatch only; `{stages: list[PlannedStage], final_stage: str|None}` |
Files:
| File | Role |
|---|---|
| [`schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py) | All public dataclasses |
| [`parser.py`](file:///home/william5lin/FastVideo/fastvideo/api/parser.py) | `from_dict`, `to_dict`, `load_yaml`, `load_json`, validation |
| [`overrides.py`](file:///home/william5lin/FastVideo/fastvideo/api/overrides.py) | Dotted override application |
| [`compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py) | Legacy kwargs translation (~370 lines, scheduled for death PRs 14-17) |
| [`presets.py`](file:///home/william5lin/FastVideo/fastvideo/api/presets.py) | Preset registry |
| [`sampling_param.py`](file:///home/william5lin/FastVideo/fastvideo/api/sampling_param.py) | Internal `SamplingParam` adapter (canonical home since PR 4) |
| [`results.py`](file:///home/william5lin/FastVideo/fastvideo/api/results.py) | `GenerationResult` / `VideoResult` |
| [`errors.py`](file:///home/william5lin/FastVideo/fastvideo/api/errors.py) | Path-aware validation errors |
## Boundary normalization rule
Every public inference entrypoint normalizes into typed config objects
before touching legacy internals (`FastVideoArgs`, `SamplingParam`).
Includes Python constructors, `generate*` calls, CLI `generate`, CLI
`serve`, OpenAI server request translation, streaming server request
translation.
Legacy internals (`FastVideoArgs`, `SamplingParam`) may remain temporarily,
but only behind a typed normalization boundary.
## Strict-by-default validation
All structured inputs are strict:
- Unknown keys → error
- Wrong types → error
- Invalid stage names → error
- Incompatible state/preset combinations → error
The only intentional escape hatches:
- `generator.pipeline.experimental` — for in-flight features without typed home
- `request.extensions` — same, request-side
These bypass validation by design. Intent: shrink as presets absorb
model-specific fields. New fields should not land in `experimental` /
`extensions` without a plan to either promote them to typed fields or
remove them within two PR cycles.
Error format includes nested path:
```
Invalid field: request.stage_overrides.refine.num_inference_steps
Expected int, got "two"
Preset: ltx2_two_stage
Stage: refine
```
## Request mutation tracking
When a `GenerationRequest` is parsed from raw dict (YAML/JSON/Python),
FastVideo tracks which fields the user explicitly provided vs. which got
schema defaults. Matters for `request_to_sampling_param()` — explicit
values override model defaults; schema defaults do NOT.
Mechanics:
- At parse time, original raw dict + baseline snapshot stored on the request.
- Dataclass field mutations (e.g. `request.sampling.seed = 7`) captured via
lightweight `__setattr__` dirty-path recording.
- Dict-typed field mutations (e.g. `del request.stage_overrides["refine"]`)
detected at access time by diffing current dict vs. baseline.
- Setting a field to its schema default value IS captured as explicit, so
it overrides model defaults.
- Raw dict reconciled lazily when `normalize_generation_request()` is called.
## Schema purity (model-specific fields still in shared schema)
Remain for back-compat during initial migration; targeted for migration
into preset-owned typed override classes:
| Field | Owner | Migration target |
|---|---|---|
| `SamplingConfig.height_sr` / `width_sr` / `num_inference_steps_sr` | Hunyuan15 SR | `HunyuanSRStageOverride` (PR 10) |
| `SamplingConfig.guidance_scale_2`, `boundary_ratio` | Wan2.2, LingBotWorld | preset-owned (per-family PR) |
| `InputConfig.mouse_cond`, `keyboard_cond`, `grid_sizes` | MatrixGame | `request.extensions` or typed input config |
| `InputConfig.c2ws_plucker_emb` | LingBotWorld | `request.extensions` or typed input config |
| `InputConfig.refine_from`, `stage1_video` | LongCat | `LongCatRefineStageOverride` inputs (PR 9) |
LTX-2 multi-modal CFG knobs (`ltx2_modality_scale_video/_audio`,
`ltx2_rescale_scale`, `ltx2_stg_scale_video/_audio`,
`ltx2_stg_blocks_video/_audio`) still leak into shared `SamplingParam` but
only LTX-2 reads them today. Migration to typed `LTX2SamplingOverride` is
deferred to per-model migration sweep.
**LTX-2 CFG-force fix landed in PR 6**: defaults moved from `3.0/7.0` to
`1.0/1.0` to stop force-enabling CFG for non-LTX-2 families.
`ltx2_base` preset still sets `3.0/7.0` explicitly. Regression guard:
`test_presets.py::TestPresetDefaultTypes::test_ltx2_cfg_defaults_are_off`.
## Continuation state
Public surface:
```python
@dataclass
class ContinuationState:
kind: str # e.g. "ltx2.v1"
payload: dict[str, Any]
```
Internally, model-specific typed subclasses (e.g. `LTX2ContinuationState`
at [`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py)).
Payload must be JSON-serializable or use opaque blob-ID indirection for
large tensors — supports both stateless OpenAI client round-trip AND
future Dynamo prefill/decode disaggregation.
Hybrid model: server-held for streaming WS, client-round-trip for
stateless HTTP. See [streaming-server.md](streaming-server.md) D-1.
## Pipeline package structure (target)
Per-family colocation under `pipelines/basic/<family>/`:
```
fastvideo/pipelines/basic/<family>/
├── <family>_pipeline.py # pipeline implementation(s)
├── presets.py # user-facing presets (DONE in PR 4)
├── pipeline_configs.py # engine/arch config (from configs/pipelines/)
└── stages/ # model-specific stages (optional, if >2 files)
```
What stays shared:
- `configs/pipelines/base.py` — `PipelineConfig` base class
- `configs/models/` — architecture defs (dits/, vaes/, encoders/)
- `pipelines/stages/` — shared stages only (denoising, encoding, decoding,
text_encoding, timestep_preparation, ...)
What's gone (PR 4):
- `fastvideo/configs/sample/` — directory removed entirely; defaults
absorbed into per-family `presets.py`.
- All 12 `*_SamplingParam` subclass files — `SamplingParam` lives at
`fastvideo/api/sampling_param.py`; defaults flow through
`SamplingParam.from_pretrained()` → `_from_preset()`.
What's pending: `configs/pipelines/<family>.py` colocation, optional
`pipelines/stages/<family>_*.py` colocation. Per-model migration PRs
(6/9/10) include the colocation step for that family.
## YAML examples
### Run config
```yaml
generator:
model_path: /models/ltx2
engine:
num_gpus: 1
parallelism: {tp_size: -1, sp_size: -1}
offload: {dit: false, text_encoder: false, vae: false, pin_cpu_memory: true}
pipeline:
workload_type: t2v
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
request:
prompt: "a fox running through snow"
sampling: {num_frames: 121, height: 1024, width: 1536, num_inference_steps: 8, seed: 42}
output: {save_video: true, return_state: true}
stage_overrides:
refine: {num_inference_steps: 2, guidance_scale: 1.0}
```
### Serve config
See [`Dreamverse/serve_configs/streaming_demo.yaml`](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml)
for a canonical example matching internal/ui defaults (LTX-2 distilled,
NVFP4, 121 frames @ 1088×1920 24fps, 5 inference steps, 2-step refine).
## Compatibility mapping (legacy → typed)
| Legacy field | New path |
|---|---|
| `model_path` | `generator.model_path` |
| `num_gpus` | `generator.engine.num_gpus` |
| `tp_size` / `sp_size` | `generator.engine.parallelism.{tp_size,sp_size}` |
| `dit_cpu_offload` | `generator.engine.offload.dit` |
| `enable_torch_compile` | `generator.engine.compile.enabled` |
| `torch_compile_kwargs` | split: `generator.engine.compile.{backend,fullgraph,mode,dynamic}` + `.extras` |
| `enable_torch_compile_text_encoder` | `generator.engine.compile.text_encoder_enabled` |
| `prompt_txt` | `request.inputs.prompt_path` |
| `image_path` / `video_path` | `request.inputs.{image_path,video_path}` |
| `output_path` / `save_video` / `return_frames` | `request.output.*` |
| `seed` / `num_frames` / `height` / `width` / `fps` / `num_inference_steps` / `guidance_scale` | `request.sampling.*` |
| `enable_teacache` / `return_trajectory_*` | `request.runtime.*` |
LTX-2 specific (private adapter, NOT public compat promise):
| Legacy LTX-2 field | New path |
|---|---|
| `config_model_path` | `generator.pipeline.components.config_root` |
| `ltx2_refine_enabled` | `generator.pipeline.preset_overrides.refine.enabled` |
| `ltx2_refine_upsampler_path` | `generator.pipeline.components.upsampler_weights` |
| `ltx2_refine_lora_path` | `generator.pipeline.components.lora_path` |
| `ltx2_refine_num_inference_steps` | `request.stage_overrides.refine.num_inference_steps` |
| `ltx2_refine_guidance_scale` | `request.stage_overrides.refine.guidance_scale` |
| `ltx2_refine_add_noise` | `generator.pipeline.preset_overrides.refine.add_noise` |
| `ltx2_image_crf` | `request.stage_overrides.refine.image_crf` |
| `return_continuation_state` | `request.output.return_state` |
LongCat:
| Legacy | New |
|---|---|
| `refine_from` / `stage1_video` | `request.inputs.{refine_from,stage1_video}` |
| `t_thresh` / `spatial_refine_only` / `num_cond_frames` | `request.stage_overrides.refine.*` |
## External inspirations (and limits)
| Source | Useful idea | Don't copy |
|---|---|---|
| Ray | YAML-first config interchange | Ray's package layout |
| SGL `multimodal_gen` | Split instance/request config; dict input parsed into typed objects; merge user overrides on model defaults | `SamplingParams._adjust(ServerArgs)` (request depending on engine config); broad weakly-typed request bags |
| vLLM-Omni | Model-owned pipeline presets; explicit stage topology; per-stage default sampling | Positional `sampling_params_list`; serving-engine stage-index semantics in primary Python API |
## Naming guidance
- Public schema names namespaced under `fastvideo.api`
- Don't export from top-level `fastvideo/__init__.py` until migration further along
- `RunConfig` / `ServeConfig` get sufficient disambiguation from training
config via the namespace
- Future rename to `EngineQuantizationConfig` reserved if a collision
arises (deferred)
## Public Python API (canonical form)
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GeneratorConfig, GenerationRequest,
EngineConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
generator = VideoGenerator.from_pretrained(
config=GeneratorConfig(
model_path="/models/ltx2",
engine=EngineConfig(num_gpus=1),
pipeline=PipelineSelection(workload_type="t2v", preset="ltx2_two_stage"),
)
)
result = generator.generate(
GenerationRequest(
prompt="a fox running through snow",
sampling=SamplingConfig(num_frames=121, height=1024, width=1536,
num_inference_steps=8, seed=42),
output=OutputConfig(save_video=True, return_state=True),
)
)
```
Accepted constructor forms:
```python
VideoGenerator.from_pretrained(config=GeneratorConfig(...))
VideoGenerator.from_config(GeneratorConfig(...))
VideoGenerator.from_file("run.yaml")
VideoGenerator.from_pretrained("model-id", num_gpus=2, ...) # stable convenience
VideoGenerator.from_pretrained(model_path, **legacy_kwargs) # compat (deprecated PR 13)
```
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,888 @@
# Integration Review — Drift Audit + Path Forward
> # ⚠️ DEPRECATED — superseded by [integration-plan.md](integration-plan.md)
>
> This document recommended **Option D** (Dreamverse stays a separate repo;
> generic backend merges into FastVideo). On 2026-05-05 the team chose
> **Option B+** instead (Dreamverse FE + product server move into FastVideo
> as `apps/dreamverse/`; generic backend stays at
> `fastvideo.entrypoints.streaming.*` per Option D's principle).
> See [decisions-log.md D-18](decisions-log.md#d-18) for the strategy
> reversal rationale and [integration-plan.md](integration-plan.md) for the
> executable migration plan.
>
> **What's still authoritative in this file:**
> - **Part 1 — Drift audit** (the 17-row drift summary table). The drift
> findings remain valid; the migration plan in `integration-plan.md`
> folds them into specific phases.
> - **OSS precedent citations** (vLLM, BentoML, Ray Serve, TGI+ChatUI,
> Transformers.js, ComfyUI, AUTOMATIC1111). Reused in `integration-plan.md`.
>
> **What's superseded:**
> - **Part 2 — Recommendation (Option D)**. Replaced by Option B+ in the
> new plan. Read `integration-plan.md` for the current decision.
> - **Part 3 — Action items**. Replaced by the phased migration plan.
>
> Kept in tree for historical reference and audit trail. Do not delete.
**Last updated:** 2026-05-05 (deprecated header added).
**Scope:** FastVideo public `will/ltx2_sr_port` at the requested audit
anchor `b36bdbc9`; Dreamverse `will/integrate-public-fastvideo` at
`ec8ef92`; FastVideo-internal `will/rebase-nbv` as read-only comparison.
**Memory-dir context:** the current integration memory snapshot tracks the
same public mega-PR lineage as `will/ltx2_sr_port`, with PRs #1257,
#1258, #1284, and #1286 already merged, #1287 closed, and #1288 open as
the consolidated landing vehicle for LTX-2 SR runtime, NVFP4,
`generate_async`, Dynamo contract, and memory-dir cleanup. Source:
[memory index](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/README.md#L8-L19)
and [D-17](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L19-L45).
**Bottom line:** Zero core typed API drift — typed construction, typed
continuation state, NVFP4 wiring, and Dynamo-facing async events are either
already public or in #1288. **Real drift remains on the realtime-runtime
contract surface (`/healthz` / `/readyz` / `/status` routes per
[cross-repo-surfaces.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L74-L88))
and on operational/product edges**: stale Dreamverse docs/scripts, a
1933-LOC Dreamverse prompt-enhancer fork, two unresolved per-session
fields (`ltx2_image_crf` D-8, `video_position_offset_sec` VPO), one
missing example config, and two internal-only utilities whose product
relevance is not yet proven.
---
## Part 1 — Drift audit
### Methodology
1. **Compared three repositories and branches.**
- FastVideo public: `/home/william5lin/FastVideo`, branch
`will/ltx2_sr_port`.
- Dreamverse: `/home/william5lin/Dreamverse`, branch
`will/integrate-public-fastvideo`.
- FastVideo-internal: `/home/william5lin/FastVideo-internal`, branch
`will/rebase-nbv`.
- Canonical repo paths are listed in the integration memory index:
[repo paths](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/README.md#L72-L79).
2. **Scoped the audit to the ultimate integration goal.**
- Dreamverse should depend on public `fastvideo`, not
`FastVideo-internal`.
- FastVideo should own the reusable backend subset that Dreamverse
currently needs from internal: streaming runtime, GPU pool, router,
prompt enhancer, NVFP4, continuation state, and typed generation.
- Dynamo should consume FastVideo through typed public Python APIs, not
through private modules.
- The three Dreamverse surfaces are documented as pipeline construction,
realtime runtime, and continuation state:
[cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L13-L20).
3. **Separated intentional refactor from drift.**
- A path rename is not drift if the public branch contains the same
responsibility under the typed design.
- A deleted file is not drift if the public design intentionally
consolidated it.
- A private alias is not drift if the public schema exposes a typed
replacement with contract tests.
- This matches the typed-public-boundary rule in
[design.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L24-L43).
4. **Used memory docs for rationale and worktree files for concrete proof.**
- API schema and public exports:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L68-L85),
[api exports](file:///home/william5lin/FastVideo/fastvideo/api/__init__.py#L49-L109).
- Streaming server current routes:
[build_app](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L88-L160).
- Dreamverse dependency state:
[pyproject server extra](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22),
[uv lock editable source](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
- Contract tests:
[Dreamverse shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26),
[Dynamo shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19),
[generate_async](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L1-L7).
5. **Did not treat product-only Dreamverse behavior as FastVideo drift.**
- Dreamverse keeps a local product server and Next.js UI today:
[README baseline](file:///home/william5lin/Dreamverse/README.md#L5-L16).
- Product-only routes, curated presets, devtools, and frontend-specific
behavior belong in Dreamverse unless a second non-Dreamverse consumer
needs them.
6. **Risk scale used below.**
- **P0:** blocks Dreamverse from running without FastVideo-internal.
- **P1:** blocks clean `BE_FLAVOR=fastvideo` or Dynamo/public API use.
- **P2:** reproducibility or maintenance drag.
- **P3:** optional parity or future memory/perf improvement.
### Findings: zero core typed API drift
The public branch is aligned with the goal on the **core typed API
surface** (construction, request, continuation state, async events).
The table below lists items that look like drift only if compared by
path name or legacy field name. They are intentional public refactors
or already guarded by tests. **Note:** the realtime-runtime _contract_
surface (FE-required health routes) is a separate matter — see "real
drift items" §4 below.
| Investigated item | Drift? | Evidence | Conclusion |
|---|---:|---|---|
| Dreamverse surface 1: pipeline construction | No | Dreamverse migrated from flat kwargs to typed `GeneratorConfig` at `d80c2a8`; mapping documented in [cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L22-L47). | Stable public surface exists. |
| Dreamverse surface 2: realtime runtime | No on architecture; some route work remains | Runtime migration target is public `streaming/`, not internal `realtime/`: [streaming upstream](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L14-L31). | Rename/refactor is intentional. |
| Dreamverse surface 3: continuation state | No | Public typed `ContinuationState` plus LTX-2 state mapping are documented in [cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L108-L146). | Public state is a superset of Dreamverse's data carrier. |
| Internal `fastvideo/entrypoints/realtime/` | No | Public design chooses parallel `fastvideo/entrypoints/streaming/`: [layout decision](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L55-L60), [current build_app](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L88-L160). | Intentional rename plus typed-config rewrite. |
| Internal `configs/sample/` presets | No | Public PR 4 intentionally deleted `configs/sample/` and moved defaults to per-family presets plus `fastvideo/api/sampling_param.py`: [design](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L194-L205), [PR roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L21-L29). | Intentional consolidation. |
| LTX-2 pipeline presets | No | Public target is model-owned named presets and per-family colocation: [design](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L175-L205). | Public layout matches design. |
| Internal `use_fp4_linear` flag | No | Public typed quant carrier is `engine.quantization.transformer_quant`; schema field exists in [schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L68-L85), compat resolves it in [compat.py](file:///home/william5lin/FastVideo/fastvideo/api/compat.py#L267-L279). | Replaced by typed NVFP4 surface. |
| Public-only `transformer_quant` field | No | Public `FastVideoArgs` pins typed quant to `dit_config.quant_config`: [fastvideo_args](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py#L220-L228), [apply logic](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py#L260-L279). | Public superset, not drift. |
| Internal `config_model_path` | No | Public typed home is `generator.pipeline.components.config_root`: [design mapping](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L261-L269), [compat mapping](file:///home/william5lin/FastVideo/fastvideo/api/compat.py#L295-L299). | Alias is covered. |
| Internal flat video request fields | No | Public `GenerationRequest` nests `inputs`, `sampling`, `runtime`, `output`, `state`, `extensions`: [schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204). Internal legacy fields live in internal protocol at [protocol.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/openai/protocol.py#L64-L82). | Intentional request refactor. |
| Dreamverse typed init kwargs | No | Contract test asserts current Dreamverse load kwargs all land on typed fields, not `experimental`: [test_dreamverse_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L44-L135). | Guard in place. |
| Dreamverse request path | No | Contract test asserts request fields round-trip through typed `GenerationRequest`: [test_dreamverse_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L153-L197). | Guard in place. |
| Dynamo native backend shape | No | FastVideo's only obligation is stable typed Python API: [cross-repo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210). | Dynamo should stay out of FastVideo. |
| `generate_async` event API | No | API exists in [video_generator](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332), event types exist in [results.py](file:///home/william5lin/FastVideo/fastvideo/api/results.py#L109-L164). | #1288 covers the async contract. |
| Dynamo request mapping | No | Authoritative source is the contract test [test_dynamo_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L90-L175) which asserts `req.prompt`, `req.sampling.{height,width,num_frames,fps,num_inference_steps,guidance_scale,seed,negative_prompt}`, and `req.inputs.{image_path,video_path}` against the actual nested [`GenerationRequest` schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204). The `streaming-server.md` Dynamo mapping table mis-cites a `prompt -> sampling.prompt` path that no longer exists; the test is correct, the doc is stale and tracked for refresh. | Guard in place; companion doc needs minor refresh. |
| Public API exports | No | `VideoEvent`, `VideoResult`, and typed schema classes are exported from [fastvideo.api](file:///home/william5lin/FastVideo/fastvideo/api/__init__.py#L49-L109). | Integration imports resolve. |
| FastVideo-internal FP4/NVFP4 paths | No | Public NVFP4 files and roles are documented in [quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L24-L35); actual `NVFP4Config` documents lazy FlashInfer and public naming in [nvfp4_config.py](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/nvfp4_config.py#L1-L19). | Public is typed superset. |
| AbsMaxFP8 refactor | No for Dreamverse | Public quant registry includes `AbsMaxFP8` and `NVFP4`: [quantization init](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/__init__.py#L1-L8). AbsMaxFP8 failure is tracked as separate tech debt: [open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L100-L117). | Not Dreamverse blocker. |
| Internal realtime API regression test | No | Public contract tests replace it: [Dreamverse contract](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26), [Dynamo contract](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19), [generate_async tests](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L91-L230). | Better scoped guards exist. |
| Dreamverse dependency declaration | No | `server` extra declares `fastvideo>=0.1.7`: [pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22). Dev lock resolves editable public `../FastVideo`: [uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722), [package source](file:///home/william5lin/Dreamverse/uv.lock#L777-L780). | Dependency is already switched in metadata/lock. |
#### Core conclusion for the zero-typed-drift section
The public typed API no longer needs to mirror `FastVideo-internal` file
paths. The correct test is whether Dreamverse and Dynamo can express their
needs through public typed objects and public entrypoints. On that test,
the **typed core** is covered (construction, request, continuation,
async events). The **runtime contract** still has health-route gaps —
see real drift §4. On the **typed core**:
- `GeneratorConfig` and `GenerationRequest` cover construction and calls:
[schema surface](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L45-L72).
- `ServeConfig.streaming` covers the server envelope:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L244-L279).
- `generate_async` covers streaming, OpenAI, and Dynamo on one substrate:
[video_generator](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332).
- Contract tests now encode the cross-repo shapes:
[Dreamverse](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L70-L214),
[Dynamo](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L170-L331),
[async events](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L91-L273).
### Findings: real drift items requiring action
#### 1. Dreamverse README and bootstrap script still point at FastVideo-internal
- **Priority:** P0 for a clean public-dependency story.
- **Effort:** Small.
- **Owner:** Dreamverse repo.
- **Evidence:** Dreamverse metadata already points at public FastVideo:
[pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22),
[uv source](file:///home/william5lin/Dreamverse/pyproject.toml#L54-L61),
[uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
- **Drift:** README still tells users that `uv` resolves from
`../FastVideo-internal` and that bootstrap expects `../FastVideo-internal`:
[README](file:///home/william5lin/Dreamverse/README.md#L76-L109).
- **Drift:** bootstrap script still defaults to cloning the private repo and
verifying imports from that clone:
[script defaults](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11),
[script clone flow](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L33-L63),
[script import assertion](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L66-L88).
- **Action:** Replace private-fork bootstrap with public FastVideo bootstrap
or delete the bootstrap once PyPI publication is the default path.
- **Do not overreach:** no FastVideo code change required.
#### 2. Dreamverse carries a 1933-line prompt-enhancer fork
- **Priority:** P1.
- **Effort:** Medium.
- **Owner:** Dreamverse repo, after public prompt enhancer is available.
- **Evidence:** Dreamverse local fork starts at
[server/prompt_enhancer.py](file:///home/william5lin/Dreamverse/server/prompt_enhancer.py#L1-L80).
- **Public replacement:** FastVideo now has provider-agnostic
`PromptEnhancer` with `enhance`, `auto_extend`, `rewrite`, and
`register_provider`:
[public enhancer](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/enhancer.py#L66-L142).
- **Provider extension point:** custom providers implement `LLMProvider`:
[provider protocol](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/providers/base.py#L63-L75).
- **Tracking:** DR-1 in open threads already defines the compat-shim shape:
[DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
- **Action:** Replace the fork with a small Dreamverse shim that adapts
public `LLMResponse` to Dreamverse's product response objects and keeps
only product-only extras.
- **Do not overreach:** do not merge Dreamverse's full prompt product layer
into FastVideo unless a second consumer needs the same semantics.
#### 3. `cerebras_ifm` provider is unresolved
- **Priority:** P1 if Dreamverse needs IFM in production; P2 otherwise.
- **Effort:** Small decision plus small/medium implementation.
- **Owner:** Team decision; implementation either Dreamverse-side or public.
- **Public state:** `PromptEnhancerConfig.provider` is currently
`Literal["cerebras", "groq"]`:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L229-L235).
- **Design note:** public Literal excludes `cerebras_ifm` today:
[streaming-server D-3](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L61-L100).
- **Tracking:** DR-2 already frames the public-vs-Dreamverse decision:
[DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
- **Recommended default:** implement IFM as a Dreamverse-side custom provider
registered through `enhancer.register_provider(...)` unless there is a
non-Dreamverse public user.
#### 4. `/healthz`, `/readyz`, and `/status` are not in public `build_app`
- **Priority:** P1 for `BE_FLAVOR=fastvideo` frontend compatibility.
- **Effort:** Medium/Large because route shapes need tests.
- **Owner:** FastVideo public.
- **Public current state:** `build_app` exposes `GET /health` and
`WS /v1/stream`:
[server.py](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L126-L160).
- **Dreamverse expected state:** Dreamverse exposes `GET /healthz`,
`GET /readyz`, and `GET /status`:
[routes/health.py](file:///home/william5lin/Dreamverse/server/routes/health.py#L34-L79).
- **Tracking:** open item #1 documents route ownership and files likely to
touch:
[open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
- **Design note:** `/curated-presets`, `/prompt-system-config`, and devtools
stay Dreamverse-side, with feature detection:
[streaming route contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L232-L259).
#### 5. `fastvideo/models/layerwise_offload.py` exists only internally
- **Priority:** P3 unless memory-tight Dreamverse deployments require it.
- **Effort:** Medium if adopted; low if documented as deferred.
- **Owner:** FastVideo public only if a concrete deployment needs it.
- **Internal evidence:** internal file defines async layerwise CPU offload
manager with pinned CPU memory and prefetch stream:
[layerwise_offload.py](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L1-L20),
[prefetch path](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L127-L180).
- **Public state:** no equivalent public file was identified in this audit.
- **Action:** defer unless Dreamverse or another public deployment hits a
memory ceiling that cannot be handled by existing offload knobs.
- **Decision rule:** if adopted, port as a generic offload utility with
tests; do not make it Dreamverse-specific.
#### 6. Standalone LTX-2 upsampler CLI exists only internally
- **Priority:** P2 for reproducibility; P3 for product runtime.
- **Effort:** Small/Medium after scope decision.
- **Owner:** FastVideo public if standalone upsampling is a supported user
workflow.
- **Internal utility:** `upscale_video_file(...)` reads an existing video,
prepares frame count/resolution, loads VAE + upsampler, and writes an mp4:
[upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L120-L180),
[write tail](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L181-L202).
- **Internal CLI:** `fastvideo upsample` wrapper exists internally:
[cli/upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35),
[CLI args](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L48-L130).
- **Public related functionality:** LTX-2 SR refine stage covers the
in-pipeline latent upsample/refine path:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22),
[upsample stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L116-L180).
- **Action:** decide whether standalone file-to-file upsampling is a public
CLI promise or whether the SR refine stage is sufficient.
#### 7. Reproducible streaming demo config lives only in Dreamverse
- **Priority:** P2.
- **Effort:** Small.
- **Owner:** FastVideo public.
- **Evidence:** canonical demo config currently lives at
[Dreamverse/serve_configs/streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L12).
- **Config content:** it documents LTX-2 distilled model, one GPU,
no offload, compile settings, NVFP4, refine overrides, default request,
and streaming settings:
[generator block](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L31-L87),
[streaming block](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L108-L149).
- **Memory pointer:** design.md already treats this as the canonical
example:
[design YAML example](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L235-L239).
- **Action:** copy/adapt it into
`examples/serving/streaming_demo.yaml` with public-safe comments.
#### 8. LTX-2 stage equivalence is a verification gap, not proven drift
- **Priority:** P2.
- **Effort:** Medium if parity checks are added; small if only manual audit.
- **Owner:** FastVideo public.
- **Public state:** model-specific LTX-2 stages are colocated under
`fastvideo/pipelines/basic/ltx2/stages/`, consistent with the target
layout in [design.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L175-L205).
- **Example public stage:** `ltx2_refine.py` explicitly says it is a
public-side port of the internal stage and describes the three-stage SR
flow:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
- **Action:** verify behavior for the six internal `ltx2_*` stage files
against public colocated stages. If a mismatch is found, file it as a
real drift item with a failing parity test.
### Findings: deferred / accepted residual
These items should not block the public-dependency transition.
1. **StepVideo residual.**
- Dreamverse's model registry is LTX-2/LTX-2.3 only:
[Dreamverse config](file:///home/william5lin/Dreamverse/server/config.py#L28-L45).
- Internal local tests even stub StepVideo modules to keep LTX registry
tests focused:
[test_ltx2_registry.py](file:///home/william5lin/FastVideo-internal/tests/local_tests/test_ltx2_registry.py#L38-L61).
- Conclusion: accepted low-priority deferral unless Dreamverse adds a
StepVideo model.
2. **Internal debug-only `FastVideoArgs` fields.**
- Internal debug fields exist around `FastVideoArgs` and stage/model sums:
[internal grep source](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203).
- They are debug-only and not a public user-facing integration surface.
- Conclusion: low-priority; do not add to public schema unless a debug
workflow requires them.
3. **Private request aliases.**
- Public request schema is nested and strict:
[GenerationRequest](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204).
- Legacy OpenAI flat fields are compatibility input, not the canonical
public API:
[internal protocol](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/openai/protocol.py#L64-L82).
- Conclusion: no action beyond current compat tests.
4. **`experimental["pipeline_config"]` escape hatch.**
- Dreamverse currently uses an explicit in-memory quant config because
typed `transformer_quant: "NVFP4"` does not expose `layer_profile`:
[quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L86-L97).
- Open thread #4 tracks `layer_profile`:
[open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260).
- Conclusion: defer broader typed carrier design; add `layer_profile`
first if Dreamverse needs base/refine profile selection.
5. **Router sticky routing and active-active semantics.**
- Public router intentionally ships active-passive first and defers
sticky/weighted routing:
[D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L118-L155).
- Follow-ups are tracked:
[D-15 action items](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201).
- Conclusion: not drift; defer until load-balancing needs are real.
6. **AbsMaxFP8 failure.**
- Pre-existing and not introduced by NVFP4:
[state](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/state.md#L154-L159),
[quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L202-L214).
- Conclusion: fix separately; not a Dreamverse public-dependency blocker.
### Drift summary table
| # | Item | Priority | Effort | Status | Tracked where | Next action |
|---:|---|---|---|---|---|---|
| 1 | Dreamverse README still names `../FastVideo-internal` | P0 | S | Real drift | [README lines](file:///home/william5lin/Dreamverse/README.md#L76-L109) | Update docs to public FastVideo / PyPI path. |
| 2 | Dreamverse private bootstrap clones internal repo | P0 | S | Real drift | [bootstrap script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11) | Replace or delete private bootstrap. |
| 3 | Dreamverse `prompt_enhancer.py` fork | P1 | M | Real drift | [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206) | Build compat shim over public enhancer. |
| 4 | `cerebras_ifm` provider path | P1/P2 | S-M | Real drift / decision | [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229) | Choose public provider vs Dreamverse custom provider. |
| 5 | Health route mismatch | P1 | M-L | Real drift | [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99) | Add `/healthz`, `/readyz`, `/status` to public build_app. |
| 6 | Missing public streaming demo config | P2 | S | Real drift | [Dreamverse config](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L12) | Add `examples/serving/streaming_demo.yaml`. |
| 7 | Standalone upsampler CLI | P2/P3 | S-M | Real drift if standalone CLI is desired | [internal CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35) | Decide CLI promise; port or defer. |
| 8 | Layerwise offload utility | P3 | M | Optional internal-only residual (no Dreamverse deployment requires it today) | [internal manager](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20) | Defer until memory-tight deployment needs it. |
| 9 | LTX-2 stage equivalence | P2 | S-M | Verification gap | [public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22) | Add targeted parity audit/test if needed. |
| 10 | StepVideo | P3 | M | Accepted residual | [Dreamverse model registry](file:///home/william5lin/Dreamverse/server/config.py#L28-L45) | No action unless Dreamverse adds StepVideo. |
| 11 | Debug-only fields | P3 | S | Accepted residual | [internal args](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203) | Do not publicize unless needed. |
| 12 | `layer_profile` typed quant knob | P2 | M | Tracked gap | [open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260) | Add typed layer profile if Dreamverse drops escape hatch. |
| 13 | `ltx2_image_crf` per-segment field flow (D-8) | P1 | S | Open verification gap — Dreamverse still passes `ltx2_image_crf=0.0` per [Dreamverse video_generation.py](file:///home/william5lin/Dreamverse/server/video_generation.py#L420-L435); needs trace-through to confirm it lands on `request.stage_overrides.refine.image_crf` rather than being silently dropped | [D-8 in open-threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L52-L68) | 10-min trace + add a Dreamverse-shape contract test pinning the field. |
| 14 | `video_position_offset_sec` semantics (VPO) | P1 | S | Open decision — persistent-vs-per-segment ambiguity unresolved | [VPO in open-threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L120-L144) | Confirm semantics with audio team; document + add test. Decision deadline was "before PR 7.6 emits state" — that PR (7.6 / #1257) is now MERGED, so the decision is overdue. |
| 15 | `GpuPool` ABC docstring missing experimental caveat (D-12-A) | P3 | trivial | Tracked gap — `GpuPool` ABC at [gpu_pool.py:74-83](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/gpu_pool.py#L74-L83) lacks the "API may change post-PR-7.10; experimental / server-internal" caveat | [D-12-A](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L301-L313) | Edit docstring; trivial. |
| 16 | `GpuPool.run_async()` migration (D-12-B) | P2 | M | Tracked gap — `GpuPool.run() -> Any` should become `run_async() -> AsyncIterator[VideoEvent]` per D-12 / D-12-B | [D-12-B](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L317-L327) | Land alongside #1288 merge or in immediate follow-up. |
| 17 | `SessionStore` / `BlobStore` lifecycle policy (SBS) | P2 | M | Tracked gap — in-memory defaults have no eviction/TTL/blob-cleanup policy | [SBS](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L278-L294) | Streaming server design pass needed before high-traffic deployment. |
| 13 | Router sticky / active-active | P3 | M | Deferred | [D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201) | Defer until reconnect/load evidence. |
| 14 | AbsMaxFP8 test failure | P2 | S | Separate tech debt | [state](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/state.md#L154-L159) | Fix outside Dreamverse migration. |
| 15 | Dynamo backend package | P1 | External | Not FastVideo drift | [Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210) | Reopen Dynamo-side PR after public API lands. |
---
## Part 2 — Integration path tradeoffs
### The four options
#### Option A — Status quo: Dreamverse stays separate and depends on `fastvideo`
**Shape**
- FastVideo remains the Python library and reusable backend runtime.
- Dreamverse remains the product repo with FastAPI product glue and Next.js
frontend.
- Dreamverse `server` extra depends on `fastvideo>=0.1.7`:
[pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22).
- Local development can keep using editable `../FastVideo` until PyPI
publication catches up:
[uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
**What it solves**
- Directly satisfies "Dreamverse depends on public FastVideo".
- Keeps frontend release cadence independent.
- Keeps product-specific prompts, routes, and UI in the product repo.
- Minimizes FastVideo packaging and CI growth.
**What it does not solve by itself**
- Does not remove Dreamverse prompt-enhancer fork unless DR-1 is executed.
- Does not give Dreamverse FE compatibility with public `build_app` until
health routes migrate.
- Does not make Dreamverse server itself reusable as a public entrypoint.
**Best fit**
- Default for the next release if the goal is to stop using
FastVideo-internal quickly and safely.
#### Option B — Dreamverse as a subfolder under FastVideo
**Shape**
- One repository: FastVideo contains `dreamverse/server/` and
`dreamverse/apps/web/`.
- Dreamverse can remain a separate package in the same repo, or FastVideo's
build can ignore Dreamverse by default.
- CI must understand Python library tests plus Next.js install/build/test.
**What it solves**
- Eliminates sibling-checkout drift.
- Makes cross-repo integration changes atomic.
- Easier for a single reviewer to see library and product changes together.
**Costs**
- Adds frontend dependency management to a Python ML library repo.
- Couples clone size, CI setup, issue tracking, and review load.
- Forces maintainers to decide whether product assets are included in source
distributions, wheels, docs, and release notes.
**Best fit**
- Only if Dreamverse becomes the primary FastVideo product surface and the
team accepts a product monorepo.
#### Option C — Full merge into `fastvideo.entrypoints.dreamverse.*`
**Shape**
- Dreamverse backend becomes FastVideo code.
- Public import becomes something like
`from fastvideo.entrypoints.dreamverse import build_app`.
- CLI becomes `fastvideo dreamverse-serve --config dreamverse.yaml`.
- Frontend either ships as static assets in the package or as a frontend
extra.
**What it solves**
- One namespace and one release train for library plus product backend.
- No dependency boundary between Dreamverse server and FastVideo internals.
- Product route contract can be tested entirely inside FastVideo CI.
**Costs**
- Maximally expands FastVideo's public/security surface.
- Locks product experiments to FastVideo release cadence.
- Makes private prompt/provider/product assumptions look like framework API.
- Has weak precedent for a Python ML library plus Next.js product being merged
into the library namespace.
**Best fit**
- Only if Dreamverse is no longer a separate product and becomes the
canonical FastVideo UI/serving mode.
#### Option D — Hybrid: backend merges, frontend stays separate
**Shape**
- Reusable backend components merge into public FastVideo.
- Frontend stays in a separate Dreamverse UI repo or Dreamverse product repo.
- The backend should be generic where possible: `fastvideo.entrypoints.streaming`,
not product-only names, unless product-only routes are intentionally
accepted as public API.
- This matches the current trajectory: streaming server, GPU pool, prompt
enhancer, safety/rewrite/session logging, router, NVFP4, and
`generate_async` are public-side work already tracked in the PR roadmap:
[pr-roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L19-L42).
**What it solves**
- Removes FastVideo-internal dependency for reusable backend pieces.
- Keeps product frontend cadence independent.
- Gives non-Dreamverse users a streaming backend and typed API without
carrying the Dreamverse app.
- Gives Dynamo a stable library API while leaving Dynamo package code in
Dynamo:
[Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210).
**Costs**
- Requires careful boundary discipline: generic streaming/server code in
FastVideo; product routes/prompts/presets in Dreamverse.
- Requires contract tests to prevent drift.
- Some Dreamverse compatibility routes may become public and need support.
**Best fit**
- Best long-term target if the team wants FastVideo to own serving/runtime
infrastructure while keeping Dreamverse as a separately evolving product.
### Comparison matrix
| Criterion | A. Separate dep | B. Subfolder monorepo | C. Full namespace merge | D. Hybrid backend merge |
|---|---|---|---|---|
| Alignment with stated goal | High: Dreamverse depends on public package | Medium: no external dep, but product becomes repo-local | Medium: dependency disappears by absorption | High: reusable backend in public, product separate |
| Time to remove `FastVideo-internal` | Fastest | Medium | Slowest | Medium-fast |
| Build complexity | Low | High: Python + Next.js in one repo | High: Python package plus static/frontend extras | Medium: Python backend only in FastVideo |
| Release cadence | Independent | Coupled clone; releases can still be separate but more friction | Fully coupled | Backend coupled to FastVideo, frontend independent |
| Security surface in FastVideo | Low | Medium/High | Highest | Medium |
| Contributor friction | Low for both repos | Higher for library contributors | Highest; product assumptions in library | Medium; clear backend boundary needed |
| Dependency management | Normal package pin | Workspace/monorepo tooling needed | FastVideo extras/static asset decisions needed | FastVideo extras for backend; FE out-of-tree |
| CI cost | Low/medium | High | High | Medium |
| Contract-test value | High; cross-repo contract tests are essential | Medium; same repo but still useful | Medium; less boundary pressure | High; generic backend vs product boundary |
| Precedent strength | Strong: library/server plus external UI patterns exist | Mixed | Weak for Python ML library + Next.js inside namespace | Strongest match: in-tree server/backend, external UI |
| Packaging risk | Low | Medium/high | High | Medium |
| Future Dynamo fit | Strong | Strong if API remains clean | Risky if product API bleeds in | Strong |
| Frontend iteration speed | Highest | Lower | Lowest | Highest |
| Risk of product-specific API leakage | Low | Medium | High | Medium; controllable with naming discipline |
| Reversibility | High | Medium | Low | Medium/high |
### OSS precedents (with citations)
| Pattern | Project | What it supports | Citation |
|---|---|---|---|
| Library plus in-tree server | vLLM | A Python ML library can ship an in-tree OpenAI-compatible server while clients remain external. | https://github.com/vllm-project/vllm/blob/bcf5cac9fb956788f649d1f5297b74c886a9d6d3/README.md#L64-L74 |
| Service packaging | BentoML | Packaging model + service + dependencies is supported, but CWD packaging creates discipline needs. | https://github.com/bentoml/BentoML/blob/32230a5276a8da8b23c4a06a9ec6272c1993451a/docs/source/build-with-bentoml/asgi.rst#L5-L18 |
| YAML-driven production serving | Ray Serve | Production updates should avoid in-place mutation; use new deployment/traffic switch. | https://docs.ray.io/en/latest/serve/advanced-guides/inplace-updates.html |
| Library/server plus external UI | TGI + ChatUI | Server can live with backend project while UI is separate. | https://github.com/huggingface/text-generation-inference/blob/b4adbf2f6e2e721280bd0ea5f91d70f7d033f5ed/docs/source/basic_tutorials/consuming_tgi.md#L182-L186 |
| Lean library plus examples elsewhere | Transformers.js | Library stays lean; demos/examples can live outside core. | https://github.com/huggingface/transformers.js/blob/f7487c737aa8cafbc106c9adf69dc9578c8f3fe0/README.md#L26-L34 |
| Product monorepo that later split frontend | ComfyUI | Product UI/server monorepo can hit release-cadence mismatch and split FE later. | https://github.com/comfyanonymous/ComfyUI/blob/fed8d5efa6b70d5b24c4c33cb643bfccc39d45b5/README.md#L131-L149 and https://github.com/Comfy-Org/ComfyUI_frontend/blob/60f789d58070a9d1d789b260f83c36d7293a39f0/README.md#L31-L60 |
| Tightly coupled UI/server product | AUTOMATIC1111 SD WebUI | Product repos can couple UI/server tightly, but security surface becomes product-sized. | https://github.com/AUTOMATIC1111/stable-diffusion-webui/blob/82a973c04367123ae98bd9abdf80d9eda9b910e2/webui.py#L48-L104 |
#### Precedent synthesis
- Strong precedents exist for a Python ML library shipping a server entrypoint.
- Strong precedents exist for keeping frontend/product UI out of the backend
library repo.
- The cited set does not contain a clean precedent for merging a Next.js
product into a Python ML library namespace.
- The most applicable pattern is **backend/server in the ML project,
product UI outside**.
### Recommendation
#### Recommend Option D, constrained: backend merges as generic FastVideo streaming; frontend stays separate
Recommendation: follow **Option D** as the long-term architecture, but keep
the backend merge generic. In practice, this means continuing the current
public FastVideo path:
- `fastvideo.entrypoints.streaming.*` owns reusable streaming runtime.
- `fastvideo.entrypoints.streaming.gpu_pool` owns generic GPU worker pools.
- `fastvideo.entrypoints.streaming.prompt.*` owns provider-agnostic prompt
operations.
- `fastvideo.entrypoints.streaming.router.*` owns FastVideo-aware routing.
- `fastvideo.api` owns typed construction, requests, results, events, and
continuation state.
- Dreamverse keeps product-only FE, curated presets, prompt UX, product
routes, and launch scripts.
This is effectively the path already underway in PRs #1257, #1258, #1284,
#1286, and #1288:
[PR roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L21-L42).
#### Why not Option A as the final answer?
Option A is the fastest near-term release posture and should be used as the
immediate migration posture. However, plain status quo is not enough for
the ultimate goal because reusable backend pieces still need to live in
public FastVideo so Dreamverse can stop reaching into internal code. That
work is already partly complete:
- GPU pool: [D-12](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L47-L116).
- Prompt enhancer: [PR roadmap 7.7](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L32-L35).
- Streaming auxiliaries: [PR roadmap 7.8](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L35-L36).
- Router: [D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L118-L155).
- `generate_async`: [streaming-server unlock PR](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L314-L345).
So the practical answer is:
- **Near term:** Option A operationally, after docs/scripts are fixed.
- **Architecture target:** Option D, with generic backend ownership in
FastVideo and product ownership in Dreamverse.
#### Why not Option B?
Option B makes cross-repo coordination easier but imports frontend build,
package, and CI complexity into FastVideo. That is unnecessary while a
normal package dependency plus contract tests can guard the integration.
FastVideo's current repo structure is a Python package with examples and
docs, not a product monorepo:
[codebase map](file:///home/william5lin/FastVideo/.agents/memory/codebase-map/README.md#L5-L75).
#### Why not Option C?
Option C makes the product backend a public FastVideo namespace. That is
only appropriate if the team wants to support Dreamverse as a first-class
FastVideo product surface. Today the known public obligations are generic:
typed requests, streaming server, GPU pool, prompt provider protocol,
router, NVFP4, and Dynamo event APIs. Product-only Dreamverse behavior does
not need to become framework API.
#### Conditions that would change the recommendation
Move from constrained D toward **C** only if all of these become true:
1. Dreamverse is declared the canonical FastVideo serving product.
2. Product routes such as curated presets and prompt-system config are
accepted as public FastVideo API.
3. FastVideo maintainers accept the security and support surface.
4. Release cadence for product UX and FastVideo core is intentionally
coupled.
5. Frontend packaging/static asset strategy is explicitly owned by
FastVideo.
Move from constrained D back toward **A** if any of these become true:
1. Prompt enhancement, router, or GPU pool turn out to be Dreamverse-only.
2. No second user appears for the streaming backend outside Dreamverse.
3. FastVideo maintainers want to minimize serving surface and publish only
Python library APIs.
4. Dreamverse needs product changes faster than FastVideo can release.
5. Security review rejects in-tree serving/router responsibilities.
### Migration sketch for the recommended path
#### Phase 0 — Land the public backend stack
- **Effort:** Large, already in flight.
- **Owner:** FastVideo public.
- **Files:** #1288 scope, especially `fastvideo/api/`,
`fastvideo/entrypoints/video_generator.py`,
`fastvideo/entrypoints/streaming/`, LTX-2 pipeline stages, NVFP4 files,
and contract tests.
- **Exit criteria:** #1288 merges; public `fastvideo.api.VideoEvent` and
`VideoGenerator.generate_async` are available:
[results.py](file:///home/william5lin/FastVideo/fastvideo/api/results.py#L109-L164),
[video_generator.py](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332).
#### Phase 1 — Fix Dreamverse dependency docs and bootstrap
- **Effort:** Small.
- **Owner:** Dreamverse.
- **Files:**
- [README.md](file:///home/william5lin/Dreamverse/README.md#L76-L109)
- [bootstrap script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11)
- [pyproject.toml](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22)
- [uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722)
- **Exit criteria:** no user-facing docs or scripts mention
`FastVideo-internal` as the expected dependency path.
#### Phase 2 — Add public health/readiness/status route compatibility
- **Effort:** Medium/Large.
- **Owner:** FastVideo public.
- **Files likely to touch:**
- `fastvideo/entrypoints/streaming/server.py::build_app`
- new `fastvideo/entrypoints/streaming/health.py`
- tests under `fastvideo/tests/entrypoints/streaming/`
- **Source route shapes:**
[Dreamverse health routes](file:///home/william5lin/Dreamverse/server/routes/health.py#L34-L79).
- **Tracking:** [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
- **Exit criteria:** Dreamverse FE can target public `build_app` for
`/healthz`, `/readyz`, `/status`, and `/v1/stream`; product-only routes
remain feature-detected.
#### Phase 3 — Replace Dreamverse prompt enhancer fork
- **Effort:** Medium.
- **Owner:** Dreamverse.
- **Files likely to touch:**
- new `Dreamverse/server/prompting/_internal_compat.py`
- `Dreamverse/server/runtime.py`
- `Dreamverse/server/main.py`
- `Dreamverse/server/prompt_enhancer.py`
- **Public API:**
[PromptEnhancer](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/enhancer.py#L66-L142),
[LLMProvider](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/providers/base.py#L63-L75).
- **Tracking:** [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
- **Exit criteria:** Dreamverse no longer carries a full local fork for the
generic prompt operations public FastVideo already owns.
#### Phase 4 — Decide and implement `cerebras_ifm`
- **Effort:** Small decision plus small/medium implementation.
- **Owner:** Team decision, then Dreamverse or FastVideo.
- **Default recommendation:** Dreamverse-side custom provider.
- **Public schema source:**
[PromptEnhancerConfig](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L229-L235).
- **Tracking:** [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
- **Exit criteria:** Dreamverse IFM provider works after prompt fork removal.
#### Phase 5 — Move streaming demo config into FastVideo examples
- **Effort:** Small.
- **Owner:** FastVideo public.
- **Source:**
[Dreamverse streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L149).
- **Target:** `examples/serving/streaming_demo.yaml`.
- **Exit criteria:** users can reproduce the typed streaming path from the
FastVideo repo without checking out Dreamverse.
#### Phase 6 — Remove `experimental["pipeline_config"]` where practical
- **Effort:** Medium for `layer_profile`; Large for a full typed
`dit_config.quant_config` carrier.
- **Owner:** FastVideo public, then Dreamverse cleanup.
- **Tracking:**
[open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260),
[quantization follow-up](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L175-L200).
- **Exit criteria:** Dreamverse can express its quant layer profile through
typed config instead of in-memory mutation.
#### Phase 7 — Decide standalone upsampler CLI
- **Effort:** Small/Medium.
- **Owner:** FastVideo public.
- **Input:** internal standalone utility
[upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L120-L180)
and internal CLI
[cli/upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L48-L130).
- **Public alternative:** SR refine stage already covers in-pipeline latent
upsampling:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L116-L180).
- **Exit criteria:** explicit decision: port CLI, document refine-stage-only
support, or defer.
#### Phase 8 — Validate LTX-2 stage parity and offload residuals
- **Effort:** Small/Medium for stage parity; Medium for layerwise offload.
- **Owner:** FastVideo public.
- **Stage source:**
[public LTX-2 stages](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/).
- **Offload source:**
[internal layerwise offload](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
- **Exit criteria:** no known behavior gap between internal and public LTX-2
stages; offload is either deliberately deferred or ported with tests.
### Open questions
1. **Which provider path for `cerebras_ifm`?**
- Public provider or Dreamverse-side custom provider?
- Default recommendation: Dreamverse-side unless there is another user.
- Source: [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
2. **Should public FastVideo support standalone LTX-2 file upsampling?**
- If yes, port internal CLI.
- If no, document that SR support is pipeline-refine only.
- Sources: [internal CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35),
[public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
3. **Does Dreamverse need layerwise CPU offload?**
- If memory-tight deployments require it, port as generic FastVideo.
- Otherwise defer.
- Source: [internal offload manager](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
4. **How much Dreamverse route surface should FastVideo own?**
- Health/readiness/status should migrate because they are part of
streaming-server compatibility.
- Curated presets and prompt-system config should stay Dreamverse-side
unless product policy changes.
- Source: [route contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L232-L259).
5. **Should `layer_profile` be the only near-term quant typed addition?**
- Adding `layer_profile` is bounded.
- A typed carrier for arbitrary mutated `PipelineConfig` is larger design
work.
- Source: [quantization follow-ups](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L175-L200).
6. **When does Option D become Option C?**
- Only if Dreamverse backend routes become public FastVideo product API.
- Until then, keep generic streaming code in FastVideo and product code in
Dreamverse.
---
## Part 3 — Action items
1. **P0 / S — Update Dreamverse README dependency notes.**
- Replace `../FastVideo-internal` with public FastVideo instructions.
- Preserve local editable `../FastVideo` dev flow where useful.
- Source: [README stale lines](file:///home/william5lin/Dreamverse/README.md#L76-L109).
2. **P0 / S — Replace or remove private FastVideo bootstrap script.**
- Current script clones `FastVideo-internal` and verifies imports from it.
- Source: [script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11).
3. **P1 / M-L — Add `/healthz`, `/readyz`, and `/status` to public `build_app`.**
- Keep `/curated-presets` and prompt-system config in Dreamverse.
- Source: [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
4. **P1 / M — Replace Dreamverse prompt-enhancer fork with compat shim.**
- Wrap public `PromptEnhancer`.
- Keep only Dreamverse-specific metadata and product fallback behavior.
- Source: [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
5. **P1 / S-M — Decide `cerebras_ifm` provider path.**
- Default: Dreamverse custom provider via `register_provider`.
- Source: [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
6. **P2 / S — Add `examples/serving/streaming_demo.yaml` to FastVideo.**
- Start from Dreamverse config and remove Dreamverse-private comments.
- Source: [streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L149).
7. **P2 / M — Verify each public LTX-2 colocated stage against internal behavior.**
- Start with refine, denoising, latent prep, image conditioning, text
encoding, and audio decoding.
- Source: [public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
8. **P2 / M — Add typed `transformer_quant_layer_profile` if Dreamverse needs it.**
- Thread schema → compat → `FastVideoArgs._apply_transformer_quant`.
- Source: [open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260).
9. **P2 / S-M — Decide standalone LTX-2 upsampler CLI support.**
- Port internal CLI only if file-to-file upsampling is a public workflow.
- Source: [internal upsample CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35).
10. **P2 / S — Fix pre-existing AbsMaxFP8 test failure separately.**
- Do not block Dreamverse migration on it.
- Source: [open item #2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L100-L117).
11. **P2 / S-M — Document public streaming install extras and dependencies.**
- Include router `websockets`, prompt enhancer provider SDKs, and optional
safety classifier extras.
- Source: [D-16 dependency note](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L224-L242).
12. **P2 / S — Keep contract tests in the FastVideo CI path.**
- Guard Dreamverse shape, Dynamo shape, and async events.
- Sources: [Dreamverse test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26),
[Dynamo test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19),
[async test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L1-L7).
13. **P3 / M — Defer layerwise offload until a deployment needs it.**
- Port only as generic FastVideo utility with tests.
- Source: [internal offload](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
14. **P3 / M — Defer StepVideo public parity for this integration.**
- Dreamverse model registry is LTX-2/LTX-2.3 only.
- Source: [Dreamverse config](file:///home/william5lin/Dreamverse/server/config.py#L28-L45).
15. **P3 / S — Do not add debug-only fields to public schema by default.**
- Keep them private unless there is a user-facing debugging workflow.
- Source: [internal debug args](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203).
16. **P3 / M — Keep router active-active and sticky routing deferred.**
- Add only when session-routing evidence justifies it.
- Source: [D-15 action items](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201).
17. **P3 / S — Preserve Dynamo as an external backend package.**
- FastVideo should expose typed API; Dynamo code lives in Dynamo.
- Source: [Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210).
18. **P3 / S — After #1288 merges, update memory-dir state.**
- Mark item D resolved and update branch tips.
- Source: [runbook post-merge steps](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/runbook.md#L51-L70).
19. **P3 / S — Remove stale split-PR mental model from follow-up docs.**
- #1288 is the current vehicle; split bookmarks are historical.
- Source: [D-17 implications](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L34-L45).
20. **P3 / S — Keep product-only Dreamverse frontend out of FastVideo unless explicitly re-scoped.**
- This preserves release cadence and avoids packaging bloat.
- Source: [Dreamverse baseline](file:///home/william5lin/Dreamverse/README.md#L5-L16).
@@ -0,0 +1,498 @@
# Open Threads — Active Follow-Ups
Live work items with priority, effort estimate, dependencies, and
recommended next action.
For why each item is open see [decisions-log.md](decisions-log.md). For
PR-level context see [pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-05 (strategy reversal — PR #1287 CLOSED, replaced
by mega-PR #1288 on `will/ltx2_sr_port` @ `b36bdbc9` covering the full
6-layer stack at once. See [decisions-log.md D-17](decisions-log.md#d-17).
Item D resolution gate is now #1288 merge instead of #1287; same content,
different vehicle.).
## Priority overview
| # | Pri | Item | Effort | Unblocks |
|---|---|---|---|---|
| **D-8** | High | Verify `ltx2_image_crf` post-`d80c2a8` | 10 min | Confirms typed stage-override path actually flows; closes a latent silent-drop bug |
| **1** | High | Migrate `/healthz`+`/readyz`+`/status` into FastVideo `build_app` | M-L | Closes BE_FLAVOR=fastvideo FE-compatibility; closes streaming-upstream contract debt |
| **2** | High | Fix pre-existing AbsMaxFP8 test failure | S | Self-contained quantization tech debt |
| **VPO** | High | Decide `video_position_offset_sec` semantics (a vs b) | 30 min | Unblocks PR 7.6 state emission |
| **D** | 🟢 in flight | Implement `generate_async` — content shipped in mega-PR **#1288** on `will/ltx2_sr_port` @ `b36bdbc9` (was #1287, CLOSED + re-routed per [D-17](decisions-log.md#d-17)) | L | Closes Q-5/Q-9/PR-7.5 TODOs simultaneously; enables Dynamo backend; unblocks audio re-encode; enables `GpuPool.run_async()` migration (D-12-B). Resolution gate: #1288 merge. |
| **DR-1** | High | Dreamverse: create `prompting/_internal_compat.py` shim + replace local `prompt_enhancer.py` (1933 LOC) — **PR #1258 has merged (`f673423b`); now actionable** | M (~150-200 LOC shim, replace upstream wiring) | Lets Dreamverse stop carrying a 1933-LOC fork |
| **DR-2** | Med | Decide `cerebras_ifm` provider path: (a) public Literal + `CerebrasIFMProvider` shipped, OR (b) Dreamverse-side custom provider via `enhancer.register_provider(...)` | S (decision) + S-M (impl) | Resolves the cerebras_ifm gap left by PR #1258. Same item as legacy #3 below; DR-2 is the Dreamverse-side framing. |
| **3** | Med | Add `cerebras_ifm` to `PromptEnhancerConfig.provider` Literal + provider | S-M | Public-side resolution if DR-2 picks (a) |
| **4** | Med | Expose `layer_profile` on typed `engine.quantization` | M | Removes Dreamverse's `experimental["pipeline_config"]` dodge for stage profiles |
| **5** | Med | Design typed `dit_config.quant_config` carrier | L design + L impl | Removes broader `experimental["pipeline_config"]` escape hatch |
| **SBS** | Med | `SessionStore` / `BlobStore` lifecycle policy | M design | Needed in PR 7.5 design pass |
| **D-12-A** | Med | Update `GpuPool` ABC docstring: mark "API may change post-PR-7.10; experimental / server-internal" | trivial | Prevents accidental promotion of streaming-internal API to framework-level |
| **D-12-B** | Med | Replace `GpuPool.run() -> Any` with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle | M | Closes the streaming-server cancellation TODO; converges with `generate_async` |
| **D-13-A** | Med | Document `fastvideo.entrypoints.streaming.prompt.*` in user-facing docs as "streaming-server scoped"; avoid framework-level framing | trivial (docs only) | Keeps future move to `fastvideo.prompt.*` cheap |
| **D-13-B** | Low | Add optional `client_factory` parameter to `LLMProvider` for `httpx.AsyncClient` pooling | S | Only if metrics show connect/TLS overhead is meaningful |
| **D-12-C** | Low | Avoid locking `PoolAssignment.gpu_id: int` as public; rename to `worker_id` (already exists) or add `device_ids: list[int]` for topology-aware pooling | S | Future multi-GPU-per-worker refactor stays cheap |
| **6** | Low | Audio attention quantization profile + test update | S | Future audio quant exploration |
| **7** | Low | Schema parity inventory cleanup (env-driven prompt fields) | S-M | Long-term consistency |
| **8** | Low | Stale `apps/web/test-results/` dir cleanup | trivial | Cosmetic |
| **11** | Low | Promote LTX-2 prompt orchestration (locked segments, segment_prompts JSON shape, rollout id/label) to `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` | M | Resolves Q-2 from decisions-log when a second LTX-2-style consumer appears |
| **12** | Low | When streaming server starts using `PromptSafetyFilter`, ensure operator-visible logging on `SafetyDecision.UNAVAILABLE` results | trivial | Surfaces degraded-safety state to operators (per D-14 Watch-Out item) |
| **13** | Low | When sticky session routing is needed, add `ReplicaRegistry.select(routing_key: str | None = None)` and document where `session_id` lives (WS URL/header preferred over first JSON frame) | M | Forward-compat from D-15 — keeps the door open without buffering/peeking |
| **14** | Low | At higher load, add `_bridge_session()` max-size + timeout limits OR recommend Envoy/HAProxy in front | S-M | The libraries' basic backpressure suffices for MVP; document the limit per D-15 |
| **15** | Low | If active-active multi-primary becomes a requirement, define behavior (round-robin within healthy primaries, weighted, sticky-by-key) | M | Currently `RouterConfig.__post_init__` rejects multi-primary; D-15 deferred until evidence |
| **~~Source-doc disposition~~** | ~~Med~~ | ~~Disposition of 7 untracked source docs~~ | ~~trivial~~ | ✅ **Resolved 2026-05-03** — moved into [source-archive/](source-archive/) |
| **~~9~~** | ~~Low~~ | ~~Commit-message cleanup: PR 8's 3 commits still have `[8/n] Improve API:` prefix~~ | ~~S~~ | ✅ **Resolved 2026-05-04** — bundled into the will/api_7.8 prep rebase. PR 8's 3 commits now read `[type] streaming: ...` |
| **~~10~~** | ~~Low~~ | ~~Commit-message cleanup: PR 7.8/7.9 commits have `streaming: streaming X` duplication~~ | ~~S~~ | ✅ **Resolved 2026-05-04** — bundled into the will/api_7.8 prep rebase. 3 commits dedup'd. |
---
## High priority
### D-8: Verify `ltx2_image_crf` typed flow post-`d80c2a8`
**Why:** Apr 26 dreamverse_review documented this field getting silently
dropped by the public `SamplingParam`. May 2 `d80c2a8` (Dreamverse)
refactored to typed `GeneratorConfig` + `preset_overrides`. Whether
`image_crf` now flows through `request.stage_overrides.refine.image_crf`
(per [design.md](design.md) mapping) or is still dropped is unverified.
**Action:**
1. Read [`Dreamverse/server/video_generation.py`](file:///home/william5lin/Dreamverse/server/video_generation.py)
post-`d80c2a8` for `image_crf` handling
2. Trace through to FastVideo's `request.stage_overrides.refine.image_crf`
3. Confirm runtime consumption in [`fastvideo/pipelines/basic/ltx2/`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/)
**Effort:** 10 min, no code changes.
**Outcome:** Either confirms working OR identifies bug → opens fix item.
### Item #1: Migrate `/healthz`+`/readyz`+`/status` into `build_app`
**Why:** Today
[`fastvideo.entrypoints.streaming.server.build_app`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py)
exposes only `/health` + `/v1/stream`. Dreamverse FE expects all of
`/healthz`, `/readyz`, `/status`, `/curated-presets`,
`/prompt-system-config`, devtools.
The streaming-server-upstream plan (line 84) explicitly lists
`/healthz`+`/readyz`+`/status` as part of the contract that the upstream
of `realtime/` → `streaming/` must preserve. They were deferred from
PR 7.5's MVP. `/curated-presets` and `/prompt-system-config` are
operator-side and stay in Dreamverse (FE feature-detects).
**Action:**
1. Read PR 7.5 (#1251) `build_app` to scope what's there
2. Read [`Dreamverse/server/routes/health.py`](file:///home/william5lin/Dreamverse/server/routes/health.py)
for the route shapes Dreamverse already consumes
3. Propose route migration as commit on top of `will/api_7.5` or as
part of PR 7.10 cycle
4. Land
**Effort:** Medium-Large (route shapes need preservation; tests).
**Dependencies:** None blocking; can land anytime.
**Files likely to touch:**
- `fastvideo/entrypoints/streaming/server.py::build_app`
- New `fastvideo/entrypoints/streaming/health.py`
- Tests in `fastvideo/tests/entrypoints/streaming/`
### Item #2: AbsMaxFP8 pre-existing test failure
**Why:** [`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_absmax_fp8.py)
fails with `AssertionError not raised`. Pre-existing on `main`; verified
NOT introduced by NVFP4 work via `git stash`.
**Action:**
1. `git log --oneline fastvideo/tests/ops/quantization/test_absmax_fp8.py`
to find when it last passed
2. Either:
- Restore the assert in `AbsMaxFP8LinearMethod.create_weights` if
intentional behavior was lost
- Drop the test if assert is no longer correct
3. Verify
**Effort:** Small.
**Dependencies:** None.
### Item VPO: `video_position_offset_sec` semantics
**Why:** Per
[`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py),
`LTX2ContinuationState.video_position_offset_sec` exists as a state
field. Two valid interpretations:
- **(a) Persistent across segments** — accumulating time offset for long
sessions; useful for time-coherent audio chaining.
- **(b) Per-segment hint that rides on the carrier** — runtime
overwrites every time; field is harmless redundancy.
Dreamverse computes `prefix_sec = float(audio_extra) / 24.0` per segment
in `apply_audio` and currently does NOT persist it on
`ContinuationState`. Field's docstring leans toward (b).
**Decision deadline:** before PR 7.6 starts emitting/consuming the
field (PR 7.6 branch is ready, not yet PR'd).
**Action:**
1. Confirm field's intended semantics with audio team
2. If (a): document the accumulation rule explicitly + add tests
3. If (b): leave docstring as-is + add test confirming overwrite
**Effort:** 30 min discussion + small implementation.
### Item D: Implement `generate_async` (PR 7.10)
**Why:** Highest leverage. Closes:
- D-5 / Q-5: audio re-encode for cross-segment continuity
- Q-9: Dynamo progress passthrough (deferred)
- PR 7.5's mid-segment cancellation TODO
- Unblocks Dynamo native backend integration
- **D-12-B**: enables `GpuPool.run() -> run_async() -> AsyncIterator[VideoEvent]` migration
**Action:** See [streaming-server.md](streaming-server.md) "PR 7.10 — the
unlock PR" section for scoping.
**Effort:** Large.
**Dependencies:** Best after PR 7.6 lands (gpu_pool upstream).
**Files:**
- `fastvideo/entrypoints/video_generator.py` — add `generate_async`,
refactor `generate_video` as wrapper
- `fastvideo/api/results.py` — add `VideoEvent`/`VideoProgressEvent`/
`VideoPartialEvent`/`VideoFinalEvent`
- `fastvideo/entrypoints/streaming/server.py` — consume `generate_async`,
remove TODO markers
- `fastvideo/entrypoints/streaming/gpu_pool.py` — add `run_async()`
forwarding events from worker to caller
- New `fastvideo/tests/entrypoints/test_generate_async.py`
- New `fastvideo/tests/contract/test_dynamo_shape.py` (already in PR 8)
### Item DR-1: Dreamverse — replace local `prompt_enhancer.py` with public + compat shim
**Why:** Today Dreamverse carries `Dreamverse/server/prompt_enhancer.py`
(1933 LOC) — a local copy/derivative of the FastVideo-internal version.
After PR #1258 merges, Dreamverse should switch to the public
`fastvideo.entrypoints.streaming.prompt.PromptEnhancer` and delete most
of the local module.
**Migration shape:**
1. **Create** `Dreamverse/server/prompting/_internal_compat.py` (~150-200 LOC):
- Wraps public `PromptEnhancer.enhance()` → returns `EnhanceResult` shape Dreamverse expects
- Wraps public `PromptEnhancer.auto_extend()` — JSON-parses `LLMResponse.content` into `{"next_prompt": "..."}`
- Wraps public `PromptEnhancer.rewrite()` — JSON-parses into `{"segment_prompts": [...]}` with lenient fallback for malformed JSON
- Layers locked-segment + rollout_id + rollout_label metadata back on top
2. **Update** `Dreamverse/server/runtime.py + main.py` — replace `from prompt_enhancer import PromptEnhancer` with `from prompting._internal_compat import PromptEnhancer`
3. **Delete most of** `Dreamverse/server/prompt_enhancer.py` (1933 LOC). Keep only the bits that don't have a public equivalent:
- Race-based parallel fallback (`_run_provider_race`) — Dreamverse-specific tail-latency optimization
- `cerebras_ifm` provider — pending DR-2 decision
- Multi-classifier prompt safety (NSFW + hate-speech chained) — public ships single classifier
4. **Tests** — verify Dreamverse session controllers still see the expected response shapes through the shim
**Effort:** Medium (~150-200 LOC shim + replace upstream wiring + delete 1700+ LOC local module + test fixture updates).
**Dependencies:**
- PR #1258 must merge first (publishes `fastvideo.entrypoints.streaming.prompt.*`)
- DR-2 informs the cerebras_ifm path
**Files:**
- New: `Dreamverse/server/prompting/_internal_compat.py`
- Modified: `Dreamverse/server/runtime.py`, `Dreamverse/server/main.py`
- Mostly deleted: `Dreamverse/server/prompt_enhancer.py`
---
## Medium priority
### Item DR-2: Decide `cerebras_ifm` provider path
**Why:** Public PR #1258's `PromptEnhancerConfig.provider` is
`Literal["cerebras", "groq"]`. Internal supports `"cerebras_ifm"` (the
Cerebras IFM API endpoint with different auth). Dreamverse needs
`cerebras_ifm` working post-migration.
Two options:
| Option | Approach | Pros | Cons |
|---|---|---|---|
| **(a) Public** | Add `"cerebras_ifm"` to public Literal + ship `CerebrasIFMProvider` in `fastvideo/entrypoints/streaming/prompt/providers/cerebras_ifm.py` | Discoverable; users with IFM access can use typed config | Adds ~50 LOC + Literal extension to public surface |
| **(b) Dreamverse-side** | Implement `CerebrasIFMProvider` Dreamverse-side as a custom `LLMProvider`, register via `enhancer.register_provider(CerebrasIFMProvider())` | Zero public surface change; private endpoint stays private | Slightly more boilerplate Dreamverse-side; not surfaced to non-Dreamverse users |
**Recommendation:** Option (b) is more contained. Option (a) is more
discoverable. Default to (b) unless there's a third-party user who needs
IFM access. The Dreamverse-side PR carrying DR-1 is the natural place to
make this decision.
**Effort:** Small (decision) + Small-Medium (implementation).
**Dependencies:** DR-1 (compat shim creation).
### Item #3: `cerebras_ifm` provider in public Literal
**Why:** Same item as DR-2 from the public-side framing. If DR-2 picks
option (a), this is the implementation. If DR-2 picks option (b), this
item is closed without implementation.
**Action:** See DR-2.
**Effort:** S-M.
### Item #4: Expose `layer_profile` on typed `engine.quantization`
**Why:** Today `transformer_quant: "NVFP4"` always constructs
`NVFP4Config()` with default `layer_profile="refine"`. Dreamverse
dodges via `experimental["pipeline_config"]`.
**Action:**
1. Add `transformer_quant_layer_profile: str | None = None` to
`QuantizationConfig` in [`schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py)
2. Thread through [`compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py)
3. Update `_apply_transformer_quant` in
[`fastvideo_args.py`](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py)
to pass profile
4. Update Dreamverse to drop the `experimental["pipeline_config"]`
dodge in favor of typed knob
5. Tests in [`test_typed_quant_flow.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_typed_quant_flow.py)
**Effort:** Medium.
**Files:** schema.py, compat.py, fastvideo_args.py, test_typed_quant_flow.py,
+ Dreamverse/server/video_generation.py.
### Item #5: Typed `dit_config.quant_config` carrier
**Why:** The `experimental["pipeline_config"]` escape hatch in
Dreamverse should eventually become a typed field. Design TBD.
**Action:** Heaviest design work. Should consult Oracle.
**Effort:** Large design + Large implementation.
**Dependencies:** #4 should land first; this is the "final form" of #4.
### Item SBS: `SessionStore` / `BlobStore` lifecycle policy
**Why:** PR 7's in-memory implementations have no eviction, no TTL, no
automatic blob cleanup on state replacement. Documented as per-deployment
policy decision.
When PR 7.5/7.6 land the live consumer, who owns:
- bounded session capacity (LRU? TTL? hard max?)
- blob `drop()` chained when state is replaced
- session expiry on websocket disconnect
**Recommendation:** streaming server's session manager. Worth stating
explicitly in PR 7.5's design.
**Effort:** Medium design + small implementation.
### Item D-12-A: Update `GpuPool` ABC docstring — mark experimental
**Why:** Per D-12 in [decisions-log.md](decisions-log.md), `GpuPool`
should be documented as "API may change post-PR-7.10; experimental /
server-internal" to prevent accidental promotion of streaming-internal
API to framework-level. PR #1257 merged without this caveat.
**Action:** Edit
[`fastvideo/entrypoints/streaming/gpu_pool.py`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/gpu_pool.py)
class docstring on `GpuPool` ABC. Add a note: "API may change post-PR-7.10
when run_async() lands; treat as server-internal for now."
**Effort:** Trivial.
**Dependencies:** None.
### Item D-12-B: Replace `GpuPool.run() -> Any` with `run_async() -> AsyncIterator[VideoEvent]`
**Why:** Per D-12, this is the canonical evolution post-PR-7.10. Closes
the streaming server's cancellation TODO and converges the streaming +
OpenAI + Dynamo consumers on a single async API.
**Action:** As part of PR 7.10 cycle:
1. Add `GpuPool.run_async(session_id, request) -> AsyncIterator[VideoEvent]`
2. Worker forwards events through `result_queue` with type discriminator
3. Streaming server replaces `await pool.run(...)` with `async for event in pool.run_async(...)`
4. Sync `run()` becomes a thin compat wrapper that collects events and returns the final
5. Cancellation propagates: client disconnect → `asyncio.CancelledError` → worker stops mid-step
**Effort:** Medium. Adds ~50-100 LOC + tests.
**Dependencies:** Item D (PR 7.10 — `generate_async` on `VideoGenerator`).
### Item D-13-A: Document `streaming/prompt/*` as streaming-scoped
**Why:** Per D-13 in [decisions-log.md](decisions-log.md), the prompt
enhancer is currently scoped to streaming-server use even though the
abstraction is general. Phrase user-facing docs as "streaming-server
prompt enhancement" to keep future move to `fastvideo.prompt.*` cheap.
**Action:** When PR 12 (docs migration) is written, the prompt enhancer
section should:
- Be titled "Streaming Server Prompt Enhancement", not "Prompt API"
- Note the 3 fixed operations (`enhance` / `auto_extend` / `rewrite`) are
shaped by LTX-2 streaming session needs
- Note that consumers wanting custom prompt operations can use
`provider.complete()` directly with their own LLMRequest
- Avoid `from fastvideo import LLMProvider` exports until a second
consumer exists
**Effort:** Trivial (docs only).
**Dependencies:** PR 12 (docs migration).
---
## Low priority
### Item D-13-B: Optional `client_factory` parameter for `httpx.AsyncClient` pooling
**Why:** Today `_openai_compat.py` instantiates `httpx.AsyncClient` per
call (no connection pooling). Reviewer flagged inefficient. Team chose
simplicity for the expected scale (~6-10 enhancer calls per LTX-2
session). If real-world metrics show connect/TLS overhead is meaningful,
add an optional `client_factory: Callable[[], httpx.AsyncClient] | None`
parameter to providers so they can share a pool.
**Action:** Only when metrics justify. Add `client_factory=None` parameter
to `CerebrasProvider` / `GroqProvider` constructors and pass through to
`complete_openai_compatible()`. Default to current per-call behavior.
**Effort:** Small.
**Dependencies:** None blocking; only act on real perf data.
### Item D-12-C: Avoid locking `PoolAssignment.gpu_id: int` as public
**Why:** Today `PoolAssignment` exposes `gpu_id: int`, assuming
one-GPU-per-worker. Future topology-aware pooling may need
`device_ids: list[int]` (one worker = group of GPUs running internal
`MultiprocExecutor`). Don't freeze the int field as public API.
**Action:**
- Treat `gpu_id` as a current-impl detail; prefer `worker_id` (already
exists, is stable identifier)
- When a worker actually spans multiple GPUs, add
`PoolAssignment.device_ids: list[int]` and let `gpu_id` be `device_ids[0]`
for backward compat
- Or rename to `gpu_id` → `device_id` with deprecation alias
**Effort:** Small (1 field rename + alias).
**Dependencies:** Driven by an actual future "one worker = many GPUs" use case. Don't act preemptively.
### Item #6: Audio attention quantization profile
**Why:** Today audio attn and FFN are bf16. If an audio-quant profile
is added to `NVFP4Config.fp4_layers`, update
[`test_basic_av_block_propagates_quant_config_to_all_children`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py).
**Effort:** Small (one test + one config field).
### Item #7: Schema parity inventory cleanup
**Why:** A few internal-only fields are not exposed publicly:
- `PROMPT_HTTP_TIMEOUT_MS`
- `PROMPT_INITIAL_STAGE_TIMEOUT_MS`
- `PROMPT_TEMPERATURE`
- `PROMPT_MAX_COMPLETION_TOKENS`
- `PROMPT_AUTO_SLEEP_MS`
- `PROMPT_AUTO_TIMEOUT_MS`
- curated-presets file paths
These flow via env vars on `dreamverse-server` today. If
`fastvideo serve --config` becomes the canonical entrypoint, they need
typed homes.
**Effort:** Small-Medium.
### Item #8: Stale `apps/web/test-results/` directory
**Why:** Cosmetic. `.gitignore` entry hides it from `git status`, but
the dir has stale `.last-run.json` (45 bytes) from a prior Playwright
run.
**Action:** `rm -rf apps/web/test-results` whenever convenient.
**Effort:** Trivial.
### ~~Item #9~~ + ~~#10~~: Commit-message cleanups — ✅ Resolved 2026-05-04
Both items resolved during the `will/api_7.8` prep rebase. A targeted
conditional script (`/tmp/opencode/cleanup_subjects_v2.sh` — only amends
when text actually changes) ran across 33 commits, modified 6:
- **#9 fix**: extended the regex from `\[\d+\.\d+/n\]` to
`\[\d+(\.\d+)?/n\]` so single-digit prefixes match. PR 8's 3 commits
now read `[type] streaming: ...` instead of `[type] [8/n] Improve API: ...`.
- **#10 fix**: added second substitution `streaming: streaming X` →
`streaming: X`. PR 7.8 / 7.9 commits no longer have the duplication.
The conditional check skipped pre-commit-hook flakiness on no-op amends
(unlike the earlier first attempt). All affected commits verified clean
post-rebase.
### Item #11: Promote LTX-2 prompt orchestration to public (when 2nd consumer exists)
**Why:** Per Q-2 in [decisions-log.md](decisions-log.md) and D-13's
"missing alternative", the LTX-2-specific orchestration (locked
segments, segment_prompts JSON shape, rollout id/label, lenient JSON
parsing) currently stays Dreamverse-side per DR-1. If a second
LTX-2-style consumer appears (e.g. another video model with multi-segment
continuation needing the same prompt orchestration), promote this layer
to `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration`.
**Action:** Wait for a second consumer to materialize. Until then, the
orchestration stays in Dreamverse's `_internal_compat.py` shim (DR-1).
**Effort:** Medium when triggered.
**Dependencies:** A second consumer.
---
## Recommended pull order
If you have unbounded time and want to maximize forward progress:
1. **D-8 verify** (10 min) — eliminates uncertainty
2. **D-12-A docstring** (trivial) — caveat the GpuPool API publicly
3. **Item #2 AbsMaxFP8** (S) — clears tech debt
4. **Item VPO video_position_offset_sec** (30 min) — unblocks PR 7.10 (since 7.6 has merged, this is now scoped to whatever consumer first reads the field)
5. **DR-1 + DR-2 Dreamverse migration** (M) — **now unblocked since PR #1258 merged**; replaces 1700+ LOC of local fork
6. **Item #4 layer_profile** (M) — closes Dreamverse quant escape hatch
7. **Item #1 build_app routes** (M-L) — closes FE-compat
8. **Item D generate_async** (L) — unlock PR; brings along D-12-B (run_async) + closes Q-5/Q-9/PR-7.5 TODOs
9. **Item #5 typed quant_config carrier** (L+L) — final form
10. **Items #6/#7/#8 + D-12-C/D-13-A/D-13-B + #11** — cleanup polish (#9, #10 resolved 2026-05-04)
If you have a specific user goal (e.g. "ship `BE_FLAVOR=fastvideo`
flavor end-to-end"), that goal dictates the order — read this list as a
menu, not a prescription.
---
## Verification gates per item
When implementing any item above, evidence required:
| Phase | Check |
|---|---|
| Build | `lsp_diagnostics` clean on changed files |
| Test | new + relevant existing tests pass; output captured |
| Manual QA | actually run the affected feature end-to-end (per AGENTS.md MANUAL_QA_MANDATE) |
| Regression | full `fastvideo/tests/api/` + `contract/` + relevant SSIM (if NVFP4 touch) |
For NVFP4 touches: re-run `test_nvfp4_ltx2_wiring.py` +
`test_typed_quant_flow.py` (CPU) + ideally a flashinfer-enabled path
test (manual, not in CI).
For Dreamverse-side items (DR-1, DR-2): re-run
`Dreamverse/apps/web/npx playwright test e2e/preset-prompt-generation.spec.ts`
end-to-end against the live BE+FE — this is the contract test that
exercises the prompt enhancer through a real session.
@@ -0,0 +1,141 @@
# PR Roadmap
Status of all 17 PRs in the FastVideo public API refactor + streaming
server upstream + Dynamo backend contract + post-deprecation cleanup.
For design rationale see [design.md](design.md). For streaming-specific
PRs (7.5-7.10) see [streaming-server.md](streaming-server.md). For NVFP4
work that runs parallel to this sequence see [quantization.md](quantization.md).
**Last updated:** 2026-05-05 (strategy reversal — single mega-PR #1288 replaces planned splits 7.10/8/LTX-2/NVFP4/post-fixes/agents_cleanup; see [decisions-log.md D-17](decisions-log.md#d-17)).
## Status legend
- ✅ **Landed on `origin/main`**
- 🟢 **Open / in flight** — branch exists, may have open PR
- 🟡 **Planned** — designed, not started
- 🔵 **Future** — deferred to post-PR-13 cleanup
## Landed PRs (0 → 7.7)
| # | PR | Status | Merge commit | Scope |
|---|---|---|---|---|
| 0 | #1218 [1/n] | ✅ | merged | Parity inventory + typed inference schema |
| 1 | #1218 [1/n] | ✅ | merged | Strict parser/validation/overrides + API tests |
| 2 | #1220 [2/n] | ✅ | merged | Typed `VideoGenerator` constructors + request path + compat |
| 3 | #1226 [3/n] | ✅ | merged | CLI/YAML-first typed config loading for `generate` and `serve` |
| 4 | #1234 [4/n] | ✅ | merged | Preset registry + presets for all 13 model families; `SamplingParam` moved to `fastvideo/api/`; `configs/sample/` deleted entirely |
| 5 | #1237 [5/n] | ✅ | merged | `ServeConfig.default_request` wired into stateless OpenAI server |
| 5.5 | (`5d1d71fc`) | ✅ | merged | Streaming server package skeleton, typed `StreamingConfig`/`GpuPoolConfig`/`PromptEnhancerConfig`/`PromptSafetyConfig`/`WarmupConfig`, `streaming-serve` CLI stub |
| 6 | #1239 [6/n] | ✅ | merged | LTX2 public preset + asset wiring + `gpu_pool.py` typed-kwarg translation |
| 7 | #1250 [7/n] | ✅ | merged | Typed LTX2 continuation state + streaming session store + blob store |
| **7.5** | **#1251** | ✅ | `95fd29e0` (merged 2026-04-26) | Streaming server skeleton (WebSocket + fMP4 + single generator). 8 commits. Deferred TODOs (per-step progress, mid-segment cancellation) carried forward to PR 7.10. |
| **7.6** | **#1257** | ✅ | `eb0a4152` (merged 2026-05-04) | GPU pool upstream + worker subprocess + two-segment warmup. 7 commits squashed. APPROVED by Eigensystem. See [decisions-log.md D-12](decisions-log.md#d-12) for the architectural review. |
| **7.7** | **#1258** | ✅ | `f673423b` (merged 2026-05-04) | Prompt enhancer with `LLMProvider` abstraction. Built-in providers: cerebras, groq. 3 commits squashed. **Public Literal does NOT include `cerebras_ifm`** — open-threads.md item DR-2 covers the gap. See [decisions-log.md D-13](decisions-log.md#d-13) for the architectural review. |
| **7.8** | **#1284** | ✅ | `eb3a3942` (merged 2026-05-04) | Streaming auxiliaries — `prompt/safety.py` (optional fasttext, lazy import), `prompt/rewrite.py`, `session_logger.py` (thread-safe JSONL), `mock_server.py` (build_mock_app + MockGenerator for FE dev). 730 LOC, 2 commits. See [decisions-log.md D-14](decisions-log.md#d-14). |
| **7.9** | **#1286** | ✅ | `2aaeee2a` (merged 2026-05-05) | Streaming router (multi-replica load balancer + WS proxy + `fastvideo router-serve` CLI). Squashed `cd76cf51 + 1ac1e732 + b0b7f59c + a152cb77` (router-polish second-pass; cherry-pick of `40e265b8` from `will/ltx2_sr_port`). See [decisions-log.md D-15](decisions-log.md#d-15) (structural review) + [D-16](decisions-log.md#d-16) (second-pass polish). |
## In flight (mega-PR #1288)
| # | PR | Status | Branch | Scope |
|---|---|---|---|---|
| **mega** | **#1288** | 🟢 OPEN, MERGEABLE | `will/ltx2_sr_port` (head `b36bdbc9`) | **Single consolidated landing of the full `will/ltx2_sr_port` chain.** Was originally planned as 6 stacked PRs (slices 1-3 / 4-6 / 7-15 / 16-21 / 22-23 / 24-34). Now landing as one PR — see [decisions-log.md D-17](decisions-log.md#d-17) for the strategy decision. **Contents** (commit-ordered): (1) streaming `generate_async` + `VideoEvent` + Dynamo backend contract (3 commits, was PR 7.10/#1287 closed); (2) server contract docs + Dreamverse/Dynamo shape tests (3 commits, was PR 8); (3) LTX-2 SR runtime port + i2v conditioning + alignment harness (9 commits); (4) NVFP4 wire-up + per-component compile + typed `transformer_quant` flow (6 commits); (5) LTX-2 post-handoff parity fixes — Gemma `to()`, list-of-generators (2 commits); (6) `.agents/memory/dreamverse-integration/` knowledge base + agents Phase 1 cleanup (11 commits). 34 commits total, 71 files, +13,074/-583 LOC. |
## Closed PRs in this scope
| # | PR | Status | Why closed |
|---|---|---|---|
| **7.10** | **#1287** | ❌ CLOSED 2026-05-05 | Superseded by mega-PR #1288 — strategy reversal to land everything in one go. Same 3 commits now form the head of #1288. |
## Deprecated split bookmarks (D-17)
`will/api_7.10` / `will/api_8` / `will/ltx2_sr_runtime` / `will/ltx2_nvfp4` / `will/ltx2_post_fixes` / `will/agents_cleanup` were the split-PR bookmarks under the abandoned 6-PR plan. They remain locally as historical references but are no longer maintained. STACK.md (top-level) is similarly deprecated.
## Planned (post-#1288 merge)
| # | Status | Branch | Scope |
|---|---|---|---|
| 9 | 🟡 | — | LongCat preset migration + colocation (9 model-specific stage files) |
| 10 | 🟡 | — | Hunyuan15 SR preset migration + colocation + SR field migration POC |
| 11 | 🟡 | — | SSIM/performance test migration off legacy `generate_video(..., **kwargs)` |
| 12 | 🟡 | — | Docs + examples migration (includes streaming server + Dynamo) |
| 13 | 🟡 | — | Deprecation cleanup (includes flat LTX2 kwargs the internal `gpu_pool.py` used to consume) |
## Future (compat.py death sequence)
After PR 13 lands deprecation warnings, `fastvideo/api/compat.py` (~370
lines) is the last translation shim between typed public API and legacy
internals (`FastVideoArgs`, `SamplingParam`).
| # | Status | Scope | Lines removed |
|---|---|---|---|
| 14 | 🔵 reachable | Strip forward translation: `legacy_from_pretrained_to_config`, `legacy_generate_call_to_request`, `_sampling_param_to_request_raw`, `_LEGACY_REQUEST_ALIASES`, `_LTX2_REFINE_FLAT_KEYS`. Depends on PRs 11/12/7.6 callers being migrated. | ~100 |
| 15 | 🔵 | `FastVideoArgs` becomes a `@dataclass` view over `GeneratorConfig` with `@property` accessors backing legacy field names. ~600-line god-object refactor. Depends on PR 14. | reverse-translation half (~150) trivial |
| 16 | 🔵 | `ForwardBatch` reads `GenerationRequest` by reference; kills `request_to_sampling_param` and the `ForwardBatch(**shallow_asdict(sampling_param), …)` spread. `SamplingParam` demoted or deleted. Depends on PR 15. | rest |
| 17 | 🔵 | Move `normalize_generator_config`, `normalize_generation_request`, `load_generator_config_from_file` to `parser.py`. Delete `compat.py`. | file gone |
PRs 15-17 touch training, distributed, and worker code in addition to
inference path; realistically 1-2 quarters beyond the current plan.
## Dependency chain
```
PR 13 (deprecation)
↓
PRs 11, 12, 7.6 (migrate callers)
↓
PR 14 (forward translation gone) ─── ~100 lines out of compat.py
↓
PR 15 (FastVideoArgs as view) ─── reverse-translation trivial
↓
PR 16 (ForwardBatch reads request) ─── SamplingParam demoted
↓
PR 17 (move normalizers, delete file)
```
## NVFP4 work (out-of-band, parallel to PR 7.5+)
NOT in the canonical PR sequence. Lives on `will/ltx2_sr_port`
(currently @ `156103b9`) — a separate stack alongside the public-API
upstreaming. See [quantization.md](quantization.md) for what each commit
locks in.
| Commit range | Topic |
|---|---|
| `cfccd292..b6ac7630` | LTX-2 i2v + SR runtime port + alignment harness |
| `a4760bae..c6c14c55` | NVFP4 LTX-2 wire-up + per-component compile + parity fixes (May 2 handoff) |
| `a5fcd19c..156103b9` | Post-handoff parity/perf fixes |
## Key landed artifacts (reference points)
- Parity inventory: [`docs/design/inference_schema_parity_inventory.yaml`](file:///home/william5lin/FastVideo/docs/design/inference_schema_parity_inventory.yaml) + guard [`fastvideo/tests/api/test_schema_parity_inventory.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_schema_parity_inventory.py)
- Typed schema: [`fastvideo/api/schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py)
- Compat layer: [`fastvideo/api/compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py)
- Preset system: [`fastvideo/api/presets.py`](file:///home/william5lin/FastVideo/fastvideo/api/presets.py) + per-family `pipelines/basic/<family>/presets.py`
- Streaming package skeleton (PR 5.5): [`fastvideo/entrypoints/streaming/`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/)
- LTX2 typed continuation state (PR 7): [`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py)
## Known notable decisions carried forward
- **Public inference boundary stays plain dataclasses + plain dict/YAML/JSON**
— not OmegaConf, not runtime config wrappers.
- **Every public entrypoint normalizes into typed config objects** before
touching legacy `FastVideoArgs` or `SamplingParam`.
- **Legacy `generate_video(..., **kwargs)` stays on direct legacy execution
path until PR 11**'s SSIM/performance migration. Prevents golden
baselines from drifting during compat period.
- **Typed requests use schema defaults**; legacy `generate_video(...)`
continues to inherit model-specific `SamplingParam` defaults during
compat period.
- **Preset registry uses explicit `_register_presets()` pattern** matching
`_register_configs()`; lookup keyed by `model_family`.
- **Stateless OpenAI server clones `ServeConfig.default_request`** and
merges user overrides; preset validation runs before legacy generation.
- **Streaming server added as sibling `fastvideo/entrypoints/streaming/`**
rather than extending `fastvideo/entrypoints/openai/` (PR 5.5).
## Per-PR commit-level detail
For per-PR commit lists, test plans, and merge criteria, the archived
source [`source-archive/PR-plan.md`](source-archive/PR-plan.md) (1145 lines)
remains the deepest reference. This file is the navigable summary.
@@ -0,0 +1,229 @@
# Quantization — NVFP4, LinearBase Fallback, Layer Profiles
What landed in the May 2 NVFP4 stack, why it's load-bearing, and what's
still owed (`layer_profile`, typed quant carrier, AbsMaxFP8 cleanup).
For overall API design see [design.md](design.md). For the open
follow-ups see [open-threads.md](open-threads.md).
**Last updated:** 2026-05-03.
## NVFP4 — what it is
NVIDIA's specific block-scaled FP4 format:
- e2m1 mantissa
- fp32 alpha
- `layout_128x4` scale layout
- group size 16
Distinct from MX-FP4 / OCP-FP4 / generic e3m0. The May 2 rename
(`94c983a2`) disambiguated the naming throughout FastVideo's public
surface.
## Files (current)
| File | Role |
|---|---|
| [`fastvideo/layers/quantization/nvfp4_config.py`](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/nvfp4_config.py) | `NVFP4Config`, `NVFP4QuantizeMethod`, `convert_model_to_nvfp4` |
| [`fastvideo/layers/quantization/__init__.py`](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/__init__.py) | `QuantizationMethods` literal includes `"NVFP4"`; `get_quantization_config` resolves it |
| [`fastvideo/layers/linear.py`](file:///home/william5lin/FastVideo/fastvideo/layers/linear.py) | `LinearBase.__init__` falls back to `UnquantizedLinearMethod` when `quant_config.get_quant_method` returns None — **load-bearing** |
| [`fastvideo/models/loader/fsdp_load.py`](file:///home/william5lin/FastVideo/fastvideo/models/loader/fsdp_load.py) | `_maybe_convert_model_to_nvfp4` helper detects via `isinstance(quant_method, NVFP4QuantizeMethod)`; calls `convert_model_to_nvfp4` to materialize buffers |
| [`fastvideo/models/dits/ltx2.py`](file:///home/william5lin/FastVideo/fastvideo/models/dits/ltx2.py) | `nn.Linear` → `ReplicatedLinear` for FP4-eligible subset; `_supports_prequantized_input` + `_linear_project_with_optional_prequant` helpers; quant_config + prefix= plumbing |
| [`fastvideo/api/compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py) | Typed `engine.quantization.transformer_quant: "NVFP4"` resolves to `NVFP4Config()` instance |
| [`fastvideo/fastvideo_args.py`](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py) | `__post_init__._apply_transformer_quant` pins `pipeline_config.dit_config.quant_config = NVFP4Config()` |
## Buffer naming (post-rename)
| Old | New |
|---|---|
| `_fp4_weight` / `_fp4_alpha` | `_nvfp4_weight` / `_nvfp4_alpha` |
| `_weight_global_sf` | unchanged |
| `convert_model_to_fp4` | `convert_model_to_nvfp4` |
| `FP4QuantizeMethod` | `NVFP4QuantizeMethod` |
| `QuantizationMethods` literal `"FP4"` | `"NVFP4"` |
Internal-scope torch op namespace `fastvideo_fp4::*` and
`_get_ltx2_fp4_stage_profile` deliberately left as-is — purely internal
naming that mirrors FastVideo-internal.
## Layer set asymmetry — by design
`NVFP4Config.fp4_layers` (default `layer_profile="refine"`) covers:
- `attn1.{to_q,to_k,to_v,to_out}` — full self-attention
- `attn2.{to_q,to_out}` — cross-attn Q + out only (text context not quantized)
- `audio_to_video_attn.{to_q,to_out}` — AV cross Q + out
- `video_to_audio_attn.{to_k,to_v}` — VA cross K + V
- `ffn.{fc_in,fc_out}` — video FFN
- `adaln_single.linear` — but this is `nn.Linear` (not `LinearBase`),
so it never actually gets FP4'd. List entry has no effect; matches
internal.
**NOT in the set:**
- audio self-attention (`audio_attn1.*`)
- audio cross-attention (`audio_attn2.*`)
- audio FFN (`audio.ffn.*`)
Audio path is cheap enough that quant overhead isn't worth it. Test
[`test_basic_av_block_propagates_quant_config_to_all_children`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py)
locks this in — if you add audio quantization later, update the test.
## `LinearBase` fallback — DO NOT REMOVE
[`fastvideo/layers/linear.py:191-202`](file:///home/william5lin/FastVideo/fastvideo/layers/linear.py#L191-L202): when `quant_config.get_quant_method` returns
`None` (layer not in the quant config's set), we fall back to
`UnquantizedLinearMethod`.
**Removing this fallback would break every non-tagged
`ReplicatedLinear` constructed with an `NVFP4Config`** — the previous
`assert quant_method is not None` would crash on unmatched layers (e.g.
text-encoder K/V projections, audio attention, etc.).
This is one of the load-bearing changes from `42b30bf9`.
## `transformer_quant` precedence rules
`FastVideoArgs._apply_transformer_quant` only writes
`dit_config.quant_config` when it's currently `None`. **If a caller has
explicitly set** `pipeline_config.dit_config.quant_config = NVFP4Config(...)`,
the explicit setter wins.
Dreamverse's `video_generation.py` relies on this precedence — it sets
`NVFP4Config()` directly via `experimental["pipeline_config"]` because
typed `transformer_quant: "NVFP4"` doesn't yet expose `layer_profile`.
See "Open follow-ups" below.
## Attention forward optimization
[`models/dits/ltx2.py`](file:///home/william5lin/FastVideo/fastvideo/models/dits/ltx2.py)
ports `_supports_prequantized_input` and
`_linear_project_with_optional_prequant`. Attention forward
pre-quantizes input once (`quantize_input`), reuses the
`(x_fp4, x_scale, x_global_sf)` tuple for k/v projections when
`context is x` — bit-matches internal's fused path.
## `prepare_for_compile` protocol
[`composed_pipeline_base._maybe_compile_pipeline_module`](file:///home/william5lin/FastVideo/fastvideo/pipelines/composed_pipeline_base.py)
calls `getattr(module, "prepare_for_compile", None)` before invoking
`torch.compile`. Defined as a duck-type protocol — no base class method.
Currently only **Gemma3** implements it (to materialize HF weights
outside Dynamo's tracer). Add to other models that have lazy external
state if you observe compile-time graph breaks.
## Per-component compile flags
`CompileConfig` (in
[`fastvideo/api/schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py))
gained per-component knobs in `221cb20a`:
```python
@dataclass
class CompileConfig:
enabled: bool = False # master DiT switch
backend: str = "inductor"
fullgraph: bool = False
mode: str | None = None
dynamic: bool | None = None
extras: dict = field(default_factory=dict)
# Per-component overlays, None = inherit master `enabled`
text_encoder_enabled: bool | None = None
vae_enabled: bool | None = None
audio_vae_enabled: bool | None = None
# Per-component kwargs override master when non-empty
dit_kwargs: dict = field(default_factory=dict)
text_encoder_kwargs: dict = field(default_factory=dict)
vae_kwargs: dict = field(default_factory=dict)
audio_vae_kwargs: dict = field(default_factory=dict)
```
**`transformer_refine` is auto-compiled with the master DiT flag.** No
separate `enable_torch_compile_refine` flag — by design, refine inherits
DiT compile state to keep typed surface small. Decoupling would add a
new flag, not repurpose existing ones.
## Quantization commit chain (`will/ltx2_sr_port`)
| Commit | Locks in |
|---|---|
| `365a66c7 feat(quantization): upstream LTX-2 FP4Config with lazy flashinfer` | Public colocation of FP4Config (resolves dreamverse_review Q-6 option 1); flashinfer lazy-imported in loader helper, no public hard-dep |
| `a4760bae fix(api): propagate generic refine_*` | `_resolve_refine_args()` copies generic `refine_*` knobs onto `ltx2_refine_*` runtime carriers; `_randn_ltx2_video_latents` reverts to `torch.randn` to bit-match internal under single-generator inference |
| `221cb20a feat(api): typed per-component CompileConfig` | `CompileConfig` per-component knobs; matching `FastVideoArgs` carriers; compat layer round-trip |
| `6da342ba feat(compile): per-component compile + transformer_refine + prepare hook` | `composed_pipeline_base.post_init` compiles `transformer_refine` alongside `transformer`/`transformer_2`; per-component compile loops; `prepare_for_compile` hook on Gemma3 |
| `42b30bf9 feat(ltx2): wire FP4 inference` (largest) | `nn.Linear` → `ReplicatedLinear` for FP4-eligible LTX2 subset; `quant_config` + `prefix=` plumbing; `_maybe_convert_model_to_nvfp4` helper; `LinearBase` fallback to `UnquantizedLinearMethod`; typed `transformer_quant` resolution |
| `94c983a2 refactor(quant): rename FP4 → NVFP4` | Mechanical rename across config, methods, buffers, tests |
| `c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow` | 6+4 tests in `test_nvfp4_ltx2_wiring.py` + `test_typed_quant_flow.py` |
| `a5fcd19c [fix]: lazy-import flash_attn 2 fallback in attention backend` | post-handoff: lazy import to avoid hard flash_attn 2 dep |
| `d4ee5be2 [fix]: avoid model.to() round-trip in Gemma encoder forward` | post-handoff: parity / perf fix |
| `156103b9 [fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep` | post-handoff: parity fix for list-of-generators (was bit-matching only single-generator path) |
## Tests
| Test | Asserts |
|---|---|
| [`fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py) (6 tests) | `LTXSelfAttention.to_q/to_k/to_v/to_out` are `ReplicatedLinear`; `NVFP4Config()` attaches `NVFP4QuantizeMethod` on the quantized subset with correct `layer_prefix`; non-tagged projections (cross-attn K/V, audio attn, audio FFN) fall back to `UnquantizedLinearMethod`; `BasicAVTransformerBlock` propagates `quant_config`+`prefix` correctly to all 4 attention modules + FFN |
| [`fastvideo/tests/api/test_typed_quant_flow.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_typed_quant_flow.py) (4 tests) | typed `engine.quantization.transformer_quant: "NVFP4"` → `NVFP4Config()` instance flow; default leaves `transformer_quant` None; explicit `dit_config.quant_config = ...` wins over typed carrier |
CPU-only by design; do NOT exercise actual FP4 kernels (no flashinfer in
CI). Real kernel coverage requires a CI run with flashinfer installed.
## Open follow-ups (quantization-specific)
### #4: Expose `layer_profile` on typed `engine.quantization`
Today `transformer_quant: "NVFP4"` always constructs `NVFP4Config()`
with default `layer_profile="refine"`. To support stage-1 profiles (no
`attn2.to_out`, no cross-modal AV) via typed config, add
`transformer_quant_layer_profile: str | None = None` and thread it
through:
- `fastvideo/api/schema.py` — `QuantizationConfig` field
- `fastvideo/api/compat.py` — typed → flat translation
- `fastvideo/fastvideo_args.py` — `_apply_transformer_quant` consumes it
Dreamverse currently dodges this by setting `NVFP4Config()` directly via
`experimental["pipeline_config"]`. Exposing `layer_profile` removes the
dodge. See [open-threads.md](open-threads.md) #4.
### #5: Typed `dit_config.quant_config` carrier (replace `experimental["pipeline_config"]`)
Long-term: design a typed home for an in-memory `PipelineConfig`
instance with mutated `dit_config`. Today `compat.py` recognizes the
`pipeline_config` key in `experimental` and threads it through to
`FastVideoArgs.from_kwargs`. This is fine for short-term but not pretty.
Heaviest design work in the open queue. May need Oracle consult.
### #2: AbsMaxFP8 pre-existing test failure
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on `will/ltx2_sr_port` with the same error
(`AssertionError not raised`). Verified via `git stash` that the
failure pre-dates NVFP4 work.
Either:
- Fix the test (`AbsMaxFP8LinearMethod.create_weights` no longer
asserts on invalid dtype — restore the assert if intentional, or drop
the test).
Self-contained tech debt; small fix.
## Don't / Cautions
- **Don't change `NVFP4Config` buffer names back to `_fp4_*`.** Rename
is intentional to disambiguate from MX-FP4 / OCP-FP4.
- **Don't remove the `LinearBase` `UnquantizedLinearMethod` fallback.**
Load-bearing for non-tagged layers when a `quant_config` is set.
- **Don't repurpose `enable_torch_compile` to mean DiT-only.** It also
drives `transformer_refine` and `transformer_2` compile.
- **Don't bypass the typed surface for new options.** New compile /
quant / refine knobs should land on the dataclass + compat.py +
parity inventory together. The existing test suite locks this in.
- **Don't merge to main without a CI run that covers FP4.** Current CI
doesn't run flashinfer-dependent paths; the wiring tests are CPU-only
by design.
@@ -0,0 +1,385 @@
# Runbook — How to Do Work in This Scope
Operational how-to for the dreamverse-integration scope. Read after
[state.md](state.md) and [open-threads.md](open-threads.md).
For design rationale see [design.md](design.md). For who to credit see
[authors.md](authors.md). For PR status see [pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-05 (strategy reversed to single mega-PR #1288 on `will/ltx2_sr_port`; #1287 closed; STACK.md split model deprecated per [decisions-log.md D-17](decisions-log.md#d-17)).
## Worktree contract
```
Repo: /home/william5lin/FastVideo
Branch: will/ltx2_sr_port
```
Other agents and the user share this worktree concurrently. If `git status`
shows changes you don't recognize, they belong to **someone else's work** —
don't revert, don't `git stash drop`, don't `git checkout -- <file>`.
Switch to `will/ltx2_sr_port` cleanly with `git checkout will/ltx2_sr_port`
(safe if your own working tree is clean) and proceed.
If your task requires a different branch (e.g. cherry-pick to
`will/api_7.9` for PR #1286 propagation), return to `will/ltx2_sr_port`
when done — that is the assumed default.
## Branch topology (single mega-PR model)
The dreamverse-integration work now ships as one PR (#1288) off
`will/ltx2_sr_port`. The split-PR model documented in earlier revisions
of this runbook (and in top-level `STACK.md`) is **abandoned** —
see [decisions-log.md D-17](decisions-log.md#d-17).
```
origin/main
↓ [public-API refactor: PRs 0..7.9 merged on main, latest #1286 = 2aaeee2a]
will/ltx2_sr_port (**PR #1288 head** — single mega-PR, 34 commits, 71 files, +13,074/-583)
```
| Branch | Role | Status |
|---|---|---|
| `will/ltx2_sr_port` | **PR #1288 head**, default working branch | OPEN, MERGEABLE |
| `will/api_7.10` / `will/api_8` / `will/ltx2_sr_runtime` / `will/ltx2_nvfp4` / `will/ltx2_post_fixes` / `will/agents_cleanup` | deprecated split-PR bookmarks | local-only historical references; safe to delete |
| `will/ltx2_sr_port-pre-1286-rebase` | safety backup | local-only; preserves the 4 commits dropped during the post-#1286 rebase |
**Sanity check:** `git merge-base --is-ancestor origin/main will/ltx2_sr_port`
should exit 0. If it doesn't, the branch is in an unexpected state — read
[state.md](state.md) before continuing.
## After PR #1288 merges
When the mega-PR squash-merges into `main`:
1. `git fetch origin main` to pull the merge commit.
2. The entire `will/ltx2_sr_port` content is now on main; the branch can
be deleted (locally + on origin) once all consumers are notified.
3. Delete deprecated split bookmarks: `git branch -D will/api_7.10
will/api_8 will/ltx2_sr_runtime will/ltx2_nvfp4 will/ltx2_post_fixes
will/agents_cleanup` (local-only, no remote).
4. Optionally remove top-level `STACK.md` (now a historical artifact).
Keep `CO-AUTHORS.md` — still the canonical roster reference.
5. Decide whether to keep `will/ltx2_sr_port-pre-1286-rebase` (safety
backup of the pre-rebase chain) — recommend deleting once #1288 is
merged and verified on main.
6. Update memory dir to reflect the post-merge state — bump
`Last reconciled` headers, mark Item D resolved in
[open-threads.md](open-threads.md), record the merge commit in
[decisions-log.md](decisions-log.md).
## Historical: split-PR re-slice protocol (deprecated)
Prior revisions of this runbook documented a 10-step re-slice protocol
for the abandoned 6-PR split model. That protocol is now obsolete.
The post-#1286 rebase (2026-05-05) was the last execution of it; details
are preserved in [state.md](state.md) "Post-#1286 rebase summary" and
git history at commit `b34d9704`.
## Verification
### Lint (pre-commit)
```bash
pre-commit run --files <changed-paths...>
```
- Binary: `/home/william5lin/miniconda3/envs/fv-main/bin/pre-commit`.
NOT `.venv/bin/pre-commit` — that doesn't exist in this worktree.
- Auto-applies yapf reformatting; re-stage modified files after.
- Hook chain: yapf → ruff → codespell → mypy → spaces-check.
- Memory dir (`.agents/memory/`) is yapf/ruff/mypy excluded — only
"spaces" runs. Memory edits don't need lint, but DO use UTF-8 and
consistent line endings.
### Tests
Router tests (PR #1286 scope):
```bash
.venv/bin/python -m pytest fastvideo/tests/entrypoints/streaming/test_router.py -v --no-header
```
Stack baseline (May 2 handoff suite — re-run when you change anything in
api/, contract/, or LTX-2 paths):
```bash
.venv/bin/python -m pytest \
fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header
```
Expected baselines:
- May 2 handoff (`156103b9`): 222 passed, 1 skipped.
- Post-D-16 (`a152cb77` / `09647a30`): +7 router tests pass on top.
### LSP
Use `lsp_diagnostics` on changed files BEFORE running build. Pre-existing
warnings to ignore (predate this work):
- `fastvideo/entrypoints/streaming/router/main.py:37` — `Task` generic.
- `fastvideo/entrypoints/cli/router_serve.py:55` — `_SubParsersAction` generic.
### gh CLI for PR status
```bash
# PR #1286 quick status
gh pr view 1286 --json headRefOid,mergeable,statusCheckRollup \
--jq '{headRefOid, mergeable, checks: [.statusCheckRollup[] | {name, status, conclusion}]}'
# All commits in a PR + co-author check
gh pr view 1286 --json commits \
--jq '.commits[] | {oid: .oid[0:8], msg: .messageHeadline, author: .authors[0].login}'
```
## Commit workflow
### Subject convention
`[type] <scope>: <imperative summary>` — keep ≤ 72 chars.
Types observed in this scope: `feat`, `fix`, `test`, `docs`, `chore`,
`refactor`. Scopes observed: `streaming`, `dreamverse-integration`,
`api`, `quant`, `ltx2`, `nvfp4`, etc.
Examples:
- `[fix] streaming: router polish — bridge cancel + state machine + deps`
- `[docs] dreamverse-integration: add authors.md + track D-16 router polish`
### Body convention
Bullet list, one bullet per file or concern. Why-before-what. Wrap at
~80 chars (yapf doesn't reformat commit messages; readability is on you).
### Co-author trailers (REQUIRED on every commit)
The 4 trailers in [authors.md](authors.md) MUST appear on every commit
in this scope. Use `--trailer` flags or write the body to a file with
`-F` — DO NOT use multiple `-m` blocks for the trailers (each `-m` is
its own paragraph and git's trailer parser only reads the LAST paragraph,
yielding 1 trailer parsed instead of 4).
**Inline `--trailer` form (preferred for short commits):**
```bash
git commit -m "subject" -m "body..." \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
```
**File form (preferred for multi-paragraph bodies):**
```bash
cat > /tmp/opencode/msg.txt <<'EOF'
[type] scope: subject
* Bullet one with rationale.
* Bullet two with rationale.
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
EOF
git commit -F /tmp/opencode/msg.txt
```
The trailers MUST be a single block at the end of the message with no
blank lines between them.
**Verify trailers parsed:**
```bash
git log -1 --format='%(trailers:key=Co-authored-by,valueonly)'
```
Should print 4 lines (one per author). If only 1 line, you have the
multi-`-m` bug — amend with `-F` to fix (allowed if commit is unpushed
and you authored it in this session per AGENTS.md amend rules).
### NEVER add to commits
Per [`AGENTS.md`](../../../AGENTS.md):
- AI co-authors (Claude, GPT, Codex, Cursor, etc.) — explicitly forbidden
- "Generated with Claude Code" footer — explicitly forbidden
- `--no-verify` to skip pre-commit — explicitly forbidden
## Push + PR propagation
### Pushing `will/ltx2_sr_port` (top of stack)
```bash
git push origin will/ltx2_sr_port # fast-forward, no force needed
```
If git wants to force-push, you've rewritten history. STOP and verify:
```bash
git log origin/will/ltx2_sr_port..will/ltx2_sr_port # local-only commits
git log will/ltx2_sr_port..origin/will/ltx2_sr_port # remote-only commits
```
Force-push requires explicit user confirmation per `AGENTS.md`.
### Propagating fixes to PR #1286 (`will/api_7.9`)
When a fix is in router code (`fastvideo/entrypoints/streaming/router/`,
`cli/router_serve.py`, `tests/entrypoints/streaming/test_router.py`,
or `pyproject.toml` router-related), it must land on BOTH branches.
Cherry-pick avoids any force-push:
```bash
# 1. Commit on will/ltx2_sr_port first (working branch)
git add <files...>
git commit -F /tmp/opencode/msg.txt # with trailers per above
# 2. Cherry-pick onto will/api_7.9 (creates a separate SHA, identical diff)
git checkout will/api_7.9
git cherry-pick <ltx2_sr_port-sha>
git push origin will/api_7.9 # fast-forward, no force
# 3. Return to working branch
git checkout will/ltx2_sr_port
# 4. Verify PR #1286 picked it up
gh pr view 1286 --json headRefOid --jq '.headRefOid'
```
Two SHAs for the same diff — they'll dedupe naturally on the next
bulk-rebase via the trailer-injection rebase command in
[authors.md](authors.md).
### When a fix is memory-dir-only
`.agents/memory/dreamverse-integration/` lives in the `agents_cleanup`
layer of the stack — it does NOT belong on `will/api_7.9`. Memory updates
stay on `will/ltx2_sr_port` only.
### When a fix is non-router code in the integration scope
Land on `will/ltx2_sr_port`. If that fix needs to ship as a separate PR
(e.g. extending PR 7.10 or starting PR 9), open a new branch off the
right base per [pr-roadmap.md](pr-roadmap.md).
## Memory dir maintenance
When state changes, update the memory dir BEFORE moving on. Every file
has a "Last updated" header — bump when you edit.
| Change | File to update |
|---|---|
| Branch tip moves | [state.md](state.md) "Branch tips" + "Last reconciled" |
| PR opens / merges | [pr-roadmap.md](pr-roadmap.md) status table |
| New decision made | [decisions-log.md](decisions-log.md) — add D-N entry, bump header |
| Open thread resolved | [open-threads.md](open-threads.md) — strikethrough + "Resolved" note |
| New open thread | [open-threads.md](open-threads.md) — priority overview + section |
| New collaborator credited | [authors.md](authors.md) roster + trailer block + bulk-rebase |
| Source doc archived | [source-archive/README.md](source-archive/README.md) + [README.md](README.md) sources table |
| Process / runbook detail changes | [runbook.md](runbook.md) (this file) |
Cross-link siblings via relative paths. Never duplicate content — link.
## Common pitfalls
### `pre-commit` not in `.venv/bin`
`pre-commit` lives at `/home/william5lin/miniconda3/envs/fv-main/bin/pre-commit`.
The `.venv` here is for the FastVideo package itself, not pre-commit.
### Trailers split across paragraphs
`git commit -m A -m B -m C` makes A, B, C separate paragraphs. Git's
trailer parser only reads the LAST paragraph — multiple `-m
"Co-authored-by: ..."` produces 1 trailer parsed, not 4. Use `--trailer`
flags or `-F` with the trailers in a single block at the end.
### Stash 0 on FastVideo IS NOT yours
`stash@{0}: WIP on main: 71bfc13d HunyuanVideo plugin` predates this work.
**DO NOT POP.** See [state.md](state.md) "Stashes — DO NOT POP".
### `AbsMaxFP8` test "failure" is pre-existing
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on every branch in this scope. NOT introduced by
integration work. See [open-threads.md](open-threads.md) item #2.
### Untracked nested clones at repo root
`dynamo/`, `ray/`, `vllm-omni/` are untracked nested git clones at the
FastVideo repo root. Reference repos for cross-repo work. **Do not
`rm -rf`** — they're someone else's working state.
### Live services on 8009 / 5274
`dreamverse-server` runs on 8009 (warmed GPU worker), Next.js dev server
on 5274. Don't start new instances on those ports without checking
[state.md](state.md) "Live services" first.
### Branch may have been switched by another agent
Other agents share this worktree. If `git branch --show-current` returns
something other than `will/ltx2_sr_port`, switch back cleanly with
`git checkout will/ltx2_sr_port` — don't disturb their work, don't
discard their uncommitted changes.
### Force-push policy
Per `AGENTS.md`: never force-push without explicit user confirmation.
For trailer fixes on already-pushed commits, prefer the bulk-rebase
command in [authors.md](authors.md) — safe to re-run.
### Two trailerless commits in PR #1286
`a152cb77` (on `will/api_7.9`) and `40e265b8` (now-superseded ancestor
on `will/ltx2_sr_port`) lack the 4 co-author trailers. **Accepted gap**
per user decision — see [authors.md](authors.md) "Known gaps".
## Self-test (verify your context is loaded)
After reading the memory dir, you should be able to answer:
1. What branch should I be on? → `will/ltx2_sr_port`
2. What's the active open PR in this scope? → #1286 on `will/api_7.9`
3. Where does PR #1286 land in the stack? → Bottom; ancestor of `will/ltx2_sr_port`
4. Who do I credit on every commit? → 4 authors per [authors.md](authors.md)
5. Where do memory updates land? → `will/ltx2_sr_port` only (NOT api_7.9)
6. What's the next-priority open thread? → See [open-threads.md](open-threads.md) "Recommended pull order" — D-8 verify is current top
7. What pre-existing failure can I ignore? → AbsMaxFP8 test (item #2)
8. What's the bulk-rebase command for adding trailers across the stack? → See [authors.md](authors.md) "How the trailers were applied"
If you can't answer one of these from the memory dir alone, the dir has
a gap — file it as a new entry in [open-threads.md](open-threads.md)
before continuing.
## First 60 seconds — copy-paste orientation
```bash
# 1. Confirm branch
cd /home/william5lin/FastVideo
git branch --show-current # should print: will/ltx2_sr_port
# If not, recover: git checkout will/ltx2_sr_port
# 2. Confirm worktree clean (untracked nested clones expected)
git status --short
# 3. Confirm PR #1286 head matches expected api_7.9 tip
gh pr view 1286 --json headRefOid --jq '.headRefOid'
git rev-parse will/api_7.9 # should match PR head
# 4. Confirm your context vs the memory dir
git log -1 --oneline
cat .agents/memory/dreamverse-integration/state.md | head -30
# 5. Confirm live services still running
curl -s http://localhost:8009/readyz | head -c 200
curl -s http://localhost:5274/ -o /dev/null -w "%{http_code}\n"
```
If any of those produce unexpected output, read [state.md](state.md)
before changing anything.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,44 @@
# Source Archive
These are the original unsynthesized design and integration docs that
predate the consolidation in
[`../`](../). They are **NOT** the source of truth — the synthesized
sibling files in the parent directory are.
Archived 2026-05-03. All previously untracked.
## Contents
| File | Original location | Date | Synthesized into |
|---|---|---|---|
| `apirefactor.md` | `FastVideo/` (repo root) | 2026-04-21 | [`../design.md`](../design.md) |
| `PR-plan.md` (was `PR plan.md` at repo root) | `FastVideo/` (repo root) | 2026-04-25 | [`../pr-roadmap.md`](../pr-roadmap.md) |
| `dreamverse_review.md` | `FastVideo/` (repo root) | 2026-04-26 | [`../decisions-log.md`](../decisions-log.md) + [`../state.md`](../state.md) |
| `handoff-nvfp4-launch-demo.md` | `.agents/exploration/` | 2026-05-02 | [`../state.md`](../state.md) + [`../quantization.md`](../quantization.md) + [`../open-threads.md`](../open-threads.md) |
| `streaming-server-upstream-plan.md` | `.agents/exploration/` | 2026-04-17 | [`../streaming-server.md`](../streaming-server.md) + [`../decisions-log.md`](../decisions-log.md) |
| `dreamverse_integration.md` | `.agents/exploration/` | 2026-04-23 | [`../cross-repo-surfaces.md`](../cross-repo-surfaces.md) |
| `video-generator-config-api-design.md` | `.agents/exploration/` | 2026-04-02 | [`../design.md`](../design.md) (early-draft material) |
## Why archived (not deleted)
- Future agents may want the **full unsynthesized rationale** for a
decision the synthesis abbreviated.
- The originals remain useful as a **time machine** for understanding
how the design evolved.
- These docs were never committed to git, so leaving them on disk costs
nothing.
## When to read the archive vs. the synthesis
- **Read the synthesis (`../*.md`)** for: current state, decision
status, action items, design rationale at the conceptual level.
- **Read the archive (here)** for: deep historical context, exact wording
of design decisions, full PR plan with all sub-PR commit details,
the original Q-1..Q-9 / D-1..D-11 prose.
## Maintenance rule
Do NOT edit files in this archive. They are point-in-time snapshots.
If new design material appears that supersedes an entry here, update the
synthesis (the parent dir) and append a note to that synthesis file —
do not mutate this archive.
@@ -0,0 +1,838 @@
# FastVideo API Refactor Design
## Related Documents
- [PR plan.md](PR%20plan.md) — PR-by-PR implementation plan for this design
- [.agents/exploration/streaming-server-upstream-plan.md](.agents/exploration/streaming-server-upstream-plan.md) — streaming-server upstream + Dynamo backend contract (shapes PRs 5.5-7.10)
- `../FastVideo-internal/.agents/exploration/rebase-upstream-fastvideo.md` — rebasing FastVideo-internal onto upstream (enables PRs 6-8)
- `../FastVideo-internal/ui/ltx2-streaming/` — source for the streaming server being upstreamed (PRs 7.5-7.9)
- `../dynamo/` — local clone of ai-dynamo/dynamo; `components/src/dynamo/sglang/` is the template for FastVideo's native backend landed in PR 7.10
- https://github.com/ai-dynamo/dynamo/pull/7544 — closed draft PR that establishes the Dynamo backend shape this design must satisfy
## Status
Design spec for the public inference API refactor. PRs 0-5.5 are landed; see [PR plan.md](PR%20plan.md) for rollout status and the PR 6+ roadmap. The typed schema, strict parser, preset system, typed VideoGenerator, typed CLI, and stateless OpenAI server default-request merge are all implemented. Streaming package skeleton + typed streaming config types are in place; live streaming server + Dynamo contract are the next milestones.
## Executive Summary
FastVideo should move to a single typed nested inference schema that is shared across:
- Python API
- CLI
- YAML/JSON config files
- OpenAI/server request translation
The core split is:
- `GeneratorConfig`: generator-instance lifetime settings
- `GenerationRequest`: per-call inputs, sampling, outputs, and continuation
- `InferencePreset`: model-owned named multi-stage defaults
The canonical user experience should be:
1. Choose a model.
2. Choose a pipeline preset.
3. Override a few typed fields.
4. Generate.
FastVideo should not make a raw free-form string dict the primary API. Dicts and YAML/JSON should be supported as serialization/interchange layers, but they must be parsed immediately into typed config objects with strict unknown-key validation.
The repo should also shift model-specific preset/default definitions closer to their pipeline implementations, while keeping the shared public schema and parsers centralized.
## Why This Refactor Is Needed
Today the public inference boundary is too flat and too forgiving.
- `VideoGenerator.from_pretrained(..., **kwargs)` mixes:
- engine/runtime settings
- pipeline init settings
- component overrides
- `VideoGenerator.generate_video(..., **kwargs)` mixes:
- prompt and inputs
- sampling parameters
- output settings
- model-specific workflow knobs
- unknown or drifting keys can be silently filtered or merely logged instead of failing fast
- model-specific multi-stage behavior is exposed through ad hoc top-level flags instead of a stable preset/stage abstraction
This is already painful in LTX2/Dreamverse, and it will get worse as more multi-stage pipelines are upstreamed.
## Design Goals
- Keep the Python API typed and editor-friendly.
- Make YAML/JSON a first-class serialization of the same schema.
- Support CLI overrides cleanly without flattening the schema into hundreds of canonical flags.
- Separate init-time config from request-time config.
- Provide a stable public abstraction for multi-stage pipelines.
- Support LTX2 two-stage and continuation behavior cleanly.
- Keep the simple case simple.
- Co-locate model-owned defaults and stage topology with the relevant pipeline.
- Protect current public/server behavior with an explicit schema parity audit before freezing the new surface.
- Preserve backward compatibility long enough to migrate examples, internal users, and servers safely.
## Non-Goals
- Do not make Ray a structural dependency or copy its package layout.
- Do not make a raw free-form dict the primary Python API.
- Do not force all models into one universal `RefineConfig`.
- Do not expose stage indices as the primary user interface.
- Do not move every shared config class into per-model directories.
## External Inspiration
### Ray
Borrow only the ergonomic idea that user-facing config can be expressed as a string-keyed dict or YAML/JSON config. Do not copy Ray's structure into FastVideo.
### SGL Multimodal Gen
Useful ideas: split instance config from request config; allow dict input at the boundary; parse dicts immediately into typed request objects; merge request overrides onto model defaults; validate request params against pipeline/task type. Do not copy: request objects depending on server/engine config; broad weakly typed request bags as the canonical API.
### vLLM-Omni
Useful ideas: model-owned pipeline presets; explicit stage topology; per-stage default sampling params; clean separation between stage topology, engine defaults, and runtime overrides. Do not copy: positional `sampling_params_list` as the primary public API; serving-engine-oriented stage index semantics in the main Python interface.
## Core Decision
FastVideo should have:
1. A shared typed public schema.
2. Model-owned named pipeline presets.
3. Semantic stage overrides by stage name.
4. Optional advanced explicit plans for power users.
5. YAML-first config loading with dotted CLI overrides.
The public API should be stable at the schema level, while model-specific behavior should be contained in preset definitions and model-specific typed override classes.
## Schema Parity Requirement
Before the new schema is declared canonical, FastVideo should build a parity inventory across all current public inference surfaces (Python `VideoGenerator` kwargs, CLI flags, YAML/JSON config inputs, OpenAI/server request models, model-specific sampling/runtime fields). Each field must be marked: kept as-is, renamed, moved to a nested path, preset-owned, private-only adapter field, or intentionally dropped. No field should disappear implicitly.
For any public field that remains supported, there should be either a normalized-config equivalence test, or an explicit parser/translation test. Fields that exist only in private Dreamverse integration code should be handled by a private adapter layer, not quietly converted into public FastVideo compatibility guarantees.
Landed artifact: [inference_schema_parity_inventory.yaml](docs/design/inference_schema_parity_inventory.yaml) + guard [test_schema_parity_inventory.py](fastvideo/tests/api/test_schema_parity_inventory.py).
## Canonical Public Schema
The typed schema is implemented in [fastvideo/api/schema.py](fastvideo/api/schema.py). Envelope types:
- `RunConfig` — offline: `generator` (GeneratorConfig) + `request` (GenerationRequest)
- `ServeConfig` — serving: `generator` + `server` (ServerConfig) + `default_request` (GenerationRequest) + optional `streaming` (StreamingConfig)
Key nested types (summary; full fields in `schema.py`):
- `GeneratorConfig` → `model_path`, `revision`, `trust_remote_code`, `engine` (EngineConfig: parallelism/offload/compile/quantization/flags), `pipeline` (PipelineSelection: workload_type, preset, preset_version, components, preset_overrides, experimental)
- `GenerationRequest` → `prompt`, `negative_prompt`, `inputs` (InputConfig), `sampling` (SamplingConfig), `runtime` (RequestRuntimeConfig), `output` (OutputConfig), `stage_overrides`, `state` (ContinuationState), `plan` (GenerationPlan), `extensions`
- `ContinuationState` → opaque `{kind: str, payload: dict[str, Any]}`
- `GenerationPlan` → `{stages: list[PlannedStage], final_stage: str | None}`; advanced/escape-hatch only
### Important Semantics
- Dataclasses are canonical for Python users.
- Dict and YAML/JSON are parsed into these dataclasses immediately.
- Unknown keys must raise validation errors.
- Typed `GenerationRequest` defaults come from the public schema, not from model-specific `SamplingParam.from_pretrained(...)` defaults.
- Legacy `generate_video(...)` continues to inherit model-specific sampling defaults until the SSIM/performance migration lands (PR 11).
- The only open-ended escape hatches are:
- `generator.pipeline.experimental`
- `request.extensions`
That keeps the public contract strict without blocking experimental work.
### Request Mutation Tracking
When a `GenerationRequest` is parsed from a raw dict (YAML, JSON, or Python mapping), FastVideo records which fields the user explicitly provided versus which received schema defaults. This matters because `request_to_sampling_param()` must distinguish user-provided values (which should override model defaults) from schema defaults (which should NOT override model defaults).
The tracking contract:
- At parse time, the original raw dict and a baseline snapshot of the parsed object are stored on the request.
- Dataclass field mutations after parsing (e.g., `request.sampling.seed = 7`) are captured via lightweight `__setattr__` dirty-path recording.
- Dict-typed field mutations (e.g., `del request.stage_overrides["refine"]`) are detected at access time by diffing the current dict against the baseline snapshot.
- Setting a field to the schema default value IS captured as explicit, so it will override model defaults.
- The raw dict is reconciled lazily when `normalize_generation_request()` is called, not on every individual mutation.
### Schema Purity and Model-Specific Fields
The shared schema currently contains fields that are specific to one or two model families. These remain for backward compatibility during the initial migration (PRs 0-3) but should migrate to preset-owned typed override classes as the preset system lands (PRs 4-10).
**SamplingConfig fields to migrate:**
- `height_sr`, `width_sr`, `num_inference_steps_sr`: Hunyuan15 SR only. Target: `HunyuanSRStageOverride` in PR 10.
- `guidance_scale_2`, `boundary_ratio`: Wan2.2 and LingBotWorld only. Target: preset-owned overrides in the relevant model migration PR.
**InputConfig fields to migrate:**
- `mouse_cond`, `keyboard_cond`, `grid_sizes`: MatrixGame action control only. Target: `request.extensions` or a typed MatrixGame input config.
- `c2ws_plucker_emb`: LingBotWorld camera control only. Target: `request.extensions` or a typed LingBotWorld input config.
- `refine_from`, `stage1_video`: LongCat refinement only. Target: `LongCatRefineStageOverride` inputs or keep in `InputConfig` if they remain a public contract.
**Universal fields that stay in the shared schema:**
- `guidance_rescale`: used by multiple denoising stages across models, default 0.0. Universally applicable.
- `true_cfg_scale`: OpenAI adapter surface. Keep for protocol compatibility.
### Escape Hatch Sunset
`generator.pipeline.experimental` and `request.extensions` are intentional escape hatches for experimental and private work. They bypass strict validation by design.
Rules for escape hatch usage:
- New fields should not be added to `experimental` or `extensions` without a plan to either promote them to typed fields or remove them within two PR cycles.
- Each model migration PR (PRs 6-10) should review and shrink escape hatch usage for that model family.
- The compatibility layer currently routes unrecognized legacy kwargs into `experimental`. This pass-through should shrink as presets absorb model-specific fields.
## Public Python API
### New Canonical API
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GeneratorConfig, GenerationRequest,
EngineConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
generator = VideoGenerator.from_pretrained(
config=GeneratorConfig(
model_path="/models/ltx2",
engine=EngineConfig(num_gpus=1),
pipeline=PipelineSelection(
workload_type="t2v",
preset="ltx2_two_stage",
),
)
)
result = generator.generate(
GenerationRequest(
prompt="a fox running through snow",
sampling=SamplingConfig(
num_frames=121, height=1024, width=1536,
num_inference_steps=8, seed=42,
),
output=OutputConfig(save_video=True, return_state=True),
)
)
```
### Accepted Construction Forms
Canonical:
```python
VideoGenerator.from_pretrained(config=GeneratorConfig(...))
VideoGenerator.from_config(GeneratorConfig(...))
VideoGenerator.from_file("run.yaml")
```
Stable convenience constructor:
```python
VideoGenerator.from_pretrained("model-id")
VideoGenerator.from_pretrained("model-id", num_gpus=2, use_fsdp_inference=False, ...)
```
Legacy compatibility:
```python
VideoGenerator.from_pretrained(model_path, **legacy_kwargs)
```
All constructor forms normalize through the same typed path. Stable convenience kwargs remain supported with no deprecation warning. Advanced model/pipeline-specific kwargs are accepted during migration but only as compatibility inputs that normalize into `GeneratorConfig`. The thing being deprecated over time is the unbounded legacy kwarg surface, not the `from_pretrained(...)` entrypoint itself.
### Generation Entry Point
Canonical: `generator.generate(request: GenerationRequest) -> GenerationResult`.
Compatibility alias: `generator.generate_video(prompt=..., **legacy_kwargs)` — converts legacy calls into a `GenerationRequest` and emits a deprecation warning.
During the compat period, `generate(request=...)` uses schema defaults while `generate_video(...)` preserves legacy model-default behavior. These paths intentionally differ until preset-owned defaults replace the remaining `SamplingParam` default logic (migrated in PR 11).
### Boundary Normalization Rule
Every public inference entrypoint normalizes into typed config objects before touching legacy internals. That includes Python constructors, generation calls, CLI `generate`, CLI `serve`, and OpenAI/server request translation. Legacy internals (`FastVideoArgs`, `SamplingParam`) may remain temporarily, but only behind a typed normalization boundary.
## Pipeline Presets
### Definition
An `InferencePreset` is a named model-owned preset that defines:
- workload selection
- stage topology
- per-stage defaults
- stage names
- allowed stage override types
- init-time feature requirements
The preset is not user-authored by default. It is supplied by the model integration.
### Why Presets Are The Right Abstraction
Users usually do not want to assemble a stage graph by hand. They want to say:
- use LongCat distill + refine
- use Hunyuan 1080p SR
- use LTX2 two-stage continuation mode
Presets provide a stable public noun for that behavior.
### Preset Naming Rules
- Use semantic names, not stage indices.
- Keep names stable across releases.
- If semantics change incompatibly, change `preset_version` or create a new preset name.
Examples: `ltx2_base`, `ltx2_two_stage`, `longcat_distill_refine`, `hunyuan15_sr_720p`, `hunyuan15_sr_1080p`.
### Preset-Owned Stage Names
Stage names are public and stable within a preset.
- LTX2: `base`, `refine`
- LongCat: `distill`, `refine`
- Hunyuan15: `base`, `sr_720p`, `sr_1080p`
Public overrides should reference these stage names, never stage indices.
## Stage Overrides
The main user override surface for multi-stage pipelines is:
```python
request.stage_overrides["refine"] = ...
```
Each model family should expose typed override classes for its stage names. Examples for the model families that land in PRs 6/9/10:
```python
@dataclass
class LTX2RefineStageOverride:
enabled: bool | None = None
num_inference_steps: int | None = None
guidance_scale: float | None = None
add_noise: bool | None = None
image_crf: int | None = None
video_position_offset_sec: float | None = None
@dataclass
class LongCatRefineStageOverride:
t_thresh: float | None = None
spatial_refine_only: bool | None = None
num_cond_frames: int | None = None
@dataclass
class HunyuanSRStageOverride:
num_inference_steps: int | None = None
guidance_scale: float | None = None
```
### Strictness Rules
- Stage names must exist in the selected preset.
- Override fields must be valid for that stage type.
- Unknown stage names and unknown fields must error.
## Advanced Explicit Plans
Presets should be the default API. `GenerationPlan` exists only for advanced composition or experimentation:
- building a custom workflow that is not yet standardized as a preset
- debugging or benchmarking stage combinations
- prototyping a future preset
Do not require `GenerationPlan` for normal users.
## Continuation State
Continuation must be a first-class part of the API.
### Public Contract
- `GenerationResult.state` may return a `ContinuationState`.
- `GenerationRequest.state` may accept a previously returned state.
- Most users should treat `state` as opaque and round-trip it back into the next request.
### Why This Matters
Dreamverse/LTX2 currently leaks continuation internals into app-level request fields like video conditions, audio clean latent, audio denoise mask, and segment offsets. Those should not remain top-level app-owned public API.
### State Design
Public surface:
```python
@dataclass
class ContinuationState:
kind: str
payload: dict[str, Any]
```
Internally, FastVideo should also define typed model-specific state subclasses, e.g. `LTX2ContinuationState` (PR 7) and `LongCatIntermediateState` if ever needed. Minimal stable surface: return state, pass state back in, validate that the state is compatible with the active preset.
Payload serialization: fields must be JSON-serializable or use an opaque blob-ID indirection for large tensors. This supports both the stateless OpenAI client round-trip AND future Dynamo prefill/decode disaggregation where prefill yields a state that decode hydrates across workers.
## YAML / JSON Design
YAML and JSON should be exact serializations of the typed schema, not a second unrelated config system. YAML is the primary documented format. JSON is accepted with the same schema.
### Run Config Example
```yaml
generator:
model_path: /models/ltx2
engine:
num_gpus: 1
parallelism: {tp_size: -1, sp_size: -1}
offload: {dit: false, text_encoder: false, vae: false, pin_cpu_memory: true}
pipeline:
workload_type: t2v
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
request:
prompt: "a fox running through snow"
sampling:
num_frames: 121
height: 1024
width: 1536
num_inference_steps: 8
seed: 42
output: {save_video: true, return_state: true}
stage_overrides:
refine: {num_inference_steps: 2, guidance_scale: 1.0}
```
### Serve Config Example
```yaml
generator:
model_path: /models/ltx2
engine: {num_gpus: 1}
pipeline: {workload_type: t2v, preset: ltx2_two_stage}
server: {host: 0.0.0.0, port: 8000, output_dir: outputs/}
default_request:
sampling: {num_frames: 121, height: 1024, width: 1536, num_inference_steps: 8}
output: {save_video: false, return_frames: false}
```
### Validation Rules
- top-level schema must match `RunConfig` or `ServeConfig`
- unknown keys must fail
- dotted CLI overrides are applied to the nested config before typed parsing
- parse errors must include the exact nested path that failed
## CLI Design
Inference CLI reuses the best parts of the current training authoring flow (YAML-first authoring, dotted nested overrides, typed parsing after merge) but stays stricter than training at the public boundary because it is a user-facing API surface for Python, CLI, YAML/JSON, and serving.
### Canonical CLI Forms
```bash
fastvideo generate --config run.yaml
fastvideo generate --config run.yaml --request.sampling.seed 42
fastvideo generate --config run.yaml --generator.engine.num_gpus 2
fastvideo serve --config serve.yaml
fastvideo serve --config serve.yaml --server.port 8090
```
The CLI is config-only. Beyond `--config`, CLI input uses dotted override paths into the nested schema rather than maintaining a second flat flag surface.
Implementation: YAML/JSON is loaded into a nested dict, dotted CLI overrides are applied to the nested dict, then the result is parsed into typed config objects. Flat CLI flags are rejected so the nested schema stays canonical.
## OpenAI / Server Mapping
`fastvideo serve` loads `ServeConfig`. Incoming HTTP requests are translated into `GenerationRequest` by:
1. cloning `default_request`
2. applying API request fields onto that request
3. validating against the selected preset
This is similar in spirit to the SGL pattern of merging user overrides onto model defaults.
Rules:
- HTTP request translation must not bypass typed validation.
- multi-stage defaults should come from the preset and `default_request`, not from ad hoc server-local logic.
- stateful continuation requests should accept and return typed `ContinuationState` payloads.
Landed in PR 5 for the stateless OpenAI server at `fastvideo/entrypoints/openai/`. The streaming/session server (PRs 7.5-7.9) uses the same preset/default_request merge through `ServeConfig.streaming`.
## Streaming Server + Dynamo Backend
The typed public API is consumed by three server-class integrations. They must share one execution substrate so we don't grow three near-duplicate progress loops.
### The three consumers
| Consumer | Transport | Request shape | State |
|---|---|---|---|
| Stateless OpenAI (`fastvideo/entrypoints/openai/`) | HTTP POST | `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; continuation via opaque payload if needed |
| Streaming WebSocket (`fastvideo/entrypoints/streaming/`) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment, session-scoped | Server-held session (per-GPU continuation cache); snapshot on demand |
| Dynamo native backend (`ai-dynamo/dynamo/components/src/dynamo/fastvideo/`) | Dynamo RPC endpoint | `NvCreateVideoRequest` ↔ adapter ↔ `GenerationRequest` | Aggregated today; disaggregated prefill/decode later via `ContinuationState` |
### Shared execution substrate: `VideoGenerator.generate_async`
The OpenAI server, streaming server, and Dynamo backend all want the same thing: a typed async API that yields progress events and a typed final result. FastVideo exposes exactly one canonical entry point:
```python
async def generate_async(
self,
request: GenerationRequest,
) -> AsyncGenerator[VideoEvent, None]: ...
```
Events:
```python
@dataclass
class VideoProgressEvent:
step: int
total_steps: int
stage: str # "denoise" | "refine" | "decode" | ...
@dataclass
class VideoPartialEvent:
frames: np.ndarray # shape: (num_frames, H, W, 3)
index: int # monotonic chunk index
@dataclass
class VideoFinalEvent:
video_bytes: bytes | None # mp4-encoded if requested
tensor: torch.Tensor | None # raw if requested
metadata: dict[str, Any]
continuation_state: ContinuationState | None
VideoEvent = VideoProgressEvent | VideoPartialEvent | VideoFinalEvent
```
The sync `generate_video(request=...) -> VideoResult` becomes a thin `asyncio.run` wrapper over `generate_async` that collects events and returns the final.
### Streaming server mapping
`fastvideo/entrypoints/streaming/` owns per-session state:
- `SessionStore.hydrate(state: ContinuationState) -> session_id`
- `SessionStore.snapshot(session_id) -> ContinuationState`
- Per-GPU implicit continuation cache (today's internal behavior) is wrapped as a `SessionStore` implementation.
Per-segment, the session writes a `GenerationRequest`, pipes the event stream to the WebSocket (progress → JSON messages, partial → fMP4 frames), and persists the final's `ContinuationState` into the session.
### Dynamo backend mapping
Dynamo's backend pattern (from `components/src/dynamo/sglang/`) is a pure Python import. FastVideo does not host a `fastvideo/entrypoints/dynamo/` subpackage; the integration lives in the Dynamo repo. FastVideo exposes a stable contract:
| Surface | Exposed as |
|---|---|
| Construction | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` |
| Execution (async) | `VideoGenerator.generate_async(request) -> AsyncGenerator[VideoEvent, None]` |
| Execution (sync) | `VideoGenerator.generate_video(request=...) -> VideoResult` |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` |
| Typed result | `fastvideo.api.VideoResult`, `VideoEvent`, `ContinuationState` |
| Health-check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` |
| Config dump | `config_to_dict(cfg)` (already exists) |
Request/response mapping the Dynamo adapter must perform:
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> seconds * nvext.fps -> sampling.num_frames
input_reference -> input.image_path | input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled at the adapter's output stage)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
uploaded URL -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> (reserved for future disaggregation)
```
All fields already exist (or will exist after PR 6's typed-kwarg expansion) on FastVideo's typed schema. **The adapter lives entirely in the Dynamo repo** at `components/src/dynamo/fastvideo/` — FastVideo does not host any Dynamo subpackage, dep, or CLI. The only FastVideo obligation is the stable public Python API listed above.
### Constraints this places on other sections
- **Continuation State** (see earlier section): `ContinuationState.payload` must be JSON-serializable or use an opaque blob-ID indirection for large tensors. This supports both the stateless OpenAI client round-trip *and* future Dynamo prefill/decode disaggregation, where prefill yields a state that decode hydrates across workers.
- **Typed GeneratorConfig** (see Public Python API): every flat legacy LTX2 kwarg currently used by the internal `gpu_pool.py` must have a typed home reachable from `GeneratorConfig`. Dynamo's `FastVideoArgGroup` builds the config from its CLI and must not have to know any legacy LTX2 name.
- **Public exports**: `from fastvideo import VideoGenerator`; `from fastvideo.api import GenerationRequest, SamplingConfig, ContinuationState, VideoResult, VideoEvent, VideoProgressEvent, VideoPartialEvent, VideoFinalEvent`.
## Repo Layout
### Shared Public API
`fastvideo/api/` contains the shared public API package. Current files:
- `schema.py` — `RunConfig`, `ServeConfig`, `ServerConfig`, `GeneratorConfig`, and all nested typed config dataclasses
- `sampling_param.py` — `SamplingParam` + `CacheParams` (canonical home since PR 4; former `configs/sample/base.py` location removed)
- `presets.py` — `InferencePreset`, `PresetStageSpec`, registry APIs
- `results.py` — `GenerationResult` / `VideoResult`
- `parser.py` — `from_dict`, `to_dict`, `load_yaml`, `load_json`, validation
- `overrides.py` — dotted override application
- `compat.py` — legacy Python kwargs translation
- `errors.py` — path-aware validation errors
May split further by concern in a future cleanup.
### Pipeline-Local Model-Owned Config
Model-owned presets and override types live next to the model pipeline:
```text
fastvideo/pipelines/basic/ltx2/
ltx2_pipeline.py, presets.py, stage_overrides.py, continuation.py
fastvideo/pipelines/basic/longcat/
longcat_pipeline.py, presets.py, stage_overrides.py
fastvideo/pipelines/basic/hunyuan15/
hunyuan15_pipeline.py, hunyuan15_sr_pipeline.py, hunyuan15_2sr_pipeline.py,
presets.py, stage_overrides.py
```
PR 4 landed `presets.py` for all 13 model families. Remaining colocation targets are `pipeline_configs.py` (moving `configs/pipelines/<family>.py`) and model-specific stages (moving `pipelines/stages/<family>_*.py`); see [PR plan.md](PR%20plan.md) "Pipeline Package Structure".
### Registry
Central registry (`fastvideo/registry.py`) registers preset providers rather than owning all model-specific defaults directly. It answers:
- which pipeline class corresponds to a model path
- which presets are available for that model family
- which override/state classes are valid for a selected preset
## Relationship To Current Internal Classes
This refactor does not require deleting current internals immediately.
- `FastVideoArgs` is an internal compatibility/input adapter, no longer the primary public inference type.
- `SamplingParam` now lives in `fastvideo/api/sampling_param.py` and gets model-specific defaults from presets via `_from_preset()`. All 12 `SamplingParam` subclasses have been removed and the former `fastvideo/configs/sample/` directory has been deleted entirely (PR 4). It remains an internal adapter between the preset system and the runtime.
- current `PipelineConfig` classes can remain temporarily as internal component config carriers
- the new public schema is the stable boundary above them
`VideoGenerator` accepts the new schema and translates down into current execution internals. Legacy `generate_video(..., **kwargs)` stays on the direct execution path during the compat period until SSIM/performance tests migrate in PR 11.
## Model-Specific Design
### LTX2 / Dreamverse
LTX2 needs both:
- init-time two-stage feature wiring
- request-time continuation/refine behavior
Expressed as:
- preset: `ltx2_two_stage`
- init-time fields: refine assets, optional config root, stage enablement
- request-time fields: stage override for refine behavior, optional returned continuation state
#### LTX2 Preset Example
```yaml
generator:
pipeline:
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
```
#### LTX2 Request Example
```yaml
request:
prompt: "continue the previous sequence"
state: ${previous_result.state}
stage_overrides:
refine:
num_inference_steps: 2
guidance_scale: 1.0
image_crf: 18
output:
return_state: true
```
#### LTX2 Explicit Decisions
- `config_model_path` becomes `generator.pipeline.components.config_root`
- `ltx2_refine_*` stops being a pile of top-level kwargs
- continuation internals move into `ContinuationState`
- app-level code should pass `state`, not raw latent/audio condition payloads
### LongCat
LongCat should expose a named preset like `longcat_distill_refine` with stage topology `distill` and `refine`.
User-facing override knobs remain model-specific (`t_thresh`, `spatial_refine_only`, `num_cond_frames`) but live under:
```yaml
request:
stage_overrides:
refine:
t_thresh: 0.5
spatial_refine_only: false
num_cond_frames: 8
```
### Hunyuan 1.5 SR
Hunyuan already behaves like an integrated multi-stage pipeline. Expose it via presets: `hunyuan15_sr_720p`, `hunyuan15_sr_1080p`. Users should not need to know the exact internal pipeline class split between base and SR stages. Per-stage override surface should stay small and mostly sampling-focused.
Hunyuan15 presets (`hunyuan15_t2v_480p`, `hunyuan15_i2v_480p_distilled`, `hunyuan15_t2v_720p`, `hunyuan15_i2v_720p_distilled`, `hunyuan15_sr_1080p`) are implemented (PR 4). The `Hunyuan15_*_SamplingParam` subclasses have been removed; defaults (including precomputed sigmas) come from preset `defaults` dicts. Remaining work: adding typed `HunyuanSRStageOverride` classes and colocating PipelineConfig (PR 10).
## Exact Compatibility Mapping
Intended translation layer for common current fields.
| Legacy Field | New Path |
| --- | --- |
| `model_path` | `generator.model_path` |
| `revision` | `generator.revision` |
| `trust_remote_code` | `generator.trust_remote_code` |
| `workload_type` | `generator.pipeline.workload_type` |
| `num_gpus` | `generator.engine.num_gpus` |
| `tp_size` | `generator.engine.parallelism.tp_size` |
| `sp_size` | `generator.engine.parallelism.sp_size` |
| `dit_cpu_offload` | `generator.engine.offload.dit` |
| `dit_layerwise_offload` | `generator.engine.offload.dit_layerwise` |
| `text_encoder_cpu_offload` | `generator.engine.offload.text_encoder` |
| `image_encoder_cpu_offload` | `generator.engine.offload.image_encoder` |
| `vae_cpu_offload` | `generator.engine.offload.vae` |
| `pin_cpu_memory` | `generator.engine.offload.pin_cpu_memory` |
| `enable_torch_compile` | `generator.engine.compile.enabled` |
| `torch_compile_kwargs` | split across `generator.engine.compile.backend`, `.fullgraph`, `.mode`, `.dynamic`; uncommon keys land in `.extras` |
| `enable_torch_compile_text_encoder` | `generator.engine.compile.text_encoder_enabled` |
| `enable_stage_verification` | `generator.engine.enable_stage_verification` |
| `prompt_txt` | `request.inputs.prompt_path` |
| `prompt` | `request.prompt` |
| `negative_prompt` | `request.negative_prompt` |
| `image_path` | `request.inputs.image_path` |
| `video_path` | `request.inputs.video_path` |
| `output_path` | `request.output.output_path` |
| `output_video_name` | `request.output.output_video_name` |
| `save_video` | `request.output.save_video` |
| `return_frames` | `request.output.return_frames` |
| `num_videos_per_prompt` | `request.sampling.num_videos_per_prompt` |
| `seed` | `request.sampling.seed` |
| `num_frames` | `request.sampling.num_frames` |
| `height` | `request.sampling.height` |
| `width` | `request.sampling.width` |
| `fps` | `request.sampling.fps` |
| `num_inference_steps` | `request.sampling.num_inference_steps` |
| `guidance_scale` | `request.sampling.guidance_scale` |
| `guidance_scale_2` | `request.sampling.guidance_scale_2` |
| `guidance_rescale` | `request.sampling.guidance_rescale` |
| `true_cfg_scale` | `request.sampling.true_cfg_scale` |
| `boundary_ratio` | `request.sampling.boundary_ratio` |
| `sigmas` | `request.sampling.sigmas` |
| `enable_teacache` | `request.runtime.enable_teacache` |
| `return_trajectory_latents` | `request.runtime.return_trajectory_latents` |
| `return_trajectory_decoded` | `request.runtime.return_trajectory_decoded` |
### Private Dreamverse Adapter Mapping
The mappings below are useful for private Dreamverse migration, but they should not be treated as a public FastVideo backward-compatibility promise unless and until those fields actually exist in the public repo surfaces.
| Private Adapter Field | New Path |
| --- | --- |
| `config_model_path` | `generator.pipeline.components.config_root` |
| `ltx2_refine_enabled` | `generator.pipeline.preset_overrides.refine.enabled` |
| `ltx2_refine_upsampler_path` | `generator.pipeline.components.upsampler_weights` |
| `ltx2_refine_lora_path` | `generator.pipeline.components.lora_path` |
| `ltx2_refine_num_inference_steps` | `request.stage_overrides.refine.num_inference_steps` |
| `ltx2_refine_guidance_scale` | `request.stage_overrides.refine.guidance_scale` |
| `ltx2_refine_add_noise` | `generator.pipeline.preset_overrides.refine.add_noise` |
| `ltx2_image_crf` | `request.stage_overrides.refine.image_crf` |
| `return_continuation_state` | `request.output.return_state` |
### LongCat Legacy Mapping
| Legacy Field | New Path |
| --- | --- |
| `refine_from` | `request.inputs.refine_from` |
| `stage1_video` | `request.inputs.stage1_video` |
| `t_thresh` | `request.stage_overrides.refine.t_thresh` |
| `spatial_refine_only` | `request.stage_overrides.refine.spatial_refine_only` |
| `num_cond_frames` | `request.stage_overrides.refine.num_cond_frames` |
## Validation and Error Handling
### Strict by Default
All structured inputs should be strict by default: unknown keys error, wrong types error, invalid stage names error, incompatible state/preset combinations error.
### Exceptions
The only intentionally open-ended fields are `generator.pipeline.experimental` and `request.extensions`. These must be clearly documented as unstable and unsupported for long-term API compatibility.
### Error Quality
Validation errors should include the full nested path, expected type or valid choices, and preset/stage context when relevant:
```text
Invalid field: request.stage_overrides.refine.num_inference_steps
Expected int, got "two"
Preset: ltx2_two_stage
Stage: refine
```
## Implementation Plan
### Phases 0-5: Landed
- Phase 0 — Schema Parity Inventory: inventory complete; field classifications live in `docs/design/inference_schema_parity_inventory.yaml`; parity test guard in `fastvideo/tests/api/test_schema_parity_inventory.py`.
- Phase 1 — Shared Schema: `fastvideo/api/` with typed dataclasses, parser, validation, dotted overrides, `RunConfig`/`ServeConfig`.
- Phase 2 — VideoGenerator Compat: `from_config`, `from_file`, `generate(request=...)`, legacy `from_pretrained(..., **kwargs)` and `generate_video(..., **kwargs)` as compat shims routed through typed normalization.
- Phase 3 — CLI Refactor: `fastvideo generate` and `fastvideo serve` parse nested YAML/JSON with training-style dotted overrides; flat flag expansion removed as the canonical path.
- Phase 4 — Preset System: shared registry + pipeline-local `presets.py` for all 13 families; all 12 `SamplingParam` subclasses removed; `SamplingParam` moved to `fastvideo/api/sampling_param.py`.
- Phase 5 — Server Request Translation: `fastvideo serve` loads `ServeConfig`; stateless OpenAI endpoint clones `default_request` and merges validated user overrides.
### Remaining Phases
- **Phase 6 — LTX2 Public Upstream Path** (PR 6): upstream `ltx2_two_stage` preset; upstream continuation-state contract; upstream only repo-visible/public LTX2 surfaces into FastVideo.
- **Phase 7 — Dreamverse Adapter Migration** (PR 7 + private repo work): translate private Dreamverse-only request/config fields in a private adapter; replace raw app-owned continuation kwargs with `state` in the private server; do not expand the public FastVideo compatibility promise just to match private adapter fields.
- **Phase 7.5-7.10 — Streaming Server and Dynamo Contract** (PRs 7.5-7.10): upstream the streaming server (skeleton, GPU pool, prompt enhancer, auxiliaries, router) consuming `generate_async`; land the Dynamo backend contract (`VideoGenerator.generate_async`, health-check helper) with the Dynamo backend package itself living in the Dynamo repo.
- **Phase 8 — Model Migration and Docs** (PRs 9-10, 12): colocate `configs/pipelines/<family>.py` with pipeline implementations; add typed stage override classes for multi-stage models; update basic examples to the new API; document YAML-first inference config and migration guidance.
- **Phase 8.5 — Golden-Test Migration** (PR 11): keep SSIM/performance regression tests on legacy Python generation while preset defaults are still settling; one dedicated migration pass after the preset system and model-default behavior are stable; complete this migration before removing legacy Python inference entrypoints or kwargs.
- **Phase 9 — Deprecation and Cleanup** (PR 13): deprecate direct public use of `FastVideoArgs`; deprecate direct public use of `SamplingParam`; gradually reduce public documentation for flat flags; eventually remove legacy kwargs after downstream migration is complete.
## Final Recommendation
The public FastVideo inference API is being rebuilt around:
- typed nested configs
- model-owned named presets
- semantic stage overrides
- first-class continuation state
- YAML-first CLI with dotted overrides
The primary abstraction is `InferencePreset`, not raw kwargs and not a fully manual stage graph.
The repo is moving model-specific defaults closer to each pipeline, while keeping the public schema and parsing logic centralized.
Regression and quality tests follow the rollout. Unit/entrypoint tests migrated to the typed API early, but SSIM/performance suites only move once the typed path can express all current knobs without compatibility exceptions and produces stable defaults through presets (PR 11).
End state:
- stable Python typing
- clean YAML/JSON support
- a much better CLI story
- a sane path for Dreamverse/LTX2
- a unified abstraction for LongCat, Hunyuan, and future multi-stage models
@@ -0,0 +1,285 @@
# Dreamverse ↔ FastVideo Integration
## Status
Working integration record. Captures how Dreamverse consumes the
FastVideo public API today, what's already shared, what's still ad
hoc, and what migrations land alongside each PR in the API refactor
sequence.
Pinned versions (last reconciled this session):
| Repo | Branch | Commit | Note |
|---|---|---|---|
| FastVideo (public) | `origin/main` | `70ee5d23` | PR 6 merged |
| FastVideo (public) | `will/api_7` | `3de5f833` | PR 7 in flight (typed continuation state) |
| FastVideo-internal | `will/rebase-nbv` | `1adc513e` | pre-PR-1 on the API refactor; has live realtime runtime |
| Dreamverse | `master` | `dc500330` | uses local + remote FastVideo runtimes via `server/runtime/` |
## Related Documents
- [PR plan.md](../../PR%20plan.md) — PR-by-PR sequence for the API refactor
- [apirefactor.md](../../apirefactor.md) — design spec
- [streaming-server-upstream-plan.md](streaming-server-upstream-plan.md) — upstream plan for `ui/ltx2-streaming/server/`
- `../../../Dreamverse/server/video_generation.py` — Dreamverse's worker + local `ContinuationState`
- `../../../Dreamverse/server/runtime/{factory,backend,gpu_pool,interfaces}.py` — runtime abstraction
- `../../../FastVideo-internal/fastvideo/entrypoints/realtime/{api_server,local_runtime}.py` — internal's realtime runtime (PR 7.5/7.6 upstream source)
## Surface Area
Dreamverse depends on FastVideo across three surfaces. Listed in order
of how stable each is.
### 1. Pipeline construction (stable)
`Dreamverse/server/video_generation.py:VideoGenerationWorker` calls
`VideoGenerator.from_pretrained(...)` with flat LTX-2 kwargs today.
After PR 6 the typed `GeneratorConfig` path exists; Dreamverse can
migrate at its own pace.
| Dreamverse usage | FastVideo public surface (post-PR 6) |
|---|---|
| `VideoGenerator.from_pretrained(model_path, ltx2_refine_enabled=…, …)` | `VideoGenerator.from_pretrained(config=GeneratorConfig(...))` |
| Flat `torch_compile_kwargs={…}` dict | `engine.compile.{backend,fullgraph,mode,dynamic,extras}` |
| `ltx2_vae_tiling=True` | `pipeline.vae_tiling=True` |
| `ltx2_refine_*` family | `pipeline.preset_overrides.refine.*` + `pipeline.components.upsampler_weights` |
| `enable_torch_compile_text_encoder` | `engine.compile.text_encoder_enabled` |
The legacy flat-kwarg path stays supported via `compat.py`; migration
is opt-in. PR 13's deprecation warnings are the eventual nudge.
### 2. Realtime runtime (in flight: PRs 7.5–7.6)
`Dreamverse/server/runtime/factory.py` selects a runtime backend at
process start:
```python
def create_runtime_pool() -> RuntimePool:
if os.getenv("FASTVIDEO_REALTIME_BASE_URL"):
return FastVideoRealtimePool(base_url=..., ws_url=..., default_model_id=...)
return GPUPool(get_available_gpus()) # in-process, wraps fastvideo.entrypoints.realtime.local_runtime
```
Both backends speak the same `RuntimePool` / `RuntimeSlot` Protocol
(`server/runtime/interfaces.py`):
- `acquire(client_id, websocket=None) -> (gpu_id, RuntimeSlot)`
- `release(client_id)`
- `RuntimeSlot.{join_user, user_step, leave_user, register_stream_queue, …}`
Today both impls reach into FastVideo-internal's
`fastvideo.entrypoints.realtime.local_runtime` (which exposes
`RealtimeRuntimeConfig`, `GPUPool`, `GPUSlot`). The remote backend
talks HTTP+WS to a separately-deployed runtime of the same shape.
**Contract that PR 7.5/7.6 must preserve:**
- `RealtimeRuntimeConfig` accepts `model_registry`, `default_model_id`,
`default_height/width/num_frames/fps/num_inference_steps/guidance_scale/seed/negative_prompt`,
`default_ltx2_image_crf`, `startup_warmup_{enabled,prompt,timeout_seconds}`.
- `GPUPool(gpu_ids: list[int], config: RealtimeRuntimeConfig)` constructor.
- `pool.initialize() / shutdown() / acquire() / release() / get_status()`.
- HTTP endpoints on the remote variant: `GET /healthz`, `GET /readyz`,
`GET /status`, `WS /ws`. (These already match what
`Dreamverse/server/routes/health.py` consumes.)
When PR 7.6 lands the upstream of `fastvideo/entrypoints/realtime/`,
Dreamverse should not need any code change unless we rename the import
path. **Open: do we rename `realtime/` → `streaming/` to match the
public package introduced in PR 5.5?** A deprecation alias module
keeps both working during transition.
### 3. Continuation state (PR 7)
`Dreamverse/server/video_generation.py:89 ContinuationState` is
Dreamverse's hand-rolled per-session state holder. PR 7 introduces
the typed equivalent at `fastvideo/pipelines/basic/ltx2/continuation.py`.
#### Field mapping
| Dreamverse | PR 7 `LTX2ContinuationState` | Notes |
|---|---|---|
| `video_images: list[PIL.Image]` | `video_frames: list[np.ndarray]` (uint8 H×W×3) | numpy is leaner; Dreamverse already round-trips PIL→numpy→PIL just to add noise |
| `audio_latents: torch.Tensor` `[B, C, T, mel]` | `audio_latents: torch.Tensor` (safetensors-serialized; bf16-safe) | unchanged shape; safetensors preserves dtype incl. `bfloat16` |
| `LTX2_VIDEO_CONDITIONING_FRAME_IDX` (env) | `video_conditioning_frame_idx: int` | env constant → per-state field |
| `LTX2_VIDEO_CONDITIONING_STRENGTH` (env) | `video_conditioning_strength: float` | env constant → per-state field |
| `AUDIO_CONDITIONING_NUM_FRAMES` (env) | `audio_conditioning_num_frames: int` | env constant → per-state field |
| `AUDIO_CONDITIONING_STRENGTH` (env) | `audio_conditioning_strength: float` | env constant → per-state field |
| `audio_lps` (passed into `apply_audio`) | `audio_sample_rate: int \| None` | analogous; rename worth confirming with audio team |
| Computed `prefix_sec` per segment | `video_position_offset_sec: float` | **see open question below** |
| `segment_idx` (param to apply_*) | `segment_index: int` | per-state field |
| `VIDEO_CONTEXT_NOISE`, `AUDIO_CONTEXT_NOISE`, `ENABLE_AUDIO_COND` | not on state | runtime policy / regularization knobs, not portable session data |
| `apply_video / apply_audio / save_video / save_audio_latents / clear` | not on PR-7 state class | state is a pure data carrier; runtime owns lifecycle policy |
PR-7 is a strict superset of Dreamverse's data model **plus** lifts
several env globals into per-session typed fields.
#### Lifecycle mapping
| Dreamverse pattern | `SessionStore` API |
|---|---|
| `self.continuation = ContinuationState()` per session | `state = session_store.snapshot(sid) or LTX2ContinuationState()` |
| `apply_video(req_kwargs, segment_idx)` + `apply_audio(req_kwargs, segment_idx, audio_lps)` | `state = session_store.snapshot(sid)`; runtime builds request from `state.video_frames` / `state.audio_latents` etc. |
| `save_video(frames)` + `save_audio_latents(latents)` | runtime constructs new `LTX2ContinuationState`, then `session_store.store(sid, new_state.to_continuation_state())` |
| `clear()` at end of session | `session_store.drop(sid)` |
`SessionStore` and `BlobStore` ABCs ship with thread-safe in-memory
defaults (`InMemorySessionStore`, `InMemoryBlobStore`). Dreamverse can
adopt them as-is for the local runtime; remote runtimes can plug in
redis-backed implementations later.
#### Wire format (HTTP/WS round-trip)
Dreamverse's `FastVideoRealtimePool` already speaks the realtime
runtime's HTTP+WS protocol. When PR 7.5/7.6 land state emission on
the server side, the on-the-wire payload is the public envelope:
```json
{
"kind": "ltx2.v1",
"payload": {
"schema_version": 1,
"segment_index": 3,
"video_conditioning_frame_idx": 9,
"video_conditioning_strength": 0.75,
"audio_sample_rate": 24000,
"audio_conditioning_num_frames": 5,
"audio_conditioning_strength": 0.5,
"video_position_offset_sec": 0.2,
"video": {"frames_b64": ["..."]},
"audio": {"safetensors_b64": "..."},
"metadata": {}
}
}
```
JSON-serializable end-to-end; safetensors blob preserves audio dtype
(incl. bf16). For payloads above the inline threshold a `BlobStore`
indirection replaces the b64-encoded body with `{"blob_id": "..."}`;
the blob itself stays inside the runtime that produced it.
## Migration Plan
Per PR landed, Dreamverse adoption is opt-in.
### After PR 7 merges
Single-file change in Dreamverse, ~50-line PR:
1. Replace `server/video_generation.py:89 ContinuationState` import
with `from fastvideo.pipelines.basic.ltx2.continuation import LTX2ContinuationState`.
2. Move `apply_video`, `apply_audio`, `save_video`, `save_audio_latents`,
`clear` off the state class onto `VideoGenerationWorker` (these are
runtime policy that uses the state, not part of the state itself).
3. Update `apply_audio` to read knobs from `state.audio_conditioning_num_frames`
and `state.audio_conditioning_strength` instead of the env globals
`AUDIO_CONDITIONING_NUM_FRAMES` / `AUDIO_CONDITIONING_STRENGTH`. The
env globals can stay as defaults that populate the state when a new
session starts.
4. Same treatment for video knobs: `state.video_conditioning_frame_idx`,
`state.video_conditioning_strength`.
5. Frame storage swaps `list[PIL.Image]` for `list[np.ndarray]` —
simpler `save_video` (no PIL conversion) and simpler `clear` (no
`.close()` loop).
### After PR 7.5 lands streaming server skeleton
Dreamverse's runtime/factory.py either:
- Continues to construct `GPUPool` from `RealtimeRuntimeConfig` (the
current path), now backed by the upstreamed `fastvideo/entrypoints/realtime/`.
- Or migrates to the upstream's `ServeConfig.streaming` shape and
invokes `fastvideo serve --config realtime.yaml` as the launch path.
Either way, `Dreamverse/server/runtime/interfaces.py` `RuntimePool` /
`RuntimeSlot` Protocol can stay in place — it was modeled after the
realtime runtime's surface. No interface change needed.
### After PR 7.6 lands the GPU pool upstream
- The `local_runtime.py` import in
`Dreamverse/server/runtime/gpu_pool.py:24` becomes a public import
with the same symbols (`RealtimeRuntimeConfig`, `GPUPool`,
`get_available_gpus`).
- Per-GPU continuation state inside the worker (`ltx2_continuation_images`,
`ltx2_continuation_audio_latents`) gets replaced by a `SessionStore`
reference. Dreamverse doesn't see this change — it's runtime-internal.
- `request.state` / `result.state` round-trip starts working end-to-end
on the local runtime. Dreamverse's worker can begin reading
`result.state` and feeding `request.state` between segments.
### After PR 7.10 lands the Dynamo backend contract
- `VideoGenerator.generate_async(...) -> AsyncGenerator[VideoEvent, None]`
is the canonical API.
- Dreamverse's per-segment `user_step` flow can migrate from the legacy
sync `generate_video(..., **kwargs)` path to consuming the typed
event stream. Optional; the sync wrapper stays.
## Open Questions
### `video_position_offset_sec` semantics
Dreamverse computes `prefix_sec = float(audio_extra) / 24.0` per
segment in `apply_audio`. Not persisted on `ContinuationState`.
PR-7 has `video_position_offset_sec` as a **state field**. Two valid
interpretations:
(a) **Persistent across segments** — accumulating time offset for
long sessions; useful for time-coherent audio chaining.
(b) **Per-segment hint that rides on the carrier** — runtime
overwrites every time; field is harmless redundancy.
Field's docstring leans toward (b). Decide before PR 7.6 starts
emitting/consuming it. If we land on (a), document the accumulation
rule explicitly.
### `BlobStore` / `SessionStore` lifecycle ownership
PR 7's in-memory implementations have no eviction, no TTL, no
automatic blob cleanup on state replacement. Documented as a
per-deployment policy decision.
When PR 7.5/7.6 land the live consumer, who owns:
- bounded session capacity (LRU? TTL? hard max?)
- blob `drop()` chained when a state is replaced
- session expiry on websocket disconnect
Probably the streaming server's session manager, but worth stating
explicitly in PR 7.5's design.
### `realtime/` vs `streaming/` package naming
Currently:
- Public PR 5.5 introduced `fastvideo/entrypoints/streaming/` (skeleton + typed config).
- Internal has `fastvideo/entrypoints/realtime/` (live runtime).
- Dreamverse imports from `fastvideo.entrypoints.realtime` (per the internal name).
PR 7.5 either picks one or ships a deprecation alias module.
Recommendation in `streaming-server-upstream-plan.md`: keep
`streaming/` (it's the post-PR-5.5 public name), provide
`realtime/__init__.py` as a re-export with a `DeprecationWarning` for
one release cycle so internal/Dreamverse can land import updates.
## Test Coverage on the FastVideo Side
PR 7 ships:
- `fastvideo/tests/api/test_ltx2_continuation.py` — typed
state round-trip (inline + blob), bf16 preservation, JSON
serializability, kind/version validation, schema_version guard.
- `fastvideo/tests/entrypoints/streaming/test_session_store.py` —
store/snapshot/hydrate/drop behavior on `InMemorySessionStore`;
put/get/drop on `InMemoryBlobStore`; thread-safety of both.
PR 7.5+ should add a contract test that exercises the round-trip via
the same wire format Dreamverse's `FastVideoRealtimePool` consumes.
## Changelog
| Date | Change |
|------|--------|
| 2026-04-23 | Initial draft. Captures PR 6 / PR 7 mapping; open questions on `video_position_offset_sec`, lifecycle ownership, and `realtime/` vs `streaming/` naming. |
@@ -0,0 +1,390 @@
# Dreamverse Integration Review Log
This document tracks design decisions, open questions, and integration-time
choices made while landing the public-side stacked PRs (7.7 → 8) and switching
Dreamverse from `FastVideo-internal` to public `FastVideo`. The user will
review this carefully — entries are deliberately verbose about *why*.
## Goal
Replace Dreamverse's dependency on `FastVideo-internal` with the public
`FastVideo` package, using the upstreamed streaming server stack
(`fastvideo.entrypoints.streaming.*`) where Dreamverse currently has local
copies or imports private modules.
## Surfaces Dreamverse currently uses from FastVideo-internal
(from `/home/william5lin/Dreamverse/server/`, scanned 2026-04-26):
| Dreamverse import | Internal path | Public replacement |
|---|---|---|
| `fastvideo.entrypoints.realtime.local_runtime.RealtimeRuntimeConfig` | `FastVideo-internal/fastvideo/entrypoints/realtime/local_runtime.py` | (none) — Dreamverse rewires through `streaming.gpu_pool.SubprocessGpuPool` |
| `fastvideo.entrypoints.realtime.local_runtime.GPUPool` | same as above | `fastvideo.entrypoints.streaming.gpu_pool.SubprocessGpuPool` (PR 7.6) |
| `fastvideo.configs.pipelines.base.PipelineConfig` | already in public | unchanged |
| `fastvideo.entrypoints.video_generator.VideoGenerator` | already in public | unchanged |
| `fastvideo.layers.quantization.fp4_config.FP4Config` | already in public | unchanged |
| `fastvideo.utils.maybe_download_model` | already in public | unchanged |
| `fastvideo.models.audio.ltx2_audio_processing.AudioProcessor` | already in public | unchanged |
| `fastvideo.models.loader.component_loader.ComponentLoader` | already in public | unchanged |
| `fastvideo.models.dits.ltx2.*` | already in public | unchanged |
| local copy: `Dreamverse/server/prompt_enhancer.py` (1933 lines) | mirrors `FastVideo-internal/.../prompt_enhancer.py` | `fastvideo.entrypoints.streaming.prompt.*` (PR 7.7) |
| local copy: `Dreamverse/server/prompt_safety.py` | mirrors `FastVideo-internal/.../prompt_safety.py` | `fastvideo.entrypoints.streaming.prompt.safety` (PR 7.8) |
| local copy: `Dreamverse/server/session_logger.py` | mirrors `FastVideo-internal/.../session_logger.py` | `fastvideo.entrypoints.streaming.session_logger` (PR 7.8) |
| local copy: `Dreamverse/server/rewrite_prompt_payload.py` | mirrors `FastVideo-internal/.../rewrite_prompt_payload.py` | `fastvideo.entrypoints.streaming.prompt.rewrite` (PR 7.8) |
| local copy: `Dreamverse/server/mock_server.py` (1200 lines) | mirrors `FastVideo-internal/.../mock_server.py` | `fastvideo.entrypoints.streaming.mock_server` (PR 7.8) |
| local copy: `Dreamverse/server/session_init_image.py` | mirrors `FastVideo-internal/.../session_init_image.py` | `fastvideo.entrypoints.streaming.session_init_image` (PR 7.5 — already public) |
## Design decisions made (auto-resolved)
### D-1: Realtime runtime → streaming GpuPool migration shape
**Context.** Dreamverse's `server/runtime/gpu_pool.py` thin-wraps
`fastvideo.entrypoints.realtime.local_runtime.GPUPool`, which takes a
`RealtimeRuntimeConfig(model_registry=…, default_model_id=…, default_height=…,
default_width=…, default_num_frames=…, default_num_inference_steps=…,
startup_warmup_*…)`. The public `streaming.gpu_pool.SubprocessGpuPool` takes a
typed `GeneratorConfig` + `GpuPoolConfig` + `WarmupConfig`.
The shapes differ in two important ways:
1. The internal version had a multi-model registry (`model_id → model_config`
dict). The public version is single-model (one `GeneratorConfig`).
2. The internal version flattened a few sampling defaults (height/width/frames/
steps) into the runtime config. The public version expects them as part of
the per-request `SamplingConfig`.
**Decision.** Dreamverse will:
1. Drop the multi-model registry on the integration branch (it is not used in
production today — Dreamverse boots one model per replica).
2. Construct a `GeneratorConfig` for the chosen model from `MODEL_REGISTRY[id]`
and pass it to `SubprocessGpuPool`.
3. Move the `default_height` / `default_width` / `default_num_frames` /
`default_num_inference_steps` defaults into a server-side
`default_request: GenerationRequest` template the session controller fills
from per-request input.
**Why.** Multi-model is feasible to add back later (one pool per model id,
acquire by `(session_id, model_id)`), but not on the migration branch — that
would couple the upstream switch to a feature redesign. Punting keeps the
upstream switch a pure mechanical refactor.
**Risk.** If a Dreamverse code path silently relied on the registry to swap
models per-session, the migration branch will surface that as a missing-model
error. The integration tests must exercise at least one segment per supported
model id before merging the Dreamverse branch.
### D-2: PR 7.7 prompt enhancer API surface narrower than the internal one
**Context.** The upstreamed `PromptEnhancer.enhance/auto_extend/rewrite` returns
`LLMResponse(content, provider, model, latency_ms, fallback_used)`. The internal
`enhance_prompt` / `generate_auto_prompt` / `rewrite_prompt_sequence` returns
`EnhanceResult(prompt, fallback_used, error, provider, model, latency_ms)` /
`RewriteResult(prompts, …, rollout_id, rollout_label, raw_response_text)`.
**Decision.** The Dreamverse integration branch will adapt at the call site:
- `enhancer.enhance_prompt(...)` → `enhancer.enhance(prompt)` + a thin shim
that maps the structured response into the existing `EnhanceResult` shape
for the session-controller code path. Move the shim to
`Dreamverse/server/prompting/_internal_compat.py`.
- The locked-segment / next-segment-index plumbing the internal version
built into the user payload becomes Dreamverse-side template logic in
the shim.
- The JSON-shaped responses the internal prompts assume (`{"next_prompt":
"..."}` / `{"segment_prompts": [...]}`) become Dreamverse-side
parsing in the shim, since the public `LLMResponse` is intentionally raw.
**Why.** The public surface stays minimal and provider-agnostic; the
LTX-2-specific orchestration (locked segments, rollout id/label, JSON
schemas) is an internal-UI concern, not something every public consumer
should wear. Dreamverse keeps its existing call shape; the public stays
clean.
**Open question for review:** Should we promote some of this into
`fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` (or similar)
once a second consumer appears? Logging here so we have the option.
### D-3: Multi-stage provider race (Dreamverse) vs sequential fallback (public)
**Context.** The internal enhancer runs all providers in a stage in parallel
and returns the first to succeed (`_run_provider_race`). The public
enhancer runs providers strictly sequentially with retryable-error fallback.
**Decision.** Public stays sequential for PR 7.7. The race-based fallback is
a Dreamverse-specific tail-latency optimization that depends on parallel API
budgets; promoting it would force every public consumer to have multiple
provider keys configured. Dreamverse can keep `_run_provider_race` as an
internal optimization on its side.
**Risk.** First-segment latency on Dreamverse may regress slightly when
Cerebras is having a bad minute (sequential fallback waits the full
20s timeout before trying Groq). If this is a real production concern,
add a public knob like `concurrency: int = 1` on `PromptEnhancer` that
gates a race path — but only after measuring.
### D-4: Skipping PR 7.9 router for the integration branch
**Context.** The internal stack ships a `router/main.py` that load-balances
across replicas with health checks. Dreamverse's deployment uses a single
replica per region (per `gpu_pool.py:_parse_requested_gpu_limit`).
**Decision.** Land PR 7.9 on the public side (so the surface is upstreamed)
but skip wiring it into the Dreamverse integration branch. Dreamverse's
`server/main.py` does not import from `router/`.
### D-5: Audio re-encode (PR 7.10) needed for streaming, deferred
**Context.** The internal streaming server's per-step path runs an audio
re-encode (`_re_encode_audio` inside `_stream_av_fmp4_events` /
`do_step_ltx2`) so each fMP4 segment ships with continuation-conditioning
audio. The whole-segment `pool.run()` path the public streaming server
currently uses doesn't need this. The PR plan defers re-encode integration
to PR 7.10 (`generate_async` / per-step streaming).
**Decision.** Land PR 7.10's `generate_async` on the public side. The
Dreamverse integration branch initially keeps using `pool.run()` (whole
segment, no re-encode); a follow-up branch swaps it to
`generate_async` + audio re-encode once that path is exercised end-to-end.
### D-6: `realtime/local_runtime.py` is *not* upstreamed
**Context.** It is the FastVideo-internal precursor to `streaming.gpu_pool`.
Upstreaming both would create two GPU pool implementations in the public
repo.
**Decision.** Don't upstream `realtime/local_runtime.py`. Dreamverse switches
to `streaming.gpu_pool.SubprocessGpuPool` on the integration branch. The
internal module can be deleted from FastVideo-internal at a follow-up.
## Open questions for user review
Each section below is a place the auto-decision could plausibly be wrong.
Please flip / annotate these in review.
### Q-1 Multi-model GPU pool (D-1)
Does any current Dreamverse production flow load multiple model ids
concurrently? If yes, we need to either (a) keep `realtime/local_runtime`
alive on the internal side until the public side gains a multi-model pool,
or (b) build the multi-model abstraction upstream as part of PR 7.6 follow-up
work.
### Q-2 Promoting LTX-2 prompt orchestration (D-2)
The locked-segments / next-segment-index / JSON-response orchestration is
LTX-2-specific. If Cosmos / Wan / Hunyuan ever grow a similar continuation
flow, we'll regret keeping the orchestration on the consumer side. Worth
promoting now?
### Q-3 Race-based provider fallback (D-3)
The sequential fallback in the public enhancer adds up to `timeout_ms` of
extra latency per failing provider before the next is tried. For Dreamverse
that's 20s. Should we land the race path now behind a `concurrency: int = 1`
knob, or wait until we have data?
### Q-4 Router upstream skip on Dreamverse branch (D-4)
We're upstreaming PR 7.9 (router) but not consuming it in the Dreamverse
integration branch. Is that right? Dreamverse currently has no router
component, so the answer is probably yes — but flagging.
### Q-5 generate_async cutover for the streaming path (D-5)
The plan leaves Dreamverse using `pool.run` (whole segment) initially.
Audio re-encode for cross-segment continuity is deferred to a follow-up.
Is that acceptable for the first switch, or does Dreamverse audio quality
regress relative to the internal path until 7.10 is wired in?
## PR-by-PR execution log
### PR 7.6 — already opened (#1257)
`will/api_7.6` rebased onto `origin/main`, with subprocess-pool robustness
review fixes pushed (boot_ok event, dead-worker detection, parallel shutdown,
reader-exit pending-job cleanup). 17/17 gpu_pool tests + 89/89 streaming
tests green at head.
### PR 7.7 — already opened (#1258)
`will/api_7.7` rebased onto the new 7.6 + LLM provider review fixes applied
locally (per-instance `retryable`, 4xx-non-retryable, json-decode wrap,
shared `_openai_compat.complete_openai_compatible`, `dataclasses.replace`
for the fallback marker). 29/29 prompt tests + 120/120 streaming tests green.
**Pending push** — the user opted to push this branch themselves.
### PR 7.8 — rebased onto new 7.7
`will/api_7.8` two commits replayed cleanly on the new 7.7. Adds
`fastvideo/entrypoints/streaming/{prompt/safety,prompt/rewrite,session_logger,
mock_server}.py` plus `test_auxiliaries.py`. 141/141 streaming tests green.
Notable gap vs internal version: the public `PromptSafetyFilter` ships one
classifier slot (`unsafe` label, single threshold) whereas the internal
version chained an NSFW filter and a hate-speech filter with marker-based
label matching. Multi-classifier composition is left to Dreamverse —
operators chain two filters explicitly. See **D-7** below.
### PR 7.9 — rebased onto new 7.8
`will/api_7.9` three commits replayed cleanly. Adds streaming router
(`router/{config,registry,main}.py`), `fastvideo router-serve` CLI
subcommand, and `test_router.py`. 151/151 streaming tests green.
Caveat: router/main.py uses the deprecated FastAPI `app.on_event("shutdown")`
hook — emits a DeprecationWarning. Migration to lifespan handlers is a
pre-merge cleanup item but not a blocker.
### PR 7.10 — rebased onto new 7.9
`will/api_7.10` three commits replayed with two trivial conflicts (line
wrap in `server.py`, redundant test in `test_cli_translation.py`). Adds
`VideoEvent` hierarchy, `VideoGenerator.generate_async`,
`default_health_check_request`, plus `test_generate_async.py` (273-line
contract test). 184/184 streaming + contract tests green.
### PR 8 — rebased onto new 7.10
`will/api_8` four commits → three (the 4th was a duplicate
`streaming.md` doc that 7.5 already shipped, dropped during rebase).
Adds `docs/design/server_contracts/{dynamo,index,openai}.md`,
`mkdocs.yml` entries, and `fastvideo/tests/contract/test_{dreamverse,
dynamo}_shape.py`. 206/206 streaming + contract tests green.
### Dreamverse `will/integrate-public-fastvideo`
Branch created from Dreamverse `master`. Single change: `pyproject.toml`
swaps `fastvideo = { path = "../FastVideo-internal", editable = true }`
to point at `../FastVideo`. Comment added linking back to this review
doc.
**Verified:** every TRACKED `from fastvideo.*` import in Dreamverse
(`server/video_generation.py` only) resolves against the public
package — except `fastvideo.layers.quantization.fp4_config.FP4Config`
(see **D-7** / Q-6 below).
**Untracked WIP** in `Dreamverse/server/{config,prompting,runtime,session}/`
imports `fastvideo.entrypoints.realtime.local_runtime` (D-6); this
branch does not migrate that WIP. The user's existing untracked work
stays untouched and will need a separate follow-up to consume
`streaming.gpu_pool.SubprocessGpuPool`.
## Test ladder (built-up to e2e per user request)
Each rung verifies the integration switch at one layer. Run from the
narrowest to the broadest before running the full e2e against real
GPU + model weights.
| # | Layer | Command | Status against the switched stack |
|---|---|---|---|
| 1 | Public FastVideo unit + contract tests | `pytest fastvideo/tests/api/ fastvideo/tests/entrypoints/streaming/ fastvideo/tests/contract/` | 358/358 passing on `will/api_8` |
| 2 | Public FastVideo FP4 lazy-import | `pytest fastvideo/tests/ops/quantization/test_fp4_config.py` | 3/3 passing |
| 3 | Dreamverse Python tests | `cd Dreamverse && uv run pytest server/tests/ -k "not stress and not benchmark and not health_endpoint"` | 73/73 passing against public FastVideo |
| 4 | Dreamverse FE unit/integration (vitest) | `cd Dreamverse/apps/web && npm test` | 54/86 passing — 32 failures are pre-existing copy-mismatches in `reducer.test.ts` etc., not caused by the switch |
| 5 | Backend HTTP smoke (Playwright) | `cd Dreamverse/apps/web && PLAYWRIGHT_SKIP_WEBSERVER=1 PLAYWRIGHT_BASE_URL=http://127.0.0.1:8009 npx playwright test e2e/backend-health.spec.ts` | 4/4 passing (5th correctly skipped because devtools-only route is off) |
| 6 | Frontend shell smoke (Playwright) | `npx playwright test e2e/frontend-shell.spec.ts` | Pending — requires Next.js dev server to be reachable; was stuck during this run, needs a clean restart |
| 7 | Full e2e preset generation | `npx playwright test e2e/preset-prompt-generation.spec.ts` | **8/8 passing** end-to-end after restart with `CUDA_VISIBLE_DEVICES=4 ENABLE_TORCH_COMPILE=0 FASTVIDEO_GPU_COUNT=1 FASTVIDEO_ENABLE_DEVTOOLS=1`. BE warmup + GPU 4 idle slot let `/readyz` flip green; the spec verifies preset → WS → backend handshake → "Generating video…" state. |
### How to reproduce e2e tier 7 from cold
```
# 1. BE — picks an idle GPU and skips torch.compile (avoids the
# aarch64 cross-compiler bug in the conda env's triton stack).
cd ~/Dreamverse
set -a; source ~/.env; set +a
CUDA_VISIBLE_DEVICES=4 ENABLE_TORCH_COMPILE=0 \
FASTVIDEO_ENABLE_DEVTOOLS=1 FASTVIDEO_GPU_COUNT=1 \
uv run dreamverse-server &
# 2. Wait for /readyz (~2 min for warmup x2 segments)
until curl -fsS http://127.0.0.1:8009/readyz >/dev/null; do sleep 5; done
# 3. FE
cd ~/Dreamverse/apps/web
BACKEND_URL=http://127.0.0.1:8009 NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
npm run dev:devtools &
# 4. Playwright
cd ~/Dreamverse/apps/web
PLAYWRIGHT_SKIP_WEBSERVER=1 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
BACKEND_URL=http://127.0.0.1:8009 \
npx playwright test --project=chromium --reporter=list
```
### Surfaced during the e2e debug pass (logged here for follow-up)
* **`SamplingParam has no field ltx2_image_crf`** — Dreamverse's
`server/video_generation.py:406` passes `ltx2_image_crf=0.0` to a
`SamplingParam(...)` constructor. The internal SamplingParam (in
`fastvideo/configs/sample/base.py`) declared this field; the public
`fastvideo.api.sampling_param.SamplingParam` does not. Currently
the BE logs an `ERROR` and silently drops the kwarg; warmup still
succeeds because the field is non-load-bearing for FP4-disabled
inference. Either re-add the field to the public schema or update
Dreamverse to stop passing it. **D-8.**
* **`aarch64-conda-linux-gnu-cc` triton compile failure** — the conda
env we boot from injects an ARM cross-compiler ahead of `gcc` on
`$PATH`, so `torch._inductor`'s triton launcher fails compilation.
Setting `ENABLE_TORCH_COMPILE=0` bypasses it. Long-term fix: clean
the conda env's compiler shadowing or add a `CC=gcc` override in
Dreamverse's worker bootstrap. **D-9.**
* **GPU pool starts but warmup OOMs on a shared GPU** — when
`CUDA_VISIBLE_DEVICES` lands on a GPU another tenant is using
(107 GiB-pegged training run on GPU 0 in this case), LTX-2 warmup
fails with OOM. Picking an idle GPU (4-7 here) is a manual step.
A pre-warm probe that checks free memory before booting the pool
would prevent this. **D-10.**
* **ffmpeg fragment write `Broken pipe`** — when the WS client closes
before the backend finishes streaming the first segment, ffmpeg
hits `[Errno 32] Broken pipe`. Currently Dreamverse's
`gpu_pool.handle_command` re-raises this as a session error,
which then propagates to "User step failed". Cosmetic for now —
swallowing pipe-broken on intentional disconnect would clean up
the logs. **D-11.**
## Additional integration gaps surfaced during the switch
### D-7: `FP4Config` is private-only
**Context.** `Dreamverse/server/video_generation.py:271` imports
`fastvideo.layers.quantization.fp4_config.FP4Config` and assigns it to
`pipeline_config.dit_config.quant_config`. The 411-line module lives only
in `FastVideo-internal/fastvideo/layers/quantization/fp4_config.py` and
hard-imports `flashinfer` at module top — it never made the public
upstream pass. Public has `base_config.py` and `absmax_fp8.py` only.
**Decision (provisional).** Don't upstream `fp4_config.py` in this
session. Reasons:
1. It introduces a new external dependency (`flashinfer`) the public
package has avoided so far.
2. The class hard-codes LTX-2 layer paths
(`ltx2.blocks.{i}.attn1.to_q` etc.) — this is "LTX-2-specific FP4",
not generic FP4. Belongs colocated with `pipelines/basic/ltx2/` if
it goes anywhere.
3. The FP4 pre-quantize/forward op surface is the kind of thing where
a careful review pass matters more than a bulk copy.
**What this means for the integration branch.** Dreamverse will boot
fine; only the FP4-quantized path inside `video_generation.py:283`
will fail (lazy import). For workflows that don't enable FP4
quantization, the integration is complete.
### Q-6 (review): how to land FP4Config publicly?
Two reasonable next steps:
1. **Colocate.** Move FP4 code to `fastvideo/pipelines/basic/ltx2/quantization.py`
with `flashinfer` as an optional extra: `pip install fastvideo[fp4]`.
Refactor `FP4QuantizeMethod` to take its layer-prefix list from a
pipeline-config field instead of hardcoding ltx2 paths so the
approach generalizes.
2. **Keep private.** Treat FP4 as a Dreamverse-side concern — Dreamverse
imports `fp4_config` from the internal repo via a thin shim. Public
FastVideo stays focused on generic surfaces. This means the
"FastVideo-internal removable" goal is partially undone.
Recommendation: option 1 once the API refactor settles — wait until
the LTX-2 colocation step (PR 9 / 10 territory) and land FP4 there.
@@ -0,0 +1,518 @@
# Handoff: LTX-2 NVFP4 wire-up + Dreamverse launch-demo skill
This document hands off in-flight work to the next coding agent. It covers
two related streams that landed across two repos:
1. **FastVideo** (`will/ltx2_sr_port`): wire NVFP4 (NVIDIA's block-scaled
FP4) inference + per-component torch.compile + supporting parity fixes
so the public package matches `FastVideo-internal` for the LTX-2
distilled streaming path used by Dreamverse.
2. **Dreamverse** (`will/integrate-public-fastvideo`): switch the GPU
worker to the typed `GeneratorConfig` API, rename `FP4Config` →
`NVFP4Config`, add a `launch-demo` skill + canonical
`serve_configs/streaming_demo.yaml` for `fastvideo serve --config`.
Stack remains green: 222/222 FastVideo unit/contract/api tests pass; 8/8
Playwright e2e tests pass against the live `dreamverse-server` + Next.js
stack.
---
## Repo + branch state
| Repo | Path | Branch | Tip |
| --- | --- | --- | --- |
| FastVideo | `/home/william5lin/FastVideo` | `will/ltx2_sr_port` | `c6c14c55` |
| Dreamverse | `/home/william5lin/Dreamverse` | `will/integrate-public-fastvideo` | `3d7fd89` |
| Reference (read-only) | `/home/william5lin/FastVideo-internal` | (their) `main` | source of truth for parity |
> **Working branch on FastVideo is `will/ltx2_sr_port`, not the default checkout.**
> The shell may report `will/uv-pip-install-everywhere` because that was
> the earlier checkout. Run `git checkout will/ltx2_sr_port` before
> picking up FastVideo work.
### Live processes (do not duplicate)
```
:8009 dreamverse-server pid 2453227 (warmed, /readyz returns 200)
:5274 next-server (dev) pid 2399103 (devtools build)
```
### Stashes
* FastVideo: `stash@{0}: WIP on main: …HunyuanVideo plugin…` — pre-existing,
unrelated to this work, do not pop.
* Dreamverse: `stash@{0}: wip: server modular refactor (split
config/prompting/runtime/session)` — 3867 lines of orphan modular split
off this branch. Do not pop on this branch; recover on a separate
feature branch if anyone wants to resurrect it.
---
## What landed (FastVideo: `cfccd292..c6c14c55`)
Six commits on top of the i2v / continuation latent port:
```
c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow
94c983a2 refactor(quant): rename FP4 → NVFP4 to disambiguate from other FP4 variants
42b30bf9 feat(ltx2): wire FP4 inference through fastvideo.layers.quantization
6da342ba feat(compile): per-component compile + transformer_refine + prepare hook
221cb20a feat(api): typed per-component CompileConfig + FastVideoArgs carriers
a4760bae fix(api): propagate generic refine_* args + match internal randn
```
Each commit message has the rationale. Highlights below.
### `a4760bae` — three small parity fixes
* `FastVideoArgs.__post_init__` now calls `_resolve_refine_args()` which
copies the public-facing generic `refine_*` knobs onto their
`ltx2_refine_*` runtime carriers. Was missing → callers that set
`refine_lora_path=...` saw "applied to 0 layers" warnings as the value
was silently dropped.
* `_randn_ltx2_video_latents` patch path reverted from `randn_tensor` →
`torch.randn` to bit-match internal under single-generator inference.
Identical for a single `torch.Generator` but diverges for
`list[Generator]` (per-sample seeds).
* Classified 19 `refine_*` / `ltx2_refine_*` / i2v / `ltx2_audio_*` /
`ltx2_conditioning_latent_*` / `ltx2_video_conditions` fields in the
schema-parity inventory yaml.
### `221cb20a` — typed CompileConfig + FastVideoArgs carriers
`CompileConfig` (in `fastvideo/api/schema.py`) gained per-component knobs:
```python
@dataclass
class CompileConfig:
enabled: bool = False # master DiT switch
backend / fullgraph / mode / dynamic / extras # master kwargs
# Per-component overlays, None = inherit master `enabled`
text_encoder_enabled: bool | None = None
vae_enabled: bool | None = None
audio_vae_enabled: bool | None = None
# Per-component kwargs override master when non-empty
dit_kwargs: dict = ...
text_encoder_kwargs: dict = ...
vae_kwargs: dict = ...
audio_vae_kwargs: dict = ...
```
Matching carrier fields on `FastVideoArgs`:
`enable_torch_compile_text_encoder/vae/audio_vae` and
`torch_compile_kwargs_dit/text_encoder/vae/audio_vae`. Compat layer
round-trips them through `legacy_from_pretrained_to_config` and
`generator_config_to_fastvideo_args`. **No behavior change yet** — these
are surface ports only; consumed in the next commit.
### `6da342ba` — refine + per-component compile + prepare_for_compile
`composed_pipeline_base.post_init` now:
* compiles `transformer_refine` alongside `transformer` and
`transformer_2` whenever the DiT compile flag is on (closes the LTX-2
stage-2 silent-eager bug);
* dispatches per-component compile loops (text encoder, VAE, audio VAE)
with per-component kwargs falling back to master when empty;
* calls `module.prepare_for_compile()` on each compiled submodule
before invoking `torch.compile` (hook protocol — model-specific).
Implemented on `Gemma3` to materialize HF weights outside Dynamo's
tracer.
### `42b30bf9` — NVFP4 LTX-2 inference wire-up *(largest)*
End-to-end:
1. `models/dits/ltx2.py` — swap `nn.Linear` → `ReplicatedLinear` for the
FP4-eligible subset (`LTXSelfAttention`, `LTXDistributedSelfAttention`,
`FeedForward`/`GELUApprox`); plumb `quant_config` and `prefix=` from
`BasicAVTransformerBlock` → `_init_transformer_blocks` → `LTXModel`
→ `LTX2Transformer3DModel`. Other linears
(`TimestepEmbedding`, `PixArtAlphaTextProjection`, `patchify_proj`,
`proj_out`, `AdaLayerNormSingle.linear`) stay `nn.Linear` —
matches internal exactly.
2. Port `_supports_prequantized_input` and
`_linear_project_with_optional_prequant` helpers. Attention forward
pre-quantizes input once (`quantize_input`), reuses the
`(x_fp4, x_scale, x_global_sf)` tuple for k/v projections when
`context is x` — bit-matches internal's fused path.
3. `models/loader/fsdp_load.py` — new `_maybe_convert_model_to_nvfp4`
helper detects via `isinstance(quant_method, NVFP4QuantizeMethod)`
(no flag); calls `convert_model_to_nvfp4` to materialize
`_nvfp4_weight*` / `_nvfp4_alpha` / `_weight_global_sf` buffers.
`flashinfer` import is lazy (inside the helper), so the loader is a
no-op on hosts without flashinfer.
4. `layers/quantization/__init__.py` — registered `"NVFP4"` in
`QuantizationMethods` literal + `get_quantization_config`.
5. `api/compat.py` + `fastvideo_args.py` — typed
`engine.quantization.transformer_quant: "NVFP4"` resolves to a
concrete `NVFP4Config()` instance, carried on `FastVideoArgs.transformer_quant`,
pinned onto `pipeline_config.dit_config.quant_config` in
`__post_init__._apply_transformer_quant`. **The explicit setter
(legacy mutation pattern) wins** if `dit_config.quant_config` is
already non-None.
6. `layers/linear.py` — `LinearBase.__init__` now falls back to
`UnquantizedLinearMethod` when `quant_config.get_quant_method` returns
`None`. `NVFP4Config` only tags a curated subset of LTX-2 layers, and
the previous `assert quant_method is not None` would crash any
non-tagged layer that received a quant_config.
### `94c983a2` — FP4 → NVFP4 rename
NVIDIA's specific block-scaled fp4 format (e2m1 mantissa, fp32 alpha,
`layout_128x4` scale layout, group size 16) — distinct from MX-FP4 /
OCP-FP4 / generic e3m0. Mechanical rename, no behavior change:
* `fp4_config.py` → `nvfp4_config.py`
* `FP4Config` → `NVFP4Config`; `get_name()` returns `"nvfp4"`
* `FP4QuantizeMethod` → `NVFP4QuantizeMethod`
* `convert_model_to_fp4` → `convert_model_to_nvfp4`
* `QuantizationMethods` literal: `"FP4"` → `"NVFP4"`
* registered buffer names: `_fp4_weight`/`_fp4_alpha` →
`_nvfp4_weight`/`_nvfp4_alpha`
* loader helper renamed
* test file rename + symbol updates
Internal-scope torch op namespace `fastvideo_fp4::*` and
`_get_ltx2_fp4_stage_profile` deliberately left as-is — purely
internal naming that mirrors FastVideo-internal.
### `c6c14c55` — contract + numerical lock-in tests
* `fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py` (6 tests):
asserts that `LTXSelfAttention.to_q/to_k/to_v/to_out` are
`ReplicatedLinear`; `NVFP4Config()` attaches `NVFP4QuantizeMethod`
on the quantized subset with the correct `layer_prefix`; non-tagged
projections (cross-attn K/V, audio attn, audio FFN) fall back to
`UnquantizedLinearMethod`; `BasicAVTransformerBlock` propagates
`quant_config` and `prefix` correctly to all 4 attention modules +
FFN at once.
* `fastvideo/tests/api/test_typed_quant_flow.py` (4 tests): asserts
typed `engine.quantization.transformer_quant: "NVFP4"` →
`NVFP4Config()` instance flow; default leaves `transformer_quant`
None; explicit `dit_config.quant_config = …` wins over typed carrier.
---
## What landed (Dreamverse: `248060b..3d7fd89`)
Three commits on top of the e2e tier:
```
3d7fd89 feat(skill): launch-demo orchestrator + fastvideo serve YAML
d80c2a8 refactor(server): drive FP4 + per-component compile via typed GeneratorConfig
4cc6b30 chore: gitignore Playwright + Next.js build artifacts under apps/web
```
### `d80c2a8` — server/video_generation.py refactor
Three coordinated changes in the GPU worker:
* Replace legacy `load_kwargs` dict + `VideoGenerator.from_pretrained(model_root, **kwargs)`
call with the typed `GeneratorConfig` (`EngineConfig` /
`OffloadConfig` / `CompileConfig` / `PipelineSelection` /
`ComponentConfig`). Refine knobs move from `ltx2_refine_*` flat
kwargs into `preset_overrides["refine"]`. **The in-memory
`pipeline_config` pin** (`dit_config.quant_config = NVFP4Config()`)
keeps using the legacy `experimental["pipeline_config"]` carrier
because typed `transformer_quant: "NVFP4"` doesn't yet support
setting `layer_profile`.
* Rename FP4 → NVFP4.
* Re-enable `"mode": "max-autotune-no-cudagraphs"` (was commented out).
Closes the last known divergence vs FastVideo-internal in the
worker-level path trace.
### `4cc6b30` — gitignore Playwright/Next.js artifacts
Added `apps/web/{node_modules,.next,test-results,playwright-report}` to
`.gitignore`. Mirror of the existing `prod-ui/` ignore set.
### `3d7fd89` — launch-demo skill
```
.agents/skills/launch-demo/
├── SKILL.md
└── scripts/
├── launch_demo.sh # orchestrator: BE + FE + health probes + Ctrl-C trap
├── launch_backend_dreamverse.sh # uv run dreamverse-server (default)
├── launch_backend_fastvideo.sh # uv run fastvideo serve --config (typed path)
└── launch_frontend.sh # next dev (devtools/dev/single5s)
serve_configs/
└── streaming_demo.yaml # canonical ServeConfig matching internal/ui
```
YAML has every field annotated with the internal source line it mirrors:
LTX-2 distilled, NVFP4, 121 frames @ 1088×1920 24fps, 5 inference steps,
2-step refine gs=1.0 add_noise=true, max-autotune-no-cudagraphs compile,
121-frame default request, 300s session timeout, 6 segment cap, av_fmp4
streaming, cinematic-drone warmup prompt, 2400s warmup timeout, 9
conditioning frames + 0 end-offset, prompt enhancer on with cerebras /
gpt-oss-120b / 20s timeout.
**Two BE flavors documented in SKILL.md:**
| `BE_FLAVOR=` | Boots | Routes served | FE compatible |
| --- | --- | --- | --- |
| `dreamverse` (default) | `dreamverse-server` | `/healthz`, `/readyz`, `/curated-presets`, `/v1/stream`, devtools, session monitor | ✓ full |
| `fastvideo` | `fastvideo serve --config <yaml>` | `/health`, `/v1/stream` | ⚠ FE will surface fetch errors for `/curated-presets`, `/readyz` until those routes migrate into FastVideo's `build_app` |
The fastvideo flavor exists today as the verifiable typed-config path
(YAML parses, streaming worker boots, dotted overrides work). It is not
yet a drop-in for the FE — see "Open follow-ups" below.
---
## Verified
* `222 passed, 1 skipped` across `fastvideo/tests/api/`,
`fastvideo/tests/contract/`,
`fastvideo/tests/ops/quantization/test_nvfp4_*`,
`tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py`.
* `8 passed` Playwright e2e (backend-health 5, frontend-shell 2,
preset-prompt-generation 1) against the live `dreamverse-server`
+ Next.js stack.
* `streaming_demo.yaml` parses cleanly against `ServeConfig`; the
validation path of `fastvideo serve --config <yaml>` runs without
error and accepts dotted overrides like `--server.port 8010`.
* FastVideo `bash -n` clean across all four launch scripts.
---
## Critical context (gotchas a successor should know)
### NVFP4 layer set is asymmetric — by design
`NVFP4Config.fp4_layers` covers:
* `attn1.{to_q,to_k,to_v,to_out}` — full self-attention
* `attn2.{to_q,to_out}` — cross-attn Q + out only (text context not quantized)
* `audio_to_video_attn.{to_q,to_out}` — AV cross Q + out
* `video_to_audio_attn.{to_k,to_v}` — VA cross K + V
* `ffn.{fc_in,fc_out}` — video FFN
* `adaln_single.linear` — but this is `nn.Linear` (not `LinearBase`),
so it never actually gets FP4'd. List entry has no effect; matches
internal.
**NOT in the set:** audio self-attention (`audio_attn1.*`), audio
cross-attention (`audio_attn2.*`), audio FFN (`audio.ffn.*`). Audio
path is cheap enough that quant overhead isn't worth it. Test
`test_basic_av_block_propagates_quant_config_to_all_children` locks
this in — if you add audio quantization later, update the test.
### `LinearBase` fallback is load-bearing
`fastvideo/layers/linear.py:191-202`: when `quant_config.get_quant_method`
returns `None` (layer not in the quant config's set), we fall back to
`UnquantizedLinearMethod`. **Do not remove this fallback** — it would
break every non-tagged `ReplicatedLinear` constructed with a
`NVFP4Config`, and `assert quant_method is not None` in
`ReplicatedLinear.__init__` would fire on unmatched layers.
### Typed `transformer_quant` precedence
`FastVideoArgs._apply_transformer_quant` only writes
`dit_config.quant_config` when it's currently `None`. If a caller has
explicitly set `pipeline_config.dit_config.quant_config = NVFP4Config(...)`,
the explicit setter wins. Dreamverse's `video_generation.py` relies on
this — it sets `NVFP4Config()` directly because the typed
`transformer_quant: "NVFP4"` doesn't expose `layer_profile`.
### Pre-existing AbsMaxFP8 test failure is NOT mine
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on this branch with the same error
("AssertionError not raised"). I confirmed via `git stash` that the
failure pre-dates my changes. Not blocking; tracked as separate tech
debt.
### `transformer_refine` is auto-compiled with the master DiT flag
Set `enable_torch_compile=True` and `transformer_refine` compiles
along with `transformer` and `transformer_2`. There is **no separate
`enable_torch_compile_refine` flag** — by design, refine inherits the
DiT compile state to keep the typed surface small. If you need them
decoupled, add a new field; don't repurpose existing ones.
### `prepare_for_compile` is a duck-type protocol, not a base class method
Defined nowhere; called via `getattr(module, "prepare_for_compile", None)`
in `composed_pipeline_base._maybe_compile_pipeline_module`. Currently
only Gemma implements it (to materialize HF weights outside Dynamo).
Add to other models that have lazy external state if you observe
compile-time graph breaks.
### Public typed `PromptEnhancerConfig.provider` is `Literal["cerebras", "groq"]`
Internal supports `"cerebras_ifm"` (config.py:143). The public typed
schema does not. The `streaming_demo.yaml` defaults to `"cerebras"`.
For agents that need `cerebras_ifm`, the `dreamverse-server` flavor
respects the `FASTVIDEO_PROMPT_PROVIDER` env var (legacy path);
`fastvideo serve --config` does not currently expose it.
### Dreamverse `pipeline_config` is still a Python object passed via `experimental`
The typed `GeneratorConfig` doesn't have a clean home for an
in-memory `PipelineConfig` instance with mutated `dit_config`. We
pass it via `pipeline.experimental["pipeline_config"]` — the
`compat.py` legacy adapter recognizes that key and threads it through
to `FastVideoArgs.from_kwargs`. This is fine but not pretty; if
someone designs a typed `dit_config` carrier later, this becomes
obsolete.
### `fastvideo serve --config` is not yet a drop-in for the FE
`fastvideo.entrypoints.streaming.server.build_app` exposes only
`/health` and `/v1/stream`. The Dreamverse Next.js shell expects
`/healthz`, `/readyz`, `/status`, `/curated-presets`,
`/curated-presets/append`, `/prompt-system-config`, and the devtools
routes. These all live in `Dreamverse/server/main.py` +
`Dreamverse/server/routes/`. Until they migrate into FastVideo's
`build_app` (or are exposed via a Dreamverse-side proxy), the
`BE_FLAVOR=fastvideo` flavor is for verifying the typed serve config
path only — not for full FE compatibility.
---
## Open follow-ups (prioritized)
### High
1. **Migrate FE-required routes into FastVideo's `build_app`.**
`/healthz`, `/readyz`, `/status` look obviously fastvideo-side
(they're streaming-server health). `/curated-presets` and
`/prompt-system-config` are operator-side surfaces and should
probably stay in Dreamverse (or migrate as opt-in routes that the
FE feature-detects). Without this, `BE_FLAVOR=fastvideo` is
permanently a "diagnostic" flavor. Closes the
`launch-demo` skill TODO.
2. **AbsMaxFP8 test failure cleanup.** Pre-existing. Either fix the
test (`AbsMaxFP8LinearMethod.create_weights` no longer asserts on
invalid dtype — restore the assert if intentional, otherwise drop
the test).
### Medium
3. **Add `cerebras_ifm` to public `PromptEnhancerConfig.provider`
Literal.** Trivial schema change; needs paired enhancer-side
provider implementation in
`fastvideo/entrypoints/streaming/prompt/providers/`.
4. **Expose `layer_profile` on typed `engine.quantization`.** Today
`transformer_quant: "NVFP4"` always constructs `NVFP4Config()`
with the default `layer_profile="refine"`. To support stage-1
profiles (no `attn2.to_out`, no cross-modal AV) via typed config,
add `transformer_quant_layer_profile: str | None = None` and
thread it through `compat.py`. Dreamverse currently dodges this
by setting `NVFP4Config()` directly via `experimental`.
5. **Typed `dit_config.quant_config` carrier.** The
`experimental["pipeline_config"]` escape hatch in Dreamverse
should eventually become a typed field. Design TBD.
### Low
6. **Audio attention quantization profile.** If an audio-quant
profile is added to `NVFP4Config.fp4_layers` (currently audio attn
and FFN are bf16), update
`test_basic_av_block_propagates_quant_config_to_all_children`.
7. **Schema parity inventory.** A few internal-only fields are not
exposed publicly (`PROMPT_HTTP_TIMEOUT_MS`,
`PROMPT_INITIAL_STAGE_TIMEOUT_MS`, `PROMPT_TEMPERATURE`,
`PROMPT_MAX_COMPLETION_TOKENS`, `PROMPT_AUTO_SLEEP_MS`,
`PROMPT_AUTO_TIMEOUT_MS`, the curated-presets file paths).
These all flow via env vars on `dreamverse-server` today; if
`fastvideo serve --config` becomes the canonical entrypoint,
they'll need typed homes.
8. **Empty `apps/web/test-results/` directory locally.** The
`.gitignore` entry I added makes it invisible to `git status`,
but the dir itself still has a stale `.last-run.json` (45 bytes)
from a prior Playwright run. Harness blocked auto-cleanup
("pre-existing files"); the user can `rm -rf
apps/web/test-results` whenever convenient.
---
## How to pick up work
### Quick orientation (run these first)
```bash
# FastVideo state
cd /home/william5lin/FastVideo
git checkout will/ltx2_sr_port
git log --oneline cfccd292..HEAD # six commits added this round
.venv/bin/python -m pytest fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header # expect 222 passed, 1 skipped
# Dreamverse state
cd /home/william5lin/Dreamverse
git log --oneline 248060b..HEAD # three commits added this round
cat serve_configs/streaming_demo.yaml | head -40
ls .agents/skills/launch-demo/
# Live stack health (already running on this host)
curl -s http://localhost:8009/readyz | head -c 200
curl -s http://localhost:5274/ | head -c 100
( cd apps/web && npx playwright test --reporter=line ) # expect 8 passed
```
### Reference docs
* **FastVideo internal/ui parity source:** `../FastVideo-internal/ui/ltx2-streaming/server/config.py`
* **NVFP4 source on internal:** `../FastVideo-internal/fastvideo/layers/quantization/fp4_config.py`
* **Worker-trace audit:** `../FastVideo/dreamverse_review.md` (D-1
multi-model, D-5 audio re-encode, prior gap inventory)
* **Schema parity inventory:** `docs/design/inference_schema_parity_inventory.yaml`
* **PR-plan for the broader migration:** `../FastVideo/PR plan.md`
### Files most likely to need touches in follow-ups
* `fastvideo/api/schema.py` — `CompileConfig`, `QuantizationConfig`,
`PromptEnhancerConfig` Literal extension.
* `fastvideo/api/compat.py` — typed → flat translation.
* `fastvideo/fastvideo_args.py` — carrier fields and
`_apply_transformer_quant`.
* `fastvideo/entrypoints/streaming/server.py::build_app` — add
`/healthz`, `/readyz`, `/status` routes for FE compatibility (high
priority follow-up #1).
* `Dreamverse/server/video_generation.py` — typed `GeneratorConfig`
builder (current).
* `Dreamverse/serve_configs/streaming_demo.yaml` — every parity
knob; edit here, not in shell scripts.
---
## Don't / Cautions
* **Don't pop the Dreamverse stash on this branch.** It's 3867 lines
of orphan modular refactor (server/{config,prompting,runtime,session}/)
with broken absolute imports. If anyone wants to resurrect it, do so
on a separate feature branch.
* **Don't remove the `LinearBase` `UnquantizedLinearMethod` fallback.**
See "Critical context" above.
* **Don't repurpose `enable_torch_compile` to mean DiT-only.** It also
drives `transformer_refine` and `transformer_2` compile. Add a new
flag if decoupling is needed.
* **Don't change `NVFP4Config` buffer names back to `_fp4_*`.** The
rename is intentional to disambiguate from MX-FP4 / OCP-FP4.
* **Don't bypass the typed surface for new options.** New compile /
quant / refine knobs should land on the dataclass + compat.py +
parity inventory together. The existing test suite locks this in.
* **Don't merge to main without a CI run that covers FP4.** Current
CI doesn't run flashinfer-dependent paths; the wiring tests in
`test_nvfp4_ltx2_wiring.py` are CPU-only by design and don't
exercise the actual FP4 kernels.
---
*Last updated: end of session that landed `c6c14c55` on FastVideo and
`3d7fd89` on Dreamverse. Stack remains green; no dirty state.*
@@ -0,0 +1,539 @@
# FastVideo Streaming Server Upstream — Design & Plan
## Status
Exploration / design draft. Captures the re-evaluation triggered by the
decision to upstream `FastVideo-internal/ui/ltx2-streaming/server/` into
the public repo. Not yet approved for execution.
## Related Documents
- [PR plan.md](../../PR%20plan.md) — PR-by-PR implementation plan for the API refactor
- [apirefactor.md](../../apirefactor.md) — design spec this plan implements
- `../../../FastVideo-internal/ui/ltx2-streaming/` — upstream source (server side)
- `../../../dynamo/` — local clone of ai-dynamo/dynamo; backend patterns at
`components/src/dynamo/{vllm,sglang,trtllm}/` and `CLAUDE.md` files
- https://github.com/ai-dynamo/dynamo/pull/7544 — draft PR that promotes
FastVideo to a native Dynamo backend (CLOSED, superseded — but establishes
the integration shape)
## Context
The internal `FastVideo-internal/ui/ltx2-streaming/` directory contains a
complete LTX2 streaming service. The user has decided:
- **Frontend clients** (`client/`, `prod-ui/`) stay in the internal repo
- **Everything server-side** — FastAPI/WebSocket server, GPU pool, prompt
enhancer, router, auxiliaries — will be upstreamed to FastVideo
In parallel, FastVideo is becoming a **first-class Dynamo backend** (same
tier as vllm, sglang, trtllm). The refactor must produce an API that
Dynamo's `components/src/dynamo/fastvideo/` package can consume as a
pure Python import, without re-introducing the legacy flat-kwarg
surface. Draft PR ai-dynamo/dynamo#7544 defines the concrete integration
shape we need to support.
This materially changes the tail of the API refactor plan. The current
PR 5 ("wire `ServeConfig.default_request` into the OpenAI-compatible
HTTP server") addresses only the stateless endpoint; the real upstream
target is a much larger, session-based stack **plus** a clean Dynamo
backend contract.
This document captures:
- what's being upstreamed and where it lands
- four design decisions that shape the upstream (continuation model,
streaming server layout, LLM provider abstraction, Dynamo backend
integration)
- a revised PR sequence for the tail of the refactor
## What's being upstreamed
| Internal path | Size | Role | Upstream target |
|---|---|---|---|
| `server/main.py` | 94KB | FastAPI + WebSocket, session lifecycle, segment orchestration | `fastvideo/entrypoints/streaming/server.py` + handlers |
| `server/gpu_pool.py` | 66KB | GPU orchestration, subprocess workers | `fastvideo/entrypoints/streaming/gpu_pool.py` |
| `server/prompt_enhancer.py` | 69KB | LLM orchestration (cerebras_ifm, cerebras, groq) | `fastvideo/entrypoints/streaming/prompt/` package |
| `server/mock_server.py` | 45KB | Mock backend for dev/tests | `fastvideo/entrypoints/streaming/mock_server.py` |
| `server/prompt_safety.py` | 7KB | Optional fasttext-gated prompt safety | `fastvideo/entrypoints/streaming/prompt/safety.py` |
| `server/session_init_image.py` | 3KB | i2v init image handling | `fastvideo/entrypoints/streaming/session_init_image.py` |
| `server/rewrite_prompt_payload.py` | 3KB | Rewrite flow payload builder | `fastvideo/entrypoints/streaming/prompt/rewrite.py` |
| `server/session_logger.py` | 1KB | Session JSONL logs | `fastvideo/entrypoints/streaming/session_logger.py` |
| `server/config.py` | 9KB | Env-driven server config | Typed `ServeConfig` extensions |
| `router/main.py` | 27KB | Multi-replica load balancer + WS proxy | `fastvideo/entrypoints/streaming/router/` (or separate package) |
| `slurm/` | — | Deployment scripts | Likely stays internal |
## FastVideo contact surface today
Direct calls from the internal stack into FastVideo, all in `gpu_pool.py`:
| Location | Call | Notes |
|---|---|---|
| `gpu_pool.py:164` | `from fastvideo.entrypoints.video_generator import VideoGenerator` | Subprocess-level import, post-`CUDA_VISIBLE_DEVICES` setup |
| `gpu_pool.py:230` | `PipelineConfig.from_pretrained(config_model_path)` | Direct access to legacy `PipelineConfig` |
| `gpu_pool.py:231` | `pipeline_config.dit_config.quant_config = FP4Config()` | Direct internals mutation |
| `gpu_pool.py:264-267` | `VideoGenerator.from_pretrained(model_root, **load_kwargs)` | Flat legacy kwargs |
| `gpu_pool.py:837` | `generator.generate_video(**request_kwargs)` | Per-segment flat kwargs |
| `gpu_pool.py:282-288` | `LTX2AudioEncoder`, `AudioProcessor`, `get_diffusers_config` | Audio re-encode path |
`load_kwargs` at `gpu_pool.py:233-260` contains:
`ltx2_refine_enabled`, `ltx2_refine_upsampler_path`, `ltx2_refine_lora_path`,
`ltx2_refine_num_inference_steps`, `ltx2_refine_guidance_scale`,
`ltx2_refine_add_noise`, `pipeline_config`, `torch_compile_kwargs`,
`dit_cpu_offload`, `dit_layerwise_offload`, `vae_cpu_offload`,
`text_encoder_cpu_offload`, `pin_cpu_memory`, `ltx2_vae_tiling`,
`use_fsdp_inference`, `enable_torch_compile`.
`request_kwargs` at `gpu_pool.py:837` includes:
`ltx2_audio_clean_latent`, `ltx2_audio_denoise_mask`,
`ltx2_video_conditions`, `video_position_offset_sec`, standard sampling
fields.
**Implication**: upstreaming `gpu_pool.py` as-is perpetuates the flat
kwarg surface inside the public server. We need a typed translation
(PR 6 expansion) at the worker boundary before, or as part of, the
gpu_pool upstream.
## Session / continuation semantics today
Per-session state (in `server/main.py`):
- `locked_segment_prompts`, `curated_prompts`, `segment_idx`,
`generated_segment_count`, `loop_iteration`
Per-**GPU** (not per-session) continuation cache (in `gpu_pool.py`):
- `ltx2_continuation_images` — last 9 decoded frames for clip conditioning
- `ltx2_continuation_audio_latents` — denoised audio latents for audio conditioning
Segment N+1 automatically conditions on segment N's trailing frames and
audio. On session reset or handoff (`USER_JOIN`), the per-GPU cache is
cleared. There is currently **no way for a client to serialize and
resume continuation state elsewhere** — it lives on the GPU only.
## Design Decision 1: Continuation model
### Options
**A. Opaque client-round-trip payload** (current plan PR 7 design)
- Server returns `ContinuationState(kind, payload)`; client sends it back.
- Pro: stateless server, trivially load-balanceable, survives disconnects.
- Con: large payloads (frames + audio latents) over every request hop;
bandwidth heavy on multi-segment WebSocket sessions.
**B. Server-held session state** (internal reality)
- Continuation lives per-GPU; implicit between adjacent segments.
- Pro: zero client bandwidth; fast; matches today.
- Con: needs GPU affinity, no resume after disconnect, harder to scale horizontally.
**C. Hybrid** (recommended)
- Server-held is the default for streaming WebSocket sessions.
- Server exposes a `snapshot_state` message that returns the opaque
payload form for migration/retry.
- Stateless HTTP endpoints always use round-trip opaque payloads.
- One serialization format underlies both surfaces.
### Decision: **C (Hybrid)**
Rationale: matches both internal streaming use (server-held, fast) and
stateless API use (client-round-trip, resumable). Cost is one serialization
layer that serves both.
### Implications
- `ContinuationState.kind` identifies the payload schema
(e.g. `"ltx2.v1"`).
- `ContinuationState.payload` must cover:
- trailing conditioning frames (or a tensor reference)
- audio latents (or a tensor reference)
- segment index / rollout position
- any model-specific conditioning metadata (e.g. audio sample rate,
`video_position_offset_sec`)
- For large tensors, payload may reference a server-side blob by ID
rather than inline everything.
- Streaming server has a `SessionStore` keyed by session ID that holds
a typed `LTX2ContinuationState` object.
- `SessionStore.snapshot(session_id) -> ContinuationState` serializes
the current state for export.
- `SessionStore.hydrate(state: ContinuationState) -> session_id` loads
a state into a new session.
- Plan PR 7 expands to cover both surfaces and define the payload schema.
## Design Decision 2: Streaming server layout
### Options
- **A. `fastvideo/entrypoints/streaming/`** — parallel to
`fastvideo/entrypoints/openai/`
- **B. `fastvideo/entrypoints/server/{stateless,streaming}/`** — reorg both
- **C. `fastvideo/streaming/`** — top-level package, not under entrypoints
### Decision: **A (parallel subpackage)**
Rationale: lowest-friction, no existing code moves, both servers share
the same `fastvideo/entrypoints/*` namespace and import style. Shared
utilities can be factored into `fastvideo/entrypoints/server_common/`
later if needed. Option B creates churn across every openai/ import for
marginal organizational win.
### Target layout
```text
fastvideo/entrypoints/
├── openai/ # existing: stateless HTTP POST
│ ├── api_server.py
│ ├── video_api.py
│ ├── image_api.py
│ ├── common_api.py
│ ├── protocol.py
│ ├── state.py
│ ├── stores.py
│ └── utils.py
├── streaming/ # NEW: session WebSocket
│ ├── server.py # FastAPI + WebSocket entry
│ ├── session.py # session lifecycle, state machine
│ ├── session_store.py # typed session state + snapshot/hydrate
│ ├── protocol.py # JSON WebSocket message schemas
│ ├── stream.py # fMP4 encoding (av_fmp4 mode)
│ ├── gpu_pool.py # subprocess workers
│ ├── worker.py # per-GPU worker loop
│ ├── continuation.py # typed LTX2 state payload
│ ├── session_init_image.py
│ ├── session_logger.py
│ ├── mock_server.py
│ ├── prompt/
│ │ ├── enhancer.py # provider-agnostic prompt ops
│ │ ├── rewrite.py
│ │ ├── safety.py # optional fasttext
│ │ ├── payload.py # rewrite payload builder
│ │ └── providers/
│ │ ├── base.py # LLMProvider protocol
│ │ ├── cerebras.py
│ │ ├── cerebras_ifm.py
│ │ └── groq.py
│ └── router/ # or separate top-level package
│ ├── main.py
│ └── registry.py
├── cli/ # existing
└── video_generator.py # existing
```
### Config integration
`ServeConfig` gets an optional `streaming: StreamingConfig | None` field:
```python
@dataclass
class StreamingConfig:
session_timeout_seconds: int = 300
generation_segment_cap: int = 6
stream_mode: Literal["av_fmp4", "legacy_jpeg"] = "av_fmp4"
warmup: WarmupConfig = field(default_factory=WarmupConfig)
pool: GpuPoolConfig = field(default_factory=GpuPoolConfig)
prompt: PromptEnhancerConfig | None = None
safety: PromptSafetyConfig | None = None
@dataclass
class GpuPoolConfig:
num_workers: int | None = None # default: CUDA_VISIBLE_DEVICES count
enable_audio_reencode: bool = True
conditioning_num_frames: int = 9
conditioning_end_offset: int = 0
@dataclass
class PromptEnhancerConfig:
provider: Literal["cerebras_ifm", "cerebras", "groq"] = "cerebras_ifm"
model: str = "gpt-oss-120b"
timeout_ms: int = 20000
system_prompt_dir: str | None = None # hot-reloadable system prompts
@dataclass
class PromptSafetyConfig:
enabled: bool = False
classifier_path: str | None = None
```
## Design Decision 3: LLM provider abstraction
### Problem
`prompt_enhancer.py` (69KB) hard-codes three providers (cerebras_ifm,
cerebras, groq) with provider-specific request/response handling
scattered throughout. Upstreaming as-is locks FastVideo to those three
providers and couples the prompt operations to their response shapes.
### Shape
Introduce an `LLMProvider` protocol:
```python
from typing import Protocol, AsyncIterator, Literal
from dataclasses import dataclass
@dataclass
class LLMMessage:
role: Literal["system", "user", "assistant"]
content: str
@dataclass
class LLMRequest:
messages: list[LLMMessage]
model: str
max_tokens: int | None = None
temperature: float | None = None
timeout_ms: int | None = None
@dataclass
class LLMResponse:
content: str
provider: str
model: str
latency_ms: float
fallback_used: bool = False
class LLMProvider(Protocol):
name: str
async def complete(self, request: LLMRequest) -> LLMResponse: ...
```
### Decision: **Protocol + built-in implementations for cerebras, cerebras_ifm, groq**
Rationale: keeps the prompt enhancer free of provider-specific branching;
users (and future OpenAI/Anthropic/local additions) can register their
own provider without modifying FastVideo. Each built-in provider is
100-200 LOC; the enhancer becomes provider-agnostic prompt orchestration.
### Implications
- `prompt_enhancer.py` splits into `enhancer.py` (prompt operations) +
`providers/` (IO).
- Config moves from scattered env vars to typed `PromptEnhancerConfig`
under `ServeConfig.streaming.prompt`.
- Hot-reloadable system prompts stay — exposed as a management endpoint
on the streaming server.
- Fallback behavior (retry across providers in priority order) moves
into the enhancer layer, orthogonal to provider implementations.
## Design Decision 4 preamble: what Dynamo expects from FastVideo
Dynamo's backend pattern (observed in
`dynamo/components/src/dynamo/sglang/` and confirmed by PR #7544) is a
**pure Python import** pattern. Dynamo owns the backend subpackage in its
own repo; FastVideo only needs to expose a stable, typed, aggregated
and (later) streaming generation surface.
### Contract surface Dynamo consumes
| Surface | Shape | Notes |
|---|---|---|
| Constructor | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` | Already exists; `typed_kwargs` must be a stable subset from `GeneratorConfig` — no flat LTX2 legacy kwargs. |
| Sync execution | `generator.generate_video(request: GenerationRequest) -> VideoResult` | Aggregated mode; Dynamo wraps in `asyncio.to_thread` under an `asyncio.Lock`. |
| Async execution | `generator.generate_async(request: GenerationRequest) -> AsyncGenerator[VideoEvent, None]` | Needed for: (a) streaming server fMP4 chunks; (b) future Dynamo disaggregation. Events: `Progress`, `Partial?`, `Final`. |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` | Stable import path; Dynamo's adapter builds this from `NvCreateVideoRequest` + `VideoNvExt`. |
| Typed result | `VideoResult` with `video_bytes` or tensor frames, plus `ContinuationState?` | Must be picklable / JSON-serializable enough for Dynamo RPC. |
| Continuation | `ContinuationState(kind, payload)` with schema-versioned payloads | Used by FastVideo's session store today; tomorrow by Dynamo disaggregated workers. |
| Health check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` | Minimal 256x256 / 8 frames / 1 step; lets Dynamo's `FastVideoHealthCheckPayload.to_dict()` produce the Dynamo `health_check_payload` kwarg without knowledge of FastVideo internals. |
| Config dump | `GeneratorConfig.to_dict()` / `ServeConfig.to_dict()` | Dynamo calls `dynamo.common.config_dump.dump_config(path, config)` at worker start; we already have `config_to_dict()`. |
### Request/response mapping (Dynamo ↔ FastVideo)
Dynamo's video protocol (`NvCreateVideoRequest` / `NvVideosResponse`):
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> (seconds * nvext.fps) -> sampling.num_frames
input_reference -> input.image_path / input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
video_url (after upload) -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
```
All fields already exist (or will exist after PR 6 expansion) on
FastVideo's typed schema. No FastVideo changes required beyond what the
rest of this plan already covers **except**:
1. `generate_async` must exist (new in PR 7.10).
2. `default_health_check_request()` helper (new in PR 7.10).
3. The sync `generate_video(request=...)` path must be reachable without
extra wrapping (exists since PR 2; confirm stability).
### Where the Dynamo subpackage lives
The Dynamo-side integration (`FastVideoHandler`, `register_fastvideo_model`,
`FastVideoHealthCheckPayload`, args parsing, main.py, Dockerfile,
request/response mapping) lives **entirely in the Dynamo repo** at
`components/src/dynamo/fastvideo/`, matching the pattern used by vllm
and sglang. FastVideo does **not** host any Dynamo-related subpackage,
Dynamo dependency, or Dynamo-specific CLI. FastVideo's only obligation
is to expose a clean, stable, typed Python API that Dynamo's backend
package can import.
## Design Decision 4: Dynamo as first-class backend target
### Problem
PR #7544 (closed) shows two frictions with the pre-refactor API:
1. **Flat legacy kwargs** — the Dynamo handler had to know about
LTX2-specific flat names.
2. **Sync-only generation** — Dynamo's async handler wrapped
`generator.generate(...)` in `asyncio.to_thread` under a lock; no
progress streaming, no disaggregation path.
The refactor's stateless OpenAI server, WebSocket streaming server, and
Dynamo backend all want the same thing: **a typed async API that yields
progress events and a typed final result**. If we build it once in
`VideoGenerator`, all three adapters become thin.
### Options
**A. Keep sync-only, each adapter wraps**
- Simple; matches PR #7544.
- Con: streaming server needs its own async runner; Dynamo loses progress
streaming; no path to disaggregation.
**B. Add async event stream to `VideoGenerator`**
- `generate_async(request) -> AsyncGenerator[VideoEvent, None]`.
- Sync `generate_video` becomes a thin `asyncio.run` wrapper internally.
- Pro: one canonical execution API; streaming server, OpenAI server,
and Dynamo all consume events directly.
- Con: larger delta in `VideoGenerator` — must thread async through the
pipeline step loop.
**C. Queue-based `generate(request, event_cb)` callback**
- Middle ground; callback receives events.
- Pro: no async rewrite needed.
- Con: callers have to invert control; awkward for Dynamo's async
handler.
### Decision: **B (async event stream)**
Rationale: one substrate serves all three consumers. The cost is a
`generate_async` implementation that runs the pipeline step loop in a
thread and bridges events back via an asyncio queue — standard pattern,
limited surface area.
### Implications
- New PR 7.10 adds `generate_async` on `VideoGenerator` with three event
types: `VideoProgressEvent(step, total_steps, stage)`,
`VideoPartialEvent(frames_ndarray, index)` (optional; emitted only in
the streaming path), `VideoFinalEvent(video_bytes_or_tensor, metadata,
continuation_state?)`.
- Sync `generate_video(request=...)` becomes `asyncio.run(...)` over
`generate_async`, collecting events and returning the final.
- Streaming server's fMP4 encoder consumes `VideoPartialEvent` frames
directly, never re-decoding through disk.
- Dynamo adapter consumes `generate_async` and yields one
`NvVideosResponse` per `VideoFinalEvent` (aggregated mode; ignores
intermediate events today; can surface progress via Dynamo's
status/progress fields in the future).
- `ContinuationState` can be attached to `VideoFinalEvent.metadata`,
giving Dynamo a first-class way to surface state for disaggregation
later.
- Stable public exports: `from fastvideo import VideoGenerator`;
`from fastvideo.api import GenerationRequest, SamplingConfig,
ContinuationState, VideoResult, VideoEvent`.
- No Dynamo subpackage, dep, or CLI lives in FastVideo. The adapter
(`NvCreateVideoRequest ↔ GenerationRequest` mapping, handler,
registration) lives entirely in the Dynamo repo at
`components/src/dynamo/fastvideo/`.
### Constraints this adds to earlier PRs
- **PR 6** (typed LTX2 kwargs): every flat kwarg must have a typed home
**reachable from `GeneratorConfig`**, so Dynamo can construct the
generator without importing internal compat paths.
- **PR 7** (continuation state): `ContinuationState.payload` must be
JSON/YAML serializable (no raw torch tensors inline; use blob
indirection) so it survives Dynamo RPC transport.
- **PR 7.5** (streaming skeleton): consume `generate_async` rather than
re-implementing a progress loop around `generate_video`.
- **PR 2/3/4 already landed**: the typed request shape is fixed and
matches Dynamo's mapping needs — no backtracking required.
## Revised PR sequence (PR 5 onwards)
PRs 0-4 are unchanged and already landed. PR 5 is narrowed; PRs 5.5-7.9
are new inserts; PRs 8-13 are reshaped or kept.
| # | Title | Change | Key deliverables |
|---|---|---|---|
| **5** | Stateless `ServeConfig.default_request` merge | **Narrowed.** Wire typed default-request into `fastvideo/entrypoints/openai/`. | `_merge_default_request` helper, validated-against-preset, tests for default+user-override precedence |
| **5.5** | Server architecture split | **NEW.** Introduce `fastvideo/entrypoints/streaming/` subpackage skeleton. No behavior change. | Empty subpackage + stub server.py; CLI subcommand `fastvideo streaming-serve` (raises NotImplementedError); doc on layout |
| **6** | LTX2 public preset + stage overrides + config colocation | **Expanded.** Also add typed replacements for every flat kwarg used by internal `gpu_pool.py`. | `ltx2_two_stage` preset, `LTX2RefineStageOverride`, `CompileConfig` field types, typed `FP4Config` integration, colocation |
| **7** | Continuation state (public + session) | **Expanded.** Define both opaque payload AND server-held session store. | `ContinuationState.payload` schema, `LTX2ContinuationState` typed subclass, `SessionStore` interface, snapshot/hydrate APIs |
| **7.5** | Streaming server skeleton | **NEW.** Minimum viable WebSocket server: session lifecycle, JSON messages, fMP4 output, single-generator. | `server.py`, `session.py`, `protocol.py`, `stream.py` (fMP4), typed `StreamingConfig` |
| **7.6** | GPU pool upstream | **NEW.** Upstream `gpu_pool.py` with typed config boundary. | `gpu_pool.py`, `worker.py`, job queue, session-to-GPU binding, session timeout handling |
| **7.7** | Prompt enhancer upstream | **NEW.** Upstream `prompt_enhancer.py` with `LLMProvider` abstraction. | `prompt/enhancer.py`, `prompt/providers/{base,cerebras,cerebras_ifm,groq}.py`, hot-reloadable system prompts |
| **7.8** | Streaming auxiliaries | **NEW.** Small, isolated. | `prompt/safety.py`, `session_init_image.py`, `prompt/rewrite.py`, `session_logger.py`, `mock_server.py` |
| **7.9** | Router upstream | **NEW.** Multi-replica load balancer + WS proxy. | `streaming/router/` (or separate top-level package), health checks, WS proxy |
| **7.10** | Dynamo backend contract | **NEW.** Add `VideoGenerator.generate_async` event stream + `default_health_check_request()` helper. FastVideo exposes the async API only; the Dynamo backend package (handler, adapter, registration) lives entirely in the Dynamo repo at `components/src/dynamo/fastvideo/`. Streaming server (PR 7.5) and Dynamo backend both consume the same async API. | `generate_async` with `VideoProgressEvent`/`VideoPartialEvent`/`VideoFinalEvent`; sync `generate_video` becomes a thin wrapper; contract tests against a mock Dynamo-style handler that imports only public FastVideo APIs |
| **8** | Internal-UI ↔ public-server contract docs & tests | **Reframed.** Was "Dreamverse Server Adaptation Layer." Also covers Dynamo integration reference. | WebSocket protocol reference, contract tests, migration examples, Dynamo adapter example that upstream PR can copy verbatim |
| **9** | LongCat preset migration + colocation | **Keep.** | Stage overrides, colocation |
| **10** | Hunyuan15 SR preset migration + colocation | **Keep.** | Stage overrides, SR field migration POC, colocation |
| **11** | SSIM / perf test migration | **Keep.** Now blocked on PR 6 expansion. | Typed API migration of golden tests |
| **12** | Docs + examples | **Keep, expand.** | Streaming server docs now part of scope |
| **13** | Deprecation + cleanup | **Keep, expand.** | Also deprecate flat kwargs that internal gpu_pool uses today |
Total PR count: 13 → ~20 (13 original + 5 streaming-upstream inserts +
1 architecture split + 1 Dynamo contract). Each new PR is small and
self-contained because the streaming components are already cleanly
separated in the internal repo, and the Dynamo contract rides on top of
the async API that the streaming server already needs.
## Open questions
1. **Router: in-repo or separate package?** — It's orthogonal to inference;
in-repo couples deploy cycles, separate leaves FastVideo cleaner.
Recommendation: separate package `fastvideo-router/` or
`fastvideo/contrib/router/`; defer final call to PR 7.9.
2. **Session ID authority** — internal uses ad-hoc client IDs.
Recommendation: server-generated UUID, accept externally provided
session ID only for resume flows.
3. **Torch compile kwargs typing** — `CompileConfig.kwargs: dict[str, Any]`
today accepts `mode`, `backend`, `fullgraph`, `dynamic`. Options: keep
as opaque dict; fully type; hybrid (type the common four + allow
extras). Recommendation: hybrid, type common fields.
4. **Prompt safety / fasttext dependency** — heavy for users who don't
need it. Recommendation: ship as optional extra
`pip install fastvideo[prompt-safety]`.
5. **Audio-specific tensor payloads** — `ltx2_audio_clean_latent`,
`ltx2_audio_denoise_mask`, `ltx2_audio_latents` are not in the current
public schema. PR 7 should classify them (probably as opaque fields
inside `LTX2ContinuationState.payload`, not top-level sampling fields).
6. **Batching behavior** — internal `test_batching.py` suggests batching
is exercised. Scope this into PR 7.5 or defer to a post-cleanup perf PR?
7. ~~**Dynamo subpackage home**~~ — **Resolved.** No Dynamo code lives
in FastVideo. The full backend package (handler, adapter,
registration, health check) is owned by the Dynamo repo at
`components/src/dynamo/fastvideo/`, same pattern as vllm/sglang.
FastVideo only guarantees the public API contract listed above.
8. **Disaggregation readiness** — PR #7544 is aggregated-only. Our
`ContinuationState` hybrid already supports a future prefill/decode
split (prefill yields state; decode hydrates it). Should PR 7.10
explicitly validate that `ContinuationState` survives round-trip
through a Dynamo-style RPC (pickle or JSON), even though Dynamo
isn't using it today? Recommendation: yes; cheap contract test that
prevents drift.
9. **Dynamo progress/status passthrough** — `NvVideosResponse` has
`status` and `progress` fields. Should PR 7.10's handler contract
emit intermediate `NvVideosResponse` chunks keyed off
`VideoProgressEvent`, or stay aggregated-final-only to match PR
#7544? Recommendation: stay aggregated-final for PR 7.10; revisit
after Dynamo clarifies their streaming/progress semantics.
## Immediate path forward
1. Land `will/api_5` cleanup commits — **done** (`e03ca7d9`, `41f93179`
force-pushed without Claude co-author).
2. Review this plan with a human — commit the doc to capture the state.
3. Execute PR 5 (narrow stateless merge) and PR 5.5 (subpackage split)
in parallel. Both small; both unblock the streaming upstream that
follows.
4. Start PR 6 expansion (typed replacements for flat LTX2 kwargs) as the
critical path for PR 7.6 (gpu_pool upstream).
@@ -0,0 +1,93 @@
# Exploration Log: Video Generator Config API Design
## Status: draft
## Context
FastVideo's Python inference API currently mixes generator-instance settings,
pipeline initialization settings, and per-request sampling/runtime settings
through broad `**kwargs` surfaces on `VideoGenerator.from_pretrained(...)` and
`VideoGenerator.generate_video(...)`.
This exploration compares the current FastVideo design with
`sglang/multimodal_gen` and examines how to upstream multi-stage LTX2 /
Dreamverse behavior without growing more ad hoc top-level flags.
## Progress
- [x] Read FastVideo onboarding, codebase map, and relevant design docs.
- [x] Inspect current FastVideo generator, args, sampling, registry, and
workflow abstractions.
- [x] Inspect internal LTX2 streaming server usage and current two-stage /
continuation requirements.
- [x] Inspect SGL diffusion generator, server args, sampling params, and
request preparation boundary.
- [x] Inspect vLLM-Omni stage config, stage metadata, request, and orchestration
surfaces for multi-stage pipeline ideas.
- [x] Inspect current FastVideo CLI/config-file loading and compare with the
training YAML-only entrypoint.
- [ ] Convert findings into a concrete implementation plan for FastVideo.
## Findings
- FastVideo already has the right internal separation points:
`FastVideoArgs`, `PipelineConfig`, `SamplingParam`, and `ForwardBatch`.
- The public boundary is the unstable part:
init-time and request-time knobs are mixed through `**kwargs`.
- Unknown init keys can be silently filtered, while unknown request keys can be
only logged rather than rejected. This makes API drift hard to detect.
- SGL's split is cleaner:
`ServerArgs` for engine/runtime, `PipelineConfig` for model-family wiring,
and `SamplingParams` for per-request settings.
- SGL also has better merge semantics for user request overrides:
it preserves model defaults, tracks explicitly provided fields, and validates
request params against pipeline task type.
- SGL still has a design smell worth avoiding in FastVideo:
`SamplingParams._adjust(...)` depends on `ServerArgs`, which leaks
engine/pipeline concerns back into the request object.
- vLLM-Omni contributes a useful extra abstraction beyond SGL:
model-owned multi-stage topology via `ModelPipeline` and `StageConfig`,
with per-stage defaults (`default_sampling_params`) and runtime override
layering.
- vLLM-Omni's best reusable idea for FastVideo is not the serving stack, but
the separation between:
1. model-defined stage topology and per-stage defaults,
2. runtime engine overrides,
3. request-time sampling/state handoff.
- vLLM-Omni also shows the downside of exposing stage-indexed request lists too
directly: `sampling_params_list` works for a serving engine, but is too
positional and low-level for FastVideo's higher-level Python API.
- FastVideo already supports YAML/JSON config files for inference CLI, but the
current mechanism flattens nested documents back into argparse flags. This
preserves backward compatibility but keeps the CLI surface as the canonical
schema instead of a typed document model.
- The training stack has a cleaner precedent: a YAML-first config loaded into a
typed schema, with dotted CLI overrides applied onto the nested document
before parsing. Inference can likely adopt a lighter variant of that pattern.
- Multi-stage generation should be unified at the orchestration layer, not by
forcing LongCat refine, Hunyuan SR, and LTX2 continuation into one leaf config.
## Mistakes / Dead Ends
- A fully free-form string-dict API would lose too much type safety and would
likely recreate the current drift problem under a different shape.
- A single universal `RefineConfig` for all models would become a sparse bag of
nullable fields and would not map cleanly to existing model families.
## Proposed Standardization
- Introduce a typed public split:
`GeneratorConfig` for instance-lifetime engine/init settings and
`GenerationRequest` for per-call inputs/sampling/output.
- Allow dict input only as an interchange layer that is parsed immediately into
typed configs with strict unknown-key validation.
- Add a typed `GenerationPlan` / multi-stage orchestration layer with
discriminated stage configs:
`SampleStageConfig`, `LongCatRefineStageConfig`,
`HunyuanSRStageConfig`, `LTX2ContinuationStageConfig`.
- Let model families own stage defaults and stage topology through named
profiles or model-defined stage plans, similar in spirit to vLLM-Omni's
pipeline YAMLs, but expose them through typed Python config objects rather
than raw stage-indexed lists in the primary API.
- Make YAML/JSON a first-class serialization of the same typed inference
schema, not just a file format that expands into CLI flags.
- Prefer a YAML-first CLI pattern for nested configs:
`fastvideo generate --config run.yaml --request.sampling.seed 42`,
while keeping a compatibility layer for existing flat flags during migration.
- Upstream LTX2 two-stage / continuation behavior as a first-class stage or
pipeline profile rather than more `ltx2_*` top-level kwargs.
@@ -0,0 +1,194 @@
# Current State — 2026-05-05 (strategy reversal — single mega-PR #1288)
Point-in-time snapshot of branches, commits, and live infrastructure.
Update whenever commits land or services restart.
For HOW to commit / push / verify see [runbook.md](runbook.md). For
roster of co-authors to credit on every commit see
[authors.md](authors.md).
## Branch tips
| Repo | Branch | Tip | Distance |
|---|---|---|---|
| FastVideo | `will/ltx2_sr_port` (**PR #1288 head**) | `b36bdbc9` | 36 commits ahead of `origin/main`; OPEN, MERGEABLE; +integration-review.md (Part 1 drift + Part 2 tradeoffs + recommended Option D) |
| FastVideo | `will/api_7.10` | `6ae7a99f` | **deprecated** — PR #1287 closed in favor of #1288. Branch can be deleted on origin and locally; kept for now as historical reference. |
| FastVideo | `will/api_8`, `will/ltx2_sr_runtime`, `will/ltx2_nvfp4`, `will/ltx2_post_fixes`, `will/agents_cleanup` | (various) | **deprecated** split bookmarks. Strategy reversed to single mega-PR (D-17). Safe to delete locally; not pushed to origin. |
| FastVideo | `will/ltx2_sr_port-pre-1286-rebase` | `1baa60bb` | **local-only safety backup** of pre-rebase chain (37 commits); keep until next slice merges |
| Dreamverse | `will/integrate-public-fastvideo` | `ec8ef92` | 10 commits ahead of `737f3c1` (the dep switch) |
| FastVideo-internal | their `main` | (read-only ref) | — |
FastVideo worktree default branch is `will/ltx2_sr_port`. Other agents
share this worktree — if `git branch --show-current` shows something
else, switch back cleanly with `git checkout will/ltx2_sr_port` (don't
disturb their uncommitted work). I observed this happen repeatedly in
the 2026-05-05 session — confirmed harmless; switching back was always
safe with a clean working tree.
## Post-#1286 rebase summary
PR #1286 merged at `2aaeee2a` (squash). `will/ltx2_sr_port` was rebased
onto new `origin/main`, dropping 4 commits whose content is now in main:
- `cd76cf51` `[feat] streaming: router (multi-replica load balancer)`
- `1ac1e732` `[feat] streaming: fastvideo router-serve CLI`
- `b0b7f59c` `[test] streaming: router registry + health loop ...`
- `40e265b8` `[fix] streaming: router polish — bridge cancel + state
machine + deps` (squashed into `2aaeee2a` via cherry-pick `a152cb77`)
Rebase was clean — no conflicts. All 33 surviving commits got new SHAs
(rebase rewrites). The pre-rebase tip `1baa60bb` is preserved on the
local backup branch `will/ltx2_sr_port-pre-1286-rebase`.
## New linearized chain (33 commits, slice indices for STACK.md)
| Slice | PR | Commits | Tip SHA | Subject |
|---|---|---|---|---|
| 1-3 | 7.10 (PR #1287) | 3 | `6ae7a99f` | `[test] streaming: generate_async coverage + refreshed streaming test` |
| 4-6 | 8 | 3 | `f32e31ec` | `[test] streaming: contract tests for Dreamverse + Dynamo shapes` |
| 7-15 | LTX-2 SR | 9 | `e7297519` | `feat(ltx2): full i2v conditioning + continuation latent port` |
| 16-21 | NVFP4 | 6 | `6793166b` | `test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow` |
| 22-23 | LTX-2 post-fixes | 2 | `25897b67` | `[fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep` |
| 24-33 | agents_cleanup | 10 | `b34d9704` | `[docs] dreamverse-integration: add runbook + fresh-context onboarding` |
5 PRs landed (7.5, 7.6, 7.7, 7.8, 7.9), 1 in flight (7.10), 5 remaining
(8 / LTX-2 SR / NVFP4 / post-fixes / agents_cleanup).
## Historical commit chain analysis (pre-#1286 rebase)
The layered chain analysis below documented the pre-rebase SHAs (LTX-2
SR layer, NVFP4 layer, post-handoff fixes layer). Those SHAs no longer
exist on `will/ltx2_sr_port` — they live only on
`will/ltx2_sr_port-pre-1286-rebase`. Content semantics are unchanged;
SHAs were rewritten by the rebase. Kept here for narrative continuity.
## FastVideo: commit chain `cfccd292..156103b9`
Three layers since LTX-2 i2v port:
### Layer 1 — LTX-2 SR port + alignment harness (5 commits)
```
365a66c7 feat(quantization): upstream LTX-2 FP4Config with lazy flashinfer
433d26b2 feat(ltx2): port LTX-2 SR runtime — upsampler, refine stages, refine args
751d05de feat(ltx2): wire SR pipeline graph + port denoising/latent-prep stages
af6bbfea test(ltx2-sr): add numerical alignment harness — public vs internal
974cd430 fix(ltx2-sr): close port gaps surfaced by alignment harness retries
b6ac7630 test(ltx2-sr): pin ltx2 sampling knobs in harness for parity diff
b043d550 fix(api): align public SamplingParam ltx2 defaults with distilled
663dda80 fix(registry): order LTX-2 detectors so distilled wins for distilled paths
cfccd292 feat(ltx2): full i2v conditioning + continuation latent port (BASE)
```
(Predates the May 2 handoff.)
### Layer 2 — NVFP4 wire-up + per-component compile (6 commits, May 2 handoff)
```
a4760bae fix(api): propagate generic refine_* args + match internal randn
221cb20a feat(api): typed per-component CompileConfig + FastVideoArgs carriers
6da342ba feat(compile): per-component compile + transformer_refine + prepare hook
42b30bf9 feat(ltx2): wire FP4 inference through fastvideo.layers.quantization
94c983a2 refactor(quant): rename FP4 → NVFP4 to disambiguate from other FP4 variants
c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow
```
See [quantization.md](quantization.md) for what each commit locks in.
### Layer 3 — Post-handoff parity/perf fixes (3 commits, since May 2)
```
a5fcd19c [fix]: lazy-import flash_attn 2 fallback in attention backend
d4ee5be2 [fix]: avoid model.to() round-trip in Gemma encoder forward
156103b9 [fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep (HEAD)
```
Three small fixes — no new features. Continued parity tightening with internal.
## Dreamverse: commit chain `737f3c1..ec8ef92`
```
737f3c1 chore: switch fastvideo dep from FastVideo-internal to public FastVideo
4cc6b30 chore: gitignore Playwright + Next.js build artifacts under apps/web
33caa92 test(e2e): align Playwright specs with the actual production composer
6fd137c test(e2e): tighten frontend-shell + preset specs to match actual UI
248060b test(e2e): add Playwright tier with backend-health smoke + preset run
d80c2a8 refactor(server): drive FP4 + per-component compile via typed GeneratorConfig
3d7fd89 feat(skill): launch-demo orchestrator + fastvideo serve YAML
72f69b9 Update ffmpeg installation instructions.
1ba5635 fix(server): block startup on GPU warmup readiness, propagate failures
ec8ef92 fix(server): detect worker death in _send_command via proc.sentinel (HEAD)
```
The post-handoff trio (`72f69b9`, `1ba5635`, `ec8ef92`) hardens server
startup robustness — ffmpeg install docs, GPU warmup readiness gate, and
worker-death detection.
## Live services (do not duplicate)
| Port | Service | PID | Status |
|---|---|---|---|
| 8009 | `dreamverse-server` | 2453227 | `/readyz` returns 200, 1 warmed GPU worker, queue 0 |
| 5274 | `next-server` (dev) | 2399103 | 200, ~13.6 KB shell |
| 8000 | unknown FastAPI | — | **Not in handoff.** Probably stray `fastvideo serve`. Verify with `lsof -i :8000` before launching a new BE on the default port. |
## Stashes — DO NOT POP
| Repo | Stash | Reason |
|---|---|---|
| FastVideo | `stash@{0}: WIP on main: 71bfc13d HunyuanVideo plugin` | Pre-existing, unrelated to integration work |
| Dreamverse | `stash@{0}: wip: server modular refactor (split config/prompting/runtime/session)` | 3867-line orphan modular split, **not part of `will/integrate-public-fastvideo`**. Recover on a separate branch if needed. |
## Test status (from May 2 handoff, not re-verified post-Layer-3)
| Suite | Status |
|---|---|
| FastVideo `fastvideo/tests/api/` + `contract/` + `nvfp4_*` + `ltx2_pipeline_smoke` | 222 passed, 1 skipped |
| Playwright e2e against live BE+FE | 8 passed (5 backend-health + 2 frontend-shell + 1 preset-prompt-generation) |
| `fastvideo serve --config streaming_demo.yaml` validation | parses cleanly; dotted overrides work |
| `bash -n` on launch-demo skill scripts | clean across all 4 |
The 3 post-handoff commits are small parity fixes; full re-verification is
recommended but not required to read this state.
## Pre-existing failures (NOT caused by this work)
| Test | Failure | Notes |
|---|---|---|
| `fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype` | `AssertionError not raised` | Pre-existing on `main`. Verified via `git stash` that NVFP4 work doesn't introduce it. See [open-threads.md](open-threads.md) item #2. |
## Source docs (archived 2026-05-03)
The 7 source docs that this memory dir consolidates have been moved into
[`source-archive/`](source-archive/) — see the
[archive README](source-archive/README.md) for the archive policy and
synthesis mapping.
Other untracked items at the FastVideo repo root:
- Nested clones: `dynamo/`, `ray/`, `vllm-omni/`
- Lock files: `uv.lock`, `fastvideo/tests/ssim/.reference_videos_download.lock`
- Skill dirs: `.agents/skills/diagnose-ssim-failure/`, `.agents/skills/review-pr-link/`
- `.agents/exploration/pr-link-review.md` (kept; already promoted to a skill)
## Quick orientation commands
```bash
# FastVideo state
cd /home/william5lin/FastVideo
git log --oneline cfccd292..HEAD # 14 commits this round
# Dreamverse state
cd /home/william5lin/Dreamverse
git log --oneline 737f3c1..HEAD # 10 commits this round
# Live stack health (already running)
curl -s http://localhost:8009/readyz | head -c 300
curl -s http://localhost:5274/ -o /dev/null -w "%{http_code}\n"
# Re-verify test suite
.venv/bin/python -m pytest fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header
```
@@ -0,0 +1,389 @@
# Streaming Server Upstream — PRs 5.5 → 7.10
The `FastVideo-internal/ui/ltx2-streaming/server/` stack is being
upstreamed into public FastVideo at `fastvideo/entrypoints/streaming/`.
In parallel, FastVideo is becoming a first-class Dynamo backend (same
tier as vllm, sglang, trtllm). This file covers both threads since they
share `generate_async` as the substrate.
For PR sequence/status see [pr-roadmap.md](pr-roadmap.md). For the
Dreamverse-side adoption see [cross-repo-surfaces.md](cross-repo-surfaces.md).
**Last updated:** 2026-05-03.
## What's being upstreamed
| Internal path | Size | Role | Public target |
|---|---|---|---|
| `server/main.py` | 94 KB | FastAPI + WebSocket, session lifecycle, segment orchestration | `fastvideo/entrypoints/streaming/server.py` + handlers |
| `server/gpu_pool.py` | 66 KB | GPU orchestration, subprocess workers | `fastvideo/entrypoints/streaming/gpu_pool.py` |
| `server/prompt_enhancer.py` | 69 KB | LLM orchestration (cerebras_ifm, cerebras, groq) | `fastvideo/entrypoints/streaming/prompt/` package |
| `server/mock_server.py` | 45 KB | Mock backend for dev/tests | `fastvideo/entrypoints/streaming/mock_server.py` |
| `server/prompt_safety.py` | 7 KB | Optional fasttext-gated prompt safety | `prompt/safety.py` |
| `server/session_init_image.py` | 3 KB | i2v init image handling | `streaming/session_init_image.py` (PR 7.5, already public) |
| `server/rewrite_prompt_payload.py` | 3 KB | Rewrite flow payload builder | `prompt/rewrite.py` |
| `server/session_logger.py` | 1 KB | Session JSONL logs | `streaming/session_logger.py` |
| `server/config.py` | 9 KB | Env-driven server config | typed `ServeConfig.streaming` extensions |
| `router/main.py` | 27 KB | Multi-replica load balancer + WS proxy | `fastvideo/entrypoints/streaming/router/` |
| `slurm/` | — | Deployment scripts | Stays internal |
Frontend clients (`client/`, `prod-ui/`) stay in the internal repo.
## Four design decisions that shape the upstream
### D-1: Continuation model — Hybrid (server-held + opaque client-round-trip)
Streaming WebSocket sessions hold continuation per-GPU (matches today's
internal behavior, fast, zero client bandwidth). Stateless HTTP endpoints
use opaque round-trip payloads. Server exposes a `snapshot_state` message
that returns the opaque form for migration/retry.
One serialization layer underlies both surfaces.
Implementation: `SessionStore` (in-memory default, pluggable for
redis/etc.) keyed by session ID, holds typed `LTX2ContinuationState`.
- `snapshot(session_id) -> ContinuationState` exports for migration
- `hydrate(state: ContinuationState) -> session_id` loads state into new session
Payload schema covers: trailing conditioning frames (or tensor-blob ID),
audio latents (or blob ID), segment index, audio sample rate,
`video_position_offset_sec`, model-specific metadata.
Landed in PR 7. See [cross-repo-surfaces.md](cross-repo-surfaces.md) for
the full wire format.
### D-2: Streaming server layout — Parallel subpackage `fastvideo/entrypoints/streaming/`
Sits next to `fastvideo/entrypoints/openai/`. No existing code moves.
Both servers share the `entrypoints/*` namespace. Shared utilities can be
factored into `fastvideo/entrypoints/server_common/` later if needed.
### D-3: LLM provider abstraction — `LLMProvider` protocol + built-in providers
`prompt_enhancer.py` (69 KB) hard-coded three providers (cerebras_ifm,
cerebras, groq) with provider-specific request/response handling
scattered throughout. Upstreaming as-is would lock FastVideo to those
providers.
Protocol shape:
```python
@dataclass
class LLMRequest:
messages: list[LLMMessage]
model: str
max_tokens: int | None = None
temperature: float | None = None
timeout_ms: int | None = None
@dataclass
class LLMResponse:
content: str
provider: str
model: str
latency_ms: float
fallback_used: bool = False
class LLMProvider(Protocol):
name: str
async def complete(self, request: LLMRequest) -> LLMResponse: ...
```
PR 7.7 ships built-in providers for cerebras, groq. **Public Literal
currently restricts to `Literal["cerebras", "groq"]`** — `cerebras_ifm`
is internal-only and remains environment-driven on `dreamverse-server`.
See [open-threads.md](open-threads.md) follow-up #3.
Hot-reloadable system prompts via management endpoint. Sequential
fallback across providers in priority order — race-based fallback (the
internal optimization) deferred per [decisions-log.md](decisions-log.md)
D-3.
### D-4: Dynamo as first-class backend target — async event stream
PR ai-dynamo/dynamo#7544 (closed draft) showed two frictions:
1. Flat legacy kwargs — Dynamo handler had to know LTX-2-specific names.
2. Sync-only generation — Dynamo wrapped `generator.generate(...)` in
`asyncio.to_thread` under a lock; no progress streaming, no
disaggregation path.
Decision: **add `generate_async`** as the canonical execution API.
```python
async def generate_async(
self,
request: GenerationRequest,
) -> AsyncGenerator[VideoEvent, None]: ...
```
Events:
```python
@dataclass
class VideoProgressEvent:
step: int
total_steps: int
stage: str # "denoise" | "refine" | "decode" | ...
@dataclass
class VideoPartialEvent:
frames: np.ndarray # (num_frames, H, W, 3)
index: int # monotonic chunk index
@dataclass
class VideoFinalEvent:
video_bytes: bytes | None
tensor: torch.Tensor | None
metadata: dict[str, Any]
continuation_state: ContinuationState | None
```
The sync `generate_video(request=...) -> VideoResult` becomes a thin
`asyncio.run` wrapper over `generate_async` that collects events and
returns the final.
**Three consumers, one substrate:**
| Consumer | Transport | Request shape | State |
|---|---|---|---|
| Stateless OpenAI (`fastvideo/entrypoints/openai/`) | HTTP POST | `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; opaque payload |
| Streaming WebSocket (`fastvideo/entrypoints/streaming/`) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment, session-scoped | Server-held; per-GPU continuation cache |
| Dynamo native backend (`ai-dynamo/dynamo/components/src/dynamo/fastvideo/`) | Dynamo RPC endpoint | `NvCreateVideoRequest` ↔ adapter ↔ `GenerationRequest` | Aggregated today; future disaggregated via `ContinuationState` |
**FastVideo does NOT host any Dynamo code.** The full backend package
(`args.py`, `main.py`, `backend.py`, `register.py`, `health_check.py`)
lives entirely in the Dynamo repo at `components/src/dynamo/fastvideo/`,
matching the vllm/sglang pattern. FastVideo's only obligation is the
stable public Python API.
PR 7.10 lands the FastVideo-side contract. Dynamo backend code lives in
ai-dynamo/dynamo (next iteration of #7544 reopens against PR 8 reference
docs).
## Target package layout
```
fastvideo/entrypoints/
├── openai/ # existing: stateless HTTP POST
├── streaming/ # NEW: session WebSocket
│ ├── server.py # FastAPI + WebSocket entry
│ ├── session.py # session lifecycle, state machine
│ ├── session_store.py # typed session state + snapshot/hydrate
│ ├── protocol.py # JSON WebSocket message schemas
│ ├── stream.py # fMP4 encoding (av_fmp4 mode)
│ ├── gpu_pool.py # subprocess workers (PR 7.6)
│ ├── worker.py # per-GPU worker loop
│ ├── continuation.py # typed LTX2 state payload
│ ├── session_init_image.py
│ ├── session_logger.py
│ ├── mock_server.py
│ ├── prompt/
│ │ ├── enhancer.py # provider-agnostic prompt ops
│ │ ├── rewrite.py
│ │ ├── safety.py # optional fasttext
│ │ └── providers/
│ │ ├── base.py # LLMProvider protocol
│ │ ├── cerebras.py
│ │ ├── cerebras_ifm.py
│ │ └── groq.py
│ └── router/
│ ├── main.py
│ └── registry.py
├── cli/
└── video_generator.py
```
## Typed config integration
`ServeConfig` gets an optional `streaming: StreamingConfig | None`:
```python
@dataclass
class StreamingConfig:
session_timeout_seconds: int = 300
generation_segment_cap: int = 6
stream_mode: Literal["av_fmp4", "legacy_jpeg"] = "av_fmp4"
warmup: WarmupConfig = field(default_factory=WarmupConfig)
pool: GpuPoolConfig = field(default_factory=GpuPoolConfig)
prompt: PromptEnhancerConfig | None = None
safety: PromptSafetyConfig | None = None
@dataclass
class GpuPoolConfig:
num_workers: int | None = None # default: CUDA_VISIBLE_DEVICES count
enable_audio_reencode: bool = True
conditioning_num_frames: int = 9
conditioning_end_offset: int = 0
@dataclass
class PromptEnhancerConfig:
provider: Literal["cerebras", "groq"] = "cerebras" # cerebras_ifm pending
model: str = "gpt-oss-120b"
timeout_ms: int = 20000
system_prompt_dir: str | None = None # hot-reloadable
@dataclass
class PromptSafetyConfig:
enabled: bool = False
classifier_path: str | None = None
```
## `build_app` route contract — open follow-up
Today `fastvideo.entrypoints.streaming.server.build_app` exposes only:
- `GET /health`
- `WS /v1/stream`
The Dreamverse Next.js shell expects these additional routes that the
upstream plan (and Dreamverse FE today) require:
| Route | Owner per upstream plan | Status |
|---|---|---|
| `GET /healthz` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /readyz` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /status` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /curated-presets` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse, FE feature-detects |
| `POST /curated-presets/append` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse |
| `GET /prompt-system-config` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse |
| Devtools routes | Dreamverse-only | 🟡 stays in Dreamverse |
Until the three health routes migrate into FastVideo's `build_app`, the
`BE_FLAVOR=fastvideo` flavor of `launch_demo.sh` is a "diagnostic" flavor
only (verifies typed serve-config path) — not FE-compatible. See
[open-threads.md](open-threads.md) follow-up #1.
The streaming-upstream plan listed `/healthz`, `/readyz`, `/status`,
`/ws` as the contract that the upstream of `realtime/` → `streaming/`
must preserve. They were deferred from PR 7.5's MVP.
## PR 7.5 status — open as #1251
Single-generator WebSocket end-to-end shipped (8 commits):
1. `feat(streaming): protocol schemas + session state machine`
2. `feat(streaming): fMP4 encoder + session init-image persistence`
3. `feat(streaming): single-generator WebSocket server entry`
4. `test(streaming): server lifecycle + protocol + fMP4 coverage`
5. `docs(streaming): server contract spec`
6. `fix(streaming): restore missing-streaming-block guard + retire stub-era test`
7. `simplify(streaming): review follow-ups (idle timeout via asyncio.wait_for, _send_error helper, _cleanup_session, Protocol-typed generator, cleanup-on-disconnect)`
8. `fix(streaming): enforce idle timeout on receive_json + flag generator-cancellation gap (TODO → PR 7.10)`
Deferred TODOs (intentionally) blocking on PR 7.10:
- **Per-step progress events** — only terminal `step_complete` today;
needs `generate_async` for per-step `VideoProgressEvent` emission.
- **Mid-segment cancellation on client disconnect** — TODO marker in
`server.py` near `pool.run`. Needs `generate_async`'s cancellation
propagation.
## PR 7.6 status — branch ready, not yet PR'd
`will/api_7.6` (5 commits, rebased on 7.5):
1. `feat [7.6/n]: GPU pool manager with typed worker boundary`
2. `refactor [7.6/n]: route streaming server through GpuPool`
3. `test [7.6/n]: GPU pool coverage (in-process + subprocess)`
4. `fix(streaming): restore missing asyncio import in server` (rebase fixup)
5. `feat(streaming): extract worker.py and add two-segment warmup`
Tests: 17/17 gpu_pool tests + 89/89 streaming tests green.
Ships:
- `GpuPool` ABC + `InProcessGpuPool` + `SubprocessGpuPool` +
`PoolAssignment` / `PoolHealth` / `PoolAcquireTimeout`
- `worker.py` — per-GPU `worker_main` and two-segment warmup helper
- Subprocess startup uses typed `GeneratorConfig`, NOT flat kwargs
- Session-to-GPU binding with timeout + queue for contention
- Two-segment startup warmup per worker (segment 1 fresh + segment 2
with returned `ContinuationState` so both compile branches are primed)
- `SessionStore` (from PR 7) wired for per-GPU continuation cache
Deferred to PR 7.10:
- **Audio re-encode (`LTX2AudioEncoder`, `AudioProcessor`)**: internal
`_re_encode_audio` runs *inside* the per-step streaming loop
(`_stream_av_fmp4_events` / `do_step_ltx2`). The whole-segment
`pool.run()` path PR 7.6 ships doesn't need it. Re-encode is a
per-step streaming concern that belongs with `generate_async`.
- **Deprecate `VideoGenerator.from_pretrained(**flat_kwargs)`**: belongs
with PR 13 cleanup.
## PR 7.10 — the unlock PR
PR 7.10 adds `generate_async` and closes three open threads
simultaneously:
- Q-5 / D-5: audio re-encode for cross-segment continuity
- Q-9: Dynamo progress passthrough
- PR 7.5's mid-segment cancellation TODO (client disconnect →
`asyncio.CancelledError` → GPU work stops)
Plus health-check helper:
```python
def default_health_check_request(self) -> GenerationRequest: ...
# Returns 256x256, 8 frames, 1 step. Lets Dynamo's
# FastVideoHealthCheckPayload.to_dict() produce a Dynamo
# health_check_payload kwarg without knowledge of FastVideo internals.
```
Stable public exports:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GenerationRequest, SamplingConfig, ContinuationState,
VideoResult, VideoEvent,
VideoProgressEvent, VideoPartialEvent, VideoFinalEvent,
)
```
Streaming server (PR 7.5) gets rewired to consume `generate_async`
directly — no wrapper duplication.
## Dynamo request/response mapping
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> seconds * nvext.fps -> sampling.num_frames
input_reference -> input.image_path | input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled at adapter's output stage)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (response_format=b64_json)
uploaded URL -> data[0].url (response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> (reserved for future disaggregation)
```
## Open questions
1. **Router placement** — in-tree at `fastvideo/entrypoints/streaming/router/`
(current implementation per PR 7.9) or separate package
`fastvideo-router/` / `fastvideo/contrib/router/`. Effectively
resolved in-tree by the PR 7.9 implementation.
2. **Session ID authority** — server-generated UUID; accept externally
provided session ID only for resume flows.
3. **Disaggregation readiness contract test** — should PR 7.10 validate
`ContinuationState` survives round-trip through Dynamo-style RPC
(pickle or JSON), even though Dynamo isn't using it today?
Recommended: yes; cheap regression guard.
4. **Dynamo progress/status passthrough** — should PR 7.10's handler
contract emit intermediate `NvVideosResponse` chunks keyed off
`VideoProgressEvent`, or stay aggregated-final-only? Recommended:
stay aggregated-final for PR 7.10; revisit after Dynamo clarifies.
5. **`video_position_offset_sec` semantics** — see [decisions-log.md](decisions-log.md)
open question; needs decision before PR 7.6 emits state.
6. **`SessionStore` / `BlobStore` lifecycle** — eviction, TTL, blob-drop
on state replacement; defer to PR 7.5 design pass.
+1
View File
@@ -2,3 +2,4 @@
{"name": "evaluation-registry", "description": "Catalog of all evaluation metrics with detailed explanations, implementation status, and usage guides", "path": "evaluation-registry/README.md", "status": "draft", "trust": "medium"}
{"name": "experiment-journal", "description": "Living log of all experiments with hypotheses, configs, metrics, and insights", "path": "experiment-journal/README.md", "status": "draft", "trust": "medium"}
{"name": "related-work", "description": "Index of related papers, repos, and blog posts with structured comparisons to FastVideo", "path": "related-work/README.md", "status": "draft", "trust": "low"}
{"name": "dreamverse-integration", "description": "Consolidated knowledge base for the FastVideo public API refactor (PRs 0-17), LTX-2 streaming server upstream, Dreamverse migration from FastVideo-internal, and NVFP4 quantization landing", "path": "dreamverse-integration/README.md", "status": "ready", "trust": "high"}
@@ -1,94 +0,0 @@
---
name: index-related-work
description: Ingest a paper or repository into the related work index
---
# Index Related Work
## Purpose
Create a structured summary of a related paper, repository, or blog post and
add it to `.agents/memory/related-work/` for future reference. This builds the
agent's knowledge base for making informed decisions about training, evaluation,
and architecture choices.
## Prerequisites
- Access to the paper/repo (URL, PDF, or local clone).
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `source` | Yes | URL, citation, or local path |
| `type` | Yes | `paper`, `repo`, or `blog` |
| `tags` | No | List of tags (default: inferred from content) |
## Steps
### 1. Extract key information
For **papers**: Read abstract, method section, experimental setup, and results.
For **repos**: Read README, key source files, and training scripts.
For **blogs**: Read the full post.
Focus on:
- What problem does it solve?
- What architecture/technique is used?
- How does it relate to FastVideo's approach?
### 2. Create the index entry
Write to `.agents/memory/related-work/<slug>.md`:
```markdown
---
title: <title>
source: <URL or citation>
type: paper | repo | blog
date_indexed: <ISO-8601>
tags: [world-model, distillation, evaluation, ...]
---
## Summary
<1-2 paragraph summary.>
## Key Differences from FastVideo
- <comparison points>
## Actionable Insights
- <what we could adopt or adapt>
```
### 3. Update the catalog
If `.agents/memory/related-work/_catalog.md` exists, append the new entry.
If not, create it:
```markdown
# Related Work Catalog
| Slug | Title | Type | Tags | Date |
|------|-------|------|------|------|
| <slug> | <title> | <type> | <tags> | <date> |
```
## Outputs
- New file in `.agents/memory/related-work/<slug>.md`.
- Updated catalog.
## Example Usage
```
Index the Self-Forcing paper:
source: https://arxiv.org/abs/2406.xxxxx
type: paper
tags: [world-model, self-forcing, distillation]
```
## References
- `.agents/memory/related-work/README.md` — schema documentation
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
-2
View File
@@ -3,7 +3,5 @@
{"name": "summarize-run", "description": "Extract a W&B run summary into a structured experiment report", "path": "summarize-run/SKILL.md", "status": "draft", "trust": "low"}
{"name": "log-experiment", "description": "Append or update an experiment entry in the experiment journal", "path": "log-experiment/SKILL.md", "status": "draft", "trust": "low"}
{"name": "evaluate-video-quality", "description": "Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)", "path": "evaluate-video-quality/SKILL.md", "status": "draft", "trust": "low"}
{"name": "index-related-work", "description": "Ingest a paper or repository into the related work index", "path": "index-related-work/SKILL.md", "status": "draft", "trust": "low"}
{"name": "search-related-work", "description": "Query the related work index for relevant papers, repos, or comparisons", "path": "search-related-work/SKILL.md", "status": "draft", "trust": "low"}
{"name": "seed-ssim-references", "description": "Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline", "path": "seed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
{"name": "reseed-ssim-references", "description": "Re-seed (overwrite) HF reference videos for an existing fastvideo/tests/ssim/ test and a single model id on Modal L40S. Always backs up current refs first, regenerates on Modal, pauses for the user to eyeball before-vs-after, then uploads with --force scoped to --model-id. Sister skill to seed-ssim-references; use when intentional code change has invalidated existing refs", "path": "reseed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
@@ -1,82 +0,0 @@
---
name: search-related-work
description: Query the related work index for relevant papers, repos, or comparisons
---
# Search Related Work
## Purpose
Search through `.agents/memory/related-work/` to find indexed papers, repos,
or blog posts relevant to a query. Use this when you need to understand how
other work compares to FastVideo's approach, or when looking for techniques
to adopt.
## Prerequisites
- The related work index has entries (`.agents/memory/related-work/*.md`).
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `query` | Yes | Natural language query |
| `tags` | No | Filter by tags (e.g., `[distillation, evaluation]`) |
| `type` | No | Filter by type (`paper`, `repo`, `blog`) |
## Steps
### 1. Search the index
Use grep-based search through `.agents/memory/related-work/`:
```bash
# Search by content
grep -rl "<query>" .agents/memory/related-work/
# Search by tags (in frontmatter)
grep -l "tags:.*<tag>" .agents/memory/related-work/*.md
```
### 2. Rank results
For each matching file:
1. Read the file.
2. Score relevance to the query based on:
- Title match
- Tag match
- Content match (summary, differences, insights)
3. Return top results.
### 3. Format output
```markdown
## Related Work Search: "<query>"
### 1. <Title> (relevance: high)
- **Source**: <URL>
- **Tags**: <tags>
- **Key insight**: <most relevant excerpt>
- **File**: `.agents/memory/related-work/<slug>.md`
### 2. <Title> (relevance: medium)
...
```
## Outputs
- Ranked list of relevant related work entries with excerpts.
## Example Usage
```
Search for work related to video quality evaluation metrics:
query: "video generation quality evaluation metrics"
tags: [evaluation]
```
## References
- `.agents/memory/related-work/README.md` — index schema
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
-67
View File
@@ -1,67 +0,0 @@
---
description: Synchronize the STATUS.md dashboard by scanning .agents/ directories
---
# Sync Dashboard
Updates `.agents/STATUS.md` by scanning the skills, workflows, memory, lessons,
and exploration directories to reflect what actually exists on disk.
## When to Use
- After adding, removing, or renaming any file in `.agents/`.
- Periodically (e.g., at end of each conversation session).
- When the dashboard feels out of date.
## Steps
### 1. Scan directories
List all files in each directory:
```bash
echo "=== Skills ==="
ls -1 .agents/skills/*.md 2>/dev/null | grep -v SKILL_TEMPLATE
echo "=== Workflows ==="
ls -1 .agents/workflows/*.md 2>/dev/null
echo "=== Memory ==="
ls -1 .agents/memory/*.md 2>/dev/null
ls -1 .agents/memory/related-work/*.md 2>/dev/null | grep -v README
echo "=== Lessons ==="
ls -1 .agents/lessons/*.md 2>/dev/null | grep -v README
echo "=== Exploration ==="
ls -1 .agents/exploration/*.md 2>/dev/null | grep -v README
```
### 2. Compare with STATUS.md
For each file found:
- If it's in STATUS.md → leave it (preserve status/trust/tested fields).
- If it's NOT in STATUS.md → add it with status `🔴 Stub`, trust `None`, tested `❌`.
For each entry in STATUS.md:
- If the file no longer exists → mark it as `❌ Removed` or delete the row.
### 3. Update counts
Recalculate the summary table at the top:
- Count files per category.
- Count by status (Ready, Draft, Stub).
### 4. Update timestamp
Set `_Last synced: <current date>_` at the top of STATUS.md.
### 5. Review
Read through the updated STATUS.md for accuracy. Flag anything that looks wrong.
## Notes
- Do NOT change trust levels during sync — those are set manually after testing.
- Do NOT change status during sync — status changes require actual validation.
- This workflow only handles structural sync (file existence), not content review.
+182
View File
@@ -0,0 +1,182 @@
# `.agents/` Cleanup Log — Phase 1 (Deletes Only)
**Status:** TEMPORARY — delete this file after the cleanup is reviewed/committed.
**Date:** 2026-05-04
**Branch:** `will/ltx2_sr_port`
**Scope:** Phase 1 of the `.agents/` cleanup plan (deletes only; no rewrites or additions).
For the full multi-phase plan, see the prior session analysis. This file tracks
exactly what got deleted, why, and what cross-references still point at deleted
content (to fix in a future phase).
---
## Deletions executed
### Files deleted
| Path | Size | Reason |
|---|---|---|
| `.agents/STATUS.md` | 3.85 KB | Stale dashboard, last synced 2026-03-02. Counts wrong (claimed 8 skills/4 workflows/4 memory; actual 9/5/5). References old snake_case filenames (`codebase_map.md`/`experiment_journal.md`) that don't exist. Hand-maintained derivative of `.agents/{memory,skills}/index.jsonl` — strictly redundant. |
| `.agents/exploration/pr-link-review.md` | 1.11 KB | Status: "promoted" to `.agents/skills/review-pr-link/`. Per `.agents/exploration/README.md` lifecycle, promoted exploration logs should not linger after the skill exists. |
| `.agents/workflows/sync-dashboard.md` | 1.87 KB | SOP for maintaining `STATUS.md` (which is also deleted). Contained obsolete file paths (`.agents/skills/launch-experiment.md` flat layout vs. actual `<skill>/SKILL.md` per-dir layout). Has never been run successfully (judging by stale dates everywhere). |
### Skill directories deleted
| Path | Size | Reason |
|---|---|---|
| `.agents/skills/index-related-work/` | 2.18 KB | Vapor-skill operating on the empty `.agents/memory/related-work/` registry. Never used (the registry has zero entries despite ~6 weeks since skill creation). Re-add when the related-work catalog gains entries. |
| `.agents/skills/search-related-work/` | 1.91 KB | Same: vapor-skill against empty registry. The skill description literally requires "The related work index has entries" as a prerequisite, and there are none. |
**Total deleted: 5 items, ~10.9 KB.**
### Registry updates
| File | Change |
|---|---|
| `.agents/skills/index.jsonl` | Removed entries for `index-related-work` and `search-related-work`. Was 9 entries; now 7. |
### Symlink hygiene
`.agents/scripts/sync-skills.sh` was run to prune now-stale symlinks under
`.claude/skills/` that pointed at the deleted skill directories. Output captured
in the run log.
---
## What was KEPT (despite being candidates)
| Path | Why kept |
|---|---|
| `.agents/scripts/sync-skills.sh` | User explicitly requested keep. **Verified**: this script is INDEPENDENT of STATUS.md / sync-dashboard.md. It mirrors `.agents/skills/` → `.claude/skills/` via symlinks for Claude Code skill discovery. Self-contained, useful, prunes its own stale symlinks. |
| `.agents/memory/related-work/README.md` | Empty placeholder, but the schema/template is reusable. Kept for when first related-work entry is added. |
| `.agents/memory/experiment-journal/README.md` | Same: empty placeholder with template; kept for when journaling begins. |
| `.agents/lessons/README.md` | Same: empty placeholder, reusable schema. |
| `.agents/exploration/README.md` | Active template for new exploration logs. Kept. |
---
## Remaining broken cross-references (FOLLOW-UP NEEDED)
These files still reference deleted content. **NOT fixed in Phase 1** — track for
the next pass (Phase 2: rewrites/dedupe).
### References to deleted `STATUS.md`
| Referencing file | Action needed |
|---|---|
| `.agents/onboarding/README.md` | Quick-reference tree (line ~65) lists `STATUS.md ← dashboard: completeness & trust of all components`. Remove that line + the `ONBOARDING.md` typo (file is `README.md`). |
### References to deleted `pr-link-review.md`
| Referencing file | Action needed |
|---|---|
| `.agents/memory/dreamverse-integration/state.md` | "Untracked but present" / "Source docs (archived)" sections still mention `pr-link-review.md` as kept. Update to reflect deletion. |
| `.agents/memory/dreamverse-integration/README.md` | Same — table row for `pr-link-review.md` says "kept in exploration dir". Update or remove the row. |
### References to deleted skills (`index-related-work`, `search-related-work`)
| Referencing file | Action needed |
|---|---|
| `.agents/memory/related-work/README.md` | Says "Use the `index-related-work` skill". Either remove that hint or note "skill removed; re-add when registry has entries". |
| `.agents/workflows/evaluation-development.md` | Step 1 says "Search `.agents/memory/related-work/` for existing evaluation approaches" — that's still valid (manual search). No change needed. |
### References to deleted `sync-dashboard.md`
| Referencing file | Action needed |
|---|---|
| `.agents/memory/evaluation-registry/README.md` | Doesn't reference sync-dashboard directly. No change. |
| `.agents/STATUS.md` | Already being deleted. |
---
## Other registry inconsistencies discovered (NOT FIXED in Phase 1)
While editing `.agents/skills/index.jsonl`, two skill directories were found
that exist on disk but **are not registered** in `index.jsonl`:
| Skill dir | Status | Why missing from index |
|---|---|---|
| `.agents/skills/diagnose-ssim-failure/` | Untracked locally; NOT on `origin/main`. 12.3 KB SKILL.md + `scripts/compare_latent_pt.py`. Recent mtime (2026-05-01). | Created in a prior session but the registration step was skipped. |
| `.agents/skills/review-pr-link/` | Untracked locally; NOT on `origin/main`. 2.9 KB SKILL.md + `scripts/prepare_pr_review.py` + `agents/openai.yaml`. The promotion target of the deleted `pr-link-review.md` exploration log. | Skipped registration when promoted from exploration log. |
Both skills are functional and exposed via `sync-skills.sh` symlinks (just verified in
`.claude/skills/`), but agents reading `index.jsonl` to discover skills will miss them.
**Action for Phase 2**: Add entries to `.agents/skills/index.jsonl` for both,
likely with `trust: medium` since they have working scripts and recent use.
---
## Skill registry parity check
After Phase 1, `.agents/skills/` contains 9 directories but `index.jsonl` lists 7:
| In `index.jsonl` | On disk |
|---|---|
| ✓ launch-experiment | ✓ launch-experiment/ |
| ✓ monitor-experiment | ✓ monitor-experiment/ |
| ✓ summarize-run | ✓ summarize-run/ |
| ✓ log-experiment | ✓ log-experiment/ |
| ✓ evaluate-video-quality | ✓ evaluate-video-quality/ |
| ✓ seed-ssim-references | ✓ seed-ssim-references/ |
| ✓ reseed-ssim-references | ✓ reseed-ssim-references/ |
| ❌ (missing) | ⚠ diagnose-ssim-failure/ |
| ❌ (missing) | ⚠ review-pr-link/ |
`.claude/skills/` symlinks (the runtime-discoverable surface) include all 9 ✓.
---
## Phase 2+ items (NOT executed in this session)
For future cleanup sessions, the prior plan identified:
**Phase 2 (rewrites)**:
- Rewrite `.agents/onboarding/worldmodel-training/README.md` to drop ~50% structural duplication with `codebase-map/README.md`
- Refresh `.agents/memory/codebase-map/README.md` (last updated 2026-03-08; missing `fastvideo/api/`, `fastvideo/entrypoints/streaming/`, etc.)
- Refresh `.agents/memory/evaluation-registry/README.md` (last updated 2026-03-02; references old `evaluation_registry.md` filename)
- Merge `.agents/workflows/experiment-journaling.md` into `experiment-lifecycle.md` (one SOP per workflow)
- Fix the broken cross-references listed above
**Phase 3 (additions)**:
- `fastvideo/api/AGENTS.md`
- `fastvideo/entrypoints/AGENTS.md`
- `tests/AGENTS.md` (top-level, distinct from `fastvideo/tests/AGENTS.md`)
- `fastvideo/distributed/AGENTS.md`
- `examples/AGENTS.md`
- `docs/AGENTS.md`
- `benchmarks/AGENTS.md`
**Phase 4 (registry)**:
- Add `.agents/workflows/index.jsonl`
- Standardize all three index.jsonl schemas
**Phase 5 (skills quality)**:
- Promote tested skills (`seed-ssim-references`, `reseed-ssim-references`, `diagnose-ssim-failure`, `review-pr-link`) from `trust: low` to `trust: medium`
- Mark untested skills (`launch-experiment`, `monitor-experiment`, `summarize-run`, `log-experiment`, `evaluate-video-quality`) explicitly with their gating prerequisite (e.g. "operates on empty registry")
---
## Recovery
All deletions are local (`will/ltx2_sr_port`, not committed). To restore any
deleted file:
```bash
git restore --source=HEAD .agents/STATUS.md
git restore --source=HEAD .agents/exploration/pr-link-review.md
git restore --source=HEAD .agents/workflows/sync-dashboard.md
git restore --source=HEAD .agents/skills/index-related-work/SKILL.md
git restore --source=HEAD .agents/skills/search-related-work/SKILL.md
```
---
## When to delete THIS file
Once:
1. The Phase 1 deletions are committed (or merged), AND
2. Phase 2 (broken cross-reference cleanup) is also committed,
remove this file. Its purpose is transient bookkeeping for a multi-phase cleanup.
+88
View File
@@ -0,0 +1,88 @@
# Co-Authors — `will/ltx2_sr_port` Stack
**Status:** PERMANENT — keep around as the source of truth for who collaborated on this work, even after the stack merges.
**Last updated:** 2026-05-04
This file documents the human co-authors credited on every commit in the
`will/ltx2_sr_port` stack and its 10 split PRs. The 4 collaborators below
worked on the FastVideo-internal precursor of this code (LTX-2 streaming
server, NVFP4 wire-up, GPU pool, prompt enhancer, etc.) and are credited as
co-authors on the public-side upstream commits via Git's standard
[`Co-authored-by`](https://docs.github.com/en/pull-requests/committing-changes-to-your-project/creating-and-editing-commits/creating-a-commit-with-multiple-authors)
trailer convention.
The trailers are added to every commit on `will/ltx2_sr_port` (see
[`STACK.md`](STACK.md)), which means GitHub will:
- Show the 4 co-authors on every commit detail page
- Show them on the merge commit / squash commit summary
- Display their avatars in the PR's "Contributors" sidebar
- Surface them in [`/contributors`](https://github.com/hao-ai-lab/FastVideo/contributors) once the stack lands
## Co-author roster
| GitHub user | Real name | GitHub ID | Trailer email |
|---|---|---|---|
| [`@Davids048`](https://github.com/Davids048) | Junda (David) Su | 90978028 | `90978028+Davids048@users.noreply.github.com` |
| [`@RandNMR73`](https://github.com/RandNMR73) | Matthew Noto | 99706358 | `99706358+RandNMR73@users.noreply.github.com` |
| [`@XOR-op`](https://github.com/XOR-op) | (unset) | 17672363 | `17672363+XOR-op@users.noreply.github.com` |
| [`@jzhang38`](https://github.com/jzhang38) | Zhang Peiyuan | 42993249 | `42993249+jzhang38@users.noreply.github.com` |
## Why no-reply emails
GitHub's `<id>+<username>@users.noreply.github.com` form is the most reliable
way to link a `Co-authored-by` trailer to a GitHub account. It:
- Always works regardless of whether the user has a public verified email
- Survives the user changing their primary email
- Doesn't expose anyone's personal email to git history
- Is the format GitHub itself produces when you click "Add co-author" in the
web UI
(All 4 collaborators have this email already used in `FastVideo-internal`
git history, verified via `git log --all` on that repo.)
## Trailer block (copy-paste ready)
The trailers added to every commit on `will/ltx2_sr_port`:
```
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
```
## How the trailers were applied
```bash
git rebase --exec '
git commit --amend --no-edit \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
' origin/main will/ltx2_sr_port
```
Git's `--trailer` flag is idempotent (it dedupes by the full `key: value`
string), so re-running the rebase is safe and won't add duplicates.
## How to add a new co-author later
1. Add the user to the roster table above.
2. Append their `Co-authored-by` line to the trailer block.
3. Re-run the rebase command above on `will/ltx2_sr_port` — git's
trailer dedupe handles the existing 4; the new one gets appended.
4. Re-slice all 10 split branches per [`STACK.md`](STACK.md).
5. Force-push `will/api_7.6`, `will/api_7.7`, and `will/ltx2_sr_port`.
## What we do NOT add
Per [`AGENTS.md`](AGENTS.md):
> Never add any coding agent or models such as Claude (or Claude Code), GPT,
> Codex or others as a co-author in commits or PRs.
So no `Co-authored-by: Claude <noreply@anthropic.com>` or similar. Only
human collaborators.
@@ -29,7 +29,15 @@ surfaces:
vae_cpu_offload: generator.engine.offload.vae
pin_cpu_memory: generator.engine.offload.pin_cpu_memory
enable_torch_compile: generator.engine.compile.enabled
enable_torch_compile_text_encoder: generator.engine.compile.text_encoder_enabled
enable_torch_compile_vae: generator.engine.compile.vae_enabled
enable_torch_compile_audio_vae: generator.engine.compile.audio_vae_enabled
torch_compile_kwargs: generator.engine.compile.backend,fullgraph,mode,dynamic,extras
torch_compile_kwargs_dit: generator.engine.compile.dit_kwargs
torch_compile_kwargs_text_encoder: generator.engine.compile.text_encoder_kwargs
torch_compile_kwargs_vae: generator.engine.compile.vae_kwargs
torch_compile_kwargs_audio_vae: generator.engine.compile.audio_vae_kwargs
transformer_quant: generator.engine.quantization.transformer_quant
disable_autocast: generator.engine.disable_autocast
enable_stage_verification: generator.engine.enable_stage_verification
prompt_txt: request.inputs.prompt_path
@@ -41,12 +49,25 @@ surfaces:
override_pipeline_cls_name: generator.pipeline.components.override_pipeline_cls_name
boundary_ratio: request.sampling.boundary_ratio
ltx2_vae_tiling: generator.pipeline.vae_tiling
refine_enabled: generator.pipeline.preset_overrides.refine.enabled
refine_upsampler_path: generator.pipeline.components.upsampler_weights
refine_lora_path: generator.pipeline.components.lora_path
refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
refine_guidance_scale: request.stage_overrides.refine.guidance_scale
refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
ltx2_refine_enabled: generator.pipeline.preset_overrides.refine.enabled
ltx2_refine_upsampler_path: generator.pipeline.components.upsampler_weights
ltx2_refine_lora_path: generator.pipeline.components.lora_path
ltx2_refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
ltx2_refine_guidance_scale: request.stage_overrides.refine.guidance_scale
ltx2_refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
preset_owned:
ltx2_vae_spatial_tile_size_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_size_in_pixels
ltx2_vae_spatial_tile_overlap_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_overlap_in_pixels
ltx2_vae_temporal_tile_size_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_size_in_frames
ltx2_vae_temporal_tile_overlap_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_overlap_in_frames
ltx2_initial_latent_path: request.extensions.ltx2.initial_latent_path
ltx2_audio_latent_path: request.extensions.ltx2.audio_latent_path
compatibility_only:
mode: "Legacy multi-mode FastVideoArgs switch; typed inference config should not expose execution mode."
inference_mode: "Legacy boolean mirror of mode; kept only through adapters while FastVideoArgs remains."
@@ -56,6 +77,12 @@ surfaces:
VSA_sparsity: "Model-specific inference optimization not yet represented in the typed public schema."
moba_config_path: "Model-specific MoBA optimization surface not yet represented in the typed public schema."
master_port: "Executor/bootstrap compatibility field; not part of the canonical inference schema."
refine_transformer_path: "Generic stage-2 refine transformer override; no typed equivalent yet."
refine_noise_path: "Generic stage-2 refine noise override; no typed equivalent yet."
refine_audio_noise_path: "Generic stage-2 refine audio noise override; no typed equivalent yet."
ltx2_refine_transformer_path: "LTX-2 refine transformer carrier; no typed equivalent yet."
ltx2_refine_noise_path: "LTX-2 refine noise carrier; no typed equivalent yet."
ltx2_refine_audio_noise_path: "LTX-2 refine audio noise carrier; no typed equivalent yet."
private_only:
ray_placement_group: "Ray deployment-only field."
ray_runtime_env: "Ray deployment-only field."
@@ -404,6 +431,11 @@ surfaces:
ltx2_stg_scale_audio: request.extensions.ltx2.stg_scale_audio
ltx2_stg_blocks_video: request.extensions.ltx2.stg_blocks_video
ltx2_stg_blocks_audio: request.extensions.ltx2.stg_blocks_audio
ltx2_images: request.extensions.ltx2.images
ltx2_image_crf: request.stage_overrides.refine.image_crf
ltx2_conditioning_latent_stage1: request.extensions.ltx2.conditioning_latent_stage1
ltx2_conditioning_latent_stage2: request.extensions.ltx2.conditioning_latent_stage2
ltx2_video_conditions: request.extensions.ltx2.video_conditions
audio_start_in_s: request.extensions.stable_audio.audio_start_in_s
audio_end_in_s: request.extensions.stable_audio.audio_end_in_s
init_audio: request.extensions.stable_audio.init_audio
+354
View File
@@ -0,0 +1,354 @@
# Dynamo Native Backend Integration
FastVideo exposes a stable Python API that the
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo) project consumes
as a pure-Python import, same tier as `vllm`, `sglang`, `trtllm`.
**FastVideo hosts no Dynamo code.** The backend subpackage
(`components/src/dynamo/fastvideo/`) lives in the Dynamo repo. This doc
is the reference integrators copy when standing up that package — it
mirrors the structure used by `dynamo/components/src/dynamo/sglang/`
and is known to satisfy the (closed) draft
[ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
pattern.
## What FastVideo provides
The public surface Dynamo imports is intentionally small:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
# Post-PR 7.10:
VideoEvent, VideoProgressEvent, VideoPartialEvent, VideoFinalEvent,
VideoResult,
)
```
| Surface | Availability | Notes |
| --- | --- | --- |
| `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` | Today | `typed_kwargs` is a stable subset from `GeneratorConfig` — no flat legacy LTX-2 kwargs (guaranteed after PR 6) |
| `VideoGenerator.generate(request: GenerationRequest) -> GenerationResult` | Today | Aggregated; Dynamo wraps in `asyncio.to_thread` under `asyncio.Lock` |
| `VideoGenerator.generate_async(request) -> AsyncGenerator[VideoEvent, None]` | **PR 7.10** | Canonical execution substrate; sync wrapper reroutes through this |
| `VideoGenerator.default_health_check_request() -> GenerationRequest` | **PR 7.10** | 256x256 / 8 frames / 1 step; lets Dynamo build its health payload without knowing any FastVideo internals |
| `fastvideo.api.GenerationRequest` / `SamplingConfig` / `InputConfig` | Today | Stable public dataclasses |
| `fastvideo.api.ContinuationState` | Today (PR 7) | JSON-safe envelope; kind-versioned payloads |
| `fastvideo.api.VideoResult` | Today | `frames`, `video_path`, `state`, `metadata` |
| `config_to_dict(cfg)` | Today | Used by Dynamo's `dump_config(path, config)` |
## Backend package layout
Modeled on `components/src/dynamo/sglang/`:
```
components/src/dynamo/fastvideo/
├── __init__.py
├── __main__.py # Entry: python -m dynamo.fastvideo
├── main.py # worker() dispatch — mirrors sglang/main.py
├── args.py # FastVideoArgGroup — CLI → GeneratorConfig
├── backend_args.py # Dynamo runtime flags (namespace, fs_url, ...)
├── init_video_generation.py # init_video_generation(runtime, config)
├── register.py # register_video_generation_model() for Dynamo
├── backend.py # VideoGenerationWorkerHandler
├── health_check.py # FastVideoHealthCheckPayload
├── protocol.py # NvCreateVideoRequest ↔ GenerationRequest adapter
├── request_handlers/
│ └── video_generation/
│ └── video_generation_handler.py # async generate(req, ctx)
├── README.md
└── CLAUDE.md # per-backend guidance
```
None of these files live in FastVideo.
## Request/response mapping
Dynamo's `NvCreateVideoRequest` / `VideoNvExt` / `NvVideosResponse` map
one-to-one onto FastVideo's typed schema:
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> request.prompt
size="WxH" -> request.sampling.width, height
seconds -> seconds * nvext.fps -> request.sampling.num_frames
input_reference -> request.inputs.image_path / video_path
nvext.fps -> request.sampling.fps
nvext.num_frames -> request.sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> request.sampling.num_inference_steps
nvext.guidance_scale -> request.sampling.guidance_scale
nvext.seed -> request.sampling.seed
nvext.negative_prompt -> request.sampling.negative_prompt
nvext.continuation_state -> request.state (opaque ContinuationState)
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
uploaded URL -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> nvext.continuation_state (reserved for disagg)
```
## Example: aggregated handler (sync wrap)
Satisfies the PR #7544 shape; works today against
`VideoGenerator.generate`, upgrades cleanly to `generate_async` after
PR 7.10.
```python
# components/src/dynamo/fastvideo/request_handlers/video_generation/
# video_generation_handler.py
from __future__ import annotations
import asyncio
import base64
import time
from typing import Any, AsyncGenerator
from fastvideo import VideoGenerator
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
class VideoGenerationWorkerHandler:
def __init__(self, generator: VideoGenerator, config, fs=None):
self.generator = generator
self.config = config
self.fs = fs
self._lock = asyncio.Lock() # aggregated = one-in-flight
async def generate(
self,
request: dict[str, Any],
context,
) -> AsyncGenerator[dict[str, Any], None]:
req = _to_fastvideo_request(request)
t0 = time.perf_counter()
async with self._lock:
result = await asyncio.to_thread(self.generator.generate, req)
elapsed = time.perf_counter() - t0
video_bytes = _materialize(result, self.fs, request.get("response_format"))
yield {
"data": [video_bytes],
"inference_time_s": elapsed,
"model": request.get("model"),
}
def _to_fastvideo_request(request: dict[str, Any]) -> GenerationRequest:
nvext = request.get("nvext") or {}
fps = nvext.get("fps", 24)
num_frames = nvext.get("num_frames") or (request.get("seconds") or 4) * fps
width, height = _parse_size(request.get("size"))
return GenerationRequest(
prompt=request["prompt"],
negative_prompt=nvext.get("negative_prompt"),
inputs=InputConfig(
image_path=request.get("input_reference"),
),
sampling=SamplingConfig(
width=width, height=height,
num_frames=num_frames, fps=fps,
num_inference_steps=nvext.get("num_inference_steps", 50),
guidance_scale=nvext.get("guidance_scale", 1.0),
seed=nvext.get("seed", 1024),
),
output=OutputConfig(save_video=False, return_frames=False),
state=nvext.get("continuation_state"), # public ContinuationState
)
```
`_parse_size` and `_materialize` are small adapter helpers owned by the
Dynamo backend package; they never appear in FastVideo.
## Example: streaming handler (post-PR 7.10)
```python
async def generate(self, request, context):
req = _to_fastvideo_request(request)
async for event in self.generator.generate_async(req):
if event.__class__.__name__ == "VideoProgressEvent":
yield {"status": "generating", "progress": event.step / event.total_steps}
elif event.__class__.__name__ == "VideoFinalEvent":
yield {
"data": [{"b64_json": base64.b64encode(event.video_bytes).decode()}],
"inference_time_s": event.metadata.get("inference_time_s"),
"nvext": {"continuation_state": _serialize_state(event.continuation_state)},
}
```
Aggregated and streaming differ only in which events the handler
forwards; both share one `generate_async` substrate.
## Example: health check
```python
# components/src/dynamo/fastvideo/health_check.py
from dynamo.health_check import HealthCheckPayload
from fastvideo import VideoGenerator
class FastVideoHealthCheckPayload(HealthCheckPayload):
def __init__(self, generator: VideoGenerator) -> None:
# Post-PR 7.10: generator.default_health_check_request() returns a
# typed GenerationRequest; dump it into the same dict shape that
# FastVideo's adapter accepts.
req = generator.default_health_check_request()
self.default_payload = {
"prompt": req.prompt or "test",
"size": f"{req.sampling.width}x{req.sampling.height}",
"response_format": "b64_json",
"nvext": {
"fps": req.sampling.fps,
"num_frames": req.sampling.num_frames,
"num_inference_steps": req.sampling.num_inference_steps,
"guidance_scale": req.sampling.guidance_scale,
},
}
super().__init__()
```
Fallback (pre-PR 7.10) — hardcoded 256×256 / 8 frames / 1 step, matching
[`VideoGenerationHealthCheckPayload`](https://github.com/ai-dynamo/dynamo/blob/main/components/src/dynamo/sglang/health_check.py#L198-L226).
## Init function sketch
```python
# components/src/dynamo/fastvideo/init_video_generation.py
async def init_video_generation(runtime, config, shutdown_endpoints):
from fastvideo import VideoGenerator
from fastvideo.api import config_to_dict
server_args, dynamo_args = config.server_args, config.dynamo_args
generator = VideoGenerator.from_pretrained(**config.fastvideo_kwargs())
dump_config(dynamo_args.dump_config_to, config)
endpoint = runtime.endpoint(
f"{dynamo_args.namespace}.{dynamo_args.component}.{dynamo_args.endpoint}"
)
shutdown_endpoints[:] = [endpoint]
handler = VideoGenerationWorkerHandler(
generator, config, fs=get_fs(dynamo_args.media_output_fs_url)
)
payload = FastVideoHealthCheckPayload(generator).to_dict()
await asyncio.gather(
endpoint.serve_endpoint(
handler.generate,
graceful_shutdown=True,
health_check_payload=payload,
),
register_video_generation_model(
generator, endpoint, server_args,
),
)
```
## Args adapter
`FastVideoArgGroup` (Dynamo-side) converts CLI flags into a typed
`GeneratorConfig` — **never** into legacy flat kwargs. Because PR 6
added typed homes for every kwarg the internal `gpu_pool.py` used,
this adapter can build the config purely from the public typed schema:
```python
def build_generator_config(args) -> "GeneratorConfig":
from fastvideo.api import (
CompileConfig, ComponentConfig, EngineConfig, GeneratorConfig,
OffloadConfig, ParallelismConfig, PipelineSelection,
)
return GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=args.num_gpus,
parallelism=ParallelismConfig(tp_size=args.tp_size, sp_size=args.sp_size),
offload=OffloadConfig(dit=args.dit_offload, text_encoder=args.te_offload),
compile=CompileConfig(enabled=args.compile, mode=args.compile_mode),
),
pipeline=PipelineSelection(
workload_type=args.workload or "t2v",
preset=args.preset, # e.g. "ltx2_two_stage"
components=ComponentConfig(
upsampler_weights=args.refine_upsampler,
lora_path=args.refine_lora,
),
),
)
```
## Registration
Dynamo's Rust side skips HuggingFace `config.json` downloads for
`ModelType::Videos`, same fast path used by image diffusion. The
Python-side registration:
```python
# components/src/dynamo/fastvideo/register.py
from dynamo.llm import ModelDeploymentCard, ModelType, register_model
async def register_video_generation_model(generator, endpoint, server_args):
mdc = ModelDeploymentCard.with_name_only(server_args.model_name or server_args.model_path)
await register_model(endpoint, mdc, ModelType.Videos, readiness_gate=asyncio.Event())
```
## Contract guarantees
These guardrails let the Dynamo backend be written once and not
re-chase FastVideo drift:
1. `GenerationRequest` field paths are stable across PR 6 onward. Any
breaking rename triggers a major bump and appears in
[`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml).
2. `ContinuationState.payload` is JSON-serializable or references
opaque blob ids. Dynamo can round-trip it through RPC without
special-casing torch tensors.
3. `VideoGenerator.from_pretrained` accepts a typed `GeneratorConfig`;
legacy flat kwargs are compatibility-only and deprecate in PR 13.
4. `generate_async` (PR 7.10+) emits events in order
`Progress* → Partial* → Final`; the final event always has exactly
one occurrence per request.
5. `default_health_check_request()` (PR 7.10+) returns a request that
passes `parse_config` and produces a non-zero-latency but bounded
workload (256×256 / 8 frames / 1 step).
FastVideo's contract tests (`fastvideo/tests/contract/`) assert these
with mocked Dynamo-style handlers that import only the public surface.
If a change to FastVideo breaks the adapter pattern, those tests fail
at FastVideo's CI — before the Dynamo-side integration even knows.
## What the Dynamo adapter MUST NOT import
* Anything under `fastvideo.pipelines.*` directly (pipelines are
internal; presets identify them by name on
`PipelineSelection.preset`).
* `fastvideo.fastvideo_args.FastVideoArgs` (legacy compat type).
* `fastvideo.api.compat.*` private helpers
(`_validate_continuation_state` etc.) — the public boundary is
`VideoGenerator` + `fastvideo.api`.
* Any flat legacy LTX-2 kwarg (`ltx2_refine_upsampler_path`,
`torch_compile_kwargs`, etc.) — all have typed homes in
`GeneratorConfig`.
## Future: disaggregated prefill/decode
PR 7's continuation state was designed to survive RPC transport, so a
future Dynamo split where prefill yields state and decode hydrates it
is expressible without changing the contract. The streaming server
(PR 7.6) already uses the `SessionStore` pattern; Dynamo's disagg
could wire a distributed `SessionStore` backend by the same interface.
## See also
* [OpenAI HTTP contract](openai.md)
* [Streaming WebSocket protocol](streaming.md)
* Draft PR reference: [ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
* Dynamo SGLang backend (template this doc is modeled on):
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo/tree/main/components/src/dynamo/sglang)
+41
View File
@@ -0,0 +1,41 @@
# Server Contracts
FastVideo's typed public API (`fastvideo.api`) is consumed by three
server-class integrations that must share one execution substrate so we
don't grow three near-duplicate progress loops:
| Consumer | Transport | Request shape | State model |
| --- | --- | --- | --- |
| [Stateless OpenAI](openai.md) | HTTP POST `/v1/videos` | `VideoGenerationsRequest` → `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; optional `ContinuationState` round-trip |
| [Streaming WebSocket](streaming.md) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment | Server-held `SessionStore`, snapshot-on-demand |
| [Dynamo native backend](dynamo.md) | Dynamo RPC | `NvCreateVideoRequest` → adapter → `GenerationRequest` | Aggregated today; disaggregated via `ContinuationState` later |
All three consume the same underlying surface:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
ServeConfig,
)
# Sync today, async after PR 7.10 lands VideoGenerator.generate_async.
result = generator.generate(request)
```
These docs lock down the request/response shapes so drift between
FastVideo, the internal UI, and Dynamo can be caught at review time.
PR 8 does not ship runtime code; it ships the contract reference and
the contract tests that guard it.
## Related
- [API refactor design](../overview.md)
- Parity inventory: [`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml)
- [Streaming server upstream plan](../../../.agents/exploration/streaming-server-upstream-plan.md)
- Draft PR (closed) that establishes the Dynamo shape:
https://github.com/ai-dynamo/dynamo/pull/7544
+116
View File
@@ -0,0 +1,116 @@
# OpenAI-compatible HTTP Contract
The stateless FastVideo HTTP server lives at
[`fastvideo/entrypoints/openai/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/entrypoints/openai).
Launch: `fastvideo serve --config serve.yaml`.
## Endpoints
| Method | Path | Description |
| --- | --- | --- |
| `POST` | `/v1/videos/generations` | Synchronous video generation |
| `GET` | `/v1/videos` | List prior jobs held in the in-memory store |
| `GET` | `/v1/videos/{id}` | Job status / result |
| `GET` | `/v1/videos/{id}/content` | Download the MP4 once ready |
| `POST` | `/v1/images/generations` | Synchronous image generation |
| `GET` | `/v1/models` | Enumerate registered models |
| `GET` | `/health` | Liveness probe |
## `VideoGenerationsRequest` shape
Mirrors the OpenAI `POST /v1/videos/generations` shape:
```json
{
"prompt": "a fox running through snow",
"size": "1024x1536",
"seconds": 5,
"fps": 24,
"num_frames": 121,
"seed": 42,
"num_inference_steps": 8,
"guidance_scale": 1.0,
"negative_prompt": "blurry, low quality",
"input_reference": "/path/to/init.png"
}
```
SGLang-compatible extensions carried today:
`num_inference_steps`, `guidance_scale`, `guidance_scale_2`,
`true_cfg_scale`, `negative_prompt`, `enable_teacache`, `output_path`.
## Merge precedence
The server builds a `GenerationRequest` each call using three layers,
highest first:
1. **Request body (client-explicit)** — only fields carried in
`request.model_fields_set` (Pydantic v2). Unset fields do not count,
even if the Pydantic model has a schema default for them.
2. **`ServeConfig.default_request` (operator-explicit)** — projected via
[`explicit_request_updates()`](../../../fastvideo/api/compat.py);
only fields the operator actually wrote into the YAML count as
defaults. Every other field inherits the schema default rather than
being pinned.
3. **Hardcoded fallback** — e.g. `fps = 24`.
The gate matters: both surfaces carry schema defaults. Without
`model_fields_set` / explicit-path tracking, schema defaults would
masquerade as intent and silently shadow the other side.
See [`video_api.py::_build_generation_kwargs`](../../../fastvideo/entrypoints/openai/video_api.py)
for the canonical implementation; the per-request assembly lives there,
not in pipeline code.
## Continuation state
The stateless surface accepts an opaque `ContinuationState` round-trip.
Clients that want continuation pass the prior `state` blob back on the
next request, and receive a new one on the response when
`request.output.return_state = true`.
Shape:
```json
{
"state": {
"kind": "ltx2.v1",
"payload": { "schema_version": 1, "segment_index": 3, ... }
}
}
```
Payload is always JSON-serializable. Large tensors may live in an
opaque blob-store reference the client simply round-trips; see
[`LTX2ContinuationState`](../../../fastvideo/pipelines/basic/ltx2/continuation.py).
Continuation is not yet wired all the way through to
`generator.generate_video(...)` — PR 7.6 (GPU pool upstream) is the
pipeline-level consumer. PR 7 locked the envelope so this surface is
stable ahead of that plumbing.
## Error codes
| HTTP | Condition |
| --- | --- |
| `400 Bad Request` | Parse/validation failure (unknown field, type mismatch, incompatible preset/state) |
| `404 Not Found` | `GET /v1/videos/{id}` for an unknown job |
| `409 Conflict` | Job id already exists |
| `500 Internal Server Error` | Pipeline raised; body mirrors upstream OpenAI error envelope |
| `503 Service Unavailable` | No generator loaded, or shutdown in progress |
Errors include a JSON body with
`{"error": {"type": "...", "message": "..."}}` matching the OpenAI
Python SDK's expectation.
## What does not cross this boundary
* Flat legacy kwargs (`ltx2_refine_enabled`, `torch_compile_kwargs`,
etc.) — these are init-time, configured via `ServeConfig.generator`,
never per-request.
* Private Dreamverse-only fields — those live in a private adapter on
the Dreamverse side; the public FastVideo surface never promises
backward compatibility for them.
* Raw tensor payloads (`ltx2_audio_clean_latent` et al.) — these are
derived by the pipeline from `ContinuationState`, never shipped as
request fields.
+13 -1
View File
@@ -46,7 +46,14 @@ from fastvideo.api.parser import (
load_serve_config,
parse_config,
)
from fastvideo.api.results import GenerationResult
from fastvideo.api.results import (
GenerationResult,
VideoEvent,
VideoFinalEvent,
VideoPartialEvent,
VideoProgressEvent,
VideoResult,
)
from fastvideo.api.sampling_param import SamplingParam
__all__ = [
@@ -76,6 +83,11 @@ __all__ = [
"ServeConfig",
"ServerConfig",
"StreamingConfig",
"VideoEvent",
"VideoFinalEvent",
"VideoPartialEvent",
"VideoProgressEvent",
"VideoResult",
"WarmupConfig",
"InferencePreset",
"PresetStageSpec",
+33 -5
View File
@@ -120,6 +120,10 @@ def legacy_from_pretrained_to_config(
compile_config["enabled"] = value
elif key == "enable_torch_compile_text_encoder":
compile_config["text_encoder_enabled"] = value
elif key == "enable_torch_compile_vae":
compile_config["vae_enabled"] = value
elif key == "enable_torch_compile_audio_vae":
compile_config["audio_vae_enabled"] = value
elif key == "torch_compile_kwargs":
remaining: dict[str, Any] = (dict(deepcopy(value)) if isinstance(value, Mapping) else {})
for first_class in _COMPILE_TYPED_KEYS:
@@ -127,6 +131,14 @@ def legacy_from_pretrained_to_config(
compile_config[first_class] = remaining.pop(first_class)
if remaining:
compile_config["extras"] = remaining
elif key in {
"torch_compile_kwargs_dit",
"torch_compile_kwargs_text_encoder",
"torch_compile_kwargs_vae",
"torch_compile_kwargs_audio_vae",
}:
compile_config[key[len("torch_compile_kwargs_"):] +
"_kwargs"] = (dict(deepcopy(value)) if isinstance(value, Mapping) else {})
elif key == "ltx2_vae_tiling":
pipeline["vae_tiling"] = value
elif key == "config_model_path":
@@ -238,17 +250,33 @@ def generator_config_to_fastvideo_args(config: GeneratorConfig | Mapping[str, An
if normalized.pipeline.vae_tiling is not None:
kwargs["ltx2_vae_tiling"] = normalized.pipeline.vae_tiling
if engine.compile.text_encoder_enabled is not None:
# ``FastVideoArgs.from_kwargs`` filters to declared fields, so
# this is a no-op on the current legacy path. Emit anyway so the
# realtime runtime (PR 7.6) — which reads from the kwargs dict
# before FastVideoArgs filtering — can pick it up once wired.
kwargs["enable_torch_compile_text_encoder"] = (engine.compile.text_encoder_enabled)
if engine.compile.vae_enabled is not None:
kwargs["enable_torch_compile_vae"] = engine.compile.vae_enabled
if engine.compile.audio_vae_enabled is not None:
kwargs["enable_torch_compile_audio_vae"] = (engine.compile.audio_vae_enabled)
if engine.compile.dit_kwargs:
kwargs["torch_compile_kwargs_dit"] = deepcopy(engine.compile.dit_kwargs)
if engine.compile.text_encoder_kwargs:
kwargs["torch_compile_kwargs_text_encoder"] = deepcopy(engine.compile.text_encoder_kwargs)
if engine.compile.vae_kwargs:
kwargs["torch_compile_kwargs_vae"] = deepcopy(engine.compile.vae_kwargs)
if engine.compile.audio_vae_kwargs:
kwargs["torch_compile_kwargs_audio_vae"] = deepcopy(engine.compile.audio_vae_kwargs)
quantization = engine.quantization
if quantization is not None and quantization.text_encoder_quant is not None:
kwargs["override_text_encoder_quant"] = quantization.text_encoder_quant
if quantization is not None and quantization.transformer_quant is not None:
kwargs["transformer_quant"] = quantization.transformer_quant
# Resolve the typed quant name to a concrete ``QuantizationConfig``
# instance and pin it on ``dit_config.quant_config``. The legacy
# path expected callers to do this themselves via
# ``pipeline_config.dit_config.quant_config = NVFP4Config()``; the
# typed surface accepts a string and does the wiring here so
# downstream code can rely on a single source of truth.
from fastvideo.layers.quantization import get_quantization_config
_resolved_quant_cls = get_quantization_config(quantization.transformer_quant)
kwargs["transformer_quant"] = _resolved_quant_cls()
components = normalized.pipeline.components
if components.pipeline_config_path is not None:
+69 -1
View File
@@ -102,4 +102,72 @@ class GenerationResult:
return result
__all__ = ["GenerationResult"]
# Alias the canonical result type; matches the public docs.
VideoResult = GenerationResult
@dataclass
class VideoProgressEvent:
"""Per-step progress event emitted by :meth:`VideoGenerator.generate_async`.
Consumers treat these as best-effort telemetry; ``total_steps`` is
the count the pipeline reported at the start of the run, not a
rolling estimate.
"""
step: int
total_steps: int
stage: str = "denoise"
"""Logical stage name (``denoise`` | ``refine`` | ``decode`` | …)."""
@dataclass
class VideoPartialEvent:
"""Chunk of decoded frames ready for streaming.
Emitted only on the streaming path; the aggregated code path never
yields partials. ``frames`` is a numpy ``(N, H, W, 3)`` uint8
ndarray; ``index`` is a monotonic chunk index starting at 0.
"""
frames: Any
index: int
@dataclass
class VideoFinalEvent:
"""Terminal event carrying the generated video and metadata.
Exactly one ``VideoFinalEvent`` is emitted per request. When
``request.output.return_state`` is True the event also carries the
:class:`ContinuationState` the caller needs to resume.
"""
video_bytes: bytes | None = None
tensor: Any | None = None
frames: Any | None = None
metadata: dict[str, Any] = field(default_factory=dict)
continuation_state: ContinuationState | None = None
result: VideoResult | None = None
"""The full :class:`VideoResult` for callers that want everything.
Streaming consumers typically only care about ``frames`` /
``continuation_state``; keeping the full result here avoids a
second code path."""
VideoEvent = VideoProgressEvent | VideoPartialEvent | VideoFinalEvent
"""Union of every event :meth:`VideoGenerator.generate_async` yields.
Consumers match by ``isinstance`` rather than ``type`` so subclasses
(e.g. a future ``VideoAudioSegmentEvent``) slot in without breaking
existing code."""
__all__ = [
"GenerationResult",
"VideoEvent",
"VideoFinalEvent",
"VideoPartialEvent",
"VideoProgressEvent",
"VideoResult",
]
+26 -9
View File
@@ -98,20 +98,37 @@ class SamplingParam:
camera_rotation: str | None = None
# LTX-2 multi-modal CFG and STG.
# cfg_scale defaults are 1.0 (CFG off) so ``ForwardBatch.__post_init__``
# doesn't force ``do_classifier_free_guidance`` on non-LTX-2 models that
# never override these fields. LTX-2 presets that need text-CFG on set
# them in their ``defaults`` dict (e.g. ``ltx2_base``).
# Class-level defaults match the *distilled* LTX-2 schedule
# (mirrors ``FastVideo-internal/.../LTX2DistilledSamplingParam``):
# the distilled model expects neutral guidance scales — modality 1,
# rescale 0, STG 0 — and explicit-CFG callers (full LTX-2) opt back
# in by selecting the ``LTX2_BASE`` preset, which overrides these
# to mod=3.0 / rescale=0.7 / stg=1.0 in its ``defaults`` dict.
# cfg_scale defaults stay at 1.0 (CFG off) so
# ``ForwardBatch.__post_init__`` doesn't force CFG on non-LTX-2
# models that never override these fields.
ltx2_cfg_scale_video: float = 1.0
ltx2_cfg_scale_audio: float = 1.0
ltx2_modality_scale_video: float = 3.0
ltx2_modality_scale_audio: float = 3.0
ltx2_rescale_scale: float = 0.7
ltx2_stg_scale_video: float = 1.0
ltx2_stg_scale_audio: float = 1.0
ltx2_modality_scale_video: float = 1.0
ltx2_modality_scale_audio: float = 1.0
ltx2_rescale_scale: float = 0.0
ltx2_stg_scale_video: float = 0.0
ltx2_stg_scale_audio: float = 0.0
ltx2_stg_blocks_video: list[int] = field(default_factory=lambda: [29])
ltx2_stg_blocks_audio: list[int] = field(default_factory=lambda: [29])
# LTX-2 image / video / continuation conditioning. These flow from
# generate_video(...) kwargs through ``sampling_param.update(kwargs)``
# onto the ForwardBatch fields of the same name. ``ltx2_image_crf``
# gates the conditioning-image H.264 re-encode; the streaming
# session controller passes ``ltx2_image_crf=0.0`` because it
# conditions on already-decoded VAE-quality frames.
ltx2_images: list[tuple[str, int, float]] | None = None
ltx2_image_crf: float = 33.0
ltx2_conditioning_latent_stage1: Any | None = None
ltx2_conditioning_latent_stage2: Any | None = None
ltx2_video_conditions: list[tuple[list[str], int, float]] | None = None
# Stable Audio (T2A): clip start/end in seconds. Honored by
# `StableAudioConditioningStage` + `StableAudioDecodingStage`. Other
# families ignore them.
+17 -6
View File
@@ -38,21 +38,32 @@ class CompileConfig:
``backend``/``fullgraph``/``mode``/``dynamic`` are the four most
common ``torch.compile`` knobs. ``extras`` holds any remaining
``torch.compile`` kwargs (e.g. ``options``, ``disable``).
The ``enabled`` switch covers the DiT transformer path (including
``transformer_2`` and the LTX-2 stage-2 ``transformer_refine``).
Per-component flags below are independent overlays — set to ``True``
to compile that component, ``None`` to leave it eager. Each
``*_kwargs`` dict overrides the master ``backend``/``fullgraph``/
``mode``/``dynamic``/``extras`` for that component when non-empty;
leaving it empty inherits the master kwargs.
"""
enabled: bool = False
text_encoder_enabled: bool | None = None
"""Whether ``torch.compile`` is applied to the text encoder. ``None``
keeps the runtime default. The public ``FastVideoArgs`` adapter does
not yet consume this flag; reserved so the realtime runtime upstream
(PR 7.6) has a typed home for its ``enable_torch_compile_text_encoder``
kwarg without routing through ``pipeline.experimental``."""
backend: str | None = None
fullgraph: bool | None = None
mode: str | None = None
dynamic: bool | None = None
extras: dict[str, Any] = field(default_factory=dict)
text_encoder_enabled: bool | None = None
vae_enabled: bool | None = None
audio_vae_enabled: bool | None = None
dit_kwargs: dict[str, Any] = field(default_factory=dict)
text_encoder_kwargs: dict[str, Any] = field(default_factory=dict)
vae_kwargs: dict[str, Any] = field(default_factory=dict)
audio_vae_kwargs: dict[str, Any] = field(default_factory=dict)
@dataclass
class QuantizationConfig:
+4
View File
@@ -37,6 +37,10 @@ class VAEConfig(ModelConfig):
use_tiling: bool = True
use_temporal_tiling: bool = True
use_parallel_tiling: bool = True
# When True, latent preparation skips the schedule shift on frames
# whose temporal index is below the model's first-frame conditioning
# threshold. LTX-2 reads this in the latent prep stage.
use_temporal_scaling_frames: bool = True
def __post_init__(self):
self.blend_num_frames = self.tile_sample_min_num_frames - self.tile_sample_stride_num_frames
+3
View File
@@ -3,6 +3,8 @@
from fastvideo.entrypoints.cli.cli_types import CLISubcommand
from fastvideo.entrypoints.cli.generate import cmd_init as generate_cmd_init
from fastvideo.utils import FlexibleArgumentParser
from fastvideo.entrypoints.cli.router_serve import (
cmd_init as router_serve_cmd_init, )
from fastvideo.entrypoints.cli.serve import cmd_init as serve_cmd_init
from fastvideo.entrypoints.cli.bench import cmd_init as bench_cmd_init
@@ -12,6 +14,7 @@ def cmd_init() -> list[CLISubcommand]:
commands = []
commands.extend(generate_cmd_init())
commands.extend(serve_cmd_init())
commands.extend(router_serve_cmd_init())
commands.extend(bench_cmd_init())
return commands
+115
View File
@@ -0,0 +1,115 @@
# SPDX-License-Identifier: Apache-2.0
"""``fastvideo router-serve`` CLI subcommand.
Launches the streaming router from a YAML config. Separate from
``fastvideo serve`` because the router is an orthogonal process: it
fronts one or more running servers rather than hosting a generator
itself.
"""
from __future__ import annotations
import argparse
import os
from typing import cast
from fastvideo.api.parser import load_raw_config
from fastvideo.entrypoints.cli.cli_types import CLISubcommand
from fastvideo.entrypoints.streaming.router.config import (
ReplicaEndpoint,
RouterConfig,
)
from fastvideo.logger import init_logger
from fastvideo.utils import FlexibleArgumentParser
logger = init_logger(__name__)
class RouterServeSubcommand(CLISubcommand):
"""Start the multi-replica WebSocket router."""
def __init__(self) -> None:
self.name = "router-serve"
super().__init__()
def cmd(self, args: argparse.Namespace) -> None:
config = _load_router_config(args.config)
logger.info(
"router listening on %s:%d (%d replicas, %d primary)",
config.host,
config.port,
len(config.replicas),
sum(1 for r in config.replicas if r.primary),
)
from fastvideo.entrypoints.streaming.router.main import run_router
run_router(config)
def validate(self, args: argparse.Namespace) -> None:
if not args.config:
raise ValueError("fastvideo router-serve requires --config PATH")
if not os.path.exists(args.config):
raise ValueError(f"Router config file not found: {args.config}")
def subparser_init(
self,
subparsers: argparse._SubParsersAction,
) -> FlexibleArgumentParser:
parser = subparsers.add_parser(
"router-serve",
help="Start the streaming router (multi-replica load balancer)",
usage="fastvideo router-serve --config ROUTER_CONFIG",
)
parser.add_argument(
"--config",
type=str,
default="",
required=False,
help="Path to a YAML/JSON router config. Required.",
)
return cast(FlexibleArgumentParser, parser)
def _load_router_config(path: str) -> RouterConfig:
raw = load_raw_config(path)
router_raw = raw.get("router") if isinstance(raw, dict) else None
if not isinstance(router_raw, dict):
raise ValueError(f"Router config {path!r} must have a top-level `router:` block")
replicas_raw = router_raw.get("replicas", [])
if not isinstance(replicas_raw, list):
raise ValueError(f"router.replicas must be a list, got {type(replicas_raw).__name__}")
replicas = []
for i, r in enumerate(replicas_raw):
if not isinstance(r, dict):
raise ValueError(f"router.replicas[{i}] must be a mapping, got {type(r).__name__}")
url = r.get("url")
if not url:
raise ValueError(f"router.replicas[{i}] is missing required key 'url'")
replicas.append(
ReplicaEndpoint(
url=url,
name=r.get("name"),
primary=bool(r.get("primary", False)),
weight=float(r.get("weight", 1.0)),
))
if not replicas:
raise ValueError("Router config must list at least one replica under `router.replicas`")
health_check = router_raw.get("health_check") or {}
return RouterConfig(
host=str(router_raw.get("host", "0.0.0.0")),
port=int(router_raw.get("port", 9000)),
replicas=replicas,
health_check_path=str(health_check.get("path", "/health")),
health_check_interval_seconds=float(health_check.get("interval_seconds", 5.0)),
health_check_timeout_seconds=float(health_check.get("timeout_seconds", 2.0)),
failure_threshold=int(health_check.get("failure_threshold", 3)),
recovery_threshold=int(health_check.get("recovery_threshold", 2)),
)
def cmd_init() -> list[CLISubcommand]:
return [RouterServeSubcommand()]
__all__ = ["RouterServeSubcommand", "cmd_init"]
@@ -11,6 +11,28 @@ from fastvideo.entrypoints.streaming.session_store import (
InMemorySessionStore,
SessionStore,
)
from fastvideo.entrypoints.streaming.gpu_pool import (
GpuPool,
InProcessGpuPool,
PoolAcquireTimeout,
SubprocessGpuPool,
)
from fastvideo.entrypoints.streaming.mock_server import (
MockGenerator,
build_mock_app,
)
from fastvideo.entrypoints.streaming.prompt import (
LLMProvider,
PromptEnhancer,
)
from fastvideo.entrypoints.streaming.prompt.safety import (
PromptSafetyFilter,
SafetyDecision,
)
from fastvideo.entrypoints.streaming.session_logger import (
SessionLogEvent,
SessionLogger,
)
from fastvideo.entrypoints.streaming.stream import (
FragmentedMP4Chunk,
FragmentedMP4Encoder,
@@ -20,12 +42,24 @@ __all__ = [
"BlobStore",
"FragmentedMP4Chunk",
"FragmentedMP4Encoder",
"GpuPool",
"InMemoryBlobStore",
"InMemorySessionStore",
"InProcessGpuPool",
"LLMProvider",
"MockGenerator",
"PoolAcquireTimeout",
"PromptEnhancer",
"PromptSafetyFilter",
"SafetyDecision",
"SessionLogEvent",
"SessionLogger",
"build_mock_app",
"Session",
"SessionManager",
"SessionState",
"SessionStore",
"SubprocessGpuPool",
"build_app",
"run_server",
]
+542
View File
@@ -0,0 +1,542 @@
# SPDX-License-Identifier: Apache-2.0
"""GPU pool manager for the streaming server.
Replaces the single-generator path in PR 7.5 with a typed pool
abstraction. Three implementations ship here:
* :class:`InProcessGpuPool` — one in-process ``VideoGenerator``; used
by tests and single-GPU dev deployments.
* :class:`SubprocessGpuPool` — one ``multiprocessing.Process`` per
GPU, each running :func:`worker_main` against a ``GeneratorConfig``.
Jobs are dispatched via ``multiprocessing.Queue``.
* :class:`GpuPool` (abstract) — the interface both use.
Session-to-GPU binding lives in the pool so continuation state stays
on the GPU that generated the previous segment (matching the internal
``gpu_pool.py``'s per-GPU cache behavior). Cross-GPU handoff is
supported via :class:`SessionStore` snapshot + hydrate, which
serializes the state before the migration and rehydrates it on the
new worker.
Typed config: workers start from a :class:`GeneratorConfig` (no flat
LTX-2 kwargs), satisfying the PR 6 + PR 7 contracts that the public
surface doesn't reintroduce the legacy kwarg bag.
"""
from __future__ import annotations
import asyncio
import multiprocessing as mp
import queue
import threading
import time
import uuid
from abc import ABC, abstractmethod
from concurrent.futures import Future
from dataclasses import dataclass, field
from typing import Any, Protocol
from fastvideo.api.schema import (
GeneratorConfig,
GenerationRequest,
GpuPoolConfig,
WarmupConfig,
)
from fastvideo.entrypoints.streaming.session_store import (
InMemorySessionStore,
SessionStore,
)
from fastvideo.entrypoints.streaming.worker import worker_main
from fastvideo.logger import init_logger
logger = init_logger(__name__)
# ---------------------------------------------------------------------------
# Public interface
# ---------------------------------------------------------------------------
class _GeneratorLike(Protocol):
"""Subset the pool calls on a worker-side generator."""
def generate(self, request: GenerationRequest) -> Any:
...
@dataclass
class PoolAssignment:
"""The worker a session is currently bound to."""
gpu_id: int
worker_id: str
pinned_at: float = field(default_factory=time.monotonic)
class GpuPool(ABC):
"""Abstract GPU pool.
``acquire`` binds a session to a worker and holds that binding
across segments so continuation state can stay hot. ``run`` submits
a single ``GenerationRequest`` for a bound session.
Acquire / release are independent of run — a session can run many
segments on one acquired worker, and must release on disconnect.
"""
@abstractmethod
async def acquire(
self,
session_id: str,
*,
timeout: float | None = None,
) -> PoolAssignment:
...
@abstractmethod
async def run(
self,
session_id: str,
request: GenerationRequest,
) -> Any:
...
@abstractmethod
async def release(self, session_id: str) -> None:
...
@abstractmethod
async def shutdown(self) -> None:
...
@abstractmethod
def health(self) -> PoolHealth:
...
@dataclass
class PoolHealth:
total_workers: int
available_workers: int
active_sessions: int
queued_sessions: int = 0
class PoolAcquireTimeout(RuntimeError):
"""Raised when ``acquire`` times out waiting for a free worker."""
# ---------------------------------------------------------------------------
# In-process implementation (single-worker, test / dev)
# ---------------------------------------------------------------------------
class InProcessGpuPool(GpuPool):
"""Single-process pool backed by one :class:`_GeneratorLike`.
This is what PR 7.5's server uses by default; PR 7.6 adds the real
``SubprocessGpuPool`` alternative but keeps this one for tests and
small deployments.
"""
def __init__(
self,
generator: _GeneratorLike,
*,
gpu_id: int = 0,
session_store: SessionStore | None = None,
) -> None:
self._generator = generator
self._gpu_id = gpu_id
self._worker_id = f"inproc-{uuid.uuid4().hex[:6]}"
self._session_store = session_store or InMemorySessionStore()
self._active: dict[str, PoolAssignment] = {}
self._lock = asyncio.Lock()
self._gen_lock = asyncio.Lock()
async def acquire(
self,
session_id: str,
*,
timeout: float | None = None,
) -> PoolAssignment:
async with self._lock:
existing = self._active.get(session_id)
if existing is not None:
return existing
assignment = PoolAssignment(gpu_id=self._gpu_id, worker_id=self._worker_id)
self._active[session_id] = assignment
return assignment
async def run(
self,
session_id: str,
request: GenerationRequest,
) -> Any:
if session_id not in self._active:
raise RuntimeError(f"session {session_id!r} is not acquired on this pool")
# Serialize generator access so one GPU runs one request at a
# time, matching the internal gpu_pool's per-GPU lock.
async with self._gen_lock:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(None, self._generator.generate, request)
async def release(self, session_id: str) -> None:
async with self._lock:
self._active.pop(session_id, None)
async def shutdown(self) -> None:
self._active.clear()
def health(self) -> PoolHealth:
return PoolHealth(
total_workers=1,
available_workers=1 if not self._active else 0,
active_sessions=len(self._active),
)
# ---------------------------------------------------------------------------
# Subprocess implementation (multi-worker, real deployment)
# ---------------------------------------------------------------------------
@dataclass
class _WorkerHandle:
process: Any # mp.Process or compatible handle with is_alive / join / kill
job_queue: mp.Queue
result_queue: mp.Queue
gpu_id: int
worker_id: str
ready: threading.Event
# ``ready`` flips on either successful boot or boot failure so the
# parent stops waiting; ``boot_ok`` is set only on a real ready
# acknowledgement and is what gates pool admission.
boot_ok: threading.Event
shutdown_event: Any # mp.Event is a factory, not a type — Any keeps mypy sane
@dataclass
class _PendingJob:
job_id: str
future: Future
session_id: str
worker_id: str
class SubprocessGpuPool(GpuPool):
"""One ``multiprocessing.Process`` per GPU.
Each worker boots :class:`fastvideo.VideoGenerator` from a typed
:class:`GeneratorConfig` inside the child process (post-
``CUDA_VISIBLE_DEVICES`` setup) and consumes jobs from an mp Queue.
This is the production shape: the parent process stays CPU-only, and
GPU state never crosses process boundaries. Continuation state is
serialized through :class:`SessionStore` for cross-GPU handoff.
PR 7.6 ships this as an opt-in; PR 7.5's in-process pool remains the
default until nightly runs validate the subprocess path.
"""
def __init__(
self,
generator_config: GeneratorConfig,
*,
pool_config: GpuPoolConfig,
warmup_config: WarmupConfig | None = None,
session_store: SessionStore | None = None,
worker_factory: WorkerFactory | None = None,
) -> None:
self._generator_config = generator_config
self._pool_config = pool_config
self._warmup_config = warmup_config or WarmupConfig()
self._session_store = session_store or InMemorySessionStore()
self._worker_factory = worker_factory or _default_worker_factory
self._workers: list[_WorkerHandle] = []
self._available: asyncio.Queue[int] = asyncio.Queue()
self._assignments: dict[str, PoolAssignment] = {}
self._worker_by_id: dict[str, _WorkerHandle] = {}
self._pending: dict[str, _PendingJob] = {}
self._lock = asyncio.Lock()
self._result_reader_tasks: list[asyncio.Task] = []
async def start(self) -> None:
"""Spawn worker processes and wait for each to report ready."""
num_workers = self._pool_config.num_workers or 1
for gpu_id in range(num_workers):
handle = self._worker_factory(
gpu_id=gpu_id,
generator_config=self._generator_config,
warmup_config=self._warmup_config,
)
self._workers.append(handle)
self._worker_by_id[handle.worker_id] = handle
# Wait for each worker's ready event in a thread to avoid
# blocking the event loop.
loop = asyncio.get_running_loop()
await asyncio.gather(*[
loop.run_in_executor(None, handle.ready.wait, self._warmup_config.timeout_seconds)
for handle in self._workers
])
# Start background result readers — one task per worker
# drains its result queue and resolves futures in _pending.
for handle in self._workers:
task = asyncio.create_task(self._drain_results(handle))
self._result_reader_tasks.append(task)
# Only admit workers that successfully booted. Anything that
# failed boot (timeout, crash, error sentinel) stays out of the
# available queue so we never assign a session to it.
for idx, handle in enumerate(self._workers):
if handle.boot_ok.is_set():
await self._available.put(idx)
else:
logger.error(
"pool: worker %s failed to boot; skipping",
handle.worker_id,
)
async def acquire(
self,
session_id: str,
*,
timeout: float | None = None,
) -> PoolAssignment:
async with self._lock:
existing = self._assignments.get(session_id)
if existing is not None:
return existing
try:
idx = await asyncio.wait_for(self._available.get(), timeout=timeout)
except asyncio.TimeoutError as exc:
raise PoolAcquireTimeout(f"no worker available after {timeout}s") from exc
handle = self._workers[idx]
assignment = PoolAssignment(gpu_id=handle.gpu_id, worker_id=handle.worker_id)
async with self._lock:
self._assignments[session_id] = assignment
return assignment
async def run(
self,
session_id: str,
request: GenerationRequest,
) -> Any:
assignment = self._assignments.get(session_id)
if assignment is None:
raise RuntimeError(f"session {session_id!r} not acquired on this pool")
handle = self._worker_by_id[assignment.worker_id]
job_id = uuid.uuid4().hex
future: Future = Future()
self._pending[job_id] = _PendingJob(
job_id=job_id,
future=future,
session_id=session_id,
worker_id=handle.worker_id,
)
# mp.Queue.put can block if the underlying pipe buffer is full;
# offload to a thread so the event loop keeps serving other
# sessions. If the put itself fails, drop the pending entry so
# _drain_results doesn't dangle a future forever.
loop = asyncio.get_running_loop()
try:
await loop.run_in_executor(
None,
handle.job_queue.put,
{
"job_id": job_id,
"request": request
},
)
except Exception:
self._pending.pop(job_id, None)
raise
return await asyncio.wrap_future(future)
async def release(self, session_id: str) -> None:
async with self._lock:
assignment = self._assignments.pop(session_id, None)
if assignment is None:
return
idx = next((i for i, h in enumerate(self._workers) if h.worker_id == assignment.worker_id), None)
if idx is None:
return
# Don't return a dead worker to the pool; otherwise the next
# acquire will hand a session to a process that can't run jobs.
if not self._workers[idx].process.is_alive():
logger.warning(
"pool: worker %s died; not returning to available queue",
self._workers[idx].worker_id,
)
return
await self._available.put(idx)
async def shutdown(self) -> None:
loop = asyncio.get_running_loop()
# Signal all workers in parallel; .put may block on a full pipe,
# so off-load it the same way run() does.
async def _signal(handle: _WorkerHandle) -> None:
try:
handle.shutdown_event.set()
await loop.run_in_executor(None, handle.job_queue.put, None)
except Exception: # pragma: no cover - best-effort cleanup
pass
await asyncio.gather(*(_signal(h) for h in self._workers))
# Join in parallel so total shutdown is bounded by the slowest
# worker, not the sum of all timeouts.
await asyncio.gather(*(loop.run_in_executor(None, handle.process.join, 5.0) for handle in self._workers))
for handle in self._workers:
if handle.process.is_alive():
handle.process.kill()
for task in self._result_reader_tasks:
task.cancel()
self._result_reader_tasks.clear()
self._workers.clear()
self._worker_by_id.clear()
def health(self) -> PoolHealth:
return PoolHealth(
total_workers=len(self._workers),
available_workers=self._available.qsize(),
active_sessions=len(self._assignments),
)
async def _drain_results(self, handle: _WorkerHandle) -> None:
loop = asyncio.get_running_loop()
try:
while not handle.shutdown_event.is_set():
try:
msg = await loop.run_in_executor(None, _safe_queue_get, handle.result_queue, 0.5)
except Exception:
logger.exception("pool: worker %s result reader failed", handle.worker_id)
return
if msg is None:
continue
job_id = msg.get("job_id")
if job_id is None:
continue
pending = self._pending.pop(job_id, None)
if pending is None:
continue
if msg.get("kind") == "error":
pending.future.set_exception(RuntimeError(msg["error"]))
else:
pending.future.set_result(msg.get("result"))
finally:
# If we exit for any reason — shutdown, exception, cancel —
# surface that to any in-flight jobs on this worker so their
# await never hangs on a future no one will resolve.
for jid in [jid for jid, job in self._pending.items() if job.worker_id == handle.worker_id]:
pending = self._pending.pop(jid, None)
if pending is not None and not pending.future.done():
pending.future.set_exception(
RuntimeError(f"worker {handle.worker_id} result reader exited "
"with pending jobs"))
def _safe_queue_get(q: mp.Queue, timeout: float) -> Any | None:
try:
return q.get(timeout=timeout)
except queue.Empty:
return None
# ---------------------------------------------------------------------------
# Worker process
# ---------------------------------------------------------------------------
class WorkerFactory(Protocol):
def __call__(
self,
*,
gpu_id: int,
generator_config: GeneratorConfig,
warmup_config: WarmupConfig,
) -> _WorkerHandle:
...
def _default_worker_factory(
*,
gpu_id: int,
generator_config: GeneratorConfig,
warmup_config: WarmupConfig,
) -> _WorkerHandle:
"""Spawn a real multiprocessing worker.
The child process calls :func:`worker_main` which constructs a
:class:`VideoGenerator` from ``generator_config`` and runs a
blocking job loop. The ``ready`` event flips after the warmup
request completes.
"""
ctx = mp.get_context("spawn")
job_queue: mp.Queue = ctx.Queue()
result_queue: mp.Queue = ctx.Queue()
ready = threading.Event()
boot_ok = threading.Event()
shutdown_event = ctx.Event()
worker_id = f"gpu{gpu_id}-{uuid.uuid4().hex[:6]}"
process = ctx.Process(
target=worker_main,
kwargs={
"gpu_id": gpu_id,
"worker_id": worker_id,
"generator_config": generator_config,
"warmup_config": warmup_config,
"job_queue": job_queue,
"result_queue": result_queue,
"shutdown_event": shutdown_event,
},
daemon=False,
)
process.start()
# Block the parent-side ``ready`` flag until the worker posts a
# ready acknowledgement on the result queue. We drain that single
# sentinel here; subsequent results belong to jobs. ``boot_ok``
# only flips on a real ready; on error we set ``ready`` to unblock
# the parent's wait but leave ``boot_ok`` clear so the pool keeps
# the worker out of the available queue.
def _await_ready() -> None:
while not shutdown_event.is_set():
try:
msg = result_queue.get(timeout=1.0)
except queue.Empty:
continue
if isinstance(msg, dict) and msg.get("kind") == "ready":
boot_ok.set()
ready.set()
return
if isinstance(msg, dict) and msg.get("kind") == "error":
logger.error("pool: worker %s failed to boot: %s", worker_id, msg.get("error"))
ready.set()
return
threading.Thread(target=_await_ready, daemon=True).start()
return _WorkerHandle(
process=process,
job_queue=job_queue,
result_queue=result_queue,
gpu_id=gpu_id,
worker_id=worker_id,
ready=ready,
boot_ok=boot_ok,
shutdown_event=shutdown_event,
)
__all__ = [
"GpuPool",
"InProcessGpuPool",
"PoolAcquireTimeout",
"PoolAssignment",
"PoolHealth",
"SubprocessGpuPool",
"WorkerFactory",
"worker_main",
]
@@ -0,0 +1,122 @@
# SPDX-License-Identifier: Apache-2.0
"""Mock streaming server — a frontend dev aid.
Boots the same FastAPI app the real streaming server uses, but backs
it with :class:`InProcessGpuPool` wrapping a synthetic generator that
emits pre-baked RGB frames. No GPU or model weights required.
Use cases:
* Frontend development without a real model loaded.
* Integration tests that exercise the WS protocol end-to-end.
* Reproducing protocol bugs locally.
Launch: ``python -m fastvideo.entrypoints.streaming.mock_server``.
"""
from __future__ import annotations
import argparse
import time
from dataclasses import dataclass
from typing import Any
import numpy as np
from fastvideo.api.schema import (
ContinuationState,
GenerationRequest,
GeneratorConfig,
SamplingConfig,
ServeConfig,
StreamingConfig,
)
from fastvideo.entrypoints.streaming.server import build_app
@dataclass
class MockGenerator:
"""Generator stand-in that returns synthetic gradient frames.
Each call produces one segment worth of frames whose pixels vary by
a constant derived from the request seed and segment index. Latency
is configurable via ``sleep_ms`` so the caller can exercise slow-
generate scenarios without spinning a GPU.
"""
sleep_ms: float = 0.0
def generate(self, request: GenerationRequest) -> dict[str, Any]:
if self.sleep_ms:
time.sleep(self.sleep_ms / 1000.0)
width = max(16, request.sampling.width)
height = max(16, request.sampling.height)
num_frames = max(1, request.sampling.num_frames)
frames = [_gradient_frame(height, width, idx, seed=request.sampling.seed) for idx in range(num_frames)]
state = ContinuationState(
kind="ltx2.v1",
payload={
"schema_version": 1,
"segment_index": 0,
"source_prompt": request.prompt,
},
)
return {
"frames": frames,
"audio_sample_rate": 24000,
"state": state,
}
def _gradient_frame(height: int, width: int, idx: int, *, seed: int) -> np.ndarray:
base = (idx * 17 + seed * 3) % 256
row = np.linspace(base, (base + 64) % 256, width, dtype=np.uint8)
frame = np.tile(row, (height, 1))
stacked = np.stack([frame, np.roll(frame, 8, axis=1), np.roll(frame, 16, axis=1)], axis=-1)
return stacked.astype(np.uint8)
def build_mock_app(*, sleep_ms: float = 0.0):
"""Build a FastAPI app backed by :class:`MockGenerator`."""
serve_config = ServeConfig(
generator=GeneratorConfig(model_path="/models/mock"),
streaming=StreamingConfig(
session_timeout_seconds=120,
generation_segment_cap=6,
),
)
serve_config.default_request.sampling = SamplingConfig(
num_frames=24,
height=256,
width=256,
fps=24,
num_inference_steps=1,
)
return build_app(serve_config, MockGenerator(sleep_ms=sleep_ms))
def main() -> None: # pragma: no cover - CLI entry
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--host", default="127.0.0.1")
parser.add_argument("--port", type=int, default=8000)
parser.add_argument(
"--sleep-ms",
type=float,
default=0.0,
help="Per-segment artificial latency for testing slow paths",
)
args = parser.parse_args()
import uvicorn
app = build_mock_app(sleep_ms=args.sleep_ms)
uvicorn.run(app, host=args.host, port=args.port)
__all__ = [
"MockGenerator",
"build_mock_app",
"main",
]
if __name__ == "__main__": # pragma: no cover - CLI entry
main()
@@ -0,0 +1,36 @@
# SPDX-License-Identifier: Apache-2.0
"""Prompt pipeline for the streaming server.
* :mod:`providers` — LLM backend abstraction + built-in adapters
* :mod:`enhancer` — provider-agnostic enhance / auto-extend / rewrite
operations on top of the provider layer
All of this is optional; the streaming server runs fine without it
(PR 7.5's skeleton never invokes the enhancer). When the operator
enables ``ServeConfig.streaming.prompt.enabled``, the server routes
each ``session_init_v2`` curated prompt through ``enhance`` before the
first segment.
"""
from fastvideo.entrypoints.streaming.prompt.enhancer import (
PromptEnhancer,
PromptOperation,
)
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMMessage,
LLMProvider,
LLMProviderError,
LLMRequest,
LLMResponse,
LLMTimeoutError,
)
__all__ = [
"LLMMessage",
"LLMProvider",
"LLMProviderError",
"LLMRequest",
"LLMResponse",
"LLMTimeoutError",
"PromptEnhancer",
"PromptOperation",
]
@@ -0,0 +1,197 @@
# SPDX-License-Identifier: Apache-2.0
"""Provider-agnostic prompt orchestration for the streaming server.
Three operations the streaming server needs:
* ``enhance`` — polish a user prompt (add cinematic detail, fix syntax)
* ``auto_extend`` — generate a follow-on prompt for loop generation
* ``rewrite`` — rewrite a seed prompt for a user-directed rewrite flow
All three share the same orchestration: pick a provider in priority
order, submit an ``LLMRequest``, fall back to the next provider on
retryable errors, and surface a structured :class:`LLMResponse` back
to the caller.
System prompts are loaded from ``system_prompt_dir`` on construction
and can be hot-reloaded via :meth:`PromptEnhancer.reload_system_prompts`.
The streaming server's management endpoint calls that method in
response to a ``rewrite_seed_prompts_started`` frame.
"""
from __future__ import annotations
import enum
import os
from collections.abc import Sequence
from dataclasses import dataclass, replace
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMMessage,
LLMProvider,
LLMProviderError,
LLMRequest,
LLMResponse,
)
from fastvideo.logger import init_logger
logger = init_logger(__name__)
class PromptOperation(enum.Enum):
ENHANCE = "enhance"
AUTO_EXTEND = "auto_extend"
REWRITE = "rewrite"
@dataclass
class _SystemPrompts:
enhance: str
auto_extend: str
rewrite: str
_DEFAULT_SYSTEM_PROMPTS = _SystemPrompts(
enhance=("You are a prompt enhancer for cinematic video generation. Given "
"a user prompt, produce an enhanced prompt that is more vivid, "
"specific, and concrete. Keep the subject intact; add lighting, "
"camera, and motion detail. Reply with just the enhanced prompt."),
auto_extend=("You are a video continuation assistant. Given the current "
"sequence of prompts, produce one new prompt that naturally "
"continues the sequence. Reply with just the next prompt."),
rewrite=("You are a creative prompt rewriter. Given a seed prompt, produce "
"a set of alternative prompts that explore different angles, "
"styles, and moods. Reply with one prompt per line."),
)
class PromptEnhancer:
"""Orchestrates prompt operations across a priority-ordered provider
list with structured fallback + hot-reloadable system prompts.
Usage::
enhancer = PromptEnhancer(
providers=[CerebrasProvider(), GroqProvider()],
model="gpt-oss-120b",
system_prompt_dir="/etc/fastvideo/prompts",
)
response = await enhancer.enhance("a fox running through snow")
"""
def __init__(
self,
*,
providers: Sequence[LLMProvider],
model: str,
timeout_ms: int = 20000,
temperature: float = 0.7,
max_tokens: int | None = 256,
system_prompt_dir: str | None = None,
) -> None:
if not providers:
raise ValueError("PromptEnhancer requires at least one LLMProvider")
self._providers = list(providers)
self._model = model
self._timeout_ms = timeout_ms
self._temperature = temperature
self._max_tokens = max_tokens
self._system_prompt_dir = system_prompt_dir
self._system_prompts = self._load_system_prompts()
@property
def providers(self) -> list[LLMProvider]:
return list(self._providers)
def register_provider(self, provider: LLMProvider, *, priority: int = -1) -> None:
"""Insert an additional provider. ``priority=0`` makes it primary;
``priority=-1`` (default) appends as a fallback."""
if priority < 0:
self._providers.append(provider)
else:
self._providers.insert(priority, provider)
def reload_system_prompts(self) -> None:
"""Re-read the system prompt files from ``system_prompt_dir``.
The streaming server exposes this via a management endpoint so
operators can iterate on prompt templates without restarting
workers.
"""
self._system_prompts = self._load_system_prompts()
logger.info("prompt enhancer: reloaded system prompts from %s", self._system_prompt_dir or "defaults")
async def enhance(self, prompt: str) -> LLMResponse:
return await self._run(
PromptOperation.ENHANCE,
system=self._system_prompts.enhance,
user=prompt,
)
async def auto_extend(self, prior_prompts: Sequence[str]) -> LLMResponse:
user = "\n".join(prior_prompts)
return await self._run(
PromptOperation.AUTO_EXTEND,
system=self._system_prompts.auto_extend,
user=user,
)
async def rewrite(self, seed_prompt: str) -> LLMResponse:
return await self._run(
PromptOperation.REWRITE,
system=self._system_prompts.rewrite,
user=seed_prompt,
)
async def _run(
self,
operation: PromptOperation,
*,
system: str,
user: str,
) -> LLMResponse:
request = LLMRequest(
messages=[
LLMMessage(role="system", content=system),
LLMMessage(role="user", content=user),
],
model=self._model,
max_tokens=self._max_tokens,
temperature=self._temperature,
timeout_ms=self._timeout_ms,
)
last_error: LLMProviderError | None = None
for idx, provider in enumerate(self._providers):
try:
response = await provider.complete(request)
if idx > 0:
# Mark the fallback flag without losing any other
# response fields the provider populated.
response = replace(response, fallback_used=True)
return response
except LLMProviderError as exc:
logger.warning("prompt %s: provider %s failed: %s; trying next", operation.value, provider.name, exc)
last_error = exc
if not exc.retryable:
break
assert last_error is not None
raise last_error
def _load_system_prompts(self) -> _SystemPrompts:
if not self._system_prompt_dir:
return _DEFAULT_SYSTEM_PROMPTS
return _SystemPrompts(
enhance=_read_prompt(self._system_prompt_dir, "enhance.txt", _DEFAULT_SYSTEM_PROMPTS.enhance),
auto_extend=_read_prompt(self._system_prompt_dir, "auto_extend.txt", _DEFAULT_SYSTEM_PROMPTS.auto_extend),
rewrite=_read_prompt(self._system_prompt_dir, "rewrite.txt", _DEFAULT_SYSTEM_PROMPTS.rewrite),
)
def _read_prompt(dirname: str, filename: str, default: str) -> str:
path = os.path.join(dirname, filename)
if not os.path.exists(path):
return default
with open(path, encoding="utf-8") as f:
content = f.read().strip()
return content or default
__all__ = ["PromptEnhancer", "PromptOperation"]
@@ -0,0 +1,24 @@
# SPDX-License-Identifier: Apache-2.0
"""LLM provider implementations used by the prompt enhancer."""
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMMessage,
LLMProvider,
LLMProviderError,
LLMRequest,
LLMResponse,
LLMTimeoutError,
)
from fastvideo.entrypoints.streaming.prompt.providers.cerebras import (
CerebrasProvider, )
from fastvideo.entrypoints.streaming.prompt.providers.groq import GroqProvider
__all__ = [
"CerebrasProvider",
"GroqProvider",
"LLMMessage",
"LLMProvider",
"LLMProviderError",
"LLMRequest",
"LLMResponse",
"LLMTimeoutError",
]
@@ -0,0 +1,101 @@
# SPDX-License-Identifier: Apache-2.0
"""Shared HTTP path for OpenAI-compatible ``/chat/completions`` providers.
Cerebras and Groq both expose the OpenAI chat-completions schema, so
the request shape, error mapping, and response decoding are identical
between them. This module centralizes that logic; the per-provider
modules stay thin (just defaults + env var wiring).
"""
from __future__ import annotations
import time
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMProviderError,
LLMRequest,
LLMResponse,
LLMTimeoutError,
)
async def complete_openai_compatible(
*,
api_key: str | None,
api_key_hint: str,
base_url: str,
provider_name: str,
request: LLMRequest,
) -> LLMResponse:
"""Issue a chat-completions call and decode the OpenAI response."""
if not api_key:
raise LLMProviderError(
f"{provider_name} provider requires {api_key_hint} "
"(or explicit api_key=...)",
retryable=False,
)
try:
import httpx
except ImportError as exc: # pragma: no cover - optional dep
raise LLMProviderError(
f"{provider_name} provider requires httpx; install httpx",
retryable=False,
) from exc
timeout_s = (request.timeout_ms or 20000) / 1000.0
t0 = time.perf_counter()
try:
async with httpx.AsyncClient(timeout=timeout_s) as client:
response = await client.post(
f"{base_url}/chat/completions",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
},
json={
"model": request.model,
"messages": [{
"role": m.role,
"content": m.content
} for m in request.messages],
"max_tokens": request.max_tokens,
"temperature": request.temperature,
},
)
except httpx.TimeoutException as exc:
raise LLMTimeoutError(f"{provider_name} timed out after {timeout_s}s") from exc
except httpx.HTTPError as exc:
raise LLMProviderError(f"{provider_name} HTTP error: {exc}") from exc
if response.status_code >= 400:
# 5xx and 429 (rate-limit) are retryable: another provider may
# succeed. 4xx (auth, bad-request, etc.) are client errors —
# the enhancer should stop fallback traversal.
retryable = (response.status_code >= 500 or response.status_code == 429)
raise LLMProviderError(
f"{provider_name} returned {response.status_code}: "
f"{response.text[:200]}",
retryable=retryable,
)
try:
data = response.json()
except Exception as exc:
# Non-JSON body usually means a proxy / load-balancer error
# page; leave it retryable so a fallback provider can try.
raise LLMProviderError(f"{provider_name} returned non-JSON body: {exc}") from exc
choices = data.get("choices") or []
if not choices:
raise LLMProviderError(f"{provider_name} returned no choices")
content = choices[0].get("message", {}).get("content") or ""
latency_ms = (time.perf_counter() - t0) * 1000.0
return LLMResponse(
content=content.strip(),
provider=provider_name,
model=request.model,
latency_ms=latency_ms,
)
__all__ = ["complete_openai_compatible"]
@@ -0,0 +1,85 @@
# SPDX-License-Identifier: Apache-2.0
"""LLM provider protocol + DTOs used by the prompt enhancer.
Third-party users add a new provider by implementing
:class:`LLMProvider` and registering it with a prompt enhancer
instance. The shipped providers live in sibling modules
(``cerebras.py``, ``groq.py``) and each is ~100-200 LOC — the
provider layer is intentionally thin so the enhancer stays
provider-agnostic.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Literal, Protocol, runtime_checkable
@dataclass
class LLMMessage:
role: Literal["system", "user", "assistant"]
content: str
@dataclass
class LLMRequest:
messages: list[LLMMessage]
model: str
max_tokens: int | None = None
temperature: float | None = None
timeout_ms: int | None = None
@dataclass
class LLMResponse:
content: str
provider: str
model: str
latency_ms: float
fallback_used: bool = False
class LLMProviderError(RuntimeError):
"""Raised when an LLM provider fails a request.
``retryable`` controls whether the enhancer falls back to the next
provider. It is settable per-instance so the same exception type
can describe retryable transport errors (5xx, 429) and
non-retryable client errors (4xx auth/bad-request) without forcing
a separate subclass for every status family.
"""
def __init__(self, message: str, *, retryable: bool = True) -> None:
super().__init__(message)
self.retryable = retryable
class LLMTimeoutError(LLMProviderError):
"""Raised when an LLM provider times out — always retryable."""
def __init__(self, message: str) -> None:
super().__init__(message, retryable=True)
@runtime_checkable
class LLMProvider(Protocol):
"""Provider interface every LLM adapter implements.
Providers are async-first because every built-in implementation
talks to an HTTP API. Synchronous providers can wrap their call in
``asyncio.to_thread`` internally.
"""
name: str
async def complete(self, request: LLMRequest) -> LLMResponse:
...
__all__ = [
"LLMMessage",
"LLMProvider",
"LLMProviderError",
"LLMRequest",
"LLMResponse",
"LLMTimeoutError",
]
@@ -0,0 +1,44 @@
# SPDX-License-Identifier: Apache-2.0
"""Cerebras LLM provider (OpenAI-compatible chat endpoint)."""
from __future__ import annotations
import os
from dataclasses import dataclass
from fastvideo.entrypoints.streaming.prompt.providers._openai_compat import (
complete_openai_compatible, )
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMRequest,
LLMResponse,
)
_DEFAULT_BASE_URL = "https://api.cerebras.ai/v1"
_API_KEY_ENV = "CEREBRAS_API_KEY"
@dataclass
class CerebrasProvider:
"""Cerebras inference adapter.
``api_key`` falls back to ``CEREBRAS_API_KEY`` when unset.
"""
api_key: str | None = None
base_url: str = _DEFAULT_BASE_URL
name: str = "cerebras"
def __post_init__(self) -> None:
if self.api_key is None:
self.api_key = os.environ.get(_API_KEY_ENV)
async def complete(self, request: LLMRequest) -> LLMResponse:
return await complete_openai_compatible(
api_key=self.api_key,
api_key_hint=_API_KEY_ENV,
base_url=self.base_url,
provider_name=self.name,
request=request,
)
__all__ = ["CerebrasProvider"]
@@ -0,0 +1,46 @@
# SPDX-License-Identifier: Apache-2.0
"""Groq LLM provider (OpenAI-compatible chat endpoint)."""
from __future__ import annotations
import os
from dataclasses import dataclass
from fastvideo.entrypoints.streaming.prompt.providers._openai_compat import (
complete_openai_compatible, )
from fastvideo.entrypoints.streaming.prompt.providers.base import (
LLMRequest,
LLMResponse,
)
_DEFAULT_BASE_URL = "https://api.groq.com/openai/v1"
_API_KEY_ENV = "GROQ_API_KEY"
@dataclass
class GroqProvider:
"""Groq inference adapter.
Identical wire format to :class:`CerebrasProvider`; both go through
:func:`complete_openai_compatible`. The two providers differ only
in base URL, env var, and model id conventions.
"""
api_key: str | None = None
base_url: str = _DEFAULT_BASE_URL
name: str = "groq"
def __post_init__(self) -> None:
if self.api_key is None:
self.api_key = os.environ.get(_API_KEY_ENV)
async def complete(self, request: LLMRequest) -> LLMResponse:
return await complete_openai_compatible(
api_key=self.api_key,
api_key_hint=_API_KEY_ENV,
base_url=self.base_url,
provider_name=self.name,
request=request,
)
__all__ = ["GroqProvider"]
@@ -0,0 +1,82 @@
# SPDX-License-Identifier: Apache-2.0
"""Rewrite payload builder.
The UI's "rewrite seed prompts" flow asks the enhancer to produce a
batch of alternative prompts given one seed. This module packages the
seed + options into the payload the enhancer expects and unpacks the
response back into a typed :class:`RewriteResult`.
Separating this from :mod:`enhancer` keeps the enhancer provider-
agnostic; anything UI-specific (how many alternatives to request, how
to split the response, temperature) lives here.
"""
from __future__ import annotations
import re
from dataclasses import dataclass
from fastvideo.entrypoints.streaming.prompt.enhancer import PromptEnhancer
_LEADING_MARKER_RE = re.compile(r"^(?:[-*•]\s*|\d+\s*[.)]\s*)+")
@dataclass
class RewriteOptions:
count: int = 3
"""Number of alternative prompts to request."""
temperature: float | None = None
@dataclass
class RewriteResult:
seed_prompt: str
alternatives: list[str]
provider: str
model: str
latency_ms: float
fallback_used: bool = False
async def build_rewrite(
enhancer: PromptEnhancer,
seed_prompt: str,
*,
options: RewriteOptions | None = None,
) -> RewriteResult:
"""Run a rewrite op through the enhancer and return a typed result."""
if not seed_prompt.strip():
raise ValueError("rewrite seed prompt must be non-empty")
options = options or RewriteOptions()
response = await enhancer.rewrite(seed_prompt)
alternatives = _split_response(response.content, limit=options.count)
return RewriteResult(
seed_prompt=seed_prompt,
alternatives=alternatives,
provider=response.provider,
model=response.model,
latency_ms=response.latency_ms,
fallback_used=response.fallback_used,
)
def _split_response(content: str, *, limit: int) -> list[str]:
"""Split the LLM response into discrete prompt candidates.
The shipped system prompt instructs the model to emit one prompt
per line; this function is forgiving about numbered lists or
leading bullets so user-supplied system prompts don't break it.
"""
lines = [line.strip() for line in content.splitlines() if line.strip()]
cleaned: list[str] = []
for line in lines:
stripped = _LEADING_MARKER_RE.sub("", line).strip()
if stripped:
cleaned.append(stripped)
return cleaned[:max(1, limit)]
__all__ = [
"RewriteOptions",
"RewriteResult",
"build_rewrite",
]
@@ -0,0 +1,146 @@
# SPDX-License-Identifier: Apache-2.0
"""Optional prompt safety filter.
Uses a fastText classifier to score prompts against a banned-content
rubric. Only loaded when ``ServeConfig.streaming.safety.enabled`` is
True and fastText is installed — users who don't need it see no
runtime cost.
Install: ``pip install fastvideo[prompt-safety]`` (ships fasttext as an
optional extra) or install fasttext directly.
"""
from __future__ import annotations
import enum
import threading
from dataclasses import dataclass
from typing import Any
from fastvideo.logger import init_logger
logger = init_logger(__name__)
class SafetyDecision(enum.Enum):
ALLOW = "allow"
BLOCK = "block"
UNAVAILABLE = "unavailable"
"""Returned when the classifier can't run (not configured, fastText
missing). Safety is opt-in; the server treats ``UNAVAILABLE`` as
``ALLOW`` but logs it so operators know the filter is off."""
@dataclass
class SafetyResult:
prompt: str
decision: SafetyDecision
score: float = 0.0
label: str | None = None
reason: str | None = None
class PromptSafetyFilter:
"""Minimal fastText-backed prompt safety filter.
Loads the classifier lazily on first use so the streaming server
can construct the filter eagerly at startup without paying the
model-load cost when safety is disabled.
"""
def __init__(
self,
*,
classifier_path: str | None,
enabled: bool = True,
block_threshold: float = 0.5,
) -> None:
self._classifier_path = classifier_path
self._enabled = enabled
self._block_threshold = block_threshold
self._model: Any | None = None
self._load_attempted = False
self._load_lock = threading.Lock()
@property
def enabled(self) -> bool:
return self._enabled and self._classifier_path is not None
def classify(self, prompt: str) -> SafetyResult:
if not self.enabled:
return SafetyResult(
prompt=prompt,
decision=SafetyDecision.UNAVAILABLE,
reason="safety filter not enabled",
)
model = self._ensure_loaded()
if model is None:
return SafetyResult(
prompt=prompt,
decision=SafetyDecision.UNAVAILABLE,
reason="fastText model unavailable",
)
try:
labels, probs = model.predict(prompt.replace("\n", " "), k=1)
except Exception as exc: # pragma: no cover - defensive
logger.warning("safety: classifier failed: %s", exc)
return SafetyResult(
prompt=prompt,
decision=SafetyDecision.UNAVAILABLE,
reason=f"classifier error: {exc}",
)
label = labels[0].removeprefix("__label__") if labels else None
score = float(probs[0]) if len(probs) else 0.0
decision = (SafetyDecision.BLOCK if
(label == "unsafe" and score >= self._block_threshold) else SafetyDecision.ALLOW)
return SafetyResult(
prompt=prompt,
decision=decision,
score=score,
label=label,
)
def _ensure_loaded(self) -> Any | None:
if self._model is not None:
return self._model
if self._load_attempted:
return None
with self._load_lock:
if self._model is not None:
return self._model
if self._load_attempted:
return None
self._load_attempted = True
if self._classifier_path is None:
return None
try:
import fasttext # type: ignore[import-not-found]
except ImportError:
logger.warning("safety: fasttext not installed; safety filter disabled. "
"Install fastvideo[prompt-safety] to enable.")
return None
try:
self._model = fasttext.load_model(self._classifier_path)
except Exception as exc: # pragma: no cover - requires real model
logger.warning("safety: failed to load %s: %s", self._classifier_path, exc)
return None
return self._model
def first_blocked(
filter_: PromptSafetyFilter,
prompts: list[str],
) -> SafetyResult | None:
"""Return the first prompt the filter blocks, or ``None``."""
for prompt in prompts:
result = filter_.classify(prompt)
if result.decision is SafetyDecision.BLOCK:
return result
return None
__all__ = [
"PromptSafetyFilter",
"SafetyDecision",
"SafetyResult",
"first_blocked",
]
@@ -0,0 +1,27 @@
# SPDX-License-Identifier: Apache-2.0
"""Multi-replica load balancer + WebSocket proxy for the streaming server.
Sits in front of one-or-more streaming-server replicas and forwards
WebSocket sessions to a healthy primary, with failover to secondaries.
Kept in-repo under ``fastvideo/entrypoints/streaming/router/`` per the
PR plan's default; the alternative (separate package) is an open
question deferred to review.
"""
from fastvideo.entrypoints.streaming.router.registry import (
Replica,
ReplicaHealth,
ReplicaRegistry,
ReplicaStatus,
)
from fastvideo.entrypoints.streaming.router.config import RouterConfig
from fastvideo.entrypoints.streaming.router.main import build_router_app, run_router
__all__ = [
"Replica",
"ReplicaHealth",
"ReplicaRegistry",
"ReplicaStatus",
"RouterConfig",
"build_router_app",
"run_router",
]
@@ -0,0 +1,88 @@
# SPDX-License-Identifier: Apache-2.0
"""Typed router configuration."""
from __future__ import annotations
from dataclasses import dataclass, field
from urllib.parse import urlparse
@dataclass
class ReplicaEndpoint:
"""One backend replica the router can route to."""
url: str
"""HTTP base URL, e.g. ``http://host:8000``. WebSocket URL is
derived automatically by replacing the scheme."""
name: str | None = None
primary: bool = False
"""``True`` = prefer this replica over others in steady state."""
weight: float = 1.0
@dataclass
class RouterConfig:
"""Typed router config loaded from a YAML file.
Example::
router:
host: 0.0.0.0
port: 9000
replicas:
- url: http://streamer-a:8000
primary: true
- url: http://streamer-b:8000
health_check:
path: /health
interval_seconds: 5
failure_threshold: 3
Validation runs in ``__post_init__``: empty replicas, non-positive
intervals/timeouts, thresholds < 1, non-http(s) URLs, and more than
one primary all raise ``ValueError`` so misconfigurations surface at
load time rather than as confusing runtime failures.
"""
host: str = "0.0.0.0"
port: int = 9000
replicas: list[ReplicaEndpoint] = field(default_factory=list)
health_check_path: str = "/health"
health_check_interval_seconds: float = 5.0
health_check_timeout_seconds: float = 2.0
failure_threshold: int = 3
recovery_threshold: int = 2
def __post_init__(self) -> None:
if not self.replicas:
raise ValueError("RouterConfig.replicas must list at least one replica")
if self.health_check_interval_seconds <= 0:
raise ValueError(f"health_check_interval_seconds must be > 0, got {self.health_check_interval_seconds}")
if self.health_check_timeout_seconds <= 0:
raise ValueError(f"health_check_timeout_seconds must be > 0, got {self.health_check_timeout_seconds}")
if self.failure_threshold < 1:
raise ValueError(f"failure_threshold must be >= 1, got {self.failure_threshold}")
if self.recovery_threshold < 1:
raise ValueError(f"recovery_threshold must be >= 1, got {self.recovery_threshold}")
seen_urls: set[str] = set()
for replica in self.replicas:
if not replica.url.startswith(("http://", "https://")):
raise ValueError(f"ReplicaEndpoint.url must start with http:// or https://, got {replica.url!r}")
parsed = urlparse(replica.url)
if parsed.path not in ("", "/"):
raise ValueError(f"ReplicaEndpoint.url must be a base host[:port] URL without a path; "
f"got {replica.url!r} with path {parsed.path!r}. The router appends "
"`/health` and `/v1/stream` itself.")
if parsed.query or parsed.fragment:
raise ValueError(f"ReplicaEndpoint.url must not include query/fragment; got {replica.url!r}")
if replica.url in seen_urls:
raise ValueError(f"Duplicate ReplicaEndpoint.url {replica.url!r}; "
"router selection keys by URL so duplicates would silently collapse")
seen_urls.add(replica.url)
primaries = sum(1 for r in self.replicas if r.primary)
if primaries > 1:
raise ValueError(f"RouterConfig allows at most one primary replica; got {primaries}. "
"Multi-primary load distribution is deferred — promote one replica to "
"primary and treat the rest as secondaries.")
__all__ = ["ReplicaEndpoint", "RouterConfig"]
@@ -0,0 +1,218 @@
# SPDX-License-Identifier: Apache-2.0
"""Router FastAPI entry point.
Exposes the same ``/v1/stream`` WebSocket path the backend servers do,
accepts a client, picks a healthy replica from the registry, and
proxies frames bidirectionally.
PR 7.9 ships the minimum-viable shape: explicit replica list, single
primary, JSON + binary passthrough in both directions, and a
``/status`` endpoint for operators. Sticky-session routing (so a
reconnect lands on the same backend) is left for a follow-up.
"""
from __future__ import annotations
import asyncio
import contextlib
from dataclasses import dataclass
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
from fastapi.responses import JSONResponse
from fastvideo.entrypoints.streaming.router.config import RouterConfig
from fastvideo.entrypoints.streaming.router.registry import (
ReplicaRegistry,
run_health_check_loop,
)
from fastvideo.logger import init_logger
logger = init_logger(__name__)
@dataclass
class _RouterState:
config: RouterConfig
registry: ReplicaRegistry
stop_event: asyncio.Event
health_task: asyncio.Task | None = None
def build_router_app(
config: RouterConfig,
*,
registry: ReplicaRegistry | None = None,
) -> FastAPI:
"""Build the router FastAPI app.
``registry`` can be injected for tests; defaults to one built from
``config.replicas``.
"""
registry = registry or ReplicaRegistry(config.replicas)
state = _RouterState(
config=config,
registry=registry,
stop_event=asyncio.Event(),
)
@contextlib.asynccontextmanager
async def _lifespan(_app: FastAPI):
state.health_task = asyncio.create_task(
run_health_check_loop(
registry=state.registry,
config=state.config,
stop_event=state.stop_event,
))
try:
yield
finally:
state.stop_event.set()
if state.health_task is not None:
with contextlib.suppress(asyncio.CancelledError):
await state.health_task
app = FastAPI(title="FastVideo Streaming Router", lifespan=_lifespan)
@app.get("/status")
async def _status() -> JSONResponse:
return JSONResponse({
"replicas": [{
"url": r.url,
"primary": r.primary,
"status": r.health.status.value,
"last_ok_at": r.health.last_ok_at,
"last_latency_ms": r.health.last_latency_ms,
"consecutive_failures": r.health.consecutive_failures,
} for r in state.registry.all()],
})
@app.websocket("/v1/stream")
async def _proxy(websocket: WebSocket) -> None:
await websocket.accept()
replica = state.registry.select()
if replica is None:
await websocket.send_json({
"type": "error",
"code": "gpu_unavailable",
"message": "router: no healthy replica available",
"retryable": True,
})
await websocket.close(code=1013, reason="no_healthy_replica")
return
ws_url = _websocket_url_for(replica.url)
try:
await _bridge_session(websocket, ws_url)
except WebSocketDisconnect:
logger.info("router: client disconnected")
except Exception as exc:
logger.exception("router: bridge failed: %s", exc)
with contextlib.suppress(RuntimeError):
await websocket.send_json({
"type": "error",
"code": "worker_failed",
"message": f"router bridge failed: {exc}",
"retryable": True,
})
with contextlib.suppress(RuntimeError):
await websocket.close(code=1011)
app.state.router_state = state
return app
def run_router(config: RouterConfig) -> None: # pragma: no cover - CLI
import uvicorn
app = build_router_app(config)
uvicorn.run(app, host=config.host, port=config.port)
async def _bridge_session(
client_ws: WebSocket,
backend_ws_url: str,
) -> None:
"""Connect to backend and shuttle messages in both directions.
Uses ``websockets`` for the backend side; imported lazily to keep
the router's import graph small for users who only want the server.
Cancellation: when either direction completes (client disconnect,
backend close, exception), the other is cancelled explicitly and
both are drained before returning. Unexpected exceptions from the
direction that completed first are re-raised; normal disconnect
paths (``WebSocketDisconnect``, ``ConnectionClosed``,
``CancelledError``) are swallowed.
"""
try:
import websockets
except ImportError as exc: # pragma: no cover - optional extra
raise RuntimeError("router requires the `websockets` package for backend proxying") from exc
async with websockets.connect(backend_ws_url + "/v1/stream") as backend_ws:
c2b = asyncio.create_task(_forward_client_to_backend(client_ws, backend_ws))
b2c = asyncio.create_task(_forward_backend_to_client(backend_ws, client_ws))
try:
done, _pending = await asyncio.wait(
{c2b, b2c},
return_when=asyncio.FIRST_COMPLETED,
)
finally:
for task in (c2b, b2c):
if not task.done():
task.cancel()
await asyncio.gather(c2b, b2c, return_exceptions=True)
for task in done:
task_exc = task.exception()
if task_exc is not None and not _is_normal_disconnect(task_exc):
raise task_exc
def _is_normal_disconnect(exc: BaseException) -> bool:
"""Whether ``exc`` is a routine WebSocket teardown vs a real bridge fault."""
if isinstance(exc, asyncio.CancelledError | WebSocketDisconnect):
return True
name = type(exc).__name__
# websockets.exceptions.ConnectionClosed{,OK,Error} all subclass
# WebSocketException; check by name to avoid the lazy-import dance.
return name.startswith("ConnectionClosed")
async def _forward_client_to_backend(client_ws: WebSocket, backend_ws) -> None:
try:
while True:
msg = await client_ws.receive()
if msg.get("type") == "websocket.disconnect":
break
if "text" in msg and msg["text"] is not None:
await backend_ws.send(msg["text"])
elif "bytes" in msg and msg["bytes"] is not None:
await backend_ws.send(msg["bytes"])
finally:
with contextlib.suppress(Exception):
await backend_ws.close()
async def _forward_backend_to_client(backend_ws, client_ws: WebSocket) -> None:
try:
async for frame in backend_ws:
if isinstance(frame, bytes):
await client_ws.send_bytes(frame)
else:
await client_ws.send_text(frame)
finally:
with contextlib.suppress(Exception):
await client_ws.close()
def _websocket_url_for(http_url: str) -> str:
if http_url.startswith("https://"):
return "wss://" + http_url[len("https://"):]
if http_url.startswith("http://"):
return "ws://" + http_url[len("http://"):]
return http_url
__all__ = [
"build_router_app",
"run_router",
]
@@ -0,0 +1,268 @@
# SPDX-License-Identifier: Apache-2.0
"""Replica registry + health-check loop.
The registry tracks the set of known backend replicas and their live
health. The router consults it for "pick a backend for this session"
decisions and a background task updates it from periodic HTTP probes.
State machine per replica::
HEALTHY ──(N consecutive failures)──▶ UNHEALTHY
▲ │
└──────(M consecutive successes)──────┘
Where N = :attr:`RouterConfig.failure_threshold` and
M = :attr:`RouterConfig.recovery_threshold`.
"""
from __future__ import annotations
import asyncio
import contextlib
import enum
import time
from collections.abc import AsyncIterator, Awaitable, Callable
from dataclasses import dataclass, field
from typing import Any
from fastvideo.entrypoints.streaming.router.config import (
ReplicaEndpoint,
RouterConfig,
)
from fastvideo.logger import init_logger
HttpProbe = Any
"""Structural alias for health-probe callables. Concrete signature is
``async def __call__(url: str, *, timeout: float) -> tuple[float,
str | None]``; typing.Callable cannot express keyword-only parameters,
so duck-typing is the pragmatic compromise."""
logger = init_logger(__name__)
class ReplicaStatus(enum.Enum):
UNKNOWN = "unknown"
HEALTHY = "healthy"
UNHEALTHY = "unhealthy"
@dataclass
class ReplicaHealth:
status: ReplicaStatus = ReplicaStatus.UNKNOWN
last_ok_at: float | None = None
last_failure_at: float | None = None
consecutive_failures: int = 0
consecutive_successes: int = 0
last_latency_ms: float | None = None
@dataclass
class Replica:
endpoint: ReplicaEndpoint
health: ReplicaHealth = field(default_factory=ReplicaHealth)
@property
def url(self) -> str:
return self.endpoint.url
@property
def primary(self) -> bool:
return self.endpoint.primary
@property
def is_healthy(self) -> bool:
return self.health.status is ReplicaStatus.HEALTHY
class ReplicaRegistry:
"""Stateful map of replica URL → :class:`Replica`.
Selection favors primary replicas when healthy; otherwise the first
healthy non-primary is returned. When none are healthy, the
registry returns ``None`` so the router can reject incoming
sessions with ``gpu_unavailable``.
"""
def __init__(self, replicas: list[ReplicaEndpoint]) -> None:
if not replicas:
raise ValueError("ReplicaRegistry requires at least one replica")
self._replicas: dict[str, Replica] = {endpoint.url: Replica(endpoint=endpoint) for endpoint in replicas}
self._lock = asyncio.Lock()
def all(self) -> list[Replica]:
return list(self._replicas.values())
def get(self, url: str) -> Replica | None:
return self._replicas.get(url)
def primaries(self) -> list[Replica]:
return [r for r in self._replicas.values() if r.primary]
def select(self) -> Replica | None:
"""Pick the best healthy replica.
Priority order:
1. The first healthy primary (insertion order).
2. The first healthy non-primary (insertion order).
3. ``None`` when nothing is healthy.
This MVP picks the first match within each tier; it does NOT
load-balance across multiple healthy replicas of the same tier.
Round-robin and weighted distribution are deferred until a real
N-way active deployment exists.
"""
healthy_primaries = [r for r in self._replicas.values() if r.primary and r.is_healthy]
if healthy_primaries:
return healthy_primaries[0]
healthy = [r for r in self._replicas.values() if r.is_healthy]
if healthy:
return healthy[0]
return None
async def record_success(
self,
replica: Replica,
*,
recovery_threshold: int,
latency_ms: float,
) -> None:
async with self._lock:
h = replica.health
h.last_ok_at = time.time()
h.last_latency_ms = latency_ms
h.consecutive_failures = 0
h.consecutive_successes += 1
# State machine: UNKNOWN -> HEALTHY is immediate; only the
# UNHEALTHY -> HEALTHY transition is gated by recovery_threshold.
if h.status is ReplicaStatus.UNKNOWN:
logger.info("router: replica %s initial probe ok, marking HEALTHY", replica.url)
h.status = ReplicaStatus.HEALTHY
h.consecutive_successes = 0
elif (h.status is ReplicaStatus.UNHEALTHY and h.consecutive_successes >= recovery_threshold):
logger.info("router: replica %s recovered to HEALTHY after %d successes", replica.url,
h.consecutive_successes)
h.status = ReplicaStatus.HEALTHY
h.consecutive_successes = 0
async def record_failure(
self,
replica: Replica,
*,
failure_threshold: int,
reason: str,
) -> None:
async with self._lock:
h = replica.health
h.last_failure_at = time.time()
h.consecutive_successes = 0
h.consecutive_failures += 1
if (h.status is not ReplicaStatus.UNHEALTHY and h.consecutive_failures >= failure_threshold):
logger.warning("router: replica %s marked UNHEALTHY after %d failures: %s", replica.url,
h.consecutive_failures, reason)
h.status = ReplicaStatus.UNHEALTHY
async def run_health_check_loop(
registry: ReplicaRegistry,
config: RouterConfig,
*,
stop_event: asyncio.Event,
http_get: HttpProbe | None = None,
) -> None:
"""Poll all replicas' health endpoints in parallel on a fixed interval.
``http_get`` is pluggable so unit tests can inject a deterministic
probe without hitting the network. The default builds a single
``httpx.AsyncClient`` shared across the loop's lifetime so the
common case (steady polling against a stable replica set) reuses
TCP/TLS connections instead of paying handshake cost per probe.
Probes within one polling cycle run concurrently via ``asyncio.gather``
so a slow replica doesn't push the cycle past
``health_check_interval_seconds``.
"""
if http_get is not None:
await _run_loop(registry, config, stop_event, http_get)
return
async with _build_default_probe(config) as probe:
await _run_loop(registry, config, stop_event, probe)
async def _run_loop(
registry: ReplicaRegistry,
config: RouterConfig,
stop_event: asyncio.Event,
http_get: Callable[..., Awaitable[tuple[float, str | None]]],
) -> None:
while not stop_event.is_set():
replicas = registry.all()
results = await asyncio.gather(
*[
http_get(replica.url + config.health_check_path, timeout=config.health_check_timeout_seconds)
for replica in replicas
],
return_exceptions=True,
)
for replica, result in zip(replicas, results, strict=True):
if isinstance(result, BaseException):
await registry.record_failure(
replica,
failure_threshold=config.failure_threshold,
reason=f"{type(result).__name__}: {result}",
)
continue
status_ms, error = result
if error is None:
await registry.record_success(
replica,
recovery_threshold=config.recovery_threshold,
latency_ms=status_ms,
)
else:
await registry.record_failure(
replica,
failure_threshold=config.failure_threshold,
reason=error,
)
try:
await asyncio.wait_for(
stop_event.wait(),
timeout=config.health_check_interval_seconds,
)
except asyncio.TimeoutError:
continue
@contextlib.asynccontextmanager
async def _build_default_probe(
config: RouterConfig, ) -> AsyncIterator[Callable[..., Awaitable[tuple[float, str | None]]]]:
try:
import httpx
except ImportError as exc: # pragma: no cover - optional extra
raise RuntimeError("router health checks require httpx; install with "
"`pip install fastvideo[streaming]` or `pip install httpx`") from exc
async with httpx.AsyncClient(timeout=config.health_check_timeout_seconds) as client:
async def probe(url: str, *, timeout: float) -> tuple[float, str | None]:
start = time.perf_counter()
try:
response = await client.get(url, timeout=timeout)
except Exception as exc:
return 0.0, f"{type(exc).__name__}: {exc}"
latency_ms = (time.perf_counter() - start) * 1000.0
if response.status_code >= 400:
return latency_ms, f"HTTP {response.status_code}"
return latency_ms, None
yield probe
__all__ = [
"HttpProbe",
"Replica",
"ReplicaHealth",
"ReplicaRegistry",
"ReplicaStatus",
"run_health_check_loop",
]
+44 -15
View File
@@ -51,6 +51,11 @@ from fastvideo.entrypoints.streaming.session import (
)
from fastvideo.entrypoints.streaming.session_init_image import (
persist_session_init_image, )
from fastvideo.entrypoints.streaming.gpu_pool import (
GpuPool,
InProcessGpuPool,
PoolAcquireTimeout,
)
from fastvideo.entrypoints.streaming.session_store import (
InMemorySessionStore,
SessionStore,
@@ -75,25 +80,37 @@ class _GeneratorProto(Protocol):
@dataclass
class ServerState:
serve_config: ServeConfig
generator: _GeneratorProto
pool: GpuPool
sessions: SessionManager
session_store: SessionStore
def build_app(
serve_config: ServeConfig,
generator: _GeneratorProto,
generator: _GeneratorProto | None = None,
*,
pool: GpuPool | None = None,
session_store: SessionStore | None = None,
) -> FastAPI:
"""Build the FastAPI app used by :func:`run_server`.
Exposed so tests can drive the WebSocket endpoint in-process via
``starlette.testclient.TestClient(app).websocket_connect(...)``.
Exactly one of ``generator`` (backed by :class:`InProcessGpuPool`)
or ``pool`` (for the subprocess-backed production shape) must be
given.
"""
if serve_config.streaming is None:
raise ValueError("ServeConfig.streaming must be set to launch the streaming "
"server; got None. Add a `streaming:` block to your serve config.")
if (generator is None) == (pool is None):
raise ValueError("build_app requires exactly one of `generator` or `pool`")
store = session_store or InMemorySessionStore()
if pool is None:
assert generator is not None
pool = InProcessGpuPool(generator, session_store=store)
sessions = SessionManager(
segment_cap=serve_config.streaming.generation_segment_cap,
@@ -101,9 +118,9 @@ def build_app(
)
state = ServerState(
serve_config=serve_config,
generator=generator,
pool=pool,
sessions=sessions,
session_store=session_store or InMemorySessionStore(),
session_store=store,
)
app = FastAPI(title="FastVideo Streaming")
@@ -135,6 +152,8 @@ def build_app(
with contextlib.suppress(InvalidSessionTransition):
session.transition(SessionState.ERROR)
finally:
with contextlib.suppress(Exception):
await state.pool.release(session.id)
_cleanup_session(session, state)
app.state.server_state = state
@@ -178,10 +197,22 @@ async def _handle_session(
await _apply_session_init(session, init, state)
await _send_json(websocket, QueueStatus(position=0, queue_depth=0))
session.transition(SessionState.GPU_BINDING)
await _send_json(websocket, GpuAssigned(
gpu_id=0,
session_timeout=state.sessions.session_timeout_seconds,
))
try:
assignment = await state.pool.acquire(
session.id,
timeout=float(state.sessions.session_timeout_seconds),
)
except PoolAcquireTimeout as exc:
await _send_error(websocket, "gpu_unavailable", str(exc), retryable=True)
with contextlib.suppress(InvalidSessionTransition):
session.transition(SessionState.TIMEOUT)
return
session.gpu_id = assignment.gpu_id
await _send_json(websocket,
GpuAssigned(
gpu_id=assignment.gpu_id,
session_timeout=state.sessions.session_timeout_seconds,
))
session.transition(SessionState.ACTIVE)
await _send_json(websocket, _build_stream_start(session, state))
@@ -327,15 +358,13 @@ async def _run_segment(
))
start = time.perf_counter()
loop = asyncio.get_running_loop()
# TODO: executor-wrapped generate() cannot be cancelled, so a
# client disconnect mid-segment leaves the GPU work running to
# completion. Real cancellation needs the generate_async API.
# TODO: pool.run() runs to completion even if the client disconnects
# mid-segment. Real cancellation needs the generate_async API.
try:
result = await loop.run_in_executor(None, state.generator.generate, request)
result = await state.pool.run(session.id, request)
except Exception as exc:
logger.exception("session %s: generator failed", session.id[:8])
await _send_error(websocket, "worker_failed", f"generator.generate failed: {exc}", retryable=True)
logger.exception("session %s: pool.run failed", session.id[:8])
await _send_error(websocket, "worker_failed", f"pool.run failed: {exc}", retryable=True)
with contextlib.suppress(InvalidSessionTransition):
session.transition(SessionState.ERROR)
return
@@ -0,0 +1,113 @@
# SPDX-License-Identifier: Apache-2.0
"""Per-session JSONL event logger.
Each session gets its own JSONL file under the configured log root so
post-hoc analytics (enhancer latency, GPU assignment, segment timings)
can be recovered without a tracing backend. The internal UI uses this
format; keeping the same shape makes log tooling portable.
"""
from __future__ import annotations
import contextlib
import json
import os
import re
import threading
import time
from dataclasses import dataclass, field
from typing import Any, TextIO
_FILENAME_SANITIZE_RE = re.compile(r"[^A-Za-z0-9._-]")
@dataclass
class SessionLogEvent:
"""One line in the session JSONL file."""
session_id: str
event: str
payload: dict[str, Any] = field(default_factory=dict)
ts: float = field(default_factory=time.time)
class SessionLogger:
"""Append-only JSONL logger keyed by session id.
Thread-safe; the server may be writing from multiple asyncio tasks
(fMP4 encoder thread + control-frame handler) for the same session.
"""
def __init__(self, log_dir: str | None) -> None:
self._log_dir = log_dir
self._files: dict[str, TextIO] = {}
self._locks: dict[str, threading.Lock] = {}
self._registry_lock = threading.Lock()
self._ensure_dir()
def log(self, event: SessionLogEvent) -> None:
if self._log_dir is None:
return
opened = self._get_file(event.session_id)
if opened is None:
return
handle, lock = opened
line = json.dumps({
"session_id": event.session_id,
"event": event.event,
"ts": event.ts,
"payload": event.payload,
})
with lock, contextlib.suppress(ValueError):
handle.write(line + "\n")
handle.flush()
def close(self, session_id: str) -> None:
with self._registry_lock:
handle = self._files.pop(session_id, None)
lock = self._locks.pop(session_id, None)
if handle is None or lock is None:
return
with lock, contextlib.suppress(Exception):
handle.close()
def close_all(self) -> None:
with self._registry_lock:
sids = list(self._files)
for sid in sids:
self.close(sid)
def _ensure_dir(self) -> None:
if self._log_dir is None:
return
os.makedirs(self._log_dir, exist_ok=True)
def _get_file(self, session_id: str) -> tuple[TextIO, threading.Lock] | None:
if self._log_dir is None:
return None
with self._registry_lock:
handle = self._files.get(session_id)
lock = self._locks.get(session_id)
if handle is not None and lock is not None:
return handle, lock
# Defense-in-depth: session_id is server-generated UUID today,
# but sanitize against path traversal in case future code paths
# allow client-supplied ids.
safe_id = _FILENAME_SANITIZE_RE.sub("_", session_id) or "unknown"
path = os.path.join(
self._log_dir,
f"session-{safe_id}.jsonl",
)
try:
handle = open(path, "a", encoding="utf-8") # noqa: SIM115
except OSError:
return None
lock = threading.Lock()
self._files[session_id] = handle
self._locks[session_id] = lock
return handle, lock
__all__ = [
"SessionLogEvent",
"SessionLogger",
]
+133
View File
@@ -0,0 +1,133 @@
# SPDX-License-Identifier: Apache-2.0
"""Per-GPU worker subprocess entry for :class:`SubprocessGpuPool`.
The pool manages binding, lifecycle, and message dispatch in the parent
process. The worker constructs its :class:`VideoGenerator` from a typed
:class:`GeneratorConfig`, runs the two-segment warmup so both
initial-segment and continuation-branch compile graphs are hot, and
then loops on the job queue.
"""
from __future__ import annotations
import multiprocessing as mp
import queue
from typing import Any
from fastvideo.api.schema import (
GeneratorConfig,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
WarmupConfig,
)
from fastvideo.logger import init_logger
logger = init_logger(__name__)
# Synthetic warmup dimensions: small enough to keep boot fast, big enough
# to exercise the real shape-dependent compile paths. Keep in sync with
# WarmupConfig if these become user-tunable.
_WARMUP_NUM_FRAMES = 8
_WARMUP_HEIGHT = 256
_WARMUP_WIDTH = 256
_WARMUP_NUM_INFERENCE_STEPS = 1
def worker_main(
*,
gpu_id: int,
worker_id: str,
generator_config: GeneratorConfig,
warmup_config: WarmupConfig,
job_queue: mp.Queue,
result_queue: mp.Queue,
shutdown_event: Any,
) -> None: # pragma: no cover - exercised via integration only
"""Per-worker subprocess entry.
Runs inside the child spawned by ``SubprocessGpuPool``. Blocking
``VideoGenerator`` construction + generation happens here, not in
the parent's event loop.
"""
import os
os.environ["CUDA_VISIBLE_DEVICES"] = str(gpu_id)
try:
from fastvideo import VideoGenerator
generator = VideoGenerator.from_pretrained(config=generator_config)
if warmup_config.enabled:
_warmup_worker(generator, warmup_config)
result_queue.put({"kind": "ready", "worker_id": worker_id})
except Exception as exc:
result_queue.put({"kind": "error", "error": repr(exc)})
return
while not shutdown_event.is_set():
try:
item = job_queue.get(timeout=0.5)
except queue.Empty:
continue
if item is None:
break
job_id = item["job_id"]
request = item["request"]
try:
result = generator.generate(request)
result_queue.put({
"kind": "result",
"job_id": job_id,
"result": result,
})
except Exception as exc:
result_queue.put({
"kind": "error",
"job_id": job_id,
"error": repr(exc),
})
def _warmup_worker(
generator: Any,
warmup_config: WarmupConfig,
) -> None:
"""Run two synthetic generations so both compile branches are primed.
Segment 1 is a fresh start (no continuation state) and exercises
the initial-segment graph. Segment 2 feeds segment 1's continuation
state back in so the conditioning branch is also compiled before
the first user request lands.
"""
sampling = SamplingConfig(
num_frames=_WARMUP_NUM_FRAMES,
height=_WARMUP_HEIGHT,
width=_WARMUP_WIDTH,
num_inference_steps=_WARMUP_NUM_INFERENCE_STEPS,
)
seg1 = GenerationRequest(
prompt=warmup_config.prompt,
sampling=sampling,
inputs=InputConfig(),
output=OutputConfig(save_video=False, return_frames=False, return_state=True),
)
seg1_result = generator.generate(seg1)
seg2 = GenerationRequest(
prompt=warmup_config.prompt,
sampling=sampling,
inputs=InputConfig(),
output=OutputConfig(save_video=False, return_frames=False),
state=_extract_continuation_state(seg1_result),
)
generator.generate(seg2)
def _extract_continuation_state(result: Any) -> Any:
state = getattr(result, "state", None)
if state is None and isinstance(result, dict):
state = result.get("state")
return state
__all__ = ["worker_main"]
+113 -2
View File
@@ -33,8 +33,18 @@ from fastvideo.api.compat import (
request_to_pipeline_overrides,
request_to_sampling_param,
)
from fastvideo.api.results import GenerationResult
from fastvideo.api.schema import GenerationRequest, GeneratorConfig
from fastvideo.api.results import (
GenerationResult,
VideoFinalEvent,
VideoProgressEvent,
)
from fastvideo.api.schema import (
GenerationRequest,
GeneratorConfig,
InputConfig,
OutputConfig,
SamplingConfig,
)
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.fastvideo_args import FastVideoArgs
from fastvideo.logger import init_logger
@@ -251,6 +261,76 @@ class VideoGenerator:
if log_queue:
self.executor.clear_log_queue()
async def generate_async(
self,
request: GenerationRequest | Mapping[str, Any],
*,
log_queue=None,
):
"""Async generation that yields typed :class:`VideoEvent`s.
Three consumers share this substrate:
* Streaming server (:mod:`fastvideo.entrypoints.streaming`) —
pipes :class:`VideoPartialEvent` frames into fMP4.
* Stateless OpenAI server — ignores progress events, forwards
:class:`VideoFinalEvent` as the HTTP response body.
* Dynamo native backend
(``components/src/dynamo/fastvideo/``) — wraps each event as
an ``NvVideosResponse`` chunk.
The aggregated code path shipped here yields a single
:class:`VideoProgressEvent` at start and one
:class:`VideoFinalEvent` at end. Future work will thread
per-step progress events through the pipeline's denoise loop
so streaming consumers don't have to wait on a materialized
final.
"""
import asyncio
normalized = normalize_generation_request(request)
total_steps = max(1, normalized.sampling.num_inference_steps)
yield VideoProgressEvent(step=0, total_steps=total_steps, stage="denoise")
if log_queue:
self.executor.set_log_queue(log_queue)
try:
result = await asyncio.to_thread(self._generate_request_impl, normalized)
finally:
if log_queue:
self.executor.clear_log_queue()
if isinstance(result, list):
# Prompt-batch expansion — emit one Final per sub-result.
for sub in result:
yield _final_event_from_result(sub)
return
yield _final_event_from_result(result)
@staticmethod
def default_health_check_request() -> GenerationRequest:
"""Return the minimal typed request Dynamo uses for probes.
256x256, 8 frames, 1 inference step -- fast enough to be a
viable liveness check, non-trivial enough to exercise the
DiT -> VAE -> decode path. Consumers adapt this shape to their
transport's health-check payload (see
``docs/design/server_contracts/dynamo.md``).
"""
return GenerationRequest(
prompt="health check",
inputs=InputConfig(),
sampling=SamplingConfig(
num_frames=8,
height=256,
width=256,
fps=24,
num_inference_steps=1,
guidance_scale=1.0,
),
output=OutputConfig(save_video=False, return_frames=False),
)
def generate_video(
self,
prompt: str | None = None,
@@ -834,3 +914,34 @@ class VideoGenerator:
"""
self.executor.shutdown()
del self.executor
def _final_event_from_result(result: GenerationResult) -> VideoFinalEvent:
"""Build a :class:`VideoFinalEvent` from a terminal result.
Streaming consumers prefer ``frames`` for the MSE-backed path;
server-side consumers prefer encoded ``video_bytes``. We carry both
— whichever the pipeline actually produced — and attach the full
:class:`GenerationResult` so callers that want the complete object
don't have to keep a second reference.
"""
video_bytes: bytes | None = None
if result.video_path and os.path.isfile(result.video_path):
try:
with open(result.video_path, "rb") as f:
video_bytes = f.read()
except OSError:
video_bytes = None
metadata = {
"generation_time": result.generation_time,
"peak_memory_mb": result.peak_memory_mb,
"video_path": result.video_path,
}
return VideoFinalEvent(
video_bytes=video_bytes,
tensor=result.samples,
frames=result.frames,
metadata=metadata,
continuation_state=result.state,
result=result,
)
+97
View File
@@ -133,8 +133,23 @@ class FastVideoArgs:
pin_cpu_memory: bool = True
# Compilation
# ``enable_torch_compile`` covers the DiT path (transformer,
# transformer_2, and the LTX-2 stage-2 transformer_refine).
# Per-component flags below let callers compile additional submodules
# independently; ``False`` leaves the component eager.
enable_torch_compile: bool = False
enable_torch_compile_text_encoder: bool = False
enable_torch_compile_vae: bool = False
enable_torch_compile_audio_vae: bool = False
# ``torch_compile_kwargs`` is the master kwargs dict (applied to every
# compiled submodule unless a per-component dict below is non-empty,
# in which case the per-component dict overrides entirely — matching
# the FastVideo-internal precedent).
torch_compile_kwargs: dict[str, Any] = field(default_factory=dict)
torch_compile_kwargs_dit: dict[str, Any] = field(default_factory=dict)
torch_compile_kwargs_text_encoder: dict[str, Any] = field(default_factory=dict)
torch_compile_kwargs_vae: dict[str, Any] = field(default_factory=dict)
torch_compile_kwargs_audio_vae: dict[str, Any] = field(default_factory=dict)
disable_autocast: bool = False
@@ -161,6 +176,36 @@ class FastVideoArgs:
ltx2_vae_temporal_tile_size_in_frames: int | None = None
ltx2_vae_temporal_tile_overlap_in_frames: int | None = None
ltx2_initial_latent_path: str | None = None
ltx2_audio_latent_path: str | None = None
# Generic stage-2 refine surface (preferred user-facing API). The
# ltx2_refine_* fields below remain the runtime carriers; these
# generic ones let CLI / typed-config callers set the same values
# without binding to a specific model family. ``None`` here means
# "fall back to the model_index.json default and/or the
# ltx2_refine_* runtime carrier".
refine_enabled: bool | None = None
refine_upsampler_path: str | None = None
refine_transformer_path: str | None = None
refine_lora_path: str | None = None
refine_num_inference_steps: int | None = None
refine_guidance_scale: float | None = None
refine_add_noise: bool | None = None
refine_noise_path: str | None = None
refine_audio_noise_path: str | None = None
# LTX-2 stage-2 spatial refinement (the SR pipeline). When enabled the
# transformer runs once at half resolution, the latents are upsampled
# by the LTX2 latent upsampler, then a short stage-2 distilled
# denoising pass refines the upsampled latents. Behaviour is opt-in
# and isolated to LTX-2 today.
ltx2_refine_enabled: bool = False
ltx2_refine_upsampler_path: str | None = None
ltx2_refine_transformer_path: str | None = None
ltx2_refine_lora_path: str | None = None
ltx2_refine_num_inference_steps: int = 3
ltx2_refine_guidance_scale: float = 1.0
ltx2_refine_add_noise: bool = True
ltx2_refine_noise_path: str | None = None
ltx2_refine_audio_noise_path: str | None = None
# model paths for correct deallocation
model_paths: dict[str, str] = field(default_factory=dict)
@@ -172,6 +217,15 @@ class FastVideoArgs:
override_text_encoder_safetensors: str | None = None # path to safetensors file for text encoder override
override_text_encoder_quant: QuantizationMethods = None
# Typed transformer quantization carrier. The typed inference API
# accepts ``engine.quantization.transformer_quant: "NVFP4"`` and the
# compat layer resolves the name to a concrete ``QuantizationConfig``
# instance (e.g. ``NVFP4Config()``); ``__post_init__`` then pins it on
# ``pipeline_config.dit_config.quant_config`` so the loader can detect
# FP4 layers via the standard ``get_quant_method`` path. ``None``
# leaves whatever value the caller already set on ``dit_config``
# untouched.
transformer_quant: Any | None = None
override_transformer_cls_name: str | None = None
init_weights_from_safetensors: str = "" # path to safetensors file for initial weight loading
@@ -199,8 +253,51 @@ class FastVideoArgs:
logger.error("Failed to load V-MoBA config from %s: %s", self.moba_config_path, e)
raise
self._apply_ltx2_vae_overrides()
self._resolve_refine_args()
self._apply_transformer_quant()
self.check_fastvideo_args()
def _apply_transformer_quant(self) -> None:
"""Pin the typed ``transformer_quant`` instance onto ``dit_config``.
``transformer_quant`` is populated by the typed compat layer when
a caller writes ``engine.quantization.transformer_quant: "NVFP4"``
in their config. We pin it here rather than at request time so
the model loader sees the quant_config when constructing the
DiT (linear layers attach their quant_method during ``__init__``).
"""
if self.transformer_quant is None or self.pipeline_config is None:
return
dit_config = getattr(self.pipeline_config, "dit_config", None)
if dit_config is None:
return
# Don't overwrite if the caller already set it explicitly on
# dit_config (e.g. via ``pipeline_config.dit_config.quant_config = NVFP4Config()``);
# the explicit setter wins.
if getattr(dit_config, "quant_config", None) is None:
dit_config.quant_config = self.transformer_quant
def _resolve_refine_args(self) -> None:
"""Map generic refine_* args to LTX-2-specific refine fields."""
if self.refine_enabled is not None:
self.ltx2_refine_enabled = self.refine_enabled
if self.refine_upsampler_path is not None:
self.ltx2_refine_upsampler_path = self.refine_upsampler_path
if self.refine_transformer_path is not None:
self.ltx2_refine_transformer_path = self.refine_transformer_path
if self.refine_lora_path is not None:
self.ltx2_refine_lora_path = self.refine_lora_path
if self.refine_num_inference_steps is not None:
self.ltx2_refine_num_inference_steps = self.refine_num_inference_steps
if self.refine_guidance_scale is not None:
self.ltx2_refine_guidance_scale = self.refine_guidance_scale
if self.refine_add_noise is not None:
self.ltx2_refine_add_noise = self.refine_add_noise
if self.refine_noise_path is not None:
self.ltx2_refine_noise_path = self.refine_noise_path
if self.refine_audio_noise_path is not None:
self.ltx2_refine_audio_noise_path = self.refine_audio_noise_path
def _apply_ltx2_vae_overrides(self) -> None:
if self.pipeline_config is None:
return
+16 -2
View File
@@ -191,7 +191,15 @@ class LinearBase(torch.nn.Module):
if quant_config is None:
self.quant_method: QuantizeMethodBase | None = (UnquantizedLinearMethod())
else:
# ``get_quant_method`` returns ``None`` for layers the config
# has decided not to quantize (e.g. ``NVFP4Config`` only tags
# a curated subset of LTX-2 attention/FFN layers). Fall back
# to ``UnquantizedLinearMethod`` so untagged layers behave
# like a plain ``nn.Linear`` instead of breaking subclass
# asserts.
self.quant_method = quant_config.get_quant_method(self, prefix=prefix)
if self.quant_method is None:
self.quant_method = UnquantizedLinearMethod()
def forward(self, x: torch.Tensor) -> tuple[torch.Tensor, Parameter | None]:
raise NotImplementedError
@@ -230,8 +238,14 @@ class ReplicatedLinear(LinearBase):
prefix=prefix,
)
# All the linear layer supports quant method.
assert self.quant_method is not None
# ``QuantizationConfig.get_quant_method`` may return ``None`` for
# layers it doesn't intend to quantize (e.g. ``NVFP4Config`` only
# tags a specific subset of LTX-2 attention/FFN layers). Fall
# back to ``UnquantizedLinearMethod`` so non-matched layers
# behave like a plain ``nn.Linear``.
if self.quant_method is None:
self.quant_method = UnquantizedLinearMethod()
self.quant_method.create_weights(
self,
self.input_size,
+3 -1
View File
@@ -2,7 +2,7 @@ from typing import Literal, get_args
from fastvideo.layers.quantization.base_config import QuantizationConfig
QuantizationMethods = Literal[None, "AbsMaxFP8"]
QuantizationMethods = Literal[None, "AbsMaxFP8", "NVFP4"]
QUANTIZATION_METHODS: list[str] = list(get_args(QuantizationMethods))
@@ -51,9 +51,11 @@ def get_quantization_config(quantization: str) -> type[QuantizationConfig]:
# lazy import to avoid triggering `torch.compile` too early
from .absmax_fp8 import AbsMaxFP8Config
from .nvfp4_config import NVFP4Config
method_to_config: dict[str, type[QuantizationConfig]] = {
"AbsMaxFP8": AbsMaxFP8Config,
"NVFP4": NVFP4Config,
}
# Update the `method_to_config` with customized quantization methods.
method_to_config.update(_CUSTOMIZED_METHOD_TO_QUANT_CONFIG)
@@ -0,0 +1,474 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2 NVFP4 quantization (FlashInfer-backed).
NVFP4 is NVIDIA's block-scaled FP4 format (e2m1 mantissa, fp32 alpha,
``layout_128x4`` scale layout, group size 16) — distinct from
generic FP4 / OCP-FP4 / MX-FP4. We name the public surface ``NVFP4``
explicitly so downstream callers don't conflate it with other FP4
variants that may land later (e.g. AMD's MX-FP4 or vendor-neutral
e3m0).
Upstreamed from ``FastVideo-internal`` so consumers that load LTX-2
weights with NVFP4 quantization can drive the public package
end-to-end.
`flashinfer` is imported lazily inside the call paths that need it.
This keeps ``import fastvideo`` cheap on hosts where flashinfer is
not installed; only the actual NVFP4 quantize / matmul ops fail at
use time, with a clear error.
"""
from __future__ import annotations
import logging
from typing import Any
import torch
import torch.nn.functional as F
from torch.nn.parameter import Parameter
from fastvideo.layers.quantization.base_config import (
QuantizationConfig,
QuantizeMethodBase,
)
from fastvideo.models.utils import set_weight_attrs
logger = logging.getLogger(__name__)
def _require_flashinfer() -> tuple[Any, Any, Any]:
"""Lazy flashinfer import — raised at use time, not import time.
Returns the bound ``(SfLayout, mm_fp4, nvfp4_quantize)`` triple from
flashinfer. Raises ``ImportError`` with an actionable hint if the
package is not available.
"""
try:
from flashinfer import ( # type: ignore[import-not-found]
SfLayout, mm_fp4, nvfp4_quantize,
)
except ImportError as exc: # pragma: no cover - depends on host env
raise ImportError("NVFP4 quantization requires flashinfer. "
"Install with `pip install flashinfer-python`.") from exc
return SfLayout, mm_fp4, nvfp4_quantize
_LTX2_REFINE_ONLY_SUFFIXES = (
".audio_to_video_attn.to_q",
".video_to_audio_attn.to_k",
".video_to_audio_attn.to_v",
)
def _is_ltx2_refine_only_prefix(prefix: str) -> bool:
return any(prefix.endswith(suffix) for suffix in _LTX2_REFINE_ONLY_SUFFIXES)
def _get_ltx2_fp4_stage_profile(default: str = "refine") -> str:
"""Read the active stage profile from the forward context.
Streaming inference flips between ``base`` and ``refine`` between
segments; the FP4 layer set differs across the two. Falls back to
``default`` whenever the context is not available — this keeps the
op safe to run outside the streaming server (e.g. during eager
tests).
"""
try:
from fastvideo.forward_context import get_forward_context
forward_ctx = get_forward_context()
forward_batch = getattr(forward_ctx, "forward_batch", None)
if forward_batch is None:
return default
extra = getattr(forward_batch, "extra", None)
if not isinstance(extra, dict):
return default
profile = extra.get("ltx2_fp4_stage_profile", default)
if profile in ("base", "refine"):
return profile
return default
except Exception:
return default
_OPS_REGISTERED = False
def _register_ops_once() -> None:
"""Register the fastvideo_fp4 torch ops on first import that needs
them. Each op binds to flashinfer at call time; this just sets up
the dispatcher entries."""
global _OPS_REGISTERED
if _OPS_REGISTERED:
return
@torch.library.custom_op(
"fastvideo_fp4::nvfp4_quantize",
mutates_args=(),
device_types="cuda",
)
def _nvfp4_quantize_op(
x: torch.Tensor,
global_sf: torch.Tensor,
sf_layout: int,
do_shuffle: bool = False,
) -> tuple[torch.Tensor, torch.Tensor]:
SfLayout, _, nvfp4_quantize = _require_flashinfer()
return nvfp4_quantize(x, global_sf, sfLayout=SfLayout(sf_layout), do_shuffle=do_shuffle)
@_nvfp4_quantize_op.register_fake
def _nvfp4_quantize_op_fake(
x: torch.Tensor,
global_sf: torch.Tensor,
sf_layout: int,
do_shuffle: bool = False,
) -> tuple[torch.Tensor, torch.Tensor]:
del global_sf, sf_layout, do_shuffle
m, k = x.shape
quantized = torch.empty((m, (k + 1) // 2), device=x.device, dtype=torch.uint8)
scales = torch.empty((m, (k + 15) // 16), device=x.device, dtype=torch.uint8)
return quantized, scales
@torch.library.custom_op(
"fastvideo_fp4::mm_fp4",
mutates_args=(),
device_types="cuda",
)
def _mm_fp4_op(
a: torch.Tensor,
b: torch.Tensor,
a_scale: torch.Tensor,
b_scale: torch.Tensor,
alpha: torch.Tensor | None,
out_dtype: torch.dtype = torch.bfloat16,
out: torch.Tensor | None = None,
block_size: int = 16,
use_8x4_sf_layout: bool = False,
backend: str = "auto",
use_nvfp4: bool = True,
) -> torch.Tensor:
_, mm_fp4, _ = _require_flashinfer()
if a.dtype == torch.float4_e2m1fn_x2:
a = a.view(torch.uint8) if a.is_contiguous() else a.contiguous().view(torch.uint8)
if b.dtype == torch.float4_e2m1fn_x2:
b = b.view(torch.uint8) if b.is_contiguous() else b.contiguous().view(torch.uint8)
return mm_fp4(
a,
b,
a_scale,
b_scale,
alpha,
out_dtype,
out,
block_size=block_size,
use_8x4_sf_layout=use_8x4_sf_layout,
backend=backend,
use_nvfp4=use_nvfp4,
)
@_mm_fp4_op.register_fake
def _mm_fp4_op_fake(
a: torch.Tensor,
b: torch.Tensor,
a_scale: torch.Tensor,
b_scale: torch.Tensor,
alpha: torch.Tensor | None,
out_dtype: torch.dtype = torch.bfloat16,
out: torch.Tensor | None = None,
block_size: int = 16,
use_8x4_sf_layout: bool = False,
backend: str = "auto",
use_nvfp4: bool = True,
) -> torch.Tensor:
del a_scale, b_scale, alpha, block_size, use_8x4_sf_layout, backend
del use_nvfp4
if out is not None:
return out
out_shape = (*a.shape[:-1], b.shape[1])
return torch.empty(out_shape, device=a.device, dtype=out_dtype)
_OPS_REGISTERED = True
def _nvfp4_quantize(
x: torch.Tensor,
global_sf: Any,
*,
sfLayout: Any,
do_shuffle: bool = False,
) -> tuple[torch.Tensor, torch.Tensor]:
_register_ops_once()
SfLayout, _, _ = _require_flashinfer()
if isinstance(sfLayout, SfLayout):
sf_layout = sfLayout.value
elif hasattr(sfLayout, "value"):
sf_layout = int(sfLayout.value)
else:
sf_layout = int(sfLayout)
if not torch.is_tensor(global_sf):
global_sf = torch.tensor(global_sf, device=x.device, dtype=torch.float32)
elif global_sf.device != x.device:
global_sf = global_sf.to(device=x.device)
if sf_layout == SfLayout.layout_linear.value:
x_for_quant = x
logical_rows = x.shape[0]
else:
# Sequence-parallel can feed either logical rows or row-padded
# rows. Normalize to the kernel tile shape for swizzled layouts
# so both paths share a stable quantization contract.
row_tile = 8 if sf_layout == SfLayout.layout_8x4.value else 128
logical_rows = x.shape[0]
pad_rows = (-logical_rows) % row_tile
x_for_quant = F.pad(x, (0, 0, 0, pad_rows))
quantized, scales = torch.ops.fastvideo_fp4.nvfp4_quantize(x_for_quant, global_sf, sf_layout, do_shuffle)
if sf_layout != SfLayout.layout_linear.value:
quantized = quantized.narrow(0, 0, logical_rows)
return quantized, scales
def _mm_fp4(
a: torch.Tensor,
b: torch.Tensor,
a_scale: torch.Tensor,
b_scale: torch.Tensor,
alpha: Any,
out_dtype: torch.dtype,
out: torch.Tensor | None,
**kwargs: Any,
) -> torch.Tensor:
_register_ops_once()
block_size = kwargs.pop("block_size", 16)
use_8x4_sf_layout = kwargs.pop("use_8x4_sf_layout", False)
backend = kwargs.pop("backend", "auto")
use_nvfp4 = kwargs.pop("use_nvfp4", True)
if kwargs:
raise TypeError(f"Unsupported kwargs for _mm_fp4: {sorted(kwargs)}")
if alpha is not None and not torch.is_tensor(alpha):
alpha = torch.tensor(alpha, device=a.device, dtype=torch.float32)
return torch.ops.fastvideo_fp4.mm_fp4(
a,
b,
a_scale,
b_scale,
alpha,
out_dtype,
out,
block_size,
use_8x4_sf_layout,
backend,
use_nvfp4,
)
class NVFP4QuantizeMethod(QuantizeMethodBase):
def __init__(self, layer_prefix: str = ""):
super().__init__()
self.weight_fp4 = None
self.weight_scale = None
self.x_global_sf = torch.tensor(1.0, device="cuda", dtype=torch.float32)
self.layer_prefix = layer_prefix
self._is_refine_only_layer = _is_ltx2_refine_only_prefix(layer_prefix)
def create_weights(self, layer: torch.nn.Module, input_size_per_partition: int, output_partition_sizes: list[int],
input_size: int, output_size: int, params_dtype: torch.dtype, **extra_weight_attrs):
weight = Parameter(torch.empty(
sum(output_partition_sizes),
input_size_per_partition,
dtype=params_dtype,
),
requires_grad=False)
set_weight_attrs(weight, {"input_dim": 1, "output_dim": 0})
layer.register_parameter("weight", weight)
set_weight_attrs(weight, extra_weight_attrs)
def quantize_input(self, x: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
SfLayout, _, _ = _require_flashinfer()
assert x.dtype == torch.bfloat16 or x.dtype == torch.float16, (
f"only allow bf16/fp16 inputs to fp4 linear, got {x.dtype}")
x_2d = x.view(-1, x.shape[-1])
x_fp4, x_scale = _nvfp4_quantize(
x_2d,
self.x_global_sf,
sfLayout=SfLayout.layout_128x4,
do_shuffle=False,
)
return x_fp4, x_scale, self.x_global_sf
def wants_prequantized_input(self) -> bool:
if not self._is_refine_only_layer:
return True
stage_profile = _get_ltx2_fp4_stage_profile(default="refine")
return stage_profile != "base"
def apply(
self,
layer: torch.nn.Module,
x: torch.Tensor,
bias: torch.Tensor | None = None,
pre_quantized: tuple[torch.Tensor, torch.Tensor, torch.Tensor]
| None = None,
) -> torch.Tensor:
SfLayout, _, _ = _require_flashinfer()
out_dim = layer.weight.shape[0]
original_shape = x.shape
# Stage-aware profile: keep refine-only FP4 layers in dense mode
# during stage-1 denoising so the base path doesn't pay the
# quantize/dequantize tax for layers it never touches.
stage_profile = _get_ltx2_fp4_stage_profile(default="refine")
if self._is_refine_only_layer and stage_profile == "base":
out = (F.linear(x, layer.weight, bias) if torch.cuda.is_available() or bias is None else F.linear(
x, layer.weight, bias.to(x.dtype)))
return out.view(*original_shape[:-1], out_dim)
if pre_quantized is not None:
x_fp4, x_scale, x_global_sf = pre_quantized
# FlashInfer fused norm+quant APIs may return 3D tensors for
# 3D inputs. mm_fp4 only accepts 2D tensors, so flatten
# batch/sequence dims here.
if x_fp4.dim() > 2:
x_fp4 = x_fp4.view(-1, x_fp4.shape[-1])
if x_scale.dim() > 2:
x_scale = x_scale.view(-1, x_scale.shape[-1])
else:
assert x.dtype == torch.bfloat16 or x.dtype == torch.float16, (
f"only allow bf16/fp16 inputs to fp4 linear, got {x.dtype}")
x = x.view(-1, x.shape[-1])
x_global_sf = self.x_global_sf
x_fp4, x_scale = _nvfp4_quantize(
x,
x_global_sf,
sfLayout=SfLayout.layout_128x4,
do_shuffle=False,
)
weight_fp4 = layer._nvfp4_weight
weight_scale = layer._nvfp4_weight_scale
weight_global_sf = layer._weight_global_sf
if hasattr(layer, "_nvfp4_alpha"):
alpha = layer._nvfp4_alpha / x_global_sf
else:
alpha = 1.0 / (x_global_sf * weight_global_sf)
out = _mm_fp4(
x_fp4,
weight_fp4.T,
x_scale,
weight_scale.T,
alpha,
torch.bfloat16,
None,
backend='auto',
)
if bias is not None:
out = out + bias
out = out.view(*original_shape[:-1], out_dim)
return out
class NVFP4Config(QuantizationConfig):
"""LTX-2-specific NVFP4 quantization configuration.
NVFP4 is NVIDIA's block-scaled FP4 (e2m1 mantissa, fp32 alpha,
``layout_128x4`` scale layout, group size 16). Today this class
hardcodes the LTX-2 layer paths it covers. When a second model
wants NVFP4, lift the layer-path list into a config field
instead of hardcoding it here.
"""
def __init__(self, layer_profile: str = "refine"):
super().__init__()
# ``base``: stage-1 set (no attn2.to_out, no cross-modal AV
# projections). ``refine``: full stage-2 set.
self.layer_profile = layer_profile
def get_name(self):
return "nvfp4"
def get_supported_act_dtypes(self):
return [torch.bfloat16, torch.float16]
@classmethod
def get_min_capability(cls):
return 100
@staticmethod
def get_config_filenames():
return []
@classmethod
def from_config(cls, config: dict[str, Any]) -> NVFP4Config:
return cls(layer_profile=config.get("layer_profile", "refine"))
def get_quant_method(self, layer: torch.nn.Module, prefix: str):
from fastvideo.layers.linear import LinearBase
# Use the superset at build/load time, then switch active subset
# dynamically in NVFP4QuantizeMethod.apply based on stage profile.
fp4_layers = [[
f"ltx2.blocks.{i}.attn1.to_q",
f"ltx2.blocks.{i}.attn1.to_k",
f"ltx2.blocks.{i}.attn1.to_v",
f"ltx2.blocks.{i}.attn1.to_out",
f"ltx2.blocks.{i}.attn2.to_q",
f"ltx2.blocks.{i}.attn2.to_out",
f"ltx2.blocks.{i}.audio_to_video_attn.to_q",
f"ltx2.blocks.{i}.audio_to_video_attn.to_out",
f"ltx2.blocks.{i}.video_to_audio_attn.to_k",
f"ltx2.blocks.{i}.video_to_audio_attn.to_v",
f"ltx2.blocks.{i}.ffn.fc_in",
f"ltx2.blocks.{i}.ffn.fc_out",
] for i in range(48)]
fp4_layers.append([
"ltx2.adaln_single.linear",
])
if isinstance(layer, LinearBase) and any(prefix in layer_names for layer_names in fp4_layers):
return NVFP4QuantizeMethod(layer_prefix=prefix)
return None
def convert_model_to_nvfp4(model: torch.nn.Module) -> None:
SfLayout, _, _ = _require_flashinfer()
from torch.distributed.tensor import DTensor # type: ignore
for mod in model.modules():
qm = getattr(mod, "quant_method", None)
if isinstance(qm, NVFP4QuantizeMethod):
weight = getattr(mod, "weight", None)
if weight is None:
continue
weight_local = weight.to_local() if isinstance(weight, DTensor) else weight # type: ignore[arg-type]
weight_global_sf = (448 * 6) / weight_local.float().abs().nan_to_num().max()
fp4_w, fp4_s = _nvfp4_quantize(
weight_local,
weight_global_sf,
sfLayout=SfLayout.layout_128x4,
do_shuffle=False,
)
weight_global_sf_t = torch.as_tensor(
weight_global_sf,
device=weight_local.device,
dtype=torch.float32,
)
mod.register_buffer("_nvfp4_weight", fp4_w, persistent=False)
mod.register_buffer("_nvfp4_weight_scale", fp4_s, persistent=False)
mod.register_buffer(
"_weight_global_sf",
weight_global_sf_t.to(dtype=torch.bfloat16),
persistent=False,
)
mod.register_buffer(
"_nvfp4_alpha",
(1.0 / weight_global_sf_t).to(dtype=torch.float32),
persistent=False,
)
__all__ = [
"NVFP4Config",
"NVFP4QuantizeMethod",
"convert_model_to_nvfp4",
]
+234 -49
View File
@@ -28,6 +28,8 @@ from fastvideo.distributed.communication_op import (
)
from fastvideo.distributed.parallel_state import get_sp_parallel_rank, get_sp_world_size
from fastvideo.forward_context import ForwardContext, get_forward_context, set_forward_context
from fastvideo.layers.linear import ReplicatedLinear
from fastvideo.layers.quantization.base_config import QuantizationConfig
from fastvideo.logger import init_logger
from fastvideo.models.dits.base import BaseDiT
from fastvideo.platforms import AttentionBackendEnum
@@ -35,6 +37,50 @@ from fastvideo.platforms import AttentionBackendEnum
logger = init_logger(__name__)
def _supports_prequantized_input(linear: ReplicatedLinear) -> bool:
"""Whether ``linear``'s quant method is willing to accept a
pre-quantized ``(x_fp4, x_scale, x_global_sf)`` input tuple.
Used by the LTX-2 attention forward path so that a single input
tensor can be quantized once and reused across q/k/v projections
when they share the same source (self-attention).
"""
quant_method = getattr(linear, "quant_method", None)
if not callable(getattr(quant_method, "quantize_input", None)):
return False
wants_prequant = getattr(quant_method, "wants_prequantized_input", None)
if callable(wants_prequant):
try:
return bool(wants_prequant())
except Exception:
return False
return True
def _linear_project_with_optional_prequant(
linear: ReplicatedLinear,
x: torch.Tensor,
pre_quantized: tuple[torch.Tensor, torch.Tensor, torch.Tensor] | None,
) -> torch.Tensor:
"""Project ``x`` through ``linear``, optionally bypassing the
in-method quantize step when ``pre_quantized`` is supplied.
When the linear's quant method does not support pre-quantized
inputs (e.g. ``UnquantizedLinearMethod``), falls back to the
standard ``linear(x)`` call and discards the bias-pass-through
tuple element.
"""
if pre_quantized is None or not _supports_prequantized_input(linear):
return linear(x)[0]
bias = linear.bias if not linear.skip_bias_add else None
return linear.quant_method.apply( # type: ignore[union-attr]
linear,
x,
bias=bias,
pre_quantized=pre_quantized,
)
def get_timestep_embedding(
timesteps: torch.Tensor,
embedding_dim: int,
@@ -172,25 +218,57 @@ class PixArtAlphaTextProjection(torch.nn.Module):
class GELUApprox(nn.Module):
"""Linear + tanh-approximate GELU used by LTX-2 FFN."""
def __init__(self, in_features: int, out_features: int):
def __init__(
self,
in_features: int,
out_features: int,
quant_config: QuantizationConfig | None = None,
prefix: str = "",
):
super().__init__()
self.proj = nn.Linear(in_features, out_features)
self.proj = ReplicatedLinear(
in_features,
out_features,
quant_config=quant_config,
prefix=f"{prefix}.fc_in",
)
self.act = nn.GELU(approximate="tanh")
def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.act(self.proj(x))
return self.act(self.proj(x)[0])
class FeedForward(nn.Module):
"""LTX-2 FFN: GELUApprox -> Identity -> Linear."""
def __init__(self, dim: int, dim_out: int, mult: int = 4) -> None:
def __init__(
self,
dim: int,
dim_out: int,
mult: int = 4,
quant_config: QuantizationConfig | None = None,
prefix: str = "",
) -> None:
super().__init__()
inner_dim = int(dim * mult)
project_in = GELUApprox(dim, inner_dim)
self.net = nn.Sequential(project_in, nn.Identity(), nn.Linear(inner_dim, dim_out))
project_in = GELUApprox(
dim,
inner_dim,
quant_config=quant_config,
prefix=f"{prefix}.ffn",
)
project_out = ReplicatedLinear(
inner_dim,
dim_out,
quant_config=quant_config,
prefix=f"{prefix}.ffn.fc_out",
)
self.net = nn.ModuleList([project_in, nn.Identity(), project_out])
def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.net(x)
x = self.net[0](x)
x = self.net[1](x)
x = self.net[2](x)[0]
return x
class VideoLatentShape(tuple):
@@ -1246,6 +1324,8 @@ class LTXSelfAttention(nn.Module):
norm_eps: float,
rope_type: LTXRopeType,
supported_attention_backends: tuple[AttentionBackendEnum, ...],
quant_config: QuantizationConfig | None = None,
prefix: str = "",
) -> None:
super().__init__()
inner_dim = dim_head * heads
@@ -1257,10 +1337,37 @@ class LTXSelfAttention(nn.Module):
self.q_norm = torch.nn.RMSNorm(inner_dim, eps=norm_eps)
self.k_norm = torch.nn.RMSNorm(inner_dim, eps=norm_eps)
self.to_q = nn.Linear(query_dim, inner_dim, bias=True)
self.to_k = nn.Linear(context_dim, inner_dim, bias=True)
self.to_v = nn.Linear(context_dim, inner_dim, bias=True)
self.to_out = nn.Sequential(nn.Linear(inner_dim, query_dim, bias=True), nn.Identity())
self.to_q = ReplicatedLinear(
query_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_q",
)
self.to_k = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_k",
)
self.to_v = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_v",
)
self.to_out = nn.ModuleList([
ReplicatedLinear(
inner_dim,
query_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_out",
),
nn.Identity(),
])
self.attn = LTXLocalAttention(
num_heads=heads,
@@ -1271,9 +1378,15 @@ class LTXSelfAttention(nn.Module):
causal=False,
supported_attention_backends=supported_attention_backends,
)
self.to_gate_compress: nn.Linear | None = None
self.to_gate_compress: ReplicatedLinear | None = None
if self.attn.backend == AttentionBackendEnum.VIDEO_SPARSE_ATTN:
self.to_gate_compress = nn.Linear(context_dim, inner_dim, bias=True)
self.to_gate_compress = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_gate_compress",
)
self.attn_masked = LTXLocalAttention(
num_heads=heads,
head_size=dim_head,
@@ -1292,11 +1405,24 @@ class LTXSelfAttention(nn.Module):
pe: tuple[torch.Tensor, torch.Tensor] | None = None,
k_pe: tuple[torch.Tensor, torch.Tensor] | None = None,
) -> torch.Tensor:
q = self.to_q(x)
context = x if context is None else context
k = self.to_k(context)
v = self.to_v(context)
gate_compress = (self.to_gate_compress(context)
q_pre_quantized = (
self.to_q.quant_method.quantize_input(x) # type: ignore[union-attr]
if _supports_prequantized_input(self.to_q) else None)
kv_pre_quantized = None
if (_supports_prequantized_input(self.to_k)
and _supports_prequantized_input(self.to_v)):
if context is x and q_pre_quantized is not None:
kv_pre_quantized = q_pre_quantized
else:
kv_pre_quantized = self.to_k.quant_method.quantize_input( # type: ignore[union-attr]
context)
q = _linear_project_with_optional_prequant(self.to_q, x, q_pre_quantized)
k = _linear_project_with_optional_prequant(self.to_k, context, kv_pre_quantized)
v = _linear_project_with_optional_prequant(self.to_v, context, kv_pre_quantized)
gate_compress = (self.to_gate_compress(context)[0]
if self.to_gate_compress is not None else None)
q = self.q_norm(q)
@@ -1346,7 +1472,8 @@ class LTXSelfAttention(nn.Module):
ltx_freqs_cis=pe,
ltx_k_freqs_cis=k_pe)
out = out.reshape(b, q_len, -1)
return self.to_out(out)
out = self.to_out[0](out)[0]
return self.to_out[1](out)
class LTXDistributedSelfAttention(nn.Module):
@@ -1361,6 +1488,7 @@ class LTXDistributedSelfAttention(nn.Module):
norm_eps: float,
rope_type: LTXRopeType,
supported_attention_backends: tuple[AttentionBackendEnum, ...],
quant_config: QuantizationConfig | None = None,
prefix: str = "",
) -> None:
super().__init__()
@@ -1373,10 +1501,37 @@ class LTXDistributedSelfAttention(nn.Module):
self.q_norm = torch.nn.RMSNorm(inner_dim, eps=norm_eps)
self.k_norm = torch.nn.RMSNorm(inner_dim, eps=norm_eps)
self.to_q = nn.Linear(query_dim, inner_dim, bias=True)
self.to_k = nn.Linear(context_dim, inner_dim, bias=True)
self.to_v = nn.Linear(context_dim, inner_dim, bias=True)
self.to_out = nn.Sequential(nn.Linear(inner_dim, query_dim, bias=True), nn.Identity())
self.to_q = ReplicatedLinear(
query_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_q",
)
self.to_k = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_k",
)
self.to_v = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_v",
)
self.to_out = nn.ModuleList([
ReplicatedLinear(
inner_dim,
query_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_out",
),
nn.Identity(),
])
self.attn = LTXDistributedAttention(
num_heads=heads,
@@ -1386,9 +1541,15 @@ class LTXDistributedSelfAttention(nn.Module):
supported_attention_backends=supported_attention_backends,
prefix=f"{prefix}.attn",
)
self.to_gate_compress: nn.Linear | None = None
self.to_gate_compress: ReplicatedLinear | None = None
if self.attn.backend == AttentionBackendEnum.VIDEO_SPARSE_ATTN:
self.to_gate_compress = nn.Linear(context_dim, inner_dim, bias=True)
self.to_gate_compress = ReplicatedLinear(
context_dim,
inner_dim,
bias=True,
quant_config=quant_config,
prefix=f"{prefix}.to_gate_compress",
)
def forward(
self,
@@ -1408,11 +1569,24 @@ class LTXDistributedSelfAttention(nn.Module):
k_pe: Rotary position embeddings for K (cos, sin), or None to use pe
original_seq_len: Original (unpadded) full sequence length
"""
q = self.to_q(x)
context = x if context is None else context
k = self.to_k(context)
v = self.to_v(context)
gate_compress = (self.to_gate_compress(context)
q_pre_quantized = (
self.to_q.quant_method.quantize_input(x) # type: ignore[union-attr]
if _supports_prequantized_input(self.to_q) else None)
kv_pre_quantized = None
if (_supports_prequantized_input(self.to_k)
and _supports_prequantized_input(self.to_v)):
if context is x and q_pre_quantized is not None:
kv_pre_quantized = q_pre_quantized
else:
kv_pre_quantized = self.to_k.quant_method.quantize_input( # type: ignore[union-attr]
context)
q = _linear_project_with_optional_prequant(self.to_q, x, q_pre_quantized)
k = _linear_project_with_optional_prequant(self.to_k, context, kv_pre_quantized)
v = _linear_project_with_optional_prequant(self.to_v, context, kv_pre_quantized)
gate_compress = (self.to_gate_compress(context)[0]
if self.to_gate_compress is not None else None)
q = self.q_norm(q)
@@ -1443,7 +1617,8 @@ class LTXDistributedSelfAttention(nn.Module):
)
out = out.reshape(b, q_len, -1)
return self.to_out(out)
out = self.to_out[0](out)[0]
return self.to_out[1](out)
class BasicAVTransformerBlock(torch.nn.Module):
@@ -1457,6 +1632,7 @@ class BasicAVTransformerBlock(torch.nn.Module):
rope_type: LTXRopeType = LTXRopeType.INTERLEAVED,
norm_eps: float = 1e-6,
use_distributed_attention: bool = False,
quant_config: QuantizationConfig | None = None,
prefix: str = "",
):
super().__init__()
@@ -1488,15 +1664,8 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=video_self_attn_backends,
prefix=f"{prefix}.blocks.{idx}.attn1" if use_distributed_attention else "",
) if use_distributed_attention else LTXSelfAttention(
query_dim=video.dim,
context_dim=None,
heads=video.heads,
dim_head=video.d_head,
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=video_self_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.attn1",
)
# Text cross-attention - always local (text is replicated)
self.attn2 = CrossAttnCls(
@@ -1507,8 +1676,15 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.attn2",
)
self.ff = FeedForward(
video.dim,
dim_out=video.dim,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}",
)
self.ff = FeedForward(video.dim, dim_out=video.dim)
self.scale_shift_table = torch.nn.Parameter(torch.empty(6, video.dim))
if audio is not None:
@@ -1521,15 +1697,8 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
prefix=f"{prefix}.blocks.{idx}.audio_attn1" if use_distributed_attention else "",
) if use_distributed_attention else LTXSelfAttention(
query_dim=audio.dim,
context_dim=None,
heads=audio.heads,
dim_head=audio.d_head,
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.audio_attn1",
)
# Text cross-attention - always local (text is replicated)
self.audio_attn2 = CrossAttnCls(
@@ -1540,8 +1709,15 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.audio_attn2",
)
self.audio_ff = FeedForward(
audio.dim,
dim_out=audio.dim,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.audio",
)
self.audio_ff = FeedForward(audio.dim, dim_out=audio.dim)
self.audio_scale_shift_table = torch.nn.Parameter(torch.empty(6, audio.dim))
if audio is not None and video is not None:
@@ -1555,6 +1731,8 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.audio_to_video_attn",
)
# Video-to-audio cross-attention
# Uses local attention - context is gathered from all SP ranks in forward()
@@ -1566,6 +1744,8 @@ class BasicAVTransformerBlock(torch.nn.Module):
norm_eps=norm_eps,
rope_type=rope_type,
supported_attention_backends=dense_attn_backends,
quant_config=quant_config,
prefix=f"{prefix}.blocks.{idx}.video_to_audio_attn",
)
self.scale_shift_table_a2v_ca_audio = torch.nn.Parameter(torch.empty(5, audio.dim))
self.scale_shift_table_a2v_ca_video = torch.nn.Parameter(torch.empty(5, video.dim))
@@ -1889,6 +2069,7 @@ class LTXModel(torch.nn.Module):
rope_type: LTXRopeType = LTXRopeType.INTERLEAVED,
double_precision_rope: bool = False,
use_distributed_attention: bool = False,
quant_config: QuantizationConfig | None = None,
prefix: str = "",
):
super().__init__()
@@ -1942,6 +2123,7 @@ class LTXModel(torch.nn.Module):
audio_cross_attention_dim=audio_cross_attention_dim,
norm_eps=norm_eps,
use_distributed_attention=use_distributed_attention,
quant_config=quant_config,
prefix=prefix,
)
@@ -2073,6 +2255,7 @@ class LTXModel(torch.nn.Module):
audio_cross_attention_dim: int,
norm_eps: float,
use_distributed_attention: bool = False,
quant_config: QuantizationConfig | None = None,
prefix: str = "",
) -> None:
video_config = (
@@ -2105,6 +2288,7 @@ class LTXModel(torch.nn.Module):
rope_type=self.rope_type,
norm_eps=norm_eps,
use_distributed_attention=use_distributed_attention,
quant_config=quant_config,
prefix=prefix,
)
for idx in range(num_layers)
@@ -2293,6 +2477,7 @@ class LTX2Transformer3DModel(BaseDiT):
audio_positional_embedding_max_pos=arch.audio_positional_embedding_max_pos,
av_ca_timestep_scale_multiplier=arch.av_ca_timestep_scale_multiplier,
use_distributed_attention=use_distributed_attention,
quant_config=config.quant_config,
prefix=config.prefix,
)
+13 -6
View File
@@ -361,6 +361,10 @@ class LTX2GemmaTextEncoderModel(TextEncoder):
continue
yield name, param
def prepare_for_compile(self) -> None:
# Load Gemma outside Dynamo so torch.compile does not trace HF file-system checks.
_ = self.gemma_model
@property
def gemma_model(self) -> Gemma3ForConditionalGeneration:
if self._gemma_model is None:
@@ -517,18 +521,21 @@ class LTX2GemmaTextEncoderModel(TextEncoder):
attention_mask = torch.ones_like(input_ids)
model = self.gemma_model
orig_device = model.device
model.to(device=get_local_torch_device())
# input_ids = input_ids.to(device=model.device)
# attention_mask = attention_mask.to(device=model.device)
target_device = get_local_torch_device()
# Do not invoke model.to() inside the compiled forward path.
# _parse_to returns a non-Tensor torch.device, which Dynamo cannot
# trace under fullgraph=True. The model is already moved to device
# when first loaded (see gemma_model property + prepare_for_compile),
# so this guard is a runtime no-op and Dynamo can DCE it.
if model.device != target_device:
model.to(device=target_device)
outputs = model(
input_ids=input_ids,
attention_mask=attention_mask,
output_hidden_states=True,
return_dict=True,
)
model.to(device=orig_device)
encoded_inputs = self._run_feature_extractor(
outputs.hidden_states,
attention_mask,
+49 -14
View File
@@ -99,6 +99,12 @@ class ComponentLoader(ABC):
# NumberConditioners; not a pure text encoder, so it gets
# its own loader.
"conditioner": (ConditionerLoader, "fastvideo"),
# LTX-2 spatial / temporal upsamplers — share the
# UpsamplerLoader path with the upsampler/upsampler_2 keys
# so the SR pipeline picks up real weights instead of the
# generic config-only loader.
"spatial_upsampler": (UpsamplerLoader, "diffusers"),
"temporal_upsampler": (UpsamplerLoader, "diffusers"),
}
if module_type in module_loaders:
@@ -1058,7 +1064,7 @@ class ConditionerLoader(ComponentLoader):
class UpsamplerLoader(ComponentLoader):
"""Loader for upsamplers."""
"""Loader for upsamplers (incl. LTX-2 spatial/temporal upsamplers)."""
def load(self, model_path: str, fastvideo_args: FastVideoArgs):
"""Load the upsampler based on the model path, and inference args."""
@@ -1068,36 +1074,65 @@ class UpsamplerLoader(ComponentLoader):
if class_name is None:
raise ValueError(
"Model config does not contain a _class_name attribute. "
"Only diffusers format is supported."
)
"Only diffusers format is supported.")
try:
upsampler_cfg = deepcopy(fastvideo_args.pipeline_config.upsampler_config[0])
upsampler_cfg.update_model_config(config_dict)
except Exception as e:
upsampler_cfg = deepcopy(fastvideo_args.pipeline_config.upsampler_config[1])
upsampler_cfg.update_model_config(config_dict)
# The base PipelineConfig declares ``upsampler_config`` as a
# single ``UpsamplerConfig`` instance, but Hunyuan15 narrows it
# to a tuple of two configs (one per SR target). We only treat
# the attribute as a multi-config when it actually is one;
# otherwise the LTX-2 branch below handles the single-class
# path that takes the diffusers config dict directly.
upsampler_config_attr = getattr(fastvideo_args.pipeline_config,
"upsampler_config", None)
if isinstance(upsampler_config_attr, list | tuple):
try:
upsampler_cfg = deepcopy(upsampler_config_attr[0])
upsampler_cfg.update_model_config(config_dict)
except Exception:
upsampler_cfg = deepcopy(upsampler_config_attr[1])
upsampler_cfg.update_model_config(config_dict)
elif class_name == "LTX2LatentUpsampler":
# LTX-2 pipeline_config does not declare upsampler_config; the
# `LTX2LatentUpsampler` wrapper takes the raw diffusers config
# dict directly via LatentUpsamplerConfigurator.
upsampler_cfg = deepcopy(config_dict)
else:
raise AttributeError(
"pipeline_config.upsampler_config is missing; cannot build "
f"upsampler config for class {class_name}")
model_cls, _ = ModelRegistry.resolve_model_cls(class_name)
model = model_cls(upsampler_cfg)
target_device = get_local_torch_device()
model = model.to(target_device, dtype=PRECISION_TO_TYPE[fastvideo_args.pipeline_config.upsampler_precision])
upsampler_precision = getattr(fastvideo_args.pipeline_config,
"upsampler_precision", "bf16")
model = model.to(target_device,
dtype=PRECISION_TO_TYPE[upsampler_precision])
# Find all safetensors files
safetensors_list = glob.glob(
os.path.join(str(model_path), "*.safetensors"))
if not safetensors_list:
raise ValueError(f"No safetensors files found in {model_path}")
if len(safetensors_list) == 1:
loaded = safetensors_load_file(safetensors_list[0])
else:
loaded = {}
for sf_file in safetensors_list:
loaded.update(safetensors_load_file(sf_file))
model.load_state_dict(loaded, strict=True)
# The LTX-2 latent upsampler wrapper exposes the actual conv
# stack at ``self.model``; checkpoint state_dicts may be saved
# without the ``model.`` prefix when the inner module was
# serialised directly. Strip / forward as needed so both layouts
# load cleanly.
target_module = getattr(model, "model", model)
if loaded and all(k.startswith("model.") for k in loaded):
stripped = {k[len("model."):]: v for k, v in loaded.items()}
target_module.load_state_dict(stripped, strict=True)
else:
target_module.load_state_dict(loaded, strict=True)
return model.eval()
+40
View File
@@ -28,6 +28,37 @@ from fastvideo.utils import set_mixed_precision_policy, is_pin_memory_available
logger = init_logger(__name__)
def _maybe_convert_model_to_nvfp4(model: nn.Module) -> None:
"""Quantize NVFP4-tagged linear layers in-place after weights are loaded.
Walks the module tree once, looking for layers whose ``quant_method``
is an :class:`NVFP4QuantizeMethod` (attached at construction time by
:meth:`NVFP4Config.get_quant_method`). When at least one such layer
exists, calls :func:`convert_model_to_nvfp4` to register the
``_nvfp4_weight*`` / ``_nvfp4_alpha`` / ``_weight_global_sf`` buffers
on each targeted layer.
The walk returns on the first NVFP4 layer found so non-NVFP4 callers
pay only an ``isinstance`` check per module. flashinfer is imported
lazily inside :func:`convert_model_to_nvfp4` so this helper is a
no-op on hosts without the NVFP4 backend.
"""
# Defer the import: nvfp4_config imports heavy diffusers /
# torch.distributed symbols at module-load time, and unconditional
# import would penalize every loader call regardless of whether
# NVFP4 is wired.
from fastvideo.layers.quantization.nvfp4_config import (
NVFP4QuantizeMethod, convert_model_to_nvfp4,
)
for mod in model.modules():
if isinstance(getattr(mod, "quant_method", None),
NVFP4QuantizeMethod):
logger.info("Converting loaded model weights for NVFP4 linear layers")
convert_model_to_nvfp4(model)
return
# TODO(PY): move this to utils elsewhere
@contextlib.contextmanager
def set_default_dtype(dtype: torch.dtype) -> Generator[None, None, None]:
@@ -158,6 +189,15 @@ def maybe_load_fsdp_model(
if isinstance(p, torch.nn.Parameter):
p.requires_grad = False
# NVFP4 weight prequantization. We detect by the registered
# ``quant_method`` on linear layers rather than by a separate flag —
# construction-time ``NVFP4Config.get_quant_method`` already attached
# ``NVFP4QuantizeMethod`` to every targeted layer, so the loader's
# responsibility is just to materialize the per-layer nvfp4 weight /
# scale buffers from the freshly-loaded bf16 weights. No-op when
# ``flashinfer`` is not installed (lazy import inside the helper).
_maybe_convert_model_to_nvfp4(model)
compile_in_loader = enable_torch_compile and training_mode
if compile_in_loader:
compile_kwargs = torch_compile_kwargs or {}
+2
View File
@@ -115,6 +115,8 @@ _SCHEDULERS = {
_UPSAMPLERS = {
"SRTo720pUpsampler": ("upsamplers", "hunyuan15", "SRTo720pUpsampler"),
"SRTo1080pUpsampler": ("upsamplers", "hunyuan15", "SRTo1080pUpsampler"),
"LTX2LatentUpsampler":
("upsamplers", "ltx2_upsampler", "LTX2LatentUpsampler"),
}
_LEGACY_FAST_VIDEO_MODELS = {
+23
View File
@@ -0,0 +1,23 @@
# SPDX-License-Identifier: Apache-2.0
from fastvideo.models.upsamplers.ltx2_upsampler import (
BlurDownsample,
LTX2LatentUpsampler,
LatentUpsampler,
LatentUpsamplerConfigurator,
PixelShuffleND,
ResBlock,
SpatialRationalResampler,
upsample_video,
)
__all__ = [
"BlurDownsample",
"LTX2LatentUpsampler",
"LatentUpsampler",
"LatentUpsamplerConfigurator",
"PixelShuffleND",
"ResBlock",
"SpatialRationalResampler",
"upsample_video",
]
@@ -0,0 +1,319 @@
# SPDX-License-Identifier: Apache-2.0
"""
LTX-2 latent upsampler (spatial/temporal) implementation.
"""
from __future__ import annotations
import math
from typing import Any, Optional, Tuple
import torch
import torch.nn as nn
from einops import rearrange
class PixelShuffleND(nn.Module):
"""N-dimensional pixel shuffle for upsampling."""
def __init__(self, dims: int, upscale_factors: Tuple[int, int, int] = (2, 2, 2)) -> None:
super().__init__()
if dims not in (1, 2, 3):
raise ValueError("dims must be 1, 2, or 3")
self.dims = dims
self.upscale_factors = upscale_factors
def forward(self, x: torch.Tensor) -> torch.Tensor:
if self.dims == 3:
return rearrange(
x,
"b (c p1 p2 p3) d h w -> b c (d p1) (h p2) (w p3)",
p1=self.upscale_factors[0],
p2=self.upscale_factors[1],
p3=self.upscale_factors[2],
)
if self.dims == 2:
return rearrange(
x,
"b (c p1 p2) h w -> b c (h p1) (w p2)",
p1=self.upscale_factors[0],
p2=self.upscale_factors[1],
)
if self.dims == 1:
return rearrange(
x,
"b (c p1) f h w -> b c (f p1) h w",
p1=self.upscale_factors[0],
)
raise ValueError(f"Unsupported dims: {self.dims}")
class BlurDownsample(nn.Module):
"""
Anti-aliased spatial downsampling by integer stride using a fixed separable binomial kernel.
Applies only on H,W. Works for dims=2 or dims=3 (per-frame).
"""
def __init__(self, dims: int, stride: int, kernel_size: int = 5) -> None:
super().__init__()
if dims not in (2, 3):
raise ValueError("dims must be 2 or 3")
if stride < 1:
raise ValueError("stride must be >= 1")
if kernel_size < 3 or kernel_size % 2 != 1:
raise ValueError("kernel_size must be an odd integer >= 3")
self.dims = dims
self.stride = stride
self.kernel_size = kernel_size
k = torch.tensor([math.comb(kernel_size - 1, idx) for idx in range(kernel_size)])
k2d = k[:, None] @ k[None, :]
k2d = (k2d / k2d.sum()).float()
self.register_buffer("kernel", k2d[None, None, :, :])
def forward(self, x: torch.Tensor) -> torch.Tensor:
if self.stride == 1:
return x
if self.dims == 2:
return self._apply_2d(x)
b, _, f, _, _ = x.shape
x = rearrange(x, "b c f h w -> (b f) c h w")
x = self._apply_2d(x)
h2, w2 = x.shape[-2:]
return rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f, h=h2, w=w2)
def _apply_2d(self, x2d: torch.Tensor) -> torch.Tensor:
c = x2d.shape[1]
weight = self.kernel.expand(c, 1, self.kernel_size, self.kernel_size)
return nn.functional.conv2d(
x2d,
weight=weight,
bias=None,
stride=self.stride,
padding=self.kernel_size // 2,
groups=c,
)
def _rational_for_scale(scale: float) -> Tuple[int, int]:
mapping = {0.75: (3, 4), 1.5: (3, 2), 2.0: (2, 1), 4.0: (4, 1)}
if float(scale) not in mapping:
raise ValueError(f"Unsupported scale {scale}. Choose from {list(mapping.keys())}")
return mapping[float(scale)]
class SpatialRationalResampler(nn.Module):
"""
Fully-learned rational spatial scaling: up by 'num' via PixelShuffle, then
anti-aliased downsample by 'den' using fixed blur + stride. Operates on H,W only.
For dims==3, work per-frame for spatial scaling (temporal axis untouched).
"""
def __init__(self, mid_channels: int, scale: float) -> None:
super().__init__()
self.scale = float(scale)
self.num, self.den = _rational_for_scale(self.scale)
self.conv = nn.Conv2d(mid_channels, (self.num**2) * mid_channels, kernel_size=3, padding=1)
self.pixel_shuffle = PixelShuffleND(2, upscale_factors=(self.num, self.num))
self.blur_down = BlurDownsample(dims=2, stride=self.den)
def forward(self, x: torch.Tensor) -> torch.Tensor:
b, _, f, _, _ = x.shape
x = rearrange(x, "b c f h w -> (b f) c h w")
x = self.conv(x)
x = self.pixel_shuffle(x)
x = self.blur_down(x)
return rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f)
class ResBlock(nn.Module):
"""Residual block with two convolutional layers, group norm, and SiLU."""
def __init__(self, channels: int, mid_channels: Optional[int] = None, dims: int = 3) -> None:
super().__init__()
if mid_channels is None:
mid_channels = channels
conv = nn.Conv2d if dims == 2 else nn.Conv3d
self.conv1 = conv(channels, mid_channels, kernel_size=3, padding=1)
self.norm1 = nn.GroupNorm(32, mid_channels)
self.conv2 = conv(mid_channels, channels, kernel_size=3, padding=1)
self.norm2 = nn.GroupNorm(32, channels)
self.activation = nn.SiLU()
def forward(self, x: torch.Tensor) -> torch.Tensor:
residual = x
x = self.conv1(x)
x = self.norm1(x)
x = self.activation(x)
x = self.conv2(x)
x = self.norm2(x)
x = self.activation(x + residual)
return x
class LatentUpsampler(nn.Module):
"""
Model to upsample VAE latents spatially and/or temporally.
"""
def __init__(
self,
in_channels: int = 128,
mid_channels: int = 512,
num_blocks_per_stage: int = 4,
dims: int = 3,
spatial_upsample: bool = True,
temporal_upsample: bool = False,
spatial_scale: float = 2.0,
rational_resampler: bool = False,
) -> None:
super().__init__()
self.in_channels = in_channels
self.mid_channels = mid_channels
self.num_blocks_per_stage = num_blocks_per_stage
self.dims = dims
self.spatial_upsample = spatial_upsample
self.temporal_upsample = temporal_upsample
self.spatial_scale = float(spatial_scale)
self.rational_resampler = rational_resampler
conv = nn.Conv2d if dims == 2 else nn.Conv3d
self.initial_conv = conv(in_channels, mid_channels, kernel_size=3, padding=1)
self.initial_norm = nn.GroupNorm(32, mid_channels)
self.initial_activation = nn.SiLU()
self.res_blocks = nn.ModuleList([ResBlock(mid_channels, dims=dims) for _ in range(num_blocks_per_stage)])
if spatial_upsample and temporal_upsample:
self.upsampler = nn.Sequential(
nn.Conv3d(mid_channels, 8 * mid_channels, kernel_size=3, padding=1),
PixelShuffleND(3),
)
elif spatial_upsample:
if rational_resampler:
self.upsampler = SpatialRationalResampler(mid_channels=mid_channels, scale=self.spatial_scale)
else:
self.upsampler = nn.Sequential(
nn.Conv2d(mid_channels, 4 * mid_channels, kernel_size=3, padding=1),
PixelShuffleND(2),
)
elif temporal_upsample:
self.upsampler = nn.Sequential(
nn.Conv3d(mid_channels, 2 * mid_channels, kernel_size=3, padding=1),
PixelShuffleND(1),
)
else:
raise ValueError("Either spatial_upsample or temporal_upsample must be True")
self.post_upsample_res_blocks = nn.ModuleList(
[ResBlock(mid_channels, dims=dims) for _ in range(num_blocks_per_stage)]
)
self.final_conv = conv(mid_channels, in_channels, kernel_size=3, padding=1)
def forward(self, latent: torch.Tensor) -> torch.Tensor:
b, _, f, _, _ = latent.shape
if self.dims == 2:
x = rearrange(latent, "b c f h w -> (b f) c h w")
x = self.initial_conv(x)
x = self.initial_norm(x)
x = self.initial_activation(x)
for block in self.res_blocks:
x = block(x)
x = self.upsampler(x)
for block in self.post_upsample_res_blocks:
x = block(x)
x = self.final_conv(x)
x = rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f)
else:
x = self.initial_conv(latent)
x = self.initial_norm(x)
x = self.initial_activation(x)
for block in self.res_blocks:
x = block(x)
if self.temporal_upsample:
x = self.upsampler(x)
x = x[:, :, 1:, :, :]
elif isinstance(self.upsampler, SpatialRationalResampler):
x = self.upsampler(x)
else:
x = rearrange(x, "b c f h w -> (b f) c h w")
x = self.upsampler(x)
x = rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f)
for block in self.post_upsample_res_blocks:
x = block(x)
x = self.final_conv(x)
return x
class LatentUpsamplerConfigurator:
"""Configurator for LatentUpsampler from a config dict."""
@classmethod
def from_config(cls, config: dict[str, Any]) -> LatentUpsampler:
cfg = dict(config)
cfg.pop("_class_name", None)
if "upsampler" in cfg and isinstance(cfg["upsampler"], dict):
cfg = cfg["upsampler"]
return LatentUpsampler(
in_channels=cfg.get("in_channels", 128),
mid_channels=cfg.get("mid_channels", 512),
num_blocks_per_stage=cfg.get("num_blocks_per_stage", 4),
dims=cfg.get("dims", 3),
spatial_upsample=cfg.get("spatial_upsample", True),
temporal_upsample=cfg.get("temporal_upsample", False),
spatial_scale=cfg.get("spatial_scale", 2.0),
rational_resampler=cfg.get("rational_resampler", False),
)
class LTX2LatentUpsampler(nn.Module):
"""Public wrapper for the LTX-2 latent upsampler."""
def __init__(self, config: dict[str, Any]):
super().__init__()
self.model: LatentUpsampler = LatentUpsamplerConfigurator.from_config(config)
def forward(self, latent: torch.Tensor) -> torch.Tensor:
return self.model(latent)
def upsample_video(latent: torch.Tensor, video_encoder: Any, upsampler: LatentUpsampler) -> torch.Tensor:
"""
Upsample a latent tensor with normalization based on the video encoder's per-channel statistics.
"""
if not hasattr(video_encoder, "per_channel_statistics"):
raise ValueError("video_encoder must expose per_channel_statistics for normalization")
stats = video_encoder.per_channel_statistics
latent = stats.un_normalize(latent)
latent = upsampler(latent)
latent = stats.normalize(latent)
return latent
__all__ = [
"PixelShuffleND",
"BlurDownsample",
"SpatialRationalResampler",
"ResBlock",
"LatentUpsampler",
"LatentUpsamplerConfigurator",
"LTX2LatentUpsampler",
"upsample_video",
]
+194 -5
View File
@@ -11,14 +11,15 @@ from transformers import AutoTokenizer
from fastvideo.fastvideo_args import FastVideoArgs
from fastvideo.logger import init_logger
from fastvideo.models.loader.component_loader import PipelineComponentLoader
from fastvideo.pipelines.composed_pipeline_base import ComposedPipelineBase
from fastvideo.pipelines.lora_pipeline import LoRAPipeline
from fastvideo.pipelines.stages import (DecodingStage, InputValidationStage, LTX2AudioDecodingStage, LTX2DenoisingStage,
LTX2LatentPreparationStage, LTX2TextEncodingStage)
LTX2LatentPreparationStage, LTX2RefineInitStage, LTX2RefineLoRAStage,
LTX2TextEncodingStage, LTX2UpsampleStage, STAGE_2_DISTILLED_SIGMA_VALUES)
logger = init_logger(__name__)
class LTX2Pipeline(ComposedPipelineBase):
class LTX2Pipeline(LoRAPipeline):
_required_config_modules = [
"text_encoder",
@@ -30,6 +31,8 @@ class LTX2Pipeline(ComposedPipelineBase):
]
def create_pipeline_stages(self, fastvideo_args: FastVideoArgs):
refine_enabled = fastvideo_args.ltx2_refine_enabled
self.add_stage(
stage_name="input_validation_stage",
stage=InputValidationStage(),
@@ -43,9 +46,18 @@ class LTX2Pipeline(ComposedPipelineBase):
),
)
if refine_enabled:
self.add_stage(
stage_name="ltx2_refine_init_stage",
stage=LTX2RefineInitStage(),
)
self.add_stage(
stage_name="latent_preparation_stage",
stage=LTX2LatentPreparationStage(transformer=self.get_module("transformer"), ),
stage=LTX2LatentPreparationStage(
transformer=self.get_module("transformer"),
vae=self.get_module("vae"),
),
)
self.add_stage(
@@ -53,6 +65,69 @@ class LTX2Pipeline(ComposedPipelineBase):
stage=LTX2DenoisingStage(transformer=self.get_module("transformer"), ),
)
if refine_enabled:
stage2_steps = fastvideo_args.ltx2_refine_num_inference_steps
# LTX-2 refine currently supports two explicitly tested step
# counts:
# - 3 steps: official distilled schedule
# - 2 steps: custom reduced schedule used for faster
# experimentation
# Other values are intentionally rejected to avoid silent
# quality regressions.
if stage2_steps == 3:
# Official distilled stage-2 refine schedule.
stage2_sigmas = STAGE_2_DISTILLED_SIGMA_VALUES
elif stage2_steps == 2:
# Reduced 2-step refine schedule (explicitly omits 0.725).
stage2_sigmas = [
STAGE_2_DISTILLED_SIGMA_VALUES[0],
STAGE_2_DISTILLED_SIGMA_VALUES[2],
STAGE_2_DISTILLED_SIGMA_VALUES[3],
]
else:
logger.warning(
"For LTX-2 refinement, "
"ltx2_refine_num_inference_steps=%s is not a tested "
"setting. Using denoising steps other than 2 or 3 "
"may cause quality degradation.",
stage2_steps,
)
raise ValueError("LTX-2 refinement supports only 2 or 3 denoising "
"steps.")
transformer_refine = self.get_module("transformer_refine", self.get_module("transformer"))
self.add_stage(
stage_name="ltx2_upsample_stage",
stage=LTX2UpsampleStage(
upsampler=self.get_module("spatial_upsampler"),
vae=self.get_module("vae"),
transformer=transformer_refine,
sigmas=stage2_sigmas,
add_noise=fastvideo_args.ltx2_refine_add_noise,
),
)
if fastvideo_args.ltx2_refine_lora_path:
self.add_stage(
stage_name="ltx2_refine_lora_stage",
stage=LTX2RefineLoRAStage(
pipeline=self,
lora_path=fastvideo_args.ltx2_refine_lora_path,
),
)
self.add_stage(
stage_name="ltx2_refine_denoising_stage",
stage=LTX2DenoisingStage(
transformer=transformer_refine,
sigmas_override=stage2_sigmas,
num_inference_steps_override=len(stage2_sigmas) - 1,
force_guidance_scale=(fastvideo_args.ltx2_refine_guidance_scale),
initial_audio_latents_key="ltx2_audio_latents",
),
)
self.add_stage(
stage_name="audio_decoding_stage",
stage=LTX2AudioDecodingStage(
@@ -67,6 +142,36 @@ class LTX2Pipeline(ComposedPipelineBase):
)
def initialize_pipeline(self, fastvideo_args: FastVideoArgs):
# Optional debug-instrumentation env vars. The internal FastVideoArgs
# carries debug_model_sums / debug_model_detail toggles; the public
# args don't define them yet, so getattr-with-default keeps the
# pipeline runnable on both shapes without forcing the public args
# to add fields a public consumer wouldn't otherwise touch.
if getattr(fastvideo_args, "debug_model_sums", False):
os.environ["LTX2_PIPELINE_DEBUG_LOG"] = "1"
sums_path = getattr(fastvideo_args, "debug_model_sums_path", None)
if sums_path:
os.environ["LTX2_PIPELINE_DEBUG_PATH"] = sums_path
else:
logger.warning("debug_model_sums is enabled but "
"debug_model_sums_path is not set; no model sums "
"will be logged.")
else:
os.environ.pop("LTX2_PIPELINE_DEBUG_LOG", None)
os.environ.pop("LTX2_PIPELINE_DEBUG_PATH", None)
if getattr(fastvideo_args, "debug_model_detail", False):
os.environ["LTX2_DEBUG_DETAIL"] = "1"
detail_path = getattr(fastvideo_args, "debug_model_detail_path", None)
if detail_path:
os.environ["LTX2_PIPELINE_DEBUG_DETAIL_PATH"] = detail_path
else:
logger.warning("debug_model_detail is enabled but "
"debug_model_detail_path is not set; no detailed "
"hooks will be logged.")
else:
os.environ.pop("LTX2_DEBUG_DETAIL", None)
os.environ.pop("LTX2_PIPELINE_DEBUG_DETAIL_PATH", None)
tokenizer = self.get_module("tokenizer")
if tokenizer is not None:
tokenizer.padding_side = "left"
@@ -81,6 +186,51 @@ class LTX2Pipeline(ComposedPipelineBase):
model_index = self._load_config(self.model_path)
logger.info("Loading pipeline modules from config: %s", model_index)
# Apply optional FastVideo-specific refine defaults embedded in
# model_index.json. These are bundled with distilled checkpoints
# so the pipeline can self-configure without explicit user kwargs.
def _resolve_refine_path(value: str | None) -> str | None:
if value is None:
return None
if os.path.isabs(value):
return value
candidate = os.path.join(self.model_path, value)
if os.path.exists(candidate):
return candidate
return value
if (model_index.get("fastvideo_refine_enabled") is True and fastvideo_args.refine_enabled is None):
fastvideo_args.ltx2_refine_enabled = True
if (fastvideo_args.refine_upsampler_path is None and fastvideo_args.ltx2_refine_upsampler_path is None):
fastvideo_args.ltx2_refine_upsampler_path = _resolve_refine_path(
model_index.get("fastvideo_refine_upsampler_path"))
if (fastvideo_args.ltx2_refine_upsampler_path is None and "spatial_upsampler" in model_index):
fastvideo_args.ltx2_refine_upsampler_path = (_resolve_refine_path("spatial_upsampler"))
if (fastvideo_args.refine_transformer_path is None and fastvideo_args.ltx2_refine_transformer_path is None):
fastvideo_args.ltx2_refine_transformer_path = (_resolve_refine_path(
model_index.get("fastvideo_refine_transformer_path")))
if (fastvideo_args.refine_lora_path is None and fastvideo_args.ltx2_refine_lora_path is None):
fastvideo_args.ltx2_refine_lora_path = _resolve_refine_path(model_index.get("fastvideo_refine_lora_path"))
if (fastvideo_args.refine_num_inference_steps is None
and fastvideo_args.ltx2_refine_num_inference_steps == FastVideoArgs.ltx2_refine_num_inference_steps
and model_index.get("fastvideo_refine_num_inference_steps") is not None):
# Only apply the model-index default when the caller didn't
# explicitly set either refine_num_inference_steps (generic)
# or ltx2_refine_num_inference_steps (LTX-2-specific). This
# prevents bundled defaults from overwriting explicit caller
# intent (e.g. requesting 2-step refinement).
fastvideo_args.ltx2_refine_num_inference_steps = int(model_index["fastvideo_refine_num_inference_steps"])
if (fastvideo_args.refine_guidance_scale is None
and model_index.get("fastvideo_refine_guidance_scale") is not None):
fastvideo_args.ltx2_refine_guidance_scale = float(model_index["fastvideo_refine_guidance_scale"])
if (fastvideo_args.refine_add_noise is None and model_index.get("fastvideo_refine_add_noise") is not None):
fastvideo_args.ltx2_refine_add_noise = bool(model_index["fastvideo_refine_add_noise"])
if (fastvideo_args.refine_noise_path is None and fastvideo_args.ltx2_refine_noise_path is None):
fastvideo_args.ltx2_refine_noise_path = _resolve_refine_path(model_index.get("fastvideo_refine_noise_path"))
if (fastvideo_args.refine_audio_noise_path is None and fastvideo_args.ltx2_refine_audio_noise_path is None):
fastvideo_args.ltx2_refine_audio_noise_path = (_resolve_refine_path(
model_index.get("fastvideo_refine_audio_noise_path")))
model_index.pop("_class_name")
model_index.pop("_diffusers_version")
model_index.pop("workload_type", None)
@@ -111,7 +261,8 @@ class LTX2Pipeline(ComposedPipelineBase):
if os.path.isdir(gemma_path):
component_model_path = gemma_path
else:
raise ValueError("Tokenizer directory missing and Gemma weights were not found.")
raise ValueError("Tokenizer directory missing and Gemma weights "
"were not found.")
module = PipelineComponentLoader.load_module(
module_name=module_name,
@@ -131,6 +282,44 @@ class LTX2Pipeline(ComposedPipelineBase):
if module_name not in modules or modules[module_name] is None:
raise ValueError(f"Required module {module_name} was not loaded properly")
if fastvideo_args.ltx2_refine_enabled:
upsampler_path = fastvideo_args.ltx2_refine_upsampler_path
if upsampler_path is None:
raise ValueError("ltx2_refine_enabled is True but "
"ltx2_refine_upsampler_path was not provided.")
if not os.path.isdir(upsampler_path):
raise ValueError("ltx2_refine_upsampler_path must be a directory "
"containing Diffusers-style upsampler weights; "
f"got {upsampler_path}")
config_path = os.path.join(upsampler_path, "config.json")
if not os.path.exists(config_path):
raise ValueError("ltx2_refine_upsampler_path must contain a Diffusers "
f"config.json; missing {config_path}")
if (loaded_modules is not None and "spatial_upsampler" in loaded_modules):
modules["spatial_upsampler"] = loaded_modules["spatial_upsampler"]
else:
modules["spatial_upsampler"] = (PipelineComponentLoader.load_module(
module_name="spatial_upsampler",
component_model_path=upsampler_path,
transformers_or_diffusers="diffusers",
fastvideo_args=fastvideo_args,
))
logger.info("Loaded module spatial_upsampler from %s", upsampler_path)
if (loaded_modules is not None and "transformer_refine" in loaded_modules):
modules["transformer_refine"] = loaded_modules["transformer_refine"]
elif fastvideo_args.ltx2_refine_transformer_path:
modules["transformer_refine"] = (PipelineComponentLoader.load_module(
module_name="transformer_refine",
component_model_path=(fastvideo_args.ltx2_refine_transformer_path),
transformers_or_diffusers="diffusers",
fastvideo_args=fastvideo_args,
))
logger.info(
"Loaded module transformer_refine from %s",
fastvideo_args.ltx2_refine_transformer_path,
)
return modules
@@ -6,6 +6,12 @@ from fastvideo.pipelines.basic.ltx2.stages.ltx2_denoising import (
LTX2DenoisingStage, )
from fastvideo.pipelines.basic.ltx2.stages.ltx2_latent_preparation import (
LTX2LatentPreparationStage, )
from fastvideo.pipelines.basic.ltx2.stages.ltx2_refine import (
STAGE_2_DISTILLED_SIGMA_VALUES,
LTX2RefineInitStage,
LTX2RefineLoRAStage,
LTX2UpsampleStage,
)
from fastvideo.pipelines.basic.ltx2.stages.ltx2_text_encoding import (
LTX2TextEncodingStage, )
@@ -13,5 +19,9 @@ __all__ = [
"LTX2AudioDecodingStage",
"LTX2DenoisingStage",
"LTX2LatentPreparationStage",
"LTX2RefineInitStage",
"LTX2RefineLoRAStage",
"LTX2TextEncodingStage",
"LTX2UpsampleStage",
"STAGE_2_DISTILLED_SIGMA_VALUES",
]
@@ -5,8 +5,11 @@ LTX-2 denoising stage using the native sigma schedule.
from __future__ import annotations
from contextlib import contextmanager
from itertools import combinations
import math
import os
from pathlib import Path
import torch
from tqdm.auto import tqdm
@@ -17,6 +20,11 @@ from fastvideo.fastvideo_args import FastVideoArgs
from fastvideo.forward_context import set_forward_context
from fastvideo.pipelines.pipeline_batch_info import ForwardBatch
from fastvideo.pipelines.stages.base import PipelineStage
from fastvideo.pipelines.basic.ltx2.stages.ltx2_image_conditioning import (LTX2_CONTINUATION_STAGE2_LAST_LATENT_KEY,
LTX2_VIDEO_CLEAN_LATENT_KEY,
LTX2_VIDEO_DENOISE_MASK_KEY,
apply_ltx2_gaussian_noiser,
post_process_ltx2_denoised)
from fastvideo.pipelines.stages.validators import StageValidators as V
from fastvideo.pipelines.stages.validators import VerificationResult
from fastvideo.logger import init_logger
@@ -25,6 +33,9 @@ from fastvideo.models.dits.ltx2 import (AudioLatentShape, DEFAULT_LTX2_AUDIO_CHA
DEFAULT_LTX2_AUDIO_SAMPLE_RATE, VideoLatentShape)
from fastvideo.utils import is_vsa_available
LTX2_AUDIO_CLEAN_LATENT_KEY = "ltx2_audio_clean_latent"
LTX2_AUDIO_DENOISE_MASK_KEY = "ltx2_audio_denoise_mask"
BASE_SHIFT_ANCHOR = 1024
MAX_SHIFT_ANCHOR = 4096
@@ -40,6 +51,18 @@ except ImportError:
vsa_available = False
@contextmanager
def _nvtx_range(name: str):
if os.getenv("FASTVIDEO_NVTX_PROFILE", "0") == "1" and torch.cuda.is_available():
torch.cuda.nvtx.range_push(name)
try:
yield
finally:
torch.cuda.nvtx.range_pop()
else:
yield
def _ltx2_sigmas(
steps: int,
latent: torch.Tensor | None,
@@ -76,12 +99,75 @@ def _ltx2_sigmas(
return sigmas
def _distilled_subset_sigmas(
steps: int,
device: torch.device,
) -> tuple[torch.Tensor, list[int]]:
"""Select distilled sigma values for a requested denoising step count.
For steps < 8, choose indices that minimize the largest adjacent sigma
drop while always preserving both endpoints.
"""
max_steps = len(DISTILLED_SIGMA_VALUES) - 1
if steps < 1 or steps > max_steps:
raise ValueError(f"Distilled subset supports steps in [1, {max_steps}], got {steps}.")
max_index = len(DISTILLED_SIGMA_VALUES) - 1
if steps == max_steps:
full_indices = list(range(max_index + 1))
full_sigmas = torch.tensor(
DISTILLED_SIGMA_VALUES,
dtype=torch.float32,
device=device,
)
return full_sigmas, full_indices
interior_count = steps - 1
best_key: tuple[float, float, float] | None = None
best_indices: tuple[int, ...] | None = None
for interior in combinations(range(1, max_index), interior_count):
candidate = (0, *interior, max_index)
gaps = [
DISTILLED_SIGMA_VALUES[lo] - DISTILLED_SIGMA_VALUES[hi]
for lo, hi in zip(candidate, candidate[1:], strict=False)
]
key = (max(gaps), gaps[-1], sum(gap * gap for gap in gaps))
if best_key is None or key < best_key:
best_key = key
best_indices = candidate
if best_indices is None:
raise RuntimeError("Failed to construct distilled subset schedule.")
base_sigmas = torch.tensor(
DISTILLED_SIGMA_VALUES,
dtype=torch.float32,
device=device,
)
subset_indices = torch.tensor(best_indices, dtype=torch.long, device=device)
subset_sigmas = base_sigmas.index_select(0, subset_indices)
return subset_sigmas, list(best_indices)
class LTX2DenoisingStage(PipelineStage):
"""Run the LTX-2 denoising loop over the sigma schedule."""
def __init__(self, transformer) -> None:
def __init__(
self,
transformer,
*,
sigmas_override: list[float] | None = None,
num_inference_steps_override: int | None = None,
force_guidance_scale: float | None = None,
initial_audio_latents_key: str | None = "ltx2_audio_latents",
) -> None:
super().__init__()
self.transformer = transformer
self.sigmas_override = sigmas_override
self.num_inference_steps_override = num_inference_steps_override
self.force_guidance_scale = force_guidance_scale
self.initial_audio_latents_key = initial_audio_latents_key
def forward(
self,
@@ -92,16 +178,48 @@ class LTX2DenoisingStage(PipelineStage):
raise ValueError("Latents must be provided before denoising.")
latents = batch.latents
video_clean_latent = batch.extra.get(LTX2_VIDEO_CLEAN_LATENT_KEY)
video_denoise_mask = batch.extra.get(LTX2_VIDEO_DENOISE_MASK_KEY)
if (video_clean_latent is None) != (video_denoise_mask is None):
raise ValueError("LTX-2 i2v conditioning state is inconsistent: clean_latent/mask "
"must both be set or both be unset.")
if video_clean_latent is not None and video_denoise_mask is not None:
if not torch.is_tensor(video_clean_latent) or not torch.is_tensor(video_denoise_mask):
raise TypeError("LTX-2 i2v conditioning tensors must be Tensors.")
video_clean_latent = video_clean_latent.to(
device=latents.device,
dtype=latents.dtype,
)
video_denoise_mask = video_denoise_mask.to(
device=latents.device,
dtype=torch.float32,
)
prompt_embeds = batch.prompt_embeds[0]
prompt_mask = None
num_inference_steps = (self.num_inference_steps_override
if self.num_inference_steps_override is not None else batch.num_inference_steps)
cfg_scale_video = batch.ltx2_cfg_scale_video
cfg_scale_audio = batch.ltx2_cfg_scale_audio
effective_guidance_scale = batch.guidance_scale
use_cfg = batch.do_classifier_free_guidance
if self.force_guidance_scale is not None:
effective_guidance_scale = float(self.force_guidance_scale)
cfg_scale_video = effective_guidance_scale
cfg_scale_audio = effective_guidance_scale
use_cfg = effective_guidance_scale > 1.0
neg_prompt_embeds = None
neg_prompt_mask = None
# Only load negative prompts if CFG is actually enabled
if batch.do_classifier_free_guidance:
if not batch.negative_prompt_embeds:
raise ValueError("CFG is enabled but negative_prompt_embeds is empty")
neg_prompt_embeds = batch.negative_prompt_embeds[0]
if use_cfg:
if batch.negative_prompt_embeds is not None and batch.negative_prompt_embeds:
neg_prompt_embeds = batch.negative_prompt_embeds[0]
else:
logger.warning("[LTX2] CFG requested but negative_prompt_embeds missing; "
"falling back to no-CFG for this stage.")
use_cfg = False
# Ensure text conditioning is on the same device as latents.
if prompt_embeds.device != latents.device:
@@ -116,22 +234,42 @@ class LTX2DenoisingStage(PipelineStage):
target_dtype = torch.bfloat16
autocast_enabled = (target_dtype != torch.float32) and not fastvideo_args.disable_autocast
# Use official distilled sigma schedule for 8 steps (distilled models)
use_distilled_sigmas = os.getenv("LTX2_USE_DISTILLED_SIGMAS", "1") == "1"
if use_distilled_sigmas and batch.num_inference_steps == 8:
if self.sigmas_override is not None:
sigmas = torch.tensor(
DISTILLED_SIGMA_VALUES,
self.sigmas_override,
device=latents.device,
dtype=torch.float32,
)
logger.info("[LTX2] Using official distilled sigma schedule")
logger.info("[LTX2] Using override sigma schedule, %s", self.sigmas_override)
else:
sigmas = _ltx2_sigmas(
steps=batch.num_inference_steps,
latent=None,
device=latents.device,
)
logger.info("[LTX2] Using computed sigma schedule")
# Use distilled hardcoded schedule (or subsets) when enabled.
use_distilled_sigmas = os.getenv("LTX2_USE_DISTILLED_SIGMAS", "1") == "1"
max_distilled_steps = len(DISTILLED_SIGMA_VALUES) - 1
if use_distilled_sigmas and num_inference_steps <= max_distilled_steps:
sigmas, distilled_indices = _distilled_subset_sigmas(
steps=num_inference_steps,
device=latents.device,
)
if num_inference_steps == max_distilled_steps:
logger.info("[LTX2] Using official distilled sigma schedule")
else:
gaps = sigmas[:-1] - sigmas[1:]
logger.info(
"[LTX2] Using distilled sigma subset for %d steps "
"(indices=%s max_gap=%.6f tail_gap=%.6f)",
num_inference_steps,
distilled_indices,
float(gaps.max().item()),
float(gaps[-1].item()),
)
else:
sigmas = _ltx2_sigmas(
steps=num_inference_steps,
latent=None,
device=latents.device,
)
logger.info("[LTX2] Using computed sigma schedule, "
"num_inference_steps=%s", num_inference_steps)
if hasattr(self.transformer, "patchifier"):
video_shape = VideoLatentShape.from_torch_shape(latents.shape)
token_count = self.transformer.patchifier.get_token_count(video_shape)
@@ -139,23 +277,53 @@ class LTX2DenoisingStage(PipelineStage):
token_count = 1
timestep_template = torch.ones(
(latents.shape[0], token_count),
(latents.shape[0], token_count, 1),
device=latents.device,
dtype=torch.float32,
)
if video_denoise_mask is not None:
patchifier = getattr(self.transformer, "patchifier", None)
if patchifier is not None:
mask_patch = patchifier.patchify(video_denoise_mask.to(dtype=latents.dtype))
timestep_template = mask_patch.mean(dim=-1).to(torch.float32)
else:
flat_mask = video_denoise_mask.reshape(video_denoise_mask.shape[0], -1).to(torch.float32)
if flat_mask.shape[1] != token_count:
raise ValueError("LTX-2 i2v timestep mask token count mismatch: "
f"expected {token_count}, got {flat_mask.shape[1]}")
timestep_template = flat_mask
audio_prompt_embeds = batch.extra.get("ltx2_audio_prompt_embeds")
audio_neg_embeds = batch.extra.get("ltx2_audio_negative_embeds")
audio_context_p = audio_prompt_embeds[0] if audio_prompt_embeds else None
audio_context_n = audio_neg_embeds[0] if audio_neg_embeds else None
audio_latents = None
audio_latents = batch.extra.get(self.initial_audio_latents_key)
if isinstance(audio_latents, torch.Tensor):
audio_latents = audio_latents.to(device=latents.device, dtype=latents.dtype)
# Audio conditioning: mirror video i2v mask approach.
audio_clean_latent = batch.extra.get(LTX2_AUDIO_CLEAN_LATENT_KEY)
audio_denoise_mask = batch.extra.get(LTX2_AUDIO_DENOISE_MASK_KEY)
if isinstance(audio_clean_latent, torch.Tensor):
audio_clean_latent = audio_clean_latent.to(device=latents.device, dtype=latents.dtype)
if isinstance(audio_denoise_mask, torch.Tensor):
audio_denoise_mask = audio_denoise_mask.to(device=latents.device, dtype=torch.float32)
audio_timestep_template = None
if audio_context_p is not None:
if audio_latents is not None:
audio_timestep_template = torch.ones(
(latents.shape[0], audio_latents.shape[2], 1),
device=latents.device,
dtype=torch.float32,
)
if audio_context_p is not None and audio_latents is None:
fps_value = batch.fps
if isinstance(fps_value, list):
fps_value = fps_value[0] if fps_value else None
if fps_value is None:
fps_value = 1.0
duration = float(batch.num_frames) / float(fps_value)
# Allow audio to span a longer duration than video so
# audio conditioning can overlap more without shrinking
# the newly-generated audio region.
audio_num_frames = batch.extra.get("audio_num_frames", batch.num_frames)
duration = float(audio_num_frames) / float(fps_value)
audio_shape = AudioLatentShape.from_duration(
batch=latents.shape[0],
duration=duration,
@@ -165,44 +333,86 @@ class LTX2DenoisingStage(PipelineStage):
hop_length=DEFAULT_LTX2_AUDIO_HOP_LENGTH,
audio_latent_downsample_factor=DEFAULT_LTX2_AUDIO_DOWNSAMPLE,
)
audio_generator = None
if fastvideo_args.ltx2_initial_latent_path and batch.seed is not None:
audio_generator = torch.Generator(device=latents.device).manual_seed(batch.seed)
elif batch.generator is not None:
audio_generator = batch.generator[0] if isinstance(batch.generator, list) else batch.generator
if audio_generator is not None and audio_generator.device.type != latents.device.type:
if batch.seed is None:
audio_generator = torch.Generator(device=latents.device)
else:
audio_generator = torch.Generator(device=latents.device).manual_seed(batch.seed)
audio_patch_shape = (
expected_shape = (
audio_shape.batch,
audio_shape.channels,
audio_shape.frames,
audio_shape.channels * audio_shape.mel_bins,
audio_shape.mel_bins,
)
audio_latents_patch = torch.randn(
audio_patch_shape,
generator=audio_generator,
audio_latent_path = fastvideo_args.ltx2_audio_latent_path
audio_latents = self._load_audio_latents(
audio_latent_path,
device=latents.device,
dtype=latents.dtype,
)
if hasattr(self.transformer, "audio_patchifier"):
audio_latents = self.transformer.audio_patchifier.unpatchify(audio_latents_patch, audio_shape)
else:
audio_latents = audio_latents_patch.view(
expected_shape=expected_shape,
) if audio_latent_path else None
if audio_latents is None:
audio_generator = None
if fastvideo_args.ltx2_initial_latent_path and batch.seed is not None:
audio_generator = torch.Generator(device=latents.device).manual_seed(batch.seed)
elif batch.generator is not None:
audio_generator = (batch.generator[0] if isinstance(batch.generator, list) else batch.generator)
if audio_generator is not None and audio_generator.device.type != latents.device.type:
if batch.seed is None:
audio_generator = torch.Generator(device=latents.device)
else:
audio_generator = torch.Generator(device=latents.device).manual_seed(batch.seed)
audio_patch_shape = (
audio_shape.batch,
audio_shape.frames,
audio_shape.channels,
audio_shape.mel_bins,
).permute(0, 2, 1, 3).contiguous()
audio_shape.channels * audio_shape.mel_bins,
)
audio_latents_patch = torch.randn(
audio_patch_shape,
generator=audio_generator,
device=latents.device,
dtype=latents.dtype,
)
if hasattr(self.transformer, "audio_patchifier"):
audio_latents = self.transformer.audio_patchifier.unpatchify(audio_latents_patch, audio_shape)
else:
audio_latents = audio_latents_patch.view(
audio_shape.batch,
audio_shape.frames,
audio_shape.channels,
audio_shape.mel_bins,
).permute(0, 2, 1, 3).contiguous()
if audio_latent_path:
self._save_audio_latents(audio_latent_path, audio_latents)
audio_timestep_template = torch.ones(
(latents.shape[0], audio_shape.frames),
(latents.shape[0], audio_shape.frames, 1),
device=latents.device,
dtype=torch.float32,
)
# Apply audio conditioning mask (mirrors video i2v approach).
if (audio_latents is not None and audio_clean_latent is not None and audio_denoise_mask is not None):
audio_T = audio_latents.shape[2]
cond_T = audio_clean_latent.shape[2]
if audio_T != cond_T:
logger.warning(
"[LTX2] Audio conditioning T mismatch: "
"latents=%d, clean=%d; skipping audio "
"conditioning.", audio_T, cond_T)
audio_clean_latent = None
audio_denoise_mask = None
else:
# audio_denoise_mask: [B, 1, T, 1] → [B, T]
# for timestep template.
audio_timestep_template = (audio_denoise_mask[:, 0, :, 0])
audio_latents = apply_ltx2_gaussian_noiser(
noise=audio_latents,
clean_latent=audio_clean_latent,
denoise_mask=audio_denoise_mask,
noise_scale=1.0,
)
# Video position offset: when audio is longer than video for
# conditioning, shift video RoPE positions forward so the
# audio prefix sits at t>=0 and video aligns with the later
# portion of audio.
video_position_offset_sec = float(batch.extra.get("video_position_offset_sec", 0.0))
# Multi-modal CFG parameters (per-stream scales).
cfg_scale_video = batch.ltx2_cfg_scale_video
cfg_scale_audio = batch.ltx2_cfg_scale_audio
modality_scale_video = batch.ltx2_modality_scale_video
modality_scale_audio = batch.ltx2_modality_scale_audio
rescale_scale = batch.ltx2_rescale_scale
@@ -211,11 +421,13 @@ class LTX2DenoisingStage(PipelineStage):
stg_scale_audio = batch.ltx2_stg_scale_audio
stg_blocks_video = batch.ltx2_stg_blocks_video
stg_blocks_audio = batch.ltx2_stg_blocks_audio
do_stg_video = stg_scale_video != 0.0
do_stg_audio = stg_scale_audio != 0.0
do_stg_video = not math.isclose(float(stg_scale_video), 0.0)
do_stg_audio = not math.isclose(float(stg_scale_audio), 0.0)
do_stg = do_stg_video or do_stg_audio
do_cfg_text = (cfg_scale_video != 1.0 or cfg_scale_audio != 1.0)
do_mod = (modality_scale_video != 1.0 or modality_scale_audio != 1.0)
do_cfg_text = use_cfg and (cfg_scale_video != 1.0 or cfg_scale_audio != 1.0)
do_modality_video = not math.isclose(float(modality_scale_video), 1.0)
do_modality_audio = not math.isclose(float(modality_scale_audio), 1.0)
do_mod = do_modality_video or do_modality_audio
do_guidance = do_cfg_text or do_mod or do_stg
if do_cfg_text and neg_prompt_embeds is None:
@@ -225,12 +437,14 @@ class LTX2DenoisingStage(PipelineStage):
logger.info(
"[LTX2] Denoising start: steps=%d dtype=%s "
"cfg=%.1f "
"cfg_video=%.1f cfg_audio=%.1f mod_video=%.1f "
"mod_audio=%.1f rescale=%.2f "
"stg_video=%.1f stg_audio=%.1f "
"sigmas_shape=%s latents_shape=%s",
batch.num_inference_steps,
num_inference_steps,
target_dtype,
effective_guidance_scale,
cfg_scale_video,
cfg_scale_audio,
modality_scale_video,
@@ -241,7 +455,20 @@ class LTX2DenoisingStage(PipelineStage):
tuple(sigmas.shape),
tuple(latents.shape),
)
use_vsa = (vsa_available and envs.FASTVIDEO_ATTENTION_BACKEND == "VIDEO_SPARSE_ATTN")
# Hint runtime FP4 layer gating (single shared transformer path):
# stage-1 denoising uses "base", stage-2 refine uses "refine".
batch.extra["ltx2_fp4_stage_profile"] = ("refine" if self.sigmas_override is not None else "base")
attention_backend = os.getenv("FASTVIDEO_ATTENTION_BACKEND", envs.FASTVIDEO_ATTENTION_BACKEND)
wants_vsa_metadata = attention_backend in (
"VIDEO_SPARSE_ATTN",
"SAGE_ATTN_THREE",
)
# VIDEO_SPARSE_ATTN requires the fastvideo-kernel VSA op.
# SAGE_ATTN_THREE VSA+QAT only needs VSA metadata and its own kernel path.
use_vsa = wants_vsa_metadata and (attention_backend != "VIDEO_SPARSE_ATTN" or vsa_available)
if attention_backend == "VIDEO_SPARSE_ATTN" and not vsa_available:
logger.warning("FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN but VSA kernel "
"is unavailable; disabling VSA metadata for this run.")
vsa_metadata_builder = (VideoSparseAttentionMetadataBuilder() if use_vsa else None)
for step_index in tqdm(range(len(sigmas) - 1)):
@@ -249,6 +476,7 @@ class LTX2DenoisingStage(PipelineStage):
sigma_next = sigmas[step_index + 1]
timestep = timestep_template * sigma
audio_timestep = (audio_timestep_template * sigma if audio_timestep_template is not None else None)
latent_model_input = latents.to(target_dtype)
attn_metadata = None
if vsa_metadata_builder is not None:
attn_metadata = vsa_metadata_builder.build(
@@ -264,26 +492,27 @@ class LTX2DenoisingStage(PipelineStage):
dtype=target_dtype,
enabled=autocast_enabled,
), set_forward_context(
current_timestep=sigma.item(),
current_timestep=sigma,
attn_metadata=attn_metadata,
forward_batch=batch,
):
# Pass 1: Full conditioning (text + cross-modal)
pos_outputs = self.transformer(
hidden_states=latents.to(target_dtype),
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_p,
audio_timestep=audio_timestep,
)
with _nvtx_range("ltx2.denoise.pass.pos"):
pos_outputs = self.transformer(
hidden_states=latent_model_input,
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_p,
audio_timestep=audio_timestep,
video_position_offset_sec=video_position_offset_sec,
)
if isinstance(pos_outputs, tuple):
pos_denoised, pos_audio = pos_outputs
else:
pos_denoised = pos_outputs
pos_audio = None
if do_guidance:
# Defaults: (pos - pos) = 0 under each scale.
neg_denoised = pos_denoised
@@ -295,56 +524,63 @@ class LTX2DenoisingStage(PipelineStage):
# Pass 2: text CFG (negative prompt)
if do_cfg_text:
neg_outputs = self.transformer(
hidden_states=latents.to(target_dtype),
encoder_hidden_states=neg_prompt_embeds,
encoder_attention_mask=neg_prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_n,
audio_timestep=audio_timestep,
)
with _nvtx_range("ltx2.denoise.pass.neg"):
neg_outputs = self.transformer(
hidden_states=latent_model_input,
encoder_hidden_states=neg_prompt_embeds,
encoder_attention_mask=neg_prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_n,
audio_timestep=audio_timestep,
video_position_offset_sec=video_position_offset_sec,
)
if isinstance(neg_outputs, tuple):
neg_denoised, neg_audio = neg_outputs
else:
neg_denoised = neg_outputs
neg_audio = None
# Pass 3: Modality-isolated (skip cross-modal
# attn)
# Pass 3: Modality-isolated (skip cross-modal attn)
if do_mod:
mod_outputs = self.transformer(
hidden_states=latents.to(target_dtype),
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_p,
audio_timestep=audio_timestep,
skip_cross_modal_attn=True,
)
with _nvtx_range("ltx2.denoise.pass.modality"):
mod_outputs = self.transformer(
hidden_states=latent_model_input,
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_p,
audio_timestep=audio_timestep,
skip_cross_modal_attn=True,
video_position_offset_sec=video_position_offset_sec,
)
if isinstance(mod_outputs, tuple):
mod_denoised, mod_audio = mod_outputs
else:
mod_denoised = mod_outputs
mod_audio = None
# Pass 4: STG perturbed (skip self-attn in
# specified blocks)
# Pass 4: STG perturbed (skip self-attn in specified blocks)
if do_stg:
ptb_outputs = self.transformer(
hidden_states=latents.to(target_dtype),
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=(audio_context_p),
audio_timestep=audio_timestep,
skip_video_self_attn_blocks=(stg_blocks_video if do_stg_video else None),
skip_audio_self_attn_blocks=(stg_blocks_audio if do_stg_audio else None),
)
with _nvtx_range("ltx2.denoise.pass.stg"):
ptb_outputs = self.transformer(
hidden_states=latent_model_input,
encoder_hidden_states=prompt_embeds,
encoder_attention_mask=prompt_mask,
timestep=timestep,
audio_hidden_states=audio_latents,
audio_encoder_hidden_states=audio_context_p,
audio_timestep=audio_timestep,
skip_video_self_attn_blocks=(stg_blocks_video if do_stg_video else None),
skip_audio_self_attn_blocks=(stg_blocks_audio if do_stg_audio else None),
video_position_offset_sec=video_position_offset_sec,
)
if isinstance(ptb_outputs, tuple):
ptb_denoised, ptb_audio = ptb_outputs
else:
ptb_denoised = ptb_outputs
ptb_audio = None
# Multi-modal guidance formula per stream.
vid = (pos_denoised + (cfg_scale_video - 1) * (pos_denoised - neg_denoised) +
@@ -369,24 +605,75 @@ class LTX2DenoisingStage(PipelineStage):
pos_denoised = vid
pos_audio = aud
if video_clean_latent is not None and video_denoise_mask is not None:
pos_denoised = post_process_ltx2_denoised(
denoised=pos_denoised,
denoise_mask=video_denoise_mask,
clean_latent=video_clean_latent,
)
if (audio_clean_latent is not None and audio_denoise_mask is not None and pos_audio is not None):
pos_audio = post_process_ltx2_denoised(
denoised=pos_audio,
denoise_mask=audio_denoise_mask,
clean_latent=audio_clean_latent,
)
sigma_value = sigma.to(torch.float32) if isinstance(sigma, torch.Tensor) else torch.tensor(
float(sigma),
device=latents.device,
dtype=torch.float32,
)
dt = sigma_next - sigma
velocity = ((latents.float() - pos_denoised.float()) / sigma_value).to(latents.dtype)
latents = (latents.float() + velocity.float() * dt).to(latents.dtype)
if pos_audio is not None and audio_latents is not None:
audio_velocity = ((audio_latents.float() - pos_audio.float()) / sigma_value).to(audio_latents.dtype)
audio_latents = (audio_latents.float() + audio_velocity.float() * dt).to(audio_latents.dtype)
with _nvtx_range("ltx2.denoise.scheduler_update"):
velocity = ((latents.float() - pos_denoised.float()) / sigma_value).to(latents.dtype)
latents = (latents.float() + velocity.float() * dt).to(latents.dtype)
if pos_audio is not None and audio_latents is not None:
audio_velocity = ((audio_latents.float() - pos_audio.float()) / sigma_value).to(audio_latents.dtype)
audio_latents = (audio_latents.float() + audio_velocity.float() * dt).to(audio_latents.dtype)
batch.latents = latents
batch.extra["ltx2_audio_latents"] = audio_latents
if (batch.return_continuation_state and self.sigmas_override is not None):
batch.extra[LTX2_CONTINUATION_STAGE2_LAST_LATENT_KEY] = (latents[:, :, -1:, :, :].detach().clone())
batch.extra[self.initial_audio_latents_key] = audio_latents
if self.initial_audio_latents_key != "ltx2_audio_latents":
batch.extra["ltx2_audio_latents"] = audio_latents
logger.info("[LTX2] Denoising done.")
return batch
def _load_audio_latents(
self,
latent_path: str | None,
*,
device: torch.device,
dtype: torch.dtype,
expected_shape: tuple[int, ...],
) -> torch.Tensor | None:
if not latent_path:
return None
path = Path(latent_path)
if not path.exists():
return None
payload = torch.load(path, map_location=device)
if isinstance(payload, dict):
latent = (payload.get("audio_latent") or payload.get("latent") or payload.get("audio"))
else:
latent = payload
if not torch.is_tensor(latent):
raise TypeError(f"Expected tensor audio latent in {path}")
if tuple(latent.shape) != tuple(expected_shape):
raise ValueError(
f"Audio latent shape mismatch for {path}: expected {expected_shape}, got {tuple(latent.shape)}")
logger.info("[LTX2] Loaded audio latent from %s", path)
return latent.to(device=device, dtype=dtype)
def _save_audio_latents(self, latent_path: str, latents: torch.Tensor) -> None:
path = Path(latent_path)
path.parent.mkdir(parents=True, exist_ok=True)
if path.exists():
return
torch.save({"audio_latent": latents.detach().cpu()}, path)
logger.info("[LTX2] Saved audio latent to %s", path)
def verify_input(self, batch: ForwardBatch, fastvideo_args: FastVideoArgs) -> VerificationResult:
result = VerificationResult()
result.add_check("latents", batch.latents, [V.is_tensor, V.with_dims(5)])
@@ -0,0 +1,499 @@
# SPDX-License-Identifier: Apache-2.0
"""FastVideo-native LTX-2 image-to-video conditioning helpers.
Public-side port of ``FastVideo-internal/.../ltx2_i2v_conditioning.py``.
The module composes a ``clean_latent`` + ``denoise_mask`` pair that the
LTX-2 latent-prep + denoising stages mix into the noise tensor, so a
generated segment can be anchored to:
* one or more conditioning images at specific latent frame indices
(``ltx2_images``),
* a multi-frame conditioning video clip jointly VAE-encoded
(``ltx2_video_conditions``),
* a continuation latent carried over from the previous segment
(``ltx2_conditioning_latent_stage1`` / ``_stage2``).
The streaming server's session controller populates the continuation
latents between segments; the legacy from_pretrained path passes
``ltx2_images`` / ``ltx2_image_crf`` through compat translation.
"""
from __future__ import annotations
import math
from dataclasses import dataclass
from io import BytesIO
from typing import TYPE_CHECKING
import numpy as np
import torch
import torch.nn.functional as F
from fastvideo.logger import init_logger
from fastvideo.models.vision_utils import load_image
if TYPE_CHECKING:
from fastvideo.pipelines.pipeline_batch_info import ForwardBatch
try:
import av
except ImportError: # pragma: no cover - optional dependency
av = None
logger = init_logger(__name__)
LTX2_VIDEO_CLEAN_LATENT_KEY = "ltx2_video_clean_latent"
LTX2_VIDEO_DENOISE_MASK_KEY = "ltx2_video_denoise_mask"
LTX2_CONTINUATION_STAGE1_LAST_LATENT_KEY = ("ltx2_continuation_stage1_last_latent")
LTX2_CONTINUATION_STAGE2_LAST_LATENT_KEY = ("ltx2_continuation_stage2_last_latent")
# NOTE: hard-coded for continuation quality experiments. The first
# latent frame of the next clip is anchored to the previous clip's
# last latent at full strength.
LTX2_CONTINUATION_TARGET_FRAME_IDX = 0
LTX2_CONTINUATION_STRENGTH = 1.0
DEFAULT_LTX2_IMAGE_CRF = 33.0
@dataclass
class LTX2ImageConditioningState:
"""Result of building image / continuation conditioning."""
clean_latent: torch.Tensor
denoise_mask: torch.Tensor
images: list[tuple[str, int, float]]
latent_conditioned: bool = False
def resolve_ltx2_images(batch: ForwardBatch) -> list[tuple[str, int, float]]:
"""Collect any LTX-2 image conditioning inputs from the batch.
Falls back to ``batch.image_path`` for the simple single-image i2v
case (anchors the first latent frame at full strength).
"""
images = batch.ltx2_images
if images is None and batch.image_path:
images = [(batch.image_path, 0, 1.0)]
if not images:
return []
resolved: list[tuple[str, int, float]] = []
for item in images:
if not isinstance(item, tuple | list) or len(item) != 3:
raise ValueError("Each ltx2_images item must be a tuple/list of "
"(path, frame_idx, strength).")
image_path, frame_idx, strength = item
frame_idx_int = int(frame_idx)
strength_float = float(strength)
if frame_idx_int < 0:
raise ValueError(f"LTX-2 frame_idx must be >= 0, got {frame_idx_int}")
if strength_float < 0.0 or strength_float > 1.0:
raise ValueError(f"LTX-2 image conditioning strength must be in [0, 1], "
f"got {strength_float}")
resolved.append((str(image_path), frame_idx_int, strength_float))
return resolved
def _resize_and_center_crop(
tensor: torch.Tensor,
height: int,
width: int,
) -> torch.Tensor:
if tensor.ndim != 3:
raise ValueError(f"Expected image tensor [H, W, C], got shape {tuple(tensor.shape)}")
tensor = tensor.permute(2, 0, 1).unsqueeze(0)
_, _, src_h, src_w = tensor.shape
scale = max(height / src_h, width / src_w)
new_h = math.ceil(src_h * scale)
new_w = math.ceil(src_w * scale)
tensor = F.interpolate(
tensor,
size=(new_h, new_w),
mode="bilinear",
align_corners=False,
)
crop_top = (new_h - height) // 2
crop_left = (new_w - width) // 2
tensor = tensor[:, :, crop_top:crop_top + height, crop_left:crop_left + width]
return tensor.unsqueeze(2)
def _encode_single_frame(output_file: BytesIO, image_array: np.ndarray, crf: float) -> None:
container = av.open(output_file, "w", format="mp4")
try:
stream = container.add_stream(
"libx264",
rate=1,
options={
"crf": str(crf),
"preset": "veryfast",
},
)
height = image_array.shape[0] // 2 * 2
width = image_array.shape[1] // 2 * 2
image_array = image_array[:height, :width]
stream.height = height
stream.width = width
frame = av.VideoFrame.from_ndarray(image_array, format="rgb24").reformat(format="yuv420p")
container.mux(stream.encode(frame))
container.mux(stream.encode())
finally:
container.close()
def _decode_single_frame(video_file: BytesIO) -> np.ndarray:
container = av.open(video_file)
try:
stream = next(s for s in container.streams if s.type == "video")
frame = next(container.decode(stream))
finally:
container.close()
return frame.to_ndarray(format="rgb24")
def _preprocess_conditioning_image(
image: np.ndarray,
image_crf: float,
) -> np.ndarray:
"""H.264 CRF re-encode the conditioning image to match the
quantization the model was trained on. ``image_crf <= 0.0`` skips
the re-encode (used by the streaming server which conditions on
already-decoded VAE-quality frames)."""
if image_crf <= 0.0:
return image
if av is None:
logger.warning("[LTX2] PyAV is unavailable; skipping CRF "
"conditioning preprocessing.")
return image
with BytesIO() as output_file:
_encode_single_frame(output_file, image, image_crf)
encoded = output_file.getvalue()
with BytesIO(encoded) as video_file:
return _decode_single_frame(video_file)
def load_ltx2_conditioning_image(
image_path: str,
*,
height: int,
width: int,
dtype: torch.dtype,
device: torch.device,
image_crf: float,
) -> torch.Tensor:
image = load_image(image_path)
image_np = np.array(image)[..., :3]
image_np = _preprocess_conditioning_image(image_np, image_crf=image_crf)
image_tensor = torch.tensor(image_np, dtype=torch.float32, device=device)
image_tensor = _resize_and_center_crop(image_tensor, height, width)
image_tensor = (image_tensor / 127.5 - 1.0).to(device=device, dtype=dtype)
return image_tensor
def load_ltx2_conditioning_video_clip(
frame_paths: list[str],
*,
height: int,
width: int,
dtype: torch.dtype,
device: torch.device,
image_crf: float,
) -> torch.Tensor:
"""Load multiple frames and stack as ``[1, C, T, H, W]`` for joint
VAE encoding so the resulting latent captures temporal/motion info.
"""
frame_tensors: list[torch.Tensor] = []
for path in frame_paths:
image = load_image(path)
image_np = np.array(image)[..., :3]
image_np = _preprocess_conditioning_image(image_np, image_crf=image_crf)
t = torch.tensor(image_np, dtype=torch.float32, device=device)
# _resize_and_center_crop returns [1, C, 1, H, W]
t = _resize_and_center_crop(t, height, width)
frame_tensors.append(t)
# Concat along T dimension -> [1, C, T, H, W]
video = torch.cat(frame_tensors, dim=2)
return (video / 127.5 - 1.0).to(device=device, dtype=dtype)
def _extract_video_latent(vae: torch.nn.Module, image: torch.Tensor) -> torch.Tensor:
with torch.no_grad():
encoded = vae.encode(image)
latent_dist = getattr(encoded, "latent_dist", None)
if latent_dist is not None:
encoded = latent_dist
latent: torch.Tensor
if torch.is_tensor(encoded):
latent = encoded
elif hasattr(encoded, "mode"):
latent = encoded.mode()
elif hasattr(encoded, "sample"):
latent = encoded.sample()
elif (isinstance(encoded, tuple | list) and encoded and torch.is_tensor(encoded[0])):
latent = encoded[0]
else:
raise TypeError(f"Unsupported VAE encode output type: {type(encoded)}")
if latent.ndim == 4:
latent = latent.unsqueeze(2)
if latent.ndim != 5:
raise ValueError(f"Expected video latent with 5 dims [B,C,T,H,W], got "
f"{tuple(latent.shape)}")
return latent
def _insert_conditioning_latent(
*,
conditioning_latent: torch.Tensor,
clean_latent: torch.Tensor,
denoise_mask: torch.Tensor,
frame_idx: int,
strength: float,
source_name: str,
) -> None:
if conditioning_latent.ndim == 4:
conditioning_latent = conditioning_latent.unsqueeze(2)
if conditioning_latent.ndim != 5:
raise ValueError(f"LTX-2 {source_name} latent must have 5 dims [B,C,T,H,W], "
f"got {tuple(conditioning_latent.shape)}")
if conditioning_latent.shape[0] == 1 and clean_latent.shape[0] > 1:
conditioned_latent = conditioning_latent.expand(
clean_latent.shape[0],
-1,
-1,
-1,
-1,
)
elif conditioning_latent.shape[0] == clean_latent.shape[0]:
conditioned_latent = conditioning_latent
else:
raise ValueError(f"LTX-2 {source_name} latent batch mismatch: "
f"{conditioning_latent.shape[0]} vs {clean_latent.shape[0]}")
if conditioned_latent.shape[1] != clean_latent.shape[1]:
raise ValueError(f"LTX-2 {source_name} latent channels mismatch: "
f"{conditioned_latent.shape[1]} vs {clean_latent.shape[1]}")
if conditioned_latent.shape[-2:] != clean_latent.shape[-2:]:
raise ValueError(f"LTX-2 {source_name} latent spatial shape "
f"mismatch: {tuple(conditioned_latent.shape[-2:])} "
f"vs {tuple(clean_latent.shape[-2:])}")
frame_count = conditioned_latent.shape[2]
end_idx = frame_idx + frame_count
if frame_idx < 0 or end_idx > clean_latent.shape[2]:
raise ValueError(f"LTX-2 {source_name} latent frame range out of bounds: "
f"frame_idx={frame_idx}, frame_count={frame_count}, "
f"latent_frames={clean_latent.shape[2]}")
clean_latent[:, :, frame_idx:end_idx] = conditioned_latent.to(
device=clean_latent.device,
dtype=clean_latent.dtype,
)
denoise_mask[:, :, frame_idx:end_idx] = 1.0 - float(strength)
def build_ltx2_image_conditioning(
*,
batch: ForwardBatch,
latents: torch.Tensor,
vae: torch.nn.Module,
height: int,
width: int,
image_crf: float | None = None,
base_clean_latent: torch.Tensor | None = None,
) -> LTX2ImageConditioningState | None:
"""Build the (clean_latent, denoise_mask) state for the next segment.
Returns ``None`` for plain T2V (no images, no continuation, no
video conditions). The denoise mask is 1 where the model should
sample fresh, 0 where it should preserve the conditioning latent
exactly. ``base_clean_latent is None`` corresponds to stage 1
(fresh half-res latent); ``base_clean_latent`` set means stage 2
(already-upsampled latent from the upsampler stage).
"""
images = resolve_ltx2_images(batch)
conditioning_latent_stage1 = getattr(batch, "ltx2_conditioning_latent_stage1", None)
conditioning_latent_stage2 = getattr(batch, "ltx2_conditioning_latent_stage2", None)
is_stage1_conditioning = base_clean_latent is None
is_stage2_conditioning = not is_stage1_conditioning
has_latent_conditioning = False
continuation_latent_to_insert: torch.Tensor | None = None
if (conditioning_latent_stage1 is not None and not torch.is_tensor(conditioning_latent_stage1)):
raise TypeError("LTX-2 stage1 continuation latent conditioning "
"expects a torch.Tensor.")
if (conditioning_latent_stage2 is not None and not torch.is_tensor(conditioning_latent_stage2)):
raise TypeError("LTX-2 stage2 continuation latent conditioning "
"expects a torch.Tensor.")
if (conditioning_latent_stage1 is None) != (conditioning_latent_stage2 is None):
raise ValueError("LTX-2 continuation expects both stage1 and stage2 "
"conditioning latents (or neither for first round).")
if is_stage1_conditioning and conditioning_latent_stage1 is not None:
has_latent_conditioning = True
continuation_latent_to_insert = conditioning_latent_stage1.to(
device=latents.device,
dtype=latents.dtype,
)
elif is_stage2_conditioning and conditioning_latent_stage2 is not None:
has_latent_conditioning = True
continuation_latent_to_insert = conditioning_latent_stage2.to(
device=latents.device,
dtype=latents.dtype,
)
video_conditions = getattr(batch, "ltx2_video_conditions", None) or []
if not images and not has_latent_conditioning and not video_conditions:
return None
clean_latent = (torch.zeros_like(latents) if base_clean_latent is None else base_clean_latent.clone())
denoise_mask = torch.ones(
(
latents.shape[0],
1,
latents.shape[2],
latents.shape[3],
latents.shape[4],
),
dtype=torch.float32,
device=latents.device,
)
if image_crf is None:
image_crf = getattr(batch, "ltx2_image_crf", DEFAULT_LTX2_IMAGE_CRF)
vae_param = next(vae.parameters(), None)
encoder_dtype = (vae_param.dtype if vae_param is not None else latents.dtype)
encoder_device = (vae_param.device if vae_param is not None else latents.device)
cache: dict[tuple[str, int, int, float], torch.Tensor] = {}
latent_conditioned = False
if has_latent_conditioning:
if continuation_latent_to_insert is None:
raise RuntimeError("LTX-2 continuation latent conditioning state is invalid.")
# NOTE: frame index and strength are intentionally hard-coded.
# We always anchor the first frame of the next clip at full
# strength to the previous clip's last latent.
_insert_conditioning_latent(
conditioning_latent=continuation_latent_to_insert,
clean_latent=clean_latent,
denoise_mask=denoise_mask,
frame_idx=LTX2_CONTINUATION_TARGET_FRAME_IDX,
strength=LTX2_CONTINUATION_STRENGTH,
source_name="continuation",
)
latent_conditioned = True
for image_path, frame_idx, strength in images:
cache_key = (image_path, height, width, float(image_crf))
image_latent = cache.get(cache_key)
if image_latent is None:
image_tensor = load_ltx2_conditioning_image(
image_path=image_path,
height=height,
width=width,
dtype=encoder_dtype,
device=encoder_device,
image_crf=float(image_crf),
)
image_latent = _extract_video_latent(vae, image_tensor).to(
device=latents.device,
dtype=latents.dtype,
)
cache[cache_key] = image_latent
_insert_conditioning_latent(
conditioning_latent=image_latent,
clean_latent=clean_latent,
denoise_mask=denoise_mask,
frame_idx=frame_idx,
strength=strength,
source_name="image",
)
for frame_paths, frame_idx, strength in video_conditions:
video_tensor = load_ltx2_conditioning_video_clip(
frame_paths,
height=height,
width=width,
dtype=encoder_dtype,
device=encoder_device,
image_crf=float(image_crf),
)
video_latent = _extract_video_latent(vae, video_tensor).to(
device=latents.device,
dtype=latents.dtype,
)
logger.info(
"[LTX2] Video-clip condition: %d frames -> "
"latent T=%d at frame_idx=%d strength=%.2f",
len(frame_paths),
video_latent.shape[2],
frame_idx,
strength,
)
_insert_conditioning_latent(
conditioning_latent=video_latent,
clean_latent=clean_latent,
denoise_mask=denoise_mask,
frame_idx=frame_idx,
strength=strength,
source_name="video_clip",
)
return LTX2ImageConditioningState(
clean_latent=clean_latent,
denoise_mask=denoise_mask,
images=images,
latent_conditioned=latent_conditioned,
)
def apply_ltx2_gaussian_noiser(
*,
noise: torch.Tensor,
clean_latent: torch.Tensor,
denoise_mask: torch.Tensor,
noise_scale: float = 1.0,
) -> torch.Tensor:
"""Mix ``noise`` into ``clean_latent`` along ``denoise_mask`` * scale.
Values close to 1 in the mask produce near-pure noise (used in a
fresh stage-2 latent), values near 0 leave the clean latent
untouched (used in conditioning regions).
"""
scaled_mask = denoise_mask * float(noise_scale)
return (noise * scaled_mask + clean_latent * (1.0 - scaled_mask)).to(noise.dtype)
def post_process_ltx2_denoised(
*,
denoised: torch.Tensor,
denoise_mask: torch.Tensor,
clean_latent: torch.Tensor,
) -> torch.Tensor:
"""Restore the conditioning regions of ``clean_latent`` outside the
denoise mask after the model has filled in the masked area."""
return (denoised * denoise_mask + clean_latent.float() * (1.0 - denoise_mask)).to(denoised.dtype)
__all__ = [
"DEFAULT_LTX2_IMAGE_CRF",
"LTX2_CONTINUATION_STAGE1_LAST_LATENT_KEY",
"LTX2_CONTINUATION_STAGE2_LAST_LATENT_KEY",
"LTX2_CONTINUATION_STRENGTH",
"LTX2_CONTINUATION_TARGET_FRAME_IDX",
"LTX2_VIDEO_CLEAN_LATENT_KEY",
"LTX2_VIDEO_DENOISE_MASK_KEY",
"LTX2ImageConditioningState",
"apply_ltx2_gaussian_noiser",
"build_ltx2_image_conditioning",
"load_ltx2_conditioning_image",
"load_ltx2_conditioning_video_clip",
"post_process_ltx2_denoised",
"resolve_ltx2_images",
]
@@ -3,6 +3,7 @@
Latent preparation stage for LTX-2 pipelines.
"""
import math
from pathlib import Path
import torch
@@ -11,28 +12,80 @@ from diffusers.utils.torch_utils import randn_tensor
from fastvideo.distributed import get_local_torch_device
from fastvideo.fastvideo_args import FastVideoArgs
from fastvideo.logger import init_logger
from fastvideo.models.dits.ltx2 import VideoLatentShape
from fastvideo.pipelines.pipeline_batch_info import ForwardBatch
from fastvideo.pipelines.stages.base import PipelineStage
from fastvideo.pipelines.basic.ltx2.stages.ltx2_image_conditioning import (LTX2_VIDEO_CLEAN_LATENT_KEY,
LTX2_VIDEO_DENOISE_MASK_KEY,
apply_ltx2_gaussian_noiser,
build_ltx2_image_conditioning)
from fastvideo.pipelines.stages.validators import StageValidators as V
from fastvideo.pipelines.stages.validators import VerificationResult
logger = init_logger(__name__)
def _randn_ltx2_video_latents(
*,
shape: tuple[int, int, int, int, int],
transformer,
generator,
device: torch.device,
dtype: torch.dtype,
) -> torch.Tensor:
"""Match official LTX-2 video noise sampling order.
Official code builds a patchified latent state first, then samples noise in
token order. FastVideo stores latents in [B, C, T, H, W], so we sample in
token order here and unpatchify back to native layout.
"""
patchifier = getattr(transformer, "patchifier", None)
if patchifier is None or not callable(getattr(patchifier, "unpatchify", None)):
return randn_tensor(
shape,
generator=generator,
device=device,
dtype=dtype,
)
video_shape = VideoLatentShape.from_torch_shape(torch.Size(shape))
patch_volume = math.prod(getattr(patchifier, "patch_size", (1, 1, 1)))
patch_shape = (
shape[0],
patchifier.get_token_count(video_shape),
shape[1] * patch_volume,
)
# `torch.randn` accepts only a single torch.Generator. Some callers
# (e.g. InputValidationStage) hand us a one-element list when
# num_videos_per_prompt == 1; unwrap it here. For batched sampling
# (>1 sample) this collapses to the first generator — match this
# against expected reproducibility semantics if that path is ever used.
if isinstance(generator, list):
generator = generator[0] if generator else None
patch_noise = torch.randn(
patch_shape,
generator=generator,
device=device,
dtype=dtype,
)
return patchifier.unpatchify(patch_noise, video_shape)
class LTX2LatentPreparationStage(PipelineStage):
"""Prepare initial LTX-2 latents without relying on a diffusers scheduler."""
def __init__(self, transformer) -> None:
def __init__(self, transformer, vae) -> None:
super().__init__()
self.transformer = transformer
self.vae = vae
def forward(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> ForwardBatch:
latent_num_frames = (batch.num_frames -
1) // fastvideo_args.pipeline_config.vae_config.arch_config.temporal_compression_ratio + 1
latent_num_frames = self._adjust_video_length(batch, fastvideo_args)
if not batch.prompt_embeds:
batch_size = 1
elif isinstance(batch.prompt, list):
@@ -61,6 +114,22 @@ class LTX2LatentPreparationStage(PipelineStage):
dtype = batch.prompt_embeds[0].dtype
device = get_local_torch_device()
generator = batch.generator
if generator is not None:
if isinstance(generator, list):
if generator and generator[0].device.type != device.type:
seeds = batch.seeds
if seeds is None and batch.seed is not None:
seeds = [batch.seed + i for i in range(len(generator))]
if seeds is not None:
generator = [torch.Generator(device=device).manual_seed(seed) for seed in seeds]
batch.generator = generator
else:
if generator.device.type != device.type:
if batch.seed is not None:
generator = torch.Generator(device=device).manual_seed(batch.seed)
else:
generator = torch.Generator(device=device)
batch.generator = generator
latents = batch.latents
num_frames = latent_num_frames if latent_num_frames is not None else batch.num_frames
height = batch.height
@@ -92,16 +161,18 @@ class LTX2LatentPreparationStage(PipelineStage):
if loaded_latents is not None:
latents = loaded_latents
else:
latents = randn_tensor(
shape,
latents = _randn_ltx2_video_latents(
shape=shape,
transformer=self.transformer,
generator=generator,
device=device,
dtype=dtype,
)
self._save_initial_latent(latent_path, latents)
else:
latents = randn_tensor(
shape,
latents = _randn_ltx2_video_latents(
shape=shape,
transformer=self.transformer,
generator=generator,
device=device,
dtype=dtype,
@@ -109,10 +180,42 @@ class LTX2LatentPreparationStage(PipelineStage):
else:
latents = latents.to(device)
image_conditioning = build_ltx2_image_conditioning(
batch=batch,
latents=latents,
vae=self.vae,
height=height,
width=width,
)
if image_conditioning is None:
batch.extra.pop(LTX2_VIDEO_CLEAN_LATENT_KEY, None)
batch.extra.pop(LTX2_VIDEO_DENOISE_MASK_KEY, None)
else:
latents = apply_ltx2_gaussian_noiser(
noise=latents,
clean_latent=image_conditioning.clean_latent,
denoise_mask=image_conditioning.denoise_mask,
noise_scale=1.0,
)
batch.extra[LTX2_VIDEO_CLEAN_LATENT_KEY] = (image_conditioning.clean_latent)
batch.extra[LTX2_VIDEO_DENOISE_MASK_KEY] = (image_conditioning.denoise_mask)
logger.info(
"[LTX2] Applied conditioning for stage-1: images=%d latent=%s.",
len(image_conditioning.images),
image_conditioning.latent_conditioned,
)
batch.latents = latents
batch.raw_latent_shape = shape
return batch
def _adjust_video_length(self, batch: ForwardBatch, fastvideo_args: FastVideoArgs) -> int | None:
if not fastvideo_args.pipeline_config.vae_config.use_temporal_scaling_frames:
return None
temporal_scale_factor = (fastvideo_args.pipeline_config.vae_config.arch_config.temporal_compression_ratio)
video_length = batch.num_frames
return int((video_length - 1) // temporal_scale_factor + 1)
def _load_initial_latent(
self,
latent_path: str,
@@ -0,0 +1,395 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2 refinement stages for 2x spatial upscaling + distilled denoising.
Public-side port of ``FastVideo-internal/.../stages/ltx2_refine.py``.
The three stages run between the stage-1 denoising pass and the stage-2
denoising pass:
* :class:`LTX2RefineInitStage` — halves the requested resolution so the
first denoise runs at ½× and stashes the original target resolution
on ``batch.extra`` so the upsample stage can recover it.
* :class:`LTX2UpsampleStage` — upsamples the stage-1 latents through
the LTX-2 latent upsampler, optionally re-applies image conditioning,
and mixes in fresh noise scaled by the stage-2 sigma so the next
denoise has something to refine.
* :class:`LTX2RefineLoRAStage` — swaps in a refinement LoRA before the
stage-2 denoise (no-op when the path is unset).
Behaviour matches the internal version 1:1 for the text-to-video path;
the i2v / continuation branches inside ``build_ltx2_image_conditioning``
defer to a NotImplementedError until the rest of the i2v conditioning
module is ported.
"""
from __future__ import annotations
import weakref
from pathlib import Path
from typing import Any
import torch
from diffusers.utils.torch_utils import randn_tensor
from fastvideo.fastvideo_args import FastVideoArgs
from fastvideo.logger import init_logger
from fastvideo.models.dits.ltx2 import AudioLatentShape, VideoLatentShape
from fastvideo.models.upsamplers import upsample_video
from fastvideo.pipelines.basic.ltx2.stages.ltx2_image_conditioning import (
LTX2_CONTINUATION_STAGE1_LAST_LATENT_KEY,
LTX2_VIDEO_CLEAN_LATENT_KEY,
LTX2_VIDEO_DENOISE_MASK_KEY,
apply_ltx2_gaussian_noiser,
build_ltx2_image_conditioning,
)
from fastvideo.pipelines.pipeline_batch_info import ForwardBatch
from fastvideo.pipelines.stages.base import PipelineStage
from fastvideo.pipelines.stages.validators import StageValidators as V
from fastvideo.pipelines.stages.validators import VerificationResult
logger = init_logger(__name__)
# Reduced schedule for super-resolution stage 2 (subset of distilled values).
# Lifted verbatim from LTX-2 upstream
# (packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py).
STAGE_2_DISTILLED_SIGMA_VALUES = [0.909375, 0.725, 0.421875, 0.0]
class LTX2RefineInitStage(PipelineStage):
"""Switch the request to half resolution before the stage-1 denoise.
Stashes the original target resolution on ``batch.extra`` so
:class:`LTX2UpsampleStage` can recover it after stage 1 runs. When
the refine path is disabled the stage is a no-op.
"""
def forward(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> ForwardBatch:
if not fastvideo_args.ltx2_refine_enabled:
return batch
height = batch.height
width = batch.width
if height is None or width is None:
raise ValueError("Height and width must be provided for LTX-2 refinement.")
if isinstance(height, list) or isinstance(width, list):
raise ValueError("LTX-2 refinement expects scalar height/width.")
if height % 2 != 0 or width % 2 != 0:
raise ValueError("LTX-2 refinement requires even height/width so stage1 can be "
"half resolution.")
spatial_ratio = (fastvideo_args.pipeline_config.vae_config.arch_config.spatial_compression_ratio)
stage1_height = height // 2
stage1_width = width // 2
if stage1_height % spatial_ratio != 0 or stage1_width % spatial_ratio != 0:
raise ValueError(f"LTX-2 refinement requires height/width divisible by "
f"{2 * spatial_ratio} (got {height}x{width}).")
batch.extra["ltx2_refine_target_height"] = height
batch.extra["ltx2_refine_target_width"] = width
batch.height = stage1_height
batch.width = stage1_width
logger.info(
"[LTX2] Refinement enabled: stage1=%dx%d stage2=%dx%d",
stage1_width,
stage1_height,
width,
height,
)
return batch
def verify_output(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> VerificationResult:
# Only meaningful checks live downstream of the upsample stage;
# the init stage just rewrites height/width which the existing
# latent-prep stage already validates.
return VerificationResult()
class LTX2UpsampleStage(PipelineStage):
"""Upsample stage-1 latents to stage-2 resolution and add refine noise."""
def __init__(
self,
*,
upsampler: Any,
vae: Any,
transformer: Any | None = None,
sigmas: list[float] | None = None,
add_noise: bool = True,
) -> None:
super().__init__()
self.upsampler = upsampler
self.vae = vae
self.transformer = transformer
self.sigmas = sigmas or STAGE_2_DISTILLED_SIGMA_VALUES
self.add_noise = add_noise
def forward(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> ForwardBatch:
if not fastvideo_args.ltx2_refine_enabled:
return batch
if batch.latents is None:
raise ValueError("Latents must be available before LTX-2 upsample stage.")
latents = batch.latents
if batch.return_continuation_state:
batch.extra[LTX2_CONTINUATION_STAGE1_LAST_LATENT_KEY] = (latents[:, :, -1:, :, :].detach().clone())
orig_dtype = latents.dtype
orig_device = latents.device
if isinstance(self.upsampler, torch.nn.Module):
first_param = next(self.upsampler.parameters(), None)
if first_param is not None:
if first_param.device != orig_device:
latents = latents.to(device=first_param.device)
if first_param.dtype != latents.dtype:
latents = latents.to(dtype=first_param.dtype)
if (latents.dtype != orig_dtype or latents.device != orig_device):
logger.info(
"[LTX2] Cast latents to %s on %s for upsampler.",
latents.dtype,
latents.device,
)
target_height = batch.extra.get("ltx2_refine_target_height")
target_width = batch.extra.get("ltx2_refine_target_width")
if target_height is None or target_width is None:
raise ValueError("Missing target resolution for LTX-2 refinement.")
video_encoder = getattr(self.vae, "encoder", None)
if video_encoder is None:
raise ValueError("LTX-2 VAE encoder is required for latent upsampling.")
upsampler_module = getattr(self.upsampler, "model", self.upsampler)
latents = upsample_video(latents, video_encoder, upsampler_module)
if latents.dtype != orig_dtype or latents.device != orig_device:
latents = latents.to(device=orig_device, dtype=orig_dtype)
image_conditioning = build_ltx2_image_conditioning(
batch=batch,
latents=latents,
vae=self.vae,
height=target_height,
width=target_width,
base_clean_latent=latents,
)
if image_conditioning is None:
batch.extra.pop(LTX2_VIDEO_CLEAN_LATENT_KEY, None)
batch.extra.pop(LTX2_VIDEO_DENOISE_MASK_KEY, None)
clean_latents = latents
denoise_mask = torch.ones(
(latents.shape[0], 1, latents.shape[2], latents.shape[3], latents.shape[4]),
dtype=torch.float32,
device=latents.device,
)
else:
clean_latents = image_conditioning.clean_latent
denoise_mask = image_conditioning.denoise_mask
batch.extra[LTX2_VIDEO_CLEAN_LATENT_KEY] = clean_latents
batch.extra[LTX2_VIDEO_DENOISE_MASK_KEY] = denoise_mask
logger.info(
"[LTX2] Applied conditioning for stage-2: images=%d latent=%s.",
len(image_conditioning.images),
image_conditioning.latent_conditioned,
)
sigma0 = float(self.sigmas[0]) if self.sigmas else 1.0
if self.add_noise:
patchifier = getattr(self.transformer, "patchifier", None)
patch_noise_shape: torch.Size | None = None
video_shape: VideoLatentShape | None = None
if patchifier is not None:
video_shape = VideoLatentShape.from_torch_shape(latents.shape)
patch_noise_shape = patchifier.patchify(latents).shape
noise_path = fastvideo_args.ltx2_refine_noise_path
noise = self._load_noise(
noise_path,
device=latents.device,
dtype=latents.dtype,
expected_shape=latents.shape,
alternate_shape=patch_noise_shape,
) if noise_path else None
if noise is None:
noise = randn_tensor(
latents.shape,
generator=batch.generator,
device=latents.device,
dtype=latents.dtype,
)
if noise_path:
self._save_noise(noise_path, noise)
elif (patchifier is not None and patch_noise_shape is not None and video_shape is not None
and noise.shape == patch_noise_shape):
noise = patchifier.unpatchify(noise, video_shape)
latents = apply_ltx2_gaussian_noiser(
noise=noise,
clean_latent=clean_latents,
denoise_mask=denoise_mask,
noise_scale=sigma0,
)
else:
latents = clean_latents
audio_latents = batch.extra.get("ltx2_audio_latents")
if audio_latents is not None:
audio_latents = audio_latents.to(device=latents.device)
if self.add_noise:
audio_patchifier = getattr(self.transformer, "audio_patchifier", None)
if audio_patchifier is not None:
audio_shape = AudioLatentShape.from_torch_shape(audio_latents.shape)
audio_patch = audio_patchifier.patchify(audio_latents)
audio_noise_shape = audio_patch.shape
audio_noise_path = (fastvideo_args.ltx2_refine_audio_noise_path)
audio_noise = self._load_noise(
audio_noise_path,
device=audio_latents.device,
dtype=audio_latents.dtype,
expected_shape=audio_noise_shape,
alternate_shape=audio_latents.shape,
) if audio_noise_path else None
if audio_noise is None:
audio_noise = randn_tensor(
audio_noise_shape,
generator=batch.generator,
device=audio_latents.device,
dtype=audio_latents.dtype,
)
if audio_noise_path:
self._save_noise(audio_noise_path, audio_noise)
elif audio_noise.shape == audio_latents.shape:
audio_noise = audio_patchifier.patchify(audio_noise)
audio_noised_patch = audio_noise * sigma0 + audio_patch * (1.0 - sigma0)
audio_latents = audio_patchifier.unpatchify(audio_noised_patch, audio_shape)
else:
audio_noise_path = (fastvideo_args.ltx2_refine_audio_noise_path)
audio_noise = self._load_noise(
audio_noise_path,
device=audio_latents.device,
dtype=audio_latents.dtype,
expected_shape=audio_latents.shape,
) if audio_noise_path else None
if audio_noise is None:
audio_noise = randn_tensor(
audio_latents.shape,
generator=batch.generator,
device=audio_latents.device,
dtype=audio_latents.dtype,
)
if audio_noise_path:
self._save_noise(audio_noise_path, audio_noise)
# Same noise mixing as video latents for the
# distilled refinement schedule.
audio_latents = audio_noise * sigma0 + audio_latents * (1.0 - sigma0)
batch.extra["ltx2_audio_latents"] = audio_latents
batch.latents = latents
batch.raw_latent_shape = latents.shape
batch.height = target_height
batch.width = target_width
return batch
def verify_input(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> VerificationResult:
result = VerificationResult()
if fastvideo_args.ltx2_refine_enabled:
result.add_check("latents", batch.latents, [V.is_tensor, V.with_dims(5)])
return result
def _load_noise(
self,
noise_path: str | None,
*,
device: torch.device,
dtype: torch.dtype,
expected_shape: torch.Size | tuple[int, ...],
alternate_shape: torch.Size | tuple[int, ...] | None = None,
) -> torch.Tensor | None:
if not noise_path:
return None
path = Path(noise_path)
if not path.exists():
return None
payload = torch.load(path, map_location=device)
if isinstance(payload, dict):
noise = (payload.get("noise") or payload.get("latent_noise") or payload.get("latent")
or payload.get("video_noise"))
else:
noise = payload
if not torch.is_tensor(noise):
raise TypeError(f"Expected tensor noise in {path}")
noise_shape = tuple(noise.shape)
if (noise_shape != tuple(expected_shape)
and (alternate_shape is None or noise_shape != tuple(alternate_shape))):
raise ValueError(f"Noise shape mismatch for {path}: expected "
f"{tuple(expected_shape)}, got {noise_shape}")
logger.info("[LTX2] Loaded refine noise from %s", path)
return noise.to(device=device, dtype=dtype)
def _save_noise(self, noise_path: str, noise: torch.Tensor) -> None:
path = Path(noise_path)
path.parent.mkdir(parents=True, exist_ok=True)
if path.exists():
return
torch.save({"noise": noise.detach().cpu()}, path)
logger.info("[LTX2] Saved refine noise to %s", path)
class LTX2RefineLoRAStage(PipelineStage):
"""Apply a refinement-specific LoRA before stage-2 denoising."""
def __init__(
self,
*,
pipeline: Any,
lora_path: str | None,
lora_nickname: str = "ltx2_refine",
) -> None:
super().__init__()
self._pipeline_ref = (weakref.ref(pipeline) if pipeline is not None else None)
self._lora_path = lora_path
self._lora_nickname = lora_nickname
self._applied = False
def forward(
self,
batch: ForwardBatch,
fastvideo_args: FastVideoArgs,
) -> ForwardBatch:
if not fastvideo_args.ltx2_refine_enabled:
return batch
lora_path = fastvideo_args.ltx2_refine_lora_path or self._lora_path
if not lora_path or self._applied:
return batch
pipeline = (self._pipeline_ref() if self._pipeline_ref is not None else None)
if pipeline is None or not hasattr(pipeline, "set_lora_adapter"):
raise ValueError("LTX2 refinement LoRA requested but pipeline does not "
"support LoRA adapters.")
pipeline.set_lora_adapter(self._lora_nickname, lora_path)
self._applied = True
logger.info("[LTX2] Applied refinement LoRA from %s", lora_path)
return batch
__all__ = [
"LTX2RefineInitStage",
"LTX2RefineLoRAStage",
"LTX2UpsampleStage",
"STAGE_2_DISTILLED_SIGMA_VALUES",
]
+62 -13
View File
@@ -131,6 +131,11 @@ class ComposedPipelineBase(ABC):
)
return
prepare_for_compile = getattr(module, "prepare_for_compile", None)
if callable(prepare_for_compile):
logger.info("Running prepare_for_compile for %s", module_name)
prepare_for_compile()
compiled_count = self._compile_with_conditions(module, compile_kwargs)
if compiled_count > 0:
logger.info(
@@ -159,7 +164,11 @@ class ComposedPipelineBase(ABC):
self.initialize_validation_pipeline(self.training_args)
self.initialize_pipeline(self.fastvideo_args)
if self.fastvideo_args.enable_torch_compile:
compile_transformer = self.fastvideo_args.enable_torch_compile
compile_text_encoder = (self.fastvideo_args.enable_torch_compile_text_encoder)
compile_vae = self.fastvideo_args.enable_torch_compile_vae
compile_audio_vae = self.fastvideo_args.enable_torch_compile_audio_vae
if (compile_transformer or compile_text_encoder or compile_vae or compile_audio_vae):
if self.fastvideo_args.training_mode:
logger.info("Torch Compile enabled via FSDP loader for training; skipping additional pipeline compile")
else:
@@ -170,18 +179,58 @@ class ComposedPipelineBase(ABC):
except Exception: # pragma: no cover - FSDP not always available
fsdp_module_cls = None
compile_kwargs = self.fastvideo_args.torch_compile_kwargs or {}
self._maybe_compile_pipeline_module(
module_name="transformer",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=compile_kwargs,
)
self._maybe_compile_pipeline_module(
module_name="transformer_2",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=compile_kwargs,
)
logger.info("Torch Compile enabled for DiT")
global_compile_kwargs = (self.fastvideo_args.torch_compile_kwargs or {})
dit_compile_kwargs = (self.fastvideo_args.torch_compile_kwargs_dit or global_compile_kwargs)
text_compile_kwargs = (self.fastvideo_args.torch_compile_kwargs_text_encoder or global_compile_kwargs)
vae_compile_kwargs = (self.fastvideo_args.torch_compile_kwargs_vae or global_compile_kwargs)
audio_vae_compile_kwargs = (self.fastvideo_args.torch_compile_kwargs_audio_vae or global_compile_kwargs)
if compile_transformer:
self._maybe_compile_pipeline_module(
module_name="transformer",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=dit_compile_kwargs,
)
self._maybe_compile_pipeline_module(
module_name="transformer_refine",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=dit_compile_kwargs,
)
self._maybe_compile_pipeline_module(
module_name="transformer_2",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=dit_compile_kwargs,
)
logger.info("Torch Compile enabled for DiT")
if compile_text_encoder:
self._maybe_compile_pipeline_module(
module_name="text_encoder",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=text_compile_kwargs,
)
self._maybe_compile_pipeline_module(
module_name="text_encoder_2",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=text_compile_kwargs,
)
logger.info("Torch Compile enabled for text encoder")
if compile_vae:
self._maybe_compile_pipeline_module(
module_name="vae",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=vae_compile_kwargs,
)
logger.info("Torch Compile enabled for VAE")
if compile_audio_vae:
self._maybe_compile_pipeline_module(
module_name="audio_vae",
fsdp_module_cls=fsdp_module_cls,
compile_kwargs=audio_vae_compile_kwargs,
)
logger.info("Torch Compile enabled for audio VAE")
if not self.fastvideo_args.training_mode:
logger.info("Creating pipeline stages...")
@@ -79,6 +79,27 @@ class ForwardBatch:
image_embeds: list[torch.Tensor] = field(default_factory=list)
pil_image: torch.Tensor | PIL.Image.Image | None = None
preprocessed_image: torch.Tensor | None = None
# LTX-2 image conditioning. Each entry is (path, frame_idx, strength)
# — strength 1.0 anchors the latent at frame_idx exactly to the
# encoded image; <1.0 mixes denoised content with the conditioning
# image. ``ltx2_image_crf`` controls the H.264 CRF re-encode applied
# to the image before VAE encoding (matches official LTX-2 quality
# adaptation; 0.0 disables the re-encode).
ltx2_images: list[tuple[str, int, float]] | None = None
ltx2_image_crf: float = 33.0
# Stage-specific continuation latents — populated by the streaming
# session controller between segments. ``stage1`` is the last latent
# from the previous segment's stage-1 denoise (half-res); ``stage2``
# is the last latent from the previous segment's stage-2 refine
# (full-res). The image-conditioning builder reads whichever the
# current pass needs (``base_clean_latent is None`` -> stage1).
ltx2_conditioning_latent_stage1: torch.Tensor | None = None
ltx2_conditioning_latent_stage2: torch.Tensor | None = None
# Video-clip conditioning: list of (frame_paths, frame_idx, strength)
# — multi-frame VAE-encoded conditioning so the resulting latent
# captures temporal/motion info (used by the streaming server's
# mid-roll continuation flow).
ltx2_video_conditions: list[tuple[list[str], int, float]] | None = None
# Text inputs
prompt: str | list[str] | None = None
+3 -5
View File
@@ -25,11 +25,9 @@ from fastvideo.pipelines.stages.latent_preparation import (Cosmos25LatentPrepara
Cosmos25AutoLatentPreparationStage,
Cosmos25T2WLatentPreparationStage,
Cosmos25V2WLatentPreparationStage, LatentPreparationStage)
from fastvideo.pipelines.basic.ltx2.stages import (
LTX2AudioDecodingStage,
LTX2DenoisingStage,
LTX2LatentPreparationStage,
LTX2TextEncodingStage,
from fastvideo.pipelines.basic.ltx2.stages import ( # noqa: F401
LTX2AudioDecodingStage, LTX2DenoisingStage, LTX2LatentPreparationStage, LTX2RefineInitStage, LTX2RefineLoRAStage,
LTX2TextEncodingStage, LTX2UpsampleStage, STAGE_2_DISTILLED_SIGMA_VALUES,
)
from fastvideo.pipelines.stages.matrixgame_denoising import (MatrixGameCausalDenoisingStage)
from fastvideo.pipelines.stages.hyworld_denoising import HYWorldDenoisingStage
+21 -14
View File
@@ -212,6 +212,27 @@ def _get_config_info(
def _register_configs() -> None:
# LTX-2 (distilled) — registered FIRST so its detector wins over
# the base detector when both fire. The detector loop in
# ``get_model_name_for_path`` ORs the path-based check with a
# pipeline-name check (``ltx2pipeline``) which the base detector's
# "distilled not in path" predicate matches as True (the
# pipeline_name string contains no "distilled" marker), so the
# less-specific BASE detector would otherwise win when the
# input is the absolute path of the distilled checkpoint.
register_configs(
sampling_param_cls=None,
pipeline_config_cls=LTX2T2VConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"FastVideo/LTX2-Distilled-Diffusers",
],
model_detectors=[
lambda path: ("ltx2" in path.lower() or "ltx-2" in path.lower()) and "distilled" in path.lower(),
],
model_family="ltx2",
default_preset="ltx2_distilled",
)
# LTX-2 (base)
register_configs(
sampling_param_cls=None,
@@ -228,20 +249,6 @@ def _register_configs() -> None:
model_family="ltx2",
default_preset="ltx2_base",
)
# LTX-2 (distilled)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=LTX2T2VConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"FastVideo/LTX2-Distilled-Diffusers",
],
model_detectors=[
lambda path: ("ltx2" in path.lower() or "ltx-2" in path.lower()) and "distilled" in path.lower(),
],
model_family="ltx2",
default_preset="ltx2_distilled",
)
# Stable Audio Open (text-to-audio). Both variants must be loaded
# from the FastVideo-curated converted Diffusers-format repos —
@@ -31,7 +31,7 @@ from fastvideo.api.compat import (
# One item from gpu_pool.py's load_kwargs is deliberately excluded:
# - ``pipeline_config=<PipelineConfig instance>`` — an opaque Python
# object; internal mutates it in place (``dit_config.quant_config =
# FP4Config()``). The typed path for quantization is tracked in
# NVFP4Config()``). The typed path for quantization is tracked in
# "Known Technical Debt" in PR plan.md; ``pipeline_config`` as an
# instance legitimately belongs in ``pipeline.experimental``.
#
+7 -1
View File
@@ -113,12 +113,18 @@ def test_load_run_config_supports_yaml_roundtrip(tmp_path) -> None:
},
"compile": {
"enabled": False,
"text_encoder_enabled": None,
"backend": None,
"fullgraph": None,
"mode": None,
"dynamic": None,
"extras": {},
"text_encoder_enabled": None,
"vae_enabled": None,
"audio_vae_enabled": None,
"dit_kwargs": {},
"text_encoder_kwargs": {},
"vae_kwargs": {},
"audio_vae_kwargs": {},
},
"enable_stage_verification": True,
"use_fsdp_inference": False,
@@ -0,0 +1,107 @@
# SPDX-License-Identifier: Apache-2.0
"""Typed quantization flow contract tests.
Locks in the path from typed
``GeneratorConfig.engine.quantization.transformer_quant: "NVFP4"``
through the compat layer to a concrete ``NVFP4Config`` instance pinned
on ``pipeline_config.dit_config.quant_config``.
The model loader detects FP4 by ``isinstance(quant_method,
NVFP4QuantizeMethod)`` rather than by a flag, so the typed surface
must reliably produce that class on the DiT config — otherwise the
loader silently runs full bf16.
"""
from __future__ import annotations
import pytest
from fastvideo.api.compat import generator_config_to_fastvideo_args
from fastvideo.api.schema import (
EngineConfig,
GeneratorConfig,
QuantizationConfig,
)
from fastvideo.layers.quantization.nvfp4_config import NVFP4Config
@pytest.fixture
def captured_kwargs(monkeypatch):
"""Replace ``FastVideoArgs.from_kwargs`` with a capturer so the
test doesn't try to download model_index.json.
"""
from fastvideo import fastvideo_args as fva
captured: dict[str, object] = {}
def _capture(**kw):
captured.update(kw)
class _Stub:
kwargs = kw
return _Stub()
monkeypatch.setattr(fva.FastVideoArgs, "from_kwargs", _capture)
return captured
def test_typed_transformer_quant_resolves_to_nvfp4_instance(
captured_kwargs) -> None:
cfg = GeneratorConfig(
model_path="FastVideo/LTX2-Distilled-Diffusers",
engine=EngineConfig(
quantization=QuantizationConfig(transformer_quant="NVFP4"), ),
)
generator_config_to_fastvideo_args(cfg)
assert "transformer_quant" in captured_kwargs, (
"compat layer must forward typed transformer_quant onto "
"FastVideoArgs.from_kwargs")
assert isinstance(captured_kwargs["transformer_quant"], NVFP4Config), (
f"Expected NVFP4Config instance, got "
f"{type(captured_kwargs['transformer_quant']).__name__}")
def test_no_typed_quant_omits_transformer_quant_kwarg(captured_kwargs) -> None:
"""Default GeneratorConfig has ``quantization=None`` — the carrier
must not be set, so the existing legacy path
(``pipeline_config.dit_config.quant_config = NVFP4Config()``)
keeps working as before.
"""
cfg = GeneratorConfig(model_path="FastVideo/LTX2-Distilled-Diffusers")
generator_config_to_fastvideo_args(cfg)
assert "transformer_quant" not in captured_kwargs
def test_apply_transformer_quant_pins_to_dit_config(monkeypatch) -> None:
"""``FastVideoArgs.__post_init__._apply_transformer_quant`` must
copy the ``transformer_quant`` instance onto
``pipeline_config.dit_config.quant_config`` so the DiT loader sees
it during construction.
"""
from fastvideo.fastvideo_args import FastVideoArgs
args = FastVideoArgs(model_path="FastVideo/LTX2-Distilled-Diffusers")
# ``transformer_quant`` defaults to None so __post_init__ leaves it.
assert args.transformer_quant is None
assert args.pipeline_config.dit_config.quant_config is None
nvfp4 = NVFP4Config()
args.transformer_quant = nvfp4
args._apply_transformer_quant()
assert args.pipeline_config.dit_config.quant_config is nvfp4
def test_apply_transformer_quant_does_not_overwrite_explicit_dit_config(
) -> None:
"""When the caller has explicitly set
``pipeline_config.dit_config.quant_config`` already, the typed
carrier defers — the explicit setter wins.
"""
from fastvideo.fastvideo_args import FastVideoArgs
explicit = NVFP4Config(layer_profile="base")
args = FastVideoArgs(model_path="FastVideo/LTX2-Distilled-Diffusers")
args.pipeline_config.dit_config.quant_config = explicit
args.transformer_quant = NVFP4Config(layer_profile="refine")
args._apply_transformer_quant()
assert args.pipeline_config.dit_config.quant_config is explicit
+9
View File
@@ -0,0 +1,9 @@
# SPDX-License-Identifier: Apache-2.0
"""Contract tests guarding FastVideo's public API against drift.
These tests run against the public surface only (`fastvideo.VideoGenerator`,
`fastvideo.api.*`) — never via private helpers. They fail at FastVideo CI
if a change breaks the shape the Dynamo backend package and the private
Dreamverse adapter depend on, so drift is caught here before it reaches
downstream integrators.
"""
@@ -0,0 +1,214 @@
# SPDX-License-Identifier: Apache-2.0
"""Contract test: Dreamverse-style inputs normalize through the public
typed API without needing any private-only compatibility promise.
The private Dreamverse UI server (``FastVideo-internal/ui/ltx2-streaming/
server/gpu_pool.py``) has historically called
``VideoGenerator.from_pretrained(**load_kwargs)`` with a flat kwarg bag
containing LTX-2-specific names (``ltx2_refine_enabled``,
``ltx2_refine_upsampler_path``, etc.). PR 6 gave every one of those
kwargs a typed home under ``GeneratorConfig``.
This test makes sure:
1. The public typed API can represent everything Dreamverse currently
passes at init time (``legacy_from_pretrained_to_config``).
2. The request-path Dreamverse uses (``generator.generate_video(**kwargs)``
with per-segment flags) round-trips through the typed ``GenerationRequest``
without reintroducing private-only fields at the public boundary.
3. Private-only Dreamverse fields that don't belong on the public
surface either go to ``pipeline.experimental`` / ``request.extensions``
(the documented escape hatch) or raise explicitly, rather than
silently becoming part of the public compatibility promise.
Regression guard for the scoping rule in ``apirefactor.md`` §"Schema
Parity Requirement".
"""
from __future__ import annotations
import pytest
from fastvideo.api import (
ComponentConfig,
CompileConfig,
GeneratorConfig,
GenerationRequest,
)
from fastvideo.api.compat import (
legacy_from_pretrained_to_config,
legacy_generate_call_to_request,
normalize_generation_request,
)
def _dreamverse_load_kwargs() -> dict:
"""The exact shape internal ``gpu_pool.py`` passes to
``VideoGenerator.from_pretrained(**load_kwargs)`` today."""
return {
"config_model_path": "/models/ltx2-config",
"ltx2_refine_enabled": True,
"ltx2_refine_upsampler_path": "/models/ltx2-refine",
"ltx2_refine_lora_path": "/models/ltx2-refine-lora",
"ltx2_refine_num_inference_steps": 2,
"ltx2_refine_guidance_scale": 1.0,
"ltx2_refine_add_noise": True,
"ltx2_vae_tiling": True,
"torch_compile_kwargs": {
"backend": "inductor",
"mode": "reduce-overhead",
"fullgraph": True,
},
"dit_cpu_offload": False,
"vae_cpu_offload": False,
"text_encoder_cpu_offload": False,
"pin_cpu_memory": True,
"use_fsdp_inference": False,
"enable_torch_compile": True,
}
class TestDreamverseLoadKwargsShape:
"""Every current Dreamverse init-time kwarg must land on a typed
field, not in the ``experimental`` escape hatch."""
def test_all_kwargs_land_on_typed_fields(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
assert isinstance(config, GeneratorConfig)
# None of the kwargs should have been routed to experimental.
assert config.pipeline.experimental == {}, (
"Dreamverse kwargs leaked into pipeline.experimental: "
f"{config.pipeline.experimental}")
def test_refine_enabled_reaches_preset_overrides(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
refine = config.pipeline.preset_overrides.get("refine") or {}
assert refine.get("enabled") is True
assert refine.get("add_noise") is True
assert refine.get("num_inference_steps") == 2
assert refine.get("guidance_scale") == 1.0
def test_refine_assets_reach_component_config(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
assert isinstance(config.pipeline.components, ComponentConfig)
assert (config.pipeline.components.upsampler_weights ==
"/models/ltx2-refine")
assert config.pipeline.components.lora_path == "/models/ltx2-refine-lora"
assert config.pipeline.components.config_root == "/models/ltx2-config"
def test_torch_compile_kwargs_reach_typed_fields(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
assert isinstance(config.engine.compile, CompileConfig)
assert config.engine.compile.enabled is True
assert config.engine.compile.backend == "inductor"
assert config.engine.compile.mode == "reduce-overhead"
assert config.engine.compile.fullgraph is True
# extras should be empty — all four common kwargs are first class.
assert config.engine.compile.extras == {}
def test_uncommon_compile_kwargs_fall_to_extras(self):
kwargs = _dreamverse_load_kwargs()
kwargs["torch_compile_kwargs"] = {
**kwargs["torch_compile_kwargs"],
"options": {"epilogue_fusion": True},
}
config = legacy_from_pretrained_to_config("/models/ltx2", kwargs)
assert config.engine.compile.extras == {
"options": {"epilogue_fusion": True},
}
def test_vae_tiling_reaches_pipeline_selection(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
assert config.pipeline.vae_tiling is True
def test_offload_fields_reach_typed_offload_config(self):
config = legacy_from_pretrained_to_config(
"/models/ltx2", _dreamverse_load_kwargs())
assert config.engine.offload.dit is False
assert config.engine.offload.vae is False
assert config.engine.offload.text_encoder is False
assert config.engine.offload.pin_cpu_memory is True
class TestDreamversePrivateOnlyFields:
"""Dreamverse carries a handful of private-only names (e.g. legacy
internal aliases). These must NOT silently turn into a public
compatibility promise — the documented contract is that unknown
fields land on ``pipeline.experimental`` so integrators see them
but FastVideo does not promise to preserve them."""
def test_unknown_kwarg_routes_to_experimental(self):
kwargs = _dreamverse_load_kwargs()
kwargs["dreamverse_internal_only_flag"] = "private"
config = legacy_from_pretrained_to_config("/models/ltx2", kwargs)
assert config.pipeline.experimental == {
"dreamverse_internal_only_flag": "private",
}
class TestDreamverseRequestShape:
"""The per-segment Dreamverse request path mirrors OpenAI's shape
plus a few LTX-2 knobs. All of them must have a typed home."""
def test_basic_request_fields_round_trip(self):
# The request path calls legacy_generate_call_to_request with a
# prompt + legacy kwargs; verify the typed shape carries them.
legacy_kwargs = {
"num_frames": 121,
"height": 1024,
"width": 1536,
"num_inference_steps": 8,
"guidance_scale": 1.0,
"seed": 42,
"fps": 24,
"negative_prompt": "blurry",
}
request = legacy_generate_call_to_request(
prompt="a fox running",
sampling_param=None,
legacy_kwargs=legacy_kwargs,
)
request = normalize_generation_request(request)
assert request.prompt == "a fox running"
assert request.negative_prompt == "blurry"
assert request.sampling.num_frames == 121
assert request.sampling.height == 1024
assert request.sampling.width == 1536
assert request.sampling.num_inference_steps == 8
assert request.sampling.guidance_scale == 1.0
assert request.sampling.seed == 42
assert request.sampling.fps == 24
def test_return_state_reaches_output_config(self):
"""PR 7 added ``output.return_state`` — must survive the legacy
translation path so Dreamverse callers can opt in."""
request = GenerationRequest(
prompt="x",
output=__import__(
"fastvideo.api", fromlist=["OutputConfig"]).OutputConfig(
return_state=True),
)
normalized = normalize_generation_request(request)
assert normalized.output.return_state is True
class TestDreamverseNoPrivateImports:
"""The public entry points must not force a Dreamverse integrator
to import from ``fastvideo.pipelines.*`` or other internal paths."""
@pytest.mark.parametrize(
"import_path",
[
"fastvideo",
"fastvideo.api",
"fastvideo.api.compat", # public in that it's re-exported
],
)
def test_public_imports_resolve(self, import_path):
import importlib
importlib.import_module(import_path)
@@ -0,0 +1,331 @@
# SPDX-License-Identifier: Apache-2.0
"""Contract test: a mock Dynamo-style handler wraps FastVideo's public API
without touching any private module.
The Dynamo backend package (``components/src/dynamo/fastvideo/`` in the
Dynamo repo) imports only these symbols:
from fastvideo import VideoGenerator
from fastvideo.api import (
ContinuationState, GenerationRequest, InputConfig, OutputConfig,
SamplingConfig,
)
If a FastVideo refactor breaks the adapter shape this test fails at
FastVideo CI — before the Dynamo-side integration knows. The plan (PR
7.10) requires the backend to be expressible without flat legacy LTX-2
kwargs or FastVideo-internal imports; this file asserts the subset that
exists today and is stable.
"""
from __future__ import annotations
import asyncio
import base64
import importlib
import sys
import time
from dataclasses import dataclass
from typing import Any, AsyncIterator
import pytest
# Public imports only — mirror what the Dynamo backend package uses.
from fastvideo.api import (
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
)
# ----------------------------------------------------------------------
# Fake Dynamo-side types (the real ones live in dynamo.llm and dynamo.*)
# We mimic their dict shape so the adapter code under test is realistic.
# ----------------------------------------------------------------------
@dataclass
class _FakeContext:
"""Stand-in for Dynamo's async RPC context."""
request_id: str = "test-req"
def _nv_create_video_request(**overrides: Any) -> dict[str, Any]:
"""Build a Dynamo ``NvCreateVideoRequest``-shaped dict."""
base: dict[str, Any] = {
"prompt": "a fox running through snow",
"size": "1024x1536",
"seconds": 5,
"response_format": "b64_json",
"model": "Lightricks/LTX-Video",
"nvext": {
"fps": 24,
"num_inference_steps": 8,
"guidance_scale": 1.0,
"seed": 42,
"negative_prompt": "blurry",
},
}
base.update(overrides)
if "nvext" in overrides:
base["nvext"] = {**base["nvext"], **overrides["nvext"]}
return base
# ----------------------------------------------------------------------
# The adapter under test — written as if it lived in
# components/src/dynamo/fastvideo/request_handlers/video_generation/.
# ----------------------------------------------------------------------
def _parse_size(size: str | None) -> tuple[int, int]:
if not size or "x" not in size:
return 1024, 1536
w, h = size.split("x", 1)
return int(w), int(h)
def nv_request_to_generation_request(
request: dict[str, Any], ) -> GenerationRequest:
"""Translate Dynamo's request shape into FastVideo's typed request.
This function is the template integrators copy into the Dynamo repo.
It uses only public FastVideo symbols.
"""
nvext = request.get("nvext") or {}
width, height = _parse_size(request.get("size"))
fps = nvext.get("fps") or 24
num_frames = nvext.get("num_frames") or (request.get("seconds") or 4) * fps
state = nvext.get("continuation_state")
if isinstance(state, dict):
state = ContinuationState(kind=state["kind"], payload=state["payload"])
return GenerationRequest(
prompt=request["prompt"],
negative_prompt=nvext.get("negative_prompt"),
inputs=InputConfig(image_path=request.get("input_reference")),
sampling=SamplingConfig(
width=width,
height=height,
num_frames=num_frames,
fps=fps,
num_inference_steps=nvext.get("num_inference_steps", 50),
guidance_scale=nvext.get("guidance_scale", 1.0),
seed=nvext.get("seed", 1024),
true_cfg_scale=nvext.get("true_cfg_scale"),
),
output=OutputConfig(save_video=False, return_frames=False),
state=state,
)
class _MockFastVideoHandler:
"""Mock ``VideoGenerationWorkerHandler`` — the Dynamo backend's
entry point. Shape verified: ``async def generate(dict, ctx) ->
AsyncGenerator[dict, None]``, the signature Dynamo's
``endpoint.serve_endpoint(...)`` expects.
"""
def __init__(self, fake_result: dict[str, Any]) -> None:
self._fake_result = fake_result
self._lock = asyncio.Lock()
async def generate(
self,
request: dict[str, Any],
context: _FakeContext,
) -> AsyncIterator[dict[str, Any]]:
req = nv_request_to_generation_request(request)
assert isinstance(req, GenerationRequest)
# The real handler would call:
# result = await asyncio.to_thread(self.generator.generate, req)
# For contract purposes we just assert the typed request reaches
# the adapter boundary and yield a synthetic final event.
t0 = time.perf_counter()
async with self._lock:
await asyncio.sleep(0) # model call stub
elapsed = time.perf_counter() - t0
yield {
"data": [{"b64_json": self._fake_result["b64_json"]}],
"inference_time_s": elapsed,
"model": request.get("model"),
"nvext": {
"continuation_state": (None if req.state is None else {
"kind": req.state.kind,
"payload": req.state.payload,
}),
},
}
# ----------------------------------------------------------------------
# Tests
# ----------------------------------------------------------------------
class TestDynamoAdapterShape:
def test_nv_request_translates_to_typed_request(self):
req = nv_request_to_generation_request(_nv_create_video_request())
assert isinstance(req, GenerationRequest)
assert req.prompt == "a fox running through snow"
assert req.sampling.width == 1024
assert req.sampling.height == 1536
assert req.sampling.num_frames == 120 # 5 seconds * 24 fps
assert req.sampling.fps == 24
assert req.sampling.num_inference_steps == 8
assert req.sampling.guidance_scale == 1.0
assert req.sampling.seed == 42
# negative_prompt lives at the request level, not inside sampling
assert req.negative_prompt == "blurry"
def test_nvext_num_frames_overrides_seconds_times_fps(self):
req = nv_request_to_generation_request(
_nv_create_video_request(nvext={"num_frames": 33}))
assert req.sampling.num_frames == 33
def test_input_reference_maps_to_image_path(self):
req = nv_request_to_generation_request(
_nv_create_video_request(input_reference="/tmp/init.png"))
assert req.inputs.image_path == "/tmp/init.png"
def test_continuation_state_round_trips_through_adapter(self):
state = {
"kind": "ltx2.v1",
"payload": {
"schema_version": 1,
"segment_index": 2,
},
}
req = nv_request_to_generation_request(
_nv_create_video_request(nvext={"continuation_state": state}))
assert isinstance(req.state, ContinuationState)
assert req.state.kind == "ltx2.v1"
assert req.state.payload["segment_index"] == 2
class TestDynamoHandlerContract:
def test_handler_generate_is_async_generator(self):
handler = _MockFastVideoHandler({"b64_json": "xyz"})
async def run():
out = []
async for chunk in handler.generate(
_nv_create_video_request(), _FakeContext()):
out.append(chunk)
return out
chunks = asyncio.run(run())
assert len(chunks) == 1
final = chunks[0]
assert final["data"][0]["b64_json"] == "xyz"
assert "inference_time_s" in final
assert final["nvext"]["continuation_state"] is None
def test_handler_serializes_state_back_to_nvext(self):
"""When the request carries state, the handler should be able to
include a matching serialized state on the response. (Dynamo's
NvVideosResponse has nvext.continuation_state reserved for this
in the pending disaggregation path.)"""
handler = _MockFastVideoHandler({"b64_json": "abc"})
state_in = {
"kind": "ltx2.v1",
"payload": {"schema_version": 1, "segment_index": 4},
}
async def run():
async for chunk in handler.generate(
_nv_create_video_request(
nvext={"continuation_state": state_in}),
_FakeContext()):
return chunk
raise AssertionError("no chunk emitted")
chunk = asyncio.run(run())
assert chunk["nvext"]["continuation_state"]["kind"] == "ltx2.v1"
assert (
chunk["nvext"]["continuation_state"]["payload"]["segment_index"]
== 4)
class TestNoInternalImports:
"""The adapter template in this file imports only the public surface.
Any change to FastVideo that requires the Dynamo adapter to reach
into a private module would make this test fail at review time.
"""
_PUBLIC_IMPORTS = frozenset({
"fastvideo",
"fastvideo.api",
})
_BANNED_PREFIXES = (
"fastvideo.pipelines.",
"fastvideo.configs.",
"fastvideo.fastvideo_args",
"fastvideo.api.compat",
"fastvideo.api.parser",
"fastvideo.api.overrides",
"fastvideo.api.errors",
"fastvideo.api.request_metadata",
"fastvideo.api.sampling_param",
)
def test_adapter_template_does_not_import_private_modules(self):
# Scan only the adapter function's source — not the whole test
# file — so the banned-prefix list itself doesn't trip the guard.
import inspect
adapter_source = inspect.getsource(nv_request_to_generation_request)
for banned in self._BANNED_PREFIXES:
assert banned not in adapter_source, (
f"Dynamo adapter template must not depend on {banned}*; "
"found a reference in the adapter function body.")
def test_public_types_are_stable_imports(self):
# These lines match the exact import snippet in docs/design/
# server_contracts/dynamo.md; any rename breaks CI.
from fastvideo import VideoGenerator # noqa: F401
from fastvideo.api import ( # noqa: F401
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
)
@pytest.mark.skipif(
"fastvideo.api" not in sys.modules,
reason="fastvideo.api not importable in this environment",
)
def test_public_api_re_exports_match_doc(self):
import fastvideo.api as api
required = {
"ContinuationState",
"GenerationRequest",
"InputConfig",
"OutputConfig",
"SamplingConfig",
}
missing = required - set(api.__all__)
assert not missing, (
f"fastvideo.api.__all__ missing Dynamo-contract symbols: {missing}")
def _read_own_source() -> str:
import pathlib
return pathlib.Path(__file__).read_text()
# Convenience export for third parties writing their own contract tests.
__all__ = [
"nv_request_to_generation_request",
]
@@ -0,0 +1,273 @@
# SPDX-License-Identifier: Apache-2.0
"""Contract tests for ``VideoGenerator.generate_async``.
These tests monkey-patch the synchronous ``_generate_request_impl`` so
the suite runs CPU-only -- the async wrapper is the piece under test,
not the pipeline.
"""
from __future__ import annotations
import asyncio
from typing import Any
import numpy as np
import pytest
from fastvideo.api.results import (
GenerationResult,
VideoFinalEvent,
VideoProgressEvent,
)
from fastvideo.api.schema import (
ContinuationState,
GenerationRequest,
OutputConfig,
SamplingConfig,
)
class _FakeExecutor:
def set_log_queue(self, q):
self.log_queue = q
def clear_log_queue(self):
self.log_queue = None
class _FakeVideoGenerator:
"""Stand-in that exposes the same generate_async contract as the
real VideoGenerator without requiring a loaded pipeline. Binds the
real method via ``__func__`` so the implementation under test is
exactly the production code."""
def __init__(self, result: GenerationResult | list[GenerationResult]):
self._result = result
self.executor = _FakeExecutor()
def _generate_request_impl(
self, request: GenerationRequest,
) -> GenerationResult | list[GenerationResult]:
return self._result
@staticmethod
def _wrap_legacy_result(result):
return GenerationResult.from_legacy_result(result)
@classmethod
def bind(cls, result):
from fastvideo.entrypoints.video_generator import VideoGenerator
instance = cls(result)
instance.generate_async = VideoGenerator.generate_async.__get__(
instance, cls)
instance.default_health_check_request = (
VideoGenerator.default_health_check_request)
return instance
def _make_request(**overrides: Any) -> GenerationRequest:
base = GenerationRequest(
prompt="hi",
sampling=SamplingConfig(num_inference_steps=4, height=64, width=64,
num_frames=8),
output=OutputConfig(save_video=False, return_frames=False),
)
for key, value in overrides.items():
setattr(base.sampling, key, value)
return base
def _make_result(state: ContinuationState | None = None) -> GenerationResult:
return GenerationResult(
prompt="hi",
frames=[np.zeros((64, 64, 3), dtype=np.uint8)],
samples=None,
generation_time=0.01,
state=state,
)
class TestEventOrdering:
def test_emits_progress_then_final(self):
gen = _FakeVideoGenerator.bind(_make_result())
async def run():
events = []
async for evt in gen.generate_async(_make_request()):
events.append(evt)
return events
events = asyncio.run(run())
assert len(events) == 2
assert isinstance(events[0], VideoProgressEvent)
assert isinstance(events[1], VideoFinalEvent)
assert events[0].step == 0
assert events[0].total_steps == 4
def test_exactly_one_final_event_per_request(self):
gen = _FakeVideoGenerator.bind(_make_result())
async def run():
events = []
async for evt in gen.generate_async(_make_request()):
events.append(evt)
return events
events = asyncio.run(run())
finals = [e for e in events if isinstance(e, VideoFinalEvent)]
assert len(finals) == 1
def test_batch_expansion_emits_one_final_per_result(self):
batch_result = [_make_result(), _make_result()]
gen = _FakeVideoGenerator.bind(batch_result)
async def run():
events = []
async for evt in gen.generate_async(_make_request()):
events.append(evt)
return events
events = asyncio.run(run())
finals = [e for e in events if isinstance(e, VideoFinalEvent)]
assert len(finals) == 2
class TestFinalEventShape:
def test_carries_frames_and_result(self):
result = _make_result()
gen = _FakeVideoGenerator.bind(result)
async def run() -> VideoFinalEvent:
async for evt in gen.generate_async(_make_request()):
if isinstance(evt, VideoFinalEvent):
return evt
raise AssertionError("no final event")
final = asyncio.run(run())
assert final.frames is result.frames
assert final.result is result
assert final.metadata["generation_time"] == 0.01
def test_carries_continuation_state_when_present(self):
state = ContinuationState(
kind="ltx2.v1", payload={"schema_version": 1, "segment_index": 3})
result = _make_result(state=state)
gen = _FakeVideoGenerator.bind(result)
async def run() -> VideoFinalEvent:
async for evt in gen.generate_async(_make_request()):
if isinstance(evt, VideoFinalEvent):
return evt
raise AssertionError("no final")
final = asyncio.run(run())
assert final.continuation_state is state
def test_no_state_when_result_has_none(self):
gen = _FakeVideoGenerator.bind(_make_result())
async def run() -> VideoFinalEvent:
async for evt in gen.generate_async(_make_request()):
if isinstance(evt, VideoFinalEvent):
return evt
raise AssertionError("no final")
final = asyncio.run(run())
assert final.continuation_state is None
class TestHealthCheckRequest:
def test_returns_minimal_workload(self):
from fastvideo.entrypoints.video_generator import VideoGenerator
req = VideoGenerator.default_health_check_request()
assert isinstance(req, GenerationRequest)
assert req.sampling.num_inference_steps == 1
assert req.sampling.num_frames == 8
assert req.sampling.width == 256
assert req.sampling.height == 256
def test_does_not_persist_output(self):
from fastvideo.entrypoints.video_generator import VideoGenerator
req = VideoGenerator.default_health_check_request()
assert req.output.save_video is False
assert req.output.return_frames is False
def test_round_trips_through_normalization(self):
from fastvideo.api.compat import normalize_generation_request
from fastvideo.entrypoints.video_generator import VideoGenerator
req = VideoGenerator.default_health_check_request()
normalized = normalize_generation_request(req)
assert normalized.prompt == "health check"
assert normalized.sampling.num_inference_steps == 1
class TestPublicExports:
def test_new_event_types_available_via_fastvideo_api(self):
import fastvideo.api as api
for name in (
"VideoEvent",
"VideoFinalEvent",
"VideoPartialEvent",
"VideoProgressEvent",
"VideoResult",
):
assert name in api.__all__
def test_video_result_is_generation_result_alias(self):
from fastvideo.api import GenerationResult as Public
from fastvideo.api import VideoResult
assert VideoResult is Public
class TestDynamoStyleHandlerIntegration:
"""Mirror the shape the Dynamo backend package uses."""
def test_async_handler_yields_typed_events_without_internal_imports(self):
# The import set is exactly what the Dynamo backend uses.
from fastvideo.api import ( # noqa: F401
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
VideoFinalEvent,
VideoProgressEvent,
)
gen = _FakeVideoGenerator.bind(_make_result())
async def handler(request_dict, context):
# Adapter: dict -> typed request.
req = GenerationRequest(
prompt=request_dict["prompt"],
sampling=SamplingConfig(
num_inference_steps=1, num_frames=8,
height=256, width=256),
output=OutputConfig(save_video=False, return_frames=False),
)
async for event in gen.generate_async(req):
if isinstance(event, VideoFinalEvent):
yield {"type": "final"}
elif isinstance(event, VideoProgressEvent):
yield {"type": "progress", "step": event.step}
async def drive():
out = []
async for chunk in handler({"prompt": "x"}, None):
out.append(chunk)
return out
chunks = asyncio.run(drive())
types = [c["type"] for c in chunks]
assert "progress" in types
assert types[-1] == "final"
@@ -0,0 +1,261 @@
# SPDX-License-Identifier: Apache-2.0
"""Tests for PR 7.8 streaming auxiliaries.
Covers:
* PromptSafetyFilter gracefully disables when fastText isn't installed
* SafetyResult semantics (allow, block, unavailable)
* RewriteOptions + _split_response parsing behavior
* SessionLogger JSONL append semantics + close lifecycle
* MockServer builds an app that drives the WS protocol end-to-end
"""
from __future__ import annotations
import json
import sys
import types
from pathlib import Path
from typing import Any
import pytest
from fastvideo.entrypoints.streaming.prompt.safety import (
PromptSafetyFilter,
SafetyDecision,
first_blocked,
)
from fastvideo.entrypoints.streaming.prompt.rewrite import (
RewriteOptions,
_split_response,
build_rewrite,
)
from fastvideo.entrypoints.streaming.session_logger import (
SessionLogEvent,
SessionLogger,
)
# ----------------------------------------------------------------------
# Safety
# ----------------------------------------------------------------------
class TestPromptSafetyFilter:
def test_disabled_by_default_when_no_path(self):
f = PromptSafetyFilter(classifier_path=None)
assert f.enabled is False
result = f.classify("hi")
assert result.decision is SafetyDecision.UNAVAILABLE
def test_disabled_when_enabled_false(self):
f = PromptSafetyFilter(classifier_path="/tmp/m.bin", enabled=False)
assert f.enabled is False
def test_unavailable_when_fasttext_missing(self, monkeypatch):
# Force `import fasttext` inside _ensure_loaded to fail.
monkeypatch.setitem(sys.modules, "fasttext", None)
f = PromptSafetyFilter(classifier_path="/tmp/m.bin", enabled=True)
result = f.classify("hi")
assert result.decision is SafetyDecision.UNAVAILABLE
def test_block_when_classifier_flags_unsafe(self, monkeypatch, tmp_path):
fake_model = types.SimpleNamespace(
predict=lambda text, k=1: (["__label__unsafe"], [0.95]))
stub = types.SimpleNamespace(load_model=lambda _p: fake_model)
monkeypatch.setitem(sys.modules, "fasttext", stub)
model_path = str(tmp_path / "m.bin")
Path(model_path).write_text("")
f = PromptSafetyFilter(classifier_path=model_path, enabled=True)
result = f.classify("please")
assert result.decision is SafetyDecision.BLOCK
assert result.label == "unsafe"
assert result.score == pytest.approx(0.95)
def test_allow_when_classifier_flags_safe(self, monkeypatch, tmp_path):
fake_model = types.SimpleNamespace(
predict=lambda text, k=1: (["__label__safe"], [0.99]))
stub = types.SimpleNamespace(load_model=lambda _p: fake_model)
monkeypatch.setitem(sys.modules, "fasttext", stub)
f = PromptSafetyFilter(classifier_path="ignored", enabled=True)
result = f.classify("hello")
assert result.decision is SafetyDecision.ALLOW
def test_below_threshold_allows_even_if_unsafe_label(self, monkeypatch):
fake_model = types.SimpleNamespace(
predict=lambda text, k=1: (["__label__unsafe"], [0.3]))
stub = types.SimpleNamespace(load_model=lambda _p: fake_model)
monkeypatch.setitem(sys.modules, "fasttext", stub)
f = PromptSafetyFilter(
classifier_path="m", enabled=True, block_threshold=0.5)
assert f.classify("x").decision is SafetyDecision.ALLOW
def test_first_blocked_returns_first_hit(self, monkeypatch):
responses = iter([
(["__label__safe"], [0.9]),
(["__label__unsafe"], [0.9]),
(["__label__safe"], [0.9]),
])
fake_model = types.SimpleNamespace(
predict=lambda text, k=1: next(responses))
stub = types.SimpleNamespace(load_model=lambda _p: fake_model)
monkeypatch.setitem(sys.modules, "fasttext", stub)
f = PromptSafetyFilter(classifier_path="m", enabled=True)
blocked = first_blocked(f, ["ok", "bad", "also ok"])
assert blocked is not None
assert blocked.prompt == "bad"
# ----------------------------------------------------------------------
# Rewrite
# ----------------------------------------------------------------------
class TestRewriteSplit:
def test_plain_lines(self):
assert _split_response("one\ntwo\nthree", limit=3) == [
"one", "two", "three",
]
def test_numbered_list(self):
assert _split_response("1. first\n2. second", limit=3) == [
"first", "second",
]
def test_bulleted_list(self):
assert _split_response("- one\n* two\n• three", limit=3) == [
"one", "two", "three",
]
def test_respects_limit(self):
assert _split_response("a\nb\nc\nd", limit=2) == ["a", "b"]
def test_limit_min_one(self):
assert _split_response("only one", limit=0) == ["only one"]
class _StubEnhancer:
async def rewrite(self, seed):
from fastvideo.entrypoints.streaming.prompt.providers.base import LLMResponse
return LLMResponse(
content="1. alpha\n2. beta\n3. gamma",
provider="stub",
model="m",
latency_ms=1.0,
)
class TestBuildRewrite:
def test_empty_seed_rejected(self):
import asyncio
with pytest.raises(ValueError):
asyncio.run(build_rewrite(_StubEnhancer(), " "))
def test_returns_limited_alternatives(self):
import asyncio
result = asyncio.run(build_rewrite(
_StubEnhancer(), "seed",
options=RewriteOptions(count=2)))
assert result.seed_prompt == "seed"
assert result.alternatives == ["alpha", "beta"]
assert result.provider == "stub"
# ----------------------------------------------------------------------
# Session logger
# ----------------------------------------------------------------------
class TestSessionLogger:
def test_no_log_dir_is_noop(self):
logger = SessionLogger(None)
logger.log(SessionLogEvent(session_id="s", event="x")) # no raise
def test_appends_jsonl(self, tmp_path):
logger = SessionLogger(str(tmp_path))
logger.log(SessionLogEvent(
session_id="s1",
event="start",
payload={"preset": "ltx2"},
ts=1.0,
))
logger.log(SessionLogEvent(
session_id="s1",
event="segment",
payload={"idx": 0},
ts=2.0,
))
logger.close("s1")
path = tmp_path / "session-s1.jsonl"
lines = path.read_text().splitlines()
assert len(lines) == 2
first = json.loads(lines[0])
assert first["event"] == "start"
assert first["payload"]["preset"] == "ltx2"
def test_separate_files_per_session(self, tmp_path):
logger = SessionLogger(str(tmp_path))
logger.log(SessionLogEvent(session_id="a", event="e"))
logger.log(SessionLogEvent(session_id="b", event="e"))
logger.close_all()
assert (tmp_path / "session-a.jsonl").exists()
assert (tmp_path / "session-b.jsonl").exists()
# ----------------------------------------------------------------------
# Mock server
# ----------------------------------------------------------------------
class TestMockServer:
def test_build_mock_app_returns_fastapi(self):
from fastapi import FastAPI
from fastvideo.entrypoints.streaming.mock_server import build_mock_app
app = build_mock_app()
assert isinstance(app, FastAPI)
def test_mock_generator_produces_frames(self):
from fastvideo.api.schema import GenerationRequest, SamplingConfig
from fastvideo.entrypoints.streaming.mock_server import MockGenerator
gen = MockGenerator()
result = gen.generate(GenerationRequest(
prompt="x",
sampling=SamplingConfig(
num_frames=3, height=32, width=32, num_inference_steps=1),
))
assert len(result["frames"]) == 3
assert result["frames"][0].shape == (32, 32, 3)
assert result["state"].kind == "ltx2.v1"
def test_mock_app_health_endpoint(self):
from starlette.testclient import TestClient
from fastvideo.entrypoints.streaming.mock_server import build_mock_app
app = build_mock_app()
client = TestClient(app)
assert client.get("/health").json()["status"] == "ok"
def test_mock_app_ws_handshake(self):
from starlette.testclient import TestClient
from fastvideo.entrypoints.streaming.mock_server import build_mock_app
app = build_mock_app()
client = TestClient(app)
with client.websocket_connect("/v1/stream") as ws:
ws.send_json({"type": "session_init_v2"})
assert ws.receive_json()["type"] == "queue_status"
assert ws.receive_json()["type"] == "gpu_assigned"
assert ws.receive_json()["type"] == "ltx2_stream_start"
@@ -0,0 +1,463 @@
# SPDX-License-Identifier: Apache-2.0
"""GPU pool tests.
InProcessGpuPool is exercised end-to-end. SubprocessGpuPool is driven
with an injected ``worker_factory`` that stands up a fake worker inside
a thread (not a subprocess) so the test suite stays CPU-only.
"""
from __future__ import annotations
import asyncio
import multiprocessing as mp
import queue
import threading
import time
from dataclasses import dataclass
from typing import Any
import pytest
from fastvideo.api.schema import (
GeneratorConfig,
GenerationRequest,
GpuPoolConfig,
WarmupConfig,
)
from fastvideo.entrypoints.streaming.gpu_pool import (
GpuPool,
InProcessGpuPool,
PoolAcquireTimeout,
SubprocessGpuPool,
_WorkerHandle,
)
# ----------------------------------------------------------------------
# In-process pool
# ----------------------------------------------------------------------
@dataclass
class _MockGenerator:
sleep_s: float = 0.0
def generate(self, request: GenerationRequest) -> dict[str, Any]:
if self.sleep_s:
time.sleep(self.sleep_s)
return {
"frames": [],
"prompt_echo": request.prompt,
}
class TestInProcessGpuPool:
def test_is_gpu_pool(self):
assert isinstance(
InProcessGpuPool(_MockGenerator()), GpuPool)
def test_acquire_returns_deterministic_assignment(self):
pool = InProcessGpuPool(_MockGenerator(), gpu_id=7)
async def run():
a = await pool.acquire("sess-a")
return a
a = asyncio.run(run())
assert a.gpu_id == 7
assert a.worker_id.startswith("inproc-")
def test_acquire_is_sticky_across_calls(self):
pool = InProcessGpuPool(_MockGenerator())
async def run():
a = await pool.acquire("sess-a")
b = await pool.acquire("sess-a")
return a, b
a, b = asyncio.run(run())
assert a == b
def test_run_without_acquire_raises(self):
pool = InProcessGpuPool(_MockGenerator())
async def run():
with pytest.raises(RuntimeError):
await pool.run(
"sess-a", GenerationRequest(prompt="hi"))
asyncio.run(run())
def test_run_returns_generator_output(self):
pool = InProcessGpuPool(_MockGenerator())
async def run():
await pool.acquire("sess-a")
return await pool.run(
"sess-a", GenerationRequest(prompt="hi"))
result = asyncio.run(run())
assert result["prompt_echo"] == "hi"
def test_release_frees_binding(self):
pool = InProcessGpuPool(_MockGenerator())
async def run():
await pool.acquire("sess-a")
await pool.release("sess-a")
with pytest.raises(RuntimeError):
await pool.run(
"sess-a", GenerationRequest(prompt="hi"))
asyncio.run(run())
def test_health_reports_active_sessions(self):
pool = InProcessGpuPool(_MockGenerator())
async def run():
await pool.acquire("sess-a")
health = pool.health()
assert health.total_workers == 1
assert health.active_sessions == 1
assert health.available_workers == 0
asyncio.run(run())
# ----------------------------------------------------------------------
# Subprocess pool (driven by a thread-backed fake worker factory)
# ----------------------------------------------------------------------
class _ThreadWorker:
"""Stand-in for a subprocess worker.
Runs a Python thread that pulls jobs from ``job_queue`` and invokes
a supplied mock generator. The control flow matches
:func:`worker_main` exactly (ready + result dict shapes) so the
parent-side pool under test exercises the same code paths.
"""
def __init__(
self,
generator: _MockGenerator,
*,
job_queue: mp.Queue,
result_queue: mp.Queue,
shutdown_event: threading.Event,
warmup_ms: float = 0.0,
) -> None:
self._generator = generator
self._job_queue = job_queue
self._result_queue = result_queue
self._shutdown_event = shutdown_event
self._warmup_ms = warmup_ms
self._thread = threading.Thread(target=self._run, daemon=True)
def start(self) -> None:
self._thread.start()
def join(self, timeout: float | None = None) -> None:
self._thread.join(timeout)
def _run(self) -> None:
if self._warmup_ms:
time.sleep(self._warmup_ms / 1000.0)
self._result_queue.put({"kind": "ready"})
while not self._shutdown_event.is_set():
try:
item = self._job_queue.get(timeout=0.1)
except queue.Empty:
continue
if item is None:
return
try:
result = self._generator.generate(item["request"])
self._result_queue.put({
"kind": "result",
"job_id": item["job_id"],
"result": result,
})
except Exception as exc: # pragma: no cover - defensive
self._result_queue.put({
"kind": "error",
"job_id": item["job_id"],
"error": repr(exc),
})
def _thread_worker_factory(generator_builder):
"""Return a WorkerFactory that uses thread workers instead of procs."""
def factory(
*,
gpu_id: int,
generator_config: GeneratorConfig,
warmup_config: WarmupConfig,
) -> _WorkerHandle:
ctx = mp.get_context("spawn")
job_queue: mp.Queue = ctx.Queue()
result_queue: mp.Queue = ctx.Queue()
shutdown_event = threading.Event()
ready = threading.Event()
boot_ok = threading.Event()
mp_shutdown = ctx.Event()
generator = generator_builder(gpu_id)
worker = _ThreadWorker(
generator,
job_queue=job_queue,
result_queue=result_queue,
shutdown_event=shutdown_event,
)
worker.start()
# Drain the ready sentinel from the queue in the same way the
# real factory does (thread waiter populates ``ready`` /
# ``boot_ok``).
def _await_ready() -> None:
while True:
try:
msg = result_queue.get(timeout=1.0)
except queue.Empty:
if shutdown_event.is_set():
return
continue
if msg.get("kind") == "ready":
boot_ok.set()
ready.set()
return
if msg.get("kind") == "error":
ready.set()
return
threading.Thread(target=_await_ready, daemon=True).start()
class _FakeProcess:
def __init__(self, stop: threading.Event, worker: _ThreadWorker):
self._stop = stop
self._worker = worker
def is_alive(self) -> bool:
return self._worker._thread.is_alive()
def join(self, timeout: float | None = None) -> None:
self._stop.set()
self._worker.join(timeout)
def kill(self) -> None:
self._stop.set()
fake_process = _FakeProcess(shutdown_event, worker)
return _WorkerHandle(
process=fake_process, # type: ignore[arg-type]
job_queue=job_queue,
result_queue=result_queue,
gpu_id=gpu_id,
worker_id=f"gpu{gpu_id}-fake",
ready=ready,
boot_ok=boot_ok,
shutdown_event=mp_shutdown,
)
return factory
@pytest.fixture
def pool_factory():
"""Provide a SubprocessGpuPool built against thread workers."""
async def _build(num_workers: int = 2):
pool = SubprocessGpuPool(
generator_config=GeneratorConfig(model_path="/models/fake"),
pool_config=GpuPoolConfig(num_workers=num_workers),
warmup_config=WarmupConfig(enabled=False),
worker_factory=_thread_worker_factory(
lambda gpu_id: _MockGenerator()),
)
await pool.start()
return pool
return _build
class TestSubprocessGpuPool:
def test_start_spawns_requested_workers(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=3)
try:
h = pool.health()
assert h.total_workers == 3
assert h.available_workers == 3
finally:
await pool.shutdown()
asyncio.run(run())
def test_acquire_decrements_available(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=2)
try:
a = await pool.acquire("sess-a")
assert a.worker_id.endswith("-fake")
assert pool.health().available_workers == 1
finally:
await pool.shutdown()
asyncio.run(run())
def test_acquire_timeout_when_all_busy(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=1)
try:
await pool.acquire("sess-a")
with pytest.raises(PoolAcquireTimeout):
await pool.acquire("sess-b", timeout=0.1)
finally:
await pool.shutdown()
asyncio.run(run())
def test_run_returns_worker_result(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=1)
try:
await pool.acquire("sess-a")
result = await pool.run(
"sess-a", GenerationRequest(prompt="hello"))
assert result["prompt_echo"] == "hello"
finally:
await pool.shutdown()
asyncio.run(run())
def test_release_returns_worker_to_pool(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=1)
try:
await pool.acquire("sess-a")
await pool.release("sess-a")
# Now a second acquire should succeed without timeout.
a = await pool.acquire("sess-b", timeout=1.0)
assert a.worker_id.endswith("-fake")
finally:
await pool.shutdown()
asyncio.run(run())
def test_sticky_binding_across_multiple_runs(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=2)
try:
a1 = await pool.acquire("sess-a")
a2 = await pool.acquire("sess-a")
assert a1.worker_id == a2.worker_id
# Two runs land on the same worker.
await pool.run("sess-a", GenerationRequest(prompt="1"))
await pool.run("sess-a", GenerationRequest(prompt="2"))
finally:
await pool.shutdown()
asyncio.run(run())
def test_run_without_acquire_raises(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=1)
try:
with pytest.raises(RuntimeError):
await pool.run(
"sess-x", GenerationRequest(prompt="x"))
finally:
await pool.shutdown()
asyncio.run(run())
def test_shutdown_is_idempotent(self, pool_factory):
async def run():
pool = await pool_factory(num_workers=2)
await pool.shutdown()
await pool.shutdown() # no raise
asyncio.run(run())
class TestSubprocessGpuPoolFailureModes:
"""Coverage for boot/runtime failures the pool has to absorb."""
def test_failed_boot_excluded_from_available(self):
"""A worker whose factory leaves ``boot_ok`` unset must not be
handed out by ``acquire``. Otherwise a session lands on a dead
worker and ``run`` blocks forever on the missing result."""
def factory_with_one_failure(*, gpu_id, generator_config, warmup_config):
ctx = mp.get_context("spawn")
job_queue: mp.Queue = ctx.Queue()
result_queue: mp.Queue = ctx.Queue()
ready = threading.Event()
boot_ok = threading.Event()
mp_shutdown = ctx.Event()
ready.set()
# gpu_id 0 boots fine; gpu_id 1 fails (boot_ok stays clear).
if gpu_id == 0:
boot_ok.set()
class _AliveProcess:
def is_alive(self) -> bool:
return True
def join(self, timeout: float | None = None) -> None:
return
def kill(self) -> None:
return
return _WorkerHandle(
process=_AliveProcess(), # type: ignore[arg-type]
job_queue=job_queue,
result_queue=result_queue,
gpu_id=gpu_id,
worker_id=f"gpu{gpu_id}",
ready=ready,
boot_ok=boot_ok,
shutdown_event=mp_shutdown,
)
async def run():
pool = SubprocessGpuPool(
generator_config=GeneratorConfig(model_path="/m"),
pool_config=GpuPoolConfig(num_workers=2),
warmup_config=WarmupConfig(enabled=False),
worker_factory=factory_with_one_failure,
)
await pool.start()
try:
# Only worker 0 booted; only one slot should be available.
assert pool.health().available_workers == 1
a = await pool.acquire("sess-a", timeout=0.1)
assert a.worker_id == "gpu0"
with pytest.raises(PoolAcquireTimeout):
await pool.acquire("sess-b", timeout=0.1)
finally:
await pool.shutdown()
asyncio.run(run())
def test_release_skips_dead_worker(self, pool_factory):
"""If a worker died while bound, releasing the session must not
return its slot to the available queue — a later acquire would
hand the dead slot to a new session."""
async def run():
pool = await pool_factory(num_workers=1)
try:
await pool.acquire("sess-a")
# Simulate the worker dying mid-session.
pool._workers[0].process._stop.set()
pool._workers[0].process._worker.join(timeout=1.0)
await pool.release("sess-a")
# Available queue must remain empty.
assert pool.health().available_workers == 0
finally:
await pool.shutdown()
asyncio.run(run())
@@ -0,0 +1,207 @@
# SPDX-License-Identifier: Apache-2.0
"""Tests for the provider-agnostic prompt enhancer."""
from __future__ import annotations
import asyncio
from dataclasses import dataclass
import pytest
from fastvideo.entrypoints.streaming.prompt import (
LLMProvider,
LLMProviderError,
LLMRequest,
LLMResponse,
PromptEnhancer,
)
@dataclass
class _StaticProvider:
name: str
content: str
async def complete(self, request: LLMRequest) -> LLMResponse:
return LLMResponse(
content=self.content,
provider=self.name,
model=request.model,
latency_ms=1.0,
)
@dataclass
class _FailingProvider:
name: str
message: str = "boom"
retryable: bool = True
async def complete(self, request: LLMRequest) -> LLMResponse:
raise LLMProviderError(self.message, retryable=self.retryable)
class TestConstruction:
def test_requires_at_least_one_provider(self):
with pytest.raises(ValueError):
PromptEnhancer(providers=[], model="m")
def test_registers_system_prompts_from_defaults(self, tmp_path):
enh = PromptEnhancer(
providers=[_StaticProvider("p", "ok")],
model="m",
)
# Defaults live inside _DEFAULT_SYSTEM_PROMPTS; no hot-reload file
# means the enhancer falls back to the shipped values.
assert "prompt enhancer" in enh._system_prompts.enhance.lower()
def test_reads_override_files_when_present(self, tmp_path):
(tmp_path / "enhance.txt").write_text("customized enhance")
(tmp_path / "auto_extend.txt").write_text("") # empty -> default
enh = PromptEnhancer(
providers=[_StaticProvider("p", "ok")],
model="m",
system_prompt_dir=str(tmp_path),
)
assert enh._system_prompts.enhance == "customized enhance"
# Empty file fell back to the default.
assert "continuation assistant" in enh._system_prompts.auto_extend
class TestEnhance:
def test_enhance_returns_primary_provider_response(self):
enh = PromptEnhancer(
providers=[_StaticProvider("primary", "enhanced!")],
model="m",
)
response = asyncio.run(enh.enhance("a fox"))
assert response.content == "enhanced!"
assert response.provider == "primary"
assert response.fallback_used is False
def test_enhance_falls_back_on_provider_error(self):
enh = PromptEnhancer(
providers=[
_FailingProvider("a"),
_StaticProvider("b", "fallback!"),
],
model="m",
)
response = asyncio.run(enh.enhance("x"))
assert response.provider == "b"
assert response.content == "fallback!"
assert response.fallback_used is True
def test_enhance_stops_on_non_retryable_error(self):
enh = PromptEnhancer(
providers=[
_FailingProvider("a", message="hard-fail", retryable=False),
_StaticProvider("b", "should-not-be-reached"),
],
model="m",
)
with pytest.raises(LLMProviderError, match="hard-fail"):
asyncio.run(enh.enhance("x"))
def test_enhance_raises_when_all_providers_fail(self):
enh = PromptEnhancer(
providers=[
_FailingProvider("a"),
_FailingProvider("b"),
],
model="m",
)
with pytest.raises(LLMProviderError):
asyncio.run(enh.enhance("x"))
class TestAutoExtendAndRewrite:
def test_auto_extend_joins_prior_prompts(self):
captured: list[str] = []
class _Capturer:
name = "cap"
async def complete(self, request: LLMRequest) -> LLMResponse:
user = next(m.content for m in request.messages
if m.role == "user")
captured.append(user)
return LLMResponse(
content="next",
provider=self.name,
model=request.model,
latency_ms=1.0,
)
enh = PromptEnhancer(providers=[_Capturer()], model="m")
asyncio.run(enh.auto_extend(["first", "second"]))
assert captured[0] == "first\nsecond"
def test_rewrite_passes_seed_through(self):
captured: list[str] = []
class _Capturer:
name = "cap"
async def complete(self, request: LLMRequest) -> LLMResponse:
user = next(m.content for m in request.messages
if m.role == "user")
captured.append(user)
return LLMResponse(
content="one\ntwo\nthree",
provider=self.name,
model=request.model,
latency_ms=1.0,
)
enh = PromptEnhancer(providers=[_Capturer()], model="m")
response = asyncio.run(enh.rewrite("seed prompt"))
assert captured == ["seed prompt"]
assert response.content.splitlines() == ["one", "two", "three"]
class TestRegisterProvider:
def test_register_appends_by_default(self):
enh = PromptEnhancer(
providers=[_StaticProvider("primary", "a")],
model="m",
)
enh.register_provider(_StaticProvider("extra", "b"))
assert [p.name for p in enh.providers] == ["primary", "extra"]
def test_register_priority_zero_makes_primary(self):
enh = PromptEnhancer(
providers=[_StaticProvider("old", "a")],
model="m",
)
enh.register_provider(
_StaticProvider("new-primary", "b"), priority=0)
assert enh.providers[0].name == "new-primary"
def test_registered_provider_is_used_in_enhance(self):
enh = PromptEnhancer(
providers=[_FailingProvider("broken")],
model="m",
)
enh.register_provider(_StaticProvider("fallback", "ok"))
response = asyncio.run(enh.enhance("x"))
assert response.content == "ok"
assert response.fallback_used is True
class TestHotReload:
def test_reload_picks_up_new_file(self, tmp_path):
(tmp_path / "enhance.txt").write_text("first version")
enh = PromptEnhancer(
providers=[_StaticProvider("p", "ok")],
model="m",
system_prompt_dir=str(tmp_path),
)
assert enh._system_prompts.enhance == "first version"
(tmp_path / "enhance.txt").write_text("second version")
enh.reload_system_prompts()
assert enh._system_prompts.enhance == "second version"
@@ -0,0 +1,298 @@
# SPDX-License-Identifier: Apache-2.0
"""Tests for the LLM provider protocol + built-in adapters.
Real Cerebras/Groq HTTP calls are stubbed with a fake ``httpx`` module
inserted into ``sys.modules`` so the unit tests don't depend on
external API availability or paid keys.
"""
from __future__ import annotations
import asyncio
import sys
import types
from dataclasses import dataclass
from typing import Any
import pytest
from fastvideo.entrypoints.streaming.prompt.providers import (
CerebrasProvider,
GroqProvider,
LLMMessage,
LLMProvider,
LLMProviderError,
LLMRequest,
LLMTimeoutError,
)
@dataclass
class _FakeResponse:
status_code: int
payload: dict[str, Any]
text: str = ""
def json(self) -> dict[str, Any]:
return self.payload
class _FakeAsyncClient:
def __init__(self, response_or_exc, *, captured: list) -> None:
self._response_or_exc = response_or_exc
self._captured = captured
async def __aenter__(self):
return self
async def __aexit__(self, exc_type, exc, tb):
return None
async def post(self, url, **kwargs):
self._captured.append({"url": url, **kwargs})
if isinstance(self._response_or_exc, Exception):
raise self._response_or_exc
return self._response_or_exc
def _install_fake_httpx(
monkeypatch,
*,
response: _FakeResponse | None = None,
exception: Exception | None = None,
) -> list[dict]:
"""Replace ``sys.modules['httpx']`` with a stub exposing the bits
providers touch: ``AsyncClient``, ``HTTPError``, ``TimeoutException``.
Returns a list the test can inspect to see what requests went out.
"""
captured: list[dict] = []
class _HTTPError(Exception):
pass
class _TimeoutException(_HTTPError):
pass
payload = response if exception is None else exception
def _client_factory(*_args, **_kwargs):
return _FakeAsyncClient(payload, captured=captured)
stub = types.SimpleNamespace(
AsyncClient=_client_factory,
HTTPError=_HTTPError,
TimeoutException=_TimeoutException,
)
monkeypatch.setitem(sys.modules, "httpx", stub)
return captured
# ----------------------------------------------------------------------
# Cerebras
# ----------------------------------------------------------------------
class TestCerebrasProvider:
def test_is_llm_provider(self):
assert isinstance(CerebrasProvider(api_key="x"), LLMProvider)
def test_requires_api_key(self, monkeypatch):
monkeypatch.delenv("CEREBRAS_API_KEY", raising=False)
provider = CerebrasProvider()
with pytest.raises(LLMProviderError, match="CEREBRAS_API_KEY"):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
def test_api_key_from_env(self, monkeypatch):
monkeypatch.setenv("CEREBRAS_API_KEY", "from-env")
provider = CerebrasProvider()
assert provider.api_key == "from-env"
def test_explicit_api_key_wins(self, monkeypatch):
monkeypatch.setenv("CEREBRAS_API_KEY", "from-env")
provider = CerebrasProvider(api_key="explicit")
assert provider.api_key == "explicit"
def test_success_path(self, monkeypatch):
captured = _install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=200,
payload={
"choices": [{"message": {"content": "enhanced"}}],
},
),
)
provider = CerebrasProvider(api_key="secret")
result = asyncio.run(provider.complete(
LLMRequest(
messages=[LLMMessage(role="user", content="a fox")],
model="gpt-oss-120b",
max_tokens=64,
temperature=0.7,
)))
assert result.content == "enhanced"
assert result.provider == "cerebras"
assert result.model == "gpt-oss-120b"
# Verify the HTTP request was shaped correctly.
body = captured[0]["json"]
assert body["model"] == "gpt-oss-120b"
assert body["messages"] == [{"role": "user", "content": "a fox"}]
assert captured[0]["headers"]["Authorization"] == "Bearer secret"
def test_http_error_wrapped(self, monkeypatch):
# Stage the stub first, then raise that stub's own HTTPError so
# the provider's ``except httpx.HTTPError`` catches it.
stub = types.SimpleNamespace()
class _HTTPError(Exception):
pass
class _TimeoutException(_HTTPError):
pass
stub.HTTPError = _HTTPError
stub.TimeoutException = _TimeoutException
stub.AsyncClient = lambda *a, **k: _FakeAsyncClient(
_HTTPError("boom"), captured=[])
monkeypatch.setitem(sys.modules, "httpx", stub)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="HTTP"):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
def test_timeout_wrapped(self, monkeypatch):
_install_fake_httpx(monkeypatch)
# Install timeout exception after stubbing.
stub = sys.modules["httpx"]
def raising_factory(*_a, **_kw):
raise stub.TimeoutException("timed out") # type: ignore[attr-defined]
stub.AsyncClient = raising_factory # type: ignore[attr-defined]
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMTimeoutError):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
def test_4xx_raises_non_retryable_provider_error(self, monkeypatch):
_install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=401,
payload={},
text="unauthorized",
),
)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="401") as excinfo:
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
assert excinfo.value.retryable is False
def test_429_raises_retryable_provider_error(self, monkeypatch):
_install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=429,
payload={},
text="rate limited",
),
)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="429") as excinfo:
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
assert excinfo.value.retryable is True
def test_5xx_raises_retryable_provider_error(self, monkeypatch):
_install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=503,
payload={},
text="service unavailable",
),
)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="503") as excinfo:
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
assert excinfo.value.retryable is True
def test_non_json_body_raises_provider_error(self, monkeypatch):
# Simulate a proxy/load-balancer HTML error page slipping in
# with a 200 status: response.json() raises, and the provider
# must wrap it in an LLMProviderError instead of bubbling.
class _BadJsonResponse:
status_code = 200
text = "<html>oops</html>"
def json(self):
raise ValueError("Expecting value: line 1 column 1 (char 0)")
_install_fake_httpx(
monkeypatch,
response=_BadJsonResponse(), # type: ignore[arg-type]
)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="non-JSON"):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
def test_empty_choices_raises(self, monkeypatch):
_install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=200, payload={"choices": []}),
)
provider = CerebrasProvider(api_key="k")
with pytest.raises(LLMProviderError, match="no choices"):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
# ----------------------------------------------------------------------
# Groq
# ----------------------------------------------------------------------
class TestGroqProvider:
def test_is_llm_provider(self):
assert isinstance(GroqProvider(api_key="x"), LLMProvider)
def test_requires_api_key(self, monkeypatch):
monkeypatch.delenv("GROQ_API_KEY", raising=False)
provider = GroqProvider()
with pytest.raises(LLMProviderError, match="GROQ_API_KEY"):
asyncio.run(provider.complete(
LLMRequest(messages=[], model="m")))
def test_api_key_from_env(self, monkeypatch):
monkeypatch.setenv("GROQ_API_KEY", "from-env")
provider = GroqProvider()
assert provider.api_key == "from-env"
def test_success_path(self, monkeypatch):
captured = _install_fake_httpx(
monkeypatch,
response=_FakeResponse(
status_code=200,
payload={
"choices": [{"message": {"content": "groq out"}}],
},
),
)
provider = GroqProvider(api_key="secret")
result = asyncio.run(provider.complete(
LLMRequest(
messages=[LLMMessage(role="user", content="a deer")],
model="llama-3.1-70b",
)))
assert result.content == "groq out"
assert result.provider == "groq"
assert captured[0]["headers"]["Authorization"] == "Bearer secret"
@@ -0,0 +1,279 @@
# SPDX-License-Identifier: Apache-2.0
"""Router tests — registry semantics + health loop behavior.
Avoids real WebSocket proxying (the bridge requires the ``websockets``
package and two running uvicorn processes); test_server.py already
covers the end-to-end WS protocol against a direct backend.
"""
from __future__ import annotations
import asyncio
import pytest
from fastvideo.entrypoints.streaming.router.config import (
ReplicaEndpoint,
RouterConfig,
)
from fastvideo.entrypoints.streaming.router.registry import (
ReplicaRegistry,
ReplicaStatus,
run_health_check_loop,
)
def _registry(
*,
num_primary: int = 1,
num_secondary: int = 1,
) -> ReplicaRegistry:
replicas = [
ReplicaEndpoint(
url=f"http://primary-{i}:8000",
primary=True,
)
for i in range(num_primary)
] + [
ReplicaEndpoint(
url=f"http://secondary-{i}:8000",
primary=False,
)
for i in range(num_secondary)
]
return ReplicaRegistry(replicas)
class TestReplicaRegistry:
def test_requires_replicas(self):
with pytest.raises(ValueError):
ReplicaRegistry([])
def test_select_none_when_all_unknown(self):
assert _registry().select() is None
def test_select_prefers_healthy_primary(self):
reg = _registry()
async def promote():
primary = reg.primaries()[0]
await reg.record_success(
primary, recovery_threshold=1, latency_ms=1.0)
# Also mark secondary healthy — primary should still win.
sec = next(r for r in reg.all() if not r.primary)
await reg.record_success(
sec, recovery_threshold=1, latency_ms=1.0)
return reg.select()
pick = asyncio.run(promote())
assert pick is not None
assert pick.primary
def test_falls_back_to_secondary_when_primary_unhealthy(self):
reg = _registry()
async def run():
primary = reg.primaries()[0]
sec = next(r for r in reg.all() if not r.primary)
# Fail primary past threshold, succeed secondary.
for _ in range(3):
await reg.record_failure(
primary, failure_threshold=3, reason="mock")
await reg.record_success(
sec, recovery_threshold=1, latency_ms=1.0)
return reg.select()
pick = asyncio.run(run())
assert pick is not None
assert not pick.primary
def test_failure_threshold_transitions_to_unhealthy(self):
reg = _registry()
primary = reg.primaries()[0]
async def run():
for _ in range(2):
await reg.record_failure(
primary, failure_threshold=3, reason="x")
assert primary.health.status is not ReplicaStatus.UNHEALTHY
await reg.record_failure(
primary, failure_threshold=3, reason="x")
return primary
result = asyncio.run(run())
assert result.health.status is ReplicaStatus.UNHEALTHY
assert result.health.consecutive_failures == 3
def test_recovery_threshold_returns_to_healthy(self):
reg = _registry()
primary = reg.primaries()[0]
async def run():
for _ in range(3):
await reg.record_failure(
primary, failure_threshold=3, reason="x")
assert primary.health.status is ReplicaStatus.UNHEALTHY
for _ in range(2):
await reg.record_success(
primary, recovery_threshold=2, latency_ms=5.0)
return primary
result = asyncio.run(run())
assert result.health.status is ReplicaStatus.HEALTHY
assert result.health.last_latency_ms == 5.0
def test_record_success_resets_failure_counter(self):
reg = _registry()
primary = reg.primaries()[0]
async def run():
await reg.record_failure(
primary, failure_threshold=10, reason="x")
await reg.record_success(
primary, recovery_threshold=1, latency_ms=1.0)
return primary
result = asyncio.run(run())
assert result.health.consecutive_failures == 0
class TestHealthCheckLoop:
def test_loop_transitions_replicas_on_probe_results(self):
config = RouterConfig(
replicas=[
ReplicaEndpoint(url="http://a", primary=True),
ReplicaEndpoint(url="http://b"),
],
health_check_interval_seconds=0.01,
failure_threshold=1,
recovery_threshold=1,
)
reg = ReplicaRegistry(config.replicas)
stop_event = asyncio.Event()
calls: list[str] = []
async def probe(url, *, timeout):
calls.append(url)
if "http://a" in url:
return 1.0, None
return 0.0, "mock failure"
async def run() -> None:
task = asyncio.create_task(run_health_check_loop(
registry=reg, config=config, stop_event=stop_event,
http_get=probe,
))
await asyncio.sleep(0.05)
stop_event.set()
await task
asyncio.run(run())
a = reg.get("http://a")
b = reg.get("http://b")
assert a is not None
assert b is not None
assert a.health.status is ReplicaStatus.HEALTHY
assert b.health.status is ReplicaStatus.UNHEALTHY
assert any("/health" in c for c in calls)
class TestRouterApp:
def test_status_endpoint_lists_replicas(self):
from starlette.testclient import TestClient
from fastvideo.entrypoints.streaming.router.main import build_router_app
config = RouterConfig(
replicas=[
ReplicaEndpoint(url="http://a", primary=True),
ReplicaEndpoint(url="http://b"),
],
health_check_interval_seconds=60, # don't actually poll
)
reg = ReplicaRegistry(config.replicas)
app = build_router_app(config, registry=reg)
client = TestClient(app)
response = client.get("/status")
body = response.json()
urls = {r["url"] for r in body["replicas"]}
assert urls == {"http://a", "http://b"}
# Initial status is UNKNOWN.
assert all(r["status"] == "unknown" for r in body["replicas"])
def test_ws_rejects_when_no_healthy_replica(self):
from starlette.testclient import TestClient
from fastvideo.entrypoints.streaming.router.main import build_router_app
config = RouterConfig(
replicas=[ReplicaEndpoint(url="http://a", primary=True)],
health_check_interval_seconds=60,
)
reg = ReplicaRegistry(config.replicas)
app = build_router_app(config, registry=reg)
client = TestClient(app)
with client.websocket_connect("/v1/stream") as ws:
err = ws.receive_json()
assert err["type"] == "error"
assert err["code"] == "gpu_unavailable"
class TestUnknownToHealthyImmediate:
"""Initial probe must promote UNKNOWN -> HEALTHY without waiting for recovery_threshold."""
def test_first_success_promotes_unknown(self):
reg = _registry(num_primary=1, num_secondary=0)
primary = reg.primaries()[0]
assert primary.health.status is ReplicaStatus.UNKNOWN
async def run():
await reg.record_success(primary, recovery_threshold=10, latency_ms=1.0)
asyncio.run(run())
assert primary.health.status is ReplicaStatus.HEALTHY
def test_unhealthy_recovery_still_gated_by_threshold(self):
reg = _registry(num_primary=1, num_secondary=0)
primary = reg.primaries()[0]
async def run():
for _ in range(3):
await reg.record_failure(primary, failure_threshold=3, reason="x")
assert primary.health.status is ReplicaStatus.UNHEALTHY
await reg.record_success(primary, recovery_threshold=2, latency_ms=1.0)
assert primary.health.status is ReplicaStatus.UNHEALTHY # 1/2 successes
await reg.record_success(primary, recovery_threshold=2, latency_ms=1.0)
assert primary.health.status is ReplicaStatus.HEALTHY # 2/2 successes
asyncio.run(run())
class TestConfigValidation:
"""RouterConfig.__post_init__ rejects malformed configs."""
def test_rejects_path_in_url(self):
with pytest.raises(ValueError, match="without a path"):
RouterConfig(replicas=[ReplicaEndpoint(url="http://host:8000/api")])
def test_rejects_query_in_url(self):
with pytest.raises(ValueError, match="query/fragment"):
RouterConfig(replicas=[ReplicaEndpoint(url="http://host:8000?x=1")])
def test_rejects_fragment_in_url(self):
with pytest.raises(ValueError, match="query/fragment"):
RouterConfig(replicas=[ReplicaEndpoint(url="http://host:8000#frag")])
def test_rejects_duplicate_urls(self):
with pytest.raises(ValueError, match="Duplicate"):
RouterConfig(replicas=[
ReplicaEndpoint(url="http://host:8000"),
ReplicaEndpoint(url="http://host:8000"),
])
def test_accepts_trailing_slash(self):
# parsed.path == "/" should be allowed
cfg = RouterConfig(replicas=[ReplicaEndpoint(url="http://host:8000/")])
assert cfg.replicas[0].url == "http://host:8000/"
@@ -0,0 +1,86 @@
# SPDX-License-Identifier: Apache-2.0
"""Unit coverage for :mod:`fastvideo.entrypoints.streaming.worker`.
The full ``worker_main`` loop runs in a subprocess and is exercised via
``test_gpu_pool.py``'s subprocess integration tests. This file covers
the in-process pieces:
* the two-segment warmup feeds segment 1's continuation state into
segment 2 so both compile branches are primed before the worker
reports ready
* result-shape extractors handle both attribute-style and dict-style
generator returns
"""
from __future__ import annotations
from dataclasses import dataclass, field
from typing import Any
from fastvideo.api.schema import (
ContinuationState,
GenerationRequest,
WarmupConfig,
)
from fastvideo.entrypoints.streaming.worker import (
_extract_continuation_state,
_warmup_worker,
)
@dataclass
class _RecordingGenerator:
"""Captures the requests passed to ``generate`` and returns a
canned :class:`ContinuationState` on the first call so the warmup
can feed it into the second call.
"""
state_to_return: ContinuationState | None = field(default_factory=lambda: ContinuationState(
kind="ltx2.v1",
payload={"schema_version": 1, "segment_index": 1},
))
requests: list[GenerationRequest] = field(default_factory=list)
def generate(self, request: GenerationRequest) -> dict[str, Any]:
self.requests.append(request)
return {"frames": [], "state": self.state_to_return}
class TestWarmupTwoSegment:
def test_warmup_runs_segment_one_then_segment_two_with_returned_state(self) -> None:
gen = _RecordingGenerator()
_warmup_worker(gen, WarmupConfig(enabled=True, prompt="warm"))
assert len(gen.requests) == 2
seg1 = gen.requests[0]
assert seg1.state is None
assert seg1.output.return_state is True
seg2 = gen.requests[1]
assert seg2.state is gen.state_to_return
def test_warmup_passes_through_when_no_state_returned(self) -> None:
gen = _RecordingGenerator(state_to_return=None)
_warmup_worker(gen, WarmupConfig(enabled=True, prompt="warm"))
assert len(gen.requests) == 2
assert gen.requests[1].state is None
class TestExtractContinuationState:
def test_extracts_from_attribute(self) -> None:
class _R:
state = ContinuationState(kind="k", payload={})
assert _extract_continuation_state(_R()).kind == "k"
def test_extracts_from_dict(self) -> None:
state = ContinuationState(kind="k", payload={})
assert _extract_continuation_state({"state": state}) is state
def test_returns_none_for_missing(self) -> None:
assert _extract_continuation_state({}) is None
assert _extract_continuation_state(object()) is None
@@ -0,0 +1,65 @@
# SPDX-License-Identifier: Apache-2.0
"""NVFP4Config import + lazy-flashinfer behavior tests.
The actual NVFP4 kernels need flashinfer + a CUDA device, so these
tests focus on the lazy-import contract: the module must load on
hosts without flashinfer, and only call sites should fail with a
clear error when flashinfer is missing.
"""
from __future__ import annotations
import sys
import types
import pytest
def test_nvfp4config_imports_without_flashinfer(monkeypatch):
"""Importing the module on a host without flashinfer must succeed.
This is the contract Dreamverse depends on — the GPU worker boots
even when flashinfer is not in the venv, because the import happens
in `video_generation.py` before any NVFP4 kernel is invoked.
"""
# Hide flashinfer from sys.modules and the import path.
monkeypatch.setitem(sys.modules, "flashinfer", None)
# Force a re-import of the target module.
sys.modules.pop("fastvideo.layers.quantization.nvfp4_config", None)
from fastvideo.layers.quantization.nvfp4_config import NVFP4Config
config = NVFP4Config()
assert config.get_name() == "nvfp4"
assert config.layer_profile == "refine"
def test_nvfp4config_layer_profile_round_trips_from_dict():
from fastvideo.layers.quantization.nvfp4_config import NVFP4Config
config = NVFP4Config.from_config({"layer_profile": "base"})
assert config.layer_profile == "base"
config = NVFP4Config.from_config({})
assert config.layer_profile == "refine"
def test_nvfp4_kernel_call_raises_clear_error_without_flashinfer(monkeypatch):
"""A call into the NVFP4 kernels must raise an actionable
ImportError when flashinfer is missing, not a confusing
AttributeError or NameError."""
# Stage a fake flashinfer that fails on import.
monkeypatch.setitem(sys.modules, "flashinfer",
_raise_module_on_import("flashinfer"))
sys.modules.pop("fastvideo.layers.quantization.nvfp4_config", None)
from fastvideo.layers.quantization.nvfp4_config import _require_flashinfer
with pytest.raises(ImportError, match="flashinfer"):
_require_flashinfer()
def _raise_module_on_import(name: str) -> types.ModuleType:
"""Build a stub module that raises ImportError on any attribute
access, so `from <name> import X` fails the way a missing package
would."""
class _RaisingModule(types.ModuleType):
def __getattr__(self, item: str):
raise ImportError(f"No module named '{name}.{item}'")
return _RaisingModule(name)
@@ -0,0 +1,185 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2 NVFP4 wiring contract tests.
These tests lock in the public-facing FP4 contract so the wire-up
between :class:`NVFP4Config`, the LTX-2 DiT (linear-class swap +
``prefix=`` plumbing), and the loader (``convert_model_to_nvfp4``
trigger via registered ``quant_method``) doesn't silently regress.
We don't need flashinfer or a GPU to assert the wiring: the layer's
``quant_method`` attachment, the prefix used by
``FP4Config.get_quant_method``, and the ``UnquantizedLinearMethod``
fallback for non-tagged layers are all CPU-only contracts.
"""
from __future__ import annotations
from fastvideo.layers.linear import ReplicatedLinear, UnquantizedLinearMethod
from fastvideo.layers.quantization.nvfp4_config import (
NVFP4Config,
NVFP4QuantizeMethod,
)
from fastvideo.models.dits.ltx2 import (
BasicAVTransformerBlock,
FeedForward,
LTXRopeType,
LTXSelfAttention,
TransformerConfig,
)
from fastvideo.platforms import AttentionBackendEnum
def _matched_attn() -> LTXSelfAttention:
"""Self-attention with a prefix in the NVFP4 layer set."""
return LTXSelfAttention(
query_dim=64,
context_dim=None,
heads=2,
dim_head=32,
norm_eps=1e-6,
rope_type=LTXRopeType.INTERLEAVED,
supported_attention_backends=(AttentionBackendEnum.TORCH_SDPA, ),
quant_config=NVFP4Config(),
prefix="ltx2.blocks.0.attn1",
)
def test_self_attention_to_q_k_v_out_get_nvfp4_method() -> None:
attn = _matched_attn()
for attr in ("to_q", "to_k", "to_v"):
linear = getattr(attn, attr)
assert isinstance(linear, ReplicatedLinear), (
f"{attr} must be ReplicatedLinear, got {type(linear).__name__}")
assert isinstance(linear.quant_method, NVFP4QuantizeMethod), (
f"{attr}.quant_method must be NVFP4QuantizeMethod, got "
f"{type(linear.quant_method).__name__}")
out_linear = attn.to_out[0]
assert isinstance(out_linear, ReplicatedLinear)
assert isinstance(out_linear.quant_method, NVFP4QuantizeMethod)
def test_self_attention_layer_prefixes_match_fp4_target_paths() -> None:
attn = _matched_attn()
assert attn.to_q.quant_method.layer_prefix == "ltx2.blocks.0.attn1.to_q"
assert attn.to_k.quant_method.layer_prefix == "ltx2.blocks.0.attn1.to_k"
assert attn.to_v.quant_method.layer_prefix == "ltx2.blocks.0.attn1.to_v"
assert attn.to_out[0].quant_method.layer_prefix == "ltx2.blocks.0.attn1.to_out"
def test_unmatched_prefix_falls_back_to_unquantized() -> None:
"""Layers whose prefix isn't in ``NVFP4Config.fp4_layers`` must
receive ``UnquantizedLinearMethod`` so the asserts in
``LinearBase`` subclasses don't fire."""
attn = LTXSelfAttention(
query_dim=64,
context_dim=128,
heads=2,
dim_head=32,
norm_eps=1e-6,
rope_type=LTXRopeType.INTERLEAVED,
supported_attention_backends=(AttentionBackendEnum.TORCH_SDPA, ),
quant_config=NVFP4Config(),
prefix="some.unrelated.prefix",
)
assert isinstance(attn.to_q.quant_method, UnquantizedLinearMethod)
def test_no_quant_config_keeps_unquantized_methods() -> None:
attn = LTXSelfAttention(
query_dim=64,
context_dim=None,
heads=2,
dim_head=32,
norm_eps=1e-6,
rope_type=LTXRopeType.INTERLEAVED,
supported_attention_backends=(AttentionBackendEnum.TORCH_SDPA, ),
quant_config=None,
prefix="ltx2.blocks.0.attn1",
)
assert isinstance(attn.to_q.quant_method, UnquantizedLinearMethod)
assert isinstance(attn.to_out[0].quant_method, UnquantizedLinearMethod)
def test_feedforward_fc_in_fc_out_get_nvfp4_method() -> None:
ff = FeedForward(
dim=64,
dim_out=64,
mult=2,
quant_config=NVFP4Config(),
prefix="ltx2.blocks.0",
)
fc_in = ff.net[0].proj
fc_out = ff.net[2]
assert isinstance(fc_in, ReplicatedLinear)
assert isinstance(fc_out, ReplicatedLinear)
assert isinstance(fc_in.quant_method, NVFP4QuantizeMethod)
assert isinstance(fc_out.quant_method, NVFP4QuantizeMethod)
assert fc_in.quant_method.layer_prefix == "ltx2.blocks.0.ffn.fc_in"
assert fc_out.quant_method.layer_prefix == "ltx2.blocks.0.ffn.fc_out"
def test_basic_av_block_propagates_quant_config_to_all_children() -> None:
"""``BasicAVTransformerBlock`` is the seam where ``quant_config``
branches into the four attention modules and FFN. Verify each
child sees the config and ends up with the right prefix.
"""
video_cfg = TransformerConfig(dim=64, heads=2, d_head=32, context_dim=128)
audio_cfg = TransformerConfig(dim=64, heads=2, d_head=32, context_dim=128)
block = BasicAVTransformerBlock(
idx=3,
video=video_cfg,
audio=audio_cfg,
rope_type=LTXRopeType.INTERLEAVED,
norm_eps=1e-6,
use_distributed_attention=False,
quant_config=NVFP4Config(),
prefix="ltx2",
)
# NVFP4 layer set is deliberately asymmetric — only the modules
# whose runtime cost dominates the DiT step are quantized:
# * attn1 (video self-attn): to_q/to_k/to_v/to_out
# * attn2 (text cross-attn for video): to_q + to_out only
# (text context isn't quantized)
# * audio_to_video_attn: to_q + to_out only
# * video_to_audio_attn: to_k + to_v only
# * ff (video FFN): fc_in + fc_out
# Audio self/cross attention and audio FFN are NOT quantized in
# the LTX-2 NVFP4 set (the audio path is much cheaper than video).
nvfp4_expectations = {
block.attn1.to_q: "ltx2.blocks.3.attn1.to_q",
block.attn1.to_out[0]: "ltx2.blocks.3.attn1.to_out",
block.attn2.to_q: "ltx2.blocks.3.attn2.to_q",
block.attn2.to_out[0]: "ltx2.blocks.3.attn2.to_out",
block.audio_to_video_attn.to_q:
"ltx2.blocks.3.audio_to_video_attn.to_q",
block.video_to_audio_attn.to_v:
"ltx2.blocks.3.video_to_audio_attn.to_v",
block.ff.net[0].proj: "ltx2.blocks.3.ffn.fc_in",
block.ff.net[2]: "ltx2.blocks.3.ffn.fc_out",
}
for linear, expected_prefix in nvfp4_expectations.items():
assert isinstance(linear, ReplicatedLinear)
assert isinstance(linear.quant_method, NVFP4QuantizeMethod), (
f"{expected_prefix} expected NVFP4QuantizeMethod, got "
f"{type(linear.quant_method).__name__}")
assert linear.quant_method.layer_prefix == expected_prefix
# Confirm the non-quantized projections still get the
# UnquantizedLinearMethod fallback (they're ReplicatedLinear, just
# not in the NVFP4 set).
unquantized_projections = (
block.attn2.to_k,
block.attn2.to_v,
block.audio_attn1.to_q,
block.audio_attn1.to_k,
block.audio_attn1.to_v,
block.audio_attn1.to_out[0],
block.audio_attn2.to_q,
block.audio_attn2.to_k,
block.audio_attn2.to_v,
block.audio_attn2.to_out[0],
block.audio_ff.net[0].proj,
block.audio_ff.net[2],
)
for linear in unquantized_projections:
assert isinstance(linear, ReplicatedLinear)
assert isinstance(linear.quant_method, UnquantizedLinearMethod)
+5
View File
@@ -161,6 +161,11 @@ nav:
- Design:
- Overview: design/overview.md
- Training Architecture: design/training_architecture.md
- Server Contracts:
- Overview: design/server_contracts/index.md
- OpenAI HTTP: design/server_contracts/openai.md
- Streaming WebSocket: design/server_contracts/streaming.md
- Dynamo Integration: design/server_contracts/dynamo.md
- Developer Guide:
- Overview: contributing/overview.md
- Developer Environment:
+14
View File
@@ -134,6 +134,20 @@ test = [
dev = [ "fastvideo[lint]", "fastvideo[test]", ]
prompt-safety = [
"fasttext",
]
prompt-enhancer = [
"httpx",
]
streaming = [
"fastvideo[prompt-enhancer]",
"fastvideo[prompt-safety]",
"websockets",
]
rocm = [
"amdsmi",
]
+405
View File
@@ -0,0 +1,405 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2 SR (refine) numerical alignment harness — public vs internal.
Runs the LTX-2 stage-2 spatial-refine pipeline with a fixed prompt and
seed and writes the decoded video frames + intermediate latents to a
torch save file. Intended to be invoked twice with different
``PYTHONPATH`` prefixes — once against ``FastVideo-internal`` and once
against the public ``FastVideo`` package — so a third pass can diff the
two outputs tensor-by-tensor.
Usage::
# 1) Internal reference run.
PYTHONPATH=/home/william5lin/FastVideo-internal \
python scripts/ltx2_sr_alignment.py --label internal \
--output /tmp/ltx2_sr_internal.pt
# 2) Public run against the new typed API.
PYTHONPATH=/home/william5lin/FastVideo \
python scripts/ltx2_sr_alignment.py --label public --use-typed-api \
--output /tmp/ltx2_sr_public.pt
# 3) Diff.
python scripts/ltx2_sr_alignment.py --diff \
--reference /tmp/ltx2_sr_internal.pt \
--candidate /tmp/ltx2_sr_public.pt
The harness deliberately matches `examples/inference/basic/basic_ltx2_upscale.py`
defaults (LTX2-Distilled, 8 base steps, 3 refine steps, 1088x1920, seed=10)
but skips FP4 and Dreamverse-specific knobs to keep the bf16 path clean.
"""
from __future__ import annotations
import argparse
import os
import sys
from dataclasses import dataclass
from pathlib import Path
from typing import Any
def _bootstrap_fastvideo_source() -> None:
"""If FASTVIDEO_FROM is set in the environment, prepend that path to
sys.path **and** strip any editable-install finders from
``sys.meta_path`` so the named source wins over a .pth-installed
sibling. This lets us point the same harness at either
``../FastVideo`` or ``../FastVideo-internal`` without juggling
venvs."""
src = os.environ.get("FASTVIDEO_FROM")
if not src:
return
src = os.path.abspath(src)
if not os.path.isdir(src):
raise SystemExit(f"FASTVIDEO_FROM={src!r} is not a directory")
# Remove editable-install finders that would otherwise win over
# PYTHONPATH. uv-installed editables register a custom finder
# named like ``__editable___fastvideo_0_1_7_finder``.
sys.meta_path = [
finder for finder in sys.meta_path
if "__editable___fastvideo" not in type(finder).__module__
]
# Drop pre-resolved fastvideo modules from any earlier import.
for name in [k for k in sys.modules if k == "fastvideo" or k.startswith("fastvideo.")]:
sys.modules.pop(name, None)
# Drop the corresponding .pth-pointed paths from sys.path so the
# editable repo doesn't shadow the explicit source.
sys.path = [p for p in sys.path if not p.endswith(".egg-info")]
sys.path.insert(0, src)
_bootstrap_fastvideo_source()
import numpy as np # noqa: E402 (after bootstrap so torch picks up the right env)
import torch # noqa: E402
# Pinned alignment fixture — matches basic_ltx2_upscale.py upstream.
PROMPT = (
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
"of a woman and a man in their 30s, facing each other with serious "
"expressions. The woman, emotional and dramatic, says softly, \"That's "
"it... Dad's lost it. And we've lost Dad.\" The man exhales, slightly "
"annoyed: \"Stop being so dramatic, Jess.\" A beat. He glances aside, "
"then mutters defensively, \"He's just having fun.\" The camera slowly "
"pans right, revealing the grandfather in the garden wearing enormous "
"butterfly wings, waving his arms in the air like he's trying to take "
"off. He shouts, \"Wheeeew!\" as he flaps his wings with full commitment. "
"The woman covers her face, on the verge of tears. The tone is deadpan, "
"absurd, and quietly tragic.")
DEFAULT_MODEL = "FastVideo/LTX2-Distilled-Diffusers"
DEFAULT_SEED = 10
DEFAULT_HEIGHT = 1088
DEFAULT_WIDTH = 1920
DEFAULT_FRAMES = 121
DEFAULT_FPS = 24
DEFAULT_BASE_STEPS = 8
DEFAULT_REFINE_STEPS = 3
@dataclass
class RunResult:
label: str
fastvideo_repo: str
output_frames: torch.Tensor
metadata: dict[str, Any]
def _detect_fastvideo_repo() -> str:
import fastvideo # noqa: PLC0415
return os.path.realpath(os.path.dirname(fastvideo.__file__))
def _run_legacy(args: argparse.Namespace) -> RunResult:
"""Legacy ``VideoGenerator.from_pretrained(**flat_kwargs)`` path —
used by the internal reference run because the internal API is
still flat-kwarg-shaped."""
from fastvideo import VideoGenerator # noqa: PLC0415
generator = VideoGenerator.from_pretrained(
args.model,
num_gpus=1,
ltx2_refine_enabled=True,
ltx2_refine_lora_path="", # disable refine LoRA — distilled needs none
ltx2_refine_num_inference_steps=args.refine_steps,
ltx2_refine_guidance_scale=1.0,
ltx2_refine_add_noise=True,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
pin_cpu_memory=True,
dit_layerwise_offload=False,
enable_torch_compile=False,
ltx2_vae_tiling=False,
)
# Pin the LTX-2-specific sampling knobs explicitly so both
# internal and public runs use identical denoising mechanics —
# otherwise the public LTX2_BASE preset's modality/rescale/stg
# defaults (3.0 / 0.7 / 1.0) would diverge from internal's
# ForwardBatch defaults (1.0 / 0.0 / 0.0).
result = generator.generate_video(
prompt=PROMPT,
output_path=str(args.output) + ".legacy.mp4",
fps=args.fps,
seed=args.seed,
num_inference_steps=args.base_steps,
guidance_scale=1.0,
save_video=False,
return_frames=True,
height=args.height,
width=args.width,
num_frames=args.frames,
ltx2_cfg_scale_video=1.0,
ltx2_cfg_scale_audio=1.0,
ltx2_modality_scale_video=1.0,
ltx2_modality_scale_audio=1.0,
ltx2_rescale_scale=0.0,
ltx2_stg_scale_video=0.0,
ltx2_stg_scale_audio=0.0,
)
generator.shutdown()
frames = _result_frames(result)
return RunResult(
label=args.label,
fastvideo_repo=_detect_fastvideo_repo(),
output_frames=frames,
metadata={
"api": "legacy_from_pretrained",
"prompt": PROMPT,
"seed": args.seed,
"height": args.height,
"width": args.width,
"num_frames": args.frames,
"fps": args.fps,
"base_steps": args.base_steps,
"refine_steps": args.refine_steps,
},
)
def _run_typed(args: argparse.Namespace) -> RunResult:
"""New typed public API: GeneratorConfig + GenerationRequest with
pipeline.preset_overrides["refine"] driving the SR path."""
from fastvideo import VideoGenerator # noqa: PLC0415
from fastvideo.api import ( # noqa: PLC0415
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig,
InputConfig, OutputConfig, PipelineSelection, SamplingConfig)
config = GeneratorConfig(
model_path=args.model,
engine=EngineConfig(num_gpus=1, ),
pipeline=PipelineSelection(
components=ComponentConfig(),
preset_overrides={
"refine": {
"enabled": True,
"add_noise": True,
"num_inference_steps": args.refine_steps,
"guidance_scale": 1.0,
},
},
),
)
generator = VideoGenerator.from_pretrained(config=config)
# The typed SamplingConfig doesn't expose the LTX-2-specific
# modality/rescale/stg knobs, so route them through experimental
# so the legacy pipeline batch builder picks them up. This keeps
# the typed run's denoising mechanics identical to the legacy
# run's (and therefore to the internal reference).
request = GenerationRequest(
prompt=PROMPT,
sampling=SamplingConfig(
num_frames=args.frames,
height=args.height,
width=args.width,
fps=args.fps,
num_inference_steps=args.base_steps,
guidance_scale=1.0,
seed=args.seed,
),
inputs=InputConfig(),
output=OutputConfig(
save_video=False,
return_frames=True,
),
)
# Set the LTX-2 sampling overrides on the request via attribute so
# the legacy SamplingParam translation picks them up alongside the
# typed fields. (Until the typed SamplingConfig grows these
# fields, this is the canonical override path on the public side.)
for attr, value in {
"ltx2_cfg_scale_video": 1.0,
"ltx2_cfg_scale_audio": 1.0,
"ltx2_modality_scale_video": 1.0,
"ltx2_modality_scale_audio": 1.0,
"ltx2_rescale_scale": 0.0,
"ltx2_stg_scale_video": 0.0,
"ltx2_stg_scale_audio": 0.0,
}.items():
setattr(request, attr, value)
result = generator.generate(request)
generator.shutdown()
frames = _result_frames(result)
return RunResult(
label=args.label,
fastvideo_repo=_detect_fastvideo_repo(),
output_frames=frames,
metadata={
"api": "typed",
"prompt": PROMPT,
"seed": args.seed,
"height": args.height,
"width": args.width,
"num_frames": args.frames,
"fps": args.fps,
"base_steps": args.base_steps,
"refine_steps": args.refine_steps,
},
)
def _result_frames(result: Any) -> torch.Tensor:
"""Coerce whatever generate() returned into ``[F, H, W, C]`` uint8."""
frames = getattr(result, "frames", None)
if frames is None and isinstance(result, dict):
frames = result.get("frames")
if frames is None:
raise RuntimeError(
"Generation returned no frames; ensure return_frames=True.")
if isinstance(frames, list):
frames = np.stack(frames, axis=0)
if isinstance(frames, np.ndarray):
return torch.from_numpy(frames)
if torch.is_tensor(frames):
return frames.detach().cpu()
raise TypeError(f"Unsupported frames container: {type(frames)!r}")
def _run(args: argparse.Namespace) -> None:
if args.use_typed_api:
result = _run_typed(args)
else:
result = _run_legacy(args)
output_path = Path(args.output)
output_path.parent.mkdir(parents=True, exist_ok=True)
torch.save(
{
"frames": result.output_frames,
"metadata": result.metadata,
"label": result.label,
"fastvideo_repo": result.fastvideo_repo,
},
output_path,
)
print(
f"[align] saved {result.output_frames.shape} frames "
f"({result.fastvideo_repo}) -> {output_path}",
flush=True,
)
def _diff(args: argparse.Namespace) -> None:
ref_path = Path(args.reference)
cand_path = Path(args.candidate)
if not ref_path.is_file() or not cand_path.is_file():
raise FileNotFoundError(
f"Need both --reference and --candidate to exist: {ref_path}, "
f"{cand_path}")
ref = torch.load(ref_path, map_location="cpu")
cand = torch.load(cand_path, map_location="cpu")
ref_frames = ref["frames"]
cand_frames = cand["frames"]
print(f"[align] reference: {ref['label']} {ref_frames.shape} from "
f"{ref['fastvideo_repo']}")
print(f"[align] candidate: {cand['label']} {cand_frames.shape} from "
f"{cand['fastvideo_repo']}")
if ref_frames.shape != cand_frames.shape:
print(f"[align] shape MISMATCH: ref={tuple(ref_frames.shape)} "
f"vs cand={tuple(cand_frames.shape)}")
return
diff = (ref_frames.float() - cand_frames.float()).abs()
max_abs = float(diff.max().item())
mean_abs = float(diff.mean().item())
rms = float(torch.sqrt((diff**2).mean()).item())
# PSNR over uint8 [0,255] data range.
if rms == 0.0:
psnr = float("inf")
else:
psnr = 20.0 * float(torch.log10(torch.tensor(255.0 / rms)).item())
print("[align] frame-level uint8 metrics (per-pixel):")
print(f" max_abs_diff = {max_abs:.4f}")
print(f" mean_abs_diff = {mean_abs:.6f}")
print(f" rms_diff = {rms:.4f}")
print(f" psnr_db = {psnr:.2f}")
# Bucket pixels by exactness so we have a sense of how close we are.
eq = (diff == 0).float().mean().item()
within_1 = (diff <= 1).float().mean().item()
within_2 = (diff <= 2).float().mean().item()
within_4 = (diff <= 4).float().mean().item()
print("[align] pixel buckets:")
print(f" exact = {eq * 100:.2f}%")
print(f" within 1 = {within_1 * 100:.2f}%")
print(f" within 2 = {within_2 * 100:.2f}%")
print(f" within 4 = {within_4 * 100:.2f}%")
def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
sub = parser.add_subparsers(dest="mode", required=False)
run = parser
run.add_argument("--label",
default="run",
help="Identifier saved alongside the output (e.g. internal/public).")
run.add_argument("--output",
default="/tmp/ltx2_sr_alignment.pt",
help="Destination .pt file.")
run.add_argument("--use-typed-api",
action="store_true",
help="Use the new typed GeneratorConfig API "
"(public side); otherwise use legacy from_pretrained.")
run.add_argument("--model", default=DEFAULT_MODEL)
run.add_argument("--seed", type=int, default=DEFAULT_SEED)
run.add_argument("--height", type=int, default=DEFAULT_HEIGHT)
run.add_argument("--width", type=int, default=DEFAULT_WIDTH)
run.add_argument("--frames", type=int, default=DEFAULT_FRAMES)
run.add_argument("--fps", type=int, default=DEFAULT_FPS)
run.add_argument("--base-steps", type=int, default=DEFAULT_BASE_STEPS)
run.add_argument("--refine-steps", type=int, default=DEFAULT_REFINE_STEPS)
run.add_argument("--diff",
action="store_true",
help="Diff two .pt outputs instead of running.")
run.add_argument("--reference")
run.add_argument("--candidate")
return parser
def main(argv: list[str] | None = None) -> int:
parser = _build_parser()
args = parser.parse_args(argv)
if args.diff:
if not (args.reference and args.candidate):
parser.error("--diff requires --reference and --candidate")
_diff(args)
return 0
_run(args)
return 0
if __name__ == "__main__":
sys.exit(main())