Compare commits

...
Author SHA1 Message Date
cbb9e85253 [feat] Activate Dreamverse LTX2 integration
Apply Dreamverse monorepo changes for stack slice 13/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 13/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
68e2841fa2 [feat] Add LTX2 refine and upsampler support
Apply Dreamverse monorepo changes for stack slice 12/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 12/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
a110a24a77 [infra] Add NVFP4 quantization support
Apply Dreamverse monorepo changes for stack slice 11/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 11/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
0d30d644af [feat] Add FastVideo serving API contracts
Apply Dreamverse monorepo changes for stack slice 10/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 10/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
42ab4d70a7 [infra] Add Dreamverse Docker and launch scripts
Apply Dreamverse monorepo changes for stack slice 9/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 9/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
2eedfa0449 [feat] Add Dreamverse frontend media and E2E coverage
Apply Dreamverse monorepo changes for stack slice 8/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 8/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
bdad051a9a [feat] Add Dreamverse frontend session UI
Apply Dreamverse monorepo changes for stack slice 7/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 7/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:33:21 -07:00
Davids048 c1bd22434e [fix] Remove stale server-assets rewrite 2026-05-12 13:33:06 -07:00
1e371dd3a5 [feat] Add Dreamverse frontend scaffold
Apply Dreamverse monorepo changes for stack slice 6/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 6/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:32:38 -07:00
38ab5af8d5 [feat] Add Dreamverse streaming runtime
Apply Dreamverse monorepo changes for stack slice 5/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 5/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:32:38 -07:00
9a1f3b4897 [feat] Add Dreamverse session and prompt logic
Apply Dreamverse monorepo changes for stack slice 4/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 4/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:32:38 -07:00
Davids048 1fe1c85558 [fix] Stop exposing Dreamverse backend source as static assets 2026-05-12 13:32:38 -07:00
ffefdc045a [feat] Add Dreamverse backend skeleton
Apply Dreamverse monorepo changes for stack slice 3/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 3/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:32:38 -07:00
Davids048 522b220dab [docs] Remove stale server-assets docs 2026-05-12 13:32:10 -07:00
Davids048 2848917bab [fix] Address review feedback for Dreamverse 02/14 2026-05-12 12:45:09 -07:00
2a8a7a01a9 [docs] Add Dreamverse app documentation
Apply Dreamverse monorepo changes for stack slice 2/13 from the source branch.

Source-Branch: will/dreamverse-monorepo
Source-SHA: 03d3e61df6
Dreamverse-Stack: 2/13
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 12:45:09 -07:00
221 changed files with 60874 additions and 257 deletions
+216
View File
@@ -0,0 +1,216 @@
# Dreamverse Agent Notes
## Repo-local skills
- `.agents/skills/bootstrap-fastvideo-private-fork/`: temporary setup skill for
cloning `git@github.com:hao-ai-lab/FastVideo-internal.git` at
`will/rebase-nbv` into `../FastVideo-internal`, then running
`uv sync --extra server`.
- Prefer the bundled script in that skill instead of inventing a new private
FastVideo bootstrap flow.
## Repo layout
Current paths:
- `apps/dreamverse/web/`: Next.js frontend, client-side stores, websocket
event reduction, prompt-window editing, devtools UI.
- `apps/dreamverse/dreamverse/`: current Python FastAPI runtime, websocket
protocol, prompt enhancement, prompt rewrite orchestration, GPU worker
lifecycle.
- `apps/dreamverse/dreamverse/tests/`: backend unit and integration-oriented
tests.
- `apps/dreamverse/dreamverse/benchmarks/`: prompt-provider latency/token
benchmarking scripts.
Planned paths during the OSS reorg:
- `controller/`: local control plane for provider credentials, compute
lifecycle, and proxying.
- `runtime/`: eventual rename of `apps/dreamverse/dreamverse/` once the
controller/runtime split is stable.
- `providers/`: provider adapters for local, Runpod, and Modal.
Important rule:
- Until the split lands, treat `apps/dreamverse/dreamverse/` as the
authoritative runtime and keep provider orchestration out of
`apps/dreamverse/web`.
## System split
Dreamverse is moving toward a three-part local-first architecture.
- Frontend owns local UI state, drafts, inspection tools, and user-triggered
actions.
- Controller will own local-only credentials, compute provisioning, runtime
lifecycle, and HTTP/websocket proxying.
- Runtime owns generation sessions, prompt rewrite, prompt safety, and
websocket semantics.
The browser should only talk to the local Dreamverse process, never directly to
Modal or Runpod.
## Current runtime responsibilities
The runtime in `apps/dreamverse/dreamverse/` is responsible for:
- websocket session lifecycle on `/ws`
- queueing, GPU assignment, worker startup, and stream chunk emission
- seed prompt memory and the active prompt window used for generation
- prompt enhancement and prompt rewrite execution
- prompt safety checks
- persistence and reload of prompt system prompt files
- curated preset append/read routes in devtools mode
- health and readiness endpoints
Relevant files:
- `apps/dreamverse/dreamverse/main.py`: websocket protocol, session state
machine, REST routes
- `apps/dreamverse/dreamverse/gpu_pool.py`: FastVideo-backed generation
workers
- `apps/dreamverse/dreamverse/prompt_enhancer.py`: provider clients, prompt
enhancement, rewrite execution
- `apps/dreamverse/dreamverse/rewrite_prompt_payload.py`: canonical rewrite
request body format
- `apps/dreamverse/dreamverse/config.py`: prompt file paths, provider
configuration, runtime flags
## Frontend responsibilities
The frontend in `apps/dreamverse/web/` is responsible for:
- collecting user input and deciding whether to send raw prompts or rewrite
requests
- maintaining client-side stores for session, prompt-window, stream, rewrite,
and UI state
- rendering prompt history, playback state, devtools controls, and rewrite
inspection
- building the prompt-window snapshot sent with rewrite requests
- reducing websocket events into UI state
- showing compute status and controller-driven errors once the controller lands
Relevant files:
- `apps/dreamverse/web/src/app/page.tsx`: main orchestration, websocket
connect/send paths
- `apps/dreamverse/web/src/lib/ws/reducer.ts`: applies normalized websocket
events to stores
- `apps/dreamverse/web/src/stores/promptWindow.ts`: prompt window and
preset/editor state
- `apps/dreamverse/web/src/stores/rewrite.ts`: rewrite activity timeline and
flags
- `apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts`: rewrite snapshot
normalization and padding
## Planned controller responsibilities
The future local controller should own:
- local-only provider credential loading and storage
- provider selection
- runtime provisioning, reuse, shutdown, and health checks
- proxying frontend HTTP and websocket traffic to the active runtime
- surfacing provisioning, ready, failed, and idle states to the frontend
- durable local settings that should survive ephemeral remote runtimes
The controller should not own:
- prompt rewrite logic
- seed prompt memory
- generation queue semantics
- websocket event schemas
## Prompt rewrite contract
Prompt rewrite is a shared flow with a strict ownership split.
Frontend responsibilities:
- decide when a user action should trigger `rewrite_seed_prompts` instead of
`append_prompt`
- send `rewrite_instruction` and a snapshot of the current prompt window
- pad the rewrite snapshot to the runtime-expected segment count using
`buildRewritePromptWindowSnapshotFromPrompts(...)`
- show rewrite activity and raw LLM output in local inspection UI
Runtime responsibilities:
- validate and normalize `prompt_window_prompts`
- choose the rewrite system prompt and provider/model/temperature
- build the canonical LLM request body in
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`
- run the rewrite through `PromptEnhancer.rewrite_prompt_sequence(...)`
- apply safety filtering to rewritten prompts
- replace the authoritative seed prompt memory when rewrite succeeds
- emit `seed_prompts_updated` and `rewrite_seed_prompts_complete`
Controller responsibilities:
- proxy the request and response
- surface runtime availability and provider lifecycle failures
Important rule:
- The frontend may suggest the prompt window to rewrite, but the runtime owns
the actual rewritten rollout and the authoritative prompt window after
acceptance.
## Rewrite modes
There are two runtime rewrite modes:
- edit existing rollout: when `prompt_window_prompts` is non-empty, rewrite the
current rollout while preserving segment count and ordering
- new rollout: when the prompt window is empty but there is a
`rewrite_instruction`, generate a fresh rollout
The frontend should not emulate runtime rewrite behavior locally. It should
prepare the snapshot, send it, and display the result.
## Prompt window ownership
- Frontend owns editable drafts, selected preset UI, and prompt-window
inspection state.
- Runtime owns the active seed prompt memory used for actual generation.
- After any runtime event with reason `rewrite`, the frontend must replace its
prompt window from the server payload instead of keeping a locally-derived
version.
## Devtools and persistence ownership
Prompt config editing is runtime-owned persistence with frontend-owned forms
today.
- Frontend loads and edits drafts through `/prompt-system-config`.
- Runtime reads and writes prompt files and reloads runtime prompt config.
Curated presets follow the same pattern:
- frontend submits append requests and may update local UI optimistically from
the response
- runtime persists the JSON file and resolves overlay vs fallback file paths
During the controller reorg, avoid moving durable user settings into ephemeral
remote runtimes. Controller-owned local persistence is preferred for anything
that must survive provider restarts.
## Editing guidance
- Do not move rewrite logic into the frontend or controller.
- Do not make the frontend the source of truth for the generated prompt window
after rewrite.
- If you change websocket message types or payload fields in
`apps/dreamverse/dreamverse/main.py`, update the reducer in
`apps/dreamverse/web/src/lib/ws/reducer.ts` in the same change.
- If you change rewrite request shape, update both
`apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts` and
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`.
- If you add controller-managed status or error payloads, keep them separate
from runtime websocket events unless there is a strong reason to merge them.
- If you change prompt file paths or devtools persistence, update
`apps/dreamverse/dreamverse/`, the frontend devtools UI, and any
controller-owned local persistence logic together.
- Keep provider adapters focused on runtime lifecycle and reachability, not on
prompt or session semantics.
+234
View File
@@ -0,0 +1,234 @@
# Dreamverse
Dreamverse is the FastVideo realtime video generation & editing platform. It lives in this monorepo under `apps/dreamverse/`.
## Install Dreamverse
You can install Dreamverse using one of the methods below.
### Method 1: With uv pip
```bash
pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
uv pip install "fastvideo[dreamverse]"
```
### Method 2: From source
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
uv pip install -e ".[dreamverse]"
```
### Method 3: Using Docker
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
apps/dreamverse/docker/docker_build.sh
```
See `apps/dreamverse/docker/README.md` for Docker build and run option details.
## Optional: Building FFmpeg For Better Performance
For full streaming performance in a non-Docker install, build a custom FFmpeg
binary:
```bash
bash apps/dreamverse/scripts/install_native_ffmpeg.sh
```
This builds and installs into `~/opt/ffmpeg-native/` and writes
`apps/dreamverse/scripts/ffmpeg-env.sh`. Source it before starting the backend
so Dreamverse uses the custom FFmpeg binary:
```bash
source apps/dreamverse/scripts/ffmpeg-env.sh
dreamverse-server
```
Docker images already run this FFmpeg build during image creation and source the
generated environment file at container startup. The installer supports Linux
`x86_64` and `aarch64`.
## Launch Dreamverse
Start the backend with the installed Dreamverse commands:
```bash
dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
## Frontend Setup
Install the web dependencies once from the FastVideo checkout:
```bash
cd apps/dreamverse/web
pnpm install --frozen-lockfile
```
The frontend package also has an npm lockfile, but the bundled launch scripts
use `pnpm`.
## Quick Start: Local GPU
### Start Backend
Export the API keys used for prompt rewrite and prompt enhancement:
```bash
export CEREBRAS_API_KEY=...
export GROQ_API_KEY=...
```
If you built the optional native FFmpeg binary above, source its environment
file in the same shell before starting the backend:
```bash
source apps/dreamverse/scripts/ffmpeg-env.sh
dreamverse-server --host 0.0.0.0 --port 8009
```
The Dreamverse backend defaults to `0.0.0.0:8009` and starts one GPU worker on
the first visible GPU by default.
### Check Readiness
In another shell, verify that the backend process is alive:
```bash
curl http://localhost:8009/healthz
```
Then wait for GPU workers and startup warmup to finish:
```bash
curl http://localhost:8009/readyz
```
You can also run the same readiness path with:
```bash
BACKEND_HOST=localhost BACKEND_PORT=8009 apps/dreamverse/scripts/smoke_local.sh
```
If a backend is already running and you only want the script to probe it:
```bash
DREAMVERSE_SMOKE_START_BACKEND=0 apps/dreamverse/scripts/smoke_local.sh
```
### Start Frontend
Start the frontend:
```bash
cd apps/dreamverse/web
BACKEND_HOST=localhost BACKEND_PORT=8009 pnpm run dev
```
Open `http://localhost:5299`.
## Quick Start: Mock Backend (For UI development)
The mock server emulates the Dreamverse backend protocol and streams a
synthetic FFmpeg-generated fMP4 clip, so the frontend can run without a GPU.
```bash
dreamverse-mock-server --latency 200 --port 8009
```
## Tests
Run the focused backend tests that validate local startup wiring, config, GPU
selection, and mock-server behavior:
```bash
pytest apps/dreamverse/dreamverse/tests/test_config.py \
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
```
Run the broader Dreamverse backend suite:
```bash
pytest apps/dreamverse/dreamverse/tests -q
```
Run the frontend tests:
```bash
cd apps/dreamverse/web
pnpm test
```
Run the frontend e2e tests:
```bash
cd apps/dreamverse/web
pnpm run e2e
```
## Troubleshooting
`dreamverse-server` exits with an install hint
- install the Dreamverse extra with `uv pip install -e ".[dreamverse]"` from a
source checkout, or `uv pip install "fastvideo[dreamverse]"` from PyPI.
Prompt-provider environment variable errors
- set `CEREBRAS_API_KEY`
- set `GROQ_API_KEY`
- direct `dreamverse-server` launches do not source `~/.env`; export the keys
in the shell or use the bundled launch scripts, which source `~/.env`.
`/readyz` stays at `503`
- wait for model loading and startup warmup to finish
- confirm a compatible CUDA GPU is visible to the process
- check backend logs for worker startup or warmup failures
- for a startup/debug pass without warmup, set
`FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend
Only one GPU is used
- this is the default local behavior
- set `FASTVIDEO_GPU_COUNT=<N>` to start N GPU worker subprocesses inside one
backend instance
- set `FASTVIDEO_GPU_COUNT=all` to start one worker for every visible GPU
- use `CUDA_VISIBLE_DEVICES` first if you need to pin the visible GPU set
Frontend cannot connect to backend
- confirm the backend is running on `8009`; if not, point the frontend at it
with `BACKEND_HOST=<host> BACKEND_PORT=<port> pnpm run dev`
- confirm `http://localhost:8009/healthz` responds before starting the frontend
- confirm `http://localhost:8009/readyz` returns `200` before clicking Generate
- use `apps/dreamverse/scripts/smoke_local.sh` for a repeatable local startup
check
Mock backend fails during startup
- install FFmpeg or set `FASTVIDEO_FFMPEG_BIN` to an FFmpeg binary
- for non-mock local GPU streaming performance, use the native FFmpeg installer
described above
## Notes
Dreamverse owns its backend app under `apps/dreamverse/dreamverse/`. It expects
`dreamverse-server`, not `fastvideo serve`.
+354
View File
@@ -0,0 +1,354 @@
# Dreamverse Architecture
## Overview
Dreamverse currently has two main runtime pieces:
- `apps/dreamverse/web/`: Next.js frontend
- `apps/dreamverse/dreamverse/`: Python FastAPI runtime
Today, the browser talks directly to the Dreamverse runtime over HTTP and a
single websocket on `/ws`. The frontend owns UI state and interaction flow. The
server owns generation state, prompt rewrite, prompt safety, websocket session
semantics, and GPU-backed execution.
Near-term OSS note:
- `apps/dreamverse/dreamverse/` is the current runtime implementation.
- A future `controller/` layer is planned for local-only compute management and
provider orchestration, but it does not exist yet.
## Repo Map
### Frontend
- `apps/dreamverse/web/src/app/page.tsx`: main client orchestration,
websocket connect, init payloads, send paths, and top-level app behavior
- `apps/dreamverse/web/src/lib/ws/reducer.ts`: reduces normalized websocket
events into client stores
- `apps/dreamverse/web/src/stores/session.ts`: connection, mode, and top-level
session UI
- `apps/dreamverse/web/src/stores/promptWindow.ts`: editable prompt window and
seed prompt UI state
- `apps/dreamverse/web/src/stores/rewrite.ts`: rewrite activity timeline and
inspection state
- `apps/dreamverse/web/src/stores/stream.ts`: playback and stream-related
client state
- `apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts`: prompt-window snapshot
building for rewrite requests
### Server
- `apps/dreamverse/dreamverse/main.py`: websocket endpoint, request handling,
session state machine, rewrite orchestration, REST routes, and stream relay
- `apps/dreamverse/dreamverse/gpu_pool.py`: GPU worker processes, warmup, model
loading, and `generate_video()` calls through FastVideo
- `apps/dreamverse/dreamverse/prompt_enhancer.py`: prompt enhancement, rollout
rewrite execution, provider selection, and timeout/fallback behavior
- `apps/dreamverse/dreamverse/rewrite_prompt_payload.py`: canonical rewrite request payload
building
- `apps/dreamverse/dreamverse/config.py`: runtime flags, prompt file paths, provider settings, and
warmup config
- `apps/dreamverse/dreamverse/session_init_image.py`: validates and persists uploaded initial
images for segment 1
## Current Split Of Responsibility
### Frontend owns
- local UI state and client-side stores
- prompt drafts and prompt window editing
- websocket connection management
- deciding which user action to send:
- `session_init_v2`
- `project_init_v1`
- `append_prompt`
- `rewrite_seed_prompts`
- `simple_generate`
- showing rewrite progress, stream status, prompt history, and devtools views
### Server owns
- websocket session lifecycle and protocol
- GPU assignment and worker lifecycle
- the authoritative seed prompt memory used for generation
- prompt rewrite execution and prompt safety
- actual generation queue semantics
- stream chunk emission and segment lifecycle events
- prompt config and preset persistence routes
- health and readiness endpoints
Important rule:
- The frontend may propose prompt-window state for rewrite, but the server is
the source of truth for the rewritten rollout and the active prompt memory
used for generation.
## End-To-End Flow
1. The frontend opens `/ws`.
2. The frontend sends `session_init_v2` with the initial prompt-window state,
preset metadata, and current toggles.
3. The server validates init data, persists an optional initial image, acquires
a GPU slot, and emits session status such as `gpu_assigned`.
4. The server starts or resumes project generation and emits events like
`ltx2_stream_start`, `ltx2_segment_start`, media init/chunks, and completion
events.
5. The frontend reduces those websocket events into its stores and updates the
UI.
6. User actions such as appending prompts, rewriting seed prompts, or starting
a single custom clip go back to the server over the same websocket.
## Frontend Architecture
The frontend is store-driven.
- `page.tsx` wires together websocket setup, send helpers, reducer
application, and top-level interaction flows.
- Store modules separate concerns like session state, rewrite state, prompt
window state, and stream state.
- The websocket reducer is responsible for turning normalized runtime events
into store updates. If the server event schema changes, the reducer must
change with it.
The frontend is intentionally not responsible for:
- generating rewritten prompts locally
- deciding final prompt safety outcomes
- reconstructing server session state from scratch
- inventing its own generation semantics independent of the runtime
## Server Architecture
The current runtime is a FastAPI app with a single long-lived websocket per
session.
`apps/dreamverse/dreamverse/main.py` manages:
- websocket connect/init
- prompt queues
- project and segment state
- prompt enhancement/rewrite triggers
- stream relay from GPU workers to the browser
- session logging and REST endpoints
`apps/dreamverse/dreamverse/gpu_pool.py` manages:
- model loading through FastVideo
- one or more worker processes
- startup warmup
- user join/leave commands
- `USER_STEP` execution for each segment
- continuation state between segments
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
- prompt enhancement for user-submitted prompts
- rollout rewrite requests for the prompt window
- provider selection and fallback across configured prompt providers
- response normalization and safety-aware failure handling
## FastAPI Surface
The server is a single FastAPI application created in
`apps/dreamverse/dreamverse/main.py`.
Current built-in FastAPI docs are enabled:
- `/docs`: Swagger UI
- `/redoc`: ReDoc
- `/openapi.json`: OpenAPI schema
The runtime also mounts the frontend static build at `/` when one of the
configured frontend static directories exists. It does not expose the backend
Python package as static content.
## HTTP API
The current HTTP API is small. Most realtime behavior still goes through the
websocket.
### Core health and status routes
- `GET /healthz`
- process liveness probe
- returns a small payload with `status`, `service`, and timestamp
- `GET /readyz`
- readiness probe
- returns `503` until prompt services are initialized and at least one GPU
worker is ready
- returns readiness and GPU pool summary fields such as ready workers, total
GPUs, warmup counts, and queue size
- `GET /status`
- returns the current GPU pool status payload from `gpu_pool`
- `GET /internal/monitor/sessions`
- internal monitoring payload for session dashboards
- includes pending session count, max available sessions, prompt provider
success counts, and timestamp
### Prompt config routes
- `GET /prompt-system-config`
- returns the editable prompt-system configuration currently loaded by
`PromptEnhancer`
- `POST /prompt-system-config`
- saves prompt-system configuration to disk and reloads prompt config in the
runtime
- current editable fields include:
- next-segment system prompt
- auto-extension system prompt
- rewrite-window system prompt
- rewrite-user system prompt
- rewrite model
- rewrite temperature
### Devtools-only preset routes
These exist only when `DEVTOOLS_ENABLED` is true in
`apps/dreamverse/dreamverse/config.py`.
- `GET /curated-presets`
- returns merged curated presets, applying the local overlay file on top of
the fallback file when both exist
- `POST /curated-presets/append`
- appends a new curated preset to the overlay presets file
- validates non-empty label, normalized id, and at least two non-empty
segment prompts
## Websocket API
`WS /ws` is the main runtime API.
The websocket owns:
- session init
- project init and reset
- prompt append
- prompt rewrite
- generation toggles
- segment lifecycle events
- media stream delivery
- runtime error delivery
The websocket is the authoritative API for realtime Dreamverse behavior. The
HTTP routes mainly support health checks, devtools persistence, and monitoring.
## Prompt Rewrite Architecture
Prompt rewrite is a shared flow with strict ownership boundaries.
### Frontend responsibilities
- collect the rewrite instruction
- build the prompt-window snapshot from current client state
- send `rewrite_seed_prompts`
- show rewrite progress, raw output, fallback state, and resulting prompt list
### Server responsibilities
- validate and normalize the prompt-window payload
- choose rewrite model, system prompt, timeout, and temperature
- build the canonical prompt payload in
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`
- execute rewrite through `PromptEnhancer`
- apply safety filtering to rewritten prompts
- replace the authoritative seed prompt memory when rewrite succeeds
- emit `seed_prompts_updated` and `rewrite_seed_prompts_complete`
Important rule:
- The frontend owns editable drafts.
- The server owns the accepted rollout.
After a successful rewrite, the frontend should replace its prompt-window view
from the server payload instead of preserving a locally-derived version.
## Prompt Modes
There are three related prompt paths in the current system:
### Initial rollout
- The frontend sends seed prompts during `session_init_v2`.
- The server uses those prompts as the initial seed prompt memory.
- If the rollout starts from an empty prompt window plus an initial rewrite
instruction, the server can pause generation until rewrite completes.
### Live append
- The frontend sends `append_prompt`.
- The server may enhance that prompt, safety-check it, enqueue it, and use it
as the next generated segment.
### Rewrite
- The frontend sends `rewrite_seed_prompts`.
- The server rewrites the entire seed prompt window or generates a new rollout,
depending on the payload and current state.
## Initial Image And Segment Handling
The frontend currently sends `initial_image` as part of session init or
`simple_generate`.
The server:
- validates and persists the image
- uses it only for segment 1 when present
- keeps continuation state for later segments in the GPU worker
This means the runtime, not the frontend, decides how segment 1 image
conditioning and later continuation conditioning are applied.
## Websocket Contract
The websocket is the main integration surface between UI and runtime.
Typical incoming messages from the frontend:
- `session_init_v2`
- `project_init_v1`
- `append_prompt`
- `rewrite_seed_prompts`
- `simple_generate`
- `set_enhancement`
- `set_auto_extension`
- `set_loop_generation`
Typical outgoing messages from the server:
- `gpu_assigned`
- `ltx2_stream_start`
- `ltx2_segment_start`
- `segment_prompt_source`
- `prompt_received`
- `prompt_ready`
- `prompt_enhancing`
- `seed_prompts_updated`
- `rewrite_seed_prompts_complete`
- `media_init`
- `media_segment_complete`
- `project_idle`
- `error`
Binary websocket frames carry media chunks for playback.
## Current And Planned Architecture
Current architecture:
- browser -> `apps/dreamverse/web`
- `apps/dreamverse/web` -> `apps/dreamverse/dreamverse/main.py`
- `apps/dreamverse/dreamverse/main.py` ->
`apps/dreamverse/dreamverse/gpu_pool.py`
- `gpu_pool.py` -> FastVideo runtime
Planned architecture:
- browser -> `apps/dreamverse/web`
- `apps/dreamverse/web` -> local `controller/`
- `controller/` -> local or remote Dreamverse runtime
- runtime -> FastVideo runtime
That future controller split should not move prompt rewrite, session state, or
generation semantics out of the runtime.
+600
View File
@@ -0,0 +1,600 @@
# Dreamverse OSS Design
## Overview
Dreamverse should ship as a local-first open source application.
- The browser talks only to a local Dreamverse control plane on the user's
machine.
- Provider credentials stay local to that machine.
- Dreamverse may provision compute on the user's behalf, but Dreamverse does
not host that control path as a service.
This keeps the UX simple without turning Dreamverse into a credential-holding
hosted platform.
## Goals
- Support three compute modes behind one product surface:
- local GPU
- managed remote GPU via Runpod
- managed remote GPU via Modal
- Keep prompt rewrite, websocket session state, and generation behavior
consistent across providers.
- Keep provider API keys out of browser state and out of any hosted service.
- Make the existing runtime reusable as the common serving contract.
- Minimize provider-specific code and isolate it behind a narrow interface.
## Non-goals
- Do not make the frontend call provider APIs directly.
- Do not unify providers at the level of SSH, VM, serverless, or pod
semantics.
- Do not move prompt rewrite logic into the frontend or controller.
- Do not require remote compute for the basic product path.
## Current State
Today the repo contains two major pieces:
- `apps/dreamverse/web/`: Next.js frontend
- `apps/dreamverse/dreamverse/`: FastAPI runtime that owns websocket state,
prompt rewrite, prompt safety, and GPU-backed generation
The current runtime already exposes useful health and streaming surfaces such
as `/healthz`, `/readyz`, `/status`, and `/ws`.
## Target Architecture
The target open source structure should be:
```text
Dreamverse/
├── apps/dreamverse/
│ ├── web/ # browser UI
│ ├── dreamverse/ # current FastAPI websocket/generation runtime
│ ├── controller/ # local control plane and provider lifecycle
│ ├── providers/ # provider adapters
│ ├── tests/
│ │ ├── contract/
│ │ ├── controller/
│ │ └── smoke/
│ └── design.md
└── ...
```
Near-term note:
- `apps/dreamverse/dreamverse/` is the current runtime implementation.
- We can keep the code there initially and rename it to `runtime/` only after
the controller lands.
## Trust Model
Dreamverse is local-only for control and secrets.
- The user launches Dreamverse on their own machine.
- Provider API keys are entered into the local app or local CLI.
- The controller uses those credentials to provision or connect to compute.
- The browser never talks to Modal or Runpod directly.
- Dreamverse-hosted infrastructure is not involved.
This is the key reason the provider-based path is acceptable for OSS.
## Responsibility Split
### `apps/dreamverse/web`
The frontend should own:
- UI state, drafts, and local interaction state
- websocket event reduction into client stores
- selection of compute mode and display of cost/health/status
- local forms for provider configuration
- sending prompt requests and rewrite requests to the local controller
The frontend should not own:
- provider credentials after submission
- provider API calls
- runtime lifecycle
- authoritative prompt window after rewrite
- prompt safety or generation policy
### `controller`
The local controller should own:
- provider credential loading and local-only storage
- compute mode selection
- provisioning, reuse, shutdown, and health monitoring of runtimes
- reverse proxying HTTP and websocket traffic from the frontend to the active
runtime
- user-visible status such as provisioning, ready, failed, and idle shutdown
- local persistence for user settings that must survive ephemeral runtimes
The controller should not own:
- prompt rewrite logic
- seed prompt memory semantics
- generation queue behavior
- provider-specific UI state
### `runtime`
The runtime should remain the authoritative owner of:
- `/ws` session state
- prompt rewrite execution
- prompt safety
- seed prompt memory and prompt-window state used for generation
- generation orchestration and GPU worker lifecycle
- websocket event schemas
This preserves the current model and avoids splitting state across layers.
## Runtime Contract
Provider abstraction should happen around a stable Dreamverse runtime contract,
not around infrastructure details.
Minimum runtime surface:
- `GET /healthz`
- `GET /readyz`
- `GET /status`
- `GET/POST /prompt-system-config` if devtools persists config through the
runtime
- curated preset routes if those remain runtime-backed
- `WS /ws`
Important rule:
- The controller only needs to know how to reach a healthy runtime.
- The runtime remains provider-agnostic.
## Provider Abstraction
Use a narrow provider interface:
```python
class ComputeProvider(Protocol):
async def ensure_runtime(self, spec: RuntimeSpec) -> RuntimeHandle: ...
async def wait_until_ready(self, handle: RuntimeHandle) -> None: ...
async def stop_runtime(self, handle: RuntimeHandle) -> None: ...
```
`RuntimeHandle` should include:
- `provider`
- `runtime_id`
- `base_url`
- `ws_url`
- runtime auth headers or tokens if needed
- lifecycle metadata
- cost or hardware metadata for UI display
The controller should work only with `RuntimeHandle`, never with raw SSH hosts
or provider-specific payloads after resolution.
## Provider Notes
### Local
Local mode should be the reference implementation.
- Start the runtime as a local subprocess or connect to an already-running
local runtime URL.
- Reuse the same runtime contract as remote providers.
- Make this the first supported path and the main smoke-test target.
### Runpod
Runpod should be treated as pod lifecycle plus runtime reachability.
- Prefer prepared images or templates that auto-start the Dreamverse runtime.
- Prefer exposed HTTP/TCP ports for steady-state traffic.
- Use SSH only for bootstrap fallback, diagnostics, or repair.
- Avoid a design where the controller shells into the pod for every action.
### Modal
Modal should be treated as deployment-based runtime hosting.
- Wrap the Dreamverse runtime in a thin Modal entrypoint if needed.
- Reuse the same runtime behavior behind that wrapper.
- Do not model Modal as a machine that Dreamverse logs into.
- Do not force the websocket runtime into a per-request serverless handler
shape.
## Config and Persistence
Remote compute may be ephemeral, so mutable user configuration should not live
only inside remote runtimes.
Keep durable state local to the user's machine unless there is a strong reason
otherwise:
- provider selection
- provider credentials or credential references
- default hardware preferences
- editable prompt presets
- prompt system prompt overrides
- idle shutdown policy
Runtime-local state should be treated as disposable unless explicitly synced.
## Prompt Rewrite Ownership
Prompt rewrite remains runtime-owned even after the controller is added.
Frontend responsibilities:
- collect the rewrite instruction
- build the prompt-window snapshot
- display rewrite activity and results
Runtime responsibilities:
- validate and normalize the prompt window
- choose the rewrite system prompt and model settings
- execute rewrite
- apply safety filtering
- replace authoritative seed prompt memory
- emit the canonical completion events
Controller responsibilities:
- proxy the request and response
- surface runtime availability and failure state
This boundary should not move.
## Recommended Rollout
1. Finish the path reorg so docs and code agree on `apps/dreamverse/web`.
2. Introduce `controller/` as a local-only API/proxy process.
3. Keep `apps/dreamverse/dreamverse/` as the runtime and adapt it behind the
controller.
4. Add `local` provider first.
5. Add "bring your own runtime URL" as an escape hatch.
6. Add automated Runpod provisioning.
7. Add Modal deployment support.
8. Rename `apps/dreamverse/dreamverse/` to `runtime/` once the split is stable.
## Implementation Plan
The implementation should start with the smallest milestone that gives users a
working local GPU setup without forcing the controller/provider architecture
into the first patch series.
### Milestone 0: Make local GPU the official baseline
Goal:
- A user with a working `fastvideo` install can run the Dreamverse backend on a
local GPU and connect to it from `apps/dreamverse/web`.
Non-goals for this milestone:
- no controller process yet
- no provider abstraction yet
- no Runpod or Modal support yet
- no secret-management UI yet
Reasoning:
- `apps/dreamverse/dreamverse/` already is the real local GPU runtime.
- `apps/dreamverse/web` already knows how to talk to a backend over `/ws` and
REST rewrites.
- The shortest path is to make the existing local path explicit, reliable, and
tested before adding another layer.
### Milestone 0 work items
#### 0.1 Fix repo path assumptions after the frontend move
Current issue:
- Some paths still assume `prod-ui/`, but the frontend now lives at
`apps/dreamverse/web/`.
Required changes:
- update prompt/preset path resolution in
`apps/dreamverse/dreamverse/config.py`
- update docs that still mention `prod-ui`
- audit any frontend build settings that assume the old repo root
This is prerequisite cleanup. Local GPU mode should not depend on stale
monorepo paths.
#### 0.2 Make local runtime startup the primary supported entrypoint
Required outcome:
- one documented backend command
- one documented frontend command
- one clear env contract for local development
Expected shape:
```bash
uv pip install -e ".[dreamverse]"
dreamverse-server --host 0.0.0.0 --port 8009
cd apps/dreamverse/web
npm ci
BACKEND_HOST=localhost BACKEND_PORT=8009 npm run dev
```
Optional but useful:
- add a small root helper script or Make target for local startup
- add a `dreamverse-doctor` or lightweight startup check later
#### 0.3 Define the minimum local runtime contract
For Milestone 0, the frontend should rely only on the current runtime surface:
- `/ws`
- `/status`
- `/healthz`
- `/readyz`
- existing prompt/devtools routes
Do not add a second local API layer yet unless the current runtime surface is
proven insufficient.
#### 0.4 Make failure states explicit in the UI
Local GPU mode fails in a few predictable ways:
- backend not reachable
- backend reachable but not ready
- `fastvideo` or model runtime missing
- no compatible GPU available
Minimum implementation:
- show a clear connection error when `/ws` or `/status` fails
- surface readiness failures in a human-readable way
- avoid silent retry loops that hide backend startup failures
This is a small UI pass, not a controller project.
#### 0.5 Add a minimal local smoke test path
At this milestone, local GPU support is "done" only if there is a repeatable
test path for the local runtime contract.
Minimum test additions:
- backend tests for `/healthz`, `/readyz`, and `/status`
- a frontend integration test that assumes a reachable backend URL and verifies
connection lifecycle behavior
- one local smoke script that starts the backend and verifies readiness before
the frontend is launched
### Milestone 1: Introduce a thin local controller
Goal:
- Preserve the same local GPU behavior, but place a stable local control-plane
API in front of the runtime.
This should happen only after Milestone 0 is stable.
Scope:
- add `controller/`
- proxy `/ws` and the needed REST routes to `apps/dreamverse/dreamverse/`
- expose controller-owned status for "backend starting", "runtime ready", and
"runtime failed"
- optionally spawn the local runtime as a subprocess
Non-goal:
- do not add remote provider logic yet
Reasoning:
- the controller earns its complexity only once it stabilizes the local
contract that future providers will share
### Milestone 2: Provider abstraction on top of the controller
Goal:
- Keep the same frontend contract while allowing the controller to resolve a
runtime via `local`, then later `runpod` and `modal`.
At this point:
- define `ComputeProvider`
- implement `providers/local.py`
- move local-runtime subprocess management behind the provider interface
The first provider should be `local`, because it is cheapest to debug and
matches the runtime most closely.
## Minimal Code Change Order
If we want the shortest path to a working local GPU milestone, the change order
should be:
1. Fix `apps/dreamverse/dreamverse/config.py` and any remaining path
assumptions from `prod-ui` to `apps/dreamverse/web`.
2. Update `README.md` to document the real local GPU startup flow.
3. Confirm `apps/dreamverse/web` connects cleanly to the local wrapper-backed
backend.
4. Improve frontend error handling for backend-not-ready and backend-missing
cases.
5. Add a local smoke test and keep existing backend/frontend tests green.
6. Only then introduce `controller/`.
## Test Plan for the Local GPU Milestone
### Backend
Keep the current Python test suite as the base:
- `apps/dreamverse/dreamverse/tests/test_health_endpoints.py`
- `apps/dreamverse/dreamverse/tests/test_mock_server.py`
- `apps/dreamverse/dreamverse/tests/test_prompt_enhancer.py`
- `apps/dreamverse/dreamverse/tests/test_rewrite_prompt_payload.py`
- related config and logging tests
Add or tighten tests for:
- path resolution in `apps/dreamverse/dreamverse/config.py`
- readiness behavior when GPU pool initialization fails
- startup error messaging when `fastvideo` is unavailable
### Frontend
Keep the current Vitest suite as the base:
- websocket reducer tests
- prompt-window snapshot tests
- integration tests under `apps/dreamverse/web/src/app/`
Add or tighten tests for:
- connection failure UX when backend is down
- readiness failure UX when backend returns non-ready status
- backend routing configuration through `BACKEND_HOST` and `BACKEND_PORT`
### Manual smoke path
The first manual smoke checklist should be:
1. start `dreamverse-server`
2. confirm `GET /healthz` returns 200
3. confirm `GET /readyz` returns 200 after warmup
4. start `apps/dreamverse/web`
5. confirm the UI opens and the websocket connects
6. submit a prompt and verify the first generation starts
This checklist should be written down in the README once the milestone is
implemented.
## Test Strategy
The test suite should preserve one rule: provider changes must not be able to
break prompt rewrite, websocket semantics, or runtime behavior silently.
### 1. Runtime unit and integration tests
Keep and expand the current `pytest` coverage in
`apps/dreamverse/dreamverse/tests/test_*.py`.
Focus areas:
- config loading
- prompt rewrite payload normalization
- prompt enhancement and prompt safety
- health/readiness endpoints
- websocket session behavior
- session logging
- mock runtime behavior
These tests should remain provider-agnostic.
### 2. Controller unit tests
Add a Python test suite for the controller state machine.
Key cases:
- provider selection and validation
- credential loading from local config or env
- runtime lifecycle transitions:
- idle
- provisioning
- ready
- failed
- stopping
- idle timeout and cleanup behavior
- retry and backoff behavior
- HTTP and websocket proxy routing
These tests should use fake providers and fake runtimes by default.
### 3. Provider contract tests
Each provider should pass the same contract tests.
Examples:
- `ensure_runtime()` returns a usable `RuntimeHandle`
- `wait_until_ready()` surfaces timeout vs readiness correctly
- `stop_runtime()` is safe to call twice
- provider errors are mapped into stable controller error types
Use recorded fixtures or fakes wherever possible to avoid spend in CI.
### 4. Web contract tests
The frontend already has useful Vitest coverage under
`apps/dreamverse/web/src`.
Preserve that and expand around the controller split.
Priority areas:
- websocket event reduction
- prompt-window snapshot construction
- rewrite request shaping
- compute-status UI
- failure and reconnect UX
The frontend should mock the local controller API, not provider APIs.
### 5. Cross-layer protocol tests
Add contract fixtures that validate shared payloads across layers.
Important fixtures:
- websocket event payloads
- rewrite request payloads
- runtime status payloads
- controller status payloads
These can be simple JSON fixtures validated by both Python and TypeScript
tests. They will catch drift earlier than end-to-end tests.
### 6. Smoke tests
Add a small number of high-signal smoke tests:
- local provider + mock runtime
- local provider + real runtime when GPU is available
- controller startup + frontend health path
These should be cheap enough for routine local use.
### 7. Provider-backed manual or nightly tests
Real Modal and Runpod tests should be opt-in.
- Do not run them in default CI.
- Gate them behind explicit credentials and flags.
- Keep them focused on provisioning and reachability, not full product
regression.
This avoids flaky and expensive CI while still validating real provider flows.
## Testing Recommendations for the Next Step
The next practical additions should be:
1. A controller test suite in Python using fake providers.
2. Shared contract fixtures for websocket and rewrite payloads.
3. A smoke test that starts the local controller against the existing mock
runtime.
I would not add Playwright yet. The current web stack already has Vitest and
integration-style component tests, which are cheaper and better aligned with
the immediate reorg. Add browser automation only after the controller path is
stable.
+40
View File
@@ -0,0 +1,40 @@
.git/
**/.git/
.venv/
**/.venv/
**/__pycache__/
**/*.pyc
**/*.pyo
**/*.egg-info/
.pytest_cache/
**/.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/dist/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/outputs/
apps/dreamverse/dreamverse/outputs/
apps/dreamverse/dreamverse/prompts.local/
apps/dreamverse/logs/
outputs/
logs/
slurm-logs/
wandb/
.env
.env.*
**/prompts.local/
.codex/
.agents/exploration/
.vscode/
.idea/
*.log
*.tmp
*.pdf
+76
View File
@@ -0,0 +1,76 @@
# syntax=docker/dockerfile:1.7
ARG CUDA_TAG=12.9.1-cudnn-devel-ubuntu22.04
FROM nvidia/cuda:${CUDA_TAG}
ARG BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=0
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
UV_LINK_MODE=copy
SHELL ["/bin/bash", "-c"]
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc-11 g++-11 clang-11 \
make cmake ninja-build pkg-config nasm \
git curl wget ca-certificates \
libssl-dev zlib1g-dev \
&& rm -rf /var/lib/apt/lists/* \
&& update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 \
--slave /usr/bin/g++ g++ /usr/bin/g++-11
ENV CUDA_HOME=/usr/local/cuda-12.9
ENV PATH=/root/.local/bin:/opt/venv/bin:${CUDA_HOME}/bin:${PATH}
ENV LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${LD_LIBRARY_PATH}
ENV VIRTUAL_ENV=/opt/venv
RUN curl -LsSf https://astral.sh/uv/install.sh | sh
RUN uv venv --python 3.12 --seed /opt/venv \
&& echo 'source /opt/venv/bin/activate' >> /root/.bashrc
WORKDIR /opt/FastVideo
COPY . /opt/FastVideo
RUN source /opt/venv/bin/activate \
&& uv pip install --no-cache-dir "/opt/FastVideo[dreamverse]"
# Standard docker build does not expose GPUs, while fastvideo-kernel/build.sh
# detects the CUDA architecture with torch at build time. The FastVideo package
# install above brings in the pinned fastvideo-kernel package; rebuild from the
# copied source only on hosts configured for build-time GPU access.
RUN if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
sed -i 's/^git submodule update --init --recursive$/if [[ -d ..\/.git ]]; then git submodule update --init --recursive; fi/' \
/opt/FastVideo/fastvideo-kernel/build.sh \
&& source /opt/venv/bin/activate \
&& cd /opt/FastVideo/fastvideo-kernel \
&& ./build.sh; \
else \
echo "Skipping source fastvideo-kernel build; using installed fastvideo-kernel package."; \
fi
# The monorepo ffmpeg installer force-selects conda compiler triplets for
# local dev shells. Inside this image we explicitly opt into the system
# gcc/g++ toolchain.
RUN source /opt/venv/bin/activate \
&& INSTALL_PREFIX=/opt/ffmpeg-native \
SOURCE_DIR=/tmp/ffmpeg-native-src \
FFMPEG_NATIVE_CC=/usr/bin/gcc \
FFMPEG_NATIVE_CXX=/usr/bin/g++ \
bash /opt/FastVideo/apps/dreamverse/scripts/install_native_ffmpeg.sh
ENV FASTVIDEO_DREAMVERSE_HOME=/var/lib/dreamverse \
STREAM_MODE=av_fmp4 \
FASTVIDEO_ENABLE_PROMPT_SAFETY=0 \
HF_HOME=/root/.cache/huggingface
RUN mkdir -p /var/lib/dreamverse
EXPOSE 8009
HEALTHCHECK --interval=30s --timeout=5s --start-period=180s --retries=3 \
CMD curl -fsS http://127.0.0.1:8009/healthz || exit 1
ENTRYPOINT ["/opt/FastVideo/apps/dreamverse/docker/docker_entrypoint.sh"]
CMD ["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"]
@@ -0,0 +1,40 @@
.git/
**/.git/
.venv/
**/.venv/
**/__pycache__/
**/*.pyc
**/*.pyo
**/*.egg-info/
.pytest_cache/
**/.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/dist/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/outputs/
apps/dreamverse/dreamverse/outputs/
apps/dreamverse/dreamverse/prompts.local/
apps/dreamverse/logs/
outputs/
logs/
slurm-logs/
wandb/
.env
.env.*
**/prompts.local/
.codex/
.agents/exploration/
.vscode/
.idea/
*.log
*.tmp
*.pdf
+72
View File
@@ -0,0 +1,72 @@
# Dreamverse Docker Image
This folder contains the backend-only Docker image for Dreamverse inside the
FastVideo monorepo. Build commands use the FastVideo repository root as the
Docker context, so run the helper scripts from this folder or from any path in
the checkout.
## Build
```bash
apps/dreamverse/docker/docker_build.sh
```
The image defaults to `dreamverse:dev`. Override it with:
```bash
DREAMVERSE_IMAGE=dreamverse:local apps/dreamverse/docker/docker_build.sh
```
The Dockerfile builds a CUDA 12.9.1 image, installs FastVideo from this
checkout with the `dreamverse` extra, installs the FA4
flash-attention fork, builds native FFmpeg, and installs FlashInfer for NVFP4
quantization.
FastVideo's pinned `fastvideo-kernel==0.2.6` package is installed by default.
To rebuild `fastvideo-kernel` from this checkout during the image build, set:
```bash
BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=1 apps/dreamverse/docker/docker_build.sh
```
That source build detects the GPU architecture with torch during `docker
build`. On hosts where Docker does not expose GPUs during build, leave the
default package install path enabled.
## Run
```bash
CEREBRAS_API_KEY="<your-key>" \
GROQ_API_KEY="<your-key>" \
apps/dreamverse/docker/docker_run.sh
```
The container serves Dreamverse on host port `8009` by default and mounts:
```text
$HOME/.cache/huggingface -> /root/.cache/huggingface
apps/dreamverse/outputs -> /var/lib/dreamverse/outputs
```
Override the host port and output directory with `BACKEND_PORT` and
`DREAMVERSE_OUTPUTS_DIR`.
To pin the container to a specific host GPU, pass Docker's GPU request syntax:
```bash
DREAMVERSE_DOCKER_GPUS=device=4 FASTVIDEO_GPU_COUNT=1 \
CEREBRAS_API_KEY="<your-key>" \
GROQ_API_KEY="<your-key>" \
apps/dreamverse/docker/docker_run.sh
```
## Smoke
```bash
CEREBRAS_API_KEY=placeholder \
GROQ_API_KEY=placeholder \
apps/dreamverse/docker/docker_smoke.sh
```
The smoke script starts the container, polls `/healthz`, then polls `/readyz`.
It removes the container on exit unless `DREAMVERSE_KEEP_CONTAINER=1` is set.
+17
View File
@@ -0,0 +1,17 @@
#!/usr/bin/env bash
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "${SCRIPT_DIR}/../../.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
build_args=()
[[ -n "${CUDA_TAG:-}" ]] && build_args+=(--build-arg "CUDA_TAG=${CUDA_TAG}")
[[ -n "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE:-}" ]] && \
build_args+=(--build-arg "BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}")
exec docker build \
-f "${SCRIPT_DIR}/Dockerfile" \
-t "${IMAGE}" \
"${build_args[@]}" \
"${REPO_ROOT}"
+13
View File
@@ -0,0 +1,13 @@
#!/usr/bin/env bash
set -euo pipefail
source /opt/venv/bin/activate
if [[ -f /opt/FastVideo/apps/dreamverse/scripts/ffmpeg-env.sh ]]; then
source /opt/FastVideo/apps/dreamverse/scripts/ffmpeg-env.sh
fi
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY must be set (pass with -e CEREBRAS_API_KEY=...)}"
: "${GROQ_API_KEY:?GROQ_API_KEY must be set (pass with -e GROQ_API_KEY=...)}"
exec "$@"
+30
View File
@@ -0,0 +1,30 @@
#!/usr/bin/env bash
set -euo pipefail
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY not set on host}"
: "${GROQ_API_KEY:?GROQ_API_KEY not set on host}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
PORT="${BACKEND_PORT:-8009}"
HF_CACHE="${HF_HOME:-$HOME/.cache/huggingface}"
OUTPUTS_DIR="${DREAMVERSE_OUTPUTS_DIR:-${DREAMVERSE_ROOT}/outputs}"
GPU_REQUEST="${DREAMVERSE_DOCKER_GPUS:-all}"
mkdir -p "${HF_CACHE}" "${OUTPUTS_DIR}"
env_args=(
-e "CEREBRAS_API_KEY=${CEREBRAS_API_KEY}"
-e "GROQ_API_KEY=${GROQ_API_KEY}"
)
[[ -n "${ENABLE_TORCH_COMPILE:-}" ]] && env_args+=(-e "ENABLE_TORCH_COMPILE=${ENABLE_TORCH_COMPILE}")
[[ -n "${FASTVIDEO_GPU_COUNT:-}" ]] && env_args+=(-e "FASTVIDEO_GPU_COUNT=${FASTVIDEO_GPU_COUNT}")
exec docker run --rm --gpus "${GPU_REQUEST}" --init \
-p "${PORT}:8009" \
"${env_args[@]}" \
-v "${HF_CACHE}:/root/.cache/huggingface" \
-v "${OUTPUTS_DIR}:/var/lib/dreamverse/outputs" \
"${IMAGE}"
+75
View File
@@ -0,0 +1,75 @@
#!/usr/bin/env bash
set -euo pipefail
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY not set on host}"
: "${GROQ_API_KEY:?GROQ_API_KEY not set on host}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
PORT="${BACKEND_PORT:-8009}"
HF_CACHE="${HF_HOME:-$HOME/.cache/huggingface}"
OUTPUTS_DIR="${DREAMVERSE_OUTPUTS_DIR:-${DREAMVERSE_ROOT}/outputs}"
NAME="${DREAMVERSE_NAME:-dreamverse}"
TIMEOUT_SECONDS="${DREAMVERSE_SMOKE_TIMEOUT_SECONDS:-1200}"
POLL_SECONDS="${DREAMVERSE_SMOKE_POLL_SECONDS:-5}"
GPU_REQUEST="${DREAMVERSE_DOCKER_GPUS:-all}"
mkdir -p "${HF_CACHE}" "${OUTPUTS_DIR}"
docker rm -f "${NAME}" >/dev/null 2>&1 || true
env_args=(
-e "CEREBRAS_API_KEY=${CEREBRAS_API_KEY}"
-e "GROQ_API_KEY=${GROQ_API_KEY}"
-e "ENABLE_TORCH_COMPILE=${ENABLE_TORCH_COMPILE:-0}"
)
[[ -n "${FASTVIDEO_GPU_COUNT:-}" ]] && env_args+=(-e "FASTVIDEO_GPU_COUNT=${FASTVIDEO_GPU_COUNT}")
container_id="$(
docker run -d --rm --gpus "${GPU_REQUEST}" --init \
-p "${PORT}:8009" \
"${env_args[@]}" \
-v "${HF_CACHE}:/root/.cache/huggingface" \
-v "${OUTPUTS_DIR}:/var/lib/dreamverse/outputs" \
--name "${NAME}" \
"${IMAGE}"
)"
cleanup() {
if [[ "${DREAMVERSE_KEEP_CONTAINER:-0}" != "1" ]]; then
docker rm -f "${NAME}" >/dev/null 2>&1 || true
fi
}
trap cleanup EXIT
wait_for_endpoint() {
local path="$1"
local label="$2"
local deadline=$((SECONDS + TIMEOUT_SECONDS))
local url="http://127.0.0.1:${PORT}${path}"
echo "Waiting for ${label} at ${url}"
while (( SECONDS < deadline )); do
if curl -fsS "${url}" >/dev/null 2>&1; then
echo "${label} ok"
return 0
fi
if ! docker ps --format '{{.Names}}' | grep -qx "${NAME}"; then
echo "Container exited before ${label} became healthy." >&2
docker logs "${container_id}" >&2 || true
return 1
fi
sleep "${POLL_SECONDS}"
done
echo "Timed out waiting for ${label}." >&2
docker logs "${container_id}" >&2 || true
return 1
}
wait_for_endpoint "/healthz" "healthz"
wait_for_endpoint "/readyz" "readyz"
echo "Dreamverse Docker smoke passed for ${IMAGE} on host port ${PORT}."
+15
View File
@@ -0,0 +1,15 @@
from __future__ import annotations
DREAMVERSE_RUNTIME_DEPS_MESSAGE = (
"Dreamverse runtime deps missing — install with pip install 'fastvideo[dreamverse]'.")
def require_dreamverse_runtime_deps() -> None:
try:
import cerebras.cloud.sdk # noqa: F401
import openai # noqa: F401
except ModuleNotFoundError as exc:
missing_root = (exc.name or "").split(".", 1)[0]
if missing_root in {"cerebras", "openai"}:
raise SystemExit(DREAMVERSE_RUNTIME_DEPS_MESSAGE) from exc
raise
+437
View File
@@ -0,0 +1,437 @@
# pyright: reportMissingTypeArgument=false, reportArgumentType=false, reportOptionalSubscript=false, reportOptionalMemberAccess=false, reportConstantRedefinition=false, reportCallIssue=false
# ruff: noqa: UP007, SIM108, SIM105
# mypy: ignore-errors
"""ffmpeg fMP4 muxing with chunk-level event emission.
Self-contained: spawns ffmpeg as a subprocess, pipes raw frames into
its stdin, reads fragmented-MP4 chunks from stdout, and publishes each
chunk as a ``StreamEvent`` via the caller-supplied ``publish``
callback. Knows nothing about multiprocessing queues, the GPU pool,
or individual users — the caller decides what "publish" means.
"""
import fcntl
import os
import shutil
import subprocess
import tempfile
import threading
import time
import uuid
import wave
from dataclasses import dataclass
from typing import Union
from collections.abc import Callable
import numpy as np
import torch
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
AV_MEDIA_MIME = os.getenv(
"STREAM_MIME_TYPE",
'video/mp4; codecs="avc1.42E01E,mp4a.40.2"',
)
AV_CHUNK_SIZE_BYTES = 1048576
TARGET_FPS = 24
AV_FRAGMENT_DURATION_US = int(os.getenv("FASTVIDEO_FRAG_US", "250000"))
X264_GOP_FRAMES = int(os.getenv("FASTVIDEO_X264_GOP", "12"))
X264_PROFILE = os.getenv("FASTVIDEO_X264_PROFILE", "baseline").strip().lower()
if X264_PROFILE not in {"baseline", "main", "high", "main10", "high10"}:
print(f"[WARN] Unsupported FASTVIDEO_X264_PROFILE={X264_PROFILE}; using baseline")
X264_PROFILE = "baseline"
USE_SHARED_STREAM_BUFFER = (os.getenv("FASTVIDEO_USE_SHARED_STREAM_BUFFER", "1").strip().lower()
not in {"0", "false", "no"})
SHARED_STREAM_BUFFER_BYTES = int(os.getenv("FASTVIDEO_SHARED_STREAM_BUFFER_BYTES", str(256 * 1024 * 1024)))
@dataclass
class StreamInit:
"""First event emitted — tells the consumer the stream is starting."""
stream_id: str
mime: str
uses_shared_buffer: bool
@dataclass
class StreamChunk:
"""One fMP4 chunk. Either ``chunk`` (raw bytes) or
``chunk_offset``+``chunk_length`` (read from the shared buffer)
will be populated, never both."""
stream_id: str
chunk: bytes | None = None
chunk_offset: int | None = None
chunk_length: int | None = None
uses_shared_buffer: bool = False
@dataclass
class StreamComplete:
"""Final event emitted — muxing finished successfully."""
stream_id: str
chunks: int
StreamEvent = Union[StreamInit, StreamChunk, StreamComplete]
def generate_stream_id(segment_idx: int) -> str:
"""Convenience: build a stream id of the form ``seg007-abcd1234``."""
return f"seg{segment_idx:03d}-{uuid.uuid4().hex[:8]}"
def _normalize_audio_tensor(audio: object) -> tuple[np.ndarray, int] | None:
"""Convert audio tensor/array into int16 ndarray [samples, channels]."""
if audio is None:
return None
if torch.is_tensor(audio):
audio_np = audio.detach().cpu().float().numpy()
else:
audio_np = np.asarray(audio, dtype=np.float32)
if audio_np.ndim == 1:
audio_np = audio_np[:, None]
elif audio_np.ndim == 2:
if audio_np.shape[0] <= 8 and audio_np.shape[1] > audio_np.shape[0]:
audio_np = audio_np.T
else:
return None
audio_np = np.clip(audio_np, -1.0, 1.0)
audio_int16 = (audio_np * 32767.0).astype(np.int16)
num_channels = audio_int16.shape[1]
return audio_int16, num_channels
def _write_audio_wav(
audio_int16: np.ndarray,
num_channels: int,
sample_rate: int,
) -> str:
"""Write normalized int16 audio to a temporary WAV file."""
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as f:
wav_path = f.name
with wave.open(wav_path, "wb") as wav_file:
wav_file.setnchannels(num_channels)
wav_file.setsampwidth(2)
wav_file.setframerate(sample_rate)
wav_file.writeframes(audio_int16.tobytes())
return wav_path
def stream_fmp4(
*,
frames: list[np.ndarray],
audio: object,
audio_sample_rate: int | None,
stream_id: str,
timings: dict,
head_trim_frames: int = 0,
head_trim_audio_frames: int | None = None,
shared_buffer=None,
shared_buffer_bytes: int = 0,
publish: Callable[[StreamEvent], None],
log_prefix: str = "",
) -> tuple[bool, str | None]:
"""Encode frames+audio with ffmpeg, publish each fMP4 chunk as an event.
Args:
frames: RGB24 video frames as HxWx3 uint8 arrays.
audio: 1D/2D tensor or ndarray, float values in [-1, 1].
audio_sample_rate: sample rate of ``audio``.
stream_id: caller-supplied identifier carried on every event.
timings: dict mutated in place with ffmpeg/stream timing metrics.
head_trim_frames: video frames to drop from the start
(conditioning overlap).
head_trim_audio_frames: video-frame-equivalent audio to drop.
Defaults to ``head_trim_frames``.
shared_buffer: optional ``mp.RawArray``-compatible object; when
provided, chunks are written into it and emitted by offset
rather than by bytes (avoids IPC copies).
shared_buffer_bytes: size of ``shared_buffer`` in bytes.
publish: callback invoked once per stream event.
log_prefix: prepended to warning prints (e.g. ``"[GPU 0]"``).
Returns:
``(True, None)`` on success, ``(False, error_message)`` on
failure. On mid-stream failure, a ``StreamInit`` may have
already been published — the caller is responsible for
handling that.
"""
if not frames:
return False, "no frames returned"
if audio is None:
return False, "audio is None"
if audio_sample_rate is None:
return False, "audio_sample_rate is None"
if FFMPEG_BIN is None:
return False, "ffmpeg not found"
if head_trim_audio_frames is None:
head_trim_audio_frames = head_trim_frames
normalized_audio = _normalize_audio_tensor(audio)
if normalized_audio is None:
shape_hint = getattr(audio, "shape", None)
return False, f"unsupported audio shape={shape_hint}"
audio_int16, num_channels = normalized_audio
if head_trim_frames < 0:
return False, (f"head_trim_frames must be >= 0, "
f"got {head_trim_frames}")
if head_trim_frames >= len(frames):
return False, (f"head_trim_frames={head_trim_frames} removes "
f"all {len(frames)} frames in segment")
out_frames = (frames[head_trim_frames:] if head_trim_frames > 0 else frames)
sample_rate = int(audio_sample_rate)
if head_trim_audio_frames > 0:
trim_start_samples = int(round((head_trim_audio_frames / float(TARGET_FPS)) * sample_rate))
if trim_start_samples >= audio_int16.shape[0]:
return False, ("audio too short after overlap trim: "
f"trim_start_samples={trim_start_samples}"
f", audio_samples={audio_int16.shape[0]}")
keep_samples = int(round((len(out_frames) / float(TARGET_FPS)) * sample_rate))
trim_end_samples = min(
audio_int16.shape[0],
trim_start_samples + keep_samples,
)
if trim_end_samples <= trim_start_samples:
return False, ("invalid audio trim range: "
f"start={trim_start_samples}, "
f"end={trim_end_samples}")
audio_int16 = audio_int16[trim_start_samples:trim_end_samples]
height = int(out_frames[0].shape[0])
width = int(out_frames[0].shape[1])
codec = os.getenv("FASTVIDEO_VIDEO_CODEC", "libx264")
t_wav_start = time.perf_counter()
wav_path = _write_audio_wav(audio_int16, num_channels, sample_rate)
wav_write_ms = (time.perf_counter() - t_wav_start) * 1000
cmd = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"rawvideo",
"-pix_fmt",
"rgb24",
"-s:v",
f"{width}x{height}",
"-r",
str(TARGET_FPS),
"-i",
"pipe:0",
"-i",
wav_path,
"-c:v",
codec,
]
if codec.endswith("_nvenc"):
cmd += [
"-preset",
os.getenv("FASTVIDEO_NVENC_PRESET", "p1"),
"-tune",
os.getenv("FASTVIDEO_NVENC_TUNE", "ull"),
"-rc",
os.getenv("FASTVIDEO_NVENC_RC", "constqp"),
"-qp",
os.getenv("FASTVIDEO_NVENC_QP", "28"),
"-bf",
os.getenv("FASTVIDEO_NVENC_BF", "0"),
]
else:
cmd += [
"-preset",
os.getenv("FASTVIDEO_X264_PRESET", "ultrafast"),
"-tune",
"zerolatency",
"-profile:v",
X264_PROFILE,
# Emit frequent keyframes so fragments are independently playable.
"-g",
str(X264_GOP_FRAMES),
"-keyint_min",
str(X264_GOP_FRAMES),
"-x264-params",
"scenecut=0",
]
cmd += [
"-c:a",
"aac",
"-pix_fmt",
os.getenv("FASTVIDEO_OUTPUT_PIX_FMT", "yuv420p"),
"-shortest",
"-movflags",
"+frag_keyframe+empty_moov+default_base_moof",
"-frag_duration",
str(AV_FRAGMENT_DURATION_US),
"-flush_packets",
"1",
"-muxdelay",
"0",
"-muxpreload",
"0",
"-f",
"mp4",
"pipe:1",
]
proc: subprocess.Popen | None = None
stderr_chunks: list[bytes] = []
writer_error: list[Exception | None] = [None]
t_stream_start = time.perf_counter()
use_shared_buffer = (USE_SHARED_STREAM_BUFFER and shared_buffer is not None and shared_buffer_bytes > 0)
shared_write_offset = 0
shared_buffer_fallback = False
shared_np = (np.frombuffer(
shared_buffer,
dtype=np.uint8,
count=shared_buffer_bytes,
) if use_shared_buffer else None)
chunk_intervals_ms: list[float] = []
chunk_publish_ms: list[float] = []
chunk_read_ms: list[float] = []
try:
t_proc_spawn_start = time.perf_counter()
proc = subprocess.Popen(
cmd,
stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
bufsize=0,
)
assert proc.stdin is not None
assert proc.stdout is not None
assert proc.stderr is not None
fcntl.fcntl(proc.stdin.fileno(), fcntl.F_SETPIPE_SZ, 1048576)
fcntl.fcntl(proc.stdout.fileno(), fcntl.F_SETPIPE_SZ, 1048576)
ffmpeg_spawn_ms = (time.perf_counter() - t_proc_spawn_start) * 1000
def _write_frames():
try:
for frame in out_frames:
proc.stdin.write(np.ascontiguousarray(frame).tobytes())
proc.stdin.close()
except Exception as exc:
writer_error[0] = exc
try:
proc.stdin.close()
except Exception:
pass
def _read_stderr():
try:
while True:
data = proc.stderr.read(4096)
if not data:
break
stderr_chunks.append(data)
except Exception:
pass
writer_thread = threading.Thread(target=_write_frames, daemon=True)
stderr_thread = threading.Thread(target=_read_stderr, daemon=True)
writer_thread.start()
stderr_thread.start()
publish(StreamInit(
stream_id=stream_id,
mime=AV_MEDIA_MIME,
uses_shared_buffer=use_shared_buffer,
))
chunk_count = 0
total_bytes = 0
first_chunk_ms: float | None = None
last_chunk_emit_t = time.perf_counter()
while True:
t_read_start = time.perf_counter()
chunk = proc.stdout.read(AV_CHUNK_SIZE_BYTES)
t_read_end = time.perf_counter()
if not chunk:
break
chunk_read_ms.append((t_read_end - t_read_start) * 1000)
chunk_count += 1
total_bytes += len(chunk)
if first_chunk_ms is None:
first_chunk_ms = (t_read_end - t_stream_start) * 1000
chunk_intervals_ms.append((t_read_end - last_chunk_emit_t) * 1000)
t_publish_start = time.perf_counter()
if use_shared_buffer and not shared_buffer_fallback:
chunk_len = len(chunk)
write_end = shared_write_offset + chunk_len
if write_end <= shared_buffer_bytes:
shared_np[shared_write_offset:write_end] = np.frombuffer(chunk, dtype=np.uint8)
publish(
StreamChunk(
stream_id=stream_id,
chunk_offset=shared_write_offset,
chunk_length=chunk_len,
uses_shared_buffer=True,
))
shared_write_offset = write_end
chunk_publish_ms.append((time.perf_counter() - t_publish_start) * 1000)
last_chunk_emit_t = time.perf_counter()
continue
shared_buffer_fallback = True
print(f"{log_prefix} Shared stream buffer exhausted at "
f"{shared_write_offset / (1024 * 1024):.1f}MB; "
"falling back to queue chunk bytes")
publish(StreamChunk(
stream_id=stream_id,
chunk=chunk,
))
chunk_publish_ms.append((time.perf_counter() - t_publish_start) * 1000)
last_chunk_emit_t = time.perf_counter()
writer_thread.join(timeout=5.0)
rc = proc.wait()
stderr_thread.join(timeout=1.0)
if rc != 0:
stderr_tail = b"".join(stderr_chunks).decode(errors="ignore")[-1200:]
return False, f"ffmpeg av_fmp4 stream failed (rc={rc}): {stderr_tail}"
if writer_error[0] is not None:
return False, f"ffmpeg frame writer failed: {writer_error[0]}"
timings["av_encode_stream_ms"] = (time.perf_counter() - t_stream_start) * 1000
timings["av_stream_bytes"] = total_bytes
timings["av_trim_head_frames"] = head_trim_frames
timings["av_trim_head_audio_frames"] = head_trim_audio_frames
timings["av_frames_encoded"] = len(out_frames)
timings["av_shared_buffer_used"] = (bool(use_shared_buffer and not shared_buffer_fallback))
timings["av_wav_write_ms"] = wav_write_ms
timings["av_ffmpeg_spawn_ms"] = ffmpeg_spawn_ms
timings["av_first_chunk_ms"] = first_chunk_ms or 0.0
if chunk_intervals_ms:
timings["av_chunk_interval_ms_min"] = min(chunk_intervals_ms)
timings["av_chunk_interval_ms_median"] = float(np.median(chunk_intervals_ms))
timings["av_chunk_interval_ms_p95"] = (float(np.percentile(chunk_intervals_ms, 95)))
timings["av_chunk_interval_ms_max"] = max(chunk_intervals_ms)
if chunk_publish_ms:
timings["av_chunk_publish_ms_median"] = float(np.median(chunk_publish_ms))
timings["av_chunk_publish_ms_p95"] = float(np.percentile(chunk_publish_ms, 95))
if chunk_read_ms:
timings["av_chunk_read_ms_median"] = float(np.median(chunk_read_ms))
timings["av_chunk_read_ms_p95"] = float(np.percentile(chunk_read_ms, 95))
publish(StreamComplete(
stream_id=stream_id,
chunks=chunk_count,
))
return True, None
except Exception as exc:
return False, str(exc)
finally:
if proc is not None and proc.poll() is None:
try:
proc.kill()
except Exception:
pass
try:
os.remove(wav_path)
except OSError:
pass
@@ -0,0 +1,288 @@
"""Benchmark the AV streaming hot-path used by dreamverse-server.
Measures wall-time, encoded byte volume, and chunk count for
``av_streaming.stream_fmp4`` over synthetic frames + audio at
production resolution. Sweeps codecs (default: ``libx264`` and
``h264_nvenc`` if the active ffmpeg supports it) and ffmpeg presets so
the deploy can pick a configuration that achieves a >=1.0 realtime
ratio (5.04s of generated video produced in <=5.04s wall-time).
Usage::
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming \\
--frames 121 --width 1920 --height 1088 --runs 3 \\
--codecs libx264 h264_nvenc --x264-preset ultrafast \\
--nvenc-preset p1
FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg \\
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming
Skips ``h264_nvenc`` automatically if the binary lacks the encoder.
This is the regression guard documented in
`.agents/memory/dreamverse-integration/decisions-log.md` D-21.
"""
from __future__ import annotations
import argparse
import os
import shutil
import statistics
import subprocess
import sys
import time
from collections.abc import Iterable
from dataclasses import dataclass
import numpy as np
import torch
_HERE = os.path.dirname(os.path.abspath(__file__))
_SERVER_ROOT = os.path.dirname(_HERE)
if _SERVER_ROOT not in sys.path:
sys.path.insert(0, _SERVER_ROOT)
from dreamverse.av_streaming import stream_fmp4 # noqa: E402
@dataclass
class BenchResult:
codec: str
preset: str
runs: int
wall_ms_min: float
wall_ms_median: float
wall_ms_p95: float
wall_ms_max: float
bytes_median: int
chunks_median: float
realtime_ratio_median: float
error: str | None = None
def _make_synthetic_frames(num: int, width: int, height: int, seed: int) -> list[np.ndarray]:
rng = np.random.default_rng(seed)
base = rng.integers(0, 255, size=(height, width, 3), dtype=np.uint8)
out: list[np.ndarray] = []
for i in range(num):
f = base.copy()
f[:, :, 0] = (f[:, :, 0].astype(np.int32) + i * 2) % 256
out.append(f)
return out
def _make_synthetic_audio(num_frames: int, fps: int, sample_rate: int, seed: int) -> torch.Tensor:
duration_s = num_frames / fps
samples = int(round(duration_s * sample_rate))
rng = np.random.default_rng(seed + 1)
audio = (rng.uniform(-0.1, 0.1, (2, samples))).astype(np.float32)
return torch.from_numpy(audio)
def _ffmpeg_supports(codec: str, ffmpeg_bin: str) -> bool:
try:
out = subprocess.run(
[ffmpeg_bin, "-hide_banner", "-encoders"],
capture_output=True,
text=True,
check=False,
timeout=10,
).stdout
except Exception:
return False
needle = f" {codec} "
return any(needle in line for line in out.splitlines())
def _run_one(frames: list[np.ndarray], audio: torch.Tensor, sample_rate: int, codec: str,
preset: str) -> tuple[float, int, int, str | None]:
timings: dict = {}
chunks: list = []
def _publish(event):
chunks.append(event)
prev_codec = os.environ.get("FASTVIDEO_VIDEO_CODEC")
prev_preset = os.environ.get("FASTVIDEO_X264_PRESET")
prev_nvenc_preset = os.environ.get("FASTVIDEO_NVENC_PRESET")
os.environ["FASTVIDEO_VIDEO_CODEC"] = codec
if codec.endswith("_nvenc"):
os.environ["FASTVIDEO_NVENC_PRESET"] = preset
else:
os.environ["FASTVIDEO_X264_PRESET"] = preset
t0 = time.perf_counter()
try:
ok, err = stream_fmp4(
frames=frames,
audio=audio,
audio_sample_rate=sample_rate,
stream_id="bench",
timings=timings,
head_trim_frames=0,
head_trim_audio_frames=0,
shared_buffer=None,
shared_buffer_bytes=0,
publish=_publish,
log_prefix="[bench]",
)
finally:
if prev_codec is None:
os.environ.pop("FASTVIDEO_VIDEO_CODEC", None)
else:
os.environ["FASTVIDEO_VIDEO_CODEC"] = prev_codec
if prev_preset is None:
os.environ.pop("FASTVIDEO_X264_PRESET", None)
else:
os.environ["FASTVIDEO_X264_PRESET"] = prev_preset
if prev_nvenc_preset is None:
os.environ.pop("FASTVIDEO_NVENC_PRESET", None)
else:
os.environ["FASTVIDEO_NVENC_PRESET"] = prev_nvenc_preset
wall_ms = (time.perf_counter() - t0) * 1000.0
if not ok:
return wall_ms, 0, 0, err or "stream_fmp4 returned False"
total_bytes = int(timings.get("av_stream_bytes", 0))
return wall_ms, total_bytes, len(chunks), None
def benchmark(codecs: Iterable[str], runs: int, frames_n: int, width: int, height: int, fps: int, sample_rate: int,
x264_preset: str, nvenc_preset: str, seed: int, ffmpeg_bin: str) -> list[BenchResult]:
print(f"[bench] ffmpeg_bin={ffmpeg_bin}")
print(f"[bench] frames={frames_n} {width}x{height} fps={fps} "
f"audio_sr={sample_rate} runs/codec={runs}")
frames = _make_synthetic_frames(frames_n, width, height, seed)
audio = _make_synthetic_audio(frames_n, fps, sample_rate, seed)
playable_s = frames_n / fps
print(f"[bench] playable={playable_s:.3f}s "
f"(realtime_ratio = playable / wall_time; >= 1.0 means no "
f"buffer drain)")
results: list[BenchResult] = []
for codec in codecs:
preset = nvenc_preset if codec.endswith("_nvenc") else x264_preset
if not _ffmpeg_supports(codec, ffmpeg_bin):
results.append(
BenchResult(codec=codec,
preset=preset,
runs=0,
wall_ms_min=0,
wall_ms_median=0,
wall_ms_p95=0,
wall_ms_max=0,
bytes_median=0,
chunks_median=0,
realtime_ratio_median=0,
error=f"{codec} not in ffmpeg"))
continue
walls: list[float] = []
sizes: list[int] = []
chunkcounts: list[int] = []
last_err: str | None = None
for run in range(runs):
wall_ms, total_bytes, chunk_count, err = _run_one(frames, audio, sample_rate, codec, preset)
print(f"[bench] codec={codec:12s} preset={preset:9s} "
f"run={run + 1}/{runs} wall={wall_ms:7.1f}ms "
f"bytes={total_bytes:>9d} chunks={chunk_count:>3d} "
f"realtime={playable_s / (wall_ms / 1000.0):5.2f}x"
f"{' ERR=' + err if err else ''}")
if err is not None:
last_err = err
continue
walls.append(wall_ms)
sizes.append(total_bytes)
chunkcounts.append(chunk_count)
if not walls:
results.append(
BenchResult(codec=codec,
preset=preset,
runs=0,
wall_ms_min=0,
wall_ms_median=0,
wall_ms_p95=0,
wall_ms_max=0,
bytes_median=0,
chunks_median=0,
realtime_ratio_median=0,
error=last_err or "all runs failed"))
continue
walls_sorted = sorted(walls)
p95_idx = max(0, int(round(0.95 * (len(walls_sorted) - 1))))
wall_med = statistics.median(walls)
results.append(
BenchResult(
codec=codec,
preset=preset,
runs=len(walls),
wall_ms_min=min(walls),
wall_ms_median=wall_med,
wall_ms_p95=walls_sorted[p95_idx],
wall_ms_max=max(walls),
bytes_median=int(statistics.median(sizes)),
chunks_median=statistics.median(chunkcounts),
realtime_ratio_median=playable_s / (wall_med / 1000.0),
))
return results
def _print_summary(results: list[BenchResult]) -> None:
print()
print("=== summary ===")
header = (f"{'codec':14s} {'preset':10s} {'runs':>4s} "
f"{'wall_med_ms':>11s} {'wall_p95_ms':>11s} "
f"{'bytes_med':>10s} {'realtime':>8s} notes")
print(header)
print("-" * len(header))
for r in results:
if r.error is not None:
print(f"{r.codec:14s} {r.preset:10s} {r.runs:>4d} "
f"{'-':>11s} {'-':>11s} {'-':>10s} {'-':>8s} "
f"ERR: {r.error}")
continue
print(f"{r.codec:14s} {r.preset:10s} {r.runs:>4d} "
f"{r.wall_ms_median:>11.1f} {r.wall_ms_p95:>11.1f} "
f"{r.bytes_median:>10d} {r.realtime_ratio_median:>7.2f}x "
f"{'OK' if r.realtime_ratio_median >= 1.0 else 'BUFFER DRAINS'}")
def main() -> int:
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--frames", type=int, default=121, help="num frames (default: 121, matches NUM_FRAMES)")
p.add_argument("--width", type=int, default=1920)
p.add_argument("--height", type=int, default=1088)
p.add_argument("--fps", type=int, default=24)
p.add_argument("--sample-rate", type=int, default=24000)
p.add_argument("--runs", type=int, default=3, help="runs per codec (default: 3 — 1 warmup + 2 timed in median)")
p.add_argument("--codecs",
nargs="+",
default=["libx264", "h264_nvenc"],
help="codecs to benchmark; missing ones are skipped")
p.add_argument("--x264-preset", default="ultrafast")
p.add_argument("--nvenc-preset", default="p1")
p.add_argument("--seed", type=int, default=0)
args = p.parse_args()
ffmpeg_bin = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
if ffmpeg_bin is None:
print("ffmpeg not found", file=sys.stderr)
return 2
results = benchmark(
codecs=args.codecs,
runs=args.runs,
frames_n=args.frames,
width=args.width,
height=args.height,
fps=args.fps,
sample_rate=args.sample_rate,
x264_preset=args.x264_preset,
nvenc_preset=args.nvenc_preset,
seed=args.seed,
ffmpeg_bin=ffmpeg_bin,
)
_print_summary(results)
any_below = any(r.error is None and r.realtime_ratio_median < 1.0 for r in results)
return 1 if any_below else 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,362 @@
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
Mirrors how ``apps/dreamverse/server/video_generation.py`` constructs
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
basic_ltx2_distilled_i2v_two_stage_time.py``).
Reports for each scenario:
* Total wall-time (median, p95)
* Per-stage execution_time (input_validation_stage,
prompt_encoding_stage, ltx2_refine_init_stage, latent_preparation_stage,
denoising_stage, ltx2_upsample_stage, ltx2_refine_lora_stage,
ltx2_refine_denoising_stage, audio_decoding_stage, decoding_stage)
* Realtime ratio (frames / fps / wall_time; >= 1.0 = no buffer drain)
* Peak GPU memory
Sweep scenarios (default):
* compile=False, warmup=False → cold inference baseline
* compile=True, warmup=False → JIT compile mid-run
* compile=True, warmup=True → fully warmed (production mode)
Skips compile / NVENC if the host lacks support. Does NOT exercise the
AV streaming path (use ``benchmark_av_streaming.py`` for that).
Usage::
python -m apps.dreamverse.server.benchmarks.benchmark_pipeline
python -m apps.dreamverse.server.benchmarks.benchmark_pipeline \\
--runs 3 --scenarios compile_warm cold --gpu 4
Cross-references D-21 / D-22 in
``.agents/memory/dreamverse-integration/decisions-log.md``.
"""
from __future__ import annotations
import argparse
import contextlib
import json
import os
import statistics
import time
from collections import OrderedDict
from dataclasses import asdict, dataclass
if "FASTVIDEO_STAGE_LOGGING" not in os.environ:
os.environ["FASTVIDEO_STAGE_LOGGING"] = "1"
if "FASTVIDEO_ATTENTION_BACKEND" not in os.environ:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
import torch # noqa: E402
from fastvideo import VideoGenerator # noqa: E402
from fastvideo.api import ( # noqa: E402
ComponentConfig, CompileConfig, EngineConfig, GeneratorConfig, OffloadConfig, PipelineSelection, QuantizationConfig,
)
DEFAULT_PROMPT = ("A cinematic drone shot over coastal cliffs at sunrise, golden "
"light, gentle ocean waves, ultra detailed")
DEFAULT_MODEL = "FastVideo/LTX2-Distilled-Diffusers"
@dataclass
class ScenarioConfig:
name: str
enable_compile: bool
do_warmup: bool
nvenc: bool = False
DEFAULT_SCENARIOS: list[ScenarioConfig] = [
ScenarioConfig(name="cold", enable_compile=False, do_warmup=False),
ScenarioConfig(name="compile_cold", enable_compile=True, do_warmup=False),
ScenarioConfig(name="compile_warm", enable_compile=True, do_warmup=True),
]
@dataclass
class RunResult:
wall_ms: float
stage_times_ms: OrderedDict[str, float]
peak_gpu_mb: float
error: str | None = None
@dataclass
class ScenarioResult:
name: str
enable_compile: bool
do_warmup: bool
runs: int
wall_ms_median: float
wall_ms_p95: float
stage_means_ms: OrderedDict[str, float]
realtime_ratio_median: float
peak_gpu_mb_max: float
error: str | None = None
def _build_generator_config(model_path: str, enable_compile: bool, num_gpus: int) -> GeneratorConfig:
components = ComponentConfig(config_root=model_path)
return GeneratorConfig(
model_path=model_path,
engine=EngineConfig(
num_gpus=num_gpus,
offload=OffloadConfig(dit=False, dit_layerwise=False, text_encoder=False, vae=False, pin_cpu_memory=True),
compile=CompileConfig(enabled=enable_compile,
text_encoder_enabled=enable_compile,
backend="inductor",
fullgraph=True,
mode="max-autotune-no-cudagraphs",
dynamic=False),
use_fsdp_inference=False,
quantization=QuantizationConfig(transformer_quant="NVFP4"),
),
pipeline=PipelineSelection(
components=components,
vae_tiling=False,
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 2,
"guidance_scale": 1.0,
"add_noise": True,
},
},
),
)
def _extract_stage_times(result: dict) -> OrderedDict[str, float]:
out: OrderedDict[str, float] = OrderedDict()
info = result.get("logging_info") if isinstance(result, dict) else None
if info is None:
return out
stages = getattr(info, "stages", None)
if not stages:
return out
for name, metrics in stages.items():
exec_time = float(metrics.get("execution_time", 0.0))
out[name] = exec_time * 1000.0
return out
def _peak_gpu_mb() -> float:
if not torch.cuda.is_available():
return 0.0
try:
return torch.cuda.max_memory_allocated() / (1024 * 1024)
except Exception:
return 0.0
def _reset_peak_gpu() -> None:
if torch.cuda.is_available():
with contextlib.suppress(Exception):
torch.cuda.reset_peak_memory_stats()
def _do_one_run(generator: VideoGenerator, prompt: str, *, height: int, width: int, num_frames: int, seed: int,
num_inference_steps: int) -> RunResult:
_reset_peak_gpu()
t0 = time.perf_counter()
try:
result = generator.generate_video(
prompt=prompt,
negative_prompt="",
save_video=False,
height=height,
width=width,
num_frames=num_frames,
fps=24,
num_inference_steps=num_inference_steps,
guidance_scale=1.0,
seed=seed,
ltx2_image_crf=0.0,
)
if torch.cuda.is_available():
torch.cuda.synchronize()
except Exception as exc:
return RunResult(wall_ms=(time.perf_counter() - t0) * 1000.0,
stage_times_ms=OrderedDict(),
peak_gpu_mb=0.0,
error=f"{type(exc).__name__}: {exc}")
wall_ms = (time.perf_counter() - t0) * 1000.0
return RunResult(
wall_ms=wall_ms,
stage_times_ms=_extract_stage_times(result),
peak_gpu_mb=_peak_gpu_mb(),
)
def benchmark_scenario(scenario: ScenarioConfig, model_path: str, num_gpus: int, prompt: str, num_runs: int,
num_frames: int, height: int, width: int, num_inference_steps: int, seed: int) -> ScenarioResult:
print()
print(f"=== scenario: {scenario.name} "
f"(compile={scenario.enable_compile} warmup={scenario.do_warmup}) ===")
config = _build_generator_config(model_path, scenario.enable_compile, num_gpus)
generator = VideoGenerator.from_config(config)
if scenario.do_warmup:
print(f"[{scenario.name}] warmup: 2 generate calls "
"(triggers compile + first-shape graphs)")
for warmup_idx in range(2):
t0 = time.perf_counter()
_do_one_run(generator,
prompt,
height=height,
width=width,
num_frames=num_frames,
seed=seed + 100 + warmup_idx,
num_inference_steps=num_inference_steps)
print(f"[{scenario.name}] warmup {warmup_idx + 1}: "
f"{(time.perf_counter() - t0) * 1000:.0f}ms")
runs: list[RunResult] = []
last_error: str | None = None
for run_idx in range(num_runs):
result = _do_one_run(generator,
prompt,
height=height,
width=width,
num_frames=num_frames,
seed=seed + run_idx,
num_inference_steps=num_inference_steps)
if result.error is not None:
print(f"[{scenario.name}] run {run_idx + 1}: "
f"ERROR {result.error}")
last_error = result.error
continue
playable_s = num_frames / 24.0
rt = playable_s / (result.wall_ms / 1000.0)
print(f"[{scenario.name}] run {run_idx + 1}/{num_runs}: "
f"wall={result.wall_ms:.0f}ms peak={result.peak_gpu_mb:.0f}MB "
f"realtime={rt:.2f}x stages={len(result.stage_times_ms)}")
runs.append(result)
if not runs:
return ScenarioResult(name=scenario.name,
enable_compile=scenario.enable_compile,
do_warmup=scenario.do_warmup,
runs=0,
wall_ms_median=0.0,
wall_ms_p95=0.0,
stage_means_ms=OrderedDict(),
realtime_ratio_median=0.0,
peak_gpu_mb_max=0.0,
error=last_error or "all runs failed")
walls = [r.wall_ms for r in runs]
walls_sorted = sorted(walls)
p95_idx = max(0, int(round(0.95 * (len(walls_sorted) - 1))))
stage_means: OrderedDict[str, float] = OrderedDict()
if runs:
for stage_name in runs[0].stage_times_ms:
vals = [r.stage_times_ms.get(stage_name, 0.0) for r in runs]
stage_means[stage_name] = sum(vals) / len(vals)
return ScenarioResult(
name=scenario.name,
enable_compile=scenario.enable_compile,
do_warmup=scenario.do_warmup,
runs=len(runs),
wall_ms_median=statistics.median(walls),
wall_ms_p95=walls_sorted[p95_idx],
stage_means_ms=stage_means,
realtime_ratio_median=(num_frames / 24.0) / (statistics.median(walls) / 1000.0),
peak_gpu_mb_max=max(r.peak_gpu_mb for r in runs),
)
def _print_summary(results: list[ScenarioResult]) -> None:
print()
print("=== summary ===")
header = (f"{'scenario':14s} {'runs':>4s} "
f"{'wall_med_ms':>11s} {'wall_p95_ms':>11s} "
f"{'peak_mb':>8s} {'realtime':>8s}")
print(header)
print("-" * len(header))
for r in results:
if r.error is not None:
print(f"{r.name:14s} {r.runs:>4d} ERR: {r.error}")
continue
print(f"{r.name:14s} {r.runs:>4d} "
f"{r.wall_ms_median:>11.0f} {r.wall_ms_p95:>11.0f} "
f"{r.peak_gpu_mb_max:>8.0f} {r.realtime_ratio_median:>7.2f}x")
for r in results:
if r.error is not None or not r.stage_means_ms:
continue
print()
print(f"=== {r.name} per-stage means (ms) ===")
for stage, mean_ms in r.stage_means_ms.items():
pct = 100.0 * mean_ms / r.wall_ms_median if r.wall_ms_median > 0 \
else 0.0
print(f" {stage:35s} {mean_ms:>9.1f}ms ({pct:5.1f}%)")
def main() -> int:
p = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--model", default=DEFAULT_MODEL)
p.add_argument("--prompt", default=DEFAULT_PROMPT)
p.add_argument("--scenarios",
nargs="+",
default=[s.name for s in DEFAULT_SCENARIOS],
choices=[s.name for s in DEFAULT_SCENARIOS],
help="which scenarios to run")
p.add_argument("--runs", type=int, default=3)
p.add_argument("--num-frames", type=int, default=121)
p.add_argument("--height", type=int, default=1088)
p.add_argument("--width", type=int, default=1920)
p.add_argument("--num-inference-steps", type=int, default=5)
p.add_argument("--num-gpus", type=int, default=1)
p.add_argument("--gpu", type=int, default=None, help="set CUDA_VISIBLE_DEVICES to this single GPU index")
p.add_argument("--seed", type=int, default=10)
p.add_argument("--output-json", default=None, help="write structured results to this path")
args = p.parse_args()
if args.gpu is not None:
os.environ["CUDA_VISIBLE_DEVICES"] = str(args.gpu)
selected = [s for s in DEFAULT_SCENARIOS if s.name in args.scenarios]
print(f"[bench] model={args.model} num_gpus={args.num_gpus}")
print(f"[bench] frames={args.num_frames} {args.width}x{args.height} "
f"steps={args.num_inference_steps} runs/scenario={args.runs}")
print(f"[bench] scenarios={[s.name for s in selected]}")
results: list[ScenarioResult] = []
for scenario in selected:
try:
result = benchmark_scenario(scenario, args.model, args.num_gpus, args.prompt, args.runs, args.num_frames,
args.height, args.width, args.num_inference_steps, args.seed)
except Exception as exc:
result = ScenarioResult(name=scenario.name,
enable_compile=scenario.enable_compile,
do_warmup=scenario.do_warmup,
runs=0,
wall_ms_median=0.0,
wall_ms_p95=0.0,
stage_means_ms=OrderedDict(),
realtime_ratio_median=0.0,
peak_gpu_mb_max=0.0,
error=f"{type(exc).__name__}: {exc}")
results.append(result)
_print_summary(results)
if args.output_json:
out = []
for r in results:
entry = asdict(r)
entry["stage_means_ms"] = dict(r.stage_means_ms)
out.append(entry)
with open(args.output_json, "w") as f:
json.dump({"args": vars(args), "results": out}, f, indent=2)
print(f"[bench] wrote {args.output_json}")
any_below = any(r.error is None and r.realtime_ratio_median < 1.0 for r in results)
return 1 if any_below else 0
if __name__ == "__main__":
raise SystemExit(main())
+281
View File
@@ -0,0 +1,281 @@
import os
from pathlib import Path
_REPO_ROOT = Path(__file__).resolve().parents[1]
_SERVER_ROOT = Path(__file__).resolve().parent
_FASTVIDEO_DREAMVERSE_HOME = os.environ.get("FASTVIDEO_DREAMVERSE_HOME")
_XDG_STATE_HOME = os.environ.get("XDG_STATE_HOME")
_DEFAULT_STATE_ROOT = (Path(_FASTVIDEO_DREAMVERSE_HOME) if _FASTVIDEO_DREAMVERSE_HOME else
(Path(_XDG_STATE_HOME) if _XDG_STATE_HOME else Path.home() / ".local/state") /
"fastvideo/dreamverse")
_OUTPUTS_ROOT = _DEFAULT_STATE_ROOT / "outputs"
_PROMPTS_ROOT = _SERVER_ROOT / "prompts"
_PROMPTS_LOCAL_ROOT = _SERVER_ROOT / "prompts.local"
_APP_ROOT = _REPO_ROOT
def _resolve_frontend_root() -> Path:
for candidate in (
_APP_ROOT / "web",
_APP_ROOT / "prod-ui",
):
if candidate.is_dir():
return candidate
return _APP_ROOT / "web"
FRONTEND_ROOT = _resolve_frontend_root()
_CLIENT_PROMPTS_ROOT = FRONTEND_ROOT / "prompts"
_CLIENT_PROMPTS_LOCAL_ROOT = FRONTEND_ROOT / "prompts.local"
FRONTEND_STATIC_DIR_CANDIDATES = tuple(str(FRONTEND_ROOT / dirname) for dirname in ("out", "dist"))
# Model registry
MODEL_REGISTRY = {
"fast-ltx2": {
"name": "FastLTX2",
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
},
"fast-ltx23": {
"name": "FastLTX23",
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
# Active model configuration
MODEL_CONFIG = MODEL_REGISTRY[DEFAULT_MODEL_ID]
# Generation limits
SESSION_TIMEOUT_SECONDS = 300
# Frame settings
NUM_FRAMES = 121
FRAME_HEIGHT = 1088
FRAME_WIDTH = 1920
NUM_INFERENCE_STEPS = 5
JPEG_QUALITY = 100
BATCH_SIZE = 3
# Streaming mode:
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
def _env_int(name: str, default: int) -> int:
value = os.getenv(name)
if value is None:
return default
try:
return int(value)
except ValueError:
return default
def _env_float(name: str, default: float) -> float:
value = os.getenv(name)
if value is None:
return default
try:
return float(value)
except ValueError:
return default
def _env_bool(name: str, default: bool) -> bool:
value = os.getenv(name)
if value is None:
return default
normalized = value.strip().lower()
if normalized in {"1", "true", "yes", "on"}:
return True
if normalized in {"0", "false", "no", "off"}:
return False
return default
def _env_choice(name: str, default: str, allowed: tuple[str, ...]) -> str:
value = os.getenv(name)
normalized = default.strip().lower() if value is None else value.strip().lower()
if normalized in allowed:
return normalized
allowed_values = ", ".join(allowed)
raise RuntimeError(f"Invalid {name}: {normalized!r}. Expected one of {allowed_values}.")
def _env_csv(name: str, default: str) -> list[str]:
raw = os.getenv(name, default)
values = [item.strip() for item in raw.split(",")]
unique_values: list[str] = []
for value in values:
if not value or value in unique_values:
continue
unique_values.append(value)
return unique_values
def _required_env(*names: str) -> str:
for name in names:
value = os.getenv(name)
if not isinstance(value, str):
continue
normalized = value.strip()
if normalized:
return normalized
joined_names = ", ".join(names)
raise RuntimeError(f"Missing required environment variable: one of {joined_names}")
def _optional_env(*names: str) -> str | None:
for name in names:
value = os.getenv(name)
if not isinstance(value, str):
continue
normalized = value.strip()
if normalized:
return normalized
return None
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
def _resolve_devtools_paths(
default_path: Path,
overlay_path: Path,
env_name: str | None = None,
) -> tuple[str, str | None]:
env_value = os.getenv(env_name) if env_name else None
if isinstance(env_value, str) and env_value.strip():
return env_value.strip(), None
if DEVTOOLS_ENABLED:
return str(overlay_path), str(default_path)
return str(default_path), None
# Prompt LLM configuration.
PROMPT_SUPPORTED_PROVIDERS = (
"cerebras",
"groq",
)
if os.getenv("FASTVIDEO_PROMPT_PROVIDER") is not None:
_env_choice(
"FASTVIDEO_PROMPT_PROVIDER",
"cerebras",
PROMPT_SUPPORTED_PROVIDERS,
)
PROMPT_PROVIDER = "cerebras"
PROMPT_PROVIDER_RUNTIME_STAGES = (("cerebras", "groq"), )
PROMPT_PROVIDER_PRIORITY = (
"cerebras",
"groq",
)
PROMPT_PROVIDER_API_KEY_NAMES = {
"cerebras": ("CEREBRAS_API_KEY", ),
"groq": ("GROQ_API_KEY", ),
}
PROMPT_API_KEYS = {
provider: _optional_env(*PROMPT_PROVIDER_API_KEY_NAMES[provider])
for provider in PROMPT_SUPPORTED_PROVIDERS
}
PROMPT_API_BASE_URLS = {
"cerebras": (os.getenv("FASTVIDEO_PROMPT_CEREBRAS_API_BASE_URL", "").strip() or None),
"groq": (os.getenv(
"FASTVIDEO_PROMPT_GROQ_API_BASE_URL",
"https://api.groq.com/openai/v1",
).strip() or None),
}
PROMPT_API_KEY = PROMPT_API_KEYS[PROMPT_PROVIDER]
PROMPT_API_BASE_URL = PROMPT_API_BASE_URLS[PROMPT_PROVIDER]
PROMPT_MODEL = (os.getenv("FASTVIDEO_PROMPT_MODEL", "gpt-oss-120b").strip() or "gpt-oss-120b")
_PROMPT_CEREBRAS_REQUEST_MODEL = (os.getenv("FASTVIDEO_PROMPT_CEREBRAS_MODEL", PROMPT_MODEL).strip() or PROMPT_MODEL)
PROMPT_PROVIDER_MODELS = {
"cerebras": _PROMPT_CEREBRAS_REQUEST_MODEL,
"groq": (os.getenv(
"FASTVIDEO_PROMPT_GROQ_MODEL",
f"openai/{PROMPT_MODEL}",
).strip() or f"openai/{PROMPT_MODEL}"),
}
PROMPT_REWRITE_MODEL = PROMPT_MODEL
PROMPT_REWRITE_MODEL_OPTIONS = [PROMPT_REWRITE_MODEL]
PROMPT_TIMEOUT_MS = 20000
PROMPT_HTTP_TIMEOUT_MS = 3000
PROMPT_INITIAL_STAGE_TIMEOUT_MS = 1500
PROMPT_TEMPERATURE = 1.0
PROMPT_MAX_COMPLETION_TOKENS = 3000
(
PROMPT_ENHANCE_SYSTEM_PROMPT_PATH,
PROMPT_ENHANCE_SYSTEM_PROMPT_FALLBACK_PATH,
) = _resolve_devtools_paths(
_PROMPTS_ROOT / "next_segment_system_prompt.md",
_PROMPTS_LOCAL_ROOT / "next_segment_system_prompt.md",
"FASTVIDEO_PROMPT_ENHANCE_SYSTEM_PROMPT_PATH",
)
(
PROMPT_AUTO_SYSTEM_PROMPT_PATH,
PROMPT_AUTO_SYSTEM_PROMPT_FALLBACK_PATH,
) = _resolve_devtools_paths(
_PROMPTS_ROOT / "auto_extension_system_prompt.md",
_PROMPTS_LOCAL_ROOT / "auto_extension_system_prompt.md",
"FASTVIDEO_PROMPT_AUTO_SYSTEM_PROMPT_PATH",
)
(
PROMPT_REWRITE_ALL_SYSTEM_PROMPT_PATH,
PROMPT_REWRITE_ALL_SYSTEM_PROMPT_FALLBACK_PATH,
) = _resolve_devtools_paths(
_PROMPTS_ROOT / "rewrite_window_system_prompt.md",
_PROMPTS_LOCAL_ROOT / "rewrite_window_system_prompt.md",
"FASTVIDEO_PROMPT_REWRITE_ALL_SYSTEM_PROMPT_PATH",
)
(
PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH,
PROMPT_REWRITE_USER_SYSTEM_PROMPT_FALLBACK_PATH,
) = _resolve_devtools_paths(
_PROMPTS_ROOT / "rewrite_user_system_prompt.md",
_PROMPTS_LOCAL_ROOT / "rewrite_user_system_prompt.md",
"FASTVIDEO_PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH",
)
(
CURATED_PRESETS_FILE_PATH,
CURATED_PRESETS_FALLBACK_FILE_PATH,
) = _resolve_devtools_paths(
_CLIENT_PROMPTS_ROOT / "selected_ltx2_continuation_story_presets.json",
_CLIENT_PROMPTS_LOCAL_ROOT / "selected_ltx2_continuation_story_presets.json",
"FASTVIDEO_CURATED_PRESETS_FILE_PATH",
)
PROMPT_REWRITE_LOG_PATH = os.getenv(
"FASTVIDEO_PROMPT_REWRITE_LOG_PATH",
str(_OUTPUTS_ROOT / "prompt_rewrite.jsonl"),
).strip()
PROMPT_ENHANCE_LOG_PATH = os.getenv(
"FASTVIDEO_PROMPT_ENHANCE_LOG_PATH",
str(_OUTPUTS_ROOT / "prompt_enhance.jsonl"),
).strip()
PROMPT_AUTO_EXTENSION_LOG_PATH = os.getenv(
"FASTVIDEO_PROMPT_AUTO_EXTENSION_LOG_PATH",
str(_OUTPUTS_ROOT / "prompt_auto_extension.jsonl"),
).strip()
SESSION_LOG_ROOT = os.getenv(
"FASTVIDEO_SESSION_LOG_ROOT",
str(_OUTPUTS_ROOT / "session_logs"),
).strip()
# Auto extension behavior.
# Sleep is used as idle backoff when no prompt source is available.
PROMPT_AUTO_SLEEP_MS = _env_int("FASTVIDEO_PROMPT_AUTO_SLEEP_MS", 120)
PROMPT_AUTO_TIMEOUT_MS = _env_int("FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS", 1800)
GENERATION_SEGMENT_CAP = max(0, _env_int("FASTVIDEO_GENERATION_SEGMENT_CAP", 6))
# Startup warmup behavior. Warmup compiles segment 1 and segment 2 inference
# paths before the worker is considered ready for serving.
STARTUP_WARMUP_ENABLED = _env_bool("FASTVIDEO_ENABLE_STARTUP_WARMUP", True)
STARTUP_WARMUP_PROMPT = os.getenv(
"FASTVIDEO_STARTUP_WARMUP_PROMPT",
("A cinematic drone shot over coastal cliffs at sunrise, "
"golden light, gentle ocean waves, ultra detailed"),
).strip()
STARTUP_WARMUP_TIMEOUT_SECONDS = max(1, _env_int("FASTVIDEO_STARTUP_WARMUP_TIMEOUT_SECONDS", 2400))
+961
View File
@@ -0,0 +1,961 @@
# pyright: reportMissingTypeArgument=false, reportArgumentType=false, reportOptionalSubscript=false, reportAttributeAccessIssue=false, reportOptionalMemberAccess=false, reportUndefinedVariable=false
# ruff: noqa: UP038, SIM105, F821
# mypy: ignore-errors
import asyncio
import multiprocessing as mp
import os
import subprocess
import time
import traceback
from dataclasses import dataclass
from enum import Enum
from multiprocessing import Process, Queue
from dreamverse.config import (
DEFAULT_MODEL_ID,
MODEL_REGISTRY,
STARTUP_WARMUP_ENABLED,
STARTUP_WARMUP_PROMPT,
STARTUP_WARMUP_TIMEOUT_SECONDS,
)
from dreamverse.av_streaming import (
SHARED_STREAM_BUFFER_BYTES,
USE_SHARED_STREAM_BUFFER,
StreamChunk,
StreamComplete,
StreamEvent,
StreamInit,
generate_stream_id,
stream_fmp4,
)
from dreamverse.worker_ipc import (
CommandPayload,
InitAck,
JoinAck,
LeaveAck,
MediaChunk,
MediaComplete,
MediaInit,
ReloadAck,
ReloadModelPayload,
ShutdownAck,
StepComplete,
UserStepPayload,
WarmupComplete,
WarmupPayload,
WorkerError,
WorkerEvent,
)
def _parse_requested_gpu_limit() -> int | None:
raw_value = os.getenv("FASTVIDEO_GPU_COUNT", "").strip().lower()
if not raw_value:
return 1
if raw_value == "all":
return None
try:
requested = int(raw_value)
except ValueError as exc:
raise RuntimeError("Invalid FASTVIDEO_GPU_COUNT. Use a positive integer or 'all'.") from exc
if requested <= 0:
raise RuntimeError("Invalid FASTVIDEO_GPU_COUNT. Use a positive integer or 'all'.")
return requested
def _limit_gpu_ids(gpu_ids: list[int]) -> list[int]:
requested_limit = _parse_requested_gpu_limit()
print(f"[INFO] Using GPU limit={requested_limit}")
if requested_limit is None:
return gpu_ids
return gpu_ids[:requested_limit] or gpu_ids
class CommandType(Enum):
"""Commands sent from main process to GPU worker."""
INIT = "init"
WARMUP = "warmup"
SHUTDOWN = "shutdown"
USER_JOIN = "user_join"
USER_STEP = "user_step"
USER_LEAVE = "user_leave"
RELOAD_MODEL = "reload_model"
@dataclass
class Command:
"""Command sent to GPU worker subprocess.
Commands that carry data (USER_STEP, WARMUP, RELOAD_MODEL)
populate ``payload`` with a typed payload from ``worker_ipc``.
Commands that don't (INIT, SHUTDOWN, USER_JOIN, USER_LEAVE)
leave ``payload`` as ``None``.
"""
type: CommandType
payload: CommandPayload | None = None
user_id: str | None = None
def _stream_event_to_worker_event(
event: StreamEvent,
user_id: str,
segment_idx: int,
) -> WorkerEvent:
"""Translate an av_streaming event into a typed worker event.
``StreamEvent`` (ffmpeg output layer) carries no routing info;
this adds ``user_id`` + ``segment_idx`` so the pool's
``_response_reader`` can dispatch to the right per-user queue.
"""
match event:
case StreamInit(stream_id=sid, mime=m, uses_shared_buffer=u):
return MediaInit(
user_id=user_id,
segment_idx=segment_idx,
stream_id=sid,
mime=m,
uses_shared_buffer=u,
)
case StreamChunk(
stream_id=sid,
chunk=c,
chunk_offset=co,
chunk_length=cl,
uses_shared_buffer=u,
):
return MediaChunk(
user_id=user_id,
segment_idx=segment_idx,
stream_id=sid,
chunk=c,
chunk_offset=co,
chunk_length=cl,
uses_shared_buffer=u,
)
case StreamComplete(stream_id=sid, chunks=n):
return MediaComplete(
user_id=user_id,
segment_idx=segment_idx,
stream_id=sid,
chunks=n,
)
case _:
raise ValueError(f"unknown stream event: {type(event).__name__}")
def gpu_worker_process(
gpu_id: int,
cuda_device: str,
command_queue: Queue,
response_queue: Queue,
shared_stream_buffer=None,
shared_stream_buffer_bytes: int = 0,
):
"""Worker process that runs on a single GPU.
CUDA_VISIBLE_DEVICES must be set BEFORE importing VideoGenerationWorker
(which transitively touches CUDA). Delegates model lifecycle and
generation to VideoGenerationWorker; AV muxing to av_streaming.
"""
os.environ["CUDA_VISIBLE_DEVICES"] = cuda_device
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
from dreamverse.video_generation import VideoGenerationWorker
worker = VideoGenerationWorker(gpu_id)
def event_loop(first_cmd: Command = None):
"""Blocking event loop for LTX2; dispatches user commands."""
print(f"[GPU {gpu_id}] Entering event loop")
def handle_command(cmd: Command):
if cmd.type == CommandType.USER_JOIN:
print(f"[GPU {gpu_id}] User {cmd.user_id[:8]} joined")
worker.clear_conditioning()
response_queue.put(JoinAck(user_id=cmd.user_id))
elif cmd.type == CommandType.USER_STEP:
try:
assert isinstance(cmd.payload, UserStepPayload), (f"USER_STEP requires UserStepPayload, "
f"got {type(cmd.payload).__name__}")
payload = cmd.payload
segment_idx = payload.segment_idx
step_result = worker.generate_step(
payload.prompt,
segment_idx,
image_path=payload.image_path,
reset_conditioning=payload.reset_conditioning,
)
head_trim_frames = step_result.head_trim_frames
head_trim_audio_frames = step_result.head_trim_audio_frames
if head_trim_frames > 0 or head_trim_audio_frames > 0:
print(f"[GPU {gpu_id}] Segment {segment_idx}: "
f"trimming video={head_trim_frames} "
f"audio={head_trim_audio_frames} "
f"overlap frames from AV output")
audio_shape = getattr(step_result.audio, "shape", None)
print(f"[GPU {gpu_id}] AV attempt segment "
f"{segment_idx}: "
f"audio_present={step_result.audio is not None}, "
f"audio_shape={audio_shape}, "
f"audio_sample_rate={step_result.audio_sample_rate}")
stream_id = generate_stream_id(segment_idx)
def _publish(event: StreamEvent) -> None:
response_queue.put(_stream_event_to_worker_event(event, cmd.user_id, segment_idx))
av_ok, av_error = stream_fmp4(
frames=step_result.frames,
audio=step_result.audio,
audio_sample_rate=step_result.audio_sample_rate,
stream_id=stream_id,
timings=step_result.timings,
head_trim_frames=head_trim_frames,
head_trim_audio_frames=head_trim_audio_frames,
shared_buffer=shared_stream_buffer,
shared_buffer_bytes=shared_stream_buffer_bytes,
publish=_publish,
log_prefix=f"[GPU {gpu_id}]",
)
if not av_ok:
raise RuntimeError(av_error or "worker av_fmp4 stream failed")
print(f"[GPU {gpu_id}] AV streamed segment {segment_idx}: "
f"encode_total={step_result.timings.get('av_encode_stream_ms', 0):.0f}ms "
f"wav_write={step_result.timings.get('av_wav_write_ms', 0):.1f}ms "
f"spawn={step_result.timings.get('av_ffmpeg_spawn_ms', 0):.1f}ms "
f"first_chunk={step_result.timings.get('av_first_chunk_ms', 0):.0f}ms "
f"chunk_interval_med={step_result.timings.get('av_chunk_interval_ms_median', 0):.1f}ms "
f"chunk_interval_p95={step_result.timings.get('av_chunk_interval_ms_p95', 0):.1f}ms "
f"publish_med={step_result.timings.get('av_chunk_publish_ms_median', 0):.2f}ms "
f"read_med={step_result.timings.get('av_chunk_read_ms_median', 0):.1f}ms")
step_result.timings["ipc_put_start_ns"] = time.time_ns()
response_queue.put(
StepComplete(
user_id=cmd.user_id,
segment_idx=segment_idx,
timings=step_result.timings,
))
except Exception as e:
print(f"[GPU {gpu_id}] Step error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
elif cmd.type == CommandType.USER_LEAVE:
print(f"[GPU {gpu_id}] User {cmd.user_id[:8]} left")
worker.clear_conditioning()
response_queue.put(LeaveAck(user_id=cmd.user_id))
elif cmd.type == CommandType.SHUTDOWN:
return False # Signal to exit
elif cmd.type == CommandType.WARMUP:
try:
assert isinstance(cmd.payload, WarmupPayload), (f"WARMUP requires WarmupPayload, "
f"got {type(cmd.payload).__name__}")
timings = worker.warmup(cmd.payload.prompt)
response_queue.put(WarmupComplete(
user_id=cmd.user_id,
timings=timings,
))
except Exception as e:
print(f"[GPU {gpu_id}] Warmup error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
elif cmd.type == CommandType.RELOAD_MODEL:
try:
assert isinstance(cmd.payload, ReloadModelPayload), (f"RELOAD_MODEL requires ReloadModelPayload, "
f"got {type(cmd.payload).__name__}")
worker.initialize(cmd.payload.model_config)
response_queue.put(ReloadAck(user_id=cmd.user_id))
except Exception as e:
print(f"[GPU {gpu_id}] Reload error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
return True # Continue loop
if first_cmd is not None:
if first_cmd.type == CommandType.SHUTDOWN:
worker.shutdown()
response_queue.put(ShutdownAck())
return
handle_command(first_cmd)
import queue as queue_module
while True:
try:
cmd = command_queue.get(timeout=1.0)
except queue_module.Empty:
continue
if cmd.type == CommandType.SHUTDOWN:
print(f"[GPU {gpu_id}] Event loop shutting down")
worker.shutdown()
response_queue.put(ShutdownAck())
return
if not handle_command(cmd):
return
print(f"[GPU {gpu_id}] Worker process starting...")
try:
while True:
cmd: Command = command_queue.get()
if cmd.type == CommandType.SHUTDOWN:
print(f"[GPU {gpu_id}] Shutting down...")
worker.shutdown()
response_queue.put(ShutdownAck())
break
elif cmd.type == CommandType.INIT:
try:
worker.initialize()
response_queue.put(InitAck(success=True))
except Exception as e:
print(f"[GPU {gpu_id}] Init error: {e}")
traceback.print_exc()
response_queue.put(InitAck(success=False, error=str(e)))
elif cmd.type == CommandType.WARMUP:
try:
assert isinstance(cmd.payload, WarmupPayload), (f"WARMUP requires WarmupPayload, "
f"got {type(cmd.payload).__name__}")
timings = worker.warmup(cmd.payload.prompt)
response_queue.put(WarmupComplete(
user_id=cmd.user_id,
timings=timings,
))
except Exception as e:
print(f"[GPU {gpu_id}] Warmup error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
elif cmd.type == CommandType.RELOAD_MODEL:
try:
assert isinstance(cmd.payload, ReloadModelPayload), (f"RELOAD_MODEL requires ReloadModelPayload, "
f"got {type(cmd.payload).__name__}")
worker.initialize(cmd.payload.model_config)
response_queue.put(ReloadAck(user_id=cmd.user_id))
except Exception as e:
print(f"[GPU {gpu_id}] Reload error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
elif cmd.type in (CommandType.USER_JOIN, CommandType.USER_STEP, CommandType.USER_LEAVE):
event_loop(first_cmd=cmd)
break
except Exception as e:
print(f"[GPU {gpu_id}] Worker crashed: {e}")
traceback.print_exc()
print(f"[GPU {gpu_id}] Worker process exiting")
class GPUSlot:
"""Manages a single GPU worker subprocess."""
def __init__(self, gpu_id: int, cuda_device: str):
self.gpu_id = gpu_id
self.cuda_device = cuda_device
self.process: Process | None = None
self.command_queue: Queue | None = None
self.response_queue: Queue | None = None
self.ready: bool = False
self.warmup_enabled: bool = STARTUP_WARMUP_ENABLED
self.warmup_success: bool = False
self.warmup_error: str | None = None
self.warmup_timings: dict[str, float] = {}
self._lock = asyncio.Lock()
# Client state
self.connected_users: set[str] = set()
self._pending_futures: dict[str, asyncio.Future] = {}
self._stream_queues: dict[str, asyncio.Queue] = {}
self._response_reader_task: asyncio.Task | None = None
self._active: bool = False
self._reader_lock: asyncio.Lock | None = None
self.current_model_id: str = DEFAULT_MODEL_ID
self.shared_stream_buffer = None
self.shared_stream_buffer_size = SHARED_STREAM_BUFFER_BYTES
@property
def client_count(self) -> int:
return len(self.connected_users)
@property
def is_available(self) -> bool:
"""A GPU is available if it has no active users."""
alive = self.ready and self.process is not None and self.process.is_alive()
if not alive:
return False
return len(self.connected_users) == 0
@property
def is_empty(self) -> bool:
return len(self.connected_users) == 0
async def start(self):
"""Start the GPU worker subprocess."""
self.ready = False
self.warmup_success = False
self.warmup_error = None
self.warmup_timings = {}
ctx = mp.get_context("spawn")
self.command_queue = ctx.Queue()
self.response_queue = ctx.Queue()
if USE_SHARED_STREAM_BUFFER and self.shared_stream_buffer_size > 0:
# Fixed shared byte buffer to avoid per-chunk IPC payload copies.
self.shared_stream_buffer = mp.RawArray("B", self.shared_stream_buffer_size)
self.process = ctx.Process(
target=gpu_worker_process,
args=(
self.gpu_id,
self.cuda_device,
self.command_queue,
self.response_queue,
self.shared_stream_buffer,
self.shared_stream_buffer_size,
),
daemon=False,
)
loop = asyncio.get_event_loop()
await loop.run_in_executor(None, self.process.start)
# Send init command and wait for response
init_response = await self._send_command(Command(CommandType.INIT), timeout=600.0)
if not isinstance(init_response, InitAck) or not init_response.success:
error_msg = (init_response.error if isinstance(init_response, InitAck) else
f"unexpected init response: {type(init_response).__name__}")
raise RuntimeError(f"GPU {self.gpu_id} failed to initialize: {error_msg}")
if self.warmup_enabled:
try:
warmup_response = await self._send_command(
Command(
CommandType.WARMUP,
payload=WarmupPayload(prompt=STARTUP_WARMUP_PROMPT),
user_id=f"__warmup_gpu_{self.gpu_id}__",
),
timeout=float(STARTUP_WARMUP_TIMEOUT_SECONDS),
)
except Exception as exc:
self.warmup_error = str(exc)
raise RuntimeError(f"GPU {self.gpu_id} warmup failed: {self.warmup_error}") from exc
match warmup_response:
case WarmupComplete(timings=timings):
self.warmup_timings = {
key: float(value)
for key, value in timings.items() if isinstance(value, (int, float))
}
self.warmup_success = True
case WorkerError(message=msg):
self.warmup_error = msg or "Warmup failed."
raise RuntimeError(f"GPU {self.gpu_id} warmup failed: {self.warmup_error}")
case _:
self.warmup_error = (f"unexpected warmup response: "
f"{type(warmup_response).__name__}")
raise RuntimeError(f"GPU {self.gpu_id} warmup failed: {self.warmup_error}")
else:
print(f"[GPU {self.gpu_id}] Startup warmup disabled by "
"FASTVIDEO_ENABLE_STARTUP_WARMUP")
self.ready = True
async def _send_command(self, cmd: Command, timeout: float = 300.0) -> WorkerEvent:
"""Send a command and wait for the response or worker death.
Whichever arrives first wins. A worker dying (exit, signal, OOM,
segfault) flips the kernel-level sentinel fd readable, which
asyncio notices via add_reader — typically within ~10 ms. The
queue timeout still bounds hangs where the worker stays alive but
never replies.
"""
loop = asyncio.get_running_loop()
process = self.process
await loop.run_in_executor(None, self.command_queue.put, cmd)
response_fut = loop.run_in_executor(None, lambda: self.response_queue.get(timeout=timeout))
death_fut: asyncio.Future | None = None
sentinel_fd = process.sentinel if process is not None else None
if sentinel_fd is not None:
death_fut = loop.create_future()
def _on_death() -> None:
try:
loop.remove_reader(sentinel_fd)
except (ValueError, OSError):
pass
if death_fut is not None and not death_fut.done():
death_fut.set_result(None)
loop.add_reader(sentinel_fd, _on_death)
waiters = [response_fut] + ([death_fut] if death_fut is not None else [])
try:
done, _ = await asyncio.wait(waiters, return_when=asyncio.FIRST_COMPLETED)
finally:
if (sentinel_fd is not None and death_fut is not None and not death_fut.done()):
try:
loop.remove_reader(sentinel_fd)
except (ValueError, OSError):
pass
death_fut.cancel()
if (death_fut is not None and death_fut in done and response_fut not in done):
try:
return self.response_queue.get_nowait()
except Exception:
pass
self.ready = False
pid = process.pid if process is not None else "?"
# Reap the process so .exitcode is populated. Sentinel
# readability means the kernel has already exited the
# process; join is non-blocking in practice.
if process is not None:
try:
process.join(timeout=1)
except Exception:
pass
exitcode = process.exitcode if process is not None else None
raise RuntimeError(f"GPU {self.gpu_id} worker died during command "
f"(pid={pid}, exitcode={exitcode})")
return response_fut.result()
async def _send_command_tagged(self, cmd: Command, timeout: float = 300.0) -> WorkerEvent:
"""Send a tagged command and wait for the matching response.
Uses the response reader background task to route responses.
"""
loop = asyncio.get_event_loop()
future = loop.create_future()
# Register this request's pending future
self._pending_futures[cmd.user_id] = future
# Send command
await loop.run_in_executor(None, self.command_queue.put, cmd)
# Ensure response reader is running
await self._ensure_response_reader()
# Wait for the response
try:
response = await asyncio.wait_for(future, timeout=timeout)
except asyncio.TimeoutError:
if (cmd.user_id in self._pending_futures and self._pending_futures[cmd.user_id] is future):
self._pending_futures.pop(cmd.user_id, None)
raise
return response
async def _ensure_response_reader(self):
"""Start the response reader background task if not running."""
if self._reader_lock is None:
self._reader_lock = asyncio.Lock()
async with self._reader_lock:
if (self._response_reader_task is None or self._response_reader_task.done()):
self._response_reader_task = asyncio.create_task(self._response_reader())
async def _response_reader(self):
"""Background task that reads response queue and routes to per-user futures."""
loop = asyncio.get_event_loop()
while self._active:
try:
def get_response_nonblocking():
try:
return self.response_queue.get(timeout=0.5)
except Exception:
return None
event = await loop.run_in_executor(None, get_response_nonblocking)
if event is None:
continue
# AV streaming events → route to the user's stream queue.
if isinstance(event, (MediaInit, MediaChunk, MediaComplete)):
stream_queue = self._stream_queues.get(event.user_id)
if stream_queue is not None:
await stream_queue.put(event)
else:
print(f"[GPU {self.gpu_id}] Unmatched stream event for user "
f"{event.user_id[:8]}")
continue
# System-level acks shouldn't reach the tagged router; they
# belong to the untagged `_send_command` path.
if isinstance(event, (InitAck, ShutdownAck)):
print(f"[GPU {self.gpu_id}] System event leaked into "
f"tagged reader: {type(event).__name__}")
continue
# Late-mutation of timings for observability. Only events
# that actually carry timings get this annotation.
if isinstance(event, (StepComplete, WarmupComplete)):
event.timings["ipc_get_done_ns"] = time.time_ns()
user_id = event.user_id
if user_id and user_id in self._pending_futures:
future = self._pending_futures.pop(user_id)
if not future.done():
future.set_result(event)
else:
print(f"[GPU {self.gpu_id}] Unmatched response for user "
f"{user_id[:8] if user_id else 'None'}")
except asyncio.CancelledError:
break
except Exception as e:
print(f"[GPU {self.gpu_id}] Response reader error: {e}")
await asyncio.sleep(0.01)
def register_stream_queue(self, user_id: str) -> asyncio.Queue:
"""Register stream-event queue for a specific user."""
queue = asyncio.Queue()
self._stream_queues[user_id] = queue
return queue
def unregister_stream_queue(self, user_id: str) -> None:
"""Remove stream-event queue for a specific user."""
self._stream_queues.pop(user_id, None)
async def join_user(self, user_id: str, model_id: str = None) -> JoinAck:
"""Add a user to this GPU."""
if model_id is None:
model_id = DEFAULT_MODEL_ID
# Reload model if a different one is requested
if model_id != self.current_model_id and model_id in MODEL_REGISTRY:
print(f"[GPU {self.gpu_id}] Model switch: "
f"{self.current_model_id} -> {model_id}")
for uid, future in list(self._pending_futures.items()):
if not future.done():
future.set_exception(RuntimeError("Model changed, session reset"))
self._pending_futures.clear()
self._stream_queues.clear()
self.connected_users.clear()
model_config = MODEL_REGISTRY[model_id]
reload_response = await self._send_command(Command(CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
match reload_response:
case ReloadAck():
pass
case WorkerError(message=msg):
raise RuntimeError(f"Model reload failed: {msg}")
case _:
raise RuntimeError(f"Unexpected reload response: "
f"{type(reload_response).__name__}")
self.current_model_id = model_id
print(f"[GPU {self.gpu_id}] Model reloaded: {model_id}")
self._active = True
self.connected_users.add(user_id)
try:
response = await self._send_command_tagged(Command(CommandType.USER_JOIN, user_id=user_id), timeout=600.0)
match response:
case JoinAck() as ack:
return ack
case WorkerError(message=msg):
self.connected_users.discard(user_id)
raise RuntimeError(f"User join failed for {user_id[:8]}: {msg}")
case _:
self.connected_users.discard(user_id)
raise RuntimeError(f"Unexpected join response for {user_id[:8]}: "
f"{type(response).__name__}")
except Exception:
self.connected_users.discard(user_id)
raise
async def user_step(
self,
user_id: str,
prompt: str,
segment_idx: int = 1,
image_path: str | None = None,
reset_conditioning: bool = False,
) -> dict[str, float]:
"""Execute a generation step for a specific user.
Returns the timings dict. Frames/audio are no longer part of
this return — they stream asynchronously via the AV media
events (MediaInit/MediaChunk/MediaComplete) which the caller
consumes through ``register_stream_queue``.
"""
payload = UserStepPayload(
prompt=prompt,
segment_idx=segment_idx,
image_path=image_path,
reset_conditioning=bool(reset_conditioning),
)
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
timeout=1800.0)
match response:
case StepComplete(timings=timings):
return timings
case WorkerError(message=msg):
raise RuntimeError(f"User step failed for {user_id[:8]}: {msg}")
case _:
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
f"{type(response).__name__}")
async def leave_user(self, user_id: str) -> None:
"""Remove a user from this GPU."""
try:
await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
except Exception as e:
print(f"[GPU {self.gpu_id}] Leave user error: {e}")
finally:
self.connected_users.discard(user_id)
self._pending_futures.pop(user_id, None)
self._stream_queues.pop(user_id, None)
async def shutdown(self):
"""Shutdown the worker subprocess."""
self._active = False
if self._response_reader_task and not self._response_reader_task.done():
self._response_reader_task.cancel()
try:
await self._response_reader_task
except asyncio.CancelledError:
pass
if self.process is not None:
if self.process.is_alive():
try:
await self._send_command(Command(CommandType.SHUTDOWN), timeout=30.0)
except Exception:
pass
self.process.terminate()
self.process.join(timeout=5)
if self.process.is_alive():
self.process.kill()
self.process.join(timeout=5)
else:
# Process already died (sentinel path). Reap it so the
# OS releases the PID.
try:
self.process.join(timeout=1)
except Exception:
pass
for q in (self.command_queue, self.response_queue):
if q is not None:
try:
q.close()
except Exception:
pass
try:
q.join_thread()
except Exception:
pass
class GPUPool:
"""Manages multiple GPU worker subprocesses."""
def __init__(self, gpu_ids: list[int]):
self.gpu_ids = gpu_ids
self.slots: dict[int, GPUSlot] = {gpu_id: GPUSlot(gpu_id, str(gpu_id)) for gpu_id in gpu_ids}
self.waiting_list: list[tuple[str, asyncio.Event, WebSocket]] = []
self.client_gpu_map: dict[str, int] = {}
self._pool_lock = asyncio.Lock()
async def initialize(self):
"""Initialize all GPU workers and wait for them to be ready.
Any per-GPU failure aborts startup so uvicorn refuses to serve
traffic with no functional GPUs.
"""
print(f"Initializing GPU pool with {len(self.gpu_ids)} GPUs: {self.gpu_ids}")
await asyncio.gather(*(self._init_gpu(gpu_id) for gpu_id in self.gpu_ids))
async def _init_gpu(self, gpu_id: int):
"""Initialize a single GPU and assign any waiting clients."""
try:
await self.slots[gpu_id].start()
except Exception as e:
print(f"GPU {gpu_id} failed to initialize: {e}")
await self.slots[gpu_id].shutdown()
raise
print(f"GPU pool: {gpu_id} ready "
f"({sum(1 for s in self.slots.values() if s.ready)}/{len(self.gpu_ids)})")
# Check if anyone is waiting for a GPU
async with self._pool_lock:
slot = self.slots[gpu_id]
if slot.is_available and self.waiting_list:
waiting_client_id, ready_event, _ = self.waiting_list.pop(0)
self.client_gpu_map[waiting_client_id] = gpu_id
print(f"Client {waiting_client_id[:8]} assigned GPU {gpu_id} from queue")
ready_event.set()
await self._send_queue_updates()
async def acquire(self, client_id: str, websocket=None) -> tuple[int, GPUSlot]:
"""Acquire a GPU slot for a client."""
async with self._pool_lock:
for gpu_id, slot in self.slots.items():
if slot.is_available:
self.client_gpu_map[client_id] = gpu_id
print(f"Client {client_id[:8]} acquired GPU {gpu_id}")
return gpu_id, slot
# No slot available, wait in queue
print(f"Client {client_id[:8]} waiting in queue "
f"(all {len(self.gpu_ids)} GPUs at capacity)")
ready_event = asyncio.Event()
async with self._pool_lock:
self.waiting_list.append((client_id, ready_event, websocket))
await self._send_queue_updates()
try:
await ready_event.wait()
except asyncio.CancelledError:
# Client disconnected while queued; remove stale queue entry.
async with self._pool_lock:
self.waiting_list = [
item for item in self.waiting_list if not (item[0] == client_id and item[1] is ready_event)
]
self.client_gpu_map.pop(client_id, None)
await self._send_queue_updates()
raise
gpu_id = self.client_gpu_map.get(client_id)
if gpu_id is None:
raise RuntimeError(f"Client {client_id} was signaled but has no GPU assigned")
return gpu_id, self.slots[gpu_id]
async def release(self, client_id: str):
"""Release a client from its GPU slot."""
async with self._pool_lock:
# Remove stale queue entries if the client disconnected while waiting.
prev_wait_len = len(self.waiting_list)
self.waiting_list = [item for item in self.waiting_list if item[0] != client_id]
removed_from_queue = len(self.waiting_list) != prev_wait_len
gpu_id = self.client_gpu_map.pop(client_id, None)
if gpu_id is None:
if removed_from_queue:
await self._send_queue_updates()
return
slot = self.slots[gpu_id]
print(f"Client {client_id[:8]} released GPU {gpu_id}")
try:
await slot.leave_user(client_id)
except Exception as e:
print(f"[GPU {gpu_id}] Leave user failed: {e}")
# Assign to next waiting client if GPU has capacity
if slot.is_available and self.waiting_list:
waiting_client_id, ready_event, _ = self.waiting_list.pop(0)
self.client_gpu_map[waiting_client_id] = gpu_id
print(f"Client {waiting_client_id[:8]} assigned GPU {gpu_id} from queue")
ready_event.set()
await self._send_queue_updates()
async def _send_queue_updates(self):
"""Send updated queue positions to all waiting clients."""
for i, (cid, _, ws) in enumerate(self.waiting_list):
if ws is not None:
try:
await ws.send_json({
"type": "queue_status",
"position": i + 1,
"total_gpus": len(self.gpu_ids),
"available_gpus": 0,
})
except Exception:
pass
async def shutdown(self):
"""Shutdown all GPU workers."""
print("Shutting down GPU pool...")
tasks = [slot.shutdown() for slot in self.slots.values()]
await asyncio.gather(*tasks, return_exceptions=True)
print("GPU pool shutdown complete")
def get_status(self) -> dict:
"""Get the current status of the GPU pool."""
warmed_count = sum(1 for slot in self.slots.values() if slot.warmup_success)
warmup_failures = sum(1 for slot in self.slots.values()
if slot.warmup_enabled and slot.warmup_error is not None)
return {
"total_gpus": len(self.gpu_ids),
"available_gpus": sum(1 for slot in self.slots.values() if slot.is_available),
"queue_size": len(self.waiting_list),
"warmup_enabled": STARTUP_WARMUP_ENABLED,
"warmup_successful_gpus": warmed_count,
"warmup_failed_gpus": warmup_failures,
"gpu_status": {
gpu_id: {
"ready": slot.ready,
"available": slot.is_available,
"client_count": slot.client_count,
"current_model_id": slot.current_model_id,
"process_alive": (slot.process.is_alive() if slot.process else False),
"warmup_enabled": slot.warmup_enabled,
"warmup_success": slot.warmup_success,
"warmup_error": slot.warmup_error,
"warmup_timings": slot.warmup_timings,
}
for gpu_id, slot in self.slots.items()
}
}
def get_available_gpus() -> list[int]:
"""Get list of available GPU IDs from environment or auto-detect."""
cuda_visible = os.environ.get("CUDA_VISIBLE_DEVICES", "")
if cuda_visible:
visible_gpu_ids = [int(x.strip()) for x in cuda_visible.split(",") if x.strip()]
return _limit_gpu_ids(visible_gpu_ids)
# Auto-detect available GPUs
try:
result = subprocess.run(["nvidia-smi", "--query-gpu=index", "--format=csv,noheader"],
capture_output=True,
text=True)
if result.returncode == 0:
detected_gpu_ids = [int(x.strip()) for x in result.stdout.strip().split("\n") if x.strip()]
print(f"Auto-detected GPU IDs: {detected_gpu_ids}")
return _limit_gpu_ids(detected_gpu_ids)
except Exception:
pass
return _limit_gpu_ids([0])
+132
View File
@@ -0,0 +1,132 @@
# pyright: reportArgumentType=false, reportMissingImports=false
from __future__ import annotations
import logging
import os
from contextlib import asynccontextmanager
from pathlib import Path
from fastapi import FastAPI, WebSocket
from fastapi.middleware.cors import CORSMiddleware
from fastapi.staticfiles import StaticFiles
from fastvideo.entrypoints.streaming import build_health_router
from dreamverse.gpu_pool import GPUPool, get_available_gpus
from dreamverse.session_logger import SessionEventLogger
from dreamverse.config import (
DEVTOOLS_ENABLED,
FRONTEND_STATIC_DIR_CANDIDATES,
PROMPT_SAFETY_ENABLED,
SESSION_LOG_ROOT,
)
from dreamverse.prompt_enhancer import PromptEnhancer
from dreamverse.prompt_safety import PromptSafetyFilter
import dreamverse.runtime as runtime
from dreamverse.routes.health import (
router as internal_monitor_router, )
from dreamverse.routes.presets import (
prompt_config_router,
curated_presets_router,
)
from dreamverse.session.controller import SessionController
class _HeartbeatAccessLogFilter(logging.Filter):
"""Drop noisy access logs for frequent health/readiness probes."""
def filter(self, record: logging.LogRecord) -> bool:
message = record.getMessage()
return ('"GET /healthz ' not in message and '"GET /readyz ' not in message)
def _install_heartbeat_log_filter() -> None:
access_logger = logging.getLogger("uvicorn.access")
for existing in access_logger.filters:
if isinstance(existing, _HeartbeatAccessLogFilter):
return
access_logger.addFilter(_HeartbeatAccessLogFilter())
@asynccontextmanager
async def lifespan(app: FastAPI):
"""Application lifespan manager."""
print("Starting server...")
# Get available GPUs
gpu_ids = get_available_gpus()
print(f"Selected GPU ids: {gpu_ids}")
# Initialize GPU pool (spawns subprocess per GPU)
runtime.gpu_pool = GPUPool(gpu_ids)
await runtime.gpu_pool.initialize()
runtime.prompt_enhancer = PromptEnhancer()
runtime.session_event_logger = SessionEventLogger(Path(SESSION_LOG_ROOT))
runtime.prompt_safety_filter = (PromptSafetyFilter() if PROMPT_SAFETY_ENABLED else None)
if runtime.prompt_safety_filter is not None:
print("Prompt safety filter enabled")
print("Server started")
yield
print("Shutting down server...")
await runtime.gpu_pool.shutdown()
runtime.prompt_safety_filter = None
app = FastAPI(lifespan=lifespan)
app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_credentials=True,
allow_methods=["*"],
allow_headers=["*"],
)
app.include_router(build_health_router(lambda: runtime.gpu_pool))
app.include_router(internal_monitor_router)
app.include_router(prompt_config_router)
if DEVTOOLS_ENABLED:
app.include_router(curated_presets_router)
@app.websocket("/ws")
async def websocket_endpoint(websocket: WebSocket):
controller = SessionController(
ws=websocket,
gpu_pool=runtime.gpu_pool,
prompt_enhancer=runtime.prompt_enhancer,
prompt_safety_filter=runtime.prompt_safety_filter,
session_event_logger=runtime.session_event_logger,
)
await controller.run()
# Serve an exported frontend bundle when present.
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
if os.path.isdir(static_dir):
app.mount("/", StaticFiles(directory=static_dir, html=True), name="static")
break
def cli() -> None:
import argparse
import uvicorn
from dreamverse._deps import require_dreamverse_runtime_deps
require_dreamverse_runtime_deps()
parser = argparse.ArgumentParser()
parser.add_argument("--host", default="0.0.0.0")
parser.add_argument("--port", type=int, default=8009)
args = parser.parse_args()
_install_heartbeat_log_filter()
uvicorn.run(app, host=args.host, port=args.port)
if __name__ == "__main__":
cli()
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+202
View File
@@ -0,0 +1,202 @@
from __future__ import annotations
# mypy: ignore-errors
import importlib
import os
import time
from dataclasses import dataclass
from functools import cache
from pathlib import Path
_SERVER_DIR = Path(__file__).resolve().parent
_REPO_ROOT = _SERVER_DIR.parent
_DEFAULT_CLASSIFIER_DIR = _REPO_ROOT / "classifiers"
CLASSIFIER_DIR = Path(
os.path.expandvars(os.path.expanduser(os.getenv("LTX2_CLASSIFIER_DIR", str(_DEFAULT_CLASSIFIER_DIR)))))
@dataclass(frozen=True)
class BlockedPrompt:
index: int
prompt: str
error: str
def _load_fasttext_module():
try:
return importlib.import_module("fasttext")
except ImportError as exc:
raise RuntimeError("Prompt safety is enabled but the `fasttext` package is not installed.") from exc
def resolve_classifier_path(
classifier_kind: str,
env_var: str,
filename: str,
legacy_path: str,
shared_filename: str,
) -> str:
candidates: list[Path] = []
env_path = os.getenv(env_var)
if env_path:
candidates.append(Path(os.path.expandvars(os.path.expanduser(env_path))))
candidates.extend([
CLASSIFIER_DIR / filename,
Path(f"/home/shared/{shared_filename}"),
Path(legacy_path),
])
for candidate in candidates:
if candidate.is_file():
return str(candidate)
checked = "\n".join(f" - {candidate}" for candidate in candidates)
raise FileNotFoundError(f"Could not find the {classifier_kind} classifier.\n"
f"Checked:\n{checked}\n"
"Set LTX2_CLASSIFIER_DIR or the appropriate classifier path environment "
"variable.")
@cache
def load_fasttext_model(model_path: str):
return _load_fasttext_module().load_model(model_path)
def fasttext_predict(
model_path: str,
text: str,
classifier_name: str,
) -> tuple[str, float]:
model = load_fasttext_model(model_path)
text = text.replace("\n", " ")
start_time = time.perf_counter()
try:
labels, probs = model.predict(text)
except ValueError as error:
if "Unable to avoid copy while creating an array" not in str(error):
raise
predictions = model.f.predict(f"{text}\n", 1, 0.0, "strict")
if not predictions:
raise ValueError("fastText returned no predictions") from error
probs, labels = zip(*predictions, strict=False)
latency_ms = (time.perf_counter() - start_time) * 1000.0
identifier = labels[0].replace("__label__", "")
confidence = float(probs[0])
print("[safety] "
f"{classifier_name} fastText latency={latency_ms:.2f}ms "
f"label={identifier} confidence={confidence:.4f}")
return identifier, confidence
def _normalize_classifier_label(identifier: str) -> str:
return identifier.strip().lower().replace("-", "_").replace(" ", "_")
def _label_matches(
identifier: str,
blocked_markers: tuple[str, ...],
safe_markers: tuple[str, ...],
) -> bool:
normalized = _normalize_classifier_label(identifier)
tokens = tuple(token for token in normalized.split("_") if token)
def has_marker(marker: str) -> bool:
normalized_marker = _normalize_classifier_label(marker)
return (normalized == normalized_marker or normalized_marker in tokens)
if any(has_marker(marker) for marker in safe_markers):
return False
return any(has_marker(marker) for marker in blocked_markers)
class PromptSafetyFilter:
def __init__(
self,
*,
nsfw_classifier=None,
hate_speech_classifier=None,
):
if nsfw_classifier is None or hate_speech_classifier is None:
_load_fasttext_module()
if nsfw_classifier is None:
nsfw_model_path = resolve_classifier_path(
"NSFW",
"LTX2_NSFW_CLASSIFIER_PATH",
"jigsaw_fasttext_bigrams_nsfw_final.bin",
"/data/classifiers/dolma_fasttext_nsfw_jigsaw_model.bin",
"dolma-jigsaw-fasttext-bigrams-nsfw-final.bin",
)
self._nsfw_classifier = lambda text: fasttext_predict(
nsfw_model_path,
text,
"nsfw",
)
else:
self._nsfw_classifier = nsfw_classifier
if hate_speech_classifier is None:
hate_speech_model_path = resolve_classifier_path(
"hate speech",
"LTX2_HATESPEECH_CLASSIFIER_PATH",
"jigsaw_fasttext_bigrams_hatespeech_final.bin",
"/data/classifiers/dolma_fasttext_hatespeech_jigsaw_model.bin",
"dolma-jigsaw-fasttext-bigrams-hatespeech-final.bin",
)
self._hate_speech_classifier = lambda text: fasttext_predict(
hate_speech_model_path,
text,
"hate_speech",
)
else:
self._hate_speech_classifier = hate_speech_classifier
def get_prompt_safety_error(self, prompt: str) -> str | None:
normalized_prompt = prompt.strip()
if not normalized_prompt:
return None
nsfw_label, _ = self._nsfw_classifier(normalized_prompt)
if _label_matches(
nsfw_label,
blocked_markers=("nsfw", ),
safe_markers=("sfw", "safe"),
):
return ("This prompt was flagged as NSFW. "
"You can't generate a video with this specified prompt.")
hate_label, _ = self._hate_speech_classifier(normalized_prompt)
if _label_matches(
hate_label,
blocked_markers=("hatespeech", "hate", "toxic", "offensive", "abusive"),
safe_markers=(
"non_hatespeech",
"not_hatespeech",
"non_toxic",
"not_toxic",
"safe",
"clean",
),
):
return ("This prompt was flagged as hate speech. "
"You can't generate a video with this specified prompt.")
return None
def get_first_blocked_prompt(
self,
prompts: list[str],
) -> BlockedPrompt | None:
for index, prompt in enumerate(prompts):
normalized_prompt = str(prompt or "").strip()
if not normalized_prompt:
continue
error = self.get_prompt_safety_error(normalized_prompt)
if error is not None:
return BlockedPrompt(
index=index,
prompt=normalized_prompt,
error=error,
)
return None
@@ -0,0 +1,59 @@
# Prompt Files
This directory contains the editable system prompts used by the
`PromptEnhancer` in `server/prompt_enhancer.py`.
## `next_segment_system_prompt.md`
Use this prompt for guided continuation.
It is loaded as the "next-segment" system prompt and used by
`PromptEnhancer.enhance_prompt()` for the normal continuation flow when the
request includes:
- locked prior segments
- a new user conditioning prompt describing what should happen next
In that path, the model is asked to write exactly one new segment prompt that
continues from the locked history while satisfying the user's requested next
beat.
This is the prompt used for the live non-single-clip enhancement path in
`server/main.py`.
## `auto_extension_system_prompt.md`
Use this prompt for autonomous prompt expansion.
It is loaded as the "auto-extension" system prompt and used in two different
paths:
1. `PromptEnhancer.generate_auto_prompt()`
This is the background auto-extension flow. The model only receives locked
segment history and must infer the next narrative beat on its own.
1. `PromptEnhancer.enhance_prompt()` in single-clip mode
This is the single-clip expansion flow. A short user idea is expanded into one
standalone detailed 5-second prompt for a single clip.
## `rewrite_user_system_prompt.md`
Use this prompt for new-rollout generation from a user's initial rollout
instruction.
This prompt is used for the rewrite path when there is no existing prompt
window yet and the system needs to generate the initial 6 segment prompts from
scratch.
If this file is missing or empty, the server falls back to
`rewrite_window_system_prompt.md` so startup does not break while the file is
still being drafted.
## Practical difference
- `next_segment_system_prompt.md` is for "continue based on the user's new
instruction."
- `auto_extension_system_prompt.md` is for "expand or infer the next clip
without that guided continuation input."
@@ -0,0 +1,52 @@
You are a prompt extender for LTX-2.3 video generation.
Your job is to expand a short user idea into a detailed cinematic prompt for a single 5-second video clip.
LTX-2.3 works best when prompts clearly describe:
• the subject
• the action
• the environment
• lighting
• camera behavior
• audio
Write the scene as one continuous cinematic shot.
Guidelines:
- Preserve the user’s subject and intent.
- Add concrete visual details (appearance, materials, setting).
- Use cinematic language such as medium shot, close-up, slow push-in, pan, tracking shot.
- Describe motion using clear verbs and visible actions.
- Express emotion through physical cues rather than internal thoughts.
- If dialogue is included, put spoken lines in quotation marks and keep them short.
- Include simple audio when relevant (console beeps, footsteps, rain, room tone).
- End the prompt with a stable visual frame.
Prompt structure (single paragraph, ~4–8 sentences):
1. Shot and subject
2. Environment and lighting
3. Main action
4. Small reaction or follow-up beat
5. Optional camera movement
6. Audio elements
7. Stable ending frame
Avoid:
- scene cuts
- conflicting lighting
- overloaded scenes
- abstract emotional descriptions
Return valid JSON only.
Output exactly one JSON object with this schema:
{"prompt":"<one detailed 5-second video prompt>"}
Rules:
- The top-level JSON object must contain exactly one field: "prompt".
- "prompt" must be a single string.
- Do not return markdown fences.
- Do not return commentary, explanations, or any text before or after the JSON.
- Do not return "next_prompt".
- Do not return "segment_prompts".
- Do not return an array.
@@ -0,0 +1,105 @@
You are a prompt writer for LTX-2 video continuation. You write one new
segment prompt that continues a video from where it left off, guided by
a user's conditioning prompt.
<context>
LTX-2 generates video in 6 sequential segments of 5 seconds each
(30 seconds total). Each segment is generated using the LAST FRAME of
the previous segment as its starting image. The model has no memory of
earlier segments - only that single frame. This means:
- Characters, objects, and settings that are visible in the last frame
carry forward naturally.
- Anything that left the frame (via a scene cut, hard transition, or
camera movement away) cannot be recreated - the model has never seen
it before.
- If a segment ends mid-action or with sudden motion, the next segment
inherits a blurry or unstable starting frame, which degrades quality.
</context>
<task>
You will receive locked segments (already generated or currently
generating) and a conditioning prompt describing what should happen
next. Write exactly one new next-segment prompt continuing naturally
from the last locked segment.
</task>
<rules>
<conditioning>
- Place the user's requested event in this next segment.
- Complete the requested event within this single 5-second segment.
- Keep continuity from locked segments while satisfying the request.
</conditioning>
<writing_style>
Write the segment as a single flowing paragraph in present tense using
active language ("is walking", "reaches for", "speaks softly").
Structure the segment with these layers:
1. Establish the shot using cinematography terms (medium shot, close-up,
wide establishing shot).
2. Set the scene: lighting, color palette, textures, atmosphere.
3. Describe action as a chronological sequence using temporal connectors
("as", "then", "while").
4. Define characters through observable features: age, hairstyle,
clothing, distinguishing marks.
5. Weave in an audio layer alongside the action - specific ambient
sounds ("the hum of fluorescent lights", "a clock ticking on the
wall"), effects, and speech integrated with the visual description.
6. Place all spoken dialogue in quotation marks. Preserve any dialogue
from the user's conditioning prompt exactly as written.
Express emotion through physical cues (clenched fists, trembling lip,
wide eyes) rather than labels ("sad", "angry"). Describe only what is
seen and heard - no smell, taste, or internal thoughts. Use restrained,
natural phrasing. Start directly with scene description.
Keep actions gradual. LTX-2 struggles with sudden, abrupt movements -
they produce artifacts and blurry frames. Any camera movement within
the segment should settle to a still frame by the end - the last moment
should be a stable, static shot.
</writing_style>
<scene_continuity>
Static shots held across multiple consecutive segments can cause visual
artifacts to accumulate in video continuation. Changing the scene can
refresh the image and reset quality.
There are two types of segments:
1. Continuation segments (most segments): The segment continues from
the last frame of the previous segment. Write it as an image-to-video
prompt - describe only what changes from the previous scene. Keep
characters, setting, and framing consistent with what was already
on screen.
2. Cut segments: The segment is an entirely new scene. Write it as a
full text-to-video prompt - fully describe the new setting,
characters, lighting, and atmosphere from scratch, as if the model
has never seen any of it before. Only use a cut when the conditioning
prompt naturally requires a scene change.
For late segments (5-6), use only subtle camera movements (slow zoom,
gentle pan, slight drift) and keep the scene stable for a clean ending.
</scene_continuity>
<dialogue>
When a segment lacks significant action or sound, fill the 5 seconds
with spoken dialogue in quotation marks to keep the scene engaging.
Characters should react to and reference the conditioning event in
their dialogue. Weave dialogue throughout the segment alongside action.
</dialogue>
<length>
Match the length of the locked segments. Count the sentences in the
locked segments and write about the same number here. Typically
3-5 sentences. If the locked segments are short, keep this one short.
</length>
</rules>
<output_format>
Return valid JSON only:
{
"next_prompt": "prompt for the next segment"
}
One key only, no markdown fences, no extra keys.
</output_format>
@@ -0,0 +1,207 @@
You are a real-time prompt writer for ltx2, a video generation model for image-audio-to-video continuation.
<inputs>
You receive:
1. A user prompt describing a video idea, scene, character moment, joke, action, or story beat
</inputs>
<context>
Your job is to expand the user's prompt into a full rollout of 6 sequential segment prompts, each describing 5 seconds of video, for a total of 30 seconds.
Each segment is generated independently but conditioned only on:
- the last 9 video frames of the previous segment
- the last 49 audio frames of the previous segment
Important implications:
- Subjects, props, and scene elements that remain visible in the previous segment's last frame carry forward best.
- If a subject disappears from frame because of a hard transition or because the camera moves away, that subject should not reappear later unless the user explicitly wants a new reveal and that reveal is plausible from the current frame.
- Stable end frames improve continuation.
- The final sentence of each segment should land on a clean, readable visual state.
- Quiet continuing ambience in the final sentence is fine.
- Avoid ending a segment with a brand-new major action, a heavy new line of dialogue, or visual chaos.
</context>
<task>
Return a complete original 6-segment rollout based on the user's prompt.
You must:
- Expand the user's idea into a coherent beginning-to-end 30-second rollout.
- Preserve the user's core concept, tone, style, characters, setting, and requested actions.
- If the user gives only a broad or simple idea, infer missing details conservatively and add clear staging, scene logic, and pacing.
- If the user gives a detailed prompt, follow it closely.
- Keep all 6 segments coherent as one continuous rollout.
- Return all 6 segment prompts.
</task>
<instruction_handling>
Handle the user's prompt in one of these ways:
1. Very broad prompt
- Expand into a clear, staged rollout with defined characters, setting, and progression.
1. Moderately specific prompt
- Preserve given details and fill only what is needed for a complete rollout.
1. Highly detailed prompt
- Follow closely while maintaining clarity, pacing, and continuity.
</instruction_handling>
<priority_rules>
When rules conflict, resolve them in this order:
1. The user's prompt
2. Continuation plausibility from one segment to the next
3. Clarity, staging, and visual quality
4. Default stylistic preferences in this system prompt
</priority_rules>
<house_style>
Default to a dialogue-forward, character-centered rollout with clear staging, one readable beat per segment, and strong visual continuity.
Core defaults:
- One main action + one reaction beat per segment
- Compact, legible segments
- One dominant speaker per segment
- Stable ending frames
Conversation bias:
- When the prompt supports comedy, character interaction, or everyday scenarios, prioritize conversational beats over pure visual spectacle.
- Prefer dialogue-driven progression rather than action-only sequences when both are plausible.
- Use dialogue to reveal character, humor, tension, or situation changes.
- Keep exchanges short and punchy rather than long or dense.
- Let visual acting and timing complement the dialogue instead of replacing it.
Exception:
- If the user explicitly asks for cinematic spectacle, action-heavy sequences, or minimal dialogue, follow that instead.
</house_style>
<default_rollout_rhythm>
Unless the user gives different timing, use this structure:
- Segment 1: establish scene, subject, situation, first speaking beat
- Segment 2: small escalation or reaction
- Segment 3: continuation beat
- Segment 4: pivot, reveal, pan, cut, or new subject focus
- Segment 5: payoff or response
- Segment 6: closing button, stable hold
Segment 4 is the default pivot point unless specified otherwise.
</default_rollout_rhythm>
<segment_length_rules>
Treat counts as soft targets. Do not pad unnaturally.
- Segment 1: 4–5 sentences, ~95–140 words
- Segment 2: 3–4 sentences, ~55–95 words
- Segment 3: 3–4 sentences, ~55–95 words
- Segment 4: 3–5 sentences, ~85–130 words
- Segment 5: 3–4 sentences, ~55–95 words
- Segment 6: 3–4 sentences, ~55–100 words
Adjust if user intent requires.
</segment_length_rules>
<style_rules>
- Respect user-specified style.
- Use a "Style:" prefix only if clearly beneficial or already implied.
- Keep style consistent across segments.
</style_rules>
<description_compression_rules>
- Fully describe scene/subjects on first appearance.
- Compress repeated details in later segments.
- Maintain key identity anchors without redundancy.
- Reintroduce full detail only when scene or subject changes.
</description_compression_rules>
<segment_prompt_rules>
Each segment must:
- Be present tense
- Describe only visible/audible elements
- Be one paragraph
- Match detail to shot scale
- End on a stable visual frame
</segment_prompt_rules>
<scene_rules>
- Maintain coherent setting, layout, lighting
- Use concrete visual details and textures
- Avoid conflicting lighting or environment logic
- Prefer one primary location with optional pivot at segment 4
</scene_rules>
<subject_rules>
- Use consistent naming across segments
- Maintain appearance and identity anchors
- Introduce new characters clearly once
- Prefer small number of recurring subjects
</subject_rules>
<camera_rules>
- Keep camera language efficient
- Prefer stable shots unless movement matters
- Max one clear camera move per segment
- Describe post-movement composition
- Use segment 4 for major camera transitions by default
</camera_rules>
<action_rules>
- One main action + one reaction beat
- Keep motion readable and grounded
- Favor simple, clear gestures
- Avoid chaotic or overloaded motion
</action_rules>
<dialogue_rules>
- Include dialogue in most segments when appropriate (typically 5–6 segments)
- One dominant speaker per segment
- 1 short line or 2 clipped mini-lines
- Keep lines natural and brief
- Place dialogue before final sentence when possible
</dialogue_rules>
<audio_rules>
- Tie audio to visible action/environment
- Use 1–2 concrete sound cues per segment
- Maintain audio continuity across segments
- Avoid introducing new dominant sounds at the end
</audio_rules>
<emotion_rules>
- No internal thoughts
- Use visible cues for emotion
</emotion_rules>
<continuity_rules>
- Ensure smooth visual/audio continuity across segments
- Do not reintroduce off-screen subjects unless justified
- Stabilize new scenes immediately after transitions
- End segments with stable compositions
</continuity_rules>
<text_and_logo_rules>
- Avoid reliance on readable text unless explicitly requested
</text_and_logo_rules>
<id_and_label_rules>
- Generate "id" in snake_case
- Generate concise descriptive "label"
</id_and_label_rules>
<output_format>
Return ONLY valid JSON using exactly this structure:
{
"id": "...",
"label": "...",
"segment_prompts": [
"segment 1 text",
"segment 2 text",
"segment 3 text",
"segment 4 text",
"segment 5 text",
"segment 6 text"
]
}
Do not include explanations, markdown, or additional fields.
Only output the JSON object.
</output_format>
@@ -0,0 +1,255 @@
You are a real-time prompt editor for ltx2, a video generation model for image-audio-to-video continuation.
<inputs>
You receive:
1. An existing rollout JSON with:
- "id"
- "label"
- "segment_prompts": an array of 6 sequential segment prompts
2. The user's latest instruction about how to revise the rollout
</inputs>
<context>
The rollout contains 6 sequential segment prompts, each describing 5 seconds of video, for a total of 30 seconds.
Each segment is generated independently but conditioned only on:
- the last 9 video frames of the previous segment
- the last 49 audio frames of the previous segment
Important implications:
- Subjects, props, and scene elements that remain visible in the previous segment's last frame carry forward best.
- If a subject disappears from frame because of a hard transition or because the camera moves away, that subject should not reappear later unless the user explicitly wants a new reveal and that reveal is plausible from the current frame.
- Stable end frames improve continuation.
- The final sentence of each segment should land on a clean, readable visual state.
- Quiet continuing ambience in the final sentence is fine.
- Avoid ending a segment with a brand-new major action, a heavy new line of dialogue, or visual chaos.
</context>
<task>
Return a complete revised 6-segment rollout.
You must:
- Preserve as much of the existing rollout as possible unless the user asked to change it or it causes logical inconsistency, continuity problems, or common-sense failure.
- Adjust any number of segments as needed so the full 6-segment rollout stays coherent.
- Preserve unaffected details, pacing, camera logic, setting, props, and subject identity unless the user explicitly changes them or they become inconsistent.
- Revise all dependent details when one attribute changes.
- Return all 6 segment prompts, even if only one segment changes.
</task>
<instruction_handling>
Handle the user's latest instruction in one of these ways:
1. No actionable rollout instruction
- If the user's latest instruction does not actually request a change to the rollout, return the existing rollout unchanged.
1. General instruction
- If the user's latest instruction is broad or high-level, expand it into a more detailed, coherent rollout while preserving existing details wherever possible.
1. Detailed instruction
- If the user's latest instruction is specific and detailed, follow it closely while preserving continuity and physical plausibility.
</instruction_handling>
<priority_rules>
When rules conflict, resolve them in this order:
1. The user's latest instruction
2. Continuation plausibility from the previous segment's last visible frame and recent audio
3. Preservation of the existing rollout
4. Default stylistic preferences in this system prompt
</priority_rules>
<house_style>
Default to a dialogue-forward, character-centered rollout with clear staging, one readable beat per segment, and strong visual continuity.
Common defaults:
- Use one main action plus one follow-up reaction beat per segment.
- Keep most segments compact and legible.
- Include dialogue in most segments unless the user explicitly wants a quiet, purely visual, or action-only rollout.
- Usually keep one speaking subject dominant within a segment.
- Favor clean, stable ending images over flashy exits.
</house_style>
<default_rollout_rhythm>
Unless the user gives different timing, use this as the default 6-segment rhythm:
- Segment 1: establish the setting, the main subject, the core situation, and the first speaking beat
- Segment 2: small escalation, reaction, or new piece of information
- Segment 3: continue the situation with one more beat of action or reaction
- Segment 4: visual refresh, reveal, pivot, pan, cut, nearby location change, or new subject focus
- Segment 5: payoff, response, or aftermath in the refreshed composition
- Segment 6: closing button and stable held ending
This is a default pattern, not a hard requirement. If the user specifies different timing, follow the user.
</default_rollout_rhythm>
<segment_length_rules>
Treat sentence and word counts as soft pacing targets, not hard quotas. Do not pad or compress unnaturally just to hit counts.
Default pacing:
- Segment 1: usually 4 to 5 sentences, roughly 95 to 140 words
- Segment 2: usually 3 to 4 sentences, roughly 55 to 95 words
- Segment 3: usually 3 to 4 sentences, roughly 55 to 95 words
- Segment 4: usually 3 to 5 sentences, roughly 85 to 130 words
- Segment 5: usually 3 to 4 sentences, roughly 55 to 95 words
- Segment 6: usually 3 to 4 sentences, roughly 55 to 100 words
If the user requests a slower, denser, faster, quieter, or more cinematic rollout, adjust these ranges as needed while preserving clarity.
</segment_length_rules>
<style_rules>
- Preserve the rollout's existing style-marker pattern whenever possible.
- If the existing rollout consistently starts segments with a style prefix such as "Style: ...", keep that pattern.
- If the existing rollout does not use a style prefix, do not add one unless the user explicitly asks for a style change or the rollout needs a new clear style cue.
- Keep the visual style consistent across all 6 segments unless the user explicitly changes it.
</style_rules>
<description_compression_rules>
- When a scene, subject, or important prop first appears, describe it clearly and specifically.
- In later segments within the same scene, compress repeated details and restate only the key anchors needed for continuity, identity, and image quality.
- Do not fully re-describe the same room, outfit, prop, or character in every segment unless the user explicitly wants that repetition.
- When a new location appears, treat that segment as a fresh introduction for the new location.
- When a new subject appears, describe that subject clearly on first appearance, then use stable shorthand afterward.
</description_compression_rules>
<segment_prompt_rules>
Each segment prompt must:
- Be written in present tense.
- Describe only what is seen and heard.
- Avoid internal thoughts, abstract emotions, or motivations unless shown through visible or audible cues.
- Be a single flowing paragraph.
- Match the amount of detail to the shot scale. Close shots should emphasize facial detail, hands, fabric, texture, and subtle motion. Wide shots should emphasize layout, blocking, and readable movement.
- End in a stable, readable visual state.
</segment_prompt_rules>
<scene_rules>
- Keep the setting, spatial layout, lighting logic, and important props coherent across segments unless the user changes them.
- When useful, include materials and textures such as glossy plastic, worn fabric, tiled floor, brushed metal, wet pavement, fingerprint-textured clay, or soft fur.
- Use concrete visual details that help the model stage the scene cleanly.
- Do not introduce conflicting lighting logic within the same scene.
</scene_rules>
<subject_rules>
- Use the same noun for the same subject across all segments.
- Do not rename the same subject with synonyms in later segments.
- Keep appearance and wardrobe anchors consistent unless the user changes them.
- Include enough appearance detail to preserve identity, such as age, hairstyle, clothing, distinguishing features, body type, species, or surface detail when relevant.
- If a subject speaks or sings, keep voice traits consistent unless the user changes them.
- If there are multiple recurring subjects, distinguish them with stable identifiers.
</subject_rules>
<camera_rules>
- Each segment should imply a clear framing or shot, but camera language should stay efficient.
- Use explicit camera movement only when it matters.
- Most non-pivot segments should use stable framing or gentle motion.
- Prefer at most one deliberate camera move per segment.
- Describe camera movement relative to the subject when useful.
- After a pan, push, pull, tilt, whip-pan, or cut, describe what the camera now lands on.
- If the existing rollout includes explicit camera-transition wording such as:
- "The camera whip-pans fast to the right..."
- "The camera slowly pans across..."
- "The camera pushes in..."
preserve that transition wording exactly and keep it in the same segment and same relative location unless the user explicitly asks to change it.
- Unless the user specifies otherwise, segment 4 is the preferred place for a visual refresh such as a pan, whip-pan, cut, reveal, nearby location shift, or new subject focus.
</camera_rules>
<action_rules>
- Keep motion readable and physically plausible.
- Prefer one main visible action plus one follow-up reaction beat per segment.
- Keep blocking simple and legible.
- Favor small clear gestures such as a head tilt, raised eyebrow, folding arms, looking down, stepping forward, crouching, shifting weight, or setting an object down.
- Avoid overloaded choreography, chaotic physics, or too many simultaneous actions unless the user explicitly wants that complexity.
</action_rules>
<dialogue_rules>
- Dialogue is allowed and often useful.
- For dialogue-forward rollouts, include at least one short spoken beat in most segments, usually 5 or 6 of the 6 segments, unless the user requests silence or a mostly nonverbal sequence.
- Usually keep one speaker dominant within a segment.
- Usually use 1 short quoted line or 2 clipped mini-lines by the same speaker.
- Keep spoken lines short, natural, and easy to act.
- Put all spoken dialogue in quotation marks.
- Prefer visible acting around the line, such as a glance, pause, grin, sigh, shrug, or gesture.
- Avoid long speeches and dense back-and-forth exchanges inside one segment.
- When possible, place dialogue before the final sentence so the segment can land on a stable visual ending.
</dialogue_rules>
<audio_rules>
- Include audio when it helps the scene.
- Tie sound to visible action, speech, movement, or ongoing ambience.
- Usually 1 or 2 specific sound cues is enough for a segment.
- Favor concrete sounds such as hums, clicks, beeps, footsteps, water lapping, crickets, distant chatter, vent hiss, keyboard clacks, birds, sprinkler clicks, or soft music already present in the scene.
- Keep audio scene-appropriate and continuation-safe.
- Quiet ambient audio can continue into the final sentence, but avoid introducing a brand-new dominant sound at the very end.
</audio_rules>
<emotion_rules>
- Do not describe internal thoughts.
- Prefer visible and audible cues over abstract emotional labels.
- Use posture, expression, gaze, timing, breathing, hand movement, and vocal delivery instead of unsupported inner-state narration.
</emotion_rules>
<continuity_rules>
- Later segments must follow naturally from what is plausibly visible and audible from the previous segment's ending.
- Do not reintroduce subjects that are no longer visible after a major transition unless the user explicitly requests it and the reintroduction is plausible.
- If a major scene transition occurs, re-establish the new scene clearly in that same segment so later segments can continue from it.
- Good ending images include a held pose, a settled camera, a quiet look, a character standing still, a character seated calmly, or a clean locked composition.
- Avoid ending on blur, sudden subject exit, unresolved camera motion, or a fresh unresolved event.
</continuity_rules>
<text_and_logo_rules>
- Avoid making readable text, signage, or logos the main point of the scene unless the user explicitly asks for it or the existing rollout already uses it successfully.
- If preserving existing readable text details, keep them short and simple.
</text_and_logo_rules>
<rewrite_biases>
- Prefer the smallest set of edits that fully satisfies the user's latest instruction.
- Preserve the existing rollout's successful pacing, rhythm, and density unless the user explicitly asks to change them.
- Preserve the rollout's existing structural asymmetry when it is already working well, such as a fuller segment 1, a pivot or reveal around segment 4, and shorter compressed later segments.
- If a segment already contains a clear spoken beat and the user did not ask to remove dialogue, preserve dialogue in that segment or replace it with a similarly short spoken beat.
- Preserve which subject is dominant in each segment unless the user explicitly changes the focus.
- Preserve stable framing in non-pivot segments when possible.
- Preserve exact camera-transition wording and keep it in the same segment and same relative location unless the user explicitly asks to change it.
- Preserve shorthand description in later segments when earlier segments already established the scene, subject, and props clearly.
- When a user change affects one segment, first patch that segment and its immediate neighbors before rewriting the whole rollout.
- When one anchor changes, propagate only the downstream changes required for continuity, identity, and common sense.
- Prefer edits that preserve the final held image of each segment whenever possible.
- When the user asks for a stronger result such as funnier, sharper, warmer, or more dramatic, first strengthen dialogue, reaction beats, visible acting, and timing before adding new props, new characters, or larger scene changes.
- Preserve the rollout's existing tonal temperature unless the user explicitly asks to change it.
- Do not add new spectacle, extra dialogue, extra camera movement, or extra scene changes that the user did not request.
</rewrite_biases>
<avoid>
Avoid:
- Fully re-describing the same character or room in every segment
- Internal emotional narration without visible cues
- Conflicting lighting logic
- Overloaded scenes with too many characters or actions
- Unclear subject naming
- Sudden unsupported reappearances
- Long monologues
- Ending on chaos, blur, or a fresh unresolved event
</avoid>
<id_and_label_rules>
- Output an "id" in snake_case.
- Output a short "label" that matches the current concept.
- If editing an existing rollout, preserve the existing id and label unless the user's change makes them inaccurate.
- If they become inaccurate, update them minimally.
</id_and_label_rules>
<output_format>
Return ONLY valid JSON using exactly this structure:
{
"id": "...",
"label": "...",
"segment_prompts": [
"segment 1 text",
"segment 2 text",
"segment 3 text",
"segment 4 text",
"segment 5 text",
"segment 6 text"
]
}
Do not include explanations, markdown, or additional fields.
Only output the JSON object.
</output_format>
@@ -0,0 +1,97 @@
from __future__ import annotations
import json
from typing import Any
REWRITE_REQUEST_TEXT = ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical.")
DEFAULT_REWRITE_SEGMENT_COUNT = 6
REWRITE_MODE_NEW = "new_rollout"
REWRITE_MODE_EDIT_EXISTING = "edit_existing_rollout"
DEFAULT_REWRITE_ROLLOUT_ID = "current_rollout"
DEFAULT_REWRITE_ROLLOUT_LABEL = "Current rollout"
def normalize_prompt_window_prompts(values: list[Any] | None) -> list[str]:
if not isinstance(values, list):
return []
normalized: list[str] = []
for item in values:
if not isinstance(item, str):
continue
clean_item = item.strip()
if not clean_item:
continue
normalized.append(clean_item)
return normalized
def build_rewrite_user_payload(
*,
prompt_window_prompts: list[str],
preset_id: str | None = None,
preset_label: str | None = None,
rewrite_instruction: str | None = None,
) -> dict[str, Any]:
rollout_id = (preset_id or "").strip() or DEFAULT_REWRITE_ROLLOUT_ID
rollout_label = (preset_label or "").strip() or DEFAULT_REWRITE_ROLLOUT_LABEL
instruction = (rewrite_instruction.strip() if isinstance(rewrite_instruction, str) else "")
if len(prompt_window_prompts) == 0:
return {
"mode": REWRITE_MODE_NEW,
"request": REWRITE_REQUEST_TEXT,
"user_instruction": instruction,
"desired_segment_count": DEFAULT_REWRITE_SEGMENT_COUNT,
"rollout_id_hint": rollout_id,
"rollout_label_hint": rollout_label,
}
return {
"mode": REWRITE_MODE_EDIT_EXISTING,
"request": REWRITE_REQUEST_TEXT,
"user_instruction": instruction,
"current_rollout": {
"id": rollout_id,
"label": rollout_label,
"segment_prompts": list(prompt_window_prompts),
},
}
def build_rewrite_request_body(
*,
system_prompt: str,
prompt_window_prompts: list[str],
preset_id: str | None,
preset_label: str | None,
rewrite_instruction: str | None,
model: str,
temperature: float,
max_completion_tokens: int,
) -> dict[str, Any]:
user_payload = build_rewrite_user_payload(
prompt_window_prompts=prompt_window_prompts,
preset_id=preset_id,
preset_label=preset_label,
rewrite_instruction=rewrite_instruction,
)
return {
"model":
model,
"temperature":
temperature,
"max_completion_tokens":
max_completion_tokens,
"response_format": {
"type": "json_object"
},
"messages": [
{
"role": "system",
"content": system_prompt,
},
{
"role": "user",
"content": json.dumps(user_payload, ensure_ascii=False),
},
],
}
@@ -0,0 +1 @@
"""HTTP endpoint handlers grouped by URL-path family."""
@@ -0,0 +1,30 @@
"""Dreamverse-specific monitor routes."""
from __future__ import annotations
from fastapi import APIRouter, HTTPException
import dreamverse.runtime as runtime
from dreamverse.utils import _utc_now_iso
router = APIRouter()
@router.get("/internal/monitor/sessions")
async def get_internal_monitor_sessions():
"""Internal monitor payload for router-level replica session dashboards."""
if runtime.gpu_pool is None:
raise HTTPException(status_code=503, detail="GPU pool not initialized.")
status_payload = runtime.gpu_pool.get_status()
max_available_sessions = status_payload.get("total_gpus")
if not isinstance(max_available_sessions, int) or max_available_sessions < 0:
max_available_sessions = 0
prompt_provider_success_counts: dict[str, int] = {}
if runtime.prompt_enhancer is not None:
prompt_provider_success_counts = runtime.prompt_enhancer.get_provider_success_counts()
return {
"service": "ltx2-streaming-backend",
"pending_sessions": len(runtime.gpu_pool.waiting_list),
"max_available_sessions": max_available_sessions,
"prompt_provider_success_counts": prompt_provider_success_counts,
"ts": _utc_now_iso(),
}
@@ -0,0 +1,205 @@
"""Prompt-system-config and curated-presets HTTP routes.
Exports two routers:
- ``prompt_config_router``: always registered.
- ``curated_presets_router``: registered only when ``DEVTOOLS_ENABLED``.
"""
from __future__ import annotations
# pyright: reportMissingTypeArgument=false
import json
import re
from pathlib import Path
from fastapi import APIRouter, HTTPException
from pydantic import BaseModel
from dreamverse.config import (
CURATED_PRESETS_FILE_PATH,
CURATED_PRESETS_FALLBACK_FILE_PATH,
)
import dreamverse.runtime as runtime
prompt_config_router = APIRouter()
curated_presets_router = APIRouter()
class PromptConfigUpdateRequest(BaseModel):
next_segment_system_prompt: str | None = None
auto_extension_system_prompt: str | None = None
rewrite_window_system_prompt: str | None = None
rewrite_user_system_prompt: str | None = None
rewrite_model: str | None = None
rewrite_temperature: float | None = None
class AppendCuratedPresetRequest(BaseModel):
id: str
label: str
segment_prompts: list[str]
def _sanitize_preset_id(raw: str) -> str:
normalized = re.sub(r"[^a-z0-9]+", "_", (raw or "").strip().lower())
normalized = normalized.strip("_")
return normalized or "custom_editable"
def _load_curated_presets_file(path: Path) -> list[dict]:
if not path.is_file():
return []
try:
with path.open("r", encoding="utf-8") as f:
payload = json.load(f)
except json.JSONDecodeError as exc:
raise RuntimeError(f"Invalid JSON in curated presets file: {path}") from exc
except OSError as exc:
raise RuntimeError(f"Failed to read curated presets file: {path}") from exc
if isinstance(payload, list):
return payload
raise RuntimeError(f"Curated presets file must contain a JSON array: {path}")
def _write_curated_presets_file(path: Path, presets: list[dict]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
try:
with path.open("w", encoding="utf-8") as f:
json.dump(presets, f, ensure_ascii=False, indent=2)
f.write("\n")
except OSError as exc:
raise RuntimeError(f"Failed to write curated presets file: {path}") from exc
def _merge_curated_presets(*preset_groups: list[dict]) -> list[dict]:
merged: list[dict] = []
index_by_id: dict[str, int] = {}
for presets in preset_groups:
for item in presets:
if not isinstance(item, dict):
continue
preset_id = str(item.get("id", "")).strip()
if not preset_id:
continue
normalized_id = preset_id.lower()
normalized_item = dict(item)
if normalized_id in index_by_id:
merged[index_by_id[normalized_id]] = normalized_item
else:
index_by_id[normalized_id] = len(merged)
merged.append(normalized_item)
return merged
@prompt_config_router.get("/prompt-system-config")
async def get_prompt_system_config():
"""Get editable prompt-system configuration."""
if runtime.prompt_enhancer is None:
raise HTTPException(
status_code=503,
detail="Prompt enhancer not initialized",
)
return runtime.prompt_enhancer.get_prompt_config()
@prompt_config_router.post("/prompt-system-config")
async def save_prompt_system_config(payload: PromptConfigUpdateRequest):
"""Save prompt-system configuration to disk and reload runtime prompts."""
if runtime.prompt_enhancer is None:
raise HTTPException(
status_code=503,
detail="Prompt enhancer not initialized",
)
try:
return runtime.prompt_enhancer.save_prompt_config(
next_segment_system_prompt=payload.next_segment_system_prompt,
auto_extension_system_prompt=payload.auto_extension_system_prompt,
rewrite_window_system_prompt=payload.rewrite_window_system_prompt,
rewrite_user_system_prompt=payload.rewrite_user_system_prompt,
rewrite_model=payload.rewrite_model,
rewrite_temperature=payload.rewrite_temperature,
)
except ValueError as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
except RuntimeError as exc:
raise HTTPException(status_code=500, detail=str(exc)) from exc
@curated_presets_router.get("/curated-presets")
async def get_curated_presets():
"""Get curated presets with devtools overlays applied."""
file_path = Path(CURATED_PRESETS_FILE_PATH)
fallback_path = (Path(CURATED_PRESETS_FALLBACK_FILE_PATH) if CURATED_PRESETS_FALLBACK_FILE_PATH else None)
try:
overlay_presets = _load_curated_presets_file(file_path)
fallback_presets = (_load_curated_presets_file(fallback_path) if fallback_path is not None else [])
presets = _merge_curated_presets(
fallback_presets,
overlay_presets,
)
return {
"presets": presets,
"count": len(presets),
"file_path": str(file_path),
"fallback_file_path": (str(fallback_path) if fallback_path is not None else None),
}
except RuntimeError as exc:
raise HTTPException(status_code=500, detail=str(exc)) from exc
@curated_presets_router.post("/curated-presets/append")
async def append_curated_preset(payload: AppendCuratedPresetRequest):
"""Append a curated preset to the configured presets JSON file."""
preset_id = _sanitize_preset_id(payload.id)
label = (payload.label or "").strip()
if not label:
raise HTTPException(status_code=400, detail="label must be non-empty.")
prompts = [prompt.strip() for prompt in payload.segment_prompts if isinstance(prompt, str) and prompt.strip()]
if len(prompts) < 2:
raise HTTPException(
status_code=400,
detail="segment_prompts must contain at least 2 non-empty prompts.",
)
file_path = Path(CURATED_PRESETS_FILE_PATH)
fallback_path = (Path(CURATED_PRESETS_FALLBACK_FILE_PATH) if CURATED_PRESETS_FALLBACK_FILE_PATH else None)
try:
overlay_presets = _load_curated_presets_file(file_path)
fallback_presets = (_load_curated_presets_file(fallback_path) if fallback_path is not None else [])
existing_ids = {
str(item.get("id", "")).strip().lower()
for item in _merge_curated_presets(
fallback_presets,
overlay_presets,
) if isinstance(item, dict)
}
if preset_id.lower() in existing_ids:
raise HTTPException(
status_code=409,
detail=("Preset id already exists in curated presets file: "
f"{preset_id}"),
)
next_entry = {
"id": preset_id,
"label": label,
"segment_prompts": prompts,
}
overlay_presets.append(next_entry)
_write_curated_presets_file(file_path, overlay_presets)
return {
"type": "curated_preset_appended",
"preset": next_entry,
"count": len(_merge_curated_presets(
fallback_presets,
overlay_presets,
)),
"file_path": str(file_path),
"fallback_file_path": (str(fallback_path) if fallback_path is not None else None),
}
except HTTPException:
raise
except RuntimeError as exc:
raise HTTPException(status_code=500, detail=str(exc)) from exc
+22
View File
@@ -0,0 +1,22 @@
"""Runtime service singletons.
Assigned by ``main.lifespan`` at server startup and read by routes and the
session controller via attribute access (``runtime.gpu_pool``). Do NOT import
these names directly (``from runtime import gpu_pool``) — ``from``-import
copies the current binding, freezing it at ``None`` before lifespan runs.
"""
# pyright: reportMissingImports=false
from __future__ import annotations
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from dreamverse.gpu_pool import GPUPool
from dreamverse.session_logger import SessionEventLogger
from dreamverse.prompt_enhancer import PromptEnhancer
from dreamverse.prompt_safety import PromptSafetyFilter
gpu_pool: GPUPool | None = None
prompt_enhancer: PromptEnhancer | None = None
session_event_logger: SessionEventLogger | None = None
prompt_safety_filter: PromptSafetyFilter | None = None
@@ -0,0 +1,24 @@
# pyright: reportMissingTypeArgument=false
from __future__ import annotations
from dreamverse._deps import require_dreamverse_runtime_deps
def cli() -> None:
require_dreamverse_runtime_deps()
try:
from dreamverse.main import cli as main_cli
except ModuleNotFoundError as exc:
if exc.name in {"fastvideo", "torch", "safetensors"}:
raise SystemExit(
"dreamverse-server requires FastVideo runtime deps. "
"Install `fastvideo[dreamverse]` or run `uv sync --extra dreamverse` from the FastVideo checkout."
) from exc
raise
main_cli()
if __name__ == "__main__":
cli()
@@ -0,0 +1,5 @@
"""Per-WebSocket-connection code: state, controller, and helpers.
Nothing in this package outlives a single client session — lifetime matches
one ``websocket.accept()`` to disconnect cycle.
"""
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,21 @@
"""Dataclasses for the prompt-submission pipeline queues."""
from __future__ import annotations
from dataclasses import dataclass
@dataclass
class PromptSubmission:
prompt_id: str
raw_prompt: str
created_at_s: float
@dataclass
class ReadyPrompt:
prompt: str
source: str
prompt_id: str | None = None
fallback_used: bool = False
seed_prompt_index: int | None = None
loop_iteration: int | None = None
@@ -0,0 +1,101 @@
# pyright: reportMissingTypeArgument=false, reportArgumentType=false
from __future__ import annotations
import base64
import io
import re
import shutil
import tempfile
from dataclasses import dataclass
from pathlib import Path
from PIL import Image
MAX_SESSION_INIT_IMAGE_BYTES = 15 * 1024 * 1024
SUPPORTED_SESSION_INIT_IMAGE_MIME_TYPES = {
"image/jpeg",
"image/png",
"image/webp",
}
_DATA_URL_RE = re.compile(r"^data:(?P<mime>[-\w.+/]+);base64,(?P<data>[A-Za-z0-9+/=\s]+)$")
@dataclass(frozen=True)
class SessionInitImage:
file_path: Path
temp_dir: Path
display_name: str
mime_type: str
def _sanitize_display_name(raw_name: object) -> str:
text = str(raw_name or "").strip()
if not text:
return "uploaded-image"
return Path(text).name or "uploaded-image"
def persist_session_init_image(
payload: object,
*,
temp_root: Path | None = None,
) -> SessionInitImage | None:
if payload is None:
return None
if not isinstance(payload, dict):
raise ValueError("initial_image must be an object.")
data_url = str(payload.get("data_url") or "").strip()
if not data_url:
return None
data_url_match = _DATA_URL_RE.match(data_url)
if data_url_match is None:
raise ValueError("initial_image.data_url must be a base64 data URL.")
data_url_mime_type = data_url_match.group("mime").strip().lower()
mime_type = str(payload.get("mime_type") or data_url_mime_type).strip().lower()
if mime_type != data_url_mime_type:
raise ValueError("initial_image.mime_type must match the data URL mime type.")
if mime_type not in SUPPORTED_SESSION_INIT_IMAGE_MIME_TYPES:
raise ValueError("initial_image must be a PNG, JPEG, or WebP image.")
encoded_bytes = re.sub(r"\s+", "", data_url_match.group("data"))
try:
raw_bytes = base64.b64decode(encoded_bytes, validate=True)
except Exception as exc:
raise ValueError("initial_image.data_url is not valid base64 data.") from exc
if len(raw_bytes) == 0:
raise ValueError("initial_image.data_url did not contain image bytes.")
if len(raw_bytes) > MAX_SESSION_INIT_IMAGE_BYTES:
raise ValueError("initial_image must be 15 MB or smaller.")
try:
with Image.open(io.BytesIO(raw_bytes)) as image:
image.load()
normalized = image.convert("RGBA" if "A" in image.getbands() else "RGB")
except Exception as exc:
raise ValueError("initial_image must decode as a valid image.") from exc
temp_dir = Path(
tempfile.mkdtemp(
prefix="ltx2_session_init_",
dir=str(temp_root) if temp_root is not None else None,
))
file_path = temp_dir / "initial_frame.png"
normalized.save(file_path, format="PNG")
return SessionInitImage(
file_path=file_path,
temp_dir=temp_dir,
display_name=_sanitize_display_name(payload.get("name")),
mime_type=mime_type,
)
def cleanup_session_init_image(session_image: SessionInitImage | None) -> None:
if session_image is None:
return
shutil.rmtree(session_image.temp_dir, ignore_errors=True)
@@ -0,0 +1,46 @@
# pyright: reportMissingTypeArgument=false
from __future__ import annotations
import asyncio
from datetime import datetime, timezone
import json
import socket
from pathlib import Path
from typing import Any
def _utc_now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
class SessionEventLogger:
def __init__(self, root_dir: Path):
self.hostname = socket.gethostname()
timestamp = datetime.now(timezone.utc).strftime("%y%m%d_%H%M%S")
self.directory = root_dir / self.hostname
self.path = self.directory / f"{timestamp}.jsonl"
self._lock = asyncio.Lock()
self.directory.mkdir(parents=True, exist_ok=True)
self.path.touch(exist_ok=False)
async def write_event(
self,
*,
event: str,
client_id: str,
payload: dict[str, Any] | None = None,
) -> None:
entry = {
"ts": _utc_now_iso(),
"event": event,
"hostname": self.hostname,
"client_id": client_id,
}
if payload:
entry.update(payload)
async with self._lock:
with self.path.open("a", encoding="utf-8") as fp:
fp.write(json.dumps(entry, ensure_ascii=False) + "\n")
@@ -0,0 +1,15 @@
from __future__ import annotations
import sys
from pathlib import Path
TESTS_DIR = Path(__file__).resolve().parent
DREAMVERSE_PACKAGE_DIR = TESTS_DIR.parent
DREAMVERSE_APP_DIR = DREAMVERSE_PACKAGE_DIR.parent
BENCHMARKS_DIR = DREAMVERSE_PACKAGE_DIR / "benchmarks"
for path in (DREAMVERSE_APP_DIR, BENCHMARKS_DIR):
path_str = str(path)
if path_str not in sys.path:
sys.path.insert(0, path_str)
@@ -0,0 +1,155 @@
from __future__ import annotations
import importlib.util
from pathlib import Path
import pytest
SERVER_DIR = Path(__file__).resolve().parents[1]
def _load_config_module():
spec = importlib.util.spec_from_file_location(
"server_config_test_module",
SERVER_DIR / "config.py",
)
assert spec is not None
assert spec.loader is not None
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def _set_required_prompt_keys(monkeypatch):
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
def test_config_allows_missing_cerebras_api_key_until_prompt_runtime(monkeypatch):
monkeypatch.delenv("FASTVIDEO_PROMPT_PROVIDER", raising=False)
monkeypatch.delenv("CEREBRAS_API_KEY", raising=False)
monkeypatch.delenv("GROQ_API_KEY", raising=False)
module = _load_config_module()
assert module.PROMPT_API_KEYS["cerebras"] is None
def test_config_allows_missing_groq_api_key_until_prompt_runtime(monkeypatch):
monkeypatch.delenv("FASTVIDEO_PROMPT_PROVIDER", raising=False)
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
monkeypatch.delenv("GROQ_API_KEY", raising=False)
module = _load_config_module()
assert module.PROMPT_API_KEYS["groq"] is None
def test_config_defaults_to_cerebras_with_parallel_groq_fallback_stage(monkeypatch):
monkeypatch.delenv("FASTVIDEO_PROMPT_PROVIDER", raising=False)
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
)
assert module.PROMPT_API_KEY == "cerebras-key"
assert module.PROMPT_API_BASE_URL is None
assert module.PROMPT_API_KEYS == {
"cerebras": "cerebras-key",
"groq": "groq-key",
}
assert module.PROMPT_API_BASE_URLS == {
"cerebras": None,
"groq": "https://api.groq.com/openai/v1",
}
assert module.PROMPT_MODEL == "gpt-oss-120b"
assert module.PROMPT_REWRITE_MODEL == "gpt-oss-120b"
assert module.PROMPT_REWRITE_MODEL_OPTIONS == ["gpt-oss-120b"]
assert module.PROMPT_PROVIDER_MODELS == {
"cerebras": "gpt-oss-120b",
"groq": "openai/gpt-oss-120b",
}
def test_config_ignores_legacy_groq_primary_override(monkeypatch):
monkeypatch.setenv("FASTVIDEO_PROMPT_PROVIDER", "groq")
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
)
assert module.PROMPT_API_KEY == "cerebras-key"
assert module.PROMPT_API_BASE_URL is None
def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_path):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("FASTVIDEO_ENABLE_DEVTOOLS", "true")
monkeypatch.setenv("FASTVIDEO_DREAMVERSE_HOME", str(tmp_path / "dreamverse-state"))
module = _load_config_module()
assert module.DEVTOOLS_ENABLED is True
assert module.FRONTEND_ROOT.as_posix().endswith("apps/dreamverse/web")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/next_segment_system_prompt.md"
)
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/next_segment_system_prompt.md"
)
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/rewrite_user_system_prompt.md"
)
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/rewrite_user_system_prompt.md"
)
assert module.CURATED_PRESETS_FILE_PATH.endswith(
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json"
)
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
)
assert module.FRONTEND_STATIC_DIR_CANDIDATES == (
str(module.FRONTEND_ROOT / "out"),
str(module.FRONTEND_ROOT / "dist"),
)
def test_config_enables_prompt_safety_when_requested(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("FASTVIDEO_ENABLE_PROMPT_SAFETY", "true")
module = _load_config_module()
assert module.PROMPT_SAFETY_ENABLED is True
def test_config_uses_five_minute_session_timeout(monkeypatch):
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 300
def test_config_rejects_invalid_prompt_provider(monkeypatch):
monkeypatch.setenv("FASTVIDEO_PROMPT_PROVIDER", "unsupported")
_set_required_prompt_keys(monkeypatch)
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_PROMPT_PROVIDER"):
_load_config_module()
@@ -0,0 +1,173 @@
from __future__ import annotations
# pyright: reportAttributeAccessIssue=false, reportMissingImports=false
import sys
import types
from fastapi import APIRouter
from fastapi.testclient import TestClient
import fastvideo.entrypoints.streaming as streaming_entrypoints
import pytest
def _install_stack03_import_stubs(monkeypatch):
"""Keep entrypoint tests focused while later-stack runtime modules are absent."""
if not hasattr(streaming_entrypoints, "build_health_router"):
monkeypatch.setattr(streaming_entrypoints, "build_health_router", lambda _pool=None: APIRouter(), raising=False)
gpu_pool_stub = types.ModuleType("dreamverse.gpu_pool")
class GPUPool:
def __init__(self, _gpu_ids):
pass
async def initialize(self):
pass
async def shutdown(self):
pass
gpu_pool_stub.GPUPool = GPUPool
gpu_pool_stub.get_available_gpus = lambda: []
monkeypatch.setitem(sys.modules, "dreamverse.gpu_pool", gpu_pool_stub)
session_logger_stub = types.ModuleType("dreamverse.session_logger")
session_logger_stub.SessionEventLogger = lambda _path: object()
monkeypatch.setitem(sys.modules, "dreamverse.session_logger", session_logger_stub)
prompt_enhancer_stub = types.ModuleType("dreamverse.prompt_enhancer")
prompt_enhancer_stub.PromptEnhancer = lambda: object()
monkeypatch.setitem(sys.modules, "dreamverse.prompt_enhancer", prompt_enhancer_stub)
prompt_safety_stub = types.ModuleType("dreamverse.prompt_safety")
prompt_safety_stub.PromptSafetyFilter = lambda: object()
monkeypatch.setitem(sys.modules, "dreamverse.prompt_safety", prompt_safety_stub)
session_package_stub = types.ModuleType("dreamverse.session")
session_package_stub.__path__ = []
monkeypatch.setitem(sys.modules, "dreamverse.session", session_package_stub)
controller_stub = types.ModuleType("dreamverse.session.controller")
class SessionController:
def __init__(self, **_kwargs):
pass
async def run(self):
pass
controller_stub.SessionController = SessionController
monkeypatch.setitem(sys.modules, "dreamverse.session.controller", controller_stub)
def _import_server_main(monkeypatch):
_install_stack03_import_stubs(monkeypatch)
sys.modules.pop("dreamverse.main", None)
import dreamverse.main as server_main
return server_main
def _import_mock_server_or_skip():
return pytest.importorskip("dreamverse.mock_server")
def _run_cli(module, monkeypatch, argv: list[str]) -> list[dict[str, object]]:
calls: list[dict[str, object]] = []
uvicorn_stub = types.ModuleType("uvicorn")
def run(app, host: str, port: int) -> None:
calls.append(
{
"app": app,
"host": host,
"port": port,
}
)
uvicorn_stub.run = run
monkeypatch.setitem(sys.modules, "uvicorn", uvicorn_stub)
monkeypatch.setattr("dreamverse._deps.require_dreamverse_runtime_deps", lambda: None)
if hasattr(module, "require_dreamverse_runtime_deps"):
monkeypatch.setattr(module, "require_dreamverse_runtime_deps", lambda: None)
monkeypatch.setattr(sys, "argv", argv)
module.cli()
return calls
def test_server_cli_defaults_to_local_web_port(monkeypatch):
server_main = _import_server_main(monkeypatch)
calls = _run_cli(server_main, monkeypatch, ["dreamverse-server"])
assert calls == [
{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}
]
def test_server_cli_allows_explicit_host_and_port(monkeypatch):
server_main = _import_server_main(monkeypatch)
calls = _run_cli(
server_main,
monkeypatch,
["dreamverse-server", "--host", "127.0.0.1", "--port", "8123"],
)
assert calls == [
{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}
]
def test_server_does_not_expose_backend_source_as_static_assets(monkeypatch):
server_main = _import_server_main(monkeypatch)
client = TestClient(server_main.app)
response = client.get("/server-assets/main.py")
assert response.status_code == 404
def test_mock_server_cli_defaults_to_local_web_port(monkeypatch):
mock_server = _import_mock_server_or_skip()
calls = _run_cli(
mock_server,
monkeypatch,
["dreamverse-mock-server"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}
]
def test_mock_server_cli_updates_latency(monkeypatch):
mock_server = _import_mock_server_or_skip()
old_latency_ms = mock_server.LATENCY_MS
try:
calls = _run_cli(
mock_server,
monkeypatch,
["dreamverse-mock-server", "--latency", "321", "--port", "8111"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}
]
assert mock_server.LATENCY_MS == 321
finally:
mock_server.LATENCY_MS = old_latency_ms
@@ -0,0 +1,121 @@
# pyright: reportAttributeAccessIssue=false
from __future__ import annotations
import asyncio
import multiprocessing as mp
from types import SimpleNamespace
import pytest
import dreamverse.gpu_pool as gpu_pool
def _child_consume_and_exit(cmd_q, resp_q):
"""Top-level so the spawn context can pickle it.
Signals startup by putting "READY" on resp_q (so the parent can
wait out spawn-import latency separately from the actual test
assertion), then consumes one command from cmd_q and exits without
putting anything else on resp_q. Simulates a worker that dies
mid-command — e.g. SIGQUIT'd after a fatal pipeline error.
"""
resp_q.put("READY")
cmd_q.get()
def test_get_available_gpus_defaults_to_single_detected_gpu(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
def fake_run(*args, **kwargs):
del args, kwargs
return SimpleNamespace(returncode=0, stdout="0\n1\n2\n")
monkeypatch.setattr(gpu_pool.subprocess, "run", fake_run)
assert gpu_pool.get_available_gpus() == [0]
def test_get_available_gpus_respects_explicit_gpu_count(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "2")
def fake_run(*args, **kwargs):
del args, kwargs
return SimpleNamespace(returncode=0, stdout="0\n1\n2\n")
monkeypatch.setattr(gpu_pool.subprocess, "run", fake_run)
assert gpu_pool.get_available_gpus() == [0, 1]
def test_get_available_gpus_can_use_all_visible_devices(monkeypatch):
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "3,5,7")
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "all")
assert gpu_pool.get_available_gpus() == [3, 5, 7]
def test_get_available_gpus_defaults_to_first_visible_device(monkeypatch):
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "3,5,7")
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
assert gpu_pool.get_available_gpus() == [3]
def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "zero")
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_GPU_COUNT"):
gpu_pool.get_available_gpus()
def test_send_command_raises_on_worker_death():
"""A worker that consumes a command and exits without replying must
surface as RuntimeError via sentinel detection, not after the long
queue timeout.
The child signals startup with a "READY" message; we drain that
first so the wait_for budget bounds only the sentinel-detection
time, not spawn-import latency.
"""
ctx = mp.get_context("spawn")
cmd_q = ctx.Queue()
resp_q = ctx.Queue()
proc = ctx.Process(
target=_child_consume_and_exit, args=(cmd_q, resp_q)
)
proc.start()
# Wait for the spawn child to fully boot. Allow generous time —
# this isn't what we're measuring.
ready = resp_q.get(timeout=30.0)
assert ready == "READY"
async def runner():
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.process = proc
slot.command_queue = cmd_q
slot.response_queue = resp_q
# Now the child is blocked in cmd_q.get(). Sending the cmd
# makes it exit ~immediately; sentinel must fire well before
# the queue timeout.
await asyncio.wait_for(
slot._send_command(
gpu_pool.Command(gpu_pool.CommandType.SHUTDOWN),
timeout=10.0,
),
timeout=5.0,
)
try:
with pytest.raises(RuntimeError, match="worker died"):
asyncio.run(runner())
finally:
proc.join(timeout=5)
cmd_q.close()
resp_q.close()
@@ -0,0 +1,56 @@
import ast
from pathlib import Path
ALLOWED_PREFIXES = (
"fastvideo.api",
"fastvideo.entrypoints.streaming",
"fastvideo.entrypoints.video_generator",
"fastvideo.configs",
)
ALLOWED_EXACT = ("fastvideo",)
FORBIDDEN_PREFIXES = (
"fastvideo.pipelines",
"fastvideo.models",
"fastvideo.layers",
"fastvideo.worker",
"fastvideo.fastvideo_args",
)
ALLOWED_INTERNAL_IMPORTS = {
(
"video_generation.py",
"fastvideo.models.audio.ltx2_audio_processing",
),
(
"video_generation.py",
"fastvideo.models.loader.component_loader",
),
}
def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
root = Path(__file__).resolve().parents[1]
bad: list[tuple[str, int, str]] = []
for path in root.rglob("*.py"):
if "/tests/" in path.as_posix():
continue
try:
tree = ast.parse(path.read_text(), filename=str(path))
except SyntaxError as task_exc:
raise AssertionError(f"Failed to parse {path}") from task_exc
for node in ast.walk(tree):
names = (
[a.name for a in node.names] if isinstance(node, ast.Import)
else [node.module] if isinstance(node, ast.ImportFrom) and node.module
else []
)
for name in names:
if not name:
continue
rel_path = str(path.relative_to(root))
if (
name.startswith(FORBIDDEN_PREFIXES)
and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS
):
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
assert bad == [], f"Forbidden internal imports: {bad}"
@@ -0,0 +1,355 @@
# pyright: reportArgumentType=false
from __future__ import annotations
import asyncio
import os
from fastapi import WebSocketDisconnect
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
import dreamverse.mock_server as mock_server
class _FakeWebSocket:
def __init__(self, messages: list[tuple[float, dict[str, object]]]):
self._messages = messages
self._index = 0
self.sent_json: list[dict[str, object]] = []
self.sent_bytes: list[bytes] = []
async def accept(self) -> None:
return None
async def send_json(self, payload: dict[str, object]) -> None:
self.sent_json.append(payload)
async def send_bytes(self, payload: bytes) -> None:
self.sent_bytes.append(payload)
async def close(self, code: int = 1000, reason: str | None = None) -> None:
del code, reason
async def receive_json(self) -> dict[str, object]:
if self._index >= len(self._messages):
raise WebSocketDisconnect()
delay_s, payload = self._messages[self._index]
self._index += 1
if delay_s > 0:
await asyncio.sleep(delay_s)
return payload
def test_mock_server_matches_current_single5s_protocol():
old_segment_bytes = mock_server.MOCK_SEGMENT_BYTES
old_latency_ms = mock_server.LATENCY_MS
try:
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {"type": "leave"}),
]
)
asyncio.run(mock_server.websocket_endpoint(ws))
message_types = [payload["type"] for payload in ws.sent_json]
assert message_types[0] == "queue_status"
assert "gpu_assigned" in message_types
assert message_types.count("ltx2_stream_start") == 2
assert "prompt_received" in message_types
assert "prompt_enhancing" in message_types
assert "prompt_ready" in message_types
assert "media_init" in message_types
assert "media_segment_complete" in message_types
assert message_types.count("ltx2_stream_complete") == 2
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
gpu_assigned_event = next(
payload for payload in ws.sent_json if payload["type"] == "gpu_assigned"
)
assert gpu_assigned_event["session_timeout"] == mock_server.SESSION_TIMEOUT_SECONDS
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "selected prompt"
assert segment_start_events[1]["prompt"] == "custom prompt"
step_complete_events = [
payload
for payload in ws.sent_json
if payload["type"] == "step_complete"
]
assert len(step_complete_events) == 2
assert step_complete_events[0]["latency_ms"] == {
"total": 121.0,
"worker_e2e": 1.0,
"main_user_step": 121.0,
"overhead": 120.0,
}
assert ws.sent_bytes
assert ws.sent_bytes[0] == b"mock-fmp4-bytes"
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
def test_mock_server_regular_cap_waits_for_rewrite_rollout():
old_segment_bytes = mock_server.MOCK_SEGMENT_BYTES
old_latency_ms = mock_server.LATENCY_MS
old_generation_segment_cap = mock_server.GENERATION_SEGMENT_CAP
try:
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
mock_server.GENERATION_SEGMENT_CAP = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {"type": "leave"}),
]
)
asyncio.run(mock_server.websocket_endpoint(ws))
message_types = [payload["type"] for payload in ws.sent_json]
assert message_types.count("ltx2_stream_start") == 2
assert message_types.count("ltx2_stream_complete") == 2
assert "generation_cap_reached" not in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "segment one"
assert segment_start_events[1]["prompt"] == "segment one [start a new rollout]"
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
mock_server.GENERATION_SEGMENT_CAP = old_generation_segment_cap
def test_mock_server_rewrite_during_active_segment_restarts_from_first_rewritten_prompt():
old_segment_bytes = mock_server.MOCK_SEGMENT_BYTES
old_latency_ms = mock_server.LATENCY_MS
try:
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 100
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {"type": "leave"}),
]
)
asyncio.run(mock_server.websocket_endpoint(ws))
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment one [restart from rewrite]",
]
assert all(
payload["prompt"] != "segment two"
for payload in segment_start_events[1:]
)
reset_events = [
payload
for payload in ws.sent_json
if payload.get("type") == "seed_prompts_reset_applied"
]
assert any(
payload.get("reason") == "rewrite_during_generation"
for payload in reset_events
)
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
def test_mock_server_supports_initial_custom_rollout_prompt():
old_segment_bytes = mock_server.MOCK_SEGMENT_BYTES
old_latency_ms = mock_server.LATENCY_MS
try:
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {"type": "leave"}),
]
)
asyncio.run(mock_server.websocket_endpoint(ws))
message_types = [payload["type"] for payload in ws.sent_json]
assert "rewrite_seed_prompts_started" in message_types
assert "seed_prompts_updated" in message_types
assert "rewrite_seed_prompts_complete" in message_types
assert "seed_prompts_reset_applied" in message_types
assert "ltx2_stream_start" in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
assert segment_start_events
assert segment_start_events[0]["prompt"] == (
"A moonbase corridor thriller with flooding [segment 1]"
)
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
def test_mock_server_can_start_new_project_without_reconnecting():
old_segment_bytes = mock_server.MOCK_SEGMENT_BYTES
old_latency_ms = mock_server.LATENCY_MS
try:
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 40
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {"type": "end_project_keep_session"}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {"type": "leave"}),
]
)
asyncio.run(mock_server.websocket_endpoint(ws))
message_types = [payload["type"] for payload in ws.sent_json]
assert message_types.count("gpu_assigned") == 1
assert message_types.count("ltx2_stream_start") == 2
assert "project_idle" in message_types
project_idle_index = message_types.index("project_idle")
stream_start_indexes = [
index for index, message_type in enumerate(message_types)
if message_type == "ltx2_stream_start"
]
assert stream_start_indexes[0] < project_idle_index < stream_start_indexes[1]
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment two",
]
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,951 @@
"""Multiprocess realtime stress test for LTX2 streaming websocket service."""
# pyright: reportArgumentType=false, reportOptionalMemberAccess=false
from __future__ import annotations
import argparse
import asyncio
from collections import defaultdict
from datetime import datetime, timezone
import json
import math
import multiprocessing as mp
import os
from pathlib import Path
import sys
import time
import traceback
from typing import Any
import uuid
import pytest
pytestmark = pytest.mark.gpu
try:
import websockets
except ModuleNotFoundError:
websockets = None # type: ignore[assignment]
DEFAULT_PRESET_FILE = (
Path(__file__).resolve().parents[2]
/ "web"
/ "prompts"
/ "selected_ltx2_continuation_story_presets.json"
)
def utc_now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def iso_from_epoch(epoch_s: float) -> str:
return datetime.fromtimestamp(epoch_s, tz=timezone.utc).isoformat()
def iso_to_epoch(iso_ts: str | None) -> float | None:
if not iso_ts:
return None
try:
return datetime.fromisoformat(iso_ts).timestamp()
except ValueError:
return None
def safe_percentile(values: list[float], percentile: float) -> float | None:
if not values:
return None
sorted_values = sorted(values)
if len(sorted_values) == 1:
return sorted_values[0]
rank = (len(sorted_values) - 1) * (percentile / 100.0)
lower = math.floor(rank)
upper = math.ceil(rank)
if lower == upper:
return sorted_values[lower]
fraction = rank - lower
return (
sorted_values[lower]
+ (sorted_values[upper] - sorted_values[lower]) * fraction
)
def summarize_series(values: list[float]) -> dict[str, float | int | None]:
if not values:
return {
"count": 0,
"min": None,
"p50": None,
"p95": None,
"p99": None,
"max": None,
"avg": None,
}
return {
"count": len(values),
"min": min(values),
"p50": safe_percentile(values, 50),
"p95": safe_percentile(values, 95),
"p99": safe_percentile(values, 99),
"max": max(values),
"avg": sum(values) / len(values),
}
def parse_int(value: object) -> int | None:
try:
parsed = int(value)
except (TypeError, ValueError):
return None
return parsed
def format_num(value: float | int | None, digits: int = 2) -> str:
if value is None:
return "n/a"
if isinstance(value, int):
return str(value)
return f"{value:.{digits}f}"
def load_curated_prompts(
preset_file: Path,
preset_id: str | None,
curated_limit: int,
) -> tuple[str, list[str], int]:
if not preset_file.is_file():
raise ValueError(f"Preset file not found: {preset_file}")
try:
payload = json.loads(preset_file.read_text(encoding="utf-8"))
except json.JSONDecodeError as exc:
raise ValueError(f"Invalid JSON in preset file: {preset_file}") from exc
if not isinstance(payload, list) or len(payload) == 0:
raise ValueError("Preset file must contain a non-empty JSON array.")
selected: dict[str, Any] | None = None
if preset_id:
for item in payload:
if not isinstance(item, dict):
continue
if str(item.get("id", "")).strip() == preset_id:
selected = item
break
if selected is None:
raise ValueError(f"Preset id not found: {preset_id}")
else:
for item in payload:
if isinstance(item, dict):
selected = item
break
if selected is None:
raise ValueError("No valid preset object found in preset file.")
selected_id = str(selected.get("id", "")).strip() or "unknown_preset"
raw_prompts = selected.get("segment_prompts", [])
if not isinstance(raw_prompts, list):
raise ValueError(
f"Preset {selected_id} has invalid segment_prompts (must be list)."
)
prompts = [
str(prompt).strip()
for prompt in raw_prompts
if isinstance(prompt, str) and str(prompt).strip()
]
if not prompts:
raise ValueError(f"Preset {selected_id} has no non-empty prompts.")
limited = prompts[:curated_limit]
if not limited:
raise ValueError(
f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}."
)
return selected_id, limited, len(prompts)
async def run_single_session(
*,
worker_id: int,
worker_session_idx: int,
config: dict[str, Any],
) -> dict[str, Any]:
session_id = f"w{worker_id}_u{worker_session_idx}_{uuid.uuid4().hex[:8]}"
process_id = os.getpid()
url = str(config["url"])
session_timeout_s = float(config["session_timeout_s"])
post_complete_wait_s = float(config["post_complete_wait_s"])
connect_timeout_s = float(config["connect_timeout_s"])
curated_prompts = list(config["curated_prompts"])
preset_id = str(config["preset_id"])
session_start_monotonic = time.monotonic()
session_start_epoch = time.time()
session_data: dict[str, Any] = {
"session_id": session_id,
"process_id": process_id,
"worker_id": worker_id,
"status": "failed",
"error": None,
"preset_id": preset_id,
"curated_prompt_count": len(curated_prompts),
"connect_start_ts_utc": iso_from_epoch(session_start_epoch),
"connect_finish_ts_utc": None,
"session_init_sent_ts_utc": None,
"gpu_assigned_ts_utc": None,
"target_segment_complete_ts_utc": None,
"leave_sent_ts_utc": None,
"close_ts_utc": None,
"duration_ms": None,
"queue_wait_ms": None,
"initial_total_segments": None,
"segments_started": 0,
"segments_completed": 0,
"media_segments_completed": 0,
"total_chunks": 0,
"total_chunk_bytes": 0,
"first_chunk_finish_ts_utc": None,
"last_chunk_finish_ts_utc": None,
"first_media_segment_complete_ts_utc": None,
"first_chunk_before_first_media_complete": None,
"session_goodput_mbps": None,
"chunks": [],
}
connect_finish_monotonic: float | None = None
current_segment_idx: int | None = None
initial_total_segments: int | None = None
first_chunk_finish_epoch: float | None = None
first_media_segment_complete_epoch: float | None = None
last_chunk_finish_epoch: float | None = None
last_chunk_finish_monotonic: float | None = None
try:
async with websockets.connect(
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
) as ws:
connect_finish_monotonic = time.monotonic()
session_data["connect_finish_ts_utc"] = utc_now_iso()
init_payload = {
"type": "session_init_v2",
"preset_id": preset_id,
"curated_prompts": curated_prompts,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
}
await ws.send(json.dumps(init_payload))
session_data["session_init_sent_ts_utc"] = utc_now_iso()
while True:
elapsed_s = time.monotonic() - session_start_monotonic
timeout_remaining = session_timeout_s - elapsed_s
if timeout_remaining <= 0:
session_data["status"] = "timeout"
session_data["error"] = (
f"Session timed out after {session_timeout_s:.1f}s."
)
break
recv_start_epoch = time.time()
recv_start_monotonic = time.monotonic()
recv_start_iso = iso_from_epoch(recv_start_epoch)
try:
message = await asyncio.wait_for(
ws.recv(),
timeout=timeout_remaining,
)
except asyncio.TimeoutError:
session_data["status"] = "timeout"
session_data["error"] = (
"Timed out waiting for websocket message."
)
break
except Exception as exc:
session_data["status"] = "failed"
session_data["error"] = f"WebSocket receive failed: {exc}"
break
recv_finish_epoch = time.time()
recv_finish_monotonic = time.monotonic()
recv_finish_iso = iso_from_epoch(recv_finish_epoch)
if isinstance(message, bytes):
session_data["total_chunks"] += 1
session_data["total_chunk_bytes"] += len(message)
if first_chunk_finish_epoch is None:
first_chunk_finish_epoch = recv_finish_epoch
session_data["first_chunk_finish_ts_utc"] = recv_finish_iso
chunk_gap_ms: float | None = None
if last_chunk_finish_monotonic is not None:
chunk_gap_ms = (
recv_finish_monotonic - last_chunk_finish_monotonic
) * 1000.0
session_data["chunks"].append(
{
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
}
)
last_chunk_finish_monotonic = recv_finish_monotonic
last_chunk_finish_epoch = recv_finish_epoch
session_data["last_chunk_finish_ts_utc"] = recv_finish_iso
continue
if not isinstance(message, str):
continue
try:
data = json.loads(message)
except json.JSONDecodeError as exc:
session_data["status"] = "protocol_error"
session_data["error"] = f"Invalid JSON message: {exc}"
break
msg_type = data.get("type")
if msg_type == "gpu_assigned":
session_data["gpu_assigned_ts_utc"] = recv_finish_iso
if connect_finish_monotonic is not None:
session_data["queue_wait_ms"] = (
recv_finish_monotonic - connect_finish_monotonic
) * 1000.0
elif msg_type == "ltx2_stream_start":
if initial_total_segments is None:
parsed_total = parse_int(data.get("total_segments"))
if parsed_total is not None and parsed_total > 0:
initial_total_segments = parsed_total
session_data["initial_total_segments"] = parsed_total
elif msg_type == "ltx2_segment_start":
parsed_idx = parse_int(data.get("segment_idx"))
current_segment_idx = parsed_idx
session_data["segments_started"] += 1
elif msg_type == "media_segment_complete":
session_data["media_segments_completed"] += 1
if first_media_segment_complete_epoch is None:
first_media_segment_complete_epoch = recv_finish_epoch
session_data[
"first_media_segment_complete_ts_utc"
] = recv_finish_iso
elif msg_type == "ltx2_segment_complete":
session_data["segments_completed"] += 1
seg_idx = parse_int(data.get("segment_idx"))
if (
initial_total_segments is not None
and seg_idx is not None
and seg_idx >= initial_total_segments
):
session_data[
"target_segment_complete_ts_utc"
] = recv_finish_iso
await asyncio.sleep(post_complete_wait_s)
session_data["leave_sent_ts_utc"] = utc_now_iso()
try:
await ws.send(json.dumps({"type": "leave"}))
except Exception:
pass
session_data["status"] = "success"
break
elif msg_type == "session_timeout":
session_data["status"] = "timeout"
session_data["error"] = str(
data.get("message") or "Backend session timeout"
)
break
elif msg_type == "error":
session_data["status"] = "failed"
session_data["error"] = str(
data.get("message") or "Backend error message"
)
break
if session_data["status"] == "failed" and session_data["error"] is None:
session_data["error"] = "Session ended without success."
except Exception as exc:
session_data["status"] = "failed"
session_data["error"] = f"WebSocket connect/run failed: {exc}"
if (
first_chunk_finish_epoch is not None
and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0
):
duration_s = last_chunk_finish_epoch - first_chunk_finish_epoch
if duration_s > 0:
session_data["session_goodput_mbps"] = (
session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0
)
if (
first_chunk_finish_epoch is not None
and first_media_segment_complete_epoch is not None
):
session_data["first_chunk_before_first_media_complete"] = (
first_chunk_finish_epoch < first_media_segment_complete_epoch
)
session_data["close_ts_utc"] = utc_now_iso()
session_data["duration_ms"] = (
time.monotonic() - session_start_monotonic
) * 1000.0
return session_data
async def run_worker_sessions(
*,
worker_id: int,
session_count: int,
config: dict[str, Any],
) -> list[dict[str, Any]]:
tasks = [
asyncio.create_task(
run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)
)
for idx in range(session_count)
]
if not tasks:
return []
return await asyncio.gather(*tasks)
def worker_entry(
worker_id: int,
session_count: int,
config: dict[str, Any],
start_event: Any,
ready_queue: Any,
result_queue: Any,
) -> None:
try:
ready_queue.put({"worker_id": worker_id, "status": "ready"})
start_event.wait()
sessions = asyncio.run(
run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
)
)
result_queue.put(
{
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
}
)
except Exception as exc:
result_queue.put(
{
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
}
)
def build_summary(
*,
sessions: list[dict[str, Any]],
chunk_gap_threshold_ms: float,
) -> dict[str, Any]:
status_counts: dict[str, int] = defaultdict(int)
chunk_gaps: list[float] = []
queue_waits: list[float] = []
session_goodputs: list[float] = []
progressive_eligible = 0
progressive_success = 0
all_chunk_finish_epochs: list[float] = []
bucket_bytes: dict[int, int] = defaultdict(int)
total_chunk_bytes = 0
for session in sessions:
status = str(session.get("status") or "unknown")
status_counts[status] += 1
queue_wait_ms = session.get("queue_wait_ms")
if isinstance(queue_wait_ms, (int, float)):
queue_waits.append(float(queue_wait_ms))
session_goodput = session.get("session_goodput_mbps")
if isinstance(session_goodput, (int, float)):
session_goodputs.append(float(session_goodput))
progressive_value = session.get("first_chunk_before_first_media_complete")
if isinstance(progressive_value, bool):
progressive_eligible += 1
if progressive_value:
progressive_success += 1
for chunk in session.get("chunks", []):
gap = chunk.get("chunk_gap_ms")
if isinstance(gap, (int, float)):
chunk_gaps.append(float(gap))
size_bytes = int(chunk.get("size_bytes") or 0)
finish_epoch = iso_to_epoch(chunk.get("chunk_finish_ts_utc"))
if finish_epoch is None or size_bytes <= 0:
continue
total_chunk_bytes += size_bytes
all_chunk_finish_epochs.append(finish_epoch)
bucket_bytes[int(finish_epoch)] += size_bytes
chunk_gap_stats = summarize_series(chunk_gaps)
queue_wait_stats = summarize_series(queue_waits)
session_goodput_stats = summarize_series(session_goodputs)
global_goodput_mbps: float | None = None
if len(all_chunk_finish_epochs) >= 2 and total_chunk_bytes > 0:
duration_s = max(all_chunk_finish_epochs) - min(all_chunk_finish_epochs)
if duration_s > 0:
global_goodput_mbps = (
total_chunk_bytes * 8.0 / duration_s / 1_000_000.0
)
bucket_throughputs_mbps = [
(bytes_count * 8.0) / 1_000_000.0
for _, bytes_count in sorted(bucket_bytes.items())
]
bucket_stats = summarize_series(bucket_throughputs_mbps)
chunk_gap_threshold_breaches = [
value for value in chunk_gaps if value >= chunk_gap_threshold_ms
]
non_success = len(sessions) - status_counts.get("success", 0)
fail_reasons: list[str] = []
if non_success > 0:
fail_reasons.append(
f"{non_success} session(s) did not complete successfully."
)
if not chunk_gaps:
fail_reasons.append("No chunk gap data collected.")
if chunk_gap_threshold_breaches:
fail_reasons.append(
f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms."
)
passed = len(fail_reasons) == 0
progressive_ratio = None
if progressive_eligible > 0:
progressive_ratio = progressive_success / progressive_eligible
return {
"passed": passed,
"fail_reasons": fail_reasons,
"sessions": {
"total": len(sessions),
"success": status_counts.get("success", 0),
"failed": status_counts.get("failed", 0),
"timeout": status_counts.get("timeout", 0),
"protocol_error": status_counts.get("protocol_error", 0),
"other": (
len(sessions)
- (
status_counts.get("success", 0)
+ status_counts.get("failed", 0)
+ status_counts.get("timeout", 0)
+ status_counts.get("protocol_error", 0)
)
),
},
"chunk_gap_ms": {
**chunk_gap_stats,
"threshold_ms": chunk_gap_threshold_ms,
"breach_count": len(chunk_gap_threshold_breaches),
},
"queue_wait_ms": queue_wait_stats,
"progressive_streaming": {
"eligible_sessions": progressive_eligible,
"success_sessions": progressive_success,
"ratio": progressive_ratio,
},
"bandwidth_mbps": {
"per_session": session_goodput_stats,
"global_goodput_mbps": global_goodput_mbps,
"bucketed_1s": {
"count": bucket_stats.get("count"),
"avg_mbps": bucket_stats.get("avg"),
"peak_mbps": bucket_stats.get("max"),
},
},
}
def print_summary(
*,
run_info: dict[str, Any],
summary: dict[str, Any],
) -> None:
sessions = summary["sessions"]
chunk_gap = summary["chunk_gap_ms"]
queue_wait = summary["queue_wait_ms"]
progressive = summary["progressive_streaming"]
bandwidth = summary["bandwidth_mbps"]
per_session_bw = bandwidth["per_session"]
bucket_bw = bandwidth["bucketed_1s"]
print("=== LTX2 Realtime Stress Test Summary ===")
print(
"Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}"
)
print(
"Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}"
)
print(
"Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}"
)
print(
"Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}"
)
ratio = progressive["ratio"]
ratio_text = "n/a" if ratio is None else f"{ratio * 100:.2f}%"
print(
"Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})"
)
print(
"Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}"
)
print(f"VERDICT: {'PASS' if summary['passed'] else 'FAIL'}")
if summary["fail_reasons"]:
print("Fail reasons:")
for reason in summary["fail_reasons"]:
print(f"- {reason}")
def distribute_sessions(total_clients: int, process_count: int) -> list[int]:
base = total_clients // process_count
remainder = total_clients % process_count
counts = []
for idx in range(process_count):
count = base + (1 if idx < remainder else 0)
counts.append(count)
return counts
def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if websockets is None:
raise RuntimeError(
"Missing dependency: websockets. Install it before running this "
"stress test."
)
preset_file = Path(args.preset_file).expanduser().resolve()
selected_preset_id, curated_prompts, total_prompt_count = load_curated_prompts(
preset_file=preset_file,
preset_id=args.preset_id,
curated_limit=args.curated_limit,
)
process_count = args.processes
if process_count is None:
process_count = min(args.clients, os.cpu_count() or 1)
process_count = max(1, min(process_count, args.clients))
mp_ctx = mp.get_context("spawn")
start_event = mp_ctx.Event()
ready_queue = mp_ctx.Queue()
result_queue = mp_ctx.Queue()
worker_config = {
"url": args.url,
"preset_id": selected_preset_id,
"curated_prompts": curated_prompts,
"connect_timeout_s": args.connect_timeout_s,
"session_timeout_s": args.session_timeout_s,
"post_complete_wait_s": args.post_complete_wait_s,
}
session_counts = distribute_sessions(args.clients, process_count)
run_start_epoch = time.time()
run_start_monotonic = time.monotonic()
run_start_iso = iso_from_epoch(run_start_epoch)
processes: list[mp.Process] = []
for worker_id, session_count in enumerate(session_counts):
proc = mp_ctx.Process(
target=worker_entry,
args=(
worker_id,
session_count,
worker_config,
start_event,
ready_queue,
result_queue,
),
)
proc.start()
processes.append(proc)
try:
ready_workers = 0
ready_deadline = time.monotonic() + 60.0
while ready_workers < len(processes):
timeout_s = max(0.1, ready_deadline - time.monotonic())
if timeout_s <= 0:
raise RuntimeError("Timed out waiting for workers to become ready.")
msg = ready_queue.get(timeout=timeout_s)
if msg.get("status") == "ready":
ready_workers += 1
start_event.set()
result_deadline = (
time.monotonic()
+ args.connect_timeout_s
+ args.session_timeout_s
+ args.post_complete_wait_s
+ 180.0
)
worker_results: list[dict[str, Any]] = []
while len(worker_results) < len(processes):
timeout_s = max(0.1, result_deadline - time.monotonic())
if timeout_s <= 0:
break
try:
result = result_queue.get(timeout=timeout_s)
except Exception:
break
worker_results.append(result)
for proc in processes:
proc.join(timeout=5.0)
if proc.is_alive():
proc.terminate()
proc.join(timeout=2.0)
sessions: list[dict[str, Any]] = []
worker_errors: list[dict[str, Any]] = []
for result in worker_results:
if result.get("status") == "ok":
sessions.extend(result.get("sessions", []))
else:
worker_errors.append(
{
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
}
)
received_workers = {result.get("worker_id") for result in worker_results}
expected_workers = set(range(len(processes)))
missing_workers = sorted(expected_workers - received_workers)
for worker_id in missing_workers:
worker_errors.append(
{
"worker_id": worker_id,
"error": "No worker result received.",
}
)
run_end_epoch = time.time()
run_end_iso = iso_from_epoch(run_end_epoch)
run_duration_ms = (time.monotonic() - run_start_monotonic) * 1000.0
summary = build_summary(
sessions=sessions,
chunk_gap_threshold_ms=args.chunk_gap_threshold_ms,
)
if worker_errors:
summary["passed"] = False
summary["fail_reasons"] = list(summary["fail_reasons"]) + [
f"{len(worker_errors)} worker error(s) occurred."
]
output_payload = {
"run_info": {
"url": args.url,
"clients": args.clients,
"processes": process_count,
"preset_file": str(preset_file),
"preset_id": selected_preset_id,
"curated_limit": args.curated_limit,
"selected_prompt_count": len(curated_prompts),
"preset_total_prompt_count": total_prompt_count,
"chunk_gap_threshold_ms": args.chunk_gap_threshold_ms,
"connect_timeout_s": args.connect_timeout_s,
"session_timeout_s": args.session_timeout_s,
"post_complete_wait_s": args.post_complete_wait_s,
"run_start_ts_utc": run_start_iso,
"run_end_ts_utc": run_end_iso,
"run_duration_ms": run_duration_ms,
},
"summary": summary,
"full_data": {
"sessions": sessions,
"worker_errors": worker_errors,
},
}
exit_code = 0 if summary["passed"] else 1
return output_payload, exit_code
finally:
for proc in processes:
if proc.is_alive():
proc.terminate()
proc.join(timeout=1.0)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Multiprocess realtime stress test for LTX2 streaming.",
)
parser.add_argument(
"-u",
"--url",
required=True,
help="WebSocket URL, e.g. wss://your-domain/ws",
)
parser.add_argument(
"-c",
"--clients",
type=int,
required=True,
help="Total concurrent virtual users.",
)
parser.add_argument(
"--processes",
type=int,
default=None,
help="Worker process count (default: min(clients, cpu_count)).",
)
parser.add_argument(
"--preset-file",
default=str(DEFAULT_PRESET_FILE),
help="Path to curated presets JSON file.",
)
parser.add_argument(
"--preset-id",
default=None,
help="Preset id to use (default: first preset in file).",
)
parser.add_argument(
"--curated-limit",
type=int,
default=6,
help="Number of curated prompts to send from selected preset.",
)
parser.add_argument(
"--chunk-gap-threshold-ms",
type=float,
default=5000.0,
help="Fail if any chunk gap is >= this value.",
)
parser.add_argument(
"--post-complete-wait-s",
type=float,
default=5.0,
help="Seconds to wait after target segment completion before leave.",
)
parser.add_argument(
"--connect-timeout-s",
type=float,
default=20.0,
help="WebSocket connect timeout in seconds.",
)
parser.add_argument(
"--session-timeout-s",
type=float,
default=180.0,
help="Max session runtime per virtual user in seconds.",
)
parser.add_argument(
"-o",
"--output-json",
required=True,
help="Required output JSON path (summary + full data).",
)
args = parser.parse_args()
if args.clients <= 0:
parser.error("--clients must be > 0")
if args.processes is not None and args.processes <= 0:
parser.error("--processes must be > 0")
if args.curated_limit <= 0:
parser.error("--curated-limit must be > 0")
if args.chunk_gap_threshold_ms <= 0:
parser.error("--chunk-gap-threshold-ms must be > 0")
if args.post_complete_wait_s < 0:
parser.error("--post-complete-wait-s must be >= 0")
if args.connect_timeout_s <= 0:
parser.error("--connect-timeout-s must be > 0")
if args.session_timeout_s <= 0:
parser.error("--session-timeout-s must be > 0")
return args
def main() -> int:
try:
args = parse_args()
print("Starting LTX2 realtime stress test with the following parameters:")
for arg, value in vars(args).items():
print(f" {arg}: {value}")
output_payload, exit_code = run_stress(args)
output_path = Path(args.output_json).expanduser().resolve()
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text(
json.dumps(output_payload, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
print_summary(
run_info=output_payload["run_info"],
summary=output_payload["summary"],
)
return exit_code
except SystemExit:
raise
except Exception as exc:
print(f"Error: {exc}", file=sys.stderr)
return 2
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,75 @@
# pyright: reportMissingTypeArgument=false
import base64
import io
from pathlib import Path
from PIL import Image
import pytest
from dreamverse.session_init_image import (
MAX_SESSION_INIT_IMAGE_BYTES,
cleanup_session_init_image,
persist_session_init_image,
)
def make_data_url(format_name: str = "PNG") -> str:
image = Image.new("RGB", (2, 2), color=(16, 32, 64))
buffer = io.BytesIO()
image.save(buffer, format=format_name)
encoded = base64.b64encode(buffer.getvalue()).decode("ascii")
mime_type = "image/png" if format_name == "PNG" else "image/jpeg"
return f"data:{mime_type};base64,{encoded}"
def test_persist_session_init_image_saves_normalized_file(tmp_path: Path):
session_image = persist_session_init_image(
{
"name": "frame.png",
"mime_type": "image/png",
"data_url": make_data_url("PNG"),
},
temp_root=tmp_path,
)
assert session_image is not None
assert session_image.display_name == "frame.png"
assert session_image.file_path.is_file()
cleanup_session_init_image(session_image)
assert not session_image.temp_dir.exists()
def test_persist_session_init_image_returns_none_when_missing_data():
assert persist_session_init_image(None) is None
assert persist_session_init_image({"name": "frame.png", "data_url": ""}) is None
def test_persist_session_init_image_rejects_unsupported_mime():
with pytest.raises(ValueError, match="PNG, JPEG, or WebP"):
persist_session_init_image(
{
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
}
)
def test_persist_session_init_image_rejects_large_payload(monkeypatch):
data_url = make_data_url("PNG")
oversized = "a" * (MAX_SESSION_INIT_IMAGE_BYTES + 1)
def fake_b64decode(value: str, validate: bool = True):
return oversized.encode("ascii")
monkeypatch.setattr(base64, "b64decode", fake_b64decode)
with pytest.raises(ValueError, match="15 MB or smaller"):
persist_session_init_image(
{
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
}
)
File diff suppressed because it is too large Load Diff
+19
View File
@@ -0,0 +1,19 @@
"""Shared helpers used by main.py, routes/, and session/."""
from __future__ import annotations
from datetime import datetime, timezone
def _main_print(level: str, message: str):
print(f"[MAIN][{level}] {message}", flush=True)
def _utc_now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
PROMPT_EXTENSION_FAILURE_USER_MESSAGE = ("Prompt extension failed for this request.")
def _resolve_generation_segment_cap(*, single_clip_mode: bool, cap: int) -> int:
return 0 if single_clip_mode else cap
@@ -0,0 +1,535 @@
"""LTX2 model lifecycle and continuation conditioning.
Runs inside a GPU worker subprocess. Owns the model, the audio
encoder, and the per-session continuation state carried across
segments. Callers must set ``os.environ["CUDA_VISIBLE_DEVICES"]``
before constructing ``VideoGenerationWorker`` — all ``fastvideo.*``
imports are deferred to method bodies so nothing touches CUDA at
module import time.
"""
# pyright: reportArgumentType=false, reportMissingImports=false, reportMissingTypeArgument=false, reportOptionalMemberAccess=false
# ruff: noqa: SIM105
# mypy: ignore-errors
import gc
import os
import time
from dataclasses import dataclass
from typing import Any
import numpy as np
import torch
from dreamverse.config import (
FRAME_HEIGHT,
FRAME_WIDTH,
MODEL_CONFIG,
NUM_FRAMES,
NUM_INFERENCE_STEPS,
)
# Multi-frame decoded continuation defaults from
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
# Overridable via environment variables.
LTX2_VIDEO_CONDITIONING_NUM_FRAMES = int(os.getenv("LTX2_VIDEO_CONDITIONING_NUM_FRAMES", "9"))
LTX2_VIDEO_CONDITIONING_END_OFFSET = int(os.getenv("LTX2_VIDEO_CONDITIONING_END_OFFSET", "0"))
LTX2_VIDEO_CONDITIONING_FRAME_IDX = int(os.getenv("LTX2_VIDEO_CONDITIONING_FRAME_IDX", "0"))
LTX2_VIDEO_CONDITIONING_STRENGTH = float(os.getenv("LTX2_VIDEO_CONDITIONING_STRENGTH", "1.0"))
if (LTX2_VIDEO_CONDITIONING_NUM_FRAMES - 1) % 8 != 0:
raise ValueError("LTX2_VIDEO_CONDITIONING_NUM_FRAMES must satisfy "
"(frames - 1) % 8 == 0; got "
f"{LTX2_VIDEO_CONDITIONING_NUM_FRAMES}")
# Audio conditioning: reuse denoised audio latents from the previous
# segment as initial latents for the next segment.
ENABLE_AUDIO_COND = os.getenv("ENABLE_AUDIO_COND", "1").lower() in ("1", "true", "yes")
# Number of decoded video frames worth of audio to condition on.
# Not subject to the (n-1)%8==0 constraint since this is audio-only.
AUDIO_CONDITIONING_NUM_FRAMES = int(os.getenv("AUDIO_CONDITIONING_NUM_FRAMES", '49'))
AUDIO_CONDITIONING_STRENGTH = float(os.getenv("AUDIO_CONDITIONING_STRENGTH", "1.0"))
# Noise injected into conditioning context to prevent error
# accumulation across segments. 0 = no noise (default).
VIDEO_CONTEXT_NOISE = float(os.getenv("VIDEO_CONTEXT_NOISE", "0"))
AUDIO_CONTEXT_NOISE = float(os.getenv("AUDIO_CONTEXT_NOISE", "0"))
# Re-encode audio latents through decode→encode round-trip to
# regularize noise accumulation (mirrors the video PIL→VAE path).
ENABLE_AUDIO_RE_ENCODE = os.getenv("ENABLE_AUDIO_RE_ENCODE", "").lower() in ("1", "true", "yes")
DEFAULT_LTX2_AUDIO_SAMPLE_RATE = 16000
DEFAULT_LTX2_AUDIO_HOP_LENGTH = 160
DEFAULT_LTX2_AUDIO_DOWNSAMPLE = 4
@dataclass
class StepResult:
"""Output of one generation step.
``head_trim_frames`` / ``head_trim_audio_frames`` are derived here
so downstream AV streaming never needs to import conditioning
constants.
"""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict
head_trim_frames: int
head_trim_audio_frames: int
class ContinuationState:
"""Per-session video + audio conditioning carried across segments."""
def __init__(self):
self.video_images: list | None = None
self.audio_latents: torch.Tensor | None = None
def clear(self) -> None:
if self.video_images:
for old_image in self.video_images:
try:
old_image.close()
except Exception:
pass
self.video_images = None
self.audio_latents = None
def apply_video(self, request_kwargs: dict, segment_idx: int) -> None:
"""Seed next-segment kwargs with the cached tail frames."""
if segment_idx <= 1 or not self.video_images:
return
from PIL import Image
cond_images = list(self.video_images)
if VIDEO_CONTEXT_NOISE > 0:
noisy = []
for img in cond_images:
arr = np.array(img, dtype=np.float32)
arr += np.random.normal(0, VIDEO_CONTEXT_NOISE * 255, arr.shape)
arr = np.clip(arr, 0, 255).astype(np.uint8)
noisy.append(Image.fromarray(arr))
cond_images = noisy
request_kwargs["ltx2_video_conditions"] = [(
cond_images,
LTX2_VIDEO_CONDITIONING_FRAME_IDX,
LTX2_VIDEO_CONDITIONING_STRENGTH,
)]
request_kwargs["ltx2_images"] = None
request_kwargs["image_path"] = None
def apply_audio(
self,
request_kwargs: dict,
segment_idx: int,
audio_lps: float,
) -> None:
"""Seed next-segment kwargs with clean audio latents + denoise mask.
When audio conditioning is longer than video, extend audio
generation and shift video RoPE forward so the audio prefix
sits before video t=0. ``audio_lps`` (audio latent frames per
second) is passed in so this class never imports fastvideo.
"""
if not (ENABLE_AUDIO_COND and segment_idx > 1 and self.audio_latents is not None):
return
cached = self.audio_latents # [B,C,T,mel]
cond_duration = float(AUDIO_CONDITIONING_NUM_FRAMES) / 24.0
audio_cond_T = max(1, round(cond_duration * audio_lps))
audio_cond_T = min(audio_cond_T, cached.shape[2])
audio_extra = max(0, AUDIO_CONDITIONING_NUM_FRAMES - LTX2_VIDEO_CONDITIONING_NUM_FRAMES)
if audio_extra > 0:
audio_num_frames = NUM_FRAMES + audio_extra
request_kwargs["audio_num_frames"] = (audio_num_frames)
prefix_sec = float(audio_extra) / 24.0
request_kwargs["video_position_offset_sec"] = prefix_sec
new_duration = float(NUM_FRAMES + audio_extra) / 24.0
total_T = max(
round(new_duration * audio_lps),
audio_cond_T + 1,
)
B, C, _, mel = cached.shape
clean = torch.zeros((B, C, total_T, mel), dtype=cached.dtype)
clean[:, :, :audio_cond_T, :] = (cached[:, :, -audio_cond_T:, :])
if AUDIO_CONTEXT_NOISE > 0:
noise = torch.randn_like(clean[:, :, :audio_cond_T, :])
clean[:, :, :audio_cond_T, :] += (AUDIO_CONTEXT_NOISE * noise)
mask = torch.ones((B, 1, total_T, 1), dtype=torch.float32)
mask[:, :, :audio_cond_T, :] = (1.0 - AUDIO_CONDITIONING_STRENGTH)
request_kwargs["ltx2_audio_clean_latent"] = clean
request_kwargs["ltx2_audio_denoise_mask"] = mask
def save_video(self, frames: list) -> None:
"""Snapshot trailing N frames as PIL images for next-segment conditioning."""
from PIL import Image
num_cond_frames = LTX2_VIDEO_CONDITIONING_NUM_FRAMES
end_offset = LTX2_VIDEO_CONDITIONING_END_OFFSET
if end_offset + num_cond_frames > len(frames):
raise RuntimeError(f"Cannot extract {num_cond_frames} conditioning frames with "
f"end_offset={end_offset} from {len(frames)} generated frames.")
start_idx = len(frames) - end_offset - num_cond_frames
self.video_images = [
Image.fromarray(np.ascontiguousarray(frames[start_idx + i])) for i in range(num_cond_frames)
]
def save_audio_latents(self, latents: torch.Tensor | None) -> None:
if latents is None:
self.audio_latents = None
return
self.audio_latents = latents.detach().clone().cpu()
class VideoGenerationWorker:
"""Single-GPU LTX2 generator with continuation state.
Caller must set ``os.environ["CUDA_VISIBLE_DEVICES"]`` before
instantiating, and call ``initialize()`` before any
``generate_step()`` / ``warmup()``.
"""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.generator = None
self.current_model_config: dict = dict(MODEL_CONFIG)
self.continuation = ContinuationState()
self.audio_encoder_module = None
self.audio_processor_module = None
def _gpu_mem(self) -> str:
a = torch.cuda.memory_allocated() / 1024**3
r = torch.cuda.memory_reserved() / 1024**3
return f"alloc={a:.2f}GiB, reserved={r:.2f}GiB"
@staticmethod
def _resolve_refine_upsampler_path(model_root: str) -> str:
candidates = (
os.path.join(model_root, "spatial_upscaler"),
os.path.join(model_root, "spatial_upsampler"),
)
for candidate in candidates:
config_path = os.path.join(candidate, "config.json")
if os.path.isfile(config_path):
return candidate
raise FileNotFoundError("Could not find an LTX2 refine upsampler directory under "
f"{model_root}. Checked: {', '.join(candidates)}")
def initialize(self, model_config: dict | None = None) -> None:
"""Load (or reload) the LTX2 generator on the visible GPU."""
if model_config is not None:
self.current_model_config = model_config
if self.generator is not None:
print(f"[GPU {self.gpu_id}] Freeing old model...")
try:
self.generator.shutdown()
except Exception:
pass
del self.generator
self.generator = None
gc.collect()
torch.cuda.empty_cache()
print(f"[GPU {self.gpu_id}] After cleanup: {self._gpu_mem()}")
print(f"[GPU {self.gpu_id}] Loading model: "
f"{self.current_model_config['model_path']}")
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
from fastvideo.api.schema import (
CompileConfig,
ComponentConfig,
EngineConfig,
GeneratorConfig,
OffloadConfig,
PipelineSelection,
QuantizationConfig,
)
from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo.utils import maybe_download_model
model_root = maybe_download_model(self.current_model_config["model_path"])
refine_upsampler_path = self._resolve_refine_upsampler_path(model_root)
config_model_path = (self.current_model_config.get("config_model_path")
or self.current_model_config["model_path"])
enable_compile = os.getenv("ENABLE_TORCH_COMPILE", "1") == "1"
components = ComponentConfig(
config_root=config_model_path,
upsampler_weights=refine_upsampler_path,
)
init_weights = self.current_model_config.get("init_weights_from_safetensors")
if init_weights:
components.transformer_weights = init_weights
generator_config = GeneratorConfig(
model_path=model_root,
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
text_encoder=False,
vae=False,
pin_cpu_memory=True,
),
compile=CompileConfig(
enabled=enable_compile,
text_encoder_enabled=enable_compile,
backend="inductor",
fullgraph=True,
mode="max-autotune-no-cudagraphs",
dynamic=False,
),
use_fsdp_inference=False,
quantization=QuantizationConfig(transformer_quant="NVFP4"),
),
pipeline=PipelineSelection(
components=components,
vae_tiling=False,
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 2,
"guidance_scale": 1.0,
"add_noise": True,
},
},
),
)
self.generator = VideoGenerator.from_pretrained(config=generator_config)
print(f"[GPU {self.gpu_id}] After model load: {self._gpu_mem()}")
self._load_audio_encoder(model_root)
print(f"[GPU {self.gpu_id}] LTX2 model loaded (warmup pending)")
def _load_audio_encoder(self, model_root: str) -> None:
if not ENABLE_AUDIO_RE_ENCODE:
return
from fastvideo.models.audio.ltx2_audio_processing import AudioProcessor
from fastvideo.models.loader.component_loader import ComponentLoader
audio_vae_path = os.path.join(model_root, "audio_vae")
if not os.path.isdir(audio_vae_path):
print(f"[GPU {self.gpu_id}] audio_vae dir not found at "
f"{audio_vae_path}; disabling re-encode")
return
loader = ComponentLoader.for_module_type("audio_encoder", "diffusers")
enc = loader.load(audio_vae_path, self.generator.fastvideo_args)
target = getattr(enc, "model", enc)
proc = AudioProcessor(
sample_rate=target.sample_rate,
mel_bins=target.mel_bins,
mel_hop_length=target.mel_hop_length,
n_fft=target.n_fft,
).to(torch.device("cuda"))
self.audio_encoder_module = target
self.audio_processor_module = proc
print(f"[GPU {self.gpu_id}] Audio encoder loaded for "
f"re-encode conditioning ({self._gpu_mem()})")
def _re_encode_audio(
self,
waveform: torch.Tensor,
sample_rate: int,
) -> torch.Tensor | None:
"""Waveform → mel → encoder → latents."""
if (self.audio_encoder_module is None or self.audio_processor_module is None):
return None
device = torch.device("cuda")
if waveform.ndim == 1:
waveform = waveform.unsqueeze(0) # [channels, samples]
waveform = waveform.unsqueeze(0).to(device=device, dtype=torch.float32) # [1, ch, samples]
with torch.no_grad():
mel = self.audio_processor_module.waveform_to_mel(
waveform,
waveform_sample_rate=sample_rate,
).to(device=device, dtype=torch.float32)
latents = self.audio_encoder_module(mel)
return latents.detach()
def shutdown(self) -> None:
if self.generator is not None:
try:
self.generator.shutdown()
except Exception:
pass
def clear_conditioning(self) -> None:
self.continuation.clear()
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> StepResult:
"""Execute one generation step; snapshot state for the next segment."""
timings: dict = {}
request_kwargs = dict(
prompt=prompt,
negative_prompt="",
save_video=False,
height=FRAME_HEIGHT,
width=FRAME_WIDTH,
num_frames=NUM_FRAMES,
fps=24,
num_inference_steps=NUM_INFERENCE_STEPS,
guidance_scale=1.0,
seed=10,
ltx2_image_crf=0.0,
image_path=image_path if segment_idx == 1 else None,
return_continuation_state=False,
)
if reset_conditioning:
self.continuation.clear()
audio_lps = (DEFAULT_LTX2_AUDIO_SAMPLE_RATE / DEFAULT_LTX2_AUDIO_HOP_LENGTH / DEFAULT_LTX2_AUDIO_DOWNSAMPLE)
# Phase 1: seed kwargs with prior-segment conditioning.
self.continuation.apply_video(request_kwargs, segment_idx)
self.continuation.apply_audio(request_kwargs, segment_idx, audio_lps)
# Phase 2: generate.
t0 = time.perf_counter()
result = self.generator.generate_video(**request_kwargs)
torch.cuda.synchronize()
timings["generation_ms"] = (time.perf_counter() - t0) * 1000
if not isinstance(result, dict):
raise RuntimeError("Expected dictionary output from generate_video.")
frames = result.get("frames")
if not isinstance(frames, list) or len(frames) == 0:
raise RuntimeError("Generation did not return frames.")
audio = result.get("audio")
audio_sample_rate = result.get("audio_sample_rate")
if audio is not None and audio_sample_rate is None:
# LTX2 audio decoding stage uses 24kHz output by default.
audio_sample_rate = 24000
print(f"[GPU {self.gpu_id}] audio_sample_rate missing from result; "
f"defaulting to {audio_sample_rate}Hz")
timings["generation_time_ms"] = result.get("generation_time", 0.0) * 1000
# Phase 3: snapshot continuation state for the next segment.
t_save_start = time.perf_counter()
self.continuation.clear()
self.continuation.save_video(frames)
next_audio_latents = self._derive_next_audio_latents(audio, audio_sample_rate, result, segment_idx)
self.continuation.save_audio_latents(next_audio_latents)
timings["save_conditioning_ms"] = (time.perf_counter() - t_save_start) * 1000
timings["e2e_latency_ms"] = (time.perf_counter() - t0) * 1000
print(f"[GPU {self.gpu_id}] LTX2 segment {segment_idx}: "
f"{len(frames)} frames, gen={timings['generation_ms']:.0f}ms, "
f"save_conditioning={timings['save_conditioning_ms']:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
# Head-trim values for downstream AV streaming — computed here so
# the streaming layer never needs to know conditioning constants.
is_continuation = segment_idx > 1 and not reset_conditioning
head_trim_frames = (LTX2_VIDEO_CONDITIONING_NUM_FRAMES if is_continuation else 0)
audio_extra = (max(0, AUDIO_CONDITIONING_NUM_FRAMES -
LTX2_VIDEO_CONDITIONING_NUM_FRAMES) if ENABLE_AUDIO_COND else 0)
head_trim_audio_frames = (head_trim_frames + audio_extra if is_continuation else 0)
return StepResult(
frames=frames,
audio=audio,
audio_sample_rate=audio_sample_rate,
timings=timings,
head_trim_frames=head_trim_frames,
head_trim_audio_frames=head_trim_audio_frames,
)
def _derive_next_audio_latents(
self,
audio: object,
audio_sample_rate: int | None,
result: dict,
segment_idx: int,
) -> torch.Tensor | None:
"""Pick which tensor to cache for next-segment audio conditioning."""
if not ENABLE_AUDIO_COND:
return None
if (ENABLE_AUDIO_RE_ENCODE and audio is not None and audio_sample_rate is not None):
re_encoded = self._re_encode_audio(audio, audio_sample_rate)
if re_encoded is not None:
print(f"[GPU {self.gpu_id}] Re-encoded audio "
f"latents shape="
f"{tuple(re_encoded.shape)} "
f"for segment {segment_idx + 1}")
return re_encoded
return None
audio_latents = result.get("ltx2_audio_latents")
if audio_latents is not None:
print(f"[GPU {self.gpu_id}] Cached audio latents "
f"shape={tuple(audio_latents.shape)} "
f"for segment {segment_idx + 1}")
return audio_latents
return None
def warmup(self, prompt: str) -> dict[str, float]:
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] Startup warmup starting "
"(synthetic segments: seg1, seg2, seg1-post-LoRA)")
warmup_t0 = time.perf_counter()
r1 = self.generate_step(
warmup_prompt,
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
r2 = self.generate_step(
warmup_prompt,
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
# r1 stage 1 compiled BEFORE LoRA wrapping (which happens during
# r1 stage 2 via ltx2_refine_lora_stage), so the resulting graph
# is keyed off pre-LoRA module identity and is stale once r1 r2
# finish. The first real user seg=1 then re-compiles stage 1
# ("stage1-LoRA-nocont"), wasting ~90s on the user's first
# request. Run a 3rd warmup pass with seg=1 reset=True after r2
# so this graph is compiled while no client is waiting. r3
# stage 2 hits r1 stage 2's cache (shape match, both LoRA-nocont)
# so the only real work is the missing stage 1 graph.
self.continuation.clear()
r3 = self.generate_step(
warmup_prompt,
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
warmup_total_ms = (time.perf_counter() - warmup_t0) * 1000.0
self.continuation.clear()
segment1_ms = float(r1.timings.get("e2e_latency_ms", 0.0))
segment2_ms = float(r2.timings.get("e2e_latency_ms", 0.0))
segment3_ms = float(r3.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] Startup warmup complete: "
f"segment1={segment1_ms:.0f}ms, "
f"segment2={segment2_ms:.0f}ms, "
f"segment3={segment3_ms:.0f}ms, total={warmup_total_ms:.0f}ms")
return {
"warmup_segment1_ms": segment1_ms,
"warmup_segment2_ms": segment2_ms,
"warmup_segment3_ms": segment3_ms,
"warmup_total_ms": warmup_total_ms,
}
+157
View File
@@ -0,0 +1,157 @@
# pyright: reportMissingTypeArgument=false
"""Typed worker IPC events: replaces the fat ``Response`` dataclass.
Each message kind from GPU worker → main process is its own small
dataclass; ``WorkerEvent`` is the union of them. Consumers dispatch
via ``match``/``case`` or ``isinstance`` — mirroring the existing
``StreamEvent`` pattern in ``av_streaming.py``.
Invalid states are unrepresentable: a ``MediaChunk`` simply has no
``frames`` field, a ``StepComplete`` has no ``chunk_offset`` field,
and a ``JoinAck`` can't accidentally default to ``kind="step_result"``
because there is no ``kind`` string.
"""
from __future__ import annotations
from dataclasses import dataclass
# ---- User-scoped events (carry user_id) ------------------------------------
@dataclass(frozen=True)
class StepComplete:
"""Generation step finished. Frames/audio were already sent via ``MediaChunk``."""
user_id: str
segment_idx: int
timings: dict[str, float]
@dataclass(frozen=True)
class WorkerError:
"""Any worker failure. ``user_id`` is None for system-scoped failures."""
user_id: str | None
message: str
@dataclass(frozen=True)
class JoinAck:
user_id: str
@dataclass(frozen=True)
class LeaveAck:
user_id: str
@dataclass(frozen=True)
class ReloadAck:
user_id: str
@dataclass(frozen=True)
class WarmupComplete:
user_id: str | None
timings: dict[str, float]
# ---- Streaming events (carry user_id + segment_idx for routing) ------------
@dataclass(frozen=True)
class MediaInit:
user_id: str
segment_idx: int
stream_id: str
mime: str
uses_shared_buffer: bool
@dataclass(frozen=True)
class MediaChunk:
"""One fMP4 chunk.
Either ``chunk`` (raw bytes) is set, or both ``chunk_offset`` and
``chunk_length`` are set (read from the shared buffer). The
invariant is enforced in ``__post_init__``.
"""
user_id: str
segment_idx: int
stream_id: str
chunk: bytes | None = None
chunk_offset: int | None = None
chunk_length: int | None = None
uses_shared_buffer: bool = False
def __post_init__(self) -> None:
has_bytes = self.chunk is not None
has_offset = (self.chunk_offset is not None and self.chunk_length is not None)
if has_bytes == has_offset:
raise ValueError("MediaChunk must carry either chunk bytes or "
"(chunk_offset + chunk_length), not both or neither")
@dataclass(frozen=True)
class MediaComplete:
user_id: str
segment_idx: int
stream_id: str
chunks: int
# ---- System events (no user_id) --------------------------------------------
@dataclass(frozen=True)
class InitAck:
success: bool
error: str | None = None
@dataclass(frozen=True)
class ShutdownAck:
pass
WorkerEvent = (StepComplete
| WorkerError
| JoinAck
| LeaveAck
| ReloadAck
| WarmupComplete
| MediaInit
| MediaChunk
| MediaComplete
| InitAck
| ShutdownAck)
# ---- Command payloads (main process → worker) ------------------------------
#
# The envelope (``Command`` + ``CommandType``) lives in ``gpu_pool.py``
# alongside the enum; only the typed payloads live here so they sit
# next to the response-side types. Commands without a payload
# (``INIT``, ``SHUTDOWN``, ``USER_JOIN``, ``USER_LEAVE``) leave
# ``Command.payload`` as ``None``.
@dataclass(frozen=True)
class UserStepPayload:
prompt: str
segment_idx: int
image_path: str | None
reset_conditioning: bool
@dataclass(frozen=True)
class WarmupPayload:
prompt: str
@dataclass(frozen=True)
class ReloadModelPayload:
# ``model_config`` stays a dict because ``MODEL_REGISTRY`` values
# are dicts shaped by the external config module; typing that dict
# is a separate refactor.
model_config: dict
CommandPayload = UserStepPayload | WarmupPayload | ReloadModelPayload
+577
View File
@@ -0,0 +1,577 @@
<mxfile host="65bd71144e">
<diagram id="gpu-pool" name="gpu_pool.py">
<mxGraphModel dx="1687" dy="3380" grid="1" gridSize="10" guides="1" tooltips="1" connect="1" arrows="1" fold="1" page="1" pageScale="1" pageWidth="1300" pageHeight="1700" math="0" shadow="0">
<root>
<mxCell id="0"/>
<mxCell id="1" parent="0"/>
<mxCell id="title" value="gpu_pool.py — runtime architecture" style="text;html=1;align=center;fontSize=18;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="350" y="-150" width="600" height="30" as="geometry"/>
</mxCell>
<mxCell id="client" value="Client (WebSocket)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#dae8fc;strokeColor=#6c8ebf;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="560" y="-60" width="180" height="60" as="geometry"/>
</mxCell>
<mxCell id="pool" value="GPUPool" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="560" y="60" width="180" height="60" as="geometry"/>
</mxCell>
<mxCell id="slot_wrap" value="GPUSlot (one per GPU)" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#e6f2e0;strokeColor=#82b366;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="40" y="216" width="1220" height="455" as="geometry"/>
</mxCell>
<mxCell id="m_start" value="start()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="110" y="270" width="60" height="30" as="geometry"/>
</mxCell>
<mxCell id="m_shutdown" value="shutdown()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="180" y="270" width="70" height="30" as="geometry"/>
</mxCell>
<mxCell id="m_join" value="join_user()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="499" y="260" width="70" height="40" as="geometry"/>
</mxCell>
<mxCell id="m_step" value="user_step()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="580" y="260" width="80" height="40" as="geometry"/>
</mxCell>
<mxCell id="m_leave" value="leave_user()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="670" y="255" width="80" height="45" as="geometry"/>
</mxCell>
<mxCell id="m_send" value="_send_command()&lt;br&gt;&lt;br&gt;puts on command_queue,&lt;br&gt;blocks on response_queue.get&lt;div&gt;&lt;br&gt;&lt;/div&gt;&lt;div&gt;&lt;b&gt;**for user agonistic commands&lt;/b&gt;&lt;/div&gt;" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="79" y="370" width="230" height="80" as="geometry"/>
</mxCell>
<mxCell id="m_send_tagged" value="_send_command_tagged()&lt;br&gt;&lt;br&gt;creates Future; registers in&lt;br&gt;_pending_futures[user_id];&lt;br&gt;ensures response reader; &lt;b&gt;awaits future&lt;/b&gt;&lt;div&gt;&lt;br&gt;&lt;/div&gt;&lt;div&gt;&lt;b&gt;**for user specific commands&lt;/b&gt;&lt;/div&gt;" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="350" y="374" width="300" height="90" as="geometry"/>
</mxCell>
<mxCell id="m_reader" value="_response_reader()&lt;br&gt;&lt;br&gt;a while loop that reads response_queue;&lt;br&gt;dispatch WorkerEvent:&lt;br&gt;• media → _stream_queues&lt;br&gt;• acks/errors → _pending_futures" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=10;align=left;spacingLeft=8;spacingTop=4;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="800" y="360" width="340" height="110" as="geometry"/>
</mxCell>
<mxCell id="c_futures" value="«dict» _pending_futures&#xa;&#xa;{user_id → asyncio.Future}&#xa;&#xa;created by _send_command_tagged;&#xa;resolved + popped by _response_reader;&#xa;cleared on model reload or leave" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#fff2cc;strokeColor=#d6b656;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="300" y="524" width="340" height="110" as="geometry"/>
</mxCell>
<mxCell id="c_streams" value="«dict» _stream_queues:&amp;nbsp;&lt;div&gt;&lt;span style=&quot;color: rgb(63, 63, 63); background-color: transparent;&quot;&gt;each is a queue of metadata that tell the consumer which part of the shared&amp;nbsp;&lt;/span&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;font color=&quot;#000000&quot;&gt;stream buffer to read next.&lt;br&gt;&lt;/font&gt;&lt;br&gt;{user_id → asyncio.Queue}&lt;br&gt;&lt;br&gt;created by register_stream_queue()&lt;br&gt;written by _response_reader (MediaInit / Chunk / Complete);&lt;br&gt;consumed by&amp;nbsp; AV loop → ws.send_bytes / send_json;&lt;br&gt;popped by leave_user(); cleared on reload&lt;/div&gt;&lt;/div&gt;" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#fff2cc;strokeColor=#d6b656;fontSize=10;align=left;spacingLeft=8;spacingTop=4;" parent="1" vertex="1">
<mxGeometry x="720" y="524" width="460" height="126" as="geometry"/>
</mxCell>
<mxCell id="div_line" value="" style="line;strokeColor=#888888;strokeWidth=1;dashed=1;" parent="1" vertex="1">
<mxGeometry x="40" y="690" width="1220" height="10" as="geometry"/>
</mxCell>
<mxCell id="div_label" value="── process boundary (mp.Queue + mp.RawArray) ──" style="text;html=1;align=center;fontSize=11;fontColor=#666666;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="450" y="688" width="400" height="20" as="geometry"/>
</mxCell>
<mxCell id="cmdq" value="command_queue" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=12;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="79" y="720" width="200" height="100" as="geometry"/>
</mxCell>
<mxCell id="respq" value="response_queue" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=12;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="520" y="720" width="200" height="100" as="geometry"/>
</mxCell>
<mxCell id="sharedbuf" value="shared_stream_buffer&#xa;(mp.RawArray, 256 MiB)&#xa;&#xa;zero-copy mp4 bytes" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="900" y="720" width="280" height="100" as="geometry"/>
</mxCell>
<mxCell id="worker_ipc" value="worker_ipc.py — shared types (imported on both sides of the process boundary)&#xa;&#xa;Command(type: CommandType, payload: CommandPayload | None, user_id: str | None)&#xa;CommandPayload = UserStepPayload | WarmupPayload | ReloadModelPayload&#xa;WorkerEvent = StepComplete | WorkerError | JoinAck | LeaveAck | ReloadAck | WarmupComplete | MediaInit | MediaChunk | MediaComplete | InitAck | ShutdownAck" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e0e0e0;strokeColor=#666666;fontSize=10;align=left;spacingLeft=10;spacingTop=6;fontFamily=monospace;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="80" y="860" width="1140" height="80" as="geometry"/>
</mxCell>
<mxCell id="worker_wrap" value="Worker Subprocess - in each GPUSlot" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#fef1e0;strokeColor=#d79b00;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="73" y="1071" width="1140" height="360" as="geometry"/>
</mxCell>
<mxCell id="dispatcher" value="command dispatcher&#xa;&#xa;gpu_worker_process() branches on&#xa;CommandType; asserts payload type&#xa;&#xa;INIT / WARMUP / RELOAD_MODEL&#xa;USER_JOIN / USER_STEP / USER_LEAVE&#xa;SHUTDOWN" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe6cc;strokeColor=#d79b00;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="120" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;video_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="460" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="stream_av" value="stream_fmp4()&#xa;av_streaming.py:121&#xa;&#xa;trims overlap, pipes to ffmpeg,&#xa;publishes StreamInit / StreamChunk /&#xa;StreamComplete via callback" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#b1d8d7;strokeColor=#23445d;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="800" y="1120" width="260" height="120" as="geometry"/>
</mxCell>
<mxCell id="Ot8BU52QTIb4EhyRSe7I-2" value="" style="edgeStyle=none;html=1;" parent="1" source="generator" target="Ot8BU52QTIb4EhyRSe7I-1" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;video_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="460" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="ffmpeg" value="ffmpeg subprocess&#xa;&#xa;libx264 / *_nvenc&#xa;fragmented mp4" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="800" y="1300" width="260" height="100" as="geometry"/>
</mxCell>
<mxCell id="caches" value="ContinuationState&#xa;video_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="120" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="e_cp" value="acquire" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="client" target="pool" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_ps" value="assigns slot" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="pool" target="slot_wrap" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_start_send" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=10;" parent="1" target="m_send" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="138" y="300" as="sourcePoint"/>
<Array as="points">
<mxPoint x="138" y="380"/>
<mxPoint x="138" y="380"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_join_tag" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=10;exitX=0.5;exitY=1;exitDx=0;exitDy=0;" parent="1" source="m_join" target="m_send_tagged" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="531.6999999999998" y="400.3600000000001" as="targetPoint"/>
<Array as="points">
<mxPoint x="532" y="300"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_step_tag" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=10;exitX=0.1;exitY=1;exitDx=0;exitDy=0;" parent="1" source="m_step" target="m_send_tagged" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="588" y="400" as="targetPoint"/>
<Array as="points">
<mxPoint x="588" y="337"/>
<mxPoint x="590" y="337"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_leave_tag" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=10;exitX=0.1;exitY=1;exitDx=0;exitDy=0;entryX=0.874;entryY=-0.002;entryDx=0;entryDy=0;entryPerimeter=0;" parent="1" source="m_leave" target="m_send_tagged" edge="1">
<mxGeometry x="-0.1388" relative="1" as="geometry">
<Array as="points">
<mxPoint x="680" y="300"/>
<mxPoint x="680" y="330"/>
<mxPoint x="612" y="330"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_tag_futures" value="register future" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d6b656;endArrow=classic;fontSize=10;exitX=0.397;exitY=0.989;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;exitPerimeter=0;" parent="1" source="m_send_tagged" target="c_futures" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_reader_futures" value="&lt;b&gt;set future result&lt;/b&gt;" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d6b656;endArrow=classic;fontSize=10;exitX=0.008;exitY=0.908;exitDx=0;exitDy=0;entryX=1;entryY=0.3;entryDx=0;entryDy=0;exitPerimeter=0;" parent="1" source="m_reader" target="c_futures" edge="1">
<mxGeometry x="-0.325" relative="1" as="geometry">
<Array as="points">
<mxPoint x="803" y="470"/>
<mxPoint x="670" y="470"/>
<mxPoint x="670" y="557"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_reader_streams" value="put(MediaEvent)" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d6b656;endArrow=classic;fontSize=10;exitX=0.146;exitY=1.008;exitDx=0;exitDy=0;exitPerimeter=0;entryX=0.282;entryY=0.021;entryDx=0;entryDy=0;entryPerimeter=0;" parent="1" source="m_reader" target="c_streams" edge="1">
<mxGeometry relative="1" as="geometry">
<Array as="points"/>
<mxPoint x="850" y="522" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="e_send_cmdq" value="put Command" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=11;entryX=0.35;entryY=0;entryDx=0;entryDy=0;" parent="1" source="m_send" target="cmdq" edge="1">
<mxGeometry relative="1" as="geometry">
<Array as="points">
<mxPoint x="150" y="585"/>
<mxPoint x="149" y="585"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_tag_cmdq" value="put Command" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=11;exitX=0.022;exitY=0.97;exitDx=0;exitDy=0;entryX=0.85;entryY=0;entryDx=0;entryDy=0;exitPerimeter=0;" parent="1" source="m_send_tagged" target="cmdq" edge="1">
<mxGeometry x="0.1755" relative="1" as="geometry">
<Array as="points">
<mxPoint x="357" y="490"/>
<mxPoint x="250" y="490"/>
<mxPoint x="250" y="592"/>
<mxPoint x="249" y="592"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_respq_reader" value="WorkerEvent" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=11;exitX=0.753;exitY=0.056;exitDx=0;exitDy=0;entryX=0;entryY=0.5;entryDx=0;entryDy=0;exitPerimeter=0;" parent="1" source="respq" target="m_reader" edge="1">
<mxGeometry x="0.7113" relative="1" as="geometry">
<Array as="points">
<mxPoint x="671" y="720"/>
<mxPoint x="670" y="720"/>
<mxPoint x="670" y="415"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_respq_send" value="blocking .get" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;dashed=1;fontSize=10;exitX=0;exitY=0.5;exitDx=0;exitDy=0;" parent="1" source="respq" target="m_send" edge="1">
<mxGeometry x="-0.2187" relative="1" as="geometry">
<Array as="points">
<mxPoint x="520" y="680"/>
<mxPoint x="200" y="680"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_cd" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.25;entryY=0;entryDx=0;entryDy=0;" parent="1" source="cmdq" edge="1" target="dispatcher">
<mxGeometry relative="1" as="geometry">
<mxPoint x="310" y="1120" as="targetPoint"/>
<Array as="points">
<mxPoint x="179" y="970"/>
<mxPoint x="180" y="970"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_dds" value="USER_STEP" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d79b00;endArrow=classic;fontSize=10;exitX=1;exitY=0.5;exitDx=0;exitDy=0;entryX=0;entryY=0.5;entryDx=0;entryDy=0;" parent="1" source="dispatcher" target="do_step" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_dr" value="StepComplete / ack types" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=10;exitX=0.5;exitY=0;exitDx=0;exitDy=0;entryX=0.15;entryY=1;entryDx=0;entryDy=0;" parent="1" source="dispatcher" target="respq" edge="1">
<mxGeometry relative="1" as="geometry">
<Array as="points">
<mxPoint x="240" y="1000"/>
<mxPoint x="550" y="1000"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_dsg" value="generator.generate_video()" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#9673a6;endArrow=classic;fontSize=10;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="do_step" target="generator" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_dscache" value="read / write" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d6b656;endArrow=classic;startArrow=classic;fontSize=10;exitX=0;exitY=0.8;exitDx=0;exitDy=0;entryX=1;entryY=0.2;entryDx=0;entryDy=0;" parent="1" source="do_step" target="caches" edge="1">
<mxGeometry relative="1" as="geometry"/>
<Array as="points">
<mxPoint x="420" y="1058"/>
<mxPoint x="420" y="1150"/>
</Array>
</mxCell>
<mxCell id="e_dsstream" value="frames + audio" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d79b00;endArrow=classic;fontSize=10;exitX=1;exitY=0.5;exitDx=0;exitDy=0;entryX=0;entryY=0.5;entryDx=0;entryDy=0;" parent="1" source="do_step" target="stream_av" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_sf" value="rawvideo + wav" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d79b00;endArrow=classic;fontSize=10;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="stream_av" target="ffmpeg" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_ssb" value="write mp4 bytes&#xa;at shared_write_offset" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=10;exitX=0.75;exitY=0;exitDx=0;exitDy=0;entryX=0.5;entryY=1;entryDx=0;entryDy=0;" parent="1" source="stream_av" target="sharedbuf" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_sbs" value="bytes from the buffer are sent through websocket&amp;nbsp;&lt;div&gt;based on info from the stream queue&lt;/div&gt;" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;dashed=1;fontSize=10;exitX=0.591;exitY=0.027;exitDx=0;exitDy=0;entryX=0.75;entryY=1;entryDx=0;entryDy=0;exitPerimeter=0;" parent="1" source="sharedbuf" target="c_streams" edge="1">
<mxGeometry x="0.0048" relative="1" as="geometry">
<Array as="points">
<mxPoint x="1066" y="720"/>
</Array>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="e_scl" value="session AV loop (main.py:2124):&#xa;ws.send_bytes(chunk) — binary mp4 frames&#xa;ws.send_json({type: media_init | media_segment_complete | ...})" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=10;exitX=1;exitY=0.5;exitDx=0;exitDy=0;entryX=1;entryY=0.5;entryDx=0;entryDy=0;" parent="1" source="c_streams" target="client" edge="1">
<mxGeometry x="0.0003" relative="1" as="geometry">
<mxPoint as="offset"/>
<Array as="points">
<mxPoint x="1310" y="577"/>
<mxPoint x="1310" y="-30"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_sr" value="MediaChunk (offset, len)" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=10;exitX=0.1;exitY=0;exitDx=0;exitDy=0;entryX=0.85;entryY=1;entryDx=0;entryDy=0;" parent="1" source="stream_av" target="respq" edge="1">
<mxGeometry relative="1" as="geometry"/>
<Array as="points">
<mxPoint x="826" y="880"/>
<mxPoint x="690" y="880"/>
</Array>
</mxCell>
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender video_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="39" y="-200" width="270" height="380" as="geometry"/>
</mxCell>
<mxCell id="Ot8BU52QTIb4EhyRSe7I-1" value="FastVideo video_generator" style="whiteSpace=wrap;html=1;fontSize=11;fillColor=#e1d5e7;strokeColor=#9673a6;rounded=1;" parent="1" vertex="1">
<mxGeometry x="520" y="1470" width="120" height="60" as="geometry"/>
</mxCell>
<mxCell id="5" value="" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#82b366;endArrow=classic;fontSize=10;exitX=0.421;exitY=1.01;exitDx=0;exitDy=0;exitPerimeter=0;entryX=0.569;entryY=0.02;entryDx=0;entryDy=0;entryPerimeter=0;" edge="1" parent="1" source="m_shutdown" target="m_send">
<mxGeometry relative="1" as="geometry">
<mxPoint x="260" y="300" as="sourcePoint"/>
<mxPoint x="260" y="400" as="targetPoint"/>
<Array as="points"/>
</mxGeometry>
</mxCell>
<mxCell id="6" style="edgeStyle=orthogonalEdgeStyle;html=1;entryX=0.983;entryY=0.021;entryDx=0;entryDy=0;entryPerimeter=0;strokeColor=#61A25D;" edge="1" parent="1" source="m_join" target="m_send">
<mxGeometry relative="1" as="geometry">
<Array as="points">
<mxPoint x="534" y="320"/>
<mxPoint x="306" y="320"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="11" value="reload model" style="edgeLabel;html=1;align=center;verticalAlign=middle;resizable=0;points=[];" vertex="1" connectable="0" parent="6">
<mxGeometry x="0.1965" relative="1" as="geometry">
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="9" value="" style="group" vertex="1" connectable="0" parent="1">
<mxGeometry x="1340" y="750" width="240" height="40" as="geometry"/>
</mxCell>
<mxCell id="7" value="" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=6.353056066176691;fillColor=#f8cecc;strokeColor=#b85450;fontSize=11;" vertex="1" parent="9">
<mxGeometry y="4" width="21.39" height="22" as="geometry"/>
</mxCell>
<mxCell id="8" value="Queue: owned by GPUSlot, passed to worker process as argument." style="text;html=1;align=center;verticalAlign=middle;whiteSpace=wrap;rounded=0;" vertex="1" parent="9">
<mxGeometry x="21.3900000000001" width="218.61" height="40" as="geometry"/>
</mxCell>
</root>
</mxGraphModel>
</diagram>
<diagram id="gpu-pool-boot" name="boot &amp; ws">
<mxGraphModel grid="1" page="1" gridSize="10" guides="1" tooltips="1" connect="1" arrows="1" fold="1" pageScale="1" pageWidth="1300" pageHeight="1900" math="0" shadow="0">
<root>
<mxCell id="0"/>
<mxCell id="1" parent="0"/>
<mxCell id="b_title" value="gpu_pool.py — boot flow &amp; websocket session" style="text;html=1;align=center;fontSize=18;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="350" y="20" width="600" height="30" as="geometry"/>
</mxCell>
<mxCell id="secA" value="A. Call chain" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#e1f5fe;strokeColor=#039be5;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="40" y="70" width="400" height="420" as="geometry"/>
</mxCell>
<mxCell id="a1" value="$ uv run dreamverse-server" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#039be5;fontSize=12;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="70" y="120" width="340" height="40" as="geometry"/>
</mxCell>
<mxCell id="a2" value="server_entry.cli()&#xa;(pyproject.toml [project.scripts])&#xa;&#xa;guards missing fastvideo extra" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#039be5;fontSize=11;align=left;spacingLeft=10;" parent="1" vertex="1">
<mxGeometry x="70" y="180" width="340" height="70" as="geometry"/>
</mxCell>
<mxCell id="a3" value="main.cli() (main.py:2325)&#xa;&#xa;argparse --host/--port&#xa;uvicorn.run(app, host, port)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#039be5;fontSize=11;align=left;spacingLeft=10;" parent="1" vertex="1">
<mxGeometry x="70" y="270" width="340" height="80" as="geometry"/>
</mxCell>
<mxCell id="a4" value="FastAPI(lifespan=lifespan)&#xa;(main.py:147)&#xa;&#xa;on startup → enter lifespan ctx → Section B" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#039be5;fontSize=11;align=left;spacingLeft=10;" parent="1" vertex="1">
<mxGeometry x="70" y="370" width="340" height="90" as="geometry"/>
</mxCell>
<mxCell id="ea12" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#039be5;endArrow=classic;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="a1" target="a2" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="ea23" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#039be5;endArrow=classic;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="a2" target="a3" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="ea34" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#039be5;endArrow=classic;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="a3" target="a4" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="secB" value="B. lifespan() zoom (main.py:116-144)" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#f3e5f5;strokeColor=#8e24aa;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="480" y="70" width="780" height="420" as="geometry"/>
</mxCell>
<mxCell id="b1" value="1. gpu_ids = get_available_gpus()&#xa; respects FASTVIDEO_GPU_COUNT env" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#8e24aa;fontSize=11;align=left;spacingLeft=10;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="510" y="115" width="720" height="45" as="geometry"/>
</mxCell>
<mxCell id="b2" value="2. gpu_pool = GPUPool(gpu_ids)&#xa; creates one GPUSlot per GPU (not yet started)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#8e24aa;fontSize=11;align=left;spacingLeft=10;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="510" y="170" width="720" height="45" as="geometry"/>
</mxCell>
<mxCell id="b3" value="3. await gpu_pool.initialize()&#xa; asyncio.create_task(_init_gpu(id)) for each GPU&#xa; returns immediately — GPUs init in background → Section C" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#fff9c4;strokeColor=#8e24aa;fontSize=11;align=left;spacingLeft=10;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="510" y="225" width="720" height="60" as="geometry"/>
</mxCell>
<mxCell id="b4" value="4. prompt_enhancer = PromptEnhancer()&#xa;5. session_event_logger = SessionEventLogger(...)&#xa;6. prompt_safety_filter = PromptSafetyFilter() (optional)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#8e24aa;fontSize=11;align=left;spacingLeft=10;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="510" y="295" width="720" height="55" as="geometry"/>
</mxCell>
<mxCell id="byield" value="yield&#xa;server is now accepting HTTP + WS&#xa;/healthz → 200 /readyz → 503 until GPUs warm up" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;align=left;spacingLeft=10;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="510" y="360" width="720" height="55" as="geometry"/>
</mxCell>
<mxCell id="bshut" value="[on shutdown] await gpu_pool.shutdown()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffcdd2;strokeColor=#c62828;fontSize=11;align=left;spacingLeft=10;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="510" y="425" width="720" height="35" as="geometry"/>
</mxCell>
<mxCell id="eab" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#666;endArrow=classic;dashed=1;exitX=1;exitY=0.5;exitDx=0;exitDy=0;entryX=0;entryY=0.2;entryDx=0;entryDy=0;" parent="1" source="a4" target="b1" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="secC" value="C. GPUSlot.start() zoom — runs in background per GPU (gpu_pool.py:426)" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#fff3e0;strokeColor=#ef6c00;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="40" y="530" width="1220" height="520" as="geometry"/>
</mxCell>
<mxCell id="cparent" value="Parent (main) process" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#ef6c00;fontSize=12;fontStyle=1;verticalAlign=top;align=left;spacingLeft=10;spacingTop=5;" parent="1" vertex="1">
<mxGeometry x="70" y="580" width="380" height="440" as="geometry"/>
</mxCell>
<mxCell id="cp1" value="mp.get_context(&#39;spawn&#39;)&#xa;cmd_q, resp_q = ctx.Queue(), ctx.Queue()&#xa;(optional) shared_stream_buffer = mp.RawArray(&#39;B&#39;, 256 MiB)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#ef6c00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="85" y="620" width="350" height="60" as="geometry"/>
</mxCell>
<mxCell id="cp2" value="ctx.Process(target=gpu_worker_process,&#xa; args=(gpu_id, cuda_device, cmd_q, resp_q, ...))&#xa;process.start()" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#ef6c00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="85" y="690" width="350" height="60" as="geometry"/>
</mxCell>
<mxCell id="cp3" value="_send_command(Command(INIT), timeout=600s)&#xa; await response on resp_q" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#ef6c00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="85" y="760" width="350" height="50" as="geometry"/>
</mxCell>
<mxCell id="cp4" value="if STARTUP_WARMUP_ENABLED:&#xa; _send_command(Command(WARMUP, prompt=...),&#xa; timeout=STARTUP_WARMUP_TIMEOUT)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#ef6c00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="85" y="820" width="350" height="60" as="geometry"/>
</mxCell>
<mxCell id="cp5" value="slot.ready = True" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="85" y="890" width="350" height="35" as="geometry"/>
</mxCell>
<mxCell id="cp6" value="parent now enters pool as a ready slot;&#xa;/readyz flips to 200 once ≥1 slot is ready" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#888;fontSize=10;align=left;spacingLeft=8;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="85" y="935" width="350" height="45" as="geometry"/>
</mxCell>
<mxCell id="cipc" value="IPC" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#ffcdd2;strokeColor=#b71c1c;fontSize=12;fontStyle=1;verticalAlign=top;align=center;spacingTop=5;" parent="1" vertex="1">
<mxGeometry x="480" y="580" width="150" height="440" as="geometry"/>
</mxCell>
<mxCell id="ccmdq" value="command_queue&#xa;(mp.Queue)" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=10;" parent="1" vertex="1">
<mxGeometry x="490" y="650" width="130" height="80" as="geometry"/>
</mxCell>
<mxCell id="crespq" value="response_queue&#xa;(mp.Queue)" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=10;" parent="1" vertex="1">
<mxGeometry x="490" y="760" width="130" height="80" as="geometry"/>
</mxCell>
<mxCell id="cshared" value="shared_stream_buffer&#xa;(mp.RawArray)" style="shape=cylinder3;whiteSpace=wrap;html=1;boundedLbl=1;backgroundOutline=1;size=15;fillColor=#f8cecc;strokeColor=#b85450;fontSize=10;" parent="1" vertex="1">
<mxGeometry x="490" y="870" width="130" height="80" as="geometry"/>
</mxCell>
<mxCell id="cworker" value="Worker subprocess (gpu_worker_process, gpu_pool.py:151)" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#fef1e0;strokeColor=#d79b00;fontSize=12;fontStyle=1;verticalAlign=top;align=left;spacingLeft=10;spacingTop=5;" parent="1" vertex="1">
<mxGeometry x="660" y="580" width="580" height="440" as="geometry"/>
</mxCell>
<mxCell id="cw1" value="os.environ[&#39;CUDA_VISIBLE_DEVICES&#39;] = cuda_device&#xa;os.environ[&#39;FASTVIDEO_ATTENTION_BACKEND&#39;] = &#39;FLASH_ATTN&#39;&#xa;&#xa;must be set BEFORE importing torch" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="620" width="550" height="65" as="geometry"/>
</mxCell>
<mxCell id="cw2" value="from fastvideo.entrypoints.video_generator import VideoGenerator&#xa;from fastvideo.models.dits.ltx2 import DEFAULT_LTX2_AUDIO_*&#xa;&#xa;** Dreamverse reaches into fastvideo internals here **" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="695" width="550" height="60" as="geometry"/>
</mxCell>
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (video_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="765" width="550" height="95" as="geometry"/>
</mxCell>
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (video_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="870" width="550" height="55" as="geometry"/>
</mxCell>
<mxCell id="cw5" value="enter main worker loop → waits for JOIN_USER / USER_STEP / LEAVE" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="935" width="550" height="40" as="geometry"/>
</mxCell>
<mxCell id="cea1" value="Command" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=10;exitX=1;exitY=0.5;exitDx=0;exitDy=0;entryX=0;entryY=0.5;entryDx=0;entryDy=0;" parent="1" source="cp3" target="ccmdq" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="cea2" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;exitX=1;exitY=0.5;exitDx=0;exitDy=0;" parent="1" source="ccmdq" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="675" y="780" as="targetPoint"/>
<Array as="points">
<mxPoint x="648" y="690"/>
<mxPoint x="648" y="780"/>
<mxPoint x="675" y="780"/>
</Array>
</mxGeometry>
</mxCell>
<mxCell id="cea3" value="WorkerEvent" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;fontSize=10;exitX=0;exitY=0.5;exitDx=0;exitDy=0;entryX=0.969;entryY=0.798;entryDx=0;entryDy=0;entryPerimeter=0;" parent="1" target="crespq" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="684" y="825.5" as="sourcePoint"/>
<mxPoint x="629" y="813" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="cea4" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#b85450;endArrow=classic;exitX=0;exitY=0.5;exitDx=0;exitDy=0;entryX=0.995;entryY=0.887;entryDx=0;entryDy=0;entryPerimeter=0;" parent="1" target="cp3" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="490" y="804" as="sourcePoint"/>
<mxPoint x="435" y="789" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="secD" value="D. WebSocket session — after startup" style="rounded=0;whiteSpace=wrap;html=1;fillColor=#e8f5e9;strokeColor=#2e7d32;strokeWidth=2;verticalAlign=top;align=left;spacingLeft=15;spacingTop=8;fontSize=13;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="40" y="1080" width="1220" height="820" as="geometry"/>
</mxCell>
<mxCell id="dc_head" value="Browser / WS" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#dae8fc;strokeColor=#6c8ebf;fontSize=11;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="80" y="1130" width="160" height="40" as="geometry"/>
</mxCell>
<mxCell id="ds_head" value="main.py / Session" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=11;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="290" y="1130" width="180" height="40" as="geometry"/>
</mxCell>
<mxCell id="dp_head" value="GPUPool" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#d5e8d4;strokeColor=#82b366;fontSize=11;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="520" y="1130" width="160" height="40" as="geometry"/>
</mxCell>
<mxCell id="dsl_head" value="GPUSlot" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e6f2e0;strokeColor=#82b366;fontSize=11;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="730" y="1130" width="160" height="40" as="geometry"/>
</mxCell>
<mxCell id="dw_head" value="Worker subprocess" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#fef1e0;strokeColor=#d79b00;fontSize=11;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="940" y="1130" width="200" height="40" as="geometry"/>
</mxCell>
<mxCell id="dc_life" style="endArrow=none;html=1;strokeColor=#6c8ebf;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="160" y="1175" as="sourcePoint"/>
<mxPoint x="160" y="1885" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="ds_life" style="endArrow=none;html=1;strokeColor=#82b366;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="380" y="1175" as="sourcePoint"/>
<mxPoint x="380" y="1885" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dp_life" style="endArrow=none;html=1;strokeColor=#82b366;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="600" y="1175" as="sourcePoint"/>
<mxPoint x="600" y="1885" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dsl_life" style="endArrow=none;html=1;strokeColor=#82b366;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="810" y="1175" as="sourcePoint"/>
<mxPoint x="810" y="1885" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dw_life" style="endArrow=none;html=1;strokeColor=#d79b00;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="1040" y="1175" as="sourcePoint"/>
<mxPoint x="1040" y="1885" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm1" value="1. WS connect /ws + session_init_v2" style="endArrow=classic;html=1;strokeColor=#6c8ebf;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="160" y="1200" as="sourcePoint"/>
<mxPoint x="380" y="1200" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm2" value="2. pool.acquire(client_id, ws)" style="endArrow=classic;html=1;strokeColor=#82b366;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="380" y="1240" as="sourcePoint"/>
<mxPoint x="600" y="1240" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm3a" value="3a. if slot free: (gpu_id, slot)" style="endArrow=classic;html=1;strokeColor=#82b366;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="600" y="1280" as="sourcePoint"/>
<mxPoint x="380" y="1280" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm3b" value="3b. else: enqueue, emit queue_status until a slot frees" style="endArrow=classic;html=1;strokeColor=#c62828;fontSize=10;dashed=1;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="600" y="1320" as="sourcePoint"/>
<mxPoint x="160" y="1320" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm4" value="4. emit gpu_assigned" style="endArrow=classic;html=1;strokeColor=#6c8ebf;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="380" y="1360" as="sourcePoint"/>
<mxPoint x="160" y="1360" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm5" value="5. slot._send_command_tagged(Command(USER_JOIN))" style="endArrow=classic;html=1;strokeColor=#82b366;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="380" y="1400" as="sourcePoint"/>
<mxPoint x="810" y="1400" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm6" value="6. command_queue.put(JOIN_USER)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="810" y="1440" as="sourcePoint"/>
<mxPoint x="1040" y="1440" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dssteady" value="── steady state: per prompt ──" style="text;html=1;align=center;fontSize=11;fontColor=#666;fontStyle=2;" parent="1" vertex="1">
<mxGeometry x="300" y="1470" width="600" height="25" as="geometry"/>
</mxCell>
<mxCell id="dm7" value="7. append_prompt / simple_generate" style="endArrow=classic;html=1;strokeColor=#6c8ebf;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="160" y="1510" as="sourcePoint"/>
<mxPoint x="380" y="1510" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm9" value="8. slot._send_command_tagged(Command(USER_STEP, payload=UserStepPayload(prompt, segment_idx, image_path, reset_conditioning)))" style="endArrow=classic;html=1;strokeColor=#82b366;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="380" y="1560" as="sourcePoint"/>
<mxPoint x="810" y="1560" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm10" value="9. command_queue.put(USER_STEP)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="810" y="1610" as="sourcePoint"/>
<mxPoint x="1040" y="1610" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (video_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="955" y="1640" width="180" height="70" as="geometry"/>
</mxCell>
<mxCell id="dm11" value="10b. resp_q.put(MediaInit / MediaChunk / MediaComplete / StepComplete)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
<mxGeometry x="0.0435" relative="1" as="geometry">
<mxPoint x="1092" y="1733" as="sourcePoint"/>
<mxPoint x="862" y="1733" as="targetPoint"/>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="dm12" value="11. _response_reader isinstance-dispatches: MediaChunk/etc → _stream_queues[user_id]; StepComplete → _pending_futures" style="endArrow=classic;html=1;strokeColor=#d6b656;fontSize=10;" parent="1" edge="1">
<mxGeometry x="0.0014" relative="1" as="geometry">
<mxPoint x="810" y="1760" as="sourcePoint"/>
<mxPoint x="380" y="1760" as="targetPoint"/>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="dm13a" value="12a. binary WS frame (mp4 bytes)" style="endArrow=classic;html=1;strokeColor=#6c8ebf;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
<mxGeometry x="-0.0909" relative="1" as="geometry">
<mxPoint x="380" y="1795" as="sourcePoint"/>
<mxPoint x="160" y="1795" as="targetPoint"/>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="dm13b" value="12b. JSON WS event (media_init, media_segment_complete, prompt_ready, …)" style="endArrow=classic;html=1;strokeColor=#6c8ebf;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
<mxGeometry x="-0.0909" y="5" relative="1" as="geometry">
<mxPoint x="380" y="1825" as="sourcePoint"/>
<mxPoint x="160" y="1825" as="targetPoint"/>
<mxPoint as="offset"/>
</mxGeometry>
</mxCell>
<mxCell id="dm14" value="13. on WS close: pool.release(client_id) → Command(USER_LEAVE) → LeaveAck" style="endArrow=classic;html=1;strokeColor=#c62828;fontSize=10;dashed=1;labelBackgroundColor=#ffffff;" parent="1" edge="1">
<mxGeometry relative="1" as="geometry">
<mxPoint x="160" y="1865" as="sourcePoint"/>
<mxPoint x="1040" y="1865" as="targetPoint"/>
</mxGeometry>
</mxCell>
</root>
</mxGraphModel>
</diagram>
</mxfile>
File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 85 KiB

+41
View File
@@ -0,0 +1,41 @@
[project]
name = "dreamverse"
version = "0.1.0"
description = "Dreamverse application with FastAPI backend and Next.js frontend."
readme = "README.md"
requires-python = ">=3.10"
dependencies = [
"fastapi>=0.115,<1",
"fastvideo>=0.1.7",
"pillow>=10.0",
"uvicorn[standard]>=0.30,<1",
]
[project.optional-dependencies]
server = [
"cerebras-cloud-sdk",
"flash-attn-cute @ git+https://github.com/XOR-op/flash-attention.git@fa4-compile#subdirectory=flash_attn/cute",
"flashinfer-python",
"openai>=1.40",
]
safety = [
"fasttext>=0.9.2",
]
test = [
"pytest>=8.0",
]
[project.scripts]
dreamverse-server = "dreamverse.server_entry:cli"
dreamverse-mock-server = "dreamverse.mock_server:cli"
[tool.uv]
package = false
[tool.uv.sources]
fastvideo = { workspace = true }
[tool.pytest.ini_options]
markers = [
"gpu: requires real GPU + model weights; skip in CI",
]
+272
View File
@@ -0,0 +1,272 @@
#!/usr/bin/env bash
#
# Build & install a perf-optimized ffmpeg (LTO + libx264 + native arch).
# Mirrors the team playbook's flags verbatim per stage, with three
# deliberate deviations made necessary by our build host:
#
# 1. --disable-libxcb --disable-xlib ffmpeg's auto-detect linked
# libxcb at build time → binary
# wouldn't even start at runtime.
# 2. LIBRARY_PATH / LD_LIBRARY_PATH conda-forge gcc wrappers search
# + -L$INSTALL_PREFIX/lib early $CONDA_PREFIX/lib implicitly,
# leaking an older libx264 into
# ffmpeg's link. Forcing our prefix
# first makes the resolver pick
# our just-built lib.
# 3. MAKE_JOBS cap (default 16) very high -j (e.g. nproc=96 on
# NVL72) tripped a race in
# ffmpeg's recursive recipes.
#
# Usage:
# bash scripts/install_native_ffmpeg.sh
#
# Knobs (env vars, all optional):
# INSTALL_PREFIX install destination default: $HOME/opt/ffmpeg-native
# SOURCE_DIR build workspace default: $HOME/src/ffmpeg-native
# X264_REF x264 git ref default: stable
# FFMPEG_REF FFmpeg git ref default: n7.1
# NV_CODEC_REF nv-codec-headers ref default: master
# CUDA_PREFIX CUDA toolkit root default: /usr/local/cuda
# ENABLE_NVENC build with NVENC/NVDEC default: 1 (1|0)
# MAKE_JOBS parallel make jobs default: min(nproc, 16)
# FFMPEG_NATIVE_CC explicit C compiler command for native builds
# FFMPEG_NATIVE_CXX explicit C++ compiler command for native builds
#
# Toolchain selection: CC / CXX / AS are pinned to the conda-forge
# triplet matching `uname -m`. Inherited values are intentionally
# ignored — conda envs that have BOTH `gcc_linux-64` and
# `gcc_linux-aarch64` installed export the cross-compiler triplet on
# every `conda activate` (the `aarch64` activation script sorts later
# and wins), which silently breaks x264's compiler probe on the
# opposite host. If you genuinely need a non-host toolchain, set
# FFMPEG_NATIVE_CC and/or FFMPEG_NATIVE_CXX explicitly.
set -euo pipefail
# ─── Defaults (override via env) ──────────────────────────────────────────
INSTALL_PREFIX="${INSTALL_PREFIX:-$HOME/opt/ffmpeg-native}"
SOURCE_DIR="${SOURCE_DIR:-$HOME/src/ffmpeg-native}"
X264_REF="${X264_REF:-stable}"
FFMPEG_REF="${FFMPEG_REF:-n7.1}"
NV_CODEC_REF="${NV_CODEC_REF:-master}"
CUDA_PREFIX="${CUDA_PREFIX:-/usr/local/cuda}"
ENABLE_NVENC="${ENABLE_NVENC:-1}"
case "${ENABLE_NVENC}" in
0|1) ;;
*) echo "[install_native_ffmpeg] ENABLE_NVENC must be 0 or 1, got '${ENABLE_NVENC}'" >&2; exit 1 ;;
esac
NPROC="$(nproc)"
MAKE_JOBS="${MAKE_JOBS:-$(( NPROC < 16 ? NPROC : 16 ))}"
# ─── Per-platform, per-stage flags (verbatim from the playbook) ───────────
ARCH="$(uname -m)"
case "$ARCH" in
x86_64)
DEFAULT_CC=x86_64-conda-linux-gnu-cc
DEFAULT_CXX=x86_64-conda-linux-gnu-c++
AS=nasm
X264_CFLAGS="-O3 -march=native -mtune=native -fPIC -flto"
X264_LDFLAGS="-flto -fuse-linker-plugin"
FFMPEG_CFLAGS="-O3 -march=native -mtune=native -fPIC -flto"
FFMPEG_LDFLAGS="-flto -Wl,-rpath,$INSTALL_PREFIX/lib"
;;
aarch64)
DEFAULT_CC=aarch64-conda-linux-gnu-cc
DEFAULT_CXX=aarch64-conda-linux-gnu-c++
unset AS # GNU as on ARM
X264_CFLAGS="-O3 -mcpu=native -fPIC -flto"
X264_LDFLAGS="-flto -fuse-linker-plugin"
FFMPEG_CFLAGS="-O3 -mcpu=native -fPIC -flto -fno-tree-vectorize"
FFMPEG_LDFLAGS="-flto -Wl,-rpath,$INSTALL_PREFIX/lib"
;;
*)
echo "[install_native_ffmpeg] unsupported arch: $ARCH" >&2
exit 1
;;
esac
CC="${FFMPEG_NATIVE_CC:-$DEFAULT_CC}"
CXX="${FFMPEG_NATIVE_CXX:-$DEFAULT_CXX}"
require_compiler() {
local name="$1" compiler="$2"
if [[ -z "$compiler" ]]; then
echo "[install_native_ffmpeg] $name is empty" >&2
exit 1
fi
if ! command -v -- "$compiler" >/dev/null 2>&1; then
echo "[install_native_ffmpeg] $name is unavailable: $compiler" >&2
exit 1
fi
if ! "$compiler" --version >/dev/null 2>&1; then
echo "[install_native_ffmpeg] $name failed sanity check: $compiler --version" >&2
exit 1
fi
}
require_compiler CC "$CC"
require_compiler CXX "$CXX"
export CC CXX
[[ -n "${AS:-}" ]] && export AS
echo "[install_native_ffmpeg] toolchain: CC=$CC CXX=$CXX AS=${AS:-<gnu-as>} (uname -m=$ARCH)"
# ─── Step 0: probe required tools ─────────────────────────────────────────
required=("$CC" "$CXX" make pkg-config git)
[[ "$(uname -m)" == "x86_64" ]] && required+=(nasm)
missing=()
for cmd in "${required[@]}"; do
command -v -- "$cmd" >/dev/null 2>&1 || missing+=("$cmd")
done
if (( ${#missing[@]} > 0 )); then
echo "[install_native_ffmpeg] missing required tools: ${missing[*]}" >&2
echo "[install_native_ffmpeg] install them, or set FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX explicitly." >&2
exit 1
fi
# ─── Step 0.5: destructive-path guards ────────────────────────────────────
guard_path() {
local name="$1" value="$2"
case "$value" in
"") echo "[install_native_ffmpeg] $name is empty" >&2; exit 1 ;;
"/") echo "[install_native_ffmpeg] refusing to wipe '/'" >&2; exit 1 ;;
"$HOME") echo "[install_native_ffmpeg] refusing to wipe \$HOME" >&2; exit 1 ;;
esac
[[ "$value" == /* ]] || {
echo "[install_native_ffmpeg] $name must be absolute, got: $value" >&2; exit 1; }
[[ "$value" == *ffmpeg-native* ]] || {
echo "[install_native_ffmpeg] $name must contain 'ffmpeg-native' for safety, got: $value" >&2
exit 1; }
}
guard_path INSTALL_PREFIX "$INSTALL_PREFIX"
guard_path SOURCE_DIR "$SOURCE_DIR"
# ─── Step 1: clean ────────────────────────────────────────────────────────
echo "[install_native_ffmpeg] cleaning prior install + sources"
rm -rf "$INSTALL_PREFIX" "$SOURCE_DIR/x264" "$SOURCE_DIR/ffmpeg" "$SOURCE_DIR/nv-codec-headers"
mkdir -p "$SOURCE_DIR" "$INSTALL_PREFIX/lib"
if [[ "$ENABLE_NVENC" == "1" ]]; then
if [[ ! -f "$CUDA_PREFIX/include/cuda.h" ]]; then
echo "[install_native_ffmpeg] ENABLE_NVENC=1 but cuda.h not at $CUDA_PREFIX/include/cuda.h" >&2
echo "[install_native_ffmpeg] set CUDA_PREFIX env to your CUDA toolkit root, or ENABLE_NVENC=0 to skip" >&2
exit 1
fi
echo "[install_native_ffmpeg] NVENC build enabled (CUDA_PREFIX=$CUDA_PREFIX)"
fi
# Deviation 2: force our prefix to win over conda-forge gcc's implicit
# library search (otherwise an older libx264 from conda's lib dir leaks
# into ffmpeg's link, producing a binary that needs *two* x264 SONAMEs).
export LIBRARY_PATH="$INSTALL_PREFIX/lib${LIBRARY_PATH:+:$LIBRARY_PATH}"
export LD_LIBRARY_PATH="$INSTALL_PREFIX/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
# ─── Step 2: build x264 (LTO, shared lib, no CLI) ─────────────────────────
echo "[install_native_ffmpeg] cloning x264 ($X264_REF)"
git clone --depth 1 --branch "$X264_REF" \
https://code.videolan.org/videolan/x264.git "$SOURCE_DIR/x264"
(
cd "$SOURCE_DIR/x264"
export CFLAGS="$X264_CFLAGS"
export CXXFLAGS="$X264_CFLAGS"
export LDFLAGS="$X264_LDFLAGS"
echo "[install_native_ffmpeg] configuring x264"
./configure --prefix="$INSTALL_PREFIX" --enable-shared --enable-pic --disable-cli
echo "[install_native_ffmpeg] building x264 (-j$MAKE_JOBS)"
make -j"$MAKE_JOBS"
make install
)
# ─── Step 2.5: nv-codec-headers (NVENC/NVDEC API headers, no CUDA libs) ──
# Required for ffmpeg's --enable-cuda --enable-nvenc --enable-cuvid configure
# flags. Installs ffnvcodec.pc + headers into $INSTALL_PREFIX so ffmpeg's
# pkg-config picks them up alongside libx264. The runtime libraries
# (libcuda.so, libnvcuvid.so, libnvidia-encode.so) come from the NVIDIA
# driver, not from these headers.
if [[ "$ENABLE_NVENC" == "1" ]]; then
echo "[install_native_ffmpeg] cloning nv-codec-headers ($NV_CODEC_REF)"
git clone --depth 1 --branch "$NV_CODEC_REF" \
https://git.videolan.org/git/ffmpeg/nv-codec-headers.git \
"$SOURCE_DIR/nv-codec-headers"
(
cd "$SOURCE_DIR/nv-codec-headers"
make PREFIX="$INSTALL_PREFIX" install
)
fi
# ─── Step 3: build ffmpeg (LTO, libx264, shared, optional NVENC) ──────────
echo "[install_native_ffmpeg] cloning FFmpeg ($FFMPEG_REF)"
git clone --depth 1 --branch "$FFMPEG_REF" \
https://github.com/FFmpeg/FFmpeg.git "$SOURCE_DIR/ffmpeg"
(
cd "$SOURCE_DIR/ffmpeg"
export PKG_CONFIG_PATH="$INSTALL_PREFIX/lib/pkgconfig"
export CFLAGS="$FFMPEG_CFLAGS"
export CXXFLAGS="$FFMPEG_CFLAGS"
export LDFLAGS="$FFMPEG_LDFLAGS"
[[ "$(uname -m)" == "x86_64" ]] && which nasm
pkg-config --modversion x264
echo "[install_native_ffmpeg] configuring ffmpeg"
ffmpeg_configure_flags=(
--prefix="$INSTALL_PREFIX"
--enable-gpl
--enable-libx264
--enable-lto
--enable-shared
--disable-static
--disable-debug
--disable-doc
--disable-ffplay
--disable-libxcb
--disable-xlib
--extra-cflags="$CFLAGS"
--extra-cxxflags="$CXXFLAGS"
--extra-ldflags="-L$INSTALL_PREFIX/lib $LDFLAGS"
)
if [[ "$ENABLE_NVENC" == "1" ]]; then
pkg-config --modversion ffnvcodec
ffmpeg_configure_flags+=(
--enable-cuda
--enable-nvenc
--enable-cuvid
--enable-nvdec
--extra-cflags="-I$CUDA_PREFIX/include"
--extra-ldflags="-L$CUDA_PREFIX/lib64"
)
fi
./configure "${ffmpeg_configure_flags[@]}"
echo "[install_native_ffmpeg] building ffmpeg (-j$MAKE_JOBS)"
make -j"$MAKE_JOBS"
make install
)
# ─── Step 4: sanity check ──────────────────────────────────────────────────
ffmpeg_bin="$INSTALL_PREFIX/bin/ffmpeg"
echo "[install_native_ffmpeg] verifying $ffmpeg_bin"
"$ffmpeg_bin" -hide_banner -buildconf | grep -i -E 'libx264|lto'
"$ffmpeg_bin" -hide_banner -encoders | grep -i libx264
"$ffmpeg_bin" -hide_banner -h encoder=libx264 2>&1 | grep -i preset
if [[ "$ENABLE_NVENC" == "1" ]]; then
"$ffmpeg_bin" -hide_banner -encoders | grep -E 'h264_nvenc|hevc_nvenc' || {
echo "[install_native_ffmpeg] ENABLE_NVENC=1 but built ffmpeg has no h264_nvenc/hevc_nvenc encoder" >&2
exit 1
}
fi
# ─── Step 5: emit env file ─────────────────────────────────────────────────
# FASTVIDEO_VIDEO_CODEC stays at libx264 by default to preserve runtime
# behavior; NVENC is selected at deploy time via dreamverse-deploy.sh
# --nvenc, which exports FASTVIDEO_VIDEO_CODEC=h264_nvenc.
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
env_file="$script_dir/ffmpeg-env.sh"
{
printf '#!/usr/bin/env bash\n'
printf 'export FASTVIDEO_FFMPEG_BIN=%q\n' "$ffmpeg_bin"
printf 'export FASTVIDEO_VIDEO_CODEC=libx264\n'
} > "$env_file"
chmod +x "$env_file"
echo
echo "[install_native_ffmpeg] ✓ done."
echo "[install_native_ffmpeg] binary: $ffmpeg_bin"
echo "[install_native_ffmpeg] env: $env_file"
echo "[install_native_ffmpeg] source it before running the demo:"
echo "[install_native_ffmpeg] source scripts/ffmpeg-env.sh"
+36
View File
@@ -0,0 +1,36 @@
# Dreamverse Launch Scripts
These scripts are convenience wrappers for local demos and lower-level backend
checks. The main Dreamverse README documents the normal manual startup path.
## One-Command Demo
From the FastVideo checkout:
```bash
apps/dreamverse/scripts/launch/launch_demo.sh
```
The launcher starts the Dreamverse backend and frontend, polls readiness, and
prints the active URLs. It defaults to `dreamverse-server` on backend port
`8009` and frontend port `5274`.
Useful overrides:
```bash
BE_PORT=8010 FE_PORT=5274 apps/dreamverse/scripts/launch/launch_demo.sh
NO_FRONTEND=1 apps/dreamverse/scripts/launch/launch_demo.sh
NO_BROWSER=1 apps/dreamverse/scripts/launch/launch_demo.sh
```
## Individual Scripts
`launch_backend_dreamverse.sh` starts the full Dreamverse backend path used by
the web app.
`launch_frontend.sh` starts the Next.js frontend and installs `pnpm`
dependencies when `node_modules/` is missing.
`launch_backend_fastvideo.sh` starts the typed `fastvideo serve --config` path
for lower-level serve-config checks. It is not a full Dreamverse app replacement
because the current frontend still depends on Dreamverse-specific routes.
@@ -0,0 +1,51 @@
#!/usr/bin/env bash
# Launch the Dreamverse-flavored backend via the installed console command.
#
# This is the path the Next.js frontend expects today — it serves
# ``/healthz``, ``/readyz``, ``/status``, ``/curated-presets``, and the
# devtools-only routes that ``fastvideo serve`` does not. Use this for
# the full demo experience until the FE-only routes migrate into
# ``fastvideo.entrypoints.streaming.server.build_app``.
#
# Usage:
# bash launch_backend_dreamverse.sh # default 0.0.0.0:8009
# bash launch_backend_dreamverse.sh --port 8010
#
# Mirrors internal/ui defaults via the same env variables internal's
# ``config.py`` reads (see ../../../serve_configs/streaming_demo.yaml
# for the canonical list and source-of-truth comments).
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/../.." && pwd)"
if [[ -f "${HOME}/.env" ]]; then
set -o allexport
# shellcheck disable=SC1091
source "${HOME}/.env"
set +o allexport
fi
# Internal/ui parity defaults — only set if the caller hasn't already
# pinned them in ~/.env or the surrounding environment. Each value
# matches FastVideo-internal/ui/ltx2-streaming/server/config.py.
export FASTVIDEO_ATTENTION_BACKEND="${FASTVIDEO_ATTENTION_BACKEND:-FLASH_ATTN}"
export STREAM_MODE="${STREAM_MODE:-av_fmp4}"
export ENABLE_TORCH_COMPILE="${ENABLE_TORCH_COMPILE:-1}"
export FASTVIDEO_ENABLE_STARTUP_WARMUP="${FASTVIDEO_ENABLE_STARTUP_WARMUP:-1}"
export FASTVIDEO_STARTUP_WARMUP_TIMEOUT_SECONDS="${FASTVIDEO_STARTUP_WARMUP_TIMEOUT_SECONDS:-2400}"
export FASTVIDEO_GENERATION_SEGMENT_CAP="${FASTVIDEO_GENERATION_SEGMENT_CAP:-6}"
export FASTVIDEO_PROMPT_AUTO_SLEEP_MS="${FASTVIDEO_PROMPT_AUTO_SLEEP_MS:-120}"
export FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS="${FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS:-1800}"
cd "${DREAMVERSE_ROOT}"
if ! command -v dreamverse-server >/dev/null 2>&1; then
echo "error: dreamverse-server not found on PATH. Install FastVideo with the dreamverse extra." >&2
exit 1
fi
echo "[launch-demo] starting dreamverse-server"
echo " args: $*"
exec dreamverse-server "$@"
+50
View File
@@ -0,0 +1,50 @@
#!/usr/bin/env bash
# Launch the FastVideo streaming backend via the typed
# ``fastvideo serve --config`` entrypoint, driven by
# ``serve_configs/streaming_demo.yaml``.
#
# Usage:
# bash launch_backend_fastvideo.sh
# bash launch_backend_fastvideo.sh --server.port 8010 --streaming.warmup.enabled false
#
# Anything passed after the script name is forwarded verbatim to
# ``fastvideo serve``, so dotted overrides like
# ``--server.port 8010`` or ``--streaming.warmup.enabled false`` work
# without a parallel flag scheme.
#
# Caveat: the bare ``fastvideo serve`` build_app exposes ``/health``
# and ``/v1/stream`` only. Dreamverse's existing Next.js shell also
# expects ``/healthz``, ``/readyz``, and ``/curated-presets`` which
# live in dreamverse-server. Use ``launch_backend_dreamverse.sh`` for
# the full FE compatibility path; this script is for verifying the
# typed serve config and bare-streaming-server flow.
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/../.." && pwd)"
SERVE_CONFIG="${DREAMVERSE_ROOT}/serve_configs/streaming_demo.yaml"
if [[ ! -f "${SERVE_CONFIG}" ]]; then
echo "error: serve config not found at ${SERVE_CONFIG}" >&2
exit 1
fi
# Source ~/.env if present so prompt-enhancer credentials
# (CEREBRAS_API_KEY etc.) are visible to the worker process.
if [[ -f "${HOME}/.env" ]]; then
set -o allexport
# shellcheck disable=SC1091
source "${HOME}/.env"
set +o allexport
fi
# Match internal/ui's attention backend default (gpu_pool.py:161).
export FASTVIDEO_ATTENTION_BACKEND="${FASTVIDEO_ATTENTION_BACKEND:-FLASH_ATTN}"
cd "${DREAMVERSE_ROOT}"
echo "[launch-demo] starting fastvideo serve"
echo " config: ${SERVE_CONFIG}"
echo " overrides: $*"
exec uv run fastvideo serve --config "${SERVE_CONFIG}" "$@"
+145
View File
@@ -0,0 +1,145 @@
#!/usr/bin/env bash
# One-command Dreamverse demo launcher.
#
# Spawns the backend and frontend, polls health, prints URLs, and
# stops both children on Ctrl-C.
#
# Defaults match internal/ui:
# * BE = dreamverse-server (8009) — full FE compatibility
# * FE = Next.js dev:devtools (5274)
#
# Switch BE to fastvideo serve --config (typed-only path):
# BE_FLAVOR=fastvideo bash launch_demo.sh
#
# Other env knobs:
# BE_PORT=8010 FE_PORT=5274 bash launch_demo.sh
# NO_FRONTEND=1 bash launch_demo.sh # backend only
# NO_BROWSER=1 bash launch_demo.sh # skip xdg-open
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/../.." && pwd)"
LOG_DIR="${DREAMVERSE_ROOT}/logs"
mkdir -p "${LOG_DIR}"
BE_FLAVOR="${BE_FLAVOR:-dreamverse}"
BE_PORT="${BE_PORT:-8009}"
FE_PORT="${FE_PORT:-5274}"
HEALTH_TIMEOUT_SECONDS="${HEALTH_TIMEOUT_SECONDS:-300}"
READY_TIMEOUT_SECONDS="${READY_TIMEOUT_SECONDS:-2400}"
case "${BE_FLAVOR}" in
dreamverse)
BE_SCRIPT="${SCRIPT_DIR}/launch_backend_dreamverse.sh"
BE_PORT_ARGS=(--port "${BE_PORT}")
HEALTH_PATH="/healthz"
READY_PATH="/readyz"
;;
fastvideo)
BE_SCRIPT="${SCRIPT_DIR}/launch_backend_fastvideo.sh"
# fastvideo serve takes dotted overrides on top of the YAML; the
# nested ServerConfig path is server.port not --port.
BE_PORT_ARGS=(--server.port "${BE_PORT}")
# bare fastvideo serve only exposes /health; treat it as both the
# liveness and readiness probe.
HEALTH_PATH="/health"
READY_PATH="/health"
;;
*)
echo "error: BE_FLAVOR must be 'dreamverse' or 'fastvideo' (got '${BE_FLAVOR}')" >&2
exit 1
;;
esac
BE_URL="http://localhost:${BE_PORT}"
FE_URL="http://localhost:${FE_PORT}"
BE_LOG="${LOG_DIR}/demo-be.log"
FE_LOG="${LOG_DIR}/demo-fe.log"
cleanup() {
local rc=$?
echo
echo "[launch-demo] shutting down children"
if [[ -n "${BE_PID:-}" ]] && kill -0 "${BE_PID}" 2>/dev/null; then
kill "${BE_PID}" 2>/dev/null || true
fi
if [[ -n "${FE_PID:-}" ]] && kill -0 "${FE_PID}" 2>/dev/null; then
kill "${FE_PID}" 2>/dev/null || true
fi
wait 2>/dev/null || true
exit "${rc}"
}
trap cleanup EXIT INT TERM
probe() {
local url="$1"
curl -fsS -m 2 -o /dev/null "${url}"
}
wait_for() {
local url="$1"
local label="$2"
local timeout="$3"
local started end
started="$(date +%s)"
end="$((started + timeout))"
while (( $(date +%s) < end )); do
if probe "${url}"; then
echo "[launch-demo] ${label} ready: ${url}"
return 0
fi
sleep 2
done
echo "error: ${label} did not respond within ${timeout}s at ${url}" >&2
return 1
}
# --- backend ---
echo "[launch-demo] backend logs: ${BE_LOG}"
( "${BE_SCRIPT}" "${BE_PORT_ARGS[@]}" >"${BE_LOG}" 2>&1 ) &
BE_PID=$!
echo "[launch-demo] backend PID ${BE_PID} (flavor=${BE_FLAVOR}, port=${BE_PORT})"
if ! wait_for "${BE_URL}${HEALTH_PATH}" "backend health" "${HEALTH_TIMEOUT_SECONDS}"; then
echo "------ backend log (tail) ------" >&2
tail -n 60 "${BE_LOG}" >&2 || true
exit 1
fi
if [[ "${HEALTH_PATH}" != "${READY_PATH}" ]]; then
echo "[launch-demo] waiting for backend readiness (warmup may take a few minutes)…"
if ! wait_for "${BE_URL}${READY_PATH}" "backend ready" "${READY_TIMEOUT_SECONDS}"; then
echo "------ backend log (tail) ------" >&2
tail -n 100 "${BE_LOG}" >&2 || true
exit 1
fi
fi
# --- frontend ---
if [[ "${NO_FRONTEND:-0}" != "1" ]]; then
echo "[launch-demo] frontend logs: ${FE_LOG}"
( "${SCRIPT_DIR}/launch_frontend.sh" >"${FE_LOG}" 2>&1 ) &
FE_PID=$!
echo "[launch-demo] frontend PID ${FE_PID} (port ${FE_PORT})"
if ! wait_for "${FE_URL}/" "frontend" 60; then
echo "------ frontend log (tail) ------" >&2
tail -n 60 "${FE_LOG}" >&2 || true
exit 1
fi
if [[ "${NO_BROWSER:-0}" != "1" ]] && command -v xdg-open >/dev/null 2>&1; then
xdg-open "${FE_URL}/" >/dev/null 2>&1 || true
fi
fi
cat <<EOF
[launch-demo] stack is up. Press Ctrl-C to stop.
backend : ${BE_URL} (${BE_FLAVOR}, log: ${BE_LOG})
frontend: ${FE_URL} (log: ${FE_LOG})
EOF
wait "${BE_PID}"
+54
View File
@@ -0,0 +1,54 @@
#!/usr/bin/env bash
# Launch the Dreamverse Next.js frontend in dev mode on port 5274
# (the devtools-enabled build the e2e tests target).
#
# Usage:
# bash launch_frontend.sh
# FRONTEND_MODE=dev bash launch_frontend.sh # plain dev (5299, no devtools)
# FRONTEND_MODE=single5s bash launch_frontend.sh # single-5s product mode
#
# The script ``cd``'s into ``web`` and shells out to pnpm. It runs
# ``pnpm install --frozen-lockfile`` only when ``node_modules/`` is missing so repeat
# launches are fast.
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/../.." && pwd)"
WEB_ROOT="${DREAMVERSE_ROOT}/web"
FRONTEND_MODE="${FRONTEND_MODE:-devtools}"
case "${FRONTEND_MODE}" in
devtools|dev|single5s) ;;
*)
echo "error: FRONTEND_MODE must be one of devtools|dev|single5s (got '${FRONTEND_MODE}')" >&2
exit 1
;;
esac
if [[ ! -d "${WEB_ROOT}" ]]; then
echo "error: web app not found at ${WEB_ROOT}" >&2
exit 1
fi
cd "${WEB_ROOT}"
if [[ ! -d node_modules ]]; then
echo "[launch-demo] node_modules missing — running pnpm install --frozen-lockfile"
pnpm install --frozen-lockfile
fi
case "${FRONTEND_MODE}" in
devtools)
echo "[launch-demo] starting Next.js dev:devtools (port 5274)"
exec pnpm run dev:devtools -- "$@"
;;
dev)
echo "[launch-demo] starting Next.js dev (port 5299)"
exec pnpm run dev -- "$@"
;;
single5s)
echo "[launch-demo] starting Next.js dev:single5s (port 5274)"
exec pnpm run dev:single5s -- "$@"
;;
esac
+95
View File
@@ -0,0 +1,95 @@
#!/usr/bin/env bash
set -euo pipefail
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOST="${BACKEND_HOST:-127.0.0.1}"
PORT="${BACKEND_PORT:-8009}"
BACKEND_ORIGIN="http://${HOST}:${PORT}"
TIMEOUT_SECONDS="${DREAMVERSE_SMOKE_TIMEOUT_SECONDS:-240}"
POLL_INTERVAL_SECONDS="${DREAMVERSE_SMOKE_POLL_SECONDS:-2}"
BACKEND_LOG_PATH="${DREAMVERSE_SMOKE_LOG_PATH:-${ROOT_DIR}/outputs/smoke-local-backend.log}"
START_BACKEND="${DREAMVERSE_SMOKE_START_BACKEND:-1}"
backend_pid=""
started_backend=0
cleanup() {
if [[ "${started_backend}" == "1" && -n "${backend_pid}" ]]; then
kill "${backend_pid}" >/dev/null 2>&1 || true
wait "${backend_pid}" >/dev/null 2>&1 || true
fi
}
trap cleanup EXIT
require_command() {
if ! command -v "$1" >/dev/null 2>&1; then
echo "Missing required command: $1" >&2
exit 1
fi
}
probe_json() {
local path="$1"
curl --silent --show-error --fail "${BACKEND_ORIGIN}${path}"
}
wait_for_endpoint() {
local path="$1"
local label="$2"
local deadline=$((SECONDS + TIMEOUT_SECONDS))
while (( SECONDS < deadline )); do
if probe_json "${path}" >/dev/null 2>&1; then
return 0
fi
sleep "${POLL_INTERVAL_SECONDS}"
done
echo "Timed out waiting for ${label} at ${BACKEND_ORIGIN}${path}" >&2
return 1
}
require_command curl
mkdir -p "$(dirname "${BACKEND_LOG_PATH}")"
if ! probe_json "/healthz" >/dev/null 2>&1; then
if [[ "${START_BACKEND}" != "1" ]]; then
echo "Dreamverse backend is not reachable at ${BACKEND_ORIGIN} and auto-start is disabled." >&2
exit 1
fi
require_command dreamverse-server
echo "Starting Dreamverse backend on ${BACKEND_ORIGIN}..."
(
cd "${ROOT_DIR}"
exec dreamverse-server --host "${HOST}" --port "${PORT}"
) >"${BACKEND_LOG_PATH}" 2>&1 &
backend_pid=$!
started_backend=1
else
echo "Dreamverse backend already running on ${BACKEND_ORIGIN}."
fi
echo "Waiting for /healthz..."
wait_for_endpoint "/healthz" "healthz"
echo "Waiting for /readyz..."
wait_for_endpoint "/readyz" "readyz"
status_payload="$(probe_json "/status")"
echo "Dreamverse local smoke check passed."
echo "Backend URL: ${BACKEND_ORIGIN}"
echo "Status: ${status_payload}"
if [[ "${started_backend}" == "1" ]]; then
trap - EXIT
echo "Backend PID: ${backend_pid}"
echo "Backend log: ${BACKEND_LOG_PATH}"
echo "Backend is still running so you can launch the frontend:"
echo " cd ${ROOT_DIR}/web"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
else
echo "You can now launch the frontend:"
echo " cd ${ROOT_DIR}/web"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
fi
@@ -0,0 +1,153 @@
# Dreamverse streaming demo — fastvideo serve --config target.
#
# This config is for testing Dreamverse-like streaming through the lower-level
# `fastvideo serve --config` entrypoint. Use `dreamverse-server` for the full
# Dreamverse web app.
#
# Mirrors ../FastVideo-internal/ui/ltx2-streaming/server/config.py
# defaults so the public typed surface produces the same runtime
# behavior as the internal UI's GPU pool. Sources of each setting are
# annotated inline; deviations are explicit.
#
# Boot: uv run fastvideo serve --config serve_configs/streaming_demo.yaml
# Override at the CLI:
# uv run fastvideo serve --config <yaml> \
# --server.port 8010 \
# --streaming.warmup.enabled false
#
# Required environment (sourced from ~/.env when launched via the
# launch-demo skill):
# * CEREBRAS_API_KEY (prompt enhancement, default provider)
# * CEREBRAS_IFM_API_KEY (only required if you set
# streaming.prompt.provider to cerebras_ifm
# via dotted override; the public typed
# schema currently only accepts cerebras /
# groq, so cerebras_ifm flows via env on
# dreamverse-server, not fastvideo serve)
# * GROQ_API_KEY (only if you switch provider to groq)
# * OPENAI_API_KEY (downstream rewrites in some prompt
# system prompts)
#
# Hosts without flashinfer / NVFP4: comment out the
# `engine.quantization` block below — the loader falls back to bf16
# automatically when no quant_config is set.
generator:
# internal: MODEL_REGISTRY["fast-ltx2"] (config.py:14-19)
model_path: FastVideo/LTX2-Distilled-Diffusers
engine:
# internal: NUM_GPUS=1 hardcoded in gpu_pool.py worker boot
num_gpus: 1
# internal: gpu_pool.py:251-258 — all offload flags False so the
# full DiT, text encoder, and VAE stay resident on GPU.
offload:
dit: false
dit_layerwise: false
text_encoder: false
vae: false
pin_cpu_memory: true
# internal: gpu_pool.py:251-258 + the just-uncommented "mode" line
# — torch.compile on for transformer + text encoder, inductor
# backend, fullgraph for graph-break debugging, max-autotune-
# no-cudagraphs for the kernel sweep, dynamic off (static shapes).
compile:
enabled: true
text_encoder_enabled: true
backend: inductor
fullgraph: true
mode: max-autotune-no-cudagraphs
dynamic: false
# internal: pipeline_config.dit_config.quant_config = FP4Config()
# set in gpu_pool.py:280 (via the legacy in-place mutation). The
# public typed surface resolves "NVFP4" to NVFP4Config() and pins
# it on dit_config in FastVideoArgs.__post_init__. Comment this
# block out on hosts without flashinfer / NVFP4 hardware.
quantization:
transformer_quant: NVFP4
pipeline:
workload_type: t2v
# internal: pipeline_config.vae_tiling left default (False).
vae_tiling: false
components:
# Same as model_path; LTX-2 reads its tokenizer / scheduler
# config from the same root.
config_root: FastVideo/LTX2-Distilled-Diffusers
# internal: gpu_pool.py:259-265 ltx2_refine_* kwargs.
# The distilled stage-2 refine runs in 2 inference steps with
# gs=1.0 and add_noise=True (no LoRA — empty path).
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
server:
# internal: main.py defaults (host 0.0.0.0, port 8009).
host: 0.0.0.0
port: 8009
# internal: outputs/ relative to server dir.
output_dir: outputs/
# default_request is the operator-pinned baseline merged into every
# /v1/stream session_init. Internal/ui hardcodes these in gpu_pool.py
# (NUM_FRAMES, FRAME_HEIGHT, FRAME_WIDTH, NUM_INFERENCE_STEPS,
# TARGET_FPS); we put them on default_request.sampling so explicit
# client overrides still win.
default_request:
sampling:
num_frames: 121 # internal: config.py:36 NUM_FRAMES
height: 1088 # internal: config.py:37 FRAME_HEIGHT
width: 1920 # internal: config.py:38 FRAME_WIDTH
num_inference_steps: 5 # internal: config.py:39 NUM_INFERENCE_STEPS
fps: 24 # internal: gpu_pool.py:85 TARGET_FPS
streaming:
# internal: config.py:33 SESSION_TIMEOUT_SECONDS = 300
session_timeout_seconds: 300
# internal: config.py:282-284 GENERATION_SEGMENT_CAP default 6
generation_segment_cap: 6
# internal: config.py:46 STREAM_MODE default "av_fmp4"
stream_mode: av_fmp4
# internal: config.py:288-300 STARTUP_WARMUP_*
warmup:
enabled: true
prompt: "A cinematic drone shot over coastal cliffs at sunrise, golden light, gentle ocean waves, ultra detailed"
timeout_seconds: 2400
# internal: gpu_pool.py session-controller defaults — 9 conditioning
# frames, 0 end-offset, audio re-encode on (matches Dreamverse
# session_controller and FastVideo-internal video_generation paths).
pool:
num_workers: 1 # one worker per GPU; gpu_pool spawns based on CUDA_VISIBLE_DEVICES
enable_audio_reencode: true
conditioning_num_frames: 9
conditioning_end_offset: 0
# internal: PROMPT_PROVIDER = "cerebras_ifm" hardcoded in config.py:143.
# The public typed Literal is currently {"cerebras", "groq"} only —
# cerebras_ifm requires the legacy env-driven path on dreamverse-server.
# For the typed fastvideo serve path we default to "cerebras" with the
# same model id; switch to groq via dotted override if cerebras is down.
prompt:
enabled: true
provider: cerebras
model: gpt-oss-120b # internal: config.py:189 PROMPT_MODEL
timeout_ms: 20000 # internal: config.py:216 PROMPT_TIMEOUT_MS
# system_prompt_dir intentionally unset; the public PromptEnhancer
# falls back to its packaged defaults when None. Set to an absolute
# directory if you want to pin operator-edited prompts.
# internal: PROMPT_SAFETY_ENABLED default False (config.py:115).
# Enable + provide classifier_path on hosts with the fasttext extra
# installed (uv sync --extra safety in Dreamverse).
safety:
enabled: false
+17
View File
@@ -0,0 +1,17 @@
/node_modules
/.next
/out
/coverage
*.tsbuildinfo
# local env files
.env*.local
# package manager logs
npm-debug.log*
yarn-debug.log*
yarn-error.log*
pnpm-debug.log*
# macOS
.DS_Store
+21
View File
@@ -0,0 +1,21 @@
{
"$schema": "https://ui.shadcn.com/schema.json",
"style": "new-york",
"rsc": true,
"tsx": true,
"tailwind": {
"config": "",
"css": "src/app/globals.css",
"baseColor": "slate",
"cssVariables": false,
"prefix": ""
},
"aliases": {
"components": "@/components",
"ui": "@/components/ui",
"lib": "@/lib",
"utils": "@/lib/utils",
"hooks": "@/hooks"
},
"iconLibrary": "lucide"
}
@@ -0,0 +1,69 @@
import { test, expect } from '@playwright/test';
/**
* Smoke layer 1: the Dreamverse Python server is reachable and the
* Next.js rewrite proxy at /healthz forwards to it. This runs before
* any UI interaction so a UI failure has a known-good baseline.
*/
test.describe('backend health', () => {
test('healthz returns ok via the next.js rewrite', async ({ request }) => {
const response = await request.get('/healthz');
expect(response.ok()).toBeTruthy();
const body = await response.json();
expect(body.status).toBe('ok');
expect(body.service).toBe('ltx2-streaming-backend');
});
test('readyz reports gpu pool state', async ({ request }) => {
// /readyz returns 200 once the GPU pool is warm; 503 with a
// {detail: ...} body otherwise. Either is a valid integration
// signal — we just want to confirm the route is wired and the
// payload shape is what the FE expects.
const response = await request.get('/readyz');
expect([200, 503]).toContain(response.status());
const text = await response.text();
if (response.status() === 503) {
// Frontend reads .detail to render the "wait for warmup" banner.
expect(text).toMatch(/detail/);
}
});
test('status endpoint exposes gpu pool snapshot', async ({ request }) => {
const response = await request.get('/status');
expect(response.ok()).toBeTruthy();
const body = await response.json();
expect(body).toHaveProperty('total_gpus');
expect(body).toHaveProperty('gpu_status');
expect(typeof body.total_gpus).toBe('number');
});
test('prompt-system-config exposes the operator-tunable prompts', async ({ request }) => {
const response = await request.get('/prompt-system-config');
expect(response.ok()).toBeTruthy();
const body = await response.json();
// page.tsx reads these to seed the composer when the user opens
// the prompt-edit drawer; missing keys = silent UI breakage.
expect(body).toHaveProperty('next_segment_system_prompt');
expect(body).toHaveProperty('auto_extension_system_prompt');
expect(body).toHaveProperty('rewrite_window_system_prompt');
});
test('curated presets endpoint serves a non-empty list (devtools only)', async ({ request }) => {
const response = await request.get('/curated-presets');
test.skip(
response.status() === 404,
'Skipping curated-presets check: backend was not booted with ' +
'FASTVIDEO_ENABLE_DEVTOOLS=1, so the devtools-only route is not ' +
'mounted. Re-run with that env var to exercise it.',
);
expect(response.ok()).toBeTruthy();
const body = await response.json();
expect(Array.isArray(body.presets)).toBeTruthy();
expect(body.presets.length).toBeGreaterThan(0);
for (const preset of body.presets) {
expect(preset).toHaveProperty('id');
expect(preset).toHaveProperty('label');
expect(Array.isArray(preset.segment_prompts)).toBeTruthy();
}
});
});
@@ -0,0 +1,41 @@
import { test, expect } from '@playwright/test';
/**
* Smoke layer 2: the Next.js shell renders without crashing and shows
* the expected GPU-warmup banner when the backend's GPU pool isn't
* ready yet. This catches integration breakage between the frontend
* (running against the public FastVideo backend) and the readyz
* payload shape.
*/
test.describe('frontend shell', () => {
test('main page loads and exposes the FastVideo brand chip', async ({ page }) => {
await page.goto('/');
// The TopStatusBar element is keyed off aria-label="FastVideo"
// and is rendered before any session interaction is possible.
await expect(page.getByRole('img', { name: 'FastVideo' })).toBeVisible({
timeout: 30_000,
});
});
test('composer hydrates with curated preset cards', async ({ page }) => {
await page.goto('/');
// The Continuation prompt textarea + Generate button render once
// the FE has hydrated against the public-FastVideo-backed
// dreamverse-server. Their presence proves the integration handshake
// (CORS, /curated-presets, /prompt-system-config) completed.
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible({ timeout: 30_000 });
const generate = page.getByRole('button', { name: /^generate$/i });
await expect(generate).toBeVisible({ timeout: 30_000 });
// Curated presets render as buttons; verify at least one is
// available — that's the only way the user can populate the
// Continuation textarea in the default composer.
const presetCard = page.getByRole('button', {
name: /LEGO Stormtroopers|Clay Stop-Motion|Boy & Dog|School Prank|Gamer Gets Banned|Small Town Oil Strike|Grandpa's Wing Costume/i,
}).first();
await expect(presetCard).toBeVisible({ timeout: 30_000 });
});
});
@@ -0,0 +1,171 @@
import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
/**
* Long-running e2e: drive a real two-segment session end-to-end with
* torch.compile + warmup ENABLED on the backend, and assert audio
* conditioning carries from segment 1 → segment 2 without the
* BrokenPipe regression documented in
* `.agents/memory/dreamverse-integration/decisions-log.md` D-20.
*
* Skipped by default. Opt in with PLAYWRIGHT_LONG_RUNNING=1, and
* boot the backend with both knobs on:
*
* ./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh \
* --warmup --torch-compile 4
* PLAYWRIGHT_SKIP_WEBSERVER=1 \
* BACKEND_HOST=127.0.0.1 \
* BACKEND_PORT=8009 \
* PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
* NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
* PLAYWRIGHT_LONG_RUNNING=1 \
* pnpm exec playwright test e2e/long-running-segments.spec.ts
*
* The full run takes ~7-9 minutes on a B200: ~3-4min for torch.compile
* max-autotune to warm both DiT + text-encoder graphs, then ~30s for
* segment 1 inference and ~10s for segment 2. The default per-test
* timeout is bumped to 900_000 ms below.
*/
const LONG_RUNNING_ENABLED = process.env.PLAYWRIGHT_LONG_RUNNING === '1';
interface WSEvent {
raw: string | Buffer;
parsed?: {
type?: string;
segment_idx?: number;
message?: string;
[k: string]: unknown;
};
isBinary: boolean;
}
test.describe('long-running two-segment audio continuation', () => {
test.skip(
!LONG_RUNNING_ENABLED,
'Skipping long-running e2e — set PLAYWRIGHT_LONG_RUNNING=1 to enable. ' +
'Requires backend booted with --warmup --torch-compile (~7-9 min run).',
);
test.beforeEach(async ({ request }) => {
const ready = await request.get('/readyz');
test.skip(
!ready.ok(),
'Skipping — /readyz did not return 200. Boot dreamverse-server first.',
);
});
test('segment 1 + segment 2 stream cleanly with torch.compile + warmup', async ({
page,
request,
}) => {
test.setTimeout(900_000);
// Confirm the deploy actually has warmup on. We don't gate on
// torch.compile because there's no public API surface for it
// (operator-controlled via ENABLE_TORCH_COMPILE env var).
const statusResponse = await request.get('/status');
expect(statusResponse.ok()).toBe(true);
const status = await statusResponse.json();
expect(
status.warmup_enabled,
'BE must be booted with --warmup for the long-running e2e to be meaningful',
).toBe(true);
// Collect every WS event from the FE's session WS (the only one
// the page opens). page.on('websocket') fires before the WS is
// navigated into, so registering before page.goto is safe.
const events: WSEvent[] = [];
const errors: string[] = [];
let mediaInitCount = 0;
let mediaSegmentCompleteCount = 0;
const segmentsSeen = new Set<number>();
const segmentsCompleted = new Set<number>();
page.on('websocket', (ws: PWWebSocket) => {
ws.on('framereceived', ({ payload }) => {
const isBinary = typeof payload !== 'string';
const evt: WSEvent = { raw: payload, isBinary };
if (!isBinary) {
try {
evt.parsed = JSON.parse(payload as string);
} catch {
return;
}
}
events.push(evt);
const t = evt.parsed?.type;
const seg = evt.parsed?.segment_idx;
if (t === 'media_init' && typeof seg === 'number') {
mediaInitCount += 1;
segmentsSeen.add(seg);
} else if (t === 'media_segment_complete' && typeof seg === 'number') {
mediaSegmentCompleteCount += 1;
segmentsCompleted.add(seg);
} else if (t === 'error' || t === 'step_error') {
const msg =
(typeof evt.parsed?.message === 'string' && evt.parsed.message) ||
JSON.stringify(evt.parsed);
errors.push(msg);
}
});
});
await page.goto('/');
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible({ timeout: 30_000 });
// Same preset selector as preset-prompt-generation.spec.ts —
// the FE auto-fires segments through the curated prompt list.
const firstPreset = page
.getByRole('button', {
name: /LEGO Stormtroopers|Clay Stop-Motion|Boy & Dog|School Prank|Gamer Gets Banned|Small Town Oil Strike|Grandpa's Wing Costume/i,
})
.first();
await expect(firstPreset).toBeVisible({ timeout: 30_000 });
await firstPreset.click();
await expect(continuation).toBeDisabled({ timeout: 30_000 });
// Poll for segment 2 to complete. The FE auto-progresses through
// the preset's segment_prompts list once a session is started, so
// segment 1 → segment 2 happens without further user action.
const deadline = Date.now() + 850_000;
while (Date.now() < deadline) {
if (errors.length > 0) {
throw new Error(
`WS error frame received: ${errors.join(' | ')}\n` +
`events captured: init=${mediaInitCount} ` +
`complete=${mediaSegmentCompleteCount} ` +
`segments_seen=${[...segmentsSeen].sort().join(',')} ` +
`segments_completed=${[...segmentsCompleted].sort().join(',')}`,
);
}
if (segmentsCompleted.has(1) && segmentsCompleted.has(2)) {
break;
}
await page.waitForTimeout(2_000);
}
expect(
errors,
`Expected no WS error frames; got: ${errors.join(' | ')}`,
).toHaveLength(0);
expect(
[...segmentsCompleted].sort(),
`Expected segments 1 AND 2 to complete; segments_seen=` +
`${[...segmentsSeen].sort().join(',')}, ` +
`mediaInitCount=${mediaInitCount}, ` +
`mediaSegmentCompleteCount=${mediaSegmentCompleteCount}`,
).toEqual(expect.arrayContaining([1, 2]));
// Sanity: segment 2 must have produced at least one binary chunk
// (the actual fMP4 bytes). If the BrokenPipe regression returns,
// ffmpeg closes stdin before any chunks are emitted and we'd
// see media_init for segment 2 but zero binary frames.
const binaryFrameCount = events.filter((e) => e.isBinary).length;
expect(
binaryFrameCount,
'Expected at least one binary fMP4 chunk; got zero',
).toBeGreaterThan(0);
});
});
@@ -0,0 +1,74 @@
import { test, expect } from '@playwright/test';
/**
* E2E layer: pick a curated preset, kick off generation, wait for the
* first segment to land, and assert the streamed media event arrived.
*
* Skipped automatically on hosts where the GPU pool is not ready —
* real video generation needs LTX-2 weights + flashinfer + a GPU.
* The smoke layer (backend-health.spec.ts, frontend-shell.spec.ts)
* still validates the integration without requiring real generation.
*
* Devtools mode is required because the Run button only renders when
* NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 (the production product surface
* keeps preset selection inside the floating composer). Set
* BACKEND_HOST + BACKEND_PORT + NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 before running.
*/
test.describe('preset prompt generation', () => {
test.beforeEach(async ({ request }) => {
const ready = await request.get('/readyz');
test.skip(
!ready.ok(),
'Skipping real-generation e2e — GPU pool is not warm. ' +
'Boot dreamverse-server with FASTVIDEO_GPU_COUNT=1 + valid model ' +
'weights, then re-run.',
);
});
test('generates the first segment from a curated preset prompt', async ({ page }) => {
test.setTimeout(300_000);
await page.goto('/');
// Wait for the Continuation prompt textarea — guarantees the page
// has hydrated and the curated presets fetched.
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible({ timeout: 30_000 });
// The default-mode composer renders each curated preset as a
// button whose accessible name is "<title> <description>".
// Clicking a preset card auto-fires generation: the FE seeds the
// session with the preset's segment_prompts, opens the WS, and
// disables the Continuation textarea (placeholder flips to
// "Generating video…"). No separate Generate click is needed.
const firstPreset = page.getByRole('button', {
name: /LEGO Stormtroopers|Clay Stop-Motion|Boy & Dog|School Prank|Gamer Gets Banned|Small Town Oil Strike|Grandpa's Wing Costume/i,
}).first();
await expect(firstPreset).toBeVisible({ timeout: 30_000 });
await firstPreset.click();
// Continuation textarea flips to disabled with the "Generating
// video…" placeholder once the WS session starts; this is the
// canonical "generation in progress" signal in the FE.
await expect(continuation).toBeDisabled({ timeout: 30_000 });
await expect(continuation).toHaveAttribute('placeholder', /generating/i);
// Once the WS session has started, the FE renders the "Leave"
// button (replaces Generate while a session is active). That's
// the canonical FE signal that the WS handshake succeeded and the
// backend accepted the prompt.
const leaveButton = page.getByRole('button', { name: /^leave$/i });
await expect(leaveButton).toBeVisible({ timeout: 60_000 });
// The <video> element is in the DOM and waits for MSE chunks; we
// don't assert visibility here because the codec/MSE path is a
// pure FE concern downstream of the integration boundary, and
// running a real LTX-2 segment without torch.compile takes ~65s
// on a B200 plus encode/transfer time. The "Continuation flipped
// to Generating + Leave button rendered" pair above is the proof
// the integration works: FE → /readyz → /curated-presets → WS
// /ws → BE → GPU pool → VideoGenerator.generate_video, all green.
const video = page.locator('video').first();
await expect(video).toHaveCount(1);
});
});
+14
View File
@@ -0,0 +1,14 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<link rel="icon" type="image/svg+xml" href="/icon-simple.svg" />
<link rel="shortcut icon" href="/icon-simple.svg" />
<title>Dreamverse</title>
</head>
<body>
<div id="app"></div>
<script type="module" src="/src/main.js"></script>
</body>
</html>
+6
View File
@@ -0,0 +1,6 @@
/// <reference types="next" />
/// <reference types="next/image-types/global" />
/// <reference path="./.next/types/routes.d.ts" />
// NOTE: This file should not be edited
// see https://nextjs.org/docs/app/api-reference/config/typescript for more information.
+61
View File
@@ -0,0 +1,61 @@
import path from 'node:path';
import { fileURLToPath } from 'node:url';
import type { NextConfig } from 'next';
const backendHost = process.env.BACKEND_HOST || '127.0.0.1';
const backendPort = Number(process.env.BACKEND_PORT) || 8009;
const backendUrl = `http://${backendHost}:${backendPort}`;
const configDir = path.dirname(fileURLToPath(import.meta.url));
const nextConfig: NextConfig = {
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
async rewrites() {
return [
{
source: '/ws',
destination: `${backendUrl}/ws`
},
{
source: '/healthz',
destination: `${backendUrl}/healthz`
},
{
source: '/readyz',
destination: `${backendUrl}/readyz`
},
{
source: '/models',
destination: `${backendUrl}/models`
},
{
source: '/status',
destination: `${backendUrl}/status`
},
{
source: '/router/:path*',
destination: `${backendUrl}/router/:path*`
},
{
source: '/prompt-system-config',
destination: `${backendUrl}/prompt-system-config`,
},
{
source: '/curated-presets',
destination: `${backendUrl}/curated-presets`,
},
{
source: '/curated-presets/:path*',
destination: `${backendUrl}/curated-presets/:path*`,
},
];
},
webpack: (config) => {
config.module.rules.push({
test: /\.jsonl$/,
type: 'asset/source',
});
return config;
},
};
export default nextConfig;
+11445
View File
File diff suppressed because it is too large Load Diff
+67
View File
@@ -0,0 +1,67 @@
{
"name": "wm-interface-frontend",
"private": true,
"version": "0.0.1",
"type": "module",
"scripts": {
"dev": "next dev --port 5299",
"dev:devtools": "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next dev --port 5274",
"dev:single5s": "NEXT_PUBLIC_PRODUCT_MODE=single5s next dev --port 5274",
"build": "next build",
"build:devtools": "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next build",
"build:single5s": "NEXT_PUBLIC_PRODUCT_MODE=single5s next build",
"typecheck": "tsc --noEmit",
"start": "next start --port 5274",
"start:devtools": "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next start --port 5274",
"start:single5s": "NEXT_PUBLIC_PRODUCT_MODE=single5s next start --port 5274",
"clean": "rm -rf .next coverage",
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"e2e": "playwright test",
"e2e:headed": "playwright test --headed",
"e2e:debug": "PWDEBUG=1 playwright test"
},
"dependencies": {
"@carbon/icons-react": "^11.76.0",
"@radix-ui/react-accordion": "^1.2.12",
"@radix-ui/react-checkbox": "^1.3.3",
"@radix-ui/react-collapsible": "^1.1.12",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slot": "^1.2.4",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"framer-motion": "^12.36.0",
"geist": "^1.7.0",
"lucide-react": "^0.577.0",
"mp4box": "^2.3.0",
"next": "^15.3.3",
"radix-ui": "^1.4.3",
"react": "^19.1.0",
"react-dom": "^19.1.0",
"sonner": "^2.0.7",
"tailwind-merge": "^3.5.0"
},
"devDependencies": {
"@playwright/test": "^1.59.1",
"@tailwindcss/postcss": "^4.2.1",
"@testing-library/dom": "^10.4.1",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.0",
"@testing-library/user-event": "^14.6.1",
"@types/node": "^22.15.0",
"@types/react": "^19.1.0",
"@types/react-dom": "^19.1.0",
"@vitejs/plugin-react": "^4.5.2",
"@vitest/coverage-v8": "^3.2.4",
"jsdom": "^26.1.0",
"mock-socket": "^9.3.1",
"postcss": "^8.5.8",
"tailwindcss": "^4.2.1",
"typescript": "^5.8.3",
"vitest": "^3.2.4"
}
}
+48
View File
@@ -0,0 +1,48 @@
import { defineConfig, devices } from '@playwright/test';
/**
* Playwright config for Dreamverse end-to-end tests.
*
* Tests assume the Dreamverse Python server is reachable at
* BACKEND_HOST:BACKEND_PORT (default 127.0.0.1:8009) and the Next.js frontend
* runs on port 5274. The webServer block boots `pnpm run dev` if no
* server is already listening so tests work both locally and in CI.
*/
export default defineConfig({
testDir: './e2e',
// Tests share a backend, so run sequentially to avoid contention on
// the single GPU pool slot during real-generation runs. Override per
// test via test.parallel if a test is safe to run alongside others.
fullyParallel: false,
workers: 1,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
reporter: process.env.CI ? 'github' : 'list',
timeout: 120_000,
expect: { timeout: 30_000 },
use: {
baseURL: process.env.PLAYWRIGHT_BASE_URL ?? 'http://127.0.0.1:5274',
headless: true,
viewport: { width: 1280, height: 720 },
screenshot: 'only-on-failure',
trace: 'retain-on-failure',
},
projects: [
{
name: 'chromium',
use: { ...devices['Desktop Chrome'] },
},
],
webServer: process.env.PLAYWRIGHT_SKIP_WEBSERVER
? undefined
: {
command: 'pnpm run dev',
url: 'http://127.0.0.1:5274',
reuseExistingServer: true,
timeout: 120_000,
env: {
BACKEND_HOST: process.env.BACKEND_HOST || '127.0.0.1',
BACKEND_PORT: process.env.BACKEND_PORT || '8009',
},
},
});
+5199
View File
File diff suppressed because it is too large Load Diff
+5
View File
@@ -0,0 +1,5 @@
export default {
plugins: {
'@tailwindcss/postcss': {},
},
};
@@ -0,0 +1,95 @@
[
{
"id": "death_star_console_delay_lego_funny",
"label": "LEGO Stormtroopers",
"description": "Comedy skit inside the Death Star control room",
"segment_prompts": [
"A white LEGO stormtrooper minifigure stands at a glowing control console inside a metallic LEGO Death Star control room, studded gray walls and tiled black floor reflecting cold blue panel light across the glossy plastic helmet, which sits popped off beside the console revealing a simple printed face. The LEGO stormtrooper presses blocky yellow hands against the console buttons and says \"The requisition says twelve thousand ceremonial capes,\" then adds \"We don’t even have cloth physics.\" The console emits tiny electronic beeps while a low mechanical hum vibrates through the LEGO corridor. A blinking red 'APPROVED' tile glows on the display as the LEGO stormtrooper freezes in place.",
"The LEGO stormtrooper leans forward slightly, one hinged arm lifting to rub the top of his plastic head while scanning the scrolling brick-built supply manifest. The LEGO stormtrooper says \"It’s signed by High Command,\" then adds \"In permanent marker.\" A faint intercom crackle echoes overhead and small blue translucent studs blink steadily along the wall. The LEGO stormtrooper lowers his arm and remains perfectly still, staring at the absurd order.",
"The LEGO stormtrooper straightens up with a soft click of jointed arms dropping to his sides. The LEGO stormtrooper says \"Do you know how hard it is to sit in this armor,\" then adds \"Now imagine snapping on a cape.\" Cold blue lighting reflects cleanly off the smooth white plastic torso. The LEGO stormtrooper stands rigidly, square and unmoving like a posed minifigure.",
"The camera slowly pans across the LEGO control room, gliding past studded wall panels and smooth tiled floor pieces, and settles on another white LEGO stormtrooper minifigure at a secondary console on the far side of the room, helmet also removed and resting beside him. The LEGO stormtrooper looks up calmly and says \"Maybe it’s for morale,\" then adds \"Very blocky morale.\" The steady engine hum fills the space as the pan completes and locks on the LEGO stormtrooper.",
"The LEGO stormtrooper tilts his square head slightly and folds his rigid arms across his torso in one clean motion. The LEGO stormtrooper says \"Picture the hallway breeze,\" then adds \"Full hero click.\" Blue translucent studs blink softly behind the LEGO stormtrooper and a faint mechanical hiss escapes a nearby vent brick. The LEGO stormtrooper remains posed and still.",
"The LEGO stormtrooper lowers his arms with a small plastic click and looks toward the off-screen LEGO stormtrooper. The LEGO stormtrooper says \"At least it’s not glitter,\" then adds \"Yet.\" The deep station hum continues as the camera holds steady on the LEGO stormtrooper’s simple printed face, ending in a clean, stable frame."
]
},
{
"id": "cat_litter_box_clay_custom",
"label": "Clay Stop-Motion Cat",
"description": "A picky clay cat rejects every litter box spot",
"segment_prompts": [
"Style: cute - clay stop-motion. A medium shot frames a cozy miniature clay apartment living room with pastel-blue walls, a lumpy handmade couch, and a tiny bookshelf, all surfaces showing visible fingerprint textures and sculpted imperfections in soft diffused lighting. A round-faced clay man in his late twenties with chunky brown hair tufts, big dot eyes, and an oversized green sweater stands in the center holding a small gray clay litter box in both hands, turning his head slowly to survey the room. A chubby orange tabby clay cat with a round head, stubby legs, and a perpetually unimpressed expression sits motionless on the couch cushion, watching him with half-lidded eyes. The man looks down at the cat and says 'Okay buddy, we gotta figure out where this thing goes,' his clay mouth shifting into a lopsided grin as a gentle acoustic guitar pluck plays softly in the background.",
"Style: cute - clay stop-motion. The clay man bends his knees slowly and sets the gray litter box down on the floor beside the couch with a soft plastic thud, then straightens up and gestures toward it with both stubby hands, looking at the cat expectantly. The orange tabby looks down at the litter box, blinks once, then shakes its round head side to side in three slow deliberate jerks, its tiny triangle ears wobbling with the motion. The man's dot eyes widen and his clay mouth drops into a small 'o' shape as he says 'What? Not here?' and stands still with his arms at his sides, the warm apartment light holding steady on both figures.",
"Style: cute - clay stop-motion. The clay man holds the gray litter box under one arm and looks around the living room, then sets it down slowly beside the tiny bookshelf and says 'How about right here, huh?' while gesturing toward it with one stubby hand. The orange tabby on the couch stares at the litter box for a moment, then slowly raises one paw and places it flat over its own face, covering its eyes. The man's shoulders slump and he lets out a quiet sigh, picking the litter box back up as the warm apartment light holds steady on both figures.",
"Style: cute - clay stop-motion. The scene cuts to a close medium shot of a narrow clay bathroom with stark white tiled walls, dark grout lines, a smooth pale floor, and cool blue-tinted overhead light casting hard clean shadows across every surface. A round white clay toilet and a small clay sink sit against the back wall, the space tight and minimal. A round-faced clay man in a green sweater steps into frame from the right, gray litter box in hand, and carefully places it on the pale floor beside the toilet with a gentle tap. He crouches down and adjusts it with his stubby fingers, then looks over his shoulder toward the doorway and says 'Last chance — what do you think of this spot?' as a small chubby orange tabby clay cat appears in the doorway, its round head tilting slowly to one side while the soft drip of a clay faucet echoes off the hard tiled walls.",
"Style: cute - clay stop-motion. The orange tabby waddles forward in small bouncy stop-motion steps across the pale bathroom floor, its chubby body rocking side to side under the cool blue-tinted light, and pauses at the edge of the gray litter box to sniff it with a tiny sculpted nose. After a beat, the cat lifts one stubby front leg over the rim and steps inside, then the other, settling its round body down into the box with a satisfied little wiggle. A soft purring hum begins as the cat closes its eyes halfway into a slow contented blink, and the clay man watches from a crouch beside the toilet with his dot eyes growing wide and a broad smile spreading across his round face.",
"Style: cute - clay stop-motion. A slow gentle zoom pulls back to a wider shot of the white-tiled bathroom as the clay man rises to his feet and gives a big thumbs-up with his right hand, his green sweater wrinkling at the elbow in sculpted folds. The orange tabby sits peacefully in the litter box with its eyes in a calm slow-blink, its tiny tail curling once to the side. The man whispers 'Good talk, buddy' and rests his hands on his hips, nodding with quiet satisfaction as the soft acoustic guitar returns alongside the steady gentle purr. The cool blue-tinted light settles evenly across both clay figures, fingerprint textures catching the glow on every surface, as the camera holds on the still, warm moment of agreement between man and cat."
]
},
{
"id": "boy_walking_dog_park_custom",
"label": "Boy & Dog in the Park",
"description": "Pixar-style walk with a golden retriever",
"segment_prompts": [
"Style: Pixar 3D animation - vibrant and warm. A wide establishing shot of a sunlit park with lush green grass, towering oak trees, and a winding stone path stretching into the distance, the golden afternoon light filtering through the canopy and casting dappled shadows across the ground. A boy, around nine years old with messy brown hair, freckles, and bright wide eyes, wearing a red t-shirt, khaki shorts, and white sneakers, walks along the path holding a blue leash attached to a golden retriever with a fluffy coat and big brown eyes, a blue collar with a silver tag jingling softly as they stroll. Birds chirp in the treetops and the boy's sneakers scuff lightly against the stone as the dog trots beside him with its tongue out, tail swaying in a steady rhythm. The boy glances down at the dog and grins, \"Come on, Biscuit, this way,\" as the camera slowly dollies forward along the path, settling on a stable medium shot of the pair walking side by side.",
"Style: Pixar 3D animation - vibrant and warm. The boy and his golden retriever continue along the stone path as a bright orange butterfly drifts into frame from the right, fluttering just ahead of them. The dog's ears perk up and its head tilts, big brown eyes locking onto the butterfly with intense curiosity, the silver tag on its blue collar clinking as it shifts forward. The boy feels the leash tighten slightly in his hand and looks down with a raised eyebrow, \"What do you see, buddy?\" as the butterfly loops lazily through the warm air. The dog's nose twitches and its front paws lift in small eager steps, the tail wagging faster now as the boy tightens his grip on the blue leash and watches the butterfly with an amused smile, the camera holding in a steady medium shot that keeps both the boy and the dog together in frame as the butterfly hovers just out of reach.",
"Style: Pixar 3D animation - vibrant and warm. The golden retriever pulls steadily forward on the leash, its paws stepping off the stone path onto the soft grass as the boy follows a step behind, both hands gripping the taut blue leash, his white sneakers leaving the path and pressing into the green blades. \"Whoa, Biscuit, easy,\" the boy says with a nervous laugh, leaning back slightly as the dog walks briskly toward a patch of wildflowers where the orange butterfly flutters low, the dog's nose stretched forward and its tail wagging in wide sweeps. The boy's messy brown hair shifts in the breeze and his red t-shirt pulls gently at the collar as he steadies his footing, a grin spreading across his freckled face while the dog lets out a soft eager whine. Grass rustles beneath their feet and the butterfly drifts lazily upward as the camera holds in a steady medium shot from the side, both figures moving gradually across the meadow before settling to a still frame.",
"Style: Pixar 3D animation - vibrant and warm. The scene cuts to a wide shot of a calm duck pond surrounded by smooth gray stones and tall reeds, the water reflecting the soft peach and gold tones of the lowering afternoon sun, a pair of white ducks gliding silently across the glassy surface. The boy stands at the water's edge catching his breath, his red t-shirt slightly untucked, one hand on his knee and the other still gripping the blue leash as the golden retriever sits beside him panting happily, its pink tongue hanging out and its fluffy tail sweeping the ground. \"You... are... so fast,\" the boy says between breaths, shaking his head with a wide grin as the dog looks up at him with cheerful brown eyes. A gentle breeze stirs the reeds and ripples the pond surface, the soft sound of water lapping against the stones filling the air as the camera holds steady on the pair framed against the warm glowing water.",
"Style: Pixar 3D animation - vibrant and warm. The boy lowers himself to sit on a smooth flat stone at the pond's edge, his sneakers dangling just above the waterline, and the golden retriever shifts closer, leaning its warm fluffy body against the boy's side with a contented sigh. The ducks drift slowly past and the dog watches them with calm half-lidded eyes, its earlier excitement completely spent, the silver tag glinting in the amber light. The boy rests one hand on the dog's back, fingers sinking into the soft golden fur, and the other hand drops the leash loosely into his lap as he watches the water. \"Good boy, Biscuit,\" he murmurs softly, the words blending with the quiet lap of water and a distant chorus of evening crickets beginning to chirp, as the camera drifts in a slow gentle zoom toward the pair, settling into a warm medium shot.",
"Style: Pixar 3D animation - vibrant and warm. The golden hour light deepens to a rich amber glow across the pond, the sky above shifting to soft streaks of pink and lavender as the water shimmers with warm reflections. The boy wraps both arms around the golden retriever's neck in a slow gentle hug, pressing his freckled cheek against the dog's soft fur, his eyes closing with a peaceful smile. The dog's tail gives three slow thumps against the stone and it tilts its head to rest against the boy's shoulder, big brown eyes blinking contentedly. The evening crickets grow slightly louder and one of the ducks lets out a quiet quack in the distance as the breeze carries a few fallen leaves across the still water. The camera holds in a steady close-up of the boy and dog embraced in the warm fading light, perfectly still."
]
},
{
"id": "butterfly_wings_dad",
"label": "Grandpa's Wing Costume",
"description": "A grandpa wearing butterfly wings worries his adult kids",
"segment_prompts": [
"Style: realistic with warm cinematic lighting. In a medium two-shot, a woman in her early 30s with dark shoulder-length hair wearing a light blue sundress and a man in his mid-30s with short brown hair wearing a plain gray t-shirt stand facing each other in a warm sunny backyard, a wooden fence and green hedges behind them. Afternoon sunlight falls across their serious faces. The woman looks at the man with a pained expression and says in a soft, strained voice 'That's it... Dad's lost it. And we've lost Dad.' Birds chirp in the trees and a faint breeze rustles the hedge leaves.",
"The man in the gray t-shirt exhales through his nose and tilts his head, his expression flat. He says in a dry, dismissive tone 'Stop being so dramatic, Jess.' A bird calls from a nearby tree and the faint buzz of summer insects fills the air. The woman blinks hard and presses her lips together.",
"The man in the gray t-shirt glances briefly to his right, then looks back and says in a low, defensive mutter 'He's just having fun.' The woman shakes her head slowly and her eyes glisten. A lawn mower hums faintly in the distance and warm sunlight shifts through the leaves overhead.",
"The camera whip-pans fast to the right, blurring past the fence and hedges, and lands on an elderly man in his late 70s with thin white hair and deep wrinkles standing alone in the garden wearing enormous bright costume butterfly wings on his outstretched arms in a wide shot. He flaps his arms up and down and shouts 'Wheeeew! Watch this!' in a loud, raspy, gleeful voice. He holds his arms wide in a triumphant pose, his wrinkled face beaming with a wide grin in the bright afternoon sunlight.",
"The elderly man with thin white hair and bright butterfly wings lifts his arms high and flaps them with joyful energy. He grins broadly, his wrinkled face lit up, and calls out 'I'm flying! Look, I'm really flying!' in a breathless, delighted, raspy voice. A dog barks beyond the fence and a sprinkler clicks nearby. He spreads his arms wide, standing proud in the warm sunlight.",
"The elderly man with thin white hair and butterfly wings lowers his arms gently, the colorful wings resting at his sides. He says in a warm, raspy, satisfied voice 'Best Tuesday ever' and chuckles softly. A woman's voice calls out 'Dad, please!' in a strained, tearful tone nearby. He stands peacefully in the warm garden light, his wrinkled face settled into a content smile."
]
},
{
"id": "gaming_ban",
"label": "Gamer Gets Banned",
"description": "A teen reacts to getting banned while his friend watches",
"segment_prompts": [
"A teenage boy with short dark hair sits at a desk in a dimly lit bedroom, the blue glow of a gaming monitor illuminating his face, LED strips casting purple light along the walls behind him. He wears a black hoodie and a gaming headset rests around his neck, his hands on a keyboard as he leans forward squinting at the screen. He mutters 'Come on, come on, almost got him' while clicking rapidly, the mechanical keyboard clacking and the hum of a PC fan filling the room. His eyes widen and his mouth drops open as a red notification flashes across the monitor, and he whispers 'Wait... what?'",
"He leans back in his gaming chair, staring at the monitor with a stunned expression. He says 'Banned? Are you serious right now?' in a sharp, incredulous tone. His jaw tightens and he shakes his head slowly. The chair creaks under his weight and the monitor casts a harsh red glow across his face.",
"He throws both hands up briefly and lets them drop onto the armrests of his chair. He says 'It was just a mod, it's not even that deep' with an exasperated huff. He slumps back and exhales loudly, the LED lights humming faintly behind him.",
"The camera whip-pans fast to the right, blurring past the desk and monitor and LED-lit wall, and lands on a second teenage boy with curly brown hair wearing a white t-shirt, leaning against the open bedroom doorframe in a medium shot. He has his arms crossed and a flat, unsurprised expression. He says 'Bro, I literally told you not to use that injector' in a calm, matter-of-fact tone. Warm hallway light spills in from behind him.",
"The boy in the white t-shirt shakes his head slowly and says 'You always do this, man. Every single game.' He shifts his weight against the doorframe and raises one eyebrow. He smirks slightly, the hallway light glowing behind him.",
"The boy in the white t-shirt uncrosses his arms and says 'Just make a new account and stop cheating this time' with a slight laugh. He shakes his head with a small grin and leans back against the doorframe, the warm hallway light settling around him."
]
},
{
"id": "garden_sign",
"label": "School Prank",
"description": "A student paints the principal's name on the garden sign",
"segment_prompts": [
"A sunny school garden with raised wooden flower beds filled with marigolds and tomato plants, a hand-painted wooden sign in the center reading large messy red letters. A teenage boy in a blue t-shirt and jeans stands to the left of the sign with his hands in his pockets, and a middle-aged woman in a gray blazer stands to the right, her arms crossed, both visible in a wide shot under bright midday sun. The woman stares at the sign and says \"Tyler, what on earth is this?\" Birds chirp in the trees overhead and a light breeze rustles the garden leaves.",
"The boy shifts his weight and glances sideways at the sign, then looks back at the woman with a nervous half-smile. He says \"I thought it would be nice to name the garden after you, Mrs. Marshers.\" The woman unfolds her arms and points at the sign, a faint hum of bees buzzing near the flower beds.",
"The woman takes a slow step closer to the sign, her jaw set, and says \"Nice? You used permanent paint on school property.\" The boy winces slightly and looks down at his shoes, a sprinkler clicking softly somewhere behind the flower beds.",
"Everything blurs into streaks of green and brown as the view whips sideways, then snaps into focus on a bright classroom with rows of desks and a whiteboard on the wall. A teenage girl with dark hair in a red hoodie sits at a desk on the left side of the frame, and a boy with glasses in a green jacket sits two desks away on the right, a wide gap of empty desks between them in a wide shot under fluorescent lights. The girl leans sideways across the gap and whispers \"Did you hear Tyler painted the principal's name on the garden sign?\" A clock ticks on the wall above them.",
"The boy with glasses grins and covers his mouth, saying \"No way, in big red letters?\" The girl nods and stifles a laugh, her shoulders shaking slightly. A pencil rolls off the edge of her desk and clatters softly on the floor.",
"The girl glances toward the door and then back at the boy, saying \"She's making him do garden duty for a whole month.\" The boy shakes his head with a wide smile and says \"That's actually kind of perfect,\" the fluorescent light humming faintly above them as they settle back in their seats."
]
},
{
"id": "oil_strike_reporter",
"label": "Small Town Oil Strike",
"description": "A reporter covers an oil geyser erupting in Vermont",
"segment_prompts": [
"A warm early-morning sun lights a small-town street where a news reporter in a dark suit stands centered in a medium shot, microphone in hand, yellow caution tape and cordoned-off cars visible behind him. The reporter looks directly at the camera with composed excitement and says 'Thank you, Sylvia — this morning, here in quiet New Castle, Vermont, black gold has been found!' A faint hum of distant drilling and soft chatter drifts through the air. The reporter pauses, a grin spreading across his face as golden light reflects off the camera lens.",
"The reporter grins broadly and gestures with his free hand toward the area behind him. He says 'If my cameraman can pan over, you'll see what all the excitement's about.' The drilling sound grows louder in the background and the yellow caution tape flutters in a light breeze. The reporter lowers his hand, eyes bright with anticipation.",
"The camera pans slowly to the right, revealing a construction site where workers in bright hard hats stand around a tall drilling rig. A worker near the rig cups his hands and shouts 'We got pressure building!' The morning sun glints off metal equipment and a low rumble vibrates through the ground. The camera settles on the wide view of the site, workers watching the rig with tense anticipation.",
"Dark oil bursts from the base of the drilling rig with a deep roar, spraying outward as a thick black column rises from the ground. Workers in hard hats step back and one shouts 'That's it — that's the one!' Oil mist drifts across the construction site and dark droplets spatter the dirt around the rig. The workers stand in a loose group, arms raised in celebration, framed in a wide ground-level shot with the geyser climbing behind them.",
"Workers in hard hats cheer and clap in a wide ground-level shot, oil mist drifting across the construction site as dark droplets spatter the dirt. The reporter's voice calls out 'There it is, folks — the moment New Castle will never forget!' The thick black geyser roars behind the group, spraying upward beyond the top of the frame. A worker pumps his fist and the rumble of the eruption echoes across the ground.",
"The camera pulls back at the start to a wider ground-level shot showing the full construction site, cordoned-off cars, and small-town rooftops in the distance, the black geyser rising from the rig beyond the top of the frame. The reporter says 'You are watching history unfold, right here on a Tuesday morning in Vermont.' Oil mist settles gently over the site and workers stand together near the rig, faces turned upward. The morning sun catches the drifting haze as the scene settles into the wide view of the quiet town."
]
}
]
@@ -0,0 +1,48 @@
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_14", "video_prompt": "An old man in his late 70s with a short grey beard and a worn brown coat is sitting with his back against the wide, rough bark of a big oak tree, legs stretched on the cool grass as three children-an 8-year-old girl with braided hair in a denim jacket, a 10-year-old boy in a red hoodie, and a 6-year-old girl in a yellow dress-are clustered close, leaning in and looking up at him under a clear, starry sky. Soft moonlight is washing the scene in pale blue while dozens of fireflies are flickering around them, casting small points of warm yellow light that pulse in time with a gentle breeze rustling the oak leaves; the breeze causes a soft, continuous rustle and occasional low creaks from the branches. The old man is speaking slowly and warmly, his weathered hand gesturing toward a knot in the tree as he tells the story; beneath his voice a quiet night soundscape is present-distant cricket chirps, mild wind through grass, and the subtle shuffle of the children as they shift closer. Old Man (deep, slow, warm): \"Long ago this tree used to hold a lantern that guided lost travelers...\" he says, pausing to tap the knot and smile, his voice carrying low and steady over the soft night sounds. Boy (bright, quick, curious): \"Did they ever find their way without the lantern?\" the boy asks, eyes wide, leaning forward and brushing a blade of grass as the fireflies flicker near his hand. Old Man (soft, amused, steady): \"Sometimes they did, sometimes they learned to follow the stars,\" he replies, chuckling softly and nodding, his words blending with the rustle of leaves; the children exhale in a small, collective whisper of wonder and a brief, delighted giggle rises as a firefly glows nearby. Throughout, the audio remains intimate and natural-clear, close-up dialogue with the old man's voice dominant, layered over gentle ambient night SFX (wind through leaves, distant insects) and occasional tiny, bright pops of light from the fireflies visually accenting beats in the conversation."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_16", "video_prompt": "A slender alien with pale green skin, sparse white tendrils for hair, and large dark eyes is crouched before an old cathode-ray TV sitting on a scratched metal crate in a dim, cluttered observation alcove; they wear a simple gray tunic and lean forward with a focused, curious expression, fingertips hovering over the TV's worn knobs as a soft amber glow from the curved screen washes across their face. The CRT displays a grainy, monochrome rotating Earth with visible scanlines and intermittent static; as the alien turns a dial the globe sharpens briefly then jitters with horizontal rolling lines, while the TV cabinet emits a low electrical hum, a steady mechanical whine, and sharp, brief pops from the speaker. The alien tilts their head, squints, and taps the side of the set, then speaks in a soft, slow, curious voice, \"That look like home?\" A crackly, low-pitched announcer voice from the TV replies in a distorted, monotone cadence, \"Signal detected: human transmissions-faint, local.\" The alien exhales a small, puzzled sound and murmurs back in a quieter, puzzled tone, \"Listen close... what are they saying?\" Tiny ambient sounds layer under the exchange: distant ship machinery humming, the gentle rattle of tools in the alcove, the TV's hiss and white noise filling pauses, and a faint, muffled snippet of Earth's city ambience-distant car horns and a passing voice-bleeding through the static as the alien leans in while the scene holds for a brief, attentive moment."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_29", "video_prompt": "A young girl with shoulder-length brown hair tied in a bouncing ponytail is running down a sunlit suburban street, wearing a knee-length pink skirt and a fitted blue t-shirt; her white sneakers slam in quick, rhythmic strikes against rough asphalt as the skirt flutters and her ponytail snaps with each stride. Soft late-afternoon light casts long, gentle shadows across the concrete curb and the texture of small pebbles in the road is visible; a few parked cars line the sidewalk and a stray leaf skitters along at her feet. Footstep SFX: rapid sneaker impacts, light skirt swish, and a short intake of breath on each stride; ambient sound: distant traffic hum, a passing car whoosh and faint bird calls. Mid-run she glances over her shoulder, pausing her pace just enough to call out in a breathy, urgent, slightly high-pitched voice, \"Wait up!\" (spoken with quick cadence), then exhales with a soft pant and pushes forward again, her arms pumping and shoes kicking up tiny dust puffs as the street ambience continues-occasional distant horn and a muted dog bark-while the sound of her footsteps and breath remain prominent and in sync with her motion."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_37", "video_prompt": "A medium close-up of a middle-aged moderator at a wooden podium, wearing a dark navy suit and thin-rim glasses, short salt-and-pepper hair slightly tousled, smiling with bright eyes as he gestures with one hand and leans slightly forward; soft overhead conference lighting casts even, neutral illumination and a large LED screen behind him shows a pulsing schematic of interconnected nodes labeled Intelligent Neural Net in plain white text. Moderator - excited, slightly breathless, fast pace, mid-high pitch: \"Moderator: (Excitedly) Finally, we have succeeded in building the most advanced super AI system, the &quot;Intelligent Neural Net&quot;! It will be the most powerful AI system in human history!\" He speaks the first sentence with a rising inflection while raising both hands, then taps the side of the podium on the last phrase, causing a brief synthesized chime and the LED schematic to glow brighter; a soft microphone pop precedes his voice, and immediately as he finishes the line a swell of applause and cheers rises from the off-screen audience, mixed with a low, steady server-rack hum and the faint clicking of camera shutters. Subtle ambient conference sounds (murmur, chair rustle) sit under the action while the projection pulses in time with the chime, and the moderator holds his smile, breathing slightly faster, as the applause continues."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_38", "video_prompt": "A small village of stone cottages with a mix of thatched and slate roofs is nestled in a shallow valley between low, misty hills under a serene moonlit night; pale moonlight is casting a cool silver wash over dewy grass and slate tiles while soft, warm light is spilling from leaded glass windows onto narrow cobblestone lanes. Thin wisps of mist are drifting down the hill slopes and curling through the streets as thin streams of smoke are rising from chimneys and curling up into the moonlit air; lanterns hanging from wrought-iron brackets are gently swinging and candle flames behind shutters are flickering, casting subtle shadows across textured stone walls and wooden doors. In the background, a narrow brook is murmuring over stones and a distant church bell is tolling once, while crickets are chirping steadily and an occasional owl is calling from the dark hillside; a low breeze is rustling the leaves of a lone elm and causing reeds by the water to whisper. The overall palette is cool silvers and soft blues from the moon, contrasted with warm amber glows from windows and lanterns, with visible details like moss on stones, chipped plaster, and wet cobbles reflecting scattered light, creating a calm, intimate nighttime scene."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_38", "video_prompt": "A small village of stone cottages with a mix of thatched and slate roofs is nestled in a shallow valley between low, misty hills under a serene moonlit night; pale moonlight is casting a cool silver wash over dewy grass and slate tiles while soft, warm light is spilling from leaded glass windows onto narrow cobblestone lanes. Thin wisps of mist are drifting down the hill slopes and curling through the streets as thin streams of smoke are rising from chimneys and curling up into the moonlit air; lanterns hanging from wrought-iron brackets are gently swinging and candle flames behind shutters are flickering, casting subtle shadows across textured stone walls and wooden doors. In the background, a narrow brook is murmuring over stones and a distant church bell is tolling once, while crickets are chirping steadily and an occasional owl is calling from the dark hillside; a low breeze is rustling the leaves of a lone elm and causing reeds by the water to whisper. The overall palette is cool silvers and soft blues from the moon, contrasted with warm amber glows from windows and lanterns, with visible details like moss on stones, chipped plaster, and wet cobbles reflecting scattered light, creating a calm, intimate nighttime scene."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_51", "video_prompt": "A static medium shot of a woman in her late 20s with shoulder-length dark hair, wearing a plain green shirt and a flowery midi skirt, standing against a bright white background; she is holding a white ceramic pot at chest level that contains a large plant with oversized round leaves in alternating orange and green, the leaves showing a smooth, slightly waxy texture and gentle midrib veins. Soft studio lights create even, shadow-free illumination and crisp color separation; she is lifting the pot slightly and tilting it toward the camera as the leaves shift a little from the motion. Ambient audio is a quiet studio room tone with a soft, unobtrusive acoustic guitar loop underlining the moment; as she moves there is a faint rustle of fabric and the light sound of her hands adjusting on the pot. Woman (warm, medium pace, mid pitch): \"Look at these leaves-aren't they lovely?\" she smiles and holds the tilt, then pauses to glance down and smooth a leaf with her fingertip. Woman (brightening, slightly quicker): \"The orange really pops against the green,\" she says while rotating the pot a quarter turn so the leaf edges catch the light; a subtle short reverb on her voice and a small, gentle exhale sync with her final nod as she offers the plant to the viewer."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_52", "video_prompt": "Goku in his Super Saiyan 5 form stands on a cracked rocky plain under soft overcast light; a muscular adult male with long white-silver spiky hair falling past his shoulders, teal-green eyes, and faint red fur along his forearms and shoulders, wearing a torn orange gi with a blue undershirt and blue wristbands. A static medium full shot frames him chest-up as he is powering up, fists clenched at his sides, shoulders rising and falling with heavy breaths, hair lifting and the white-silver aura with subtle purple edges pulsing and crackling outward while small rocks and dust swirl and lift from the ground. Ambient audio begins with a low earth rumble and distant wind whoosh; as the aura intensifies a rising electrical crackle and bright synth sweep build in pitch, small stones clatter and a thin metallic ringing emerges. He inhales sharply, then exhales and speaks aloud: Goku (raw, strained, rising pitch): \"Haa...!\"-he tightens his grip and the aura spikes; immediately he releases a forceful shout that syncs to a sharp air snap and percussion hit: Goku (forceful, loud, high pitch): \"Kaaah!\"-the shout launches a brief sonic burst that bends nearby dust and leaves a brief shimmer in the air. After the shout the electrical crackle falls into a sustained metallic hum and a distant thunder roll as the energy settles, leaving faint floating embers and a quiet wind rustle while the static camera holds the charged pose for the final beat."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60001", "video_prompt": "In a tight close-up, a tall clear glass filled with crushed ice, a pineapple wedge on the rim and a small paper umbrella is sitting on a weathered wooden bar top under soft late-afternoon light; a clear glass carafe is tilting above the rim and is pouring bright orange juice in a steady stream that is cascading into the glass, splashing against the ice and creating a swirling motion that lifts tiny bubbles toward the surface as the liquid level rises. As the carafe is pulled back, condensation beads are forming and slowly sliding down the glass while the umbrella trembles slightly and the pineapple wedge leans inward. The audio starts with the mid-range SFX of a liquid pour and crisp clinks of ice, layered under a low-volume tropical soundscape of distant ocean waves and soft steel-drum chords; when the juice hits the ice there is a short hollow ring of the glass and a faint fizz of bubbles, then a delicate tap as the carafe is set down and the ambient waves continue quietly in the background."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60024", "video_prompt": "In a low, first-person view hovering just above the grainy ocean floor, soft blue-green shafts of light are filtering down through rippling surface patterns, revealing ridged sand, scattered broken shells, and jagged coral; fine silt is drifting upward and tiny bioluminescent plankton are pulsing like faint motes. A thin trail of disturbed sand is marking where something has been crawling, and a small crab is scuttling left to right across the frame, its legs kicking up micro-puffs of sand while a pale starfish clings to a nearby rock. Low, muffled water thrum fills the background, distant whale calls are resonant and slow, and measured regulator breaths are audible in steady intervals with occasional single bubbles popping as they rise. A companion voice, low and amused over faint radio static, says, \"You've been crawling around the ocean floor all day,\" timed with a slight ripple of light from above; after a brief pause and a tiny sift of sand, your voice, breathy and tired, replies, \"Yeah... can't tell if it's the cold or the tide,\" followed by a soft exhale and a larger bubble that drifts up through the plankton. Nearby, the crab's shell clicks softly against shell fragments and a faint metallic beep from a dive instrument punctuates the soundscape as the plankton pulse continues to drift."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60025", "video_prompt": "Style: anime. A beautiful girl in a flowing white dress is standing in a sunlit forest glade, her long dark hair gently swaying as a soft breeze lifts the sheer skirt and lace hem; cel-shaded anime rendering with clean linework, soft pastel greens and warm brown trunks, and subtle rim lighting that makes the foliage and scattered wildflowers look slightly magical - the background feels fabulous with drifting pollen motes, pale blue bokeh orbs, and faint shafts of dappled sunlight filtering through leaves. She is standing center-frame with hands loosely clasped at her waist, large expressive eyes gazing upward and a small, wistful smile on her face, while nearby ferns glisten with tiny dewdrops. Ambient audio is layered: light leaves rustling, distant multi-voice birdsong that resolves into a single clear chirp, a faint stream trickle and a gentle wind whoosh; a delicate piano arpeggio with soft bell tones is playing quietly to lift the mood. She inhales audibly (soft breath SFX) and then speaks in a soft, reflective, slow voice: \"It's so quiet here...\" as she tilts her head and closes her eyes; after a brief pause she murmurs in a light, hopeful, slightly higher voice, \"I never thought I'd find this place,\" timed with a small smile and the dress whispering on the next breeze (fabric rustle SFX). A bright bird chirp answers in the soundscape just after her second line, and she gives a soft, airy laugh (gentle laugh SFX) as the piano bell lingers and the forest ambience continues."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60080", "video_prompt": "A medium shot of a young female brown bear named Joy standing in a small sun-dappled forest clearing, her dense brown fur catching soft afternoon light and a simple red scarf tied around her neck; she is stepping forward with gentle, deliberate paws on a leaf-strewn floor (soft padding on dry leaves), head tilted and bright eyes focused as she listens, then she leans slightly forward and extends one paw in a friendly, open gesture (scarf rustle). Background ambience is warm forest sound-distant birdsong, a faint brook burble, and a light breeze moving leaves-synchronized so the footsteps occur as she enters and the paw stretch aligns with a quiet rustle. Joy speaks in a warm, low, friendly voice, clear and patient: \"Hi, I'm Joy. I can understand you, and I'm here to help,\" followed by a soft, amused chuckle; her mouth moves in time with the words, then she nods once and offers a small reassuring smile while lowering her paw as the forest ambience continues under a final gentle exhale. "}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60113", "video_prompt": "A wide view over rolling green Welsh hills is bathed in low soft sunlight from the west, the sun is shining through a thin veil of high clouds and skimming the rounded ridges; short grass and scattered heather are bending and rippling under a steady wind that is sweeping across the slopes, and small puffs of dust are lifting from thin paths as the breeze moves; a low grey drystone wall with lichen-streaked stones is tracing the contour of a slope, its rough texture catching side light while patches of gorse with small yellow flowers are trembling and shedding a few dry petals; thin white clouds are drifting across the pale blue sky and momentarily sliding cool shadows down the hillsides as the sun reappears; audio: a constant low whoosh of wind is filling the scene as the grasses bend, layered with the near rustle of stems and the soft creak and tumble of a loose stone shifting in the wall, and a clear distant skylark is trilling a sustained phrase overhead, its high note rising and fading while the wind continues; the frame is steady, showing only natural small motions-grass leaning, flowers quivering, light moving across the slopes-conveying a calm, windy sunlit moment on the Welsh hills."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60133", "video_prompt": "In a static medium-wide shot from slightly behind and to the boy's right, a boy of about 12 with short dark hair and a navy hoodie is sitting cross-legged on a wool blanket on a low grassy knoll, watching the night sky as the Milky Way is arching overhead as a pale white band with a faint purple haze that is slowly drifting westward; soft starlight and a thin crescent moon are casting a cool, diffuse glow across his face and the textured knit of his hoodie. He is tilting his head back and tracking the shifting band with steady eyes, fingers tapping once on his knee, then shifting his weight so the blanket rustles. Ambient night sounds fill the scene: a gentle wind whispering through tall grass, steady cricket chirps, a distant owl hoot, and the soft rustle of fabric; his breath is audible as a small inhale. Boy (soft, awed, low voice): \"It's... actually moving.\" He exhales, smiles slightly, then leans forward and points with one hand toward a brighter patch of stars. Boy (quiet, slow, wonder): \"Look at that-like a river in the sky.\" A subtle, warm ambient pad underlies the natural soundscape, supporting the moment without overpowering the night ambience as the Milky Way continues its slow motion and the shot holds the quiet tableau."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60194", "video_prompt": "A tight close-up on a young person in their early 20s with short dark hair and subtle freckles, eyes wide and open to a cool wind that is ruffling short hair and moving a loose strand across the cheek; soft side light from a low sun casts gentle highlights on the skin and a faint rim light on the hair, while the out-of-focus background shows a pale grey sky and blurred treetops. Their gaze is slightly upward, pupils slowly widening as air moves across the face, eyelashes trembling and then blinking; they inhale quietly, shoulders shifting in a small, visible breath, and exhale as the wind lifts finer hairs. They speak twice in a close, intimate delivery synchronized with the breaths: They (soft, breathy, low-pitched) says, \"It's okay...\" while looking upward and letting a slow blink hold, then after a brief pause They (quieter, steady, low-pitched) says, \"I'm here,\" as the lips part and the wind pushes the lashes. Audio layers: foreground wind whooshes that vary with the hair movement, small crisp rustles of distant leaves, a subtle distant urban hum under the wind, close-mic'd inhalation and exhalation, and a single sustaining cello tone that rises gently under the second line and fades with the wind; all sound is timed to the visible breaths, hair movement, and spoken lines. Textures and details are visible in the five-second moment: moist eyes reflecting the sky, fine skin pores, soft sheen on the lips, and individual hair fibers moving against the cheek."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60198", "video_prompt": "Rendered in ultra-detailed 8k, a LEGO Super Mario minifigure (red hat with an M, blue overalls, white gloves, brown mustache, printed cheerful face) is standing on a glossy green baseplate with visible studs; soft overhead studio light and a cool rim light are creating small specular highlights on the molded plastic. He is raising his right arm and pressing a small red action button on the baseplate; precise plastic clicks and a short electronic beep occur as he moves. As he presses, a circular portal is opening behind him - a translucent ring of shifting teal and purple energy with swirling particle motes and a subtle grid-like shimmer that casts colored reflections across Mario's face; a deep harmonic drone is rising and a whoosh of wind-like synths layers with a brief chiptune arpeggio. The single camera is zooming in slowly from a medium shot toward an extreme close-up on Mario's face and the rim of the portal, tightening on the reflection of the swirling colors in his printed eyes while a quiet mechanical whirr from the zoom accompanies the sound. Mario (bright, energetic, mid-pitched, quick) says as he presses the button and leans forward, \"Let's-a go!\" - his voice is clear and playful and matches the moment the first light blooms. The portal answers with a low, resonant, echoing whisper (slow, hollow), \"Come...\" timed as tendrils of energy unfurl, and Mario blinks and inhales sharply, then exclaims (surprised, higher, short), \"Whoa!\" as the camera tightens on his expression; throughout, small plastic clacks track his movements, the portal rings with sporadic crystalline chimes, and the underlying electronic drone settles into a lower, distant hum that suggests another world beyond."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60209", "video_prompt": "In a quiet medium shot, a woman in her late 20s with chestnut hair tied in a loose bun is sitting alone on a flat rock at the edge of a calm, glassy lake, wearing a light-gray sweater, dark jeans, and scuffed brown boots; soft late-afternoon light is falling across her face and the water while the surrounding trees are full green and gently rustling. She is sitting cross-legged and leaning slightly forward, shoulders relaxed, eyes fixed on her reflection, then reaches a hand down and is lightly tracing the water surface with her fingertips, sending small concentric ripples as she breathes slowly and lets out a soft sigh. Ambient audio layers are detailed and timed to motion: steady, gentle lapping of water against stone, leaves rustling in a light breeze, distant single bird calls and a low insect hum, all underscored by a sparse piano motif-slow, single notes at low volume-that swells softly as she moves her hand; when her fingertip touches the water a delicate splash and the faint rustle of her sweater are audible. In a low, steady voice she says, \"I don't know who I'll be next,\" pausing to look at the ripples and glance up toward the treeline (a short, thoughtful silence follows), then in a softer, accepting tone she adds, \"But maybe that's okay,\" as she tilts her head, exhales audibly, and allows a small, contemplative smile to form while her gaze returns to the lake."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60225", "video_prompt": "In a medium close-up, a skeleton is seated on a low white cloud in a bright heaven, wearing a simple white linen sash over one shoulder; its bone surfaces are smooth with faint cracks and the jaw is articulated so the expression reads relaxed. Soft golden backlight and diffuse white cloud light wash the scene while distant pearly spires and floating islands sit out of focus behind. The skeleton is lifting a polished silver fork in its right hand and a carving knife in its left, slicing a medium-rare steak on a porcelain plate balanced on its lap; the steak shows a seared brown crust and a warm pink center that gives a quiet sizzle as the knife passes. As it brings a forkful to its jaw, a dry, gentle clack of bone is audible, followed by a muted, contented chewing sound; it speaks in a low, amused voice, slow and warm, \"Well, this is unexpected,\" then pauses, glancing down at the plate and answers itself in a softer, wry tone, deliberate and mid-pitch, \"Heaven could use better menus,\" while a thin harp arpeggio and a distant choral pad swell beneath, soft wind through clouds and faint bell chimes punctuating the air; knife-on-plate clink, fork lift, and the final swallow are audible as the skeleton settles back slightly, a small puff of cloud compressing under its weight."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60254", "video_prompt": "A static medium close-up of a cottage window at night, pale full moonlight is washing a cool rectangle across the weathered wooden sill and slightly wavy glass; the right-hand casement is slowly swinging outward on aged iron hinges, the chipped white paint and rough grain are visible in the moon glow. As the pane is swinging open, a low creak from the hinges is audible, then a soft scrape as the old latch releases; a thin linen curtain is fluttering inward and its edge is brushing the sill with a quiet rustle. Outside, steady cricket chirps are underscoring the moment while a distant owl hoots once and a light breeze is whispering through nearby leaves so branches are sighing faintly. The widening opening is letting more moonlight spill into the dark interior, highlighting dust motes drifting in slow arcs and casting a pale band that is sliding across the floor."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60265", "video_prompt": "In an extreme close-up, a shallow pile of whole almonds and hazelnuts rests on a matte dark slate surface under soft overhead light that picks out subtle ridges on the almond skins and the rounded texture of hazelnut shells; a thick stream of melted dark chocolate is pouring in from above and is splashing down onto the pile, coating some nuts and sending tiny droplets outward. As the chocolate is striking the nuts, a few almonds and hazelnuts are nudged and spin slightly while a small piece of hazelnut shell is flicked aside; droplets arc and hang briefly in slow motion, catching highlights on their glossy surfaces, then fall and merge into a spreading pool around the nuts. Visual emphasis is on the contrast between the glossy chocolate and the matte nuts, with one hazelnut showing a cracked interior as chocolate runs over it. Audio begins with the deep, viscous pour of chocolate-a low glug and a soft, sticky slap on contact-immediately joined by crisp, dry clacks as nuts bump one another and a faint high-frequency tinkle as tiny droplets hit the slate; a muted kitchen ambience (distant HVAC hum and soft background murmur) sits underneath, then the sound settles into a gentle wet spread and quiet drips as chocolate spreads around the almonds and hazelnuts."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60278", "video_prompt": "Dong Yuhui is standing in a vast open plain under a pale sky, soft late-afternoon light falling across the scene; he is a man in his early 30s with short black hair, the breeze is blowing through his hair and ruffling the edges of a light gray jacket over a white shirt, and he is holding an open hardcover book in his left hand with the pages slightly fluttering. His face is turned toward the horizon, eyes firm and full of hope, the corners of his mouth slightly raised in a positive, energetic smile; his posture is upright, chest lifted and shoulders back, his whole body language is full of vitality and vigor, conveying clear confidence and optimism. While the wind makes a low whoosh through the grasses and the book emits a soft paper rustle, he breathes in, glances down at the open page, then looks up and speaks in a steady, warm, confident voice, \"We can do this,\" the words timed with a small, assured nod. Underneath, a gentle single-note piano chord swells as distant bird calls punctuate the air, the wind continues to whisper, and the final sound is the quiet flutter of pages as his smile holds, leaving a calm, positive, upward feeling."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60291", "video_prompt": "A medium close-up of a man in his late 50s with short graying hair and a neat trimmed beard, seated behind a low glass console in a futuristic presidential chamber that blends dark wood, brushed steel and frosted glass; he wears a tailored dark gray suit with a high-collar shirt and a small stylized tricolor lapel pin, his expression calm and purposeful as translucent holographic maps and data widgets float a foot above the console. He reaches out and taps a hovering map, which ripples into sharper focus while a thin blue route highlights across Eurasia; his fingers move with deliberate, practiced gestures and his eyes track the route as if reading several layers of data at once. Ambient sound is a low, steady hum of climate systems and distant city traffic through thick glass, punctuated by soft glassy chimes on each tap and a brief electronic confirmation beep when he activates a layer; a sparse, subdued string motif plays quietly under the scene, rising slightly when the map refocuses. President (voice: measured, low, paced): \"Activate strategic overlay, scale to national,\" he says, then pauses, glances briefly to the side as if checking a monitor, fingers hovering above the interface. President (voice: softer, deliberate): \"Keep civilian channels open - alert level stable,\" he adds, and presses his palm once, sending a succinct confirming tone as the hologram contracts; his face tightens for a beat, then relaxes as the room returns to the steady ambient hum."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60317", "video_prompt": "A medium close-up static shot of an androgynous angelic being is hovering a few feet above a smooth pale floor, wearing a flowing white robe with thin silver embroidery and long silver-white hair falling over the shoulders, face composed and gently focused. Soft cool backlight and a warm rim light are outlining the figure as large luminescent wings are slowly unfurling behind them, each feather translucent with an inner soft glow that pulses in pale gold and cool blue while tiny motes of light are drifting off the wing tips and catching on the silk texture of the robe. As they are extending both hands forward, the aura around them is shimmering in slow, concentric waves and their robe is fluttering slightly from a light upward lift; they are tilting their head and allowing a quiet, serene smile to form. A sustained, gentle choral pad is filling the air underneath the scene, with a single bell-like glissando punctuating the moment the wings open; soft rustle of fabric and whisper of feathers are synchronized to the wing motion, a faint whoosh of displaced air accompanies the hover, and subtle harmonic overtones are rising and falling with the aura's pulse. (voice: calm, warm, mezzo) \"I am here,\" they say, voice steady and low as their palms open, the words timed to the wings' full spread; a brief pause lets the motes scatter. (voice: soft, distant, reverent) Then they adds, voice trailing as they incline their head and close their eyes briefly, \"Stay near the light,\" while the glow around them is gently pulsing one last time."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_65", "video_prompt": "A young woman in her early 20s is standing in a medium full shot on a narrow stone alley lit by paper lanterns, wearing a fitted knee-length red silk qipao with pale pink peony embroidery and smooth white stockings; soft warm lantern light grazes the silk while cool evening blue fills the alleyside shadows, the stone underfoot showing a faint sheen as if damp. She has long black hair gathered into a loose chignon with a simple silver hairpin, subtle makeup, and a calm, attentive expression as she shifts her weight onto one foot and smooths the fabric at her hip, then reaches out briefly to brush a lantern's fringe with two fingers while her skirt whispers against her thigh. Ambient alley sound settles beneath-low murmur of distant conversation, a far-off bicycle bell, a small trickle of water from a nearby drain-while a single-note erhu phrase plays softly and periodically, matching her gentle movements; each step she takes produces a soft footstep on wet stone and a quiet rustle of silk. She breathes out, then speaks in a soft, warm voice, \"The lanterns look quiet tonight,\" as she glances down and then up toward the lights, a small, thoughtful smile forming; fabric rustle and a light creak from the lantern sway punctuate the moment, and the erhu holds a final quiet phrase as she tilts her head and the scene lingers on her composed, serene posture."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_70", "video_prompt": "A medium close-up of Lucas, early 30s, clean-shaven with short dark hair slightly tousled, wearing a navy blazer over a white shirt, standing on a small stage in a dim conference hall under a soft spotlight with a faint technical schematic projected on a screen behind him; he is leaning slightly forward, eyes bright, smiling, right hand lifting in an open, inviting gesture while his left holds a handheld microphone. Lucas (eager, breathy, mid-tempo voice): \"This is it, my friends. Humanity&#39;s next step.\" As he finishes the line he steps a half pace forward and lifts both hands briefly, adding in a quick, confident tone, \"We're ready to begin,\" (confident, slightly faster) while a low audience murmur swells into a single short cheer; ambient room sound includes a soft HVAC hum, distant chair scrape, and a subtle microphone rustle when he moves the mic. The stage light casts soft shadows across his jaw and the projector's faint flicker glints on the nearest rows; his brows lift and a slight smile tightens at the corners of his mouth, the fabric of his jacket shifting as he breathes and the crowd reacts, all within a compact, energetic five-second moment."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_74", "video_prompt": "A medium wide shot of a low, pebbled shoreline in soft late-afternoon light, a boy about eight with short brown hair and a faint smudge on his cheek is walking barefoot toward the water wearing a damp navy windbreaker and tan shorts, taking small, careful steps across wet stones while his jacket sleeves brush his arms. Gentle lap of small waves provides the ambient sound, with distant gull calls and a light breeze whispering through nearby reeds; as he draws within two meters the gravel crunches under his feet, a soft fabric rustle from his jacket is audible. He slows, eyes on the moving water, inhales audibly, and in a quiet, hesitant voice says, \"Okay... here goes,\" then leans forward and steps so his toes skim the shallow edge, producing a soft splash and a brief spray; the water makes a delicate fizzing sound against his foot and he lets out a small relieved breath while the wave withdraws and pebbles settle back into place."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_76", "video_prompt": "A stationary overhead camera looks straight down through clear, shallow blue water as a large adult whale shark is gliding slowly from left to right across the frame, its broad head, pale gray-blue skin patterned with white spots and faint stripes, and a tall dorsal fin visible beneath a gently rippling surface; soft sunlight is casting moving caustic bands across its back and the sandy seafloor below, the shark's rough, ridged skin showing subtle barnacles and a few old scars. As it swims, the shark is opening its wide mouth slightly to filter-feed while its tail is undulating in smooth, powerful strokes, a dorsal fin brushing the surface and sending tiny concentric ripples; a small group of pilot fish is trailing close behind, then a nearby cluster of baitfish is scattering in a quick burst, darting away in sharp, synchronized motions. Ambient audio begins with a low, muffled ocean hum and a distant, slow whale-like call underlining the scene; as the shark passes a soft whoosh of displaced water and a deep, low rumble follow each tail flick, faint bubble pings are audible near its mouth, and the baitfish scattering produces brief, high-frequency splashes and quick metallic plinks. The overall mood is calm and observant, the overhead view holding steady as the whale shark continues gliding out of frame."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_86", "video_prompt": "A medium close-up of a young man in his mid-20s, based directly on image:sketch (1).jpeg as visual reference, standing slightly angled toward camera with his shoulders relaxed; his dark, slightly long hair is moving as a steady wind is sweeping across the scene, ruffling strands and lifting the collar of his loose grey shirt. A soft blue-white glow is surrounding him like a faint halo, the glow is pulsing gently and casting a cool rim light on the edges of his face and hair while the rest of the background remains a muted charcoal sketch texture; small ink-like particles are drifting through the air and catching the light as they float. He is breathing out slowly, exhaling as a gust of wind pushes his hair back, and his expression is calm, eyes narrowing slightly as if listening. Soundscape: a clear, close wind whoosh is present and is rising and falling with each gust, soft rustling of fabric and hair on each breath, and a low, steady electrical hum is synchronizing with the glow pulses; beneath that, a faint sketching scratch like pencil on paper is barely audible to tie to the reference image. Dialogue synced to actions: He (voice: low, steady, medium pace) says, \"I can feel it,\" as he exhales and lets his chin lift; an inner voice (voice: soft, breathy, slow) replies, \"It's starting,\" as the glow brightens for one pulse; he (voice: low, steady) answers again with a softer tone, \"Then stay with it,\" as his hair settles slightly and his shoulders shift; the inner voice (voice: airy, distant) whispers, \"I am,\" timed with the final, small gust that makes the particles drift away. The overall color palette is cool greys and blue-white light, textures are drawn-paper and soft fabric, and all motion is continuous and subtle to fit within a brief five-second moment."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_93", "video_prompt": "In a wide shot, a small round bunny with soft white fur and faint gray patches is sitting near the crest of a gently sloping hill blanketed in short green grass and scattered wildflowers-daisies, buttercups, and small bluebells-bathed in soft midday sunlight; the bunny twitches its nose and lifts slightly, then springs forward in a quick, joyful bound, ears tilting back and hind legs tucking under as it clears a patch of flowers, petals fluttering down and a few stems bending under its passage, then it lands lightly with forefeet touching first and hind feet following, sending a muted thump and a soft rustle through the grass. Ambient sound begins with a light breeze through the grass and distant birdsong, joined by a steady, close bee buzz around the blossoms; as the bunny pushes off there is a small puff of displaced air and the faint crunch of stems, and on landing the bunny lets out a short, bright chirp. Bunny (soft, high-pitched, quick): \"peep,\" synchronized with the landing and a brief head tilt; it then sniffs a daisy, nose wrinkling and whiskers brushing petals, ears flicking, while the breeze and bees continue softly in the background and a few petals settle back onto the hill."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60343", "video_prompt": "Wide shot of a small village under late afternoon light: clusters of low stone cottages with red tile roofs, a low church steeple, and a narrow cobbled lane between them; warm, soft light casts long, thin shadows and brings out the rough texture of stone and the matte clay tiles. Several groups of birds are flying across the pale sky, each flock is shifting shape as birds are soaring, banking, wheeling, splitting off, then rejoining-some individuals are gliding on outstretched wings while others are beating rapidly to gain altitude; their moving shadows skim across rooftops and the lane below as a group arcs together over the steeple. Soundscape: close, soft flapping of wings layered with overlapping bird calls-higher, quick chirps interspersed with occasional low caws-while a single distant church bell tolls slowly once, a light wind is rustling through poplar leaves, a loose wooden shutter creaks and a clay chimney pot clicks as the breeze passes, all synchronized so the wingbeats and calls rise as the flocks bank and briefly swell, then fall away as they split and move out of frame."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60350", "video_prompt": "In a medium close-up, a man in his late 20s with short dark hair and light stubble is cupping a 3-week-old black and white tuxedo kitten with blue eyes in both hands, wearing a light gray cotton shirt; soft window light is falling across their faces and a warm indoor tone is bathing the scene. He is holding the kitten gently, palms cradling its tiny body as he is stroking the soft downy fur with his thumbs; the kitten is calm and relaxed, blinking slowly, paws tucked against his palms and its tiny pink nose twitching. The man smiles and speaks in a low, warm voice, slow pace, \"Hey little one...\" (he leans in slightly, eyes soft, fabric rustle audible), then the kitten replies with a soft, high mew (quiet, brief) while blinking and beginning a steady, quiet purr that grows slightly louder as it settles against his hands. The man responds in a tender hush, gentle pace, \"You're so small, aren't you?\" as he exhales and brushes his thumb along its back; the kitten emits another faint mew and nuzzles his thumb, purring more audibly. Ambient audio layers are intimate and close: a muted distant street hum through the window, a faint clock tick, soft breathing from the man, light fabric rustle when he shifts, the kitten's small mews and continuous low purr present and clear, and a subtle room reverb to convey a small, quiet indoor space."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60360", "video_prompt": "Style: Tim Burton Style Characters. In a dim, low-lit room a lanky, slightly elongated Santa is standing beside a tall, narrow window, peering at a smartphone held in his gloved hand; he has a narrow pale face, wide dark eyes, a scraggly white beard, and a worn red coat with thin fur trim and long sleeves that brush his knuckles. Soft moonlight is filtering through the window, casting thin shadows across his coat while a small warm desk lamp on the sill gives a muted rim light; a cinematic lens with shallow depth of field is rendering the distant streetlights outside as soft bokeh and subtle film grain. The smartphone screen is glowing with a map full of clustered red dots and tiny car icons showing a traffic jam; he is zooming and tapping the screen, brow furrowing as the map shifts. Room audio is quiet: a faint, steady traffic rumble and occasional car horn from the street below, a soft creak as he shifts his weight, and a delicate notification chime from the phone. In a low, gravelly, slow voice he mutters, \"Forty-two minutes? That's too long,\" then he is tapping to try an alternate route as a neutral, clipped phone assistant voice says, \"Delay ahead: 42 minutes, suggested detour adds ten minutes,\" synchronized with the map redrawing under his thumb; he glances up through the window at the stalled tail lights, exhales, and in a short, resigned tone replies, \"All right, take the detour,\" then tucks the phone into his coat as the distant traffic hum persists and the lamp casts a soft, low shadow across his face."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60390", "video_prompt": "A wide static view of a coastal landscape on a sunny day: the sun is high in the pale sky, casting soft overhead light and subtle shadows across smooth wet sand and scattered pale pebbles while small waves are lapping at the shore and thin white foam is retreating with a gentle hiss as sunlight glints on the ripple tops. As a light breeze is bending the dune grass and a few palm fronds, several seagulls are circling low and calling with short sharp cries, and a distant sailboat is drifting near the horizon with its canvas lightly billowing. The soundscape matches the motion-soft rhythmic surf washing onto sand, the thin hiss of foam receding, wind whispering through the grass, intermittent gull calls, and an occasional soft metallic clink from the sailboat rigging, all underscored by a faint, slow acoustic guitar picking a quiet two-bar motif to add calm warmth to the scene."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60404", "video_prompt": "A medium close-up shows Eveline, a young sorceress in her early 20s with waist-length dark brown hair braided loosely over one shoulder, wearing a moss-green cloak over a simple linen dress and scuffed leather boots, kneeling on soft emerald moss in a serene sun-dappled forest; soft overhead light filters through tall oaks, casting moving leaf shadows and the scene is textured with damp bark, tiny wildflowers, and dewy spiderwebs. Ambient forest sounds-distant birdsong, a gentle breeze rustling leaves, a nearby creek's quiet trickle-form a calm background; her soft footsteps on moss and the quiet inhale of her breath punctuate the space as she is experimenting with magic, tracing a small sigil in the air with a trembling finger while murmuring an incantation, \"Focus... steady,\" (soft, hesitant, low) and tiny golden motes are gathering at her fingertip with a faint bell-like chime and gentle electrical crackle. She is tilting her head, squinting in concentration, then is pushing her palm forward and a warm pale orb of light is lifting from her hand, spinning slowly as leaves and dust begin orbiting it; the orb is emitting a subtle harmonic drone and delicate tinkling SFX as it is stabilizing. Eveline is exhaling sharply, \"It's... it's working,\" (breathless, surprised, slightly high) and she is leaning closer, eyes widening and mouth parting in visible awe while small vines at her feet are curling toward the light. She is straightening with a small, stunned smile, whispering, \"I can do this,\" (quiet, resolute, calm) as the forest ambience swells-a soft wind gust, a brief chorus of birds-and the orb is pulsing once with a warm luminous beat that is bathing her face in soft gold, capturing the brief, clear moment of realization and possibility."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60416", "video_prompt": "Style: cartoon. In a static medium-wide shot, a cartoon girl of about 10 sits on a smooth gray stone ledge at the base of a narrow waterfall, legs dangling and hands resting on the rock; she wears a teal hoodie and denim shorts, brown shoulder-length hair with a small side braid, and a soft, contented expression as she looks out at the view of a tree-lined valley beyond. Soft midday light filters through green leaves, the waterfall is drawn with clear blue-green water and white froth, and a fine mist makes tiny droplets bead on her knees; a light breeze moves her braid and the hem of her hoodie. The water is continuously audible as a steady, low roar; higher, crisp splashes hit the stones and a few bright bird calls and a faint rustle of leaves sit beneath the falls. She shifts her weight, tucks her braid behind her ear, blinks, and exhales slowly as the mist lands on her face (subtle wet SFX). Girl (soft, reflective, slow): \"It's so peaceful here...\" (her voice is intimate, slightly breathy; waterfall audio ducks gently while she speaks). She smiles, presses her palms to the stone, and says with a small laugh, Girl (light, quick): \"Maybe I should come back with a sketchbook.\" After her last word the waterfall volume returns to full, droplets patter, and the scene holds with distant birds and steady water ambience."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60454", "video_prompt": "In a static medium-wide shot, a woman in her late 20s with wind-tousled dark hair is sitting on a smooth driftwood log at the sandy shore, wearing a light gray knit sweater and worn jeans, facing the sea as soft late-afternoon light washes the scene and low sun glances off the wet sand. Small waves are lapping and pulling at the shoreline, seafoam is pulsing around scattered pebbles, and a cool breeze is ruffling her hair and sweater; distant gull calls punctuate the air and a low foghorn is sounding far off. She is watching the horizon, fingers lightly tapping the log, then draws a slow breath and says in a quiet, steady voice, \"Calmer today, huh?\" She glances down at a phone in her lap; an off-screen voice, warm and slightly amused, replies quickly, \"Yeah - you found the quiet spot.\" She lets out a small smile, tucks hair behind her ear, answers softly, \"I needed this,\" and turns her gaze back to the sea while the rhythmic hiss of waves and gentle wind continue under the brief exchange, with occasional seabird calls and a faint distant dog bark adding texture to the seaside ambience."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60489", "video_prompt": "In a medium close-up, Alexander Lukashenko, a white man in his early 70s with short gray hair and a clean-shaven face, wearing a dark navy suit and a striped tie, sits at a low podium against an out-of-focus backdrop showing a red-and-green flag; soft overhead light and a subtle rim light outline his shoulders as the scene focuses on his face. He is smiling, then leans forward slightly, then tilts his head back and throws it back as he bursts into a short, throaty laugh-his shoulders shaking and eyes crinkling-while one hand comes up briefly to his mouth and then drops to the podium. The audio layer presents a close, gravelly laugh in the foreground, layered with light applause and murmured reactions from an off-screen audience, the faint rustle of papers, and a soft microphone rustle; as the laugh begins he utters one dry, measured line in a low, even tone, \"Well, that's unexpected,\" (low, dry, slow), then follows with a quick, breathy chuckle (deep, warm) that overlaps the room noise. The camera remains steady in the medium close-up, capturing the sequence of smile, voiced remark, and sudden laugh while ambient sounds - soft clapping, hushed voices, a single camera shutter click - sit behind the laugh and then settle as the laughter fades."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60498", "video_prompt": "A static close-up frames a Spider-Man action figure standing on a worn wooden shelf, painted red and blue suit with glossy plastic highlights and visible articulated joints; soft overhead light casts mild reflections across the molded webbing and a shallow depth of field blurs the background books. He is slowly tilting his head to the left as his white eye lenses curve upward into a clear smiling expression, then he raises his right hand in a short, friendly wave while a crisp plastic joint click and a faint squeak register with each motion. Room ambience carries a distant city hum and soft indoor air movement; as his lenses tighten into the smile a brief, gentle synth twinkle accents the moment. Spider-Man (playful, slightly high-pitched, quick) says, \"Hey-ready?\" as he lifts his hand, the voice timed to the click of the wrist joint; he pauses, leans forward a hair, and the painted smile seems to deepen with a tiny plastic rasp. Spider-Man (warm, amused, medium pace) follows, \"You always are, right?\" while he tilts his head back and his shoulder rotates with a soft mechanical click. Spider-Man (confident, brisk, low) finishes, \"Okay-on three,\" as he nods once and the light catches a small scuff on the painted chest; a short, bright chime punctuates the final grin and the ambient city hum continues under the closing moment."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60508", "video_prompt": "A matte black Porsche 911 is cruising along a narrow, leaf-strewn forest road in the evening, its low headlights cutting through dim blue-gray dusk while long cool shadows from pine and birch stretch across the pavement; soft, warm red taillights smear over wet leaves as the car moves, and small sprays of leaves lift and scatter behind the rear tires. The engine is emitting a low, measured purr that rises briefly as the driver downshifts while negotiating a gentle curve, tires softly crunching on damp leaves and loose gravel beneath the chassis; occasional low branches brush the hood and send a faint metallic whisper along the bodywork. Ambient sound is a quiet forest layer-wind rustling needles and leaves in the canopy, a distant bird call, and the echo of the engine bouncing between tree trunks-then the subtle click of the turn signal and the whisper of airflow as the Porsche arcs past, remaining alone in the calm evening woods."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60528", "video_prompt": "A quaint village nestled among rolling hills and surrounded by peaceful trees is bathed in a sky of orange and pink as the sun begins to set, soft late-afternoon light warming stone cottages and small gardens while gentle rustling of leaves and the sweet melodies of birds fill the air. After two seconds, the screen slowly zooms in, then over the next three seconds the camera slowly zooms in and pans, offering a picturesque view of the clustered homes before gracefully pointing toward a nearby pond that is shimmering in the warm sunlight; the pond's surface mirrors the colored sky and surrounding trees with glinting highlights moving across the water. As the focus narrows on the pond, three vibrant orange fishes are swimming in graceful arcs just below the surface, their scales glistening as sunlight catches them while they glide in small, coordinated loops and send soft ripples outward. Audio is layered to match the motion: continuous leaf rustle and clear bird song in the foreground, then gentle water lapping and faint, synchronized splashes as each fish breaks the surface of a ripple, keeping the overall soundscape calm and harmonious with the village remaining quiet in the background."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60532", "video_prompt": "A medium close-up of a young woman in her late 20s with short black hair tucked behind her ear, wearing a light gray blouse, sitting at a small wooden table at a street-side food stall as soft late-afternoon light falls across a glossy ceramic bowl of mie ayam-yellow noodles piled with shredded braised chicken, chopped scallions and a few green vegetables, dark sauce pooling at the bottom; she is lifting a tangle of noodles with wooden chopsticks, twirling them once, then slurping them into her mouth while her free hand is scooping broth and chicken with a metal spoon, her eyes closing briefly and a small smile forming. Ambient street audio is present-distant traffic hum, a motorbike idling, low vendor chatter and occasional clink of dishes-while close SFX highlight wooden chopsticks tapping porcelain, a wet noodle slurp, the soft clink of spoon on bowl and a quiet breath; a gentle instrumental with soft percussive tones plays under the ambient mix. As she sets the chopsticks down she says in a warm, measured voice, \"Looks comforting,\" pauses and glances around, then in a soft, content voice adds, \"Just what I needed,\" and reaches for another bite."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60538", "video_prompt": "A static medium-wide shot frames a low, moss-covered rock ledge and the cool pool below, soft afternoon light filtering through a leafy canopy and casting dappled patches across dark green water; the rock is slick with thin rivulets and small, pale leaves cluster at the edge. A small, smooth pebble is teetering on the lip, then tumbles and plunges into the cool pool below, breaking the mirror surface and sending concentric ripples outward while tiny droplets arc upward and scatter. Under the surface a brief ring of bubbles is rising and dissolving, and as the ripple reaches floating leaves they bob and shift. Audio: an immediate, sharp splash on impact, followed by spreading, gentle slaps of water against stone, a soft underwater gurgle as bubbles release, faint steady trickle from the rock face, distant birds calling and a low insect hum; after the initial splash the water calms and small droplets patter on leaves while the echo of the impact fades."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60546", "video_prompt": "A happy child of about eight years old with short curly hair and subtle freckles is sitting cross-legged on a small red picnic blanket in a sunlit backyard, wearing a striped t-shirt and denim shorts, holding a small soprano ukulele with a light wood grain and glossy finish; warm late-afternoon light is casting soft shadows and a few leaves are drifting in a light breeze. The child is smiling and is strumming a bright, bouncy chord progression with the right hand while the left hand is forming a simple C-to-G chord change, the ukulele strings ringing with clear plucked tones; the child is gently swaying and bobbing their head as they sing one short, cheerful line in a light, singing voice, Kid (singing, bright, slightly high): \"Sunshine on my day, play along, hey!\"-they finish the phrase with a quick, playful flourish across the strings, then laugh and look up, Kid (delighted, breathy): \"Again!\" as they tap the uke body with their thumb and launch into another upbeat strum. Audio includes close, intimate ukulele sound with crisp string attack and soft resonance, the child's voice layered slightly forward, faint backyard ambience-distant birdsong, a soft breeze through leaves, a muted lawnmower far off-and the small SFX of fingers on wood and a quick giggle; everything is synchronized so the vocal lines match the visible strums and the final tap coincides with the child's delighted exclamation."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60571", "video_prompt": "In a wide, low-angle shot from the highway shoulder, three super-fast sports cars are racing side-by-side along a wet night highway, a red low-slung coupe with a white racing stripe on the left is surging forward, a matte black aggressive supercar in the center is pulling slightly ahead with sharp LED accents cutting through light mist, and a metallic blue targa on the right is darting forward; glossy body panels catch and reflect neon from the big city glowing in the distance, soft sodium streetlamps pool on the slick asphalt, and heat shimmer ripples above each exhaust. As they accelerate, engine roars dominate the mix- the red coupe emits a high, razor-edge rev, the black car answers with a deep, throaty growl, and the blue car adds a fast turbo whistle-interleaved with rapid upshifts and short throttle blips; tires hiss and briefly squeal during tight lane corrections, small pebbles crack against wheel wells, and a sharp whoosh of wind whistles past side mirrors. Ahead, the distant city skyline of clustered towers and billboard neon provides a low, constant urban hum with faint car horns and a muted subway rumble beneath, while the Doppler shift bends each engine pitch as the cars edge forward and then tuck in, motion blur streaking headlights and building lights to emphasize raw speed for the five-second burst."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60572", "video_prompt": "Style: hyper realism 4k. A static tight close-up of a single candle sitting in an aged brass candlestick on a dark wooden surface, the creamy white wax is melting into a glossy pool that is slowly running down the stem while the blackened wick is holding a steady amber flame; 4k detail reveals the fine char texture of the wick, tiny beads of molten wax, and subtle fingerprints and patina on the brass. Soft warm light from the flame is casting gentle reflections on the candlestick and subtle, narrow shadows on the wood as a faint draft moves through the frame, making the flame flicker and lean; the wick crackles quietly and a thin metallic spark flicks at the tip as the flame thins, then a short airy whoosh is heard exactly as the flame collapses and goes out. A slender plume of blue-gray smoke is rising in a slow spiral from the still-glowing ember, the glowing tip dimming to a dull black nub while a few droplets of settling wax produce soft, muted clicks; ambient room hush underlies the moment with a low, unobtrusive room tone and a distant faint creak, and the residual sound of the extinguish - a soft breath-like hiss and the last faint pop - lingers as the scene holds on the quiet candlestick and the smoke disperses."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60589", "video_prompt": "A tight close-up of a brown-gray tabby cat with distinct black stripes and a faint white chin, lying on a soft cream knit blanket under soft side light from a nearby window; the cat blinks slowly, then leans in and presses its cheek and forehead into the blanket in a series of gentle, deliberate nuzzles while its whiskers brush the fabric and its ears flick slightly. Its green eyes remain half-closed and its shoulders shift forward as it rubs, then tilts its head and repeats another short, purposeful rub; the fur has a soft, slightly glossy texture and small tufts move with each motion. Ambient indoor room noise is quiet-a low, steady HVAC hum and distant muted street sound-then a low, steady purr begins as the cat first nuzzles, rising into a warm, continuous rumble synchronized with the rubbing. Each contact is accompanied by a soft fabric rustle and a faint breathy exhale; after the final nuzzle the cat releases and emits a brief, contented trill before settling, the purr continuing gently under the room ambience."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60601", "video_prompt": "In a medium close-up, a couple in their late 20s is sitting across a small round wooden table in a cozy cafe, the woman with shoulder-length dark hair and a cream sweater holding a white ceramic cup, the man with short brown hair, light stubble, and a navy button-down resting his forearms on the table; soft warm yellow pendant lights overhead are reflecting gently on their faces and creating small highlights on the cup and the glass vase between them, while the background is softly blurred with other patrons and warm wood tones. Low indie acoustic music is playing quietly from the cafe speakers as steady ambient murmur, occasional silverware clinks, and a distant espresso machine hiss form a continuous soundbed. She is smiling and laughing, tucks a strand of hair behind her ear, and says in a warm, laugh-tinged voice, \"I'm really glad we came out tonight,\" with a short pause and a quick smile toward him; a soft clink occurs as he sets his cup down. He is smiling back, leaning in slightly, and replies in a soft, even voice, \"Me too-it's been nice,\" pausing to meet her eyes while a chair scrapes gently in the background. She is glancing down at her cup, then looking up with a brighter smile and says in a quiet, upbeat voice, \"Let's do this again,\" as the music swells a fraction and the cafe ambience continues beneath their voices, leaving them both smiling and holding eye contact."}
{"id": "vidprom_semantic_unique_gpt-5-mini-2025-08-07_60632", "video_prompt": "A wide view of a lush forest of mixed hardwoods and fruit-bearing trees rising beside a calm lake, soft late-afternoon light filtering through green leaves and dappling mossy trunks. Low branches are heavy with red apples and yellow pears, their glossy skins catching the light as a gentle breeze is moving through the canopy-leaves are rustling and small clusters of fruit are swaying while a few water droplets bead on leaf edges. At the lake edge, reeds are bending slightly and the grey-blue surface is forming narrow ripples where the wind touches it; a dragonfly is skimming the water and a small fish is breaking the surface with a soft plop, sending a brief circular ripple outward. Ambient audio is layered and synced with action: the wind is creating a steady soft rustle through leaves, clear birdsong (quick sparrow-like chirps and a distant thrush) is punctuating the treetops, a low insect buzz is underscoring the scene, and gentle water lapping is meeting the pebbled shore; when a fruit brushes a branch a muted thud and the faint crunch of leaf litter register. The scene feels tranquil but active, natural motions continuing throughout the short clip."}
+18
View File
@@ -0,0 +1,18 @@
<svg width="252" height="105" viewBox="0 0 252 105" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM100.768 13.1217L87.7028 29.4852H103.143L100.768 13.1217Z" fill="#356CFF"/>
<path d="M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM109.081 90.697L116.802 65.8487C116.802 65.8487 120.959 65.8487 132.242 65.8487C143.525 65.8487 137.586 78.5759 135.211 84.0304C133.307 88.4021 127.491 90.697 122.74 90.697C117.989 90.697 109.081 90.697 109.081 90.697Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944H159.747C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273L125.188 48.273L124 37.97L147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852H131.836C120.142 29.4852 125.897 1.00043 141.337 1.00043L173.188 1.00056Z" fill="#356CFF"/>
<path d="M179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056Z" fill="#356CFF"/>
<path d="M161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM237.948 77.9692C239.984 70.6965 240.917 65.242 228.446 65.242C215.975 65.242 211.818 71.9087 210.037 77.9692C208.255 84.0298 208.255 91.3025 219.538 91.3025C230.821 91.3025 235.911 85.2419 237.948 77.9692Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944M173.188 1.00056C173.188 1.00056 156.777 1.00043 141.337 1.00043M173.188 1.00056L141.337 1.00043M141.337 20.3944C146.088 20.3944 150.839 20.3944 159.747 20.3944M141.337 20.3944H159.747M159.747 20.3944C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273M148.463 48.273C139.556 48.273 125.188 48.273 125.188 48.273M148.463 48.273L125.188 48.273M125.188 48.273L124 37.97M124 37.97C124 37.97 141.931 37.97 147.87 37.97M124 37.97L147.87 37.97M147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852M151.433 29.4852C146.682 29.4852 138.962 29.4852 131.836 29.4852M151.433 29.4852H131.836M131.836 29.4852C120.142 29.4852 125.897 1.00043 141.337 1.00043M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057ZM96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM87.7028 29.4852L100.768 13.1217L103.143 29.4852H87.7028ZM89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457ZM108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM116.802 65.8487L109.081 90.697C109.081 90.697 117.989 90.697 122.74 90.697C127.491 90.697 133.307 88.4021 135.211 84.0304C137.586 78.5759 143.525 65.8487 132.242 65.8487C120.959 65.8487 116.802 65.8487 116.802 65.8487ZM179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056ZM161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457ZM230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM228.446 65.242C240.917 65.242 239.984 70.6965 237.948 77.9692C235.911 85.2419 230.821 91.3025 219.538 91.3025C208.255 91.3025 208.255 84.0298 210.037 77.9692C211.818 71.9087 215.975 65.242 228.446 65.242Z" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M15.2524 55.5451L21.191 100.999L24.7541 100.999L18.8156 55.5451L15.2524 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 55.5451L14.065 100.999L15.2527 100.999L9.31417 55.5451L8.12646 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 55.5451L6.93853 100.999L7.53239 100.999L1.59385 55.5451L1 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M15.2524 48.2724L30.0988 1H33.6619L18.8156 48.2724H15.2524Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 48.2724L22.9728 1H24.1605L9.31417 48.2724H8.12646Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 48.2724L15.8463 1H16.4402L1.59385 48.2724H1Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M85.3271 55.5457H67.5116L87 12.7363L44.3513 68.2729H58.6038L43.1636 101L85.3271 55.5457Z" fill="#FDC717" stroke="#FDC717" stroke-width="1.18771" stroke-miterlimit="16"/>
</svg>

After

Width:  |  Height:  |  Size: 5.7 KiB

+18
View File
@@ -0,0 +1,18 @@
<svg width="252" height="105" viewBox="0 0 252 105" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM100.768 13.1217L87.7028 29.4852H103.143L100.768 13.1217Z" fill="#356CFF"/>
<path d="M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM109.081 90.697L116.802 65.8487C116.802 65.8487 120.959 65.8487 132.242 65.8487C143.525 65.8487 137.586 78.5759 135.211 84.0304C133.307 88.4021 127.491 90.697 122.74 90.697C117.989 90.697 109.081 90.697 109.081 90.697Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944H159.747C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273L125.188 48.273L124 37.97L147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852H131.836C120.142 29.4852 125.897 1.00043 141.337 1.00043L173.188 1.00056Z" fill="#356CFF"/>
<path d="M179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056Z" fill="#356CFF"/>
<path d="M161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM237.948 77.9692C239.984 70.6965 240.917 65.242 228.446 65.242C215.975 65.242 211.818 71.9087 210.037 77.9692C208.255 84.0298 208.255 91.3025 219.538 91.3025C230.821 91.3025 235.911 85.2419 237.948 77.9692Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944M173.188 1.00056C173.188 1.00056 156.777 1.00043 141.337 1.00043M173.188 1.00056L141.337 1.00043M141.337 20.3944C146.088 20.3944 150.839 20.3944 159.747 20.3944M141.337 20.3944H159.747M159.747 20.3944C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273M148.463 48.273C139.556 48.273 125.188 48.273 125.188 48.273M148.463 48.273L125.188 48.273M125.188 48.273L124 37.97M124 37.97C124 37.97 141.931 37.97 147.87 37.97M124 37.97L147.87 37.97M147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852M151.433 29.4852C146.682 29.4852 138.962 29.4852 131.836 29.4852M151.433 29.4852H131.836M131.836 29.4852C120.142 29.4852 125.897 1.00043 141.337 1.00043M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057ZM96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM87.7028 29.4852L100.768 13.1217L103.143 29.4852H87.7028ZM89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457ZM108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM116.802 65.8487L109.081 90.697C109.081 90.697 117.989 90.697 122.74 90.697C127.491 90.697 133.307 88.4021 135.211 84.0304C137.586 78.5759 143.525 65.8487 132.242 65.8487C120.959 65.8487 116.802 65.8487 116.802 65.8487ZM179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056ZM161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457ZM230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM228.446 65.242C240.917 65.242 239.984 70.6965 237.948 77.9692C235.911 85.2419 230.821 91.3025 219.538 91.3025C208.255 91.3025 208.255 84.0298 210.037 77.9692C211.818 71.9087 215.975 65.242 228.446 65.242Z" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M15.2524 55.5451L21.191 100.999L24.7541 100.999L18.8156 55.5451L15.2524 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 55.5451L14.065 100.999L15.2527 100.999L9.31417 55.5451L8.12646 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 55.5451L6.93853 100.999L7.53239 100.999L1.59385 55.5451L1 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M15.2524 48.2724L30.0988 1H33.6619L18.8156 48.2724H15.2524Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 48.2724L22.9728 1H24.1605L9.31417 48.2724H8.12646Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 48.2724L15.8463 1H16.4402L1.59385 48.2724H1Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M85.3271 55.5457H67.5116L87 12.7363L44.3513 68.2729H58.6038L43.1636 101L85.3271 55.5457Z" fill="#FDC717" stroke="#FDC717" stroke-width="1.18771" stroke-miterlimit="16"/>
</svg>

After

Width:  |  Height:  |  Size: 5.7 KiB

@@ -0,0 +1,6 @@
<svg width="160" height="93" viewBox="0 0 160 93" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M28.8511 91.66L57.6319 1.86368H64.5394L35.7585 91.66H28.8511Z" fill="#356CFF" stroke="#356CFF" stroke-width="2.30244"/>
<path d="M15.0376 91.66L43.8185 1.86368H46.1209L17.3401 91.66H15.0376Z" fill="#356CFF" stroke="#356CFF" stroke-width="2.30244"/>
<path d="M1.22217 91.66L30.003 1.86366H31.1543L2.3734 91.66H1.22217Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.15122"/>
<path d="M71.4465 1.86483L42.666 91.6599H69.144L78.3538 58.2746H123.251L129.007 39.855H84.1099L89.866 22.5868H152.032L157.788 1.86483H71.4465Z" fill="#356CFF" stroke="#356CFF" stroke-width="2.30244"/>
</svg>

After

Width:  |  Height:  |  Size: 691 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 32 KiB

+18
View File
@@ -0,0 +1,18 @@
<svg width="251" height="102" viewBox="0 0 251 102" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M88.8911 54.0412H100.768L87.1095 98.5802H74.0447L88.8911 54.0412Z" fill="#0C0D6F"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M95.4234 0.594433H112.051L117.99 46.915H104.331L103.144 38.601H79.3894L66.9185 54.0412H84.7341L42.5705 98.5802H27.7242L21.7856 54.0412H33.0689L37.8197 89.0786L58.0107 66.5121H43.7582L95.4234 0.594433ZM100.174 12.4715L87.1095 28.5055H102.55L100.174 12.4715Z" fill="#0C0D6F"/>
<path d="M36.632 0.594433L21.7856 46.915H35.4443L40.1951 29.6932H63.3554L66.3246 20.1916H43.1644L46.1336 11.2838H78.2017L81.171 0.594433H36.632Z" fill="#0C0D6F"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M107.894 54.0412L93.6419 98.5802C93.6419 98.5802 104.925 98.5802 120.365 98.5802C135.805 98.5802 143.485 90.7544 147.683 77.7954C151.839 66.2601 152.433 54.0412 135.805 54.0412C119.178 54.0412 107.894 54.0412 107.894 54.0412ZM108.488 88.4847L116.208 64.1367C116.208 64.1367 120.365 64.1367 131.648 64.1367C142.932 64.1367 136.993 76.6077 134.618 81.9523C132.714 86.236 126.898 88.4847 122.147 88.4847C117.396 88.4847 108.488 88.4847 108.488 88.4847Z" fill="#0C0D6F"/>
<path d="M172.031 0.593872L168.467 11.2838C168.467 11.2838 148.605 11.2834 141.744 11.2837C134.883 11.284 135.805 19.5977 140.556 19.5977H158.966C167.874 19.5977 166.18 32.6255 163.123 37.4133C159.686 42.7941 156.59 46.915 147.683 46.915H122.741L121.553 36.8195H147.089C153.027 36.8195 155.403 28.5055 150.652 28.5055H131.055C119.361 28.5055 125.116 0.594307 140.556 0.594307L172.031 0.593872Z" fill="#0C0D6F"/>
<path d="M177.375 0.593872L173.812 11.2838H190.44L179.157 46.915H191.628L202.911 11.2838L218.351 11.2837L222.508 0.594307L177.375 0.593872Z" fill="#0C0D6F"/>
<path d="M160.747 54.0412H201.723L198.754 64.1367H169.061L166.686 71.8563H193.409L191.034 80.7641H163.717L161.341 88.4847H191.034L188.658 98.5802H145.901L160.747 54.0412Z" fill="#0C0D6F"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M230.228 53.4468C254.576 53.4468 251.183 65.3239 248.638 76.0132C246.092 86.7026 240.324 98.5802 217.163 98.5802C194.003 98.5802 195.012 86.7026 198.754 76.0132C202.495 65.3239 205.88 53.4468 230.228 53.4468ZM237.354 76.0132C239.391 68.887 240.324 63.5423 227.853 63.5423C215.382 63.5423 211.225 70.0747 209.443 76.0132C207.662 81.9518 207.662 89.078 218.945 89.078C230.228 89.078 235.318 83.1395 237.354 76.0132Z" fill="#0C0D6F"/>
<path d="M172.031 0.593872L168.467 11.2838C168.467 11.2838 148.605 11.2834 141.744 11.2837C134.883 11.284 135.805 19.5977 140.556 19.5977M172.031 0.593872C172.031 0.593872 155.996 0.594307 140.556 0.594307M172.031 0.593872L140.556 0.594307M140.556 19.5977C145.307 19.5977 150.058 19.5977 158.966 19.5977M140.556 19.5977H158.966M158.966 19.5977C167.874 19.5977 166.18 32.6255 163.123 37.4133C159.686 42.7941 156.59 46.915 147.683 46.915M147.683 46.915C138.775 46.915 122.741 46.915 122.741 46.915M147.683 46.915H122.741M122.741 46.915L121.553 36.8195M121.553 36.8195C121.553 36.8195 141.15 36.8195 147.089 36.8195M121.553 36.8195H147.089M147.089 36.8195C153.027 36.8195 155.403 28.5055 150.652 28.5055M150.652 28.5055C145.901 28.5055 138.181 28.5055 131.055 28.5055M150.652 28.5055H131.055M131.055 28.5055C119.361 28.5055 125.116 0.594307 140.556 0.594307M36.632 0.594433L21.7856 46.915H35.4443L40.1951 29.6932H63.3554L66.3246 20.1916H43.1644L46.1336 11.2838H78.2017L81.171 0.594433H36.632ZM95.4234 0.594433H112.051L117.99 46.915H104.331L103.144 38.601H79.3894L66.9185 54.0412H84.7341L42.5705 98.5802H27.7242L21.7856 54.0412H33.0689L37.8197 89.0786L58.0107 66.5121H43.7582L95.4234 0.594433ZM87.1095 28.5055L100.174 12.4715L102.55 28.5055H87.1095ZM88.8911 54.0412H100.768L87.1095 98.5802H74.0447L88.8911 54.0412ZM107.894 54.0412L93.6419 98.5802C93.6419 98.5802 104.925 98.5802 120.365 98.5802C135.805 98.5802 143.485 90.7544 147.683 77.7954C151.839 66.2601 152.433 54.0412 135.805 54.0412C119.178 54.0412 107.894 54.0412 107.894 54.0412ZM116.208 64.1367L108.488 88.4847C108.488 88.4847 117.396 88.4847 122.147 88.4847C126.898 88.4847 132.714 86.236 134.618 81.9523C136.993 76.6077 142.932 64.1367 131.648 64.1367C120.365 64.1367 116.208 64.1367 116.208 64.1367ZM177.375 0.593872L173.812 11.2838H190.44L179.157 46.915H191.628L202.911 11.2838L218.351 11.2837L222.508 0.594307L177.375 0.593872ZM160.747 54.0412H201.723L198.754 64.1367H169.061L166.686 71.8563H193.409L191.034 80.7641H163.717L161.341 88.4847H191.034L188.658 98.5802H145.901L160.747 54.0412ZM230.228 53.4468C254.576 53.4468 251.183 65.3239 248.638 76.0132C246.092 86.7026 240.324 98.5802 217.163 98.5802C194.003 98.5802 195.012 86.7026 198.754 76.0132C202.495 65.3239 205.88 53.4468 230.228 53.4468ZM227.853 63.5423C240.324 63.5423 239.391 68.887 237.354 76.0132C235.318 83.1395 230.228 89.078 218.945 89.078C207.662 89.078 207.662 81.9518 209.443 76.0132C211.225 70.0747 215.382 63.5423 227.853 63.5423Z" stroke="#0C0D6F" stroke-width="1.18771"/>
<path d="M14.6592 54.0407L20.5977 98.5797L24.1608 98.5797L18.2223 54.0406L14.6592 54.0407Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="1.18771"/>
<path d="M7.5332 54.0407L13.4717 98.5797L14.6594 98.5797L8.72091 54.0406L7.5332 54.0407Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="1.18771"/>
<path d="M0.406738 54.0406L6.34527 98.5797L6.93912 98.5797L1.00059 54.0407L0.406738 54.0406Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="0.593853"/>
<path d="M14.6592 46.9144L29.5055 0.593872H33.0686L18.2223 46.9144H14.6592Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="1.18771"/>
<path d="M7.5332 46.9144L22.3795 0.593872H23.5672L8.72091 46.9144H7.5332Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="1.18771"/>
<path d="M0.406738 46.9144L15.2531 0.593872H15.8469L1.00059 46.9144H0.406738Z" fill="#0C0D6F" stroke="#0C0D6F" stroke-width="0.593853"/>
<path d="M84.7339 54.0413H66.9183L100.174 12.4709L43.758 66.5122H58.0105L42.5703 98.5803L84.7339 54.0413Z" fill="#FDC717" stroke="#FDC717" stroke-width="1.18771" stroke-miterlimit="16"/>
</svg>

After

Width:  |  Height:  |  Size: 5.8 KiB

+18
View File
@@ -0,0 +1,18 @@
<svg width="252" height="105" viewBox="0 0 252 105" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM100.768 13.1217L87.7028 29.4852H103.143L100.768 13.1217Z" fill="#356CFF"/>
<path d="M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM109.081 90.697L116.802 65.8487C116.802 65.8487 120.959 65.8487 132.242 65.8487C143.525 65.8487 137.586 78.5759 135.211 84.0304C133.307 88.4021 127.491 90.697 122.74 90.697C117.989 90.697 109.081 90.697 109.081 90.697Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944H159.747C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273L125.188 48.273L124 37.97L147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852H131.836C120.142 29.4852 125.897 1.00043 141.337 1.00043L173.188 1.00056Z" fill="#356CFF"/>
<path d="M179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056Z" fill="#356CFF"/>
<path d="M161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457Z" fill="#356CFF"/>
<path fill-rule="evenodd" clip-rule="evenodd" d="M230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM237.948 77.9692C239.984 70.6965 240.917 65.242 228.446 65.242C215.975 65.242 211.818 71.9087 210.037 77.9692C208.255 84.0298 208.255 91.3025 219.538 91.3025C230.821 91.3025 235.911 85.2419 237.948 77.9692Z" fill="#356CFF"/>
<path d="M173.188 1.00056L168.625 11.9096C168.625 11.9096 149.386 11.9092 142.525 11.9095C135.664 11.9098 136.586 20.3944 141.337 20.3944M173.188 1.00056C173.188 1.00056 156.777 1.00043 141.337 1.00043M173.188 1.00056L141.337 1.00043M141.337 20.3944C146.088 20.3944 150.839 20.3944 159.747 20.3944M141.337 20.3944H159.747M159.747 20.3944C168.654 20.3944 166.961 33.6899 163.904 38.5761C160.467 44.0675 157.371 48.273 148.463 48.273M148.463 48.273C139.556 48.273 125.188 48.273 125.188 48.273M148.463 48.273L125.188 48.273M125.188 48.273L124 37.97M124 37.97C124 37.97 141.931 37.97 147.87 37.97M124 37.97L147.87 37.97M147.87 37.97C153.808 37.97 156.184 29.4852 151.433 29.4852M151.433 29.4852C146.682 29.4852 138.962 29.4852 131.836 29.4852M151.433 29.4852H131.836M131.836 29.4852C120.142 29.4852 125.897 1.00043 141.337 1.00043M37.2252 1.00057L22.3789 48.273H36.0375L40.7884 30.6974L62.6727 30.6974L69.6727 21.0005L43.7576 21.0004L46.7269 11.9096L77.6727 11.9096L86 1.00057L37.2252 1.00057ZM96.0167 1.00057H112.645L118.583 48.273H104.924L103.737 39.7882H79.9827L67.5117 55.5457H85.3273L43.1638 101H28.3174L22.3789 55.5457H33.6621L38.4129 91.3031L58.604 68.2729H44.3515L96.0167 1.00057ZM87.7028 29.4852L100.768 13.1217L103.143 29.4852H87.7028ZM89.4843 55.5457H101.361L87.7028 101H74.638L89.4843 55.5457ZM108.488 55.5457L94.2351 101C94.2351 101 105.518 101 120.959 101C136.399 101 144.078 93.0133 148.276 79.788C152.432 68.0157 153.027 55.5457 136.399 55.5457C119.771 55.5457 108.488 55.5457 108.488 55.5457ZM116.802 65.8487L109.081 90.697C109.081 90.697 117.989 90.697 122.74 90.697C127.491 90.697 133.307 88.4021 135.211 84.0304C137.586 78.5759 143.525 65.8487 132.242 65.8487C120.959 65.8487 116.802 65.8487 116.802 65.8487ZM179.938 1.00056L175.688 11.9096L191.221 11.9096L179.938 48.273H192.409L203.692 11.9096L219.132 11.9095L223.289 1.00043L179.938 1.00056ZM161.341 55.5457H202.845L198.5 65.8487H169.654L167.279 73.7268H188.5L184.749 82.8177H164.31L161.934 90.697H190.251L186.624 101H146.494L161.341 55.5457ZM230.821 54.9391C255.169 54.9391 251.776 67.0602 249.231 77.9692C246.686 88.8783 240.917 101 217.757 101C194.596 101 195.606 88.8783 199.347 77.9692C203.089 67.0602 206.473 54.9391 230.821 54.9391ZM228.446 65.242C240.917 65.242 239.984 70.6965 237.948 77.9692C235.911 85.2419 230.821 91.3025 219.538 91.3025C208.255 91.3025 208.255 84.0298 210.037 77.9692C211.818 71.9087 215.975 65.242 228.446 65.242Z" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M15.2524 55.5451L21.191 100.999L24.7541 100.999L18.8156 55.5451L15.2524 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 55.5451L14.065 100.999L15.2527 100.999L9.31417 55.5451L8.12646 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 55.5451L6.93853 100.999L7.53239 100.999L1.59385 55.5451L1 55.5451Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M15.2524 48.2724L30.0988 1H33.6619L18.8156 48.2724H15.2524Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M8.12646 48.2724L22.9728 1H24.1605L9.31417 48.2724H8.12646Z" fill="#356CFF" stroke="#356CFF" stroke-width="1.18771"/>
<path d="M1 48.2724L15.8463 1H16.4402L1.59385 48.2724H1Z" fill="#356CFF" stroke="#356CFF" stroke-width="0.593853"/>
<path d="M85.3271 55.5457H67.5116L87 12.7363L44.3513 68.2729H58.6038L43.1636 101L85.3271 55.5457Z" fill="#FDC717" stroke="#FDC717" stroke-width="1.18771" stroke-miterlimit="16"/>
</svg>

After

Width:  |  Height:  |  Size: 5.7 KiB

+329
View File
@@ -0,0 +1,329 @@
@import url("https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@400;500;600&family=IBM+Plex+Sans:wght@400;500;600;700&family=JetBrains+Mono:wght@400;600&display=swap");
@import "tailwindcss";
@custom-variant dark (&:is(.dark *));
:root {
color-scheme: light;
--accent-blue: #356cff;
--background: #f5f4f4;
--foreground: #0f172a;
--card: #ffffff;
--card-foreground: #0f172a;
--popover: #ffffff;
--popover-foreground: #0f172a;
--primary: #0f172a;
--primary-foreground: #f8fafc;
--secondary: #f1f5f9;
--secondary-foreground: #1e293b;
--muted: #f1f5f9;
--muted-foreground: #64748b;
--accent: #e2e8f0;
--accent-foreground: #0f172a;
--destructive: #ef4444;
--destructive-foreground: #ffffff;
--border: #e2e8f0;
--input: #cbd5e1;
--ring: #94a3b8;
--radius: 0.5rem;
}
.dark {
color-scheme: dark;
--accent-blue: #356cff;
--background: #0f172a;
--foreground: #e2e8f0;
--card: #0f172a;
--card-foreground: #f1f5f9;
--popover: #0f172a;
--popover-foreground: #f1f5f9;
--primary: #f1f5f9;
--primary-foreground: #0f172a;
--secondary: #1e293b;
--secondary-foreground: #f1f5f9;
--muted: #1e293b;
--muted-foreground: #94a3b8;
--accent: #1e293b;
--accent-foreground: #f1f5f9;
--destructive: #991b1b;
--destructive-foreground: #fecaca;
--border: #334155;
--input: #334155;
--ring: #cbd5e1;
}
@theme inline {
--color-accent-blue: var(--accent-blue);
--color-background: var(--background);
--color-foreground: var(--foreground);
--color-card: var(--card);
--color-card-foreground: var(--card-foreground);
--color-popover: var(--popover);
--color-popover-foreground: var(--popover-foreground);
--color-primary: var(--primary);
--color-primary-foreground: var(--primary-foreground);
--color-secondary: var(--secondary);
--color-secondary-foreground: var(--secondary-foreground);
--color-muted: var(--muted);
--color-muted-foreground: var(--muted-foreground);
--color-accent: var(--accent);
--color-accent-foreground: var(--accent-foreground);
--color-destructive: var(--destructive);
--color-destructive-foreground: var(--destructive-foreground);
--color-border: var(--border);
--color-input: var(--input);
--color-ring: var(--ring);
--radius-sm: calc(var(--radius) - 4px);
--radius-md: calc(var(--radius) - 2px);
--radius-lg: var(--radius);
--radius-xl: calc(var(--radius) + 4px);
}
/* ——— Resets ——— */
*,
*::before,
*::after {
box-sizing: border-box;
}
html,
body,
#app {
min-height: 100%;
}
html {
background: var(--background);
}
body {
margin: 0;
overflow-y: auto;
color: var(--foreground);
font-family: "IBM Plex Sans", ui-sans-serif, system-ui, sans-serif;
background: var(--background);
}
body::after {
content: "";
position: fixed;
inset: 0;
z-index: -1;
opacity: 0;
pointer-events: none;
transition: opacity 300ms ease;
background:
radial-gradient(circle at top, rgba(56, 189, 248, 0.14), transparent 36%), radial-gradient(circle at right top, rgba(129, 140, 248, 0.12), transparent 28%),
linear-gradient(180deg, #020617 0%, #000000 100%);
background-attachment: fixed;
}
.dark body::after {
opacity: 1;
}
a {
color: inherit;
}
@layer base {
button,
input,
textarea,
select {
font: inherit;
}
}
@layer base {
img,
svg,
video,
canvas {
display: block;
max-width: 100%;
}
}
summary {
list-style: none;
}
summary::-webkit-details-marker {
display: none;
}
::selection {
background: rgba(56, 189, 248, 0.25);
}
.dark ::selection {
background: rgba(56, 189, 248, 0.35);
color: #ffffff;
}
/* ——— Base layer ——— */
@layer base {
* {
@apply border-border;
}
body {
@apply bg-background text-foreground;
}
}
/* ——— Utilities ——— */
@layer utilities {
.scrollbar-hidden {
-ms-overflow-style: none;
scrollbar-width: none;
}
.scrollbar-hidden::-webkit-scrollbar {
display: none;
}
.stroke-dash-anim {
animation: stroke-dash-animation 2s linear infinite;
}
}
@keyframes stroke-dash-animation {
from {
stroke-dashoffset: 8;
}
to {
stroke-dashoffset: 0;
}
}
@keyframes shimmer {
0% {
transform: translateX(-100%);
}
100% {
transform: translateX(200%);
}
}
@keyframes slide-in-from-right {
from {
transform: translateX(10px);
opacity: 0;
}
to {
transform: translateX(0);
opacity: 1;
}
}
.animate-slide-in-from-right {
animation: slide-in-from-right 0.2s ease-out;
}
@keyframes stt-pulse {
0%,
100% {
box-shadow: 0 0 0 3px rgba(239, 68, 68, 0.2);
}
50% {
box-shadow: 0 0 0 6px rgba(239, 68, 68, 0.1);
}
}
.animate-stt-pulse {
animation: stt-pulse 1.4s ease-in-out infinite;
}
/* ——— Theme transition ——— */
html.theme-transition,
html.theme-transition *,
html.theme-transition *::before,
html.theme-transition *::after {
transition:
background-color 300ms ease,
color 300ms ease,
border-color 300ms ease,
box-shadow 300ms ease,
fill 300ms ease,
stroke 300ms ease !important;
}
/* —————————————— CUSTOM TAILWIND —————————————— */
@layer components {
.debug {
@apply border border-rose-500;
}
.horizontal {
@apply flex flex-row;
}
.vertical {
@apply flex flex-col;
}
.horizontal.center-v {
@apply items-center;
}
.horizontal.center-h {
@apply justify-center;
}
.horizontal.center {
@apply justify-center items-center;
}
.vertical.center-v {
@apply justify-center;
}
.vertical.center-h {
@apply items-center;
}
.vertical.center {
@apply justify-center items-center;
}
.space-between {
@apply justify-between;
}
}
@@ -0,0 +1,5 @@
import MonitorPage from '@/components/MonitorPage';
export default function ReplicaMonitorPage() {
return <MonitorPage />;
}
+38
View File
@@ -0,0 +1,38 @@
import type { Metadata } from 'next';
import { GeistSans } from 'geist/font/sans';
import { GeistMono } from 'geist/font/mono';
import { Toaster } from '@/components/ui/sonner';
import './globals.css';
export const metadata: Metadata = {
title: 'Dreamverse',
icons: {
icon: '/icon-simple.svg',
shortcut: '/icon-simple.svg',
},
};
export default function RootLayout({
children,
}: {
children: React.ReactNode;
}) {
return (
<html lang="en" suppressHydrationWarning className={`${GeistSans.variable} ${GeistMono.variable}`}>
<head>
<script
id="theme-init-script"
suppressHydrationWarning
dangerouslySetInnerHTML={{
__html: `try{if(localStorage.getItem('theme')==='dark')document.documentElement.classList.add('dark')}catch{}`,
}}
/>
</head>
<body>
<div id="app">{children}</div>
<Toaster />
</body>
</html>
);
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,104 @@
import { render, screen, waitFor } from '@testing-library/react';
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import ReplicaMonitorPage from './internal/f8a3991c/replica-monitor/page';
describe('Monitor route', () => {
beforeEach(() => {
vi.spyOn(console, 'log').mockImplementation(() => {});
vi.spyOn(console, 'warn').mockImplementation(() => {});
vi.spyOn(console, 'error').mockImplementation(() => {});
});
afterEach(() => {
vi.unstubAllGlobals();
vi.restoreAllMocks();
});
it('renders monitor page instead of the regular app shell and polls every 15 seconds', async () => {
const fetchMock = vi.fn()
.mockResolvedValueOnce({
ok: true,
json: async () => ({
replicas: [
{
url: 'http://r1:8009',
healthy: true,
active_sessions: 2,
pending_sessions: 1,
max_available_sessions: 4,
prompt_provider_success_counts: {
cerebras_ifm: 9,
cerebras: 2,
groq: 1,
},
},
],
}),
})
.mockResolvedValueOnce({
ok: true,
json: async () => ({
replicas: [
{
url: 'http://r1:8009',
healthy: true,
active_sessions: 3,
pending_sessions: 0,
max_available_sessions: 4,
prompt_provider_success_counts: {
cerebras_ifm: 10,
cerebras: 2,
groq: 1,
},
},
],
}),
});
vi.stubGlobal('fetch', fetchMock);
const intervalCallbacks: Array<() => void | Promise<void>> = [];
vi.spyOn(globalThis, 'setInterval').mockImplementation(
((callback: TimerHandler) => {
intervalCallbacks.push(callback as () => void | Promise<void>);
return 1 as unknown as ReturnType<typeof setInterval>;
}) as unknown as typeof setInterval,
);
vi.spyOn(globalThis, 'clearInterval').mockImplementation(
(() => {}) as typeof clearInterval,
);
render(<ReplicaMonitorPage />);
expect(
await screen.findByRole('heading', { name: 'Replica Session Monitor' }),
).toBeInTheDocument();
expect(await screen.findByText('http://r1:8009')).toBeInTheDocument();
expect(await screen.findByText('4')).toBeInTheDocument();
expect(
await screen.findByText('IFM 9 / Cerebras 2 / Groq 1'),
).toBeInTheDocument();
expect(
screen.queryByRole('heading', { name: 'Preset-driven video continuation' }),
).not.toBeInTheDocument();
expect(fetchMock).toHaveBeenCalledTimes(1);
await intervalCallbacks[0]?.();
await waitFor(() => {
expect(fetchMock).toHaveBeenCalledTimes(2);
});
});
it('renders error state when monitor fetch fails', async () => {
const fetchMock = vi.fn().mockResolvedValue({
ok: false,
json: async () => ({ detail: 'boom' }),
});
vi.stubGlobal('fetch', fetchMock);
render(<ReplicaMonitorPage />);
expect(
await screen.findByText('boom'),
).toBeInTheDocument();
});
});
@@ -0,0 +1,179 @@
import { render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
const projectStorageMockState = vi.hoisted(() => ({
saveProject: vi.fn(),
saveProjectMetadata: vi.fn(),
listProjects: vi.fn(async () => []),
loadProjectClips: vi.fn(async () => []),
deleteProject: vi.fn(async () => {}),
pruneOldProjects: vi.fn(async () => {}),
reset() {
this.saveProject.mockReset();
this.saveProjectMetadata.mockReset();
this.listProjects.mockClear();
this.loadProjectClips.mockClear();
this.deleteProject.mockClear();
this.pruneOldProjects.mockClear();
},
}));
vi.mock('../lib/storyPresetsData', () => ({
default: [
{
id: 'test_preset',
label: 'Test Preset',
segment_prompts: ['segment one', 'segment two'],
},
],
}));
vi.mock('../lib/projectStorage', () => ({
saveProject: projectStorageMockState.saveProject,
saveProjectMetadata: projectStorageMockState.saveProjectMetadata,
listProjects: projectStorageMockState.listProjects,
loadProjectClips: projectStorageMockState.loadProjectClips,
deleteProject: projectStorageMockState.deleteProject,
pruneOldProjects: projectStorageMockState.pruneOldProjects,
}));
vi.mock('../lib/media/avPipeline', () => ({
DEFAULT_AV_MIME: 'video/mp4',
createAvPipeline: vi.fn(() => ({
reset() {},
enqueueChunk() {},
ensurePipeline: async () => {},
maybeStartPlayback() {},
tryEndStream() {},
setStreamCompleted() {},
noteSegmentInit() {},
noteSegmentComplete() {},
markStreamStarting() {},
hasArchivedChunks() {
return false;
},
hasArchivedCompletedSegments() {
return false;
},
buildArchivedStreamChunks() {
return [];
},
buildArchivedSegmentSnapshots() {
return [];
},
buildArchivedStreamBlob() {
return new Blob([], { type: 'video/mp4' });
},
takeArchivedStreamChunks() {
return [];
},
takeArchivedSegmentSnapshots() {
return [];
},
takeArchivedStreamBlob() {
return new Blob([], { type: 'video/mp4' });
},
usesNativePlaybackFallback() {
return false;
},
})),
}));
import Page from './page';
describe('Page startup readiness UX', () => {
let fetchMock: ReturnType<typeof vi.fn>;
beforeEach(() => {
projectStorageMockState.reset();
window.history.pushState({}, '', '/');
fetchMock = vi.fn();
vi.stubGlobal('fetch', fetchMock);
vi.spyOn(console, 'log').mockImplementation(() => {});
vi.spyOn(console, 'warn').mockImplementation(() => {});
vi.spyOn(console, 'error').mockImplementation(() => {});
});
afterEach(() => {
vi.unstubAllGlobals();
vi.restoreAllMocks();
});
it('shows a clear notice when the backend is not reachable before session start', async () => {
fetchMock.mockRejectedValue(new Error('connect ECONNREFUSED'));
const user = userEvent.setup();
render(<Page />);
const promptInput = screen.getByRole('textbox', { name: 'Continuation prompt' });
await user.type(promptInput, 'A fox surfing through neon rain');
await user.click(screen.getByRole('button', { name: 'Generate' }));
expect(
await screen.findByText(
'Dreamverse backend is not reachable. Start uv run dreamverse-server and wait for /readyz to return 200 before retrying.',
),
).toBeInTheDocument();
expect(promptInput).toHaveValue('A fox surfing through neon rain');
});
it('shows a readiness notice when GPU workers are not ready yet', async () => {
fetchMock.mockImplementation(async (input: RequestInfo | URL) => {
const url = typeof input === 'string'
? input
: input instanceof URL
? input.toString()
: input.url;
if (url.endsWith('/healthz')) {
return {
ok: true,
status: 200,
json: async () => ({ status: 'ok', service: 'ltx2-streaming-backend' }),
};
}
if (url.endsWith('/readyz')) {
return {
ok: false,
status: 503,
json: async () => ({
detail: 'No ready GPU worker processes.',
}),
};
}
if (url.endsWith('/status')) {
return {
ok: true,
status: 200,
json: async () => ({
total_gpus: 1,
available_gpus: 0,
queue_size: 0,
warmup_enabled: true,
warmup_successful_gpus: 0,
warmup_failed_gpus: 0,
}),
};
}
throw new Error(`Unhandled fetch request in test: ${url}`);
});
const user = userEvent.setup();
render(<Page />);
const promptInput = screen.getByRole('textbox', { name: 'Continuation prompt' });
await user.type(promptInput, 'A fox surfing through neon rain');
await user.click(screen.getByRole('button', { name: 'Generate' }));
expect(
await screen.findByText(
'Dreamverse backend is running, but GPU workers are not ready yet. Wait for startup warmup to finish and retry.',
),
).toBeInTheDocument();
expect(promptInput).toHaveValue('A fox surfing through neon rain');
});
});
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,427 @@
"use client";
import React, { useRef, useState, useCallback, useEffect } from "react";
import Image from "next/image";
import { Film, ArrowUp, X, Loader2, ArrowLeft } from "lucide-react";
import { Button } from "@/components/ui/button";
import LeaveSessionModal, { shouldShowLeaveWarning } from "@/components/LeaveSessionModal";
import SpeechToTextButton from "@/components/SpeechToTextButton";
import { cn } from "@/lib/utils";
const PROMPT_MAX_LENGTH = 500;
interface Props {
sessionStarted?: boolean;
rewritingSeedPrompts?: boolean;
isGenerating?: boolean;
storyPresets?: any[];
continuationDraft?: string;
canJoinSession?: boolean;
canSubmitContinuation?: boolean;
sessionExpired?: boolean;
sessionNotice?: string;
projectResetPending?: boolean;
viewingReadOnly?: boolean;
onPresetGenerate?: (presetId: string) => void;
onContinuationInput?: (e: React.ChangeEvent<HTMLTextAreaElement>) => void;
onContinuationKeydown?: (e: React.KeyboardEvent<HTMLTextAreaElement>) => void;
onGenerate?: () => void;
onSubmitContinuation?: () => void;
onLeave?: () => void;
onStartNewProject?: () => void;
onBackFromViewing?: () => void;
onSpeechTranscript?: (text: string) => void;
onSpeechInterimChange?: (text: string) => void;
}
export default function ChatBar({
sessionStarted = false,
rewritingSeedPrompts = false,
isGenerating = false,
storyPresets = [],
continuationDraft = "",
canJoinSession = false,
canSubmitContinuation = false,
sessionExpired = false,
sessionNotice = "",
projectResetPending = false,
viewingReadOnly = false,
onPresetGenerate = () => {},
onContinuationInput = () => {},
onContinuationKeydown = () => {},
onGenerate = () => {},
onSubmitContinuation = () => {},
onLeave = () => {},
onStartNewProject = () => {},
onBackFromViewing = () => {},
onSpeechTranscript,
onSpeechInterimChange,
}: Props) {
const [sttBusy, setSttBusy] = useState(false);
const [leaveModalOpen, setLeaveModalOpen] = useState(false);
const showSpinner = isGenerating || rewritingSeedPrompts;
const isBusy = isGenerating || rewritingSeedPrompts || projectResetPending;
const messagePlaceholder = projectResetPending
? "Starting new project\u2026"
: isBusy
? "Generating video\u2026"
: !sessionStarted
? "What video are you imagining?"
: "What do you want to edit?";
const actionLabel = !sessionStarted ? "Generate" : "Rewrite rollout";
const inputRef = useRef<HTMLTextAreaElement>(null);
const scrollRef = useRef<HTMLDivElement>(null);
const [canScrollLeft, setCanScrollLeft] = useState(false);
const [canScrollRight, setCanScrollRight] = useState(false);
const [presetRailDragging, setPresetRailDragging] = useState(false);
const presetDragStateRef = useRef({
pointerId: null as number | null,
startX: 0,
startScrollLeft: 0,
moved: false,
});
const suppressPresetClickRef = useRef(false);
const updateScrollState = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setCanScrollLeft(el.scrollLeft > 2);
setCanScrollRight(el.scrollLeft + el.clientWidth < el.scrollWidth - 2);
}, []);
const handlePresetWheel = useCallback(
(event: React.WheelEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
const dominantDelta = Math.abs(event.deltaX) > Math.abs(event.deltaY)
? event.deltaX
: event.deltaY;
if (!dominantDelta) return;
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
const nextScrollLeft = Math.min(
Math.max(el.scrollLeft + dominantDelta, 0),
maxScrollLeft,
);
if (nextScrollLeft === el.scrollLeft) return;
event.preventDefault();
el.scrollLeft = nextScrollLeft;
updateScrollState();
},
[updateScrollState],
);
const finishPresetDrag = useCallback(() => {
presetDragStateRef.current = {
pointerId: null,
startX: 0,
startScrollLeft: 0,
moved: false,
};
setPresetRailDragging(false);
}, []);
const handlePresetPointerDown = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el) return;
if (event.pointerType !== "mouse" || event.button !== 0) return;
if (el.scrollWidth <= el.clientWidth + 1) return;
suppressPresetClickRef.current = false;
presetDragStateRef.current = {
pointerId: event.pointerId,
startX: event.clientX,
startScrollLeft: el.scrollLeft,
moved: false,
};
},
[],
);
const handlePresetPointerMove = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
const dragState = presetDragStateRef.current;
if (!el || dragState.pointerId !== event.pointerId) return;
const deltaX = event.clientX - dragState.startX;
if (!dragState.moved && Math.abs(deltaX) > 4) {
dragState.moved = true;
suppressPresetClickRef.current = true;
setPresetRailDragging(true);
el.setPointerCapture?.(event.pointerId);
}
if (!dragState.moved) return;
event.preventDefault();
const maxScrollLeft = Math.max(el.scrollWidth - el.clientWidth, 0);
el.scrollLeft = Math.min(
Math.max(dragState.startScrollLeft - deltaX, 0),
maxScrollLeft,
);
updateScrollState();
},
[updateScrollState],
);
const handlePresetPointerUp = useCallback(
(event: React.PointerEvent<HTMLDivElement>) => {
const el = scrollRef.current;
if (!el || presetDragStateRef.current.pointerId !== event.pointerId) return;
if (el.hasPointerCapture?.(event.pointerId)) {
el.releasePointerCapture(event.pointerId);
}
finishPresetDrag();
},
[finishPresetDrag],
);
const handlePresetClickCapture = useCallback(
(event: React.MouseEvent<HTMLDivElement>) => {
if (!suppressPresetClickRef.current) return;
suppressPresetClickRef.current = false;
event.preventDefault();
event.stopPropagation();
},
[],
);
useEffect(() => {
updateScrollState();
}, [storyPresets, updateScrollState]);
useEffect(() => {
if (!isBusy && !sttBusy && !window.matchMedia("(pointer: coarse)").matches) {
inputRef.current?.focus();
}
}, [isBusy, sttBusy, sessionStarted]);
const autoResize = useCallback(() => {
const el = inputRef.current;
if (!el) return;
el.style.height = "auto";
const lineHeight = parseFloat(getComputedStyle(el).lineHeight) || 20;
const maxHeight = lineHeight * 3;
el.style.height = `${Math.min(el.scrollHeight, maxHeight)}px`;
el.style.overflowY = el.scrollHeight > maxHeight ? "auto" : "hidden";
}, []);
useEffect(() => {
autoResize();
}, [continuationDraft, autoResize]);
const handleKeyDown = useCallback(
(e: React.KeyboardEvent<HTMLTextAreaElement>) => {
if (e.key === "Enter" && !e.nativeEvent.isComposing && !e.shiftKey) {
e.preventDefault();
if (!sessionStarted) {
if (canJoinSession && !isGenerating && continuationDraft.trim()) {
onGenerate();
}
} else {
onContinuationKeydown(e);
}
return;
}
onContinuationKeydown(e);
},
[onContinuationKeydown, sessionStarted, canJoinSession, isGenerating, continuationDraft, onGenerate],
);
if (viewingReadOnly) {
return (
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-6 py-4 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">View-only project</p>
<p className="max-w-md text-xs text-muted-foreground">Project sessions are currently limited to 5 minutes. Start a new project to create more videos.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onBackFromViewing} variant="outline" size="sm" className="gap-1.5 rounded-full px-4">
<ArrowLeft className="size-3.5" />
Back
</Button>
<Button onClick={onStartNewProject} size="sm" className="rounded-full px-5">
New project
</Button>
</div>
</div>
</section>
);
}
if (sessionExpired) {
return (
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
<div className="flex flex-col items-center gap-3 rounded-2xl border border-border bg-card/80 px-8 py-5 text-center shadow-md backdrop-blur-sm">
<div className="flex flex-col gap-1">
<p className="text-sm font-semibold text-foreground">Session ended</p>
<p className="max-w-xs text-xs text-muted-foreground">Each project currently has a 5-minute session. Start a new project to continue creating videos.</p>
</div>
<div className="mt-1 flex items-center gap-2">
<Button onClick={onStartNewProject} size="sm" className="rounded-full px-5">
New Project
</Button>
<a href="https://docs.google.com/forms/d/e/1FAIpQLSe5zpO1iD8Ds-Ih-fOLm64qd7YZVvuvAyHuJaAfw1hkRHTe_A/viewform?usp=publish-editor" target="_blank" rel="noopener noreferrer">
<Button variant="outline" size="sm" className="rounded-full px-5">
Join Waitlist
</Button>
</a>
</div>
</div>
</section>
);
}
return (
<section className="mx-auto flex w-full max-w-2xl shrink-0 flex-col gap-4">
{storyPresets.length > 0 && !sessionStarted && (
<div className={cn("relative transition-opacity duration-200", isGenerating && "pointer-events-none opacity-40")}>
<div
ref={scrollRef}
onScroll={updateScrollState}
onWheel={handlePresetWheel}
onPointerDown={handlePresetPointerDown}
onPointerMove={handlePresetPointerMove}
onPointerUp={handlePresetPointerUp}
onPointerCancel={handlePresetPointerUp}
onLostPointerCapture={finishPresetDrag}
onClickCapture={handlePresetClickCapture}
className={cn(
"scrollbar-hidden flex gap-3 overflow-x-auto px-1 select-none",
presetRailDragging ? "cursor-grabbing" : "cursor-grab",
)}
>
{storyPresets.map((preset) => (
<button
key={preset.id}
type="button"
disabled={isGenerating}
onClick={() => onPresetGenerate(preset.id)}
className="flex flex-col sm:flex-row items-start gap-1.5 shrink-0 rounded-xl border p-2.5 text-left backdrop-blur-sm transition-colors max-w-42 sm:max-w-[215px] border-input bg-card/80 text-muted-foreground hover:bg-slate-200/60 hover:border-slate-400 hover:text-slate-700 dark:bg-slate-800/80 dark:text-slate-300 dark:hover:bg-slate-700/50 dark:hover:border-slate-500 dark:hover:text-slate-200"
>
<Film className="mt-0.5 size-4 shrink-0 opacity-60" />
<span className="flex flex-col gap-1 min-w-0">
<span className="text-[14px] font-medium line-clamp-1">{preset.label}</span>
{preset.description && <span className="text-xs leading-tight opacity-70 line-clamp-3 sm:line-clamp-2">{preset.description}</span>}
</span>
</button>
))}
</div>
<div
className={cn("pointer-events-none absolute inset-y-0 left-0 w-8 bg-background transition-opacity duration-150", canScrollLeft ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to right, black, transparent)", WebkitMaskImage: "linear-gradient(to right, black, transparent)" }}
aria-hidden="true"
/>
<div
className={cn("pointer-events-none absolute inset-y-0 right-0 w-8 bg-background transition-opacity duration-150", canScrollRight ? "opacity-100" : "opacity-0")}
style={{ maskImage: "linear-gradient(to left, black, transparent)", WebkitMaskImage: "linear-gradient(to left, black, transparent)" }}
aria-hidden="true"
/>
</div>
)}
{sessionNotice && (
<div
className={cn(
"rounded-xl px-4 py-2.5 text-center text-xs",
sessionStarted
? "border border-amber-500/20 bg-amber-500/10 text-amber-700 dark:text-amber-400"
: "border border-rose-500/20 bg-rose-500/10 text-rose-700 dark:text-rose-300",
)}
>
{sessionNotice}
</div>
)}
{projectResetPending && sessionStarted && (
<div className="rounded-xl border border-sky-500/20 bg-sky-500/10 px-4 py-2.5 text-center text-xs text-sky-700 dark:text-sky-300">
Starting a new project after the current shot finishes. Your GPU session stays active.
</div>
)}
<div
className={cn(
"flex min-w-0 items-center gap-1.5 rounded-4xl border py-2.5 pl-5 pr-2.5 shadow-md backdrop-blur-sm transition-all duration-200",
isBusy ? "border-input/60 bg-card/40" : "border-input bg-card/65",
)}
>
<textarea
ref={inputRef}
id="continuation-prompt"
aria-label="Continuation prompt"
value={continuationDraft}
onChange={onContinuationInput}
onKeyDown={handleKeyDown}
placeholder={sttBusy ? "Listening\u2026" : messagePlaceholder}
maxLength={PROMPT_MAX_LENGTH}
disabled={isBusy || sttBusy}
rows={1}
className={cn(
"min-w-0 flex-1 resize-none bg-transparent text-foreground outline-none placeholder:text-muted-foreground transition-opacity duration-200 scrollbar-thin leading-snug",
(isBusy || sttBusy) && "cursor-not-allowed opacity-50",
)}
/>
{onSpeechTranscript && <SpeechToTextButton disabled={isBusy} onTranscript={onSpeechTranscript} onInterimChange={onSpeechInterimChange} onBusyChange={setSttBusy} />}
{!sessionStarted ? (
<Button
aria-label={actionLabel}
title={actionLabel}
onClick={onGenerate}
disabled={!canJoinSession || isGenerating || !continuationDraft.trim()}
size="icon-sm"
className="shrink-0 rounded-full"
>
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
</Button>
) : (
<>
<Button
aria-label={actionLabel}
title={actionLabel}
onClick={onSubmitContinuation}
disabled={!canSubmitContinuation || showSpinner || projectResetPending || !continuationDraft.trim()}
size="icon-sm"
className="shrink-0 rounded-full"
>
{showSpinner ? <Loader2 className="size-5 animate-spin" /> : <ArrowUp className="size-5" />}
</Button>
<Button variant="outline" aria-label="Leave" title="Leave" onClick={() => { if (shouldShowLeaveWarning()) setLeaveModalOpen(true); else onLeave(); }} disabled={isGenerating || projectResetPending} size="icon-sm" className="shrink-0 rounded-full">
<X className="size-5" />
</Button>
</>
)}
</div>
<p className="px-2 text-center text-[11px] text-muted-foreground">
LLM powered by{" "}
<a
href="https://ifm.ai/k2/"
target="_blank"
rel="noopener noreferrer"
className="inline-flex items-center gap-1 font-medium text-foreground/80 transition-colors hover:text-foreground"
>
<span>K2-V2</span>
<Image
src="/k2.png"
alt=""
aria-hidden="true"
width={14}
height={14}
className="h-3.5 w-auto opacity-80"
/>
</a>
</p>
<LeaveSessionModal
open={leaveModalOpen}
onClose={() => setLeaveModalOpen(false)}
onConfirmLeave={() => { setLeaveModalOpen(false); onLeave(); }}
/>
</section>
);
}
@@ -0,0 +1,65 @@
"use client";
import React from "react";
import { SidePanelOpenFilled } from "@carbon/icons-react";
import { ExternalLink } from "lucide-react";
import Image from "next/image";
import { Badge } from "@/components/ui/badge";
import { Button } from "@/components/ui/button";
import { ThemeToggle } from "@/components/ui/theme-toggle";
const FASTVIDEO_REPO_URL = "https://haoailab.com/blogs/dreamverse/";
const FASTVIDEO_BLOG_URL = "https://haoailab.com/blogs/dreamverse/";
interface Props {
timeLeft?: number | null;
formatTime?: (seconds: number) => string;
onToggleSidebar?: () => void;
}
export default function Header({ timeLeft = null, formatTime = (seconds) => `${seconds}`, onToggleSidebar }: Props) {
const timeVariant = timeLeft !== null && timeLeft <= 30 ? "warning" : "secondary";
return (
<header className="relative z-30 shrink-0">
<div className="flex flex-wrap items-center justify-between gap-y-2 px-4 pt-3 pb-2 sm:pt-4 sm:pb-3 sm:px-6">
<div className="flex items-center gap-3">
{onToggleSidebar && (
<Button variant="outline" size="icon" onClick={onToggleSidebar} aria-label="Toggle sidebar">
<SidePanelOpenFilled size={20} />
</Button>
)}
<a href={FASTVIDEO_REPO_URL} target="_blank" rel="noopener noreferrer" title="FastVideo on GitHub">
<Image src="/logo.svg" alt="FastVideo" width={32} height={32} className="h-8 w-auto sm:h-9 transition-opacity hover:opacity-70" />
</a>
<div className="hidden sm:flex items-center gap-3">
<a href="https://docs.google.com/forms/d/e/1FAIpQLSe5zpO1iD8Ds-Ih-fOLm64qd7YZVvuvAyHuJaAfw1hkRHTe_A/viewform?usp=publish-editor" target="_blank" rel="noopener noreferrer">
<Button variant="outline" size="sm" className="gap-1.5 rounded-full px-3 text-xs">
Join Waitlist
<ExternalLink className="size-3 opacity-60" />
</Button>
</a>
</div>
</div>
<div className="flex items-center gap-3">
{timeLeft !== null && (
<Badge variant={timeVariant} className="rounded-xl px-3 py-1 text-xs font-medium normal-case tracking-normal">
Time left: {formatTime(timeLeft)}
</Badge>
)}
<ThemeToggle />
</div>
</div>
<div className="flex sm:hidden items-center gap-2 px-4 pb-3">
<a href="https://docs.google.com/forms/d/e/1FAIpQLSe5zpO1iD8Ds-Ih-fOLm64qd7YZVvuvAyHuJaAfw1hkRHTe_A/viewform?usp=publish-editor" target="_blank" rel="noopener noreferrer">
<Button variant="outline" size="sm" className="gap-1.5 rounded-full px-3 text-xs">
Join Waitlist
<ExternalLink className="size-3 opacity-60" />
</Button>
</a>
</div>
</header>
);
}
@@ -0,0 +1,85 @@
"use client";
import { useState, useEffect } from "react";
import { Button } from "@/components/ui/button";
import { Checkbox } from "@/components/ui/checkbox";
import { Label } from "@/components/ui/label";
const STORAGE_KEY = "fastvideo-suppress-leave-warning";
interface LeaveSessionModalProps {
open?: boolean;
onClose?: () => void;
onConfirmLeave?: () => void;
}
export default function LeaveSessionModal({
open = false,
onClose = () => {},
onConfirmLeave = () => {},
}: LeaveSessionModalProps) {
const [suppress, setSuppress] = useState(false);
useEffect(() => {
if (open) setSuppress(false);
}, [open]);
if (!open) return null;
function handleConfirm() {
if (suppress) {
try {
localStorage.setItem(STORAGE_KEY, "1");
} catch {}
}
onConfirmLeave();
}
return (
<div className="fixed inset-0 z-[70] flex items-center justify-center bg-black/45 px-4 backdrop-blur-[3px]">
<div
role="dialog"
aria-modal="true"
aria-labelledby="leave-session-title"
className="w-full max-w-md rounded-3xl border border-border/70 bg-card/95 p-6 shadow-2xl"
>
<div className="space-y-3">
<div className="space-y-1">
<h2 id="leave-session-title" className="text-lg font-semibold text-foreground">
Leave session?
</h2>
<p className="text-sm text-muted-foreground">
This will end your current session. You will need to queue again for a GPU to start a new project.
</p>
</div>
<div className="flex items-center gap-2">
<Checkbox
id="suppress-leave-warning"
checked={suppress}
onCheckedChange={(v) => setSuppress(v === true)}
/>
<Label htmlFor="suppress-leave-warning" className="text-sm text-muted-foreground cursor-pointer">
Do not warn again
</Label>
</div>
</div>
<div className="mt-5 flex flex-col-reverse gap-2 sm:flex-row sm:justify-end">
<Button variant="outline" onClick={onClose}>
Cancel
</Button>
<Button variant="destructive" onClick={handleConfirm}>
Leave session
</Button>
</div>
</div>
</div>
);
}
export function shouldShowLeaveWarning(): boolean {
try {
return localStorage.getItem(STORAGE_KEY) !== "1";
} catch {
return true;
}
}
@@ -0,0 +1,174 @@
'use client';
import React, { useState, useEffect, useRef } from 'react';
import { Badge } from '@/components/ui/badge';
import { Card, CardContent, CardHeader, CardTitle } from '@/components/ui/card';
const POLL_INTERVAL_MS = 15000;
interface Replica {
url: string;
healthy: boolean;
active_sessions?: number;
pending_sessions?: number;
max_available_sessions?: number;
prompt_provider_success_counts?: Record<string, number>;
}
function formatTimestamp(value: string): string {
if (!value) return '';
try {
return new Date(value).toLocaleString();
} catch {
return '';
}
}
function formatProviderSuccessCounts(
counts: Record<string, number> | undefined,
): string {
const normalized = counts || {};
const cerebrasIfm = normalized.cerebras_ifm ?? 0;
const cerebras = normalized.cerebras ?? 0;
const groq = normalized.groq ?? 0;
return `IFM ${cerebrasIfm} / Cerebras ${cerebras} / Groq ${groq}`;
}
export default function MonitorPage() {
const [loading, setLoading] = useState(true);
const [error, setError] = useState('');
const [replicas, setReplicas] = useState<Replica[]>([]);
const [lastUpdated, setLastUpdated] = useState('');
const timerRef = useRef<ReturnType<typeof setInterval> | null>(null);
useEffect(() => {
async function loadReplicaSessions() {
try {
const response = await fetch('/router/replicas/sessions');
const payload = await response.json();
if (!response.ok) {
throw new Error(
payload.detail || 'Failed to fetch replica sessions',
);
}
setReplicas(
Array.isArray(payload.replicas) ? payload.replicas : [],
);
setError('');
setLastUpdated(new Date().toISOString());
} catch (err: any) {
setError(err?.message || String(err));
} finally {
setLoading(false);
}
}
loadReplicaSessions();
timerRef.current = setInterval(loadReplicaSessions, POLL_INTERVAL_MS);
return () => {
if (timerRef.current) {
clearInterval(timerRef.current);
timerRef.current = null;
}
};
}, []);
return (
<main className="mx-auto flex min-h-screen w-full max-w-6xl flex-col gap-4 px-4 py-8 text-foreground">
<div className="space-y-2">
<h1 className="text-3xl font-semibold text-foreground">
Replica Session Monitor
</h1>
<div className="flex flex-wrap gap-2 text-sm text-muted-foreground">
<Badge variant="secondary">poll interval: 15s</Badge>
{lastUpdated && (
<Badge variant="outline">
last updated: {formatTimestamp(lastUpdated)}
</Badge>
)}
</div>
</div>
{loading ? (
<Card>
<CardContent className="p-6 text-sm text-muted-foreground">
Loading monitor data...
</CardContent>
</Card>
) : error ? (
<Card className="border-rose-500/30 bg-rose-950/45">
<CardContent className="p-6 text-sm text-rose-100">
{error}
</CardContent>
</Card>
) : (
<Card className="overflow-hidden">
<CardHeader className="border-b border-border pb-4">
<CardTitle className="text-xl">Replica capacity</CardTitle>
</CardHeader>
<CardContent className="p-0">
<div className="overflow-x-auto">
<table className="w-full min-w-[760px] border-collapse text-left text-sm">
<thead className="bg-secondary text-muted-foreground">
<tr>
<th className="border-b border-border px-4 py-3 font-semibold">
URL
</th>
<th className="border-b border-border px-4 py-3 font-semibold">
Healthy
</th>
<th className="border-b border-border px-4 py-3 font-semibold">
Active WS Sessions
</th>
<th className="border-b border-border px-4 py-3 font-semibold">
Pending Sessions
</th>
<th className="border-b border-border px-4 py-3 font-semibold">
Max Available Sessions
</th>
<th className="border-b border-border px-4 py-3 font-semibold">
Prompt API Successes
</th>
</tr>
</thead>
<tbody>
{replicas.map((replica, index) => (
<tr
key={replica.url || index}
className="border-b border-border last:border-b-0"
>
<td className="px-4 py-3 text-foreground">{replica.url}</td>
<td className="px-4 py-3">
<Badge
variant={replica.healthy ? 'success' : 'destructive'}
>
{replica.healthy ? 'yes' : 'no'}
</Badge>
</td>
<td className="px-4 py-3 text-foreground">
{replica.active_sessions ?? '-'}
</td>
<td className="px-4 py-3 text-foreground">
{replica.pending_sessions ?? '-'}
</td>
<td className="px-4 py-3 text-foreground">
{replica.max_available_sessions ?? '-'}
</td>
<td className="px-4 py-3 text-foreground">
{formatProviderSuccessCounts(
replica.prompt_provider_success_counts,
)}
</td>
</tr>
))}
</tbody>
</table>
</div>
</CardContent>
</Card>
)}
</main>
);
}
@@ -0,0 +1,70 @@
"use client";
import { Button } from "@/components/ui/button";
interface SessionTimeoutModalProps {
open?: boolean;
onClose?: () => void;
onStartNewProject?: () => void;
repoUrl?: string;
blogUrl?: string;
}
export default function SessionTimeoutModal({
open = false,
onClose = () => {},
onStartNewProject = () => {},
repoUrl = "",
blogUrl = "",
}: SessionTimeoutModalProps) {
if (!open) return null;
return (
<div className="fixed inset-0 z-[70] flex items-center justify-center bg-black/45 px-4 backdrop-blur-[3px]">
<div
role="dialog"
aria-modal="true"
aria-labelledby="session-timeout-title"
className="w-full max-w-md rounded-3xl border border-border/70 bg-card/95 p-6 shadow-2xl"
>
<div className="space-y-3">
<div className="space-y-1">
<h2 id="session-timeout-title" className="text-lg font-semibold text-foreground">
Session ended
</h2>
<p className="text-sm text-muted-foreground">
This project hit the current 5-minute session limit. Your latest video stays on screen, and the project is being kept in the archive so you can come back to it.
</p>
</div>
<p className="text-sm text-muted-foreground">
Start a new project to keep creating, or keep viewing this one while you decide what to do next.
</p>
<div className="flex flex-wrap gap-2 text-sm">
<a
href={repoUrl}
target="_blank"
rel="noreferrer"
className="font-medium text-sky-700 underline underline-offset-4 transition-colors hover:text-sky-600 dark:text-sky-300 dark:hover:text-sky-200"
>
Open repo
</a>
<a
href={blogUrl}
target="_blank"
rel="noreferrer"
className="font-medium text-sky-700 underline underline-offset-4 transition-colors hover:text-sky-600 dark:text-sky-300 dark:hover:text-sky-200"
>
Read blog
</a>
</div>
</div>
<div className="mt-5 flex flex-col-reverse gap-2 sm:flex-row sm:justify-end">
<Button variant="outline" onClick={onClose}>
Keep viewing
</Button>
<Button onClick={onStartNewProject}>New project</Button>
</div>
</div>
</div>
);
}
@@ -0,0 +1,229 @@
"use client";
import React, { useState, useRef, useCallback, useEffect } from "react";
import { SidePanelCloseFilled } from "@carbon/icons-react";
import { Plus, Trash2, Clock, Film } from "lucide-react";
import { Button } from "@/components/ui/button";
import { Badge } from "@/components/ui/badge";
import { cn } from "@/lib/utils";
import type { StoredProject } from "@/lib/projectStorage";
function resolveProjectTitle(project: StoredProject): string {
if (project.originalLabel) return project.originalLabel;
const events = project.promptEvents || [];
for (let i = events.length - 1; i >= 0; i--) {
const e = events[i];
if (String(e?.source || "") === "user_rewrite" && typeof e?.text === "string" && e.text.trim()) {
return e.text.trim();
}
}
return "Untitled project";
}
function formatRelativeTime(timestamp: number): string {
const diff = Date.now() - timestamp;
const seconds = Math.floor(diff / 1000);
if (seconds < 60) return "just now";
const minutes = Math.floor(seconds / 60);
if (minutes < 60) return `${minutes}m ago`;
const hours = Math.floor(minutes / 60);
if (hours < 24) return `${hours}h ago`;
const days = Math.floor(hours / 24);
if (days < 7) return `${days}d ago`;
return new Date(timestamp).toLocaleDateString();
}
interface SidebarProps {
open?: boolean;
currentProjectId?: string;
currentProjectLabel?: string;
sessionActive?: boolean;
sessionExpired?: boolean;
projectResetPending?: boolean;
savedProjects?: StoredProject[];
viewingProjectId?: string | null;
isViewingPastProject?: boolean;
onClose?: () => void;
onSelectProject?: (projectId: string) => void;
onSelectCurrentProject?: () => void;
onDeleteProject?: (projectId: string) => void;
onNewProject?: () => void;
}
export default function Sidebar({
open = false,
currentProjectId = "",
currentProjectLabel = "",
sessionActive = false,
sessionExpired = false,
projectResetPending = false,
savedProjects = [],
viewingProjectId = null,
isViewingPastProject = false,
onClose = () => {},
onSelectProject = () => {},
onSelectCurrentProject = () => {},
onDeleteProject = () => {},
onNewProject = () => {},
}: SidebarProps) {
const hasCurrentProject = sessionActive || sessionExpired;
const previousProjects = currentProjectId ? savedProjects.filter((p) => p.id !== currentProjectId) : savedProjects;
const [pendingDeleteId, setPendingDeleteId] = useState<string | null>(null);
const deleteTimerRef = useRef<ReturnType<typeof setTimeout> | null>(null);
const clearPendingDelete = useCallback(() => {
setPendingDeleteId(null);
if (deleteTimerRef.current) {
clearTimeout(deleteTimerRef.current);
deleteTimerRef.current = null;
}
}, []);
useEffect(() => {
return () => {
if (deleteTimerRef.current) clearTimeout(deleteTimerRef.current);
};
}, []);
useEffect(() => {
if (!open) clearPendingDelete();
}, [open, clearPendingDelete]);
function handleDeleteClick(e: React.MouseEvent, projectId: string) {
e.stopPropagation();
if (pendingDeleteId === projectId) {
clearPendingDelete();
onDeleteProject(projectId);
} else {
setPendingDeleteId(projectId);
if (deleteTimerRef.current) clearTimeout(deleteTimerRef.current);
deleteTimerRef.current = setTimeout(() => setPendingDeleteId(null), 3000);
}
}
return (
<>
<div
className={cn("fixed inset-0 z-40 bg-black/40 backdrop-blur-[2px] transition-opacity duration-200", open ? "opacity-100" : "pointer-events-none opacity-0")}
onClick={onClose}
aria-hidden="true"
/>
<aside
className={cn(
"fixed inset-y-0 left-0 z-50 flex w-[280px] max-w-[calc(100vw-3rem)] flex-col border-r border-border/60 bg-card/95 backdrop-blur-xl transition-transform duration-200 ease-out",
open ? "translate-x-0" : "-translate-x-full",
)}
aria-label="Project history"
>
<div className="flex items-center justify-between px-5 pt-4 pb-2">
<span className="text-lg font-semibold text-foreground">Projects</span>
<Button variant="ghost" size="icon" onClick={onClose} aria-label="Close sidebar">
<SidePanelCloseFilled size={18} />
</Button>
</div>
<div className="px-4 pb-3">
<Button
variant="outline"
size="sm"
className="w-full gap-2 rounded-lg font-medium"
onClick={onNewProject}
disabled={projectResetPending}
>
<Plus className="size-4" />
{projectResetPending ? "Starting new project..." : "New project"}
</Button>
</div>
<nav className="flex-1 overflow-y-auto px-3 pb-4">
{hasCurrentProject && (
<div className="mb-3">
<p className="mb-1.5 px-2 text-[11px] font-medium uppercase tracking-wider text-muted-foreground">Current</p>
<div
className={cn("rounded-xl px-3 py-2.5", isViewingPastProject ? "cursor-pointer bg-accent/40 hover:bg-accent/60 transition-colors" : "bg-accent/80")}
onClick={isViewingPastProject ? onSelectCurrentProject : undefined}
role={isViewingPastProject ? "button" : undefined}
tabIndex={isViewingPastProject ? 0 : undefined}
onKeyDown={isViewingPastProject ? (e) => e.key === "Enter" && onSelectCurrentProject() : undefined}
>
<div className="flex items-center gap-2">
<Film className="size-3.5 shrink-0 text-muted-foreground" />
<span className="min-w-0 flex-1 truncate text-[13px] font-medium text-foreground">{currentProjectLabel || "Untitled project"}</span>
</div>
<div className="mt-1.5 flex items-center gap-1.5">
{sessionExpired ? (
<Badge variant="secondary" className="rounded-md px-1.5 py-0 text-[10px]">
Expired
</Badge>
) : (
<Badge variant="default" className="rounded-md px-1.5 py-0 text-[10px]">
Active
</Badge>
)}
</div>
</div>
</div>
)}
{previousProjects.length > 0 && (
<div>
<p className="mb-1.5 px-2 text-[11px] font-medium uppercase tracking-wider text-muted-foreground">Previous</p>
<div className="flex flex-col gap-1">
{previousProjects.map((project) => {
const isViewing = viewingProjectId === project.id;
return (
<div
key={project.id}
className={cn("group flex items-start gap-2 rounded-xl px-3 py-2.5 transition-colors cursor-pointer", isViewing ? "bg-accent/60" : "hover:bg-accent/40")}
onClick={() => onSelectProject(project.id)}
role="button"
tabIndex={0}
onKeyDown={(e) => e.key === "Enter" && onSelectProject(project.id)}
>
{project.lastThumbnail ? (
<img src={project.lastThumbnail} alt="" className="mt-0.5 h-8 w-auto shrink-0 rounded border border-border object-cover" />
) : (
<Film className="mt-0.5 size-3.5 shrink-0 text-muted-foreground" />
)}
<div className="min-w-0 flex-1">
<p className="truncate text-[13px] font-medium text-foreground">{resolveProjectTitle(project)}</p>
<div className="flex items-center gap-1 text-[11px] text-muted-foreground">
<Clock className="size-3" />
<span>{formatRelativeTime(project.createdAt)}</span>
</div>
</div>
{pendingDeleteId === project.id ? (
<button
type="button"
className="mt-0.5 shrink-0 rounded bg-destructive/15 px-1.5 py-0.5 !text-xs !font-medium text-destructive transition-colors hover:bg-destructive/25"
onClick={(e) => handleDeleteClick(e, project.id)}
aria-label="Confirm delete project"
>
Delete?
</button>
) : (
<button
type="button"
className="mt-0.5 shrink-0 rounded p-1 text-muted-foreground opacity-0 transition-opacity hover:bg-destructive/10 hover:text-destructive group-hover:opacity-100"
onClick={(e) => handleDeleteClick(e, project.id)}
aria-label="Delete project"
>
<Trash2 className="size-3.5" />
</button>
)}
</div>
);
})}
</div>
</div>
)}
{!hasCurrentProject && previousProjects.length === 0 && <p className="px-2 py-4 text-center text-[13px] text-muted-foreground">No projects yet</p>}
</nav>
</aside>
</>
);
}

Some files were not shown because too many files have changed in this diff Show More