Compare commits

...
Author SHA1 Message Date
SolitaryThinker cb46d63f07 [bugfix]: unset GenerationRequest fields inherit model presets
Directly-constructed GenerationRequests marked every schema default as
explicit (normalize_generation_request bound the full serialization when
_fastvideo_explicit_paths was absent), so request_to_sampling_param
overwrote model preset values with schema defaults: FastWan DMD ran
3->50 steps, TurboDiffusion 4->50, stable-audio 100 steps/CFG7 ->
50/CFG1, and 1.3B Wan examples rendered 720x1280x125f with CFG off.

Fix at the root: SamplingConfig fields now default to None = "inherit
the model preset". None leaves are never bound as explicit paths — for
directly-constructed requests (normalize_generation_request) and for
parsed YAML/JSON (parse_config treats explicit null as unset, so
widening the types cannot let nulls stomp presets either). Dicts
emptied by the pruning are dropped too, so None-valued extensions
cannot resurface via the live-attribute read in
explicit_request_updates. The _SCHEMA_DEFAULT_UPDATES tolerance hack is
deleted (an explicitly-set field unsupported by a model's SamplingParam
now always fails loudly), and return_state is skipped in the update
loop since it is already translated to return_continuation_state.

Consumers of raw (pre-resolution) sampling fields are guarded: the
streaming server validates operator-pinned default_request.sampling
before the multi-minute model boot (run_server) and in build_app; its
segment negative_prompt merge preserves an explicit "" clear; the mock
server supplies concrete fallbacks; generate_async resolves the preset
so its advertised total_steps matches the actual run; the ray-serve
gradio demo passes "" to keep its no-negative-prompt default. Stale
schema-default claims in video_api/ServeConfig/openai.md docs updated.

Tests: test_preset_inheritance.py locks the contract (bare request
inherits preset; partial override; explicit value equal to an old
schema default wins; negative_prompt None-inherits/""-clears; parsed
nulls inherit; None-valued extensions are dropped; unsupported explicit
field raises) plus a build_app pin-validation test.
2026-07-11 03:26:33 -07:00
SolitaryThinker bbbb7ab021 [refactor]: migrate examples/scripts to typed VideoGenerator API; rename api/compat.py to api/translation.py
Move all examples/inference and scripts off the legacy
VideoGenerator.from_pretrained(model, **kwargs) +
generate_video(prompt, **kwargs) surface onto the typed
from_config(GeneratorConfig(...)) + generate(GenerationRequest(...))
convention, including the Kandinsky-5 and DreamX-World examples added
on main after the original migration. git-mv api/compat.py to
api/translation.py ('compat' implied a temporary shim; the forward
translation is its honest permanent role) and repoint all importers,
tests, and doc links.
2026-07-11 03:26:24 -07:00
Satyam Srivastava 19a51a1fe6 [ci] Trigger performance benchmarks for performance code changes (#1583) 2026-07-10 20:21:33 -07:00
William Lin d3232cea5a [ci]: gate the full-suite trigger on pre-commit and docs build (#1572) 2026-07-11 02:56:49 +00:00
Raghav KandSolitaryThinker 0c90c8c24d [bugfix] nvfp4: cast fp32 inputs to bf16 instead of asserting (#1488)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-10 21:39:17 +00:00
Mingjia HuoandClaude Fable 5 4c08ffce49 [feat] World model training using third person games (#1443)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 05:09:07 +00:00
Atharv Ramesh af4a77553c [ci]: add SSIM reference bootstrap flow (#1522) (#1547) 2026-07-10 01:49:38 +00:00
alexzmsandSolitaryThinker c096fda1eb [docs] Add LTX-2.3 distilled inference run configs (t2v/i2v × 5+2/8+3 × resolutions) (#1568)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-09 18:53:12 +00:00
William Lin 8f47e85be0 [bugfix]: retry remote image downloads in load_image (#1570) 2026-07-09 07:33:24 -07:00
William Lin 90d3bd19eb [infra] Deliver per-job Buildkite env to Modal CI at runtime, not as image layers (#1569) 2026-07-09 06:33:37 -07:00
Mac LeeandSolitaryThinker afb4f7d3c5 [ci]: emit v2 performance result schema (#1551)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-09 06:11:15 +00:00
02e1143f22 [feat] Add Kandinsky-5 T2V/I2V pipeline support (#1471)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: leffff <levnovitskiy@gmail.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 14:43:30 -07:00
KaredandSolitaryThinker e2f4d1a7b5 [feat]: add SwanLab tracker (#1461)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 08:35:38 +00:00
Kaiqin KongandSolitaryThinker f037351146 [feat] Add Clean-history Teacher Forcing and Causal Consistency Distillation (#1505)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 08:13:11 +00:00
Mac LeeandSolitaryThinker 1ee11e08dc [ci]: add performance fingerprint cohorts (#1546)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 07:24:00 +00:00
595f0ea60e [feat] Add DreamX-World 5B Cam and AR pipelines (#1538)
Co-authored-by: Suckl <Suckl@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 06:16:18 +00:00
280 changed files with 20741 additions and 1989 deletions
@@ -1,9 +1,9 @@
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"config_schema_version": 2,
"workload_id": "wan-t2v-1.3b",
"variant_id": "canonical",
"benchmark_version": 1,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"description": "Wan2.1 T2V 1.3B inference performance",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
+1
View File
@@ -481,6 +481,7 @@ steps:
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/performance/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
+23 -1
View File
@@ -80,6 +80,23 @@ MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUI
POST_RUN_HOOK=""
is_truthy() {
case "${1:-}" in
1|true|TRUE|yes|YES|on|ON) return 0 ;;
*) return 1 ;;
esac
}
ssim_bootstrap_args() {
local title="${PR_TITLE:-}"
local message="${BUILDKITE_MESSAGE:-}"
if is_truthy "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-}" \
|| [[ "$title" == *"[new-model]"* ]] \
|| [[ "$message" == *"[new-model]"* ]]; then
printf ' --bootstrap-mode'
fi
}
upload_performance_artifacts() {
SHORT_SHA=${BUILDKITE_COMMIT:0:7}
LOCAL_DIR="downloaded_reports"
@@ -172,7 +189,12 @@ case "$TEST_TYPE" in
;;
"ssim")
log "Running SSIM tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_SSIM_TEST_FILE::run_ssim_tests"
SSIM_BOOTSTRAP_ARGS=$(ssim_bootstrap_args)
if [ -n "$SSIM_BOOTSTRAP_ARGS" ]; then
log "SSIM bootstrap mode enabled for new-model reference draft generation"
fi
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run "
MODAL_COMMAND+="$MODAL_SSIM_TEST_FILE::run_ssim_tests$SSIM_BOOTSTRAP_ARGS"
;;
"training")
log "Running training tests..."
+133
View File
@@ -0,0 +1,133 @@
#!/usr/bin/env bash
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
# already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no full suite; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
# - watched runs pending -> poll until done
# - docs run absent -> not applicable after a
# short grace period ('Deploy Documentation' is path-filtered on PRs)
# - pre-commit run absent -> keep polling: pre-commit
# is never path-filtered, so its absence is always anomalous
# - 'ready' label removed while waiting -> exit 1 (fail CLOSED:
# un-labeling is a deliberate maintainer action)
# - GitHub API unreachable or timeout -> exit 0 (fail OPEN,
# loud warning: never brick CI on a GitHub outage)
#
# Required env: PR_SHA (PR head commit), PR_NUMBER, GITHUB_REPOSITORY, GH_TOKEN.
set -euo pipefail
: "${PR_SHA:?PR_SHA (PR head commit) is required}"
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the full suite
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
WATCHED_REGEX='^(pre-commit|Deploy Documentation)$'
POLL_SECS="${POLL_SECS:-20}"
GRACE_SECS="${GRACE_SECS:-60}"
MAX_WAIT_SECS="${MAX_WAIT_SECS:-1500}"
# Bound each API call so a hung connection hits the 3-strike fail-open path
# instead of pinning the loop until the job timeout (which would fail closed
# on exactly the GitHub-outage case this script is meant to survive).
if command -v timeout >/dev/null 2>&1; then
gh_api() { timeout 30 gh api "$@"; }
else
gh_api() { gh api "$@"; } # macOS dev boxes; CI always has coreutils timeout
fi
# The workflow checked the label before starting the gate, but the wait can
# last ~25 min: re-check once before any exit 0 and fail closed if 'ready'
# was removed in the meantime. An API error here proceeds (the label was
# present when the gate started; never brick CI on an outage).
recheck_ready_label() {
local pr_json
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
exit 1
fi
else
echo "::warning::Could not re-check the 'ready' label on PR #${PR_NUMBER}; proceeding (it was present when the gate started)."
fi
}
start=$(date +%s)
api_fails=0
missing=""
while true; do
elapsed=$(( $(date +%s) - start ))
if runs_json=$(gh_api "repos/${GITHUB_REPOSITORY}/actions/runs?head_sha=${PR_SHA}&per_page=100" 2>/dev/null) \
&& state=$(jq --arg re "$WATCHED_REGEX" '
[.workflow_runs[]? | select(.name // "" | test($re))]
| group_by(.name) | map(max_by(.id))
| map({name, status, conclusion})' <<<"$runs_json" 2>/dev/null); then
api_fails=0
echo "t+${elapsed}s watched checks: $(jq -c . <<<"$state")"
failed=$(jq -r '[.[] | select(.status == "completed"
and (.conclusion | IN("success", "skipped", "neutral", "cancelled") | not))]
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
fi
# 'cancelled' counts as pending: wait for a re-run to reach a real verdict
# (bounded by MAX_WAIT, then the fail-open below).
pending=$(jq '[.[] | select(.status != "completed" or .conclusion == "cancelled")] | length' <<<"$state")
missing=$(jq -r --argjson watched "$WATCHED_NAMES" '($watched - map(.name)) | join(", ")' <<<"$state")
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — full suite may proceed."
exit 0
fi
case "$missing" in
*pre-commit*)
echo "pre-commit run not found for ${PR_SHA} yet; waiting (pre-commit is never path-filtered, so its absence is anomalous)."
;;
*)
if [ "$elapsed" -ge "$GRACE_SECS" ]; then
recheck_ready_label
echo "::warning::Watched run(s) never appeared for ${PR_SHA}: ${missing} (path-filtered, likely not applicable). Proceeding on the checks that did run."
exit 0
fi
echo "Waiting up to ${GRACE_SECS}s grace for path-filtered run(s) to appear: ${missing}."
;;
esac
fi
else
api_fails=$(( api_fails + 1 ))
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
exit 0
fi
sleep "$POLL_SECS"
done
+122
View File
@@ -0,0 +1,122 @@
#!/usr/bin/env bash
# Self-test for gate_full_suite.sh using a mocked `gh`. No network, runs on
# any dev box: bash .github/scripts/test_gate_full_suite.sh
set -u
here=$(cd "$(dirname "$0")" && pwd)
tmp=$(mktemp -d)
trap 'rm -rf "$tmp"' EXIT
# Mock gh. Asserts the exact endpoint (including head_sha) it is called
# with — an endpoint typo in the gate script fails the test rather than
# silently serving canned data. On the runs endpoint it serves
# $MOCK_DIR/response_<call#>.json, sticking on the highest existing file,
# and exits 1 if none exist (simulates a GitHub API outage). On the pulls
# endpoint it serves $MOCK_DIR/pr.json, defaulting to a 'ready'-labeled PR.
cat > "$tmp/gh" <<'EOF'
#!/usr/bin/env bash
if [ "${1:-}" != "api" ]; then
echo "unexpected gh invocation: $*" >> "$MOCK_DIR/endpoint_error"
exit 2
fi
case "${2:-}" in
"repos/o/r/actions/runs?head_sha=deadbeef&per_page=100")
n=$(( $(cat "$MOCK_DIR/count" 2>/dev/null || echo 0) + 1 ))
echo "$n" > "$MOCK_DIR/count"
while [ "$n" -gt 0 ]; do
if [ -f "$MOCK_DIR/response_$n.json" ]; then
cat "$MOCK_DIR/response_$n.json"
exit 0
fi
n=$(( n - 1 ))
done
echo "api outage" >&2
exit 1
;;
"repos/o/r/pulls/42")
if [ -f "$MOCK_DIR/pr.json" ]; then
cat "$MOCK_DIR/pr.json"
else
echo '{"labels": [{"name": "ready"}]}'
fi
;;
*)
echo "unexpected gh endpoint: $2" >> "$MOCK_DIR/endpoint_error"
exit 2
;;
esac
EOF
chmod +x "$tmp/gh"
PC_OK='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "success"}'
PC_BAD='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "failure"}'
PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion": null}'
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
fails=0
want_log="" # optional: expect() also greps out.log for this regex, then resets
pr_json="" # optional: served for the pulls (label re-check) endpoint, then resets
raw_body="" # optional: serve responses verbatim instead of wrapping in workflow_runs
expect() { # <name> <expected-exit> <response json>...
local name=$1 want=$2 dir i=1
shift 2
dir=$(mktemp -d "$tmp/test_XXXXXX")
for body in "$@"; do
if [ -n "$raw_body" ]; then
printf '%s' "$body" > "$dir/response_$i.json"
else
printf '{"workflow_runs": [%s]}' "$body" > "$dir/response_$i.json"
fi
i=$(( i + 1 ))
done
[ -n "$pr_json" ] && printf '%s' "$pr_json" > "$dir/pr.json"
( export PATH="$tmp:$PATH" MOCK_DIR="$dir" PR_SHA=deadbeef PR_NUMBER=42 \
GITHUB_REPOSITORY=o/r POLL_SECS=0 GRACE_SECS=1 MAX_WAIT_SECS=3
bash "$here/gate_full_suite.sh" > "$dir/out.log" 2>&1 )
local rc=$?
if [ "$rc" -ne "$want" ]; then
echo "FAIL: $name (exit $rc, want $want)"
cat "$dir/out.log"
fails=1
elif [ -f "$dir/endpoint_error" ]; then
echo "FAIL: $name (mock gh got an unexpected call)"
cat "$dir/endpoint_error"
fails=1
elif [ -n "$want_log" ] && ! grep -Eq "$want_log" "$dir/out.log"; then
echo "FAIL: $name (log does not match: $want_log)"
cat "$dir/out.log"
fails=1
else
echo "ok: $name"
fi
want_log="" pr_json="" raw_body=""
}
expect "both green -> proceed" 0 "$PC_OK, $DOCS_OK, $OTHER, $NULL_NAME"
expect "docs build failed -> blocked" 1 "$PC_OK, $DOCS_BAD"
expect "pre-commit failed -> blocked" 1 "$PC_BAD"
expect "pending then green -> proceed" 0 "$PC_PENDING" "$PC_OK, $DOCS_OK"
want_log="never appeared.*Deploy Documentation"
expect "docs run absent (path-filtered) -> proceed after grace" 0 "$PC_OK"
expect "API outage -> fail open" 0
want_log="FAILING OPEN"
expect "pending past MAX_WAIT -> fail open" 0 "$PC_PENDING"
want_log="FAILING OPEN"
expect "unrelated runs only -> no grace, fail open at MAX_WAIT" 0 "$OTHER"
expect "cancelled docs then green -> proceed" 0 \
"$PC_OK, $DOCS_CANCELLED" "$PC_OK, $DOCS_OK"
want_log="FAILING OPEN"
expect "cancelled docs forever -> fail open at MAX_WAIT" 0 "$PC_OK, $DOCS_CANCELLED"
want_log="FAILING OPEN"
expect "pre-commit absent -> no grace, fail open at MAX_WAIT" 0 "$DOCS_OK"
expect "duplicate run names -> latest wins" 0 "$PC_BAD, $PC_OK_RERUN, $DOCS_OK"
raw_body=1
expect "garbage response body -> fail open" 0 "this is not json"
pr_json='{"labels": [{"name": "other"}]}'
expect "ready label removed mid-gate -> blocked" 1 "$PC_OK, $DOCS_OK"
exit "$fails"
+3
View File
@@ -47,3 +47,6 @@ jobs:
- uses: pre-commit/action@v3.0.1
with:
extra_args: --all-files --hook-stage manual
# After pre-commit so a self-test failure cannot mask lint failures.
- name: Full-suite gate self-test
run: bash .github/scripts/test_gate_full_suite.sh
+17 -10
View File
@@ -52,6 +52,7 @@ jobs:
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
core.setOutput('pr_title', pr.title);
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
@@ -60,6 +61,7 @@ jobs:
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
PR_TITLE: ${{ steps.label.outputs.pr_title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -71,6 +73,7 @@ jobs:
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -80,11 +83,12 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
}
}')"
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
parse-command:
if: >-
@@ -241,6 +245,7 @@ jobs:
TEST_SCOPE: ${{ needs.parse-command.outputs.test_scope }}
FULL_SUITE: ${{ needs.parse-command.outputs.full_suite }}
TEST_TYPE: ${{ needs.parse-command.outputs.test_type }}
PR_TITLE: ${{ github.event.issue.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -257,6 +262,7 @@ jobs:
--arg full_suite "$FULL_SUITE" \
--arg test_type "$TEST_TYPE" \
--arg pr_number "$PR_NUMBER" \
--arg pr_title "$PR_TITLE" \
'{
commit: $commit,
branch: $branch,
@@ -266,8 +272,9 @@ jobs:
pull_request_base_branch: "main",
env: {
TEST_SCOPE: $test_scope,
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number
}
}')"
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number,
PR_TITLE: $pr_title
}
}')"
+21 -1
View File
@@ -7,6 +7,7 @@ on:
permissions:
contents: read
pull-requests: read
actions: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
@@ -18,6 +19,8 @@ jobs:
(github.event.action == 'labeled' && github.event.label.name == 'ready')
|| github.event.action == 'synchronize'
runs-on: ubuntu-latest
# Gate below may wait for cheap checks (up to MAX_WAIT_SECS = 25 min).
timeout-minutes: 35
steps:
- name: Check ready label
id: check
@@ -49,6 +52,20 @@ jobs:
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
# Checks out the BASE branch (default for pull_request_target), so PR
# authors cannot tamper with the gate script.
- name: Checkout gate script
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite Full Suite
if: steps.check.outputs.has_ready == 'true'
env:
@@ -56,6 +73,7 @@ jobs:
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -67,6 +85,7 @@ jobs:
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -78,6 +97,7 @@ jobs:
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
@@ -1,6 +1,7 @@
import { useEffect, useMemo, useState } from "react";
import { fetchSummary, fetchTrends, refreshData, RunSource, SummaryResponse, TrendGroup, TrendPoint } from "./api";
import { fetchSummary, fetchTrends, refreshData } from "./api";
import type { CohortValue, RunSource, SummaryResponse, TrendGroup, TrendPoint } from "./api";
const METRIC_KEYS = ["latency", "throughput", "memory", "text_encoder_time_s", "dit_time_s", "vae_decode_time_s"];
const RUN_SOURCES: Array<{ value: "" | RunSource; label: string }> = [
@@ -109,6 +110,61 @@ function metricLabel(metricKey: string) {
return METRIC_DEFINITIONS[metricKey]?.label ?? metricKey;
}
type CohortFields = {
model_id: string;
gpu_type: string;
workload_id: CohortValue;
variant_id: CohortValue;
benchmark_version: CohortValue;
recipe_fingerprint: CohortValue;
hardware_profile_id: CohortValue;
software_profile_id: CohortValue;
};
function cohortValue(value: CohortValue) {
if (value === null || value === undefined || value === "") {
return "legacy";
}
return String(value);
}
function shortCohortValue(value: CohortValue) {
const text = cohortValue(value);
if (text === "legacy" || text.length <= 14) {
return text;
}
return text.slice(0, 12);
}
function cohortKey(cohort: CohortFields) {
return [
cohort.model_id,
cohort.gpu_type,
cohortValue(cohort.workload_id),
cohortValue(cohort.variant_id),
cohortValue(cohort.benchmark_version),
cohortValue(cohort.recipe_fingerprint),
cohortValue(cohort.hardware_profile_id),
cohortValue(cohort.software_profile_id)
].join("|");
}
function cohortTitle(cohort: CohortFields) {
const workload = cohortValue(cohort.workload_id);
const variant = cohortValue(cohort.variant_id);
const version = cohortValue(cohort.benchmark_version);
const versionLabel = version === "legacy" ? version : `v${version}`;
return `${workload} / ${variant} / ${versionLabel}`;
}
function cohortDetail(cohort: CohortFields) {
return [
`recipe ${shortCohortValue(cohort.recipe_fingerprint)}`,
shortCohortValue(cohort.hardware_profile_id),
shortCohortValue(cohort.software_profile_id)
].join(" | ");
}
function formatMetricValue(metricKey: string, value: number | null | undefined, tooltip = false) {
const definition = METRIC_DEFINITIONS[metricKey];
if (!definition) {
@@ -171,7 +227,9 @@ function TrendChart({ group, metricKey }: { group: TrendGroup; metricKey: string
top: `${(activePoint.y / height) * 100}%`
}
: undefined;
const ariaLabel = `${metricLabel(metricKey)} trend for ${group.model_id} on ${group.gpu_type}`;
const ariaLabel = `${metricLabel(metricKey)} trend for ${group.model_id} on ${group.gpu_type}, ${cohortTitle(
group
)}`;
return (
<div className="chart-shell">
@@ -419,7 +477,7 @@ export default function App() {
<section className="panel">
<div className="panel-header">
<h2>Latest Status</h2>
<span>{latestRows.length} model/GPU groups</span>
<span>{latestRows.length} comparison cohorts</span>
</div>
{latestRows.length === 0 ? (
<div className="empty">No records match the selected filters.</div>
@@ -432,6 +490,7 @@ export default function App() {
<th>Recomputed</th>
<th>Model</th>
<th>GPU</th>
<th>Cohort</th>
<th>Commit</th>
<th>Source</th>
<th>Baseline</th>
@@ -446,7 +505,7 @@ export default function App() {
</thead>
<tbody>
{latestRows.map((row) => (
<tr key={`${row.model_id}-${row.gpu_type}`}>
<tr key={cohortKey(row)}>
<td>
<span className={`badge ${row.status}`}>{row.status}</span>
</td>
@@ -457,6 +516,12 @@ export default function App() {
</td>
<td>{row.model_id}</td>
<td>{row.gpu_type}</td>
<td>
<div className="cohort-cell">
<strong>{cohortTitle(row)}</strong>
<span>{cohortDetail(row)}</span>
</div>
</td>
<td>{shortSha(row.commit_sha)}</td>
<td>
<span className={`source-badge source-${row.run_source}`}>{runSourceLabel(row.run_source)}</span>
@@ -495,11 +560,13 @@ export default function App() {
) : (
trends.map((group) =>
METRIC_KEYS.map((metricKey) => (
<article className="trend-card" key={`${group.model_id}-${group.gpu_type}-${metricKey}`}>
<article className="trend-card" key={`${cohortKey(group)}-${metricKey}`}>
<div>
<h3>{metricLabel(metricKey)}</h3>
<p>
{group.model_id} | {group.gpu_type}
<span>{cohortTitle(group)}</span>
<span>{cohortDetail(group)}</span>
</p>
</div>
<TrendChart group={group} metricKey={metricKey} />
+14 -3
View File
@@ -13,6 +13,17 @@ export type MetricValue = {
precision: number;
};
export type CohortValue = string | number | null;
export type ComparisonCohort = {
workload_id: CohortValue;
variant_id: CohortValue;
benchmark_version: CohortValue;
recipe_fingerprint: CohortValue;
hardware_profile_id: CohortValue;
software_profile_id: CohortValue;
};
export type SummaryRow = {
model_id: string;
gpu_type: string;
@@ -34,7 +45,7 @@ export type SummaryRow = {
build_id: string;
job_id: string;
metrics: Record<string, MetricValue>;
};
} & ComparisonCohort;
export type RunSource = "pr" | "local" | "scheduled_main" | "unknown";
@@ -68,13 +79,13 @@ export type TrendPoint = {
build_id: string;
job_id: string;
metrics: Record<string, number | null>;
};
} & ComparisonCohort;
export type TrendGroup = {
model_id: string;
gpu_type: string;
points: TrendPoint[];
};
} & ComparisonCohort;
export type TrendsResponse = {
groups: TrendGroup[];
@@ -149,6 +149,11 @@ h3 {
font-size: 0.82rem;
}
.trend-card p {
display: grid;
gap: 2px;
}
.stat strong {
display: block;
margin-top: 8px;
@@ -186,7 +191,7 @@ h3 {
table {
width: 100%;
min-width: 1120px;
min-width: 1260px;
border-collapse: collapse;
}
@@ -209,6 +214,25 @@ td {
font-size: 0.9rem;
}
.cohort-cell {
display: grid;
gap: 2px;
}
.cohort-cell strong,
.trend-card p span {
color: #1b2836;
font-size: 0.78rem;
font-weight: 700;
}
.cohort-cell span,
.trend-card p span + span {
color: #607080;
font-family: ui-monospace, SFMono-Regular, Menlo, Monaco, Consolas, "Liberation Mono", monospace;
font-size: 0.72rem;
}
.badge {
display: inline-flex;
align-items: center;
@@ -0,0 +1,22 @@
{
"alpha_yaw": 0.08734091699186919,
"alpha_pitch": 0.08169667696275307,
"alpha_turn": 5.724587470723463e-17,
"beta_fwd": 0.02842768078408099,
"beta_strafe": 0.022531015077067108,
"focal_length": 457.0,
"frame_shape": [
352,
640
],
"calibrated_from": [
"1_wasd_only",
"camera",
"camera4hold_alpha1",
"fully_random",
"wasdonly_alpha1",
"wasd4holdrandview_simple_1key1mouse1"
],
"residual_rms": 15.890399609478676,
"n_equations": 4125232
}
+7
View File
@@ -99,6 +99,13 @@ status.
Full Suite is also path-filtered. It validates broader behavior before Mergify
can merge a PR.
A `ready`-labeled PR does not hit Buildkite immediately:
`ci-trigger-full-suite.yml` first runs `.github/scripts/gate_full_suite.sh`,
which waits for the cheap Tier-1 checks (pre-commit, docs build) on the PR
head. A red cheap check blocks the suite (fail closed; the next push re-arms
it), while a GitHub outage or a >25 min wait lets it run anyway (fail open).
`/test full` bypasses the gate.
| Buildkite label | `TEST_TYPE` | Main watched paths |
|---|---|---|
| SSIM Tests | `ssim` | `fastvideo/**/*.py`, `pyproject.toml`, `docker/Dockerfile` |
+131 -40
View File
@@ -79,10 +79,11 @@ fastvideo/performance/
```
The HF dataset (`FastVideo/performance-tracking` by default) holds one
normalized JSON per `(model_id, gpu_type, run)` tuple. The rolling baseline is
the median of the last 5 successful, baseline-eligible records for that
model+GPU. PR and local records are visible in the dashboard but are not
baseline eligible.
normalized JSON per run. For v2 records, the rolling baseline is the median of
the last 5 successful, baseline-eligible records in the same comparison cohort:
`model_id`, `gpu_type`, `workload_id`, `variant_id`, `benchmark_version`,
`recipe_fingerprint`, `hardware_profile_id`, and `software_profile_id`. PR and
local records are visible in the dashboard but are not baseline eligible.
## Planned Coverage
@@ -158,13 +159,16 @@ unrealistic memory growth, and optionally large component-specific slowdowns
even when the rolling baseline is empty. They are hand-set with generous
headroom and almost never need touching.
### Rolling baseline (per `(model_id, gpu_type)`)
### Rolling baseline (per comparison cohort)
`compare_baseline.py` loads the last 5 successful, baseline-eligible records
for the same `(model_id, gpu_type)` from the HF dataset, computes the median
for each available metric, and evaluates the current run with the metric's
rolling regression policy. For latency, memory, and component times, higher
values are regressions. For throughput, lower values are regressions.
for the same comparison cohort from the HF dataset, computes the median for
each available metric, and evaluates the current run with the metric's
rolling regression policy. For v2 records, that cohort is `model_id`,
`gpu_type`, `workload_id`, `variant_id`, `benchmark_version`,
`recipe_fingerprint`, `hardware_profile_id`, and `software_profile_id`. For
latency, memory, and component times, higher values are regressions. For
throughput, lower values are regressions.
A metric exceeds its rolling threshold when both of these are true:
@@ -201,9 +205,9 @@ configs and remain loadable. New or migrated configs should use
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"config_schema_version": 2,
"workload_id": "wan-t2v-1.3b",
"variant_id": "canonical",
"benchmark_version": 1
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2
}
```
@@ -214,21 +218,27 @@ metadata that make the measured workload explicit:
| Field | Purpose |
|---|---|
| `workload_id` | Stable benchmark family, such as `wan-t2v-1.3b`. |
| `variant_id` | Intentional recipe family, such as `canonical`. |
| `workload_id` | Stable benchmark family, such as `wan-t2v`. |
| `variant_id` | Intentional recipe family, including model size and parallelism config, such as `1.3b-sp2`. |
| `benchmark_version` | Version of the measurement protocol and comparison policy. |
If a config declares `config_schema_version: 2`, loading fails clearly when any
required v2 identity field is missing. If v2 identity or metadata fields are
added without `config_schema_version: 2`, loading also fails so partial
migrations do not silently run as v1 configs. Optional v2 metadata fields
reserved for follow-up work, such as `recipe`, `metric_threshold_policy`, and
`quality_metadata`, must be JSON objects when present.
reserved for follow-up work, such as `metric_threshold_policy` and
`quality_metadata`, must be JSON objects when present. (`recipe` is emitted
by the harness and is not config-declarable.)
Recipe fingerprinting, hardware/software profile IDs, exact-identity
comparison, metric-specific threshold policy behavior, promoted baselines, and
dashboard regrouping are separate follow-up changes. Until those land, rolling
baseline comparison remains keyed by `(model_id, gpu_type)`.
comparison, and dashboard cohort grouping land with this change: v2 records
compare only within their identity cohort, and a record that opens a NEW
cohort is marked `baseline_status: "initialized_new_cohort"` (regression
gating starts once that cohort accumulates history). Legacy v1 configs still
run and are normalized for reporting, but their records skip rolling-baseline
comparison entirely (`baseline_status: "skipped_missing_identity"`, never
baseline eligible); only static thresholds gate them. Metric-specific
threshold policies and promoted baselines remain separate follow-ups.
### Raw record (`results/perf_*.json`)
@@ -237,10 +247,10 @@ Written by `test_inference_performance.py`. One file per benchmark run.
```jsonc
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"config_schema_version": 2,
"workload_id": "wan-t2v-1.3b",
"variant_id": "canonical",
"benchmark_version": 1,
"result_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"model_short_name": "Wan2.1-T2V-1.3B-Diffusers",
"device": "NVIDIA L40S",
"num_gpus": 2,
@@ -266,11 +276,54 @@ Written by `test_inference_performance.py`. One file per benchmark run.
}
},
"commit": "<full sha>",
"run_source": "pr",
"branch": "feature/perf-change",
"pr_number": "1234",
"test_scope": "direct",
"build_url": "https://buildkite.example/build",
"build_id": "<buildkite-build-id>",
"job_id": "<buildkite-job-id>",
"timestamp": "2026-05-08T22:00:00+00:00",
"quality_metadata": { "quality_status": "canonical" },
"text_encoder_time_s": 2.141,
"dit_time_s": 8.437,
"vae_decode_time_s": 3.208
"vae_decode_time_s": 3.208,
"recipe": {
"recipe_schema_version": 1,
"benchmark": {
"benchmark_id": "wan-t2v-1.3b-2gpu",
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2
},
"model": { "model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" },
"init_kwargs": { "num_gpus": 2, "sp_size": 2, "tp_size": 1 },
"generation_kwargs": { "height": 480, "width": 832, "num_frames": 45 },
"inputs": { "prompt_count": 1, "prompt_sha256": ["<measured-prompt-sha256>"] },
"attention": { "requested_backend": "FLASH_ATTN", "resolved_backend": "FLASH_ATTN" }
},
"recipe_fingerprint": "<sha256>",
"hardware_profile": {
"device_type": "cuda",
"gpu_count": 2,
"gpus": [{ "name": "NVIDIA L40S", "memory_gb": 48, "compute_capability": "8.9" }],
"interconnect": "none_or_partial"
},
"hardware_profile_id": "hw-<sha256-prefix>",
"software_profile": {
"python": "3.12",
"pytorch": "2.12",
"cuda": "13.0",
"packages": {
"fastvideo_kernel": "0.3.2",
"flashinfer": "0.2.11",
"nvidia_cutlass_dsl": "4.5.0",
"triton": "3.4.1"
}
},
"software_profile_id": "sw-<sha256-prefix>",
"environment_metadata": { "env": { "IMAGE_VERSION": "py3.12-cuda13.0.0" } },
"environment_fingerprint": "env-<sha256-prefix>"
}
```
@@ -282,6 +335,10 @@ result, used as the rolling-baseline source of truth.
```jsonc
{
"model_id": "wan-t2v-1.3b-2gpu",
"result_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"timestamp": "2026-05-08T22:00:00+00:00",
"commit_sha": "<full sha>",
"gpu_type": "NVIDIA L40S",
@@ -298,17 +355,46 @@ result, used as the rolling-baseline source of truth.
"gated": true
}
},
"recipe_fingerprint": "<sha256>",
"hardware_profile_id": "hw-<sha256-prefix>",
"software_profile_id": "sw-<sha256-prefix>",
"environment_fingerprint": "env-<sha256-prefix>",
"run_source": "pr",
"branch": "feature/perf-change",
"pr_number": "1234",
"test_scope": "direct",
"build_url": "https://buildkite.example/build",
"build_id": "<buildkite-build-id>",
"job_id": "<buildkite-job-id>",
"quality_metadata": { "quality_status": "canonical" },
"success": true
}
```
### Compatibility with legacy records
Older records in the HF dataset may not have component timing fields. The
comparator ignores missing or `null` metrics when computing a median, and the
dashboard lists skipped plots for metric series that have no non-null values.
Records missing both `run_source` and `baseline_eligible` are treated as legacy
successful main/full-suite uploads and remain eligible for rolling baselines.
Older records in the HF dataset may not have `result_schema_version`,
component timing fields, or v2 identity/profile fields. Records without
`result_schema_version` are treated as v1. The comparator ignores missing or
`null` metrics when computing a median, and the dashboard lists skipped plots
for metric series that have no non-null values. Records missing both
`run_source` and `baseline_eligible` are treated as legacy successful
main/full-suite uploads and remain eligible for rolling baselines.
Current `perf_*.json` artifacts that lack the v2 comparison identity are
normalized for reporting but skip rolling-baseline comparison and are not marked
baseline eligible.
New records compare only against the same `model_id`, `gpu_type`,
`workload_id`, `variant_id`, `benchmark_version`, `recipe_fingerprint`,
`hardware_profile_id`, and `software_profile_id` cohort.
`environment_metadata` and `environment_fingerprint` are audit data and are not
part of the comparison key.
The recipe prompt digests describe the prompts actually measured by the
benchmark run; extra configured prompts are ignored unless the benchmark runner
executes them.
Software profile package cohorts keep exact versions for relevant
attention/kernel packages, including FastVideo kernels, FlashAttention,
FlashInfer, Cutlass DSL, SageAttention, Triton, and xFormers when installed.
## Environment variable reference
@@ -318,7 +404,7 @@ successful main/full-suite uploads and remain eligible for rolling baselines.
| `PERF_REPORTS_DIR` | `/root/data/perf_reports` | `compare_baseline.py`, `dashboard.py` | Where the Markdown summary and Plotly HTML get written for Buildkite to pick up. |
| `HF_REPO_ID` | `FastVideo/performance-tracking` | `fastvideo/performance/hf_store.py` | HF dataset repo holding rolling-baseline records. |
| `HF_API_KEY`, `HUGGINGFACE_HUB_TOKEN`, `HF_TOKEN` | unset | `fastvideo/performance/hf_store.py` | Required for upload or private dataset reads. |
| `PERF_RUN_SOURCE` | inferred | `compare_baseline.py` | Source metadata for uploaded records: `pr`, `local`, `scheduled_main`, or `unknown`. |
| `PERF_RUN_SOURCE` | inferred | `compare_baseline.py`, `test_inference_performance.py` | Source metadata for uploaded records: `pr`, `local`, `scheduled_main`, or `unknown`. |
| `PERF_UPLOAD_POLICY` | `never` | `compare_baseline.py` | Upload policy: `never`, `pass`, or `always`. |
| `PERF_PYTEST_RC` | unset | `compare_baseline.py` | Static-threshold pytest exit code, used so scheduled-main failures can be uploaded with `success=false`. |
| `TEST_SCOPE` | unset | `compare_baseline.py` | CI context used to infer scheduled-main runs together with `BUILDKITE_BRANCH=main`. |
@@ -335,11 +421,14 @@ point is `fastvideo/tests/modal/pr_test.py:run_performance_tests` and the
Buildkite artifact upload is in
`.buildkite/scripts/pr_test.sh:upload_performance_artifacts`.
Each performance build runs pytest first. If that fixed-threshold phase fails,
PR/direct runs skip `compare_baseline.py` because they only upload passing
records. Scheduled-main runs still execute `compare_baseline.py` with
`PERF_PYTEST_RC` set so the failed canonical attempt is visible in normalized
JSON and dashboard history. The dashboard runs best-effort for observability.
Each performance build runs pytest first. PR and direct runs only continue to
`compare_baseline.py` when that fixed-threshold phase passes; if pytest fails,
Markdown summaries and normalized JSON artifacts are not emitted. Scheduled
main runs set `PERF_UPLOAD_POLICY=always`, so they still run
`compare_baseline.py` (with `PERF_PYTEST_RC` set) after a fixed-threshold
failure. Those failed scheduled main runs emit summaries and normalized
records, upload records with `success=false`, and are excluded from future
rolling baselines. The dashboard still runs best-effort for observability.
When the rolling-baseline phase runs, it emits:
* **Markdown summary** — appended to `$GITHUB_STEP_SUMMARY` when that variable
@@ -347,7 +436,7 @@ When the rolling-baseline phase runs, it emits:
per-benchmark row with current vs. baseline values for latency, throughput,
memory, text encoder time, DiT time, and VAE decode time.
* **Plotly dashboard** — `dashboard_<sha>_<ts>.html` showing time-series for
each metric grouped by `(model_id, gpu_type)`.
each metric grouped by comparison cohort.
* **Normalized records** — `normalized_perf_*.json`, one per benchmark.
Useful as input to the
[`reseed-performance-baseline`](https://github.com/hao-ai-lab/FastVideo/blob/main/.agents/skills/reseed-performance-baseline/SKILL.md)
@@ -364,7 +453,7 @@ When the rolling-baseline phase runs, it emits:
"benchmark_id": "<unique-id>",
"config_schema_version": 2,
"workload_id": "<stable-workload-id>",
"variant_id": "canonical",
"variant_id": "<variant, e.g. 1.3b-sp2>",
"benchmark_version": 1,
"model": { "model_path": "...", "model_short_name": "..." },
"init_kwargs": { "num_gpus": 1, ... },
@@ -391,7 +480,9 @@ When the rolling-baseline phase runs, it emits:
Legacy v1 configs without `config_schema_version` still load, but should not
gain v2 identity or metadata fields until they are migrated to
`config_schema_version: 2`.
`config_schema_version: 2`. For v2 configs, `workload_id`, `variant_id`,
and `benchmark_version` are part of the comparison key; benchmark runs
fail if any of these identity fields are missing.
2. The pytest test auto-discovers all configs — no test code needed. CI
picks it up on the next `/test performance` run.
@@ -419,8 +510,8 @@ When the rolling-baseline phase runs, it emits:
## Troubleshooting
**"No baseline for ... Initializing"** — first run for this `(model_id,
gpu_type)`. Run will pass and (if persisting) seed the first record.
**"No baseline for ... Initializing"** — first run for this comparison cohort.
Run will pass and (if persisting) seed the first record.
**Persistent failure right after a torch / kernel / image upgrade** —
genuine regression *or* baseline drift. Compare the failing normalized record
+24
View File
@@ -180,6 +180,30 @@ python fastvideo/tests/ssim/reference_videos_cli.py copy-local \
--device-folder L40S_reference_videos
```
### SSIM Bootstrap Mode
Normal SSIM runs are strict: if a reference video or latent is missing, the
test fails. For new-model PRs, CI can run SSIM in bootstrap mode so missing
references are uploaded as draft artifacts for review instead of immediately
blocking on a missing canonical reference.
Buildkite enables SSIM bootstrap mode when either condition is true:
- the PR title or Buildkite message contains `[new-model]`;
- `FASTVIDEO_SSIM_BOOTSTRAP_MODE=1` is set for the Buildkite job.
Bootstrap mode passes `--ssim-bootstrap-mode` to pytest. When a generated
artifact is available, the test uploads it under the `drafts/...` namespace in
the SSIM reference repo and marks that case as expected-failed. After reviewing
the draft, promote it into the canonical reference layout:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py promote-draft \
--quality-tier default \
--device-folder L40S_reference_videos \
--model-id <model_id>
```
## CI Integration
FastVideo CI tests are orchestrated by Buildkite and run on Modal GPU
@@ -191,6 +191,9 @@ surfaces:
sources:
- fastvideo.configs.pipelines.gen3c.Gen3CConfig
- fastvideo.configs.pipelines.gen3c.Gen3CInferenceConfig
color_correction_strength:
sources:
- fastvideo.configs.pipelines.dreamx_world.DreamXWorld5BARPipelineConfig
default_camera_rotation:
sources:
- fastvideo.configs.pipelines.gen3c.Gen3CConfig
+1 -1
View File
@@ -330,7 +330,7 @@ at FastVideo's CI — before the Dynamo-side integration even knows.
internal; presets identify them by name on
`PipelineSelection.preset`).
* `fastvideo.fastvideo_args.FastVideoArgs` (legacy compat type).
* `fastvideo.api.compat.*` private helpers
* `fastvideo.api.translation.*` private helpers
(`_validate_continuation_state` etc.) — the public boundary is
`VideoGenerator` + `fastvideo.api`.
* Any flat legacy LTX-2 kwarg (`ltx2_refine_upsampler_path`,
+8 -6
View File
@@ -48,15 +48,17 @@ highest first:
`request.model_fields_set` (Pydantic v2). Unset fields do not count,
even if the Pydantic model has a schema default for them.
2. **`ServeConfig.default_request` (operator-explicit)** — projected via
[`explicit_request_updates()`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/api/compat.py);
[`explicit_request_updates()`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/api/translation.py);
only fields the operator actually wrote into the YAML count as
defaults. Every other field inherits the schema default rather than
being pinned.
defaults (an explicit `null` counts as unset). Every other sampling
field stays `None` — "inherit the model preset" — and other sections
keep their schema defaults without being pinned.
3. **Hardcoded fallback** — e.g. `fps = 24`.
The gate matters: both surfaces carry schema defaults. Without
`model_fields_set` / explicit-path tracking, schema defaults would
masquerade as intent and silently shadow the other side.
The gate matters: the Pydantic surface carries schema defaults and the
dataclass surface carries non-None defaults outside `sampling`. Without
`model_fields_set` / explicit-path tracking, defaults would masquerade
as intent and silently shadow the other side.
See [`video_api.py::_build_generation_kwargs`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/entrypoints/openai/video_api.py)
for the canonical implementation; the per-request assembly lives there,
+2
View File
@@ -58,6 +58,8 @@ pipeline initialization and sampling.
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| Wan2.2 TI2V 5B | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | 720P | ⭕ | ⭕ | ✅ | ⭕ | ⭕ |
| DreamX-World 5B Cam | `FastVideo/DreamX-World-5B-Cam-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| DreamX-World 5B AR | `FastVideo/DreamX-World-5B-Diffusers` | 704px1280p | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Lucy Edit Dev 5B*** | `decart-ai/Lucy-Edit-Dev` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Wan2.2 T2V A14B | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
| Wan2.2 I2V A14B | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
+91
View File
@@ -0,0 +1,91 @@
# Training Trackers
FastVideo can send training metrics and validation media to Weights & Biases
or SwanLab. Tracking runs only on global rank 0, and local tracker files are
stored under `<output_dir>/tracker`.
## Supported Trackers
| Value | Backend | Installation |
|-------|---------|--------------|
| `wandb` | Weights & Biases | Included with FastVideo |
| `swanlab` | SwanLab | Install the optional `swanlab` dependency |
| `none` | Disable external tracking | No additional package |
You can enable more than one backend, for example `trackers: [wandb, swanlab]`.
Metrics and validation media are converted to the artifact type required by
each backend.
## Install SwanLab
For a published FastVideo installation, install the SwanLab extra:
```bash
uv pip install "fastvideo[swanlab]"
```
For an editable source checkout, include the same extra during installation:
```bash
uv pip install -e ".[swanlab]"
```
If FastVideo is already installed, you can install the compatible SDK directly:
```bash
uv pip install "swanlab>=0.6.7"
```
Authenticate once before starting a training run:
```bash
swanlab login
```
See the [SwanLab login documentation](https://docs.swanlab.cn/en/api/cli-swanlab-login.html)
for non-interactive and self-hosted setups.
## Configure Tracking
Select SwanLab in the YAML config used by the modular training framework:
```yaml
training:
checkpoint:
output_dir: outputs/my_run
tracker:
trackers: [swanlab]
project_name: my_project
run_name: my_run
```
To log to both supported services:
```yaml
training:
tracker:
trackers: [wandb, swanlab]
project_name: my_project
run_name: my_run
```
An empty or omitted `trackers` list selects W&B when `project_name` is set.
Use an explicit `none` entry to disable external tracking:
```yaml
training:
tracker:
trackers: [none]
```
## Validation Videos
SwanLab currently accepts GIF video artifacts. FastVideo converts validation
MP4 files and in-memory video arrays to GIF automatically before logging them.
For video files, FastVideo uses the sampling frame rate supplied by the caller,
or the source file's frame rate when no value is supplied. In-memory arrays use
the frame rate supplied by the caller. Both forms fall back to 16 FPS when no
frame rate is available.
For details about configuring validation callbacks, see
[Training Infrastructure](train_infra.md#callbacks-pluggable-hooks).
+49
View File
@@ -161,6 +161,21 @@ training:
decay_interval_steps: 0
```
`training.data.data_path` can also mix multiple preprocessed datasets by using a mapping from dataset path to repeat count:
```yaml
training:
data:
data_path:
data/zeldam2-clean: 1
data/multi3d_games: 2
```
The repeat count duplicates that dataset's parquet file list before shuffling/sampling, so the example above trains with roughly twice as much `multi3d_games` exposure as `zeldam2-clean`. Paths are just suggested locations; use any local path that contains a FastVideo preprocessed parquet dataset.
See [Training Trackers](trackers.md) to configure Weights & Biases or SwanLab,
including SwanLab installation and authentication.
### `callbacks` — Pluggable hooks
Callbacks run at specific points in the training loop (before/after optimizer
@@ -323,6 +338,40 @@ Self-Forcing inherits all DMD2 parameters, plus:
| `enable_gradient_in_rollout` | `true` | Enable backprop through rollout |
| `start_gradient_frame` | `0` | Frame index where gradients begin |
### Streaming Long Tuning
`StreamingLongTuningMethod` extends Self-Forcing for LongLive-style rollouts. It
keeps a streaming state, generates overlapping chunks, and trains only the new
frames while preserving context from earlier chunks.
For the MatrixGame2/Zelda world-model example, self-forcing and long tuning are
separate runs: first train or load the 1k-step self-forcing checkpoint using
`examples/train/scenario/worldmodel/zelda/self_forcing_causal_i2v.yaml`,
then run
`examples/train/scenario/worldmodel/zelda/streaming_long_tuning_causal_i2v.yaml`
from that checkpoint for the 3k-step streaming long-tuning stage.
```yaml
method:
_target_: fastvideo.train.methods.distribution_matching.streaming_long_tuning.StreamingLongTuningMethod
streaming_chunk_size: 9
streaming_max_length: 39
streaming_fixed_overlap_latents: 3
streaming_reencode_overlap_anchor: true
streaming_anchor_inject_k: 1
streaming_require_full_blocks: true
multi_phased_distill_schedule:
- stage: streaming_long
start_step: 0
end_step: 3000
num_latent_t: 39
streaming_training: true
```
See
`examples/train/scenario/worldmodel/zelda/streaming_long_tuning_causal_i2v.yaml`
for a complete MatrixGame2/Zelda configuration.
---
## Callbacks
+7 -4
View File
@@ -39,17 +39,20 @@ All you need to generate videos using multi-gpus from state-of-the-art diffusion
```python
from fastvideo import VideoGenerator
from fastvideo.api import EngineConfig, GenerationRequest, GeneratorConfig
def main():
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
engine=EngineConfig(num_gpus=1),
)
)
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones.")
video = generator.generate_video(prompt)
result = generator.generate(GenerationRequest(prompt=prompt))
if __name__ == "__main__":
main()
+23 -18
View File
@@ -1,6 +1,7 @@
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig, OutputConfig,
)
OUTPUT_PATH = "video_samples"
def main():
@@ -8,29 +9,32 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder=False,
),
),
)
)
# sampling_param = SamplingParam.from_pretrained("Wan-AI/Wan2.1-T2V-1.3B-Diffusers")
# sampling_param.num_frames = 45
# sampling_param.image_path = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/astronaut.jpg"
# Generate videos with the same simple API, regardless of GPU count
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True)
# video = generator.generate_video(prompt, sampling_param=sampling_param, output_path="wan_t2v_videos/")
video = generator.generate(
GenerationRequest(prompt=prompt, output=OutputConfig(output_path=OUTPUT_PATH, save_video=True)))
# Generate another video with a different prompt, without reloading the
# model!
@@ -40,7 +44,8 @@ def main():
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True)
video2 = generator.generate(
GenerationRequest(prompt=prompt2, output=OutputConfig(output_path=OUTPUT_PATH, save_video=True)))
if __name__ == "__main__":
+27 -19
View File
@@ -1,24 +1,30 @@
# SPDX-License-Identifier: Apache-2.0
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig,
)
def main():
# Point this to your local diffusers model dir (or replace with a HF model ID).
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
)
)
sampling_param = SamplingParam.from_pretrained(model_path)
# image2world example from official repo
image_path = "assets/images/bus_terminal.jpg"
@@ -33,13 +39,16 @@ def main():
"Overhead signage in Chinese characters remains illuminated, enhancing the vibrant, urban night scene."
)
generator.generate_video(
prompt,
sampling_param=sampling_param,
image_path=str(image_path),
num_cond_frames=1,
output_path="outputs_video/cosmos2_5_i2w.mp4",
save_video=True,
generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=str(image_path)),
output=OutputConfig(
output_path="outputs_video/cosmos2_5_i2w.mp4",
save_video=True,
),
extensions={"num_cond_frames": 1},
)
)
generator.shutdown()
@@ -47,4 +56,3 @@ def main():
if __name__ == "__main__":
main()
+25 -20
View File
@@ -1,24 +1,29 @@
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig, OutputConfig,
)
def main():
# Point this to your local diffusers model dir (or replace with a HF model ID).
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
)
)
# Load default sampling parameters (negative_prompt, resolution, steps, etc.)
sampling_param = SamplingParam.from_pretrained(model_path)
prompt = (
"A high-definition video captures the precision of robotic welding in an industrial setting. "
"The first frame showcases a robotic arm, equipped with a welding torch, positioned over a large metal structure. "
@@ -34,11 +39,14 @@ def main():
"underscoring the ongoing nature of the welding operation."
)
generator.generate_video(
prompt,
sampling_param=sampling_param,
output_path="outputs_video/cosmos2_5_t2w.mp4",
save_video=True,
generator.generate(
GenerationRequest(
prompt=prompt,
output=OutputConfig(
output_path="outputs_video/cosmos2_5_t2w.mp4",
save_video=True,
),
)
)
generator.shutdown()
@@ -46,6 +54,3 @@ def main():
if __name__ == "__main__":
main()
+28 -21
View File
@@ -1,23 +1,29 @@
# SPDX-License-Identifier: Apache-2.0
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig,
OffloadConfig, OutputConfig,
)
def main():
# Point this to your local diffusers model dir (or replace with a HF model ID).
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
)
sampling_param = SamplingParam.from_pretrained(model_path)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
))
# video2world example from official repo
video_path = "assets/videos/robot_pouring.mp4"
@@ -36,18 +42,19 @@ def main():
"The final frame captures the robotic arm with the pitcher finishing the pour, with the glass now filled to a higher level, while the pitcher is slightly tilted but still held securely by the gripper."
)
generator.generate_video(
prompt,
sampling_param=sampling_param,
video_path=str(video_path),
num_cond_frames=1,
output_path="outputs_video/cosmos2_5_v2w.mp4",
save_video=True,
)
generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(video_path=str(video_path)),
output=OutputConfig(
output_path="outputs_video/cosmos2_5_v2w.mp4",
save_video=True,
),
extensions={"num_cond_frames": 1},
))
generator.shutdown()
if __name__ == "__main__":
main()
+32 -19
View File
@@ -2,7 +2,9 @@ import os
import time
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OffloadConfig, OutputConfig, PipelineSelection,
SamplingConfig)
OUTPUT_PATH = "video_samples_dmd2"
def main():
@@ -10,30 +12,36 @@ def main():
load_start_time = time.perf_counter()
model_name = "FastVideo/FastWan2.1-T2V-1.3B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_name,
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
# Adjust these offload parameters if you have < 32GB of VRAM
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
dit_cpu_offload=False,
vae_cpu_offload=False,
VSA_sparsity=0.8,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_name,
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
# Adjust these offload parameters if you have < 32GB of VRAM
offload=OffloadConfig(
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
dit=False,
vae=False,
),
),
pipeline=PipelineSelection(experimental={"VSA_sparsity": 0.8}),
))
load_end_time = time.perf_counter()
load_time = load_end_time - load_start_time
sampling_param = SamplingParam.from_pretrained(model_name)
sampling_param.num_frames = 81
prompt = (
"A neon-lit alley in futuristic Tokyo during a heavy rainstorm at night. The puddles reflect glowing signs in kanji, advertising ramen, karaoke, and VR arcades. A woman in a translucent raincoat walks briskly with an LED umbrella. Steam rises from a street food cart, and a cat darts across the screen. Raindrops are visible on the camera lens, creating a cinematic bokeh effect."
)
start_time = time.perf_counter()
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, sampling_param=sampling_param)
video = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
end_time = time.perf_counter()
gen_time = end_time - start_time
@@ -46,7 +54,12 @@ def main():
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
start_time = time.perf_counter()
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True, num_frames=81)
video2 = generator.generate(
GenerationRequest(
prompt=prompt2,
sampling=SamplingConfig(num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
end_time = time.perf_counter()
gen_time2 = end_time - start_time
@@ -0,0 +1,79 @@
import os
from fastvideo import VideoGenerator
from fastvideo.api import (ComponentConfig, EngineConfig, GenerationRequest,
GeneratorConfig, InputConfig, OffloadConfig,
OutputConfig, PipelineSelection, SamplingConfig)
OUTPUT_PATH = os.getenv("DREAMX_WORLD_OUTPUT_PATH", "video_samples_dreamx_world")
def _env_int(name: str, default: int) -> int:
return int(os.getenv(name, str(default)))
def _env_float(name: str, default: float) -> float:
return float(os.getenv(name, str(default)))
def main():
model_name = os.getenv("DREAMX_WORLD_MODEL_DIR", "FastVideo/DreamX-World-5B-Cam-Diffusers")
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_name,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(override_pipeline_cls_name="DreamXWorldPipeline"), ),
))
prompt = os.getenv(
"DREAMX_WORLD_PROMPT",
"A cinematic first-person drive through a futuristic coastal city at "
"sunrise, reflective glass towers, clean streets, soft volumetric light.",
)
image_path = os.getenv(
"DREAMX_WORLD_IMAGE_PATH",
"https://huggingface.co/datasets/YiYiXu/testing-images/resolve/main/wan_i2v_input.JPG",
)
request = GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=image_path or None),
sampling=SamplingConfig(
height=_env_int("DREAMX_WORLD_HEIGHT", 480),
width=_env_int("DREAMX_WORLD_WIDTH", 832),
num_frames=_env_int("DREAMX_WORLD_NUM_FRAMES", 161),
num_inference_steps=_env_int("DREAMX_WORLD_STEPS", 30),
guidance_scale=_env_float("DREAMX_WORLD_GUIDANCE", 5.0),
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=os.getenv("DREAMX_WORLD_SAVE_VIDEO", "1") != "0",
),
extensions={
"action_list": os.getenv("DREAMX_WORLD_ACTIONS", "w,d,w").split(","),
"action_speed_list": [
float(value)
for value in os.getenv("DREAMX_WORLD_ACTION_SPEEDS", "4.0,2.0,4.0").split(",")
],
},
)
try:
generator.generate(request)
finally:
generator.shutdown()
if __name__ == "__main__":
main()
+40 -21
View File
@@ -26,6 +26,15 @@ import os
import torch
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
InputConfig,
OffloadConfig,
OutputConfig,
SamplingConfig,
)
from fastvideo.models.camera import create_camera_trajectory
# Model configuration (use GAMECRAFT_MODEL_PATH for local weights)
@@ -55,14 +64,20 @@ OUTPUT_PATH = "video_samples_gamecraft"
def main():
# Initialize generator
# FastVideo will automatically download weights from HuggingFace
generator = VideoGenerator.from_pretrained(
MODEL_PATH,
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=MODEL_PATH,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=True,
),
),
)
)
# Video parameters
@@ -96,23 +111,27 @@ def main():
prompt = DEFAULT_I2V_PROMPT if is_i2v else DEFAULT_PROMPTS["temple"]
print(f"Mode: {'I2V' if is_i2v else 'T2V'}, prompt: {prompt[:60]}...")
gen_kw = dict(
request = GenerationRequest(
prompt=prompt,
negative_prompt="",
camera_states=camera_states,
height=height,
width=width,
num_frames=num_frames,
num_inference_steps=50,
guidance_scale=6.0,
seed=42,
fps=24,
output_path=OUTPUT_PATH,
save_video=True,
sampling=SamplingConfig(
height=height,
width=width,
num_frames=num_frames,
num_inference_steps=50,
guidance_scale=6.0,
seed=42,
fps=24,
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
extensions={"camera_states": camera_states},
)
if is_i2v:
gen_kw["image_path"] = image_path
generator.generate_video(**gen_kw)
request.inputs = InputConfig(image_path=image_path)
generator.generate(request)
if __name__ == "__main__":
+44 -26
View File
@@ -22,6 +22,10 @@ Requirements:
import argparse
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig,
OffloadConfig, OutputConfig, SamplingConfig,
)
def main():
@@ -74,33 +78,47 @@ def main():
parser.add_argument("--seed", type=int, default=42)
args = parser.parse_args()
generator = VideoGenerator.from_pretrained(
args.model_path,
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=True,
),
),
))
video = generator.generate_video(
args.prompt,
negative_prompt=args.negative_prompt,
image_path=args.image_path,
trajectory_type=args.trajectory,
movement_distance=args.movement_distance,
camera_rotation=args.camera_rotation,
height=args.height,
width=args.width,
num_frames=args.num_frames,
num_inference_steps=args.num_inference_steps,
guidance_scale=args.guidance_scale,
fps=24,
seed=args.seed,
output_path=args.output_path,
save_video=True,
)
video = generator.generate(
GenerationRequest(
prompt=args.prompt,
negative_prompt=args.negative_prompt,
inputs=InputConfig(
image_path=args.image_path,
),
sampling=SamplingConfig(
height=args.height,
width=args.width,
num_frames=args.num_frames,
num_inference_steps=args.num_inference_steps,
guidance_scale=args.guidance_scale,
fps=24,
seed=args.seed,
),
output=OutputConfig(
output_path=args.output_path,
save_video=True,
),
extensions={
"trajectory_type": args.trajectory,
"movement_distance": args.movement_distance,
"camera_rotation": args.camera_rotation,
},
))
generator.shutdown()
+35 -14
View File
@@ -1,6 +1,13 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
SamplingConfig,
)
import json
# from fastvideo.api.sampling_param import SamplingParam
OUTPUT_PATH = "video_samples_hy15"
def main():
@@ -8,17 +15,21 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-480p_t2v",
generator = VideoGenerator.from_config(GeneratorConfig(
model_path="hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-480p_t2v",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder_cpu_offload=False,
)
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder=False,
),
),
))
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
@@ -26,7 +37,12 @@ def main():
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, negative_prompt="", num_frames=81, fps=16)
generator.generate(GenerationRequest(
prompt=prompt,
negative_prompt="",
sampling=SamplingConfig(num_frames=81, fps=16),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
prompt2 = (
"A majestic lion strides across the golden savanna, its powerful frame "
@@ -35,8 +51,13 @@ def main():
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True, negative_prompt="", num_frames=81, fps=16)
generator.generate(GenerationRequest(
prompt=prompt2,
negative_prompt="",
sampling=SamplingConfig(num_frames=81, fps=16),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
if __name__ == "__main__":
main()
main()
+38 -14
View File
@@ -1,6 +1,12 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
)
import json
# from fastvideo.api.sampling_param import SamplingParam
OUTPUT_PATH = "video_samples_hy15_1080p"
def main():
@@ -8,17 +14,23 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"weizhou03/HunyuanVideo-1.5-Diffusers-1080p-2SR", # 480p -> 720p -> 1080p
# or "weizhou03/HunyuanVideo-1.5-Diffusers-1080p" # 720p -> 1080p
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="weizhou03/HunyuanVideo-1.5-Diffusers-1080p-2SR", # 480p -> 720p -> 1080p
# or "weizhou03/HunyuanVideo-1.5-Diffusers-1080p" # 720p -> 1080p
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder=False,
),
),
)
)
prompt = (
@@ -27,7 +39,13 @@ def main():
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, negative_prompt="")
video = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt="",
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
prompt2 = (
"A majestic lion strides across the golden savanna, its powerful frame "
@@ -36,7 +54,13 @@ def main():
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True, negative_prompt="")
video2 = generator.generate(
GenerationRequest(
prompt=prompt2,
negative_prompt="",
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
+37 -23
View File
@@ -1,4 +1,6 @@
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig,
SamplingConfig)
from fastvideo.models.dits.hyworld.resolution_utils import get_resolution_from_image
# Default prompt from HY-WorldPlay run.sh
@@ -31,33 +33,45 @@ def main():
# Initialize generator
print("\nInitializing VideoGenerator for HYWorld...")
generator = VideoGenerator.from_pretrained(
"FastVideo/HY-WorldPlay-Bidirectional-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
image_encoder_cpu_offload=True,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/HY-WorldPlay-Bidirectional-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=True,
image_encoder=True,
),
),
))
# Generate video
# The pose string is automatically converted to camera matrices by the pipeline
print("\nGenerating video...")
generator.generate_video(
prompt=args.prompt,
image_path=args.image,
pose=args.pose, # Camera trajectory control
output_path=args.output_path,
save_video=True,
negative_prompt="",
num_frames=args.num_frames,
fps=24,
height=HEIGHT,
width=WIDTH,
seed=args.seed,
)
generator.generate(
GenerationRequest(
prompt=args.prompt,
negative_prompt="",
inputs=InputConfig(
image_path=args.image,
pose=args.pose, # Camera trajectory control
),
sampling=SamplingConfig(
num_frames=args.num_frames,
fps=24,
height=HEIGHT,
width=WIDTH,
seed=args.seed,
),
output=OutputConfig(
output_path=args.output_path,
save_video=True,
),
))
print(f"\nVideo saved to: {args.output_path}")
@@ -0,0 +1,44 @@
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
InputConfig, OffloadConfig, OutputConfig,
SamplingConfig)
OUTPUT_PATH = "video_samples_kandinsky5_i2v"
IMAGE_PATH = "assets/girl.png"
def main():
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="kandinskylab/Kandinsky-5.0-I2V-Pro-distilled-5s-Diffusers",
# "kandinskylab/Kandinsky-5.0-I2V-Pro-sft-5s-Diffusers"
# "kandinskylab/Kandinsky-5.0-I2V-Lite-5s-Diffusers"
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
# image_encoder=False,
),
),
))
prompt = (
"A woman stands up and walks away"
)
_ = generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=IMAGE_PATH),
sampling=SamplingConfig(height=1024, width=1024, num_frames=121),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
if __name__ == "__main__":
main()
@@ -0,0 +1,55 @@
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OffloadConfig, OutputConfig, SamplingConfig)
OUTPUT_PATH = "video_samples_kandinsky5_t2v"
def main():
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="kandinskylab/Kandinsky-5.0-T2V-Lite-sft-5s-Diffusers",
# "kandinskylab/Kandinsky-5.0-T2V-Pro-sft-5s-Diffusers"
# "kandinskylab/Kandinsky-5.0-T2V-Lite-distilled16steps-5s-Diffusers"
# "kandinskylab/Kandinsky-5.0-T2V-Pro-distilled-5s-Diffusers"
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
# image_encoder=False,
),
),
))
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
_ = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(height=512, width=768, num_frames=121),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
prompt2 = (
"A majestic lion strides across the golden savanna, its powerful frame "
"glistening under the warm afternoon sun. The tall grass ripples gently in "
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
_ = generator.generate(
GenerationRequest(
prompt=prompt2,
sampling=SamplingConfig(height=512, width=768, num_frames=121),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
if __name__ == "__main__":
main()
@@ -1,24 +1,31 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig, SamplingConfig,
)
from fastvideo.models.dits.lingbotworld.cam_utils import prepare_camera_embedding
# from fastvideo.api.sampling_param import SamplingParam
OUTPUT_PATH = "video_samples_lingbotworld"
def main():
# FastVideo will automatically use the optimal default arguments for the
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"FastVideo/LingBot-World-Base-Cam-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LingBot-World-Base-Cam-Diffusers",
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
),
),
)
)
num_frames = 81
@@ -33,15 +40,23 @@ def main():
spatial_scale=8,
)
generator.generate_video(
prompt,
image_path=image_path,
output_path=OUTPUT_PATH,
save_video=True,
num_frames=num_frames,
height=480,
width=832,
c2ws_plucker_emb=c2ws_plucker_emb,
generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(
image_path=image_path,
c2ws_plucker_emb=c2ws_plucker_emb,
),
sampling=SamplingConfig(
num_frames=num_frames,
height=480,
width=832,
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
)
+140 -95
View File
@@ -19,6 +19,10 @@ import glob
import os
from fastvideo import VideoGenerator
from fastvideo.api import (
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig,
InputConfig, OffloadConfig, OutputConfig, PipelineSelection, SamplingConfig,
)
# Common prompts and settings matching the shell script examples
PROMPT = (
@@ -45,41 +49,50 @@ SEED = 42
def basic_generation():
"""
Run basic LongCat I2V generation (50 steps at 480p).
This uses the full 50-step denoising process for highest quality.
"""
print("=" * 60)
print("LongCat I2V: Basic Generation (50 steps, 480p)")
print("=" * 60)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-I2V-Diffusers",
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-I2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(experimental={"enable_bsa": False}),
)
)
output_path = "outputs_video/longcat_i2v_basic"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
image_path=IMAGE_PATH,
output_path=output_path,
save_video=True,
height=480,
width=480, # Square
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(image_path=IMAGE_PATH),
sampling=SamplingConfig(
height=480,
width=480, # Square
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
),
output=OutputConfig(output_path=output_path, save_video=True),
)
)
print(f"\nBasic generation complete! Video saved to: {output_path}")
generator.shutdown()
@@ -87,55 +100,70 @@ def basic_generation():
def distill_refine_generation():
"""
Run LongCat I2V with distill+refine pipeline (16 steps + refinement to 768p).
This uses the distilled LoRA for fast 480p generation (16 steps),
then refines to 768p using the refinement LoRA with BSA enabled.
"""
print("\n" + "=" * 60)
print("LongCat I2V: Distill + Refine Pipeline")
print("=" * 60)
# Stage 1: Distilled generation (16 steps at 480p)
print("\n[Stage 1] Distilled generation (16 steps, 480p)")
print("-" * 40)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-I2V-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
lora_nickname="distilled",
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-I2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
),
experimental={
"enable_bsa": False,
"lora_nickname": "distilled",
},
),
)
)
distill_output_path = "outputs_video/longcat_i2v_distill"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
image_path=IMAGE_PATH,
output_path=distill_output_path,
save_video=True,
height=480,
width=480, # Square
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(image_path=IMAGE_PATH),
sampling=SamplingConfig(
height=480,
width=480, # Square
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
),
output=OutputConfig(output_path=distill_output_path, save_video=True),
)
)
print(f"Distilled generation complete! Video saved to: {distill_output_path}")
generator.shutdown()
# Stage 2: Refinement (480p -> 768p)
print("\n[Stage 2] Refinement (480p -> 768p with BSA)")
print("-" * 40)
# Find the actual saved video file from stage 1
video_files = glob.glob(os.path.join(distill_output_path, "*.mp4"))
if not video_files:
@@ -143,46 +171,63 @@ def distill_refine_generation():
# Use the most recently created video file
distill_video_path = max(video_files, key=os.path.getmtime)
print(f"Using stage 1 video: {distill_video_path}")
# Create a new generator with refinement LoRA and BSA enabled
# Note: Refinement uses the T2V model (not I2V) since it's upscaling the generated video
# For BSA [4, 4, 8]: latent must be divisible by 8
# 768x768: latent 48x48, 48%8=0 ✓
refine_generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-T2V-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=True,
bsa_sparsity=0.875,
bsa_chunk_q=[4, 4, 4],
bsa_chunk_k=[4, 4, 4],
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
lora_nickname="refinement",
refine_generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-T2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
),
experimental={
"enable_bsa": True,
"bsa_sparsity": 0.875,
"bsa_chunk_q": [4, 4, 4],
"bsa_chunk_k": [4, 4, 4],
"lora_nickname": "refinement",
},
),
)
)
refine_output_path = "outputs_video/longcat_i2v_refine_720p"
refine_generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output_path=refine_output_path,
save_video=True,
refine_from=distill_video_path,
t_thresh=0.5,
spatial_refine_only=False,
num_cond_frames=0,
height=720,
width=720,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
refine_generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(refine_from=distill_video_path),
sampling=SamplingConfig(
height=720,
width=720,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
),
output=OutputConfig(output_path=refine_output_path, save_video=True),
extensions={
"t_thresh": 0.5,
"spatial_refine_only": False,
"num_cond_frames": 0,
},
)
)
print(f"Refinement complete! Video saved to: {refine_output_path}")
refine_generator.shutdown()
@@ -192,13 +237,13 @@ def main():
print("\n" + "=" * 60)
print("LongCat Image-to-Video Example")
print("=" * 60 + "\n")
# Run basic generation
basic_generation()
# Run distill+refine pipeline
distill_refine_generation()
print("\n" + "=" * 60)
print("All generations complete!")
print("=" * 60)
+151 -95
View File
@@ -13,6 +13,10 @@ import glob
import os
from fastvideo import VideoGenerator
from fastvideo.api import (
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig,
InputConfig, OffloadConfig, OutputConfig, PipelineSelection, SamplingConfig,
)
# Common prompts and settings matching the shell script examples
PROMPT = (
@@ -38,40 +42,54 @@ SEED = 42
def basic_generation():
"""
Run basic LongCat T2V generation (50 steps at 480p).
This uses the full 50-step denoising process for highest quality.
"""
print("=" * 60)
print("LongCat T2V: Basic Generation (50 steps, 480p)")
print("=" * 60)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-T2V-Diffusers",
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-T2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
experimental={"enable_bsa": False},
),
)
)
output_path = "outputs_video/longcat_t2v_basic"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output_path=output_path,
save_video=True,
height=480,
width=832,
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output=OutputConfig(
output_path=output_path,
save_video=True,
),
sampling=SamplingConfig(
height=480,
width=832,
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
),
)
)
print(f"\nBasic generation complete! Video saved to: {output_path}")
generator.shutdown()
@@ -79,54 +97,72 @@ def basic_generation():
def distill_refine_generation():
"""
Run LongCat T2V with distill+refine pipeline (16 steps + refinement to 720p).
This uses the distilled LoRA for fast 480p generation (16 steps),
then refines to 720p using the refinement LoRA with BSA enabled.
"""
print("\n" + "=" * 60)
print("LongCat T2V: Distill + Refine Pipeline")
print("=" * 60)
# Stage 1: Distilled generation (16 steps at 480p)
print("\n[Stage 1] Distilled generation (16 steps, 480p)")
print("-" * 40)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-T2V-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
lora_nickname="distilled",
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-T2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
),
experimental={
"enable_bsa": False,
"lora_nickname": "distilled",
},
),
)
)
distill_output_path = "outputs_video/longcat_t2v_distill"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output_path=distill_output_path,
save_video=True,
height=480,
width=832,
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output=OutputConfig(
output_path=distill_output_path,
save_video=True,
),
sampling=SamplingConfig(
height=480,
width=832,
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
),
)
)
print(f"Distilled generation complete! Video saved to: {distill_output_path}")
generator.shutdown()
# Stage 2: Refinement (480p -> 720p)
print("\n[Stage 2] Refinement (480p -> 720p with BSA)")
print("-" * 40)
# Find the actual saved video file from stage 1
video_files = glob.glob(os.path.join(distill_output_path, "*.mp4"))
if not video_files:
@@ -134,43 +170,65 @@ def distill_refine_generation():
# Use the most recently created video file
distill_video_path = max(video_files, key=os.path.getmtime)
print(f"Using stage 1 video: {distill_video_path}")
# Create a new generator with refinement LoRA and BSA enabled
refine_generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-T2V-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=True,
bsa_sparsity=0.875,
bsa_chunk_q=[4, 4, 8],
bsa_chunk_k=[4, 4, 8],
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
lora_nickname="refinement",
refine_generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-T2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
),
experimental={
"enable_bsa": True,
"bsa_sparsity": 0.875,
"bsa_chunk_q": [4, 4, 8],
"bsa_chunk_k": [4, 4, 8],
"lora_nickname": "refinement",
},
),
)
)
refine_output_path = "outputs_video/longcat_t2v_refine_720p"
refine_generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output_path=refine_output_path,
save_video=True,
refine_from=distill_video_path,
t_thresh=0.5,
spatial_refine_only=False,
num_cond_frames=0,
height=720,
width=1280,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
refine_generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output=OutputConfig(
output_path=refine_output_path,
save_video=True,
),
inputs=InputConfig(
refine_from=distill_video_path,
),
sampling=SamplingConfig(
height=720,
width=1280,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
),
extensions={
"t_thresh": 0.5,
"spatial_refine_only": False,
"num_cond_frames": 0,
},
)
)
print(f"Refinement complete! Video saved to: {refine_output_path}")
refine_generator.shutdown()
@@ -180,13 +238,13 @@ def main():
print("\n" + "=" * 60)
print("LongCat Text-to-Video Example")
print("=" * 60 + "\n")
# Run basic generation
basic_generation()
# Run distill+refine pipeline
distill_refine_generation()
print("\n" + "=" * 60)
print("All generations complete!")
print("=" * 60)
@@ -194,5 +252,3 @@ def main():
if __name__ == "__main__":
main()
+142 -86
View File
@@ -19,6 +19,10 @@ import glob
import os
from fastvideo import VideoGenerator
from fastvideo.api import (
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
# Common prompts and settings matching the shell script examples
PROMPT = (
@@ -63,35 +67,49 @@ def basic_generation():
"Please provide a valid video path."
)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-VC-Diffusers",
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-VC-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
experimental={"enable_bsa": False},
),
)
)
output_path = "outputs_video/longcat_vc_basic"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
video_path=VIDEO_PATH,
num_cond_frames=NUM_COND_FRAMES,
output_path=output_path,
save_video=True,
height=480,
width=832,
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(video_path=VIDEO_PATH),
sampling=SamplingConfig(
height=480,
width=832,
num_frames=93,
num_inference_steps=50,
fps=15,
guidance_scale=4.0,
seed=SEED,
),
output=OutputConfig(
output_path=output_path,
save_video=True,
),
extensions={"num_cond_frames": NUM_COND_FRAMES},
)
)
print(f"\nBasic generation complete! Video saved to: {output_path}")
generator.shutdown()
@@ -118,37 +136,55 @@ def distill_refine_generation():
print("\n[Stage 1] Distilled generation (16 steps, 480p)")
print("-" * 40)
generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-VC-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=False,
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
lora_nickname="distilled",
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-VC-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Distilled-LoRA",
),
experimental={
"enable_bsa": False,
"lora_nickname": "distilled",
},
),
)
)
distill_output_path = "outputs_video/longcat_vc_distill"
generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
video_path=VIDEO_PATH,
num_cond_frames=NUM_COND_FRAMES,
output_path=distill_output_path,
save_video=True,
height=480,
width=832,
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(video_path=VIDEO_PATH),
sampling=SamplingConfig(
height=480,
width=832,
num_frames=93,
num_inference_steps=16,
fps=15,
guidance_scale=1.0,
seed=SEED,
),
output=OutputConfig(
output_path=distill_output_path,
save_video=True,
),
extensions={"num_cond_frames": NUM_COND_FRAMES},
)
)
print(f"Distilled generation complete! Video saved to: {distill_output_path}")
generator.shutdown()
@@ -166,41 +202,61 @@ def distill_refine_generation():
# Create a new generator with refinement LoRA and BSA enabled
# Note: Refinement uses the T2V model (not VC) since it's upscaling the generated video
refine_generator = VideoGenerator.from_pretrained(
"FastVideo/LongCat-Video-T2V-Diffusers",
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=False,
enable_bsa=True,
bsa_sparsity=0.875,
bsa_chunk_q=[4, 4, 8],
bsa_chunk_k=[4, 4, 8],
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
lora_nickname="refinement",
refine_generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LongCat-Video-T2V-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="FastVideo/LongCat-Video-T2V-Refinement-LoRA",
),
experimental={
"enable_bsa": True,
"bsa_sparsity": 0.875,
"bsa_chunk_q": [4, 4, 8],
"bsa_chunk_k": [4, 4, 8],
"lora_nickname": "refinement",
},
),
)
)
refine_output_path = "outputs_video/longcat_vc_refine_720p"
refine_generator.generate_video(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
output_path=refine_output_path,
save_video=True,
refine_from=distill_video_path,
t_thresh=0.5,
spatial_refine_only=False,
num_cond_frames=0, # For refinement, no conditioning frames
height=720,
width=1280,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
refine_generator.generate(
GenerationRequest(
prompt=PROMPT,
negative_prompt=NEGATIVE_PROMPT,
inputs=InputConfig(refine_from=distill_video_path),
sampling=SamplingConfig(
height=720,
width=1280,
num_inference_steps=50,
fps=30,
guidance_scale=1.0,
seed=SEED,
),
output=OutputConfig(
output_path=refine_output_path,
save_video=True,
),
extensions={
"t_thresh": 0.5,
"spatial_refine_only": False,
"num_cond_frames": 0, # For refinement, no conditioning frames
},
)
)
print(f"Refinement complete! Video saved to: {refine_output_path}")
refine_generator.shutdown()
+28 -11
View File
@@ -1,4 +1,11 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OutputConfig,
SamplingConfig,
)
PROMPT = (
@@ -18,22 +25,32 @@ PROMPT = (
def main() -> None:
# Uses FastVideo default sampling settings for LTX2 base.
generator = VideoGenerator.from_pretrained(
"Davids048/LTX2-Base-Diffusers",
num_gpus=1,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Davids048/LTX2-Base-Diffusers",
engine=EngineConfig(
num_gpus=1,
),
)
)
output_path = "outputs_video/ltx2_basic/output_ltx2_base_t2v_1088_1920_1.1.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
num_frames=121,
height=1088,
width=1920,
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(
output_path=output_path,
save_video=True,
),
sampling=SamplingConfig(
num_frames=121,
height=1088,
width=1920,
),
)
)
generator.shutdown()
if __name__ == "__main__":
main()
main()
@@ -49,6 +49,11 @@ from pathlib import Path
import torch._inductor.config as _inductor
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig, ComponentConfig, EngineConfig, GenerationRequest,
GeneratorConfig, OffloadConfig, OutputConfig, PipelineSelection,
SamplingConfig,
)
from fastvideo.configs.pipelines.base import PipelineConfig
from fastvideo.utils import maybe_download_model
@@ -86,9 +91,9 @@ PROMPT = os.getenv("LTX23_I2V_PROMPT", DEFAULT_PROMPT)
# Per-stage timing helpers --------------------------------------------------
def _print_stage_breakdown(result: dict, label: str) -> float | None:
def _print_stage_breakdown(result, label: str) -> float | None:
"""Print stage execution times and return the sum, or None if missing."""
logging_info = result.get("logging_info")
logging_info = result.logging_info
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
print(f" [{label}] stage breakdown unavailable")
@@ -104,11 +109,11 @@ def _print_stage_breakdown(result: dict, label: str) -> float | None:
def _collect_stage_times(
result: dict,
result,
stage_times: dict[str, list[float]],
stage_order: OrderedDict[str, None],
) -> None:
logging_info = result.get("logging_info")
logging_info = result.logging_info
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
return
@@ -169,34 +174,54 @@ def main() -> None:
pipeline_config = PipelineConfig.from_pretrained(model_root)
pipeline_config.dit_config.quant_config = None
generator = VideoGenerator.from_pretrained(
model_root,
num_gpus=1,
# LTX-2.3 distilled uses the two-stage refine pipeline; the refine
# LoRA is intentionally empty for the distilled student.
ltx2_refine_enabled=True,
ltx2_refine_upsampler_path=str(refine_upsampler_path),
ltx2_refine_lora_path="",
ltx2_refine_num_inference_steps=3,
ltx2_refine_guidance_scale=1.0,
ltx2_refine_add_noise=True,
pipeline_config=pipeline_config,
enable_torch_compile=True,
enable_torch_compile_text_encoder=True,
# Compile the VAE codec submodules (encoder / decoder) too. The
# `LTX2CausalVideoAutoencoder` declares `_compile_conditions` so
# `_compile_with_conditions` targets just those submodules and
# leaves the surrounding tiling control flow eager — needed for
# fullgraph + dynamic=False to succeed. VAE eager decode is
# ~1.0s; compiling it brings the stage to ~0.3s.
enable_torch_compile_vae=True,
torch_compile_kwargs=torch_compile_kwargs,
torch_compile_kwargs_vae=torch_compile_kwargs,
# Keep everything resident — no CPU offload for serving-style runs.
dit_cpu_offload=False,
text_encoder_cpu_offload=False,
vae_cpu_offload=False,
ltx2_vae_tiling=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_root,
engine=EngineConfig(
num_gpus=1,
compile=CompileConfig(
enabled=True,
text_encoder_enabled=True,
# Compile the VAE codec submodules (encoder / decoder)
# too. The `LTX2CausalVideoAutoencoder` declares
# `_compile_conditions` so `_compile_with_conditions`
# targets just those submodules and leaves the
# surrounding tiling control flow eager — needed for
# fullgraph + dynamic=False to succeed. VAE eager decode
# is ~1.0s; compiling it brings the stage to ~0.3s.
vae_enabled=True,
backend=torch_compile_kwargs["backend"],
fullgraph=torch_compile_kwargs["fullgraph"],
mode=torch_compile_kwargs["mode"],
dynamic=torch_compile_kwargs["dynamic"],
vae_kwargs=torch_compile_kwargs,
),
# Keep everything resident — no CPU offload for serving runs.
offload=OffloadConfig(
dit=False,
text_encoder=False,
vae=False,
),
),
pipeline=PipelineSelection(
vae_tiling=False,
# LTX-2.3 distilled uses the two-stage refine pipeline; the
# refine LoRA is intentionally empty for the distilled
# student.
components=ComponentConfig(
upsampler_weights=str(refine_upsampler_path),
),
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 3,
"guidance_scale": 1.0,
"add_noise": True,
}
},
experimental={"pipeline_config": pipeline_config},
),
)
)
common_kwargs = dict(
@@ -206,12 +231,15 @@ def main() -> None:
height=1280, width=832, # portrait runway aspect
num_frames=121, fps=24, # ~5s clip
num_inference_steps=8, # distilled denoise steps
# i2v: anchor the input image at frame 0 with full strength.
# `ltx2_image_crf=0.0` skips an extra JPEG re-encode of an already
# JPEG conditioning image.
)
# i2v: anchor the input image at frame 0 with full strength.
# `ltx2_image_crf=0.0` skips an extra JPEG re-encode of an already
# JPEG conditioning image. These are model-specific knobs routed through
# the request extensions escape hatch.
common_extensions = dict(
ltx2_images=[(I2V_IMAGE, 0, 1.0)],
ltx2_image_crf=0.0,
save_video=True,
)
warmup_runs = 2
@@ -227,10 +255,25 @@ def main() -> None:
for w in range(warmup_runs):
t0 = time.perf_counter()
print(f"\n[warmup {w + 1}/{warmup_runs}] compiling + generating…")
generator.generate_video(
output_path=str(OUTPUT_DIR / f"_warmup_{w + 1}.mp4"),
seed=7,
**common_kwargs,
generator.generate(
GenerationRequest(
prompt=common_kwargs["prompt"],
negative_prompt=common_kwargs["negative_prompt"],
sampling=SamplingConfig(
guidance_scale=common_kwargs["guidance_scale"],
height=common_kwargs["height"],
width=common_kwargs["width"],
num_frames=common_kwargs["num_frames"],
fps=common_kwargs["fps"],
num_inference_steps=common_kwargs["num_inference_steps"],
seed=7,
),
output=OutputConfig(
output_path=str(OUTPUT_DIR / f"_warmup_{w + 1}.mp4"),
save_video=True,
),
extensions=common_extensions,
)
)
dt = time.perf_counter() - t0
warmup_secs.append(dt)
@@ -245,19 +288,31 @@ def main() -> None:
out_path = OUTPUT_DIR / f"output_ltx2_3_distilled_i2v_run_{m + 1}.mp4"
print(f"\n[measured {m + 1}/{measured_runs}] generating: {out_path}")
t0 = time.perf_counter()
result = generator.generate_video(
output_path=str(out_path),
seed=2002 + m,
**common_kwargs,
result = generator.generate(
GenerationRequest(
prompt=common_kwargs["prompt"],
negative_prompt=common_kwargs["negative_prompt"],
sampling=SamplingConfig(
guidance_scale=common_kwargs["guidance_scale"],
height=common_kwargs["height"],
width=common_kwargs["width"],
num_frames=common_kwargs["num_frames"],
fps=common_kwargs["fps"],
num_inference_steps=common_kwargs["num_inference_steps"],
seed=2002 + m,
),
output=OutputConfig(
output_path=str(out_path),
save_video=True,
),
extensions=common_extensions,
)
)
wall = time.perf_counter() - t0
e2e = (
result.get("e2e_latency")
if isinstance(result, dict) else None
) or wall
e2e = (result.extra.get("e2e_latency") if result is not None else None) or wall
measured_secs.append(e2e)
print(f"[measured {m + 1}/{measured_runs}] e2e={e2e:.2f}s wall={wall:.2f}s")
if isinstance(result, dict):
if result is not None:
_print_stage_breakdown(result, f"measured {m + 1}")
_collect_stage_times(result, stage_times, stage_order)
@@ -1,4 +1,5 @@
from fastvideo import VideoGenerator
from fastvideo.api import EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig
PROMPT = (
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
@@ -17,16 +18,19 @@ import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/LTX2-Distilled-Diffusers",
num_gpus=4,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/LTX2-Distilled-Diffusers",
engine=EngineConfig(num_gpus=4),
)
)
output_path = "outputs_video/ltx2_basic/output_ltx2_distilled_t2v.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(output_path=output_path, save_video=True),
)
)
generator.shutdown()
@@ -8,6 +8,11 @@ from pathlib import Path
import torch
import torch._inductor.config
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig, ComponentConfig, EngineConfig, GenerationRequest,
GenerationResult, GeneratorConfig, OffloadConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
from fastvideo.configs.pipelines.base import PipelineConfig
from fastvideo.layers.quantization.nvfp4_config import NVFP4Config
from fastvideo.utils import maybe_download_model
@@ -45,11 +50,11 @@ def load_validation_entries(path: Path) -> list[dict]:
def print_stage_breakdown(
result: dict,
result: GenerationResult,
run_idx: int,
num_runs: int,
) -> float | None:
logging_info = result.get("logging_info")
logging_info = result.logging_info
if logging_info is None:
print(f"[{run_idx}/{num_runs}] Stage breakdown unavailable: no logging_info")
return None
@@ -70,9 +75,9 @@ def print_stage_breakdown(
def extract_sr_forward_latency(
result: dict,
result: GenerationResult,
) -> tuple[float | None, list[tuple[str, float]], list[str]]:
logging_info = result.get("logging_info")
logging_info = result.logging_info
if logging_info is None:
return None, [], []
@@ -106,11 +111,11 @@ def extract_sr_forward_latency(
def collect_stage_times(
result: dict,
result: GenerationResult,
stage_times: dict[str, list[float]],
stage_order: OrderedDict[str, None],
) -> None:
logging_info = result.get("logging_info")
logging_info = result.logging_info
if logging_info is None:
return
stages = getattr(logging_info, "stages", None)
@@ -202,26 +207,45 @@ def main() -> None:
"dynamic": False,
}
generator = VideoGenerator.from_pretrained(
model_root,
num_gpus=1,
ltx2_refine_enabled=True,
ltx2_refine_upsampler_path=str(refine_upsampler_path),
refine_lora_path="", # keep refine LoRA disabled in this repo's typed adapter
ltx2_refine_lora_path="", # keep refine LoRA disabled for distilled model
ltx2_refine_num_inference_steps=2,
ltx2_refine_guidance_scale=1.0,
ltx2_refine_add_noise=True,
pipeline_config=pipeline_config,
enable_torch_compile=True,
enable_torch_compile_text_encoder=True,
enable_torch_compile_vae=True,
torch_compile_kwargs=torch_compile_kwargs,
torch_compile_kwargs_vae=torch_compile_kwargs,
dit_cpu_offload=False,
text_encoder_cpu_offload=False,
vae_cpu_offload=False,
ltx2_vae_tiling=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_root,
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
text_encoder=False,
vae=False,
),
compile=CompileConfig(
enabled=True,
text_encoder_enabled=True,
vae_enabled=True,
backend="inductor",
fullgraph=True,
dynamic=False,
vae_kwargs=torch_compile_kwargs,
),
),
pipeline=PipelineSelection(
vae_tiling=False,
components=ComponentConfig(
upsampler_weights=str(refine_upsampler_path),
),
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 2,
"guidance_scale": 1.0,
"add_noise": True,
}
},
experimental={
"refine_lora_path": "", # keep refine LoRA disabled in this repo's typed adapter
"pipeline_config": pipeline_config,
},
),
)
)
run_times: list[float] = []
@@ -243,25 +267,31 @@ def main() -> None:
torch.cuda.synchronize()
start = time.perf_counter()
result = generator.generate_video(
prompt=prompt,
output_path=str(output_path),
fps=24,
seed=10,
save_video=True,
guidance_scale=1.0,
height=benchmark_entry.get("height", 1088),
width=benchmark_entry.get("width", 1920),
num_frames=121,
num_inference_steps=5,
# image_path="examples/inference/basic/prompt1.png",
# ltx2_image_crf=0.0
result = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(
fps=24,
seed=10,
guidance_scale=1.0,
height=benchmark_entry.get("height", 1088),
width=benchmark_entry.get("width", 1920),
num_frames=121,
num_inference_steps=5,
),
output=OutputConfig(
output_path=str(output_path),
save_video=True,
),
# inputs=InputConfig(image_path="examples/inference/basic/prompt1.png"),
# extensions={"ltx2_image_crf": 0.0},
)
)
if os.environ.get("FASTVIDEO_STAGE_LOGGING") == "0":
torch.cuda.synchronize()
elapsed = result.get("generation_time") if isinstance(result, dict) else None
e2e_elapsed = result.get("e2e_latency") if isinstance(result, dict) else None
elapsed = result.generation_time if isinstance(result, GenerationResult) else None
e2e_elapsed = result.extra.get("e2e_latency") if isinstance(result, GenerationResult) else None
if elapsed is None:
elapsed = time.perf_counter() - start
if e2e_elapsed is None:
@@ -272,7 +302,7 @@ def main() -> None:
print(f"[{i + 1}/{num_runs}] Generation time: {elapsed:.2f}s")
print(f"[{i + 1}/{num_runs}] End-to-end latency: {e2e_elapsed:.2f}s")
if isinstance(result, dict):
if isinstance(result, GenerationResult):
stage_sum = print_stage_breakdown(result, i + 1, num_runs)
if stage_sum is not None:
non_stage_overhead = e2e_elapsed - stage_sum
+32 -21
View File
@@ -1,18 +1,27 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig,
OffloadConfig, OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples_lucy_edit"
def main():
generator = VideoGenerator.from_pretrained(
"decart-ai/Lucy-Edit-Dev",
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=True,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="decart-ai/Lucy-Edit-Dev",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=True,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
))
prompt = ("Change the apron and blouse to a classic clown costume: satin "
"polka-dot jumpsuit in bright primary colors, ruffled white collar, "
@@ -20,18 +29,20 @@ def main():
"foam nose; soft window light from left, eye-level medium shot.")
video_path = "https://d2drjpuinn46lb.cloudfront.net/painter_original_edit.mp4"
generator.generate_video(
prompt,
negative_prompt="",
video_path=video_path,
output_path=OUTPUT_PATH,
save_video=True,
height=480,
width=832,
num_frames=81,
fps=24,
guidance_scale=5.0,
)
generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt="",
inputs=InputConfig(video_path=video_path),
sampling=SamplingConfig(
height=480,
width=832,
num_frames=81,
fps=24,
guidance_scale=5.0,
),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
))
if __name__ == "__main__":
+38 -23
View File
@@ -1,4 +1,6 @@
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig,
SamplingConfig)
from fastvideo.models.dits.matrixgame2.utils import create_action_presets
import torch
@@ -38,35 +40,48 @@ def main():
# attempt to identify the optimal arguments.
config = VARIANT_CONFIG[MODEL_VARIANT]
generator = VideoGenerator.from_pretrained(
config["model_path"],
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=config["model_path"],
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
),
),
)
)
num_frames = 597
actions = create_action_presets(num_frames, keyboard_dim=config["keyboard_dim"])
grid_sizes = torch.tensor([150, 44, 80])
generator.generate_video(
prompt="",
image_path=config["image_url"],
mouse_cond=actions["mouse"].unsqueeze(0),
keyboard_cond=actions["keyboard"].unsqueeze(0),
grid_sizes=grid_sizes,
num_frames=num_frames,
height=352,
width=640,
num_inference_steps=50,
output_path=OUTPUT_PATH,
save_video=True,
generator.generate(
GenerationRequest(
prompt="",
inputs=InputConfig(
image_path=config["image_url"],
mouse_cond=actions["mouse"].unsqueeze(0),
keyboard_cond=actions["keyboard"].unsqueeze(0),
grid_sizes=grid_sizes,
),
sampling=SamplingConfig(
num_frames=num_frames,
height=352,
width=640,
num_inference_steps=50,
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
)
@@ -1,5 +1,6 @@
from fastvideo.entrypoints.streaming_generator import StreamingVideoGenerator
from fastvideo.models.dits.matrixgame2.utils import get_current_action_async, expand_action_to_frames
from fastvideo.api import EngineConfig, GeneratorConfig, OffloadConfig
import torch
import asyncio
@@ -42,17 +43,23 @@ async def main():
# attempt to identify the optimal arguments.
config = VARIANT_CONFIG[MODEL_VARIANT]
generator = StreamingVideoGenerator.from_pretrained(
config["model_path"],
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = StreamingVideoGenerator.from_config(
GeneratorConfig(
model_path=config["model_path"],
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder=False,
),
),
)
)
max_blocks = 50
+34 -21
View File
@@ -1,4 +1,7 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig, SamplingConfig,
)
MODEL_PATH = "FastVideo/Matrix-Game-3.0-Base-Distilled-Diffusers"
IMAGE_URL = "https://raw.githubusercontent.com/SkyworkAI/Matrix-Game/main/Matrix-Game-3/demo_images/001/image.png"
@@ -7,28 +10,38 @@ OUTPUT_PATH = "video_samples_matrixgame3"
def main():
generator = VideoGenerator.from_pretrained(
MODEL_PATH,
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=MODEL_PATH,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
))
generator.generate_video(
prompt=PROMPT,
image_path=IMAGE_URL,
height=720,
width=1280,
num_frames=57,
num_inference_steps=3,
guidance_scale=1.0,
seed=42,
output_path=OUTPUT_PATH,
save_video=True,
)
generator.generate(
GenerationRequest(
prompt=PROMPT,
inputs=InputConfig(image_path=IMAGE_URL),
sampling=SamplingConfig(
height=720,
width=1280,
num_frames=57,
num_inference_steps=3,
guidance_scale=1.0,
seed=42,
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
))
if __name__ == "__main__":
+36 -20
View File
@@ -1,40 +1,56 @@
from fastvideo import VideoGenerator, PipelineConfig
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
PipelineSelection,
SamplingConfig,
)
def main():
config = PipelineConfig.from_pretrained("Wan-AI/Wan2.1-T2V-1.3B-Diffusers")
config.text_encoder_precisions = ["fp16"]
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
pipeline_config=config,
use_fsdp_inference=False, # Disable FSDP for MPS
dit_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
disable_autocast=False,
num_gpus=1,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # Disable FSDP for MPS
disable_autocast=False,
offload=OffloadConfig(
dit=True,
text_encoder=True,
pin_cpu_memory=True,
),
),
pipeline=PipelineSelection(
experimental={"pipeline_config": config},
),
)
)
# Create sampling parameters with reduced number of frames
sampling_param = SamplingParam.from_pretrained("Wan-AI/Wan2.1-T2V-1.3B-Diffusers")
sampling_param.num_frames = 25 # Reduce from default 81 to 25 frames bc we have to use the SDPA attn backend for mps
sampling_param.height = 256
sampling_param.width = 256
# Reduce from default 81 to 25 frames bc we have to use the SDPA attn backend for mps
sampling = SamplingConfig(
num_frames=25,
height=256,
width=256,
)
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones.")
video = generator.generate_video(prompt, sampling_param=sampling_param)
video = generator.generate(GenerationRequest(prompt=prompt, sampling=sampling))
prompt2 = ("A majestic lion strides across the golden savanna, its powerful frame "
"glistening under the warm afternoon sun. The tall grass ripples gently in "
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, sampling_param=sampling_param)
video2 = generator.generate(GenerationRequest(prompt=prompt2, sampling=sampling))
if __name__ == "__main__":
main()
+25 -15
View File
@@ -1,6 +1,8 @@
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig,
OutputConfig,
)
OUTPUT_PATH = "video_samples"
def main():
@@ -8,17 +10,23 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=2,
use_fsdp_inference=True,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
distributed_executor_backend="ray",
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=2,
use_fsdp_inference=True,
execution_backend="ray",
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder=False,
),
),
)
)
# Generate videos with the same simple API, regardless of GPU count
@@ -27,7 +35,8 @@ def main():
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True)
video = generator.generate(
GenerationRequest(prompt=prompt, output=OutputConfig(output_path=OUTPUT_PATH, save_video=True)))
# Generate another video with a different prompt, without reloading the
# model!
@@ -37,7 +46,8 @@ def main():
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True)
video2 = generator.generate(
GenerationRequest(prompt=prompt2, output=OutputConfig(output_path=OUTPUT_PATH, save_video=True)))
if __name__ == "__main__":
+45 -29
View File
@@ -85,24 +85,35 @@ def main() -> None:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig, OutputConfig,
ParallelismConfig, PipelineSelection, SamplingConfig,
)
os.makedirs(args.out_dir, exist_ok=True)
init_kwargs = {
"num_gpus": args.num_gpus,
"workload_type": "t2i",
"sp_size": 1,
"tp_size": 1,
"dit_cpu_offload": False,
"dit_layerwise_offload": False,
"text_encoder_cpu_offload": False,
"vae_cpu_offload": False,
"image_encoder_cpu_offload": False,
"pin_cpu_memory": False,
"use_fsdp_inference": False,
}
generator = VideoGenerator.from_pretrained(model_path=args.model_path, **init_kwargs)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=args.num_gpus,
use_fsdp_inference=False,
parallelism=ParallelismConfig(
sp_size=1,
tp_size=1,
),
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
text_encoder=False,
vae=False,
image_encoder=False,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(workload_type="t2i"),
)
)
try:
for i, prompt in enumerate(prompts):
seed = args.seed + i
@@ -113,20 +124,25 @@ def main() -> None:
output_path = os.path.join(args.out_dir, f"{filename_base}.png")
print(f"[sd35] prompt_idx={i} seed={seed} output_path={output_path}")
generation_kwargs = {
"output_path": output_path,
"height": args.height,
"width": args.width,
"num_frames": 1,
"fps": 1,
"num_inference_steps": args.steps,
"guidance_scale": args.guidance,
"seed": seed,
"negative_prompt": args.negative,
"save_video": True,
}
generator.generate_video(prompt, **generation_kwargs)
generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=args.negative,
sampling=SamplingConfig(
height=args.height,
width=args.width,
num_frames=1,
fps=1,
num_inference_steps=args.steps,
guidance_scale=args.guidance,
seed=seed,
),
output=OutputConfig(
output_path=output_path,
save_video=True,
),
)
)
print(f"[sd35] done. outputs written to: {args.out_dir}")
finally:
@@ -1,6 +1,14 @@
import os
import time
from fastvideo import VideoGenerator, SamplingParam
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
)
OUTPUT_PATH = "video_samples_causal"
def main():
@@ -9,23 +17,33 @@ def main():
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
model_name = "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_name,
generator_config = GeneratorConfig(
model_path=model_name,
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
text_encoder_cpu_offload=False,
dit_cpu_offload=False,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
text_encoder=False,
dit=False,
),
),
)
sampling_param = SamplingParam.from_pretrained(model_name)
generator = VideoGenerator.from_config(generator_config)
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, sampling_param=sampling_param)
request = GenerationRequest(
prompt=prompt,
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
video = generator.generate(request)
if __name__ == "__main__":
main()
@@ -1,8 +1,17 @@
# NOTE: This is still a work in progress, and the checkpoints are not released yet.
from fastvideo import VideoGenerator, SamplingParam
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
InputConfig,
OffloadConfig,
OutputConfig,
PipelineSelection,
SamplingConfig,
)
import json
# from fastvideo.api.sampling_param import SamplingParam
OUTPUT_PATH = "video_samples_self_forcing_causal_wan2_2_14B_i2v"
def main():
@@ -10,26 +19,37 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"FastVideo/SFWan2.2-I2V-A14B-Preview-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
dit_precision="fp32",
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
dmd_denoising_steps=[1000, 850, 700, 550, 350, 275, 200, 125],
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/SFWan2.2-I2V-A14B-Preview-Diffusers",
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder=False,
),
),
pipeline=PipelineSelection(
experimental={
"dit_precision": "fp32",
"dmd_denoising_steps": [1000, 850, 700, 550, 350, 275, 200, 125],
},
),
)
)
sampling_param = SamplingParam.from_pretrained("FastVideo/SFWan2.2-I2V-A14B-Preview-Diffusers")
sampling_param.num_frames = 81
sampling_param.width = 832
sampling_param.height = 480
sampling_param.seed = 1000
sampling = SamplingConfig(
num_frames=81,
width=832,
height=480,
seed=1000,
)
with open("assets/prompts/mixkit_i2v.jsonl", "r") as f:
prompt_image_pairs = json.load(f)
@@ -37,7 +57,14 @@ def main():
for prompt_image_pair in prompt_image_pairs:
prompt = prompt_image_pair["prompt"]
image_path = prompt_image_pair["image_path"]
_ = generator.generate_video(prompt, image_path=image_path, output_path=OUTPUT_PATH, save_video=True, sampling_param=sampling_param)
_ = generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=image_path),
sampling=sampling,
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
@@ -1,8 +1,10 @@
# NOTE: This is still a work in progress, and the checkpoints are not released yet.
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig,
OffloadConfig, OutputConfig, PipelineSelection, SamplingConfig,
)
OUTPUT_PATH = "video_samples_self_forcing_causal_wan2_2_14B_t2v"
def main():
@@ -10,34 +12,49 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"rand0nmr/SFWan2.2-T2V-A14B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
dmd_denoising_steps=[1000, 850, 700, 550, 350, 275, 200, 125],
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
init_weights_from_safetensors="/mnt/sharefs/users/hao.zhang/wei/SFwan2.2_distill_self_forcing_release_cfg2/checkpoint-246_weight_only/generator_inference_transformer/",
init_weights_from_safetensors_2="/mnt/sharefs/users/hao.zhang/wei/SFwan2.2_distill_self_forcing_release_cfg2/checkpoint-246_weight_only/generator_2_inference_transformer/",
num_frame_per_block=7,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="rand0nmr/SFWan2.2-T2V-A14B-Diffusers",
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
transformer_weights="/mnt/sharefs/users/hao.zhang/wei/SFwan2.2_distill_self_forcing_release_cfg2/checkpoint-246_weight_only/generator_inference_transformer/",
transformer_2_weights="/mnt/sharefs/users/hao.zhang/wei/SFwan2.2_distill_self_forcing_release_cfg2/checkpoint-246_weight_only/generator_2_inference_transformer/",
),
experimental={
"dmd_denoising_steps": [1000, 850, 700, 550, 350, 275, 200, 125],
"num_frame_per_block": 7,
},
),
# image_encoder_cpu_offload=False,
)
)
# sampling_param = SamplingParam.from_pretrained("Wan-AI/Wan2.1-T2V-1.3B-Diffusers")
# sampling_param.num_frames = 45
# sampling_param.image_path = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/astronaut.jpg"
# Generate videos with the same simple API, regardless of GPU count
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
_ = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, num_frames=81)
_ = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
main()
main()
+19 -13
View File
@@ -51,25 +51,31 @@ Prerequisites:
uv pip install k_diffusion einops_exts alias_free_torch torchsde
"""
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OutputConfig)
PROMPT = "Lo-fi hip hop instrumental with vinyl crackle and gentle piano."
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/stable-audio-open-1.0-Diffusers",
num_gpus=1,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/stable-audio-open-1.0-Diffusers",
engine=EngineConfig(num_gpus=1),
))
output_path = "outputs_audio/stable_audio_basic/output_stable_audio.wav"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
# 6-second clip; the model max is ~47.5s.
audio_end_in_s=6.0,
# The registered preset gives 100 steps + CFG=7.0 by default;
# override num_inference_steps / guidance_scale here for QA.
)
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(
output_path=output_path,
save_video=True,
),
# 6-second clip; the model max is ~47.5s.
extensions={"audio_end_in_s": 6.0},
# The registered preset gives 100 steps + CFG=7.0 by default;
# override num_inference_steps / guidance_scale here for QA.
))
generator.shutdown()
@@ -48,6 +48,12 @@ Picking `init_audio_strength` (0.0 to 1.0):
Prerequisites: same as `basic_stable_audio.py`.
"""
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OutputConfig,
)
PROMPT = "Change the piano to a cello playing the same notes"
# Path to any audio-bearing file (wav, mp3, mp4, m4a, flac, ...).
@@ -58,18 +64,24 @@ INIT_AUDIO_STRENGTH = 0.6
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/stable-audio-open-1.0-Diffusers",
num_gpus=1,
)
generator.generate_video(
prompt=PROMPT,
output_path="outputs_audio/stable_audio_a2a/output_a2a.wav",
save_video=True,
audio_end_in_s=6.0,
init_audio=INIT_AUDIO_PATH,
init_audio_strength=INIT_AUDIO_STRENGTH,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/stable-audio-open-1.0-Diffusers",
engine=EngineConfig(num_gpus=1),
))
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(
output_path="outputs_audio/stable_audio_a2a/output_a2a.wav",
save_video=True,
),
extensions={
"audio_end_in_s": 6.0,
"init_audio": INIT_AUDIO_PATH,
"init_audio_strength": INIT_AUDIO_STRENGTH,
},
))
generator.shutdown()
@@ -48,6 +48,9 @@ Prerequisites: same as `basic_stable_audio.py`.
import os
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GeneratorConfig, GenerationRequest, OutputConfig,
)
PROMPT = "Steady lo-fi hip hop drum loop with vinyl crackle."
# Required: path to the reference audio file (wav, mp3, mp4, m4a, flac,
@@ -64,19 +67,25 @@ def main() -> None:
f"REFERENCE_AUDIO_PATH={REFERENCE_AUDIO_PATH!r} does not exist. "
"Edit this script to point at a real audio file (wav/mp3/mp4/"
"m4a/flac) before running.")
generator = VideoGenerator.from_pretrained(
"FastVideo/stable-audio-open-1.0-Diffusers",
num_gpus=1,
)
generator.generate_video(
prompt=PROMPT,
output_path="outputs_audio/stable_audio_inpaint/output_inpaint.wav",
save_video=True,
audio_end_in_s=TOTAL_SECONDS,
inpaint_audio=REFERENCE_AUDIO_PATH,
# Tuple form: keep first KEEP_SECONDS, regenerate the rest.
inpaint_mask=(KEEP_SECONDS, TOTAL_SECONDS),
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/stable-audio-open-1.0-Diffusers",
engine=EngineConfig(num_gpus=1),
))
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(
output_path="outputs_audio/stable_audio_inpaint/output_inpaint.wav",
save_video=True,
),
extensions={
"audio_end_in_s": TOTAL_SECONDS,
"inpaint_audio": REFERENCE_AUDIO_PATH,
# Tuple form: keep first KEEP_SECONDS, regenerate the rest.
"inpaint_mask": (KEEP_SECONDS, TOTAL_SECONDS),
},
))
generator.shutdown()
@@ -28,24 +28,27 @@ Prerequisites: same as `basic_stable_audio.py`. The converted repo is
public so no gated-access flow is required.
"""
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OutputConfig)
PROMPT = "Lo-fi hip hop instrumental with vinyl crackle and gentle piano."
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/stable-audio-open-small-Diffusers",
num_gpus=1,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="FastVideo/stable-audio-open-small-Diffusers",
engine=EngineConfig(num_gpus=1),
))
output_path = "outputs_audio/stable_audio_small/output_stable_audio_small.wav"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
# Small variant trains on a ~11.9s window — keep `audio_end_in_s`
# at or below that.
audio_end_in_s=6.0,
)
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(output_path=output_path, save_video=True),
# Small variant trains on a ~11.9s window — keep `audio_end_in_s`
# at or below that.
extensions={"audio_end_in_s": 6.0},
))
generator.shutdown()
@@ -4,6 +4,9 @@ import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "SLA_ATTN"
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples_turbodiffusion"
@@ -11,14 +14,17 @@ OUTPUT_PATH = "video_samples_turbodiffusion"
def main() -> None:
# TurboDiffusion: 1-4 step video generation using RCM scheduler + SLA attention
# FastVideo will automatically use TurboDiffusionPipeline when specified
generator = VideoGenerator.from_pretrained(
"loayrashid/TurboWan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
# set to false if using RTX 4090
# pin_cpu_memory=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="loayrashid/TurboWan2.1-T2V-1.3B-Diffusers",
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
),
# set to false if using RTX 4090
# pin_cpu_memory=False,
)
)
# Generate videos with the same simple API, regardless of GPU count
@@ -28,11 +34,17 @@ def main() -> None:
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(
prompt,
output_path=OUTPUT_PATH,
save_video=True,
seed=42,
video = generator.generate(
GenerationRequest(
prompt=prompt,
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
sampling=SamplingConfig(
seed=42,
),
)
)
# Generate another video with a different prompt, without reloading the model!
@@ -43,11 +55,17 @@ def main() -> None:
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic."
)
video2 = generator.generate_video(
prompt2,
output_path=OUTPUT_PATH,
save_video=True,
seed=42,
video2 = generator.generate(
GenerationRequest(
prompt=prompt2,
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
sampling=SamplingConfig(
seed=42,
),
)
)
@@ -4,6 +4,9 @@ import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "SLA_ATTN"
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples_turbodiffusion_14B"
@@ -11,10 +14,12 @@ OUTPUT_PATH = "video_samples_turbodiffusion_14B"
def main() -> None:
# TurboDiffusion 14B: 1-4 step video generation using RCM scheduler + SLA attention
# FastVideo will automatically use TurboDiffusionPipeline when specified
generator = VideoGenerator.from_pretrained(
"loayrashid/TurboWan2.1-T2V-14B-Diffusers",
# 14B model needs more GPUs
num_gpus=2,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="loayrashid/TurboWan2.1-T2V-14B-Diffusers",
# 14B model needs more GPUs
engine=EngineConfig(num_gpus=2),
)
)
prompt = (
@@ -22,11 +27,12 @@ def main() -> None:
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
video = generator.generate_video(
prompt,
output_path=OUTPUT_PATH,
save_video=True,
seed=42,
video = generator.generate(
GenerationRequest(
prompt=prompt,
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
sampling=SamplingConfig(seed=42),
)
)
# Generate another video with a different prompt, without reloading the model!
@@ -37,11 +43,12 @@ def main() -> None:
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic."
)
video2 = generator.generate_video(
prompt2,
output_path=OUTPUT_PATH,
save_video=True,
seed=42,
video2 = generator.generate(
GenerationRequest(
prompt=prompt2,
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
sampling=SamplingConfig(seed=42),
)
)
@@ -4,6 +4,10 @@ import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "SLA_ATTN"
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig,
OutputConfig, SamplingConfig,
)
# Use local model path
MODEL_PATH = "loayrashid/TurboWan2.2-I2V-A14B-Diffusers"
@@ -12,9 +16,11 @@ OUTPUT_PATH = "video_samples_turbodiffusion_i2v"
def main() -> None:
# TurboDiffusion I2V: 1-4 step image-to-video generation
generator = VideoGenerator.from_pretrained(
MODEL_PATH,
num_gpus=2,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=MODEL_PATH,
engine=EngineConfig(num_gpus=2),
)
)
# Example prompt and image for I2V
@@ -24,12 +30,13 @@ def main() -> None:
image_path = "https://huggingface.co/datasets/YiYiXu/testing-images/resolve/main/wan_i2v_input.JPG"
video = generator.generate_video(
prompt,
image_path=image_path,
output_path=OUTPUT_PATH,
save_video=True,
seed=42,
video = generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=image_path),
sampling=SamplingConfig(seed=42),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
+35 -20
View File
@@ -1,6 +1,8 @@
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig,
OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples_wan2_2_14B_t2v"
def main():
@@ -8,30 +10,37 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.2-T2V-A14B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=2,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.2-T2V-A14B-Diffusers",
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=2,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
),
),
)
)
# sampling_param = SamplingParam.from_pretrained("Wan-AI/Wan2.1-T2V-1.3B-Diffusers")
# sampling_param.num_frames = 45
# sampling_param.image_path = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/astronaut.jpg"
# Generate videos with the same simple API, regardless of GPU count
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
_ = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, height=720, width=1280, num_frames=81)
# video = generator.generate_video(prompt, sampling_param=sampling_param, output_path="wan_t2v_videos/")
_ = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(height=720, width=1280, num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
# Generate another video with a different prompt, without reloading the
# model!
@@ -41,8 +50,14 @@ def main():
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
_ = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True, height=720, width=1280, num_frames=81)
_ = generator.generate(
GenerationRequest(
prompt=prompt2,
sampling=SamplingConfig(height=720, width=1280, num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
main()
main()
+35 -16
View File
@@ -1,6 +1,12 @@
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
InputConfig,
OffloadConfig,
OutputConfig,
)
OUTPUT_PATH = "video_samples_wan2_1_Fun"
OUTPUT_NAME = "wan2.1_test"
@@ -9,18 +15,24 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"IRMChen/Wan2.1-Fun-1.3B-Control-Diffusers",
# "alibaba-pai/Wan2.2-Fun-A14B-Control",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="IRMChen/Wan2.1-Fun-1.3B-Control-Diffusers",
# "alibaba-pai/Wan2.2-Fun-A14B-Control",
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder=False,
),
),
)
)
prompt = "一位年轻女性穿着一件粉色的连衣裙,裙子上有白色的装饰和粉色的纽扣。她的头发是紫色的,头上戴着一个红色的大蝴蝶结,显得非常可爱和精致。她还戴着一个红色的领结,整体造型充满了少女感和活力。她的表情温柔,双手轻轻交叉放在身前,姿态优雅。背景是简单的灰色,没有任何多余的装饰,使得人物更加突出。她的妆容清淡自然,突显了她的清新气质。整体画面给人一种甜美、梦幻的感觉,仿佛置身于童话世界中。"
@@ -30,7 +42,14 @@ def main():
image_path = "https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/wan_fun/asset_Wan2_2/v1.0/8.png"
control_video_path = "https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/wan_fun/asset_Wan2_2/v1.0/pose.mp4"
video = generator.generate_video(prompt, negative_prompt=negative_prompt, image_path=image_path, video_path=control_video_path, output_path=OUTPUT_PATH, output_video_name=OUTPUT_NAME, save_video=True)
video = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=negative_prompt,
inputs=InputConfig(image_path=image_path, video_path=control_video_path),
output=OutputConfig(output_path=OUTPUT_PATH, output_video_name=OUTPUT_NAME, save_video=True),
)
)
if __name__ == "__main__":
main()
main()
+30 -15
View File
@@ -1,6 +1,8 @@
from fastvideo import VideoGenerator
# from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig,
OffloadConfig, OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples_wan2_2_14B_i2v"
def main():
@@ -8,23 +10,36 @@ def main():
# model.
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.2-I2V-A14B-Diffusers",
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True, # DiT need to be offloaded for MoE
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.2-I2V-A14B-Diffusers",
# FastVideo will automatically handle distributed setup
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True, # DiT need to be offloaded for MoE
vae=False,
text_encoder=True,
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
pin_cpu_memory=True,
# image_encoder=False,
),
),
)
)
prompt = "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard. The fluffy-furred feline gazes directly at the camera with a relaxed expression. Blurred beach scenery forms the background featuring crystal-clear waters, distant green hills, and a blue sky dotted with white clouds. The cat assumes a naturally relaxed posture, as if savoring the sea breeze and warm sunlight. A close-up shot highlights the feline's intricate details and the refreshing atmosphere of the seaside."
image_path = "https://huggingface.co/datasets/YiYiXu/testing-images/resolve/main/wan_i2v_input.JPG"
video = generator.generate_video(prompt, image_path=image_path, output_path=OUTPUT_PATH, save_video=True, height=832, width=480, num_frames=81)
video = generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=image_path),
sampling=SamplingConfig(height=832, width=480, num_frames=81),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
main()
main()
+33 -13
View File
@@ -1,4 +1,7 @@
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, InputConfig, OffloadConfig, OutputConfig,
)
OUTPUT_PATH = "video_samples_wan2_2_5B_ti2v"
def main():
@@ -7,22 +10,34 @@ def main():
# If a local path is provided, FastVideo will make a best effort
# attempt to identify the optimal arguments.
model_name = "Wan-AI/Wan2.2-TI2V-5B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_name,
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
dit_cpu_offload=True,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder_cpu_offload=False,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=model_name,
engine=EngineConfig(
# FastVideo will automatically handle distributed setup
num_gpus=1,
use_fsdp_inference=False, # set to True if GPU is out of memory
offload=OffloadConfig(
dit=True,
vae=False,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
# image_encoder=False,
),
),
)
)
# I2V is triggered just by passing in an image_path argument
prompt = "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard. The fluffy-furred feline gazes directly at the camera with a relaxed expression. Blurred beach scenery forms the background featuring crystal-clear waters, distant green hills, and a blue sky dotted with white clouds. The cat assumes a naturally relaxed posture, as if savoring the sea breeze and warm sunlight. A close-up shot highlights the feline's intricate details and the refreshing atmosphere of the seaside."
image_path = "https://huggingface.co/datasets/YiYiXu/testing-images/resolve/main/wan_i2v_input.JPG"
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, image_path=image_path)
video = generator.generate(
GenerationRequest(
prompt=prompt,
inputs=InputConfig(image_path=image_path),
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
# Generate another video with a different prompt, without reloading the
# model!
@@ -34,8 +49,13 @@ def main():
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
"cinematic.")
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True)
video2 = generator.generate(
GenerationRequest(
prompt=prompt2,
output=OutputConfig(output_path=OUTPUT_PATH, save_video=True),
)
)
if __name__ == "__main__":
main()
main()
+21 -11
View File
@@ -50,23 +50,33 @@ N_DUP = 4 # how many times to duplicate the video for the gen/ref corpora
def generate_one_ltx2_video() -> str:
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "FLASH_ATTN")
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OutputConfig, SamplingConfig)
Path(OUTPUT_PATH).parent.mkdir(parents=True, exist_ok=True)
# Davids048/LTX2-Base-Diffusers is the audio-capable LTX-2 checkpoint
# (the Distilled variant ships without the audio VAE, so its mp4
# audio track is silence/noise — unusable for audio.* metrics).
generator = VideoGenerator.from_pretrained(
"Davids048/LTX2-Base-Diffusers",
num_gpus=1,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Davids048/LTX2-Base-Diffusers",
engine=EngineConfig(num_gpus=1),
)
)
generator.generate_video(
prompt=PROMPT,
output_path=OUTPUT_PATH,
save_video=True,
num_frames=121, # ~5s @ 24 fps — long enough for audio.desync (Synchformer ≥14 segments)
height=480,
width=832,
fps=24,
generator.generate(
GenerationRequest(
prompt=PROMPT,
sampling=SamplingConfig(
num_frames=121, # ~5s @ 24 fps — long enough for audio.desync (Synchformer ≥14 segments)
height=480,
width=832,
fps=24,
),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
)
generator.shutdown()
torch.cuda.empty_cache()
@@ -21,6 +21,10 @@ Install: ``uv pip install -e .[eval-audio]`` covers both metrics here
import torch
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig,
SamplingConfig,
)
from fastvideo.eval import create_evaluator
PROMPT = (
@@ -39,20 +43,26 @@ METRICS = [
def main() -> None:
generator = VideoGenerator.from_pretrained(
"Davids048/LTX2-Base-Diffusers",
num_gpus=1,
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Davids048/LTX2-Base-Diffusers",
engine=EngineConfig(num_gpus=1),
))
output_path = "outputs_video/ltx2_audio_eval/output.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
num_frames=121,
height=1088,
width=1920,
)
generator.generate(
GenerationRequest(
prompt=PROMPT,
sampling=SamplingConfig(
num_frames=121,
height=1088,
width=1920,
),
output=OutputConfig(
output_path=output_path,
save_video=True,
),
))
generator.shutdown()
torch.cuda.empty_cache()
+15 -10
View File
@@ -22,6 +22,10 @@ sharing, or run on a smaller-resolution generation.
import torch
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig,
SamplingConfig,
)
from fastvideo.eval import Evaluator
from fastvideo.eval.io import build_eval_kwargs
@@ -58,19 +62,20 @@ METRICS = [
def main() -> None:
# ----- generation (matches examples/inference/basic/basic_ltx2.py) -----
generator = VideoGenerator.from_pretrained(
"Davids048/LTX2-Base-Diffusers",
num_gpus=1,
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Davids048/LTX2-Base-Diffusers",
engine=EngineConfig(num_gpus=1),
)
)
output_path = "outputs_video/ltx2_basic/output_ltx2_base_t2v_1088_1920_1.1.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
num_frames=121,
height=1088,
width=1920,
generator.generate(
GenerationRequest(
prompt=PROMPT,
output=OutputConfig(output_path=output_path, save_video=True),
sampling=SamplingConfig(num_frames=121, height=1088, width=1920),
)
)
generator.shutdown()
# Free residual CUDA memory the generator left behind so the
+11 -5
View File
@@ -45,6 +45,9 @@ def _generate_videos(rows: list[dict], videos_dir: Path,
model: str, num_gpus: int,
num_frames: int, height: int, width: int) -> None:
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OutputConfig, SamplingConfig,
)
videos_dir.mkdir(parents=True, exist_ok=True)
todo = [(row, videos_dir / _expected_filename(row)) for row in rows]
@@ -55,13 +58,16 @@ def _generate_videos(rows: list[dict], videos_dir: Path,
print(f"[gen] {len(todo)}/{len(rows)} scenarios to render with {model} "
f"({num_frames}x{height}x{width})...")
gen = VideoGenerator.from_pretrained(model, num_gpus=num_gpus)
gen = VideoGenerator.from_config(GeneratorConfig(
model_path=model, engine=EngineConfig(num_gpus=num_gpus),
))
try:
for row, out_path in todo:
gen.generate_video(
prompt=row["prompt"], output_path=str(out_path), save_video=True,
num_frames=num_frames, height=height, width=width,
)
gen.generate(GenerationRequest(
prompt=row["prompt"],
sampling=SamplingConfig(num_frames=num_frames, height=height, width=width),
output=OutputConfig(output_path=str(out_path), save_video=True),
))
finally:
gen.shutdown()
+9 -5
View File
@@ -43,6 +43,8 @@ def _generate_videos(prompts: list[str], videos_dir: Path,
model: str, num_gpus: int,
num_frames: int, height: int, width: int) -> None:
from fastvideo import VideoGenerator
from fastvideo.api import (EngineConfig, GenerationRequest, GeneratorConfig,
OutputConfig, SamplingConfig)
videos_dir.mkdir(parents=True, exist_ok=True)
todo = [(p, videos_dir / f"{_slugify(p)}.mp4") for p in prompts]
@@ -53,13 +55,15 @@ def _generate_videos(prompts: list[str], videos_dir: Path,
print(f"[gen] {len(todo)}/{len(prompts)} prompts to render with {model} "
f"({num_frames}x{height}x{width})...")
gen = VideoGenerator.from_pretrained(model, num_gpus=num_gpus)
gen = VideoGenerator.from_config(GeneratorConfig(
model_path=model, engine=EngineConfig(num_gpus=num_gpus)))
try:
for prompt, out_path in todo:
gen.generate_video(
prompt=prompt, output_path=str(out_path), save_video=True,
num_frames=num_frames, height=height, width=width,
)
gen.generate(GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(num_frames=num_frames, height=height, width=width),
output=OutputConfig(output_path=str(out_path), save_video=True),
))
finally:
gen.shutdown()
+26 -8
View File
@@ -33,6 +33,13 @@ import json
from pathlib import Path
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OutputConfig,
SamplingConfig,
)
from fastvideo.eval import create_evaluator
from fastvideo.eval.io import load_video
@@ -99,16 +106,27 @@ def generate(args: argparse.Namespace) -> Path:
out.parent.mkdir(parents=True, exist_ok=True)
print(f"[gen] loading {args.model} ({args.num_gpus} GPU)...")
generator = VideoGenerator.from_pretrained(args.model, num_gpus=args.num_gpus)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=args.model,
engine=EngineConfig(num_gpus=args.num_gpus),
)
)
try:
print(f"[gen] generating to {out}...")
generator.generate_video(
prompt=args.prompt,
output_path=str(out),
save_video=True,
num_frames=args.num_frames,
height=args.height,
width=args.width,
generator.generate(
GenerationRequest(
prompt=args.prompt,
sampling=SamplingConfig(
num_frames=args.num_frames,
height=args.height,
width=args.width,
),
output=OutputConfig(
output_path=str(out),
save_video=True,
),
)
)
finally:
generator.shutdown()
+1 -1
View File
@@ -33,7 +33,7 @@ This demo initializes a `VideoGenerator` with the minimum required arguments for
The core functionality is in the `generate_video` function, which:
1. Processes user inputs
2. Uses the FastVideo VideoGenerator from earlier to run inference (`generator.generate_video()`)
2. Uses the FastVideo VideoGenerator from earlier to run inference (`generator.generate(GenerationRequest(...))`)
## Gradio Interface
@@ -5,7 +5,13 @@ import time
import gradio as gr
from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
GenerationRequest,
GeneratorConfig,
OutputConfig,
SamplingConfig,
SamplingParam,
)
from copy import deepcopy
@@ -129,9 +135,22 @@ def create_gradio_interface(default_params: dict[str, SamplingParam], generators
output_dir = "outputs/"
os.makedirs(output_dir, exist_ok=True)
start_time = time.time()
result = generator.generate_video(prompt=prompt, sampling_param=params, save_video=True, return_frames=False)
result = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=params.negative_prompt,
sampling=SamplingConfig(
seed=int(params.seed),
guidance_scale=params.guidance_scale,
num_frames=int(params.num_frames),
height=int(params.height),
width=int(params.width),
),
output=OutputConfig(save_video=True, return_frames=False),
)
)
inference_time = time.time() - start_time
logging_info = result.get("logging_info", None)
logging_info = result.logging_info
if logging_info:
stage_names = logging_info.get_execution_order()
stage_execution_times = [
@@ -550,7 +569,7 @@ def main():
for model_path in model_paths:
print(f"Loading model: {model_path}")
setup_model_environment(model_path)
generators[model_path] = VideoGenerator.from_pretrained(model_path)
generators[model_path] = VideoGenerator.from_config(GeneratorConfig(model_path=model_path))
default_params[model_path] = SamplingParam.from_pretrained(model_path)
demo = create_gradio_interface(default_params, generators)
print(f"Starting Gradio frontend at http://{args.host}:{args.port}")
@@ -55,10 +55,11 @@ demo can actually boot:
`fastvideo/fastvideo_args.py` currently wires only `ltx2_vae_tiling`.
The backing stages (`ltx2_refine.py`, `ltx2_i2v_conditioning.py`) are
also missing from `fastvideo/pipelines/stages/`.
3. **`fastvideo.configs.sample.base.SamplingParam`** — the import path used
by this demo. Upstream moved sampling params to
`fastvideo.api.sampling_param`. A re-export shim at the old path, or an
import update here once the other two prereqs land, will resolve it.
3. **`SamplingParam`** — now imported from `fastvideo.api` (the public
re-export of `fastvideo.api.sampling_param`); the old
`fastvideo.configs.sample.base` path was removed upstream. `SamplingParam`
here only sources model-default slider values — generation itself runs
through the typed `GenerationRequest` / `generator.generate(...)` path.
## Environment variables
@@ -4,8 +4,16 @@ from pathlib import Path
import gradio as gr
from fastvideo.api import (
CompileConfig,
ComponentConfig,
EngineConfig,
GeneratorConfig,
OffloadConfig,
PipelineSelection,
SamplingParam,
)
from fastvideo.configs.pipelines.base import PipelineConfig
from fastvideo.configs.sample.base import SamplingParam
from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo.layers.quantization.fp4_config import FP4Config
from fastvideo.utils import maybe_download_model
@@ -48,28 +56,44 @@ def main():
refine_upsampler_path = resolve_refine_upsampler_path(resolved_model_path)
print(f"Using refine upsampler: {refine_upsampler_path}")
generators[model_path] = VideoGenerator.from_pretrained(
str(resolved_model_path),
num_gpus=1,
ltx2_refine_enabled=True,
ltx2_refine_upsampler_path=str(refine_upsampler_path),
ltx2_refine_lora_path="", # disable refine LoRA for distilled model
ltx2_refine_num_inference_steps=2,
ltx2_refine_guidance_scale=1.0,
ltx2_refine_add_noise=True,
pipeline_config=pipeline_config,
enable_torch_compile=True,
enable_torch_compile_text_encoder=True,
torch_compile_kwargs={
"backend": "inductor",
"fullgraph": True,
"mode": "max-autotune-no-cudagraphs",
"dynamic": False,
},
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
ltx2_vae_tiling=False,
generators[model_path] = VideoGenerator.from_config(
GeneratorConfig(
model_path=str(resolved_model_path),
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=False,
),
compile=CompileConfig(
enabled=True,
text_encoder_enabled=True,
backend="inductor",
fullgraph=True,
mode="max-autotune-no-cudagraphs",
dynamic=False,
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
upsampler_weights=str(refine_upsampler_path),
# Empty refine LoRA path (distilled needs none) -> omit.
),
vae_tiling=False,
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 2,
"guidance_scale": 1.0,
"add_noise": True,
},
},
# PipelineConfig object (with FP4 quant wired on above) has
# no first-class typed field; route via experimental.
experimental={"pipeline_config": pipeline_config},
),
)
)
default_params[model_path] = apply_ltx2_defaults(
SamplingParam.from_pretrained(str(resolved_model_path))
@@ -4,7 +4,7 @@ from pathlib import Path
import torch
import torch._inductor.config
from fastvideo.configs.sample.base import SamplingParam
from fastvideo.api import SamplingParam
LOCAL_DEMO_DIR = Path(__file__).resolve().parent
CLASSIFIER_DIR = Path(
@@ -5,8 +5,14 @@ from copy import deepcopy
import gradio as gr
from fastvideo.configs.sample.base import SamplingParam
from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo import VideoGenerator
from fastvideo.api import (
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
SamplingParam,
)
from .config import (
DEFAULT_FPS,
@@ -69,40 +75,38 @@ def create_gradio_interface(default_params: dict[str, SamplingParam], generators
output_path = str(OUTPUT_DIR / video_filename)
params.output_path = output_path
start_time = time.perf_counter()
result = generator.generate_video(
prompt=prompt,
output_path=output_path,
fps=DEFAULT_FPS,
seed=int(params.seed),
save_video=True,
return_frames=False,
guidance_scale=float(params.guidance_scale),
height=int(params.height),
width=int(params.width),
num_frames=int(params.num_frames),
num_inference_steps=DEFAULT_NUM_INFERENCE_STEPS,
negative_prompt=params.negative_prompt,
image_path=params.image_path,
ltx2_image_crf=0.0
result = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=params.negative_prompt,
inputs=InputConfig(image_path=params.image_path),
sampling=SamplingConfig(
seed=int(params.seed),
fps=DEFAULT_FPS,
guidance_scale=float(params.guidance_scale),
height=int(params.height),
width=int(params.width),
num_frames=int(params.num_frames),
num_inference_steps=DEFAULT_NUM_INFERENCE_STEPS,
),
output=OutputConfig(
output_path=output_path,
save_video=True,
return_frames=False,
),
# LTX-2 i2v knob without a first-class typed field yet.
extensions={"ltx2_image_crf": 0.0},
)
)
wall_time = time.perf_counter() - start_time
generation_time = (
result.get("generation_time")
if isinstance(result, dict) else None
)
e2e_latency = (
result.get("e2e_latency")
if isinstance(result, dict) else None
)
generation_time = result.generation_time
e2e_latency = result.extra.get("e2e_latency")
if generation_time is None:
generation_time = wall_time
if e2e_latency is None:
e2e_latency = wall_time
resolved_output_path = (
result.get("output_path", output_path)
if isinstance(result, dict) else output_path
)
logging_info = result.get("logging_info", None) if isinstance(result, dict) else None
resolved_output_path = result.video_path or output_path
logging_info = result.logging_info
if logging_info:
stage_names = logging_info.get_execution_order()
stage_execution_times = [
@@ -9,6 +9,7 @@ import uvicorn
from fastapi import FastAPI, Request, HTTPException
from fastapi.responses import HTMLResponse, FileResponse
from fastvideo.api import EngineConfig, GeneratorConfig, OffloadConfig
from fastvideo.entrypoints.streaming_generator import StreamingVideoGenerator
from fastvideo.models.dits.matrixgame2.utils import expand_action_to_frames
@@ -572,14 +573,20 @@ def main():
print(f"Loading model: {model_path}")
setup_model_environment(model_path)
generator = StreamingVideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=True,
dit_cpu_offload=True,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
generator = StreamingVideoGenerator.from_config(
GeneratorConfig(
model_path=model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
offload=OffloadConfig(
dit=True,
vae=False,
text_encoder=True,
pin_cpu_memory=True,
),
),
)
)
generators = {model_path: generator}
@@ -3,7 +3,6 @@ import os
import torch
import base64
import io
from copy import deepcopy
from typing import Dict, Any, Optional, List
import signal
import sys
@@ -20,6 +19,17 @@ import imageio
from ray.serve.handle import DeploymentHandle
from prometheus_client import Counter, Histogram, generate_latest
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
InputConfig,
OffloadConfig,
OutputConfig,
PipelineSelection,
SamplingConfig,
)
NUM_GPUS = 16
DEFAULT_FPS = 16
SEED_RANGE_MAX = 1_000_000
@@ -136,10 +146,10 @@ def setup_model_environment(model_path: str) -> None:
def process_generation_result(result: Any) -> tuple[List[np.ndarray], float, List[str], List[float]]:
frames = result if isinstance(result, list) else result.get("frames", [])
generation_time = result.get("generation_time", 0.0) if isinstance(result, dict) else 0.0
logging_info = result.get("logging_info", None)
frames = result.frames or []
generation_time = result.generation_time or 0.0
logging_info = result.logging_info
if logging_info:
stage_names = logging_info.get_execution_order()
stage_execution_times = [
@@ -153,24 +163,29 @@ def process_generation_result(result: Any) -> tuple[List[np.ndarray], float, Lis
return frames, generation_time, stage_names, stage_execution_times
def prepare_sampling_params(video_request: VideoGenerationRequest, default_params: Any) -> Any:
params = deepcopy(default_params)
params.prompt = video_request.prompt
if video_request.use_negative_prompt:
params.negative_prompt = video_request.negative_prompt
def prepare_generation_request(video_request: VideoGenerationRequest, image_path: Optional[str] = None) -> Any:
seed = (video_request.seed if not video_request.randomize_seed
else torch.randint(0, SEED_RANGE_MAX, (1,)).item())
params.seed = (video_request.seed if not video_request.randomize_seed
else torch.randint(0, SEED_RANGE_MAX, (1,)).item())
params.randomize_seed = video_request.randomize_seed
params.guidance_scale = video_request.guidance_scale
params.num_frames = video_request.num_frames
params.height = video_request.height
params.width = video_request.width
params.save_video = False
params.return_frames = True
return params
# "" explicitly clears the model preset's negative prompt (None would
# inherit it, changing this demo's long-standing behavior).
negative_prompt = video_request.negative_prompt if video_request.use_negative_prompt else ""
request = GenerationRequest(
prompt=video_request.prompt,
negative_prompt=negative_prompt,
inputs=InputConfig(image_path=image_path),
sampling=SamplingConfig(
seed=seed,
guidance_scale=video_request.guidance_scale,
num_frames=video_request.num_frames,
height=video_request.height,
width=video_request.width,
),
output=OutputConfig(save_video=False, return_frames=True),
)
return request, seed
class BaseModelDeployment:
@@ -185,31 +200,34 @@ class BaseModelDeployment:
def _initialize_generator(self, config: Dict[str, Any]) -> None:
from fastvideo.entrypoints.video_generator import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
print(f"Initializing model: {self.model_path}")
self.generator = VideoGenerator.from_pretrained(
model_path=self.model_path,
num_gpus=1,
use_fsdp_inference=True,
text_encoder_cpu_offload=config["text_encoder_cpu_offload"],
dmd_denoising_steps=[1000, 850, 700, 550, 350, 275, 200, 125], # TODO: hardocde for I2V
dit_precision="fp32", # TODO: hardocde for I2V
dit_cpu_offload=config["dit_cpu_offload"],
vae_cpu_offload=config["vae_cpu_offload"],
VSA_sparsity=config["VSA_sparsity"],
enable_stage_verification=False,
self.generator = VideoGenerator.from_config(
GeneratorConfig(
model_path=self.model_path,
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=True,
enable_stage_verification=False,
offload=OffloadConfig(
text_encoder=config["text_encoder_cpu_offload"],
dit=config["dit_cpu_offload"],
vae=config["vae_cpu_offload"],
),
),
pipeline=PipelineSelection(
# I2V knobs without first-class typed fields yet.
experimental={
"dmd_denoising_steps": [1000, 850, 700, 550, 350, 275, 200, 125],
"dit_precision": "fp32",
"VSA_sparsity": config["VSA_sparsity"],
},
),
)
)
self.default_params = SamplingParam.from_pretrained(self.model_path)
self.default_params.seed = 1000
self.default_params.num_frames = 73
self.default_params.width = 832
self.default_params.height = 480
def generate_video(self, video_request: VideoGenerationRequest) -> VideoGenerationResponse:
total_start_time = time.time()
params = prepare_sampling_params(video_request, self.default_params)
# Save image if provided (for I2V)
image_path = None
@@ -218,19 +236,15 @@ class BaseModelDeployment:
if image_path is None:
return VideoGenerationResponse(
video_data=None,
seed=params.seed,
seed=video_request.seed,
success=False,
error_message="Failed to save input image",
)
request, seed = prepare_generation_request(video_request, image_path)
inference_start_time = time.time()
result = self.generator.generate_video(
prompt=video_request.prompt,
sampling_param=params,
image_path=image_path,
save_video=False,
return_frames=True,
)
result = self.generator.generate(request)
inference_time = time.time() - inference_start_time
frames, generation_time, stage_names, stage_execution_times = process_generation_result(result)
@@ -250,7 +264,7 @@ class BaseModelDeployment:
return VideoGenerationResponse(
video_data=video_data,
seed=params.seed,
seed=seed,
success=True,
generation_time=generation_time,
inference_time=inference_time,
+45 -25
View File
@@ -1,18 +1,31 @@
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (
ComponentConfig, EngineConfig, GenerationRequest, GeneratorConfig,
OffloadConfig, OutputConfig, PipelineSelection, SamplingConfig,
)
OUTPUT_PATH = "./lora_out"
def main():
# Initialize VideoGenerator with the Wan model
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
lora_path="benjamin-paine/steamboat-willie-1.3b",
lora_nickname="steamboat"
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="benjamin-paine/steamboat-willie-1.3b",
),
experimental={"lora_nickname": "steamboat"},
),
)
)
kwargs = {
"height": 480,
@@ -26,25 +39,32 @@ def main():
prompt = "steamboat willie style, golden era animation, close-up of a short fluffy monster kneeling beside a melting red candle. the mood is one of wonder and curiosity, as the monster gazes at the flame with wide eyes and open mouth. Its pose and expression convey a sense of innocence and playfulness, as if it is exploring the world around it for the first time. The use of warm colors and dramatic lighting further enhances the cozy atmosphere of the image."
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
video = generator.generate_video(
prompt,
# sampling_param=sampling_param,
output_path=OUTPUT_PATH,
save_video=True,
negative_prompt=negative_prompt,
**kwargs
video = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=negative_prompt,
sampling=SamplingConfig(**kwargs),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
)
generator.set_lora_adapter(lora_nickname="flat_color", lora_path="motimalu/wan-flat-color-1.3b-v2")
prompt = "flat color, no lineart, blending, negative space, artist:[john kafka|ponsuke kaikai|hara id 21|yoneyama mai|fuzichoco], 1girl, sakura miko, pink hair, cowboy shot, white shirt, floral print, off shoulder, outdoors, cherry blossom, tree shade, wariza, looking up, falling petals, half-closed eyes, white sky, clouds, live2d animation, upper body, high quality cinematic video of a woman sitting under a sakura tree. Dreamy and lonely, the camera close-ups on the face of the woman as she turns towards the viewer. The Camera is steady, This is a cowboy shot. The animation is smooth and fluid."
negative_prompt = "bad quality video,色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
video = generator.generate_video(
prompt,
output_path=OUTPUT_PATH,
save_video=True,
negative_prompt=negative_prompt,
**kwargs
video = generator.generate(
GenerationRequest(
prompt=prompt,
negative_prompt=negative_prompt,
sampling=SamplingConfig(**kwargs),
output=OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
),
)
)
if __name__ == "__main__":
main()
main()
@@ -2,45 +2,60 @@
Inference using a LoRA checkpoint from FastVideo trainer.
"""
from fastvideo import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.api import (ComponentConfig, EngineConfig, GenerationRequest,
GeneratorConfig, OffloadConfig, OutputConfig,
PipelineSelection, SamplingConfig)
OUTPUT_PATH = "./lora_out"
def main():
# Initialize VideoGenerator with the Wan model
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
lora_path="checkpoints/wan_t2v_finetune_lora/checkpoint-160/transformer",
lora_nickname="crush_smol"
)
generator = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
),
),
pipeline=PipelineSelection(
components=ComponentConfig(
lora_path="checkpoints/wan_t2v_finetune_lora/checkpoint-160/transformer",
),
experimental={"lora_nickname": "crush_smol"},
),
))
generator.unmerge_lora_weights()
kwargs = {
"height": 480,
"width": 832,
"num_frames": 77,
"guidance_scale": 6.0,
"num_inference_steps": 50,
"seed": 42,
}
sampling = SamplingConfig(
height=480,
width=832,
num_frames=77,
guidance_scale=6.0,
num_inference_steps=50,
seed=42,
)
output = OutputConfig(
output_path=OUTPUT_PATH,
save_video=True,
)
# Generate video with LoRA style
prompt = "A large metal cylinder is seen pressing down on a pile of Oreo cookies, flattening them as if they were under a hydraulic press."
video = generator.generate_video(
prompt,
output_path=OUTPUT_PATH,
save_video=True,
**kwargs
)
video = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=sampling,
output=output,
))
prompt = "A large metal cylinder is seen compressing colorful clay into a compact shape, demonstrating the power of a hydraulic press."
video = generator.generate_video(
prompt,
output_path=OUTPUT_PATH,
save_video=True,
**kwargs
)
video = generator.generate(
GenerationRequest(
prompt=prompt,
sampling=sampling,
output=output,
))
if __name__ == "__main__":
main()
main()
+69
View File
@@ -0,0 +1,69 @@
# LTX-2.3 distilled inference configs
Ready-to-run `fastvideo generate` run configs for the LTX-2.3
distilled model (`FastVideo/LTX-2.3-Distilled-Diffusers`), covering both
workloads (t2v / i2v), both two-stage step schedules (`5+2`, `8+3` = denoise
+ refine), and four resolutions.
```bash
fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_1280x832.yaml
```
Each config is self-contained (no preset registry needed): the two-stage
refine is wired via `generator.pipeline.preset_overrides.refine`, and the
base sampling knobs live under `request.sampling`. The refine upsampler
auto-resolves from the model's `spatial_upscaler`.
## Configs
| workload | schedule | resolution (HxW) | file |
|---|---|---|---|
| t2v | 5+2 | 1280x832 | `t2v_5s2_1280x832.yaml` |
| t2v | 5+2 | 1024x1536 | `t2v_5s2_1024x1536.yaml` |
| t2v | 5+2 | 768x1280 | `t2v_5s2_768x1280.yaml` |
| t2v | 5+2 | 512x768 | `t2v_5s2_512x768.yaml` |
| t2v | 8+3 | 1280x832 | `t2v_8s3_1280x832.yaml` |
| t2v | 8+3 | 1024x1536 | `t2v_8s3_1024x1536.yaml` |
| t2v | 8+3 | 768x1280 | `t2v_8s3_768x1280.yaml` |
| t2v | 8+3 | 512x768 | `t2v_8s3_512x768.yaml` |
| i2v | 5+2 | 1280x832 | `i2v_5s2_1280x832.yaml` |
| i2v | 5+2 | 1024x1536 | `i2v_5s2_1024x1536.yaml` |
| i2v | 5+2 | 768x1280 | `i2v_5s2_768x1280.yaml` |
| i2v | 5+2 | 512x768 | `i2v_5s2_512x768.yaml` |
| i2v | 8+3 | 1280x832 | `i2v_8s3_1280x832.yaml` |
| i2v | 8+3 | 1024x1536 | `i2v_8s3_1024x1536.yaml` |
| i2v | 8+3 | 768x1280 | `i2v_8s3_768x1280.yaml` |
| i2v | 8+3 | 512x768 | `i2v_8s3_512x768.yaml` |
## Overriding without editing a file
Dotted overrides (prefixes `generator.` / `request.`) let you tweak any field:
```bash
# swap prompt
fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_1280x832.yaml \
--request.prompt "a red fox running through fresh snow"
# change output path / gpu count
fastvideo generate --config examples/inference/ltx2_3/t2v_5s2_512x768.yaml \
--request.output.output_path outputs/preview.mp4 \
--generator.engine.num_gpus 4
```
## i2v
The `i2v_*` configs take a first-frame image via
`request.extensions.ltx2_images` (`[[path, frame_offset, weight]]`). Edit the
path in the file, or override it:
```bash
fastvideo generate --config examples/inference/ltx2_3/i2v_8s3_1280x832.yaml \
--request.extensions.ltx2_images '[["/data/portrait.jpg", 0, 1.0]]'
```
## Schedules
`5+2` is the fast preview schedule; `8+3` is the higher-quality distilled
recipe. Refine (`preset_overrides.refine.num_inference_steps`) only accepts 2
or 3 steps.
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 5+2 two-stage at 1024x1536.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_5s2_1024x1536.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 1024
width: 1536
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_5s2_1024x1536.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 5+2 two-stage at 1280x832.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_5s2_1280x832.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 1280
width: 832
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_5s2_1280x832.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 5+2 two-stage at 512x768.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_5s2_512x768.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 512
width: 768
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_5s2_512x768.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 5+2 two-stage at 768x1280.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_5s2_768x1280.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 768
width: 1280
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_5s2_768x1280.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 8+3 two-stage at 1024x1536.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_8s3_1024x1536.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 1024
width: 1536
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_8s3_1024x1536.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 8+3 two-stage at 1280x832.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_8s3_1280x832.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 1280
width: 832
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_8s3_1280x832.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 8+3 two-stage at 512x768.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_8s3_512x768.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 512
width: 768
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_8s3_512x768.mp4
save_video: true
@@ -0,0 +1,38 @@
# LTX-2.3 distilled i2v — 8+3 two-stage at 768x1280.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/i2v_8s3_768x1280.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: i2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
The subject slowly turns toward the camera with a soft, natural expression, hair and clothing swaying gently, shallow depth of field, subtle cinematic motion.
sampling:
height: 768
width: 1280
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
# i2v conditioning — replace the path with your own first-frame image.
extensions:
ltx2_images:
- ["/path/to/your/first_frame.jpg", 0, 1.0]
ltx2_image_crf: 0.0
output:
output_path: outputs/ltx2_3_i2v_8s3_768x1280.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 5+2 two-stage at 1024x1536.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_5s2_1024x1536.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 1024
width: 1536
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
output:
output_path: outputs/ltx2_3_t2v_5s2_1024x1536.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 5+2 two-stage at 1280x832.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_5s2_1280x832.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 1280
width: 832
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
output:
output_path: outputs/ltx2_3_t2v_5s2_1280x832.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 5+2 two-stage at 512x768.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_5s2_512x768.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 512
width: 768
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
output:
output_path: outputs/ltx2_3_t2v_5s2_512x768.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 5+2 two-stage at 768x1280.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_5s2_768x1280.yaml
#
# Stage 1 denoises for 5 steps at half resolution; the latents are then
# spatially upsampled and refined for 2 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 2
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 768
width: 1280
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 5
output:
output_path: outputs/ltx2_3_t2v_5s2_768x1280.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 8+3 two-stage at 1024x1536.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_1024x1536.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 1024
width: 1536
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
output:
output_path: outputs/ltx2_3_t2v_8s3_1024x1536.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 8+3 two-stage at 1280x832.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_1280x832.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 1280
width: 832
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
output:
output_path: outputs/ltx2_3_t2v_8s3_1280x832.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 8+3 two-stage at 512x768.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_512x768.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 512
width: 768
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
output:
output_path: outputs/ltx2_3_t2v_8s3_512x768.mp4
save_video: true
@@ -0,0 +1,33 @@
# LTX-2.3 distilled t2v — 8+3 two-stage at 768x1280.
#
# Run:
# fastvideo generate --config examples/inference/ltx2_3/t2v_8s3_768x1280.yaml
#
# Stage 1 denoises for 8 steps at half resolution; the latents are then
# spatially upsampled and refined for 3 steps (refine only supports 2 or 3).
# The refine upsampler auto-resolves from the model's `spatial_upscaler`.
generator:
model_path: FastVideo/LTX-2.3-Distilled-Diffusers
engine:
num_gpus: 1
pipeline:
workload_type: t2v
preset_overrides:
refine:
enabled: true
num_inference_steps: 3
guidance_scale: 1.0
add_noise: true
request:
prompt: >-
A cinematic drone shot flying over dramatic coastal cliffs at sunrise, golden light spilling across the water, gentle waves breaking on the rocks below, ultra-detailed, smooth camera motion.
sampling:
height: 768
width: 1280
num_frames: 121
fps: 24
guidance_scale: 1.0
num_inference_steps: 8
output:
output_path: outputs/ltx2_3_t2v_8s3_768x1280.mp4
save_video: true
@@ -37,6 +37,17 @@ import imageio
import torch
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig,
ComponentConfig,
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
PipelineSelection,
SamplingConfig,
)
from fastvideo.configs.pipelines.base import PipelineConfig
from fastvideo.layers.quantization.nvfp4_qat_config import NVFP4QATConfig
@@ -148,35 +159,49 @@ def build_generator(args: argparse.Namespace) -> VideoGenerator:
compile_enabled = not args.no_compile
extra_kwargs = {}
# ``pipeline_config`` is a PipelineConfig object (not a string path) and
# ``output_type`` has no first-class typed field, so both are routed through
# the pipeline experimental escape hatch.
experimental = {"pipeline_config": pipeline_config}
components = ComponentConfig()
if args.distilled_model:
weights_path = resolve_distilled_weights(args.distilled_model)
print(f"Using distilled weights: {args.distilled_model} -> {weights_path}")
extra_kwargs["init_weights_from_safetensors"] = weights_path
components.transformer_weights = weights_path
if args.taehv:
# Skip the in-pipeline VAE decode entirely: the pipeline returns raw
# latents, the Wan VAE is offloaded to CPU (and not compiled) since we
# decode with TAEHV in this script instead.
extra_kwargs["output_type"] = "latent"
experimental["output_type"] = "latent"
generator = VideoGenerator.from_pretrained(
model_id,
pipeline_config=pipeline_config,
num_gpus=args.num_gpus,
# Keep everything resident on the GPU -- no offloading, except the
# unused Wan VAE when TAEHV handles decoding.
use_fsdp_inference=False,
dit_cpu_offload=False,
dit_layerwise_offload=False,
vae_cpu_offload=args.taehv,
text_encoder_cpu_offload=False,
pin_cpu_memory=False,
enable_torch_compile=compile_enabled,
enable_torch_compile_text_encoder=compile_enabled,
enable_torch_compile_vae=compile_enabled and not args.taehv,
**extra_kwargs,
generator_config = GeneratorConfig(
model_path=model_id,
engine=EngineConfig(
num_gpus=args.num_gpus,
# Keep everything resident on the GPU -- no offloading, except the
# unused Wan VAE when TAEHV handles decoding.
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
vae=args.taehv,
text_encoder=False,
pin_cpu_memory=False,
),
compile=CompileConfig(
enabled=compile_enabled,
text_encoder_enabled=compile_enabled,
vae_enabled=compile_enabled and not args.taehv,
),
),
pipeline=PipelineSelection(
components=components,
experimental=experimental,
),
)
generator = VideoGenerator.from_config(generator_config)
return generator
@@ -237,11 +262,11 @@ def main() -> None:
# runs below measure steady-state latency only.
with silence_request_log():
for _ in range(args.warmups):
warm = generator.generate(request={
"prompt": PROMPT,
"sampling": {"num_inference_steps": 2, "guidance_scale": args.guidance_scale},
"output": {"save_video": False, "return_frames": args.taehv},
})
warm = generator.generate(GenerationRequest(
prompt=PROMPT,
sampling=SamplingConfig(num_inference_steps=2, guidance_scale=args.guidance_scale),
output=OutputConfig(save_video=False, return_frames=args.taehv),
))
if args.taehv:
taehv.decode(warm.samples)
@@ -257,18 +282,18 @@ def main() -> None:
frames = None
with silence_request_log():
for i in range(args.benchmark_runs):
result = generator.generate(request={
"prompt": PROMPT,
"sampling": {
"num_inference_steps": args.infer_steps,
"guidance_scale": args.guidance_scale,
},
"output": {
"save_video": False,
"return_frames": args.taehv,
"output_path": output_path,
},
})
result = generator.generate(GenerationRequest(
prompt=PROMPT,
sampling=SamplingConfig(
num_inference_steps=args.infer_steps,
guidance_scale=args.guidance_scale,
),
output=OutputConfig(
save_video=False,
return_frames=args.taehv,
output_path=output_path,
),
))
denoise_elapsed = result.generation_time
denoise_times.append(denoise_elapsed)
@@ -2,31 +2,41 @@ import os
import time
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig, GenerationRequest, GeneratorConfig, OffloadConfig,
OutputConfig, SamplingConfig,
)
def main():
# set the attention backend
# set the attention backend
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
start_time = time.perf_counter()
gen = VideoGenerator.from_pretrained(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
)
gen = VideoGenerator.from_config(
GeneratorConfig(
model_path="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
engine=EngineConfig(
num_gpus=1,
offload=OffloadConfig(
dit=False,
vae=False,
text_encoder=True,
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
),
),
))
load_time = time.perf_counter() - start_time
print(f"Model loading time: {load_time:.2f} seconds")
gen_start_time = time.perf_counter()
gen.generate_video(
prompt=
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting.",
seed=1024,
output_path="example_outputs/")
gen.generate(
GenerationRequest(
prompt=
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting.",
sampling=SamplingConfig(seed=1024),
output=OutputConfig(output_path="example_outputs/")))
generation_time = time.perf_counter() - gen_start_time
print(f"Video generation time: {generation_time:.2f} seconds")
@@ -18,6 +18,10 @@ import os
import time
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig, EngineConfig, GenerationRequest, GeneratorConfig,
OffloadConfig, OutputConfig, SamplingConfig,
)
OUTPUT_PATH = "video_samples"
@@ -39,17 +43,24 @@ def main():
mode += "_compile"
print(f"Mode: {mode.upper()}")
generator = VideoGenerator.from_pretrained(
args.model,
num_gpus=args.num_gpus,
nvfp4_fa4=args.nvfp4_fa4,
use_fsdp_inference=not args.nvfp4_fa4,
dit_cpu_offload=False,
dit_layerwise_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
enable_torch_compile=args.compile,
)
if args.nvfp4_fa4:
os.environ["FASTVIDEO_NVFP4_FA4"] = "1"
os.environ.setdefault("CUTE_DSL_ENABLE_TVM_FFI", "1")
generator = VideoGenerator.from_config(GeneratorConfig(
model_path=args.model,
engine=EngineConfig(
num_gpus=args.num_gpus,
use_fsdp_inference=not args.nvfp4_fa4,
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
vae=True,
text_encoder=True,
),
compile=CompileConfig(enabled=args.compile),
),
))
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
@@ -59,16 +70,22 @@ def main():
n_warmup = 2 if args.compile else 1
for i in range(n_warmup):
generator.generate(request={"prompt": prompt, "sampling": {"num_inference_steps": 2},
"output": {"save_video": False}})
generator.generate(GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(num_inference_steps=2),
output=OutputConfig(save_video=False),
))
os.makedirs(OUTPUT_PATH, exist_ok=True)
start = time.time()
generator.generate(request={
"prompt": prompt,
"sampling": {"num_inference_steps": args.infer_steps},
"output": {"save_video": True, "output_path": os.path.join(OUTPUT_PATH, f"raccoon_{mode}.mp4")},
})
generator.generate(GenerationRequest(
prompt=prompt,
sampling=SamplingConfig(num_inference_steps=args.infer_steps),
output=OutputConfig(
save_video=True,
output_path=os.path.join(OUTPUT_PATH, f"raccoon_{mode}.mp4"),
),
))
elapsed = time.time() - start
print(f"[{mode.upper()}] {args.infer_steps} steps in {elapsed:.2f}s "
f"({args.infer_steps / elapsed:.2f} it/s)")

Some files were not shown because too many files have changed in this diff Show More