Pipelines path → tensor decode behind GPU metric compute via a new
VideoPool, so multi-sample eval runs no longer serialize disk I/O and
metric work. The Evaluator owns one pool per evaluate(samples=...)
call; each EvalWorker is a single-GPU consumer that grabs decoded
samples from the shared queue (work-stealing across replicas when
num_gpus > 1).
Worker pre-uploads video/reference to its device once per sample so
every metric in the loop consumes the same GPU-resident tensor (no
per-metric .to(device) traffic).
SSIM and PSNR move to the GPU — at 1080p × 121 frames the CPU path
was both slow (5–10 s/pair) and contended with the loader thread for
DDR bandwidth. LPIPS gains a chunk_size knob (default 8) that caps
peak from ~60 GB to ~5 GB with bit-identical output. Optical-flow
metrics drop chunk_size to 1 because DPFlow's cost volume is ~4 GB
per frame pair at 1080p (matches mhuo/ptlflow upstream).
physics_iq and a handful of vbench metrics drop their list-batch
shims — the per-sample contract is now uniform across the suite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CI runs `pre-commit run --all-files` which surfaces every latent issue
across the eval suite, not just the ones in changed files. Before this
commit, that surfaced 9 ruff errors and 47 mypy errors that had built
up over earlier PRs (last touched in `[style] eval: ruff/yapf pass` and
related). Cleaning them in one sweep so subsequent eval PRs land green.
Ruff
- yapf reformat: 3 files (entrypoints/cli/eval.py, optical_flow/_shared.py,
physics_iq/utils.py).
- B024: drop ABC inheritance from PromptDataset (it has no abstract
methods; subclasses just populate self._rows).
- B027: BaseMetric.setup is intentionally an optional no-op override
(metrics with no eager state inherit it). Document and noqa instead
of forcing every native metric to declare an empty override.
- SIM105 / SIM115: contextlib.suppress for ipc_collect, mkstemp for
the temp video file in vbench.scene.
- UP038: isinstance(x, (A, B)) -> isinstance(x, A | B).
Mypy
- Optional-narrowing pattern across 12 metric files: each metric class
initialised self._model (and sometimes _processor / _tokenizer /
_head) to None and reassigned in setup(), but mypy narrows the type
to None and flags every later attribute access. Annotate the slots
as Any. Same fix already applied to physics_iq + human_action +
videoscore2 in earlier commits; this extends it to the rest of the
suite.
- Untyped helpers: annotate _safe_amt_forward (motion_smoothness),
_patch_detectron2_registries (_grit_helper), _Loader (vbench/__init__).
- _grit_helper: function-attribute writes (`func._patched_idempotent`)
type: ignore'd at the assignment site.
- imaging_quality: rename `all_scores` rebinding to `chunks` / `per_frame`
so mypy doesn't carry the list[Any] type into the cat'd tensor.
- worker._resolve_video_input: add Any -> Any annotation.
No runtime behaviour changes. 33/33 eval pytest still pass.
Voice pass over the three eval-related docs to bring them in line with
the project's existing voice (see docs/contributing/pull_requests.md,
docs/contributing/testing.md, docs/getting_started/installation/gpu.md):
* fastvideo/eval/README.md
* docs/contributing/eval-metrics.md
* .agents/workflows/evaluation-development.md
Concretely: cut em-dashes from ~30 across the three files to 1 (left
in a table cell where it reads naturally), removed "no X, no Y"
negation patterns, dropped buzzy phrasing ("first-class", "drop a
file", "out of the box"), and demoted bold imperatives ("Do not X")
to plain prose where the surrounding context already conveys the
instruction.
Two stale path references in evaluation-development.md fixed while
the file was open:
* "Update .agents/skills/evaluate-video-quality.md" pointed at a flat
file; skills in this repo are directories. Corrected to
.agents/skills/evaluate-video-quality/SKILL.md.
* "Check the evaluation_registry.md" referenced a bare filename that
does not exist; aligned to .agents/memory/evaluation-registry/README.md
to match the other references in the same doc.
No code changes; tests still 33/4 on a 1-GPU borrow.
create_evaluator(metrics='vbench') now scores the 11 vbench sub-metrics
that don't need detectron2 instead of crashing on construction. Explicit
metric names (e.g. metrics=['vbench.color']) still raise ImportError
with the full two-step install command.
Group selectors filter missing deps
- evaluator._resolve_metric_names: when expanding a group prefix or
'all', drop metrics whose declared dependencies aren't importable.
One warning per skipped metric. Explicit names pass through unchanged
so the missing dep surfaces as ImportError -- the friendly contract
for "user asked for this specific metric."
- registry.missing_dependencies(name): new helper exposing per-metric
dep status without instantiation.
Install hint actually satisfies the dep
- registry._install_hint(metric, dep) replaces _extra_for() in the
ImportError formatter. detectron2 needs the base extra PLUS a git+
install with build-isolation off; the old hint sent users in a
circle. All other deps still resolve via the existing extra map.
- README install table aligned to use uv pip install verbatim
(matches docs/getting_started/installation/gpu.md).
Other VBench/VLM fixes uncovered by running the group end-to-end
- vbench.human_action: filename was 'l16_25m.pth' which 404s on the
OpenGVLab/VBench_Used_Models HF repo. The actual UMT-L Kinetics-400
checkpoint there is 'l16_ptk710_ftk710_ftk400_f16_res224.pth'
(matches the metric's expected vit_large_patch16_224 / num_classes=400
/ all_frames=16 shape exactly).
- videoscore2: switch to AutoModelForImageTextToText with a
AutoModelForVision2Seq fallback (the legacy alias is being phased
out in transformers 4.45+); pass dtype=torch.bfloat16 explicitly
(transformers 4.57 deprecated torch_dtype= in favor of dtype=).
Mypy hygiene scoped to the two metric files I touched
- self._model / _processor / _tokenizer typed as Any to silence the
pre-existing None-narrowing errors that surface only when these files
are staged. Same pattern used across the eval suite; fixing it
branch-wide is out of scope for this change.
Tested
- pytest fastvideo/tests/eval/ -> 33 passed (single-GPU subset).
- create_evaluator(metrics=['vbench.color']) -> raises ImportError with
the full uv pip install ... && uv pip install --no-build-isolation
'git+https://github.com/facebookresearch/detectron2.git' command.
- _resolve_metric_names('vbench') -> 11 of 16 vbench metrics, with
4 detectron2-deps + 1 qwen_omni_utils-dep correctly filtered.
- vbench.human_action loads against the correct checkpoint name
(verified end-to-end on a noise tensor; metric runs forward pass).
`get_dataset("physics_iq")` now works with no kwargs. The manifest CSV is
vendored under fastvideo/eval/metrics/physics_iq/_vendored/; per-scenario
videos, masks, and switch-frames auto-fetch on first use from the public
DeepMind bucket (https://storage.googleapis.com/physics-iq-benchmark) into
${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq/, sibling to the existing
models/torch/clip/ subdirs.
Why
- Old default dataset_root was /root/physics-IQ-benchmark, which doesn't
exist on shared hosts and isn't documented anywhere.
- Examples crashed with FileNotFoundError before any user-facing message
pointed at where to get the data.
- The official upstream download script needs gcloud SDK; the bucket is
also reachable over plain HTTPS (verified) so no SDK dependency is
needed.
Behavior
- dataset_root= is now an opt-in override (mirroring vbench's
full_info_path=); defaults to get_cache_dir() / "datasets" /
"physics_iq".
- auto_download=True by default; flip to False for air-gapped runs.
- FASTVIDEO_PHYSICS_IQ_BUCKET_URL env var redirects to internal mirrors.
- Atomic .part -> rename for safe concurrent SLURM-rank fetches.
Bench script fix
- bench_physics_iq.py: --limit was applied as a post-construction
list slice, after the dataset module had already eagerly resolved
every scenario's on-disk paths. With auto-download that means we'd
pull all 198 scenarios on every smoke run. Pass limit=args.limit
through to get_dataset() so partial-data smoke runs only fetch what
they need. --dataset-root is now optional (defaults to the cache
path).
Convention for future vendored files
- _vendored/ subdirs hold upstream-provenance content. They're
auto-skipped by metric discovery (the leading _) and by codespell
(single */_vendored/* glob in [tool.codespell].skip), so future
vendored files require no further config.
- docs/contributing/eval-metrics.md updated to point at this convention
for new metrics that ship paired-reference datasets.
Tested
- pytest fastvideo/tests/eval/ -> 33 passed (1-GPU subset); 4 multi-GPU
tests skipped on 1-GPU borrow as expected.
- Smoke: get_dataset("physics_iq", limit=1) auto-fetches the 5 expected
assets; physics_iq.{mse,spatial_iou} on take-1 self-pair returns
0.0 / 1.0.
**Per-benchmark eval extras (Option B layout)**
Replace the single ``[eval]`` rollup with per-benchmark groups so users
can install only what they need:
* ``[eval-vbench]`` — ``openai-clip``, ``pyiqa``, ``easydict`` (covers
12 of the 16 vbench sub-metrics; the four GRiT ones still need
``detectron2`` installed manually).
* ``[eval-physics-iq]`` — empty group; documents intent (all
physics_iq metrics already use base fastvideo deps).
* ``[eval]`` — sensible default rollup: common.lpips, optical_flow.*,
videoscore2, plus eval-vbench + eval-physics-iq. The 80% case.
* ``[eval-full]`` — adds ``qwen-omni-utils`` for the AVoCaDO-based
``vbench.scene`` metric. Detectron2 still manual.
Most deps are *not* in any of these — they're already in base
fastvideo's pinned dependencies (transformers, timm, einops, scipy,
omegaconf, opencv-python, imageio). Per-metric runtime patching keeps
versions consistent across all metrics.
Registry's missing-dep ImportError now points at the right extra
(e.g. ``pip install 'fastvideo[eval-vbench]'`` for vbench metrics)
instead of the generic ``[eval]`` it used to say.
**Drop audio metrics from this PR**
Five ``audio.*`` metrics (clap_score, frechet_distance, kl_divergence,
wer, audiobox_aesthetics; 452 lines) came in with the original wm-eval
port and have not been iterated, tested, or exercised end-to-end since.
None of the test suite or example scripts touched them. They would
have forced 7 untested deps into a public extra.
Pull them out of this PR. Code is preserved on the side branch
``shao/eval-audio`` (pushed to origin) for a follow-up audio-eval PR
that lands them with proper tests + a real end-to-end run.
**Also: actually remove the ``_assets/`` gitignore rule**
The previous commit (``381a7aae``) claimed to revert the
``fastvideo/eval/_assets/`` gitignore carve-out but the change was
unstaged at commit time, so the rule remained. This commit removes it
for real. Untracked content under that path is now visible to ``git
status`` again, which is what we wanted — that scratch dir is not the
kind of carve-out the project root ``.gitignore`` should carry.
Net registry: 32 metrics → 27; tests: 33 pass / 4 multi-GPU skipped
(unchanged); ``shao/eval-audio`` branch pushed for the follow-up.
* Delete ``scripts/eval/score_folder.py`` and ``scripts/eval/run_vbench_e2e.py``.
Both pre-dated the simpler ``examples/inference/eval/score_folder.py`` /
``bench_vbench.py`` versions and now overlap (one even name-collides).
Users follow the ``examples/`` path going forward.
* Drop the ``fastvideo/eval/_assets/`` ignore rule. It was specific to a
branch-local synthetic-flow scratch dir that no longer ships in this
PR; not the kind of carve-out the project root ``.gitignore`` should
carry.
Three loose ends from earlier hygiene passes that the final PR review
caught:
* Empty out 27 sub-metric ``__init__.py`` files that still re-exported
the metric class. The earlier ``54f26bea`` commit only cleaned 2 of
them; the rest still carried ``from .metric import FooMetric # noqa:
F401`` lines copied from the original wm-eval port. Auto-discovery
doesn't need them, and ``get_metric("group.name")`` is the canonical
access pattern. Now uniformly empty across all sub-metric inits.
* Fix stale layout references in ``fastvideo/eval/README.md`` and
``.agents/workflows/evaluation-development.md`` that still described
the abandoned ``<bench>/external/upstream/`` per-metric layout. The
actual layout is the flat ``fastvideo/third_party/eval/<bench>/`` we
switched to during the port. Also drops the stale ``_third_party``
example from ``fastvideo/eval/metrics/__init__.py``'s docstring (the
underscore-skip rule is unchanged; the example was just wrong).
* Drop ``open_clip_torch`` and ``torchmetrics`` from the
``[eval]`` extra. Neither is imported anywhere under
``fastvideo/eval/`` (or in the vbench submodule we vendor); they were
carried over from the wm-eval port's heavier dep graph.
Verified: 32 metrics still register, 33/4 tests pass/skip on 1 GPU.
``Evaluator.evaluate`` and the one-shot ``evaluate()`` helper now accept
``video`` / ``reference`` as either a pre-loaded ``(T, C, H, W)`` tensor
or a path-like (``str`` / ``Path``). Paths are decoded inside the
worker thread that picks the sample up — see
``EvalWorker._resolve_video_input``.
Why this matters for batch eval: the dispatcher in
``Evaluator.evaluate(samples=[...])`` submits every sample to the
``ThreadPoolExecutor`` upfront. With pre-loaded tensors that meant
every video in the batch was resident in CPU RAM at once (≈ 3 GB per
video at 1088×1920×121); a full benchmark of hundreds of clips would
OOM before scoring started.
With paths, the queued futures hold cheap strings; only ``num_gpus``
videos are decoded concurrently. Peak resident memory becomes
``O(num_gpus)`` instead of ``O(len(samples))``. No dispatcher rewrite
needed — the change is ~10 lines at the worker boundary.
* ``EvalWorker.evaluate`` now normalizes ``sample["video"]`` and
``sample["reference"]`` through ``_resolve_video_input`` (path →
``load_video`` decode; ``(1, T, C, H, W)`` → squeeze; tensor →
passthrough). Metrics keep the existing contract: by the time
``compute()`` runs, ``sample["video"]`` is always a 4-D tensor.
* ``score_folder.py`` and ``bench_vbench.py`` simplified to pass paths
directly. ``bench_physics_iq.py`` already did.
* ``fastvideo.eval.api.evaluate`` signature widened to ``Tensor | str |
Path``; the worker handles the decode either way.
Tests: new ``test_evaluator_paths.py`` covers the kwargs form, the
samples-list form, mixing paths and tensors in the same batch, parity
between path-form and tensor-form scores, and the missing-path
exception path. End-to-end ``score_folder.py`` smoke run confirmed.
Two small fixes in ``fastvideo/entrypoints/cli/eval.py``:
* ``_expand_paths`` deduplicated through ``[x for x in out if not (x in
seen or seen.add(x))]`` — a known idiom that ruff flags because
``set.add`` returns ``None`` (B023/func-returns-value). Replace with
an explicit loop so the side effect doesn't ride on the boolean
expression.
* ``_cmd_run`` constructed ``metrics_arg`` through an
``if/else``-block where SIM108 (project policy) prefers a ternary.
Fold to a single conditional expression.
Pure mechanical cleanups; ``ruff check`` is now clean on the file and
the eval CLI smoke test (``score_video.py`` against a self-paired
mp4 → ``common.psnr=100``) still passes.
Pure formatting pass — no behavior changes. Most edits are one of:
* yapf splitting / re-flowing kwarg-heavy callsites (argparse setup,
``MetricResult(...)`` construction).
* ruff auto-fixes around imports (collapsing ``Iterable`` to
``collections.abc``, removing the ``typing.Tuple`` shim where
``tuple[...]`` works, dropping unused imports).
* ``zip(..., strict=False)`` added to every two-iterable zip call.
* ``getattr(np, "trapz")`` → ``np.trapz`` where the fallback can read
the attribute directly (the ``trapezoid`` lookup still uses
``getattr`` because the name itself is the moving target).
Verified the test suite still passes after the pass.
The two end-to-end runners called ``VideoGenerator.generate_video``
without specifying generation dimensions, leaving the model to fall
back on whatever default sampling shape it carries internally — which
isn't necessarily what the user wants for a benchmark run.
Mirror the basic ``examples/inference/basic/basic_ltx2.py`` script:
default to ``121 x 1088 x 1920`` (LTX2's intended sampling shape) and
expose ``--num-frames`` / ``--height`` / ``--width`` so users can
downscale for smoke runs without editing the script.
Verified end-to-end on a borrowed H200: ``bench_vbench.py --limit 1
--num-frames 49 --height 480 --width 768`` generates an mp4 in ~97s
and scores it with ``vbench.aesthetic_quality`` (auto-downloads
``Davids048/LTX2-Base-Diffusers`` weights through ``from_pretrained``).
Add four small example scripts under ``examples/inference/eval/``,
each driving the public eval API in a different real-world shape:
* ``score_video.py`` — score one mp4 on one GPU. Smallest possible
use of ``create_evaluator`` + ``Evaluator.evaluate``. Optional
``--reference``, ``--text-prompt``, ``--fps`` for paired / prompt-
aware / fps-aware metrics.
* ``score_folder.py`` — score every mp4 in a directory across ``--num-gpus``
replicas via ``Evaluator.evaluate(samples=[...])``. Pair each
generated video with a same-name reference under ``--reference-dir``
if you want paired metrics.
* ``bench_vbench.py`` — full VBench end-to-end: ``get_dataset("vbench",
dimensions=...)`` → ``VideoGenerator`` (LTX2 by default) → score
with the matching ``vbench.*`` sub-metrics → print per-metric
averages. ``--skip-generation`` re-uses existing mp4s under
``--videos-dir`` so you can iterate on metric selection without
re-paying generation cost.
* ``bench_physics_iq.py`` — full Physics-IQ end-to-end:
``get_dataset("physics_iq", dataset_root=...)`` → ``VideoGenerator``
→ score with the composite ``physics_iq`` metric → aggregate via the
upstream's ``aggregate_components`` recipe.
The pre-existing ``basic_ltx2_eval.py`` and ``eval_ltx2_vbench.py`` are
left in place — they target the single-prompt generate-and-score case
and complement (rather than overlap with) the new full-dataset runners.
Verified end-to-end on a borrowed GPU node:
* ``score_video.py``: PSNR=100, SSIM=1.0 on a self-paired mp4 (sanity).
* ``score_folder.py``: produces well-formed scores.json.
Add four end-to-end test modules under ``fastvideo/tests/eval/``,
all driving the public API only — no reaches into ``EvalWorker``,
``BaseMetric``, or the ``_REGISTRY`` dict. Each test mirrors a real
caller flow.
* ``test_registry.py`` — ``list_metrics`` invariants, group-prefix
resolution (verified against the no-model ``physics_iq`` group so
the test stays cheap), unknown-metric error path.
* ``test_evaluator_single.py`` — one-shot ``evaluate(...)``,
long-lived ``Evaluator`` with both kwargs and ``samples=[...]``
shapes, deterministic scoring, the legacy ``(1, T, C, H, W)``
back-compat unwrap. Runs on CPU using ``common.psnr`` /
``common.ssim`` so it's free in CI.
* ``test_evaluator_multi_gpu.py`` — auto-skips when fewer than 2 CUDA
devices visible. Pins the round-robin contract by computing a
single-GPU baseline and asserting bit-equivalent scores under
multi-GPU dispatch with the same input list, plus the kwargs-form
→ worker-0 invariant and ``release_cuda_memory`` no-crash check.
* ``test_evaluator_with_dataset.py`` — full ``get_dataset("vbench")
→ Evaluator.evaluate(**row)`` flow. Synthesizes random video tensors
per row (no diffusion model needed) and verifies that extra dataset
keys (``prompt``, ``n_samples``, ``dimensions``, ``auxiliary_info``)
flow through unused metrics without breaking them.
29 tests total: 25 pass on a single GPU, all 29 pass on 2 GPUs.
``fastvideo eval run --output scores.json`` would crash with
``TypeError: Object of type float32 is not JSON serializable`` whenever
the chosen metric set landed numpy or torch values in
``MetricResult.details``. The optical-flow metrics in particular
populate ``per_frame_metrics`` with ``np.float64`` scalars, and pretty
much any metric is one ``np.percentile`` call away from the same crash.
Pass a ``default=`` callback to ``json.dumps`` that walks the unknown
leaves and coerces:
* ``np.integer`` / ``np.floating`` / ``np.bool_`` → native Python scalars
* ``np.ndarray`` → ``.tolist()``
* ``torch.Tensor`` → detached CPU ``.tolist()``
* ``pathlib.Path`` → ``str``
Anything else still raises ``TypeError`` — the goal is to handle the
known metric outputs cleanly, not to silently coerce arbitrary objects.
The Physics-IQ metric package was carrying a ~200-line dataset loader
(``PhysicsIQDataLoader`` + ``PhysicsIQScenario`` + an FPS-conversion
helper + manifest-walking constants) inside its metric directory, with
``__init__.py`` re-exporting the loader as a public name. Two distinct
concerns were mixed: dataset traversal (a user-of-the-metric concern)
and metric-pipeline configuration (an internal concern).
Split them:
* New ``fastvideo/eval/datasets/physics_iq.py`` registers a
``PhysicsIQPromptDataset(PromptDataset)`` under
``@register_dataset("physics_iq")``. It walks ``descriptions.csv``,
resolves per-take video and real-mask paths, FPS-converts the source
release on cache miss, and yields one sample dict per take-1
scenario shaped to drop straight into ``Evaluator.evaluate(**row)``.
The previously-public ``PhysicsIQScenario`` dataclass moves with it.
* Metric defaults (``DEFAULT_TARGET_FPS=30``,
``DEFAULT_DURATION_SECONDS=5``) move into
``metrics/physics_iq/utils.py``, where they were already used as
default kwargs.
* ``metrics/physics_iq/models.py`` is deleted; ``__init__.py`` is
emptied to match the rest of the metric directories.
Users now do ``get_dataset("physics_iq", dataset_root=...)`` to walk
the corpus instead of reaching for ``PhysicsIQDataLoader`` directly.
The old import path is dropped without a back-compat shim — the eval
suite hasn't shipped yet, so there's no API contract to honor.
The eval-suite README (``fastvideo/eval/README.md``) gets a short
``Prompt datasets`` section showing the ``get_dataset`` workflow. The
project root README is unchanged.
Two sub-metric ``__init__.py`` files re-exported their metric class
(``VBClapScoreMetric``, ``AestheticQualityMetric``) while every other
sub-metric leaves ``__init__.py`` empty. Auto-discovery imports each
``metric.py`` module by full path, so the re-export was redundant —
and the inconsistency obscured the contract for new contributors
(adding a metric should not require touching ``__init__.py``).
Standardize on empty sub-metric ``__init__.py``; users instantiate
metrics through ``get_metric("group.name")`` rather than reaching for
the class object directly.
The ``physics_iq`` group's ``__init__.py`` also re-exports a dataset
loader and scenario dataclass; that one's left alone in this commit
pending a separate move of those helpers into ``fastvideo/eval/datasets/``.
The ``EvalWorker`` always invokes metrics on a single video and only
ever reads ``result[0]`` from the returned list, so every metric had a
dead ``for b in range(B)`` loop. Tighten the contract:
* ``BaseMetric.compute(sample) -> MetricResult`` — return one result,
not a one-element list.
* ``BaseMetric._skip(sample, reason) -> MetricResult`` likewise.
* Sample-side: ``video`` and ``reference`` are ``(T, C, H, W)`` —
no leading batch dim. The worker still unwraps a ``(1, T, C, H, W)``
caller for back-compat, so existing user code keeps working.
* ``Evaluator.metric_names`` now reads through a public
``EvalWorker.metric_names`` property instead of poking the private
``_metrics`` dict.
All 28 metrics — common.{ssim,psnr,lpips}, optical_flow.*, audio.*,
vbench.* (16), physics_iq.* (5 incl. composite), videoscore2 — drop
their ``for b in range(B)`` loop and return a single ``MetricResult``.
List-shaped optional inputs (``text_prompt``, ``audio``,
``auxiliary_info``, ``actions``) are still accepted: each metric
unwraps a single-element list before use, so callers can keep the
existing ``[prompt]`` convention or pass a scalar — both work.
Update the layout snippets in both ``fastvideo/eval/README.md`` and
``docs/contributing/eval-metrics.md`` to show the new
``optical_flow/`` group sibling to ``common/``, with the two
sub-metrics (``gt_optical_flow``, ``synthetic_optical_flow``) listed
explicitly.
Move ``common.optical_flow`` to a dedicated ``optical_flow`` group so
flow-based comparisons can register additional reference modes without
piling into ``common``. Two metrics live under the new group:
* ``optical_flow.gt_optical_flow`` — the existing video-vs-video
comparison, ported verbatim minus a thin shared-helper extraction.
* ``optical_flow.synthetic_optical_flow`` — new metric that takes a
per-frame action stream + a ``ThirdPersonCalibration`` JSON and
predicts the reference flow analytically (Longuet-Higgins + off-pivot
translation, no depth) instead of extracting it from a GT video.
Both metrics share ``optical_flow/_shared.py`` for ptlflow loading,
per-frame metric computation, and temporal aggregation, so scores are
directly comparable across the two reference modes.
The third-person predictor is vendored at
``optical_flow/synthetic_optical_flow/_thirdperson.py`` to keep the
metric self-contained; the calibration *fitter* (``calibrate.py`` etc.)
is intentionally not part of the eval suite — it's an offline fitting
tool.
Several scaffolds that pre-date the runner refactor still pointed at
removed APIs:
* ``EvalResult`` (and its ``Evaluator.evaluate_dataset`` docstring)
exposed a class that is never returned anywhere; drop it from the
public API.
* ``Evaluator``'s module docstring still referenced the removed
``EvalRunner`` layer; trim to the current Evaluator → EvalWorker
shape.
* ``WM_EVAL_CACHE`` survived as a fallback env var from the wm-eval
port; remove and keep only ``FASTVIDEO_EVAL_CACHE``.
* The contributor docs documented ``batch_unit`` and
``trial_forward``/``Evaluator.calibrate()`` which were dropped from
``BaseMetric`` in earlier refactors; update to the current contract.
The metric resolved the Kinetics-400 label file via a wm-eval-era
``_third_party/umt/kinetics_400_categories.txt`` path that no longer
exists after the vbench-as-submodule port. ``os.path.exists`` was
silently False, leaving the label dict empty; every prediction then
mapped to the empty string and every score returned 0.0.
Resolve the path through ``vbench.third_party.umt.__file__`` instead,
which always points at the pinned upstream submodule.
Two changes that together let motion_smoothness score 1088×1920×121
videos on shared GPUs:
1. _get_scale() now queries torch.cuda.mem_get_info() free memory on
every call instead of caching total_memory at setup() time.
Adapts to whatever's actually available — other metric replicas
already loaded, residual generator allocations, another process
sharing the GPU. Upstream cached total_memory once which on a
shared/loaded GPU lets AMT attempt 30+ GB correlation reshapes.
2. _safe_amt_forward() wraps the model call with two-axis OOM retry:
- Halve batch until size 1 (per-pair memory dominates).
- Halve scale_factor until 1/16 (AMT's internal feature-map
resolution and therefore correlation volume size).
- Bottom out at scale=1/16, batch=1; if still OOM, the resolution
is genuinely impossible at this headroom and we re-raise.
The second change is needed because upstream's autoscale formula in
_get_scale mis-extrapolates at high resolution: it scales linearly
in pixel count, but AMT's correlation volume grows quadratically.
Rather than rewriting upstream's formula, the retry path makes the
metric robust to whatever the formula picks.
Verified on fs-mbz-gpu-085 (shared with another job holding 45 GB):
Before: motion_smoothness OOM at 31.75 GB allocation on 1088×1920×121.
After: motion_smoothness=0.9897 across all 8 metrics, no OOM, all
other scores (aesthetic_quality=0.5430, subject_consistency=
0.8629, etc.) unchanged.
Parity against runs/vbench_smoke/results_v3.json byte-identical.
Two new entrypoints showcasing the post-refactor eval surface:
- examples/inference/eval/basic_ltx2_eval.py — generate one LTX2
video using the same parameters as basic_ltx2.py (same prompt,
model, 1088×1920×121), then score it with the prompt-aware VBench
subset. Builds Evaluator directly, calls evaluate(**kwargs) per
the single-sample API; no runner, no argparse.
- scripts/eval/score_folder.py — bulk-score every video in a folder
with prompt-free VBench metrics. Optional --prompts-json maps
filenames to prompts to enable prompt-aware metrics. Uses
Evaluator.evaluate(samples=[...]) for multi-GPU fan-out.
Tested on fs-mbz-gpu-085:
- score_folder.py: 3 duplicate mp4s × 4 metrics → 3 identical score
sets, summary aggregated, scores.json written. Clean run.
- basic_ltx2_eval.py: generation succeeded; scoring path verified
for 7 of 8 metrics on the LTX2 output (aesthetic_quality,
subject_consistency, background_consistency, imaging_quality,
temporal_flickering, dynamic_degree, overall_consistency).
vbench.motion_smoothness OOMs at 1088×1920 on a shared GPU because
VBench's AMT memory autoscale reads total_memory rather than
mem_get_info() free memory, so it underestimates the required
resolution scale-down when another process holds 45 GB. Documented
in the script docstring; not refactor-related (same behavior at
HEAD pre-refactor). Drop motion_smoothness from METRICS or run on
a dedicated GPU to score it.
Added a torch.cuda.empty_cache() between generator.shutdown() and
evaluator construction to free residual generation memory.
Completes the d1bc1413 cleanup pass on the five files that were
skipped because dataloader had pending modifications. With those
modifications now landed (commits 7cdbcc94 + c2a22560), the dead
code in optical_flow + the four aux-using vbench metrics
(color/multiple_objects/object_class/spatial_relationship) can go.
Mechanical removal of batch_unit class attrs and trial_forward
method overrides — neither is read anywhere since calibrate() was
deleted in 502f061d.
Parity verified on fs-mbz-gpu-085: byte-identical to v3 baseline.
Replace the single mean-EPE port with mhuo's complete validation set
(see fastvideo/training/ptlflow_validation.py in mhuo's tree):
Per-frame: mf_epe, mf_angle_err, mf_cosine, mf_mag_ratio,
pixel_epe_mean/max, px_angle_rmse, grid_epe_mean/max,
fl_all, foe_dist, flow_kl_2d
Aggregated over time: <name>_mean / _std / _max / _auc, plus
divergence_onset_frame / divergence_threshold
Added supporting math: least-squares Focus-of-Expansion estimation,
2D KL divergence over (angle, log-magnitude) histogram, trapezoid
integral with numpy 2.x compat shim.
Headline ``score`` is pixel_epe_mean_mean (lower is better);
everything else lives in MetricResult.details so downstream
consumers can pick whichever scalar they care about.
Pairs gen and ref videos across all B inputs and batches the model
through them together (chunk_size=16) for GPU efficiency.
Authored by dataloader.
The four aux-using vbench metrics (color, multiple_objects,
object_class, spatial_relationship) used to bail at the top of
compute() if auxiliary_info was None, then unconditionally indexed
aux[b]["<key>"] for every row. This crashed with KeyError on rows
that had aux but lacked the metric's specific key — common once
the dataset yields heterogeneous rows under "vbench" / multi-dim
runs (each row carries flat aux populated only for its own
dimensions).
Switch to per-row skip: each row checks for its own required key
and emits MetricResult(score=None, details={"skipped": "..."}) if
absent. Lets a single evaluator.evaluate(samples=[...]) call score
heterogeneous-aux rows in one pass.
Also handle two structural sub-cases:
- multiple_objects requires "<a> and <b>" in aux["object"]; rows
without the separator (single-object data leaking in) skip.
- object_class is the inverse: rows whose "object" contains " and "
belong to multiple_objects' territory and skip here.
- spatial_relationship's nested {object_a/object_b/relationship}
sub-dict is read defensively; missing keys skip with a reason.
Pairs with the flat-aux-at-load behavior in VBenchPromptDataset
(commit 06735161). Without these per-row skips, dimensions="all"
runs crash on the first row that lacks the active metric's key.
Authored by dataloader.
Two small additions paired with the in-progress eval orchestration:
- types.py: EvalResult dataclass with summary/per_video fields plus
from_raw / save / print helpers. Used by scripts and any future
callers that aggregate per-sample MetricResults into a corpus-level
summary (e.g. scripts/eval/run_vbench_e2e.py).
- .gitignore: exclude fastvideo/eval/_assets/ which holds calibration
data, downloaded videos, and run dumps used by the synthetic-flow
evaluator that we don't want to track.
Authored by dataloader.
The flatten loop in VBenchPromptDataset strips exactly one level of
upstream's {dim_name: ...} wrapper. For most VBench dims this leaves
a flat scalar dict ({"color": "red"}, {"object": "person"}). For
spatial_relationship, upstream double-wraps, so one level of unwrap
leaves {"spatial_relationship": {object_a, object_b, relationship}} —
which is exactly what SpatialRelationshipMetric reads.
That last case looks accidental; it isn't. Documenting all four shapes
inline so a future reader doesn't "simplify" the wrapping away and
silently break spatial_relationship scoring.
Empirically verified output for all four aux dims matches the
documented shapes.
The Evaluator's calibrate() is gone, so batch_unit class attrs and
trial_forward() method overrides have no callers — pure cruft from the
old auto-calibration design. Mechanical removal across 15 metric files.
Internal time-dim chunking (the part of batching that actually does
work) lives on as metric-owned _chunk_size constants set in __init__,
unaffected by this change.
Five files (color/multiple_objects/object_class/spatial_relationship/
optical_flow metrics) skipped — they have dataloader's in-flight
modifications uncommitted; left for that branch's cleanup pass.
Parity verified on fs-mbz-gpu-085: results byte-identical to v3
baseline.
Match FastVideo's existing depth: VideoGenerator is the top-level
inference object with no Runner above it; loops live in scripts.
Eval should mirror that — Evaluator is the top-level scoring object,
EvalWorker × N is the layer below, and end-to-end pipelines (prompts
→ generate → score) are scripts, not classes.
Removed:
- fastvideo/eval/runner.py (EvalRunner + classmethod constructors).
- EvalRunner export from fastvideo.eval.
Added:
- fastvideo/eval/io/paths.py with sanitize_prompt, default_filename,
glob_videos, build_eval_kwargs as free functions. The reusable bits
of the runner survive; the class wrapping them does not.
Rewritten:
- scripts/eval/run_vbench_e2e.py is now top-to-bottom procedural:
parse_args → dataset → optional generate() loop → optional score()
loop. Reads in one pass; no classmethod indirection to chase.
Parity verified on fs-mbz-gpu-085 against runs/vbench_smoke/videos/
(2 videos × 3 metrics, dimensions=subject_consistency):
- 1-GPU run (results_v4.json) is byte-identical to v3 baseline.
- 2-GPU run (results_v4_2gpu.json) is byte-identical to v3 baseline
(order-preserving fan-out via Evaluator.evaluate(samples=[...])).
Continues the eval refactor. Datasets now yield plain dicts (matches
ValidationDataset / VideoGenerator / Evaluator's kwargs-flowing-through
pattern). End-to-end orchestration moves into EvalRunner so neither the
dataset nor the Evaluator owns videos_dir / fps / filename / generator
state.
Layering:
EvalRunner dataset / generator / file conventions / manifest
└── Evaluator thin dispatcher (single evaluate(), 1 or N samples)
└── EvalWorker × N single-GPU metric replicas
Dataset (fastvideo/eval/datasets/):
- PromptDataset is just an Iterable[dict]. _rows holds dicts with
prompt / n_samples / dimensions / auxiliary_info / ... — unused
fields stay absent rather than living as Optional[None] on a
dataclass. BasePromptDataset alias kept for back-compat.
- BenchmarkSample dataclass deleted. Sample TypedDict documents the
recognized keys without forcing the schema.
- VBench flattens its nested {dim: {key: val}} aux into {key: val} at
load time. Aux-using metrics already read flat keys; no metric edits.
Runner (fastvideo/eval/runner.py, new):
- Three named constructors: from_dataset, from_videos, from_samples.
- generate() drives the generator over the dataset (writes manifest);
score() loads videos and dispatches via Evaluator.evaluate(list).
run() = generate then score.
- Multi-GPU lives in the underlying Evaluator; the runner just hands
num_gpus through. No double-orchestration.
- File naming convention (filename_fn) and eval-kwargs assembly
(eval_kwargs_fn) are runner-side hooks, not dataset overrides.
Script (scripts/eval/run_vbench_e2e.py):
- Rewritten against EvalRunner; ~95 LOC, was 149.
Includes prior scaffolding from dataloader (datasets/registry.py,
the test file, the script) carried over so the refactor lands as one
coherent state.
Mirrors FastVideo's VideoGenerator → Worker layering, in-process:
Evaluator (user-facing dispatcher)
└── EvalWorker × N (single-GPU, owns metric replicas)
EvalWorker (new) holds metric replicas on one device and scores one
sample at a time. Evaluator builds num_gpus workers eagerly in __init__
(every metric loaded on every replica) and exposes a single evaluate()
method that handles both shapes:
ev.evaluate(video=..., ...) → one sample, runs on worker 0
ev.evaluate([s1, s2, ...]) → fan out across workers
Removed (dead code or moved to runner in a follow-up):
- Evaluator.calibrate, _compute_chunked, _evaluate_multi_gpu,
_gpu_metrics, evaluate_dataset, B>1 input path
- BaseMetric.batch_unit and trial_forward (kept _chunk_size as a plain
default class attr; metrics use it for internal time-dim chunking)
- memory.is_batch_too_large, slice_sample (clear_cache stays)
Per-metric batch_unit/trial_forward overrides remain as harmless
class-attr cruft pending a mechanical removal pass.
Net: evaluator.py 447 → 135 LOC; metrics/base.py 92 → 66; memory.py
40 → 14; +78 LOC for worker.py.
Adds an in-process evaluation suite covering native (SSIM/PSNR/LPIPS/
optical_flow), audio, physics_iq, vbench (16 sub-metrics), and
videoscore2. Public API: create_evaluator/evaluate, BaseMetric +
@register, ensure_checkpoint, get_cache_dir. CLI: fastvideo eval
list/run.
VBench upstream pinned as a git submodule at
fastvideo/third_party/eval/vbench (Vchitect/VBench@45e79ec); modern-dep
compat is achieved via runtime shims in vbench/__init__.py rather than
on-disk patches. CLIP/torch.hub caches are routed through
${FASTVIDEO_EVAL_CACHE}; HF cache stays at the system default.
ensure_checkpoint delegates to fastvideo.utils.get_lock + huggingface_hub
primitives. Evaluator supports release_cuda_memory(), unload(),
reload() for training-time eval that frees GPU between calls.
Verified parity against upstream vbench (5/8 metrics bit-exact, others
within 1% drift driven by transformers/torch version skew) and against
upstream VideoScore2 (regex anchored on the actual model's output
format, soft-score formula matches upstream's argmax*max_prob/total).
Includes docs/contributing/eval-metrics.md as the porting guide.
Out of scope (deferred follow-ups): MIND, VBench-2.0, FVD as a
registered metric, training-time EvalCallback.
## Summary
Follow-up to #1187. Two small changes:
1. **Merge Protections expanded** — adds `#approved-reviews-by>=1` and `check-success~=pre-commit` to `merge_protections` so the Mergify check shows a unified requirements checklist on every PR (title format + approval + pre-commit), instead of only showing the title format.
2. **Buildkite pipeline comment fix** — updates the outdated Full Suite section comment from "Triggered by adding the 'ready' label via GitHub Actions → Buildkite API" to reflect the new Merge Queue trigger path.
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Delete Open Sora Plan Modeling
amend log validation
Change name to fast video
Remove OSP modeling
clean up name changing
remove files
Deleted unnecessary files
commit first
commit first
training
ok
Update generate_synthetic.sh and deepspeed_zero2_config.yaml
14
debug
overfitting ....
debugging ..
Debug successful!
Add zero3
OK
update optimizer
load
small bug
random seed args
sp enable & still has bug
rename
update inference sp code / can run / still has bug / don't output normal mp4
switch to deepspeed dummyoptim
fix some bugs; still output green
SP inference done!
fix typos in readme and latent dataset debug file
Update dependencies; Setup code for debugging
Delete unused files and code
Remove all npu code
Update torchvision imports
Update EMA model and t2v_debug.sh script
Delete npu related stuff and remove inpaint module
Remove compress kv
Fix warning with dataset handling and model loading
Update PyTorch index URLs and video length tolerance range
Remove UDIT and inpaint
fix typo for dataset download
include pretrained open-sora
Add pretrained model for OpenSoraT2V-ROPE-L
Update max height and width for video processing
{"name":"codebase-map","description":"High-level structural index of the FastVideo-WorldModel repository","path":"codebase-map/README.md","status":"ready","trust":"high"}
{"name":"evaluation-registry","description":"Catalog of all evaluation metrics with detailed explanations, implementation status, and usage guides","path":"evaluation-registry/README.md","status":"draft","trust":"medium"}
{"name":"experiment-journal","description":"Living log of all experiments with hypotheses, configs, metrics, and insights","path":"experiment-journal/README.md","status":"draft","trust":"medium"}
{"name":"related-work","description":"Index of related papers, repos, and blog posts with structured comparisons to FastVideo","path":"related-work/README.md","status":"draft","trust":"low"}
{"name":"launch-experiment","description":"Generate and execute a training launch command for FastVideo models","path":"launch-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"monitor-experiment","description":"Poll a running W&B training run for progress and emit structured alerts","path":"monitor-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"summarize-run","description":"Extract a W&B run summary into a structured experiment report","path":"summarize-run/SKILL.md","status":"draft","trust":"low"}
{"name":"log-experiment","description":"Append or update an experiment entry in the experiment journal","path":"log-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"evaluate-video-quality","description":"Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)","path":"evaluate-video-quality/SKILL.md","status":"draft","trust":"low"}
{"name":"index-related-work","description":"Ingest a paper or repository into the related work index","path":"index-related-work/SKILL.md","status":"draft","trust":"low"}
{"name":"search-related-work","description":"Query the related work index for relevant papers, repos, or comparisons","path":"search-related-work/SKILL.md","status":"draft","trust":"low"}
{"name":"seed-ssim-references","description":"Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline","path":"seed-ssim-references/SKILL.md","status":"draft","trust":"low"}
description:Seed HF reference videos for a single newly-added SSIM test. Runs the test on Modal L40S, downloads the generated mp4s via `modal volume get`, pauses for the user to eyeball quality, then uploads only that test's files to `FastVideo/ssim-reference-videos`. Use when a new `fastvideo/tests/ssim/test_*_similarity.py` has just been added and has no references on HF yet.
---
# Seed SSIM Reference Videos
## Purpose
A brand-new SSIM test in `fastvideo/tests/ssim/` fails forever until its
reference videos exist on the HF dataset (`FastVideo/ssim-reference-videos`).
This skill:
1. Runs the test on Modal's L40S pool to generate the videos.
2. Downloads them to the local repo via `modal volume get`.
3. Pauses so the user can eyeball the mp4s and confirm quality.
4. Uploads only the new test's files to HF, with a guard that refuses to
overwrite anything already present.
The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an mp4 without crashing. The skill does not re-test locally; it
goes straight to Modal L40S (which is what CI uses).
## When to use
- A new `test_*_similarity.py` file has been added in `fastvideo/tests/ssim/`
and the HF dataset has no `reference_videos/default/L40S_reference_videos/<model_id>/`
subtree for it yet.
## When not to use
- Regular CI runs — once refs exist, `pytest fastvideo/tests/ssim/` downloads
them automatically.
- Re-seeding an existing test. That requires `--force` on the upload step, and
is out of scope here; treat as a separate, deliberate operation.
## Inputs
The skill has **one required input**: the path to the new SSIM test file.
Prompt the user for it if they didn't supply it.
| Parameter | Required | Description |
|-----------|----------|-------------|
| `test_file` | Yes | e.g. `fastvideo/tests/ssim/test_ltx2_similarity.py`. The skill's first action is to ask for this if missing. |
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
The extra `generated_videos/` level comes from the volume layout in
`_sync_generated_videos_to_volume` (`ssim_test.py`) — the command copies
`<repo>/fastvideo/tests/ssim/generated_videos/<tier>` to
`ssim_generated_videos/<tier>/<SUBDIR>/generated_videos/`, and `modal volume
get` preserves that trailing `generated_videos/` segment.
### 4. PAUSE — user reviews quality
Print the list of downloaded mp4s and their paths, then stop. Tell the user:
> "Generated videos downloaded to `./generated_videos_modal/default/generated_videos/L40S_reference_videos/`. Please open them and confirm the quality looks correct. Reply **`upload`** to continue, or anything else to abort."
Do not proceed until the user explicitly says `upload`. If they abort, leave
everything on disk so they can inspect further — no cleanup.
### 5. Copy into the local reference layout
Scoped copy — only the new test's mp4s. Loop over each `<model_id>` extracted
| 2026-04-21 | Post-first-run fixes: `modal volume get` needs `--force` when parent exists; download tree has an extra `generated_videos/` level so `--generated-dir` must reflect it. |
"neg_prompt":"Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
},
"test_prompts":[
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
VALID="encoder vae transformer kernel unit ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
- Naming: `snake_case` for functions/files, `PascalCase` for classes, `UPPER_SNAKE_CASE` for constants.
## Testing Guidelines
- Use `pytest` and place tests near relevant domains (e.g., `fastvideo/tests/encoders/`).
- Prefer descriptive names like `test_<feature>_<expected_behavior>.py`.
- For new pipelines/backends, include at least one regression-oriented test; add SSIM coverage when output quality must be preserved.
- Document GPU assumptions in tests that require specific hardware.
## Commit & Pull Request Guidelines
- Follow existing commit style: short subject with optional tag prefix, e.g. `[bugfix]: ...`, `[feat]: ...`, `[misc]: ...`, and include PR reference like `(#1234)` when applicable.
- Keep commits focused by concern (feature, refactor, fix).
- PRs should include:
- clear problem/solution summary,
- test evidence (`pytest`/SSIM outputs or rationale if skipped),
- linked issue/PR context,
- screenshots or sample outputs for UI/demo/docs changes.
## Agent Infrastructure
This repository is agent-friendly. Before doing any work, read:
1.`.agents/onboarding/README.md` — full onboarding guide with step-by-step instructions.
2.`.agents/memory/codebase-map/README.md` — structural index of the entire repository.
3.`.agents/skills/` — available agent skills (check if one exists before writing code).
4.`.agents/workflows/` — SOPs for common procedures (experiment lifecycle, evaluation, etc.).
5.`.agents/lessons/` — known pitfalls and their documented fixes.
If you are exploring a new procedure that has no existing SOP, document your
progress in `.agents/exploration/` and flag it for review at the end of your
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
-`2026/03/17`: Release Live demo: [Into the Dreamverse: Vibe Directing in FastVideo](https://dreamverse.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/dreamverse/).
-`2026/03/13`: Release Live demo: [Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU](https://1080p.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
-`2025/08/04`: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
### More News
-`2025/06/14`: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389).
-`2025/04/24`: [FastVideo V1](https://hao-ai-lab.github.io/blogs/fastvideo/) is released!
-`2025/02/18`: Release the inference code for [Sliding Tile Attention](https://hao-ai-lab.github.io/blogs/sta/).
## Key Features
FastVideo has the following features:
- End-to-end post-training support for bidirectional and autoregressive models:
- Support full finetuning and LoRA finetuning for state-of-the-art open video DiTs
- Data preprocessing pipeline for video, image, and text data
- Distribution Matching Distillation (DMD2) stepwise distillation.
- Sparse attention with [Video Sparse Attention](https://arxiv.org/pdf/2505.13389)
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
- User-friendly CLI and Python API
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/optimizations/) for full list of supported optimizations.
- Diverse hardware and OS support
- Support H100, A100, 4090
- Support Linux, Windows, MacOS
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
## Getting Started
We recommend using [uv](https://docs.astral.sh/uv/) to create a clean environment. If you previously used Conda, switching to uv generally gives faster and more stable installs.
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
## Sparse Distillation
For our sparse distillation techniques, please see our [distillation docs](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/) and check out our [blog](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/attention/vsa/#installation). Create a file called `example.py` with the following code:
# Create a video generator with a pre-trained model
generator=VideoGenerator.from_pretrained(
"FastVideo/FastWan2.1-T2V-1.3B-Diffusers",
num_gpus=1,# Adjust based on your hardware
)
# Define a prompt for your video
prompt="A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest."
# Generate the video
video=generator.generate_video(
prompt,
output_path="my_videos/",# Controls where videos are saved
save_video=True
)
if__name__=='__main__':
main()
```
## Prepare Data & Models
We've prepared some debug data to facilitate development. To make sure the training pipeline is correct, train on the debug data and make sure the model overfit on it (feed it the same text prompt and see if the output video is the same as the training data)
## Awesome work using FastVideo or our research projects
- [SGLang](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen): SGLang's diffusion inference functionality is based on a fork of FastVideo on Sept. 24, 2025.
- [DanceGRPO](https://github.com/XueZeyue/DanceGRPO): A unified framework to adapt Group Relative Policy Optimization (GRPO) to visual generation paradigms. Code based on FastVideo.
- [SRPO](https://github.com/Tencent-Hunyuan/SRPO): A method to directly align the full diffusion trajectory with fine-grained human preference. Code based on FastVideo.
- [DCM](https://github.com/Vchitect/DCM): Dual-expert consistency model for efficient and high-quality video generation. Code based on FastVideo.
- [HY-WorldPlay](https://github.com/Tencent-Hunyuan/HY-WorldPlay): An action-conditioned world model model trained using FastVideo framework.
- [Hunyuan Video 1.5](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5): A leading lightweight video generation model, where they proposed SSTA based on Sliding Tile Attention.
- [Kandinsky-5.0](https://github.com/kandinskylab/kandinsky-5): A family of diffusion models for video & image generation, where their NABLA attention includes a Sliding Tile Attention branch.
- [LongCat Video](https://github.com/meituan-longcat/LongCat-Video): A foundational video generation model with 13.6B parameters with block-sparse attention similar to Video Sparse Attention.
17. no cfg, validation no cfg, pcm_linear_quadratic, euler_steps 50, 0.1, linear_range 0.75
We welcome all contributions. Please check out our guide [here](https://hao-ai-lab.github.io/FastVideo/contributing/overview/).
See details in [development roadmap](https://github.com/hao-ai-lab/FastVideo/issues/899).
## Acknowledgement
We learned the design and reused code from the following projects: [Wan-Video](https://github.com/Wan-Video), [ThunderKittens](https://github.com/HazyResearch/ThunderKittens), [DMD2](https://github.com/tianweiy/DMD2), [diffusers](https://github.com/huggingface/diffusers), [xDiT](https://github.com/xdit-project/xDiT), [vLLM](https://github.com/vllm-project/vllm), [SGLang](https://github.com/sgl-project/sglang). We thank [MBZUAI](https://ifm.mbzuai.ac.ae/), [Anyscale](https://www.anyscale.com/), and [GMI Cloud](https://www.gmicloud.ai/) for their support throughout this project.
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.