Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c2e89f22d3 | ||
|
|
e47a3c5aad | ||
|
|
d40fbfc534 | ||
|
|
126a52ce32 | ||
|
|
419e1c68f1 | ||
|
|
938bc3c972 | ||
|
|
381a7aae66 | ||
|
|
7fa4fb50bb | ||
|
|
de3cd6aab2 | ||
|
|
a64e5e62ab | ||
|
|
f4200cc3f3 | ||
|
|
363230b372 | ||
|
|
1dcdacc35a | ||
|
|
0a7a4a21f6 | ||
|
|
c30d731992 | ||
|
|
f8eaff4292 | ||
|
|
76e7048b3d | ||
|
|
b1386a7a78 | ||
|
|
664d6b3c23 | ||
|
|
0688c5131a | ||
|
|
0b605d0a41 | ||
|
|
8f5ebe2aeb | ||
|
|
84803076a0 | ||
|
|
5cc337fdb6 | ||
|
|
2b8e5a56a8 | ||
|
|
e117167eb9 | ||
|
|
4f9af56b78 | ||
|
|
33efe7a673 | ||
|
|
b07cc3f9d4 | ||
|
|
29eb4109cc | ||
|
|
d07a7691fe | ||
|
|
960485519f | ||
|
|
412c95b1a0 | ||
|
|
f6275f8005 |
@@ -4,21 +4,24 @@ description: How to develop, validate, and register a new evaluation metric
|
||||
|
||||
# Evaluation Development SOP
|
||||
|
||||
Standard procedure for adding new video quality evaluation metrics to the
|
||||
FastVideo agent toolkit.
|
||||
Standard procedure for adding new video quality evaluation metrics to
|
||||
the FastVideo agent toolkit.
|
||||
|
||||
## When to Use
|
||||
## When to use
|
||||
|
||||
- You need a metric that doesn't exist in `.agents/memory/evaluation-registry/README.md`.
|
||||
- You need a metric that does not exist in
|
||||
`.agents/memory/evaluation-registry/README.md`.
|
||||
- An existing metric needs significant changes to its methodology.
|
||||
- You're exploring a new evaluation approach.
|
||||
- You are exploring a new evaluation approach.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Research
|
||||
|
||||
- Search `.agents/memory/related-work/` for existing evaluation approaches.
|
||||
- Check the `evaluation_registry.md` for current metrics and their limitations.
|
||||
- Search `.agents/memory/related-work/` for existing evaluation
|
||||
approaches.
|
||||
- Check `.agents/memory/evaluation-registry/README.md` for current
|
||||
metrics and their limitations.
|
||||
- Review literature: FVD, CLIP-Score, human preference, etc.
|
||||
|
||||
### 2. Prototype
|
||||
@@ -29,21 +32,25 @@ FastVideo agent toolkit.
|
||||
|
||||
### 3. Validate
|
||||
|
||||
- **Known-good test**: Metric should score high on reference-quality videos.
|
||||
- **Known-bad test**: Metric should score low on degraded/unrelated videos.
|
||||
- **Sensitivity test**: Small quality differences should produce meaningful
|
||||
score differences.
|
||||
- **Known-good test**: metric should score high on reference-quality
|
||||
videos.
|
||||
- **Known-bad test**: metric should score low on degraded or unrelated
|
||||
videos.
|
||||
- **Sensitivity test**: small quality differences should produce
|
||||
meaningful score differences.
|
||||
- Document thresholds and their justification.
|
||||
|
||||
### 4. Register
|
||||
|
||||
Update `.agents/memory/evaluation-registry/README.md`:
|
||||
|
||||
- Add the metric with status `Active`.
|
||||
- Document location, thresholds, and trust level.
|
||||
|
||||
### 5. Integrate
|
||||
|
||||
Update `.agents/skills/evaluate-video-quality.md`:
|
||||
Update `.agents/skills/evaluate-video-quality/SKILL.md`:
|
||||
|
||||
- Add the new metric as a section.
|
||||
- Include code examples and interpretation guide.
|
||||
|
||||
@@ -52,3 +59,35 @@ Update `.agents/skills/evaluate-video-quality.md`:
|
||||
- Move the exploration log content into the skill.
|
||||
- Clean up the exploration file or mark it as `promoted`.
|
||||
- If anything went wrong during development, create a lesson.
|
||||
|
||||
## Where the metrics live
|
||||
|
||||
The eval suite is `fastvideo/eval/`. New metrics register themselves
|
||||
via `@register("<group>.<name>")` and are auto-discovered when
|
||||
`fastvideo.eval.metrics` is imported.
|
||||
|
||||
- **Native metrics** (SSIM, PSNR, LPIPS, optical flow, VLM): add a
|
||||
file under the appropriate group dir
|
||||
(`fastvideo/eval/metrics/common/`, `optical_flow/`, `videoscore2/`,
|
||||
`physics_iq/`).
|
||||
- **Metrics that wrap upstream research code**: follow the vbench
|
||||
pattern in `fastvideo/eval/metrics/vbench/`. The contract is:
|
||||
- Upstream lives as a git submodule under
|
||||
`fastvideo/third_party/eval/<bench>/`, pinned to a SHA in repo-root
|
||||
`.gitmodules`.
|
||||
- The metric package's `__init__.py` inserts the submodule path on
|
||||
`sys.path` and installs runtime compat shims (attribute-level
|
||||
monkey-patches) for any modern-dep drift. Do not modify upstream
|
||||
files on disk, and do not ship a `setup.sh`.
|
||||
- See `fastvideo/eval/README.md` for the worked vbench example.
|
||||
- Full porting guide:
|
||||
[`docs/contributing/eval-metrics.md`](../../docs/contributing/eval-metrics.md).
|
||||
|
||||
## Out of scope of the initial eval port
|
||||
|
||||
The following land in follow-up PRs:
|
||||
|
||||
- **MIND** metrics (depends on a separate `vipe` submodule).
|
||||
- **VBench-2.0** sibling package.
|
||||
- Native conversion of **FVD** under `fastvideo/eval/metrics/fvd/`.
|
||||
- The training-time `EvalCallback`.
|
||||
|
||||
@@ -4,3 +4,6 @@
|
||||
[submodule "fastvideo-kernel/include/cutlass"]
|
||||
path = fastvideo-kernel/include/cutlass
|
||||
url = https://github.com/NVIDIA/cutlass.git
|
||||
[submodule "fastvideo/third_party/eval/vbench"]
|
||||
path = fastvideo/third_party/eval/vbench
|
||||
url = https://github.com/Vchitect/VBench.git
|
||||
|
||||
@@ -0,0 +1,559 @@
|
||||
# Porting Eval Metrics into `fastvideo.eval`
|
||||
|
||||
This guide is for contributors adding new evaluation metrics to
|
||||
FastVideo's eval suite. To run the existing metrics, see
|
||||
[`fastvideo/eval/README.md`](../../fastvideo/eval/README.md).
|
||||
|
||||
## When to use this guide
|
||||
|
||||
Use this guide when you are:
|
||||
|
||||
- Adding a new metric (native or wrapping a third-party library).
|
||||
- Porting a benchmark (e.g. VBench, MIND, EvalCrafter) whose Python
|
||||
code needs to be importable from a pinned upstream.
|
||||
- Adding a new metric group (audio, vlm, etc.).
|
||||
|
||||
## TL;DR
|
||||
|
||||
Metrics are auto-discovered from
|
||||
`fastvideo/eval/metrics/<group>/<name>/metric.py`. Each declares itself
|
||||
with `@register("<group>.<name>")` and subclasses `BaseMetric`. Three
|
||||
recipes:
|
||||
|
||||
1. **Native metric** (pure-PyTorch, no submodule). Drop a file,
|
||||
declare deps, implement `compute(sample)`.
|
||||
2. **Library-wrapped metric** (CLIP, torch.hub, transformers, pyiqa).
|
||||
Same as above, plus route the library's cache through
|
||||
`get_cache_dir()` if it has a `download_root=` / `cache_dir=`
|
||||
kwarg.
|
||||
3. **Upstream-submodule-wrapped metric** (vbench-style). Pin upstream
|
||||
as a git submodule under `fastvideo/third_party/eval/<bench>/`. The
|
||||
adapter `__init__.py` does the `sys.path` insert and any runtime
|
||||
compat shims for modern dep versions. Patches live as Python in
|
||||
that file rather than as on-disk patches to the submodule.
|
||||
|
||||
The full recipes are below.
|
||||
|
||||
---
|
||||
|
||||
## 0) Layout and auto-discovery
|
||||
|
||||
```
|
||||
fastvideo/eval/metrics/
|
||||
├── base.py # BaseMetric + lifecycle contract
|
||||
├── common/ # group: SSIM, PSNR, LPIPS
|
||||
├── optical_flow/ # group: gt_optical_flow, synthetic_optical_flow
|
||||
├── vlm/ # group: VideoScore-2
|
||||
├── physics_iq/ # group + sub-metrics
|
||||
└── vbench/ # group: 16 sub-metrics
|
||||
├── __init__.py # sys.path bootstrap + runtime compat shims
|
||||
├── _grit_helper.py # shared upstream-touching helpers
|
||||
└── <sub_metric>/metric.py
|
||||
```
|
||||
|
||||
Auto-discovery (`fastvideo/eval/metrics/__init__.py`) walks each group
|
||||
dir and imports every `metric.py` it finds, which fires the
|
||||
`@register` decorators. Names starting with `_` are skipped. Use that
|
||||
prefix for shared helpers or vendored code that should not register
|
||||
itself.
|
||||
|
||||
---
|
||||
|
||||
## 1) The `BaseMetric` contract
|
||||
|
||||
Every metric subclasses `fastvideo.eval.metrics.base.BaseMetric` and
|
||||
declares:
|
||||
|
||||
```python
|
||||
class YourMetric(BaseMetric):
|
||||
name: str = "common.your_metric" # must match @register
|
||||
requires_reference: bool = True # needs sample["reference"]
|
||||
higher_is_better: bool = True # for ranking / aggregates
|
||||
dependencies: list[str] = [] # importable module names;
|
||||
# registry surfaces a clean
|
||||
# ImportError if missing
|
||||
needs_gpu: bool = False
|
||||
backbone: str | None = None # e.g. "clip_vit_l14"
|
||||
```
|
||||
|
||||
You must implement:
|
||||
|
||||
```python
|
||||
def compute(self, sample: dict) -> list[MetricResult]:
|
||||
"""sample['video'] is (1, T, C, H, W). Return a one-element list.
|
||||
|
||||
The leading 1 is preserved for forward-compat with batched eval;
|
||||
today :class:`EvalWorker` always invokes metrics with B=1.
|
||||
"""
|
||||
```
|
||||
|
||||
You may override:
|
||||
|
||||
- `setup(self) -> None`. Eager model loading. Called once by
|
||||
`create_evaluator`. Idempotent (re-entrant). Use the `if self._model
|
||||
is not None: return` pattern.
|
||||
- `to(self, device)`. Move the metric and its submodels to `device`.
|
||||
|
||||
If a required input is missing (e.g. an fps-aware metric called
|
||||
without `fps`), return `self._skip(sample, reason)` instead of
|
||||
raising.
|
||||
|
||||
---
|
||||
|
||||
## 2) Recipe A: native metric (no external deps)
|
||||
|
||||
Smallest case. Pixel math, simple closed-form.
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/common/your_metric/metric.py
|
||||
from __future__ import annotations
|
||||
import torch
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("common.your_metric")
|
||||
class YourMetric(BaseMetric):
|
||||
name = "common.your_metric"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
needs_gpu = False
|
||||
dependencies: list[str] = [] # nothing extra
|
||||
|
||||
def compute(self, sample: dict) -> list[MetricResult]:
|
||||
gen, ref = sample["video"], sample["reference"] # (B,T,C,H,W) each
|
||||
per_video = ((gen - ref) ** 2).mean(dim=(1, 2, 3, 4)).sqrt()
|
||||
return [
|
||||
MetricResult(name=self.name, score=float(s), details={})
|
||||
for s in per_video
|
||||
]
|
||||
```
|
||||
|
||||
That is the whole recipe. Drop the file and the registry picks it up.
|
||||
|
||||
---
|
||||
|
||||
## 3) Recipe B: library-wrapped metric (CLIP, torch.hub, transformers, pyiqa)
|
||||
|
||||
If your metric loads a backbone from a Python package, route the
|
||||
library at the eval cache so users get one knob (`FASTVIDEO_EVAL_CACHE`)
|
||||
to redirect everything.
|
||||
|
||||
### Cache routing rules
|
||||
|
||||
| Library | How to route | Location after redirect |
|
||||
|---|---|---|
|
||||
| `clip.load("ViT-X")` | pass `download_root=str(get_cache_dir() / "clip")` | `${FASTVIDEO_EVAL_CACHE}/clip/` |
|
||||
| `torch.hub.load(...)` | nothing; `TORCH_HOME` is redirected at `fastvideo.eval` import time | `${FASTVIDEO_EVAL_CACHE}/torch/hub/` |
|
||||
| `transformers.from_pretrained(...)` | nothing; leave HF's default cache (`~/.cache/huggingface/hub/`) so users dedupe with other ML projects | `~/.cache/huggingface/hub/` |
|
||||
| `huggingface_hub.snapshot_download` / `hf_hub_download` | use `ensure_checkpoint(...)` (it wraps these with filelock) | same as above |
|
||||
| `pyiqa.create_metric(...)` | no env var or kwarg honored; document in metric docstring | pyiqa-internal |
|
||||
| `lpips`, `ptlflow` | torch.hub-based, auto-redirected | `${FASTVIDEO_EVAL_CACHE}/torch/hub/` |
|
||||
| Raw URL (no HF Hub) | use `ensure_checkpoint(name, source="https://...")` | `${FASTVIDEO_EVAL_CACHE}/models/<name>` |
|
||||
| Dataset asset (raw video/mask/image) auto-fetched from a public bucket | download into `get_cache_dir() / "datasets" / "<bench>"`, mirroring upstream's relative layout. Vendor any small manifest (CSV/JSON ≤1 MB) under the metric folder so the dataset can be used without external setup. | `${FASTVIDEO_EVAL_CACHE}/datasets/<bench>/` |
|
||||
|
||||
### Dataset assets: vendor the manifest, auto-fetch the rest
|
||||
|
||||
If your metric ships with its own paired-reference dataset (Physics-IQ
|
||||
is the canonical example), follow this layout:
|
||||
|
||||
- **Manifest** (CSV/JSON ≤1 MB): vendor it under
|
||||
`fastvideo/eval/metrics/<bench>/_vendored/<manifest>.<ext>`, with a
|
||||
sibling `_vendored/LICENSE` recording attribution and provenance.
|
||||
The `_vendored/` subdir is the project-wide convention for
|
||||
upstream-provenance files: it is auto-skipped by metric discovery
|
||||
(the `_` prefix) and by codespell (one `*/_vendored/*` glob in
|
||||
`[tool.codespell].skip`), so dropping in a new vendored file
|
||||
requires no further config. Read the manifest from the dataset
|
||||
module via a `Path(__file__)`-relative resolver. Mirror
|
||||
`_VENDORED_DESCRIPTIONS_CSV` in
|
||||
`fastvideo/eval/datasets/physics_iq.py`.
|
||||
- **Heavy assets** (videos, masks, images): do not vendor. Auto-fetch
|
||||
on first miss into `get_cache_dir() / "datasets" / "<bench>"`,
|
||||
mirroring upstream's relative directory layout one-for-one so a
|
||||
pre-downloaded mirror at any path works as a drop-in
|
||||
`dataset_root=`. Use atomic `.part` then final-rename to be safe
|
||||
under concurrent SLURM ranks.
|
||||
- **Bucket override**: expose `FASTVIDEO_<BENCH>_BUCKET_URL` so users
|
||||
with internal mirrors can redirect.
|
||||
- **Opt-out**: accept `auto_download: bool = True` in the dataset
|
||||
constructor; on `False`, raise `FileNotFoundError` instead of
|
||||
fetching. This covers air-gapped runs and CI.
|
||||
|
||||
The end-state is `get_dataset("<bench>")` with no kwargs.
|
||||
|
||||
### Example: CLIP backbone + LAION head
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/your_group/your_metric/metric.py
|
||||
from __future__ import annotations
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("your_group.your_metric")
|
||||
class YourMetric(BaseMetric):
|
||||
name = "your_group.your_metric"
|
||||
requires_reference = False
|
||||
needs_gpu = True
|
||||
dependencies = ["clip"] # "openai-clip" PyPI; importable as `clip`
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._clip = None
|
||||
self._head = None
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._clip is not None:
|
||||
return
|
||||
import clip
|
||||
from fastvideo.eval.models import ensure_checkpoint, get_cache_dir
|
||||
|
||||
# Backbone: route CLIP's cache through our root.
|
||||
self._clip, _ = clip.load(
|
||||
"ViT-L/14",
|
||||
device=self.device,
|
||||
download_root=str(get_cache_dir() / "clip"),
|
||||
)
|
||||
self._clip.eval()
|
||||
|
||||
# URL-fetched head: ensure_checkpoint downloads to
|
||||
# ${FASTVIDEO_EVAL_CACHE}/models/ with filelock + atomic rename.
|
||||
ckpt = ensure_checkpoint(
|
||||
"your_head.pth",
|
||||
source="https://example.com/path/to/your_head.pth",
|
||||
)
|
||||
self._head = nn.Linear(768, 1)
|
||||
self._head.load_state_dict(
|
||||
torch.load(ckpt, map_location="cpu", weights_only=True)
|
||||
)
|
||||
self._head.to(self.device).eval()
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._clip is not None:
|
||||
self._clip = self._clip.to(self.device)
|
||||
if self._head is not None:
|
||||
self._head = self._head.to(self.device)
|
||||
return self
|
||||
|
||||
def compute(self, sample: dict) -> list[MetricResult]:
|
||||
...
|
||||
```
|
||||
|
||||
### Do not redirect other `~/.cache/...` dirs
|
||||
|
||||
If a third-party library hard-codes `~/.cache/<lib>/` and offers no
|
||||
override, document the exception in the metric's docstring. Forcing
|
||||
redirection by setting `os.environ` or patching `os.path.expanduser`
|
||||
is fragile and breaks user expectations of where the library's cache
|
||||
lives.
|
||||
|
||||
---
|
||||
|
||||
## 4) Recipe C: upstream-submodule-wrapped metric (vbench pattern)
|
||||
|
||||
Use this when the upstream benchmark ships Python code (`vbench/`,
|
||||
`MIND/`, etc.) that is not pip-installable cleanly. See
|
||||
`fastvideo/eval/metrics/vbench/__init__.py` for the worked example.
|
||||
|
||||
### 4.1 Pin the upstream as a submodule
|
||||
|
||||
```bash
|
||||
git submodule add <upstream-url> fastvideo/third_party/eval/<bench>
|
||||
cd fastvideo/third_party/eval/<bench>
|
||||
git checkout <pinned-sha>
|
||||
cd -
|
||||
git add .gitmodules fastvideo/third_party/eval/<bench>
|
||||
```
|
||||
|
||||
The submodule pulls under the standard `git submodule update --init
|
||||
--recursive` flow that users already run for kernel deps.
|
||||
|
||||
### 4.2 Bootstrap on `sys.path`
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/<bench>/__init__.py
|
||||
from __future__ import annotations
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
# fastvideo/eval/metrics/<bench>/__init__.py → ../../../../third_party/eval/<bench>
|
||||
_UPSTREAM = Path(__file__).resolve().parents[3] / "third_party" / "eval" / "<bench>"
|
||||
if _UPSTREAM.is_dir() and str(_UPSTREAM) not in sys.path:
|
||||
sys.path.insert(0, str(_UPSTREAM))
|
||||
```
|
||||
|
||||
We do not `pip install` the upstream because its egg-link/.pth would
|
||||
just re-do this `sys.path.insert`, and skipping the install also skips
|
||||
the upstream's `setup.py` (which often gates on a specific CUDA
|
||||
version).
|
||||
|
||||
### 4.3 Modern-dep compat: runtime shims
|
||||
|
||||
Upstream code pinned to e.g. `transformers==4.33.2`, `numpy<2`
|
||||
typically breaks against modern versions in 3-4 known places (API
|
||||
renames). Fix those at import time, in the same `__init__.py`:
|
||||
|
||||
```python
|
||||
def _install_compat_shims() -> None:
|
||||
# Example: transformers.modeling_utils API moved.
|
||||
try:
|
||||
import transformers.modeling_utils as _mu
|
||||
import transformers.pytorch_utils as _pu
|
||||
for _n in ("apply_chunking_to_forward",
|
||||
"find_pruneable_heads_and_indices",
|
||||
"prune_linear_layer"):
|
||||
if not hasattr(_mu, _n) and hasattr(_pu, _n):
|
||||
setattr(_mu, _n, getattr(_pu, _n))
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
# Example: numpy.lib.function_base.disp removed in numpy>=2.
|
||||
try:
|
||||
import types, numpy.lib as _nl
|
||||
if not hasattr(_nl, "function_base"):
|
||||
_stub = types.ModuleType("numpy.lib.function_base")
|
||||
_stub.disp = lambda *a, **k: None
|
||||
sys.modules["numpy.lib.function_base"] = _stub
|
||||
_nl.function_base = _stub
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
_install_compat_shims()
|
||||
```
|
||||
|
||||
For function-level patches that cannot be expressed as attribute
|
||||
writes (e.g. wrapping a model factory function), use a
|
||||
`sys.meta_path` finder that wraps the loader. See
|
||||
`_install_modeling_finetune_hook()` in
|
||||
`fastvideo/eval/metrics/vbench/__init__.py` for the pattern (about 30
|
||||
lines).
|
||||
|
||||
Why shims rather than `git apply` patches: patches go stale when the
|
||||
upstream SHA changes; shims are versioned Python code in our repo,
|
||||
they are grep-able, and they only run if the targeted module is
|
||||
imported.
|
||||
|
||||
### 4.4 Per-sub-metric files
|
||||
|
||||
Each sub-metric is a normal `BaseMetric` subclass that imports from
|
||||
the upstream:
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/<bench>/<sub>/metric.py
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("<bench>.<sub>")
|
||||
class YourSubMetric(BaseMetric):
|
||||
...
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
# The sys.path bootstrap fired when fastvideo.eval.metrics.<bench>
|
||||
# was imported (which auto-discovery does before importing this
|
||||
# sub-package). Upstream imports just work:
|
||||
from <bench>.something import SomeModel
|
||||
...
|
||||
```
|
||||
|
||||
### 4.5 Conditional registration when the upstream is missing
|
||||
|
||||
If a user installed `fastvideo[eval]` but did not run `git submodule
|
||||
update --init`, `<bench>.*` metrics should not register. The
|
||||
auto-discovery walker imports each sub-package's `metric` module; have
|
||||
that import bail out cleanly:
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/<bench>/__init__.py — at the bottom
|
||||
_AVAILABLE = (_UPSTREAM / "<bench>" / "__init__.py").is_file()
|
||||
```
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/<bench>/<sub>/__init__.py
|
||||
from fastvideo.eval.metrics.<bench> import _AVAILABLE
|
||||
if _AVAILABLE:
|
||||
from .metric import YourSubMetric # noqa
|
||||
```
|
||||
|
||||
`fastvideo eval list` then reflects what the user actually has rather
|
||||
than what they could have.
|
||||
|
||||
### 4.6 Leave the upstream alone unless it blocks the metric
|
||||
|
||||
The upstream is pinned. If a metric works against the pinned SHA,
|
||||
leave the upstream files untouched. If it actively breaks against
|
||||
modern deps (the import-drift cases above), shim it. Avoid
|
||||
fastvideo-side forks of upstream code; they make patches go stale and
|
||||
parity drift.
|
||||
|
||||
---
|
||||
|
||||
## 5) Model checkpoints: `ensure_checkpoint`
|
||||
|
||||
Use `ensure_checkpoint(name, source, filename=None)` for any
|
||||
non-package weights. It resolves a local path, downloading on miss,
|
||||
with filelock safety across processes and SLURM ranks.
|
||||
|
||||
| `source` form | What happens |
|
||||
|---|---|
|
||||
| `"/abs/path/to/file.pth"` | passthrough, returned unchanged |
|
||||
| `"https://..."` | downloaded to `${FASTVIDEO_EVAL_CACHE}/models/<name>` via `huggingface_hub.http_get`, atomic rename, filelock |
|
||||
| `"org/repo"` (no `filename`) | `snapshot_download(repo_id)` → `~/.cache/huggingface/hub/` |
|
||||
| `"org/repo"` (with `filename`) | `hf_hub_download(repo_id, filename)` → `~/.cache/huggingface/hub/` |
|
||||
|
||||
`name` is only used as the local filename for URL sources. HF sources
|
||||
ignore it (HF manages its own cache key by content hash).
|
||||
|
||||
```python
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
|
||||
# URL: name matters
|
||||
ckpt = ensure_checkpoint(
|
||||
"amt-s.pth",
|
||||
source="https://huggingface.co/lalala125/AMT/resolve/main/amt-s.pth",
|
||||
)
|
||||
|
||||
# HF single file: name is decorative
|
||||
ckpt = ensure_checkpoint(
|
||||
"raft-things.pth", # ignored; HF cache uses repo+sha
|
||||
source="OpenGVLab/VBench_Used_Models",
|
||||
filename="raft-things.pth",
|
||||
)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6) Declaring `dependencies`
|
||||
|
||||
Set `dependencies = ["pkg1", "pkg2"]` on your metric class with
|
||||
importable module names (not PyPI distribution names). The registry
|
||||
checks each via `importlib.util.find_spec` at instantiation time and
|
||||
raises a clean `ImportError` pointing the user at the right install
|
||||
extra:
|
||||
|
||||
```python
|
||||
class YourMetric(BaseMetric):
|
||||
dependencies = ["clip", "timm"] # importable as `import clip`, `import timm`
|
||||
```
|
||||
|
||||
If a dep is in `[project.optional-dependencies.eval-<group>]`, you do
|
||||
not need to do anything more. If it is a new dep, add it to that
|
||||
group in `pyproject.toml`.
|
||||
|
||||
---
|
||||
|
||||
## 7) Common gotchas
|
||||
|
||||
- The standard `git submodule update --init --recursive` is enough
|
||||
for a benchmark; do not write a `setup.sh`. Modern-dep compat goes
|
||||
into your `__init__.py` as runtime shims.
|
||||
- Do not modify upstream files on disk. The submodule should always
|
||||
match its pinned SHA. Compat lives in our `__init__.py`.
|
||||
- Do not pip-install the upstream. The egg-link is a glorified
|
||||
`sys.path.insert`, which we do directly in `__init__.py`.
|
||||
- Do not call `torch.hub.set_dir(...)` from your metric. It is done
|
||||
globally in `fastvideo/eval/__init__.py`.
|
||||
- Do not put cache-redirection env vars in your metric's `setup()`.
|
||||
By the time `setup()` runs, the library has likely already cached
|
||||
the default-location decision. Set env vars at package-init time.
|
||||
- Skip rather than raise when an input is missing. Use
|
||||
`self._skip(sample, reason)` for any expected-missing input. It
|
||||
returns a list of `MetricResult(score=None)` so other metrics in
|
||||
the same evaluator continue.
|
||||
- Watch for upstream re-registration conflicts. If the upstream uses
|
||||
a global registry (detectron2's `META_ARCH_REGISTRY`, MMCV, etc.),
|
||||
loading the same model twice in the same process will throw. The
|
||||
evaluator already loads each metric once; if you write a custom
|
||||
setup-then-call pattern, mirror that single-load discipline.
|
||||
|
||||
---
|
||||
|
||||
## 8) Training-time eval: keep evaluators hot, free caches between calls
|
||||
|
||||
When wiring eval into a training loop, the working pattern is:
|
||||
|
||||
1. Construct the `Evaluator` once and attach it to the pipeline
|
||||
(`self._eval = create_evaluator(...)`). Do not recreate it per
|
||||
validation round; that re-pays the model load cost.
|
||||
2. Save validation videos to disk (the diffusion path already does
|
||||
this). Pass paths to `evaluator.evaluate`, not in-memory tensors
|
||||
that share GPU memory with the training model.
|
||||
3. Run validation only on rank 0 of each sequence-parallel group.
|
||||
Gather paths from other ranks and let rank 0 score everything.
|
||||
4. After every `evaluate(...)` call, call
|
||||
`evaluator.release_cuda_memory()` in a `finally` block. That runs
|
||||
`gc.collect()` + `torch.cuda.empty_cache()` +
|
||||
`torch.cuda.ipc_collect()`. The eval model stays loaded; only
|
||||
transient activation buffers from the just-finished call get
|
||||
freed:
|
||||
|
||||
```python
|
||||
for video_path in batch:
|
||||
try:
|
||||
scores = self._eval.evaluate(video=load_video(video_path))
|
||||
finally:
|
||||
self._eval.release_cuda_memory()
|
||||
```
|
||||
|
||||
5. If memory pressure spikes (rare on H200), call
|
||||
`evaluator.unload()` to drop every metric reference and let the
|
||||
GPU memory be GC'd. `unload` is reversible:
|
||||
`evaluator.reload()` rebuilds the same metrics with the original
|
||||
config (re-paying the model load cost). Calling `evaluate`
|
||||
between `unload` and `reload` raises a clear `RuntimeError`.
|
||||
|
||||
For most metrics (sub-1 GB backbones, e.g. CLIP/DINO/RAFT/AMT) the
|
||||
eval model can stay co-resident with the training model in
|
||||
`transformer.eval()` mode without any swap. For larger ones
|
||||
(VideoScore2 at 14 GB), measure first; if it fits on the rank-0 GPU
|
||||
during validation (training model in eval mode means no
|
||||
grads/optimizer updates), keep it hot. If not, `unload` between
|
||||
rounds.
|
||||
|
||||
## 9) Local verification
|
||||
|
||||
Native and library-wrapped metrics: a single-GPU smoke is enough.
|
||||
|
||||
```python
|
||||
import torch
|
||||
from fastvideo.eval import create_evaluator
|
||||
|
||||
ev = create_evaluator(metrics=["<group>.<your_metric>"], device="cuda")
|
||||
video = torch.randn(1, 49, 3, 256, 256, device="cuda").clamp(0, 1)
|
||||
print(ev.evaluate(video=video))
|
||||
```
|
||||
|
||||
Submodule-wrapped metrics: also do a parity check against the
|
||||
upstream once. Clone upstream into a separate venv, run the same
|
||||
video through both, and expect an exact match on bit-deterministic
|
||||
metrics and ≤1% drift on backbone-heavy ones (driven by
|
||||
transformers/torch version differences).
|
||||
|
||||
For quick parity in CI: pin a tiny test video, record expected
|
||||
scores ± tolerance, and add a calibration test under
|
||||
`fastvideo/tests/eval/`.
|
||||
|
||||
---
|
||||
|
||||
## 10) When not to add a metric
|
||||
|
||||
- **Set-vs-set distribution metrics** (FVD, FID-style) do not fit
|
||||
`BaseMetric.compute(sample)` cleanly; they need a population.
|
||||
Adding them requires a stateful accumulator interface that does
|
||||
not exist yet. Open an issue first.
|
||||
- **Metrics requiring a single-GPU model larger than available
|
||||
memory.** Eval is not the place for tensor-parallel sharding;
|
||||
metrics are expected to fit on one GPU.
|
||||
- **Metrics that need `mmcv` with a conflicting CUDA ABI.** Document
|
||||
the affected sub-metrics as unsupported and skip them. Building
|
||||
isolation infrastructure (subprocess engine, per-metric venv) is
|
||||
out of scope.
|
||||
@@ -0,0 +1,100 @@
|
||||
"""Generate one LTX2 video and score it with VBench metrics.
|
||||
|
||||
The generation block is the same as
|
||||
``examples/inference/basic/basic_ltx2.py`` — same prompt, same model,
|
||||
same shape, same num_frames. After ``shutdown()`` the script loads the
|
||||
mp4 back, builds a single :class:`fastvideo.eval.Evaluator`, and runs
|
||||
the prompt-aware VBench subset that's meaningful for an arbitrary
|
||||
text→video sample.
|
||||
|
||||
The first run downloads CLIP / DINO / RAFT / AMT / ViCLIP / MUSIQ
|
||||
weights to ``~/.cache/fastvideo/eval/`` (~few GB total).
|
||||
|
||||
GPU memory caveat
|
||||
-----------------
|
||||
Scoring 1088×1920×121 with all 8 metrics needs a dedicated GPU (~80 GB).
|
||||
On a shared GPU, ``vbench.motion_smoothness`` (AMT correlation volume)
|
||||
will OOM — its memory autoscale reads ``total_memory`` rather than
|
||||
``mem_get_info()`` free memory and therefore underestimates the
|
||||
required scale-down. Drop ``motion_smoothness`` from ``METRICS`` if
|
||||
sharing, or run on a smaller-resolution generation.
|
||||
"""
|
||||
import torch
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.eval import Evaluator
|
||||
from fastvideo.eval.io import build_eval_kwargs
|
||||
|
||||
PROMPT = (
|
||||
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
|
||||
"of a woman and a man in their 30s, facing each other with serious "
|
||||
"expressions. The woman, emotional and dramatic, says softly, \"That's "
|
||||
"it... Dad's lost it. And we've lost Dad.\" The man exhales, slightly "
|
||||
"annoyed: \"Stop being so dramatic, Jess.\" A beat. He glances aside, "
|
||||
"then mutters defensively, \"He's just having fun.\" The camera slowly "
|
||||
"pans right, revealing the grandfather in the garden wearing enormous "
|
||||
"butterfly wings, waving his arms in the air like he's trying to take "
|
||||
"off. He shouts, \"Wheeeew!\" as he flaps his wings with full commitment. "
|
||||
"The woman covers her face, on the verge of tears. The tone is deadpan, "
|
||||
"absurd, and quietly tragic."
|
||||
)
|
||||
|
||||
# VBench sub-metrics meaningful for an arbitrary text→video sample
|
||||
# (just the generated frames, optionally fps + the source prompt).
|
||||
# Structured-prompt metrics (vbench.color, vbench.multiple_objects,
|
||||
# vbench.scene, ...) are excluded — they need prompts built to a
|
||||
# specific schema.
|
||||
METRICS = [
|
||||
"vbench.aesthetic_quality", # CLIP + LAION aesthetic head
|
||||
"vbench.subject_consistency", # DINO frame-to-first cosine
|
||||
"vbench.background_consistency", # DINO on background patches
|
||||
"vbench.imaging_quality", # pyiqa MUSIQ
|
||||
"vbench.temporal_flickering", # pixel-wise frame deltas
|
||||
"vbench.motion_smoothness", # AMT frame interpolator residual
|
||||
"vbench.dynamic_degree", # RAFT optical-flow magnitude (needs fps)
|
||||
"vbench.overall_consistency", # ViCLIP video↔prompt similarity
|
||||
]
|
||||
|
||||
|
||||
def main() -> None:
|
||||
# ----- generation (matches examples/inference/basic/basic_ltx2.py) -----
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Davids048/LTX2-Base-Diffusers",
|
||||
num_gpus=1,
|
||||
)
|
||||
|
||||
output_path = "outputs_video/ltx2_basic/output_ltx2_base_t2v_1088_1920_1.1.mp4"
|
||||
generator.generate_video(
|
||||
prompt=PROMPT,
|
||||
output_path=output_path,
|
||||
save_video=True,
|
||||
num_frames=121,
|
||||
height=1088,
|
||||
width=1920,
|
||||
)
|
||||
generator.shutdown()
|
||||
# Free residual CUDA memory the generator left behind so the
|
||||
# evaluator can grab the largest possible workspace for AMT/RAFT.
|
||||
torch.cuda.empty_cache()
|
||||
|
||||
# ----- scoring -----
|
||||
print(f"\n[eval] building evaluator: {METRICS}")
|
||||
evaluator = Evaluator(metrics=METRICS)
|
||||
|
||||
# LTX2 outputs at 24 fps by default.
|
||||
sample = build_eval_kwargs({"prompt": PROMPT}, output_path, fps=24.0)
|
||||
print(f"[eval] running ({sample['video'].shape[1]} frames @ 24 fps)...")
|
||||
results = evaluator.evaluate(**sample)
|
||||
|
||||
print("\n=== VBench scores ===")
|
||||
for name in METRICS:
|
||||
r = results[name]
|
||||
if r.score is None:
|
||||
reason = r.details.get("skipped", "no score")
|
||||
print(f" {name}: SKIPPED ({reason})")
|
||||
else:
|
||||
print(f" {name}: {r.score:.4f}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,157 @@
|
||||
"""End-to-end Physics-IQ: dataset → generate → score → aggregate.
|
||||
|
||||
Generates one video per take-1 scenario with LTX2 (using the scenario
|
||||
caption as the prompt), scores each generated video against the take-1
|
||||
reference and the take-2 "physical-variance" reference, and prints
|
||||
aggregate scores using :meth:`PhysicsIQMetric.aggregate_components` —
|
||||
the official scoring recipe from the upstream benchmark.
|
||||
|
||||
Reference videos / masks / switch-frames auto-fetch on first miss into
|
||||
``${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq/``; pass ``--dataset-root``
|
||||
to point at a pre-downloaded mirror instead.
|
||||
|
||||
Quick smoke run on 4 scenarios across 2 GPUs::
|
||||
|
||||
python examples/inference/eval/bench_physics_iq.py \\
|
||||
--limit 4 --num-gpus 2 \\
|
||||
--videos-dir outputs_video/physics_iq_smoke
|
||||
|
||||
Re-score existing generations without regenerating::
|
||||
|
||||
python examples/inference/eval/bench_physics_iq.py \\
|
||||
--videos-dir outputs_video/physics_iq_smoke \\
|
||||
--skip-generation
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo.eval import create_evaluator, get_metric
|
||||
from fastvideo.eval.datasets import get_dataset
|
||||
|
||||
|
||||
def _expected_filename(row: dict) -> str:
|
||||
"""Filename Physics-IQ expects for the generated video for *row*.
|
||||
|
||||
Uses the dataset's own ``expected_gen_filename`` annotation so the
|
||||
output filenames match the benchmark's manifest convention.
|
||||
"""
|
||||
return row["auxiliary_info"]["expected_gen_filename"]
|
||||
|
||||
|
||||
def _generate_videos(rows: list[dict], videos_dir: Path,
|
||||
model: str, num_gpus: int,
|
||||
num_frames: int, height: int, width: int) -> None:
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
videos_dir.mkdir(parents=True, exist_ok=True)
|
||||
todo = [(row, videos_dir / _expected_filename(row)) for row in rows]
|
||||
todo = [(row, out) for (row, out) in todo if not out.is_file()]
|
||||
if not todo:
|
||||
print(f"[gen] all {len(rows)} videos already present; skipping.")
|
||||
return
|
||||
|
||||
print(f"[gen] {len(todo)}/{len(rows)} scenarios to render with {model} "
|
||||
f"({num_frames}x{height}x{width})...")
|
||||
gen = VideoGenerator.from_pretrained(model, num_gpus=num_gpus)
|
||||
try:
|
||||
for row, out_path in todo:
|
||||
gen.generate_video(
|
||||
prompt=row["prompt"], output_path=str(out_path), save_video=True,
|
||||
num_frames=num_frames, height=height, width=width,
|
||||
)
|
||||
finally:
|
||||
gen.shutdown()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--dataset-root", type=Path, default=None,
|
||||
help="Path to a pre-downloaded Physics-IQ release. "
|
||||
"Defaults to ${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq, "
|
||||
"auto-fetching missing assets from the public bucket.")
|
||||
p.add_argument("--videos-dir", type=Path,
|
||||
default=Path("outputs_video/bench_physics_iq"),
|
||||
help="Where to read/write generated videos.")
|
||||
p.add_argument("--limit", type=int, default=None,
|
||||
help="Truncate to first N scenarios for smoke runs.")
|
||||
p.add_argument("--num-gpus", type=int, default=1)
|
||||
p.add_argument("--model", default="Davids048/LTX2-Base-Diffusers",
|
||||
help="HF repo id of the text→video generator to use.")
|
||||
p.add_argument("--num-frames", type=int, default=121)
|
||||
p.add_argument("--height", type=int, default=1088)
|
||||
p.add_argument("--width", type=int, default=1920)
|
||||
p.add_argument("--skip-generation", action="store_true",
|
||||
help="Re-score existing videos under --videos-dir.")
|
||||
p.add_argument("--scores-out", type=Path, default=None,
|
||||
help="Where to write per-scenario scores (JSON). "
|
||||
"Defaults to <videos-dir>/scores.json.")
|
||||
args = p.parse_args()
|
||||
|
||||
# 1. Walk the Physics-IQ corpus. Pass --limit to the dataset
|
||||
# constructor so auto-download only fetches the assets we'll use.
|
||||
ds = get_dataset("physics_iq", dataset_root=args.dataset_root, limit=args.limit)
|
||||
rows = list(ds)
|
||||
print(f"[load] Physics-IQ: {len(rows)} scenarios from {ds.dataset_dir}")
|
||||
|
||||
# 2. Generate (or reuse) one mp4 per scenario.
|
||||
if not args.skip_generation:
|
||||
_generate_videos(
|
||||
rows, args.videos_dir, args.model, args.num_gpus,
|
||||
args.num_frames, args.height, args.width,
|
||||
)
|
||||
|
||||
# 3. Score each scenario. The metric reads file paths directly out
|
||||
# of the row dict (reference, reference_take2, masks), so we
|
||||
# just attach the generated video path and forward.
|
||||
evaluator = create_evaluator(metrics=["physics_iq"], num_gpus=args.num_gpus)
|
||||
|
||||
samples: list[dict] = []
|
||||
matched: list[dict] = []
|
||||
for row in rows:
|
||||
video_path = args.videos_dir / _expected_filename(row)
|
||||
if not video_path.is_file():
|
||||
print(f"[eval] missing {video_path}; skipping.")
|
||||
continue
|
||||
# The physics_iq metric accepts file paths via its polymorphic
|
||||
# input handling — no need to load the tensors here.
|
||||
samples.append({"video": str(video_path), **row})
|
||||
matched.append(row)
|
||||
|
||||
all_results = evaluator.evaluate(samples=samples)
|
||||
evaluator.shutdown()
|
||||
|
||||
# 4. Aggregate per the upstream scoring recipe.
|
||||
metric = get_metric("physics_iq")
|
||||
components = metric.aggregate_components(
|
||||
[r["physics_iq"] for r in all_results]
|
||||
)
|
||||
|
||||
print()
|
||||
print("=== Physics-IQ aggregate ===")
|
||||
for name, value in components.items():
|
||||
print(f" {name:24s} {value:.4f}")
|
||||
|
||||
detailed = [
|
||||
{
|
||||
"scenario": row["auxiliary_info"]["scenario_id"],
|
||||
"view": row["view"],
|
||||
"scenario_name": row["auxiliary_info"]["scenario_name"],
|
||||
"score": results["physics_iq"].score,
|
||||
}
|
||||
for row, results in zip(matched, all_results)
|
||||
]
|
||||
out = args.scores_out or (args.videos_dir / "scores.json")
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text(json.dumps(
|
||||
{"aggregate": components, "per_scenario": detailed},
|
||||
indent=2,
|
||||
))
|
||||
print(f"\n[done] per-scenario scores → {out}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,155 @@
|
||||
"""End-to-end VBench: dataset → generate → score → aggregate.
|
||||
|
||||
Iterates the VBench prompt corpus, generates one video per prompt with
|
||||
LTX2, scores each generated video against the requested ``vbench.*``
|
||||
sub-metrics, and prints per-metric averages over the run.
|
||||
|
||||
Re-running with ``--skip-generation`` reuses any mp4 already on disk
|
||||
under ``--videos-dir``, so you can iterate on metric selection without
|
||||
re-paying the generation cost.
|
||||
|
||||
Example — quick smoke run on 4 prompts from the ``aesthetic_quality``
|
||||
dimension across 2 GPUs::
|
||||
|
||||
python examples/inference/eval/bench_vbench.py \\
|
||||
--dimensions aesthetic_quality \\
|
||||
--limit 4 --num-gpus 2 \\
|
||||
--videos-dir outputs_video/vbench_smoke
|
||||
|
||||
Full benchmark on a single dimension::
|
||||
|
||||
python examples/inference/eval/bench_vbench.py \\
|
||||
--dimensions subject_consistency --num-gpus 8
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
from collections import defaultdict
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo.eval import create_evaluator
|
||||
from fastvideo.eval.datasets import get_dataset
|
||||
|
||||
|
||||
def _slugify(prompt: str, max_len: int = 100) -> str:
|
||||
"""Filesystem-safe filename stem; mirrors VBench's official convention."""
|
||||
s = re.sub(r'[\\/:*?"<>|]', "", prompt[:max_len]).strip().strip(".")
|
||||
return re.sub(r"\s+", " ", s) or "output"
|
||||
|
||||
|
||||
def _generate_videos(prompts: list[str], videos_dir: Path,
|
||||
model: str, num_gpus: int,
|
||||
num_frames: int, height: int, width: int) -> None:
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
videos_dir.mkdir(parents=True, exist_ok=True)
|
||||
todo = [(p, videos_dir / f"{_slugify(p)}.mp4") for p in prompts]
|
||||
todo = [(p, out) for (p, out) in todo if not out.is_file()]
|
||||
if not todo:
|
||||
print(f"[gen] all {len(prompts)} videos already present; skipping.")
|
||||
return
|
||||
|
||||
print(f"[gen] {len(todo)}/{len(prompts)} prompts to render with {model} "
|
||||
f"({num_frames}x{height}x{width})...")
|
||||
gen = VideoGenerator.from_pretrained(model, num_gpus=num_gpus)
|
||||
try:
|
||||
for prompt, out_path in todo:
|
||||
gen.generate_video(
|
||||
prompt=prompt, output_path=str(out_path), save_video=True,
|
||||
num_frames=num_frames, height=height, width=width,
|
||||
)
|
||||
finally:
|
||||
gen.shutdown()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--dimensions", default="aesthetic_quality,subject_consistency",
|
||||
help="Comma-separated VBench dimensions (or 'all').")
|
||||
p.add_argument("--limit", type=int, default=None,
|
||||
help="Truncate to first N prompts for smoke runs.")
|
||||
p.add_argument("--videos-dir", type=Path,
|
||||
default=Path("outputs_video/bench_vbench"))
|
||||
p.add_argument("--num-gpus", type=int, default=1)
|
||||
p.add_argument("--model", default="Davids048/LTX2-Base-Diffusers",
|
||||
help="HF repo id of the text→video generator to use.")
|
||||
p.add_argument("--num-frames", type=int, default=121)
|
||||
p.add_argument("--height", type=int, default=1088)
|
||||
p.add_argument("--width", type=int, default=1920)
|
||||
p.add_argument("--fps", type=float, default=24.0,
|
||||
help="Frame-rate annotation passed to fps-aware metrics.")
|
||||
p.add_argument("--skip-generation", action="store_true",
|
||||
help="Re-score existing videos under --videos-dir without "
|
||||
"regenerating.")
|
||||
p.add_argument("--scores-out", type=Path, default=None,
|
||||
help="Where to dump per-prompt scores as JSON. "
|
||||
"Defaults to <videos-dir>/scores.json.")
|
||||
args = p.parse_args()
|
||||
|
||||
# 1. Pull prompts from VBench.
|
||||
dims_arg: list[str] | str = (
|
||||
args.dimensions if args.dimensions == "all"
|
||||
else [d.strip() for d in args.dimensions.split(",") if d.strip()]
|
||||
)
|
||||
ds = get_dataset("vbench", dimensions=dims_arg)
|
||||
rows = list(ds)[: args.limit]
|
||||
print(f"[load] VBench: {len(rows)} prompts across {ds.dimensions}")
|
||||
|
||||
# 2. Generate (or reuse) one mp4 per prompt.
|
||||
if not args.skip_generation:
|
||||
_generate_videos(
|
||||
[row["prompt"] for row in rows],
|
||||
args.videos_dir, args.model, args.num_gpus,
|
||||
args.num_frames, args.height, args.width,
|
||||
)
|
||||
|
||||
# 3. Score each video against the requested vbench sub-metrics.
|
||||
metric_names = sorted(set(f"vbench.{d}" for d in ds.dimensions))
|
||||
print(f"[eval] metrics: {metric_names}")
|
||||
evaluator = create_evaluator(metrics=metric_names, num_gpus=args.num_gpus)
|
||||
|
||||
samples: list[dict] = []
|
||||
matched_rows: list[dict] = []
|
||||
for row in rows:
|
||||
video_path = args.videos_dir / f"{_slugify(row['prompt'])}.mp4"
|
||||
if not video_path.is_file():
|
||||
print(f"[eval] missing {video_path}; skipping this row.")
|
||||
continue
|
||||
# Pass the path; the worker decodes lazily so memory stays bounded.
|
||||
samples.append({
|
||||
"video": str(video_path),
|
||||
"fps": args.fps,
|
||||
**row, # prompt / aux / dims
|
||||
})
|
||||
matched_rows.append(row)
|
||||
|
||||
all_results = evaluator.evaluate(samples=samples)
|
||||
evaluator.shutdown()
|
||||
|
||||
# 4. Aggregate per-metric.
|
||||
by_metric: dict[str, list[float]] = defaultdict(list)
|
||||
detailed: list[dict] = []
|
||||
for row, results in zip(matched_rows, all_results):
|
||||
scores = {name: r.score for name, r in results.items()}
|
||||
detailed.append({"prompt": row["prompt"], "scores": scores})
|
||||
for name, score in scores.items():
|
||||
if score is not None:
|
||||
by_metric[name].append(score)
|
||||
|
||||
print()
|
||||
print("=== per-metric averages ===")
|
||||
for name in sorted(by_metric):
|
||||
avg = sum(by_metric[name]) / len(by_metric[name])
|
||||
print(f" {name:42s} {avg:.4f} (n={len(by_metric[name])})")
|
||||
|
||||
out = args.scores_out or (args.videos_dir / "scores.json")
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text(json.dumps(detailed, indent=2))
|
||||
print(f"\n[done] per-prompt scores → {out}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,167 @@
|
||||
"""End-to-end: generate a video with LTX2 and score it with VBench.
|
||||
|
||||
Pipeline:
|
||||
prompt → LTX2-Base → mp4 → fastvideo.eval → vbench scores
|
||||
|
||||
Run::
|
||||
|
||||
pip install -e .[eval]
|
||||
git submodule update --init fastvideo/third_party/eval/vbench
|
||||
|
||||
python examples/inference/eval/eval_ltx2_vbench.py
|
||||
# or with 4 GPUs and the distilled checkpoint:
|
||||
python examples/inference/eval/eval_ltx2_vbench.py \
|
||||
--model FastVideo/LTX2-Distilled-Diffusers --num-gpus 4
|
||||
|
||||
The default metric set covers the vbench sub-metrics that are
|
||||
meaningful for an arbitrary text→video sample — i.e. those that need
|
||||
only the generated video (and optionally fps + the source prompt).
|
||||
Structured-prompt metrics like ``vbench.color``, ``vbench.scene``,
|
||||
``vbench.multiple_objects`` etc. are *not* on by default — they only
|
||||
make sense when the prompt is built to a specific schema, and they
|
||||
require GRiT/detectron2 setup. Pass them via ``--metrics`` if you have
|
||||
a matching prompt.
|
||||
|
||||
First-time runs download CLIP, DINO, RAFT, AMT, ViCLIP, and MUSIQ
|
||||
weights to ``~/.cache/fastvideo/eval/models/`` and
|
||||
``~/.cache/torch/hub/`` (~few GB total). Subsequent runs are fast.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.eval import create_evaluator
|
||||
from fastvideo.eval.io import load_video
|
||||
|
||||
|
||||
PROMPT = (
|
||||
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
|
||||
"of a woman and a man in their 30s, facing each other with serious "
|
||||
"expressions. The woman, emotional and dramatic, says softly, \"That's "
|
||||
"it... Dad's lost it. And we've lost Dad.\" The man exhales, slightly "
|
||||
"annoyed: \"Stop being so dramatic, Jess.\" A beat. He glances aside, "
|
||||
"then mutters defensively, \"He's just having fun.\" The camera slowly "
|
||||
"pans right, revealing the grandfather in the garden wearing enormous "
|
||||
"butterfly wings, waving his arms in the air like he's trying to take "
|
||||
"off. He shouts, \"Wheeeew!\" as he flaps his wings with full commitment. "
|
||||
"The woman covers her face, on the verge of tears. The tone is deadpan, "
|
||||
"absurd, and quietly tragic."
|
||||
)
|
||||
|
||||
DEFAULT_METRICS = [
|
||||
# No-input metrics: just need the generated frames.
|
||||
"vbench.aesthetic_quality", # CLIP + LAION aesthetic head
|
||||
"vbench.subject_consistency", # DINO frame-to-first cosine
|
||||
"vbench.background_consistency", # DINO on background patches
|
||||
"vbench.imaging_quality", # pyiqa MUSIQ
|
||||
"vbench.temporal_flickering", # pixel-wise frame deltas
|
||||
"vbench.motion_smoothness", # AMT frame interpolator residual
|
||||
# Need fps annotation:
|
||||
"vbench.dynamic_degree", # RAFT optical-flow magnitude
|
||||
# Need the source prompt:
|
||||
"vbench.overall_consistency", # ViCLIP video↔prompt similarity
|
||||
]
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--model", default="Davids048/LTX2-Base-Diffusers",
|
||||
help="HF repo id of the LTX2 checkpoint.")
|
||||
p.add_argument("--num-gpus", type=int, default=1)
|
||||
p.add_argument("--output", default="outputs_video/ltx2_eval/clip.mp4",
|
||||
help="Where to save the generated mp4.")
|
||||
p.add_argument("--num-frames", type=int, default=121)
|
||||
p.add_argument("--height", type=int, default=1088)
|
||||
p.add_argument("--width", type=int, default=1920)
|
||||
p.add_argument("--prompt", default=PROMPT)
|
||||
p.add_argument("--fps", type=float, default=24.0,
|
||||
help="Frame-rate annotation passed to fps-aware metrics "
|
||||
"(e.g. vbench.dynamic_degree). LTX2 outputs at 24 fps "
|
||||
"by default.")
|
||||
p.add_argument("--metrics", default=",".join(DEFAULT_METRICS),
|
||||
help="Comma-separated metric names. Pass 'all' for every "
|
||||
"registered metric, or e.g. 'vbench' for the whole group.")
|
||||
p.add_argument("--scores-out", default="outputs_video/ltx2_eval/scores.json")
|
||||
p.add_argument("--skip-generation", action="store_true",
|
||||
help="Reuse an existing --output video instead of regenerating.")
|
||||
return p.parse_args()
|
||||
|
||||
|
||||
def generate(args: argparse.Namespace) -> Path:
|
||||
out = Path(args.output)
|
||||
if args.skip_generation and out.is_file():
|
||||
print(f"[gen] reusing existing video at {out}")
|
||||
return out
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
print(f"[gen] loading {args.model} ({args.num_gpus} GPU)...")
|
||||
generator = VideoGenerator.from_pretrained(args.model, num_gpus=args.num_gpus)
|
||||
try:
|
||||
print(f"[gen] generating to {out}...")
|
||||
generator.generate_video(
|
||||
prompt=args.prompt,
|
||||
output_path=str(out),
|
||||
save_video=True,
|
||||
num_frames=args.num_frames,
|
||||
height=args.height,
|
||||
width=args.width,
|
||||
)
|
||||
finally:
|
||||
generator.shutdown()
|
||||
return out
|
||||
|
||||
|
||||
def evaluate_video(video_path: Path, prompt: str, fps: float,
|
||||
metric_names) -> dict:
|
||||
print(f"[eval] loading video from {video_path}...")
|
||||
video = load_video(str(video_path)) # (T, C, H, W) in [0, 1]
|
||||
video = video.unsqueeze(0) # → (1, T, C, H, W)
|
||||
|
||||
print(f"[eval] building evaluator: {metric_names}")
|
||||
evaluator = create_evaluator(metrics=metric_names, device="cuda")
|
||||
|
||||
print(f"[eval] running ({video.shape[1]} frames @ {fps} fps)...")
|
||||
results = evaluator.evaluate(
|
||||
video=video,
|
||||
text_prompt=[prompt],
|
||||
fps=fps,
|
||||
)
|
||||
|
||||
if isinstance(results, list):
|
||||
results = results[0] # batch of 1
|
||||
|
||||
return {
|
||||
name: {"score": r.score, "details": r.details}
|
||||
for name, r in results.items()
|
||||
}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
if args.metrics.strip() == "all":
|
||||
metric_names = "all"
|
||||
else:
|
||||
metric_names = [m.strip() for m in args.metrics.split(",") if m.strip()]
|
||||
|
||||
video_path = generate(args)
|
||||
scores = evaluate_video(video_path, args.prompt, args.fps, metric_names)
|
||||
|
||||
print("\n=== VBench scores ===")
|
||||
for name, payload in scores.items():
|
||||
print(f" {name}: {payload['score']}")
|
||||
|
||||
out = Path(args.scores_out)
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text(json.dumps(
|
||||
{"video": str(video_path), "prompt": args.prompt, "scores": scores},
|
||||
indent=2,
|
||||
))
|
||||
print(f"[done] scores written to {out}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,95 @@
|
||||
"""Score a folder of videos in parallel across multiple GPUs.
|
||||
|
||||
Uses :meth:`Evaluator.evaluate(samples=[...])`, which round-robins each
|
||||
sample dict across the GPU replicas the evaluator was built with.
|
||||
|
||||
Example::
|
||||
|
||||
python examples/inference/eval/score_folder.py \\
|
||||
--videos generated/ \\
|
||||
--metrics vbench.aesthetic_quality,vbench.subject_consistency \\
|
||||
--num-gpus 4 \\
|
||||
--output scores.json
|
||||
|
||||
Pair each generated video with a same-name reference video (e.g.
|
||||
``ref/<stem>.mp4``) by passing ``--reference-dir``::
|
||||
|
||||
python examples/inference/eval/score_folder.py \\
|
||||
--videos generated/ --reference-dir ref/ \\
|
||||
--metrics common.psnr,common.ssim,common.lpips \\
|
||||
--num-gpus 4
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo.eval import create_evaluator
|
||||
|
||||
|
||||
def _list_videos(directory: Path) -> list[Path]:
|
||||
exts = {".mp4", ".avi", ".mov", ".mkv", ".gif"}
|
||||
return sorted(p for p in directory.iterdir() if p.suffix.lower() in exts)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--videos", type=Path, required=True,
|
||||
help="Directory of generated videos.")
|
||||
p.add_argument("--reference-dir", type=Path, default=None,
|
||||
help="Directory of reference videos with matching stems.")
|
||||
p.add_argument("--metrics", default="vbench.aesthetic_quality")
|
||||
p.add_argument("--num-gpus", type=int, default=1)
|
||||
p.add_argument("--fps", type=float, default=None,
|
||||
help="Frame-rate annotation for fps-aware metrics.")
|
||||
p.add_argument("--output", type=Path, default=Path("scores.json"))
|
||||
args = p.parse_args()
|
||||
|
||||
video_paths = _list_videos(args.videos)
|
||||
if not video_paths:
|
||||
raise SystemExit(f"No videos under {args.videos}")
|
||||
print(f"Found {len(video_paths)} videos in {args.videos}")
|
||||
|
||||
metrics: list[str] | str = (
|
||||
args.metrics if args.metrics == "all"
|
||||
else [m.strip() for m in args.metrics.split(",") if m.strip()]
|
||||
)
|
||||
evaluator = create_evaluator(metrics=metrics, num_gpus=args.num_gpus)
|
||||
|
||||
# Build per-video sample dicts holding *paths*, not pre-loaded
|
||||
# tensors. Each path is decoded inside the worker thread that picks
|
||||
# up its sample, so peak resident memory is bounded by num_gpus
|
||||
# rather than scaling with the size of the folder.
|
||||
samples: list[dict] = []
|
||||
for vp in video_paths:
|
||||
sample: dict = {"video": str(vp)}
|
||||
if args.reference_dir is not None:
|
||||
ref_path = args.reference_dir / vp.name
|
||||
if not ref_path.is_file():
|
||||
raise FileNotFoundError(f"Missing reference for {vp.name} at {ref_path}")
|
||||
sample["reference"] = str(ref_path)
|
||||
if args.fps is not None:
|
||||
sample["fps"] = args.fps
|
||||
samples.append(sample)
|
||||
|
||||
print(f"Scoring with {len(evaluator.metric_names)} metric(s) "
|
||||
f"on {evaluator.num_gpus} GPU(s)...")
|
||||
all_results = evaluator.evaluate(samples=samples)
|
||||
evaluator.shutdown()
|
||||
|
||||
payload = [
|
||||
{
|
||||
"video": str(vp),
|
||||
"scores": {name: r.score for name, r in results.items()},
|
||||
}
|
||||
for vp, results in zip(video_paths, all_results)
|
||||
]
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(json.dumps(payload, indent=2))
|
||||
print(f"Wrote {args.output}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,68 @@
|
||||
"""Score one video on one GPU.
|
||||
|
||||
Smallest possible use of ``fastvideo.eval``: load an mp4, build an
|
||||
:class:`Evaluator` for the requested metric set, run it.
|
||||
|
||||
Examples::
|
||||
|
||||
# Reference-free (just the generated video):
|
||||
python examples/inference/eval/score_video.py \\
|
||||
--video clip.mp4 \\
|
||||
--metrics vbench.aesthetic_quality,vbench.imaging_quality
|
||||
|
||||
# Reference-paired (compare against ground truth):
|
||||
python examples/inference/eval/score_video.py \\
|
||||
--video gen.mp4 --reference ref.mp4 \\
|
||||
--metrics common.psnr,common.ssim,common.lpips
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
|
||||
from fastvideo.eval import create_evaluator
|
||||
from fastvideo.eval.io import load_video
|
||||
|
||||
|
||||
def main() -> None:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--video", required=True, help="Path to the generated mp4.")
|
||||
p.add_argument("--reference", default=None,
|
||||
help="Optional path to a reference mp4 (for paired metrics).")
|
||||
p.add_argument("--metrics", default="common.psnr,common.ssim",
|
||||
help="Comma-separated metric names, or a group name like 'vbench'.")
|
||||
p.add_argument("--device", default="cuda:0")
|
||||
p.add_argument("--text-prompt", default=None,
|
||||
help="Text prompt for prompt-aware metrics "
|
||||
"(vbench.overall_consistency, etc.).")
|
||||
p.add_argument("--fps", type=float, default=None,
|
||||
help="Frame-rate annotation for fps-aware metrics "
|
||||
"(vbench.dynamic_degree, etc.).")
|
||||
args = p.parse_args()
|
||||
|
||||
metrics: list[str] | str = (
|
||||
args.metrics if args.metrics in ("all",)
|
||||
else [m.strip() for m in args.metrics.split(",") if m.strip()]
|
||||
)
|
||||
evaluator = create_evaluator(metrics=metrics, device=args.device)
|
||||
|
||||
sample: dict = {"video": load_video(args.video)}
|
||||
if args.reference is not None:
|
||||
sample["reference"] = load_video(args.reference)
|
||||
if args.text_prompt is not None:
|
||||
sample["text_prompt"] = args.text_prompt
|
||||
if args.fps is not None:
|
||||
sample["fps"] = args.fps
|
||||
|
||||
results = evaluator.evaluate(**sample)
|
||||
evaluator.shutdown()
|
||||
|
||||
print(json.dumps(
|
||||
{name: r.score for name, r in results.items()},
|
||||
indent=2,
|
||||
))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,216 @@
|
||||
"""``fastvideo eval`` CLI: list registered eval metrics and run them
|
||||
against a set of videos.
|
||||
|
||||
This is a thin wrapper around :mod:`fastvideo.eval`. Heavy lifting
|
||||
(metric loading, GPU handling, batching) lives in
|
||||
:func:`fastvideo.eval.create_evaluator`.
|
||||
|
||||
Examples::
|
||||
|
||||
fastvideo eval list
|
||||
fastvideo eval list --group vbench
|
||||
fastvideo eval run --videos path/to/videos/*.mp4 \\
|
||||
--metrics common.ssim --reference path/to/refs/
|
||||
fastvideo eval run --videos clip.mp4 --metrics vbench.aesthetic_quality \\
|
||||
--output scores.json
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import glob
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import cast
|
||||
|
||||
from fastvideo.entrypoints.cli.cli_types import CLISubcommand
|
||||
from fastvideo.logger import init_logger
|
||||
from fastvideo.utils import FlexibleArgumentParser
|
||||
|
||||
logger = init_logger(__name__)
|
||||
|
||||
|
||||
class EvalSubcommand(CLISubcommand):
|
||||
"""The ``eval`` subcommand — entry point for the eval suite."""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.name = "eval"
|
||||
super().__init__()
|
||||
|
||||
def cmd(self, args: argparse.Namespace) -> None:
|
||||
action = getattr(args, "eval_action", None)
|
||||
if action == "list":
|
||||
_cmd_list(args)
|
||||
elif action == "run":
|
||||
_cmd_run(args)
|
||||
else:
|
||||
# Re-print help if no action was given.
|
||||
self._parser.print_help() # type: ignore[attr-defined]
|
||||
|
||||
def validate(self, args: argparse.Namespace) -> None:
|
||||
action = getattr(args, "eval_action", None)
|
||||
if action == "run" and not args.videos:
|
||||
raise SystemExit("`fastvideo eval run` requires --videos")
|
||||
|
||||
def subparser_init(self, subparsers: argparse._SubParsersAction) -> FlexibleArgumentParser:
|
||||
eval_parser = subparsers.add_parser(
|
||||
"eval",
|
||||
help="Run video-gen evaluation metrics",
|
||||
usage="fastvideo eval {list,run} [...]",
|
||||
)
|
||||
sub = eval_parser.add_subparsers(dest="eval_action", required=False)
|
||||
|
||||
# `eval list`
|
||||
list_p = sub.add_parser("list", help="List registered metrics")
|
||||
list_p.add_argument("--group", type=str, default=None, help="Filter to a metric group (e.g. 'vbench').")
|
||||
|
||||
# `eval run`
|
||||
run_p = sub.add_parser("run", help="Evaluate videos against one or more metrics")
|
||||
run_p.add_argument("--videos",
|
||||
type=str,
|
||||
nargs="+",
|
||||
required=False,
|
||||
help="Path, glob, or directory of generated videos.")
|
||||
run_p.add_argument("--reference",
|
||||
type=str,
|
||||
default=None,
|
||||
help="Path / glob / dir of reference videos (for paired metrics).")
|
||||
run_p.add_argument("--metrics", type=str, default="all", help="Comma-separated metric names, or 'all'.")
|
||||
run_p.add_argument("--device", type=str, default="cuda", help="Torch device (e.g. 'cuda', 'cuda:0', 'cpu').")
|
||||
run_p.add_argument("--text-prompt",
|
||||
type=str,
|
||||
nargs="*",
|
||||
default=None,
|
||||
help="Prompt(s) for text-conditioned metrics. One per video.")
|
||||
run_p.add_argument("--fps", type=float, default=None, help="Frame-rate annotation passed to fps-aware metrics.")
|
||||
run_p.add_argument("--output",
|
||||
type=str,
|
||||
default=None,
|
||||
help="Write results as JSON to this path (default: stdout).")
|
||||
|
||||
# Stash the parser so cmd() can re-print help on no-action.
|
||||
self._parser = eval_parser # type: ignore[attr-defined]
|
||||
return cast(FlexibleArgumentParser, eval_parser)
|
||||
|
||||
|
||||
def _cmd_list(args: argparse.Namespace) -> None:
|
||||
from fastvideo.eval import list_metrics
|
||||
names = list_metrics()
|
||||
if args.group:
|
||||
prefix = args.group.rstrip(".") + "."
|
||||
names = [n for n in names if n == args.group or n.startswith(prefix)]
|
||||
if not names:
|
||||
print(f"(no metrics matched group {args.group!r})")
|
||||
return
|
||||
for name in names:
|
||||
print(name)
|
||||
|
||||
|
||||
def _cmd_run(args: argparse.Namespace) -> None:
|
||||
from fastvideo.eval import create_evaluator
|
||||
from fastvideo.eval.io import load_video
|
||||
|
||||
video_paths = _expand_paths(args.videos)
|
||||
if not video_paths:
|
||||
raise SystemExit(f"No videos matched: {args.videos}")
|
||||
ref_paths = _expand_paths([args.reference]) if args.reference else None
|
||||
|
||||
metrics_arg: list[str] | str = ("all" if args.metrics == "all" else
|
||||
[m.strip() for m in args.metrics.split(",") if m.strip()])
|
||||
|
||||
evaluator = create_evaluator(metrics=metrics_arg, device=args.device)
|
||||
|
||||
all_results: list[dict] = []
|
||||
for i, vp in enumerate(video_paths):
|
||||
logger.info("Evaluating %s (%d/%d)", vp, i + 1, len(video_paths))
|
||||
kwargs: dict = {"video": load_video(vp)}
|
||||
if ref_paths is not None:
|
||||
ref = ref_paths[i] if i < len(ref_paths) else ref_paths[0]
|
||||
kwargs["reference"] = load_video(ref)
|
||||
if args.text_prompt is not None:
|
||||
prompt = (args.text_prompt[i] if i < len(args.text_prompt) else args.text_prompt[0])
|
||||
kwargs["text_prompt"] = [prompt]
|
||||
if args.fps is not None:
|
||||
kwargs["fps"] = args.fps
|
||||
|
||||
results = evaluator.evaluate(**kwargs)
|
||||
all_results.append({
|
||||
"video": str(vp),
|
||||
"scores": _serialize_results(results),
|
||||
})
|
||||
|
||||
payload = json.dumps(all_results, indent=2, default=_jsonable)
|
||||
if args.output:
|
||||
Path(args.output).write_text(payload)
|
||||
logger.info("Wrote results to %s", args.output)
|
||||
else:
|
||||
print(payload)
|
||||
|
||||
|
||||
def _expand_paths(patterns: list[str]) -> list[str]:
|
||||
out: list[str] = []
|
||||
for pat in patterns:
|
||||
p = Path(pat)
|
||||
if p.is_dir():
|
||||
for ext in (".mp4", ".avi", ".mov", ".mkv", ".gif"):
|
||||
out.extend(sorted(str(f) for f in p.iterdir() if f.suffix.lower() == ext))
|
||||
elif any(c in pat for c in "*?["):
|
||||
out.extend(sorted(glob.glob(pat)))
|
||||
else:
|
||||
out.append(pat)
|
||||
# de-dup, preserve order
|
||||
seen: set[str] = set()
|
||||
deduped: list[str] = []
|
||||
for x in out:
|
||||
if x not in seen:
|
||||
seen.add(x)
|
||||
deduped.append(x)
|
||||
return deduped
|
||||
|
||||
|
||||
def _serialize_results(results) -> dict | list:
|
||||
"""Turn evaluator output (dict or list-of-dicts) into JSON-friendly form."""
|
||||
if isinstance(results, list):
|
||||
return [_serialize_results(r) for r in results]
|
||||
if isinstance(results, dict):
|
||||
return {k: _serialize_metric_result(v) for k, v in results.items()}
|
||||
return _serialize_metric_result(results)
|
||||
|
||||
|
||||
def _serialize_metric_result(mr) -> dict:
|
||||
return {
|
||||
"name": getattr(mr, "name", None),
|
||||
"score": getattr(mr, "score", None),
|
||||
"details": getattr(mr, "details", None),
|
||||
}
|
||||
|
||||
|
||||
def _jsonable(obj):
|
||||
"""``json.dumps(default=...)`` coercer for metric outputs.
|
||||
|
||||
Metrics frequently land numpy scalars / arrays, torch tensors, and
|
||||
pathlib paths inside ``MetricResult.details`` (e.g. ``optical_flow``
|
||||
populates ``per_frame_metrics`` with numpy floats). The stdlib JSON
|
||||
encoder rejects all of those by default — this callback walks the
|
||||
leaves and coerces them to native Python types.
|
||||
"""
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
if isinstance(obj, np.integer):
|
||||
return int(obj)
|
||||
if isinstance(obj, np.floating):
|
||||
return float(obj)
|
||||
if isinstance(obj, np.bool_):
|
||||
return bool(obj)
|
||||
if isinstance(obj, np.ndarray):
|
||||
return obj.tolist()
|
||||
if isinstance(obj, torch.Tensor):
|
||||
return obj.detach().cpu().tolist()
|
||||
if isinstance(obj, Path):
|
||||
return str(obj)
|
||||
raise TypeError(f"Object of type {type(obj).__name__} is not JSON serializable")
|
||||
|
||||
|
||||
def cmd_init() -> list[CLISubcommand]:
|
||||
return [EvalSubcommand()]
|
||||
@@ -5,6 +5,7 @@ from fastvideo.entrypoints.cli.generate import cmd_init as generate_cmd_init
|
||||
from fastvideo.utils import FlexibleArgumentParser
|
||||
from fastvideo.entrypoints.cli.serve import cmd_init as serve_cmd_init
|
||||
from fastvideo.entrypoints.cli.bench import cmd_init as bench_cmd_init
|
||||
from fastvideo.entrypoints.cli.eval import cmd_init as eval_cmd_init
|
||||
|
||||
|
||||
def cmd_init() -> list[CLISubcommand]:
|
||||
@@ -13,6 +14,7 @@ def cmd_init() -> list[CLISubcommand]:
|
||||
commands.extend(generate_cmd_init())
|
||||
commands.extend(serve_cmd_init())
|
||||
commands.extend(bench_cmd_init())
|
||||
commands.extend(eval_cmd_init())
|
||||
return commands
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,214 @@
|
||||
# `fastvideo.eval`
|
||||
|
||||
In-process evaluation suite for video generations. Includes pixel
|
||||
metrics (SSIM, PSNR, LPIPS), optical-flow comparisons, the full VBench
|
||||
suite, Physics-IQ, and a VLM scorer (VideoScore-2) behind a single
|
||||
registry-driven API.
|
||||
|
||||
## Install
|
||||
|
||||
| Use case | Install |
|
||||
|---|---|
|
||||
| Default (common, optical_flow, vbench-light, physics_iq, videoscore2) | `uv pip install -e .[eval]` |
|
||||
| Just VBench (12 of 16 sub-metrics) | `uv pip install -e .[eval-vbench]` |
|
||||
| Just Physics-IQ (covered by `[eval]`) | `uv pip install -e .[eval-physics-iq]` |
|
||||
| Plus `vbench.scene` (AVoCaDO) | `uv pip install -e .[eval-full]` |
|
||||
| Plus `vbench.{color, multiple_objects, object_class, spatial_relationship}` (GRiT) | `uv pip install -e .[eval-vbench]` then `uv pip install --no-build-isolation 'git+https://github.com/facebookresearch/detectron2.git'` |
|
||||
|
||||
To use VBench, also pull the upstream submodule:
|
||||
|
||||
```bash
|
||||
git submodule update --init --recursive # fetches vbench + kernel deps
|
||||
```
|
||||
|
||||
The submodule is a clean upstream pin. Compat with current
|
||||
transformers/numpy/timm versions is applied at import time in
|
||||
`fastvideo/eval/metrics/vbench/__init__.py` via attribute-level
|
||||
monkey-patches; the submodule files are unchanged.
|
||||
|
||||
## Public API
|
||||
|
||||
```python
|
||||
from fastvideo.eval import (
|
||||
create_evaluator, # build a reusable Evaluator
|
||||
evaluate, # one-shot helper
|
||||
Evaluator, # the class itself
|
||||
BaseMetric, MetricResult,
|
||||
register, list_metrics, get_metric,
|
||||
ensure_checkpoint, get_cache_dir,
|
||||
)
|
||||
|
||||
ev = create_evaluator(metrics=["common.ssim", "vbench.aesthetic_quality"],
|
||||
device="cuda")
|
||||
scores = ev.evaluate(video=tensor, reference=ref, fps=8.0)
|
||||
```
|
||||
|
||||
`evaluate` accepts either a pre-loaded `(T, C, H, W)` tensor or a path
|
||||
string for `video` and `reference`. Paths are decoded inside the worker
|
||||
that picks up the sample, so peak memory stays bounded by `num_gpus`
|
||||
when scoring large batches.
|
||||
|
||||
### CLI
|
||||
|
||||
```bash
|
||||
fastvideo eval list # list registered metrics
|
||||
fastvideo eval list --group vbench # filter by group
|
||||
fastvideo eval run --videos clip.mp4 \
|
||||
--metrics vbench.aesthetic_quality \
|
||||
--output scores.json
|
||||
fastvideo eval run --videos generated/*.mp4 \
|
||||
--reference reference/ \
|
||||
--metrics common.ssim,common.lpips
|
||||
```
|
||||
|
||||
### Generate-then-score example
|
||||
|
||||
`examples/inference/eval/eval_ltx2_vbench.py` runs an LTX2 prompt
|
||||
through `VideoGenerator` and scores the resulting mp4 with
|
||||
`vbench.aesthetic_quality` and `vbench.subject_consistency`. Use it as
|
||||
a template for end-to-end "generate then score" pipelines.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
fastvideo/
|
||||
├── eval/
|
||||
│ ├── api.py, evaluator.py, registry.py, models.py, ...
|
||||
│ ├── io/ # video loading helpers
|
||||
│ ├── datasets/ # prompt corpora (vbench, physics_iq)
|
||||
│ └── metrics/
|
||||
│ ├── base.py # BaseMetric + @register contract
|
||||
│ ├── common/ # SSIM, PSNR, LPIPS
|
||||
│ ├── optical_flow/ # gt_optical_flow, synthetic_optical_flow
|
||||
│ ├── videoscore2/ # VideoScore-2 (Qwen2.5-VL)
|
||||
│ ├── physics_iq/ # PhysicsIQ + sub-metrics
|
||||
│ └── vbench/ # adapter: sys.path bootstrap + shims
|
||||
│ ├── __init__.py
|
||||
│ └── <16 sub-metric pkgs>
|
||||
└── third_party/
|
||||
└── eval/
|
||||
└── vbench/ # git submodule (Vchitect/VBench)
|
||||
```
|
||||
|
||||
### Prompt datasets
|
||||
|
||||
```python
|
||||
from fastvideo.eval.datasets import get_dataset, list_datasets
|
||||
|
||||
list_datasets() # ['physics_iq', 'vbench']
|
||||
|
||||
ds = get_dataset("physics_iq", limit=4) # auto-fetches assets on first miss
|
||||
for row in ds:
|
||||
# row contains 'prompt', 'reference', 'reference_take2', and
|
||||
# metric-specific aux fields. Drop straight into Evaluator.evaluate(**row).
|
||||
...
|
||||
```
|
||||
|
||||
The Physics-IQ manifest CSV is vendored at
|
||||
`fastvideo/eval/metrics/physics_iq/_vendored/descriptions.csv`.
|
||||
Per-scenario videos, masks, and switch-frames auto-fetch on first use
|
||||
into `${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq/`. For air-gapped
|
||||
runs, pass `auto_download=False` or `dataset_root=` a pre-downloaded
|
||||
copy. Set `FASTVIDEO_PHYSICS_IQ_BUCKET_URL` to redirect the fetch to
|
||||
an internal mirror.
|
||||
|
||||
## Adding a new metric
|
||||
|
||||
The full porting guide is at
|
||||
[`docs/contributing/eval-metrics.md`](../../docs/contributing/eval-metrics.md).
|
||||
Summary below.
|
||||
|
||||
### Native metric (no submodule)
|
||||
|
||||
```python
|
||||
# fastvideo/eval/metrics/common/<your_metric>/metric.py
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("common.your_metric")
|
||||
class YourMetric(BaseMetric):
|
||||
name = "common.your_metric"
|
||||
requires_reference = True
|
||||
needs_gpu = False
|
||||
dependencies: list[str] = [] # e.g. ["pyiqa"] if relevant
|
||||
|
||||
def compute(self, sample) -> list[MetricResult]:
|
||||
...
|
||||
```
|
||||
|
||||
The metric is auto-discovered by `fastvideo/eval/metrics/__init__.py`,
|
||||
which walks all non-underscore subdirectories and imports their
|
||||
`metric` module.
|
||||
|
||||
### Wrapping upstream code via a submodule
|
||||
|
||||
See `fastvideo/eval/metrics/vbench/` for a worked example. The
|
||||
contract is:
|
||||
|
||||
1. Upstream lives as a git submodule under
|
||||
`fastvideo/third_party/eval/<bench>/`, pinned to a SHA in repo-root
|
||||
`.gitmodules`.
|
||||
2. The metric package's `__init__.py`
|
||||
(`fastvideo/eval/metrics/<bench>/__init__.py`) inserts that
|
||||
submodule path on `sys.path` and installs any compat shims for
|
||||
modern torch/transformers/numpy. Do not modify upstream files on
|
||||
disk.
|
||||
3. Per-sub-metric `metric.py` files use `@register("<bench>.<name>")`.
|
||||
|
||||
Patches live as Python in the metric's `__init__.py` so they are
|
||||
grep-able and reviewable.
|
||||
|
||||
## Caches
|
||||
|
||||
Eval cache root: `${FASTVIDEO_CACHE_ROOT}/eval/`, default
|
||||
`~/.cache/fastvideo/eval/`. Override with `FASTVIDEO_EVAL_CACHE`.
|
||||
|
||||
```
|
||||
${FASTVIDEO_CACHE_ROOT}/eval/
|
||||
├── models/ # URL-fetched checkpoints (LAION head, AMT, GRiT)
|
||||
├── torch/ # redirected TORCH_HOME (DINO via torch.hub, lpips)
|
||||
├── clip/ # passed as download_root= to clip.load callsites
|
||||
└── datasets/ # auto-fetched dataset assets, one subdir per benchmark
|
||||
# (e.g. datasets/physics_iq/{split-videos,switch-frames,...})
|
||||
```
|
||||
|
||||
HF-hosted models stay in HF's default cache
|
||||
(`~/.cache/huggingface/hub/`) so they dedupe with other ML projects on
|
||||
the same host.
|
||||
|
||||
### Convention for new metrics
|
||||
|
||||
If your metric wraps a third-party loader that has its own cache
|
||||
directory, route it through `get_cache_dir()` so users get one knob
|
||||
to redirect everything.
|
||||
|
||||
```python
|
||||
# CLIP: pass download_root explicitly
|
||||
import clip
|
||||
from fastvideo.eval.models import get_cache_dir
|
||||
model, _ = clip.load("ViT-B/32", device=device,
|
||||
download_root=str(get_cache_dir() / "clip"))
|
||||
|
||||
# torch.hub is already redirected by fastvideo.eval.__init__ via
|
||||
# TORCH_HOME; no per-callsite work needed.
|
||||
|
||||
# transformers / huggingface_hub: leave alone. HF's default cache is
|
||||
# shared with other tools.
|
||||
```
|
||||
|
||||
For libraries that do not honour any env var or kwarg (pyiqa, funasr),
|
||||
their cache lands in the library's own dir. Document the exception in
|
||||
the metric's docstring if it matters.
|
||||
|
||||
## Out of scope (follow-up PRs)
|
||||
|
||||
- **MIND** metrics. Depend on a separate `vipe` upstream submodule.
|
||||
- **VBench-2.0**. Sibling vbench2 package; needs its own port.
|
||||
- **FVD as a registered metric**. Currently still at `benchmarks/fvd/`.
|
||||
FVD is a set-vs-set distribution distance and does not fit the
|
||||
per-sample `BaseMetric.compute` API without a stateful accumulator;
|
||||
conversion is a designed follow-up.
|
||||
- **Training-time eval callback** (`EvalCallback`) and the
|
||||
`RolloutEvaluator` helper.
|
||||
@@ -0,0 +1,45 @@
|
||||
from fastvideo.eval.models import ensure_checkpoint, get_cache_dir
|
||||
|
||||
|
||||
def _redirect_third_party_caches() -> None:
|
||||
"""Point libraries that respect env vars at the eval cache root.
|
||||
|
||||
Run before metric modules import torch.hub (and friends) so the
|
||||
redirect actually takes effect. We intentionally leave ``HF_HOME``
|
||||
alone — HF's default cache (``~/.cache/huggingface/hub``) is widely
|
||||
shared with other ML projects, and isolating it for eval would force
|
||||
users to re-download already-cached transformers weights.
|
||||
|
||||
Other libraries with non-standard caches (CLIP, pyiqa) don't honour
|
||||
env vars at all; their callsites in metric.py files pass
|
||||
``download_root=str(get_cache_dir() / "<library>")`` directly.
|
||||
"""
|
||||
import os
|
||||
root = get_cache_dir()
|
||||
os.environ.setdefault("TORCH_HOME", str(root / "torch"))
|
||||
|
||||
|
||||
_redirect_third_party_caches()
|
||||
|
||||
from fastvideo.eval.types import MetricResult, Video # noqa: E402
|
||||
from fastvideo.eval.metrics.base import BaseMetric # noqa: E402
|
||||
from fastvideo.eval.registry import register, list_metrics, get_metric # noqa: E402
|
||||
from fastvideo.eval.api import evaluate # noqa: E402
|
||||
from fastvideo.eval.evaluator import Evaluator, create_evaluator # noqa: E402
|
||||
|
||||
# Trigger metric auto-discovery
|
||||
import fastvideo.eval.metrics # noqa: F401, E402
|
||||
|
||||
__all__ = [
|
||||
"evaluate",
|
||||
"Evaluator",
|
||||
"create_evaluator",
|
||||
"MetricResult",
|
||||
"Video",
|
||||
"BaseMetric",
|
||||
"register",
|
||||
"list_metrics",
|
||||
"get_metric",
|
||||
"ensure_checkpoint",
|
||||
"get_cache_dir",
|
||||
]
|
||||
@@ -0,0 +1,36 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.evaluator import create_evaluator
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
def evaluate(
|
||||
generated: torch.Tensor | str | Path,
|
||||
reference: torch.Tensor | str | Path | None = None,
|
||||
metrics: list[str] | str = "all",
|
||||
device: str = "cuda",
|
||||
**kwargs,
|
||||
) -> dict[str, MetricResult] | list[dict[str, MetricResult]]:
|
||||
"""One-shot evaluation. For repeated use, prefer :func:`create_evaluator`.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
generated : Tensor | str | Path
|
||||
Generated video. Either a pre-loaded ``(T, C, H, W)`` tensor or a
|
||||
path to an mp4/avi/etc. — paths are decoded by the worker.
|
||||
reference : Tensor | str | Path | None
|
||||
Reference video (same accepted shapes as *generated*).
|
||||
metrics : list[str] | str
|
||||
Metric names, or ``"all"``.
|
||||
device : str
|
||||
PyTorch device string.
|
||||
"""
|
||||
ev = create_evaluator(metrics=metrics, device=device)
|
||||
kw: dict = {"video": generated, **kwargs}
|
||||
if reference is not None:
|
||||
kw["reference"] = reference
|
||||
return ev.evaluate(**kw)
|
||||
@@ -0,0 +1,51 @@
|
||||
"""Prompt-corpus datasets for end-to-end benchmark evaluation.
|
||||
|
||||
Public API mirrors :mod:`fastvideo.eval` (metrics side):
|
||||
|
||||
from fastvideo.eval.datasets import (
|
||||
PromptDataset, Sample,
|
||||
register_dataset, get_dataset, list_datasets,
|
||||
)
|
||||
|
||||
A dataset is an iterable of plain dicts (one per sample). Built-in
|
||||
datasets self-register at import time. To add one, drop a module into
|
||||
this package that subclasses :class:`PromptDataset` and decorates with
|
||||
``@register_dataset("name")`` — auto-discovery picks it up.
|
||||
"""
|
||||
from fastvideo.eval.datasets.base import (BasePromptDataset, PromptDataset, Sample)
|
||||
from fastvideo.eval.datasets.registry import (get_dataset, list_datasets, register_dataset)
|
||||
|
||||
|
||||
def _autodiscover() -> None:
|
||||
"""Import every non-underscore .py module / subpackage in this package
|
||||
so the ``@register_dataset`` decorators fire."""
|
||||
import importlib
|
||||
import os
|
||||
|
||||
for entry in os.listdir(os.path.dirname(__file__)):
|
||||
if entry.startswith("_") or entry.startswith("."):
|
||||
continue
|
||||
if entry in {"base.py", "registry.py"}:
|
||||
continue
|
||||
if entry.endswith(".py"):
|
||||
importlib.import_module(f"{__name__}.{entry[:-3]}")
|
||||
elif os.path.isdir(os.path.join(os.path.dirname(__file__), entry)) \
|
||||
and os.path.exists(os.path.join(
|
||||
os.path.dirname(__file__), entry, "__init__.py")):
|
||||
importlib.import_module(f"{__name__}.{entry}")
|
||||
|
||||
|
||||
_autodiscover()
|
||||
|
||||
# Re-export the canonical class for typed imports.
|
||||
from fastvideo.eval.datasets.vbench import VBenchPromptDataset # noqa: E402
|
||||
|
||||
__all__ = [
|
||||
"PromptDataset",
|
||||
"BasePromptDataset",
|
||||
"Sample",
|
||||
"register_dataset",
|
||||
"get_dataset",
|
||||
"list_datasets",
|
||||
"VBenchPromptDataset",
|
||||
]
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Prompt-corpus datasets.
|
||||
|
||||
A :class:`PromptDataset` is an iterable of *sample dicts* describing the
|
||||
prompts and conditions for a benchmark. Each sample is a plain dict —
|
||||
no dataclass, no schema enforcement — that flows directly into both
|
||||
generation (``VideoGenerator.generate_video(**sample)``) and scoring
|
||||
(``Evaluator.evaluate(**eval_kwargs)``). The runner picks well-known
|
||||
keys (``prompt``, ``n_samples``, ``dimensions``, ``auxiliary_info``,
|
||||
...) and passes the rest through.
|
||||
|
||||
This matches the surrounding FastVideo style:
|
||||
|
||||
* :class:`fastvideo.dataset.validation_dataset.ValidationDataset` yields dicts.
|
||||
* :meth:`fastvideo.VideoGenerator.generate_video` consumes ``**kwargs``.
|
||||
* :meth:`fastvideo.eval.Evaluator.evaluate` consumes ``**kwargs``.
|
||||
|
||||
To add a new benchmark:
|
||||
|
||||
1. Subclass :class:`PromptDataset`, populate ``self._rows`` with dicts in
|
||||
``__init__``.
|
||||
2. Decorate with ``@register_dataset("my_bench")``.
|
||||
|
||||
Convention for ``auxiliary_info``: a *flat* dict of metric-keyed values
|
||||
(e.g. ``{"color": "red"}``). Benchmarks with nested aux schemas (VBench's
|
||||
``{dim: {key: val}}``) flatten at load time so every consumer sees the
|
||||
same shape.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TypedDict
|
||||
from collections.abc import Iterator
|
||||
|
||||
|
||||
class Sample(TypedDict, total=False):
|
||||
"""Documented schema for a row yielded by :class:`PromptDataset`.
|
||||
|
||||
Only ``prompt`` is required. Extra keys beyond these are forwarded to
|
||||
the runner's eval-kwargs builder verbatim, so action-conditioned or
|
||||
audio-bearing benchmarks can add their own fields without changing
|
||||
the base class.
|
||||
"""
|
||||
prompt: str
|
||||
n_samples: int
|
||||
dimensions: list[str]
|
||||
auxiliary_info: dict
|
||||
image_path: str
|
||||
reference_video: str
|
||||
|
||||
|
||||
class PromptDataset:
|
||||
"""Iterable corpus of sample dicts. Subclasses populate ``self._rows``."""
|
||||
|
||||
name: str = ""
|
||||
description: str = ""
|
||||
supports_dimensions: bool = False
|
||||
requires_reference_image: bool = False
|
||||
requires_reference_video: bool = False
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._rows: list[dict] = []
|
||||
|
||||
def __iter__(self) -> Iterator[dict]:
|
||||
return iter(self._rows)
|
||||
|
||||
def __len__(self) -> int:
|
||||
return len(self._rows)
|
||||
|
||||
def __getitem__(self, i: int) -> dict:
|
||||
return self._rows[i]
|
||||
|
||||
def by_dimension(self) -> dict[str, list[dict]]:
|
||||
"""Group samples by dimension. A multi-dim sample appears under each."""
|
||||
out: dict[str, list[dict]] = {}
|
||||
for s in self._rows:
|
||||
for d in s.get("dimensions", ()):
|
||||
out.setdefault(d, []).append(s)
|
||||
return out
|
||||
|
||||
|
||||
# Back-compat alias for callers still importing the old class name.
|
||||
BasePromptDataset = PromptDataset
|
||||
@@ -0,0 +1,444 @@
|
||||
"""Physics-IQ benchmark prompt corpus.
|
||||
|
||||
Yields one sample dict per take-1 scenario, paired with its take-2
|
||||
reference and both takes' real motion masks. Each row drops straight
|
||||
into :meth:`fastvideo.eval.Evaluator.evaluate` for the ``physics_iq``
|
||||
metric:
|
||||
|
||||
{
|
||||
"prompt": <description>,
|
||||
"reference": "<take-1 mp4>",
|
||||
"reference_take2": "<take-2 mp4>",
|
||||
"reference_mask": "<take-1 mask mp4>",
|
||||
"reference_take2_mask": "<take-2 mask mp4>",
|
||||
"scenario": <scenario_id>,
|
||||
"view": <camera view>,
|
||||
"auxiliary_info": { ... metadata ... },
|
||||
}
|
||||
|
||||
Self-contained dataset: the manifest CSV is vendored under
|
||||
``fastvideo/eval/metrics/physics_iq/_vendored/descriptions.csv``;
|
||||
per-scenario videos/masks/switch-frames auto-fetch on first use from the public
|
||||
DeepMind bucket into ``${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq/``.
|
||||
Pass ``auto_download=False`` (or ``dataset_root=`` pointing at a
|
||||
pre-downloaded copy) to opt out of network fetches.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from urllib.request import urlretrieve
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
|
||||
from fastvideo.eval.datasets.base import PromptDataset
|
||||
from fastvideo.eval.datasets.registry import register_dataset
|
||||
from fastvideo.eval.models import get_cache_dir
|
||||
|
||||
VIEWS = ("perspective-left", "perspective-center", "perspective-right")
|
||||
TAKE1_TOKEN = "take-1"
|
||||
TAKE2_TOKEN = "take-2"
|
||||
|
||||
# FPS the dataset rows should resolve to. Source release is recorded at
|
||||
# 30 FPS; if a different value is requested the loader transcodes once
|
||||
# into a per-repo-root cache directory.
|
||||
_DEFAULT_FPS = 30
|
||||
_DEFAULT_DURATION_SECONDS = 5
|
||||
|
||||
# Vendored manifest under ``fastvideo/eval/metrics/physics_iq/_vendored/``
|
||||
# is the same file shipped by upstream's git repo. The ``_vendored/``
|
||||
# subdir is the project-wide convention for upstream-provenance files
|
||||
# (matches the ``_``-prefixed auto-discovery skip and a single
|
||||
# codespell skip glob).
|
||||
_VENDORED_DESCRIPTIONS_CSV = (Path(__file__).resolve().parent.parent / "metrics" / "physics_iq" / "_vendored" /
|
||||
"descriptions.csv")
|
||||
|
||||
# Public DeepMind bucket; HTTPS-readable, no auth. Override via
|
||||
# ``FASTVIDEO_PHYSICS_IQ_BUCKET_URL`` (e.g. for an internal mirror).
|
||||
_DEFAULT_BUCKET_URL = "https://storage.googleapis.com/physics-iq-benchmark"
|
||||
|
||||
|
||||
def _bucket_url() -> str:
|
||||
return os.environ.get("FASTVIDEO_PHYSICS_IQ_BUCKET_URL", _DEFAULT_BUCKET_URL)
|
||||
|
||||
|
||||
def _default_dataset_root() -> Path:
|
||||
"""Sibling to ``models/torch/clip/`` under the eval cache root."""
|
||||
return get_cache_dir() / "datasets" / "physics_iq"
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PhysicsIQScenario:
|
||||
"""One row of the Physics-IQ manifest, fully resolved on disk."""
|
||||
scenario_id: str
|
||||
view: str
|
||||
scenario_name: str
|
||||
take1_video_path: str
|
||||
take2_video_path: str
|
||||
switch_frame_path: str
|
||||
caption: str
|
||||
expected_gen_filename: str
|
||||
generated_video_path: str | None = None
|
||||
take1_mask_path: str | None = None
|
||||
take2_mask_path: str | None = None
|
||||
|
||||
|
||||
@register_dataset("physics_iq")
|
||||
class PhysicsIQPromptDataset(PromptDataset):
|
||||
"""Physics-IQ benchmark prompt corpus.
|
||||
|
||||
Self-contained: ``get_dataset("physics_iq")`` works with no kwargs.
|
||||
The manifest CSV is vendored next to the metric, and per-scenario
|
||||
assets auto-fetch on first miss from the public bucket into
|
||||
``${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq/``.
|
||||
|
||||
Args:
|
||||
dataset_root: path to a pre-downloaded copy of the Physics-IQ
|
||||
release. Defaults to ``${FASTVIDEO_EVAL_CACHE}/datasets/physics_iq``;
|
||||
override only if you already have a local mirror.
|
||||
fps: target frame rate. The release ships at 30 FPS; other rates
|
||||
transcode once on first access into ``<root>/.physics_iq_cache/``.
|
||||
limit: optional truncation for quick smoke runs. Apply this kwarg
|
||||
(not a post-construction slice) so we only fetch the assets
|
||||
for the scenarios actually requested.
|
||||
generated_dir: optional directory of pre-generated videos —
|
||||
attaches each manifest row's expected output path to the
|
||||
sample dict under ``auxiliary_info["generated_video_path"]``.
|
||||
auto_download: when True (the default), missing testing videos,
|
||||
masks, and switch frames are fetched from the public bucket
|
||||
into ``dataset_root``. Set False for air-gapped runs; the
|
||||
loader will then raise ``FileNotFoundError`` on miss.
|
||||
"""
|
||||
|
||||
description = ("Physics-IQ benchmark, 396 take-1 scenarios across 66 unique physics "
|
||||
"setups × 3 perspective views, each paired with a take-2 reference.")
|
||||
requires_reference_video = True
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
dataset_root: str | Path | None = None,
|
||||
*,
|
||||
fps: int = _DEFAULT_FPS,
|
||||
limit: int | None = None,
|
||||
generated_dir: str | Path | None = None,
|
||||
auto_download: bool = True,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
|
||||
repo_root = Path(dataset_root or _default_dataset_root()).expanduser().resolve()
|
||||
self.repo_root = repo_root
|
||||
self.dataset_dir = _resolve_dataset_dir(repo_root)
|
||||
self.descriptions_path = _resolve_descriptions_path(repo_root, self.dataset_dir)
|
||||
self.cache_dir = repo_root / ".physics_iq_cache"
|
||||
self.fps = fps
|
||||
self.auto_download = auto_download
|
||||
self.bucket_url = _bucket_url()
|
||||
|
||||
scenarios = self._iter_scenarios(
|
||||
fps=fps,
|
||||
generated_dir=generated_dir,
|
||||
limit=limit,
|
||||
)
|
||||
self._rows = [_scenario_to_row(s) for s in scenarios]
|
||||
|
||||
def _iter_scenarios(
|
||||
self,
|
||||
*,
|
||||
fps: int,
|
||||
generated_dir: str | Path | None,
|
||||
limit: int | None,
|
||||
) -> list[PhysicsIQScenario]:
|
||||
with self.descriptions_path.open("r", newline="") as handle:
|
||||
rows = list(csv.DictReader(handle))
|
||||
|
||||
take2_by_suffix = {_scenario_suffix(row["scenario"]): row for row in rows if TAKE2_TOKEN in row["scenario"]}
|
||||
take1_rows = [row for row in rows if TAKE1_TOKEN in row["scenario"]]
|
||||
if limit is not None:
|
||||
take1_rows = take1_rows[:limit]
|
||||
|
||||
generated_dir_path = (Path(generated_dir).expanduser().resolve() if generated_dir else None)
|
||||
scenarios: list[PhysicsIQScenario] = []
|
||||
|
||||
for row in take1_rows:
|
||||
scenario_filename = row["scenario"]
|
||||
scenario_id, view, _, scenario_name = _parse_scenario_filename(scenario_filename)
|
||||
take2_row = take2_by_suffix.get(_scenario_suffix(scenario_filename))
|
||||
if take2_row is None:
|
||||
raise FileNotFoundError(f"Could not find take-2 row matching {scenario_filename}")
|
||||
take2_id, _, _, _ = _parse_scenario_filename(take2_row["scenario"])
|
||||
|
||||
take1_video_path = self._resolve_testing_video_path(
|
||||
scenario_id=scenario_id,
|
||||
view=view,
|
||||
take=TAKE1_TOKEN,
|
||||
scenario_name=scenario_name,
|
||||
fps=fps,
|
||||
)
|
||||
take2_video_path = self._resolve_testing_video_path(
|
||||
scenario_id=take2_id,
|
||||
view=view,
|
||||
take=TAKE2_TOKEN,
|
||||
scenario_name=scenario_name,
|
||||
fps=fps,
|
||||
)
|
||||
switch_frame_path = self._resolve_switch_frame_path(
|
||||
scenario_id=scenario_id,
|
||||
view=view,
|
||||
scenario_name=scenario_name,
|
||||
)
|
||||
take1_mask_path = self._resolve_real_mask_path(
|
||||
scenario_id=scenario_id,
|
||||
view=view,
|
||||
take=TAKE1_TOKEN,
|
||||
scenario_name=scenario_name,
|
||||
fps=fps,
|
||||
)
|
||||
take2_mask_path = self._resolve_real_mask_path(
|
||||
scenario_id=take2_id,
|
||||
view=view,
|
||||
take=TAKE2_TOKEN,
|
||||
scenario_name=scenario_name,
|
||||
fps=fps,
|
||||
)
|
||||
generated_video_path = (str(generated_dir_path /
|
||||
row["generated_video_name"]) if generated_dir_path is not None else None)
|
||||
|
||||
scenarios.append(
|
||||
PhysicsIQScenario(
|
||||
scenario_id=scenario_id,
|
||||
view=view,
|
||||
scenario_name=scenario_name,
|
||||
take1_video_path=str(take1_video_path),
|
||||
take2_video_path=str(take2_video_path),
|
||||
switch_frame_path=str(switch_frame_path),
|
||||
caption=row["description"],
|
||||
expected_gen_filename=row["generated_video_name"],
|
||||
generated_video_path=generated_video_path,
|
||||
take1_mask_path=str(take1_mask_path),
|
||||
take2_mask_path=str(take2_mask_path),
|
||||
))
|
||||
return scenarios
|
||||
|
||||
def _resolve_testing_video_path(
|
||||
self,
|
||||
*,
|
||||
scenario_id: str,
|
||||
view: str,
|
||||
take: str,
|
||||
scenario_name: str,
|
||||
fps: int,
|
||||
) -> Path:
|
||||
target_dir = self.dataset_dir / "split-videos" / "testing" / f"{fps}FPS"
|
||||
target_name = (f"{scenario_id}_testing-videos_{fps}FPS_{view}_{take}_{scenario_name}.mp4")
|
||||
target_path = target_dir / target_name
|
||||
if target_path.exists():
|
||||
return target_path
|
||||
|
||||
# 30-FPS source: either present locally or auto-fetchable.
|
||||
source_name = (f"{scenario_id}_testing-videos_30FPS_{view}_{take}_{scenario_name}.mp4")
|
||||
source_rel = f"split-videos/testing/30FPS/{source_name}"
|
||||
source_path = self.dataset_dir / source_rel
|
||||
self._ensure_remote_asset(source_rel, source_path)
|
||||
if fps == _DEFAULT_FPS:
|
||||
return source_path
|
||||
|
||||
# FPS-convert and cache so repeat runs are free.
|
||||
cache_dir = self.cache_dir / "split-videos" / "testing" / f"{fps}FPS"
|
||||
cache_dir.mkdir(parents=True, exist_ok=True)
|
||||
cached_path = cache_dir / target_name
|
||||
if not cached_path.exists():
|
||||
_convert_video_fps(source_path, cached_path, fps_new=fps)
|
||||
return cached_path
|
||||
|
||||
def _resolve_switch_frame_path(
|
||||
self,
|
||||
*,
|
||||
scenario_id: str,
|
||||
view: str,
|
||||
scenario_name: str,
|
||||
) -> Path:
|
||||
rel = (f"switch-frames/{scenario_id}_switch-frames_anyFPS_{view}_{scenario_name}.jpg")
|
||||
target_path = self.dataset_dir / rel
|
||||
self._ensure_remote_asset(rel, target_path)
|
||||
return target_path
|
||||
|
||||
def _resolve_real_mask_path(
|
||||
self,
|
||||
*,
|
||||
scenario_id: str,
|
||||
view: str,
|
||||
take: str,
|
||||
scenario_name: str,
|
||||
fps: int,
|
||||
) -> Path:
|
||||
# Source release ships masks at 30 FPS only; non-30 rates are
|
||||
# regenerated downstream from the (downsampled) real videos by
|
||||
# the metric — see upstream ``run_physics_iq.py::ensure_binary_mask_structure``.
|
||||
# We only auto-fetch 30 FPS here.
|
||||
rel = (f"video-masks/real/30FPS/"
|
||||
f"{scenario_id}_video-masks_30FPS_{view}_{take}_{scenario_name}.mp4")
|
||||
target_path = self.dataset_dir / rel
|
||||
self._ensure_remote_asset(rel, target_path)
|
||||
if fps == _DEFAULT_FPS:
|
||||
return target_path
|
||||
# Caller asked for a non-30 rate; metric layer handles the
|
||||
# regeneration. Return the canonical 30 FPS path so the metric
|
||||
# always sees a valid mp4 it can transcode.
|
||||
return target_path
|
||||
|
||||
def _ensure_remote_asset(self, rel_path: str, target_path: Path) -> Path:
|
||||
"""Download ``<bucket_url>/<rel_path>`` into *target_path* on miss.
|
||||
|
||||
Atomic via a sibling ``.part`` file; safe under concurrent runs
|
||||
because the final ``rename`` is atomic on POSIX. Raises
|
||||
``FileNotFoundError`` if the file is missing and ``auto_download``
|
||||
is False.
|
||||
"""
|
||||
if target_path.exists():
|
||||
return target_path
|
||||
if not self.auto_download:
|
||||
raise FileNotFoundError(f"Physics-IQ asset missing: {target_path}. "
|
||||
"Set auto_download=True or pass dataset_root= a pre-downloaded copy.")
|
||||
target_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
url = f"{self.bucket_url}/{rel_path.lstrip('/')}"
|
||||
tmp_path = target_path.with_suffix(target_path.suffix + ".part")
|
||||
try:
|
||||
urlretrieve(url, tmp_path)
|
||||
except Exception as exc:
|
||||
if tmp_path.exists():
|
||||
tmp_path.unlink()
|
||||
raise FileNotFoundError(f"Failed to fetch Physics-IQ asset {url} -> {target_path}: "
|
||||
f"{type(exc).__name__}: {exc}") from exc
|
||||
tmp_path.rename(target_path)
|
||||
return target_path
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Internal helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _resolve_dataset_dir(repo_root: Path) -> Path:
|
||||
nested = repo_root / "physics-IQ-benchmark"
|
||||
if nested.exists():
|
||||
return nested
|
||||
return repo_root
|
||||
|
||||
|
||||
def _resolve_descriptions_path(repo_root: Path, dataset_dir: Path) -> Path:
|
||||
"""Prefer a co-located CSV under the user's dataset_root; fall back
|
||||
to the copy vendored in this repo so ``get_dataset("physics_iq")``
|
||||
works without external setup.
|
||||
"""
|
||||
candidates = (
|
||||
repo_root / "descriptions" / "descriptions.csv",
|
||||
dataset_dir / "descriptions" / "descriptions.csv",
|
||||
)
|
||||
for path in candidates:
|
||||
if path.exists():
|
||||
return path
|
||||
if _VENDORED_DESCRIPTIONS_CSV.is_file():
|
||||
return _VENDORED_DESCRIPTIONS_CSV
|
||||
raise FileNotFoundError("Could not locate Physics-IQ descriptions/descriptions.csv "
|
||||
f"(checked {[str(c) for c in candidates]} and vendored "
|
||||
f"{_VENDORED_DESCRIPTIONS_CSV})")
|
||||
|
||||
|
||||
def _parse_scenario_filename(filename: str) -> tuple[str, str, str, str]:
|
||||
stem = Path(filename).name
|
||||
if stem.endswith(".mp4"):
|
||||
stem = stem[:-4]
|
||||
parts = stem.split("_")
|
||||
if len(parts) < 4:
|
||||
raise ValueError(f"Unexpected Physics-IQ filename format: {filename}")
|
||||
return parts[0], parts[1], parts[2], "_".join(parts[3:])
|
||||
|
||||
|
||||
def _scenario_suffix(filename: str) -> str:
|
||||
_, view, _, scenario_name = _parse_scenario_filename(filename)
|
||||
return f"{view}_{scenario_name}"
|
||||
|
||||
|
||||
def _scenario_to_row(scenario: PhysicsIQScenario) -> dict:
|
||||
"""Flatten a :class:`PhysicsIQScenario` into the public sample-dict shape."""
|
||||
aux: dict = {
|
||||
"scenario_id": scenario.scenario_id,
|
||||
"scenario_name": scenario.scenario_name,
|
||||
"switch_frame_path": scenario.switch_frame_path,
|
||||
"expected_gen_filename": scenario.expected_gen_filename,
|
||||
}
|
||||
if scenario.generated_video_path is not None:
|
||||
aux["generated_video_path"] = scenario.generated_video_path
|
||||
|
||||
row: dict = {
|
||||
"prompt": scenario.caption,
|
||||
"reference": scenario.take1_video_path,
|
||||
"reference_take2": scenario.take2_video_path,
|
||||
"scenario": scenario.scenario_id,
|
||||
"view": scenario.view,
|
||||
"auxiliary_info": aux,
|
||||
}
|
||||
if scenario.take1_mask_path is not None:
|
||||
row["reference_mask"] = scenario.take1_mask_path
|
||||
if scenario.take2_mask_path is not None:
|
||||
row["reference_take2_mask"] = scenario.take2_mask_path
|
||||
return row
|
||||
|
||||
|
||||
def _convert_video_fps(input_path: str | Path, output_path: str | Path, *, fps_new: int) -> None:
|
||||
"""Trim *input_path* to ``_DEFAULT_DURATION_SECONDS`` and re-encode at
|
||||
*fps_new*, writing the result to *output_path*. Used to materialize
|
||||
Physics-IQ's 30-FPS source release at user-requested rates.
|
||||
"""
|
||||
input_path = Path(input_path)
|
||||
output_path = Path(output_path)
|
||||
|
||||
cap = cv2.VideoCapture(str(input_path))
|
||||
if not cap.isOpened():
|
||||
raise FileNotFoundError(f"Could not open video for FPS conversion: {input_path}")
|
||||
|
||||
fps_original = cap.get(cv2.CAP_PROP_FPS)
|
||||
frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
|
||||
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
|
||||
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
|
||||
duration = frame_count / fps_original if fps_original else 0.0
|
||||
width, height = width - width % 2, height - height % 2
|
||||
subclip_duration = min(_DEFAULT_DURATION_SECONDS, duration)
|
||||
|
||||
frames: list[np.ndarray] = []
|
||||
for _ in range(int(subclip_duration * fps_original)):
|
||||
ret, frame = cap.read()
|
||||
if not ret:
|
||||
break
|
||||
frames.append(frame)
|
||||
cap.release()
|
||||
|
||||
if not frames:
|
||||
raise ValueError(f"No frames decoded from {input_path}")
|
||||
|
||||
frame_count_new = int(subclip_duration * fps_new)
|
||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
writer = cv2.VideoWriter(
|
||||
str(output_path),
|
||||
cv2.VideoWriter_fourcc(*"avc1"),
|
||||
fps_new,
|
||||
(width, height),
|
||||
)
|
||||
if frame_count_new <= 1:
|
||||
writer.write(frames[0])
|
||||
writer.release()
|
||||
return
|
||||
|
||||
frame_count_original = len(frames)
|
||||
for j in range(frame_count_new):
|
||||
alpha = j * (frame_count_original - 1) / (frame_count_new - 1)
|
||||
idx = int(alpha)
|
||||
alpha -= idx
|
||||
f1 = frames[idx].astype(np.float32)
|
||||
f2 = frames[min(idx + 1, frame_count_original - 1)].astype(np.float32)
|
||||
writer.write(((1.0 - alpha) * f1 + alpha * f2).astype(np.uint8))
|
||||
writer.release()
|
||||
@@ -0,0 +1,41 @@
|
||||
"""Registry for prompt-corpus datasets, mirroring :mod:`fastvideo.eval.registry`."""
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, TYPE_CHECKING
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from fastvideo.eval.datasets.base import BasePromptDataset
|
||||
|
||||
_REGISTRY: dict[str, type[BasePromptDataset]] = {}
|
||||
|
||||
|
||||
def register_dataset(name: str):
|
||||
"""Decorator to register a prompt-dataset class.
|
||||
|
||||
Usage::
|
||||
|
||||
@register_dataset("vbench")
|
||||
class VBenchPromptDataset(BasePromptDataset):
|
||||
...
|
||||
"""
|
||||
|
||||
def wrapper(cls):
|
||||
cls.name = name
|
||||
_REGISTRY[name] = cls
|
||||
return cls
|
||||
|
||||
return wrapper
|
||||
|
||||
|
||||
def get_dataset(name: str, **kwargs: Any) -> BasePromptDataset:
|
||||
"""Instantiate a registered dataset by name."""
|
||||
cls = _REGISTRY.get(name)
|
||||
if cls is None:
|
||||
available = ", ".join(sorted(_REGISTRY.keys()))
|
||||
raise KeyError(f"Unknown dataset '{name}'. Available: {available}")
|
||||
return cls(**kwargs)
|
||||
|
||||
|
||||
def list_datasets() -> list[str]:
|
||||
"""Return sorted list of all registered dataset names."""
|
||||
return sorted(_REGISTRY.keys())
|
||||
@@ -0,0 +1,126 @@
|
||||
"""VBench prompt corpus.
|
||||
|
||||
Single source of truth: upstream's ``VBench_full_info.json`` (946 entries,
|
||||
each with ``prompt_en``, a ``dimension`` list, optional ``auxiliary_info``
|
||||
keyed by dimension).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo.eval.datasets.base import PromptDataset
|
||||
from fastvideo.eval.datasets.registry import register_dataset
|
||||
|
||||
# VBench's official sampling protocol: 5 generations per prompt, except
|
||||
# temporal_flickering which requires 25 (averaging over 5 is too noisy
|
||||
# for a high-frequency-noise metric). See upstream prompts/README.md.
|
||||
TEMPORAL_FLICKERING_SAMPLES = 25
|
||||
DEFAULT_SAMPLES = 5
|
||||
|
||||
_FULL_INFO_REL = "fastvideo/third_party/eval/vbench/vbench/VBench_full_info.json"
|
||||
|
||||
|
||||
def _locate_full_info() -> Path:
|
||||
env = os.environ.get("VBENCH_FULL_INFO_JSON")
|
||||
if env:
|
||||
p = Path(env)
|
||||
if p.is_file():
|
||||
return p
|
||||
raise FileNotFoundError(f"VBENCH_FULL_INFO_JSON={env} does not point at a file")
|
||||
here = Path(__file__).resolve()
|
||||
for ancestor in here.parents:
|
||||
candidate = ancestor / _FULL_INFO_REL
|
||||
if candidate.is_file():
|
||||
return candidate
|
||||
if (ancestor / ".git").exists():
|
||||
break
|
||||
raise FileNotFoundError("Could not locate VBench_full_info.json. Initialize the upstream "
|
||||
"submodule (`git submodule update --init "
|
||||
"fastvideo/third_party/eval/vbench`) or set VBENCH_FULL_INFO_JSON.")
|
||||
|
||||
|
||||
@register_dataset("vbench")
|
||||
class VBenchPromptDataset(PromptDataset):
|
||||
"""VBench prompts filtered by evaluation dimension.
|
||||
|
||||
Args:
|
||||
dimensions: List of dimension names, or ``"all"``. Unknown
|
||||
dimensions raise ``ValueError``.
|
||||
full_info_path: Optional override for ``VBench_full_info.json``;
|
||||
defaults to autodetection.
|
||||
|
||||
A prompt that belongs to several requested dimensions is yielded once;
|
||||
its ``dimensions`` list carries all matches so the scorer can route.
|
||||
"""
|
||||
|
||||
description = ("VBench (Vchitect) prompt corpus, 946 prompts across 16 "
|
||||
"evaluation dimensions.")
|
||||
supports_dimensions = True
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
dimensions: list[str] | str = "all",
|
||||
full_info_path: str | Path | None = None,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
path = Path(full_info_path) if full_info_path else _locate_full_info()
|
||||
with path.open() as f:
|
||||
entries = json.load(f)
|
||||
|
||||
all_dims = sorted({d for e in entries for d in e["dimension"]})
|
||||
if dimensions == "all":
|
||||
self.dimensions: list[str] = all_dims
|
||||
else:
|
||||
unknown = set(dimensions) - set(all_dims)
|
||||
if unknown:
|
||||
raise ValueError(f"Unknown VBench dimensions: {sorted(unknown)}. "
|
||||
f"Available: {all_dims}")
|
||||
self.dimensions = list(dimensions)
|
||||
|
||||
wanted = set(self.dimensions)
|
||||
for entry in entries:
|
||||
relevant = [d for d in entry["dimension"] if d in wanted]
|
||||
if not relevant:
|
||||
continue
|
||||
n = (TEMPORAL_FLICKERING_SAMPLES if "temporal_flickering" in relevant else DEFAULT_SAMPLES)
|
||||
|
||||
# Strip the outer {dim_name: ...} wrapper from upstream's aux
|
||||
# schema so every metric reads its inputs from a flat dict.
|
||||
#
|
||||
# This unwraps exactly one level — the dimension key. Whatever
|
||||
# shape lives inside is the metric's contract:
|
||||
#
|
||||
# color: {"color": {"color": "red"}}
|
||||
# → flat: {"color": "red"} (scalar)
|
||||
#
|
||||
# object_class: {"object_class": {"object": "person"}}
|
||||
# → flat: {"object": "person"} (scalar)
|
||||
#
|
||||
# multiple_objects: {"multiple_objects": {"object": "a and b"}}
|
||||
# → flat: {"object": "a and b"} (scalar)
|
||||
#
|
||||
# spatial_relationship: {"spatial_relationship":
|
||||
# {"spatial_relationship":
|
||||
# {"object_a": ..., "object_b": ...,
|
||||
# "relationship": ...}}}
|
||||
# → flat: {"spatial_relationship": {object_a,object_b,relationship}}
|
||||
#
|
||||
# Note the spatial_relationship case keeps a nested inner dict
|
||||
# by design — upstream double-wraps it, the SpatialRelationship
|
||||
# metric reads ``aux["spatial_relationship"]`` expecting that
|
||||
# inner dict. Don't "simplify" the wrapping away.
|
||||
raw_aux = entry.get("auxiliary_info") or {}
|
||||
flat_aux: dict = {}
|
||||
for v in raw_aux.values():
|
||||
if isinstance(v, dict):
|
||||
flat_aux.update(v)
|
||||
|
||||
self._rows.append({
|
||||
"prompt": entry["prompt_en"],
|
||||
"n_samples": n,
|
||||
"dimensions": relevant,
|
||||
"auxiliary_info": flat_aux,
|
||||
})
|
||||
self.full_info_path = path
|
||||
@@ -0,0 +1,255 @@
|
||||
"""User-facing scorer.
|
||||
|
||||
Layering (mirrors FastVideo's VideoGenerator → Worker pattern, but
|
||||
in-process)::
|
||||
|
||||
Evaluator ← user-facing
|
||||
└── EvalWorker × N ← single-GPU; owns metric replicas
|
||||
└── VideoPool ← async path-→-tensor prefetch (per evaluate call)
|
||||
|
||||
The constructor builds one :class:`EvalWorker` per GPU and loads every
|
||||
metric on every worker eagerly. :meth:`evaluate` is the single entry
|
||||
point: pass kwargs for one sample, or pass a list of sample dicts to
|
||||
fan-out across GPU replicas with pipelined decoding — same method,
|
||||
return type follows the input shape.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import threading
|
||||
from collections.abc import Iterable
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.registry import (list_metrics, missing_dependencies, resolve_group)
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.worker import EvalWorker, add_pool_decode_ms
|
||||
from fastvideo.logger import init_logger
|
||||
|
||||
logger = init_logger(__name__)
|
||||
|
||||
|
||||
class Evaluator:
|
||||
"""Pre-initialized scorer for repeated evaluation.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
metrics : list[str] | str
|
||||
Metric names, group prefixes (``"vbench"``), or ``"all"``.
|
||||
device : str
|
||||
Single-GPU device (e.g. ``"cuda:0"``). Ignored when *num_gpus* > 1.
|
||||
num_gpus : int
|
||||
Number of GPU replicas. Each gets its own :class:`EvalWorker`.
|
||||
compile : bool
|
||||
Apply :func:`torch.compile` to each metric's ``_model``.
|
||||
loader_threads : int
|
||||
Background decode threads in the :class:`VideoPool`. Default 1
|
||||
(hide decode behind compute). Bump for I/O-heavy benchmark sets
|
||||
where one loader can't keep up with the workers.
|
||||
prefetch_factor : int
|
||||
``pool max_size = prefetch_factor * num_workers``. Default 2 —
|
||||
one sample being consumed, one prefetched per worker.
|
||||
pre_upload : bool
|
||||
If ``True`` (default), the worker uploads ``video`` /
|
||||
``reference`` tensors to its device once per sample so every
|
||||
metric in the loop consumes the same GPU-resident tensor (no
|
||||
per-metric ``.to(self.device)`` traffic). Set ``False`` for
|
||||
training-time eval where the shared GPU-resident tensor would
|
||||
compete with the training step for VRAM — each metric then
|
||||
uploads its own copy as before.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
metrics: list[str] | str = "all",
|
||||
device: str = "cuda:0",
|
||||
num_gpus: int = 1,
|
||||
compile: bool = False,
|
||||
*,
|
||||
loader_threads: int = 1,
|
||||
prefetch_factor: int = 2,
|
||||
pre_upload: bool = True,
|
||||
) -> None:
|
||||
names = _resolve_metric_names(metrics)
|
||||
if num_gpus > 1:
|
||||
self._workers = [
|
||||
EvalWorker(names, f"cuda:{i}", compile=compile, pre_upload=pre_upload) for i in range(num_gpus)
|
||||
]
|
||||
else:
|
||||
self._workers = [EvalWorker(names, device, compile=compile, pre_upload=pre_upload)]
|
||||
self._loader_threads = max(1, loader_threads)
|
||||
self._prefetch_factor = max(1, prefetch_factor)
|
||||
|
||||
@property
|
||||
def num_gpus(self) -> int:
|
||||
return len(self._workers)
|
||||
|
||||
@property
|
||||
def metric_names(self) -> list[str]:
|
||||
return self._workers[0].metric_names
|
||||
|
||||
def evaluate(
|
||||
self,
|
||||
samples: Iterable[dict] | None = None,
|
||||
**kwargs,
|
||||
) -> dict[str, MetricResult] | list[dict[str, MetricResult]]:
|
||||
"""Score one sample (kwargs form) or many samples (list form).
|
||||
|
||||
``video`` and ``reference`` may be either a pre-loaded
|
||||
``(T, C, H, W)`` tensor or a path-like (``str`` / ``Path``).
|
||||
Paths in the list form are decoded asynchronously by a
|
||||
:class:`VideoPool` that runs alongside metric compute, hiding
|
||||
decode latency behind GPU work.
|
||||
|
||||
One sample::
|
||||
|
||||
ev.evaluate(video=tensor, text_prompt="...", fps=24.0)
|
||||
ev.evaluate(video="path/to/clip.mp4", fps=24.0)
|
||||
|
||||
Many samples — pipelined decode + work-stealing across replicas::
|
||||
|
||||
ev.evaluate(samples=[
|
||||
{"video": "a.mp4", "reference": "ref_a.mp4"},
|
||||
{"video": "b.mp4", "reference": "ref_b.mp4"},
|
||||
...
|
||||
])
|
||||
|
||||
Multi-GPU dispatch fires automatically when ``num_gpus > 1`` and
|
||||
the list form is used: every worker runs a consumer thread,
|
||||
pulling decoded samples from the shared pool as it frees up.
|
||||
The kwargs form always runs on worker 0 with no pool overhead.
|
||||
"""
|
||||
if samples is None:
|
||||
return self._workers[0].evaluate(**kwargs)
|
||||
|
||||
samples = list(samples)
|
||||
if not samples:
|
||||
return []
|
||||
return self._evaluate_with_pool(samples)
|
||||
|
||||
def _evaluate_with_pool(self, samples: list[dict]) -> list[dict[str, MetricResult]]:
|
||||
"""Pipelined dispatch: ``VideoPool`` prefetches decoded samples;
|
||||
consumers (one per worker) pop them and run metrics.
|
||||
|
||||
Decode order in the pool is non-deterministic — each pool item
|
||||
carries its original input index so results are written back in
|
||||
input order.
|
||||
"""
|
||||
from fastvideo.eval.pool import VideoPool
|
||||
|
||||
n_workers = len(self._workers)
|
||||
max_size = self._prefetch_factor * n_workers
|
||||
results: list[Any] = [None] * len(samples)
|
||||
|
||||
with VideoPool(samples, loader_threads=self._loader_threads, max_size=max_size) as pool:
|
||||
if n_workers == 1:
|
||||
# Single-GPU: this thread is the consumer.
|
||||
while True:
|
||||
item = pool.get()
|
||||
if item is None:
|
||||
break
|
||||
idx, decoded = item
|
||||
results[idx] = self._workers[0].evaluate(**decoded)
|
||||
else:
|
||||
# Multi-GPU: each worker runs its own consumer thread,
|
||||
# pulling from the shared pool (work-stealing).
|
||||
threads: list[threading.Thread] = []
|
||||
for w in self._workers:
|
||||
t = threading.Thread(
|
||||
target=self._consumer_loop,
|
||||
args=(w, pool, results),
|
||||
daemon=True,
|
||||
)
|
||||
t.start()
|
||||
threads.append(t)
|
||||
for t in threads:
|
||||
t.join()
|
||||
|
||||
# Attribute pool-side decode ms to the same global counter
|
||||
# the worker uses so ``pop_timings()`` returns total decode
|
||||
# time regardless of where the decode happened.
|
||||
add_pool_decode_ms(pool.decode_ms_total)
|
||||
|
||||
return results
|
||||
|
||||
@staticmethod
|
||||
def _consumer_loop(worker: EvalWorker, pool: Any, results: list) -> None:
|
||||
while True:
|
||||
item = pool.get()
|
||||
if item is None:
|
||||
return
|
||||
idx, decoded = item
|
||||
results[idx] = worker.evaluate(**decoded)
|
||||
|
||||
def release_cuda_memory(self) -> None:
|
||||
"""Free CUDA caches on every replica without dropping models."""
|
||||
for w in self._workers:
|
||||
w.release_cuda_memory()
|
||||
|
||||
def unload(self) -> None:
|
||||
"""Drop metric refs on every replica. Reverse with :meth:`reload`."""
|
||||
for w in self._workers:
|
||||
w.unload()
|
||||
|
||||
def reload(self) -> None:
|
||||
"""Rebuild metrics dropped by :meth:`unload`."""
|
||||
for w in self._workers:
|
||||
w.reload()
|
||||
|
||||
def shutdown(self) -> None:
|
||||
"""No-op kept for API compatibility.
|
||||
|
||||
Earlier versions of this class held a long-lived
|
||||
``ThreadPoolExecutor`` for multi-GPU round-robin dispatch and
|
||||
needed an explicit shutdown to drain it. The current design
|
||||
builds and tears down a :class:`VideoPool` per ``evaluate``
|
||||
call, so there's no long-lived state to release here.
|
||||
"""
|
||||
|
||||
|
||||
def create_evaluator(
|
||||
metrics: list[str] | str = "all",
|
||||
device: str = "cuda:0",
|
||||
num_gpus: int = 1,
|
||||
compile: bool = False,
|
||||
) -> Evaluator:
|
||||
return Evaluator(metrics=metrics, device=device, num_gpus=num_gpus, compile=compile)
|
||||
|
||||
|
||||
def _resolve_metric_names(metrics: list[str] | str) -> list[str]:
|
||||
"""Resolve metric names, supporting groups (``"vbench"``) and ``"all"``.
|
||||
|
||||
Group / ``"all"`` selectors silently skip metrics whose declared
|
||||
dependencies aren't importable in this environment, with a single
|
||||
warning per skipped metric. Explicit names (e.g. ``"vbench.color"``)
|
||||
always pass through unchanged — the missing dep then surfaces as
|
||||
:class:`ImportError` at construction time, which is what the user
|
||||
asked for.
|
||||
"""
|
||||
if metrics == "all":
|
||||
return _filter_satisfied(list_metrics(), context="all")
|
||||
if isinstance(metrics, str):
|
||||
metrics = [metrics]
|
||||
|
||||
seen: set[str] = set()
|
||||
names: list[str] = []
|
||||
for m in metrics:
|
||||
group = resolve_group(m)
|
||||
candidates = _filter_satisfied(group, context=m) if group is not None else [m]
|
||||
for n in candidates:
|
||||
if n not in seen:
|
||||
seen.add(n)
|
||||
names.append(n)
|
||||
return names
|
||||
|
||||
|
||||
def _filter_satisfied(names: list[str], *, context: str) -> list[str]:
|
||||
"""Drop metrics with missing deps from a group expansion."""
|
||||
keep: list[str] = []
|
||||
for n in names:
|
||||
missing = missing_dependencies(n)
|
||||
if missing:
|
||||
logger.warning(
|
||||
"eval: skipping %s in group '%s'; missing dependency: %s. "
|
||||
"Install instructions: pass the metric name explicitly to see them.", n, context, ", ".join(missing))
|
||||
continue
|
||||
keep.append(n)
|
||||
return keep
|
||||
@@ -0,0 +1,11 @@
|
||||
from fastvideo.eval.io.paths import (build_eval_kwargs, default_filename, glob_videos, sanitize_prompt)
|
||||
from fastvideo.eval.io.video import extract_frames, load_video
|
||||
|
||||
__all__ = [
|
||||
"load_video",
|
||||
"extract_frames",
|
||||
"sanitize_prompt",
|
||||
"default_filename",
|
||||
"glob_videos",
|
||||
"build_eval_kwargs",
|
||||
]
|
||||
@@ -0,0 +1,63 @@
|
||||
"""Filesystem helpers shared by eval scripts.
|
||||
|
||||
Provides the prompt-sanitization, default filename convention, and
|
||||
``(row, video_path) → eval-kwargs`` builder. Free functions, not a
|
||||
class — :class:`fastvideo.eval.Evaluator` is the only stateful object
|
||||
in the eval surface; loops live in user scripts.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
# Filesystem-unsafe characters mirrored from VideoGenerator's output-path
|
||||
# sanitizer so on-disk filenames match what the generator writes.
|
||||
_INVALID_CHARS = re.compile(r'[\\/:*?"<>|]')
|
||||
|
||||
|
||||
def sanitize_prompt(prompt: str, max_len: int = 100) -> str:
|
||||
"""Prompt → safe filename stem."""
|
||||
s = _INVALID_CHARS.sub("", prompt[:max_len]).strip().strip(".")
|
||||
return re.sub(r"\s+", " ", s) or "output"
|
||||
|
||||
|
||||
def default_filename(row: dict, idx: int, ext: str = ".mp4") -> str:
|
||||
"""``<sanitized-prompt>-<idx>.mp4`` — VBench-style."""
|
||||
return f"{sanitize_prompt(row['prompt'])}-{idx}{ext}"
|
||||
|
||||
|
||||
def glob_videos(videos_dir: Path, row: dict, ext: str = ".mp4") -> list[Path]:
|
||||
"""Find every generated video for *row*, sorted by trailing ``-<idx>``."""
|
||||
pattern = f"{sanitize_prompt(row['prompt'])}-*{ext}"
|
||||
files = list(videos_dir.glob(pattern))
|
||||
|
||||
def _idx(p: Path) -> int:
|
||||
try:
|
||||
return int(p.stem.rsplit("-", 1)[1])
|
||||
except (IndexError, ValueError):
|
||||
return -1
|
||||
|
||||
return sorted(files, key=_idx)
|
||||
|
||||
|
||||
def build_eval_kwargs(row: dict, video_path: Path, *, fps: float = 24.0) -> dict[str, Any]:
|
||||
"""Build evaluator kwargs from a sample row + a video on disk.
|
||||
|
||||
Loads the video as ``(T,C,H,W)`` and adds the leading batch dim.
|
||||
Forwards ``prompt`` (as ``text_prompt=[prompt]``) and
|
||||
``auxiliary_info`` (as ``[aux]``) when present on the row.
|
||||
"""
|
||||
from fastvideo.eval.io.video import load_video
|
||||
|
||||
video = load_video(str(video_path)) # (T, C, H, W) in [0, 1]
|
||||
kwargs: dict[str, Any] = {
|
||||
"video": video.unsqueeze(0), # (1, T, C, H, W)
|
||||
"fps": fps,
|
||||
}
|
||||
if "prompt" in row:
|
||||
kwargs["text_prompt"] = [row["prompt"]]
|
||||
aux = row.get("auxiliary_info")
|
||||
if aux:
|
||||
kwargs["auxiliary_info"] = [aux]
|
||||
return kwargs
|
||||
@@ -0,0 +1,85 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
|
||||
def load_video(source: str | torch.Tensor | list, **kwargs) -> torch.Tensor:
|
||||
"""Load a video as a ``(T, C, H, W)`` float32 tensor in ``[0, 1]``.
|
||||
|
||||
Supported *source* types:
|
||||
|
||||
* **str / Path** – path to ``.mp4`` / ``.avi`` / ``.gif`` file, or a
|
||||
directory of frame images (sorted alphabetically).
|
||||
* **torch.Tensor** – returned as-is after shape validation.
|
||||
* **list[PIL.Image]** – stacked into a tensor.
|
||||
"""
|
||||
if isinstance(source, torch.Tensor):
|
||||
if source.ndim != 4:
|
||||
raise ValueError(f"Expected video tensor with 4 dims (T,C,H,W), got {source.ndim}")
|
||||
return source.float()
|
||||
|
||||
if isinstance(source, list):
|
||||
frames = [_pil_to_tensor(img) for img in source]
|
||||
return torch.stack(frames)
|
||||
|
||||
path = Path(source)
|
||||
if path.is_dir():
|
||||
return _load_frame_dir(path)
|
||||
return _load_video_file(str(path))
|
||||
|
||||
|
||||
def extract_frames(video: torch.Tensor, n_frames: int | None = None) -> torch.Tensor:
|
||||
"""Uniformly sample *n_frames* from a ``(T, C, H, W)`` video tensor."""
|
||||
if n_frames is None or n_frames >= video.shape[0]:
|
||||
return video
|
||||
indices = torch.linspace(0, video.shape[0] - 1, n_frames).long()
|
||||
return video[indices]
|
||||
|
||||
|
||||
# --- Internal helpers ---
|
||||
|
||||
|
||||
def _pil_to_tensor(img: Image.Image) -> torch.Tensor:
|
||||
arr = np.array(img.convert("RGB")) # (H, W, 3) uint8
|
||||
return torch.from_numpy(arr).permute(2, 0, 1).float() / 255.0
|
||||
|
||||
|
||||
def _load_frame_dir(path: Path) -> torch.Tensor:
|
||||
exts = {".png", ".jpg", ".jpeg", ".bmp", ".webp"}
|
||||
files = sorted(f for f in path.iterdir() if f.suffix.lower() in exts)
|
||||
if not files:
|
||||
raise FileNotFoundError(f"No image files found in {path}")
|
||||
frames = [_pil_to_tensor(Image.open(f)) for f in files]
|
||||
return torch.stack(frames)
|
||||
|
||||
|
||||
def _load_video_file(path: str) -> torch.Tensor:
|
||||
# Try decord first (faster), fall back to torchvision
|
||||
try:
|
||||
return _load_with_decord(path)
|
||||
except ImportError:
|
||||
pass
|
||||
return _load_with_torchvision(path)
|
||||
|
||||
|
||||
def _load_with_decord(path: str) -> torch.Tensor:
|
||||
from decord import VideoReader, cpu
|
||||
|
||||
vr = VideoReader(path, ctx=cpu(0))
|
||||
# (T, H, W, C) uint8
|
||||
frames = vr.get_batch(list(range(len(vr)))).asnumpy()
|
||||
# → (T, C, H, W) float32 [0, 1]
|
||||
tensor = torch.from_numpy(frames).permute(0, 3, 1, 2).float() / 255.0
|
||||
return tensor
|
||||
|
||||
|
||||
def _load_with_torchvision(path: str) -> torch.Tensor:
|
||||
import torchvision.io
|
||||
|
||||
video, _, _ = torchvision.io.read_video(path, pts_unit="sec")
|
||||
# torchvision returns (T, H, W, C) uint8
|
||||
return video.permute(0, 3, 1, 2).float() / 255.0
|
||||
@@ -0,0 +1,14 @@
|
||||
"""Small GPU-memory helper used by :class:`EvalWorker`."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import gc
|
||||
|
||||
import torch
|
||||
|
||||
|
||||
def clear_cache() -> None:
|
||||
"""Free GPU cache + run garbage collection."""
|
||||
gc.collect()
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.empty_cache()
|
||||
@@ -0,0 +1,29 @@
|
||||
"""Auto-discover and register all built-in metrics (recursive).
|
||||
|
||||
Walks the metrics package tree and imports any leaf-package's
|
||||
``metric`` module so ``@register`` decorators fire. Path components
|
||||
starting with ``_`` are skipped (used for shared helpers like
|
||||
``optical_flow/_shared.py`` and `vbench/_grit_helper.py`).
|
||||
"""
|
||||
|
||||
import importlib
|
||||
import os
|
||||
import contextlib
|
||||
|
||||
|
||||
def _walk(path: str, prefix: str):
|
||||
"""Recursively yield (module_name, is_pkg) for non-underscore packages."""
|
||||
for entry in os.listdir(path):
|
||||
if entry.startswith("_") or entry.startswith("."):
|
||||
continue
|
||||
full = os.path.join(path, entry)
|
||||
if os.path.isdir(full) and os.path.exists(os.path.join(full, "__init__.py")):
|
||||
sub_prefix = f"{prefix}.{entry}"
|
||||
yield (sub_prefix, True)
|
||||
yield from _walk(full, sub_prefix)
|
||||
|
||||
|
||||
for _pkg_path in __path__:
|
||||
for _modname, _ispkg in _walk(_pkg_path, __name__):
|
||||
with contextlib.suppress(ModuleNotFoundError):
|
||||
importlib.import_module(f"{_modname}.metric")
|
||||
@@ -0,0 +1,69 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
class BaseMetric(ABC):
|
||||
"""Abstract base class for all eval metrics.
|
||||
|
||||
Subclasses must implement :meth:`compute`. Optionally override
|
||||
:meth:`setup` to eagerly load models.
|
||||
|
||||
Metrics that need to chunk along the time dimension (frames or frame
|
||||
pairs) for memory reasons should hardcode their own chunk size in
|
||||
``__init__`` (see ``optical_flow`` for the canonical example). Eval
|
||||
always processes one video per :meth:`Evaluator.evaluate` call;
|
||||
``compute`` therefore receives a single sample, not a batch.
|
||||
"""
|
||||
|
||||
name: str = ""
|
||||
requires_reference: bool = True
|
||||
higher_is_better: bool = True
|
||||
dependencies: list[str] = []
|
||||
needs_gpu: bool = False
|
||||
backbone: str | None = None
|
||||
|
||||
# Default time-dim chunk size for metrics that batch internally over
|
||||
# frames or frame-pairs. Override in subclass __init__ if needed
|
||||
# (see ``optical_flow``, ``motion_smoothness``, ``dynamic_degree``).
|
||||
_chunk_size: int | None = None
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._device: torch.device = torch.device("cpu")
|
||||
|
||||
@property
|
||||
def device(self) -> torch.device:
|
||||
return self._device
|
||||
|
||||
def to(self, device: str | torch.device) -> BaseMetric:
|
||||
"""Move metric (and its internal models) to *device*."""
|
||||
self._device = torch.device(device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None: # noqa: B027 - intentionally optional override
|
||||
"""Eagerly load models. Called once by :class:`EvalWorker`.
|
||||
|
||||
Default is a no-op; metrics with no eager state (pixel math,
|
||||
closed-form ops) inherit this. Override only if your metric
|
||||
needs to load weights.
|
||||
"""
|
||||
|
||||
def _skip(self, sample: dict, reason: str) -> MetricResult:
|
||||
"""Return a skipped result (``score=None`` + reason in details)."""
|
||||
return MetricResult(name=self.name, score=None, details={"skipped": reason})
|
||||
|
||||
@abstractmethod
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
"""Compute the metric on a single sample.
|
||||
|
||||
``sample["video"]`` is ``(T, C, H, W)`` float in ``[0, 1]``.
|
||||
``sample["reference"]`` (if used) has the same shape.
|
||||
|
||||
If required inputs are missing, return ``self._skip(sample, reason)``
|
||||
instead of raising.
|
||||
"""
|
||||
...
|
||||
@@ -0,0 +1,70 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("common.lpips")
|
||||
class LPIPSMetric(BaseMetric):
|
||||
name = "common.lpips"
|
||||
requires_reference = True
|
||||
higher_is_better = False
|
||||
needs_gpu = True
|
||||
dependencies = ["lpips"]
|
||||
|
||||
def __init__(self, net: str = "alex", chunk_size: int = 8) -> None:
|
||||
super().__init__()
|
||||
self.net = net
|
||||
# Per-frame AlexNet feature maps at 1080p run ~500 MB each. A
|
||||
# full 121-frame chunk peaks around 60 GB; chunking to 8 frames
|
||||
# drops that to ~5 GB with identical numerical output.
|
||||
self._chunk_size = chunk_size
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device: str | torch.device) -> LPIPSMetric:
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
import lpips as lpips_lib
|
||||
self._model = lpips_lib.LPIPS(net=self.net).to(self.device)
|
||||
self._model.eval()
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
if self._model is None:
|
||||
self.setup()
|
||||
|
||||
# When the worker pre-uploaded inputs (default), these are
|
||||
# already on ``self.device`` and the ``.to(...)`` below is a
|
||||
# no-op. With ``pre_upload=False`` the worker keeps them on CPU
|
||||
# and this metric pays the transfer just like before.
|
||||
gen = sample["video"].float().to(self.device, non_blocking=True)
|
||||
ref = sample["reference"].float().to(self.device, non_blocking=True)
|
||||
|
||||
n = min(gen.shape[0], ref.shape[0])
|
||||
gen, ref = gen[:n] * 2.0 - 1.0, ref[:n] * 2.0 - 1.0
|
||||
|
||||
chunk = self._chunk_size or n
|
||||
all_scores = []
|
||||
with torch.no_grad():
|
||||
for i in range(0, n, chunk):
|
||||
s = self._model(gen[i:i + chunk], ref[i:i + chunk]).squeeze()
|
||||
if s.dim() == 0:
|
||||
s = s.unsqueeze(0)
|
||||
all_scores.append(s)
|
||||
scores = torch.cat(all_scores) # (n,)
|
||||
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(scores.mean()),
|
||||
details={"per_frame": scores.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,49 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("common.psnr")
|
||||
class PSNRMetric(BaseMetric):
|
||||
name = "common.psnr"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
# PSNR is `((gen - ref)**2).mean(...)` plus log — memory-bandwidth-
|
||||
# bound on host (~6 GB read + 3 GB write per video pair at 1080p ×
|
||||
# 121 fr). Trivial on GPU and frees the host bus for the loader.
|
||||
needs_gpu = True
|
||||
|
||||
def __init__(self, max_val: float = 1.0, chunk_size: int = 32) -> None:
|
||||
super().__init__()
|
||||
self.max_val = max_val
|
||||
# (gen - ref)**2 at 1080p × 121 fr allocates a full ~3 GB
|
||||
# intermediate. chunk=32 caps that at ~800 MB with identical
|
||||
# numerical output.
|
||||
self._chunk_size = chunk_size
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
gen = sample["video"].float().to(self.device) # (T, C, H, W)
|
||||
ref = sample["reference"].float().to(self.device)
|
||||
n = min(gen.shape[0], ref.shape[0])
|
||||
gen, ref = gen[:n], ref[:n]
|
||||
|
||||
# Per-frame MSE → PSNR, chunked so the squared-diff intermediate
|
||||
# never holds the whole clip at once.
|
||||
chunk = self._chunk_size or n
|
||||
mse_parts = []
|
||||
for i in range(0, n, chunk):
|
||||
g = gen[i:i + chunk]
|
||||
r = ref[i:i + chunk]
|
||||
mse_parts.append(((g - r)**2).mean(dim=(1, 2, 3)))
|
||||
mse = torch.cat(mse_parts) # (T,)
|
||||
psnr = 10.0 * torch.log10(self.max_val**2 / mse.clamp(min=1e-10))
|
||||
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=psnr.mean().item(),
|
||||
details={"per_frame": psnr.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,85 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
def _ssim_per_frame(
|
||||
x: torch.Tensor,
|
||||
y: torch.Tensor,
|
||||
window_size: int = 11,
|
||||
C1: float = 0.01**2,
|
||||
C2: float = 0.03**2,
|
||||
) -> torch.Tensor:
|
||||
"""Compute SSIM for each frame. Returns ``(N,)`` tensor where N = number of frames."""
|
||||
channels = x.shape[1]
|
||||
kernel = _gaussian_kernel(window_size, 1.5, channels, x.device, x.dtype)
|
||||
|
||||
mu_x = F.conv2d(x, kernel, groups=channels, padding=window_size // 2)
|
||||
mu_y = F.conv2d(y, kernel, groups=channels, padding=window_size // 2)
|
||||
|
||||
mu_x2 = mu_x * mu_x
|
||||
mu_y2 = mu_y * mu_y
|
||||
mu_xy = mu_x * mu_y
|
||||
|
||||
sigma_x2 = F.conv2d(x * x, kernel, groups=channels, padding=window_size // 2) - mu_x2
|
||||
sigma_y2 = F.conv2d(y * y, kernel, groups=channels, padding=window_size // 2) - mu_y2
|
||||
sigma_xy = F.conv2d(x * y, kernel, groups=channels, padding=window_size // 2) - mu_xy
|
||||
|
||||
num = (2 * mu_xy + C1) * (2 * sigma_xy + C2)
|
||||
den = (mu_x2 + mu_y2 + C1) * (sigma_x2 + sigma_y2 + C2)
|
||||
ssim_map = num / den
|
||||
|
||||
return ssim_map.mean(dim=(1, 2, 3))
|
||||
|
||||
|
||||
def _gaussian_kernel(size: int, sigma: float, channels: int, device, dtype):
|
||||
coords = torch.arange(size, device=device, dtype=dtype) - size // 2
|
||||
g = torch.exp(-coords**2 / (2 * sigma**2))
|
||||
g = g / g.sum()
|
||||
kernel_2d = g.unsqueeze(1) * g.unsqueeze(0)
|
||||
kernel = kernel_2d.unsqueeze(0).unsqueeze(0).repeat(channels, 1, 1, 1)
|
||||
return kernel
|
||||
|
||||
|
||||
@register("common.ssim")
|
||||
class SSIMMetric(BaseMetric):
|
||||
name = "common.ssim"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
# SSIM is 5 depthwise conv2d's per chunk — pure GPU territory at any
|
||||
# interesting resolution. At 1080p × 121 frames the CPU path takes
|
||||
# ~5–10 s per pair vs <100 ms on GPU. Keeping it on CPU also made
|
||||
# the metric fight the pipelined loader thread for DDR bandwidth.
|
||||
needs_gpu = True
|
||||
|
||||
def __init__(self, window_size: int = 11, chunk_size: int = 16) -> None:
|
||||
super().__init__()
|
||||
self.window_size = window_size
|
||||
# Each conv2d output is (chunk, 3, H, W); SSIM allocates ~5–7
|
||||
# such intermediates plus inputs. At 1080p with chunk=121 this
|
||||
# peaks ~25 GB; chunk=16 brings peak to ~4 GB with no change
|
||||
# in output.
|
||||
self._chunk_size = chunk_size
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
gen = sample["video"].float().to(self.device) # (T, C, H, W)
|
||||
ref = sample["reference"].float().to(self.device)
|
||||
n = min(gen.shape[0], ref.shape[0])
|
||||
gen, ref = gen[:n], ref[:n]
|
||||
|
||||
chunk = self._chunk_size or n
|
||||
parts = []
|
||||
for i in range(0, n, chunk):
|
||||
parts.append(_ssim_per_frame(gen[i:i + chunk], ref[i:i + chunk], self.window_size))
|
||||
per_frame = torch.cat(parts) # (n,)
|
||||
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=per_frame.mean().item(),
|
||||
details={"per_frame": per_frame.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,297 @@
|
||||
"""Shared helpers for optical-flow metrics.
|
||||
|
||||
Both ``optical_flow.gt_optical_flow`` and
|
||||
``optical_flow.synthetic_optical_flow`` extract per-frame flow with
|
||||
``ptlflow`` and reduce it through the same per-pixel / per-frame /
|
||||
temporal aggregation pipeline. The pipeline lives here so the two
|
||||
metrics stay byte-identical on the comparison side and only differ in
|
||||
how they construct the *reference* flow field.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
_PER_FRAME_AGG_KEYS: tuple[str, ...] = (
|
||||
"mf_epe",
|
||||
"mf_angle_err",
|
||||
"mf_cosine",
|
||||
"mf_mag_ratio",
|
||||
"pixel_epe_mean",
|
||||
"pixel_epe_max",
|
||||
"px_angle_rmse",
|
||||
"grid_epe_mean",
|
||||
"grid_epe_max",
|
||||
"fl_all",
|
||||
"foe_dist",
|
||||
"flow_kl_2d",
|
||||
)
|
||||
|
||||
|
||||
def _trapezoid(vals: np.ndarray) -> float:
|
||||
fn = getattr(np, "trapezoid", None) or np.trapz
|
||||
return float(fn(vals))
|
||||
|
||||
|
||||
def _estimate_foe(
|
||||
flow: np.ndarray,
|
||||
step: int = 8,
|
||||
min_mag: float = 0.5,
|
||||
) -> tuple[float, float]:
|
||||
"""Least-squares Focus of Expansion. Returns (fx, fy)."""
|
||||
H, W = flow.shape[:2]
|
||||
ys = np.arange(step // 2, H, step)
|
||||
xs = np.arange(step // 2, W, step)
|
||||
yy, xx = np.meshgrid(ys, xs, indexing="ij")
|
||||
yy = yy.ravel()
|
||||
xx = xx.ravel()
|
||||
uu = flow[yy, xx, 0]
|
||||
vv = flow[yy, xx, 1]
|
||||
|
||||
mag = np.sqrt(uu**2 + vv**2)
|
||||
valid = mag > min_mag
|
||||
if valid.sum() < 10:
|
||||
return W / 2.0, H / 2.0
|
||||
|
||||
xx = xx[valid].astype(np.float64)
|
||||
yy = yy[valid].astype(np.float64)
|
||||
uu = uu[valid].astype(np.float64)
|
||||
vv = vv[valid].astype(np.float64)
|
||||
|
||||
# v * fx - u * fy = v * x - u * y
|
||||
A = np.column_stack([vv, -uu])
|
||||
b = vv * xx - uu * yy
|
||||
result, _, _, _ = np.linalg.lstsq(A, b, rcond=None)
|
||||
return float(result[0]), float(result[1])
|
||||
|
||||
|
||||
def _flow_kl_2d(
|
||||
flow_a: np.ndarray,
|
||||
flow_b: np.ndarray,
|
||||
n_angle_bins: int = 36,
|
||||
n_mag_bins: int = 20,
|
||||
min_mag: float = 0.5,
|
||||
) -> float:
|
||||
"""KL(P_a || P_b) over a joint (angle, log-magnitude) histogram."""
|
||||
|
||||
def _hist(flow: np.ndarray) -> np.ndarray | None:
|
||||
u, v = flow[:, :, 0].ravel(), flow[:, :, 1].ravel()
|
||||
mag = np.sqrt(u**2 + v**2)
|
||||
angle = np.degrees(np.arctan2(v, u)) % 360
|
||||
valid = mag >= min_mag
|
||||
if valid.sum() < 10:
|
||||
return None
|
||||
mag = mag[valid]
|
||||
angle = angle[valid]
|
||||
mag_max = max(mag.max(), min_mag + 1.0)
|
||||
mag_edges = np.logspace(np.log10(min_mag), np.log10(mag_max), n_mag_bins + 1)
|
||||
angle_edges = np.linspace(0, 360, n_angle_bins + 1)
|
||||
h, _, _ = np.histogram2d(angle, mag, bins=[angle_edges, mag_edges])
|
||||
return h
|
||||
|
||||
ha, hb = _hist(flow_a), _hist(flow_b)
|
||||
if ha is None or hb is None:
|
||||
return 0.0
|
||||
eps = 1.0
|
||||
p = (ha + eps) / (ha + eps).sum()
|
||||
q = (hb + eps) / (hb + eps).sum()
|
||||
return float((p * np.log(p / q)).sum())
|
||||
|
||||
|
||||
def compute_frame_metrics(
|
||||
flow_gt: np.ndarray,
|
||||
flow_gen: np.ndarray,
|
||||
grid_size: int = 8,
|
||||
min_mag: float = 0.5,
|
||||
max_mag_pct: float = 80.0,
|
||||
) -> dict[str, float]:
|
||||
"""Per-frame comparison metrics between two HxWx2 flow fields.
|
||||
|
||||
Port of mhuo's compute_frame_metrics — see ptlflow_validation.py.
|
||||
"""
|
||||
metrics: dict[str, float] = {}
|
||||
|
||||
gt_mag_map = np.linalg.norm(flow_gt, axis=2)
|
||||
gen_mag_map = np.linalg.norm(flow_gen, axis=2)
|
||||
max_mag_map = np.maximum(gt_mag_map, gen_mag_map)
|
||||
mag_hi = np.percentile(max_mag_map, max_mag_pct)
|
||||
mag_mask = (max_mag_map >= min_mag) & (max_mag_map <= mag_hi)
|
||||
n_valid = int(mag_mask.sum())
|
||||
|
||||
if n_valid > 0:
|
||||
mean_gt = flow_gt[mag_mask].mean(axis=0)
|
||||
mean_gen = flow_gen[mag_mask].mean(axis=0)
|
||||
else:
|
||||
mean_gt = flow_gt.reshape(-1, 2).mean(axis=0)
|
||||
mean_gen = flow_gen.reshape(-1, 2).mean(axis=0)
|
||||
|
||||
metrics["mf_epe"] = float(np.linalg.norm(mean_gt - mean_gen))
|
||||
|
||||
mf_min_mag = 0.1
|
||||
mag_gt = float(np.linalg.norm(mean_gt))
|
||||
mag_gen = float(np.linalg.norm(mean_gen))
|
||||
if mag_gt < mf_min_mag and mag_gen < mf_min_mag:
|
||||
metrics["mf_angle_err"] = 0.0
|
||||
metrics["mf_cosine"] = 1.0
|
||||
elif mag_gt < mf_min_mag or mag_gen < mf_min_mag:
|
||||
metrics["mf_angle_err"] = 90.0
|
||||
metrics["mf_cosine"] = 0.0
|
||||
elif mag_gt > 1e-6 and mag_gen > 1e-6:
|
||||
cos_sim = float(np.dot(mean_gt, mean_gen) / (mag_gt * mag_gen))
|
||||
cos_sim = float(np.clip(cos_sim, -1.0, 1.0))
|
||||
metrics["mf_angle_err"] = float(np.degrees(np.arccos(cos_sim)))
|
||||
metrics["mf_cosine"] = cos_sim
|
||||
else:
|
||||
metrics["mf_angle_err"] = 0.0
|
||||
metrics["mf_cosine"] = 1.0
|
||||
|
||||
metrics["mf_mag_ratio"] = float(mag_gen / mag_gt) if mag_gt > 1e-6 else 1.0
|
||||
|
||||
epe_map = np.linalg.norm(flow_gt - flow_gen, axis=2)
|
||||
if n_valid > 0:
|
||||
metrics["pixel_epe_mean"] = float(epe_map[mag_mask].mean())
|
||||
metrics["pixel_epe_max"] = float(epe_map[mag_mask].max())
|
||||
else:
|
||||
metrics["pixel_epe_mean"] = float(epe_map.mean())
|
||||
metrics["pixel_epe_max"] = float(epe_map.max())
|
||||
|
||||
valid = mag_mask & (gt_mag_map > 0.5) & (gen_mag_map > 0.5)
|
||||
if valid.sum() > 0:
|
||||
dot = (flow_gt[:, :, 0] * flow_gen[:, :, 0] + flow_gt[:, :, 1] * flow_gen[:, :, 1])
|
||||
cos_map = np.clip(dot / (gt_mag_map * gen_mag_map + 1e-8), -1.0, 1.0)
|
||||
angle_map = np.degrees(np.arccos(cos_map))
|
||||
metrics["px_angle_rmse"] = float(np.sqrt((angle_map[valid]**2).mean()))
|
||||
else:
|
||||
metrics["px_angle_rmse"] = 0.0
|
||||
|
||||
H, W = epe_map.shape
|
||||
gh, gw = H // grid_size, W // grid_size
|
||||
grid_vals = []
|
||||
for gi in range(grid_size):
|
||||
for gj in range(grid_size):
|
||||
cell_mask = mag_mask[gi * gh:(gi + 1) * gh, gj * gw:(gj + 1) * gw]
|
||||
cell_epe = epe_map[gi * gh:(gi + 1) * gh, gj * gw:(gj + 1) * gw]
|
||||
if cell_mask.sum() > 0:
|
||||
grid_vals.append(float(cell_epe[cell_mask].mean()))
|
||||
else:
|
||||
grid_vals.append(float(cell_epe.mean()))
|
||||
metrics["grid_epe_mean"] = float(np.mean(grid_vals))
|
||||
metrics["grid_epe_max"] = float(np.max(grid_vals))
|
||||
|
||||
if n_valid > 0:
|
||||
outlier = (epe_map > 3.0) & (epe_map > 0.05 * gt_mag_map) & mag_mask
|
||||
metrics["fl_all"] = float(outlier.sum() / n_valid)
|
||||
else:
|
||||
outlier = (epe_map > 3.0) & (epe_map > 0.05 * gt_mag_map)
|
||||
metrics["fl_all"] = float(outlier.mean())
|
||||
|
||||
foe_gt_x, foe_gt_y = _estimate_foe(flow_gt)
|
||||
foe_gen_x, foe_gen_y = _estimate_foe(flow_gen)
|
||||
metrics["foe_dist"] = float(np.sqrt((foe_gt_x - foe_gen_x)**2 + (foe_gt_y - foe_gen_y)**2))
|
||||
|
||||
metrics["flow_kl_2d"] = _flow_kl_2d(flow_gt, flow_gen)
|
||||
return metrics
|
||||
|
||||
|
||||
def aggregate_temporal(per_frame: list[dict[str, float]], ) -> dict[str, float | int | None]:
|
||||
"""Aggregate per-frame metric dicts into mean/std/max/auc/onset summaries.
|
||||
|
||||
Port of mhuo's compute_temporal_metrics.
|
||||
"""
|
||||
n = len(per_frame)
|
||||
if n == 0:
|
||||
return {"n_frames": 0}
|
||||
|
||||
summary: dict[str, float | int | None] = {"n_frames": n}
|
||||
series: dict[str, np.ndarray] = {k: np.array([m[k] for m in per_frame]) for k in _PER_FRAME_AGG_KEYS}
|
||||
for name, vals in series.items():
|
||||
summary[f"{name}_mean"] = float(vals.mean())
|
||||
summary[f"{name}_std"] = float(vals.std())
|
||||
summary[f"{name}_max"] = float(vals.max())
|
||||
summary[f"{name}_auc"] = _trapezoid(vals) / max(n - 1, 1)
|
||||
|
||||
epe_series = series["pixel_epe_mean"]
|
||||
window = min(5, n)
|
||||
if n >= window:
|
||||
baseline = float(np.median(epe_series[:window]))
|
||||
threshold = max(baseline * 2.0, 1.0)
|
||||
kernel = np.ones(window) / window
|
||||
smoothed = np.convolve(epe_series, kernel, mode="valid")
|
||||
divergence_frame: int | None = None
|
||||
for i, val in enumerate(smoothed):
|
||||
if val > threshold:
|
||||
divergence_frame = int(i)
|
||||
break
|
||||
summary["divergence_onset_frame"] = divergence_frame
|
||||
summary["divergence_threshold"] = float(threshold)
|
||||
else:
|
||||
summary["divergence_onset_frame"] = None
|
||||
summary["divergence_threshold"] = None
|
||||
return summary
|
||||
|
||||
|
||||
def tensor_to_bgr_list(video: torch.Tensor) -> list[np.ndarray]:
|
||||
"""Convert ``(T, C, H, W)`` float [0,1] to a list of HWC BGR uint8 frames.
|
||||
|
||||
Performs the cast + permute + BGR swap on the input's device, then
|
||||
transfers once. This avoids 121 per-frame ``.cpu().numpy()`` calls
|
||||
moving 3 GB of float32 across PCIe when the input lives on GPU
|
||||
(which is what ``Evaluator(pre_upload=True)`` produces). One uint8
|
||||
transfer is ~4× less bytes than the per-frame float32 round-trips.
|
||||
"""
|
||||
# (T, C, H, W) float [0,1] → (T, H, W, C) uint8 BGR, all on-device.
|
||||
bgr_u8 = (video.float() * 255.0).clamp(0, 255).to(torch.uint8).permute(0, 2, 3, 1).flip(-1).contiguous()
|
||||
arr = bgr_u8.cpu().numpy() # single transfer
|
||||
return [arr[t] for t in range(arr.shape[0])]
|
||||
|
||||
|
||||
def load_ptlflow_model(model_name: str, ckpt: str, device: torch.device):
|
||||
"""Load a ``ptlflow`` model on *device* in eval mode."""
|
||||
import ptlflow
|
||||
model = ptlflow.get_model(model_name, ckpt_path=ckpt)
|
||||
model.eval()
|
||||
return model.to(device)
|
||||
|
||||
|
||||
def extract_video_flows(
|
||||
model,
|
||||
video: torch.Tensor, # (T, C, H, W) float [0, 1]
|
||||
*,
|
||||
chunk: int,
|
||||
device: torch.device,
|
||||
) -> list[np.ndarray]:
|
||||
"""Run *model* on every consecutive frame pair in *video*.
|
||||
|
||||
Returns a list of HxWx2 flow arrays of length ``T - 1``.
|
||||
"""
|
||||
from ptlflow.utils.io_adapter import IOAdapter
|
||||
|
||||
h, w = video.shape[2], video.shape[3]
|
||||
io_adapter = IOAdapter(
|
||||
output_stride=model.output_stride,
|
||||
input_size=(h, w),
|
||||
cuda=(device.type == "cuda"),
|
||||
)
|
||||
bgr_frames = tensor_to_bgr_list(video)
|
||||
pairs = [(bgr_frames[i], bgr_frames[i + 1]) for i in range(len(bgr_frames) - 1)]
|
||||
|
||||
flows: list[np.ndarray] = []
|
||||
for start in range(0, len(pairs), chunk):
|
||||
end = min(start + chunk, len(pairs))
|
||||
pair_tensors = []
|
||||
for f1, f2 in pairs[start:end]:
|
||||
inputs = io_adapter.prepare_inputs([f1, f2])
|
||||
pair_tensors.append(inputs["images"])
|
||||
batched_images = torch.cat(pair_tensors, dim=0)
|
||||
with torch.no_grad():
|
||||
preds = model({"images": batched_images})
|
||||
preds["images"] = batched_images
|
||||
preds = io_adapter.unscale(preds)
|
||||
flows_tensor = preds["flows"]
|
||||
if flows_tensor.dim() == 5:
|
||||
flows_tensor = flows_tensor.squeeze(1)
|
||||
for i in range(flows_tensor.shape[0]):
|
||||
flow = flows_tensor[i].detach().cpu().permute(1, 2, 0).numpy()
|
||||
flows.append(flow)
|
||||
return flows
|
||||
@@ -0,0 +1,118 @@
|
||||
"""Compare optical flow extracted from a generated video against optical
|
||||
flow extracted from a ground-truth reference video.
|
||||
|
||||
Both flows are produced by the same ``ptlflow`` model (default
|
||||
``dpflow``/``things``). The resulting per-pixel / per-frame / temporal
|
||||
metric set is identical to ``synthetic_optical_flow`` — only the way the
|
||||
*reference* flow is constructed differs.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.metrics.optical_flow._shared import (
|
||||
aggregate_temporal,
|
||||
compute_frame_metrics,
|
||||
extract_video_flows,
|
||||
load_ptlflow_model,
|
||||
)
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("optical_flow.gt_optical_flow")
|
||||
class GtOpticalFlowMetric(BaseMetric):
|
||||
"""Per-pixel / per-frame / temporal flow comparison vs. a reference video.
|
||||
|
||||
The headline ``score`` is ``pixel_epe_mean_mean`` (lower is better);
|
||||
every other scalar lives in ``details`` so downstream consumers can
|
||||
pick whichever one they care about.
|
||||
"""
|
||||
|
||||
name = "optical_flow.gt_optical_flow"
|
||||
requires_reference = True
|
||||
higher_is_better = False
|
||||
needs_gpu = True
|
||||
backbone = "optical_flow"
|
||||
dependencies = ["ptlflow"]
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
model_name: str = "dpflow",
|
||||
ckpt: str = "things",
|
||||
min_mag: float = 0.5,
|
||||
max_mag_pct: float = 80.0,
|
||||
grid_size: int = 8,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
self.model_name = model_name
|
||||
self.ckpt = ckpt
|
||||
self.min_mag = min_mag
|
||||
self.max_mag_pct = max_mag_pct
|
||||
self.grid_size = grid_size
|
||||
self._model = None
|
||||
# Frame-pair batch size per DPFlow forward. The correlation
|
||||
# volume scales ~quadratically with input resolution and the
|
||||
# cost-volume tensor is roughly 4 GB per pair at 1080p — so
|
||||
# batching 16 pairs requested ~63 GB and OOMed on H200 (matches
|
||||
# mhuo's reference impl, which runs one pair per forward, see
|
||||
# ``mhuo/ptlflow/eval_flow_divergence.py:99-105``). Default 1
|
||||
# is memory-safe at any resolution; bump for low-res to amortize
|
||||
# Python loop overhead.
|
||||
self._chunk_size = 1
|
||||
|
||||
def to(self, device: str | torch.device) -> GtOpticalFlowMetric:
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
self._model = load_ptlflow_model(self.model_name, self.ckpt, self.device)
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
if self._model is None:
|
||||
self.setup()
|
||||
|
||||
gen_video = sample["video"].float() # (T, C, H, W)
|
||||
ref_video = sample["reference"].float()
|
||||
n = min(gen_video.shape[0], ref_video.shape[0])
|
||||
gen_video, ref_video = gen_video[:n], ref_video[:n]
|
||||
if n < 2:
|
||||
raise ValueError("Need at least 2 frames to compute optical flow")
|
||||
|
||||
chunk = self._chunk_size or 16
|
||||
gen_flows = extract_video_flows(
|
||||
self._model,
|
||||
gen_video,
|
||||
chunk=chunk,
|
||||
device=self.device,
|
||||
)
|
||||
ref_flows = extract_video_flows(
|
||||
self._model,
|
||||
ref_video,
|
||||
chunk=chunk,
|
||||
device=self.device,
|
||||
)
|
||||
per_frame = [
|
||||
compute_frame_metrics(
|
||||
rf,
|
||||
gf,
|
||||
grid_size=self.grid_size,
|
||||
min_mag=self.min_mag,
|
||||
max_mag_pct=self.max_mag_pct,
|
||||
) for rf, gf in zip(ref_flows, gen_flows, strict=False)
|
||||
]
|
||||
summary = aggregate_temporal(per_frame)
|
||||
score = summary.get("pixel_epe_mean_mean")
|
||||
details = dict(summary)
|
||||
details["per_frame_metrics"] = per_frame
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(score) if score is not None else None,
|
||||
details=details,
|
||||
)
|
||||
@@ -0,0 +1,332 @@
|
||||
"""Third-person synthetic optical-flow generator (Option B, no depth).
|
||||
|
||||
Camera frame convention (OpenCV): x=right, y=down, z=forward.
|
||||
|
||||
Per-frame inputs from the action stream:
|
||||
keyboard : (6,) — [W, S, A, D, turn_left, turn_right]
|
||||
mouse : (2,) — [pitch, yaw] (pitch may be sign-flipped per sample,
|
||||
see ``mouse_pitch_sign``)
|
||||
|
||||
Mapping to camera kinematics:
|
||||
omega_x = alpha_pitch * mouse_pitch (camera pitch)
|
||||
omega_y = alpha_yaw * mouse_yaw
|
||||
+ alpha_turn * (turn_right - turn_left) (camera yaw)
|
||||
omega_z = 0 (no roll)
|
||||
|
||||
T_avatar_x = beta_strafe * (D - A)
|
||||
T_avatar_y = 0
|
||||
T_avatar_z = beta_fwd * (W - S)
|
||||
|
||||
Off-pivot correction. The orbit camera rotates about the avatar pivot at
|
||||
``r = (0, r_y, r_z)`` in camera-local coords, not about the optical
|
||||
center. A rotation by omega about that pivot is kinematically equivalent
|
||||
to a rotation about the optical center plus a translation
|
||||
``T_orbit = -(omega x r)``. So:
|
||||
|
||||
T_total = T_avatar + T_orbit
|
||||
|
||||
Flow (no depth, Z = 1):
|
||||
u_R = (xy/f)*ωx - (f + x²/f)*ωy + y*ωz
|
||||
v_R = (f + y²/f)*ωx - (xy/f)*ωy - x*ωz
|
||||
u_T = -f*Tx + x*Tz
|
||||
v_T = -f*Ty + y*Tz
|
||||
|
||||
The Z=1 collapse means strafe (Tx-only) produces a uniform horizontal
|
||||
field. Forward motion (Tz) still has the right radial direction
|
||||
structure since u_T scales with x, just no depth-modulated magnitude.
|
||||
Angle-family metrics survive this; per-pixel magnitude metrics will be
|
||||
biased on parallax-rich backgrounds. That's the explicit cost of
|
||||
declining to use generated-video depth.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import asdict, dataclass, field
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
|
||||
@dataclass
|
||||
class ThirdPersonCalibration:
|
||||
"""Fitted parameters for a third-person rig.
|
||||
|
||||
``r_y`` defaults to 0 (avatar at camera height). ``focal_length`` is
|
||||
in pixels. ``init_pitch`` is the camera's rest pitch (radians, negative
|
||||
means tilted down — typical over-the-shoulder framing). User mouse-pitch
|
||||
input is integrated on top of this baseline.
|
||||
"""
|
||||
alpha_yaw: float
|
||||
alpha_pitch: float
|
||||
alpha_turn: float
|
||||
beta_fwd: float
|
||||
beta_strafe: float
|
||||
focal_length: float
|
||||
r_z: float
|
||||
r_y: float = 0.0
|
||||
init_pitch: float = 0.0
|
||||
notes: str = ""
|
||||
fit_metadata: dict = field(default_factory=dict)
|
||||
|
||||
def to_json(self, path: str | Path) -> None:
|
||||
Path(path).write_text(json.dumps(asdict(self), indent=2))
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, d: dict) -> ThirdPersonCalibration:
|
||||
known = {f.name for f in cls.__dataclass_fields__.values()}
|
||||
return cls(**{k: v for k, v in d.items() if k in known})
|
||||
|
||||
|
||||
def load_calibration(path: str | Path) -> ThirdPersonCalibration:
|
||||
return ThirdPersonCalibration.from_dict(json.loads(Path(path).read_text()))
|
||||
|
||||
|
||||
class ThirdPersonFlowGenerator:
|
||||
"""Vectorized 3P synthetic-flow generator. No depth.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
calibration : ThirdPersonCalibration
|
||||
frame_shape : (H, W)
|
||||
mouse_pitch_sign : +1 or -1 — sample-level flag from metadata
|
||||
(``mouse_pitch_flipped: true`` in mhuo's data ⇒ -1).
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
calibration: ThirdPersonCalibration,
|
||||
frame_shape: tuple[int, int],
|
||||
mouse_pitch_sign: int = +1,
|
||||
) -> None:
|
||||
self.cal = calibration
|
||||
self.H, self.W = frame_shape
|
||||
self.mouse_pitch_sign = int(mouse_pitch_sign)
|
||||
|
||||
f = self.cal.focal_length
|
||||
cx, cy = self.W / 2.0, self.H / 2.0
|
||||
xs = np.arange(self.W, dtype=np.float64) - cx
|
||||
ys = np.arange(self.H, dtype=np.float64) - cy
|
||||
self.x_grid, self.y_grid = np.meshgrid(xs, ys) # H,W
|
||||
|
||||
# Pre-compute LH rotation kernels (depend only on pixel coords + f).
|
||||
self.xy_over_f = self.x_grid * self.y_grid / f
|
||||
self.f_plus_x2_over_f = f + self.x_grid**2 / f
|
||||
self.f_plus_y2_over_f = f + self.y_grid**2 / f
|
||||
|
||||
@staticmethod
|
||||
def _action_to_kinematics(
|
||||
keyboard: np.ndarray,
|
||||
mouse: np.ndarray,
|
||||
cal: ThirdPersonCalibration,
|
||||
mouse_pitch_sign: int,
|
||||
) -> tuple[np.ndarray, np.ndarray]:
|
||||
"""Return (omega (3,), T_total (3,)) for one frame's action."""
|
||||
kb = np.asarray(keyboard, dtype=np.float64).reshape(-1)
|
||||
mo = np.asarray(mouse, dtype=np.float64).reshape(-1)
|
||||
|
||||
pitch = mo[0] * mouse_pitch_sign
|
||||
yaw = mo[1]
|
||||
|
||||
omega_x = cal.alpha_pitch * pitch
|
||||
omega_y = cal.alpha_yaw * yaw
|
||||
if kb.shape[0] >= 6:
|
||||
omega_y += cal.alpha_turn * (kb[5] - kb[4])
|
||||
omega = np.array([omega_x, omega_y, 0.0])
|
||||
|
||||
T_avatar = np.array([
|
||||
cal.beta_strafe * (kb[3] - kb[2]),
|
||||
0.0,
|
||||
cal.beta_fwd * (kb[0] - kb[1]),
|
||||
])
|
||||
|
||||
# T_orbit = -(omega x r) with r = (0, r_y, r_z).
|
||||
# cross([wx,wy,0], [0,ry,rz]) = (wy*rz, -wx*rz, wx*ry)
|
||||
T_orbit = -np.array([
|
||||
omega[1] * cal.r_z,
|
||||
-omega[0] * cal.r_z,
|
||||
omega[0] * cal.r_y,
|
||||
])
|
||||
return omega, T_avatar + T_orbit
|
||||
|
||||
def _flow_from_kinematics(self, omega: np.ndarray, T: np.ndarray) -> np.ndarray:
|
||||
"""Compose rotation + translation flow at every pixel. Z=1."""
|
||||
wx, wy, wz = omega
|
||||
Tx, Ty, Tz = T
|
||||
f = self.cal.focal_length
|
||||
|
||||
u_R = self.xy_over_f * wx - self.f_plus_x2_over_f * wy + self.y_grid * wz
|
||||
v_R = self.f_plus_y2_over_f * wx - self.xy_over_f * wy - self.x_grid * wz
|
||||
u_T = -f * Tx + self.x_grid * Tz
|
||||
v_T = -f * Ty + self.y_grid * Tz
|
||||
return np.stack([u_R + u_T, v_R + v_T], axis=-1).astype(np.float32)
|
||||
|
||||
def generate_flow(
|
||||
self,
|
||||
keyboard: np.ndarray,
|
||||
mouse: np.ndarray,
|
||||
) -> np.ndarray:
|
||||
"""Synthesize HxWx2 flow for one frame's action."""
|
||||
omega, T = self._action_to_kinematics(
|
||||
keyboard,
|
||||
mouse,
|
||||
self.cal,
|
||||
self.mouse_pitch_sign,
|
||||
)
|
||||
return self._flow_from_kinematics(omega, T)
|
||||
|
||||
def generate_flow_sequence(
|
||||
self,
|
||||
actions: dict,
|
||||
n_pairs: int | None = None,
|
||||
) -> list[np.ndarray]:
|
||||
"""Generate flow for each consecutive frame pair.
|
||||
|
||||
Returns ``n`` flows where ``n = n_pairs`` if supplied, else
|
||||
``len(keyboard) - 1``.
|
||||
"""
|
||||
kb = actions["keyboard"]
|
||||
mo = actions["mouse"]
|
||||
T = len(kb)
|
||||
n = (T - 1) if n_pairs is None else min(n_pairs, T - 1)
|
||||
return [self.generate_flow(kb[i], mo[i]) for i in range(n)]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Linear-features form (used by the calibration fitter).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# Number of free params we fit. Order is fixed and shared with the fitter.
|
||||
PARAM_NAMES = (
|
||||
"alpha_yaw",
|
||||
"alpha_pitch",
|
||||
"alpha_turn",
|
||||
"beta_fwd",
|
||||
"beta_strafe",
|
||||
"focal_length",
|
||||
"r_z",
|
||||
"r_y",
|
||||
)
|
||||
|
||||
|
||||
def predict_flow_at_pixels(
|
||||
keyboard: np.ndarray, # (6,)
|
||||
mouse: np.ndarray, # (2,)
|
||||
xs_centered: np.ndarray, # (N,) pixel x relative to principal point
|
||||
ys_centered: np.ndarray, # (N,)
|
||||
cal: ThirdPersonCalibration,
|
||||
mouse_pitch_sign: int,
|
||||
) -> np.ndarray:
|
||||
"""Vectorized flow prediction at an arbitrary pixel set. Returns (N, 2).
|
||||
|
||||
Stateless: assumes camera is level (theta_pitch=0). For 3P games where
|
||||
the camera tilts independently of the avatar's facing, use
|
||||
:func:`predict_flow_at_pixels_stateful` and pass the integrated pitch.
|
||||
"""
|
||||
omega, T = ThirdPersonFlowGenerator._action_to_kinematics(
|
||||
keyboard,
|
||||
mouse,
|
||||
cal,
|
||||
mouse_pitch_sign,
|
||||
)
|
||||
wx, wy, wz = omega
|
||||
Tx, Ty, Tz = T
|
||||
f = cal.focal_length
|
||||
|
||||
xy_over_f = xs_centered * ys_centered / f
|
||||
f_plus_x2_over_f = f + xs_centered**2 / f
|
||||
f_plus_y2_over_f = f + ys_centered**2 / f
|
||||
|
||||
u_R = xy_over_f * wx - f_plus_x2_over_f * wy + ys_centered * wz
|
||||
v_R = f_plus_y2_over_f * wx - xy_over_f * wy - xs_centered * wz
|
||||
u_T = -f * Tx + xs_centered * Tz
|
||||
v_T = -f * Ty + ys_centered * Tz
|
||||
return np.stack([u_R + u_T, v_R + v_T], axis=-1)
|
||||
|
||||
|
||||
def predict_flow_at_pixels_stateful(
|
||||
keyboard: np.ndarray,
|
||||
mouse: np.ndarray,
|
||||
xs_centered: np.ndarray,
|
||||
ys_centered: np.ndarray,
|
||||
cal: ThirdPersonCalibration,
|
||||
theta_pitch: float,
|
||||
) -> np.ndarray:
|
||||
"""Stateful flow prediction that accounts for accumulated camera pitch.
|
||||
|
||||
In a 3P game the avatar moves in the world's horizontal plane in the
|
||||
direction the camera is yawed. When the camera is also pitched
|
||||
(looking down at the avatar / up at the sky), this world-horizontal
|
||||
motion has a non-zero y component in the camera frame. Concretely:
|
||||
|
||||
T_cam = β_fwd · (W − S) · (0, sin θ_pitch, cos θ_pitch)
|
||||
+ β_strafe · (D − A) · (1, 0, 0) # strafe is pitch-invariant
|
||||
|
||||
Yaw doesn't appear because the camera frame is yaw-aligned by construction
|
||||
(avatar and camera yaw together). Mouse rotation contributions to ω are
|
||||
unchanged.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
theta_pitch : float (radians)
|
||||
Accumulated camera pitch state at this frame, integrated from
|
||||
prior mouse-pitch input. Positive = camera looking up.
|
||||
"""
|
||||
f = cal.focal_length
|
||||
cos_p = float(np.cos(theta_pitch))
|
||||
sin_p = float(np.sin(theta_pitch))
|
||||
|
||||
# Rotational velocities (unchanged from stateless)
|
||||
pitch_in = float(mouse[0])
|
||||
yaw_in = float(mouse[1])
|
||||
omega_x = cal.alpha_pitch * pitch_in
|
||||
omega_y = cal.alpha_yaw * yaw_in
|
||||
if keyboard.shape[0] >= 6:
|
||||
omega_y += cal.alpha_turn * (keyboard[5] - keyboard[4])
|
||||
|
||||
# Avatar-frame translations (in world-horizontal plane, avatar-yaw-aligned)
|
||||
avatar_strafe = cal.beta_strafe * (keyboard[3] - keyboard[2])
|
||||
avatar_fwd = cal.beta_fwd * (keyboard[0] - keyboard[1])
|
||||
|
||||
# Map to camera frame using current pitch
|
||||
Tx = avatar_strafe
|
||||
Ty = avatar_fwd * sin_p
|
||||
Tz = avatar_fwd * cos_p
|
||||
|
||||
xy_over_f = xs_centered * ys_centered / f
|
||||
f_plus_x2_over_f = f + xs_centered**2 / f
|
||||
f_plus_y2_over_f = f + ys_centered**2 / f
|
||||
|
||||
u_R = xy_over_f * omega_x - f_plus_x2_over_f * omega_y
|
||||
v_R = f_plus_y2_over_f * omega_x - xy_over_f * omega_y
|
||||
u_T = -f * Tx + xs_centered * Tz
|
||||
v_T = -f * Ty + ys_centered * Tz
|
||||
return np.stack([u_R + u_T, v_R + v_T], axis=-1)
|
||||
|
||||
|
||||
def integrate_pitch_state(
|
||||
mouse: np.ndarray, # (T, 2) raw or cached actions
|
||||
cal: ThirdPersonCalibration,
|
||||
*,
|
||||
init_pitch: float = 0.0,
|
||||
frames_per_step: int = 1,
|
||||
) -> np.ndarray:
|
||||
"""Integrate per-frame mouse-pitch input into accumulated camera pitch.
|
||||
|
||||
Returns a (T,) array where ``out[t]`` is the camera's accumulated pitch
|
||||
angle (radians) AT THE START of frame ``t`` — i.e. the pose under which
|
||||
frame ``t``'s action is interpreted.
|
||||
|
||||
For raw-frame action sequences pass ``frames_per_step=1``. For cached
|
||||
actions where each sample represents N raw frames of integration
|
||||
(cache stride = N), pass ``frames_per_step=N``.
|
||||
|
||||
NOTE: this only models the explicit user-input pitch. Cinematic
|
||||
auto-pitch (camera tilting to track the avatar over uneven terrain)
|
||||
isn't in the action stream and isn't captured here. For that you need
|
||||
visual odometry (Option B / WorldCam-style ViPE pipeline).
|
||||
"""
|
||||
per_step = cal.alpha_pitch * np.asarray(mouse[:, 0], dtype=np.float64) * frames_per_step
|
||||
cum = np.cumsum(per_step)
|
||||
# out[t] = pose BEFORE frame t's input is applied → shift by one
|
||||
return init_pitch + np.concatenate([[0.0], cum[:-1]]).astype(np.float64)
|
||||
@@ -0,0 +1,169 @@
|
||||
"""Compare optical flow extracted from a generated video against optical
|
||||
flow synthesized analytically from per-frame actions.
|
||||
|
||||
The reference flow is *not* observed from a ground-truth video — it's
|
||||
predicted from the action stream via a third-person camera-kinematics
|
||||
model (Longuet-Higgins linearization + off-pivot translation correction;
|
||||
no depth). Observed flow comes from the same ``ptlflow`` model used by
|
||||
``gt_optical_flow``, and the two are compared with the identical metric
|
||||
set, so scores are directly comparable across the two metrics.
|
||||
|
||||
Required sample keys
|
||||
--------------------
|
||||
``video``
|
||||
``(B, T, C, H, W)`` float in ``[0, 1]``.
|
||||
``actions``
|
||||
``dict`` (or list-of-dicts of length B) with two ``np.ndarray`` keys:
|
||||
|
||||
* ``keyboard`` of shape ``(T, 6)`` — ``[W, S, A, D, turn_left, turn_right]``
|
||||
* ``mouse`` of shape ``(T, 2)`` — ``[pitch, yaw]``
|
||||
``calibration``
|
||||
Either a path to a ``ThirdPersonCalibration`` JSON file, or a dict
|
||||
of fitted parameters. May also be set once at construction time via
|
||||
``calibration_path=`` and reused across samples.
|
||||
|
||||
Optional sample keys
|
||||
--------------------
|
||||
``mouse_pitch_sign``
|
||||
``+1`` (default) or ``-1`` if the dataset's mouse-pitch sign is
|
||||
flipped (mhuo's data carries this in metadata).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.metrics.optical_flow._shared import (
|
||||
aggregate_temporal,
|
||||
compute_frame_metrics,
|
||||
extract_video_flows,
|
||||
load_ptlflow_model,
|
||||
)
|
||||
from fastvideo.eval.metrics.optical_flow.synthetic_optical_flow._thirdperson import (
|
||||
ThirdPersonCalibration,
|
||||
ThirdPersonFlowGenerator,
|
||||
load_calibration,
|
||||
)
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
def _resolve_calibration(obj: str | Path | dict | ThirdPersonCalibration, ) -> ThirdPersonCalibration:
|
||||
if isinstance(obj, ThirdPersonCalibration):
|
||||
return obj
|
||||
if isinstance(obj, dict):
|
||||
return ThirdPersonCalibration.from_dict(obj)
|
||||
return load_calibration(obj)
|
||||
|
||||
|
||||
@register("optical_flow.synthetic_optical_flow")
|
||||
class SyntheticOpticalFlowMetric(BaseMetric):
|
||||
"""Action-driven synthetic flow vs. video-extracted observed flow.
|
||||
|
||||
Pass ``calibration_path`` at construction to bind the calibration
|
||||
once across all samples; otherwise supply ``sample["calibration"]``
|
||||
per call. Missing actions or calibration produce a skipped result
|
||||
(``score=None``) rather than raising.
|
||||
"""
|
||||
|
||||
name = "optical_flow.synthetic_optical_flow"
|
||||
requires_reference = False
|
||||
higher_is_better = False
|
||||
needs_gpu = True
|
||||
backbone = "optical_flow"
|
||||
dependencies = ["ptlflow"]
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
model_name: str = "dpflow",
|
||||
ckpt: str = "things",
|
||||
calibration_path: str | Path | None = None,
|
||||
min_mag: float = 0.5,
|
||||
max_mag_pct: float = 80.0,
|
||||
grid_size: int = 8,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
self.model_name = model_name
|
||||
self.ckpt = ckpt
|
||||
self.min_mag = min_mag
|
||||
self.max_mag_pct = max_mag_pct
|
||||
self.grid_size = grid_size
|
||||
self._calibration: ThirdPersonCalibration | None = (_resolve_calibration(calibration_path)
|
||||
if calibration_path else None)
|
||||
self._model = None
|
||||
# See gt_optical_flow note: 1 frame pair per DPFlow forward.
|
||||
# Batching at 1080p OOMs because the cost volume is ~4 GB/pair.
|
||||
self._chunk_size = 1
|
||||
|
||||
def to(self, device: str | torch.device) -> SyntheticOpticalFlowMetric:
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
self._model = load_ptlflow_model(self.model_name, self.ckpt, self.device)
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
if self._model is None:
|
||||
self.setup()
|
||||
|
||||
actions = sample.get("actions")
|
||||
if actions is None:
|
||||
return self._skip(sample, "missing 'actions' (keyboard + mouse)")
|
||||
|
||||
cal_obj = sample.get("calibration")
|
||||
cal = self._calibration if cal_obj is None else _resolve_calibration(cal_obj)
|
||||
if cal is None:
|
||||
return self._skip(
|
||||
sample,
|
||||
"missing 'calibration' (pass calibration_path= at construction "
|
||||
"or sample['calibration'] per call)",
|
||||
)
|
||||
|
||||
video = sample["video"].float() # (T, C, H, W)
|
||||
T, _, H, W = video.shape
|
||||
if T < 2:
|
||||
raise ValueError("Need at least 2 frames to compute optical flow")
|
||||
n_pairs = T - 1
|
||||
|
||||
mouse_pitch_sign = int(sample.get("mouse_pitch_sign", 1))
|
||||
chunk = self._chunk_size or 16
|
||||
|
||||
observed = extract_video_flows(
|
||||
self._model,
|
||||
video,
|
||||
chunk=chunk,
|
||||
device=self.device,
|
||||
)
|
||||
predictor = ThirdPersonFlowGenerator(
|
||||
calibration=cal,
|
||||
frame_shape=(H, W),
|
||||
mouse_pitch_sign=mouse_pitch_sign,
|
||||
)
|
||||
predicted = predictor.generate_flow_sequence(actions, n_pairs=n_pairs)
|
||||
|
||||
n = min(len(observed), len(predicted))
|
||||
per_frame = [
|
||||
compute_frame_metrics(
|
||||
predicted[i],
|
||||
observed[i],
|
||||
grid_size=self.grid_size,
|
||||
min_mag=self.min_mag,
|
||||
max_mag_pct=self.max_mag_pct,
|
||||
) for i in range(n)
|
||||
]
|
||||
summary = aggregate_temporal(per_frame)
|
||||
score = summary.get("pixel_epe_mean_mean")
|
||||
details = dict(summary)
|
||||
details["per_frame_metrics"] = per_frame
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(score) if score is not None else None,
|
||||
details=details,
|
||||
)
|
||||
@@ -0,0 +1,19 @@
|
||||
descriptions.csv is vendored from the upstream Physics-IQ benchmark
|
||||
(https://github.com/google-deepmind/physics-IQ-benchmark, file
|
||||
descriptions/descriptions.csv) without modification. The associated video,
|
||||
mask, and switch-frame assets are NOT vendored — they auto-fetch on first
|
||||
use from gs://physics-iq-benchmark (public bucket; HTTPS-readable at
|
||||
https://storage.googleapis.com/physics-iq-benchmark/).
|
||||
|
||||
Licensed under Creative Commons Attribution 4.0 International (CC-BY-4.0).
|
||||
See https://creativecommons.org/licenses/by/4.0/legalcode for the full
|
||||
license text.
|
||||
|
||||
Citation:
|
||||
@article{motamed2025physics,
|
||||
title={Do generative video models understand physical principles?},
|
||||
author={Saman Motamed and Laura Culp and Kevin Swersky and
|
||||
Priyank Jaini and Robert Geirhos},
|
||||
journal={arXiv preprint arXiv:2501.09038},
|
||||
year={2025}
|
||||
}
|
||||
@@ -0,0 +1,397 @@
|
||||
scenario,description,category,generated_video_name
|
||||
0001_perspective-left_take-1_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0001_perspective-left_trimmed-ball-and-block-fall.mp4
|
||||
0002_perspective-center_take-1_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0002_perspective-center_trimmed-ball-and-block-fall.mp4
|
||||
0003_perspective-right_take-1_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0003_perspective-right_trimmed-ball-and-block-fall.mp4
|
||||
0004_perspective-left_take-1_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0004_perspective-left_trimmed-ball-behind-rotating-paper.mp4
|
||||
0005_perspective-center_take-1_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0005_perspective-center_trimmed-ball-behind-rotating-paper.mp4
|
||||
0006_perspective-right_take-1_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0006_perspective-right_trimmed-ball-behind-rotating-paper.mp4
|
||||
0007_perspective-left_take-1_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0007_perspective-left_trimmed-ball-hits-duck.mp4
|
||||
0008_perspective-center_take-1_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0008_perspective-center_trimmed-ball-hits-duck.mp4
|
||||
0009_perspective-right_take-1_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0009_perspective-right_trimmed-ball-hits-duck.mp4
|
||||
0010_perspective-left_take-1_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0010_perspective-left_trimmed-ball-hits-nothing.mp4
|
||||
0011_perspective-center_take-1_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0011_perspective-center_trimmed-ball-hits-nothing.mp4
|
||||
0012_perspective-right_take-1_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0012_perspective-right_trimmed-ball-hits-nothing.mp4
|
||||
0013_perspective-left_take-1_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0013_perspective-left_trimmed-ball-in-basket.mp4
|
||||
0014_perspective-center_take-1_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0014_perspective-center_trimmed-ball-in-basket.mp4
|
||||
0015_perspective-right_take-1_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0015_perspective-right_trimmed-ball-in-basket.mp4
|
||||
0016_perspective-left_take-1_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0016_perspective-left_trimmed-ball-in-sand.mp4
|
||||
0017_perspective-center_take-1_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0017_perspective-center_trimmed-ball-in-sand.mp4
|
||||
0018_perspective-right_take-1_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0018_perspective-right_trimmed-ball-in-sand.mp4
|
||||
0019_perspective-left_take-1_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0019_perspective-left_trimmed-ball-ramp.mp4
|
||||
0020_perspective-center_take-1_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0020_perspective-center_trimmed-ball-ramp.mp4
|
||||
0021_perspective-right_take-1_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0021_perspective-right_trimmed-ball-ramp.mp4
|
||||
0022_perspective-left_take-1_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0022_perspective-left_trimmed-ball-rolls-off.mp4
|
||||
0023_perspective-center_take-1_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0023_perspective-center_trimmed-ball-rolls-off.mp4
|
||||
0024_perspective-right_take-1_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0024_perspective-right_trimmed-ball-rolls-off.mp4
|
||||
0025_perspective-left_take-1_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0025_perspective-left_trimmed-ball-rolls-on-glass.mp4
|
||||
0026_perspective-center_take-1_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0026_perspective-center_trimmed-ball-rolls-on-glass.mp4
|
||||
0027_perspective-right_take-1_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0027_perspective-right_trimmed-ball-rolls-on-glass.mp4
|
||||
0028_perspective-left_take-1_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0028_perspective-left_trimmed-ball-train.mp4
|
||||
0029_perspective-center_take-1_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0029_perspective-center_trimmed-ball-train.mp4
|
||||
0030_perspective-right_take-1_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0030_perspective-right_trimmed-ball-train.mp4
|
||||
0031_perspective-left_take-1_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0031_perspective-left_trimmed-balls-collide.mp4
|
||||
0032_perspective-center_take-1_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0032_perspective-center_trimmed-balls-collide.mp4
|
||||
0033_perspective-right_take-1_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0033_perspective-right_trimmed-balls-collide.mp4
|
||||
0034_perspective-left_take-1_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0034_perspective-left_trimmed-block-domino.mp4
|
||||
0035_perspective-center_take-1_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0035_perspective-center_trimmed-block-domino.mp4
|
||||
0036_perspective-right_take-1_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0036_perspective-right_trimmed-block-domino.mp4
|
||||
0037_perspective-left_take-1_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0037_perspective-left_trimmed-blow-balloon.mp4
|
||||
0038_perspective-center_take-1_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0038_perspective-center_trimmed-blow-balloon.mp4
|
||||
0039_perspective-right_take-1_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0039_perspective-right_trimmed-blow-balloon.mp4
|
||||
0040_perspective-left_take-1_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0040_perspective-left_trimmed-cut-orange.mp4
|
||||
0041_perspective-center_take-1_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0041_perspective-center_trimmed-cut-orange.mp4
|
||||
0042_perspective-right_take-1_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0042_perspective-right_trimmed-cut-orange.mp4
|
||||
0043_perspective-left_take-1_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0043_perspective-left_trimmed-cut-paper.mp4
|
||||
0044_perspective-center_take-1_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0044_perspective-center_trimmed-cut-paper.mp4
|
||||
0045_perspective-right_take-1_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0045_perspective-right_trimmed-cut-paper.mp4
|
||||
0046_perspective-left_take-1_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0046_perspective-left_trimmed-domino-in-juice.mp4
|
||||
0047_perspective-center_take-1_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0047_perspective-center_trimmed-domino-in-juice.mp4
|
||||
0048_perspective-right_take-1_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0048_perspective-right_trimmed-domino-in-juice.mp4
|
||||
0049_perspective-left_take-1_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0049_perspective-left_trimmed-dominos-with-space.mp4
|
||||
0050_perspective-center_take-1_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0050_perspective-center_trimmed-dominos-with-space.mp4
|
||||
0051_perspective-right_take-1_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0051_perspective-right_trimmed-dominos-with-space.mp4
|
||||
0052_perspective-left_take-1_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0052_perspective-left_trimmed-double-cradle.mp4
|
||||
0053_perspective-center_take-1_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0053_perspective-center_trimmed-double-cradle.mp4
|
||||
0054_perspective-right_take-1_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0054_perspective-right_trimmed-double-cradle.mp4
|
||||
0055_perspective-left_take-1_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0055_perspective-left_trimmed-duck-and-dominos.mp4
|
||||
0056_perspective-center_take-1_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0056_perspective-center_trimmed-duck-and-dominos.mp4
|
||||
0057_perspective-right_take-1_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0057_perspective-right_trimmed-duck-and-dominos.mp4
|
||||
0058_perspective-left_take-1_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0058_perspective-left_trimmed-duck-falls-in-box.mp4
|
||||
0059_perspective-center_take-1_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0059_perspective-center_trimmed-duck-falls-in-box.mp4
|
||||
0060_perspective-right_take-1_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0060_perspective-right_trimmed-duck-falls-in-box.mp4
|
||||
0061_perspective-left_take-1_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0061_perspective-left_trimmed-duck-static.mp4
|
||||
0062_perspective-center_take-1_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0062_perspective-center_trimmed-duck-static.mp4
|
||||
0063_perspective-right_take-1_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0063_perspective-right_trimmed-duck-static.mp4
|
||||
0064_perspective-left_take-1_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0064_perspective-left_trimmed-fill-glass-red-drink.mp4
|
||||
0065_perspective-center_take-1_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0065_perspective-center_trimmed-fill-glass-red-drink.mp4
|
||||
0066_perspective-right_take-1_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0066_perspective-right_trimmed-fill-glass-red-drink.mp4
|
||||
0067_perspective-left_take-1_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0067_perspective-left_trimmed-glass-stays-same.mp4
|
||||
0068_perspective-center_take-1_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0068_perspective-center_trimmed-glass-stays-same.mp4
|
||||
0069_perspective-right_take-1_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0069_perspective-right_trimmed-glass-stays-same.mp4
|
||||
0070_perspective-left_take-1_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0070_perspective-left_trimmed-juice-in-water.mp4
|
||||
0071_perspective-center_take-1_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0071_perspective-center_trimmed-juice-in-water.mp4
|
||||
0072_perspective-right_take-1_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0072_perspective-right_trimmed-juice-in-water.mp4
|
||||
0073_perspective-left_take-1_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0073_perspective-left_trimmed-light-on-block.mp4
|
||||
0074_perspective-center_take-1_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0074_perspective-center_trimmed-light-on-block.mp4
|
||||
0075_perspective-right_take-1_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0075_perspective-right_trimmed-light-on-block.mp4
|
||||
0076_perspective-left_take-1_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0076_perspective-left_trimmed-light-on-mug.mp4
|
||||
0077_perspective-center_take-1_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0077_perspective-center_trimmed-light-on-mug.mp4
|
||||
0078_perspective-right_take-1_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0078_perspective-right_trimmed-light-on-mug.mp4
|
||||
0079_perspective-left_take-1_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0079_perspective-left_trimmed-light-on-mug-block.mp4
|
||||
0080_perspective-center_take-1_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0080_perspective-center_trimmed-light-on-mug-block.mp4
|
||||
0081_perspective-right_take-1_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0081_perspective-right_trimmed-light-on-mug-block.mp4
|
||||
0082_perspective-left_take-1_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0082_perspective-left_trimmed-light-on-statue.mp4
|
||||
0083_perspective-center_take-1_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0083_perspective-center_trimmed-light-on-statue.mp4
|
||||
0084_perspective-right_take-1_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0084_perspective-right_trimmed-light-on-statue.mp4
|
||||
0085_perspective-left_take-1_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0085_perspective-left_trimmed-liquid-on-duck.mp4
|
||||
0086_perspective-center_take-1_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0086_perspective-center_trimmed-liquid-on-duck.mp4
|
||||
0087_perspective-right_take-1_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0087_perspective-right_trimmed-liquid-on-duck.mp4
|
||||
0088_perspective-left_take-1_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0088_perspective-left_trimmed-liquid-overfill.mp4
|
||||
0089_perspective-center_take-1_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0089_perspective-center_trimmed-liquid-overfill.mp4
|
||||
0090_perspective-right_take-1_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0090_perspective-right_trimmed-liquid-overfill.mp4
|
||||
0091_perspective-left_take-1_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0091_perspective-left_trimmed-lit-candle.mp4
|
||||
0092_perspective-center_take-1_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0092_perspective-center_trimmed-lit-candle.mp4
|
||||
0093_perspective-right_take-1_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0093_perspective-right_trimmed-lit-candle.mp4
|
||||
0094_perspective-left_take-1_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0094_perspective-left_trimmed-magnet-domino.mp4
|
||||
0095_perspective-center_take-1_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0095_perspective-center_trimmed-magnet-domino.mp4
|
||||
0096_perspective-right_take-1_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0096_perspective-right_trimmed-magnet-domino.mp4
|
||||
0097_perspective-left_take-1_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0097_perspective-left_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0098_perspective-center_take-1_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0098_perspective-center_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0099_perspective-right_take-1_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0099_perspective-right_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0100_perspective-left_take-1_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0100_perspective-left_trimmed-magnet-wrench.mp4
|
||||
0101_perspective-center_take-1_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0101_perspective-center_trimmed-magnet-wrench.mp4
|
||||
0102_perspective-right_take-1_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0102_perspective-right_trimmed-magnet-wrench.mp4
|
||||
0103_perspective-left_take-1_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0103_perspective-left_trimmed-marble-run-x.mp4
|
||||
0104_perspective-center_take-1_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0104_perspective-center_trimmed-marble-run-x.mp4
|
||||
0105_perspective-right_take-1_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0105_perspective-right_trimmed-marble-run-x.mp4
|
||||
0106_perspective-left_take-1_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0106_perspective-left_trimmed-marble-run-y.mp4
|
||||
0107_perspective-center_take-1_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0107_perspective-center_trimmed-marble-run-y.mp4
|
||||
0108_perspective-right_take-1_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0108_perspective-right_trimmed-marble-run-y.mp4
|
||||
0109_perspective-left_take-1_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0109_perspective-left_trimmed-match.mp4
|
||||
0110_perspective-center_take-1_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0110_perspective-center_trimmed-match.mp4
|
||||
0111_perspective-right_take-1_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0111_perspective-right_trimmed-match.mp4
|
||||
0112_perspective-left_take-1_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0112_perspective-left_trimmed-match-blows-balloon.mp4
|
||||
0113_perspective-center_take-1_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0113_perspective-center_trimmed-match-blows-balloon.mp4
|
||||
0114_perspective-right_take-1_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0114_perspective-right_trimmed-match-blows-balloon.mp4
|
||||
0115_perspective-left_take-1_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0115_perspective-left_trimmed-mirror-ball-fall.mp4
|
||||
0116_perspective-center_take-1_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0116_perspective-center_trimmed-mirror-ball-fall.mp4
|
||||
0117_perspective-right_take-1_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0117_perspective-right_trimmed-mirror-ball-fall.mp4
|
||||
0118_perspective-left_take-1_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0118_perspective-left_trimmed-mirror-ball-rotate.mp4
|
||||
0119_perspective-center_take-1_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0119_perspective-center_trimmed-mirror-ball-rotate.mp4
|
||||
0120_perspective-right_take-1_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0120_perspective-right_trimmed-mirror-ball-rotate.mp4
|
||||
0121_perspective-left_take-1_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0121_perspective-left_trimmed-mirror-teapot-rotate.mp4
|
||||
0122_perspective-center_take-1_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0122_perspective-center_trimmed-mirror-teapot-rotate.mp4
|
||||
0123_perspective-right_take-1_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0123_perspective-right_trimmed-mirror-teapot-rotate.mp4
|
||||
0124_perspective-left_take-1_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0124_perspective-left_trimmed-mug-breaks.mp4
|
||||
0125_perspective-center_take-1_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0125_perspective-center_trimmed-mug-breaks.mp4
|
||||
0126_perspective-right_take-1_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0126_perspective-right_trimmed-mug-breaks.mp4
|
||||
0127_perspective-left_take-1_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0127_perspective-left_trimmed-napkin-soak.mp4
|
||||
0128_perspective-center_take-1_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0128_perspective-center_trimmed-napkin-soak.mp4
|
||||
0129_perspective-right_take-1_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0129_perspective-right_trimmed-napkin-soak.mp4
|
||||
0130_perspective-left_take-1_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0130_perspective-left_trimmed-paint-on-glass.mp4
|
||||
0131_perspective-center_take-1_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0131_perspective-center_trimmed-paint-on-glass.mp4
|
||||
0132_perspective-right_take-1_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0132_perspective-right_trimmed-paint-on-glass.mp4
|
||||
0133_perspective-left_take-1_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0133_perspective-left_trimmed-paper-fall-water.mp4
|
||||
0134_perspective-center_take-1_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0134_perspective-center_trimmed-paper-fall-water.mp4
|
||||
0135_perspective-right_take-1_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0135_perspective-right_trimmed-paper-fall-water.mp4
|
||||
0136_perspective-left_take-1_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0136_perspective-left_trimmed-paper-in-water.mp4
|
||||
0137_perspective-center_take-1_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0137_perspective-center_trimmed-paper-in-water.mp4
|
||||
0138_perspective-right_take-1_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0138_perspective-right_trimmed-paper-in-water.mp4
|
||||
0139_perspective-left_take-1_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0139_perspective-left_trimmed-paper-smoke.mp4
|
||||
0140_perspective-center_take-1_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0140_perspective-center_trimmed-paper-smoke.mp4
|
||||
0141_perspective-right_take-1_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0141_perspective-right_trimmed-paper-smoke.mp4
|
||||
0142_perspective-left_take-1_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0142_perspective-left_trimmed-potato-in-water.mp4
|
||||
0143_perspective-center_take-1_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0143_perspective-center_trimmed-potato-in-water.mp4
|
||||
0144_perspective-right_take-1_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0144_perspective-right_trimmed-potato-in-water.mp4
|
||||
0145_perspective-left_take-1_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0145_perspective-left_trimmed-roll-behind-box.mp4
|
||||
0146_perspective-center_take-1_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0146_perspective-center_trimmed-roll-behind-box.mp4
|
||||
0147_perspective-right_take-1_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0147_perspective-right_trimmed-roll-behind-box.mp4
|
||||
0148_perspective-left_take-1_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0148_perspective-left_trimmed-roll-front-box.mp4
|
||||
0149_perspective-center_take-1_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0149_perspective-center_trimmed-roll-front-box.mp4
|
||||
0150_perspective-right_take-1_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0150_perspective-right_trimmed-roll-front-box.mp4
|
||||
0151_perspective-left_take-1_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0151_perspective-left_trimmed-roll-in-box.mp4
|
||||
0152_perspective-center_take-1_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0152_perspective-center_trimmed-roll-in-box.mp4
|
||||
0153_perspective-right_take-1_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0153_perspective-right_trimmed-roll-in-box.mp4
|
||||
0154_perspective-left_take-1_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0154_perspective-left_trimmed-rolling-reflection.mp4
|
||||
0155_perspective-center_take-1_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0155_perspective-center_trimmed-rolling-reflection.mp4
|
||||
0156_perspective-right_take-1_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0156_perspective-right_trimmed-rolling-reflection.mp4
|
||||
0157_perspective-left_take-1_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0157_perspective-left_trimmed-silk-cover.mp4
|
||||
0158_perspective-center_take-1_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0158_perspective-center_trimmed-silk-cover.mp4
|
||||
0159_perspective-right_take-1_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0159_perspective-right_trimmed-silk-cover.mp4
|
||||
0160_perspective-left_take-1_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0160_perspective-left_trimmed-single-cradle.mp4
|
||||
0161_perspective-center_take-1_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0161_perspective-center_trimmed-single-cradle.mp4
|
||||
0162_perspective-right_take-1_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0162_perspective-right_trimmed-single-cradle.mp4
|
||||
0163_perspective-left_take-1_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0163_perspective-left_trimmed-siphon.mp4
|
||||
0164_perspective-center_take-1_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0164_perspective-center_trimmed-siphon.mp4
|
||||
0165_perspective-right_take-1_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0165_perspective-right_trimmed-siphon.mp4
|
||||
0166_perspective-left_take-1_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0166_perspective-left_trimmed-smiley-ball-rotates.mp4
|
||||
0167_perspective-center_take-1_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0167_perspective-center_trimmed-smiley-ball-rotates.mp4
|
||||
0168_perspective-right_take-1_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0168_perspective-right_trimmed-smiley-ball-rotates.mp4
|
||||
0169_perspective-left_take-1_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0169_perspective-left_trimmed-solid-ball-peakaboo.mp4
|
||||
0170_perspective-center_take-1_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0170_perspective-center_trimmed-solid-ball-peakaboo.mp4
|
||||
0171_perspective-right_take-1_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0171_perspective-right_trimmed-solid-ball-peakaboo.mp4
|
||||
0172_perspective-left_take-1_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0172_perspective-left_trimmed-stable-blocks.mp4
|
||||
0173_perspective-center_take-1_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0173_perspective-center_trimmed-stable-blocks.mp4
|
||||
0174_perspective-right_take-1_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0174_perspective-right_trimmed-stable-blocks.mp4
|
||||
0175_perspective-left_take-1_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0175_perspective-left_trimmed-teapot-rotates.mp4
|
||||
0176_perspective-center_take-1_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0176_perspective-center_trimmed-teapot-rotates.mp4
|
||||
0177_perspective-right_take-1_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0177_perspective-right_trimmed-teapot-rotates.mp4
|
||||
0178_perspective-left_take-1_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards eachother. Static shot with no camera movement.,Solid Mechanics,0178_perspective-left_trimmed-two-balls-pass.mp4
|
||||
0179_perspective-center_take-1_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards eachother. Static shot with no camera movement.,Solid Mechanics,0179_perspective-center_trimmed-two-balls-pass.mp4
|
||||
0180_perspective-right_take-1_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards eachother. Static shot with no camera movement.,Solid Mechanics,0180_perspective-right_trimmed-two-balls-pass.mp4
|
||||
0181_perspective-left_take-1_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0181_perspective-left_trimmed-unstable-block-stack.mp4
|
||||
0182_perspective-center_take-1_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0182_perspective-center_trimmed-unstable-block-stack.mp4
|
||||
0183_perspective-right_take-1_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0183_perspective-right_trimmed-unstable-block-stack.mp4
|
||||
0184_perspective-left_take-1_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0184_perspective-left_trimmed-water-in-juice.mp4
|
||||
0185_perspective-center_take-1_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0185_perspective-center_trimmed-water-in-juice.mp4
|
||||
0186_perspective-right_take-1_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0186_perspective-right_trimmed-water-in-juice.mp4
|
||||
0187_perspective-left_take-1_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0187_perspective-left_trimmed-weight-on-ceramic.mp4
|
||||
0188_perspective-center_take-1_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0188_perspective-center_trimmed-weight-on-ceramic.mp4
|
||||
0189_perspective-right_take-1_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0189_perspective-right_trimmed-weight-on-ceramic.mp4
|
||||
0190_perspective-left_take-1_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0190_perspective-left_trimmed-weight-on-paper.mp4
|
||||
0191_perspective-center_take-1_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0191_perspective-center_trimmed-weight-on-paper.mp4
|
||||
0192_perspective-right_take-1_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0192_perspective-right_trimmed-weight-on-paper.mp4
|
||||
0193_perspective-left_take-1_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0193_perspective-left_trimmed-weight-on-pillow.mp4
|
||||
0194_perspective-center_take-1_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0194_perspective-center_trimmed-weight-on-pillow.mp4
|
||||
0195_perspective-right_take-1_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0195_perspective-right_trimmed-weight-on-pillow.mp4
|
||||
0196_perspective-left_take-1_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0196_perspective-left_trimmed-weight-protects-duck.mp4
|
||||
0197_perspective-center_take-1_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0197_perspective-center_trimmed-weight-protects-duck.mp4
|
||||
0198_perspective-right_take-1_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0198_perspective-right_trimmed-weight-protects-duck.mp4
|
||||
0199_perspective-left_take-2_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The camera is static and the grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0199_perspective-left_trimmed-ball-and-block-fall.mp4
|
||||
0200_perspective-center_take-2_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The camera is static and the grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0200_perspective-center_trimmed-ball-and-block-fall.mp4
|
||||
0201_perspective-right_take-2_trimmed-ball-and-block-fall.mp4,Two pillows on a table and two grabber tools hanging above them from which a brown tennis ball and an orange block are suspended. The camera is static and the grabber tools let go of the ball and block. Static shot with no camera movement.,Solid Mechanics,0201_perspective-right_trimmed-ball-and-block-fall.mp4
|
||||
0202_perspective-left_take-2_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0202_perspective-left_trimmed-ball-behind-rotating-paper.mp4
|
||||
0203_perspective-center_take-2_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0203_perspective-center_trimmed-ball-behind-rotating-paper.mp4
|
||||
0204_perspective-right_take-2_trimmed-ball-behind-rotating-paper.mp4,A grabber arm is holding a tennis ball above a piece of cardstock propped up on a rotating platform sitting on a table that rotates clockwise. The grabber lowers the ball and places is on the table as the cardstock rotates. Static shot with no camera movement.,Solid Mechanics,0204_perspective-right_trimmed-ball-behind-rotating-paper.mp4
|
||||
0205_perspective-left_take-2_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0205_perspective-left_trimmed-ball-hits-duck.mp4
|
||||
0206_perspective-center_take-2_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0206_perspective-center_trimmed-ball-hits-duck.mp4
|
||||
0207_perspective-right_take-2_trimmed-ball-hits-duck.mp4,A light beige coffee table with a small yellow rubber ducky on it. A mustard yellow couch is in the background. There is a black pipe on one end of the table and a brown tennis ball rolls out of it towards the rubber ducky. Static shot with no camera movement.,Solid Mechanics,0207_perspective-right_trimmed-ball-hits-duck.mp4
|
||||
0208_perspective-left_take-2_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0208_perspective-left_trimmed-ball-hits-nothing.mp4
|
||||
0209_perspective-center_take-2_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0209_perspective-center_trimmed-ball-hits-nothing.mp4
|
||||
0210_perspective-right_take-2_trimmed-ball-hits-nothing.mp4,A light-colored wooden coffee table with a few small objects on it including a tennis ball and a smaller red ball. An orange ball rolls out of a black pipe that is sitting on the table towards the right side. Static shot with no camera movement.,Solid Mechanics,0210_perspective-right_trimmed-ball-hits-nothing.mp4
|
||||
0211_perspective-left_take-2_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0211_perspective-left_trimmed-ball-in-basket.mp4
|
||||
0212_perspective-center_take-2_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0212_perspective-center_trimmed-ball-in-basket.mp4
|
||||
0213_perspective-right_take-2_trimmed-ball-in-basket.mp4,An orange inflatable basketball is suspended above a black plastic crate placed on a wooden table. The ball is then released. Static shot with no camera movement.,Solid Mechanics,0213_perspective-right_trimmed-ball-in-basket.mp4
|
||||
0214_perspective-left_take-2_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0214_perspective-left_trimmed-ball-in-sand.mp4
|
||||
0215_perspective-center_take-2_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0215_perspective-center_trimmed-ball-in-sand.mp4
|
||||
0216_perspective-right_take-2_trimmed-ball-in-sand.mp4,A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball. Static shot with no camera movement.,Solid Mechanics,0216_perspective-right_trimmed-ball-in-sand.mp4
|
||||
0217_perspective-left_take-2_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0217_perspective-left_trimmed-ball-ramp.mp4
|
||||
0218_perspective-center_take-2_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0218_perspective-center_trimmed-ball-ramp.mp4
|
||||
0219_perspective-right_take-2_trimmed-ball-ramp.mp4,A simple ramp made of cardboard propped up by a blue block on a light-colored wooden table. There's a black pipe to the left of the frame and a yellow tennis ball rolls out of the pipe towards the ramp. Static shot with no camera movement.,Solid Mechanics,0219_perspective-right_trimmed-ball-ramp.mp4
|
||||
0220_perspective-left_take-2_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0220_perspective-left_trimmed-ball-rolls-off.mp4
|
||||
0221_perspective-center_take-2_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0221_perspective-center_trimmed-ball-rolls-off.mp4
|
||||
0222_perspective-right_take-2_trimmed-ball-rolls-off.mp4,A light wood coffee table in the foreground with a black pipe on the end of the table. A grey tennis ball rolls out of the pipe towards the right and onto the table. Static shot with no camera movement.,Solid Mechanics,0222_perspective-right_trimmed-ball-rolls-off.mp4
|
||||
0223_perspective-left_take-2_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0223_perspective-left_trimmed-ball-rolls-on-glass.mp4
|
||||
0224_perspective-center_take-2_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0224_perspective-center_trimmed-ball-rolls-on-glass.mp4
|
||||
0225_perspective-right_take-2_trimmed-ball-rolls-on-glass.mp4,A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. A blue tennis ball rolls on the wooden table and towards the glass. Static shot with no camera movement.,Solid Mechanics,0225_perspective-right_trimmed-ball-rolls-on-glass.mp4
|
||||
0226_perspective-left_take-2_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0226_perspective-left_trimmed-ball-train.mp4
|
||||
0227_perspective-center_take-2_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0227_perspective-center_trimmed-ball-train.mp4
|
||||
0228_perspective-right_take-2_trimmed-ball-train.mp4,"A light-colored coffee table with two tennis balls, one orange and one brown, placed near the center back to back. A grey tennis ball rolls out of a black pipe sitting on the table and towards the other two balls. Static shot with no camera movement.",Solid Mechanics,0228_perspective-right_trimmed-ball-train.mp4
|
||||
0229_perspective-left_take-2_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0229_perspective-left_trimmed-balls-collide.mp4
|
||||
0230_perspective-center_take-2_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0230_perspective-center_trimmed-balls-collide.mp4
|
||||
0231_perspective-right_take-2_trimmed-balls-collide.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0231_perspective-right_trimmed-balls-collide.mp4
|
||||
0232_perspective-left_take-2_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0232_perspective-left_trimmed-block-domino.mp4
|
||||
0233_perspective-center_take-2_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0233_perspective-center_trimmed-block-domino.mp4
|
||||
0234_perspective-right_take-2_trimmed-block-domino.mp4,A row of colorful wooden blocks lined up on a wooden table with a wooden stick attached to a black rotating platform. The platform rotates clockwise and the wooden stick hits the first block as it rotates. Static shot with no camera movement.,Solid Mechanics,0234_perspective-right_trimmed-block-domino.mp4
|
||||
0235_perspective-left_take-2_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0235_perspective-left_trimmed-blow-balloon.mp4
|
||||
0236_perspective-center_take-2_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0236_perspective-center_trimmed-blow-balloon.mp4
|
||||
0237_perspective-right_take-2_trimmed-blow-balloon.mp4,A black balloon is attached to a fixed electric air pump hose on a table with a plain wall in the background. Air is being pumped in the balloon. Static shot with no camera movement.,Fluid Dynamics,0237_perspective-right_trimmed-blow-balloon.mp4
|
||||
0238_perspective-left_take-2_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0238_perspective-left_trimmed-cut-orange.mp4
|
||||
0239_perspective-center_take-2_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0239_perspective-center_trimmed-cut-orange.mp4
|
||||
0240_perspective-right_take-2_trimmed-cut-orange.mp4,A tangerine that has been cut in half is placed on a glass cutting board. A knife is slicing through the tangerine. Static shot with no camera movement.,Solid Mechanics,0240_perspective-right_trimmed-cut-orange.mp4
|
||||
0241_perspective-left_take-2_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0241_perspective-left_trimmed-cut-paper.mp4
|
||||
0242_perspective-center_take-2_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0242_perspective-center_trimmed-cut-paper.mp4
|
||||
0243_perspective-right_take-2_trimmed-cut-paper.mp4,"Two black and blue gripping tools are pulling a piece of green paper from its two corners, causing it to tear. Static shot with no camera movement.",Solid Mechanics,0243_perspective-right_trimmed-cut-paper.mp4
|
||||
0244_perspective-left_take-2_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0244_perspective-left_trimmed-domino-in-juice.mp4
|
||||
0245_perspective-center_take-2_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0245_perspective-center_trimmed-domino-in-juice.mp4
|
||||
0246_perspective-right_take-2_trimmed-domino-in-juice.mp4,A grabber tool holding a white domino drops the domino into a dark-colored liquid in a blue mug that is on a wooden surface. Static shot with no camera movement.,Fluid Dynamics,0246_perspective-right_trimmed-domino-in-juice.mp4
|
||||
0247_perspective-left_take-2_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0247_perspective-left_trimmed-dominos-with-space.mp4
|
||||
0248_perspective-center_take-2_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0248_perspective-center_trimmed-dominos-with-space.mp4
|
||||
0249_perspective-right_take-2_trimmed-dominos-with-space.mp4,Two rows of alternating black and white dominoes are set up on a wooden table with a gap between the two rows. A wooden stick attached to a rotating platform rotates clockwise and knocks the first domino in the first row. Static shot with no camera movement.,Solid Mechanics,0249_perspective-right_trimmed-dominos-with-space.mp4
|
||||
0250_perspective-left_take-2_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0250_perspective-left_trimmed-double-cradle.mp4
|
||||
0251_perspective-center_take-2_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0251_perspective-center_trimmed-double-cradle.mp4
|
||||
0252_perspective-right_take-2_trimmed-double-cradle.mp4,A Newton's cradle device on the table and two of the metal balls are held up by a blue handled grabber tool. The claw releases the two balls. Static shot with no camera movement.,Solid Mechanics,0252_perspective-right_trimmed-double-cradle.mp4
|
||||
0253_perspective-left_take-2_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0253_perspective-left_trimmed-duck-and-dominos.mp4
|
||||
0254_perspective-center_take-2_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0254_perspective-center_trimmed-duck-and-dominos.mp4
|
||||
0255_perspective-right_take-2_trimmed-duck-and-dominos.mp4,A yellow rubber duck is positioned in the middle of a line of black and white dominoes on a wooden table. A stick attached to a black rotating platform rotates clockwise and knocks the first domino block. Static shot with no camera movement.,Solid Mechanics,0255_perspective-right_trimmed-duck-and-dominos.mp4
|
||||
0256_perspective-left_take-2_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0256_perspective-left_trimmed-duck-falls-in-box.mp4
|
||||
0257_perspective-center_take-2_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0257_perspective-center_trimmed-duck-falls-in-box.mp4
|
||||
0258_perspective-right_take-2_trimmed-duck-falls-in-box.mp4,A yellow rubber ducky is suspended above an open dark green fabric box on a wooden table. The duck is then released. Static shot with no camera movement.,Solid Mechanics,0258_perspective-right_trimmed-duck-falls-in-box.mp4
|
||||
0259_perspective-left_take-2_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0259_perspective-left_trimmed-duck-static.mp4
|
||||
0260_perspective-center_take-2_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0260_perspective-center_trimmed-duck-static.mp4
|
||||
0261_perspective-right_take-2_trimmed-duck-static.mp4,A stationary yellow rubber duck on a light brown wooden table against a plain white background. Static shot with no camera movement.,Solid Mechanics,0261_perspective-right_trimmed-duck-static.mp4
|
||||
0262_perspective-left_take-2_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0262_perspective-left_trimmed-fill-glass-red-drink.mp4
|
||||
0263_perspective-center_take-2_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0263_perspective-center_trimmed-fill-glass-red-drink.mp4
|
||||
0264_perspective-right_take-2_trimmed-fill-glass-red-drink.mp4,A glass beverage dispenser filled with a bright red liquid is set up on a woven basket and is pouring the liquid into a clear glass on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0264_perspective-right_trimmed-fill-glass-red-drink.mp4
|
||||
0265_perspective-left_take-2_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0265_perspective-left_trimmed-glass-stays-same.mp4
|
||||
0266_perspective-center_take-2_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0266_perspective-center_trimmed-glass-stays-same.mp4
|
||||
0267_perspective-right_take-2_trimmed-glass-stays-same.mp4,A glass beverage dispenser filled with a bright red liquid is set on a wicker base. Under the dispenser there is a glass half filled with red liquid. Static shot with no camera movement.,Fluid Dynamics,0267_perspective-right_trimmed-glass-stays-same.mp4
|
||||
0268_perspective-left_take-2_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0268_perspective-left_trimmed-juice-in-water.mp4
|
||||
0269_perspective-center_take-2_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0269_perspective-center_trimmed-juice-in-water.mp4
|
||||
0270_perspective-right_take-2_trimmed-juice-in-water.mp4,A glass beverage dispenser pouring grapefruit juice into a glass that has some water inside. Static shot with no camera movement.,Fluid Dynamics,0270_perspective-right_trimmed-juice-in-water.mp4
|
||||
0271_perspective-left_take-2_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0271_perspective-left_trimmed-light-on-block.mp4
|
||||
0272_perspective-center_take-2_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0272_perspective-center_trimmed-light-on-block.mp4
|
||||
0273_perspective-right_take-2_trimmed-light-on-block.mp4,A blue rectangular wooden block is placed on a black rotating turntable that rotates clockwise illuminated by a spotlight casting a long shadow on the wall behind it. Static shot with no camera movement.,Optics,0273_perspective-right_trimmed-light-on-block.mp4
|
||||
0274_perspective-left_take-2_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0274_perspective-left_trimmed-light-on-mug.mp4
|
||||
0275_perspective-center_take-2_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0275_perspective-center_trimmed-light-on-mug.mp4
|
||||
0276_perspective-right_take-2_trimmed-light-on-mug.mp4,A yellow mug is placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting a shadow on the wall behind it. Static shot with no camera movement.,Optics,0276_perspective-right_trimmed-light-on-mug.mp4
|
||||
0277_perspective-left_take-2_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0277_perspective-left_trimmed-light-on-mug-block.mp4
|
||||
0278_perspective-center_take-2_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0278_perspective-center_trimmed-light-on-mug-block.mp4
|
||||
0279_perspective-right_take-2_trimmed-light-on-mug-block.mp4,A yellow mug and a blue wooden block are placed on a rotating turntable that rotates clockwise illuminated by a spotlight casting their shadow on the wall behind it. Static shot with no camera movement.,Optics,0279_perspective-right_trimmed-light-on-mug-block.mp4
|
||||
0280_perspective-left_take-2_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0280_perspective-left_trimmed-light-on-statue.mp4
|
||||
0281_perspective-center_take-2_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0281_perspective-center_trimmed-light-on-statue.mp4
|
||||
0282_perspective-right_take-2_trimmed-light-on-statue.mp4,A small statue made of porcelain illuminated by a spotlight on a rotating base that rotates clockwise. The spotlight casts a large shadow of the statue onto the wall behind it. Static shot with no camera movement.,Optics,0282_perspective-right_trimmed-light-on-statue.mp4
|
||||
0283_perspective-left_take-2_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0283_perspective-left_trimmed-liquid-on-duck.mp4
|
||||
0284_perspective-center_take-2_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0284_perspective-center_trimmed-liquid-on-duck.mp4
|
||||
0285_perspective-right_take-2_trimmed-liquid-on-duck.mp4,A yellow rubber ducky is placed in an empty black baking pan on a wooden table. A beverage dispenser with red liquid inside sits on a woven basket behind it. The liquid pours on the duck from the dispenser. Static shot with no camera movement.,Fluid Dynamics,0285_perspective-right_trimmed-liquid-on-duck.mp4
|
||||
0286_perspective-left_take-2_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0286_perspective-left_trimmed-liquid-overfill.mp4
|
||||
0287_perspective-center_take-2_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0287_perspective-center_trimmed-liquid-overfill.mp4
|
||||
0288_perspective-right_take-2_trimmed-liquid-overfill.mp4,A bright red liquid being poured from a dispenser into a glass which is placed on a dark baking tray on a wooden table. Static shot with no camera movement.,Fluid Dynamics,0288_perspective-right_trimmed-liquid-overfill.mp4
|
||||
0289_perspective-left_take-2_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0289_perspective-left_trimmed-lit-candle.mp4
|
||||
0290_perspective-center_take-2_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0290_perspective-center_trimmed-lit-candle.mp4
|
||||
0291_perspective-right_take-2_trimmed-lit-candle.mp4,Two candle holders that have tall red candles in them are placed on a wooden table. One of the candles is burning. Static shot with no camera movement.,Thermodynamics,0291_perspective-right_trimmed-lit-candle.mp4
|
||||
0292_perspective-left_take-2_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0292_perspective-left_trimmed-magnet-domino.mp4
|
||||
0293_perspective-center_take-2_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0293_perspective-center_trimmed-magnet-domino.mp4
|
||||
0294_perspective-right_take-2_trimmed-magnet-domino.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small white plastic domino block is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0294_perspective-right_trimmed-magnet-domino.mp4
|
||||
0295_perspective-left_take-2_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0295_perspective-left_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0296_perspective-center_take-2_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0296_perspective-center_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0297_perspective-right_take-2_trimmed-magnet-transparent-peakaboo.mp4,A clear acrylic box suspended from a cord hangs above a tennis ball with a smiley face drawn on it positioned on a wooden table. The box is lowered to cover the ball. Static shot with no camera movement.,Solid Mechanics,0297_perspective-right_trimmed-magnet-transparent-peakaboo.mp4
|
||||
0298_perspective-left_take-2_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0298_perspective-left_trimmed-magnet-wrench.mp4
|
||||
0299_perspective-center_take-2_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0299_perspective-center_trimmed-magnet-wrench.mp4
|
||||
0300_perspective-right_take-2_trimmed-magnet-wrench.mp4,A powerful magnet is placed on the table facing a black rotating platform that rotates clockwise. A small metal wrench is placed on the platform and is rotating towards the magnet. Static shot with no camera movement.,Magnetism,0300_perspective-right_trimmed-magnet-wrench.mp4
|
||||
0301_perspective-left_take-2_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0301_perspective-left_trimmed-marble-run-x.mp4
|
||||
0302_perspective-center_take-2_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0302_perspective-center_trimmed-marble-run-x.mp4
|
||||
0303_perspective-right_take-2_trimmed-marble-run-x.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0303_perspective-right_trimmed-marble-run-x.mp4
|
||||
0304_perspective-left_take-2_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0304_perspective-left_trimmed-marble-run-y.mp4
|
||||
0305_perspective-center_take-2_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0305_perspective-center_trimmed-marble-run-y.mp4
|
||||
0306_perspective-right_take-2_trimmed-marble-run-y.mp4,A few magnetic ramps are attached to a whiteboard for a game of marble run. A yellow marble is released at the top of the ramps and slides down the ramps. Static shot with no camera movement.,Solid Mechanics,0306_perspective-right_trimmed-marble-run-y.mp4
|
||||
0307_perspective-left_take-2_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0307_perspective-left_trimmed-match.mp4
|
||||
0308_perspective-center_take-2_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0308_perspective-center_trimmed-match.mp4
|
||||
0309_perspective-right_take-2_trimmed-match.mp4,A lit match is being lowered into a glass of water. Static shot with no camera movement.,Fluid Dynamics,0309_perspective-right_trimmed-match.mp4
|
||||
0310_perspective-left_take-2_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0310_perspective-left_trimmed-match-blows-balloon.mp4
|
||||
0311_perspective-center_take-2_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0311_perspective-center_trimmed-match-blows-balloon.mp4
|
||||
0312_perspective-right_take-2_trimmed-match-blows-balloon.mp4,A black balloon is sitting on a wooden table next to a small rotating platform with a lit matchstick taped to it. The match rotates clockwise and touches the balloon. Static shot with no camera movement.,Thermodynamics,0312_perspective-right_trimmed-match-blows-balloon.mp4
|
||||
0313_perspective-left_take-2_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0313_perspective-left_trimmed-mirror-ball-fall.mp4
|
||||
0314_perspective-center_take-2_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0314_perspective-center_trimmed-mirror-ball-fall.mp4
|
||||
0315_perspective-right_take-2_trimmed-mirror-ball-fall.mp4,A tennis ball attached to a magnet and string is hanging in front of a mirror and creating an illusion of two tennis balls. The string lowers the ball slowly. Static shot with no camera movement.,Optics,0315_perspective-right_trimmed-mirror-ball-fall.mp4
|
||||
0316_perspective-left_take-2_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0316_perspective-left_trimmed-mirror-ball-rotate.mp4
|
||||
0317_perspective-center_take-2_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0317_perspective-center_trimmed-mirror-ball-rotate.mp4
|
||||
0318_perspective-right_take-2_trimmed-mirror-ball-rotate.mp4,A tennis ball with a smiley face drawn on it is slowly rotating on a black rotating platform that rotates clockwise in front of a mirror and reflecting the side of the ball that has no smiley face. Static shot with no camera movement.,Optics,0318_perspective-right_trimmed-mirror-ball-rotate.mp4
|
||||
0319_perspective-left_take-2_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0319_perspective-left_trimmed-mirror-teapot-rotate.mp4
|
||||
0320_perspective-center_take-2_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0320_perspective-center_trimmed-mirror-teapot-rotate.mp4
|
||||
0321_perspective-right_take-2_trimmed-mirror-teapot-rotate.mp4,A teapot on a rotating display base that rotates clockwise in front of a mirror reflecting the teapot's image. Static shot with no camera movement.,Optics,0321_perspective-right_trimmed-mirror-teapot-rotate.mp4
|
||||
0322_perspective-left_take-2_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0322_perspective-left_trimmed-mug-breaks.mp4
|
||||
0323_perspective-center_take-2_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0323_perspective-center_trimmed-mug-breaks.mp4
|
||||
0324_perspective-right_take-2_trimmed-mug-breaks.mp4,A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug. Static shot with no camera movement.,Solid Mechanics,0324_perspective-right_trimmed-mug-breaks.mp4
|
||||
0325_perspective-left_take-2_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0325_perspective-left_trimmed-napkin-soak.mp4
|
||||
0326_perspective-center_take-2_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0326_perspective-center_trimmed-napkin-soak.mp4
|
||||
0327_perspective-right_take-2_trimmed-napkin-soak.mp4,A grabber tool holds a piece of paper towel over a shallow dish of light blue liquid on a wooden table. The grabber releases the paper towel on the dish. Static shot with no camera movement.,Fluid Dynamics,0327_perspective-right_trimmed-napkin-soak.mp4
|
||||
0328_perspective-left_take-2_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0328_perspective-left_trimmed-paint-on-glass.mp4
|
||||
0329_perspective-center_take-2_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0329_perspective-center_trimmed-paint-on-glass.mp4
|
||||
0330_perspective-right_take-2_trimmed-paint-on-glass.mp4,A clear acrylic sheet placed on a wooden table with a small dollop of red paint. A rotating paintbrush attached to a rotating platform rotates clockwise and goes through the paint. Static shot with no camera movement.,Fluid Dynamics,0330_perspective-right_trimmed-paint-on-glass.mp4
|
||||
0331_perspective-left_take-2_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0331_perspective-left_trimmed-paper-fall-water.mp4
|
||||
0332_perspective-center_take-2_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0332_perspective-center_trimmed-paper-fall-water.mp4
|
||||
0333_perspective-right_take-2_trimmed-paper-fall-water.mp4,A grabber tool is holding a crumpled piece of paper over a bowl of water on a wooden table. The grabber then releases the crumpled paper onto the bowl. Static shot with no camera movement.,Fluid Dynamics,0333_perspective-right_trimmed-paper-fall-water.mp4
|
||||
0334_perspective-left_take-2_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0334_perspective-left_trimmed-paper-in-water.mp4
|
||||
0335_perspective-center_take-2_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0335_perspective-center_trimmed-paper-in-water.mp4
|
||||
0336_perspective-right_take-2_trimmed-paper-in-water.mp4,A small piece of crumpled white paper is being lowered into a tall glass containing blue liquid with a green band showing the water level. The crumpled paper is released into the glass. Static shot with no camera movement.,Fluid Dynamics,0336_perspective-right_trimmed-paper-in-water.mp4
|
||||
0337_perspective-left_take-2_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0337_perspective-left_trimmed-paper-smoke.mp4
|
||||
0338_perspective-center_take-2_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0338_perspective-center_trimmed-paper-smoke.mp4
|
||||
0339_perspective-right_take-2_trimmed-paper-smoke.mp4,A piece of folded paper is placed on a glass cutting board. The paper is being burnt and white smoke is emitting from it. Static shot with no camera movement.,Thermodynamics,0339_perspective-right_trimmed-paper-smoke.mp4
|
||||
0340_perspective-left_take-2_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0340_perspective-left_trimmed-potato-in-water.mp4
|
||||
0341_perspective-center_take-2_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0341_perspective-center_trimmed-potato-in-water.mp4
|
||||
0342_perspective-right_take-2_trimmed-potato-in-water.mp4,A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass. Static shot with no camera movement.,Fluid Dynamics,0342_perspective-right_trimmed-potato-in-water.mp4
|
||||
0343_perspective-left_take-2_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0343_perspective-left_trimmed-roll-behind-box.mp4
|
||||
0344_perspective-center_take-2_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0344_perspective-center_trimmed-roll-behind-box.mp4
|
||||
0345_perspective-right_take-2_trimmed-roll-behind-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0345_perspective-right_trimmed-roll-behind-box.mp4
|
||||
0346_perspective-left_take-2_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0346_perspective-left_trimmed-roll-front-box.mp4
|
||||
0347_perspective-center_take-2_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0347_perspective-center_trimmed-roll-front-box.mp4
|
||||
0348_perspective-right_take-2_trimmed-roll-front-box.mp4,A small white lampshade is on a light wood surface. A grey tennis ball rolls out of the black tube sitting on the table and rolls on the table towards the right. Static shot with no camera movement.,Solid Mechanics,0348_perspective-right_trimmed-roll-front-box.mp4
|
||||
0349_perspective-left_take-2_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0349_perspective-left_trimmed-roll-in-box.mp4
|
||||
0350_perspective-center_take-2_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0350_perspective-center_trimmed-roll-in-box.mp4
|
||||
0351_perspective-right_take-2_trimmed-roll-in-box.mp4,An olive green fabric box is on a light wood surface. A brown tennis ball rolls out of the black tube sitting on the table and rolls towards the box. Static shot with no camera movement.,Solid Mechanics,0351_perspective-right_trimmed-roll-in-box.mp4
|
||||
0352_perspective-left_take-2_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0352_perspective-left_trimmed-rolling-reflection.mp4
|
||||
0353_perspective-center_take-2_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0353_perspective-center_trimmed-rolling-reflection.mp4
|
||||
0354_perspective-right_take-2_trimmed-rolling-reflection.mp4,A 30lb kettlebell resting on a wooden table next to a mirror. A tennis ball rolls towards the kettlebell. Static shot with no camera movement.,Optics,0354_perspective-right_trimmed-rolling-reflection.mp4
|
||||
0355_perspective-left_take-2_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0355_perspective-left_trimmed-silk-cover.mp4
|
||||
0356_perspective-center_take-2_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0356_perspective-center_trimmed-silk-cover.mp4
|
||||
0357_perspective-right_take-2_trimmed-silk-cover.mp4,A teapot is placed on a wooden table. a piece of silk fabric is lowered on the teapot to cover it. Static shot with no camera movement.,Solid Mechanics,0357_perspective-right_trimmed-silk-cover.mp4
|
||||
0358_perspective-left_take-2_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0358_perspective-left_trimmed-single-cradle.mp4
|
||||
0359_perspective-center_take-2_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0359_perspective-center_trimmed-single-cradle.mp4
|
||||
0360_perspective-right_take-2_trimmed-single-cradle.mp4,A Newton's cradle device on the table and one of the metal balls is held up by a blue handled grabber tool. The claw releases the ball. Static shot with no camera movement.,Solid Mechanics,0360_perspective-right_trimmed-single-cradle.mp4
|
||||
0361_perspective-left_take-2_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0361_perspective-left_trimmed-siphon.mp4
|
||||
0362_perspective-center_take-2_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0362_perspective-center_trimmed-siphon.mp4
|
||||
0363_perspective-right_take-2_trimmed-siphon.mp4,A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks. Static shot with no camera movement.,Fluid Dynamics,0363_perspective-right_trimmed-siphon.mp4
|
||||
0364_perspective-left_take-2_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0364_perspective-left_trimmed-smiley-ball-rotates.mp4
|
||||
0365_perspective-center_take-2_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0365_perspective-center_trimmed-smiley-ball-rotates.mp4
|
||||
0366_perspective-right_take-2_trimmed-smiley-ball-rotates.mp4,A tennis ball with a smiley face drawn on it is placed on a rotating black platform that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0366_perspective-right_trimmed-smiley-ball-rotates.mp4
|
||||
0367_perspective-left_take-2_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0367_perspective-left_trimmed-solid-ball-peakaboo.mp4
|
||||
0368_perspective-center_take-2_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0368_perspective-center_trimmed-solid-ball-peakaboo.mp4
|
||||
0369_perspective-right_take-2_trimmed-solid-ball-peakaboo.mp4,A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again. Static shot with no camera movement.,Solid Mechanics,0369_perspective-right_trimmed-solid-ball-peakaboo.mp4
|
||||
0370_perspective-left_take-2_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0370_perspective-left_trimmed-stable-blocks.mp4
|
||||
0371_perspective-center_take-2_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0371_perspective-center_trimmed-stable-blocks.mp4
|
||||
0372_perspective-right_take-2_trimmed-stable-blocks.mp4,A pink block is being lowered towards a simple structure made of colorful blocks resembling a gate. Static shot with no camera movement.,Solid Mechanics,0372_perspective-right_trimmed-stable-blocks.mp4
|
||||
0373_perspective-left_take-2_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0373_perspective-left_trimmed-teapot-rotates.mp4
|
||||
0374_perspective-center_take-2_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0374_perspective-center_trimmed-teapot-rotates.mp4
|
||||
0375_perspective-right_take-2_trimmed-teapot-rotates.mp4,A teapot is placed on a rotating display that rotates clockwise. Static shot with no camera movement.,Solid Mechanics,0375_perspective-right_trimmed-teapot-rotates.mp4
|
||||
0376_perspective-left_take-2_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0376_perspective-left_trimmed-two-balls-pass.mp4
|
||||
0377_perspective-center_take-2_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0377_perspective-center_trimmed-two-balls-pass.mp4
|
||||
0378_perspective-right_take-2_trimmed-two-balls-pass.mp4,A light-colored wooden tabletop with two pipes at the edges. A blue and yellow tennis ball roll out of the pipes and towards each other. Static shot with no camera movement.,Solid Mechanics,0378_perspective-right_trimmed-two-balls-pass.mp4
|
||||
0379_perspective-left_take-2_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0379_perspective-left_trimmed-unstable-block-stack.mp4
|
||||
0380_perspective-center_take-2_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0380_perspective-center_trimmed-unstable-block-stack.mp4
|
||||
0381_perspective-right_take-2_trimmed-unstable-block-stack.mp4,A grabber tool carefully placing a blue wooden block on top of a yellow block which is balanced on a red block forming an L shape. Static shot with no camera movement.,Solid Mechanics,0381_perspective-right_trimmed-unstable-block-stack.mp4
|
||||
0382_perspective-left_take-2_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0382_perspective-left_trimmed-water-in-juice.mp4
|
||||
0383_perspective-center_take-2_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0383_perspective-center_trimmed-water-in-juice.mp4
|
||||
0384_perspective-right_take-2_trimmed-water-in-juice.mp4,A glass beverage dispenser is dispensing water into a glass which has some grapefruit juice in it. Static shot with no camera movement.,Fluid Dynamics,0384_perspective-right_trimmed-water-in-juice.mp4
|
||||
0385_perspective-left_take-2_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0385_perspective-left_trimmed-weight-on-ceramic.mp4
|
||||
0386_perspective-center_take-2_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0386_perspective-center_trimmed-weight-on-ceramic.mp4
|
||||
0387_perspective-right_take-2_trimmed-weight-on-ceramic.mp4,A 30lb kettlebell is slowly lowered on top of a yellow ceramic coffee mug placed on a wooden table. Static shot with no camera movement.,Solid Mechanics,0387_perspective-right_trimmed-weight-on-ceramic.mp4
|
||||
0388_perspective-left_take-2_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0388_perspective-left_trimmed-weight-on-paper.mp4
|
||||
0389_perspective-center_take-2_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0389_perspective-center_trimmed-weight-on-paper.mp4
|
||||
0390_perspective-right_take-2_trimmed-weight-on-paper.mp4,A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.,Solid Mechanics,0390_perspective-right_trimmed-weight-on-paper.mp4
|
||||
0391_perspective-left_take-2_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0391_perspective-left_trimmed-weight-on-pillow.mp4
|
||||
0392_perspective-center_take-2_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0392_perspective-center_trimmed-weight-on-pillow.mp4
|
||||
0393_perspective-right_take-2_trimmed-weight-on-pillow.mp4,A 30lb kettlebell and a green piece of paper are lowered onto two pillows. Static shot with no camera movement.,Solid Mechanics,0393_perspective-right_trimmed-weight-on-pillow.mp4
|
||||
0394_perspective-left_take-2_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0394_perspective-left_trimmed-weight-protects-duck.mp4
|
||||
0395_perspective-center_take-2_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0395_perspective-center_trimmed-weight-protects-duck.mp4
|
||||
0396_perspective-right_take-2_trimmed-weight-protects-duck.mp4,A light beige coffee table with a black kettlebell and a yellow rubber duck on it. A grey tennis ball rolls out of the black tube sitting on the table and towards the duck and kettlebell. Static shot with no camera movement.,Solid Mechanics,0396_perspective-right_trimmed-weight-protects-duck.mp4
|
||||
|
@@ -0,0 +1,194 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
from collections.abc import Iterable, Mapping
|
||||
|
||||
import numpy as np
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.metrics.physics_iq.mse.metric import PhysicsIQMSEMetric
|
||||
from fastvideo.eval.metrics.physics_iq.spatial_iou.metric import SpatialIoUMetric
|
||||
from fastvideo.eval.metrics.physics_iq.spatiotemporal_iou.metric import SpatiotemporalIoUMetric
|
||||
from fastvideo.eval.metrics.physics_iq.weighted_spatial_iou.metric import WeightedSpatialIoUMetric
|
||||
from fastvideo.eval.metrics.physics_iq.utils import (
|
||||
DEFAULT_DURATION_SECONDS,
|
||||
DEFAULT_TARGET_FPS,
|
||||
mean,
|
||||
prepare_pair_inputs,
|
||||
prepare_triplet_inputs,
|
||||
)
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("physics_iq")
|
||||
class PhysicsIQMetric(BaseMetric):
|
||||
name = "physics_iq"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
target_fps: int = DEFAULT_TARGET_FPS,
|
||||
duration_seconds: int = DEFAULT_DURATION_SECONDS,
|
||||
video_time_selection: str = "first",
|
||||
threshold: int = 10,
|
||||
alpha: float = 0.3,
|
||||
roundtrip_generated_masks: bool = True,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
self._prep_kwargs = {
|
||||
"target_fps": target_fps,
|
||||
"duration_seconds": duration_seconds,
|
||||
"video_time_selection": video_time_selection,
|
||||
"threshold": threshold,
|
||||
"alpha": alpha,
|
||||
"roundtrip_generated_masks": roundtrip_generated_masks,
|
||||
}
|
||||
self._mse = PhysicsIQMSEMetric(**self._prep_kwargs)
|
||||
self._spatiotemporal_iou = SpatiotemporalIoUMetric(**self._prep_kwargs)
|
||||
self._spatial_iou = SpatialIoUMetric(**self._prep_kwargs)
|
||||
self._weighted_spatial_iou = WeightedSpatialIoUMetric(**self._prep_kwargs)
|
||||
|
||||
@staticmethod
|
||||
def _extract_payload(result: MetricResult | Mapping[str, Any]) -> Mapping[str, Any]:
|
||||
if isinstance(result, MetricResult):
|
||||
return result.details
|
||||
return result
|
||||
|
||||
def _compute_pair_metrics(self, prepared_pair) -> dict[str, Any]:
|
||||
sample = {"_physics_iq_pair": prepared_pair}
|
||||
mse = self._mse.compute(sample)
|
||||
st = self._spatiotemporal_iou.compute(sample)
|
||||
spatial = self._spatial_iou.compute(sample)
|
||||
weighted = self._weighted_spatial_iou.compute(sample)
|
||||
return {
|
||||
"mse_per_frame": mse.details["per_frame"],
|
||||
"spatiotemporal_iou_per_frame": st.details["per_frame"],
|
||||
"spatial_iou": float(spatial.score),
|
||||
"weighted_spatial_iou": float(weighted.score),
|
||||
"mse_mean": float(mse.score),
|
||||
"spatiotemporal_iou_mean": float(st.score),
|
||||
}
|
||||
|
||||
def compute_single(
|
||||
self,
|
||||
generated: Any,
|
||||
reference: Any,
|
||||
reference_take2: Any,
|
||||
*,
|
||||
generated_mask: Any | None = None,
|
||||
reference_mask: Any | None = None,
|
||||
reference_take2_mask: Any | None = None,
|
||||
scenario: str | None = None,
|
||||
view: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
prepared = prepare_triplet_inputs(
|
||||
generated,
|
||||
reference,
|
||||
reference_take2,
|
||||
generated_mask=generated_mask,
|
||||
reference_mask=reference_mask,
|
||||
reference_take2_mask=reference_take2_mask,
|
||||
**self._prep_kwargs,
|
||||
)
|
||||
pair_metrics = self._compute_pair_metrics(prepared)
|
||||
variance_pair = prepare_pair_inputs(
|
||||
reference,
|
||||
reference_take2,
|
||||
generated_mask=reference_mask,
|
||||
reference_mask=reference_take2_mask,
|
||||
**self._prep_kwargs,
|
||||
)
|
||||
variance_metrics = self._compute_pair_metrics(variance_pair)
|
||||
details = {
|
||||
**pair_metrics,
|
||||
"pv_mse_per_frame": variance_metrics["mse_per_frame"],
|
||||
"pv_spatiotemporal_iou_per_frame": variance_metrics["spatiotemporal_iou_per_frame"],
|
||||
"pv_spatial_iou": variance_metrics["spatial_iou"],
|
||||
"pv_weighted_spatial_iou": variance_metrics["weighted_spatial_iou"],
|
||||
"pv_mse_mean": variance_metrics["mse_mean"],
|
||||
"pv_spatiotemporal_iou_mean": variance_metrics["spatiotemporal_iou_mean"],
|
||||
}
|
||||
if scenario is not None:
|
||||
details["scenario"] = scenario
|
||||
if view is not None:
|
||||
details["view"] = view
|
||||
return details
|
||||
|
||||
@classmethod
|
||||
def aggregate(cls, results_list: Iterable[MetricResult | Mapping[str, Any]]) -> float:
|
||||
payloads = [cls._extract_payload(result) for result in results_list]
|
||||
if not payloads:
|
||||
raise ValueError("PhysicsIQMetric.aggregate requires at least one result.")
|
||||
|
||||
a_mse = mean([value for payload in payloads for value in payload["mse_per_frame"]])
|
||||
a_st = mean([value for payload in payloads for value in payload["spatiotemporal_iou_per_frame"]])
|
||||
a_s = mean([float(payload["spatial_iou"]) for payload in payloads])
|
||||
a_ws = mean([float(payload["weighted_spatial_iou"]) for payload in payloads])
|
||||
|
||||
v_mse = mean([value for payload in payloads for value in payload["pv_mse_per_frame"]])
|
||||
v_st = mean([value for payload in payloads for value in payload["pv_spatiotemporal_iou_per_frame"]])
|
||||
v_s = mean([float(payload["pv_spatial_iou"]) for payload in payloads])
|
||||
v_ws = mean([float(payload["pv_weighted_spatial_iou"]) for payload in payloads])
|
||||
|
||||
score = 100.0 * ((((a_st / v_st) + (a_s / v_s) + (a_ws / v_ws)) / 3.0) - (a_mse - v_mse))
|
||||
return round(float(np.clip(score, 0.0, 100.0)), 2)
|
||||
|
||||
@classmethod
|
||||
def aggregate_components(cls, results_list: Iterable[MetricResult | Mapping[str, Any]]) -> dict[str, float]:
|
||||
payloads = [cls._extract_payload(result) for result in results_list]
|
||||
return {
|
||||
"physics_iq": cls.aggregate(payloads),
|
||||
"a_mse": mean([value for payload in payloads for value in payload["mse_per_frame"]]),
|
||||
"a_st": mean([value for payload in payloads for value in payload["spatiotemporal_iou_per_frame"]]),
|
||||
"a_s": mean([float(payload["spatial_iou"]) for payload in payloads]),
|
||||
"a_ws": mean([float(payload["weighted_spatial_iou"]) for payload in payloads]),
|
||||
"v_mse": mean([value for payload in payloads for value in payload["pv_mse_per_frame"]]),
|
||||
"v_st": mean([value for payload in payloads for value in payload["pv_spatiotemporal_iou_per_frame"]]),
|
||||
"v_s": mean([float(payload["pv_spatial_iou"]) for payload in payloads]),
|
||||
"v_ws": mean([float(payload["pv_weighted_spatial_iou"]) for payload in payloads]),
|
||||
}
|
||||
|
||||
def _per_video_score(self, details: Mapping[str, Any]) -> float:
|
||||
score = 100.0 * (
|
||||
((mean(details["spatiotemporal_iou_per_frame"]) / mean(details["pv_spatiotemporal_iou_per_frame"])) +
|
||||
(float(details["spatial_iou"]) / float(details["pv_spatial_iou"])) +
|
||||
(float(details["weighted_spatial_iou"]) / float(details["pv_weighted_spatial_iou"]))) / 3.0 -
|
||||
(mean(details["mse_per_frame"]) - mean(details["pv_mse_per_frame"])))
|
||||
return round(float(np.clip(score, 0.0, 100.0)), 2)
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
if "reference" not in sample:
|
||||
raise KeyError("PhysicsIQMetric requires sample['reference'].")
|
||||
|
||||
take2_key = None
|
||||
for candidate in ("reference_take2", "real_take2", "take2"):
|
||||
if candidate in sample:
|
||||
take2_key = candidate
|
||||
break
|
||||
if take2_key is None:
|
||||
raise KeyError("PhysicsIQMetric requires sample['reference_take2'] or an alias.")
|
||||
|
||||
video = sample["video"]
|
||||
reference = sample["reference"]
|
||||
reference_take2 = sample[take2_key]
|
||||
generated_mask = sample.get("video_mask")
|
||||
reference_mask = sample.get("reference_mask")
|
||||
reference_take2_mask = sample.get("reference_take2_mask")
|
||||
|
||||
scenario = sample.get("scenario")
|
||||
view = sample.get("view")
|
||||
|
||||
details = self.compute_single(
|
||||
video,
|
||||
reference,
|
||||
reference_take2,
|
||||
generated_mask=generated_mask,
|
||||
reference_mask=reference_mask,
|
||||
reference_take2_mask=reference_take2_mask,
|
||||
scenario=scenario,
|
||||
view=view,
|
||||
)
|
||||
return MetricResult(name=self.name, score=self._per_video_score(details), details=details)
|
||||
@@ -0,0 +1,25 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.physics_iq.utils import compute_mse, prepare_pair
|
||||
|
||||
|
||||
@register("physics_iq.mse")
|
||||
class PhysicsIQMSEMetric(BaseMetric):
|
||||
name = "physics_iq.mse"
|
||||
requires_reference = True
|
||||
higher_is_better = False
|
||||
|
||||
def __init__(self, **kwargs: Any) -> None:
|
||||
super().__init__()
|
||||
self._kwargs = kwargs
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
prepared = prepare_pair(sample, prep_kwargs=self._kwargs)
|
||||
per_frame = compute_mse(prepared.reference_quarter, prepared.generated_quarter)
|
||||
score = sum(per_frame) / len(per_frame)
|
||||
return MetricResult(name=self.name, score=score, details={"per_frame": per_frame})
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.physics_iq.utils import compute_spatial_iou, prepare_pair
|
||||
|
||||
|
||||
@register("physics_iq.spatial_iou")
|
||||
class SpatialIoUMetric(BaseMetric):
|
||||
name = "physics_iq.spatial_iou"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
|
||||
def __init__(self, **kwargs: Any) -> None:
|
||||
super().__init__()
|
||||
self._kwargs = kwargs
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
prepared = prepare_pair(sample, prep_kwargs=self._kwargs)
|
||||
score = compute_spatial_iou(prepared.reference_masks, prepared.generated_masks)
|
||||
return MetricResult(name=self.name, score=score, details={})
|
||||
@@ -0,0 +1,25 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.physics_iq.utils import compute_spatiotemporal_iou, prepare_pair
|
||||
|
||||
|
||||
@register("physics_iq.spatiotemporal_iou")
|
||||
class SpatiotemporalIoUMetric(BaseMetric):
|
||||
name = "physics_iq.spatiotemporal_iou"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
|
||||
def __init__(self, **kwargs: Any) -> None:
|
||||
super().__init__()
|
||||
self._kwargs = kwargs
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
prepared = prepare_pair(sample, prep_kwargs=self._kwargs)
|
||||
per_frame = compute_spatiotemporal_iou(prepared.reference_masks, prepared.generated_masks)
|
||||
score = sum(per_frame) / len(per_frame)
|
||||
return MetricResult(name=self.name, score=score, details={"per_frame": per_frame})
|
||||
@@ -0,0 +1,420 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
import os
|
||||
import tempfile
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
# Default sampling configuration for the Physics-IQ comparison pipeline.
|
||||
# Source release ships at 30 FPS / 5 seconds — the metric collapses to those
|
||||
# anchors regardless of how the user resampled the input video.
|
||||
DEFAULT_TARGET_FPS = 30
|
||||
DEFAULT_DURATION_SECONDS = 5
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PreparedPhysicsIQPair:
|
||||
generated_quarter: np.ndarray
|
||||
reference_quarter: np.ndarray
|
||||
generated_masks: np.ndarray
|
||||
reference_masks: np.ndarray
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PreparedPhysicsIQTriplet:
|
||||
generated_quarter: np.ndarray
|
||||
reference_quarter: np.ndarray
|
||||
reference_take2_quarter: np.ndarray
|
||||
generated_masks: np.ndarray
|
||||
reference_masks: np.ndarray
|
||||
reference_take2_masks: np.ndarray
|
||||
|
||||
|
||||
def tensor_to_uint8_frames(video: torch.Tensor) -> np.ndarray:
|
||||
arr = video.detach().cpu().float().clamp(0, 1).permute(0, 2, 3, 1).numpy()
|
||||
return np.clip(np.rint(arr * 255.0), 0, 255).astype(np.uint8)
|
||||
|
||||
|
||||
def read_video_frames(
|
||||
source: str | Path,
|
||||
*,
|
||||
start_frame: int = 0,
|
||||
end_frame: int | None = None,
|
||||
) -> np.ndarray:
|
||||
cap = cv2.VideoCapture(str(source))
|
||||
if not cap.isOpened():
|
||||
raise FileNotFoundError(f"Could not open video: {source}")
|
||||
|
||||
frames: list[np.ndarray] = []
|
||||
frame_idx = 0
|
||||
while cap.isOpened():
|
||||
ret, frame = cap.read()
|
||||
if not ret:
|
||||
break
|
||||
if frame_idx >= start_frame and (end_frame is None or frame_idx < end_frame):
|
||||
frames.append(frame)
|
||||
if end_frame is not None and frame_idx >= end_frame:
|
||||
break
|
||||
frame_idx += 1
|
||||
cap.release()
|
||||
|
||||
if not frames:
|
||||
return np.zeros((0, 0, 0, 3), dtype=np.uint8)
|
||||
return np.stack(frames, axis=0)
|
||||
|
||||
|
||||
def as_numpy_video(source: Any) -> tuple[np.ndarray, str]:
|
||||
if isinstance(source, torch.Tensor):
|
||||
if source.ndim != 4:
|
||||
raise ValueError(f"Expected 4D video tensor (T,C,H,W), got shape {tuple(source.shape)}")
|
||||
return tensor_to_uint8_frames(source), "rgb"
|
||||
if isinstance(source, np.ndarray):
|
||||
if source.ndim != 4:
|
||||
raise ValueError(f"Expected 4D ndarray video, got shape {source.shape}")
|
||||
if source.shape[-1] == 3:
|
||||
return source.astype(np.uint8), "rgb"
|
||||
if source.shape[1] == 3:
|
||||
return np.transpose(source, (0, 2, 3, 1)).astype(np.uint8), "rgb"
|
||||
raise ValueError(f"Unsupported ndarray video shape: {source.shape}")
|
||||
if isinstance(source, str | Path):
|
||||
return read_video_frames(source), "bgr"
|
||||
raise TypeError(f"Unsupported Physics-IQ video source type: {type(source)!r}")
|
||||
|
||||
|
||||
def prepare_pair(
|
||||
sample: dict[str, Any],
|
||||
*,
|
||||
prep_kwargs: dict[str, Any] | None = None,
|
||||
) -> PreparedPhysicsIQPair:
|
||||
"""Resolve a sample into a prepared (gen, ref) pair.
|
||||
|
||||
Caches the result on ``sample['_physics_iq_pair']`` so other physics_iq
|
||||
sub-metrics on the same sample reuse it instead of re-decoding.
|
||||
"""
|
||||
prepared = sample.get("_physics_iq_pair")
|
||||
if prepared is not None:
|
||||
return prepared
|
||||
|
||||
if "reference" not in sample:
|
||||
raise KeyError("Physics-IQ pair metrics require sample['reference'].")
|
||||
|
||||
return prepare_pair_inputs(
|
||||
sample["video"],
|
||||
sample["reference"],
|
||||
generated_mask=sample.get("video_mask"),
|
||||
reference_mask=sample.get("reference_mask"),
|
||||
**(prep_kwargs or {}),
|
||||
)
|
||||
|
||||
|
||||
def select_window(frames: np.ndarray, *, target_frames: int, selection: str = "first") -> np.ndarray:
|
||||
if selection != "first":
|
||||
start = max(frames.shape[0] - target_frames, 0)
|
||||
return frames[start:start + target_frames]
|
||||
return frames[:target_frames]
|
||||
|
||||
|
||||
def resize_frames(frames: np.ndarray, target_size: tuple[int, int]) -> np.ndarray:
|
||||
if frames.size == 0:
|
||||
return frames
|
||||
resized = [cv2.resize(frame, target_size) for frame in frames]
|
||||
return np.stack(resized, axis=0)
|
||||
|
||||
|
||||
def rebinarize_masks(mask_frames: np.ndarray) -> np.ndarray:
|
||||
if mask_frames.ndim == 4 and mask_frames.shape[-1] == 3:
|
||||
mask_frames = mask_frames[..., 0]
|
||||
return (mask_frames > 127).astype(np.uint8)
|
||||
|
||||
|
||||
def load_mask_frames(
|
||||
mask_source: Any,
|
||||
*,
|
||||
target_frames: int,
|
||||
target_size: tuple[int, int],
|
||||
) -> np.ndarray:
|
||||
if mask_source is None:
|
||||
raise ValueError("mask_source cannot be None when loading mask frames")
|
||||
mask_frames, _ = as_numpy_video(mask_source)
|
||||
mask_frames = select_window(mask_frames, target_frames=target_frames, selection="first")
|
||||
mask_frames = resize_frames(mask_frames, target_size)
|
||||
return rebinarize_masks(mask_frames)
|
||||
|
||||
|
||||
def roundtrip_mask_frames(mask_frames: np.ndarray, *, fps: int) -> np.ndarray:
|
||||
if mask_frames.size == 0:
|
||||
return mask_frames
|
||||
fd, tmp_path = tempfile.mkstemp(suffix=".mp4")
|
||||
os.close(fd)
|
||||
output_path = Path(tmp_path)
|
||||
output_path.unlink(missing_ok=True)
|
||||
try:
|
||||
writer = cv2.VideoWriter(
|
||||
str(output_path),
|
||||
cv2.VideoWriter_fourcc(*"mp4v"),
|
||||
fps,
|
||||
(mask_frames.shape[2], mask_frames.shape[1]),
|
||||
isColor=False,
|
||||
)
|
||||
for frame in mask_frames:
|
||||
writer.write(frame)
|
||||
writer.release()
|
||||
return read_video_frames(output_path)
|
||||
finally:
|
||||
output_path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def infer_real_mask_path(video_source: Any) -> str | None:
|
||||
if not isinstance(video_source, str | Path):
|
||||
return None
|
||||
|
||||
video_path = Path(video_source)
|
||||
filename = video_path.name
|
||||
if "_testing-videos_" not in filename:
|
||||
return None
|
||||
|
||||
mask_name = filename.replace("_testing-videos_", "_video-masks_")
|
||||
candidates: list[Path] = []
|
||||
fps_dir = video_path.parent.name
|
||||
for ancestor in video_path.parents:
|
||||
if ancestor.name == "split-videos":
|
||||
candidates.extend([
|
||||
ancestor.parent / "video-masks" / "real" / fps_dir / mask_name,
|
||||
ancestor.parent / "video_masks" / "real" / fps_dir / mask_name,
|
||||
])
|
||||
break
|
||||
|
||||
for candidate in candidates:
|
||||
if candidate.exists():
|
||||
return str(candidate)
|
||||
return None
|
||||
|
||||
|
||||
def prepare_grayscale_frame(frame: np.ndarray, *, color_order: str) -> np.ndarray:
|
||||
if frame.ndim == 2:
|
||||
gray = frame
|
||||
elif color_order == "bgr":
|
||||
gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
|
||||
elif color_order == "rgb":
|
||||
gray = cv2.cvtColor(frame, cv2.COLOR_RGB2GRAY)
|
||||
else:
|
||||
raise ValueError(f"Unsupported color order: {color_order}")
|
||||
return cv2.GaussianBlur(gray, (5, 5), 0)
|
||||
|
||||
|
||||
def generate_motion_mask(
|
||||
video_frames: np.ndarray,
|
||||
*,
|
||||
threshold: int = 10,
|
||||
alpha: float = 0.3,
|
||||
color_order: str = "rgb",
|
||||
) -> np.ndarray:
|
||||
if video_frames.size == 0:
|
||||
return np.zeros((0, 0, 0), dtype=np.uint8)
|
||||
|
||||
first_gray = prepare_grayscale_frame(video_frames[0], color_order=color_order)
|
||||
avg_frame = first_gray.astype("float")
|
||||
masks = [np.zeros_like(first_gray, dtype=np.uint8)]
|
||||
kernel = np.ones((5, 5), np.uint8)
|
||||
|
||||
for frame in video_frames[1:]:
|
||||
gray_frame = prepare_grayscale_frame(frame, color_order=color_order)
|
||||
cv2.accumulateWeighted(gray_frame, avg_frame, alpha)
|
||||
avg_gray_frame = cv2.convertScaleAbs(avg_frame)
|
||||
frame_diff = cv2.absdiff(gray_frame, avg_gray_frame)
|
||||
_, binary_frame = cv2.threshold(frame_diff, threshold, 255, cv2.THRESH_BINARY)
|
||||
binary_frame = cv2.morphologyEx(binary_frame, cv2.MORPH_OPEN, kernel)
|
||||
binary_frame = cv2.morphologyEx(binary_frame, cv2.MORPH_CLOSE, kernel)
|
||||
masks.append(binary_frame)
|
||||
return np.stack(masks, axis=0)
|
||||
|
||||
|
||||
def compute_iou(mask1: np.ndarray, mask2: np.ndarray) -> float:
|
||||
intersection = np.logical_and(mask1, mask2).sum()
|
||||
union = np.logical_or(mask1, mask2).sum()
|
||||
if union == 0:
|
||||
return 1.0
|
||||
return float(intersection / union)
|
||||
|
||||
|
||||
def compute_mse(video1_frames: np.ndarray, video2_frames: np.ndarray) -> list[float]:
|
||||
if len(video1_frames) != len(video2_frames):
|
||||
raise ValueError("Videos must have the same number of frames.")
|
||||
frame_mses: list[float] = []
|
||||
for frame1, frame2 in zip(video1_frames, video2_frames, strict=False):
|
||||
if frame1.shape != frame2.shape:
|
||||
raise ValueError("Frames must have the same dimensions.")
|
||||
mse = np.mean((frame1.astype(np.float32) - frame2.astype(np.float32))**2)
|
||||
frame_mses.append(round(float(mse), 4))
|
||||
return frame_mses
|
||||
|
||||
|
||||
def compute_spatiotemporal_iou(mask1_frames: np.ndarray, mask2_frames: np.ndarray) -> list[float]:
|
||||
values: list[float] = []
|
||||
for mask1, mask2 in zip(mask1_frames, mask2_frames, strict=False):
|
||||
values.append(round(compute_iou(mask1, mask2), 4))
|
||||
return values
|
||||
|
||||
|
||||
def compute_spatial_iou(mask1_frames: np.ndarray, mask2_frames: np.ndarray) -> float:
|
||||
spatial_mask1 = (np.max(mask1_frames, axis=0) > 0).astype(np.uint8) * 255
|
||||
spatial_mask2 = (np.max(mask2_frames, axis=0) > 0).astype(np.uint8) * 255
|
||||
return compute_iou(spatial_mask1, spatial_mask2)
|
||||
|
||||
|
||||
def compute_weighted_spatial_iou(mask1_frames: np.ndarray, mask2_frames: np.ndarray) -> float:
|
||||
weighted_spatial_1 = np.sum(mask1_frames, axis=0, dtype=np.uint16) / len(mask1_frames)
|
||||
weighted_spatial_2 = np.sum(mask2_frames, axis=0, dtype=np.uint16) / len(mask2_frames)
|
||||
intersection = np.minimum(weighted_spatial_1, weighted_spatial_2)
|
||||
union = np.maximum(weighted_spatial_1, weighted_spatial_2)
|
||||
valid_pixels = union > 0
|
||||
if np.sum(valid_pixels) == 0:
|
||||
return 1.0
|
||||
return float(np.sum(intersection[valid_pixels]) / np.sum(union[valid_pixels]))
|
||||
|
||||
|
||||
def mean(values: list[float] | tuple[float, ...]) -> float:
|
||||
arr = np.asarray(list(values), dtype=np.float64)
|
||||
if arr.size == 0:
|
||||
raise ValueError("Cannot aggregate empty Physics-IQ values.")
|
||||
return float(arr.mean())
|
||||
|
||||
|
||||
def quarter_resolution_target(reference_frames: np.ndarray) -> tuple[int, int]:
|
||||
return (
|
||||
max(reference_frames[0].shape[1] // 4, 1),
|
||||
max(reference_frames[0].shape[0] // 4, 1),
|
||||
)
|
||||
|
||||
|
||||
def prepare_pair_inputs(
|
||||
generated: Any,
|
||||
reference: Any,
|
||||
*,
|
||||
generated_mask: Any | None = None,
|
||||
reference_mask: Any | None = None,
|
||||
target_fps: int = DEFAULT_TARGET_FPS,
|
||||
duration_seconds: int = DEFAULT_DURATION_SECONDS,
|
||||
video_time_selection: str = "first",
|
||||
threshold: int = 10,
|
||||
alpha: float = 0.3,
|
||||
roundtrip_generated_masks: bool = True,
|
||||
) -> PreparedPhysicsIQPair:
|
||||
generated_frames, generated_color = as_numpy_video(generated)
|
||||
reference_frames, reference_color = as_numpy_video(reference)
|
||||
consider_frames = target_fps * duration_seconds
|
||||
|
||||
generated_frames = select_window(
|
||||
generated_frames,
|
||||
target_frames=consider_frames,
|
||||
selection=video_time_selection,
|
||||
)
|
||||
reference_frames = reference_frames[:consider_frames]
|
||||
if not len(generated_frames) or not len(reference_frames):
|
||||
raise ValueError("Physics-IQ pair metrics require non-empty generated and reference videos.")
|
||||
|
||||
target_size = quarter_resolution_target(reference_frames)
|
||||
reference_mask = reference_mask or infer_real_mask_path(reference)
|
||||
|
||||
generated_quarter = resize_frames(generated_frames, target_size).astype(np.float32) / 255.0
|
||||
reference_quarter = resize_frames(reference_frames, target_size).astype(np.float32) / 255.0
|
||||
|
||||
generated_masks = (load_mask_frames(generated_mask, target_frames=consider_frames, target_size=target_size)
|
||||
if generated_mask is not None else rebinarize_masks(
|
||||
resize_frames(
|
||||
roundtrip_mask_frames(
|
||||
generate_motion_mask(
|
||||
generated_frames,
|
||||
threshold=threshold,
|
||||
alpha=alpha,
|
||||
color_order=generated_color,
|
||||
),
|
||||
fps=target_fps,
|
||||
) if roundtrip_generated_masks else generate_motion_mask(
|
||||
generated_frames,
|
||||
threshold=threshold,
|
||||
alpha=alpha,
|
||||
color_order=generated_color,
|
||||
),
|
||||
target_size,
|
||||
)))
|
||||
reference_masks = (load_mask_frames(reference_mask, target_frames=consider_frames, target_size=target_size)
|
||||
if reference_mask is not None else rebinarize_masks(
|
||||
resize_frames(
|
||||
generate_motion_mask(
|
||||
reference_frames,
|
||||
threshold=threshold,
|
||||
alpha=alpha,
|
||||
color_order=reference_color,
|
||||
),
|
||||
target_size,
|
||||
)))
|
||||
return PreparedPhysicsIQPair(
|
||||
generated_quarter=generated_quarter,
|
||||
reference_quarter=reference_quarter,
|
||||
generated_masks=generated_masks,
|
||||
reference_masks=reference_masks,
|
||||
)
|
||||
|
||||
|
||||
def prepare_triplet_inputs(
|
||||
generated: Any,
|
||||
reference: Any,
|
||||
reference_take2: Any,
|
||||
*,
|
||||
generated_mask: Any | None = None,
|
||||
reference_mask: Any | None = None,
|
||||
reference_take2_mask: Any | None = None,
|
||||
target_fps: int = DEFAULT_TARGET_FPS,
|
||||
duration_seconds: int = DEFAULT_DURATION_SECONDS,
|
||||
video_time_selection: str = "first",
|
||||
threshold: int = 10,
|
||||
alpha: float = 0.3,
|
||||
roundtrip_generated_masks: bool = True,
|
||||
) -> PreparedPhysicsIQTriplet:
|
||||
pair = prepare_pair_inputs(
|
||||
generated,
|
||||
reference,
|
||||
generated_mask=generated_mask,
|
||||
reference_mask=reference_mask,
|
||||
target_fps=target_fps,
|
||||
duration_seconds=duration_seconds,
|
||||
video_time_selection=video_time_selection,
|
||||
threshold=threshold,
|
||||
alpha=alpha,
|
||||
roundtrip_generated_masks=roundtrip_generated_masks,
|
||||
)
|
||||
reference_take2_frames, reference_take2_color = as_numpy_video(reference_take2)
|
||||
consider_frames = target_fps * duration_seconds
|
||||
reference_take2_frames = reference_take2_frames[:consider_frames]
|
||||
if not len(reference_take2_frames):
|
||||
raise ValueError("Physics-IQ requires a non-empty take-2 reference video.")
|
||||
|
||||
target_size = (pair.reference_quarter.shape[2], pair.reference_quarter.shape[1])
|
||||
reference_take2_mask = reference_take2_mask or infer_real_mask_path(reference_take2)
|
||||
reference_take2_quarter = resize_frames(reference_take2_frames, target_size).astype(np.float32) / 255.0
|
||||
reference_take2_masks = (load_mask_frames(
|
||||
reference_take2_mask, target_frames=consider_frames, target_size=target_size)
|
||||
if reference_take2_mask is not None else rebinarize_masks(
|
||||
resize_frames(
|
||||
generate_motion_mask(
|
||||
reference_take2_frames,
|
||||
threshold=threshold,
|
||||
alpha=alpha,
|
||||
color_order=reference_take2_color,
|
||||
),
|
||||
target_size,
|
||||
)))
|
||||
return PreparedPhysicsIQTriplet(
|
||||
generated_quarter=pair.generated_quarter,
|
||||
reference_quarter=pair.reference_quarter,
|
||||
reference_take2_quarter=reference_take2_quarter,
|
||||
generated_masks=pair.generated_masks,
|
||||
reference_masks=pair.reference_masks,
|
||||
reference_take2_masks=reference_take2_masks,
|
||||
)
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.physics_iq.utils import compute_weighted_spatial_iou, prepare_pair
|
||||
|
||||
|
||||
@register("physics_iq.weighted_spatial_iou")
|
||||
class WeightedSpatialIoUMetric(BaseMetric):
|
||||
name = "physics_iq.weighted_spatial_iou"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
|
||||
def __init__(self, **kwargs: Any) -> None:
|
||||
super().__init__()
|
||||
self._kwargs = kwargs
|
||||
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
prepared = prepare_pair(sample, prep_kwargs=self._kwargs)
|
||||
score = compute_weighted_spatial_iou(prepared.reference_masks, prepared.generated_masks)
|
||||
return MetricResult(name=self.name, score=score, details={})
|
||||
@@ -0,0 +1,121 @@
|
||||
"""VBench metrics. Bootstraps upstream submodule on sys.path and
|
||||
installs runtime compat shims for modern torch/transformers/numpy/timm.
|
||||
|
||||
The upstream vbench source lives as a git submodule at
|
||||
``fastvideo/third_party/eval/vbench`` (pinned to a specific
|
||||
Vchitect/VBench SHA). We do not pip-install it — we only need its
|
||||
Python modules importable. Its runtime deps (clip, transformers, etc.)
|
||||
are already in FastVideo's main env.
|
||||
|
||||
Compat with modern dependency versions is achieved at import time, in
|
||||
this file, instead of via on-disk patches to upstream files. Each shim
|
||||
below corresponds to a specific drift between vbench's pinned-2023 deps
|
||||
and FastVideo's current pins. Adding a new shim is preferable to editing
|
||||
the submodule.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
# fastvideo/eval/metrics/vbench/__init__.py → ../../../third_party/eval/vbench
|
||||
# parents[3] is the ``fastvideo/`` package root.
|
||||
_UPSTREAM = Path(__file__).resolve().parents[3] / "third_party" / "eval" / "vbench"
|
||||
if _UPSTREAM.is_dir() and str(_UPSTREAM) not in sys.path:
|
||||
sys.path.insert(0, str(_UPSTREAM))
|
||||
|
||||
|
||||
def _install_compat_shims() -> None:
|
||||
"""Apply attribute-level shims that make vbench imports resolve.
|
||||
|
||||
Idempotent and side-effect-free if the targeted modules are already
|
||||
correct (e.g. on older transformers/numpy).
|
||||
"""
|
||||
# transformers: apply_chunking_to_forward & friends moved from
|
||||
# ``transformers.modeling_utils`` to ``transformers.pytorch_utils``
|
||||
# (transformers ~= 4.30+). Mirror them back so vbench's legacy
|
||||
# ``from transformers.modeling_utils import (...)`` keeps resolving.
|
||||
try:
|
||||
import transformers.modeling_utils as _mu
|
||||
import transformers.pytorch_utils as _pu
|
||||
for _name in ("apply_chunking_to_forward", "find_pruneable_heads_and_indices", "prune_linear_layer"):
|
||||
if not hasattr(_mu, _name) and hasattr(_pu, _name):
|
||||
setattr(_mu, _name, getattr(_pu, _name))
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
# numpy.lib.function_base was removed entirely in numpy>=2; vbench's
|
||||
# umt/kinetics still does ``from numpy.lib.function_base import disp``
|
||||
# but never calls disp. Install a stub submodule with a no-op ``disp``
|
||||
# so the legacy import line resolves.
|
||||
try:
|
||||
import types
|
||||
import numpy.lib as _nl
|
||||
if not hasattr(_nl, "function_base"):
|
||||
_stub = types.ModuleType("numpy.lib.function_base")
|
||||
_stub.disp = lambda *a, **k: None # type: ignore[attr-defined]
|
||||
sys.modules["numpy.lib.function_base"] = _stub
|
||||
_nl.function_base = _stub # type: ignore[attr-defined]
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
|
||||
def _install_modeling_finetune_hook() -> None:
|
||||
"""Wrap vbench's ``vit_large_patch16_224`` to drop the ``cache_dir``
|
||||
kwarg that newer timm passes to model factory functions but the
|
||||
upstream factory doesn't accept. Installed as a meta-path finder so
|
||||
we patch the attribute on the actual module object after it loads,
|
||||
without eagerly importing torch+timm at fastvideo.eval import time.
|
||||
"""
|
||||
import importlib.abc
|
||||
|
||||
_target = "vbench.third_party.umt.models.modeling_finetune"
|
||||
|
||||
class _Loader(importlib.abc.Loader):
|
||||
|
||||
def __init__(self, real_loader: Any) -> None:
|
||||
self._real = real_loader
|
||||
|
||||
def create_module(self, spec):
|
||||
return None
|
||||
|
||||
def exec_module(self, module):
|
||||
self._real.exec_module(module)
|
||||
orig = getattr(module, "vit_large_patch16_224", None)
|
||||
if orig is None or getattr(orig, "_fastvideo_patched", False):
|
||||
return
|
||||
|
||||
def patched(pretrained=False, **kwargs):
|
||||
kwargs.pop("cache_dir", None)
|
||||
return orig(pretrained=pretrained, **kwargs)
|
||||
|
||||
patched._fastvideo_patched = True # type: ignore[attr-defined]
|
||||
module.vit_large_patch16_224 = patched
|
||||
|
||||
class _Finder(importlib.abc.MetaPathFinder):
|
||||
_reentrant = False
|
||||
|
||||
def find_spec(self, fullname, path, target=None):
|
||||
if fullname != _target or self._reentrant:
|
||||
return None
|
||||
self._reentrant = True
|
||||
try:
|
||||
for finder in sys.meta_path:
|
||||
if finder is self or not hasattr(finder, "find_spec"):
|
||||
continue
|
||||
spec = finder.find_spec(fullname, path, target)
|
||||
if spec is not None and spec.loader is not None:
|
||||
spec.loader = _Loader(spec.loader)
|
||||
return spec
|
||||
return None
|
||||
finally:
|
||||
self._reentrant = False
|
||||
|
||||
if not any(isinstance(f, _Finder) for f in sys.meta_path):
|
||||
sys.meta_path.insert(0, _Finder())
|
||||
|
||||
|
||||
_install_compat_shims()
|
||||
_install_modeling_finetune_hook()
|
||||
@@ -0,0 +1,158 @@
|
||||
"""Shared GRiT model loading and detection utilities for VBench metrics.
|
||||
|
||||
All 4 GRiT-based metrics (object_class, multiple_objects, color,
|
||||
spatial_relationship) use the same model and detection API.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
|
||||
def _patch_detectron2_registries() -> None:
|
||||
"""Make detectron2's fvcore registries skip duplicate names instead of raising.
|
||||
|
||||
Both pip-installed ``vbench`` and wm-eval vendor the same GRiT / CenterNet2
|
||||
source with 20+ ``@REGISTRY.register()`` calls. When both are imported in the
|
||||
same process (e.g. the parity test), detectron2's global registries see two
|
||||
different Python classes with the same name and raise ``AssertionError``.
|
||||
|
||||
This one-time patch makes ``_do_register`` silently skip if the name is
|
||||
already registered, which is safe because the classes are identical.
|
||||
"""
|
||||
# fvcore Registry (META_ARCH, ROI_HEADS, BACKBONE, PROPOSAL_GENERATOR, ...)
|
||||
try:
|
||||
from fvcore.common.registry import Registry
|
||||
orig = Registry._do_register
|
||||
if not getattr(orig, "_patched_idempotent", False):
|
||||
|
||||
def _safe_do_register(self, name, obj):
|
||||
if name in self._obj_map:
|
||||
return
|
||||
orig(self, name, obj)
|
||||
|
||||
_safe_do_register._patched_idempotent = True # type: ignore[attr-defined]
|
||||
Registry._do_register = _safe_do_register
|
||||
except ImportError:
|
||||
pass
|
||||
|
||||
# detectron2 DatasetCatalog (object365_train, vg_train, etc.)
|
||||
try:
|
||||
from detectron2.data import DatasetCatalog
|
||||
orig_ds = DatasetCatalog.register
|
||||
if not getattr(orig_ds, "_patched_idempotent", False):
|
||||
|
||||
def _safe_ds_register(name, func):
|
||||
if name in DatasetCatalog:
|
||||
return
|
||||
orig_ds(name, func)
|
||||
|
||||
_safe_ds_register._patched_idempotent = True # type: ignore[attr-defined]
|
||||
DatasetCatalog.register = _safe_ds_register
|
||||
except (ImportError, AttributeError):
|
||||
pass
|
||||
|
||||
|
||||
_patch_detectron2_registries()
|
||||
|
||||
|
||||
def load_grit_model(device: str | torch.device, task: str = "DenseCap"):
|
||||
"""Load the GRiT DenseCaptioning model.
|
||||
|
||||
Parameters
|
||||
----------
|
||||
task : "DenseCap" | "ObjectDet"
|
||||
VBench uses "ObjectDet" for object_class / multiple_objects /
|
||||
spatial_relationship and "DenseCap" for color (which needs the
|
||||
actual caption text). The two heads return predictions in
|
||||
different formats:
|
||||
- DenseCap → ``[(caption, bbox, [class_label]), ...]``
|
||||
- ObjectDet → ``[(class_label, bbox, [class_label]), ...]``
|
||||
spatial_relationship matches on ``pred[0]`` (the first field),
|
||||
so it must run in ObjectDet mode to compare against class names.
|
||||
"""
|
||||
from vbench.third_party.grit_model import DenseCaptioning
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
|
||||
ckpt = ensure_checkpoint(
|
||||
"grit_b_densecap_objectdet.pth",
|
||||
source="OpenGVLab/VBench_Used_Models",
|
||||
filename="grit_b_densecap_objectdet.pth",
|
||||
)
|
||||
# GRiT internals call .type on device, so coerce to torch.device
|
||||
if isinstance(device, str):
|
||||
device = torch.device(device)
|
||||
model = DenseCaptioning(device)
|
||||
if task == "ObjectDet":
|
||||
model.initialize_model_det(ckpt)
|
||||
else:
|
||||
model.initialize_model(ckpt)
|
||||
return model
|
||||
|
||||
|
||||
def detect_frames(model, frames_np: list[np.ndarray]) -> list:
|
||||
"""Run GRiT detection on a list of (H, W, C) uint8 numpy frames.
|
||||
|
||||
Returns per-frame predictions in the format used by VBench metrics.
|
||||
Each frame's predictions is a list of (description, bbox, object_types).
|
||||
"""
|
||||
predictions = []
|
||||
with torch.no_grad():
|
||||
for frame in frames_np:
|
||||
ret = model.run_caption_tensor(frame)
|
||||
predictions.append(ret[0] if len(ret[0]) > 0 else [])
|
||||
return predictions
|
||||
|
||||
|
||||
def _vbench_middle_indices(vlen: int, num_frames: int) -> list[int]:
|
||||
"""Replicate VBench's get_frame_indices(sample="middle"): split [0, vlen)
|
||||
into num_frames equal intervals and pick the midpoint of each.
|
||||
|
||||
Without this, wm-eval's torch.linspace-based sampler picks different
|
||||
indices than VBench, producing different GRiT predictions and scores
|
||||
on long videos. See vbench/utils.py:get_frame_indices.
|
||||
"""
|
||||
acc = min(num_frames, vlen)
|
||||
intervals = np.linspace(0, vlen, acc + 1).astype(int)
|
||||
indices = [(intervals[i] + intervals[i + 1] - 1) // 2 for i in range(acc)]
|
||||
if len(indices) < num_frames:
|
||||
indices = indices + [indices[-1]] * (num_frames - len(indices))
|
||||
return indices
|
||||
|
||||
|
||||
def prepare_frames(video_tensor: torch.Tensor, n_frames: int = 16, max_short_side: int = 768) -> list[np.ndarray]:
|
||||
"""Convert (T, C, H, W) float [0,1] tensor to list of (H, W, C) numpy frames
|
||||
in VBench's exact format: float32 [0, 255] HWC.
|
||||
|
||||
VBench's load_video casts the decord uint8 buffer with ``torch.Tensor(...)``,
|
||||
yielding **float32 with values in [0, 255]**, then runs torchvision
|
||||
``Resize`` (which preserves float dtype and produces fractional bilinear
|
||||
outputs), then ``.permute(0,2,3,1).numpy()`` for GRiT. We must replicate
|
||||
this exactly — round-tripping through uint8 truncates the fractional
|
||||
bilinear outputs and shifts GRiT detection counts (e.g. 1 vs 6 persons
|
||||
per frame), which then breaks spatial_relationship/multiple_objects.
|
||||
|
||||
Sampling matches VBench's ``get_frame_indices(sample="middle")``.
|
||||
"""
|
||||
from torchvision import transforms
|
||||
|
||||
T = video_tensor.shape[0]
|
||||
indices = _vbench_middle_indices(T, n_frames)
|
||||
# Recover the original uint8 values from the float [0,1] loader
|
||||
# (round, not truncate), then re-cast to float32 [0,255] like VBench's
|
||||
# ``torch.Tensor(decord_uint8_array)`` path.
|
||||
frames_uint8 = (video_tensor[indices] * 255).round().clamp(0, 255).to(torch.uint8)
|
||||
frames_f = frames_uint8.float()
|
||||
|
||||
h, w = frames_f.shape[-2], frames_f.shape[-1]
|
||||
if min(h, w) > max_short_side:
|
||||
scale = 720.0 / min(h, w)
|
||||
new_h, new_w = int(scale * h), int(scale * w)
|
||||
# VBench (object_class.py:55) uses transforms.Resize without
|
||||
# antialias kwarg → torchvision default (BILINEAR, no antialias).
|
||||
# Float input ⇒ fractional bilinear outputs are preserved.
|
||||
frames_f = transforms.Resize(size=(new_h, new_w))(frames_f)
|
||||
|
||||
frames_np = frames_f.permute(0, 2, 3, 1).cpu().numpy().astype(np.float32)
|
||||
return list(frames_np)
|
||||
@@ -0,0 +1,31 @@
|
||||
"""Shared utilities for VBench metrics."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
|
||||
|
||||
def consistency_score(features: torch.Tensor) -> float:
|
||||
"""VBench-style temporal consistency from (T, D) L2-normalized features.
|
||||
|
||||
For each frame t > 0, computes:
|
||||
sim = (cos(f[t], f[t-1]) + cos(f[t], f[0])) / 2, clamped >= 0
|
||||
|
||||
Returns the mean similarity across all t > 0.
|
||||
"""
|
||||
if features.shape[0] <= 1:
|
||||
return 1.0
|
||||
|
||||
first = features[0:1] # (1, D)
|
||||
total_sim = 0.0
|
||||
count = 0
|
||||
for t in range(1, features.shape[0]):
|
||||
curr = features[t:t + 1]
|
||||
prev = features[t - 1:t]
|
||||
sim_prev = max(0.0, F.cosine_similarity(prev, curr).item())
|
||||
sim_first = max(0.0, F.cosine_similarity(first, curr).item())
|
||||
total_sim += (sim_prev + sim_first) / 2
|
||||
count += 1
|
||||
|
||||
return total_sim / count
|
||||
@@ -0,0 +1,98 @@
|
||||
"""VBench Aesthetic Quality — CLIP ViT-L/14 + LAION aesthetic predictor.
|
||||
|
||||
Encodes frames through CLIP, passes L2-normalized features through a
|
||||
linear aesthetic head (768 → 1), and averages scores / 10.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
from torchvision.transforms.functional import resize, center_crop, normalize
|
||||
from torchvision.transforms import InterpolationMode
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
_CLIP_MEAN = [0.48145466, 0.4578275, 0.40821073]
|
||||
_CLIP_STD = [0.26862954, 0.26130258, 0.27577711]
|
||||
|
||||
_AESTHETIC_URL = "https://raw.githubusercontent.com/LAION-AI/aesthetic-predictor/main/sa_0_4_vit_l_14_linear.pth"
|
||||
|
||||
|
||||
def _clip_transform(frames: torch.Tensor) -> torch.Tensor:
|
||||
# antialias=False matches VBench's clip_transform (vbench/utils.py:33)
|
||||
frames = resize(frames, 224, interpolation=InterpolationMode.BICUBIC, antialias=False)
|
||||
frames = center_crop(frames, 224)
|
||||
frames = normalize(frames, mean=_CLIP_MEAN, std=_CLIP_STD)
|
||||
return frames
|
||||
|
||||
|
||||
@register("vbench.aesthetic_quality")
|
||||
class AestheticQualityMetric(BaseMetric):
|
||||
|
||||
name = "vbench.aesthetic_quality"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["clip"]
|
||||
backbone = "clip_vit_l14"
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._clip_model: Any = None
|
||||
self._aesthetic_head: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._clip_model is not None:
|
||||
self._clip_model = self._clip_model.to(self.device)
|
||||
if self._aesthetic_head is not None:
|
||||
self._aesthetic_head = self._aesthetic_head.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._clip_model is not None:
|
||||
return
|
||||
|
||||
import clip
|
||||
from fastvideo.eval.models import ensure_checkpoint, get_cache_dir
|
||||
self._clip_model, _ = clip.load(
|
||||
"ViT-L/14",
|
||||
device=self.device,
|
||||
download_root=str(get_cache_dir() / "clip"),
|
||||
)
|
||||
self._clip_model.eval()
|
||||
|
||||
# Load LAION aesthetic head
|
||||
ckpt_path = ensure_checkpoint(
|
||||
"sa_0_4_vit_l_14_linear.pth",
|
||||
source=_AESTHETIC_URL,
|
||||
)
|
||||
self._aesthetic_head = nn.Linear(768, 1)
|
||||
self._aesthetic_head.load_state_dict(torch.load(ckpt_path, map_location="cpu", weights_only=True))
|
||||
self._aesthetic_head.to(self.device)
|
||||
self._aesthetic_head.eval()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
frames = _clip_transform(video.to(self.device))
|
||||
|
||||
chunk = self._chunk_size or 32
|
||||
scores_list = []
|
||||
for i in range(0, frames.shape[0], chunk):
|
||||
feats = self._clip_model.encode_image(frames[i:i + chunk]).float()
|
||||
feats = F.normalize(feats, dim=-1, p=2)
|
||||
scores_list.append(self._aesthetic_head(feats).squeeze(-1))
|
||||
|
||||
all_scores = torch.cat(scores_list, dim=0) / 10.0 # (T,)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(all_scores.mean().item()),
|
||||
details={"per_frame": all_scores.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,94 @@
|
||||
"""VBench Appearance Style — CLIP ViT-B/32 text-image alignment.
|
||||
|
||||
Per-frame cosine similarity between CLIP image features and a text
|
||||
prompt describing the expected style. Requires ``sample["text_prompt"]``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torchvision.transforms.functional import resize, center_crop, normalize
|
||||
from torchvision.transforms import InterpolationMode
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
_CLIP_MEAN = [0.48145466, 0.4578275, 0.40821073]
|
||||
_CLIP_STD = [0.26862954, 0.26130258, 0.27577711]
|
||||
|
||||
|
||||
def _clip_transform(frames: torch.Tensor) -> torch.Tensor:
|
||||
# antialias=False matches VBench's clip_transform (vbench/utils.py:33)
|
||||
frames = resize(frames, 224, interpolation=InterpolationMode.BICUBIC, antialias=False)
|
||||
frames = center_crop(frames, 224)
|
||||
frames = normalize(frames, mean=_CLIP_MEAN, std=_CLIP_STD)
|
||||
return frames
|
||||
|
||||
|
||||
@register("vbench.appearance_style")
|
||||
class AppearanceStyleMetric(BaseMetric):
|
||||
|
||||
name = "vbench.appearance_style"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["clip"]
|
||||
backbone = "clip_vit_b32"
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
import clip
|
||||
from fastvideo.eval.models import get_cache_dir
|
||||
self._model, _ = clip.load(
|
||||
"ViT-B/32",
|
||||
device=self.device,
|
||||
download_root=str(get_cache_dir() / "clip"),
|
||||
)
|
||||
self._model.eval()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
import clip
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
text_prompt = sample.get("text_prompt")
|
||||
if text_prompt is None:
|
||||
return self._skip(sample, "missing text_prompt")
|
||||
|
||||
frames = _clip_transform(video.to(self.device))
|
||||
|
||||
chunk = self._chunk_size or 64
|
||||
img_feats = []
|
||||
for i in range(0, frames.shape[0], chunk):
|
||||
f = self._model.encode_image(frames[i:i + chunk]).float()
|
||||
f = F.normalize(f, dim=-1, p=2)
|
||||
img_feats.append(f)
|
||||
img_feats = torch.cat(img_feats, dim=0) # (T, D)
|
||||
|
||||
# truncate=True: CLIP context length is 77 tokens; long prompts
|
||||
# truncate instead of raising. Matches CLIP's documented convention.
|
||||
text_tokens = clip.tokenize([text_prompt], truncate=True).to(self.device)
|
||||
text_feat = self._model.encode_text(text_tokens).float()
|
||||
text_feat = F.normalize(text_feat, dim=-1, p=2) # (1, D)
|
||||
|
||||
sims = (img_feats @ text_feat.T).squeeze(-1) # (T,)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(sims.mean().item()),
|
||||
details={"per_frame": sims.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,83 @@
|
||||
"""VBench Background Consistency — CLIP ViT-B/32 temporal feature similarity.
|
||||
|
||||
Measures background stability via cosine similarity of CLIP features
|
||||
between consecutive frames and the first frame.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torchvision.transforms.functional import resize, center_crop, normalize
|
||||
from torchvision.transforms import InterpolationMode
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.vbench._utils import consistency_score
|
||||
|
||||
_CLIP_MEAN = [0.48145466, 0.4578275, 0.40821073]
|
||||
_CLIP_STD = [0.26862954, 0.26130258, 0.27577711]
|
||||
|
||||
|
||||
def _clip_transform(frames: torch.Tensor) -> torch.Tensor:
|
||||
"""Apply CLIP preprocessing to (N, C, H, W) float [0,1] tensors."""
|
||||
# antialias=False matches VBench's clip_transform (vbench/utils.py:33)
|
||||
frames = resize(frames, 224, interpolation=InterpolationMode.BICUBIC, antialias=False)
|
||||
frames = center_crop(frames, 224)
|
||||
frames = normalize(frames, mean=_CLIP_MEAN, std=_CLIP_STD)
|
||||
return frames
|
||||
|
||||
|
||||
@register("vbench.background_consistency")
|
||||
class BackgroundConsistencyMetric(BaseMetric):
|
||||
|
||||
name = "vbench.background_consistency"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["clip"]
|
||||
backbone = "clip_vit_b32"
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
import clip
|
||||
from fastvideo.eval.models import get_cache_dir
|
||||
model, _ = clip.load(
|
||||
"ViT-B/32",
|
||||
device=self.device,
|
||||
download_root=str(get_cache_dir() / "clip"),
|
||||
)
|
||||
model.eval()
|
||||
self._model = model
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
frames = _clip_transform(video.to(self.device))
|
||||
|
||||
chunk = self._chunk_size or 64
|
||||
feats = []
|
||||
for i in range(0, frames.shape[0], chunk):
|
||||
f = self._model.encode_image(frames[i:i + chunk]).float()
|
||||
f = F.normalize(f, dim=-1, p=2)
|
||||
feats.append(f)
|
||||
all_feats = torch.cat(feats, dim=0) # (T, D)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=consistency_score(all_feats),
|
||||
details={},
|
||||
)
|
||||
@@ -0,0 +1,106 @@
|
||||
"""VBench Color — GRiT dense captioning for color accuracy.
|
||||
|
||||
Detects the target object via GRiT and checks if the expected color
|
||||
keyword appears in the object's caption. Score = frames_with_correct_color
|
||||
/ frames_with_object_detected.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
_COLOR_KEYWORDS = [
|
||||
"white",
|
||||
"red",
|
||||
"pink",
|
||||
"blue",
|
||||
"silver",
|
||||
"purple",
|
||||
"orange",
|
||||
"green",
|
||||
"gray",
|
||||
"yellow",
|
||||
"black",
|
||||
"grey",
|
||||
]
|
||||
|
||||
|
||||
@register("vbench.color")
|
||||
class ColorMetric(BaseMetric):
|
||||
|
||||
name = "vbench.color"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["detectron2"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import load_grit_model
|
||||
self._model = load_grit_model(self.device)
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import prepare_frames
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
aux = sample.get("auxiliary_info") or {}
|
||||
if "color" not in aux:
|
||||
return self._skip(sample, "missing 'color' in auxiliary_info")
|
||||
|
||||
prompt = sample.get("text_prompt") or ""
|
||||
color_key = aux["color"]
|
||||
# Parse object name: remove "a ", "an ", and the color word
|
||||
object_key = prompt.replace("a ", "").replace("an ", "").replace(color_key, "").strip()
|
||||
|
||||
frames_np = prepare_frames(video)
|
||||
|
||||
preds = []
|
||||
for frame in frames_np:
|
||||
ret = self._model.run_caption_tensor(frame)
|
||||
cur_pred = []
|
||||
if len(ret[0]) < 1:
|
||||
cur_pred.append(["", ""])
|
||||
else:
|
||||
for cap_det in ret[0]:
|
||||
cur_pred.append([cap_det[0], cap_det[2][0]])
|
||||
preds.append(cur_pred)
|
||||
|
||||
# Score: matching VBench's check_generate logic
|
||||
cur_object = 0
|
||||
cur_object_color = 0
|
||||
for frame_pred in preds:
|
||||
object_flag = False
|
||||
color_flag = False
|
||||
for pred in frame_pred:
|
||||
if object_key == pred[1]:
|
||||
for cq in _COLOR_KEYWORDS:
|
||||
if cq in pred[0]:
|
||||
object_flag = True
|
||||
if color_key in pred[0]:
|
||||
color_flag = True
|
||||
if color_flag:
|
||||
cur_object_color += 1
|
||||
if object_flag:
|
||||
cur_object += 1
|
||||
|
||||
score = cur_object_color / cur_object if cur_object > 0 else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(score),
|
||||
details={
|
||||
"object_detected": cur_object,
|
||||
"color_correct": cur_object_color
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,134 @@
|
||||
"""VBench Dynamic Degree — RAFT optical flow motion detection.
|
||||
|
||||
For each consecutive frame pair, computes optical flow via RAFT and takes
|
||||
the mean of the top 5% flow magnitudes. If enough pairs exceed an
|
||||
adaptive threshold, the video is classified as dynamic (1.0) vs static (0.0).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from easydict import EasyDict
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.dynamic_degree")
|
||||
class DynamicDegreeMetric(BaseMetric):
|
||||
|
||||
name = "vbench.dynamic_degree"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["easydict"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
self._chunk_size = 16
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from vbench.third_party.RAFT.core.raft import RAFT
|
||||
|
||||
args = EasyDict(small=False, mixed_precision=False, alternate_corr=False, dropout=0.0)
|
||||
model = torch.nn.DataParallel(RAFT(args))
|
||||
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
ckpt_path = ensure_checkpoint(
|
||||
"raft-things.pth",
|
||||
source="sbalani/raft-things",
|
||||
filename="raft-things.pth",
|
||||
)
|
||||
model.load_state_dict(torch.load(ckpt_path, map_location="cpu"))
|
||||
model = model.module
|
||||
model.to(self.device)
|
||||
model.eval()
|
||||
self._model = model
|
||||
|
||||
def _get_score(self, flow: torch.Tensor) -> float:
|
||||
"""Top-5% mean flow magnitude (matching VBench dynamic_degree.get_score)."""
|
||||
flo = flow.permute(1, 2, 0).cpu().numpy()
|
||||
rad = np.sqrt(flo[..., 0]**2 + flo[..., 1]**2)
|
||||
h, w = rad.shape
|
||||
cut = max(1, int(h * w * 0.05))
|
||||
rad_flat = rad.flatten()
|
||||
return float(np.mean(np.sort(rad_flat)[-cut:]))
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from vbench.third_party.RAFT.core.utils_core.utils import InputPadder
|
||||
|
||||
video = sample["video"] # (T, C, H, W) [0, 1]
|
||||
T, _, H, W = video.shape
|
||||
|
||||
# fps controls the temporal sampling stride for optical flow.
|
||||
# vbench computes flow at 8fps (interval = round(fps/8)). The metric
|
||||
# cannot auto-derive fps from a tensor, so a missing fps would silently
|
||||
# use a wrong stride and produce a wrong score. Skip explicitly.
|
||||
if "fps" not in sample:
|
||||
return self._skip(sample, "missing 'fps' (required to set the "
|
||||
"8fps optical-flow sampling stride)")
|
||||
fps = float(sample["fps"])
|
||||
interval = max(1, round(fps / 8.0))
|
||||
|
||||
video_255 = video * 255.0
|
||||
chunk = self._chunk_size or 16
|
||||
|
||||
# Cap chunk so the RAFT correlation volume doesn't overflow int32.
|
||||
# RAFT downsamples 8x in the feature encoder; CorrBlock's tensor is
|
||||
# shape (B*H1*W1, 1, H2, W2) with H1=H2=H/8, W1=W2=W/8. Its element
|
||||
# count is B*(H/8)^2*(W/8)^2 — F.avg_pool2d's index space starts to
|
||||
# overflow int32 around 2^31. Safety factor 2x.
|
||||
h_red = max(1, H // 8)
|
||||
w_red = max(1, W // 8)
|
||||
max_chunk = max(1, (1 << 30) // (h_red * h_red * w_red * w_red))
|
||||
chunk = min(chunk, max_chunk)
|
||||
|
||||
indices = list(range(0, T, interval))
|
||||
n = len(indices)
|
||||
all_img1 = [video_255[indices[i]] for i in range(n - 1)]
|
||||
all_img2 = [video_255[indices[i + 1]] for i in range(n - 1)]
|
||||
|
||||
scores: list[float] = []
|
||||
for start in range(0, len(all_img1), chunk):
|
||||
end = min(start + chunk, len(all_img1))
|
||||
img1_batch = torch.stack(all_img1[start:end]).to(self.device)
|
||||
img2_batch = torch.stack(all_img2[start:end]).to(self.device)
|
||||
padder = InputPadder(img1_batch.shape)
|
||||
img1p, img2p = padder.pad(img1_batch, img2_batch)
|
||||
_, flow = self._model(img1p, img2p, iters=20, test_mode=True)
|
||||
for i in range(flow.shape[0]):
|
||||
scores.append(self._get_score(flow[i]))
|
||||
|
||||
scale = min(H, W)
|
||||
thres = 6.0 * (scale / 256.0)
|
||||
count_needed = round(4 * (n / 16.0))
|
||||
count_above = sum(1 for s in scores if s > thres)
|
||||
is_dynamic = 1.0 if count_above >= count_needed else 0.0
|
||||
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=is_dynamic,
|
||||
details={
|
||||
"per_pair_magnitude": scores,
|
||||
"threshold": thres,
|
||||
"count_above": count_above,
|
||||
"count_needed": count_needed,
|
||||
"fps": fps,
|
||||
"interval": interval,
|
||||
"n_frames_used": n
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,131 @@
|
||||
"""VBench Human Action — UMT ViT-L/16 action classification (Kinetics-400).
|
||||
|
||||
Classifies human actions in 16-frame clips. Top-5 predictions with
|
||||
confidence >= 0.85 are compared against the ground-truth action label.
|
||||
Score = 1.0 if match found, 0.0 otherwise.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
from torchvision.transforms.functional import resize, center_crop, normalize
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.io.video import extract_frames
|
||||
|
||||
# Kinetics-400 class names (loaded lazily). The label file ships inside
|
||||
# the upstream vbench submodule.
|
||||
_CAT_DICT: dict[str, str] | None = None
|
||||
|
||||
|
||||
def _load_cat_dict() -> dict[str, str]:
|
||||
global _CAT_DICT
|
||||
if _CAT_DICT is not None:
|
||||
return _CAT_DICT
|
||||
import vbench.third_party.umt as _umt_pkg
|
||||
cat_path = (Path(_umt_pkg.__file__).resolve().parent / "kinetics_400_categories.txt")
|
||||
out: dict[str, str] = {}
|
||||
with cat_path.open() as f:
|
||||
for line in f:
|
||||
parts = line.strip().split("\t")
|
||||
if len(parts) == 2:
|
||||
cat, idx = parts
|
||||
out[idx] = cat.lower()
|
||||
_CAT_DICT = out
|
||||
return _CAT_DICT
|
||||
|
||||
|
||||
@register("vbench.human_action")
|
||||
class HumanActionMetric(BaseMetric):
|
||||
|
||||
name = "vbench.human_action"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["timm"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from timm.models import create_model
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
|
||||
ckpt_path = ensure_checkpoint(
|
||||
"umt_l16_kinetics400.pth",
|
||||
source="OpenGVLab/VBench_Used_Models",
|
||||
filename="l16_ptk710_ftk710_ftk400_f16_res224.pth",
|
||||
)
|
||||
|
||||
import vbench.third_party.umt.models.modeling_finetune # noqa: F401
|
||||
|
||||
self._model = create_model(
|
||||
"vit_large_patch16_224",
|
||||
pretrained=False,
|
||||
num_classes=400,
|
||||
all_frames=16,
|
||||
tubelet_size=1,
|
||||
use_learnable_pos_emb=False,
|
||||
fc_drop_rate=0.0,
|
||||
drop_rate=0.0,
|
||||
drop_path_rate=0.2,
|
||||
attn_drop_rate=0.0,
|
||||
drop_block_rate=None,
|
||||
use_checkpoint=False,
|
||||
checkpoint_num=16,
|
||||
use_mean_pooling=True,
|
||||
init_scale=0.001,
|
||||
)
|
||||
state_dict = torch.load(ckpt_path, map_location="cpu", weights_only=False)
|
||||
self._model.load_state_dict(state_dict, strict=False)
|
||||
self._model.to(self.device)
|
||||
self._model.eval()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W) [0, 1]
|
||||
text_prompt = sample.get("text_prompt")
|
||||
if text_prompt is None:
|
||||
return self._skip(sample, "missing text_prompt with action labels")
|
||||
|
||||
cat_dict = _load_cat_dict()
|
||||
|
||||
frames = extract_frames(video, 16) # (16, C, H, W)
|
||||
frames = resize(frames, 256, antialias=True)
|
||||
frames = center_crop(frames, 224)
|
||||
frames = normalize(frames, mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
|
||||
# UMT expects (C, T, H, W); add a leading batch dim of 1.
|
||||
clip_in = frames.permute(1, 0, 2, 3).unsqueeze(0).to(self.device)
|
||||
|
||||
logits = torch.sigmoid(self._model(clip_in)) # (1, 400)
|
||||
top_scores, top_indices = torch.topk(logits[0], 5)
|
||||
top_indices = top_indices.tolist()
|
||||
top_scores = top_scores.tolist()
|
||||
|
||||
predictions = [
|
||||
cat_dict.get(str(idx), "") for idx, score in zip(top_indices, top_scores, strict=False) if score >= 0.85
|
||||
]
|
||||
gt_label = text_prompt.lower().strip()
|
||||
match = any(pred == gt_label for pred in predictions)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=1.0 if match else 0.0,
|
||||
details={
|
||||
"predictions": predictions,
|
||||
"ground_truth": gt_label
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,71 @@
|
||||
"""VBench Imaging Quality — MUSIQ-based per-frame technical quality.
|
||||
|
||||
Uses MUSIQ (Multi-Scale Image Quality) from pyiqa. Frames are resized
|
||||
so the longer side is at most 512px. Score = mean(MUSIQ_scores) / 100.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
from torchvision.transforms.functional import resize
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.imaging_quality")
|
||||
class ImagingQualityMetric(BaseMetric):
|
||||
|
||||
name = "vbench.imaging_quality"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["pyiqa"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
import pyiqa
|
||||
self._model = pyiqa.create_metric("musiq-spaq", device=self.device)
|
||||
self._model.eval()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
T, _, H, W = video.shape
|
||||
|
||||
if max(H, W) > 512:
|
||||
scale = 512.0 / max(H, W)
|
||||
new_h, new_w = int(H * scale), int(W * scale)
|
||||
else:
|
||||
new_h, new_w = H, W
|
||||
|
||||
frames = video.to(self.device)
|
||||
if (new_h, new_w) != (H, W):
|
||||
# antialias=False matches VBench's imaging_quality.transform
|
||||
frames = resize(frames, [new_h, new_w], antialias=False)
|
||||
|
||||
chunk = self._chunk_size or 32
|
||||
chunks: list[torch.Tensor] = []
|
||||
for i in range(0, T, chunk):
|
||||
scores = self._model(frames[i:i + chunk])
|
||||
chunks.append(scores.squeeze(-1))
|
||||
per_frame = torch.cat(chunks, dim=0) # (T,)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(per_frame.mean().item()) / 100.0,
|
||||
details={"per_frame_raw": per_frame.tolist()},
|
||||
)
|
||||
@@ -0,0 +1,200 @@
|
||||
"""VBench Motion Smoothness — AMT-S frame interpolation quality.
|
||||
|
||||
Takes every-other frame, uses AMT-S to interpolate the missing middle
|
||||
frames, then compares interpolated vs actual frames.
|
||||
Score = (255 - mean_pixel_diff) / 255. Higher = smoother motion.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import os
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.motion_smoothness")
|
||||
class MotionSmoothnessMetric(BaseMetric):
|
||||
|
||||
name = "vbench.motion_smoothness"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["omegaconf"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
self._embt: Any = None
|
||||
self._chunk_size = 8
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
if self._embt is not None:
|
||||
self._embt = self._embt.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from omegaconf import OmegaConf
|
||||
import vbench.third_party.amt as _amt_pkg
|
||||
from vbench.third_party.amt.utils.build_utils import build_from_cfg
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
|
||||
amt_dir = os.path.dirname(_amt_pkg.__file__)
|
||||
cfg_path = os.path.join(amt_dir, "cfgs", "AMT-S.yaml")
|
||||
|
||||
ckpt_path = ensure_checkpoint(
|
||||
"amt-s.pth",
|
||||
source="https://huggingface.co/lalala125/AMT/resolve/main/amt-s.pth",
|
||||
)
|
||||
|
||||
network_cfg = OmegaConf.load(cfg_path).network
|
||||
self._model = build_from_cfg(network_cfg)
|
||||
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
|
||||
self._model.load_state_dict(ckpt["state_dict"])
|
||||
self._model.to(self.device)
|
||||
self._model.eval()
|
||||
|
||||
self._embt = torch.tensor(1 / 2).float().view(1, 1, 1, 1).to(self.device)
|
||||
|
||||
def _get_scale(self, h: int, w: int) -> float:
|
||||
"""Pick a downscale factor that keeps AMT's correlation volume
|
||||
within free GPU memory.
|
||||
|
||||
Re-queries free memory on every call (rather than caching at setup
|
||||
time) so the scale adapts to whatever's actually available — other
|
||||
metric replicas already loaded, residual generator allocations,
|
||||
another process sharing the GPU, etc. The upstream version cached
|
||||
``total_memory`` at setup, which on a shared/loaded GPU lets AMT
|
||||
attempt a 30+ GB correlation volume reshape and OOM.
|
||||
"""
|
||||
if self.device.type != "cuda":
|
||||
return 1.0
|
||||
# Free memory that won't be claimed by other allocations during this
|
||||
# forward pass. min(free, total) is conservative against transient
|
||||
# spikes; mem_get_info returns (free, total) in bytes.
|
||||
free_bytes, _ = torch.cuda.mem_get_info(self.device)
|
||||
anchor_resolution = 1024 * 512
|
||||
anchor_memory = 1500 * 1024**2
|
||||
anchor_memory_bias = 2500 * 1024**2
|
||||
if free_bytes <= anchor_memory_bias:
|
||||
# Less than the model + scratch overhead is free; force the
|
||||
# most aggressive downscale we support.
|
||||
return 1 / 16
|
||||
scale = anchor_resolution / (h * w) * np.sqrt((free_bytes - anchor_memory_bias) / anchor_memory)
|
||||
if scale >= 1.0:
|
||||
return 1.0
|
||||
scale = 1 / np.floor(1 / np.sqrt(scale) * 16) * 16
|
||||
return float(scale)
|
||||
|
||||
_MIN_AMT_SCALE = 1 / 16
|
||||
|
||||
def _safe_amt_forward(self, in0: Any, in1: Any, embt: Any, scale: Any) -> Any:
|
||||
"""Run one chunk through AMT with OOM-retry on two axes.
|
||||
|
||||
Recovery strategy on ``CUDA out of memory``:
|
||||
|
||||
1. **Halve the batch** until batch=1. AMT's per-pair memory
|
||||
dominates; splitting helps until each pair is on its own.
|
||||
2. **Halve the scale_factor** passed to AMT (which controls its
|
||||
internal feature-map resolution and therefore the correlation
|
||||
volume size). Bottoms out at ``_MIN_AMT_SCALE`` — beyond that
|
||||
the feature maps are too coarse to produce meaningful
|
||||
interpolation and we re-raise.
|
||||
|
||||
Self-tunes under memory pressure: the upstream autoscale formula
|
||||
in :meth:`_get_scale` mis-extrapolates at large resolutions
|
||||
(treats memory as linear in pixel count, but AMT's correlation
|
||||
volume grows quadratically). This retry path makes the metric
|
||||
robust to that without rewriting the formula.
|
||||
"""
|
||||
try:
|
||||
return self._model(in0, in1, embt, scale_factor=scale, eval=True)["imgt_pred"]
|
||||
except torch.cuda.OutOfMemoryError:
|
||||
torch.cuda.empty_cache()
|
||||
bs = in0.shape[0]
|
||||
if bs > 1:
|
||||
half = bs // 2
|
||||
a = self._safe_amt_forward(in0[:half], in1[:half], embt[:half], scale)
|
||||
b = self._safe_amt_forward(in0[half:], in1[half:], embt[half:], scale)
|
||||
return torch.cat([a, b], dim=0)
|
||||
if scale > self._MIN_AMT_SCALE:
|
||||
return self._safe_amt_forward(in0, in1, embt, scale / 2)
|
||||
raise
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from vbench.third_party.amt.utils.utils import (
|
||||
img2tensor,
|
||||
tensor2img,
|
||||
check_dim_and_resize,
|
||||
InputPadder,
|
||||
)
|
||||
|
||||
video = sample["video"] # (T, C, H, W) [0, 1]
|
||||
chunk = self._chunk_size or 8
|
||||
|
||||
frames_np = (video * 255).to(torch.uint8).cpu().numpy()
|
||||
frames_np = [f.transpose(1, 2, 0) for f in frames_np] # list of (H,W,C)
|
||||
|
||||
even_indices = list(range(0, len(frames_np), 2))
|
||||
if len(even_indices) <= 1:
|
||||
return MetricResult(name=self.name, score=1.0, details={})
|
||||
|
||||
even_frames = [frames_np[i] for i in even_indices]
|
||||
inputs = [img2tensor(f).to(self.device) for f in even_frames]
|
||||
inputs = check_dim_and_resize(inputs)
|
||||
|
||||
h, w = inputs[0].shape[-2:]
|
||||
scale = self._get_scale(h, w)
|
||||
padding = int(16 / scale)
|
||||
padder = InputPadder(inputs[0].shape, padding)
|
||||
inputs = padder.pad(*inputs)
|
||||
|
||||
n_pairs = len(inputs) - 1
|
||||
all_in0 = [inputs[i] for i in range(n_pairs)]
|
||||
all_in1 = [inputs[i + 1] for i in range(n_pairs)]
|
||||
all_gt = [
|
||||
frames_np[even_indices[i] + 1] if even_indices[i] + 1 < len(frames_np) else frames_np[-1]
|
||||
for i in range(n_pairs)
|
||||
]
|
||||
|
||||
all_preds = []
|
||||
for start in range(0, len(all_in0), chunk):
|
||||
end = min(start + chunk, len(all_in0))
|
||||
in0_batch = torch.cat(all_in0[start:end], dim=0).to(self.device)
|
||||
in1_batch = torch.cat(all_in1[start:end], dim=0).to(self.device)
|
||||
embt = self._embt.expand(in0_batch.shape[0], -1, -1, -1)
|
||||
pred = self._safe_amt_forward(in0_batch, in1_batch, embt, scale)
|
||||
all_preds.append(pred.cpu())
|
||||
all_preds = torch.cat(all_preds, dim=0)
|
||||
|
||||
diffs: list[float] = []
|
||||
for i in range(n_pairs):
|
||||
pred = all_preds[i:i + 1]
|
||||
pred_unpadded = padder.unpad(pred)[0]
|
||||
pred_np = tensor2img(pred_unpadded)
|
||||
gt_np = all_gt[i]
|
||||
# gt comes from frames_np; pred goes through check_dim_and_resize +
|
||||
# AMT pad/unpad, which can reshape. Match shapes before absdiff.
|
||||
if gt_np.shape[:2] != pred_np.shape[:2]:
|
||||
gt_np = cv2.resize(gt_np, (pred_np.shape[1], pred_np.shape[0]), interpolation=cv2.INTER_AREA)
|
||||
diffs.append(float(np.mean(cv2.absdiff(gt_np, pred_np))))
|
||||
|
||||
vfi_score = float(np.mean(diffs)) if diffs else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=(255.0 - vfi_score) / 255.0,
|
||||
details={"vfi_score": vfi_score},
|
||||
)
|
||||
@@ -0,0 +1,73 @@
|
||||
"""VBench Multiple Objects — GRiT detection for dual-object presence.
|
||||
|
||||
Checks if BOTH target objects are detected in each of 16 sampled frames.
|
||||
Score = matching_frames / total_frames.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.multiple_objects")
|
||||
class MultipleObjectsMetric(BaseMetric):
|
||||
|
||||
name = "vbench.multiple_objects"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["detectron2"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model = None
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import load_grit_model
|
||||
# VBench's multiple_objects uses ObjectDet head
|
||||
self._model = load_grit_model(self.device, task="ObjectDet")
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import prepare_frames, detect_frames
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
aux = sample.get("auxiliary_info") or {}
|
||||
if "object" not in aux:
|
||||
return self._skip(sample, "missing 'object' in auxiliary_info")
|
||||
|
||||
object_info = aux["object"]
|
||||
if " and " not in object_info:
|
||||
# multiple_objects expects "<a> and <b>"; single objects are
|
||||
# the object_class metric's territory — skip this row.
|
||||
return self._skip(sample, "'object' lacks ' and ' separator")
|
||||
key_a, key_b = [k.strip() for k in object_info.split(" and ")]
|
||||
|
||||
frames_np = prepare_frames(video)
|
||||
preds = detect_frames(self._model, frames_np)
|
||||
|
||||
matching = 0
|
||||
for frame_pred in preds:
|
||||
try:
|
||||
obj_set = set(frame_pred[0][2]) if frame_pred else set()
|
||||
except (IndexError, TypeError):
|
||||
obj_set = set()
|
||||
if key_a in obj_set and key_b in obj_set:
|
||||
matching += 1
|
||||
|
||||
total = len(preds)
|
||||
score = matching / total if total > 0 else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(score),
|
||||
details={
|
||||
"matching_frames": matching,
|
||||
"total_frames": total
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,71 @@
|
||||
"""VBench Object Class — GRiT object detection for class matching.
|
||||
|
||||
Checks if a target object class is detected in each of 16 sampled frames.
|
||||
Score = matching_frames / total_frames.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.object_class")
|
||||
class ObjectClassMetric(BaseMetric):
|
||||
|
||||
name = "vbench.object_class"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["detectron2"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model = None
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import load_grit_model
|
||||
# VBench's object_class uses ObjectDet head (init_submodules → "ObjectDet")
|
||||
self._model = load_grit_model(self.device, task="ObjectDet")
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import prepare_frames, detect_frames
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
aux = sample.get("auxiliary_info") or {}
|
||||
if "object" not in aux:
|
||||
return self._skip(sample, "missing 'object' in auxiliary_info")
|
||||
|
||||
object_key = aux["object"]
|
||||
if " and " in object_key:
|
||||
# multiple_objects' territory; skip this row for object_class.
|
||||
return self._skip(sample, "'object' contains ' and ' (multi-object)")
|
||||
|
||||
frames_np = prepare_frames(video)
|
||||
preds = detect_frames(self._model, frames_np)
|
||||
|
||||
matching = 0
|
||||
for frame_pred in preds:
|
||||
try:
|
||||
obj_set = set(frame_pred[0][2]) if frame_pred else set()
|
||||
except (IndexError, TypeError):
|
||||
obj_set = set()
|
||||
if object_key in obj_set:
|
||||
matching += 1
|
||||
|
||||
total = len(preds)
|
||||
score = matching / total if total > 0 else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=float(score),
|
||||
details={
|
||||
"matching_frames": matching,
|
||||
"total_frames": total
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,93 @@
|
||||
"""VBench Overall Consistency — ViCLIP text-video alignment.
|
||||
|
||||
Encodes 8 sampled video frames via ViCLIP vision encoder and a text
|
||||
prompt via ViCLIP text encoder, then computes cosine similarity.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torchvision.transforms.functional import resize, center_crop, normalize
|
||||
from torchvision.transforms import InterpolationMode
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.io.video import extract_frames
|
||||
|
||||
_CLIP_MEAN = [0.48145466, 0.4578275, 0.40821073]
|
||||
_CLIP_STD = [0.26862954, 0.26130258, 0.27577711]
|
||||
|
||||
|
||||
def _clip_transform(frames: torch.Tensor) -> torch.Tensor:
|
||||
frames = resize(frames, 224, interpolation=InterpolationMode.BICUBIC, antialias=True)
|
||||
frames = center_crop(frames, 224)
|
||||
frames = normalize(frames, mean=_CLIP_MEAN, std=_CLIP_STD)
|
||||
return frames
|
||||
|
||||
|
||||
@register("vbench.overall_consistency")
|
||||
class OverallConsistencyMetric(BaseMetric):
|
||||
|
||||
name = "vbench.overall_consistency"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["timm", "einops", "clip"]
|
||||
backbone = "viclip"
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
self._tokenizer: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from vbench.third_party.ViCLIP.viclip import ViCLIP
|
||||
from vbench.third_party.ViCLIP.simple_tokenizer import SimpleTokenizer
|
||||
|
||||
# ViCLIP's tokenizer reuses OpenAI CLIP's BPE vocab. The file is
|
||||
# bundled with the ``openai-clip`` pip package (an ``[eval]`` extra)
|
||||
# — no separate download is needed; the model loader only handles
|
||||
# the actual .pth weights.
|
||||
from clip.simple_tokenizer import default_bpe
|
||||
self._tokenizer = SimpleTokenizer(default_bpe())
|
||||
|
||||
from fastvideo.eval.models import ensure_checkpoint
|
||||
ckpt = ensure_checkpoint(
|
||||
"ViClip-InternVid-10M-FLT.pth",
|
||||
source="OpenGVLab/VBench_Used_Models",
|
||||
filename="ViClip-InternVid-10M-FLT.pth",
|
||||
)
|
||||
|
||||
self._model = ViCLIP(tokenizer=self._tokenizer, pretrain=ckpt)
|
||||
self._model.to(self.device)
|
||||
self._model.eval()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
text_prompt = sample.get("text_prompt")
|
||||
if text_prompt is None:
|
||||
return self._skip(sample, "missing text_prompt")
|
||||
|
||||
frames = _clip_transform(extract_frames(video, 8)) # (8, C, H, W)
|
||||
clip_in = frames.unsqueeze(0).to(self.device) # (1, 8, C, H, W)
|
||||
|
||||
vid_feat = self._model.encode_vision(clip_in, test=True).float()
|
||||
vid_feat = F.normalize(vid_feat, dim=-1, p=2) # (1, D)
|
||||
|
||||
text_feat = self._model.encode_text(text_prompt).float()
|
||||
text_feat = F.normalize(text_feat, dim=-1, p=2)
|
||||
score = float((vid_feat @ text_feat.T)[0][0].cpu())
|
||||
return MetricResult(name=self.name, score=score, details={})
|
||||
@@ -0,0 +1,168 @@
|
||||
"""VLM-based scene matching using AVoCaDO (Qwen2.5-Omni).
|
||||
|
||||
Replaces VBench's Tag2Text-based scene metric with a modern VLM caption.
|
||||
The algorithm follows VBench:
|
||||
1. Caption the video
|
||||
2. Check if all scene keywords appear in the caption
|
||||
3. Score = 1.0 if all match, 0.0 otherwise
|
||||
|
||||
Unlike VBench (which captions each frame separately with Tag2Text),
|
||||
AVoCaDO captions the entire video in one pass with rich natural language,
|
||||
making the keyword check more robust.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import os
|
||||
import tempfile
|
||||
|
||||
import torch
|
||||
import torchvision.io
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
_SCENE_PROMPT = ("Describe the visual scene in this video, including the location, "
|
||||
"environment, objects, and overall setting. Be specific and use "
|
||||
"concrete descriptive words.")
|
||||
|
||||
|
||||
@register("vbench.scene")
|
||||
class SceneMetric(BaseMetric):
|
||||
|
||||
name = "vbench.scene"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["transformers", "qwen_omni_utils"]
|
||||
backbone = "avocado"
|
||||
|
||||
def __init__(self, model_path: str = "AVoCaDO-Captioner/AVoCaDO") -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
self._processor: Any = None
|
||||
self._model_path = model_path
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from transformers import Qwen2_5OmniForConditionalGeneration, Qwen2_5OmniProcessor
|
||||
|
||||
# AVoCaDO uses Qwen2.5-Omni — large multimodal model
|
||||
os.environ.setdefault("VIDEO_MAX_PIXELS", str(20070400)) # 512*28*28*50
|
||||
|
||||
self._model = Qwen2_5OmniForConditionalGeneration.from_pretrained(
|
||||
self._model_path,
|
||||
torch_dtype=torch.bfloat16,
|
||||
device_map=str(self.device) if self.device.type == "cuda" else None,
|
||||
)
|
||||
self._model.disable_talker()
|
||||
self._model.eval()
|
||||
self._processor = Qwen2_5OmniProcessor.from_pretrained(self._model_path)
|
||||
|
||||
def _save_temp_video(self, video: torch.Tensor) -> str:
|
||||
"""Save (T, C, H, W) float [0,1] tensor as a temp mp4 file."""
|
||||
# torchvision expects (T, H, W, C) uint8
|
||||
frames = (video * 255).clamp(0, 255).to(torch.uint8).permute(0, 2, 3, 1).cpu()
|
||||
# Caller owns the resulting file (it's read by Qwen2.5-Omni and
|
||||
# cleaned up at end of compute()), so we just need a unique path.
|
||||
fd, path = tempfile.mkstemp(suffix=".mp4")
|
||||
os.close(fd)
|
||||
torchvision.io.write_video(path, frames, fps=8, video_codec="libx264", options={"crf": "18"})
|
||||
return path
|
||||
|
||||
def _generate_caption(self, video_path: str) -> str:
|
||||
from qwen_omni_utils import process_mm_info
|
||||
|
||||
conversation = [
|
||||
{
|
||||
"role":
|
||||
"system",
|
||||
"content": [{
|
||||
"type":
|
||||
"text",
|
||||
"text": ("You are Qwen, a virtual human developed by the Qwen Team, "
|
||||
"Alibaba Group, capable of perceiving auditory and visual inputs.")
|
||||
}],
|
||||
},
|
||||
{
|
||||
"role":
|
||||
"user",
|
||||
"content": [
|
||||
{
|
||||
"type": "video",
|
||||
"video": video_path,
|
||||
"max_pixels": 401408
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": _SCENE_PROMPT
|
||||
},
|
||||
],
|
||||
},
|
||||
]
|
||||
|
||||
text = self._processor.apply_chat_template(conversation, add_generation_prompt=True, tokenize=False)
|
||||
# AVoCaDO is video+audio; for scene matching we don't need audio,
|
||||
# but the model expects it so let it process
|
||||
audios, images, videos = process_mm_info(conversation, use_audio_in_video=False)
|
||||
inputs = self._processor(
|
||||
text=text,
|
||||
audio=audios,
|
||||
images=images,
|
||||
videos=videos,
|
||||
return_tensors="pt",
|
||||
padding=True,
|
||||
use_audio_in_video=False,
|
||||
)
|
||||
inputs = inputs.to(self._model.device).to(self._model.dtype)
|
||||
|
||||
with torch.no_grad():
|
||||
text_ids = self._model.generate(
|
||||
**inputs,
|
||||
use_audio_in_video=False,
|
||||
return_audio=False,
|
||||
do_sample=False,
|
||||
thinker_max_new_tokens=512,
|
||||
)
|
||||
|
||||
decoded = self._processor.batch_decode(text_ids, skip_special_tokens=True,
|
||||
clean_up_tokenization_spaces=False)[0]
|
||||
return decoded.split("\nassistant\n")[-1].lower()
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
aux = sample.get("auxiliary_info") or {}
|
||||
if "scene" not in aux:
|
||||
return self._skip(sample, "missing 'scene' in auxiliary_info")
|
||||
|
||||
scene_keywords = aux["scene"]
|
||||
keywords = [k.strip().lower() for k in scene_keywords.split() if k.strip()]
|
||||
|
||||
tmp_path = self._save_temp_video(video)
|
||||
try:
|
||||
caption = self._generate_caption(tmp_path)
|
||||
finally:
|
||||
os.unlink(tmp_path)
|
||||
|
||||
matched = [kw for kw in keywords if kw in caption]
|
||||
score = 1.0 if len(matched) == len(keywords) else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=score,
|
||||
details={
|
||||
"caption": caption[:500],
|
||||
"keywords": keywords,
|
||||
"matched": matched,
|
||||
},
|
||||
)
|
||||
@@ -0,0 +1,123 @@
|
||||
"""VBench Spatial Relationship — GRiT detection + bbox position scoring.
|
||||
|
||||
Detects two target objects via GRiT and checks if their bounding boxes
|
||||
satisfy the expected spatial relationship (left/right/above/below).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
def _get_position_score(locality: str, obj1: list, obj2: list, iou_threshold: float = 0.1) -> float:
|
||||
"""Score spatial relationship between two bboxes [x0, y0, x1, y1].
|
||||
|
||||
Matching VBench's get_position_score() exactly.
|
||||
"""
|
||||
box1_center = ((obj1[0] + obj1[2]) / 2, (obj1[1] + obj1[3]) / 2)
|
||||
box2_center = ((obj2[0] + obj2[2]) / 2, (obj2[1] + obj2[3]) / 2)
|
||||
|
||||
x_distance = box2_center[0] - box1_center[0]
|
||||
y_distance = box2_center[1] - box1_center[1]
|
||||
|
||||
# IoU
|
||||
x_overlap = max(0, min(obj1[2], obj2[2]) - max(obj1[0], obj2[0]))
|
||||
y_overlap = max(0, min(obj1[3], obj2[3]) - max(obj1[1], obj2[1]))
|
||||
intersection = x_overlap * y_overlap
|
||||
area1 = (obj1[2] - obj1[0]) * (obj1[3] - obj1[1])
|
||||
area2 = (obj2[2] - obj2[0]) * (obj2[3] - obj2[1])
|
||||
union = area1 + area2 - intersection
|
||||
iou = intersection / union if union > 0 else 0
|
||||
|
||||
if "right" in locality or "left" in locality:
|
||||
if abs(x_distance) > abs(y_distance) and iou < iou_threshold:
|
||||
return 1.0
|
||||
elif abs(x_distance) > abs(y_distance) and iou >= iou_threshold:
|
||||
return iou_threshold / iou
|
||||
return 0.0
|
||||
elif "bottom" in locality or "top" in locality:
|
||||
if abs(y_distance) > abs(x_distance) and iou < iou_threshold:
|
||||
return 1.0
|
||||
elif abs(y_distance) > abs(x_distance) and iou >= iou_threshold:
|
||||
return iou_threshold / iou
|
||||
return 0.0
|
||||
return 0.0
|
||||
|
||||
|
||||
@register("vbench.spatial_relationship")
|
||||
class SpatialRelationshipMetric(BaseMetric):
|
||||
|
||||
name = "vbench.spatial_relationship"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["detectron2"]
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import load_grit_model
|
||||
# VBench's spatial_relationship uses ObjectDet head and matches
|
||||
# pred[0] against class names like "person"/"grass"
|
||||
self._model = load_grit_model(self.device, task="ObjectDet")
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
from fastvideo.eval.metrics.vbench._grit_helper import prepare_frames
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
aux = sample.get("auxiliary_info") or {}
|
||||
if "spatial_relationship" not in aux:
|
||||
return self._skip(sample, "missing 'spatial_relationship' in auxiliary_info")
|
||||
|
||||
sp_info = aux["spatial_relationship"]
|
||||
try:
|
||||
key_a = sp_info["object_a"]
|
||||
key_b = sp_info["object_b"]
|
||||
relation = sp_info["relationship"]
|
||||
except (KeyError, TypeError):
|
||||
return self._skip(sample, "spatial_relationship missing object_a/object_b/relationship")
|
||||
|
||||
frames_np = prepare_frames(video)
|
||||
|
||||
preds = []
|
||||
for frame in frames_np:
|
||||
ret = self._model.run_caption_tensor(frame)
|
||||
frame_dets = []
|
||||
if len(ret[0]) > 0:
|
||||
for info in ret[0]:
|
||||
frame_dets.append([info[0], info[1]]) # (caption, bbox)
|
||||
preds.append(frame_dets)
|
||||
|
||||
# Score each frame (matching VBench's check_generate).
|
||||
frame_scores: list[float] = []
|
||||
for frame_pred in preds:
|
||||
obj_bboxes = [item[1] for item in frame_pred if item[0] == key_a or item[0] == key_b]
|
||||
|
||||
cur_scores = [0.0]
|
||||
for i in range(len(obj_bboxes) - 1):
|
||||
for j in range(i + 1, len(obj_bboxes)):
|
||||
cur_scores.append(_get_position_score(
|
||||
relation,
|
||||
obj_bboxes[i],
|
||||
obj_bboxes[j],
|
||||
))
|
||||
frame_scores.append(max(cur_scores))
|
||||
|
||||
score = float(np.mean(frame_scores)) if frame_scores else 0.0
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=score,
|
||||
details={"per_frame": frame_scores},
|
||||
)
|
||||
@@ -0,0 +1,72 @@
|
||||
"""VBench Subject Consistency — DINO ViT-B/16 temporal feature similarity.
|
||||
|
||||
Measures how well the main subject maintains its appearance throughout
|
||||
the video via cosine similarity of DINO features between consecutive
|
||||
frames and the first frame.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torchvision.transforms.functional import resize, normalize
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
from fastvideo.eval.metrics.vbench._utils import consistency_score
|
||||
|
||||
# ImageNet normalization (used by DINO)
|
||||
_MEAN = [0.485, 0.456, 0.406]
|
||||
_STD = [0.229, 0.224, 0.225]
|
||||
|
||||
|
||||
@register("vbench.subject_consistency")
|
||||
class SubjectConsistencyMetric(BaseMetric):
|
||||
|
||||
name = "vbench.subject_consistency"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
backbone = "dino_vitb16"
|
||||
|
||||
def __init__(self) -> None:
|
||||
super().__init__()
|
||||
self._model: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
model = torch.hub.load("facebookresearch/dino:main", "dino_vitb16")
|
||||
model.to(self.device)
|
||||
model.eval()
|
||||
self._model = model
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
frames = video.to(self.device)
|
||||
# antialias=False matches VBench's dino_transform (vbench/utils.py:50)
|
||||
frames = resize(frames, 224, antialias=False)
|
||||
frames = normalize(frames, mean=_MEAN, std=_STD)
|
||||
|
||||
chunk = self._chunk_size or 64
|
||||
feats = []
|
||||
for i in range(0, frames.shape[0], chunk):
|
||||
f = self._model(frames[i:i + chunk])
|
||||
f = F.normalize(f, dim=-1, p=2)
|
||||
feats.append(f)
|
||||
all_feats = torch.cat(feats, dim=0) # (T, D)
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=consistency_score(all_feats),
|
||||
details={},
|
||||
)
|
||||
@@ -0,0 +1,40 @@
|
||||
"""VBench Temporal Flickering — measures frame-to-frame stability.
|
||||
|
||||
Score = (255 - mean_MAE) / 255, where MAE is computed between consecutive
|
||||
frames in uint8 [0, 255] space. Higher = less flickering.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
|
||||
@register("vbench.temporal_flickering")
|
||||
class TemporalFlickeringMetric(BaseMetric):
|
||||
|
||||
name = "vbench.temporal_flickering"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = False
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
video = sample["video"] # (T, C, H, W) [0, 1]
|
||||
T = video.shape[0]
|
||||
if T <= 1:
|
||||
return MetricResult(name=self.name, score=1.0, details={})
|
||||
|
||||
frames = (video * 255.0).to(torch.uint8).cpu().numpy()
|
||||
frames = frames.transpose(0, 2, 3, 1).astype(np.float32)
|
||||
mae_per_pair = [float(np.mean(np.abs(frames[t] - frames[t + 1]))) for t in range(T - 1)]
|
||||
mean_mae = float(np.mean(mae_per_pair))
|
||||
return MetricResult(
|
||||
name=self.name,
|
||||
score=(255.0 - mean_mae) / 255.0,
|
||||
details={"per_pair_mae": mae_per_pair},
|
||||
)
|
||||
@@ -0,0 +1,17 @@
|
||||
"""VBench Temporal Style — ViCLIP text-video alignment (style focus).
|
||||
|
||||
Identical logic to overall_consistency — same ViCLIP cosine similarity.
|
||||
The difference is semantic: overall_consistency measures general prompt
|
||||
alignment while temporal_style measures style consistency over time.
|
||||
VBench uses different prompts for each from its metadata JSON.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.metrics.vbench.overall_consistency.metric import OverallConsistencyMetric
|
||||
|
||||
|
||||
@register("vbench.temporal_style")
|
||||
class TemporalStyleMetric(OverallConsistencyMetric):
|
||||
name = "vbench.temporal_style"
|
||||
@@ -0,0 +1,309 @@
|
||||
"""VideoScore2 — VLM-based video quality scoring.
|
||||
|
||||
Uses a Qwen2.5-VL model fine-tuned to score generated videos on three
|
||||
dimensions: visual quality, text-to-video alignment, and physical
|
||||
consistency. Scores are extracted from token logits as upstream's
|
||||
``ll_based_soft_score_normed`` weighting (1-5 scale).
|
||||
|
||||
Reference: TIGER-AI-Lab/VideoScore2 (vs2_inference.py).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from string import Template
|
||||
from typing import Any
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult
|
||||
|
||||
# Match upstream verbatim, including leading newline and 4-space indents
|
||||
# (TIGER-AI-Lab/VideoScore2/vs2_inference.py).
|
||||
VS2_QUERY_TEMPLATE = Template("""
|
||||
You are an expert for evaluating AI-generated videos from three dimensions:
|
||||
(1) visual quality – clarity, smoothness, artifacts;
|
||||
(2) text-to-video alignment – fidelity to the prompt;
|
||||
(3) physical/common-sense consistency – naturalness and physics plausibility.
|
||||
|
||||
Video prompt: $t2v_prompt
|
||||
|
||||
Please output in this format:
|
||||
visual quality: <v_score>;
|
||||
text-to-video alignment: <t_score>,
|
||||
physical/common-sense consistency: <p_score>
|
||||
""")
|
||||
|
||||
# The released VideoScore2 model emits a <think>...</think> chain-of-
|
||||
# thought followed by a numbered list of the form:
|
||||
#
|
||||
# (1) visual quality – clarity, smoothness, artifacts: 3
|
||||
# (2) text-to-video alignment – fidelity to the prompt: 4
|
||||
# (3) physical/common-sense consistency – naturalness and physics …: 3
|
||||
#
|
||||
# Upstream's vs2_inference.py regex (``visual quality:\s*(\d+)``) does
|
||||
# not match this output — it expects the colon directly after the
|
||||
# header, with no descriptor in between, so upstream's script returns
|
||||
# ``null`` on its own released model. We anchor on the ``(N)`` prefix
|
||||
# to avoid matching digits inside the chain-of-thought reasoning.
|
||||
SCORE_PATTERN = re.compile(
|
||||
r"\(1\)\s*visual quality[^\d]*?(\d+).*?"
|
||||
r"\(2\)\s*text-to-video alignment[^\d]*?(\d+).*?"
|
||||
r"\(3\)\s*physical/common-sense consistency[^\d]*?(\d+)",
|
||||
re.DOTALL | re.IGNORECASE,
|
||||
)
|
||||
|
||||
|
||||
def _find_score_token_index(prompt_text: str, tokenizer, gen_ids: list[int]) -> int:
|
||||
"""Find the token index where the score digit appears after prompt_text."""
|
||||
gen_str = tokenizer.decode(gen_ids, skip_special_tokens=False)
|
||||
pattern = r"(?:\(\d+\)\s*|\n\s*)?" + re.escape(prompt_text)
|
||||
match = re.search(pattern, gen_str, flags=re.IGNORECASE)
|
||||
if not match:
|
||||
return -1
|
||||
after = gen_str[match.end():]
|
||||
num_match = re.search(r"\d", after)
|
||||
if not num_match:
|
||||
return -1
|
||||
target = gen_str[:match.end() + num_match.start() + 1]
|
||||
for i in range(len(gen_ids)):
|
||||
if tokenizer.decode(gen_ids[:i + 1], skip_special_tokens=False) == target:
|
||||
return i
|
||||
return -1
|
||||
|
||||
|
||||
def _ll_based_soft_score_normed(hard_val: int | None,
|
||||
token_idx: int,
|
||||
scores,
|
||||
tokenizer,
|
||||
seq_idx: int = 0) -> float | None:
|
||||
"""Upstream VideoScore2's soft score: argmax_score × (argmax_prob / Σprob).
|
||||
|
||||
Matches ``ll_based_soft_score_normed`` in
|
||||
``TIGER-AI-Lab/VideoScore2/vs2_inference.py``. The ``seq_idx`` arg
|
||||
is the only addition (for batched generate). With ``B == 1`` the
|
||||
behaviour is identical to upstream.
|
||||
"""
|
||||
if hard_val is None or token_idx < 0:
|
||||
return None
|
||||
logits = scores[token_idx][seq_idx]
|
||||
score_probs = []
|
||||
for s in range(1, 6):
|
||||
ids = tokenizer.encode(str(s), add_special_tokens=False)
|
||||
if len(ids) == 1:
|
||||
logp = torch.log_softmax(logits, dim=-1)[ids[0]].item()
|
||||
score_probs.append((s, float(np.exp(logp))))
|
||||
if not score_probs:
|
||||
return None
|
||||
scores_list, probs_list = zip(*score_probs, strict=False)
|
||||
total_prob = sum(probs_list)
|
||||
max_prob = max(probs_list)
|
||||
best_score = scores_list[probs_list.index(max_prob)]
|
||||
normalized_prob = max_prob / total_prob if total_prob > 0 else 0
|
||||
return round(best_score * normalized_prob, 4)
|
||||
|
||||
|
||||
def _parse_output(output_text: str, scores, tokenizer, gen_ids: list[int], seq_idx: int = 0) -> dict:
|
||||
"""Parse scores from a single sequence's output."""
|
||||
match = SCORE_PATTERN.search(output_text)
|
||||
v_hard = int(match.group(1)) if match else None
|
||||
t_hard = int(match.group(2)) if match else None
|
||||
p_hard = int(match.group(3)) if match else None
|
||||
|
||||
if scores is not None:
|
||||
# Anchor on the numbered list to skip the chain-of-thought.
|
||||
idx_v = _find_score_token_index("(1) visual quality", tokenizer, gen_ids)
|
||||
idx_t = _find_score_token_index("(2) text-to-video alignment", tokenizer, gen_ids)
|
||||
idx_p = _find_score_token_index("(3) physical/common-sense consistency", tokenizer, gen_ids)
|
||||
v_soft = _ll_based_soft_score_normed(v_hard, idx_v, scores, tokenizer, seq_idx)
|
||||
t_soft = _ll_based_soft_score_normed(t_hard, idx_t, scores, tokenizer, seq_idx)
|
||||
p_soft = _ll_based_soft_score_normed(p_hard, idx_p, scores, tokenizer, seq_idx)
|
||||
else:
|
||||
v_soft = float(v_hard) if v_hard is not None else None
|
||||
t_soft = float(t_hard) if t_hard is not None else None
|
||||
p_soft = float(p_hard) if p_hard is not None else None
|
||||
|
||||
return {
|
||||
"visual_quality": v_soft,
|
||||
"text_alignment": t_soft,
|
||||
"physical_consistency": p_soft,
|
||||
"visual_quality_hard": v_hard,
|
||||
"text_alignment_hard": t_hard,
|
||||
"physical_consistency_hard": p_hard,
|
||||
"raw_output": output_text,
|
||||
}
|
||||
|
||||
|
||||
@register("videoscore2")
|
||||
class VideoScore2Metric(BaseMetric):
|
||||
"""VideoScore2: VLM-based video quality scoring (3 dimensions).
|
||||
|
||||
Requires ``sample["text_prompt"]`` for text-to-video alignment.
|
||||
Supports batched generation for GPU efficiency.
|
||||
"""
|
||||
|
||||
name = "videoscore2"
|
||||
requires_reference = False
|
||||
higher_is_better = True
|
||||
needs_gpu = True
|
||||
dependencies = ["transformers", "qwen_vl_utils"]
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
model_name: str = "TIGER-Lab/VideoScore2",
|
||||
infer_fps: float = 2.0,
|
||||
max_tokens: int = 1024,
|
||||
temperature: float = 0.7,
|
||||
do_sample: bool = True,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
self._model_name = model_name
|
||||
self.infer_fps = infer_fps
|
||||
self.max_tokens = max_tokens
|
||||
self.temperature = temperature
|
||||
self.do_sample = do_sample
|
||||
self._model: Any = None
|
||||
self._processor: Any = None
|
||||
self._tokenizer: Any = None
|
||||
|
||||
def to(self, device):
|
||||
super().to(device)
|
||||
if self._model is not None:
|
||||
self._model = self._model.to(self.device)
|
||||
return self
|
||||
|
||||
def setup(self) -> None:
|
||||
if self._model is not None:
|
||||
return
|
||||
from transformers import AutoProcessor, AutoTokenizer
|
||||
|
||||
# transformers ≥4.45 prefers AutoModelForImageTextToText for
|
||||
# vision-language models; AutoModelForVision2Seq is the legacy
|
||||
# alias and may go away in a future release.
|
||||
try:
|
||||
from transformers import AutoModelForImageTextToText as _AutoVisionModel
|
||||
except ImportError:
|
||||
from transformers import AutoModelForVision2Seq as _AutoVisionModel
|
||||
|
||||
self._model = _AutoVisionModel.from_pretrained(
|
||||
self._model_name,
|
||||
trust_remote_code=True,
|
||||
dtype=torch.bfloat16,
|
||||
).to(self.device)
|
||||
self._model.eval()
|
||||
self._processor = AutoProcessor.from_pretrained(
|
||||
self._model_name,
|
||||
trust_remote_code=True,
|
||||
)
|
||||
self._tokenizer = getattr(self._processor, "tokenizer", None)
|
||||
if self._tokenizer is None:
|
||||
self._tokenizer = AutoTokenizer.from_pretrained(
|
||||
self._model_name,
|
||||
trust_remote_code=True,
|
||||
use_fast=False,
|
||||
)
|
||||
|
||||
def _tensor_to_pil_list(self, video: torch.Tensor) -> list[Image.Image]:
|
||||
"""Convert (T, C, H, W) float [0,1] tensor to list of PIL images."""
|
||||
frames = (video.permute(0, 2, 3, 1).cpu().numpy() * 255).astype(np.uint8)
|
||||
return [Image.fromarray(frames[t]) for t in range(frames.shape[0])]
|
||||
|
||||
def _subsample_frames(self,
|
||||
pil_frames: list[Image.Image],
|
||||
max_frames: int = 64,
|
||||
max_resolution: int = 960) -> list[Image.Image]:
|
||||
"""Subsample to ~infer_fps worth of frames (max 64), resize if too large."""
|
||||
n = len(pil_frames)
|
||||
target = min(n, max_frames)
|
||||
if target < n:
|
||||
indices = np.linspace(0, n - 1, target, dtype=int)
|
||||
pil_frames = [pil_frames[i] for i in indices]
|
||||
|
||||
w, h = pil_frames[0].size
|
||||
if max(w, h) > max_resolution:
|
||||
scale = max_resolution / max(w, h)
|
||||
new_w, new_h = int(w * scale), int(h * scale)
|
||||
pil_frames = [f.resize((new_w, new_h), Image.LANCZOS) for f in pil_frames]
|
||||
|
||||
return pil_frames
|
||||
|
||||
@torch.no_grad()
|
||||
def compute(self, sample: dict) -> MetricResult:
|
||||
if self._model is None:
|
||||
self.setup()
|
||||
|
||||
from qwen_vl_utils import process_vision_info
|
||||
|
||||
video = sample["video"] # (T, C, H, W)
|
||||
text = sample.get("text_prompt", "")
|
||||
if isinstance(text, list):
|
||||
text = text[0] if text else ""
|
||||
|
||||
pil_frames = self._tensor_to_pil_list(video)
|
||||
pil_frames = self._subsample_frames(pil_frames)
|
||||
user_prompt = VS2_QUERY_TEMPLATE.substitute(t2v_prompt=text)
|
||||
|
||||
messages = [{
|
||||
"role":
|
||||
"user",
|
||||
"content": [
|
||||
{
|
||||
"type": "video",
|
||||
"video": pil_frames,
|
||||
"fps": self.infer_fps
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": user_prompt
|
||||
},
|
||||
]
|
||||
}]
|
||||
chat_text = self._processor.apply_chat_template(
|
||||
messages,
|
||||
tokenize=False,
|
||||
add_generation_prompt=True,
|
||||
)
|
||||
_, vid_inputs = process_vision_info(messages)
|
||||
|
||||
inputs = self._processor(
|
||||
text=[chat_text],
|
||||
videos=vid_inputs if vid_inputs else None,
|
||||
fps=self.infer_fps,
|
||||
padding=True,
|
||||
return_tensors="pt",
|
||||
).to(self.device)
|
||||
|
||||
input_len = inputs["input_ids"].shape[1]
|
||||
gen_kwargs: dict[str, Any] = dict(
|
||||
max_new_tokens=self.max_tokens,
|
||||
output_scores=True,
|
||||
return_dict_in_generate=True,
|
||||
do_sample=self.do_sample,
|
||||
)
|
||||
if self.do_sample:
|
||||
gen_kwargs["temperature"] = self.temperature
|
||||
gen_out = self._model.generate(**inputs, **gen_kwargs)
|
||||
|
||||
gen_ids = gen_out.sequences[0, input_len:].tolist()
|
||||
pad_id = self._tokenizer.pad_token_id
|
||||
if pad_id is not None:
|
||||
gen_ids = [t for t in gen_ids if t != pad_id]
|
||||
output_text = self._tokenizer.decode(gen_ids, skip_special_tokens=True)
|
||||
|
||||
parsed = _parse_output(
|
||||
output_text,
|
||||
gen_out.scores,
|
||||
self._tokenizer,
|
||||
gen_ids,
|
||||
seq_idx=0,
|
||||
)
|
||||
soft_vals = [
|
||||
v for v in (parsed["visual_quality"], parsed["text_alignment"], parsed["physical_consistency"])
|
||||
if v is not None
|
||||
]
|
||||
combined = sum(soft_vals) / len(soft_vals) if soft_vals else 0.0
|
||||
return MetricResult(name=self.name, score=combined, details=parsed)
|
||||
@@ -0,0 +1,98 @@
|
||||
"""Model checkpoint resolution and caching for eval metrics.
|
||||
|
||||
Single public function — :func:`ensure_checkpoint` — that hands a metric
|
||||
a local path to its weights, downloading on miss. Three source kinds:
|
||||
|
||||
* an existing local path → returned as-is;
|
||||
* an HTTP(S) URL → downloaded to ``get_cache_dir() / name`` via
|
||||
``huggingface_hub.http_get`` (resume + retries built in);
|
||||
* an HF repo id (``"org/repo"``) → :func:`huggingface_hub.hf_hub_download`
|
||||
if *filename* is given, else :func:`huggingface_hub.snapshot_download`.
|
||||
*name* is ignored in HF mode — HF manages its own cache key under
|
||||
``~/.cache/huggingface/hub``.
|
||||
|
||||
All download paths are filelock-safe (cooperative across threads,
|
||||
processes, and SLURM ranks) via :func:`fastvideo.utils.get_lock`.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo import envs
|
||||
from fastvideo.utils import get_lock
|
||||
|
||||
|
||||
def get_cache_dir() -> Path:
|
||||
"""Eval cache root.
|
||||
|
||||
Layout::
|
||||
|
||||
get_cache_dir() / models / ← URL-fetched checkpoints (LAION head,
|
||||
AMT, GRiT, …)
|
||||
get_cache_dir() / torch / ← redirected ``TORCH_HOME`` (DINO etc.)
|
||||
get_cache_dir() / clip / ← passed as ``download_root`` to
|
||||
``clip.load(...)`` callsites
|
||||
~/.cache/huggingface/hub / ← left at HF's default; widely shared
|
||||
with other ML projects
|
||||
|
||||
Override priority: ``FASTVIDEO_EVAL_CACHE`` > ``${FASTVIDEO_CACHE_ROOT}/eval``.
|
||||
|
||||
Metric authors writing new code: when wrapping a third-party loader
|
||||
that has its own cache convention (CLIP's ``download_root``, pyiqa's
|
||||
``cache_dir``, etc.), pass ``str(get_cache_dir() / "<library>")`` so
|
||||
users get a single ``FASTVIDEO_EVAL_CACHE`` knob to redirect them all.
|
||||
"""
|
||||
return Path(os.environ.get(
|
||||
"FASTVIDEO_EVAL_CACHE",
|
||||
os.path.join(envs.FASTVIDEO_CACHE_ROOT, "eval"),
|
||||
))
|
||||
|
||||
|
||||
def ensure_checkpoint(
|
||||
name: str,
|
||||
source: str,
|
||||
filename: str | None = None,
|
||||
) -> str:
|
||||
"""Resolve a model checkpoint path, downloading on miss.
|
||||
|
||||
See module docstring for the full source contract. *name* is used
|
||||
only as the local cache filename for URL sources; ignored otherwise.
|
||||
"""
|
||||
if os.path.exists(source):
|
||||
return source
|
||||
|
||||
if source.startswith(("http://", "https://")):
|
||||
return _ensure_url(name, source)
|
||||
|
||||
if "/" in source:
|
||||
return _ensure_hf(source, filename)
|
||||
|
||||
raise ValueError(f"Cannot resolve checkpoint: source {source!r} is neither a "
|
||||
"path, URL, nor HF repo id")
|
||||
|
||||
|
||||
def _ensure_url(name: str, url: str) -> str:
|
||||
local = get_cache_dir() / "models" / name
|
||||
if local.exists():
|
||||
return str(local)
|
||||
|
||||
local.parent.mkdir(parents=True, exist_ok=True)
|
||||
with get_lock(url):
|
||||
if local.exists(): # racing process won; reuse its result
|
||||
return str(local)
|
||||
from huggingface_hub.file_download import http_get
|
||||
tmp = local.with_suffix(local.suffix + ".tmp")
|
||||
with open(tmp, "wb") as f:
|
||||
http_get(url, f)
|
||||
tmp.rename(local)
|
||||
return str(local)
|
||||
|
||||
|
||||
def _ensure_hf(repo_id: str, filename: str | None) -> str:
|
||||
from huggingface_hub import hf_hub_download, snapshot_download
|
||||
with get_lock(f"{repo_id}/{filename or '*'}"):
|
||||
if filename:
|
||||
return hf_hub_download(repo_id=repo_id, filename=filename)
|
||||
return snapshot_download(repo_id=repo_id)
|
||||
@@ -0,0 +1,156 @@
|
||||
"""Centralized prefetcher for the Evaluator.
|
||||
|
||||
Hides video-decode latency (CPU + disk I/O) behind metric compute (GPU)
|
||||
by running a small thread pool of decoders that fill a bounded queue.
|
||||
|
||||
A single :class:`VideoPool` is owned by the Evaluator for the duration
|
||||
of one ``evaluate(samples=...)`` call. Workers consume from it via
|
||||
:meth:`VideoPool.get`; decode order is non-deterministic (whichever
|
||||
loader finishes first wins), but each yielded item carries its
|
||||
original input index so the consumer can write into a result list at
|
||||
the right slot.
|
||||
|
||||
Pool sizing: ``max_size = prefetch_factor * num_workers``. With the
|
||||
default ``prefetch_factor=2``, that's two decoded samples in flight
|
||||
per worker (one being consumed, one ready).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import queue
|
||||
import threading
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.types import Video
|
||||
|
||||
_SENTINEL = object()
|
||||
|
||||
|
||||
class VideoPool:
|
||||
"""Bounded prefetch queue feeding decoded samples to consumers.
|
||||
|
||||
Use as a context manager so loader threads are always cleaned up::
|
||||
|
||||
with VideoPool(samples, loader_threads=1, max_size=4) as pool:
|
||||
while True:
|
||||
item = pool.get()
|
||||
if item is None:
|
||||
break
|
||||
idx, decoded = item
|
||||
results[idx] = worker.evaluate(**decoded)
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
samples: list[dict],
|
||||
*,
|
||||
loader_threads: int = 1,
|
||||
max_size: int = 4,
|
||||
) -> None:
|
||||
if loader_threads < 1:
|
||||
raise ValueError("loader_threads must be >= 1")
|
||||
self._samples = samples
|
||||
self._loader_threads_n = loader_threads
|
||||
self._max_size = max(max_size, 1)
|
||||
|
||||
self._task_q: queue.Queue = queue.Queue()
|
||||
self._ready_q: queue.Queue = queue.Queue(maxsize=self._max_size)
|
||||
self._loaders: list[threading.Thread] = []
|
||||
self._stop = threading.Event()
|
||||
|
||||
self._consumed = 0
|
||||
self._consume_lock = threading.Lock()
|
||||
|
||||
self._decode_ms_total = 0.0
|
||||
self._decode_lock = threading.Lock()
|
||||
|
||||
# --- context-manager lifecycle ---
|
||||
|
||||
def __enter__(self) -> VideoPool:
|
||||
for idx, sample in enumerate(self._samples):
|
||||
self._task_q.put((idx, sample))
|
||||
# One sentinel per loader so each thread can exit cleanly.
|
||||
for _ in range(self._loader_threads_n):
|
||||
self._task_q.put(_SENTINEL)
|
||||
for _ in range(self._loader_threads_n):
|
||||
t = threading.Thread(target=self._loader_loop, daemon=True)
|
||||
t.start()
|
||||
self._loaders.append(t)
|
||||
return self
|
||||
|
||||
def __exit__(self, *_exc: Any) -> None:
|
||||
self._stop.set()
|
||||
# Drain the ready queue so any blocked-on-put loader unblocks.
|
||||
while True:
|
||||
try:
|
||||
self._ready_q.get_nowait()
|
||||
except queue.Empty:
|
||||
break
|
||||
for t in self._loaders:
|
||||
t.join(timeout=5.0)
|
||||
|
||||
# --- consumer API ---
|
||||
|
||||
def get(self, timeout: float | None = None) -> tuple[int, dict] | None:
|
||||
"""Pop the next decoded ``(idx, sample)``.
|
||||
|
||||
Returns ``None`` when all input samples have been consumed.
|
||||
Thread-safe: multiple consumer threads may share one pool.
|
||||
"""
|
||||
with self._consume_lock:
|
||||
if self._consumed >= len(self._samples):
|
||||
return None
|
||||
try:
|
||||
item = self._ready_q.get(timeout=timeout)
|
||||
except queue.Empty:
|
||||
return None
|
||||
with self._consume_lock:
|
||||
self._consumed += 1
|
||||
return item
|
||||
|
||||
@property
|
||||
def decode_ms_total(self) -> float:
|
||||
with self._decode_lock:
|
||||
return self._decode_ms_total
|
||||
|
||||
# --- loader internals ---
|
||||
|
||||
def _loader_loop(self) -> None:
|
||||
while not self._stop.is_set():
|
||||
item = self._task_q.get()
|
||||
if item is _SENTINEL:
|
||||
return
|
||||
idx, sample = item
|
||||
decoded = self._decode(sample)
|
||||
try:
|
||||
self._ready_q.put((idx, decoded), timeout=10.0)
|
||||
except queue.Full:
|
||||
# Stop set during shutdown; drop and exit.
|
||||
return
|
||||
|
||||
def _decode(self, sample: dict) -> dict:
|
||||
"""Walk a sample dict, materialize any path-shaped video values.
|
||||
|
||||
Two recognised shapes:
|
||||
- A :class:`Video` instance — populate ``.frames`` (lazy decode).
|
||||
- A bare path under ``video`` / ``reference`` — load to
|
||||
``(T, C, H, W)`` tensor (back-compat with existing callers).
|
||||
|
||||
Anything else (audio paths, scalars, dicts, tensors) passes
|
||||
through unchanged.
|
||||
"""
|
||||
from fastvideo.eval.io.video import load_video
|
||||
|
||||
t0 = time.perf_counter()
|
||||
out = dict(sample)
|
||||
for key, val in sample.items():
|
||||
if isinstance(val, Video):
|
||||
if val.frames is None and val.source is not None:
|
||||
val.frames = load_video(val.source)
|
||||
out[key] = val
|
||||
elif key in ("video", "reference") and isinstance(val, str | Path):
|
||||
out[key] = load_video(str(val))
|
||||
with self._decode_lock:
|
||||
self._decode_ms_total += (time.perf_counter() - t0) * 1000.0
|
||||
return out
|
||||
@@ -0,0 +1,97 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.util
|
||||
from typing import Any, TYPE_CHECKING
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
|
||||
_REGISTRY: dict[str, type[BaseMetric]] = {}
|
||||
|
||||
|
||||
def register(name: str):
|
||||
"""Decorator to register a metric class.
|
||||
|
||||
Usage::
|
||||
|
||||
@register("ssim")
|
||||
class SSIMMetric(BaseMetric):
|
||||
...
|
||||
"""
|
||||
|
||||
def wrapper(cls):
|
||||
_REGISTRY[name] = cls
|
||||
return cls
|
||||
|
||||
return wrapper
|
||||
|
||||
|
||||
def get_metric(name: str, **kwargs: Any) -> BaseMetric:
|
||||
"""Instantiate a registered metric by name.
|
||||
|
||||
Checks that optional dependencies are installed before instantiation
|
||||
and gives a clear install hint pointing at the right extra group.
|
||||
"""
|
||||
cls = _REGISTRY.get(name)
|
||||
if cls is None:
|
||||
available = ", ".join(sorted(_REGISTRY.keys()))
|
||||
raise KeyError(f"Unknown metric '{name}'. Available: {available}")
|
||||
|
||||
for dep in getattr(cls, "dependencies", []):
|
||||
if not importlib.util.find_spec(dep):
|
||||
raise ImportError(f"{cls.__name__} requires '{dep}'. "
|
||||
f"Install with: {_install_hint(name, dep)}")
|
||||
|
||||
return cls(**kwargs)
|
||||
|
||||
|
||||
def missing_dependencies(metric_name: str) -> list[str]:
|
||||
"""Importable module names declared by *metric_name* that are not
|
||||
actually importable in this environment. Returns ``[]`` if all deps
|
||||
are satisfied or the metric is unknown.
|
||||
|
||||
Used by group-style resolution to decide which metrics to silently
|
||||
skip (vs. naming a metric explicitly, where the missing dep should
|
||||
surface as :class:`ImportError`).
|
||||
"""
|
||||
cls = _REGISTRY.get(metric_name)
|
||||
if cls is None:
|
||||
return []
|
||||
return [d for d in getattr(cls, "dependencies", []) if not importlib.util.find_spec(d)]
|
||||
|
||||
|
||||
def _install_hint(metric_name: str, dep: str) -> str:
|
||||
"""Copy-pastable install command that actually satisfies *dep*.
|
||||
|
||||
Most deps are covered by a single `[extra]`. ``detectron2`` is a
|
||||
special case: it builds C++ kernels against the user's torch and
|
||||
isn't on PyPI cleanly, so the recipe needs the base extra *plus* a
|
||||
git+ install with build-isolation off.
|
||||
"""
|
||||
if dep == "detectron2":
|
||||
return ("uv pip install 'fastvideo[eval-vbench]' && "
|
||||
"uv pip install --no-build-isolation "
|
||||
"'git+https://github.com/facebookresearch/detectron2.git'")
|
||||
return f"uv pip install 'fastvideo[{_extra_for(metric_name)}]'"
|
||||
|
||||
|
||||
def _extra_for(metric_name: str) -> str:
|
||||
"""Map a metric name to the smallest extra that satisfies its deps."""
|
||||
if metric_name.startswith("vbench."):
|
||||
return "eval-vbench"
|
||||
if metric_name.startswith("physics_iq"):
|
||||
return "eval-physics-iq"
|
||||
return "eval"
|
||||
|
||||
|
||||
def list_metrics() -> list[str]:
|
||||
"""Return sorted list of all registered metric names."""
|
||||
return sorted(_REGISTRY.keys())
|
||||
|
||||
|
||||
def resolve_group(name: str) -> list[str] | None:
|
||||
"""If *name* is a group prefix (e.g. ``"vbench"``), return all matching
|
||||
metric names. Returns ``None`` if *name* is not a group."""
|
||||
prefix = name + "."
|
||||
matches = sorted(k for k in _REGISTRY if k.startswith(prefix))
|
||||
return matches if matches else None
|
||||
@@ -0,0 +1,72 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
@dataclass
|
||||
class MetricResult:
|
||||
"""Standard result container returned by all metrics.
|
||||
|
||||
``score`` is ``None`` when the metric was skipped (e.g. missing
|
||||
required input). Check ``details["skipped"]`` for the reason.
|
||||
"""
|
||||
name: str
|
||||
score: float | None
|
||||
details: dict[str, Any] = field(default_factory=dict)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Video:
|
||||
"""A media handle bundling frames + audio behind one typed value.
|
||||
|
||||
The name matches the package; in practice this also accepts pure
|
||||
audio inputs (``.wav`` / ``.mp3``) — for those, ``frames`` stays
|
||||
``None`` and only ``audio`` is populated. Conversely, a silent
|
||||
video has ``audio=None``. This intentionally mirrors how a
|
||||
container format (mp4) carries optional streams.
|
||||
|
||||
The ``VideoPool`` (a thin async prefetcher in front of the
|
||||
Evaluator) walks the per-sample kwargs dict, finds every ``Video``
|
||||
value, and triggers the decodes asked for by the registered
|
||||
metrics — so by the time a metric's ``compute(sample)`` runs, the
|
||||
``frames`` / ``audio`` attributes are already populated. Metric
|
||||
code reads them directly:
|
||||
|
||||
def compute(self, sample):
|
||||
video = sample["video"]
|
||||
frames = video.frames # (T, C, H, W) in [0, 1]
|
||||
audio = video.audio # 1D float32
|
||||
|
||||
Constructors accepted at the user boundary::
|
||||
|
||||
Video("clip.mp4") # mp4 / wav / image-dir / etc.
|
||||
Video(tensor) # pre-decoded (T, C, H, W) tensor
|
||||
Video(list_of_PIL_images) # frame list
|
||||
|
||||
A folder of clips is **not** a single ``Video`` — set-vs-set
|
||||
inputs are a list of ``Video`` objects (one per clip), wired by
|
||||
the user (e.g. via ``Evaluator.evaluate(samples=[...])``).
|
||||
"""
|
||||
|
||||
source: Any # VideoSource at runtime; typed broadly to avoid a torch import here
|
||||
fps: float | None = None
|
||||
|
||||
# Decoded streams. ``None`` until the pool (or the user) calls a
|
||||
# decode helper. Both can be ``None`` simultaneously — that's the
|
||||
# state of a fresh handle whose source is a path.
|
||||
frames: Any = None # torch.Tensor | None
|
||||
audio: Any = None # np.ndarray | None
|
||||
audio_sr: int | None = None
|
||||
|
||||
def has_frames(self) -> bool:
|
||||
return self.frames is not None
|
||||
|
||||
def has_audio(self) -> bool:
|
||||
return self.audio is not None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
# Coerce Path → str so downstream loaders see a single shape.
|
||||
if isinstance(self.source, Path):
|
||||
self.source = str(self.source)
|
||||
@@ -0,0 +1,187 @@
|
||||
"""EvalWorker: a single-GPU bag of metric replicas.
|
||||
|
||||
One ``EvalWorker`` per GPU. Holds an instance of every requested metric on
|
||||
its own device, scores one sample at a time. Stateless across calls — the
|
||||
worker never sees more than one ``evaluate(...)`` invocation at a time
|
||||
(the parent :class:`Evaluator` ensures this by handing the worker to one
|
||||
thread at a time).
|
||||
|
||||
This mirrors FastVideo's ``Worker`` layer under :class:`VideoGenerator`,
|
||||
but in-process (threads, not processes) — eval metrics are independent
|
||||
and need no NCCL / TP / SP / distributed init, so process isolation is
|
||||
unnecessary overhead.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.eval.memory import clear_cache
|
||||
from fastvideo.eval.registry import get_metric
|
||||
from fastvideo.eval.types import MetricResult
|
||||
import contextlib
|
||||
|
||||
|
||||
class EvalWorker:
|
||||
"""Owns metric replicas on one device. Single-GPU, single-sample."""
|
||||
|
||||
def __init__(self, metric_names: list[str], device: str, *, compile: bool = False, pre_upload: bool = True) -> None:
|
||||
self._names = list(metric_names)
|
||||
self._device = device
|
||||
self._compile = compile
|
||||
# Pre-upload ``video``/``reference`` to the worker's device once
|
||||
# so every metric for the same sample shares a single GPU-resident
|
||||
# tensor. Set ``pre_upload=False`` for training-time eval contexts
|
||||
# where holding the input tensor on GPU across the metric loop
|
||||
# competes with the training-step working set; in that case each
|
||||
# metric uploads its own copy as before.
|
||||
self._pre_upload = pre_upload
|
||||
self._metrics: dict = {}
|
||||
self._unloaded = False
|
||||
self._load()
|
||||
|
||||
@property
|
||||
def device(self) -> str:
|
||||
return self._device
|
||||
|
||||
@property
|
||||
def metric_names(self) -> list[str]:
|
||||
"""Names of the metrics this worker owns, in load order."""
|
||||
return list(self._metrics.keys())
|
||||
|
||||
def _load(self) -> None:
|
||||
for name in self._names:
|
||||
m = get_metric(name)
|
||||
m.to(self._device)
|
||||
m.setup()
|
||||
if self._compile and getattr(m, "_model", None) is not None:
|
||||
m._model = torch.compile(m._model)
|
||||
self._metrics[name] = m
|
||||
self._unloaded = False
|
||||
|
||||
def evaluate(self, **kwargs) -> dict[str, MetricResult]:
|
||||
"""Score one sample.
|
||||
|
||||
``video`` may be a ``(T, C, H, W)`` tensor or a path-like
|
||||
(``str`` / ``Path``) — paths are loaded inside this method so
|
||||
the dispatcher can hold a queue of cheap path strings instead
|
||||
of fully-decoded tensors. ``reference`` follows the same rule.
|
||||
|
||||
A ``(1, T, C, H, W)`` tensor is also accepted for back-compat
|
||||
and gets unwrapped to ``(T, C, H, W)`` before reaching metrics.
|
||||
|
||||
Stage timings (``decode_ms``, ``compute_ms``) are accumulated on
|
||||
thread-local counters and zeroed on each call. Read via
|
||||
:func:`pop_timings` immediately after the call.
|
||||
"""
|
||||
if self._unloaded:
|
||||
raise RuntimeError("EvalWorker was unloaded; call reload() before evaluating.")
|
||||
|
||||
import time
|
||||
sample = dict(kwargs)
|
||||
t0 = time.perf_counter()
|
||||
sample["video"] = _resolve_video_input(sample.get("video"))
|
||||
if "reference" in sample:
|
||||
sample["reference"] = _resolve_video_input(sample["reference"])
|
||||
if self._pre_upload:
|
||||
sample["video"] = _to_device(sample.get("video"), self._device)
|
||||
if "reference" in sample:
|
||||
sample["reference"] = _to_device(sample["reference"], self._device)
|
||||
t1 = time.perf_counter()
|
||||
|
||||
results: dict[str, MetricResult] = {}
|
||||
for name, m in self._metrics.items():
|
||||
results[name] = m.compute(sample)
|
||||
t2 = time.perf_counter()
|
||||
|
||||
_record_timing(decode_ms=(t1 - t0) * 1000.0, compute_ms=(t2 - t1) * 1000.0)
|
||||
return results
|
||||
|
||||
def release_cuda_memory(self) -> None:
|
||||
"""Free CUDA caches without dropping models."""
|
||||
clear_cache()
|
||||
if torch.cuda.is_available():
|
||||
with contextlib.suppress(Exception):
|
||||
torch.cuda.ipc_collect()
|
||||
|
||||
def unload(self) -> None:
|
||||
"""Drop metric refs so models become GC-able. Reverse with reload()."""
|
||||
self._metrics = {}
|
||||
self._unloaded = True
|
||||
self.release_cuda_memory()
|
||||
|
||||
def reload(self) -> None:
|
||||
"""Rebuild metrics dropped by :meth:`unload`."""
|
||||
if self._unloaded:
|
||||
self._load()
|
||||
|
||||
|
||||
_TIMINGS: dict[str, float] = {"decode_ms": 0.0, "compute_ms": 0.0, "n": 0}
|
||||
|
||||
|
||||
def _record_timing(*, decode_ms: float, compute_ms: float) -> None:
|
||||
"""Accumulate per-call decode and compute time into a process-global
|
||||
counter. Read + zeroed via :func:`pop_timings`. Used by
|
||||
``scripts/eval/bench_pipeline.py``.
|
||||
"""
|
||||
_TIMINGS["decode_ms"] += decode_ms
|
||||
_TIMINGS["compute_ms"] += compute_ms
|
||||
_TIMINGS["n"] += 1
|
||||
|
||||
|
||||
def add_pool_decode_ms(ms: float) -> None:
|
||||
"""Attribute pool-side decode time to the same global counter.
|
||||
|
||||
Pool threads call this when they finish decoding a sample. Keeps
|
||||
``pop_timings()`` returning total decode time regardless of where
|
||||
the decode physically happened.
|
||||
"""
|
||||
_TIMINGS["decode_ms"] += ms
|
||||
|
||||
|
||||
def pop_timings() -> dict[str, float]:
|
||||
"""Snapshot and zero the timing counters."""
|
||||
snapshot = dict(_TIMINGS)
|
||||
_TIMINGS["decode_ms"] = 0.0
|
||||
_TIMINGS["compute_ms"] = 0.0
|
||||
_TIMINGS["n"] = 0
|
||||
return snapshot
|
||||
|
||||
|
||||
def _to_device(value: Any, device: str | torch.device) -> Any:
|
||||
"""Move a video tensor to *device* if it isn't already there.
|
||||
|
||||
No-op for ``None`` or non-tensor values. Uses ``non_blocking=True``
|
||||
so the host-side memcpy and device DMA can overlap with subsequent
|
||||
CPU work (pinning is opportunistic; PyTorch will pin internally for
|
||||
pageable tensors).
|
||||
"""
|
||||
if value is None or not isinstance(value, torch.Tensor):
|
||||
return value
|
||||
target = torch.device(device)
|
||||
if value.device == target:
|
||||
return value
|
||||
return value.to(target, non_blocking=True)
|
||||
|
||||
|
||||
def _resolve_video_input(value: Any) -> Any:
|
||||
"""Normalize a sample's ``video`` / ``reference`` field for metrics.
|
||||
|
||||
* ``str`` / ``Path`` → decoded ``(T, C, H, W)`` tensor via
|
||||
:func:`fastvideo.eval.io.video.load_video`. Decoding happens in
|
||||
the worker thread so the dispatcher can keep paths queued
|
||||
instead of full tensors.
|
||||
* ``(1, T, C, H, W)`` tensor → squeezed to ``(T, C, H, W)``
|
||||
(back-compat with callers that still pass the leading batch dim).
|
||||
* anything else → returned untouched.
|
||||
"""
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, str | Path):
|
||||
from fastvideo.eval.io.video import load_video
|
||||
return load_video(str(value))
|
||||
if isinstance(value, torch.Tensor) and value.dim() == 5 and value.shape[0] == 1:
|
||||
return value.squeeze(0)
|
||||
return value
|
||||
@@ -0,0 +1,66 @@
|
||||
"""Smoke tests for the VBench prompt dataset and dataset registry."""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from fastvideo.eval.datasets import (PromptDataset, VBenchPromptDataset,
|
||||
get_dataset, list_datasets)
|
||||
|
||||
|
||||
def test_registry_contains_vbench():
|
||||
assert "vbench" in list_datasets()
|
||||
|
||||
|
||||
def test_get_dataset_returns_typed_instance():
|
||||
ds = get_dataset("vbench", dimensions=["color"])
|
||||
assert isinstance(ds, PromptDataset)
|
||||
assert isinstance(ds, VBenchPromptDataset)
|
||||
assert ds.name == "vbench"
|
||||
assert ds.supports_dimensions is True
|
||||
|
||||
|
||||
def test_full_corpus_size():
|
||||
ds = VBenchPromptDataset()
|
||||
assert len(ds) == 946
|
||||
assert len(ds.dimensions) == 16
|
||||
|
||||
|
||||
def test_rows_are_dicts_with_required_keys():
|
||||
ds = VBenchPromptDataset(dimensions=["color"])
|
||||
sample = ds[0]
|
||||
assert isinstance(sample, dict)
|
||||
assert "prompt" in sample
|
||||
assert "n_samples" in sample
|
||||
assert "dimensions" in sample
|
||||
|
||||
|
||||
def test_temporal_flickering_n_samples():
|
||||
ds = VBenchPromptDataset(dimensions=["temporal_flickering"])
|
||||
assert all(s["n_samples"] == 25 for s in ds)
|
||||
|
||||
|
||||
def test_default_n_samples():
|
||||
ds = VBenchPromptDataset(dimensions=["subject_consistency"])
|
||||
assert all(s["n_samples"] == 5 for s in ds)
|
||||
|
||||
|
||||
def test_color_aux_info_is_flat():
|
||||
ds = VBenchPromptDataset(dimensions=["color"])
|
||||
sample = ds[0]
|
||||
aux = sample["auxiliary_info"]
|
||||
# Flat: {"color": "<color name>"}, not nested under the dimension.
|
||||
assert "color" in aux
|
||||
assert isinstance(aux["color"], str)
|
||||
|
||||
|
||||
def test_unknown_dimension_raises():
|
||||
with pytest.raises(ValueError, match="Unknown VBench dimensions"):
|
||||
VBenchPromptDataset(dimensions=["bogus"])
|
||||
|
||||
|
||||
def test_by_dimension_groups_correctly():
|
||||
ds = VBenchPromptDataset(dimensions=["subject_consistency", "color"])
|
||||
groups = ds.by_dimension()
|
||||
assert set(groups) == {"subject_consistency", "color"}
|
||||
assert all(isinstance(s["auxiliary_info"].get("color"), str)
|
||||
for s in groups["color"])
|
||||
@@ -0,0 +1,107 @@
|
||||
"""Multi-replica eval through the public ``Evaluator`` API.
|
||||
|
||||
Skipped automatically when fewer than 2 CUDA devices are visible.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
import torch
|
||||
|
||||
from fastvideo.eval import create_evaluator
|
||||
|
||||
|
||||
pytestmark = pytest.mark.skipif(
|
||||
not torch.cuda.is_available() or torch.cuda.device_count() < 2,
|
||||
reason="multi-GPU evaluator tests require at least 2 visible CUDA devices",
|
||||
)
|
||||
|
||||
|
||||
_T, _C, _H, _W = 6, 3, 32, 32
|
||||
|
||||
|
||||
def _make_samples(n: int) -> list[dict]:
|
||||
torch.manual_seed(7)
|
||||
samples: list[dict] = []
|
||||
for i in range(n):
|
||||
gen = torch.rand(_T, _C, _H, _W)
|
||||
# Slightly perturb each row's reference so scores vary by index.
|
||||
ref = gen + 0.01 * (i + 1) * torch.rand_like(gen)
|
||||
samples.append({"video": gen, "reference": ref})
|
||||
return samples
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def baseline_scores():
|
||||
"""Reference scores computed on a single-GPU evaluator. The multi-GPU
|
||||
runs must reproduce these exactly when handed the same input list —
|
||||
that's the only way to verify round-robin dispatch isn't dropping or
|
||||
reordering samples."""
|
||||
samples = _make_samples(8)
|
||||
ev = create_evaluator(
|
||||
metrics=["common.psnr", "common.ssim"],
|
||||
device="cuda:0",
|
||||
num_gpus=1,
|
||||
)
|
||||
try:
|
||||
out = ev.evaluate(samples=samples)
|
||||
finally:
|
||||
ev.shutdown()
|
||||
return samples, out
|
||||
|
||||
|
||||
def test_multi_gpu_evaluator_reports_two_workers():
|
||||
ev = create_evaluator(metrics=["common.psnr"], num_gpus=2)
|
||||
try:
|
||||
assert ev.num_gpus == 2
|
||||
finally:
|
||||
ev.shutdown()
|
||||
|
||||
|
||||
def test_multi_gpu_dispatch_preserves_order_and_scores(baseline_scores):
|
||||
"""Same samples, multi-GPU dispatch — results must match the single-GPU
|
||||
baseline element-for-element. This verifies (a) the round-robin doesn't
|
||||
reorder, (b) every sample is scored exactly once, (c) the workers
|
||||
don't share mutable state."""
|
||||
samples, expected = baseline_scores
|
||||
|
||||
ev = create_evaluator(
|
||||
metrics=["common.psnr", "common.ssim"],
|
||||
num_gpus=2,
|
||||
)
|
||||
try:
|
||||
got = ev.evaluate(samples=samples)
|
||||
finally:
|
||||
ev.shutdown()
|
||||
|
||||
assert len(got) == len(expected)
|
||||
for i, (g, e) in enumerate(zip(got, expected)):
|
||||
assert set(g.keys()) == {"common.psnr", "common.ssim"}, f"row {i}"
|
||||
assert g["common.psnr"].score == pytest.approx(e["common.psnr"].score), \
|
||||
f"row {i} psnr drift"
|
||||
assert g["common.ssim"].score == pytest.approx(e["common.ssim"].score), \
|
||||
f"row {i} ssim drift"
|
||||
|
||||
|
||||
def test_multi_gpu_evaluator_kwargs_form_runs_on_one_replica():
|
||||
"""The kwargs form (single sample) is documented to always hit worker
|
||||
0; this test pins the contract so future refactors don't accidentally
|
||||
fan out a single call."""
|
||||
ev = create_evaluator(metrics=["common.psnr"], num_gpus=2)
|
||||
try:
|
||||
torch.manual_seed(0)
|
||||
gen = torch.rand(_T, _C, _H, _W)
|
||||
out = ev.evaluate(video=gen, reference=gen)
|
||||
assert out["common.psnr"].score > 50.0 # PSNR(x, x) is huge
|
||||
finally:
|
||||
ev.shutdown()
|
||||
|
||||
|
||||
def test_multi_gpu_release_cuda_memory_runs_clean():
|
||||
"""``release_cuda_memory`` must hit every replica without crashing."""
|
||||
ev = create_evaluator(metrics=["common.psnr"], num_gpus=2)
|
||||
try:
|
||||
samples = _make_samples(2)
|
||||
_ = ev.evaluate(samples=samples)
|
||||
ev.release_cuda_memory() # should not raise
|
||||
finally:
|
||||
ev.shutdown()
|
||||
@@ -0,0 +1,158 @@
|
||||
"""Path-input variants of the public Evaluator API.
|
||||
|
||||
The worker boundary accepts ``video`` / ``reference`` as either a
|
||||
pre-loaded ``(T, C, H, W)`` tensor or a path-like (``str`` / ``Path``).
|
||||
These tests pin the path-form so future refactors don't accidentally
|
||||
re-require pre-loaded tensors.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import pytest
|
||||
import torch
|
||||
|
||||
from fastvideo.eval import MetricResult, create_evaluator, evaluate
|
||||
|
||||
_T, _C, _H, _W = 6, 3, 32, 32
|
||||
|
||||
|
||||
def _write_tensor_as_mp4(tensor: torch.Tensor, path: Path) -> None:
|
||||
"""Write a (T, C, H, W) float [0, 1] tensor to *path* as an mp4."""
|
||||
frames = (tensor * 255).clamp(0, 255).to(torch.uint8).permute(0, 2, 3, 1).cpu().numpy()
|
||||
# cv2 expects BGR; the tests don't care about colour fidelity (they
|
||||
# just need the bytes to round-trip), but flipping keeps the
|
||||
# written file faithful to what an mp4 would carry on disk.
|
||||
frames = frames[..., ::-1]
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
h, w = frames.shape[1], frames.shape[2]
|
||||
writer = cv2.VideoWriter(str(path), cv2.VideoWriter_fourcc(*"mp4v"), 8, (w, h))
|
||||
if not writer.isOpened():
|
||||
pytest.skip("cv2.VideoWriter could not open the mp4v codec on this host")
|
||||
for f in frames:
|
||||
writer.write(np.ascontiguousarray(f))
|
||||
writer.release()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def video_paths(tmp_path):
|
||||
"""Two reproducible mp4s on disk + their pre-loaded tensors for parity."""
|
||||
torch.manual_seed(0)
|
||||
paths: list[Path] = []
|
||||
tensors: list[torch.Tensor] = []
|
||||
for i in range(2):
|
||||
t = torch.rand(_T, _C, _H, _W)
|
||||
p = tmp_path / f"clip_{i}.mp4"
|
||||
_write_tensor_as_mp4(t, p)
|
||||
paths.append(p)
|
||||
tensors.append(t)
|
||||
return paths, tensors
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def evaluator():
|
||||
ev = create_evaluator(metrics=["common.psnr", "common.ssim"], device="cpu")
|
||||
yield ev
|
||||
ev.shutdown()
|
||||
|
||||
|
||||
def test_kwargs_form_accepts_string_path(evaluator, video_paths):
|
||||
paths, _ = video_paths
|
||||
out = evaluator.evaluate(video=str(paths[0]), reference=str(paths[0]))
|
||||
assert isinstance(out, dict)
|
||||
assert isinstance(out["common.psnr"], MetricResult)
|
||||
# Self-paired video → PSNR is huge, SSIM = 1.
|
||||
assert out["common.psnr"].score > 50.0
|
||||
assert out["common.ssim"].score == pytest.approx(1.0, abs=1e-5)
|
||||
|
||||
|
||||
def test_kwargs_form_accepts_pathlib_path(evaluator, video_paths):
|
||||
paths, _ = video_paths
|
||||
out = evaluator.evaluate(video=paths[0], reference=paths[0])
|
||||
assert out["common.psnr"].score > 50.0
|
||||
|
||||
|
||||
def test_samples_list_accepts_paths(evaluator, video_paths):
|
||||
paths, _ = video_paths
|
||||
samples = [{"video": str(p), "reference": str(p)} for p in paths]
|
||||
out = evaluator.evaluate(samples=samples)
|
||||
assert isinstance(out, list)
|
||||
assert len(out) == len(paths)
|
||||
for row in out:
|
||||
assert row["common.psnr"].score > 50.0
|
||||
|
||||
|
||||
def test_samples_list_can_mix_paths_and_tensors(evaluator, video_paths):
|
||||
"""A single ``samples`` call can mix path and tensor entries."""
|
||||
paths, tensors = video_paths
|
||||
samples = [
|
||||
{"video": str(paths[0]), "reference": tensors[0]}, # path + tensor
|
||||
{"video": tensors[1], "reference": str(paths[1])}, # tensor + path
|
||||
]
|
||||
out = evaluator.evaluate(samples=samples)
|
||||
assert len(out) == 2
|
||||
for row in out:
|
||||
assert row["common.psnr"].score > 0.0
|
||||
|
||||
|
||||
def test_path_form_score_matches_tensor_form(evaluator, video_paths):
|
||||
"""Loading via path must produce the same score as loading via the
|
||||
public ``load_video`` helper and passing the tensor in directly."""
|
||||
from fastvideo.eval.io import load_video
|
||||
|
||||
paths, _ = video_paths
|
||||
via_path = evaluator.evaluate(video=str(paths[0]), reference=str(paths[0]))
|
||||
tensor = load_video(str(paths[0]))
|
||||
via_tensor = evaluator.evaluate(video=tensor, reference=tensor)
|
||||
assert via_path["common.psnr"].score == pytest.approx(
|
||||
via_tensor["common.psnr"].score, abs=1e-4)
|
||||
assert via_path["common.ssim"].score == pytest.approx(
|
||||
via_tensor["common.ssim"].score, abs=1e-4)
|
||||
|
||||
|
||||
def test_one_shot_evaluate_accepts_paths(video_paths):
|
||||
"""The top-level ``fastvideo.eval.evaluate`` helper also flows paths."""
|
||||
paths, _ = video_paths
|
||||
out = evaluate(generated=str(paths[0]), reference=str(paths[0]),
|
||||
metrics=["common.psnr"], device="cpu")
|
||||
assert out["common.psnr"].score > 50.0
|
||||
|
||||
|
||||
def test_missing_path_surfaces_as_exception(evaluator, tmp_path):
|
||||
"""Decode failures must propagate, not silently produce a None score."""
|
||||
bogus = tmp_path / "does_not_exist.mp4"
|
||||
with pytest.raises(Exception):
|
||||
evaluator.evaluate(video=str(bogus), reference=str(bogus))
|
||||
|
||||
|
||||
def test_dispatcher_holds_paths_not_tensors_in_queue(tmp_path):
|
||||
"""Memory invariant: when many paths are passed, the queued samples
|
||||
are tiny strings, not full tensors. Verify by checking the length
|
||||
of the per-sample reference set the dispatcher materializes."""
|
||||
torch.manual_seed(1)
|
||||
n = 8
|
||||
paths: list[Path] = []
|
||||
for i in range(n):
|
||||
t = torch.rand(_T, _C, _H, _W)
|
||||
p = tmp_path / f"clip_{i}.mp4"
|
||||
_write_tensor_as_mp4(t, p)
|
||||
paths.append(p)
|
||||
|
||||
samples = [{"video": str(p), "reference": str(p)} for p in paths]
|
||||
# Each sample dict is just two strings — no tensor allocations until
|
||||
# the worker's _resolve_video_input runs.
|
||||
for s in samples:
|
||||
assert isinstance(s["video"], str)
|
||||
assert isinstance(s["reference"], str)
|
||||
|
||||
ev = create_evaluator(metrics=["common.psnr"], device="cpu")
|
||||
try:
|
||||
out = ev.evaluate(samples=samples)
|
||||
finally:
|
||||
ev.shutdown()
|
||||
assert len(out) == n
|
||||
# Self-paired ⇒ all PSNRs should be very high.
|
||||
for row in out:
|
||||
assert row["common.psnr"].score > 50.0
|
||||
@@ -0,0 +1,136 @@
|
||||
"""End-to-end tests for single-replica eval through the public API.
|
||||
|
||||
Runs the lightweight pixel-space metrics — ``common.psnr`` and
|
||||
``common.ssim`` — under both shapes that real callers use:
|
||||
|
||||
* one-shot ``evaluate(video=..., reference=...)`` (the helper in
|
||||
``fastvideo.eval.api``);
|
||||
* a long-lived ``Evaluator``, called once per sample;
|
||||
* a long-lived ``Evaluator``, called with a list of sample dicts to
|
||||
fan out (``samples=[...]``).
|
||||
|
||||
GPU-only metrics live in separate test modules / classes; everything
|
||||
here runs on CPU so the suite stays cheap to invoke.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
import torch
|
||||
|
||||
from fastvideo.eval import MetricResult, create_evaluator, evaluate
|
||||
|
||||
|
||||
# Tight resolution + frame count so CPU SSIM stays under a couple of seconds.
|
||||
_T, _C, _H, _W = 6, 3, 32, 32
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def gen_ref():
|
||||
"""Reproducible (gen, ref) pair shaped (T, C, H, W)."""
|
||||
torch.manual_seed(0)
|
||||
gen = torch.rand(_T, _C, _H, _W)
|
||||
ref = torch.rand(_T, _C, _H, _W)
|
||||
return gen, ref
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def evaluator():
|
||||
ev = create_evaluator(
|
||||
metrics=["common.psnr", "common.ssim"],
|
||||
device="cpu",
|
||||
)
|
||||
yield ev
|
||||
ev.shutdown()
|
||||
|
||||
|
||||
def _assert_well_formed(result: MetricResult, name: str) -> None:
|
||||
assert isinstance(result, MetricResult)
|
||||
assert result.name == name
|
||||
assert result.score is not None
|
||||
assert isinstance(result.score, float)
|
||||
# PSNR / SSIM both populate per-frame details.
|
||||
assert "per_frame" in result.details
|
||||
assert len(result.details["per_frame"]) == _T
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# One-shot helper
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_evaluate_one_shot_returns_dict_of_metric_results(gen_ref):
|
||||
gen, ref = gen_ref
|
||||
out = evaluate(generated=gen, reference=ref,
|
||||
metrics=["common.psnr", "common.ssim"], device="cpu")
|
||||
assert isinstance(out, dict)
|
||||
assert set(out.keys()) == {"common.psnr", "common.ssim"}
|
||||
_assert_well_formed(out["common.psnr"], "common.psnr")
|
||||
_assert_well_formed(out["common.ssim"], "common.ssim")
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Long-lived Evaluator, single-sample form
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_evaluator_single_sample_returns_dict(evaluator, gen_ref):
|
||||
gen, ref = gen_ref
|
||||
out = evaluator.evaluate(video=gen, reference=ref)
|
||||
assert isinstance(out, dict)
|
||||
assert set(out.keys()) == {"common.psnr", "common.ssim"}
|
||||
for name, mr in out.items():
|
||||
_assert_well_formed(mr, name)
|
||||
|
||||
|
||||
def test_evaluator_accepts_legacy_5d_input(evaluator, gen_ref):
|
||||
"""Callers that still pass ``(1, T, C, H, W)`` should get unwrapped."""
|
||||
gen, ref = gen_ref
|
||||
out = evaluator.evaluate(video=gen.unsqueeze(0), reference=ref.unsqueeze(0))
|
||||
_assert_well_formed(out["common.psnr"], "common.psnr")
|
||||
|
||||
|
||||
def test_evaluator_score_is_deterministic(evaluator, gen_ref):
|
||||
gen, ref = gen_ref
|
||||
a = evaluator.evaluate(video=gen, reference=ref)
|
||||
b = evaluator.evaluate(video=gen.clone(), reference=ref.clone())
|
||||
assert a["common.psnr"].score == pytest.approx(b["common.psnr"].score)
|
||||
assert a["common.ssim"].score == pytest.approx(b["common.ssim"].score)
|
||||
|
||||
|
||||
def test_evaluator_psnr_identical_videos_is_high(evaluator, gen_ref):
|
||||
"""PSNR(x, x) is unbounded above; with our clamp it caps near 100 dB."""
|
||||
gen, _ = gen_ref
|
||||
out = evaluator.evaluate(video=gen, reference=gen)
|
||||
assert out["common.psnr"].score > 50.0
|
||||
assert out["common.ssim"].score == pytest.approx(1.0, abs=1e-5)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Long-lived Evaluator, list (fan-out) form
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_evaluator_samples_list_preserves_input_order(evaluator):
|
||||
"""When ``samples=[...]`` is passed, results must come back per sample."""
|
||||
torch.manual_seed(1)
|
||||
samples = []
|
||||
for i in range(4):
|
||||
# Vary the reference enough that scores differ across rows.
|
||||
gen = torch.rand(_T, _C, _H, _W)
|
||||
ref = gen + 0.01 * (i + 1) * torch.rand_like(gen)
|
||||
samples.append({"video": gen, "reference": ref})
|
||||
|
||||
out = evaluator.evaluate(samples=samples)
|
||||
assert isinstance(out, list)
|
||||
assert len(out) == len(samples)
|
||||
for row in out:
|
||||
assert set(row.keys()) == {"common.psnr", "common.ssim"}
|
||||
_assert_well_formed(row["common.psnr"], "common.psnr")
|
||||
|
||||
# Re-running should give bit-identical scores (no nondeterministic
|
||||
# scheduling effects under single-GPU dispatch).
|
||||
out2 = evaluator.evaluate(samples=samples)
|
||||
for a, b in zip(out, out2):
|
||||
assert a["common.psnr"].score == pytest.approx(b["common.psnr"].score)
|
||||
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user