Compare commits

...
Author SHA1 Message Date
SolitaryThinker e54106110e [feat]: Activate MagiHuman pipeline (registry + examples + SSIM regression)
The activation switch. After this PR merges, MagiHuman is publicly
loadable via:

    from fastvideo import VideoGenerator
    gen = VideoGenerator.from_pretrained('FastVideo/MagiHuman-Diffusers/base')

Files:
- fastvideo/registry.py: 153 lines of register_configs/register_presets
  calls covering all 4 variants x 2 modes (8 total entrypoints).
- examples/inference/basic/basic_magi_human{_,_ti2v,_distill,_distill_ti2v,
  _sr540p,_sr540p_ti2v,_sr1080p,_sr1080p_ti2v}.py: 8 user-facing scripts.
- fastvideo/tests/ssim/test_magi_human_similarity.py: CI-eligible SSIM
  regression test against the umbrella HF repo.
- .agents/memory/codebase-map/models/magi_human.md: codebase-map entry
  (first per-model entry under the models/ subdir; sets the convention).
- fastvideo/pipelines/basic/magi_human/AGENTS.md: provenance section
  finalized with all 8 PR numbers + the will/magi source SHA.

Verification:
- Existing tests on main: should still pass (no shared component changes
  in this PR; the umbrella loader infra landed in PR-B / loader-infra).
- pre-commit run --all-files: clean.
- E2E smoke after activation: any of the 8 examples emits expected mp4
  with hash dcf5f2bf6534c7c0d91e7353e42b23db on the base T2V variant.

Magi-Stack: 8/8
2026-07-13 03:18:28 -07:00
12 changed files with 668 additions and 5 deletions
@@ -0,0 +1,77 @@
# MagiHuman — Codebase Map Entry
**Family:** daVinci-MagiHuman (joint audio-visual generative model)
**Reference:** [GAIR-NLP/daVinci-MagiHuman](https://github.com/GAIR-NLP/daVinci-MagiHuman)
**Architecture:** 15B-param single-stream DiT, 40 layers, hidden=5120,
head_dim=128, GQA num_query_groups=8. Joint AV denoising in a unified token
sequence; **no cross-attention**.
## Variant Matrix
| Variant | T2V | TI2V | DiT | Steps | CFG | Resolution |
|---|---|---|---|---|---|---|
| `base` | yes | yes | base | 32 | 2 | 480x256 |
| `distill` | yes | yes | distill (DMD-2) | 8 | 1 (no CFG) | 480x256 |
| `sr_540p` | yes | yes | base + sr_540p | 32 + 5 | 2 + cfg-trick | 896x512 |
| `sr_1080p` | yes | yes | base + sr_1080p | 32 + 5 | 2 + cfg-trick | 1920x1056 |
SR-1080p uses block-sparse video→video local-window attention on 32 of 40
SR DiT layers (`frame_receptive_field=11`), implemented as a 3-block SDPA
accumulator that mirrors upstream `flex_flash_attn_func`.
## File Locations
| Role | Path |
|---|---|
| Pipeline class | `fastvideo/pipelines/basic/magi_human/magi_human_pipeline.py` |
| Pipeline package AGENTS.md | `fastvideo/pipelines/basic/magi_human/AGENTS.md` |
| Stages | `fastvideo/pipelines/basic/magi_human/stages/*.py` |
| DiT | `fastvideo/models/dits/magi_human.py` |
| Text encoder (T5-Gemma) | `fastvideo/models/encoders/t5gemma.py` |
| Audio VAE wrapper | `fastvideo/models/vaes/sa_audio.py` (shared with `stable_audio` pipeline) |
| Conversion script | `scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py` |
| Examples | `examples/inference/basic/basic_magi_human*.py` (8 files, one per variant × mode) |
| SSIM regression | `fastvideo/tests/ssim/test_magi_human_similarity.py` |
| Local parity battery | `tests/local_tests/magi_human/` (14 tests, GPU-gated) |
| Port journal | `fastvideo/pipelines/basic/magi_human/JOURNAL.md` |
## Canonical HF Repo
[FastVideo/MagiHuman-Diffusers](https://huggingface.co/FastVideo/MagiHuman-Diffusers)
— umbrella repo with sibling subfolders per variant.
```python
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained("FastVideo/MagiHuman-Diffusers/base")
gen.generate_video(prompt="...", output_path="out.mp4", save_video=True)
```
Four shared upstream components are lazy-loaded by `MagiHumanPipeline.load_modules`:
| Component | Upstream repo |
|---|---|
| Wan 2.2 VAE | `Wan-AI/Wan2.2-TI2V-5B` |
| T5-Gemma 9B UL2 | `google/t5gemma-9b-9b-ul2` |
| Stable Audio VAE | `stabilityai/stable-audio-open-1.0` |
| MagiHuman DiT weights | `GAIR/daVinci-MagiHuman` (gated) |
## Parity Invariants
Three load-bearing invariants. See `fastvideo/pipelines/basic/magi_human/AGENTS.md`
for the full discussion.
1. **Channel-major video token packing** (`stages/latent_preparation.py`)
2. **DiT dtype boundary**: residual stream stays fp32 across blocks
3. **Conversion `_FP32_KEEP_SUFFIXES`** (`scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py`)
## Lessons
- `.agents/lessons/2026-05-07_silent-channel-major-packing-bugs.md`
- `.agents/lessons/2026-05-07_dit-dtype-boundary-with-flash-attn.md`
- `.agents/lessons/2026-05-07_conversion-cast-bf16-suffix-allowlist.md`
## Provenance
Decomposed from PR [#1280](https://github.com/hao-ai-lab/FastVideo/pull/1280)
(`will/magi` @ `4e1603634d27c8e1b5c4cc5d9387f046547f5c49`). See the package
AGENTS.md for the full PR-stack table.
@@ -0,0 +1,51 @@
# SPDX-License-Identifier: Apache-2.0
"""Minimal user-runnable example for the daVinci-MagiHuman base AV pipeline.
Produces an mp4 with both video (Wan 2.2 TI2V-5B VAE) and audio (Stable
Audio Open 1.0 VAE, first-class FastVideo port in
`fastvideo/models/vaes/oobleck.py`) muxed together via PyAV.
Prerequisites (one-off):
# Accept terms of use on the gated HF repos with your HF_TOKEN:
# - https://huggingface.co/google/t5gemma-9b-9b-ul2
# - https://huggingface.co/stabilityai/stable-audio-open-1.0
# All four cross-variant shared components (Wan 2.2 VAE, T5-Gemma
# encoder + tokenizer, Stable Audio VAE) are lazy-loaded from their
# canonical upstream HF repos on first build, so a single ~25 GB
# cache is shared across every MagiHuman variant.
The umbrella HF repo `FastVideo/MagiHuman-Diffusers` holds all four
variants (base / distill / sr_540p / sr_1080p) under sibling subfolders
and FastVideo will download just the requested subfolder. Local
conversion via `scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py`
is also supported.
"""
from fastvideo import VideoGenerator
PROMPT = (
"A warm afternoon scene: a person sits on a park bench reading a book, "
"surrounded by softly swaying trees."
)
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/base",
num_gpus=1,
)
output_path = "outputs_video/magi_human_basic/output_magi_human.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
# Defaults pulled from the registered preset (magi_human_base):
# height=256, width=448, fps=25, num_inference_steps=32, seed=42.
# Override here only if you have a specific QA scenario.
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,53 @@
# SPDX-License-Identifier: Apache-2.0
"""Minimal user-runnable example for the daVinci-MagiHuman DMD-2 distilled
text-to-AV pipeline.
Same arch as the base model (`basic_magi_human.py`) but with DMD-2 distilled
weights: 8 denoising steps, no classifier-free guidance. ~4x faster than
base at the same 256x480 resolution. Mirrors upstream
`daVinci-MagiHuman/example/distill/run_T2V.sh`.
Prerequisites (one-off):
# 1) Accept terms on the gated HF repos with your HF_TOKEN:
# - https://huggingface.co/google/t5gemma-9b-9b-ul2
# - https://huggingface.co/stabilityai/stable-audio-open-1.0
# Cross-variant shared components (Wan 2.2 VAE + T5-Gemma + Stable
# Audio VAE) are lazy-loaded from their canonical upstream HF repos
# and shared with the base variant cache.
# 2) Convert the distill subfolder of GAIR/daVinci-MagiHuman:
python scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py \\
--source GAIR/daVinci-MagiHuman \\
--subfolder distill \\
--output converted_weights/magi_human_distill \\
--cast-bf16
# `--cast-bf16` is recommended (61 GB fp32 -> 30 GB bf16); the FV pipeline
# loads bf16 anyway, and the conversion keeps norms / RoPE bands fp32.
"""
from fastvideo import VideoGenerator
PROMPT = (
"A warm afternoon scene: a person sits on a park bench reading a book, "
"surrounded by softly swaying trees."
)
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/distill",
num_gpus=1,
)
output_path = "outputs_video/magi_human_basic/output_magi_human_distill.mp4"
generator.generate_video(
prompt=PROMPT,
output_path=output_path,
save_video=True,
# Defaults pulled from the registered preset (magi_human_distill):
# height=256, width=480, fps=25, num_inference_steps=8, cfg=1, seed=42.
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,34 @@
# SPDX-License-Identifier: Apache-2.0
"""Minimal daVinci-MagiHuman DMD-2 distilled text+image-to-AV example."""
from fastvideo import VideoGenerator
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanDistillI2VConfig,
)
PROMPT = (
"A cheerful saxophonist performs a short line with expressive facial "
"motion, natural head movement, and synchronized audio in a small jazz club."
)
IMAGE_PATH = "assets/images/saxophonist.jpg"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/distill",
num_gpus=1,
workload_type="i2v",
override_pipeline_cls_name="MagiHumanI2VPipeline",
pipeline_config=MagiHumanDistillI2VConfig(),
)
generator.generate_video(
prompt=PROMPT,
image_path=IMAGE_PATH,
output_path="outputs_video/magi_human_distill_ti2v/output_magi_human_distill_ti2v.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,45 @@
# SPDX-License-Identifier: Apache-2.0
"""Run daVinci-MagiHuman SR-1080p text-to-AV in FastVideo.
Build the converted repo on large local storage, then symlink it into the
workspace:
python scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py \
--source GAIR/daVinci-MagiHuman \
--subfolder base \
--sr-source GAIR/daVinci-MagiHuman \
--sr-subfolder 1080p_sr \
--output /raid/william5lin_converted_weights/magi_human_sr_1080p \
--cast-bf16
ln -s /raid/william5lin_converted_weights/magi_human_sr_1080p \
converted_weights/magi_human_sr_1080p
"""
from fastvideo import VideoGenerator
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanSR1080pConfig,
)
PROMPT = (
"A warm afternoon scene: a person sits on a park bench reading a book, "
"surrounded by softly swaying trees."
)
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/sr_1080p",
num_gpus=1,
override_pipeline_cls_name="MagiHumanSR1080pPipeline",
pipeline_config=MagiHumanSR1080pConfig(),
)
generator.generate_video(
prompt=PROMPT,
output_path="outputs_video/magi_human_sr1080p/output_magi_human_sr1080p.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,34 @@
# SPDX-License-Identifier: Apache-2.0
"""Run daVinci-MagiHuman SR-1080p text+image-to-AV in FastVideo."""
from fastvideo import VideoGenerator
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanSR1080pI2VConfig,
)
PROMPT = (
"A cheerful saxophonist performs a short line with expressive facial "
"motion, natural head movement, and synchronized audio in a small jazz club."
)
IMAGE_PATH = "assets/images/saxophonist.jpg"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/sr_1080p",
num_gpus=1,
workload_type="i2v",
override_pipeline_cls_name="MagiHumanSR1080pI2VPipeline",
pipeline_config=MagiHumanSR1080pI2VConfig(),
)
generator.generate_video(
prompt=PROMPT,
image_path=IMAGE_PATH,
output_path="outputs_video/magi_human_sr1080p_ti2v/output_magi_human_sr1080p_ti2v.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,37 @@
# SPDX-License-Identifier: Apache-2.0
"""Run daVinci-MagiHuman SR-540p text-to-AV in FastVideo.
The converted repo must contain both ``transformer/`` (base DiT) and
``sr_transformer/`` (540p SR DiT). Build it with:
python scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py \
--source GAIR/daVinci-MagiHuman \
--subfolder base \
--sr-source GAIR/daVinci-MagiHuman \
--sr-subfolder 540p_sr \
--output converted_weights/magi_human_sr_540p
"""
from fastvideo import VideoGenerator
PROMPT = (
"A warm afternoon scene: a person sits on a park bench reading a book, "
"surrounded by softly swaying trees."
)
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/sr_540p",
num_gpus=1,
)
generator.generate_video(
prompt=PROMPT,
output_path="outputs_video/magi_human_sr540p/output_magi_human_sr540p.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,34 @@
# SPDX-License-Identifier: Apache-2.0
"""Run daVinci-MagiHuman SR-540p text+image-to-AV in FastVideo."""
from fastvideo import VideoGenerator
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanSR540pI2VConfig,
)
PROMPT = (
"A cheerful saxophonist performs a short line with expressive facial "
"motion, natural head movement, and synchronized audio in a small jazz club."
)
IMAGE_PATH = "assets/images/saxophonist.jpg"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/sr_540p",
num_gpus=1,
workload_type="i2v",
override_pipeline_cls_name="MagiHumanSRI2VPipeline",
pipeline_config=MagiHumanSR540pI2VConfig(),
)
generator.generate_video(
prompt=PROMPT,
image_path=IMAGE_PATH,
output_path="outputs_video/magi_human_sr540p_ti2v/output_magi_human_sr540p_ti2v.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -0,0 +1,34 @@
# SPDX-License-Identifier: Apache-2.0
"""Minimal daVinci-MagiHuman base text+image-to-AV example."""
from fastvideo import VideoGenerator
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanBaseI2VConfig,
)
PROMPT = (
"A cheerful saxophonist performs a short line with expressive facial "
"motion, natural head movement, and synchronized audio in a small jazz club."
)
IMAGE_PATH = "assets/images/saxophonist.jpg"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"FastVideo/MagiHuman-Diffusers/base",
num_gpus=1,
workload_type="i2v",
override_pipeline_cls_name="MagiHumanI2VPipeline",
pipeline_config=MagiHumanBaseI2VConfig(),
)
generator.generate_video(
prompt=PROMPT,
image_path=IMAGE_PATH,
output_path="outputs_video/magi_human_ti2v/output_magi_human_ti2v.mp4",
save_video=True,
)
generator.shutdown()
if __name__ == "__main__":
main()
@@ -183,8 +183,6 @@ focused PRs:
| 3/8 | [#1297](https://github.com/hao-ai-lab/FastVideo/pull/1297) | `will/magi-03-dit` | DiT + parity tests |
| 4/8 | [#1298](https://github.com/hao-ai-lab/FastVideo/pull/1298) | `will/magi-04a-stages` | Pipeline stages + sr_transformer alias |
| 5/8 | [#1299](https://github.com/hao-ai-lab/FastVideo/pull/1299) | `will/magi-04b-orchestrator` | Pipeline orchestrator + parity battery |
| 6/8 | this PR | `will/magi-04c-provenance` | This AGENTS.md, JOURNAL.md, lessons, parent AGENTS.md hook |
| 7/8 | (next) | `will/magi-05-conversion` | Checkpoint conversion + 3rd lesson |
| 8/8 | (last) | `will/magi-06-activate` | Registry + examples + SSIM + codebase-map |
Provenance section will be finalized in 8/8 with the actual PR numbers for 7/8 and 8/8.
| 6/8 | [#1300](https://github.com/hao-ai-lab/FastVideo/pull/1300) | `will/magi-04c-provenance` | This AGENTS.md, JOURNAL.md, lessons, parent AGENTS.md hook |
| 7/8 | [#1301](https://github.com/hao-ai-lab/FastVideo/pull/1301) | `will/magi-05-conversion` | Checkpoint conversion + 3rd lesson |
| 8/8 | this PR | `will/magi-06-activate` | Registry + examples + SSIM + codebase-map |
+153
View File
@@ -38,6 +38,16 @@ from fastvideo.configs.pipelines.flux_2 import (
)
from fastvideo.configs.pipelines.matrixgame2 import MatrixGame2I2V480PConfig
from fastvideo.configs.pipelines.matrixgame3 import MatrixGame3I2V720PConfig
from fastvideo.pipelines.basic.magi_human.pipeline_configs import (
MagiHumanBaseConfig,
MagiHumanBaseI2VConfig,
MagiHumanDistillConfig,
MagiHumanDistillI2VConfig,
MagiHumanSR1080pConfig,
MagiHumanSR1080pI2VConfig,
MagiHumanSR540pConfig,
MagiHumanSR540pI2VConfig,
)
from fastvideo.configs.pipelines.turbodiffusion import (
TurboDiffusionI2V_A14B_Config,
TurboDiffusionT2V_14B_Config,
@@ -379,6 +389,146 @@ def _register_configs() -> None:
default_preset="flux2_dev",
)
# daVinci-MagiHuman SR-1080p (two-stage base + local-window SR text-to-AV).
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanSR1080pConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-SR-1080p-Diffusers",
"FastVideo/MagiHuman-Diffusers/sr_1080p",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()) and
("sr_1080p" in path.lower() or "sr-1080p" in path.lower() or "1080p_sr" in path.lower() or "sr1080p" in
path.lower()) and "ti2v" not in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_sr_1080p",
)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanSR1080pI2VConfig,
workload_types=(WorkloadType.I2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-SR-1080p-TI2V-Diffusers",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()) and
("sr_1080p" in path.lower() or "sr-1080p" in path.lower() or "1080p_sr" in path.lower() or "sr1080p" in
path.lower()) and "ti2v" in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_sr_1080p_ti2v",
)
# daVinci-MagiHuman SR-540p (two-stage base + SR text-to-AV).
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanSR540pConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-SR-540p-Diffusers",
"FastVideo/MagiHuman-Diffusers/sr_540p",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()) and
("sr_540p" in path.lower() or "sr-540p" in path.lower() or "540p_sr" in path.lower() or "srpipeline" in
path.lower()) and "1080" not in path.lower() and "ti2v" not in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_sr_540p",
)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanSR540pI2VConfig,
workload_types=(WorkloadType.I2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-SR-540p-TI2V-Diffusers",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()) and
("sr_540p" in path.lower() or "sr-540p" in path.lower() or "540p_sr" in path.lower() or "srpipeline" in
path.lower()) and "1080" not in path.lower() and "ti2v" in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_sr_540p_ti2v",
)
# daVinci-MagiHuman (base text-to-AV).
# NOTE: WorkloadType has no T2AV variant yet; using T2V as the
# placeholder until the enum is extended (same as Stable Audio).
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanBaseConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"GAIR/daVinci-MagiHuman",
"FastVideo/MagiHuman-Base-Diffusers",
"FastVideo/MagiHuman-Diffusers/base",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()
) and "distill" not in path.lower() and "ti2v" not in path.lower() and "sr_540p" not in path.lower() and
"sr-540p" not in path.lower() and "540p_sr" not in path.lower() and "sr_1080p" not in path.lower() and
"sr-1080p" not in path.lower() and "1080p_sr" not in path.lower() and "srpipeline" not in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_base",
)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanBaseI2VConfig,
workload_types=(WorkloadType.I2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-Base-TI2V-Diffusers",
],
model_detectors=[
lambda path:
(("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower()
) and "ti2v" in path.lower() and "distill" not in path.lower() and "sr_540p" not in path.lower() and
"sr-540p" not in path.lower() and "540p_sr" not in path.lower() and "sr_1080p" not in path.lower() and
"sr-1080p" not in path.lower() and "1080p_sr" not in path.lower() and "srpipeline" not in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_base_ti2v",
)
# daVinci-MagiHuman (DMD-2 distilled text-to-AV)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanDistillConfig,
workload_types=(WorkloadType.T2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-Distilled-Diffusers",
"FastVideo/MagiHuman-Diffusers/distill",
],
model_detectors=[
lambda path: (("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower())
and "distill" in path.lower() and "ti2v" not in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_distill",
)
register_configs(
sampling_param_cls=None,
pipeline_config_cls=MagiHumanDistillI2VConfig,
workload_types=(WorkloadType.I2V, ),
hf_model_paths=[
"FastVideo/MagiHuman-Distilled-TI2V-Diffusers",
],
model_detectors=[
lambda path: (("magihuman" in path.lower() or "magi_human" in path.lower() or "magi-human" in path.lower())
and "ti2v" in path.lower() and "distill" in path.lower()),
],
model_family="magi_human",
default_preset="magi_human_distill_ti2v",
)
# Hunyuan 1.5 (specific)
register_configs(
sampling_param_cls=None,
@@ -1214,6 +1364,8 @@ def _register_presets() -> None:
ALL_PRESETS as LONGCAT_PRESETS, )
from fastvideo.pipelines.basic.ltx2.presets import (
ALL_PRESETS as LTX2_PRESETS, )
from fastvideo.pipelines.basic.magi_human.presets import (
ALL_PRESETS as MAGI_HUMAN_PRESETS, )
from fastvideo.pipelines.basic.matrixgame2.presets import (
ALL_PRESETS as MATRIXGAME2_PRESETS, )
from fastvideo.pipelines.basic.matrixgame3.presets import (
@@ -1242,6 +1394,7 @@ def _register_presets() -> None:
LINGBOTWORLD_PRESETS,
LONGCAT_PRESETS,
LTX2_PRESETS,
MAGI_HUMAN_PRESETS,
MATRIXGAME2_PRESETS,
MATRIXGAME3_PRESETS,
SD35_PRESETS,
@@ -0,0 +1,113 @@
# SPDX-License-Identifier: Apache-2.0
"""SSIM-based similarity test for daVinci-MagiHuman base text-to-AV.
Reference videos for this test are seeded separately via the
`.agents/skills/seed-ssim-references/` skill on Modal L40S and uploaded
to `FastVideo/ssim-reference-videos`. Until refs exist, the first run
will fail downloading; run the seed skill once and commit the URLs.
Resolution + steps kept small enough for a CI budget; the full-quality
variant falls back to the registered preset defaults.
"""
import os
import pytest
from fastvideo.api.sampling_param import SamplingParam
from fastvideo.logger import init_logger
from fastvideo.tests.ssim.inference_similarity_utils import (
resolve_inference_device_reference_folder,
run_text_to_video_similarity_test,
)
logger = init_logger(__name__)
# 15B DiT + T5-Gemma 9B + Wan VAE + Stable-Audio VAE doesn't fit on a
# single L40S (44 GB). Shard across 2 ranks via FSDP.
REQUIRED_GPUS = 2
device_reference_folder = resolve_inference_device_reference_folder(logger)
# Umbrella HF repo holds all four variants under sibling subfolders;
# `maybe_download_model` parses "org/repo/subfolder" and only fetches
# the selected subfolder. Override via `MAGI_HUMAN_MODEL_PATH` to point
# at a local converted_weights/ dir.
_MAGI_HUMAN_MODEL_PATH = os.getenv(
"MAGI_HUMAN_MODEL_PATH",
"FastVideo/MagiHuman-Diffusers/base",
)
MAGI_HUMAN_BASE_PARAMS = {
"num_gpus": 2,
"model_path": _MAGI_HUMAN_MODEL_PATH,
# height/width/guidance_scale/seed/fps mirror the registered
# `magi_human_base` preset defaults (see
# `fastvideo/pipelines/basic/magi_human/presets.py::MAGI_HUMAN_BASE`)
# so the SSIM test exercises the same code path as
# `examples/inference/basic/basic_magi_human.py`. Only the budget
# knobs (num_frames, num_inference_steps, sp_size) differ for CI fit.
"height": 256,
"width": 480,
"num_frames": 26, # seconds=1 at fps=25 + 1; preset = 101
"num_inference_steps": 8, # CI budget; preset = 32
"guidance_scale": 5.0,
"seed": 42,
"sp_size": 2,
"tp_size": 1,
"fps": 25,
}
try:
_MAGI_HUMAN_FULL_DEFAULTS = SamplingParam.from_pretrained(_MAGI_HUMAN_MODEL_PATH)
MAGI_HUMAN_FULL_PARAMS = {
"num_gpus": MAGI_HUMAN_BASE_PARAMS["num_gpus"],
"model_path": MAGI_HUMAN_BASE_PARAMS["model_path"],
"height": _MAGI_HUMAN_FULL_DEFAULTS.height,
"width": _MAGI_HUMAN_FULL_DEFAULTS.width,
"num_frames": _MAGI_HUMAN_FULL_DEFAULTS.num_frames,
"num_inference_steps": _MAGI_HUMAN_FULL_DEFAULTS.num_inference_steps,
"guidance_scale": _MAGI_HUMAN_FULL_DEFAULTS.guidance_scale,
"seed": _MAGI_HUMAN_FULL_DEFAULTS.seed,
"sp_size": MAGI_HUMAN_BASE_PARAMS["sp_size"],
"tp_size": MAGI_HUMAN_BASE_PARAMS["tp_size"],
"fps": _MAGI_HUMAN_FULL_DEFAULTS.fps,
}
except Exception:
# Model not registered / accessible on this machine — fall back to the
# quick params as the full-quality map too; the test will skip anyway
# when the model path is unavailable.
MAGI_HUMAN_FULL_PARAMS = MAGI_HUMAN_BASE_PARAMS
MAGI_HUMAN_MODEL_TO_PARAMS = {
"MagiHuman-Base-Diffusers": MAGI_HUMAN_BASE_PARAMS,
}
FULL_QUALITY_MAGI_HUMAN_MODEL_TO_PARAMS = {
"MagiHuman-Base-Diffusers": MAGI_HUMAN_FULL_PARAMS,
}
MAGI_HUMAN_TEST_PROMPTS = [
"A person sitting by a window, softly lit by afternoon sun, waving at "
"the camera with a gentle smile.",
]
@pytest.mark.parametrize("prompt", MAGI_HUMAN_TEST_PROMPTS)
@pytest.mark.parametrize("attention_backend_name", ["FLASH_ATTN"])
@pytest.mark.parametrize("model_id", list(MAGI_HUMAN_MODEL_TO_PARAMS.keys()))
def test_magi_human_base_inference_similarity(
prompt: str,
attention_backend_name: str,
model_id: str,
) -> None:
run_text_to_video_similarity_test(
logger=logger,
script_dir=os.path.dirname(os.path.abspath(__file__)),
device_reference_folder=device_reference_folder,
prompt=prompt,
attention_backend_name=attention_backend_name,
model_id=model_id,
default_params_map=MAGI_HUMAN_MODEL_TO_PARAMS,
full_quality_params_map=FULL_QUALITY_MAGI_HUMAN_MODEL_TO_PARAMS,
min_acceptable_ssim=0.60,
)