Pure logic test (no GPU, no model load, no upstream daVinci-MagiHuman
clone) that asserts MagiHumanSRLatentPreparationStage.forward() clears
batch.magi_static_packed_layout. Catches the f1eeb630 regression that
the existing SR-540p pipeline-parity test missed because that test
bypasses the production stage composition and re-implements the SR
denoise loop with simplified inline helpers, so the cross-stage state
transfer via ForwardBatch is never exercised.
Verified the test fails against HEAD~1 (without the f1eeb630 fix) with:
AssertionError: assert <sentinel object> is None
and passes against HEAD.
C4 (4190c720) added a precompute_static_packed_layout call in the base
latent prep stage that stashes coords/modality-maps/max_ch on
batch.magi_static_packed_layout, sized for the BASE-resolution latent.
The SR latent prep stage upsamples batch.latents to the SR grid (e.g.
256x480 -> 512x896 for SR-540p), changing video_token_num and the
shapes of video_coords / video_mm — but it didn't invalidate the
precomputed layout. The SR denoising loop then passed the stale
base-sized layout to build_static_packed_inputs, which produced a
modality_mapping whose first-dim mismatched the SR-sized token tensor,
crashing in MagiHumanDiT.adapter at the text_mask scatter:
IndexError: The shape of the mask [3243] at index 0 does not match
the shape of the indexed tensor [11771, 3584] at index 0
Fix: clear batch.magi_static_packed_layout in MagiHumanSRLatentPrep so
the SR denoising loop falls back to the slow path of
build_static_packed_inputs (which rebuilds from the current latent
shape). Base C4 perf win is preserved (32 base steps); SR has only ~5
steps so the meshgrid recompute cost is negligible.
Repro: examples/inference/basic/basic_magi_human_sr540p.py now runs
end-to-end (34s on B200). SR-540p parity tests t2v + ti2v + DiT parity
+ distill DiT parity all pass.
The MagiAttention LocalAttention layer was hardcoded to TORCH_SDPA. The
attention dispatch already routes through FastVideo's selector, so
adding FLASH_ATTN to the supported list lets bf16 inference paths pick
up Hopper FA-3/FA-4 (or FA-2 elsewhere) automatically. Falls back to
SDPA for fp32 (parity tests) which the selector handles cleanly.
Switch the SSIM test to FLASH_ATTN since that's the production
inference backend; pin the parametrize list to a single backend so the
seeded reference videos correspond to the path users actually run.
All 8 runnable magi-human parity tests still pass under
FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN.
Match basic_magi_human.py code path exactly except for CI budget knobs:
- width 448 -> 480 (preset default, was a typo'd test value)
- num_inference_steps 4 -> 8 (4 was too low to produce a stable
regression baseline; 8 mirrors the distill preset and gives a
recognizable but cheap sample)
- All other knobs (height, guidance_scale, seed, fps, model_path,
cfg_number from pipeline_config, negative_prompt from preset) match
the registered magi_human_base preset and basic_magi_human.py.
Single L40S (44 GB) OOMs during MagiHuman model load: 15B DiT +
T5-Gemma 9B encoder + Wan2.2 VAE + Stable-Audio VAE total ~56 GB bf16.
FSDP across 2 L40S shards the DiT + text encoder so each rank stays
under 44 GB. Mirrors test_ltx2_similarity.py and test_wan_t2v_similarity.py
which use the same 2-GPU FSDP layout for 5-15B-class models on L40S.
The latent preparation stage was deriving num_frames from
batch.num_seconds * fps + 1 unconditionally and ignoring
batch.num_frames. num_seconds was never set anywhere (no plumbing in
SamplingParam or ForwardBatch), so every generation defaulted to 4
seconds = 101 frames, regardless of what the caller passed via
SamplingParam.num_frames.
Now we prefer batch.num_frames when it's a sensible video length (>1)
and fall back to the seconds-based derivation when the caller explicitly
passes num_seconds. Backward-compatible with the production preset
(num_frames=101 == 4*25+1). Fixes the SSIM test budget — it was running
at 101 frames despite asking for 26.
All 8 runnable magi-human parity tests still pass.
Aligns the SSIM test with the rest of the magi-human ports which now use
FastVideo/MagiHuman-Diffusers/base via maybe_download_model's umbrella-
repo support (org/repo/subfolder). The standalone -Base-Diffusers repo
no longer exists publicly; the umbrella repo is the canonical source.
build_static_packed_inputs was called every step inside the denoising
loop, redoing the meshgrid/torch.full() coords + modality-map work even
though those depend only on latent shape, audio length, and channel
widths — all fixed for a single generation.
Add StaticPackedLayout + precompute_static_packed_layout. The latent prep
stage stashes the layout on the batch; the denoise / sr-denoise loops
pass it back through the new layout= arg, which short-circuits the
invariant work and only rebuilds the per-step token tensors. The slow
path (layout=None) is kept bit-exact for build_packed_inputs callers in
parity tests.
Bit-exact verified slow vs fast path. dit / distill_dit / pipeline_smoke
/ sr540p / sr1080p / vae parity tests pass.
Replace F.interpolate(mode='linear') with scipy.signal.resample to match
upstream video_process.resample_audio_sinc which uses the same FFT-based
polyphase resampler. Removes high-frequency aliasing and roll-off the
linear path introduced. scipy is already a direct fastvideo dep.
Pre-detect bundled state by reading model_index.json upfront and only
defer-remove non-bundled keys from required_config_modules. After
super().load_modules, prefer modules.get(key) over loaded_modules so
super-loaded bundled or caller-provided overrides aren't silently
clobbered with a fresh upstream lazy-load.
Phase 5: with all 4 weight variants now uploaded under
FastVideo/MagiHuman-Diffusers (one HF repo, four sibling subfolders
base / distill / sr_540p / sr_1080p), all magi-human examples now
default to the umbrella string. Local conversion via the
checkpoint_conversion script remains supported and is documented in
the example docstrings.
Files updated:
* basic_magi_human.py: base T2V -> base
* basic_magi_human_ti2v.py: base TI2V -> base
* basic_magi_human_distill.py: distill T2V -> distill
* basic_magi_human_distill_ti2v.py: distill TI2V-> distill
* basic_magi_human_sr540p.py: sr-540p T2V -> sr_540p
* basic_magi_human_sr540p_ti2v.py: sr-540p TI2V-> sr_540p
* basic_magi_human_sr1080p.py: sr-1080p T2V-> sr_1080p
* basic_magi_human_sr1080p_ti2v.py: sr-1080p TI2V-> sr_1080p
* fastvideo/registry.py: hf_model_paths for the base T2V config now
includes 'FastVideo/MagiHuman-Diffusers/base' and the distill T2V
config includes 'FastVideo/MagiHuman-Diffusers/distill', alongside
the existing per-variant repo names. The SR umbrella paths
(sr_540p / sr_1080p) were already registered by Phases 3/4. TI2V
variants reuse the same weight subfolders and are selected at
load time via override_pipeline_cls_name + pipeline_config (see
basic_magi_human_ti2v.py for the pattern).
Verified end-to-end:
* pytest tests/local_tests/magi_human/ -v -s: 14 passed, 0 failed.
All bit-exact (diff_max=0.0, diff_mean=0.0).
* basic_magi_human.py with local converted_weights/magi_human_base
moved out, forcing snapshot_download from the umbrella repo:
mp4 byte-identical md5 dcf5f2bf6534c7c0d91e7353e42b23db (matches
pre-upload local-path output exactly).
Notes:
* fastvideo/utils.py:maybe_download_model already supports the
'org/repo/subfolder' umbrella form (committed in e2ef3234), and
fastvideo/pipelines/basic/magi_human/magi_human_pipeline.py
lazy-loads the four shared components (Wan VAE, T5-Gemma encoder
+ tokenizer, Stable Audio VAE) from their canonical upstream HF
repos so each umbrella subfolder only ships transformer/ +
scheduler/ + (sr_transformer/) + model_index.json.
* Total HF Hub footprint after upload: ~165 GB raw, server-side
deduped. User cache footprint per variant is ~5-30 GB transformer
(+ ~30 GB sr_transformer for SR) plus a single ~25 GB upstream
cache shared across all variants.
Ports the daVinci-MagiHuman SR-1080p inference flow to FastVideo. Builds
on the SR-540p two-stage pipeline plus block-sparse local-window
video->video attention on 32 of 40 SR DiT layers. Mirrors upstream
SR2_1080 config override at inference/common/config.py:225-244 and
calc_local_qk_range at inference/pipeline/data_proxy.py:31-79.
Files added:
* tests/local_tests/magi_human/test_magi_human_sr1080p_pipeline_parity.py:
parametric parity test for both T2V and TI2V modes. Both pass
diff_max=0.0000 / diff_mean=0.0000 -- BIT-EXACT.
Files modified (key changes):
* fastvideo/models/dits/magi_human.py:
- AttentionSubConfig.use_local_attn / frame_receptive_field
- MagiAttention.configure_local_attention(): per-layer toggle
- MagiAttention._sdpa(): thin SDPA wrapper for [L,H,D] tensors
- MagiAttention._local_window_attention(): 3-block accumulator
mirroring upstream FFA semantics with vanilla SDPA segments
(per-frame video window + all video->audio+text + audio+text->all).
- MagiAttention.forward() dispatches to _local_window_attention
when use_local_attn flag is set.
- MagiTransformerLayer wires use_local_attn from arch.local_attn_layers.
- MagiHumanDiT.configure_local_attention() top-level toggle.
* pipeline_configs.py: MagiHumanSR1080pConfig + I2V variant with
sr_local_attn_layers populated to upstream's 32 indices.
* presets.py: MAGI_HUMAN_SR_1080P + MAGI_HUMAN_SR_1080P_TI2V presets.
* registry.py: SR-1080p config entries with detectors.
* magi_human_pipeline.py: SR-1080p pipeline classes activating
local_attn_layers on the SR DiT at construction.
* scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py:
--sr-subfolder 1080p_sr support.
* tests/local_tests/helpers/magi_human_upstream.py: arch override
support for SR-1080p parity test.
* test_magi_human_pipeline_smoke.py: preset set expanded.
* basic_magi_human_sr1080p{,_ti2v}.py: stubs -> runnable.
Verification:
* pytest tests/local_tests/magi_human/ -v -s: 14 passed, 0 failed.
All bit-exact including SR-1080p T2V and TI2V parity. The
block-sparse SDPA-segmented implementation matches upstream FFA's
q_ranges/k_ranges accumulator semantics exactly for the 3-block
layout (overlap accumulation handled by explicit '+' at
magi_human.py:421-425).
* basic_magi_human.py (T2V regression check): mp4 byte-identical
md5 dcf5f2bf6534c7c0d91e7353e42b23db.
* Base/Distill/TI2V/SR-540p flows untouched.
Notes:
* Converted SR-1080p artifact at /raid/.../magi_human_sr_1080p
(~58 GB), symlinked at converted_weights/magi_human_sr_1080p
(gitignored).
* Upstream's flex_flash_attn_func via SandAI-org/MagiAttention is
NOT a dependency; FV's pure-SDPA segmented implementation is
mathematically equivalent for this 3-block layout.
Ports the daVinci-MagiHuman SR-540p inference flow to FastVideo for
both T2V and TI2V modes. SR-540p is a TWO-STAGE pipeline: the base
model produces a 256x480 latent, then a separate SR DiT (same arch as
base, different weights) refines it to 896x512. Mirrors upstream
MagiEvaluator.evaluate at video_generate.py:300-360.
Files added:
* fastvideo/pipelines/basic/magi_human/stages/sr_latent_preparation.py
* fastvideo/pipelines/basic/magi_human/stages/sr_denoising.py
* tests/local_tests/magi_human/test_magi_human_sr540p_pipeline_parity.py
Files modified (key changes):
* pipeline_configs.py: SR config classes with sr_* knobs sourced
from upstream EvaluationConfig
* presets.py: MAGI_HUMAN_SR_540P + MAGI_HUMAN_SR_540P_TI2V
* registry.py: SR-540p config entries
* magi_human_pipeline.py: MagiHumanSRPipeline + MagiHumanSRI2VPipeline
classes wiring the 9-stage chain (base denoise -> sr latent prep
-> sr denoise -> decode)
* fastvideo/models/loader/component_loader.py: registers
'sr_transformer' alongside transformer / transformer_2
* scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py:
--sr-source / --sr-subfolder flags for SR DiT into sr_transformer/
* examples/inference/basic/basic_magi_human_sr540p{,_ti2v}.py:
runnable
Verification:
* pytest tests/local_tests/magi_human/ -v -s: 12 passed, 0 failed.
All bit-exact (diff_max=0.0, diff_mean=0.0) including new SR-540p
T2V + TI2V parity tests.
* basic_magi_human.py (T2V regression check): mp4 byte-identical
md5 dcf5f2bf6534c7c0d91e7353e42b23db.
* basic_magi_human_sr540p.py: 896x512 mp4, coherent reading-on-
park-bench scene, much higher quality than base 480x256.
* basic_magi_human_sr540p_ti2v.py: 896x512 mp4 with reference-image-
conditioned saxophonist; reference conditioning preserved through
SR upscale.
Notes:
* Converted SR-540p artifact at converted_weights/magi_human_sr_540p
(~86 GB; both transformer/ and sr_transformer/) is gitignored.
* Base T2V/TI2V/distill flows untouched and bit-exact.
Ports the daVinci-MagiHuman TI2V branch to FastVideo for both base
and distill variants. The TI2V case takes a reference image, encodes
it through the Wan VAE, and overwrites the first frame's video latent
with the encoded image latent at every denoise step (mirrors upstream
inference/pipeline/video_generate.py:300-360 evaluate + 424-425
per-step overwrite).
Files added:
* fastvideo/pipelines/basic/magi_human/stages/reference_image.py:
new MagiHumanReferenceImageStage. Loads PIL image (or path),
resizecrops to (height, width) matching upstream resizecrop,
runs VideoProcessor.preprocess at vae_scale_factor=16, encodes
via the Wan VAE (uses .mean for deterministic latent), applies
shift_factor / scaling_factor normalization, stashes on
batch.image_latent.
* tests/local_tests/magi_human/test_magi_human_ti2v_pipeline_parity.py:
bit-exact parity test against upstream MagiEvaluator's TI2V denoise
(with reference image conditioning). Passes
ti2v video diff_max=0.0000 diff_mean=0.0000.
Files modified:
* pipeline_configs.py: MagiHumanBaseI2VConfig keeps the VAE encoder
loaded (load_encoder=True) so the reference image path can encode.
* presets.py: MAGI_HUMAN_BASE_TI2V and MAGI_HUMAN_DISTILL_TI2V presets
with workload_type=i2v.
* registry.py: TI2V config entries for both base and distill variants.
* magi_human_pipeline.py: MagiHumanI2VPipeline subclass that inserts
the MagiHumanReferenceImageStage between prompt encoding and latent
preparation. Reuses the lazy-load path for shared components.
* stages/latent_preparation.py: pre-loop overwrite of
latent_video[:, :, :1] with batch.image_latent[:, :, :1] when
image conditioning is present (matches upstream
evaluate_with_latent line 425 first-iteration overwrite).
* stages/denoising.py: per-step _overwrite_first_frame helper that
applies the same overwrite at the start of every denoise step
(matches upstream evaluate_with_latent line 424 in-loop overwrite).
static_packed rebuild moved inside the loop after the overwrite
so packed video tokens reflect the conditioned latent.
* examples/inference/basic/basic_magi_human_ti2v.py: rewritten from
NotImplementedError stub to runnable. Uses local
converted_weights/magi_human_base + the existing example
saxophonist reference image; produces a coherent mp4 with the
image-conditioned subject.
* examples/inference/basic/basic_magi_human_distill_ti2v.py: same
pattern against converted_weights/magi_human_distill (runnable
once distill weights are converted).
Verification:
* pytest tests/local_tests/magi_human/ -v -s: 10 passed, 0 failed.
All bit-exact (diff_max=0.0, diff_mean=0.0) including new
test_magi_human_distill_dit_parity (Phase 1) and
test_magi_human_ti2v_pipeline_parity (Phase 2) tests.
* basic_magi_human.py (T2V regression check): mp4 byte-identical
md5 dcf5f2bf6534c7c0d91e7353e42b23db.
* basic_magi_human_ti2v.py: produces coherent saxophonist scene
matching the reference image conditioning. Frames at
/tmp/opencode/ti2v_frame_*.png.
T2V flow unchanged: both T2V example and base parity test produce
identical output to pre-change. TI2V is purely additive on top.
The conversion script's _FP32_KEEP_SUFFIXES list was missing 8 keys
that the BASE checkpoint stores as fp32:
adapter.video_embedder.{weight,bias}
adapter.text_embedder.{weight,bias}
adapter.audio_embedder.{weight,bias}
final_linear_video.weight
final_linear_audio.weight
For the BASE conversion this never surfaced because the BASE checkpoint
already ships these as fp32 (no --cast-bf16 needed; conversion was
identity for these weights). For the DISTILL conversion (which ships
ALL 331 weights as fp32 and relies on --cast-bf16 to produce a 30 GB
bf16 artifact), the omission caused these 8 fp32 layers to be cast
to bf16, which mismatched FV's MagiAdapter/final_linear modules
(declared dtype=torch.float32 at magi_human.py:519-527 and 645-648,
mirroring upstream Adapter at dit_module.py:721-723 and DiTModel at
dit_module.py:896-900).
Symptom: distill DiT parity vs upstream had video diff_mean=0.114
(19% relative error) instead of the expected 0.0. Adding the 8 keys
to _FP32_KEEP_SUFFIXES restores bit-exact parity.
Also adds tests/local_tests/magi_human/test_magi_human_distill_parity.py
mirroring the base DiT parity test but pointing at the distill shards
and converted weights. Bit-exact diff_max=0.0, diff_mean=0.0 vs
upstream daVinci-MagiHuman/inference/model/dit/dit_module.py:DiTModel
loaded from the distill subfolder.
Verified with --cast-bf16 reconversion of GAIR/daVinci-MagiHuman/distill
into converted_weights/magi_human_distill (29 GB).
Recognises an 'umbrella' HF repo layout where a single repo holds
multiple pipeline variants under sibling subfolders, e.g.
FastVideo/MagiHuman-Diffusers/
base/{model_index.json, transformer/, scheduler/}
distill/{...}
sr_540p/{...}
sr_1080p/{...}
Users can pass 'org/repo/subfolder' as the model path and the loader
downloads only that subfolder's blobs (allow_patterns=['<sub>/**'])
and returns the local subfolder snapshot path:
generator = VideoGenerator.from_pretrained(
'FastVideo/MagiHuman-Diffusers/base',
)
Detection is structural: HF Hub repo ids are always two
slash-separated components; a path with 3+ components that does not
exist locally and is not posix-absolute or relative-prefixed is
treated as an umbrella reference. The existing single-repo-per-variant
layout ('FastVideo/MagiHuman-Base-Diffusers') still works unchanged.
Combined with the lazy-load of the four cross-variant shared
components landed in 53ac1985, an umbrella MagiHuman repo only needs
to ship transformer/+scheduler/+model_index.json per variant and the
user's local cache stays at ~75 GB total for all 4 variants instead
of ~400 GB.
Documented in tests/local_tests/magi-human.md under 'Design notes'.
Verified existing converted_weights/magi_human_base local-path flow
still produces a byte-identical mp4 (md5 dcf5f2bf...) to the pre-
refactor reference.
Extends the existing lazy-load pattern for text_encoder, tokenizer,
and audio_vae to the video VAE: each MagiHuman variant's converted
repo no longer needs to bundle a copy of the Wan 2.2 TI2V-5B VAE.
The four cross-variant shared components are now all fetched from
their canonical upstream HF repos at first build:
* text_encoder, tokenizer -> google/t5gemma-9b-9b-ul2 (gated)
* audio_vae -> stabilityai/stable-audio-open-1.0 (gated)
* vae -> Wan-AI/Wan2.2-TI2V-5B-Diffusers
Per-variant converted repo shrinks to transformer/ + scheduler/ +
model_index.json (~5 GB for base bf16, ~30 GB for distill bf16). All
variants share the same ~25 GB cache of upstream weights, so a user
running 4 variants ends up with ~75 GB total instead of ~400 GB.
Implementation:
* fastvideo/utils.py:verify_model_config_and_directory now treats
the contents of model_index.json as authoritative for which
component subfolders must exist locally. Pipelines that emit a
minimal model_index.json (omitting vae / text_encoder / etc.)
pass verification; pipelines that DO declare a component must
still ship its subfolder. transformer/ remains mandatory.
* fastvideo/pipelines/basic/magi_human/magi_human_pipeline.py adds
vae to the deferred list in load_modules and a new
_load_video_vae helper that prefers a bundled vae/ subfolder
(legacy converted repos) and falls back to snapshot_download +
the standard FV VAELoader. Both paths produce the same FV
AutoencoderKLWan, so production behavior is unchanged.
* convert_magi_human_to_diffusers.py docstring updated; --bundle-vae
flag is unchanged (still optional) but the README example now
omits it so new converted repos default to the minimal layout.
Verification:
* basic_magi_human.py with bundled vae/: byte-identical mp4 (md5
dcf5f2bf...) to pre-refactor output.
* basic_magi_human.py with vae/ moved out and removed from
model_index.json: same byte-identical mp4 via the lazy-load path.
* 4/4 magi-human parity tests pass with diff_max=0.0 / diff_mean=0.0
(DiT, pipeline, smoke + typed surface preflight).
Upstream daVinci-MagiHuman ships 4 model variants x 2 input modes = 8
inference entrypoints (base / distill / sr_540p / sr_1080p, each in
T2V and TI2V mode). FastVideo currently has working code for the base
T2V path (basic_magi_human.py) and a registered preset for distill
T2V (magi_human_distill, but no example until now).
Add example files for the 7 remaining variants:
* basic_magi_human_distill.py -- runnable T2V example for the
DMD-2 distilled model. Just point conversion at the distill
subfolder and the existing magi_human_distill preset takes over.
* basic_magi_human_ti2v.py
basic_magi_human_distill_ti2v.py -- not-yet-ported TI2V (image
conditioning) variants. Each docstring lists the pipeline-side
work that is missing in FastVideo (VAE encoder load, reference
image stage, latent_video[..., :1] overwrite at every denoise
step, new I2V config + preset). main() raises NotImplementedError
with a pointer to magi-human.md.
* basic_magi_human_sr540p.py
basic_magi_human_sr540p_ti2v.py -- not-yet-ported super-resolution
to 540p. Docstring describes the upstream two-stage flow (base ->
SR latent prep with trilinear up + ZeroSNR noise -> SR DiT -> Wan
VAE) and lists the FV components needed (MagiHumanSR540pConfig,
SR latent prep stage, SR denoise stage with cfg-trick guidance
tensor, conversion script invocation for the 540p_sr subfolder,
new preset + registry entry).
* basic_magi_human_sr1080p.py
basic_magi_human_sr1080p_ti2v.py -- not-yet-ported SR-1080p.
Same SR-540p scaffolding plus a block-sparse local-window
attention path (32 of 40 SR DiT layers). The docstring points at
upstream's FFAHandler q_ranges/k_ranges blocks in dit_module.py
and notes that MagiAttention currently always runs full SDPA.
All stubs follow the existing example file convention (SPDX header,
focused docstring, single main()). The not-yet-ported stubs exit with
NotImplementedError so they fail loudly rather than silently misbehaving;
each error message points at the docstring for the missing-component
checklist.
Drop the file-local apply_rotary_emb / _rotate_half helpers and call
fastvideo.layers.rotary_embedding._apply_rotary_emb with
is_neox_style=True instead. Magi uses partial RoPE (rotate first
6 * (head_dim // 8) = 96 of 128 head_dim positions, leave the trailing
32 unrotated), which the FV primitive does not handle directly, so
the partial-RoPE slicing stays in the call site:
q_rot = _apply_rotary_emb(q[..., :rot_dim], cos, sin, is_neox_style=True)
q = torch.cat([q_rot, q[..., rot_dim:]], dim=-1)
The math is identical to the previous local impl (Magi's
'rotate_half + doubled cos/sin' expands to upstream's 'chunk + cat(o1,
o2)' Neox formulation). Drops the einops dependency from this file.
Bit-exact preservation verified at both scales:
* All 7 parity tests still pass with 0.0/0.0 diff vs upstream.
* Production E2E mp4 is byte-identical (same md5 hash) to the
post-LocalAttention output, so basic_magi_human.py output is
unchanged.
Replace bare F.scaled_dot_product_attention call inside MagiAttention
with FastVideo's LocalAttention layer so the backend selection (SDPA /
FlashAttn / SLA / SageAttn) flows through the standard configurable
path. Also drops the manual GQA repeat_interleave: SDPABackend's
enable_gqa=True handles num_heads_q != num_heads_kv directly.
LocalAttention requires a forward_context, so add set_forward_context
wrapping at the two call sites:
* Production: MagiHumanDenoisingStage's per-step DiT calls now run
inside set_forward_context(current_timestep=t, attn_metadata=None),
matching the pattern used by the generic DenoisingStage and other
custom denoise stages (ltx2, stable_audio, longcat).
* Parity tests: test_magi_human_parity.py and
test_magi_human_pipeline_parity.py wrap their direct DiT calls in
the same context so LocalAttention's get_forward_context() doesn't
assert.
Parity tests stay bit-exact (all 7 magi-human parity tests pass with
0.0/0.0 diff vs upstream). Production E2E still produces a coherent
video matching the prompt; the mp4 output is no longer byte-identical
to the pre-refactor version because SDPA dispatches to a different
kernel when enable_gqa=True at production sequence length (~3000
tokens), with mean per-pixel drift ~3/255 (~1.3%) -- visually
indistinguishable, just a different bf16 quantization noise pattern.
The MagiAttention block stays as a custom nn.Module rather than being
fully replaced by LocalAttention because of the per-modality packed
linears (PackedExpertLinear), per-head sigmoid gating, and per-modality
RMSNorm pattern that the FV layer abstractions don't model. The bare
SDPA call inside it is now the only piece reused from FV.
Wave 14b's bit-exact dtype boundary fix introduced two violations of
Wave 10's dtype-agnostic pattern in MagiAttention:
1. q/k/v.to(torch.bfloat16) hardcoded before SDPA
2. out.float() explicit upcast after attention to fp32 the gating
Both turn out to be unnecessary:
1. orig_dtype = self.linear_qkv.weight.dtype already evaluates to
bf16 in production (loader sets bf16) AND in the parity test
(PackedExpertLinear's __init__ default is bf16, matching upstream
BaseLinear at dit_module.py:330). So q.to(orig_dtype) gives the
same bf16 cast as the hardcode without the dtype lock-in.
2. PyTorch's float type-promotion rules already handle the
bf16 -> fp32 transition implicitly: bf16_out * sigmoid(fp32_g)
promotes to fp32, exactly matching upstream's intentional
bf16 * fp32 -> fp32 boundary at dit_module.py:649. The explicit
.float() was redundant.
Result: 11 insertions, 14 deletions, attention block reads more
cleanly without sacrificing any of the parity work:
- All 7 magi-human local parity tests pass with bit-exact (0.0/0.0)
or near-bit-exact (VAE 8e-4) outputs vs upstream
- basic_magi_human.py E2E produces byte-identical mp4 (same md5
hash) as the pre-refactor Wave 14b state, so production
behavior is unchanged
The MagiAttention block stays custom (rather than reusing FV's
LocalAttention) because LocalAttention requires a forward_context
that the parity test does not establish; switching would force
either expanded test scaffolding or a parity-validation regression.
Three cumulative dtype-boundary divergences caused FV parity tests
to fail after the Wave 14 channel-major fix exposed real-signal
processing (vs the prior garbage-in-garbage-out kernel cancellation):
1. MagiAttention forward: hardcode bf16 for SDPA q/k/v inputs to
match upstream flash_attn_with_cp's hardcoded bf16 cast at
dit_module.py:508. Upcast attention output to fp32 before the
per-head gating multiply to match upstream's bf16*fp32 promotion
at dit_module.py:649. Cast back to bf16 only for linear_proj.
This is an intentional exception to Wave 10's dtype-agnostic
pattern: upstream is not dtype-agnostic at this boundary.
2. MagiHumanDiT forward: drop the x.to(linear_qkv.weight.dtype)
cast that ran the residual stream in bf16 across all 40 layers.
Upstream casts to params_dtype which defaults to fp32, keeping
the cross-layer accumulator fp32 with bf16 internal compute.
The bf16 residual cast was compounding ~6-7 bits of mantissa
loss per layer and was visible in pipeline parity diff.
3. Parity test _build_fastvideo_schedulers: switch to single-shift
construction (default shift=1 in __init__, then set_timesteps
with the real shift) to match production migration in Wave 11.
The stale double-shift was applying a non-trivial shift twice
versus upstream's single-shift, leaving FV with a different
timestep schedule.
Result: all 7 magi-human local parity tests pass with bit-exact
or near-bit-exact (VAE 8e-4 max) outputs vs upstream:
test_magi_human_dit_parity diff=0.0
test_magi_human_t5gemma_parity diff=0.0
test_magi_human_sa_audio_parity diff=0.0
test_magi_human_sa_audio_official_parity diff=0.0
test_magi_human_vae_parity max=8e-4
test_magi_human_pipeline_smoke passes
test_magi_human_pipeline_latent_parity diff=0.0
Production E2E re-validated: examples/inference/basic/basic_magi_human.py
still produces coherent video at the standard 480x256 / 32-step /
seed-42 prompt, with unchanged 23.5s runtime.
Closes OQ-6 (full resolution including dtype boundaries).
FV's _img2tokens was packing video latent patches as spatial-major
'(pT pH pW C)' (channels innermost), but the DiT's video_embedder
Linear weight was trained on upstream's channel-major '(C pT pH pW)'
layout. Upstream's MagiDataProxy uses UnfoldNd, which is implemented
via a grouped-channel conv whose output reshape is channel-major
(in_channels * kernel_size_numel ordering, channels slowest). The
spatial-major layout silently permuted the in-features of every
video token, scrambling the entire video feature representation
and producing pure colorful-blob noise from the basic example.
The pipeline parity test could not catch this because it imports
FV's build_packed_inputs for the upstream side too, so both sides
consumed equally-permuted tokens and agreed on garbage at production
scale (T=26 frames, ~3120 video tokens), while reporting healthy
~0.5%/step drift on the tiny synthetic (2,6,6) latent. Fix is a
one-character change in the einops rearrange pattern.
unpack_tokens stays spatial-major because the DiT's final_linear_video
is trained to emit (pT pH pW C), matching upstream's
SingleData.depack_token_sequence at data_proxy.py:220-228.
Validated end-to-end: examples/inference/basic/basic_magi_human.py
now produces coherent video at the standard 480x256 / 32-step / seed
42 prompt -- woman on a park bench reading a book under green trees,
matching prompt. Output mp4 size dropped from ~932KB (incompressible
noise) to ~222KB (coherent compressible video).
Closes magi-human OQ-6.
Notes for follow-up:
- OQ-11 (NEW): test_magi_human_pipeline_parity.py:222 should drive
the upstream side through real MagiDataProxy.process_input so the
parity test catches packing-layout regressions natively.
Match the canonical FastVideo dtype pattern (per WanVideo DiT). Remove
all 7 hardcoded `.to(torch.bfloat16)` casts that were verbatim copies of
upstream daVinci-MagiHuman/inference/model/dit/dit_module.py. Replace
with `orig_dtype = self.linear_qkv.weight.dtype` (or equivalent
loader-owned dtype) + `.to(orig_dtype)` pattern, mirroring
fastvideo/models/dits/wanvideo.py:325-392.
Refactored sites: attention pre_norm output, q/k/v post-RoPE, attention
output, MLP pre_norm, MLP activation output, top-level block-input cast.
Model dtype is now loader-owned via pipeline_config.dit_precision, not
hardcoded. Production bf16 behavior is bit-identical (loader sets
default_dtype=bf16, all params/inputs naturally bf16, orig_dtype=bf16,
outputs preserved as bf16). fp32 parity now works end-to-end on the FV
side; remaining drift in fp32 parity tests is upstream's still-hardcoded
bf16 casts (tracked as OQ-9 in tests/local_tests/magi-human.md).
Add docs/contributing/activation_trace.md: a 210-line contributor guide
covering the env-gated activation trace infrastructure (Extension 0),
its five configuration env vars, the trace_step() context manager, and
the JSONL output format.
The doc also sketches Extensions 1-3 (FX graph capture, AST-level
instrumentation, dispatch-table interception) as future work, giving
contributors a clear design ladder to climb without requiring them to
implement everything at once.
Update tests/local_tests/magi-human.md with the Wave 9 investigation
entry: records the E2E smoke result (28,864 JSONL records, 451 hooked
modules), the zero-overhead-when-off confirmation, and the new
fastvideo/tests/hooks/test_activation_trace.py test row in the Phase 11
status table. Last-verified date bumped to 2026-05-01 (Wave 9).
Wire the new trace_step(step_idx) context manager around each DiT
forward call in the magi-human denoising stage so that per-step
activation dumps are correctly indexed.
The context manager sets a thread-local step index that hook callbacks
read when deciding whether to emit a record (controlled by
FASTVIDEO_TRACE_STEPS). Without this wiring, all records would carry
step_idx=None and step-filtered traces would be empty.
The change is a no-op when FASTVIDEO_TRACE_ACTIVATIONS is unset: the
context manager is a lightweight nullcontext in that path.
Add an env-gated activation trace mode that registers PyTorch forward
hooks on selected modules, computes per-tensor stats (abs_mean, sum,
min, max, mean, std, shape, dtype), and writes JSONL records to a
configurable sink path.
Designed for parity-debug across model ports: enable trace on both
FastVideo's path and the upstream reference, then `diff` the two JSONL
files to find the first divergent layer. Inspired by SGLang's
--debug-tensor-dump-output-folder pattern and TransformerEngine's
DumpTensors selective-dump infra.
Zero-overhead-when-off guarantee: the master toggle
FASTVIDEO_TRACE_ACTIVATIONS is checked ONCE at pipeline startup. When
unset/false, attach_activation_trace() returns None and no hooks are
ever registered. The production forward path is untouched.
Configuration via 5 env vars (FASTVIDEO_TRACE_LAYERS regex filter,
FASTVIDEO_TRACE_STATS, FASTVIDEO_TRACE_OUTPUT path with <pid> templating,
FASTVIDEO_TRACE_STEPS step-index filter). Step indexing via the
trace_step(step_idx) context manager wires step into thread-local for
hook callbacks.
Six unit tests cover off/on/filter/stats/step-filter/tuple-flattening
/cleanup paths. End-to-end smoke against magi-human's basic example
generated 28,864 JSONL records on 451 hooked modules; trace OFF does
not create the output file.
test_magi_human_pipeline_parity.py previously used random txt_feat and
neg_txt_feat tensors (identical on both FV and upstream sides). This
meant the parity test never exercised the actual text-encoding path,
and production-facing preset values (the real positive/negative prompt
strings) were never validated end-to-end.
Lines 59-153 now encode the real preset prompts via T5-Gemma on both
sides before the denoising loop. Lines 388-395 wire the encoded
embeddings into the FV pipeline call. This validates that:
- The preset prompt strings flow through T5-Gemma correctly.
- The encoded embeddings are numerically consistent between FV and
upstream for the same input text.
- Production-facing preset values are exercised in the parity path.
Note: Wave 8 production fixes (tokenizer pre-padding, resolution
defaults) do NOT change parity numbers because both sides use the same
encoder/decoder. The residual drift is the inherent bf16+CFG
amplification floor (Wave 8 audit conclusion).
MagiHumanAudioDecodingStage.forward() previously returned silently
(no audio output) when batch.audio_latents was None or missing. In a
joint audio-video pipeline this is a real bug: the caller expects audio
output and gets nothing, with no indication of why.
audio_decoding.py:90-96 now raises ValueError with a descriptive
message when audio_latents is absent. This converts a silent wrong
result into a loud, actionable error.
This is Wave 8 fix#10. Classified AMBIGUOUS→HARMFUL: silent return
is acceptable for optional audio in a T2V-only pipeline, but MagiHuman
is a joint AV model where missing audio_latents indicates a real
upstream failure (e.g. denoising stage dropped the audio modality).
FV's magi_human_base and magi_human_distill presets used width=448,
height=256 as defaults. Upstream daVinci-MagiHuman uses width=480,
height=272 (snapped to 256 by the latent-preparation stage). Production
users running FV with default settings got a different aspect ratio
(448/256=1.75) than upstream (480/256=1.875), causing visual composition
differences.
Fix: presets.py:50-84 now uses width=480 for both base and distill
presets. latent_preparation.py:130-133 fallback also updated to 480.
This is Wave 8 fix#3. Upstream reference: daVinci-MagiHuman/
video_generate.py default resolution args. Fixes OQ-6 production root
cause: aspect ratio mismatch between FV default and upstream default.
The T5GemmaEncoderModel._encode() call was passing truncation=True,
padding='max_length', max_length=640 to the HF tokenizer, causing the
tokenizer to pad every sequence to 640 tokens BEFORE encoding. This
means pad-token hidden states were fed into the DiT as real content,
and magi_original_text_lens reported the pre-padded length (640) rather
than the actual token count.
Upstream (daVinci-MagiHuman/models/text_encoder.py) does NOT pre-pad:
it tokenizes without padding/truncation and lets the downstream
MagiHumanLatentPreparationStage._pad_or_trim_dim1 handle length
alignment. FV now matches this: t5gemma.py:57-64 no longer passes
truncation/padding/max_length to the tokenizer call.
This is Wave 8 fix#1. Fixes OQ-6 production root cause: pad-token
hidden states were polluting DiT cross-attention input for every
inference call.
Document Wave 7 CFG + negative-prompt investigation findings in the
numerical-alignment investigation section. Key results: CFG math is
identical on both sides; audio VAE is bit-exact vs official (new parity
test confirms); root cause of production audio drift was the incomplete
negative prompt in FV's preset (missing audio-quality + speech-delivery
blocks from upstream video_generate.py:222-224).
Update Phase 11 status table with the new SA-official parity test row
(PASS, diff_max=0, diff_mean=0). Add test_magi_human_sa_audio_official_parity.py
to the running-tests command block and the What-each-test-covers section.
Mark OQ-6 as PARTIALLY-RESOLVED: production neg-prompt fix shipped in
this wave; parity-test 4-step compounding is a separate inherent
FlowUniPC scheduler phenomenon, not a code bug. Update Last-verified
to 2026-05-01 (Wave 7).
Add parity test comparing FastVideo's SAAudioVAEModel against the official
daVinci-MagiHuman SAAudioFeatureExtractor.decode() path. The sibling
test_magi_human_sa_audio_parity.py validates against Diffusers
AutoencoderOobleck; this test validates against the upstream repo's custom
integration layer, which rebuilds AudioAutoencoder from model.pretransform.config
and filters pretransform.model.* weights.
Test passes at machine-eps (diff_max=0, diff_mean=0) in fp32, confirming
FV's SAAudioVAEModel is bit-identical to the official decode path. This
rules out the audio VAE as a contributor to OQ-6 (pipeline compounding).
Requires the daVinci-MagiHuman upstream clone and the gated
stabilityai/stable-audio-open-1.0 repo. Skips cleanly when either is absent.
Complete the magi-human negative prompt to match upstream's three-block
concatenation (video + audio-quality + speech-delivery negatives). FV's
preset previously included only the video-side block, leaving audio CFG
to amplify the missing-block delta 5x via `v = uncond + 5 * (cond - uncond)`.
This explains the audio-side amplification observed in the pipeline-trace
investigation (step 1 v_cond_audio: 4.6% drift; v_cfg_audio: 13.8%).
Upstream reference: daVinci-MagiHuman/inference/pipeline/video_generate.py
:222-224.
Also harden CFG path: missing negative embeds previously silently fell
back to zeros, which is a hidden CFG amplifier. Now raises ValueError
explicitly when CFG=2 is set without negative embeds. Tracks OQ-6
production fix in tests/local_tests/magi-human.md.
Document Wave 1-5 numerical-alignment investigation results in
tests/local_tests/magi-human.md (417 lines total):
- OQ-6: DiT bf16 noise-floor drift (diff_max=0.057 at atol=0.03);
pipeline compounding at 4 steps (18.85x ratio vs 1-step baseline);
root cause under investigation via _debug_magi_human_block_parity.py
- OQ-7: Wan VAE fp32 op-order drift (z*std+mean vs z/(1/std)+mean);
shared Wan-family bug, deferred pending upstream fix
Also records Wave 1-5 methodology: loader fix, per-side log emission,
tolerance tightening, scheduler split, and bisect confirmation that
the compounding bug predates Wave 1.
Surface known bf16 numerical alignment bugs by tightening parity bounds:
- DiT parity atol=0.1 -> atol=0.03 (test FAILs at observed diff_max=0.057;
bf16-noise-floor; tracked as OQ-6 root-cause investigation)
- Pipeline parity num_inference_steps=1 -> 4 (test FAILs at video
diff_max=15.10, diff_mean=1.30 = 18.85x compounding ratio vs 1-step
baseline 0.069; pre-existing compounding bug in original magi port,
NOT introduced by Wave 1 -- bisect confirmed; tracked as OQ-6)
- Wan VAE parity atol=5e-2 -> atol=1e-3 (Wan VAE shared fp32 op-order
drift, FV `z * std + mean` vs upstream `z / (1/std) + mean`; affects
all Wan-family pipelines, deferred per OQ-7)
Also split _build_schedulers into _build_fastvideo_schedulers (double-shift,
matching MagiHumanDenoisingStage production) and _build_upstream_schedulers
(single-shift, matching MagiEvaluator.eval_with_text) to faithfully mirror
each side's production scheduler init pattern.
These tests SHIP IN A FAILING STATE intentionally as a forcing function
for follow-up investigation. See OQ-6 + OQ-7 in magi-human.md.
Add _debug_magi_human_block_parity.py, a standalone script that runs
forward hooks on both the upstream DiTModel and FastVideo MagiHumanDiT,
then writes per-side layer activation logs to:
/tmp/opencode/magi_dit_up_layers.log
/tmp/opencode/magi_dit_fv_layers.log
The logs record (name, shape, abs_mean, sum, min, max) for every hooked
activation, sorted by layer order. This enables side-by-side diff via
`diff /tmp/opencode/magi_dit_up_layers.log /tmp/opencode/magi_dit_fv_layers.log`
to pinpoint the first block where numerical divergence appears — a key
diagnostic step for OQ-6 root-cause investigation.
Replace hf_hub_download (index-only canary) with snapshot_download using
allow_patterns=["base/*.safetensors", "base/model.safetensors.index.json"].
The old approach returned the parent of the index file, which could be a
symlink-resolved HF cache path that lacked the actual shard blobs. The new
approach verifies that at least one .safetensors shard is present before
returning the candidate directory, preventing false-positive skip decisions.
Applies to test_magi_human_parity.py, test_magi_human_pipeline_parity.py,
and the new _debug_magi_human_weight_diff.py debug script (which uses the
same loader pattern for weight-diff analysis).
Drop the magi-introduced 'encoder-style' Stable Audio VAE wrapper now
that main has merged a first-class Oobleck VAE port plus a shared
SAAudioVAEModel pipeline-glue lazy-loader (#1260). MagiHuman now
reuses that infrastructure instead of carrying duplicates:
- audio VAE config: fastvideo.configs.models.vaes.OobleckVAEConfig
- audio VAE wrapper: fastvideo.models.vaes.sa_audio.SAAudioVAEModel
Removes:
- fastvideo/configs/models/encoders/sa_audio.py (SAAudioVAEConfig)
- fastvideo/models/encoders/sa_audio.py (SAAudioVAEModel dup)
Updates:
- magi-human pipeline + config to construct OobleckVAEConfig and set
pretrained_path instead of arch_config.sa_audio_model_path.
- encoders/__init__.py drops the SAAudioVAEConfig export.
- parity test moves from tests/local_tests/encoders/ to .../vaes/
(matches main's classification) and switches to the shared wrapper;
explicit pretrained_dtype='float32' override keeps the fp32 ref
parity check intact (default is fp16 to match official
stable_audio_tools).
audio_decoding.py docstring still mentions 'sa_audio_vae_model' — main's
wrapper exposes that name as a back-compat alias for oobleck_vae, so
the existing comment is still accurate.
- Drop misleading "T2V" framing — base MagiHuman is a joint audio-visual
generator. Rename MagiHumanT2VConfig→MagiHumanBaseConfig,
MagiHumanDistillT2VConfig→MagiHumanDistillConfig, presets
magi_human_base_t2v→magi_human_base, magi_human_distill_t2v→
magi_human_distill. Update comments / docstrings / example output
filename. Keep workload_type="t2v" string (framework enum has no
T2AV variant yet — same placeholder Stable Audio uses for T2A).
- Pipeline parity test: tighten atol/rtol from 0.5/0.5 to 0.35/0.05.
atol absorbs the observed worst-element drift (~0.31 on signal abs_mean
~2.4); the tight rtol still flags gross structural bugs.
- Wan VAE parity: import FastVideo's AutoencoderKLWan from
fastvideo.models.vaes.wanvae instead of diffusers. The test now
actually validates the class MagiHumanBaseConfig.vae_config resolves
to in production.
- Pipeline parity scheduler init: split per-side so upstream mirrors
MagiEvaluator's single-shift pattern (fresh FlowUniPCMultistepScheduler
+ shift only in set_timesteps) and FastVideo mirrors
MagiHumanDenoisingStage's double-shift pattern (constructor + set_timesteps
both with shift). Surfaces a real production divergence between
FastVideo's magi denoise loop and the official one for N>1 steps.
Ports GAIR-NLP/daVinci-MagiHuman's 15B-param joint-AV DiT into FastVideo:
text + video + audio joint denoise in one flat token stream, 32-step
FlowUniPC w/ CFG=2, Wan 2.2 TI2V-5B video VAE, Stable Audio Open 1.0
audio VAE (first-class port at fastvideo/models/vaes/oobleck.py, no
runtime diffusers import), T5-Gemma 9B UL2 text encoder, auto-muxed
mp4 with h264 video + aac stereo audio.
Parity tests all pass on converted weights: DiT (diff_median=2.6e-3),
text encoder (exact), video VAE (8e-4), audio VAE (exact), pipeline
latent (5e-2).
See .agents/skills/add-model/REVIEW.md for porting procedure + open
review items.
## Summary
Follow-up to #1187. Two small changes:
1. **Merge Protections expanded** — adds `#approved-reviews-by>=1` and `check-success~=pre-commit` to `merge_protections` so the Mergify check shows a unified requirements checklist on every PR (title format + approval + pre-commit), instead of only showing the title format.
2. **Buildkite pipeline comment fix** — updates the outdated Full Suite section comment from "Triggered by adding the 'ready' label via GitHub Actions → Buildkite API" to reflect the new Merge Queue trigger path.
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Delete Open Sora Plan Modeling
amend log validation
Change name to fast video
Remove OSP modeling
clean up name changing
remove files
Deleted unnecessary files
commit first
commit first
training
ok
Update generate_synthetic.sh and deepspeed_zero2_config.yaml
14
debug
overfitting ....
debugging ..
Debug successful!
Add zero3
OK
update optimizer
load
small bug
random seed args
sp enable & still has bug
rename
update inference sp code / can run / still has bug / don't output normal mp4
switch to deepspeed dummyoptim
fix some bugs; still output green
SP inference done!
fix typos in readme and latent dataset debug file
Update dependencies; Setup code for debugging
Delete unused files and code
Remove all npu code
Update torchvision imports
Update EMA model and t2v_debug.sh script
Delete npu related stuff and remove inpaint module
Remove compress kv
Fix warning with dataset handling and model loading
Update PyTorch index URLs and video length tolerance range
Remove UDIT and inpaint
fix typo for dataset download
include pretrained open-sora
Add pretrained model for OpenSoraT2V-ROPE-L
Update max height and width for video processing
{"name":"codebase-map","description":"High-level structural index of the FastVideo-WorldModel repository","path":"codebase-map/README.md","status":"ready","trust":"high"}
{"name":"evaluation-registry","description":"Catalog of all evaluation metrics with detailed explanations, implementation status, and usage guides","path":"evaluation-registry/README.md","status":"draft","trust":"medium"}
{"name":"experiment-journal","description":"Living log of all experiments with hypotheses, configs, metrics, and insights","path":"experiment-journal/README.md","status":"draft","trust":"medium"}
{"name":"related-work","description":"Index of related papers, repos, and blog posts with structured comparisons to FastVideo","path":"related-work/README.md","status":"draft","trust":"low"}
{"name":"launch-experiment","description":"Generate and execute a training launch command for FastVideo models","path":"launch-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"monitor-experiment","description":"Poll a running W&B training run for progress and emit structured alerts","path":"monitor-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"summarize-run","description":"Extract a W&B run summary into a structured experiment report","path":"summarize-run/SKILL.md","status":"draft","trust":"low"}
{"name":"log-experiment","description":"Append or update an experiment entry in the experiment journal","path":"log-experiment/SKILL.md","status":"draft","trust":"low"}
{"name":"evaluate-video-quality","description":"Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)","path":"evaluate-video-quality/SKILL.md","status":"draft","trust":"low"}
{"name":"index-related-work","description":"Ingest a paper or repository into the related work index","path":"index-related-work/SKILL.md","status":"draft","trust":"low"}
{"name":"search-related-work","description":"Query the related work index for relevant papers, repos, or comparisons","path":"search-related-work/SKILL.md","status":"draft","trust":"low"}
{"name":"seed-ssim-references","description":"Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline","path":"seed-ssim-references/SKILL.md","status":"draft","trust":"low"}
{"name":"reseed-ssim-references","description":"Re-seed (overwrite) HF reference videos for an existing fastvideo/tests/ssim/ test and a single model id on Modal L40S. Always backs up current refs first, regenerates on Modal, pauses for the user to eyeball before-vs-after, then uploads with --force scoped to --model-id. Sister skill to seed-ssim-references; use when intentional code change has invalidated existing refs","path":"reseed-ssim-references/SKILL.md","status":"draft","trust":"low"}
{"name":"add-model","description":"Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.","path":"add-model/SKILL.md","status":"draft","trust":"low"}
description:Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
## Purpose
Replace the existing SSIM reference videos for a single `(test_file, model_id)`
pair on the HF dataset (`FastVideo/ssim-reference-videos`). This is **destructive**
on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
Pairs with `seed-ssim-references`, which is the inverse (first-time seeding
only, refuses to overwrite). Re-seeding is intentionally a separate, more
ceremonial operation because mistakenly clobbering production refs is much
harder to recover from than failing closed.
## When to use
- An intentional code change (model port fix, kernel upgrade, attention
backend swap, hyperparameter change in the test itself) has shifted the
expected SSIM output and the existing refs no longer represent the new
ground truth.
- A test is failing in CI **for the right reason** (the new code is correct,
the old refs are stale).
## When not to use
- A test is failing for the **wrong** reason (the port is buggy, not the
refs). Fix the port; re-seeding hides the bug.
- A brand-new test that has no refs on HF yet. Use `seed-ssim-references`.
- "Just to clean up drift" without a concrete code change to point at. The
PR description has to justify *why* refs changed; without a concrete
change, there's nothing to write.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `test_file` | Yes | Path to the SSIM test, e.g. `fastvideo/tests/ssim/test_matrixgame_similarity.py`. Validated against `fastvideo/tests/ssim/test_*_similarity.py`. |
| `model_id` | Yes | Single model id from the test's `*_MODEL_TO_PARAMS`, e.g. `Matrix-Game-2.0-Diffusers-Base`. Re-seed runs are **per model**. For multi-model tests, invoke the skill once per model. |
| `intent_rationale` | Yes | One-line explanation of *why* refs are being regenerated (e.g. "Relax FA-2 head_size whitelist to include 80 — matrix_game now uses FLASH_ATTN instead of TORCH_SDPA"). Recorded in the backup directory and reused in the PR description. |
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
| 2026-05-02 | Initial version. Sister skill to `seed-ssim-references`, scoped to single `(test_file, model_id)` re-seeds, with mandatory backup and two-token confirm. |
description:Seed HF reference artefacts for a single newly-added SSIM test (pixel `.mp4` for `run_text_to_video_similarity_test`-style tests, or latent `.pt` for `run_text_to_latent_similarity_test`-style tests). Runs the test on Modal L40S, downloads the generated artefacts via `modal volume get`, pauses for the user to verify (visual eyeball for mp4, numerics dump for pt), then uploads only that test's files to `FastVideo/ssim-reference-videos`. Use when a new `fastvideo/tests/ssim/test_*_similarity.py` has just been added and has no references on HF yet.
---
# Seed SSIM Reference Artefacts (mp4 or pt)
## Purpose
A brand-new SSIM test in `fastvideo/tests/ssim/` fails forever until its
reference artefacts exist on the HF dataset
(`FastVideo/ssim-reference-videos`). The dataset hosts two kinds of artefacts
side-by-side per `(model_id, backend, prompt)`:
- **`.mp4`** — pixel ground-truth for tests that call
The extra `generated_videos/` level comes from the volume layout in
`_sync_generated_videos_to_volume` (`ssim_test.py`) — the command copies
`<repo>/fastvideo/tests/ssim/generated_videos/<tier>` to
`ssim_generated_videos/<tier>/<SUBDIR>/generated_videos/`, and `modal volume
get` preserves that trailing `generated_videos/` segment.
### 4. PAUSE — user reviews quality
Type-aware verification.
**For `ARTEFACT_TYPE = pixel`** — list the downloaded mp4s and ask the user to
open them in a video player:
> "Generated videos downloaded to `./generated_videos_modal/default/generated_videos/L40S_reference_videos/`. Please open them and confirm the quality looks correct. Reply **`upload`** to continue, or anything else to abort."
**For `ARTEFACT_TYPE = latent`** — `.pt` files are not human-watchable. Print
a numerics dump for each `.pt` so the user can sanity-check shape, distribution,
| 2026-04-21 | Post-first-run fixes: `modal volume get` needs `--force` when parent exists; download tree has an extra `generated_videos/` level so `--generated-dir` must reflect it. |
| 2026-05-01 | Latent (`*.pt`) artefact support: artefact-type detection in step 1, type-aware verification (visual eyeball for mp4, numerics dump for pt) in step 4, FSDP+inference_mode failure-mode added, design notes for the unified Modal flow. Triggered by PR #1253 (LTX-2 latent migration + Stable Audio latent test). |
"neg_prompt":"Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
},
"test_prompts":[
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
VALID="encoder vae transformer kernel unit ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
- Lint via `pre-commit run --files <changed paths>` (or `pre-commit run --all-files` for a full sweep) before committing. Do not shell out to `yapf`/`ruff`/`codespell`/`mypy` directly — pre-commit chains them with the project's config and respects the `.pre-commit-config.yaml` excludes (e.g. `fastvideo/tests/` is intentionally skipped). If pre-commit reports `(no files to check)` for your paths, that exclude is deliberate — don't bypass it.
- Target line length is 120 (configured in `pyproject.toml` for ruff, yapf, and isort).
- Naming: `snake_case` for functions/files, `PascalCase` for classes, `UPPER_SNAKE_CASE` for constants.
## Testing Guidelines
- Use `pytest` and place tests near relevant domains (e.g., `fastvideo/tests/encoders/`).
- Prefer descriptive names like `test_<feature>_<expected_behavior>.py`.
- For new pipelines/backends, include at least one regression-oriented test; add SSIM coverage when output quality must be preserved.
- Document GPU assumptions in tests that require specific hardware.
## Commit & Pull Request Guidelines
- Follow existing commit style: short subject with optional tag prefix, e.g. `[bugfix]: ...`, `[feat]: ...`, `[misc]: ...`, and include PR reference like `(#1234)` when applicable.
- Keep commits focused by concern (feature, refactor, fix).
- PRs should include:
- clear problem/solution summary,
- test evidence (`pytest`/SSIM outputs or rationale if skipped),
- linked issue/PR context,
- screenshots or sample outputs for UI/demo/docs changes.
## Agent Infrastructure
This repository is agent-friendly. Before doing any work, read:
1.`.agents/onboarding/README.md` — full onboarding guide with step-by-step instructions.
2.`.agents/memory/codebase-map/README.md` — structural index of the entire repository.
3.`.agents/skills/` — available agent skills (check if one exists before writing code).
4.`.agents/workflows/` — SOPs for common procedures (experiment lifecycle, evaluation, etc.).
5.`.agents/lessons/` — known pitfalls and their documented fixes.
If you are exploring a new procedure that has no existing SOP, document your
progress in `.agents/exploration/` and flag it for review at the end of your
session.
## Per-Directory AGENTS.md
Local guidance lives next to the code. Read the in-scope file before editing:
| Directory | What it covers |
|-----------|----------------|
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
-`2026/03/17`: Release Live demo: [Into the Dreamverse: Vibe Directing in FastVideo](https://dreamverse.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/dreamverse/).
-`2026/03/13`: Release Live demo: [Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU](https://1080p.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
-`2025/08/04`: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
### More News
-`2025/06/14`: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389).
-`2025/04/24`: [FastVideo V1](https://hao-ai-lab.github.io/blogs/fastvideo/) is released!
-`2025/02/18`: Release the inference code for [Sliding Tile Attention](https://hao-ai-lab.github.io/blogs/sta/).
## Key Features
FastVideo has the following features:
- End-to-end post-training support for bidirectional and autoregressive models:
- Support full finetuning and LoRA finetuning for state-of-the-art open video DiTs
- Data preprocessing pipeline for video, image, and text data
- Distribution Matching Distillation (DMD2) stepwise distillation.
- Sparse attention with [Video Sparse Attention](https://arxiv.org/pdf/2505.13389)
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
- User-friendly CLI and Python API
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/optimizations/) for full list of supported optimizations.
- Diverse hardware and OS support
- Support H100, A100, 4090
- Support Linux, Windows, MacOS
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
## Getting Started
We recommend using [uv](https://docs.astral.sh/uv/) to create a clean environment. If you previously used Conda, switching to uv generally gives faster and more stable installs.
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
## Sparse Distillation
For our sparse distillation techniques, please see our [distillation docs](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/) and check out our [blog](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/attention/vsa/#installation). Create a file called `example.py` with the following code:
# Create a video generator with a pre-trained model
generator=VideoGenerator.from_pretrained(
"FastVideo/FastWan2.1-T2V-1.3B-Diffusers",
num_gpus=1,# Adjust based on your hardware
)
# Define a prompt for your video
prompt="A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest."
# Generate the video
video=generator.generate_video(
prompt,
output_path="my_videos/",# Controls where videos are saved
save_video=True
)
if__name__=='__main__':
main()
```
## Prepare Data & Models
We've prepared some debug data to facilitate development. To make sure the training pipeline is correct, train on the debug data and make sure the model overfit on it (feed it the same text prompt and see if the output video is the same as the training data)
## Awesome work using FastVideo or our research projects
- [SGLang](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen): SGLang's diffusion inference functionality is based on a fork of FastVideo on Sept. 24, 2025.
- [DanceGRPO](https://github.com/XueZeyue/DanceGRPO): A unified framework to adapt Group Relative Policy Optimization (GRPO) to visual generation paradigms. Code based on FastVideo.
- [SRPO](https://github.com/Tencent-Hunyuan/SRPO): A method to directly align the full diffusion trajectory with fine-grained human preference. Code based on FastVideo.
- [DCM](https://github.com/Vchitect/DCM): Dual-expert consistency model for efficient and high-quality video generation. Code based on FastVideo.
- [HY-WorldPlay](https://github.com/Tencent-Hunyuan/HY-WorldPlay): An action-conditioned world model model trained using FastVideo framework.
- [Hunyuan Video 1.5](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5): A leading lightweight video generation model, where they proposed SSTA based on Sliding Tile Attention.
- [Kandinsky-5.0](https://github.com/kandinskylab/kandinsky-5): A family of diffusion models for video & image generation, where their NABLA attention includes a Sliding Tile Attention branch.
- [LongCat Video](https://github.com/meituan-longcat/LongCat-Video): A foundational video generation model with 13.6B parameters with block-sparse attention similar to Video Sparse Attention.
17. no cfg, validation no cfg, pcm_linear_quadratic, euler_steps 50, 0.1, linear_range 0.75
We welcome all contributions. Please check out our guide [here](https://hao-ai-lab.github.io/FastVideo/contributing/overview/).
See details in [development roadmap](https://github.com/hao-ai-lab/FastVideo/issues/899).
## Acknowledgement
We learned the design and reused code from the following projects: [Wan-Video](https://github.com/Wan-Video), [ThunderKittens](https://github.com/HazyResearch/ThunderKittens), [DMD2](https://github.com/tianweiy/DMD2), [diffusers](https://github.com/huggingface/diffusers), [xDiT](https://github.com/xdit-project/xDiT), [vLLM](https://github.com/vllm-project/vllm), [SGLang](https://github.com/sgl-project/sglang). We thank [MBZUAI](https://ifm.mbzuai.ac.ae/), [Anyscale](https://www.anyscale.com/), and [GMI Cloud](https://www.gmicloud.ai/) for their support throughout this project.
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.