Compare commits

...
956 Commits
Author SHA1 Message Date
SolitaryThinker 60054536d8 [evidence]: RL (DiffusionNFT) + VideoAlign + dataloader gates PASS — goldens, ledger, report 2026-07-23 14:27:33 -07:00
SolitaryThinker 4dcdb0d42c [docs]: PORT_NOTES — RL roles are FSDP2 mp bf16-compute over fp32 storage 2026-07-23 12:53:57 -07:00
SolitaryThinker b01bde8dca [fix]: NFTStep — modular stack IS fp32-master/bf16-compute (FSDP2 mp param_dtype=bf16); restore _MasterOpt chain + master-space old-EMA 2026-07-23 12:52:41 -07:00
SolitaryThinker eaedd5d39d [misc]: probe v3 — weight dtype/hash + op-input hashes for to_q/norm_q 2026-07-23 12:49:49 -07:00
SolitaryThinker 47a0bb6f2f [misc]: probe v2 — full tuple hashes + block0 call-sequence bisection 2026-07-23 12:42:22 -07:00
SolitaryThinker 641b8b210e [misc]: RL forward bisection probes (main vs v2.1, per-module hashes) 2026-07-23 12:34:05 -07:00
SolitaryThinker 1e9234c31b [docs]: PORT_NOTES — RL measured facts, tfv5 CLIP finding, dataloader gate 2026-07-23 12:18:51 -07:00
SolitaryThinker 3866db276b [feat]: vendor VideoAlign reward stack from PR #1476 (byte-identical) + MQ/VQ/TA scorers + gates 2026-07-23 12:15:18 -07:00
SolitaryThinker d95e17f1d9 [feat]: DiffusionNFT NFTStep aligned to modular stack (plain AdamW, autocast fwd, live-param EMA) + rl anchor mode 2026-07-23 12:06:47 -07:00
SolitaryThinker 635a9472ec [fix]: RL capture — staticmethod-safe decay, per-call row/timestep/adv records, param-dtype probe 2026-07-23 12:02:12 -07:00
SolitaryThinker 05060e8cd0 [feat]: v2.1 parquet dataloader port + order gate; RL capture PickScore-only CLIP unwrap 2026-07-23 11:55:22 -07:00
SolitaryThinker 14e0fc7a3a [fix]: RL capture — unwrap tfv5 CLIP features, disable step-0 validation 2026-07-23 11:50:54 -07:00
SolitaryThinker 1308740e1b [feat]: DiffusionNFT RL — NFTStep math module + modular-stack capture shim 2026-07-23 11:33:23 -07:00
SolitaryThinker ab3aed1dca [docs]: DiffusionNFT RL build facts 2026-07-23 11:30:25 -07:00
SolitaryThinker b0808d0935 [evidence]: dreamverse runtime gate pass — full protocol session, fMP4 streaming, live steps 2026-07-23 11:28:52 -07:00
SolitaryThinker fdddc5976b [feat]: Dreamverse runtime on fastvideo2 — session WS protocol, fMP4 segments, live per-step events, dummy-key enhancer; protocol gate 2026-07-23 11:26:47 -07:00
SolitaryThinker 8cc6abc107 [evidence]: serving identity gate pass — served latents bitwise vs offline, live per-step WS 2026-07-23 11:23:59 -07:00
SolitaryThinker b376207026 [fix]: register the stream route at starlette level (bypass fastapi ws DI) 2026-07-23 11:20:02 -07:00
SolitaryThinker d144277615 [debug]: ws entry prints 2026-07-23 11:18:21 -07:00
SolitaryThinker 7c12dba373 [fix]: surface ws setup exceptions to the client 2026-07-23 11:17:16 -07:00
SolitaryThinker e2fa9397f0 [fix]: drop return annotation on ws route 2026-07-23 11:10:46 -07:00
SolitaryThinker f40ae0c405 [fix]: pydantic-free request mapping (engine Request is the schema) 2026-07-23 10:58:38 -07:00
SolitaryThinker cbf1b8b54c [feat]: serving identity anchor — served latents bitwise vs offline + live WS step contract 2026-07-23 10:56:38 -07:00
SolitaryThinker 6a273121a1 [feat]: online serving — async-job REST + live per-step WebSocket over the engine (on_step hook); latents-sha identity 2026-07-23 10:55:58 -07:00
SolitaryThinker d708a3ffda [evidence]: self-forcing training parity — 121/121 rollout forwards bitwise; autocast-LN root cause; upstream grad-ckpt finding 2026-07-23 10:52:52 -07:00
SolitaryThinker 93153f8cbc [fix]: SF rollout forwards under autocast (plain-LN causal blocks are autocast-sensitive) 2026-07-23 10:51:08 -07:00
SolitaryThinker 1ba818b8b8 [fix]: register SF forward hook inside the distillation branch 2026-07-23 10:45:52 -07:00
SolitaryThinker 32bb106a0a [fix]: hoist fwd_seq before branch use 2026-07-23 10:41:52 -07:00
SolitaryThinker 7fe15c6368 [feat]: per-forward hash triage for SF gate 2026-07-23 10:37:33 -07:00
SolitaryThinker 9cbadec200 [fix]: neg_embeds optional in SF anchor (critic-only gate) 2026-07-23 10:35:20 -07:00
SolitaryThinker 9ea86848f6 [fix]: SF gate v1 critic-only (generator bwd requires grad-checkpointing upstream; mutated-cache recompute documented) 2026-07-23 10:30:41 -07:00
SolitaryThinker f748b617e4 [fix]: multi-element-safe randint recorders; SF train_one_step wrapper 2026-07-23 06:18:48 -07:00
SolitaryThinker 9120b535ed [feat]: self-forcing training gate — sequential draw replay through the causal serving model 2026-07-23 06:13:42 -07:00
SolitaryThinker aa2621ad3f [evidence]: QAD training parity — QAT rollout bitwise, first-run pass under dense band 2026-07-23 06:08:49 -07:00
SolitaryThinker a8205ea73b [feat]: QAD gate — QAT student + dense scorers + guidance 2.0 (legacy loader-gated recipe) 2026-07-23 06:04:56 -07:00
SolitaryThinker 8808976943 [evidence]: attn-QAT training parity — fake-quant forward bitwise, measured band 2026-07-23 06:03:53 -07:00
SolitaryThinker 088b9b45eb [feat]: attn-QAT port — vendored kernel wrapper (fail-closed), WanModelFVQAT, qat gate mode 2026-07-23 05:57:26 -07:00
SolitaryThinker d3ee4fb372 [evidence]: VSA+DMD2 parity — sparse student/dense scorers, measured band 2026-07-23 05:55:46 -07:00
SolitaryThinker 420ed38418 [fix]: vsa_dmd2 roles — VSA student on FastWan weights, dense scorers on base 2026-07-23 05:48:04 -07:00
SolitaryThinker 77afe2d747 [feat]: VSA+DMD2 gate — sparse student rollout, sparsity-0 VSA scoring (main's dense-copy semantics) 2026-07-23 05:45:52 -07:00
SolitaryThinker 26fe3b4226 [evidence]: DMD2 training parity — rollout bitwise pre-update, losses in 3-run band 2026-07-23 05:43:55 -07:00
SolitaryThinker a95f36ee6d [fix]: x0 hash rows informational after first student update 2026-07-23 05:35:02 -07:00
SolitaryThinker ed82305192 [fix]: __main__ after run_dmd2 definition 2026-07-23 05:32:46 -07:00
SolitaryThinker 4edfab8001 [feat]: DMD2 anchor — replay recorded draws through DMD2Step 2026-07-23 05:32:09 -07:00
SolitaryThinker c55eb09099 [fix]: load a text encoder in the shim for neg embeds (distillation has none) 2026-07-23 05:27:38 -07:00
SolitaryThinker 70facf77ae [fix]: use resolved training_args for neg-embed encoding 2026-07-23 05:25:10 -07:00
SolitaryThinker 64f17fde2f [fix]: pin negative embeds via main TextEncodingStage in dmd2 capture 2026-07-23 05:23:38 -07:00
SolitaryThinker c096554d36 [fix]: rec_dmd keyword signature 2026-07-23 05:20:24 -07:00
SolitaryThinker 60400b67ee [fix]: do not install finetune recorders in dmd2 mode 2026-07-23 05:18:24 -07:00
SolitaryThinker 271b6faa0c [fix]: record uncond embeds from batch dict (negative_prompt_embeds attr is optional) 2026-07-23 05:07:16 -07:00
SolitaryThinker 825cfa344a [feat]: DMD2 capture mode — per-phase RNG draw recording via method wrappers 2026-07-23 05:04:54 -07:00
SolitaryThinker fac7ac51e5 [docs]: DMD2 build facts in PORT_NOTES 2026-07-23 05:02:14 -07:00
SolitaryThinker afe9c077be [evidence]: VSA training parity — step0 bitwise, band from measured VSA self-noise 2026-07-23 04:59:54 -07:00
SolitaryThinker c73ccc75ed [feat]: VSA training gate — capture/anchor vsa mode, per-step ramped metadata through FinetuneStep 2026-07-23 04:50:49 -07:00
SolitaryThinker d96fd59525 [evidence]: finetune training parity vs main — step0 bitwise, later steps within measured self-noise band 2026-07-23 04:48:32 -07:00
SolitaryThinker e14afd78bf [feat]: train anchor gate — exact-where-deterministic, measured self-noise band elsewhere 2026-07-23 04:46:11 -07:00
SolitaryThinker c0fb861f0b [fix]: record post-normalization latents in train capture (normalize runs between fetch and prepare) 2026-07-23 04:42:29 -07:00
SolitaryThinker 1f43859e56 [feat]: train gate triage — pred hashes in capture, noisy/pred rows in anchor 2026-07-23 04:39:06 -07:00
SolitaryThinker cd22ac03db [fix]: gather DTensor param slice in train capture 2026-07-23 04:32:15 -07:00
SolitaryThinker fae777ec4b [feat]: finetune math trainer (fp32 masters + bf16 compute, main's exact chain) + train anchor; capture records all-step batches 2026-07-23 04:27:53 -07:00
SolitaryThinker 71aa5276a7 [feat]: training finetune capture shim (per-step goldens from main legacy pipeline) 2026-07-23 04:25:39 -07:00
SolitaryThinker d59f77787b [evidence]: SFWan bitwise vs fastvideo-main — 35/35 rollout forwards, goldens + pass record 2026-07-23 04:22:31 -07:00
SolitaryThinker 8b7285a845 [fix]: per-class accepted checkpoint _class_name (causal) 2026-07-23 04:17:33 -07:00
SolitaryThinker 1b52944f5a [fix]: sfwan golden dir name in anchor; [docs] training port authority map 2026-07-23 04:07:49 -07:00
SolitaryThinker bd00d273ad [feat]: SFWan port — vendored causal block/model (per-frame temb, KV+cross caches, fp64 rope at absolute positions), chunked causal DMD loop, card, sfwan capture+anchor 2026-07-23 03:43:22 -07:00
SolitaryThinker 9b83b7b385 [evidence]: FastWan QAD+VSA bitwise vs fastvideo-main — goldens, anchor pass records, numerics report 2026-07-22 20:21:13 -07:00
SolitaryThinker 0b332ff4e3 [fix]: renoise sigma must be non-0-dim so mixing runs in fp32 like main (0-dim tensors are scalars in promotion); parameterize fastwan anchor for vsa 2026-07-22 20:11:13 -07:00
SolitaryThinker f0b9c1773a [fix]: quantize fp8 weights lazily on the serving device (main converts post-GPU; CPU fp8 casts can round differently); ledger GateResult 2026-07-22 20:02:56 -07:00
SolitaryThinker 70beba1bc8 [feat]: FastWan VSA port — vendored tile math + kernel wrapper (fail-closed), VSA block/model, card; capture shim covers qad+vsa 2026-07-22 19:55:19 -07:00
SolitaryThinker 17a5d8400f [fix]: capture shim must not recast fp8-quantized model to bf16 2026-07-22 19:48:25 -07:00
SolitaryThinker 574bdc44dc [fix]: fp8 device-placement check; capture shim disables fsdp/offload like the released example; fastwan anchor CLI 2026-07-22 19:45:52 -07:00
SolitaryThinker 3b134a41fb [feat]: FastWan-QAD-FP8 port — vendored fastvideo-main forward, fp8 layer, main-faithful DMD loop, card, capture shim 2026-07-22 19:41:34 -07:00
SolitaryThinker 23176a48b6 [misc]: installable package — console script, evidence as package data, editable-install convention (no PYTHONPATH) 2026-07-22 18:59:31 -07:00
SolitaryThinker 9d7daad178 [feat]: WanDMDLoop — FastWan few-step x0+renoise sampler with main's exact sigma-table semantics (wan.dmd.x0renoise/v1) 2026-07-22 18:47:28 -07:00
SolitaryThinker 89ffc4b68d [docs]: examples/ — canonical SDK text-to-video script 2026-07-22 18:18:41 -07:00
SolitaryThinker a3aa600bd9 [misc]: wan21/gates/ — separate the measurement side (goldens, anchor, capture, compare, probe) from logic code 2026-07-22 16:25:43 -07:00
SolitaryThinker d1e2d2c974 [misc]: backend-band evidence — SDPA-policy anchor records alongside bitwise AUTO records 2026-07-22 16:20:03 -07:00
SolitaryThinker de2743d443 [feat]: attention backend registry — FLASH_ATTN/SDPA, FASTVIDEO2_ATTENTION_BACKEND override, fail-closed selection, policy in env fingerprint 2026-07-22 16:16:42 -07:00
SolitaryThinker dcec5644c2 [misc]: model_id is the only load key — HF repo strings are card ingredients, never resolvable 2026-07-22 16:06:59 -07:00
SolitaryThinker ac11912063 [feat]: identity policy — model_id primary, HF repos as fail-closed aliases (weights + component sources) 2026-07-22 16:04:47 -07:00
SolitaryThinker 97d653456e [feat]: offline SDK — Model/Result (modality-neutral handle over the gated engine path) 2026-07-22 15:54:30 -07:00
SolitaryThinker ec1e32e070 [misc]: layers-extraction equivalence evidence (anchor bitwise 0.0 on unchanged baseline) 2026-07-22 15:44:32 -07:00
SolitaryThinker 3a709f78cc [feat]: shared layers/ (norms, MLP, embeddings, rotary, attention) extracted from the Wan model — key- and cast-compatible 2026-07-22 15:42:14 -07:00
SolitaryThinker 0c824beb7b [misc]: aligned evidence — bitwise DiT anchor (0.0 bf16+fp32), full ladder green, aligned 50-step sample 2026-07-22 15:34:45 -07:00
SolitaryThinker a37198afb7 [fix]: attention fallback keyed on autocast state, not q dtype — official casts fp32-promoted q/k to bf16 and runs flash 2026-07-22 15:28:31 -07:00
SolitaryThinker ee18433a6e [fix]: split storage vs compute dtype — DiT stores fp32 like official, computes bf16 via provenance.precision autocast 2026-07-22 15:25:50 -07:00
SolitaryThinker 2000e8a0ff [feat]: align the DiT with official — vendored WanModel (pinned 9737cba) + official weights + official autocast regime 2026-07-22 15:20:45 -07:00
SolitaryThinker 3333d959c0 [misc]: fp32 exact-math golden + final 3-way numerics report + decomposition ledger records 2026-07-22 15:03:17 -07:00
SolitaryThinker 7db862e6ff [misc]: dit_fp32 gate bound above the measured 2.6e-3 diffusers-vs-official fp32 delta (suspected RoPE; growth fails) 2026-07-22 15:01:38 -07:00
SolitaryThinker 787e17f500 [fix]: sdpa-fp32 exact-math golden (flash helper force-casts bf16); isolated anchor probes; robust report over record unions 2026-07-22 14:59:23 -07:00
SolitaryThinker 93f80b5c4b [feat]: fp32 exact-math anchor row (dit_fp32 golden + adapters); document load-bearing 512 zero-pad context 2026-07-22 14:55:43 -07:00
SolitaryThinker b57f3ad47d [fix]: probe sdpa shim k_lens device 2026-07-22 14:52:52 -07:00
SolitaryThinker a31f19ab5a [misc]: DiT delta decomposition probe (MATH/PAD/CAST/ATTN) + fp32 exact-math golden capture 2026-07-22 14:49:15 -07:00
SolitaryThinker 194d27220f [misc]: official Wan2.1 goldens (commit 9737cba) + 3-way numerics report + anchor ledger records 2026-07-22 14:39:11 -07:00
SolitaryThinker b28b01e405 [fix]: official text cleaning (ftfy width-fold retokenizes CJK; anchor caught rel 0.42 on negative prompt); fastvideo_main VAE device; resilient compare 2026-07-22 14:35:54 -07:00
SolitaryThinker 8e17148bd1 [misc]: pin py3.12/torch2.12; goldens captured in the fastvideo env by policy (env_policy in manifest) 2026-07-22 14:25:34 -07:00
SolitaryThinker 1387a9c5e7 [feat]: official-goldens anchor — capture shim, per-impl adapters, verify --anchor, 3-way numerics comparison 2026-07-22 14:18:57 -07:00
SolitaryThinker ad1dd41c99 [docs]: tighten card docstring to standard terms (serializable, content-addressed) 2026-07-22 13:54:56 -07:00
SolitaryThinker ceb5a24a30 [docs]: spell out what JSON round-trip and the content digest guarantee (and don't) on cards 2026-07-22 13:51:01 -07:00
SolitaryThinker ae99a82ccf [misc]: GB200 evidence — blessed fingerprints, T0-T3 ledger records, 50-step sample (md5-identical to reference) 2026-07-22 07:45:26 -07:00
SolitaryThinker 89a55be5f5 [fix]: card-declared compute dtypes — preserve loader fp32 islands, never introspect mixed-dtype modules 2026-07-22 07:39:09 -07:00
SolitaryThinker 964769efb7 [fix]: force card-declared dtypes after loading (torch_dtype renamed upstream; loader default must not win) 2026-07-22 07:34:39 -07:00
SolitaryThinker e10d45d8be [feat]: fastvideo2 MVP — data cards, driven loop, enforced pipelines, wan2.1 reference oracle, tiered verifier + evidence ledger 2026-07-22 07:24:24 -07:00
SolitaryThinker 7cd5d02104 [misc]: strip repo for the v2.1 clean-slate MVP branch 2026-07-22 07:24:14 -07:00
pkisfaludi-nvandClaude Opus 4.8 9fb74b9732 Make LTX-2 RMSNorm out-of-place so torch_tensorrt + Ulysses SP compiles (#1623)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 04:32:04 -07:00
Adriel FungandSolitaryThinker 521dee0e82 [perf]: enable per-block torch.compile for LTX2 with persistent cache (#1602)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-20 15:59:13 -07:00
William Lin 229419208e [ci]: opt in to fork-PR head checkout after actions/checkout guard change (#1625) 2026-07-20 13:12:22 -07:00
William Lin 65f3b946b9 [feat] Add LTX-2 and LTX-2.3 fine-tuning to the modular trainer (#1624) 2026-07-20 11:58:39 -07:00
Zhang Peiyuan 191fcbf46c [feat] add qat docs (#1621) 2026-07-19 18:37:47 -07:00
Junda Su 755a4e4470 [new-model] Add LingBot-Video Dense and MoE/refiner T2V inference (#1595) 2026-07-18 20:18:16 -07:00
9709b7513b [feat] Port NVFP4 QAT/QAD to modular train framework (#1619)
Co-authored-by: Peiyuan Zhang <email>

Co-authored-by: Peiyuan <a>
2026-07-18 15:01:09 -07:00
William Lin 32cd603515 Revert docs trusted-branch-only workflow (#1618) 2026-07-17 00:22:56 -07:00
Junda Su d4bdd3621a [new-model] Port LingBot-World-v2 (#1579) 2026-07-16 19:51:53 -07:00
William Lin e2f8322842 [ci]: skip unused Buildkite submodule checkout (#1614) 2026-07-16 18:37:57 -07:00
6966f9e0bc [fix] Z-Image (#1236) draft port: rebase + strict-load contract + bf16 encoder parity + PORT_STATUS (#1339)
Co-authored-by: Mrinaal Dogra <mdogra@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-16 18:18:25 -07:00
Mac Lee 743f4ed5f9 [ci]: harden Modal repository checkout (#1590) 2026-07-16 18:05:50 -07:00
Mac Lee 1c04ace573 [ci]: pin VSA training regression to H100 (#1591) 2026-07-15 22:04:57 -07:00
Mac LeeandSatyam Srivastava 6cbff73687 [bugfix]: skip unused output materialization (#1567)
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
2026-07-16 04:03:07 +00:00
William Lin dec8b10939 [docs]: document automatic Docker image builds (#1608) 2026-07-15 18:52:21 -07:00
William Lin 133a5278af [ci] Run docs only for trusted PR branches (#1610) 2026-07-15 18:51:59 -07:00
William Lin da856274cc [feat]: FastVideo Studio — SvelteKit → Next.js port + review fixes (#1612) 2026-07-15 18:51:33 -07:00
Mac LeeandSolitaryThinker a253856147 [ci] Add exact identity performance statuses (#1560)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-15 14:59:09 -07:00
Satyam Srivastavaandgemini-code-assist[bot] 6e25d94ebc [ci]: enable scheduled perf runs to update rolling baseline (#1599)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-07-14 16:55:12 -07:00
Mac Lee cae8fa18dc [bugfix]: propagate Qwen2.5-VL visual dtype (#1580) 2026-07-13 18:37:11 -07:00
William Lin 821e5a0832 [bugfix]: fix FlashAttention resolver tests after tuple return (#1597) 2026-07-13 16:29:03 -07:00
William Lin c1abc42782 [bugfix]: allow unrestricted head sizes in SDPA (#1596) 2026-07-13 16:04:51 -07:00
William Lin ef15ea2391 [bugfix]: keep LTX2 rms_norm outputs bf16 under torch 2.12 autocast (#1587) 2026-07-13 16:04:21 -07:00
Mac Lee 1ea2517e22 [ci]: extend LoRA training CI timeout (#1589) 2026-07-13 02:47:28 -07:00
MookandSolitaryThinker 0c63528c59 [perf] Cache RoPE position-embedding tables across denoising steps (#1442)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-13 02:34:51 -07:00
b063f8ca41 [feat] Fix FLUX.1-dev port: native RoPE, parity tests, SSIM reference (#1321)
Co-authored-by: Ishan Vaish <ivaish@ucsd.edu>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 01:49:42 -07:00
William Lin e7fff0173a [bugfix]: benchmark_weight_loading_comparison.py — iterate safe_open via .keys() (#1378) 2026-07-12 22:50:45 -07:00
Shreejith SGandH1yori233 d82abc271e [feat] Add GLM-Image inference support (#1030)
Co-authored-by: H1yori233 <k1kong@ucsd.edu>
2026-07-12 22:42:42 -07:00
Guian FangandSolitaryThinker 970409962f [feat] Add AnyFlow any-step video distillation (pretrain + on-policy) (#1371)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-12 02:32:08 +00:00
Raghav K 055586703d [perf]: register a real backward for FA2 default + masked/varlen custom ops (training-under-compile) (#1388) 2026-07-12 00:45:27 +00:00
Mac Lee 5d89f86675 [ci] Stop forcing FA4 in model-load lanes (#1561) 2026-07-11 14:10:36 -07:00
Satyam Srivastava 19a51a1fe6 [ci] Trigger performance benchmarks for performance code changes (#1583) 2026-07-10 20:21:33 -07:00
William Lin d3232cea5a [ci]: gate the full-suite trigger on pre-commit and docs build (#1572) 2026-07-11 02:56:49 +00:00
Raghav KandSolitaryThinker 0c90c8c24d [bugfix] nvfp4: cast fp32 inputs to bf16 instead of asserting (#1488)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-10 21:39:17 +00:00
Mingjia HuoandClaude Fable 5 4c08ffce49 [feat] World model training using third person games (#1443)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 05:09:07 +00:00
Atharv Ramesh af4a77553c [ci]: add SSIM reference bootstrap flow (#1522) (#1547) 2026-07-10 01:49:38 +00:00
alexzmsandSolitaryThinker c096fda1eb [docs] Add LTX-2.3 distilled inference run configs (t2v/i2v × 5+2/8+3 × resolutions) (#1568)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-09 18:53:12 +00:00
William Lin 8f47e85be0 [bugfix]: retry remote image downloads in load_image (#1570) 2026-07-09 07:33:24 -07:00
William Lin 90d3bd19eb [infra] Deliver per-job Buildkite env to Modal CI at runtime, not as image layers (#1569) 2026-07-09 06:33:37 -07:00
Mac LeeandSolitaryThinker afb4f7d3c5 [ci]: emit v2 performance result schema (#1551)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-09 06:11:15 +00:00
02e1143f22 [feat] Add Kandinsky-5 T2V/I2V pipeline support (#1471)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: leffff <levnovitskiy@gmail.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 14:43:30 -07:00
KaredandSolitaryThinker e2f4d1a7b5 [feat]: add SwanLab tracker (#1461)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 08:35:38 +00:00
Kaiqin KongandSolitaryThinker f037351146 [feat] Add Clean-history Teacher Forcing and Causal Consistency Distillation (#1505)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 08:13:11 +00:00
Mac LeeandSolitaryThinker 1ee11e08dc [ci]: add performance fingerprint cohorts (#1546)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 07:24:00 +00:00
595f0ea60e [feat] Add DreamX-World 5B Cam and AR pipelines (#1538)
Co-authored-by: Suckl <Suckl@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-07 06:16:18 +00:00
William Lin d921832cd2 [misc]: reserve CPU/memory for timing-sensitive Modal CI lanes (#1566) 2026-07-06 22:13:53 -07:00
William Lin 629697629a [bugfix]: free CUDA memory between train-framework model tests (LongCat OOM on L40S) (#1565) 2026-07-06 21:27:30 -07:00
William Lin a25313beec [ci]: wire fastvideo/tests/ops/ into the unit-test lane (#1559) 2026-07-06 12:07:26 -07:00
William Lin dbde64385b [bugfix]: bump FA4 pin to the CuTe DSL 4.6 compatible rev (#1564) 2026-07-06 12:06:59 -07:00
William Lin 9d909f5f04 [test]: remove dead and duplicate tests (-489 lines) (#1556) 2026-07-05 15:53:40 -07:00
William Lin 76b0550c15 [ci]: run pre-commit on fork PRs without manual approval (#1555) 2026-07-05 14:18:16 -07:00
William Lin 384c1e9493 [misc]: update reseed-performance-baseline skill for the hf_store move (#1545 follow-up) (#1553) 2026-07-05 14:16:55 -07:00
William Lin b1dbcc93f6 [misc]: reformat fastvideo/performance to the repo yapf config (#1554) 2026-07-05 14:16:20 -07:00
William Lin b93833772e [ci]: guard against test directories no CI lane collects (#1552) 2026-07-05 14:07:38 -07:00
Mac Lee 30b523edd6 [ci] Normalize performance stage component metrics (#1475) (#1550) 2026-07-05 14:05:26 -07:00
Mac LeeandSolitaryThinker 6aab7f3832 [ci] cover Hunyuan 1.5 chat-list text preprocessing (#1518)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-05 12:05:25 -07:00
Mac Lee 9cd53fe5f8 [ci] Add metric-specific performance thresholds (#1545) 2026-07-05 12:04:33 -07:00
Mac Lee 6a32cf3a5e [ci]: expose LoRA extraction slash command (#1542) 2026-07-05 06:45:31 -07:00
William Lin 98be9b3da2 [bugfix]: address the three remaining #1447 review findings (#1549) 2026-07-05 06:43:27 -07:00
Mac Lee c53e85b767 [ci] Add v2 performance benchmark config identity fields (#1544) 2026-07-05 06:18:15 -07:00
William Lin 40a8bd2d3b [bugfix]: skip ThunderKittens kernels on aarch64 and document the kernel build matrix (#1548) 2026-07-05 06:13:11 -07:00
Mac LeeandSolitaryThinker 31aa115611 [bugfix]: preserve FSDP hooks for RMSNorm qk norms (#1513)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-05 04:15:49 +00:00
zainnhandSolitaryThinker 98ac10a528 [infra] Auto-rebuild CUDA images when docker/Dockerfile changes on main (#1526)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-04 16:25:46 -07:00
William Lin a5a6d171e5 [attn] Make FA4 explicit opt-in via FASTVIDEO_FA4 and delete the runtime fallback machinery (#1540) 2026-07-04 14:51:00 -07:00
William Lin 51ed1ea423 [build] Bump fastvideo-kernel pin to 0.3.2 (#1541) 2026-07-03 14:48:15 -07:00
Mac Lee 00ec3e7388 [bugfix]: compute VSA topk from padded blocks (#1517) 2026-07-03 14:16:51 -07:00
William Lin 0d626ef2d1 [kernel] Bump fastvideo-kernel pin to 0.3.1 and version the FA4 tile_mn port as 0.3.2 (#1539) 2026-07-03 14:11:08 -07:00
William Lin b36d0ef085 [infra] Add DGX Spark and multi-architecture CUDA support (#1447) 2026-07-03 13:41:43 -07:00
William Lin 10c8c5df4d [misc] reorg: relocate ui/ and performance_dashboard/ under apps/ (#1537) 2026-07-02 20:29:17 -07:00
William Lin 31e26abec4 [chore] release fastvideo-kernel 0.3.1 (#1520) 2026-06-30 12:25:27 -07:00
sumyyyyyandSolitaryThinker 9e83ba630c [kernel] Extract VSA utility functions into fastvideo_kernel (#1408)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-30 12:19:25 -07:00
William Lin fc02a9ce8e [ci] aarch64 kernel wheel: build for Blackwell (sm_100a/sm_120a), not Hopper (#1516) 2026-06-29 16:10:42 -07:00
William Lin 3d36160fc4 [bugfix] -fsigned-char so ThunderKittens compiles on aarch64 (Grace Hopper) (#1515) 2026-06-29 14:05:38 -07:00
William Lin a11ec43de2 [ci] build + publish aarch64 (Grace Hopper) kernel wheels alongside x86_64 (#1514) 2026-06-29 12:28:34 -07:00
William Lin 2656d6530c [bugfix] compile FP4 (attn_qat_infer) kernels for sm_120a only via per-arch split (#1508) 2026-06-29 11:07:27 -07:00
Kevin Lin e3f54e7169 [bugfix] Fix causal attention mask for Blackwell FP4 MMA column layout (#1506) 2026-06-28 20:36:40 -07:00
William Lin 4ba5681307 [bugfix] guard ThunderKittens Hopper kernels for non-sm_90a device passes (#1507) 2026-06-28 20:07:43 -07:00
alexzms 8c23c86994 [docs]: add raw-video preprocess script for the QAD MixKit recipe (#1487) 2026-06-28 18:01:13 -07:00
William Lin 78d606a3a7 [feat]: make VSA tile cache configurable for training (#1444) 2026-06-28 02:18:31 -07:00
William Lin c4e108de78 [docs]: modernize documentation build (#1503) 2026-06-28 02:02:09 -07:00
Kaiqin Kong 16bf2eaf77 [feat] Relativistic RoPE re-indexing for long causal rollouts (#1454) 2026-06-28 01:43:00 -07:00
William Lin 1d15d974fa [misc]: cleanup outdated or unneeded agent infrastructure (#1504) 2026-06-27 19:37:49 -07:00
alexzms 5e57868b76 [ci] train-framework model coverage: Cosmos/MatrixGame2 finetune grad-norm + Cosmos/LongCat/MatrixGame2 loading smokes (#1497) 2026-06-27 17:16:00 -07:00
William Lin 6be280c914 [bugfix]: fix docs build (#1502) 2026-06-27 13:00:09 -07:00
KyleNeverGivesUpandClaude Opus 4.8 3ccdec9798 [bugfix] Make enable_torch_compile_vae actually compile the Wan VAE (#1498)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 12:30:59 -07:00
Satyam Srivastava 8658f774f3 [docs] Restructure contributing CI/CD and testing docs (#1501) 2026-06-27 11:46:18 -07:00
Raghav K 4f3ad3f6df [perf] Default Wan VAE decode to bf16 (lossless, faster) (#1472) 2026-06-26 12:42:29 -07:00
William Lin b454aa56c3 [misc] fix pre-commit (#1500) 2026-06-26 12:39:28 -07:00
William Lin f9b8e30ff3 Update README.md 2026-06-26 09:23:30 -07:00
William Lin 0356205b84 [bugfix] revert fastvideo kernel version to 0.2.6 (#1495) 2026-06-25 03:52:43 -07:00
Kevin Lin 719a1879bd [bugfix] Remove incorrect V-row permutation in scaled_fp4_quant_trans_kernel (#1493) 2026-06-25 03:05:12 -07:00
Utkarsh RanjanandClaude Opus 4.8 b57180bf97 [bugfix] Warn when a requested attention backend is unsupported by a layer (#1254) (#1486)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 15:56:39 -07:00
William Lin 7cebf5f82c [ci] cap kernel wheel build parallelism to avoid runner OOM (#1483) 2026-06-23 14:17:56 -07:00
William Lin dd0f4b6753 [misc] cleanup misc files (#1484) 2026-06-23 14:17:22 -07:00
William Lin d303b4e03a [docs] update README (#1482) 2026-06-23 12:07:21 -07:00
xsankandmergify[bot] 31b719ae49 [perf] optimize compress & topk kernel (#1421)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 18:51:22 +00:00
William Lin 887aaf3d3e [ci] bump cuda-toolkit action to v0.2.35 to fix kernel cu130 publish (#1481) 2026-06-23 10:33:31 -07:00
William Lin b1d89eba1f [chore] release fastvideo-kernel 0.3.0 (#1478) 2026-06-23 10:02:27 -07:00
Loay RashidandSolitaryThinker 995a5fdf97 [bugfix] fixing denoising time in the fastwan script (#1480)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 10:01:23 -07:00
William Lin 70a70b689e [ci] mergify: stop auto-syncing ready PRs (#1477) 2026-06-23 03:23:58 -07:00
Shao Duan 4d6ac89b43 [ci] eval: add metric regression + identity-invariant ci tests (#1451) 2026-06-23 01:27:34 -07:00
Raghav K 4171cacd93 [bugfix] Wire Cosmos-Predict2.5 2B to its sampling preset (#1468) 2026-06-23 01:25:35 -07:00
b2ade71467 [Bugfix] QAD 5090: Torch.compile and other optimizations (15/12) (#1466)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 01:23:57 -07:00
Kevin Lin 82ed9fe58d [feat] QAD 5090: FP8 linear layer inference (#1465) 2026-06-22 18:25:06 -07:00
Satyam Srivastava 3d8cc4f0a0 [bugfix] Fix performance component timing extraction (#1473) 2026-06-22 13:05:35 -07:00
Satyam Srivastava 0557f7a7d9 [ci] Add performance dashboard metadata and visualizations (#1470) 2026-06-19 14:27:01 -07:00
dc66cd97ef [feat] QAD 5090: FP8 QAT linear training (14/12) (#1464)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-19 01:03:51 -07:00
6da206e196 [feat] QAD 5090: FP4 QAT linear STE for training (13/12) (#1463)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:55:47 -07:00
sumyyyyy 87f98c9b8b [kernel] Add varlen support for block-sparse attention (#1319) 2026-06-18 01:06:00 +00:00
e60601df7f [feat] QAD 5090: QAT training recipe — finetune + DMD distillation (12/12) (#1462)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 15:18:52 -07:00
eed9c4bfbf [kernel] QAD 5090: Add Attn-QAT training Triton kernels (11/12) (#1460)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
2026-06-17 13:35:12 -07:00
1dee77f4a4 [feat] QAD 5090: Wire the Attn-QAT training attention backend (10/12) (#1459)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 13:32:56 -07:00
Satyam Srivastava b80148819c [ci] Add performance dashboard visualisation scripts (#1469) 2026-06-16 23:29:29 -07:00
c3b971488e [docs] QAD 5090: Add NVFP4 + Attn-QAT inference example and how-to (9/12) (#1458)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 14:36:50 -07:00
88e753f281 [feat] QAD 5090: Wire the Attn-QAT inference attention backend (8/12) (#1457)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 13:12:33 -07:00
77832059cc [kernel] QAD 5090: Add modified SageAttention3 FP4 inference kernels (7/12) (#1455)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
2026-06-16 12:02:30 -07:00
alexzmsandmergify[bot] 633d393568 [ci] layer-0 grad-norm regression for per-method training tests (5a-ii) (#1396)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 04:45:07 +00:00
Junda SuandPeiyuan Zhang 5854aec2ce [feat] Add Wan RL DiffusionNFT training (#1450)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2026-06-11 21:18:59 -07:00
Mook 30e45c2411 [bugfix] Classify new config/sampling fields in schema parity inventory (#1446) 2026-06-10 13:38:36 -07:00
alexzms 2a4fe697a6 [docs] LTX-2.3 distilled i2v: typed-API example (from_config + generate) (#1448) 2026-06-10 10:58:14 -07:00
alexzms 921db7479d [perf] LTX-2.3 distilled i2v: drop max-autotune from compile kwargs (#1445) 2026-06-10 10:06:41 -07:00
Aryan KumarandAryan Kumar 7f539424cb [feat]: add Lucy Edit inference scaffold (#1363)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-06-09 15:21:43 -07:00
19a838f54f [bugfix]: release VSA tile cache during training (#1434)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 15:01:22 -07:00
d922ab2cbc [model] Flux2 Klein Port (#1349)
Co-authored-by: Gnav3852 <63612880+Gnav3852@users.noreply.github.com>
Co-authored-by: Mac Lee <macthecadillac@gmail.com>
2026-06-09 14:55:55 -07:00
Kaiqin Kongandmergify[bot] 9ea77d37f3 [bugfix] EMA shadow on resume and EMA under MoE path (#1441)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 00:55:43 +00:00
2e35b0c6bd [refactor]: linear/mlp FP4 path additions for Wan-2.1 (Attn-QAT 6/12) (#1390)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
Co-authored-by: Matthew Noto <notomatthew31@gmail.com>
2026-06-08 14:51:57 -07:00
William Lin 1c627a3f98 [bugfix]: build fastvideo-kernel on GPU-less Docker runners (#1437) 2026-06-06 17:52:45 -07:00
Kaiqin Kong a931efe33a [bugfix] EMA in distillation pipeline (#1440) 2026-06-06 17:38:03 -07:00
William Lin 041e5e9029 [bugfix]: unblock PyPI publish (flash-attn-cute direct dep) (#1436) 2026-06-05 10:07:12 -07:00
Kaiqin Kongandmergify[bot] efcc245c2e [feat] VLM as judge for WM (#1429)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 04:25:54 +00:00
alexzmsandmergify[bot] 922e7e0813 [docs] LTX-2.3 distilled i2v example with compile + timing breakdown (#1430)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 01:26:50 +00:00
William Lin c62a8514b0 [chore]: release v0.2.0 (#1432) 2026-06-04 14:22:49 -07:00
William Lin 3eb8081801 [chore]: unpin runtime deps in pyproject.toml (#1431) 2026-06-04 13:46:49 -07:00
Raghav K 3505d09564 [bugfix] tests: include ltx2_3_base in expected LTX2 preset set (#1427) (#1428) 2026-06-04 12:17:18 -07:00
Shao DuanandSolitaryThinker 570607945c [feat] dreamverse: sequence parallelism for serving (#1424)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 23:13:49 -07:00
Kaiqin KongandSolitaryThinker d3a821cdcf [feat] LoRA controls and integration for Dreamverse (#1420)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 18:54:11 -07:00
Kaiqin Kong 3f24578139 [bugfix] LTX2: honor video_position_offset_sec in the DiT (#1422) 2026-06-01 16:59:15 -07:00
Kevin Lin 89fcf08378 [bugfix] Fix STFT dtype mismatch (#1419) 2026-05-31 20:56:07 -07:00
c2930b2aa1 [ci] Add additional Dreamverse UI tests (#1417)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-31 19:53:29 -07:00
KUAN-HAO HUANGandSolitaryThinker 019239690b [perf] Add Adaptive Guidance (CFG gating) for stale-uncond reuse (#1372)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-30 13:58:55 -07:00
alexzms d6119c1f82 [bugfix]: dreamverse modal bypasses ENTRYPOINT — set ffmpeg env + key check (#1413) 2026-05-29 16:50:01 -07:00
alexzms 84214c80bb [feat] LTX-2.3 audio: BWE vocoder path (#1398) 2026-05-29 16:43:26 -07:00
alexzmsandmergify[bot] c2d7143c72 [feat] LTX-2.3 transformer support (config-gated extension of LTX-2) (#1397)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 16:42:37 -07:00
Kaiqin KongandSolitaryThinker afdb6fbfa5 [feat] Add MatrixGame3.0 (#1201)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-27 21:07:01 +00:00
alexzms ba4c02d883 [docs]: highlight Dreamverse deployment paths + add Server B200 (SSH) guide (#1409) 2026-05-27 09:30:58 -07:00
Junda Su 2c137931f3 [bugfix] Fix Dreamverse Modal compile warmup latency (#1394) 2026-05-26 17:58:38 -07:00
Shao Duan 36682797a0 [feat] eval: input ergonomics + Evaluator features + bug fixes (#1392) 2026-05-26 13:59:45 -07:00
William Lin 0ef1357a77 [docs]: surface activation-trace utility in add-model skills (#1399) 2026-05-26 13:45:12 -07:00
William Lin a75d19786a [docs]: Wire activation trace into mkdocs nav + perf/troubleshooting (#1304) 2026-05-26 13:14:36 -07:00
alexzms be548a78ea [feat] VSA-256 fastpath on Blackwell via FA4 CuTe block-sparse attention (#1354) 2026-05-26 12:58:29 -07:00
alexzms 6a610e2bc9 [ci] add per-method single-step training tests for fastvideo.train (#1343) 2026-05-25 17:09:54 -07:00
Shao Duanandabaghyangor 321d5112b4 [refactor] eval: consolidate FVD into common.fvd, remove benchmarks/fvd (#1380)
Co-authored-by: abaghyangor <abaghyangor@gmail.com>
2026-05-24 12:09:47 -07:00
William Lin ba75ad82db [refactor]: shared attention infra additions for QAT-compat (Attn-QAT 5/12) (#1383) 2026-05-23 16:19:24 -07:00
Junda Su 58caa5109f [ci] Add DreamVerse app CI tests (#1386) 2026-05-23 14:28:26 -07:00
2f3ca8aaad [perf]: register FA2/FA3 default flash_attn_func as a torch.library custom op (#1373)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-23 00:39:06 -07:00
Junda Su 3a67319cb6 [infra] Use npm for Dreamverse web builds (#1385) 2026-05-22 18:50:10 -07:00
Junda Su 266fa044b3 [infra] Add Dreamverse Modal UI image build (#1381) 2026-05-22 11:36:53 -07:00
William Lin f3398db868 chore: pin dreamverse npm deps to address Dependabot alerts (#1359) 2026-05-22 11:29:33 -07:00
fda02036bc [feat]: Attn-QAT inference + training backends (deadcode) (Attn-QAT 4/12) (#1358)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-22 02:59:25 -07:00
Raghav K 68179cd752 [docs] Document enable_torch_compile (+ A/B example) (#1366) 2026-05-22 00:25:42 -07:00
Satyam SrivastavaandSolitaryThinker 2dd5760291 [docs] Document performance benchmark workflow (#1376)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 20:25:58 -07:00
Wenxuan TanandSolitaryThinker af2ee9c78a [feat] Optimize distributed weight loading in multi-node training (#572)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 16:59:48 +00:00
Satyam SrivastavaandSatyam Srivastava 1c80371b27 [ci] Component time performance + reseed hf baseline skill (#1292)
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
2026-05-20 13:51:11 -07:00
Junda Su eef473225d [ci] Add Dreamverse Docker image workflow (#1369) 2026-05-20 13:50:02 -07:00
Junda Su 44fb84ef6a [bugfix]: shrink Dreamverse Docker context (#1368) 2026-05-20 13:38:52 -07:00
e8597b7448 [Bugfix] FP4 FA4 installation fix (#1367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-19 22:45:24 -07:00
Raghav Kandmergify[bot] e2252c0a5e [perf] Mark LayerwiseOffloadHook entry points torch.compiler.disable (remove per-layer graph break) (#1365)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-20 03:09:23 +00:00
Junda Su 72cb427cd9 [feat]: add FastLTX-2.3 Gradio demo package (draft) (#1247) 2026-05-17 23:03:16 -07:00
Mingjia Huo 63030cf6ec [fix] Fix causal self-forcing attention settings (#1355) 2026-05-17 22:59:47 -07:00
Kaiqin Kongandmergify[bot] 773d44b875 [misc] Rename MatrixGame to MatrixGame2 (#1357)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 22:57:03 -07:00
1df513922f [feat] Add minimal LoRA finetuning support to the YAML training stack (#1242)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 14:49:13 -07:00
William LinandDavids048 30c45620a2 [infra] [dreamverse]: add instruction to install nasm and update ffmpeg installer to work in plain venv (#1361)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
2026-05-17 01:10:55 +00:00
Shao Duanandklhhhhh 6b2c731596 [feat] eval: add audio metrics (#1352)
Co-authored-by: klhhhhh <1412841649@qq.com>
2026-05-16 14:47:37 -07:00
William Lin e6022c20b2 [misc]: demote ROCm-unavailable startup message to DEBUG (#1360) 2026-05-15 22:12:16 -07:00
460f6e398e [feat]: Add NVFP4QAT linear layer (Attn-QAT 3/12) (#1350)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-15 18:42:02 -07:00
d2ffec5cce [perf] Dreamverse 14/14: Add LTX2 profile speedups (#1337)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 18:10:47 -07:00
alexzmsandmergify[bot] cb12e88713 [perf] shallow-copy VSA attn_metadata in train model plugins (#1342)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-15 16:10:12 -07:00
0958c344b8 [feat] Dreamverse 13/14: Activate LTX2 integration (#1336)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 15:35:05 -07:00
Junda Su 71dc27ea7b [misc]: Add Dreamverse deploy skill frontmatter (#1353) 2026-05-15 15:23:17 -07:00
Junda SuandSolitaryThinker 1263449d2a [docs] Add copy page action (#1351)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-15 14:59:04 -07:00
b4de5a9f1b [feat] Dreamverse 12/14: Add LTX2 refine and upsampler support (#1335)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 12:21:40 -07:00
8acd8e21f9 [feat]: Add NVFP4QAT quantization config (Attn-QAT 2/12) (#1348)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-14 16:53:48 -07:00
d45d82334f [infra] Dreamverse 11/14: Add NVFP4 quantization support (#1334)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 16:26:39 -07:00
6392bd40a9 [feat] Dreamverse 10/14: Add serving API contracts (#1333)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 10:07:01 -07:00
Shao Duan 17f07bc313 [feat] eval: async VideoPool + metric streamlines (#1320) 2026-05-13 15:57:37 -07:00
William Lin 325861fb99 [misc]: PR-1225 sync — housekeeping (1/12) (#1347) 2026-05-13 15:21:52 -07:00
William Lin c5088670c8 [misc]: empty __init__.py files with no logic (#1346) 2026-05-13 13:09:40 -07:00
0403c6f47e [infra] Dreamverse 09/14: Add Docker and launch scripts (#1332)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 18:01:11 -07:00
a55b1cdde4 [feat] Dreamverse 08/14: Add frontend media and E2E coverage (#1331)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 17:45:05 -07:00
aec9a20a16 [feat] Dreamverse 07/14: Add frontend session UI (#1330)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 17:22:45 -07:00
b4458e5bae feat: FP4 Flash Attention 4 for Blackwell GPUs (#1221)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 17:04:39 -07:00
Kaiqin Kongandmergify[bot] a790153705 [bugfix] MatrixGame2 SF distillation under gradient checkpointing (#1340)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 17:01:04 -07:00
alexzmsandmergify[bot] 0df1445d0d [ci] add GPU model loading tests for fastvideo.train (PR 4/9) (#1274)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 16:54:37 -07:00
038da6e02b [feat] Dreamverse 06/14: Add frontend scaffold (#1329)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 16:44:13 -07:00
William Lin 3fb2fbe1a2 [infra]: MagiHuman checkpoint conversion + push scripts (7/8) (#1301) 2026-05-12 15:45:15 -07:00
William Lin acea9d23e1 [docs]: MagiHuman provenance - AGENTS.md, JOURNAL.md, lessons (6/8) (#1300) 2026-05-12 15:40:47 -07:00
William Lin b7a448cf5b [feat]: MagiHuman pipeline orchestrator + 10-test parity battery (5/8) (#1299) 2026-05-12 15:12:54 -07:00
04e32991c6 [feat] Dreamverse 05/14: Add streaming runtime (#1328)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 15:11:18 -07:00
William Lin 424c643b6e [feat]: MagiHuman pipeline stages (4/8) (#1298) 2026-05-12 14:59:21 -07:00
William Lin de803cb250 [feat]: MagiHuman DiT (transformer) port + parity tests (3/8) (#1297) 2026-05-12 14:49:03 -07:00
effc1d3492 [feat] Dreamverse 04/14: Add session and prompt logic (#1327)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 14:45:56 -07:00
William Lin b1ddb1ba33 [feat]: T5-Gemma encoder for MagiHuman pipeline (2/8) (#1296) 2026-05-12 13:59:55 -07:00
9ff65c83b6 [feat] Dreamverse 03/14: Add backend skeleton (#1326)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:58:23 -07:00
15473f3c77 [docs] Dreamverse 02/14: Add app documentation (#1325)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:34:42 -07:00
William Lin 490641cf47 [infra]: MagiHuman housekeeping (gitignore, codespell, skills index) (1/8) (#1295) 2026-05-12 12:50:01 -07:00
4ad6880ca6 [docs] Dreamverse 01/14: Add integration provenance (#1324)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 11:42:42 -07:00
Raghav K 9d0af28307 [feat] Add Cosmos 2.5 T2W training pipeline (LoRA + full fine-tune) (#1227) 2026-05-11 19:15:00 -07:00
d6dfe95466 [feat] FastVideo World Model Training (#1179)
Co-authored-by: mignonjia <mhuo@ucsd.edu>
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 05:13:12 +00:00
alexzmsandmergify[bot] 636d3b743e [misc] attention hot-path cleanup + denoising loop hoists (#1272)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 11:51:10 -07:00
William LinandRaghav e3a5c6954f [misc]: import add-model skill stack to .agents/skills/ (#1308)
Co-authored-by: Raghav <ragg04@gmail.com>
2026-05-09 13:06:54 -07:00
f633e30ebb [feat]: add LongCat bidirectional finetuning support (#1244)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-09 12:20:43 -07:00
William Lin 5ce4947ac2 [ci] mergify: accept [skill]/[skills] and [infra] PR title tags (#1309) 2026-05-08 17:14:21 -07:00
William Lin 323d74c0d2 [skills]: New skill - decompose-pipeline-pr (#1303) 2026-05-08 16:05:58 -07:00
William Lin d98aeafc86 [feat]: Loader umbrella-repo support + optional component dirs (#1294) 2026-05-08 15:24:26 -07:00
Shao Duan f6396fb8c6 [feat] Add fastvideo.eval video evaluation suite (#1305) 2026-05-07 20:31:15 -07:00
William Lin 6300329cd5 [infra]: Add activation trace hooks for pipeline debugging (#1293) 2026-05-07 16:38:13 -07:00
MookandSolitaryThinker c17d33bf33 [ci] Replace flaky LTX-2 pixel SSIM with latent-slice cosine regression (#1253)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-05 03:36:58 -07:00
2aaeee2ab8 [feat] Improve API: streaming router (multi-replica load balancer + ws proxy) (#1286)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:00:08 -07:00
eb3a394224 [feat] Improve API: streaming auxiliaries (safety, rewrite, logger, mock) (#1284)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 00:14:34 -07:00
f673423b51 [feat] Improve API: streaming prompt enhancer with LLMProvider abstraction (#1258)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 13:44:40 -07:00
eb0a41528a [feat] Improve API: streaming server GpuPool + worker subprocess (#1257)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 12:56:31 -07:00
William Lin 140bd1a6cf [misc]: standardize install instructions on uv pip install (#1279) 2026-05-02 12:45:50 -07:00
William Lin 11f5a8e582 [misc] pin torch to 2.11.0 (#1277) 2026-05-02 11:48:07 -07:00
71b3cb8c34 [ci] Add CI Performance Regression Tracking Changes (#1248)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-02 03:29:53 -07:00
William Lin c85f6a477f [docs] add hierarchical AGENTS.md per-directory guidance (#1278) 2026-05-02 03:28:18 -07:00
Junda Su 40d4930d73 [bugfix] Update fa import (#1271) 2026-05-02 01:25:22 -07:00
William Lin f9be085243 [ci] pre-commit: drop stale excludes + document agent lint flow (#1276) 2026-05-02 01:19:06 -07:00
William Lin 36b53ff350 [bugfix]: classify stable_audio fields in schema parity inventory (#1275) 2026-05-02 00:12:10 -07:00
William Lin 9801037c3d [refactor] tests/local_tests: organize by model family (#1269) 2026-05-01 01:49:54 -07:00
alexzmsandmergify[bot] 74d09b0efd [misc] cleanup: grad-norm asserts, dead offload file, callback names (#1268)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 01:16:13 -07:00
alexzms 38dc8820ac [ci] add CPU unit tests for train callback system in fastvideo.train (#1267) 2026-05-01 00:53:29 -07:00
William Lin c77a76c6af [feat] Stable Audio Open 1.0: T2A + A2A + RePaint inpainting (native) (#1260) 2026-05-01 00:07:11 -07:00
alexzms d14d5aadea [feat] Cosmos 2.5 training support in fastvideo.train (#1224) 2026-05-01 01:15:02 +00:00
alexzms 4c915b7742 [ci] add CPU unit tests for train checkpoint utilities in fastvideo.train (#1265) 2026-04-29 18:55:39 +00:00
alexzms 9a8bbe18fa [bugfix]: fix SP deadlock in negative prompt encoding during training (#1178) 2026-04-28 01:06:49 +00:00
alexzms ea25441ef0 [ci] add CPU unit tests for fastvideo.train load_run_config (#1264) 2026-04-28 01:06:18 +00:00
48957fcde1 [bugfix] Fix modal remote functions crash container on sys exit in CI remote functions (#1261)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-27 21:50:55 +00:00
Mook 7b872cc41e [Perf] Skip bool-mask round-trip in block-sparse VSA attention (#1243) 2026-04-26 15:14:37 -07:00
alexzms 37418946c8 [docs]: clarify real_score_guidance_scale CFG parameterization (#1256) 2026-04-26 16:38:00 +08:00
William Lin 95fd29e0cb [feat] Streaming WebSocket server skeleton (single generator + fMP4) (#1251) 2026-04-26 00:33:49 -07:00
Junda Suandmergify[bot] e17cd2633c [bugfix]: normalize uint8 pil_image in I2V VAE encoding (#1249)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-24 09:16:01 +00:00
William Lin e0dc5f2b0c [feat] Add typed LTX-2 continuation state and streaming session store (#1250) 2026-04-24 01:28:07 -07:00
William Lin 70ee5d230c [feat] [6/n] Improve API: LTX-2 public preset + asset wiring + gpu_pool translation (#1239) 2026-04-23 11:36:45 -07:00
William Lin 24ced500f5 [test] add LTX-2 distilled T2V SSIM regression test (#1240) 2026-04-21 12:03:38 -07:00
William Lin 4ddcdf541f [feat] [5.5/n] Improve API: streaming server config surface + serve dispatch (#1238) 2026-04-17 15:36:21 -07:00
William Lin 0e3529869c [feat] [5/n] Improve API: wire ServeConfig.default_request into OpenAI serving (#1237) 2026-04-17 13:26:18 -07:00
William Lin e1e0d91c00 [misc] small cleanup for API handling (#1235) 2026-04-16 16:21:21 -07:00
William Lin 145a3f166b [feat] [4/n] Improve API: refactor sampling param and merge with presets (#1234) 2026-04-16 14:10:02 -07:00
William Lin 88a5a933ab [feat] [3/n] Improve API: extend support to cli (#1226) 2026-04-14 15:20:47 -07:00
William Lin c591d6d2a6 [feat] [2/n] Improve API: add initial support in video_generator (#1220) 2026-04-06 10:33:54 -07:00
Kun Linandmergify[bot] 65dff806a8 [bugfix]Fixing Lora distillation training distributed checkpointing bug (#1192)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-06 02:20:26 +00:00
KUAN-HAO HUANGandmergify[bot] b85f0f4c2a [perf]: Eliminate CPU-GPU synchronization bottlenecks in training pipeline (#1217)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-06 02:03:46 +00:00
William Lin 76c62d7a00 [feat] [1/n] API improvements: add intial files for new fastvideo public API (#1218) 2026-04-05 18:13:19 -07:00
f6e65ff668 [Feature] Add BSA (Bidirectional Sparse Attention) inference backend (#1174)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-05 05:00:33 +00:00
mergify[bot] c220aa8000 [ci](mergify): upgrade configuration to current format (#1216)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-04 23:09:17 +00:00
Jinzhe PanandDarren Sadr 4713fc17ed [feat] Job Runner UI (#1189)
Co-authored-by: Darren Sadr <darrensadr@gmail.com>
2026-04-02 16:07:24 -07:00
vishruthb 5789955bbe [feat] add gen3c (cosmos-7b) model and pipeline support (#1059) 2026-04-01 11:42:02 +00:00
Jinzhe Pan 2ad84a3b78 [ci] Use update instead of rebase for auto branch sync (#1215) 2026-04-01 19:16:59 +08:00
Jinzhe Pan 12d699cd78 [ci] Add direct test retry with check overwrite and aggregate status refresh (#1214) 2026-04-01 17:21:28 +08:00
Jinzhe Pan 34f14ded21 [ci] Use pull_request_target for Full Suite trigger (#1213) 2026-04-01 03:01:07 +08:00
Jinzhe Pan 71d1ab411f [ci] Fix jq crash when Buildkite build env is null (#1212) 2026-04-01 02:35:01 +08:00
Jinzhe Pan 805e487773 [ci] Ignore legacy reference videos when checking for HF download (#1211) 2026-04-01 02:12:09 +08:00
Jinzhe Pan 8803b4547e [ci] Add retry for flaky tests and fix stale SSIM references (#1210) 2026-04-01 01:11:49 +08:00
Jinzhe Pan 3b3806b3f6 [ci] Fix /merge to directly trigger Full Suite + simplify rebase conditions (#1209) 2026-03-31 23:17:09 +08:00
Jinzhe Pan 38d962e89d [ci] Remove Mergify ready-label race condition (#1208) 2026-03-31 20:59:13 +08:00
Jinzhe Pan 3966a365d0 [ci] Add statuses:write permission for /test pre-commit (#1207) 2026-03-31 20:33:18 +08:00
Jinzhe Pan d73fd14af0 [ci] Post pre-commit status to PR commit SHA (#1206) 2026-03-31 20:21:21 +08:00
Jinzhe Pan a87cc89916 [ci] Trigger pre-commit on /test slash commands (#1205) 2026-03-31 20:12:57 +08:00
Jinzhe Pan 81fd80c8ee [ci] Add TEST_SCOPE routing for clean single-test execution (#1203) 2026-03-31 19:40:59 +08:00
Jinzhe Pan ff22439f28 [ci] Fix fork PR checkout for /test and Full Suite triggers (#1202) 2026-03-31 13:42:13 +08:00
Jinzhe Pan de0de04212 [ci] Replace Merge Queue with auto-merge — reduce CI complexity (#1200) 2026-03-31 10:09:33 +08:00
mergify[bot] 7f2c3e1f64 [ci](mergify): upgrade configuration to current format (#1194)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-03-31 00:47:48 +08:00
Jinzhe Pan ab55e57c22 [ci] Fix Merge Queue requeue and draft PR pre-commit skip (#1197) 2026-03-30 22:35:13 +08:00
Jinzhe Pan 46f6b43a53 [ci] Fix Merge Queue immediate dequeue (#1196) 2026-03-30 21:33:45 +08:00
Jinzhe Pan 833a33b663 [ci] CI follow-up: gate checks, issue label unification, draft PR skip (#1193) 2026-03-30 20:19:46 +08:00
Jinzhe Pan 9ea1307cd4 [ci] Add approval and pre-commit checks to merge protections (#1190)
## Summary

Follow-up to #1187. Two small changes:

1. **Merge Protections expanded** — adds `#approved-reviews-by>=1` and `check-success~=pre-commit` to `merge_protections` so the Mergify check shows a unified requirements checklist on every PR (title format + approval + pre-commit), instead of only showing the title format.

2. **Buildkite pipeline comment fix** — updates the outdated Full Suite section comment from "Triggered by adding the 'ready' label via GitHub Actions → Buildkite API" to reflect the new Merge Queue trigger path.
2026-03-30 05:14:49 +00:00
Jinzhe Pan be35003cb1 [ci] Merge Queue, label system overhaul, and slash commands (2/2) (#1187) 2026-03-30 08:22:35 +08:00
Jinzhe Pan 26bd4db253 [ci] CI infrastructure cleanup and workflow reorganization (1/2) (#1186) 2026-03-29 17:01:09 -07:00
Jinzhe PanandWill Lin e294ca011c [feat]: overhaul SSIM test infrastructure — partition scheduling, helper migration, CI fixes (#1185)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-03-29 23:59:25 +00:00
alexzms 2085a4fc4a [bugfix]: fix VAE temporal tiling blend corruption in tiled_encode (#1181) 2026-03-29 23:12:42 +00:00
Jinzhe Pan c0c8e39c04 Revert "[feat] Job Runner UI" (#1188) 2026-03-29 16:46:27 +08:00
Darren f72618dafb [feat] Job Runner UI (#1172) 2026-03-29 16:17:56 +08:00
alexzms b3edfacdd8 [bugfix]: fix I2V preprocessing crash for models without CLIP (Wan2.2 I2V) (#1184) 2026-03-28 10:28:50 +08:00
alexzms 30129a3350 [misc]: reorganize training configs and add documentation (#1177) 2026-03-26 16:38:31 -07:00
jaisurya27 4d49f7b0aa Kandinsky5 lite dit clean (#1088) 2026-03-26 08:07:11 +08:00
alexzms 71bfc13d75 [feat]: add HunyuanVideo model plugin for fastvideo/train framework (#1175) 2026-03-24 16:49:51 -07:00
Kaiqin Kong 74db6e18d1 [misc] update action loading in validation and preprocess (#1143) 2026-03-24 15:10:03 -07:00
Kaiqin Kong 7d263c6a36 [bugfix] self-forcing train/validation step mismatch (#1173) 2026-03-20 00:52:45 -07:00
Zhang Peiyuan 454c32d1d1 Update README.md 2026-03-17 14:26:07 -07:00
Jinzhe Pan d1240b9238 [CI] add contributor interaction automation (#1170) 2026-03-17 12:05:40 +08:00
Hao Zhang 4105094fa5 [docs] Update README with realtime demo announcement (#1169) 2026-03-13 15:52:02 -07:00
alexzms f036469d3d [feat]: Knowledge Distillation training method for ODE-init (KDMethod + KDCausalMethod) (#1166) 2026-03-11 20:43:24 -07:00
alexzms 14261bc98c [feat] pre-commit support 120 col num (#1167) 2026-03-11 20:19:30 -07:00
alexzms d92858659d [feat] Self-Forcing methods in refactored training infra (#1164) 2026-03-09 20:20:59 -07:00
1a383f3f66 [refactor] train v1 clean up
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-215.slurm-compute.tenant-slurm.svc.cluster.local>
2026-03-09 18:58:56 -07:00
alexzmsandPeiyuan Zhang bc27a032c5 [feat] Refactor training framework into fastvideo/train (#1159)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2026-03-09 15:16:42 -07:00
alexzms 2b13e117f0 [Feat] Add causal Wan pipeline with multi-step denoising (#1161) 2026-03-08 13:32:04 -07:00
Junda Chen 99c166c381 feat: Building agent friendly repo (#1151) 2026-03-07 17:46:29 -08:00
XOR-op 95066245db [misc] FlashAttention 4 support (#1114) 2026-03-07 16:53:43 -08:00
Jinzhe Panandgemini-code-assist[bot] 6dcaac768b [CI] PR template (#1157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-07 11:51:28 -08:00
Ajay Anubolu 02c1c49b75 [CI] Add inference performance regression tests (#1140) 2026-03-07 08:26:54 +08:00
Zhang Peiyuan cd1b7cf139 [Refactor] SP Mask --> original seq len; HunyuanVideo 1.5 does not need mask (#1142) 2026-03-04 11:44:47 +08:00
Ajay Anubolu e63b7d8ac4 [Feat] Added OpenAI-compatible API server and benchmark script (#1109) 2026-03-02 17:12:32 -05:00
Jinzhe Pan 5190c1bb1e [Doc] add doc for inference architecture (#1147) 2026-03-02 13:44:27 -08:00
Darren 2cb3bba658 [bugfix]: fix a bug where collect_env was not running properly... (#1145) 2026-03-02 10:57:29 -08:00
Jinzhe Pan f9e1c46c3c [CI][Feat] launch 2 instance to run ssim (#1137) 2026-03-01 01:49:29 -08:00
Peiyuan Zhang e1eda47589 remove temporal frame adjustment 2026-02-27 20:47:16 +00:00
Zhang Peiyuan d902967208 Py/fix sp (#1138) 2026-02-27 12:14:44 +08:00
Zhang Peiyuan fea556269b [Misc] Fix memory leakage in VideoGenerator (#1132) 2026-02-26 19:51:32 -08:00
William Lin 69dd3c68f6 [bugfix] fix matrix game kv indexing and CI (#1135) 2026-02-26 01:24:42 -08:00
Jinzhe Pan 5433f6e80b [fix] preprocessing issue (#1134) 2026-02-25 21:49:52 -08:00
Junda (David) Su e315657066 [docs] [kernel] Migrate to uv (#1127) 2026-02-25 14:13:50 -08:00
Zhang Peiyuan f8d9a0c57f [misc] fix hunyuan (#1125) 2026-02-25 08:26:29 +08:00
Jinzhe PanandWill Lin fa6d276925 [Feat] Improved CI (#1119)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-24 12:53:02 -08:00
Zhang PeiyuanandWill Lin fc80d95d7e [Misc] Remove STA (#1124)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-23 15:14:42 -08:00
Shao DuanandSolitaryThinker 37cab18780 [bugfix] Added ltx2 guidance missing modulation term (#1100)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-02-23 14:22:31 -08:00
Zhang Peiyuan 128d0b7fc5 [Misc] Remove Teacache (#1121) 2026-02-22 16:53:07 -08:00
Matthew Noto 8092f02e6d small refactor in post-processing to improve efficiency (#1123) 2026-02-22 16:45:44 -08:00
Zhang Peiyuan 03d9ce2edb [Misc] Remove StepVideo (#1118) 2026-02-21 17:15:42 -08:00
10fc92dba5 Upstream LTX2 Training (#1116)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-239.slurm-compute.tenant-slurm.svc.cluster.local>
2026-02-21 16:06:55 -08:00
6736dc06a5 Improve Docs (#1112)
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-227.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local>
2026-02-19 14:23:32 -08:00
William Lin 8c002c62af [misc] add hy-world link to readme (#1113) 2026-02-18 12:01:10 -08:00
Darren 7061313d04 [bugfix] get_torch_device and other device calls were being made on non-cuda platforms (#1107) 2026-02-18 11:43:46 -08:00
Zhang Peiyuanandgemini-code-assist[bot] 76d3ba69e0 [Misc] clean up VSA finetuning examples. (#1111)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-18 11:37:20 -08:00
8e39ce38c9 [Feat] Native dit implementation for SD3.5 (#1093)
Co-authored-by: Ishan Vaish <vaish.ishan@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 10:19:40 +08:00
Darren d4bd8bf2c0 Update README.md (#1110) 2026-02-17 14:54:34 -08:00
Darren e83d7bc50c [bugfix] fix import PreTrainedModel in stepllm.py (#1108) 2026-02-16 21:45:08 -08:00
Jinzhe Pan ff3d5aff75 [Fix] hunyuan postprecessing issue (#1104) 2026-02-15 12:26:35 -08:00
XOR-op 959dbcc8a2 [perf] causal MatrixGame optimization (#1078) 2026-02-15 09:51:00 +08:00
William Lin 36bf37e9ba [bugfix] Fix failed kernel publish and SFT regressions (#1103) 2026-02-14 16:19:17 -08:00
Mihir Jagtap 7a83e0e6fc [feature] Add Hunyuan-GameCraft model support (#1071) 2026-02-14 08:07:44 +08:00
William Lin 8be1313b86 [kernel] add torch 2.10 to package build matrix (#1099) 2026-02-13 13:03:05 -08:00
alexzms d925ad05f3 [bugfix] fastvideo-kernel: fix VSA Triton padding NaNs and support q/kv length mismatch (#1094) 2026-02-13 12:39:49 -08:00
Shao DuanandWill Lin 7f795600c8 [bugfix] Fixed ltx2 base cfg guidance (#1095)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-13 11:53:11 -08:00
William Lin ec16b6b01d [misc] update wechat group link (#1098) 2026-02-13 01:19:11 -08:00
Kaiqin Kong 31c0f1b341 [feat] Port LingBot-World-Base (Cam) (#1081) 2026-02-10 11:12:33 -08:00
William Lin 4bee0fa199 [misc] cleanup assets/ and demo/ (#1091) 2026-02-10 02:26:09 -08:00
530e6b8363 [Model] LTX 2 Base (#1064)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-10 01:11:17 -08:00
Jinzhe Pan 9ab2725db1 [ci] CI Transformer Tests (#1089) 2026-02-10 01:08:59 -08:00
IshanandJinzhe Pan 0aff68f51d [Feat] Add Stable Diffusion 3.5 (#1075)
Co-authored-by: Jinzhe Pan <eigensystem1318@gmail.com>
2026-02-10 14:31:36 +08:00
ad58f802f3 [Feat] Port LTX2 trainer (#1074)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
2026-02-09 17:32:57 -08:00
Wei Zhou 04fa356ee3 [Misc] [Training] Fixed a bunch of bugs in current training pipeline (#1084) 2026-02-09 16:01:05 -08:00
Matthew Notoandgemini-code-assist[bot] f9c076fe2b [misc] add AGENTS.md file (#1085)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-09 01:03:49 -08:00
XOR-op f76efe798e [bugfix]: _compile_conditions regression (#1077) 2026-02-06 18:57:43 -08:00
Zhang Peiyuan 09f455233e [misc] readme small fix (#1076) 2026-02-06 16:57:48 -08:00
XOR-op aea300f690 [perf]: use CUDA IPC in multiproc executor to avoid serialization overhead (#1061) 2026-02-06 19:43:18 -05:00
Hao Zhang b92219f6a6 more fix and relocate STA arguments to pipeline config (#1073) 2026-02-06 13:38:35 -08:00
Jinzhe Pan a321b95a8a [Fix] remove video ratio limitation (#1069) 2026-02-05 20:33:08 -08:00
Wei Zhou 98308db7e0 [Feature] [Hy1.5] Support HY1.5 super-resolution pipeline for 1080p videos (#1046) 2026-02-05 16:39:20 -08:00
Hao Zhang c1e18f6722 Some minor fixes (#1068) 2026-02-05 16:28:00 -08:00
William Lin 75e193a2c9 [core] Refactor and centralize our registry for models, pipelines, and sampling params (#1066) 2026-02-05 14:40:30 -08:00
XOR-op d6e0a7d0dd [refactor] Action module (#1065) 2026-02-05 13:56:35 -08:00
William Linandgemini-code-assist[bot] 7fc5f241da [misc] Fix naming instruction in runpod.md (#1067)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-05 13:45:49 -08:00
Mingjia Huo aae48a7e90 [feat] HYworld VAE with cache (#1057) 2026-02-05 04:18:17 -08:00
William Lin 88f38eb0f4 [misc] upgrade torch to 2.10 (#1048) 2026-02-05 04:15:22 -08:00
Shao Duan d750b463dc Added Sequence Parallelism for LTX-2 Distilled (#1036) 2026-02-04 15:25:30 -08:00
KyleShaoandWill Lin e10b26a3d8 [feat] Add Cosmos 2.5 I2W/V2W support (staged pipeline + examples) (#1021)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-04 14:19:30 -08:00
XOR-op 74636ba246 [chore]: use higher precision timestamp in logging (#1062) 2026-02-04 13:22:59 -08:00
Wei Zhou 7e2f3f14e7 [Bugfix] [Wan I2V] Fix CLIP Image encoder config (#1063) 2026-02-04 13:20:19 -08:00
Kaiqin Kong caa1c402ba [bugfix] Double Normalization in Preprocessing Dataset (#1055) 2026-01-31 11:11:59 -08:00
XOR-op 38a6bd93d3 [chore]: update sageattn3 installation instructions (#1050) 2026-01-29 15:45:52 -08:00
alexzms b867ef7e7c [SP Sharding] Fix SP loss sharding on token axis (thw) with padding; add distributed correctness tests (#1045)
Fixes the sequence parallel sharding on t, now SP shards on t*h*w
2026-01-27 22:53:04 -08:00
William Lin 3ae58c277a [docs] Update design overview and add agents tutorial (#1044) 2026-01-27 15:56:58 -08:00
Kaiqin Kong 0c6862ca55 [feature] Add Matrix Game 2.0 training (#1017)
The CI tests are quite unstable, but since multiple CI tests indicates that each individual tests are passed, I think we can merge this.
2026-01-26 19:21:43 -08:00
XOR-opandWill Lin e8c854bcf1 [docs] Offloading instruction (#1022)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-01-26 13:40:28 -08:00
William Lin 06860e96fe [docs] Update runpod instructions (#1043) 2026-01-26 13:13:54 -08:00
Matthew Noto 1b503554d1 [bugfix] fix torchvision import (#1039) 2026-01-24 22:37:14 -08:00
Shreejith SGandgemini-code-assist[bot] 351ceb7c59 [bugfix]: handle architectural differences while lora extraction (#1035)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-24 15:42:28 -08:00
KyleShao 10875e0d7b [bugfix] Fix NCCL all_gather contiguity + correct ParallelTiledVAE decode tiling threshold (#1037) 2026-01-24 15:37:35 -08:00
alexzms 1eaae8a10b [ci] Increase ci test error threshold (#1038) 2026-01-24 15:36:10 -08:00
Mingjia Huo 59e00f6164 [feat] Add HY-World1.5-Bidirectional-480P-I2V (#1027)
VAE requires further improvement, will raise PR in near future.
2026-01-23 14:18:04 -08:00
745cc05b10 [bugfix] Allow update timesteps for hy1.5 model. (#1033)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-22 22:04:53 -08:00
William Lin c5dc244871 [bugfix] add omegaconf as dep. (#1032) 2026-01-22 11:59:28 -08:00
alexzms dbf3917bf4 [fastvideo-kernel] replace map to index with Triton implementation + add vsa benchmark (#1029) 2026-01-22 11:35:02 -08:00
XOR-op 050f189c95 fix: SP for hunyuanvideo 1.5 (#1026) 2026-01-21 14:40:06 -08:00
Shao DuanandWill Lin 029216029f Added LTX-2 Distilled T2V Generation (#1016)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-01-21 14:11:39 -08:00
alexzmsandWilliam Lin 31f44110b5 [kernel] [bugfix] [ci] bump v0.2.4. Fix STA output handling, TurboDiffusion CUDA norm dtypes for fastvideo-kernel unit tests. (#1020)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
2026-01-19 17:42:42 -08:00
William Lin 21f3ce6577 [kernel] Fix fastvideo-kernel release workflow (#1019) 2026-01-17 15:57:46 -08:00
XOR-op 785d123e36 [feat] Hooks API and layerwise offloading for all DiTs (#1006) 2026-01-17 11:22:02 -08:00
William Lin d58c551c11 [chore] release fastvideo-kernel 0.2.3 (#1018) 2026-01-17 02:24:23 -08:00
alexzms 560628709c [Bug Fix] Add autograd wrapper for block-sparse attention in fastvideo-kernel + fix CMake extension linking (#1015) 2026-01-16 21:16:43 -08:00
William Lin 0f53b51e6c [CI] Fix OOM issues in ssim tests (#1011) 2026-01-16 21:15:20 -08:00
alexzmsandWill Lin 06093a9c4e [CI] SSIM tests optimization: load all model weights from Modal persistent Volume (#958)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-01-16 11:50:43 -08:00
KyleShao dbddfab6d2 [feat] Introduce Cosmos 2.5 Text2World pipeline (#974) 2026-01-15 15:09:05 -08:00
William Lin 7188170277 [misc] [bugfix] unpin 'av' in pyproject (#1009) 2026-01-13 15:40:46 -08:00
XOR-op b7f69c2c1d [feat!] Disable FSDP inference by default (#1001) 2026-01-13 14:20:05 -08:00
Loay Rashid 23a4531491 [CI] Fixed Turbodiffusion I2V CI (#1002) 2026-01-13 01:08:58 -08:00
William Linandgemini-code-assist[bot] 7d52ad0118 [ci] temporarily disable turbodiffusion ssim test (#1000)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-08 15:43:50 -08:00
Will Lin 4d7bf35fa3 Revert "dit"
This reverts commit a6a9c9ca07.
2026-01-07 03:22:48 -08:00
Will Lin a6a9c9ca07 dit 2026-01-07 03:18:48 -08:00
f4704847c2 [bugfix] Add configs for TurboDiffusion T2V/I2V models (#993)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-01-06 16:36:45 -06:00
Shreejith SGandWill Lin d9c996310b [docs]: add LoRA extraction utilities documentation (#992)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-01-06 16:36:26 -06:00
Shao Duan d6651afd2e [examples] Added longcat-video python api examples (#994) 2026-01-06 15:03:42 -06:00
William Lin cf67618cad [chore] release 0.1.7 (real) (#980) 2026-01-05 15:47:05 -06:00
William Lin 2f0a2b3c57 [misc] add pin_cpu_memory false for RTX 4090 (#990) 2026-01-05 15:45:35 -06:00
Loay Rashid e7748d9952 [feat] add Turbodiffusion I2V pipeline (#984) 2026-01-05 15:41:23 -06:00
William Lin 8eb3140b2f [misc] pin fastvideo-kernel in .toml file (#989) 2026-01-05 13:42:32 -06:00
Shao Duan d6ddcea682 Add LongCat-Video I2V and Video Continuation (Base, Distillation and Refinement) Support to FastVideo (#953) 2026-01-04 22:20:09 -06:00
William Lin 3559ba2377 [chore] update wechat QR code (#988) 2026-01-04 21:59:38 -06:00
William Lin 61e63ea0d7 [chore] release fastvideo-kernel 0.2.2 (#986) 2026-01-04 21:21:06 -06:00
William Lin 4ce4ac4734 [ci] increase ssim and lora inference test timeout (#985) 2026-01-04 15:08:20 -06:00
William Lin e7f6db9bd1 [docs] Update docs and README (#975) 2026-01-04 14:59:19 -06:00
Ohm-Rishabh d83f45a6a0 Layer offloading (#966) 2026-01-03 21:46:00 -08:00
XOR-op dd91542cd1 [feat] Support text encoder weight override and quantization (#983) 2026-01-03 15:33:33 -06:00
Kaiqin Kong 581e8115fe [feat] support Matrix-Game 2.0 streaming generation (#957) 2026-01-02 19:14:38 -06:00
Loay Rashid dea69cf651 [New Model] Turbodiffusion (#971) 2026-01-02 17:55:56 -06:00
XOR-op 60ac6537df [feat] Support absmax style quantization for FP8 (#981) 2026-01-02 16:00:18 -06:00
Qi Jia 5285116e73 [docs]: fix various broken links across the documentation (#979) 2026-01-01 20:02:39 -06:00
William Lin 40ce2d72f5 [kernel] add turbodiffusion kernels (#972) 2025-12-30 04:23:10 -06:00
William Lin 704bc9aaf9 [misc] Add util script to create diffuser HF repo from custom component weights (#970) 2025-12-29 19:38:30 -06:00
RoyWangandroywang de264fcc99 [fix]: fix STA trition kernel for AMD RDNA archs (#969)
Co-authored-by: roywang <roywang@amd.com>
2025-12-29 14:25:32 -06:00
RoyWangandroywang 7b952e4673 [fix]: fix fastvideo-kernel Rocm build and Dockerfile for Rocm (#968)
Co-authored-by: roywang <roywang@amd.com>
2025-12-29 14:24:46 -06:00
551b2d2048 [fix]: fix sliding_tile_attn with sdpa(without flash_attn) (#967)
Co-authored-by: roywang <roywang@amd.com>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-12-29 14:19:40 -06:00
Ketaki Tank 7bfaf82fd7 [feat] Add new feature extractors for fvd (#954) 2025-12-27 05:08:26 -06:00
William Lin 9cd6a86b95 [chore] release v0.1.7 (#955) 2025-12-27 05:05:37 -06:00
William Lin 16e9552778 [kernel] Fix docker release build for kernel (#965) 2025-12-26 21:42:55 -06:00
William Lin 87f8a2782d [docs] refactor attention docs (#964) 2025-12-26 15:21:49 -06:00
William Lin cbbb09d7b8 [kernel] Release fastvideo-kernel v0.2.1 (#963) 2025-12-26 13:59:03 -06:00
William LinandShreejithSG 2f6230abcf [kernel] Reorg and fix fastvideo-kernel (#962)
Co-authored-by: ShreejithSG <shreejithsg@gmail.com>
2025-12-26 01:50:24 -06:00
Shreejith SGandWilliam Lin f8bfc76015 feat: consolidate attention kernels into unified fastvideo-kernel package (#946)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
2025-12-24 01:39:51 -06:00
alexzmsandShao Duan 8f1e6c3336 Add LongCat T2V (Base, Distillation and Refinement) Support to FastVideo (#883)
Co-authored-by: Shao Duan <shaoxiongduan@gmail.com>
2025-12-23 01:11:18 -06:00
William Lin 8e7d2e7879 [bugfix] [dmd2] allow dmd2 simulate_student_forward to use text-only dataset (#951) 2025-12-23 00:41:31 -06:00
William Lin 6ab2870942 [rocm] Add rocm fastvideo docker image (#952) 2025-12-22 18:21:04 -06:00
RoyWang e0ad145152 [feat] add sliding_tile attention triton kernel and ROCM support (#916) 2025-12-22 18:02:51 -06:00
Matthew Noto da04d08426 [docs] small fixes (#947) 2025-12-22 15:23:55 -06:00
Wei Zhou 1f70032af5 [New Model] Hunyuan1.5 (#943) 2025-12-21 00:57:52 -06:00
William Lin 7f71994653 [misc] Allow manual override of Pipeline class through override_pipeline_cls_name (#945) 2025-12-20 14:39:17 -06:00
Loay Rashid 2bb3349da1 [bugfix] Added VSA Padding logic (#944) 2025-12-20 14:29:11 -06:00
Kaiqin Kong 8fe1689968 [feat] Add Matrix-Game 2.0 (#938) 2025-12-20 14:09:12 -06:00
Loay Rashid e53730f324 [docs] Minor Fixes (#942) 2025-12-19 16:48:16 -06:00
Loay Rashid 7a4fe9086a [feat] Support sequence packing and shard after pachification for USP (#894) 2025-12-19 16:19:46 -06:00
Ohm-Rishabh d277361aae [misc] add schedule configurations to pytorch profiler (#934) 2025-12-18 01:45:23 -06:00
alexzms 734a54e7a9 [ci]: Use pre-built docker image & skip VSA compilation (#939) 2025-12-16 23:14:11 -08:00
alexzms 91364982df [Feature] Support for Variable Q/KV Sequence Lengths in VSA ThunderKittens kernel (#911) 2025-12-16 20:08:15 -08:00
William Lin 50145e4fcb [CI] Fix CI tests (#935) 2025-12-16 04:59:43 -08:00
William Lin 4112507e99 [misc] upgrade pytorch version to 2.9.0 (#928) 2025-12-15 04:12:43 -08:00
William Lin 424fc2b4ae [bugfix] [lora] [distillation] Fix lora distillation bug (#933) 2025-12-15 04:12:02 -08:00
William Lin e6066223e6 [bugfix] [VSA] [distillation] Various bugfixes for VSA and distillation and nightly tests (#932) 2025-12-12 16:51:54 -08:00
William Lin b6fa3d24d8 [misc] update wechat image (#931) 2025-12-11 21:22:21 -08:00
Ketaki Tank 55c2e7cd76 [feat] Add fvd implementation (#923) 2025-12-11 19:06:19 -08:00
Tuyabei 5a549af823 [bugfix] [VSA] Fix block_size computation in backward kernel (#925) 2025-12-10 14:36:40 -08:00
Shreejith SG 92fb660c2e Add LoRA extraction, verification, and comparison scripts (#865) 2025-12-08 16:07:58 -08:00
William Lin 3ff640b2e6 [bigfix] [distillation] Fix DMD inference pipeline noise initialization shape (#921) 2025-12-08 13:00:48 -08:00
William Lin c722429ab5 [docs] fix testing.md visibility (#920) 2025-12-08 00:44:53 -08:00
KyleShaoandKyleS1016 e04a192de6 [feat]: add COSMOS 2.5 DiT implementation (#897)
Co-authored-by: KyleS1016 <kyle.s@gmicloud.ai>
2025-12-07 21:48:32 -08:00
William Lin c9ca6d1298 [docs] add docs for ssim testing (#918) 2025-12-06 18:20:04 -08:00
Wenxuan TanandSolitaryThinker 754292c419 Use assert_close in tests (#429)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-12-06 18:18:25 -08:00
Qi Jia 0082bc66fc fix: correct mp backend GPU assignment on multi-GPU systems (#912) 2025-11-30 23:00:22 -08:00
Ohm-Rishabh 8b1937422e [feat] training mfu calculation scripts (#871) 2025-11-27 16:54:17 -08:00
fb6cbf23e6 Fix the docs (#905)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
2025-11-27 00:37:03 -08:00
Mihir Jagtap c8fdd5ed7b [docs] modified the .github/workflows/docs.yml file to include path filtering (#906) 2025-11-26 17:34:21 -08:00
Loay Rashid 1c19a6a00c [Bugfix] Minor bugfixes (#889) 2025-11-26 17:20:45 -08:00
William Lin d44409c704 [CI] fix VSA training CI (#900) 2025-11-24 17:47:59 -08:00
Zhang Peiyuan 5d1c7852b7 + Awesome work using FastVideo or our research projects (#898) 2025-11-23 22:22:27 -08:00
Wenxuan Tan 77a211d006 [misc] Update wechat link (#893) 2025-11-20 19:59:05 -08:00
Wei Zhou bef8169bb1 [Feat] [I2V] resize all image sizes to below 480*832 (#890) 2025-11-20 00:08:36 -08:00
William Lin 681f1583f9 [readme] update link to inference code (#887) 2025-11-19 13:24:13 -08:00
e3b4564d5a [feat] Add inference for MoE SF (#880)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-11-19 13:16:24 -08:00
Shao Duan c0d03fc43d [bugfix] [lora] [CI] Fix LoRA alpha scaling factor & Fix LoRA Inference CI (#870) 2025-11-19 01:02:01 -08:00
Wei Zhou 404ee8538e [Bugfix] [DMD Distillation] Each rank should have its own timestep sampled (#885) 2025-11-18 14:03:25 -08:00
Shao Duan e57ac59462 Fix mp worker busy loop to handle all string RPC methods (#881) 2025-11-16 13:26:44 -08:00
Mihir Jagtap 8c55fdaf7e [docs] add favicon (#878) 2025-11-15 13:44:16 -08:00
Y-aang c30779184f fix: incorrect dv in vsa Triton kernel causing test_vsa error (#879) 2025-11-14 22:00:39 -08:00
William Lin 9d188c0b6c [misc] update wechat and slack invite links (#875) 2025-11-12 23:03:56 -08:00
Mihir Jagtap 9dd7c54221 [docs] Update Home Readme.md with fixed links (#873) 2025-11-12 13:32:44 -08:00
William Lin 62b95d8287 [feat] prepare for wan2.2 SF (#861) 2025-11-04 18:06:48 -08:00
Kaiqin Kong fdf21702f5 [Docs] add diagrams to docs (#863) 2025-11-04 16:29:07 -08:00
Ohm-Rishabh 2972fc9449 Improve FSDP loading with size-based filtering (#853) 2025-11-04 15:31:07 -08:00
Mihir Jagtap 8f5712629f [docs] port to mkdocs (#855) 2025-11-04 14:31:56 -08:00
Kevin Lin 436c701b9f [bugfix] Add Cosmos2 sampling params to registry (#862) 2025-11-02 00:09:17 -07:00
Kevin Lin 543fea88e3 [Feature] Add Cosmos2 i2v pipeline (#837) 2025-10-30 20:03:57 -07:00
Kaiqin Kong bdec816b31 move STA_configuration.py to fastvideo/attention/backends (#856) 2025-10-29 13:54:13 -07:00
William Lin 2cd2e57d2e [ci] fix causal ssim test (#848) 2025-10-26 19:33:07 -07:00
William Linandainsley 9370234294 [feat] Add gradio local inference demo (#847)
Co-authored-by: ainsley <jzhang2765@wisc.edu>
2025-10-26 07:01:33 -07:00
Jinzhe Pan 50da62e722 [bugfix] always force spawn instead of fork (#852) 2025-10-23 16:36:50 -07:00
William Lin 4f3e8751db [bugfix] [misc] Use training_state_checkpointing_steps in scripts/ (#846) 2025-10-19 20:20:53 -07:00
Jinzhe PanandXingyu Long f4c58894d9 [Feat] add ray support (#838)
Co-authored-by: Xingyu Long <xingyulong97@gmail.com>
2025-10-16 23:17:54 -07:00
Ohm-Rishabh 01c94ef385 [feat] unified trainer logging (#841) 2025-10-16 23:16:16 -07:00
Zhang Peiyuan 2415226d25 Update WeChat Link 2025-10-13 21:02:46 -07:00
Jiali Chen 404314d00f [Feature]Add video-to-video (V2V) pipeline (#829) 2025-10-12 21:53:05 -07:00
zyang6andkiritorl 87489f0872 Add wan2.1 functionality support for Ascend NPU platform (#810)
Co-authored-by: kiritorl <1021709528@qq.com>
2025-10-09 16:25:08 -07:00
Zhang Peiyuan 9ce7c8039e Update Wechat link 2025-10-06 15:01:19 -07:00
William Lin e1e25e95f9 [feature] Add torch profiler (#827) 2025-10-06 07:59:46 -07:00
William Lin 490bde90e1 [bugfix] Allow overriding dit checkpoint for inference and Lower VSA LR in example scripts (#831) 2025-10-05 01:44:49 -07:00
dc7596b973 [self-forcing][8/n] Self-Forcing For Wan2.2-A14B + torch.compile training and distillation support (#818)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-02 15:01:45 -07:00
William Lin 335afa4457 [bugfix] Use training_state_checkpointing_steps instead of checkpointing_steps (#821) 2025-09-28 15:22:43 -07:00
Yongqi Chen 3f77a6805a [Feature]Update count trainable param for FSDP2 (#820) 2025-09-28 15:22:04 -07:00
RandNMR73 13d0aae706 Add Sage Attention 3 Backend (#815) 2025-09-24 15:11:38 -07:00
William Lin 404cbf4f3c [self-forcing] [6/n] Add Ode Init training (#811) 2025-09-22 17:58:19 -07:00
William Lin 958ffec844 [bugfix] Update learning rates for sparse distillation recipe (#812) 2025-09-22 12:07:03 -07:00
31f000d1cc [self-forcing] [5/n] Add Self-Forcing distillation pipeline (#808)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-09-20 19:32:10 -07:00
Yongqi Chen cd32b3e02f Update example files and readme (#809) 2025-09-20 18:15:59 -07:00
Zhang Peiyuan bf27908095 Update WeChat Link 2025-09-20 14:16:20 -07:00
William Lin c5f9ea53b2 [self-forcing] [4/n] Preprocessing for collecting ODE trajectory (#788) 2025-09-15 17:54:42 -07:00
William Lin d32a7184da [bugfix] Wan2.2 Boundary ratio (#804) 2025-09-15 11:17:35 -07:00
Wenxuan Tanandgemini-code-assist[bot] 2930abe456 [Bugfix] Fix VMoba requirements (#802)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-09-14 18:28:52 -07:00
William Lin b93ef4289d [bugfix] Fix empty PipelineConfigs for Wan2.2 A14B (#800) 2025-09-13 17:31:38 -07:00
401bdbd316 [self-forcing] [3/n] Text embed only preprocessing (#797)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-09-13 14:03:53 -07:00
William Lin 1048d79cf8 [bugfix] pin gradio version and set current_vsa_sparsity in TrainingPipeline (#798) 2025-09-11 17:04:47 -07:00
1e8406162d [bugfix] Fix delta calculation (#796)
Co-authored-by: zbchu2 <zbchu2@iflytek.com>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
2025-09-11 16:31:23 -07:00
William Lin 03edd35c83 [preprocessing] [self-forcing] [2/n] Improve preprocessing and add ode trajectory dataset schema (#794) 2025-09-10 17:33:57 -07:00
William LinandRandNMR73 ac11127397 [Self-forcing] [1/n] Handle extra dim in time embedding and add timestep warping (#792)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
2025-09-09 02:52:02 -07:00
Eric LiangandEricLiang e028dcc7c0 [Backend][Vmoba] Add implementation of VMoba (#778)
Co-authored-by: EricLiang <https://github.com/EricLina>
2025-09-08 23:53:25 -07:00
Wenxuan Tanandgemini-code-assist[bot] 076f45c1ee [Feature] Support Lora for DMD (#755)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-09-08 14:18:21 -07:00
85eb7265db fix: lora_B init zeros (#781)
Co-authored-by: zbchu2 <zbchu2@iflytek.com>
Co-authored-by: Wenxuan Tan <wenxuan.tan@wisc.edu>
2025-09-05 22:56:52 -07:00
William Lin d3ceb67e66 [misc] Update Slack invite link (#786) 2025-09-05 12:16:18 -07:00
Zhang Peiyuan 7ac153a5ca Update WeChat Link 2025-09-05 11:40:47 -07:00
William Lin d1e7aa0abd [CI] Add ssim test for causal inference (#784) 2025-09-05 01:23:01 -07:00
William Lin 2d846c55a1 [misc] Improve text encoding stage (#774) 2025-09-04 17:51:27 -07:00
Jinzhe Pan b318063c0a [Preprocess][Fix] video quality issue (#773) 2025-09-03 20:47:33 -07:00
Jinzhe Pan 4aa307be55 [Preprocess][Feat] support torchvision to load video in new preprocessing (#761) 2025-09-01 23:37:01 -07:00
William Lin 055e52e5ea [misc] [VSA] [STA] fix tk_root in setup.py for VSA and STA (#772) 2025-08-29 01:13:37 -07:00
William Lin 7d2069596b [bugfix] [VSA] [STA] Fix MANIFEST.in for VSA and STA; Move tk into both directories (#771) 2025-08-29 00:51:05 -07:00
William Lin c45009c9a4 [bugfix] fix STA install setup.py import (#770) 2025-08-28 23:02:53 -07:00
William LinandPeiyuan Zhang b91020b407 [VSA] [STA] Fix directory structure for pypi publishing (#769)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-08-28 22:34:03 -07:00
William Lin 2dcc5ea4f6 [chore] Release 0.1.6 (#768) 2025-08-28 20:56:21 -07:00
Wei ZhouandSolitaryThinker 359151d9a0 [Feature] Add wan2.2 5b i2v (#760)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-08-28 18:15:59 -07:00
Wei ZhouandSolitaryThinker ce67cd3729 [Feat] Support Self-Forcing's Causal Inference for Wan2.1 T2V 1.3B (#766)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-08-28 16:47:49 -07:00
Zhang Peiyuan 7c554e5da8 Update Community Link (#765) 2025-08-27 16:12:47 -07:00
William Lin 663ea33ff1 [bugfix] Fix wrong HF model string for FastWan2.2 5B (#763) 2025-08-26 22:05:40 -07:00
William Lin 3ef04f1654 [misc] [docs] Various fixes for logging and docs (#758) 2025-08-23 21:13:50 -07:00
Jinzhe Pan 0eced76a41 [Feat][Preprocess] support multi-gpus (#753) 2025-08-23 11:34:42 +08:00
Jinzhe Pan 3ab6470d1a [Feat][Preprocess] support merged dataset (#752) 2025-08-22 15:29:33 -07:00
Wenxuan Tan 989a03532c Optionally use unmerged weights for inference (#745) 2025-08-22 15:20:31 -07:00
William Lin fa15369a02 [bugfix] Check that model_index.json module is in required_modules list before removing (#756) 2025-08-22 14:36:44 -07:00
Zhang Peiyuan 78a9cb88d8 [Fix] fix seed in dmd denoising loop (#736) 2025-08-21 18:06:16 -07:00
Peng Xiaoand肖鹏 a0bff12746 [bugfix] [dmd] Align backward simulation with dmd2 sample back (#744)
Co-authored-by: 肖鹏 <xiaopeng1@aishi.ai>
2025-08-20 22:25:33 -07:00
William Lin 98f2af94e5 [bugfix] Missing Docker file for cuda12.9 (#750) 2025-08-20 15:34:31 -07:00
William Lin 46f7b6d574 [Docker] add 12.9 docker image and also fix py3.10 and py3.11 dockerfile (#749) 2025-08-20 15:31:15 -07:00
Jinzhe Pan 911a6a6a35 [Feat][Preprocessing] i2v preprocessing workflow (#737) 2025-08-14 20:47:25 -07:00
Zhang Peiyuan 38c7949d5c Update WeChat group link (#739) 2025-08-14 15:03:35 -07:00
Jinzhe Pan 7e7a0dba9d feat: preprocess validation dataset only when exist (#734) 2025-08-12 02:16:31 -07:00
Zhang Peiyuan f62e210ae6 Fix vsa backward gQ (#735) 2025-08-11 21:43:13 -07:00
William Lin 6ceb4942a0 [bugfix] [dmd] Fix backward simulation and also naming in wan_i2v_dmd_pipeline (#731) 2025-08-10 21:13:30 -07:00
William LinandRandNMR73 8cae5e4708 [feature] add Gradio live serving demo code (#727)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
2025-08-10 15:34:03 -07:00
William Lin 2a773fa34e [bugfix] [distill] remove i2v validation schema import in distill (#728) 2025-08-09 20:47:42 -07:00
Wenxuan Tan 5357f63327 Fix LoRA load from training checkpoint (#719) 2025-08-09 20:46:00 -05:00
William Lin 60f61c8101 [bugfix] fix pyproject install and VSA precision test (#726) 2025-08-08 18:45:03 -07:00
Jiali Chen 3d75ba8251 update version selection for VSA workflow (#725) 2025-08-08 13:05:16 -07:00
Wenxuan Tan 6c6bcd914d Remove all empty_cache (#713) 2025-08-07 22:50:38 -07:00
Jiali Chen f79b08de81 add cicd workflow for publishing VSA kernel (#723) 2025-08-07 18:53:05 -07:00
Jinzhe Pan f2bc037fff [Fix] training pipeline pin_cpu_memory issue (#692) 2025-08-07 02:31:20 -07:00
Jinzhe Pan 86604a684b [3/3][Preprocess] add preprocessing workflows (#645) 2025-08-07 01:49:07 -07:00
Zhang Peiyuan 47bd1e0178 [Misc] change installation logic of vsa (#721) 2025-08-06 21:54:09 -07:00
Wei ZhouandSolitaryThinker c41305ad18 [Feat] Add Wan2.2 14B MoE (#688)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-08-06 20:31:03 -07:00
Zhang Peiyuan 98ce9034f0 [Chore] Include our demo in the readme. (#720) 2025-08-06 19:29:40 -07:00
William Lin 0ceff110da [chore] Release 0.1.5 (#717) 2025-08-06 13:07:52 -07:00
Yongqi Chen 1d018acb3e [Feature]Add Data-free distillation readme (#710) 2025-08-05 14:27:39 -04:00
Yongqi Chen 7d8cf38dbe Fix typo (#709) 2025-08-04 20:21:21 -07:00
Yongqi Chen 8d483fe4aa [Bugfix] Fix neg_prompt bug when training from local cp (#708) 2025-08-04 15:54:06 -07:00
Zhang Peiyuan c1191250bf Add WeChat group link (#707) 2025-08-04 15:19:01 -07:00
Wenxuan Tan 4b7266349a [misc] Remove allow_tf32 in scripts (#705) 2025-08-04 15:37:56 -05:00
Yongqi Chen 22f9b7681f [Feature]Update Wan2.2+DMD doc example (#706) 2025-08-04 16:14:22 -04:00
Yongqi Chen 589d32cc39 [Feature] Update Readme and scripts (#703) 2025-08-04 15:02:32 -04:00
Hao Zhang 89199837db Update readme pre-release (#704) 2025-08-04 11:54:31 -07:00
William Lin d6ebaf1b49 [Docs] Fix README (#701) 2025-08-04 11:27:27 -07:00
Yongqi Chen fac927777c [Feature] Update readme (#702) 2025-08-04 14:27:18 -04:00
William Lin ecbd697dae [misc] Readme fixes (#699) 2025-08-04 10:20:57 -07:00
Yongqi Chen 7d4acef64d [Feature] Update sparse distill readme and doc (#700) 2025-08-04 10:16:20 -07:00
William Lin 9f0ce517cf [Docs] Update README and docs for FastWan (#698) 2025-08-04 09:05:18 -07:00
Yongqi Chen c718e56b0d [Feature] Remove unused args (#695) 2025-08-03 23:01:54 -04:00
Yongqi ChenandSolitaryThinker b65f0316d1 [Feature] Add Wan2.2 DMD example files; Update lr scheduler (#694)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-08-03 22:55:01 -04:00
William Lin 8d8bcb76b0 [config] Add config for FastWan2.2 ti2v 5B (#693) 2025-08-03 19:09:30 -07:00
Yongqi ChenandSolitaryThinker 5f42748ed1 [Feature] Add Wan2.2-TI2V-5B Sparse Distill (#690)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-08-03 01:29:57 -04:00
Yongqi Chen c9005045dc [Feature[[Readme] Add VSA/DMD doc (#673) 2025-08-02 02:35:46 -04:00
Wenxuan Tan 6c81befc87 [Feature] Optionally enable torch compile (#684) 2025-08-01 20:17:40 -07:00
Yongqi Chen dfe0b288e1 [Bugfix] Add i2v vae loading (#686) 2025-08-01 23:15:37 -04:00
Wenxuan Tanandgemini-code-assist[bot] 31200fbb83 [Misc] Fix training scripts (#683)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-08-01 15:54:31 -05:00
Yongqi Chen 9185978c55 [Bugfix] Fix multi-gpu training lr_scheduler (#682) 2025-08-01 15:54:03 -04:00
MartinPernus fcba463553 [Bugfix] fix _normalize_dit_input (#681) 2025-08-01 05:10:27 -04:00
Yongqi Chen 2c53d3eecf [Feature]Add DMD visualization for debugging (#674) 2025-07-31 05:54:47 -04:00
Zhang Peiyuan 516ecd374a [Misc] Update examples/ and other misc (#672) 2025-07-30 19:10:27 -07:00
Wei Zhou 3b1b54a74d Modify args to make sure the scripts are runnable on 4090 (#671) 2025-07-30 14:55:08 -07:00
Yongqi Chen 6914e7c904 [Bugfix]Fix DMD pipeline registry (#670) 2025-07-30 13:21:35 -07:00
Sopiko Kurdadze 5452369749 [Feature] [Inference]Add ROCm platform support for single-gpu inference (#669) 2025-07-30 12:57:02 -07:00
Yongqi Chen a113311e77 [Bugfix][Training]Fix Wan2.2 training vae config issue (#668) 2025-07-30 12:24:45 -07:00
Kevin Lin 44da97da92 [chore] Release 0.1.4 (#667) 2025-07-30 01:02:36 -07:00
Yongqi Chen f759980a58 [Feature]Add VSA slurm training example scripts (#666) 2025-07-30 01:27:54 -04:00
Zhang Peiyuan 37e0f8c236 [BUG] Fix distillation + vsa (#665) 2025-07-29 19:48:11 -07:00
Kevin Lin 51711d5906 [ComfyUI] Add __init__.py for node discovery (#663) 2025-07-29 18:21:09 -07:00
William LinandJerryZhou54 6375223b16 [Feature] Add wan2.2 5B T2V (#658)
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
2025-07-29 17:16:03 -07:00
Yongqi Chen 4cb046768d [Feature]Add DMD distillation training resume checkpoint; Update DMD CI test (#662) 2025-07-29 19:06:05 -04:00
Yongqi Chen 3322542444 [Feature] Add DMD CI test (#661) 2025-07-29 03:14:00 -04:00
Yongqi Chen 65f707354b [Bugfix]Fix mdoel inference checkpoint saving when enabling HSDP (#660) 2025-07-28 22:39:06 -07:00
Yongqi Chen 109e2e7e9d [Bugfix]Fix DMD wan pipeline (#659) 2025-07-28 21:50:44 -07:00
Zhang Peiyuan cbc3a6bb9d [Feat] Support VSA with any resolution. (#650) 2025-07-28 20:14:40 -07:00
Yongqi Chen 2fa8d4ae6d [Feature][Distill]Add 14B 480p T2V distill example scripts (#655) 2025-07-28 18:36:31 -04:00
Jinzhe Pan 7b6c8aee99 [2/3][Preprocess] refactor pipeline registry & file structure (#639) 2025-07-27 23:30:17 -07:00
Yongqi Chen 6284eaa363 [Feature][Distill]Add DMD+VSA joint training example (#654) 2025-07-27 18:18:01 -04:00
Yongqi Chen 636524e87f [Feature] Add Wan-14B-T2V-VSA CLI inference; add master port args (#653) 2025-07-27 07:13:44 -04:00
Yongqi Chen 202b2f3972 [Feature] Ignore [union-attr] and [override] mypy check and remove from training (#652) 2025-07-27 04:34:38 -04:00
Yongqi Chen 247fe273d8 [Feature] Add DMD T2V training pipeline (#651) 2025-07-27 03:35:51 -04:00
William Lin cb320dfa3a [bugfix] VideoGenerator improperly extracts output_video_name (#649) 2025-07-26 19:46:26 -07:00
Kevin Lin d8bb5abc46 [CI] Fix ComfyUI publisher ID (#648) 2025-07-25 19:02:52 -07:00
Kevin Lin cc703eca51 [CI] Add publish workflow for ComfyUI (#647) 2025-07-25 18:39:05 -07:00
William Lin 81c9df629c [core] Add offloading for vae and image encoder and rename offloading args (#643) 2025-07-25 17:55:03 -07:00
Yongqi Chen d3c0c52208 [Feature] Add prompt_txt support for CLI inference; Add DMD CLI inference (#646) 2025-07-25 19:44:10 -04:00
William Lin 744e0555c0 [misc] Use FASTVIDEO_STAGE_LOGGING for perf timing of stage (#644) 2025-07-25 16:20:53 -07:00
Jinzhe Pan 3a38f7dfdc [1/3][Preprocess] refactor preprocessing configs (#638) 2025-07-25 14:27:12 -07:00
William Lin f572319bd9 [Feature] Remove V1 folder (#642) 2025-07-24 22:43:12 -07:00
Wenxuan TanandWei Feng 48528f468c [Feature] Multi-lora inference (#640)
Co-authored-by: Wei (Will) Feng <134637289+weifengpy@users.noreply.github.com>
2025-07-24 21:01:28 -07:00
Yongqi ChenandSolitaryThinker 4264a80ca9 [Feature] Add DMD inference pipeline (#637)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-07-24 21:00:47 -07:00
William Lin 8573d4f05e [Docs] Docs update for Training and MPS (#641) 2025-07-24 19:23:58 -07:00
Wenxuan Tan 210a733515 [Bugfix] Fix LoRA trainable params and training ckpt loading (#630) 2025-07-23 20:01:40 -07:00
William Lin 0aef0e6f63 [bugfix] Fix preprocessing pipelines and nightly tests (#633) 2025-07-22 22:44:25 -07:00
Kevin Lin dd022ad9be [CI] Fix CI for pull request targets other than main (#632) 2025-07-22 21:00:03 -07:00
William Lin 832ad61e5b [bugfix] fa3 no longer returns lse (#631) 2025-07-22 18:31:34 -07:00
Wenxuan Tan 9419c04ee3 Fix lora train steps (#627) 2025-07-21 23:29:30 -05:00
Zhang Peiyuanandroot a37b39d83c Py/add triton block sparse (#593)
Co-authored-by: root <a1286225768@gmail,com>
2025-07-17 16:50:40 -07:00
Wenxuan Tan bb8c769c8e [LoRA] Support v1 LoRA training (#576) 2025-07-17 15:52:09 -05:00
William Lin 576c214f28 [v0] Remove V0 code (#621) 2025-07-15 22:05:16 -07:00
RandNMR73 b79d1fc15b video gen working on apple silicon (addressed issues from prior pr) (#595) 2025-07-15 22:04:27 -07:00
William Lin eb66e1c18d [chore] release 0.1.2 (#622) 2025-07-15 14:42:54 -07:00
William Lin 616d43c1cf [bugfix] [training] use separate generator for validation (#610) 2025-07-15 13:19:38 -07:00
Wenxuan Tan 7244a4b27f [CI] Add LoRA inference tests (#546) 2025-07-15 15:06:44 -05:00
Yongqi Chen 7e5ebb4582 [Feature][Training]Update example fine-tuning scripts to enable gradient checkpointing (#618) 2025-07-15 11:22:56 -07:00
Wenxuan Tan 65ed588570 Set encoder TP size to 1 by default (#569) 2025-07-09 17:44:24 -05:00
Wenxuan Tan 14adfe2edc Remove all unnecessary torch.cuda.empty_cache (#606) 2025-07-09 16:45:20 -05:00
William Lin 6198c6a640 [docs] update dev guide runpod image to py3.12 (#602) 2025-07-07 19:56:07 -05:00
William Lin e6b71b531b [docs] Update slack invite (#601) 2025-07-07 14:20:42 -05:00
William Lin ae1d112c6a [bugfix] [training] fix deadlock in latent datasets and init error in multi-node training (#598) 2025-07-06 01:34:15 -05:00
William Lin bf4de1f38f [chore] Upgrade min Python version from 3.8 to 3.10 (#597) 2025-07-04 22:08:31 -05:00
William Lin 66fdcc8e76 [Training] Use inference pipeline for training validation (#585) 2025-07-04 17:23:55 -05:00
Wenxuan Tan ed1e8d6bad [Feature] Offload all text encoders by default (#594) 2025-07-03 19:30:14 -05:00
Kevin Lin b9423ca3f8 Add ComfyUI custom node for inference (#596) 2025-07-03 16:09:03 -05:00
Wenxuan Tan ad16289871 [LoRA] Fix lora merge weights (#579) 2025-07-02 00:12:52 -05:00
Wenxuan Tan 2a41da1e6b Fix VAE precisions (#588) 2025-07-01 14:49:04 -05:00
William Lin 32133171da [chore] Release 0.1.1 (#592) 2025-07-01 01:43:12 -05:00
Kevin Lin 19674c6f29 [CI] Fix fork builds (#590) 2025-07-01 01:03:34 -05:00
Yongqi Chen 508afb7002 [docs] Update Readme (#591) 2025-06-30 23:48:45 -05:00
Yongqi Chen 288ea88105 [Feat][Training] Rename weight conversion function and update gradient checkpoint in scripts (#589) 2025-07-01 00:20:02 -04:00
Jinzhe Pan eb0f1318f3 [Feat] activation checkpointing (#584) 2025-06-30 15:24:29 -05:00
William Lin ce9b5910cc [Training] add caption to validation log (#582) 2025-06-30 02:42:17 -05:00
William Lin d0e5a6214a [misc] [training] Add --video_length_tolerance_range 10 to preprocessing scripts (#581) 2025-06-30 02:21:22 -05:00
Wenxuan Tan 834562b2db [CI] Fix pre-commit CI (#578) 2025-06-29 16:52:29 -05:00
Wei (Will) Feng 060cc7b9ba fully_shard usage on RMSNorm (#577) 2025-06-29 16:35:24 -05:00
Yongqi Chen 6c58a5ba62 [Bugfix]Fix VSA sp for training/inference (#574) 2025-06-29 13:44:33 -05:00
William Lin 48d9f61f86 [ci] [misc] fix training test threshold (#573) 2025-06-28 22:17:18 -05:00
William Lin 5f938b5844 [Revert] "[Feature] Load weights from distributed" (#571) 2025-06-28 20:55:14 -05:00
Wenxuan Tan 74da2a7370 Fix CLIP config (#568) 2025-06-28 19:01:23 -05:00
Kevin Lin 580d6dfe1f [CI] Add tests to Modal (#562) 2025-06-28 14:02:16 -05:00
Wenxuan Tan 344e43006a [CI] Fix SSIM and transformers CI (#564) 2025-06-28 00:26:20 -05:00
Wenxuan Tan c5155b256e [Feature] Load weights from distributed (#470) 2025-06-27 22:52:40 -05:00
William Lin e005c7f3ac [Docs] [Training] add readme for example training (#563) 2025-06-27 14:50:42 -05:00
Yongqi Chen ff5a79ef60 [Feature][Inference] Add VSA inference script (#561) 2025-06-27 02:19:23 -05:00
William Lin ab01dc4ba5 [Feature] [Training] Add i2v training (#559) 2025-06-27 01:56:50 -05:00
William Lin 285a950c1b [CI] fix vae and ssim tests (#557) 2025-06-26 23:53:01 -05:00
William Lin 46a0a85d85 [Training] Fixes SP for training; Improve Datasets and schema (#555) 2025-06-26 21:13:28 -05:00
Yongqi Chen 4aeabbc629 [Feature][Training] Add cfg rate for dataset loader (#556) 2025-06-26 18:22:37 -04:00
Wenxuan Tan 949bb5c835 [CI] Fix CI checks (#553) 2025-06-25 14:07:51 -05:00
Wenxuan Tan aab74c1271 [Kernel] Remove all syncs from STA & VSA kernels (#517) 2025-06-23 13:13:09 -07:00
Yongqi Chen f89d86944f [Feature][Training]Add diffusers format checkpoint saving for inference (#542) 2025-06-22 01:23:41 -04:00
William Lin 8741d204a5 [Training] Refactor and improve validation datasets (#539) 2025-06-21 17:58:35 -07:00
Wenxuan Tan cdc85f58a8 [chore] Bump torch to 2.7.1 to support Blackwell (#483) 2025-06-20 22:10:56 -07:00
William Lin 0262d2f089 [misc] [training] Reorganize training pipeline (#533) 2025-06-20 20:42:25 -07:00
William Lin 62c0343465 [bugfix] [VSA] Fix layernorm type for VSA Wan2.1 TransformerBlock (#534) 2025-06-20 00:24:51 -07:00
William Lin 1e1a023fb0 [bugfix] Fix stage validator for multi text encoder models (#535) 2025-06-19 22:49:16 -07:00
William Lin 1d2517ad8e [misc] Remove gradient checking code (#532) 2025-06-18 23:29:25 -07:00
William Lin d41186cb4a [Feat] Add Stage input and output verification (#523) 2025-06-18 23:29:11 -07:00
78e0c7eec9 Specify cu128 Pytorch installation (#530)
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
Co-authored-by: Wenxuan Tan <wenxuan.tan@wisc.edu>
2025-06-18 20:02:50 -05:00
Wenxuan Tan 1c41a94b62 [Refactor] Move dict_to_3d_list under utils (#507) 2025-06-18 13:34:37 -07:00
Yongqi Chen 2e66aafe20 [Bugfix][Readme]Fix readme website bugs and add VSA finetune docs (#531) 2025-06-17 22:48:29 -07:00
Yongqi ChenandWill Lin 55074bda76 [CI] Add STA-inference/VSA-training test (#527)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-06-17 21:13:06 -07:00
William Lin de65bec2b7 [Ci] add sta and vsa install to docker image (#528) 2025-06-17 18:09:48 -07:00
Yongqi Chen 7664dd0de3 [Bugfix][Inference]Fix envs.attn_backend (#525) 2025-06-17 18:38:06 -05:00
William Linandkevin314 019a88ced4 [CI][bugfix] Use new 3.12 docker image (#526)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-06-17 15:37:08 -07:00
Kevin Lin 72de11abcc [CI] Add current PR test workflow to Buildkite/Modal (#512) 2025-06-17 13:29:22 -07:00
Kevin Lin d71a4ebffc [CI] Update Docker image to flash-attn 2.8.0 / CUDA 12.8 (#524) 2025-06-16 17:48:23 -07:00
William Lin 1089ab43bf [bugfix] [Training] use diffusers fp32layernorm for wan2.1 (#490) 2025-06-15 22:45:48 -07:00
William Lin 97d4b984c9 [misc] [ci] fix e2e preprocess+training data path (#521) 2025-06-14 22:37:51 -07:00
Wenxuan Tan 2a8953d74d [Refactor] Fix attn backend selection not correctly setting env variable (#516) 2025-06-15 00:04:54 -05:00
Yongqi Chen 8801b10da7 [Bugfix][Preprocess]fix mini dataset name (#520) 2025-06-14 22:03:22 -07:00
William Lin 6b413f2ec4 [CI] [Training] drop negative prompt in validation dataset and CI test for preprocess + training overfit (#519) 2025-06-14 18:50:17 -07:00
Yongqi Chen 28b72694aa [Feature][Preprocess]Add Readme doc for preprocess (#518) 2025-06-14 20:41:13 -04:00
Yongqi Chen 4afb0cfe4f [Feature][Training]vsa for t2v training ready (#513) 2025-06-14 01:08:00 -04:00
Zhang Peiyuan 3eec1281cf [misc] Fix preprocessing and dataloader extra padding (#514) 2025-06-13 15:15:33 -07:00
Wenxuan Tan 0660489e38 [CI] Restrict training CI to v1 (#508) 2025-06-12 15:26:05 -07:00
Zhang Peiyuan dd871a17bf fix logging (#509) 2025-06-12 15:24:12 -07:00
dc11529862 [Refactor][Configurations] clean config orgnization (#505)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-06-12 13:27:08 -07:00
Zhang Peiyuan ffabf85e31 [feat] Add parquet iterable dataset. (#506) 2025-06-12 04:30:56 -04:00
William Lin c0026ca5ba [CI] [Training] Initial e2e small training test (#504) 2025-06-11 13:53:36 -07:00
Zhang Peiyuan 0f2bbe71ac [misc] rename dp_size to hdsp_replicate_dim (#491) 2025-06-10 16:36:56 -07:00
Yongqi ChenandJerryZhou54 2a46902ecb [Feature][VSA]Update STA publish workflow (#498)
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
2025-06-10 19:33:34 -04:00
Zhang Peiyuan 66012d3a4c [Feat][Dataloader] 1/n Refactor parquet map-style dataloader (#492) 2025-06-10 16:00:13 -07:00
William Lin f666b9de41 [misc] Add missing license headers (#499) 2025-06-10 14:25:32 -07:00
Yongqi Chen 7e3c073b55 [Feature] Adding VSA inference (#478) 2025-06-10 16:03:53 -04:00
Wei Zhou a6aa21bd07 [bugfix][Cli Inference] Resolve runtime errors when running fastvideo generate (#495) 2025-06-10 02:50:21 -04:00
Wenxuan Tan 6519b57aab [chore] Fix main pre-commit CI failure (#494) 2025-06-10 00:54:37 -05:00
Wei Zhou 675aea6ece [bugfix][Cli Inference] Resolve runtime errors when running fastvideo generate (#493) 2025-06-09 19:49:34 -07:00
Zhang Peiyuan 46e7a15e0d [misc] Improve distributed related env variables and setup (#487) 2025-06-08 09:14:48 -07:00
Yongqi ChenandJerryZhou54 e4f702d7ec [Bug] Fix multi gpus issues in v1 scripts (#489)
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
2025-06-07 21:55:49 -07:00
Wenxuan Tan bb68fcc809 Revert "Add torch.compile for all small ops" (#484) 2025-06-07 07:21:17 -05:00
Wenxuan Tan b392e6a874 Add torch.compile for all small ops (#432) 2025-06-06 21:10:42 -07:00
Zhang PeiyuanandWill Lin 0991003905 [bugfix] [misc] fix denoising stage init; rename distributed env function; fix logging. (#481)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-06-06 20:01:10 -07:00
Zhang PeiyuanandWill Lin 8f8ce6d9e1 [bugfix] [training] Add negative prompt to preprocessing and validation (#479)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-06-06 11:21:21 -07:00
KevinzzandBrianChen1129 e3d0cbe185 [STA] Implement mask search for V1's Wan2.1 (#415)
Co-authored-by: BrianChen1129 <yongqich@umich.edu>
2025-06-05 15:03:31 -04:00
eeecho d5ec468d43 [V0] [Distill] support distill in V0 for wan (#444) 2025-06-04 16:20:14 -07:00
Yongqi Chen e55fa6e5dc [Preprocess] I2V dataset (#473) 2025-06-04 16:19:20 -07:00
Wenxuan Tan 61b6ddeee1 [Issue template] Move env report to the end for readability (#476) 2025-06-04 16:18:22 -07:00
William Lin a9a000f45d [bugfix] fix bz >1 for training (#477) 2025-06-04 16:03:39 -07:00
Wenxuan Tan 66b8b8561e [LoRA] Support V1 LoRA inference (#451) 2025-06-04 16:19:29 -05:00
Wenxuan Tan 6684872616 [misc] Find unused port in distributed init (#475) 2025-06-04 14:15:12 -05:00
Wenxuan Tan 7f654e3332 [misc] Polish V1 training code (#469) 2025-06-03 22:02:03 -05:00
William Lin 8631c1b806 [Training] Support Multi-Node training with FSDP + SP (#459) 2025-06-03 17:40:22 -07:00
Wei Zhouand“BrianChen1129” 5357e12b5a Update v1 inference scripts (#467)
Co-authored-by: “BrianChen1129” <yongqich@umich.edu>
2025-06-03 02:54:12 -04:00
Kevin Lin bdfdf1dfee [Training] Add distributed checkpointing (#458) 2025-06-02 15:15:45 -07:00
Wei Zhou d156461785 Fix WanVideo (#461) 2025-06-01 23:08:11 -07:00
Yongqi Chen 6edf113838 [Misc] Bring back mask files under asset/ and update new Wan mask strategy file (#462) 2025-06-01 19:46:11 -07:00
William Lin dcf7738cbc [Misc] disable cast_forward_inputs (#460) 2025-05-31 22:06:29 -07:00
Wenxuan Tan b2ebaaf865 [Misc] Remove InferenceEngine (#455) 2025-05-31 11:12:05 -07:00
Wenxuan Tan 7768bb80f6 misc: add remote pdb for debugging workers (#456) 2025-05-30 20:44:27 -05:00
William Lin a335811869 [Training] [8/n] SP Training (#450) 2025-05-29 17:02:26 -07:00
William LinandZihang-He 357b0533fe [Training] [7/n] gradient clipping (#449)
Co-authored-by: Zihang-He <z6he@ucsd.edu>
2025-05-29 15:07:29 -07:00
William Lin 2ec3732758 [Training] [6/n]Mixed precision training (#448) 2025-05-29 14:34:28 -07:00
Wei Zhouand“BrianChen1129” a004408a93 [Training] [0/n] Add preprocessing pipeline (#442)
Co-authored-by: “BrianChen1129” <yongqich@umich.edu>
2025-05-29 14:30:09 -07:00
007e237e69 [Training] [5/n] Add single gpu training pipeline (#447)
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
Co-authored-by: Wei Zhou <69577934+JerryZhou54@users.noreply.github.com>
Co-authored-by: Kevin Lin <42618777+kevin314@users.noreply.github.com>
Co-authored-by: “BrianChen1129” <yongqich@umich.edu>
2025-05-29 11:49:46 -07:00
Yongqi Chen 8e18dc9f71 Update STA mask strategy downloading (#445) 2025-05-28 12:49:22 -07:00
7ab32539af [Training] [1/n] Add latent datasets (#438)
Co-authored-by: Wei Zhou <wzhou322@gatech.edu>
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
Co-authored-by: “BrianChen1129” <yongqich@umich.edu>
2025-05-28 11:01:52 -07:00
William Lin 6ef8fcb61d [Training] [4/n] add training save checkpoint (#441) 2025-05-27 17:53:53 -07:00
William Lin 016e24da63 [Training] [3/n] Add training args and dependencies (#440) 2025-05-27 17:53:39 -07:00
William Lin 85b8717545 [Training] [2/n] add bwd for all2all and all_gather (#439) 2025-05-27 14:27:54 -07:00
Wenxuan Tan 657fd745e1 misc: Trigger transformers CI for layers and attention code change (#434) 2025-05-27 11:43:23 -07:00
applesaucethebunandBrayden Zhong 12647457a7 [Misc] Small fixes to Torch code (#395)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2025-05-23 14:40:24 -07:00
Kevin Lin 298f74f956 Set device for encode (#420) 2025-05-23 14:19:45 -07:00
Wenxuan Tan ee8babb298 Unify env report script in issue template (#423) 2025-05-23 14:19:18 -07:00
Wenxuan Tan 60295cc03f Use version.py (#424) 2025-05-23 14:17:33 -07:00
William Lin 1572e13b6e [Tests] don't run 3.10 and 3.11 for SSIM (#427) 2025-05-23 12:59:50 -07:00
Wenxuan Tan a157275b4c Fix version number (#422) 2025-05-22 12:37:34 -07:00
William Lin c4dbe7dac3 [bug] fix bs > 1 (#418) 2025-05-21 21:07:01 -07:00
Kevin Lin d39591108e Fulfill worker response on interrupt (#417) 2025-05-21 20:59:11 -07:00
William Lin ace6e971e5 [V1] Remove vLLM dependency (#413) 2025-05-18 00:57:28 -07:00
William Lin b4f6758253 [Teacache] allow None for forward_context batch when using teacache (#412) 2025-05-17 18:43:14 -07:00
William Lin 535d29b392 [Docs] Fix image (#407) 2025-05-12 14:30:05 -07:00
William Lin b4255517e0 [Docs] Add CLI docs (#406) 2025-05-12 14:16:48 -07:00
William Lin 6eeb60613f Release 0.1.0 (#405) 2025-05-12 11:54:23 -07:00
William Lin 53d2c7791f [V1] Update where num_frame rounding is done (#403) 2025-05-12 11:52:49 -07:00
William Lin 53cb693dca [V1] Docs Update (#402) 2025-05-12 11:52:09 -07:00
Kevin Lin 6f72d24876 [CLI] Default to pipeline config (#401) 2025-05-12 00:22:53 -07:00
William Lin d1459e9976 [V1] Update README (#400) 2025-05-11 22:32:03 -07:00
William Lin 59ab481eb1 release 0.0.5 (#399) 2025-05-11 16:23:11 -07:00
William Lin 0cf001986a [Docs] More docs update (#394) 2025-05-11 16:19:47 -07:00
applesaucethebunandBrayden Zhong 51956369a5 [Misc] Replace instances of time.time() with time.perf_counter() (#396)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2025-05-11 15:56:39 -07:00
Kevin Lin 4b0970cbbf [CI] Set volume_size required to false (#398) 2025-05-11 15:55:34 -07:00
Kevin Lin 94bf47a572 [CI] Set default disk size (#397) 2025-05-11 15:34:17 -07:00
River (Zihang He)andWill Lin 1a3ac9074b Zihang stepvideo v1 (#389)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-05-11 14:41:42 -07:00
Kevin Lin fb0581d5b0 [CLI] Update cli to support new api/model config (#384) 2025-05-11 13:50:33 -07:00
Kevin Lin 6c74ab4132 [CI] Use python 3.10/3.11 for SSIM test (#392) 2025-05-10 13:48:29 -07:00
William Lin 9a91021c56 [Docs] Add collect_env.py and various docs update (#393) 2025-05-10 13:42:42 -07:00
Wei Zhou 51c94d6a73 [Misc] Small Fixes & Features (#390) 2025-05-07 15:40:49 -07:00
William Lin dba38dbc03 Release 0.0.4 (#388) 2025-05-07 02:40:42 -07:00
William Lin 2034cc3c4f [misc] Improve worker cleanup (#387) 2025-05-06 22:45:22 -07:00
William Lin c69afce2f6 [Docs] Update for V1 (#381) 2025-05-06 16:34:44 -07:00
Wei Zhou b08e758eb3 Small Fixes & Features (#378) 2025-05-06 15:52:34 -07:00
William Lin 3f3462d7ce Cleanup Teacache params (#386) 2025-05-06 15:52:05 -07:00
William Lin c9c47dd89c Add Teacache to V1 (#371) 2025-05-06 11:58:56 -07:00
William Lin 9c4ef7c2f1 [Lint] fix (#382) 2025-05-06 00:42:55 -07:00
Kevin Lin f25eb4b905 [CI] Add write permissions to build-image workflow (#379) 2025-05-05 19:21:03 -07:00
Kevin Lin 048d55ccbb [CI] Add new images for different Python versions (#377) 2025-05-05 10:28:33 -07:00
Wei Zhou a271c55fe4 Fix FSDP issues when using cpu_offload flag (#376) 2025-05-02 15:54:52 -07:00
William Lin f663ae0d8a change gradio example to use model configs (#375) 2025-05-02 10:42:03 -07:00
William Lin c0911aa3dd release 0.0.3 (#374) 2025-05-02 03:50:41 -07:00
William Lin 5f59687ae7 Fix model config for python 3.11+ (#373) 2025-05-02 03:18:56 -07:00
Wei Zhou 5adbc81cdc Refactor encoder (#370) 2025-05-02 02:41:12 -07:00
Kevin Lin f26d5c37c1 Update SSIM tests to use new API (#369) 2025-05-01 04:14:24 -07:00
Wei Zhou 6a4ef42378 [V1] Model config (#358) 2025-04-30 14:48:22 -07:00
William Lin f1098c77dc [Attn] Add SageAttention Backend (#366) 2025-04-28 12:40:07 -06:00
William Lin 0405b618f8 [Docs] Docs for design and adding new pipeline (#363) 2025-04-24 01:04:10 -07:00
Kevin Lin eac79b753f [V1] Worker improvements/cleanup (#361) 2025-04-22 00:32:43 -07:00
William Lin 4d58cf20d0 chore: Release FastVideo 0.0.2 and update python requirements (#360) 2025-04-21 14:10:48 -07:00
Kevin LinandWill Lin 52c93ecc9d [V1] Gradio demo with new API (#357)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-04-19 18:14:24 -07:00
William Lin 42d63166ac [V1] Process aware logging; improve logging msg (#356) 2025-04-19 15:03:29 -07:00
William Lin 6db20345a2 [V1] Worker cleanup; Logging clean up; enables isort again (#355) 2025-04-18 19:26:47 -07:00
William Lin ad27ea596c [sta] release 0.0.4 (#354) 2025-04-18 14:54:40 -07:00
William Lin 9aadb4bf8c [1/n] [v1] Add Worker abstractions for User API (#336) 2025-04-18 14:38:46 -07:00
Kevin Lin bd941df271 [Docs] Fix developer guide images (#353) 2025-04-17 22:32:19 -07:00
Yongqi Chen 8a73876d3b add STA to Wan v1 (#349) 2025-04-17 16:35:19 -07:00
Kevin Lin 1483a1138a [CLI] Fix duplicate --num-gpus (#352) 2025-04-17 13:01:48 -07:00
Wei Zhou 5e243d8292 Default to using original WanVAE's encoding/decoding algorithm (#351) 2025-04-17 13:00:25 -07:00
Kevin Lin b0c66d3200 [CI] Docker image improvements (#350) 2025-04-17 12:27:10 -07:00
Wei ZhouandWill Lin c86da2c736 [core] Pipeline config (#343)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-04-15 15:43:10 -07:00
Kevin Lin bae2a19dcf [CI] Add manual trigger to sta-publish and fastvideo-publish (#346) 2025-04-15 15:39:11 -07:00
Kevin Lin 057686f59d [CI] Free up runner disk for sta-publish (#345) 2025-04-15 15:27:09 -07:00
William Lin 67da56628b [STA] Sta release 0.0.3 (#344) 2025-04-15 13:06:03 -07:00
Zhang PeiyuanandSolitaryThinker 2325adffa2 Add STA to V1 (#312)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2025-04-15 02:39:58 -07:00
Kevin Lin 20cf836ef1 [CI] Support custom Docker image (#342) 2025-04-14 20:08:53 -07:00
William Lin 2bf69b6f92 [Docs] Initial examples setup and more docs (#332) 2025-04-14 16:53:57 -07:00
William Lin 13583f5ffb [Model] Remove RMSNorm's forward_native hardcode from Wan (#339) 2025-04-14 16:46:43 -07:00
William LinandJerryZhou54 008ee2099a V1 wan rebased (#335)
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
2025-04-11 15:34:54 -07:00
Kevin Lin 137f61f2fe Port tests to v1 (#333) 2025-04-11 01:40:01 -07:00
William Lin ccb262974e [Docs] Add dev guide and doc building CI (#330) 2025-04-09 13:09:00 -07:00
Kevin Lin 30966e3bc9 [CI] Set allowedCudaVersions (#329) 2025-04-09 10:16:05 -07:00
William Lin 7b4272d6b7 [Docs] Fix doc lint (#325) 2025-04-09 10:14:53 -07:00
William Linandkevin314 15553f7706 [CI] Use pre-commit to run linter (#321)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-08 11:38:47 -07:00
William LinandPorridgeSwim 60eeea50bb [Docs] Initial Docs Build (#322)
Co-authored-by: PorridgeSwim <yz3883@columbia.edu>
2025-04-08 11:38:36 -07:00
Kevin Linandkevin314 927b3a40b9 [CI] Add manual triggers for PR workflow (#320)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-07 14:25:39 -07:00
William Linandkevin314 c64f826ae2 Add torch sdpa backend to ssim test (#316)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-07 13:50:46 -07:00
Zhang Peiyuan 55c1040f0b Fix sdpa (#315) 2025-04-06 12:00:34 -07:00
Kevin Linandkevin314 8a3e7aa761 Add ssim test (#314)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-05 17:35:45 -07:00
You ZhouandWill Lin 4324c1c21d refactor the env setup and install of fastvideo (#309)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2025-04-04 16:33:03 -07:00
Kevin Linandkevin314 2c342ee37f [CI] Add test workflow improvements (#311)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-02 23:14:05 -07:00
Kevin Linandkevin314 708201f531 Set up text encoder tests to work with pytest and Github Actions (#302)
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
2025-04-01 17:56:21 -07:00
1fee098f10 [do not merge] Rebased refactor (#270)
Signed-off-by: <>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
Co-authored-by: Zhou, Wei <wzhou322@gatech.edu>
Co-authored-by: Kevin Lin <42618777+kevin314@users.noreply.github.com>
Co-authored-by: kevin314 <kevin.lin.cs1@gmail.com>
Co-authored-by: JerryZhou54 <69577934+JerryZhou54@users.noreply.github.com>
Co-authored-by: Yongqi Chen <144848849+BrianChen1129@users.noreply.github.com>
Co-authored-by: Peiyuan Zhang <m2deng@ucsd.edu>
2025-03-29 17:43:47 -05:00
You Zhou 8a77cf22c9 Establish cicd workflow to build and publish FastVideo and STA Kernel (#227) 2025-03-11 20:27:36 -07:00
Yongqi ChenandPeiyuan Zhang d869d90d12 fix training mask strategy issue (#248)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-03-05 20:00:16 -08:00
Zhang Peiyuan 554ee17de5 [BUG] update cfg bug? (#223) 2025-02-27 16:02:44 -08:00
Yongqi ChenandPeiyuan Zhang 0be4fc62c9 fix train/distill issue (#215)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-25 08:11:17 -08:00
Yongqi ChenandPeiyuan Zhang 1e08893546 Added multi-GPU support for Hunyuan STA (#211)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-21 14:16:28 -08:00
Zhang Peiyuan 09ab452610 Update STA README.md (#206) 2025-02-20 22:26:26 -08:00
Yongqi ChenandPeiyuan Zhang e768b5ec5b Update readme (#202)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-20 13:16:25 -08:00
Zhang Peiyuan 59ec42f40e [FIX] Make STA optinal (#204) 2025-02-20 13:09:50 -08:00
rlsu9 5ae5b247b3 [FIX] fix isort format (#203) 2025-02-20 12:20:20 -08:00
ead6c62be4 [Feat] Add STA for StepVideo (#200)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
Co-authored-by: BrianChen1129 <yongqich@umich.edu>
2025-02-20 11:33:58 -08:00
Yongqi ChenandPeiyuan Zhang 6805eaa06c [bug]: fix ori hunyuan inference issue (#199)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-19 14:18:15 -08:00
Zhang Peiyuan c39a15551c Update typo (#198) 2025-02-18 19:34:45 -08:00
Zhang Peiyuan e6dda263b0 Update Cite (#195) 2025-02-18 21:01:46 -05:00
Zhang Peiyuan f9482d113c update env (#194) 2025-02-18 20:45:08 -05:00
rlsu9 a3ec969397 [feat]: fix readme demo and add video to readme (#191) 2025-02-18 17:46:32 -05:00
Yongqi ChenandPeiyuan Zhang 76a12cc8a1 Infer sta tea with torch.compile (#190)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-18 11:29:36 -08:00
Yongqi ChenandPeiyuan Zhang ac490399c6 fix kernel issue (#185)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-16 21:35:56 -08:00
Yongqi ChenandPeiyuan Zhang 9ea39cee57 Add STA and teacache forward (#184)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2025-02-15 16:22:01 -08:00
Zhang Peiyuanandrlsu9 52e6e612a2 add sliding tile attn (#182)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2025-02-15 15:44:34 -08:00
Hangliang Ding 9aebc4ada1 Create config.yml (#152) 2025-01-20 20:11:01 -08:00
Yongqi Chen b53cf7425c Lora README update (#155) 2025-01-18 12:30:53 -08:00
Zhang Peiyuan d9ce056901 [typo] 2025-01-13 20:05:57 -08:00
Brian Chen 218449c54d adding hunyuan hf (support lora finetuning); unified hunyuan hf inference with quantization (#135) 2025-01-13 19:47:42 -08:00
Hangliang Ding 221958bcde Update README.md (#131) 2025-01-08 09:02:40 -08:00
Yuzhou Nieand“Peiyuan Zhang” 4a1f1e35bb add parallel for vae decoding (#134)
Co-authored-by: “Peiyuan Zhang” <a1286225768@gmail.com>
2025-01-07 17:14:21 -08:00
rlsu9 e0e05f97f2 [feat]: Add tests for FastVideo (#127) 2025-01-06 12:27:39 -08:00
Zhang Peiyuan dd75ee8509 [Fix] Save CK, Dataset bug fix (#125) 2024-12-31 22:19:10 -08:00
rlsu9 0aed1868df [feat]: Add format auto fixer to main branch (#124) 2024-12-31 15:23:17 -08:00
Hangliang Ding d467c7cd35 [Minor] Adding issue template. (#114) 2024-12-25 21:50:57 -08:00
Zhang Peiyuanandrlsu9 88b2583c2c [feat]:Single 4090 inference for fasthunyuan (#104)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-25 12:40:16 -08:00
rlsu9 a730e43d5f Update README.md layout 2024-12-19 13:36:43 -08:00
Brian Chen edf116fa46 fix lora checkpoint saving issue (#97) 2024-12-19 08:42:59 -08:00
Luis Catacora de3cefb5e5 Add Replicate demo and API (#93) 2024-12-18 19:56:09 -08:00
Hangliang Ding e087e85e09 Adding Development plan 2024-12-18 16:46:14 +08:00
Your Name e1b998b6ef merge 2024-12-17 12:48:16 -08:00
rlsu9 fb49c93dbc Update README.md 2024-12-17 12:29:03 -08:00
rlsu9 172f4802b4 Update README.md 2024-12-17 12:28:08 -08:00
rlsu9 24e57fafc9 Update README.md 2024-12-17 12:26:17 -08:00
Your Name 6debd46482 merge docs 2024-12-17 12:20:42 -08:00
rlsu9 f7dc36f7ea Update README.md 2024-12-17 12:13:33 -08:00
Brian Chen a0fb954f56 Update README.md
fix typo
2024-12-17 15:09:10 -05:00
rlsu9 053106922c Update README.md 2024-12-17 11:43:49 -08:00
a57122c519 Rlsu lora readme (#86)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
Co-authored-by: rlsu9 <147024991+rlsu9@users.noreply.github.com>
2024-12-17 11:37:07 -08:00
Zhang Peiyuanandrlsu9 b393570e45 Update README (#85)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-16 17:06:14 -08:00
Zhang Peiyuanandrlsu9 285635e8c0 Clean up (#84)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-15 20:33:11 -08:00
Zhang Peiyuanandrlsu9 58cfd71b5e Cleanup
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-15 17:03:29 -08:00
Hangliang Dingandrlsu9 3bf892b6ab update release readme (#81)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-15 22:24:13 +08:00
Zhang Peiyuan 85639d1101 [feat] add hunyuan adv (#79) 2024-12-13 11:52:57 -08:00
Zhang Peiyuanandforeverpiano 6ab2263f3a [Feat] Add HunyuanVideo (#78)
Co-authored-by: foreverpiano <pianoqwz@qq.com>
2024-12-12 14:14:09 -08:00
Zhang Peiyuanandrlsu9 b421c2e183 Cleanup (#77)
Co-authored-by: rlsu9 <r3su@ucsd.edu>
2024-12-12 14:04:02 -08:00
233 changed files with 15797 additions and 9273 deletions
+112 -29
View File
@@ -1,17 +1,12 @@
ucf101_stride4x4x4
__pycache__
*.mp4
.ipynb_checkpoints
*.pth
UCF-101/
results/
build/
fastvideo.egg-info/
wandb/
.idea
*.ipynb
*.jpg
*.mp3
!examples/dataset/lingbotworld2/image.jpg
*.safetensors
*.mp4
*.png
@@ -20,35 +15,123 @@ wandb/
*.pt
cache_dir/
wandb/
test*
sample_video*
sample_image*
512*
720*
1024*
debug*
private*
caption*
*deepspeed*
revised*
129f*
all*
read*
YSH*
*pick*
*ysh*
hw*
257f*
513f*
taming*
221hw*
65x512x512
venv/
.venv/
runs/
samples/
Miniconda3-latest-Linux-x86_64.sh
*validation/
data/
outputs/
outputs_video
checkpoints/
sbatch.sh
*.out
env
*.o
**/build/
**.pyc
**.txt
*.log
weights/
logs/
/Z-Image/
official_weights/
converted_weights/
# SSIM test outputs
fastvideo/tests/ssim/generated_videos/
**/.cache/**
# Distribution / packaging
build/
dist/
*.egg-info/
*.egg
eggs/
.eggs/
# MkDocs documentation
site/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
docs/training/examples/
docs/distillation/examples/
!requirements-mkdocs.txt
# VSCode
.vscode/
# DS Store
.DS_Store
# vim swap files
*.swo
*.swp
# Python pickle files
*.pkl
# Reference videos (negations must come after the catch-all on line below)
# Static images
!docs/assets/images/**/*.png
!comfyui/assets/**/*.png
!comfyui/assets/**/*.gif
!assets/images/**/*.png
!assets/images/**/*.jpg
!assets/images/**/*.jpeg
!assets/images/**/*.gif
!assets/videos/**/*.mp4
dmd_t2v_output/
preprocess_output_text/
# SvelteKit / Node artifacts under apps/fastvideo_studio/: see apps/fastvideo_studio/.gitignore
# Next.js / Node artifacts under apps/dreamverse/web/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/coverage/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/web/.env.local
apps/dreamverse/web/.env.development.local
apps/dreamverse/web/.env.test.local
# Generated by apps/dreamverse/scripts/install_native_ffmpeg.sh — host-specific
apps/dreamverse/scripts/ffmpeg-env.sh
apps/dreamverse/web/.env.production.local
# Unignore migrated Dreamverse product assets — root .gitignore globally
# ignores *.png/*.jpg/*.mp4/*.gif, but apps/dreamverse/web/public/ MUST
# be tracked (logo, icons, k2.png, etc.).
!apps/dreamverse/web/public/**/*.png
!apps/dreamverse/web/public/**/*.jpg
!apps/dreamverse/web/public/**/*.jpeg
!apps/dreamverse/web/public/**/*.mp4
!apps/dreamverse/web/public/**/*.gif
!apps/dreamverse/web/prompts/**/*.png
!apps/dreamverse/web/prompts/**/*.jpg
!apps/dreamverse/web/prompts/**/*.jpeg
!apps/dreamverse/web/prompts/**/*.mp4
!apps/dreamverse/web/prompts/**/*.gif
!apps/dreamverse/gpu-pool.svg
!apps/dreamverse/gpu-pool.drawio
.claude/
.codex/
.agents/tmp/
.sisyphus/
openspec/
fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
.python-version
+10
View File
@@ -0,0 +1,10 @@
# fastvideo2/rl_rewards is vendored byte-identical from upstream (gated by
# sha256 in tests) — formatters must not touch it
exclude: ^fastvideo2/rl_rewards/
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v5.0.0
hooks:
- id: trailing-whitespace
- id: end-of-file-fixer
- id: check-merge-conflict
+89
View File
@@ -0,0 +1,89 @@
# Repository guidelines (branch `will/v2.1`)
This branch is the fastvideo2 MVP — one package, one model (Wan2.1), four
surfaces. Read `README.md` first.
## Layout
| Path | Role |
|---|---|
| `fastvideo2/card.py` | Frozen data cards, `derive()`, digest, validation (stdlib-only) |
| `fastvideo2/loop.py` | Driven-loop protocol + `LoopRunner` (stdlib; NVTX lazily) |
| `fastvideo2/pipeline.py` | Stage list with enforced `reads`/`writes` (stdlib-only) |
| `fastvideo2/loading.py` | Checkpoint → modules, standalone; component fingerprints |
| `fastvideo2/layers/` | Shared model layers (norms, MLP, rotary, attention) — torch-only, checkpoint-key compatible, cast semantics preserved (anchor-proven) |
| `fastvideo2/engine.py` | One-shot runner: request → outputs + identity-chained trace |
| `fastvideo2/verify.py` | Gates T0–T3 + evidence ledger |
| `fastvideo2/registry.py` | The only catalog: name → (card, pipeline builder) |
| `fastvideo2/wan21/` | The family's logic: card constant, loop, pipeline, vendored `model.py`, `reference.py` (the executable spec) |
| `fastvideo2/wan21/gates/` | The family's measurement side: goldens, anchor adapters, official capture shim, comparison CLI, diagnostics — never imported by logic code |
| `fastvideo2/evidence/` | Append-only ledger + blessed baselines (see its README) |
| `fastvideo2/tests/` | T0 contract tests — CPU, no torch, no weights |
## Invariants (enforced by review; violating them is the bug)
1. **Cards are pure data.** No callables, no live objects, no deploy-local
paths. If it can't round-trip through JSON, it doesn't belong on a card.
2. **Import direction is one-way:** `card` → `loop`/`pipeline` → `engine` →
`verify`. Family packages depend on core, never the reverse.
`wan21/reference.py` imports only the vendored official model file
(`wan21/model.py`, itself standalone) — never core/runtime modules — and
nothing outside `verify.py` may import the reference.
3. **Loop modules import torch-free** (torch inside methods) so contracts
validate anywhere; `import fastvideo2` must work without torch installed.
4. **Model-specific inputs are typed** (`WanForwardInputs`); never add an
untyped passthrough kwarg to a forward call.
5. **Evidence is append-only and human-owned.** Agents run `verify` and commit
the records; agents do not edit tolerances, re-bless baselines, or touch
`reference.py` to make a failing gate pass — say so instead.
6. **One catalog.** New servable ⇒ card constant + registry entry. No parallel
model lists.
7. **Official implementations are the numerics authority.** Where fidelity is
the requirement, run the authors' modeling code: the Wan DiT is vendored
from the pinned official commit (`wan21/model.py`, provenance in its
header). Restructuring vendored code (e.g. extracting `layers/`) is
allowed ONLY when the anchor stays bitwise 0.0 — the gate, not "verbatim",
is the equivalence guarantee. Two invariants for any extraction:
checkpoint keys unchanged (Sequential indices preserved) and cast/dtype
semantics unchanged (fp32 islands, promotion order, fp64 RoPE). Ports
(diffusers, etc.) may serve components only with anchor certification,
never on trust; the official repo never becomes a dependency or submodule;
when conventions conflict, official wins.
8. **One environment for goldens and gates.** The supported env is python 3.12
+ torch 2.12 (the fastvideo cluster venv). Goldens are captured with that
same env — official code rides `PYTHONPATH`, its extra deps go to a pip
`--target` dir — so anchor deltas measure implementation differences, never
torch/kernel version differences.
## Commands
```bash
pytest # T0, runs on a laptop
python -m fastvideo2 verify <model> --tier N # gates; appends evidence
python -m fastvideo2 describe <model> # card JSON + digest
python -m fastvideo2 generate <model> --prompt ... # one request
```
GPU work runs on dlcluster via the `run-fastvideo-dlcluster` skill from the
main FastVideo checkout (sync this branch with `git push origin HEAD`, then
run inside the branch clone at `/mnt/fv21` — do not disturb `/mnt/FastVideo`'s
checkout). One-time per environment, install the package editable with no
dependency changes (torch etc. already live in the venv) — after this,
scripts and `fastvideo2 <cmd>` work from any directory, no PYTHONPATH:
```bash
/mnt/FastVideo/.venv/bin/pip install -e /mnt/fv21 --no-deps -q
```
Cluster runs append to `fastvideo2/evidence/` and those files get
fetched and committed locally, so before every cluster `git pull`, reset that
tree or the pull conflicts:
```bash
git checkout -- fastvideo2/evidence; git clean -qfd fastvideo2/evidence; git pull
```
## Commit style
Short subject with a tag prefix (`[feat]: ...`, `[fix]: ...`, `[docs]: ...`).
Do not add AI co-author trailers.
+1 -15
View File
@@ -184,18 +184,4 @@
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [2023] Lightning AI
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
identification within third-party archives.
+59 -149
View File
@@ -1,162 +1,72 @@
# FastVideo
# fastvideo2 — the v2.1 MVP (branch `will/v2.1`)
<div align="center">
<a href=""><img src="https://img.shields.io/static/v1?label=API:H100&message=Replicate&color=pink"></a> &ensp;
<a href=""><img src="https://img.shields.io/static/v1?label=Discuss&message=Discord&color=purple&logo=discord"></a> &ensp;
</div>
<br>
<div align="center">
<img src=assets/logo.png width="50%"/>
</div>
A from-scratch, deliberately small substrate for the FastVideo big bet:
**post-training → inference-optimized serving**, designed so that both kinds of
agents — the ones that *build* the framework and the ones that will *operate*
video models inside products — get inspectable contracts, a ground-truth
oracle, machine-readable verification, and an identity-chained runtime.
FastVideo is a scalable framework for post-training video diffusion models, addressing the growing challenges of fine-tuning, distillation, and inference as model sizes and sequence lengths increase. As a first step, it provides an efficient script for distilling and fine-tuning the 10B Mochi model, with plans to expand features and support for more models.
This branch is a clean slate: everything except `LICENSE` was removed, and the
MVP supports exactly one model, **Wan2.1-T2V-1.3B**, end to end.
### Features
## The four surfaces
- FastMochi, a distilled Mochi model that can generate videos with merely 8 sampling steps.
- Finetuning with FSDP (both master weight and ema weight), sequence parallelism, and selective gradient checkpointing.
- LoRA coupled with pecomputed the latents and text embedding for minumum memory consumption.
- Finetuning with both image and videos.
| Surface | Where | What it guarantees |
|---|---|---|
| **Contracts** | `fastvideo2/card.py`, `pipeline.py`, `loop.py` | Cards are frozen *data* (no callables): JSON round-trip, stable content digest — the identity used by deploy configs, trainers, and RL environment manifests alike. Pipeline stage edges (`reads`/`writes`) are enforced at run time. Loop classes carry a `semantics` id and provenance pins it: distilled weights cannot silently run under a base sampler. |
| **Reference** | `fastvideo2/wan21/reference.py` | The complete model in one standalone eager file — the textbook an agent copies, and the oracle the production path is measured against. Never imported by production code. |
| **Verifier** | `fastvideo2/verify.py`, `fastvideo2/evidence/` | Tiered gates: T0 contracts (CPU, seconds) → T1 component fingerprints vs a blessed baseline → T2 trajectory parity vs the reference, with tolerance calibrated by measured run-to-run self-noise (the determinism contract) → T3 decoded-output parity + anti-degeneracy anchors. Every run appends typed records (card digest + env fingerprint) to the evidence ledger. |
| **Trace** | `fastvideo2/engine.py`, `loop.py` | Every unit of work is named `request/stage/loop.step`; the same identity chain lands in the returned trace (typed timings) and in nested NVTX ranges, so Nsight correlates kernels to model-level identity for free. |
## Change Log
- ```2024/12/06```: `FastMochi` v0.0.1 is released.
## Fast and High-Quality Text-to-video Generation
### 8-Step Results of FastMochi
<table class="center">
<td><img src=assets/8steps/1.gif width="320"></td></td>
<td><img src=assets/8steps/2.gif width="320"></td></td></td>
<tr>
<td style="text-align:center;" width="320">tmp</td>
<td style="text-align:center;" width="320">tmp</td>
<tr>
</table >
## Table of Contents
Jump to a specific section:
- [🔧 Installation](#-installation)
- [🚀 Inference](#-inference)
- [🎯 Distill](#-distill)
- [⚡ Finetune](#-lora-finetune)
## 🔧 Installation
```
conda create -n fastmochi python=3.10.0 -y && conda activate fastmochi
pip3 install torch==2.5.0 torchvision --index-url https://download.pytorch.org/whl/cu121
pip install packaging ninja && pip install flash-attn==2.7.0.post2 --no-build-isolation
pip install "git+https://github.com/huggingface/diffusers.git@bf64b32652a63a1865a0528a73a13652b201698b"
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo && pip install -e .
```
## 🚀 Inference
Use [scripts/download_hf.py](scripts/download_hf.py) to download the hugging-face style model to a local directory. Use it like this:
```bash
python scripts/download_hf.py --repo_id=FastVideo/FastMochi --local_dir=data/FastMochi --repo_type=model
```
Start the gradio UI with
```
python fastvideo/demo/gradio_web_demo.py --model_path data/FastMochi
```
We also provide CLI inference script featured with sequence parallelism.
```
export NUM_GPUS=4
torchrun --nnodes=1 --nproc_per_node=$NUM_GPUS \
fastvideo/sample/sample_t2v_mochi.py \
--model_path data/FastMochi \
--prompt_path assets/prompt.txt \
--num_frames 163 \
--height 480 \
--width 848 \
--num_inference_steps 8 \
--guidance_scale 1.5 \
--output_path outputs_video/demo_video \
--seed 12345 \
--scheduler_type "pcm_linear_quadratic" \
--linear_threshold 0.1 \
--linear_range 0.75
```
For the mochi style, simply following the scripts list in mochi repo.
```
git clone https://github.com/genmoai/mochi.git
cd mochi
# install env
...
python3 ./demos/cli.py --model_dir weights/ --cpu_offload
```
## 🎯 Distill
## 💰Hardware requirement
- VRAM is required for both distill 10B mochi model
To launch distillation, you will first need to prepare data in the following formats
## Quickstart
```bash
asset/example_data
├── AAA.txt
├── AAA.png
├── BCC.txt
├── BCC.png
├── ......
├── CCC.txt
└── CCC.png
# machine-readable capability discovery
python -m fastvideo2 describe wan2.1-t2v-1.3b
# contracts only — CPU, no weights, no torch
python -m fastvideo2 verify wan2.1-t2v-1.3b --tier 0
pytest # the same contracts, as tests
# GPU: bless the component baseline once, then gate against it
python -m fastvideo2 verify wan2.1-t2v-1.3b --tier 1 --bless
python -m fastvideo2 verify wan2.1-t2v-1.3b --tier 3
# generate — CLI or the SDK (the handle is the loaded card, modality-neutral)
python -m fastvideo2 generate wan2.1-t2v-1.3b --prompt "a cat surfing a wave" --out cat.mp4
python -c '
import fastvideo2 as fv2
model = fv2.load("wan2.1-t2v-1.3b") # -> Model (capabilities from the card)
model.generate("a cat surfing a wave", seed=7).save("cat.mp4")'
# the oracle, standalone (this file works copied out of the repo)
python -m fastvideo2.wan21.reference --prompt "a cat surfing a wave" --out ref.mp4
```
We provide a dataset example here. First download testing data. Use [scripts/download_hf.py](scripts/download_hf.py) to download the data to a local directory. Use it like this:
```bash
python scripts/download_hf.py --repo_id=Stealths-Video/Merge-425-Data --local_dir=data/Merge-425-Data --repo_type=dataset
python scripts/download_hf.py --repo_id=Stealths-Video/validation_embeddings --local_dir=data/validation_embeddings --repo_type=dataset
```
Weights resolve from the HF cache (`Wan-AI/Wan2.1-T2V-1.3B-Diffusers`) or an
explicit `--root`; components are stock diffusers/transformers modules, so
there is no conversion step and `load_component()` works standalone in a REPL.
Then the distillation can be launched by:
## Design lineage (what this MVP encodes)
```
bash scripts/distill_t2v.sh
```
- **Cards as declared constants; variants as `derive()` diffs** — no builder
functions, no factory bags, no toy backends welded into production cards.
- **The card digest is the axle artifact**: the same identity a deploy config
points at, a trainer stamps provenance into, and an RL environment manifest
pins (`substitution: exact | bounded | quality-changing` is already on
`Provenance` for the post-training flywheel).
- **Typed conditioning** (`WanForwardInputs`): a new control channel is a new
field the forward must consume — never a silently dropped kwarg.
- **Verification is the product**: gates fail closed, evidence is append-only
data, baselines and tolerances are human-owned.
- **One loop, runtime-visible**: the driven-loop contract is what sessions,
interleaved serving, and RL rollout branching will consume next; the engine
stays a deliberately dumb one-shot runner until those consumers land.
## Scope and non-goals (MVP)
## ⚡ Lora Finetune
## 💰Hardware requirement
- VRAM is required for both distill 10B mochi model
To launch finetuning, you will first need to prepare data in the following formats.
Then the finetuning can be launched by:
```
bash scripts/lora_finetune.sh
```
## Acknowledgement
We learned from and reused code from the following projects: [PCM](https://github.com/G-U-N/Phased-Consistency-Model), [diffusers](https://github.com/huggingface/diffusers), and [OpenSoraPlan](https://github.com/PKU-YuanGroup/Open-Sora-Plan).
In: Wan2.1 T2V, bidirectional, single GPU, one-shot generation, tiers T0–T3.
Out (next, in order): causal/self-forcing students + sessions with forkable
state, the post-training flywheel emitting derived cards + evidence, the RL
environment server (`reset/step/branch`) over the same contracts, additional
model families via `derive()` and new recipe packages.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.3 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.3 MiB

BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 380 KiB

-9
View File
@@ -1,9 +0,0 @@
A hand enters the frame, pulling a sheet of plastic wrap over three balls of dough placed on a wooden surface. The plastic wrap is stretched to cover the dough more securely. The hand adjusts the wrap, ensuring that it is tight and smooth over the dough. The scene focuses on the hand's movements as it secures the edges of the plastic wrap. No new objects appear, and the camera remains stationary, focusing on the action of covering the dough.
A vintage train snakes through the mountains, its plume of white steam rising dramatically against the jagged peaks. The cars glint in the late afternoon sun, their deep crimson and gold accents lending a touch of elegance. The tracks carve a precarious path along the cliffside, revealing glimpses of a roaring river far below. Inside, passengers peer out the large windows, their faces lit with awe as the landscape unfolds.
A crowded rooftop bar buzzes with energy, the city skyline twinkling like a field of stars in the background. Strings of fairy lights hang above, casting a warm, golden glow over the scene. Groups of people gather around high tables, their laughter blending with the soft rhythm of live jazz. The aroma of freshly mixed cocktails and charred appetizers wafts through the air, mingling with the cool night breeze.
En "The Matrix", Neo, interpretado por Keanu Reeves, personifica la lucha contra un sistema opresor a través de su icónica imagen, que incluye unos anteojos oscuros. Estos lentes no son solo un accesorio de moda; representan una barrera entre la realidad y la percepción. Al usar estos anteojos, Neo se sumerge en un mundo donde la verdad se oculta detrás de ilusiones y engaños. La oscuridad de los lentes simboliza la ignorancia y el control que las máquinas tienen sobre la humanidad, mientras que su propia búsqueda de la verdad lo lleva a descubrir sus auténticos poderes. La escena en que se los pone se convierte en un momento crucial, marcando su transformación de un simple programador a "El Elegido". Esta imagen se ha convertido en un ícono cultural, encapsulando el mensaje de que, al enfrentar la oscuridad, podemos encontrar la luz que nos guía hacia la libertad. Así, los anteojos de Neo se convierten en un símbolo de resistencia y autoconocimiento en un mundo manipulado.
Medium close up. Low-angle shot. A woman in a 1950s retro dress sits in a diner bathed in neon light, surrounded by classic decor and lively chatter. The camera starts with a medium shot of her sitting at the counter, then slowly zooms in as she blows a shiny pink bubblegum bubble. The bubble swells dramatically before popping with a soft, playful burst. The scene is vibrant and nostalgic, evoking the fun and carefree spirit of the 1950s.
Will Smith eats noodles.
A short clip of the blonde woman taking a sip from her whiskey glass, her eyes locking with the camera as she smirks playfully. The background shows a group of people laughing and enjoying the party, with vibrant neon signs illuminating the space. The shot is taken in a way that conveys the feeling of a tipsy, carefree night out. The camera then zooms in on her face as she winks, creating a cheeky, flirtatious vibe.
A superintelligent humanoid robot waking up. The robot has a sleek metallic body with futuristic design features. Its glowing red eyes are the focal point, emanating a sharp, intense light as it powers on. The scene is set in a dimly lit, high-tech laboratory filled with glowing control panels, robotic arms, and holographic screens. The setting emphasizes advanced technology and an atmosphere of mystery. The ambiance is eerie and dramatic, highlighting the moment of awakening and the robot's immense intelligence. Photorealistic style with a cinematic, dark sci-fi aesthetic. Aspect ratio: 16:9 --v 6.1
A chimpanzee lead vocalist singing into a microphone on stage. The camera zooms in to show him singing. There is a spotlight on him.
+10
View File
@@ -0,0 +1,10 @@
# Examples
- `generate_t2v.py` — text-to-video via the SDK (`fv2.load(...)` →
`model.generate(...)` → `result.save(...)`). Card defaults, everything
overridable by flag. The standalone reference implementation (no SDK, no
runtime) lives at `fastvideo2/wan21/reference.py`.
The `--model` flag takes any catalog id — e.g. the 3-step FastWan students
(`fastwan-qad-fp8-1.3b`, `fastwan-t2v-1.3b`); their step count, sampler, and
sparsity/quant recipe come from the card, so no other flags change.
+47
View File
@@ -0,0 +1,47 @@
#!/usr/bin/env python3
"""Text-to-video with the fastvideo2 SDK — the canonical example.
python examples/generate_t2v.py --prompt "a cat surfing a wave" --out cat.mp4
Loads the card resident once, generates with card defaults (50 steps, 81
frames, 480x832 — override anything via flags), saves an mp4. Requires a CUDA
box; weights resolve from the HF cache on first use.
"""
from __future__ import annotations
import argparse
import fastvideo2 as fv2
def main() -> None:
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--prompt",
default="a golden retriever puppy running through a sprinkler "
"on a sunny lawn, water droplets sparkling, slow motion, cinematic")
p.add_argument("--model", default="wan2.1-t2v-1.3b",
help="a model id from the catalog (see `python -m fastvideo2 describe`)")
p.add_argument("--out", default="out.mp4")
p.add_argument("--seed", type=int, default=42)
p.add_argument("--num-steps", dest="num_steps", type=int, default=None,
help="unset -> the card's default")
p.add_argument("--num-frames", dest="num_frames", type=int, default=None)
p.add_argument("--guidance-scale", dest="guidance_scale", type=float, default=None)
args = p.parse_args()
model = fv2.load(args.model)
print(model)
overrides = {k: getattr(args, k) for k in ("seed", "num_steps", "num_frames", "guidance_scale")
if getattr(args, k) is not None}
result = model.generate(args.prompt, **overrides)
steps = [t for t in result.trace if "/denoise." in t["label"]]
print(f"video {result.video.shape} | {len(steps)} denoise steps, "
f"{sum(t['seconds'] for t in steps) / max(len(steps), 1):.2f}s/step, "
f"{result.seconds:.1f}s total")
print(f"saved -> {result.save(args.out)}")
if __name__ == "__main__":
main()
-111
View File
@@ -1,111 +0,0 @@
from transformers import AutoTokenizer
from torchvision import transforms
from torchvision.transforms import Lambda
from fastvideo.dataset.t2v_datasets import T2V_dataset
from fastvideo.dataset.latent_datasets import LatentDataset
from fastvideo.dataset.transform import (
Normalize255,
TemporalRandomCrop,
CenterCropResizeVideo,
)
def getdataset(args):
temporal_sample = TemporalRandomCrop(args.num_frames) # 16 x
norm_fun = Lambda(lambda x: 2.0 * x - 1.0)
resize_topcrop = [
CenterCropResizeVideo((args.max_height, args.max_width), top_crop=True),
]
resize = [
CenterCropResizeVideo((args.max_height, args.max_width)),
]
transform = transforms.Compose(
[
# Normalize255(),
*resize,
# RandomHorizontalFlipVideo(p=0.5), # in case their caption have position decription
# norm_fun
]
)
transform_topcrop = transforms.Compose(
[
Normalize255(),
*resize_topcrop,
# RandomHorizontalFlipVideo(p=0.5), # in case their caption have position decription
norm_fun,
]
)
# tokenizer = AutoTokenizer.from_pretrained("/storage/ongoing/new/Open-Sora-Plan/cache_dir/mt5-xxl", cache_dir=args.cache_dir)
tokenizer = AutoTokenizer.from_pretrained(
args.text_encoder_name, cache_dir=args.cache_dir
)
if args.dataset == "t2v":
return T2V_dataset(
args,
transform=transform,
temporal_sample=temporal_sample,
tokenizer=tokenizer,
transform_topcrop=transform_topcrop,
)
raise NotImplementedError(args.dataset)
if __name__ == "__main__":
from accelerate import Accelerator
from fastvideo.dataset.t2v_datasets import dataset_prog
import random
from tqdm import tqdm
args = type(
"args",
(),
{
"ae": "CausalVAEModel_4x8x8",
"dataset": "t2v",
"attention_mode": "xformers",
"use_rope": True,
"text_max_length": 300,
"max_height": 320,
"max_width": 240,
"num_frames": 1,
"use_image_num": 0,
"interpolation_scale_t": 1,
"interpolation_scale_h": 1,
"interpolation_scale_w": 1,
"cache_dir": "../cache_dir",
"image_data": "/storage/ongoing/new/Open-Sora-Plan-bak/7.14bak/scripts/train_data/image_data.txt",
"video_data": "1",
"train_fps": 24,
"drop_short_ratio": 1.0,
"use_img_from_vid": False,
"speed_factor": 1.0,
"cfg": 0.1,
"text_encoder_name": "google/mt5-xxl",
"dataloader_num_workers": 10,
},
)
accelerator = Accelerator()
dataset = getdataset(args)
num = len(dataset_prog.img_cap_list)
zero = 0
for idx in tqdm(range(num)):
image_data = dataset_prog.img_cap_list[idx]
caps = [
i["cap"] if isinstance(i["cap"], list) else [i["cap"]] for i in image_data
]
try:
caps = [[random.choice(i)] for i in caps]
except Exception as e:
print(e)
# import ipdb;ipdb.set_trace()
print(image_data)
zero += 1
continue
assert caps[0] is not None and len(caps[0]) > 0
print(num, zero)
import ipdb
ipdb.set_trace()
print("end")
-126
View File
@@ -1,126 +0,0 @@
import torch
from torch.utils.data import Dataset
import json
import os
import random
class LatentDataset(Dataset):
def __init__(
self,
json_path,
num_latent_t,
cfg_rate,
):
# data_merge_path: video_dir, latent_dir, prompt_embed_dir, json_path
self.json_path = json_path
self.cfg_rate = cfg_rate
self.datase_dir_path = os.path.dirname(json_path)
self.video_dir = os.path.join(self.datase_dir_path, "video")
self.latent_dir = os.path.join(self.datase_dir_path, "latent")
self.prompt_embed_dir = os.path.join(self.datase_dir_path, "prompt_embed")
self.prompt_attention_mask_dir = os.path.join(
self.datase_dir_path, "prompt_attention_mask"
)
with open(self.json_path, "r") as f:
self.data_anno = json.load(f)
# json.load(f) already keeps the order
# self.data_anno = sorted(self.data_anno, key=lambda x: x['latent_path'])
self.num_latent_t = num_latent_t
# just zero embeddings [256, 4096]
self.uncond_prompt_embed = torch.zeros(256, 4096).to(torch.float32)
# 256 zeros
self.uncond_prompt_mask = torch.zeros(256).bool()
self.lengths = [
data_item["length"] if "length" in data_item else 1
for data_item in self.data_anno
]
def __getitem__(self, idx):
latent_file = self.data_anno[idx]["latent_path"]
prompt_embed_file = self.data_anno[idx]["prompt_embed_path"]
prompt_attention_mask_file = self.data_anno[idx]["prompt_attention_mask"]
# load
latent = torch.load(
os.path.join(self.latent_dir, latent_file),
map_location="cpu",
weights_only=True,
)
latent = latent.squeeze(0)[:, -self.num_latent_t :]
if random.random() < self.cfg_rate:
prompt_embed = self.uncond_prompt_embed
prompt_attention_mask = self.uncond_prompt_mask
else:
prompt_embed = torch.load(
os.path.join(self.prompt_embed_dir, prompt_embed_file),
map_location="cpu",
weights_only=True,
)
prompt_attention_mask = torch.load(
os.path.join(
self.prompt_attention_mask_dir, prompt_attention_mask_file
),
map_location="cpu",
weights_only=True,
)
return latent, prompt_embed, prompt_attention_mask
def __len__(self):
return len(self.data_anno)
def latent_collate_function(batch):
# return latent, prompt, latent_attn_mask, text_attn_mask
# latent_attn_mask: # b t h w
# text_attn_mask: b 1 l
# needs to check if the latent/prompt' size and apply padding & attn mask
latents, prompt_embeds, prompt_attention_masks = zip(*batch)
# calculate max shape
max_t = max([latent.shape[1] for latent in latents])
max_h = max([latent.shape[2] for latent in latents])
max_w = max([latent.shape[3] for latent in latents])
# padding
latents = [
torch.nn.functional.pad(
latent,
(
0,
max_t - latent.shape[1],
0,
max_h - latent.shape[2],
0,
max_w - latent.shape[3],
),
)
for latent in latents
]
# attn mask
latent_attn_mask = torch.ones(len(latents), max_t, max_h, max_w)
# set to 0 if padding
for i, latent in enumerate(latents):
latent_attn_mask[i, latent.shape[1] :, :, :] = 0
latent_attn_mask[i, :, latent.shape[2] :, :] = 0
latent_attn_mask[i, :, :, latent.shape[3] :] = 0
prompt_embeds = torch.stack(prompt_embeds, dim=0)
prompt_attention_masks = torch.stack(prompt_attention_masks, dim=0)
latents = torch.stack(latents, dim=0)
return latents, prompt_embeds, latent_attn_mask, prompt_attention_masks
if __name__ == "__main__":
dataset = LatentDataset("data/Mochi-Synthetic-Data/merge.txt", num_latent_t=28)
dataloader = torch.utils.data.DataLoader(
dataset, batch_size=2, shuffle=False, collate_fn=latent_collate_function
)
for latent, prompt_embed, latent_attn_mask, prompt_attention_mask in dataloader:
print(
latent.shape,
prompt_embed.shape,
latent_attn_mask.shape,
prompt_attention_mask.shape,
)
import pdb
pdb.set_trace()
-357
View File
@@ -1,357 +0,0 @@
import json
import os, io, csv, math, random
import numpy as np
from einops import rearrange
from decord import VideoReader
from os.path import join as opj
from collections import Counter
import torch
from torch.utils.data.dataset import Dataset
from torch.utils.data import DataLoader, Dataset, get_worker_info
from tqdm import tqdm
from PIL import Image
from accelerate.logging import get_logger
from fastvideo.utils.dataset_utils import DecordInit
import torchvision
logger = get_logger(__name__)
class SingletonMeta(type):
_instances = {}
def __call__(cls, *args, **kwargs):
if cls not in cls._instances:
instance = super().__call__(*args, **kwargs)
cls._instances[cls] = instance
return cls._instances[cls]
class DataSetProg(metaclass=SingletonMeta):
def __init__(self):
self.cap_list = []
self.elements = []
self.num_workers = 1
self.n_elements = 0
self.worker_elements = dict()
self.n_used_elements = dict()
def set_cap_list(self, num_workers, cap_list, n_elements):
self.num_workers = num_workers
self.cap_list = cap_list
self.n_elements = n_elements
self.elements = list(range(n_elements))
random.shuffle(self.elements)
print(f"n_elements: {len(self.elements)}", flush=True)
for i in range(self.num_workers):
self.n_used_elements[i] = 0
per_worker = int(math.ceil(len(self.elements) / float(self.num_workers)))
start = i * per_worker
end = min(start + per_worker, len(self.elements))
self.worker_elements[i] = self.elements[start:end]
def get_item(self, work_info):
if work_info is None:
worker_id = 0
else:
worker_id = work_info.id
idx = self.worker_elements[worker_id][
self.n_used_elements[worker_id] % len(self.worker_elements[worker_id])
]
self.n_used_elements[worker_id] += 1
return idx
dataset_prog = DataSetProg()
def filter_resolution(h, w, max_h_div_w_ratio=17 / 16, min_h_div_w_ratio=8 / 16):
if h / w <= max_h_div_w_ratio and h / w >= min_h_div_w_ratio:
return True
return False
class T2V_dataset(Dataset):
def __init__(self, args, transform, temporal_sample, tokenizer, transform_topcrop):
self.data = args.data_merge_path
self.num_frames = args.num_frames
self.train_fps = args.train_fps
self.use_image_num = args.use_image_num
self.transform = transform
self.transform_topcrop = transform_topcrop
self.temporal_sample = temporal_sample
self.tokenizer = tokenizer
self.text_max_length = args.text_max_length
self.cfg = args.cfg
self.speed_factor = args.speed_factor
self.max_height = args.max_height
self.max_width = args.max_width
self.drop_short_ratio = args.drop_short_ratio
assert self.speed_factor >= 1
self.v_decoder = DecordInit()
self.video_length_tolerance_range = args.video_length_tolerance_range
self.support_Chinese = True
if not ("mt5" in args.text_encoder_name):
self.support_Chinese = False
cap_list = self.get_cap_list()
assert len(cap_list) > 0
cap_list, self.sample_num_frames = self.define_frame_index(cap_list)
self.lengths = self.sample_num_frames
n_elements = len(cap_list)
dataset_prog.set_cap_list(args.dataloader_num_workers, cap_list, n_elements)
print(f"video length: {len(dataset_prog.cap_list)}", flush=True)
def set_checkpoint(self, n_used_elements):
for i in range(len(dataset_prog.n_used_elements)):
dataset_prog.n_used_elements[i] = n_used_elements
def __len__(self):
return dataset_prog.n_elements
def __getitem__(self, idx):
try:
data = self.get_data(idx)
return data
except Exception as e:
logger.info(f"Error with {e}")
if idx in dataset_prog.cap_list:
logger.info(f"Caught an exception! {dataset_prog.cap_list[idx]}")
return self.__getitem__(random.randint(0, self.__len__() - 1))
def get_data(self, idx):
path = dataset_prog.cap_list[idx]["path"]
if path.endswith(".mp4"):
return self.get_video(idx)
else:
return self.get_image(idx)
def get_video(self, idx):
video_path = dataset_prog.cap_list[idx]["path"]
assert os.path.exists(video_path), f"file {video_path} do not exist!"
frame_indices = dataset_prog.cap_list[idx]["sample_frame_index"]
torchvision_video, _, metadata = torchvision.io.read_video(
video_path, output_format="TCHW"
)
video = torchvision_video[frame_indices]
video = self.transform(video)
video = rearrange(video, "t c h w -> c t h w")
video = video.to(torch.uint8)
assert video.dtype == torch.uint8
h, w = video.shape[-2:]
assert (
h / w <= 17 / 16 and h / w >= 8 / 16
), f"Only videos with a ratio (h/w) less than 17/16 and more than 8/16 are supported. But video ({video_path}) found ratio is {round(h / w, 2)} with the shape of {video.shape}"
video = video.float() / 127.5 - 1.0
text = dataset_prog.cap_list[idx]["cap"]
if not isinstance(text, list):
text = [text]
text = [random.choice(text)]
text = text[0] if random.random() > self.cfg else ""
text_tokens_and_mask = self.tokenizer(
text,
max_length=self.text_max_length,
padding="max_length",
truncation=True,
return_attention_mask=True,
add_special_tokens=True,
return_tensors="pt",
)
input_ids = text_tokens_and_mask["input_ids"]
cond_mask = text_tokens_and_mask["attention_mask"]
return dict(
pixel_values=video,
text=text,
input_ids=input_ids,
cond_mask=cond_mask,
path=video_path,
)
def get_image(self, idx):
image_data = dataset_prog.cap_list[idx] # [{'path': path, 'cap': cap}, ...]
image = Image.open(image_data["path"]).convert("RGB") # [h, w, c]
image = torch.from_numpy(np.array(image)) # [h, w, c]
image = rearrange(image, "h w c -> c h w").unsqueeze(0) # [1 c h w]
# for i in image:
# h, w = i.shape[-2:]
# assert h / w <= 17 / 16 and h / w >= 8 / 16, f'Only image with a ratio (h/w) less than 17/16 and more than 8/16 are supported. But found ratio is {round(h / w, 2)} with the shape of {i.shape}'
image = (
self.transform_topcrop(image)
if "human_images" in image_data["path"]
else self.transform(image)
) # [1 C H W] -> num_img [1 C H W]
image = image.transpose(0, 1) # [1 C H W] -> [C 1 H W]
image = image.float() / 127.5 - 1.0
caps = (
image_data["cap"]
if isinstance(image_data["cap"], list)
else [image_data["cap"]]
)
caps = [random.choice(caps)]
text = caps
input_ids, cond_mask = [], []
text = text if random.random() > self.cfg else ""
text_tokens_and_mask = self.tokenizer(
text,
max_length=self.text_max_length,
padding="max_length",
truncation=True,
return_attention_mask=True,
add_special_tokens=True,
return_tensors="pt",
)
input_ids = text_tokens_and_mask["input_ids"] # 1, l
cond_mask = text_tokens_and_mask["attention_mask"] # 1, l
return dict(
pixel_values=image,
text=text,
input_ids=input_ids,
cond_mask=cond_mask,
path=image_data["path"],
)
def define_frame_index(self, cap_list):
new_cap_list = []
sample_num_frames = []
cnt_too_long = 0
cnt_too_short = 0
cnt_no_cap = 0
cnt_no_resolution = 0
cnt_resolution_mismatch = 0
cnt_movie = 0
cnt_img = 0
for i in cap_list:
path = i["path"]
cap = i.get("cap", None)
# ======no caption=====
if cap is None:
cnt_no_cap += 1
continue
if path.endswith(".mp4"):
# ======no fps and duration=====
duration = i.get("duration", None)
fps = i.get("fps", None)
if fps is None or duration is None:
continue
# ======resolution mismatch=====
resolution = i.get("resolution", None)
if resolution is None:
cnt_no_resolution += 1
continue
else:
if (
resolution.get("height", None) is None
or resolution.get("width", None) is None
):
cnt_no_resolution += 1
continue
height, width = i["resolution"]["height"], i["resolution"]["width"]
aspect = self.max_height / self.max_width
hw_aspect_thr = 1.5
is_pick = filter_resolution(
height,
width,
max_h_div_w_ratio=hw_aspect_thr * aspect,
min_h_div_w_ratio=1 / hw_aspect_thr * aspect,
)
if not is_pick:
print("resolution mismatch")
cnt_resolution_mismatch += 1
continue
# import ipdb;ipdb.set_trace()
i["num_frames"] = math.ceil(fps * duration)
# max 5.0 and min 1.0 are just thresholds to filter some videos which have suitable duration.
if (
i["num_frames"] / fps
> self.video_length_tolerance_range
* (self.num_frames / self.train_fps * self.speed_factor)
): # too long video is not suitable for this training stage (self.num_frames)
cnt_too_long += 1
continue
# resample in case high fps, such as 50/60/90/144 -> train_fps(e.g, 24)
frame_interval = fps / self.train_fps
start_frame_idx = 0
frame_indices = np.arange(
start_frame_idx, i["num_frames"], frame_interval
).astype(int)
# comment out it to enable dynamic frames training
if (
len(frame_indices) < self.num_frames
and random.random() < self.drop_short_ratio
):
cnt_too_short += 1
continue
# too long video will be temporal-crop randomly
if len(frame_indices) > self.num_frames:
begin_index, end_index = self.temporal_sample(len(frame_indices))
frame_indices = frame_indices[begin_index:end_index]
# frame_indices = frame_indices[:self.num_frames] # head crop
i["sample_frame_index"] = frame_indices.tolist()
new_cap_list.append(i)
i["sample_num_frames"] = len(
i["sample_frame_index"]
) # will use in dataloader(group sampler)
sample_num_frames.append(i["sample_num_frames"])
elif path.endswith(".jpg"): # image
cnt_img += 1
new_cap_list.append(i)
i["sample_num_frames"] = 1
sample_num_frames.append(i["sample_num_frames"])
else:
raise NameError(
f"Unknown file extention {path.split('.')[-1]}, only support .mp4 for video and .jpg for image"
)
# import ipdb;ipdb.set_trace()
logger.info(
f"no_cap: {cnt_no_cap}, too_long: {cnt_too_long}, too_short: {cnt_too_short}, "
f"no_resolution: {cnt_no_resolution}, resolution_mismatch: {cnt_resolution_mismatch}, "
f"Counter(sample_num_frames): {Counter(sample_num_frames)}, cnt_movie: {cnt_movie}, cnt_img: {cnt_img}, "
f"before filter: {len(cap_list)}, after filter: {len(new_cap_list)}"
)
return new_cap_list, sample_num_frames
def decord_read(self, path, frame_indices):
decord_vr = self.v_decoder(path)
video_data = decord_vr.get_batch(frame_indices).asnumpy()
video_data = torch.from_numpy(video_data)
video_data = video_data.permute(0, 3, 1, 2) # (T, H, W, C) -> (T C H W)
return video_data
def read_jsons(self, data):
cap_lists = []
with open(data, "r") as f:
folder_anno = [
i.strip().split(",") for i in f.readlines() if len(i.strip()) > 0
]
print(folder_anno)
for folder, anno in folder_anno:
with open(anno, "r") as f:
sub_list = json.load(f)
logger.info(f"Building {anno}...")
for i in range(len(sub_list)):
sub_list[i]["path"] = opj(folder, sub_list[i]["path"])
cap_lists += sub_list
return cap_lists
def get_cap_list(self):
cap_lists = self.read_jsons(self.data)
return cap_lists
-645
View File
@@ -1,645 +0,0 @@
import torch
import random
import numbers
from torchvision.transforms import RandomCrop, RandomResizedCrop
def _is_tensor_video_clip(clip):
if not torch.is_tensor(clip):
raise TypeError("clip should be Tensor. Got %s" % type(clip))
if not clip.ndimension() == 4:
raise ValueError("clip should be 4D. Got %dD" % clip.dim())
return True
def center_crop_arr(pil_image, image_size):
"""
Center cropping implementation from ADM.
https://github.com/openai/guided-diffusion/blob/8fb3ad9197f16bbc40620447b2742e13458d2831/guided_diffusion/image_datasets.py#L126
"""
while min(*pil_image.size) >= 2 * image_size:
pil_image = pil_image.resize(
tuple(x // 2 for x in pil_image.size), resample=Image.BOX
)
scale = image_size / min(*pil_image.size)
pil_image = pil_image.resize(
tuple(round(x * scale) for x in pil_image.size), resample=Image.BICUBIC
)
arr = np.array(pil_image)
crop_y = (arr.shape[0] - image_size) // 2
crop_x = (arr.shape[1] - image_size) // 2
return Image.fromarray(
arr[crop_y : crop_y + image_size, crop_x : crop_x + image_size]
)
def crop(clip, i, j, h, w):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
"""
if len(clip.size()) != 4:
raise ValueError("clip should be a 4D tensor")
return clip[..., i : i + h, j : j + w]
def resize(clip, target_size, interpolation_mode):
if len(target_size) != 2:
raise ValueError(
f"target size should be tuple (height, width), instead got {target_size}"
)
return torch.nn.functional.interpolate(
clip,
size=target_size,
mode=interpolation_mode,
align_corners=True,
antialias=True,
)
def resize_scale(clip, target_size, interpolation_mode):
if len(target_size) != 2:
raise ValueError(
f"target size should be tuple (height, width), instead got {target_size}"
)
H, W = clip.size(-2), clip.size(-1)
scale_ = target_size[0] / min(H, W)
return torch.nn.functional.interpolate(
clip,
scale_factor=scale_,
mode=interpolation_mode,
align_corners=True,
antialias=True,
)
def resized_crop(clip, i, j, h, w, size, interpolation_mode="bilinear"):
"""
Do spatial cropping and resizing to the video clip
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
i (int): i in (i,j) i.e coordinates of the upper left corner.
j (int): j in (i,j) i.e coordinates of the upper left corner.
h (int): Height of the cropped region.
w (int): Width of the cropped region.
size (tuple(int, int)): height and width of resized clip
Returns:
clip (torch.tensor): Resized and cropped clip. Size is (T, C, H, W)
"""
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
clip = crop(clip, i, j, h, w)
clip = resize(clip, size, interpolation_mode)
return clip
def center_crop(clip, crop_size):
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
h, w = clip.size(-2), clip.size(-1)
th, tw = crop_size
if h < th or w < tw:
raise ValueError("height and width must be no smaller than crop_size")
i = int(round((h - th) / 2.0))
j = int(round((w - tw) / 2.0))
return crop(clip, i, j, th, tw)
def center_crop_using_short_edge(clip):
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
h, w = clip.size(-2), clip.size(-1)
if h < w:
th, tw = h, h
i = 0
j = int(round((w - tw) / 2.0))
else:
th, tw = w, w
i = int(round((h - th) / 2.0))
j = 0
return crop(clip, i, j, th, tw)
def center_crop_th_tw(clip, th, tw, top_crop):
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
# import ipdb;ipdb.set_trace()
h, w = clip.size(-2), clip.size(-1)
tr = th / tw
if h / w > tr:
new_h = int(w * tr)
new_w = w
else:
new_h = h
new_w = int(h / tr)
i = 0 if top_crop else int(round((h - new_h) / 2.0))
j = int(round((w - new_w) / 2.0))
return crop(clip, i, j, new_h, new_w)
def random_shift_crop(clip):
"""
Slide along the long edge, with the short edge as crop size
"""
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
h, w = clip.size(-2), clip.size(-1)
if h <= w:
long_edge = w
short_edge = h
else:
long_edge = h
short_edge = w
th, tw = short_edge, short_edge
i = torch.randint(0, h - th + 1, size=(1,)).item()
j = torch.randint(0, w - tw + 1, size=(1,)).item()
return crop(clip, i, j, th, tw)
def normalize_video(clip):
"""
Convert tensor data type from uint8 to float, divide value by 255.0 and
permute the dimensions of clip tensor
Args:
clip (torch.tensor, dtype=torch.uint8): Size is (T, C, H, W)
Return:
clip (torch.tensor, dtype=torch.float): Size is (T, C, H, W)
"""
_is_tensor_video_clip(clip)
if not clip.dtype == torch.uint8:
raise TypeError(
"clip tensor should have data type uint8. Got %s" % str(clip.dtype)
)
# return clip.float().permute(3, 0, 1, 2) / 255.0
return clip.float() / 255.0
def normalize(clip, mean, std, inplace=False):
"""
Args:
clip (torch.tensor): Video clip to be normalized. Size is (T, C, H, W)
mean (tuple): pixel RGB mean. Size is (3)
std (tuple): pixel standard deviation. Size is (3)
Returns:
normalized clip (torch.tensor): Size is (T, C, H, W)
"""
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
if not inplace:
clip = clip.clone()
mean = torch.as_tensor(mean, dtype=clip.dtype, device=clip.device)
# print(mean)
std = torch.as_tensor(std, dtype=clip.dtype, device=clip.device)
clip.sub_(mean[:, None, None, None]).div_(std[:, None, None, None])
return clip
def hflip(clip):
"""
Args:
clip (torch.tensor): Video clip to be normalized. Size is (T, C, H, W)
Returns:
flipped clip (torch.tensor): Size is (T, C, H, W)
"""
if not _is_tensor_video_clip(clip):
raise ValueError("clip should be a 4D torch.tensor")
return clip.flip(-1)
class RandomCropVideo:
def __init__(self, size):
if isinstance(size, numbers.Number):
self.size = (int(size), int(size))
else:
self.size = size
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: randomly cropped video clip.
size is (T, C, OH, OW)
"""
i, j, h, w = self.get_params(clip)
return crop(clip, i, j, h, w)
def get_params(self, clip):
h, w = clip.shape[-2:]
th, tw = self.size
if h < th or w < tw:
raise ValueError(
f"Required crop size {(th, tw)} is larger than input image size {(h, w)}"
)
if w == tw and h == th:
return 0, 0, h, w
i = torch.randint(0, h - th + 1, size=(1,)).item()
j = torch.randint(0, w - tw + 1, size=(1,)).item()
return i, j, th, tw
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size})"
class SpatialStrideCropVideo:
def __init__(self, stride):
self.stride = stride
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: cropped video clip by stride.
size is (T, C, OH, OW)
"""
i, j, h, w = self.get_params(clip)
return crop(clip, i, j, h, w)
def get_params(self, clip):
h, w = clip.shape[-2:]
th, tw = h // self.stride * self.stride, w // self.stride * self.stride
return 0, 0, th, tw # from top-left
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size})"
class LongSideResizeVideo:
"""
First use the long side,
then resize to the specified size
"""
def __init__(
self,
size,
skip_low_resolution=False,
interpolation_mode="bilinear",
):
self.size = size
self.skip_low_resolution = skip_low_resolution
self.interpolation_mode = interpolation_mode
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: scale resized video clip.
size is (T, C, 512, *) or (T, C, *, 512)
"""
_, _, h, w = clip.shape
if self.skip_low_resolution and max(h, w) <= self.size:
return clip
if h > w:
w = int(w * self.size / h)
h = self.size
else:
h = int(h * self.size / w)
w = self.size
resize_clip = resize(
clip, target_size=(h, w), interpolation_mode=self.interpolation_mode
)
return resize_clip
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size}, interpolation_mode={self.interpolation_mode}"
class CenterCropResizeVideo:
"""
First use the short side for cropping length,
center crop video, then resize to the specified size
"""
def __init__(
self,
size,
top_crop=False,
interpolation_mode="bilinear",
):
if len(size) != 2:
raise ValueError(
f"size should be tuple (height, width), instead got {size}"
)
self.size = size
self.top_crop = top_crop
self.interpolation_mode = interpolation_mode
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: scale resized / center cropped video clip.
size is (T, C, crop_size, crop_size)
"""
# clip_center_crop = center_crop_using_short_edge(clip)
clip_center_crop = center_crop_th_tw(
clip, self.size[0], self.size[1], top_crop=self.top_crop
)
# import ipdb;ipdb.set_trace()
clip_center_crop_resize = resize(
clip_center_crop,
target_size=self.size,
interpolation_mode=self.interpolation_mode,
)
return clip_center_crop_resize
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size}, interpolation_mode={self.interpolation_mode}"
class UCFCenterCropVideo:
"""
First scale to the specified size in equal proportion to the short edge,
then center cropping
"""
def __init__(
self,
size,
interpolation_mode="bilinear",
):
if isinstance(size, tuple):
if len(size) != 2:
raise ValueError(
f"size should be tuple (height, width), instead got {size}"
)
self.size = size
else:
self.size = (size, size)
self.interpolation_mode = interpolation_mode
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: scale resized / center cropped video clip.
size is (T, C, crop_size, crop_size)
"""
clip_resize = resize_scale(
clip=clip, target_size=self.size, interpolation_mode=self.interpolation_mode
)
clip_center_crop = center_crop(clip_resize, self.size)
return clip_center_crop
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size}, interpolation_mode={self.interpolation_mode}"
class KineticsRandomCropResizeVideo:
"""
Slide along the long edge, with the short edge as crop size. And resie to the desired size.
"""
def __init__(
self,
size,
interpolation_mode="bilinear",
):
if isinstance(size, tuple):
if len(size) != 2:
raise ValueError(
f"size should be tuple (height, width), instead got {size}"
)
self.size = size
else:
self.size = (size, size)
self.interpolation_mode = interpolation_mode
def __call__(self, clip):
clip_random_crop = random_shift_crop(clip)
clip_resize = resize(clip_random_crop, self.size, self.interpolation_mode)
return clip_resize
class CenterCropVideo:
def __init__(
self,
size,
interpolation_mode="bilinear",
):
if isinstance(size, tuple):
if len(size) != 2:
raise ValueError(
f"size should be tuple (height, width), instead got {size}"
)
self.size = size
else:
self.size = (size, size)
self.interpolation_mode = interpolation_mode
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Video clip to be cropped. Size is (T, C, H, W)
Returns:
torch.tensor: center cropped video clip.
size is (T, C, crop_size, crop_size)
"""
clip_center_crop = center_crop(clip, self.size)
return clip_center_crop
def __repr__(self) -> str:
return f"{self.__class__.__name__}(size={self.size}, interpolation_mode={self.interpolation_mode}"
class Normalize:
"""
Normalize the video clip by mean subtraction and division by standard deviation
Args:
mean (3-tuple): pixel RGB mean
std (3-tuple): pixel RGB standard deviation
inplace (boolean): whether do in-place normalization
"""
def __init__(self, mean, std, inplace=False):
self.mean = mean
self.std = std
self.inplace = inplace
def __call__(self, clip):
"""
Args:
clip (torch.tensor): video clip must be normalized. Size is (C, T, H, W)
"""
return normalize(clip, self.mean, self.std, self.inplace)
def __repr__(self) -> str:
return f"{self.__class__.__name__}(mean={self.mean}, std={self.std}, inplace={self.inplace})"
class Normalize255:
"""
Convert tensor data type from uint8 to float, divide value by 255.0 and
"""
def __init__(self):
pass
def __call__(self, clip):
"""
Args:
clip (torch.tensor, dtype=torch.uint8): Size is (T, C, H, W)
Return:
clip (torch.tensor, dtype=torch.float): Size is (T, C, H, W)
"""
return normalize_video(clip)
def __repr__(self) -> str:
return self.__class__.__name__
class RandomHorizontalFlipVideo:
"""
Flip the video clip along the horizontal direction with a given probability
Args:
p (float): probability of the clip being flipped. Default value is 0.5
"""
def __init__(self, p=0.5):
self.p = p
def __call__(self, clip):
"""
Args:
clip (torch.tensor): Size is (T, C, H, W)
Return:
clip (torch.tensor): Size is (T, C, H, W)
"""
if random.random() < self.p:
clip = hflip(clip)
return clip
def __repr__(self) -> str:
return f"{self.__class__.__name__}(p={self.p})"
# ------------------------------------------------------------
# --------------------- Sampling ---------------------------
# ------------------------------------------------------------
class TemporalRandomCrop(object):
"""Temporally crop the given frame indices at a random location.
Args:
size (int): Desired length of frames will be seen in the model.
"""
def __init__(self, size):
self.size = size
def __call__(self, total_frames):
rand_end = max(0, total_frames - self.size - 1)
begin_index = random.randint(0, rand_end)
end_index = min(begin_index + self.size, total_frames)
return begin_index, end_index
class DynamicSampleDuration(object):
"""Temporally crop the given frame indices at a random location.
Args:
size (int): Desired length of frames will be seen in the model.
"""
def __init__(self, t_stride, extra_1):
self.t_stride = t_stride
self.extra_1 = extra_1
def __call__(self, t, h, w):
if self.extra_1:
t = t - 1
truncate_t_list = list(range(t + 1))[t // 2 :][
:: self.t_stride
] # need half at least
truncate_t = random.choice(truncate_t_list)
if self.extra_1:
truncate_t = truncate_t + 1
return 0, truncate_t
if __name__ == "__main__":
from torchvision import transforms
import torchvision.io as io
import numpy as np
from torchvision.utils import save_image
import os
vframes, aframes, info = io.read_video(
filename="./v_Archery_g01_c03.avi", pts_unit="sec", output_format="TCHW"
)
trans = transforms.Compose(
[
Normalize255(),
RandomHorizontalFlipVideo(),
UCFCenterCropVideo(512),
# NormalizeVideo(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True),
transforms.Normalize(
mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True
),
]
)
target_video_len = 32
frame_interval = 1
total_frames = len(vframes)
print(total_frames)
temporal_sample = TemporalRandomCrop(target_video_len * frame_interval)
# Sampling video frames
start_frame_ind, end_frame_ind = temporal_sample(total_frames)
# print(start_frame_ind)
# print(end_frame_ind)
assert end_frame_ind - start_frame_ind >= target_video_len
frame_indice = np.linspace(
start_frame_ind, end_frame_ind - 1, target_video_len, dtype=int
)
print(frame_indice)
select_vframes = vframes[frame_indice]
print(select_vframes.shape)
print(select_vframes.dtype)
select_vframes_trans = trans(select_vframes)
print(select_vframes_trans.shape)
print(select_vframes_trans.dtype)
select_vframes_trans_int = ((select_vframes_trans * 0.5 + 0.5) * 255).to(
dtype=torch.uint8
)
print(select_vframes_trans_int.dtype)
print(select_vframes_trans_int.permute(0, 2, 3, 1).shape)
io.write_video("./test.avi", select_vframes_trans_int.permute(0, 2, 3, 1), fps=8)
for i in range(target_video_len):
save_image(
select_vframes_trans[i],
os.path.join("./test000", "%04d.png" % i),
normalize=True,
value_range=(-1, 1),
)
-142
View File
@@ -1,142 +0,0 @@
import gradio as gr
import torch
from fastvideo.model.pipeline_mochi import MochiPipeline
from fastvideo.model.modeling_mochi import MochiTransformer3DModel
from diffusers import FlowMatchEulerDiscreteScheduler
from diffusers.utils import export_to_video
from fastvideo.distill.solver import PCMFMScheduler
import tempfile
import os
import argparse
def init_args():
parser = argparse.ArgumentParser()
parser.add_argument("--prompts", nargs='+', default=[])
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument("--height", type=int, default=480)
parser.add_argument("--width", type=int, default=848)
parser.add_argument("--num_inference_steps", type=int, default=64)
parser.add_argument("--guidance_scale", type=float, default=4.5)
parser.add_argument("--model_path", type=str, default="data/mochi")
parser.add_argument("--seed", type=int, default=42)
parser.add_argument("--transformer_path", type=str, default=None)
parser.add_argument("--scheduler_type", type=str, default="euler")
parser.add_argument("--lora_checkpoint_dir", type=str, default=None)
parser.add_argument("--shift", type=float, default=8.0)
parser.add_argument("--num_euler_timesteps", type=int, default=100)
parser.add_argument("--linear_threshold", type=float, default=0.025)
parser.add_argument("--linear_range", type=float, default=0.5)
parser.add_argument("--cpu_offload", action="store_true")
return parser.parse_args()
def load_model(args):
device = "cuda" if torch.cuda.is_available() else "cpu"
if args.scheduler_type == "euler":
scheduler = FlowMatchEulerDiscreteScheduler()
else:
scheduler = PCMFMScheduler(1000, args.shift, args.num_euler_timesteps, False, args.linear_threshold, args.linear_range)
if args.transformer_path:
transformer = MochiTransformer3DModel.from_pretrained(args.transformer_path)
else:
transformer = MochiTransformer3DModel.from_pretrained(args.model_path, subfolder='transformer/')
pipe = MochiPipeline.from_pretrained(args.model_path, transformer=transformer, scheduler=scheduler)
pipe.enable_vae_tiling()
pipe.to(device)
if args.cpu_offload:
pipe.enable_model_cpu_offload()
return pipe
def generate_video(prompt, negative_prompt, use_negative_prompt, seed, guidance_scale,
num_frames, height, width, num_inference_steps, randomize_seed=False):
if randomize_seed:
seed = torch.randint(0, 1000000, (1,)).item()
pipe = load_model(args)
print("load model successfully")
generator = torch.Generator(device="cuda").manual_seed(seed)
if not use_negative_prompt:
negative_prompt = None
with torch.autocast("cuda", dtype=torch.bfloat16):
output = pipe(
prompt=[prompt],
negative_prompt=negative_prompt,
height=height,
width=width,
num_frames=num_frames,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
generator=generator,
).frames[0]
output_path = os.path.join(tempfile.mkdtemp(), "output.mp4")
export_to_video(output, output_path, fps=30)
return output_path, seed
examples = [
"A hand enters the frame, pulling a sheet of plastic wrap over three balls of dough placed on a wooden surface. The plastic wrap is stretched to cover the dough more securely. The hand adjusts the wrap, ensuring that it is tight and smooth over the dough. The scene focuses on the hand’s movements as it secures the edges of the plastic wrap. No new objects appear, and the camera remains stationary, focusing on the action of covering the dough.",
"A vintage train snakes through the mountains, its plume of white steam rising dramatically against the jagged peaks. The cars glint in the late afternoon sun, their deep crimson and gold accents lending a touch of elegance. The tracks carve a precarious path along the cliffside, revealing glimpses of a roaring river far below. Inside, passengers peer out the large windows, their faces lit with awe as the landscape unfolds.",
"A crowded rooftop bar buzzes with energy, the city skyline twinkling like a field of stars in the background. Strings of fairy lights hang above, casting a warm, golden glow over the scene. Groups of people gather around high tables, their laughter blending with the soft rhythm of live jazz. The aroma of freshly mixed cocktails and charred appetizers wafts through the air, mingling with the cool night breeze."
]
args = init_args()
with gr.Blocks() as demo:
gr.Markdown("# Mochi Video Generation Demo")
with gr.Group():
with gr.Row():
prompt = gr.Text(
label="Prompt",
show_label=False,
max_lines=1,
placeholder="Enter your prompt",
container=False,
)
run_button = gr.Button("Run", scale=0)
result = gr.Video(label="Result", show_label=False)
with gr.Accordion("Advanced options", open=False):
with gr.Group():
with gr.Row():
height = gr.Slider(label="Height", minimum=256, maximum=1024, step=32, value=args.height)
width = gr.Slider(label="Width", minimum=256, maximum=1024, step=32, value=args.width)
with gr.Row():
num_frames = gr.Slider(label="Number of Frames", minimum=8, maximum=256, value=args.num_frames)
guidance_scale = gr.Slider(label="Guidance Scale", minimum=1, maximum=20, value=args.guidance_scale)
num_inference_steps = gr.Slider(label="Inference Steps", minimum=10, maximum=100, value=args.num_inference_steps)
with gr.Row():
use_negative_prompt = gr.Checkbox(label="Use negative prompt", value=False)
negative_prompt = gr.Text(
label="Negative prompt",
max_lines=1,
placeholder="Enter a negative prompt",
visible=False
)
seed = gr.Slider(label="Seed", minimum=0, maximum=1000000, step=1, value=args.seed)
randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
seed_output = gr.Number(label="Used Seed")
gr.Examples(examples=examples, inputs=prompt)
use_negative_prompt.change(
fn=lambda x: gr.update(visible=x),
inputs=use_negative_prompt,
outputs=negative_prompt,
)
run_button.click(
fn=generate_video,
inputs=[prompt, negative_prompt, use_negative_prompt, seed, guidance_scale,
num_frames, height, width, num_inference_steps, randomize_seed],
outputs=[result, seed_output]
)
if __name__ == "__main__":
demo.queue(max_size=20).launch(server_name="0.0.0.0", server_port=7860)
-901
View File
@@ -1,901 +0,0 @@
import argparse
import math
import os
from fastvideo.utils.parallel_states import (
initialize_sequence_parallel_state,
destroy_sequence_parallel_group,
get_sequence_parallel_state,
nccl_info,
)
from fastvideo.utils.communications import sp_parallel_dataloader_wrapper, broadcast
from fastvideo.models.mochi_hf.mochi_latents_utils import normalize_mochi_dit_input
from fastvideo.utils.validation import log_validation
import time
from torch.utils.data import DataLoader
import torch
from torch.distributed.fsdp import ShardingStrategy
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
StateDictType,
FullStateDictConfig,
)
from fastvideo.models.mochi_hf.pipeline_mochi import linear_quadratic_schedule
import json
from torch.utils.data.distributed import DistributedSampler
from fastvideo.utils.dataset_utils import LengthGroupedSampler
import wandb
from accelerate.utils import set_seed
from tqdm.auto import tqdm
from fastvideo.fsdp_util import get_dit_fsdp_kwargs, apply_fsdp_checkpointing
from diffusers import (
FlowMatchEulerDiscreteScheduler,
)
from fastvideo.distill.solver import EulerSolver, extract_into_tensor
from copy import deepcopy
from diffusers.optimization import get_scheduler
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
from diffusers.utils import check_min_version
from fastvideo.dataset.latent_datasets import LatentDataset, latent_collate_function
import torch.distributed as dist
from safetensors.torch import save_file
from peft import LoraConfig
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
)
from fastvideo.utils.checkpoint import (
save_checkpoint,
save_lora_checkpoint,
resume_lora_optimizer,
)
# Will error if the minimal version of diffusers is not installed. Remove at your own risks.
check_min_version("0.31.0")
import time
from collections import deque
def main_print(content):
if int(os.environ["LOCAL_RANK"]) <= 0:
print(content)
def save_checkpoint(transformer: MochiTransformer3DModel, rank, output_dir, step):
main_print(f"--> saving checkpoint at step {step}")
with FSDP.state_dict_type(
transformer,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
cpu_state = transformer.state_dict()
# todo move to get_state_dict
if rank <= 0:
save_dir = os.path.join(output_dir, f"checkpoint-{step}")
os.makedirs(save_dir, exist_ok=True)
# save using safetensors
weight_path = os.path.join(save_dir, "diffusion_pytorch_model.safetensors")
save_file(cpu_state, weight_path)
config_dict = dict(transformer.config)
config_path = os.path.join(save_dir, "config.json")
# save dict as json
with open(config_path, "w") as f:
json.dump(config_dict, f, indent=4)
main_print(f"--> checkpoint saved at step {step}")
def reshard_fsdp(model):
for m in FSDP.fsdp_modules(model):
if m._has_params and m.sharding_strategy is not ShardingStrategy.NO_SHARD:
torch.distributed.fsdp._runtime_utils._reshard(m, m._handle, True)
def get_norm(model_pred, norms, gradient_accumulation_steps):
fro_norm = (
torch.linalg.matrix_norm(model_pred, ord="fro") / gradient_accumulation_steps
)
largest_singular_value = (
torch.linalg.matrix_norm(model_pred, ord=2) / gradient_accumulation_steps
)
absolute_mean = torch.mean(torch.abs(model_pred)) / gradient_accumulation_steps
absolute_max = torch.max(torch.abs(model_pred)) / gradient_accumulation_steps
dist.all_reduce(fro_norm, op=dist.ReduceOp.AVG)
dist.all_reduce(largest_singular_value, op=dist.ReduceOp.AVG)
dist.all_reduce(absolute_mean, op=dist.ReduceOp.AVG)
norms["fro"] += torch.mean(fro_norm).item()
norms["largest singular value"] += torch.mean(largest_singular_value).item()
norms["absolute mean"] += absolute_mean.item()
norms["absolute max"] += absolute_max.item()
def train_one_step_mochi(
transformer,
teacher_transformer,
ema_transformer,
optimizer,
lr_scheduler,
loader,
noise_scheduler,
solver,
noise_random_generator,
gradient_accumulation_steps,
sp_size,
max_grad_norm,
uncond_prompt_embed,
uncond_prompt_mask,
num_euler_timesteps,
multiphase,
not_apply_cfg_solver,
distill_cfg,
ema_decay,
pred_decay_weight,
pred_decay_type,
):
total_loss = 0.0
optimizer.zero_grad()
model_pred_norm = {
"fro": 0.0,
"largest singular value": 0.0,
"absolute mean": 0.0,
"absolute max": 0.0,
}
for _ in range(gradient_accumulation_steps):
(
latents,
encoder_hidden_states,
latents_attention_mask,
encoder_attention_mask,
) = next(loader)
model_input = normalize_mochi_dit_input(latents)
noise = torch.randn_like(model_input)
bsz = model_input.shape[0]
index = torch.randint(
0, num_euler_timesteps, (bsz,), device=model_input.device
).long()
if sp_size > 1:
broadcast(index)
# Add noise according to flow matching.
# sigmas = get_sigmas(start_timesteps, n_dim=model_input.ndim, dtype=model_input.dtype)
sigmas = extract_into_tensor(solver.sigmas, index, model_input.shape)
sigmas_prev = extract_into_tensor(solver.sigmas_prev, index, model_input.shape)
timesteps = (sigmas * noise_scheduler.config.num_train_timesteps).view(-1)
# if squeeze to [], unsqueeze to [1]
timesteps_prev = (
sigmas_prev * noise_scheduler.config.num_train_timesteps
).view(-1)
noisy_model_input = sigmas * noise + (1.0 - sigmas) * model_input
# Predict the noise residual
with torch.autocast("cuda", dtype=torch.bfloat16):
model_pred = transformer(
noisy_model_input,
encoder_hidden_states,
timesteps,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
# if accelerator.is_main_process:
model_pred, end_index = solver.euler_style_multiphase_pred(
noisy_model_input, model_pred, index, multiphase
)
with torch.no_grad():
w = distill_cfg
with torch.autocast("cuda", dtype=torch.bfloat16):
cond_teacher_output = teacher_transformer(
noisy_model_input,
encoder_hidden_states,
timesteps,
encoder_attention_mask, # B, L
return_dict=False,
)[0].float()
if not_apply_cfg_solver:
uncond_teacher_output = cond_teacher_output
else:
# Get teacher model prediction on noisy_latents and unconditional embedding
with torch.autocast("cuda", dtype=torch.bfloat16):
uncond_teacher_output = teacher_transformer(
noisy_model_input,
uncond_prompt_embed.unsqueeze(0).expand(bsz, -1, -1),
timesteps,
uncond_prompt_mask.unsqueeze(0).expand(bsz, -1),
return_dict=False,
)[0].float()
teacher_output = cond_teacher_output + w * (
cond_teacher_output - uncond_teacher_output
)
x_prev = solver.euler_step(noisy_model_input, teacher_output, index)
# 20.4.12. Get target LCM prediction on x_prev, w, c, t_n
with torch.no_grad():
with torch.autocast("cuda", dtype=torch.bfloat16):
if ema_transformer is not None:
target_pred = ema_transformer(
x_prev.float(),
encoder_hidden_states,
timesteps_prev,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
else:
target_pred = transformer(
x_prev.float(),
encoder_hidden_states,
timesteps_prev,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
target, end_index = solver.euler_style_multiphase_pred(
x_prev, target_pred, index, multiphase, True
)
huber_c = 0.001
# loss = loss.mean()
loss = (
torch.mean(
torch.sqrt((model_pred.float() - target.float()) ** 2 + huber_c**2)
- huber_c
)
/ gradient_accumulation_steps
)
if pred_decay_weight > 0:
if pred_decay_type == "l1":
pred_decay_loss = (
torch.mean(torch.sqrt(model_pred.float() ** 2))
* pred_decay_weight
/ gradient_accumulation_steps
)
loss += pred_decay_loss
elif pred_decay_type == "l2":
# essnetially k2?
pred_decay_loss = (
torch.mean(model_pred.float() ** 2)
* pred_decay_weight
/ gradient_accumulation_steps
)
loss += pred_decay_loss
else:
assert NotImplementedError("pred_decay_type is not implemented")
# calculate model_pred norm and mean
get_norm(
model_pred.detach().float(), model_pred_norm, gradient_accumulation_steps
)
loss.backward()
avg_loss = loss.detach().clone()
dist.all_reduce(avg_loss, op=dist.ReduceOp.AVG)
dist.all_reduce(pred_decay_loss.detach(), op=dist.ReduceOp.AVG)
total_loss += avg_loss.item()
# update ema
if ema_transformer is not None:
reshard_fsdp(ema_transformer)
for p_averaged, p_model in zip(
ema_transformer.parameters(), transformer.parameters()
):
with torch.no_grad():
p_averaged.copy_(
torch.lerp(p_averaged.detach(), p_model.detach(), 1 - ema_decay)
)
grad_norm = transformer.clip_grad_norm_(max_grad_norm)
optimizer.step()
lr_scheduler.step()
return total_loss, grad_norm.item(), model_pred_norm, pred_decay_loss.item()
def main(args):
torch.backends.cuda.matmul.allow_tf32 = True
local_rank = int(os.environ["LOCAL_RANK"])
rank = int(os.environ["RANK"])
world_size = int(os.environ["WORLD_SIZE"])
dist.init_process_group("nccl")
torch.cuda.set_device(local_rank)
device = torch.cuda.current_device()
initialize_sequence_parallel_state(args.sp_size)
# If passed along, set the training seed now. On GPU...
if args.seed is not None:
# TODO: t within the same seq parallel group should be the same. Noise should be different.
set_seed(args.seed + rank)
# We use different seeds for the noise generation in each process to ensure that the noise is different in a batch.
noise_random_generator = None
# Handle the repository creation
if rank <= 0 and args.output_dir is not None:
os.makedirs(args.output_dir, exist_ok=True)
# For mixed precision training we cast all non-trainable weigths to half-precision
# as these weights are only used for inference, keeping weights in full precision is not required.
# Create model:
main_print(f"--> loading model from {args.pretrained_model_name_or_path}")
# keep the master weight to float32
if args.dit_model_name_or_path:
transformer = transformer = MochiTransformer3DModel.from_pretrained(
args.dit_model_name_or_path,
torch_dtype=torch.float32,
# torch_dtype=torch.bfloat16 if args.use_lora else torch.float32,
)
else:
transformer = MochiTransformer3DModel.from_pretrained(
args.pretrained_model_name_or_path,
subfolder="transformer",
torch_dtype=torch.float32,
# torch_dtype=torch.bfloat16 if args.use_lora else torch.float32,
)
teacher_transformer = deepcopy(transformer)
if args.use_ema:
ema_transformer = deepcopy(transformer)
else:
ema_transformer = None
if args.use_lora:
transformer.requires_grad_(False)
transformer_lora_config = LoraConfig(
r=args.lora_rank,
lora_alpha=args.lora_alpha,
init_lora_weights=True,
target_modules=["to_k", "to_q", "to_v", "to_out.0"],
)
transformer.add_adapter(transformer_lora_config)
main_print(
f" Total training parameters = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e6} M"
)
main_print(
f"--> Initializing FSDP with sharding strategy: {args.fsdp_sharding_startegy}"
)
fsdp_kwargs = get_dit_fsdp_kwargs(
args.fsdp_sharding_startegy,
args.use_lora,
args.use_cpu_offload,
args.master_weight_type,
)
if args.use_lora:
transformer.config.lora_rank = args.lora_rank
transformer.config.lora_alpha = args.lora_alpha
transformer.config.lora_target_modules = ["to_k", "to_q", "to_v", "to_out.0"]
transformer._no_split_modules = ["MochiTransformerBlock"]
fsdp_kwargs["auto_wrap_policy"] = fsdp_kwargs["auto_wrap_policy"](transformer)
transformer = FSDP(
transformer,
**fsdp_kwargs,
)
teacher_transformer = FSDP(
teacher_transformer,
**fsdp_kwargs,
)
if args.use_ema:
ema_transformer = FSDP(
ema_transformer,
**fsdp_kwargs,
)
main_print(f"--> model loaded")
if args.gradient_checkpointing:
apply_fsdp_checkpointing(transformer, args.selective_checkpointing)
apply_fsdp_checkpointing(teacher_transformer, args.selective_checkpointing)
if args.use_ema:
apply_fsdp_checkpointing(ema_transformer, args.selective_checkpointing)
# Set model as trainable.
transformer.train()
teacher_transformer.requires_grad_(False)
if args.use_ema:
ema_transformer.requires_grad_(False)
noise_scheduler = FlowMatchEulerDiscreteScheduler(shift=args.shift)
if args.scheduler_type == "pcm_linear_quadratic":
linear_steps = int(
noise_scheduler.config.num_train_timesteps * args.linear_range
)
sigmas = linear_quadratic_schedule(
noise_scheduler.config.num_train_timesteps,
args.linear_quadratic_threshold,
linear_steps,
)
sigmas = torch.tensor(sigmas).to(dtype=torch.float32)
else:
sigmas = noise_scheduler.sigmas
solver = EulerSolver(
sigmas.numpy()[::-1],
noise_scheduler.config.num_train_timesteps,
euler_timesteps=args.num_euler_timesteps,
)
solver.to(device)
params_to_optimize = transformer.parameters()
params_to_optimize = list(filter(lambda p: p.requires_grad, params_to_optimize))
optimizer = torch.optim.AdamW(
params_to_optimize,
lr=args.learning_rate,
betas=(0.9, 0.999),
weight_decay=args.weight_decay,
eps=1e-8,
)
init_steps = 0
if args.resume_from_lora_checkpoint:
transformer, optimizer, init_steps = resume_lora_optimizer(
transformer, args.resume_from_lora_checkpoint, optimizer
)
main_print(f"optimizer: {optimizer}")
# todo add lr scheduler
lr_scheduler = get_scheduler(
args.lr_scheduler,
optimizer=optimizer,
num_warmup_steps=args.lr_warmup_steps * world_size,
num_training_steps=args.max_train_steps * world_size,
num_cycles=args.lr_num_cycles,
power=args.lr_power,
last_epoch=init_steps - 1,
)
train_dataset = LatentDataset(args.data_json_path, args.num_latent_t, args.cfg)
uncond_prompt_embed = train_dataset.uncond_prompt_embed
uncond_prompt_mask = train_dataset.uncond_prompt_mask
sampler = (
LengthGroupedSampler(
args.train_batch_size,
rank=rank,
world_size=world_size,
lengths=train_dataset.lengths,
group_frame=args.group_frame,
group_resolution=args.group_resolution,
)
if (args.group_frame or args.group_resolution)
else DistributedSampler(
train_dataset, rank=rank, num_replicas=world_size, shuffle=False
)
)
train_dataloader = DataLoader(
train_dataset,
sampler=sampler,
collate_fn=latent_collate_function,
pin_memory=True,
batch_size=args.train_batch_size,
num_workers=args.dataloader_num_workers,
drop_last=True,
)
num_update_steps_per_epoch = math.ceil(
len(train_dataloader)
/ args.gradient_accumulation_steps
* args.sp_size
/ args.train_sp_batch_size
)
args.num_train_epochs = math.ceil(args.max_train_steps / num_update_steps_per_epoch)
if rank <= 0:
project = args.tracker_project_name or "fastvideo"
wandb.init(project=project, config=args)
# Train!
total_batch_size = (
args.train_batch_size
* world_size
* args.gradient_accumulation_steps
/ args.sp_size
* args.train_sp_batch_size
)
main_print("***** Running training *****")
main_print(f" Num examples = {len(train_dataset)}")
main_print(f" Dataloader size = {len(train_dataloader)}")
main_print(f" Num Epochs = {args.num_train_epochs}")
main_print(f" Resume training from step {init_steps}")
main_print(f" Instantaneous batch size per device = {args.train_batch_size}")
main_print(
f" Total train batch size (w. data & sequence parallel, accumulation) = {total_batch_size}"
)
main_print(f" Gradient Accumulation steps = {args.gradient_accumulation_steps}")
main_print(f" Total optimization steps = {args.max_train_steps}")
main_print(
f" Total training parameters per FSDP shard = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e9} B"
)
# print dtype
main_print(f" Master weight dtype: {transformer.parameters().__next__().dtype}")
# Potentially load in the weights and states from a previous save
if args.resume_from_checkpoint:
assert NotImplementedError("resume_from_checkpoint is not supported now.")
# TODO
progress_bar = tqdm(
range(0, args.max_train_steps),
initial=init_steps,
desc="Steps",
# Only show the progress bar once on each machine.
disable=local_rank > 0,
)
loader = sp_parallel_dataloader_wrapper(
train_dataloader,
device,
args.train_batch_size,
args.sp_size,
args.train_sp_batch_size,
)
step_times = deque(maxlen=100)
# todo future
for i in range(init_steps):
next(loader)
# log_validation(args, transformer, device,
# torch.bfloat16, 0, scheduler_type=args.scheduler_type, shift=args.shift, num_euler_timesteps=args.num_euler_timesteps, linear_quadratic_threshold=args.linear_quadratic_threshold,ema=False)
def get_num_phases(multi_phased_distill_schedule, step):
# step-phase,step-phase
multi_phases = multi_phased_distill_schedule.split(",")
phase = multi_phases[-1].split("-")[-1]
for step_phases in multi_phases:
phase_step, phase = step_phases.split("-")
if step <= int(phase_step):
return int(phase)
return phase
for step in range(init_steps + 1, args.max_train_steps + 1):
start_time = time.time()
assert args.multi_phased_distill_schedule is not None
num_phases = get_num_phases(args.multi_phased_distill_schedule, step)
loss, grad_norm, pred_norm, aux_loss = train_one_step_mochi(
transformer,
teacher_transformer,
ema_transformer,
optimizer,
lr_scheduler,
loader,
noise_scheduler,
solver,
noise_random_generator,
args.gradient_accumulation_steps,
args.sp_size,
args.max_grad_norm,
uncond_prompt_embed,
uncond_prompt_mask,
args.num_euler_timesteps,
num_phases,
args.not_apply_cfg_solver,
args.distill_cfg,
args.ema_decay,
args.pred_decay_weight,
args.pred_decay_type,
)
step_time = time.time() - start_time
step_times.append(step_time)
avg_step_time = sum(step_times) / len(step_times)
progress_bar.set_postfix(
{
"loss": f"{loss:.4f}",
"step_time": f"{step_time:.2f}s",
"grad_norm": grad_norm,
"phases": num_phases,
}
)
progress_bar.update(1)
if rank <= 0:
wandb.log(
{
"train_loss": loss,
"learning_rate": lr_scheduler.get_last_lr()[0],
"step_time": step_time,
"avg_step_time": avg_step_time,
"grad_norm": grad_norm,
"pred_fro_norm": pred_norm["fro"],
"pred_largest_singular_value": pred_norm["largest singular value"],
"pred_absolute_mean": pred_norm["absolute mean"],
"pred_absolute_max": pred_norm["absolute max"],
"aux_loss": aux_loss,
},
step=step,
)
if step % args.checkpointing_steps == 0:
if args.use_lora:
# Save LoRA weights
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, step
)
else:
# Your existing checkpoint saving code
if args.use_ema:
save_checkpoint(ema_transformer, rank, args.output_dir, step)
else:
save_checkpoint(transformer, rank, args.output_dir, step)
dist.barrier()
if args.log_validation and step % args.validation_steps == 0:
log_validation(
args,
transformer,
device,
torch.bfloat16,
step,
scheduler_type=args.scheduler_type,
shift=args.shift,
num_euler_timesteps=args.num_euler_timesteps,
linear_quadratic_threshold=args.linear_quadratic_threshold,
linear_range=args.linear_range,
ema=False,
)
if args.use_ema:
log_validation(
args,
ema_transformer,
device,
torch.bfloat16,
step,
scheduler_type=args.scheduler_type,
shift=args.shift,
num_euler_timesteps=args.num_euler_timesteps,
linear_quadratic_threshold=args.linear_quadratic_threshold,
linear_range=args.linear_range,
ema=True,
)
if args.use_lora:
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, args.max_train_steps
)
else:
save_checkpoint(transformer, rank, args.output_dir, args.max_train_steps)
if get_sequence_parallel_state():
destroy_sequence_parallel_group()
if __name__ == "__main__":
parser = argparse.ArgumentParser()
# dataset & dataloader
parser.add_argument("--data_json_path", type=str, required=True)
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument(
"--dataloader_num_workers",
type=int,
default=10,
help="Number of subprocesses to use for data loading. 0 means that the data will be loaded in the main process.",
)
parser.add_argument(
"--train_batch_size",
type=int,
default=16,
help="Batch size (per device) for the training dataloader.",
)
parser.add_argument(
"--num_latent_t", type=int, default=28, help="Number of latent timesteps."
)
parser.add_argument("--group_frame", action="store_true") # TODO
parser.add_argument("--group_resolution", action="store_true") # TODO
# text encoder & vae & diffusion model
parser.add_argument("--pretrained_model_name_or_path", type=str)
parser.add_argument("--dit_model_name_or_path", type=str)
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
# diffusion setting
parser.add_argument("--ema_decay", type=float, default=0.95)
parser.add_argument("--ema_start_step", type=int, default=0)
parser.add_argument("--cfg", type=float, default=0.1)
# validation & logs
parser.add_argument("--validation_prompt_dir", type=str)
parser.add_argument("--validation_sampling_steps", type=str, default="64")
parser.add_argument("--validation_guidance_scale", type=str, default="4.5")
parser.add_argument("--validation_steps", type=float, default=64)
parser.add_argument("--log_validation", action="store_true")
parser.add_argument("--tracker_project_name", type=str, default=None)
parser.add_argument(
"--seed", type=int, default=None, help="A seed for reproducible training."
)
parser.add_argument(
"--output_dir",
type=str,
default=None,
help="The output directory where the model predictions and checkpoints will be written.",
)
parser.add_argument(
"--checkpoints_total_limit",
type=int,
default=None,
help=("Max number of checkpoints to store."),
)
parser.add_argument(
"--checkpointing_steps",
type=int,
default=500,
help=(
"Save a checkpoint of the training state every X updates. These checkpoints can be used both as final"
" checkpoints in case they are better than the last checkpoint, and are also suitable for resuming"
" training using `--resume_from_checkpoint`."
),
)
parser.add_argument("--shift", type=float, default=1.0)
parser.add_argument(
"--resume_from_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--resume_from_lora_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous lora checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--logging_dir",
type=str,
default="logs",
help=(
"[TensorBoard](https://www.tensorflow.org/tensorboard) log directory. Will default to"
" *output_dir/runs/**CURRENT_DATETIME_HOSTNAME***."
),
)
# optimizer & scheduler & Training
parser.add_argument("--num_train_epochs", type=int, default=100)
parser.add_argument(
"--max_train_steps",
type=int,
default=None,
help="Total number of training steps to perform. If provided, overrides num_train_epochs.",
)
parser.add_argument(
"--gradient_accumulation_steps",
type=int,
default=1,
help="Number of updates steps to accumulate before performing a backward/update pass.",
)
parser.add_argument(
"--learning_rate",
type=float,
default=1e-4,
help="Initial learning rate (after the potential warmup period) to use.",
)
parser.add_argument(
"--scale_lr",
action="store_true",
default=False,
help="Scale the learning rate by the number of GPUs, gradient accumulation steps, and batch size.",
)
parser.add_argument(
"--lr_warmup_steps",
type=int,
default=10,
help="Number of steps for the warmup in the lr scheduler.",
)
parser.add_argument(
"--max_grad_norm", default=1.0, type=float, help="Max gradient norm."
)
parser.add_argument(
"--gradient_checkpointing",
action="store_true",
help="Whether or not to use gradient checkpointing to save memory at the expense of slower backward pass.",
)
parser.add_argument("--selective_checkpointing", type=float, default=1.0)
parser.add_argument(
"--allow_tf32",
action="store_true",
help=(
"Whether or not to allow TF32 on Ampere GPUs. Can be used to speed up training. For more information, see"
" https://pytorch.org/docs/stable/notes/cuda.html#tensorfloat-32-tf32-on-ampere-devices"
),
)
parser.add_argument(
"--mixed_precision",
type=str,
default=None,
choices=["no", "fp16", "bf16"],
help=(
"Whether to use mixed precision. Choose between fp16 and bf16 (bfloat16). Bf16 requires PyTorch >="
" 1.10.and an Nvidia Ampere GPU. Default to the value of accelerate config of the current system or the"
" flag passed with the `accelerate.launch` command. Use this argument to override the accelerate config."
),
)
parser.add_argument(
"--use_cpu_offload",
action="store_true",
help="Whether to use CPU offload for param & gradient & optimizer states.",
)
parser.add_argument("--sp_size", type=int, default=1, help="For sequence parallel")
parser.add_argument(
"--train_sp_batch_size",
type=int,
default=1,
help="Batch size for sequence parallel training",
)
parser.add_argument(
"--use_lora",
action="store_true",
default=False,
help="Whether to use LoRA for finetuning.",
)
parser.add_argument(
"--lora_alpha", type=int, default=256, help="Alpha parameter for LoRA."
)
parser.add_argument(
"--lora_rank", type=int, default=128, help="LoRA rank parameter. "
)
parser.add_argument("--fsdp_sharding_startegy", default="full")
# lr_scheduler
parser.add_argument(
"--lr_scheduler",
type=str,
default="constant",
help=(
'The scheduler type to use. Choose between ["linear", "cosine", "cosine_with_restarts", "polynomial",'
' "constant", "constant_with_warmup"]'
),
)
parser.add_argument("--num_euler_timesteps", type=int, default=100)
parser.add_argument(
"--lr_num_cycles",
type=int,
default=1,
help="Number of cycles in the learning rate scheduler.",
)
parser.add_argument(
"--lr_power",
type=float,
default=1.0,
help="Power factor of the polynomial scheduler.",
)
parser.add_argument(
"--not_apply_cfg_solver",
action="store_true",
help="Whether to apply the cfg_solver.",
)
parser.add_argument(
"--distill_cfg", type=float, default=3.0, help="Distillation coefficient."
)
# ["euler_linear_quadratic", "pcm", "pcm_linear_qudratic"]
parser.add_argument(
"--scheduler_type", type=str, default="pcm", help="The scheduler type to use."
)
parser.add_argument(
"--linear_quadratic_threshold",
type=float,
default=0.025,
help="Threshold for linear quadratic scheduler.",
)
parser.add_argument(
"--linear_range",
type=float,
default=0.5,
help="Range for linear quadratic scheduler.",
)
parser.add_argument(
"--weight_decay", type=float, default=0.001, help="Weight decay to apply."
)
parser.add_argument("--use_ema", action="store_true", help="Whether to use EMA.")
parser.add_argument("--multi_phased_distill_schedule", type=str, default=None)
parser.add_argument("--finetune_weight", type=float, default=0.0)
parser.add_argument("--pred_decay_weight", type=float, default=0.0)
parser.add_argument("--pred_decay_type", default="l1")
parser.add_argument(
"--master_weight_type",
type=str,
default="fp32",
help="Weight type to use - fp32 or bf16.",
)
args = parser.parse_args()
main(args)
-102
View File
@@ -1,102 +0,0 @@
from typing import Any, Dict, Optional, Union
import torch
import torch.nn as nn
from diffusers.configuration_utils import ConfigMixin, register_to_config
from diffusers.loaders import FromOriginalModelMixin, PeftAdapterMixin
from diffusers.models.attention import JointTransformerBlock
from diffusers.models.attention_processor import Attention, AttentionProcessor
from diffusers.models.modeling_utils import ModelMixin
from diffusers.models.normalization import AdaLayerNormContinuous
from diffusers.utils import (
USE_PEFT_BACKEND,
is_torch_version,
logging,
scale_lora_layers,
unscale_lora_layers,
)
from diffusers.models.embeddings import CombinedTimestepTextProjEmbeddings, PatchEmbed
from diffusers.models.transformers.transformer_2d import Transformer2DModelOutput
from diffusers.models.transformers.transformer_sd3 import SD3Transformer2DModel
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
class DiscriminatorHead(nn.Module):
def __init__(self, input_channel, output_channel=1):
super().__init__()
inner_channel = 1024
self.conv1 = nn.Sequential(
nn.Conv2d(input_channel, inner_channel, 1, 1, 0),
nn.GroupNorm(32, inner_channel),
nn.LeakyReLU(
inplace=True
), # use LeakyReLu instead of GELU shown in the paper to save memory
)
self.conv2 = nn.Sequential(
nn.Conv2d(inner_channel, inner_channel, 1, 1, 0),
nn.GroupNorm(32, inner_channel),
nn.LeakyReLU(
inplace=True
), # use LeakyReLu instead of GELU shown in the paper to save memory
)
self.conv_out = nn.Conv2d(inner_channel, output_channel, 1, 1, 0)
def forward(self, x):
b, twh, c = x.shape
t = twh // (30 * 53)
x = x.view(-1, 30 * 53, c)
x = x.permute(0, 2, 1)
x = x.view(b * t, c, 30, 53)
x = self.conv1(x)
x = self.conv2(x) + x
x = self.conv_out(x)
return x
class Discriminator(nn.Module):
def __init__(
self,
stride=8,
num_h_per_head=1,
adapter_channel_dims=[3072],
):
super().__init__()
adapter_channel_dims = adapter_channel_dims * (48 // stride)
self.stride = stride
self.num_h_per_head = num_h_per_head
self.head_num = len(adapter_channel_dims)
self.heads = nn.ModuleList(
[
nn.ModuleList(
[
DiscriminatorHead(adapter_channel)
for _ in range(self.num_h_per_head)
]
)
for adapter_channel in adapter_channel_dims
]
)
def forward(self, features):
outputs = []
def create_custom_forward(module):
def custom_forward(*inputs):
return module(*inputs)
return custom_forward
assert len(features) // self.stride == len(self.heads)
for i in range(0, len(features), self.stride):
for h in self.heads[i // self.stride]:
# out = torch.utils.checkpoint.checkpoint(
# create_custom_forward(h),
# features[i],
# use_reentrant=False
# )
out = h(features[i])
outputs.append(out)
return outputs
-309
View File
@@ -1,309 +0,0 @@
from dataclasses import dataclass
from typing import Optional, Tuple, Union
import numpy as np
import torch
from diffusers.configuration_utils import ConfigMixin, register_to_config
from diffusers.utils import BaseOutput, logging
from diffusers.utils.torch_utils import randn_tensor
from diffusers.schedulers.scheduling_utils import SchedulerMixin
from fastvideo.models.mochi_hf.pipeline_mochi import linear_quadratic_schedule
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
@dataclass
class PCMFMSchedulerOutput(BaseOutput):
prev_sample: torch.FloatTensor
def extract_into_tensor(a, t, x_shape):
b, *_ = t.shape
out = a.gather(-1, t)
return out.reshape(b, *((1,) * (len(x_shape) - 1)))
class PCMFMScheduler(SchedulerMixin, ConfigMixin):
_compatibles = []
order = 1
@register_to_config
def __init__(
self,
num_train_timesteps: int = 1000,
shift: float = 1.0,
pcm_timesteps: int = 50,
linear_quadratic=False,
linear_quadratic_threshold=0.025,
linear_range=0.5,
):
if linear_quadratic:
linear_steps = int(num_train_timesteps * linear_range)
sigmas = linear_quadratic_schedule(
num_train_timesteps, linear_quadratic_threshold, linear_steps
)
sigmas = torch.tensor(sigmas).to(dtype=torch.float32)
else:
timesteps = np.linspace(
1, num_train_timesteps, num_train_timesteps, dtype=np.float32
)[::-1].copy()
timesteps = torch.from_numpy(timesteps).to(dtype=torch.float32)
sigmas = timesteps / num_train_timesteps
sigmas = shift * sigmas / (1 + (shift - 1) * sigmas)
self.euler_timesteps = (
np.arange(1, pcm_timesteps + 1) * (num_train_timesteps // pcm_timesteps)
).round().astype(np.int64) - 1
self.sigmas = sigmas.numpy()[::-1][self.euler_timesteps]
self.sigmas = torch.from_numpy((self.sigmas[::-1].copy()))
self.timesteps = self.sigmas * num_train_timesteps
self._step_index = None
self._begin_index = None
self.sigmas = self.sigmas.to("cpu") # to avoid too much CPU/GPU communication
self.sigma_min = self.sigmas[-1].item()
self.sigma_max = self.sigmas[0].item()
@property
def step_index(self):
"""
The index counter for current timestep. It will increase 1 after each scheduler step.
"""
return self._step_index
@property
def begin_index(self):
"""
The index for the first timestep. It should be set from pipeline with `set_begin_index` method.
"""
return self._begin_index
# Copied from diffusers.schedulers.scheduling_dpmsolver_multistep.DPMSolverMultistepScheduler.set_begin_index
def set_begin_index(self, begin_index: int = 0):
"""
Sets the begin index for the scheduler. This function should be run from pipeline before the inference.
Args:
begin_index (`int`):
The begin index for the scheduler.
"""
self._begin_index = begin_index
def scale_noise(
self,
sample: torch.FloatTensor,
timestep: Union[float, torch.FloatTensor],
noise: Optional[torch.FloatTensor] = None,
) -> torch.FloatTensor:
"""
Forward process in flow-matching
Args:
sample (`torch.FloatTensor`):
The input sample.
timestep (`int`, *optional*):
The current timestep in the diffusion chain.
Returns:
`torch.FloatTensor`:
A scaled input sample.
"""
if self.step_index is None:
self._init_step_index(timestep)
sigma = self.sigmas[self.step_index]
sample = sigma * noise + (1.0 - sigma) * sample
return sample
def _sigma_to_t(self, sigma):
return sigma * self.config.num_train_timesteps
def set_timesteps(
self, num_inference_steps: int, device: Union[str, torch.device] = None
):
"""
Sets the discrete timesteps used for the diffusion chain (to be run before inference).
Args:
num_inference_steps (`int`):
The number of diffusion steps used when generating samples with a pre-trained model.
device (`str` or `torch.device`, *optional*):
The device to which the timesteps should be moved to. If `None`, the timesteps are not moved.
"""
self.num_inference_steps = num_inference_steps
inference_indices = np.linspace(
0, self.config.pcm_timesteps, num=num_inference_steps, endpoint=False
)
inference_indices = np.floor(inference_indices).astype(np.int64)
inference_indices = torch.from_numpy(inference_indices).long()
self.sigmas_ = self.sigmas[inference_indices]
timesteps = self.sigmas_ * self.config.num_train_timesteps
self.timesteps = timesteps.to(device=device)
self.sigmas_ = torch.cat(
[self.sigmas_, torch.zeros(1, device=self.sigmas_.device)]
)
self._step_index = None
self._begin_index = None
def index_for_timestep(self, timestep, schedule_timesteps=None):
if schedule_timesteps is None:
schedule_timesteps = self.timesteps
indices = (schedule_timesteps == timestep).nonzero()
# The sigma index that is taken for the **very** first `step`
# is always the second index (or the last index if there is only 1)
# This way we can ensure we don't accidentally skip a sigma in
# case we start in the middle of the denoising schedule (e.g. for image-to-image)
pos = 1 if len(indices) > 1 else 0
return indices[pos].item()
def _init_step_index(self, timestep):
if self.begin_index is None:
if isinstance(timestep, torch.Tensor):
timestep = timestep.to(self.timesteps.device)
self._step_index = self.index_for_timestep(timestep)
else:
self._step_index = self._begin_index
def step(
self,
model_output: torch.FloatTensor,
timestep: Union[float, torch.FloatTensor],
sample: torch.FloatTensor,
generator: Optional[torch.Generator] = None,
return_dict: bool = True,
) -> Union[PCMFMSchedulerOutput, Tuple]:
"""
Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion
process from the learned model outputs (most often the predicted noise).
Args:
model_output (`torch.FloatTensor`):
The direct output from learned diffusion model.
timestep (`float`):
The current discrete timestep in the diffusion chain.
sample (`torch.FloatTensor`):
A current instance of a sample created by the diffusion process.
s_churn (`float`):
s_tmin (`float`):
s_tmax (`float`):
s_noise (`float`, defaults to 1.0):
Scaling factor for noise added to the sample.
generator (`torch.Generator`, *optional*):
A random number generator.
return_dict (`bool`):
Whether or not to return a [`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] or
tuple.
Returns:
[`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] or `tuple`:
If return_dict is `True`, [`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] is
returned, otherwise a tuple is returned where the first element is the sample tensor.
"""
if (
isinstance(timestep, int)
or isinstance(timestep, torch.IntTensor)
or isinstance(timestep, torch.LongTensor)
):
raise ValueError(
(
"Passing integer indices (e.g. from `enumerate(timesteps)`) as timesteps to"
" `EulerDiscreteScheduler.step()` is not supported. Make sure to pass"
" one of the `scheduler.timesteps` as a timestep."
),
)
if self.step_index is None:
self._init_step_index(timestep)
sample = sample.to(torch.float32)
sigma = self.sigmas_[self.step_index]
denoised = sample - model_output * sigma
derivative = (sample - denoised) / sigma
dt = self.sigmas_[self.step_index + 1] - sigma
prev_sample = sample + derivative * dt
prev_sample = prev_sample.to(model_output.dtype)
self._step_index += 1
if not return_dict:
return (prev_sample,)
return PCMFMSchedulerOutput(prev_sample=prev_sample)
def __len__(self):
return self.config.num_train_timesteps
class EulerSolver:
def __init__(self, sigmas, timesteps=1000, euler_timesteps=50):
self.step_ratio = timesteps // euler_timesteps
self.euler_timesteps = (
np.arange(1, euler_timesteps + 1) * self.step_ratio
).round().astype(np.int64) - 1
self.euler_timesteps_prev = np.asarray([0] + self.euler_timesteps[:-1].tolist())
self.sigmas = sigmas[self.euler_timesteps]
self.sigmas_prev = np.asarray(
[sigmas[0]] + sigmas[self.euler_timesteps[:-1]].tolist()
) # either use sigma0 or 0
self.euler_timesteps = torch.from_numpy(self.euler_timesteps).long()
self.euler_timesteps_prev = torch.from_numpy(self.euler_timesteps_prev).long()
self.sigmas = torch.from_numpy(self.sigmas)
self.sigmas_prev = torch.from_numpy(self.sigmas_prev)
def to(self, device):
self.euler_timesteps = self.euler_timesteps.to(device)
self.euler_timesteps_prev = self.euler_timesteps_prev.to(device)
self.sigmas = self.sigmas.to(device)
self.sigmas_prev = self.sigmas_prev.to(device)
return self
def euler_step(self, sample, model_pred, timestep_index):
sigma = extract_into_tensor(self.sigmas, timestep_index, model_pred.shape)
sigma_prev = extract_into_tensor(
self.sigmas_prev, timestep_index, model_pred.shape
)
x_prev = sample + (sigma_prev - sigma) * model_pred
return x_prev
def euler_style_multiphase_pred(
self,
sample,
model_pred,
timestep_index,
multiphase,
is_target=False,
):
inference_indices = np.linspace(
0, len(self.euler_timesteps), num=multiphase, endpoint=False
)
inference_indices = np.floor(inference_indices).astype(np.int64)
inference_indices = (
torch.from_numpy(inference_indices).long().to(self.euler_timesteps.device)
)
expanded_timestep_index = timestep_index.unsqueeze(1).expand(
-1, inference_indices.size(0)
)
valid_indices_mask = expanded_timestep_index >= inference_indices
last_valid_index = valid_indices_mask.flip(dims=[1]).long().argmax(dim=1)
last_valid_index = inference_indices.size(0) - 1 - last_valid_index
timestep_index_end = inference_indices[last_valid_index]
if is_target:
sigma = extract_into_tensor(self.sigmas_prev, timestep_index, sample.shape)
else:
sigma = extract_into_tensor(self.sigmas, timestep_index, sample.shape)
sigma_prev = extract_into_tensor(
self.sigmas_prev, timestep_index_end, sample.shape
)
x_prev = sample + (sigma_prev - sigma) * model_pred
return x_prev, timestep_index_end
-936
View File
@@ -1,936 +0,0 @@
import argparse
from email.policy import strict
import logging
import math
import os
import shutil
from pathlib import Path
from fastvideo.utils.parallel_states import (
initialize_sequence_parallel_state,
destroy_sequence_parallel_group,
get_sequence_parallel_state,
nccl_info,
)
from fastvideo.utils.communications import sp_parallel_dataloader_wrapper, broadcast
from fastvideo.models.mochi_hf.mochi_latents_utils import normalize_mochi_dit_input
from fastvideo.utils.validation import log_validation
import time
from torch.utils.data import DataLoader
import torch
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
StateDictType,
FullStateDictConfig,
)
from fastvideo.models.mochi_hf.pipeline_mochi import linear_quadratic_schedule
import json
from torch.utils.data.distributed import DistributedSampler
from fastvideo.utils.dataset_utils import LengthGroupedSampler
import wandb
from accelerate.utils import set_seed
from tqdm.auto import tqdm
from fastvideo.fsdp_util import (
get_dit_fsdp_kwargs,
apply_fsdp_checkpointing,
get_discriminator_fsdp_kwargs,
)
import diffusers
from diffusers import (
FlowMatchEulerDiscreteScheduler,
)
from fastvideo.distill.discriminator import Discriminator
from fastvideo.distill.solver import EulerSolver, extract_into_tensor
from copy import deepcopy
from diffusers.optimization import get_scheduler
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
from diffusers.utils import check_min_version
from fastvideo.dataset.latent_datasets import LatentDataset, latent_collate_function
import torch.distributed as dist
from peft import LoraConfig
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
)
from fastvideo.utils.checkpoint import (
save_checkpoint,
save_lora_checkpoint,
resume_lora_optimizer,
resume_training,
save_checkpoint_generator_discriminator,
resume_training_generator_discriminator,
)
from fastvideo.utils.logging import main_print
# Will error if the minimal version of diffusers is not installed. Remove at your own risks.
check_min_version("0.31.0")
import time
from collections import deque
def gan_d_loss(
discriminator,
teacher_transformer,
sample_fake,
sample_real,
timestep,
encoder_hidden_states,
encoder_attention_mask,
weight,
):
loss = 0.0
# collate sample_fake and sample_real
with torch.no_grad():
fake_features = teacher_transformer(
sample_fake,
encoder_hidden_states,
timestep,
encoder_attention_mask,
output_attn=True,
return_dict=False,
)[1]
real_features = teacher_transformer(
sample_real,
encoder_hidden_states,
timestep,
encoder_attention_mask,
output_attn=True,
return_dict=False,
)[1]
fake_outputs = discriminator(fake_features)
real_outputs = discriminator(real_features)
for fake_output, real_output in zip(fake_outputs, real_outputs):
loss += (
torch.mean(weight * torch.relu(fake_output.float() + 1))
+ torch.mean(weight * torch.relu(1 - real_output.float()))
) / (discriminator.head_num * discriminator.num_h_per_head)
return loss
def gan_g_loss(
discriminator,
teacher_transformer,
sample_fake,
timestep,
encoder_hidden_states,
encoder_attention_mask,
weight,
):
loss = 0.0
features = teacher_transformer(
sample_fake,
encoder_hidden_states,
timestep,
encoder_attention_mask,
output_attn=True,
return_dict=False,
)[1]
fake_outputs = discriminator(
features,
)
for fake_output in fake_outputs:
loss += torch.mean(weight * torch.relu(1 - fake_output.float())) / (
discriminator.head_num * discriminator.num_h_per_head
)
return loss
def train_one_step_mochi(
transformer,
teacher_transformer,
optimizer,
discriminator,
discriminator_optimizer,
global_step,
lr_scheduler,
loader,
noise_scheduler,
solver,
noise_random_generator,
sp_size,
precondition_outputs,
max_grad_norm,
uncond_prompt_embed,
uncond_prompt_mask,
num_euler_timesteps,
multiphase,
not_apply_cfg_solver,
distill_cfg,
adv_weight,
):
optimizer.zero_grad()
discriminator_optimizer.zero_grad()
(
latents,
encoder_hidden_states,
latents_attention_mask,
encoder_attention_mask,
) = next(loader)
model_input = normalize_mochi_dit_input(latents)
noise = torch.randn_like(model_input)
bsz = model_input.shape[0]
index = torch.randint(
0, num_euler_timesteps, (bsz,), device=model_input.device
).long()
if sp_size > 1:
broadcast(index)
# Add noise according to flow matching.
# sigmas = get_sigmas(start_timesteps, n_dim=model_input.ndim, dtype=model_input.dtype)
sigmas = extract_into_tensor(solver.sigmas, index, model_input.shape)
sigmas_prev = extract_into_tensor(solver.sigmas_prev, index, model_input.shape)
timesteps = (sigmas * noise_scheduler.config.num_train_timesteps).view(-1)
# if squeeze to [], unsqueeze to [1]
timesteps_prev = (sigmas_prev * noise_scheduler.config.num_train_timesteps).view(-1)
noisy_model_input = sigmas * noise + (1.0 - sigmas) * model_input
# Predict the noise residual
with torch.autocast("cuda", dtype=torch.bfloat16):
model_pred = transformer(
noisy_model_input,
encoder_hidden_states,
timesteps,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
# if accelerator.is_main_process:
model_pred, end_index = solver.euler_style_multiphase_pred(
noisy_model_input, model_pred, index, multiphase
)
weighting = 1.0
# # simplified flow matching aka 0-rectified flow matching loss
# # target = model_input - noise
# target = model_input
adv_index = torch.empty_like(end_index)
for i in range(end_index.size(0)):
adv_index[i] = torch.randint(
end_index[i].item(),
end_index[i].item() + num_euler_timesteps // multiphase,
(1,),
dtype=end_index.dtype,
device=end_index.device,
)
sigmas_end = extract_into_tensor(solver.sigmas_prev, end_index, model_input.shape)
sigmas_adv = extract_into_tensor(solver.sigmas_prev, adv_index, model_input.shape)
timesteps_end = (sigmas_end * noise_scheduler.config.num_train_timesteps).view(-1)
timesteps_adv = (sigmas_adv * noise_scheduler.config.num_train_timesteps).view(-1)
with torch.no_grad():
w = distill_cfg
with torch.autocast("cuda", dtype=torch.bfloat16):
cond_teacher_output = teacher_transformer(
noisy_model_input,
encoder_hidden_states,
timesteps,
encoder_attention_mask, # B, L
return_dict=False,
)[0].float()
if not_apply_cfg_solver:
uncond_teacher_output = cond_teacher_output
else:
# Get teacher model prediction on noisy_latents and unconditional embedding
with torch.autocast("cuda", dtype=torch.bfloat16):
uncond_teacher_output = teacher_transformer(
noisy_model_input,
uncond_prompt_embed.unsqueeze(0).expand(bsz, -1, -1),
timesteps,
uncond_prompt_mask.unsqueeze(0).expand(bsz, -1),
return_dict=False,
)[0].float()
teacher_output = cond_teacher_output + w * (
cond_teacher_output - uncond_teacher_output
)
x_prev = solver.euler_step(noisy_model_input, teacher_output, index)
# 20.4.12. Get target LCM prediction on x_prev, w, c, t_n
with torch.no_grad():
with torch.autocast("cuda", dtype=torch.bfloat16):
target_pred = transformer(
x_prev.float(),
encoder_hidden_states,
timesteps_prev,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
target, end_index = solver.euler_style_multiphase_pred(
x_prev, target_pred, index, multiphase, True
)
real_adv = (
(1 - sigmas_adv) * target + (sigmas_adv - sigmas_end) * torch.randn_like(target)
) / (1 - sigmas_end)
fake_adv = (
(1 - sigmas_adv) * model_pred
+ (sigmas_adv - sigmas_end) * torch.randn_like(model_pred)
) / (1 - sigmas_end)
huber_c = 0.001
g_loss = torch.mean(
torch.sqrt((model_pred.float() - target.float()) ** 2 + huber_c**2) - huber_c
)
discriminator.requires_grad_(False)
with torch.autocast("cuda", dtype=torch.bfloat16):
g_gan_loss = adv_weight * gan_g_loss(
discriminator,
teacher_transformer,
fake_adv.float(),
timesteps_adv,
encoder_hidden_states.float(),
encoder_attention_mask,
1.0,
)
g_loss += g_gan_loss
g_loss.backward()
g_loss = g_loss.detach().clone()
dist.all_reduce(g_loss, op=dist.ReduceOp.AVG)
g_grad_norm = transformer.clip_grad_norm_(max_grad_norm).item()
optimizer.step()
lr_scheduler.step()
optimizer.zero_grad()
discriminator_optimizer.zero_grad()
discriminator.requires_grad_(True)
with torch.autocast("cuda", dtype=torch.bfloat16):
d_loss = gan_d_loss(
discriminator,
teacher_transformer,
fake_adv.detach(),
real_adv.detach(),
timesteps_adv,
encoder_hidden_states,
encoder_attention_mask,
1.0,
)
d_loss.backward()
d_grad_norm = discriminator.clip_grad_norm_(max_grad_norm).item()
discriminator_optimizer.step()
discriminator_optimizer.zero_grad()
return g_loss, g_grad_norm, d_loss, d_grad_norm
def main(args):
torch.backends.cuda.matmul.allow_tf32 = True
local_rank = int(os.environ["LOCAL_RANK"])
rank = int(os.environ["RANK"])
world_size = int(os.environ["WORLD_SIZE"])
dist.init_process_group("nccl")
torch.cuda.set_device(local_rank)
device = torch.cuda.current_device()
initialize_sequence_parallel_state(args.sp_size)
# If passed along, set the training seed now. On GPU...
if args.seed is not None:
# TODO: t within the same seq parallel group should be the same. Noise should be different.
set_seed(args.seed + rank)
# We use different seeds for the noise generation in each process to ensure that the noise is different in a batch.
noise_random_generator = None
# Handle the repository creation
if rank <= 0 and args.output_dir is not None:
os.makedirs(args.output_dir, exist_ok=True)
# For mixed precision training we cast all non-trainable weigths to half-precision
# as these weights are only used for inference, keeping weights in full precision is not required.
# Create model:
main_print(f"--> loading model from {args.pretrained_model_name_or_path}")
# keep the master weight to float32
if args.dit_model_name_or_path:
transformer = transformer = MochiTransformer3DModel.from_pretrained(
args.dit_model_name_or_path,
torch_dtype=torch.float32,
# torch_dtype=torch.bfloat16 if args.use_lora else torch.float32,
)
else:
transformer = MochiTransformer3DModel.from_pretrained(
args.pretrained_model_name_or_path,
subfolder="transformer",
torch_dtype=torch.float32,
# torch_dtype=torch.bfloat16 if args.use_lora else torch.float32,
)
teacher_transformer = deepcopy(transformer)
discriminator = Discriminator(args.discriminator_head_stride)
if args.use_lora:
transformer.requires_grad_(False)
transformer_lora_config = LoraConfig(
r=args.lora_rank,
lora_alpha=args.lora_alpha,
init_lora_weights=True,
target_modules=["to_k", "to_q", "to_v", "to_out.0"],
)
transformer.add_adapter(transformer_lora_config)
main_print(
f" Total transformer parameters = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e6} M"
)
# discriminator
main_print(
f" Total discriminator parameters = {sum(p.numel() for p in discriminator.parameters() if p.requires_grad) / 1e6} M"
)
main_print(
f"--> Initializing FSDP with sharding strategy: {args.fsdp_sharding_startegy}"
)
fsdp_kwargs = get_dit_fsdp_kwargs(
args.fsdp_sharding_startegy, args.use_lora, args.use_cpu_offload
)
discriminator_fsdp_kwargs = get_discriminator_fsdp_kwargs(args.master_weight_type)
if args.use_lora:
transformer.config.lora_rank = args.lora_rank
transformer.config.lora_alpha = args.lora_alpha
transformer.config.lora_target_modules = ["to_k", "to_q", "to_v", "to_out.0"]
transformer._no_split_modules = ["MochiTransformerBlock"]
fsdp_kwargs["auto_wrap_policy"] = fsdp_kwargs["auto_wrap_policy"](transformer)
transformer = FSDP(
transformer,
**fsdp_kwargs,
)
teacher_transformer = FSDP(
teacher_transformer,
**fsdp_kwargs,
)
discriminator = FSDP(
discriminator,
**discriminator_fsdp_kwargs,
)
main_print(f"--> model loaded")
if args.gradient_checkpointing:
apply_fsdp_checkpointing(transformer, args.selective_checkpointing)
apply_fsdp_checkpointing(teacher_transformer, args.selective_checkpointing)
# Set model as trainable.
transformer.train()
teacher_transformer.requires_grad_(False)
noise_scheduler = FlowMatchEulerDiscreteScheduler(shift=args.shift)
if args.scheduler_type == "pcm_linear_quadratic":
sigmas = linear_quadratic_schedule(
noise_scheduler.config.num_train_timesteps, args.linear_quadratic_threshold
)
sigmas = torch.tensor(sigmas).to(dtype=torch.float32)
else:
sigmas = noise_scheduler.sigmas
solver = EulerSolver(
sigmas.numpy()[::-1],
noise_scheduler.config.num_train_timesteps,
euler_timesteps=args.num_euler_timesteps,
)
solver.to(device)
params_to_optimize = transformer.parameters()
params_to_optimize = list(filter(lambda p: p.requires_grad, params_to_optimize))
optimizer = torch.optim.AdamW(
params_to_optimize,
lr=args.learning_rate,
betas=(0.9, 0.999),
weight_decay=1e-3,
eps=1e-8,
)
discriminator_optimizer = torch.optim.AdamW(
discriminator.parameters(),
lr=args.discriminator_learning_rate,
betas=(0, 0.999),
weight_decay=1e-3,
eps=1e-8,
)
init_steps = 0
if args.resume_from_lora_checkpoint:
transformer, optimizer, init_steps = resume_lora_optimizer(
transformer, args.resume_from_lora_checkpoint, optimizer
)
elif args.resume_from_checkpoint:
(
transformer,
optimizer,
discriminator,
discriminator_optimizer,
init_steps,
) = resume_training_generator_discriminator(
transformer,
optimizer,
discriminator,
discriminator_optimizer,
args.resume_from_checkpoint,
rank,
)
main_print(f"optimizer: {optimizer}")
lr_scheduler = get_scheduler(
args.lr_scheduler,
optimizer=optimizer,
num_warmup_steps=args.lr_warmup_steps * world_size,
num_training_steps=args.max_train_steps * world_size,
num_cycles=args.lr_num_cycles,
power=args.lr_power,
last_epoch=init_steps - 1,
)
train_dataset = LatentDataset(args.data_json_path, args.num_latent_t, args.cfg)
uncond_prompt_embed = train_dataset.uncond_prompt_embed
uncond_prompt_mask = train_dataset.uncond_prompt_mask
sampler = (
LengthGroupedSampler(
args.train_batch_size,
rank=rank,
world_size=world_size,
lengths=train_dataset.lengths,
group_frame=args.group_frame,
group_resolution=args.group_resolution,
)
if (args.group_frame or args.group_resolution)
else DistributedSampler(
train_dataset, rank=rank, num_replicas=world_size, shuffle=False
)
)
train_dataloader = DataLoader(
train_dataset,
sampler=sampler,
collate_fn=latent_collate_function,
pin_memory=True,
batch_size=args.train_batch_size,
num_workers=args.dataloader_num_workers,
drop_last=True,
)
assert args.gradient_accumulation_steps == 1
num_update_steps_per_epoch = math.ceil(
len(train_dataloader)
/ args.gradient_accumulation_steps
* args.sp_size
/ args.train_sp_batch_size
)
args.num_train_epochs = math.ceil(args.max_train_steps / num_update_steps_per_epoch)
if rank <= 0:
project = args.tracker_project_name or "fastvideo"
wandb.init(project=project, config=args)
# Train!
total_batch_size = (
args.train_batch_size
* world_size
* args.gradient_accumulation_steps
/ args.sp_size
* args.train_sp_batch_size
)
main_print("***** Running training *****")
main_print(f" Num examples = {len(train_dataset)}")
main_print(f" Dataloader size = {len(train_dataloader)}")
main_print(f" Num Epochs = {args.num_train_epochs}")
main_print(f" Resume training from step {init_steps}")
main_print(f" Instantaneous batch size per device = {args.train_batch_size}")
main_print(
f" Total train batch size (w. data & sequence parallel, accumulation) = {total_batch_size}"
)
main_print(f" Gradient Accumulation steps = {args.gradient_accumulation_steps}")
main_print(f" Total optimization steps = {args.max_train_steps}")
main_print(
f" Total training parameters per FSDP shard = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e9} B"
)
# print dtype
main_print(f" Master weight dtype: {transformer.parameters().__next__().dtype}")
progress_bar = tqdm(
range(0, args.max_train_steps),
initial=init_steps,
desc="Steps",
# Only show the progress bar once on each machine.
disable=local_rank > 0,
)
loader = sp_parallel_dataloader_wrapper(
train_dataloader,
device,
args.train_batch_size,
args.sp_size,
args.train_sp_batch_size,
)
step_times = deque(maxlen=100)
# log_validation(args, transformer, device,
# torch.bfloat16, init_steps, scheduler_type=args.scheduler_type, shift=args.shift, num_euler_timesteps=args.num_euler_timesteps, linear_quadratic_threshold=args.linear_quadratic_threshold, ema=False)
for i in range(init_steps):
_ = next(loader)
for step in range(init_steps + 1, args.max_train_steps + 1):
start_time = time.time()
(
generator_loss,
generator_grad_norm,
discriminator_loss,
discriminator_grad_norm,
) = train_one_step_mochi(
transformer,
teacher_transformer,
optimizer,
discriminator,
discriminator_optimizer,
step,
lr_scheduler,
loader,
noise_scheduler,
solver,
noise_random_generator,
args.sp_size,
args.precondition_outputs,
args.max_grad_norm,
uncond_prompt_embed,
uncond_prompt_mask,
args.num_euler_timesteps,
args.validation_sampling_steps,
args.not_apply_cfg_solver,
args.distill_cfg,
args.adv_weight,
)
step_time = time.time() - start_time
step_times.append(step_time)
avg_step_time = sum(step_times) / len(step_times)
progress_bar.set_postfix(
{
"g_loss": f"{generator_loss:.4f}",
"d_loss": f"{discriminator_loss:.4f}",
"g_grad_norm": generator_grad_norm,
"d_grad_norm": discriminator_grad_norm,
"step_time": f"{step_time:.2f}s",
}
)
progress_bar.update(1)
if rank <= 0:
wandb.log(
{
"generator_loss": generator_loss,
"discriminator_loss": discriminator_loss,
"generator_grad_norm": generator_grad_norm,
"discriminator_grad_norm": discriminator_grad_norm,
"learning_rate": lr_scheduler.get_last_lr()[0],
"step_time": step_time,
"avg_step_time": avg_step_time,
},
step=step,
)
if step % args.checkpointing_steps == 0:
main_print(f"--> saving checkpoint at step {step}")
if args.use_lora:
# Save LoRA weights
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, step
)
else:
# Your existing checkpoint saving code
save_checkpoint_generator_discriminator(
transformer,
optimizer,
discriminator,
discriminator_optimizer,
rank,
args.output_dir,
step,
)
main_print(f"--> checkpoint saved at step {step}")
dist.barrier()
if args.log_validation and step % args.validation_steps == 0:
log_validation(
args,
transformer,
device,
torch.bfloat16,
step,
scheduler_type=args.scheduler_type,
shift=args.shift,
num_euler_timesteps=args.num_euler_timesteps,
linear_quadratic_threshold=args.linear_quadratic_threshold,
ema=False,
)
if args.use_lora:
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, args.max_train_steps
)
else:
save_checkpoint(
transformer, optimizer, rank, args.output_dir, args.max_train_steps
)
save_checkpoint(
discriminator,
discriminator_optimizer,
rank,
args.output_dir,
step,
discriminator=True,
)
if get_sequence_parallel_state():
destroy_sequence_parallel_group()
if __name__ == "__main__":
parser = argparse.ArgumentParser()
# dataset & dataloader
parser.add_argument("--data_json_path", type=str, required=True)
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument(
"--dataloader_num_workers",
type=int,
default=10,
help="Number of subprocesses to use for data loading. 0 means that the data will be loaded in the main process.",
)
parser.add_argument(
"--train_batch_size",
type=int,
default=16,
help="Batch size (per device) for the training dataloader.",
)
parser.add_argument(
"--num_latent_t", type=int, default=28, help="Number of latent timesteps."
)
parser.add_argument("--group_frame", action="store_true") # TODO
parser.add_argument("--group_resolution", action="store_true") # TODO
# text encoder & vae & diffusion model
parser.add_argument("--pretrained_model_name_or_path", type=str)
parser.add_argument("--dit_model_name_or_path", type=str)
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
# diffusion setting
parser.add_argument("--ema_decay", type=float, default=0.999)
parser.add_argument("--ema_start_step", type=int, default=0)
parser.add_argument("--cfg", type=float, default=0.1)
parser.add_argument(
"--precondition_outputs",
action="store_true",
help="Whether to precondition the outputs of the model.",
)
# validation & logs
parser.add_argument("--validation_prompt_dir", type=str)
parser.add_argument("--validation_sampling_steps", type=int, default=64)
parser.add_argument("--validation_guidance_scale", type=float, default=4.5)
parser.add_argument("--validation_steps", type=float, default=64)
parser.add_argument("--log_validation", action="store_true")
parser.add_argument("--tracker_project_name", type=str, default=None)
parser.add_argument(
"--seed", type=int, default=None, help="A seed for reproducible training."
)
parser.add_argument(
"--output_dir",
type=str,
default=None,
help="The output directory where the model predictions and checkpoints will be written.",
)
parser.add_argument(
"--checkpoints_total_limit",
type=int,
default=None,
help=("Max number of checkpoints to store."),
)
parser.add_argument(
"--checkpointing_steps",
type=int,
default=500,
help=(
"Save a checkpoint of the training state every X updates. These checkpoints can be used both as final"
" checkpoints in case they are better than the last checkpoint, and are also suitable for resuming"
" training using `--resume_from_checkpoint`."
),
)
parser.add_argument("--shift", type=float, default=1.0)
parser.add_argument(
"--resume_from_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--resume_from_lora_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous lora checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--logging_dir",
type=str,
default="logs",
help=(
"[TensorBoard](https://www.tensorflow.org/tensorboard) log directory. Will default to"
" *output_dir/runs/**CURRENT_DATETIME_HOSTNAME***."
),
)
# optimizer & scheduler & Training
parser.add_argument("--num_train_epochs", type=int, default=100)
parser.add_argument(
"--max_train_steps",
type=int,
default=None,
help="Total number of training steps to perform. If provided, overrides num_train_epochs.",
)
parser.add_argument(
"--learning_rate",
type=float,
default=1e-4,
help="Initial learning rate (after the potential warmup period) to use.",
)
parser.add_argument(
"--discriminator_learning_rate",
type=float,
default=1e-5,
help="Initial learning rate (after the potential warmup period) to use.",
)
parser.add_argument(
"--scale_lr",
action="store_true",
default=False,
help="Scale the learning rate by the number of GPUs, gradient accumulation steps, and batch size.",
)
parser.add_argument(
"--lr_warmup_steps",
type=int,
default=10,
help="Number of steps for the warmup in the lr scheduler.",
)
parser.add_argument(
"--max_grad_norm", default=1.0, type=float, help="Max gradient norm."
)
parser.add_argument(
"--gradient_checkpointing",
action="store_true",
help="Whether or not to use gradient checkpointing to save memory at the expense of slower backward pass.",
)
parser.add_argument("--selective_checkpointing", type=float, default=1.0)
parser.add_argument(
"--allow_tf32",
action="store_true",
help=(
"Whether or not to allow TF32 on Ampere GPUs. Can be used to speed up training. For more information, see"
" https://pytorch.org/docs/stable/notes/cuda.html#tensorfloat-32-tf32-on-ampere-devices"
),
)
parser.add_argument(
"--mixed_precision",
type=str,
default=None,
choices=["no", "fp16", "bf16"],
help=(
"Whether to use mixed precision. Choose between fp16 and bf16 (bfloat16). Bf16 requires PyTorch >="
" 1.10.and an Nvidia Ampere GPU. Default to the value of accelerate config of the current system or the"
" flag passed with the `accelerate.launch` command. Use this argument to override the accelerate config."
),
)
parser.add_argument(
"--use_cpu_offload",
action="store_true",
help="Whether to use CPU offload for param & gradient & optimizer states.",
)
parser.add_argument("--sp_size", type=int, default=1, help="For sequence parallel")
parser.add_argument(
"--train_sp_batch_size",
type=int,
default=1,
help="Batch size for sequence parallel training",
)
parser.add_argument(
"--use_lora",
action="store_true",
default=False,
help="Whether to use LoRA for finetuning.",
)
parser.add_argument(
"--lora_alpha", type=int, default=256, help="Alpha parameter for LoRA."
)
parser.add_argument(
"--lora_rank", type=int, default=128, help="LoRA rank parameter. "
)
parser.add_argument("--fsdp_sharding_startegy", default="full")
parser.add_argument(
"--gradient_accumulation_steps",
type=int,
default=1,
help="Number of updates steps to accumulate before performing a backward/update pass.",
)
# lr_scheduler
parser.add_argument(
"--lr_scheduler",
type=str,
default="constant",
help=(
'The scheduler type to use. Choose between ["linear", "cosine", "cosine_with_restarts", "polynomial",'
' "constant", "constant_with_warmup"]'
),
)
parser.add_argument("--num_euler_timesteps", type=int, default=100)
parser.add_argument(
"--lr_num_cycles",
type=int,
default=1,
help="Number of cycles in the learning rate scheduler.",
)
parser.add_argument(
"--lr_power",
type=float,
default=1.0,
help="Power factor of the polynomial scheduler.",
)
parser.add_argument(
"--not_apply_cfg_solver",
action="store_true",
help="Whether to apply the cfg_solver.",
)
parser.add_argument(
"--distill_cfg", type=float, default=3.0, help="Distillation coefficient."
)
# ["euler_linear_quadratic", "pcm", "pcm_linear_qudratic"]
parser.add_argument(
"--scheduler_type", type=str, default="pcm", help="The scheduler type to use."
)
parser.add_argument(
"--adv_weight",
type=float,
default=0.1,
help="The weight of the adversarial loss.",
)
parser.add_argument(
"--discriminator_head_stride",
type=int,
default=2,
help="The stride of the discriminator head.",
)
parser.add_argument(
"--linear_quadratic_threshold",
type=float,
default=0.025,
help="The threshold of the linear quadratic scheduler.",
)
parser.add_argument(
"--master_weight_type",
type=str,
default="fp32",
help="Weight type to use - fp32 or bf16.",
)
args = parser.parse_args()
main(args)
-149
View File
@@ -1,149 +0,0 @@
from sympy import use
import torch
import os
import torch.distributed as dist
from torch.distributed.algorithms._checkpoint.checkpoint_wrapper import (
checkpoint_wrapper,
CheckpointImpl,
apply_activation_checkpointing,
)
from peft.utils.other import fsdp_auto_wrap_policy
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
StateDictType,
FullStateDictConfig, # general model non-sharded, non-flattened params
LocalStateDictConfig, # flattened params, usable only by FSDP
# ShardedStateDictConfig, # un-flattened param but shards, usable by other parallel schemes.
)
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformerBlock
from functools import partial
from torch.distributed.fsdp.wrap import transformer_auto_wrap_policy
from torch.distributed.fsdp import MixedPrecision, ShardingStrategy
import functools
non_reentrant_wrapper = partial(
checkpoint_wrapper,
checkpoint_impl=CheckpointImpl.NO_REENTRANT,
)
check_fn = lambda submodule: isinstance(submodule, MochiTransformerBlock)
def apply_fsdp_checkpointing(model, p=1):
# https://github.com/foundation-model-stack/fms-fsdp/blob/408c7516d69ea9b6bcd4c0f5efab26c0f64b3c2d/fms_fsdp/policies/ac_handler.py#L16
"""apply activation checkpointing to model
returns None as model is updated directly
"""
print(f"--> applying fdsp activation checkpointing...")
block_idx = 0
cut_off = 1 / 2
# when passing p as a fraction number (e.g. 1/3), it will be interpreted
# as a string in argv, thus we need eval("1/3") here for fractions.
p = eval(p) if isinstance(p, str) else p
def selective_checkpointing(submodule):
nonlocal block_idx
nonlocal cut_off
if isinstance(submodule, MochiTransformerBlock):
block_idx += 1
if block_idx * p >= cut_off:
cut_off += 1
return True
return False
apply_activation_checkpointing(
model,
checkpoint_wrapper_fn=non_reentrant_wrapper,
check_fn=selective_checkpointing,
)
def get_mixed_precision(master_weight_type="fp32"):
weight_type = torch.float32 if master_weight_type == "fp32" else torch.bfloat16
mixed_precision = MixedPrecision(
param_dtype=weight_type,
# Gradient communication precision.
reduce_dtype=weight_type,
# Buffer precision.
buffer_dtype=weight_type,
cast_forward_inputs=False,
)
return mixed_precision
def get_dit_fsdp_kwargs(
sharding_strategy, use_lora=False, cpu_offload=False, master_weight_type="fp32"
):
if use_lora:
auto_wrap_policy = fsdp_auto_wrap_policy
else:
auto_wrap_policy = functools.partial(
transformer_auto_wrap_policy,
transformer_layer_cls={
MochiTransformerBlock,
},
)
# we use float32 for fsdp but autocast during training
mixed_precision = get_mixed_precision(master_weight_type)
if sharding_strategy == "full":
sharding_strategy = ShardingStrategy.FULL_SHARD
elif sharding_strategy == "hybrid_full":
sharding_strategy = ShardingStrategy.HYBRID_SHARD
elif sharding_strategy == "none":
sharding_strategy = ShardingStrategy.NO_SHARD
auto_wrap_policy = None
elif sharding_strategy == "hybrid_zero2":
sharding_strategy = ShardingStrategy._HYBRID_SHARD_ZERO2
device_id = torch.cuda.current_device()
cpu_offload = (
torch.distributed.fsdp.CPUOffload(offload_params=True) if cpu_offload else None
)
fsdp_kwargs = {
"auto_wrap_policy": auto_wrap_policy,
"mixed_precision": mixed_precision,
"sharding_strategy": sharding_strategy,
"device_id": device_id,
"limit_all_gathers": True,
"cpu_offload": cpu_offload,
}
# Add LoRA-specific settings when LoRA is enabled
if use_lora:
fsdp_kwargs.update(
{
"use_orig_params": False, # Required for LoRA memory savings
"sync_module_states": True,
}
)
return fsdp_kwargs
def get_discriminator_fsdp_kwargs(master_weight_type="fp32"):
auto_wrap_policy = None
# Use existing mixed precision settings
mixed_precision = get_mixed_precision(master_weight_type)
sharding_strategy = ShardingStrategy.NO_SHARD
device_id = torch.cuda.current_device()
fsdp_kwargs = {
"auto_wrap_policy": auto_wrap_policy,
"mixed_precision": mixed_precision,
"sharding_strategy": sharding_strategy,
"device_id": device_id,
"limit_all_gathers": True,
}
return fsdp_kwargs
@@ -1,431 +0,0 @@
import torch
import argparse
from safetensors.torch import save_file
import os
parser = argparse.ArgumentParser()
parser.add_argument("--diffusers_path", required=True, type=str)
parser.add_argument("--transformer_path", type=str, default=None, help="Path to save transformer model")
parser.add_argument("--vae_encoder_path", type=str, default=None, help="Path to save VAE encoder model")
parser.add_argument("--vae_decoder_path", type=str, default=None, help="Path to save VAE decoder model")
args = parser.parse_args()
def reverse_scale_shift(weight, dim):
scale, shift = weight.chunk(2, dim=0)
new_weight = torch.cat([shift, scale], dim=0)
return new_weight
def reverse_proj_gate(weight):
gate, proj = weight.chunk(2, dim=0)
new_weight = torch.cat([proj, gate], dim=0)
return new_weight
def convert_diffusers_transformer_to_mochi(state_dict):
original_state_dict = state_dict.copy()
new_state_dict = {}
# Convert patch_embed
new_state_dict["x_embedder.proj.weight"] = original_state_dict.pop("patch_embed.proj.weight")
new_state_dict["x_embedder.proj.bias"] = original_state_dict.pop("patch_embed.proj.bias")
# Convert time_embed
new_state_dict["t_embedder.mlp.0.weight"] = original_state_dict.pop("time_embed.timestep_embedder.linear_1.weight")
new_state_dict["t_embedder.mlp.0.bias"] = original_state_dict.pop("time_embed.timestep_embedder.linear_1.bias")
new_state_dict["t_embedder.mlp.2.weight"] = original_state_dict.pop("time_embed.timestep_embedder.linear_2.weight")
new_state_dict["t_embedder.mlp.2.bias"] = original_state_dict.pop("time_embed.timestep_embedder.linear_2.bias")
new_state_dict["t5_y_embedder.to_kv.weight"] = original_state_dict.pop("time_embed.pooler.to_kv.weight")
new_state_dict["t5_y_embedder.to_kv.bias"] = original_state_dict.pop("time_embed.pooler.to_kv.bias")
new_state_dict["t5_y_embedder.to_q.weight"] = original_state_dict.pop("time_embed.pooler.to_q.weight")
new_state_dict["t5_y_embedder.to_q.bias"] = original_state_dict.pop("time_embed.pooler.to_q.bias")
new_state_dict["t5_y_embedder.to_out.weight"] = original_state_dict.pop("time_embed.pooler.to_out.weight")
new_state_dict["t5_y_embedder.to_out.bias"] = original_state_dict.pop("time_embed.pooler.to_out.bias")
new_state_dict["t5_yproj.weight"] = original_state_dict.pop("time_embed.caption_proj.weight")
new_state_dict["t5_yproj.bias"] = original_state_dict.pop("time_embed.caption_proj.bias")
# Convert transformer blocks
num_layers = 48
for i in range(num_layers):
block_prefix = f"transformer_blocks.{i}."
new_prefix = f"blocks.{i}."
# norm1
new_state_dict[new_prefix + "mod_x.weight"] = original_state_dict.pop(block_prefix + "norm1.linear.weight")
new_state_dict[new_prefix + "mod_x.bias"] = original_state_dict.pop(block_prefix + "norm1.linear.bias")
if i < num_layers - 1:
new_state_dict[new_prefix + "mod_y.weight"] = original_state_dict.pop(
block_prefix + "norm1_context.linear.weight"
)
new_state_dict[new_prefix + "mod_y.bias"] = original_state_dict.pop(
block_prefix + "norm1_context.linear.bias"
)
else:
new_state_dict[new_prefix + "mod_y.weight"] = original_state_dict.pop(
block_prefix + "norm1_context.linear_1.weight"
)
new_state_dict[new_prefix + "mod_y.bias"] = original_state_dict.pop(
block_prefix + "norm1_context.linear_1.bias"
)
# Visual attention
q = original_state_dict.pop(block_prefix + "attn1.to_q.weight")
k = original_state_dict.pop(block_prefix + "attn1.to_k.weight")
v = original_state_dict.pop(block_prefix + "attn1.to_v.weight")
qkv_weight = torch.cat([q, k, v], dim=0)
new_state_dict[new_prefix + "attn.qkv_x.weight"] = qkv_weight
new_state_dict[new_prefix + "attn.q_norm_x.weight"] = original_state_dict.pop(
block_prefix + "attn1.norm_q.weight"
)
new_state_dict[new_prefix + "attn.k_norm_x.weight"] = original_state_dict.pop(
block_prefix + "attn1.norm_k.weight"
)
new_state_dict[new_prefix + "attn.proj_x.weight"] = original_state_dict.pop(
block_prefix + "attn1.to_out.0.weight"
)
new_state_dict[new_prefix + "attn.proj_x.bias"] = original_state_dict.pop(
block_prefix + "attn1.to_out.0.bias"
)
# Context attention
q = original_state_dict.pop(block_prefix + "attn1.add_q_proj.weight")
k = original_state_dict.pop(block_prefix + "attn1.add_k_proj.weight")
v = original_state_dict.pop(block_prefix + "attn1.add_v_proj.weight")
qkv_weight = torch.cat([q, k, v], dim=0)
new_state_dict[new_prefix + "attn.qkv_y.weight"] = qkv_weight
new_state_dict[new_prefix + "attn.q_norm_y.weight"] = original_state_dict.pop(
block_prefix + "attn1.norm_added_q.weight"
)
new_state_dict[new_prefix + "attn.k_norm_y.weight"] = original_state_dict.pop(
block_prefix + "attn1.norm_added_k.weight"
)
if i < num_layers - 1:
new_state_dict[new_prefix + "attn.proj_y.weight"] = original_state_dict.pop(
block_prefix + "attn1.to_add_out.weight"
)
new_state_dict[new_prefix + "attn.proj_y.bias"] = original_state_dict.pop(
block_prefix + "attn1.to_add_out.bias"
)
# MLP
new_state_dict[new_prefix + "mlp_x.w1.weight"] = reverse_proj_gate(
original_state_dict.pop(block_prefix + "ff.net.0.proj.weight")
)
new_state_dict[new_prefix + "mlp_x.w2.weight"] = original_state_dict.pop(block_prefix + "ff.net.2.weight")
if i < num_layers - 1:
new_state_dict[new_prefix + "mlp_y.w1.weight"] = reverse_proj_gate(
original_state_dict.pop(block_prefix + "ff_context.net.0.proj.weight")
)
new_state_dict[new_prefix + "mlp_y.w2.weight"] = original_state_dict.pop(
block_prefix + "ff_context.net.2.weight"
)
# Output layers
new_state_dict["final_layer.mod.weight"] = reverse_scale_shift(
original_state_dict.pop("norm_out.linear.weight"), dim=0
)
new_state_dict["final_layer.mod.bias"] = reverse_scale_shift(
original_state_dict.pop("norm_out.linear.bias"), dim=0
)
new_state_dict["final_layer.linear.weight"] = original_state_dict.pop("proj_out.weight")
new_state_dict["final_layer.linear.bias"] = original_state_dict.pop("proj_out.bias")
new_state_dict["pos_frequencies"] = original_state_dict.pop("pos_frequencies")
print("Remaining Keys:", original_state_dict.keys())
return new_state_dict
def convert_diffusers_vae_to_mochi(state_dict):
original_state_dict = state_dict.copy()
encoder_state_dict = {}
decoder_state_dict = {}
# Convert encoder
prefix = "encoder."
encoder_state_dict["layers.0.weight"] = original_state_dict.pop(f"{prefix}proj_in.weight")
encoder_state_dict["layers.0.bias"] = original_state_dict.pop(f"{prefix}proj_in.bias")
# Convert block_in
for i in range(3):
encoder_state_dict[f"layers.{i+1}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm1.norm_layer.weight"
)
encoder_state_dict[f"layers.{i+1}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm1.norm_layer.bias"
)
encoder_state_dict[f"layers.{i+1}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv1.conv.weight"
)
encoder_state_dict[f"layers.{i+1}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv1.conv.bias"
)
encoder_state_dict[f"layers.{i+1}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm2.norm_layer.weight"
)
encoder_state_dict[f"layers.{i+1}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm2.norm_layer.bias"
)
encoder_state_dict[f"layers.{i+1}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv2.conv.weight"
)
encoder_state_dict[f"layers.{i+1}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv2.conv.bias"
)
# Convert down_blocks
down_block_layers = [3, 4, 6]
for block in range(3):
encoder_state_dict[f"layers.{block+4}.layers.0.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.conv_in.conv.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.0.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.conv_in.conv.bias"
)
for i in range(down_block_layers[block]):
# Convert resnets
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.conv1.conv.bias"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.resnets.{i}.conv2.conv.bias"
)
# Convert attentions
q = original_state_dict.pop(f"{prefix}down_blocks.{block}.attentions.{i}.to_q.weight")
k = original_state_dict.pop(f"{prefix}down_blocks.{block}.attentions.{i}.to_k.weight")
v = original_state_dict.pop(f"{prefix}down_blocks.{block}.attentions.{i}.to_v.weight")
qkv_weight = torch.cat([q, k, v], dim=0)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.attn_block.attn.qkv.weight"] = qkv_weight
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.attn_block.attn.out.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.attentions.{i}.to_out.0.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.attn_block.attn.out.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.attentions.{i}.to_out.0.bias"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.attn_block.norm.weight"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.norms.{i}.norm_layer.weight"
)
encoder_state_dict[f"layers.{block+4}.layers.{i+1}.attn_block.norm.bias"] = original_state_dict.pop(
f"{prefix}down_blocks.{block}.norms.{i}.norm_layer.bias"
)
# Convert block_out
for i in range(3):
encoder_state_dict[f"layers.{i+7}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm1.norm_layer.weight"
)
encoder_state_dict[f"layers.{i+7}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm1.norm_layer.bias"
)
encoder_state_dict[f"layers.{i+7}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv1.conv.weight"
)
encoder_state_dict[f"layers.{i+7}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv1.conv.bias"
)
encoder_state_dict[f"layers.{i+7}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm2.norm_layer.weight"
)
encoder_state_dict[f"layers.{i+7}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm2.norm_layer.bias"
)
encoder_state_dict[f"layers.{i+7}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv2.conv.weight"
)
encoder_state_dict[f"layers.{i+7}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv2.conv.bias"
)
q = original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_q.weight")
k = original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_k.weight")
v = original_state_dict.pop(f"{prefix}block_out.attentions.{i}.to_v.weight")
qkv_weight = torch.cat([q, k, v], dim=0)
encoder_state_dict[f"layers.{i+7}.attn_block.attn.qkv.weight"] = qkv_weight
encoder_state_dict[f"layers.{i+7}.attn_block.attn.out.weight"] = original_state_dict.pop(
f"{prefix}block_out.attentions.{i}.to_out.0.weight"
)
encoder_state_dict[f"layers.{i+7}.attn_block.attn.out.bias"] = original_state_dict.pop(
f"{prefix}block_out.attentions.{i}.to_out.0.bias"
)
encoder_state_dict[f"layers.{i+7}.attn_block.norm.weight"] = original_state_dict.pop(
f"{prefix}block_out.norms.{i}.norm_layer.weight"
)
encoder_state_dict[f"layers.{i+7}.attn_block.norm.bias"] = original_state_dict.pop(
f"{prefix}block_out.norms.{i}.norm_layer.bias"
)
# Convert output layers
encoder_state_dict["output_norm.weight"] = original_state_dict.pop(f"{prefix}norm_out.norm_layer.weight")
encoder_state_dict["output_norm.bias"] = original_state_dict.pop(f"{prefix}norm_out.norm_layer.bias")
encoder_state_dict["output_proj.weight"] = original_state_dict.pop(f"{prefix}proj_out.weight")
# Convert decoder
prefix = "decoder."
decoder_state_dict["blocks.0.0.weight"] = original_state_dict.pop(f"{prefix}conv_in.weight")
decoder_state_dict["blocks.0.0.bias"] = original_state_dict.pop(f"{prefix}conv_in.bias")
# Convert block_in
for i in range(3):
decoder_state_dict[f"blocks.0.{i+1}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm1.norm_layer.weight"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm1.norm_layer.bias"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv1.conv.weight"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv1.conv.bias"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm2.norm_layer.weight"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.norm2.norm_layer.bias"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv2.conv.weight"
)
decoder_state_dict[f"blocks.0.{i+1}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}block_in.resnets.{i}.conv2.conv.bias"
)
# Convert up_blocks
up_block_layers = [6, 4, 3]
for block in range(3):
for i in range(up_block_layers[block]):
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.weight"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.norm1.norm_layer.bias"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.weight"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.conv1.conv.bias"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.weight"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.norm2.norm_layer.bias"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.weight"
)
decoder_state_dict[f"blocks.{block+1}.blocks.{i}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.resnets.{i}.conv2.conv.bias"
)
decoder_state_dict[f"blocks.{block+1}.proj.weight"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.proj.weight"
)
decoder_state_dict[f"blocks.{block+1}.proj.bias"] = original_state_dict.pop(
f"{prefix}up_blocks.{block}.proj.bias"
)
# Convert block_out
for i in range(3):
decoder_state_dict[f"blocks.4.{i}.stack.0.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm1.norm_layer.weight"
)
decoder_state_dict[f"blocks.4.{i}.stack.0.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm1.norm_layer.bias"
)
decoder_state_dict[f"blocks.4.{i}.stack.2.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv1.conv.weight"
)
decoder_state_dict[f"blocks.4.{i}.stack.2.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv1.conv.bias"
)
decoder_state_dict[f"blocks.4.{i}.stack.3.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm2.norm_layer.weight"
)
decoder_state_dict[f"blocks.4.{i}.stack.3.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.norm2.norm_layer.bias"
)
decoder_state_dict[f"blocks.4.{i}.stack.5.weight"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv2.conv.weight"
)
decoder_state_dict[f"blocks.4.{i}.stack.5.bias"] = original_state_dict.pop(
f"{prefix}block_out.resnets.{i}.conv2.conv.bias"
)
# Convert output layers
decoder_state_dict["output_proj.weight"] = original_state_dict.pop(f"{prefix}proj_out.weight")
decoder_state_dict["output_proj.bias"] = original_state_dict.pop(f"{prefix}proj_out.bias")
return encoder_state_dict, decoder_state_dict
def ensure_safetensors_extension(path):
if not path.endswith('.safetensors'):
path = path + '.safetensors'
return path
def ensure_directory_exists(path):
directory = os.path.dirname(path)
if directory:
os.makedirs(directory, exist_ok=True)
def main(args):
from diffusers import MochiPipeline
pipe = MochiPipeline.from_pretrained(args.diffusers_path)
if args.transformer_path:
transformer_path = ensure_safetensors_extension(args.transformer_path)
ensure_directory_exists(transformer_path)
print(f"Converting transformer model...")
transformer_state_dict = convert_diffusers_transformer_to_mochi(pipe.transformer.state_dict())
save_file(transformer_state_dict, transformer_path)
print(f"Saved transformer to {transformer_path}")
if args.vae_encoder_path and args.vae_decoder_path:
encoder_path = ensure_safetensors_extension(args.vae_encoder_path)
decoder_path = ensure_safetensors_extension(args.vae_decoder_path)
ensure_directory_exists(encoder_path)
ensure_directory_exists(decoder_path)
print(f"Converting VAE models...")
encoder_state_dict, decoder_state_dict = convert_diffusers_vae_to_mochi(pipe.vae.state_dict())
save_file(encoder_state_dict, encoder_path)
print(f"Saved VAE encoder to {encoder_path}")
save_file(decoder_state_dict, decoder_path)
print(f"Saved VAE decoder to {decoder_path}")
elif args.vae_encoder_path or args.vae_decoder_path:
print("Warning: Both VAE encoder and decoder paths must be specified to convert VAE models.")
if __name__ == "__main__":
main(args)
@@ -1,42 +0,0 @@
import torch
mochi_latents_mean = torch.tensor(
[
-0.06730895953510081,
-0.038011381506090416,
-0.07477820912866141,
-0.05565264470995561,
0.012767231469026969,
-0.04703542746246419,
0.043896967884726704,
-0.09346305707025976,
-0.09918314763016893,
-0.008729793427399178,
-0.011931556316503654,
-0.0321993391887285,
]
).view(1, 12, 1, 1, 1)
mochi_latents_std = torch.tensor(
[
0.9263795028493863,
0.9248894543193766,
0.9393059390890617,
0.959253732819592,
0.8244560132752793,
0.917259975397747,
0.9294154431013696,
1.3720942357788521,
0.881393668867029,
0.9168315692124348,
0.9185249279345552,
0.9274757570805041,
]
).view(1, 12, 1, 1, 1)
mochi_scaling_factor = 1.0
def normalize_mochi_dit_input(latents):
latents_mean = mochi_latents_mean.to(latents.device, latents.dtype)
latents_std = mochi_latents_std.to(latents.device, latents.dtype)
latents = (latents - latents_mean) / latents_std
return latents
-766
View File
@@ -1,766 +0,0 @@
# Copyright 2024 The Genmo team and The HuggingFace Team.
# All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
from typing import Any, Dict, Optional, Tuple
import torch
import torch.nn as nn
import diffusers
from diffusers.configuration_utils import ConfigMixin, register_to_config
from diffusers.utils import is_torch_version, logging
from diffusers.utils import (
USE_PEFT_BACKEND,
is_torch_version,
logging,
scale_lora_layers,
unscale_lora_layers,
)
from diffusers.utils.torch_utils import maybe_allow_in_graph
from diffusers.models.attention import FeedForward as HF_FeedForward
from diffusers.models.attention_processor import Attention
from diffusers.models.embeddings import (
MochiCombinedTimestepCaptionEmbedding,
PatchEmbed,
)
from diffusers.models.modeling_outputs import Transformer2DModelOutput
from diffusers.models.modeling_utils import ModelMixin
from diffusers.loaders import PeftAdapterMixin
from fastvideo.models.mochi_hf.norm import (
MochiLayerNormContinuous,
MochiRMSNormZero,
MochiModulatedRMSNorm,
MochiRMSNorm,
)
from diffusers.models.normalization import AdaLayerNormContinuous
from fastvideo.utils.parallel_states import get_sequence_parallel_state, nccl_info
from fastvideo.utils.communications import all_gather, all_to_all_4D
import torch.nn.functional as F
from diffusers.utils.torch_utils import is_torch_version, maybe_allow_in_graph
from einops import rearrange
import numbers
from flash_attn import flash_attn_varlen_qkvpacked_func
from flash_attn.bert_padding import pad_input, unpad_input
from liger_kernel.ops.swiglu import LigerSiLUMulFunction
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
class FeedForward(HF_FeedForward):
def __init__(
self,
dim: int,
dim_out: Optional[int] = None,
mult: int = 4,
dropout: float = 0.0,
activation_fn: str = "geglu",
final_dropout: bool = False,
inner_dim=None,
bias: bool = True,
):
super().__init__(
dim, dim_out, mult, dropout, activation_fn, final_dropout, inner_dim, bias
)
assert activation_fn == "swiglu"
def forward(self, hidden_states: torch.Tensor) -> torch.Tensor:
hidden_states = self.net[0].proj(hidden_states)
hidden_states, gate = hidden_states.chunk(2, dim=-1)
return self.net[2](LigerSiLUMulFunction.apply(gate, hidden_states))
def flash_attn_no_pad(
qkv, key_padding_mask, causal=False, dropout_p=0.0, softmax_scale=None
):
# adapted from https://github.com/Dao-AILab/flash-attention/blob/13403e81157ba37ca525890f2f0f2137edf75311/flash_attn/flash_attention.py#L27
batch_size = qkv.shape[0]
seqlen = qkv.shape[1]
nheads = qkv.shape[-2]
x = rearrange(qkv, "b s three h d -> b s (three h d)")
x_unpad, indices, cu_seqlens, max_s, used_seqlens_in_batch = unpad_input(
x, key_padding_mask
)
x_unpad = rearrange(x_unpad, "nnz (three h d) -> nnz three h d", three=3, h=nheads)
output_unpad = flash_attn_varlen_qkvpacked_func(
x_unpad,
cu_seqlens,
max_s,
dropout_p,
softmax_scale=softmax_scale,
causal=causal,
)
output = rearrange(
pad_input(
rearrange(output_unpad, "nnz h d -> nnz (h d)"), indices, batch_size, seqlen
),
"b s (h d) -> b s h d",
h=nheads,
)
return output
class MochiAttention(nn.Module):
def __init__(
self,
query_dim: int,
processor: "MochiAttnProcessor2_0",
heads: int = 8,
dim_head: int = 64,
dropout: float = 0.0,
bias: bool = False,
added_kv_proj_dim: Optional[int] = None,
added_proj_bias: Optional[bool] = True,
out_dim: int = None,
out_context_dim: int = None,
out_bias: bool = True,
context_pre_only: bool = False,
eps: float = 1e-5,
):
super().__init__()
self.inner_dim = out_dim if out_dim is not None else dim_head * heads
self.out_dim = out_dim if out_dim is not None else query_dim
self.out_context_dim = out_context_dim if out_context_dim else query_dim
self.context_pre_only = context_pre_only
self.heads = out_dim // dim_head if out_dim is not None else heads
self.norm_q = MochiRMSNorm(dim_head, eps)
self.norm_k = MochiRMSNorm(dim_head, eps)
self.norm_added_q = MochiRMSNorm(dim_head, eps)
self.norm_added_k = MochiRMSNorm(dim_head, eps)
self.to_q = nn.Linear(query_dim, self.inner_dim, bias=bias)
self.to_k = nn.Linear(query_dim, self.inner_dim, bias=bias)
self.to_v = nn.Linear(query_dim, self.inner_dim, bias=bias)
self.add_k_proj = nn.Linear(
added_kv_proj_dim, self.inner_dim, bias=added_proj_bias
)
self.add_v_proj = nn.Linear(
added_kv_proj_dim, self.inner_dim, bias=added_proj_bias
)
if self.context_pre_only is not None:
self.add_q_proj = nn.Linear(
added_kv_proj_dim, self.inner_dim, bias=added_proj_bias
)
self.to_out = nn.ModuleList([])
self.to_out.append(nn.Linear(self.inner_dim, self.out_dim, bias=out_bias))
self.to_out.append(nn.Dropout(dropout))
if not self.context_pre_only:
self.to_add_out = nn.Linear(
self.inner_dim, self.out_context_dim, bias=out_bias
)
self.processor = processor
def forward(
self,
hidden_states: torch.Tensor,
encoder_hidden_states: Optional[torch.Tensor] = None,
attention_mask: Optional[torch.Tensor] = None,
**kwargs,
):
return self.processor(
self,
hidden_states,
encoder_hidden_states=encoder_hidden_states,
attention_mask=attention_mask,
**kwargs,
)
class MochiAttnProcessor2_0:
"""Attention processor used in Mochi."""
def __init__(self):
if not hasattr(F, "scaled_dot_product_attention"):
raise ImportError(
"MochiAttnProcessor2_0 requires PyTorch 2.0. To use it, please upgrade PyTorch to 2.0."
)
def __call__(
self,
attn: Attention,
hidden_states: torch.Tensor,
encoder_hidden_states: torch.Tensor,
encoder_attention_mask: torch.Tensor,
attention_mask: Optional[torch.Tensor] = None,
image_rotary_emb: Optional[torch.Tensor] = None,
) -> torch.Tensor:
# [b, s, h * d]
query = attn.to_q(hidden_states)
key = attn.to_k(hidden_states)
value = attn.to_v(hidden_states)
# [b, s, h=24, d=128]
query = query.unflatten(2, (attn.heads, -1))
key = key.unflatten(2, (attn.heads, -1))
value = value.unflatten(2, (attn.heads, -1))
if attn.norm_q is not None:
query = attn.norm_q(query)
if attn.norm_k is not None:
key = attn.norm_k(key)
# [b, 256, h * d]
encoder_query = attn.add_q_proj(encoder_hidden_states)
encoder_key = attn.add_k_proj(encoder_hidden_states)
encoder_value = attn.add_v_proj(encoder_hidden_states)
# [b, 256, h=24, d=128]
encoder_query = encoder_query.unflatten(2, (attn.heads, -1))
encoder_key = encoder_key.unflatten(2, (attn.heads, -1))
encoder_value = encoder_value.unflatten(2, (attn.heads, -1))
if attn.norm_added_q is not None:
encoder_query = attn.norm_added_q(encoder_query)
if attn.norm_added_k is not None:
encoder_key = attn.norm_added_k(encoder_key)
if image_rotary_emb is not None:
freqs_cos, freqs_sin = image_rotary_emb[0], image_rotary_emb[1]
# shard the head dimension
if get_sequence_parallel_state():
# B, S, H, D to (S, B,) H, D
# batch_size, seq_len, attn_heads, head_dim
query = all_to_all_4D(query, scatter_dim=2, gather_dim=1)
key = all_to_all_4D(key, scatter_dim=2, gather_dim=1)
value = all_to_all_4D(value, scatter_dim=2, gather_dim=1)
def shrink_head(encoder_state, dim):
local_heads = encoder_state.shape[dim] // nccl_info.sp_size
return encoder_state.narrow(
dim, nccl_info.rank_within_group * local_heads, local_heads
)
encoder_query = shrink_head(encoder_query, dim=2)
encoder_key = shrink_head(encoder_key, dim=2)
encoder_value = shrink_head(encoder_value, dim=2)
if image_rotary_emb is not None:
freqs_cos = shrink_head(freqs_cos, dim=1)
freqs_sin = shrink_head(freqs_sin, dim=1)
if image_rotary_emb is not None:
def apply_rotary_emb(x, freqs_cos, freqs_sin):
x_even = x[..., 0::2].float()
x_odd = x[..., 1::2].float()
cos = (x_even * freqs_cos - x_odd * freqs_sin).to(x.dtype)
sin = (x_even * freqs_sin + x_odd * freqs_cos).to(x.dtype)
return torch.stack([cos, sin], dim=-1).flatten(-2)
query = apply_rotary_emb(query, freqs_cos, freqs_sin)
key = apply_rotary_emb(key, freqs_cos, freqs_sin)
# query, key, value = query.transpose(1, 2), key.transpose(1, 2), value.transpose(1, 2)
# encoder_query, encoder_key, encoder_value = (
# encoder_query.transpose(1, 2),
# encoder_key.transpose(1, 2),
# encoder_value.transpose(1, 2),
# )
# [b, s, h, d]
sequence_length = query.size(1)
encoder_sequence_length = encoder_query.size(1)
# H
query = torch.cat([query, encoder_query], dim=1).unsqueeze(2)
key = torch.cat([key, encoder_key], dim=1).unsqueeze(2)
value = torch.cat([value, encoder_value], dim=1).unsqueeze(2)
# B, S, 3, H, D
qkv = torch.cat([query, key, value], dim=2)
attn_mask = encoder_attention_mask[:, :].bool()
attn_mask = F.pad(attn_mask, (sequence_length, 0), value=True)
hidden_states = flash_attn_no_pad(
qkv, attn_mask, causal=False, dropout_p=0.0, softmax_scale=None
)
# hidden_states = F.scaled_dot_product_attention(query, key, value, attn_mask = None, dropout_p=0.0, is_causal=False)
# valid_lengths = encoder_attention_mask.sum(dim=1) + sequence_length
# def no_padding_mask(score, b, h, q_idx, kv_idx):
# return torch.where(kv_idx < valid_lengths[b],score, -float("inf"))
# hidden_states = flex_attention(query, key, value, score_mod=no_padding_mask)
if get_sequence_parallel_state():
hidden_states, encoder_hidden_states = hidden_states.split_with_sizes(
(sequence_length, encoder_sequence_length), dim=1
)
# B, S, H, D
hidden_states = all_to_all_4D(hidden_states, scatter_dim=1, gather_dim=2)
encoder_hidden_states = all_gather(
encoder_hidden_states, dim=2
).contiguous()
hidden_states = hidden_states.flatten(2, 3)
hidden_states = hidden_states.to(query.dtype)
encoder_hidden_states = encoder_hidden_states.flatten(2, 3)
encoder_hidden_states = encoder_hidden_states.to(query.dtype)
else:
hidden_states = hidden_states.flatten(2, 3)
hidden_states = hidden_states.to(query.dtype)
hidden_states, encoder_hidden_states = hidden_states.split_with_sizes(
(sequence_length, encoder_sequence_length), dim=1
)
# linear proj
hidden_states = attn.to_out[0](hidden_states)
# dropout
hidden_states = attn.to_out[1](hidden_states)
if hasattr(attn, "to_add_out"):
encoder_hidden_states = attn.to_add_out(encoder_hidden_states)
return hidden_states, encoder_hidden_states
@maybe_allow_in_graph
class MochiTransformerBlock(nn.Module):
r"""
Transformer block used in [Mochi](https://huggingface.co/genmo/mochi-1-preview).
Args:
dim (`int`):
The number of channels in the input and output.
num_attention_heads (`int`):
The number of heads to use for multi-head attention.
attention_head_dim (`int`):
The number of channels in each head.
qk_norm (`str`, defaults to `"rms_norm"`):
The normalization layer to use.
activation_fn (`str`, defaults to `"swiglu"`):
Activation function to use in feed-forward.
context_pre_only (`bool`, defaults to `False`):
Whether or not to process context-related conditions with additional layers.
eps (`float`, defaults to `1e-6`):
Epsilon value for normalization layers.
"""
def __init__(
self,
dim: int,
num_attention_heads: int,
attention_head_dim: int,
pooled_projection_dim: int,
qk_norm: str = "rms_norm",
activation_fn: str = "swiglu",
context_pre_only: bool = False,
eps: float = 1e-6,
) -> None:
super().__init__()
self.context_pre_only = context_pre_only
self.ff_inner_dim = (4 * dim * 2) // 3
self.ff_context_inner_dim = (4 * pooled_projection_dim * 2) // 3
self.norm1 = MochiRMSNormZero(dim, 4 * dim, eps=eps, elementwise_affine=False)
if not context_pre_only:
self.norm1_context = MochiRMSNormZero(
dim, 4 * pooled_projection_dim, eps=eps, elementwise_affine=False
)
else:
self.norm1_context = MochiLayerNormContinuous(
embedding_dim=pooled_projection_dim,
conditioning_embedding_dim=dim,
eps=eps,
)
self.attn1 = MochiAttention(
query_dim=dim,
heads=num_attention_heads,
dim_head=attention_head_dim,
bias=False,
added_kv_proj_dim=pooled_projection_dim,
added_proj_bias=False,
out_dim=dim,
out_context_dim=pooled_projection_dim,
context_pre_only=context_pre_only,
processor=MochiAttnProcessor2_0(),
eps=1e-5,
)
# TODO(aryan): norm_context layers are not needed when `context_pre_only` is True
self.norm2 = MochiModulatedRMSNorm(eps=eps)
self.norm2_context = (
MochiModulatedRMSNorm(eps=eps) if not self.context_pre_only else None
)
self.norm3 = MochiModulatedRMSNorm(eps)
self.norm3_context = (
MochiModulatedRMSNorm(eps=eps) if not self.context_pre_only else None
)
self.ff = FeedForward(
dim, inner_dim=self.ff_inner_dim, activation_fn=activation_fn, bias=False
)
self.ff_context = None
if not context_pre_only:
self.ff_context = FeedForward(
pooled_projection_dim,
inner_dim=self.ff_context_inner_dim,
activation_fn=activation_fn,
bias=False,
)
self.norm4 = MochiModulatedRMSNorm(eps=eps)
self.norm4_context = MochiModulatedRMSNorm(eps=eps)
def forward(
self,
hidden_states: torch.Tensor,
encoder_hidden_states: torch.Tensor,
encoder_attention_mask: torch.Tensor,
temb: torch.Tensor,
image_rotary_emb: Optional[torch.Tensor] = None,
output_attn=False,
) -> Tuple[torch.Tensor, torch.Tensor]:
norm_hidden_states, gate_msa, scale_mlp, gate_mlp = self.norm1(
hidden_states, temb
)
if not self.context_pre_only:
(
norm_encoder_hidden_states,
enc_gate_msa,
enc_scale_mlp,
enc_gate_mlp,
) = self.norm1_context(encoder_hidden_states, temb)
else:
norm_encoder_hidden_states = self.norm1_context(encoder_hidden_states, temb)
attn_hidden_states, context_attn_hidden_states = self.attn1(
hidden_states=norm_hidden_states,
encoder_hidden_states=norm_encoder_hidden_states,
image_rotary_emb=image_rotary_emb,
encoder_attention_mask=encoder_attention_mask,
)
hidden_states = hidden_states + self.norm2(
attn_hidden_states, torch.tanh(gate_msa).unsqueeze(1)
)
norm_hidden_states = self.norm3(
hidden_states, (1 + scale_mlp.unsqueeze(1).to(torch.float32))
)
ff_output = self.ff(norm_hidden_states)
hidden_states = hidden_states + self.norm4(
ff_output, torch.tanh(gate_mlp).unsqueeze(1)
)
if not self.context_pre_only:
encoder_hidden_states = encoder_hidden_states + self.norm2_context(
context_attn_hidden_states, torch.tanh(enc_gate_msa).unsqueeze(1)
)
norm_encoder_hidden_states = self.norm3_context(
encoder_hidden_states,
(1 + enc_scale_mlp.unsqueeze(1).to(torch.float32)),
)
context_ff_output = self.ff_context(norm_encoder_hidden_states)
encoder_hidden_states = encoder_hidden_states + self.norm4_context(
context_ff_output, torch.tanh(enc_gate_mlp).unsqueeze(1)
)
if not output_attn:
attn_hidden_states = None
return hidden_states, encoder_hidden_states, attn_hidden_states
class MochiRoPE(nn.Module):
r"""
RoPE implementation used in [Mochi](https://huggingface.co/genmo/mochi-1-preview).
Args:
base_height (`int`, defaults to `192`):
Base height used to compute interpolation scale for rotary positional embeddings.
base_width (`int`, defaults to `192`):
Base width used to compute interpolation scale for rotary positional embeddings.
"""
def __init__(self, base_height: int = 192, base_width: int = 192) -> None:
super().__init__()
self.target_area = base_height * base_width
def _centers(self, start, stop, num, device, dtype) -> torch.Tensor:
edges = torch.linspace(start, stop, num + 1, device=device, dtype=dtype)
return (edges[:-1] + edges[1:]) / 2
def _get_positions(
self,
num_frames: int,
height: int,
width: int,
device: Optional[torch.device] = None,
dtype: Optional[torch.dtype] = None,
) -> torch.Tensor:
scale = (self.target_area / (height * width)) ** 0.5
t = torch.arange(num_frames * nccl_info.sp_size, device=device, dtype=dtype)
h = self._centers(
-height * scale / 2, height * scale / 2, height, device, dtype
)
w = self._centers(-width * scale / 2, width * scale / 2, width, device, dtype)
grid_t, grid_h, grid_w = torch.meshgrid(t, h, w, indexing="ij")
positions = torch.stack([grid_t, grid_h, grid_w], dim=-1).view(-1, 3)
return positions
def _create_rope(self, freqs: torch.Tensor, pos: torch.Tensor) -> torch.Tensor:
with torch.autocast(freqs.device.type, enabled=False):
# Always run ROPE freqs computation in FP32
freqs = torch.einsum(
"nd,dhf->nhf", pos.to(torch.float32), freqs.to(torch.float32)
)
freqs_cos = torch.cos(freqs)
freqs_sin = torch.sin(freqs)
return freqs_cos, freqs_sin
def forward(
self,
pos_frequencies: torch.Tensor,
num_frames: int,
height: int,
width: int,
device: Optional[torch.device] = None,
dtype: Optional[torch.dtype] = None,
) -> Tuple[torch.Tensor, torch.Tensor]:
pos = self._get_positions(num_frames, height, width, device, dtype)
rope_cos, rope_sin = self._create_rope(pos_frequencies, pos)
return rope_cos, rope_sin
@maybe_allow_in_graph
class MochiTransformer3DModel(ModelMixin, ConfigMixin, PeftAdapterMixin):
r"""
A Transformer model for video-like data introduced in [Mochi](https://huggingface.co/genmo/mochi-1-preview).
Args:
patch_size (`int`, defaults to `2`):
The size of the patches to use in the patch embedding layer.
num_attention_heads (`int`, defaults to `24`):
The number of heads to use for multi-head attention.
attention_head_dim (`int`, defaults to `128`):
The number of channels in each head.
num_layers (`int`, defaults to `48`):
The number of layers of Transformer blocks to use.
in_channels (`int`, defaults to `12`):
The number of channels in the input.
out_channels (`int`, *optional*, defaults to `None`):
The number of channels in the output.
qk_norm (`str`, defaults to `"rms_norm"`):
The normalization layer to use.
text_embed_dim (`int`, defaults to `4096`):
Input dimension of text embeddings from the text encoder.
time_embed_dim (`int`, defaults to `256`):
Output dimension of timestep embeddings.
activation_fn (`str`, defaults to `"swiglu"`):
Activation function to use in feed-forward.
max_sequence_length (`int`, defaults to `256`):
The maximum sequence length of text embeddings supported.
"""
_supports_gradient_checkpointing = True
@register_to_config
def __init__(
self,
patch_size: int = 2,
num_attention_heads: int = 24,
attention_head_dim: int = 128,
num_layers: int = 48,
pooled_projection_dim: int = 1536,
in_channels: int = 12,
out_channels: Optional[int] = None,
qk_norm: str = "rms_norm",
text_embed_dim: int = 4096,
time_embed_dim: int = 256,
activation_fn: str = "swiglu",
max_sequence_length: int = 256,
) -> None:
super().__init__()
inner_dim = num_attention_heads * attention_head_dim
out_channels = out_channels or in_channels
self.patch_embed = PatchEmbed(
patch_size=patch_size,
in_channels=in_channels,
embed_dim=inner_dim,
pos_embed_type=None,
)
self.time_embed = MochiCombinedTimestepCaptionEmbedding(
embedding_dim=inner_dim,
pooled_projection_dim=pooled_projection_dim,
text_embed_dim=text_embed_dim,
time_embed_dim=time_embed_dim,
num_attention_heads=8,
)
self.pos_frequencies = nn.Parameter(
torch.full((3, num_attention_heads, attention_head_dim // 2), 0.0)
)
self.rope = MochiRoPE()
self.transformer_blocks = nn.ModuleList(
[
MochiTransformerBlock(
dim=inner_dim,
num_attention_heads=num_attention_heads,
attention_head_dim=attention_head_dim,
pooled_projection_dim=pooled_projection_dim,
qk_norm=qk_norm,
activation_fn=activation_fn,
context_pre_only=i == num_layers - 1,
)
for i in range(num_layers)
]
)
self.norm_out = AdaLayerNormContinuous(
inner_dim,
inner_dim,
elementwise_affine=False,
eps=1e-6,
norm_type="layer_norm",
)
self.proj_out = nn.Linear(inner_dim, patch_size * patch_size * out_channels)
self.gradient_checkpointing = False
def _set_gradient_checkpointing(self, module, value=False):
if hasattr(module, "gradient_checkpointing"):
module.gradient_checkpointing = value
def forward(
self,
hidden_states: torch.Tensor,
encoder_hidden_states: torch.Tensor,
timestep: torch.LongTensor,
encoder_attention_mask: torch.Tensor,
output_attn=False,
attention_kwargs: Optional[Dict[str, Any]] = None,
return_dict: bool = False,
) -> torch.Tensor:
assert (
return_dict is False
), "return_dict is not supported in MochiTransformer3DModel"
if attention_kwargs is not None:
attention_kwargs = attention_kwargs.copy()
lora_scale = attention_kwargs.pop("scale", 1.0)
else:
lora_scale = 1.0
if USE_PEFT_BACKEND:
# weight the lora layers by setting `lora_scale` for each PEFT layer
scale_lora_layers(self, lora_scale)
else:
if (
attention_kwargs is not None
and attention_kwargs.get("scale", None) is not None
):
logger.warning(
"Passing `scale` via `attention_kwargs` when not using the PEFT backend is ineffective."
)
batch_size, num_channels, num_frames, height, width = hidden_states.shape
p = self.config.patch_size
post_patch_height = height // p
post_patch_width = width // p
# Peiyuan: This is hacked to force mochi to follow the behaviour of SD3 and Flux
timestep = 1000 - timestep
temb, encoder_hidden_states = self.time_embed(
timestep,
encoder_hidden_states,
encoder_attention_mask,
hidden_dtype=hidden_states.dtype,
)
hidden_states = hidden_states.permute(0, 2, 1, 3, 4).flatten(0, 1)
hidden_states = self.patch_embed(hidden_states)
hidden_states = hidden_states.unflatten(0, (batch_size, -1)).flatten(1, 2)
image_rotary_emb = self.rope(
self.pos_frequencies,
num_frames,
post_patch_height,
post_patch_width,
device=hidden_states.device,
dtype=torch.float32,
)
attn_outputs_list = []
for i, block in enumerate(self.transformer_blocks):
if self.gradient_checkpointing:
def create_custom_forward(module):
def custom_forward(*inputs):
return module(*inputs)
return custom_forward
ckpt_kwargs: Dict[str, Any] = (
{"use_reentrant": False} if is_torch_version(">=", "1.11.0") else {}
)
(
hidden_states,
encoder_hidden_states,
attn_outputs,
) = torch.utils.checkpoint.checkpoint(
create_custom_forward(block),
hidden_states,
encoder_hidden_states,
encoder_attention_mask,
temb,
image_rotary_emb,
output_attn,
**ckpt_kwargs,
)
else:
hidden_states, encoder_hidden_states, attn_outputs = block(
hidden_states=hidden_states,
encoder_hidden_states=encoder_hidden_states,
encoder_attention_mask=encoder_attention_mask,
temb=temb,
image_rotary_emb=image_rotary_emb,
output_attn=output_attn,
)
attn_outputs_list.append(attn_outputs)
hidden_states = self.norm_out(hidden_states, temb)
hidden_states = self.proj_out(hidden_states)
hidden_states = hidden_states.reshape(
batch_size, num_frames, post_patch_height, post_patch_width, p, p, -1
)
hidden_states = hidden_states.permute(0, 6, 1, 2, 4, 3, 5)
output = hidden_states.reshape(batch_size, -1, num_frames, height, width)
if USE_PEFT_BACKEND:
# remove `lora_scale` from each PEFT layer
unscale_lora_layers(self, lora_scale)
if not output_attn:
attn_outputs_list = None
else:
attn_outputs_list = torch.stack(attn_outputs_list, dim=0)
# Peiyuan: This is hacked to force mochi to follow the behaviour of SD3 and Flux
return (-output, attn_outputs_list)
-130
View File
@@ -1,130 +0,0 @@
# Copyright 2024 The Genmo team and The HuggingFace Team.
# All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import numbers
from typing import Dict, Optional, Tuple
import torch
import torch.nn as nn
import torch.nn.functional as F
class MochiModulatedRMSNorm(nn.Module):
def __init__(self, eps: float):
super().__init__()
self.eps = eps
def forward(self, hidden_states, scale=None):
hidden_states_dtype = hidden_states.dtype
hidden_states = hidden_states.to(torch.float32)
variance = hidden_states.pow(2).mean(-1, keepdim=True)
hidden_states = hidden_states * torch.rsqrt(variance + self.eps)
if scale is not None:
hidden_states = hidden_states * scale
hidden_states = hidden_states.to(hidden_states_dtype)
return hidden_states
class MochiRMSNorm(nn.Module):
def __init__(self, dim, eps: float, elementwise_affine=True):
super().__init__()
self.eps = eps
if elementwise_affine:
self.weight = nn.Parameter(torch.ones(dim))
else:
self.weight = None
def forward(self, hidden_states):
hidden_states_dtype = hidden_states.dtype
hidden_states = hidden_states.to(torch.float32)
variance = hidden_states.pow(2).mean(-1, keepdim=True)
hidden_states = hidden_states * torch.rsqrt(variance + self.eps)
if self.weight is not None:
# convert into half-precision if necessary
if self.weight.dtype in [torch.float16, torch.bfloat16]:
hidden_states = hidden_states.to(self.weight.dtype)
hidden_states = hidden_states * self.weight
hidden_states = hidden_states.to(hidden_states_dtype)
return hidden_states
class MochiLayerNormContinuous(nn.Module):
def __init__(
self,
embedding_dim: int,
conditioning_embedding_dim: int,
eps=1e-5,
bias=True,
):
super().__init__()
# AdaLN
self.silu = nn.SiLU()
self.linear_1 = nn.Linear(conditioning_embedding_dim, embedding_dim, bias=bias)
self.norm = MochiModulatedRMSNorm(eps=eps)
def forward(
self,
x: torch.Tensor,
conditioning_embedding: torch.Tensor,
) -> torch.Tensor:
input_dtype = x.dtype
# convert back to the original dtype in case `conditioning_embedding`` is upcasted to float32 (needed for hunyuanDiT)
scale = self.linear_1(self.silu(conditioning_embedding).to(x.dtype))
x = self.norm(x, (1 + scale.unsqueeze(1).to(torch.float32)))
return x.to(input_dtype)
class MochiRMSNormZero(nn.Module):
r"""
Adaptive RMS Norm used in Mochi.
Parameters:
embedding_dim (`int`): The size of each embedding vector.
"""
def __init__(
self,
embedding_dim: int,
hidden_dim: int,
eps: float = 1e-5,
elementwise_affine: bool = False,
) -> None:
super().__init__()
self.silu = nn.SiLU()
self.linear = nn.Linear(embedding_dim, hidden_dim)
self.norm = MochiModulatedRMSNorm(eps=eps)
def forward(
self, hidden_states: torch.Tensor, emb: torch.Tensor
) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor, torch.Tensor]:
hidden_states_dtype = hidden_states.dtype
emb = self.linear(self.silu(emb))
scale_msa, gate_msa, scale_mlp, gate_mlp = emb.chunk(4, dim=1)
hidden_states = self.norm(
hidden_states, (1 + scale_msa[:, None].to(torch.float32))
)
hidden_states = hidden_states.to(hidden_states_dtype)
return hidden_states, gate_msa, scale_mlp, gate_mlp
-849
View File
@@ -1,849 +0,0 @@
# Copyright 2024 Black Forest Labs and The HuggingFace Team. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import inspect
from typing import Callable, Dict, List, Optional, Union, Any
import copy
import numpy as np
import torch
from transformers import T5EncoderModel, T5TokenizerFast
from diffusers.callbacks import MultiPipelineCallbacks, PipelineCallback
from diffusers.models.autoencoders import AutoencoderKL
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
from diffusers.schedulers import FlowMatchEulerDiscreteScheduler
from diffusers.utils import (
is_torch_xla_available,
logging,
replace_example_docstring,
)
from diffusers.utils.torch_utils import randn_tensor
from diffusers.video_processor import VideoProcessor
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
from diffusers.pipelines.mochi.pipeline_output import MochiPipelineOutput
from einops import rearrange
from fastvideo.utils.parallel_states import get_sequence_parallel_state, nccl_info
from fastvideo.utils.communications import all_gather
from diffusers.loaders import Mochi1LoraLoaderMixin
if is_torch_xla_available():
import torch_xla.core.xla_model as xm
XLA_AVAILABLE = True
else:
XLA_AVAILABLE = False
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
EXAMPLE_DOC_STRING = """
Examples:
```py
>>> import torch
>>> from diffusers import MochiPipeline
>>> from diffusers.utils import export_to_video
>>> pipe = MochiPipeline.from_pretrained("genmo/mochi-1-preview", torch_dtype=torch.bfloat16)
>>> pipe.to("cuda")
>>> prompt = "Close-up of a chameleon's eye, with its scaly skin changing color. Ultra high resolution 4k."
>>> frames = pipe(prompt, num_inference_steps=28, guidance_scale=3.5).frames[0]
>>> export_to_video(frames, "mochi.mp4")
```
"""
def calculate_shift(
image_seq_len,
base_seq_len: int = 256,
max_seq_len: int = 4096,
base_shift: float = 0.5,
max_shift: float = 1.16,
):
m = (max_shift - base_shift) / (max_seq_len - base_seq_len)
b = base_shift - m * base_seq_len
mu = image_seq_len * m + b
return mu
# from: https://github.com/genmoai/models/blob/075b6e36db58f1242921deff83a1066887b9c9e1/src/mochi_preview/infer.py#L77
def linear_quadratic_schedule(num_steps, threshold_noise, linear_steps=None):
if linear_steps is None:
linear_steps = num_steps // 2
linear_sigma_schedule = [
i * threshold_noise / linear_steps for i in range(linear_steps)
]
threshold_noise_step_diff = linear_steps - threshold_noise * num_steps
quadratic_steps = num_steps - linear_steps
quadratic_coef = threshold_noise_step_diff / (linear_steps * quadratic_steps**2)
linear_coef = threshold_noise / linear_steps - 2 * threshold_noise_step_diff / (
quadratic_steps**2
)
const = quadratic_coef * (linear_steps**2)
quadratic_sigma_schedule = [
quadratic_coef * (i**2) + linear_coef * i + const
for i in range(linear_steps, num_steps)
]
sigma_schedule = linear_sigma_schedule + quadratic_sigma_schedule
sigma_schedule = [1.0 - x for x in sigma_schedule]
return sigma_schedule
# Copied from diffusers.pipelines.stable_diffusion.pipeline_stable_diffusion.retrieve_timesteps
def retrieve_timesteps(
scheduler,
num_inference_steps: Optional[int] = None,
device: Optional[Union[str, torch.device]] = None,
timesteps: Optional[List[int]] = None,
sigmas: Optional[List[float]] = None,
**kwargs,
):
r"""
Calls the scheduler's `set_timesteps` method and retrieves timesteps from the scheduler after the call. Handles
custom timesteps. Any kwargs will be supplied to `scheduler.set_timesteps`.
Args:
scheduler (`SchedulerMixin`):
The scheduler to get timesteps from.
num_inference_steps (`int`):
The number of diffusion steps used when generating samples with a pre-trained model. If used, `timesteps`
must be `None`.
device (`str` or `torch.device`, *optional*):
The device to which the timesteps should be moved to. If `None`, the timesteps are not moved.
timesteps (`List[int]`, *optional*):
Custom timesteps used to override the timestep spacing strategy of the scheduler. If `timesteps` is passed,
`num_inference_steps` and `sigmas` must be `None`.
sigmas (`List[float]`, *optional*):
Custom sigmas used to override the timestep spacing strategy of the scheduler. If `sigmas` is passed,
`num_inference_steps` and `timesteps` must be `None`.
Returns:
`Tuple[torch.Tensor, int]`: A tuple where the first element is the timestep schedule from the scheduler and the
second element is the number of inference steps.
"""
if timesteps is not None and sigmas is not None:
raise ValueError(
"Only one of `timesteps` or `sigmas` can be passed. Please choose one to set custom values"
)
if timesteps is not None:
accepts_timesteps = "timesteps" in set(
inspect.signature(scheduler.set_timesteps).parameters.keys()
)
if not accepts_timesteps:
raise ValueError(
f"The current scheduler class {scheduler.__class__}'s `set_timesteps` does not support custom"
f" timestep schedules. Please check whether you are using the correct scheduler."
)
scheduler.set_timesteps(timesteps=timesteps, device=device, **kwargs)
timesteps = scheduler.timesteps
num_inference_steps = len(timesteps)
elif sigmas is not None:
accept_sigmas = "sigmas" in set(
inspect.signature(scheduler.set_timesteps).parameters.keys()
)
if not accept_sigmas:
raise ValueError(
f"The current scheduler class {scheduler.__class__}'s `set_timesteps` does not support custom"
f" sigmas schedules. Please check whether you are using the correct scheduler."
)
scheduler.set_timesteps(sigmas=sigmas, device=device, **kwargs)
timesteps = scheduler.timesteps
num_inference_steps = len(timesteps)
else:
scheduler.set_timesteps(num_inference_steps, device=device, **kwargs)
timesteps = scheduler.timesteps
return timesteps, num_inference_steps
class MochiPipeline(DiffusionPipeline, Mochi1LoraLoaderMixin):
r"""
The mochi pipeline for text-to-video generation.
Reference: https://github.com/genmoai/models
Args:
transformer ([`MochiTransformer3DModel`]):
Conditional Transformer architecture to denoise the encoded video latents.
scheduler ([`FlowMatchEulerDiscreteScheduler`]):
A scheduler to be used in combination with `transformer` to denoise the encoded image latents.
vae ([`AutoencoderKL`]):
Variational Auto-Encoder (VAE) Model to encode and decode images to and from latent representations.
text_encoder ([`T5EncoderModel`]):
[T5](https://huggingface.co/docs/transformers/en/model_doc/t5#transformers.T5EncoderModel), specifically
the [google/t5-v1_1-xxl](https://huggingface.co/google/t5-v1_1-xxl) variant.
tokenizer (`CLIPTokenizer`):
Tokenizer of class
[CLIPTokenizer](https://huggingface.co/docs/transformers/en/model_doc/clip#transformers.CLIPTokenizer).
tokenizer (`T5TokenizerFast`):
Second Tokenizer of class
[T5TokenizerFast](https://huggingface.co/docs/transformers/en/model_doc/t5#transformers.T5TokenizerFast).
"""
model_cpu_offload_seq = "text_encoder->transformer->vae"
_optional_components = []
_callback_tensor_inputs = ["latents", "prompt_embeds", "negative_prompt_embeds"]
def __init__(
self,
scheduler: FlowMatchEulerDiscreteScheduler,
vae: AutoencoderKL,
text_encoder: T5EncoderModel,
tokenizer: T5TokenizerFast,
transformer: MochiTransformer3DModel,
):
super().__init__()
self.register_modules(
vae=vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
transformer=transformer,
scheduler=scheduler,
)
self.vae_spatial_scale_factor = 8
self.vae_temporal_scale_factor = 6
self.patch_size = 2
self.video_processor = VideoProcessor(
vae_scale_factor=self.vae_spatial_scale_factor
)
self.tokenizer_max_length = (
self.tokenizer.model_max_length
if hasattr(self, "tokenizer") and self.tokenizer is not None
else 77
)
self.default_height = 480
self.default_width = 848
# Adapted from diffusers.pipelines.cogvideo.pipeline_cogvideox.CogVideoXPipeline._get_t5_prompt_embeds
def _get_t5_prompt_embeds(
self,
prompt: Union[str, List[str]] = None,
num_videos_per_prompt: int = 1,
max_sequence_length: int = 256,
device: Optional[torch.device] = None,
dtype: Optional[torch.dtype] = None,
):
device = device or self._execution_device
dtype = dtype or self.text_encoder.dtype
prompt = [prompt] if isinstance(prompt, str) else prompt
batch_size = len(prompt)
text_inputs = self.tokenizer(
prompt,
padding="max_length",
max_length=max_sequence_length,
truncation=True,
add_special_tokens=True,
return_tensors="pt",
)
text_input_ids = text_inputs.input_ids
prompt_attention_mask = text_inputs.attention_mask
prompt_attention_mask = prompt_attention_mask.bool().to(device)
untruncated_ids = self.tokenizer(
prompt, padding="longest", return_tensors="pt"
).input_ids
if untruncated_ids.shape[-1] >= text_input_ids.shape[-1] and not torch.equal(
text_input_ids, untruncated_ids
):
removed_text = self.tokenizer.batch_decode(
untruncated_ids[:, max_sequence_length - 1 : -1]
)
logger.warning(
"The following part of your input was truncated because `max_sequence_length` is set to "
f" {max_sequence_length} tokens: {removed_text}"
)
prompt_embeds = self.text_encoder(
text_input_ids.to(device), attention_mask=prompt_attention_mask
)[0]
prompt_embeds = prompt_embeds.to(dtype=dtype, device=device)
# duplicate text embeddings for each generation per prompt, using mps friendly method
_, seq_len, _ = prompt_embeds.shape
prompt_embeds = prompt_embeds.repeat(1, num_videos_per_prompt, 1)
prompt_embeds = prompt_embeds.view(
batch_size * num_videos_per_prompt, seq_len, -1
)
prompt_attention_mask = prompt_attention_mask.view(batch_size, -1)
prompt_attention_mask = prompt_attention_mask.repeat(num_videos_per_prompt, 1)
return prompt_embeds, prompt_attention_mask
# Adapted from diffusers.pipelines.cogvideo.pipeline_cogvideox.CogVideoXPipeline.encode_prompt
def encode_prompt(
self,
prompt: Union[str, List[str]],
negative_prompt: Optional[Union[str, List[str]]] = None,
do_classifier_free_guidance: bool = True,
num_videos_per_prompt: int = 1,
prompt_embeds: Optional[torch.Tensor] = None,
negative_prompt_embeds: Optional[torch.Tensor] = None,
prompt_attention_mask: Optional[torch.Tensor] = None,
negative_prompt_attention_mask: Optional[torch.Tensor] = None,
max_sequence_length: int = 256,
device: Optional[torch.device] = None,
dtype: Optional[torch.dtype] = None,
):
r"""
Encodes the prompt into text encoder hidden states.
Args:
prompt (`str` or `List[str]`, *optional*):
prompt to be encoded
negative_prompt (`str` or `List[str]`, *optional*):
The prompt or prompts not to guide the image generation. If not defined, one has to pass
`negative_prompt_embeds` instead. Ignored when not using guidance (i.e., ignored if `guidance_scale` is
less than `1`).
do_classifier_free_guidance (`bool`, *optional*, defaults to `True`):
Whether to use classifier free guidance or not.
num_videos_per_prompt (`int`, *optional*, defaults to 1):
Number of videos that should be generated per prompt. torch device to place the resulting embeddings on
prompt_embeds (`torch.Tensor`, *optional*):
Pre-generated text embeddings. Can be used to easily tweak text inputs, *e.g.* prompt weighting. If not
provided, text embeddings will be generated from `prompt` input argument.
negative_prompt_embeds (`torch.Tensor`, *optional*):
Pre-generated negative text embeddings. Can be used to easily tweak text inputs, *e.g.* prompt
weighting. If not provided, negative_prompt_embeds will be generated from `negative_prompt` input
argument.
device: (`torch.device`, *optional*):
torch device
dtype: (`torch.dtype`, *optional*):
torch dtype
"""
device = device or self._execution_device
prompt = [prompt] if isinstance(prompt, str) else prompt
if prompt is not None:
batch_size = len(prompt)
else:
batch_size = prompt_embeds.shape[0]
if prompt_embeds is None:
prompt_embeds, prompt_attention_mask = self._get_t5_prompt_embeds(
prompt=prompt,
num_videos_per_prompt=num_videos_per_prompt,
max_sequence_length=max_sequence_length,
device=device,
dtype=dtype,
)
if do_classifier_free_guidance and negative_prompt_embeds is None:
negative_prompt = negative_prompt or ""
negative_prompt = (
batch_size * [negative_prompt]
if isinstance(negative_prompt, str)
else negative_prompt
)
if prompt is not None and type(prompt) is not type(negative_prompt):
raise TypeError(
f"`negative_prompt` should be the same type to `prompt`, but got {type(negative_prompt)} !="
f" {type(prompt)}."
)
elif batch_size != len(negative_prompt):
raise ValueError(
f"`negative_prompt`: {negative_prompt} has batch size {len(negative_prompt)}, but `prompt`:"
f" {prompt} has batch size {batch_size}. Please make sure that passed `negative_prompt` matches"
" the batch size of `prompt`."
)
(
negative_prompt_embeds,
negative_prompt_attention_mask,
) = self._get_t5_prompt_embeds(
prompt=negative_prompt,
num_videos_per_prompt=num_videos_per_prompt,
max_sequence_length=max_sequence_length,
device=device,
dtype=dtype,
)
return (
prompt_embeds,
prompt_attention_mask,
negative_prompt_embeds,
negative_prompt_attention_mask,
)
def check_inputs(
self,
prompt,
height,
width,
callback_on_step_end_tensor_inputs=None,
prompt_embeds=None,
negative_prompt_embeds=None,
prompt_attention_mask=None,
negative_prompt_attention_mask=None,
):
if height % 8 != 0 or width % 8 != 0:
raise ValueError(
f"`height` and `width` have to be divisible by 8 but are {height} and {width}."
)
if callback_on_step_end_tensor_inputs is not None and not all(
k in self._callback_tensor_inputs
for k in callback_on_step_end_tensor_inputs
):
raise ValueError(
f"`callback_on_step_end_tensor_inputs` has to be in {self._callback_tensor_inputs}, but found {[k for k in callback_on_step_end_tensor_inputs if k not in self._callback_tensor_inputs]}"
)
if prompt is not None and prompt_embeds is not None:
raise ValueError(
f"Cannot forward both `prompt`: {prompt} and `prompt_embeds`: {prompt_embeds}. Please make sure to"
" only forward one of the two."
)
elif prompt is None and prompt_embeds is None:
raise ValueError(
"Provide either `prompt` or `prompt_embeds`. Cannot leave both `prompt` and `prompt_embeds` undefined."
)
elif prompt is not None and (
not isinstance(prompt, str) and not isinstance(prompt, list)
):
raise ValueError(
f"`prompt` has to be of type `str` or `list` but is {type(prompt)}"
)
if prompt_embeds is not None and prompt_attention_mask is None:
raise ValueError(
"Must provide `prompt_attention_mask` when specifying `prompt_embeds`."
)
if (
negative_prompt_embeds is not None
and negative_prompt_attention_mask is None
):
raise ValueError(
"Must provide `negative_prompt_attention_mask` when specifying `negative_prompt_embeds`."
)
if prompt_embeds is not None and negative_prompt_embeds is not None:
if prompt_embeds.shape != negative_prompt_embeds.shape:
raise ValueError(
"`prompt_embeds` and `negative_prompt_embeds` must have the same shape when passed directly, but"
f" got: `prompt_embeds` {prompt_embeds.shape} != `negative_prompt_embeds`"
f" {negative_prompt_embeds.shape}."
)
if prompt_attention_mask.shape != negative_prompt_attention_mask.shape:
raise ValueError(
"`prompt_attention_mask` and `negative_prompt_attention_mask` must have the same shape when passed directly, but"
f" got: `prompt_attention_mask` {prompt_attention_mask.shape} != `negative_prompt_attention_mask`"
f" {negative_prompt_attention_mask.shape}."
)
def enable_vae_slicing(self):
r"""
Enable sliced VAE decoding. When this option is enabled, the VAE will split the input tensor in slices to
compute decoding in several steps. This is useful to save some memory and allow larger batch sizes.
"""
self.vae.enable_slicing()
def disable_vae_slicing(self):
r"""
Disable sliced VAE decoding. If `enable_vae_slicing` was previously enabled, this method will go back to
computing decoding in one step.
"""
self.vae.disable_slicing()
def enable_vae_tiling(self):
r"""
Enable tiled VAE decoding. When this option is enabled, the VAE will split the input tensor into tiles to
compute decoding and encoding in several steps. This is useful for saving a large amount of memory and to allow
processing larger images.
"""
self.vae.enable_tiling()
def disable_vae_tiling(self):
r"""
Disable tiled VAE decoding. If `enable_vae_tiling` was previously enabled, this method will go back to
computing decoding in one step.
"""
self.vae.disable_tiling()
def prepare_latents(
self,
batch_size,
num_channels_latents,
height,
width,
num_frames,
dtype,
device,
generator,
latents=None,
):
height = height // self.vae_spatial_scale_factor
width = width // self.vae_spatial_scale_factor
num_frames = (num_frames - 1) // self.vae_temporal_scale_factor + 1
shape = (batch_size, num_channels_latents, num_frames, height, width)
if latents is not None:
return latents.to(device=device, dtype=dtype)
if isinstance(generator, list) and len(generator) != batch_size:
raise ValueError(
f"You have passed a list of generators of length {len(generator)}, but requested an effective batch"
f" size of {batch_size}. Make sure the batch size matches the length of the generators."
)
latents = randn_tensor(shape, generator=generator, device=device, dtype=dtype)
return latents
@property
def guidance_scale(self):
return self._guidance_scale
@property
def do_classifier_free_guidance(self):
return self._guidance_scale > 1.0
@property
def num_timesteps(self):
return self._num_timesteps
@property
def attention_kwargs(self):
return self._attention_kwargs
@property
def interrupt(self):
return self._interrupt
@torch.no_grad()
@replace_example_docstring(EXAMPLE_DOC_STRING)
def __call__(
self,
prompt: Union[str, List[str]] = None,
negative_prompt: Optional[Union[str, List[str]]] = None,
height: Optional[int] = None,
width: Optional[int] = None,
num_frames: int = 16,
num_inference_steps: int = 28,
timesteps: List[int] = None,
guidance_scale: float = 4.5,
num_videos_per_prompt: Optional[int] = 1,
generator: Optional[Union[torch.Generator, List[torch.Generator]]] = None,
latents: Optional[torch.Tensor] = None,
prompt_embeds: Optional[torch.Tensor] = None,
prompt_attention_mask: Optional[torch.Tensor] = None,
negative_prompt_embeds: Optional[torch.Tensor] = None,
negative_prompt_attention_mask: Optional[torch.Tensor] = None,
output_type: Optional[str] = "pil",
return_dict: bool = True,
attention_kwargs: Optional[Dict[str, Any]] = None,
callback_on_step_end: Optional[Callable[[int, int, Dict], None]] = None,
callback_on_step_end_tensor_inputs: List[str] = ["latents"],
max_sequence_length: int = 256,
return_all_states=False,
):
r"""
Function invoked when calling the pipeline for generation.
Args:
prompt (`str` or `List[str]`, *optional*):
The prompt or prompts to guide the image generation. If not defined, one has to pass `prompt_embeds`.
instead.
height (`int`, *optional*, defaults to self.unet.config.sample_size * self.vae_scale_factor):
The height in pixels of the generated image. This is set to 1024 by default for the best results.
width (`int`, *optional*, defaults to self.unet.config.sample_size * self.vae_scale_factor):
The width in pixels of the generated image. This is set to 1024 by default for the best results.
num_frames (`int`, defaults to 16):
The number of video frames to generate
num_inference_steps (`int`, *optional*, defaults to 50):
The number of denoising steps. More denoising steps usually lead to a higher quality image at the
expense of slower inference.
timesteps (`List[int]`, *optional*):
Custom timesteps to use for the denoising process with schedulers which support a `timesteps` argument
in their `set_timesteps` method. If not defined, the default behavior when `num_inference_steps` is
passed will be used. Must be in descending order.
guidance_scale (`float`, defaults to `4.5`):
Guidance scale as defined in [Classifier-Free Diffusion Guidance](https://arxiv.org/abs/2207.12598).
`guidance_scale` is defined as `w` of equation 2. of [Imagen
Paper](https://arxiv.org/pdf/2205.11487.pdf). Guidance scale is enabled by setting `guidance_scale >
1`. Higher guidance scale encourages to generate images that are closely linked to the text `prompt`,
usually at the expense of lower image quality.
num_videos_per_prompt (`int`, *optional*, defaults to 1):
The number of videos to generate per prompt.
generator (`torch.Generator` or `List[torch.Generator]`, *optional*):
One or a list of [torch generator(s)](https://pytorch.org/docs/stable/generated/torch.Generator.html)
to make generation deterministic.
latents (`torch.Tensor`, *optional*):
Pre-generated noisy latents, sampled from a Gaussian distribution, to be used as inputs for image
generation. Can be used to tweak the same generation with different prompts. If not provided, a latents
tensor will ge generated by sampling using the supplied random `generator`.
prompt_embeds (`torch.Tensor`, *optional*):
Pre-generated text embeddings. Can be used to easily tweak text inputs, *e.g.* prompt weighting. If not
provided, text embeddings will be generated from `prompt` input argument.
prompt_attention_mask (`torch.Tensor`, *optional*):
Pre-generated attention mask for text embeddings.
negative_prompt_embeds (`torch.FloatTensor`, *optional*):
Pre-generated negative text embeddings. For PixArt-Sigma this negative prompt should be "". If not
provided, negative_prompt_embeds will be generated from `negative_prompt` input argument.
negative_prompt_attention_mask (`torch.FloatTensor`, *optional*):
Pre-generated attention mask for negative text embeddings.
output_type (`str`, *optional*, defaults to `"pil"`):
The output format of the generate image. Choose between
[PIL](https://pillow.readthedocs.io/en/stable/): `PIL.Image.Image` or `np.array`.
return_dict (`bool`, *optional*, defaults to `True`):
Whether or not to return a [`~pipelines.mochi.MochiPipelineOutput`] instead of a plain tuple.
attention_kwargs (`dict`, *optional*):
A kwargs dictionary that if specified is passed along to the `AttentionProcessor` as defined under
`self.processor` in
[diffusers.models.attention_processor](https://github.com/huggingface/diffusers/blob/main/src/diffusers/models/attention_processor.py).
callback_on_step_end (`Callable`, *optional*):
A function that calls at the end of each denoising steps during the inference. The function is called
with the following arguments: `callback_on_step_end(self: DiffusionPipeline, step: int, timestep: int,
callback_kwargs: Dict)`. `callback_kwargs` will include a list of all tensors as specified by
`callback_on_step_end_tensor_inputs`.
callback_on_step_end_tensor_inputs (`List`, *optional*):
The list of tensor inputs for the `callback_on_step_end` function. The tensors specified in the list
will be passed as `callback_kwargs` argument. You will only be able to include variables listed in the
`._callback_tensor_inputs` attribute of your pipeline class.
max_sequence_length (`int` defaults to `256`):
Maximum sequence length to use with the `prompt`.
Examples:
Returns:
[`~pipelines.mochi.MochiPipelineOutput`] or `tuple`:
If `return_dict` is `True`, [`~pipelines.mochi.MochiPipelineOutput`] is returned, otherwise a `tuple`
is returned where the first element is a list with the generated images.
"""
if isinstance(callback_on_step_end, (PipelineCallback, MultiPipelineCallbacks)):
callback_on_step_end_tensor_inputs = callback_on_step_end.tensor_inputs
height = height or self.default_height
width = width or self.default_width
# 1. Check inputs. Raise error if not correct
self.check_inputs(
prompt=prompt,
height=height,
width=width,
callback_on_step_end_tensor_inputs=callback_on_step_end_tensor_inputs,
prompt_embeds=prompt_embeds,
negative_prompt_embeds=negative_prompt_embeds,
prompt_attention_mask=prompt_attention_mask,
negative_prompt_attention_mask=negative_prompt_attention_mask,
)
self._guidance_scale = guidance_scale
self._attention_kwargs = attention_kwargs
self._interrupt = False
# 2. Define call parameters
if prompt is not None and isinstance(prompt, str):
batch_size = 1
elif prompt is not None and isinstance(prompt, list):
batch_size = len(prompt)
else:
batch_size = prompt_embeds.shape[0]
device = self._execution_device
# 3. Prepare text embeddings
(
prompt_embeds,
prompt_attention_mask,
negative_prompt_embeds,
negative_prompt_attention_mask,
) = self.encode_prompt(
prompt=prompt,
negative_prompt=negative_prompt,
do_classifier_free_guidance=self.do_classifier_free_guidance,
num_videos_per_prompt=num_videos_per_prompt,
prompt_embeds=prompt_embeds,
negative_prompt_embeds=negative_prompt_embeds,
prompt_attention_mask=prompt_attention_mask,
negative_prompt_attention_mask=negative_prompt_attention_mask,
max_sequence_length=max_sequence_length,
device=device,
)
if self.do_classifier_free_guidance:
prompt_embeds = torch.cat([negative_prompt_embeds, prompt_embeds], dim=0)
prompt_attention_mask = torch.cat(
[negative_prompt_attention_mask, prompt_attention_mask], dim=0
)
# 4. Prepare latent variables
num_channels_latents = self.transformer.config.in_channels
latents = self.prepare_latents(
batch_size * num_videos_per_prompt,
num_channels_latents,
height,
width,
num_frames,
prompt_embeds.dtype,
device,
generator,
latents,
)
world_size, rank = nccl_info.sp_size, nccl_info.rank_within_group
if get_sequence_parallel_state():
latents = rearrange(
latents, "b t (n s) h w -> b t n s h w", n=world_size
).contiguous()
latents = latents[:, :, rank, :, :, :]
original_noise = copy.deepcopy(latents)
# 5. Prepare timestep
# from https://github.com/genmoai/models/blob/075b6e36db58f1242921deff83a1066887b9c9e1/src/mochi_preview/infer.py#L77
threshold_noise = 0.025
sigmas = linear_quadratic_schedule(num_inference_steps, threshold_noise)
sigmas = np.array(sigmas)
# check if of type FlowMatchEulerDiscreteScheduler
if isinstance(self.scheduler, FlowMatchEulerDiscreteScheduler):
timesteps, num_inference_steps = retrieve_timesteps(
self.scheduler,
num_inference_steps,
device,
timesteps,
sigmas,
)
else:
timesteps, num_inference_steps = retrieve_timesteps(
self.scheduler,
num_inference_steps,
device,
)
num_warmup_steps = max(
len(timesteps) - num_inference_steps * self.scheduler.order, 0
)
self._num_timesteps = len(timesteps)
# 6. Denoising loop
with self.progress_bar(total=num_inference_steps) as progress_bar:
for i, t in enumerate(timesteps):
if self.interrupt:
continue
latent_model_input = (
torch.cat([latents] * 2)
if self.do_classifier_free_guidance
else latents
)
# broadcast to batch dimension in a way that's compatible with ONNX/Core ML
timestep = t.expand(latent_model_input.shape[0]).to(latents.dtype)
noise_pred = self.transformer(
hidden_states=latent_model_input,
encoder_hidden_states=prompt_embeds,
timestep=timestep,
encoder_attention_mask=prompt_attention_mask,
attention_kwargs=attention_kwargs,
return_dict=False,
)[0]
# Mochi CFG + Sampling runs in FP32
noise_pred = noise_pred.to(torch.float32)
if self.do_classifier_free_guidance:
noise_pred_uncond, noise_pred_text = noise_pred.chunk(2)
noise_pred = noise_pred_uncond + self.guidance_scale * (
noise_pred_text - noise_pred_uncond
)
# compute the previous noisy sample x_t -> x_t-1
latents_dtype = latents.dtype
latents = self.scheduler.step(
noise_pred, t, latents.to(torch.float32), return_dict=False
)[0]
latents = latents.to(latents_dtype)
if latents.dtype != latents_dtype:
if torch.backends.mps.is_available():
# some platforms (eg. apple mps) misbehave due to a pytorch bug: https://github.com/pytorch/pytorch/pull/99272
latents = latents.to(latents_dtype)
if callback_on_step_end is not None:
callback_kwargs = {}
for k in callback_on_step_end_tensor_inputs:
callback_kwargs[k] = locals()[k]
callback_outputs = callback_on_step_end(self, i, t, callback_kwargs)
latents = callback_outputs.pop("latents", latents)
prompt_embeds = callback_outputs.pop("prompt_embeds", prompt_embeds)
# call the callback, if provided
if i == len(timesteps) - 1 or (
(i + 1) > num_warmup_steps and (i + 1) % self.scheduler.order == 0
):
progress_bar.update()
if XLA_AVAILABLE:
xm.mark_step()
if get_sequence_parallel_state():
latents = all_gather(latents, dim=2)
# latents_shape = list(latents.shape)
# full_shape = [latents_shape[0] * world_size] + latents_shape[1:]
# all_latents = torch.zeros(full_shape, dtype=latents.dtype, device=latents.device)
# torch.distributed.all_gather_into_tensor(all_latents, latents)
# latents_list = list(all_latents.chunk(world_size, dim=0))
# latents = torch.cat(latents_list, dim=2)
if output_type == "latent":
video = latents
else:
# unscale/denormalize the latents
# denormalize with the mean and std if available and not None
has_latents_mean = (
hasattr(self.vae.config, "latents_mean")
and self.vae.config.latents_mean is not None
)
has_latents_std = (
hasattr(self.vae.config, "latents_std")
and self.vae.config.latents_std is not None
)
if has_latents_mean and has_latents_std:
latents_mean = (
torch.tensor(self.vae.config.latents_mean)
.view(1, 12, 1, 1, 1)
.to(latents.device, latents.dtype)
)
latents_std = (
torch.tensor(self.vae.config.latents_std)
.view(1, 12, 1, 1, 1)
.to(latents.device, latents.dtype)
)
latents = (
latents * latents_std / self.vae.config.scaling_factor
+ latents_mean
)
else:
latents = latents / self.vae.config.scaling_factor
video = self.vae.decode(latents, return_dict=False)[0]
video = self.video_processor.postprocess_video(
video, output_type=output_type
)
# Offload all models
self.maybe_free_model_hooks()
if return_all_states:
# Pay extra attention here:
# prompt_embeds with shape torch.Size([2, 256]), where prompt_embeds[1] is the prompt_embeds for the actual prompt
# prompt_embeds[0] is for negative prompt
return original_noise, video, latents, prompt_embeds, prompt_attention_mask
if not return_dict:
return (video,)
return MochiPipelineOutput(frames=video)
-133
View File
@@ -1,133 +0,0 @@
import json
import torch.distributed as dist
import torch
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
import os
from diffusers.utils import export_to_video
import argparse
def generate_video_and_latent(
pipe, prompt, height, width, num_frames, num_inference_steps, guidance_scale
):
# Set the random seed for reproducibility
generator = torch.Generator("cuda").manual_seed(12345)
# Generate videos from the input prompt
noise, video, latent, prompt_embed, prompt_attention_mask = pipe(
prompt=prompt,
height=height,
width=width,
num_frames=num_frames,
generator=generator,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
output_type="latent_and_video",
)
# prompt_embed has negative prompt at index 0
return noise[0], video[0], latent[0], prompt_embed[1], prompt_attention_mask[1]
# return dummy tensor to debug first
# return torch.zeros(1, 3, 480, 848), torch.zeros(1, 256, 16, 16)
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument("--height", type=int, default=480)
parser.add_argument("--width", type=int, default=848)
parser.add_argument("--num_inference_steps", type=int, default=64)
parser.add_argument("--guidance_scale", type=float, default=4.5)
parser.add_argument("--model_path", type=str, default="data/mochi")
parser.add_argument(
"--prompt_path", type=str, default="data/dummyVid/videos2caption.json"
)
parser.add_argument("--dataset_output_dir", type=str, default="data/dummySynthetic")
args = parser.parse_args()
local_rank = int(os.getenv("RANK", 0))
world_size = int(os.getenv("WORLD_SIZE", 1))
print("world_size", world_size, "local rank", local_rank)
torch.cuda.set_device(local_rank)
dist.init_process_group(
backend="nccl", init_method="env://", world_size=world_size, rank=local_rank
)
if not isinstance(args.prompt_path, list):
args.prompt_path = [args.prompt_path]
if len(args.prompt_path) == 1 and args.prompt_path[0].endswith("txt"):
text_prompt = open(args.prompt_path[0], "r").readlines()
text_prompt = [i.strip() for i in text_prompt]
pipe = MochiPipeline.from_pretrained(args.model_path, torch_dtype=torch.bfloat16)
pipe.enable_vae_tiling()
pipe.enable_model_cpu_offload(gpu_id=local_rank)
# make dir if not exist
os.makedirs(args.dataset_output_dir, exist_ok=True)
os.makedirs(os.path.join(args.dataset_output_dir, "noise"), exist_ok=True)
os.makedirs(os.path.join(args.dataset_output_dir, "video"), exist_ok=True)
os.makedirs(os.path.join(args.dataset_output_dir, "latent"), exist_ok=True)
os.makedirs(os.path.join(args.dataset_output_dir, "prompt_embed"), exist_ok=True)
os.makedirs(
os.path.join(args.dataset_output_dir, "prompt_attention_mask"), exist_ok=True
)
data = []
for i, prompt in enumerate(text_prompt):
if i % world_size != local_rank:
continue
(
noise,
video,
latent,
prompt_embed,
prompt_attention_mask,
) = generate_video_and_latent(
pipe,
prompt,
args.height,
args.width,
args.num_frames,
args.num_inference_steps,
args.guidance_scale,
)
# save latent
video_name = str(i)
noise_path = os.path.join(args.dataset_output_dir, "noise", video_name + ".pt")
latent_path = os.path.join(
args.dataset_output_dir, "latent", video_name + ".pt"
)
prompt_embed_path = os.path.join(
args.dataset_output_dir, "prompt_embed", video_name + ".pt"
)
video_path = os.path.join(args.dataset_output_dir, "video", video_name + ".mp4")
prompt_attention_mask_path = os.path.join(
args.dataset_output_dir, "prompt_attention_mask", video_name + ".pt"
)
# save latent
torch.save(noise, noise_path)
torch.save(latent, latent_path)
torch.save(prompt_embed, prompt_embed_path)
torch.save(prompt_attention_mask, prompt_attention_mask_path)
export_to_video(video, video_path, fps=30)
item = {}
item["cap"] = prompt
item["video"] = video_name + ".mp4"
item["noise"] = video_name + ".pt"
item["latent_path"] = video_name + ".pt"
item["prompt_embed_path"] = video_name + ".pt"
item["prompt_attention_mask"] = video_name + ".pt"
data.append(item)
dist.barrier()
local_data = data
gathered_data = [None] * world_size
dist.all_gather_object(gathered_data, local_data)
# save json
if local_rank == 0:
all_data = [item for sublist in gathered_data for item in sublist]
with open(
os.path.join(args.dataset_output_dir, "videos2caption.json"), "w"
) as f:
json.dump(all_data, f, indent=4)
-177
View File
@@ -1,177 +0,0 @@
import torch
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
import torch.distributed as dist
from diffusers.utils import export_to_video
from fastvideo.utils.parallel_states import (
initialize_sequence_parallel_state,
nccl_info,
)
import argparse
import os
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
import json
from typing import Optional
from safetensors.torch import save_file, load_file
from peft import set_peft_model_state_dict, inject_adapter_in_model, load_peft_weights
from peft import LoraConfig
import sys
import pdb
import copy
from typing import Dict
from diffusers import FlowMatchEulerDiscreteScheduler
from diffusers.utils import convert_unet_state_dict_to_peft
from fastvideo.distill.solver import PCMFMScheduler
def initialize_distributed():
local_rank = int(os.getenv("RANK", 0))
world_size = int(os.getenv("WORLD_SIZE", 1))
print("world_size", world_size)
torch.cuda.set_device(local_rank)
dist.init_process_group(
backend="nccl", init_method="env://", world_size=world_size, rank=local_rank
)
initialize_sequence_parallel_state(world_size)
def main(args):
initialize_distributed()
print(nccl_info.sp_size)
device = torch.cuda.current_device()
generator = torch.Generator(device).manual_seed(args.seed)
weight_dtype = torch.bfloat16
if args.scheduler_type == "euler":
scheduler = FlowMatchEulerDiscreteScheduler()
else:
linear_quadratic = True if "linear_quadratic" in args.scheduler_type else False
scheduler = PCMFMScheduler(
1000,
args.shift,
args.num_euler_timesteps,
linear_quadratic,
args.linear_threshold,
args.linear_range,
)
if args.transformer_path is not None:
transformer = MochiTransformer3DModel.from_pretrained(args.transformer_path)
else:
transformer = MochiTransformer3DModel.from_pretrained(
args.model_path, subfolder="transformer/"
)
pipe = MochiPipeline.from_pretrained(
args.model_path, transformer=transformer, scheduler=scheduler
)
pipe.enable_vae_tiling()
if args.lora_checkpoint_dir is not None:
print(f"Loading LoRA weights from {args.lora_checkpoint_dir}")
config_path = os.path.join(args.lora_checkpoint_dir, "lora_config.json")
with open(config_path, "r") as f:
lora_config_dict = json.load(f)
rank = lora_config_dict["lora_params"]["lora_rank"]
lora_alpha = lora_config_dict["lora_params"]["lora_alpha"]
lora_scaling = lora_alpha / rank
pipe.load_lora_weights(args.lora_checkpoint_dir, adapter_name="default")
pipe.set_adapters(["default"], [lora_scaling])
print(f"Successfully Loaded LoRA weights from {args.lora_checkpoint_dir}")
# pipe.to(device)
pipe.enable_model_cpu_offload(device)
# Generate videos from the input prompt
if args.prompt_embed_path is not None:
prompt_embeds = (
torch.load(args.prompt_embed_path, map_location="cpu", weights_only=True)
.to(device)
.unsqueeze(0)
)
encoder_attention_mask = (
torch.load(
args.encoder_attention_mask_path, map_location="cpu", weights_only=True
)
.to(device)
.unsqueeze(0)
)
prompts = None
elif args.prompt_path is not None:
prompts = [line.strip() for line in open(args.prompt_path, "r")]
prompt_embeds = None
encoder_attention_mask = None
else:
prompts = args.prompts
prompt_embeds = None
encoder_attention_mask = None
if prompts is not None:
videos = []
with torch.autocast("cuda", dtype=torch.bfloat16):
for prompt in prompts:
video = pipe(
prompt=[prompt],
height=args.height,
width=args.width,
num_frames=args.num_frames,
num_inference_steps=args.num_inference_steps,
guidance_scale=args.guidance_scale,
generator=generator,
).frames
videos.append(video[0])
else:
with torch.autocast("cuda", dtype=torch.bfloat16):
videos = pipe(
prompt_embeds=prompt_embeds,
prompt_attention_mask=encoder_attention_mask,
height=args.height,
width=args.width,
num_frames=args.num_frames,
num_inference_steps=args.num_inference_steps,
guidance_scale=args.guidance_scale,
generator=generator,
).frames
if nccl_info.global_rank <= 0:
if prompts is not None:
# mkdir
os.makedirs(args.output_path, exist_ok=True)
for video, prompt in zip(videos, prompts):
suffix = prompt.split(".")[0]
export_to_video(
video, os.path.join(args.output_path, f"{suffix}.mp4"), fps=30
)
else:
export_to_video(videos[0], args.output_path + ".mp4", fps=30)
if __name__ == "__main__":
# arg parse
parser = argparse.ArgumentParser()
parser.add_argument("--prompts", nargs="+", default=[])
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument("--height", type=int, default=480)
parser.add_argument("--width", type=int, default=848)
parser.add_argument("--num_inference_steps", type=int, default=64)
parser.add_argument("--guidance_scale", type=float, default=4.5)
parser.add_argument("--model_path", type=str, default="data/mochi")
parser.add_argument("--seed", type=int, default=42)
parser.add_argument("--output_path", type=str, default="./outputs.mp4")
parser.add_argument("--transformer_path", type=str, default=None)
parser.add_argument("--prompt_embed_path", type=str, default=None)
parser.add_argument("--prompt_path", type=str, default=None)
parser.add_argument("--scheduler_type", type=str, default="euler")
parser.add_argument("--encoder_attention_mask_path", type=str, default=None)
parser.add_argument(
"--lora_checkpoint_dir",
type=str,
default=None,
help="Path to the directory containing LoRA checkpoints",
)
parser.add_argument("--shift", type=float, default=8.0)
parser.add_argument("--num_euler_timesteps", type=int, default=100)
parser.add_argument("--linear_threshold", type=float, default=0.025)
parser.add_argument("--linear_range", type=float, default=0.5)
args = parser.parse_args()
main(args)
@@ -1,57 +0,0 @@
import torch
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
from diffusers.utils import export_to_video, load_image, load_video
import argparse
from diffusers import FlowMatchEulerDiscreteScheduler
def main(args):
# Set the random seed for reproducibility
generator = torch.Generator("cuda").manual_seed(args.seed)
# do not invert
scheduler = FlowMatchEulerDiscreteScheduler()
if args.transformer_path is not None:
transformer = MochiTransformer3DModel.from_pretrained(args.transformer_path)
else:
transformer = MochiTransformer3DModel.from_pretrained(
args.model_path, subfolder="transformer/"
)
pipe = MochiPipeline.from_pretrained(
args.model_path, transformer=transformer, scheduler=scheduler
)
pipe.enable_vae_tiling()
# pipe.to("cuda:1")
pipe.enable_model_cpu_offload()
# Generate videos from the input prompt
with torch.autocast("cuda", dtype=torch.bfloat16):
videos = pipe(
prompt=args.prompts,
height=args.height,
width=args.width,
num_frames=args.num_frames,
generator=generator,
num_inference_steps=args.num_inference_steps,
guidance_scale=args.guidance_scale,
).frames
for prompt, video in zip(args.prompts, videos):
export_to_video(video, args.output_path + f"_{prompt}.mp4", fps=30)
if __name__ == "__main__":
# arg parse
parser = argparse.ArgumentParser()
parser.add_argument("--prompts", nargs="+", default=[])
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument("--height", type=int, default=480)
parser.add_argument("--width", type=int, default=848)
parser.add_argument("--num_inference_steps", type=int, default=64)
parser.add_argument("--guidance_scale", type=float, default=4.5)
parser.add_argument("--model_path", type=str, default="data/mochi")
parser.add_argument("--seed", type=int, default=12345)
parser.add_argument("--transformer_path", type=str, default=None)
parser.add_argument("--output_path", type=str, default="./outputs.mp4")
args = parser.parse_args()
main(args)
-733
View File
@@ -1,733 +0,0 @@
import argparse
from email.policy import strict
import logging
import math
import os
import shutil
from pathlib import Path
from fastvideo.utils.parallel_states import (
initialize_sequence_parallel_state,
destroy_sequence_parallel_group,
get_sequence_parallel_state,
nccl_info,
)
from fastvideo.utils.communications import sp_parallel_dataloader_wrapper, broadcast
from fastvideo.models.mochi_hf.mochi_latents_utils import normalize_mochi_dit_input
from fastvideo.utils.validation import log_validation
import time
from torch.utils.data import DataLoader
import torch
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
StateDictType,
FullStateDictConfig,
)
import json
from torch.utils.data.distributed import DistributedSampler
from fastvideo.utils.dataset_utils import LengthGroupedSampler
import wandb
from accelerate.utils import set_seed
from tqdm.auto import tqdm
from fastvideo.fsdp_util import get_dit_fsdp_kwargs, apply_fsdp_checkpointing
from diffusers.utils import convert_unet_state_dict_to_peft
from diffusers import (
FlowMatchEulerDiscreteScheduler,
)
from diffusers.optimization import get_scheduler
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
from diffusers.utils import check_min_version
from fastvideo.dataset.latent_datasets import LatentDataset, latent_collate_function
import torch.distributed as dist
from safetensors.torch import save_file, load_file
from peft import LoraConfig, get_peft_model_state_dict, set_peft_model_state_dict
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
)
from fastvideo.utils.checkpoint import (
save_checkpoint,
save_lora_checkpoint,
resume_lora_optimizer,
)
from fastvideo.utils.logging import main_print
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
# Will error if the minimal version of diffusers is not installed. Remove at your own risks.
check_min_version("0.31.0")
import time
from collections import deque
def compute_density_for_timestep_sampling(
weighting_scheme: str,
batch_size: int,
generator,
logit_mean: float = None,
logit_std: float = None,
mode_scale: float = None,
):
"""
Compute the density for sampling the timesteps when doing SD3 training.
Courtesy: This was contributed by Rafie Walker in https://github.com/huggingface/diffusers/pull/8528.
SD3 paper reference: https://arxiv.org/abs/2403.03206v1.
"""
if weighting_scheme == "logit_normal":
# See 3.1 in the SD3 paper ($rf/lognorm(0.00,1.00)$).
u = torch.normal(
mean=logit_mean,
std=logit_std,
size=(batch_size,),
device="cpu",
generator=generator,
)
u = torch.nn.functional.sigmoid(u)
elif weighting_scheme == "mode":
u = torch.rand(size=(batch_size,), device="cpu", generator=generator)
u = 1 - u - mode_scale * (torch.cos(math.pi * u / 2) ** 2 - 1 + u)
else:
u = torch.rand(size=(batch_size,), device="cpu", generator=generator)
return u
def get_sigmas(noise_scheduler, device, timesteps, n_dim=4, dtype=torch.float32):
sigmas = noise_scheduler.sigmas.to(device=device, dtype=dtype)
schedule_timesteps = noise_scheduler.timesteps.to(device)
timesteps = timesteps.to(device)
step_indices = [(schedule_timesteps == t).nonzero().item() for t in timesteps]
sigma = sigmas[step_indices].flatten()
while len(sigma.shape) < n_dim:
sigma = sigma.unsqueeze(-1)
return sigma
def train_one_step_mochi(
transformer,
optimizer,
lr_scheduler,
loader,
noise_scheduler,
noise_random_generator,
gradient_accumulation_steps,
sp_size,
precondition_outputs,
max_grad_norm,
weighting_scheme,
logit_mean,
logit_std,
mode_scale,
):
total_loss = 0.0
optimizer.zero_grad()
for _ in range(gradient_accumulation_steps):
(
latents,
encoder_hidden_states,
latents_attention_mask,
encoder_attention_mask,
) = next(loader)
latents = normalize_mochi_dit_input(latents)
batch_size = latents.shape[0]
noise = torch.randn_like(latents)
u = compute_density_for_timestep_sampling(
weighting_scheme=weighting_scheme,
batch_size=batch_size,
generator=noise_random_generator,
logit_mean=logit_mean,
logit_std=logit_std,
mode_scale=mode_scale,
)
indices = (u * noise_scheduler.config.num_train_timesteps).long()
timesteps = noise_scheduler.timesteps[indices].to(device=latents.device)
if sp_size > 1:
# Make sure that the timesteps are the same across all sp processes.
broadcast(timesteps)
sigmas = get_sigmas(
noise_scheduler,
latents.device,
timesteps,
n_dim=latents.ndim,
dtype=latents.dtype,
)
noisy_model_input = (1.0 - sigmas) * latents + sigmas * noise
# if rank<=0:
# print("2222222222222222222222222222222222222222222222")
# print(type(latents_attention_mask))
# print(latents_attention_mask)
with torch.autocast("cuda", torch.bfloat16):
model_pred = transformer(
noisy_model_input,
encoder_hidden_states,
timesteps,
encoder_attention_mask, # B, L
return_dict=False,
)[0]
# if rank<=0:
# print("333333333333333333333333333333333333333333333333")
if precondition_outputs:
model_pred = noisy_model_input - model_pred * sigmas
if precondition_outputs:
target = latents
else:
target = noise - latents
loss = (
torch.mean((model_pred.float() - target.float()) ** 2)
/ gradient_accumulation_steps
)
loss.backward()
avg_loss = loss.detach().clone()
dist.all_reduce(avg_loss, op=dist.ReduceOp.AVG)
total_loss += avg_loss.item()
grad_norm = transformer.clip_grad_norm_(max_grad_norm)
optimizer.step()
lr_scheduler.step()
return total_loss, grad_norm.item()
def main(args):
torch.backends.cuda.matmul.allow_tf32 = True
local_rank = int(os.environ["LOCAL_RANK"])
rank = int(os.environ["RANK"])
world_size = int(os.environ["WORLD_SIZE"])
dist.init_process_group("nccl")
torch.cuda.set_device(local_rank)
device = torch.cuda.current_device()
initialize_sequence_parallel_state(args.sp_size)
# If passed along, set the training seed now. On GPU...
if args.seed is not None:
# TODO: t within the same seq parallel group should be the same. Noise should be different.
set_seed(args.seed + rank)
# We use different seeds for the noise generation in each process to ensure that the noise is different in a batch.
noise_random_generator = None
# Handle the repository creation
if rank <= 0 and args.output_dir is not None:
os.makedirs(args.output_dir, exist_ok=True)
# For mixed precision training we cast all non-trainable weigths to half-precision
# as these weights are only used for inference, keeping weights in full precision is not required.
# Create model:
main_print(f"--> loading model from {args.pretrained_model_name_or_path}")
# keep the master weight to float32
transformer = MochiTransformer3DModel.from_pretrained(
args.pretrained_model_name_or_path,
subfolder="transformer",
torch_dtype=torch.float32
if args.master_weight_type == "fp32"
else torch.bfloat16,
)
if args.use_lora:
transformer.requires_grad_(False)
transformer_lora_config = LoraConfig(
r=args.lora_rank,
lora_alpha=args.lora_alpha,
init_lora_weights=True,
target_modules=["to_k", "to_q", "to_v", "to_out.0"],
)
transformer.add_adapter(transformer_lora_config)
if args.resume_from_lora_checkpoint:
lora_state_dict = MochiPipeline.lora_state_dict(
args.resume_from_lora_checkpoint
)
transformer_state_dict = {
f'{k.replace("transformer.", "")}': v
for k, v in lora_state_dict.items()
if k.startswith("transformer.")
}
transformer_state_dict = convert_unet_state_dict_to_peft(transformer_state_dict)
incompatible_keys = set_peft_model_state_dict(
transformer, transformer_state_dict, adapter_name="default"
)
if incompatible_keys is not None:
# check only for unexpected keys
unexpected_keys = getattr(incompatible_keys, "unexpected_keys", None)
if unexpected_keys:
main_print(
f"Loading adapter weights from state_dict led to unexpected keys not found in the model: "
f" {unexpected_keys}. "
)
main_print(
f" Total training parameters = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e6} M"
)
main_print(
f"--> Initializing FSDP with sharding strategy: {args.fsdp_sharding_startegy}"
)
fsdp_kwargs = get_dit_fsdp_kwargs(
args.fsdp_sharding_startegy,
args.use_lora,
args.use_cpu_offload,
args.master_weight_type,
)
if args.use_lora:
transformer.config.lora_rank = args.lora_rank
transformer.config.lora_alpha = args.lora_alpha
transformer.config.lora_target_modules = ["to_k", "to_q", "to_v", "to_out.0"]
transformer._no_split_modules = ["MochiTransformerBlock"]
fsdp_kwargs["auto_wrap_policy"] = fsdp_kwargs["auto_wrap_policy"](transformer)
transformer = FSDP(
transformer,
**fsdp_kwargs,
)
main_print(f"--> model loaded")
if args.gradient_checkpointing:
apply_fsdp_checkpointing(transformer, args.selective_checkpointing)
# Set model as trainable.
transformer.train()
noise_scheduler = FlowMatchEulerDiscreteScheduler()
params_to_optimize = transformer.parameters()
params_to_optimize = list(filter(lambda p: p.requires_grad, params_to_optimize))
optimizer = torch.optim.AdamW(
params_to_optimize,
lr=args.learning_rate,
betas=(0.9, 0.999),
weight_decay=args.weight_decay,
eps=1e-8,
)
init_steps = 0
if args.resume_from_lora_checkpoint:
transformer, optimizer, init_steps = resume_lora_optimizer(
transformer, args.resume_from_lora_checkpoint, optimizer
)
main_print(f"optimizer: {optimizer}")
lr_scheduler = get_scheduler(
args.lr_scheduler,
optimizer=optimizer,
num_warmup_steps=args.lr_warmup_steps,
num_training_steps=args.max_train_steps,
num_cycles=args.lr_num_cycles,
power=args.lr_power,
last_epoch=init_steps - 1,
)
train_dataset = LatentDataset(args.data_json_path, args.num_latent_t, args.cfg)
sampler = (
LengthGroupedSampler(
args.train_batch_size,
rank=rank,
world_size=world_size,
lengths=train_dataset.lengths,
group_frame=args.group_frame,
group_resolution=args.group_resolution,
)
if (args.group_frame or args.group_resolution)
else DistributedSampler(
train_dataset, rank=rank, num_replicas=world_size, shuffle=False
)
)
train_dataloader = DataLoader(
train_dataset,
sampler=sampler,
collate_fn=latent_collate_function,
pin_memory=True,
batch_size=args.train_batch_size,
num_workers=args.dataloader_num_workers,
drop_last=True,
)
num_update_steps_per_epoch = math.ceil(
len(train_dataloader)
/ args.gradient_accumulation_steps
* args.sp_size
/ args.train_sp_batch_size
)
args.num_train_epochs = math.ceil(args.max_train_steps / num_update_steps_per_epoch)
if rank <= 0:
project = args.tracker_project_name or "fastvideo"
wandb.init(project=project, config=args)
# Train!
total_batch_size = (
args.train_batch_size
* world_size
* args.gradient_accumulation_steps
/ args.sp_size
* args.train_sp_batch_size
)
main_print("***** Running training *****")
main_print(f" Num examples = {len(train_dataset)}")
main_print(f" Dataloader size = {len(train_dataloader)}")
main_print(f" Num Epochs = {args.num_train_epochs}")
main_print(f" Resume training from step {init_steps}")
main_print(f" Instantaneous batch size per device = {args.train_batch_size}")
main_print(
f" Total train batch size (w. data & sequence parallel, accumulation) = {total_batch_size}"
)
main_print(f" Gradient Accumulation steps = {args.gradient_accumulation_steps}")
main_print(f" Total optimization steps = {args.max_train_steps}")
main_print(
f" Total training parameters per FSDP shard = {sum(p.numel() for p in transformer.parameters() if p.requires_grad) / 1e9} B"
)
# print dtype
main_print(f" Master weight dtype: {transformer.parameters().__next__().dtype}")
# Potentially load in the weights and states from a previous save
if args.resume_from_checkpoint:
assert NotImplementedError("resume_from_checkpoint is not supported now.")
# TODO
progress_bar = tqdm(
range(0, args.max_train_steps),
initial=init_steps,
desc="Steps",
# Only show the progress bar once on each machine.
disable=local_rank > 0,
)
loader = sp_parallel_dataloader_wrapper(
train_dataloader,
device,
args.train_batch_size,
args.sp_size,
args.train_sp_batch_size,
)
step_times = deque(maxlen=100)
# todo future
for i in range(init_steps):
next(loader)
for step in range(init_steps + 1, args.max_train_steps + 1):
start_time = time.time()
loss, grad_norm = train_one_step_mochi(
transformer,
optimizer,
lr_scheduler,
loader,
noise_scheduler,
noise_random_generator,
args.gradient_accumulation_steps,
args.sp_size,
args.precondition_outputs,
args.max_grad_norm,
args.weighting_scheme,
args.logit_mean,
args.logit_std,
args.mode_scale,
)
step_time = time.time() - start_time
step_times.append(step_time)
avg_step_time = sum(step_times) / len(step_times)
progress_bar.set_postfix(
{
"loss": f"{loss:.4f}",
"step_time": f"{step_time:.2f}s",
"grad_norm": grad_norm,
}
)
progress_bar.update(1)
if rank <= 0:
wandb.log(
{
"train_loss": loss,
"learning_rate": lr_scheduler.get_last_lr()[0],
"step_time": step_time,
"avg_step_time": avg_step_time,
"grad_norm": grad_norm,
},
step=step,
)
if step % args.checkpointing_steps == 0:
if args.use_lora:
# Save LoRA weights
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, step
)
else:
# Your existing checkpoint saving code
save_checkpoint(transformer, optimizer, rank, args.output_dir, step)
dist.barrier()
if args.log_validation and step % args.validation_steps == 0:
log_validation(args, transformer, device, torch.bfloat16, step)
if args.use_lora:
save_lora_checkpoint(
transformer, optimizer, rank, args.output_dir, args.max_train_steps
)
else:
save_checkpoint(
transformer, optimizer, rank, args.output_dir, args.max_train_steps
)
if get_sequence_parallel_state():
destroy_sequence_parallel_group()
if __name__ == "__main__":
parser = argparse.ArgumentParser()
# dataset & dataloader
parser.add_argument("--data_json_path", type=str, required=True)
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument(
"--dataloader_num_workers",
type=int,
default=10,
help="Number of subprocesses to use for data loading. 0 means that the data will be loaded in the main process.",
)
parser.add_argument(
"--train_batch_size",
type=int,
default=16,
help="Batch size (per device) for the training dataloader.",
)
parser.add_argument(
"--num_latent_t", type=int, default=28, help="Number of latent timesteps."
)
parser.add_argument("--group_frame", action="store_true") # TODO
parser.add_argument("--group_resolution", action="store_true") # TODO
# text encoder & vae & diffusion model
parser.add_argument("--pretrained_model_name_or_path", type=str)
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
# diffusion setting
parser.add_argument("--ema_decay", type=float, default=0.999)
parser.add_argument("--ema_start_step", type=int, default=0)
parser.add_argument("--cfg", type=float, default=0.1)
parser.add_argument(
"--precondition_outputs",
action="store_true",
help="Whether to precondition the outputs of the model.",
)
# validation & logs
parser.add_argument("--validation_prompt_dir", type=str)
parser.add_argument("--uncond_prompt_dir", type=str)
parser.add_argument(
"--validation_sampling_steps",
type=str,
default="64",
help="use ',' to split multi sampling steps",
)
parser.add_argument(
"--validation_guidance_scale",
type=str,
default="4.5",
help="use ',' to split multi scale",
)
parser.add_argument("--validation_steps", type=int, default=50)
parser.add_argument("--log_validation", action="store_true")
parser.add_argument("--tracker_project_name", type=str, default=None)
parser.add_argument(
"--seed", type=int, default=None, help="A seed for reproducible training."
)
parser.add_argument(
"--output_dir",
type=str,
default=None,
help="The output directory where the model predictions and checkpoints will be written.",
)
parser.add_argument(
"--checkpoints_total_limit",
type=int,
default=None,
help=("Max number of checkpoints to store."),
)
parser.add_argument(
"--checkpointing_steps",
type=int,
default=500,
help=(
"Save a checkpoint of the training state every X updates. These checkpoints can be used both as final"
" checkpoints in case they are better than the last checkpoint, and are also suitable for resuming"
" training using `--resume_from_checkpoint`."
),
)
parser.add_argument(
"--resume_from_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--resume_from_lora_checkpoint",
type=str,
default=None,
help=(
"Whether training should be resumed from a previous lora checkpoint. Use a path saved by"
' `--checkpointing_steps`, or `"latest"` to automatically select the last available checkpoint.'
),
)
parser.add_argument(
"--logging_dir",
type=str,
default="logs",
help=(
"[TensorBoard](https://www.tensorflow.org/tensorboard) log directory. Will default to"
" *output_dir/runs/**CURRENT_DATETIME_HOSTNAME***."
),
)
# optimizer & scheduler & Training
parser.add_argument("--num_train_epochs", type=int, default=100)
parser.add_argument(
"--max_train_steps",
type=int,
default=None,
help="Total number of training steps to perform. If provided, overrides num_train_epochs.",
)
parser.add_argument(
"--gradient_accumulation_steps",
type=int,
default=1,
help="Number of updates steps to accumulate before performing a backward/update pass.",
)
parser.add_argument(
"--learning_rate",
type=float,
default=1e-4,
help="Initial learning rate (after the potential warmup period) to use.",
)
parser.add_argument(
"--scale_lr",
action="store_true",
default=False,
help="Scale the learning rate by the number of GPUs, gradient accumulation steps, and batch size.",
)
parser.add_argument(
"--lr_warmup_steps",
type=int,
default=10,
help="Number of steps for the warmup in the lr scheduler.",
)
parser.add_argument(
"--max_grad_norm", default=1.0, type=float, help="Max gradient norm."
)
parser.add_argument(
"--gradient_checkpointing",
action="store_true",
help="Whether or not to use gradient checkpointing to save memory at the expense of slower backward pass.",
)
parser.add_argument("--selective_checkpointing", type=float, default=1.0)
parser.add_argument(
"--allow_tf32",
action="store_true",
help=(
"Whether or not to allow TF32 on Ampere GPUs. Can be used to speed up training. For more information, see"
" https://pytorch.org/docs/stable/notes/cuda.html#tensorfloat-32-tf32-on-ampere-devices"
),
)
parser.add_argument(
"--mixed_precision",
type=str,
default=None,
choices=["no", "fp16", "bf16"],
help=(
"Whether to use mixed precision. Choose between fp16 and bf16 (bfloat16). Bf16 requires PyTorch >="
" 1.10.and an Nvidia Ampere GPU. Default to the value of accelerate config of the current system or the"
" flag passed with the `accelerate.launch` command. Use this argument to override the accelerate config."
),
)
parser.add_argument(
"--use_cpu_offload",
action="store_true",
help="Whether to use CPU offload for param & gradient & optimizer states.",
)
parser.add_argument("--sp_size", type=int, default=1, help="For sequence parallel")
parser.add_argument(
"--train_sp_batch_size",
type=int,
default=1,
help="Batch size for sequence parallel training",
)
parser.add_argument(
"--use_lora",
action="store_true",
default=False,
help="Whether to use LoRA for finetuning.",
)
parser.add_argument(
"--lora_alpha", type=int, default=256, help="Alpha parameter for LoRA."
)
parser.add_argument(
"--lora_rank", type=int, default=128, help="LoRA rank parameter. "
)
parser.add_argument("--fsdp_sharding_startegy", default="full")
parser.add_argument(
"--weighting_scheme",
type=str,
default="uniform",
choices=["sigma_sqrt", "logit_normal", "mode", "cosmap", "uniform"],
)
parser.add_argument(
"--logit_mean",
type=float,
default=0.0,
help="mean to use when using the `'logit_normal'` weighting scheme.",
)
parser.add_argument(
"--logit_std",
type=float,
default=1.0,
help="std to use when using the `'logit_normal'` weighting scheme.",
)
parser.add_argument(
"--mode_scale",
type=float,
default=1.29,
help="Scale of mode weighting scheme. Only effective when using the `'mode'` as the `weighting_scheme`.",
)
# lr_scheduler
parser.add_argument(
"--lr_scheduler",
type=str,
default="constant",
help=(
'The scheduler type to use. Choose between ["linear", "cosine", "cosine_with_restarts", "polynomial",'
' "constant", "constant_with_warmup"]'
),
)
parser.add_argument(
"--lr_num_cycles",
type=int,
default=1,
help="Number of cycles in the learning rate scheduler.",
)
parser.add_argument(
"--lr_power",
type=float,
default=1.0,
help="Power factor of the polynomial scheduler.",
)
parser.add_argument(
"--weight_decay", type=float, default=0.01, help="Weight decay to apply."
)
parser.add_argument(
"--master_weight_type",
type=str,
default="fp32",
help="Weight type to use - fp32 or bf16.",
)
args = parser.parse_args()
main(args)
-282
View File
@@ -1,282 +0,0 @@
# import
import os
import json
import torch
from fastvideo.utils.logging import main_print
from torch.distributed.fsdp import (
FullyShardedDataParallel as FSDP,
StateDictType,
FullStateDictConfig,
)
from safetensors.torch import save_file, load_file
import torch.distributed.checkpoint as dist_cp
from torch.distributed.checkpoint.default_planner import (
DefaultSavePlanner,
DefaultLoadPlanner,
)
from torch.distributed.checkpoint.optimizer import load_sharded_optimizer_state_dict
from torch.distributed.fsdp import FullOptimStateDictConfig
from peft import LoraConfig, get_peft_model_state_dict, set_peft_model_state_dict
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
def save_checkpoint(model, optimizer, rank, output_dir, step, discriminator=False):
with FSDP.state_dict_type(
model,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
FullOptimStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
cpu_state = model.state_dict()
optim_state = FSDP.optim_state_dict(
model,
optimizer,
)
# todo move to get_state_dict
save_dir = os.path.join(output_dir, f"checkpoint-{step}")
os.makedirs(save_dir, exist_ok=True)
# save using safetensors
if rank <= 0 and not discriminator:
weight_path = os.path.join(save_dir, "diffusion_pytorch_model.safetensors")
save_file(cpu_state, weight_path)
config_dict = dict(model.config)
config_path = os.path.join(save_dir, "config.json")
# save dict as json
with open(config_path, "w") as f:
json.dump(config_dict, f, indent=4)
optimizer_path = os.path.join(save_dir, "optimizer.pt")
torch.save(optim_state, optimizer_path)
else:
weight_path = os.path.join(save_dir, "discriminator_pytorch_model.safetensors")
save_file(cpu_state, weight_path)
optimizer_path = os.path.join(save_dir, "discriminator_optimizer.pt")
torch.save(optim_state, optimizer_path)
def save_checkpoint_generator_discriminator(
model,
optimizer,
discriminator,
discriminator_optimizer,
rank,
output_dir,
step,
):
with FSDP.state_dict_type(
model,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
cpu_state = model.state_dict()
# todo move to get_state_dict
save_dir = os.path.join(output_dir, f"checkpoint-{step}")
os.makedirs(save_dir, exist_ok=True)
hf_weight_dir = os.path.join(save_dir, "hf_weights")
os.makedirs(hf_weight_dir, exist_ok=True)
# save using safetensors
if rank <= 0:
config_dict = dict(model.config)
config_path = os.path.join(hf_weight_dir, "config.json")
# save dict as json
with open(config_path, "w") as f:
json.dump(config_dict, f, indent=4)
weight_path = os.path.join(hf_weight_dir, "diffusion_pytorch_model.safetensors")
save_file(cpu_state, weight_path)
main_print(f"--> saved HF weight checkpoint at path {hf_weight_dir}")
model_weight_dir = os.path.join(save_dir, "model_weights_state")
os.makedirs(model_weight_dir, exist_ok=True)
model_optimizer_dir = os.path.join(save_dir, "model_optimizer_state")
os.makedirs(model_optimizer_dir, exist_ok=True)
with FSDP.state_dict_type(model, StateDictType.SHARDED_STATE_DICT):
optim_state = FSDP.optim_state_dict(model, optimizer)
model_state = model.state_dict()
weight_state_dict = {"model": model_state}
dist_cp.save_state_dict(
state_dict=weight_state_dict,
storage_writer=dist_cp.FileSystemWriter(model_weight_dir),
planner=DefaultSavePlanner(),
)
optimizer_state_dict = {"optimizer": optim_state}
dist_cp.save_state_dict(
state_dict=optimizer_state_dict,
storage_writer=dist_cp.FileSystemWriter(model_optimizer_dir),
planner=DefaultSavePlanner(),
)
discriminator_fsdp_state_dir = os.path.join(save_dir, "discriminator_fsdp_state")
os.makedirs(discriminator_fsdp_state_dir, exist_ok=True)
with FSDP.state_dict_type(
discriminator,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
FullOptimStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
optim_state = FSDP.optim_state_dict(discriminator, discriminator_optimizer)
model_state = discriminator.state_dict()
state_dict = {"optimizer": optim_state, "model": model_state}
if rank <= 0:
discriminator_fsdp_state_fil = os.path.join(
discriminator_fsdp_state_dir, "discriminator_state.pt"
)
torch.save(state_dict, discriminator_fsdp_state_fil)
main_print("--> saved FSDP state checkpoint")
def load_sharded_model(model, optimizer, model_dir, optimizer_dir):
with FSDP.state_dict_type(model, StateDictType.SHARDED_STATE_DICT):
weight_state_dict = {"model": model.state_dict()}
optim_state = load_sharded_optimizer_state_dict(
model_state_dict=weight_state_dict["model"],
optimizer_key="optimizer",
storage_reader=dist_cp.FileSystemReader(optimizer_dir),
)
optim_state = optim_state["optimizer"]
flattened_osd = FSDP.optim_state_dict_to_load(
model=model, optim=optimizer, optim_state_dict=optim_state
)
optimizer.load_state_dict(flattened_osd)
dist_cp.load_state_dict(
state_dict=weight_state_dict,
storage_reader=dist_cp.FileSystemReader(model_dir),
planner=DefaultLoadPlanner(),
)
model_state = weight_state_dict["model"]
model.load_state_dict(model_state)
main_print(f"--> loaded model and optimizer from path {model_dir}")
return model, optimizer
def load_full_state_model(model, optimizer, checkpoint_file, rank):
with FSDP.state_dict_type(
model,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
FullOptimStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
discriminator_state = torch.load(checkpoint_file)
model_state = discriminator_state["model"]
if rank <= 0:
optim_state = discriminator_state["optimizer"]
else:
optim_state = None
model.load_state_dict(model_state)
discriminator_optim_state = FSDP.optim_state_dict_to_load(
model=model, optim=optimizer, optim_state_dict=optim_state
)
optimizer.load_state_dict(discriminator_optim_state)
main_print(
f"--> loaded discriminator and discriminator optimizer from path {checkpoint_file}"
)
return model, optimizer
def resume_training_generator_discriminator(
model, optimizer, discriminator, discriminator_optimizer, checkpoint_dir, rank
):
step = int(checkpoint_dir.split("-")[-1])
model_weight_dir = os.path.join(checkpoint_dir, "model_weights_state")
model_optimizer_dir = os.path.join(checkpoint_dir, "model_optimizer_state")
model, optimizer = load_sharded_model(
model, optimizer, model_weight_dir, model_optimizer_dir
)
discriminator_ckpt_file = os.path.join(
checkpoint_dir, "discriminator_fsdp_state", "discriminator_state.pt"
)
discriminator, discriminator_optimizer = load_full_state_model(
discriminator, discriminator_optimizer, discriminator_ckpt_file, rank
)
return model, optimizer, discriminator, discriminator_optimizer, step
def resume_training(model, optimizer, checkpoint_dir, discriminator=False):
weight_path = os.path.join(checkpoint_dir, "diffusion_pytorch_model.safetensors")
if discriminator:
weight_path = os.path.join(
checkpoint_dir, "discriminator_pytorch_model.safetensors"
)
model_weights = load_file(weight_path)
with FSDP.state_dict_type(
model,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
FullOptimStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
current_state = model.state_dict()
current_state.update(model_weights)
model.load_state_dict(current_state, strict=False)
if discriminator:
optim_path = os.path.join(checkpoint_dir, "discriminator_optimizer.pt")
else:
optim_path = os.path.join(checkpoint_dir, "optimizer.pt")
optimizer_state_dict = torch.load(optim_path, weights_only=False)
optim_state = FSDP.optim_state_dict_to_load(
model=model, optim=optimizer, optim_state_dict=optimizer_state_dict
)
optimizer.load_state_dict(optim_state)
step = int(checkpoint_dir.split("-")[-1])
return model, optimizer, step
def save_lora_checkpoint(transformer, optimizer, rank, output_dir, step):
with FSDP.state_dict_type(
transformer,
StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
):
full_state_dict = transformer.state_dict()
lora_optim_state = FSDP.optim_state_dict(
transformer,
optimizer,
)
if rank <= 0:
save_dir = os.path.join(output_dir, f"lora-checkpoint-{step}")
os.makedirs(save_dir, exist_ok=True)
# save optimizer
optim_path = os.path.join(save_dir, "lora_optimizer.pt")
torch.save(lora_optim_state, optim_path)
# save lora weight
main_print(f"--> saving LoRA checkpoint at step {step}")
transformer_lora_layers = get_peft_model_state_dict(
model=transformer, state_dict=full_state_dict
)
MochiPipeline.save_lora_weights(
save_directory=save_dir,
transformer_lora_layers=transformer_lora_layers,
is_main_process=True,
)
# save config
lora_config = {
"step": step,
"lora_params": {
"lora_rank": transformer.config.lora_rank,
"lora_alpha": transformer.config.lora_alpha,
"target_modules": transformer.config.lora_target_modules,
},
}
config_path = os.path.join(save_dir, "lora_config.json")
with open(config_path, "w") as f:
json.dump(lora_config, f, indent=4)
main_print(f"--> LoRA checkpoint saved at step {step}")
def resume_lora_optimizer(transformer, checkpoint_dir, optimizer):
config_path = os.path.join(checkpoint_dir, "lora_config.json")
with open(config_path, "r") as f:
config_dict = json.load(f)
optim_path = os.path.join(checkpoint_dir, "lora_optimizer.pt")
optimizer_state_dict = torch.load(optim_path, weights_only=False)
optim_state = FSDP.optim_state_dict_to_load(
model=transformer, optim=optimizer, optim_state_dict=optimizer_state_dict
)
optimizer.load_state_dict(optim_state)
step = config_dict["step"]
main_print(f"--> Successfully resuming LoRA optimizer from step {step}")
return transformer, optimizer, step
-334
View File
@@ -1,334 +0,0 @@
# Copyright (c) Microsoft Corporation.
# SPDX-License-Identifier: Apache-2.0
# DeepSpeed Team
import torch
import torch.distributed as dist
from fastvideo.utils.parallel_states import nccl_info
from typing import Any, Tuple
from torch import Tensor
from torch.nn import Module
def broadcast(input_: torch.Tensor):
src = nccl_info.group_id * nccl_info.sp_size
dist.broadcast(input_, src=src, group=nccl_info.group)
def _all_to_all_4D(
input: torch.tensor, scatter_idx: int = 2, gather_idx: int = 1, group=None
) -> torch.tensor:
"""
all-to-all for QKV
Args:
input (torch.tensor): a tensor sharded along dim scatter dim
scatter_idx (int): default 1
gather_idx (int): default 2
group : torch process group
Returns:
torch.tensor: resharded tensor (bs, seqlen/P, hc, hs)
"""
assert (
input.dim() == 4
), f"input must be 4D tensor, got {input.dim()} and shape {input.shape}"
seq_world_size = dist.get_world_size(group)
if scatter_idx == 2 and gather_idx == 1:
# input (torch.tensor): a tensor sharded along dim 1 (bs, seqlen/P, hc, hs) output: (bs, seqlen, hc/P, hs)
bs, shard_seqlen, hc, hs = input.shape
seqlen = shard_seqlen * seq_world_size
shard_hc = hc // seq_world_size
# transpose groups of heads with the seq-len parallel dimension, so that we can scatter them!
# (bs, seqlen/P, hc, hs) -reshape-> (bs, seq_len/P, P, hc/P, hs) -transpose(0,2)-> (P, seq_len/P, bs, hc/P, hs)
input_t = (
input.reshape(bs, shard_seqlen, seq_world_size, shard_hc, hs)
.transpose(0, 2)
.contiguous()
)
output = torch.empty_like(input_t)
# https://pytorch.org/docs/stable/distributed.html#torch.distributed.all_to_all_single
# (P, seq_len/P, bs, hc/P, hs) scatter seqlen -all2all-> (P, seq_len/P, bs, hc/P, hs) scatter head
if seq_world_size > 1:
dist.all_to_all_single(output, input_t, group=group)
torch.cuda.synchronize()
else:
output = input_t
# if scattering the seq-dim, transpose the heads back to the original dimension
output = output.reshape(seqlen, bs, shard_hc, hs)
# (seq_len, bs, hc/P, hs) -reshape-> (bs, seq_len, hc/P, hs)
output = output.transpose(0, 1).contiguous().reshape(bs, seqlen, shard_hc, hs)
return output
elif scatter_idx == 1 and gather_idx == 2:
# input (torch.tensor): a tensor sharded along dim 1 (bs, seqlen, hc/P, hs) output: (bs, seqlen/P, hc, hs)
bs, seqlen, shard_hc, hs = input.shape
hc = shard_hc * seq_world_size
shard_seqlen = seqlen // seq_world_size
seq_world_size = dist.get_world_size(group)
# transpose groups of heads with the seq-len parallel dimension, so that we can scatter them!
# (bs, seqlen, hc/P, hs) -reshape-> (bs, P, seq_len/P, hc/P, hs) -transpose(0, 3)-> (hc/P, P, seqlen/P, bs, hs) -transpose(0, 1) -> (P, hc/P, seqlen/P, bs, hs)
input_t = (
input.reshape(bs, seq_world_size, shard_seqlen, shard_hc, hs)
.transpose(0, 3)
.transpose(0, 1)
.contiguous()
.reshape(seq_world_size, shard_hc, shard_seqlen, bs, hs)
)
output = torch.empty_like(input_t)
# https://pytorch.org/docs/stable/distributed.html#torch.distributed.all_to_all_single
# (P, bs x hc/P, seqlen/P, hs) scatter seqlen -all2all-> (P, bs x seq_len/P, hc/P, hs) scatter head
if seq_world_size > 1:
dist.all_to_all_single(output, input_t, group=group)
torch.cuda.synchronize()
else:
output = input_t
# if scattering the seq-dim, transpose the heads back to the original dimension
output = output.reshape(hc, shard_seqlen, bs, hs)
# (hc, seqlen/N, bs, hs) -tranpose(0,2)-> (bs, seqlen/N, hc, hs)
output = output.transpose(0, 2).contiguous().reshape(bs, shard_seqlen, hc, hs)
return output
else:
raise RuntimeError("scatter_idx must be 1 or 2 and gather_idx must be 1 or 2")
class SeqAllToAll4D(torch.autograd.Function):
@staticmethod
def forward(
ctx: Any,
group: dist.ProcessGroup,
input: Tensor,
scatter_idx: int,
gather_idx: int,
) -> Tensor:
ctx.group = group
ctx.scatter_idx = scatter_idx
ctx.gather_idx = gather_idx
return _all_to_all_4D(input, scatter_idx, gather_idx, group=group)
@staticmethod
def backward(ctx: Any, *grad_output: Tensor) -> Tuple[None, Tensor, None, None]:
return (
None,
SeqAllToAll4D.apply(
ctx.group, *grad_output, ctx.gather_idx, ctx.scatter_idx
),
None,
None,
)
def all_to_all_4D(
input_: torch.Tensor,
scatter_dim: int = 2,
gather_dim: int = 1,
):
return SeqAllToAll4D.apply(nccl_info.group, input_, scatter_dim, gather_dim)
def _all_to_all(
input_: torch.Tensor,
world_size: int,
group: dist.ProcessGroup,
scatter_dim: int,
gather_dim: int,
):
input_list = [
t.contiguous() for t in torch.tensor_split(input_, world_size, scatter_dim)
]
output_list = [torch.empty_like(input_list[0]) for _ in range(world_size)]
dist.all_to_all(output_list, input_list, group=group)
return torch.cat(output_list, dim=gather_dim).contiguous()
class _AllToAll(torch.autograd.Function):
"""All-to-all communication.
Args:
input_: input matrix
process_group: communication group
scatter_dim: scatter dimension
gather_dim: gather dimension
"""
@staticmethod
def forward(ctx, input_, process_group, scatter_dim, gather_dim):
ctx.process_group = process_group
ctx.scatter_dim = scatter_dim
ctx.gather_dim = gather_dim
ctx.world_size = dist.get_world_size(process_group)
output = _all_to_all(
input_, ctx.world_size, process_group, scatter_dim, gather_dim
)
return output
@staticmethod
def backward(ctx, grad_output):
grad_output = _all_to_all(
grad_output,
ctx.world_size,
ctx.process_group,
ctx.gather_dim,
ctx.scatter_dim,
)
return (
grad_output,
None,
None,
None,
)
def all_to_all(
input_: torch.Tensor,
scatter_dim: int = 2,
gather_dim: int = 1,
):
return _AllToAll.apply(input_, nccl_info.group, scatter_dim, gather_dim)
class _AllGather(torch.autograd.Function):
"""All-gather communication with autograd support.
Args:
input_: input tensor
dim: dimension along which to concatenate
"""
@staticmethod
def forward(ctx, input_, dim):
ctx.dim = dim
world_size = nccl_info.sp_size
group = nccl_info.group
input_size = list(input_.size())
ctx.input_size = input_size[dim]
tensor_list = [torch.empty_like(input_) for _ in range(world_size)]
input_ = input_.contiguous()
dist.all_gather(tensor_list, input_, group=group)
output = torch.cat(tensor_list, dim=dim)
return output
@staticmethod
def backward(ctx, grad_output):
world_size = nccl_info.sp_size
rank = nccl_info.rank_within_group
dim = ctx.dim
input_size = ctx.input_size
sizes = [input_size] * world_size
grad_input_list = torch.split(grad_output, sizes, dim=dim)
grad_input = grad_input_list[rank]
return grad_input, None
def all_gather(input_: torch.Tensor, dim: int = 1):
"""Performs an all-gather operation on the input tensor along the specified dimension.
Args:
input_ (torch.Tensor): Input tensor of shape [B, H, S, D].
dim (int, optional): Dimension along which to concatenate. Defaults to 1.
Returns:
torch.Tensor: Output tensor after all-gather operation, concatenated along 'dim'.
"""
return _AllGather.apply(input_, dim)
def prepare_sequence_parallel_data(
hidden_states, encoder_hidden_states, attention_mask, encoder_attention_mask
):
if nccl_info.sp_size == 1:
return (
hidden_states,
encoder_hidden_states,
attention_mask,
encoder_attention_mask,
)
def prepare(
hidden_states, encoder_hidden_states, attention_mask, encoder_attention_mask
):
hidden_states = all_to_all(hidden_states, scatter_dim=2, gather_dim=0)
encoder_hidden_states = all_to_all(
encoder_hidden_states, scatter_dim=1, gather_dim=0
)
attention_mask = all_to_all(attention_mask, scatter_dim=1, gather_dim=0)
encoder_attention_mask = all_to_all(
encoder_attention_mask, scatter_dim=1, gather_dim=0
)
return (
hidden_states,
encoder_hidden_states,
attention_mask,
encoder_attention_mask,
)
sp_size = nccl_info.sp_size
frame = hidden_states.shape[2]
assert frame % sp_size == 0, "frame should be a multiple of sp_size"
(
hidden_states,
encoder_hidden_states,
attention_mask,
encoder_attention_mask,
) = prepare(
hidden_states,
encoder_hidden_states.repeat(1, sp_size, 1),
attention_mask.repeat(1, sp_size, 1, 1),
encoder_attention_mask.repeat(1, sp_size),
)
return hidden_states, encoder_hidden_states, attention_mask, encoder_attention_mask
def sp_parallel_dataloader_wrapper(
dataloader, device, train_batch_size, sp_size, train_sp_batch_size
):
while True:
for data_item in dataloader:
latents, cond, attn_mask, cond_mask = data_item
latents = latents.to(device)
cond = cond.to(device)
attn_mask = attn_mask.to(device)
cond_mask = cond_mask.to(device)
frame = latents.shape[2]
if frame == 1:
yield latents, cond, attn_mask, cond_mask
else:
latents, cond, attn_mask, cond_mask = prepare_sequence_parallel_data(
latents, cond, attn_mask, cond_mask
)
assert (
train_batch_size * sp_size >= train_sp_batch_size
), "train_batch_size * sp_size should be greater than train_sp_batch_size"
for iter in range(train_batch_size * sp_size // train_sp_batch_size):
st_idx = iter * train_sp_batch_size
ed_idx = (iter + 1) * train_sp_batch_size
encoder_hidden_states = cond[st_idx:ed_idx]
attention_mask = attn_mask[st_idx:ed_idx]
encoder_attention_mask = cond_mask[st_idx:ed_idx]
yield (
latents[st_idx:ed_idx],
encoder_hidden_states,
attention_mask,
encoder_attention_mask,
)
@@ -1,152 +0,0 @@
import argparse
import torch
from accelerate.logging import get_logger
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
from diffusers.utils import export_to_video
import json
import os
import torch.distributed as dist
logger = get_logger(__name__)
from torch.utils.data import Dataset
from torch.utils.data.distributed import DistributedSampler
from torch.utils.data import DataLoader
class T5dataset(Dataset):
def __init__(
self,
json_path,
vae_debug,
):
self.json_path = json_path
self.vae_debug = vae_debug
with open(self.json_path, "r") as f:
train_dataset = json.load(f)
self.train_dataset = sorted(train_dataset, key=lambda x: x["latent_path"])
def __getitem__(self, idx):
caption = self.train_dataset[idx]["caption"]
filename = self.train_dataset[idx]["latent_path"].split(".")[0]
length = self.train_dataset[idx]["length"]
if self.vae_debug:
latents = torch.load(
os.path.join(
args.output_dir, "latent", self.train_dataset[idx]["latent_path"]
),
map_location="cpu",
)
else:
latents = []
return dict(caption=caption, latents=latents, filename=filename, length=length)
def __len__(self):
return len(self.train_dataset)
def main(args):
local_rank = int(os.getenv("RANK", 0))
world_size = int(os.getenv("WORLD_SIZE", 1))
print("world_size", world_size, "local rank", local_rank)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
torch.cuda.set_device(local_rank)
if not dist.is_initialized():
dist.init_process_group(
backend="nccl", init_method="env://", world_size=world_size, rank=local_rank
)
pipe = MochiPipeline.from_pretrained(args.model_path).to(device)
pipe.vae.enable_tiling()
os.makedirs(args.output_dir, exist_ok=True)
os.makedirs(os.path.join(args.output_dir, "video"), exist_ok=True)
os.makedirs(os.path.join(args.output_dir, "latent"), exist_ok=True)
os.makedirs(os.path.join(args.output_dir, "prompt_embed"), exist_ok=True)
os.makedirs(os.path.join(args.output_dir, "prompt_attention_mask"), exist_ok=True)
latents_json_path = os.path.join(args.output_dir, "videos2caption_temp.json")
train_dataset = T5dataset(latents_json_path, args.vae_debug)
sampler = DistributedSampler(
train_dataset, rank=local_rank, num_replicas=world_size, shuffle=True
)
train_dataloader = DataLoader(
train_dataset,
sampler=sampler,
batch_size=args.train_batch_size,
num_workers=args.dataloader_num_workers,
)
json_data = []
for _, data in enumerate(train_dataloader):
with torch.inference_mode():
with torch.autocast("cuda", dtype=torch.bfloat16):
prompt_embeds, prompt_attention_mask, _, _ = pipe.encode_prompt(
prompt=data["caption"],
)
if args.vae_debug:
latents = data["latents"]
video = pipe.vae.decode(latents.to(device), return_dict=False)[0]
video = pipe.video_processor.postprocess_video(video)
for idx, video_name in enumerate(data["filename"]):
prompt_embed_path = os.path.join(
args.output_dir, "prompt_embed", video_name + ".pt"
)
video_path = os.path.join(
args.output_dir, "video", video_name + ".mp4"
)
prompt_attention_mask_path = os.path.join(
args.output_dir, "prompt_attention_mask", video_name + ".pt"
)
# save latent
torch.save(prompt_embeds[idx], prompt_embed_path)
torch.save(prompt_attention_mask[idx], prompt_attention_mask_path)
print(f"sample {video_name} saved")
if args.vae_debug:
export_to_video(video[idx], video_path, fps=30)
item = {}
item["length"] = int(data["length"][idx])
item["latent_path"] = video_name + ".pt"
item["prompt_embed_path"] = video_name + ".pt"
item["prompt_attention_mask"] = video_name + ".pt"
item["caption"] = data["caption"][idx]
json_data.append(item)
dist.barrier()
local_data = json_data
gathered_data = [None] * world_size
dist.all_gather_object(gathered_data, local_data)
if local_rank == 0:
# os.remove(latents_json_path)
all_json_data = [item for sublist in gathered_data for item in sublist]
with open(os.path.join(args.output_dir, "videos2caption.json"), "w") as f:
json.dump(all_json_data, f, indent=4)
if __name__ == "__main__":
parser = argparse.ArgumentParser()
# dataset & dataloader
parser.add_argument("--model_path", type=str, default="data/mochi")
# text encoder & vae & diffusion model
parser.add_argument(
"--dataloader_num_workers",
type=int,
default=1,
help="Number of subprocesses to use for data loading. 0 means that the data will be loaded in the main process.",
)
parser.add_argument(
"--train_batch_size",
type=int,
default=1,
help="Batch size (per device) for the training dataloader.",
)
parser.add_argument("--text_encoder_name", type=str, default="google/t5-v1_1-xxl")
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
parser.add_argument(
"--output_dir",
type=str,
default=None,
help="The output directory where the model predictions and checkpoints will be written.",
)
parser.add_argument("--vae_debug", action="store_true")
args = parser.parse_args()
main(args)
@@ -1,143 +0,0 @@
from fastvideo.dataset import getdataset
from torch.utils.data import DataLoader
from fastvideo.utils.dataset_utils import Collate
import argparse
import torch
from accelerate import Accelerator
from accelerate.logging import get_logger
from accelerate.utils import ProjectConfiguration
import json
import os
from diffusers import AutoencoderKLMochi
import torch.distributed as dist
from torch.utils.data.distributed import DistributedSampler
logger = get_logger(__name__)
def main(args):
local_rank = int(os.getenv("RANK", 0))
world_size = int(os.getenv("WORLD_SIZE", 1))
print("world_size", world_size, "local rank", local_rank)
args.ae_stride_t, args.ae_stride_h, args.ae_stride_w = 4, 8, 8
args.ae_stride = args.ae_stride_h
patch_size_t, patch_size_h, patch_size_w = 1, 2, 2
args.patch_size = patch_size_h
args.patch_size_t, args.patch_size_h, args.patch_size_w = (
patch_size_t,
patch_size_h,
patch_size_w,
)
accelerator_project_config = ProjectConfiguration(
project_dir=args.output_dir, logging_dir=args.logging_dir
)
accelerator = Accelerator(
project_config=accelerator_project_config,
)
train_dataset = getdataset(args)
sampler = DistributedSampler(
train_dataset, rank=local_rank, num_replicas=world_size, shuffle=True
)
train_dataloader = DataLoader(
train_dataset,
sampler=sampler,
batch_size=args.train_batch_size,
num_workers=args.dataloader_num_workers,
)
encoder_device = torch.device(f"cuda" if torch.cuda.is_available() else "cpu")
torch.cuda.set_device(local_rank)
if not dist.is_initialized():
dist.init_process_group(
backend="nccl", init_method="env://", world_size=world_size, rank=local_rank
)
vae = AutoencoderKLMochi.from_pretrained(args.model_path, subfolder="vae").to(
"cuda"
)
vae.enable_tiling()
os.makedirs(args.output_dir, exist_ok=True)
os.makedirs(os.path.join(args.output_dir, "latent"), exist_ok=True)
json_data = []
for _, data in enumerate(train_dataloader):
with torch.inference_mode():
with torch.autocast("cuda", dtype=torch.bfloat16):
latents = vae.encode(data["pixel_values"].to(encoder_device))[
"latent_dist"
].sample()
for idx, video_path in enumerate(data["path"]):
video_name = os.path.basename(video_path).split(".")[0]
latent_path = os.path.join(
args.output_dir, "latent", video_name + ".pt"
)
torch.save(latents[idx].to(torch.bfloat16), latent_path)
item = {}
item["length"] = latents[idx].shape[1]
item["latent_path"] = video_name + ".pt"
item["caption"] = data["text"][idx]
json_data.append(item)
print(f"{video_name} processed")
dist.barrier()
local_data = json_data
gathered_data = [None] * world_size
dist.all_gather_object(gathered_data, local_data)
if local_rank == 0:
all_json_data = [item for sublist in gathered_data for item in sublist]
with open(os.path.join(args.output_dir, "videos2caption_temp.json"), "w") as f:
json.dump(all_json_data, f, indent=4)
if __name__ == "__main__":
parser = argparse.ArgumentParser()
# dataset & dataloader
parser.add_argument("--model_path", type=str, default="data/mochi")
parser.add_argument("--data_merge_path", type=str, required=True)
parser.add_argument("--num_frames", type=int, default=163)
parser.add_argument(
"--dataloader_num_workers",
type=int,
default=1,
help="Number of subprocesses to use for data loading. 0 means that the data will be loaded in the main process.",
)
parser.add_argument(
"--train_batch_size",
type=int,
default=16,
help="Batch size (per device) for the training dataloader.",
)
parser.add_argument(
"--num_latent_t", type=int, default=28, help="Number of latent timesteps."
)
parser.add_argument("--max_height", type=int, default=480)
parser.add_argument("--max_width", type=int, default=848)
parser.add_argument("--video_length_tolerance_range", type=int, default=2.0)
parser.add_argument("--group_frame", action="store_true") # TODO
parser.add_argument("--group_resolution", action="store_true") # TODO
parser.add_argument("--dataset", default="t2v")
parser.add_argument("--train_fps", type=int, default=30)
parser.add_argument("--use_image_num", type=int, default=0)
parser.add_argument("--text_max_length", type=int, default=256)
parser.add_argument("--speed_factor", type=float, default=1.0)
parser.add_argument("--drop_short_ratio", type=float, default=1.0)
# text encoder & vae & diffusion model
parser.add_argument("--text_encoder_name", type=str, default="google/t5-v1_1-xxl")
parser.add_argument("--cache_dir", type=str, default="./cache_dir")
parser.add_argument("--cfg", type=float, default=0.0)
parser.add_argument(
"--output_dir",
type=str,
default=None,
help="The output directory where the model predictions and checkpoints will be written.",
)
parser.add_argument(
"--logging_dir",
type=str,
default="logs",
help=(
"[TensorBoard](https://www.tensorflow.org/tensorboard) log directory. Will default to"
" *output_dir/runs/**CURRENT_DATETIME_HOSTNAME***."
),
)
args = parser.parse_args()
main(args)
-377
View File
@@ -1,377 +0,0 @@
import math
from einops import rearrange
import decord
from torch.nn import functional as F
import torch
from typing import Optional
import torch.utils
import torch.utils.data
import torch
from torch.utils.data import Sampler
from typing import List
from collections import Counter
import random
IMG_EXTENSIONS = [".jpg", ".JPG", ".jpeg", ".JPEG", ".png", ".PNG"]
def is_image_file(filename):
return any(filename.endswith(extension) for extension in IMG_EXTENSIONS)
class DecordInit(object):
"""Using Decord(https://github.com/dmlc/decord) to initialize the video_reader."""
def __init__(self, num_threads=1):
self.num_threads = num_threads
self.ctx = decord.cpu(0)
def __call__(self, filename):
"""Perform the Decord initialization.
Args:
results (dict): The resulting dict to be modified and passed
to the next transform in pipeline.
"""
reader = decord.VideoReader(
filename, ctx=self.ctx, num_threads=self.num_threads
)
return reader
def __repr__(self):
repr_str = (
f"{self.__class__.__name__}("
f"sr={self.sr},"
f"num_threads={self.num_threads})"
)
return repr_str
def pad_to_multiple(number, ds_stride):
remainder = number % ds_stride
if remainder == 0:
return number
else:
padding = ds_stride - remainder
return number + padding
class Collate:
def __init__(self, args):
self.batch_size = args.train_batch_size
self.group_frame = args.group_frame
self.group_resolution = args.group_resolution
self.max_height = args.max_height
self.max_width = args.max_width
self.ae_stride = args.ae_stride
self.ae_stride_t = args.ae_stride_t
self.ae_stride_thw = (self.ae_stride_t, self.ae_stride, self.ae_stride)
self.patch_size = args.patch_size
self.patch_size_t = args.patch_size_t
self.num_frames = args.num_frames
self.use_image_num = args.use_image_num
self.max_thw = (self.num_frames, self.max_height, self.max_width)
def package(self, batch):
batch_tubes = [i["pixel_values"] for i in batch] # b [c t h w]
input_ids = [i["input_ids"] for i in batch] # b [1 l]
cond_mask = [i["cond_mask"] for i in batch] # b [1 l]
return batch_tubes, input_ids, cond_mask
def __call__(self, batch):
batch_tubes, input_ids, cond_mask = self.package(batch)
ds_stride = self.ae_stride * self.patch_size
t_ds_stride = self.ae_stride_t * self.patch_size_t
pad_batch_tubes, attention_mask, input_ids, cond_mask = self.process(
batch_tubes,
input_ids,
cond_mask,
t_ds_stride,
ds_stride,
self.max_thw,
self.ae_stride_thw,
)
assert not torch.any(torch.isnan(pad_batch_tubes)), "after pad_batch_tubes"
return pad_batch_tubes, attention_mask, input_ids, cond_mask
def process(
self,
batch_tubes,
input_ids,
cond_mask,
t_ds_stride,
ds_stride,
max_thw,
ae_stride_thw,
):
# pad to max multiple of ds_stride
batch_input_size = [i.shape for i in batch_tubes] # [(c t h w), (c t h w)]
assert len(batch_input_size) == self.batch_size
if self.group_frame or self.group_resolution or self.batch_size == 1: #
len_each_batch = batch_input_size
idx_length_dict = dict([*zip(list(range(self.batch_size)), len_each_batch)])
count_dict = Counter(len_each_batch)
if len(count_dict) != 1:
sorted_by_value = sorted(count_dict.items(), key=lambda item: item[1])
pick_length = sorted_by_value[-1][0] # the highest frequency
candidate_batch = [
idx
for idx, length in idx_length_dict.items()
if length == pick_length
]
random_select_batch = [
random.choice(candidate_batch)
for _ in range(len(len_each_batch) - len(candidate_batch))
]
print(
batch_input_size,
idx_length_dict,
count_dict,
sorted_by_value,
pick_length,
candidate_batch,
random_select_batch,
)
pick_idx = candidate_batch + random_select_batch
batch_tubes = [batch_tubes[i] for i in pick_idx]
batch_input_size = [
i.shape for i in batch_tubes
] # [(c t h w), (c t h w)]
input_ids = [input_ids[i] for i in pick_idx] # b [1, l]
cond_mask = [cond_mask[i] for i in pick_idx] # b [1, l]
for i in range(1, self.batch_size):
assert batch_input_size[0] == batch_input_size[i]
max_t = max([i[1] for i in batch_input_size])
max_h = max([i[2] for i in batch_input_size])
max_w = max([i[3] for i in batch_input_size])
else:
max_t, max_h, max_w = max_thw
pad_max_t, pad_max_h, pad_max_w = (
pad_to_multiple(max_t - 1 + self.ae_stride_t, t_ds_stride),
pad_to_multiple(max_h, ds_stride),
pad_to_multiple(max_w, ds_stride),
)
pad_max_t = pad_max_t + 1 - self.ae_stride_t
each_pad_t_h_w = [
[pad_max_t - i.shape[1], pad_max_h - i.shape[2], pad_max_w - i.shape[3]]
for i in batch_tubes
]
pad_batch_tubes = [
F.pad(im, (0, pad_w, 0, pad_h, 0, pad_t), value=0)
for (pad_t, pad_h, pad_w), im in zip(each_pad_t_h_w, batch_tubes)
]
pad_batch_tubes = torch.stack(pad_batch_tubes, dim=0)
max_tube_size = [pad_max_t, pad_max_h, pad_max_w]
max_latent_size = [
((max_tube_size[0] - 1) // ae_stride_thw[0] + 1),
max_tube_size[1] // ae_stride_thw[1],
max_tube_size[2] // ae_stride_thw[2],
]
valid_latent_size = [
[
int(math.ceil((i[1] - 1) / ae_stride_thw[0])) + 1,
int(math.ceil(i[2] / ae_stride_thw[1])),
int(math.ceil(i[3] / ae_stride_thw[2])),
]
for i in batch_input_size
]
attention_mask = [
F.pad(
torch.ones(i, dtype=pad_batch_tubes.dtype),
(
0,
max_latent_size[2] - i[2],
0,
max_latent_size[1] - i[1],
0,
max_latent_size[0] - i[0],
),
value=0,
)
for i in valid_latent_size
]
attention_mask = torch.stack(attention_mask) # b t h w
if self.batch_size == 1 or self.group_frame or self.group_resolution:
assert torch.all(attention_mask.bool())
input_ids = torch.stack(input_ids) # b 1 l
cond_mask = torch.stack(cond_mask) # b 1 l
return pad_batch_tubes, attention_mask, input_ids, cond_mask
def split_to_even_chunks(indices, lengths, num_chunks, batch_size):
"""
Split a list of indices into `chunks` chunks of roughly equal lengths.
"""
if len(indices) % num_chunks != 0:
chunks = [indices[i::num_chunks] for i in range(num_chunks)]
else:
num_indices_per_chunk = len(indices) // num_chunks
chunks = [[] for _ in range(num_chunks)]
chunks_lengths = [0 for _ in range(num_chunks)]
for index in indices:
shortest_chunk = chunks_lengths.index(min(chunks_lengths))
chunks[shortest_chunk].append(index)
chunks_lengths[shortest_chunk] += lengths[index]
if len(chunks[shortest_chunk]) == num_indices_per_chunk:
chunks_lengths[shortest_chunk] = float("inf")
# return chunks
pad_chunks = []
for idx, chunk in enumerate(chunks):
if batch_size != len(chunk):
assert batch_size > len(chunk)
if len(chunk) != 0:
chunk = chunk + [
random.choice(chunk) for _ in range(batch_size - len(chunk))
]
else:
chunk = random.choice(pad_chunks)
print(chunks[idx], "->", chunk)
pad_chunks.append(chunk)
return pad_chunks
def group_frame_fun(indices, lengths):
# sort by num_frames
indices.sort(key=lambda i: lengths[i], reverse=True)
return indices
def megabatch_frame_alignment(megabatches, lengths):
aligned_magabatches = []
for _, megabatch in enumerate(megabatches):
assert len(megabatch) != 0
len_each_megabatch = [lengths[i] for i in megabatch]
idx_length_dict = dict([*zip(megabatch, len_each_megabatch)])
count_dict = Counter(len_each_megabatch)
# mixed frame length, align megabatch inside
if len(count_dict) != 1:
sorted_by_value = sorted(count_dict.items(), key=lambda item: item[1])
pick_length = sorted_by_value[-1][0] # the highest frequency
candidate_batch = [
idx for idx, length in idx_length_dict.items() if length == pick_length
]
random_select_batch = [
random.choice(candidate_batch)
for i in range(len(idx_length_dict) - len(candidate_batch))
]
aligned_magabatch = candidate_batch + random_select_batch
aligned_magabatches.append(aligned_magabatch)
# already aligned megabatches
else:
aligned_magabatches.append(megabatch)
return aligned_magabatches
def get_length_grouped_indices(
lengths,
batch_size,
world_size,
generator=None,
group_frame=False,
group_resolution=False,
seed=42,
):
# We need to use torch for the random part as a distributed sampler will set the random seed for torch.
if generator is None:
generator = torch.Generator().manual_seed(
seed
) # every rank will generate a fixed order but random index
indices = torch.randperm(len(lengths), generator=generator).tolist()
# sort dataset according to frame
indices = group_frame_fun(indices, lengths)
# chunk dataset to megabatches
megabatch_size = world_size * batch_size
megabatches = [
indices[i : i + megabatch_size] for i in range(0, len(lengths), megabatch_size)
]
# make sure the length in each magabatch is align with each other
megabatches = megabatch_frame_alignment(megabatches, lengths)
# aplit aligned megabatch into batches
megabatches = [
split_to_even_chunks(megabatch, lengths, world_size, batch_size)
for megabatch in megabatches
]
# random megabatches to do video-image mix training
indices = torch.randperm(len(megabatches), generator=generator).tolist()
shuffled_megabatches = [megabatches[i] for i in indices]
# expand indices and return
return [
i for megabatch in shuffled_megabatches for batch in megabatch for i in batch
]
class LengthGroupedSampler(Sampler):
r"""
Sampler that samples indices in a way that groups together features of the dataset of roughly the same length while
keeping a bit of randomness.
"""
def __init__(
self,
batch_size: int,
rank: int,
world_size: int,
lengths: Optional[List[int]] = None,
group_frame=False,
group_resolution=False,
generator=None,
):
if lengths is None:
raise ValueError("Lengths must be provided.")
self.batch_size = batch_size
self.rank = rank
self.world_size = world_size
self.lengths = lengths
self.group_frame = group_frame
self.group_resolution = group_resolution
self.generator = generator
def __len__(self):
return len(self.lengths)
def __iter__(self):
indices = get_length_grouped_indices(
self.lengths,
self.batch_size,
self.world_size,
group_frame=self.group_frame,
group_resolution=self.group_resolution,
generator=self.generator,
)
def distributed_sampler(lst, rank, batch_size, world_size):
result = []
index = rank * batch_size
while index < len(lst):
result.extend(lst[index : index + batch_size])
index += batch_size * world_size
return result
indices = distributed_sampler(
indices, self.rank, self.batch_size, self.world_size
)
return iter(indices)
-24
View File
@@ -1,24 +0,0 @@
import sys
import pdb
import os
def main_print(content):
if int(os.environ["LOCAL_RANK"]) <= 0:
print(content)
# ForkedPdb().set_trace()
class ForkedPdb(pdb.Pdb):
"""A Pdb subclass that may be used
from a forked multiprocessing child
"""
def interaction(self, *args, **kwargs):
_stdin = sys.stdin
try:
sys.stdin = open("/dev/stdin")
pdb.Pdb.interaction(self, *args, **kwargs)
finally:
sys.stdin = _stdin
-77
View File
@@ -1,77 +0,0 @@
from accelerate.logging import get_logger
import torch
logger = get_logger(__name__)
def get_optimizer(args, params_to_optimize, use_deepspeed: bool = False):
# Optimizer creation
supported_optimizers = ["adam", "adamw", "prodigy"]
if args.optimizer not in supported_optimizers:
logger.warning(
f"Unsupported choice of optimizer: {args.optimizer}. Supported optimizers include {supported_optimizers}. Defaulting to AdamW"
)
args.optimizer = "adamw"
if args.use_8bit_adam and not (args.optimizer.lower() not in ["adam", "adamw"]):
logger.warning(
f"use_8bit_adam is ignored when optimizer is not set to 'Adam' or 'AdamW'. Optimizer was "
f"set to {args.optimizer.lower()}"
)
if args.use_8bit_adam:
try:
import bitsandbytes as bnb
except ImportError:
raise ImportError(
"To use 8-bit Adam, please install the bitsandbytes library: `pip install bitsandbytes`."
)
if args.optimizer.lower() == "adamw":
optimizer_class = (
bnb.optim.AdamW8bit if args.use_8bit_adam else torch.optim.AdamW
)
optimizer = optimizer_class(
params_to_optimize,
betas=(args.adam_beta1, args.adam_beta2),
eps=args.adam_epsilon,
weight_decay=args.adam_weight_decay,
)
elif args.optimizer.lower() == "adam":
optimizer_class = bnb.optim.Adam8bit if args.use_8bit_adam else torch.optim.Adam
optimizer = optimizer_class(
params_to_optimize,
betas=(args.adam_beta1, args.adam_beta2),
eps=args.adam_epsilon,
weight_decay=args.adam_weight_decay,
)
elif args.optimizer.lower() == "prodigy":
try:
import prodigyopt
except ImportError:
raise ImportError(
"To use Prodigy, please install the prodigyopt library: `pip install prodigyopt`"
)
optimizer_class = prodigyopt.Prodigy
if args.learning_rate <= 0.1:
logger.warning(
"Learning rate is too low. When using prodigy, it's generally better to set learning rate around 1.0"
)
optimizer = optimizer_class(
params_to_optimize,
lr=args.learning_rate,
betas=(args.adam_beta1, args.adam_beta2),
beta3=args.prodigy_beta3,
weight_decay=args.adam_weight_decay,
eps=args.adam_epsilon,
decouple=args.prodigy_decouple,
use_bias_correction=args.prodigy_use_bias_correction,
safeguard_warmup=args.prodigy_safeguard_warmup,
)
return optimizer
-63
View File
@@ -1,63 +0,0 @@
import torch
import torch.distributed as dist
import os
class COMM_INFO:
def __init__(self):
self.group = None
self.sp_size = 1
self.global_rank = 0
self.rank_within_group = 0
self.group_id = 0
nccl_info = COMM_INFO()
_SEQUENCE_PARALLEL_STATE = False
def initialize_sequence_parallel_state(sequence_parallel_size):
global _SEQUENCE_PARALLEL_STATE
if sequence_parallel_size > 1:
_SEQUENCE_PARALLEL_STATE = True
initialize_sequence_parallel_group(sequence_parallel_size)
else:
nccl_info.sp_size = 1
nccl_info.global_rank = int(os.getenv("RANK", "0"))
nccl_info.rank_within_group = 0
nccl_info.group_id = int(os.getenv("RANK", "0"))
def set_sequence_parallel_state(state):
global _SEQUENCE_PARALLEL_STATE
_SEQUENCE_PARALLEL_STATE = state
def get_sequence_parallel_state():
return _SEQUENCE_PARALLEL_STATE
def initialize_sequence_parallel_group(sequence_parallel_size):
"""Initialize the sequence parallel group."""
rank = int(os.getenv("RANK", "0"))
world_size = int(os.getenv("WORLD_SIZE", "1"))
assert (
world_size % sequence_parallel_size == 0
), "world_size must be divisible by sequence_parallel_size, but got world_size: {}, sequence_parallel_size: {}".format(
world_size, sequence_parallel_size
)
nccl_info.sp_size = sequence_parallel_size
nccl_info.global_rank = rank
num_sequence_parallel_groups: int = world_size // sequence_parallel_size
for i in range(num_sequence_parallel_groups):
ranks = range(i * sequence_parallel_size, (i + 1) * sequence_parallel_size)
group = dist.new_group(ranks)
if rank in ranks:
nccl_info.group = group
nccl_info.rank_within_group = rank - i * sequence_parallel_size
nccl_info.group_id = i
def destroy_sequence_parallel_group():
"""Destroy the sequence parallel group."""
dist.destroy_process_group()
-341
View File
@@ -1,341 +0,0 @@
from typing import Optional, Union, List
import numpy as np
import torch
from einops import rearrange
from fastvideo.utils.parallel_states import get_sequence_parallel_state, nccl_info
from fastvideo.utils.communications import all_gather
from diffusers.utils.torch_utils import randn_tensor
from fastvideo.models.mochi_hf.pipeline_mochi import (
linear_quadratic_schedule,
retrieve_timesteps,
)
from tqdm import tqdm
from diffusers.video_processor import VideoProcessor
from diffusers import (
FlowMatchEulerDiscreteScheduler,
AutoencoderKLMochi,
)
from fastvideo.utils.logging import main_print
from fastvideo.distill.solver import PCMFMScheduler
from diffusers.utils import export_to_video
import os
import wandb
import gc
def prepare_latents(
batch_size,
num_channels_latents,
height,
width,
num_frames,
dtype,
device,
generator,
vae_spatial_scale_factor,
vae_temporal_scale_factor,
):
height = height // vae_spatial_scale_factor
width = width // vae_spatial_scale_factor
num_frames = (num_frames - 1) // vae_temporal_scale_factor + 1
shape = (batch_size, num_channels_latents, num_frames, height, width)
latents = randn_tensor(shape, generator=generator, device=device, dtype=dtype)
return latents
def sample_validation_video(
transformer,
vae,
scheduler,
scheduler_type="euler",
height: Optional[int] = None,
width: Optional[int] = None,
num_frames: int = 16,
num_inference_steps: int = 28,
timesteps: List[int] = None,
guidance_scale: float = 4.5,
num_videos_per_prompt: Optional[int] = 1,
generator: Optional[Union[torch.Generator, List[torch.Generator]]] = None,
prompt_embeds: Optional[torch.Tensor] = None,
prompt_attention_mask: Optional[torch.Tensor] = None,
negative_prompt_embeds: Optional[torch.Tensor] = None,
negative_prompt_attention_mask: Optional[torch.Tensor] = None,
output_type: Optional[str] = "pil",
vae_spatial_scale_factor=8,
vae_temporal_scale_factor=6,
):
device = vae.device
batch_size = prompt_embeds.shape[0]
do_classifier_free_guidance = guidance_scale > 1.0
if do_classifier_free_guidance:
prompt_embeds = torch.cat([negative_prompt_embeds, prompt_embeds], dim=0)
prompt_attention_mask = torch.cat(
[negative_prompt_attention_mask, prompt_attention_mask], dim=0
)
# 4. Prepare latent variables
# TODO: Remove hardcore
num_channels_latents = 12
latents = prepare_latents(
batch_size * num_videos_per_prompt,
num_channels_latents,
height,
width,
num_frames,
prompt_embeds.dtype,
device,
generator,
vae_spatial_scale_factor,
vae_temporal_scale_factor,
)
world_size, rank = nccl_info.sp_size, nccl_info.rank_within_group
if get_sequence_parallel_state():
latents = rearrange(
latents, "b t (n s) h w -> b t n s h w", n=world_size
).contiguous()
latents = latents[:, :, rank, :, :, :]
# 5. Prepare timestep
# from https://github.com/genmoai/models/blob/075b6e36db58f1242921deff83a1066887b9c9e1/src/mochi_preview/infer.py#L77
threshold_noise = 0.025
sigmas = linear_quadratic_schedule(num_inference_steps, threshold_noise)
sigmas = np.array(sigmas)
if scheduler_type == "euler":
timesteps, num_inference_steps = retrieve_timesteps(
scheduler,
num_inference_steps,
device,
timesteps,
sigmas,
)
else:
timesteps, num_inference_steps = retrieve_timesteps(
scheduler,
num_inference_steps,
device,
)
num_warmup_steps = max(len(timesteps) - num_inference_steps * scheduler.order, 0)
# 6. Denoising loop
# with self.progress_bar(total=num_inference_steps) as progress_bar:
# write with tqdm instead
# only enable if nccl_info.global_rank == 0
with tqdm(
total=num_inference_steps,
disable=nccl_info.rank_within_group != 0,
desc="Validation sampling...",
) as progress_bar:
for i, t in enumerate(timesteps):
latent_model_input = (
torch.cat([latents] * 2) if do_classifier_free_guidance else latents
)
# broadcast to batch dimension in a way that's compatible with ONNX/Core ML
timestep = t.expand(latent_model_input.shape[0]).to(latents.dtype)
noise_pred = transformer(
hidden_states=latent_model_input,
encoder_hidden_states=prompt_embeds,
timestep=timestep,
encoder_attention_mask=prompt_attention_mask,
return_dict=False,
)[0]
# Mochi CFG + Sampling runs in FP32
noise_pred = noise_pred.to(torch.float32)
if do_classifier_free_guidance:
noise_pred_uncond, noise_pred_text = noise_pred.chunk(2)
noise_pred = noise_pred_uncond + guidance_scale * (
noise_pred_text - noise_pred_uncond
)
# compute the previous noisy sample x_t -> x_t-1
latents_dtype = latents.dtype
latents = scheduler.step(
noise_pred, t, latents.to(torch.float32), return_dict=False
)[0]
latents = latents.to(latents_dtype)
if latents.dtype != latents_dtype:
if torch.backends.mps.is_available():
# some platforms (eg. apple mps) misbehave due to a pytorch bug: https://github.com/pytorch/pytorch/pull/99272
latents = latents.to(latents_dtype)
if i == len(timesteps) - 1 or (
(i + 1) > num_warmup_steps and (i + 1) % scheduler.order == 0
):
progress_bar.update()
if get_sequence_parallel_state():
latents = all_gather(latents, dim=2)
if output_type == "latent":
video = latents
else:
# unscale/denormalize the latents
# denormalize with the mean and std if available and not None
has_latents_mean = (
hasattr(vae.config, "latents_mean") and vae.config.latents_mean is not None
)
has_latents_std = (
hasattr(vae.config, "latents_std") and vae.config.latents_std is not None
)
if has_latents_mean and has_latents_std:
latents_mean = (
torch.tensor(vae.config.latents_mean)
.view(1, 12, 1, 1, 1)
.to(latents.device, latents.dtype)
)
latents_std = (
torch.tensor(vae.config.latents_std)
.view(1, 12, 1, 1, 1)
.to(latents.device, latents.dtype)
)
latents = latents * latents_std / vae.config.scaling_factor + latents_mean
else:
latents = latents / vae.config.scaling_factor
video = vae.decode(latents, return_dict=False)[0]
video_processor = VideoProcessor(vae_scale_factor=vae_spatial_scale_factor)
video = video_processor.postprocess_video(video, output_type=output_type)
return (video,)
@torch.no_grad()
@torch.autocast("cuda", dtype=torch.bfloat16)
def log_validation(
args,
transformer,
device,
weight_dtype,
global_step,
scheduler_type="euler",
shift=1.0,
num_euler_timesteps=100,
linear_quadratic_threshold=0.025,
linear_range=0.5,
ema=False,
):
# TODO
print(f"Running validation....\n")
vae = AutoencoderKLMochi.from_pretrained(
args.pretrained_model_name_or_path, subfolder="vae", torch_dtype=weight_dtype
).to("cuda")
vae.enable_tiling()
if scheduler_type == "euler":
scheduler = FlowMatchEulerDiscreteScheduler()
else:
linear_quadraic = True if scheduler_type == "pcm_linear_quadratic" else False
scheduler = PCMFMScheduler(
1000,
shift,
num_euler_timesteps,
linear_quadraic,
linear_quadratic_threshold,
linear_range,
)
# args.validation_prompt_dir
validation_guidance_scale_ls = args.validation_guidance_scale.split(",")
validation_guidance_scale_ls = [
float(scale) for scale in validation_guidance_scale_ls
]
for validation_sampling_step in args.validation_sampling_steps.split(","):
validation_sampling_step = int(validation_sampling_step)
for validation_guidance_scale in validation_guidance_scale_ls:
videos = []
# prompt_embed are named embed0 to embedN
# check how many embeds are there
num_embeds = len(
[f for f in os.listdir(args.validation_prompt_dir) if "embed" in f]
)
validation_prompt_ids = list(range(num_embeds))
num_sp_groups = int(os.getenv("WORLD_SIZE", "1")) // nccl_info.sp_size
# pad to multiple of groups
validation_prompt_ids += [0] * (num_sp_groups - num_embeds % num_sp_groups)
num_embeds_per_group = len(validation_prompt_ids) // num_sp_groups
local_prompt_ids = validation_prompt_ids[
nccl_info.group_id * num_embeds_per_group : (nccl_info.group_id + 1)
* num_embeds_per_group
]
for i in local_prompt_ids:
prompt_embed_path = os.path.join(
args.validation_prompt_dir, f"embed{i}.pt"
)
prompt_mask_path = os.path.join(
args.validation_prompt_dir, f"mask{i}.pt"
)
prompt_embeds = (
torch.load(prompt_embed_path, map_location="cpu", weights_only=True)
.to(device)
.to(weight_dtype)
.unsqueeze(0)
)
prompt_attention_mask = (
torch.load(prompt_mask_path, map_location="cpu", weights_only=True)
.to(device)
.to(weight_dtype)
.unsqueeze(0)
)
negative_prompt_embeds = (
torch.zeros(256, 4096).to(device).to(weight_dtype).unsqueeze(0)
)
negative_prompt_attention_mask = (
torch.zeros(256).bool().to(device).unsqueeze(0)
)
generator = torch.Generator(device="cuda").manual_seed(12345)
video = sample_validation_video(
transformer,
vae,
scheduler,
scheduler_type=scheduler_type,
num_frames=args.num_frames,
# Peiyuan TODO: remove hardcode
height=480,
width=848,
num_inference_steps=validation_sampling_step,
guidance_scale=validation_guidance_scale,
generator=generator,
prompt_embeds=prompt_embeds,
prompt_attention_mask=prompt_attention_mask,
negative_prompt_embeds=negative_prompt_embeds,
negative_prompt_attention_mask=negative_prompt_attention_mask,
)[0]
if nccl_info.rank_within_group == 0:
videos.append(video[0])
# collect videos from all process to process zero
gc.collect()
torch.cuda.empty_cache()
# log if main process
torch.distributed.barrier()
all_videos = [
None for i in range(int(os.getenv("WORLD_SIZE", "1")))
] # remove padded videos
torch.distributed.all_gather_object(all_videos, videos)
if nccl_info.global_rank == 0:
# remove padding
videos = [video for videos in all_videos for video in videos]
videos = videos[:num_embeds]
# linearize all videos
video_filenames = []
for i, video in enumerate(videos):
filename = os.path.join(
args.output_dir,
f"validation_step_{global_step}_sample_{validation_sampling_step}_guidance_{validation_guidance_scale}_video_{i}.mp4",
)
export_to_video(video, filename, fps=30)
video_filenames.append(filename)
logs = {
f"{'ema_' if ema else ''}validation_sample_{validation_sampling_step}_guidance_{validation_guidance_scale}": [
wandb.Video(filename)
for i, filename in enumerate(video_filenames)
]
}
wandb.log(logs, step=global_step)
+31
View File
@@ -0,0 +1,31 @@
"""fastvideo2 — a post-training-to-serving substrate for video models, MVP.
Four surfaces (see README.md at the repo root):
contracts fastvideo2.card / pipeline / loop — frozen data cards, enforced
stage edges, the driven-loop protocol
reference fastvideo2.<family>.reference — the standalone eager oracle
verifier fastvideo2.verify — tiered gates + evidence ledger
trace engine identity chain — request/stage/loop.step -> NVTX
``import fastvideo2`` is dependency-light: torch / diffusers / transformers
load lazily, only when weights are actually touched.
"""
from fastvideo2.card import ModelCard, derive
from fastvideo2.engine import Instance, Output, Request, run
from fastvideo2.registry import resolve
from fastvideo2.sdk import Model, Result, load
__version__ = "0.1.0"
__all__ = ["ModelCard", "derive", "Model", "Result", "load", "resolve",
"Instance", "Output", "Request", "run", "generate", "__version__"]
def generate(model: str, prompt: str, *, root: str | None = None,
device: str | None = None, **request_kwargs) -> Result:
"""One-call convenience over the SDK: load then generate (loads per call —
hold a :class:`Model` via :func:`load` to amortize residency).
>>> result = fastvideo2.generate("wan2.1-t2v-1.3b", "a cat surfing", seed=7)
>>> result.video.shape # [T, H, W, C] uint8
"""
return load(model, root=root, device=device).generate(prompt, **request_kwargs)
+3
View File
@@ -0,0 +1,3 @@
from fastvideo2.cli import main
raise SystemExit(main())
+213
View File
@@ -0,0 +1,213 @@
"""Model cards — the contract surface.
A card is a frozen, pure-data description of one servable artifact: its
components, the loops its weights assume, and the sampling defaults that are
part of the trained artifact. Components and loops are declared as
``"module:attr"`` reference strings and loop params must be plain JSON values
— no callables — so a card is:
* **serializable** — ``to_json``/``from_dict`` are lossless (T0-gated), so a
card ships as ``card.json`` beside a checkpoint or inside a deploy config;
* **content-addressed** — ``digest()`` hashes the canonical JSON (think git
object id) and is the card's identity in evidence records, T1 baselines,
and environment manifests. It names the declaration only; weights, code,
and environment drift are checked separately (T1, T2, env fingerprint).
Variants are expressed as diffs against a base card via :func:`derive` — never
as builder functions with keyword arguments. Derivation is additive: a variant
that needs to *remove* something picked the wrong base and should be declared
fresh.
Import discipline: this module is stdlib-only. ``validate()`` imports the
declared loop modules to check their ``semantics`` ids, so loop modules must be
importable without torch.
"""
from __future__ import annotations
import hashlib
import importlib
import json
from dataclasses import asdict, dataclass, field, fields, is_dataclass, replace
from typing import Any
class CardError(ValueError):
pass
def resolve_ref(ref: str) -> Any:
"""Resolve a ``"module:attr"`` reference string to the live object."""
mod, _, attr = ref.partition(":")
if not mod or not attr:
raise CardError(f"bad reference {ref!r} (expected 'module:attr')")
return getattr(importlib.import_module(mod), attr)
@dataclass(frozen=True)
class ComponentSpec:
"""One weight-bearing (or processing) component of the artifact."""
component_id: str
kind: str # dit | vae | text_encoder | tokenizer
module: str # loader reference, e.g. "fastvideo2.wan21.model:WanModel"
subfolder: str # subfolder in the checkpoint layout ("" = repo root)
dtype: str = "bf16" # bf16 | fp32 | "" (dtype-less, e.g. tokenizer)
source: str = "" # weights repo override; "" = the card-level `weights`
@dataclass(frozen=True)
class LoopSpec:
"""One iterative computation the card can run.
``loop`` names the implementation class; ``params`` are plain JSON values
passed to its constructor. The class carries a ``semantics`` id that
provenance pins (see ``Provenance.assumes_loop``).
"""
loop_id: str
loop: str # "fastvideo2.wan21.loop:WanDenoiseLoop"
params: dict[str, Any] = field(default_factory=dict)
@dataclass(frozen=True)
class SamplingDefaults:
"""Per-model generation defaults that are part of the trained artifact."""
num_steps: int
guidance_scale: float
height: int
width: int
num_frames: int
fps: int
shift: float
negative_prompt: str = ""
@dataclass(frozen=True)
class Provenance:
"""Where the weights came from and what they assume.
``assumes_loop`` is a *semantics id* (e.g. ``"wan.flow_euler.cfg/v1"``),
not a loop_id: validation resolves every declared loop class and requires
one whose ``semantics`` matches. A distilled student that requires a
different sampler therefore cannot validate against a base card.
``substitution`` classifies this artifact relative to ``parents``:
``exact`` | ``bounded`` | ``quality-changing``.
"""
method: str = "base"
parents: tuple[str, ...] = ()
assumes_loop: str = ""
precision: str = "bf16"
substitution: str = "exact"
tolerances: dict[str, float] = field(default_factory=dict)
@dataclass(frozen=True)
class ModelCard:
model_id: str
family: str
weights: str # canonical source (HF repo id)
components: dict[str, ComponentSpec]
loops: dict[str, LoopSpec]
capabilities: tuple[str, ...]
provenance: Provenance
sampling_defaults: SamplingDefaults
determinism: str = "tolerance" # bitwise | tolerance
# --- identity ---------------------------------------------------------- #
def to_dict(self) -> dict:
return asdict(self)
def to_json(self) -> str:
return json.dumps(self.to_dict(), sort_keys=True, indent=2)
def digest(self) -> str:
"""Content digest over the canonical JSON — the card's identity in the
evidence ledger and every environment manifest."""
canon = json.dumps(self.to_dict(), sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canon.encode()).hexdigest()[:16]
@classmethod
def from_dict(cls, d: dict) -> "ModelCard":
return cls(
model_id=d["model_id"],
family=d["family"],
weights=d["weights"],
components={k: ComponentSpec(**v) for k, v in d["components"].items()},
loops={k: LoopSpec(**v) for k, v in d["loops"].items()},
capabilities=tuple(d["capabilities"]),
provenance=Provenance(**{**d["provenance"], "parents": tuple(d["provenance"]["parents"])}),
sampling_defaults=SamplingDefaults(**d["sampling_defaults"]),
determinism=d.get("determinism", "tolerance"),
)
# --- validation -------------------------------------------------------- #
def validate(self) -> "ModelCard":
errs: list[str] = []
if not self.components:
errs.append("card declares no components")
if not self.loops:
errs.append("card declares no loops")
for cid, spec in self.components.items():
if cid != spec.component_id:
errs.append(f"component key {cid!r} != component_id {spec.component_id!r}")
semantics_seen: list[str] = []
for lid, spec in self.loops.items():
if lid != spec.loop_id:
errs.append(f"loop key {lid!r} != loop_id {spec.loop_id!r}")
try:
cls = resolve_ref(spec.loop)
except Exception as e: # unresolvable ref is a contract violation
errs.append(f"loop {lid!r}: cannot resolve {spec.loop!r} ({e})")
continue
sem = getattr(cls, "semantics", None)
if not sem:
errs.append(f"loop {lid!r}: class {spec.loop!r} declares no `semantics` id")
else:
semantics_seen.append(sem)
try:
json.dumps(spec.params)
except TypeError:
errs.append(f"loop {lid!r}: params are not plain JSON values")
# the teeth: weights may only be served under a loop whose semantics
# they were trained for.
if self.provenance.assumes_loop and self.provenance.assumes_loop not in semantics_seen:
errs.append(
f"provenance.assumes_loop={self.provenance.assumes_loop!r} matches no declared "
f"loop semantics (have {semantics_seen}) — these weights cannot run on this card")
if self.determinism not in ("bitwise", "tolerance"):
errs.append(f"unknown determinism class {self.determinism!r}")
if errs:
raise CardError(f"ModelCard {self.model_id!r} failed validation:\n - " + "\n - ".join(errs))
return self
def _merge_field(old: Any, patch: Any) -> Any:
"""One-level structural merge used by :func:`derive`.
dict field + dict patch -> merge by key (spec values replace; dict
values patch the existing spec/dict)
dataclass field + dict patch -> replace() with recursively merged fields
anything else -> the patch value wins
"""
if is_dataclass(old) and isinstance(patch, dict):
merged = {k: _merge_field(getattr(old, k), v) for k, v in patch.items()}
return replace(old, **merged)
if isinstance(old, dict) and isinstance(patch, dict):
out = dict(old)
for k, v in patch.items():
out[k] = _merge_field(old[k], v) if k in old else v
return out
return patch
def derive(base: ModelCard, **delta: Any) -> ModelCard:
"""A variant as an explicit diff against a base card.
Additive only: keys merge, nothing is deleted. A variant that must remove a
component or loop is a different architecture — declare it fresh. The
derived card re-validates, so an invalid diff fails at declaration.
"""
valid = {f.name for f in fields(ModelCard)}
unknown = set(delta) - valid
if unknown:
raise CardError(f"derive: unknown card fields {sorted(unknown)}")
merged = {k: _merge_field(getattr(base, k), v) for k, v in delta.items()}
return replace(base, **merged).validate()
+92
View File
@@ -0,0 +1,92 @@
"""CLI: describe / generate / verify — the three agent-facing verbs.
python -m fastvideo2 describe wan2.1-t2v-1.3b
python -m fastvideo2 generate wan2.1-t2v-1.3b --prompt "a cat surfing" --out cat.mp4
python -m fastvideo2 verify wan2.1-t2v-1.3b --tier 2 [--bless]
``describe`` prints the card as JSON plus its digest — machine-readable
capability discovery. ``verify`` appends typed results to the evidence ledger
and exits non-zero on any failed gate.
"""
from __future__ import annotations
import argparse
import json
import sys
def _describe(args) -> int:
from fastvideo2.registry import resolve
card, _ = resolve(args.model)
print(card.to_json())
print(f'// digest: {card.digest()}', file=sys.stderr)
return 0
def _generate(args) -> int:
import fastvideo2
kwargs = {k: getattr(args, k) for k in
("seed", "num_steps", "guidance_scale", "height", "width", "num_frames", "shift")
if getattr(args, k) is not None}
model = fastvideo2.load(args.model, root=args.root, device=args.device)
result = model.generate(args.prompt, **kwargs)
result.save(args.out, fps=args.fps)
steps = [t for t in result.trace if "/denoise." in t["label"]]
print(f"video {result.video.shape} -> {args.out}")
print(f"total {result.seconds:.1f}s; denoise steps {len(steps)}, "
f"mean {sum(t['seconds'] for t in steps) / max(len(steps), 1):.2f}s/step")
return 0
def _verify(args) -> int:
from fastvideo2.verify import LEDGER, verify
results = verify(args.model, tier=args.tier, root=args.root, device=args.device,
bless=args.bless, anchor=args.anchor)
for r in results:
mark = {"pass": "PASS ", "blessed": "BLESS", "fail": "FAIL "}[r.status]
print(f" {mark} {r.gate:14s} {r.detail or json.dumps(r.metrics)[:120]}")
print(f"ledger: {LEDGER}")
return 0 if all(r.ok for r in results) else 1
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(prog="fastvideo2", description=__doc__)
sub = p.add_subparsers(dest="cmd", required=True)
d = sub.add_parser("describe", help="print a card as JSON + digest")
d.add_argument("model")
d.set_defaults(fn=_describe)
g = sub.add_parser("generate", help="run one request, save an mp4")
g.add_argument("model")
g.add_argument("--prompt", required=True)
g.add_argument("--out", default="out.mp4")
g.add_argument("--root", default=None, help="local checkpoint dir (else HF cache)")
g.add_argument("--device", default=None)
g.add_argument("--seed", type=int, default=0)
g.add_argument("--num-steps", dest="num_steps", type=int, default=None)
g.add_argument("--guidance-scale", dest="guidance_scale", type=float, default=None)
g.add_argument("--height", type=int, default=None)
g.add_argument("--width", type=int, default=None)
g.add_argument("--num-frames", dest="num_frames", type=int, default=None)
g.add_argument("--shift", type=float, default=None)
g.add_argument("--fps", type=int, default=16)
g.set_defaults(fn=_generate)
v = sub.add_parser("verify", help="run tiered gates; append to the evidence ledger")
v.add_argument("model")
v.add_argument("--tier", type=int, default=3, choices=(0, 1, 2, 3))
v.add_argument("--root", default=None)
v.add_argument("--device", default=None)
v.add_argument("--bless", action="store_true",
help="write the T1 fingerprint baseline for this environment")
v.add_argument("--anchor", action="store_true",
help="also certify components against the official Wan2.1 goldens")
v.set_defaults(fn=_verify)
args = p.parse_args(argv)
return args.fn(args)
if __name__ == "__main__":
raise SystemExit(main())
+4
View File
@@ -0,0 +1,4 @@
"""Dreamverse runtime on fastvideo2 — session WS + fMP4 segments. See server.py."""
from fastvideo2.dreamverse.server import build_app, main
__all__ = ["build_app", "main"]
+3
View File
@@ -0,0 +1,3 @@
from fastvideo2.dreamverse.server import main
main()
@@ -0,0 +1,100 @@
"""Dreamverse runtime anchor: boot the server with DUMMY prompt keys, drive
one full session over the protocol, and assert the streaming contract:
segment_start -> live step_complete x3 -> media_init -> binary fMP4 chunks
(first chunk carries an ISO-BMFF `ftyp` box) -> media_segment_complete ->
segment_complete with a latents sha.
Usage (cluster): python -m fastvideo2.dreamverse.gates.dreamverse_anchor
"""
from __future__ import annotations
import asyncio
import json
import os
import subprocess
import sys
import time
import urllib.request
MODEL = "fastwan-qad-fp8-1.3b"
PORT = 8019
PROMPT = "A raccoon in a field of sunflowers, warm light, mid-shot."
def main() -> int:
env = dict(os.environ, CEREBRAS_API_KEY="dummy", GROQ_API_KEY="dummy")
server = subprocess.Popen(
[sys.executable, "-m", "fastvideo2.dreamverse", "--model", MODEL,
"--port", str(PORT)], env=env,
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
try:
for _ in range(360):
try:
with urllib.request.urlopen(
f"http://127.0.0.1:{PORT}/health", timeout=5) as r:
if json.loads(r.read())["model"] == MODEL:
break
except Exception:
time.sleep(2)
else:
raise RuntimeError("server never became healthy")
import websockets
async def session() -> dict:
counts = {"steps": 0, "chunks": 0, "ftyp": False}
async with websockets.connect(
f"ws://127.0.0.1:{PORT}/ws", max_size=None) as ws:
await ws.send(json.dumps({"type": "session_init_v2",
"enhancement": True}))
for expected in ("queue_status", "gpu_assigned", "stream_start"):
got = json.loads(await ws.recv())["type"]
assert got == expected, (got, expected)
await ws.send(json.dumps({"type": "segment_prompt_source",
"prompt": PROMPT, "seed": 7}))
while True:
raw = await asyncio.wait_for(ws.recv(), timeout=900)
if isinstance(raw, bytes):
if counts["chunks"] == 0:
counts["ftyp"] = b"ftyp" in raw[:64]
counts["chunks"] += 1
continue
msg = json.loads(raw)
t = msg["type"]
if t == "step_complete":
counts["steps"] += 1
elif t == "segment_complete":
counts["latents_sha"] = msg["latents_sha"]
counts["frames"] = msg["frames"]
break
elif t == "error":
raise RuntimeError(msg)
await ws.send(json.dumps({"type": "leave"}))
assert json.loads(await ws.recv())["type"] == "stream_complete"
return counts
c = asyncio.run(session())
finally:
server.terminate()
server.wait(timeout=30)
ok = (c["steps"] == 3 and c["chunks"] >= 1 and c["ftyp"]
and c.get("frames", 0) == 81)
print(f"steps={c['steps']} chunks={c['chunks']} ftyp={c['ftyp']} "
f"frames={c.get('frames')} sha={c.get('latents_sha')} "
f"{'OK' if ok else 'FAIL'}")
from fastvideo2.verify import GateResult, append_ledger, env_fingerprint
append_ledger([GateResult(gate="anchor.dreamverse-runtime",
status="pass" if ok else "fail", model_id=MODEL,
card_digest="-",
metrics={"steps": float(c["steps"]),
"chunks": float(c["chunks"]),
"ftyp": 1.0 if c["ftyp"] else 0.0},
tolerances={}, env=env_fingerprint(),
detail=f"latents {c.get('latents_sha')}")])
return 0 if ok else 1
if __name__ == "__main__":
sys.exit(main())
+269
View File
@@ -0,0 +1,269 @@
"""Dreamverse runtime on fastvideo2 — the realtime video session server.
Ported from ``apps/dreamverse`` (fastvideo-main): the WebSocket session
protocol (``session_init_v2`` → per-segment prompts → fMP4 fragments over
the socket), the ffmpeg fragmented-MP4 encoder (verbatim flags from
``entrypoints/streaming/stream.py``: libx264, zerolatency,
``empty_moov+default_base_moof+frag_keyframe+faststart``), and an optional
Cerebras/Groq prompt enhancer (boots with dummy keys — enhancement simply
stays off, the same bring-up shortcut the GB200 deploys used).
Deliberately re-based for v2.1 (the original is LTX2-specific — audio,
refine stage, continuation-state, LoRA stack): segments generate through the
fastvideo2 SDK on the FastWan 3-step DMD student (seconds per segment on
GB200), and per-step progress is LIVE via the engine's ``on_step`` hook
(``step_complete`` per denoise step — the original emits one terminal event
per segment). Message names follow the upstream protocol so their web client
schema maps directly; LTX2-only fields are ignored.
Run: python -m fastvideo2.dreamverse --model fastwan-qad-fp8-1.3b --port 8009
ffmpeg: FASTVIDEO_FFMPEG_BIN, or PATH, or the imageio-ffmpeg bundled binary.
"""
from __future__ import annotations
import asyncio
import hashlib
import json
import os
import shutil
import subprocess
import threading
import urllib.request
import uuid
from typing import Any
def find_ffmpeg() -> str:
p = os.environ.get("FASTVIDEO_FFMPEG_BIN") or shutil.which("ffmpeg")
if p:
return p
try: # the GB200 bring-up shortcut: pip-installed bundled binary
import imageio_ffmpeg
return imageio_ffmpeg.get_ffmpeg_exe()
except ImportError as e:
raise RuntimeError("no ffmpeg (set FASTVIDEO_FFMPEG_BIN, install "
"ffmpeg, or `pip install imageio-ffmpeg`)") from e
def fmp4_encode(frames: Any, *, fps: int, ffmpeg: str) -> list[bytes]:
"""One segment -> fragmented-MP4 byte chunks (upstream's exact flags)."""
t, h, w, _ = frames.shape
args = [ffmpeg, "-hide_banner", "-loglevel", "error",
"-f", "rawvideo", "-pix_fmt", "rgb24", "-s", f"{w}x{h}",
"-r", str(fps), "-i", "-",
"-c:v", "libx264", "-preset", "ultrafast", "-tune", "zerolatency",
"-pix_fmt", "yuv420p",
"-movflags", "empty_moov+default_base_moof+frag_keyframe+faststart",
"-f", "mp4", "-"]
proc = subprocess.Popen(args, stdin=subprocess.PIPE, stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL, bufsize=0)
out: list[bytes] = []
def _read() -> None:
while True:
chunk = proc.stdout.read(65536)
if not chunk:
break
out.append(chunk)
reader = threading.Thread(target=_read, daemon=True)
reader.start()
for i in range(t):
proc.stdin.write(frames[i].tobytes())
proc.stdin.close()
proc.wait(timeout=120)
reader.join(timeout=30)
return out
class PromptEnhancer:
"""Cerebras-or-Groq chat call (upstream's provider pair, gpt-oss-120b).
Dummy/missing keys or any failure -> pass the prompt through unchanged."""
def __init__(self) -> None:
self.cerebras = os.environ.get("CEREBRAS_API_KEY", "")
self.groq = os.environ.get("GROQ_API_KEY", "")
self.enabled = any(k and k != "dummy" for k in (self.cerebras, self.groq))
def enhance(self, prompt: str, history: list[str]) -> str:
if not self.enabled:
return prompt
targets = []
if self.cerebras and self.cerebras != "dummy":
targets.append(("https://api.cerebras.ai/v1/chat/completions",
self.cerebras, "gpt-oss-120b"))
if self.groq and self.groq != "dummy":
targets.append(("https://api.groq.com/openai/v1/chat/completions",
self.groq, "openai/gpt-oss-120b"))
system = ("Rewrite the user's next-video-segment prompt into one vivid, "
"concrete shot description. Prior segments: "
+ " | ".join(history[-3:]))
for url, key, model_name in targets:
try:
req = urllib.request.Request(
url, method="POST",
headers={"Authorization": f"Bearer {key}",
"Content-Type": "application/json"},
data=json.dumps({"model": model_name, "temperature": 1.0,
"messages": [{"role": "system", "content": system},
{"role": "user", "content": prompt}]
}).encode())
with urllib.request.urlopen(req, timeout=20) as r:
return json.loads(r.read())["choices"][0]["message"]["content"].strip()
except Exception:
continue
return prompt
def build_app(model: Any) -> Any:
from fastapi import FastAPI
from starlette.routing import WebSocketRoute
from starlette.websockets import WebSocketDisconnect
from fastvideo2.engine import Request
from fastvideo2.engine import run as engine_run
from fastvideo2.sdk import Result
app = FastAPI(title="dreamverse-fv2", version="0.1")
ffmpeg = find_ffmpeg()
enhancer = PromptEnhancer()
gen_lock = threading.Lock()
@app.get("/health")
def health() -> dict:
return {"status": "ok", "model": model.model_id,
"enhancer": enhancer.enabled, "ffmpeg": ffmpeg}
@app.get("/readyz")
def readyz() -> dict:
return {"ready": True}
async def ws_session(ws) -> None:
await ws.accept()
try:
init = await ws.receive_json()
except WebSocketDisconnect:
return
if init.get("type") != "session_init_v2":
await ws.send_json({"type": "error", "code": "bad_init",
"error": "expected session_init_v2"})
await ws.close()
return
session_id = uuid.uuid4().hex[:12]
enhancement_on = bool(init.get("enhancement", False)) and enhancer.enabled
history: list[str] = []
segment_idx = 0
await ws.send_json({"type": "queue_status", "position": 0})
await ws.send_json({"type": "gpu_assigned", "session_id": session_id})
await ws.send_json({"type": "stream_start", "session_id": session_id,
"model": model.model_id})
loop = asyncio.get_running_loop()
while True:
try:
msg = await ws.receive_json()
except WebSocketDisconnect:
return
mtype = msg.get("type")
if mtype == "leave":
await ws.send_json({"type": "stream_complete",
"segments": segment_idx})
await ws.close()
return
if mtype == "enhancement_updated":
enhancement_on = bool(msg.get("enabled")) and enhancer.enabled
continue
if mtype != "segment_prompt_source":
await ws.send_json({"type": "error", "code": "bad_message",
"error": f"unsupported type {mtype!r}"})
continue
prompt = str(msg.get("prompt", ""))
if not prompt:
await ws.send_json({"type": "error", "code": "bad_prompt",
"error": "prompt required"})
continue
if enhancement_on:
prompt = await loop.run_in_executor(
None, enhancer.enhance, prompt, history)
history.append(prompt)
await ws.send_json({"type": "segment_start", "segment": segment_idx,
"prompt": prompt})
q: asyncio.Queue = asyncio.Queue()
def on_step(label: str, seconds: float, meta: dict) -> None:
loop.call_soon_threadsafe(
q.put_nowait, {"type": "step_complete", "label": label,
"seconds": round(seconds, 4)})
def generate(p: str = prompt, seed: Any = msg.get("seed", 0),
steps: Any = msg.get("num_steps")) -> None:
try:
req = Request(prompt=p, request_id=f"{session_id}-{segment_idx}",
seed=int(seed), num_steps=steps)
with gen_lock:
out = engine_run(model.instance, model.pipeline, req,
on_step=on_step)
result = Result(outputs=out.outputs, trace=out.trace,
request=req.resolve(model.card),
model_id=model.model_id,
card_digest=model.card.digest(),
fps=model.card.sampling_defaults.fps)
chunks = fmp4_encode(result.video, fps=result.fps,
ffmpeg=ffmpeg)
import torch
sha = hashlib.sha256(result.latents.detach().to(
torch.float32).cpu().numpy().tobytes()).hexdigest()[:16]
loop.call_soon_threadsafe(
q.put_nowait, {"__chunks": chunks, "latents_sha": sha,
"frames": int(result.video.shape[0])})
except Exception as e:
loop.call_soon_threadsafe(
q.put_nowait, {"type": "error", "code": "generation",
"error": f"{type(e).__name__}: {e}"})
threading.Thread(target=generate, daemon=True).start()
while True:
ev = await q.get()
if "__chunks" in ev:
await ws.send_json({"type": "media_init",
"segment": segment_idx,
"mime": 'video/mp4; codecs="avc1"'})
for chunk in ev["__chunks"]:
await ws.send_bytes(chunk)
await ws.send_json({"type": "media_segment_complete",
"segment": segment_idx})
await ws.send_json({"type": "segment_complete",
"segment": segment_idx,
"frames": ev["frames"],
"latents_sha": ev["latents_sha"]})
segment_idx += 1
break
await ws.send_json(ev)
if ev.get("type") == "error":
break
app.router.routes.append(WebSocketRoute("/ws", ws_session))
return app
def main(argv: list[str] | None = None) -> None:
import argparse
import uvicorn
import fastvideo2 as fv2
p = argparse.ArgumentParser("fastvideo2.dreamverse")
p.add_argument("--model", default="fastwan-qad-fp8-1.3b")
p.add_argument("--host", default="127.0.0.1")
p.add_argument("--port", type=int, default=8009)
p.add_argument("--device", default=None)
args = p.parse_args(argv)
model = fv2.load(args.model, device=args.device)
uvicorn.run(build_app(model), host=args.host, port=args.port)
if __name__ == "__main__":
main()
+178
View File
@@ -0,0 +1,178 @@
"""The engine — drives a pipeline over one resident instance, one request at a
time, with the identity chain attached.
Identity chain: every unit of work is named ``request/stage`` for one-shot
stages and ``request/stage/loop.step`` for loop steps. The same name goes to
(a) the returned trace (typed timings, machine-readable) and (b) NVTX ranges
when CUDA is present — so Nsight correlates kernels to model-level identity
with no extra instrumentation.
Deliberately absent (this is the one-shot MVP): queueing, admission, batching,
sessions, cancellation. Sessions with forkable state are the next consumer of
the loop contract, not a reason to grow this file now.
"""
from __future__ import annotations
import contextlib
from dataclasses import dataclass, field, replace
from typing import Any
from fastvideo2.card import ModelCard
from fastvideo2.loading import load_component, resolve_weights
from fastvideo2.loop import LoopRunner, build_loop
from fastvideo2.pipeline import ComponentStage, LoopStage, Pipeline, run_component_stage
@dataclass(frozen=True)
class Request:
"""One generation request. ``None`` fields resolve from the card's
sampling defaults via :meth:`resolve`."""
prompt: str
request_id: str = "req0"
negative_prompt: str | None = None
seed: int = 0
num_steps: int | None = None
guidance_scale: float | None = None
height: int | None = None
width: int | None = None
num_frames: int | None = None
shift: float | None = None
capture_trajectory: bool = False
def resolve(self, card: ModelCard) -> "Request":
d = card.sampling_defaults
fill = {
"negative_prompt": d.negative_prompt,
"num_steps": d.num_steps,
"guidance_scale": d.guidance_scale,
"height": d.height,
"width": d.width,
"num_frames": d.num_frames,
"shift": d.shift,
}
patch = {k: v for k, v in fill.items() if getattr(self, k) is None}
return replace(self, **patch)
@dataclass
class Output:
request_id: str
outputs: dict[str, Any]
trace: list[dict] = field(default_factory=list) # [{label, seconds, ...meta}]
@property
def seconds(self) -> float:
return sum(t["seconds"] for t in self.trace)
class Instance:
"""A resident, loaded card: components materialize lazily and are shared
by reference; loops are built from the card's declared specs."""
def __init__(self, card: ModelCard, root: str | None = None, device: str = "cpu"):
self.card = card
self.device = device
self.root = resolve_weights(card, root)
self._components: dict[str, Any] = {}
self._source_roots: dict[str, str] = {}
self._loops: dict[str, Any] = {}
def component(self, component_id: str) -> Any:
if component_id not in self._components:
spec = self.card.components.get(component_id)
if spec is None:
raise KeyError(f"component {component_id!r} not declared on card {self.card.model_id!r}")
root = self._source_root(spec.source) if spec.source else self.root
self._components[component_id] = load_component(spec, root, self.device)
return self._components[component_id]
def _source_root(self, source: str) -> str:
"""Resolve a per-component weights source (e.g. the official-layout
transformer repo) through the same snapshot cache as card weights."""
if source not in self._source_roots:
from huggingface_hub import snapshot_download
self._source_roots[source] = snapshot_download(source)
return self._source_roots[source]
def loop(self, loop_id: str) -> Any:
if loop_id not in self._loops:
spec = self.card.loops.get(loop_id)
if spec is None:
raise KeyError(f"loop {loop_id!r} not declared on card {self.card.model_id!r}")
self._loops[loop_id] = build_loop(spec)
return self._loops[loop_id]
def load(card: ModelCard, root: str | None = None, device: str | None = None) -> Instance:
"""The public entrypoint: card + weights root -> resident instance."""
card.validate()
if device is None:
device = _detect_device()
return Instance(card, root=root, device=device)
def _detect_device() -> str:
try:
import torch
if torch.cuda.is_available():
return "cuda"
except ImportError:
pass
return "cpu"
@contextlib.contextmanager
def _nvtx(name: str):
"""NVTX range when CUDA is live; free otherwise."""
pushed = False
try:
import torch
if torch.cuda.is_available():
torch.cuda.nvtx.range_push(name)
pushed = True
except ImportError:
pass
try:
yield
finally:
if pushed:
import torch
torch.cuda.nvtx.range_pop()
def run(instance: Instance, pipeline: Pipeline, request: Request,
on_step: Any = None) -> Output:
"""Run one request through a pipeline to completion. ``on_step(label,
seconds, meta)``, when given, fires live after every loop step (serving
streams progress through it)."""
import time
req = request.resolve(instance.card)
trace: list[dict] = []
slots: dict[str, Any] = {}
for name in pipeline.inputs: # request-provided slots, by attribute name
slots[name] = getattr(req, name)
for stage in pipeline.stages:
chain = f"{req.request_id}/{stage.stage_id}"
if isinstance(stage, ComponentStage):
with _nvtx(chain):
t0 = time.perf_counter()
run_component_stage(stage, instance, slots, req)
trace.append({"label": chain, "seconds": time.perf_counter() - t0})
elif isinstance(stage, LoopStage):
loop = instance.loop(stage.loop_id)
def observe(label: str, seconds: float, meta: dict, _chain: str = chain) -> None:
trace.append({"label": f"{_chain}/{label}", "seconds": seconds, **meta})
if on_step is not None:
on_step(f"{_chain}/{label}", seconds, meta)
inputs = {k: slots[k] for k in stage.reads}
with _nvtx(chain):
runner = LoopRunner(loop, req, instance, inputs, observe=observe)
slots[stage.writes[0]] = runner.run()
else:
raise TypeError(f"unknown stage kind {type(stage).__name__}")
outputs = {name: slots[slot] for name, slot in pipeline.outputs.items()}
return Output(request_id=req.request_id, outputs=outputs, trace=trace)
+18
View File
@@ -0,0 +1,18 @@
# Evidence ledger
Typed verification records, committed with the code they vouch for.
- `ledger.jsonl` — append-only `GateResult` records: gate, status, card digest,
metrics, tolerances, environment fingerprint, timestamp. Written only by
`python -m fastvideo2 verify`; never edited by hand.
- `<model_id>.fingerprints.json` — the blessed T1 component baseline for one
card digest in one environment.
- `sample_*.mp4` — eyeballable artifacts from full-scale runs (e.g.
`sample_wan21_seed7.mp4`, 50 steps / 81 frames on GB200, byte-identical
between the production pipeline and `reference.py` at the same seed). Re-bless deliberately (`verify --bless`)
when the card or environment legitimately changes; a digest mismatch is a
failure, not a skip.
Ownership rule: baselines and gate tolerances are human-owned. Agents run the
gates and append evidence; they do not re-bless baselines to make a failure
disappear.
@@ -0,0 +1,82 @@
# FastWan variants — bitwise alignment vs fastvideo-main
Authority for FastWan artifacts is **fastvideo main** (they were distilled in
that stack); the alignment target was bit-exactness against main's own serving
path, pinned to main's exposed knobs so the goldens measure the artifact, not
the accelerator stack: `FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN` (QAD) /
`VIDEO_SPARSE_ATTN` (VSA), FP8 per-tensor dynamic quant (QAD), no torch.compile,
no FSDP, single GB200.
Goldens captured at main commit `c459a1897899ffcec3be7765534d81000b9bb9c1`;
the vendored forward (`wan21/model_fv.py`) was read at `e3f47dc2de2d…` — all
10 numerics-relevant source files verified byte-identical between the two.
## Final anchor results (all rows target 0.0 — bitwise)
| row | fastwan-qad-fp8-1.3b | fastwan-t2v-1.3b (VSA) |
|---|---|---|
| dit bf16 probes t∈{1000,757,522} | 0.0 / 0.0 / 0.0 | — |
| dit fp8 probes t∈{1000,757,522} | 0.0 / 0.0 / 0.0 | — |
| dit vsa probes t∈{1000,757,522} | — | 0.0 / 0.0 / 0.0 |
| text_encoder (e2e + probe prompts) | 0.0 / 0.0 | 0.0 / 0.0 |
| e2e step-1 / step-2 latent chain | 0.0 / 0.0 | 0.0 / 0.0 |
| e2e final latents (81f, 480×832, 3 steps) | 0.0 | 0.0 |
Ledger: `anchor.fastwan-qad-main` and `anchor.fastwan-vsa-main`, both `pass`
(card digests `1c8e6f7d1380552d`, `0bd5c7771e10ce44`).
## Root causes found by the gates (in discovery order)
1. **fp8 quantization is device-sensitive.** main converts weights to fp8 on
the GPU (post-materialization); quantizing the *identical* bf16 weights on
CPU produces different fp8 codes often enough to move a full forward by
~4e-2 rel (fp8's coarse grid amplifies conversion-tie differences).
Fix: `FP8Linear` defers quantization to first forward on the serving
device (`layers/fp8.py`).
2. **0-dim sigma tensors demote the renoise mixing to bf16.** In torch type
promotion 0-dim tensors act as scalars, so `(1-σ)*x0 + σ*ε` with a 0-dim
fp32 σ ran in bf16 (two roundings); main's `[B,1,1,1]` fp32 σ promotes the
arithmetic to fp32 with one final bf16 cast. 3.2e-3 per step, compounding
to 8.2e-2 over 3 steps. Fix: non-0-dim σ in `WanDMDLoop`.
3. **main's DMD sigma table is NOT the one the code appears to prepare.**
`DmdDenoisingStage.__init__` hardcodes a fresh internal
`FlowMatchEulerDiscreteScheduler(shift=8.0)`; the pipeline scheduler that
`TimestepPreparationStage.set_timesteps(n)` configured is never consulted,
and the config `flow_shift` is ignored. Lookups run against the 1000-entry
warped **init** table: σ(1000)=1.0, σ(757)=0.7567567, σ(522)=0.5217391
(confirmed in the capture manifests). `dmd_inference_table` reproduces
this exactly; a canary T0 test guards it.
Also confirmed en route: the CPU-generator RNG stream (initial fp32 draw +
bf16 renoise draws), the fp64 x0 math, and the flash/dense forward at full
81-frame geometry are each independently bitwise (triage decomposition in
session evidence).
## Caveats
- Text parity holds for ASCII prompts; main's ftfy cleaning diverges from
official's on CJK width-folding (see wan21 report — main measured 4.19e-1
vs official on the Chinese negative prompt). FastWan cards reuse the wan21
text stage; DMD uses no negative prompt. Add main's clean fn + a CJK golden
before serving non-ASCII prompts against these cards.
- VAE decode is not bitwise-gated (shared component; wan21 anchors cover it);
golden videos are in the goldens dirs for SSIM-level comparison.
- Committed goldens are trimmed (per-step model outputs and the reproducible
step-0 input dropped); the full set regenerates via
`gates/capture_fastvideo_main.py {qad,vsa}` — one command, pinned config.
## SFWan (self-forcing causal) — added 2026-07-23
`sfwan-t2v-1.3b` (wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers) anchored bitwise on
the FIRST complete run: all **35/35 chunk-rollout forwards** (7 blocks x
4 warped DMD steps + context pass) hash-match main's CausalDMDDenosingStage
exactly, text 0.0, e2e final latents 0.0 (`anchor.sfwan-main` pass).
Causal-specific semantics vendored (each different from BOTH other Wan
forwards): per-frame temb `[B, T_temb, 6, dim]`; ALL-bf16 modulation (no
fp32 promotion anywhere); plain bf16 LayerNorms; fp64 RoPE multipliers at
absolute positions (start_frame offsets); block-causal KV cache
(21-frame global window, `.detach()` on writes — training rollout reuses
this same module); cached text cross-attention; warp table =
SelfForcingFlowMatchScheduler(shift 5, extra_one_step) rows
`[1000, 937.5, 833.33, 625]` self-indexing their own sigmas.
@@ -0,0 +1,51 @@
{
"repo": "FastVideo/FastWan-QAD-FP8-1.3B",
"snapshot": "3de0eec0e2562923d38a87344127a86a35a3c11d",
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"fastvideo_src": "/mnt/FastVideo",
"torch": "2.12.0+cu130",
"flash_attn": "2.8.3",
"python": "3.12.13",
"gpu": "NVIDIA GB200",
"attention_backend": "FLASH_ATTN",
"quant": "FP8 per-tensor (dynamic act, post-load weight quant from bf16)",
"vsa_sparsity": null,
"seed": 1234,
"e2e_prompt": "A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest. The playful yet serene atmosphere is complemented by soft natural light filtering through the petals. Mid-shot, warm and cheerful tones.",
"probe_prompt": "A cat and a dog baking a cake together in a kitchen.",
"probe_timesteps": [
1000,
757,
522
],
"probe_latent_bcfhw": [
1,
16,
5,
60,
104
],
"e2e_latent_btchw": [
1,
21,
16,
60,
104
],
"dmd_denoising_steps": [
1000,
757,
522
],
"scheduler": {
"class": "FlowMatchEulerDiscreteScheduler",
"shift": 8.0,
"table_len": 1000,
"sigma_lookup": {
"1000": 1.0,
"757": 0.7567567229270935,
"522": 0.52173912525177
}
},
"notes": "no compile, no fsdp, single GPU; DmdDenoisingStage's INTERNAL scheduler (hardcoded shift 8.0) is the sigma authority; committed goldens trimmed: e2e_step0 dropped (input is the seeded draw, reproducible) and per-step outputs dropped (triage-only) \u2014 full set regenerable via capture_fastvideo_main.py"
}
@@ -0,0 +1,51 @@
{
"repo": "FastVideo/FastWan2.1-T2V-1.3B-Diffusers",
"snapshot": "25e7ed7f41fd8ce2fdd108688c65e8caf0ce3aef",
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"fastvideo_src": "/mnt/FastVideo",
"torch": "2.12.0+cu130",
"flash_attn": "2.8.3",
"python": "3.12.13",
"gpu": "NVIDIA GB200",
"attention_backend": "VIDEO_SPARSE_ATTN",
"quant": "none (bf16, VSA sparsity 0.80)",
"vsa_sparsity": 0.8,
"seed": 1234,
"e2e_prompt": "A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest. The playful yet serene atmosphere is complemented by soft natural light filtering through the petals. Mid-shot, warm and cheerful tones.",
"probe_prompt": "A cat and a dog baking a cake together in a kitchen.",
"probe_timesteps": [
1000,
757,
522
],
"probe_latent_bcfhw": [
1,
16,
5,
60,
104
],
"e2e_latent_btchw": [
1,
21,
16,
60,
104
],
"dmd_denoising_steps": [
1000,
757,
522
],
"scheduler": {
"class": "FlowMatchEulerDiscreteScheduler",
"shift": 8.0,
"table_len": 1000,
"sigma_lookup": {
"1000": 1.0,
"757": 0.7567567229270935,
"522": 0.52173912525177
}
},
"notes": "no compile, no fsdp, single GPU; DmdDenoisingStage's INTERNAL scheduler (hardcoded shift 8.0) is the sigma authority; committed goldens trimmed: e2e_step0 dropped (input is the seeded draw, reproducible) and per-step outputs dropped (triage-only) \u2014 full set regenerable via capture_fastvideo_main.py"
}
@@ -0,0 +1,282 @@
[
{
"x_hash": "03ceb634cfbdc8a5",
"out_hash": "658efa11e9a88e87",
"t": [
1000.0
],
"start": 0
},
{
"x_hash": "b986baa8885d6cb9",
"out_hash": "dd84656c50a4b92b",
"t": [
937.5
],
"start": 0
},
{
"x_hash": "c08abf4de0fc5fc2",
"out_hash": "4001c9596842592d",
"t": [
833.3333129882812
],
"start": 0
},
{
"x_hash": "a27dd2144f6a3c91",
"out_hash": "2a0f48499caa3ca5",
"t": [
625.0
],
"start": 0
},
{
"x_hash": "29769211170d71a2",
"out_hash": "a4fe211f3fda4514",
"t": [
0.0
],
"start": 0
},
{
"x_hash": "d96a639192b5aa24",
"out_hash": "4b1f9dd86d718170",
"t": [
1000.0
],
"start": 4680
},
{
"x_hash": "13c644ac8937ac8e",
"out_hash": "d804417cc5fcc794",
"t": [
937.5
],
"start": 4680
},
{
"x_hash": "2a9fd4b8c4362a57",
"out_hash": "3f996d20f583eb87",
"t": [
833.3333129882812
],
"start": 4680
},
{
"x_hash": "c0bc578106311f69",
"out_hash": "1a2bf4336c45ebc8",
"t": [
625.0
],
"start": 4680
},
{
"x_hash": "cd0c54679935f14d",
"out_hash": "3fcda958d74bf79a",
"t": [
0.0
],
"start": 4680
},
{
"x_hash": "cd20f035b46125c5",
"out_hash": "9a976bbb313fc9b9",
"t": [
1000.0
],
"start": 9360
},
{
"x_hash": "3332d96e01ae4004",
"out_hash": "ddb5cde97146b967",
"t": [
937.5
],
"start": 9360
},
{
"x_hash": "1f02a448f8c27f80",
"out_hash": "86a4ac2b3e6624f7",
"t": [
833.3333129882812
],
"start": 9360
},
{
"x_hash": "c34398a75af94118",
"out_hash": "e19727744ee978ee",
"t": [
625.0
],
"start": 9360
},
{
"x_hash": "b8229bafed68f655",
"out_hash": "8a07b4778c138a91",
"t": [
0.0
],
"start": 9360
},
{
"x_hash": "570ff151d9d1444e",
"out_hash": "1a66ace7a1d60b51",
"t": [
1000.0
],
"start": 14040
},
{
"x_hash": "61bd37b3e6382f8b",
"out_hash": "6705d1c5591a147b",
"t": [
937.5
],
"start": 14040
},
{
"x_hash": "5aaccd0e0c06d59c",
"out_hash": "e43adfb21fa1bc9b",
"t": [
833.3333129882812
],
"start": 14040
},
{
"x_hash": "65de10d92f0d12bd",
"out_hash": "800293d62b41f4ae",
"t": [
625.0
],
"start": 14040
},
{
"x_hash": "637bc60ba3c0b7a1",
"out_hash": "63fb4744b7d48b10",
"t": [
0.0
],
"start": 14040
},
{
"x_hash": "ca0cccb30274219c",
"out_hash": "2fe6ebe902961d57",
"t": [
1000.0
],
"start": 18720
},
{
"x_hash": "d4e78ae4fa1bbc1d",
"out_hash": "24694df9ba77685c",
"t": [
937.5
],
"start": 18720
},
{
"x_hash": "2b8ddf173f274fc8",
"out_hash": "1156a77610e6dc1b",
"t": [
833.3333129882812
],
"start": 18720
},
{
"x_hash": "6bbd2511192162cc",
"out_hash": "dcb87decbf4d126f",
"t": [
625.0
],
"start": 18720
},
{
"x_hash": "b144c64592c830fd",
"out_hash": "09b00b37c56dc714",
"t": [
0.0
],
"start": 18720
},
{
"x_hash": "3f11cee5b1830f29",
"out_hash": "ac56b4506da4e278",
"t": [
1000.0
],
"start": 23400
},
{
"x_hash": "a7bc9435cd702102",
"out_hash": "c238bb0f2fa1b957",
"t": [
937.5
],
"start": 23400
},
{
"x_hash": "6d9a5dcc9af79188",
"out_hash": "1b7759348b2e0a06",
"t": [
833.3333129882812
],
"start": 23400
},
{
"x_hash": "42fd32612cfc36b8",
"out_hash": "d12e11bb03d918bb",
"t": [
625.0
],
"start": 23400
},
{
"x_hash": "a7583660de097f48",
"out_hash": "90473cb353fbb62d",
"t": [
0.0
],
"start": 23400
},
{
"x_hash": "e5c85230c382e2f4",
"out_hash": "6ed5d13fa54d6f85",
"t": [
1000.0
],
"start": 28080
},
{
"x_hash": "f60119f243f5bbb2",
"out_hash": "52741bbb32a4fa39",
"t": [
937.5
],
"start": 28080
},
{
"x_hash": "c12f401fafe1e333",
"out_hash": "a574f69e6525862d",
"t": [
833.3333129882812
],
"start": 28080
},
{
"x_hash": "f68f10bfa7c33eae",
"out_hash": "1e2aa2baa76fdfbb",
"t": [
625.0
],
"start": 28080
},
{
"x_hash": "503b0c92faec0367",
"out_hash": "21ea40e0d7d642e3",
"t": [
0.0
],
"start": 28080
}
]
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,23 @@
{
"repo": "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers",
"snapshot": "4b44356635ae5e927ca552a220f768022be76004",
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"torch": "2.12.0+cu130",
"python": "3.12.13",
"gpu": "NVIDIA GB200",
"attention_backend": "FLASH_ATTN",
"seed": 1234,
"e2e_prompt": "A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest. The playful yet serene atmosphere is complemented by soft natural light filtering through the petals. Mid-shot, warm and cheerful tones.",
"probe_prompt": "A cat and a dog baking a cake together in a kitchen.",
"dmd_denoising_steps": [
1000,
750,
500,
250
],
"warp_denoising_step": true,
"scheduler": "SelfForcingFlowMatchScheduler(shift=5, extra_one_step, sigma_min=0)",
"num_frames_per_block": 3,
"context_noise": 0,
"notes": "causal chunk rollout via main's CausalDMDDenosingStage; per-forward hashes for all 35 forwards, full tensors for the first two chunks; no compile/fsdp; FLASH_ATTN"
}
@@ -0,0 +1,71 @@
{
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"mode": "dmd2",
"seed": 42,
"gen_losses": [
0.0,
0.33539342880249023,
0.0,
0.31753748655319214,
0.0
],
"fake_losses": [
0.003410761244595051,
0.00596819119527936,
0.013151273131370544,
0.022480690851807594,
0.0037461533211171627
],
"torch": "2.12.0+cu130",
"gpu": "NVIDIA GB200",
"config": "dmd2 legacy: interval2 gw3.5 lr2e-6 shift8 steps[1000,757,522] simulate nlt4 1gpu",
"self_noise_runs": {
"gen": [
[
0.0,
0.3357574939727783,
0.0,
0.3176053762435913,
0.0
],
[
0.0,
0.33570683002471924,
0.0,
0.3168887794017792,
0.0
],
[
0.0,
0.33539342880249023,
0.0,
0.31753748655319214,
0.0
]
],
"fake": [
[
0.003410761244595051,
0.0059582265093922615,
0.013146793469786644,
0.022459683939814568,
0.0037503200583159924
],
[
0.003410761244595051,
0.005963137373328209,
0.013137918896973133,
0.022351054474711418,
0.0037392042577266693
],
[
0.003410761244595051,
0.00596819119527936,
0.013151273131370544,
0.022480690851807594,
0.0037461533211171627
]
]
},
"self_noise_max": 0.0007165968418121338
}
@@ -0,0 +1,54 @@
[
{
"targets": [
2
],
"dmd_t": null,
"critic_t": 737.588623046875,
"gen_loss": 0.0,
"fake_loss": 0.003410761244595051,
"x0_student_hash": "96f01be58584097b"
},
{
"targets": [
1,
0
],
"dmd_t": 386.4990234375,
"critic_t": 716.4179077148438,
"gen_loss": 0.33539342880249023,
"fake_loss": 0.00596819119527936,
"x0_student_hash": "34f51407f1b630e5"
},
{
"targets": [
2
],
"dmd_t": null,
"critic_t": 967.0870971679688,
"gen_loss": 0.0,
"fake_loss": 0.013151273131370544,
"x0_student_hash": "2d65be80d333fd5d"
},
{
"targets": [
0,
1
],
"dmd_t": 662.4631958007812,
"critic_t": 980.0,
"gen_loss": 0.31753748655319214,
"fake_loss": 0.022480690851807594,
"x0_student_hash": "822e22773ec52d1e"
},
{
"targets": [
2
],
"dmd_t": null,
"critic_t": 456.45648193359375,
"gen_loss": 0.0,
"fake_loss": 0.0037461533211171627,
"x0_student_hash": "472a8020d513b4ea"
}
]
@@ -0,0 +1,41 @@
{
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"dataset": "wlsaidhi/crush-smol_processed_t2v",
"seed": 42,
"steps": 5,
"num_latent_t": 8,
"config": "1gpu bs1 accum1 lr5e-5 wd1e-4 betas(0.9,0.999) clip1.0 uniform-t cfg_rate0 dit_fp32 mixed_bf16 flow-match target=noise-latents flash_attn",
"torch": "2.12.0+cu130",
"gpu": "NVIDIA GB200",
"losses": [
0.1940414160490036,
0.94484943151474,
0.1001732274889946,
0.9069435000419617,
0.10525074601173401
],
"self_noise_runs": [
[
0.1940414160490036,
0.9445295333862305,
0.10048552602529526,
0.9095064997673035,
0.10758557915687561
],
[
0.1940414160490036,
0.9446913599967957,
0.10070198774337769,
0.9106038808822632,
0.10818696022033691
],
[
0.1940414160490036,
0.94484943151474,
0.1001732274889946,
0.9069435000419617,
0.10525074601173401
]
],
"self_noise_max": 0.0036603808403015137
}
@@ -0,0 +1,82 @@
[
{
"caption": "A large metal cylinder is seen pressing down on a pile of colorful candies, flattening them as if they were under a hydraulic press. The candies are crushed and broken into small pieces, creating a mess on the table.",
"latents_hash": "93ae01bc47c0817b",
"embeds_hash": "53f384dfed477e28",
"noise_hash": "e1f135db4ff97ff8",
"noisy_hash": "6a0cabf3c2e98794",
"timesteps": [
118.0
],
"sigmas": [
0.1181640625
],
"pred_hash": "2903c217cc9279df",
"loss": 0.1940414160490036,
"grad_norm": 0.23514027893543243
},
{
"caption": "The video shows a colorful sponge being flattened as if it were under a hydraulic press, with the sponge being compressed and eventually flattened into a thin layer.",
"latents_hash": "b9d5abc838e815da",
"embeds_hash": "82c3f45fec89f678",
"noise_hash": "ea2e0427b4584180",
"noisy_hash": "5ff0bac18dd8e80e",
"timesteps": [
85.0
],
"sigmas": [
0.0849609375
],
"pred_hash": "688def082dcbfa60",
"loss": 0.94484943151474,
"grad_norm": 10.1902437210083
},
{
"caption": "The video shows a hydraulic press in action, flattening objects as if they were under a hydraulic press. The press is composed of a large, cylindrical metal cylinder with yellow and black stripes, and a metal base. The objects being flattened are two cylindrical blocks of cotton candy, one pink and one blue. The press is positioned on a metal table, and the background features a green wall with a yellow and red sign.",
"latents_hash": "840ddc5d78da4b7b",
"embeds_hash": "ffce0c1bf664fb3d",
"noise_hash": "2f512661cebc8a51",
"noisy_hash": "05b1746f31f8d83d",
"timesteps": [
618.0
],
"sigmas": [
0.6171875
],
"pred_hash": "963b9731aafbe2ea",
"loss": 0.1001732274889946,
"grad_norm": 0.523212730884552
},
{
"caption": "A watermelon wearing a helmet is crushed by a hydraulic press, causing it to flatten and burst open.",
"latents_hash": "c5d8005785df08ab",
"embeds_hash": "88eea058d58c26f7",
"noise_hash": "1022b2e68b615e5f",
"noisy_hash": "ee1750fb9fdb7094",
"timesteps": [
41.0
],
"sigmas": [
0.041015625
],
"pred_hash": "5fd1a77f3716b63b",
"loss": 0.9069435000419617,
"grad_norm": 3.6921098232269287
},
{
"caption": "The video shows a large, industrial press flattening objects as if they were under a hydraulic press. The press is shown in action, compressing a pile of pink objects into a pile of crumbs. The press is large and metallic, with a yellow and black striped pattern on its side. The background is a green wall with a yellow warning sign.",
"latents_hash": "57045eb31ae1e81a",
"embeds_hash": "71cee36d5ca3083d",
"noise_hash": "cdc1cbc4a1fed9ac",
"noisy_hash": "53ada1e537af47ac",
"timesteps": [
610.0
],
"sigmas": [
0.609375
],
"pred_hash": "315edc5acc47767a",
"loss": 0.10525074601173401,
"grad_norm": 0.9913093447685242
}
]
@@ -0,0 +1,22 @@
{
"fastvideo_commit": "c459a1897899ffcec3be7765534d81000b9bb9c1",
"mode": "qad",
"seed": 42,
"gen_losses": [
0.0,
0.1784808486700058,
0.0,
0.17068849503993988,
0.0
],
"fake_losses": [
0.0031898675952106714,
0.005877670831978321,
0.015015869401395321,
0.023572081699967384,
0.0025676116347312927
],
"torch": "2.12.0+cu130",
"gpu": "NVIDIA GB200",
"config": "dmd2 legacy: interval2 gw3.5 lr2e-6 shift8 steps[1000,757,522] simulate nlt4 1gpu"
}
Binary file not shown.
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More