Author SHA1 Message Date
bubbliiiing 8c6428593b Merge remote-tracking branch 'origin/main' into gpu_memory_mode 2026-09-29 19:12:29 +08:00
bubbliiiing 66591d338d delete useless import 2026-09-27 16:11:17 +08:00
bubbliiiing 8f76c66d67 Update qwenimage21 control yaml 2026-09-27 15:38:20 +08:00
bubbliiiing f5338a4aa7 Update comments 2026-09-27 15:27:00 +08:00
bubbliiiing 28e60b1182 Update gpu memory mode and fix bug in fsdp between cpu offload in minimax h3 2026-09-27 15:06:06 +08:00
dellenovo f3d5491459 Fix a typo in xFuserLongContextAttention (#520) 2026-09-27 09:13:59 +08:00
Bubbliiiing 67e3b4b818 Update Qwen Image 2.1 Control and Flex Forcing (#519) 2026-09-24 15:52:16 +08:00
Bubbliiiing bce42b301f Update Readme (#517) 2026-09-22 21:06:38 +08:00
Bubbliiiing 2beb171099 Support Qwen-Image2.1 && Support new Control for Minimax-H3 && Support Wan2.2 with 2.2VAE && Update new tqdm for All Training Code (#511) 2026-09-22 20:41:39 +08:00
hkz 968f0e2192 Add PDD for MiniMax-H3 (#515) 2026-09-04 17:02:21 +08:00
谢翊凡and谢翊凡 43739895a1 fix: use dtype instead of deprecated torch_dtype (#514)
* fix: use dtype instead of deprecated torch_dtype for transformers >= 4.56

config.torch_dtype and the torch_dtype keyword argument were deprecated in transformers 4.56 (PR #39782). Use dtype when accepted and fall back to torch_dtype via try/except TypeError for older versions.

* fix: use dtype instead of deprecated torch_dtype for transformers >= 4.56

config.torch_dtype and the torch_dtype keyword argument were deprecated in transformers 4.56 (PR #39782). Pass dtype based on the installed transformers version (packaging.version), falling back to torch_dtype on older versions.

---------

Co-authored-by: 谢翊凡 <xyf5432@users.noreply.github.com>
2026-09-02 16:47:39 +08:00
Bubbliiiing 6f3fb60dad Update scale fp8 with fsdp (#508) 2026-08-25 12:35:36 +08:00
Bubbliiiing b0acf916c2 Update Lingbot Worl, Lingbot Video, Minimax-H3 and Minimax-H3 Control (#506) 2026-08-24 16:47:38 +08:00
hkz 6787dc8ed4 Add DFD for Wan2.1 and Wan2.2 (#505) 2026-08-13 16:56:53 +08:00
hkz 248ab0ac0e Add Causal-Focing for Wan2.1 (#500) 2026-07-22 17:37:05 +08:00
Bubbliiiing 403f1f7b78 Fix Bug in Self Forcing in Multi-Gpus Infernece && Update Training Code and Docs && Update Reward Models (#498) 2026-07-14 10:30:54 +08:00
Bubbliiiing 1fd9ed9208 Ode training && Update Lens model && Update LTX2 upsampler (#497) 2026-06-09 15:20:04 +08:00
Bubbliiiing 2b5596b8e6 Update Self-Forcing and Ernie Image (#490) 2026-05-25 17:39:17 +08:00
Bubbliiiing 804a4258e2 Update Flash Head && Update Readmes (#491) 2026-05-12 16:48:47 +08:00
Tyler Luan 8fb0bb165c Align with the official FlashHead implementation (#488)
- Update predict_s2v.py with improved inference configuration
- Align pipeline_flashhead.py with official streaming, scheduling, conditioning, and motion-cache behavior
2026-05-08 14:04:09 +08:00
Bubbliiiing 0266bab98b Update Model Group Offload and Readmes (#487) 2026-05-06 10:48:08 +08:00
Bubbliiiing 199a43b544 Update READMEs, add LongCatVideo multi-GPU inference, and fix import bugs (#485) 2026-04-24 15:54:22 +08:00
Bubbliiiing 34036517a8 Added LTX-2.3 support, optimized memory usage by casting e and e0 types in Wan-based models, and updated READMEs for image training and digital human models. (#483) 2026-04-16 20:45:16 +08:00
Bubbliiiing 54bf97ab66 Fix bug in FA3 and LTX2 && Update FA4 support && Update Infinitalk && Update MOVA && Update FlashHead (#479) 2026-04-07 14:08:39 +08:00
Bubbliiiing 4a86483cc2 Reformat S2V models && Update LTX-2 (#476) 2026-03-20 10:45:57 +08:00
Bubbliiiing ad72867c0f Fix Bug in Group Offload when diffusers version is low. (#474) 2026-03-10 15:02:16 +08:00
Bubbliiiing 5202421e7c Update Z Image Turbo Control 2602 (#460) 2026-03-04 16:17:04 +08:00
Sense_wangandhaosenwang1018 745acc1f47 Fix: Replace 74 bare excepts with except Exception (#458)
Co-authored-by: haosenwang1018 <haosenwang1018@users.noreply.github.com>
2026-02-26 16:24:51 +08:00
Bubbliiiing a4b60f40bf Update Z Image Control Tile (#457) 2026-02-25 11:12:57 +08:00
Bubbliiiing 7a085c9535 Update Longcat Avatar (#453) 2026-02-13 10:14:14 +08:00
Bubbliiiing 1ca4162447 Support validations in all models && Update Saving Code && Preparing Code (#452) 2026-02-10 17:23:23 +08:00
Bubbliiiing ee44fbc950 Lora Loading Log and Z Image Distill (#450) 2026-02-04 16:55:24 +08:00
Bubbliiiing a6b026526f Update Qwen Image Layered, Turbowan Training Code and Z Image predict Code (#448) 2026-02-03 18:23:45 +08:00
Bubbliiiing 0f0e2bd5ab Update Flux2 Control Cfg Distill && Fix Bug in Lora Training Register Hook (#445) 2026-02-03 10:23:10 +08:00
700 changed files with 267914 additions and 18394 deletions
+12 -10
View File
@@ -1,17 +1,19 @@
# Byte-compiled / optimized / DLL files
models*
output*
logs*
taming*
samples*
datasets*
asset*
# Used in VideoX-Fun
_*
logs*
/models*
/output*
/logs*
/taming*
/samples*
/datasets*
/asset*
/repo*
/scripts_demo*
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[cod]
*$py.class
scripts_demo*
# C extensions
*.so
+128
View File
@@ -0,0 +1,128 @@
---
name: integrating-models
description: Guides adding, porting, or onboarding a diffusion model (transformer/VAE/encoder, inference pipeline, training script, config) into the VideoX-Fun repository by mirroring the closest existing model family and maximizing reuse of the repository's existing code and shared infrastructure. Use when integrating a new model/architecture, or when creating predict_*.py inference scripts, scripts/*/train*.py training scripts, pipeline_*.py, config/*.yaml, or model definitions under videox_fun/models/.
---
# Integrating Models into VideoX-Fun
## Core rule: maximize reuse of existing repo code — mirror, extend, never reinvent
**Prime directive: reuse this repository's existing code to the maximum.** Nearly every building block you need already exists in `videox_fun/` or in a sibling model family. Your job is to **find it, import it, and extend it** — not to write a parallel implementation. A new file should be mostly reused structure plus the genuinely model-specific delta; the less new code you write, the better.
**Reuse-first protocol — before writing ANY new function / class / util:**
1. **Search the repo first.** Grep `videox_fun/` and the closest family for an existing equivalent (weight loader, scheduler, sampler, offload, attention, LoRA, fp8, dataset, dist helper, save/metric util). If one exists → **import and reuse it**. If it is 80% right → **extend / parameterize it**, do not fork it.
2. **Only if nothing exists** may you add new code — and then put it in the shared layer (`videox_fun/utils`, `videox_fun/data`, `videox_fun/dist`) so the next model reuses it too, instead of burying it in a family folder.
3. **Never copy-paste** a util into a new file (that creates drift); import the single source of truth.
**Mirror the closest family.** Every model follows the **same layered template**. Integrating a model means finding the closest existing family and mirroring its structure, changing only what genuinely differs:
1. Pick the closest existing family by task type (t2v / i2v / v2v-control / s2v / t2i / edit / distill): `wan2.1`, `wan2.1_fun`, `wan2.2`, `qwenimage`, `flux2`, `minimax_h3`, `ltx2`, `longcatvideo`, `cogvideox_fun`, `z_image`, etc.
2. Read that family end-to-end across all layers:
- `examples/<family>/predict_*.py` (inference entry)
- `scripts/<family>/train*.py` + `*.sh` + `README_TRAIN*.md` (training)
- `videox_fun/pipeline/pipeline_<family>*.py` (pipeline)
- `videox_fun/models/<family>_*.py` (model definitions)
- `config/<family>/*.yaml` (config)
3. Copy that structure and adapt. Keep names, argument sets, control flow, and reuse points identical in shape.
Writing a bespoke pipeline, weight loader, trainer, sampler, dataset, or offload scheme from scratch is a **failure mode**. If you are tempted to, **stop** and check the Reuse inventory below first.
## Repository layout (where each layer lives)
| Layer | Location | What it is |
|-------|----------|------------|
| Model definitions | `videox_fun/models/<family>_*.py` | Transformer / VAE / text-audio-image encoders. Diffusers `ModelMixin`+`ConfigMixin`, `@register_to_config`, custom `from_pretrained`. |
| Model registry | `videox_fun/models/__init__.py` | Imports every model class. **Must be updated** for a new model. |
| Inference pipelines | `videox_fun/pipeline/pipeline_<family>*.py` | `<Family>Pipeline(DiffusionPipeline)` with `__call__`. |
| Pipeline registry | `videox_fun/pipeline/__init__.py` | Imports every pipeline + aliases. **Must be updated.** |
| Configs (optional) | `config/<family>/*.yaml` | OmegaConf YAML for civitai/custom layouts; a standard diffusers-layout checkpoint can load without one. |
| Inference entry scripts | `examples/<family>/predict_*.py` | User-facing, config-block-at-top runnable scripts. |
| Inference services | `examples/<family>/{app.py,launch_api.py,post_infer*.py}` | Gradio UI / API server / batch inference. |
| Training scripts | `scripts/<family>/train*.py` | `train.py`, `train_lora.py`, `train_control.py`, `train_distill.py`, ... |
| Training launchers | `scripts/<family>/train*.sh` | `accelerate launch` / DeepSpeed command with full arg list. |
| Training docs | `scripts/<family>/README_TRAIN*.md` | Bilingual pairs: `README_TRAIN.md` + `README_TRAIN_zh-CN.md`. |
| Shared: schedulers/utils | `videox_fun/utils/` | `fm_solvers`, `fm_solvers_unipc`, `lora_utils`, `fp8_optimization`, `group_offload`, `utils.py`. |
| Shared: distributed | `videox_fun/dist/` | `fsdp.shard_model`, `fuser.set_multi_gpus_devices`, `<family>_xfuser` sequence-parallel attention. |
| Shared: data | `videox_fun/data/` | Datasets (`ImageVideoDataset`, `VideoDataset`, ...) + bucket/aspect-ratio samplers. |
| Demo / test datasets | `datasets/X-Fun-*-Demo/` | Ready-made smoke-test data, downloaded via `modelscope download --dataset PAI/<name>`; each ships several `metadata*.json` variants. **The only test data to use** (see reference.md §8). |
| Preprocessing (data gen) | `scripts/<family>/generate_*.py` / `train_preprocess.py` (+ `.sh`) | Offline multi-GPU generation of cached training data (latents / ODE pairs / embeddings) → per-sample `.safetensors` + `outputs.json`, loaded by `ImageVideoSafetensorsDataset`. |
| ComfyUI nodes | `comfyui/<family>/nodes.py` | Optional node integration mirroring the pipeline. |
## Integration workflow
Copy this checklist and track progress:
```
Integration Progress:
- [ ] Step 0: Choose the closest family to mirror; read it across all layers
- [ ] Step 1: Model definitions in videox_fun/models/ + register in models/__init__.py
- [ ] Step 2: Pipeline in videox_fun/pipeline/ + register in pipeline/__init__.py
- [ ] Step 3: Config YAML in config/<family>/
- [ ] Step 4: Inference script(s) in examples/<family>/predict_*.py
- [ ] Step 5: Training script(s) in scripts/<family>/train*.py + .sh
- [ ] Step 6: Training docs README_TRAIN.md + README_TRAIN_zh-CN.md
- [ ] Step 7: Reuse audit + verification (incl. smoke test on the matching demo dataset)
```
**Step 0 — Choose the mirror.** Match by task and architecture. A new control model mirrors an existing `*_fun`/`*_control` family; a new audio/talking model mirrors `minimax_h3`/`longcatvideo`/`infinitetalk`; a new image model mirrors `qwenimage`/`flux2`/`z_image`.
**Step 1 — Model.** Create `videox_fun/models/<family>_transformer3d.py` (or `2d`), `<family>_vae.py`, encoders as needed. Mirror the class shape: `class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin)`, `_supports_gradient_checkpointing = True`, `@register_to_config __init__`, and a `from_pretrained` that supports `transformer_additional_kwargs`, `dict_mapping`, `low_cpu_mem_usage`, and missing-key init. Add imports to `videox_fun/models/__init__.py`.
**Step 2 — Pipeline.** Create `videox_fun/pipeline/pipeline_<family>.py`. Mirror `pipeline_wan.py`: module-level `retrieve_timesteps`, a `<Family>PipelineOutput(BaseOutput)` dataclass, `<Family>Pipeline(DiffusionPipeline)` with `model_cpu_offload_seq`, `_callback_tensor_inputs`, `__init__(vae, tokenizer, text_encoder, transformer, scheduler, ...)`, `encode_prompt`, and `__call__`. Add imports/aliases to `videox_fun/pipeline/__init__.py`.
**Step 3 — Config (optional).** A YAML under `config/<family>/` is **not always required**. It is needed mainly for **civitai-format / custom single-file layouts** — to supply `transformer_additional_kwargs`, `dict_mapping` (civitai key → `__init__` kwarg), component subpaths, and `vae_kwargs`/`text_encoder_kwargs`/`scheduler_kwargs`/`image_encoder_kwargs`. For a **standard diffusers-layout** checkpoint (`model_index.json` + per-subfolder `config.json`), load directly via `from_pretrained(model_name, subfolder=...)` with no YAML — mirror `examples/minimax_h3_fun/predict_v2v_control.py`, which guards `if config_path is not None:`. When you do add a YAML, load it via `OmegaConf.load(config_path)` and spread into `from_pretrained` instead of hardcoding those values.
**Step 4 — Inference script.** Create `examples/<family>/predict_<task>.py` following the exact template (config block at top → component loading → scheduler dict → pipeline construction → multi-GPU/FSDP/compile → `GPU_memory_mode` branching → TeaCache → LoRA merge → inference → `save_results`). See [examples.md](examples.md).
**Step 5 — Training script.** Create `scripts/<family>/train.py` (+ `train_lora.py` etc.). Mirror the shared structure: license header, `sys.path` bootstrap, imports from `videox_fun`, `log_validation()` that **reuses the inference Pipeline**, `parse_args()` (reuse the existing shared argument set), `main()`. Add a `train.sh` launcher. Reuse `videox_fun.data` datasets/samplers — do not write a new dataset.
**Step 6 — Docs.** Write `README_TRAIN.md` and `README_TRAIN_zh-CN.md` as an aligned bilingual pair (same structure, same commands/params, matching section order).
**Step 7 — Reuse audit + verification.** Confirm you reused shared infra (below), smoke-test the new train/predict path on the **matching official demo dataset** under `datasets/X-Fun-*-Demo/` (pick by task and metadata variant — see reference.md §8), then run the verification checklist. Never invent an ad-hoc test set and never leave `datasets/internal_datasets/` placeholders in shipped scripts/docs.
## Reuse inventory (use these, do not reimplement)
**Reuse-first catalog: import from here instead of reimplementing. If a helper you need is not listed, grep `videox_fun/` and the closest family before writing your own.**
- **Schedulers**: `FlowMatchEulerDiscreteScheduler`, `videox_fun.utils.fm_solvers.FlowDPMSolverMultistepScheduler`, `fm_solvers_unipc.FlowUniPCMultistepScheduler`. Selected via a `sampler_name` dict.
- **LoRA**: `videox_fun.utils.lora_utils` — `merge_lora`, `unmerge_lora`, `create_network`, `convert_peft_lora_to_kohya_lora`.
- **FP8 / quantization**: `videox_fun.utils.fp8_optimization` — `convert_model_weight_to_float8`, `convert_weight_dtype_wrapper`, `replace_parameters_by_name`.
- **Offloading**: `videox_fun.utils.group_offload` — `register_auto_device_hook`, `safe_enable_group_offload`; plus pipeline `enable_sequential_cpu_offload` / `enable_model_cpu_offload` / `.to(device)`.
- **Distributed**: `videox_fun.dist` — `set_multi_gpus_devices`, `shard_model` (FSDP), `<family>_xfuser` sequence-parallel attention processors, `enable_multi_gpus_inference()`.
- **IO / helpers**: `videox_fun.utils.utils` — `save_videos_grid`, `save_videos_with_audio_grid`, `get_image_to_video_latent`, `get_video_to_video_latent`, `get_image_latent`, `filter_kwargs`, `calculate_dimensions`.
- **Data**: `videox_fun.data` — `ImageVideoDataset`, `VideoDataset`, `ImageVideoControlDataset`, `VideoSpeechDataset`, bucket/aspect-ratio samplers, `get_closest_ratio`, `get_random_mask`.
- **Caching / speedups**: TeaCache (`models/cache_utils`, `get_teacache_coefficients`, `transformer.enable_teacache`), `enable_cfg_skip`, Riflex (`enable_riflex`), `torch.compile` on `transformer.blocks`.
- **Preprocessing (data gen, multi-GPU)**: mirror `scripts/wan2.1_self_forcing/generate_ode_pairs.py` — `accelerate launch` + `Accelerator` (interleaved rank sharding), config-driven `from_pretrained` for the teacher/VAE/text-encoder, `safetensors.torch.save_file` per sample + `outputs.json` index, consumed by `videox_fun.data.ImageVideoSafetensorsDataset`. Store as **safetensors only — never LMDB or `.pt`** (see reference.md §10).
## Non-negotiable conventions
- **Maximize reuse of existing repo code**: import existing `videox_fun/` helpers and mirror the closest family; never fork or copy-paste a util, and never write a parallel pipeline / loader / scheduler / sampler / offload. Genuinely-new shared code goes in `videox_fun/{utils,data,dist}` (so the next model reuses it), not buried in a family folder.
- **`sys.path` bootstrap**: every runnable script starts with the 3-level `project_roots` loop inserting into `sys.path` before importing `videox_fun`.
- **Config-driven loading (YAML optional)**: a `config/<family>/*.yaml` is required for civitai-format/custom layouts (it supplies `transformer_additional_kwargs`/`dict_mapping`/subpaths); it is **optional for standard diffusers-layout checkpoints**, which load directly via `from_pretrained(model_name, subfolder=...)`. When a YAML is used, don't hardcode the values it provides.
- **`GPU_memory_mode`**: support the standard six modes — `model_full_load`, `model_full_load_and_qfloat8`, `model_cpu_offload`, `model_cpu_offload_and_qfloat8`, `model_group_offload`, `sequential_cpu_offload` — with the exact branching order used in existing `predict_*.py`.
- **Naming**: files `<family>_transformer3d.py` / `<family>_vae.py` / `pipeline_<family>.py`; classes `<Family>Transformer3DModel` / `AutoencoderKL<Family>` / `<Family>Pipeline`.
- **Resolution args**: drive canvas size with a single square `--video_sample_size` (`type=int`, height = width); never `--video_sample_height` / `--video_sample_width`. For a fixed non-square shape add `--fix_sample_size` (`nargs=2, type=int`, `[height, width]`) that overrides the square size, and derive the effective height/width once in `parse_args()` (see reference.md §5).
- **Registries**: a model is not integrated until it is imported in BOTH `videox_fun/models/__init__.py` and `videox_fun/pipeline/__init__.py`.
- **Two weight formats**: support `civitai` and `diffusers` via config `format` + `dict_mapping` (maps civitai keys such as `in_dim`→`in_channels`, `dim`→`hidden_size`).
- **Bilingual docs**: training READMEs ship as EN + `_zh-CN` pairs with aligned structure and identical commands/params.
- **Test data = official demo datasets**: smoke tests, `log_validation` checks, launcher `.sh` defaults, and doc examples all point at `datasets/X-Fun-*-Demo/` (ModelScope `PAI/<name>`), with the metadata variant matching the task — `metadata_add_width_height.json` by default, `_add_objects.json` for VACE/subject-reference, `_add_wav.json` for audio-visual joint models, `metadata_lingbot_video_add_width_height.json` for `lingbot_video`. Selection matrix: reference.md §8.
- **Preprocessing = offline data generation, multi-GPU + safetensors**: cached training data (latents / ODE pairs / embeddings) is produced by `accelerate launch` scripts like `generate_ode_pairs.py` (interleaved rank sharding, resume by skipping existing files, `wait_for_everyone`, rank-0 JSON index) and saved with `safetensors.torch.save_file` + an `outputs.json` index for `ImageVideoSafetensorsDataset`. **Never single-GPU / `cuda:0`; never LMDB or `.pt`/`torch.save` pickles for preprocessed data.** See reference.md §10.
## Verification checklist
- [ ] New model classes imported in `videox_fun/models/__init__.py`
- [ ] New pipeline(s) imported in `videox_fun/pipeline/__init__.py`
- [ ] Config YAML present **only if** the checkpoint is civitai-format/custom-layout; a diffusers-layout model may load directly via `from_pretrained(model_name, subfolder=...)` with no YAML. When a YAML is used, it drives component loading (no hardcoded kwargs)
- [ ] `predict_*.py` mirrors an existing script: `sys.path` bootstrap, config block, scheduler dict, `GPU_memory_mode` branching, LoRA merge, `save_results`
- [ ] `train*.py` reuses `videox_fun.data` + shared args, and `log_validation()` reuses the inference Pipeline
- [ ] `train*.sh` launcher provided (`accelerate launch` / DeepSpeed)
- [ ] Shared infra reused (schedulers / lora_utils / fp8 / group_offload / dist / utils / data) — nothing reimplemented
- [ ] Any offline data-generation/preprocessing script runs multi-GPU (`accelerate launch` + `Accelerator`) and saves cached tensors as **safetensors + `outputs.json`** for `ImageVideoSafetensorsDataset` — never LMDB or `.pt`
- [ ] `README_TRAIN.md` + `README_TRAIN_zh-CN.md` aligned pair present
- [ ] Smoke test / doc examples use the matching `datasets/X-Fun-*-Demo` dataset and the correct `metadata*.json` variant — no `internal_datasets` placeholders (reference.md §8)
- [ ] Optional: ComfyUI node in `comfyui/<family>/nodes.py` mirrors the pipeline
## Additional resources
- Detailed file-by-file conventions, class/method shapes, and the model-loading internals: [reference.md](reference.md)
- **Dataset & sampler selection matrix** (which `videox_fun.data` dataset/loader each training task uses), **demo-dataset / metadata-variant selection matrix** (which `datasets/X-Fun-*-Demo` to smoke-test with), **inference task matrix** (which pipeline each `predict_<task>.py` uses), and **multi-GPU preprocessing patterns**: [reference.md](reference.md) §8–§10
- Concrete skeletons (config YAML, `predict_*.py`, pipeline class, training script + DataLoader): [examples.md](examples.md)
+510
View File
@@ -0,0 +1,510 @@
# VideoX-Fun Integration Skeletons
Starting templates. **Always open the mirrored family's real file and adapt it** — these skeletons show shape and required reuse points, not full implementations. Replace `<family>` / `<Family>` / `<task>`.
## Config — `config/<family>/<variant>.yaml` (optional)
> **Not always required.** Author a YAML only for civitai-format / custom single-file layouts. A standard diffusers-layout checkpoint (`model_index.json` + per-subfolder `config.json`) loads directly via `from_pretrained(model_name, subfolder=...)` with no YAML — set `config_path = None` and guard `if config_path is not None:` (see `examples/minimax_h3_fun/predict_v2v_control.py`).
```yaml
format: civitai
pipeline: <Family>
transformer_additional_kwargs:
transformer_subpath: ./
dict_mapping:
in_dim: in_channels
dim: hidden_size
vae_kwargs:
vae_subpath: <Family>_VAE.pth
temporal_compression_ratio: 4
spatial_compression_ratio: 8
text_encoder_kwargs:
text_encoder_subpath: <text_encoder>.pth
tokenizer_subpath: <tokenizer_id>
text_length: 512
scheduler_kwargs:
scheduler_subpath: null
num_train_timesteps: 1000
shift: 5.0
# Only for i2v / models with a CLIP image encoder:
image_encoder_kwargs:
image_encoder_subpath: <image_encoder>.pth
```
## Inference — `examples/<family>/predict_<task>.py`
```python
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
from transformers import AutoTokenizer
# --- sys.path bootstrap (required, before importing videox_fun) ---
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKL<Family>, <Family>TextEncoder,
<Family>Transformer3DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import <Family>Pipeline
from videox_fun.utils import register_auto_device_hook, safe_enable_group_offload
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper,
replace_parameters_by_name)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
save_videos_grid)
# --- user config block (keep the conventional order + comments) ---
GPU_memory_mode = "sequential_cpu_offload"
ulysses_degree = 1
ring_degree = 1
fsdp_dit = False
fsdp_text_encoder = True
compile_dit = False
enable_teacache = True
teacache_threshold = 0.10
num_skip_start_steps = 5
teacache_offload = False
cfg_skip_ratio = 0
enable_riflex = False
riflex_k = 6
config_path = "config/<family>/<variant>.yaml"
model_name = "models/Diffusion_Transformer/<Family>-Model"
sampler_name = "Flow"
shift = 3
transformer_path = None
vae_path = None
lora_path = None
sample_size = [480, 832]
video_length = 81
fps = 16
weight_dtype = torch.bfloat16
prompt = "..."
negative_prompt = "..."
guidance_scale = 6.0
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/<family>-<task>"
# --- device + config (config_path may be None for a diffusers-layout checkpoint) ---
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
config = OmegaConf.load(config_path) # or guard: if config_path is not None: ... (then load components via subfolder=...)
# --- components (when a YAML is used, paths/kwargs come from config; otherwise pass subfolder=... directly) ---
transformer = <Family>Transformer3DModel.from_pretrained(
os.path.join(model_name, config['transformer_additional_kwargs'].get('transformer_subpath', 'transformer')),
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
low_cpu_mem_usage=True, torch_dtype=weight_dtype,
)
# optional transformer_path / vae_path override -> load_state_dict(strict=False) + print missing/unexpected
vae = AutoencoderKL<Family>.from_pretrained(
os.path.join(model_name, config['vae_kwargs'].get('vae_subpath', 'vae')),
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
).to(weight_dtype)
tokenizer = AutoTokenizer.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('tokenizer_subpath', 'tokenizer')))
text_encoder = <Family>TextEncoder.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('text_encoder_subpath', 'text_encoder')),
additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']),
low_cpu_mem_usage=True, torch_dtype=weight_dtype).eval()
# --- scheduler selection dict ---
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler(**filter_kwargs(Chosen_Scheduler, OmegaConf.to_container(config['scheduler_kwargs'])))
# --- pipeline ---
pipeline = <Family>Pipeline(vae=vae, tokenizer=tokenizer, text_encoder=text_encoder,
transformer=transformer, scheduler=scheduler)
# --- multi-gpu / fsdp / compile ---
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
pipeline.transformer = partial(shard_model, device_id=device, param_dtype=weight_dtype)(pipeline.transformer)
if fsdp_text_encoder:
pipeline.text_encoder = partial(shard_model, device_id=device, param_dtype=weight_dtype)(pipeline.text_encoder)
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
# --- GPU_memory_mode branching (keep this exact order) ---
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# --- teacache / cfg_skip / riflex / lora ---
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
if coefficients is not None:
pipeline.transformer.enable_teacache(coefficients, num_inference_steps, teacache_threshold,
num_skip_start_steps=num_skip_start_steps, offload=teacache_offload)
if cfg_skip_ratio is not None:
pipeline.transformer.enable_cfg_skip(cfg_skip_ratio, num_inference_steps)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
# --- inference ---
with torch.no_grad():
video_length = int((video_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio) + 1 if video_length != 1 else 1
if enable_riflex:
pipeline.transformer.enable_riflex(k=riflex_k, L_test=(video_length - 1) // vae.config.temporal_compression_ratio + 1)
sample = pipeline(prompt, num_frames=video_length, negative_prompt=negative_prompt,
height=sample_size[0], width=sample_size[1], generator=generator,
guidance_scale=guidance_scale, num_inference_steps=num_inference_steps,
shift=shift).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
# --- save (rank 0 only when multi-gpu) ---
def save_results():
os.makedirs(save_path, exist_ok=True)
prefix = str(len(os.listdir(save_path)) + 1).zfill(8)
if video_length == 1:
image = (sample[0, :, 0].transpose(0, 1).transpose(1, 2) * 255).numpy().astype(np.uint8)
Image.fromarray(image).save(os.path.join(save_path, prefix + ".png"))
else:
save_videos_grid(sample, os.path.join(save_path, prefix + ".mp4"), fps=fps)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
```
For i2v, gate the CLIP image encoder and pass `video`/`mask_video`:
```python
if transformer.config.in_channels != vae.config.latent_channels:
clip_image_encoder = CLIPModel.from_pretrained(
os.path.join(model_name, config['image_encoder_kwargs'].get('image_encoder_subpath', 'image_encoder'))).to(weight_dtype).eval()
input_video, input_video_mask, _ = get_image_to_video_latent(start_image, None, video_length=video_length, sample_size=sample_size)
# pipeline = <Family>InpaintPipeline(..., clip_image_encoder=clip_image_encoder)
# sample = pipeline(..., video=input_video, mask_video=input_video_mask).videos
```
## Pipeline class — `videox_fun/pipeline/pipeline_<family>.py`
```python
from dataclasses import dataclass
from typing import List, Optional, Union
import torch
from diffusers.callbacks import MultiPipelineCallbacks, PipelineCallback
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
from diffusers.utils import BaseOutput, logging, replace_example_docstring
from diffusers.utils.torch_utils import randn_tensor
from ..models import AutoencoderKL<Family>, <Family>Transformer3DModel
from ..utils.fm_solvers import FlowDPMSolverMultistepScheduler, get_sampling_sigmas
from ..utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
logger = logging.get_logger(__name__)
EXAMPLE_DOC_STRING = """Examples:\n```python\npass\n```"""
# reuse retrieve_timesteps verbatim from pipeline_wan.py
@dataclass
class <Family>PipelineOutput(BaseOutput):
videos: torch.Tensor
class <Family>Pipeline(DiffusionPipeline):
model_cpu_offload_seq = "text_encoder->transformer->vae"
_callback_tensor_inputs = ["latents", "prompt_embeds", "negative_prompt_embeds"]
def __init__(self, tokenizer, text_encoder, vae, transformer, scheduler):
super().__init__()
self.register_modules(tokenizer=tokenizer, text_encoder=text_encoder, vae=vae,
transformer=transformer, scheduler=scheduler)
# video_processor / vae_scale_factor / etc. as in pipeline_wan.py
def encode_prompt(self, prompt, negative_prompt, device, num_videos_per_prompt=1, ...):
... # mirror pipeline_wan.py
def prepare_latents(self, batch_size, num_channels_latents, height, width, num_frames, dtype, device, generator, latents=None):
...
@torch.no_grad()
@replace_example_docstring(EXAMPLE_DOC_STRING)
def __call__(self, prompt, negative_prompt=None, height=480, width=832, num_frames=81,
num_inference_steps=50, guidance_scale=6.0, generator=None, shift=1.0,
callback_on_step_end=None, return_dict=True, **kwargs) -> Union[<Family>PipelineOutput, tuple]:
# 1. encode_prompt 2. prepare_latents 3. retrieve_timesteps
# 4. denoising loop with guidance 5. vae.decode 6. return <Family>PipelineOutput(videos=...)
...
```
Then register in `videox_fun/pipeline/__init__.py`:
```python
from .pipeline_<family> import <Family>Pipeline
```
## Model class — `videox_fun/models/<family>_transformer3d.py`
```python
from diffusers.configuration_utils import ConfigMixin, register_to_config
from diffusers.loaders.single_file_model import FromOriginalModelMixin
from diffusers.models.modeling_utils import ModelMixin
from .attention_utils import attention # unified FA/SDPA backend — do not hand-roll SDPA
class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin):
_supports_gradient_checkpointing = True
@register_to_config
def __init__(self, model_type='t2v', in_dim=16, dim=2048, ffn_dim=8192,
num_heads=16, num_layers=32, in_channels=16, hidden_size=2048, ...):
super().__init__()
...
def _set_gradient_checkpointing(self, *args, **kwargs):
self.gradient_checkpointing = True
def enable_multi_gpus_inference(self): ... # route attn through dist/<family>_xfuser.py
def enable_teacache(self, ...): ...
def enable_cfg_skip(self, ...): ...
def forward(self, x, timestep, context, ...): ...
@classmethod
def from_pretrained(cls, pretrained_model_path, subfolder=None,
transformer_additional_kwargs=None, low_cpu_mem_usage=False,
torch_dtype=torch.bfloat16):
... # mirror wan_transformer3d.py: config.json -> dict_mapping -> init_empty_weights
# -> load .bin/.safetensors -> shape-filter -> initialize missing keys -> load
```
Then register in `videox_fun/models/__init__.py`:
```python
from .<family>_transformer3d import <Family>Transformer3DModel
from .<family>_vae import AutoencoderKL<Family>
```
## Training — `scripts/<family>/train.py` (key reuse points)
```python
"""Modified from https://github.com/huggingface/diffusers/.../train_text_to_image.py"""
import argparse, gc, logging, math, os, sys
import accelerate, diffusers, torch, transformers
from accelerate import Accelerator
from diffusers.optimization import get_scheduler
from omegaconf import OmegaConf
# same sys.path bootstrap as predict scripts
from videox_fun.data import (ASPECT_RATIO_512, AspectRatioBatchImageVideoSampler,
ImageVideoDataset, ImageVideoSampler, RandomSampler,
get_closest_ratio, get_random_mask)
from videox_fun.models import AutoencoderKL<Family>, <Family>Transformer3DModel
from videox_fun.pipeline import <Family>Pipeline # REUSED for validation
from videox_fun.utils.lora_utils import create_network # for train_lora
from videox_fun.utils.utils import save_videos_grid, get_image_to_video_latent
def log_validation(vae, text_encoder, tokenizer, transformer3d, args, config,
accelerator, weight_dtype, global_step):
# build <Family>Pipeline from accelerator.unwrap_model(transformer3d),
# run validation_prompts, save_videos_grid to output_dir/sample/. Reuse the pipeline.
...
def parse_args():
parser = argparse.ArgumentParser(...)
# reuse the shared arg surface: --config_path, --pretrained_model_name_or_path,
# --train_data_dir, --train_data_meta, --video_sample_n_frames, --train_batch_size,
# --gradient_accumulation_steps, --learning_rate, --lr_scheduler, --checkpointing_steps,
# --output_dir, --mixed_precision, --gradient_checkpointing, --enable_bucket,
# --train_mode, --trainable_modules, --validation_prompts ... (add only what's needed)
return parser.parse_args()
def main():
args = parse_args()
accelerator = Accelerator(mixed_precision=args.mixed_precision, ...)
config = OmegaConf.load(args.config_path)
# load transformer/vae/text_encoder via config
# --- Dataset: pick by task (see reference.md §8) ---
# T2V/I2V base + inpaint -> ImageVideoDataset(enable_inpaint = args.train_mode != "normal")
# Control -> ImageVideoControlDataset(enable_camera_info = ...)
# Image edit -> ImageEditDataset
# Speech/audio (S2V) -> VideoSpeechDataset / VideoSpeechControlDataset
# Animate -> VideoAnimateDataset
# Distill text / GRPO / DPO -> TextDataset
# Smoke-test on the matching official demo dataset (reference.md §8), e.g.
# datasets/X-Fun-Videos-Demo + metadata_add_width_height.json for T2V/I2V.
train_dataset = ImageVideoDataset(
args.train_data_meta, args.train_data_dir,
video_sample_size=args.video_sample_size, video_sample_stride=args.video_sample_stride,
video_sample_n_frames=args.video_sample_n_frames, video_repeat=args.video_repeat,
image_sample_size=args.image_sample_size, enable_bucket=args.enable_bucket,
enable_inpaint=True if args.train_mode != "normal" else False)
# --- Sampler + DataLoader: branch on enable_bucket (see reference.md §8) ---
batch_sampler_generator = torch.Generator().manual_seed(args.seed)
if args.enable_bucket:
aspect_ratio_sample_size = {k: [x / 512 * args.video_sample_size for x in ASPECT_RATIO_512[k]] for k in ASPECT_RATIO_512}
batch_sampler = AspectRatioBatchImageVideoSampler(
sampler=RandomSampler(train_dataset, generator=batch_sampler_generator), dataset=train_dataset.dataset,
batch_size=args.train_batch_size, train_folder=args.train_data_dir, drop_last=True,
aspect_ratios=aspect_ratio_sample_size)
def collate_fn(examples):
new_examples = {"pixel_values": [], "text": []}
if args.train_mode != "normal":
new_examples.update({"mask_pixel_values": [], "mask": [], "clip_pixel_values": []})
# get_closest_ratio -> Resize/CenterCrop/Normalize -> stack; masks via get_random_mask
return new_examples
train_dataloader = torch.utils.data.DataLoader(
train_dataset, batch_sampler=batch_sampler, collate_fn=collate_fn,
num_workers=args.dataloader_num_workers,
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
else:
batch_sampler = ImageVideoSampler(RandomSampler(train_dataset, generator=batch_sampler_generator), train_dataset, args.train_batch_size)
train_dataloader = torch.utils.data.DataLoader(
train_dataset, batch_sampler=batch_sampler, num_workers=args.dataloader_num_workers,
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
# trainable-module filtering or create_network for LoRA
# optimizer + get_scheduler; accelerator.prepare; checkpoint hooks
# training loop: timestep sampling -> transformer forward -> loss -> backward
# periodic log_validation(...); final save weights / LoRA
...
if __name__ == "__main__":
main()
```
## Launcher — `scripts/<family>/train.sh`
```bash
export MODEL_NAME="models/Diffusion_Transformer/<Family>-Model"
# Test data = the official demo dataset matching the task (reference.md §8). Download once, e.g.:
# modelscope download --dataset PAI/X-Fun-Videos-Demo --local_dir ./datasets/X-Fun-Videos-Demo
# T2I -> X-Fun-Images-Demo | control -> X-Fun-{Videos,Images}-Controls-Demo
# S2V -> X-Fun-Videos-Audios-Demo | image edit -> X-Fun-Images-Edit-Demo
export DATASET_NAME="datasets/X-Fun-Videos-Demo/" # = train_data_dir (data_root); media live under train/
export DATASET_META_NAME="datasets/X-Fun-Videos-Demo/metadata_add_width_height.json" # = train_data_meta: [{"file_path","text","type","width","height"}] — see reference.md §8
# Metadata variants: VACE/subject-ref -> metadata_add_width_height_add_objects.json (X-Fun-Videos-Controls-Demo);
# audio-visual joint -> metadata_add_width_height_add_wav.json; lingbot_video -> metadata_lingbot_video_add_width_height.json
NCCL_DEBUG=INFO
accelerate launch --mixed_precision="bf16" scripts/<family>/train.py \
--config_path="config/<family>/<variant>.yaml" \
--pretrained_model_name_or_path=$MODEL_NAME \
--train_data_dir=$DATASET_NAME \
--train_data_meta=$DATASET_META_NAME \
--video_sample_n_frames=81 \
--train_batch_size=1 \
--gradient_accumulation_steps=1 \
--learning_rate=2e-05 \
--lr_scheduler="constant_with_warmup" \
--lr_warmup_steps=100 \
--checkpointing_steps=50 \
--output_dir="output_dir_<family>" \
--gradient_checkpointing \
--mixed_precision="bf16" \
--enable_bucket \
--low_vram \
--train_mode="normal" \
--trainable_modules "."
```
## Preprocessing (data gen) — `scripts/<family>/generate_<...>.py`
Offline generation of cached training data (latents / ODE-trajectory pairs / prompt embeddings). **Always multi-GPU** (`accelerate launch` + `Accelerator`) and **always safetensors** (`safetensors.torch.save_file` + an `outputs.json` index for `ImageVideoSafetensorsDataset`) — never LMDB, never `.pt`. Mirror `scripts/wan2.1_self_forcing/generate_ode_pairs.py`:
```python
# ...license header + sys.path bootstrap...
import argparse, json, math, os, torch
from accelerate import Accelerator
from omegaconf import OmegaConf
from safetensors.torch import save_file
from tqdm import tqdm
from videox_fun.models import AutoencoderKLWan, WanT5EncoderModel, WanTransformer3DModel # reuse repo models
from videox_fun.utils.utils import save_videos_grid # reuse repo IO
def main():
args = parse_args() # --pretrained_model_name_or_path --config_path --caption_path --output_folder
# --num_inference_steps --guidance_scale --shift --mixed_precision ...
accelerator = Accelerator(mixed_precision=args.mixed_precision)
device, world_size, rank = accelerator.device, accelerator.num_processes, accelerator.process_index
torch.set_grad_enabled(False) # inference-only
torch.backends.cuda.matmul.allow_tf32 = True
config = OmegaConf.load(args.config_path) # config-driven loading (Section 3)
weight_dtype = {"fp16": torch.float16, "bf16": torch.bfloat16}.get(accelerator.mixed_precision, torch.float32)
text_encoder = WanT5EncoderModel.from_pretrained(..., additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']), torch_dtype=weight_dtype).to(device).eval()
vae = AutoencoderKLWan.from_pretrained(..., additional_kwargs=OmegaConf.to_container(config['vae_kwargs'])).to(device, dtype=weight_dtype).eval()
transformer = WanTransformer3DModel.from_pretrained(..., transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs'])).to(device, dtype=weight_dtype).eval()
prompts = [l.rstrip() for l in open(args.caption_path, encoding="utf-8") if l.strip()]
os.makedirs(args.output_folder, exist_ok=True)
total_per_rank = math.ceil(len(prompts) / world_size)
for index in tqdm(range(total_per_rank), disable=rank != 0, desc="Generating"):
prompt_index = index * world_size + rank # interleaved multi-GPU shard
if prompt_index >= len(prompts):
continue
out_path = os.path.join(args.output_folder, f"{prompt_index:05d}.safetensors")
if os.path.exists(out_path): # resume: skip already-done samples
continue
prompt = prompts[prompt_index]
# ... encode prompt, sample noise, run the teacher ODE (CFG), collect latents ...
save_file( # safetensors ONLY (no lmdb / no .pt)
{"latents": latents.cpu(), "prompt_embeds": text_embeds.cpu(), "prompt_attention_mask": mask.cpu()},
out_path, metadata={"prompt": prompt},
)
accelerator.wait_for_everyone()
if accelerator.is_main_process: # rank-0 writes the JSON index
entries = [{"file_path": os.path.join(args.output_folder, f"{i:05d}.safetensors")}
for i in range(len(prompts))
if os.path.exists(os.path.join(args.output_folder, f"{i:05d}.safetensors"))]
json.dump(entries, open(os.path.join(args.output_folder, "outputs.json"), "w"), ensure_ascii=False, indent=4)
if __name__ == "__main__":
main()
```
Launcher (`generate_<...>.sh`) — `accelerate launch` uses every visible GPU:
```bash
export MODEL_NAME="models/Diffusion_Transformer/Wan2.1-T2V-1.3B"
accelerate launch --mixed_precision="bf16" scripts/<family>/generate_<...>.py \
--pretrained_model_name_or_path=$MODEL_NAME \
--config_path="config/<family>/*.yaml" \
--caption_path="datasets/prompts.txt" \
--output_folder="datasets/<family>_ode_pairs" \
--num_inference_steps=48 --guidance_scale=6.0 --shift=8.0
```
Training then reads the cache with `ImageVideoSafetensorsDataset(ann_path=".../outputs.json")` (single-file mode `{"file_path": ...}`, or per-tensor mode via `--save_per_tensor`). See reference.md §10.
> Dataset *curation* (scoring/filtering/captioning under `videox_fun/video_caption/`) is a different activity: also multi-GPU (accelerate `PartialState.split_between_processes`/`gather_object`, or vLLM tensor-parallel) but writes csv/jsonl metadata, not safetensors. See reference.md §10 “Related but different”.
+414
View File
@@ -0,0 +1,414 @@
# VideoX-Fun Integration Reference
Detailed conventions per layer. Read the mirrored family's real files alongside this — the existing code is always the source of truth.
## 1. Model definitions — `videox_fun/models/<family>_*.py`
### File naming
- Transformer / DiT: `<family>_transformer3d.py` (video) or `<family>_transformer2d.py` (image). Variants append a suffix: `_control`, `_s2v`, `_vace`, `_animate`, `_self_forcing`, `_avatar`.
- VAE: `<family>_vae.py` → class `AutoencoderKL<Family>`.
- Encoders: `<family>_text_encoder.py`, `<family>_audio_encoder.py`, `<family>_image_encoder.py`.
### Class shape (mirror `wan_transformer3d.py`)
```python
from diffusers.configuration_utils import ConfigMixin, register_to_config
from diffusers.loaders.single_file_model import FromOriginalModelMixin
from diffusers.models.modeling_utils import ModelMixin
class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin):
_supports_gradient_checkpointing = True
@register_to_config
def __init__(self, model_type='t2v', patch_size=(1,2,2), in_dim=16, dim=2048,
ffn_dim=8192, num_heads=16, num_layers=32, in_channels=16,
hidden_size=2048, ...):
super().__init__()
...
```
- Keep BOTH civitai names (`in_dim`, `dim`, `ffn_dim`) and diffusers aliases (`in_channels`, `hidden_size`) in `__init__` so either format maps cleanly.
- Implement `_set_gradient_checkpointing(self, *args, **kwargs)`.
- Attention must go through `videox_fun.models.attention_utils.attention` (backend-agnostic), not a hand-rolled `scaled_dot_product_attention`.
- Multi-GPU: expose `enable_multi_gpus_inference()` and route attention through the family's `dist/<family>_xfuser.py` processor.
- Speedups live on the model: `enable_teacache(...)`, `enable_cfg_skip(...)`, `enable_riflex(...)`.
### `from_pretrained` internals (do not simplify)
The custom classmethod must keep these behaviors (see `wan_transformer3d.py::from_pretrained`):
1. Accept `transformer_additional_kwargs`, `subfolder`, `low_cpu_mem_usage`, `torch_dtype`.
2. Read `config.json`; auto-convert foreign configs (e.g. diffsynth `has_image_input`) via a `_convert_from_*_config` helper.
3. Apply `dict_mapping`: pop it from kwargs, then for each `key: target` set `kwargs[target] = config[key]`.
4. Under `low_cpu_mem_usage`, build with `accelerate.init_empty_weights()`, load `.bin`/`.safetensors` (single file or glob all shards), and **filter by exact shape match** before loading.
5. Initialize missing keys deliberately: zero-init control/audio projections (`after_proj`, `before_proj`, `processor.k_proj/v_proj`, `audio_injector`, `cond_encoder`, ...), ones for norms, xavier for ≥2D weights, so new branches start as no-ops.
### Registry — `videox_fun/models/__init__.py`
Add an import line for every new public class, grouped with the family. Wrap optional-dependency imports in `try/except` with a helpful upgrade message (see the Qwen2.5-VL / Mistral3 blocks at the top).
## 2. Pipelines — `videox_fun/pipeline/pipeline_<family>*.py`
Mirror `pipeline_wan.py`. Required pieces:
- Module-level `retrieve_timesteps(scheduler, num_inference_steps, device, timesteps, sigmas, **kwargs)` (copied from diffusers) — reuse verbatim.
- `EXAMPLE_DOC_STRING` for the `@replace_example_docstring` decorator.
- Output dataclass:
```python
@dataclass
class <Family>PipelineOutput(BaseOutput):
videos: torch.Tensor
```
- Pipeline class:
```python
class <Family>Pipeline(DiffusionPipeline):
_optional_component = [...]
model_cpu_offload_seq = "text_encoder->transformer->vae" # order matters for offload
_callback_tensor_inputs = ["latents", "prompt_embeds", "negative_prompt_embeds"]
def __init__(self, tokenizer, text_encoder, vae, transformer, scheduler, ...): ...
def encode_prompt(...): ...
def prepare_latents(...): ...
@torch.no_grad()
@replace_example_docstring(EXAMPLE_DOC_STRING)
def __call__(self, prompt, negative_prompt=..., height=..., width=...,
num_frames=..., num_inference_steps=..., guidance_scale=...,
generator=None, ..., return_dict=True) -> Union[<Family>PipelineOutput, Tuple]: ...
```
- Import schedulers from `..utils.fm_solvers` / `..utils.fm_solvers_unipc`, models from `..models`.
- Separate pipelines per task: base (`pipeline_<family>.py`), inpaint/i2v (`_inpaint`), control (`_control`), s2v, etc. Register all in `videox_fun/pipeline/__init__.py`, adding convenience aliases (e.g. `WanI2VPipeline = WanFunInpaintPipeline`) where existing code expects them.
## 3. Config — `config/<family>/<name>.yaml` (optional)
**The YAML is not mandatory.** Decide by checkpoint layout:
- **Required** for civitai-format / custom single-file layouts, where weights and key names are not diffusers-native. The YAML supplies `transformer_additional_kwargs` (incl. `dict_mapping` mapping civitai config keys → model `__init__` kwargs), component `*_subpath`s, and `vae/text_encoder/scheduler/image_encoder` kwargs.
- **Optional** for a standard diffusers-layout checkpoint (`model_index.json` + each subfolder carrying its own `config.json`). Load components directly: `<Family>Transformer3DModel.from_pretrained(model_name, subfolder="transformer", low_cpu_mem_usage=True, torch_dtype=...)`, `AutoencoderKL<Family>.from_pretrained(model_name, subfolder="vae")`, etc. Guard the config path exactly like `examples/minimax_h3_fun/predict_v2v_control.py`:
```python
transformer_load_kwargs = {}
if config_path is not None:
from omegaconf import OmegaConf
config = OmegaConf.load(config_path)
transformer_load_kwargs.update(OmegaConf.to_container(config["transformer_additional_kwargs"], resolve=True))
transformer = <Family>Transformer3DModel.from_pretrained(model_name, subfolder="transformer", **transformer_load_kwargs, ...)
```
When you do use a YAML, the canonical schema is below (see `config/wan2.1/wan_civitai.yaml`):
```yaml
format: civitai # or diffusers — selects weight-key handling
pipeline: Wan # family label consumed by API/ComfyUI loaders
transformer_additional_kwargs:
transformer_subpath: ./ # subfolder under model_name holding the DiT
dict_mapping: # civitai config key -> model __init__ kwarg
in_dim: in_channels
dim: hidden_size
vae_kwargs:
vae_subpath: Wan2.1_VAE.pth
temporal_compression_ratio: 4
spatial_compression_ratio: 8
text_encoder_kwargs:
text_encoder_subpath: models_t5_umt5-xxl-enc-bf16.pth
tokenizer_subpath: google/umt5-xxl
text_length: 512
...
scheduler_kwargs:
scheduler_subpath: null
num_train_timesteps: 1000
shift: 5.0
...
image_encoder_kwargs: # only for i2v / models with a CLIP image encoder
image_encoder_subpath: models_clip_...pth
```
Every `*_subpath` is joined onto `model_name` in scripts. Load with `OmegaConf.load` and pass `OmegaConf.to_container(config['<section>'])` into `from_pretrained`. Use `filter_kwargs(Cls, OmegaConf.to_container(config['scheduler_kwargs']))` to build schedulers.
## 4. Inference scripts — `examples/<family>/predict_<task>.py`
Anatomy, top to bottom (see `examples/wan2.1_fun/predict_t2v.py`):
1. **`sys.path` bootstrap** (before importing `videox_fun`):
```python
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
```
2. **User config block** as top-level variables with explanatory comments, in the conventional order: `GPU_memory_mode`, `ulysses_degree`/`ring_degree`, `fsdp_dit`/`fsdp_text_encoder`, `compile_dit`, TeaCache (`enable_teacache`, `teacache_threshold`, `num_skip_start_steps`, `teacache_offload`), `cfg_skip_ratio`, Riflex (`enable_riflex`, `riflex_k`), `config_path`, `model_name`, `sampler_name`, `shift`, `transformer_path`/`vae_path`/`lora_path`, `sample_size`, `video_length`, `fps`, `weight_dtype`, `prompt`/`negative_prompt`, `guidance_scale`, `seed`, `num_inference_steps`, `lora_weight`, `save_path`.
3. **Device + config**: `device = set_multi_gpus_devices(ulysses_degree, ring_degree)`; then either `config = OmegaConf.load(config_path)` (civitai/custom layout) **or** guard `if config_path is not None:` and load components directly from a diffusers-layout checkpoint (see §3).
4. **Component loading**: transformer (`from_pretrained(..., transformer_additional_kwargs=...)`), optional `transformer_path`/`vae_path` override with `load_state_dict(strict=False)` + missing/unexpected key print, vae, tokenizer, text_encoder, and clip image encoder gated by `transformer.config.in_channels != vae.config.latent_channels`.
5. **Scheduler selection dict**: `{"Flow": FlowMatchEulerDiscreteScheduler, "Flow_Unipc": FlowUniPCMultistepScheduler, "Flow_DPM++": FlowDPMSolverMultistepScheduler}[sampler_name]`; build with `filter_kwargs`.
6. **Pipeline construction**: choose base vs inpaint/i2v/control pipeline by the model's channel condition.
7. **Multi-GPU / FSDP / compile**: if `ulysses_degree>1 or ring_degree>1` call `transformer.enable_multi_gpus_inference()` and optionally `shard_model`; if `compile_dit`, `torch.compile` each `transformer.blocks[i]`.
8. **`GPU_memory_mode` branching** — keep this exact order:
```python
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
transformer.freqs = transformer.freqs.to(device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
```
9. **TeaCache / cfg_skip / Riflex** enablement, `generator = torch.Generator(device).manual_seed(seed)`, LoRA `merge_lora`.
10. **Inference** under `torch.no_grad()`; align `video_length` to `vae.config.temporal_compression_ratio`; pass `video`/`mask_video` for i2v via `get_image_to_video_latent`.
11. **`save_results()`**: `save_videos_grid(sample, path, fps=fps)` for video, PIL save for a single frame; only rank 0 saves when multi-GPU. LoRA `unmerge_lora` after.
Other entry points to mirror when needed: `app.py` (Gradio), `launch_api.py` (API server backed by `videox_fun/api`), `post_infer*.py` (batch/queue inference).
## 5. Training scripts — `scripts/<family>/train*.py`
Mirror `scripts/wan2.1_fun/train.py`. Structure:
1. Diffusers-derived license header + `"""Modified from ..."""` note.
2. Third-party imports, then the **same `sys.path` bootstrap**, then `from videox_fun.data/models/pipeline/utils import ...`.
3. Helper funcs: `filter_kwargs`, `resize_mask`, `linear_decay`, `generate_timestep_with_lognorm`.
4. **`log_validation(vae, text_encoder, tokenizer, clip_image_encoder, transformer3d, args, config, accelerator, weight_dtype, global_step)`** — builds the **inference Pipeline** from the live (unwrapped) transformer and runs it to produce sample videos under `output_dir/sample/`. Wrapped in try/except; handles DeepSpeed (`transformer3d.config` swap) and restores VAE/text-encoder placement (`low_vram`). **Reuse the pipeline; never write a separate sampler.**
5. **`parse_args()`** — reuse the shared argument surface: `--config_path`, `--pretrained_model_name_or_path`, `--train_data_dir`, `--train_data_meta`, `--image_sample_size`/`--video_sample_size`/`--token_sample_size`, `--video_sample_n_frames`, `--video_sample_stride`, `--train_batch_size`, `--gradient_accumulation_steps`, `--learning_rate`, `--lr_scheduler`, `--lr_warmup_steps`, `--checkpointing_steps`, `--output_dir`, `--mixed_precision`, `--gradient_checkpointing`, `--enable_bucket`, `--random_hw_adapt`, `--training_with_video_token_length`, `--uniform_sampling`, `--low_vram`, `--train_mode`, `--trainable_modules`, LoRA args (`--use_lora`, `--rank`, ...), `--validation_prompts`/`--validation_paths`. Add new args only when the family genuinely needs them.
6. **`main()`** — Accelerator setup, DeepSpeed/FSDP zero-stage handling (auto-sets `save_state`), model loading via config, dataset + bucket sampler from `videox_fun.data`, trainable-module filtering / LoRA network via `create_network`, optimizer + `get_scheduler`, `accelerator.prepare`, checkpoint save/load hooks, training loop with timestep sampling, loss, `log_validation` at intervals, and final weight/LoRA save.
### Resolution args — `--video_sample_size` (+ `--fix_sample_size`)
Canvas resolution is always driven by a **single square** `--video_sample_size` (`type=int`, height = width) — never by separate `--video_sample_height` / `--video_sample_width`. When a **fixed non-square shape** is required, add `--fix_sample_size` (`nargs=2, type=int, default=None`, `[height, width]`) that overrides the square size; mirror `scripts/wan2.2_fun/train_lora.py`, `scripts/z_image/train_distill.py`. Derive the effective `height` / `width` once in `parse_args()` and reuse them everywhere downstream:
```python
parser.add_argument("--video_sample_size", type=int, default=1280)
parser.add_argument("--fix_sample_size", nargs=2, type=int, default=None,
help="Fix Sample size [height, width] to override `--video_sample_size` with a fixed non-square shape.")
...
if args.fix_sample_size is not None:
args.video_sample_height, args.video_sample_width = args.fix_sample_size
else:
args.video_sample_height = args.video_sample_width = args.video_sample_size
```
In bucket datasets `--fix_sample_size` also forces `random_hw_adapt=False` / `training_with_video_token_length=False` and bumps `video_sample_size = max(max(fix_sample_size), video_sample_size)`; in data-free scripts (e.g. `scripts/minimax_h3/train_pdd_lora.py`) it simply pins the generation canvas. Always validate the size against the patch/VAE constraint (minimax_h3: `% 32`). The `.sh` launcher passes it space-separated (`nargs=2`): `--fix_sample_size 768 1344`.
### Launcher — `scripts/<family>/train*.sh`
`export MODEL_NAME/DATASET_NAME/DATASET_META_NAME`, then `accelerate launch --mixed_precision="bf16" scripts/<family>/train.py --config_path=... <full arg list>`. Include commented I2V/control variants and DeepSpeed/NCCL notes as the existing scripts do.
### Docs — `README_TRAIN.md` + `README_TRAIN_zh-CN.md`
Aligned bilingual pair: identical section order, identical commands and parameter tables; only the prose language differs. Follow the top-level section order used across existing training READMEs.
## 6. Shared infrastructure map (reuse, never reimplement)
| Need | Import from |
|------|-------------|
| Flow/DPM/UniPC schedulers | `diffusers`, `videox_fun.utils.fm_solvers`, `videox_fun.utils.fm_solvers_unipc` |
| LoRA create/merge/unmerge/convert | `videox_fun.utils.lora_utils` |
| FP8 quantization | `videox_fun.utils.fp8_optimization` |
| Group / leaf offload hooks | `videox_fun.utils.group_offload` |
| Multi-GPU device + FSDP shard + seq-parallel attn | `videox_fun.dist` |
| Save video/audio, image→video latents, kwarg filter, dimension calc | `videox_fun.utils.utils` |
| Datasets + bucket/aspect-ratio samplers + masks | `videox_fun.data` |
| TeaCache coefficients | `videox_fun.models.cache_utils` |
## 7. Naming quick reference
| Concept | Convention | Example |
|---------|-----------|---------|
| Model file | `<family>_transformer3d.py` | `wan_transformer3d.py` |
| Model class | `<Family>Transformer3DModel` | `WanTransformer3DModel` |
| VAE class | `AutoencoderKL<Family>` | `AutoencoderKLWan` |
| Pipeline file | `pipeline_<family>.py` | `pipeline_wan.py` |
| Pipeline class | `<Family>Pipeline` | `WanPipeline` / `WanFunInpaintPipeline` |
| Config | `config/<family>/<variant>.yaml` | `config/wan2.1/wan_civitai.yaml` |
| Inference | `examples/<family>/predict_<task>.py` | `predict_t2v.py`, `predict_i2v.py`, `predict_v2v_control.py` |
| Training | `scripts/<family>/train[_<variant>].py` | `train.py`, `train_lora.py`, `train_control.py`, `train_distill.py` |
## 8. Training data pipeline — dataset & sampler selection
Pick the dataset by **task / `train_mode`**, then the sampler by **`enable_bucket`** and dataset type. All datasets/samplers come from `videox_fun.data` — never write a new one.
### Annotation format — the `train_data_meta` file (`metadata.json` / `.csv`)
Every dataset class reads an annotation file (`args.train_data_meta`) that indexes the media under `args.train_data_dir` (`data_root`). `ImageVideoDataset` accepts **`.json`** (a top-level array of records) or **`.csv`** (`csv.DictReader`; the header row is the field names). Each record for ordinary image/video training:
| Field | Required | Meaning |
|-------|----------|---------|
| `file_path` | yes | Media path, resolved **relative to `train_data_dir`** via `os.path.join(data_root, file_path)`. If `data_root is None`, `file_path` is used as-is. |
| `text` | yes | Caption / prompt. Dropped to `""` with probability `text_drop_ratio` (default `0.1`) for classifier-free guidance. |
| `type` | no | `"video"` or `"image"`; **defaults to `"image"`** when the key is absent (`data_info.get('type', 'image')`). |
```json
[
{"file_path": "train/00000000.mp4", "text": "A young woman gently turns her head to the right ...", "type": "video"},
{"file_path": "train/00000001.jpg", "text": "a dog running on the beach", "type": "image"}
]
```
The directory layout matches the index — media in a `train/` subdir, the annotation file beside it. Ready-made examples ship in `datasets/X-Fun-Videos-Demo/` (`train/*.mp4` + `metadata.json`) and `datasets/X-Fun-Images-Demo/`. The equivalent `.csv`:
```csv
file_path,text,type
train/00000000.mp4,"A young woman gently turns her head to the right ...",video
train/00000001.jpg,"a dog running on the beach",image
```
**Variant datasets append extra fields to this same record shape**, each consumed by its own class (see the table below) — e.g. camera-pose adds `action_path` (`LingbotImageVideoDataset`), object/VACE/S2V variants add object fields (`object_file_path` / `objects`). The demo folders also ship several augmented metadata variants (next subsection). Always read the target class's `get_batch` for the exact fields it consumes.
### Ready-made demo datasets — the standard test data (never invent a test set)
Smoke tests, `log_validation` checks, and doc examples all run on the official demo datasets under `datasets/`, downloaded from ModelScope as `PAI/<name>`:
```bash
modelscope download --dataset PAI/X-Fun-Videos-Demo --local_dir ./datasets/X-Fun-Videos-Demo
```
Pick the demo by **task**, matching the dataset class in the table below:
| Demo dataset (`datasets/...`) | Contents | Extra metadata fields | Task it tests | Dataset class |
|-------------------------------|----------|----------------------|---------------|---------------|
| `X-Fun-Videos-Demo` | 16 videos (832×480) in `train/` | — | T2V / I2V base + inpaint, distill | `ImageVideoDataset` |
| `X-Fun-Videos-Controls-Demo` | 16 videos in `train/` + `canny/` + `object/<video_id>/` + `wav/` | `control_file_path`, `object_file_path` (list), `audio_path` | V2V control, VACE, S2V-with-control | `ImageVideoControlDataset`, `VideoSpeechControlDataset` |
| `X-Fun-Videos-Audios-Demo` | 17 video/audio pairs: `train/` (1280×720) + `wav/` (16 kHz mono) + `pose/` | `audio_path`, `control_file_path` | Speech-driven S2V / avatar / talking-head | `VideoSpeechDataset` |
| `X-Fun-Images-Demo` | 19 images in `train/` | — | T2I full fine-tune + LoRA (z_image / flux2 / qwenimage / lens / ernie) | `ImageVideoDataset` |
| `X-Fun-Images-Controls-Demo` | 19 images in `train/` + `canny/` | `control_file_path` | Image control / ControlNet / i2i inpaint | `ImageVideoControlDataset` |
| `X-Fun-Images-Edit-Demo` | 21 records: `source/souce-<id>/` (multi-source supported) → `train/` | `source_file_path` (**list**) | Image edit (Qwen-Image-Edit family) | `ImageEditDataset` |
| `X-Fun-Videos-Lingbot-Demo` | video + `intrinsics.npy` / `poses.npy` | camera pose / action | Camera-pose world model (`lingbot_world`) | `LingbotImageVideoDataset` |
**Which metadata file to point `--train_data_meta` at** (each demo ships several variants beside the media):
| Metadata file | Use when |
|---------------|----------|
| `metadata.json` | Base format only (`file_path` / `text` / `type`) — fine for a minimal check |
| `metadata_add_width_height.json` | **Default choice.** Adds `width` / `height` so bucketing doesn't decode media (matters on slow storage such as OSS). Used by non-VACE control / S2V training too |
| `metadata_add_width_height_add_objects.json` | VACE / subject-reference training (`object_file_path` list → `object/<video_id>/`; shuffled at train time) |
| `metadata_add_width_height_add_wav.json` | Audio-visual joint models (e.g. `minimax_h3_fun` control training): `audio_path` → `wav/`. Keep the `.sh` launcher and the README on the same file |
| `metadata_lingbot_video_add_width_height.json` | `lingbot_video` — `text` is already a structured JSON caption (lives in `X-Fun-Videos-Demo`) |
| `metadata_origin.json` | Pre-processing original kept for reference; not used for training |
Regenerate the width/height variant with the shipped helper when adding your own media:
`python scripts/process_json_add_width_and_height.py --input_file datasets/<Demo>/metadata.json --output_file datasets/<Demo>/metadata_add_width_height.json`.
`audio_path` optionality differs per class (`videox_fun/data/dataset_video.py`): `VideoSpeechDataset` reads `video_dict['audio_path']` directly, so it is **required**; `VideoSpeechControlDataset` uses `.get('audio_path')` and **falls back to the video file's own audio track** when the field is absent.
### Dataset by task (all take `train_data_meta, train_data_dir, ...`)
| Task / mode | Dataset class | Used by | Key kwargs |
|-------------|--------------|---------|-----------|
| T2V / I2V base (`normal` + inpaint) | `ImageVideoDataset` | `train.py`, `train_lora.py`, t2i `train.py` | `enable_inpaint = train_mode != "normal"`, `video_sample_size/stride/n_frames`, `image_sample_size`, `video_repeat` |
| Image T2I (qwenimage/flux/z_image) | `ImageVideoDataset` | `scripts/<img>/train.py` | `image_sample_size` |
| Control (canny/pose/depth/camera) | `ImageVideoControlDataset` | `train_control*.py`, `train_control_distill.py` | `enable_camera_info = train_mode == "control_camera_ref"` |
| Image Edit (source→target) | `ImageEditDataset` | `qwenimage/train_edit*.py` | `image_sample_size` |
| Speech/audio-driven (S2V, avatar, talking) | `VideoSpeechDataset` | `mova`, `ltx2`, `minimax_h3`, `fantasytalking`, `infinitetalk`, `flashhead`, `longcatvideo/train_avatar*` | audio + video fields |
| S2V **with control** | `VideoSpeechControlDataset` | `wan2.2/train_s2v*.py`, `minimax_h3_fun/train_control*` | audio + control |
| Motion/pose animate | `VideoAnimateDataset` | `wan2.2/train_animate*.py` | motion/pose driven |
| Distill text-only branch, GRPO, DPO | `TextDataset` | `train_distill*.py` (text branch), `z_image/train_grpo_lora.py`, `train_dpo_lora.py` | reads only the `text` field; `text_drop_ratio` |
| Precomputed latents (ODE pairs) | `ImageVideoSafetensorsDataset` | `wan2.1_self_forcing/train_ode.py` | `data_root` |
| Camera-pose conditioning | `LingbotImageVideoDataset` | `lingbot_world/train.py` | `intrinsics.npy` / `poses.npy` |
| Video-only (VAE/TAEHV distill) | `VideoDataset` | `taehv/train_taehv.py` | `sample_size/stride/n_frames`, `enable_inpaint=False` |
### Sampler by condition
| Condition | Sampler | Shape |
|-----------|---------|-------|
| `enable_bucket=True` (default; image+video) | `AspectRatioBatchImageVideoSampler` | `sampler=RandomSampler(ds, generator=g), dataset=train_dataset.dataset, batch_size, train_folder=args.train_data_dir, drop_last=True, aspect_ratios=aspect_ratio_sample_size` |
| `enable_bucket=False` | `ImageVideoSampler` | `ImageVideoSampler(RandomSampler(ds, generator=g), train_dataset, batch_size)` |
| `TextDataset` (distill text branch / GRPO / DPO) | `BatchSampler` (plain) | `BatchSampler(RandomSampler(ds, generator=g), batch_size, drop_last=True)`; GRPO adds `k_repeat=args.num_image_per_prompt` |
| video-only bucket (available, not used by current scripts) | `AspectRatioBatchSampler` | — |
| image-only bucket (available, not used by current scripts) | `AspectRatioBatchImageSampler` | — |
`aspect_ratio_sample_size` is built from `ASPECT_RATIO_512` scaled by `args.video_sample_size`; `get_closest_ratio` picks the bucket inside `collate_fn`.
### Universal DataLoader creation pattern
```python
batch_sampler_generator = torch.Generator().manual_seed(args.seed)
if args.enable_bucket:
aspect_ratio_sample_size = {k: [x / 512 * args.video_sample_size for x in ASPECT_RATIO_512[k]] for k in ASPECT_RATIO_512}
batch_sampler = AspectRatioBatchImageVideoSampler(
sampler=RandomSampler(train_dataset, generator=batch_sampler_generator), dataset=train_dataset.dataset,
batch_size=args.train_batch_size, train_folder=args.train_data_dir, drop_last=True,
aspect_ratios=aspect_ratio_sample_size)
def collate_fn(examples):
new_examples = {"pixel_values": [], "text": []}
if args.train_mode != "normal": # inpaint/i2v adds mask fields
new_examples.update({"mask_pixel_values": [], "mask": [], "clip_pixel_values": []})
# bucket via get_closest_ratio -> transform (Resize/CenterCrop/Normalize) -> stack
# masked branch uses get_random_mask(...)
return new_examples
train_dataloader = torch.utils.data.DataLoader(
train_dataset, batch_sampler=batch_sampler, collate_fn=collate_fn,
persistent_workers=args.dataloader_num_workers != 0, num_workers=args.dataloader_num_workers,
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
else:
batch_sampler = ImageVideoSampler(RandomSampler(train_dataset, generator=batch_sampler_generator), train_dataset, args.train_batch_size)
train_dataloader = torch.utils.data.DataLoader(
train_dataset, batch_sampler=batch_sampler,
persistent_workers=args.dataloader_num_workers != 0, num_workers=args.dataloader_num_workers,
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
```
`collate_fn` receives the `examples` **list** (not a `batch` dict); build every batch-level field (`text`, `pixel_values`, masks) explicitly from `examples` into `new_examples`. When `--enable_text_encoder_in_dataloader`, encode prompts inside `collate_fn` and emit `encoder_hidden_states` / `encoder_attention_mask`.
## 9. Inference task matrix — predict script → pipeline → inputs
Pick the pipeline by **task**; the `predict_<task>.py` name and its inputs follow the same convention across families.
| Task | `predict_<task>.py` | Pipeline (family example) | Extra `__call__` inputs | Input helper |
|------|--------------------|---------------------------|-------------------------|--------------|
| Text→Video | `predict_t2v.py` | `WanPipeline`, `Wan2_2Pipeline`, `CogVideoXFunPipeline`, `LongCatVideoPipeline`, `LTX2Pipeline` | `prompt` only | — |
| Image→Video | `predict_i2v.py` | `WanI2VPipeline`(=`WanFunInpaintPipeline`), `Wan2_2FunInpaintPipeline`, `Wan2_2I2VPipeline`, `HunyuanVideoI2VPipeline` | `video`, `mask_video` | `get_image_to_video_latent(start_image, end_image, video_length, sample_size)` |
| Text+Image→Video (5B) | `predict_ti2v.py` | `Wan2_2TI2VPipeline` | `prompt` (+ optional image) | `get_image_to_video_latent` |
| Video→Video Control | `predict_v2v_control.py` | `WanFunControlPipeline`, `Wan2_2FunControlPipeline` | `control_video` | `get_video_to_video_latent(control_video, ...)` |
| Control + reference | `predict_v2v_control_ref.py` | `WanFunControlPipeline` | `control_video` + `ref_image` | `get_video_to_video_latent` + `get_image_latent` |
| Control + camera | `predict_v2v_control_camera.py` | `WanFunControlPipeline` | `control_video` + camera pose | — |
| VACE (control/mask/i2v/s2v) | `predict_v2v_control.py`, `predict_v2v_mask.py`, `predict_s2v.py`, `predict_i2v.py` | `WanVacePipeline`, `Wan2_2VaceFunPipeline` | control/mask/ref | — |
| Speech→Video (audio) | `predict_s2v.py` | `Wan2_2S2VPipeline`, `MiniMaxH3Pipeline`, `InfiniteTalkPipeline`, `FantasyTalkingPipeline`, `FlashHeadPipeline`, `MOVAPipeline`, `LongCatVideoAvatarPipeline` | `audio` + reference image | — |
| Animate (motion/pose) | `predict_animate.py` | `Wan2_2AnimatePipeline` | motion/pose video + ref | — |
| Subject reference | `predict_s2v.py` (phantom) | `WanFunPhantomPipeline` | reference images | — |
| Text→Image | `predict_t2i.py` | `QwenImagePipeline`, `Flux2Pipeline`, `ZImagePipeline`, `LensPipeline`, `ErnieImagePipeline` | `prompt` | — |
| Image Control (t2i) | `predict_t2i_control.py` | `QwenImageControlPipeline`, `ZImageControlPipeline`, `Flux2ControlPipeline`, `QwenImageControlNetPipeline` | `control_image` | — |
| Inpaint (i2i) | `predict_i2i_inpaint.py` | `QwenImageControlPipeline`, `ZImageControlPipeline`, `Flux2ControlPipeline` | `image` + `mask` | — |
| Image Edit | `predict_t2i_edit.py`, `predict_t2i_edit_plus.py` | `QwenImageEditPipeline`, `QwenImageEditPlusPipeline` | source image + instruction | — |
| Layered edit | `predict_i2i_layered.py` | `QwenImageLayeredPipeline` | image | — |
| Camera-pose world | `predict_i2v.py` (lingbot_world) | `Wan2_2I2VPipeline`, `WanFunLingbotWorldFastPipeline` | image + camera pose | — |
| Latent upsample | `predict_i2v_upsample.py` | `LTX2LatentUpsamplePipeline`, `WanLatentUpsamplePipeline` | low-res latent/video | — |
| AR / streaming distill | `predict_t2v_stream.py` | `WanSelfForcingPipeline` | prompt (streamed) | — |
### Predict-script variant suffixes (same task, different backend/model)
| Suffix | Meaning |
|--------|---------|
| `_tae` | Fast decode via `AutoencoderTinyWan` (TAEHV) instead of the full VAE |
| `_2.2vae` | Uses the Wan2.2 VAE (`AutoencoderKLWan3_8`) |
| `_5b` | 5B-parameter model variant |
| `turbo` / distill | Distilled model, few-step inference (e.g. `predict_turbo_*.py`) |
| `_refine` | Two-stage refine pass |
| `_ref` / `_camera` | Adds reference-image / camera conditioning |
All variants keep the identical config block, `GPU_memory_mode` branching, and `save_results()` from Section 4 — only the loaded VAE/transformer and pipeline class change.
## 10. Preprocessing — offline training-data generation (multi-GPU + safetensors)
Here "preprocessing" means **generating/caching training data offline** with the teacher / VAE / text-encoder — latents, ODE-trajectory pairs, prompt/text embeddings — so training just reads cached tensors instead of re-encoding every step. Canonical example: `scripts/wan2.1_self_forcing/generate_ode_pairs.py` (+ `generate_ode_pairs.sh`); the loader-side contract is `ImageVideoSafetensorsDataset` in `videox_fun/data/dataset_image_video.py`. Two rules are non-negotiable.
### Rule 1 — multi-GPU is mandatory
Never a single-GPU / hardcoded `cuda:0` loop. Launch with `accelerate launch` and shard work across ranks by interleaving:
```python
from accelerate import Accelerator
accelerator = Accelerator(mixed_precision=args.mixed_precision)
device, world_size, rank = accelerator.device, accelerator.num_processes, accelerator.process_index
torch.set_grad_enabled(False) # inference-only
total_per_rank = math.ceil(len(prompts) / world_size)
for index in tqdm(range(total_per_rank), disable=rank != 0):
prompt_index = index * world_size + rank # interleaved shard
if prompt_index >= len(prompts):
continue
out_path = os.path.join(args.output_folder, f"{prompt_index:05d}.safetensors")
if os.path.exists(out_path): # resume-friendly
continue
... # encode prompt / run teacher ODE / collect latents
accelerator.wait_for_everyone()
if accelerator.is_main_process: # write the JSON index once, on rank 0
json.dump([{"file_path": p} for p in all_safetensor_paths],
open(os.path.join(args.output_folder, "outputs.json"), "w"), ensure_ascii=False, indent=4)
```
Launcher (`.sh`): `accelerate launch --mixed_precision="bf16" scripts/<family>/generate_<...>.py --pretrained_model_name_or_path=... --config_path=config/<family>/*.yaml --output_folder=datasets/<...> ...`. Reuse `videox_fun.models` + config-driven `from_pretrained` (Section 3) and `videox_fun.utils.utils.save_videos_grid` for sample previews — do not write a new loader.
### Rule 2 — store as safetensors; do NOT use LMDB or `.pt`
Save every cached tensor with `safetensors.torch.save_file`, one `.safetensors` per sample (or per tensor), plus a JSON index of `{"file_path": ...}` entries:
```python
from safetensors.torch import save_file
save_file(
{"latents": latents.cpu(), "prompt_embeds": text_embeds.cpu(), "prompt_attention_mask": mask.cpu()},
out_path, # f"{prompt_index:05d}.safetensors"
metadata={"prompt": prompt},
)
```
`ImageVideoSafetensorsDataset(ann_path, data_root=None)` reads that JSON and supports two layouts:
- **Single-file (default)**: `{"file_path": "scene.safetensors"}` — whole state dict in one archive.
- **Per-tensor (`--save_per_tensor`)**: `{"file_path": "scene_dir", "latents": ".../latents.safetensors", "prompt_embeds": ".../prompt_embeds.safetensors"}` — each key loaded and merged.
**Do not** cache preprocessed data in **LMDB** or as **`.pt`/`.pth` `torch.save` pickles**. safetensors is the repo-wide standard (also used for LoRA/weight saving), is pickle-free/safe, memory-maps fast, and is exactly what `ImageVideoSafetensorsDataset` loads. (Scope: this governs cached *data tensors*; accelerate optimizer/scheduler/scaler `.pt` states written during training checkpoints are a separate mechanism and unaffected.)
### Related but different — dataset curation
Scoring / filtering / captioning under `videox_fun/video_caption/` (`compute_*.py`, `internvl2_video_recaptioning.py`) is dataset *curation*, not latent caching. It is also multi-GPU (accelerate `PartialState.split_between_processes`/`gather_object`, or vLLM `tensor_parallel_size=device_count()`), but writes csv/jsonl **metadata** (not tensors), so Rule 2 does not apply there.
+419 -450
View File
@@ -11,52 +11,59 @@ Wan-Fun:
English | [简体中文](./README_zh-CN.md) | [日本語](./README_ja-JP.md)
# Table of Contents
- [Table of Contents](#table-of-contents)
- [Introduction](#introduction)
- [Quick Start](#quick-start)
- [Video Result](#video-result)
- [How to use](#how-to-use)
- [Model zoo](#model-zoo)
- [Reference](#reference)
- [License](#license)
- [I. Introduction](#i-introduction)
- [II. Quick Start and Usage](#ii-quick-start-and-usage)
- [1. Environment Preparation](#1-environment-preparation)
- [2. Inference Generation](#2-inference-generation)
- [3. Model Training](#3-model-training)
- [III. Supported Models](#iii-supported-models)
- [IV. Video Works](#iv-video-works)
- [V. References](#v-references)
- [VI. Citation](#vi-citation)
- [VII. Limitations and Risks](#vii-limitations-and-risks)
- [VIII. License](#viii-license)
# Introduction
# I. Introduction
VideoX-Fun is a video generation pipeline that can be used to generate AI images and videos, as well as to train baseline and Lora models for Diffusion Transformer. We support direct prediction from pre-trained baseline models to generate videos with different resolutions, durations, and FPS. Additionally, we also support users in training their own baseline and Lora models to perform specific style transformations.
We will support quick pull-ups from different platforms, refer to [Quick Start](#quick-start).
# II. Quick Start and Usage
What's New:
- Added support for Wan 2.2 series models, Wan-VACE control model, Fantasy Talking digital human model, Qwen-Image, Flux image generation models, and more. [2025.10.16]
- Update Wan2.1-Fun-V1.1: Support for 14B and 1.3B model Control + Reference Image models, support for camera control, and the Inpaint model has been retrained for improved performance. [2025.04.25]
- Update Wan2.1-Fun-V1.0: Support I2V and Control models for 14B and 1.3B models, with support for start and end frame prediction. [2025.03.26]
- Update CogVideoX-Fun-V1.5: Upload I2V model and related training/prediction code. [2024.12.16]
- Reward Lora Support: Train Lora using reward backpropagation techniques to optimize generated videos, making them better aligned with human preferences. [More Information](scripts/README_TRAIN_REWARD.md). New version of the control model supports various control conditions such as Canny, Depth, Pose, MLSD, etc. [2024.11.21]
- Diffusers Support: CogVideoX-Fun Control is now supported in diffusers. Thanks to [a-r-r-o-w](https://github.com/a-r-r-o-w) for contributing support in this [PR](https://github.com/huggingface/diffusers/pull/9671). Check out the [documentation](https://huggingface.co/docs/diffusers/main/en/api/pipelines/cogvideox) for more details. [2024.10.16]
- Update CogVideoX-Fun-V1.1: Retrain i2v model, add Noise to increase the motion amplitude of the video. Upload control model training code and Control model. [2024.09.29]
- Update CogVideoX-Fun-V1.0: Initial code release! Now supports Windows and Linux. Supports video generation at arbitrary resolutions from 256x256x49 to 1024x1024x49 for 2B and 5B models. [2024.09.18]
<a id="quick-start"></a>
Function:
- [Data Preprocessing](#data-preprocess)
- [Train DiT](#dit-train)
- [Video Generation](#video-gen)
## 1. Environment Preparation
Our UI interface is as follows:
![ui](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/ui.jpg)
### 1.1 Cloud Usage: AliyunDSW
# Quick Start
### 1. Cloud usage: AliyunDSW/Docker
#### a. From AliyunDSW
DSW has free GPU time, which can be applied once by a user and is valid for 3 months after applying.
Aliyun provide free GPU time in [Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1), get it and use in Aliyun PAI-DSW to start CogVideoX-Fun within 5min!
[![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
#### b. From ComfyUI
Our ComfyUI is as follows, please refer to [ComfyUI README](comfyui/README.md) for details.
![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.jpg)
### 1.2 Local Dependency Installation
We have verified this repo execution on the following environment:
The detailed of Windows:
- OS: Windows 10
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU: Nvidia-3060 12G & Nvidia-3090 24G
The detailed of Linux:
- OS: Ubuntu 20.04, CentOS
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
We need about 60GB available on disk (for saving weights), please check!
### 1.3 Using Docker
#### c. From docker
If you are using docker, please make sure that the graphics card driver and CUDA environment have been installed correctly in your machine.
Then execute the following commands in this way:
@@ -88,30 +95,9 @@ mkdir models/Personalized_Model
# https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP
```
### 2. Local install: Environment Check/Downloading/Installation
#### a. Environment Check
We have verified this repo execution on the following environment:
### 1.4 Weight Placement
The detailed of Windows:
- OS: Windows 10
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU: Nvidia-3060 12G & Nvidia-3090 24G
The detailed of Linux:
- OS: Ubuntu 20.04, CentOS
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
We need about 60GB available on disk (for saving weights), please check!
#### b. Weights
We'd better place the [weights](#model-zoo) along the specified path:
We'd better place the [weights](#iii-supported-models) along the specified path:
**Via ComfyUI**:
Put the models into the ComfyUI weights folder `ComfyUI/models/Fun_Models/`:
@@ -137,265 +123,24 @@ Put the models into the ComfyUI weights folder `ComfyUI/models/Fun_Models/`:
│ └── your trained trainformer model / your trained lora model (for UI load)
```
# Video Result
## 2. Inference Generation
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
<a id="video-gen"></a>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
</td>
</tr>
</table>
Video and image models share the exact same inference entry, provided by scripts or UI under `examples/{model_name}/`.
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### 2.1 Entry Selection
### Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control
| Entry | Suitable Scenario | Config Granularity |
|--|--|--|
| Python file | Batch generation, parameter debugging | Full parameters |
| WebUI | Interactive experience | Common parameters only |
| ComfyUI | Existing ComfyUI workflow | Node parameters |
Generic Control Video + Reference Image:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Reference Image
</td>
<td>
Control Video
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
</td>
<tr>
</table>
Table: inference entry selection
### 2.2 GPU Memory Saving Options
Generic Control Video (Canny, Pose, Depth, etc.) and Trajectory Control:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
</td>
<tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Pan Up
</td>
<td>
Pan Left
</td>
<td>
Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload loop></video>
</td>
<tr>
<td>
Pan Down
</td>
<td>
Pan Up + Pan Left
</td>
<td>
Pan Up + Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B
Resolution-1024
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload loop></video>
</td>
</tr>
</table>
Resolution-768
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload loop></video>
</td>
</tr>
</table>
Resolution-512
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B-Control
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload loop></video>
</td>
<tr>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young bear.
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload loop></video>
</td>
</tr>
</table>
# How to Use
<h3 id="video-gen">1. Generation</h3>
#### a. GPU Memory Optimization
Since Wan2.1 has a very large number of parameters, we need to consider memory optimization strategies to adapt to consumer-grade GPUs. We provide `GPU_memory_mode` for each prediction file, allowing you to choose between `model_cpu_offload`, `model_cpu_offload_and_qfloat8`, and `sequential_cpu_offload`. This solution is also applicable to CogVideoX-Fun generation.
- `model_cpu_offload`: The entire model is moved to the CPU after use, saving some GPU memory.
@@ -404,14 +149,11 @@ Since Wan2.1 has a very large number of parameters, we need to consider memory o
`qfloat8` may slightly reduce model performance but saves more GPU memory. If you have sufficient GPU memory, it is recommended to use `model_cpu_offload`.
#### b. Using ComfyUI
For details, refer to [ComfyUI README](comfyui/README.md).
#### c. Running Python Files
### 2.3 Via Python Files
##### i. Single-GPU Inference:
- **Step 1**: Download the corresponding [weights](#model-zoo) and place them in the `models` folder.
- **Step 1**: Download the corresponding [weights](#iii-supported-models) and place them in the `models` folder.
- **Step 2**: Use different files for prediction based on the weights and prediction goals. This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:
- **Text-to-Video**:
- Modify `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_t2v.py`.
@@ -456,21 +198,29 @@ After setting the parameters, run the following command for parallel inference:
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
```
#### d. Using the Web UI
### 2.4 Via the Web UI
The web UI supports text-to-video, image-to-video, video-to-video, and controlled video generation (Canny, Pose, Depth, etc.). This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:
- **Step 1**: Download the corresponding [weights](#model-zoo) and place them in the `models` folder.
- **Step 1**: Download the corresponding [weights](#iii-supported-models) and place them in the `models` folder.
- **Step 2**: Run the file `examples/cogvideox_fun/app.py` to access the Gradio interface.
- **Step 3**: Select the generation model on the page, fill in `prompt`, `neg_prompt`, `guidance_scale`, and `seed`, click "Generate," and wait for the results. The generated videos will be saved in the `sample` folder.
### 2. Model Training
A complete model training pipeline should include data preprocessing and Video DiT training. The training process for different models is similar, and the data formats are also similar:
### 2.5 Via ComfyUI
<h4 id="data-preprocess">a. data preprocessing</h4>
For details, refer to [ComfyUI README](comfyui/README.md).
We have provided a simple demo of training the Lora model through image data, which can be found in the [wiki](https://github.com/aigc-apps/CogVideoX-Fun/wiki/Training-Lora) for details.
A complete data preprocessing link for long video segmentation, cleaning, and description can refer to [README](cogvideox/video_caption/README.md) in the video captions section.
## 3. Model Training
A complete model training pipeline consists of data preprocessing and Video DiT training.
### 3.1 Data Preprocessing
<a id="data-preprocess"></a>
Training documents for each model are unified under `scripts/{model_name}/`. For details, see [3.3 Training Documents per Model](#33-training-documents-per-model).
A complete data preprocessing link for long video segmentation, cleaning, and description can refer to [README](videox_fun/video_caption/README.md) in the video captions section.
If you want to train a text to image and video generation model. You need to arrange the dataset in this format.
@@ -519,168 +269,362 @@ You can also set the path as absolute path as follow:
]
```
<h4 id="dit-train">b. Video DiT training </h4>
### 3.2 Video DiT Training
<a id="dit-train"></a>
The training scripts and launch sh files for each model are located under `scripts/{model_name}/`. The sh file names vary by task, such as `train.sh`, `train_lora.sh`, `train_control.sh`, `train_control_distill.sh`, etc.; refer to the actual files in the directory.
If the data format is relative path during data preprocessing, please set ```scripts/{model_name}/train.sh``` as follow.
```
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
```
If the data format is absolute path during data preprocessing, please set ```scripts/train.sh``` as follow.
If the data format is absolute path during data preprocessing, please set ```scripts/{model_name}/train.sh``` as follow (`DATASET_NAME` is left empty so the dataset directory prefix is no longer concatenated).
```
export DATASET_NAME=""
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
```
Then, we run scripts/train.sh.
Finally, run the corresponding script.
```sh
sh scripts/train.sh
sh scripts/{model_name}/train.sh
```
For details on some parameter settings:
Wan2.1-Fun can be found in [Readme Train](scripts/wan2.1_fun/README_TRAIN.md) and [Readme Lora](scripts/wan2.1_fun/README_TRAIN_LORA.md).
Wan2.1 can be found in [Readme Train](scripts/wan2.1/README_TRAIN.md) and [Readme Lora](scripts/wan2.1/README_TRAIN_LORA.md).
CogVideoX-Fun can be found in [Readme Train](scripts/cogvideox_fun/README_TRAIN.md) and [Readme Lora](scripts/cogvideox_fun/README_TRAIN_LORA.md).
### 3.3 Training Documents per Model
For parameter details, training documents for each model are unified under `scripts/{model_name}/`.
# Model zoo
## 1. Wan2.2-Fun
| Name | Storage Size | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Wan2.2-Fun-A14B-InP | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP) | Wan2.2-Fun-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction. |
| Wan2.2-Fun-A14B-Control | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control)| Wan2.2-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support. |
| Wan2.2-Fun-A14B-Control-Camera | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera)| Wan2.2-Fun-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction. |
| Wan2.2-VACE-Fun-A14B | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B) | Control weights for Wan2.2 trained using the VACE scheme (based on the base model Wan2.2-T2V-A14B), supporting various control conditions such as Canny, Depth, Pose, MLSD, trajectory control, etc. It supports video generation by specifying the subject. It supports multi-resolution (512, 768, 1024) video prediction, and is trained with 81 frames at 16 FPS. It also supports multi-language prediction. |
| Wan2.2-Fun-5B-InP | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP) | Wan2.2-Fun-5B text-to-video weights trained at 121 frames, 24 FPS, supporting first/last frame prediction. |
| Wan2.2-Fun-5B-Control | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control)| Wan2.2-Fun-5B video control weights, supporting control conditions like Canny, Depth, Pose, MLSD, and trajectory control. Trained at 121 frames, 24 FPS, with multilingual prediction support. |
| Wan2.2-Fun-5B-Control-Camera | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera)| Wan2.2-Fun-5B camera lens control weights. Trained at 121 frames, 24 FPS, with multilingual prediction support. |
## 2. Wan2.2
| Name | Hugging Face | Model Scope | Description |
| Model | Baseline Training | LoRA Training | Others |
|--|--|--|--|
| Wan2.2-TI2V-5B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B) | Wan2.2-5B Text-to-Video Weights |
| Wan2.2-T2V-14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B) | Wan2.2-14B Text-to-Video Weights |
| Wan2.2-I2V-A14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B) | Wan2.2-I2V-A14B Image-to-Video Weights |
| Wan2.1-Fun | [EN](scripts/wan2.1_fun/README_TRAIN.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) | [Control EN](scripts/wan2.1_fun/README_TRAIN_CONTROL.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
| Wan2.2 | [EN](scripts/wan2.2/README_TRAIN.md) / [ZH](scripts/wan2.2/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) | [Distill EN](scripts/wan2.2/README_TRAIN_DISTILL.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
| Wan2.2-Fun | [EN](scripts/wan2.2_fun/README_TRAIN.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) | [Control LoRA EN](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md) |
| CogVideoX-Fun | [EN](scripts/cogvideox_fun/README_TRAIN.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) | [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) | [Control EN](scripts/cogvideox_fun/README_TRAIN_CONTROL.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
| Qwen-Image | [EN](scripts/qwenimage/README_TRAIN.md) / [ZH](scripts/qwenimage/README_TRAIN_zh-CN.md) | [EN](scripts/qwenimage/README_TRAIN_LORA.md) / [ZH](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) | [Edit EN](scripts/qwenimage/README_TRAIN_EDIT.md) |
| Qwen-Image-2.1 | [EN](scripts/qwenimage21/README_TRAIN.md) / [ZH](scripts/qwenimage21/README_TRAIN_zh-CN.md) | - | [Control EN](scripts/qwenimage21_fun/README_TRAIN.md) / [ZH](scripts/qwenimage21_fun/README_TRAIN_zh-CN.md) |
| Z-Image | [EN](scripts/z_image/README_TRAIN.md) / [ZH](scripts/z_image/README_TRAIN_zh-CN.md) | [EN](scripts/z_image/README_TRAIN_LORA.md) / [ZH](scripts/z_image/README_TRAIN_LORA_zh-CN.md) | [GRPO LoRA EN](scripts/z_image/README_TRAIN_GRPO_LORA.md) |
## 3. Wan2.1-Fun
For other models, check the READMEs under `scripts/{model_name}/`.
V1.1:
| Name | Storage Size | Hugging Face | Model Scope | Description |
|------|--------------|--------------|-------------|-------------|
| Wan2.1-Fun-V1.1-1.3B-InP | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP) | Wan2.1-Fun-V1.1-1.3B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction. |
| Wan2.1-Fun-V1.1-14B-InP | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP) | Wan2.1-Fun-V1.1-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction. |
| Wan2.1-Fun-V1.1-1.3B-Control | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control) | Wan2.1-Fun-V1.1-1.3B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction. |
| Wan2.1-Fun-V1.1-14B-Control | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control) | Wan2.1-Fun-V1.1-14B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction. |
| Wan2.1-Fun-V1.1-1.3B-Control-Camera | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera) | Wan2.1-Fun-V1.1-1.3B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction. |
| Wan2.1-Fun-V1.1-14B-Control-Camera | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera) | Wan2.1-Fun-V1.1-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction. |
# III. Supported Models
V1.0:
| Name | Storage Space | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Wan2.1-Fun-1.3B-InP | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP) | Wan2.1-Fun-1.3B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction. |
| Wan2.1-Fun-14B-InP | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP) | Wan2.1-Fun-14B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction. |
| Wan2.1-Fun-1.3B-Control | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control) | Wan2.1-Fun-1.3B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support. |
| Wan2.1-Fun-14B-Control | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control) | Wan2.1-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support. |
The table below summarizes currently supported model families and weights. Video and image models share the same inference and training entry. Each row represents one model family; the fourth column is an embedded four-column HTML table (Weight, Hugging Face, ModelScope, Description). 🤗 is Hugging Face, 🤖 is ModelScope (recommended for users in mainland China), and `-` means the corresponding channel has no public repo or requires authentication. For training docs of each model, see [3.3 Training Documents per Model](#33-training-documents-per-model).
## 4. Wan2.1
| Name | Hugging Face | Model Scope | Description |
| Model Family | Modality | Supported Tasks | Weight / Download / Description |
|--|--|--|--|
| Wan2.1-T2V-1.3B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B) | Wanxiang 2.1-1.3B text-to-video weights |
| Wan2.1-T2V-14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B) | Wanxiang 2.1-14B text-to-video weights |
| Wan2.1-I2V-14B-480P | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P) | Wanxiang 2.1-14B-480P image-to-video weights |
| Wan2.1-I2V-14B-720P| [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | Wanxiang 2.1-14B-720P image-to-video weights |
| Wan2.2-Fun | Video | Series trained by this project on Wan2.2, covering T2V, I2V, first/last frame, controlled generation, and camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B text-to-video weights trained at 121 frames, 24 FPS, supporting first/last frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B video control weights, supporting control conditions like Canny, Depth, Pose, MLSD, and trajectory control. Trained at 121 frames, 24 FPS, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B camera lens control weights. Trained at 121 frames, 24 FPS, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Reward LoRAs that optimize Wan2.2-Fun generated videos via reward backpropagation</td></tr></table> |
| Wan2.2-VACE-Fun | Video | Series trained by this project with the VACE scheme, covering controlled generation and subject reference | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Control weights for Wan2.2 trained using the VACE scheme (based on the base model Wan2.2-T2V-A14B), supporting various control conditions such as Canny, Depth, Pose, MLSD, trajectory control, etc. It supports video generation by specifying the subject. It supports multi-resolution (512, 768, 1024) video prediction, and is trained with 81 frames at 16 FPS. It also supports multi-language prediction.</td></tr></table> |
| Wan2.2 | Video | Official Wan weights covering T2V, I2V, audio-driven, and character animation; can be used as training baseline for Wan2.2-Fun | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-5B text/image-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B text-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B image-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B audio-to-video weights, speaker-driven digital human</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B character replacement and motion transfer weights; repo contains multiple precision files</td></tr></table> |
| Wan2.1-Fun V1.1 | Video | V1.1 series trained by this project on Wan2.1, multi-resolution (512/768/1024), 81 frames at 16fps, covering T2V, I2V, first/last frame, controlled generation, and camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr></table> |
| Wan2.1-Fun V1.0 | Video | V1.0 series trained by this project on Wan2.1; same capabilities as V1.1 but without camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Alignment LoRAs trained with reward backpropagation</td></tr></table> |
| Wan2.1 | Video | Official Wan weights covering T2V, I2V, audio-driven, and controlled generation; can be used as training baseline for Wan2.1-Fun | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">Wan 2.1-14B-720P image-to-video model weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE control and subject reference</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE control and subject reference</td></tr></table> |
| Self-Forcing / Causal-Forcing / Flex-Forcing | Video | Autoregressive distillation schemes covering streaming, interactive generation, and flexible chunked attention | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">Autoregressive distillation weights, use with Wan2.1-T2V for streaming and interactive generation; Flex-Forcing (chunk-wise causal/bidirectional attention) weights are produced by `scripts/wan2.1_flex_forcing`</td></tr></table> |
| TurboWan / TurboDiffusion | Video | Distilled few-step weights publicly released by the TurboDiffusion scheme | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B text-to-video distilled weights; officially released as .pth, the repo also ships a quantised version</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B image-to-video distilled weights; the repo contains low/high noise variants (plus quantised). Place them in Personalized_Model and reference via transformer_path / transformer_high_path</td></tr></table> |
| CogVideoX-Fun V1.5 | Video | Official CogVideoX-Fun V1.5 weights, multi-resolution (512/768/1024), 85 frames at 8fps, covering I2V and reward alignment | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024) and has been trained on 85 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
| CogVideoX-Fun V1.1 | Video | Official CogVideoX-Fun V1.1 weights, multi-resolution (512/768/1024/1280), 49 frames at 8fps, covering I2V, pose control, controlled generation, and reward alignment | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Noise has been added to the reference image, and the amplitude of motion is greater compared to V1.0.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
| CogVideoX-Fun V1.0 | Video | Legacy weights trained at 49 frames 8fps, superseded by V1.1/V1.5 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr></table> |
| HunyuanVideo | Video | Official diffusers-format weights; this project directly supports inference and LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
| MiniMax-H3 | Video | Official video generation weights and the ControlNet trained by this project | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">Official MiniMax-H3 T2V/I2V weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet trained by this project, supports multiple control conditions and trajectory control</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet trained by this project (2.0), supporting multiple control conditions, trajectory control, and inpaint checkpoints</td></tr></table> |
| TaoMate-H3 | Video+Audio | Official streaming audio-video generation adapter built on MiniMax-H3 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">Official rank-128 adapter (step-3000 EMA) with a built-in 3-step distilled schedule for streaming speech-driven generation; requires the MiniMax-H3 base weights</td></tr></table> |
| LTX-2 | Video+Audio | Official DiT audio-video joint generation weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">Official audio-video joint generation weights; repo contains multiple precision files</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">v2.3 requires community-converted diffusers weights; see Lightricks/LTX-2.3 for official weights</td></tr></table> |
| LongCat-Video | Video | Official long-video generation weights; supports LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">Official LongCat-Video T2V weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">Official LongCat-Video avatar/digital-human weights</td></tr></table> |
| FantasyTalking | Audio-driven Video | Audio-conditioned incremental weights; requires base video weights and audio encoder | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
| InfiniteTalk | Audio-driven Video | Audio-conditioned incremental weights; requires base video weights and audio encoder | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">Official InfiniteTalk audio-driven weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">Chinese audio encoder</td></tr></table> |
| FlashHead | Audio-driven Video | Official high-fidelity audio-driven head weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">SoulX FlashHead 1.3B audio-driven head weights; requires wav2vec audio encoder</td></tr></table> |
| MOVA | Video+Audio | Official MOVA weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">Image-to-video and audio-video joint generation</td></tr></table> |
| LingBot | Video | Camera-controllable world model; directory structure matches Wan2.2-I2V-A14B | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">Camera-control baseline weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA; use with Qwen3.6-27B generated structured captions</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B dense video generation weights; trainable on 1-2 GPUs</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE (3B active) video generation weights; training requires 8x80GB or more</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">Distilled few-step world model checkpoint (16 transformer shards); its VAE/T5 are reused from lingbot-world-base-cam, and inference must use the Flow_Unipc sampler</td></tr></table> |
| Phantom | Video | Incremental weights for multi-subject reference video generation; based on Wan2.1-T2V | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B version. Officially released as .pth; place in Personalized_Model and reference via transformer_path in predict file</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B version. Officially released as sharded safetensors</td></tr></table> |
| Qwen-Image | Image | Official text-to-image and image-editing weights; supports baseline and LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">Updated text-to-image version</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">Image layer-decomposition weights; splits an image into multiple editable RGBA layers</td></tr></table> |
| Qwen-Image-2.1 | Image | Official next-generation text-to-image weights; single-stream block-causal transformer with prefix KV cache | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Single-stream block-causal transformer; supports full-parameter training, prefix KV cache speeds up inference</td></tr></table> |
| Qwen-Image ControlNet | Image | Image controlled generation; supports Canny, Depth, Pose, MLSD, and Scribble | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Qwen-Image-2512, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, Scribble, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2.1-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet-Union weights for Qwen-Image-2.1 trained by this project, supporting control conditions such as Canny, Depth, Pose, MLSD, and image inpainting</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Equivalent ControlNet provided by InstantX</td></tr></table> |
| Z-Image | Image | Official text-to-image weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
| Z-Image-Fun | Image | ControlNet and distillation LoRA trained by this project on Z-Image; supports Canny, Depth, Pose, MLSD, Scribble, and Gray | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image. Compared to the first version, it adds to more layers and has been trained for a longer period. It supports multiple control conditions including Canny, Depth, Pose, MLSD, Scribble and Gray.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image-Turbo, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image-Turbo. Compared to the first version, it adds to more layers and has been trained for a longer period. It supports multiple control conditions including Canny, Depth, Pose, MLSD, and more.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">This is a Distill LoRA for Z-Image that distills both steps and CFG. This model does not require CFG and uses 8 steps for inference.</td></tr></table> |
| Flux | Image | Official FLUX.1/FLUX.2 weights and the ControlNet trained by this project | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for FLUX.2-dev</td></tr></table> |
| ERNIE-Image | Image | Official Baidu text-to-image weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">Official ERNIE-Image text-to-image weights</td></tr></table> |
| Lens | Image | Official Microsoft camera-control weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">Official Lens camera-control weights</td></tr></table> |
| Auxiliary Models | - | Non-generative models used for reward alignment, data annotation, and fast decoding | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">Scoring model used in reward backpropagation</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">Multimodal encoder used in the video captioning pipeline</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoders (~20 MB) sharing the latent spaces of the Wan2.1 / Wan2.2 VAEs, ~100x faster decoding for previews and low-memory generation; weights from <a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a></td></tr></table> |
## 5. FantasyTalking
> Notes:
> - Audio-driven and reference models (FantasyTalking, InfiniteTalk, Phantom, TaoMate-H3) are incremental weights and must be used together with the corresponding base video weights and audio encoder.
> - The TurboWan weights released by the TurboDiffusion scheme are listed above; other distillation schemes such as Flex-Forcing and PDD have no publicly released weights — train them following `scripts/{model_name}/README_TRAIN*.md` and then fill the resulting path into `transformer_path`.
> - Weight names map one-to-one to folder names under `models/Diffusion_Transformer/`. Weights within the same family are not interchangeable; choose according to the inference task. If a weight is not listed here, it is either produced by this project or should be obtained from the upstream official repository.
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Wan2.1-I2V-14B-720P | - | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | Wan 2.1-14B-720P image-to-video model weights |
| Wav2Vec | - | [🤗Link](https://huggingface.co/facebook/wav2vec2-base-960h) | [😄Link](https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h) | Wav2Vec model; place inside the Wan2.1-I2V-14B-720P folder and rename to `audio_encoder` |
| FantasyTalking model | - | [🤗Link](https://huggingface.co/acvlab/FantasyTalking/) | [😄Link](https://www.modelscope.cn/models/amap_cvlab/FantasyTalking/) | Official audio-conditioned weights |
# IV. Video Works
## 6. Qwen-Image
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Qwen-Image | [🤗Link](https://huggingface.co/Qwen/Qwen-Image) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image) | Official Qwen-Image weights |
| Qwen-Image-Edit | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit) | Official Qwen-Image-Edit weights |
| Qwen-Image-Edit-2509 | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit-2509) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509) | Official Qwen-Image-Edit-2509 weights |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
## 7. Qwen-Image-Fun
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Qwen-Image-2512-Fun-Controlnet-Union | - | [🤗Link](https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union) | [😄Link](https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union) | ControlNet weights for Qwen-Image-2512, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, Scribble, etc. |
## 8. Z-Image
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Z-Image-Turbo | [🤗Link](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [😄Link](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo) | Official weights for Z-Image-Turbo |
## 9. Z-Image-Fun
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Z-Image-Turbo-Fun-Controlnet-Union | - | [🤗Link](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union) | [😄Link](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union) | ControlNet weights for Z-Image-Turbo, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, etc. |
| Z-Image-Turbo-Fun-Controlnet-Union-2.1 | - | [🤗Link](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | [😄Link](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | ControlNet weights for Z-Image-Turbo. Compared to the first version, it adds to more layers and has been trained for a longer period. It supports multiple control conditions including Canny, Depth, Pose, MLSD, and more. |
## 10. Flux
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| FLUX.1-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.1-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev) | Official FLUX.1-dev weights |
| FLUX.2-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.2-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev) | Official FLUX.2-dev weights |
## 11. Flux-Fun
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| Flux.2-dev-Fun-Controlnet-Union | - | [🤗Link](https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union) | [😄Link](https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union) | Flux.2-dev control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc. |
## 12. HunyuanVideo
| Name | Storage | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| HunyuanVideo | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo) | - | HunyuanVideo-diffusers weights |
| HunyuanVideo-I2V | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | - | HunyuanVideo-I2V-diffusers weights |
## 13. CogVideoX-Fun
V1.5:
| Name | Storage Space | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| CogVideoX-Fun-V1.5-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP) | Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024) and has been trained on 85 frames at a rate of 8 frames per second. |
| CogVideoX-Fun-V1.5-Reward-LoRAs | - | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs) | The official reward backpropagation technology model optimizes the videos generated by CogVideoX-Fun-V1.5 to better match human preferences. |
V1.1:
| Name | Storage Space | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| CogVideoX-Fun-V1.1-2b-InP | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP) | Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. |
| CogVideoX-Fun-V1.1-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP) | Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Noise has been added to the reference image, and the amplitude of motion is greater compared to V1.0. |
| CogVideoX-Fun-V1.1-2b-Pose | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose) | Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.|
| CogVideoX-Fun-V1.1-2b-Control | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control) | Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.|
| CogVideoX-Fun-V1.1-5b-Pose | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose) | Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.|
| CogVideoX-Fun-V1.1-5b-Control | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control) | Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.|
| CogVideoX-Fun-V1.1-Reward-LoRAs | - | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs) | The official reward backpropagation technology model optimizes the videos generated by CogVideoX-Fun-V1.1 to better match human preferences. |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<details>
<summary>(Obsolete) V1.0:</summary>
<summary><b>Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control</b></summary>
Generic Control Video + Reference Image:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Reference Image
</td>
<td>
Control Video
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
</tr>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload="none"></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
Generic Control Video (Canny, Pose, Depth, etc.) and Trajectory Control:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
| Name | Storage Space | Hugging Face | Model Scope | Description |
|--|--|--|--|--|
| CogVideoX-Fun-2b-InP | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP) | Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. |
| CogVideoX-Fun-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP)| [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP)| Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. |
</details>
# Reference
<details>
<summary><b>Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera</b></summary>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Pan Up
</td>
<td>
Pan Left
</td>
<td>
Pan Right
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
Pan Down
</td>
<td>
Pan Up + Pan Left
</td>
<td>
Pan Up + Pan Right
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
<details>
<summary><b>CogVideoX-Fun-V1.1-5B</b></summary>
Resolution-1024
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
Resolution-768
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
Resolution-512
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
<details>
<summary><b>CogVideoX-Fun-V1.1-5B-Control</b></summary>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young bear.
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
# V. References
- CogVideo: https://github.com/THUDM/CogVideo/
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
@@ -696,7 +640,32 @@ V1.1:
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
- CameraCtrl: https://github.com/hehao13/CameraCtrl
# License
# VI. Citation
If you use VideoX-Fun in your research or project, please cite it as follows:
```bibtex
@misc{aigc_apps_VideoX_Fun_2026,
author = {aigc-apps},
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/aigc-apps/VideoX-Fun}
}
```
# VII. Limitations and Risks
- Generated videos may have artifacts or quality issues, especially in complex scenes.
- The model may struggle with fine details, text rendering, or specific artistic styles.
- Performance varies with input prompt quality, resolution, and other parameters.
- The technology could be misused to create misleading content (e.g., deepfakes). Users are responsible for ethical use.
- The model may reflect biases present in the training data.
- Users should respect privacy and copyright when using real people's images or videos.
We encourage responsible use and recommend implementing safeguards in production environments.
# VIII. License
This project is licensed under the [Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE).
The CogVideoX-2B model (including its corresponding Transformers module and VAE module) is released under the [Apache 2.0 License](LICENSE).
+417 -448
View File
@@ -11,52 +11,59 @@ Wan-Fun:
[English](./README.md) | [简体中文](./README_zh-CN.md) | 日本語
# 目次
- [目次](#目次)
- [紹介](#紹介)
- [クイックスタート](#クイックスタート)
- [ビデオ結果](#ビデオ結果)
- [使用方法](#使用方法)
- [モデルの場所](#モデルの場所)
- [参考文献](#参考文献)
- [ライセンス](#ライセンス)
- [一、紹介](#一紹介)
- [二、クイックスタートと使用](#二クイックスタートと使用)
- [1. 環境準備](#1-環境準備)
- [2. 推論生成](#2-推論生成)
- [3. モデルのトレーニング](#3-モデルのトレーニング)
- [三、サポート済みモデル](#三サポート済みモデル)
- [四、ビデオ作品](#四ビデオ作品)
- [五、参考文献](#五参考文献)
- [六、引用](#六引用)
- [七、制限とリスク](#七制限とリスク)
- [八、ライセンス](#八ライセンス)
# 紹介
# 一、紹介
VideoX-Funはビデオ生成のパイプラインであり、AI画像やビデオの生成、Diffusion TransformerのベースラインモデルとLoraモデルのトレーニングに使用できます。我々は、すでに学習済みのベースラインモデルから直接予測を行い、異なる解像度、秒数、FPSのビデオを生成することをサポートしています。また、ユーザーが独自のベースラインモデルやLoraモデルをトレーニングし、特定のスタイル変換を行うこともサポートしています。
異なるプラットフォームからのクイックスタートをサポートします。詳細は[クイックスタート](#クイックスタート)を参照してください。
# 二、クイックスタートと使用
新機能:
- Wan 2.2シリーズモデル、Wan-VACE制御モデル、Fantasy Talkingデジタルヒューマンモデル、Qwen-Image、Flux画像生成モデルなどのサポートを追加しました。[2025.10.16]
- Wan2.1-Fun-V1.1バージョンを更新:14Bと1.3BモデルのControl+参照画像モデルをサポート、カメラ制御にも対応。さらに、Inpaintモデルを再訓練し、性能が向上しました。[2025.04.25]
- Wan2.1-Fun-V1.0の更新:14Bおよび1.3BのI2V(画像からビデオ)モデルとControlモデルをサポートし、開始フレームと終了フレームの予測に対応。[2025.03.26]
- CogVideoX-Fun-V1.5の更新:I2Vモデルと関連するトレーニング・予測コードをアップロード。[2024.12.16]
- 報酬Loraのサポート:報酬逆伝播技術を使用してLoraをトレーニングし、生成された動画を最適化し、人間の好みによりよく一致させる。[詳細情報](scripts/README_TRAIN_REWARD.md)。新しいバージョンの制御モデルでは、Canny、Depth、Pose、MLSDなどの異なる制御条件に対応。[2024.11.21]
- diffusersのサポート:CogVideoX-Fun Controlがdiffusersでサポートされるようになりました。[a-r-r-o-w](https://github.com/a-r-r-o-w)がこの[PR](https://github.com/huggingface/diffusers/pull/9671)でサポートを提供してくれたことに感謝します。詳細は[ドキュメント](https://huggingface.co/docs/diffusers/main/en/api/pipelines/cogvideox)をご覧ください。[2024.10.16]
- CogVideoX-Fun-V1.1の更新:i2vモデルを再トレーニングし、Noiseを追加して動画の動きの範囲を拡大。制御モデルのトレーニングコードとControlモデルをアップロード。[2024.09.29]
- CogVideoX-Fun-V1.0の更新:コードを作成!WindowsとLinuxに対応しました。2Bおよび5Bモデルでの最大256x256x49から1024x1024x49までの任意の解像度の動画生成をサポート。[2024.09.18]
<a id="quick-start"></a>
機能:
- [データ前処理](#data-preprocess)
- [DiTのトレーニング](#dit-train)
- [ビデオ生成](#video-gen)
## 1. 環境準備
私たちのUIインターフェースは次のとおりです:
![ui](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/ui.jpg)
### 1.1 クラウド使用: AliyunDSW
# クイックスタート
### 1. クラウド使用: AliyunDSW/Docker
#### a. AliyunDSWから
DSWには無料のGPU時間があり、ユーザーは一度申請でき、申請後3か月間有効です。
Aliyunは[Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1)で無料のGPU時間を提供しています。取得してAliyun PAI-DSWで使用し、5分以内にCogVideoX-Funを開始できます!
[![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
#### b. ComfyUIから
私たちのComfyUIは次のとおりです。詳細は[ComfyUI README](comfyui/README.md)を参照してください。
![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.jpg)
### 1.2 ローカル依存のインストール
以下の環境でこのライブラリの実行を確認しています:
Windowsの詳細:
- OS: Windows 10
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU: Nvidia-3060 12G & Nvidia-3090 24G
Linuxの詳細:
- OS: Ubuntu 20.04, CentOS
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
重みを保存するために約60GBのディスクスペースが必要です。確認してください!
### 1.3 Dockerの使用
#### c. Dockerから
Dockerを使用する場合、マシンにグラフィックスカードドライバとCUDA環境が正しくインストールされていることを確認してください。
次のコマンドをこの方法で実行します:
@@ -88,30 +95,9 @@ mkdir models/Personalized_Model
# https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP
```
### 2. ローカルインストール: 環境チェック/ダウンロード/インストール
#### a. 環境チェック
以下の環境でこのライブラリの実行を確認しています:
### 1.4 重みの配置
Windowsの詳細:
- OS: Windows 10
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU: Nvidia-3060 12G & Nvidia-3090 24G
Linuxの詳細:
- OS: Ubuntu 20.04, CentOS
- python: python3.10 & python3.11
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
重みを保存するために約60GBのディスクスペースが必要です。確認してください!
#### b. 重み
[重み](#model-zoo)を指定されたパスに配置することをお勧めします:
[重み](#三サポート済みモデル)を指定されたパスに配置することをお勧めします:
**ComfyUIを通じて**:
モデルをComfyUIの重みフォルダ `ComfyUI/models/Fun_Models/` に入れます:
@@ -137,265 +123,24 @@ Linuxの詳細:
│ └── あなたのトレーニング済みのトランスフォーマーモデル / あなたのトレーニング済みのLoraモデル(UIロード用)
```
# ビデオ結果
## 2. 推論生成
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
<a id="video-gen"></a>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
</td>
</tr>
</table>
ビデオモデルと画像モデルの推論入口は完全に一致しており、`examples/{model_name}/`下のスクリプトまたはUIから実行します。
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### 2.1 入口の選択
### Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control
| 使用入口 | 適用シーン | 設定粒度 |
|--|--|--|
| Pythonファイル | バッチ生成、スクリプト内でパラメータを調整 | 全パラメータ |
| WebUI | 対話的な体験、モデルの迅速な切り替え | よく使うパラメータのみ |
| ComfyUI | 既存のComfyUIワークフロー | ノードパラメータ |
Generic Control Video + Reference Image:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Reference Image
</td>
<td>
Control Video
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
</td>
<tr>
</table>
表:推論入口の選択
### 2.2 顕存節約方案
Generic Control Video (Canny, Pose, Depth, etc.) and Trajectory Control:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
</td>
<tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Pan Up
</td>
<td>
Pan Left
</td>
<td>
Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload loop></video>
</td>
<tr>
<td>
Pan Down
</td>
<td>
Pan Up + Pan Left
</td>
<td>
Pan Up + Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B
解像度-1024
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload loop></video>
</td>
</tr>
</table>
解像度-768
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload loop></video>
</td>
</tr>
</table>
解像度-512
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B-Control
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload loop></video>
</td>
<tr>
<td>
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
</td>
<td>
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
</td>
<td>
若いクマ。
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload loop></video>
</td>
</tr>
</table>
# 使い方
<h3 id="video-gen">1. 生成</h3>
#### a. GPUメモリ節約方法
Wan2.1のパラメータが非常に大きいため、GPUメモリを節約し、コンシューマー向けGPUに適応させる必要があります。各予測ファイルには`GPU_memory_mode`を提供しており、`model_cpu_offload`、`model_cpu_offload_and_qfloat8`、`sequential_cpu_offload`の中から選択できます。この方法はCogVideoX-Funの生成にも適用されます。
- `model_cpu_offload`: モデル全体が使用後にCPUに移動し、一部のGPUメモリを節約します。
@@ -404,14 +149,11 @@ Wan2.1のパラメータが非常に大きいため、GPUメモリを節約し
`qfloat8`はモデルの性能を部分的に低下させる可能性がありますが、より多くのGPUメモリを節約できます。十分なGPUメモリがある場合は、`model_cpu_offload`の使用をお勧めします。
#### b. ComfyUIを使用する
詳細は[ComfyUI README](comfyui/README.md)をご覧ください。
#### c. Pythonファイルを実行する
### 2.3 Pythonファイルから
##### i. 単一GPUでの推論:
- ステップ1: 対応する[重み](#model-zoo)をダウンロードし、`models`フォルダに配置します。
- ステップ1: 対応する[重み](#三サポート済みモデル)をダウンロードし、`models`フォルダに配置します。
- ステップ2: 異なる重みと予測目標に基づいて、異なるファイルを使用して予測を行います。現在、このライブラリはCogVideoX-Fun、Wan2.1、およびWan2.1-Funをサポートしています。`examples`フォルダ内のフォルダ名で区別され、異なるモデルがサポートする機能が異なりますので、状況に応じて区別してください。以下はCogVideoX-Funを例として説明します。
- テキストからビデオ:
- `examples/cogvideox_fun/predict_t2v.py`ファイルで`prompt`、`neg_prompt`、`guidance_scale`、`seed`を変更します。
@@ -456,22 +198,29 @@ pip install yunchang==0.6.2 --progress-bar off -i https://mirrors.aliyun.com/pyp
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
```
#### d. UIインターフェースを使用する
### 2.4 UIインターフェースから
WebUIは、テキストからビデオ、画像からビデオ、ビデオからビデオ、および通常の制御付きビデオ生成(Canny、Pose、Depthなど)をサポートします。現在、このライブラリはCogVideoX-Fun、Wan2.1、およびWan2.1-Funをサポートしており、`examples`フォルダ内のフォルダ名で区別されています。異なるモデルがサポートする機能が異なるため、状況に応じて区別してください。以下はCogVideoX-Funを例として説明します。
- ステップ1: 対応する[重み](#model-zoo)をダウンロードし、`models`フォルダに配置します。
- ステップ1: 対応する[重み](#三サポート済みモデル)をダウンロードし、`models`フォルダに配置します。
- ステップ2: `examples/cogvideox_fun/app.py`ファイルを実行し、Gradioページに入ります。
- ステップ3: ページ上で生成モデルを選択し、`prompt`、`neg_prompt`、`guidance_scale`、`seed`などを入力し、「生成」をクリックして結果が生成されるのを待ちます。結果は`sample`フォルダに保存されます。
### 2. モデルのトレーニング
完全なモデルトレーニングの流れには、データの前処理とVideo DiTのトレーニングが含まれるべきです。異なるモデルのトレーニングプロセスは類似しており、データ形式も類似しています:
### 2.5 ComfyUIから
<h4 id="data-preprocess">a. データ前処理</h4>
詳細は[ComfyUI README](comfyui/README.md)をご覧ください。
画像データを使用してLoraモデルをトレーニングする簡単なデモを提供しました。詳細は[wiki](https://github.com/aigc-apps/CogVideoX-Fun/wiki/Training-Lora)をご覧ください。
長いビデオのセグメンテーション、クリーニング、説明のための完全なデータ前処理リンクは、ビデオキャプションセクションの[README](cogvideox/video_caption/README.md)を参照してください。
## 3. モデルのトレーニング
完全なモデルトレーニングパイプラインは、データ前処理とVideo DiTトレーニングで構成されます。
### 3.1 データ前処理
<a id="data-preprocess"></a>
各モデルの訓練ドキュメントは`scripts/{model_name}/`下に統一されています。詳細は[3.3 各モデルの訓練ドキュメント](#33-各モデルの訓練ドキュメント)を参照してください。
長いビデオのセグメンテーション、クリーニング、説明のための完全なデータ前処理リンクは、ビデオキャプションセクションの[README](videox_fun/video_caption/README.md)を参照してください。
テキストから画像およびビデオ生成モデルをトレーニングしたい場合。この形式でデータセットを配置する必要があります。
@@ -520,7 +269,10 @@ json_of_internal_datasets.jsonは標準のJSONファイルです。json内のfil
]
```
<h4 id="dit-train">b. Video DiTトレーニング </h4>
### 3.2 Video DiTのトレーニング
<a id="dit-train"></a>
各モデルの訓練スクリプトと起動shは`scripts/{model_name}/`下にあり、shの名称はタスクによって異なります(例:`train.sh`、`train_lora.sh`、`train_control.sh`、`train_control_distill.sh`など)。ディレクトリ内の実際のファイルを基準としてください。
データ前処理時にデータ形式が相対パスの場合、```scripts/{model_name}/train.sh```を次のように設定します。
```
@@ -528,159 +280,351 @@ export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
```
データ形式が絶対パスの場合、```scripts/train.sh```を次のように設定します。
データ形式が絶対パスの場合、同じスクリプトで次のように設定します(このとき`DATASET_NAME`は空にし、データセットディレクトリのプレフィックスを連結しません)。
```
export DATASET_NAME=""
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
```
次に、scripts/train.shを実行します。
最後に対応するスクリプトを実行します。
```sh
sh scripts/train.sh
sh scripts/{model_name}/train.sh
```
いくつかのパラメータ設定の詳細について:
Wan2.1-Funは[Readme Train](scripts/wan2.1_fun/README_TRAIN.md)と[Readme Lora](scripts/wan2.1_fun/README_TRAIN_LORA.md)を参照してください。
Wan2.1は[Readme Train](scripts/wan2.1/README_TRAIN.md)と[Readme Lora](scripts/wan2.1/README_TRAIN_LORA.md)を参照してください。
CogVideoX-Funは[Readme Train](scripts/cogvideox_fun/README_TRAIN.md)と[Readme Lora](scripts/cogvideox_fun/README_TRAIN_LORA.md)を参照してください。
# モデルの場所
### 3.3 各モデルの訓練ドキュメント
## 1. Wan2.2-Fun
パラメータ設定の詳細について、各モデルの訓練ドキュメントは`scripts/{model_name}/`下に統一されています。
| 名前 | ストレージ容量 | Hugging Face | Model Scope | 説明 |
|------|----------------|------------|-------------|------|
| Wan2.2-Fun-A14B-InP | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP) | Wan2.2-Fun-14Bのテキスト・画像から動画を生成するモデルの重み。複数の解像度で学習されており、動画の最初と最後のフレームの予測をサポートしています。 |
| Wan2.2-Fun-A14B-Control | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control) | Wan2.2-Fun-14Bの動画制御用重み。Canny、Depth、Pose、MLSDなどのさまざまな制御条件に対応しており、軌跡制御もサポートしています。512、768、1024の複数解像度での動画生成が可能で、81フレーム、16fpsで学習されています。多言語対応の予測もサポートしています。 |
| Wan2.2-Fun-A14B-Contro-Camera | 64.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera)| Wan2.2-Fun-14Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。 |
| Wan2.2-VACE-Fun-A14B | 64.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B) | VACE方式でトレーニングされたWan2.2の制御ウェイト(ベースモデルはWan2.2-T2V-A14B)。Canny、Depth、Pose、MLSD、軌道制御などの異なる制御条件をサポートします。対象を指定して動画生成が可能です。多解像度(512、768、1024)の動画予測をサポートし、81フレームで16FPSでトレーニングされています。多言語予測にも対応しています。 |
| Wan2.2-Fun-5B-InP | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP) | Wan2.2-Fun-5B テキストから動画生成用の重み。121フレーム、24 FPSで学習され、先頭/末尾フレーム予測をサポート。 |
| Wan2.2-Fun-5B-Control | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control)| Wan2.2-Fun-5B 動画制御用重み。Canny、Depth、Pose、MLSDなどの制御条件や軌道制御をサポート。121フレーム、24 FPSで学習され、多言語予測に対応。 |
| Wan2.2-Fun-5B-Control-Camera | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera)| Wan2.2-Fun-5B カメラレンズ制御用重み。121フレーム、24 FPSで学習され、多言語予測に対応。 |
## 2. Wan2.2
| モデル名 | Hugging Face | Model Scope | 説明 |
| モデル | ベーストレーニング | LoRAトレーニング | その他 |
|--|--|--|--|
| Wan2.2-TI2V-5B | [🤗リンク](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B) | [😄リンク](https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B) | 万象2.2-5B テキストから動画生成重み |
| Wan2.2-T2V-A14B | [🤗リンク](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B) | [😄リンク](https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B) | 万象2.2-14B テキストから動画生成重み |
| Wan2.2-I2V-A14B | [🤗リンク](https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B) | [😄リンク](https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B) | 万象2.2-14B 画像から動画生成重み |
| Wan2.1-Fun | [EN](scripts/wan2.1_fun/README_TRAIN.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) | [Control ZH](scripts/wan2.1_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
| Wan2.2 | [EN](scripts/wan2.2/README_TRAIN.md) / [ZH](scripts/wan2.2/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) | [Distill ZH](scripts/wan2.2/README_TRAIN_DISTILL_zh-CN.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
| Wan2.2-Fun | [EN](scripts/wan2.2_fun/README_TRAIN.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) | [Control LoRA ZH](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA_zh-CN.md) |
| CogVideoX-Fun | [EN](scripts/cogvideox_fun/README_TRAIN.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) | [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) | [Control ZH](scripts/cogvideox_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
| Qwen-Image | [EN](scripts/qwenimage/README_TRAIN.md) / [ZH](scripts/qwenimage/README_TRAIN_zh-CN.md) | [EN](scripts/qwenimage/README_TRAIN_LORA.md) / [ZH](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) | [Edit ZH](scripts/qwenimage/README_TRAIN_EDIT_zh-CN.md) |
| Qwen-Image-2.1 | [EN](scripts/qwenimage21/README_TRAIN.md) / [ZH](scripts/qwenimage21/README_TRAIN_zh-CN.md) | - | [Control EN](scripts/qwenimage21_fun/README_TRAIN.md) / [ZH](scripts/qwenimage21_fun/README_TRAIN_zh-CN.md) |
| Z-Image | [EN](scripts/z_image/README_TRAIN.md) / [ZH](scripts/z_image/README_TRAIN_zh-CN.md) | [EN](scripts/z_image/README_TRAIN_LORA.md) / [ZH](scripts/z_image/README_TRAIN_LORA_zh-CN.md) | [GRPO LoRA](scripts/z_image/README_TRAIN_GRPO_LORA.md) |
## 3. Wan2.1-Fun
その他のモデルも同様に、対応する`scripts/{model_name}/`下のREADMEを参照してください。
V1.1:
| 名称 | ストレージ容量 | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Wan2.1-Fun-V1.1-1.3B-InP | 19.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP) | Wan2.1-Fun-V1.1-1.3Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。 |
| Wan2.1-Fun-V1.1-14B-InP | 47.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP) | Wan2.1-Fun-V1.1-14Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。 |
| Wan2.1-Fun-V1.1-1.3B-Control | 19.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control)| Wan2.1-Fun-V1.1-1.3Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。 |
| Wan2.1-Fun-V1.1-14B-Control | 47.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control)| Wan2.1-Fun-V1.1-14Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。 |
| Wan2.1-Fun-V1.1-1.3B-Control-Camera | 19.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera)| Wan2.1-Fun-V1.1-1.3Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。 |
| Wan2.1-Fun-V1.1-14B-Control-Camera | 47.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera) | [😄リンク](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera)| Wan2.1-Fun-V1.1-14Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。 |
# 三、サポート済みモデル
下表は、現在サポートされているモデル系列と重みをまとめたものです。ビデオモデルと画像モデルは同じ推論・訓練インターフェースを共有しています。各行は1つのモデル系列を表し、第4列は4列のHTML埋め込みテーブル(重み、Hugging Face、ModelScope、説明)です。🤗 は Hugging Face、🤖 は ModelScope(中国国内ネットワーク向け)、`-` は該当チャネルに対応リポジトリがないか、ログイン認証が必要なことを示します。各モデルの訓練ドキュメントについては[3.3 各モデルの訓練ドキュメント](#33-各モデルの訓練ドキュメント)を参照してください。
V1.0:
| 名称 | ストレージ容量 | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Wan2.1-Fun-1.3B-InP | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP) | Wan2.1-Fun-1.3Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。 |
| Wan2.1-Fun-14B-InP | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP) | Wan2.1-Fun-14Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。 |
| Wan2.1-Fun-1.3B-Control | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control) | Wan2.1-Fun-1.3Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。 |
| Wan2.1-Fun-14B-Control | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control) | Wan2.1-Fun-14Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。 |
## 4. Wan2.1
| 名称 | Hugging Face | Model Scope | 説明 |
| モデル系列 | モダリティ | サポートタスク | 重み / ダウンロード / 説明 |
|--|--|--|--|
| Wan2.1-T2V-1.3B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B) | 万象2.1-1.3Bのテキストから動画生成する重み |
| Wan2.1-T2V-14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B) | 万象2.1-14Bのテキストから動画生成する重み |
| Wan2.1-I2V-14B-480P | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P) | 万象2.1-14B-480Pの画像から動画生成する重み |
| Wan2.1-I2V-14B-720P| [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | 万象2.1-14B-720Pの画像から動画生成する重み |
| Wan2.2-Fun | ビデオ | 本プロジェクトがWan2.2で訓練した系列。テキスト/画像から動画、首尾画像、制御生成、カメラ制御をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14Bのテキスト・画像から動画を生成するモデルの重み。複数の解像度で学習されており、動画の最初と最後のフレームの予測をサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14Bの動画制御用重み。Canny、Depth、Pose、MLSDなどのさまざまな制御条件に対応しており、軌跡制御もサポートしています。512、768、1024の複数解像度での動画生成が可能で、81フレーム、16fpsで学習されています。多言語対応の予測もサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">14B Controlにカメラモーション制御を追加</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B テキストから動画生成用の重み。121フレーム、24 FPSで学習され、先頭/末尾フレーム予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B 動画制御用重み。Canny、Depth、Pose、MLSDなどの制御条件や軌道制御をサポート。121フレーム、24 FPSで学習され、多言語予測に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B カメラレンズ制御用重み。121フレーム、24 FPSで学習され、多言語予測に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun生成動画を報酬逆伝播で最適化するReward LoRA集合</td></tr></table> |
| Wan2.2-VACE-Fun | ビデオ | 本プロジェクトがVACE方式で訓練した系列。制御生成と主題参照をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">VACE方式でトレーニングされたWan2.2の制御ウェイト(ベースモデルはWan2.2-T2V-A14B)。Canny、Depth、Pose、MLSD、軌道制御などの異なる制御条件をサポートします。対象を指定して動画生成が可能です。多解像度(512、768、1024)の動画予測をサポートし、81フレームで16FPSでトレーニングされています。多言語予測にも対応しています。</td></tr></table> |
| Wan2.2 | ビデオ | Wan公式重み。テキスト/画像から動画、音声駆動、キャラクターアニメーションをカバー。Wan2.2-Fun系列の訓練基線としても使用可能 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-5B テキスト/画像から動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B テキストから動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B 画像から動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B 音声から動画生成重み、話者駆動デジタルヒューマン</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B キャラクター置換・モーション転移重み。リポジトリに複数精度ファイルを含む</td></tr></table> |
| Wan2.1-Fun V1.1 | ビデオ | 本プロジェクトがWan2.1で訓練したV1.1系列。マルチ解像度(512/768/1024)、81フレーム16fps、テキスト/画像から動画、首尾画像、制御生成、カメラ制御をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr></table> |
| Wan2.1-Fun V1.0 | ビデオ | 本プロジェクトがWan2.1で訓練したV1.0系列。V1.1と同じ能力だがカメラ制御は非対応 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">報酬逆伝播で訓練された整列LoRA</td></tr></table> |
| Wan2.1 | ビデオ | Wan公式重み。テキスト/画像から動画、音声駆動、制御生成をカバー。Wan2.1-Fun系列の訓練基線としても使用可能 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">万象2.1-14B-720P 画像→動画モデルの重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE制御と主題参照</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE制御と主題参照</td></tr></table> |
| Self-Forcing / Causal-Forcing / Flex-Forcing | ビデオ | 自己回帰蒸留方案。ストリーミング生成、インタラクティブ生成、チャンク単位の注意をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">自己回帰蒸留重み、Wan2.1-T2Vと組み合わせて流式・インタラクティブ生成に対応;Flex-Forcing(チャンク単位の因果/双方向注意)の重みは`scripts/wan2.1_flex_forcing`で訓練して生成</td></tr></table> |
| TurboWan / TurboDiffusion | ビデオ | TurboDiffusion方案が公開した少ステップ蒸留重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3Bテキストから動画生成の蒸留重み。公式は.pth形式で公開、リポジトリに量子化版も含む</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B画像から動画生成の蒸留重み。リポジトリにlow/highの2種類のノイズモデル(量子化版も含む)を含み、Personalized_Modelに配置しpredictファイルのtransformer_path/transformer_high_pathで指定</td></tr></table> |
| CogVideoX-Fun V1.5 | ビデオ | 公式CogVideoX-Fun V1.5重み。マルチ解像度(512/768/1024)、85フレーム8fps、画像から動画と報酬整列をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024)でビデオを予測できます。85フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.5が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。</td></tr></table> |
| CogVideoX-Fun V1.1 | ビデオ | 公式CogVideoX-Fun V1.1重み。マルチ解像度(512/768/1024/1280)、49フレーム8fps、画像から動画、ポーズ制御、制御生成、報酬整列をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.1が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。</td></tr></table> |
| CogVideoX-Fun V1.0 | ビデオ | 旧版重み。49フレーム8fpsで訓練。V1.1/V1.5に置き換え済み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr></table> |
| HunyuanVideo | ビデオ | 公式diffusers形式重み。本プロジェクトは推論とLoRA訓練を直接サポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
| MiniMax-H3 | ビデオ | 公式動画生成重みと本プロジェクトが訓練したControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">MiniMax-H3公式T2V/I2V重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本プロジェクトが訓練したControlNet。複数制御条件と軌跡制御をサポート</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">本プロジェクトが訓練したControlNet(2.0版)。複数制御条件、軌跡制御、inpaint重みをサポート</td></tr></table> |
| TaoMate-H3 | ビデオ+音声 | MiniMax-H3をベースにした公式ストリーミング音声・動画生成アダプタ | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">公式rank 128アダプタ(step-3000 EMA)。3ステップ蒸留サンプリングスケジュールを内蔵し、ストリーミング音声駆動生成に対応;MiniMax-H3基盤重みと組み合わせて使用</td></tr></table> |
| LTX-2 | ビデオ+音声 | 公式DiT音声・動画共同生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">音声・動画共同生成の公式重み。リポジトリに複数精度ファイルを含む</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">v2.3はコミュニティ変換のdiffusers形式重みを使用。公式重みはLightricks/LTX-2.3を参照</td></tr></table> |
| LongCat-Video | ビデオ | 公式長尺動画生成重み。LoRA訓練をサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">LongCat-Video公式T2V重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">LongCat-Video公式アバター/デジタルヒューマン重み</td></tr></table> |
| FantasyTalking | 音声駆動ビデオ | 音声条件付き増分重み。基盤ビデオ重みと音声エンコーダが必要 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
| InfiniteTalk | 音声駆動ビデオ | 音声条件付き増分重み。基盤ビデオ重みと音声エンコーダが必要 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">InfiniteTalk公式オーディオ駆動重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">中国語音声エンコーダ</td></tr></table> |
| FlashHead | 音声駆動ビデオ | 公式高品質頭部動作デジタルヒューマン重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">SoulX FlashHead 1.3B 音声駆動頭部重み。wav2vec音声エンコーダが必要</td></tr></table> |
| MOVA | ビデオ+音声 | 公式MOVA重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">画像から動画と音声・動画共同生成</td></tr></table> |
| LingBot | ビデオ | カメラ制御可能なワールドモデル。ディレクトリ構造はWan2.2-I2V-A14Bと一致 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">カメラ制御基線重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA。Qwen3.6-27Bで構造化キャプションを生成して使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B dense版動画生成重み。1〜2枚のGPUで訓練可能</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE(3Bアクティブ)動画生成重み。訓練には8×80GB以上を推奨</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">蒸留少ステップワールドモデル(transformerは16シャード)。VAE/T5はlingbot-world-base-camを再利用し、推論にはFlow_Unipcサンプラーを使用</td></tr></table> |
| Phantom | ビデオ | 複数主体参照による動画生成の増分重み。Wan2.1-T2Vベース | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B版。公式は.pth形式で公開。Personalized_Modelに配置しpredictファイルのtransformer_pathで指定</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B版。公式は分割safetensors形式で公開</td></tr></table> |
| Qwen-Image | 画像 | 公式テキストから画像生成・画像編集重み。基線とLoRA訓練をサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">テキストから画像生成の更新版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">画像レイヤー分解重み。画像を複数の編集可能なRGBAレイヤーに分解可能</td></tr></table> |
| Qwen-Image-2.1 | 画像 | 公式次世代テキストから画像生成重み。シングルストリームblock-causal構造、プレフィックスKV cacheに対応 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">シングルストリームblock-causal構造。全パラメータ訓練をサポート、プレフィックスKV cacheで推論を高速化</td></tr></table> |
| Qwen-Image ControlNet | 画像 | 画像制御生成。Canny、Depth、Pose、MLSD、Scribbleをサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Qwen-Image-2512のControlNet重み。Canny、Depth、Pose、MLSD、Scribbleなど、複数の制御条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2.1-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本プロジェクトがQwen-Image 2.1向けに訓練したControlNet-Union。Canny、Depth、Pose、MLSDなどの制御条件と画像補完(inpaint)をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">InstantX提供の同種ControlNet</td></tr></table> |
| Z-Image | 画像 | 公式テキストから画像生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
| Z-Image-Fun | 画像 | 本プロジェクトがZ-Imageで訓練したControlNetと蒸留LoRA。Canny、Depth、Pose、MLSD、Scribble、Grayをサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Z-ImageのControlNet重み、Canny、Depth、Pose、MLSD、ScribbleおよびGrayなど複数の制御条件に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Z-Image-Turbo用のControlNet重み。Canny、Depth、Pose、MLSDなど複数の制御条件をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Z-Image-TurboのControlNet重み。第1版と比較して、より多くの層に追加され、より長時間トレーニングされています。Canny、Depth、Pose、MLSDなど、複数の制御条件をサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">これはZ-Image用の蒸留LoRAで、ステップ数とCFGの両方を蒸留します。このモデルはCFGを必要とせず、推論には8ステップを使用します。</td></tr></table> |
| Flux | 画像 | 公式FLUX.1/FLUX.2重みと本プロジェクトが訓練したControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">FLUX.2-dev用ControlNet重み</td></tr></table> |
| ERNIE-Image | 画像 | Baidu公式テキストから画像生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">ERNIE-Image公式画像生成重み</td></tr></table> |
| Lens | 画像 | Microsoft公式カメラ制御重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">Lens公式カメラ制御重み</td></tr></table> |
| 補助モデル | - | 生成モデルではなく、報酬整列、データアノテーション、高速デコードに使用 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">報酬逆伝播で使用されるスコアリングモデル</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">動画キャプション生成パイプラインで使用されるマルチモーダルエンコーダ</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoder(約20MB)。Wan2.1/Wan2.2 VAEと同じlatent空間を共有し、デコード速度は完全なVAEの約100倍で、高速プレビューや低メモリ生成に使用;重みは<a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a>より</td></tr></table> |
## 5. FantasyTalking
> 補足説明:
> - 音声駆動・参照系モデル(FantasyTalking、InfiniteTalk、Phantom、TaoMate-H3)は増分重みであり、対応する基盤ビデオ重みと音声エンコーダを同時にダウンロードする必要があります。
> - TurboDiffusion方案はTurboWan系列の蒸留重みを公開済みです(上表参照)。Flex-Forcing、PDDなどその他の蒸留方案は公開重みがなく、`scripts/{model_name}/README_TRAIN*.md`で訓練後、`transformer_path`に指定して使用できます。
> - 重み名は`models/Diffusion_Transformer/`下のフォルダ名と一対一で対応します。同じ系列内の各重みは互換性がないため、推論タスクに応じて選択してください。ここに掲載されていない重みは、本プロジェクトの訓練成果物、または上流の公式リポジトリから取得する必要があります。
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Wan2.1-I2V-14B-720P | - | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | 万象2.1-14B-720P 画像→動画モデルの重み |
| Wav2Vec | - | [🤗Link](https://huggingface.co/facebook/wav2vec2-base-960h) | [😄Link](https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h) | Wav2Vecモデル。Wan2.1-I2V-14B-720Pフォルダ内に配置し、`audio_encoder` という名前に変更してください |
| FantasyTalking model | - | [🤗Link](https://huggingface.co/acvlab/FantasyTalking/) | [😄Link](https://www.modelscope.cn/models/amap_cvlab/FantasyTalking/) | 公式Audio Condition重み |
# 四、ビデオ作品
## 6. Qwen-Image
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Qwen-Image | [🤗Link](https://huggingface.co/Qwen/Qwen-Image) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image) | Qwen-Image 公式重み |
| Qwen-Image-Edit | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit) | Qwen-Image-Edit 公式重み |
| Qwen-Image-Edit-2509 | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit-2509) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509) | Qwen-Image-Edit-2509 公式重み |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
## 7. Qwen-Image-Fun
| 名前 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Qwen-Image-2512-Fun-Controlnet-Union | - | [🤗リンク](https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union) | [😄リンク](https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union) | Qwen-Image-2512のControlNet重み。Canny、Depth、Pose、MLSD、Scribbleなど、複数の制御条件をサポートします。 |
## 8. Z-Image
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Z-Image-Turbo | [🤗リンク](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [😄リンク](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo) | Z-Image-Turboの公式重み |
## 9. Z-Image-Fun
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| Z-Image-Turbo-Fun-Controlnet-Union | - | [🤗リンク](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union) | [😄リンク](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union) | Z-Image-Turbo用のControlNet重み。Canny、Depth、Pose、MLSDなど複数の制御条件をサポート。 |
| Z-Image-Turbo-Fun-Controlnet-Union-2.1 | - | [🤗リンク](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | [😄リンク](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | Z-Image-TurboのControlNet重み。第1版と比較して、より多くの層に追加され、より長時間トレーニングされています。Canny、Depth、Pose、MLSDなど、複数の制御条件をサポートしています。 |
## 10. Flux
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| FLUX.1-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.1-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev)| FLUX.1-dev 公式重み |
| FLUX.2-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.2-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev) | FLUX.2-dev 公式重み |
## 11. Flux-Fun
| 名前 | ストレージ | Hugging Face | ModelScope | 説明 |
|--|--|--|--|--|
| Flux.2-dev-Fun-Controlnet-Union | - | [🤗リンク](https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union) | [😄リンク](https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union) | Flux.2-dev 用の ControlNet 重みで、Canny、Depth、Pose、MLSD など様々な制御条件をサポートします。 |
## 12. HunyuanVideo
| 名称 | ストレージ | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| HunyuanVideo | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo) | - | HunyuanVideo-diffusers 公式重み |
| HunyuanVideo-I2V | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | - | HunyuanVideo-I2V-diffusers 公式重み |
## 13. CogVideoX-Fun
V1.5:
| 名称 | ストレージスペース | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| CogVideoX-Fun-V1.5-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP) | 公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024)でビデオを予測できます。85フレーム、8フレーム/秒でトレーニングされています。 |
| CogVideoX-Fun-V1.5-Reward-LoRAs | - | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs) | 公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.5が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。 |
V1.1:
| 名称 | ストレージスペース | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| CogVideoX-Fun-V1.1-2b-InP | 13.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP) | 公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。 |
| CogVideoX-Fun-V1.1-5b-InP | 20.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP) | 公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。 |
| CogVideoX-Fun-V1.1-2b-Pose | 13.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose) | 公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。|
| CogVideoX-Fun-V1.1-2b-Control | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control) | 公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。|
| CogVideoX-Fun-V1.1-5b-Pose | 20.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose) | 公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。|
| CogVideoX-Fun-V1.1-5b-Control | 20.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control) | 公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。|
| CogVideoX-Fun-V1.1-Reward-LoRAs | - | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs) | 公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.1が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。 |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<details>
<summary>(Obsolete) V1.0:</summary>
<summary><b>Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control</b></summary>
汎用制御動画 + 参照画像:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
参照画像
</td>
<td>
制御動画
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
</tr>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload="none"></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
汎用制御動画(Canny、Pose、Depth など)と軌跡制御:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
| 名称 | ストレージスペース | Hugging Face | Model Scope | 説明 |
|--|--|--|--|--|
| CogVideoX-Fun-2b-InP | 13.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP) | [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP) | 公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。 |
| CogVideoX-Fun-5b-InP | 20.0 GB | [🤗リンク](https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP)| [😄リンク](https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP)| 公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。|
</details>
# 参考文献
<details>
<summary><b>Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera</b></summary>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Pan Up
</td>
<td>
Pan Left
</td>
<td>
Pan Right
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
Pan Down
</td>
<td>
Pan Up + Pan Left
</td>
<td>
Pan Up + Pan Right
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
<details>
<summary><b>CogVideoX-Fun-V1.1-5B</b></summary>
解像度-1024
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
解像度-768
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
解像度-512
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
<details>
<summary><b>CogVideoX-Fun-V1.1-5B-Control</b></summary>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
</td>
<td>
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
</td>
<td>
若いクマ。
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
</details>
# 五、参考文献
- CogVideo: https://github.com/THUDM/CogVideo/
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
@@ -696,7 +640,32 @@ V1.1:
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
- CameraCtrl: https://github.com/hehao13/CameraCtrl
# ライセンス
# 六、引用
研究やプロジェクトでVideoX-Funを使用する場合は、以下の形式で引用してください:
```bibtex
@misc{aigc_apps_VideoX_Fun_2026,
author = {aigc-apps},
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/aigc-apps/VideoX-Fun}
}
```
# 七、制限とリスク
- 生成された動画には、特に複雑なシーンでアーティファクトや品質の問題がある場合があります。
- モデルは、細かい詳細、テキストのレンダリング、または特定の芸術スタイルで苦労する場合があります。
- パフォーマンスは、入力プロンプトの品質、解像度、その他のパラメータによって異なります。
- この技術は、誤解を招くコンテンツ(例:ディープフェイク)を作成するために悪用される可能性があります。ユーザーは倫理的な使用に責任を持ちます。
- モデルは、トレーニングデータに存在するバイアスを反映する可能性があります。
- ユーザーは、実在の人物の画像や動画を使用する際、プライバシーと著作権を尊重する必要があります。
責任ある使用を推奨し、本番環境でのセーフガードの実装をお勧めします。
# 八、ライセンス
このプロジェクトは[Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE)の下でライセンスされています。
CogVideoX-2Bモデル(対応するTransformersモジュール、VAEモジュールを含む)は、[Apache 2.0ライセンス](LICENSE)の下でリリースされています。
+330 -502
View File
@@ -11,83 +11,37 @@ Wan-Fun:
[English](./README.md) | 简体中文 | [日本語](./README_ja-JP.md)
# 目录
- [目录](#目录)
- [简介](#简介)
- [快速启动](#快速启动)
- [视频作品](#视频作品)
- [如何使用](#如何使用)
- [模型地址](#模型地址)
- [参考文献](#参考文献)
- [许可证](#许可证)
- [一、简介](#一简介)
- [二、快速开始与使用](#二快速开始与使用)
- [1. 环境准备](#1-环境准备)
- [2. 推理生成](#2-推理生成)
- [3. 模型训练](#3-模型训练)
- [三、已支持的模型](#三已支持的模型)
- [四、视频作品](#四视频作品)
- [五、参考文献](#五参考文献)
- [六、引用](#六引用)
- [七、限制与风险](#七限制与风险)
- [八、许可证](#八许可证)
# 简介
VideoX-Fun是一个视频生成的pipeline,可用于生成AI图片与视频、训练Diffusion Transformer的基线模型与Lora模型,我们支持从已经训练好的基线模型直接进行预测,生成不同分辨率,不同秒数、不同FPS的视频,也支持用户训练自己的基线模型与Lora模型,进行一定的风格变换。
# 一、简介
VideoX-Fun是一个图片与视频生成的pipeline,可用于生成AI图片与视频、训练Diffusion Transformer的基线模型与Lora模型。我们同时支持视频与图片两类Diffusion Transformer模型:视频侧涵盖Wan2.1/Wan2.2(含Fun、VACE、Animate、S2V等变体)、CogVideoX-Fun、HunyuanVideo、MiniMax-H3、LTX-2、LongCat-Video、FantasyTalking与LingBot等,图片侧涵盖Qwen-Image(含Edit)、Z-Image(含Turbo)、Flux/Flux2与ERNIE-Image等,完整列表见[已支持的模型](#三已支持的模型)。在此基础上,我们支持从已经训练好的基线模型直接进行预测,生成不同分辨率、不同秒数、不同FPS的视频与不同分辨率的图片,也支持用户训练自己的基线模型与Lora模型,进行一定的风格变换。
我们会逐渐支持从不同平台快速启动,请参阅 [快速启动](#快速启动)。
新特性:
- 更新支持Wan2.2系列模型、Wan-VACE控制模型、支持Fantasy Talking数字人模型、Qwen-Image和Flux图片生成模型等。[2025.10.16]。
- 更新Wan2.1-Fun-V1.1版本:支持14B与1.3B模型Control+参考图模型,支持镜头控制,另外Inpaint模型重新训练,性能更佳。[2025.04.25]
- 更新Wan2.1-Fun-V1.0版本:支持14B与1.3B模型的I2V和Control模型,支持首尾图预测。[2025.03.26]
- 更新CogVideoX-Fun-V1.5版本:上传I2V模型与相关训练预测代码。[2024.12.16]
- 奖励Lora支持:通过奖励反向传播技术训练Lora,以优化生成的视频,使其更好地与人类偏好保持一致,[更多信息](scripts/README_TRAIN_REWARD.md)。新版本的控制模型,支持不同的控制条件,如Canny、Depth、Pose、MLSD等。[2024.11.21]
- diffusers支持:CogVideoX-Fun Control现在在diffusers中得到了支持。感谢 [a-r-r-o-w](https://github.com/a-r-r-o-w)在这个 [PR](https://github.com/huggingface/diffusers/pull/9671)中贡献了支持。查看[文档](https://huggingface.co/docs/diffusers/main/en/api/pipelines/cogvideox)以了解更多信息。[2024.10.16]
- 更新CogVideoX-Fun-V1.1版本:重新训练i2v模型,添加Noise,使得视频的运动幅度更大。上传控制模型训练代码与Control模型。[2024.09.29]
- 更新CogVideoX-Fun-V1.0版本:创建代码!现在支持 Windows 和 Linux。支持2b与5b最大256x256x49到1024x1024x49的任意分辨率的视频生成。[2024.09.18]
# 二、快速开始与使用
功能概览:
- [数据预处理](#data-preprocess)
- [训练DiT](#dit-train)
- [模型生成](#video-gen)
<a id="quick-start"></a>
我们的ui界面如下:
![ui](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/ui.jpg)
## 1. 环境准备
# 快速启动
### 1. 云使用: AliyunDSW/Docker
#### a. 通过阿里云 DSW
### 1.1 云使用: AliyunDSW
DSW 有免费 GPU 时间,用户可申请一次,申请后3个月内有效。
阿里云在[Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1)提供免费GPU时间,获取并在阿里云PAI-DSW中使用,5分钟内即可启动CogVideoX-Fun。
阿里云在[Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1)提供免费GPU时间,获取并在阿里云PAI-DSW中使用,5分钟内即可启动VideoX-Fun。
[![DSW Notebook](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/easyanimate/asset/dsw.png)](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
#### b. 通过ComfyUI
我们的ComfyUI界面如下,具体查看[ComfyUI README](comfyui/README.md)。
![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.jpg)
### 1.2 本地依赖安装
#### c. 通过docker
使用docker的情况下,请保证机器中已经正确安装显卡驱动与CUDA环境,然后以此执行以下命令:
```
# pull image
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
# enter image
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
# clone code
git clone https://github.com/aigc-apps/VideoX-Fun.git
# enter VideoX-Fun's dir
cd VideoX-Fun
# download weights
mkdir models/Diffusion_Transformer
mkdir models/Personalized_Model
# Please use the hugginface link or modelscope link to download the model.
# CogVideoX-Fun
# https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
# https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP
# Wan
# https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP
# https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP
```
### 2. 本地安装: 环境检查/下载/安装
#### a. 环境检查
我们已验证该库可在以下环境中执行:
Windows 的详细信息:
@@ -104,12 +58,71 @@ Linux 的详细信息:
- pytorch: torch2.2.0
- CUDA: 11.8 & 12.1
- CUDNN: 8+
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G & Nvidia-H800 80G
我们需要大约 60GB 的可用磁盘空间,请检查!
**方式一:使用requirements.txt**
#### b. 权重放置
我们最好将[权重](#model-zoo)按照指定路径进行放置:
```bash
pip install -r requirements.txt
```
**方式二:手动安装依赖**
```bash
# 核心依赖,与requirements.txt保持一致
pip install Pillow einops safetensors timm tomesd albumentations librosa "torch>=2.1.2" torchdiffeq torchsde decord datasets numpy scikit-image
pip install omegaconf SentencePiece imageio[ffmpeg] imageio[pyav] tensorboard beautifulsoup4 ftfy func_timeout onnxruntime
pip install "peft>=0.17.0" "accelerate>=0.25.0" "gradio>=3.41.2" "diffusers>=0.30.1" "transformers>=4.46.2"
# 权重下载
pip install modelscope
# 多卡并行推理需要,推荐固定版本,单卡可跳过
pip install "xfuser==0.4.2"
# opencv统一使用headless版本,避免部分环境下的GUI依赖
pip uninstall opencv-python opencv-contrib-python opencv-python-headless -y
pip install opencv-python-headless
# 训练可选:DeepSpeed训练需要,固定numpy版本以避免兼容性问题
pip install deepspeed==0.17.0 numpy==1.26.4
# 加速可选:安装后注意力自动使用Flash Attention后端,未安装时回退到SDPA
pip install flash-attn --no-build-isolation
```
> 说明:`torch`与`flash-attn`建议按照本机的CUDA版本从官方渠道安装指定版本,国内网络可追加`-i https://mirrors.aliyun.com/pypi/simple/`加速,具体依赖请以[requirements.txt](requirements.txt)为准。
### 1.3 使用Docker
使用docker的情况下,请保证机器中已经正确安装显卡驱动与CUDA环境,然后以此执行以下命令:
```
# pull image
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
# enter image
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
# clone code
git clone https://github.com/aigc-apps/VideoX-Fun.git
# enter VideoX-Fun's dir
cd VideoX-Fun
```
### 1.4 权重放置
我们最好将[权重](#三已支持的模型)按照指定路径进行放置:
**运行自身的python文件或ui界面**:
```
📦 models/
├── 📂 Diffusion_Transformer/
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
│ ├── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
│ ├── 📂 Z-Image/
│ └── 📂 Qwen-Image/
├── 📂 Personalized_Model/
│ └── your trained trainformer model / your trained lora model (for UI load)
```
视频模型与图片模型的权重均统一放在`models/Diffusion_Transformer/`下,文件夹名与[已支持的模型](#三已支持的模型)中的权重名保持一致。
**通过comfyui**:
将模型放入Comfyui的权重文件夹`ComfyUI/models/Fun_Models/`:
@@ -123,297 +136,42 @@ Linux 的详细信息:
│ └── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
```
**运行自身的python文件或ui界面**:
```
📦 models/
├── 📂 Diffusion_Transformer/
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
│ └── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
├── 📂 Personalized_Model/
│ └── your trained trainformer model / your trained lora model (for UI load)
```
## 2. 推理生成
# 视频作品
<a id="video-gen"></a>
视频模型与图片模型的推理入口完全一致,均由`examples/{model_name}/`下的脚本或界面提供,模型清单见[已支持的模型](#三已支持的模型)。
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
### 2.1 入口选择
| 使用入口 | 适合场景 | 可配置粒度 |
|--|--|--|
| python文件 | 批量生成、参数写在脚本里调试 | 全量参数,含`GPU_memory_mode`、`transformer_path`、`lora_path` |
| webui | 交互体验、快速切换模型 | 常见参数,显存方案仅4档,见2.2 |
| ComfyUI | 已有ComfyUI工作流、节点化组合 | 节点参数,权重放置见1.5 |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### 2.2 显存节省方案
基线模型的参数量普遍很大,为适应消费级显卡,每个预测文件都提供了GPU_memory_mode,视频模型与图片模型通用。可选项按省显存程度从高到低排列,与代码中的判断顺序一致:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control
Generic Control Video + Reference Image:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Reference Image
</td>
<td>
Control Video
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
</td>
<tr>
</table>
Generic Control Video (Canny, Pose, Depth, etc.) and Trajectory Control:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
</td>
<tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
Pan Up
</td>
<td>
Pan Left
</td>
<td>
Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload loop></video>
</td>
<tr>
<td>
Pan Down
</td>
<td>
Pan Up + Pan Left
</td>
<td>
Pan Up + Pan Right
</td>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B
Resolution-1024
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload loop></video>
</td>
</tr>
</table>
Resolution-768
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload loop></video>
</td>
</tr>
</table>
Resolution-512
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload loop></video>
</td>
</tr>
</table>
### CogVideoX-Fun-V1.1-5B-Control
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload loop></video>
</td>
<tr>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
</td>
<td>
A young bear.
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload loop></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload loop></video>
</td>
</tr>
</table>
# 如何使用
<h3 id="video-gen">1. 生成 </h3>
#### a、显存节省方案
由于Wan2.1的参数非常大,我们需要考虑显存节省方案,以节省显存适应消费级显卡。我们给每个预测文件都提供了GPU_memory_mode,可以在model_cpu_offload,model_cpu_offload_and_qfloat8,sequential_cpu_offload中进行选择。该方案同样适用于CogVideoX-Fun的生成。
- model_cpu_offload代表整个模型在使用后会进入cpu,可以节省部分显存。
- model_cpu_offload_and_qfloat8代表整个模型在使用后会进入cpu,并且对transformer模型进行了float8的量化,可以节省更多的显存。
- sequential_cpu_offload代表模型的每一层在使用后会进入cpu,速度较慢,节省大量显存。
- sequential_cpu_offload:模型的每一层在使用后会进入cpu,速度较慢,节省大量显存。
- model_group_offload:以leaf层级在cpu与gpu之间搬运权重,并借助stream异步预取,兼顾速度与显存。
- model_cpu_offload_and_qfloat8:整个模型在使用后会进入cpu,并且对transformer模型进行了float8的量化,可以节省更多的显存。
- model_cpu_offload:整个模型在使用后会进入cpu,可以节省部分显存。
- model_full_load_and_qfloat8:模型常驻gpu,仅对transformer做float8量化,显存临界且对速度要求较高时可选。
- 默认(传入model_full_load或其他取值):模型全部进入gpu,速度最快,显存需求最高。
qfloat8会部分降低模型的性能,但可以节省更多的显存。如果显存足够,推荐使用model_cpu_offload。
#### b、通过comfyui
具体查看[ComfyUI README](comfyui/README.md)。
> 注意:`app.py`中仅提供model_full_load、model_cpu_offload、model_cpu_offload_and_qfloat8、sequential_cpu_offload四种模式,`model_group_offload`与`model_full_load_and_qfloat8`需在python预测文件中使用;另外compile类加速与`sequential_cpu_offload`、fsdp_dit不兼容。
#### c、运行python文件
### 2.3 通过python文件
推理脚本统一命名为`predict_{任务}.py`,在脚本内修改`model_name`、prompt等参数后直接运行,结果保存到脚本中`save_path`指定的目录。视频模型与图片模型的差别只在任务后缀,例如`examples/cogvideox_fun/predict_t2v.py`、`examples/wan2.2_fun/predict_i2v.py`与`examples/z_image/predict_t2i.py`、`examples/qwenimage/predict_t2i_edit.py`。具体某个模型支持哪些任务,以`examples/{model_name}/`下实际存在的脚本为准。
##### i、单卡运行:
**i、单卡运行**:以CogVideoX-Fun为例。
- 步骤1:下载对应[权重](#model-zoo)放入models文件夹。
- 步骤2:根据不同的权重与预测目标使用不同的文件进行预测。当前该库支持CogVideoX-Fun、Wan2.1和Wan2.1-Fun,在examples文件夹下用文件夹名以区分,不同模型支持的功能不同,请视具体情况予以区分。以CogVideoX-Fun为例。
- 步骤1:下载对应[权重](#三已支持的模型)并按1.5放入models文件夹。
- 步骤2:根据不同的权重与预测目标使用不同的文件进行预测。
- 文生视频:
- 使用examples/cogvideox_fun/predict_t2v.py文件中修改prompt、neg_prompt、guidance_scale和seed。
- 而后运行examples/cogvideox_fun/predict_t2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos文件夹中。
- 而后运行examples/cogvideox_fun/predict_t2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos-t2v文件夹中。
- 图生视频:
- 使用examples/cogvideox_fun/predict_i2v.py文件中修改validation_image_start、validation_image_end、prompt、neg_prompt、guidance_scale和seed。
- validation_image_start是视频的开始图片,validation_image_end是视频的结尾图片。
@@ -425,15 +183,11 @@ qfloat8会部分降低模型的性能,但可以节省更多的显存。如果
- 普通控制生视频(Canny、Pose、Depth等):
- 使用examples/cogvideox_fun/predict_v2v_control.py文件中修改control_video、validation_image_end、prompt、neg_prompt、guidance_scale和seed。
- control_video是控制生视频的控制视频,是使用Canny、Pose、Depth等算子提取后的视频。您可以使用以下视频运行演示:[演示视频](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1.1/pose.mp4)
- 而后运行examples/cogvideox_fun/predict_v2v_control.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos_v2v_control文件夹中。
- 步骤3:如果想结合自己训练的其他backbone与Lora,则看情况修改examples/{model_name}/predict_t2v.py中的examples/{model_name}/predict_i2v.py和lora_path。
- 而后运行examples/cogvideox_fun/predict_v2v_control.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos_control文件夹中。
- 步骤3:如果想结合自己训练的其他backbone与Lora,则在对应的`examples/{model_name}/predict_*.py`中设置`transformer_path`与`lora_path`(Wan2.2双 Transformer模型另有`transformer_high_path`与`lora_high_path`,分别对应high noise阶段)。
##### ii、多卡运行:
在使用多卡预测时请注意安装xfuser仓库,推荐安装xfuser==0.4.2和yunchang==0.6.2。
```
pip install xfuser==0.4.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
pip install yunchang==0.6.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
```
**ii、多卡运行**:
多卡并行推理所需的`xfuser`已列入1.3,推荐固定为`xfuser==0.4.2`。
请确保ulysses_degree和ring_degree的乘积等于使用的GPU数量。例如,如果您使用8个GPU,则可以设置ulysses_degree=2和ring_degree=4,也可以设置ulysses_degree=4和ring_degree=2。
@@ -448,21 +202,25 @@ ulysses_degree是在head进行切分后并行生成,ring_degree是在sequence
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
```
#### d、通过ui界面
### 2.4 通过ui界面
webui支持文生视频、图生视频、视频生视频和普通控制生视频(Canny、Pose、Depth等)。当前提供`app.py`的是CogVideoX-Fun、Wan2.1、Wan2.1-Fun、Wan2.2、Wan2.2-Fun(界面实现位于`videox_fun/ui/`),其余模型(包含图片模型)请使用python文件进行预测。以CogVideoX-Fun为例。
webui支持文生视频、图生视频、视频生视频和普通控制生视频(Canny、Pose、Depth等)。当前该库支持CogVideoX-Fun、Wan2.1和Wan2.1-Fun,在examples文件夹下用文件夹名以区分,不同模型支持的功能不同,请视具体情况予以区分。以CogVideoX-Fun为例。
- 步骤1:下载对应[权重](#model-zoo)放入models文件夹。
- 步骤1:下载对应[权重](#三已支持的模型)并按1.5放入models文件夹。
- 步骤2:运行examples/cogvideox_fun/app.py文件,进入gradio页面。
- 步骤3:根据页面选择生成模型,填入prompt、neg_prompt、guidance_scale和seed等,点击生成,等待生成结果,结果保存在sample文件夹中。
### 2. 模型训练
一个完整的模型训练链路应该包括数据预处理和Video DiT训练。不同模型的训练流程类似,数据格式也类似:
### 2.5 通过ComfyUI
具体查看[ComfyUI README](comfyui/README.md),我们的ComfyUI界面如下:
![workflow graph](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.jpg)
<h4 id="data-preprocess">a.数据预处理</h4>
我们给出了一个简单的demo通过图片数据训练lora模型,详情可以查看[wiki](https://github.com/aigc-apps/CogVideoX-Fun/wiki/Training-Lora)。
## 3. 模型训练
一个完整的模型训练链路应该包括数据预处理和Video DiT训练。不同模型的训练流程类似,数据格式也类似。
一个完整的长视频切分、清洗、描述的数据预处理链路可以参考video caption部分的[README](cogvideox/video_caption/README.md)进行。
<a id="data-preprocess"></a>
### 3.1 数据预处理
各模型的 LoRA 训练文档统一放在 `scripts/{model_name}/` 下,中文版以 `_zh-CN` 结尾,详情见[3.3 各模型训练文档](#33-各模型训练文档)。
一个完整的长视频切分、清洗、描述的数据预处理链路可以参考video caption部分的[README](videox_fun/video_caption/README_zh-CN.md)进行。
如果期望训练一个文生图视频的生成模型,您需要以这种格式排列数据集。
```
@@ -509,184 +267,254 @@ json_of_internal_datasets.json是一个标准的json文件。json中的file_path
.....
]
```
<h4 id="dit-train">b. Video DiT训练 </h4>
如果数据预处理时,数据的格式为相对路径,则进入scripts/{model_name}/train.sh进行如下设置。
<a id="dit-train"></a>
### 3.2 Video DiT训练
各模型的训练脚本与启动sh均位于`scripts/{model_name}/`下,sh的命名随任务而变,如`train.sh`、`train_lora.sh`、`train_control.sh`、`train_control_distill.sh`等,以目录内实际文件为准。
如果数据预处理时,数据的格式为相对路径,则进入对应的`scripts/{model_name}/train.sh`进行如下设置。
```
export DATASET_NAME="datasets/internal_datasets/"
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
```
如果数据的格式为绝对路径,则进入scripts/train.sh进行如下设置。
如果数据的格式为绝对路径,则在同一个脚本中设置如下(此时`DATASET_NAME`置空,不再拼接数据集目录前缀)。
```
export DATASET_NAME=""
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
```
最后运行scripts/train.sh。
最后运行对应的脚本。
```sh
sh scripts/train.sh
sh scripts/{model_name}/train.sh
```
关于一些参数的设置细节:
Wan2.1-Fun可以查看[Readme Train](scripts/wan2.1_fun/README_TRAIN.md)与[Readme Lora](scripts/wan2.1_fun/README_TRAIN_LORA.md)。
Wan2.1可以查看[Readme Train](scripts/wan2.1/README_TRAIN.md)与[Readme Lora](scripts/wan2.1/README_TRAIN_LORA.md)。
CogVideoX-Fun可以查看[Readme Train](scripts/cogvideox_fun/README_TRAIN.md)与[Readme Lora](scripts/cogvideox_fun/README_TRAIN_LORA.md)。
### 3.3 各模型训练文档
关于参数设置细节,各模型的训练文档统一放在`scripts/{model_name}/`下,`README_TRAIN*`为基线训练,`README_TRAIN_LORA*`为LoRA训练,`README_TRAIN_CONTROL*`为控制训练,中文版以`_zh-CN`结尾。常用模型如下:
# 模型地址
## 1.Wan2.2-Fun
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Wan2.2-Fun-A14B-InP | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP) | Wan2.2-Fun-14B文图生视频权重,以多分辨率训练,支持首尾图预测。 |
| Wan2.2-Fun-A14B-Control | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control)| Wan2.2-Fun-14B视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.2-Fun-A14B-Control-Camera | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera)| Wan2.2-Fun-14B相机镜头控制权重。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.2-VACE-Fun-A14B | 64.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B)| 以VACE方案训练的Wan2.2控制权重,基础模型为Wan2.2-T2V-A14B,支持不同的控制条件,如Canny、Depth、Pose、MLSD、轨迹控制等。支持通过主体指定生视频。支持多分辨率(512,768,1024)的视频预测,支持多分辨率(512,768,1024)的视频预测,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.2-Fun-5B-InP | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP) | Wan2.2-Fun-5B文图生视频权重,以121帧、每秒24帧进行训练支持首尾图预测。 |
| Wan2.2-Fun-5B-Control | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control)| Wan2.2-Fun-5B视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。以121帧、每秒24帧进行训练,支持多语言预测 |
| Wan2.2-Fun-5B-Control-Camera | 23.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera)| Wan2.2-Fun-5B相机镜头控制权重。以121帧、每秒24帧进行训练,支持多语言预测 |
## 2. Wan2.2
| 名称 | Hugging Face | Model Scope | 描述 |
| 模型 | 基线训练 | LoRA训练 | 其他 |
|--|--|--|--|
| Wan2.2-TI2V-5B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B) | 万象2.2-5B文生视频权重 |
| Wan2.2-T2V-A14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B) | 万象2.2-14B文生视频权重 |
| Wan2.2-I2V-A14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B) | 万象2.2-14B图生视频权重 |
| Wan2.1-Fun | [中文](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.1_fun/README_TRAIN.md) | [中文](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) | [Control 中文](scripts/wan2.1_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
| Wan2.2 | [中文](scripts/wan2.2/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.2/README_TRAIN.md) | [中文](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.2/README_TRAIN_LORA.md) | [蒸馏 中文](scripts/wan2.2/README_TRAIN_DISTILL_zh-CN.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V_zh-CN.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
| Wan2.2-Fun | [中文](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.2_fun/README_TRAIN.md) | [中文](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) | [Control LoRA 中文](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA_zh-CN.md) |
| CogVideoX-Fun | [中文](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) / [EN](scripts/cogvideox_fun/README_TRAIN.md) | [中文](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) | [Control 中文](scripts/cogvideox_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
| Qwen-Image | [中文](scripts/qwenimage/README_TRAIN_zh-CN.md) / [EN](scripts/qwenimage/README_TRAIN.md) | [中文](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/qwenimage/README_TRAIN_LORA.md) | [Edit 中文](scripts/qwenimage/README_TRAIN_EDIT_zh-CN.md) |
| Qwen-Image-2.1 | [中文](scripts/qwenimage21/README_TRAIN_zh-CN.md) / [EN](scripts/qwenimage21/README_TRAIN.md) | - | [Control 中文](scripts/qwenimage21_fun/README_TRAIN_zh-CN.md) / [EN](scripts/qwenimage21_fun/README_TRAIN.md) |
| Z-Image | [中文](scripts/z_image/README_TRAIN_zh-CN.md) / [EN](scripts/z_image/README_TRAIN.md) | [中文](scripts/z_image/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/z_image/README_TRAIN_LORA.md) | [GRPO LoRA 中文](scripts/z_image/README_TRAIN_GRPO_LORA_zh-CN.md) |
## 3. Wan2.1-Fun
其余模型(如HunyuanVideo、MiniMax-H3、Flux2-Fun、InfiniteTalk、LingBot等)同理,直接查看对应`scripts/{model_name}/`下的README即可。
V1.1:
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Wan2.1-Fun-V1.1-1.3B-InP | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP) | Wan2.1-Fun-V1.1-1.3B文图生视频权重,以多分辨率训练,支持首尾图预测。 |
| Wan2.1-Fun-V1.1-14B-InP | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP) | Wan2.1-Fun-V1.1-14B文图生视频权重,以多分辨率训练,支持首尾图预测。 |
| Wan2.1-Fun-V1.1-1.3B-Control | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control)| Wan2.1-Fun-V1.1-1.3B视频控制权重支持不同的控制条件,如Canny、Depth、Pose、MLSD等,支持参考图 + 控制条件进行控制,支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.1-Fun-V1.1-14B-Control | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control)| Wan2.1-Fun-V1.1-14B视视频控制权重支持不同的控制条件,如Canny、Depth、Pose、MLSD等,支持参考图 + 控制条件进行控制,支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.1-Fun-V1.1-1.3B-Control-Camera | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera)| Wan2.1-Fun-V1.1-1.3B相机镜头控制权重。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.1-Fun-V1.1-14B-Control-Camera | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera)| Wan2.1-Fun-V1.1-14B相机镜头控制权重。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
V1.0:
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Wan2.1-Fun-1.3B-InP | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP) | Wan2.1-Fun-1.3B文图生视频权重,以多分辨率训练,支持首尾图预测。 |
| Wan2.1-Fun-14B-InP | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP) | Wan2.1-Fun-14B文图生视频权重,以多分辨率训练,支持首尾图预测。 |
| Wan2.1-Fun-1.3B-Control | 19.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control)| Wan2.1-Fun-1.3B视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
| Wan2.1-Fun-14B-Control | 47.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control) | [😄Link](https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control)| Wan2.1-Fun-14B视频控制权重,支持不同的控制条件,如Canny、Depth、Pose、MLSD等,同时支持使用轨迹控制。支持多分辨率(512,768,1024)的视频预测,,以81帧、每秒16帧进行训练,支持多语言预测 |
# 三、已支持的模型
下表按模型系列汇总目前已支持的权重,视频模型与图片模型共用同一套推理与训练入口。每个系列一行,第四列为内嵌的四列表格,依次为权重、Hugging Face、ModelScope、对应说明;🤗 为 Hugging Face、🤖 为 ModelScope(国内网络推荐),`-` 表示该渠道确认无对应仓库或需登录授权。各模型训练文档见[3.3 各模型训练文档](#33-各模型训练文档)。
## 4. Wan2.1
| 名称 | Hugging Face | Model Scope | 描述 |
| 模型系列 | 模态 | 支持任务 | 权重 / 下载 / 说明 |
|--|--|--|--|
| Wan2.1-T2V-1.3B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B) | 万象2.1-1.3B文生视频权重 |
| Wan2.1-T2V-14B | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-T2V-14B) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B) | 万象2.1-14B文生视频权重 |
| Wan2.1-I2V-14B-480P | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P) | 万象2.1-14B-480P图生视频权重 |
| Wan2.1-I2V-14B-720P| [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | 万象2.1-14B-720P图生视频权重 |
| Wan2.2-Fun | 视频 | 本项目在Wan2.2上训练的系列,覆盖文生视频、图生视频、首尾图、控制生成、相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE双阶段文/图生视频,多分辨率训练、81帧16fps,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">14B控制生成,支持Canny、Depth、Pose、MLSD与轨迹控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">在14B Control基础上增加相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5B统一VAE文/图生视频,121帧24fps,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">5B控制生成,控制条件与14B一致</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">5B相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA,叠加在上述权重上使用</td></tr></table> |
| Wan2.2-VACE-Fun | 视频 | 本项目以VACE方案训练的系列,覆盖控制生成、主体参考 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">以Wan2.2-T2V-A14B为基础,支持Canny、Depth、Pose、MLSD、轨迹控制与主体参考生视频</td></tr></table> |
| Wan2.2 | 视频 | 万象官方权重,覆盖文生视频、图生视频、音频驱动、角色动画,可作为Wan2.2-Fun系列的训练基线 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">5B统一VAE,文生图生视频通用权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">语音驱动数字人</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">角色替换与动作迁移,仓库含多精度文件</td></tr></table> |
| Wan2.1-Fun V1.1 | 视频 | 本项目在Wan2.1上训练的V1.1版本,多分辨率(512/768/1024)、81帧16fps,覆盖文生视频、图生视频、首尾图、控制生成、相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B轻量文/图生视频,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">14B文/图生视频,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B控制生成,同时支持参考图+控制条件组合</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">14B控制生成,同时支持参考图+控制条件组合</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">14B相机运动控制</td></tr></table> |
| Wan2.1-Fun V1.0 | 视频 | 本项目在Wan2.1上训练的V1.0版本,能力与V1.1相同但无相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的1.3B文/图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的14B文/图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的1.3B控制生成</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的14B控制生成</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
| Wan2.1 | 视频 | 万象官方权重,覆盖文生视频、图生视频、控制生成,可作为Wan2.1-Fun系列的训练基线 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">720P图生视频,是FantasyTalking的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE控制与主体参考</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE控制与主体参考</td></tr></table> |
| Self-Forcing / Causal-Forcing / Flex-Forcing | 视频 | 自回归蒸馏方案,覆盖流式生成、交互式生成与分块注意力 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">官方发布的蒸馏权重,配合Wan2.1-T2V使用;也可由`scripts/wan2.1_self_forcing`与`scripts/wan2.1_causal_forcing`自行训练得到;Flex-Forcing(分块因果/双向注意力)权重由`scripts/wan2.1_flex_forcing`训练产出</td></tr></table> |
| TurboWan / TurboDiffusion | 视频 | TurboDiffusion方案公开发布的少步蒸馏权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频蒸馏权重,官方以.pth发布,仓库另含量化版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B图生视频蒸馏权重,仓库含low/high两档噪声模型(另含量化版),放入Personalized_Model后按预测脚本的transformer_path/transformer_high_path引用</td></tr></table> |
| CogVideoX-Fun V1.5 | 视频 | V1.5官方权重,多分辨率(512/768/1024)、85帧8fps,覆盖图生视频、奖励对齐 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5b图生视频权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
| CogVideoX-Fun V1.1 | 视频 | V1.1官方权重,多分辨率(512/768/1024/1280)、49帧8fps,覆盖图生视频、姿态控制、控制生成、奖励对齐 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">2b图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5b图生视频,添加Noise,运动幅度大于V1.0</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">2b姿态控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">5b姿态控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">2b控制生成,支持Canny、Depth、Pose、MLSD</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">5b控制生成,支持Canny、Depth、Pose、MLSD</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
| CogVideoX-Fun V1.0 | 视频 | 旧版权重,仍以49帧8fps训练,已被V1.1/V1.5取代 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的2b图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的5b图生视频</td></tr></table> |
| HunyuanVideo | 视频 | 官方diffusers格式权重,本项目直接支持预测与LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
| MiniMax-H3 | 视频 | 官方视频生成权重与本项目训练的ControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">官方基线权重,仓库含多种精度与组件,可按需下载</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet,支持多种控制条件与轨迹控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet(2.0版),支持多种控制条件、轨迹控制与inpaint权重</td></tr></table> |
| TaoMate-H3 | 视频+音频 | 基于MiniMax-H3的官方流式音视频生成适配器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">官方rank 128适配器(step-3000 EMA),内置3步蒸馏采样调度,支持流式语音驱动生成;需搭配MiniMax-H3基座权重使用</td></tr></table> |
| LTX-2 | 视频+音频 | 官方DiT音视频联合生成权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">音视频联合生成的官方权重,仓库含多种精度</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">2.3版本需使用社区转换的diffusers格式权重,官方原始权重见<a href="https://huggingface.co/Lightricks/LTX-2.3">Lightricks/LTX-2.3</a></td></tr></table> |
| LongCat-Video | 视频 | 官方长视频生成权重,支持LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">文/图生长视频基线</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">数字人权重</td></tr></table> |
| FantasyTalking | 音频驱动视频 | 音频条件增量权重,需搭配基础视频权重与音频编码器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
| InfiniteTalk | 音频驱动视频 | 音频条件增量权重,需搭配基础视频权重与音频编码器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-480P使用,仓库含多个版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">中文音频编码器</td></tr></table> |
| FlashHead | 音频驱动视频 | 官方头部动作数字人权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">语音驱动头部数字人,同样需要wav2vec音频编码器</td></tr></table> |
| MOVA | 视频+音频 | 官方MOVA权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频与音视频联合生成</td></tr></table> |
| LingBot | 视频 | 相机可控世界模型,目录结构与Wan2.2-I2V-A14B一致 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">相机控制基线权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA,搭配Qwen3.6-27B生成结构化caption</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B稠密版视频生成权重,1-2卡即可训练</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE(3B激活)视频生成权重,训练建议8×80GB及以上</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">蒸馏少步世界模型(transformer共16个分片),VAE/T5复用lingbot-world-base-cam,推理需使用Flow_Unipc采样器</td></tr></table> |
| Phantom | 视频 | 多主体参考生视频的增量权重,基于Wan2.1-T2V | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B版,官方以.pth发布,放入Personalized_Model后按预测脚本的transformer_path引用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B版,官方以分片safetensors发布</td></tr></table> |
| Qwen-Image | 图片 | 官方文生图与图像编辑权重,支持基线与LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">文生图更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">图像图层分解权重,可将图像拆分为多个可编辑的RGBA图层</td></tr></table> |
| Qwen-Image-2.1 | 图片 | 官方新一代文生图权重,单流block-causal结构,支持前缀KV cache | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">单流block-causal结构,支持全参数训练;前缀KV cache可加速推理</td></tr></table> |
| Qwen-Image ControlNet | 图片 | 图片控制生成,支持Canny、Depth、Pose、MLSD、Scribble | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2.1-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目为 Qwen-Image-2.1 训练的 ControlNet-Union,支持 Canny、Depth、Pose、MLSD 等控制条件与图像修复(inpaint)</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">InstantX提供的同类型ControlNet</td></tr></table> |
| Z-Image | 图片 | 官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
| Z-Image-Fun | 图片 | 本项目在Z-Image上训练的ControlNet与蒸馏LoRA,控制条件支持Canny、Depth、Pose、MLSD、Scribble、Gray | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">基于基础版的ControlNet,2.1版层数更多、训练更充分</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">基于Turbo的ControlNet</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">基于Turbo的2.1版ControlNet,仓库含多精度文件</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">同时蒸馏步数与CFG,推理仅需8步</td></tr></table> |
| Flux | 图片 | 官方FLUX.1/FLUX.2权重与本项目训练的ControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目为FLUX.2-dev训练的ControlNet,支持Canny、Depth、Pose、MLSD等</td></tr></table> |
| ERNIE-Image | 图片 | 百度官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">单流DiT文生图,Hugging Face为baidu组织、ModelScope为PaddlePaddle组织</td></tr></table> |
| Lens | 图片 | 微软官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">3.8B文生图,仓库内含GPT-OSS文本编码器;Hugging Face侧无公开下载仓库,请从ModelScope获取</td></tr></table> |
| 辅助模型 | - | 非生成模型,服务于奖励对齐、数据打标与快速解码 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播使用的打分模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">视频打标流程使用的多模态编码器</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoder(约20MB),与Wan2.1/Wan2.2 VAE共享同一latent空间,解码速度约为完整VAE的100倍,用于快速预览与低显存生成;权重来自<a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a></td></tr></table> |
## 5. FantasyTalking
> 补充说明:
> - 音频驱动与参考类模型(FantasyTalking、InfiniteTalk、Phantom、TaoMate-H3)本身只是增量权重,必须同时下载表中对应的基础视频权重与音频编码器。
> - TurboDiffusion方案已公开发布TurboWan系列蒸馏权重(见上表);Flex-Forcing、PDD等其余蒸馏方案没有公开发布的权重,按`scripts/{model_name}/README_TRAIN*.md`训练后即可得到,可直接填入预测文件中的`transformer_path`。
> - 权重名与`models/Diffusion_Transformer/`下的文件夹名一一对应;同一系列内各权重互不通用,需按预测任务选择,若某个权重未在此列出,说明它由本项目训练产出或需从上游官方仓库获取。
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Wan2.1-I2V-14B-720P | - | [🤗Link](https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P) | [😄Link](https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P) | 万象2.1-14B-720P图生视频权重 |
| Wav2Vec | - | [🤗Link](https://huggingface.co/facebook/wav2vec2-base-960h) | [😄Link](https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h) | Wav2Vec模型,请放在Wan2.1-I2V-14B-720P文件夹下,命名为audio_encoder |
| FantasyTalking model | - | [🤗Link](https://huggingface.co/acvlab/FantasyTalking/) | [😄Link](https://www.modelscope.cn/models/amap_cvlab/FantasyTalking/) | 官方Audio Condition的权重。 |
# 四、视频作品
## 6. Qwen-Image
图生视频:
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Qwen-Image | [🤗Link](https://huggingface.co/Qwen/Qwen-Image) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image) | Qwen-Image官方权重 |
| Qwen-Image-Edit | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit) | Qwen-Image-Edit官方权重 |
| Qwen-Image-Edit-2509 | [🤗Link](https://huggingface.co/Qwen/Qwen-Image-Edit-2509) | [😄Link](https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509) | Qwen-Image-Edit-2509官方权重 |
## 7. Qwen-Image-Fun
| 名称 | 存储 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Qwen-Image-2512-Fun-Controlnet-Union | - | [🤗链接](https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union) | [😄链接](https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union) | Qwen-Image-2512的ControlNet权重,支持多种控制条件,如Canny、Depth、Pose、MLSD、Scribble等。 |
## 8. Z-Image
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Z-Image-Turbo | [🤗Link](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [😄Link](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo) | Z-Image-Turbo官方权重 |
## 9. Z-Image-Fun
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| Z-Image-Turbo-Fun-Controlnet-Union | - | [🤗链接](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union) | [😄链接](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union) | Z-Image-Turbo 的 ControlNet 权重,支持 Canny、Depth、Pose、MLSD 等多种控制条件。 |
| Z-Image-Turbo-Fun-Controlnet-Union-2.1 | - | [🤗链接](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | [😄链接](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | Z-Image-Turbo 的 ControlNet 权重,相比第一版在更多层进行添加,也训练了更长时间,支持 Canny、Depth、Pose、MLSD 等多种控制条件。 |
## 10. Flux
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| FLUX.1-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.1-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev) | FLUX.1-dev官方权重 |
| FLUX.2-dev | [🤗Link](https://huggingface.co/black-forest-labs/FLUX.2-dev) | [😄Link](https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev) | FLUX.2-dev官方权重 |
## 11. Flux-Fun
| 名称 | 存储 | Hugging Face | 魔搭社区(ModelScope) | 描述 |
|--|--|--|--|--|
| Flux.2-dev-Fun-Controlnet-Union | - | [🤗链接](https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union) | [😄链接](https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union) | Flux.2-dev 的 ControlNet 权重,支持 Canny、Depth、Pose、MLSD 等多种控制条件。 |
## 12. HunyuanVideo
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| HunyuanVideo | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo) | - | HunyuanVideo-diffusers权重 |
| HunyuanVideo-I2V | [🤗Link](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | - | HunyuanVideo-I2V-diffusers权重 |
## 13. CogVideoX-Fun
V1.5:
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| CogVideoX-Fun-V1.5-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP) | 官方的图生视频权重。支持多分辨率(512,768,1024)的视频预测,以85帧、每秒8帧进行训练 |
| CogVideoX-Fun-V1.5-Reward-LoRAs | - | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs) | 官方的奖励反向传播技术模型,优化CogVideoX-Fun-V1.5生成的视频,使其更好地符合人类偏好。 |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
V1.1:
通用控制视频 + 参考图像:
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| CogVideoX-Fun-V1.1-2b-InP | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP) | 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
| CogVideoX-Fun-V1.1-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP) | 官方的图生视频权重。添加了Noise,运动幅度相比于V1.0更大。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
| CogVideoX-Fun-V1.1-2b-Pose | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose) | 官方的姿态控制生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
| CogVideoX-Fun-V1.1-2b-Control | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control) | 官方的控制生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练。支持不同的控制条件,如Canny、Depth、Pose、MLSD等 |
| CogVideoX-Fun-V1.1-5b-Pose | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose) | 官方的姿态控制生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
| CogVideoX-Fun-V1.1-5b-Control | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control) | 官方的控制生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练。支持不同的控制条件,如Canny、Depth、Pose、MLSD等 |
| CogVideoX-Fun-V1.1-Reward-LoRAs | - | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs) | 官方的奖励反向传播技术模型,优化CogVideoX-Fun-V1.1生成的视频,使其更好地符合人类偏好。 |
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
参考图像
</td>
<td>
控制视频
</td>
<td>
Wan2.1-Fun-V1.1-14B-Control
</td>
<td>
Wan2.1-Fun-V1.1-1.3B-Control
</td>
</tr>
<tr>
<td>
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload="none"></image>
</td>
<td>
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<details>
<summary>(Obsolete) V1.0:</summary>
| 名称 | 存储空间 | Hugging Face | Model Scope | 描述 |
|--|--|--|--|--|
| CogVideoX-Fun-2b-InP | 13.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP) | 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
| CogVideoX-Fun-5b-InP | 20.0 GB | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP) | [😄Link](https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP) | 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以49帧、每秒8帧进行训练 |
</details>
通用控制视频(Canny、Pose、Depth 等)与轨迹控制:
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
<tr>
<td>
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload="none"></video>
</td>
</tr>
<tr>
<td>
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload="none"></video>
</td>
<td>
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload="none"></video>
</td>
</tr>
</table>
# 五、参考文献
本节列出[已支持的模型](#三已支持的模型)中各模型系列的官方仓库,以及本项目在实现与流程中参考的代码来源,感谢这些开源工作。
# 参考文献
- CogVideo: https://github.com/THUDM/CogVideo/
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
- Wan2.2: https://github.com/Wan-Video/Wan2.2/
- Diffusers: https://github.com/huggingface/diffusers
- HunyuanVideo: https://github.com/Tencent-Hunyuan/HunyuanVideo
- HunyuanVideo-I2V: https://github.com/Tencent-Hunyuan/HunyuanVideo-I2V
- MiniMax-H3: https://github.com/MiniMax-AI/MiniMax-H3
- LTX-Video: https://github.com/Lightricks/LTX-Video
- LTX-2: https://github.com/Lightricks/LTX-2
- LongCat-Video: https://github.com/meituan-longcat/LongCat-Video
- FantasyTalking: https://github.com/Fantasy-AMAP/fantasy-talking
- InfiniteTalk: https://github.com/MeiGen-AI/InfiniteTalk
- FlashHead: https://github.com/Soul-AILab/SoulX-FlashHead
- MOVA: https://github.com/OpenMOSS/MOVA
- LingBot-Video: https://github.com/Robbyant/lingbot-video
- LingBot-World: https://github.com/Robbyant/lingbot-world
- Phantom: https://github.com/Phantom-video/Phantom
- Qwen-Image: https://github.com/QwenLM/Qwen-Image
- Self-Forcing: https://github.com/guandeh17/Self-Forcing
- Z-Image: https://github.com/Tongyi-MAI/Z-Image
- Flux: https://github.com/black-forest-labs/flux
- Flux2: https://github.com/black-forest-labs/flux2
- HunyuanVideo: https://github.com/Tencent-Hunyuan/HunyuanVideo
- ERNIE-Image: https://github.com/baidu/ernie-image
- Lens: https://www.microsoft.com/en-us/research/publication/lens-rethinking-training-efficiency-for-foundational-text-to-image-models/
- VACE: https://github.com/ali-vilab/VACE
- CameraCtrl: https://github.com/hehao13/CameraCtrl
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
- DWPose: https://github.com/IDEA-Research/DWPose
- MiDaS: https://github.com/isl-org/MiDaS
- Self-Forcing: https://github.com/guandeh17/Self-Forcing
- Causal-Forcing: https://github.com/thu-ml/Causal-Forcing
- TurboDiffusion: https://github.com/thu-ml/TurboDiffusion
- TAEHV: https://github.com/madebyollin/taehv
- HPS v2: https://github.com/tgxs002/HPSv2
- HPSv3: https://github.com/MizzenAI/HPSv3
- MPS: https://github.com/Kwai-Kolors/MPS
- Qwen2-VL: https://github.com/QwenLM/Qwen2-VL
- AnimateDiff: https://github.com/guoyww/AnimateDiff
- ComfyUI-KJNodes: https://github.com/kijai/ComfyUI-KJNodes
- ComfyUI-EasyAnimateWrapper: https://github.com/kijai/ComfyUI-EasyAnimateWrapper
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
- CameraCtrl: https://github.com/hehao13/CameraCtrl
- Diffusers: https://github.com/huggingface/diffusers
# 许可证
# 六、引用
如果您在研究或项目中使用了 VideoX-Fun,请按以下格式引用:
```bibtex
@misc{aigc_apps_VideoX_Fun_2026,
author = {aigc-apps},
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
year = {2026},
publisher = {GitHub},
url = {https://github.com/aigc-apps/VideoX-Fun}
}
```
# 七、限制与风险
- 生成的视频可能存在伪影或质量问题,尤其在复杂场景中。
- 模型在处理精细细节、文字渲染或特定艺术风格时可能有困难。
- 性能因输入提示词质量、分辨率等参数而异。
- 该技术可能被滥用于创建误导性内容(如深度伪造)。用户需对道德使用负责。
- 模型可能反映训练数据中存在的偏见。
- 用户在使用真人图片或视频时应尊重隐私和版权。
我们鼓励负责任地使用该技术,并建议在生产环境中实施安全措施。
# 八、许可证
本项目采用 [Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE).
CogVideoX-2B 模型 (包括其对应的Transformers模块,VAE模块) 根据 [Apache 2.0 协议](LICENSE) 许可证发布。
BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 349 KiB

Binary file not shown.
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 150 KiB

Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,9 @@
# This file intentionally exists (empty) to make `midas` a regular package.
#
# Without it, `midas` is only a namespace-package portion, and Python's import
# machinery lets ANY regular package named `midas` found elsewhere on sys.path
# (e.g. comfyui_controlnet_aux's .../src/custom_controlnet_aux/midas/) win the
# resolution, even though torch.hub.load inserts this midas_repo dir at
# sys.path[0]. That collision produces:
# ImportError: attempted relative import beyond top-level package
# See https://github.com/aigc-apps/VideoX-Fun/issues/502
+2 -2
View File
@@ -15,8 +15,7 @@ from diffusers import EulerDiscreteScheduler
from einops import rearrange
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
@@ -121,6 +120,7 @@ class LoadCogVideoXFunModel:
transformer = CogVideoXTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=torch.float8_e4m3fn if GPU_memory_mode == "model_cpu_offload_and_qfloat8" else weight_dtype,
).to(weight_dtype)
# Update pbar
+1 -1
View File
@@ -37,7 +37,7 @@ For chunked loading, it is recommended to directly download the FLUX.2-dev weigh
│ ├── 📂 Fun_Models/
│ │ └── flux2_tokenizer/
│ └── 📂 model_patches/
│ └── FLUX.2-dev-Fun-Controlnet-Union.safetensors
│ └── FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors
```
### 2. Preprocessing Weights (Optional)
+5 -6
View File
@@ -26,8 +26,7 @@ else:
from diffusers.models.modeling_utils import \
load_model_dict_into_meta
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKLFlux2,
Flux2ControlTransformer2DModel,
Flux2Transformer2DModel,
@@ -598,10 +597,10 @@ class LoadFlux2TextEncoderModel:
[os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "models/Diffusion_Transformer")] # Possible folder names to check
try:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="flux2_tokenizer")
except:
except Exception:
try:
tokenizer_path = os.path.join(search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="FLUX.2-dev"), "tokenizer")
except:
except Exception:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Mistral-Nemo-Instruct-2407")
tokenizer = PixtralProcessor.from_pretrained(tokenizer_path)
@@ -875,7 +874,7 @@ class LoadFlux2ControlNetInPipeline:
),
"model_name": (
folder_paths.get_filename_list("model_patches"),
{"default": "FLUX.2-dev-Fun-Controlnet-Union.safetensors", },
{"default": "FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors", },
),
"funmodels": ("FunModels",),
},
@@ -1017,7 +1016,7 @@ class LoadFlux2ControlNetInModel:
),
"model_name": (
folder_paths.get_filename_list("model_patches"),
{"default": "FLUX.2-dev-Fun-Controlnet-Union.safetensors", },
{"default": "FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors", },
),
"transformer": ("TransformerModel",),
},
@@ -194,7 +194,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
@@ -316,7 +316,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
@@ -194,7 +194,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
@@ -186,7 +186,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
@@ -149,7 +149,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
@@ -122,7 +122,7 @@
},
"widgets_values": [
"flux2/flux2_control.yaml",
"FLUX.2-dev-Fun-Controlnet-Union.safetensors"
"FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
]
},
{
+5 -6
View File
@@ -25,8 +25,7 @@ else:
from diffusers.models.modeling_utils import \
load_model_dict_into_meta
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKLQwenImage, Qwen2_5_VLConfig,
Qwen2_5_VLForConditionalGeneration,
Qwen2Tokenizer, Qwen2VLProcessor,
@@ -476,10 +475,10 @@ class LoadQwenImageTextEncoderModel:
[os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "models/Diffusion_Transformer")] # Possible folder names to check
try:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="qwen2_tokenizer")
except:
except Exception:
try:
tokenizer_path = os.path.join(search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Qwen-Image"), "tokenizer")
except:
except Exception:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Qwen2.5-VL-7B-Instruct")
tokenizer = Qwen2Tokenizer.from_pretrained(tokenizer_path)
@@ -504,10 +503,10 @@ class LoadQwenImageProcessor:
[os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "models/Diffusion_Transformer")] # Possible folder names to check
try:
processor_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="qwen2_processor")
except:
except Exception:
try:
processor_path = os.path.join(search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Qwen-Image-Edit"), "processor")
except:
except Exception:
processor_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Qwen2.5-VL-7B-Instruct")
# Get processor
+1 -2
View File
@@ -17,8 +17,7 @@ from einops import rearrange
from omegaconf import OmegaConf
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, CLIPModel, WanT5EncoderModel,
WanTransformer3DModel)
+2 -3
View File
@@ -16,9 +16,8 @@ from einops import rearrange
from omegaconf import OmegaConf
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data.dataset_image_video import process_pose_params
from ...videox_fun.data import (ASPECT_RATIO_512, get_closest_ratio,
process_pose_params)
from ...videox_fun.models import (AutoencoderKLWan, AutoTokenizer, CLIPModel,
WanT5EncoderModel, WanTransformer3DModel)
from ...videox_fun.models.cache_utils import get_teacache_coefficients
+1 -2
View File
@@ -17,8 +17,7 @@ from einops import rearrange
from omegaconf import OmegaConf
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, CLIPModel,
Wan2_2Transformer3DModel, WanT5EncoderModel)
+2 -3
View File
@@ -16,9 +16,8 @@ from einops import rearrange
from omegaconf import OmegaConf
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data.dataset_image_video import process_pose_params
from ...videox_fun.data import (ASPECT_RATIO_512, get_closest_ratio,
process_pose_params)
from ...videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, CLIPModel,
Wan2_2Transformer3DModel, WanT5EncoderModel)
+2 -3
View File
@@ -17,9 +17,8 @@ from einops import rearrange
from omegaconf import OmegaConf
from PIL import Image
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data.dataset_image_video import process_pose_params
from ...videox_fun.data import (ASPECT_RATIO_512, get_closest_ratio,
process_pose_params)
from ...videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, CLIPModel,
VaceWanTransformer3DModel, WanT5EncoderModel)
+30 -11
View File
@@ -1,4 +1,4 @@
# Z-Image-Turbo Model Setup Guide
# Z-Image Model Setup Guide
## a. Model Links and Storage Locations
@@ -13,7 +13,7 @@ For chunked loading, it is recommended to directly download the Z-Image weights
| Component | File Name |
|-----------|-----------|
| Text Encoder | [`qwen_3_4b.safetensors`](https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/text_encoders/qwen_3_4b.safetensors) |
| Diffusion Model | [`z_image_turbo_bf16.safetensors`](https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/diffusion_models/z_image_turbo_bf16.safetensors) |
| Diffusion Model | [`z_image_turbo_bf16.safetensors`](https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/diffusion_models/z_image_turbo_bf16.safetensors) and [`z_image_bf16.safetensors`](https://huggingface.co/Comfy-Org/z_image/resolve/main/split_files/diffusion_models/z_image_bf16.safetensors) |
| VAE | [`ae.safetensors`](https://huggingface.co/Comfy-Org/z_image_turbo/resolve/main/split_files/vae/ae.safetensors) |
| tokenizer(Qwen3-4B) | [`tokenizer`](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo/tree/main/tokenizer) |
@@ -23,6 +23,7 @@ For chunked loading, it is recommended to directly download the Z-Image weights
|------|--------------|-------------|-------------|
| Z-Image-Turbo-Fun-Controlnet-Union | [🤗Link](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union) | [😄Link](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union) | ControlNet weights for Z-Image-Turbo, supporting multiple control conditions including Canny, Depth, Pose, MLSD, etc. |
| Z-Image-Turbo-Fun-Controlnet-Union-2.1 | [🤗Link](https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | [😄Link](https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1) | Upgraded ControlNet weights for Z-Image-Turbo with additions at more layers and longer training time, supporting multiple control conditions including Canny, Depth, Pose, MLSD, etc. |
| Z-Image-Fun-Controlnet-Union-2.1 | [🤗Link](https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1) | [😄Link](https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1) | Upgraded ControlNet weights for Z-Image with additions at more layers and longer training time, supporting multiple control conditions including Canny, Depth, Pose, MLSD, Scribble, Hed and Gray. |
**Storage Location:**
@@ -74,8 +75,9 @@ If you prefer full model loading, you can directly download the diffusers weight
| Name | Hugging Face | Model Scope | Description |
|------|--------------|-------------|-------------|
| Z-Image-Turbo | [🤗Link](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | [😄Link](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo) | Official full weights for Z-Image-Turbo |
| Z-Image | [🤗Link](https://huggingface.co/Tongyi-MAI/Z-Image) | [😄Link](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image) | Official full weights for Z-Image |
For full model loading, use the diffusers version of Z-Image Turbo and place the model in `ComfyUI/models/Fun_Models/`.
For full model loading, use the diffusers version of Z-Image and place the model in `ComfyUI/models/Fun_Models/`.
**Storage Location:**
@@ -83,6 +85,7 @@ For full model loading, use the diffusers version of Z-Image Turbo and place the
📂 ComfyUI/
├── 📂 models/
│ └── 📂 Fun_Models/
│ ├── 📂 Z-Image
│ └── 📂 Z-Image-Turbo
```
@@ -90,20 +93,36 @@ For full model loading, use the diffusers version of Z-Image Turbo and place the
### 1. Chunked Loading (Recommended)
[Z Image Turbo Text to Image](v1/z_image_chunked_loading_workflow_t2i.json)
[Z Image Text to Image](v1/z_image_chunked_loading_workflow_t2i.json)
[Z Image Turbo Text to Image and Control](v1/z_image_chunked_loading_workflow_t2i_control.json)
[Z Image Text to Image and Control](v1/z_image_chunked_loading_workflow_t2i_control.json)
[Z Image Turbo Text to Image and Control with Pose Detect](v1/z_image_chunked_loading_workflow_t2i_control_pose_process.json)
[Z Image Text to Image and Control with Pose Detect](v1/z_image_chunked_loading_workflow_t2i_control_pose_process.json)
[Z Image Turbo Text to Image and Control with Depth Detect](v1/z_image_chunked_loading_workflow_t2i_control_depth_process.json)
[Z Image Text to Image and Control with Depth Detect](v1/z_image_chunked_loading_workflow_t2i_control_depth_process.json)
[Z Image Turbo Text to Image and Control with Canny Detect](v1/z_image_chunked_loading_workflow_t2i_control_canny_process.json)
[Z Image Text to Image and Control with Canny Detect](v1/z_image_chunked_loading_workflow_t2i_control_canny_process.json)
[Z Image Turbo Image to Image with Inpaint](v1/z_image_chunked_loading_workflow_i2i_inpaint.json)
[Z Image Image to Image with Inpaint](v1/z_image_chunked_loading_workflow_i2i_inpaint.json)
[Z Image Turbo Text to Image](v1/z_image_turbo_chunked_loading_workflow_t2i.json)
[Z Image Turbo Text to Image and Control](v1/z_image_turbo_chunked_loading_workflow_t2i_control.json)
[Z Image Turbo Text to Image and Control with Pose Detect](v1/z_image_turbo_chunked_loading_workflow_t2i_control_pose_process.json)
[Z Image Turbo Text to Image and Control with Depth Detect](v1/z_image_turbo_chunked_loading_workflow_t2i_control_depth_process.json)
[Z Image Turbo Text to Image and Control with Canny Detect](v1/z_image_turbo_chunked_loading_workflow_t2i_control_canny_process.json)
[Z Image Turbo Image to Image with Inpaint](v1/z_image_turbo_chunked_loading_workflow_i2i_inpaint.json)
### 2. Full Model Loading (Optional)
[Z Image Turbo Text to Image](v1/z_image_workflow_t2i.json)
[Z Image Text to Image](v1/z_image_workflow_t2i.json)
[Z Image Turbo Text to Image and Control](v1/z_image_workflow_t2i_control.json)
[Z Image Text to Image and Control](v1/z_image_workflow_t2i_control.json)
[Z Image Turbo Text to Image](v1/z_image_turbo_workflow_t2i.json)
[Z Image Turbo Text to Image and Control](v1/z_image_turbo_workflow_t2i_control.json)
+31 -14
View File
@@ -18,8 +18,7 @@ from omegaconf import OmegaConf
from PIL import Image
from safetensors.torch import load_file
from ...videox_fun.data.bucket_sampler import (ASPECT_RATIO_512,
get_closest_ratio)
from ...videox_fun.data import ASPECT_RATIO_512, get_closest_ratio
from ...videox_fun.models import (AutoencoderKL, AutoTokenizer,
Qwen2VLProcessor, Qwen3Config,
Qwen3ForCausalLM,
@@ -27,6 +26,9 @@ from ...videox_fun.models import (AutoencoderKL, AutoTokenizer,
ZImageTransformer2DModel)
from ...videox_fun.models.cache_utils import get_teacache_coefficients
from ...videox_fun.pipeline import ZImageControlPipeline, ZImagePipeline
from ...videox_fun.utils import (register_auto_device_hook,
safe_enable_group_offload,
safe_remove_group_offloading)
from ...videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from ...videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from ...videox_fun.utils.fp8_optimization import (
@@ -432,10 +434,10 @@ class LoadZImageTextEncoderModel:
[os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "models/Diffusion_Transformer")] # Possible folder names to check
try:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="qwen3_tokenizer")
except:
except Exception:
try:
tokenizer_path = os.path.join(search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Z-Image-Turbo"), "tokenizer")
except:
except Exception:
tokenizer_path = search_sub_dir_in_possible_folders(possible_folders, sub_dir_name="Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained(tokenizer_path)
@@ -453,7 +455,9 @@ class CombineZImagePipeline:
"tokenizer": ("Tokenizer",),
"model_name": ("STRING",),
"GPU_memory_mode":(
["model_full_load", "model_full_load_and_qfloat8","model_cpu_offload", "model_cpu_offload_and_qfloat8", "sequential_cpu_offload"],
[
"model_full_load", "model_full_load_and_qfloat8", "model_cpu_offload",
"model_cpu_offload_and_qfloat8", "model_group_offload", "sequential_cpu_offload"],
{
"default": "model_cpu_offload",
}
@@ -499,20 +503,23 @@ class CombineZImagePipeline:
)
pipeline.remove_all_hooks()
safe_remove_group_offloading(pipeline)
undo_convert_weight_dtype_wrapper(transformer)
pipeline.to(device=offload_device)
transformer = transformer.to(weight_dtype)
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device=offload_device, offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
@@ -536,14 +543,17 @@ class LoadZImageModel:
"required": {
"model": (
[
"Z-Image-Turbo"
"Z-Image-Turbo",
"Z-Image"
],
{
"default": 'Z-Image-Turbo',
}
),
"GPU_memory_mode":(
["model_full_load", "model_full_load_and_qfloat8","model_cpu_offload", "model_cpu_offload_and_qfloat8", "sequential_cpu_offload"],
[
"model_full_load", "model_full_load_and_qfloat8", "model_cpu_offload",
"model_cpu_offload_and_qfloat8", "model_group_offload", "sequential_cpu_offload"],
{
"default": "model_cpu_offload",
}
@@ -637,14 +647,17 @@ class LoadZImageModel:
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device=offload_device, offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
@@ -733,6 +746,7 @@ class LoadZImageControlNetInPipeline:
# Remove hooks
funmodels["pipeline"].remove_all_hooks()
safe_remove_group_offloading(funmodels["pipeline"])
# Load config
config_path = f"{script_directory}/config/{config}"
@@ -800,14 +814,17 @@ class LoadZImageControlNetInPipeline:
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device=offload_device, offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(control_transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(control_transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(control_transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(control_transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_model_weight_to_float8(control_transformer, exclude_module_name=["x_pad_token", "cap_pad_token"], device=device)
convert_weight_dtype_wrapper(control_transformer, weight_dtype)
pipeline.to(device=device)
else:
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -223,7 +223,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -261,7 +261,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -300,7 +300,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1.safetensors"
]
},
{
@@ -369,11 +369,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -223,7 +223,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -261,7 +261,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -331,11 +331,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -535,7 +535,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-lite-2601-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1-lite.safetensors"
]
}
],
@@ -104,8 +104,8 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3
]
@@ -145,7 +145,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -245,7 +245,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -350,7 +350,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -255,7 +255,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -357,7 +357,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -396,7 +396,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1.safetensors"
]
},
{
@@ -465,11 +465,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
}
],
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -223,7 +223,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -261,7 +261,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -300,7 +300,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1.safetensors"
]
},
{
@@ -369,11 +369,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -223,7 +223,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -261,7 +261,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -300,7 +300,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1.safetensors"
]
},
{
@@ -369,11 +369,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -290,7 +290,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -360,11 +360,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -431,7 +431,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -469,7 +469,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-lite-2601-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1-lite.safetensors"
]
}
],
@@ -131,7 +131,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -223,7 +223,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"model_group_offload"
]
},
{
@@ -261,7 +261,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -300,7 +300,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Union-2.1.safetensors"
]
},
{
@@ -369,11 +369,11 @@
1568,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -255,7 +255,7 @@
},
"widgets_values": [
"",
"model_cpu_offload"
"sequential_cpu_offload"
]
},
{
@@ -357,7 +357,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -396,7 +396,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Tile-2.1-8steps.safetensors"
"Z-Image-Fun-Controlnet-Tile-2.1.safetensors"
]
},
{
@@ -465,11 +465,11 @@
2416,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
}
],
@@ -290,7 +290,7 @@
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"z_image_bf16.safetensors",
"bf16"
]
},
@@ -360,11 +360,11 @@
2416,
43,
"fixed",
8,
0,
25,
4.5,
"Flow",
3,
0.8
0.85
]
},
{
@@ -402,7 +402,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Tile-2.1-lite-2601-8steps.safetensors"
"Z-Image-Fun-Controlnet-Tile-2.1-lite.safetensors"
]
},
{
@@ -0,0 +1,703 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 107,
"last_link_id": 113,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": 110
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": 112
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 104,
"type": "PreviewImage",
"pos": [
911.4377073728218,
418.9988584275326
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 113
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 107,
"type": "MaskToImage",
"pos": [
658.3044275669289,
511.8717365608471
],
"size": [
140,
26
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "mask",
"type": "MASK",
"link": 111
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
112,
113
]
}
],
"properties": {
"Node name for S&R": "MaskToImage"
}
},
{
"id": 100,
"type": "LoadImage",
"pos": [
323.5443519405259,
489.8702544908114
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
110
]
},
{
"name": "MASK",
"type": "MASK",
"links": [
111
]
}
],
"properties": {
"Node name for S&R": "LoadImage",
"image": "clipspace/clipspace-painted-masked-1766731857414.png [input]"
},
"widgets_values": [
"clipspace/clipspace-painted-masked-1766731857414.png [input]",
"image"
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
],
[
110,
100,
0,
99,
4,
"IMAGE"
],
[
111,
100,
1,
107,
0,
"MASK"
],
[
112,
107,
0,
99,
5,
"IMAGE"
],
[
113,
107,
0,
104,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7772383863288174,
"offset": [
-164.75871381690226,
267.25116947501067
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "93aa7b2530dccd1e91c625eee439a5e24f8ffa04",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,704 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 107,
"last_link_id": 113,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": 110
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": 112
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 104,
"type": "PreviewImage",
"pos": [
911.4377073728218,
418.9988584275326
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 113
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 107,
"type": "MaskToImage",
"pos": [
658.3044275669289,
511.8717365608471
],
"size": [
140,
26
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "mask",
"type": "MASK",
"link": 111
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
112,
113
]
}
],
"properties": {
"Node name for S&R": "MaskToImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
323.5443519405259,
489.8702544908114
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
110
]
},
{
"name": "MASK",
"type": "MASK",
"links": [
111
]
}
],
"properties": {
"Node name for S&R": "LoadImage",
"image": "clipspace/clipspace-painted-masked-1766731857414.png [input]"
},
"widgets_values": [
"clipspace/clipspace-painted-masked-1766731857414.png [input]",
"image"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-lite-2602-8steps.safetensors"
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
],
[
110,
100,
0,
99,
4,
"IMAGE"
],
[
111,
100,
1,
107,
0,
"MASK"
],
[
112,
107,
0,
99,
5,
"IMAGE"
],
[
113,
107,
0,
104,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7772383863288174,
"offset": [
248.90036994067935,
733.942567031054
]
},
"frontendVersion": "1.34.9",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "07dd34b942f866d5f95e8b812b6082d359079260",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,504 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 97,
"last_link_id": 83,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1070.207763671875,
-73.63389587402344
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 77
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 95,
"type": "ZImageT2ISampler",
"pos": [
719.3201904296875,
-72.24609375
],
"size": [
280.724609375,
386
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 83
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 75
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 76
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
77
]
}
],
"properties": {
"Node name for S&R": "ZImageT2ISampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
78
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
775.0554809570312,
-470.7688293457031
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 78
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
83
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
75
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
76
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
}
],
"links": [
[
75,
75,
0,
95,
1,
"STRING_PROMPT"
],
[
76,
73,
0,
95,
2,
"STRING_PROMPT"
],
[
77,
95,
0,
88,
0,
"IMAGE"
],
[
78,
92,
0,
96,
0,
"TransformerModel"
],
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
83,
96,
0,
95,
0,
"FunModels"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
985.5581665039062,
393.7902526855469
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7513148009015777,
"offset": [
598.8157343153007,
736.7507277278353
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"comfy-core": "0.6.0",
"CogVideoX-Fun": "244f11053106af1a58ac25fd0bbb508fd7b89c0f"
}
},
"version": 0.4
}
@@ -0,0 +1,614 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 102,
"last_link_id": 96,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
396.48551767952716,
419.158511044569
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
86
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"a7kXeQ5l9Dhspes7q3x3G (1).png",
"image"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 86
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
86,
100,
0,
99,
3,
"IMAGE"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.9594420649225085,
"offset": [
96.69529077203181,
609.5309015132086
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "244f11053106af1a58ac25fd0bbb508fd7b89c0f",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,696 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 106,
"last_link_id": 109,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 109
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
281.12436800744575,
417.89046761218935
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
107
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"z-image-turbo_00004_.png",
"image"
]
},
{
"id": 104,
"type": "PreviewImage",
"pos": [
911.4377073728218,
418.9988584275326
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 108
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 106,
"type": "ImageToCanny",
"pos": [
600.6679574832448,
416.8636458672411
],
"size": [
270,
82
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "input_image",
"type": "IMAGE",
"link": 107
}
],
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
108,
109
]
}
],
"properties": {
"Node name for S&R": "ImageToCanny"
},
"widgets_values": [
100,
200
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
],
[
107,
100,
0,
106,
0,
"IMAGE"
],
[
108,
106,
0,
104,
0,
"IMAGE"
],
[
109,
106,
0,
99,
3,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.5785382099684587,
"offset": [
793.7261950497102,
754.1754224231845
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "93aa7b2530dccd1e91c625eee439a5e24f8ffa04",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,692 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 105,
"last_link_id": 104,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 104
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 104,
"type": "PreviewImage",
"pos": [
844.9628845788926,
422.400403774606
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 103
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
281.12436800744575,
417.89046761218935
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
102
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"z-image-turbo_00004_.png",
"image"
]
},
{
"id": 105,
"type": "ImageToDepth",
"pos": [
611.889314416795,
425.1418418025041
],
"size": [
164.7039856092299,
28.074046747697935
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "input_image",
"type": "IMAGE",
"link": 102
}
],
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
103,
104
]
}
],
"properties": {
"Node name for S&R": "ImageToDepth"
}
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
],
[
102,
100,
0,
105,
0,
"IMAGE"
],
[
103,
105,
0,
104,
0,
"IMAGE"
],
[
104,
105,
0,
99,
3,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7154376913110498,
"offset": [
365.0760840892105,
597.0698647794829
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "93aa7b2530dccd1e91c625eee439a5e24f8ffa04",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,614 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 102,
"last_link_id": 96,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
396.48551767952716,
419.158511044569
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
86
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"a7kXeQ5l9Dhspes7q3x3G (1).png",
"image"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 86
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-lite-2602-8steps.safetensors"
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
86,
100,
0,
99,
3,
"IMAGE"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7208430239838531,
"offset": [
460.4224504078358,
644.7360602102879
]
},
"frontendVersion": "1.34.9",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "07dd34b942f866d5f95e8b812b6082d359079260",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,692 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 104,
"last_link_id": 99,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"model_cpu_offload"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 98
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 100,
"type": "LoadImage",
"pos": [
281.12436800744575,
417.89046761218935
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
97
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"z-image-turbo_00004_.png",
"image"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 103,
"type": "ImageToPose",
"pos": [
582.6473087858849,
424.6900692435995
],
"size": [
226.09371582945823,
27.59660529340124
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "input_image",
"type": "IMAGE",
"link": 97
}
],
"outputs": [
{
"name": "image",
"type": "IMAGE",
"links": [
98,
99
]
}
],
"properties": {
"Node name for S&R": "ImageToPose"
}
},
{
"id": 104,
"type": "PreviewImage",
"pos": [
844.9628845788926,
422.400403774606
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 99
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
],
[
97,
100,
0,
103,
0,
"IMAGE"
],
[
98,
103,
0,
99,
3,
"IMAGE"
],
[
99,
103,
0,
104,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7608917213421558,
"offset": [
217.08766227758116,
385.1943830579828
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "93aa7b2530dccd1e91c625eee439a5e24f8ffa04",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,614 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 102,
"last_link_id": 96,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"sequential_cpu_offload"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
396.48551767952716,
419.158511044569
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
86
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"z-image-turbo_00004_.png",
"image"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Tile-2.1-2601-8steps.safetensors"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 86
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1824,
2416,
43,
"fixed",
8,
0,
"Flow",
3,
0.8
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
86,
100,
0,
99,
3,
"IMAGE"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.6477940671634007,
"offset": [
411.8286437242131,
622.7162385081058
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "93aa7b2530dccd1e91c625eee439a5e24f8ffa04",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,614 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 102,
"last_link_id": 96,
"nodes": [
{
"id": 78,
"type": "Note",
"pos": [
18,
-46
],
"size": [
210,
88
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 91,
"type": "LoadZImageTextEncoderModel",
"pos": [
283.53765869140625,
-280.6837463378906
],
"size": [
407.4130859375,
102
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "text_encoder",
"type": "TextEncoderModel",
"links": [
80
]
},
{
"name": "tokenizer",
"type": "Tokenizer",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTextEncoderModel"
},
"widgets_values": [
"qwen_3_4b.safetensors",
"bf16"
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
88
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 97,
"type": "Note",
"pos": [
-354.4680507215508,
-433.6570714778354
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 93,
"type": "LoadZImageVAEModel",
"pos": [
1168.1599508804559,
-314.8650476692169
],
"size": [
377.8583984375,
84.69844055175781
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "vae",
"type": "VAEModel",
"links": [
79
]
}
],
"properties": {
"Node name for S&R": "LoadZImageVAEModel"
},
"widgets_values": [
"ae.safetensors",
"bf16"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1049.1402001998824,
-52.699751832945104
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 90
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 100,
"type": "LoadImage",
"pos": [
396.48551767952716,
419.158511044569
],
"size": [
270,
314.00000000000006
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
86
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"z-image-turbo_00004_.png",
"image"
]
},
{
"id": 92,
"type": "LoadZImageTransformerModel",
"pos": [
275.9798278808594,
-465.2391052246094
],
"size": [
416.3677673339844,
106.13789367675781
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
95
]
},
{
"name": "model_name",
"type": "STRING",
"links": [
82
]
}
],
"properties": {
"Node name for S&R": "LoadZImageTransformerModel"
},
"widgets_values": [
"z_image_turbo_bf16.safetensors",
"bf16"
]
},
{
"id": 99,
"type": "ZImageControlSampler",
"pos": [
727.2482831521481,
-52.41010988674983
],
"size": [
270,
350
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 87
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 88
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 86
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
90
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1824,
2416,
43,
"fixed",
8,
0,
"Flow",
3,
0.8
]
},
{
"id": 102,
"type": "LoadZImageControlNetInModel",
"pos": [
779.793189390101,
-457.3825558553134
],
"size": [
589.8698159570357,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 95
}
],
"outputs": [
{
"name": "transformer",
"type": "TransformerModel",
"links": [
96
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInModel"
},
"widgets_values": [
"z_image/z_image_control_2.1_lite.yaml",
"Z-Image-Turbo-Fun-Controlnet-Tile-2.1-lite-2601-8steps.safetensors"
]
},
{
"id": 96,
"type": "CombineZImagePipeline",
"pos": [
790.3572998046875,
-328.7134094238281
],
"size": [
342.5804748535156,
162
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "transformer",
"type": "TransformerModel",
"link": 96
},
{
"name": "vae",
"type": "VAEModel",
"link": 79
},
{
"name": "text_encoder",
"type": "TextEncoderModel",
"link": 80
},
{
"name": "tokenizer",
"type": "Tokenizer",
"link": 81
},
{
"name": "processor",
"shape": 7,
"type": "Processor",
"link": null
},
{
"name": "model_name",
"type": "STRING",
"widget": {
"name": "model_name"
},
"link": 82
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
87
]
}
],
"properties": {
"Node name for S&R": "CombineZImagePipeline"
},
"widgets_values": [
"",
"sequential_cpu_offload"
]
}
],
"links": [
[
79,
93,
0,
96,
1,
"VAEModel"
],
[
80,
91,
0,
96,
2,
"TextEncoderModel"
],
[
81,
91,
1,
96,
3,
"Tokenizer"
],
[
82,
92,
1,
96,
5,
"STRING"
],
[
86,
100,
0,
99,
3,
"IMAGE"
],
[
87,
96,
0,
99,
0,
"FunModels"
],
[
88,
75,
0,
99,
1,
"STRING_PROMPT"
],
[
89,
73,
0,
99,
2,
"STRING_PROMPT"
],
[
90,
99,
0,
88,
0,
"IMAGE"
],
[
95,
92,
0,
102,
0,
"TransformerModel"
],
[
96,
102,
0,
96,
0,
"TransformerModel"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
227.96267700195312,
-546.4359741210938,
1350.4793699732413,
404.87677206390265
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.6477940671634007,
"offset": [
475.03594149909674,
811.7170336004714
]
},
"frontendVersion": "1.34.9",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "07dd34b942f866d5f95e8b812b6082d359079260",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,320 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 91,
"last_link_id": 70,
"nodes": [
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
69
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
68
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1070.207763671875,
-73.63389587402344
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 70
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 80,
"type": "Note",
"pos": [
-424.31362772254084,
-372.4056987182365
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 86,
"type": "LoadZImageModel",
"pos": [
241.10867359909312,
-295.7069265122969
],
"size": [
428.9360739181402,
120.18527437036698
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
67
]
}
],
"properties": {
"Node name for S&R": "LoadZImageModel"
},
"widgets_values": [
"Z-Image-Turbo",
"model_cpu_offload",
"bf16"
]
},
{
"id": 91,
"type": "ZImageT2ISampler",
"pos": [
719.3201904296875,
-72.24609375
],
"size": [
280.724609375,
386
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 67
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 68
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 69
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
70
]
}
],
"properties": {
"Node name for S&R": "ZImageT2ISampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3
]
},
{
"id": 78,
"type": "Note",
"pos": [
17.042007499433538,
-27.473822111441212
],
"size": [
210,
88
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
}
],
"links": [
[
67,
86,
0,
91,
0,
"FunModels"
],
[
68,
75,
0,
91,
1,
"STRING_PROMPT"
],
[
69,
73,
0,
91,
2,
"STRING_PROMPT"
],
[
70,
91,
0,
88,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
220,
-380,
472,
232
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.6378658119731967,
"offset": [
856.870045195949,
740.927030543633
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "244f11053106af1a58ac25fd0bbb508fd7b89c0f",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
@@ -0,0 +1,431 @@
{
"id": "dcf2fcac-6293-4a86-b30b-f63e420177f2",
"revision": 0,
"last_node_id": 101,
"last_link_id": 92,
"nodes": [
{
"id": 73,
"type": "FunTextBox",
"pos": [
250,
160
],
"size": [
383.7149963378906,
183.83506774902344
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
90
]
}
],
"title": "Negtive Prompt(反向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
]
},
{
"id": 78,
"type": "Note",
"pos": [
17.042007499433538,
-27.473822111441212
],
"size": [
210,
88
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"You can write prompt here\n(你可以在此填写提示词)"
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 75,
"type": "FunTextBox",
"pos": [
250,
-50
],
"size": [
383.54010009765625,
156.71620178222656
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "prompt",
"type": "STRING_PROMPT",
"slot_index": 0,
"links": [
89
]
}
],
"title": "Positive Prompt(正向提示词)",
"properties": {
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
"A photo of Sakura, a 17-year-old high school student from Japan, captured in a candid, high-fidelity cinematic moment on a rainy evening. She is squatting low on the rain-slicked asphalt of an urban sidewalk, holding a transparent vinyl umbrella with a white handle resting over her shoulder in one hand, her other hand resting on her knee. The clear plastic canopy is streaked with rivulets of water and beaded with droplets that catch the ambient city light. A profound, silent interaction defines the scene: Sakura is looking directly downward, her expression gentle and focused, locking eyes with a small black cat sitting on the wet ground in front of her.\n\nSakura has long, lustrous black hair styled in a precise hime cut with blunt bangs across her forehead and sidelocks framing her cheeks, damp strands clinging subtly to her jacket, with a single red ribbon tied on the left side. Her visible pores on her nose, and a soft sheen of moisture on her cheeks. She wears a dark navy sailor-style school uniform (seifuku) featuring a white collar with red linear detailing and a bright red necktie loosely knotted at the chest; a simple black choker encircles her neck. The uniform jacket has oversized sleeves. Her lower body features a short, dark pleated miniskirt that fans slightly over clean white ankle socks that provide a stark contrast to the wet asphalt, ending in dark leather loafers that gleam with moisture.\n\nThe black cat sits upright in a shallow puddle, its short fur slicked by the rain, tilting its head back to stare intently up into Sakura's face, establishing a clear line of sight. The background is anchored by a large, illuminated red vending machine standing against the darkness, its cool bluish-white interior light spilling onto Sakura's profile and the umbrella. The ground reflects the red chassis and the neon streetlights in distorted patches on the wet pavement. Additional cool rain streaks fall through the frame, some caught in sharp focus and others blurred into vertical lines against the background lights. The scene is rendered with a wide-aperture lens creating a shallow depth of field, keeping the girl and cat in sharp focus while softening the background into gentle bokeh, with the texture of fine-grain 35mm film stock.\n"
]
},
{
"id": 96,
"type": "LoadImage",
"pos": [
429.45065831038295,
448.4639992924031
],
"size": [
270,
314
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
91
]
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"a7kXeQ5l9Dhspes7q3x3G (1).png",
"image"
]
},
{
"id": 86,
"type": "LoadZImageModel",
"pos": [
158.6376150119142,
-352.0410222978062
],
"size": [
428.9360739181402,
120.18527437036698
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
81
]
}
],
"properties": {
"Node name for S&R": "LoadZImageModel"
},
"widgets_values": [
"Z-Image-Turbo",
"model_cpu_offload",
"bf16"
]
},
{
"id": 80,
"type": "Note",
"pos": [
-486.1347709940621,
-416.6940105696127
],
"size": [
598.1623727144193,
233.48501180428053
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {
"text": ""
},
"widgets_values": [
"GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].\nmodel_full_load means that the entire model will be moved to the GPU.\n\nmodel_full_load_and_qfloat8 means that the entire model will be moved to the GPU,\nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nmodel_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.\n\nmodel_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use, \nand the transformer model has been quantized to float8, which can save more GPU memory. \n\nsequential_cpu_offload means that each layer of the model will be moved to the CPU after use, \nresulting in slower speeds but saving a large amount of GPU memory."
],
"color": "#432",
"bgcolor": "#653"
},
{
"id": 99,
"type": "LoadZImageControlNetInPipeline",
"pos": [
662.479156335491,
-343.8489854292294
],
"size": [
556.7720385739738,
107.30044919237525
],
"flags": {},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 81
}
],
"outputs": [
{
"name": "funmodels",
"type": "FunModels",
"links": [
88
]
}
],
"properties": {
"Node name for S&R": "LoadZImageControlNetInPipeline"
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-2602-8steps.safetensors",
"transformer"
]
},
{
"id": 88,
"type": "PreviewImage",
"pos": [
1070.207763671875,
-73.63389587402344
],
"size": [
366.56134033203125,
415.4429626464844
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 92
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 101,
"type": "ZImageControlSampler",
"pos": [
752.6720492832281,
-61.24351896170549
],
"size": [
270,
350
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "funmodels",
"type": "FunModels",
"link": 88
},
{
"name": "prompt",
"type": "STRING_PROMPT",
"link": 89
},
{
"name": "negative_prompt",
"type": "STRING_PROMPT",
"link": 90
},
{
"name": "control_image",
"shape": 7,
"type": "IMAGE",
"link": 91
},
{
"name": "inpaint_image",
"shape": 7,
"type": "IMAGE",
"link": null
},
{
"name": "mask_image",
"shape": 7,
"type": "IMAGE",
"link": null
}
],
"outputs": [
{
"name": "images",
"type": "IMAGE",
"links": [
92
]
}
],
"properties": {
"Node name for S&R": "ZImageControlSampler"
},
"widgets_values": [
1184,
1568,
43,
"fixed",
8,
0,
"Flow",
3,
0.85
]
}
],
"links": [
[
81,
86,
0,
99,
0,
"FunModels"
],
[
88,
99,
0,
101,
0,
"FunModels"
],
[
89,
75,
0,
101,
1,
"STRING_PROMPT"
],
[
90,
73,
0,
101,
2,
"STRING_PROMPT"
],
[
91,
96,
0,
101,
3,
"IMAGE"
],
[
92,
101,
0,
88,
0,
"IMAGE"
]
],
"groups": [
{
"id": 1,
"title": "Load Model",
"bounding": [
137.52894141282107,
-436.3340957855093,
472,
232
],
"color": "#b06634",
"font_size": 24,
"flags": {}
},
{
"id": 2,
"title": "Prompts",
"bounding": [
218,
-127,
450,
483
],
"color": "#3f789e",
"font_size": 24,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7396517503305734,
"offset": [
584.364342724736,
515.7058946883703
]
},
"frontendVersion": "1.36.11",
"workflowRendererVersion": "LG",
"workspace_info": {
"id": "776b62b4-bd17-4ed3-9923-b7aad000b1ea"
},
"node_versions": {
"CogVideoX-Fun": "244f11053106af1a58ac25fd0bbb508fd7b89c0f",
"comfy-core": "0.6.0"
}
},
"version": 0.4
}
+3 -3
View File
@@ -34,7 +34,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -150,8 +150,8 @@
"Node name for S&R": "LoadZImageModel"
},
"widgets_values": [
"Z-Image-Turbo",
"model_cpu_offload",
"Z-Image",
"model_group_offload",
"bf16"
]
},
@@ -34,7 +34,7 @@
"Node name for S&R": "FunTextBox"
},
"widgets_values": [
""
"低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
]
},
{
@@ -160,8 +160,8 @@
"Node name for S&R": "LoadZImageModel"
},
"widgets_values": [
"Z-Image-Turbo",
"model_cpu_offload",
"Z-Image",
"model_group_offload",
"bf16"
]
},
@@ -225,7 +225,7 @@
},
"widgets_values": [
"z_image/z_image_control_2.1.yaml",
"Z-Image-Turbo-Fun-Controlnet-Union-2.1-8steps.safetensors",
"Z-Image-Fun-Controlnet-Union-2.1.safetensors",
"transformer"
]
},
@@ -326,7 +326,7 @@
0,
"Flow",
3,
0.8
0.85
]
}
],
@@ -0,0 +1,6 @@
format: diffusers
pipeline: minimax-h3
transformer_additional_kwargs:
control_blocks_places: [0, 10, 20, 30, 40]
control_in_dim: 49
control_apply_audio: false
@@ -0,0 +1,7 @@
format: diffusers
pipeline: minimax-h3
transformer_additional_kwargs:
control_blocks_places: [0, 5, 10, 15, 20, 25, 30, 35, 40, 45]
control_in_dim: 49
control_apply_audio: false
inpaint_masked_pixel_mode: post_norm
@@ -0,0 +1,6 @@
format: diffusers
pipeline: minimax-h3
transformer_additional_kwargs:
control_blocks_places: [0, 10, 20, 30, 40]
control_in_dim: 24
control_apply_audio: false
@@ -0,0 +1,7 @@
format: diffusers
pipeline: qwenimage21
transformer_additional_kwargs:
# Dense control: inject a skip at every 2nd block (16 of 32 layers), matching the model's built-in
# `control_layers=None` default. Must contain 0. Dial back (e.g. [0, 4, 8, ...]) to shrink the trainable adapter.
control_layers: [0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30]
control_in_dim: 129
+52
View File
@@ -0,0 +1,52 @@
format: civitai
pipeline: Wan
transformer_additional_kwargs:
transformer_low_noise_model_subpath: ./low_noise_model
transformer_high_noise_model_subpath: ./high_noise_model
transformer_combination_type: "moe"
boundary: 0.900
dict_mapping:
in_dim: in_channels
dim: hidden_size
vae_kwargs:
vae_type: "AutoencoderKLWan3_8"
vae_subpath: Wan2.2_VAE.pth
temporal_compression_ratio: 4
spatial_compression_ratio: 16
latent_upsampler_kwargs:
mid_channels: 512
num_blocks_per_stage: 4
dims: 3
spatial_upsample: true
temporal_upsample: false
rational_spatial_scale: 1.5
use_rational_resampler: true
text_encoder_kwargs:
text_encoder_subpath: models_t5_umt5-xxl-enc-bf16.pth
tokenizer_subpath: google/umt5-xxl
text_length: 512
vocab: 256384
dim: 4096
dim_attn: 4096
dim_ffn: 10240
num_heads: 64
num_layers: 24
num_buckets: 32
shared_pos: False
dropout: 0.0
scheduler_kwargs:
scheduler_subpath: null
num_train_timesteps: 1000
shift: 5.0
use_dynamic_shifting: false
base_shift: 0.5
max_shift: 1.15
base_image_seq_len: 256
max_image_seq_len: 4096
image_encoder_kwargs:
image_encoder_subpath: models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth
+52
View File
@@ -0,0 +1,52 @@
format: civitai
pipeline: Wan
transformer_additional_kwargs:
transformer_low_noise_model_subpath: ./low_noise_model
transformer_high_noise_model_subpath: ./high_noise_model
transformer_combination_type: "moe"
boundary: 0.875
dict_mapping:
in_dim: in_channels
dim: hidden_size
vae_kwargs:
vae_type: "AutoencoderKLWan3_8"
vae_subpath: Wan2.2_VAE.pth
temporal_compression_ratio: 4
spatial_compression_ratio: 16
latent_upsampler_kwargs:
mid_channels: 512
num_blocks_per_stage: 4
dims: 3
spatial_upsample: true
temporal_upsample: false
rational_spatial_scale: 1.5
use_rational_resampler: true
text_encoder_kwargs:
text_encoder_subpath: models_t5_umt5-xxl-enc-bf16.pth
tokenizer_subpath: google/umt5-xxl
text_length: 512
vocab: 256384
dim: 4096
dim_attn: 4096
dim_ffn: 10240
num_heads: 64
num_layers: 24
num_buckets: 32
shared_pos: False
dropout: 0.0
scheduler_kwargs:
scheduler_subpath: null
num_train_timesteps: 1000
shift: 12.0
use_dynamic_shifting: false
base_shift: 0.5
max_shift: 1.15
base_image_seq_len: 256
max_image_seq_len: 4096
image_encoder_kwargs:
image_encoder_subpath: models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth
+2 -2
View File
@@ -10,9 +10,9 @@ for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.api.api import (infer_forward_api,
update_diffusion_transformer_api)
from videox_fun.ui.controller import ddpm_scheduler_dict
update_diffusion_transformer_api)
from videox_fun.ui.cogvideox_fun_ui import ui, ui_client, ui_host
from videox_fun.ui.controller import ddpm_scheduler_dict
if __name__ == "__main__":
# Choose the ui mode
+3 -3
View File
@@ -4,7 +4,6 @@ import sys
import time
import gradio as gr
import ray
import torch
current_file_path = os.path.abspath(__file__)
@@ -13,9 +12,10 @@ for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.api.api_multi_nodes import (MultiNodesEngine,
multi_nodes_infer_forward_api)
from videox_fun.ui.controller import flow_scheduler_dict
multi_nodes_infer_forward_api)
from videox_fun.ui.cogvideox_fun_ui import CogVideoXFunController
from videox_fun.ui.controller import flow_scheduler_dict
def main():
parser = argparse.ArgumentParser(description='xDiT HTTP Service')
+19 -26
View File
@@ -16,16 +16,14 @@ for project_root in project_roots:
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunInpaintPipeline,
CogVideoXFunPipeline)
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import get_image_to_video_latent, save_videos_grid
CogVideoXFunPipeline)
from videox_fun.utils import (apply_gpu_memory_mode, get_image_to_video_latent,
merge_lora, save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -36,6 +34,9 @@ from videox_fun.utils.utils import get_image_to_video_latent, save_videos_grid
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload_and_qfloat8"
@@ -75,7 +76,7 @@ partial_video_length = None
overlap_video_length = 4
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None and validation_image_end = None
validation_image_start = "asset/1.png"
@@ -102,7 +103,7 @@ transformer = CogVideoXTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -120,7 +121,7 @@ vae = AutoencoderKLCogVideoX.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -138,7 +139,7 @@ text_encoder = T5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Euler": EulerDiscreteScheduler,
"Euler A": EulerAncestralDiscreteScheduler,
"DPM++": DPMSolverMultistepScheduler,
@@ -184,20 +185,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -211,6 +202,7 @@ if partial_video_length is not None:
additional_frames = transformer.config.patch_size_t - latent_frames % transformer.config.patch_size_t
partial_video_length += additional_frames * vae.config.temporal_compression_ratio
validation_image = validation_image_start
init_frames = 0
last_frames = init_frames + partial_video_length
while init_frames < video_length:
@@ -308,6 +300,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+20 -28
View File
@@ -15,18 +15,16 @@ project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dir
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunPipeline,
CogVideoXFunInpaintPipeline)
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import get_image_to_video_latent, save_videos_grid
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunInpaintPipeline,
CogVideoXFunPipeline)
from videox_fun.utils import (apply_gpu_memory_mode, get_image_to_video_latent,
merge_lora, save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -37,6 +35,9 @@ from videox_fun.dist import set_multi_gpus_devices, shard_model
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload_and_qfloat8"
@@ -72,7 +73,7 @@ video_length = 49
fps = 8
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "A young woman with beautiful and clear eyes and blonde hair standing and white dress in a forest wearing a crown. She seems to be lost in thought, and the camera focuses on her face. The video is of high quality, and the view is very clear. High quality, masterpiece, best quality, highres, ultra-detailed, fantastic."
negative_prompt = "The video is not of a high quality, it has a low resolution. Watermark present in each frame. The background is solid. Strange body and strange trajectory. Distortion. "
@@ -94,7 +95,7 @@ transformer = CogVideoXTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -112,7 +113,7 @@ vae = AutoencoderKLCogVideoX.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -130,7 +131,7 @@ text_encoder = T5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Euler": EulerDiscreteScheduler,
"Euler A": EulerAncestralDiscreteScheduler,
"DPM++": DPMSolverMultistepScheduler,
@@ -176,20 +177,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -248,6 +239,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+20 -28
View File
@@ -14,18 +14,16 @@ project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dir
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunPipeline,
CogVideoXFunInpaintPipeline)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.utils import get_video_to_video_latent, save_videos_grid
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunInpaintPipeline,
CogVideoXFunPipeline)
from videox_fun.utils import (apply_gpu_memory_mode, get_video_to_video_latent,
merge_lora, save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -36,6 +34,9 @@ from videox_fun.dist import set_multi_gpus_devices, shard_model
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload_and_qfloat8"
@@ -70,7 +71,7 @@ video_length = 49
fps = 8
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you are preparing to redraw the reference video, set validation_video and validation_video_mask.
# If you do not use validation_video_mask, the entire video will be redrawn;
@@ -101,7 +102,7 @@ transformer = CogVideoXTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -119,7 +120,7 @@ vae = AutoencoderKLCogVideoX.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -137,7 +138,7 @@ text_encoder = T5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Euler": EulerDiscreteScheduler,
"Euler A": EulerAncestralDiscreteScheduler,
"DPM++": DPMSolverMultistepScheduler,
@@ -183,20 +184,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -243,6 +234,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+19 -29
View File
@@ -1,7 +1,6 @@
import os
import sys
import cv2
import numpy as np
import torch
from diffusers import (CogVideoXDDIMScheduler, DDIMScheduler,
@@ -16,18 +15,15 @@ project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dir
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import (CogVideoXFunControlPipeline,
CogVideoXFunInpaintPipeline)
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import get_video_to_video_latent, save_videos_grid
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLCogVideoX,
CogVideoXTransformer3DModel, T5EncoderModel,
T5Tokenizer)
from videox_fun.pipeline import CogVideoXFunControlPipeline
from videox_fun.utils import (apply_gpu_memory_mode, get_video_to_video_latent,
merge_lora, save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -38,6 +34,9 @@ from videox_fun.dist import set_multi_gpus_devices, shard_model
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload_and_qfloat8"
@@ -72,7 +71,7 @@ video_length = 49
fps = 8
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
control_video = "asset/pose.mp4"
@@ -97,7 +96,7 @@ transformer = CogVideoXTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -115,7 +114,7 @@ vae = AutoencoderKLCogVideoX.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -133,7 +132,7 @@ text_encoder = T5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Euler": EulerDiscreteScheduler,
"Euler A": EulerAncestralDiscreteScheduler,
"DPM++": DPMSolverMultistepScheduler,
@@ -170,20 +169,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=[], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -228,6 +217,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+194
View File
@@ -0,0 +1,194 @@
import os
import sys
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2, AutoTokenizer,
ErnieImageTransformer2DModel, Mistral3Model)
from videox_fun.pipeline import ErnieImagePipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/ERNIE-Image"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [1728, 992]
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
negative_prompt = "低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲。"
guidance_scale = 4.5
seed = 43
num_inference_steps = 40
lora_weight = 0.55
save_path = "samples/ernie-image-t2i"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = ErnieImageTransformer2DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
).to(weight_dtype)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLFlux2.from_pretrained(
model_name,
subfolder="vae"
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get tokenizer and text_encoder
tokenizer = AutoTokenizer.from_pretrained(
model_name, subfolder="tokenizer"
)
text_encoder = Mistral3Model.from_pretrained(
model_name, subfolder="text_encoder", torch_dtype=weight_dtype
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = ErnieImagePipeline(
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
transformer=transformer,
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=list(transformer.layers))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if compile_dit:
for i in range(len(pipeline.transformer.layers)):
pipeline.transformer.layers[i] = torch.compile(pipeline.transformer.layers[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
sample = pipeline(
prompt,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
num_inference_steps = num_inference_steps,
).images
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+29 -45
View File
@@ -15,22 +15,19 @@ for project_root in project_roots:
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, CLIPModel,
FantasyTalkingTransformer3DModel, FantasyTalkingAudioEncoder,
FantasyTalkingAudioEncoder,
FantasyTalkingTransformer3DModel,
WanT5EncoderModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import FantasyTalkingPipeline
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper,
replace_parameters_by_name)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_latent,
get_image_to_video_latent,
get_video_to_video_latent,
merge_video_audio, save_videos_grid)
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs,
get_image_to_video_latent, merge_lora,
merge_video_audio, save_videos_grid,
unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -41,6 +38,9 @@ from videox_fun.utils.utils import (filter_kwargs, get_image_latent,
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
@@ -72,10 +72,6 @@ num_skip_start_steps = 5
# Whether to offload TeaCache tensors to cpu to save a little bit of GPU memory.
teacache_offload = False
# Skip some cfg steps in inference
# Recommended to be set between 0.00 and 0.25
cfg_skip_ratio = 0
# Riflex config
enable_riflex = False
# Index of intrinsic frequency
@@ -85,8 +81,10 @@ riflex_k = 6
config_path = "config/wan2.1/wan_civitai.yaml"
# model path
# Please Download https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h/summary
# to models/Diffusion_Transformer/Wan2.1-I2V-14B-720P/audio_encoder for encoding audio.
# to models/Diffusion_Transformer/wav2vec2-base-960h for encoding audio.
model_name = "models/Diffusion_Transformer/Wan2.1-I2V-14B-720P"
# audio encoder model path. If None, will use os.path.join(model_name, "audio_encoder")
model_name_audio = "models/Diffusion_Transformer/wav2vec2-base-960h"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
@@ -97,7 +95,7 @@ shift = 5
# Load pretrained model if need
# The transformer_path is used for low noise model, the transformer_high_path is used for high noise model.
# The fantasytalking_model.ckpt can be downloaded in https://www.modelscope.cn/models/amap_cvlab/FantasyTalking/
transformer_path = "models/Personalized_Model/fantasytalking_model.ckpt"
transformer_path = "models/Personalized_Model/FantasyTalking/fantasytalking_model.ckpt"
vae_path = None
# Load lora model if need
lora_path = None
@@ -108,7 +106,7 @@ video_length = 81
fps = 23
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None
validation_image_start = "asset/8.png"
@@ -118,6 +116,7 @@ audio_path = "asset/talk.wav"
prompt = "一个女孩在海边说话。"
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
guidance_scale = 4.5
audio_guide_scale = 4.0
seed = 43
num_inference_steps = 40
lora_weight = 0.55
@@ -136,7 +135,7 @@ transformer = FantasyTalkingTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -169,7 +168,7 @@ vae = Chosen_AutoencoderKL.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -198,12 +197,11 @@ clip_image_encoder = CLIPModel.from_pretrained(
).to(weight_dtype)
clip_image_encoder = clip_image_encoder.eval()
audio_encoder = FantasyTalkingAudioEncoder(
os.path.join(model_name, "audio_encoder")
)
audio_encoder_path = model_name_audio if model_name_audio is not None else os.path.join(model_name, "audio_encoder")
audio_encoder = FantasyTalkingAudioEncoder(audio_encoder_path)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -241,22 +239,10 @@ if compile_dit:
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
transformer.freqs = transformer.freqs.to(device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
if coefficients is not None:
@@ -265,10 +251,6 @@ if coefficients is not None:
coefficients, num_inference_steps, teacache_threshold, num_skip_start_steps=num_skip_start_steps, offload=teacache_offload
)
if cfg_skip_ratio is not None:
print(f"Enable cfg_skip_ratio {cfg_skip_ratio}.")
pipeline.transformer.enable_cfg_skip(cfg_skip_ratio, num_inference_steps)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
@@ -291,6 +273,7 @@ with torch.no_grad():
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
audio_guide_scale = audio_guide_scale,
num_inference_steps = num_inference_steps,
video = input_video,
@@ -318,6 +301,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+242
View File
@@ -0,0 +1,242 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLWan, FlashHeadAudioEncoder,
FlashHeadTransformer3DModel)
from videox_fun.pipeline import FlashHeadPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs,
get_image_latent, merge_lora, merge_video_audio,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_full_load"
# Multi GPUs config
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# Config and model path
config_path = "config/wan2.1/wan_civitai.yaml"
# model path
# Please Download https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h/summary
model_name = "models/Diffusion_Transformer/SoulX-FlashHead-1_3B"
model_name_audio = "models/Diffusion_Transformer/wav2vec2-base-960h"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
shift = 5.0
stochastic_sampling = True
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [512, 512]
segment_frame_length = 33
fps = 25
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# The path of the reference image
ref_image = "asset/9.png"
# The path of the audio
audio_path = "asset/talk.wav"
# Audio guidance scale (FlashHead does not use text encoder, only audio conditioning)
audio_guide_scale = 1.0
seed = 42
num_inference_steps = 4
lora_weight = 0.55
save_path = "samples/flashhead-videos"
# FlashHead specific parameters
max_frames_num = 500
color_correction_strength = 1.0
use_apg = False
apg_momentum = 0.5
apg_norm_threshold = 1.0
audio_encode_mode = "stream"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
config = OmegaConf.load(config_path)
transformer = FlashHeadTransformer3DModel.from_pretrained(
os.path.join(model_name, "Model_Pro", config['transformer_additional_kwargs'].get('transformer_subpath', 'transformer')),
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLWan.from_pretrained(
os.path.join(model_name, "VAE_Wan/Wan2.1_VAE.pth"),
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Initialize FlashHead audio encoder for real-time audio encoding
# Uses Wav2Vec2Model (not Wav2Vec2ForCTC) matching original FlashHead implementation
audio_encoder = FlashHeadAudioEncoder(
model_name_audio, "cpu"
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
if sampler_name == "Flow_Unipc" or sampler_name == "Flow_DPM++":
config['scheduler_kwargs']['shift'] = 1
scheduler = Chosen_Scheduler(
**filter_kwargs(Chosen_Scheduler, OmegaConf.to_container(config['scheduler_kwargs']))
)
# Get Pipeline (FlashHead does not use text encoder or clip image encoder)
pipeline = FlashHeadPipeline(
transformer=transformer,
vae=vae,
scheduler=scheduler,
audio_encoder=audio_encoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
# For FlashHead, (segment_frame_length - 1) must be divisible by 4
segment_frame_length = (segment_frame_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio + 1 if segment_frame_length != 1 else 1
latent_frames = (segment_frame_length - 1) // vae.config.temporal_compression_ratio + 1
# Prepare ref_image latent for FlashHead (no clip_image needed)
ref_image = get_image_latent(ref_image, sample_size=sample_size)
sample = pipeline(
segment_frame_length = segment_frame_length,
height = sample_size[0],
width = sample_size[1],
generator = generator,
audio_guide_scale = audio_guide_scale,
num_inference_steps = num_inference_steps,
ref_image = ref_image,
audio_path = audio_path,
audio_encode_mode = audio_encode_mode,
shift = shift,
fps = fps,
max_frames_num = max_frames_num,
color_correction_strength = color_correction_strength,
use_apg = use_apg,
apg_momentum = apg_momentum,
apg_norm_threshold = apg_norm_threshold,
stochastic_sampling = stochastic_sampling,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if sample.size()[2] == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
merge_video_audio(video_path=video_path, audio_path=audio_path)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+18 -26
View File
@@ -14,13 +14,11 @@ from videox_fun.models import (AutoencoderKL, CLIPTextModel, CLIPTokenizer,
FluxTransformer2DModel, T5EncoderModel,
T5TokenizerFast)
from videox_fun.pipeline import FluxPipeline
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -31,6 +29,9 @@ from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload_and_qfloat8"
@@ -62,10 +63,10 @@ lora_path = None
sample_size = [1344, 768]
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
negative_prompt = "The video is not of a high quality, it has a low resolution. Watermark present in each frame. The background is solid. Strange body and strange trajectory. Distortion. "
negative_prompt = " "
guidance_scale = 1.0
seed = 43
num_inference_steps = 50
@@ -84,7 +85,7 @@ transformer = FluxTransformer2DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -102,7 +103,7 @@ vae = AutoencoderKL.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -127,7 +128,7 @@ text_encoder_2 = T5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -156,7 +157,7 @@ if ulysses_degree > 1 or ring_degree > 1:
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.text_model.encoder.layers)
text_encoder = shard_fn(text_encoder)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
@@ -164,20 +165,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -207,6 +198,7 @@ def save_results():
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
+17 -35
View File
@@ -2,8 +2,7 @@ import os
import sys
import torch
from diffusers import (FlowMatchEulerDiscreteScheduler)
from diffusers import FlowMatchEulerDiscreteScheduler
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
@@ -11,20 +10,15 @@ for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2,
from videox_fun.models import (AutoencoderKLFlux2, Flux2Transformer2DModel,
Mistral3ForConditionalGeneration,
PixtralProcessor, Flux2Transformer2DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
PixtralProcessor)
from videox_fun.pipeline import Flux2Pipeline
from videox_fun.utils import (register_auto_device_hook,
safe_enable_group_offload)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -69,7 +63,7 @@ lora_path = None
sample_size = [1344, 768]
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Please use as detailed a prompt as possible to describe the object that needs to be generated.
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
@@ -92,7 +86,7 @@ transformer = Flux2Transformer2DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -110,7 +104,7 @@ vae = AutoencoderKLFlux2.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -129,7 +123,7 @@ text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -156,7 +150,7 @@ if ulysses_degree > 1 or ring_degree > 1:
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.language_model.layers)
text_encoder = shard_fn(text_encoder)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
@@ -164,23 +158,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -209,6 +190,7 @@ def save_results():
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
+18 -39
View File
@@ -1,11 +1,9 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
@@ -14,23 +12,16 @@ for project_root in project_roots:
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2,
Flux2ControlTransformer2DModel,
Mistral3ForConditionalGeneration,
PixtralProcessor, Flux2ControlTransformer2DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
PixtralProcessor)
from videox_fun.pipeline import Flux2ControlPipeline
from videox_fun.utils import (register_auto_device_hook,
safe_enable_group_offload)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image, get_image_latent,
get_image_to_video_latent,
get_video_to_video_latent,
save_videos_grid)
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image,
get_image_latent, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -69,7 +60,7 @@ model_name = "models/Diffusion_Transformer/FLUX.2-dev"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union.safetensors"
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
vae_path = None
lora_path = None
@@ -77,7 +68,7 @@ lora_path = None
sample_size = [1728, 992]
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
image = None
control_image = None
@@ -108,7 +99,7 @@ transformer = Flux2ControlTransformer2DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -126,7 +117,7 @@ vae = AutoencoderKLFlux2.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -145,7 +136,7 @@ text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -172,7 +163,7 @@ if ulysses_degree > 1 or ring_degree > 1:
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.language_model.layers, ignored_modules=[text_encoder.language_model.embed_tokens], transformer_layer_cls_to_wrap=["MistralDecoderLayer", "PixtralTransformer"])
text_encoder = shard_fn(text_encoder)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
@@ -180,23 +171,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -249,6 +227,7 @@ def save_results():
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
+18 -39
View File
@@ -1,11 +1,9 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
@@ -14,23 +12,16 @@ for project_root in project_roots:
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2,
Flux2ControlTransformer2DModel,
Mistral3ForConditionalGeneration,
PixtralProcessor, Flux2ControlTransformer2DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
PixtralProcessor)
from videox_fun.pipeline import Flux2ControlPipeline
from videox_fun.utils import (register_auto_device_hook,
safe_enable_group_offload)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image, get_image_latent,
get_image_to_video_latent,
get_video_to_video_latent,
save_videos_grid)
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image,
get_image_latent, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -69,7 +60,7 @@ model_name = "models/Diffusion_Transformer/FLUX.2-dev"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union.safetensors"
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
vae_path = None
lora_path = None
@@ -77,7 +68,7 @@ lora_path = None
sample_size = [1728, 992]
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
image = None
control_image = "asset/pose.jpg"
@@ -108,7 +99,7 @@ transformer = Flux2ControlTransformer2DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -126,7 +117,7 @@ vae = AutoencoderKLFlux2.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -145,7 +136,7 @@ text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -172,7 +163,7 @@ if ulysses_degree > 1 or ring_degree > 1:
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.language_model.layers, ignored_modules=[text_encoder.language_model.embed_tokens], transformer_layer_cls_to_wrap=["MistralDecoderLayer", "PixtralTransformer"])
text_encoder = shard_fn(text_encoder)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
@@ -180,23 +171,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -249,6 +227,7 @@ def save_results():
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
+18 -39
View File
@@ -1,11 +1,9 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
@@ -14,23 +12,16 @@ for project_root in project_roots:
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2,
Flux2ControlTransformer2DModel,
Mistral3ForConditionalGeneration,
PixtralProcessor, Flux2ControlTransformer2DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
PixtralProcessor)
from videox_fun.pipeline import Flux2ControlPipeline
from videox_fun.utils import (register_auto_device_hook,
safe_enable_group_offload)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image, get_image_latent,
get_image_to_video_latent,
get_video_to_video_latent,
save_videos_grid)
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image,
get_image_latent, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -69,7 +60,7 @@ model_name = "models/Diffusion_Transformer/FLUX.2-dev"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union.safetensors"
transformer_path = "models/Personalized_Model/FLUX.2-dev-Fun-Controlnet-Union-2602.safetensors"
vae_path = None
lora_path = None
@@ -77,7 +68,7 @@ lora_path = None
sample_size = [1728, 992]
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
image = "asset/8.png"
control_image = "asset/pose.jpg"
@@ -108,7 +99,7 @@ transformer = Flux2ControlTransformer2DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -126,7 +117,7 @@ vae = AutoencoderKLFlux2.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -145,7 +136,7 @@ text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -172,7 +163,7 @@ if ulysses_degree > 1 or ring_degree > 1:
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.language_model.layers, ignored_modules=[text_encoder.language_model.embed_tokens], transformer_layer_cls_to_wrap=["MistralDecoderLayer", "PixtralTransformer"])
text_encoder = shard_fn(text_encoder)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
@@ -180,23 +171,10 @@ if compile_dit:
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_group_offload":
register_auto_device_hook(pipeline.transformer)
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["img_in", "txt_in", "timestep"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -249,6 +227,7 @@ def save_results():
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
+23 -37
View File
@@ -4,8 +4,6 @@ import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from diffusers.utils import export_to_video
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
@@ -13,26 +11,20 @@ project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dir
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from diffusers.schedulers.scheduling_unipc_multistep import \
UniPCMultistepScheduler
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLHunyuanVideo, CLIPTextModel, CLIPImageProcessor,
CLIPTokenizer, HunyuanVideoTransformer3DModel,
LlavaForConditionalGeneration, LlamaTokenizerFast)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import HunyuanVideoPipeline, HunyuanVideoI2VPipeline
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper,
replace_parameters_by_name)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
save_videos_grid)
from videox_fun.utils.utils import get_image
from videox_fun.models import (AutoencoderKLHunyuanVideo, CLIPImageProcessor,
CLIPTextModel, CLIPTokenizer,
HunyuanVideoTransformer3DModel,
LlamaTokenizerFast,
LlavaForConditionalGeneration)
from videox_fun.pipeline import HunyuanVideoI2VPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -43,6 +35,9 @@ from videox_fun.utils.utils import get_image
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
@@ -76,7 +71,7 @@ video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None and validation_image_end = None
validation_image_start = "asset/1.png"
@@ -101,7 +96,7 @@ transformer = HunyuanVideoTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -118,7 +113,7 @@ vae = AutoencoderKLHunyuanVideo.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -157,7 +152,7 @@ image_processor = CLIPImageProcessor.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -194,20 +189,10 @@ if compile_dit:
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["x_embedder", "context_embedder", "time_text_embed", "rope", "proj_out"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["x_embedder", "context_embedder", "time_text_embed", "rope", "proj_out"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['x_embedder', 'context_embedder', 'time_text_embed', 'rope', 'proj_out'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -250,6 +235,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+17 -32
View File
@@ -4,8 +4,6 @@ import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from diffusers.utils import export_to_video
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
@@ -13,25 +11,18 @@ project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dir
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from diffusers.schedulers.scheduling_unipc_multistep import \
UniPCMultistepScheduler
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLHunyuanVideo, CLIPTextModel,
CLIPTokenizer, HunyuanVideoTransformer3DModel,
LlamaModel, LlamaTokenizerFast)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import HunyuanVideoPipeline
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper,
replace_parameters_by_name)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
save_videos_grid)
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -42,6 +33,9 @@ from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
@@ -75,7 +69,7 @@ video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
negative_prompt = "The video is not of a high quality, it has a low resolution. Watermark present in each frame. The background is solid. Strange body and strange trajectory. Distortion. "
@@ -96,7 +90,7 @@ transformer = HunyuanVideoTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -113,7 +107,7 @@ vae = AutoencoderKLHunyuanVideo.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -147,7 +141,7 @@ text_encoder_2 = CLIPTextModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -183,20 +177,10 @@ if compile_dit:
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["x_embedder", "context_embedder", "time_text_embed", "rope", "proj_out"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["x_embedder", "context_embedder", "time_text_embed", "rope", "proj_out"], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['x_embedder', 'context_embedder', 'time_text_embed', 'rope', 'proj_out'])
generator = torch.Generator(device=device).manual_seed(seed)
@@ -235,6 +219,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
+299
View File
@@ -0,0 +1,299 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLWan, AutoTokenizer, CLIPModel,
InfiniteTalkAudioEncoder,
InfiniteTalkTransformer3DModel,
WanT5EncoderModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import InfiniteTalkPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs, get_image,
get_image_latent, merge_lora, merge_video_audio,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
# Multi GPUs config
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# Support TeaCache.
enable_teacache = True
# Recommended to be set between 0.05 and 0.30. A larger threshold can cache more steps, speeding up the inference process,
# but it may cause slight differences between the generated content and the original content.
# # --------------------------------------------------------------------------------------------------- #
# | Model Name | threshold | Model Name | threshold | Model Name | threshold |
# | Wan2.1-T2V-1.3B | 0.05~0.10 | Wan2.1-T2V-14B | 0.10~0.15 | Wan2.1-I2V-14B-720P | 0.20~0.30 |
# | Wan2.1-I2V-14B-480P | 0.20~0.25 | Wan2.1-Fun-*-1.3B-* | 0.05~0.10 | Wan2.1-Fun-*-14B-* | 0.20~0.30 |
# # --------------------------------------------------------------------------------------------------- #
teacache_threshold = 0.20
# The number of steps to skip TeaCache at the beginning of the inference process, which can
# reduce the impact of TeaCache on generated video quality.
num_skip_start_steps = 5
# Whether to offload TeaCache tensors to cpu to save a little bit of GPU memory.
teacache_offload = False
# Config and model path
config_path = "config/wan2.1/wan_civitai.yaml"
# model path
model_name = "models/Diffusion_Transformer/Wan2.1-I2V-14B-480P"
model_name_audio = "models/Diffusion_Transformer/chinese-wav2vec2-base/"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
shift = 5.0
# Load pretrained model if need
transformer_path = "models/Personalized_Model/infinitetalk.safetensors"
vae_path = None
lora_path = None
# Other params
sample_size = [832, 480]
segment_frame_length = 81
fps = 25
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# The path of the reference image
ref_image = "asset/8.png"
# The path of the audio
audio_path = "asset/talk.wav"
# prompts
prompt = "一个人在说话。"
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
guidance_scale = 5.0
audio_guide_scale = 4.0
seed = 43
num_inference_steps = 40
lora_weight = 0.55
save_path = "samples/infitetalk-videos"
# InfiniteTalk specific parameters
max_frames_num = 500 # Total frames to generate
color_correction_strength = 1 # Color correction strength (0.0-1.0)
use_apg = False # Use Adaptive Projected Guidance
apg_momentum = 0.5
apg_norm_threshold = 1.0
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
config = OmegaConf.load(config_path)
transformer = InfiniteTalkTransformer3DModel.from_pretrained(
os.path.join(model_name, config['transformer_additional_kwargs'].get('transformer_subpath', 'transformer')),
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLWan.from_pretrained(
os.path.join(model_name, config['vae_kwargs'].get('vae_subpath', 'vae')),
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Tokenizer
tokenizer = AutoTokenizer.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('tokenizer_subpath', 'tokenizer')),
)
# Get Text encoder
text_encoder = WanT5EncoderModel.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('text_encoder_subpath', 'text_encoder')),
additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Initialize InfiniteTalk audio encoder for real-time audio encoding
# Uses Wav2Vec2Model (not Wav2Vec2ForCTC) matching original InfiniteTalk implementation
audio_encoder = InfiniteTalkAudioEncoder(
model_name_audio, "cpu"
)
# Get Clip Image Encoder
clip_image_encoder = CLIPModel.from_pretrained(
os.path.join(model_name, config['image_encoder_kwargs'].get('image_encoder_subpath', 'image_encoder')),
).to(weight_dtype)
clip_image_encoder = clip_image_encoder.eval()
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
if sampler_name == "Flow_Unipc" or sampler_name == "Flow_DPM++":
config['scheduler_kwargs']['shift'] = 1
scheduler = Chosen_Scheduler(
**filter_kwargs(Chosen_Scheduler, OmegaConf.to_container(config['scheduler_kwargs']))
)
# Get Pipeline
pipeline = InfiniteTalkPipeline(
transformer=transformer,
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
scheduler=scheduler,
audio_encoder=audio_encoder,
clip_image_encoder=clip_image_encoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
if coefficients is not None:
print(f"Enable TeaCache with threshold {teacache_threshold} and skip the first {num_skip_start_steps} steps.")
pipeline.transformer.enable_teacache(
coefficients, num_inference_steps, teacache_threshold, num_skip_start_steps=num_skip_start_steps, offload=teacache_offload
)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
# For InfiniteTalk, (segment_frame_length - 1) must be divisible by 4
segment_frame_length = (segment_frame_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio + 1 if segment_frame_length != 1 else 1
latent_frames = (segment_frame_length - 1) // vae.config.temporal_compression_ratio + 1
# Prepare clip_image from original ref_image path
clip_image = get_image(ref_image)
ref_image = get_image_latent(ref_image, sample_size=sample_size)
sample = pipeline(
prompt,
segment_frame_length = segment_frame_length,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
audio_guide_scale = audio_guide_scale,
num_inference_steps = num_inference_steps,
ref_image = ref_image,
clip_image = clip_image, # Pass clip_image
audio_path = audio_path,
shift = shift,
fps = fps,
max_frames_num = max_frames_num,
color_correction_strength = color_correction_strength,
use_apg = use_apg,
apg_momentum = apg_momentum,
apg_norm_threshold = apg_norm_threshold,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if sample.size()[2] == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
merge_video_audio(video_path=video_path, audio_path=audio_path)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+210
View File
@@ -0,0 +1,210 @@
import os
import sys
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLFlux2, AutoTokenizer,
LensGptOssEncoder, LensTransformer2DModel)
from videox_fun.pipeline import LensPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/Lens"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [1728, 992]
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Set to True on A100/V100 to dequantize MXFP4 GPT-OSS weights.
dequantize_mxfp4 = False
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
negative_prompt = " "
guidance_scale = 4.5
seed = 43
num_inference_steps = 40
lora_weight = 0.55
save_path = "samples/lens-t2i"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Get transformer
transformer = LensTransformer2DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
).to(weight_dtype)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLFlux2.from_pretrained(
model_name,
subfolder="vae",
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get tokenizer and text_encoder
tokenizer = AutoTokenizer.from_pretrained(
model_name, subfolder="tokenizer"
)
text_encoder_kwargs = {"subfolder": "text_encoder", "torch_dtype": weight_dtype}
try:
from transformers import Mxfp4Config
text_encoder_kwargs["quantization_config"] = Mxfp4Config(
dequantize=dequantize_mxfp4
)
except ImportError:
pass # Older transformers without Mxfp4Config
text_encoder = LensGptOssEncoder.from_pretrained(
model_name, **text_encoder_kwargs
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LensPipeline(
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
transformer=transformer,
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=list(text_encoder.model.layers))
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['img_in', 'txt_in', 'timestep'])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
sample = pipeline(
prompt,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
num_inference_steps = num_inference_steps,
).images
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0]
image.save(video_path)
print(f"Saved image to: {video_path}")
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+238
View File
@@ -0,0 +1,238 @@
import os
import sys
import numpy as np
import torch
from PIL import Image
from transformers import AutoProcessor
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLQwenImage,
LingBotVideoTransformer3DModel,
Qwen3VLForConditionalGeneration)
from videox_fun.models.lingbot_video_rewriter import ensure_json_caption
from videox_fun.pipeline import LingBotVideoI2VPipeline
from videox_fun.pipeline.pipeline_lingbot_video import DEFAULT_NEGATIVE_PROMPT
from videox_fun.utils import (FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
# Sequence parallelism shards the video tokens across ranks and keeps the text tokens
# replicated, so the video token count (T/pF * H/16 * W/16) must be divisible by
# ulysses_degree * ring_degree.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
# Config and model path
# model path
model_name = "models/Diffusion_Transformer/lingbot-video-dense-1.3b"
# Rewriter weights: the base VLM and the rewriter LoRA used to rewrite the
# plain prompt into the structured JSON caption the DiT expects.
rewriter_base_model = "models/Diffusion_Transformer/Qwen3.6-27B"
rewriter_lora_path = "models/Diffusion_Transformer/lingbot-video-rewriter-lora"
# Only "Flow_Unipc" is supported: LingBot-Video ships and was trained with FlowUniPCMultistepScheduler.
sampler_name = "Flow_Unipc"
# Flow shift. 3.0 is the officially recommended value for both dense and MoE models.
shift = 3.0
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [480, 832]
video_length = 81
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# The condition image is used twice: as Qwen3-VL visual input and as a clean
# first-frame latent injected into the diffusion latent (ti2v).
validation_image = "asset/1.png"
# prompts
# Write a plain natural-language prompt: it is ALWAYS rewritten into the
# structured JSON caption the DiT expects by the official prompt rewriter
# (EXPAND -> MAP, Qwen3.6-27B base + rewriter LoRA). For ti2v the same first
# frame is fed to the rewriter. Direct JSON/hand-written input is not a
# supported path; the rewrite result is cached under save_path.
prompt = "一只棕色的狗摇着头,坐在舒适房间里的浅色沙发上。在狗的后面,架子上有一幅镶框的画,周围是粉红色的花朵。房间里柔和温暖的灯光营造出舒适的氛围。"
negative_prompt = DEFAULT_NEGATIVE_PROMPT
guidance_scale = 3.0
seed = 43
num_inference_steps = 40
lora_weight = 0.55
save_path = "samples/lingbot-video-i2v"
# Rewrite the prompt before loading any generation model (the rewriter's 27B
# base VLM is freed right after, so it never coexists with the DiT on GPU).
# ti2v: the same first frame is fed to the rewriter.
prompt = ensure_json_caption(
prompt, mode="ti2v", duration=round(video_length / fps, 2),
first_frame=validation_image,
cache_file=os.path.join(save_path, "caption_cache.json"),
base=rewriter_base_model, adapter=rewriter_lora_path,
)
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = LingBotVideoTransformer3DModel.from_pretrained(
os.path.join(model_name, "transformer"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Re-apply the fp32-sensitive-module cast (norm / router / modulation stay fp32).
transformer = transformer.to(weight_dtype)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae (diffusers-format QwenImage VAE, Wan-style 16ch causal VAE)
vae = AutoencoderKLQwenImage.from_pretrained(
model_name,
subfolder="vae",
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Processor (Qwen3-VL tokenizer + image processor)
processor = AutoProcessor.from_pretrained(
os.path.join(model_name, "processor"),
)
# Get Text encoder (Qwen3-VL)
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow_Unipc": FlowUniPCMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
# Get Pipeline
pipeline = LingBotVideoI2VPipeline(
transformer=transformer,
vae=vae,
text_encoder=text_encoder,
processor=processor,
scheduler=scheduler,
)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['time_embedder', 'time_modulation', 'text_embedder', 'norm', 'router', 'scale_shift_table', 'proj_out'])
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
video_length = int((video_length - 1) // pipeline.vae_scale_factor_temporal * pipeline.vae_scale_factor_temporal) + 1 if video_length != 1 else 1
image = Image.open(validation_image).convert("RGB")
sample = pipeline(
prompt,
image = image,
num_frames = video_length,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
shift = shift,
num_inference_steps = num_inference_steps,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
# count outputs only: caption_cache.json must not shift the index
index = len([path for path in os.listdir(save_path) if path.endswith((".mp4", ".png"))]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+238
View File
@@ -0,0 +1,238 @@
import os
import sys
import numpy as np
import torch
from PIL import Image
from transformers import AutoProcessor
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLQwenImage,
LingBotVideoTransformer3DModel,
Qwen3VLForConditionalGeneration)
from videox_fun.models.lingbot_video_rewriter import ensure_json_caption
from videox_fun.pipeline import LingBotVideoPipeline
from videox_fun.pipeline.pipeline_lingbot_video import DEFAULT_NEGATIVE_PROMPT
from videox_fun.utils import (FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
# Sequence parallelism shards the video tokens across ranks and keeps the text tokens
# replicated, so the video token count (T/pF * H/16 * W/16) must be divisible by
# ulysses_degree * ring_degree.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
# Config and model path
# model path
model_name = "models/Diffusion_Transformer/lingbot-video-dense-1.3b/"
# Rewriter weights: the base VLM and the rewriter LoRA used to rewrite the
# plain prompt into the structured JSON caption the DiT expects.
rewriter_base_model = "models/Diffusion_Transformer/Qwen3.6-27B"
rewriter_lora_path = "models/Diffusion_Transformer/lingbot-video-rewriter-lora"
# Only "Flow_Unipc" is supported: LingBot-Video ships and was trained with FlowUniPCMultistepScheduler.
sampler_name = "Flow_Unipc"
# Flow shift. 3.0 is the officially recommended value for both dense and MoE models.
shift = 3.0
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [480, 832]
# video_length 1 generates a still image (t2i); videos must be 4n+1 frames.
video_length = 81
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# prompts
# Write a plain natural-language prompt: it is ALWAYS rewritten into the
# structured JSON caption the DiT expects by the official prompt rewriter
# (EXPAND -> MAP, Qwen3.6-27B base + rewriter LoRA). Direct JSON/hand-written
# input is not a supported path; the rewrite result is cached under save_path
# so re-runs with the same prompt skip the rewrite.
prompt = (
"A young musician sits on a weathered wooden stool in a sunlit rehearsal room, "
"steadily strumming an acoustic guitar. Warm golden-hour light streams through "
"tall windows, dust motes drifting slowly in the air; the wooden floor shows "
"visible wear and the walls carry acoustic panels. The camera slowly orbits "
"from a side profile to a frontal view at eye level, keeping the musician "
"centered in the frame."
)
negative_prompt = DEFAULT_NEGATIVE_PROMPT
guidance_scale = 3.0
seed = 43
num_inference_steps = 40
lora_weight = 0.55
save_path = "samples/lingbot-video-t2v"
# Rewrite the prompt before loading any generation model (the rewriter's 27B
# base VLM is freed right after, so it never coexists with the DiT on GPU).
prompt = ensure_json_caption(
prompt, mode="t2v", duration=round(video_length / fps, 2),
cache_file=os.path.join(save_path, "caption_cache.json"),
base=rewriter_base_model, adapter=rewriter_lora_path,
)
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = LingBotVideoTransformer3DModel.from_pretrained(
os.path.join(model_name, "transformer"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Re-apply the fp32-sensitive-module cast (norm / router / modulation stay fp32).
transformer = transformer.to(weight_dtype)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae (diffusers-format QwenImage VAE, Wan-style 16ch causal VAE)
vae = AutoencoderKLQwenImage.from_pretrained(
model_name,
subfolder="vae",
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Processor (Qwen3-VL tokenizer + image processor)
processor = AutoProcessor.from_pretrained(
os.path.join(model_name, "processor"),
)
# Get Text encoder (Qwen3-VL)
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow_Unipc": FlowUniPCMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
# Get Pipeline
pipeline = LingBotVideoPipeline(
transformer=transformer,
vae=vae,
text_encoder=text_encoder,
processor=processor,
scheduler=scheduler,
)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['time_embedder', 'time_modulation', 'text_embedder', 'norm', 'router', 'scale_shift_table', 'proj_out'])
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
video_length = int((video_length - 1) // pipeline.vae_scale_factor_temporal * pipeline.vae_scale_factor_temporal) + 1 if video_length != 1 else 1
sample = pipeline(
prompt,
num_frames = video_length,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
shift = shift,
num_inference_steps = num_inference_steps,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
# count outputs only: caption_cache.json must not shift the index
index = len([path for path in os.listdir(save_path) if path.endswith((".mp4", ".png"))]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
@@ -0,0 +1,335 @@
import gc
import os
import sys
import torch
import torch.nn.functional as F
from transformers import AutoProcessor
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLQwenImage,
LingBotVideoTransformer3DModel,
Qwen3VLForConditionalGeneration)
from videox_fun.models.lingbot_video_rewriter import ensure_json_caption
from videox_fun.pipeline import LingBotVideoPipeline
from videox_fun.pipeline.pipeline_lingbot_video import (
DEFAULT_NEGATIVE_PROMPT, prepare_refiner_latent)
from videox_fun.utils import (FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, save_videos_grid)
# Two-stage LingBot-Video t2v: the base DiT samples at a low resolution, then the
# "refiner" DiT re-noises the upsampled latent to sigma = refiner_t_thresh and
# denoises it at the target resolution.
#
# The two DiTs are loaded and freed one at a time, so a single GPU only ever holds
# one 30B transformer (the MoE base and refiner are ~60GB each in bfloat16).
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
# Sequence parallelism shards the video tokens across ranks and keeps the text tokens
# replicated, so the video token count (T/pF * H/16 * W/16) must be divisible by
# ulysses_degree * ring_degree, at both the base and the refiner resolution.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
# Config and model path
# model path
# The refiner ships only with the MoE 30B-A3B model, as its "refiner" subfolder.
model_name = "models/Diffusion_Transformer/lingbot-video-moe-30b-a3b"
refiner_model_name = model_name
# Subfolders of the base and refiner DiT inside the model root.
transformer_subpath = "transformer"
refiner_subpath = "refiner"
# Rewriter weights: the base VLM and the rewriter LoRA used to rewrite the
# plain prompt into the structured JSON caption the DiT expects.
rewriter_base_model = "models/Diffusion_Transformer/Qwen3.6-27B"
rewriter_lora_path = "models/Diffusion_Transformer/lingbot-video-rewriter-lora"
# Only "Flow_Unipc" is supported: LingBot-Video ships and was trained with FlowUniPCMultistepScheduler.
sampler_name = "Flow_Unipc"
# Flow shift. 3.0 is the officially recommended value for both dense and MoE models.
shift = 3.0
# Load pretrained model if need
transformer_path = None
refiner_path = None
vae_path = None
# Base stage params. video_length must be 1 or 4n+1; 121 frames is 5s at 24 fps.
sample_size = [480, 832]
video_length = 81
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# prompts
# Write a plain natural-language prompt: it is ALWAYS rewritten into the
# structured JSON caption the DiT expects by the official prompt rewriter
# (EXPAND -> MAP, Qwen3.6-27B base + rewriter LoRA). Direct JSON/hand-written
# input is not a supported path; the rewrite result is cached under save_path.
prompt = (
"A young musician sits on a weathered wooden stool in a sunlit rehearsal room, "
"steadily strumming an acoustic guitar. Warm golden-hour light streams through "
"tall windows, dust motes drifting slowly in the air. The camera slowly orbits "
"from a side profile to a frontal view at eye level, keeping the musician "
"centered in the frame."
)
negative_prompt = DEFAULT_NEGATIVE_PROMPT
guidance_scale = 3.0
seed = 43
num_inference_steps = 40
save_path = "samples/lingbot-video-t2v-refine"
# Refiner stage params (official defaults). The refiner attends over the full
# high-resolution latent, so 1088x1920 is heavy on a single GPU: prefer fewer
# frames there, or shard the DiT across GPUs.
refiner_sample_size = [1088, 1920]
refiner_steps = 8
refiner_guidance_scale = 3.0
refiner_shift = 3.0
# Re-noise level: the refiner only walks the schedule from this sigma down to 0.
refiner_t_thresh = 0.85
# Extra low-noise steps appended after the truncated schedule.
refiner_sigma_tail_steps = 2
# Rewrite the prompt before loading any generation model (the rewriter's 27B
# base VLM is freed right after, so it never coexists with the DiT on GPU).
prompt = ensure_json_caption(
prompt, mode="t2v", duration=round(video_length / fps, 2),
cache_file=os.path.join(save_path, "caption_cache.json"),
base=rewriter_base_model, adapter=rewriter_lora_path,
)
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
def load_transformer(root, subpath, checkpoint_path):
transformer = LingBotVideoTransformer3DModel.from_pretrained(
os.path.join(root, subpath),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Re-apply the fp32-sensitive-module cast (norm / router / modulation stay fp32).
transformer = transformer.to(weight_dtype)
if checkpoint_path is not None:
print(f"From checkpoint: {checkpoint_path}")
if checkpoint_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(checkpoint_path)
else:
state_dict = torch.load(checkpoint_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
return transformer
def next_index():
if not os.path.exists(save_path):
return 1
return len([path for path in os.listdir(save_path) if path.endswith("_base.mp4")]) + 1
def save_results(sample, index, prefix, save_fps):
# Both stages of a run share an index: 00000001_base.mp4 / 00000001_refined.mp4.
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
video_path = os.path.join(save_path, f"{str(index).zfill(8)}_{prefix}.mp4")
save_videos_grid(sample, video_path, fps=save_fps)
return video_path
# Get Vae (diffusers-format QwenImage VAE, Wan-style 16ch causal VAE), shared by both stages
vae = AutoencoderKLQwenImage.from_pretrained(
model_name,
subfolder="vae",
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Processor (Qwen3-VL tokenizer + image processor)
processor = AutoProcessor.from_pretrained(
os.path.join(model_name, "processor"),
)
# Get Scheduler
Chosen_Scheduler = {
"Flow_Unipc": FlowUniPCMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
# Stage 0: encode the prompts once. Both stages condition on the same text, so the
# text encoder is released before either 30B DiT is loaded.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
encode_pipeline = LingBotVideoPipeline(
transformer=None,
vae=vae,
text_encoder=text_encoder,
processor=processor,
scheduler=scheduler,
)
encode_pipeline.text_encoder.to(device)
with torch.no_grad():
prompt_embeds, prompt_mask = encode_pipeline.encode_prompt(prompt, device=device)
negative_prompt_embeds, negative_prompt_mask = encode_pipeline.encode_prompt(negative_prompt, device=device)
del encode_pipeline, text_encoder
gc.collect()
torch.cuda.empty_cache()
# Stage 1: base sampling at sample_size
transformer = load_transformer(model_name, transformer_subpath, transformer_path)
pipeline = LingBotVideoPipeline(
transformer=transformer,
vae=vae,
text_encoder=None,
processor=processor,
scheduler=scheduler,
)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['time_embedder', 'time_modulation', 'text_embedder', 'norm', 'router', 'scale_shift_table', 'proj_out'])
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
generator = torch.Generator(device=device).manual_seed(seed)
with torch.no_grad():
video_length = int((video_length - 1) // pipeline.vae_scale_factor_temporal * pipeline.vae_scale_factor_temporal) + 1 if video_length != 1 else 1
base_sample = pipeline(
prompt,
num_frames = video_length,
prompt_embeds = prompt_embeds,
prompt_mask = prompt_mask,
negative_prompt_embeds = negative_prompt_embeds,
negative_prompt_mask = negative_prompt_mask,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
shift = shift,
num_inference_steps = num_inference_steps,
).videos
save_index = next_index()
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
base_path = save_results(base_sample, save_index, "base", fps)
else:
base_path = save_results(base_sample, save_index, "base", fps)
del pipeline, transformer
gc.collect()
torch.cuda.empty_cache()
# Stage 2: refinement at refiner_sample_size. Unlike the reference runner, the base
# frames are refined in memory instead of being re-read from the saved mp4, which
# skips a lossy encode/decode round trip.
refiner = load_transformer(refiner_model_name, refiner_subpath, refiner_path)
refiner_pipeline = LingBotVideoPipeline(
transformer=refiner,
vae=vae,
text_encoder=None,
processor=processor,
scheduler=scheduler,
)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(refiner_pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=["time_embedder", "time_modulation", "text_embedder", "norm", "router", "scale_shift_table", "proj_out"])
if ulysses_degree > 1 or ring_degree > 1:
refiner.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
refiner_pipeline.transformer = shard_fn(refiner_pipeline.transformer)
print("Add FSDP DIT")
refiner_generator = torch.Generator(device=device).manual_seed(seed)
with torch.no_grad():
# video: [B, C, T, H, W] in [0, 1]
bsz, channels, frames, _height, _width = base_sample.shape
flat = base_sample.permute(0, 2, 1, 3, 4).reshape(bsz * frames, channels, _height, _width)
resized = F.interpolate(flat, size=(refiner_sample_size[0], refiner_sample_size[1]), mode="bicubic", align_corners=False).clamp(0.0, 1.0)
lowres_video = resized.reshape(bsz, frames, channels, refiner_sample_size[0], refiner_sample_size[1]).permute(0, 2, 1, 3, 4).contiguous()
x_up = refiner_pipeline.encode_video_latent(lowres_video, generator=refiner_generator)
noise = torch.randn(x_up.shape, device=x_up.device, dtype=x_up.dtype, generator=refiner_generator)
initial_latent = prepare_refiner_latent(x_up, noise, refiner_t_thresh)
del lowres_video, x_up, noise
# The refiner's unconditional branch uses zeroed conditions rather than the
# negative prompt (null_cond_clone_zero in the reference implementation).
refiner_sample = refiner_pipeline(
prompt,
num_frames = video_length,
prompt_embeds = prompt_embeds,
prompt_mask = prompt_mask,
negative_prompt_embeds = torch.zeros_like(prompt_embeds),
negative_prompt_mask = prompt_mask.clone(),
height = refiner_sample_size[0],
width = refiner_sample_size[1],
latents = initial_latent,
generator = refiner_generator,
guidance_scale = refiner_guidance_scale,
shift = refiner_shift,
num_inference_steps = refiner_steps,
t_thresh = refiner_t_thresh,
refiner_sigma_tail_steps = refiner_sigma_tail_steps,
).videos
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
refined_path = save_results(refiner_sample, save_index, "refined", fps)
print(f"base: {base_path}\nrefined: {refined_path}")
else:
refined_path = save_results(refiner_sample, save_index, "refined", fps)
print(f"base: {base_path}\nrefined: {refined_path}")
+343
View File
@@ -0,0 +1,343 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.data.utils import prepare_lingbot_dit_cond_dict
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLWan, AutoencoderKLWan3_8,
AutoTokenizer, WanT5EncoderModel,
WanTransformer3DModel_LingbotWorld)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import Wan2_2I2VPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs,
get_image_to_video_latent, save_videos_grid)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
# TeaCache config
enable_teacache = True
# Recommended to be set between 0.05 and 0.30. A larger threshold can cache more steps, speeding up the inference process,
# but it may cause slight differences between the generated content and the original content.
teacache_threshold = 0.10
# The number of steps to skip TeaCache at the beginning of the inference process, which can
# reduce the impact of TeaCache on generated video quality.
num_skip_start_steps = 5
# Whether to offload TeaCache tensors to cpu to save a little bit of GPU memory.
teacache_offload = False
# Skip some cfg steps in inference
# Recommended to be set between 0.00 and 0.25
cfg_skip_ratio = 0
# Config and model path (the lingbot model reuses the Wan2.2 I2V layout).
config_path = "config/wan2.2/wan_civitai_i2v.yaml"
model_name = "models/Diffusion_Transformer/lingbot-world-base-cam"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow_Unipc"
# [NOTE]: Noise schedule shift parameter. Affects temporal dynamics.
# For 480p generation, a shift of 3.0 is recommended.
shift = 5
# Load pretrained model if need
transformer_path = None
transformer_high_path = None
vae_path = None
lora_path = None
lora_high_path = None
# Other params
sample_size = [480, 832]
video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Camera trajectory (poses.npy / intrinsics.npy) + reference image + prompt.
action_path = "asset/lingbot_demo"
validation_image_start = "asset/lingbot_demo/image.jpg"
# prompts
prompt = "The video presents a soaring journey through a fantasy jungle. The wind whips past the rider's blue hands gripping the reins, causing the leather straps to vibrate. The ancient gothic castle approaches steadily, its stone details becoming clearer against the backdrop of floating islands and distant waterfalls."
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
guidance_scale = 5.0
seed = 43
num_inference_steps = 40
lora_weight = 0.55
lora_high_weight = 0.55
save_path = "samples/lingbot-world-i2v"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
config = OmegaConf.load(config_path)
boundary = config['transformer_additional_kwargs'].get('boundary', 0.900)
transformer = WanTransformer3DModel_LingbotWorld.from_pretrained(
os.path.join(model_name, config['transformer_additional_kwargs'].get('transformer_low_noise_model_subpath', 'transformer')),
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if config['transformer_additional_kwargs'].get('transformer_combination_type', 'single') == "moe":
transformer_2 = WanTransformer3DModel_LingbotWorld.from_pretrained(
os.path.join(model_name, config['transformer_additional_kwargs'].get('transformer_high_noise_model_subpath', 'transformer')),
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
else:
transformer_2 = None
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
if transformer_2 is not None:
if transformer_high_path is not None:
print(f"From checkpoint: {transformer_high_path}")
if transformer_high_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_high_path)
else:
state_dict = torch.load(transformer_high_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer_2.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
Chosen_AutoencoderKL = {
"AutoencoderKLWan": AutoencoderKLWan,
"AutoencoderKLWan3_8": AutoencoderKLWan3_8
}[config['vae_kwargs'].get('vae_type', 'AutoencoderKLWan')]
vae = Chosen_AutoencoderKL.from_pretrained(
os.path.join(model_name, config['vae_kwargs'].get('vae_subpath', 'vae')),
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Tokenizer
tokenizer = AutoTokenizer.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('tokenizer_subpath', 'tokenizer')),
)
# Get Text encoder
text_encoder = WanT5EncoderModel.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('text_encoder_subpath', 'text_encoder')),
additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
if sampler_name == "Flow_Unipc" or sampler_name == "Flow_DPM++":
config['scheduler_kwargs']['shift'] = 1
scheduler = Chosen_Scheduler(
**filter_kwargs(Chosen_Scheduler, OmegaConf.to_container(config['scheduler_kwargs']))
)
# Get Pipeline (reuse the standard Wan2.2 I2V pipeline unchanged).
pipeline = Wan2_2I2VPipeline(
transformer=transformer,
transformer_2=transformer_2,
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if transformer_2 is not None:
transformer_2.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
if transformer_2 is not None:
pipeline.transformer_2 = shard_fn(pipeline.transformer_2)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
if transformer_2 is not None:
for i in range(len(pipeline.transformer_2.blocks)):
pipeline.transformer_2.blocks[i] = torch.compile(pipeline.transformer_2.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed, and both transformers of this MoE setup are handled in one call, which
# is exactly the bookkeeping the old 30-line if/elif chain repeated per script.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
if coefficients is not None:
print(f"Enable TeaCache with threshold {teacache_threshold} and skip the first {num_skip_start_steps} steps.")
pipeline.transformer.enable_teacache(
coefficients, num_inference_steps, teacache_threshold, num_skip_start_steps=num_skip_start_steps, offload=teacache_offload
)
if transformer_2 is not None:
pipeline.transformer_2.share_teacache(transformer=pipeline.transformer)
if cfg_skip_ratio is not None:
print(f"Enable cfg_skip_ratio {cfg_skip_ratio}.")
pipeline.transformer.enable_cfg_skip(cfg_skip_ratio, num_inference_steps)
if transformer_2 is not None:
pipeline.transformer_2.share_cfg_skip(transformer=pipeline.transformer)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
from videox_fun.utils import merge_lora, unmerge_lora
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
if lora_high_path is not None and transformer_2 is not None:
from videox_fun.utils import merge_lora, unmerge_lora
pipeline = merge_lora(pipeline, lora_high_path, lora_high_weight, device=device, dtype=weight_dtype, sub_transformer_name="transformer_2")
with torch.no_grad():
video_length = int((video_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio) + 1 if video_length != 1 else 1
# Prepare the camera condition. The trajectory length may shrink video_length.
dit_cond_dict, video_length = prepare_lingbot_dit_cond_dict(
action_path=action_path,
frame_num=video_length,
height=sample_size[0],
width=sample_size[1],
device=device,
dtype=weight_dtype,
control_type=getattr(transformer, "control_type", "cam"),
vae_stride=(vae.config.temporal_compression_ratio, vae.config.spatial_compression_ratio, vae.config.spatial_compression_ratio),
patch_size=transformer.config.patch_size,
)
# Feed the camera condition to the transformers so the standard pipeline can
# be reused without any modification.
pipeline.transformer.dit_cond_dict = dit_cond_dict
if transformer_2 is not None:
pipeline.transformer_2.dit_cond_dict = dit_cond_dict
latent_frames = (video_length - 1) // vae.config.temporal_compression_ratio + 1
input_video, input_video_mask, clip_image = get_image_to_video_latent(validation_image_start, None, video_length=video_length, sample_size=sample_size)
sample = pipeline(
prompt,
num_frames = video_length,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
num_inference_steps = num_inference_steps,
boundary = boundary,
video = input_video,
mask_video = input_video_mask,
shift = shift,
).videos
# Clear the camera condition after generation.
pipeline.transformer.dit_cond_dict = None
if transformer_2 is not None:
pipeline.transformer_2.dit_cond_dict = None
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
if lora_high_path is not None and transformer_2 is not None:
pipeline = unmerge_lora(pipeline, lora_high_path, lora_high_weight, device=device, dtype=weight_dtype, sub_transformer_name="transformer_2")
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+298
View File
@@ -0,0 +1,298 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLWan, AutoTokenizer,
WanT5EncoderModel,
WanTransformer3DModel_LingbotWorldFast)
from videox_fun.pipeline import WanFunLingbotWorldFastPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
# Config and model path.
# The lingbot-world fast checkpoint only ships the transformer (16 shards at the
# repo root); its VAE / T5 / tokenizer are sourced from the base-cam repo, which
# keeps the raw Wan2.1 layout used by ``config/wan2.1/wan_civitai.yaml``.
config_path = "config/wan2.1/wan_civitai.yaml"
transformer_name = "models/Diffusion_Transformer/lingbot-world-fast"
model_name = "models/Diffusion_Transformer/lingbot-world-base-cam"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++".
# The distilled fast model is trained on the FlowUniPC schedule (int64 timesteps,
# shift applied inside set_timesteps); the tuned ``timesteps_index`` only matches
# that grid, so "Flow_Unipc" is required to reproduce the reference quality.
sampler_name = "Flow_Unipc"
# [NOTE]: Noise schedule shift parameter. Affects temporal dynamics of the
# distilled few-step flow-matching schedule. The official lingbot-world fast
# model is calibrated with sample_shift=10.0 (wan_i2v_A14B.py); generate_fast.py
# passes cfg.sample_shift (=10.0), NOT the generate() signature default of 5.0.
# Using 5.0 here builds the wrong sigma grid ([999,957,899,702] instead of
# [999,978,947,825]), so every renoise step is off-distribution and the decoded
# frames look grainy/painterly. 10.0 reproduces the reference quality.
shift = 10.0
# stochastic_sampling=True selects the native calibrated few-step schedule
# (fixed-index FlowUniPC grid the fast model was distilled on) — the correct
# path. False falls back to the generic scheduler dispatch (off-distribution
# for the fast weights).
stochastic_sampling = True
# Load pretrained transformer weights (optional override).
transformer_path = None
vae_path = None
# LoRA path (optional). The fast model is a single distilled transformer, so only
# one LoRA is used here (no high-noise counterpart like the MoE predict_i2v.py).
lora_path = None
# Camera control type - 'cam' (6-dim plücker).
control_type = "cam"
# Self-Forcing causal inference config.
# `num_frame_per_block`: 3 = chunk-wise:
# 1 = frame-wise:
num_frame_per_block = 3
# Local attention window size (-1 for global attention).
local_attn_size = -1
sink_size = 0
# Other params
# The reference derives the output resolution from a pixel-area budget and the
# input image aspect ratio (image2video_fast.py), rather than forcing a fixed
# size. ``sample_size`` here is only used as the area budget (max_area =
# sample_size[0] * sample_size[1]); the actual height/width are computed per
# image below so the aspect ratio is preserved (no stretch).
sample_size = [480, 832]
video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Camera trajectory (poses.npy / intrinsics.npy) + reference image + prompt.
action_path = "asset/lingbot_demo"
validation_image_start = "asset/lingbot_demo/image.jpg"
# prompts
prompt = "The video presents a soaring journey through a fantasy jungle. The wind whips past the rider's blue hands gripping the reins, causing the leather straps to vibrate. The ancient gothic castle approaches steadily, its stone details becoming clearer against the backdrop of floating islands and distant waterfalls."
# negative_prompt / guidance_scale mirror predict_i2v.py. The distilled few-step
# model is trained WITHOUT classifier-free guidance, so guidance_scale is left at
# 1.0 (CFG disabled, native behavior). Set it > 1.0 to enable CFG with the
# negative prompt (off-distribution, may degrade quality).
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
guidance_scale = 1.0
seed = 43
num_inference_steps = 4
lora_weight = 0.55
save_path = "samples/lingbot-world-i2v-fast"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
config = OmegaConf.load(config_path)
# Load transformer with causal inference + camera control support.
transformer_additional_kwargs = OmegaConf.to_container(config['transformer_additional_kwargs'])
transformer_additional_kwargs['local_attn_size'] = local_attn_size
transformer_additional_kwargs['sink_size'] = sink_size
transformer_additional_kwargs['control_type'] = control_type
transformer_additional_kwargs['cross_attn_type'] = 'cross_attn'
transformer = WanTransformer3DModel_LingbotWorldFast.from_pretrained(
transformer_name,
transformer_additional_kwargs=transformer_additional_kwargs,
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLWan.from_pretrained(
os.path.join(model_name, config['vae_kwargs'].get('vae_subpath', 'vae')),
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Tokenizer
tokenizer = AutoTokenizer.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('tokenizer_subpath', 'tokenizer')),
)
# Get Text encoder
text_encoder = WanT5EncoderModel.from_pretrained(
os.path.join(model_name, config['text_encoder_kwargs'].get('text_encoder_subpath', 'text_encoder')),
additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler_kwargs = OmegaConf.to_container(config['scheduler_kwargs'])
# The lingbot-world fast reference builds FlowUniPCMultistepScheduler with shift=1
# and applies the real shift only inside set_timesteps. The shared config carries
# shift=5.0 (used by other pipelines), which would double-shift the sigma grid
# here, so pin the constructor shift to 1 for the UniPC path; the runtime shift
# is forwarded to set_timesteps by the pipeline instead.
if Chosen_Scheduler is FlowUniPCMultistepScheduler:
scheduler_kwargs['shift'] = 1
scheduler = Chosen_Scheduler(
**filter_kwargs(Chosen_Scheduler, scheduler_kwargs)
)
# Get Pipeline
pipeline = WanFunLingbotWorldFastPipeline(
transformer=transformer,
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
video_length = int((video_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio) + 1 if video_length != 1 else 1
image = Image.open(validation_image_start).convert("RGB")
# Aspect-preserving resolution from the area budget (matches the reference):
# keep width*height close to max_area while snapping to the VAE stride *
# patch-size grid, so the input aspect ratio is preserved instead of forcing
# a fixed sample_size (which would stretch a 16:9 frame).
max_area = sample_size[0] * sample_size[1]
vae_stride = vae.config.spatial_compression_ratio
patch = pipeline.transformer.config.patch_size[1]
aspect_ratio = image.height / image.width
lat_h = round(np.sqrt(max_area * aspect_ratio) // vae_stride // patch * patch)
lat_w = round(np.sqrt(max_area / aspect_ratio) // vae_stride // patch * patch)
height = lat_h * vae_stride
width = lat_w * vae_stride
sample = pipeline(
prompt,
image = image,
negative_prompt = negative_prompt,
action_path = action_path,
control_type = control_type,
height = height,
width = width,
num_frames = video_length,
num_frame_per_block = num_frame_per_block,
num_inference_steps = num_inference_steps,
guidance_scale = guidance_scale,
stochastic_sampling = stochastic_sampling,
shift = shift,
generator = generator,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+49 -36
View File
@@ -4,7 +4,6 @@ import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
@@ -13,19 +12,16 @@ for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLongCatVideo, UMT5EncoderModel, AutoTokenizer,
LongCatVideoTransformer3DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.models import (AutoencoderKLLongCatVideo, AutoTokenizer,
LongCatVideoTransformer3DModel,
UMT5EncoderModel)
from videox_fun.pipeline import LongCatVideoPipeline
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
save_videos_grid)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image_to_video_latent,
merge_lora, save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -36,9 +32,21 @@ from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
@@ -60,7 +68,7 @@ video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
validation_image_start = "asset/1.png"
@@ -69,11 +77,11 @@ prompt = "The dog is shaking head. The video is of high quality, an
negative_prompt = "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
guidance_scale = 4.0
seed = 43
num_inference_steps = 50
num_inference_steps = 25
lora_weight = 0.55
save_path = "samples/longcat-videos-i2v"
device = set_multi_gpus_devices(1, 1)
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = LongCatVideoTransformer3DModel.from_pretrained(
os.path.join(model_name, "dit"),
@@ -84,7 +92,7 @@ transformer = LongCatVideoTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -101,7 +109,7 @@ vae = AutoencoderKLLongCatVideo.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -123,7 +131,7 @@ text_encoder = UMT5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -141,28 +149,27 @@ pipeline = LongCatVideoPipeline(
text_encoder=text_encoder,
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.encoder.block)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
transformer.freqs = transformer.freqs.to(device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
@@ -208,8 +215,14 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
save_results()
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+273
View File
@@ -0,0 +1,273 @@
import os
import sys
from pathlib import Path
import numpy as np
import torch
from audio_separator.separator import Separator
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLongCatVideo, AutoTokenizer,
LongCatVideoAudioEncoder,
LongCatVideoAvatarTransformer3DModel,
UMT5EncoderModel)
from videox_fun.pipeline import LongCatVideoAvatarPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, get_image_to_video_latent,
merge_lora, merge_video_audio, save_videos_grid,
unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LongCat-Video"
model_name_avatar = "models/Diffusion_Transformer/LongCat-Video-Avatar"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [832, 480]
video_length = 81
fps = 16
# Start Image
validation_image_start = "asset/8.png"
# Audio params
audio_path = "asset/talk.wav"
use_audio_vocal_separator = False
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Prompt
prompt = "A young woman with long flowing purple hair stands by the seaside on a sunny day, singing. Wearing a white sleeveless dress with a navy blue bow at the collar, her hair gently sways in the ocean breeze. The sparkling sea, blue sky with white clouds, and pink wildflowers along the shore create a beautiful and vibrant scene."
negative_prompt = "Close-up, Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
guidance_scale = 4.5
seed = 43
num_inference_steps = 25
lora_weight = 0.55
save_path = "samples/longcat-avatar-videos-t2v"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = LongCatVideoAvatarTransformer3DModel.from_pretrained(
os.path.join(model_name_avatar, "avatar_single"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype, cp_split_hw=[1, 1]
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Vae
vae = AutoencoderKLLongCatVideo.from_pretrained(
os.path.join(model_name, "vae"),
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Tokenizer
tokenizer = AutoTokenizer.from_pretrained(
os.path.join(model_name, "tokenizer"),
)
# Get Text encoder
text_encoder = UMT5EncoderModel.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Get Audio encoder (for avatar mode)
audio_encoder = LongCatVideoAudioEncoder(
os.path.join(model_name_avatar, 'chinese-wav2vec2-base')
)
audio_encoder.audio_encoder.feature_extractor._freeze_parameters()
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
# Get Pipeline
pipeline = LongCatVideoAvatarPipeline(
transformer=transformer,
vae=vae,
tokenizer=tokenizer,
text_encoder=text_encoder,
scheduler=scheduler,
audio_encoder=audio_encoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.encoder.block)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
# Get Vocal separator
if use_audio_vocal_separator:
vocal_separator_path = os.path.join(model_name_avatar, 'vocal_separator/Kim_Vocal_2.onnx')
audio_output_dir_temp = Path("./audio_temp_file")
audio_output_dir_temp.mkdir(parents=True, exist_ok=True)
vocal_separator = Separator(
output_dir=audio_output_dir_temp / "vocals",
output_single_stem="vocals",
model_file_dir=os.path.dirname(vocal_separator_path),
)
vocal_separator.load_model(os.path.basename(vocal_separator_path))
# Process audio if provided
audio_emb = None
if audio_path is not None:
# Extract vocal from audio
outputs = vocal_separator.separate(audio_path)
if len(outputs) > 0:
temp_vocal_path = audio_output_dir_temp / "vocals" / outputs[0]
temp_vocal_path = temp_vocal_path.resolve().as_posix()
audio_path = temp_vocal_path
with torch.no_grad():
video_length = int((video_length - 1) // vae.scale_factor_temporal * vae.scale_factor_temporal) + 1 if video_length != 1 else 1
latent_frames = (video_length - 1) // vae.scale_factor_temporal + 1
if validation_image_start is not None:
input_video, input_video_mask, clip_image = get_image_to_video_latent(validation_image_start, None, video_length=video_length, sample_size=sample_size)
else:
input_video, input_video_mask, clip_image = None, None, None
sample = pipeline(
prompt = prompt,
num_frames = video_length,
negative_prompt = negative_prompt,
height = sample_size[0],
width = sample_size[1],
generator = generator,
guidance_scale = guidance_scale,
num_inference_steps = num_inference_steps,
audio_path = audio_path,
video = input_video,
mask_video = input_video_mask,
fps = fps,
).videos
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
merge_video_audio(video_path=video_path, audio_path=audio_path)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+50 -36
View File
@@ -4,7 +4,6 @@ import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from omegaconf import OmegaConf
from PIL import Image
current_file_path = os.path.abspath(__file__)
@@ -13,19 +12,16 @@ for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLongCatVideo, UMT5EncoderModel, AutoTokenizer,
LongCatVideoTransformer3DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.models import (AutoencoderKLLongCatVideo, AutoTokenizer,
LongCatVideoTransformer3DModel,
UMT5EncoderModel)
from videox_fun.pipeline import LongCatVideoPipeline
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8, replace_parameters_by_name,
convert_weight_dtype_wrapper)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
save_videos_grid)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -36,9 +32,21 @@ from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with the fsdp_dit and sequential_cpu_offload.
compile_dit = False
@@ -60,18 +68,18 @@ video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Prompt
prompt = "1girl, black_hair, brown_eyes, earrings, freckles, grey_background, jewelry, lips, long_hair, looking_at_viewer, nose, piercing, realistic, red_lips, solo, upper_body"
negative_prompt = "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
guidance_scale = 4.0
seed = 43
num_inference_steps = 50
num_inference_steps = 25
lora_weight = 0.55
save_path = "samples/longcat-videos-t2v"
device = set_multi_gpus_devices(1, 1)
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
transformer = LongCatVideoTransformer3DModel.from_pretrained(
os.path.join(model_name, "dit"),
@@ -82,7 +90,7 @@ transformer = LongCatVideoTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -99,7 +107,7 @@ vae = AutoencoderKLLongCatVideo.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -121,7 +129,7 @@ text_encoder = UMT5EncoderModel.from_pretrained(
)
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -140,27 +148,27 @@ pipeline = LongCatVideoPipeline(
scheduler=scheduler,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.encoder.block)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
transformer.freqs = transformer.freqs.to(device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
@@ -199,8 +207,14 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)
save_results()
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+286
View File
@@ -0,0 +1,286 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLTX2Audio, AutoencoderKLLTX2Video,
Gemma3ForConditionalGeneration, Gemma3Processor,
LTX2TextConnectors, LTX2VideoTransformer3DModel,
LTX2VocoderWithBWE)
from videox_fun.pipeline import LTX2I2VPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LTX-2.3-Diffusers"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [512, 768]
video_length = 121
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None and validation_image_end = None
validation_image_start = "asset/1.png"
# prompts
prompt = "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room. Behind the dog, there is a framed painting on a shelf, surrounded by pink flowers. "
negative_prompt = "worst quality, inconsistent motion, blurry, jittery, distorted, static, low quality, artifacts"
# CFG guidance scale for video and audio modality
guidance_scale = 3.0
audio_guidance_scale = 7.0
# Spatio-Temporal Guidance (STG) scale for video and audio
stg_scale = 1.0
audio_stg_scale = 1.0
# Modality isolation guidance scale for video and audio
modality_scale = 3.0
audio_modality_scale = 3.0
# Guidance rescale factor for video and audio to prevent overexposure
guidance_rescale = 0.7
audio_guidance_rescale = 0.7
spatio_temporal_guidance_blocks = [28]
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/ltx2-videos-i2v"
# Audio sample rate will be read from vocoder config
audio_sample_rate = 24000
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Transformer
transformer = LTX2VideoTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
vae = AutoencoderKLLTX2Video.from_pretrained(
model_name,
subfolder="vae",
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE
audio_vae = AutoencoderKLLTX2Audio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=weight_dtype,
)
# Get Processor
processor = Gemma3Processor.from_pretrained(
model_name,
subfolder="processor",
)
# Get Tokenizer
tokenizer = processor.tokenizer
# Get Text encoder
text_encoder = Gemma3ForConditionalGeneration.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Connectors
connectors = LTX2TextConnectors.from_pretrained(
model_name,
subfolder="connectors",
torch_dtype=weight_dtype,
)
# Vocoder
vocoder = LTX2VocoderWithBWE.from_pretrained(
model_name,
subfolder="vocoder",
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LTX2I2VPipeline(
scheduler=scheduler,
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
connectors=connectors,
transformer=transformer,
vocoder=vocoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=text_encoder.language_model.layers)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['scale_shift_table', 'audio_scale_shift_table', 'video_a2v_cross_attn_scale_shift_table', 'audio_a2v_cross_attn_scale_shift_table', ''])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
output = pipeline(
image=Image.open(validation_image_start),
prompt=prompt,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
stg_scale=stg_scale,
modality_scale=modality_scale,
guidance_rescale=guidance_rescale,
audio_guidance_scale=audio_guidance_scale,
audio_stg_scale=audio_stg_scale,
audio_modality_scale=audio_modality_scale,
audio_guidance_rescale=audio_guidance_rescale,
spatio_temporal_guidance_blocks=spatio_temporal_guidance_blocks,
generator=generator,
output_type="pt",
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.vocoder.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+281
View File
@@ -0,0 +1,281 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLTX2Audio, AutoencoderKLLTX2Video,
Gemma3ForConditionalGeneration, Gemma3Processor,
LTX2TextConnectors, LTX2VideoTransformer3DModel,
LTX2VocoderWithBWE)
from videox_fun.pipeline import LTX2Pipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LTX-2.3-Diffusers"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [512, 768]
video_length = 121
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room. Behind the dog, there is a framed painting on a shelf, surrounded by pink flowers. "
negative_prompt = "worst quality, inconsistent motion, blurry, jittery, distorted, static, low quality, artifacts"
# CFG guidance scale for video and audio modality
guidance_scale = 3.0
audio_guidance_scale = 7.0
# Spatio-Temporal Guidance (STG) scale for video and audio
stg_scale = 1.0
audio_stg_scale = 1.0
# Modality isolation guidance scale for video and audio
modality_scale = 3.0
audio_modality_scale = 3.0
# Guidance rescale factor for video and audio to prevent overexposure
guidance_rescale = 0.7
audio_guidance_rescale = 0.7
spatio_temporal_guidance_blocks = [28]
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/ltx2-videos-t2v"
# Audio sample rate will be read from vocoder config
audio_sample_rate = 24000
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Transformer
transformer = LTX2VideoTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
vae = AutoencoderKLLTX2Video.from_pretrained(
model_name,
subfolder="vae",
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE
audio_vae = AutoencoderKLLTX2Audio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=weight_dtype,
)
# Get Processor
processor = Gemma3Processor.from_pretrained(
model_name,
subfolder="processor",
)
# Get Tokenizer
tokenizer = processor.tokenizer
# Get Text encoder
text_encoder = Gemma3ForConditionalGeneration.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Connectors
connectors = LTX2TextConnectors.from_pretrained(
model_name,
subfolder="connectors",
torch_dtype=weight_dtype,
)
# Vocoder
vocoder = LTX2VocoderWithBWE.from_pretrained(
model_name,
subfolder="vocoder",
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LTX2Pipeline(
scheduler=scheduler,
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
connectors=connectors,
transformer=transformer,
vocoder=vocoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=text_encoder.language_model.layers)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['scale_shift_table', 'audio_scale_shift_table', 'video_a2v_cross_attn_scale_shift_table', 'audio_a2v_cross_attn_scale_shift_table', ''])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
output = pipeline(
prompt=prompt,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
stg_scale=stg_scale,
modality_scale=modality_scale,
guidance_rescale=guidance_rescale,
audio_guidance_scale=audio_guidance_scale,
audio_stg_scale=audio_stg_scale,
audio_modality_scale=audio_modality_scale,
audio_guidance_rescale=audio_guidance_rescale,
spatio_temporal_guidance_blocks=spatio_temporal_guidance_blocks,
generator=generator,
output_type="pt",
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.vocoder.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+262
View File
@@ -0,0 +1,262 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLTX2Audio, AutoencoderKLLTX2Video,
Gemma3ForConditionalGeneration,
GemmaTokenizerFast, LTX2TextConnectors,
LTX2VideoTransformer3DModel, LTX2Vocoder)
from videox_fun.pipeline import LTX2I2VPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LTX-2"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [480, 832]
video_length = 121
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None and validation_image_end = None
validation_image_start = "asset/1.png"
# prompts
prompt = "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room. Behind the dog, there is a framed painting on a shelf, surrounded by pink flowers. "
negative_prompt = "worst quality, inconsistent motion, blurry, jittery, distorted, static, low quality, artifacts"
guidance_scale = 6.0
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/ltx2-videos-i2v"
# Audio sample rate will be read from vocoder config
audio_sample_rate = 24000
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Transformer
transformer = LTX2VideoTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
vae = AutoencoderKLLTX2Video.from_pretrained(
model_name,
subfolder="vae",
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE
audio_vae = AutoencoderKLLTX2Audio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=weight_dtype,
)
# Get Tokenizer
tokenizer = GemmaTokenizerFast.from_pretrained(
model_name,
subfolder="tokenizer",
)
# Get Text encoder
text_encoder = Gemma3ForConditionalGeneration.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Connectors
connectors = LTX2TextConnectors.from_pretrained(
model_name,
subfolder="connectors",
torch_dtype=weight_dtype,
)
# Vocoder
vocoder = LTX2Vocoder.from_pretrained(
model_name,
subfolder="vocoder",
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LTX2I2VPipeline(
scheduler=scheduler,
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
connectors=connectors,
transformer=transformer,
vocoder=vocoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=text_encoder.language_model.layers)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['scale_shift_table', 'audio_scale_shift_table', 'video_a2v_cross_attn_scale_shift_table', 'audio_a2v_cross_attn_scale_shift_table', ''])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
output = pipeline(
image=Image.open(validation_image_start),
prompt=prompt,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
generator=generator,
output_type="pt",
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.vocoder.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+307
View File
@@ -0,0 +1,307 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLTX2Audio, AutoencoderKLLTX2Video,
Gemma3ForConditionalGeneration,
GemmaTokenizerFast, LTX2LatentUpsamplerModel,
LTX2TextConnectors, LTX2VideoTransformer3DModel,
LTX2Vocoder)
from videox_fun.pipeline import LTX2I2VPipeline, LTX2LatentUpsamplePipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LTX-2"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
latent_upsampler_path = None
# Other params
sample_size = [480, 832]
video_length = 121
fps = 24
# Latent upsampler config
enable_latent_upsample = True
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# If you want to generate from text, please set the validation_image_start = None and validation_image_end = None
validation_image_start = "asset/1.png"
# prompts
prompt = "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room. Behind the dog, there is a framed painting on a shelf, surrounded by pink flowers. "
negative_prompt = "worst quality, inconsistent motion, blurry, jittery, distorted, static, low quality, artifacts"
guidance_scale = 6.0
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/ltx2-videos-i2v"
# Audio sample rate will be read from vocoder config
audio_sample_rate = 24000
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Transformer
transformer = LTX2VideoTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
vae = AutoencoderKLLTX2Video.from_pretrained(
model_name,
subfolder="vae",
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE
audio_vae = AutoencoderKLLTX2Audio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=weight_dtype,
)
# Get Tokenizer
tokenizer = GemmaTokenizerFast.from_pretrained(
model_name,
subfolder="tokenizer",
)
# Get Text encoder
text_encoder = Gemma3ForConditionalGeneration.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Connectors
connectors = LTX2TextConnectors.from_pretrained(
model_name,
subfolder="connectors",
torch_dtype=weight_dtype,
)
# Vocoder
vocoder = LTX2Vocoder.from_pretrained(
model_name,
subfolder="vocoder",
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LTX2I2VPipeline(
scheduler=scheduler,
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
connectors=connectors,
transformer=transformer,
vocoder=vocoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=text_encoder.language_model.layers)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['scale_shift_table', 'audio_scale_shift_table', 'video_a2v_cross_attn_scale_shift_table', 'audio_a2v_cross_attn_scale_shift_table', ''])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
output = pipeline(
image=Image.open(validation_image_start),
prompt=prompt,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
generator=generator,
output_type="latent" if enable_latent_upsample else "pt",
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
if enable_latent_upsample:
# Load latent upsampler model
latent_upsampler = LTX2LatentUpsamplerModel.from_pretrained(
model_name, subfolder="latent_upsampler", torch_dtype=weight_dtype,
)
if latent_upsampler_path is not None:
print(f"From latent_upsampler checkpoint: {latent_upsampler_path}")
if latent_upsampler_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(latent_upsampler_path)
else:
state_dict = torch.load(latent_upsampler_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = latent_upsampler.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
upsample_pipeline = LTX2LatentUpsamplePipeline(
vae=pipeline.vae,
latent_upsampler=latent_upsampler,
)
upsample_pipeline.vae.enable_tiling()
upsample_pipeline.to(device=device, dtype=weight_dtype)
# output_type="latent" returns denormalized (raw) video latents [B, C, F, H, W]
# and raw audio latents [B, C, L, M]; decode audio manually
audio_latents = output.audio.to(device=device, dtype=pipeline.audio_vae.dtype)
mel = pipeline.audio_vae.decode(audio_latents, return_dict=False)[0]
audio = pipeline.vocoder(mel).cpu().float()
# Pass video latents directly to upsample pipeline (skip decode→re-encode roundtrip)
with torch.no_grad():
upsampled = upsample_pipeline(
latents=output.videos,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
output_type="pt",
return_dict=False,
)
sample = upsampled[0]
else:
sample = output.videos
audio = output.audio
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.vocoder.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+257
View File
@@ -0,0 +1,257 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLLTX2Audio, AutoencoderKLLTX2Video,
Gemma3ForConditionalGeneration,
GemmaTokenizerFast, LTX2TextConnectors,
LTX2VideoTransformer3DModel, LTX2Vocoder)
from videox_fun.pipeline import LTX2Pipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = False
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/LTX-2"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
# Load pretrained model if need
transformer_path = None
vae_path = None
lora_path = None
# Other params
sample_size = [512, 768]
video_length = 121
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "A brown dog barks on a sofa, sitting on a light-colored couch in a cozy room. Behind the dog, there is a framed painting on a shelf, surrounded by pink flowers. "
negative_prompt = "worst quality, inconsistent motion, blurry, jittery, distorted, static, low quality, artifacts"
guidance_scale = 6.0
seed = 43
num_inference_steps = 50
lora_weight = 0.55
save_path = "samples/ltx2-videos-t2v"
# Audio sample rate will be read from vocoder config
audio_sample_rate = 24000
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# Transformer
transformer = LTX2VideoTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
vae = AutoencoderKLLTX2Video.from_pretrained(
model_name,
subfolder="vae",
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE
audio_vae = AutoencoderKLLTX2Audio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=weight_dtype,
)
# Get Tokenizer
tokenizer = GemmaTokenizerFast.from_pretrained(
model_name,
subfolder="tokenizer",
)
# Get Text encoder
text_encoder = Gemma3ForConditionalGeneration.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Connectors
connectors = LTX2TextConnectors.from_pretrained(
model_name,
subfolder="connectors",
torch_dtype=weight_dtype,
)
# Vocoder
vocoder = LTX2Vocoder.from_pretrained(
model_name,
subfolder="vocoder",
torch_dtype=weight_dtype,
)
# Get Scheduler
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
pipeline = LTX2Pipeline(
scheduler=scheduler,
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
connectors=connectors,
transformer=transformer,
vocoder=vocoder,
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(transformer.transformer_blocks))
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=text_encoder.language_model.layers)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype, exclude_module_name=['scale_shift_table', 'audio_scale_shift_table', 'video_a2v_cross_attn_scale_shift_table', 'audio_a2v_cross_attn_scale_shift_table', ''])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
with torch.no_grad():
output = pipeline(
prompt=prompt,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
generator=generator,
output_type="pt",
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.vocoder.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+295
View File
@@ -0,0 +1,295 @@
import os
import sys
import torch
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMiniMaxH3,
AutoencoderKLMiniMaxH3Audio,
MiniMaxH3Transformer3DModel, Qwen2TokenizerFast,
Qwen3VLForConditionalGeneration,
Qwen3VLProcessor)
from videox_fun.pipeline import MiniMaxH3Pipeline
from videox_fun.utils import (MiniMaxH3Scheduler, apply_gpu_memory_mode,
convert_model_weight_to_float8, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus. The Qwen3-VL conditioner is ~62 GB, so with fsdp_dit alone every
# rank still replicates it; fsdp_text_encoder shards it too. Note it must wrap the inner `text_encoder.model`
# (Qwen3VLModel): encode_prompt calls that submodule directly, so a wrap on the top-level module would never fire.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MiniMax-H3"
# Load pretrained model if need
# A full finetune goes in `transformer_path`, either as the `transformer` folder a training checkpoint writes
# (`output_dir_minimax_h3/checkpoint-N/transformer`, config.json included) or as a single safetensors file. A LoRA
# goes in `lora_path`: handed to `transformer_path` it would match no key at all and load nothing. A PDD LoRA
# (parallel decoder) goes in `pdd_lora_path` and cannot be combined with `lora_path`.
transformer_path = None
vae_path = None
lora_path = None
pdd_lora_path = None
# Other params
# MiniMax-H3 generates at a fixed 24 fps, only accepts multiples of 32 as height / width, and snaps video_length up
# to the next 17 * n + 5 the video VAE can decode (the duration has to stay between 5 and 15 seconds).
# Leave sample_size as None to follow the aspect ratio of the first keyframe, which is what the model was released
# for; the first keyframe is stretched onto that canvas.
sample_size = [704, 1280]
video_length = 124
fps = 24
# The keyframe the video starts from, and the one it ends on. Either can be left as None: only an end frame generates
# *up to* that frame, and neither of them is a plain text-to-video request (see predict_t2v.py).
validation_image_start = "asset/1.png"
validation_image_end = None
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "一只棕色的狗摇着头,坐在舒适房间里的浅色沙发上。在狗的后面,架子上有一幅镶框的画,周围是粉红色的花朵。房间里柔和温暖的灯光营造出舒适的氛围。"
seed = 43
# Number of denoising steps, i.e. of model evaluations: num_inference_steps = 40 runs 40 of them.
num_inference_steps = 40
# The released checkpoint is guidance-distilled: leave guidance_scale at 1 to run one forward pass per step
# with no CFG. A value above 1 enables classifier-free guidance with a negative_prompt, running two passes.
guidance_scale = 1
# The exponential sigma shifts of the two schedules. None keeps the ones of the checkpoint (12.0 video, 3.0 audio).
flow_shift = None
audio_flow_shift = None
lora_weight = 0.55
save_path = "samples/minimax-h3-videos-i2v"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# `model_name` may point either at a converted diffusers layout or at an *original* MiniMax-H3 partition (e.g.
# `MiniMax-H3/FL2VA`); the original shards are converted on the fly while loading, no intermediate copy on disk.
# Transformer
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if os.path.isdir(transformer_path):
# A training checkpoint's `transformer` folder carries its own config.json, so the loader restores the
# mixed-precision contract of the checkpoint (`_keep_in_fp32_modules`) by itself.
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
transformer_path,
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
else:
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# `strict=False` accepts a file whose keys belong to another model — a LoRA checkpoint, say — by loading
# nothing at all and silently generating with the base weights, so an unexpected key is a hard error.
assert len(u) == 0, (
f"{transformer_path} holds {len(u)} key(s) the transformer does not have, e.g. {u[:3]}. A LoRA "
"checkpoint belongs in `lora_path`, not `transformer_path`."
)
pdd_config = None
if pdd_lora_path is not None:
if lora_path is not None:
raise ValueError("`lora_path` and `pdd_lora_path` cannot be used together.")
from videox_fun.models.minimax_h3_pdd import (load_pdd_lora,
pdd_num_inference_steps,
pdd_step_callback)
pdd_config = load_pdd_lora(transformer, pdd_lora_path)
num_inference_steps = pdd_num_inference_steps(pdd_config, num_inference_steps, teacher_default=40)
# Video VAE. The released weights are float32 and the decode runs under float16 autocast, so the VAE is not
# downcast even when the rest of the pipeline is bfloat16 (this is also how the training scripts load it).
vae = AutoencoderKLMiniMaxH3.from_pretrained(
model_name,
subfolder="vae",
low_cpu_mem_usage=True,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE, waveform in / waveform out: MiniMax-H3 has no separate vocoder. Float32 as released, like the video VAE.
audio_vae = AutoencoderKLMiniMaxH3Audio.from_pretrained(
model_name,
subfolder="audio_vae",
low_cpu_mem_usage=True,
)
# Get Tokenizer and Processor
tokenizer = Qwen2TokenizerFast.from_pretrained(os.path.join(model_name, "tokenizer"))
processor = Qwen3VLProcessor.from_pretrained(os.path.join(model_name, "processor"))
# Get Text encoder. MiniMax-H3 reads the unnormalized hidden state after the 50th decoder layer of Qwen3-VL.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Schedulers. MiniMax-H3 steps the video and the audio latents down two schedules inside one transformer call.
scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="scheduler")
audio_scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="audio_scheduler")
pipeline = MiniMaxH3Pipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
transformer=transformer,
scheduler=scheduler,
audio_scheduler=audio_scheduler,
)
# The float32 modules of the mixed-precision checkpoint stay untouched by the float8 quantization.
fp8_exclude_module_name = [
"proj_in", "audio_proj_in", "context_embedder", "time_embedder", "time_proj",
"token_refiner", "norm_out", "proj_out", "audio_proj_out",
]
use_qfloat8 = "qfloat8" in GPU_memory_mode
if use_qfloat8:
convert_model_weight_to_float8(transformer, exclude_module_name=fp8_exclude_module_name, device=device)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
fp32_modules = [m for m in transformer.modules()
if any(p.dtype == torch.float32 for p in m.parameters(recurse=False))]
shard_fn = partial(shard_model, device_id=device, param_dtype=None, cast_dtype=False,
module_to_wrapper=list(transformer.transformer_blocks),
ignored_modules=fp32_modules)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(text_encoder.model.language_model.layers))
pipeline.text_encoder.model = shard_fn(pipeline.text_encoder.model)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
# The FP8 conversion above has already run (before the FSDP sharding, on purpose); only the dequant
# wrapper and the memory placement are left, which is what the preconverted tag installs.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype,
quant_tag="qfloat8_preconverted" if GPU_memory_mode.endswith("_and_qfloat8") else None,
exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
image_start = None if validation_image_start is None else Image.open(validation_image_start)
image_end = None if validation_image_end is None else Image.open(validation_image_end)
pdd_callback = None if pdd_config is None else pdd_step_callback(
transformer, scheduler, audio_scheduler, pdd_config, num_inference_steps
)
with torch.no_grad():
output = pipeline(
prompt=prompt,
image=image_start,
last_image=image_end,
height=None if sample_size is None else sample_size[0],
width=None if sample_size is None else sample_size[1],
num_frames=video_length,
num_inference_steps=num_inference_steps,
flow_shift=flow_shift,
audio_flow_shift=audio_flow_shift,
guidance_scale=guidance_scale,
generator=generator,
output_type="pt",
callback_on_step_end=pdd_callback,
)
print(f"[{os.environ.get('RANK', '0')}] generation done, decoding", flush=True)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
audio_sample_rate = output.sampling_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=audio_sample_rate)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
# Keep every rank alive until the saving rank finishes; an early exit of one rank makes the elastic launcher
# terminate the others.
dist.barrier()
else:
save_results()
+315
View File
@@ -0,0 +1,315 @@
import os
import sys
import torch
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMiniMaxH3,
AutoencoderKLMiniMaxH3Audio,
MiniMaxH3Transformer3DModel, Qwen2TokenizerFast,
Qwen3VLForConditionalGeneration,
Qwen3VLProcessor)
from videox_fun.pipeline import (MiniMaxH3AudioReference,
MiniMaxH3ImageReference, MiniMaxH3Pipeline,
MiniMaxH3VideoReference)
from videox_fun.utils import (MiniMaxH3Scheduler, apply_gpu_memory_mode,
convert_model_weight_to_float8, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus. The Qwen3-VL conditioner is ~62 GB, so with fsdp_dit alone every
# rank still replicates it; fsdp_text_encoder shards it too. Note it must wrap the inner `text_encoder.model`
# (Qwen3VLModel): encode_prompt calls that submodule directly, so a wrap on the top-level module would never fire.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MiniMax-H3"
# Load pretrained model if need
# The `ref2va` weights ship in their own subfolder, same architecture as the base transformer. A full finetune goes
# in `transformer_path`, either as the `transformer` folder a training checkpoint writes (config.json included) or
# as a single safetensors file, overriding the `transformer_ref` subfolder. A LoRA goes in `lora_path`: handed to
# `transformer_path` it would match no key at all and load nothing. A PDD LoRA (parallel decoder) goes in
# `pdd_lora_path` and cannot be combined with `lora_path`; use a checkpoint trained with `--train_mode=ref2va`.
transformer_subfolder = "transformer_ref"
transformer_path = None
vae_path = None
lora_path = None
pdd_lora_path = None
# Other params
# MiniMax-H3 generates at a fixed 24 fps, only accepts multiples of 32 as height / width, and snaps video_length up
# to the next 17 * n + 5 the video VAE can decode (the duration has to stay between 5 and 15 seconds). References
# never bind the generated geometry: leaving height / width unset resolves MiniMax-H3's own 16:9 canvas.
sample_size = [1280, 704]
video_length = 124
fps = 24
# The references to condition on, **in the order the model should read them**: the order labels them in the prompt
# presentation and lays them out on the shared rotary clock. One entry per reference, `image=path`, `video=path` or
# `audio=path`; a video's own soundtrack is conditioned on with it. Budgets of the released checkpoint: at most 9
# images, 3 videos, 3 audios and 12 references in total, and an audio reference cannot stand alone.
references = [
"video=asset/ref2va_video.mp4",
"audio=asset/ref2va_audio.wav",
]
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "参考视频中的角色与场景,生成一段动作连贯、镜头流畅的续写视频,环境音与画面同步。"
seed = 43
# Number of denoising steps, i.e. of model evaluations: num_inference_steps = 50 runs 50 of them.
num_inference_steps = 50
# The released `ref2va` checkpoint is guidance-distilled with no unconditional branch, so `references` runs one
# forward pass per step and needs guidance_scale of 1 — the pipeline raises on anything above.
guidance_scale = 1.0
# The exponential sigma shifts of the two schedules. None keeps the ones of the checkpoint (12.0 video, 3.0 audio).
flow_shift = None
audio_flow_shift = None
lora_weight = 0.55
save_path = "samples/minimax-h3-videos-ref2va"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# `model_name` may point either at a converted diffusers layout or at an *original* MiniMax-H3 partition; the
# original shards are converted on the fly while loading, no intermediate copy on disk. The transformer comes from
# the `transformer_ref` subfolder — the released `ref2va` weights, same architecture as the base model.
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
model_name,
subfolder=transformer_subfolder,
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if os.path.isdir(transformer_path):
# A training checkpoint's `transformer` folder carries its own config.json, so the loader restores the
# mixed-precision contract of the checkpoint (`_keep_in_fp32_modules`) by itself.
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
transformer_path,
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
else:
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# `strict=False` accepts a file whose keys belong to another model — a LoRA checkpoint, say — by loading
# nothing at all and silently generating with the base weights, so an unexpected key is a hard error.
assert len(u) == 0, (
f"{transformer_path} holds {len(u)} key(s) the transformer does not have, e.g. {u[:3]}. A LoRA "
"checkpoint belongs in `lora_path`, not `transformer_path`."
)
pdd_config = None
if pdd_lora_path is not None:
if lora_path is not None:
raise ValueError("`lora_path` and `pdd_lora_path` cannot be used together.")
from videox_fun.models.minimax_h3_pdd import (load_pdd_lora,
pdd_num_inference_steps,
pdd_step_callback)
pdd_config = load_pdd_lora(transformer, pdd_lora_path)
num_inference_steps = pdd_num_inference_steps(pdd_config, num_inference_steps, teacher_default=50)
# Video VAE. The released weights are float32 and the decode runs under float16 autocast, so the VAE is not
# downcast even when the rest of the pipeline is bfloat16 (this is also how the training scripts load it).
vae = AutoencoderKLMiniMaxH3.from_pretrained(
model_name,
subfolder="vae",
low_cpu_mem_usage=True,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE, waveform in / waveform out: MiniMax-H3 has no separate vocoder. Float32 as released, like the video VAE.
audio_vae = AutoencoderKLMiniMaxH3Audio.from_pretrained(
model_name,
subfolder="audio_vae",
low_cpu_mem_usage=True,
)
# Get Tokenizer and Processor
tokenizer = Qwen2TokenizerFast.from_pretrained(os.path.join(model_name, "tokenizer"))
processor = Qwen3VLProcessor.from_pretrained(os.path.join(model_name, "processor"))
# Get Text encoder. MiniMax-H3 reads the unnormalized hidden state after the 50th decoder layer of Qwen3-VL.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Schedulers. MiniMax-H3 steps the video and the audio latents down two schedules inside one transformer call.
scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="scheduler")
audio_scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="audio_scheduler")
pipeline = MiniMaxH3Pipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
transformer=transformer,
scheduler=scheduler,
audio_scheduler=audio_scheduler,
)
# The float32 modules of the mixed-precision checkpoint stay untouched by the float8 quantization.
fp8_exclude_module_name = [
"proj_in", "audio_proj_in", "context_embedder", "time_embedder", "time_proj",
"token_refiner", "norm_out", "proj_out", "audio_proj_out",
]
use_qfloat8 = "qfloat8" in GPU_memory_mode
if use_qfloat8:
convert_model_weight_to_float8(transformer, exclude_module_name=fp8_exclude_module_name, device=device)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
fp32_modules = [m for m in transformer.modules()
if any(p.dtype == torch.float32 for p in m.parameters(recurse=False))]
shard_fn = partial(shard_model, device_id=device, param_dtype=None, cast_dtype=False,
module_to_wrapper=list(transformer.transformer_blocks),
ignored_modules=fp32_modules)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(text_encoder.model.language_model.layers))
pipeline.text_encoder.model = shard_fn(pipeline.text_encoder.model)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
# The FP8 conversion above has already run (before the FSDP sharding, on purpose); only the dequant
# wrapper and the memory placement are left, which is what the preconverted tag installs.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype,
quant_tag="qfloat8_preconverted" if GPU_memory_mode.endswith("_and_qfloat8") else None,
exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def parse_reference(entry: str):
kind, _, media = entry.partition("=")
kind, media = kind.strip().lower(), media.strip()
if not media:
raise ValueError(f"A reference entry must be `image=path`, `video=path` or `audio=path`, got {entry!r}.")
if kind == "image":
return MiniMaxH3ImageReference.from_file(media)
if kind == "video":
return MiniMaxH3VideoReference.from_file(media)
if kind == "audio":
return MiniMaxH3AudioReference.from_file(media)
raise ValueError(f"A reference entry must start with `image=`, `video=` or `audio=`, got {entry!r}.")
# Decode every reference at the rate its container carries, which the pipeline's setup resamples onto MiniMax-H3's
# own 24 fps and the audio VAE's sample rate.
parsed_references = [parse_reference(entry) for entry in references]
pdd_callback = None if pdd_config is None else pdd_step_callback(
transformer, scheduler, audio_scheduler, pdd_config, num_inference_steps
)
with torch.no_grad():
output = pipeline(
prompt=prompt,
references=parsed_references,
height=None if sample_size is None else sample_size[0],
width=None if sample_size is None else sample_size[1],
num_frames=video_length,
num_inference_steps=num_inference_steps,
flow_shift=flow_shift,
audio_flow_shift=audio_flow_shift,
guidance_scale=guidance_scale,
generator=generator,
output_type="pt",
callback_on_step_end=pdd_callback,
)
print(f"[{os.environ.get('RANK', '0')}] generation done, decoding", flush=True)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
audio_sample_rate = output.sampling_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=audio_sample_rate)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
# Keep every rank alive until the saving rank finishes; an early exit of one rank makes the elastic launcher
# terminate the others.
dist.barrier()
else:
save_results()
+282
View File
@@ -0,0 +1,282 @@
import os
import sys
import torch
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMiniMaxH3,
AutoencoderKLMiniMaxH3Audio,
MiniMaxH3Transformer3DModel, Qwen2TokenizerFast,
Qwen3VLForConditionalGeneration,
Qwen3VLProcessor)
from videox_fun.pipeline import MiniMaxH3Pipeline
from videox_fun.utils import (MiniMaxH3Scheduler, apply_gpu_memory_mode,
convert_model_weight_to_float8, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus. The Qwen3-VL conditioner is ~62 GB, so with fsdp_dit alone every
# rank still replicates it; fsdp_text_encoder shards it too. Note it must wrap the inner `text_encoder.model`
# (Qwen3VLModel): encode_prompt calls that submodule directly, so a wrap on the top-level module would never fire.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MiniMax-H3"
# Load pretrained model if need
# A full finetune goes in `transformer_path`, either as the `transformer` folder a training checkpoint writes
# (`output_dir_minimax_h3/checkpoint-N/transformer`, config.json included) or as a single safetensors file. A LoRA
# goes in `lora_path`: handed to `transformer_path` it would match no key at all and load nothing. A PDD LoRA
# (parallel decoder) goes in `pdd_lora_path` and cannot be combined with `lora_path`.
transformer_path = None
vae_path = None
lora_path = None
pdd_lora_path = None
# Other params
# MiniMax-H3 generates at a fixed 24 fps, only accepts multiples of 32 as height / width, and snaps video_length up
# to the next 17 * n + 5 the video VAE can decode (the duration has to stay between 5 and 15 seconds).
# Leave sample_size as None to use MiniMax-H3's own 16:9 canvas (768x1344).
sample_size = [704, 1280]
video_length = 124
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
prompt = "A red fox trotting through a snowy pine forest, snow crunching underfoot"
seed = 43
# Number of denoising steps, i.e. of model evaluations: num_inference_steps = 40 runs 40 of them.
num_inference_steps = 40
# The released checkpoint is guidance-distilled: leave guidance_scale at 1 to run one forward pass per step
# with no CFG. A value above 1 enables classifier-free guidance with a negative_prompt, running two passes.
guidance_scale = 1
# The exponential sigma shifts of the two schedules. None keeps the ones of the checkpoint (12.0 video, 3.0 audio).
flow_shift = None
audio_flow_shift = None
lora_weight = 0.55
save_path = "samples/minimax-h3-videos-t2v"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# `model_name` may point either at a converted diffusers layout or at an *original* MiniMax-H3 partition (e.g.
# `MiniMax-H3/FL2VA`); the original shards are converted on the fly while loading, no intermediate copy on disk.
# Transformer
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if os.path.isdir(transformer_path):
# A training checkpoint's `transformer` folder carries its own config.json, so the loader restores the
# mixed-precision contract of the checkpoint (`_keep_in_fp32_modules`) by itself.
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
transformer_path,
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
else:
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# `strict=False` accepts a file whose keys belong to another model — a LoRA checkpoint, say — by loading
# nothing at all and silently generating with the base weights, so an unexpected key is a hard error.
assert len(u) == 0, (
f"{transformer_path} holds {len(u)} key(s) the transformer does not have, e.g. {u[:3]}. A LoRA "
"checkpoint belongs in `lora_path`, not `transformer_path`."
)
pdd_config = None
if pdd_lora_path is not None:
if lora_path is not None:
raise ValueError("`lora_path` and `pdd_lora_path` cannot be used together.")
from videox_fun.models.minimax_h3_pdd import (load_pdd_lora,
pdd_num_inference_steps,
pdd_step_callback)
pdd_config = load_pdd_lora(transformer, pdd_lora_path)
num_inference_steps = pdd_num_inference_steps(pdd_config, num_inference_steps, teacher_default=40)
# Video VAE. The released weights are float32 and the decode runs under float16 autocast, so the VAE is not
# downcast even when the rest of the pipeline is bfloat16 (this is also how the training scripts load it).
vae = AutoencoderKLMiniMaxH3.from_pretrained(
model_name,
subfolder="vae",
low_cpu_mem_usage=True,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE, waveform in / waveform out: MiniMax-H3 has no separate vocoder. Float32 as released, like the video VAE.
audio_vae = AutoencoderKLMiniMaxH3Audio.from_pretrained(
model_name,
subfolder="audio_vae",
low_cpu_mem_usage=True,
)
# Get Tokenizer and Processor
tokenizer = Qwen2TokenizerFast.from_pretrained(os.path.join(model_name, "tokenizer"))
processor = Qwen3VLProcessor.from_pretrained(os.path.join(model_name, "processor"))
# Get Text encoder. MiniMax-H3 reads the unnormalized hidden state after the 50th decoder layer of Qwen3-VL.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Schedulers. MiniMax-H3 steps the video and the audio latents down two schedules inside one transformer call.
scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="scheduler")
audio_scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="audio_scheduler")
pipeline = MiniMaxH3Pipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
transformer=transformer,
scheduler=scheduler,
audio_scheduler=audio_scheduler,
)
# The float32 modules of the mixed-precision checkpoint stay untouched by the float8 quantization.
fp8_exclude_module_name = [
"proj_in", "audio_proj_in", "context_embedder", "time_embedder", "time_proj",
"token_refiner", "norm_out", "proj_out", "audio_proj_out",
]
use_qfloat8 = "qfloat8" in GPU_memory_mode
if use_qfloat8:
convert_model_weight_to_float8(transformer, exclude_module_name=fp8_exclude_module_name, device=device)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
fp32_modules = [m for m in transformer.modules()
if any(p.dtype == torch.float32 for p in m.parameters(recurse=False))]
shard_fn = partial(shard_model, device_id=device, param_dtype=None, cast_dtype=False,
module_to_wrapper=list(transformer.transformer_blocks),
ignored_modules=fp32_modules)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(text_encoder.model.language_model.layers))
pipeline.text_encoder.model = shard_fn(pipeline.text_encoder.model)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
# The FP8 conversion above has already run (before the FSDP sharding, on purpose); only the dequant
# wrapper and the memory placement are left, which is what the preconverted tag installs.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype,
quant_tag="qfloat8_preconverted" if GPU_memory_mode.endswith("_and_qfloat8") else None,
exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
pdd_callback = None if pdd_config is None else pdd_step_callback(
transformer, scheduler, audio_scheduler, pdd_config, num_inference_steps
)
with torch.no_grad():
output = pipeline(
prompt=prompt,
height=None if sample_size is None else sample_size[0],
width=None if sample_size is None else sample_size[1],
num_frames=video_length,
num_inference_steps=num_inference_steps,
flow_shift=flow_shift,
audio_flow_shift=audio_flow_shift,
guidance_scale=guidance_scale,
generator=generator,
output_type="pt",
callback_on_step_end=pdd_callback,
)
print(f"[{os.environ.get('RANK', '0')}] generation done, decoding", flush=True)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
audio_sample_rate = output.sampling_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=audio_sample_rate)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
# Keep every rank alive until the saving rank finishes; an early exit of one rank makes the elastic launcher
# terminate the others.
dist.barrier()
else:
save_results()
@@ -0,0 +1,328 @@
import os
import sys
import torch
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMiniMaxH3,
AutoencoderKLMiniMaxH3Audio,
MiniMaxH3ControlTransformer3DModel,
Qwen2TokenizerFast,
Qwen3VLForConditionalGeneration,
Qwen3VLProcessor)
from videox_fun.pipeline import MiniMaxH3ControlPipeline
from videox_fun.utils import (MiniMaxH3Scheduler, apply_gpu_memory_mode,
convert_model_weight_to_float8,
get_video_to_video_latent, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
# Multi-GPU runs through the xfuser sequence-parallel path and must be launched with torchrun, e.g.
# `torchrun --nproc_per_node=2 examples/minimax_h3_fun/predict_v2v_control.py` for ulysses_degree=2, ring_degree=1.
# It is incompatible with the *cpu_offload* memory modes (accelerate offload hooks own a single device);
# use model_full_load / model_full_load_and_qfloat8 there, with fsdp_dit to save memory.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus. The Qwen3-VL conditioner is ~62 GB, so with fsdp_dit alone every
# rank still replicates it; fsdp_text_encoder shards it too. Note it must wrap the inner `text_encoder.model`
# (Qwen3VLModel): encode_prompt calls that submodule directly, so a wrap on the top-level module would never fire.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MiniMax-H3"
# Control branch layout, must match the yaml `train_control.py` ran with: `control_blocks_places` selects the
# layers the control blocks attach to and `control_in_dim` the channels the control rows carry (49 for an
# `--enable_inpaint` checkpoint, whose `control_proj_in` is widened with the mask channels). Leaving it None
# builds the default 24-channel branch, which cannot load an inpaint checkpoint.
config_path = "config/minimax_h3/minimax_h3_control_inpaint_post_norm.yaml"
# Load pretrained model if need. The control branch is not part of the released MiniMax-H3 weights, so a base
# `model_name` starts the side branch as an identity (`after_proj` is zero) and the c ontrol video has no effect;
# point `transformer_path` at a control checkpoint trained by `scripts/minimax_h3_fun/train_control.py`.
transformer_path = "models/Diffusion_Transformer/MiniMax-H3-Fun-Controlnet-Union-2.0/MiniMax-H3-Fun-Controlnet-Union-2.0.safetensors"
vae_path = None
lora_path = None
# Other params
# MiniMax-H3 generates at a fixed 24 fps, only accepts multiples of 32 as height / width, and the generation
# follows the control video's actual length — snapped down to the largest 17 * n + 5 the video VAE can decode so
# a short control video is never padded (the duration has to stay under 15 seconds), capped by video_length.
# Control inference fits the control video onto this canvas with the training's resize + crop geometry, so
# sample_size must be set (it cannot be None).
sample_size = [1280, 704]
video_length = 243
fps = 24
# Scale applied to every control skip before it is added to the main branch. 0.0 switches the control branch off,
# values below 1.0 weaken the guidance of the control video.
control_context_scale = 1.00
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
control_video = "asset/pose.mp4"
# Inpaint inputs, only read by checkpoints trained with `--enable_inpaint` (control_in_dim widened, e.g. 49):
# `inpaint_video` is the source video behind the mask and `inpaint_video_mask` marks the regions to regenerate
# (white = repaint, black = keep). With an inpaint checkpoint but no inpaint inputs given, the pipeline zero-pads
# the mask channels and the run degrades to pure generation; a mask-less checkpoint rejects them outright.
inpaint_video = None
inpaint_video_mask = None
prompt = "视频中,一位年轻女性站在阳光洒满的沙滩上,背景是无垠碧蓝的大海与澄澈如洗的天空,构成一幅充满夏日度假氛围的画面。她身穿一件深海军蓝吊带泳衣,线条简约贴身,凸显健康匀称的身材曲线;外搭一条纯白色背带短裙,裙摆轻盈飘逸,随风微微扬起,增添了几分俏皮与少女感。她的长发柔顺披肩,发梢微卷,在阳光下泛着自然光泽,耳畔垂挂着一对小巧精致的珍珠吊坠耳环,为整体造型注入一丝温柔优雅的气息。她面带甜美笑容,嘴角上扬,露出整齐洁白的牙齿,眼神清澈明亮,直视镜头时流露出真诚与自信,仿佛在与观众分享此刻的快乐。"
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
seed = 43
# Number of denoising steps, i.e. of model evaluations: num_inference_steps = 40 runs 40 of them.
num_inference_steps = 40
# The released checkpoint is guidance-distilled: leave guidance_scale at 1 to run one forward pass per step
# with no CFG — the distill checkpoints of train_control_distill.py already bake the teacher's CFG target into
# the weights, so any value above 1 applies guidance twice and degrades the output. A value above 1 enables
# classifier-free guidance with a negative_prompt, running two passes.
guidance_scale = 1.0
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
# The exponential sigma shifts of the two schedules. None keeps the ones of the checkpoint (12.0 video, 3.0 audio).
flow_shift = None
audio_flow_shift = None
lora_weight = 0.55
save_path = "samples/minimax-h3-videos-v2v-control"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# The yaml pins the control branch layout exactly as in training (scripts/minimax_h3_fun/train_control.py), where
# `transformer_additional_kwargs` is spread into `from_pretrained` the same way.
transformer_load_kwargs = {}
if config_path is not None:
from omegaconf import OmegaConf
config = OmegaConf.load(config_path)
transformer_load_kwargs.update(
OmegaConf.to_container(config["transformer_additional_kwargs"], resolve=True)
)
# `model_name` may point either at a converted diffusers layout or at an *original* MiniMax-H3 partition (e.g.
# `MiniMax-H3/FL2VA`); the original shards are converted on the fly while loading, no intermediate copy on disk.
# Transformer. `from_pretrained` fills the control branch the released checkpoint does not carry: every control
# block is initialised from the main block it is attached to and `control_proj_in` from `proj_in`, with
# before_proj / after_proj zeroed, so a freshly loaded model is numerically identical to the base MiniMax-H3 model.
transformer = MiniMaxH3ControlTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
**transformer_load_kwargs,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE. The released weights are float32 and the decode runs under float16 autocast, so the VAE is not
# downcast even when the rest of the pipeline is bfloat16.
vae = AutoencoderKLMiniMaxH3.from_pretrained(
model_name,
subfolder="vae",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE, waveform in / waveform out: MiniMax-H3 has no separate vocoder.
audio_vae = AutoencoderKLMiniMaxH3Audio.from_pretrained(
model_name,
subfolder="audio_vae",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Get Tokenizer and Processor
tokenizer = Qwen2TokenizerFast.from_pretrained(os.path.join(model_name, "tokenizer"))
processor = Qwen3VLProcessor.from_pretrained(os.path.join(model_name, "processor"))
# Get Text encoder. MiniMax-H3 reads the unnormalized hidden state after the 50th decoder layer of Qwen3-VL.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Schedulers. MiniMax-H3 steps the video and the audio latents down two schedules inside one transformer call.
scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="scheduler")
audio_scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="audio_scheduler")
pipeline = MiniMaxH3ControlPipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
transformer=transformer,
scheduler=scheduler,
audio_scheduler=audio_scheduler,
)
# The float32 modules of the mixed-precision checkpoint stay untouched by the float8 quantization. The `proj_in`
# entry also covers the control patch projection `control_proj_in`, which shares the video patch projection's dtype.
fp8_exclude_module_name = [
"proj_in", "audio_proj_in", "context_embedder", "time_embedder", "time_proj",
"token_refiner", "norm_out", "proj_out", "audio_proj_out",
]
use_qfloat8 = "qfloat8" in GPU_memory_mode
if use_qfloat8:
convert_model_weight_to_float8(transformer, exclude_module_name=fp8_exclude_module_name, device=device)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
fp32_modules = [m for m in transformer.modules()
if any(p.dtype == torch.float32 for p in m.parameters(recurse=False))]
shard_fn = partial(shard_model, device_id=device, param_dtype=None, cast_dtype=False,
module_to_wrapper=list(transformer.transformer_blocks) + list(transformer.control_blocks),
ignored_modules=fp32_modules)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(text_encoder.model.language_model.layers))
pipeline.text_encoder.model = shard_fn(pipeline.text_encoder.model)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
# The FP8 conversion above has already run (before the FSDP sharding, on purpose); only the dequant
# wrapper and the memory placement are left, which is what the preconverted tag installs.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype,
quant_tag="qfloat8_preconverted" if GPU_memory_mode.endswith("_and_qfloat8") else None,
exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def snap_num_frames(actual_num_frames, max_num_frames):
"""
Pick the generation length from the control video instead of padding a short one: the largest `17 * n + 5`
the video VAE can decode that does not exceed the frames actually read (capped by `max_num_frames`), snapping
down so no tail frame is ever repeated. A control video below 5 frames is raised to 5, the smallest count
the video VAE can encode.
"""
num_frames = min(actual_num_frames, max_num_frames)
num_frames = (num_frames - 5) // 17 * 17 + 5
return max(num_frames, 5)
with torch.no_grad():
control_video, _, _, _ = get_video_to_video_latent(control_video, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
# Generate at the control video's actual length, never padding; only control videos below the 5 frames the
# video VAE can encode are raised to 5.
num_frames = snap_num_frames(control_video.shape[2], video_length)
if num_frames != video_length:
print(f"[{os.environ.get('RANK', '0')}] control video holds {control_video.shape[2]} frames, generating "
f"{num_frames} instead of {video_length}", flush=True)
mask_video = None
if inpaint_video is not None:
if inpaint_video_mask is None:
raise ValueError("inpaint_video_mask is required when inpaint_video is provided")
inpaint_video, _, _, _ = get_video_to_video_latent(inpaint_video, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
inpaint_video_mask, _, _, _ = get_video_to_video_latent(inpaint_video_mask, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
# Binarize the grayscale mask onto one channel: 1 marks the regions to regenerate, mirroring the training
# `get_random_mask` convention the visibility map `1 - mask` is built from.
mask_video = (inpaint_video_mask[:, :1] > 0.5).to(inpaint_video_mask.dtype)
output = pipeline(
prompt=prompt,
control_video=control_video,
control_context_scale=control_context_scale,
mask_video=mask_video,
inpaint_video=inpaint_video,
height=None if sample_size is None else sample_size[0],
width=None if sample_size is None else sample_size[1],
num_frames=num_frames,
num_inference_steps=num_inference_steps,
flow_shift=flow_shift,
audio_flow_shift=audio_flow_shift,
guidance_scale=guidance_scale,
negative_prompt=negative_prompt,
generator=generator,
output_type="pt",
)
print(f"[{os.environ.get('RANK', '0')}] generation done, decoding", flush=True)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
audio_sample_rate = output.sampling_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=audio_sample_rate)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
# Keep every rank alive until the saving rank finishes; an early exit of one rank makes the elastic launcher
# terminate the others.
dist.barrier()
else:
save_results()
@@ -0,0 +1,351 @@
import os
import sys
import torch
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMiniMaxH3,
AutoencoderKLMiniMaxH3Audio,
MiniMaxH3ControlTransformer3DModel,
Qwen2TokenizerFast,
Qwen3VLForConditionalGeneration,
Qwen3VLProcessor)
from videox_fun.pipeline import MiniMaxH3ControlPipeline
from videox_fun.utils import (MiniMaxH3Scheduler, apply_gpu_memory_mode,
convert_model_weight_to_float8,
get_video_to_video_latent, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "model_group_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
# Multi-GPU runs through the xfuser sequence-parallel path and must be launched with torchrun, e.g.
# `torchrun --nproc_per_node=2 examples/minimax_h3_fun/predict_v2v_control.py` for ulysses_degree=2, ring_degree=1.
# It is incompatible with the *cpu_offload* memory modes (accelerate offload hooks own a single device);
# use model_full_load / model_full_load_and_qfloat8 there, with fsdp_dit to save memory.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus. The Qwen3-VL conditioner is ~62 GB, so with fsdp_dit alone every
# rank still replicates it; fsdp_text_encoder shards it too. Note it must wrap the inner `text_encoder.model`
# (Qwen3VLModel): encode_prompt calls that submodule directly, so a wrap on the top-level module would never fire.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MiniMax-H3"
# Control branch layout, must match the yaml `train_control.py` ran with: `control_blocks_places` selects the
# layers the control blocks attach to and `control_in_dim` the channels the control rows carry (49 for an
# `--enable_inpaint` checkpoint, whose `control_proj_in` is widened with the mask channels). Leaving it None
# builds the default 24-channel branch, which cannot load an inpaint checkpoint.
config_path = "config/minimax_h3/minimax_h3_control_inpaint_post_norm.yaml"
# Load pretrained model if need. The control branch is not part of the released MiniMax-H3 weights, so a base
# `model_name` starts the side branch as an identity (`after_proj` is zero) and the c ontrol video has no effect;
# point `transformer_path` at a control checkpoint trained by `scripts/minimax_h3_fun/train_control.py`.
transformer_path = "models/Diffusion_Transformer/MiniMax-H3-Fun-Controlnet-Union-2.0/MiniMax-H3-Fun-Controlnet-Union-2.0.safetensors"
vae_path = None
lora_path = None
# Other params
# MiniMax-H3 generates at a fixed 24 fps, only accepts multiples of 32 as height / width, and the generation
# follows the control video's actual length — snapped down to the largest 17 * n + 5 the video VAE can decode so
# a short control video is never padded (the duration has to stay under 15 seconds), capped by video_length.
# Control inference fits the control video onto this canvas with the training's resize + crop geometry, so
# sample_size must be set (it cannot be None).
sample_size = [704, 1280]
video_length = 124
fps = 24
# Scale applied to every control skip before it is added to the main branch. 0.0 switches the control branch off,
# values below 1.0 weaken the guidance of the control video.
control_context_scale = 1.00
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Path of the control (e.g. pose) video; leaving it None zeroes the control channels of the side branch. With
# inpaint inputs given the mask then guides the run on its own (the layout training reaches when it drops the
# control rows); without them the run degrades to plain base-pipeline generation at `video_length` frames.
control_video = None
# Inpaint inputs, only read by checkpoints trained with `--enable_inpaint` (control_in_dim widened, e.g. 49):
# `inpaint_video` is the source video behind the mask and `inpaint_video_mask` marks the regions to regenerate
# (white = repaint, black = keep). With an inpaint checkpoint but no inpaint inputs given, the pipeline zero-pads
# the mask channels and the run degrades to pure generation; a mask-less checkpoint rejects them outright.
inpaint_video = "asset/inpaint_video.mp4"
inpaint_video_mask = "asset/inpaint_video_mask.mp4"
prompt = "一只狗在沙发上摇头"
seed = 43
# Number of denoising steps, i.e. of model evaluations: num_inference_steps = 40 runs 40 of them.
num_inference_steps = 40
# The released checkpoint is guidance-distilled: leave guidance_scale at 1 to run one forward pass per step
# with no CFG — the distill checkpoints of train_control_distill.py already bake the teacher's CFG target into
# the weights, so any value above 1 applies guidance twice and degrades the output. A value above 1 enables
# classifier-free guidance with a negative_prompt, running two passes.
guidance_scale = 1.0
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走"
# The exponential sigma shifts of the two schedules. None keeps the ones of the checkpoint (12.0 video, 3.0 audio).
flow_shift = None
audio_flow_shift = None
lora_weight = 0.55
save_path = "samples/minimax-h3-videos-v2v-control"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# The yaml pins the control branch layout exactly as in training (scripts/minimax_h3_fun/train_control.py), where
# `transformer_additional_kwargs` is spread into `from_pretrained` the same way.
transformer_load_kwargs = {}
if config_path is not None:
from omegaconf import OmegaConf
config = OmegaConf.load(config_path)
transformer_load_kwargs.update(
OmegaConf.to_container(config["transformer_additional_kwargs"], resolve=True)
)
# `model_name` may point either at a converted diffusers layout or at an *original* MiniMax-H3 partition (e.g.
# `MiniMax-H3/FL2VA`); the original shards are converted on the fly while loading, no intermediate copy on disk.
# Transformer. `from_pretrained` fills the control branch the released checkpoint does not carry: every control
# block is initialised from the main block it is attached to and `control_proj_in` from `proj_in`, with
# before_proj / after_proj zeroed, so a freshly loaded model is numerically identical to the base MiniMax-H3 model.
transformer = MiniMaxH3ControlTransformer3DModel.from_pretrained(
model_name,
subfolder="transformer",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
**transformer_load_kwargs,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE. The released weights are float32 and the decode runs under float16 autocast, so the VAE is not
# downcast even when the rest of the pipeline is bfloat16.
vae = AutoencoderKLMiniMaxH3.from_pretrained(
model_name,
subfolder="vae",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio VAE, waveform in / waveform out: MiniMax-H3 has no separate vocoder.
audio_vae = AutoencoderKLMiniMaxH3Audio.from_pretrained(
model_name,
subfolder="audio_vae",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
# Get Tokenizer and Processor
tokenizer = Qwen2TokenizerFast.from_pretrained(os.path.join(model_name, "tokenizer"))
processor = Qwen3VLProcessor.from_pretrained(os.path.join(model_name, "processor"))
# Get Text encoder. MiniMax-H3 reads the unnormalized hidden state after the 50th decoder layer of Qwen3-VL.
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
os.path.join(model_name, "text_encoder"),
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Schedulers. MiniMax-H3 steps the video and the audio latents down two schedules inside one transformer call.
scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="scheduler")
audio_scheduler = MiniMaxH3Scheduler.from_pretrained(model_name, subfolder="audio_scheduler")
pipeline = MiniMaxH3ControlPipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
processor=processor,
transformer=transformer,
scheduler=scheduler,
audio_scheduler=audio_scheduler,
)
# The float32 modules of the mixed-precision checkpoint stay untouched by the float8 quantization. The `proj_in`
# entry also covers the control patch projection `control_proj_in`, which shares the video patch projection's dtype.
fp8_exclude_module_name = [
"proj_in", "audio_proj_in", "context_embedder", "time_embedder", "time_proj",
"token_refiner", "norm_out", "proj_out", "audio_proj_out",
]
use_qfloat8 = "qfloat8" in GPU_memory_mode
if use_qfloat8:
# Scale-aware fp8 must run before the FSDP wrapping below so the flat buffers hold the fp8 values.
convert_model_weight_to_float8(transformer, exclude_module_name=fp8_exclude_module_name, device=device)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
transformer.enable_multi_gpus_inference()
if fsdp_dit:
# The mixed-precision checkpoint pins the patch embedders / timestep MLP / output heads to float32;
# FSDP keeps them replicated via ignored_states so the flat buffers stay uniform-dtype.
#
# Root cause of the temporal flicker, verified by per-step / per-block instrumentation: with
# `MixedPrecision(param_dtype=...)` the root FSDP unit applies `cast_root_forward_inputs` (default
# True), so the whole root forward runs in `param_dtype`. That casts the root forward inputs — the
# sinusoidal timestep embedding, the packed latents, the context — to bfloat16 and forces the fp32-
# pinned heads (proj_in / time_embedder / audio_proj_in) to compute on coarsely rounded inputs in
# bfloat16 instead of their native fp32; the deviation compounds over the sampling steps and flips
# trajectories that sit on the numerical-stability edge into coherent flicker at fixed latent-time
# positions, seed-independently.
# Sharding with `param_dtype=None` + `cast_dtype=False` casts nothing (no MixedPrecision compute
# dtype, no root input cast), keeps the native fp32 hidden path and matches the non-FSDP numerics.
fp32_modules = [m for m in transformer.modules()
if any(p.dtype == torch.float32 for p in m.parameters(recurse=False))]
shard_fn = partial(shard_model, device_id=device, param_dtype=None, cast_dtype=False,
module_to_wrapper=list(transformer.transformer_blocks) + list(transformer.control_blocks),
ignored_modules=fp32_modules)
pipeline.transformer = shard_fn(pipeline.transformer)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype,
module_to_wrapper=list(text_encoder.model.language_model.layers))
pipeline.text_encoder.model = shard_fn(pipeline.text_encoder.model)
print("Add FSDP TEXT ENCODER")
if compile_dit:
for i in range(len(pipeline.transformer.transformer_blocks)):
pipeline.transformer.transformer_blocks[i] = torch.compile(pipeline.transformer.transformer_blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
# The FP8 conversion above has already run (before the FSDP sharding, on purpose); only the dequant
# wrapper and the memory placement are left, which is what the preconverted tag installs.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype,
quant_tag="qfloat8_preconverted" if GPU_memory_mode.endswith("_and_qfloat8") else None,
exclude_module_name=[])
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
def snap_num_frames(actual_num_frames, max_num_frames):
"""
Pick the generation length from the control video instead of padding a short one: the largest `17 * n + 5`
the video VAE can decode that does not exceed the frames actually read (capped by `max_num_frames`), snapping
down so no tail frame is ever repeated. A control video below 5 frames is raised to 5, the smallest count
the video VAE can encode.
"""
num_frames = min(actual_num_frames, max_num_frames)
num_frames = (num_frames - 5) // 17 * 17 + 5
return max(num_frames, 5)
with torch.no_grad():
control_video, _, _, _ = get_video_to_video_latent(control_video, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
# Generate at the control video's actual length, never padding; only control videos below the 5 frames the
# video VAE can encode are raised to 5. Without a control video the request keeps `video_length`, which the
# pipeline snaps up to the next 17 * n + 5 itself (control branch off = plain base-pipeline generation).
if control_video is None:
num_frames = video_length
print(f"[{os.environ.get('RANK', '0')}] no control video given, "
+ (f"running inpaint alone at {num_frames} frames" if inpaint_video is not None
else f"generating {num_frames} frames without the control branch"), flush=True)
else:
num_frames = snap_num_frames(control_video.shape[2], video_length)
if num_frames != video_length:
print(f"[{os.environ.get('RANK', '0')}] control video holds {control_video.shape[2]} frames, generating "
f"{num_frames} instead of {video_length}", flush=True)
mask_video = None
if inpaint_video is not None:
if inpaint_video_mask is None:
raise ValueError("inpaint_video_mask is required when inpaint_video is provided")
inpaint_video, _, _, _ = get_video_to_video_latent(inpaint_video, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
inpaint_video_mask, _, _, _ = get_video_to_video_latent(inpaint_video_mask, video_length=video_length, sample_size=sample_size, fps=fps, ref_image=None, keep_aspect_ratio=True)
# Binarize the grayscale mask onto one channel: 1 marks the regions to regenerate, mirroring the training
# `get_random_mask` convention the visibility map `1 - mask` is built from.
mask_video = (inpaint_video_mask[:, :1] > 0.5).to(inpaint_video_mask.dtype)
output = pipeline(
prompt=prompt,
control_video=control_video,
control_context_scale=control_context_scale,
mask_video=mask_video,
inpaint_video=inpaint_video,
height=None if sample_size is None else sample_size[0],
width=None if sample_size is None else sample_size[1],
num_frames=num_frames,
num_inference_steps=num_inference_steps,
flow_shift=flow_shift,
audio_flow_shift=audio_flow_shift,
guidance_scale=guidance_scale,
negative_prompt=negative_prompt,
generator=generator,
output_type="pt",
)
print(f"[{os.environ.get('RANK', '0')}] generation done, decoding", flush=True)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
sample = output.videos
audio = output.audio
audio_sample_rate = output.sampling_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=audio_sample_rate)
if ulysses_degree * ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
# Keep every rank alive until the saving rank finishes; an early exit of one rank makes the elastic launcher
# terminate the others.
dist.barrier()
else:
save_results()
+356
View File
@@ -0,0 +1,356 @@
import os
import sys
import numpy as np
import torch
from diffusers import FlowMatchEulerDiscreteScheduler
from PIL import Image
current_file_path = os.path.abspath(__file__)
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
for project_root in project_roots:
sys.path.insert(0, project_root) if project_root not in sys.path else None
from videox_fun.dist import set_multi_gpus_devices, shard_model
from videox_fun.models import (AutoencoderKLMOVAAudio, AutoencoderKLWan,
AutoTokenizer, MOVADualTowerConditionalBridge,
UMT5EncoderModel, WanAudioTransformer3DModel,
WanTransformer3DModel)
from videox_fun.pipeline import MOVAPipeline
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, merge_lora,
save_videos_with_audio_grid, unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_cpu_offload means that the entire model will be moved to the CPU after use, which can save some GPU memory.
#
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
# Multi GPUs config
# Please ensure that the product of ulysses_degree and ring_degree equals the number of GPUs used.
# For example, if you are using 8 GPUs, you can set ulysses_degree = 2 and ring_degree = 4.
# If you are using 1 GPU, you can set ulysses_degree = 1 and ring_degree = 1.
ulysses_degree = 1
ring_degree = 1
# Use FSDP to save more GPU memory in multi gpus.
fsdp_dit = False
fsdp_text_encoder = True
# Compile will give a speedup in fixed resolution and need a little GPU memory.
# The compile_dit is not compatible with sequential_cpu_offload.
compile_dit = False
# model path
model_name = "models/Diffusion_Transformer/MOVA-360p"
# Choose the sampler in "Flow", "Flow_Unipc", "Flow_DPM++"
sampler_name = "Flow"
boundary_ratio = 0.9
# Load pretrained model if need
# The transformer_path is used for low noise model, the transformer_high_path is used for high noise model.
transformer_path = None
transformer_high_path = None
transformer_audio_path = None
bridge_path = None
vae_path = None
audio_vae_path = None
# Load lora model if need
# The lora_path is used for low noise model, the lora_high_path is used for high noise model.
lora_path = None
lora_high_path = None
# Other params
sample_size = [640, 352]
video_length = 81
fps = 24
# Use torch.float16 if GPU does not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
# Input image for I2V
validation_image = "asset/8.png"
# prompts
prompt = "Medium shot of a girl by the ocean. She starts with a bright smile, then gently nods her head while speaking. Her mouth moves naturally to say: \"Hi, nice to meet you.\" She maintains eye contact throughout. The background shows calm waves. Smooth motion, cinematic quality, realistic facial expressions."
negative_prompt = "色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指"
guidance_scale = 5.0
seed = 43
num_inference_steps = 50
# The lora_weight is used for low noise model, the lora_high_weight is used for high noise model.
lora_weight = 0.55
lora_high_weight = 0.55
save_path = "samples/mova-videos-i2v"
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
# The from_pretrained method automatically converts WanModel config to WanTransformer3DModel config
print("Loading Video DiT (High Noise) with WanTransformer3DModel...")
transformer = WanTransformer3DModel.from_pretrained(
model_name,
subfolder="video_dit_2",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video DiT 2 (Low Noise) - Using WanTransformer3DModel
print("Loading Video DiT 2 (Low Noise) with WanTransformer3DModel...")
transformer_2 = WanTransformer3DModel.from_pretrained(
model_name,
subfolder="video_dit",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_high_path is not None:
print(f"From checkpoint: {transformer_high_path}")
if transformer_high_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_high_path)
else:
state_dict = torch.load(transformer_high_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer_2.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Audio DiT - Using WanAudioTransformer3DModel
print("Loading Audio DiT with WanAudioTransformer3DModel...")
transformer_audio = WanAudioTransformer3DModel.from_pretrained(
model_name,
subfolder="audio_dit",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if transformer_audio_path is not None:
print(f"From checkpoint: {transformer_audio_path}")
if transformer_audio_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(transformer_audio_path)
else:
state_dict = torch.load(transformer_audio_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = transformer_audio.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Dual Tower Bridge
print("Loading Dual Tower Bridge...")
dual_tower_bridge = MOVADualTowerConditionalBridge.from_pretrained(
model_name,
subfolder="dual_tower_bridge",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
if bridge_path is not None:
print(f"From checkpoint: {bridge_path}")
if bridge_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(bridge_path)
else:
state_dict = torch.load(bridge_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = dual_tower_bridge.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Video VAE
print("Loading Video VAE...")
vae = AutoencoderKLWan.from_pretrained(
os.path.join(model_name, "video_vae/diffusion_pytorch_model.safetensors")
).to(weight_dtype)
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
audio_vae = AutoencoderKLMOVAAudio.from_pretrained(
model_name,
subfolder="audio_vae",
torch_dtype=torch.float32,
)
if audio_vae_path is not None:
print(f"From checkpoint: {audio_vae_path}")
if audio_vae_path.endswith("safetensors"):
from safetensors.torch import load_file
state_dict = load_file(audio_vae_path)
else:
state_dict = torch.load(audio_vae_path, map_location="cpu")
state_dict = state_dict["state_dict"] if "state_dict" in state_dict else state_dict
m, u = audio_vae.load_state_dict(state_dict, strict=False)
print(f"missing keys: {len(m)}, unexpected keys: {len(u)}")
# Get Tokenizer
print("Loading Tokenizer...")
tokenizer = AutoTokenizer.from_pretrained(
model_name,
subfolder="tokenizer",
)
# Get Text Encoder
print("Loading Text Encoder...")
text_encoder = UMT5EncoderModel.from_pretrained(
model_name,
subfolder="text_encoder",
low_cpu_mem_usage=True,
torch_dtype=weight_dtype,
)
text_encoder = text_encoder.eval()
# Get Scheduler
print("Loading Scheduler...")
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
}[sampler_name]
scheduler = Chosen_Scheduler.from_pretrained(
model_name,
subfolder="scheduler"
)
# Build Pipeline
print("Building MOVAPipeline Pipeline...")
pipeline = MOVAPipeline(
vae=vae,
audio_vae=audio_vae,
text_encoder=text_encoder,
tokenizer=tokenizer,
scheduler=scheduler,
transformer=transformer,
transformer_2=transformer_2,
transformer_audio=transformer_audio,
dual_tower_bridge=dual_tower_bridge,
audio_vae_type="dac",
)
if ulysses_degree > 1 or ring_degree > 1:
from functools import partial
# Enable multi-GPU inference for visual transformers
transformer.enable_multi_gpus_inference()
transformer_2.enable_multi_gpus_inference()
if fsdp_dit:
# Apply FSDP to visual transformer blocks
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype)
pipeline.transformer = shard_fn(pipeline.transformer)
pipeline.transformer_2 = shard_fn(pipeline.transformer_2)
print("Add FSDP DIT")
if fsdp_text_encoder:
shard_fn = partial(shard_model, device_id=device, param_dtype=weight_dtype, module_to_wrapper=text_encoder.encoder.block)
pipeline.text_encoder = shard_fn(pipeline.text_encoder)
print("Add FSDP TEXT ENCODER")
if compile_dit:
# Compile MOVAModel blocks
# NOTE: compile_dit is not compatible with fsdp_dit
if fsdp_dit:
print("WARNING: compile_dit is not compatible with fsdp_dit. Disabling compile.")
else:
for i in range(len(pipeline.transformer.blocks)):
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
for i in range(len(pipeline.transformer_2.blocks)):
pipeline.transformer_2.blocks[i] = torch.compile(pipeline.transformer_2.blocks[i])
for i in range(len(pipeline.transformer_audio.blocks)):
pipeline.transformer_audio.blocks[i] = torch.compile(pipeline.transformer_audio.blocks[i])
print("Add Compile")
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed, and both transformers of this MoE setup are handled in one call, which
# is exactly the bookkeeping the old 30-line if/elif chain repeated per script.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
generator = torch.Generator(device=device).manual_seed(seed)
if lora_path is not None:
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
pipeline = merge_lora(pipeline, lora_high_path, lora_high_weight, device=device, dtype=weight_dtype, sub_transformer_name="transformer_2")
# Run inference
print("Running inference...")
with torch.no_grad():
image = Image.open(validation_image).convert("RGB")
output = pipeline(
prompt=prompt,
image=image,
negative_prompt=negative_prompt,
height=sample_size[0],
width=sample_size[1],
num_frames=video_length,
frame_rate=fps,
num_inference_steps=num_inference_steps,
guidance_scale=guidance_scale,
generator=generator,
boundary=boundary_ratio,
)
if lora_path is not None:
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
pipeline = unmerge_lora(pipeline, lora_high_path, lora_high_weight, device=device, dtype=weight_dtype, sub_transformer_name="transformer_2")
sample = output.videos
audio = output.audio
# Get audio sample rate from pipeline
audio_sample_rate = pipeline.audio_sample_rate
def save_results():
if not os.path.exists(save_path):
os.makedirs(save_path, exist_ok=True)
index = len([path for path in os.listdir(save_path)]) + 1
prefix = str(index).zfill(8)
if video_length == 1:
video_path = os.path.join(save_path, prefix + ".png")
image = sample[0, :, 0]
image = image.transpose(0, 1).transpose(1, 2)
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
sr = getattr(pipeline.audio_vae.config, "output_sampling_rate", audio_sample_rate)
save_videos_with_audio_grid(sample, audio, video_path, fps=fps, audio_sample_rate=sr)
if ulysses_degree > 1 or ring_degree > 1:
import torch.distributed as dist
if dist.get_rank() == 0:
save_results()
else:
save_results()
+18 -29
View File
@@ -18,16 +18,13 @@ from videox_fun.models import (AutoencoderKLWan, AutoTokenizer,
WanT5EncoderModel, WanTransformer3DModel)
from videox_fun.models.cache_utils import get_teacache_coefficients
from videox_fun.pipeline import WanFunPhantomPipeline
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
convert_weight_dtype_wrapper,
replace_parameters_by_name)
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
from videox_fun.utils.utils import (filter_kwargs, get_image_latent,
save_videos_grid)
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
from videox_fun.utils import (FlowDPMSolverMultistepScheduler,
FlowUniPCMultistepScheduler,
apply_gpu_memory_mode, filter_kwargs,
get_image_latent, merge_lora, save_videos_grid,
unmerge_lora)
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, sequential_cpu_offload].
# GPU memory mode, which can be chosen in [model_full_load, model_full_load_and_qfloat8, model_cpu_offload, model_cpu_offload_and_qfloat8, model_group_offload, sequential_cpu_offload].
# model_full_load means that the entire model will be moved to the GPU.
#
# model_full_load_and_qfloat8 means that the entire model will be moved to the GPU,
@@ -38,6 +35,9 @@ from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
# model_cpu_offload_and_qfloat8 indicates that the entire model will be moved to the CPU after use,
# and the transformer model has been quantized to float8, which can save more GPU memory.
#
# model_group_offload transfers internal layer groups between CPU/CUDA,
# balancing memory efficiency and speed between full-module and leaf-level offloading methods.
#
# sequential_cpu_offload means that each layer of the model will be moved to the CPU after use,
# resulting in slower speeds but saving a large amount of GPU memory.
GPU_memory_mode = "sequential_cpu_offload"
@@ -103,7 +103,7 @@ video_length = 81
fps = 16
# Use torch.float16 if GPU does not support torch.bfloat16
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
# Some graphics cards, such as v100, 2080ti, do not support torch.bfloat16
weight_dtype = torch.bfloat16
subject_ref_images = ["asset/ref_1.png", "asset/ref_2.png"]
@@ -135,7 +135,7 @@ transformer = WanTransformer3DModel.from_pretrained(
if transformer_path is not None:
print(f"From checkpoint: {transformer_path}")
if transformer_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(transformer_path)
else:
state_dict = torch.load(transformer_path, map_location="cpu")
@@ -153,7 +153,7 @@ vae = AutoencoderKLWan.from_pretrained(
if vae_path is not None:
print(f"From checkpoint: {vae_path}")
if vae_path.endswith("safetensors"):
from safetensors.torch import load_file, safe_open
from safetensors.torch import load_file
state_dict = load_file(vae_path)
else:
state_dict = torch.load(vae_path, map_location="cpu")
@@ -177,7 +177,7 @@ text_encoder = WanT5EncoderModel.from_pretrained(
text_encoder = text_encoder.eval()
# Get Scheduler
Chosen_Scheduler = scheduler_dict = {
Chosen_Scheduler = {
"Flow": FlowMatchEulerDiscreteScheduler,
"Flow_Unipc": FlowUniPCMultistepScheduler,
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
@@ -213,22 +213,10 @@ if compile_dit:
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
print("Add Compile")
if GPU_memory_mode == "sequential_cpu_offload":
replace_parameters_by_name(transformer, ["modulation",], device=device)
transformer.freqs = transformer.freqs.to(device=device)
pipeline.enable_sequential_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_cpu_offload":
pipeline.enable_model_cpu_offload(device=device)
elif GPU_memory_mode == "model_full_load_and_qfloat8":
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
convert_weight_dtype_wrapper(transformer, weight_dtype)
pipeline.to(device=device)
else:
pipeline.to(device=device)
# Quantize (when the mode carries an "_and_<quant>" suffix) and then place the pipeline.
# The order lives inside the helper: quantization has to happen before the offload hooks
# are installed.
apply_gpu_memory_mode(pipeline, GPU_memory_mode, device, weight_dtype)
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
if coefficients is not None:
@@ -288,6 +276,7 @@ def save_results():
image = (image * 255).numpy().astype(np.uint8)
image = Image.fromarray(image)
image.save(video_path)
print(f"Saved image to: {video_path}")
else:
video_path = os.path.join(save_path, prefix + ".mp4")
save_videos_grid(sample, video_path, fps=fps)

Some files were not shown because too many files have changed in this diff Show More