Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
efcc128fd4 | ||
|
|
ed0a9dd9b0 | ||
|
|
13c0091109 | ||
|
|
a49ebb962c |
@@ -1,28 +0,0 @@
|
||||
name: Publish to Comfy registry
|
||||
on:
|
||||
workflow_dispatch:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
- master
|
||||
paths:
|
||||
- "pyproject.toml"
|
||||
|
||||
permissions:
|
||||
issues: write
|
||||
|
||||
jobs:
|
||||
publish-node:
|
||||
name: Publish Custom Node to registry
|
||||
runs-on: ubuntu-latest
|
||||
if: ${{ github.repository_owner == 'aigc-apps' }}
|
||||
steps:
|
||||
- name: Check out code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
submodules: true
|
||||
- name: Publish Custom Node
|
||||
uses: Comfy-Org/publish-node-action@v1
|
||||
with:
|
||||
## Add your own personal access token to your Github Repository secrets and reference it here.
|
||||
personal_access_token: ${{ secrets.REGISTRY_ACCESS_TOKEN }}
|
||||
@@ -1,16 +1,11 @@
|
||||
# Used in VideoX-Fun
|
||||
_*
|
||||
/models*
|
||||
/output*
|
||||
/logs*
|
||||
/taming*
|
||||
/samples*
|
||||
/datasets*
|
||||
/asset*
|
||||
/repo*
|
||||
/scripts_demo*
|
||||
|
||||
# Byte-compiled / optimized / DLL files
|
||||
models*
|
||||
output*
|
||||
logs*
|
||||
taming*
|
||||
samples*
|
||||
datasets*
|
||||
asset*
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
*$py.class
|
||||
|
||||
@@ -1,128 +0,0 @@
|
||||
---
|
||||
name: integrating-models
|
||||
description: Guides adding, porting, or onboarding a diffusion model (transformer/VAE/encoder, inference pipeline, training script, config) into the VideoX-Fun repository by mirroring the closest existing model family and maximizing reuse of the repository's existing code and shared infrastructure. Use when integrating a new model/architecture, or when creating predict_*.py inference scripts, scripts/*/train*.py training scripts, pipeline_*.py, config/*.yaml, or model definitions under videox_fun/models/.
|
||||
---
|
||||
|
||||
# Integrating Models into VideoX-Fun
|
||||
|
||||
## Core rule: maximize reuse of existing repo code — mirror, extend, never reinvent
|
||||
|
||||
**Prime directive: reuse this repository's existing code to the maximum.** Nearly every building block you need already exists in `videox_fun/` or in a sibling model family. Your job is to **find it, import it, and extend it** — not to write a parallel implementation. A new file should be mostly reused structure plus the genuinely model-specific delta; the less new code you write, the better.
|
||||
|
||||
**Reuse-first protocol — before writing ANY new function / class / util:**
|
||||
1. **Search the repo first.** Grep `videox_fun/` and the closest family for an existing equivalent (weight loader, scheduler, sampler, offload, attention, LoRA, fp8, dataset, dist helper, save/metric util). If one exists → **import and reuse it**. If it is 80% right → **extend / parameterize it**, do not fork it.
|
||||
2. **Only if nothing exists** may you add new code — and then put it in the shared layer (`videox_fun/utils`, `videox_fun/data`, `videox_fun/dist`) so the next model reuses it too, instead of burying it in a family folder.
|
||||
3. **Never copy-paste** a util into a new file (that creates drift); import the single source of truth.
|
||||
|
||||
**Mirror the closest family.** Every model follows the **same layered template**. Integrating a model means finding the closest existing family and mirroring its structure, changing only what genuinely differs:
|
||||
1. Pick the closest existing family by task type (t2v / i2v / v2v-control / s2v / t2i / edit / distill): `wan2.1`, `wan2.1_fun`, `wan2.2`, `qwenimage`, `flux2`, `minimax_h3`, `ltx2`, `longcatvideo`, `cogvideox_fun`, `z_image`, etc.
|
||||
2. Read that family end-to-end across all layers:
|
||||
- `examples/<family>/predict_*.py` (inference entry)
|
||||
- `scripts/<family>/train*.py` + `*.sh` + `README_TRAIN*.md` (training)
|
||||
- `videox_fun/pipeline/pipeline_<family>*.py` (pipeline)
|
||||
- `videox_fun/models/<family>_*.py` (model definitions)
|
||||
- `config/<family>/*.yaml` (config)
|
||||
3. Copy that structure and adapt. Keep names, argument sets, control flow, and reuse points identical in shape.
|
||||
|
||||
Writing a bespoke pipeline, weight loader, trainer, sampler, dataset, or offload scheme from scratch is a **failure mode**. If you are tempted to, **stop** and check the Reuse inventory below first.
|
||||
|
||||
## Repository layout (where each layer lives)
|
||||
|
||||
| Layer | Location | What it is |
|
||||
|-------|----------|------------|
|
||||
| Model definitions | `videox_fun/models/<family>_*.py` | Transformer / VAE / text-audio-image encoders. Diffusers `ModelMixin`+`ConfigMixin`, `@register_to_config`, custom `from_pretrained`. |
|
||||
| Model registry | `videox_fun/models/__init__.py` | Imports every model class. **Must be updated** for a new model. |
|
||||
| Inference pipelines | `videox_fun/pipeline/pipeline_<family>*.py` | `<Family>Pipeline(DiffusionPipeline)` with `__call__`. |
|
||||
| Pipeline registry | `videox_fun/pipeline/__init__.py` | Imports every pipeline + aliases. **Must be updated.** |
|
||||
| Configs (optional) | `config/<family>/*.yaml` | OmegaConf YAML for civitai/custom layouts; a standard diffusers-layout checkpoint can load without one. |
|
||||
| Inference entry scripts | `examples/<family>/predict_*.py` | User-facing, config-block-at-top runnable scripts. |
|
||||
| Inference services | `examples/<family>/{app.py,launch_api.py,post_infer*.py}` | Gradio UI / API server / batch inference. |
|
||||
| Training scripts | `scripts/<family>/train*.py` | `train.py`, `train_lora.py`, `train_control.py`, `train_distill.py`, ... |
|
||||
| Training launchers | `scripts/<family>/train*.sh` | `accelerate launch` / DeepSpeed command with full arg list. |
|
||||
| Training docs | `scripts/<family>/README_TRAIN*.md` | Bilingual pairs: `README_TRAIN.md` + `README_TRAIN_zh-CN.md`. |
|
||||
| Shared: schedulers/utils | `videox_fun/utils/` | `fm_solvers`, `fm_solvers_unipc`, `lora_utils`, `fp8_optimization`, `group_offload`, `utils.py`. |
|
||||
| Shared: distributed | `videox_fun/dist/` | `fsdp.shard_model`, `fuser.set_multi_gpus_devices`, `<family>_xfuser` sequence-parallel attention. |
|
||||
| Shared: data | `videox_fun/data/` | Datasets (`ImageVideoDataset`, `VideoDataset`, ...) + bucket/aspect-ratio samplers. |
|
||||
| Demo / test datasets | `datasets/X-Fun-*-Demo/` | Ready-made smoke-test data, downloaded via `modelscope download --dataset PAI/<name>`; each ships several `metadata*.json` variants. **The only test data to use** (see reference.md §8). |
|
||||
| Preprocessing (data gen) | `scripts/<family>/generate_*.py` / `train_preprocess.py` (+ `.sh`) | Offline multi-GPU generation of cached training data (latents / ODE pairs / embeddings) → per-sample `.safetensors` + `outputs.json`, loaded by `ImageVideoSafetensorsDataset`. |
|
||||
| ComfyUI nodes | `comfyui/<family>/nodes.py` | Optional node integration mirroring the pipeline. |
|
||||
|
||||
## Integration workflow
|
||||
|
||||
Copy this checklist and track progress:
|
||||
|
||||
```
|
||||
Integration Progress:
|
||||
- [ ] Step 0: Choose the closest family to mirror; read it across all layers
|
||||
- [ ] Step 1: Model definitions in videox_fun/models/ + register in models/__init__.py
|
||||
- [ ] Step 2: Pipeline in videox_fun/pipeline/ + register in pipeline/__init__.py
|
||||
- [ ] Step 3: Config YAML in config/<family>/
|
||||
- [ ] Step 4: Inference script(s) in examples/<family>/predict_*.py
|
||||
- [ ] Step 5: Training script(s) in scripts/<family>/train*.py + .sh
|
||||
- [ ] Step 6: Training docs README_TRAIN.md + README_TRAIN_zh-CN.md
|
||||
- [ ] Step 7: Reuse audit + verification (incl. smoke test on the matching demo dataset)
|
||||
```
|
||||
|
||||
**Step 0 — Choose the mirror.** Match by task and architecture. A new control model mirrors an existing `*_fun`/`*_control` family; a new audio/talking model mirrors `minimax_h3`/`longcatvideo`/`infinitetalk`; a new image model mirrors `qwenimage`/`flux2`/`z_image`.
|
||||
|
||||
**Step 1 — Model.** Create `videox_fun/models/<family>_transformer3d.py` (or `2d`), `<family>_vae.py`, encoders as needed. Mirror the class shape: `class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin)`, `_supports_gradient_checkpointing = True`, `@register_to_config __init__`, and a `from_pretrained` that supports `transformer_additional_kwargs`, `dict_mapping`, `low_cpu_mem_usage`, and missing-key init. Add imports to `videox_fun/models/__init__.py`.
|
||||
|
||||
**Step 2 — Pipeline.** Create `videox_fun/pipeline/pipeline_<family>.py`. Mirror `pipeline_wan.py`: module-level `retrieve_timesteps`, a `<Family>PipelineOutput(BaseOutput)` dataclass, `<Family>Pipeline(DiffusionPipeline)` with `model_cpu_offload_seq`, `_callback_tensor_inputs`, `__init__(vae, tokenizer, text_encoder, transformer, scheduler, ...)`, `encode_prompt`, and `__call__`. Add imports/aliases to `videox_fun/pipeline/__init__.py`.
|
||||
|
||||
**Step 3 — Config (optional).** A YAML under `config/<family>/` is **not always required**. It is needed mainly for **civitai-format / custom single-file layouts** — to supply `transformer_additional_kwargs`, `dict_mapping` (civitai key → `__init__` kwarg), component subpaths, and `vae_kwargs`/`text_encoder_kwargs`/`scheduler_kwargs`/`image_encoder_kwargs`. For a **standard diffusers-layout** checkpoint (`model_index.json` + per-subfolder `config.json`), load directly via `from_pretrained(model_name, subfolder=...)` with no YAML — mirror `examples/minimax_h3_fun/predict_v2v_control.py`, which guards `if config_path is not None:`. When you do add a YAML, load it via `OmegaConf.load(config_path)` and spread into `from_pretrained` instead of hardcoding those values.
|
||||
|
||||
**Step 4 — Inference script.** Create `examples/<family>/predict_<task>.py` following the exact template (config block at top → component loading → scheduler dict → pipeline construction → multi-GPU/FSDP/compile → `GPU_memory_mode` branching → TeaCache → LoRA merge → inference → `save_results`). See [examples.md](examples.md).
|
||||
|
||||
**Step 5 — Training script.** Create `scripts/<family>/train.py` (+ `train_lora.py` etc.). Mirror the shared structure: license header, `sys.path` bootstrap, imports from `videox_fun`, `log_validation()` that **reuses the inference Pipeline**, `parse_args()` (reuse the existing shared argument set), `main()`. Add a `train.sh` launcher. Reuse `videox_fun.data` datasets/samplers — do not write a new dataset.
|
||||
|
||||
**Step 6 — Docs.** Write `README_TRAIN.md` and `README_TRAIN_zh-CN.md` as an aligned bilingual pair (same structure, same commands/params, matching section order).
|
||||
|
||||
**Step 7 — Reuse audit + verification.** Confirm you reused shared infra (below), smoke-test the new train/predict path on the **matching official demo dataset** under `datasets/X-Fun-*-Demo/` (pick by task and metadata variant — see reference.md §8), then run the verification checklist. Never invent an ad-hoc test set and never leave `datasets/internal_datasets/` placeholders in shipped scripts/docs.
|
||||
|
||||
## Reuse inventory (use these, do not reimplement)
|
||||
|
||||
**Reuse-first catalog: import from here instead of reimplementing. If a helper you need is not listed, grep `videox_fun/` and the closest family before writing your own.**
|
||||
|
||||
- **Schedulers**: `FlowMatchEulerDiscreteScheduler`, `videox_fun.utils.fm_solvers.FlowDPMSolverMultistepScheduler`, `fm_solvers_unipc.FlowUniPCMultistepScheduler`. Selected via a `sampler_name` dict.
|
||||
- **LoRA**: `videox_fun.utils.lora_utils` — `merge_lora`, `unmerge_lora`, `create_network`, `convert_peft_lora_to_kohya_lora`.
|
||||
- **FP8 / quantization**: `videox_fun.utils.fp8_optimization` — `convert_model_weight_to_float8`, `convert_weight_dtype_wrapper`, `replace_parameters_by_name`.
|
||||
- **Offloading**: `videox_fun.utils.group_offload` — `register_auto_device_hook`, `safe_enable_group_offload`; plus pipeline `enable_sequential_cpu_offload` / `enable_model_cpu_offload` / `.to(device)`.
|
||||
- **Distributed**: `videox_fun.dist` — `set_multi_gpus_devices`, `shard_model` (FSDP), `<family>_xfuser` sequence-parallel attention processors, `enable_multi_gpus_inference()`.
|
||||
- **IO / helpers**: `videox_fun.utils.utils` — `save_videos_grid`, `save_videos_with_audio_grid`, `get_image_to_video_latent`, `get_video_to_video_latent`, `get_image_latent`, `filter_kwargs`, `calculate_dimensions`.
|
||||
- **Data**: `videox_fun.data` — `ImageVideoDataset`, `VideoDataset`, `ImageVideoControlDataset`, `VideoSpeechDataset`, bucket/aspect-ratio samplers, `get_closest_ratio`, `get_random_mask`.
|
||||
- **Caching / speedups**: TeaCache (`models/cache_utils`, `get_teacache_coefficients`, `transformer.enable_teacache`), `enable_cfg_skip`, Riflex (`enable_riflex`), `torch.compile` on `transformer.blocks`.
|
||||
- **Preprocessing (data gen, multi-GPU)**: mirror `scripts/wan2.1_self_forcing/generate_ode_pairs.py` — `accelerate launch` + `Accelerator` (interleaved rank sharding), config-driven `from_pretrained` for the teacher/VAE/text-encoder, `safetensors.torch.save_file` per sample + `outputs.json` index, consumed by `videox_fun.data.ImageVideoSafetensorsDataset`. Store as **safetensors only — never LMDB or `.pt`** (see reference.md §10).
|
||||
|
||||
## Non-negotiable conventions
|
||||
|
||||
- **Maximize reuse of existing repo code**: import existing `videox_fun/` helpers and mirror the closest family; never fork or copy-paste a util, and never write a parallel pipeline / loader / scheduler / sampler / offload. Genuinely-new shared code goes in `videox_fun/{utils,data,dist}` (so the next model reuses it), not buried in a family folder.
|
||||
- **`sys.path` bootstrap**: every runnable script starts with the 3-level `project_roots` loop inserting into `sys.path` before importing `videox_fun`.
|
||||
- **Config-driven loading (YAML optional)**: a `config/<family>/*.yaml` is required for civitai-format/custom layouts (it supplies `transformer_additional_kwargs`/`dict_mapping`/subpaths); it is **optional for standard diffusers-layout checkpoints**, which load directly via `from_pretrained(model_name, subfolder=...)`. When a YAML is used, don't hardcode the values it provides.
|
||||
- **`GPU_memory_mode`**: support the standard six modes — `model_full_load`, `model_full_load_and_qfloat8`, `model_cpu_offload`, `model_cpu_offload_and_qfloat8`, `model_group_offload`, `sequential_cpu_offload` — with the exact branching order used in existing `predict_*.py`.
|
||||
- **Naming**: files `<family>_transformer3d.py` / `<family>_vae.py` / `pipeline_<family>.py`; classes `<Family>Transformer3DModel` / `AutoencoderKL<Family>` / `<Family>Pipeline`.
|
||||
- **Resolution args**: drive canvas size with a single square `--video_sample_size` (`type=int`, height = width); never `--video_sample_height` / `--video_sample_width`. For a fixed non-square shape add `--fix_sample_size` (`nargs=2, type=int`, `[height, width]`) that overrides the square size, and derive the effective height/width once in `parse_args()` (see reference.md §5).
|
||||
- **Registries**: a model is not integrated until it is imported in BOTH `videox_fun/models/__init__.py` and `videox_fun/pipeline/__init__.py`.
|
||||
- **Two weight formats**: support `civitai` and `diffusers` via config `format` + `dict_mapping` (maps civitai keys such as `in_dim`→`in_channels`, `dim`→`hidden_size`).
|
||||
- **Bilingual docs**: training READMEs ship as EN + `_zh-CN` pairs with aligned structure and identical commands/params.
|
||||
- **Test data = official demo datasets**: smoke tests, `log_validation` checks, launcher `.sh` defaults, and doc examples all point at `datasets/X-Fun-*-Demo/` (ModelScope `PAI/<name>`), with the metadata variant matching the task — `metadata_add_width_height.json` by default, `_add_objects.json` for VACE/subject-reference, `_add_wav.json` for audio-visual joint models, `metadata_lingbot_video_add_width_height.json` for `lingbot_video`. Selection matrix: reference.md §8.
|
||||
- **Preprocessing = offline data generation, multi-GPU + safetensors**: cached training data (latents / ODE pairs / embeddings) is produced by `accelerate launch` scripts like `generate_ode_pairs.py` (interleaved rank sharding, resume by skipping existing files, `wait_for_everyone`, rank-0 JSON index) and saved with `safetensors.torch.save_file` + an `outputs.json` index for `ImageVideoSafetensorsDataset`. **Never single-GPU / `cuda:0`; never LMDB or `.pt`/`torch.save` pickles for preprocessed data.** See reference.md §10.
|
||||
|
||||
## Verification checklist
|
||||
|
||||
- [ ] New model classes imported in `videox_fun/models/__init__.py`
|
||||
- [ ] New pipeline(s) imported in `videox_fun/pipeline/__init__.py`
|
||||
- [ ] Config YAML present **only if** the checkpoint is civitai-format/custom-layout; a diffusers-layout model may load directly via `from_pretrained(model_name, subfolder=...)` with no YAML. When a YAML is used, it drives component loading (no hardcoded kwargs)
|
||||
- [ ] `predict_*.py` mirrors an existing script: `sys.path` bootstrap, config block, scheduler dict, `GPU_memory_mode` branching, LoRA merge, `save_results`
|
||||
- [ ] `train*.py` reuses `videox_fun.data` + shared args, and `log_validation()` reuses the inference Pipeline
|
||||
- [ ] `train*.sh` launcher provided (`accelerate launch` / DeepSpeed)
|
||||
- [ ] Shared infra reused (schedulers / lora_utils / fp8 / group_offload / dist / utils / data) — nothing reimplemented
|
||||
- [ ] Any offline data-generation/preprocessing script runs multi-GPU (`accelerate launch` + `Accelerator`) and saves cached tensors as **safetensors + `outputs.json`** for `ImageVideoSafetensorsDataset` — never LMDB or `.pt`
|
||||
- [ ] `README_TRAIN.md` + `README_TRAIN_zh-CN.md` aligned pair present
|
||||
- [ ] Smoke test / doc examples use the matching `datasets/X-Fun-*-Demo` dataset and the correct `metadata*.json` variant — no `internal_datasets` placeholders (reference.md §8)
|
||||
- [ ] Optional: ComfyUI node in `comfyui/<family>/nodes.py` mirrors the pipeline
|
||||
|
||||
## Additional resources
|
||||
|
||||
- Detailed file-by-file conventions, class/method shapes, and the model-loading internals: [reference.md](reference.md)
|
||||
- **Dataset & sampler selection matrix** (which `videox_fun.data` dataset/loader each training task uses), **demo-dataset / metadata-variant selection matrix** (which `datasets/X-Fun-*-Demo` to smoke-test with), **inference task matrix** (which pipeline each `predict_<task>.py` uses), and **multi-GPU preprocessing patterns**: [reference.md](reference.md) §8–§10
|
||||
- Concrete skeletons (config YAML, `predict_*.py`, pipeline class, training script + DataLoader): [examples.md](examples.md)
|
||||
@@ -1,510 +0,0 @@
|
||||
# VideoX-Fun Integration Skeletons
|
||||
|
||||
Starting templates. **Always open the mirrored family's real file and adapt it** — these skeletons show shape and required reuse points, not full implementations. Replace `<family>` / `<Family>` / `<task>`.
|
||||
|
||||
## Config — `config/<family>/<variant>.yaml` (optional)
|
||||
|
||||
> **Not always required.** Author a YAML only for civitai-format / custom single-file layouts. A standard diffusers-layout checkpoint (`model_index.json` + per-subfolder `config.json`) loads directly via `from_pretrained(model_name, subfolder=...)` with no YAML — set `config_path = None` and guard `if config_path is not None:` (see `examples/minimax_h3_fun/predict_v2v_control.py`).
|
||||
|
||||
```yaml
|
||||
format: civitai
|
||||
pipeline: <Family>
|
||||
transformer_additional_kwargs:
|
||||
transformer_subpath: ./
|
||||
dict_mapping:
|
||||
in_dim: in_channels
|
||||
dim: hidden_size
|
||||
|
||||
vae_kwargs:
|
||||
vae_subpath: <Family>_VAE.pth
|
||||
temporal_compression_ratio: 4
|
||||
spatial_compression_ratio: 8
|
||||
|
||||
text_encoder_kwargs:
|
||||
text_encoder_subpath: <text_encoder>.pth
|
||||
tokenizer_subpath: <tokenizer_id>
|
||||
text_length: 512
|
||||
|
||||
scheduler_kwargs:
|
||||
scheduler_subpath: null
|
||||
num_train_timesteps: 1000
|
||||
shift: 5.0
|
||||
|
||||
# Only for i2v / models with a CLIP image encoder:
|
||||
image_encoder_kwargs:
|
||||
image_encoder_subpath: <image_encoder>.pth
|
||||
```
|
||||
|
||||
## Inference — `examples/<family>/predict_<task>.py`
|
||||
|
||||
```python
|
||||
import os
|
||||
import sys
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from diffusers import FlowMatchEulerDiscreteScheduler
|
||||
from omegaconf import OmegaConf
|
||||
from PIL import Image
|
||||
from transformers import AutoTokenizer
|
||||
|
||||
# --- sys.path bootstrap (required, before importing videox_fun) ---
|
||||
current_file_path = os.path.abspath(__file__)
|
||||
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
|
||||
for project_root in project_roots:
|
||||
sys.path.insert(0, project_root) if project_root not in sys.path else None
|
||||
|
||||
from videox_fun.dist import set_multi_gpus_devices, shard_model
|
||||
from videox_fun.models import (AutoencoderKL<Family>, <Family>TextEncoder,
|
||||
<Family>Transformer3DModel)
|
||||
from videox_fun.models.cache_utils import get_teacache_coefficients
|
||||
from videox_fun.pipeline import <Family>Pipeline
|
||||
from videox_fun.utils import register_auto_device_hook, safe_enable_group_offload
|
||||
from videox_fun.utils.fm_solvers import FlowDPMSolverMultistepScheduler
|
||||
from videox_fun.utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
|
||||
from videox_fun.utils.fp8_optimization import (convert_model_weight_to_float8,
|
||||
convert_weight_dtype_wrapper,
|
||||
replace_parameters_by_name)
|
||||
from videox_fun.utils.lora_utils import merge_lora, unmerge_lora
|
||||
from videox_fun.utils.utils import (filter_kwargs, get_image_to_video_latent,
|
||||
save_videos_grid)
|
||||
|
||||
# --- user config block (keep the conventional order + comments) ---
|
||||
GPU_memory_mode = "sequential_cpu_offload"
|
||||
ulysses_degree = 1
|
||||
ring_degree = 1
|
||||
fsdp_dit = False
|
||||
fsdp_text_encoder = True
|
||||
compile_dit = False
|
||||
enable_teacache = True
|
||||
teacache_threshold = 0.10
|
||||
num_skip_start_steps = 5
|
||||
teacache_offload = False
|
||||
cfg_skip_ratio = 0
|
||||
enable_riflex = False
|
||||
riflex_k = 6
|
||||
config_path = "config/<family>/<variant>.yaml"
|
||||
model_name = "models/Diffusion_Transformer/<Family>-Model"
|
||||
sampler_name = "Flow"
|
||||
shift = 3
|
||||
transformer_path = None
|
||||
vae_path = None
|
||||
lora_path = None
|
||||
sample_size = [480, 832]
|
||||
video_length = 81
|
||||
fps = 16
|
||||
weight_dtype = torch.bfloat16
|
||||
prompt = "..."
|
||||
negative_prompt = "..."
|
||||
guidance_scale = 6.0
|
||||
seed = 43
|
||||
num_inference_steps = 50
|
||||
lora_weight = 0.55
|
||||
save_path = "samples/<family>-<task>"
|
||||
|
||||
# --- device + config (config_path may be None for a diffusers-layout checkpoint) ---
|
||||
device = set_multi_gpus_devices(ulysses_degree, ring_degree)
|
||||
config = OmegaConf.load(config_path) # or guard: if config_path is not None: ... (then load components via subfolder=...)
|
||||
|
||||
# --- components (when a YAML is used, paths/kwargs come from config; otherwise pass subfolder=... directly) ---
|
||||
transformer = <Family>Transformer3DModel.from_pretrained(
|
||||
os.path.join(model_name, config['transformer_additional_kwargs'].get('transformer_subpath', 'transformer')),
|
||||
transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs']),
|
||||
low_cpu_mem_usage=True, torch_dtype=weight_dtype,
|
||||
)
|
||||
# optional transformer_path / vae_path override -> load_state_dict(strict=False) + print missing/unexpected
|
||||
vae = AutoencoderKL<Family>.from_pretrained(
|
||||
os.path.join(model_name, config['vae_kwargs'].get('vae_subpath', 'vae')),
|
||||
additional_kwargs=OmegaConf.to_container(config['vae_kwargs']),
|
||||
).to(weight_dtype)
|
||||
tokenizer = AutoTokenizer.from_pretrained(
|
||||
os.path.join(model_name, config['text_encoder_kwargs'].get('tokenizer_subpath', 'tokenizer')))
|
||||
text_encoder = <Family>TextEncoder.from_pretrained(
|
||||
os.path.join(model_name, config['text_encoder_kwargs'].get('text_encoder_subpath', 'text_encoder')),
|
||||
additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']),
|
||||
low_cpu_mem_usage=True, torch_dtype=weight_dtype).eval()
|
||||
|
||||
# --- scheduler selection dict ---
|
||||
Chosen_Scheduler = {
|
||||
"Flow": FlowMatchEulerDiscreteScheduler,
|
||||
"Flow_Unipc": FlowUniPCMultistepScheduler,
|
||||
"Flow_DPM++": FlowDPMSolverMultistepScheduler,
|
||||
}[sampler_name]
|
||||
scheduler = Chosen_Scheduler(**filter_kwargs(Chosen_Scheduler, OmegaConf.to_container(config['scheduler_kwargs'])))
|
||||
|
||||
# --- pipeline ---
|
||||
pipeline = <Family>Pipeline(vae=vae, tokenizer=tokenizer, text_encoder=text_encoder,
|
||||
transformer=transformer, scheduler=scheduler)
|
||||
|
||||
# --- multi-gpu / fsdp / compile ---
|
||||
if ulysses_degree > 1 or ring_degree > 1:
|
||||
from functools import partial
|
||||
transformer.enable_multi_gpus_inference()
|
||||
if fsdp_dit:
|
||||
pipeline.transformer = partial(shard_model, device_id=device, param_dtype=weight_dtype)(pipeline.transformer)
|
||||
if fsdp_text_encoder:
|
||||
pipeline.text_encoder = partial(shard_model, device_id=device, param_dtype=weight_dtype)(pipeline.text_encoder)
|
||||
if compile_dit:
|
||||
for i in range(len(pipeline.transformer.blocks)):
|
||||
pipeline.transformer.blocks[i] = torch.compile(pipeline.transformer.blocks[i])
|
||||
|
||||
# --- GPU_memory_mode branching (keep this exact order) ---
|
||||
if GPU_memory_mode == "sequential_cpu_offload":
|
||||
replace_parameters_by_name(transformer, ["modulation",], device=device)
|
||||
pipeline.enable_sequential_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_group_offload":
|
||||
register_auto_device_hook(pipeline.transformer)
|
||||
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
|
||||
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
|
||||
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
|
||||
convert_weight_dtype_wrapper(transformer, weight_dtype)
|
||||
pipeline.enable_model_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_cpu_offload":
|
||||
pipeline.enable_model_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_full_load_and_qfloat8":
|
||||
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
|
||||
convert_weight_dtype_wrapper(transformer, weight_dtype)
|
||||
pipeline.to(device=device)
|
||||
else:
|
||||
pipeline.to(device=device)
|
||||
|
||||
# --- teacache / cfg_skip / riflex / lora ---
|
||||
coefficients = get_teacache_coefficients(model_name) if enable_teacache else None
|
||||
if coefficients is not None:
|
||||
pipeline.transformer.enable_teacache(coefficients, num_inference_steps, teacache_threshold,
|
||||
num_skip_start_steps=num_skip_start_steps, offload=teacache_offload)
|
||||
if cfg_skip_ratio is not None:
|
||||
pipeline.transformer.enable_cfg_skip(cfg_skip_ratio, num_inference_steps)
|
||||
generator = torch.Generator(device=device).manual_seed(seed)
|
||||
if lora_path is not None:
|
||||
pipeline = merge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
|
||||
|
||||
# --- inference ---
|
||||
with torch.no_grad():
|
||||
video_length = int((video_length - 1) // vae.config.temporal_compression_ratio * vae.config.temporal_compression_ratio) + 1 if video_length != 1 else 1
|
||||
if enable_riflex:
|
||||
pipeline.transformer.enable_riflex(k=riflex_k, L_test=(video_length - 1) // vae.config.temporal_compression_ratio + 1)
|
||||
sample = pipeline(prompt, num_frames=video_length, negative_prompt=negative_prompt,
|
||||
height=sample_size[0], width=sample_size[1], generator=generator,
|
||||
guidance_scale=guidance_scale, num_inference_steps=num_inference_steps,
|
||||
shift=shift).videos
|
||||
if lora_path is not None:
|
||||
pipeline = unmerge_lora(pipeline, lora_path, lora_weight, device=device, dtype=weight_dtype)
|
||||
|
||||
# --- save (rank 0 only when multi-gpu) ---
|
||||
def save_results():
|
||||
os.makedirs(save_path, exist_ok=True)
|
||||
prefix = str(len(os.listdir(save_path)) + 1).zfill(8)
|
||||
if video_length == 1:
|
||||
image = (sample[0, :, 0].transpose(0, 1).transpose(1, 2) * 255).numpy().astype(np.uint8)
|
||||
Image.fromarray(image).save(os.path.join(save_path, prefix + ".png"))
|
||||
else:
|
||||
save_videos_grid(sample, os.path.join(save_path, prefix + ".mp4"), fps=fps)
|
||||
|
||||
if ulysses_degree * ring_degree > 1:
|
||||
import torch.distributed as dist
|
||||
if dist.get_rank() == 0:
|
||||
save_results()
|
||||
else:
|
||||
save_results()
|
||||
```
|
||||
|
||||
For i2v, gate the CLIP image encoder and pass `video`/`mask_video`:
|
||||
```python
|
||||
if transformer.config.in_channels != vae.config.latent_channels:
|
||||
clip_image_encoder = CLIPModel.from_pretrained(
|
||||
os.path.join(model_name, config['image_encoder_kwargs'].get('image_encoder_subpath', 'image_encoder'))).to(weight_dtype).eval()
|
||||
input_video, input_video_mask, _ = get_image_to_video_latent(start_image, None, video_length=video_length, sample_size=sample_size)
|
||||
# pipeline = <Family>InpaintPipeline(..., clip_image_encoder=clip_image_encoder)
|
||||
# sample = pipeline(..., video=input_video, mask_video=input_video_mask).videos
|
||||
```
|
||||
|
||||
## Pipeline class — `videox_fun/pipeline/pipeline_<family>.py`
|
||||
|
||||
```python
|
||||
from dataclasses import dataclass
|
||||
from typing import List, Optional, Union
|
||||
import torch
|
||||
from diffusers.callbacks import MultiPipelineCallbacks, PipelineCallback
|
||||
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
|
||||
from diffusers.utils import BaseOutput, logging, replace_example_docstring
|
||||
from diffusers.utils.torch_utils import randn_tensor
|
||||
|
||||
from ..models import AutoencoderKL<Family>, <Family>Transformer3DModel
|
||||
from ..utils.fm_solvers import FlowDPMSolverMultistepScheduler, get_sampling_sigmas
|
||||
from ..utils.fm_solvers_unipc import FlowUniPCMultistepScheduler
|
||||
|
||||
logger = logging.get_logger(__name__)
|
||||
EXAMPLE_DOC_STRING = """Examples:\n```python\npass\n```"""
|
||||
|
||||
# reuse retrieve_timesteps verbatim from pipeline_wan.py
|
||||
|
||||
@dataclass
|
||||
class <Family>PipelineOutput(BaseOutput):
|
||||
videos: torch.Tensor
|
||||
|
||||
class <Family>Pipeline(DiffusionPipeline):
|
||||
model_cpu_offload_seq = "text_encoder->transformer->vae"
|
||||
_callback_tensor_inputs = ["latents", "prompt_embeds", "negative_prompt_embeds"]
|
||||
|
||||
def __init__(self, tokenizer, text_encoder, vae, transformer, scheduler):
|
||||
super().__init__()
|
||||
self.register_modules(tokenizer=tokenizer, text_encoder=text_encoder, vae=vae,
|
||||
transformer=transformer, scheduler=scheduler)
|
||||
# video_processor / vae_scale_factor / etc. as in pipeline_wan.py
|
||||
|
||||
def encode_prompt(self, prompt, negative_prompt, device, num_videos_per_prompt=1, ...):
|
||||
... # mirror pipeline_wan.py
|
||||
|
||||
def prepare_latents(self, batch_size, num_channels_latents, height, width, num_frames, dtype, device, generator, latents=None):
|
||||
...
|
||||
|
||||
@torch.no_grad()
|
||||
@replace_example_docstring(EXAMPLE_DOC_STRING)
|
||||
def __call__(self, prompt, negative_prompt=None, height=480, width=832, num_frames=81,
|
||||
num_inference_steps=50, guidance_scale=6.0, generator=None, shift=1.0,
|
||||
callback_on_step_end=None, return_dict=True, **kwargs) -> Union[<Family>PipelineOutput, tuple]:
|
||||
# 1. encode_prompt 2. prepare_latents 3. retrieve_timesteps
|
||||
# 4. denoising loop with guidance 5. vae.decode 6. return <Family>PipelineOutput(videos=...)
|
||||
...
|
||||
```
|
||||
Then register in `videox_fun/pipeline/__init__.py`:
|
||||
```python
|
||||
from .pipeline_<family> import <Family>Pipeline
|
||||
```
|
||||
|
||||
## Model class — `videox_fun/models/<family>_transformer3d.py`
|
||||
|
||||
```python
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.loaders.single_file_model import FromOriginalModelMixin
|
||||
from diffusers.models.modeling_utils import ModelMixin
|
||||
from .attention_utils import attention # unified FA/SDPA backend — do not hand-roll SDPA
|
||||
|
||||
class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
_supports_gradient_checkpointing = True
|
||||
|
||||
@register_to_config
|
||||
def __init__(self, model_type='t2v', in_dim=16, dim=2048, ffn_dim=8192,
|
||||
num_heads=16, num_layers=32, in_channels=16, hidden_size=2048, ...):
|
||||
super().__init__()
|
||||
...
|
||||
|
||||
def _set_gradient_checkpointing(self, *args, **kwargs):
|
||||
self.gradient_checkpointing = True
|
||||
|
||||
def enable_multi_gpus_inference(self): ... # route attn through dist/<family>_xfuser.py
|
||||
def enable_teacache(self, ...): ...
|
||||
def enable_cfg_skip(self, ...): ...
|
||||
|
||||
def forward(self, x, timestep, context, ...): ...
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(cls, pretrained_model_path, subfolder=None,
|
||||
transformer_additional_kwargs=None, low_cpu_mem_usage=False,
|
||||
torch_dtype=torch.bfloat16):
|
||||
... # mirror wan_transformer3d.py: config.json -> dict_mapping -> init_empty_weights
|
||||
# -> load .bin/.safetensors -> shape-filter -> initialize missing keys -> load
|
||||
```
|
||||
Then register in `videox_fun/models/__init__.py`:
|
||||
```python
|
||||
from .<family>_transformer3d import <Family>Transformer3DModel
|
||||
from .<family>_vae import AutoencoderKL<Family>
|
||||
```
|
||||
|
||||
## Training — `scripts/<family>/train.py` (key reuse points)
|
||||
|
||||
```python
|
||||
"""Modified from https://github.com/huggingface/diffusers/.../train_text_to_image.py"""
|
||||
import argparse, gc, logging, math, os, sys
|
||||
import accelerate, diffusers, torch, transformers
|
||||
from accelerate import Accelerator
|
||||
from diffusers.optimization import get_scheduler
|
||||
from omegaconf import OmegaConf
|
||||
|
||||
# same sys.path bootstrap as predict scripts
|
||||
from videox_fun.data import (ASPECT_RATIO_512, AspectRatioBatchImageVideoSampler,
|
||||
ImageVideoDataset, ImageVideoSampler, RandomSampler,
|
||||
get_closest_ratio, get_random_mask)
|
||||
from videox_fun.models import AutoencoderKL<Family>, <Family>Transformer3DModel
|
||||
from videox_fun.pipeline import <Family>Pipeline # REUSED for validation
|
||||
from videox_fun.utils.lora_utils import create_network # for train_lora
|
||||
from videox_fun.utils.utils import save_videos_grid, get_image_to_video_latent
|
||||
|
||||
def log_validation(vae, text_encoder, tokenizer, transformer3d, args, config,
|
||||
accelerator, weight_dtype, global_step):
|
||||
# build <Family>Pipeline from accelerator.unwrap_model(transformer3d),
|
||||
# run validation_prompts, save_videos_grid to output_dir/sample/. Reuse the pipeline.
|
||||
...
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(...)
|
||||
# reuse the shared arg surface: --config_path, --pretrained_model_name_or_path,
|
||||
# --train_data_dir, --train_data_meta, --video_sample_n_frames, --train_batch_size,
|
||||
# --gradient_accumulation_steps, --learning_rate, --lr_scheduler, --checkpointing_steps,
|
||||
# --output_dir, --mixed_precision, --gradient_checkpointing, --enable_bucket,
|
||||
# --train_mode, --trainable_modules, --validation_prompts ... (add only what's needed)
|
||||
return parser.parse_args()
|
||||
|
||||
def main():
|
||||
args = parse_args()
|
||||
accelerator = Accelerator(mixed_precision=args.mixed_precision, ...)
|
||||
config = OmegaConf.load(args.config_path)
|
||||
# load transformer/vae/text_encoder via config
|
||||
|
||||
# --- Dataset: pick by task (see reference.md §8) ---
|
||||
# T2V/I2V base + inpaint -> ImageVideoDataset(enable_inpaint = args.train_mode != "normal")
|
||||
# Control -> ImageVideoControlDataset(enable_camera_info = ...)
|
||||
# Image edit -> ImageEditDataset
|
||||
# Speech/audio (S2V) -> VideoSpeechDataset / VideoSpeechControlDataset
|
||||
# Animate -> VideoAnimateDataset
|
||||
# Distill text / GRPO / DPO -> TextDataset
|
||||
# Smoke-test on the matching official demo dataset (reference.md §8), e.g.
|
||||
# datasets/X-Fun-Videos-Demo + metadata_add_width_height.json for T2V/I2V.
|
||||
train_dataset = ImageVideoDataset(
|
||||
args.train_data_meta, args.train_data_dir,
|
||||
video_sample_size=args.video_sample_size, video_sample_stride=args.video_sample_stride,
|
||||
video_sample_n_frames=args.video_sample_n_frames, video_repeat=args.video_repeat,
|
||||
image_sample_size=args.image_sample_size, enable_bucket=args.enable_bucket,
|
||||
enable_inpaint=True if args.train_mode != "normal" else False)
|
||||
|
||||
# --- Sampler + DataLoader: branch on enable_bucket (see reference.md §8) ---
|
||||
batch_sampler_generator = torch.Generator().manual_seed(args.seed)
|
||||
if args.enable_bucket:
|
||||
aspect_ratio_sample_size = {k: [x / 512 * args.video_sample_size for x in ASPECT_RATIO_512[k]] for k in ASPECT_RATIO_512}
|
||||
batch_sampler = AspectRatioBatchImageVideoSampler(
|
||||
sampler=RandomSampler(train_dataset, generator=batch_sampler_generator), dataset=train_dataset.dataset,
|
||||
batch_size=args.train_batch_size, train_folder=args.train_data_dir, drop_last=True,
|
||||
aspect_ratios=aspect_ratio_sample_size)
|
||||
def collate_fn(examples):
|
||||
new_examples = {"pixel_values": [], "text": []}
|
||||
if args.train_mode != "normal":
|
||||
new_examples.update({"mask_pixel_values": [], "mask": [], "clip_pixel_values": []})
|
||||
# get_closest_ratio -> Resize/CenterCrop/Normalize -> stack; masks via get_random_mask
|
||||
return new_examples
|
||||
train_dataloader = torch.utils.data.DataLoader(
|
||||
train_dataset, batch_sampler=batch_sampler, collate_fn=collate_fn,
|
||||
num_workers=args.dataloader_num_workers,
|
||||
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
|
||||
else:
|
||||
batch_sampler = ImageVideoSampler(RandomSampler(train_dataset, generator=batch_sampler_generator), train_dataset, args.train_batch_size)
|
||||
train_dataloader = torch.utils.data.DataLoader(
|
||||
train_dataset, batch_sampler=batch_sampler, num_workers=args.dataloader_num_workers,
|
||||
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
|
||||
|
||||
# trainable-module filtering or create_network for LoRA
|
||||
# optimizer + get_scheduler; accelerator.prepare; checkpoint hooks
|
||||
# training loop: timestep sampling -> transformer forward -> loss -> backward
|
||||
# periodic log_validation(...); final save weights / LoRA
|
||||
...
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
## Launcher — `scripts/<family>/train.sh`
|
||||
|
||||
```bash
|
||||
export MODEL_NAME="models/Diffusion_Transformer/<Family>-Model"
|
||||
# Test data = the official demo dataset matching the task (reference.md §8). Download once, e.g.:
|
||||
# modelscope download --dataset PAI/X-Fun-Videos-Demo --local_dir ./datasets/X-Fun-Videos-Demo
|
||||
# T2I -> X-Fun-Images-Demo | control -> X-Fun-{Videos,Images}-Controls-Demo
|
||||
# S2V -> X-Fun-Videos-Audios-Demo | image edit -> X-Fun-Images-Edit-Demo
|
||||
export DATASET_NAME="datasets/X-Fun-Videos-Demo/" # = train_data_dir (data_root); media live under train/
|
||||
export DATASET_META_NAME="datasets/X-Fun-Videos-Demo/metadata_add_width_height.json" # = train_data_meta: [{"file_path","text","type","width","height"}] — see reference.md §8
|
||||
# Metadata variants: VACE/subject-ref -> metadata_add_width_height_add_objects.json (X-Fun-Videos-Controls-Demo);
|
||||
# audio-visual joint -> metadata_add_width_height_add_wav.json; lingbot_video -> metadata_lingbot_video_add_width_height.json
|
||||
NCCL_DEBUG=INFO
|
||||
|
||||
accelerate launch --mixed_precision="bf16" scripts/<family>/train.py \
|
||||
--config_path="config/<family>/<variant>.yaml" \
|
||||
--pretrained_model_name_or_path=$MODEL_NAME \
|
||||
--train_data_dir=$DATASET_NAME \
|
||||
--train_data_meta=$DATASET_META_NAME \
|
||||
--video_sample_n_frames=81 \
|
||||
--train_batch_size=1 \
|
||||
--gradient_accumulation_steps=1 \
|
||||
--learning_rate=2e-05 \
|
||||
--lr_scheduler="constant_with_warmup" \
|
||||
--lr_warmup_steps=100 \
|
||||
--checkpointing_steps=50 \
|
||||
--output_dir="output_dir_<family>" \
|
||||
--gradient_checkpointing \
|
||||
--mixed_precision="bf16" \
|
||||
--enable_bucket \
|
||||
--low_vram \
|
||||
--train_mode="normal" \
|
||||
--trainable_modules "."
|
||||
```
|
||||
|
||||
## Preprocessing (data gen) — `scripts/<family>/generate_<...>.py`
|
||||
|
||||
Offline generation of cached training data (latents / ODE-trajectory pairs / prompt embeddings). **Always multi-GPU** (`accelerate launch` + `Accelerator`) and **always safetensors** (`safetensors.torch.save_file` + an `outputs.json` index for `ImageVideoSafetensorsDataset`) — never LMDB, never `.pt`. Mirror `scripts/wan2.1_self_forcing/generate_ode_pairs.py`:
|
||||
|
||||
```python
|
||||
# ...license header + sys.path bootstrap...
|
||||
import argparse, json, math, os, torch
|
||||
from accelerate import Accelerator
|
||||
from omegaconf import OmegaConf
|
||||
from safetensors.torch import save_file
|
||||
from tqdm import tqdm
|
||||
from videox_fun.models import AutoencoderKLWan, WanT5EncoderModel, WanTransformer3DModel # reuse repo models
|
||||
from videox_fun.utils.utils import save_videos_grid # reuse repo IO
|
||||
|
||||
def main():
|
||||
args = parse_args() # --pretrained_model_name_or_path --config_path --caption_path --output_folder
|
||||
# --num_inference_steps --guidance_scale --shift --mixed_precision ...
|
||||
accelerator = Accelerator(mixed_precision=args.mixed_precision)
|
||||
device, world_size, rank = accelerator.device, accelerator.num_processes, accelerator.process_index
|
||||
torch.set_grad_enabled(False) # inference-only
|
||||
torch.backends.cuda.matmul.allow_tf32 = True
|
||||
|
||||
config = OmegaConf.load(args.config_path) # config-driven loading (Section 3)
|
||||
weight_dtype = {"fp16": torch.float16, "bf16": torch.bfloat16}.get(accelerator.mixed_precision, torch.float32)
|
||||
text_encoder = WanT5EncoderModel.from_pretrained(..., additional_kwargs=OmegaConf.to_container(config['text_encoder_kwargs']), torch_dtype=weight_dtype).to(device).eval()
|
||||
vae = AutoencoderKLWan.from_pretrained(..., additional_kwargs=OmegaConf.to_container(config['vae_kwargs'])).to(device, dtype=weight_dtype).eval()
|
||||
transformer = WanTransformer3DModel.from_pretrained(..., transformer_additional_kwargs=OmegaConf.to_container(config['transformer_additional_kwargs'])).to(device, dtype=weight_dtype).eval()
|
||||
|
||||
prompts = [l.rstrip() for l in open(args.caption_path, encoding="utf-8") if l.strip()]
|
||||
os.makedirs(args.output_folder, exist_ok=True)
|
||||
total_per_rank = math.ceil(len(prompts) / world_size)
|
||||
|
||||
for index in tqdm(range(total_per_rank), disable=rank != 0, desc="Generating"):
|
||||
prompt_index = index * world_size + rank # interleaved multi-GPU shard
|
||||
if prompt_index >= len(prompts):
|
||||
continue
|
||||
out_path = os.path.join(args.output_folder, f"{prompt_index:05d}.safetensors")
|
||||
if os.path.exists(out_path): # resume: skip already-done samples
|
||||
continue
|
||||
prompt = prompts[prompt_index]
|
||||
# ... encode prompt, sample noise, run the teacher ODE (CFG), collect latents ...
|
||||
save_file( # safetensors ONLY (no lmdb / no .pt)
|
||||
{"latents": latents.cpu(), "prompt_embeds": text_embeds.cpu(), "prompt_attention_mask": mask.cpu()},
|
||||
out_path, metadata={"prompt": prompt},
|
||||
)
|
||||
|
||||
accelerator.wait_for_everyone()
|
||||
if accelerator.is_main_process: # rank-0 writes the JSON index
|
||||
entries = [{"file_path": os.path.join(args.output_folder, f"{i:05d}.safetensors")}
|
||||
for i in range(len(prompts))
|
||||
if os.path.exists(os.path.join(args.output_folder, f"{i:05d}.safetensors"))]
|
||||
json.dump(entries, open(os.path.join(args.output_folder, "outputs.json"), "w"), ensure_ascii=False, indent=4)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
Launcher (`generate_<...>.sh`) — `accelerate launch` uses every visible GPU:
|
||||
```bash
|
||||
export MODEL_NAME="models/Diffusion_Transformer/Wan2.1-T2V-1.3B"
|
||||
accelerate launch --mixed_precision="bf16" scripts/<family>/generate_<...>.py \
|
||||
--pretrained_model_name_or_path=$MODEL_NAME \
|
||||
--config_path="config/<family>/*.yaml" \
|
||||
--caption_path="datasets/prompts.txt" \
|
||||
--output_folder="datasets/<family>_ode_pairs" \
|
||||
--num_inference_steps=48 --guidance_scale=6.0 --shift=8.0
|
||||
```
|
||||
|
||||
Training then reads the cache with `ImageVideoSafetensorsDataset(ann_path=".../outputs.json")` (single-file mode `{"file_path": ...}`, or per-tensor mode via `--save_per_tensor`). See reference.md §10.
|
||||
|
||||
> Dataset *curation* (scoring/filtering/captioning under `videox_fun/video_caption/`) is a different activity: also multi-GPU (accelerate `PartialState.split_between_processes`/`gather_object`, or vLLM tensor-parallel) but writes csv/jsonl metadata, not safetensors. See reference.md §10 “Related but different”.
|
||||
@@ -1,414 +0,0 @@
|
||||
# VideoX-Fun Integration Reference
|
||||
|
||||
Detailed conventions per layer. Read the mirrored family's real files alongside this — the existing code is always the source of truth.
|
||||
|
||||
## 1. Model definitions — `videox_fun/models/<family>_*.py`
|
||||
|
||||
### File naming
|
||||
- Transformer / DiT: `<family>_transformer3d.py` (video) or `<family>_transformer2d.py` (image). Variants append a suffix: `_control`, `_s2v`, `_vace`, `_animate`, `_self_forcing`, `_avatar`.
|
||||
- VAE: `<family>_vae.py` → class `AutoencoderKL<Family>`.
|
||||
- Encoders: `<family>_text_encoder.py`, `<family>_audio_encoder.py`, `<family>_image_encoder.py`.
|
||||
|
||||
### Class shape (mirror `wan_transformer3d.py`)
|
||||
```python
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.loaders.single_file_model import FromOriginalModelMixin
|
||||
from diffusers.models.modeling_utils import ModelMixin
|
||||
|
||||
class <Family>Transformer3DModel(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
_supports_gradient_checkpointing = True
|
||||
|
||||
@register_to_config
|
||||
def __init__(self, model_type='t2v', patch_size=(1,2,2), in_dim=16, dim=2048,
|
||||
ffn_dim=8192, num_heads=16, num_layers=32, in_channels=16,
|
||||
hidden_size=2048, ...):
|
||||
super().__init__()
|
||||
...
|
||||
```
|
||||
- Keep BOTH civitai names (`in_dim`, `dim`, `ffn_dim`) and diffusers aliases (`in_channels`, `hidden_size`) in `__init__` so either format maps cleanly.
|
||||
- Implement `_set_gradient_checkpointing(self, *args, **kwargs)`.
|
||||
- Attention must go through `videox_fun.models.attention_utils.attention` (backend-agnostic), not a hand-rolled `scaled_dot_product_attention`.
|
||||
- Multi-GPU: expose `enable_multi_gpus_inference()` and route attention through the family's `dist/<family>_xfuser.py` processor.
|
||||
- Speedups live on the model: `enable_teacache(...)`, `enable_cfg_skip(...)`, `enable_riflex(...)`.
|
||||
|
||||
### `from_pretrained` internals (do not simplify)
|
||||
The custom classmethod must keep these behaviors (see `wan_transformer3d.py::from_pretrained`):
|
||||
1. Accept `transformer_additional_kwargs`, `subfolder`, `low_cpu_mem_usage`, `torch_dtype`.
|
||||
2. Read `config.json`; auto-convert foreign configs (e.g. diffsynth `has_image_input`) via a `_convert_from_*_config` helper.
|
||||
3. Apply `dict_mapping`: pop it from kwargs, then for each `key: target` set `kwargs[target] = config[key]`.
|
||||
4. Under `low_cpu_mem_usage`, build with `accelerate.init_empty_weights()`, load `.bin`/`.safetensors` (single file or glob all shards), and **filter by exact shape match** before loading.
|
||||
5. Initialize missing keys deliberately: zero-init control/audio projections (`after_proj`, `before_proj`, `processor.k_proj/v_proj`, `audio_injector`, `cond_encoder`, ...), ones for norms, xavier for ≥2D weights, so new branches start as no-ops.
|
||||
|
||||
### Registry — `videox_fun/models/__init__.py`
|
||||
Add an import line for every new public class, grouped with the family. Wrap optional-dependency imports in `try/except` with a helpful upgrade message (see the Qwen2.5-VL / Mistral3 blocks at the top).
|
||||
|
||||
## 2. Pipelines — `videox_fun/pipeline/pipeline_<family>*.py`
|
||||
|
||||
Mirror `pipeline_wan.py`. Required pieces:
|
||||
- Module-level `retrieve_timesteps(scheduler, num_inference_steps, device, timesteps, sigmas, **kwargs)` (copied from diffusers) — reuse verbatim.
|
||||
- `EXAMPLE_DOC_STRING` for the `@replace_example_docstring` decorator.
|
||||
- Output dataclass:
|
||||
```python
|
||||
@dataclass
|
||||
class <Family>PipelineOutput(BaseOutput):
|
||||
videos: torch.Tensor
|
||||
```
|
||||
- Pipeline class:
|
||||
```python
|
||||
class <Family>Pipeline(DiffusionPipeline):
|
||||
_optional_component = [...]
|
||||
model_cpu_offload_seq = "text_encoder->transformer->vae" # order matters for offload
|
||||
_callback_tensor_inputs = ["latents", "prompt_embeds", "negative_prompt_embeds"]
|
||||
def __init__(self, tokenizer, text_encoder, vae, transformer, scheduler, ...): ...
|
||||
def encode_prompt(...): ...
|
||||
def prepare_latents(...): ...
|
||||
@torch.no_grad()
|
||||
@replace_example_docstring(EXAMPLE_DOC_STRING)
|
||||
def __call__(self, prompt, negative_prompt=..., height=..., width=...,
|
||||
num_frames=..., num_inference_steps=..., guidance_scale=...,
|
||||
generator=None, ..., return_dict=True) -> Union[<Family>PipelineOutput, Tuple]: ...
|
||||
```
|
||||
- Import schedulers from `..utils.fm_solvers` / `..utils.fm_solvers_unipc`, models from `..models`.
|
||||
- Separate pipelines per task: base (`pipeline_<family>.py`), inpaint/i2v (`_inpaint`), control (`_control`), s2v, etc. Register all in `videox_fun/pipeline/__init__.py`, adding convenience aliases (e.g. `WanI2VPipeline = WanFunInpaintPipeline`) where existing code expects them.
|
||||
|
||||
## 3. Config — `config/<family>/<name>.yaml` (optional)
|
||||
|
||||
**The YAML is not mandatory.** Decide by checkpoint layout:
|
||||
- **Required** for civitai-format / custom single-file layouts, where weights and key names are not diffusers-native. The YAML supplies `transformer_additional_kwargs` (incl. `dict_mapping` mapping civitai config keys → model `__init__` kwargs), component `*_subpath`s, and `vae/text_encoder/scheduler/image_encoder` kwargs.
|
||||
- **Optional** for a standard diffusers-layout checkpoint (`model_index.json` + each subfolder carrying its own `config.json`). Load components directly: `<Family>Transformer3DModel.from_pretrained(model_name, subfolder="transformer", low_cpu_mem_usage=True, torch_dtype=...)`, `AutoencoderKL<Family>.from_pretrained(model_name, subfolder="vae")`, etc. Guard the config path exactly like `examples/minimax_h3_fun/predict_v2v_control.py`:
|
||||
```python
|
||||
transformer_load_kwargs = {}
|
||||
if config_path is not None:
|
||||
from omegaconf import OmegaConf
|
||||
config = OmegaConf.load(config_path)
|
||||
transformer_load_kwargs.update(OmegaConf.to_container(config["transformer_additional_kwargs"], resolve=True))
|
||||
transformer = <Family>Transformer3DModel.from_pretrained(model_name, subfolder="transformer", **transformer_load_kwargs, ...)
|
||||
```
|
||||
|
||||
When you do use a YAML, the canonical schema is below (see `config/wan2.1/wan_civitai.yaml`):
|
||||
```yaml
|
||||
format: civitai # or diffusers — selects weight-key handling
|
||||
pipeline: Wan # family label consumed by API/ComfyUI loaders
|
||||
transformer_additional_kwargs:
|
||||
transformer_subpath: ./ # subfolder under model_name holding the DiT
|
||||
dict_mapping: # civitai config key -> model __init__ kwarg
|
||||
in_dim: in_channels
|
||||
dim: hidden_size
|
||||
vae_kwargs:
|
||||
vae_subpath: Wan2.1_VAE.pth
|
||||
temporal_compression_ratio: 4
|
||||
spatial_compression_ratio: 8
|
||||
text_encoder_kwargs:
|
||||
text_encoder_subpath: models_t5_umt5-xxl-enc-bf16.pth
|
||||
tokenizer_subpath: google/umt5-xxl
|
||||
text_length: 512
|
||||
...
|
||||
scheduler_kwargs:
|
||||
scheduler_subpath: null
|
||||
num_train_timesteps: 1000
|
||||
shift: 5.0
|
||||
...
|
||||
image_encoder_kwargs: # only for i2v / models with a CLIP image encoder
|
||||
image_encoder_subpath: models_clip_...pth
|
||||
```
|
||||
Every `*_subpath` is joined onto `model_name` in scripts. Load with `OmegaConf.load` and pass `OmegaConf.to_container(config['<section>'])` into `from_pretrained`. Use `filter_kwargs(Cls, OmegaConf.to_container(config['scheduler_kwargs']))` to build schedulers.
|
||||
|
||||
## 4. Inference scripts — `examples/<family>/predict_<task>.py`
|
||||
|
||||
Anatomy, top to bottom (see `examples/wan2.1_fun/predict_t2v.py`):
|
||||
1. **`sys.path` bootstrap** (before importing `videox_fun`):
|
||||
```python
|
||||
current_file_path = os.path.abspath(__file__)
|
||||
project_roots = [os.path.dirname(current_file_path), os.path.dirname(os.path.dirname(current_file_path)), os.path.dirname(os.path.dirname(os.path.dirname(current_file_path)))]
|
||||
for project_root in project_roots:
|
||||
sys.path.insert(0, project_root) if project_root not in sys.path else None
|
||||
```
|
||||
2. **User config block** as top-level variables with explanatory comments, in the conventional order: `GPU_memory_mode`, `ulysses_degree`/`ring_degree`, `fsdp_dit`/`fsdp_text_encoder`, `compile_dit`, TeaCache (`enable_teacache`, `teacache_threshold`, `num_skip_start_steps`, `teacache_offload`), `cfg_skip_ratio`, Riflex (`enable_riflex`, `riflex_k`), `config_path`, `model_name`, `sampler_name`, `shift`, `transformer_path`/`vae_path`/`lora_path`, `sample_size`, `video_length`, `fps`, `weight_dtype`, `prompt`/`negative_prompt`, `guidance_scale`, `seed`, `num_inference_steps`, `lora_weight`, `save_path`.
|
||||
3. **Device + config**: `device = set_multi_gpus_devices(ulysses_degree, ring_degree)`; then either `config = OmegaConf.load(config_path)` (civitai/custom layout) **or** guard `if config_path is not None:` and load components directly from a diffusers-layout checkpoint (see §3).
|
||||
4. **Component loading**: transformer (`from_pretrained(..., transformer_additional_kwargs=...)`), optional `transformer_path`/`vae_path` override with `load_state_dict(strict=False)` + missing/unexpected key print, vae, tokenizer, text_encoder, and clip image encoder gated by `transformer.config.in_channels != vae.config.latent_channels`.
|
||||
5. **Scheduler selection dict**: `{"Flow": FlowMatchEulerDiscreteScheduler, "Flow_Unipc": FlowUniPCMultistepScheduler, "Flow_DPM++": FlowDPMSolverMultistepScheduler}[sampler_name]`; build with `filter_kwargs`.
|
||||
6. **Pipeline construction**: choose base vs inpaint/i2v/control pipeline by the model's channel condition.
|
||||
7. **Multi-GPU / FSDP / compile**: if `ulysses_degree>1 or ring_degree>1` call `transformer.enable_multi_gpus_inference()` and optionally `shard_model`; if `compile_dit`, `torch.compile` each `transformer.blocks[i]`.
|
||||
8. **`GPU_memory_mode` branching** — keep this exact order:
|
||||
```python
|
||||
if GPU_memory_mode == "sequential_cpu_offload":
|
||||
replace_parameters_by_name(transformer, ["modulation",], device=device)
|
||||
transformer.freqs = transformer.freqs.to(device=device)
|
||||
pipeline.enable_sequential_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_group_offload":
|
||||
register_auto_device_hook(pipeline.transformer)
|
||||
safe_enable_group_offload(pipeline, onload_device=device, offload_device="cpu", offload_type="leaf_level", use_stream=True)
|
||||
elif GPU_memory_mode == "model_cpu_offload_and_qfloat8":
|
||||
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
|
||||
convert_weight_dtype_wrapper(transformer, weight_dtype)
|
||||
pipeline.enable_model_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_cpu_offload":
|
||||
pipeline.enable_model_cpu_offload(device=device)
|
||||
elif GPU_memory_mode == "model_full_load_and_qfloat8":
|
||||
convert_model_weight_to_float8(transformer, exclude_module_name=["modulation",], device=device)
|
||||
convert_weight_dtype_wrapper(transformer, weight_dtype)
|
||||
pipeline.to(device=device)
|
||||
else:
|
||||
pipeline.to(device=device)
|
||||
```
|
||||
9. **TeaCache / cfg_skip / Riflex** enablement, `generator = torch.Generator(device).manual_seed(seed)`, LoRA `merge_lora`.
|
||||
10. **Inference** under `torch.no_grad()`; align `video_length` to `vae.config.temporal_compression_ratio`; pass `video`/`mask_video` for i2v via `get_image_to_video_latent`.
|
||||
11. **`save_results()`**: `save_videos_grid(sample, path, fps=fps)` for video, PIL save for a single frame; only rank 0 saves when multi-GPU. LoRA `unmerge_lora` after.
|
||||
|
||||
Other entry points to mirror when needed: `app.py` (Gradio), `launch_api.py` (API server backed by `videox_fun/api`), `post_infer*.py` (batch/queue inference).
|
||||
|
||||
## 5. Training scripts — `scripts/<family>/train*.py`
|
||||
|
||||
Mirror `scripts/wan2.1_fun/train.py`. Structure:
|
||||
1. Diffusers-derived license header + `"""Modified from ..."""` note.
|
||||
2. Third-party imports, then the **same `sys.path` bootstrap**, then `from videox_fun.data/models/pipeline/utils import ...`.
|
||||
3. Helper funcs: `filter_kwargs`, `resize_mask`, `linear_decay`, `generate_timestep_with_lognorm`.
|
||||
4. **`log_validation(vae, text_encoder, tokenizer, clip_image_encoder, transformer3d, args, config, accelerator, weight_dtype, global_step)`** — builds the **inference Pipeline** from the live (unwrapped) transformer and runs it to produce sample videos under `output_dir/sample/`. Wrapped in try/except; handles DeepSpeed (`transformer3d.config` swap) and restores VAE/text-encoder placement (`low_vram`). **Reuse the pipeline; never write a separate sampler.**
|
||||
5. **`parse_args()`** — reuse the shared argument surface: `--config_path`, `--pretrained_model_name_or_path`, `--train_data_dir`, `--train_data_meta`, `--image_sample_size`/`--video_sample_size`/`--token_sample_size`, `--video_sample_n_frames`, `--video_sample_stride`, `--train_batch_size`, `--gradient_accumulation_steps`, `--learning_rate`, `--lr_scheduler`, `--lr_warmup_steps`, `--checkpointing_steps`, `--output_dir`, `--mixed_precision`, `--gradient_checkpointing`, `--enable_bucket`, `--random_hw_adapt`, `--training_with_video_token_length`, `--uniform_sampling`, `--low_vram`, `--train_mode`, `--trainable_modules`, LoRA args (`--use_lora`, `--rank`, ...), `--validation_prompts`/`--validation_paths`. Add new args only when the family genuinely needs them.
|
||||
6. **`main()`** — Accelerator setup, DeepSpeed/FSDP zero-stage handling (auto-sets `save_state`), model loading via config, dataset + bucket sampler from `videox_fun.data`, trainable-module filtering / LoRA network via `create_network`, optimizer + `get_scheduler`, `accelerator.prepare`, checkpoint save/load hooks, training loop with timestep sampling, loss, `log_validation` at intervals, and final weight/LoRA save.
|
||||
|
||||
### Resolution args — `--video_sample_size` (+ `--fix_sample_size`)
|
||||
Canvas resolution is always driven by a **single square** `--video_sample_size` (`type=int`, height = width) — never by separate `--video_sample_height` / `--video_sample_width`. When a **fixed non-square shape** is required, add `--fix_sample_size` (`nargs=2, type=int, default=None`, `[height, width]`) that overrides the square size; mirror `scripts/wan2.2_fun/train_lora.py`, `scripts/z_image/train_distill.py`. Derive the effective `height` / `width` once in `parse_args()` and reuse them everywhere downstream:
|
||||
```python
|
||||
parser.add_argument("--video_sample_size", type=int, default=1280)
|
||||
parser.add_argument("--fix_sample_size", nargs=2, type=int, default=None,
|
||||
help="Fix Sample size [height, width] to override `--video_sample_size` with a fixed non-square shape.")
|
||||
...
|
||||
if args.fix_sample_size is not None:
|
||||
args.video_sample_height, args.video_sample_width = args.fix_sample_size
|
||||
else:
|
||||
args.video_sample_height = args.video_sample_width = args.video_sample_size
|
||||
```
|
||||
In bucket datasets `--fix_sample_size` also forces `random_hw_adapt=False` / `training_with_video_token_length=False` and bumps `video_sample_size = max(max(fix_sample_size), video_sample_size)`; in data-free scripts (e.g. `scripts/minimax_h3/train_pdd_lora.py`) it simply pins the generation canvas. Always validate the size against the patch/VAE constraint (minimax_h3: `% 32`). The `.sh` launcher passes it space-separated (`nargs=2`): `--fix_sample_size 768 1344`.
|
||||
|
||||
### Launcher — `scripts/<family>/train*.sh`
|
||||
`export MODEL_NAME/DATASET_NAME/DATASET_META_NAME`, then `accelerate launch --mixed_precision="bf16" scripts/<family>/train.py --config_path=... <full arg list>`. Include commented I2V/control variants and DeepSpeed/NCCL notes as the existing scripts do.
|
||||
|
||||
### Docs — `README_TRAIN.md` + `README_TRAIN_zh-CN.md`
|
||||
Aligned bilingual pair: identical section order, identical commands and parameter tables; only the prose language differs. Follow the top-level section order used across existing training READMEs.
|
||||
|
||||
## 6. Shared infrastructure map (reuse, never reimplement)
|
||||
|
||||
| Need | Import from |
|
||||
|------|-------------|
|
||||
| Flow/DPM/UniPC schedulers | `diffusers`, `videox_fun.utils.fm_solvers`, `videox_fun.utils.fm_solvers_unipc` |
|
||||
| LoRA create/merge/unmerge/convert | `videox_fun.utils.lora_utils` |
|
||||
| FP8 quantization | `videox_fun.utils.fp8_optimization` |
|
||||
| Group / leaf offload hooks | `videox_fun.utils.group_offload` |
|
||||
| Multi-GPU device + FSDP shard + seq-parallel attn | `videox_fun.dist` |
|
||||
| Save video/audio, image→video latents, kwarg filter, dimension calc | `videox_fun.utils.utils` |
|
||||
| Datasets + bucket/aspect-ratio samplers + masks | `videox_fun.data` |
|
||||
| TeaCache coefficients | `videox_fun.models.cache_utils` |
|
||||
|
||||
## 7. Naming quick reference
|
||||
|
||||
| Concept | Convention | Example |
|
||||
|---------|-----------|---------|
|
||||
| Model file | `<family>_transformer3d.py` | `wan_transformer3d.py` |
|
||||
| Model class | `<Family>Transformer3DModel` | `WanTransformer3DModel` |
|
||||
| VAE class | `AutoencoderKL<Family>` | `AutoencoderKLWan` |
|
||||
| Pipeline file | `pipeline_<family>.py` | `pipeline_wan.py` |
|
||||
| Pipeline class | `<Family>Pipeline` | `WanPipeline` / `WanFunInpaintPipeline` |
|
||||
| Config | `config/<family>/<variant>.yaml` | `config/wan2.1/wan_civitai.yaml` |
|
||||
| Inference | `examples/<family>/predict_<task>.py` | `predict_t2v.py`, `predict_i2v.py`, `predict_v2v_control.py` |
|
||||
| Training | `scripts/<family>/train[_<variant>].py` | `train.py`, `train_lora.py`, `train_control.py`, `train_distill.py` |
|
||||
|
||||
## 8. Training data pipeline — dataset & sampler selection
|
||||
|
||||
Pick the dataset by **task / `train_mode`**, then the sampler by **`enable_bucket`** and dataset type. All datasets/samplers come from `videox_fun.data` — never write a new one.
|
||||
|
||||
### Annotation format — the `train_data_meta` file (`metadata.json` / `.csv`)
|
||||
Every dataset class reads an annotation file (`args.train_data_meta`) that indexes the media under `args.train_data_dir` (`data_root`). `ImageVideoDataset` accepts **`.json`** (a top-level array of records) or **`.csv`** (`csv.DictReader`; the header row is the field names). Each record for ordinary image/video training:
|
||||
|
||||
| Field | Required | Meaning |
|
||||
|-------|----------|---------|
|
||||
| `file_path` | yes | Media path, resolved **relative to `train_data_dir`** via `os.path.join(data_root, file_path)`. If `data_root is None`, `file_path` is used as-is. |
|
||||
| `text` | yes | Caption / prompt. Dropped to `""` with probability `text_drop_ratio` (default `0.1`) for classifier-free guidance. |
|
||||
| `type` | no | `"video"` or `"image"`; **defaults to `"image"`** when the key is absent (`data_info.get('type', 'image')`). |
|
||||
|
||||
```json
|
||||
[
|
||||
{"file_path": "train/00000000.mp4", "text": "A young woman gently turns her head to the right ...", "type": "video"},
|
||||
{"file_path": "train/00000001.jpg", "text": "a dog running on the beach", "type": "image"}
|
||||
]
|
||||
```
|
||||
The directory layout matches the index — media in a `train/` subdir, the annotation file beside it. Ready-made examples ship in `datasets/X-Fun-Videos-Demo/` (`train/*.mp4` + `metadata.json`) and `datasets/X-Fun-Images-Demo/`. The equivalent `.csv`:
|
||||
```csv
|
||||
file_path,text,type
|
||||
train/00000000.mp4,"A young woman gently turns her head to the right ...",video
|
||||
train/00000001.jpg,"a dog running on the beach",image
|
||||
```
|
||||
|
||||
**Variant datasets append extra fields to this same record shape**, each consumed by its own class (see the table below) — e.g. camera-pose adds `action_path` (`LingbotImageVideoDataset`), object/VACE/S2V variants add object fields (`object_file_path` / `objects`). The demo folders also ship several augmented metadata variants (next subsection). Always read the target class's `get_batch` for the exact fields it consumes.
|
||||
|
||||
### Ready-made demo datasets — the standard test data (never invent a test set)
|
||||
Smoke tests, `log_validation` checks, and doc examples all run on the official demo datasets under `datasets/`, downloaded from ModelScope as `PAI/<name>`:
|
||||
```bash
|
||||
modelscope download --dataset PAI/X-Fun-Videos-Demo --local_dir ./datasets/X-Fun-Videos-Demo
|
||||
```
|
||||
Pick the demo by **task**, matching the dataset class in the table below:
|
||||
|
||||
| Demo dataset (`datasets/...`) | Contents | Extra metadata fields | Task it tests | Dataset class |
|
||||
|-------------------------------|----------|----------------------|---------------|---------------|
|
||||
| `X-Fun-Videos-Demo` | 16 videos (832×480) in `train/` | — | T2V / I2V base + inpaint, distill | `ImageVideoDataset` |
|
||||
| `X-Fun-Videos-Controls-Demo` | 16 videos in `train/` + `canny/` + `object/<video_id>/` + `wav/` | `control_file_path`, `object_file_path` (list), `audio_path` | V2V control, VACE, S2V-with-control | `ImageVideoControlDataset`, `VideoSpeechControlDataset` |
|
||||
| `X-Fun-Videos-Audios-Demo` | 17 video/audio pairs: `train/` (1280×720) + `wav/` (16 kHz mono) + `pose/` | `audio_path`, `control_file_path` | Speech-driven S2V / avatar / talking-head | `VideoSpeechDataset` |
|
||||
| `X-Fun-Images-Demo` | 19 images in `train/` | — | T2I full fine-tune + LoRA (z_image / flux2 / qwenimage / lens / ernie) | `ImageVideoDataset` |
|
||||
| `X-Fun-Images-Controls-Demo` | 19 images in `train/` + `canny/` | `control_file_path` | Image control / ControlNet / i2i inpaint | `ImageVideoControlDataset` |
|
||||
| `X-Fun-Images-Edit-Demo` | 21 records: `source/souce-<id>/` (multi-source supported) → `train/` | `source_file_path` (**list**) | Image edit (Qwen-Image-Edit family) | `ImageEditDataset` |
|
||||
| `X-Fun-Videos-Lingbot-Demo` | video + `intrinsics.npy` / `poses.npy` | camera pose / action | Camera-pose world model (`lingbot_world`) | `LingbotImageVideoDataset` |
|
||||
|
||||
**Which metadata file to point `--train_data_meta` at** (each demo ships several variants beside the media):
|
||||
|
||||
| Metadata file | Use when |
|
||||
|---------------|----------|
|
||||
| `metadata.json` | Base format only (`file_path` / `text` / `type`) — fine for a minimal check |
|
||||
| `metadata_add_width_height.json` | **Default choice.** Adds `width` / `height` so bucketing doesn't decode media (matters on slow storage such as OSS). Used by non-VACE control / S2V training too |
|
||||
| `metadata_add_width_height_add_objects.json` | VACE / subject-reference training (`object_file_path` list → `object/<video_id>/`; shuffled at train time) |
|
||||
| `metadata_add_width_height_add_wav.json` | Audio-visual joint models (e.g. `minimax_h3_fun` control training): `audio_path` → `wav/`. Keep the `.sh` launcher and the README on the same file |
|
||||
| `metadata_lingbot_video_add_width_height.json` | `lingbot_video` — `text` is already a structured JSON caption (lives in `X-Fun-Videos-Demo`) |
|
||||
| `metadata_origin.json` | Pre-processing original kept for reference; not used for training |
|
||||
|
||||
Regenerate the width/height variant with the shipped helper when adding your own media:
|
||||
`python scripts/process_json_add_width_and_height.py --input_file datasets/<Demo>/metadata.json --output_file datasets/<Demo>/metadata_add_width_height.json`.
|
||||
|
||||
`audio_path` optionality differs per class (`videox_fun/data/dataset_video.py`): `VideoSpeechDataset` reads `video_dict['audio_path']` directly, so it is **required**; `VideoSpeechControlDataset` uses `.get('audio_path')` and **falls back to the video file's own audio track** when the field is absent.
|
||||
|
||||
### Dataset by task (all take `train_data_meta, train_data_dir, ...`)
|
||||
| Task / mode | Dataset class | Used by | Key kwargs |
|
||||
|-------------|--------------|---------|-----------|
|
||||
| T2V / I2V base (`normal` + inpaint) | `ImageVideoDataset` | `train.py`, `train_lora.py`, t2i `train.py` | `enable_inpaint = train_mode != "normal"`, `video_sample_size/stride/n_frames`, `image_sample_size`, `video_repeat` |
|
||||
| Image T2I (qwenimage/flux/z_image) | `ImageVideoDataset` | `scripts/<img>/train.py` | `image_sample_size` |
|
||||
| Control (canny/pose/depth/camera) | `ImageVideoControlDataset` | `train_control*.py`, `train_control_distill.py` | `enable_camera_info = train_mode == "control_camera_ref"` |
|
||||
| Image Edit (source→target) | `ImageEditDataset` | `qwenimage/train_edit*.py` | `image_sample_size` |
|
||||
| Speech/audio-driven (S2V, avatar, talking) | `VideoSpeechDataset` | `mova`, `ltx2`, `minimax_h3`, `fantasytalking`, `infinitetalk`, `flashhead`, `longcatvideo/train_avatar*` | audio + video fields |
|
||||
| S2V **with control** | `VideoSpeechControlDataset` | `wan2.2/train_s2v*.py`, `minimax_h3_fun/train_control*` | audio + control |
|
||||
| Motion/pose animate | `VideoAnimateDataset` | `wan2.2/train_animate*.py` | motion/pose driven |
|
||||
| Distill text-only branch, GRPO, DPO | `TextDataset` | `train_distill*.py` (text branch), `z_image/train_grpo_lora.py`, `train_dpo_lora.py` | reads only the `text` field; `text_drop_ratio` |
|
||||
| Precomputed latents (ODE pairs) | `ImageVideoSafetensorsDataset` | `wan2.1_self_forcing/train_ode.py` | `data_root` |
|
||||
| Camera-pose conditioning | `LingbotImageVideoDataset` | `lingbot_world/train.py` | `intrinsics.npy` / `poses.npy` |
|
||||
| Video-only (VAE/TAEHV distill) | `VideoDataset` | `taehv/train_taehv.py` | `sample_size/stride/n_frames`, `enable_inpaint=False` |
|
||||
|
||||
### Sampler by condition
|
||||
| Condition | Sampler | Shape |
|
||||
|-----------|---------|-------|
|
||||
| `enable_bucket=True` (default; image+video) | `AspectRatioBatchImageVideoSampler` | `sampler=RandomSampler(ds, generator=g), dataset=train_dataset.dataset, batch_size, train_folder=args.train_data_dir, drop_last=True, aspect_ratios=aspect_ratio_sample_size` |
|
||||
| `enable_bucket=False` | `ImageVideoSampler` | `ImageVideoSampler(RandomSampler(ds, generator=g), train_dataset, batch_size)` |
|
||||
| `TextDataset` (distill text branch / GRPO / DPO) | `BatchSampler` (plain) | `BatchSampler(RandomSampler(ds, generator=g), batch_size, drop_last=True)`; GRPO adds `k_repeat=args.num_image_per_prompt` |
|
||||
| video-only bucket (available, not used by current scripts) | `AspectRatioBatchSampler` | — |
|
||||
| image-only bucket (available, not used by current scripts) | `AspectRatioBatchImageSampler` | — |
|
||||
|
||||
`aspect_ratio_sample_size` is built from `ASPECT_RATIO_512` scaled by `args.video_sample_size`; `get_closest_ratio` picks the bucket inside `collate_fn`.
|
||||
|
||||
### Universal DataLoader creation pattern
|
||||
```python
|
||||
batch_sampler_generator = torch.Generator().manual_seed(args.seed)
|
||||
if args.enable_bucket:
|
||||
aspect_ratio_sample_size = {k: [x / 512 * args.video_sample_size for x in ASPECT_RATIO_512[k]] for k in ASPECT_RATIO_512}
|
||||
batch_sampler = AspectRatioBatchImageVideoSampler(
|
||||
sampler=RandomSampler(train_dataset, generator=batch_sampler_generator), dataset=train_dataset.dataset,
|
||||
batch_size=args.train_batch_size, train_folder=args.train_data_dir, drop_last=True,
|
||||
aspect_ratios=aspect_ratio_sample_size)
|
||||
def collate_fn(examples):
|
||||
new_examples = {"pixel_values": [], "text": []}
|
||||
if args.train_mode != "normal": # inpaint/i2v adds mask fields
|
||||
new_examples.update({"mask_pixel_values": [], "mask": [], "clip_pixel_values": []})
|
||||
# bucket via get_closest_ratio -> transform (Resize/CenterCrop/Normalize) -> stack
|
||||
# masked branch uses get_random_mask(...)
|
||||
return new_examples
|
||||
train_dataloader = torch.utils.data.DataLoader(
|
||||
train_dataset, batch_sampler=batch_sampler, collate_fn=collate_fn,
|
||||
persistent_workers=args.dataloader_num_workers != 0, num_workers=args.dataloader_num_workers,
|
||||
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
|
||||
else:
|
||||
batch_sampler = ImageVideoSampler(RandomSampler(train_dataset, generator=batch_sampler_generator), train_dataset, args.train_batch_size)
|
||||
train_dataloader = torch.utils.data.DataLoader(
|
||||
train_dataset, batch_sampler=batch_sampler,
|
||||
persistent_workers=args.dataloader_num_workers != 0, num_workers=args.dataloader_num_workers,
|
||||
worker_init_fn=worker_init_fn(args.seed + accelerator.process_index))
|
||||
```
|
||||
`collate_fn` receives the `examples` **list** (not a `batch` dict); build every batch-level field (`text`, `pixel_values`, masks) explicitly from `examples` into `new_examples`. When `--enable_text_encoder_in_dataloader`, encode prompts inside `collate_fn` and emit `encoder_hidden_states` / `encoder_attention_mask`.
|
||||
|
||||
## 9. Inference task matrix — predict script → pipeline → inputs
|
||||
|
||||
Pick the pipeline by **task**; the `predict_<task>.py` name and its inputs follow the same convention across families.
|
||||
|
||||
| Task | `predict_<task>.py` | Pipeline (family example) | Extra `__call__` inputs | Input helper |
|
||||
|------|--------------------|---------------------------|-------------------------|--------------|
|
||||
| Text→Video | `predict_t2v.py` | `WanPipeline`, `Wan2_2Pipeline`, `CogVideoXFunPipeline`, `LongCatVideoPipeline`, `LTX2Pipeline` | `prompt` only | — |
|
||||
| Image→Video | `predict_i2v.py` | `WanI2VPipeline`(=`WanFunInpaintPipeline`), `Wan2_2FunInpaintPipeline`, `Wan2_2I2VPipeline`, `HunyuanVideoI2VPipeline` | `video`, `mask_video` | `get_image_to_video_latent(start_image, end_image, video_length, sample_size)` |
|
||||
| Text+Image→Video (5B) | `predict_ti2v.py` | `Wan2_2TI2VPipeline` | `prompt` (+ optional image) | `get_image_to_video_latent` |
|
||||
| Video→Video Control | `predict_v2v_control.py` | `WanFunControlPipeline`, `Wan2_2FunControlPipeline` | `control_video` | `get_video_to_video_latent(control_video, ...)` |
|
||||
| Control + reference | `predict_v2v_control_ref.py` | `WanFunControlPipeline` | `control_video` + `ref_image` | `get_video_to_video_latent` + `get_image_latent` |
|
||||
| Control + camera | `predict_v2v_control_camera.py` | `WanFunControlPipeline` | `control_video` + camera pose | — |
|
||||
| VACE (control/mask/i2v/s2v) | `predict_v2v_control.py`, `predict_v2v_mask.py`, `predict_s2v.py`, `predict_i2v.py` | `WanVacePipeline`, `Wan2_2VaceFunPipeline` | control/mask/ref | — |
|
||||
| Speech→Video (audio) | `predict_s2v.py` | `Wan2_2S2VPipeline`, `MiniMaxH3Pipeline`, `InfiniteTalkPipeline`, `FantasyTalkingPipeline`, `FlashHeadPipeline`, `MOVAPipeline`, `LongCatVideoAvatarPipeline` | `audio` + reference image | — |
|
||||
| Animate (motion/pose) | `predict_animate.py` | `Wan2_2AnimatePipeline` | motion/pose video + ref | — |
|
||||
| Subject reference | `predict_s2v.py` (phantom) | `WanFunPhantomPipeline` | reference images | — |
|
||||
| Text→Image | `predict_t2i.py` | `QwenImagePipeline`, `Flux2Pipeline`, `ZImagePipeline`, `LensPipeline`, `ErnieImagePipeline` | `prompt` | — |
|
||||
| Image Control (t2i) | `predict_t2i_control.py` | `QwenImageControlPipeline`, `ZImageControlPipeline`, `Flux2ControlPipeline`, `QwenImageControlNetPipeline` | `control_image` | — |
|
||||
| Inpaint (i2i) | `predict_i2i_inpaint.py` | `QwenImageControlPipeline`, `ZImageControlPipeline`, `Flux2ControlPipeline` | `image` + `mask` | — |
|
||||
| Image Edit | `predict_t2i_edit.py`, `predict_t2i_edit_plus.py` | `QwenImageEditPipeline`, `QwenImageEditPlusPipeline` | source image + instruction | — |
|
||||
| Layered edit | `predict_i2i_layered.py` | `QwenImageLayeredPipeline` | image | — |
|
||||
| Camera-pose world | `predict_i2v.py` (lingbot_world) | `Wan2_2I2VPipeline`, `WanFunLingbotWorldFastPipeline` | image + camera pose | — |
|
||||
| Latent upsample | `predict_i2v_upsample.py` | `LTX2LatentUpsamplePipeline`, `WanLatentUpsamplePipeline` | low-res latent/video | — |
|
||||
| AR / streaming distill | `predict_t2v_stream.py` | `WanSelfForcingPipeline` | prompt (streamed) | — |
|
||||
|
||||
### Predict-script variant suffixes (same task, different backend/model)
|
||||
| Suffix | Meaning |
|
||||
|--------|---------|
|
||||
| `_tae` | Fast decode via `AutoencoderTinyWan` (TAEHV) instead of the full VAE |
|
||||
| `_2.2vae` | Uses the Wan2.2 VAE (`AutoencoderKLWan3_8`) |
|
||||
| `_5b` | 5B-parameter model variant |
|
||||
| `turbo` / distill | Distilled model, few-step inference (e.g. `predict_turbo_*.py`) |
|
||||
| `_refine` | Two-stage refine pass |
|
||||
| `_ref` / `_camera` | Adds reference-image / camera conditioning |
|
||||
|
||||
All variants keep the identical config block, `GPU_memory_mode` branching, and `save_results()` from Section 4 — only the loaded VAE/transformer and pipeline class change.
|
||||
|
||||
## 10. Preprocessing — offline training-data generation (multi-GPU + safetensors)
|
||||
|
||||
Here "preprocessing" means **generating/caching training data offline** with the teacher / VAE / text-encoder — latents, ODE-trajectory pairs, prompt/text embeddings — so training just reads cached tensors instead of re-encoding every step. Canonical example: `scripts/wan2.1_self_forcing/generate_ode_pairs.py` (+ `generate_ode_pairs.sh`); the loader-side contract is `ImageVideoSafetensorsDataset` in `videox_fun/data/dataset_image_video.py`. Two rules are non-negotiable.
|
||||
|
||||
### Rule 1 — multi-GPU is mandatory
|
||||
Never a single-GPU / hardcoded `cuda:0` loop. Launch with `accelerate launch` and shard work across ranks by interleaving:
|
||||
```python
|
||||
from accelerate import Accelerator
|
||||
accelerator = Accelerator(mixed_precision=args.mixed_precision)
|
||||
device, world_size, rank = accelerator.device, accelerator.num_processes, accelerator.process_index
|
||||
torch.set_grad_enabled(False) # inference-only
|
||||
|
||||
total_per_rank = math.ceil(len(prompts) / world_size)
|
||||
for index in tqdm(range(total_per_rank), disable=rank != 0):
|
||||
prompt_index = index * world_size + rank # interleaved shard
|
||||
if prompt_index >= len(prompts):
|
||||
continue
|
||||
out_path = os.path.join(args.output_folder, f"{prompt_index:05d}.safetensors")
|
||||
if os.path.exists(out_path): # resume-friendly
|
||||
continue
|
||||
... # encode prompt / run teacher ODE / collect latents
|
||||
accelerator.wait_for_everyone()
|
||||
if accelerator.is_main_process: # write the JSON index once, on rank 0
|
||||
json.dump([{"file_path": p} for p in all_safetensor_paths],
|
||||
open(os.path.join(args.output_folder, "outputs.json"), "w"), ensure_ascii=False, indent=4)
|
||||
```
|
||||
Launcher (`.sh`): `accelerate launch --mixed_precision="bf16" scripts/<family>/generate_<...>.py --pretrained_model_name_or_path=... --config_path=config/<family>/*.yaml --output_folder=datasets/<...> ...`. Reuse `videox_fun.models` + config-driven `from_pretrained` (Section 3) and `videox_fun.utils.utils.save_videos_grid` for sample previews — do not write a new loader.
|
||||
|
||||
### Rule 2 — store as safetensors; do NOT use LMDB or `.pt`
|
||||
Save every cached tensor with `safetensors.torch.save_file`, one `.safetensors` per sample (or per tensor), plus a JSON index of `{"file_path": ...}` entries:
|
||||
```python
|
||||
from safetensors.torch import save_file
|
||||
save_file(
|
||||
{"latents": latents.cpu(), "prompt_embeds": text_embeds.cpu(), "prompt_attention_mask": mask.cpu()},
|
||||
out_path, # f"{prompt_index:05d}.safetensors"
|
||||
metadata={"prompt": prompt},
|
||||
)
|
||||
```
|
||||
`ImageVideoSafetensorsDataset(ann_path, data_root=None)` reads that JSON and supports two layouts:
|
||||
- **Single-file (default)**: `{"file_path": "scene.safetensors"}` — whole state dict in one archive.
|
||||
- **Per-tensor (`--save_per_tensor`)**: `{"file_path": "scene_dir", "latents": ".../latents.safetensors", "prompt_embeds": ".../prompt_embeds.safetensors"}` — each key loaded and merged.
|
||||
|
||||
**Do not** cache preprocessed data in **LMDB** or as **`.pt`/`.pth` `torch.save` pickles**. safetensors is the repo-wide standard (also used for LoRA/weight saving), is pickle-free/safe, memory-maps fast, and is exactly what `ImageVideoSafetensorsDataset` loads. (Scope: this governs cached *data tensors*; accelerate optimizer/scheduler/scaler `.pt` states written during training checkpoints are a separate mechanism and unaffected.)
|
||||
|
||||
### Related but different — dataset curation
|
||||
Scoring / filtering / captioning under `videox_fun/video_caption/` (`compute_*.py`, `internvl2_video_recaptioning.py`) is dataset *curation*, not latent caching. It is also multi-GPU (accelerate `PartialState.split_between_processes`/`gather_object`, or vLLM `tensor_parallel_size=device_count()`), but writes csv/jsonl **metadata** (not tensors), so Rule 2 does not apply there.
|
||||
@@ -1,52 +0,0 @@
|
||||
FROM nvidia/cuda:12.1.0-cudnn8-devel-ubuntu22.04
|
||||
ENV DEBIAN_FRONTEND noninteractive
|
||||
|
||||
RUN rm -r /etc/apt/sources.list.d/
|
||||
|
||||
RUN apt-get update -y && apt-get install -y \
|
||||
libgl1 libglib2.0-0 google-perftools \
|
||||
sudo wget git git-lfs vim tig pkg-config libcairo2-dev \
|
||||
aria2 telnet curl net-tools iputils-ping jq \
|
||||
python3-pip python-is-python3 python3.10-venv tzdata lsof zip tmux
|
||||
RUN apt-get update && \
|
||||
apt-get install -y software-properties-common && \
|
||||
add-apt-repository ppa:ubuntuhandbook1/ffmpeg6 && \
|
||||
apt-get update && \
|
||||
apt-get install -y ffmpeg
|
||||
|
||||
RUN pip3 install --upgrade pip -i https://mirrors.aliyun.com/pypi/simple/
|
||||
|
||||
# add all extensions
|
||||
RUN pip install wandb tqdm GitPython==3.1.32 Pillow==9.5.0 setuptools --upgrade -i https://mirrors.aliyun.com/pypi/simple/
|
||||
|
||||
RUN pip install torch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 --index-url https://download.pytorch.org/whl/cu118
|
||||
RUN pip install xformers==0.0.27.post2 --index-url https://download.pytorch.org/whl/cu118
|
||||
|
||||
# install vllm (video-caption)
|
||||
RUN pip install vllm==0.6.3
|
||||
|
||||
# install requirements (video-caption)
|
||||
WORKDIR /root/
|
||||
COPY easyanimate/video_caption/requirements.txt /root/requirements-video_caption.txt
|
||||
RUN pip install -r /root/requirements-video_caption.txt
|
||||
RUN rm /root/requirements-video_caption.txt
|
||||
|
||||
RUN pip install -U http://eas-data.oss-cn-shanghai.aliyuncs.com/sdk/allspark-0.15-py2.py3-none-any.whl
|
||||
RUN pip install -e git+https://github.com/CompVis/taming-transformers.git@master#egg=taming-transformers
|
||||
RUN pip install came-pytorch deepspeed pytorch_lightning==1.9.4 func_timeout -i https://mirrors.aliyun.com/pypi/simple/
|
||||
|
||||
# install requirements
|
||||
RUN pip install bitsandbytes mamba-ssm causal-conv1d>=1.4.0 -i https://mirrors.aliyun.com/pypi/simple/
|
||||
RUN pip install ipykernel -i https://mirrors.aliyun.com/pypi/simple/
|
||||
COPY ./requirements.txt /root/requirements.txt
|
||||
RUN pip install -r /root/requirements.txt -i https://mirrors.aliyun.com/pypi/simple/
|
||||
RUN rm -rf /root/requirements.txt
|
||||
|
||||
# install package patches (video-caption)
|
||||
COPY easyanimate/video_caption/package_patches/easyocr_detection_patched.py /usr/local/lib/python3.10/dist-packages/easyocr/detection.py
|
||||
COPY easyanimate/video_caption/package_patches/vila_siglip_encoder_patched.py /usr/local/lib/python3.10/dist-packages/llava/model/multimodal_encoder/siglip_encoder.py
|
||||
|
||||
ENV PYTHONUNBUFFERED 1
|
||||
ENV NVIDIA_DISABLE_REQUIRE 1
|
||||
|
||||
WORKDIR /root/
|
||||
@@ -1,66 +1,78 @@
|
||||
# VideoX-Fun
|
||||
# CogVideoX-Fun
|
||||
|
||||
😊 Welcome!
|
||||
|
||||
CogVideoX-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/CogVideoX-Fun-5b)
|
||||
|
||||
Wan-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/Wan2.1-Fun-1.3B-InP)
|
||||
|
||||
English | [简体中文](./README_zh-CN.md) | [日本語](./README_ja-JP.md)
|
||||
English | [简体中文](./README_zh-CN.md)
|
||||
|
||||
# Table of Contents
|
||||
- [I. Introduction](#i-introduction)
|
||||
- [II. Quick Start and Usage](#ii-quick-start-and-usage)
|
||||
- [1. Environment Preparation](#1-environment-preparation)
|
||||
- [2. Inference Generation](#2-inference-generation)
|
||||
- [3. Model Training](#3-model-training)
|
||||
- [III. Supported Models](#iii-supported-models)
|
||||
- [IV. Video Works](#iv-video-works)
|
||||
- [V. References](#v-references)
|
||||
- [VI. Citation](#vi-citation)
|
||||
- [VII. Limitations and Risks](#vii-limitations-and-risks)
|
||||
- [VIII. License](#viii-license)
|
||||
- [Table of Contents](#table-of-contents)
|
||||
- [Introduction](#introduction)
|
||||
- [Quick Start](#quick-start)
|
||||
- [How to use](#how-to-use)
|
||||
- [Model zoo](#model-zoo)
|
||||
- [TODO List](#todo-list)
|
||||
- [Reference](#reference)
|
||||
- [License](#license)
|
||||
|
||||
# I. Introduction
|
||||
VideoX-Fun is a video generation pipeline that can be used to generate AI images and videos, as well as to train baseline and Lora models for Diffusion Transformer. We support direct prediction from pre-trained baseline models to generate videos with different resolutions, durations, and FPS. Additionally, we also support users in training their own baseline and Lora models to perform specific style transformations.
|
||||
# Introduction
|
||||
CogVideoX-Fun is a modified pipeline based on the CogVideoX structure, designed to provide more flexibility in generation. It can be used to create AI images and videos, as well as to train baseline models and Lora models for Diffusion Transformer. We support predictions directly from the already trained CogVideoX-Fun model, allowing the generation of videos at different resolutions, approximately 6 seconds long with 8 fps (1 to 49 frames). Users can also train their own baseline models and Lora models to achieve certain style transformations.
|
||||
|
||||
We will support quick pull-ups from different platforms, refer to [Quick Start](#quick-start).
|
||||
|
||||
What's New:
|
||||
- Added support for Wan 2.2 series models, Wan-VACE control model, Fantasy Talking digital human model, Qwen-Image, Flux image generation models, and more. [2025.10.16]
|
||||
- Update Wan2.1-Fun-V1.1: Support for 14B and 1.3B model Control + Reference Image models, support for camera control, and the Inpaint model has been retrained for improved performance. [2025.04.25]
|
||||
- Update Wan2.1-Fun-V1.0: Support I2V and Control models for 14B and 1.3B models, with support for start and end frame prediction. [2025.03.26]
|
||||
- Update CogVideoX-Fun-V1.5: Upload I2V model and related training/prediction code. [2024.12.16]
|
||||
- Reward Lora Support: Train Lora using reward backpropagation techniques to optimize generated videos, making them better aligned with human preferences. [More Information](scripts/README_TRAIN_REWARD.md). New version of the control model supports various control conditions such as Canny, Depth, Pose, MLSD, etc. [2024.11.21]
|
||||
- Diffusers Support: CogVideoX-Fun Control is now supported in diffusers. Thanks to [a-r-r-o-w](https://github.com/a-r-r-o-w) for contributing support in this [PR](https://github.com/huggingface/diffusers/pull/9671). Check out the [documentation](https://huggingface.co/docs/diffusers/main/en/api/pipelines/cogvideox) for more details. [2024.10.16]
|
||||
- Update CogVideoX-Fun-V1.1: Retrain i2v model, add Noise to increase the motion amplitude of the video. Upload control model training code and Control model. [2024.09.29]
|
||||
- Update CogVideoX-Fun-V1.0: Initial code release! Now supports Windows and Linux. Supports video generation at arbitrary resolutions from 256x256x49 to 1024x1024x49 for 2B and 5B models. [2024.09.18]
|
||||
- Create code! Now supporting Windows and Linux. Supports video generation at any resolution from 256x256x49 to 1024x1024x49. [ 2024.09.09 ]
|
||||
|
||||
Function:
|
||||
- [Data Preprocessing](#data-preprocess)
|
||||
- [Train DiT](#dit-train)
|
||||
- [Video Generation](#video-gen)
|
||||
|
||||
These are our generated results [GALLERY](scripts/Result_Gallery.md) (Click the image below to see the video):
|
||||
|
||||
Our UI interface is as follows:
|
||||

|
||||
|
||||
# II. Quick Start and Usage
|
||||
# Quick Start
|
||||
### 1. Cloud usage: AliyunDSW/Docker
|
||||
#### a. From AliyunDSW
|
||||
On the way.
|
||||
|
||||
<a id="quick-start"></a>
|
||||
#### b. From ComfyUI
|
||||
Our ComfyUI is as follows, please refer to [ComfyUI README](comfyui/README.md) for details.
|
||||

|
||||
|
||||
## 1. Environment Preparation
|
||||
#### c. From docker
|
||||
If you are using docker, please make sure that the graphics card driver and CUDA environment have been installed correctly in your machine.
|
||||
|
||||
### 1.1 Cloud Usage: AliyunDSW
|
||||
Then execute the following commands in this way:
|
||||
|
||||
DSW has free GPU time, which can be applied once by a user and is valid for 3 months after applying.
|
||||
```
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
Aliyun provide free GPU time in [Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1), get it and use in Aliyun PAI-DSW to start CogVideoX-Fun within 5min!
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
[](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/CogVideoX-Fun.git
|
||||
|
||||
### 1.2 Local Dependency Installation
|
||||
# enter CogVideoX-Fun's dir
|
||||
cd CogVideoX-Fun
|
||||
|
||||
We have verified this repo execution on the following environment:
|
||||
# download weights
|
||||
mkdir models/Diffusion_Transformer
|
||||
mkdir models/Personalized_Model
|
||||
|
||||
wget https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz -O models/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz
|
||||
|
||||
cd models/Diffusion_Transformer/
|
||||
tar -xvf CogVideoX-Fun-2b-InP.tar.gz
|
||||
cd ../../
|
||||
```
|
||||
|
||||
### 2. Local install: Environment Check/Downloading/Installation
|
||||
#### a. Environment Check
|
||||
We have verified CogVideoX-Fun execution on the following environment:
|
||||
|
||||
The detailed of Windows:
|
||||
- OS: Windows 10
|
||||
@@ -80,165 +92,43 @@ The detailed of Linux:
|
||||
|
||||
We need about 60GB available on disk (for saving weights), please check!
|
||||
|
||||
### 1.3 Using Docker
|
||||
#### b. Weights
|
||||
We'd better place the [weights](#model-zoo) along the specified path:
|
||||
|
||||
If you are using docker, please make sure that the graphics card driver and CUDA environment have been installed correctly in your machine.
|
||||
|
||||
Then execute the following commands in this way:
|
||||
|
||||
```
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/VideoX-Fun.git
|
||||
|
||||
# enter VideoX-Fun's dir
|
||||
cd VideoX-Fun
|
||||
|
||||
# download weights
|
||||
mkdir models/Diffusion_Transformer
|
||||
mkdir models/Personalized_Model
|
||||
|
||||
# Please use the hugginface link or modelscope link to download the model.
|
||||
# CogVideoX-Fun
|
||||
# https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
|
||||
# https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP
|
||||
|
||||
# Wan
|
||||
# https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP
|
||||
# https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP
|
||||
```
|
||||
|
||||
### 1.4 Weight Placement
|
||||
|
||||
We'd better place the [weights](#iii-supported-models) along the specified path:
|
||||
|
||||
**Via ComfyUI**:
|
||||
Put the models into the ComfyUI weights folder `ComfyUI/models/Fun_Models/`:
|
||||
```
|
||||
📦 ComfyUI/
|
||||
├── 📂 models/
|
||||
│ └── 📂 Fun_Models/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-14B-InP
|
||||
│ └── 📂 Wan2.1-Fun-1.3B-InP/
|
||||
```
|
||||
|
||||
**Run its own python file or UI interface**:
|
||||
```
|
||||
📦 models/
|
||||
├── 📂 Diffusion_Transformer/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-14B-InP
|
||||
│ └── 📂 Wan2.1-Fun-1.3B-InP/
|
||||
│ └── 📂 CogVideoX-Fun-2b-InP/
|
||||
├── 📂 Personalized_Model/
|
||||
│ └── your trained trainformer model / your trained lora model (for UI load)
|
||||
```
|
||||
|
||||
## 2. Inference Generation
|
||||
# How to use
|
||||
|
||||
<a id="video-gen"></a>
|
||||
<h3 id="video-gen">1. Inference </h3>
|
||||
|
||||
Video and image models share the exact same inference entry, provided by scripts or UI under `examples/{model_name}/`.
|
||||
#### a. Using Python Code
|
||||
- Step 1: Download the corresponding [weights](#model-zoo) and place them in the models folder.
|
||||
- Step 2: Modify prompt, neg_prompt, guidance_scale, and seed in the predict_t2v.py file.
|
||||
- Step 3: Run the predict_t2v.py file, wait for the generated results, and save the results in the samples/cogvideox-fun-videos-t2v folder.
|
||||
- Step 4: If you want to combine other backbones you have trained with Lora, modify the predict_t2v.py and Lora_path in predict_t2v.py depending on the situation.
|
||||
|
||||
### 2.1 Entry Selection
|
||||
#### b. Using webui
|
||||
- Step 1: Download the corresponding [weights](#model-zoo) and place them in the models folder.
|
||||
- Step 2: Run the app.py file to enter the graph page.
|
||||
- Step 3: Select the generated model based on the page, fill in prompt, neg_prompt, guidance_scale, and seed, click on generate, wait for the generated result, and save the result in the samples folder.
|
||||
|
||||
| Entry | Suitable Scenario | Config Granularity |
|
||||
|--|--|--|
|
||||
| Python file | Batch generation, parameter debugging | Full parameters |
|
||||
| WebUI | Interactive experience | Common parameters only |
|
||||
| ComfyUI | Existing ComfyUI workflow | Node parameters |
|
||||
#### c. From ComfyUI
|
||||
Please refer to [ComfyUI README](comfyui/README.md) for details.
|
||||
|
||||
Table: inference entry selection
|
||||
### 2. Model Training
|
||||
A complete CogVideoX-Fun training pipeline should include data preprocessing, and Video DiT training.
|
||||
|
||||
### 2.2 GPU Memory Saving Options
|
||||
<h4 id="data-preprocess">a. data preprocessing</h4>
|
||||
|
||||
Since Wan2.1 has a very large number of parameters, we need to consider memory optimization strategies to adapt to consumer-grade GPUs. We provide `GPU_memory_mode` for each prediction file, allowing you to choose between `model_cpu_offload`, `model_cpu_offload_and_qfloat8`, and `sequential_cpu_offload`. This solution is also applicable to CogVideoX-Fun generation.
|
||||
We have provided a simple demo of training the Lora model through image data, which can be found in the [wiki](https://github.com/aigc-apps/CogVideoX-Fun/wiki/Training-Lora) for details.
|
||||
|
||||
- `model_cpu_offload`: The entire model is moved to the CPU after use, saving some GPU memory.
|
||||
- `model_cpu_offload_and_qfloat8`: The entire model is moved to the CPU after use, and the transformer model is quantized to float8, saving more GPU memory.
|
||||
- `sequential_cpu_offload`: Each layer of the model is moved to the CPU after use. It is slower but saves a significant amount of GPU memory.
|
||||
|
||||
`qfloat8` may slightly reduce model performance but saves more GPU memory. If you have sufficient GPU memory, it is recommended to use `model_cpu_offload`.
|
||||
|
||||
### 2.3 Via Python Files
|
||||
|
||||
##### i. Single-GPU Inference:
|
||||
|
||||
- **Step 1**: Download the corresponding [weights](#iii-supported-models) and place them in the `models` folder.
|
||||
- **Step 2**: Use different files for prediction based on the weights and prediction goals. This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:
|
||||
- **Text-to-Video**:
|
||||
- Modify `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_t2v.py`.
|
||||
- Run the file `examples/cogvideox_fun/predict_t2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos`.
|
||||
- **Image-to-Video**:
|
||||
- Modify `validation_image_start`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_i2v.py`.
|
||||
- `validation_image_start` is the starting image of the video, and `validation_image_end` is the ending image of the video.
|
||||
- Run the file `examples/cogvideox_fun/predict_i2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_i2v`.
|
||||
- **Video-to-Video**:
|
||||
- Modify `validation_video`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_v2v.py`.
|
||||
- `validation_video` is the reference video for video-to-video generation. You can use the following demo video: [Demo Video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/play_guitar.mp4).
|
||||
- Run the file `examples/cogvideox_fun/predict_v2v.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_v2v`.
|
||||
- **Controlled Video Generation (Canny, Pose, Depth, etc.)**:
|
||||
- Modify `control_video`, `validation_image_end`, `prompt`, `neg_prompt`, `guidance_scale`, and `seed` in the file `examples/cogvideox_fun/predict_v2v_control.py`.
|
||||
- `control_video` is the control video extracted using operators such as Canny, Pose, or Depth. You can use the following demo video: [Demo Video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1.1/pose.mp4).
|
||||
- Run the file `examples/cogvideox_fun/predict_v2v_control.py` and wait for the results. The generated videos will be saved in the folder `samples/cogvideox-fun-videos_v2v_control`.
|
||||
- **Step 3**: If you want to integrate other backbones or Loras trained by yourself, modify `lora_path` and relevant paths in `examples/{model_name}/predict_t2v.py` or `examples/{model_name}/predict_i2v.py` as needed.
|
||||
|
||||
##### ii. Multi-GPU Inference:
|
||||
When using multi-GPU inference, please make sure to install the xfuser. We recommend installing xfuser==0.4.2 and yunchang==0.6.2.
|
||||
```
|
||||
pip install xfuser==0.4.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
|
||||
pip install yunchang==0.6.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
|
||||
```
|
||||
|
||||
Please ensure that the product of `ulysses_degree` and `ring_degree` equals the number of GPUs being used. For example, if you are using 8 GPUs, you can set `ulysses_degree=2` and `ring_degree=4`, or alternatively `ulysses_degree=4` and `ring_degree=2`.
|
||||
|
||||
- `ulysses_degree` performs parallelization after splitting across the heads.
|
||||
- `ring_degree` performs parallelization after splitting across the sequence.
|
||||
|
||||
Compared to `ulysses_degree`, `ring_degree` incurs higher communication costs. Therefore, when setting these parameters, you should take into account both the sequence length and the number of heads in the model.
|
||||
|
||||
Let’s take 8-GPU parallel inference as an example:
|
||||
|
||||
- **For Wan2.1-Fun-V1.1-14B-InP**, which has 40 heads, `ulysses_degree` should be set to a divisor of 40 (e.g., 2, 4, 8, etc.). Thus, when using 8 GPUs for parallel inference, you can set `ulysses_degree=8` and `ring_degree=1`.
|
||||
|
||||
- **For Wan2.1-Fun-V1.1-1.3B-InP**, which has 12 heads, `ulysses_degree` should be set to a divisor of 12 (e.g., 2, 4, etc.). Thus, when using 8 GPUs for parallel inference, you can set `ulysses_degree=4` and `ring_degree=2`.
|
||||
|
||||
After setting the parameters, run the following command for parallel inference:
|
||||
|
||||
```sh
|
||||
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
|
||||
```
|
||||
|
||||
### 2.4 Via the Web UI
|
||||
|
||||
The web UI supports text-to-video, image-to-video, video-to-video, and controlled video generation (Canny, Pose, Depth, etc.). This library currently supports CogVideoX-Fun, Wan2.1, and Wan2.1-Fun. Different models are distinguished by folder names under the `examples` folder, and their supported features vary. Use them accordingly. Below is an example using CogVideoX-Fun:
|
||||
|
||||
- **Step 1**: Download the corresponding [weights](#iii-supported-models) and place them in the `models` folder.
|
||||
- **Step 2**: Run the file `examples/cogvideox_fun/app.py` to access the Gradio interface.
|
||||
- **Step 3**: Select the generation model on the page, fill in `prompt`, `neg_prompt`, `guidance_scale`, and `seed`, click "Generate," and wait for the results. The generated videos will be saved in the `sample` folder.
|
||||
|
||||
### 2.5 Via ComfyUI
|
||||
|
||||
For details, refer to [ComfyUI README](comfyui/README.md).
|
||||
|
||||
|
||||
## 3. Model Training
|
||||
|
||||
A complete model training pipeline consists of data preprocessing and Video DiT training.
|
||||
|
||||
### 3.1 Data Preprocessing
|
||||
|
||||
<a id="data-preprocess"></a>
|
||||
Training documents for each model are unified under `scripts/{model_name}/`. For details, see [3.3 Training Documents per Model](#33-training-documents-per-model).
|
||||
|
||||
A complete data preprocessing link for long video segmentation, cleaning, and description can refer to [README](videox_fun/video_caption/README.md) in the video captions section.
|
||||
A complete data preprocessing link for long video segmentation, cleaning, and description can refer to [README](cogvideox/video_caption/README.md) in the video captions section.
|
||||
|
||||
If you want to train a text to image and video generation model. You need to arrange the dataset in this format.
|
||||
|
||||
@@ -287,393 +177,42 @@ You can also set the path as absolute path as follow:
|
||||
]
|
||||
```
|
||||
|
||||
### 3.2 Video DiT Training
|
||||
|
||||
<a id="dit-train"></a>
|
||||
The training scripts and launch sh files for each model are located under `scripts/{model_name}/`. The sh file names vary by task, such as `train.sh`, `train_lora.sh`, `train_control.sh`, `train_control_distill.sh`, etc.; refer to the actual files in the directory.
|
||||
|
||||
If the data format is relative path during data preprocessing, please set ```scripts/{model_name}/train.sh``` as follow.
|
||||
<h4 id="dit-train">b. Video DiT training </h4>
|
||||
|
||||
If the data format is relative path during data preprocessing, please set ```scripts/train.sh``` as follow.
|
||||
```
|
||||
export DATASET_NAME="datasets/internal_datasets/"
|
||||
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
|
||||
```
|
||||
|
||||
If the data format is absolute path during data preprocessing, please set ```scripts/{model_name}/train.sh``` as follow (`DATASET_NAME` is left empty so the dataset directory prefix is no longer concatenated).
|
||||
If the data format is absolute path during data preprocessing, please set ```scripts/train.sh``` as follow.
|
||||
```
|
||||
export DATASET_NAME=""
|
||||
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
|
||||
```
|
||||
|
||||
Finally, run the corresponding script.
|
||||
Then, we run scripts/train.sh.
|
||||
```sh
|
||||
sh scripts/{model_name}/train.sh
|
||||
sh scripts/train.sh
|
||||
```
|
||||
|
||||
### 3.3 Training Documents per Model
|
||||
|
||||
For parameter details, training documents for each model are unified under `scripts/{model_name}/`.
|
||||
|
||||
| Model | Baseline Training | LoRA Training | Others |
|
||||
|--|--|--|--|
|
||||
| Wan2.1-Fun | [EN](scripts/wan2.1_fun/README_TRAIN.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) | [Control EN](scripts/wan2.1_fun/README_TRAIN_CONTROL.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
|
||||
| Wan2.2 | [EN](scripts/wan2.2/README_TRAIN.md) / [ZH](scripts/wan2.2/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) | [Distill EN](scripts/wan2.2/README_TRAIN_DISTILL.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
|
||||
| Wan2.2-Fun | [EN](scripts/wan2.2_fun/README_TRAIN.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) | [Control LoRA EN](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA.md) |
|
||||
| CogVideoX-Fun | [EN](scripts/cogvideox_fun/README_TRAIN.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) | [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) | [Control EN](scripts/cogvideox_fun/README_TRAIN_CONTROL.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
|
||||
| Qwen-Image | [EN](scripts/qwenimage/README_TRAIN.md) / [ZH](scripts/qwenimage/README_TRAIN_zh-CN.md) | [EN](scripts/qwenimage/README_TRAIN_LORA.md) / [ZH](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) | [Edit EN](scripts/qwenimage/README_TRAIN_EDIT.md) |
|
||||
| Qwen-Image-2.1 | [EN](scripts/qwenimage21/README_TRAIN.md) / [ZH](scripts/qwenimage21/README_TRAIN_zh-CN.md) | - | - |
|
||||
| Z-Image | [EN](scripts/z_image/README_TRAIN.md) / [ZH](scripts/z_image/README_TRAIN_zh-CN.md) | [EN](scripts/z_image/README_TRAIN_LORA.md) / [ZH](scripts/z_image/README_TRAIN_LORA_zh-CN.md) | [GRPO LoRA EN](scripts/z_image/README_TRAIN_GRPO_LORA.md) |
|
||||
|
||||
For other models, check the READMEs under `scripts/{model_name}/`.
|
||||
|
||||
# III. Supported Models
|
||||
|
||||
The table below summarizes currently supported model families and weights. Video and image models share the same inference and training entry. Each row represents one model family; the fourth column is an embedded four-column HTML table (Weight, Hugging Face, ModelScope, Description). 🤗 is Hugging Face, 🤖 is ModelScope (recommended for users in mainland China), and `-` means the corresponding channel has no public repo or requires authentication. For training docs of each model, see [3.3 Training Documents per Model](#33-training-documents-per-model).
|
||||
|
||||
| Model Family | Modality | Supported Tasks | Weight / Download / Description |
|
||||
|--|--|--|--|
|
||||
| Wan2.2-Fun | Video | Series trained by this project on Wan2.2, covering T2V, I2V, first/last frame, controlled generation, and camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B text-to-video weights trained at 121 frames, 24 FPS, supporting first/last frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B video control weights, supporting control conditions like Canny, Depth, Pose, MLSD, and trajectory control. Trained at 121 frames, 24 FPS, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B camera lens control weights. Trained at 121 frames, 24 FPS, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Reward LoRAs that optimize Wan2.2-Fun generated videos via reward backpropagation</td></tr></table> |
|
||||
| Wan2.2-VACE-Fun | Video | Series trained by this project with the VACE scheme, covering controlled generation and subject reference | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Control weights for Wan2.2 trained using the VACE scheme (based on the base model Wan2.2-T2V-A14B), supporting various control conditions such as Canny, Depth, Pose, MLSD, trajectory control, etc. It supports video generation by specifying the subject. It supports multi-resolution (512, 768, 1024) video prediction, and is trained with 81 frames at 16 FPS. It also supports multi-language prediction.</td></tr></table> |
|
||||
| Wan2.2 | Video | Official Wan weights covering T2V, I2V, audio-driven, and character animation; can be used as training baseline for Wan2.2-Fun | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-5B text/image-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B text-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B image-to-video weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B audio-to-video weights, speaker-driven digital human</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B character replacement and motion transfer weights; repo contains multiple precision files</td></tr></table> |
|
||||
| Wan2.1-Fun V1.1 | Video | V1.1 series trained by this project on Wan2.1, multi-resolution (512/768/1024), 81 frames at 16fps, covering T2V, I2V, first/last frame, controlled generation, and camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B text-to-video generation weights, trained at multiple resolutions, supports start-end image prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B video control weights support various control conditions such as Canny, Depth, Pose, MLSD, etc., supports reference image + control condition-based control, and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14B camera lens control weights. Supports multi-resolution (512, 768, 1024) video prediction, trained with 81 frames at 16 FPS, supports multilingual prediction.</td></tr></table> |
|
||||
| Wan2.1-Fun V1.0 | Video | V1.0 series trained by this project on Wan2.1; same capabilities as V1.1 but without camera control | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14B text-to-video weights, trained at multiple resolutions, supporting start and end frame prediction.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14B video control weights, supporting various control conditions such as Canny, Depth, Pose, MLSD, etc., and trajectory control. Supports multi-resolution (512, 768, 1024) video prediction at 81 frames, trained at 16 frames per second, with multilingual prediction support.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Alignment LoRAs trained with reward backpropagation</td></tr></table> |
|
||||
| Wan2.1 | Video | Official Wan weights covering T2V, I2V, audio-driven, and controlled generation; can be used as training baseline for Wan2.1-Fun | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">Wan 2.1-14B-720P image-to-video model weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE control and subject reference</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE control and subject reference</td></tr></table> |
|
||||
| Self-Forcing / Causal-Forcing / Flex-Forcing | Video | Autoregressive distillation schemes covering streaming, interactive generation, and flexible chunked attention | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">Autoregressive distillation weights, use with Wan2.1-T2V for streaming and interactive generation; Flex-Forcing (chunk-wise causal/bidirectional attention) weights are produced by `scripts/wan2.1_flex_forcing`</td></tr></table> |
|
||||
| TurboWan / TurboDiffusion | Video | Distilled few-step weights publicly released by the TurboDiffusion scheme | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B text-to-video distilled weights; officially released as .pth, the repo also ships a quantised version</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B image-to-video distilled weights; the repo contains low/high noise variants (plus quantised). Place them in Personalized_Model and reference via transformer_path / transformer_high_path</td></tr></table> |
|
||||
| CogVideoX-Fun V1.5 | Video | Official CogVideoX-Fun V1.5 weights, multi-resolution (512/768/1024), 85 frames at 8fps, covering I2V and reward alignment | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024) and has been trained on 85 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
|
||||
| CogVideoX-Fun V1.1 | Video | Official CogVideoX-Fun V1.1 weights, multi-resolution (512/768/1024/1280), 49 frames at 8fps, covering I2V, pose control, controlled generation, and reward alignment | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Noise has been added to the reference image, and the amplitude of motion is greater compared to V1.0.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">Our official pose-control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Our official control video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second. Supporting various control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
|
||||
| CogVideoX-Fun V1.0 | Video | Legacy weights trained at 49 frames 8fps, superseded by V1.1/V1.5 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 49 frames at a rate of 8 frames per second.</td></tr></table> |
|
||||
| HunyuanVideo | Video | Official diffusers-format weights; this project directly supports inference and LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
|
||||
| MiniMax-H3 | Video | Official video generation weights and the ControlNet trained by this project | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">Official MiniMax-H3 T2V/I2V weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet trained by this project, supports multiple control conditions and trajectory control</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet trained by this project (2.0), supporting multiple control conditions, trajectory control, and inpaint checkpoints</td></tr></table> |
|
||||
| TaoMate-H3 | Video+Audio | Official streaming audio-video generation adapter built on MiniMax-H3 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">Official rank-128 adapter (step-3000 EMA) with a built-in 3-step distilled schedule for streaming speech-driven generation; requires the MiniMax-H3 base weights</td></tr></table> |
|
||||
| LTX-2 | Video+Audio | Official DiT audio-video joint generation weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">Official audio-video joint generation weights; repo contains multiple precision files</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">v2.3 requires community-converted diffusers weights; see Lightricks/LTX-2.3 for official weights</td></tr></table> |
|
||||
| LongCat-Video | Video | Official long-video generation weights; supports LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">Official LongCat-Video T2V weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">Official LongCat-Video avatar/digital-human weights</td></tr></table> |
|
||||
| FantasyTalking | Audio-driven Video | Audio-conditioned incremental weights; requires base video weights and audio encoder | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
|
||||
| InfiniteTalk | Audio-driven Video | Audio-conditioned incremental weights; requires base video weights and audio encoder | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">Official InfiniteTalk audio-driven weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">Chinese audio encoder</td></tr></table> |
|
||||
| FlashHead | Audio-driven Video | Official high-fidelity audio-driven head weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">SoulX FlashHead 1.3B audio-driven head weights; requires wav2vec audio encoder</td></tr></table> |
|
||||
| MOVA | Video+Audio | Official MOVA weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">Image-to-video and audio-video joint generation</td></tr></table> |
|
||||
| LingBot | Video | Camera-controllable world model; directory structure matches Wan2.2-I2V-A14B | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">Camera-control baseline weights</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA; use with Qwen3.6-27B generated structured captions</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B dense video generation weights; trainable on 1-2 GPUs</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE (3B active) video generation weights; training requires 8x80GB or more</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">Distilled few-step world model checkpoint (16 transformer shards); its VAE/T5 are reused from lingbot-world-base-cam, and inference must use the Flow_Unipc sampler</td></tr></table> |
|
||||
| Phantom | Video | Incremental weights for multi-subject reference video generation; based on Wan2.1-T2V | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B version. Officially released as .pth; place in Personalized_Model and reference via transformer_path in predict file</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B version. Officially released as sharded safetensors</td></tr></table> |
|
||||
| Qwen-Image | Image | Official text-to-image and image-editing weights; supports baseline and LoRA training | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">Updated text-to-image version</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">Image layer-decomposition weights; splits an image into multiple editable RGBA layers</td></tr></table> |
|
||||
| Qwen-Image-2.1 | Image | Official next-generation text-to-image weights; single-stream block-causal transformer with prefix KV cache | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Single-stream block-causal transformer; supports full-parameter training, prefix KV cache speeds up inference</td></tr></table> |
|
||||
| Qwen-Image ControlNet | Image | Image controlled generation; supports Canny, Depth, Pose, MLSD, and Scribble | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Qwen-Image-2512, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, Scribble, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Equivalent ControlNet provided by InstantX</td></tr></table> |
|
||||
| Z-Image | Image | Official text-to-image weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
|
||||
| Z-Image-Fun | Image | ControlNet and distillation LoRA trained by this project on Z-Image; supports Canny, Depth, Pose, MLSD, Scribble, and Gray | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image. Compared to the first version, it adds to more layers and has been trained for a longer period. It supports multiple control conditions including Canny, Depth, Pose, MLSD, Scribble and Gray.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image-Turbo, supporting multiple control conditions such as Canny, Depth, Pose, MLSD, etc.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for Z-Image-Turbo. Compared to the first version, it adds to more layers and has been trained for a longer period. It supports multiple control conditions including Canny, Depth, Pose, MLSD, and more.</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">This is a Distill LoRA for Z-Image that distills both steps and CFG. This model does not require CFG and uses 8 steps for inference.</td></tr></table> |
|
||||
| Flux | Image | Official FLUX.1/FLUX.2 weights and the ControlNet trained by this project | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">ControlNet weights for FLUX.2-dev</td></tr></table> |
|
||||
| ERNIE-Image | Image | Official Baidu text-to-image weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">Official ERNIE-Image text-to-image weights</td></tr></table> |
|
||||
| Lens | Image | Official Microsoft camera-control weights | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">Official Lens camera-control weights</td></tr></table> |
|
||||
| Auxiliary Models | - | Non-generative models used for reward alignment, data annotation, and fast decoding | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">Scoring model used in reward backpropagation</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">Multimodal encoder used in the video captioning pipeline</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoders (~20 MB) sharing the latent spaces of the Wan2.1 / Wan2.2 VAEs, ~100x faster decoding for previews and low-memory generation; weights from <a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a></td></tr></table> |
|
||||
|
||||
> Notes:
|
||||
> - Audio-driven and reference models (FantasyTalking, InfiniteTalk, Phantom, TaoMate-H3) are incremental weights and must be used together with the corresponding base video weights and audio encoder.
|
||||
> - The TurboWan weights released by the TurboDiffusion scheme are listed above; other distillation schemes such as Flex-Forcing and PDD have no publicly released weights — train them following `scripts/{model_name}/README_TRAIN*.md` and then fill the resulting path into `transformer_path`.
|
||||
> - Weight names map one-to-one to folder names under `models/Diffusion_Transformer/`. Weights within the same family are not interchangeable; choose according to the inference task. If a weight is not listed here, it is either produced by this project or should be obtained from the upstream official repository.
|
||||
|
||||
# IV. Video Works
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control
|
||||
|
||||
Generic Control Video + Reference Image:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
Reference Image
|
||||
</td>
|
||||
<td>
|
||||
Control Video
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-14B-Control
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-1.3B-Control
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
For details on setting some parameters, please refer to [Readme Train](scripts/README_TRAIN.md) and [Readme Lora](scripts/README_TRAIN_LORA.md).
|
||||
|
||||
|
||||
Generic Control Video (Canny, Pose, Depth, etc.) and Trajectory Control:
|
||||
# Model zoo
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
| Name | Storage Space | Url | Hugging Face | Description |
|
||||
|--|--|--|--|--|
|
||||
| CogVideoX-Fun-2b-InP.tar.gz | Before extraction:9.69 GB \/ After extraction: 13.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP)| Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 144 frames at a rate of 24 frames per second. |
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
# TODO List
|
||||
- Support CogVideoX-5b.
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
Pan Up
|
||||
</td>
|
||||
<td>
|
||||
Pan Left
|
||||
</td>
|
||||
<td>
|
||||
Pan Right
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
Pan Down
|
||||
</td>
|
||||
<td>
|
||||
Pan Up + Pan Left
|
||||
</td>
|
||||
<td>
|
||||
Pan Up + Pan Right
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### CogVideoX-Fun-V1.1-5B
|
||||
|
||||
Resolution-1024
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
Resolution-768
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
Resolution-512
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### CogVideoX-Fun-V1.1-5B-Control
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
|
||||
</td>
|
||||
<td>
|
||||
A young woman with beautiful clear eyes and blonde hair, wearing white clothes and twisting her body, with the camera focused on her face. High quality, masterpiece, best quality, high resolution, ultra-fine, dreamlike.
|
||||
</td>
|
||||
<td>
|
||||
A young bear.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
# V. References
|
||||
# Reference
|
||||
- CogVideo: https://github.com/THUDM/CogVideo/
|
||||
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
|
||||
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
|
||||
- Wan2.2: https://github.com/Wan-Video/Wan2.2/
|
||||
- Diffusers: https://github.com/huggingface/diffusers
|
||||
- Qwen-Image: https://github.com/QwenLM/Qwen-Image
|
||||
- Self-Forcing: https://github.com/guandeh17/Self-Forcing
|
||||
- Flux: https://github.com/black-forest-labs/flux
|
||||
- Flux2: https://github.com/black-forest-labs/flux2
|
||||
- HunyuanVideo: https://github.com/Tencent-Hunyuan/HunyuanVideo
|
||||
- ComfyUI-KJNodes: https://github.com/kijai/ComfyUI-KJNodes
|
||||
- ComfyUI-EasyAnimateWrapper: https://github.com/kijai/ComfyUI-EasyAnimateWrapper
|
||||
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
|
||||
- CameraCtrl: https://github.com/hehao13/CameraCtrl
|
||||
|
||||
# VI. Citation
|
||||
|
||||
If you use VideoX-Fun in your research or project, please cite it as follows:
|
||||
|
||||
```bibtex
|
||||
@misc{aigc_apps_VideoX_Fun_2026,
|
||||
author = {aigc-apps},
|
||||
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
|
||||
year = {2026},
|
||||
publisher = {GitHub},
|
||||
url = {https://github.com/aigc-apps/VideoX-Fun}
|
||||
}
|
||||
```
|
||||
|
||||
# VII. Limitations and Risks
|
||||
|
||||
- Generated videos may have artifacts or quality issues, especially in complex scenes.
|
||||
- The model may struggle with fine details, text rendering, or specific artistic styles.
|
||||
- Performance varies with input prompt quality, resolution, and other parameters.
|
||||
- The technology could be misused to create misleading content (e.g., deepfakes). Users are responsible for ethical use.
|
||||
- The model may reflect biases present in the training data.
|
||||
- Users should respect privacy and copyright when using real people's images or videos.
|
||||
|
||||
We encourage responsible use and recommend implementing safeguards in production environments.
|
||||
|
||||
# VIII. License
|
||||
# License
|
||||
This project is licensed under the [Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE).
|
||||
|
||||
The CogVideoX-2B model (including its corresponding Transformers module and VAE module) is released under the [Apache 2.0 License](LICENSE).
|
||||
|
||||
The CogVideoX-5B model (Transformers module) is released under the [CogVideoX LICENSE](https://huggingface.co/THUDM/CogVideoX-5b/blob/main/LICENSE).
|
||||
The CogVideoX-2B model (including its corresponding Transformers module and VAE module) is released under the [Apache 2.0 License](LICENSE).
|
||||
@@ -1,679 +0,0 @@
|
||||
# VideoX-Fun
|
||||
|
||||
😊 ようこそ!
|
||||
|
||||
CogVideoX-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/CogVideoX-Fun-5b)
|
||||
|
||||
Wan-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/Wan2.1-Fun-1.3B-InP)
|
||||
|
||||
[English](./README.md) | [简体中文](./README_zh-CN.md) | 日本語
|
||||
|
||||
# 目次
|
||||
- [一、紹介](#一紹介)
|
||||
- [二、クイックスタートと使用](#二クイックスタートと使用)
|
||||
- [1. 環境準備](#1-環境準備)
|
||||
- [2. 推論生成](#2-推論生成)
|
||||
- [3. モデルのトレーニング](#3-モデルのトレーニング)
|
||||
- [三、サポート済みモデル](#三サポート済みモデル)
|
||||
- [四、ビデオ作品](#四ビデオ作品)
|
||||
- [五、参考文献](#五参考文献)
|
||||
- [六、引用](#六引用)
|
||||
- [七、制限とリスク](#七制限とリスク)
|
||||
- [八、ライセンス](#八ライセンス)
|
||||
|
||||
# 一、紹介
|
||||
VideoX-Funはビデオ生成のパイプラインであり、AI画像やビデオの生成、Diffusion TransformerのベースラインモデルとLoraモデルのトレーニングに使用できます。我々は、すでに学習済みのベースラインモデルから直接予測を行い、異なる解像度、秒数、FPSのビデオを生成することをサポートしています。また、ユーザーが独自のベースラインモデルやLoraモデルをトレーニングし、特定のスタイル変換を行うこともサポートしています。
|
||||
|
||||
新機能:
|
||||
- Wan 2.2シリーズモデル、Wan-VACE制御モデル、Fantasy Talkingデジタルヒューマンモデル、Qwen-Image、Flux画像生成モデルなどのサポートを追加しました。[2025.10.16]
|
||||
- Wan2.1-Fun-V1.1バージョンを更新:14Bと1.3BモデルのControl+参照画像モデルをサポート、カメラ制御にも対応。さらに、Inpaintモデルを再訓練し、性能が向上しました。[2025.04.25]
|
||||
- Wan2.1-Fun-V1.0の更新:14Bおよび1.3BのI2V(画像からビデオ)モデルとControlモデルをサポートし、開始フレームと終了フレームの予測に対応。[2025.03.26]
|
||||
- CogVideoX-Fun-V1.5の更新:I2Vモデルと関連するトレーニング・予測コードをアップロード。[2024.12.16]
|
||||
- 報酬Loraのサポート:報酬逆伝播技術を使用してLoraをトレーニングし、生成された動画を最適化し、人間の好みによりよく一致させる。[詳細情報](scripts/README_TRAIN_REWARD.md)。新しいバージョンの制御モデルでは、Canny、Depth、Pose、MLSDなどの異なる制御条件に対応。[2024.11.21]
|
||||
- diffusersのサポート:CogVideoX-Fun Controlがdiffusersでサポートされるようになりました。[a-r-r-o-w](https://github.com/a-r-r-o-w)がこの[PR](https://github.com/huggingface/diffusers/pull/9671)でサポートを提供してくれたことに感謝します。詳細は[ドキュメント](https://huggingface.co/docs/diffusers/main/en/api/pipelines/cogvideox)をご覧ください。[2024.10.16]
|
||||
- CogVideoX-Fun-V1.1の更新:i2vモデルを再トレーニングし、Noiseを追加して動画の動きの範囲を拡大。制御モデルのトレーニングコードとControlモデルをアップロード。[2024.09.29]
|
||||
- CogVideoX-Fun-V1.0の更新:コードを作成!WindowsとLinuxに対応しました。2Bおよび5Bモデルでの最大256x256x49から1024x1024x49までの任意の解像度の動画生成をサポート。[2024.09.18]
|
||||
|
||||
機能:
|
||||
- [データ前処理](#data-preprocess)
|
||||
- [DiTのトレーニング](#dit-train)
|
||||
- [ビデオ生成](#video-gen)
|
||||
|
||||
私たちのUIインターフェースは次のとおりです:
|
||||

|
||||
|
||||
# 二、クイックスタートと使用
|
||||
|
||||
<a id="quick-start"></a>
|
||||
|
||||
## 1. 環境準備
|
||||
|
||||
### 1.1 クラウド使用: AliyunDSW
|
||||
|
||||
DSWには無料のGPU時間があり、ユーザーは一度申請でき、申請後3か月間有効です。
|
||||
|
||||
Aliyunは[Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1)で無料のGPU時間を提供しています。取得してAliyun PAI-DSWで使用し、5分以内にCogVideoX-Funを開始できます!
|
||||
|
||||
[](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
|
||||
|
||||
### 1.2 ローカル依存のインストール
|
||||
|
||||
以下の環境でこのライブラリの実行を確認しています:
|
||||
|
||||
Windowsの詳細:
|
||||
- OS: Windows 10
|
||||
- python: python3.10 & python3.11
|
||||
- pytorch: torch2.2.0
|
||||
- CUDA: 11.8 & 12.1
|
||||
- CUDNN: 8+
|
||||
- GPU: Nvidia-3060 12G & Nvidia-3090 24G
|
||||
|
||||
Linuxの詳細:
|
||||
- OS: Ubuntu 20.04, CentOS
|
||||
- python: python3.10 & python3.11
|
||||
- pytorch: torch2.2.0
|
||||
- CUDA: 11.8 & 12.1
|
||||
- CUDNN: 8+
|
||||
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
|
||||
|
||||
重みを保存するために約60GBのディスクスペースが必要です。確認してください!
|
||||
|
||||
### 1.3 Dockerの使用
|
||||
|
||||
Dockerを使用する場合、マシンにグラフィックスカードドライバとCUDA環境が正しくインストールされていることを確認してください。
|
||||
|
||||
次のコマンドをこの方法で実行します:
|
||||
|
||||
```
|
||||
# イメージをプル
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# イメージに入る
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# コードをクローン
|
||||
git clone https://github.com/aigc-apps/VideoX-Fun.git
|
||||
|
||||
# VideoX-Funのディレクトリに入る
|
||||
cd VideoX-Fun
|
||||
|
||||
# 重みをダウンロード
|
||||
mkdir models/Diffusion_Transformer
|
||||
mkdir models/Personalized_Model
|
||||
|
||||
# Please use the hugginface link or modelscope link to download the model.
|
||||
# CogVideoX-Fun
|
||||
# https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
|
||||
# https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP
|
||||
|
||||
# Wan
|
||||
# https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP
|
||||
# https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP
|
||||
```
|
||||
|
||||
### 1.4 重みの配置
|
||||
|
||||
[重み](#三サポート済みモデル)を指定されたパスに配置することをお勧めします:
|
||||
|
||||
**ComfyUIを通じて**:
|
||||
モデルをComfyUIの重みフォルダ `ComfyUI/models/Fun_Models/` に入れます:
|
||||
```
|
||||
📦 ComfyUI/
|
||||
├── 📂 models/
|
||||
│ └── 📂 Fun_Models/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
|
||||
│ └── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
|
||||
```
|
||||
|
||||
**独自のpythonファイルまたはUIインターフェースを実行**:
|
||||
```
|
||||
📦 models/
|
||||
├── 📂 Diffusion_Transformer/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
|
||||
│ └── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
|
||||
├── 📂 Personalized_Model/
|
||||
│ └── あなたのトレーニング済みのトランスフォーマーモデル / あなたのトレーニング済みのLoraモデル(UIロード用)
|
||||
```
|
||||
|
||||
## 2. 推論生成
|
||||
|
||||
<a id="video-gen"></a>
|
||||
|
||||
ビデオモデルと画像モデルの推論入口は完全に一致しており、`examples/{model_name}/`下のスクリプトまたはUIから実行します。
|
||||
|
||||
### 2.1 入口の選択
|
||||
|
||||
| 使用入口 | 適用シーン | 設定粒度 |
|
||||
|--|--|--|
|
||||
| Pythonファイル | バッチ生成、スクリプト内でパラメータを調整 | 全パラメータ |
|
||||
| WebUI | 対話的な体験、モデルの迅速な切り替え | よく使うパラメータのみ |
|
||||
| ComfyUI | 既存のComfyUIワークフロー | ノードパラメータ |
|
||||
|
||||
表:推論入口の選択
|
||||
|
||||
### 2.2 顕存節約方案
|
||||
|
||||
Wan2.1のパラメータが非常に大きいため、GPUメモリを節約し、コンシューマー向けGPUに適応させる必要があります。各予測ファイルには`GPU_memory_mode`を提供しており、`model_cpu_offload`、`model_cpu_offload_and_qfloat8`、`sequential_cpu_offload`の中から選択できます。この方法はCogVideoX-Funの生成にも適用されます。
|
||||
|
||||
- `model_cpu_offload`: モデル全体が使用後にCPUに移動し、一部のGPUメモリを節約します。
|
||||
- `model_cpu_offload_and_qfloat8`: モデル全体が使用後にCPUに移動し、Transformerモデルに対してfloat8の量子化を行い、より多くのGPUメモリを節約します。
|
||||
- `sequential_cpu_offload`: モデルの各層が使用後にCPUに移動します。速度は遅くなりますが、大量のGPUメモリを節約します。
|
||||
|
||||
`qfloat8`はモデルの性能を部分的に低下させる可能性がありますが、より多くのGPUメモリを節約できます。十分なGPUメモリがある場合は、`model_cpu_offload`の使用をお勧めします。
|
||||
|
||||
### 2.3 Pythonファイルから
|
||||
|
||||
##### i. 単一GPUでの推論:
|
||||
|
||||
- ステップ1: 対応する[重み](#三サポート済みモデル)をダウンロードし、`models`フォルダに配置します。
|
||||
- ステップ2: 異なる重みと予測目標に基づいて、異なるファイルを使用して予測を行います。現在、このライブラリはCogVideoX-Fun、Wan2.1、およびWan2.1-Funをサポートしています。`examples`フォルダ内のフォルダ名で区別され、異なるモデルがサポートする機能が異なりますので、状況に応じて区別してください。以下はCogVideoX-Funを例として説明します。
|
||||
- テキストからビデオ:
|
||||
- `examples/cogvideox_fun/predict_t2v.py`ファイルで`prompt`、`neg_prompt`、`guidance_scale`、`seed`を変更します。
|
||||
- 次に、`examples/cogvideox_fun/predict_t2v.py`ファイルを実行し、結果が生成されるのを待ちます。結果は`samples/cogvideox-fun-videos`フォルダに保存されます。
|
||||
- 画像からビデオ:
|
||||
- `examples/cogvideox_fun/predict_i2v.py`ファイルで`validation_image_start`、`validation_image_end`、`prompt`、`neg_prompt`、`guidance_scale`、`seed`を変更します。
|
||||
- `validation_image_start`はビデオの開始画像、`validation_image_end`はビデオの終了画像です。
|
||||
- 次に、`examples/cogvideox_fun/predict_i2v.py`ファイルを実行し、結果が生成されるのを待ちます。結果は`samples/cogvideox-fun-videos_i2v`フォルダに保存されます。
|
||||
- ビデオからビデオ:
|
||||
- `examples/cogvideox_fun/predict_v2v.py`ファイルで`validation_video`、`validation_image_end`、`prompt`、`neg_prompt`、`guidance_scale`、`seed`を変更します。
|
||||
- `validation_video`はビデオ生成のための参照ビデオです。以下のデモビデオを使用して実行できます:[デモビデオ](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/play_guitar.mp4)
|
||||
- 次に、`examples/cogvideox_fun/predict_v2v.py`ファイルを実行し、結果が生成されるのを待ちます。結果は`samples/cogvideox-fun-videos_v2v`フォルダに保存されます。
|
||||
- 通常の制御付きビデオ生成(Canny、Pose、Depthなど):
|
||||
- `examples/cogvideox_fun/predict_v2v_control.py`ファイルで`control_video`、`validation_image_end`、`prompt`、`neg_prompt`、`guidance_scale`、`seed`を変更します。
|
||||
- `control_video`は、Canny、Pose、Depthなどの演算子で抽出された制御用ビデオです。以下のデモビデオを使用して実行できます:[デモビデオ](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1.1/pose.mp4)
|
||||
- 次に、`examples/cogvideox_fun/predict_v2v_control.py`ファイルを実行し、結果が生成されるのを待ちます。結果は`samples/cogvideox-fun-videos_v2v_control`フォルダに保存されます。
|
||||
- ステップ3: 自分でトレーニングした他のバックボーンやLoraを組み合わせたい場合は、必要に応じて`examples/{model_name}/predict_t2v.py`や`examples/{model_name}/predict_i2v.py`、`lora_path`を修正します。
|
||||
|
||||
##### ii. 複数GPUでの推論:
|
||||
多カードでの推論を行う際は、xfuserリポジトリのインストールに注意してください。xfuser==0.4.2 と yunchang==0.6.2 のインストールが推奨されます。
|
||||
```
|
||||
pip install xfuser==0.4.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
|
||||
pip install yunchang==0.6.2 --progress-bar off -i https://mirrors.aliyun.com/pypi/simple/
|
||||
```
|
||||
|
||||
`ulysses_degree` と `ring_degree` の積が使用する GPU 数と一致することを確認してください。たとえば、8つのGPUを使用する場合、`ulysses_degree=2` と `ring_degree=4`、または `ulysses_degree=4` と `ring_degree=2` を設定することができます。
|
||||
|
||||
- `ulysses_degree` はヘッド(head)に分割した後の並列化を行います。
|
||||
- `ring_degree` はシーケンスに分割した後の並列化を行います。
|
||||
|
||||
`ring_degree` は `ulysses_degree` よりも通信コストが高いため、これらのパラメータを設定する際には、シーケンス長とモデルのヘッド数を考慮する必要があります。
|
||||
|
||||
8GPUでの並列推論を例に挙げます:
|
||||
|
||||
- **Wan2.1-Fun-V1.1-14B-InP** はヘッド数が40あります。この場合、`ulysses_degree` は40で割り切れる値(例:2, 4, 8など)に設定する必要があります。したがって、8GPUを使用して並列推論を行う場合、`ulysses_degree=8` と `ring_degree=1` を設定できます。
|
||||
|
||||
- **Wan2.1-Fun-V1.1-1.3B-InP** はヘッド数が12あります。この場合、`ulysses_degree` は12で割り切れる値(例:2, 4など)に設定する必要があります。したがって、8GPUを使用して並列推論を行う場合、`ulysses_degree=4` と `ring_degree=2` を設定できます。
|
||||
|
||||
パラメータの設定が完了したら、以下のコマンドで並列推論を実行してください:
|
||||
|
||||
```sh
|
||||
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
|
||||
```
|
||||
|
||||
### 2.4 UIインターフェースから
|
||||
|
||||
WebUIは、テキストからビデオ、画像からビデオ、ビデオからビデオ、および通常の制御付きビデオ生成(Canny、Pose、Depthなど)をサポートします。現在、このライブラリはCogVideoX-Fun、Wan2.1、およびWan2.1-Funをサポートしており、`examples`フォルダ内のフォルダ名で区別されています。異なるモデルがサポートする機能が異なるため、状況に応じて区別してください。以下はCogVideoX-Funを例として説明します。
|
||||
|
||||
- ステップ1: 対応する[重み](#三サポート済みモデル)をダウンロードし、`models`フォルダに配置します。
|
||||
- ステップ2: `examples/cogvideox_fun/app.py`ファイルを実行し、Gradioページに入ります。
|
||||
- ステップ3: ページ上で生成モデルを選択し、`prompt`、`neg_prompt`、`guidance_scale`、`seed`などを入力し、「生成」をクリックして結果が生成されるのを待ちます。結果は`sample`フォルダに保存されます。
|
||||
|
||||
### 2.5 ComfyUIから
|
||||
|
||||
詳細は[ComfyUI README](comfyui/README.md)をご覧ください。
|
||||
|
||||
|
||||
## 3. モデルのトレーニング
|
||||
|
||||
完全なモデルトレーニングパイプラインは、データ前処理とVideo DiTトレーニングで構成されます。
|
||||
|
||||
### 3.1 データ前処理
|
||||
|
||||
<a id="data-preprocess"></a>
|
||||
各モデルの訓練ドキュメントは`scripts/{model_name}/`下に統一されています。詳細は[3.3 各モデルの訓練ドキュメント](#33-各モデルの訓練ドキュメント)を参照してください。
|
||||
|
||||
長いビデオのセグメンテーション、クリーニング、説明のための完全なデータ前処理リンクは、ビデオキャプションセクションの[README](videox_fun/video_caption/README.md)を参照してください。
|
||||
|
||||
テキストから画像およびビデオ生成モデルをトレーニングしたい場合。この形式でデータセットを配置する必要があります。
|
||||
|
||||
```
|
||||
📦 project/
|
||||
├── 📂 datasets/
|
||||
│ ├── 📂 internal_datasets/
|
||||
│ ├── 📂 train/
|
||||
│ │ ├── 📄 00000001.mp4
|
||||
│ │ ├── 📄 00000002.jpg
|
||||
│ │ └── 📄 .....
|
||||
│ └── 📄 json_of_internal_datasets.json
|
||||
```
|
||||
|
||||
json_of_internal_datasets.jsonは標準のJSONファイルです。json内のfile_pathは相対パスとして設定できます。以下のように:
|
||||
```json
|
||||
[
|
||||
{
|
||||
"file_path": "train/00000001.mp4",
|
||||
"text": "スーツとサングラスを着た若い男性のグループが街の通りを歩いている。",
|
||||
"type": "video"
|
||||
},
|
||||
{
|
||||
"file_path": "train/00000002.jpg",
|
||||
"text": "スーツとサングラスを着た若い男性のグループが街の通りを歩いている。",
|
||||
"type": "image"
|
||||
},
|
||||
.....
|
||||
]
|
||||
```
|
||||
|
||||
次のように絶対パスとして設定することもできます:
|
||||
```json
|
||||
[
|
||||
{
|
||||
"file_path": "/mnt/data/videos/00000001.mp4",
|
||||
"text": "スーツとサングラスを着た若い男性のグループが街の通りを歩いている。",
|
||||
"type": "video"
|
||||
},
|
||||
{
|
||||
"file_path": "/mnt/data/train/00000001.jpg",
|
||||
"text": "スーツとサングラスを着た若い男性のグループが街の通りを歩いている。",
|
||||
"type": "image"
|
||||
},
|
||||
.....
|
||||
]
|
||||
```
|
||||
|
||||
### 3.2 Video DiTのトレーニング
|
||||
|
||||
<a id="dit-train"></a>
|
||||
各モデルの訓練スクリプトと起動shは`scripts/{model_name}/`下にあり、shの名称はタスクによって異なります(例:`train.sh`、`train_lora.sh`、`train_control.sh`、`train_control_distill.sh`など)。ディレクトリ内の実際のファイルを基準としてください。
|
||||
|
||||
データ前処理時にデータ形式が相対パスの場合、```scripts/{model_name}/train.sh```を次のように設定します。
|
||||
```
|
||||
export DATASET_NAME="datasets/internal_datasets/"
|
||||
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
|
||||
```
|
||||
|
||||
データ形式が絶対パスの場合、同じスクリプトで次のように設定します(このとき`DATASET_NAME`は空にし、データセットディレクトリのプレフィックスを連結しません)。
|
||||
```
|
||||
export DATASET_NAME=""
|
||||
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
|
||||
```
|
||||
|
||||
最後に対応するスクリプトを実行します。
|
||||
```sh
|
||||
sh scripts/{model_name}/train.sh
|
||||
```
|
||||
|
||||
### 3.3 各モデルの訓練ドキュメント
|
||||
|
||||
パラメータ設定の詳細について、各モデルの訓練ドキュメントは`scripts/{model_name}/`下に統一されています。
|
||||
|
||||
| モデル | ベーストレーニング | LoRAトレーニング | その他 |
|
||||
|--|--|--|--|
|
||||
| Wan2.1-Fun | [EN](scripts/wan2.1_fun/README_TRAIN.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) | [Control ZH](scripts/wan2.1_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
|
||||
| Wan2.2 | [EN](scripts/wan2.2/README_TRAIN.md) / [ZH](scripts/wan2.2/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) | [Distill ZH](scripts/wan2.2/README_TRAIN_DISTILL_zh-CN.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
|
||||
| Wan2.2-Fun | [EN](scripts/wan2.2_fun/README_TRAIN.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) | [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) / [ZH](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) | [Control LoRA ZH](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA_zh-CN.md) |
|
||||
| CogVideoX-Fun | [EN](scripts/cogvideox_fun/README_TRAIN.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) | [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) / [ZH](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) | [Control ZH](scripts/cogvideox_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
|
||||
| Qwen-Image | [EN](scripts/qwenimage/README_TRAIN.md) / [ZH](scripts/qwenimage/README_TRAIN_zh-CN.md) | [EN](scripts/qwenimage/README_TRAIN_LORA.md) / [ZH](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) | [Edit ZH](scripts/qwenimage/README_TRAIN_EDIT_zh-CN.md) |
|
||||
| Qwen-Image-2.1 | [EN](scripts/qwenimage21/README_TRAIN.md) / [ZH](scripts/qwenimage21/README_TRAIN_zh-CN.md) | - | - |
|
||||
| Z-Image | [EN](scripts/z_image/README_TRAIN.md) / [ZH](scripts/z_image/README_TRAIN_zh-CN.md) | [EN](scripts/z_image/README_TRAIN_LORA.md) / [ZH](scripts/z_image/README_TRAIN_LORA_zh-CN.md) | [GRPO LoRA](scripts/z_image/README_TRAIN_GRPO_LORA.md) |
|
||||
|
||||
その他のモデルも同様に、対応する`scripts/{model_name}/`下のREADMEを参照してください。
|
||||
|
||||
# 三、サポート済みモデル
|
||||
|
||||
下表は、現在サポートされているモデル系列と重みをまとめたものです。ビデオモデルと画像モデルは同じ推論・訓練インターフェースを共有しています。各行は1つのモデル系列を表し、第4列は4列のHTML埋め込みテーブル(重み、Hugging Face、ModelScope、説明)です。🤗 は Hugging Face、🤖 は ModelScope(中国国内ネットワーク向け)、`-` は該当チャネルに対応リポジトリがないか、ログイン認証が必要なことを示します。各モデルの訓練ドキュメントについては[3.3 各モデルの訓練ドキュメント](#33-各モデルの訓練ドキュメント)を参照してください。
|
||||
|
||||
| モデル系列 | モダリティ | サポートタスク | 重み / ダウンロード / 説明 |
|
||||
|--|--|--|--|
|
||||
| Wan2.2-Fun | ビデオ | 本プロジェクトがWan2.2で訓練した系列。テキスト/画像から動画、首尾画像、制御生成、カメラ制御をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14Bのテキスト・画像から動画を生成するモデルの重み。複数の解像度で学習されており、動画の最初と最後のフレームの予測をサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-14Bの動画制御用重み。Canny、Depth、Pose、MLSDなどのさまざまな制御条件に対応しており、軌跡制御もサポートしています。512、768、1024の複数解像度での動画生成が可能で、81フレーム、16fpsで学習されています。多言語対応の予測もサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">14B Controlにカメラモーション制御を追加</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B テキストから動画生成用の重み。121フレーム、24 FPSで学習され、先頭/末尾フレーム予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B 動画制御用重み。Canny、Depth、Pose、MLSDなどの制御条件や軌道制御をサポート。121フレーム、24 FPSで学習され、多言語予測に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun-5B カメラレンズ制御用重み。121フレーム、24 FPSで学習され、多言語予測に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-Fun生成動画を報酬逆伝播で最適化するReward LoRA集合</td></tr></table> |
|
||||
| Wan2.2-VACE-Fun | ビデオ | 本プロジェクトがVACE方式で訓練した系列。制御生成と主題参照をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">VACE方式でトレーニングされたWan2.2の制御ウェイト(ベースモデルはWan2.2-T2V-A14B)。Canny、Depth、Pose、MLSD、軌道制御などの異なる制御条件をサポートします。対象を指定して動画生成が可能です。多解像度(512、768、1024)の動画予測をサポートし、81フレームで16FPSでトレーニングされています。多言語予測にも対応しています。</td></tr></table> |
|
||||
| Wan2.2 | ビデオ | Wan公式重み。テキスト/画像から動画、音声駆動、キャラクターアニメーションをカバー。Wan2.2-Fun系列の訓練基線としても使用可能 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-5B テキスト/画像から動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B テキストから動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B 画像から動画生成重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B 音声から動画生成重み、話者駆動デジタルヒューマン</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.2-14B キャラクター置換・モーション転移重み。リポジトリに複数精度ファイルを含む</td></tr></table> |
|
||||
| Wan2.1-Fun V1.1 | ビデオ | 本プロジェクトがWan2.1で訓練したV1.1系列。マルチ解像度(512/768/1024)、81フレーム16fps、テキスト/画像から動画、首尾画像、制御生成、カメラ制御をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのテキスト・画像から動画生成の重み。マルチ解像度で訓練され、最初と最後の画像予測をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのビデオ制御重み。Canny、Depth、Pose、MLSDなどの異なる制御条件に対応し、参照画像+制御条件を使用した制御や軌跡制御をサポートします。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-1.3Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-V1.1-14Bのカメラレンズ制御重み。512、768、1024のマルチ解像度での動画予測をサポートし、81フレーム、毎秒16フレームで訓練されています。多言語予測に対応しています。</td></tr></table> |
|
||||
| Wan2.1-Fun V1.0 | ビデオ | 本プロジェクトがWan2.1で訓練したV1.0系列。V1.1と同じ能力だがカメラ制御は非対応 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14Bのテキスト・画像から動画生成する重み。マルチ解像度で学習され、開始・終了画像予測をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-1.3Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">Wan2.1-Fun-14Bのビデオ制御ウェイト。Canny、Depth、Pose、MLSDなどの異なる制御条件をサポートし、トラジェクトリ制御も利用可能。512、768、1024のマルチ解像度でのビデオ予測をサポートし、81フレーム(1秒間に16フレーム)でトレーニング済みで、多言語予測にも対応しています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">報酬逆伝播で訓練された整列LoRA</td></tr></table> |
|
||||
| Wan2.1 | ビデオ | Wan公式重み。テキスト/画像から動画、音声駆動、制御生成をカバー。Wan2.1-Fun系列の訓練基線としても使用可能 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">万象2.1-14B-720P 画像→動画モデルの重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE制御と主題参照</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE制御と主題参照</td></tr></table> |
|
||||
| Self-Forcing / Causal-Forcing / Flex-Forcing | ビデオ | 自己回帰蒸留方案。ストリーミング生成、インタラクティブ生成、チャンク単位の注意をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">自己回帰蒸留重み、Wan2.1-T2Vと組み合わせて流式・インタラクティブ生成に対応;Flex-Forcing(チャンク単位の因果/双方向注意)の重みは`scripts/wan2.1_flex_forcing`で訓練して生成</td></tr></table> |
|
||||
| TurboWan / TurboDiffusion | ビデオ | TurboDiffusion方案が公開した少ステップ蒸留重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3Bテキストから動画生成の蒸留重み。公式は.pth形式で公開、リポジトリに量子化版も含む</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B画像から動画生成の蒸留重み。リポジトリにlow/highの2種類のノイズモデル(量子化版も含む)を含み、Personalized_Modelに配置しpredictファイルのtransformer_path/transformer_high_pathで指定</td></tr></table> |
|
||||
| CogVideoX-Fun V1.5 | ビデオ | 公式CogVideoX-Fun V1.5重み。マルチ解像度(512/768/1024)、85フレーム8fps、画像から動画と報酬整列をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024)でビデオを予測できます。85フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.5が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。</td></tr></table> |
|
||||
| CogVideoX-Fun V1.1 | ビデオ | 公式CogVideoX-Fun V1.1重み。マルチ解像度(512/768/1024/1280)、49フレーム8fps、画像から動画、ポーズ制御、制御生成、報酬整列をカバー | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。参照画像にノイズが追加され、V1.0と比較して動きの幅が広がっています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">公式のポーズコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">公式のコントロールビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。Canny、Depth、Pose、MLSDなどのさまざまなコントロール条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">公式の報酬逆伝播技術モデルで、CogVideoX-Fun-V1.1が生成するビデオを最適化し、人間の嗜好によりよく合うようにする。</td></tr></table> |
|
||||
| CogVideoX-Fun V1.0 | ビデオ | 旧版重み。49フレーム8fpsで訓練。V1.1/V1.5に置き換え済み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">公式のグラフ生成ビデオモデルは、複数の解像度(512、768、1024、1280)でビデオを予測できます。49フレーム、8フレーム/秒でトレーニングされています。</td></tr></table> |
|
||||
| HunyuanVideo | ビデオ | 公式diffusers形式重み。本プロジェクトは推論とLoRA訓練を直接サポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
|
||||
| MiniMax-H3 | ビデオ | 公式動画生成重みと本プロジェクトが訓練したControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">MiniMax-H3公式T2V/I2V重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本プロジェクトが訓練したControlNet。複数制御条件と軌跡制御をサポート</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">本プロジェクトが訓練したControlNet(2.0版)。複数制御条件、軌跡制御、inpaint重みをサポート</td></tr></table> |
|
||||
| TaoMate-H3 | ビデオ+音声 | MiniMax-H3をベースにした公式ストリーミング音声・動画生成アダプタ | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">公式rank 128アダプタ(step-3000 EMA)。3ステップ蒸留サンプリングスケジュールを内蔵し、ストリーミング音声駆動生成に対応;MiniMax-H3基盤重みと組み合わせて使用</td></tr></table> |
|
||||
| LTX-2 | ビデオ+音声 | 公式DiT音声・動画共同生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">音声・動画共同生成の公式重み。リポジトリに複数精度ファイルを含む</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">v2.3はコミュニティ変換のdiffusers形式重みを使用。公式重みはLightricks/LTX-2.3を参照</td></tr></table> |
|
||||
| LongCat-Video | ビデオ | 公式長尺動画生成重み。LoRA訓練をサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">LongCat-Video公式T2V重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">LongCat-Video公式アバター/デジタルヒューマン重み</td></tr></table> |
|
||||
| FantasyTalking | 音声駆動ビデオ | 音声条件付き増分重み。基盤ビデオ重みと音声エンコーダが必要 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
|
||||
| InfiniteTalk | 音声駆動ビデオ | 音声条件付き増分重み。基盤ビデオ重みと音声エンコーダが必要 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">InfiniteTalk公式オーディオ駆動重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">中国語音声エンコーダ</td></tr></table> |
|
||||
| FlashHead | 音声駆動ビデオ | 公式高品質頭部動作デジタルヒューマン重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">SoulX FlashHead 1.3B 音声駆動頭部重み。wav2vec音声エンコーダが必要</td></tr></table> |
|
||||
| MOVA | ビデオ+音声 | 公式MOVA重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">画像から動画と音声・動画共同生成</td></tr></table> |
|
||||
| LingBot | ビデオ | カメラ制御可能なワールドモデル。ディレクトリ構造はWan2.2-I2V-A14Bと一致 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">カメラ制御基線重み</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA。Qwen3.6-27Bで構造化キャプションを生成して使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B dense版動画生成重み。1〜2枚のGPUで訓練可能</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE(3Bアクティブ)動画生成重み。訓練には8×80GB以上を推奨</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">蒸留少ステップワールドモデル(transformerは16シャード)。VAE/T5はlingbot-world-base-camを再利用し、推論にはFlow_Unipcサンプラーを使用</td></tr></table> |
|
||||
| Phantom | ビデオ | 複数主体参照による動画生成の増分重み。Wan2.1-T2Vベース | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B版。公式は.pth形式で公開。Personalized_Modelに配置しpredictファイルのtransformer_pathで指定</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B版。公式は分割safetensors形式で公開</td></tr></table> |
|
||||
| Qwen-Image | 画像 | 公式テキストから画像生成・画像編集重み。基線とLoRA訓練をサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">テキストから画像生成の更新版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">画像レイヤー分解重み。画像を複数の編集可能なRGBAレイヤーに分解可能</td></tr></table> |
|
||||
| Qwen-Image-2.1 | 画像 | 公式次世代テキストから画像生成重み。シングルストリームblock-causal構造、プレフィックスKV cacheに対応 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">シングルストリームblock-causal構造。全パラメータ訓練をサポート、プレフィックスKV cacheで推論を高速化</td></tr></table> |
|
||||
| Qwen-Image ControlNet | 画像 | 画像制御生成。Canny、Depth、Pose、MLSD、Scribbleをサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Qwen-Image-2512のControlNet重み。Canny、Depth、Pose、MLSD、Scribbleなど、複数の制御条件をサポートします。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">InstantX提供の同種ControlNet</td></tr></table> |
|
||||
| Z-Image | 画像 | 公式テキストから画像生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
|
||||
| Z-Image-Fun | 画像 | 本プロジェクトがZ-Imageで訓練したControlNetと蒸留LoRA。Canny、Depth、Pose、MLSD、Scribble、Grayをサポート | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Z-ImageのControlNet重み、Canny、Depth、Pose、MLSD、ScribbleおよびGrayなど複数の制御条件に対応。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">Z-Image-Turbo用のControlNet重み。Canny、Depth、Pose、MLSDなど複数の制御条件をサポート。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">Z-Image-TurboのControlNet重み。第1版と比較して、より多くの層に追加され、より長時間トレーニングされています。Canny、Depth、Pose、MLSDなど、複数の制御条件をサポートしています。</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">これはZ-Image用の蒸留LoRAで、ステップ数とCFGの両方を蒸留します。このモデルはCFGを必要とせず、推論には8ステップを使用します。</td></tr></table> |
|
||||
| Flux | 画像 | 公式FLUX.1/FLUX.2重みと本プロジェクトが訓練したControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">FLUX.2-dev用ControlNet重み</td></tr></table> |
|
||||
| ERNIE-Image | 画像 | Baidu公式テキストから画像生成重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">ERNIE-Image公式画像生成重み</td></tr></table> |
|
||||
| Lens | 画像 | Microsoft公式カメラ制御重み | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">Lens公式カメラ制御重み</td></tr></table> |
|
||||
| 補助モデル | - | 生成モデルではなく、報酬整列、データアノテーション、高速デコードに使用 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">報酬逆伝播で使用されるスコアリングモデル</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">動画キャプション生成パイプラインで使用されるマルチモーダルエンコーダ</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoder(約20MB)。Wan2.1/Wan2.2 VAEと同じlatent空間を共有し、デコード速度は完全なVAEの約100倍で、高速プレビューや低メモリ生成に使用;重みは<a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a>より</td></tr></table> |
|
||||
|
||||
> 補足説明:
|
||||
> - 音声駆動・参照系モデル(FantasyTalking、InfiniteTalk、Phantom、TaoMate-H3)は増分重みであり、対応する基盤ビデオ重みと音声エンコーダを同時にダウンロードする必要があります。
|
||||
> - TurboDiffusion方案はTurboWan系列の蒸留重みを公開済みです(上表参照)。Flex-Forcing、PDDなどその他の蒸留方案は公開重みがなく、`scripts/{model_name}/README_TRAIN*.md`で訓練後、`transformer_path`に指定して使用できます。
|
||||
> - 重み名は`models/Diffusion_Transformer/`下のフォルダ名と一対一で対応します。同じ系列内の各重みは互換性がないため、推論タスクに応じて選択してください。ここに掲載されていない重みは、本プロジェクトの訓練成果物、または上流の公式リポジトリから取得する必要があります。
|
||||
|
||||
# 四、ビデオ作品
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-InP && Wan2.1-Fun-V1.1-1.3B-InP
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/747b6ab8-9617-4ba2-84a0-b51c0efbd4f8" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ae94dcda-9d5e-4bae-a86f-882c4282a367" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/a4aa1a82-e162-4ab5-8f05-72f79568a191" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/83c005b8-ccbc-44a0-a845-c0472763119c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-Control && Wan2.1-Fun-V1.1-1.3B-Control
|
||||
|
||||
汎用制御動画 + 参照画像:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
参照画像
|
||||
</td>
|
||||
<td>
|
||||
制御動画
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-14B-Control
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-1.3B-Control
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
汎用制御動画(Canny、Pose、Depth など)と軌跡制御:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### Wan2.1-Fun-V1.1-14B-Control-Camera && Wan2.1-Fun-V1.1-1.3B-Control-Camera
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
Pan Up
|
||||
</td>
|
||||
<td>
|
||||
Pan Left
|
||||
</td>
|
||||
<td>
|
||||
Pan Right
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/869fe2ef-502a-484e-8656-fe9e626b9f63" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/2d4185c8-d6ec-4831-83b4-b1dbfc3616fa" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7dfb7cad-ed24-4acc-9377-832445a07ec7" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
Pan Down
|
||||
</td>
|
||||
<td>
|
||||
Pan Up + Pan Left
|
||||
</td>
|
||||
<td>
|
||||
Pan Up + Pan Right
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/3ea3a08d-f2df-43a2-976e-bf2659345373" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/4a85b028-4120-4293-886b-b8afe2d01713" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ad0d58c1-13ef-450c-b658-4fed7ff5ed36" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### CogVideoX-Fun-V1.1-5B
|
||||
|
||||
解像度-1024
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/34e7ec8f-293e-4655-bb14-5e1ee476f788" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7809c64f-eb8c-48a9-8bdc-ca9261fd5434" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8e76aaa4-c602-44ac-bcb4-8b24b72c386c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/19dba894-7c35-4f25-b15c-384167ab3b03" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
解像度-768
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/0bc339b9-455b-44fd-8917-80272d702737" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/70a043b9-6721-4bd9-be47-78b7ec5c27e9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d5dd6c09-14f3-40f8-8b6d-91e26519b8ac" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/9327e8bc-4f17-46b0-b50d-38c250a9483a" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
解像度-512
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ef407030-8062-454d-aba3-131c21e6b58c" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7610f49e-38b6-4214-aa48-723ae4d1b07e" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/1fff0567-1e15-415c-941e-53ee8ae2c841" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/bcec48da-b91b-43a0-9d50-cf026e00fa4f" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
### CogVideoX-Fun-V1.1-5B-Control
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/53002ce2-dd18-4d4f-8135-b6f68364cabd" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/a1a07cf8-d86d-4cd2-831f-18a6c1ceee1d" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/3224804f-342d-4947-918d-d9fec8e3d273" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
|
||||
</td>
|
||||
<td>
|
||||
美しい澄んだ目と金髪の若い女性が白い服を着て体をひねり、カメラは彼女の顔に焦点を合わせています。高品質、傑作、最高品質、高解像度、超微細、夢のような。
|
||||
</td>
|
||||
<td>
|
||||
若いクマ。
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ea908454-684b-4d60-b562-3db229a250a9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ffb7c6fc-8b69-453b-8aad-70dfae3899b9" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d3f757a3-3551-4dcb-9372-7a61469813f5" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
# 五、参考文献
|
||||
- CogVideo: https://github.com/THUDM/CogVideo/
|
||||
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
|
||||
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
|
||||
- Wan2.2: https://github.com/Wan-Video/Wan2.2/
|
||||
- Diffusers: https://github.com/huggingface/diffusers
|
||||
- Qwen-Image: https://github.com/QwenLM/Qwen-Image
|
||||
- Self-Forcing: https://github.com/guandeh17/Self-Forcing
|
||||
- Flux: https://github.com/black-forest-labs/flux
|
||||
- Flux2: https://github.com/black-forest-labs/flux2
|
||||
- HunyuanVideo: https://github.com/Tencent-Hunyuan/HunyuanVideo
|
||||
- ComfyUI-KJNodes: https://github.com/kijai/ComfyUI-KJNodes
|
||||
- ComfyUI-EasyAnimateWrapper: https://github.com/kijai/ComfyUI-EasyAnimateWrapper
|
||||
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
|
||||
- CameraCtrl: https://github.com/hehao13/CameraCtrl
|
||||
|
||||
# 六、引用
|
||||
|
||||
研究やプロジェクトでVideoX-Funを使用する場合は、以下の形式で引用してください:
|
||||
|
||||
```bibtex
|
||||
@misc{aigc_apps_VideoX_Fun_2026,
|
||||
author = {aigc-apps},
|
||||
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
|
||||
year = {2026},
|
||||
publisher = {GitHub},
|
||||
url = {https://github.com/aigc-apps/VideoX-Fun}
|
||||
}
|
||||
```
|
||||
|
||||
# 七、制限とリスク
|
||||
|
||||
- 生成された動画には、特に複雑なシーンでアーティファクトや品質の問題がある場合があります。
|
||||
- モデルは、細かい詳細、テキストのレンダリング、または特定の芸術スタイルで苦労する場合があります。
|
||||
- パフォーマンスは、入力プロンプトの品質、解像度、その他のパラメータによって異なります。
|
||||
- この技術は、誤解を招くコンテンツ(例:ディープフェイク)を作成するために悪用される可能性があります。ユーザーは倫理的な使用に責任を持ちます。
|
||||
- モデルは、トレーニングデータに存在するバイアスを反映する可能性があります。
|
||||
- ユーザーは、実在の人物の画像や動画を使用する際、プライバシーと著作権を尊重する必要があります。
|
||||
|
||||
責任ある使用を推奨し、本番環境でのセーフガードの実装をお勧めします。
|
||||
|
||||
# 八、ライセンス
|
||||
このプロジェクトは[Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE)の下でライセンスされています。
|
||||
|
||||
CogVideoX-2Bモデル(対応するTransformersモジュール、VAEモジュールを含む)は、[Apache 2.0ライセンス](LICENSE)の下でリリースされています。
|
||||
|
||||
CogVideoX-5Bモデル(Transformersモジュール)は、[CogVideoXライセンス](https://huggingface.co/THUDM/CogVideoX-5b/blob/main/LICENSE)の下でリリースされています。
|
||||
@@ -1,48 +1,76 @@
|
||||
# VideoX-Fun
|
||||
# CogVideoX-Fun
|
||||
|
||||
😊 Welcome!
|
||||
|
||||
CogVideoX-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/CogVideoX-Fun-5b)
|
||||
|
||||
Wan-Fun:
|
||||
[](https://huggingface.co/spaces/alibaba-pai/Wan2.1-Fun-1.3B-InP)
|
||||
|
||||
[English](./README.md) | 简体中文 | [日本語](./README_ja-JP.md)
|
||||
[English](./README.md) | 简体中文
|
||||
|
||||
# 目录
|
||||
- [一、简介](#一简介)
|
||||
- [二、快速开始与使用](#二快速开始与使用)
|
||||
- [1. 环境准备](#1-环境准备)
|
||||
- [2. 推理生成](#2-推理生成)
|
||||
- [3. 模型训练](#3-模型训练)
|
||||
- [三、已支持的模型](#三已支持的模型)
|
||||
- [四、视频作品](#四视频作品)
|
||||
- [五、参考文献](#五参考文献)
|
||||
- [六、引用](#六引用)
|
||||
- [七、限制与风险](#七限制与风险)
|
||||
- [八、许可证](#八许可证)
|
||||
- [目录](#目录)
|
||||
- [简介](#简介)
|
||||
- [快速启动](#快速启动)
|
||||
- [如何使用](#如何使用)
|
||||
- [模型地址](#模型地址)
|
||||
- [未来计划](#未来计划)
|
||||
- [参考文献](#参考文献)
|
||||
- [许可证](#许可证)
|
||||
|
||||
# 一、简介
|
||||
VideoX-Fun是一个图片与视频生成的pipeline,可用于生成AI图片与视频、训练Diffusion Transformer的基线模型与Lora模型。我们同时支持视频与图片两类Diffusion Transformer模型:视频侧涵盖Wan2.1/Wan2.2(含Fun、VACE、Animate、S2V等变体)、CogVideoX-Fun、HunyuanVideo、MiniMax-H3、LTX-2、LongCat-Video、FantasyTalking与LingBot等,图片侧涵盖Qwen-Image(含Edit)、Z-Image(含Turbo)、Flux/Flux2与ERNIE-Image等,完整列表见[已支持的模型](#三已支持的模型)。在此基础上,我们支持从已经训练好的基线模型直接进行预测,生成不同分辨率、不同秒数、不同FPS的视频与不同分辨率的图片,也支持用户训练自己的基线模型与Lora模型,进行一定的风格变换。
|
||||
# 简介
|
||||
CogVideoX-Fun是一个基于CogVideoX结构修改后的的pipeline,是一个生成条件更自由的CogVideoX,可用于生成AI图片与视频、训练Diffusion Transformer的基线模型与Lora模型,我们支持从已经训练好的CogVideoX-Fun模型直接进行预测,生成不同分辨率,6秒左右、fps8的视频(1 ~ 49帧),也支持用户训练自己的基线模型与Lora模型,进行一定的风格变换。
|
||||
|
||||
我们会逐渐支持从不同平台快速启动,请参阅 [快速启动](#快速启动)。
|
||||
|
||||
# 二、快速开始与使用
|
||||
新特性:
|
||||
- 创建代码!现在支持 Windows 和 Linux。支持最大256x256x49到1024x1024x49的任意分辨率的视频生成。[ 2024.09.09 ]
|
||||
|
||||
<a id="quick-start"></a>
|
||||
功能概览:
|
||||
- [数据预处理](#data-preprocess)
|
||||
- [训练DiT](#dit-train)
|
||||
- [模型生成](#video-gen)
|
||||
|
||||
## 1. 环境准备
|
||||
这些是我们的生成结果 [GALLERY](scripts/Result_Gallery.md) (点击下方的图片可查看视频):
|
||||
|
||||
### 1.1 云使用: AliyunDSW
|
||||
DSW 有免费 GPU 时间,用户可申请一次,申请后3个月内有效。
|
||||
我们的ui界面如下:
|
||||

|
||||
|
||||
阿里云在[Freetier](https://free.aliyun.com/?product=9602825&crowd=enterprise&spm=5176.28055625.J_5831864660.1.e939154aRgha4e&scm=20140722.M_9974135.P_110.MO_1806-ID_9974135-MID_9974135-CID_30683-ST_8512-V_1)提供免费GPU时间,获取并在阿里云PAI-DSW中使用,5分钟内即可启动VideoX-Fun。
|
||||
# 快速启动
|
||||
### 1. 云使用: AliyunDSW/Docker
|
||||
#### a. 通过阿里云 DSW
|
||||
正在路上
|
||||
|
||||
[](https://gallery.pai-ml.com/#/preview/deepLearning/cv/cogvideox_fun)
|
||||
#### b. 通过ComfyUI
|
||||
我们的ComfyUI界面如下,具体查看[ComfyUI README](comfyui/README.md)。
|
||||

|
||||
|
||||
### 1.2 本地依赖安装
|
||||
#### c. 通过docker
|
||||
使用docker的情况下,请保证机器中已经正确安装显卡驱动与CUDA环境,然后以此执行以下命令:
|
||||
|
||||
我们已验证该库可在以下环境中执行:
|
||||
```
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/CogVideoX-Fun.git
|
||||
|
||||
# enter CogVideoX-Fun's dir
|
||||
cd CogVideoX-Fun
|
||||
|
||||
# download weights
|
||||
mkdir models/Diffusion_Transformer
|
||||
mkdir models/Personalized_Model
|
||||
|
||||
wget https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz -O models/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz
|
||||
|
||||
cd models/Diffusion_Transformer/
|
||||
tar -xvf CogVideoX-Fun-2b-InP.tar.gz
|
||||
cd ../../
|
||||
```
|
||||
|
||||
### 2. 本地安装: 环境检查/下载/安装
|
||||
#### a. 环境检查
|
||||
我们已验证CogVideoX-Fun可在以下环境中执行:
|
||||
|
||||
Windows 的详细信息:
|
||||
- 操作系统 Windows 10
|
||||
@@ -58,169 +86,47 @@ Linux 的详细信息:
|
||||
- pytorch: torch2.2.0
|
||||
- CUDA: 11.8 & 12.1
|
||||
- CUDNN: 8+
|
||||
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G & Nvidia-H800 80G
|
||||
- GPU:Nvidia-V100 16G & Nvidia-A10 24G & Nvidia-A100 40G & Nvidia-A100 80G
|
||||
|
||||
**方式一:使用requirements.txt**
|
||||
我们需要大约 60GB 的可用磁盘空间,请检查!
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
#### b. 权重放置
|
||||
我们最好将[权重](#model-zoo)按照指定路径进行放置:
|
||||
|
||||
**方式二:手动安装依赖**
|
||||
|
||||
```bash
|
||||
# 核心依赖,与requirements.txt保持一致
|
||||
pip install Pillow einops safetensors timm tomesd albumentations librosa "torch>=2.1.2" torchdiffeq torchsde decord datasets numpy scikit-image
|
||||
pip install omegaconf SentencePiece imageio[ffmpeg] imageio[pyav] tensorboard beautifulsoup4 ftfy func_timeout onnxruntime
|
||||
pip install "peft>=0.17.0" "accelerate>=0.25.0" "gradio>=3.41.2" "diffusers>=0.30.1" "transformers>=4.46.2"
|
||||
# 权重下载
|
||||
pip install modelscope
|
||||
# 多卡并行推理需要,推荐固定版本,单卡可跳过
|
||||
pip install "xfuser==0.4.2"
|
||||
# opencv统一使用headless版本,避免部分环境下的GUI依赖
|
||||
pip uninstall opencv-python opencv-contrib-python opencv-python-headless -y
|
||||
pip install opencv-python-headless
|
||||
# 训练可选:DeepSpeed训练需要,固定numpy版本以避免兼容性问题
|
||||
pip install deepspeed==0.17.0 numpy==1.26.4
|
||||
# 加速可选:安装后注意力自动使用Flash Attention后端,未安装时回退到SDPA
|
||||
pip install flash-attn --no-build-isolation
|
||||
```
|
||||
|
||||
> 说明:`torch`与`flash-attn`建议按照本机的CUDA版本从官方渠道安装指定版本,国内网络可追加`-i https://mirrors.aliyun.com/pypi/simple/`加速,具体依赖请以[requirements.txt](requirements.txt)为准。
|
||||
|
||||
### 1.3 使用Docker
|
||||
使用docker的情况下,请保证机器中已经正确安装显卡驱动与CUDA环境,然后以此执行以下命令:
|
||||
|
||||
```
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/VideoX-Fun.git
|
||||
|
||||
# enter VideoX-Fun's dir
|
||||
cd VideoX-Fun
|
||||
```
|
||||
|
||||
### 1.4 权重放置
|
||||
我们最好将[权重](#三已支持的模型)按照指定路径进行放置:
|
||||
|
||||
**运行自身的python文件或ui界面**:
|
||||
```
|
||||
📦 models/
|
||||
├── 📂 Diffusion_Transformer/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
|
||||
│ ├── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
|
||||
│ ├── 📂 Z-Image/
|
||||
│ └── 📂 Qwen-Image/
|
||||
│ └── 📂 CogVideoX-Fun-2b-InP/
|
||||
├── 📂 Personalized_Model/
|
||||
│ └── your trained trainformer model / your trained lora model (for UI load)
|
||||
```
|
||||
|
||||
视频模型与图片模型的权重均统一放在`models/Diffusion_Transformer/`下,文件夹名与[已支持的模型](#三已支持的模型)中的权重名保持一致。
|
||||
# 如何使用
|
||||
|
||||
**通过comfyui**:
|
||||
将模型放入Comfyui的权重文件夹`ComfyUI/models/Fun_Models/`:
|
||||
```
|
||||
📦 ComfyUI/
|
||||
├── 📂 models/
|
||||
│ └── 📂 Fun_Models/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-2b-InP/
|
||||
│ ├── 📂 CogVideoX-Fun-V1.1-5b-InP/
|
||||
│ ├── 📂 Wan2.1-Fun-V1.1-14B-InP
|
||||
│ └── 📂 Wan2.1-Fun-V1.1-1.3B-InP/
|
||||
```
|
||||
<h3 id="video-gen">1. 生成 </h3>
|
||||
|
||||
## 2. 推理生成
|
||||
#### a. 视频生成
|
||||
##### i、运行python文件
|
||||
- 步骤1:下载对应[权重](#model-zoo)放入models文件夹。
|
||||
- 步骤2:在predict_t2v.py文件中修改prompt、neg_prompt、guidance_scale和seed。
|
||||
- 步骤3:运行predict_t2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos-t2v文件夹中。
|
||||
- 步骤4:如果想结合自己训练的其他backbone与Lora,则看情况修改predict_t2v.py中的predict_t2v.py和lora_path。
|
||||
|
||||
<a id="video-gen"></a>
|
||||
视频模型与图片模型的推理入口完全一致,均由`examples/{model_name}/`下的脚本或界面提供,模型清单见[已支持的模型](#三已支持的模型)。
|
||||
|
||||
### 2.1 入口选择
|
||||
| 使用入口 | 适合场景 | 可配置粒度 |
|
||||
|--|--|--|
|
||||
| python文件 | 批量生成、参数写在脚本里调试 | 全量参数,含`GPU_memory_mode`、`transformer_path`、`lora_path` |
|
||||
| webui | 交互体验、快速切换模型 | 常见参数,显存方案仅4档,见2.2 |
|
||||
| ComfyUI | 已有ComfyUI工作流、节点化组合 | 节点参数,权重放置见1.5 |
|
||||
|
||||
### 2.2 显存节省方案
|
||||
基线模型的参数量普遍很大,为适应消费级显卡,每个预测文件都提供了GPU_memory_mode,视频模型与图片模型通用。可选项按省显存程度从高到低排列,与代码中的判断顺序一致:
|
||||
|
||||
- sequential_cpu_offload:模型的每一层在使用后会进入cpu,速度较慢,节省大量显存。
|
||||
- model_group_offload:以leaf层级在cpu与gpu之间搬运权重,并借助stream异步预取,兼顾速度与显存。
|
||||
- model_cpu_offload_and_qfloat8:整个模型在使用后会进入cpu,并且对transformer模型进行了float8的量化,可以节省更多的显存。
|
||||
- model_cpu_offload:整个模型在使用后会进入cpu,可以节省部分显存。
|
||||
- model_full_load_and_qfloat8:模型常驻gpu,仅对transformer做float8量化,显存临界且对速度要求较高时可选。
|
||||
- 默认(传入model_full_load或其他取值):模型全部进入gpu,速度最快,显存需求最高。
|
||||
|
||||
qfloat8会部分降低模型的性能,但可以节省更多的显存。如果显存足够,推荐使用model_cpu_offload。
|
||||
|
||||
> 注意:`app.py`中仅提供model_full_load、model_cpu_offload、model_cpu_offload_and_qfloat8、sequential_cpu_offload四种模式,`model_group_offload`与`model_full_load_and_qfloat8`需在python预测文件中使用;另外compile类加速与`sequential_cpu_offload`、fsdp_dit不兼容。
|
||||
|
||||
### 2.3 通过python文件
|
||||
推理脚本统一命名为`predict_{任务}.py`,在脚本内修改`model_name`、prompt等参数后直接运行,结果保存到脚本中`save_path`指定的目录。视频模型与图片模型的差别只在任务后缀,例如`examples/cogvideox_fun/predict_t2v.py`、`examples/wan2.2_fun/predict_i2v.py`与`examples/z_image/predict_t2i.py`、`examples/qwenimage/predict_t2i_edit.py`。具体某个模型支持哪些任务,以`examples/{model_name}/`下实际存在的脚本为准。
|
||||
|
||||
**i、单卡运行**:以CogVideoX-Fun为例。
|
||||
|
||||
- 步骤1:下载对应[权重](#三已支持的模型)并按1.5放入models文件夹。
|
||||
- 步骤2:根据不同的权重与预测目标使用不同的文件进行预测。
|
||||
- 文生视频:
|
||||
- 使用examples/cogvideox_fun/predict_t2v.py文件中修改prompt、neg_prompt、guidance_scale和seed。
|
||||
- 而后运行examples/cogvideox_fun/predict_t2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos-t2v文件夹中。
|
||||
- 图生视频:
|
||||
- 使用examples/cogvideox_fun/predict_i2v.py文件中修改validation_image_start、validation_image_end、prompt、neg_prompt、guidance_scale和seed。
|
||||
- validation_image_start是视频的开始图片,validation_image_end是视频的结尾图片。
|
||||
- 而后运行examples/cogvideox_fun/predict_i2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos_i2v文件夹中。
|
||||
- 视频生视频:
|
||||
- 使用examples/cogvideox_fun/predict_v2v.py文件中修改validation_video、validation_image_end、prompt、neg_prompt、guidance_scale和seed。
|
||||
- validation_video是视频生视频的参考视频。您可以使用以下视频运行演示:[演示视频](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/play_guitar.mp4)
|
||||
- 而后运行examples/cogvideox_fun/predict_v2v.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos_v2v文件夹中。
|
||||
- 普通控制生视频(Canny、Pose、Depth等):
|
||||
- 使用examples/cogvideox_fun/predict_v2v_control.py文件中修改control_video、validation_image_end、prompt、neg_prompt、guidance_scale和seed。
|
||||
- control_video是控制生视频的控制视频,是使用Canny、Pose、Depth等算子提取后的视频。您可以使用以下视频运行演示:[演示视频](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1.1/pose.mp4)
|
||||
- 而后运行examples/cogvideox_fun/predict_v2v_control.py文件,等待生成结果,结果保存在samples/cogvideox-fun-videos_control文件夹中。
|
||||
- 步骤3:如果想结合自己训练的其他backbone与Lora,则在对应的`examples/{model_name}/predict_*.py`中设置`transformer_path`与`lora_path`(Wan2.2双 Transformer模型另有`transformer_high_path`与`lora_high_path`,分别对应high noise阶段)。
|
||||
|
||||
**ii、多卡运行**:
|
||||
多卡并行推理所需的`xfuser`已列入1.3,推荐固定为`xfuser==0.4.2`。
|
||||
|
||||
请确保ulysses_degree和ring_degree的乘积等于使用的GPU数量。例如,如果您使用8个GPU,则可以设置ulysses_degree=2和ring_degree=4,也可以设置ulysses_degree=4和ring_degree=2。
|
||||
|
||||
ulysses_degree是在head进行切分后并行生成,ring_degree是在sequence上进行切分后并行生成。ring_degree相比ulysses_degree有更大的通信成本,在设置参数时需要结合序列长度和模型的head数进行设置。
|
||||
|
||||
以8卡并行预测为例。
|
||||
- 以Wan2.1-Fun-V1.1-14B-InP为例,其head数为40,ulysses_degree需要设置为其可以整除的数如2、4、8等。因此在使用8卡并行预测时,可以设置ulysses_degree=8和ring_degree=1.
|
||||
- 以Wan2.1-Fun-V1.1-1.3B-InP为例,其head数为12,ulysses_degree需要设置为其可以整除的数如2、4等。因此在使用8卡并行预测时,可以设置ulysses_degree=4和ring_degree=2.
|
||||
|
||||
设置完成后,使用如下指令进行并行预测:
|
||||
```sh
|
||||
torchrun --nproc-per-node=8 examples/wan2.1_fun/predict_t2v.py
|
||||
```
|
||||
|
||||
### 2.4 通过ui界面
|
||||
webui支持文生视频、图生视频、视频生视频和普通控制生视频(Canny、Pose、Depth等)。当前提供`app.py`的是CogVideoX-Fun、Wan2.1、Wan2.1-Fun、Wan2.2、Wan2.2-Fun(界面实现位于`videox_fun/ui/`),其余模型(包含图片模型)请使用python文件进行预测。以CogVideoX-Fun为例。
|
||||
|
||||
- 步骤1:下载对应[权重](#三已支持的模型)并按1.5放入models文件夹。
|
||||
- 步骤2:运行examples/cogvideox_fun/app.py文件,进入gradio页面。
|
||||
##### ii、通过ui界面
|
||||
- 步骤1:下载对应[权重](#model-zoo)放入models文件夹。
|
||||
- 步骤2:运行app.py文件,进入gradio页面。
|
||||
- 步骤3:根据页面选择生成模型,填入prompt、neg_prompt、guidance_scale和seed等,点击生成,等待生成结果,结果保存在sample文件夹中。
|
||||
|
||||
### 2.5 通过ComfyUI
|
||||
具体查看[ComfyUI README](comfyui/README.md),我们的ComfyUI界面如下:
|
||||

|
||||
##### iii、通过comfyui
|
||||
具体查看[ComfyUI README](comfyui/README.md)。
|
||||
|
||||
## 3. 模型训练
|
||||
一个完整的模型训练链路应该包括数据预处理和Video DiT训练。不同模型的训练流程类似,数据格式也类似。
|
||||
### 2. 模型训练
|
||||
一个完整的CogVideoX-Fun训练链路应该包括数据预处理和Video DiT训练。
|
||||
|
||||
<a id="data-preprocess"></a>
|
||||
### 3.1 数据预处理
|
||||
各模型的 LoRA 训练文档统一放在 `scripts/{model_name}/` 下,中文版以 `_zh-CN` 结尾,详情见[3.3 各模型训练文档](#33-各模型训练文档)。
|
||||
<h4 id="data-preprocess">a.数据预处理</h4>
|
||||
我们给出了一个简单的demo通过图片数据训练lora模型,详情可以查看[wiki](https://github.com/aigc-apps/CogVideoX-Fun/wiki/Training-Lora)。
|
||||
|
||||
一个完整的长视频切分、清洗、描述的数据预处理链路可以参考video caption部分的[README](videox_fun/video_caption/README_zh-CN.md)进行。
|
||||
一个完整的长视频切分、清洗、描述的数据预处理链路可以参考video caption部分的[README](cogvideox/video_caption/README.md)进行。
|
||||
|
||||
如果期望训练一个文生图视频的生成模型,您需要以这种格式排列数据集。
|
||||
```
|
||||
@@ -267,256 +173,44 @@ json_of_internal_datasets.json是一个标准的json文件。json中的file_path
|
||||
.....
|
||||
]
|
||||
```
|
||||
<h4 id="dit-train">b. Video DiT训练 </h4>
|
||||
|
||||
<a id="dit-train"></a>
|
||||
### 3.2 Video DiT训练
|
||||
各模型的训练脚本与启动sh均位于`scripts/{model_name}/`下,sh的命名随任务而变,如`train.sh`、`train_lora.sh`、`train_control.sh`、`train_control_distill.sh`等,以目录内实际文件为准。
|
||||
|
||||
如果数据预处理时,数据的格式为相对路径,则进入对应的`scripts/{model_name}/train.sh`进行如下设置。
|
||||
如果数据预处理时,数据的格式为相对路径,则进入scripts/train.sh进行如下设置。
|
||||
```
|
||||
export DATASET_NAME="datasets/internal_datasets/"
|
||||
export DATASET_META_NAME="datasets/internal_datasets/json_of_internal_datasets.json"
|
||||
|
||||
...
|
||||
|
||||
train_data_format="normal"
|
||||
```
|
||||
|
||||
如果数据的格式为绝对路径,则在同一个脚本中设置如下(此时`DATASET_NAME`置空,不再拼接数据集目录前缀)。
|
||||
如果数据的格式为绝对路径,则进入scripts/train.sh进行如下设置。
|
||||
```
|
||||
export DATASET_NAME=""
|
||||
export DATASET_META_NAME="/mnt/data/json_of_internal_datasets.json"
|
||||
```
|
||||
|
||||
最后运行对应的脚本。
|
||||
最后运行scripts/train.sh。
|
||||
```sh
|
||||
sh scripts/{model_name}/train.sh
|
||||
sh scripts/train.sh
|
||||
```
|
||||
|
||||
### 3.3 各模型训练文档
|
||||
关于参数设置细节,各模型的训练文档统一放在`scripts/{model_name}/`下,`README_TRAIN*`为基线训练,`README_TRAIN_LORA*`为LoRA训练,`README_TRAIN_CONTROL*`为控制训练,中文版以`_zh-CN`结尾。常用模型如下:
|
||||
关于一些参数的设置细节,可以查看[Readme Train](scripts/README_TRAIN.md)与[Readme Lora](scripts/README_TRAIN_LORA.md)
|
||||
|
||||
| 模型 | 基线训练 | LoRA训练 | 其他 |
|
||||
|--|--|--|--|
|
||||
| Wan2.1-Fun | [中文](scripts/wan2.1_fun/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.1_fun/README_TRAIN.md) | [中文](scripts/wan2.1_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.1_fun/README_TRAIN_LORA.md) | [Control 中文](scripts/wan2.1_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/wan2.1_fun/README_TRAIN_REWARD.md) |
|
||||
| Wan2.2 | [中文](scripts/wan2.2/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.2/README_TRAIN.md) | [中文](scripts/wan2.2/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.2/README_TRAIN_LORA.md) | [蒸馏 中文](scripts/wan2.2/README_TRAIN_DISTILL_zh-CN.md)、[S2V](scripts/wan2.2/README_TRAIN_S2V_zh-CN.md)、[Animate](scripts/wan2.2/README_TRAIN_ANIMATE.md) |
|
||||
| Wan2.2-Fun | [中文](scripts/wan2.2_fun/README_TRAIN_zh-CN.md) / [EN](scripts/wan2.2_fun/README_TRAIN.md) | [中文](scripts/wan2.2_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/wan2.2_fun/README_TRAIN_LORA.md) | [Control LoRA 中文](scripts/wan2.2_fun/README_TRAIN_CONTROL_LORA_zh-CN.md) |
|
||||
| CogVideoX-Fun | [中文](scripts/cogvideox_fun/README_TRAIN_zh-CN.md) / [EN](scripts/cogvideox_fun/README_TRAIN.md) | [中文](scripts/cogvideox_fun/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/cogvideox_fun/README_TRAIN_LORA.md) | [Control 中文](scripts/cogvideox_fun/README_TRAIN_CONTROL_zh-CN.md)、[Reward LoRA](scripts/cogvideox_fun/README_TRAIN_REWARD.md) |
|
||||
| Qwen-Image | [中文](scripts/qwenimage/README_TRAIN_zh-CN.md) / [EN](scripts/qwenimage/README_TRAIN.md) | [中文](scripts/qwenimage/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/qwenimage/README_TRAIN_LORA.md) | [Edit 中文](scripts/qwenimage/README_TRAIN_EDIT_zh-CN.md) |
|
||||
| Qwen-Image-2.1 | [中文](scripts/qwenimage21/README_TRAIN_zh-CN.md) / [EN](scripts/qwenimage21/README_TRAIN.md) | - | - |
|
||||
| Z-Image | [中文](scripts/z_image/README_TRAIN_zh-CN.md) / [EN](scripts/z_image/README_TRAIN.md) | [中文](scripts/z_image/README_TRAIN_LORA_zh-CN.md) / [EN](scripts/z_image/README_TRAIN_LORA.md) | [GRPO LoRA 中文](scripts/z_image/README_TRAIN_GRPO_LORA_zh-CN.md) |
|
||||
# 模型地址
|
||||
| 名称 | 存储空间 | 下载地址 | Hugging Face | 描述 |
|
||||
|--|--|--|--|--|
|
||||
| CogVideoX-Fun-2b-InP.tar.gz | 解压前 9.69 GB / 解压后 13.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP)| 官方的图生视频权重。支持多分辨率(512,768,1024,1280)的视频预测,以144帧、每秒24帧进行训练 |
|
||||
|
||||
其余模型(如HunyuanVideo、MiniMax-H3、Flux2-Fun、InfiniteTalk、LingBot等)同理,直接查看对应`scripts/{model_name}/`下的README即可。
|
||||
|
||||
|
||||
# 三、已支持的模型
|
||||
下表按模型系列汇总目前已支持的权重,视频模型与图片模型共用同一套推理与训练入口。每个系列一行,第四列为内嵌的四列表格,依次为权重、Hugging Face、ModelScope、对应说明;🤗 为 Hugging Face、🤖 为 ModelScope(国内网络推荐),`-` 表示该渠道确认无对应仓库或需登录授权。各模型训练文档见[3.3 各模型训练文档](#33-各模型训练文档)。
|
||||
|
||||
| 模型系列 | 模态 | 支持任务 | 权重 / 下载 / 说明 |
|
||||
|--|--|--|--|
|
||||
| Wan2.2-Fun | 视频 | 本项目在Wan2.2上训练的系列,覆盖文生视频、图生视频、首尾图、控制生成、相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE双阶段文/图生视频,多分辨率训练、81帧16fps,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">14B控制生成,支持Canny、Depth、Pose、MLSD与轨迹控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-A14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-A14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-A14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">在14B Control基础上增加相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5B统一VAE文/图生视频,121帧24fps,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">5B控制生成,控制条件与14B一致</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-5B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-5B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-5B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">5B相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA,叠加在上述权重上使用</td></tr></table> |
|
||||
| Wan2.2-VACE-Fun | 视频 | 本项目以VACE方案训练的系列,覆盖控制生成、主体参考 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-VACE-Fun-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.2-VACE-Fun-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.2-VACE-Fun-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">以Wan2.2-T2V-A14B为基础,支持Canny、Depth、Pose、MLSD、轨迹控制与主体参考生视频</td></tr></table> |
|
||||
| Wan2.2 | 视频 | 万象官方权重,覆盖文生视频、图生视频、音频驱动、角色动画,可作为Wan2.2-Fun系列的训练基线 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-TI2V-5B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-TI2V-5B">🤖</a></td><td valign="top" style="padding:2px 0;">5B统一VAE,文生图生视频通用权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-T2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-T2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-I2V-A14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.2-I2V-A14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B MoE图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-S2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-S2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-S2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">语音驱动数字人</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.2-Animate-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.2-Animate-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.2-Animate-14B">🤖</a></td><td valign="top" style="padding:2px 0;">角色替换与动作迁移,仓库含多精度文件</td></tr></table> |
|
||||
| Wan2.1-Fun V1.1 | 视频 | 本项目在Wan2.1上训练的V1.1版本,多分辨率(512/768/1024)、81帧16fps,覆盖文生视频、图生视频、首尾图、控制生成、相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B轻量文/图生视频,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">14B文/图生视频,支持首尾图</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B控制生成,同时支持参考图+控制条件组合</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">14B控制生成,同时支持参考图+控制条件组合</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-1.3B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-1.3B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B相机运动控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-V1.1-14B-Control-Camera</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-14B-Control-Camera">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-V1.1-14B-Control-Camera">🤖</a></td><td valign="top" style="padding:2px 0;">14B相机运动控制</td></tr></table> |
|
||||
| Wan2.1-Fun V1.0 | 视频 | 本项目在Wan2.1上训练的V1.0版本,能力与V1.1相同但无相机控制 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的1.3B文/图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的14B文/图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-1.3B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-1.3B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-1.3B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的1.3B控制生成</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-14B-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-14B-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-14B-Control">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的14B控制生成</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-Fun-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Wan2.1-Fun-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Wan2.1-Fun-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
|
||||
| Wan2.1 | 视频 | 万象官方权重,覆盖文生视频、图生视频、控制生成,可作为Wan2.1-Fun系列的训练基线 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-T2V-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-T2V-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-T2V-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">480P图生视频,是InfiniteTalk的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-I2V-14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-I2V-14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Wan-AI/Wan2.1-I2V-14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">720P图生视频,是FantasyTalking的基础模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-1.3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-1.3B">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B VACE控制与主体参考</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Wan2.1-VACE-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Wan-AI/Wan2.1-VACE-14B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Wan-AI/Wan2.1-VACE-14B">🤖</a></td><td valign="top" style="padding:2px 0;">14B VACE控制与主体参考</td></tr></table> |
|
||||
| Self-Forcing / Causal-Forcing / Flex-Forcing | 视频 | 自回归蒸馏方案,覆盖流式生成、交互式生成与分块注意力 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Self-Forcing</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/gdhe17/Self-Forcing">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/Self-Forcing">🤖</a></td><td valign="top" style="padding:2px 0;">官方发布的蒸馏权重,配合Wan2.1-T2V使用;也可由`scripts/wan2.1_self_forcing`与`scripts/wan2.1_causal_forcing`自行训练得到;Flex-Forcing(分块因果/双向注意力)权重由`scripts/wan2.1_flex_forcing`训练产出</td></tr></table> |
|
||||
| TurboWan / TurboDiffusion | 视频 | TurboDiffusion方案公开发布的少步蒸馏权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.1-T2V-1.3B-480P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.1-T2V-1.3B-480P">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B文生视频蒸馏权重,官方以.pth发布,仓库另含量化版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">TurboWan2.2-I2V-A14B-720P</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TurboDiffusion/TurboWan2.2-I2V-A14B-720P">🤖</a></td><td valign="top" style="padding:2px 0;">14B图生视频蒸馏权重,仓库含low/high两档噪声模型(另含量化版),放入Personalized_Model后按预测脚本的transformer_path/transformer_high_path引用</td></tr></table> |
|
||||
| CogVideoX-Fun V1.5 | 视频 | V1.5官方权重,多分辨率(512/768/1024)、85帧8fps,覆盖图生视频、奖励对齐 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5b图生视频权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.5-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.5-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.5-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
|
||||
| CogVideoX-Fun V1.1 | 视频 | V1.1官方权重,多分辨率(512/768/1024/1280)、49帧8fps,覆盖图生视频、姿态控制、控制生成、奖励对齐 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">2b图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">5b图生视频,添加Noise,运动幅度大于V1.0</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">2b姿态控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Pose</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Pose">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Pose">🤖</a></td><td valign="top" style="padding:2px 0;">5b姿态控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-2b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-2b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-2b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">2b控制生成,支持Canny、Depth、Pose、MLSD</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-5b-Control</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-5b-Control">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-5b-Control">🤖</a></td><td valign="top" style="padding:2px 0;">5b控制生成,支持Canny、Depth、Pose、MLSD</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-V1.1-Reward-LoRAs</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-V1.1-Reward-LoRAs">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-V1.1-Reward-LoRAs">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播训练的对齐LoRA</td></tr></table> |
|
||||
| CogVideoX-Fun V1.0 | 视频 | 旧版权重,仍以49帧8fps训练,已被V1.1/V1.5取代 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-2b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-2b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的2b图生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">CogVideoX-Fun-5b-InP</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/CogVideoX-Fun-5b-InP">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/CogVideoX-Fun-5b-InP">🤖</a></td><td valign="top" style="padding:2px 0;">V1.0的5b图生视频</td></tr></table> |
|
||||
| HunyuanVideo | 视频 | 官方diffusers格式权重,本项目直接支持预测与LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo">🤖</a></td><td valign="top" style="padding:2px 0;">文生视频</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">HunyuanVideo-I2V</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Tencent-Hunyuan/HunyuanVideo-I2V">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频</td></tr></table> |
|
||||
| MiniMax-H3 | 视频 | 官方视频生成权重与本项目训练的ControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MiniMax/MiniMax-H3">🤖</a></td><td valign="top" style="padding:2px 0;">官方基线权重,仓库含多种精度与组件,可按需下载</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet,支持多种控制条件与轨迹控制</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">MiniMax-H3-Fun-Controlnet-Union-2.0</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/MiniMax-H3-Fun-Controlnet-Union-2.0">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet(2.0版),支持多种控制条件、轨迹控制与inpaint权重</td></tr></table> |
|
||||
| TaoMate-H3 | 视频+音频 | 基于MiniMax-H3的官方流式音视频生成适配器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">TaoMate-H3-Adapter</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TaoLiveAIGC/TaoMate-H3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TaoLiveAIGC/TaoMate-H3">🤖</a></td><td valign="top" style="padding:2px 0;">官方rank 128适配器(step-3000 EMA),内置3步蒸馏采样调度,支持流式语音驱动生成;需搭配MiniMax-H3基座权重使用</td></tr></table> |
|
||||
| LTX-2 | 视频+音频 | 官方DiT音视频联合生成权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Lightricks/LTX-2">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Lightricks/LTX-2">🤖</a></td><td valign="top" style="padding:2px 0;">音视频联合生成的官方权重,仓库含多种精度</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LTX-2.3-Diffusers</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/dg845/LTX-2.3-Diffusers">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">2.3版本需使用社区转换的diffusers格式权重,官方原始权重见<a href="https://huggingface.co/Lightricks/LTX-2.3">Lightricks/LTX-2.3</a></td></tr></table> |
|
||||
| LongCat-Video | 视频 | 官方长视频生成权重,支持LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video">🤖</a></td><td valign="top" style="padding:2px 0;">文/图生长视频基线</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">LongCat-Video-Avatar</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/meituan-longcat/LongCat-Video-Avatar">🤖</a></td><td valign="top" style="padding:2px 0;">数字人权重</td></tr></table> |
|
||||
| FantasyTalking | 音频驱动视频 | 音频条件增量权重,需搭配基础视频权重与音频编码器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FantasyTalking</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/acvlab/FantasyTalking">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/amap_cvlab/FantasyTalking">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-720P使用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">wav2vec2-base-960h</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/facebook/wav2vec2-base-960h">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/AI-ModelScope/wav2vec2-base-960h">🤖</a></td><td valign="top" style="padding:2px 0;">音频编码器,放入基础权重目录并命名为audio_encoder</td></tr></table> |
|
||||
| InfiniteTalk | 音频驱动视频 | 音频条件增量权重,需搭配基础视频权重与音频编码器 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">InfiniteTalk</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MeiGen-AI/InfiniteTalk">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MeiGen-AI/InfiniteTalk">🤖</a></td><td valign="top" style="padding:2px 0;">需搭配Wan2.1-I2V-14B-480P使用,仓库含多个版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">chinese-wav2vec2-base</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/TencentGameMate/chinese-wav2vec2-base">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base">🤖</a></td><td valign="top" style="padding:2px 0;">中文音频编码器</td></tr></table> |
|
||||
| FlashHead | 音频驱动视频 | 官方头部动作数字人权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">SoulX-FlashHead-1_3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Soul-AILab/SoulX-FlashHead-1_3B">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Soul-AILab/SoulX-FlashHead-1_3B">🤖</a></td><td valign="top" style="padding:2px 0;">语音驱动头部数字人,同样需要wav2vec音频编码器</td></tr></table> |
|
||||
| MOVA | 视频+音频 | 官方MOVA权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">MOVA-360p</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/OpenMOSS-Team/MOVA-360p">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/OpenMOSS/MOVA-360p">🤖</a></td><td valign="top" style="padding:2px 0;">图生视频与音视频联合生成</td></tr></table> |
|
||||
| LingBot | 视频 | 相机可控世界模型,目录结构与Wan2.2-I2V-A14B一致 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-base-cam</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-base-cam">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-base-cam">🤖</a></td><td valign="top" style="padding:2px 0;">相机控制基线权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-rewriter-lora</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-rewriter-lora">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-rewriter-lora">🤖</a></td><td valign="top" style="padding:2px 0;">rewriter LoRA,搭配Qwen3.6-27B生成结构化caption</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-dense-1.3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-dense-1.3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-dense-1.3b">🤖</a></td><td valign="top" style="padding:2px 0;">1.3B稠密版视频生成权重,1-2卡即可训练</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-video-moe-30b-a3b</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-video-moe-30b-a3b">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-video-moe-30b-a3b">🤖</a></td><td valign="top" style="padding:2px 0;">30B MoE(3B激活)视频生成权重,训练建议8×80GB及以上</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">lingbot-world-fast</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Robbyant/lingbot-world-fast">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Robbyant/lingbot-world-fast">🤖</a></td><td valign="top" style="padding:2px 0;">蒸馏少步世界模型(transformer共16个分片),VAE/T5复用lingbot-world-base-cam,推理需使用Flow_Unipc采样器</td></tr></table> |
|
||||
| Phantom | 视频 | 多主体参考生视频的增量权重,基于Wan2.1-T2V | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-1.3B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">1.3B版,官方以.pth发布,放入Personalized_Model后按预测脚本的transformer_path引用</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Phantom-Wan-14B</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/bytedance-research/Phantom">🤗</a></td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">14B版,官方以分片safetensors发布</td></tr></table> |
|
||||
| Qwen-Image | 图片 | 官方文生图与图像编辑权重,支持基线与LoRA训练 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image">🤖</a></td><td valign="top" style="padding:2px 0;">文生图基础权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2512">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2512">🤖</a></td><td valign="top" style="padding:2px 0;">文生图更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Edit-2509</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Edit-2509">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Edit-2509">🤖</a></td><td valign="top" style="padding:2px 0;">图像编辑更新版本</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-Layered</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-Layered">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-Layered">🤖</a></td><td valign="top" style="padding:2px 0;">图像图层分解权重,可将图像拆分为多个可编辑的RGBA图层</td></tr></table> |
|
||||
| Qwen-Image-2.1 | 图片 | 官方新一代文生图权重,单流block-causal结构,支持前缀KV cache | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen-Image-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">单流block-causal结构,支持全参数训练;前缀KV cache可加速推理</td></tr></table> |
|
||||
| Qwen-Image ControlNet | 图片 | 图片控制生成,支持Canny、Depth、Pose、MLSD、Scribble | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-2512-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Qwen-Image-2512-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Qwen-Image-2512-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目训练的ControlNet</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen-Image-ControlNet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/InstantX/Qwen-Image-ControlNet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/InstantX/Qwen-Image-ControlNet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">InstantX提供的同类型ControlNet</td></tr></table> |
|
||||
| Z-Image | 图片 | 官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image">🤖</a></td><td valign="top" style="padding:2px 0;">基础版</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Tongyi-MAI/Z-Image-Turbo">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo">🤖</a></td><td valign="top" style="padding:2px 0;">加速版</td></tr></table> |
|
||||
| Z-Image-Fun | 图片 | 本项目在Z-Image上训练的ControlNet与蒸馏LoRA,控制条件支持Canny、Depth、Pose、MLSD、Scribble、Gray | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">基于基础版的ControlNet,2.1版层数更多、训练更充分</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">基于Turbo的ControlNet</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Turbo-Fun-Controlnet-Union-2.1</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Turbo-Fun-Controlnet-Union-2.1">🤖</a></td><td valign="top" style="padding:2px 0;">基于Turbo的2.1版ControlNet,仓库含多精度文件</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Z-Image-Fun-Lora-Distill</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/Z-Image-Fun-Lora-Distill">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/Z-Image-Fun-Lora-Distill">🤖</a></td><td valign="top" style="padding:2px 0;">同时蒸馏步数与CFG,推理仅需8步</td></tr></table> |
|
||||
| Flux | 图片 | 官方FLUX.1/FLUX.2权重与本项目训练的ControlNet | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.1-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.1-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.1-dev">🤖</a></td><td valign="top" style="padding:2px 0;">文生图与图像编辑</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/black-forest-labs/FLUX.2-dev">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://www.modelscope.cn/models/black-forest-labs/FLUX.2-dev">🤖</a></td><td valign="top" style="padding:2px 0;">第二代官方权重</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">FLUX.2-dev-Fun-Controlnet-Union</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/alibaba-pai/FLUX.2-dev-Fun-Controlnet-Union">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PAI/FLUX.2-dev-Fun-Controlnet-Union">🤖</a></td><td valign="top" style="padding:2px 0;">本项目为FLUX.2-dev训练的ControlNet,支持Canny、Depth、Pose、MLSD等</td></tr></table> |
|
||||
| ERNIE-Image | 图片 | 百度官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">ERNIE-Image</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/baidu/ERNIE-Image">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/PaddlePaddle/ERNIE-Image">🤖</a></td><td valign="top" style="padding:2px 0;">单流DiT文生图,Hugging Face为baidu组织、ModelScope为PaddlePaddle组织</td></tr></table> |
|
||||
| Lens | 图片 | 微软官方文生图权重 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">Lens</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/microsoft/Lens">🤖</a></td><td valign="top" style="padding:2px 0;">3.8B文生图,仓库内含GPT-OSS文本编码器;Hugging Face侧无公开下载仓库,请从ModelScope获取</td></tr></table> |
|
||||
| 辅助模型 | - | 非生成模型,服务于奖励对齐、数据打标与快速解码 | <table style="width:100%;border-collapse:collapse;"><tr><td valign="top" style="padding:2px 8px 2px 0;">HPSv3</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/MizzenAI/HPSv3">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/MizzenAI/HPSv3">🤖</a></td><td valign="top" style="padding:2px 0;">奖励反向传播使用的打分模型</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">Qwen2-VL-7B-Instruct</td><td valign="top" style="padding:2px 8px;"><a href="https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct">🤗</a></td><td valign="top" style="padding:2px 8px;"><a href="https://modelscope.cn/models/Qwen/Qwen2-VL-7B-Instruct">🤖</a></td><td valign="top" style="padding:2px 0;">视频打标流程使用的多模态编码器</td></tr><tr><td valign="top" style="padding:2px 8px 2px 0;">taew2_1 / taew2_2</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 8px;">-</td><td valign="top" style="padding:2px 0;">Tiny AutoEncoder(约20MB),与Wan2.1/Wan2.2 VAE共享同一latent空间,解码速度约为完整VAE的100倍,用于快速预览与低显存生成;权重来自<a href="https://github.com/madebyollin/taehv">madebyollin/taehv</a></td></tr></table> |
|
||||
|
||||
> 补充说明:
|
||||
> - 音频驱动与参考类模型(FantasyTalking、InfiniteTalk、Phantom、TaoMate-H3)本身只是增量权重,必须同时下载表中对应的基础视频权重与音频编码器。
|
||||
> - TurboDiffusion方案已公开发布TurboWan系列蒸馏权重(见上表);Flex-Forcing、PDD等其余蒸馏方案没有公开发布的权重,按`scripts/{model_name}/README_TRAIN*.md`训练后即可得到,可直接填入预测文件中的`transformer_path`。
|
||||
> - 权重名与`models/Diffusion_Transformer/`下的文件夹名一一对应;同一系列内各权重互不通用,需按预测任务选择,若某个权重未在此列出,说明它由本项目训练产出或需从上游官方仓库获取。
|
||||
|
||||
# 四、视频作品
|
||||
|
||||
图生视频:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d6a46051-8fe6-4174-be12-95ee52c96298" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8572c656-8548-4b1f-9ec8-8107c6236cb1" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/d3411c95-483d-4e30-bc72-483c2b288918" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/b2f5addc-06bd-49d9-b925-973090a32800" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
通用控制视频 + 参考图像:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
参考图像
|
||||
</td>
|
||||
<td>
|
||||
控制视频
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-14B-Control
|
||||
</td>
|
||||
<td>
|
||||
Wan2.1-Fun-V1.1-1.3B-Control
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<image src="https://github.com/user-attachments/assets/221f2879-3b1b-4fbd-84f9-c3e0b0b3533e" width="100%" controls preload loop></image>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f361af34-b3b3-4be4-9d03-cd478cb3dfc5" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85e2f00b-6ef0-4922-90ab-4364afb2c93d" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/1f3fe763-2754-4215-bc9a-ae804950d4b3" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
通用控制视频(Canny、Pose、Depth 等)与轨迹控制:
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/f35602c4-9f0a-4105-9762-1e3a88abbac6" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/8b0f0e87-f1be-4915-bb35-2d53c852333e" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/972012c1-772b-427a-bce6-ba8b39edcfad" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table border="0" style="width: 100%; text-align: left; margin-top: 20px;">
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/ce62d0bd-82c0-4d7b-9c49-7e0e4b605745" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/89dfbffb-c4a6-4821-bcef-8b1489a3ca00" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/72a43e33-854f-4349-861b-c959510d1a84" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/bb0ce13d-dee0-4049-9eec-c92f3ebc1358" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/7840c333-7bec-4582-ba63-20a39e1139c4" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
<td>
|
||||
<video src="https://github.com/user-attachments/assets/85147d30-ae09-4f36-a077-2167f7a578c0" width="100%" controls preload loop></video>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
|
||||
# 五、参考文献
|
||||
本节列出[已支持的模型](#三已支持的模型)中各模型系列的官方仓库,以及本项目在实现与流程中参考的代码来源,感谢这些开源工作。
|
||||
# 未来计划
|
||||
- 支持CogVideoX-5b。
|
||||
|
||||
# 参考文献
|
||||
- CogVideo: https://github.com/THUDM/CogVideo/
|
||||
- EasyAnimate: https://github.com/aigc-apps/EasyAnimate
|
||||
- Wan2.1: https://github.com/Wan-Video/Wan2.1/
|
||||
- Wan2.2: https://github.com/Wan-Video/Wan2.2/
|
||||
- HunyuanVideo: https://github.com/Tencent-Hunyuan/HunyuanVideo
|
||||
- HunyuanVideo-I2V: https://github.com/Tencent-Hunyuan/HunyuanVideo-I2V
|
||||
- MiniMax-H3: https://github.com/MiniMax-AI/MiniMax-H3
|
||||
- LTX-Video: https://github.com/Lightricks/LTX-Video
|
||||
- LTX-2: https://github.com/Lightricks/LTX-2
|
||||
- LongCat-Video: https://github.com/meituan-longcat/LongCat-Video
|
||||
- FantasyTalking: https://github.com/Fantasy-AMAP/fantasy-talking
|
||||
- InfiniteTalk: https://github.com/MeiGen-AI/InfiniteTalk
|
||||
- FlashHead: https://github.com/Soul-AILab/SoulX-FlashHead
|
||||
- MOVA: https://github.com/OpenMOSS/MOVA
|
||||
- LingBot-Video: https://github.com/Robbyant/lingbot-video
|
||||
- LingBot-World: https://github.com/Robbyant/lingbot-world
|
||||
- Phantom: https://github.com/Phantom-video/Phantom
|
||||
- Qwen-Image: https://github.com/QwenLM/Qwen-Image
|
||||
- Z-Image: https://github.com/Tongyi-MAI/Z-Image
|
||||
- Flux: https://github.com/black-forest-labs/flux
|
||||
- Flux2: https://github.com/black-forest-labs/flux2
|
||||
- ERNIE-Image: https://github.com/baidu/ernie-image
|
||||
- Lens: https://www.microsoft.com/en-us/research/publication/lens-rethinking-training-efficiency-for-foundational-text-to-image-models/
|
||||
- VACE: https://github.com/ali-vilab/VACE
|
||||
- CameraCtrl: https://github.com/hehao13/CameraCtrl
|
||||
- ComfyUI-CameraCtrl-Wrapper: https://github.com/chaojie/ComfyUI-CameraCtrl-Wrapper
|
||||
- DWPose: https://github.com/IDEA-Research/DWPose
|
||||
- MiDaS: https://github.com/isl-org/MiDaS
|
||||
- Self-Forcing: https://github.com/guandeh17/Self-Forcing
|
||||
- Causal-Forcing: https://github.com/thu-ml/Causal-Forcing
|
||||
- TurboDiffusion: https://github.com/thu-ml/TurboDiffusion
|
||||
- TAEHV: https://github.com/madebyollin/taehv
|
||||
- HPS v2: https://github.com/tgxs002/HPSv2
|
||||
- HPSv3: https://github.com/MizzenAI/HPSv3
|
||||
- MPS: https://github.com/Kwai-Kolors/MPS
|
||||
- Qwen2-VL: https://github.com/QwenLM/Qwen2-VL
|
||||
- AnimateDiff: https://github.com/guoyww/AnimateDiff
|
||||
- ComfyUI-KJNodes: https://github.com/kijai/ComfyUI-KJNodes
|
||||
- ComfyUI-EasyAnimateWrapper: https://github.com/kijai/ComfyUI-EasyAnimateWrapper
|
||||
- Diffusers: https://github.com/huggingface/diffusers
|
||||
|
||||
# 六、引用
|
||||
|
||||
如果您在研究或项目中使用了 VideoX-Fun,请按以下格式引用:
|
||||
|
||||
```bibtex
|
||||
@misc{aigc_apps_VideoX_Fun_2026,
|
||||
author = {aigc-apps},
|
||||
title = {VideoX-Fun: A Video Generation Pipeline for Diffusion Transformer},
|
||||
year = {2026},
|
||||
publisher = {GitHub},
|
||||
url = {https://github.com/aigc-apps/VideoX-Fun}
|
||||
}
|
||||
```
|
||||
|
||||
# 七、限制与风险
|
||||
|
||||
- 生成的视频可能存在伪影或质量问题,尤其在复杂场景中。
|
||||
- 模型在处理精细细节、文字渲染或特定艺术风格时可能有困难。
|
||||
- 性能因输入提示词质量、分辨率等参数而异。
|
||||
- 该技术可能被滥用于创建误导性内容(如深度伪造)。用户需对道德使用负责。
|
||||
- 模型可能反映训练数据中存在的偏见。
|
||||
- 用户在使用真人图片或视频时应尊重隐私和版权。
|
||||
|
||||
我们鼓励负责任地使用该技术,并建议在生产环境中实施安全措施。
|
||||
|
||||
# 八、许可证
|
||||
# 许可证
|
||||
本项目采用 [Apache License (Version 2.0)](https://github.com/modelscope/modelscope/blob/master/LICENSE).
|
||||
|
||||
CogVideoX-2B 模型 (包括其对应的Transformers模块,VAE模块) 根据 [Apache 2.0 协议](LICENSE) 许可证发布。
|
||||
|
||||
CogVideoX-5B 模型(Transformer 模块)在[CogVideoX许可证](https://huggingface.co/THUDM/CogVideoX-5b/blob/main/LICENSE)下发布.
|
||||
CogVideoX-2B 模型 (包括其对应的Transformers模块,VAE模块) 根据 [Apache 2.0 协议](LICENSE) 许可证发布。
|
||||
@@ -0,0 +1,46 @@
|
||||
import time
|
||||
import torch
|
||||
|
||||
from cogvideox.api.api import infer_forward_api, update_diffusion_transformer_api, update_edition_api
|
||||
from cogvideox.ui.ui import ui_modelscope, ui_eas, ui
|
||||
|
||||
if __name__ == "__main__":
|
||||
# Choose the ui mode
|
||||
ui_mode = "normal"
|
||||
|
||||
# Low gpu memory mode, this is used when the GPU memory is under 16GB
|
||||
low_gpu_memory_mode = False
|
||||
# Use torch.float16 if GPU does not support torch.bfloat16
|
||||
# ome graphics cards, such as v100, 2080ti, do not support torch.bfloat16
|
||||
weight_dtype = torch.bfloat16
|
||||
|
||||
# Server ip
|
||||
server_name = "0.0.0.0"
|
||||
server_port = 7860
|
||||
|
||||
# Params below is used when ui_mode = "modelscope"
|
||||
model_name = "models/Diffusion_Transformer/CogVideoX-Fun-2b-InP"
|
||||
savedir_sample = "samples"
|
||||
|
||||
if ui_mode == "modelscope":
|
||||
demo, controller = ui_modelscope(model_name, savedir_sample, low_gpu_memory_mode, weight_dtype)
|
||||
elif ui_mode == "eas":
|
||||
demo, controller = ui_eas(model_name, savedir_sample)
|
||||
else:
|
||||
demo, controller = ui(low_gpu_memory_mode, weight_dtype)
|
||||
|
||||
# launch gradio
|
||||
app, _, _ = demo.queue(status_update_rate=1).launch(
|
||||
server_name=server_name,
|
||||
server_port=server_port,
|
||||
prevent_thread_lock=True
|
||||
)
|
||||
|
||||
# launch api
|
||||
infer_forward_api(None, app, controller)
|
||||
update_diffusion_transformer_api(None, app, controller)
|
||||
update_edition_api(None, app, controller)
|
||||
|
||||
# not close the python
|
||||
while True:
|
||||
time.sleep(5)
|
||||
|
Before Width: | Height: | Size: 128 KiB |
|
Before Width: | Height: | Size: 477 KiB |
|
Before Width: | Height: | Size: 569 KiB |
|
Before Width: | Height: | Size: 349 KiB |
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9999164298554373 -0.012928004685808521 0.0 0.0 0.012928004685808521 0.9999164298554373 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9996657333896874 -0.025853848581176044 0.0 0.0 0.025853848581176044 0.9996657333896874 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.99924795250423 -0.03877537125681671 0.0 0.0 0.03877537125681671 0.99924795250423 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9986631570270832 -0.051690413005694553 0.0 0.0 0.051690413005694553 0.9986631570270832 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.997911444701132 -0.06459681520399763 0.0 0.0 0.06459681520399763 0.997911444701132 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.996992941167792 -0.07749242067193093 0.0 0.0 0.07749242067193093 0.996992941167792 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9959077999460093 -0.0903750740342681 0.0 0.0 0.0903750740342681 0.9959077999460093 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9946562024066015 -0.10324262208060146 0.0 0.0 0.10324262208060146 0.9946562024066015 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.993238357741943 -0.11609291412523022 0.0 0.0 0.11609291412523022 0.993238357741943 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9916545029310012 -0.12892380236662665 0.0 0.0 0.12892380236662665 0.9916545029310012 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.989904902699727 -0.14173314224642042 0.0 0.0 0.14173314224642042 0.989904902699727 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.987989849476809 -0.15451879280784048 0.0 0.0 0.15451879280784048 0.987989849476809 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9859096633447965 -0.16727861705355532 0.0 0.0 0.16727861705355532 0.9859096633447965 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9836646919866011 -0.18001048230285133 0.0 0.0 0.18001048230285133 0.9836646919866011 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9812553106273847 -0.19271226054808965 0.0 0.0 0.19271226054808965 0.9812553106273847 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9786819219718442 -0.20538182881038197 0.0 0.0 0.20538182881038197 0.9786819219718442 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9759449561369036 -0.21801706949442584 0.0 0.0 0.21801706949442584 0.9759449561369036 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9730448705798238 -0.23061587074244014 0.0 0.0 0.23061587074244014 0.9730448705798238 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9699821500217435 -0.24317612678714165 0.0 0.0 0.24317612678714165 0.9699821500217435 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.966757306366662 -0.25569573830370357 0.0 0.0 0.25569573830370357 0.966757306366662 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9633708786158803 -0.26817261276063736 0.0 0.0 0.26817261276063736 0.9633708786158803 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9598234327779119 -0.2806046647695387 0.0 0.0 0.2806046647695387 0.9598234327779119 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9561155617738797 -0.29298981643364064 0.0 0.0 0.29298981643364064 0.9561155617738797 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9522478853384153 -0.3053259976951131 0.0 0.0 0.3053259976951131 0.9522478853384153 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9482210499160765 -0.3176111466810532 0.0 0.0 0.3176111466810532 0.9482210499160765 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9440357285533 -0.3298432100481077 0.0 0.0 0.3298432100481077 0.9440357285533 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9396926207859084 -0.34202014332566866 0.0 0.0 0.34202014332566866 0.9396926207859084 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9351924525221897 -0.35413991125758754 0.0 0.0 0.35413991125758754 0.9351924525221897 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9305359759215686 -0.36620048814234796 0.0 0.0 0.36620048814234796 0.9305359759215686 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9257239692688904 -0.3781998581716424 0.0 0.0 0.3781998581716424 0.9257239692688904 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9207572368443384 -0.3901360157672949 0.0 0.0 0.3901360157672949 0.9207572368443384 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9156366087890059 -0.40200696591647384 0.0 0.0 0.40200696591647384 0.9156366087890059 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9103629409661467 -0.4138107245051391 0.0 0.0 0.4138107245051391 0.9103629409661467 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9049371148181253 -0.4255453186496674 0.0 0.0 0.4255453186496674 0.9049371148181253 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8993600372190931 -0.4372087870266006 0.0 0.0 0.4372087870266006 0.8993600372190931 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8936326403234123 -0.4487991802004621 0.0 0.0 0.4487991802004621 0.8936326403234123 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.887755881409856 -0.4603145609495856 0.0 0.0 0.4603145609495856 0.887755881409856 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.881730742721608 -0.4717530045899035 0.0 0.0 0.4717530045899035 0.881730742721608 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8755582313020909 -0.4831125992966384 0.0 0.0 0.4831125992966384 0.8755582313020909 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8692393788266478 -0.4943914464238468 0.0 0.0 0.4943914464238468 0.8692393788266478 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8627752414301085 -0.5055876608217589 0.0 0.0 0.5055876608217589 0.8627752414301085 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8561668995302665 -0.5166993711518628 0.0 0.0 0.5166993711518628 0.8561668995302665 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8494154576472973 -0.5277247201996818 0.0 0.0 0.5277247201996818 0.8494154576472973 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8425220442191496 -0.5386618651851877 0.0 0.0 0.5386618651851877 0.8425220442191496 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8354878114129365 -0.549508978070806 0.0 0.0 0.549508978070806 0.8354878114129365 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.828313934932363 -0.5602642458669524 0.0 0.0 0.5602642458669524 0.828313934932363 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8210016138212185 -0.5709258709350582 0.0 0.0 0.5709258709350582 0.8210016138212185 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8135520702629676 -0.5814920712880266 0.0 0.0 0.5814920712880266 0.8135520702629676 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8059665493764744 -0.5919610808880759 0.0 0.0 0.5919610808880759 0.8059665493764744 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.798246319007893 -0.6023311499419145 0.0 0.0 0.6023311499419145 0.798246319007893 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7903926695187593 -0.6126005451932028 0.0 0.0 0.6126005451932028 0.7903926695187593 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7824069135703198 -0.622767550212249 0.0 0.0 0.622767550212249 0.7824069135703198 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7742903859041323 -0.632830465682895 0.0 0.0 0.632830465682895 0.7742903859041323 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7660444431189781 -0.6427876096865393 0.0 0.0 0.6427876096865393 0.7660444431189781 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7576704634441179 -0.6526373179832544 0.0 0.0 0.6526373179832544 0.7576704634441179 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.749169846508936 -0.6623779442899478 0.0 0.0 0.6623779442899478 0.749169846508936 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7405440131090046 -0.6720078605555224 0.0 0.0 0.6720078605555224 0.7405440131090046 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7317944049686121 -0.6815254572329892 0.0 0.0 0.6815254572329892 0.7317944049686121 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7229224844997929 -0.6909291435484877 0.0 0.0 0.6909291435484877 0.7229224844997929 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7139297345578991 -0.7002173477671684 0.0 0.0 0.7002173477671684 0.7139297345578991 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7048176581937561 -0.7093885174558928 0.0 0.0 0.7093885174558928 0.7048176581937561 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.695587778402442 -0.7184411197427074 0.0 0.0 0.7184411197427074 0.695587778402442 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6862416378687336 -0.7273736415730486 0.0 0.0 0.7273736415730486 0.6862416378687336 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6767807987092621 -0.7361845899626351 0.0 0.0 0.7361845899626351 0.6767807987092621 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6672068422114197 -0.7448724922470058 0.0 0.0 0.7448724922470058 0.6672068422114197 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6575213685690637 -0.7534358963276606 0.0 0.0 0.7534358963276606 0.6575213685690637 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.647725996615059 -0.761873370914766 0.0 0.0 0.761873370914766 0.647725996615059 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6378223635507061 -0.7701835057663796 0.0 0.0 0.7701835057663796 0.6378223635507061 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6278121246720987 -0.7783649119241599 0.0 0.0 0.7783649119241599 0.6278121246720987 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6176969530934572 -0.7864162219455162 0.0 0.0 0.7864162219455162 0.6176969530934572 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6074785394674836 -0.7943360901321637 0.0 0.0 0.7943360901321637 0.6074785394674836 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5971585917027863 -0.8021231927550437 0.0 0.0 0.8021231927550437 0.5971585917027863 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5867388346784178 -0.8097762282755726 0.0 0.0 0.8097762282755726 0.5867388346784178 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5762210099555805 -0.8172939175631805 0.0 0.0 0.8172939175631805 0.5762210099555805 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5656068754865388 -0.8246750041091067 0.0 0.0 0.8246750041091067 0.5656068754865388 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.554898205320797 -0.8319182542364115 0.0 0.0 0.8319182542364115 0.554898205320797 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5440967893085827 -0.8390224573061746 0.0 0.0 0.8390224573061746 0.5440967893085827 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5332044328016914 -0.845986425919841 0.0 0.0 0.845986425919841 0.5332044328016914 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5222229563517384 -0.852808996117683 0.0 0.0 0.852808996117683 0.5222229563517384 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5111541954058733 -0.8594890275733451 0.0 0.0 0.8594890275733451 0.5111541954058733 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9999164298554373 0.012928004685808521 0.0 0.0 -0.012928004685808521 0.9999164298554373 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9996657333896874 0.025853848581176044 0.0 0.0 -0.025853848581176044 0.9996657333896874 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.99924795250423 0.03877537125681671 0.0 0.0 -0.03877537125681671 0.99924795250423 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9986631570270832 0.051690413005694553 0.0 0.0 -0.051690413005694553 0.9986631570270832 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.997911444701132 0.06459681520399763 0.0 0.0 -0.06459681520399763 0.997911444701132 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.996992941167792 0.07749242067193093 0.0 0.0 -0.07749242067193093 0.996992941167792 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9959077999460093 0.0903750740342681 0.0 0.0 -0.0903750740342681 0.9959077999460093 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9946562024066015 0.10324262208060146 0.0 0.0 -0.10324262208060146 0.9946562024066015 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.993238357741943 0.11609291412523022 0.0 0.0 -0.11609291412523022 0.993238357741943 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9916545029310012 0.12892380236662665 0.0 0.0 -0.12892380236662665 0.9916545029310012 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.989904902699727 0.14173314224642042 0.0 0.0 -0.14173314224642042 0.989904902699727 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.987989849476809 0.15451879280784048 0.0 0.0 -0.15451879280784048 0.987989849476809 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9859096633447965 0.16727861705355532 0.0 0.0 -0.16727861705355532 0.9859096633447965 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9836646919866011 0.18001048230285133 0.0 0.0 -0.18001048230285133 0.9836646919866011 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9812553106273847 0.19271226054808965 0.0 0.0 -0.19271226054808965 0.9812553106273847 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9786819219718442 0.20538182881038197 0.0 0.0 -0.20538182881038197 0.9786819219718442 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9759449561369036 0.21801706949442584 0.0 0.0 -0.21801706949442584 0.9759449561369036 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9730448705798238 0.23061587074244014 0.0 0.0 -0.23061587074244014 0.9730448705798238 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9699821500217435 0.24317612678714165 0.0 0.0 -0.24317612678714165 0.9699821500217435 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.966757306366662 0.25569573830370357 0.0 0.0 -0.25569573830370357 0.966757306366662 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9633708786158803 0.26817261276063736 0.0 0.0 -0.26817261276063736 0.9633708786158803 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9598234327779119 0.2806046647695387 0.0 0.0 -0.2806046647695387 0.9598234327779119 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9561155617738797 0.29298981643364064 0.0 0.0 -0.29298981643364064 0.9561155617738797 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9522478853384153 0.3053259976951131 0.0 0.0 -0.3053259976951131 0.9522478853384153 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9482210499160765 0.3176111466810532 0.0 0.0 -0.3176111466810532 0.9482210499160765 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9440357285533 0.3298432100481077 0.0 0.0 -0.3298432100481077 0.9440357285533 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9396926207859084 0.34202014332566866 0.0 0.0 -0.34202014332566866 0.9396926207859084 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9351924525221897 0.35413991125758754 0.0 0.0 -0.35413991125758754 0.9351924525221897 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9305359759215686 0.36620048814234796 0.0 0.0 -0.36620048814234796 0.9305359759215686 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9257239692688904 0.3781998581716424 0.0 0.0 -0.3781998581716424 0.9257239692688904 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9207572368443384 0.3901360157672949 0.0 0.0 -0.3901360157672949 0.9207572368443384 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9156366087890059 0.40200696591647384 0.0 0.0 -0.40200696591647384 0.9156366087890059 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9103629409661467 0.4138107245051391 0.0 0.0 -0.4138107245051391 0.9103629409661467 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.9049371148181253 0.4255453186496674 0.0 0.0 -0.4255453186496674 0.9049371148181253 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8993600372190931 0.4372087870266006 0.0 0.0 -0.4372087870266006 0.8993600372190931 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8936326403234123 0.4487991802004621 0.0 0.0 -0.4487991802004621 0.8936326403234123 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.887755881409856 0.4603145609495856 0.0 0.0 -0.4603145609495856 0.887755881409856 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.881730742721608 0.4717530045899035 0.0 0.0 -0.4717530045899035 0.881730742721608 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8755582313020909 0.4831125992966384 0.0 0.0 -0.4831125992966384 0.8755582313020909 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8692393788266478 0.4943914464238468 0.0 0.0 -0.4943914464238468 0.8692393788266478 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8627752414301085 0.5055876608217589 0.0 0.0 -0.5055876608217589 0.8627752414301085 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8561668995302665 0.5166993711518628 0.0 0.0 -0.5166993711518628 0.8561668995302665 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8494154576472973 0.5277247201996818 0.0 0.0 -0.5277247201996818 0.8494154576472973 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8425220442191496 0.5386618651851877 0.0 0.0 -0.5386618651851877 0.8425220442191496 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8354878114129365 0.549508978070806 0.0 0.0 -0.549508978070806 0.8354878114129365 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.828313934932363 0.5602642458669524 0.0 0.0 -0.5602642458669524 0.828313934932363 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8210016138212185 0.5709258709350582 0.0 0.0 -0.5709258709350582 0.8210016138212185 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8135520702629676 0.5814920712880266 0.0 0.0 -0.5814920712880266 0.8135520702629676 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.8059665493764744 0.5919610808880759 0.0 0.0 -0.5919610808880759 0.8059665493764744 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.798246319007893 0.6023311499419145 0.0 0.0 -0.6023311499419145 0.798246319007893 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7903926695187593 0.6126005451932028 0.0 0.0 -0.6126005451932028 0.7903926695187593 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7824069135703198 0.622767550212249 0.0 0.0 -0.622767550212249 0.7824069135703198 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7742903859041323 0.632830465682895 0.0 0.0 -0.632830465682895 0.7742903859041323 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7660444431189781 0.6427876096865393 0.0 0.0 -0.6427876096865393 0.7660444431189781 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7576704634441179 0.6526373179832544 0.0 0.0 -0.6526373179832544 0.7576704634441179 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.749169846508936 0.6623779442899478 0.0 0.0 -0.6623779442899478 0.749169846508936 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7405440131090046 0.6720078605555224 0.0 0.0 -0.6720078605555224 0.7405440131090046 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7317944049686121 0.6815254572329892 0.0 0.0 -0.6815254572329892 0.7317944049686121 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7229224844997929 0.6909291435484877 0.0 0.0 -0.6909291435484877 0.7229224844997929 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7139297345578991 0.7002173477671684 0.0 0.0 -0.7002173477671684 0.7139297345578991 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.7048176581937561 0.7093885174558928 0.0 0.0 -0.7093885174558928 0.7048176581937561 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.695587778402442 0.7184411197427074 0.0 0.0 -0.7184411197427074 0.695587778402442 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6862416378687336 0.7273736415730486 0.0 0.0 -0.7273736415730486 0.6862416378687336 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6767807987092621 0.7361845899626351 0.0 0.0 -0.7361845899626351 0.6767807987092621 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6672068422114197 0.7448724922470058 0.0 0.0 -0.7448724922470058 0.6672068422114197 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6575213685690637 0.7534358963276606 0.0 0.0 -0.7534358963276606 0.6575213685690637 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.647725996615059 0.761873370914766 0.0 0.0 -0.761873370914766 0.647725996615059 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6378223635507061 0.7701835057663796 0.0 0.0 -0.7701835057663796 0.6378223635507061 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6278121246720987 0.7783649119241599 0.0 0.0 -0.7783649119241599 0.6278121246720987 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6176969530934572 0.7864162219455162 0.0 0.0 -0.7864162219455162 0.6176969530934572 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.6074785394674836 0.7943360901321637 0.0 0.0 -0.7943360901321637 0.6074785394674836 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5971585917027863 0.8021231927550437 0.0 0.0 -0.8021231927550437 0.5971585917027863 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5867388346784178 0.8097762282755726 0.0 0.0 -0.8097762282755726 0.5867388346784178 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5762210099555805 0.8172939175631805 0.0 0.0 -0.8172939175631805 0.5762210099555805 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5656068754865388 0.8246750041091067 0.0 0.0 -0.8246750041091067 0.5656068754865388 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.554898205320797 0.8319182542364115 0.0 0.0 -0.8319182542364115 0.554898205320797 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5440967893085827 0.8390224573061746 0.0 0.0 -0.8390224573061746 0.5440967893085827 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5332044328016914 0.845986425919841 0.0 0.0 -0.845986425919841 0.5332044328016914 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5222229563517384 0.852808996117683 0.0 0.0 -0.852808996117683 0.5222229563517384 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 0.5111541954058733 0.8594890275733451 0.0 0.0 -0.8594890275733451 0.5111541954058733 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.018518518518518517 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.037037037037037035 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.05555555555555555 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.07407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.09259259259259259 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.1111111111111111 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.12962962962962962 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.14814814814814814 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.16666666666666666 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.18518518518518517 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.24074074074074073 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.25925925925925924 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2777777777777778 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.31481481481481477 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.3333333333333333 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.35185185185185186 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.37037037037037035 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.38888888888888884 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.42592592592592593 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4629629629629629 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.48148148148148145 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5185185185185185 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.537037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5555555555555556 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5740740740740741 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5925925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.611111111111111 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6296296296296295 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6481481481481481 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6666666666666666 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6851851851851851 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7592592592592593 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7777777777777777 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8148148148148148 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8333333333333334 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8518518518518519 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8703703703703705 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8888888888888888 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9259259259259258 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9629629629629629 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9814814814814815 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0185185185185186 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0555555555555556 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0925925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1111111111111112 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1296296296296298 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1481481481481481 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1666666666666667 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1851851851851851 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.259259259259259 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2777777777777777 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3148148148148149 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3333333333333333 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3518518518518519 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3703703703703702 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3888888888888888 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.425925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.462962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4814814814814814 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 -0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.018518518518518517 0.0 1.0 0.0 -0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.037037037037037035 0.0 1.0 0.0 -0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.05555555555555555 0.0 1.0 0.0 -0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.07407407407407407 0.0 1.0 0.0 -0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.09259259259259259 0.0 1.0 0.0 -0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.1111111111111111 0.0 1.0 0.0 -0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.12962962962962962 0.0 1.0 0.0 -0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.14814814814814814 0.0 1.0 0.0 -0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.16666666666666666 0.0 1.0 0.0 -0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.18518518518518517 0.0 1.0 0.0 -0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2037037037037037 0.0 1.0 0.0 -0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2222222222222222 0.0 1.0 0.0 -0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.24074074074074073 0.0 1.0 0.0 -0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.25925925925925924 0.0 1.0 0.0 -0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2777777777777778 0.0 1.0 0.0 -0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2962962962962963 0.0 1.0 0.0 -0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.31481481481481477 0.0 1.0 0.0 -0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.3333333333333333 0.0 1.0 0.0 -0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.35185185185185186 0.0 1.0 0.0 -0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.37037037037037035 0.0 1.0 0.0 -0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.38888888888888884 0.0 1.0 0.0 -0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4074074074074074 0.0 1.0 0.0 -0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.42592592592592593 0.0 1.0 0.0 -0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4444444444444444 0.0 1.0 0.0 -0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4629629629629629 0.0 1.0 0.0 -0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.48148148148148145 0.0 1.0 0.0 -0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5 0.0 1.0 0.0 -0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5185185185185185 0.0 1.0 0.0 -0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.537037037037037 0.0 1.0 0.0 -0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5555555555555556 0.0 1.0 0.0 -0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5740740740740741 0.0 1.0 0.0 -0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5925925925925926 0.0 1.0 0.0 -0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.611111111111111 0.0 1.0 0.0 -0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6296296296296295 0.0 1.0 0.0 -0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6481481481481481 0.0 1.0 0.0 -0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6666666666666666 0.0 1.0 0.0 -0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6851851851851851 0.0 1.0 0.0 -0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7037037037037037 0.0 1.0 0.0 -0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7222222222222222 0.0 1.0 0.0 -0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7407407407407407 0.0 1.0 0.0 -0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7592592592592593 0.0 1.0 0.0 -0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7777777777777777 0.0 1.0 0.0 -0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7962962962962963 0.0 1.0 0.0 -0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8148148148148148 0.0 1.0 0.0 -0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8333333333333334 0.0 1.0 0.0 -0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8518518518518519 0.0 1.0 0.0 -0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8703703703703705 0.0 1.0 0.0 -0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8888888888888888 0.0 1.0 0.0 -0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9074074074074074 0.0 1.0 0.0 -0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9259259259259258 0.0 1.0 0.0 -0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9444444444444444 0.0 1.0 0.0 -0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9629629629629629 0.0 1.0 0.0 -0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9814814814814815 0.0 1.0 0.0 -0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0 0.0 1.0 0.0 -1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0185185185185186 0.0 1.0 0.0 -1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.037037037037037 0.0 1.0 0.0 -1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0555555555555556 0.0 1.0 0.0 -1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.074074074074074 0.0 1.0 0.0 -1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0925925925925926 0.0 1.0 0.0 -1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1111111111111112 0.0 1.0 0.0 -1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1296296296296298 0.0 1.0 0.0 -1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1481481481481481 0.0 1.0 0.0 -1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1666666666666667 0.0 1.0 0.0 -1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1851851851851851 0.0 1.0 0.0 -1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2037037037037037 0.0 1.0 0.0 -1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.222222222222222 0.0 1.0 0.0 -1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2407407407407407 0.0 1.0 0.0 -1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.259259259259259 0.0 1.0 0.0 -1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2777777777777777 0.0 1.0 0.0 -1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2962962962962963 0.0 1.0 0.0 -1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3148148148148149 0.0 1.0 0.0 -1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3333333333333333 0.0 1.0 0.0 -1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3518518518518519 0.0 1.0 0.0 -1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3703703703703702 0.0 1.0 0.0 -1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3888888888888888 0.0 1.0 0.0 -1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4074074074074074 0.0 1.0 0.0 -1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.425925925925926 0.0 1.0 0.0 -1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4444444444444444 0.0 1.0 0.0 -1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.462962962962963 0.0 1.0 0.0 -1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4814814814814814 0.0 1.0 0.0 -1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.018518518518518517 0.0 1.0 0.0 0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.037037037037037035 0.0 1.0 0.0 0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.05555555555555555 0.0 1.0 0.0 0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.07407407407407407 0.0 1.0 0.0 0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.09259259259259259 0.0 1.0 0.0 0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.1111111111111111 0.0 1.0 0.0 0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.12962962962962962 0.0 1.0 0.0 0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.14814814814814814 0.0 1.0 0.0 0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.16666666666666666 0.0 1.0 0.0 0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.18518518518518517 0.0 1.0 0.0 0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2037037037037037 0.0 1.0 0.0 0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2222222222222222 0.0 1.0 0.0 0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.24074074074074073 0.0 1.0 0.0 0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.25925925925925924 0.0 1.0 0.0 0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2777777777777778 0.0 1.0 0.0 0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.2962962962962963 0.0 1.0 0.0 0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.31481481481481477 0.0 1.0 0.0 0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.3333333333333333 0.0 1.0 0.0 0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.35185185185185186 0.0 1.0 0.0 0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.37037037037037035 0.0 1.0 0.0 0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.38888888888888884 0.0 1.0 0.0 0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4074074074074074 0.0 1.0 0.0 0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.42592592592592593 0.0 1.0 0.0 0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4444444444444444 0.0 1.0 0.0 0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.4629629629629629 0.0 1.0 0.0 0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.48148148148148145 0.0 1.0 0.0 0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5 0.0 1.0 0.0 0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5185185185185185 0.0 1.0 0.0 0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.537037037037037 0.0 1.0 0.0 0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5555555555555556 0.0 1.0 0.0 0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5740740740740741 0.0 1.0 0.0 0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.5925925925925926 0.0 1.0 0.0 0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.611111111111111 0.0 1.0 0.0 0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6296296296296295 0.0 1.0 0.0 0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6481481481481481 0.0 1.0 0.0 0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6666666666666666 0.0 1.0 0.0 0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.6851851851851851 0.0 1.0 0.0 0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7037037037037037 0.0 1.0 0.0 0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7222222222222222 0.0 1.0 0.0 0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7407407407407407 0.0 1.0 0.0 0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7592592592592593 0.0 1.0 0.0 0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7777777777777777 0.0 1.0 0.0 0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.7962962962962963 0.0 1.0 0.0 0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8148148148148148 0.0 1.0 0.0 0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8333333333333334 0.0 1.0 0.0 0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8518518518518519 0.0 1.0 0.0 0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8703703703703705 0.0 1.0 0.0 0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.8888888888888888 0.0 1.0 0.0 0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9074074074074074 0.0 1.0 0.0 0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9259259259259258 0.0 1.0 0.0 0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9444444444444444 0.0 1.0 0.0 0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9629629629629629 0.0 1.0 0.0 0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.9814814814814815 0.0 1.0 0.0 0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0 0.0 1.0 0.0 1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0185185185185186 0.0 1.0 0.0 1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.037037037037037 0.0 1.0 0.0 1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0555555555555556 0.0 1.0 0.0 1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.074074074074074 0.0 1.0 0.0 1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.0925925925925926 0.0 1.0 0.0 1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1111111111111112 0.0 1.0 0.0 1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1296296296296298 0.0 1.0 0.0 1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1481481481481481 0.0 1.0 0.0 1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1666666666666667 0.0 1.0 0.0 1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.1851851851851851 0.0 1.0 0.0 1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2037037037037037 0.0 1.0 0.0 1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.222222222222222 0.0 1.0 0.0 1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2407407407407407 0.0 1.0 0.0 1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.259259259259259 0.0 1.0 0.0 1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2777777777777777 0.0 1.0 0.0 1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.2962962962962963 0.0 1.0 0.0 1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3148148148148149 0.0 1.0 0.0 1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3333333333333333 0.0 1.0 0.0 1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3518518518518519 0.0 1.0 0.0 1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3703703703703702 0.0 1.0 0.0 1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.3888888888888888 0.0 1.0 0.0 1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4074074074074074 0.0 1.0 0.0 1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.425925925925926 0.0 1.0 0.0 1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4444444444444444 0.0 1.0 0.0 1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.462962962962963 0.0 1.0 0.0 1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 1.4814814814814814 0.0 1.0 0.0 1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.018518518518518517 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.037037037037037035 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.05555555555555555 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.07407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.09259259259259259 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.1111111111111111 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.12962962962962962 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.14814814814814814 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.16666666666666666 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.18518518518518517 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.24074074074074073 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.25925925925925924 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2777777777777778 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.31481481481481477 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.3333333333333333 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.35185185185185186 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.37037037037037035 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.38888888888888884 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.42592592592592593 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4629629629629629 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.48148148148148145 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5185185185185185 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.537037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5555555555555556 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5740740740740741 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5925925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.611111111111111 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6296296296296295 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6481481481481481 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6666666666666666 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6851851851851851 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7592592592592593 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7777777777777777 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8148148148148148 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8333333333333334 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8518518518518519 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8703703703703705 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8888888888888888 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9259259259259258 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9629629629629629 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9814814814814815 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0185185185185186 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0555555555555556 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0925925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1111111111111112 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1296296296296298 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1481481481481481 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1666666666666667 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1851851851851851 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2037037037037037 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.222222222222222 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2407407407407407 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.259259259259259 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2777777777777777 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2962962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3148148148148149 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3333333333333333 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3518518518518519 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3703703703703702 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3888888888888888 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4074074074074074 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.425925925925926 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4444444444444444 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.462962962962963 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4814814814814814 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.0 0.0 1.0 0.0 -0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.018518518518518517 0.0 1.0 0.0 -0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.037037037037037035 0.0 1.0 0.0 -0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.05555555555555555 0.0 1.0 0.0 -0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.07407407407407407 0.0 1.0 0.0 -0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.09259259259259259 0.0 1.0 0.0 -0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.1111111111111111 0.0 1.0 0.0 -0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.12962962962962962 0.0 1.0 0.0 -0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.14814814814814814 0.0 1.0 0.0 -0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.16666666666666666 0.0 1.0 0.0 -0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.18518518518518517 0.0 1.0 0.0 -0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2037037037037037 0.0 1.0 0.0 -0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2222222222222222 0.0 1.0 0.0 -0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.24074074074074073 0.0 1.0 0.0 -0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.25925925925925924 0.0 1.0 0.0 -0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2777777777777778 0.0 1.0 0.0 -0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2962962962962963 0.0 1.0 0.0 -0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.31481481481481477 0.0 1.0 0.0 -0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.3333333333333333 0.0 1.0 0.0 -0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.35185185185185186 0.0 1.0 0.0 -0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.37037037037037035 0.0 1.0 0.0 -0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.38888888888888884 0.0 1.0 0.0 -0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4074074074074074 0.0 1.0 0.0 -0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.42592592592592593 0.0 1.0 0.0 -0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4444444444444444 0.0 1.0 0.0 -0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4629629629629629 0.0 1.0 0.0 -0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.48148148148148145 0.0 1.0 0.0 -0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5 0.0 1.0 0.0 -0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5185185185185185 0.0 1.0 0.0 -0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.537037037037037 0.0 1.0 0.0 -0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5555555555555556 0.0 1.0 0.0 -0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5740740740740741 0.0 1.0 0.0 -0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5925925925925926 0.0 1.0 0.0 -0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.611111111111111 0.0 1.0 0.0 -0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6296296296296295 0.0 1.0 0.0 -0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6481481481481481 0.0 1.0 0.0 -0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6666666666666666 0.0 1.0 0.0 -0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6851851851851851 0.0 1.0 0.0 -0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7037037037037037 0.0 1.0 0.0 -0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7222222222222222 0.0 1.0 0.0 -0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7407407407407407 0.0 1.0 0.0 -0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7592592592592593 0.0 1.0 0.0 -0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7777777777777777 0.0 1.0 0.0 -0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7962962962962963 0.0 1.0 0.0 -0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8148148148148148 0.0 1.0 0.0 -0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8333333333333334 0.0 1.0 0.0 -0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8518518518518519 0.0 1.0 0.0 -0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8703703703703705 0.0 1.0 0.0 -0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8888888888888888 0.0 1.0 0.0 -0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9074074074074074 0.0 1.0 0.0 -0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9259259259259258 0.0 1.0 0.0 -0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9444444444444444 0.0 1.0 0.0 -0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9629629629629629 0.0 1.0 0.0 -0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9814814814814815 0.0 1.0 0.0 -0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0 0.0 1.0 0.0 -1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0185185185185186 0.0 1.0 0.0 -1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.037037037037037 0.0 1.0 0.0 -1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0555555555555556 0.0 1.0 0.0 -1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.074074074074074 0.0 1.0 0.0 -1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0925925925925926 0.0 1.0 0.0 -1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1111111111111112 0.0 1.0 0.0 -1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1296296296296298 0.0 1.0 0.0 -1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1481481481481481 0.0 1.0 0.0 -1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1666666666666667 0.0 1.0 0.0 -1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1851851851851851 0.0 1.0 0.0 -1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2037037037037037 0.0 1.0 0.0 -1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.222222222222222 0.0 1.0 0.0 -1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2407407407407407 0.0 1.0 0.0 -1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.259259259259259 0.0 1.0 0.0 -1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2777777777777777 0.0 1.0 0.0 -1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2962962962962963 0.0 1.0 0.0 -1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3148148148148149 0.0 1.0 0.0 -1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3333333333333333 0.0 1.0 0.0 -1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3518518518518519 0.0 1.0 0.0 -1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3703703703703702 0.0 1.0 0.0 -1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3888888888888888 0.0 1.0 0.0 -1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4074074074074074 0.0 1.0 0.0 -1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.425925925925926 0.0 1.0 0.0 -1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4444444444444444 0.0 1.0 0.0 -1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.462962962962963 0.0 1.0 0.0 -1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4814814814814814 0.0 1.0 0.0 -1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.018518518518518517 0.0 1.0 0.0 0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.037037037037037035 0.0 1.0 0.0 0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.05555555555555555 0.0 1.0 0.0 0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.07407407407407407 0.0 1.0 0.0 0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.09259259259259259 0.0 1.0 0.0 0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.1111111111111111 0.0 1.0 0.0 0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.12962962962962962 0.0 1.0 0.0 0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.14814814814814814 0.0 1.0 0.0 0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.16666666666666666 0.0 1.0 0.0 0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.18518518518518517 0.0 1.0 0.0 0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2037037037037037 0.0 1.0 0.0 0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2222222222222222 0.0 1.0 0.0 0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.24074074074074073 0.0 1.0 0.0 0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.25925925925925924 0.0 1.0 0.0 0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2777777777777778 0.0 1.0 0.0 0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.2962962962962963 0.0 1.0 0.0 0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.31481481481481477 0.0 1.0 0.0 0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.3333333333333333 0.0 1.0 0.0 0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.35185185185185186 0.0 1.0 0.0 0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.37037037037037035 0.0 1.0 0.0 0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.38888888888888884 0.0 1.0 0.0 0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4074074074074074 0.0 1.0 0.0 0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.42592592592592593 0.0 1.0 0.0 0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4444444444444444 0.0 1.0 0.0 0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.4629629629629629 0.0 1.0 0.0 0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.48148148148148145 0.0 1.0 0.0 0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5 0.0 1.0 0.0 0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5185185185185185 0.0 1.0 0.0 0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.537037037037037 0.0 1.0 0.0 0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5555555555555556 0.0 1.0 0.0 0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5740740740740741 0.0 1.0 0.0 0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.5925925925925926 0.0 1.0 0.0 0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.611111111111111 0.0 1.0 0.0 0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6296296296296295 0.0 1.0 0.0 0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6481481481481481 0.0 1.0 0.0 0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6666666666666666 0.0 1.0 0.0 0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.6851851851851851 0.0 1.0 0.0 0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7037037037037037 0.0 1.0 0.0 0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7222222222222222 0.0 1.0 0.0 0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7407407407407407 0.0 1.0 0.0 0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7592592592592593 0.0 1.0 0.0 0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7777777777777777 0.0 1.0 0.0 0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.7962962962962963 0.0 1.0 0.0 0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8148148148148148 0.0 1.0 0.0 0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8333333333333334 0.0 1.0 0.0 0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8518518518518519 0.0 1.0 0.0 0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8703703703703705 0.0 1.0 0.0 0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.8888888888888888 0.0 1.0 0.0 0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9074074074074074 0.0 1.0 0.0 0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9259259259259258 0.0 1.0 0.0 0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9444444444444444 0.0 1.0 0.0 0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9629629629629629 0.0 1.0 0.0 0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -0.9814814814814815 0.0 1.0 0.0 0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0 0.0 1.0 0.0 1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0185185185185186 0.0 1.0 0.0 1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.037037037037037 0.0 1.0 0.0 1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0555555555555556 0.0 1.0 0.0 1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.074074074074074 0.0 1.0 0.0 1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.0925925925925926 0.0 1.0 0.0 1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1111111111111112 0.0 1.0 0.0 1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1296296296296298 0.0 1.0 0.0 1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1481481481481481 0.0 1.0 0.0 1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1666666666666667 0.0 1.0 0.0 1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.1851851851851851 0.0 1.0 0.0 1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2037037037037037 0.0 1.0 0.0 1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.222222222222222 0.0 1.0 0.0 1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2407407407407407 0.0 1.0 0.0 1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.259259259259259 0.0 1.0 0.0 1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2777777777777777 0.0 1.0 0.0 1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.2962962962962963 0.0 1.0 0.0 1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3148148148148149 0.0 1.0 0.0 1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3333333333333333 0.0 1.0 0.0 1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3518518518518519 0.0 1.0 0.0 1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3703703703703702 0.0 1.0 0.0 1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.3888888888888888 0.0 1.0 0.0 1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4074074074074074 0.0 1.0 0.0 1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.425925925925926 0.0 1.0 0.0 1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4444444444444444 0.0 1.0 0.0 1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.462962962962963 0.0 1.0 0.0 1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 -1.4814814814814814 0.0 1.0 0.0 1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.018518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.037037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.05555555555555555 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.07407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.09259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.1111111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.12962962962962962 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.14814814814814814 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.16666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.18518518518518517 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.2222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.24074074074074073 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.25925925925925924 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.2777777777777778 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.31481481481481477 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.35185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.37037037037037035 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.38888888888888884 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.42592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.4629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.48148148148148145 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.5 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.5185185185185185 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.537037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.5555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.5740740740740741 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.5925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.611111111111111 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.6296296296296295 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.6481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.6666666666666666 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.6851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7592592592592593 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.7962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.8148148148148148 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.8333333333333334 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.8518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.8703703703703705 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.8888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.9074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.9259259259259258 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.9444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.9629629629629629 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.9814814814814815 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.0185185185185186 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.0555555555555556 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.0925925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.1111111111111112 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.1296296296296298 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.1481481481481481 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.1666666666666667 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.1851851851851851 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.2037037037037037 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.222222222222222 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.2407407407407407 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.259259259259259 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.2777777777777777 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.2962962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.3148148148148149 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.3333333333333333 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.3518518518518519 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.3703703703703702 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.3888888888888888 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.4074074074074074 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.425925925925926 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.4444444444444444 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.462962962962963 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 1.4814814814814814 0.0 0.0 1.0 0.0
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.037037037037037035
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.07407407407407407
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.1111111111111111
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.14814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.18518518518518517
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.2222222222222222
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.25925925925925924
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.3333333333333333
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.37037037037037035
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.4444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.48148148148148145
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.5185185185185185
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.5555555555555556
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.6296296296296295
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.6666666666666666
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.7407407407407407
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.8148148148148148
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.8518518518518519
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.8888888888888888
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.9259259259259258
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -0.9629629629629629
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.1111111111111112
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.1481481481481481
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.1851851851851851
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.222222222222222
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.259259259259259
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.3333333333333333
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.3703703703703702
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.4444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.4814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.5185185185185186
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.5555555555555554
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.6296296296296295
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.6666666666666667
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.740740740740741
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.8148148148148149
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.8518518518518516
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.8888888888888888
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.9259259259259258
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -1.962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.111111111111111
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.148148148148148
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.185185185185185
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.2222222222222223
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.2592592592592595
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.3333333333333335
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.3703703703703702
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.4814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.518518518518518
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.5555555555555554
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.6296296296296298
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.6666666666666665
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.7407407407407405
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.814814814814815
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.851851851851852
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.888888888888889
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 -2.962962962962963
|
||||
@@ -1,82 +0,0 @@
|
||||
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.037037037037037035
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.07407407407407407
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.1111111111111111
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.14814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.18518518518518517
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.2222222222222222
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.25925925925925924
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.3333333333333333
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.37037037037037035
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.4444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.48148148148148145
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.5185185185185185
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.5555555555555556
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.6296296296296295
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.6666666666666666
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.7407407407407407
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.8148148148148148
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.8518518518518519
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.8888888888888888
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.9259259259259258
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 0.9629629629629629
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.1111111111111112
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.1481481481481481
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.1851851851851851
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.222222222222222
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.259259259259259
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.3333333333333333
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.3703703703703702
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.4444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.4814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.5185185185185186
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.5555555555555554
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.6296296296296295
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.6666666666666667
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.740740740740741
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.8148148148148149
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.8518518518518516
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.8888888888888888
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.9259259259259258
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 1.962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.0
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.111111111111111
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.148148148148148
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.185185185185185
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.2222222222222223
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.2592592592592595
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.2962962962962963
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.3333333333333335
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.3703703703703702
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.4074074074074074
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.444444444444444
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.4814814814814814
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.518518518518518
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.5555555555555554
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.5925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.6296296296296298
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.6666666666666665
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.7037037037037037
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.7407407407407405
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.7777777777777777
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.814814814814815
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.851851851851852
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.888888888888889
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.925925925925926
|
||||
0 0.532139961 0.946026558 0.5 0.5 0 0 1.0 0.0 0.0 0.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0 2.962962962962963
|
||||
|
Before Width: | Height: | Size: 150 KiB |
|
Before Width: | Height: | Size: 43 KiB |
|
Before Width: | Height: | Size: 11 KiB |
|
Before Width: | Height: | Size: 42 KiB |
|
Before Width: | Height: | Size: 129 KiB |
|
Before Width: | Height: | Size: 226 KiB |
@@ -0,0 +1,149 @@
|
||||
import io
|
||||
import gc
|
||||
import base64
|
||||
import torch
|
||||
import gradio as gr
|
||||
import tempfile
|
||||
import hashlib
|
||||
import os
|
||||
|
||||
from fastapi import FastAPI
|
||||
from io import BytesIO
|
||||
from PIL import Image
|
||||
|
||||
# Function to encode a file to Base64
|
||||
def encode_file_to_base64(file_path):
|
||||
with open(file_path, "rb") as file:
|
||||
# Encode the data to Base64
|
||||
file_base64 = base64.b64encode(file.read())
|
||||
return file_base64
|
||||
|
||||
def update_edition_api(_: gr.Blocks, app: FastAPI, controller):
|
||||
@app.post("/cogvideox_fun/update_edition")
|
||||
def _update_edition_api(
|
||||
datas: dict,
|
||||
):
|
||||
edition = datas.get('edition', 'v2')
|
||||
|
||||
try:
|
||||
controller.update_edition(
|
||||
edition
|
||||
)
|
||||
comment = "Success"
|
||||
except Exception as e:
|
||||
torch.cuda.empty_cache()
|
||||
comment = f"Error. error information is {str(e)}"
|
||||
|
||||
return {"message": comment}
|
||||
|
||||
def update_diffusion_transformer_api(_: gr.Blocks, app: FastAPI, controller):
|
||||
@app.post("/cogvideox_fun/update_diffusion_transformer")
|
||||
def _update_diffusion_transformer_api(
|
||||
datas: dict,
|
||||
):
|
||||
diffusion_transformer_path = datas.get('diffusion_transformer_path', 'none')
|
||||
|
||||
try:
|
||||
controller.update_diffusion_transformer(
|
||||
diffusion_transformer_path
|
||||
)
|
||||
comment = "Success"
|
||||
except Exception as e:
|
||||
torch.cuda.empty_cache()
|
||||
comment = f"Error. error information is {str(e)}"
|
||||
|
||||
return {"message": comment}
|
||||
|
||||
def save_base64_video(base64_string):
|
||||
video_data = base64.b64decode(base64_string)
|
||||
|
||||
md5_hash = hashlib.md5(video_data).hexdigest()
|
||||
filename = f"{md5_hash}.mp4"
|
||||
|
||||
temp_dir = tempfile.gettempdir()
|
||||
file_path = os.path.join(temp_dir, filename)
|
||||
|
||||
with open(file_path, 'wb') as video_file:
|
||||
video_file.write(video_data)
|
||||
|
||||
return file_path
|
||||
|
||||
def infer_forward_api(_: gr.Blocks, app: FastAPI, controller):
|
||||
@app.post("/cogvideox_fun/infer_forward")
|
||||
def _infer_forward_api(
|
||||
datas: dict,
|
||||
):
|
||||
base_model_path = datas.get('base_model_path', 'none')
|
||||
lora_model_path = datas.get('lora_model_path', 'none')
|
||||
lora_alpha_slider = datas.get('lora_alpha_slider', 0.55)
|
||||
prompt_textbox = datas.get('prompt_textbox', None)
|
||||
negative_prompt_textbox = datas.get('negative_prompt_textbox', 'The video is not of a high quality, it has a low resolution, and the audio quality is not clear. Strange motion trajectory, a poor composition and deformed video, low resolution, duplicate and ugly, strange body structure, long and strange neck, bad teeth, bad eyes, bad limbs, bad hands, rotating camera, blurry camera, shaking camera. Deformation, low-resolution, blurry, ugly, distortion.')
|
||||
sampler_dropdown = datas.get('sampler_dropdown', 'Euler')
|
||||
sample_step_slider = datas.get('sample_step_slider', 30)
|
||||
resize_method = datas.get('resize_method', "Generate by")
|
||||
width_slider = datas.get('width_slider', 672)
|
||||
height_slider = datas.get('height_slider', 384)
|
||||
base_resolution = datas.get('base_resolution', 512)
|
||||
is_image = datas.get('is_image', False)
|
||||
generation_method = datas.get('generation_method', False)
|
||||
length_slider = datas.get('length_slider', 144)
|
||||
overlap_video_length = datas.get('overlap_video_length', 4)
|
||||
partial_video_length = datas.get('partial_video_length', 72)
|
||||
cfg_scale_slider = datas.get('cfg_scale_slider', 6)
|
||||
start_image = datas.get('start_image', None)
|
||||
end_image = datas.get('end_image', None)
|
||||
validation_video = datas.get('validation_video', None)
|
||||
denoise_strength = datas.get('denoise_strength', 0.70)
|
||||
seed_textbox = datas.get("seed_textbox", 43)
|
||||
|
||||
generation_method = "Image Generation" if is_image else generation_method
|
||||
|
||||
if start_image is not None:
|
||||
start_image = base64.b64decode(start_image)
|
||||
start_image = [Image.open(BytesIO(start_image))]
|
||||
|
||||
if end_image is not None:
|
||||
end_image = base64.b64decode(end_image)
|
||||
end_image = [Image.open(BytesIO(end_image))]
|
||||
|
||||
if validation_video is not None:
|
||||
validation_video = save_base64_video(validation_video)
|
||||
|
||||
try:
|
||||
save_sample_path, comment = controller.generate(
|
||||
"",
|
||||
base_model_path,
|
||||
lora_model_path,
|
||||
lora_alpha_slider,
|
||||
prompt_textbox,
|
||||
negative_prompt_textbox,
|
||||
sampler_dropdown,
|
||||
sample_step_slider,
|
||||
resize_method,
|
||||
width_slider,
|
||||
height_slider,
|
||||
base_resolution,
|
||||
generation_method,
|
||||
length_slider,
|
||||
overlap_video_length,
|
||||
partial_video_length,
|
||||
cfg_scale_slider,
|
||||
start_image,
|
||||
end_image,
|
||||
validation_video,
|
||||
denoise_strength,
|
||||
seed_textbox,
|
||||
is_api = True,
|
||||
)
|
||||
except Exception as e:
|
||||
gc.collect()
|
||||
torch.cuda.empty_cache()
|
||||
torch.cuda.ipc_collect()
|
||||
save_sample_path = ""
|
||||
comment = f"Error. error information is {str(e)}"
|
||||
return {"message": comment}
|
||||
|
||||
if save_sample_path != "":
|
||||
return {"message": comment, "save_sample_path": save_sample_path, "base64_encoding": encode_file_to_base64(save_sample_path)}
|
||||
else:
|
||||
return {"message": comment, "save_sample_path": save_sample_path}
|
||||
@@ -0,0 +1,96 @@
|
||||
import base64
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime
|
||||
from io import BytesIO
|
||||
|
||||
import cv2
|
||||
import requests
|
||||
import base64
|
||||
|
||||
|
||||
def post_diffusion_transformer(diffusion_transformer_path, url='http://127.0.0.1:7860'):
|
||||
datas = json.dumps({
|
||||
"diffusion_transformer_path": diffusion_transformer_path
|
||||
})
|
||||
r = requests.post(f'{url}/cogvideox_fun/update_diffusion_transformer', data=datas, timeout=1500)
|
||||
data = r.content.decode('utf-8')
|
||||
return data
|
||||
|
||||
def post_update_edition(edition, url='http://0.0.0.0:7860'):
|
||||
datas = json.dumps({
|
||||
"edition": edition
|
||||
})
|
||||
r = requests.post(f'{url}/cogvideox_fun/update_edition', data=datas, timeout=1500)
|
||||
data = r.content.decode('utf-8')
|
||||
return data
|
||||
|
||||
def post_infer(generation_method, length_slider, url='http://127.0.0.1:7860'):
|
||||
datas = json.dumps({
|
||||
"base_model_path": "none",
|
||||
"motion_module_path": "none",
|
||||
"lora_model_path": "none",
|
||||
"lora_alpha_slider": 0.55,
|
||||
"prompt_textbox": "This video shows Mount saint helens, washington - the stunning scenery of a rocky mountains during golden hours - wide shot. A soaring drone footage captures the majestic beauty of a coastal cliff, its red and yellow stratified rock faces rich in color and against the vibrant turquoise of the sea.",
|
||||
"negative_prompt_textbox": "Strange motion trajectory, a poor composition and deformed video, worst quality, normal quality, low quality, low resolution, duplicate and ugly, strange body structure, long and strange neck, bad teeth, bad eyes, bad limbs, bad hands, rotating camera, blurry camera, shaking camera",
|
||||
"sampler_dropdown": "Euler",
|
||||
"sample_step_slider": 30,
|
||||
"width_slider": 672,
|
||||
"height_slider": 384,
|
||||
"generation_method": "Video Generation",
|
||||
"length_slider": length_slider,
|
||||
"cfg_scale_slider": 6,
|
||||
"seed_textbox": 43,
|
||||
})
|
||||
r = requests.post(f'{url}/cogvideox_fun/infer_forward', data=datas, timeout=1500)
|
||||
data = r.content.decode('utf-8')
|
||||
return data
|
||||
|
||||
if __name__ == '__main__':
|
||||
# initiate time
|
||||
now_date = datetime.now()
|
||||
time_start = time.time()
|
||||
|
||||
# -------------------------- #
|
||||
# Step 1: update edition
|
||||
# -------------------------- #
|
||||
edition = "v3"
|
||||
outputs = post_update_edition(edition)
|
||||
print('Output update edition: ', outputs)
|
||||
|
||||
# -------------------------- #
|
||||
# Step 2: update edition
|
||||
# -------------------------- #
|
||||
diffusion_transformer_path = "models/Diffusion_Transformer/cogvideox_funV3-XL-2-512x512"
|
||||
outputs = post_diffusion_transformer(diffusion_transformer_path)
|
||||
print('Output update edition: ', outputs)
|
||||
|
||||
# -------------------------- #
|
||||
# Step 3: infer
|
||||
# -------------------------- #
|
||||
# "Video Generation" and "Image Generation"
|
||||
generation_method = "Video Generation"
|
||||
length_slider = 72
|
||||
outputs = post_infer(generation_method, length_slider)
|
||||
|
||||
# Get decoded data
|
||||
outputs = json.loads(outputs)
|
||||
base64_encoding = outputs["base64_encoding"]
|
||||
decoded_data = base64.b64decode(base64_encoding)
|
||||
|
||||
is_image = True if generation_method == "Image Generation" else False
|
||||
if is_image or length_slider == 1:
|
||||
file_path = "1.png"
|
||||
else:
|
||||
file_path = "1.mp4"
|
||||
with open(file_path, "wb") as file:
|
||||
file.write(decoded_data)
|
||||
|
||||
# End of record time
|
||||
# The calculated time difference is the execution time of the program, expressed in seconds / s
|
||||
time_end = time.time()
|
||||
time_sum = (time_end - time_start) % 60
|
||||
print('# --------------------------------------------------------- #')
|
||||
print(f'# Total expenditure: {time_sum}s')
|
||||
print('# --------------------------------------------------------- #')
|
||||
@@ -9,7 +9,6 @@ import torch
|
||||
from PIL import Image
|
||||
from torch.utils.data import BatchSampler, Dataset, Sampler
|
||||
|
||||
|
||||
ASPECT_RATIO_512 = {
|
||||
'0.25': [256.0, 1024.0], '0.26': [256.0, 992.0], '0.27': [256.0, 960.0], '0.28': [256.0, 928.0],
|
||||
'0.32': [288.0, 896.0], '0.33': [288.0, 864.0], '0.35': [288.0, 832.0], '0.4': [320.0, 800.0],
|
||||
@@ -38,18 +37,15 @@ ASPECT_RATIO_RANDOM_CROP_PROB = [
|
||||
]
|
||||
ASPECT_RATIO_RANDOM_CROP_PROB = np.array(ASPECT_RATIO_RANDOM_CROP_PROB) / sum(ASPECT_RATIO_RANDOM_CROP_PROB)
|
||||
|
||||
|
||||
def get_closest_ratio(height: float, width: float, ratios: dict = ASPECT_RATIO_512):
|
||||
aspect_ratio = height / width
|
||||
closest_ratio = min(ratios.keys(), key=lambda ratio: abs(float(ratio) - aspect_ratio))
|
||||
return ratios[closest_ratio], float(closest_ratio)
|
||||
|
||||
|
||||
def get_image_size_without_loading(path):
|
||||
with Image.open(path) as img:
|
||||
return img.size # (width, height)
|
||||
|
||||
|
||||
class RandomSampler(Sampler[int]):
|
||||
r"""Samples elements randomly. If without replacement, then sample from a shuffled dataset.
|
||||
|
||||
@@ -60,22 +56,18 @@ class RandomSampler(Sampler[int]):
|
||||
replacement (bool): samples are drawn on-demand with replacement if ``True``, default=``False``
|
||||
num_samples (int): number of samples to draw, default=`len(dataset)`.
|
||||
generator (Generator): Generator used in sampling.
|
||||
k_repeat (int): number of times to repeat each sampled index consecutively, default=1.
|
||||
When k_repeat > 1, each index is yielded k_repeat times in a row,
|
||||
so a batch of size B will contain B // k_repeat unique samples.
|
||||
"""
|
||||
|
||||
data_source: Sized
|
||||
replacement: bool
|
||||
|
||||
def __init__(self, data_source: Sized, replacement: bool = False,
|
||||
num_samples: Optional[int] = None, generator=None, k_repeat: int = 1) -> None:
|
||||
num_samples: Optional[int] = None, generator=None) -> None:
|
||||
self.data_source = data_source
|
||||
self.replacement = replacement
|
||||
self._num_samples = num_samples
|
||||
self.generator = generator
|
||||
self._pos_start = 0
|
||||
self.k_repeat = k_repeat
|
||||
|
||||
if not isinstance(self.replacement, bool):
|
||||
raise TypeError(f"replacement should be a boolean value, but got replacement={self.replacement}")
|
||||
@@ -101,12 +93,8 @@ class RandomSampler(Sampler[int]):
|
||||
|
||||
if self.replacement:
|
||||
for _ in range(self.num_samples // 32):
|
||||
for idx in torch.randint(high=n, size=(32,), dtype=torch.int64, generator=generator).tolist():
|
||||
for _ in range(self.k_repeat):
|
||||
yield idx
|
||||
for idx in torch.randint(high=n, size=(self.num_samples % 32,), dtype=torch.int64, generator=generator).tolist():
|
||||
for _ in range(self.k_repeat):
|
||||
yield idx
|
||||
yield from torch.randint(high=n, size=(32,), dtype=torch.int64, generator=generator).tolist()
|
||||
yield from torch.randint(high=n, size=(self.num_samples % 32,), dtype=torch.int64, generator=generator).tolist()
|
||||
else:
|
||||
for _ in range(self.num_samples // n):
|
||||
xx = torch.randperm(n, generator=generator).tolist()
|
||||
@@ -114,17 +102,13 @@ class RandomSampler(Sampler[int]):
|
||||
self._pos_start = 0
|
||||
print("xx top 10", xx[:10], self._pos_start)
|
||||
for idx in range(self._pos_start, n):
|
||||
for _ in range(self.k_repeat):
|
||||
yield xx[idx]
|
||||
yield xx[idx]
|
||||
self._pos_start = (self._pos_start + 1) % n
|
||||
self._pos_start = 0
|
||||
for idx in torch.randperm(n, generator=generator).tolist()[:self.num_samples % n]:
|
||||
for _ in range(self.k_repeat):
|
||||
yield idx
|
||||
yield from torch.randperm(n, generator=generator).tolist()[:self.num_samples % n]
|
||||
|
||||
def __len__(self) -> int:
|
||||
return self.num_samples * self.k_repeat
|
||||
|
||||
return self.num_samples
|
||||
|
||||
class AspectRatioBatchImageSampler(BatchSampler):
|
||||
"""A sampler wrapper for grouping images with similar aspect ratio into a same batch.
|
||||
@@ -246,7 +230,7 @@ class AspectRatioBatchSampler(BatchSampler):
|
||||
for idx in self.sampler:
|
||||
try:
|
||||
video_dict = self.dataset[idx]
|
||||
width, height = video_dict.get("width", None), video_dict.get("height", None)
|
||||
width, more = video_dict.get("width", None), video_dict.get("height", None)
|
||||
|
||||
if width is None or height is None:
|
||||
if self.train_data_format == "normal":
|
||||
@@ -260,9 +244,9 @@ class AspectRatioBatchSampler(BatchSampler):
|
||||
video_dir = os.path.join(self.video_folder, f"{videoid}.mp4")
|
||||
cap = cv2.VideoCapture(video_dir)
|
||||
|
||||
# Get video dimensions
|
||||
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # Convert float to integer
|
||||
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)) # Convert float to integer
|
||||
# 获取视频尺寸
|
||||
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # 浮点数转换为整数
|
||||
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)) # 浮点数转换为整数
|
||||
|
||||
ratio = height / width # self.dataset[idx]
|
||||
else:
|
||||
@@ -270,7 +254,7 @@ class AspectRatioBatchSampler(BatchSampler):
|
||||
width = int(width)
|
||||
ratio = height / width # self.dataset[idx]
|
||||
except Exception as e:
|
||||
print(e, self.dataset[idx], "This item is error, please check it.")
|
||||
print(e)
|
||||
continue
|
||||
# find the closest aspect ratio
|
||||
closest_ratio = min(self.aspect_ratios.keys(), key=lambda r: abs(float(r) - ratio))
|
||||
@@ -332,19 +316,11 @@ class AspectRatioBatchImageVideoSampler(BatchSampler):
|
||||
|
||||
width, height = image_dict.get("width", None), image_dict.get("height", None)
|
||||
if width is None or height is None:
|
||||
image_id = image_dict['file_path']
|
||||
# Handle multiview: file_path can be list or str
|
||||
if isinstance(image_id, list):
|
||||
image_dir = image_id[0]
|
||||
else:
|
||||
image_dir = image_id
|
||||
|
||||
image_id, name = image_dict['file_path'], image_dict['text']
|
||||
if self.train_folder is None:
|
||||
pass # image_dir is already absolute path
|
||||
elif isinstance(self.train_folder, list):
|
||||
pass # train_folder is list, use image_dir directly
|
||||
image_dir = image_id
|
||||
else:
|
||||
image_dir = os.path.join(self.train_folder, image_dir)
|
||||
image_dir = os.path.join(self.train_folder, image_id)
|
||||
|
||||
width, height = get_image_size_without_loading(image_dir)
|
||||
|
||||
@@ -354,7 +330,7 @@ class AspectRatioBatchImageVideoSampler(BatchSampler):
|
||||
width = int(width)
|
||||
ratio = height / width # self.dataset[idx]
|
||||
except Exception as e:
|
||||
print(e, self.dataset[idx], "This item is error, please check it.")
|
||||
print(e)
|
||||
continue
|
||||
# find the closest aspect ratio
|
||||
closest_ratio = min(self.aspect_ratios.keys(), key=lambda r: abs(float(r) - ratio))
|
||||
@@ -372,24 +348,16 @@ class AspectRatioBatchImageVideoSampler(BatchSampler):
|
||||
width, height = video_dict.get("width", None), video_dict.get("height", None)
|
||||
|
||||
if width is None or height is None:
|
||||
video_id = video_dict['file_path']
|
||||
# Handle multiview: file_path can be list or str
|
||||
if isinstance(video_id, list):
|
||||
video_dir = video_id[0] # Use first view for aspect ratio
|
||||
else:
|
||||
video_dir = video_id
|
||||
|
||||
video_id, name = video_dict['file_path'], video_dict['text']
|
||||
if self.train_folder is None:
|
||||
pass # video_dir is already absolute path
|
||||
elif isinstance(self.train_folder, list):
|
||||
pass # train_folder is list, use video_dir directly
|
||||
video_dir = video_id
|
||||
else:
|
||||
video_dir = os.path.join(self.train_folder, video_dir)
|
||||
video_dir = os.path.join(self.train_folder, video_id)
|
||||
cap = cv2.VideoCapture(video_dir)
|
||||
|
||||
# Get video dimensions
|
||||
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # Convert float to integer
|
||||
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)) # Convert float to integer
|
||||
# 获取视频尺寸
|
||||
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) # 浮点数转换为整数
|
||||
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)) # 浮点数转换为整数
|
||||
|
||||
ratio = height / width # self.dataset[idx]
|
||||
else:
|
||||
@@ -397,7 +365,7 @@ class AspectRatioBatchImageVideoSampler(BatchSampler):
|
||||
width = int(width)
|
||||
ratio = height / width # self.dataset[idx]
|
||||
except Exception as e:
|
||||
print(e, self.dataset[idx], "This item is error, please check it.")
|
||||
print(e)
|
||||
continue
|
||||
# find the closest aspect ratio
|
||||
closest_ratio = min(self.aspect_ratios.keys(), key=lambda r: abs(float(r) - ratio))
|
||||
@@ -0,0 +1,76 @@
|
||||
import json
|
||||
import os
|
||||
import random
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torchvision.transforms as transforms
|
||||
from PIL import Image
|
||||
from torch.utils.data.dataset import Dataset
|
||||
|
||||
|
||||
class CC15M(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
json_path,
|
||||
video_folder=None,
|
||||
resolution=512,
|
||||
enable_bucket=False,
|
||||
):
|
||||
print(f"loading annotations from {json_path} ...")
|
||||
self.dataset = json.load(open(json_path, 'r'))
|
||||
self.length = len(self.dataset)
|
||||
print(f"data scale: {self.length}")
|
||||
|
||||
self.enable_bucket = enable_bucket
|
||||
self.video_folder = video_folder
|
||||
|
||||
resolution = tuple(resolution) if not isinstance(resolution, int) else (resolution, resolution)
|
||||
self.pixel_transforms = transforms.Compose([
|
||||
transforms.Resize(resolution[0]),
|
||||
transforms.CenterCrop(resolution),
|
||||
transforms.ToTensor(),
|
||||
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True),
|
||||
])
|
||||
|
||||
def get_batch(self, idx):
|
||||
video_dict = self.dataset[idx]
|
||||
video_id, name = video_dict['file_path'], video_dict['text']
|
||||
|
||||
if self.video_folder is None:
|
||||
video_dir = video_id
|
||||
else:
|
||||
video_dir = os.path.join(self.video_folder, video_id)
|
||||
|
||||
pixel_values = Image.open(video_dir).convert("RGB")
|
||||
return pixel_values, name
|
||||
|
||||
def __len__(self):
|
||||
return self.length
|
||||
|
||||
def __getitem__(self, idx):
|
||||
while True:
|
||||
try:
|
||||
pixel_values, name = self.get_batch(idx)
|
||||
break
|
||||
except Exception as e:
|
||||
print(e)
|
||||
idx = random.randint(0, self.length-1)
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = self.pixel_transforms(pixel_values)
|
||||
else:
|
||||
pixel_values = np.array(pixel_values)
|
||||
|
||||
sample = dict(pixel_values=pixel_values, text=name)
|
||||
return sample
|
||||
|
||||
if __name__ == "__main__":
|
||||
dataset = CC15M(
|
||||
csv_path="/mnt_wg/zhoumo.xjq/CCUtils/cc15m_add_index.json",
|
||||
resolution=512,
|
||||
)
|
||||
|
||||
dataloader = torch.utils.data.DataLoader(dataset, batch_size=4, num_workers=0,)
|
||||
for idx, batch in enumerate(dataloader):
|
||||
print(batch["pixel_values"].shape, len(batch["text"]))
|
||||
@@ -0,0 +1,324 @@
|
||||
import csv
|
||||
import io
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import random
|
||||
from threading import Thread
|
||||
|
||||
import albumentations
|
||||
import cv2
|
||||
import gc
|
||||
import numpy as np
|
||||
import torch
|
||||
import torchvision.transforms as transforms
|
||||
|
||||
from func_timeout import func_timeout, FunctionTimedOut
|
||||
from decord import VideoReader
|
||||
from PIL import Image
|
||||
from torch.utils.data import BatchSampler, Sampler
|
||||
from torch.utils.data.dataset import Dataset
|
||||
from contextlib import contextmanager
|
||||
|
||||
VIDEO_READER_TIMEOUT = 20
|
||||
|
||||
def get_random_mask(shape):
|
||||
f, c, h, w = shape
|
||||
|
||||
if f != 1:
|
||||
mask_index = np.random.choice([0, 1, 2, 3, 4], p = [0.05, 0.3, 0.3, 0.3, 0.05]) # np.random.randint(0, 5)
|
||||
else:
|
||||
mask_index = np.random.choice([0, 1], p = [0.2, 0.8]) # np.random.randint(0, 2)
|
||||
mask = torch.zeros((f, 1, h, w), dtype=torch.uint8)
|
||||
|
||||
if mask_index == 0:
|
||||
center_x = torch.randint(0, w, (1,)).item()
|
||||
center_y = torch.randint(0, h, (1,)).item()
|
||||
block_size_x = torch.randint(w // 4, w // 4 * 3, (1,)).item() # 方块的宽度范围
|
||||
block_size_y = torch.randint(h // 4, h // 4 * 3, (1,)).item() # 方块的高度范围
|
||||
|
||||
start_x = max(center_x - block_size_x // 2, 0)
|
||||
end_x = min(center_x + block_size_x // 2, w)
|
||||
start_y = max(center_y - block_size_y // 2, 0)
|
||||
end_y = min(center_y + block_size_y // 2, h)
|
||||
mask[:, :, start_y:end_y, start_x:end_x] = 1
|
||||
elif mask_index == 1:
|
||||
mask[:, :, :, :] = 1
|
||||
elif mask_index == 2:
|
||||
mask_frame_index = np.random.randint(1, 5)
|
||||
mask[mask_frame_index:, :, :, :] = 1
|
||||
elif mask_index == 3:
|
||||
mask_frame_index = np.random.randint(1, 5)
|
||||
mask[mask_frame_index:-mask_frame_index, :, :, :] = 1
|
||||
elif mask_index == 4:
|
||||
center_x = torch.randint(0, w, (1,)).item()
|
||||
center_y = torch.randint(0, h, (1,)).item()
|
||||
block_size_x = torch.randint(w // 4, w // 4 * 3, (1,)).item() # 方块的宽度范围
|
||||
block_size_y = torch.randint(h // 4, h // 4 * 3, (1,)).item() # 方块的高度范围
|
||||
|
||||
start_x = max(center_x - block_size_x // 2, 0)
|
||||
end_x = min(center_x + block_size_x // 2, w)
|
||||
start_y = max(center_y - block_size_y // 2, 0)
|
||||
end_y = min(center_y + block_size_y // 2, h)
|
||||
|
||||
mask_frame_before = np.random.randint(0, f // 2)
|
||||
mask_frame_after = np.random.randint(f // 2, f)
|
||||
mask[mask_frame_before:mask_frame_after, :, start_y:end_y, start_x:end_x] = 1
|
||||
else:
|
||||
raise ValueError(f"The mask_index {mask_index} is not define")
|
||||
return mask
|
||||
|
||||
class ImageVideoSampler(BatchSampler):
|
||||
"""A sampler wrapper for grouping images with similar aspect ratio into a same batch.
|
||||
|
||||
Args:
|
||||
sampler (Sampler): Base sampler.
|
||||
dataset (Dataset): Dataset providing data information.
|
||||
batch_size (int): Size of mini-batch.
|
||||
drop_last (bool): If ``True``, the sampler will drop the last batch if
|
||||
its size would be less than ``batch_size``.
|
||||
aspect_ratios (dict): The predefined aspect ratios.
|
||||
"""
|
||||
|
||||
def __init__(self,
|
||||
sampler: Sampler,
|
||||
dataset: Dataset,
|
||||
batch_size: int,
|
||||
drop_last: bool = False
|
||||
) -> None:
|
||||
if not isinstance(sampler, Sampler):
|
||||
raise TypeError('sampler should be an instance of ``Sampler``, '
|
||||
f'but got {sampler}')
|
||||
if not isinstance(batch_size, int) or batch_size <= 0:
|
||||
raise ValueError('batch_size should be a positive integer value, '
|
||||
f'but got batch_size={batch_size}')
|
||||
self.sampler = sampler
|
||||
self.dataset = dataset
|
||||
self.batch_size = batch_size
|
||||
self.drop_last = drop_last
|
||||
|
||||
# buckets for each aspect ratio
|
||||
self.bucket = {'image':[], 'video':[]}
|
||||
|
||||
def __iter__(self):
|
||||
for idx in self.sampler:
|
||||
content_type = self.dataset.dataset[idx].get('type', 'image')
|
||||
self.bucket[content_type].append(idx)
|
||||
|
||||
# yield a batch of indices in the same aspect ratio group
|
||||
if len(self.bucket['video']) == self.batch_size:
|
||||
bucket = self.bucket['video']
|
||||
yield bucket[:]
|
||||
del bucket[:]
|
||||
elif len(self.bucket['image']) == self.batch_size:
|
||||
bucket = self.bucket['image']
|
||||
yield bucket[:]
|
||||
del bucket[:]
|
||||
|
||||
@contextmanager
|
||||
def VideoReader_contextmanager(*args, **kwargs):
|
||||
vr = VideoReader(*args, **kwargs)
|
||||
try:
|
||||
yield vr
|
||||
finally:
|
||||
del vr
|
||||
gc.collect()
|
||||
|
||||
def get_video_reader_batch(video_reader, batch_index):
|
||||
frames = video_reader.get_batch(batch_index).asnumpy()
|
||||
return frames
|
||||
|
||||
def resize_frame(frame, target_short_side):
|
||||
h, w, _ = frame.shape
|
||||
if h < w:
|
||||
if target_short_side > h:
|
||||
return frame
|
||||
new_h = target_short_side
|
||||
new_w = int(target_short_side * w / h)
|
||||
else:
|
||||
if target_short_side > w:
|
||||
return frame
|
||||
new_w = target_short_side
|
||||
new_h = int(target_short_side * h / w)
|
||||
|
||||
resized_frame = cv2.resize(frame, (new_w, new_h))
|
||||
return resized_frame
|
||||
|
||||
class ImageVideoDataset(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
ann_path, data_root=None,
|
||||
video_sample_size=512, video_sample_stride=4, video_sample_n_frames=16,
|
||||
image_sample_size=512,
|
||||
video_repeat=0,
|
||||
text_drop_ratio=-1,
|
||||
enable_bucket=False,
|
||||
video_length_drop_start=0.1,
|
||||
video_length_drop_end=0.9,
|
||||
enable_inpaint=False,
|
||||
):
|
||||
# Loading annotations from files
|
||||
print(f"loading annotations from {ann_path} ...")
|
||||
if ann_path.endswith('.csv'):
|
||||
with open(ann_path, 'r') as csvfile:
|
||||
dataset = list(csv.DictReader(csvfile))
|
||||
elif ann_path.endswith('.json'):
|
||||
dataset = json.load(open(ann_path))
|
||||
|
||||
self.data_root = data_root
|
||||
|
||||
# It's used to balance num of images and videos.
|
||||
self.dataset = []
|
||||
for data in dataset:
|
||||
if data.get('type', 'image') != 'video':
|
||||
self.dataset.append(data)
|
||||
if video_repeat > 0:
|
||||
for _ in range(video_repeat):
|
||||
for data in dataset:
|
||||
if data.get('type', 'image') == 'video':
|
||||
self.dataset.append(data)
|
||||
del dataset
|
||||
|
||||
self.length = len(self.dataset)
|
||||
print(f"data scale: {self.length}")
|
||||
# TODO: enable bucket training
|
||||
self.enable_bucket = enable_bucket
|
||||
self.text_drop_ratio = text_drop_ratio
|
||||
self.enable_inpaint = enable_inpaint
|
||||
|
||||
self.video_length_drop_start = video_length_drop_start
|
||||
self.video_length_drop_end = video_length_drop_end
|
||||
|
||||
# Video params
|
||||
self.video_sample_stride = video_sample_stride
|
||||
self.video_sample_n_frames = video_sample_n_frames
|
||||
self.video_sample_size = tuple(video_sample_size) if not isinstance(video_sample_size, int) else (video_sample_size, video_sample_size)
|
||||
self.video_transforms = transforms.Compose(
|
||||
[
|
||||
transforms.Resize(min(self.video_sample_size)),
|
||||
transforms.CenterCrop(self.video_sample_size),
|
||||
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True),
|
||||
]
|
||||
)
|
||||
|
||||
# Image params
|
||||
self.image_sample_size = tuple(image_sample_size) if not isinstance(image_sample_size, int) else (image_sample_size, image_sample_size)
|
||||
self.image_transforms = transforms.Compose([
|
||||
transforms.Resize(min(self.image_sample_size)),
|
||||
transforms.CenterCrop(self.image_sample_size),
|
||||
transforms.ToTensor(),
|
||||
transforms.Normalize([0.5, 0.5, 0.5],[0.5, 0.5, 0.5])
|
||||
])
|
||||
|
||||
self.larger_side_of_image_and_video = max(min(self.image_sample_size), min(self.video_sample_size))
|
||||
|
||||
def get_batch(self, idx):
|
||||
data_info = self.dataset[idx % len(self.dataset)]
|
||||
|
||||
if data_info.get('type', 'image')=='video':
|
||||
video_id, text = data_info['file_path'], data_info['text']
|
||||
|
||||
if self.data_root is None:
|
||||
video_dir = video_id
|
||||
else:
|
||||
video_dir = os.path.join(self.data_root, video_id)
|
||||
|
||||
with VideoReader_contextmanager(video_dir, num_threads=2) as video_reader:
|
||||
min_sample_n_frames = min(
|
||||
self.video_sample_n_frames,
|
||||
int(len(video_reader) * (self.video_length_drop_end - self.video_length_drop_start) // self.video_sample_stride)
|
||||
)
|
||||
if min_sample_n_frames == 0:
|
||||
raise ValueError(f"No Frames in video.")
|
||||
|
||||
video_length = int(self.video_length_drop_end * len(video_reader))
|
||||
clip_length = min(video_length, (min_sample_n_frames - 1) * self.video_sample_stride + 1)
|
||||
start_idx = random.randint(int(self.video_length_drop_start * video_length), video_length - clip_length) if video_length != clip_length else 0
|
||||
batch_index = np.linspace(start_idx, start_idx + clip_length - 1, min_sample_n_frames, dtype=int)
|
||||
|
||||
try:
|
||||
sample_args = (video_reader, batch_index)
|
||||
pixel_values = func_timeout(
|
||||
VIDEO_READER_TIMEOUT, get_video_reader_batch, args=sample_args
|
||||
)
|
||||
resized_frames = []
|
||||
for i in range(len(pixel_values)):
|
||||
frame = pixel_values[i]
|
||||
resized_frame = resize_frame(frame, self.larger_side_of_image_and_video)
|
||||
resized_frames.append(resized_frame)
|
||||
pixel_values = np.array(resized_frames)
|
||||
except FunctionTimedOut:
|
||||
raise ValueError(f"Read {idx} timeout.")
|
||||
except Exception as e:
|
||||
raise ValueError(f"Failed to extract frames from video. Error is {e}.")
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = torch.from_numpy(pixel_values).permute(0, 3, 1, 2).contiguous()
|
||||
pixel_values = pixel_values / 255.
|
||||
del video_reader
|
||||
else:
|
||||
pixel_values = pixel_values
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = self.video_transforms(pixel_values)
|
||||
|
||||
# Random use no text generation
|
||||
if random.random() < self.text_drop_ratio:
|
||||
text = ''
|
||||
return pixel_values, text, 'video'
|
||||
else:
|
||||
image_path, text = data_info['file_path'], data_info['text']
|
||||
if self.data_root is not None:
|
||||
image_path = os.path.join(self.data_root, image_path)
|
||||
image = Image.open(image_path).convert('RGB')
|
||||
if not self.enable_bucket:
|
||||
image = self.image_transforms(image).unsqueeze(0)
|
||||
else:
|
||||
image = np.expand_dims(np.array(image), 0)
|
||||
if random.random() < self.text_drop_ratio:
|
||||
text = ''
|
||||
return image, text, 'image'
|
||||
|
||||
def __len__(self):
|
||||
return self.length
|
||||
|
||||
def __getitem__(self, idx):
|
||||
data_info = self.dataset[idx % len(self.dataset)]
|
||||
data_type = data_info.get('type', 'image')
|
||||
while True:
|
||||
sample = {}
|
||||
try:
|
||||
data_info_local = self.dataset[idx % len(self.dataset)]
|
||||
data_type_local = data_info_local.get('type', 'image')
|
||||
if data_type_local != data_type:
|
||||
raise ValueError("data_type_local != data_type")
|
||||
|
||||
pixel_values, name, data_type = self.get_batch(idx)
|
||||
sample["pixel_values"] = pixel_values
|
||||
sample["text"] = name
|
||||
sample["data_type"] = data_type
|
||||
sample["idx"] = idx
|
||||
|
||||
if len(sample) > 0:
|
||||
break
|
||||
except Exception as e:
|
||||
print(e, self.dataset[idx % len(self.dataset)])
|
||||
idx = random.randint(0, self.length-1)
|
||||
|
||||
if self.enable_inpaint and not self.enable_bucket:
|
||||
mask = get_random_mask(pixel_values.size())
|
||||
mask_pixel_values = pixel_values * (1 - mask) + torch.ones_like(pixel_values) * -1 * mask
|
||||
sample["mask_pixel_values"] = mask_pixel_values
|
||||
sample["mask"] = mask
|
||||
|
||||
clip_pixel_values = sample["pixel_values"][0].permute(1, 2, 0).contiguous()
|
||||
clip_pixel_values = (clip_pixel_values * 0.5 + 0.5) * 255
|
||||
sample["clip_pixel_values"] = clip_pixel_values
|
||||
|
||||
ref_pixel_values = sample["pixel_values"][0].unsqueeze(0)
|
||||
if (mask == 1).all():
|
||||
ref_pixel_values = torch.ones_like(ref_pixel_values) * -1
|
||||
sample["ref_pixel_values"] = ref_pixel_values
|
||||
|
||||
return sample
|
||||
|
||||
@@ -0,0 +1,262 @@
|
||||
import csv
|
||||
import gc
|
||||
import io
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import random
|
||||
from contextlib import contextmanager
|
||||
from threading import Thread
|
||||
|
||||
import albumentations
|
||||
import cv2
|
||||
import numpy as np
|
||||
import torch
|
||||
import torchvision.transforms as transforms
|
||||
from decord import VideoReader
|
||||
from einops import rearrange
|
||||
from func_timeout import FunctionTimedOut, func_timeout
|
||||
from PIL import Image
|
||||
from torch.utils.data import BatchSampler, Sampler
|
||||
from torch.utils.data.dataset import Dataset
|
||||
|
||||
VIDEO_READER_TIMEOUT = 20
|
||||
|
||||
def get_random_mask(shape):
|
||||
f, c, h, w = shape
|
||||
|
||||
mask_index = np.random.randint(0, 4)
|
||||
mask = torch.zeros((f, 1, h, w), dtype=torch.uint8)
|
||||
if mask_index == 0:
|
||||
mask[1:, :, :, :] = 1
|
||||
elif mask_index == 1:
|
||||
mask_frame_index = 1
|
||||
mask[mask_frame_index:-mask_frame_index, :, :, :] = 1
|
||||
elif mask_index == 2:
|
||||
center_x = torch.randint(0, w, (1,)).item()
|
||||
center_y = torch.randint(0, h, (1,)).item()
|
||||
block_size_x = torch.randint(w // 4, w // 4 * 3, (1,)).item() # 方块的宽度范围
|
||||
block_size_y = torch.randint(h // 4, h // 4 * 3, (1,)).item() # 方块的高度范围
|
||||
|
||||
start_x = max(center_x - block_size_x // 2, 0)
|
||||
end_x = min(center_x + block_size_x // 2, w)
|
||||
start_y = max(center_y - block_size_y // 2, 0)
|
||||
end_y = min(center_y + block_size_y // 2, h)
|
||||
mask[:, :, start_y:end_y, start_x:end_x] = 1
|
||||
elif mask_index == 3:
|
||||
center_x = torch.randint(0, w, (1,)).item()
|
||||
center_y = torch.randint(0, h, (1,)).item()
|
||||
block_size_x = torch.randint(w // 4, w // 4 * 3, (1,)).item() # 方块的宽度范围
|
||||
block_size_y = torch.randint(h // 4, h // 4 * 3, (1,)).item() # 方块的高度范围
|
||||
|
||||
start_x = max(center_x - block_size_x // 2, 0)
|
||||
end_x = min(center_x + block_size_x // 2, w)
|
||||
start_y = max(center_y - block_size_y // 2, 0)
|
||||
end_y = min(center_y + block_size_y // 2, h)
|
||||
|
||||
mask_frame_before = np.random.randint(0, f // 2)
|
||||
mask_frame_after = np.random.randint(f // 2, f)
|
||||
mask[mask_frame_before:mask_frame_after, :, start_y:end_y, start_x:end_x] = 1
|
||||
else:
|
||||
raise ValueError(f"The mask_index {mask_index} is not define")
|
||||
return mask
|
||||
|
||||
|
||||
@contextmanager
|
||||
def VideoReader_contextmanager(*args, **kwargs):
|
||||
vr = VideoReader(*args, **kwargs)
|
||||
try:
|
||||
yield vr
|
||||
finally:
|
||||
del vr
|
||||
gc.collect()
|
||||
|
||||
|
||||
def get_video_reader_batch(video_reader, batch_index):
|
||||
frames = video_reader.get_batch(batch_index).asnumpy()
|
||||
return frames
|
||||
|
||||
|
||||
class WebVid10M(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
csv_path, video_folder,
|
||||
sample_size=256, sample_stride=4, sample_n_frames=16,
|
||||
enable_bucket=False, enable_inpaint=False, is_image=False,
|
||||
):
|
||||
print(f"loading annotations from {csv_path} ...")
|
||||
with open(csv_path, 'r') as csvfile:
|
||||
self.dataset = list(csv.DictReader(csvfile))
|
||||
self.length = len(self.dataset)
|
||||
print(f"data scale: {self.length}")
|
||||
|
||||
self.video_folder = video_folder
|
||||
self.sample_stride = sample_stride
|
||||
self.sample_n_frames = sample_n_frames
|
||||
self.enable_bucket = enable_bucket
|
||||
self.enable_inpaint = enable_inpaint
|
||||
self.is_image = is_image
|
||||
|
||||
sample_size = tuple(sample_size) if not isinstance(sample_size, int) else (sample_size, sample_size)
|
||||
self.pixel_transforms = transforms.Compose([
|
||||
transforms.Resize(sample_size[0]),
|
||||
transforms.CenterCrop(sample_size),
|
||||
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True),
|
||||
])
|
||||
|
||||
def get_batch(self, idx):
|
||||
video_dict = self.dataset[idx]
|
||||
videoid, name, page_dir = video_dict['videoid'], video_dict['name'], video_dict['page_dir']
|
||||
|
||||
video_dir = os.path.join(self.video_folder, f"{videoid}.mp4")
|
||||
video_reader = VideoReader(video_dir)
|
||||
video_length = len(video_reader)
|
||||
|
||||
if not self.is_image:
|
||||
clip_length = min(video_length, (self.sample_n_frames - 1) * self.sample_stride + 1)
|
||||
start_idx = random.randint(0, video_length - clip_length)
|
||||
batch_index = np.linspace(start_idx, start_idx + clip_length - 1, self.sample_n_frames, dtype=int)
|
||||
else:
|
||||
batch_index = [random.randint(0, video_length - 1)]
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = torch.from_numpy(video_reader.get_batch(batch_index).asnumpy()).permute(0, 3, 1, 2).contiguous()
|
||||
pixel_values = pixel_values / 255.
|
||||
del video_reader
|
||||
else:
|
||||
pixel_values = video_reader.get_batch(batch_index).asnumpy()
|
||||
|
||||
if self.is_image:
|
||||
pixel_values = pixel_values[0]
|
||||
return pixel_values, name
|
||||
|
||||
def __len__(self):
|
||||
return self.length
|
||||
|
||||
def __getitem__(self, idx):
|
||||
while True:
|
||||
try:
|
||||
pixel_values, name = self.get_batch(idx)
|
||||
break
|
||||
|
||||
except Exception as e:
|
||||
print("Error info:", e)
|
||||
idx = random.randint(0, self.length-1)
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = self.pixel_transforms(pixel_values)
|
||||
if self.enable_inpaint:
|
||||
mask = get_random_mask(pixel_values.size())
|
||||
mask_pixel_values = pixel_values * (1 - mask) + torch.ones_like(pixel_values) * -1 * mask
|
||||
sample = dict(pixel_values=pixel_values, mask_pixel_values=mask_pixel_values, mask=mask, text=name)
|
||||
else:
|
||||
sample = dict(pixel_values=pixel_values, text=name)
|
||||
return sample
|
||||
|
||||
|
||||
class VideoDataset(Dataset):
|
||||
def __init__(
|
||||
self,
|
||||
json_path, video_folder=None,
|
||||
sample_size=256, sample_stride=4, sample_n_frames=16,
|
||||
enable_bucket=False, enable_inpaint=False
|
||||
):
|
||||
print(f"loading annotations from {json_path} ...")
|
||||
self.dataset = json.load(open(json_path, 'r'))
|
||||
self.length = len(self.dataset)
|
||||
print(f"data scale: {self.length}")
|
||||
|
||||
self.video_folder = video_folder
|
||||
self.sample_stride = sample_stride
|
||||
self.sample_n_frames = sample_n_frames
|
||||
self.enable_bucket = enable_bucket
|
||||
self.enable_inpaint = enable_inpaint
|
||||
|
||||
sample_size = tuple(sample_size) if not isinstance(sample_size, int) else (sample_size, sample_size)
|
||||
self.pixel_transforms = transforms.Compose(
|
||||
[
|
||||
transforms.Resize(sample_size[0]),
|
||||
transforms.CenterCrop(sample_size),
|
||||
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5], inplace=True),
|
||||
]
|
||||
)
|
||||
|
||||
def get_batch(self, idx):
|
||||
video_dict = self.dataset[idx]
|
||||
video_id, name = video_dict['file_path'], video_dict['text']
|
||||
|
||||
if self.video_folder is None:
|
||||
video_dir = video_id
|
||||
else:
|
||||
video_dir = os.path.join(self.video_folder, video_id)
|
||||
|
||||
with VideoReader_contextmanager(video_dir, num_threads=2) as video_reader:
|
||||
video_length = len(video_reader)
|
||||
|
||||
clip_length = min(video_length, (self.sample_n_frames - 1) * self.sample_stride + 1)
|
||||
start_idx = random.randint(0, video_length - clip_length)
|
||||
batch_index = np.linspace(start_idx, start_idx + clip_length - 1, self.sample_n_frames, dtype=int)
|
||||
|
||||
try:
|
||||
sample_args = (video_reader, batch_index)
|
||||
pixel_values = func_timeout(
|
||||
VIDEO_READER_TIMEOUT, get_video_reader_batch, args=sample_args
|
||||
)
|
||||
except FunctionTimedOut:
|
||||
raise ValueError(f"Read {idx} timeout.")
|
||||
except Exception as e:
|
||||
raise ValueError(f"Failed to extract frames from video. Error is {e}.")
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = torch.from_numpy(pixel_values).permute(0, 3, 1, 2).contiguous()
|
||||
pixel_values = pixel_values / 255.
|
||||
del video_reader
|
||||
else:
|
||||
pixel_values = pixel_values
|
||||
|
||||
return pixel_values, name
|
||||
|
||||
def __len__(self):
|
||||
return self.length
|
||||
|
||||
def __getitem__(self, idx):
|
||||
while True:
|
||||
try:
|
||||
pixel_values, name = self.get_batch(idx)
|
||||
break
|
||||
|
||||
except Exception as e:
|
||||
print("Error info:", e)
|
||||
idx = random.randint(0, self.length-1)
|
||||
|
||||
if not self.enable_bucket:
|
||||
pixel_values = self.pixel_transforms(pixel_values)
|
||||
if self.enable_inpaint:
|
||||
mask = get_random_mask(pixel_values.size())
|
||||
mask_pixel_values = pixel_values * (1 - mask) + torch.ones_like(pixel_values) * -1 * mask
|
||||
sample = dict(pixel_values=pixel_values, mask_pixel_values=mask_pixel_values, mask=mask, text=name)
|
||||
else:
|
||||
sample = dict(pixel_values=pixel_values, text=name)
|
||||
return sample
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if 1:
|
||||
dataset = VideoDataset(
|
||||
json_path="/home/zhoumo.xjq/disk3/datasets/webvidval/results_2M_val.json",
|
||||
sample_size=256,
|
||||
sample_stride=4, sample_n_frames=16,
|
||||
)
|
||||
|
||||
if 0:
|
||||
dataset = WebVid10M(
|
||||
csv_path="/mnt/petrelfs/guoyuwei/projects/datasets/webvid/results_2M_val.csv",
|
||||
video_folder="/mnt/petrelfs/guoyuwei/projects/datasets/webvid/2M_val",
|
||||
sample_size=256,
|
||||
sample_stride=4, sample_n_frames=16,
|
||||
is_image=False,
|
||||
)
|
||||
|
||||
dataloader = torch.utils.data.DataLoader(dataset, batch_size=4, num_workers=0,)
|
||||
for idx, batch in enumerate(dataloader):
|
||||
print(batch["pixel_values"].shape, len(batch["text"]))
|
||||
@@ -0,0 +1,110 @@
|
||||
from typing import Optional
|
||||
import torch
|
||||
from flash_attn import flash_attn_func
|
||||
from einops import rearrange
|
||||
|
||||
from diffusers.models.attention import Attention
|
||||
from diffusers.models.embeddings import apply_rotary_emb
|
||||
|
||||
|
||||
class CogVideoXSWAAttnProcessor2_0:
|
||||
r"""
|
||||
Processor for implementing scaled dot-product attention for the CogVideoX model. It applies a rotary embedding on
|
||||
query and key vectors, but does not include spatial normalization.
|
||||
"""
|
||||
|
||||
def __init__(self, window_size=1024):
|
||||
self.window_size = window_size
|
||||
|
||||
def __call__(
|
||||
self,
|
||||
attn: Attention,
|
||||
hidden_states: torch.Tensor,
|
||||
encoder_hidden_states: torch.Tensor,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
image_rotary_emb: Optional[torch.Tensor] = None,
|
||||
num_frames: int = None,
|
||||
height: int = None,
|
||||
width: int = None,
|
||||
) -> torch.Tensor:
|
||||
text_seq_length = encoder_hidden_states.size(1)
|
||||
|
||||
hidden_states = torch.cat([encoder_hidden_states, hidden_states], dim=1)
|
||||
|
||||
batch_size, sequence_length, _ = (
|
||||
hidden_states.shape if encoder_hidden_states is None else encoder_hidden_states.shape
|
||||
)
|
||||
|
||||
if attention_mask is not None:
|
||||
attention_mask = attn.prepare_attention_mask(attention_mask, sequence_length, batch_size)
|
||||
attention_mask = attention_mask.view(batch_size, attn.heads, -1, attention_mask.shape[-1])
|
||||
|
||||
query = attn.to_q(hidden_states)
|
||||
key = attn.to_k(hidden_states)
|
||||
value = attn.to_v(hidden_states)
|
||||
|
||||
inner_dim = key.shape[-1]
|
||||
head_dim = inner_dim // attn.heads
|
||||
|
||||
query = query.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
|
||||
key = key.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
|
||||
value = value.view(batch_size, -1, attn.heads, head_dim) # .transpose(1, 2)
|
||||
|
||||
if attn.norm_q is not None:
|
||||
query = attn.norm_q(query)
|
||||
if attn.norm_k is not None:
|
||||
key = attn.norm_k(key)
|
||||
|
||||
# Apply RoPE if needed
|
||||
if image_rotary_emb is not None:
|
||||
|
||||
query[:, :, text_seq_length:] = apply_rotary_emb(query[:, :, text_seq_length:], image_rotary_emb)
|
||||
if not attn.is_cross_attention:
|
||||
key[:, :, text_seq_length:] = apply_rotary_emb(key[:, :, text_seq_length:], image_rotary_emb)
|
||||
|
||||
query = query.transpose(1, 2).to(value)
|
||||
key = key.transpose(1, 2).to(value)
|
||||
|
||||
interval = max((query.size(1) - text_seq_length) // (self.window_size - 256), 1)
|
||||
cross_key = torch.cat([key[:, :text_seq_length], key[:, text_seq_length::interval]], dim=1)
|
||||
cross_val = torch.cat([value[:, :text_seq_length], value[:, text_seq_length::interval]], dim=1)
|
||||
cross_hidden_states = flash_attn_func(query, cross_key, cross_val, dropout_p=0.0, causal=False)
|
||||
query_txt = query[:, :text_seq_length]
|
||||
key_txt = key[:, :text_seq_length]
|
||||
value_txt = value[:, :text_seq_length]
|
||||
querys = torch.tensor_split(query[:, text_seq_length:], 6, 2)
|
||||
keys = torch.tensor_split(key[:, text_seq_length:], 6, 2)
|
||||
values = torch.tensor_split(value[:, text_seq_length:], 6, 2)
|
||||
new_querys = [querys[0]]
|
||||
new_keys = [keys[0]]
|
||||
new_values = [values[0]]
|
||||
for index, mode in enumerate(["bs (f h w) hn hd -> bs (f w h) hn hd", "bs (f h w) hn hd -> bs (h f w) hn hd", "bs (f h w) hn hd -> bs (h w f) hn hd",
|
||||
"bs (f h w) hn hd -> bs (w f h) hn hd", "bs (f h w) hn hd -> bs (w h f) hn hd"]):
|
||||
new_querys.append(rearrange(querys[index + 1], mode, f=num_frames, h=height, w=width))
|
||||
new_keys.append(rearrange(keys[index + 1], mode, f=num_frames, h=height, w=width))
|
||||
new_values.append(rearrange(values[index + 1], mode, f=num_frames, h=height, w=width))
|
||||
query = torch.cat([query_txt, torch.cat(new_querys, dim=2)], dim=1)
|
||||
key = torch.cat([key_txt, torch.cat(new_keys, dim=2)], dim=1)
|
||||
value = torch.cat([value_txt, torch.cat(new_values, dim=2)], dim=1)
|
||||
|
||||
hidden_states = flash_attn_func(query, key, value, dropout_p=0.0, causal=False, window_size=(self.window_size, self.window_size))
|
||||
hidden_states_txt = hidden_states[:, :text_seq_length]
|
||||
hidden_states = torch.tensor_split(hidden_states[:, text_seq_length:], 6, 2)
|
||||
new_hidden_states = [hidden_states[0]]
|
||||
for index, mode in enumerate(["bs (f w h) hn hd -> bs (f h w) hn hd", "bs (h f w) hn hd -> bs (f h w) hn hd", "bs (h w f) hn hd -> bs (f h w) hn hd",
|
||||
"bs (w f h) hn hd -> bs (f h w) hn hd", "bs (w h f) hn hd -> bs (f h w) hn hd"]):
|
||||
new_hidden_states.append(rearrange(hidden_states[index + 1], mode, f=num_frames, h=height, w=width))
|
||||
hidden_states = torch.cat([hidden_states_txt, torch.cat(new_hidden_states, dim=2)], dim=1) + cross_hidden_states
|
||||
|
||||
|
||||
hidden_states = hidden_states.reshape(batch_size, -1, attn.heads * head_dim)
|
||||
|
||||
# linear proj
|
||||
hidden_states = attn.to_out[0](hidden_states)
|
||||
# dropout
|
||||
hidden_states = attn.to_out[1](hidden_states)
|
||||
|
||||
encoder_hidden_states, hidden_states = hidden_states.split(
|
||||
[text_seq_length, hidden_states.size(1) - text_seq_length], dim=1
|
||||
)
|
||||
return hidden_states, encoder_hidden_states
|
||||
@@ -13,14 +13,12 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
from typing import Dict, Optional, Tuple, Union
|
||||
from typing import Optional, Tuple, Union
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
import json
|
||||
import os
|
||||
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.loaders.single_file_model import FromOriginalModelMixin
|
||||
@@ -43,9 +41,7 @@ class CogVideoXSafeConv3d(nn.Conv3d):
|
||||
"""
|
||||
|
||||
def forward(self, input: torch.Tensor) -> torch.Tensor:
|
||||
memory_count = (
|
||||
(input.shape[0] * input.shape[1] * input.shape[2] * input.shape[3] * input.shape[4]) * 2 / 1024**3
|
||||
)
|
||||
memory_count = torch.prod(torch.tensor(input.shape)).item() * 2 / 1024**3
|
||||
|
||||
# Set to 2GB, suitable for CuDNN
|
||||
if memory_count > 2:
|
||||
@@ -96,13 +92,11 @@ class CogVideoXCausalConv3d(nn.Module):
|
||||
|
||||
time_kernel_size, height_kernel_size, width_kernel_size = kernel_size
|
||||
|
||||
# TODO(aryan): configure calculation based on stride and dilation in the future.
|
||||
# Since CogVideoX does not use it, it is currently tailored to "just work" with Mochi
|
||||
time_pad = time_kernel_size - 1
|
||||
height_pad = (height_kernel_size - 1) // 2
|
||||
width_pad = (width_kernel_size - 1) // 2
|
||||
|
||||
self.pad_mode = pad_mode
|
||||
time_pad = dilation * (time_kernel_size - 1) + (1 - stride)
|
||||
height_pad = height_kernel_size // 2
|
||||
width_pad = width_kernel_size // 2
|
||||
|
||||
self.height_pad = height_pad
|
||||
self.width_pad = width_pad
|
||||
self.time_pad = time_pad
|
||||
@@ -111,7 +105,7 @@ class CogVideoXCausalConv3d(nn.Module):
|
||||
self.temporal_dim = 2
|
||||
self.time_kernel_size = time_kernel_size
|
||||
|
||||
stride = stride if isinstance(stride, tuple) else (stride, 1, 1)
|
||||
stride = (stride, 1, 1)
|
||||
dilation = (dilation, 1, 1)
|
||||
self.conv = CogVideoXSafeConv3d(
|
||||
in_channels=in_channels,
|
||||
@@ -121,30 +115,34 @@ class CogVideoXCausalConv3d(nn.Module):
|
||||
dilation=dilation,
|
||||
)
|
||||
|
||||
def fake_context_parallel_forward(
|
||||
self, inputs: torch.Tensor, conv_cache: Optional[torch.Tensor] = None
|
||||
) -> torch.Tensor:
|
||||
if self.pad_mode == "replicate":
|
||||
inputs = F.pad(inputs, self.time_causal_padding, mode="replicate")
|
||||
else:
|
||||
kernel_size = self.time_kernel_size
|
||||
if kernel_size > 1:
|
||||
cached_inputs = [conv_cache] if conv_cache is not None else [inputs[:, :, :1]] * (kernel_size - 1)
|
||||
inputs = torch.cat(cached_inputs + [inputs], dim=2)
|
||||
self.conv_cache = None
|
||||
|
||||
def fake_context_parallel_forward(self, inputs: torch.Tensor) -> torch.Tensor:
|
||||
kernel_size = self.time_kernel_size
|
||||
if kernel_size > 1:
|
||||
cached_inputs = (
|
||||
[self.conv_cache] if self.conv_cache is not None else [inputs[:, :, :1]] * (kernel_size - 1)
|
||||
)
|
||||
inputs = torch.cat(cached_inputs + [inputs], dim=2)
|
||||
return inputs
|
||||
|
||||
def forward(self, inputs: torch.Tensor, conv_cache: Optional[torch.Tensor] = None) -> torch.Tensor:
|
||||
inputs = self.fake_context_parallel_forward(inputs, conv_cache)
|
||||
def _clear_fake_context_parallel_cache(self):
|
||||
del self.conv_cache
|
||||
self.conv_cache = None
|
||||
|
||||
if self.pad_mode == "replicate":
|
||||
conv_cache = None
|
||||
else:
|
||||
padding_2d = (self.width_pad, self.width_pad, self.height_pad, self.height_pad)
|
||||
conv_cache = inputs[:, :, -self.time_kernel_size + 1 :].clone()
|
||||
inputs = F.pad(inputs, padding_2d, mode="constant", value=0)
|
||||
def forward(self, inputs: torch.Tensor) -> torch.Tensor:
|
||||
inputs = self.fake_context_parallel_forward(inputs)
|
||||
|
||||
self._clear_fake_context_parallel_cache()
|
||||
# Note: we could move these to the cpu for a lower maximum memory usage but its only a few
|
||||
# hundred megabytes and so let's not do it for now
|
||||
self.conv_cache = inputs[:, :, -self.time_kernel_size + 1 :].clone()
|
||||
|
||||
padding_2d = (self.width_pad, self.width_pad, self.height_pad, self.height_pad)
|
||||
inputs = F.pad(inputs, padding_2d, mode="constant", value=0)
|
||||
|
||||
output = self.conv(inputs)
|
||||
return output, conv_cache
|
||||
return output
|
||||
|
||||
|
||||
class CogVideoXSpatialNorm3D(nn.Module):
|
||||
@@ -174,12 +172,7 @@ class CogVideoXSpatialNorm3D(nn.Module):
|
||||
self.conv_y = CogVideoXCausalConv3d(zq_channels, f_channels, kernel_size=1, stride=1)
|
||||
self.conv_b = CogVideoXCausalConv3d(zq_channels, f_channels, kernel_size=1, stride=1)
|
||||
|
||||
def forward(
|
||||
self, f: torch.Tensor, zq: torch.Tensor, conv_cache: Optional[Dict[str, torch.Tensor]] = None
|
||||
) -> torch.Tensor:
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
def forward(self, f: torch.Tensor, zq: torch.Tensor) -> torch.Tensor:
|
||||
if f.shape[2] > 1 and f.shape[2] % 2 == 1:
|
||||
f_first, f_rest = f[:, :, :1], f[:, :, 1:]
|
||||
f_first_size, f_rest_size = f_first.shape[-3:], f_rest.shape[-3:]
|
||||
@@ -190,87 +183,9 @@ class CogVideoXSpatialNorm3D(nn.Module):
|
||||
else:
|
||||
zq = F.interpolate(zq, size=f.shape[-3:])
|
||||
|
||||
conv_y, new_conv_cache["conv_y"] = self.conv_y(zq, conv_cache=conv_cache.get("conv_y"))
|
||||
conv_b, new_conv_cache["conv_b"] = self.conv_b(zq, conv_cache=conv_cache.get("conv_b"))
|
||||
|
||||
norm_f = self.norm_layer(f)
|
||||
new_f = norm_f * conv_y + conv_b
|
||||
return new_f, new_conv_cache
|
||||
|
||||
|
||||
class CogVideoXUpsample3D(nn.Module):
|
||||
r"""
|
||||
A 3D Upsample layer using in CogVideoX by Tsinghua University & ZhipuAI # Todo: Wait for paper relase.
|
||||
|
||||
Args:
|
||||
in_channels (`int`):
|
||||
Number of channels in the input image.
|
||||
out_channels (`int`):
|
||||
Number of channels produced by the convolution.
|
||||
kernel_size (`int`, defaults to `3`):
|
||||
Size of the convolving kernel.
|
||||
stride (`int`, defaults to `1`):
|
||||
Stride of the convolution.
|
||||
padding (`int`, defaults to `1`):
|
||||
Padding added to all four sides of the input.
|
||||
compress_time (`bool`, defaults to `False`):
|
||||
Whether or not to compress the time dimension.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
kernel_size: int = 3,
|
||||
stride: int = 1,
|
||||
padding: int = 1,
|
||||
compress_time: bool = False,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
|
||||
self.conv = nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size, stride=stride, padding=padding)
|
||||
self.compress_time = compress_time
|
||||
|
||||
self.auto_split_process = True
|
||||
self.first_frame_flag = False
|
||||
|
||||
def forward(self, inputs: torch.Tensor) -> torch.Tensor:
|
||||
if self.compress_time:
|
||||
if self.auto_split_process:
|
||||
if inputs.shape[2] > 1 and inputs.shape[2] % 2 == 1:
|
||||
# split first frame
|
||||
x_first, x_rest = inputs[:, :, 0], inputs[:, :, 1:]
|
||||
|
||||
x_first = F.interpolate(x_first, scale_factor=2.0)
|
||||
x_rest = F.interpolate(x_rest, scale_factor=2.0)
|
||||
x_first = x_first[:, :, None, :, :]
|
||||
inputs = torch.cat([x_first, x_rest], dim=2)
|
||||
elif inputs.shape[2] > 1:
|
||||
inputs = F.interpolate(inputs, scale_factor=2.0)
|
||||
else:
|
||||
inputs = inputs.squeeze(2)
|
||||
inputs = F.interpolate(inputs, scale_factor=2.0)
|
||||
inputs = inputs[:, :, None, :, :]
|
||||
else:
|
||||
if self.first_frame_flag:
|
||||
inputs = inputs.squeeze(2)
|
||||
inputs = F.interpolate(inputs, scale_factor=2.0)
|
||||
inputs = inputs[:, :, None, :, :]
|
||||
else:
|
||||
inputs = F.interpolate(inputs, scale_factor=2.0)
|
||||
else:
|
||||
# only interpolate 2D
|
||||
b, c, t, h, w = inputs.shape
|
||||
inputs = inputs.permute(0, 2, 1, 3, 4).reshape(b * t, c, h, w)
|
||||
inputs = F.interpolate(inputs, scale_factor=2.0)
|
||||
inputs = inputs.reshape(b, t, c, *inputs.shape[2:]).permute(0, 2, 1, 3, 4)
|
||||
|
||||
b, c, t, h, w = inputs.shape
|
||||
inputs = inputs.permute(0, 2, 1, 3, 4).reshape(b * t, c, h, w)
|
||||
inputs = self.conv(inputs)
|
||||
inputs = inputs.reshape(b, t, *inputs.shape[1:]).permute(0, 2, 1, 3, 4)
|
||||
|
||||
return inputs
|
||||
new_f = norm_f * self.conv_y(zq) + self.conv_b(zq)
|
||||
return new_f
|
||||
|
||||
|
||||
class CogVideoXResnetBlock3D(nn.Module):
|
||||
@@ -321,7 +236,6 @@ class CogVideoXResnetBlock3D(nn.Module):
|
||||
self.out_channels = out_channels
|
||||
self.nonlinearity = get_activation(non_linearity)
|
||||
self.use_conv_shortcut = conv_shortcut
|
||||
self.spatial_norm_dim = spatial_norm_dim
|
||||
|
||||
if spatial_norm_dim is None:
|
||||
self.norm1 = nn.GroupNorm(num_channels=in_channels, num_groups=groups, eps=eps)
|
||||
@@ -365,43 +279,34 @@ class CogVideoXResnetBlock3D(nn.Module):
|
||||
inputs: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
zq: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
hidden_states = inputs
|
||||
|
||||
if zq is not None:
|
||||
hidden_states, new_conv_cache["norm1"] = self.norm1(hidden_states, zq, conv_cache=conv_cache.get("norm1"))
|
||||
hidden_states = self.norm1(hidden_states, zq)
|
||||
else:
|
||||
hidden_states = self.norm1(hidden_states)
|
||||
|
||||
hidden_states = self.nonlinearity(hidden_states)
|
||||
hidden_states, new_conv_cache["conv1"] = self.conv1(hidden_states, conv_cache=conv_cache.get("conv1"))
|
||||
hidden_states = self.conv1(hidden_states)
|
||||
|
||||
if temb is not None:
|
||||
hidden_states = hidden_states + self.temb_proj(self.nonlinearity(temb))[:, :, None, None, None]
|
||||
|
||||
if zq is not None:
|
||||
hidden_states, new_conv_cache["norm2"] = self.norm2(hidden_states, zq, conv_cache=conv_cache.get("norm2"))
|
||||
hidden_states = self.norm2(hidden_states, zq)
|
||||
else:
|
||||
hidden_states = self.norm2(hidden_states)
|
||||
|
||||
hidden_states = self.nonlinearity(hidden_states)
|
||||
hidden_states = self.dropout(hidden_states)
|
||||
hidden_states, new_conv_cache["conv2"] = self.conv2(hidden_states, conv_cache=conv_cache.get("conv2"))
|
||||
hidden_states = self.conv2(hidden_states)
|
||||
|
||||
if self.in_channels != self.out_channels:
|
||||
if self.use_conv_shortcut:
|
||||
inputs, new_conv_cache["conv_shortcut"] = self.conv_shortcut(
|
||||
inputs, conv_cache=conv_cache.get("conv_shortcut")
|
||||
)
|
||||
else:
|
||||
inputs = self.conv_shortcut(inputs)
|
||||
inputs = self.conv_shortcut(inputs)
|
||||
|
||||
hidden_states = hidden_states + inputs
|
||||
return hidden_states, new_conv_cache
|
||||
return hidden_states
|
||||
|
||||
|
||||
class CogVideoXDownBlock3D(nn.Module):
|
||||
@@ -487,17 +392,9 @@ class CogVideoXDownBlock3D(nn.Module):
|
||||
hidden_states: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
zq: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
r"""Forward method of the `CogVideoXDownBlock3D` class."""
|
||||
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
for i, resnet in enumerate(self.resnets):
|
||||
conv_cache_key = f"resnet_{i}"
|
||||
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
for resnet in self.resnets:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def create_forward(*inputs):
|
||||
@@ -505,23 +402,17 @@ class CogVideoXDownBlock3D(nn.Module):
|
||||
|
||||
return create_forward
|
||||
|
||||
hidden_states, new_conv_cache[conv_cache_key] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet),
|
||||
hidden_states,
|
||||
temb,
|
||||
zq,
|
||||
conv_cache.get(conv_cache_key),
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet), hidden_states, temb, zq
|
||||
)
|
||||
else:
|
||||
hidden_states, new_conv_cache[conv_cache_key] = resnet(
|
||||
hidden_states, temb, zq, conv_cache=conv_cache.get(conv_cache_key)
|
||||
)
|
||||
hidden_states = resnet(hidden_states, temb, zq)
|
||||
|
||||
if self.downsamplers is not None:
|
||||
for downsampler in self.downsamplers:
|
||||
hidden_states = downsampler(hidden_states)
|
||||
|
||||
return hidden_states, new_conv_cache
|
||||
return hidden_states
|
||||
|
||||
|
||||
class CogVideoXMidBlock3D(nn.Module):
|
||||
@@ -589,17 +480,9 @@ class CogVideoXMidBlock3D(nn.Module):
|
||||
hidden_states: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
zq: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
r"""Forward method of the `CogVideoXMidBlock3D` class."""
|
||||
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
for i, resnet in enumerate(self.resnets):
|
||||
conv_cache_key = f"resnet_{i}"
|
||||
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
for resnet in self.resnets:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def create_forward(*inputs):
|
||||
@@ -607,15 +490,13 @@ class CogVideoXMidBlock3D(nn.Module):
|
||||
|
||||
return create_forward
|
||||
|
||||
hidden_states, new_conv_cache[conv_cache_key] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet), hidden_states, temb, zq, conv_cache.get(conv_cache_key)
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet), hidden_states, temb, zq
|
||||
)
|
||||
else:
|
||||
hidden_states, new_conv_cache[conv_cache_key] = resnet(
|
||||
hidden_states, temb, zq, conv_cache=conv_cache.get(conv_cache_key)
|
||||
)
|
||||
hidden_states = resnet(hidden_states, temb, zq)
|
||||
|
||||
return hidden_states, new_conv_cache
|
||||
return hidden_states
|
||||
|
||||
|
||||
class CogVideoXUpBlock3D(nn.Module):
|
||||
@@ -703,17 +584,10 @@ class CogVideoXUpBlock3D(nn.Module):
|
||||
hidden_states: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
zq: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
r"""Forward method of the `CogVideoXUpBlock3D` class."""
|
||||
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
for i, resnet in enumerate(self.resnets):
|
||||
conv_cache_key = f"resnet_{i}"
|
||||
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
for resnet in self.resnets:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def create_forward(*inputs):
|
||||
@@ -721,23 +595,17 @@ class CogVideoXUpBlock3D(nn.Module):
|
||||
|
||||
return create_forward
|
||||
|
||||
hidden_states, new_conv_cache[conv_cache_key] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet),
|
||||
hidden_states,
|
||||
temb,
|
||||
zq,
|
||||
conv_cache.get(conv_cache_key),
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(resnet), hidden_states, temb, zq
|
||||
)
|
||||
else:
|
||||
hidden_states, new_conv_cache[conv_cache_key] = resnet(
|
||||
hidden_states, temb, zq, conv_cache=conv_cache.get(conv_cache_key)
|
||||
)
|
||||
hidden_states = resnet(hidden_states, temb, zq)
|
||||
|
||||
if self.upsamplers is not None:
|
||||
for upsampler in self.upsamplers:
|
||||
hidden_states = upsampler(hidden_states)
|
||||
|
||||
return hidden_states, new_conv_cache
|
||||
return hidden_states
|
||||
|
||||
|
||||
class CogVideoXEncoder3D(nn.Module):
|
||||
@@ -837,20 +705,11 @@ class CogVideoXEncoder3D(nn.Module):
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
def forward(
|
||||
self,
|
||||
sample: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
def forward(self, sample: torch.Tensor, temb: Optional[torch.Tensor] = None) -> torch.Tensor:
|
||||
r"""The forward method of the `CogVideoXEncoder3D` class."""
|
||||
hidden_states = self.conv_in(sample)
|
||||
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
hidden_states, new_conv_cache["conv_in"] = self.conv_in(sample, conv_cache=conv_cache.get("conv_in"))
|
||||
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def custom_forward(*inputs):
|
||||
@@ -859,44 +718,28 @@ class CogVideoXEncoder3D(nn.Module):
|
||||
return custom_forward
|
||||
|
||||
# 1. Down
|
||||
for i, down_block in enumerate(self.down_blocks):
|
||||
conv_cache_key = f"down_block_{i}"
|
||||
hidden_states, new_conv_cache[conv_cache_key] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(down_block),
|
||||
hidden_states,
|
||||
temb,
|
||||
None,
|
||||
conv_cache.get(conv_cache_key),
|
||||
for down_block in self.down_blocks:
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(down_block), hidden_states, temb, None
|
||||
)
|
||||
|
||||
# 2. Mid
|
||||
hidden_states, new_conv_cache["mid_block"] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block),
|
||||
hidden_states,
|
||||
temb,
|
||||
None,
|
||||
conv_cache.get("mid_block"),
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block), hidden_states, temb, None
|
||||
)
|
||||
else:
|
||||
# 1. Down
|
||||
for i, down_block in enumerate(self.down_blocks):
|
||||
conv_cache_key = f"down_block_{i}"
|
||||
hidden_states, new_conv_cache[conv_cache_key] = down_block(
|
||||
hidden_states, temb, None, conv_cache=conv_cache.get(conv_cache_key)
|
||||
)
|
||||
for down_block in self.down_blocks:
|
||||
hidden_states = down_block(hidden_states, temb, None)
|
||||
|
||||
# 2. Mid
|
||||
hidden_states, new_conv_cache["mid_block"] = self.mid_block(
|
||||
hidden_states, temb, None, conv_cache=conv_cache.get("mid_block")
|
||||
)
|
||||
hidden_states = self.mid_block(hidden_states, temb, None)
|
||||
|
||||
# 3. Post-process
|
||||
hidden_states = self.norm_out(hidden_states)
|
||||
hidden_states = self.conv_act(hidden_states)
|
||||
|
||||
hidden_states, new_conv_cache["conv_out"] = self.conv_out(hidden_states, conv_cache=conv_cache.get("conv_out"))
|
||||
|
||||
return hidden_states, new_conv_cache
|
||||
hidden_states = self.conv_out(hidden_states)
|
||||
return hidden_states
|
||||
|
||||
|
||||
class CogVideoXDecoder3D(nn.Module):
|
||||
@@ -1003,20 +846,11 @@ class CogVideoXDecoder3D(nn.Module):
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
def forward(
|
||||
self,
|
||||
sample: torch.Tensor,
|
||||
temb: Optional[torch.Tensor] = None,
|
||||
conv_cache: Optional[Dict[str, torch.Tensor]] = None,
|
||||
) -> torch.Tensor:
|
||||
def forward(self, sample: torch.Tensor, temb: Optional[torch.Tensor] = None) -> torch.Tensor:
|
||||
r"""The forward method of the `CogVideoXDecoder3D` class."""
|
||||
hidden_states = self.conv_in(sample)
|
||||
|
||||
new_conv_cache = {}
|
||||
conv_cache = conv_cache or {}
|
||||
|
||||
hidden_states, new_conv_cache["conv_in"] = self.conv_in(sample, conv_cache=conv_cache.get("conv_in"))
|
||||
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def custom_forward(*inputs):
|
||||
@@ -1025,45 +859,28 @@ class CogVideoXDecoder3D(nn.Module):
|
||||
return custom_forward
|
||||
|
||||
# 1. Mid
|
||||
hidden_states, new_conv_cache["mid_block"] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block),
|
||||
hidden_states,
|
||||
temb,
|
||||
sample,
|
||||
conv_cache.get("mid_block"),
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block), hidden_states, temb, sample
|
||||
)
|
||||
|
||||
# 2. Up
|
||||
for i, up_block in enumerate(self.up_blocks):
|
||||
conv_cache_key = f"up_block_{i}"
|
||||
hidden_states, new_conv_cache[conv_cache_key] = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(up_block),
|
||||
hidden_states,
|
||||
temb,
|
||||
sample,
|
||||
conv_cache.get(conv_cache_key),
|
||||
for up_block in self.up_blocks:
|
||||
hidden_states = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(up_block), hidden_states, temb, sample
|
||||
)
|
||||
else:
|
||||
# 1. Mid
|
||||
hidden_states, new_conv_cache["mid_block"] = self.mid_block(
|
||||
hidden_states, temb, sample, conv_cache=conv_cache.get("mid_block")
|
||||
)
|
||||
hidden_states = self.mid_block(hidden_states, temb, sample)
|
||||
|
||||
# 2. Up
|
||||
for i, up_block in enumerate(self.up_blocks):
|
||||
conv_cache_key = f"up_block_{i}"
|
||||
hidden_states, new_conv_cache[conv_cache_key] = up_block(
|
||||
hidden_states, temb, sample, conv_cache=conv_cache.get(conv_cache_key)
|
||||
)
|
||||
for up_block in self.up_blocks:
|
||||
hidden_states = up_block(hidden_states, temb, sample)
|
||||
|
||||
# 3. Post-process
|
||||
hidden_states, new_conv_cache["norm_out"] = self.norm_out(
|
||||
hidden_states, sample, conv_cache=conv_cache.get("norm_out")
|
||||
)
|
||||
hidden_states = self.norm_out(hidden_states, sample)
|
||||
hidden_states = self.conv_act(hidden_states)
|
||||
hidden_states, new_conv_cache["conv_out"] = self.conv_out(hidden_states, conv_cache=conv_cache.get("conv_out"))
|
||||
|
||||
return hidden_states, new_conv_cache
|
||||
hidden_states = self.conv_out(hidden_states)
|
||||
return hidden_states
|
||||
|
||||
|
||||
class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
@@ -1134,7 +951,6 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
force_upcast: float = True,
|
||||
use_quant_conv: bool = False,
|
||||
use_post_quant_conv: bool = False,
|
||||
invert_scale_latents: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
@@ -1165,7 +981,6 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
|
||||
self.use_slicing = False
|
||||
self.use_tiling = False
|
||||
self.auto_split_process = False
|
||||
|
||||
# Can be increased to decode more latent frames at once, but comes at a reasonable memory cost and it is not
|
||||
# recommended because the temporal parts of the VAE, here, are tricky to understand.
|
||||
@@ -1184,7 +999,6 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
# setting it to anything other than 2 would give poor results because the VAE hasn't been trained to be adaptive with different
|
||||
# number of temporal frames.
|
||||
self.num_latent_frames_batch_size = 2
|
||||
self.num_sample_frames_batch_size = 8
|
||||
|
||||
# We make the minimum height and width of sample for tiling half that of the generally supported
|
||||
self.tile_sample_min_height = sample_height // 2
|
||||
@@ -1204,6 +1018,12 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
if isinstance(module, (CogVideoXEncoder3D, CogVideoXDecoder3D)):
|
||||
module.gradient_checkpointing = value
|
||||
|
||||
def _clear_fake_context_parallel_cache(self):
|
||||
for name, module in self.named_modules():
|
||||
if isinstance(module, CogVideoXCausalConv3d):
|
||||
logger.debug(f"Clearing fake Context Parallel cache for layer: {name}")
|
||||
module._clear_fake_context_parallel_cache()
|
||||
|
||||
def enable_tiling(
|
||||
self,
|
||||
tile_sample_min_height: Optional[int] = None,
|
||||
@@ -1260,53 +1080,6 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
decoding in one step.
|
||||
"""
|
||||
self.use_slicing = False
|
||||
|
||||
def _set_first_frame(self):
|
||||
for name, module in self.named_modules():
|
||||
if isinstance(module, CogVideoXUpsample3D):
|
||||
module.auto_split_process = False
|
||||
module.first_frame_flag = True
|
||||
|
||||
def _set_rest_frame(self):
|
||||
for name, module in self.named_modules():
|
||||
if isinstance(module, CogVideoXUpsample3D):
|
||||
module.auto_split_process = False
|
||||
module.first_frame_flag = False
|
||||
|
||||
def enable_auto_split_process(self) -> None:
|
||||
self.auto_split_process = True
|
||||
for name, module in self.named_modules():
|
||||
if isinstance(module, CogVideoXUpsample3D):
|
||||
module.auto_split_process = True
|
||||
|
||||
def disable_auto_split_process(self) -> None:
|
||||
self.auto_split_process = False
|
||||
|
||||
def _encode(self, x: torch.Tensor) -> torch.Tensor:
|
||||
batch_size, num_channels, num_frames, height, width = x.shape
|
||||
|
||||
if self.use_tiling and (width > self.tile_sample_min_width or height > self.tile_sample_min_height):
|
||||
return self.tiled_encode(x)
|
||||
|
||||
frame_batch_size = self.num_sample_frames_batch_size
|
||||
# Note: We expect the number of frames to be either `1` or `frame_batch_size * k` or `frame_batch_size * k + 1` for some k.
|
||||
# As the extra single frame is handled inside the loop, it is not required to round up here.
|
||||
num_batches = max(num_frames // frame_batch_size, 1)
|
||||
conv_cache = None
|
||||
enc = []
|
||||
|
||||
for i in range(num_batches):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * i + (0 if i == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (i + 1) + remaining_frames
|
||||
x_intermediate = x[:, :, start_frame:end_frame]
|
||||
x_intermediate, conv_cache = self.encoder(x_intermediate, conv_cache=conv_cache)
|
||||
if self.quant_conv is not None:
|
||||
x_intermediate = self.quant_conv(x_intermediate)
|
||||
enc.append(x_intermediate)
|
||||
|
||||
enc = torch.cat(enc, dim=2)
|
||||
return enc
|
||||
|
||||
@apply_forward_hook
|
||||
def encode(
|
||||
@@ -1321,17 +1094,30 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
Whether to return a [`~models.autoencoder_kl.AutoencoderKLOutput`] instead of a plain tuple.
|
||||
|
||||
Returns:
|
||||
The latent representations of the encoded videos. If `return_dict` is True, a
|
||||
The latent representations of the encoded images. If `return_dict` is True, a
|
||||
[`~models.autoencoder_kl.AutoencoderKLOutput`] is returned, otherwise a plain `tuple` is returned.
|
||||
"""
|
||||
if self.use_slicing and x.shape[0] > 1:
|
||||
encoded_slices = [self._encode(x_slice) for x_slice in x.split(1)]
|
||||
h = torch.cat(encoded_slices)
|
||||
batch_size, num_channels, num_frames, height, width = x.shape
|
||||
if num_frames == 1:
|
||||
h = self.encoder(x)
|
||||
if self.quant_conv is not None:
|
||||
h = self.quant_conv(h)
|
||||
posterior = DiagonalGaussianDistribution(h)
|
||||
else:
|
||||
h = self._encode(x)
|
||||
|
||||
posterior = DiagonalGaussianDistribution(h)
|
||||
|
||||
frame_batch_size = 4
|
||||
h = []
|
||||
for i in range(num_frames // frame_batch_size):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * i + (0 if i == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (i + 1) + remaining_frames
|
||||
z_intermediate = x[:, :, start_frame:end_frame]
|
||||
z_intermediate = self.encoder(z_intermediate)
|
||||
if self.quant_conv is not None:
|
||||
z_intermediate = self.quant_conv(z_intermediate)
|
||||
h.append(z_intermediate)
|
||||
self._clear_fake_context_parallel_cache()
|
||||
h = torch.cat(h, dim=2)
|
||||
posterior = DiagonalGaussianDistribution(h)
|
||||
if not return_dict:
|
||||
return (posterior,)
|
||||
return AutoencoderKLOutput(latent_dist=posterior)
|
||||
@@ -1342,47 +1128,27 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
if self.use_tiling and (width > self.tile_latent_min_width or height > self.tile_latent_min_height):
|
||||
return self.tiled_decode(z, return_dict=return_dict)
|
||||
|
||||
if self.auto_split_process:
|
||||
frame_batch_size = self.num_latent_frames_batch_size
|
||||
num_batches = max(num_frames // frame_batch_size, 1)
|
||||
conv_cache = None
|
||||
if num_frames == 1:
|
||||
dec = []
|
||||
|
||||
for i in range(num_batches):
|
||||
z_intermediate = z
|
||||
if self.post_quant_conv is not None:
|
||||
z_intermediate = self.post_quant_conv(z_intermediate)
|
||||
z_intermediate = self.decoder(z_intermediate)
|
||||
dec.append(z_intermediate)
|
||||
else:
|
||||
frame_batch_size = self.num_latent_frames_batch_size
|
||||
dec = []
|
||||
for i in range(num_frames // frame_batch_size):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * i + (0 if i == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (i + 1) + remaining_frames
|
||||
z_intermediate = z[:, :, start_frame:end_frame]
|
||||
if self.post_quant_conv is not None:
|
||||
z_intermediate = self.post_quant_conv(z_intermediate)
|
||||
z_intermediate, conv_cache = self.decoder(z_intermediate, conv_cache=conv_cache)
|
||||
z_intermediate = self.decoder(z_intermediate)
|
||||
dec.append(z_intermediate)
|
||||
else:
|
||||
conv_cache = None
|
||||
start_frame = 0
|
||||
end_frame = 1
|
||||
dec = []
|
||||
|
||||
self._set_first_frame()
|
||||
z_intermediate = z[:, :, start_frame:end_frame]
|
||||
if self.post_quant_conv is not None:
|
||||
z_intermediate = self.post_quant_conv(z_intermediate)
|
||||
z_intermediate, conv_cache = self.decoder(z_intermediate, conv_cache=conv_cache)
|
||||
dec.append(z_intermediate)
|
||||
|
||||
self._set_rest_frame()
|
||||
start_frame = end_frame
|
||||
end_frame += self.num_latent_frames_batch_size
|
||||
|
||||
while start_frame < num_frames:
|
||||
z_intermediate = z[:, :, start_frame:end_frame]
|
||||
if self.post_quant_conv is not None:
|
||||
z_intermediate = self.post_quant_conv(z_intermediate)
|
||||
z_intermediate, conv_cache = self.decoder(z_intermediate, conv_cache=conv_cache)
|
||||
dec.append(z_intermediate)
|
||||
start_frame = end_frame
|
||||
end_frame += self.num_latent_frames_batch_size
|
||||
|
||||
self._clear_fake_context_parallel_cache()
|
||||
dec = torch.cat(dec, dim=2)
|
||||
|
||||
if not return_dict:
|
||||
@@ -1431,80 +1197,6 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
)
|
||||
return b
|
||||
|
||||
def tiled_encode(self, x: torch.Tensor) -> torch.Tensor:
|
||||
r"""Encode a batch of images using a tiled encoder.
|
||||
|
||||
When this option is enabled, the VAE will split the input tensor into tiles to compute encoding in several
|
||||
steps. This is useful to keep memory use constant regardless of image size. The end result of tiled encoding is
|
||||
different from non-tiled encoding because each tile uses a different encoder. To avoid tiling artifacts, the
|
||||
tiles overlap and are blended together to form a smooth output. You may still see tile-sized changes in the
|
||||
output, but they should be much less noticeable.
|
||||
|
||||
Args:
|
||||
x (`torch.Tensor`): Input batch of videos.
|
||||
|
||||
Returns:
|
||||
`torch.Tensor`:
|
||||
The latent representation of the encoded videos.
|
||||
"""
|
||||
# For a rough memory estimate, take a look at the `tiled_decode` method.
|
||||
batch_size, num_channels, num_frames, height, width = x.shape
|
||||
|
||||
overlap_height = int(self.tile_sample_min_height * (1 - self.tile_overlap_factor_height))
|
||||
overlap_width = int(self.tile_sample_min_width * (1 - self.tile_overlap_factor_width))
|
||||
blend_extent_height = int(self.tile_latent_min_height * self.tile_overlap_factor_height)
|
||||
blend_extent_width = int(self.tile_latent_min_width * self.tile_overlap_factor_width)
|
||||
row_limit_height = self.tile_latent_min_height - blend_extent_height
|
||||
row_limit_width = self.tile_latent_min_width - blend_extent_width
|
||||
frame_batch_size = self.num_sample_frames_batch_size
|
||||
|
||||
# Split x into overlapping tiles and encode them separately.
|
||||
# The tiles have an overlap to avoid seams between tiles.
|
||||
rows = []
|
||||
for i in range(0, height, overlap_height):
|
||||
row = []
|
||||
for j in range(0, width, overlap_width):
|
||||
# Note: We expect the number of frames to be either `1` or `frame_batch_size * k` or `frame_batch_size * k + 1` for some k.
|
||||
# As the extra single frame is handled inside the loop, it is not required to round up here.
|
||||
num_batches = max(num_frames // frame_batch_size, 1)
|
||||
conv_cache = None
|
||||
time = []
|
||||
|
||||
for k in range(num_batches):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * k + (0 if k == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (k + 1) + remaining_frames
|
||||
tile = x[
|
||||
:,
|
||||
:,
|
||||
start_frame:end_frame,
|
||||
i : i + self.tile_sample_min_height,
|
||||
j : j + self.tile_sample_min_width,
|
||||
]
|
||||
tile, conv_cache = self.encoder(tile, conv_cache=conv_cache)
|
||||
if self.quant_conv is not None:
|
||||
tile = self.quant_conv(tile)
|
||||
time.append(tile)
|
||||
|
||||
row.append(torch.cat(time, dim=2))
|
||||
rows.append(row)
|
||||
|
||||
result_rows = []
|
||||
for i, row in enumerate(rows):
|
||||
result_row = []
|
||||
for j, tile in enumerate(row):
|
||||
# blend the above tile and the left tile
|
||||
# to the current tile and add the current tile to the result row
|
||||
if i > 0:
|
||||
tile = self.blend_v(rows[i - 1][j], tile, blend_extent_height)
|
||||
if j > 0:
|
||||
tile = self.blend_h(row[j - 1], tile, blend_extent_width)
|
||||
result_row.append(tile[:, :, :, :row_limit_height, :row_limit_width])
|
||||
result_rows.append(torch.cat(result_row, dim=4))
|
||||
|
||||
enc = torch.cat(result_rows, dim=3)
|
||||
return enc
|
||||
|
||||
def tiled_decode(self, z: torch.Tensor, return_dict: bool = True) -> Union[DecoderOutput, torch.Tensor]:
|
||||
r"""
|
||||
Decode a batch of images using a tiled decoder.
|
||||
@@ -1545,34 +1237,11 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
for i in range(0, height, overlap_height):
|
||||
row = []
|
||||
for j in range(0, width, overlap_width):
|
||||
if self.auto_split_process:
|
||||
num_batches = max(num_frames // frame_batch_size, 1)
|
||||
conv_cache = None
|
||||
time = []
|
||||
|
||||
for k in range(num_batches):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * k + (0 if k == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (k + 1) + remaining_frames
|
||||
tile = z[
|
||||
:,
|
||||
:,
|
||||
start_frame:end_frame,
|
||||
i : i + self.tile_latent_min_height,
|
||||
j : j + self.tile_latent_min_width,
|
||||
]
|
||||
if self.post_quant_conv is not None:
|
||||
tile = self.post_quant_conv(tile)
|
||||
tile, conv_cache = self.decoder(tile, conv_cache=conv_cache)
|
||||
time.append(tile)
|
||||
|
||||
row.append(torch.cat(time, dim=2))
|
||||
else:
|
||||
conv_cache = None
|
||||
start_frame = 0
|
||||
end_frame = 1
|
||||
dec = []
|
||||
|
||||
time = []
|
||||
for k in range(num_frames // frame_batch_size):
|
||||
remaining_frames = num_frames % frame_batch_size
|
||||
start_frame = frame_batch_size * k + (0 if k == 0 else remaining_frames)
|
||||
end_frame = frame_batch_size * (k + 1) + remaining_frames
|
||||
tile = z[
|
||||
:,
|
||||
:,
|
||||
@@ -1580,33 +1249,12 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
i : i + self.tile_latent_min_height,
|
||||
j : j + self.tile_latent_min_width,
|
||||
]
|
||||
|
||||
self._set_first_frame()
|
||||
if self.post_quant_conv is not None:
|
||||
tile = self.post_quant_conv(tile)
|
||||
tile, conv_cache = self.decoder(tile, conv_cache=conv_cache)
|
||||
dec.append(tile)
|
||||
|
||||
self._set_rest_frame()
|
||||
start_frame = end_frame
|
||||
end_frame += self.num_latent_frames_batch_size
|
||||
|
||||
while start_frame < num_frames:
|
||||
tile = z[
|
||||
:,
|
||||
:,
|
||||
start_frame:end_frame,
|
||||
i : i + self.tile_latent_min_height,
|
||||
j : j + self.tile_latent_min_width,
|
||||
]
|
||||
if self.post_quant_conv is not None:
|
||||
tile = self.post_quant_conv(tile)
|
||||
tile, conv_cache = self.decoder(tile, conv_cache=conv_cache)
|
||||
dec.append(tile)
|
||||
start_frame = end_frame
|
||||
end_frame += self.num_latent_frames_batch_size
|
||||
|
||||
row.append(torch.cat(dec, dim=2))
|
||||
tile = self.decoder(tile)
|
||||
time.append(tile)
|
||||
self._clear_fake_context_parallel_cache()
|
||||
row.append(torch.cat(time, dim=2))
|
||||
rows.append(row)
|
||||
|
||||
result_rows = []
|
||||
@@ -1646,30 +1294,3 @@ class AutoencoderKLCogVideoX(ModelMixin, ConfigMixin, FromOriginalModelMixin):
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
return dec
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(cls, pretrained_model_path, subfolder=None, **vae_additional_kwargs):
|
||||
if subfolder is not None:
|
||||
pretrained_model_path = os.path.join(pretrained_model_path, subfolder)
|
||||
|
||||
config_file = os.path.join(pretrained_model_path, 'config.json')
|
||||
if not os.path.isfile(config_file):
|
||||
raise RuntimeError(f"{config_file} does not exist")
|
||||
with open(config_file, "r") as f:
|
||||
config = json.load(f)
|
||||
|
||||
model = cls.from_config(config, **vae_additional_kwargs)
|
||||
from diffusers.utils import WEIGHTS_NAME
|
||||
model_file = os.path.join(pretrained_model_path, WEIGHTS_NAME)
|
||||
model_file_safetensors = model_file.replace(".bin", ".safetensors")
|
||||
if os.path.exists(model_file_safetensors):
|
||||
from safetensors.torch import load_file, safe_open
|
||||
state_dict = load_file(model_file_safetensors)
|
||||
else:
|
||||
if not os.path.isfile(model_file):
|
||||
raise RuntimeError(f"{model_file} does not exist")
|
||||
state_dict = torch.load(model_file, map_location="cpu", weights_only=True)
|
||||
m, u = model.load_state_dict(state_dict, strict=False)
|
||||
print(f"### missing keys: {len(m)}; \n### unexpected keys: {len(u)};")
|
||||
print(m, u)
|
||||
return model
|
||||
@@ -13,235 +13,29 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
import glob
|
||||
import json
|
||||
import os
|
||||
from typing import Any, Dict, Optional, Tuple, Union
|
||||
|
||||
import os
|
||||
import json
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torch import nn
|
||||
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.utils import is_torch_version, logging
|
||||
from diffusers.utils.torch_utils import maybe_allow_in_graph
|
||||
from diffusers.models.attention import Attention, FeedForward
|
||||
from diffusers.models.attention_processor import (
|
||||
AttentionProcessor, FusedCogVideoXAttnProcessor2_0)
|
||||
from diffusers.models.embeddings import (CogVideoXPatchEmbed,
|
||||
TimestepEmbedding, Timesteps,
|
||||
get_3d_sincos_pos_embed)
|
||||
from diffusers.models.attention_processor import AttentionProcessor, CogVideoXAttnProcessor2_0, FusedCogVideoXAttnProcessor2_0
|
||||
from diffusers.models.embeddings import CogVideoXPatchEmbed, TimestepEmbedding, Timesteps, get_3d_sincos_pos_embed
|
||||
from diffusers.models.modeling_outputs import Transformer2DModelOutput
|
||||
from diffusers.models.modeling_utils import ModelMixin
|
||||
from diffusers.models.normalization import AdaLayerNorm, CogVideoXLayerNormZero
|
||||
from diffusers.utils import is_torch_version, logging
|
||||
from diffusers.utils.torch_utils import maybe_allow_in_graph
|
||||
from torch import nn
|
||||
|
||||
from ..dist import (get_sequence_parallel_rank,
|
||||
get_sequence_parallel_world_size, get_sp_group,
|
||||
xFuserLongContextAttention)
|
||||
from ..dist.cogvideox_xfuser import CogVideoXMultiGPUsAttnProcessor2_0
|
||||
from .attention_utils import attention
|
||||
from cogvideox.models.attention_processor import CogVideoXSWAAttnProcessor2_0
|
||||
|
||||
|
||||
class CogVideoXAttnProcessor2_0:
|
||||
r"""
|
||||
Processor for implementing scaled dot-product attention for the CogVideoX model. It applies a rotary embedding on
|
||||
query and key vectors, but does not include spatial normalization.
|
||||
"""
|
||||
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
def __init__(self):
|
||||
if not hasattr(F, "scaled_dot_product_attention"):
|
||||
raise ImportError("CogVideoXAttnProcessor requires PyTorch 2.0, to use it, please upgrade PyTorch to 2.0.")
|
||||
|
||||
def __call__(
|
||||
self,
|
||||
attn,
|
||||
hidden_states: torch.Tensor,
|
||||
encoder_hidden_states: torch.Tensor,
|
||||
attention_mask: torch.Tensor = None,
|
||||
image_rotary_emb: torch.Tensor = None,
|
||||
) -> torch.Tensor:
|
||||
text_seq_length = encoder_hidden_states.size(1)
|
||||
|
||||
hidden_states = torch.cat([encoder_hidden_states, hidden_states], dim=1)
|
||||
|
||||
batch_size, sequence_length, _ = hidden_states.shape
|
||||
|
||||
if attention_mask is not None:
|
||||
attention_mask = attn.prepare_attention_mask(attention_mask, sequence_length, batch_size)
|
||||
attention_mask = attention_mask.view(batch_size, attn.heads, -1, attention_mask.shape[-1])
|
||||
|
||||
query = attn.to_q(hidden_states)
|
||||
key = attn.to_k(hidden_states)
|
||||
value = attn.to_v(hidden_states)
|
||||
|
||||
inner_dim = key.shape[-1]
|
||||
head_dim = inner_dim // attn.heads
|
||||
|
||||
query = query.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
|
||||
key = key.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
|
||||
value = value.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
|
||||
|
||||
if attn.norm_q is not None:
|
||||
query = attn.norm_q(query)
|
||||
if attn.norm_k is not None:
|
||||
key = attn.norm_k(key)
|
||||
|
||||
# Apply RoPE if needed
|
||||
if image_rotary_emb is not None:
|
||||
from diffusers.models.embeddings import apply_rotary_emb
|
||||
|
||||
query[:, :, text_seq_length:] = apply_rotary_emb(query[:, :, text_seq_length:], image_rotary_emb)
|
||||
if not attn.is_cross_attention:
|
||||
key[:, :, text_seq_length:] = apply_rotary_emb(key[:, :, text_seq_length:], image_rotary_emb)
|
||||
|
||||
query = query.transpose(1, 2)
|
||||
key = key.transpose(1, 2)
|
||||
value = value.transpose(1, 2)
|
||||
|
||||
hidden_states = attention(
|
||||
query, key, value, attn_mask=attention_mask, dropout_p=0.0, causal=False
|
||||
)
|
||||
hidden_states = hidden_states.reshape(batch_size, -1, attn.heads * head_dim)
|
||||
|
||||
# linear proj
|
||||
hidden_states = attn.to_out[0](hidden_states)
|
||||
# dropout
|
||||
hidden_states = attn.to_out[1](hidden_states)
|
||||
|
||||
encoder_hidden_states, hidden_states = hidden_states.split(
|
||||
[text_seq_length, hidden_states.size(1) - text_seq_length], dim=1
|
||||
)
|
||||
return hidden_states, encoder_hidden_states
|
||||
|
||||
|
||||
class CogVideoXPatchEmbed(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
patch_size: int = 2,
|
||||
patch_size_t: Optional[int] = None,
|
||||
in_channels: int = 16,
|
||||
embed_dim: int = 1920,
|
||||
text_embed_dim: int = 4096,
|
||||
bias: bool = True,
|
||||
sample_width: int = 90,
|
||||
sample_height: int = 60,
|
||||
sample_frames: int = 49,
|
||||
temporal_compression_ratio: int = 4,
|
||||
max_text_seq_length: int = 226,
|
||||
spatial_interpolation_scale: float = 1.875,
|
||||
temporal_interpolation_scale: float = 1.0,
|
||||
use_positional_embeddings: bool = True,
|
||||
use_learned_positional_embeddings: bool = True,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
|
||||
post_patch_height = sample_height // patch_size
|
||||
post_patch_width = sample_width // patch_size
|
||||
post_time_compression_frames = (sample_frames - 1) // temporal_compression_ratio + 1
|
||||
self.num_patches = post_patch_height * post_patch_width * post_time_compression_frames
|
||||
self.post_patch_height = post_patch_height
|
||||
self.post_patch_width = post_patch_width
|
||||
self.post_time_compression_frames = post_time_compression_frames
|
||||
self.patch_size = patch_size
|
||||
self.patch_size_t = patch_size_t
|
||||
self.embed_dim = embed_dim
|
||||
self.sample_height = sample_height
|
||||
self.sample_width = sample_width
|
||||
self.sample_frames = sample_frames
|
||||
self.temporal_compression_ratio = temporal_compression_ratio
|
||||
self.max_text_seq_length = max_text_seq_length
|
||||
self.spatial_interpolation_scale = spatial_interpolation_scale
|
||||
self.temporal_interpolation_scale = temporal_interpolation_scale
|
||||
self.use_positional_embeddings = use_positional_embeddings
|
||||
self.use_learned_positional_embeddings = use_learned_positional_embeddings
|
||||
|
||||
if patch_size_t is None:
|
||||
# CogVideoX 1.0 checkpoints
|
||||
self.proj = nn.Conv2d(
|
||||
in_channels, embed_dim, kernel_size=(patch_size, patch_size), stride=patch_size, bias=bias
|
||||
)
|
||||
else:
|
||||
# CogVideoX 1.5 checkpoints
|
||||
self.proj = nn.Linear(in_channels * patch_size * patch_size * patch_size_t, embed_dim)
|
||||
|
||||
self.text_proj = nn.Linear(text_embed_dim, embed_dim)
|
||||
|
||||
if use_positional_embeddings or use_learned_positional_embeddings:
|
||||
persistent = use_learned_positional_embeddings
|
||||
pos_embedding = self._get_positional_embeddings(sample_height, sample_width, sample_frames)
|
||||
self.register_buffer("pos_embedding", pos_embedding, persistent=persistent)
|
||||
|
||||
def _get_positional_embeddings(self, sample_height: int, sample_width: int, sample_frames: int) -> torch.Tensor:
|
||||
post_patch_height = sample_height // self.patch_size
|
||||
post_patch_width = sample_width // self.patch_size
|
||||
post_time_compression_frames = (sample_frames - 1) // self.temporal_compression_ratio + 1
|
||||
num_patches = post_patch_height * post_patch_width * post_time_compression_frames
|
||||
|
||||
pos_embedding = get_3d_sincos_pos_embed(
|
||||
self.embed_dim,
|
||||
(post_patch_width, post_patch_height),
|
||||
post_time_compression_frames,
|
||||
self.spatial_interpolation_scale,
|
||||
self.temporal_interpolation_scale,
|
||||
output_type="pt",
|
||||
)
|
||||
pos_embedding = pos_embedding.flatten(0, 1)
|
||||
joint_pos_embedding = torch.zeros(
|
||||
1, self.max_text_seq_length + num_patches, self.embed_dim, requires_grad=False
|
||||
)
|
||||
joint_pos_embedding.data[:, self.max_text_seq_length :].copy_(pos_embedding)
|
||||
|
||||
return joint_pos_embedding
|
||||
|
||||
def forward(self, text_embeds: torch.Tensor, image_embeds: torch.Tensor):
|
||||
r"""
|
||||
Args:
|
||||
text_embeds (`torch.Tensor`):
|
||||
Input text embeddings. Expected shape: (batch_size, seq_length, embedding_dim).
|
||||
image_embeds (`torch.Tensor`):
|
||||
Input image embeddings. Expected shape: (batch_size, num_frames, channels, height, width).
|
||||
"""
|
||||
text_embeds = self.text_proj(text_embeds)
|
||||
|
||||
text_batch_size, text_seq_length, text_channels = text_embeds.shape
|
||||
batch_size, num_frames, channels, height, width = image_embeds.shape
|
||||
|
||||
if self.patch_size_t is None:
|
||||
image_embeds = image_embeds.reshape(-1, channels, height, width)
|
||||
image_embeds = self.proj(image_embeds)
|
||||
image_embeds = image_embeds.view(batch_size, num_frames, *image_embeds.shape[1:])
|
||||
image_embeds = image_embeds.flatten(3).transpose(2, 3) # [batch, num_frames, height x width, channels]
|
||||
image_embeds = image_embeds.flatten(1, 2) # [batch, num_frames x height x width, channels]
|
||||
else:
|
||||
p = self.patch_size
|
||||
p_t = self.patch_size_t
|
||||
|
||||
image_embeds = image_embeds.permute(0, 1, 3, 4, 2)
|
||||
# b, f, h, w, c => b, f // 2, 2, h // 2, 2, w // 2, 2, c
|
||||
image_embeds = image_embeds.reshape(
|
||||
batch_size, num_frames // p_t, p_t, height // p, p, width // p, p, channels
|
||||
)
|
||||
# b, f // 2, 2, h // 2, 2, w // 2, 2, c => b, f // 2, h // 2, w // 2, c, 2, 2, 2
|
||||
image_embeds = image_embeds.permute(0, 1, 3, 5, 7, 2, 4, 6).flatten(4, 7).flatten(1, 3)
|
||||
image_embeds = self.proj(image_embeds)
|
||||
|
||||
embeds = torch.cat(
|
||||
[text_embeds, image_embeds], dim=1
|
||||
).contiguous() # [batch, seq_length + num_frames x height x width, channels]
|
||||
|
||||
if self.use_positional_embeddings or self.use_learned_positional_embeddings:
|
||||
seq_length = height * width * num_frames // (self.patch_size**2)
|
||||
# pos_embeds = self.pos_embedding[:, : text_seq_length + seq_length]
|
||||
pos_embeds = self.pos_embedding
|
||||
emb_size = embeds.size()[-1]
|
||||
pos_embeds_without_text = pos_embeds[:, text_seq_length: ].view(1, self.post_time_compression_frames, self.post_patch_height, self.post_patch_width, emb_size)
|
||||
pos_embeds_without_text = pos_embeds_without_text.permute([0, 4, 1, 2, 3])
|
||||
pos_embeds_without_text = F.interpolate(pos_embeds_without_text,size=[self.post_time_compression_frames, height // self.patch_size, width // self.patch_size], mode='trilinear', align_corners=False)
|
||||
pos_embeds_without_text = pos_embeds_without_text.permute([0, 2, 3, 4, 1]).view(1, -1, emb_size)
|
||||
pos_embeds = torch.cat([pos_embeds[:, :text_seq_length], pos_embeds_without_text], dim = 1)
|
||||
pos_embeds = pos_embeds[:, : text_seq_length + seq_length]
|
||||
embeds = embeds + pos_embeds
|
||||
|
||||
return embeds
|
||||
|
||||
@maybe_allow_in_graph
|
||||
class CogVideoXBlock(nn.Module):
|
||||
@@ -295,6 +89,8 @@ class CogVideoXBlock(nn.Module):
|
||||
ff_inner_dim: Optional[int] = None,
|
||||
ff_bias: bool = True,
|
||||
attention_out_bias: bool = True,
|
||||
swa: bool = False,
|
||||
window_size: int = 1024,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
@@ -309,7 +105,7 @@ class CogVideoXBlock(nn.Module):
|
||||
eps=1e-6,
|
||||
bias=attention_bias,
|
||||
out_bias=attention_out_bias,
|
||||
processor=CogVideoXAttnProcessor2_0(),
|
||||
processor=CogVideoXSWAAttnProcessor2_0(window_size) if swa else CogVideoXAttnProcessor2_0(),
|
||||
)
|
||||
|
||||
# 2. Feed Forward
|
||||
@@ -330,6 +126,9 @@ class CogVideoXBlock(nn.Module):
|
||||
encoder_hidden_states: torch.Tensor,
|
||||
temb: torch.Tensor,
|
||||
image_rotary_emb: Optional[Tuple[torch.Tensor, torch.Tensor]] = None,
|
||||
num_frames: int = None,
|
||||
height: int = None,
|
||||
width: int = None,
|
||||
) -> torch.Tensor:
|
||||
text_seq_length = encoder_hidden_states.size(1)
|
||||
|
||||
@@ -343,6 +142,9 @@ class CogVideoXBlock(nn.Module):
|
||||
hidden_states=norm_hidden_states,
|
||||
encoder_hidden_states=norm_encoder_hidden_states,
|
||||
image_rotary_emb=image_rotary_emb,
|
||||
num_frames = num_frames,
|
||||
height = height,
|
||||
width = width,
|
||||
)
|
||||
|
||||
hidden_states = hidden_states + gate_msa * attn_hidden_states
|
||||
@@ -437,7 +239,6 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
sample_height: int = 60,
|
||||
sample_frames: int = 49,
|
||||
patch_size: int = 2,
|
||||
patch_size_t: Optional[int] = None,
|
||||
temporal_compression_ratio: int = 4,
|
||||
max_text_seq_length: int = 226,
|
||||
activation_fn: str = "gelu-approximate",
|
||||
@@ -447,45 +248,43 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
spatial_interpolation_scale: float = 1.875,
|
||||
temporal_interpolation_scale: float = 1.0,
|
||||
use_rotary_positional_embeddings: bool = False,
|
||||
use_learned_positional_embeddings: bool = False,
|
||||
patch_bias: bool = True,
|
||||
add_noise_in_inpaint_model: bool = False,
|
||||
swa: bool = False,
|
||||
window_size: int = 1024,
|
||||
):
|
||||
super().__init__()
|
||||
inner_dim = num_attention_heads * attention_head_dim
|
||||
self.patch_size_t = patch_size_t
|
||||
if not use_rotary_positional_embeddings and use_learned_positional_embeddings:
|
||||
raise ValueError(
|
||||
"There are no CogVideoX checkpoints available with disable rotary embeddings and learned positional "
|
||||
"embeddings. If you're using a custom model and/or believe this should be supported, please open an "
|
||||
"issue at https://github.com/huggingface/diffusers/issues."
|
||||
)
|
||||
|
||||
post_patch_height = sample_height // patch_size
|
||||
post_patch_width = sample_width // patch_size
|
||||
post_time_compression_frames = (sample_frames - 1) // temporal_compression_ratio + 1
|
||||
self.num_patches = post_patch_height * post_patch_width * post_time_compression_frames
|
||||
self.post_patch_height = post_patch_height
|
||||
self.post_patch_width = post_patch_width
|
||||
self.post_time_compression_frames = post_time_compression_frames
|
||||
self.patch_size = patch_size
|
||||
|
||||
# 1. Patch embedding
|
||||
self.patch_embed = CogVideoXPatchEmbed(
|
||||
patch_size=patch_size,
|
||||
patch_size_t=patch_size_t,
|
||||
in_channels=in_channels,
|
||||
embed_dim=inner_dim,
|
||||
text_embed_dim=text_embed_dim,
|
||||
bias=patch_bias,
|
||||
sample_width=sample_width,
|
||||
sample_height=sample_height,
|
||||
sample_frames=sample_frames,
|
||||
temporal_compression_ratio=temporal_compression_ratio,
|
||||
max_text_seq_length=max_text_seq_length,
|
||||
spatial_interpolation_scale=spatial_interpolation_scale,
|
||||
temporal_interpolation_scale=temporal_interpolation_scale,
|
||||
use_positional_embeddings=not use_rotary_positional_embeddings,
|
||||
use_learned_positional_embeddings=use_learned_positional_embeddings,
|
||||
)
|
||||
self.patch_embed = CogVideoXPatchEmbed(patch_size, in_channels, inner_dim, text_embed_dim, bias=True)
|
||||
self.embedding_dropout = nn.Dropout(dropout)
|
||||
|
||||
# 2. Time embeddings
|
||||
# 2. 3D positional embeddings
|
||||
spatial_pos_embedding = get_3d_sincos_pos_embed(
|
||||
inner_dim,
|
||||
(post_patch_width, post_patch_height),
|
||||
post_time_compression_frames,
|
||||
spatial_interpolation_scale,
|
||||
temporal_interpolation_scale,
|
||||
)
|
||||
spatial_pos_embedding = torch.from_numpy(spatial_pos_embedding).flatten(0, 1)
|
||||
pos_embedding = torch.zeros(1, max_text_seq_length + self.num_patches, inner_dim, requires_grad=False)
|
||||
pos_embedding.data[:, max_text_seq_length:].copy_(spatial_pos_embedding)
|
||||
self.register_buffer("pos_embedding", pos_embedding, persistent=False)
|
||||
|
||||
# 3. Time embeddings
|
||||
self.time_proj = Timesteps(inner_dim, flip_sin_to_cos, freq_shift)
|
||||
self.time_embedding = TimestepEmbedding(inner_dim, time_embed_dim, timestep_activation_fn)
|
||||
|
||||
# 3. Define spatio-temporal transformers blocks
|
||||
# 4. Define spatio-temporal transformers blocks
|
||||
self.transformer_blocks = nn.ModuleList(
|
||||
[
|
||||
CogVideoXBlock(
|
||||
@@ -498,13 +297,15 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
attention_bias=attention_bias,
|
||||
norm_elementwise_affine=norm_elementwise_affine,
|
||||
norm_eps=norm_eps,
|
||||
swa=swa,
|
||||
window_size=window_size,
|
||||
)
|
||||
for _ in range(num_layers)
|
||||
]
|
||||
)
|
||||
self.norm_final = nn.LayerNorm(inner_dim, norm_eps, norm_elementwise_affine)
|
||||
|
||||
# 4. Output blocks
|
||||
# 5. Output blocks
|
||||
self.norm_out = AdaLayerNorm(
|
||||
embedding_dim=time_embed_dim,
|
||||
output_dim=2 * inner_dim,
|
||||
@@ -512,32 +313,12 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
norm_eps=norm_eps,
|
||||
chunk_dim=1,
|
||||
)
|
||||
|
||||
if patch_size_t is None:
|
||||
# For CogVideox 1.0
|
||||
output_dim = patch_size * patch_size * out_channels
|
||||
else:
|
||||
# For CogVideoX 1.5
|
||||
output_dim = patch_size * patch_size * patch_size_t * out_channels
|
||||
|
||||
self.proj_out = nn.Linear(inner_dim, output_dim)
|
||||
self.proj_out = nn.Linear(inner_dim, patch_size * patch_size * out_channels)
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
self.sp_world_size = 1
|
||||
self.sp_world_rank = 0
|
||||
|
||||
def _set_gradient_checkpointing(self, *args, **kwargs):
|
||||
if "value" in kwargs:
|
||||
self.gradient_checkpointing = kwargs["value"]
|
||||
elif "enable" in kwargs:
|
||||
self.gradient_checkpointing = kwargs["enable"]
|
||||
else:
|
||||
raise ValueError("Invalid set gradient checkpointing")
|
||||
|
||||
def enable_multi_gpus_inference(self,):
|
||||
self.sp_world_size = get_sequence_parallel_world_size()
|
||||
self.sp_world_rank = get_sequence_parallel_rank()
|
||||
self.set_attn_processor(CogVideoXMultiGPUsAttnProcessor2_0())
|
||||
def _set_gradient_checkpointing(self, module, value=False):
|
||||
self.gradient_checkpointing = value
|
||||
|
||||
@property
|
||||
# Copied from diffusers.models.unets.unet_2d_condition.UNet2DConditionModel.attn_processors
|
||||
@@ -646,20 +427,11 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
timestep: Union[int, float, torch.LongTensor],
|
||||
timestep_cond: Optional[torch.Tensor] = None,
|
||||
inpaint_latents: Optional[torch.Tensor] = None,
|
||||
control_latents: Optional[torch.Tensor] = None,
|
||||
image_rotary_emb: Optional[Tuple[torch.Tensor, torch.Tensor]] = None,
|
||||
return_dict: bool = True,
|
||||
):
|
||||
batch_size, num_frames, channels, height, width = hidden_states.shape
|
||||
if num_frames == 1 and self.patch_size_t is not None:
|
||||
hidden_states = torch.cat([hidden_states, torch.zeros_like(hidden_states)], dim=1)
|
||||
if inpaint_latents is not None:
|
||||
inpaint_latents = torch.concat([inpaint_latents, torch.zeros_like(inpaint_latents)], dim=1)
|
||||
if control_latents is not None:
|
||||
control_latents = torch.concat([control_latents, torch.zeros_like(control_latents)], dim=1)
|
||||
local_num_frames = num_frames + 1
|
||||
else:
|
||||
local_num_frames = num_frames
|
||||
p = self.config.patch_size
|
||||
|
||||
# 1. Time embedding
|
||||
timesteps = timestep
|
||||
@@ -674,27 +446,30 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
# 2. Patch embedding
|
||||
if inpaint_latents is not None:
|
||||
hidden_states = torch.concat([hidden_states, inpaint_latents], 2)
|
||||
if control_latents is not None:
|
||||
hidden_states = torch.concat([hidden_states, control_latents], 2)
|
||||
hidden_states = self.patch_embed(encoder_hidden_states, hidden_states)
|
||||
hidden_states = self.embedding_dropout(hidden_states)
|
||||
|
||||
# 3. Position embedding
|
||||
text_seq_length = encoder_hidden_states.shape[1]
|
||||
if not self.config.use_rotary_positional_embeddings:
|
||||
seq_length = height * width * num_frames // (self.config.patch_size**2)
|
||||
# pos_embeds = self.pos_embedding[:, : text_seq_length + seq_length]
|
||||
pos_embeds = self.pos_embedding
|
||||
emb_size = hidden_states.size()[-1]
|
||||
pos_embeds_without_text = pos_embeds[:, text_seq_length: ].view(1, self.post_time_compression_frames, self.post_patch_height, self.post_patch_width, emb_size)
|
||||
pos_embeds_without_text = pos_embeds_without_text.permute([0, 4, 1, 2, 3])
|
||||
pos_embeds_without_text = F.interpolate(pos_embeds_without_text,size=[self.post_time_compression_frames, height // self.config.patch_size, width // self.config.patch_size],mode='trilinear',align_corners=False)
|
||||
pos_embeds_without_text = pos_embeds_without_text.permute([0, 2, 3, 4, 1]).view(1, -1, emb_size)
|
||||
pos_embeds = torch.cat([pos_embeds[:, :text_seq_length], pos_embeds_without_text], dim = 1)
|
||||
pos_embeds = pos_embeds[:, : text_seq_length + seq_length]
|
||||
hidden_states = hidden_states + pos_embeds
|
||||
hidden_states = self.embedding_dropout(hidden_states)
|
||||
|
||||
encoder_hidden_states = hidden_states[:, :text_seq_length]
|
||||
hidden_states = hidden_states[:, text_seq_length:]
|
||||
|
||||
# Context Parallel
|
||||
if self.sp_world_size > 1:
|
||||
hidden_states = torch.chunk(hidden_states, self.sp_world_size, dim=1)[self.sp_world_rank]
|
||||
if image_rotary_emb is not None:
|
||||
image_rotary_emb = (
|
||||
torch.chunk(image_rotary_emb[0], self.sp_world_size, dim=0)[self.sp_world_rank],
|
||||
torch.chunk(image_rotary_emb[1], self.sp_world_size, dim=0)[self.sp_world_rank]
|
||||
)
|
||||
|
||||
# 3. Transformer blocks
|
||||
# 4. Transformer blocks
|
||||
for i, block in enumerate(self.transformer_blocks):
|
||||
if torch.is_grad_enabled() and self.gradient_checkpointing:
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def custom_forward(*inputs):
|
||||
@@ -709,6 +484,9 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
encoder_hidden_states,
|
||||
emb,
|
||||
image_rotary_emb,
|
||||
num_frames,
|
||||
height // p,
|
||||
width // p,
|
||||
**ckpt_kwargs,
|
||||
)
|
||||
else:
|
||||
@@ -717,6 +495,9 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
encoder_hidden_states=encoder_hidden_states,
|
||||
temb=emb,
|
||||
image_rotary_emb=image_rotary_emb,
|
||||
num_frames=num_frames,
|
||||
height=height // p,
|
||||
width=width // p,
|
||||
)
|
||||
|
||||
if not self.config.use_rotary_positional_embeddings:
|
||||
@@ -728,39 +509,20 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
hidden_states = self.norm_final(hidden_states)
|
||||
hidden_states = hidden_states[:, text_seq_length:]
|
||||
|
||||
# 4. Final block
|
||||
# 5. Final block
|
||||
hidden_states = self.norm_out(hidden_states, temb=emb)
|
||||
hidden_states = self.proj_out(hidden_states)
|
||||
|
||||
if self.sp_world_size > 1:
|
||||
hidden_states = get_sp_group().all_gather(hidden_states, dim=1)
|
||||
|
||||
# 5. Unpatchify
|
||||
p = self.config.patch_size
|
||||
p_t = self.config.patch_size_t
|
||||
|
||||
if p_t is None:
|
||||
output = hidden_states.reshape(batch_size, local_num_frames, height // p, width // p, -1, p, p)
|
||||
output = output.permute(0, 1, 4, 2, 5, 3, 6).flatten(5, 6).flatten(3, 4)
|
||||
else:
|
||||
output = hidden_states.reshape(
|
||||
batch_size, (local_num_frames + p_t - 1) // p_t, height // p, width // p, -1, p_t, p, p
|
||||
)
|
||||
output = output.permute(0, 1, 5, 4, 2, 6, 3, 7).flatten(6, 7).flatten(4, 5).flatten(1, 2)
|
||||
|
||||
if num_frames == 1:
|
||||
output = output[:, :num_frames, :]
|
||||
# 6. Unpatchify
|
||||
output = hidden_states.reshape(batch_size, num_frames, height // p, width // p, channels, p, p)
|
||||
output = output.permute(0, 1, 4, 2, 5, 3, 6).flatten(5, 6).flatten(3, 4)
|
||||
|
||||
if not return_dict:
|
||||
return (output,)
|
||||
return Transformer2DModelOutput(sample=output)
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(
|
||||
cls, pretrained_model_path, subfolder=None, transformer_additional_kwargs=None,
|
||||
low_cpu_mem_usage=False, torch_dtype=torch.bfloat16
|
||||
):
|
||||
transformer_additional_kwargs = {} if transformer_additional_kwargs is None else dict(transformer_additional_kwargs)
|
||||
def from_pretrained_2d(cls, pretrained_model_path, subfolder=None, transformer_additional_kwargs={}):
|
||||
if subfolder is not None:
|
||||
pretrained_model_path = os.path.join(pretrained_model_path, subfolder)
|
||||
print(f"loaded 3D transformer's pretrained weights from {pretrained_model_path} ...")
|
||||
@@ -772,155 +534,47 @@ class CogVideoXTransformer3DModel(ModelMixin, ConfigMixin):
|
||||
config = json.load(f)
|
||||
|
||||
from diffusers.utils import WEIGHTS_NAME
|
||||
model = cls.from_config(config, **transformer_additional_kwargs)
|
||||
model_file = os.path.join(pretrained_model_path, WEIGHTS_NAME)
|
||||
model_file_safetensors = model_file.replace(".bin", ".safetensors")
|
||||
|
||||
if "dict_mapping" in transformer_additional_kwargs.keys():
|
||||
dict_mapping = transformer_additional_kwargs.pop("dict_mapping")
|
||||
for key in dict_mapping:
|
||||
transformer_additional_kwargs[dict_mapping[key]] = config[key]
|
||||
|
||||
if low_cpu_mem_usage:
|
||||
try:
|
||||
import re
|
||||
|
||||
from diffusers import __version__ as diffusers_version
|
||||
from packaging import version as pkg_version
|
||||
if pkg_version.parse(diffusers_version) >= pkg_version.parse("0.33.0"):
|
||||
from diffusers.models.model_loading_utils import \
|
||||
load_model_dict_into_meta
|
||||
else:
|
||||
from diffusers.models.modeling_utils import \
|
||||
load_model_dict_into_meta
|
||||
from diffusers.utils import is_accelerate_available
|
||||
if is_accelerate_available():
|
||||
import accelerate
|
||||
|
||||
# Instantiate model with empty weights
|
||||
with accelerate.init_empty_weights():
|
||||
model = cls.from_config(config, **transformer_additional_kwargs)
|
||||
|
||||
param_device = "cpu"
|
||||
if os.path.exists(model_file):
|
||||
state_dict = torch.load(model_file, map_location="cpu", weights_only=True)
|
||||
elif os.path.exists(model_file_safetensors):
|
||||
from safetensors.torch import load_file, safe_open
|
||||
state_dict = load_file(model_file_safetensors)
|
||||
else:
|
||||
from safetensors.torch import load_file, safe_open
|
||||
model_files_safetensors = glob.glob(os.path.join(pretrained_model_path, "*.safetensors"))
|
||||
state_dict = {}
|
||||
for _model_file_safetensors in model_files_safetensors:
|
||||
_state_dict = load_file(_model_file_safetensors)
|
||||
for key in _state_dict:
|
||||
state_dict[key] = _state_dict[key]
|
||||
if len(state_dict) == 0:
|
||||
raise FileNotFoundError(f"No weights found in {pretrained_model_path}")
|
||||
model._convert_deprecated_attention_blocks(state_dict)
|
||||
|
||||
if pkg_version.parse(diffusers_version) >= pkg_version.parse("0.33.0"):
|
||||
# Diffusers has refactored `load_model_dict_into_meta` since version 0.33.0 in this commit:
|
||||
# https://github.com/huggingface/diffusers/commit/f5929e03060d56063ff34b25a8308833bec7c785.
|
||||
load_model_dict_into_meta(
|
||||
model,
|
||||
state_dict,
|
||||
dtype=torch_dtype,
|
||||
model_name_or_path=pretrained_model_path,
|
||||
)
|
||||
else:
|
||||
# move the params from meta device to cpu
|
||||
missing_keys = set(model.state_dict().keys()) - set(state_dict.keys())
|
||||
if len(missing_keys) > 0:
|
||||
raise ValueError(
|
||||
f"Cannot load {cls} from {pretrained_model_path} because the following keys are"
|
||||
f" missing: \n {', '.join(missing_keys)}. \n Please make sure to pass"
|
||||
" `low_cpu_mem_usage=False` and `device_map=None` if you want to randomly initialize"
|
||||
" those weights or else make sure your checkpoint file is correct."
|
||||
)
|
||||
|
||||
unexpected_keys = load_model_dict_into_meta(
|
||||
model,
|
||||
state_dict,
|
||||
device=param_device,
|
||||
dtype=torch_dtype,
|
||||
model_name_or_path=pretrained_model_path,
|
||||
)
|
||||
|
||||
if cls._keys_to_ignore_on_load_unexpected is not None:
|
||||
for pat in cls._keys_to_ignore_on_load_unexpected:
|
||||
unexpected_keys = [k for k in unexpected_keys if re.search(pat, k) is None]
|
||||
|
||||
if len(unexpected_keys) > 0:
|
||||
print(
|
||||
f"Some weights of the model checkpoint were not used when initializing {cls.__name__}: \n {[', '.join(unexpected_keys)]}"
|
||||
)
|
||||
|
||||
return model
|
||||
except Exception as e:
|
||||
import traceback
|
||||
traceback.print_exc()
|
||||
print(
|
||||
f"The low_cpu_mem_usage mode is not work because {e}. Use low_cpu_mem_usage=False instead."
|
||||
)
|
||||
|
||||
model = cls.from_config(config, **transformer_additional_kwargs)
|
||||
if os.path.exists(model_file):
|
||||
state_dict = torch.load(model_file, map_location="cpu", weights_only=True)
|
||||
elif os.path.exists(model_file_safetensors):
|
||||
if os.path.exists(model_file_safetensors):
|
||||
from safetensors.torch import load_file, safe_open
|
||||
state_dict = load_file(model_file_safetensors)
|
||||
else:
|
||||
from safetensors.torch import load_file, safe_open
|
||||
model_files_safetensors = glob.glob(os.path.join(pretrained_model_path, "*.safetensors"))
|
||||
state_dict = {}
|
||||
for _model_file_safetensors in model_files_safetensors:
|
||||
_state_dict = load_file(_model_file_safetensors)
|
||||
for key in _state_dict:
|
||||
state_dict[key] = _state_dict[key]
|
||||
if len(state_dict) == 0:
|
||||
raise FileNotFoundError(f"No weights found in {pretrained_model_path}")
|
||||
if not os.path.isfile(model_file):
|
||||
raise RuntimeError(f"{model_file} does not exist")
|
||||
state_dict = torch.load(model_file, map_location="cpu")
|
||||
|
||||
model_state_dict = model.state_dict()
|
||||
if model_state_dict['patch_embed.proj.weight'].size() != state_dict['patch_embed.proj.weight'].size():
|
||||
new_shape = model_state_dict['patch_embed.proj.weight'].size()
|
||||
if model.state_dict()['patch_embed.proj.weight'].size() != state_dict['patch_embed.proj.weight'].size():
|
||||
new_shape = model.state_dict()['patch_embed.proj.weight'].size()
|
||||
if len(new_shape) == 5:
|
||||
state_dict['patch_embed.proj.weight'] = state_dict['patch_embed.proj.weight'].unsqueeze(2).expand(new_shape).clone()
|
||||
state_dict['patch_embed.proj.weight'][:, :, :-1] = 0
|
||||
elif len(new_shape) == 2:
|
||||
if model_state_dict['patch_embed.proj.weight'].size()[1] > state_dict['patch_embed.proj.weight'].size()[1]:
|
||||
model_state_dict['patch_embed.proj.weight'][:, :state_dict['patch_embed.proj.weight'].size()[1]] = state_dict['patch_embed.proj.weight']
|
||||
model_state_dict['patch_embed.proj.weight'][:, state_dict['patch_embed.proj.weight'].size()[1]:] = 0
|
||||
state_dict['patch_embed.proj.weight'] = model_state_dict['patch_embed.proj.weight']
|
||||
else:
|
||||
model_state_dict['patch_embed.proj.weight'][:, :] = state_dict['patch_embed.proj.weight'][:, :model_state_dict['patch_embed.proj.weight'].size()[1]]
|
||||
state_dict['patch_embed.proj.weight'] = model_state_dict['patch_embed.proj.weight']
|
||||
else:
|
||||
if model_state_dict['patch_embed.proj.weight'].size()[1] > state_dict['patch_embed.proj.weight'].size()[1]:
|
||||
model_state_dict['patch_embed.proj.weight'][:, :state_dict['patch_embed.proj.weight'].size()[1], :, :] = state_dict['patch_embed.proj.weight']
|
||||
model_state_dict['patch_embed.proj.weight'][:, state_dict['patch_embed.proj.weight'].size()[1]:, :, :] = 0
|
||||
state_dict['patch_embed.proj.weight'] = model_state_dict['patch_embed.proj.weight']
|
||||
if model.state_dict()['patch_embed.proj.weight'].size()[1] > state_dict['patch_embed.proj.weight'].size()[1]:
|
||||
model.state_dict()['patch_embed.proj.weight'][:, :state_dict['patch_embed.proj.weight'].size()[1], :, :] = state_dict['patch_embed.proj.weight']
|
||||
model.state_dict()['patch_embed.proj.weight'][:, state_dict['patch_embed.proj.weight'].size()[1]:, :, :] = 0
|
||||
state_dict['patch_embed.proj.weight'] = model.state_dict()['patch_embed.proj.weight']
|
||||
else:
|
||||
model_state_dict['patch_embed.proj.weight'][:, :, :, :] = state_dict['patch_embed.proj.weight'][:, :model_state_dict['patch_embed.proj.weight'].size()[1], :, :]
|
||||
state_dict['patch_embed.proj.weight'] = model_state_dict['patch_embed.proj.weight']
|
||||
model.state_dict()['patch_embed.proj.weight'][:, :, :, :] = state_dict['patch_embed.proj.weight'][:, :model.state_dict()['patch_embed.proj.weight'].size()[1], :, :]
|
||||
state_dict['patch_embed.proj.weight'] = model.state_dict()['patch_embed.proj.weight']
|
||||
|
||||
tmp_state_dict = {}
|
||||
for key in state_dict:
|
||||
if key in model_state_dict.keys() and model_state_dict[key].size() == state_dict[key].size():
|
||||
if key in model.state_dict().keys() and model.state_dict()[key].size() == state_dict[key].size():
|
||||
tmp_state_dict[key] = state_dict[key]
|
||||
else:
|
||||
print(key, "Size don't match, skip")
|
||||
|
||||
state_dict = tmp_state_dict
|
||||
|
||||
m, u = model.load_state_dict(state_dict, strict=False)
|
||||
print(f"### missing keys: {len(m)}; \n### unexpected keys: {len(u)};")
|
||||
print(m)
|
||||
|
||||
params = [p.numel() if "." in n else 0 for n, p in model.named_parameters()]
|
||||
print(f"### All Parameters: {sum(params) / 1e6} M")
|
||||
params = [p.numel() if "mamba" in n else 0 for n, p in model.named_parameters()]
|
||||
print(f"### Mamba Parameters: {sum(params) / 1e6} M")
|
||||
|
||||
params = [p.numel() if "attn1." in n else 0 for n, p in model.named_parameters()]
|
||||
print(f"### attn1 Parameters: {sum(params) / 1e6} M")
|
||||
|
||||
model = model.to(torch_dtype)
|
||||
return model
|
||||
@@ -16,21 +16,20 @@
|
||||
import inspect
|
||||
import math
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Callable, Dict, List, Optional, Tuple, Union
|
||||
from typing import Callable, Dict, List, Optional, Tuple, Union
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from transformers import T5EncoderModel, T5Tokenizer
|
||||
|
||||
from diffusers.callbacks import MultiPipelineCallbacks, PipelineCallback
|
||||
from diffusers.models.embeddings import get_1d_rotary_pos_embed
|
||||
from diffusers.models import AutoencoderKLCogVideoX, CogVideoXTransformer3DModel
|
||||
from diffusers.models.embeddings import get_3d_rotary_pos_embed
|
||||
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
|
||||
from diffusers.schedulers import CogVideoXDDIMScheduler, CogVideoXDPMScheduler
|
||||
from diffusers.utils import BaseOutput, logging, replace_example_docstring
|
||||
from diffusers.utils.torch_utils import randn_tensor
|
||||
from diffusers.video_processor import VideoProcessor
|
||||
|
||||
from ..models import (AutoencoderKLCogVideoX,
|
||||
CogVideoXTransformer3DModel, T5EncoderModel,
|
||||
T5Tokenizer)
|
||||
|
||||
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
@@ -38,106 +37,26 @@ logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
EXAMPLE_DOC_STRING = """
|
||||
Examples:
|
||||
```python
|
||||
pass
|
||||
>>> import torch
|
||||
>>> from diffusers import CogVideoX_Fun_Pipeline
|
||||
>>> from diffusers.utils import export_to_video
|
||||
|
||||
>>> # Models: "THUDM/CogVideoX-2b" or "THUDM/CogVideoX-5b"
|
||||
>>> pipe = CogVideoX_Fun_Pipeline.from_pretrained("THUDM/CogVideoX-2b", torch_dtype=torch.float16).to("cuda")
|
||||
>>> prompt = (
|
||||
... "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. "
|
||||
... "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other "
|
||||
... "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, "
|
||||
... "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. "
|
||||
... "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical "
|
||||
... "atmosphere of this unique musical performance."
|
||||
... )
|
||||
>>> video = pipe(prompt=prompt, guidance_scale=6, num_inference_steps=50).frames[0]
|
||||
>>> export_to_video(video, "output.mp4", fps=8)
|
||||
```
|
||||
"""
|
||||
|
||||
|
||||
# Copied from diffusers.models.embeddings.get_3d_rotary_pos_embed
|
||||
def get_3d_rotary_pos_embed(
|
||||
embed_dim,
|
||||
crops_coords,
|
||||
grid_size,
|
||||
temporal_size,
|
||||
theta: int = 10000,
|
||||
use_real: bool = True,
|
||||
grid_type: str = "linspace",
|
||||
max_size: Optional[Tuple[int, int]] = None,
|
||||
) -> Union[torch.Tensor, Tuple[torch.Tensor, torch.Tensor]]:
|
||||
"""
|
||||
RoPE for video tokens with 3D structure.
|
||||
|
||||
Args:
|
||||
embed_dim: (`int`):
|
||||
The embedding dimension size, corresponding to hidden_size_head.
|
||||
crops_coords (`Tuple[int]`):
|
||||
The top-left and bottom-right coordinates of the crop.
|
||||
grid_size (`Tuple[int]`):
|
||||
The grid size of the spatial positional embedding (height, width).
|
||||
temporal_size (`int`):
|
||||
The size of the temporal dimension.
|
||||
theta (`float`):
|
||||
Scaling factor for frequency computation.
|
||||
grid_type (`str`):
|
||||
Whether to use "linspace" or "slice" to compute grids.
|
||||
|
||||
Returns:
|
||||
`torch.Tensor`: positional embedding with shape `(temporal_size * grid_size[0] * grid_size[1], embed_dim/2)`.
|
||||
"""
|
||||
if use_real is not True:
|
||||
raise ValueError(" `use_real = False` is not currently supported for get_3d_rotary_pos_embed")
|
||||
|
||||
if grid_type == "linspace":
|
||||
start, stop = crops_coords
|
||||
grid_size_h, grid_size_w = grid_size
|
||||
grid_h = np.linspace(start[0], stop[0], grid_size_h, endpoint=False, dtype=np.float32)
|
||||
grid_w = np.linspace(start[1], stop[1], grid_size_w, endpoint=False, dtype=np.float32)
|
||||
grid_t = np.arange(temporal_size, dtype=np.float32)
|
||||
grid_t = np.linspace(0, temporal_size, temporal_size, endpoint=False, dtype=np.float32)
|
||||
elif grid_type == "slice":
|
||||
max_h, max_w = max_size
|
||||
grid_size_h, grid_size_w = grid_size
|
||||
grid_h = np.arange(max_h, dtype=np.float32)
|
||||
grid_w = np.arange(max_w, dtype=np.float32)
|
||||
grid_t = np.arange(temporal_size, dtype=np.float32)
|
||||
else:
|
||||
raise ValueError("Invalid value passed for `grid_type`.")
|
||||
|
||||
# Compute dimensions for each axis
|
||||
dim_t = embed_dim // 4
|
||||
dim_h = embed_dim // 8 * 3
|
||||
dim_w = embed_dim // 8 * 3
|
||||
|
||||
# Temporal frequencies
|
||||
freqs_t = get_1d_rotary_pos_embed(dim_t, grid_t, use_real=True)
|
||||
# Spatial frequencies for height and width
|
||||
freqs_h = get_1d_rotary_pos_embed(dim_h, grid_h, use_real=True)
|
||||
freqs_w = get_1d_rotary_pos_embed(dim_w, grid_w, use_real=True)
|
||||
|
||||
# BroadCast and concatenate temporal and spaial frequencie (height and width) into a 3d tensor
|
||||
def combine_time_height_width(freqs_t, freqs_h, freqs_w):
|
||||
freqs_t = freqs_t[:, None, None, :].expand(
|
||||
-1, grid_size_h, grid_size_w, -1
|
||||
) # temporal_size, grid_size_h, grid_size_w, dim_t
|
||||
freqs_h = freqs_h[None, :, None, :].expand(
|
||||
temporal_size, -1, grid_size_w, -1
|
||||
) # temporal_size, grid_size_h, grid_size_2, dim_h
|
||||
freqs_w = freqs_w[None, None, :, :].expand(
|
||||
temporal_size, grid_size_h, -1, -1
|
||||
) # temporal_size, grid_size_h, grid_size_2, dim_w
|
||||
|
||||
freqs = torch.cat(
|
||||
[freqs_t, freqs_h, freqs_w], dim=-1
|
||||
) # temporal_size, grid_size_h, grid_size_w, (dim_t + dim_h + dim_w)
|
||||
freqs = freqs.view(
|
||||
temporal_size * grid_size_h * grid_size_w, -1
|
||||
) # (temporal_size * grid_size_h * grid_size_w), (dim_t + dim_h + dim_w)
|
||||
return freqs
|
||||
|
||||
t_cos, t_sin = freqs_t # both t_cos and t_sin has shape: temporal_size, dim_t
|
||||
h_cos, h_sin = freqs_h # both h_cos and h_sin has shape: grid_size_h, dim_h
|
||||
w_cos, w_sin = freqs_w # both w_cos and w_sin has shape: grid_size_w, dim_w
|
||||
|
||||
if grid_type == "slice":
|
||||
t_cos, t_sin = t_cos[:temporal_size], t_sin[:temporal_size]
|
||||
h_cos, h_sin = h_cos[:grid_size_h], h_sin[:grid_size_h]
|
||||
w_cos, w_sin = w_cos[:grid_size_w], w_sin[:grid_size_w]
|
||||
|
||||
cos = combine_time_height_width(t_cos, h_cos, w_cos)
|
||||
sin = combine_time_height_width(t_sin, h_sin, w_sin)
|
||||
return cos, sin
|
||||
|
||||
|
||||
# Similar to diffusers.pipelines.hunyuandit.pipeline_hunyuandit.get_resize_crop_region_for_grid
|
||||
def get_resize_crop_region_for_grid(src, tgt_width, tgt_height):
|
||||
tw = tgt_width
|
||||
@@ -218,7 +137,7 @@ def retrieve_timesteps(
|
||||
|
||||
|
||||
@dataclass
|
||||
class CogVideoXFunPipelineOutput(BaseOutput):
|
||||
class CogVideoX_Fun_PipelineOutput(BaseOutput):
|
||||
r"""
|
||||
Output class for CogVideo pipelines.
|
||||
|
||||
@@ -232,7 +151,7 @@ class CogVideoXFunPipelineOutput(BaseOutput):
|
||||
videos: torch.Tensor
|
||||
|
||||
|
||||
class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
class CogVideoX_Fun_Pipeline(DiffusionPipeline):
|
||||
r"""
|
||||
Pipeline for text-to-video generation using CogVideoX_Fun.
|
||||
|
||||
@@ -412,12 +331,6 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
def prepare_latents(
|
||||
self, batch_size, num_channels_latents, num_frames, height, width, dtype, device, generator, latents=None
|
||||
):
|
||||
if isinstance(generator, list) and len(generator) != batch_size:
|
||||
raise ValueError(
|
||||
f"You have passed a list of generators of length {len(generator)}, but requested an effective batch"
|
||||
f" size of {batch_size}. Make sure the batch size matches the length of the generators."
|
||||
)
|
||||
|
||||
shape = (
|
||||
batch_size,
|
||||
(num_frames - 1) // self.vae_scale_factor_temporal + 1,
|
||||
@@ -425,6 +338,11 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
height // self.vae_scale_factor_spatial,
|
||||
width // self.vae_scale_factor_spatial,
|
||||
)
|
||||
if isinstance(generator, list) and len(generator) != batch_size:
|
||||
raise ValueError(
|
||||
f"You have passed a list of generators of length {len(generator)}, but requested an effective batch"
|
||||
f" size of {batch_size}. Make sure the batch size matches the length of the generators."
|
||||
)
|
||||
|
||||
if latents is None:
|
||||
latents = randn_tensor(shape, generator=generator, device=device, dtype=dtype)
|
||||
@@ -537,36 +455,19 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
grid_height = height // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
grid_width = width // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
base_size_width = 720 // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
base_size_height = 480 // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
|
||||
p = self.transformer.config.patch_size
|
||||
p_t = self.transformer.config.patch_size_t
|
||||
|
||||
base_size_width = self.transformer.config.sample_width // p
|
||||
base_size_height = self.transformer.config.sample_height // p
|
||||
|
||||
if p_t is None:
|
||||
# CogVideoX 1.0
|
||||
grid_crops_coords = get_resize_crop_region_for_grid(
|
||||
(grid_height, grid_width), base_size_width, base_size_height
|
||||
)
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=grid_crops_coords,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=num_frames,
|
||||
)
|
||||
else:
|
||||
# CogVideoX 1.5
|
||||
base_num_frames = (num_frames + p_t - 1) // p_t
|
||||
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=None,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=base_num_frames,
|
||||
grid_type="slice",
|
||||
max_size=(base_size_height, base_size_width),
|
||||
)
|
||||
grid_crops_coords = get_resize_crop_region_for_grid(
|
||||
(grid_height, grid_width), base_size_width, base_size_height
|
||||
)
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=grid_crops_coords,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=num_frames,
|
||||
use_real=True,
|
||||
)
|
||||
|
||||
freqs_cos = freqs_cos.to(device=device)
|
||||
freqs_sin = freqs_sin.to(device=device)
|
||||
@@ -580,10 +481,6 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
def num_timesteps(self):
|
||||
return self._num_timesteps
|
||||
|
||||
@property
|
||||
def attention_kwargs(self):
|
||||
return self._attention_kwargs
|
||||
|
||||
@property
|
||||
def interrupt(self):
|
||||
return self._interrupt
|
||||
@@ -607,15 +504,14 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
latents: Optional[torch.FloatTensor] = None,
|
||||
prompt_embeds: Optional[torch.FloatTensor] = None,
|
||||
negative_prompt_embeds: Optional[torch.FloatTensor] = None,
|
||||
output_type: str = "pil",
|
||||
return_dict: bool = True,
|
||||
output_type: str = "numpy",
|
||||
return_dict: bool = False,
|
||||
callback_on_step_end: Optional[
|
||||
Union[Callable[[int, int, Dict], None], PipelineCallback, MultiPipelineCallbacks]
|
||||
] = None,
|
||||
attention_kwargs: Optional[Dict[str, Any]] = None,
|
||||
callback_on_step_end_tensor_inputs: List[str] = ["latents"],
|
||||
max_sequence_length: int = 226,
|
||||
) -> Union[CogVideoXFunPipelineOutput, Tuple]:
|
||||
) -> Union[CogVideoX_Fun_PipelineOutput, Tuple]:
|
||||
"""
|
||||
Function invoked when calling the pipeline for generation.
|
||||
|
||||
@@ -687,18 +583,21 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
Examples:
|
||||
|
||||
Returns:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoXFunPipelineOutput`] or `tuple`:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoXFunPipelineOutput`] if `return_dict` is True, otherwise a
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoX_Fun_PipelineOutput`] or `tuple`:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoX_Fun_PipelineOutput`] if `return_dict` is True, otherwise a
|
||||
`tuple`. When returning a tuple, the first element is a list with the generated images.
|
||||
"""
|
||||
|
||||
if num_frames > 49:
|
||||
raise ValueError(
|
||||
"The number of frames must be less than 49 for now due to static positional embeddings. This will be updated in the future to remove this limitation."
|
||||
)
|
||||
|
||||
if isinstance(callback_on_step_end, (PipelineCallback, MultiPipelineCallbacks)):
|
||||
callback_on_step_end_tensor_inputs = callback_on_step_end.tensor_inputs
|
||||
|
||||
height = height or self.transformer.config.sample_height * self.vae_scale_factor_spatial
|
||||
width = width or self.transformer.config.sample_width * self.vae_scale_factor_spatial
|
||||
num_frames = num_frames or self.transformer.config.sample_frames
|
||||
|
||||
height = height or self.transformer.config.sample_size * self.vae_scale_factor_spatial
|
||||
width = width or self.transformer.config.sample_size * self.vae_scale_factor_spatial
|
||||
num_videos_per_prompt = 1
|
||||
|
||||
# 1. Check inputs. Raise error if not correct
|
||||
@@ -712,7 +611,6 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
negative_prompt_embeds,
|
||||
)
|
||||
self._guidance_scale = guidance_scale
|
||||
self._attention_kwargs = attention_kwargs
|
||||
self._interrupt = False
|
||||
|
||||
# 2. Default call parameters
|
||||
@@ -748,16 +646,7 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
timesteps, num_inference_steps = retrieve_timesteps(self.scheduler, num_inference_steps, device, timesteps)
|
||||
self._num_timesteps = len(timesteps)
|
||||
|
||||
# 5. Prepare latents
|
||||
latent_frames = (num_frames - 1) // self.vae_scale_factor_temporal + 1
|
||||
|
||||
# For CogVideoX 1.5, the latent frames should be padded to make it divisible by patch_size_t
|
||||
patch_size_t = self.transformer.config.patch_size_t
|
||||
additional_frames = 0
|
||||
if num_frames != 1 and patch_size_t is not None and latent_frames % patch_size_t != 0:
|
||||
additional_frames = patch_size_t - latent_frames % patch_size_t
|
||||
num_frames += additional_frames * self.vae_scale_factor_temporal
|
||||
|
||||
# 5. Prepare latents.
|
||||
latent_channels = self.transformer.config.in_channels
|
||||
latents = self.prepare_latents(
|
||||
batch_size * num_videos_per_prompt,
|
||||
@@ -845,9 +734,11 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
if i == len(timesteps) - 1 or ((i + 1) > num_warmup_steps and (i + 1) % self.scheduler.order == 0):
|
||||
progress_bar.update()
|
||||
|
||||
if output_type == "pil":
|
||||
if output_type == "numpy":
|
||||
video = self.decode_latents(latents)
|
||||
video = torch.from_numpy(video)
|
||||
elif not output_type == "latent":
|
||||
video = self.decode_latents(latents)
|
||||
video = self.video_processor.postprocess_video(video=video, output_type=output_type)
|
||||
else:
|
||||
video = latents
|
||||
|
||||
@@ -855,6 +746,6 @@ class CogVideoXFunPipeline(DiffusionPipeline):
|
||||
self.maybe_free_model_hooks()
|
||||
|
||||
if not return_dict:
|
||||
return video
|
||||
video = torch.from_numpy(video)
|
||||
|
||||
return CogVideoXFunPipelineOutput(videos=video)
|
||||
return CogVideoX_Fun_PipelineOutput(videos=video)
|
||||
@@ -16,24 +16,24 @@
|
||||
import inspect
|
||||
import math
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Callable, Dict, List, Optional, Tuple, Union
|
||||
from typing import Callable, Dict, List, Optional, Tuple, Union
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from einops import rearrange
|
||||
from transformers import T5EncoderModel, T5Tokenizer
|
||||
|
||||
from diffusers.callbacks import MultiPipelineCallbacks, PipelineCallback
|
||||
from diffusers.image_processor import VaeImageProcessor
|
||||
from diffusers.models.embeddings import get_1d_rotary_pos_embed
|
||||
from diffusers.models import AutoencoderKLCogVideoX, CogVideoXTransformer3DModel
|
||||
from diffusers.models.embeddings import get_3d_rotary_pos_embed
|
||||
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
|
||||
from diffusers.schedulers import CogVideoXDDIMScheduler, CogVideoXDPMScheduler
|
||||
from diffusers.utils import BaseOutput, logging, replace_example_docstring
|
||||
from diffusers.utils.torch_utils import randn_tensor
|
||||
from diffusers.video_processor import VideoProcessor
|
||||
from diffusers.image_processor import VaeImageProcessor
|
||||
from einops import rearrange
|
||||
|
||||
from ..models import (AutoencoderKLCogVideoX,
|
||||
CogVideoXTransformer3DModel, T5EncoderModel,
|
||||
T5Tokenizer)
|
||||
|
||||
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
@@ -41,104 +41,25 @@ logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
EXAMPLE_DOC_STRING = """
|
||||
Examples:
|
||||
```python
|
||||
pass
|
||||
>>> import torch
|
||||
>>> from diffusers import CogVideoX_Fun_Pipeline
|
||||
>>> from diffusers.utils import export_to_video
|
||||
|
||||
>>> # Models: "THUDM/CogVideoX-2b" or "THUDM/CogVideoX-5b"
|
||||
>>> pipe = CogVideoX_Fun_Pipeline.from_pretrained("THUDM/CogVideoX-2b", torch_dtype=torch.float16).to("cuda")
|
||||
>>> prompt = (
|
||||
... "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. "
|
||||
... "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other "
|
||||
... "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, "
|
||||
... "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. "
|
||||
... "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical "
|
||||
... "atmosphere of this unique musical performance."
|
||||
... )
|
||||
>>> video = pipe(prompt=prompt, guidance_scale=6, num_inference_steps=50).frames[0]
|
||||
>>> export_to_video(video, "output.mp4", fps=8)
|
||||
```
|
||||
"""
|
||||
|
||||
# Copied from diffusers.models.embeddings.get_3d_rotary_pos_embed
|
||||
def get_3d_rotary_pos_embed(
|
||||
embed_dim,
|
||||
crops_coords,
|
||||
grid_size,
|
||||
temporal_size,
|
||||
theta: int = 10000,
|
||||
use_real: bool = True,
|
||||
grid_type: str = "linspace",
|
||||
max_size: Optional[Tuple[int, int]] = None,
|
||||
) -> Union[torch.Tensor, Tuple[torch.Tensor, torch.Tensor]]:
|
||||
"""
|
||||
RoPE for video tokens with 3D structure.
|
||||
|
||||
Args:
|
||||
embed_dim: (`int`):
|
||||
The embedding dimension size, corresponding to hidden_size_head.
|
||||
crops_coords (`Tuple[int]`):
|
||||
The top-left and bottom-right coordinates of the crop.
|
||||
grid_size (`Tuple[int]`):
|
||||
The grid size of the spatial positional embedding (height, width).
|
||||
temporal_size (`int`):
|
||||
The size of the temporal dimension.
|
||||
theta (`float`):
|
||||
Scaling factor for frequency computation.
|
||||
grid_type (`str`):
|
||||
Whether to use "linspace" or "slice" to compute grids.
|
||||
|
||||
Returns:
|
||||
`torch.Tensor`: positional embedding with shape `(temporal_size * grid_size[0] * grid_size[1], embed_dim/2)`.
|
||||
"""
|
||||
if use_real is not True:
|
||||
raise ValueError(" `use_real = False` is not currently supported for get_3d_rotary_pos_embed")
|
||||
|
||||
if grid_type == "linspace":
|
||||
start, stop = crops_coords
|
||||
grid_size_h, grid_size_w = grid_size
|
||||
grid_h = np.linspace(start[0], stop[0], grid_size_h, endpoint=False, dtype=np.float32)
|
||||
grid_w = np.linspace(start[1], stop[1], grid_size_w, endpoint=False, dtype=np.float32)
|
||||
grid_t = np.arange(temporal_size, dtype=np.float32)
|
||||
grid_t = np.linspace(0, temporal_size, temporal_size, endpoint=False, dtype=np.float32)
|
||||
elif grid_type == "slice":
|
||||
max_h, max_w = max_size
|
||||
grid_size_h, grid_size_w = grid_size
|
||||
grid_h = np.arange(max_h, dtype=np.float32)
|
||||
grid_w = np.arange(max_w, dtype=np.float32)
|
||||
grid_t = np.arange(temporal_size, dtype=np.float32)
|
||||
else:
|
||||
raise ValueError("Invalid value passed for `grid_type`.")
|
||||
|
||||
# Compute dimensions for each axis
|
||||
dim_t = embed_dim // 4
|
||||
dim_h = embed_dim // 8 * 3
|
||||
dim_w = embed_dim // 8 * 3
|
||||
|
||||
# Temporal frequencies
|
||||
freqs_t = get_1d_rotary_pos_embed(dim_t, grid_t, use_real=True)
|
||||
# Spatial frequencies for height and width
|
||||
freqs_h = get_1d_rotary_pos_embed(dim_h, grid_h, use_real=True)
|
||||
freqs_w = get_1d_rotary_pos_embed(dim_w, grid_w, use_real=True)
|
||||
|
||||
# BroadCast and concatenate temporal and spaial frequencie (height and width) into a 3d tensor
|
||||
def combine_time_height_width(freqs_t, freqs_h, freqs_w):
|
||||
freqs_t = freqs_t[:, None, None, :].expand(
|
||||
-1, grid_size_h, grid_size_w, -1
|
||||
) # temporal_size, grid_size_h, grid_size_w, dim_t
|
||||
freqs_h = freqs_h[None, :, None, :].expand(
|
||||
temporal_size, -1, grid_size_w, -1
|
||||
) # temporal_size, grid_size_h, grid_size_2, dim_h
|
||||
freqs_w = freqs_w[None, None, :, :].expand(
|
||||
temporal_size, grid_size_h, -1, -1
|
||||
) # temporal_size, grid_size_h, grid_size_2, dim_w
|
||||
|
||||
freqs = torch.cat(
|
||||
[freqs_t, freqs_h, freqs_w], dim=-1
|
||||
) # temporal_size, grid_size_h, grid_size_w, (dim_t + dim_h + dim_w)
|
||||
freqs = freqs.view(
|
||||
temporal_size * grid_size_h * grid_size_w, -1
|
||||
) # (temporal_size * grid_size_h * grid_size_w), (dim_t + dim_h + dim_w)
|
||||
return freqs
|
||||
|
||||
t_cos, t_sin = freqs_t # both t_cos and t_sin has shape: temporal_size, dim_t
|
||||
h_cos, h_sin = freqs_h # both h_cos and h_sin has shape: grid_size_h, dim_h
|
||||
w_cos, w_sin = freqs_w # both w_cos and w_sin has shape: grid_size_w, dim_w
|
||||
|
||||
if grid_type == "slice":
|
||||
t_cos, t_sin = t_cos[:temporal_size], t_sin[:temporal_size]
|
||||
h_cos, h_sin = h_cos[:grid_size_h], h_sin[:grid_size_h]
|
||||
w_cos, w_sin = w_cos[:grid_size_w], w_sin[:grid_size_w]
|
||||
|
||||
cos = combine_time_height_width(t_cos, h_cos, w_cos)
|
||||
sin = combine_time_height_width(t_sin, h_sin, w_sin)
|
||||
return cos, sin
|
||||
|
||||
|
||||
# Similar to diffusers.pipelines.hunyuandit.pipeline_hunyuandit.get_resize_crop_region_for_grid
|
||||
def get_resize_crop_region_for_grid(src, tgt_width, tgt_height):
|
||||
@@ -256,21 +177,8 @@ def resize_mask(mask, latent, process_first_frame_only=True):
|
||||
return resized_mask
|
||||
|
||||
|
||||
def add_noise_to_reference_video(image, ratio=None):
|
||||
if ratio is None:
|
||||
sigma = torch.normal(mean=-3.0, std=0.5, size=(image.shape[0],)).to(image.device)
|
||||
sigma = torch.exp(sigma).to(image.dtype)
|
||||
else:
|
||||
sigma = torch.ones((image.shape[0],)).to(image.device, image.dtype) * ratio
|
||||
|
||||
image_noise = torch.randn_like(image) * sigma[:, None, None, None, None]
|
||||
image_noise = torch.where(image==-1, torch.zeros_like(image), image_noise)
|
||||
image = image + image_noise
|
||||
return image
|
||||
|
||||
|
||||
@dataclass
|
||||
class CogVideoXFunPipelineOutput(BaseOutput):
|
||||
class CogVideoX_Fun_PipelineOutput(BaseOutput):
|
||||
r"""
|
||||
Output class for CogVideo pipelines.
|
||||
|
||||
@@ -284,7 +192,7 @@ class CogVideoXFunPipelineOutput(BaseOutput):
|
||||
videos: torch.Tensor
|
||||
|
||||
|
||||
class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
class CogVideoX_Fun_Pipeline_Inpaint(DiffusionPipeline):
|
||||
r"""
|
||||
Pipeline for text-to-video generation using CogVideoX.
|
||||
|
||||
@@ -308,7 +216,7 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
"""
|
||||
|
||||
_optional_components = []
|
||||
model_cpu_offload_seq = "text_encoder->transformer->vae"
|
||||
model_cpu_offload_seq = "text_encoder->vae->transformer->vae"
|
||||
|
||||
_callback_tensor_inputs = [
|
||||
"latents",
|
||||
@@ -536,7 +444,7 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
return outputs
|
||||
|
||||
def prepare_mask_latents(
|
||||
self, mask, masked_image, batch_size, height, width, dtype, device, generator, do_classifier_free_guidance, noise_aug_strength
|
||||
self, mask, masked_image, batch_size, height, width, dtype, device, generator, do_classifier_free_guidance
|
||||
):
|
||||
# resize the mask to latents shape as we concatenate the mask to the latents
|
||||
# we do that before converting to dtype to avoid breaking in case we're using cpu_offload
|
||||
@@ -555,8 +463,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
mask = mask * self.vae.config.scaling_factor
|
||||
|
||||
if masked_image is not None:
|
||||
if self.transformer.config.add_noise_in_inpaint_model:
|
||||
masked_image = add_noise_to_reference_video(masked_image, ratio=noise_aug_strength)
|
||||
masked_image = masked_image.to(device=device, dtype=self.vae.dtype)
|
||||
bs = 1
|
||||
new_mask_pixel_values = []
|
||||
@@ -674,36 +580,19 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
grid_height = height // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
grid_width = width // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
base_size_width = 720 // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
base_size_height = 480 // (self.vae_scale_factor_spatial * self.transformer.config.patch_size)
|
||||
|
||||
p = self.transformer.config.patch_size
|
||||
p_t = self.transformer.config.patch_size_t
|
||||
|
||||
base_size_width = self.transformer.config.sample_width // p
|
||||
base_size_height = self.transformer.config.sample_height // p
|
||||
|
||||
if p_t is None:
|
||||
# CogVideoX 1.0
|
||||
grid_crops_coords = get_resize_crop_region_for_grid(
|
||||
(grid_height, grid_width), base_size_width, base_size_height
|
||||
)
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=grid_crops_coords,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=num_frames,
|
||||
)
|
||||
else:
|
||||
# CogVideoX 1.5
|
||||
base_num_frames = (num_frames + p_t - 1) // p_t
|
||||
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=None,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=base_num_frames,
|
||||
grid_type="slice",
|
||||
max_size=(base_size_height, base_size_width),
|
||||
)
|
||||
grid_crops_coords = get_resize_crop_region_for_grid(
|
||||
(grid_height, grid_width), base_size_width, base_size_height
|
||||
)
|
||||
freqs_cos, freqs_sin = get_3d_rotary_pos_embed(
|
||||
embed_dim=self.transformer.config.attention_head_dim,
|
||||
crops_coords=grid_crops_coords,
|
||||
grid_size=(grid_height, grid_width),
|
||||
temporal_size=num_frames,
|
||||
use_real=True,
|
||||
)
|
||||
|
||||
freqs_cos = freqs_cos.to(device=device)
|
||||
freqs_sin = freqs_sin.to(device=device)
|
||||
@@ -717,10 +606,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
def num_timesteps(self):
|
||||
return self._num_timesteps
|
||||
|
||||
@property
|
||||
def attention_kwargs(self):
|
||||
return self._attention_kwargs
|
||||
|
||||
@property
|
||||
def interrupt(self):
|
||||
return self._interrupt
|
||||
@@ -757,18 +642,16 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
latents: Optional[torch.FloatTensor] = None,
|
||||
prompt_embeds: Optional[torch.FloatTensor] = None,
|
||||
negative_prompt_embeds: Optional[torch.FloatTensor] = None,
|
||||
output_type: str = "pil",
|
||||
return_dict: bool = True,
|
||||
output_type: str = "numpy",
|
||||
return_dict: bool = False,
|
||||
callback_on_step_end: Optional[
|
||||
Union[Callable[[int, int, Dict], None], PipelineCallback, MultiPipelineCallbacks]
|
||||
] = None,
|
||||
attention_kwargs: Optional[Dict[str, Any]] = None,
|
||||
callback_on_step_end_tensor_inputs: List[str] = ["latents"],
|
||||
max_sequence_length: int = 226,
|
||||
strength: float = 1,
|
||||
noise_aug_strength: float = 0.0563,
|
||||
comfyui_progressbar: bool = False,
|
||||
) -> Union[CogVideoXFunPipelineOutput, Tuple]:
|
||||
) -> Union[CogVideoX_Fun_PipelineOutput, Tuple]:
|
||||
"""
|
||||
Function invoked when calling the pipeline for generation.
|
||||
|
||||
@@ -840,18 +723,21 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
Examples:
|
||||
|
||||
Returns:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoXFunPipelineOutput`] or `tuple`:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoXFunPipelineOutput`] if `return_dict` is True, otherwise a
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoX_Fun_PipelineOutput`] or `tuple`:
|
||||
[`~pipelines.cogvideo.pipeline_cogvideox.CogVideoX_Fun_PipelineOutput`] if `return_dict` is True, otherwise a
|
||||
`tuple`. When returning a tuple, the first element is a list with the generated images.
|
||||
"""
|
||||
|
||||
if num_frames > 49:
|
||||
raise ValueError(
|
||||
"The number of frames must be less than 49 for now due to static positional embeddings. This will be updated in the future to remove this limitation."
|
||||
)
|
||||
|
||||
if isinstance(callback_on_step_end, (PipelineCallback, MultiPipelineCallbacks)):
|
||||
callback_on_step_end_tensor_inputs = callback_on_step_end.tensor_inputs
|
||||
|
||||
height = height or self.transformer.config.sample_height * self.vae_scale_factor_spatial
|
||||
width = width or self.transformer.config.sample_width * self.vae_scale_factor_spatial
|
||||
num_frames = num_frames or self.transformer.config.sample_frames
|
||||
|
||||
height = height or self.transformer.config.sample_size * self.vae_scale_factor_spatial
|
||||
width = width or self.transformer.config.sample_size * self.vae_scale_factor_spatial
|
||||
num_videos_per_prompt = 1
|
||||
|
||||
# 1. Check inputs. Raise error if not correct
|
||||
@@ -865,7 +751,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
negative_prompt_embeds,
|
||||
)
|
||||
self._guidance_scale = guidance_scale
|
||||
self._attention_kwargs = attention_kwargs
|
||||
self._interrupt = False
|
||||
|
||||
# 2. Default call parameters
|
||||
@@ -920,23 +805,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
else:
|
||||
init_video = None
|
||||
|
||||
# Magvae needs the number of frames to be 4n + 1.
|
||||
local_latent_length = (num_frames - 1) // self.vae_scale_factor_temporal + 1
|
||||
# For CogVideoX 1.5, the latent frames should be clipped to make it divisible by patch_size_t
|
||||
patch_size_t = self.transformer.config.patch_size_t
|
||||
additional_frames = 0
|
||||
if patch_size_t is not None and local_latent_length % patch_size_t != 0:
|
||||
additional_frames = local_latent_length % patch_size_t
|
||||
num_frames -= additional_frames * self.vae_scale_factor_temporal
|
||||
if num_frames <= 0:
|
||||
num_frames = 1
|
||||
if video_length > num_frames:
|
||||
logger.warning("The length of condition video is not right, the latent frames should be clipped to make it divisible by patch_size_t. ")
|
||||
video_length = num_frames
|
||||
video = video[:, :, :video_length]
|
||||
init_video = init_video[:, :, :video_length]
|
||||
mask_video = mask_video[:, :, :video_length]
|
||||
|
||||
num_channels_latents = self.vae.config.latent_channels
|
||||
num_channels_transformer = self.transformer.config.in_channels
|
||||
return_image_latents = num_channels_transformer == num_channels_latents
|
||||
@@ -998,7 +866,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
device,
|
||||
generator,
|
||||
do_classifier_free_guidance,
|
||||
noise_aug_strength=noise_aug_strength,
|
||||
)
|
||||
mask_latents = resize_mask(1 - mask_condition, masked_video_latents)
|
||||
mask_latents = mask_latents.to(masked_video_latents.device) * self.vae.config.scaling_factor
|
||||
@@ -1119,9 +986,11 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
if comfyui_progressbar:
|
||||
pbar.update(1)
|
||||
|
||||
if output_type == "pil":
|
||||
if output_type == "numpy":
|
||||
video = self.decode_latents(latents)
|
||||
video = torch.from_numpy(video)
|
||||
elif not output_type == "latent":
|
||||
video = self.decode_latents(latents)
|
||||
video = self.video_processor.postprocess_video(video=video, output_type=output_type)
|
||||
else:
|
||||
video = latents
|
||||
|
||||
@@ -1129,6 +998,6 @@ class CogVideoXFunInpaintPipeline(DiffusionPipeline):
|
||||
self.maybe_free_model_hooks()
|
||||
|
||||
if not return_dict:
|
||||
return video
|
||||
video = torch.from_numpy(video)
|
||||
|
||||
return CogVideoXFunPipelineOutput(videos=video)
|
||||
return CogVideoX_Fun_PipelineOutput(videos=video)
|
||||
@@ -0,0 +1,477 @@
|
||||
# LoRA network module
|
||||
# reference:
|
||||
# https://github.com/microsoft/LoRA/blob/main/loralib/layers.py
|
||||
# https://github.com/cloneofsimo/lora/blob/master/lora_diffusion/lora.py
|
||||
# https://github.com/bmaltais/kohya_ss
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
import os
|
||||
from collections import defaultdict
|
||||
from io import BytesIO
|
||||
from typing import List, Optional, Type, Union
|
||||
|
||||
import safetensors.torch
|
||||
import torch
|
||||
import torch.utils.checkpoint
|
||||
from diffusers.models.lora import LoRACompatibleConv, LoRACompatibleLinear
|
||||
from safetensors.torch import load_file
|
||||
from transformers import T5EncoderModel
|
||||
|
||||
|
||||
class LoRAModule(torch.nn.Module):
|
||||
"""
|
||||
replaces forward method of the original Linear, instead of replacing the original Linear module.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
lora_name,
|
||||
org_module: torch.nn.Module,
|
||||
multiplier=1.0,
|
||||
lora_dim=4,
|
||||
alpha=1,
|
||||
dropout=None,
|
||||
rank_dropout=None,
|
||||
module_dropout=None,
|
||||
):
|
||||
"""if alpha == 0 or None, alpha is rank (no scaling)."""
|
||||
super().__init__()
|
||||
self.lora_name = lora_name
|
||||
|
||||
if org_module.__class__.__name__ == "Conv2d":
|
||||
in_dim = org_module.in_channels
|
||||
out_dim = org_module.out_channels
|
||||
else:
|
||||
in_dim = org_module.in_features
|
||||
out_dim = org_module.out_features
|
||||
|
||||
self.lora_dim = lora_dim
|
||||
if org_module.__class__.__name__ == "Conv2d":
|
||||
kernel_size = org_module.kernel_size
|
||||
stride = org_module.stride
|
||||
padding = org_module.padding
|
||||
self.lora_down = torch.nn.Conv2d(in_dim, self.lora_dim, kernel_size, stride, padding, bias=False)
|
||||
self.lora_up = torch.nn.Conv2d(self.lora_dim, out_dim, (1, 1), (1, 1), bias=False)
|
||||
else:
|
||||
self.lora_down = torch.nn.Linear(in_dim, self.lora_dim, bias=False)
|
||||
self.lora_up = torch.nn.Linear(self.lora_dim, out_dim, bias=False)
|
||||
|
||||
if type(alpha) == torch.Tensor:
|
||||
alpha = alpha.detach().float().numpy() # without casting, bf16 causes error
|
||||
alpha = self.lora_dim if alpha is None or alpha == 0 else alpha
|
||||
self.scale = alpha / self.lora_dim
|
||||
self.register_buffer("alpha", torch.tensor(alpha))
|
||||
|
||||
# same as microsoft's
|
||||
torch.nn.init.kaiming_uniform_(self.lora_down.weight, a=math.sqrt(5))
|
||||
torch.nn.init.zeros_(self.lora_up.weight)
|
||||
|
||||
self.multiplier = multiplier
|
||||
self.org_module = org_module # remove in applying
|
||||
self.dropout = dropout
|
||||
self.rank_dropout = rank_dropout
|
||||
self.module_dropout = module_dropout
|
||||
|
||||
def apply_to(self):
|
||||
self.org_forward = self.org_module.forward
|
||||
self.org_module.forward = self.forward
|
||||
del self.org_module
|
||||
|
||||
def forward(self, x, *args, **kwargs):
|
||||
weight_dtype = x.dtype
|
||||
org_forwarded = self.org_forward(x)
|
||||
|
||||
# module dropout
|
||||
if self.module_dropout is not None and self.training:
|
||||
if torch.rand(1) < self.module_dropout:
|
||||
return org_forwarded
|
||||
|
||||
lx = self.lora_down(x.to(self.lora_down.weight.dtype))
|
||||
|
||||
# normal dropout
|
||||
if self.dropout is not None and self.training:
|
||||
lx = torch.nn.functional.dropout(lx, p=self.dropout)
|
||||
|
||||
# rank dropout
|
||||
if self.rank_dropout is not None and self.training:
|
||||
mask = torch.rand((lx.size(0), self.lora_dim), device=lx.device) > self.rank_dropout
|
||||
if len(lx.size()) == 3:
|
||||
mask = mask.unsqueeze(1) # for Text Encoder
|
||||
elif len(lx.size()) == 4:
|
||||
mask = mask.unsqueeze(-1).unsqueeze(-1) # for Conv2d
|
||||
lx = lx * mask
|
||||
|
||||
# scaling for rank dropout: treat as if the rank is changed
|
||||
scale = self.scale * (1.0 / (1.0 - self.rank_dropout)) # redundant for readability
|
||||
else:
|
||||
scale = self.scale
|
||||
|
||||
lx = self.lora_up(lx)
|
||||
|
||||
return org_forwarded.to(weight_dtype) + lx.to(weight_dtype) * self.multiplier * scale
|
||||
|
||||
|
||||
def addnet_hash_legacy(b):
|
||||
"""Old model hash used by sd-webui-additional-networks for .safetensors format files"""
|
||||
m = hashlib.sha256()
|
||||
|
||||
b.seek(0x100000)
|
||||
m.update(b.read(0x10000))
|
||||
return m.hexdigest()[0:8]
|
||||
|
||||
|
||||
def addnet_hash_safetensors(b):
|
||||
"""New model hash used by sd-webui-additional-networks for .safetensors format files"""
|
||||
hash_sha256 = hashlib.sha256()
|
||||
blksize = 1024 * 1024
|
||||
|
||||
b.seek(0)
|
||||
header = b.read(8)
|
||||
n = int.from_bytes(header, "little")
|
||||
|
||||
offset = n + 8
|
||||
b.seek(offset)
|
||||
for chunk in iter(lambda: b.read(blksize), b""):
|
||||
hash_sha256.update(chunk)
|
||||
|
||||
return hash_sha256.hexdigest()
|
||||
|
||||
|
||||
def precalculate_safetensors_hashes(tensors, metadata):
|
||||
"""Precalculate the model hashes needed by sd-webui-additional-networks to
|
||||
save time on indexing the model later."""
|
||||
|
||||
# Because writing user metadata to the file can change the result of
|
||||
# sd_models.model_hash(), only retain the training metadata for purposes of
|
||||
# calculating the hash, as they are meant to be immutable
|
||||
metadata = {k: v for k, v in metadata.items() if k.startswith("ss_")}
|
||||
|
||||
bytes = safetensors.torch.save(tensors, metadata)
|
||||
b = BytesIO(bytes)
|
||||
|
||||
model_hash = addnet_hash_safetensors(b)
|
||||
legacy_hash = addnet_hash_legacy(b)
|
||||
return model_hash, legacy_hash
|
||||
|
||||
|
||||
class LoRANetwork(torch.nn.Module):
|
||||
TRANSFORMER_TARGET_REPLACE_MODULE = ["CogVideoXTransformer3DModel"]
|
||||
TEXT_ENCODER_TARGET_REPLACE_MODULE = ["T5LayerSelfAttention", "T5LayerFF", "BertEncoder"]
|
||||
LORA_PREFIX_TRANSFORMER = "lora_unet"
|
||||
LORA_PREFIX_TEXT_ENCODER = "lora_te"
|
||||
def __init__(
|
||||
self,
|
||||
text_encoder: Union[List[T5EncoderModel], T5EncoderModel],
|
||||
unet,
|
||||
multiplier: float = 1.0,
|
||||
lora_dim: int = 4,
|
||||
alpha: float = 1,
|
||||
dropout: Optional[float] = None,
|
||||
module_class: Type[object] = LoRAModule,
|
||||
add_lora_in_attn_temporal: bool = False,
|
||||
varbose: Optional[bool] = False,
|
||||
) -> None:
|
||||
super().__init__()
|
||||
self.multiplier = multiplier
|
||||
|
||||
self.lora_dim = lora_dim
|
||||
self.alpha = alpha
|
||||
self.dropout = dropout
|
||||
|
||||
print(f"create LoRA network. base dim (rank): {lora_dim}, alpha: {alpha}")
|
||||
print(f"neuron dropout: p={self.dropout}")
|
||||
|
||||
# create module instances
|
||||
def create_modules(
|
||||
is_unet: bool,
|
||||
root_module: torch.nn.Module,
|
||||
target_replace_modules: List[torch.nn.Module],
|
||||
) -> List[LoRAModule]:
|
||||
prefix = (
|
||||
self.LORA_PREFIX_TRANSFORMER
|
||||
if is_unet
|
||||
else self.LORA_PREFIX_TEXT_ENCODER
|
||||
)
|
||||
loras = []
|
||||
skipped = []
|
||||
for name, module in root_module.named_modules():
|
||||
if module.__class__.__name__ in target_replace_modules:
|
||||
for child_name, child_module in module.named_modules():
|
||||
is_linear = child_module.__class__.__name__ == "Linear" or child_module.__class__.__name__ == "LoRACompatibleLinear"
|
||||
is_conv2d = child_module.__class__.__name__ == "Conv2d" or child_module.__class__.__name__ == "LoRACompatibleConv"
|
||||
is_conv2d_1x1 = is_conv2d and child_module.kernel_size == (1, 1)
|
||||
|
||||
if not add_lora_in_attn_temporal:
|
||||
if "attn_temporal" in child_name:
|
||||
continue
|
||||
|
||||
if is_linear or is_conv2d:
|
||||
lora_name = prefix + "." + name + "." + child_name
|
||||
lora_name = lora_name.replace(".", "_")
|
||||
|
||||
dim = None
|
||||
alpha = None
|
||||
|
||||
if is_linear or is_conv2d_1x1:
|
||||
dim = self.lora_dim
|
||||
alpha = self.alpha
|
||||
|
||||
if dim is None or dim == 0:
|
||||
if is_linear or is_conv2d_1x1:
|
||||
skipped.append(lora_name)
|
||||
continue
|
||||
|
||||
lora = module_class(
|
||||
lora_name,
|
||||
child_module,
|
||||
self.multiplier,
|
||||
dim,
|
||||
alpha,
|
||||
dropout=dropout,
|
||||
)
|
||||
loras.append(lora)
|
||||
return loras, skipped
|
||||
|
||||
text_encoders = text_encoder if type(text_encoder) == list else [text_encoder]
|
||||
|
||||
self.text_encoder_loras = []
|
||||
skipped_te = []
|
||||
for i, text_encoder in enumerate(text_encoders):
|
||||
if text_encoder is not None:
|
||||
text_encoder_loras, skipped = create_modules(False, text_encoder, LoRANetwork.TEXT_ENCODER_TARGET_REPLACE_MODULE)
|
||||
self.text_encoder_loras.extend(text_encoder_loras)
|
||||
skipped_te += skipped
|
||||
print(f"create LoRA for Text Encoder: {len(self.text_encoder_loras)} modules.")
|
||||
|
||||
self.unet_loras, skipped_un = create_modules(True, unet, LoRANetwork.TRANSFORMER_TARGET_REPLACE_MODULE)
|
||||
print(f"create LoRA for U-Net: {len(self.unet_loras)} modules.")
|
||||
|
||||
# assertion
|
||||
names = set()
|
||||
for lora in self.text_encoder_loras + self.unet_loras:
|
||||
assert lora.lora_name not in names, f"duplicated lora name: {lora.lora_name}"
|
||||
names.add(lora.lora_name)
|
||||
|
||||
def apply_to(self, text_encoder, unet, apply_text_encoder=True, apply_unet=True):
|
||||
if apply_text_encoder:
|
||||
print("enable LoRA for text encoder")
|
||||
else:
|
||||
self.text_encoder_loras = []
|
||||
|
||||
if apply_unet:
|
||||
print("enable LoRA for U-Net")
|
||||
else:
|
||||
self.unet_loras = []
|
||||
|
||||
for lora in self.text_encoder_loras + self.unet_loras:
|
||||
lora.apply_to()
|
||||
self.add_module(lora.lora_name, lora)
|
||||
|
||||
def set_multiplier(self, multiplier):
|
||||
self.multiplier = multiplier
|
||||
for lora in self.text_encoder_loras + self.unet_loras:
|
||||
lora.multiplier = self.multiplier
|
||||
|
||||
def load_weights(self, file):
|
||||
if os.path.splitext(file)[1] == ".safetensors":
|
||||
from safetensors.torch import load_file
|
||||
|
||||
weights_sd = load_file(file)
|
||||
else:
|
||||
weights_sd = torch.load(file, map_location="cpu")
|
||||
info = self.load_state_dict(weights_sd, False)
|
||||
return info
|
||||
|
||||
def prepare_optimizer_params(self, text_encoder_lr, unet_lr, default_lr):
|
||||
self.requires_grad_(True)
|
||||
all_params = []
|
||||
|
||||
def enumerate_params(loras):
|
||||
params = []
|
||||
for lora in loras:
|
||||
params.extend(lora.parameters())
|
||||
return params
|
||||
|
||||
if self.text_encoder_loras:
|
||||
param_data = {"params": enumerate_params(self.text_encoder_loras)}
|
||||
if text_encoder_lr is not None:
|
||||
param_data["lr"] = text_encoder_lr
|
||||
all_params.append(param_data)
|
||||
|
||||
if self.unet_loras:
|
||||
param_data = {"params": enumerate_params(self.unet_loras)}
|
||||
if unet_lr is not None:
|
||||
param_data["lr"] = unet_lr
|
||||
all_params.append(param_data)
|
||||
|
||||
return all_params
|
||||
|
||||
def enable_gradient_checkpointing(self):
|
||||
pass
|
||||
|
||||
def get_trainable_params(self):
|
||||
return self.parameters()
|
||||
|
||||
def save_weights(self, file, dtype, metadata):
|
||||
if metadata is not None and len(metadata) == 0:
|
||||
metadata = None
|
||||
|
||||
state_dict = self.state_dict()
|
||||
|
||||
if dtype is not None:
|
||||
for key in list(state_dict.keys()):
|
||||
v = state_dict[key]
|
||||
v = v.detach().clone().to("cpu").to(dtype)
|
||||
state_dict[key] = v
|
||||
|
||||
if os.path.splitext(file)[1] == ".safetensors":
|
||||
from safetensors.torch import save_file
|
||||
|
||||
# Precalculate model hashes to save time on indexing
|
||||
if metadata is None:
|
||||
metadata = {}
|
||||
model_hash, legacy_hash = precalculate_safetensors_hashes(state_dict, metadata)
|
||||
metadata["sshs_model_hash"] = model_hash
|
||||
metadata["sshs_legacy_hash"] = legacy_hash
|
||||
|
||||
save_file(state_dict, file, metadata)
|
||||
else:
|
||||
torch.save(state_dict, file)
|
||||
|
||||
def create_network(
|
||||
multiplier: float,
|
||||
network_dim: Optional[int],
|
||||
network_alpha: Optional[float],
|
||||
text_encoder: Union[T5EncoderModel, List[T5EncoderModel]],
|
||||
transformer,
|
||||
neuron_dropout: Optional[float] = None,
|
||||
add_lora_in_attn_temporal: bool = False,
|
||||
**kwargs,
|
||||
):
|
||||
if network_dim is None:
|
||||
network_dim = 4 # default
|
||||
if network_alpha is None:
|
||||
network_alpha = 1.0
|
||||
|
||||
network = LoRANetwork(
|
||||
text_encoder,
|
||||
transformer,
|
||||
multiplier=multiplier,
|
||||
lora_dim=network_dim,
|
||||
alpha=network_alpha,
|
||||
dropout=neuron_dropout,
|
||||
add_lora_in_attn_temporal=add_lora_in_attn_temporal,
|
||||
varbose=True,
|
||||
)
|
||||
return network
|
||||
|
||||
def merge_lora(pipeline, lora_path, multiplier, device='cpu', dtype=torch.float32, state_dict=None, transformer_only=False):
|
||||
LORA_PREFIX_TRANSFORMER = "lora_unet"
|
||||
LORA_PREFIX_TEXT_ENCODER = "lora_te"
|
||||
if state_dict is None:
|
||||
state_dict = load_file(lora_path, device=device)
|
||||
else:
|
||||
state_dict = state_dict
|
||||
updates = defaultdict(dict)
|
||||
for key, value in state_dict.items():
|
||||
layer, elem = key.split('.', 1)
|
||||
updates[layer][elem] = value
|
||||
|
||||
for layer, elems in updates.items():
|
||||
|
||||
if "lora_te" in layer:
|
||||
if transformer_only:
|
||||
continue
|
||||
else:
|
||||
layer_infos = layer.split(LORA_PREFIX_TEXT_ENCODER + "_")[-1].split("_")
|
||||
curr_layer = pipeline.text_encoder
|
||||
else:
|
||||
layer_infos = layer.split(LORA_PREFIX_TRANSFORMER + "_")[-1].split("_")
|
||||
curr_layer = pipeline.transformer
|
||||
|
||||
temp_name = layer_infos.pop(0)
|
||||
while len(layer_infos) > -1:
|
||||
try:
|
||||
curr_layer = curr_layer.__getattr__(temp_name)
|
||||
if len(layer_infos) > 0:
|
||||
temp_name = layer_infos.pop(0)
|
||||
elif len(layer_infos) == 0:
|
||||
break
|
||||
except Exception:
|
||||
if len(layer_infos) == 0:
|
||||
print('Error loading layer')
|
||||
if len(temp_name) > 0:
|
||||
temp_name += "_" + layer_infos.pop(0)
|
||||
else:
|
||||
temp_name = layer_infos.pop(0)
|
||||
|
||||
weight_up = elems['lora_up.weight'].to(dtype)
|
||||
weight_down = elems['lora_down.weight'].to(dtype)
|
||||
if 'alpha' in elems.keys():
|
||||
alpha = elems['alpha'].item() / weight_up.shape[1]
|
||||
else:
|
||||
alpha = 1.0
|
||||
|
||||
curr_layer.weight.data = curr_layer.weight.data.to(device)
|
||||
if len(weight_up.shape) == 4:
|
||||
curr_layer.weight.data += multiplier * alpha * torch.mm(weight_up.squeeze(3).squeeze(2),
|
||||
weight_down.squeeze(3).squeeze(2)).unsqueeze(
|
||||
2).unsqueeze(3)
|
||||
else:
|
||||
curr_layer.weight.data += multiplier * alpha * torch.mm(weight_up, weight_down)
|
||||
|
||||
return pipeline
|
||||
|
||||
# TODO: Refactor with merge_lora.
|
||||
def unmerge_lora(pipeline, lora_path, multiplier=1, device="cpu", dtype=torch.float32):
|
||||
"""Unmerge state_dict in LoRANetwork from the pipeline in diffusers."""
|
||||
LORA_PREFIX_UNET = "lora_unet"
|
||||
LORA_PREFIX_TEXT_ENCODER = "lora_te"
|
||||
state_dict = load_file(lora_path, device=device)
|
||||
|
||||
updates = defaultdict(dict)
|
||||
for key, value in state_dict.items():
|
||||
layer, elem = key.split('.', 1)
|
||||
updates[layer][elem] = value
|
||||
|
||||
for layer, elems in updates.items():
|
||||
|
||||
if "lora_te" in layer:
|
||||
layer_infos = layer.split(LORA_PREFIX_TEXT_ENCODER + "_")[-1].split("_")
|
||||
curr_layer = pipeline.text_encoder
|
||||
else:
|
||||
layer_infos = layer.split(LORA_PREFIX_UNET + "_")[-1].split("_")
|
||||
curr_layer = pipeline.transformer
|
||||
|
||||
temp_name = layer_infos.pop(0)
|
||||
while len(layer_infos) > -1:
|
||||
try:
|
||||
curr_layer = curr_layer.__getattr__(temp_name)
|
||||
if len(layer_infos) > 0:
|
||||
temp_name = layer_infos.pop(0)
|
||||
elif len(layer_infos) == 0:
|
||||
break
|
||||
except Exception:
|
||||
if len(layer_infos) == 0:
|
||||
print('Error loading layer')
|
||||
if len(temp_name) > 0:
|
||||
temp_name += "_" + layer_infos.pop(0)
|
||||
else:
|
||||
temp_name = layer_infos.pop(0)
|
||||
|
||||
weight_up = elems['lora_up.weight'].to(dtype)
|
||||
weight_down = elems['lora_down.weight'].to(dtype)
|
||||
if 'alpha' in elems.keys():
|
||||
alpha = elems['alpha'].item() / weight_up.shape[1]
|
||||
else:
|
||||
alpha = 1.0
|
||||
|
||||
curr_layer.weight.data = curr_layer.weight.data.to(device)
|
||||
if len(weight_up.shape) == 4:
|
||||
curr_layer.weight.data -= multiplier * alpha * torch.mm(weight_up.squeeze(3).squeeze(2),
|
||||
weight_down.squeeze(3).squeeze(2)).unsqueeze(2).unsqueeze(3)
|
||||
else:
|
||||
curr_layer.weight.data -= multiplier * alpha * torch.mm(weight_up, weight_down)
|
||||
|
||||
return pipeline
|
||||
@@ -0,0 +1,189 @@
|
||||
import os
|
||||
import gc
|
||||
import imageio
|
||||
import numpy as np
|
||||
import torch
|
||||
import torchvision
|
||||
import cv2
|
||||
from einops import rearrange
|
||||
from PIL import Image
|
||||
|
||||
def get_width_and_height_from_image_and_base_resolution(image, base_resolution):
|
||||
target_pixels = int(base_resolution) * int(base_resolution)
|
||||
original_width, original_height = Image.open(image).size
|
||||
ratio = (target_pixels / (original_width * original_height)) ** 0.5
|
||||
width_slider = round(original_width * ratio)
|
||||
height_slider = round(original_height * ratio)
|
||||
return height_slider, width_slider
|
||||
|
||||
def color_transfer(sc, dc):
|
||||
"""
|
||||
Transfer color distribution from of sc, referred to dc.
|
||||
|
||||
Args:
|
||||
sc (numpy.ndarray): input image to be transfered.
|
||||
dc (numpy.ndarray): reference image
|
||||
|
||||
Returns:
|
||||
numpy.ndarray: Transferred color distribution on the sc.
|
||||
"""
|
||||
|
||||
def get_mean_and_std(img):
|
||||
x_mean, x_std = cv2.meanStdDev(img)
|
||||
x_mean = np.hstack(np.around(x_mean, 2))
|
||||
x_std = np.hstack(np.around(x_std, 2))
|
||||
return x_mean, x_std
|
||||
|
||||
sc = cv2.cvtColor(sc, cv2.COLOR_RGB2LAB)
|
||||
s_mean, s_std = get_mean_and_std(sc)
|
||||
dc = cv2.cvtColor(dc, cv2.COLOR_RGB2LAB)
|
||||
t_mean, t_std = get_mean_and_std(dc)
|
||||
img_n = ((sc - s_mean) * (t_std / s_std)) + t_mean
|
||||
np.putmask(img_n, img_n > 255, 255)
|
||||
np.putmask(img_n, img_n < 0, 0)
|
||||
dst = cv2.cvtColor(cv2.convertScaleAbs(img_n), cv2.COLOR_LAB2RGB)
|
||||
return dst
|
||||
|
||||
def save_videos_grid(videos: torch.Tensor, path: str, rescale=False, n_rows=6, fps=12, imageio_backend=True, color_transfer_post_process=False):
|
||||
videos = rearrange(videos, "b c t h w -> t b c h w")
|
||||
outputs = []
|
||||
for x in videos:
|
||||
x = torchvision.utils.make_grid(x, nrow=n_rows)
|
||||
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
|
||||
if rescale:
|
||||
x = (x + 1.0) / 2.0 # -1,1 -> 0,1
|
||||
x = (x * 255).numpy().astype(np.uint8)
|
||||
outputs.append(Image.fromarray(x))
|
||||
|
||||
if color_transfer_post_process:
|
||||
for i in range(1, len(outputs)):
|
||||
outputs[i] = Image.fromarray(color_transfer(np.uint8(outputs[i]), np.uint8(outputs[0])))
|
||||
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
if imageio_backend:
|
||||
if path.endswith("mp4"):
|
||||
imageio.mimsave(path, outputs, fps=fps)
|
||||
else:
|
||||
imageio.mimsave(path, outputs, duration=(1000 * 1/fps))
|
||||
else:
|
||||
if path.endswith("mp4"):
|
||||
path = path.replace('.mp4', '.gif')
|
||||
outputs[0].save(path, format='GIF', append_images=outputs, save_all=True, duration=100, loop=0)
|
||||
|
||||
def get_image_to_video_latent(validation_image_start, validation_image_end, video_length, sample_size):
|
||||
if validation_image_start is not None and validation_image_end is not None:
|
||||
if type(validation_image_start) is str and os.path.isfile(validation_image_start):
|
||||
image_start = clip_image = Image.open(validation_image_start).convert("RGB")
|
||||
image_start = image_start.resize([sample_size[1], sample_size[0]])
|
||||
clip_image = clip_image.resize([sample_size[1], sample_size[0]])
|
||||
else:
|
||||
image_start = clip_image = validation_image_start
|
||||
image_start = [_image_start.resize([sample_size[1], sample_size[0]]) for _image_start in image_start]
|
||||
clip_image = [_clip_image.resize([sample_size[1], sample_size[0]]) for _clip_image in clip_image]
|
||||
|
||||
if type(validation_image_end) is str and os.path.isfile(validation_image_end):
|
||||
image_end = Image.open(validation_image_end).convert("RGB")
|
||||
image_end = image_end.resize([sample_size[1], sample_size[0]])
|
||||
else:
|
||||
image_end = validation_image_end
|
||||
image_end = [_image_end.resize([sample_size[1], sample_size[0]]) for _image_end in image_end]
|
||||
|
||||
if type(image_start) is list:
|
||||
clip_image = clip_image[0]
|
||||
start_video = torch.cat(
|
||||
[torch.from_numpy(np.array(_image_start)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0) for _image_start in image_start],
|
||||
dim=2
|
||||
)
|
||||
input_video = torch.tile(start_video[:, :, :1], [1, 1, video_length, 1, 1])
|
||||
input_video[:, :, :len(image_start)] = start_video
|
||||
|
||||
input_video_mask = torch.zeros_like(input_video[:, :1])
|
||||
input_video_mask[:, :, len(image_start):] = 255
|
||||
else:
|
||||
input_video = torch.tile(
|
||||
torch.from_numpy(np.array(image_start)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0),
|
||||
[1, 1, video_length, 1, 1]
|
||||
)
|
||||
input_video_mask = torch.zeros_like(input_video[:, :1])
|
||||
input_video_mask[:, :, 3:] = 255
|
||||
|
||||
if type(image_end) is list:
|
||||
image_end = [_image_end.resize(image_start[0].size if type(image_start) is list else image_start.size) for _image_end in image_end]
|
||||
end_video = torch.cat(
|
||||
[torch.from_numpy(np.array(_image_end)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0) for _image_end in image_end],
|
||||
dim=2
|
||||
)
|
||||
input_video[:, :, -len(end_video):] = end_video
|
||||
|
||||
input_video_mask[:, :, -len(image_end):] = 0
|
||||
else:
|
||||
image_end = image_end.resize(image_start[0].size if type(image_start) is list else image_start.size)
|
||||
input_video[:, :, -3:] = torch.from_numpy(np.array(image_end)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0)
|
||||
input_video_mask[:, :, -3:] = 0
|
||||
|
||||
input_video = input_video / 255
|
||||
|
||||
elif validation_image_start is not None:
|
||||
if type(validation_image_start) is str and os.path.isfile(validation_image_start):
|
||||
image_start = clip_image = Image.open(validation_image_start).convert("RGB")
|
||||
image_start = image_start.resize([sample_size[1], sample_size[0]])
|
||||
clip_image = clip_image.resize([sample_size[1], sample_size[0]])
|
||||
else:
|
||||
image_start = clip_image = validation_image_start
|
||||
image_start = [_image_start.resize([sample_size[1], sample_size[0]]) for _image_start in image_start]
|
||||
clip_image = [_clip_image.resize([sample_size[1], sample_size[0]]) for _clip_image in clip_image]
|
||||
image_end = None
|
||||
|
||||
if type(image_start) is list:
|
||||
clip_image = clip_image[0]
|
||||
start_video = torch.cat(
|
||||
[torch.from_numpy(np.array(_image_start)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0) for _image_start in image_start],
|
||||
dim=2
|
||||
)
|
||||
input_video = torch.tile(start_video[:, :, :1], [1, 1, video_length, 1, 1])
|
||||
input_video[:, :, :len(image_start)] = start_video
|
||||
input_video = input_video / 255
|
||||
|
||||
input_video_mask = torch.zeros_like(input_video[:, :1])
|
||||
input_video_mask[:, :, len(image_start):] = 255
|
||||
else:
|
||||
input_video = torch.tile(
|
||||
torch.from_numpy(np.array(image_start)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0),
|
||||
[1, 1, video_length, 1, 1]
|
||||
) / 255
|
||||
input_video_mask = torch.zeros_like(input_video[:, :1])
|
||||
input_video_mask[:, :, 3:, ] = 255
|
||||
else:
|
||||
image_start = None
|
||||
image_end = None
|
||||
input_video = torch.zeros([1, 3, video_length, sample_size[0], sample_size[1]])
|
||||
input_video_mask = torch.ones([1, 1, video_length, sample_size[0], sample_size[1]]) * 255
|
||||
clip_image = None
|
||||
|
||||
del image_start
|
||||
del image_end
|
||||
gc.collect()
|
||||
|
||||
return input_video, input_video_mask, clip_image
|
||||
|
||||
def get_video_to_video_latent(input_video_path, video_length, sample_size):
|
||||
if type(input_video_path) is str:
|
||||
cap = cv2.VideoCapture(input_video_path)
|
||||
input_video = []
|
||||
while True:
|
||||
ret, frame = cap.read()
|
||||
if not ret:
|
||||
break
|
||||
frame = cv2.resize(frame, (sample_size[1], sample_size[0]))
|
||||
input_video.append(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB))
|
||||
cap.release()
|
||||
else:
|
||||
input_video = input_video_path
|
||||
|
||||
input_video = torch.from_numpy(np.array(input_video))[:video_length]
|
||||
input_video = input_video.permute([3, 0, 1, 2]).unsqueeze(0) / 255
|
||||
|
||||
input_video_mask = torch.zeros_like(input_video[:, :1])
|
||||
input_video_mask[:, :, :] = 255
|
||||
|
||||
return input_video, input_video_mask, None
|
||||
@@ -1,7 +1,7 @@
|
||||
# Video Caption
|
||||
English | [简体中文](./README_zh-CN.md)
|
||||
|
||||
The folder contains codes for dataset preprocessing (i.e., video splitting, filtering, and recaptioning), and beautiful prompt used by EasyAnimate.
|
||||
The folder contains codes for dataset preprocessing (i.e., video splitting, filtering, and recaptioning), and beautiful prompt used by CogVideoX-Fun.
|
||||
The entire process supports distributed parallel processing, capable of handling large-scale datasets.
|
||||
|
||||
Meanwhile, we are collaborating with [Data-Juicer](https://github.com/modelscope/data-juicer/blob/main/docs/DJ_SORA.md),
|
||||
@@ -17,7 +17,7 @@ allowing you to easily perform video data processing on [Aliyun PAI-DLC](https:/
|
||||
- [Video Splitting](#video-splitting)
|
||||
- [Video Filtering](#video-filtering)
|
||||
- [Video Recaptioning](#video-recaptioning)
|
||||
- [Beautiful Prompt (For EasyAnimate Inference)](#beautiful-prompt-for-easyanimate-inference)
|
||||
- [Beautiful Prompt (For CogVideoX-Fun Inference)](#beautiful-prompt-for-cogvideox-inference)
|
||||
- [Batched Inference](#batched-inference)
|
||||
- [OpenAI Server](#openai-server)
|
||||
|
||||
@@ -27,18 +27,21 @@ allowing you to easily perform video data processing on [Aliyun PAI-DLC](https:/
|
||||
AliyunDSW or Docker is recommended to setup the environment, please refer to [Quick Start](../../README.md#quick-start).
|
||||
You can also refer to the image build process in the [Dockerfile](../../Dockerfile.ds) to configure the conda environment and other dependencies locally.
|
||||
|
||||
Since the video recaptioning depends on [llm-awq](https://github.com/mit-han-lab/llm-awq) for faster and memory efficient inference,
|
||||
the minimum GPU requirment should be RTX 3060 or A2 (CUDA Compute Capability >= 8.0).
|
||||
|
||||
```shell
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:easyanimate
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:easyanimate
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/EasyAnimate.git
|
||||
git clone https://github.com/aigc-apps/CogVideoX-Fun.git
|
||||
|
||||
# enter video_caption
|
||||
cd EasyAnimate/easyanimate/video_caption
|
||||
cd CogVideoX-Fun/cogvideox/video_caption
|
||||
```
|
||||
|
||||
### Data Preprocessing
|
||||
@@ -55,7 +58,7 @@ Taking Panda-70M as an example, the entire dataset directory structure is shown
|
||||
```
|
||||
|
||||
#### Video Splitting
|
||||
EasyAnimate utilizes [PySceneDetect](https://github.com/Breakthrough/PySceneDetect) to identify scene changes within the video
|
||||
CogVideoX-Fun utilizes [PySceneDetect](https://github.com/Breakthrough/PySceneDetect) to identify scene changes within the video
|
||||
and performs video splitting via FFmpeg based on certain threshold values to ensure consistency of the video clip.
|
||||
Video clips shorter than 3 seconds will be discarded, and those longer than 10 seconds will be splitted recursively.
|
||||
|
||||
@@ -64,14 +67,12 @@ After running
|
||||
```shell
|
||||
sh scripts/stage_1_video_splitting.sh
|
||||
```
|
||||
the video clips are obtained in `easyanimate/video_caption/datasets/panda_70m/videos_clips/data/`.
|
||||
the video clips are obtained in `cogvideox/video_caption/datasets/panda_70m/videos_clips/data/`.
|
||||
|
||||
#### Video Filtering
|
||||
Based on the videos obtained in the previous step, EasyAnimate provides a simple yet effective pipeline to filter out high-quality videos for recaptioning.
|
||||
Based on the videos obtained in the previous step, CogVideoX-Fun provides a simple yet effective pipeline to filter out high-quality videos for recaptioning.
|
||||
The overall process is as follows:
|
||||
|
||||
- Scene transition filtering: Filter out videos with scene transition introduced by missing or superfluous splitting of PySceneDetect by calculating the semantic similarity
|
||||
accoss the beginning frame, the last frame, and the keyframes via [CLIP](https://github.com/openai/CLIP) or [DINOv2](https://github.com/facebookresearch/dinov2).
|
||||
- Aesthetic filtering: Filter out videos with poor content (blurry, dim, etc.) by calculating the average aesthetic score of uniformly sampled 4 frames via [aesthetic-predictor-v2-5](https://github.com/discus0434/aesthetic-predictor-v2-5).
|
||||
- Text filtering: Use [EasyOCR](https://github.com/JaidedAI/EasyOCR) to calculate the text area proportion of the middle frame to filter out videos with a large area of text.
|
||||
- Motion filtering: Calculate interframe optical flow differences to filter out videos that move too slowly or too quickly.
|
||||
@@ -81,25 +82,23 @@ After running
|
||||
```shell
|
||||
sh scripts/stage_2_video_filtering.sh
|
||||
```
|
||||
the semantic consistency score, aesthetic score, text score, and motion score of videos will be saved
|
||||
in the corresponding meta files in the folder `easyanimate/video_caption/datasets/panda_70m/videos_clips/`.
|
||||
the aesthetic score, text score, and motion score of videos will be saved in the corresponding meta files in the folder `cogvideox/video_caption/datasets/panda_70m/videos_clips/`.
|
||||
|
||||
> [!NOTE]
|
||||
> The computation of semantic consistency score depends on the [openai/clip-vit-large-patch14-336](https://huggingface.co/openai/clip-vit-large-patch14-336).
|
||||
Meanwhile, the aesthetic score depends on the [google/siglip-so400m-patch14-384 model](https://huggingface.co/google/siglip-so400m-patch14-384).
|
||||
> The computation of the aesthetic score depends on the [google/siglip-so400m-patch14-384 model](https://huggingface.co/google/siglip-so400m-patch14-384).
|
||||
Please run `HF_ENDPOINT=https://hf-mirror.com sh scripts/stage_2_video_filtering.sh` if you cannot access to huggingface.com.
|
||||
|
||||
|
||||
#### Video Recaptioning
|
||||
After obtaining the aboved high-quality filtered videos, EasyAnimate utilizes [InternVL2](https://internvl.readthedocs.io/en/latest/internvl2.0/introduction.html) to perform video recaptioning.
|
||||
Subsequently, the recaptioning results are rewritten by LLMs to better meet with the requirements of video generation tasks.
|
||||
Finally, an advanced [VideoCLIP-XL](https://arxiv.org/abs/2410.00741) model is used to filter out (video, long caption) pairs with poor alignment, resulting in the final training dataset.
|
||||
After obtaining the aboved high-quality filtered videos, CogVideoX-Fun utilizes [VILA1.5](https://github.com/NVlabs/VILA) to perform video recaptioning.
|
||||
Subsequently, the recaptioning results are rewritten by LLMs to better meet with the requirements of video generation tasks.
|
||||
Finally, an advanced VideoCLIPXL model is developed to filter out video-caption pairs with poor alignment, resulting in the final training dataset.
|
||||
|
||||
Please download the video caption model from [InternVL2](https://huggingface.co/collections/OpenGVLab/internvl-20-667d3961ab5eb12c7ed1463e) of the appropriate size based on the GPU memory of your machine.
|
||||
For A100 with 40G VRAM, you can download [InternVL2-40B-AWQ](https://huggingface.co/OpenGVLab/InternVL2-40B-AWQ) by running
|
||||
Please download the video caption model from [VILA1.5](https://huggingface.co/collections/Efficient-Large-Model/vila-on-pre-training-for-visual-language-models-65d8022a3a52cd9bcd62698e) of the appropriate size based on the GPU memory of your machine.
|
||||
For A100 with 40G VRAM, you can download [VILA1.5-40b-AWQ](https://huggingface.co/Efficient-Large-Model/VILA1.5-40b-AWQ) by running
|
||||
```shell
|
||||
# Add HF_ENDPOINT=https://hf-mirror.com before the command if you cannot access to huggingface.com
|
||||
huggingface-cli download OpenGVLab/InternVL2-40B-AWQ --local-dir-use-symlinks False --local-dir /PATH/TO/INTERNVL2_MODEL
|
||||
huggingface-cli download Efficient-Large-Model/VILA1.5-40b-AWQ --local-dir-use-symlinks False --local-dir /PATH/TO/VILA_MODEL
|
||||
```
|
||||
|
||||
Optionally, you can prepare local LLMs to rewrite the recaption results.
|
||||
@@ -112,18 +111,18 @@ huggingface-cli download NousResearch/Meta-Llama-3-8B-Instruct --local-dir-use-s
|
||||
The entire workflow of video recaption is in the [stage_3_video_recaptioning.sh](./scripts/stage_3_video_recaptioning.sh).
|
||||
After running
|
||||
```shell
|
||||
CAPTION_MODEL_PATH=/PATH/TO/INTERNVL2_MODEL REWRITE_MODEL_PATH=/PATH/TO/REWRITE_MODEL sh scripts/stage_3_video_recaptioning.sh
|
||||
VILA_MODEL_PATH=/PATH/TO/VILA_MODEL REWRITE_MODEL_PATH=/PATH/TO/REWRITE_MODEL sh scripts/stage_3_video_recaptioning.sh
|
||||
```
|
||||
the final train file is obtained in `easyanimate/video_caption/datasets/panda_70m/videos_clips/meta_train_info.json`.
|
||||
the final train file is obtained in `cogvideox/video_caption/datasets/panda_70m/videos_clips/meta_train_info.json`.
|
||||
|
||||
|
||||
### Beautiful Prompt (For EasyAnimate Inference)
|
||||
Beautiful Prompt aims to rewrite and beautify the user-uploaded prompt via LLMs, mapping it to the style of EasyAnimate's training captions,
|
||||
### Beautiful Prompt (For CogVideoX-Fun Inference)
|
||||
Beautiful Prompt aims to rewrite and beautify the user-uploaded prompt via LLMs, mapping it to the style of CogVideoX-Fun's training captions,
|
||||
making it more suitable as the inference prompt and thus improving the quality of the generated videos.
|
||||
We support batched inference with local LLMs or OpenAI compatible server based on [vLLM](https://github.com/vllm-project/vllm) for beautiful prompt.
|
||||
|
||||
#### Batched Inference
|
||||
1. Prepare original prompts in a jsonl file `easyanimate/video_caption/datasets/original_prompt.jsonl` with the following format:
|
||||
1. Prepare original prompts in a jsonl file `cogvideox/video_caption/datasets/original_prompt.jsonl` with the following format:
|
||||
```json
|
||||
{"prompt": "A stylish woman in a black leather jacket, red dress, and boots walks confidently down a damp Tokyo street."}
|
||||
{"prompt": "An underwater world with realistic fish and other creatures of the sea."}
|
||||
@@ -140,13 +139,10 @@ We support batched inference with local LLMs or OpenAI compatible server based o
|
||||
python caption_rewrite.py \
|
||||
--video_metadata_path datasets/original_prompt.jsonl \
|
||||
--caption_column "prompt" \
|
||||
--beautiful_prompt_column "beautiful_prompt" \
|
||||
--batch_size 1 \
|
||||
--model_name /path/to/your_llm \
|
||||
--prompt prompt/beautiful_prompt.txt \
|
||||
--prefix '"detailed description": ' \
|
||||
--answer_template "your detailed description here" \
|
||||
--max_retry_count 10 \
|
||||
--saved_path datasets/beautiful_prompt.jsonl \
|
||||
--saved_freq 1
|
||||
```
|
||||
@@ -1,7 +1,7 @@
|
||||
# 数据预处理
|
||||
[English](./README.md) | 简体中文
|
||||
|
||||
该文件夹包含 EasyAnimate 使用的数据集预处理(即视频切分、过滤和生成描述)和提示词美化的代码。整个过程支持分布式并行处理,能够处理大规模数据集。
|
||||
该文件夹包含 CogVideoX-Fun 使用的数据集预处理(即视频切分、过滤和生成描述)和提示词美化的代码。整个过程支持分布式并行处理,能够处理大规模数据集。
|
||||
|
||||
此外,我们和 [Data-Juicer](https://github.com/modelscope/data-juicer/blob/main/docs/DJ_SORA.md) 合作,能让你在 [Aliyun PAI-DLC](https://help.aliyun.com/zh/pai/user-guide/video-preprocessing/) 轻松进行视频数据的处理。
|
||||
|
||||
@@ -22,20 +22,22 @@
|
||||
|
||||
## 快速开始
|
||||
### 安装
|
||||
推荐使用阿里云 DSW 和 Docker 来安装环境,请参考 [快速开始](../../README_zh-CN.md#quick-start). 你也可以参考 [Dockerfile](../../Dockerfile.ds) 中的镜像构建流程在本地安装对应的 conda 环境和其余依赖。
|
||||
推荐使用阿里云 DSW 和 Docker 来安装环境,请参考 [快速开始](../../README_zh-CN.md#1-云使用-aliyundswdocker). 你也可以参考 [Dockerfile](../../Dockerfile.ds) 中的镜像构建流程在本地安装对应的 conda 环境和其余依赖。
|
||||
|
||||
为了提高推理速度和节省推理的显存,生成视频描述依赖于 [llm-awq](https://github.com/mit-han-lab/llm-awq)。因此,需要 RTX 3060 或者 A2 及以上的显卡 (CUDA Compute Capability >= 8.0)。
|
||||
|
||||
```shell
|
||||
# pull image
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:easyanimate
|
||||
docker pull mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# enter image
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:easyanimate
|
||||
docker run -it -p 7860:7860 --network host --gpus all --security-opt seccomp:unconfined --shm-size 200g mybigpai-public-registry.cn-beijing.cr.aliyuncs.com/easycv/torch_cuda:cogvideox_fun
|
||||
|
||||
# clone code
|
||||
git clone https://github.com/aigc-apps/EasyAnimate.git
|
||||
git clone https://github.com/aigc-apps/CogVideoX-Fun.git
|
||||
|
||||
# enter video_caption
|
||||
cd EasyAnimate/easyanimate/video_caption
|
||||
cd CogVideoX-Fun/cogvideox/video_caption
|
||||
```
|
||||
|
||||
### 数据集预处理
|
||||
@@ -51,7 +53,7 @@ cd EasyAnimate/easyanimate/video_caption
|
||||
```
|
||||
|
||||
#### 视频切分
|
||||
EasyAnimate 使用 [PySceneDetect](https://github.com/Breakthrough/PySceneDetect) 来识别视频中的场景变化
|
||||
CogVideoX-Fun 使用 [PySceneDetect](https://github.com/Breakthrough/PySceneDetect) 来识别视频中的场景变化
|
||||
并根据某些阈值通过 FFmpeg 执行视频分割,以确保视频片段的一致性。
|
||||
短于 3 秒的视频片段将被丢弃,长于 10 秒的视频片段将被递归切分。
|
||||
|
||||
@@ -59,12 +61,11 @@ EasyAnimate 使用 [PySceneDetect](https://github.com/Breakthrough/PySceneDetect
|
||||
```shell
|
||||
sh scripts/stage_1_video_splitting.sh
|
||||
```
|
||||
后,切分后的视频位于 `easyanimate/video_caption/datasets/panda_70m/videos_clips/data/`。
|
||||
后,切分后的视频位于 `cogvideox/video_caption/datasets/panda_70m/videos_clips/data/`。
|
||||
|
||||
#### 视频过滤
|
||||
基于上一步获得的视频,EasyAnimate 提供了一个简单而有效的流程来过滤出高质量的视频。总体流程如下:
|
||||
基于上一步获得的视频,CogVideoX-Fun 提供了一个简单而有效的流程来过滤出高质量的视频。总体流程如下:
|
||||
|
||||
- 场景跳变过滤:通过 [CLIP](https://github.com/openai/CLIP) 或者 [DINOv2](https://github.com/facebookresearch/dinov2) 来计算关键帧和首尾帧的语义相似度,从而过滤掉由于 PySceneDetect 缺失或多余分割引入的场景跳变的视频。
|
||||
- 美学过滤:通过 [aesthetic-predictor-v2-5](https://github.com/discus0434/aesthetic-predictor-v2-5) 计算均匀采样的 4 帧视频的平均美学分数,从而筛选出内容不佳(模糊、昏暗等)的视频。
|
||||
- 文本过滤:使用 [EasyOCR](https://github.com/JaidedAI/EasyOCR) 计算中间帧的文本区域比例,过滤掉含有大面积文本的视频。
|
||||
- 运动过滤:计算帧间光流差,过滤掉移动太慢或太快的视频。
|
||||
@@ -73,19 +74,19 @@ sh scripts/stage_1_video_splitting.sh
|
||||
```shell
|
||||
sh scripts/stage_2_video_filtering.sh
|
||||
```
|
||||
后,视频的美学得分、文本得分和运动得分对应的元文件保存在 `easyanimate/video_caption/datasets/panda_70m/videos_clips/`。
|
||||
后,视频的美学得分、文本得分和运动得分对应的元文件保存在 `cogvideox/video_caption/datasets/panda_70m/videos_clips/`。
|
||||
|
||||
> [!NOTE]
|
||||
> 美学得分的计算依赖于 [google/siglip-so400m-patch14-384 model](https://huggingface.co/google/siglip-so400m-patch14-384).
|
||||
请执行 `HF_ENDPOINT=https://hf-mirror.com sh scripts/stage_2_video_filtering.sh` 如果你无法访问 huggingface.com.
|
||||
|
||||
#### 视频描述
|
||||
在获得上述高质量的过滤视频后,EasyAnimate 利用 [InternVL2](https://internvl.readthedocs.io/en/latest/internvl2.0/introduction.html) 来生成视频描述。随后,使用 LLMs 对生成的视频描述进行重写,以更好地满足视频生成任务的要求。最后,使用自研的 [VideoCLIP-XL](https://arxiv.org/abs/2410.00741) 模型来过滤掉描述和视频内容不一致的数据,从而得到最终的训练数据集。
|
||||
在获得上述高质量的过滤视频后,CogVideoX-Fun 利用 [VILA1.5](https://github.com/NVlabs/VILA) 来生成视频描述。随后,使用 LLMs 对生成的视频描述进行重写,以更好地满足视频生成任务的要求。最后,使用自研的 VideoCLIPXL 模型来过滤掉描述和视频内容不一致的数据,从而得到最终的训练数据集。
|
||||
|
||||
请根据机器的显存从 [InternVL2](https://huggingface.co/collections/OpenGVLab/internvl-20-667d3961ab5eb12c7ed1463e) 下载合适大小的模型。对于 A100 40G,你可以执行下面的命令来下载 [InternVL2-40B-AWQ](https://huggingface.co/OpenGVLab/InternVL2-40B-AWQ)
|
||||
请根据机器的显存从 [VILA1.5](https://huggingface.co/collections/Efficient-Large-Model/vila-on-pre-training-for-visual-language-models-65d8022a3a52cd9bcd62698e) 下载合适大小的模型。对于 A100 40G,你可以执行下面的命令来下载 [VILA1.5-40b-AWQ](https://huggingface.co/Efficient-Large-Model/VILA1.5-40b-AWQ)
|
||||
```shell
|
||||
# Add HF_ENDPOINT=https://hf-mirror.com before the command if you cannot access to huggingface.com
|
||||
huggingface-cli download OpenGVLab/InternVL2-40B-AWQ --local-dir-use-symlinks False --local-dir /PATH/TO/INTERNVL2_MODEL
|
||||
huggingface-cli download Efficient-Large-Model/VILA1.5-40b-AWQ --local-dir-use-symlinks False --local-dir /PATH/TO/VILA_MODEL
|
||||
```
|
||||
|
||||
你可以选择性地准备 LLMs 来改写上述视频描述的结果。例如,你执行下面的命令来下载 [Meta-Llama-3-8B-Instruct](https://huggingface.co/NousResearch/Meta-Llama-3-8B-Instruct)
|
||||
@@ -97,18 +98,18 @@ huggingface-cli download NousResearch/Meta-Llama-3-8B-Instruct --local-dir-use-s
|
||||
视频描述的完整流程在 [stage_3_video_recaptioning.sh](./scripts/stage_3_video_recaptioning.sh).
|
||||
执行
|
||||
```shell
|
||||
CAPTION_MODEL_PATH=/PATH/TO/INTERNVL2_MODEL REWRITE_MODEL_PATH=/PATH/TO/REWRITE_MODEL sh scripts/stage_3_video_recaptioning.sh
|
||||
VILA_MODEL_PATH=/PATH/TO/VILA_MODEL REWRITE_MODEL_PATH=/PATH/TO/REWRITE_MODEL sh scripts/stage_3_video_recaptioning.sh
|
||||
```
|
||||
后,最后的训练文件会保存在 `easyanimate/video_caption/datasets/panda_70m/videos_clips/meta_train_info.json`。
|
||||
后,最后的训练文件会保存在 `cogvideox/video_caption/datasets/panda_70m/videos_clips/meta_train_info.json`。
|
||||
|
||||
### 提示词美化
|
||||
提示词美化旨在通过 LLMs 重写和美化用户上传的提示,将其映射为 EasyAnimate 训练所使用的视频描述风格、
|
||||
提示词美化旨在通过 LLMs 重写和美化用户上传的提示,将其映射为 CogVideoX-Fun 训练所使用的视频描述风格、
|
||||
使其更适合用作推理提示词,从而提高生成视频的质量。
|
||||
|
||||
基于 [vLLM](https://github.com/vllm-project/vllm),我们支持使用本地 LLM 进行批量推理或请求 OpenAI 服务器的方式,以进行提示词美化。
|
||||
|
||||
#### 批量推理
|
||||
1. 将原始的提示词以下面的格式准备在文件 `easyanimate/video_caption/datasets/original_prompt.jsonl` 中:
|
||||
1. 将原始的提示词以下面的格式准备在文件 `cogvideox/video_caption/datasets/original_prompt.jsonl` 中:
|
||||
```json
|
||||
{"prompt": "A stylish woman in a black leather jacket, red dress, and boots walks confidently down a damp Tokyo street."}
|
||||
{"prompt": "An underwater world with realistic fish and other creatures of the sea."}
|
||||
@@ -125,13 +126,10 @@ CAPTION_MODEL_PATH=/PATH/TO/INTERNVL2_MODEL REWRITE_MODEL_PATH=/PATH/TO/REWRITE_
|
||||
python caption_rewrite.py \
|
||||
--video_metadata_path datasets/original_prompt.jsonl \
|
||||
--caption_column "prompt" \
|
||||
--beautiful_prompt_column "beautiful_prompt" \
|
||||
--batch_size 1 \
|
||||
--model_name /path/to/your_llm \
|
||||
--prompt prompt/beautiful_prompt.txt \
|
||||
--prefix '"detailed description": ' \
|
||||
--answer_template "your detailed description here" \
|
||||
--max_retry_count 10 \
|
||||
--saved_path datasets/beautiful_prompt.jsonl \
|
||||
--saved_freq 1
|
||||
```
|
||||
@@ -1,5 +1,5 @@
|
||||
"""
|
||||
This script (optional) can rewrite and beautify the user-uploaded prompt via LLMs, mapping it to the style of EasyAnimate's training captions,
|
||||
This script (optional) can rewrite and beautify the user-uploaded prompt via LLMs, mapping it to the style of cogvideox's training captions,
|
||||
making it more suitable as the inference prompt and thus improving the quality of the generated videos.
|
||||
|
||||
Usage:
|
||||
@@ -32,7 +32,7 @@ import os
|
||||
|
||||
from openai import OpenAI
|
||||
|
||||
from easyanimate.video_caption.caption_rewrite import extract_output
|
||||
from cogvideox.video_caption.caption_rewrite import extract_output
|
||||
|
||||
|
||||
def parse_args():
|
||||
@@ -42,7 +42,7 @@ def parse_args():
|
||||
parser.add_argument(
|
||||
"--template",
|
||||
type=str,
|
||||
default="easyanimate/video_caption/prompt/beautiful_prompt.txt",
|
||||
default="cogvideox/video_caption/prompt/beautiful_prompt.txt",
|
||||
help="A string or a txt file contains the template for beautiful prompt."
|
||||
)
|
||||
parser.add_argument(
|
||||
@@ -1,14 +1,13 @@
|
||||
import argparse
|
||||
import os
|
||||
import re
|
||||
from copy import deepcopy
|
||||
import os
|
||||
from tqdm import tqdm
|
||||
|
||||
import pandas as pd
|
||||
import torch
|
||||
from natsort import index_natsorted
|
||||
from tqdm import tqdm
|
||||
from transformers import AutoTokenizer
|
||||
from vllm import LLM, SamplingParams
|
||||
from transformers import AutoTokenizer
|
||||
|
||||
from utils.logger import logger
|
||||
|
||||
@@ -33,26 +32,6 @@ def extract_output(s, prefix='"rewritten description": '):
|
||||
logger.warning(f"{output} does not start with {prefix}. Return None.")
|
||||
return None
|
||||
|
||||
"""The file unifies the following two tasks:
|
||||
1. Caption Rewrite: rewrite the video recaption results by LLMs.
|
||||
2. Beautiful Prompt: rewrite and beautify the user-uploaded prompt via LLMs.
|
||||
|
||||
For the caption rewrite task, the input video_metadata_path should have the following format:
|
||||
```jsonl
|
||||
{"video_path_column": "1.mp4", "caption_column": "a man is running in the street."}
|
||||
...
|
||||
{"video_path_column": "100.mp4", "caption_column": "a dog is chasing a cat."}
|
||||
```
|
||||
The video_path_column in the argparse must be specified.
|
||||
|
||||
For the beautiful prompt task, the input video_metadata_path should have the following format:
|
||||
```jsonl
|
||||
{"caption_column": "a man is running in the street."}
|
||||
...
|
||||
{"caption_column": "a dog is chasing a cat."}
|
||||
```
|
||||
The beautiful_prompt_column in the argparse must be specified for the saving purpose.
|
||||
"""
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="Rewrite the video caption by LLMs.")
|
||||
@@ -63,10 +42,7 @@ def parse_args():
|
||||
"--video_path_column",
|
||||
type=str,
|
||||
default=None,
|
||||
help=(
|
||||
"The column contains the video path (an absolute path or a relative path w.r.t the video_folder)."
|
||||
"It is conflicted with the beautiful_prompt_column."
|
||||
),
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--caption_column",
|
||||
@@ -74,12 +50,6 @@ def parse_args():
|
||||
default="caption",
|
||||
help="The column contains the video caption.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--beautiful_prompt_column",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The column name for the beautiful prompt column. It is conflicted with the video_path_column.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--batch_size",
|
||||
type=int,
|
||||
@@ -104,18 +74,6 @@ def parse_args():
|
||||
required=True,
|
||||
help="The prefix to extract the output from LLMs.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--answer_template",
|
||||
type=str,
|
||||
default="",
|
||||
help="The anwer template in the prompt. If specified, rewritten results same as the answer template will be removed.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--max_retry_count",
|
||||
type=int,
|
||||
default=1,
|
||||
help="The maximum retry count to ensure outputs with the valid format from LLMs.",
|
||||
)
|
||||
parser.add_argument("--saved_path", type=str, required=True, help="The save path to the output results (csv/jsonl).")
|
||||
parser.add_argument("--saved_freq", type=int, default=1, help="The frequency to save the output results.")
|
||||
|
||||
@@ -138,34 +96,20 @@ def main():
|
||||
saved_suffix = os.path.splitext(args.saved_path)[1]
|
||||
if saved_suffix not in set([".csv", ".jsonl", ".json"]):
|
||||
raise ValueError(f"The saved_path must end with .csv, .jsonl or .json.")
|
||||
|
||||
if args.video_path_column is None and args.beautiful_prompt_column is None:
|
||||
raise ValueError("Either video_path_column or beautiful_prompt_column should be specified in the arguments.")
|
||||
if args.video_path_column is not None and args.beautiful_prompt_column is not None:
|
||||
raise ValueError(
|
||||
"Both video_path_column and beautiful_prompt_column can not be specified in the arguments at the same time."
|
||||
)
|
||||
|
||||
if os.path.exists(args.saved_path):
|
||||
if os.path.exists(args.saved_path) and args.video_path_column is not None:
|
||||
if args.saved_path.endswith(".csv"):
|
||||
saved_metadata_df = pd.read_csv(args.saved_path)
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
saved_metadata_df = pd.read_json(args.saved_path, lines=True)
|
||||
|
||||
if args.video_path_column is not None:
|
||||
# Filter out the unprocessed video-caption pairs by setting the indicator=True.
|
||||
merged_df = video_metadata_df.merge(saved_metadata_df, on=args.video_path_column, how="outer", indicator=True)
|
||||
video_metadata_df = merged_df[merged_df["_merge"] == "left_only"]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_metadata_df = video_metadata_df.iloc[index_natsorted(video_metadata_df[args.video_path_column])]
|
||||
video_metadata_df = video_metadata_df.reset_index(drop=True)
|
||||
if args.beautiful_prompt_column is not None:
|
||||
# Filter out the unprocessed caption-beautifil_prompt pairs by setting the indicator=True.
|
||||
merged_df = video_metadata_df.merge(saved_metadata_df, on=args.caption_column, how="outer", indicator=True)
|
||||
video_metadata_df = merged_df[merged_df["_merge"] == "left_only"]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_metadata_df = video_metadata_df.iloc[index_natsorted(video_metadata_df[args.caption_column])]
|
||||
video_metadata_df = video_metadata_df.reset_index(drop=True)
|
||||
# Filter out the unprocessed video-caption pairs by setting the indicator=True.
|
||||
merged_df = video_metadata_df.merge(saved_metadata_df, on=args.video_path_column, how="outer", indicator=True)
|
||||
video_metadata_df = merged_df[merged_df["_merge"] == "left_only"]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_metadata_df = video_metadata_df.iloc[index_natsorted(video_metadata_df[args.video_path_column])].reset_index(
|
||||
drop=True
|
||||
)
|
||||
logger.info(
|
||||
f"Resume from {args.saved_path}: {len(saved_metadata_df)} processed and {len(video_metadata_df)} to be processed."
|
||||
)
|
||||
@@ -175,9 +119,6 @@ def main():
|
||||
args.prompt = "".join(f.readlines())
|
||||
logger.info(f"Prompt: {args.prompt}")
|
||||
|
||||
if args.max_retry_count < 1:
|
||||
raise ValueError(f"The max_retry_count {args.max_retry_count} must be greater than 0.")
|
||||
|
||||
if args.video_path_column is not None:
|
||||
video_path_list = video_metadata_df[args.video_path_column].tolist()
|
||||
if args.caption_column in video_metadata_df.columns:
|
||||
@@ -204,10 +145,9 @@ def main():
|
||||
tokenizer = AutoTokenizer.from_pretrained(args.model_name)
|
||||
sampling_params = SamplingParams(temperature=0.7, top_p=1, max_tokens=1024)
|
||||
|
||||
result_dict = {args.caption_column: []}
|
||||
if args.video_path_column is not None:
|
||||
result_dict = {args.video_path_column: [], args.caption_column: []}
|
||||
if args.beautiful_prompt_column is not None:
|
||||
result_dict = {args.caption_column: [], args.beautiful_prompt_column: []}
|
||||
|
||||
for i in tqdm(range(0, len(sampled_frame_caption_list), args.batch_size)):
|
||||
if args.video_path_column is not None:
|
||||
@@ -222,77 +162,63 @@ def main():
|
||||
]
|
||||
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
||||
batch_prompt.append(text)
|
||||
|
||||
cur_retry_count = 0
|
||||
while cur_retry_count < args.max_retry_count:
|
||||
if len(batch_prompt) == 0:
|
||||
break
|
||||
|
||||
batch_result = []
|
||||
batch_output = llm.generate(batch_prompt, sampling_params)
|
||||
batch_output = [output.outputs[0].text.rstrip() for output in batch_output]
|
||||
if args.prefix is not None:
|
||||
batch_output = [extract_output(output, args.prefix) for output in batch_output]
|
||||
batch_output = llm.generate(batch_prompt, sampling_params)
|
||||
batch_output = [output.outputs[0].text.rstrip() for output in batch_output]
|
||||
batch_output = [extract_output(output, prefix=args.prefix) for output in batch_output]
|
||||
|
||||
if args.video_path_column is not None:
|
||||
retry_batch_video_path, retry_batch_prompt = [], []
|
||||
for (video_path, prompt, output) in zip(batch_video_path, batch_prompt, batch_output):
|
||||
# Filter out data that does not meet the output format to retry.
|
||||
if output is not None and output != args.answer_template:
|
||||
batch_result.append((video_path, output))
|
||||
else:
|
||||
retry_batch_video_path.append(video_path)
|
||||
retry_batch_prompt.append(prompt)
|
||||
if len(batch_result) != 0:
|
||||
batch_video_path, batch_output = zip(*batch_result)
|
||||
result_dict[args.video_path_column].extend(deepcopy(batch_video_path))
|
||||
result_dict[args.caption_column].extend(deepcopy(batch_output))
|
||||
|
||||
batch_video_path, batch_prompt = retry_batch_video_path, retry_batch_prompt
|
||||
if args.beautiful_prompt_column is not None:
|
||||
retry_batch_caption, retry_batch_prompt = [], []
|
||||
for (caption, prompt, output) in zip(batch_caption, batch_prompt, batch_output):
|
||||
# Filter out data that does not meet the output format to retry.
|
||||
if output is not None and output != args.answer_template:
|
||||
batch_result.append((caption, output))
|
||||
else:
|
||||
retry_batch_caption.append(caption)
|
||||
retry_batch_prompt.append(prompt)
|
||||
if len(batch_result) != 0:
|
||||
batch_caption, batch_output = zip(*batch_result)
|
||||
result_dict[args.caption_column].extend(deepcopy(batch_caption))
|
||||
result_dict[args.beautiful_prompt_column].extend(deepcopy(batch_output))
|
||||
|
||||
batch_caption, batch_prompt = retry_batch_caption, retry_batch_prompt
|
||||
|
||||
cur_retry_count += 1
|
||||
logger.info(
|
||||
f"Current retry count/Maximum retry count: {cur_retry_count}/{args.max_retry_count}.: "
|
||||
f"Retrying {len(batch_prompt)} prompts with invalid output format."
|
||||
)
|
||||
# Filter out data that does not meet the output format.
|
||||
batch_result = []
|
||||
if args.video_path_column is not None:
|
||||
for video_path, output in zip(batch_video_path, batch_output):
|
||||
if output is not None:
|
||||
batch_result.append((video_path, output))
|
||||
batch_video_path, batch_output = zip(*batch_result)
|
||||
|
||||
result_dict[args.video_path_column].extend(batch_video_path)
|
||||
else:
|
||||
for output in batch_output:
|
||||
if output is not None:
|
||||
batch_result.append(output)
|
||||
|
||||
result_dict[args.caption_column].extend(batch_result)
|
||||
|
||||
# Save the metadata every args.saved_freq.
|
||||
if (i // args.batch_size) % args.saved_freq == 0 or (i + 1) * args.batch_size >= len(sampled_frame_caption_list):
|
||||
if i != 0 and ((i // args.batch_size) % args.saved_freq) == 0:
|
||||
if len(result_dict[args.caption_column]) > 0:
|
||||
result_df = pd.DataFrame(result_dict)
|
||||
# Append is not supported (oss).
|
||||
if args.saved_path.endswith(".csv"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_csv(args.saved_path)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_csv(args.saved_path, index=False)
|
||||
header = True if not os.path.exists(args.saved_path) else False
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
elif args.saved_path.endswith(".json"):
|
||||
# Append is not supported.
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_json(args.saved_path, orient="records", lines=True)
|
||||
saved_df = pd.read_json(args.saved_path, orient="records")
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, force_ascii=False)
|
||||
result_df.to_json(args.saved_path, orient="records", indent=4, force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
|
||||
result_dict = {args.caption_column: []}
|
||||
if args.video_path_column is not None:
|
||||
result_dict = {args.video_path_column: [], args.caption_column: []}
|
||||
if args.beautiful_prompt_column is not None:
|
||||
result_dict = {args.caption_column: [], args.beautiful_prompt_column: []}
|
||||
|
||||
if len(result_dict[args.caption_column]) > 0:
|
||||
result_df = pd.DataFrame(result_dict)
|
||||
if args.saved_path.endswith(".csv"):
|
||||
header = True if not os.path.exists(args.saved_path) else False
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a")
|
||||
elif args.saved_path.endswith(".json"):
|
||||
# Append is not supported.
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_json(args.saved_path, orient="records")
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_json(args.saved_path, orient="records", indent=4, force_ascii=False)
|
||||
logger.info(f"Save the final result to {args.saved_path}.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,7 +1,9 @@
|
||||
import ast
|
||||
import argparse
|
||||
import gc
|
||||
import os
|
||||
from contextlib import contextmanager
|
||||
from pathlib import Path
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
@@ -10,8 +12,8 @@ from joblib import Parallel, delayed
|
||||
from natsort import natsorted
|
||||
from tqdm import tqdm
|
||||
|
||||
from utils.filter import filter
|
||||
from utils.logger import logger
|
||||
from utils.filter import filter
|
||||
|
||||
|
||||
@contextmanager
|
||||
@@ -74,11 +76,11 @@ def compute_motion_score(video_path):
|
||||
video_motion_scores.append(frame_motion_score)
|
||||
prev_frame = gray_frame
|
||||
|
||||
motion_score_result = {
|
||||
"video_path": video_path,
|
||||
video_meta_info = {
|
||||
"video_path": Path(video_path).name,
|
||||
"motion_score": round(float(np.mean(video_motion_scores)), 5),
|
||||
}
|
||||
return motion_score_result
|
||||
return video_meta_info
|
||||
|
||||
except Exception as e:
|
||||
print(f"Compute motion score for video {video_path} with error: {e}.")
|
||||
@@ -97,35 +99,27 @@ def parse_args():
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument("--saved_path", type=str, required=True, help="The save path to the output results (csv/jsonl).")
|
||||
parser.add_argument("--saved_freq", type=int, default=1, help="The frequency to save the output results.")
|
||||
parser.add_argument("--saved_freq", type=int, default=100, help="The frequency to save the output results.")
|
||||
parser.add_argument("--n_jobs", type=int, default=1, help="The number of concurrent processes.")
|
||||
|
||||
parser.add_argument("--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl).")
|
||||
parser.add_argument(
|
||||
"--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_resolution", type=float, default=0, help="The resolution threshold.")
|
||||
parser.add_argument("--min_duration", type=float, default=-1, help="The minimum duration.")
|
||||
parser.add_argument("--max_duration", type=float, default=-1, help="The maximum duration.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
"--asethetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score", type=float, default=4.0, help="The aesthetic score threshold.")
|
||||
parser.add_argument("--min_asethetic_score", type=float, default=4.0, help="The asethetic score threshold.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
"--asethetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score_siglip", type=float, default=4.0, help="The aesthetic score (SigLIP) threshold.")
|
||||
parser.add_argument("--min_asethetic_score_siglip", type=float, default=4.0, help="The asethetic score (SigLIP) threshold.")
|
||||
parser.add_argument(
|
||||
"--text_score_metadata_path", type=str, default=None, help="The path to the video text score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_text_score", type=float, default=0.02, help="The text threshold.")
|
||||
parser.add_argument(
|
||||
"--semantic_consistency_score_metadata_path",
|
||||
nargs="+",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The path to the semantic consistency metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--min_semantic_consistency_score", type=float, default=0.80, help="The semantic consistency score threshold."
|
||||
)
|
||||
|
||||
args = parser.parse_args()
|
||||
return args
|
||||
@@ -160,49 +154,33 @@ def main():
|
||||
min_resolution=args.min_resolution,
|
||||
min_duration=args.min_duration,
|
||||
max_duration=args.max_duration,
|
||||
aesthetic_score_metadata_path=args.aesthetic_score_metadata_path,
|
||||
min_aesthetic_score=args.min_aesthetic_score,
|
||||
aesthetic_score_siglip_metadata_path=args.aesthetic_score_siglip_metadata_path,
|
||||
min_aesthetic_score_siglip=args.min_aesthetic_score_siglip,
|
||||
asethetic_score_metadata_path=args.asethetic_score_metadata_path,
|
||||
min_asethetic_score=args.min_asethetic_score,
|
||||
asethetic_score_siglip_metadata_path=args.asethetic_score_siglip_metadata_path,
|
||||
min_asethetic_score_siglip=args.min_asethetic_score_siglip,
|
||||
text_score_metadata_path=args.text_score_metadata_path,
|
||||
min_text_score=args.min_text_score,
|
||||
semantic_consistency_score_metadata_path=args.semantic_consistency_score_metadata_path,
|
||||
min_semantic_consistency_score=args.min_semantic_consistency_score,
|
||||
video_path_column=args.video_path_column
|
||||
)
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_path_list = natsorted(video_path_list)
|
||||
logger.info(f"{len(video_path_list)} videos are to be processed.")
|
||||
|
||||
for i in tqdm(range(0, len(video_path_list), args.saved_freq)):
|
||||
# Get motion score result for each video asynchronously.
|
||||
motion_score_result_list = Parallel(n_jobs=args.n_jobs)(
|
||||
result_list = Parallel(n_jobs=args.n_jobs)(
|
||||
delayed(compute_motion_score)(video_path) for video_path in tqdm(video_path_list[i: i + args.saved_freq])
|
||||
)
|
||||
result_list = []
|
||||
for motion_score_result in motion_score_result_list:
|
||||
if motion_score_result is not None:
|
||||
video_path = motion_score_result["video_path"]
|
||||
if args.video_folder != "":
|
||||
video_path = os.path.relpath(video_path, args.video_folder)
|
||||
result_list.append({args.video_path_column: video_path, "motion_score": motion_score_result["motion_score"]})
|
||||
result_list = [result for result in result_list if result is not None]
|
||||
if len(result_list) == 0:
|
||||
continue
|
||||
|
||||
result_df = pd.DataFrame(result_list)
|
||||
# Append is not supported (oss).
|
||||
if args.saved_path.endswith(".csv"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_csv(args.saved_path)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_csv(args.saved_path, index=False)
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_json(args.saved_path, orient="records", lines=True)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, force_ascii=False)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,5 +1,6 @@
|
||||
import argparse
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import easyocr
|
||||
import numpy as np
|
||||
@@ -10,9 +11,9 @@ from natsort import natsorted
|
||||
from tqdm import tqdm
|
||||
from torchvision.datasets.utils import download_url
|
||||
|
||||
from utils.filter import filter
|
||||
from utils.logger import logger
|
||||
from utils.video_utils import extract_frames
|
||||
from utils.filter import filter
|
||||
|
||||
|
||||
def init_ocr_reader(root: str = "~/.cache/easyocr", device: str = "gpu"):
|
||||
@@ -46,8 +47,8 @@ def triangle_area(p1, p2, p3):
|
||||
return tri_area
|
||||
|
||||
|
||||
def compute_text_score(video_path, ocr_reader, sample_method="mid", num_sampled_frames=1):
|
||||
_, images = extract_frames(video_path, sample_method=sample_method, num_sampled_frames=num_sampled_frames)
|
||||
def compute_text_score(video_path, ocr_reader):
|
||||
_, images = extract_frames(video_path, sample_method="mid")
|
||||
images = [np.array(image) for image in images]
|
||||
|
||||
frame_ocr_area_ratios = []
|
||||
@@ -71,15 +72,18 @@ def compute_text_score(video_path, ocr_reader, sample_method="mid", num_sampled_
|
||||
quad_area += triangle_area(*triangle1)
|
||||
triangle2 = points[3:] + [points[0]]
|
||||
quad_area += triangle_area(*triangle2)
|
||||
except Exception:
|
||||
except:
|
||||
quad_area = 0
|
||||
text_area = rect_area + quad_area
|
||||
|
||||
frame_ocr_area_ratios.append(text_area / total_area)
|
||||
|
||||
text_score = round(np.mean(frame_ocr_area_ratios), 5)
|
||||
video_meta_info = {
|
||||
"video_path": Path(video_path).name,
|
||||
"text_score": round(np.mean(frame_ocr_area_ratios), 5),
|
||||
}
|
||||
|
||||
return text_score
|
||||
return video_meta_info
|
||||
|
||||
|
||||
def parse_args():
|
||||
@@ -94,47 +98,27 @@ def parse_args():
|
||||
default="video_path",
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--frame_sample_method",
|
||||
type=str,
|
||||
default="mid",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--num_sampled_frames",
|
||||
type=int,
|
||||
default=1,
|
||||
help="num_sampled_frames",
|
||||
)
|
||||
parser.add_argument("--saved_path", type=str, required=True, help="The save path to the output results (csv/jsonl).")
|
||||
parser.add_argument("--saved_freq", type=int, default=1, help="The frequency to save the output results.")
|
||||
parser.add_argument("--saved_freq", type=int, default=100, help="The frequency to save the output results.")
|
||||
|
||||
parser.add_argument("--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl).")
|
||||
parser.add_argument(
|
||||
"--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_resolution", type=float, default=0, help="The resolution threshold.")
|
||||
parser.add_argument("--min_duration", type=float, default=-1, help="The minimum duration.")
|
||||
parser.add_argument("--max_duration", type=float, default=-1, help="The maximum duration.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
"--asethetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score", type=float, default=4.0, help="The aesthetic score threshold.")
|
||||
parser.add_argument("--min_asethetic_score", type=float, default=4.0, help="The asethetic score threshold.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
"--asethetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score_siglip", type=float, default=4.0, help="The aesthetic score (SigLIP) threshold.")
|
||||
parser.add_argument("--min_asethetic_score_siglip", type=float, default=4.0, help="The asethetic score (SigLIP) threshold.")
|
||||
parser.add_argument(
|
||||
"--motion_score_metadata_path", type=str, default=None, help="The path to the video motion score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_motion_score", type=float, default=2, help="The minimum motion threshold.")
|
||||
parser.add_argument("--max_motion_score", type=float, default=999999, help="The maximum motion threshold.")
|
||||
parser.add_argument(
|
||||
"--semantic_consistency_score_metadata_path",
|
||||
nargs="+",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The path to the semantic consistency metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--min_semantic_consistency_score", type=float, default=0.80, help="The semantic consistency score threshold."
|
||||
)
|
||||
parser.add_argument("--min_motion_score", type=float, default=2, help="The motion threshold.")
|
||||
|
||||
args = parser.parse_args()
|
||||
return args
|
||||
@@ -169,16 +153,12 @@ def main():
|
||||
min_resolution=args.min_resolution,
|
||||
min_duration=args.min_duration,
|
||||
max_duration=args.max_duration,
|
||||
aesthetic_score_metadata_path=args.aesthetic_score_metadata_path,
|
||||
min_aesthetic_score=args.min_aesthetic_score,
|
||||
aesthetic_score_siglip_metadata_path=args.aesthetic_score_siglip_metadata_path,
|
||||
min_aesthetic_score_siglip=args.min_aesthetic_score_siglip,
|
||||
asethetic_score_metadata_path=args.asethetic_score_metadata_path,
|
||||
min_asethetic_score=args.min_asethetic_score,
|
||||
asethetic_score_siglip_metadata_path=args.asethetic_score_siglip_metadata_path,
|
||||
min_asethetic_score_siglip=args.min_asethetic_score_siglip,
|
||||
motion_score_metadata_path=args.motion_score_metadata_path,
|
||||
min_motion_score=args.min_motion_score,
|
||||
max_motion_score=args.max_motion_score,
|
||||
semantic_consistency_score_metadata_path=args.semantic_consistency_score_metadata_path,
|
||||
min_semantic_consistency_score=args.min_semantic_consistency_score,
|
||||
video_path_column=args.video_path_column
|
||||
)
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
@@ -193,10 +173,7 @@ def main():
|
||||
|
||||
index = len(video_path_list) - len(video_path_list) % state.num_processes
|
||||
# Avoid the NCCL timeout in the final gather operation.
|
||||
logger.info(
|
||||
f"Drop the last {len(video_path_list) % state.num_processes} videos to "
|
||||
"ensure each process handles the same number of videos."
|
||||
)
|
||||
logger.info(f"Drop {len(video_path_list) % state.num_processes} videos to ensure each process handles the same number of videos.")
|
||||
video_path_list = video_path_list[:index]
|
||||
logger.info(f"{len(video_path_list)} videos are to be processed.")
|
||||
|
||||
@@ -204,39 +181,34 @@ def main():
|
||||
with state.split_between_processes(video_path_list) as splitted_video_path_list:
|
||||
for i, video_path in enumerate(tqdm(splitted_video_path_list)):
|
||||
try:
|
||||
text_score = compute_text_score(
|
||||
video_path,
|
||||
ocr_reader,
|
||||
sample_method=args.frame_sample_method,
|
||||
num_sampled_frames=args.num_sampled_frames,
|
||||
)
|
||||
video_meta_info = {}
|
||||
if args.video_folder == "":
|
||||
video_meta_info[args.video_path_column] = video_path
|
||||
else:
|
||||
video_meta_info[args.video_path_column] = os.path.relpath(video_path, args.video_folder)
|
||||
video_meta_info["text_score"] = text_score
|
||||
video_meta_info = compute_text_score(video_path, ocr_reader)
|
||||
result_list.append(video_meta_info)
|
||||
except Exception as e:
|
||||
logger.warning(f"Compute text score for video {video_path} with error: {e}.")
|
||||
if i % args.saved_freq == 0 or i == len(splitted_video_path_list) - 1:
|
||||
if i != 0 and i % args.saved_freq == 0:
|
||||
state.wait_for_everyone()
|
||||
gathered_result_list = gather_object(result_list)
|
||||
if state.is_main_process and len(gathered_result_list) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_list)
|
||||
# Append is not supported (oss).
|
||||
if args.saved_path.endswith(".csv"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_csv(args.saved_path)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_csv(args.saved_path, index=False)
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_json(args.saved_path, orient="records", lines=True)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, force_ascii=False)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
result_list = []
|
||||
|
||||
state.wait_for_everyone()
|
||||
gathered_result_list = gather_object(result_list)
|
||||
if state.is_main_process and len(gathered_result_list) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_list)
|
||||
if args.saved_path.endswith(".csv"):
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save the final result to {args.saved_path}.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -10,7 +10,6 @@ from torch.utils.data import DataLoader
|
||||
|
||||
import utils.image_evaluator as image_evaluator
|
||||
import utils.video_evaluator as video_evaluator
|
||||
from utils.filter import filter
|
||||
from utils.logger import logger
|
||||
from utils.video_dataset import VideoDataset, collate_fn
|
||||
|
||||
@@ -27,43 +26,41 @@ def parse_args():
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument("--video_folder", type=str, default="", help="The video folder.")
|
||||
parser.add_argument("--caption_column", type=str, default=None, help="The column contains the caption.")
|
||||
parser.add_argument(
|
||||
"--caption_column",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The column contains the caption.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--frame_sample_method",
|
||||
type=str,
|
||||
choices=["mid", "uniform", "image"],
|
||||
default="uniform",
|
||||
)
|
||||
parser.add_argument("--num_sampled_frames", type=int, default=8, help="The number of sampled frames.")
|
||||
parser.add_argument(
|
||||
"--num_sampled_frames",
|
||||
type=int,
|
||||
default=8,
|
||||
help="num_sampled_frames",
|
||||
)
|
||||
parser.add_argument("--metrics", nargs="+", type=str, required=True, help="The evaluation metric(s) for generated images.")
|
||||
parser.add_argument("--batch_size", type=int, default=1, help="The batch size for the video dataset.")
|
||||
parser.add_argument("--num_workers", type=int, default=1, help="The number of workers for the video dataset.")
|
||||
parser.add_argument(
|
||||
"--batch_size",
|
||||
type=int,
|
||||
default=10,
|
||||
required=False,
|
||||
help="The batch size for the video dataset.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--num_workers",
|
||||
type=int,
|
||||
default=4,
|
||||
required=False,
|
||||
help="The number of workers for the video dataset.",
|
||||
)
|
||||
parser.add_argument("--saved_path", type=str, required=True, help="The save path to the output results (csv/jsonl).")
|
||||
parser.add_argument("--saved_freq", type=int, default=1, help="The frequency to save the output results.")
|
||||
|
||||
parser.add_argument("--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl).")
|
||||
parser.add_argument("--min_resolution", type=float, default=0, help="The resolution threshold.")
|
||||
parser.add_argument("--min_duration", type=float, default=-1, help="The minimum duration.")
|
||||
parser.add_argument("--max_duration", type=float, default=-1, help="The maximum duration.")
|
||||
parser.add_argument(
|
||||
"--text_score_metadata_path", type=str, default=None, help="The path to the video text score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_text_score", type=float, default=0.02, help="The text threshold.")
|
||||
parser.add_argument(
|
||||
"--motion_score_metadata_path", type=str, default=None, help="The path to the video motion score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_motion_score", type=float, default=2, help="The minimum motion threshold.")
|
||||
parser.add_argument("--max_motion_score", type=float, default=999999, help="The maximum motion threshold.")
|
||||
parser.add_argument(
|
||||
"--semantic_consistency_score_metadata_path",
|
||||
nargs="+",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The path to the semantic consistency metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--min_semantic_consistency_score", type=float, default=0.80, help="The semantic consistency score threshold."
|
||||
)
|
||||
parser.add_argument("--saved_freq", type=int, default=1000, help="The frequency to save the output results.")
|
||||
|
||||
args = parser.parse_args()
|
||||
return args
|
||||
@@ -89,34 +86,16 @@ def main():
|
||||
saved_metadata_df = pd.read_json(args.saved_path, lines=True)
|
||||
|
||||
# Filter out the unprocessed video-caption pairs by setting the indicator=True.
|
||||
merged_df = video_metadata_df.merge(saved_metadata_df, on=args.video_path_column, how="outer", indicator=True)
|
||||
merged_df = video_metadata_df.merge(saved_metadata_df, on="video_path", how="outer", indicator=True)
|
||||
video_metadata_df = merged_df[merged_df["_merge"] == "left_only"]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_metadata_df = video_metadata_df.iloc[index_natsorted(video_metadata_df[args.video_path_column])].reset_index(drop=True)
|
||||
video_metadata_df = video_metadata_df.iloc[index_natsorted(video_metadata_df["video_path"])].reset_index(drop=True)
|
||||
if args.caption_column is None:
|
||||
video_metadata_df = video_metadata_df[[args.video_path_column]]
|
||||
else:
|
||||
video_metadata_df = video_metadata_df[[args.video_path_column, args.caption_column + "_x"]]
|
||||
video_metadata_df.rename(columns={args.caption_column + "_x": args.caption_column}, inplace=True)
|
||||
logger.info(f"Resume from {args.saved_path}: {len(saved_metadata_df)} processed and {len(video_metadata_df)} to be processed.")
|
||||
|
||||
video_path_list = video_metadata_df[args.video_path_column].tolist()
|
||||
video_path_list = filter(
|
||||
video_path_list,
|
||||
basic_metadata_path=args.basic_metadata_path,
|
||||
min_resolution=args.min_resolution,
|
||||
min_duration=args.min_duration,
|
||||
max_duration=args.max_duration,
|
||||
text_score_metadata_path=args.text_score_metadata_path,
|
||||
min_text_score=args.min_text_score,
|
||||
motion_score_metadata_path=args.motion_score_metadata_path,
|
||||
min_motion_score=args.min_motion_score,
|
||||
max_motion_score=args.max_motion_score,
|
||||
semantic_consistency_score_metadata_path=args.semantic_consistency_score_metadata_path,
|
||||
min_semantic_consistency_score=args.min_semantic_consistency_score,
|
||||
video_path_column=args.video_path_column
|
||||
)
|
||||
video_metadata_df = video_metadata_df[video_metadata_df[args.video_path_column].isin(video_path_list)]
|
||||
|
||||
state = PartialState()
|
||||
metric_fns = []
|
||||
@@ -148,10 +127,7 @@ def main():
|
||||
|
||||
index = len(video_metadata_df) - len(video_metadata_df) % state.num_processes
|
||||
# Avoid the NCCL timeout in the final gather operation.
|
||||
logger.info(
|
||||
f"Drop the last {len(video_metadata_df) % state.num_processes} videos "
|
||||
"to ensure each process handles the same number of videos."
|
||||
)
|
||||
logger.info(f"Drop {len(video_metadata_df) % state.num_processes} videos to ensure each process handles the same number of videos.")
|
||||
video_metadata_df = video_metadata_df.iloc[:index]
|
||||
logger.info(f"{len(video_metadata_df)} videos are to be processed.")
|
||||
|
||||
@@ -160,7 +136,6 @@ def main():
|
||||
video_dataset = VideoDataset(
|
||||
dataset_inputs=splitted_video_metadata,
|
||||
video_folder=args.video_folder,
|
||||
video_path_column=args.video_path_column,
|
||||
text_column=args.caption_column,
|
||||
sample_method=args.frame_sample_method,
|
||||
num_sampled_frames=args.num_sampled_frames
|
||||
@@ -195,25 +170,32 @@ def main():
|
||||
result_dict[args.video_path_column].extend(saved_video_path_list)
|
||||
|
||||
# Save the metadata in the main process every saved_freq.
|
||||
if (idx % args.saved_freq) == 0 or idx == len(video_loader) - 1:
|
||||
if (idx != 0) and (idx % args.saved_freq == 0):
|
||||
state.wait_for_everyone()
|
||||
gathered_result_dict = {k: gather_object(v) for k, v in result_dict.items()}
|
||||
if state.is_main_process and len(gathered_result_dict[args.video_path_column]) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_dict)
|
||||
# Append is not supported (oss).
|
||||
if args.saved_path.endswith(".csv"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_csv(args.saved_path)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_csv(args.saved_path, index=False)
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
if os.path.exists(args.saved_path):
|
||||
saved_df = pd.read_json(args.saved_path, orient="records", lines=True)
|
||||
result_df = pd.concat([saved_df, result_df], ignore_index=True)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, force_ascii=False)
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
for k in result_dict.keys():
|
||||
result_dict[k] = []
|
||||
|
||||
# Wait for all processes to finish and gather the final result.
|
||||
state.wait_for_everyone()
|
||||
gathered_result_dict = {k: gather_object(v) for k, v in result_dict.items()}
|
||||
# Save the metadata in the main process.
|
||||
if state.is_main_process and len(gathered_result_dict[args.video_path_column]) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_dict)
|
||||
if args.saved_path.endswith(".csv"):
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save the final result to {args.saved_path}.")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,13 +1,14 @@
|
||||
import argparse
|
||||
import os
|
||||
from copy import deepcopy
|
||||
from multiprocessing import Pool
|
||||
from pathlib import Path
|
||||
from multiprocessing import Pool
|
||||
|
||||
import pandas as pd
|
||||
from scenedetect import SceneManager, open_video
|
||||
from scenedetect import open_video, SceneManager
|
||||
from scenedetect.detectors import ContentDetector
|
||||
from tqdm import tqdm
|
||||
|
||||
from utils.logger import logger
|
||||
|
||||
|
||||
@@ -15,14 +16,14 @@ def cutscene_detection_star(args):
|
||||
return cutscene_detection(*args)
|
||||
|
||||
|
||||
def cutscene_detection(video_path, video_folder, saved_path, cutscene_threshold=27, min_scene_len=15):
|
||||
def cutscene_detection(video_path, saved_path, cutscene_threshold=27, min_scene_len=15):
|
||||
try:
|
||||
if os.path.exists(saved_path):
|
||||
logger.info(f"{video_path} has been processed.")
|
||||
return
|
||||
# Use PyAV as the backend to avoid (to some exent) containing the last frame of the previous scene.
|
||||
# https://github.com/Breakthrough/PySceneDetect/issues/279#issuecomment-2152596761.
|
||||
video = open_video(os.path.join(video_folder, video_path), backend="pyav")
|
||||
video = open_video(video_path, backend="pyav")
|
||||
frame_rate, frame_size = video.frame_rate, video.frame_size
|
||||
duration = deepcopy(video.duration)
|
||||
|
||||
@@ -50,14 +51,12 @@ def cutscene_detection(video_path, video_folder, saved_path, cutscene_threshold=
|
||||
|
||||
timecode_list = [(frame_timecode_tuple[0].get_timecode(), frame_timecode_tuple[1].get_timecode()) for frame_timecode_tuple in output_scene_list]
|
||||
meta_scene = [{
|
||||
"video_path": video_path,
|
||||
"video_path": Path(video_path).name,
|
||||
"timecode_list": timecode_list,
|
||||
"fram_rate": frame_rate,
|
||||
"frame_size": frame_size,
|
||||
"duration": str(duration) # __repr__
|
||||
}]
|
||||
if not os.path.exists(Path(saved_path).parent):
|
||||
os.makedirs(Path(saved_path).parent, exist_ok=True)
|
||||
pd.DataFrame(meta_scene).to_json(saved_path, orient="records", lines=True)
|
||||
except Exception as e:
|
||||
logger.warning(f"Cutscene detection with {video_path} failed. Error is: {e}.")
|
||||
@@ -76,18 +75,20 @@ if __name__ == "__main__":
|
||||
)
|
||||
parser.add_argument("--video_folder", type=str, default="", help="The video folder.")
|
||||
parser.add_argument("--saved_folder", type=str, required=True, help="The save path to the output results (csv/jsonl).")
|
||||
parser.add_argument("--cutscene_threshold", type=int, default=27, help="The threshold of ContentDetector.")
|
||||
parser.add_argument("--n_jobs", type=int, default=1, help="The number of processes.")
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
metadata_df = pd.read_json(args.video_metadata_path, lines=True)
|
||||
video_path_list = metadata_df[args.video_path_column].tolist()
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
|
||||
if not os.path.exists(args.saved_folder):
|
||||
os.makedirs(args.saved_folder, exist_ok=True)
|
||||
# The glob can be slow when there are many small jsonl files.
|
||||
saved_path_list = [os.path.join(args.saved_folder, Path(video_path).with_suffix(".jsonl")) for video_path in video_path_list]
|
||||
saved_path_list = [os.path.join(args.saved_folder, Path(video_path).stem + ".jsonl") for video_path in video_path_list]
|
||||
args_list = [
|
||||
(video_path, args.video_folder, saved_path, args.cutscene_threshold)
|
||||
(video_path, saved_path)
|
||||
for video_path, saved_path in zip(video_path_list, saved_path_list)
|
||||
]
|
||||
# Since the length of the video is not uniform, the gather operation is not performed.
|
||||
@@ -3,8 +3,9 @@ import os
|
||||
|
||||
import pandas as pd
|
||||
from natsort import natsorted
|
||||
from utils.filter import filter
|
||||
|
||||
from utils.logger import logger
|
||||
from utils.filter import filter
|
||||
|
||||
|
||||
def parse_args():
|
||||
@@ -18,12 +19,6 @@ def parse_args():
|
||||
default="video_path",
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--caption_column",
|
||||
type=str,
|
||||
default="caption",
|
||||
help="The column contains the caption.",
|
||||
)
|
||||
parser.add_argument("--video_folder", type=str, default="", help="The video folder.")
|
||||
parser.add_argument(
|
||||
"--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl)."
|
||||
@@ -32,13 +27,13 @@ def parse_args():
|
||||
parser.add_argument("--min_duration", type=float, default=-1, help="The minimum duration.")
|
||||
parser.add_argument("--max_duration", type=float, default=-1, help="The maximum duration.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
"--asethetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score", type=float, default=4.0, help="The aesthetic score threshold.")
|
||||
parser.add_argument("--min_asethetic_score", type=float, default=4.0, help="The asethetic score threshold.")
|
||||
parser.add_argument(
|
||||
"--aesthetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality (SigLIP) metadata (csv/jsonl)."
|
||||
"--asethetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality (SigLIP) metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_aesthetic_score_siglip", type=float, default=4.0, help="The aesthetic score (SigLIP) threshold.")
|
||||
parser.add_argument("--min_asethetic_score_siglip", type=float, default=4.0, help="The asethetic score (SigLIP) threshold.")
|
||||
parser.add_argument(
|
||||
"--text_score_metadata_path", type=str, default=None, help="The path to the video text score metadata (csv/jsonl)."
|
||||
)
|
||||
@@ -68,10 +63,10 @@ def main():
|
||||
min_resolution=args.min_resolution,
|
||||
min_duration=args.min_duration,
|
||||
max_duration=args.max_duration,
|
||||
aesthetic_score_metadata_path=args.aesthetic_score_metadata_path,
|
||||
min_aesthetic_score=args.min_aesthetic_score,
|
||||
aesthetic_score_siglip_metadata_path=args.aesthetic_score_siglip_metadata_path,
|
||||
min_aesthetic_score_siglip=args.min_aesthetic_score_siglip,
|
||||
asethetic_score_metadata_path=args.asethetic_score_metadata_path,
|
||||
min_asethetic_score=args.min_asethetic_score,
|
||||
asethetic_score_siglip_metadata_path=args.asethetic_score_siglip_metadata_path,
|
||||
min_asethetic_score_siglip=args.min_asethetic_score_siglip,
|
||||
text_score_metadata_path=args.text_score_metadata_path,
|
||||
min_text_score=args.min_text_score,
|
||||
motion_score_metadata_path=args.motion_score_metadata_path,
|
||||
@@ -82,7 +77,7 @@ def main():
|
||||
)
|
||||
filtered_video_path_list = natsorted(filtered_video_path_list)
|
||||
filtered_caption_df = raw_caption_df[raw_caption_df[args.video_path_column].isin(filtered_video_path_list)]
|
||||
train_df = filtered_caption_df.rename(columns={args.video_path_column: "file_path", args.caption_column: "text"})
|
||||
train_df = filtered_caption_df.rename(columns={"video_path": "file_path", "caption": "text"})
|
||||
train_df["file_path"] = train_df["file_path"].map(lambda x: os.path.join(args.video_folder, x))
|
||||
train_df["type"] = "video"
|
||||
train_df.to_json(args.saved_path, orient="records", force_ascii=False, indent=2)
|
||||
@@ -1,19 +1,17 @@
|
||||
"""Modified from https://github.com/JaidedAI/EasyOCR/blob/803b907/easyocr/detection.py.
|
||||
1. Disable DataParallel.
|
||||
"""
|
||||
import torch
|
||||
import torch.backends.cudnn as cudnn
|
||||
from torch.autograd import Variable
|
||||
from PIL import Image
|
||||
from collections import OrderedDict
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.backends.cudnn as cudnn
|
||||
from PIL import Image
|
||||
from torch.autograd import Variable
|
||||
|
||||
from .craft_utils import getDetBoxes, adjustResultCoordinates
|
||||
from .imgproc import resize_aspect_ratio, normalizeMeanVariance
|
||||
from .craft import CRAFT
|
||||
from .craft_utils import adjustResultCoordinates, getDetBoxes
|
||||
from .imgproc import normalizeMeanVariance, resize_aspect_ratio
|
||||
|
||||
|
||||
def copyStateDict(state_dict):
|
||||
if list(state_dict.keys())[0].startswith("module"):
|
||||
@@ -84,7 +82,7 @@ def get_detector(trained_model, device='cpu', quantize=True, cudnn_benchmark=Fal
|
||||
if quantize:
|
||||
try:
|
||||
torch.quantization.quantize_dynamic(net, dtype=torch.qint8, inplace=True)
|
||||
except Exception:
|
||||
except:
|
||||
pass
|
||||
else:
|
||||
net.load_state_dict(copyStateDict(torch.load(trained_model, map_location=device)))
|
||||
@@ -0,0 +1,42 @@
|
||||
# Modified from https://github.com/NVlabs/VILA/blob/1c88211/llava/model/multimodal_encoder/siglip_encoder.py
|
||||
# 1. Support transformers >= 4.36.2.
|
||||
import torch
|
||||
import transformers
|
||||
from packaging import version
|
||||
from transformers import AutoConfig, AutoModel, PretrainedConfig
|
||||
|
||||
from llava.model.multimodal_encoder.vision_encoder import VisionTower, VisionTowerS2
|
||||
|
||||
if version.parse(transformers.__version__) > version.parse("4.36.2"):
|
||||
from transformers import SiglipImageProcessor, SiglipVisionConfig, SiglipVisionModel
|
||||
else:
|
||||
from .siglip import SiglipImageProcessor, SiglipVisionConfig, SiglipVisionModel
|
||||
|
||||
|
||||
class SiglipVisionTower(VisionTower):
|
||||
def __init__(self, model_name_or_path: str, config: PretrainedConfig, state_dict=None):
|
||||
super().__init__(model_name_or_path, config)
|
||||
self.image_processor = SiglipImageProcessor.from_pretrained(model_name_or_path)
|
||||
self.vision_tower = SiglipVisionModel.from_pretrained(
|
||||
# TODO(ligeng): why pass config here leading to errors?
|
||||
model_name_or_path, torch_dtype=eval(config.model_dtype), state_dict=state_dict
|
||||
)
|
||||
self.is_loaded = True
|
||||
|
||||
|
||||
class SiglipVisionTowerS2(VisionTowerS2):
|
||||
def __init__(self, model_name_or_path: str, config: PretrainedConfig):
|
||||
super().__init__(model_name_or_path, config)
|
||||
self.image_processor = SiglipImageProcessor.from_pretrained(model_name_or_path)
|
||||
self.vision_tower = SiglipVisionModel.from_pretrained(
|
||||
model_name_or_path, torch_dtype=eval(config.model_dtype)
|
||||
)
|
||||
|
||||
# Make sure it crops/resizes the image to the largest scale in self.scales to maintain high-res information
|
||||
self.image_processor.size['height'] = self.image_processor.size['width'] = self.scales[-1]
|
||||
|
||||
self.is_loaded = True
|
||||
|
||||
if version.parse(transformers.__version__) <= version.parse("4.36.2"):
|
||||
AutoConfig.register("siglip_vision_model", SiglipVisionConfig)
|
||||
AutoModel.register(SiglipVisionConfig, SiglipVisionModel)
|
||||
@@ -2,10 +2,8 @@ Please rewrite the video description to be useful for AI to re-generate the vide
|
||||
1. Do not start with something similar to 'The video/scene/frame shows' or "In this video/scene/frame".
|
||||
2. Remove the subjective content deviates from describing the visual content of the video. For instance, a sentence like "It gives a feeling of ease and tranquility and makes people feel comfortable" is considered subjective.
|
||||
3. Remove the non-existent description that does not in the visual content of the video, For instance, a sentence like "There is no visible detail that could be used to identify the individual beyond what is shown." is considered as the non-existent description.
|
||||
4. The rewritten description should include the main subject (person, object, animal, or none) actions and their attributes or status sequence, the background (the objects, location, weather, and time).
|
||||
5. If the original description includes the view shot, camera movement and the video style, the rewritten description should also include them. If not, there is no need to invent them on your own.
|
||||
6. Here are some examples of good descriptions: 1) A stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is damp and reflective, creating a mirror effect of the colorful lights. Many pedestrians walk about. 2) A large orange octopus is seen resting on the bottom of the ocean floor, blending in with the sandy and rocky terrain. Its tentacles are spread out around its body, and its eyes are closed. The octopus is unaware of a king crab that is crawling towards it from behind a rock, its claws raised and ready to attack. The crab is brown and spiny, with long legs and antennae. The scene is captured from a wide angle, showing the vastness and depth of the ocean. The water is clear and blue, with rays of sunlight filtering through. The shot is sharp and crisp, with a high dynamic range. The octopus and the crab are in focus, while the background is slightly blurred, creating a depth of field effect.
|
||||
7. Output with the following json format:
|
||||
4. Here are some examples of good descriptions: 1) A stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is damp and reflective, creating a mirror effect of the colorful lights. Many pedestrians walk about. 2) A large orange octopus is seen resting on the bottom of the ocean floor, blending in with the sandy and rocky terrain. Its tentacles are spread out around its body, and its eyes are closed. The octopus is unaware of a king crab that is crawling towards it from behind a rock, its claws raised and ready to attack. The crab is brown and spiny, with long legs and antennae. The scene is captured from a wide angle, showing the vastness and depth of the ocean. The water is clear and blue, with rays of sunlight filtering through. The shot is sharp and crisp, with a high dynamic range. The octopus and the crab are in focus, while the background is slightly blurred, creating a depth of field effect.
|
||||
5. Output with the following json format:
|
||||
{"rewritten description": "your rewritten description here"}
|
||||
|
||||
Here is the video description:
|
||||
@@ -4,4 +4,6 @@ git+https://github.com/openai/CLIP.git
|
||||
natsort
|
||||
joblib
|
||||
scenedetect
|
||||
av
|
||||
av
|
||||
# https://github.com/NVlabs/VILA/issues/78#issuecomment-2195568292
|
||||
numpy<2.0.0
|
||||
@@ -0,0 +1,41 @@
|
||||
META_FILE_PATH="datasets/panda_70m/videos_clips/data/meta_file_info.jsonl"
|
||||
VIDEO_FOLDER="datasets/panda_70m/videos_clips/data/"
|
||||
VIDEO_QUALITY_SAVED_PATH="datasets/panda_70m/videos_clips/meta_quality_info_siglip.jsonl"
|
||||
MIN_ASETHETIC_SCORE_SIGLIP=4.0
|
||||
TEXT_SAVED_PATH="datasets/panda_70m/videos_clips/meta_text_info.jsonl"
|
||||
MIN_TEXT_SCORE=0.02
|
||||
MOTION_SAVED_PATH="datasets/panda_70m/videos_clips/meta_motion_info.jsonl"
|
||||
|
||||
python -m utils.get_meta_file \
|
||||
--video_folder $VIDEO_FOLDER \
|
||||
--saved_path $META_FILE_PATH
|
||||
|
||||
# Get the asethetic score (SigLIP) of all videos
|
||||
accelerate launch compute_video_quality.py \
|
||||
--video_metadata_path $META_FILE_PATH \
|
||||
--video_folder $VIDEO_FOLDER \
|
||||
--metrics "AestheticScoreSigLIP" \
|
||||
--frame_sample_method uniform \
|
||||
--num_sampled_frames 4 \
|
||||
--saved_freq 10 \
|
||||
--saved_path $VIDEO_QUALITY_SAVED_PATH \
|
||||
--batch_size 4
|
||||
|
||||
# Get the text score of all videos filtered by the video quality score.
|
||||
accelerate launch compute_text_score.py \
|
||||
--video_metadata_path $META_FILE_PATH \
|
||||
--video_folder $VIDEO_FOLDER \
|
||||
--saved_freq 10 \
|
||||
--saved_path $TEXT_SAVED_PATH \
|
||||
--asethetic_score_siglip_metadata_path $VIDEO_QUALITY_SAVED_PATH \
|
||||
--min_asethetic_score_siglip $MIN_ASETHETIC_SCORE_SIGLIP
|
||||
|
||||
# Get the motion score of all videos filtered by the video quality score and text score.
|
||||
python compute_motion_score.py \
|
||||
--video_metadata_path $META_FILE_PATH \
|
||||
--video_folder $VIDEO_FOLDER \
|
||||
--saved_freq 10 \
|
||||
--saved_path $MOTION_SAVED_PATH \
|
||||
--n_jobs 8 \
|
||||
--text_score_metadata_path $TEXT_SAVED_PATH \
|
||||
--min_text_score $MIN_TEXT_SCORE
|
||||
@@ -0,0 +1,52 @@
|
||||
META_FILE_PATH="datasets/panda_70m/videos_clips/data/meta_file_info.jsonl"
|
||||
VIDEO_FOLDER="datasets/panda_70m/videos_clips/data/"
|
||||
MOTION_SAVED_PATH="datasets/panda_70m/videos_clips/meta_motion_info.jsonl"
|
||||
MIN_MOTION_SCORE=2
|
||||
VIDEO_CAPTION_SAVED_PATH="datasets/panda_70m/meta_caption_info_vila_8b.jsonl"
|
||||
REWRITTEN_VIDEO_CAPTION_SAVED_PATH="datasets/panda_70m/meta_caption_info_vila_8b_rewritten.jsonl"
|
||||
VIDEOCLIPXL_SCORE_SAVED_PATH="datasets/panda_70m/meta_caption_info_vila_8b_rewritten_videoclipxl.jsonl"
|
||||
MIN_VIDEOCLIPXL_SCORE=0.20
|
||||
TRAIN_SAVED_PATH="datasets/panda_70m/train_panda_70m.json"
|
||||
# Manually download Efficient-Large-Model/Llama-3-VILA1.5-8b-AWQ to VILA_MODEL_PATH.
|
||||
# Manually download meta-llama/Meta-Llama-3-8B-Instruct to REWRITE_MODEL_PATH.
|
||||
|
||||
# Use VILA1.5-AWQ to perform recaptioning.
|
||||
accelerate launch vila_video_recaptioning.py \
|
||||
--video_metadata_path ${META_FILE_PATH} \
|
||||
--video_folder ${VIDEO_FOLDER} \
|
||||
--model_path ${VILA_MODEL_PATH} \
|
||||
--precision "W4A16" \
|
||||
--saved_path $VIDEO_CAPTION_SAVED_PATH \
|
||||
--saved_freq 1 \
|
||||
--motion_score_metadata_path $MOTION_SAVED_PATH \
|
||||
--min_motion_score $MIN_MOTION_SCORE
|
||||
|
||||
# Rewrite video captions (optional).
|
||||
python caption_rewrite.py \
|
||||
--video_metadata_path $VIDEO_CAPTION_SAVED_PATH \
|
||||
--batch_size 4096 \
|
||||
--model_name $REWRITE_MODEL_PATH \
|
||||
--prompt prompt/rewrite.txt \
|
||||
--prefix '"rewritten description": ' \
|
||||
--saved_path $REWRITTEN_VIDEO_CAPTION_SAVED_PATH \
|
||||
--saved_freq 1
|
||||
|
||||
# Compute caption-video alignment (optional).
|
||||
accelerate launch compute_video_quality.py \
|
||||
--video_metadata_path $REWRITTEN_VIDEO_CAPTION_SAVED_PATH \
|
||||
--caption_column caption \
|
||||
--video_folder $VIDEO_FOLDER \
|
||||
--frame_sample_method uniform \
|
||||
--num_sampled_frames 8 \
|
||||
--metrics VideoCLIPXLScore \
|
||||
--batch_size 4 \
|
||||
--saved_path $VIDEOCLIPXL_SCORE_SAVED_PATH \
|
||||
--saved_freq 10
|
||||
|
||||
# Get the final train file.
|
||||
python filter_meta_train.py \
|
||||
--caption_metadata_path $REWRITTEN_VIDEO_CAPTION_SAVED_PATH \
|
||||
--video_folder=$VIDEO_FOLDER \
|
||||
--videoclipxl_score_metadata_path $VIDEOCLIPXL_SCORE_SAVED_PATH \
|
||||
--min_videoclipxl_score $MIN_VIDEOCLIPXL_SCORE \
|
||||
--saved_path=$TRAIN_SAVED_PATH
|
||||
@@ -1,42 +1,41 @@
|
||||
import ast
|
||||
from typing import Optional
|
||||
import os
|
||||
|
||||
import pandas as pd
|
||||
|
||||
from .logger import logger
|
||||
|
||||
|
||||
# Ensure each item in the video_path_list matches the paths in the video_path column of the metadata.
|
||||
def filter(
|
||||
video_path_list: list[str],
|
||||
basic_metadata_path: Optional[str] = None,
|
||||
min_resolution: float = 720*1280,
|
||||
min_duration: float = -1,
|
||||
max_duration: float = -1,
|
||||
aesthetic_score_metadata_path: Optional[str] = None,
|
||||
min_aesthetic_score: float = 4,
|
||||
aesthetic_score_siglip_metadata_path: Optional[str] = None,
|
||||
min_aesthetic_score_siglip: float = 4,
|
||||
text_score_metadata_path: Optional[str] = None,
|
||||
min_text_score: float = 0.02,
|
||||
motion_score_metadata_path: Optional[str] = None,
|
||||
min_motion_score: float = 2,
|
||||
max_motion_score: float = 999999,
|
||||
videoclipxl_score_metadata_path: Optional[str] = None,
|
||||
min_videoclipxl_score: float = 0.20,
|
||||
semantic_consistency_score_metadata_path: Optional[list[str]] = None,
|
||||
min_semantic_consistency_score: float = 0.80,
|
||||
video_path_column: str = "video_path"
|
||||
video_path_list,
|
||||
basic_metadata_path=None,
|
||||
min_resolution=0,
|
||||
min_duration=-1,
|
||||
max_duration=-1,
|
||||
asethetic_score_metadata_path=None,
|
||||
min_asethetic_score=4,
|
||||
asethetic_score_siglip_metadata_path=None,
|
||||
min_asethetic_score_siglip=4,
|
||||
text_score_metadata_path=None,
|
||||
min_text_score=0.02,
|
||||
motion_score_metadata_path=None,
|
||||
min_motion_score=2,
|
||||
videoclipxl_score_metadata_path=None,
|
||||
min_videoclipxl_score=0.20,
|
||||
video_path_column="video_path",
|
||||
):
|
||||
video_path_list = [os.path.basename(video_path) for video_path in video_path_list]
|
||||
|
||||
if basic_metadata_path is not None:
|
||||
if basic_metadata_path.endswith(".csv"):
|
||||
basic_df = pd.read_csv(basic_metadata_path)
|
||||
elif basic_metadata_path.endswith(".jsonl"):
|
||||
basic_df = pd.read_json(basic_metadata_path, lines=True)
|
||||
|
||||
|
||||
basic_df["resolution"] = basic_df["frame_size"].apply(lambda x: x[0] * x[1])
|
||||
filtered_basic_df = basic_df[basic_df["resolution"] < min_resolution]
|
||||
filtered_video_path_list = filtered_basic_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
@@ -47,6 +46,7 @@ def filter(
|
||||
if min_duration != -1:
|
||||
filtered_basic_df = basic_df[basic_df["duration"] < min_duration]
|
||||
filtered_video_path_list = filtered_basic_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
@@ -57,6 +57,7 @@ def filter(
|
||||
if max_duration != -1:
|
||||
filtered_basic_df = basic_df[basic_df["duration"] > max_duration]
|
||||
filtered_video_path_list = filtered_basic_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
@@ -64,48 +65,50 @@ def filter(
|
||||
f"with duration greater than {max_duration}."
|
||||
)
|
||||
|
||||
if aesthetic_score_metadata_path is not None:
|
||||
if aesthetic_score_metadata_path.endswith(".csv"):
|
||||
aesthetic_score_df = pd.read_csv(aesthetic_score_metadata_path)
|
||||
elif aesthetic_score_metadata_path.endswith(".jsonl"):
|
||||
aesthetic_score_df = pd.read_json(aesthetic_score_metadata_path, lines=True)
|
||||
if asethetic_score_metadata_path is not None:
|
||||
if asethetic_score_metadata_path.endswith(".csv"):
|
||||
asethetic_score_df = pd.read_csv(asethetic_score_metadata_path)
|
||||
elif asethetic_score_metadata_path.endswith(".jsonl"):
|
||||
asethetic_score_df = pd.read_json(asethetic_score_metadata_path, lines=True)
|
||||
|
||||
# In pandas, csv will save lists as strings, whereas jsonl will not.
|
||||
aesthetic_score_df["aesthetic_score"] = aesthetic_score_df["aesthetic_score"].apply(
|
||||
asethetic_score_df["aesthetic_score"] = asethetic_score_df["aesthetic_score"].apply(
|
||||
lambda x: ast.literal_eval(x) if isinstance(x, str) else x
|
||||
)
|
||||
aesthetic_score_df["aesthetic_score_mean"] = aesthetic_score_df["aesthetic_score"].apply(lambda x: sum(x) / len(x))
|
||||
filtered_aesthetic_score_df = aesthetic_score_df[aesthetic_score_df["aesthetic_score_mean"] < min_aesthetic_score]
|
||||
filtered_video_path_list = filtered_aesthetic_score_df[video_path_column].tolist()
|
||||
asethetic_score_df["aesthetic_score_mean"] = asethetic_score_df["aesthetic_score"].apply(lambda x: sum(x) / len(x))
|
||||
filtered_asethetic_score_df = asethetic_score_df[asethetic_score_df["aesthetic_score_mean"] < min_asethetic_score]
|
||||
filtered_video_path_list = filtered_asethetic_score_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {aesthetic_score_metadata_path} ({len(aesthetic_score_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with aesthetic score less than {min_aesthetic_score}."
|
||||
f"Load {asethetic_score_metadata_path} ({len(asethetic_score_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with aesthetic score less than {min_asethetic_score}."
|
||||
)
|
||||
|
||||
if aesthetic_score_siglip_metadata_path is not None:
|
||||
if aesthetic_score_siglip_metadata_path.endswith(".csv"):
|
||||
aesthetic_score_siglip_df = pd.read_csv(aesthetic_score_siglip_metadata_path)
|
||||
elif aesthetic_score_siglip_metadata_path.endswith(".jsonl"):
|
||||
aesthetic_score_siglip_df = pd.read_json(aesthetic_score_siglip_metadata_path, lines=True)
|
||||
|
||||
if asethetic_score_siglip_metadata_path is not None:
|
||||
if asethetic_score_siglip_metadata_path.endswith(".csv"):
|
||||
asethetic_score_siglip_df = pd.read_csv(asethetic_score_siglip_metadata_path)
|
||||
elif asethetic_score_siglip_metadata_path.endswith(".jsonl"):
|
||||
asethetic_score_siglip_df = pd.read_json(asethetic_score_siglip_metadata_path, lines=True)
|
||||
|
||||
# In pandas, csv will save lists as strings, whereas jsonl will not.
|
||||
aesthetic_score_siglip_df["aesthetic_score_siglip"] = aesthetic_score_siglip_df["aesthetic_score_siglip"].apply(
|
||||
asethetic_score_siglip_df["aesthetic_score_siglip"] = asethetic_score_siglip_df["aesthetic_score_siglip"].apply(
|
||||
lambda x: ast.literal_eval(x) if isinstance(x, str) else x
|
||||
)
|
||||
aesthetic_score_siglip_df["aesthetic_score_siglip_mean"] = aesthetic_score_siglip_df["aesthetic_score_siglip"].apply(
|
||||
asethetic_score_siglip_df["aesthetic_score_siglip_mean"] = asethetic_score_siglip_df["aesthetic_score_siglip"].apply(
|
||||
lambda x: sum(x) / len(x)
|
||||
)
|
||||
filtered_aesthetic_score_siglip_df = aesthetic_score_siglip_df[
|
||||
aesthetic_score_siglip_df["aesthetic_score_siglip_mean"] < min_aesthetic_score_siglip
|
||||
filtered_asethetic_score_siglip_df = asethetic_score_siglip_df[
|
||||
asethetic_score_siglip_df["aesthetic_score_siglip_mean"] < min_asethetic_score_siglip
|
||||
]
|
||||
filtered_video_path_list = filtered_aesthetic_score_siglip_df[video_path_column].tolist()
|
||||
filtered_video_path_list = filtered_asethetic_score_siglip_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {aesthetic_score_siglip_metadata_path} ({len(aesthetic_score_siglip_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with aesthetic score (SigLIP) less than {min_aesthetic_score_siglip}."
|
||||
f"Load {asethetic_score_siglip_metadata_path} ({len(asethetic_score_siglip_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with aesthetic score (SigLIP) less than {min_asethetic_score_siglip}."
|
||||
)
|
||||
|
||||
if text_score_metadata_path is not None:
|
||||
@@ -116,6 +119,7 @@ def filter(
|
||||
|
||||
filtered_text_score_df = text_score_df[text_score_df["text_score"] > min_text_score]
|
||||
filtered_video_path_list = filtered_text_score_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
@@ -128,24 +132,16 @@ def filter(
|
||||
motion_score_df = pd.read_csv(motion_score_metadata_path)
|
||||
elif motion_score_metadata_path.endswith(".jsonl"):
|
||||
motion_score_df = pd.read_json(motion_score_metadata_path, lines=True)
|
||||
|
||||
|
||||
filtered_motion_score_df = motion_score_df[motion_score_df["motion_score"] < min_motion_score]
|
||||
filtered_video_path_list = filtered_motion_score_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {motion_score_metadata_path} ({len(motion_score_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with motion score smaller than {min_motion_score}."
|
||||
)
|
||||
|
||||
filtered_motion_score_df = motion_score_df[motion_score_df["motion_score"] > max_motion_score]
|
||||
filtered_video_path_list = filtered_motion_score_df[video_path_column].tolist()
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {motion_score_metadata_path} ({len(motion_score_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with motion score greater than {min_motion_score}."
|
||||
)
|
||||
|
||||
if videoclipxl_score_metadata_path is not None:
|
||||
if videoclipxl_score_metadata_path.endswith(".csv"):
|
||||
@@ -155,28 +151,12 @@ def filter(
|
||||
|
||||
filtered_videoclipxl_score_df = videoclipxl_score_df[videoclipxl_score_df["videoclipxl_score"] < min_videoclipxl_score]
|
||||
filtered_video_path_list = filtered_videoclipxl_score_df[video_path_column].tolist()
|
||||
filtered_video_path_list = [os.path.basename(video_path) for video_path in filtered_video_path_list]
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {videoclipxl_score_metadata_path} ({len(videoclipxl_score_df)}) and "
|
||||
f"filter {len(filtered_video_path_list)} videos with mixclip score smaller than {min_videoclipxl_score}."
|
||||
)
|
||||
|
||||
if semantic_consistency_score_metadata_path is not None:
|
||||
for f in semantic_consistency_score_metadata_path:
|
||||
if f.endswith(".csv"):
|
||||
semantic_consistency_score_df = pd.read_csv(f)
|
||||
elif f.endswith(".jsonl"):
|
||||
semantic_consistency_score_df = pd.read_json(f, lines=True)
|
||||
filtered_semantic_consistency_score_df = semantic_consistency_score_df[
|
||||
semantic_consistency_score_df["similarity_cross_frame"].apply(lambda x: min(x) < min_semantic_consistency_score)
|
||||
]
|
||||
filtered_video_path_list = filtered_semantic_consistency_score_df[video_path_column].tolist()
|
||||
|
||||
video_path_list = list(set(video_path_list).difference(set(filtered_video_path_list)))
|
||||
logger.info(
|
||||
f"Load {f} ({len(semantic_consistency_score_df)}) and filter {len(filtered_video_path_list)} videos "
|
||||
f"with the minimum semantic consistency score smaller than {min_semantic_consistency_score}."
|
||||
)
|
||||
|
||||
return video_path_list
|
||||
@@ -1,13 +1,12 @@
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
from multiprocessing import Manager, Pool
|
||||
from pathlib import Path
|
||||
import glob
|
||||
import json
|
||||
from multiprocessing import Pool, Manager
|
||||
|
||||
import pandas as pd
|
||||
from natsort import index_natsorted
|
||||
|
||||
from .get_meta_file import parallel_rglob
|
||||
from .logger import logger
|
||||
|
||||
|
||||
@@ -21,10 +20,9 @@ def process_file(file_path, shared_list):
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="Gather all jsonl files in a folder (meta_folder) to a single jsonl file (meta_file_path).")
|
||||
parser.add_argument("--meta_folder", type=str, required=True)
|
||||
parser.add_argument("--video_path_column", type=str, default="video_path")
|
||||
parser.add_argument("--meta_file_path", type=str, required=True)
|
||||
parser.add_argument("--video_path_column", type=str, default="video_path")
|
||||
parser.add_argument("--n_jobs", type=int, default=1)
|
||||
parser.add_argument("--recursive", action="store_true", help="Whether to search sub-folders recursively.")
|
||||
|
||||
args = parser.parse_args()
|
||||
return args
|
||||
@@ -33,13 +31,7 @@ def parse_args():
|
||||
def main():
|
||||
args = parse_args()
|
||||
|
||||
if not os.path.exists(args.meta_folder):
|
||||
raise ValueError(f"The meta_folder {args.meta_folder} does not exist.")
|
||||
meta_folder = Path(args.meta_folder)
|
||||
if args.recursive:
|
||||
jsonl_files = [str(file) for file in parallel_rglob(meta_folder, f"*.jsonl", max_workers=args.n_jobs)]
|
||||
else:
|
||||
jsonl_files = [str(file) for file in meta_folder.glob(f"*.jsonl")]
|
||||
jsonl_files = glob.glob(os.path.join(args.meta_folder, "*.jsonl"))
|
||||
|
||||
with Manager() as manager:
|
||||
shared_list = manager.list()
|
||||
@@ -1,40 +1,17 @@
|
||||
import argparse
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import pandas as pd
|
||||
from natsort import natsorted
|
||||
from tqdm import tqdm
|
||||
|
||||
from .logger import logger
|
||||
|
||||
ALL_VIDEO_EXT = set(["mp4", "webm", "mkv", "avi", "flv", "mov", "rmvb"])
|
||||
|
||||
ALL_VIDEO_EXT = set(["mp4", "webm", "mkv", "avi", "flv", "mov"])
|
||||
ALL_IMGAE_EXT = set(["png", "webp", "jpg", "jpeg", "bmp", "gif"])
|
||||
|
||||
|
||||
def get_relative_file_paths(directory, recursive=False, ext_set=None):
|
||||
"""Get the relative paths of subfiles (recursively) in the directory that match the extension set.
|
||||
"""
|
||||
if not recursive:
|
||||
for entry in os.scandir(directory):
|
||||
if entry.is_file():
|
||||
file_name = entry.name
|
||||
if ext_set is not None:
|
||||
ext = os.path.splitext(file_name)[1][1:].lower()
|
||||
if ext in ext_set:
|
||||
yield file_name
|
||||
else:
|
||||
yield file_name
|
||||
else:
|
||||
for root, _, files in os.walk(directory):
|
||||
for file in files:
|
||||
relative_path = os.path.relpath(os.path.join(root, file), directory)
|
||||
if ext_set is not None:
|
||||
ext = os.path.splitext(file)[1][1:].lower()
|
||||
if ext in ext_set:
|
||||
yield relative_path
|
||||
else:
|
||||
yield relative_path
|
||||
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="Compute scores of uniform sampled frames from videos.")
|
||||
parser.add_argument(
|
||||
@@ -65,19 +42,27 @@ def main():
|
||||
raise ValueError("Either video_folder or image_folder should be specified in the arguments.")
|
||||
if args.video_folder is not None and args.image_folder is not None:
|
||||
raise ValueError("Both video_folder and image_folder can not be specified in the arguments at the same time.")
|
||||
if args.image_folder is None and not os.path.exists(args.video_folder):
|
||||
raise ValueError(f"The video_folder {args.video_folder} does not exist.")
|
||||
if args.video_folder is None and not os.path.exists(args.image_folder):
|
||||
raise ValueError(f"The image_folder {args.image_folder} does not exist.")
|
||||
|
||||
# Use the path name instead of the file name as video_path/image_path (unique ID).
|
||||
if args.video_folder is not None:
|
||||
video_path_list = list(get_relative_file_paths(args.video_folder, recursive=args.recursive, ext_set=ALL_VIDEO_EXT))
|
||||
video_path_list = []
|
||||
video_folder = Path(args.video_folder)
|
||||
for ext in tqdm(list(ALL_VIDEO_EXT)):
|
||||
if args.recursive:
|
||||
video_path_list += [str(file.relative_to(video_folder)) for file in video_folder.rglob(f"*.{ext}")]
|
||||
else:
|
||||
video_path_list += [str(file.relative_to(video_folder)) for file in video_folder.glob(f"*.{ext}")]
|
||||
video_path_list = natsorted(video_path_list)
|
||||
meta_file_df = pd.DataFrame({args.video_path_column: video_path_list})
|
||||
|
||||
if args.image_folder is not None:
|
||||
image_path_list = list(get_relative_file_paths(args.image_folder, recursive=args.recursive, ext_set=ALL_IMGAE_EXT))
|
||||
image_path_list = []
|
||||
image_folder = Path(args.image_folder)
|
||||
for ext in tqdm(list(ALL_IMGAE_EXT)):
|
||||
if args.recursive:
|
||||
image_path_list += [str(file.relative_to(image_folder)) for file in image_folder.rglob(f"*.{ext}")]
|
||||
else:
|
||||
image_path_list += [str(file.relative_to(image_folder)) for file in image_folder.glob(f"*.{ext}")]
|
||||
image_path_list = natsorted(image_path_list)
|
||||
meta_file_df = pd.DataFrame({args.image_path_column: image_path_list})
|
||||
|
||||
@@ -223,7 +223,6 @@ class CLIPScore:
|
||||
if __name__ == "__main__":
|
||||
from torch.utils.data import DataLoader
|
||||
from tqdm import tqdm
|
||||
|
||||
from .video_dataset import VideoDataset, collate_fn
|
||||
|
||||
aesthetic_score = AestheticScore(device="cuda")
|
||||
@@ -2,14 +2,12 @@ import hashlib
|
||||
import os
|
||||
import urllib
|
||||
import warnings
|
||||
from typing import Any, List, Union
|
||||
|
||||
import torch
|
||||
from PIL import Image
|
||||
from typing import Any, Union, List
|
||||
from pkg_resources import packaging
|
||||
from torch import nn
|
||||
from torchvision.transforms import (CenterCrop, Compose, Normalize, Resize,
|
||||
ToTensor)
|
||||
import torch
|
||||
from PIL import Image
|
||||
from torchvision.transforms import Compose, Resize, CenterCrop, ToTensor, Normalize
|
||||
from tqdm import tqdm
|
||||
|
||||
from .model_longclip import build_model
|
||||
@@ -98,7 +98,7 @@ class SimpleTokenizer(object):
|
||||
j = word.index(first, i)
|
||||
new_word.extend(word[i:j])
|
||||
i = j
|
||||
except Exception:
|
||||
except:
|
||||
new_word.extend(word[i:])
|
||||
break
|
||||
|
||||
@@ -6,8 +6,12 @@ from typing import Final
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from transformers import (SiglipImageProcessor, SiglipVisionConfig,
|
||||
SiglipVisionModel, logging)
|
||||
from transformers import (
|
||||
SiglipImageProcessor,
|
||||
SiglipVisionConfig,
|
||||
SiglipVisionModel,
|
||||
logging,
|
||||
)
|
||||
from transformers.image_processing_utils import BatchFeature
|
||||
from transformers.modeling_outputs import ImageClassifierOutputWithNoAttention
|
||||
|
||||
@@ -1,11 +1,9 @@
|
||||
import os
|
||||
|
||||
import cv2
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from .simple_tokenizer import SimpleTokenizer as _Tokenizer
|
||||
from .viclip import ViCLIP
|
||||
import torch
|
||||
import numpy as np
|
||||
import cv2
|
||||
import os
|
||||
|
||||
|
||||
def get_viclip(size='l',
|
||||
@@ -101,7 +101,7 @@ class SimpleTokenizer(object):
|
||||
j = word.index(first, i)
|
||||
new_word.extend(word[i:j])
|
||||
i = j
|
||||
except Exception:
|
||||
except:
|
||||
new_word.extend(word[i:])
|
||||
break
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
import logging
|
||||
import math
|
||||
import os
|
||||
import logging
|
||||
|
||||
import torch
|
||||
from einops import rearrange
|
||||
from torch import nn
|
||||
import math
|
||||
|
||||
# from .criterions import VTC_VTM_Loss
|
||||
from .simple_tokenizer import SimpleTokenizer as _Tokenizer
|
||||
from .viclip_text import clip_text_b16, clip_text_l14
|
||||
from .viclip_vision import clip_joint_b16, clip_joint_l14
|
||||
from .viclip_vision import clip_joint_l14, clip_joint_b16
|
||||
from .viclip_text import clip_text_l14, clip_text_b16
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -1,16 +1,15 @@
|
||||
import functools
|
||||
import logging
|
||||
import os
|
||||
import logging
|
||||
from collections import OrderedDict
|
||||
from pkg_resources import packaging
|
||||
from .simple_tokenizer import SimpleTokenizer as _Tokenizer
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
import torch.utils.checkpoint as checkpoint
|
||||
from pkg_resources import packaging
|
||||
from torch import nn
|
||||
|
||||
from .simple_tokenizer import SimpleTokenizer as _Tokenizer
|
||||
import torch.utils.checkpoint as checkpoint
|
||||
import functools
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -1,14 +1,15 @@
|
||||
#!/usr/bin/env python
|
||||
import logging
|
||||
import os
|
||||
import logging
|
||||
from collections import OrderedDict
|
||||
|
||||
import torch
|
||||
import torch.utils.checkpoint as checkpoint
|
||||
from torch import nn
|
||||
from einops import rearrange
|
||||
from timm.models.layers import DropPath
|
||||
from timm.models.registry import register_model
|
||||
from torch import nn
|
||||
|
||||
import torch.utils.checkpoint as checkpoint
|
||||
|
||||
# from models.utils import load_temp_embed_with_mismatch
|
||||
|
||||
@@ -339,9 +340,9 @@ def interpolate_pos_embed_vit(state_dict, new_model):
|
||||
|
||||
if __name__ == '__main__':
|
||||
import time
|
||||
|
||||
from fvcore.nn import FlopCountAnalysis
|
||||
from fvcore.nn import flop_count_table
|
||||
import numpy as np
|
||||
from fvcore.nn import FlopCountAnalysis, flop_count_table
|
||||
|
||||
seed = 4217
|
||||
np.random.seed(seed)
|
||||
@@ -1,16 +1,16 @@
|
||||
import os
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
from pathlib import Path
|
||||
|
||||
from func_timeout import FunctionTimedOut, func_timeout
|
||||
from func_timeout import func_timeout, FunctionTimedOut
|
||||
from PIL import Image
|
||||
from torch.utils.data import DataLoader, Dataset
|
||||
from torch.utils.data import Dataset, DataLoader
|
||||
|
||||
from .logger import logger
|
||||
from .video_utils import extract_frames
|
||||
|
||||
|
||||
ALL_VIDEO_EXT = set([".mp4", ".webm", ".mkv", ".avi", ".flv", ".mov", ".ts"])
|
||||
ALL_VIDEO_EXT = set(["mp4", "webm", "mkv", "avi", "flv", "mov"])
|
||||
VIDEO_READER_TIMEOUT = 300
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ class VideoDataset(Dataset):
|
||||
text_column: Optional[str] = None,
|
||||
sample_method: str = "mid",
|
||||
num_sampled_frames: int = 1,
|
||||
sample_stride: Optional[int] = None
|
||||
num_sample_stride: Optional[int] = None
|
||||
):
|
||||
length = len(dataset_inputs[list(dataset_inputs.keys())[0]])
|
||||
if not all(len(v) == length for v in dataset_inputs.values()):
|
||||
@@ -46,7 +46,7 @@ class VideoDataset(Dataset):
|
||||
|
||||
self.sample_method = sample_method
|
||||
self.num_sampled_frames = num_sampled_frames
|
||||
self.sample_stride = sample_stride
|
||||
self.num_sample_stride = num_sample_stride
|
||||
|
||||
def __getitem__(self, index):
|
||||
video_path = self.video_path_list[index]
|
||||
@@ -61,7 +61,7 @@ class VideoDataset(Dataset):
|
||||
else:
|
||||
# It is a trick to deal with decord hanging when reading some abnormal videos.
|
||||
try:
|
||||
sample_args = (video_path, self.sample_method, self.num_sampled_frames, self.sample_stride)
|
||||
sample_args = (video_path, self.sample_method, self.num_sampled_frames, self.num_sample_stride)
|
||||
sampled_frame_idx_list, sampled_frame_list = func_timeout(
|
||||
VIDEO_READER_TIMEOUT, extract_frames, args=sample_args
|
||||
)
|
||||
@@ -0,0 +1,44 @@
|
||||
import gc
|
||||
import random
|
||||
from contextlib import contextmanager
|
||||
from typing import List, Tuple, Optional
|
||||
|
||||
import numpy as np
|
||||
from decord import VideoReader
|
||||
from PIL import Image
|
||||
|
||||
|
||||
@contextmanager
|
||||
def video_reader(*args, **kwargs):
|
||||
"""A context manager to solve the memory leak of decord.
|
||||
"""
|
||||
vr = VideoReader(*args, **kwargs)
|
||||
try:
|
||||
yield vr
|
||||
finally:
|
||||
del vr
|
||||
gc.collect()
|
||||
|
||||
|
||||
def extract_frames(
|
||||
video_path: str,
|
||||
sample_method: str = "mid",
|
||||
num_sampled_frames: int = -1,
|
||||
sample_stride: int = -1,
|
||||
**kwargs
|
||||
) -> Optional[Tuple[List[int], List[Image.Image]]]:
|
||||
with video_reader(video_path, num_threads=2, **kwargs) as vr:
|
||||
if sample_method == "mid":
|
||||
sampled_frame_idx_list = [len(vr) // 2]
|
||||
elif sample_method == "uniform":
|
||||
sampled_frame_idx_list = np.linspace(0, len(vr), num_sampled_frames, endpoint=False, dtype=int)
|
||||
elif sample_method == "random":
|
||||
clip_length = min(len(vr), (num_sampled_frames - 1) * sample_stride + 1)
|
||||
start_idx = random.randint(0, len(vr) - clip_length)
|
||||
sampled_frame_idx_list = np.linspace(start_idx, start_idx + clip_length - 1, num_sampled_frames, dtype=int)
|
||||
else:
|
||||
raise ValueError(f"The sample_method {sample_method} must be mid, uniform or random.")
|
||||
sampled_frame_list = vr.get_batch(sampled_frame_idx_list).asnumpy()
|
||||
sampled_frame_list = [Image.fromarray(frame) for frame in sampled_frame_list]
|
||||
|
||||
return list(sampled_frame_idx_list), sampled_frame_list
|
||||
@@ -2,13 +2,15 @@ import argparse
|
||||
import os
|
||||
import subprocess
|
||||
from datetime import datetime, timedelta
|
||||
from multiprocessing import Pool
|
||||
from pathlib import Path
|
||||
from multiprocessing import Pool
|
||||
|
||||
import pandas as pd
|
||||
from tqdm import tqdm
|
||||
|
||||
from utils.logger import logger
|
||||
|
||||
|
||||
MIN_SECONDS = int(os.getenv("MIN_SECONDS", 3))
|
||||
MAX_SECONDS = int(os.getenv("MAX_SECONDS", 10))
|
||||
|
||||
@@ -38,11 +40,10 @@ def clip_video_star(args):
|
||||
|
||||
def clip_video(video_path, timecode_list, output_folder, video_duration):
|
||||
"""Recursively clip the video within the range of [MIN_SECONDS, MAX_SECONDS],
|
||||
according to the timecode obtained from easyanimate/video_caption/cutscene_detect.py.
|
||||
according to the timecode obtained from cogvideox/video_caption/cutscene_detect.py.
|
||||
"""
|
||||
try:
|
||||
os.makedirs(output_folder, exist_ok=True)
|
||||
video_stem = Path(video_path).stem
|
||||
video_name = Path(video_path).stem
|
||||
|
||||
if len(timecode_list) == 0: # The video of a single scene.
|
||||
splitted_timecode_list = []
|
||||
@@ -58,7 +59,7 @@ def clip_video(video_path, timecode_list, output_folder, video_duration):
|
||||
splitted_index += 1
|
||||
continue
|
||||
splitted_timecode_list.append([cur_start.strftime("%H:%M:%S.%f")[:-3], cur_end.strftime("%H:%M:%S.%f")[:-3]])
|
||||
output_path = os.path.join(output_folder, video_stem + f"_{splitted_index}.mp4")
|
||||
output_path = os.path.join(output_folder, video_name + f"_{splitted_index}.mp4")
|
||||
if os.path.exists(output_path):
|
||||
logger.info(f"The clipped video {output_path} exists.")
|
||||
cur_start = cur_end
|
||||
@@ -78,7 +79,7 @@ def clip_video(video_path, timecode_list, output_folder, video_duration):
|
||||
start_time = datetime.strptime(timecode[0], "%H:%M:%S.%f")
|
||||
end_time = datetime.strptime(timecode[1], "%H:%M:%S.%f")
|
||||
video_duration = (end_time - start_time).total_seconds()
|
||||
output_path = os.path.join(output_folder, video_stem + f"_{i}.mp4")
|
||||
output_path = os.path.join(output_folder, video_name + f"_{i}.mp4")
|
||||
if os.path.exists(output_path):
|
||||
logger.info(f"The clipped video {output_path} exists.")
|
||||
continue
|
||||
@@ -94,7 +95,7 @@ def clip_video(video_path, timecode_list, output_folder, video_duration):
|
||||
if cur_video_duration < MIN_SECONDS:
|
||||
break
|
||||
splitted_timecode_list.append([cur_start.strftime("%H:%M:%S.%f")[:-3], cur_end.strftime("%H:%M:%S.%f")[:-3]])
|
||||
splitted_output_path = os.path.join(output_folder, video_stem + f"_{i}_{splitted_index}.mp4")
|
||||
splitted_output_path = os.path.join(output_folder, video_name + f"_{i}_{splitted_index}.mp4")
|
||||
if os.path.exists(splitted_output_path):
|
||||
logger.info(f"The clipped video {splitted_output_path} exists.")
|
||||
cur_start = cur_end
|
||||
@@ -146,23 +147,19 @@ if __name__ == "__main__":
|
||||
video_metadata_df = video_metadata_df[video_metadata_df["resolution"] >= args.resolution_threshold]
|
||||
logger.info(f"Filter {num_videos - len(video_metadata_df)} videos with resolution smaller than {args.resolution_threshold}.")
|
||||
video_path_list = video_metadata_df[args.video_path_column].to_list()
|
||||
video_id_list = [Path(video_path).stem for video_path in video_path_list]
|
||||
if len(video_id_list) != len(list(set(video_id_list))):
|
||||
logger.warning("Duplicate file names exist in the input video path list.")
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
video_timecode_list = video_metadata_df["timecode_list"].to_list()
|
||||
video_duration_list = video_metadata_df["duration"].to_list()
|
||||
|
||||
if args.video_folder == "":
|
||||
output_folder_list = [args.output_folder] * len(video_path_list)
|
||||
video_name_list = [Path(video_path).name for video_path in video_path_list]
|
||||
# We only check the unique video name with the absolute video path.
|
||||
if len(video_name_list) != len(set(video_name_list)):
|
||||
logger.error(f"The video path in {args.video_metadata_path} should has an unique video name.")
|
||||
else:
|
||||
output_folder_list = [os.path.join(args.output_folder, os.path.dirname(video_path)) for video_path in video_path_list]
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
|
||||
assert len(video_path_list) == len(video_timecode_list)
|
||||
os.makedirs(args.output_folder, exist_ok=True)
|
||||
args_list = [
|
||||
(video_path, timecode_list, output_folder, video_duration)
|
||||
for video_path, timecode_list, output_folder, video_duration in zip(
|
||||
video_path_list, video_timecode_list, output_folder_list, video_duration_list
|
||||
(video_path, timecode_list, args.output_folder, video_duration)
|
||||
for video_path, timecode_list, video_duration in zip(
|
||||
video_path_list, video_timecode_list, video_duration_list
|
||||
)
|
||||
]
|
||||
with Pool(args.n_jobs) as pool:
|
||||
@@ -0,0 +1,354 @@
|
||||
# Modified from https://github.com/mit-han-lab/llm-awq/blob/main/tinychat/vlm_demo_new.py.
|
||||
import argparse
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
import torch
|
||||
from accelerate import load_checkpoint_and_dispatch, PartialState
|
||||
from accelerate.utils import gather_object
|
||||
from decord import VideoReader
|
||||
from PIL import Image
|
||||
from natsort import natsorted
|
||||
from tqdm import tqdm
|
||||
from transformers import AutoConfig, AutoTokenizer
|
||||
|
||||
import tinychat.utils.constants
|
||||
# from tinychat.models.llava_llama import LlavaLlamaForCausalLM
|
||||
from tinychat.models.vila_llama import VilaLlamaForCausalLM
|
||||
from tinychat.stream_generators.llava_stream_gen import LlavaStreamGenerator
|
||||
from tinychat.utils.conversation_utils import gen_params
|
||||
from tinychat.utils.llava_image_processing import process_images
|
||||
from tinychat.utils.prompt_templates import (
|
||||
get_image_token,
|
||||
get_prompter,
|
||||
get_stop_token_ids,
|
||||
)
|
||||
from tinychat.utils.tune import (
|
||||
device_warmup,
|
||||
tune_llava_patch_embedding,
|
||||
)
|
||||
|
||||
from utils.filter import filter
|
||||
from utils.logger import logger
|
||||
|
||||
gen_params.seed = 1
|
||||
gen_params.temp = 1.0
|
||||
gen_params.top_p = 1.0
|
||||
|
||||
|
||||
def extract_uniform_frames(video_path: str, num_sampled_frames: int = 8):
|
||||
vr = VideoReader(video_path)
|
||||
sampled_frame_idx_list = np.linspace(0, len(vr), num_sampled_frames, endpoint=False, dtype=int)
|
||||
sampled_frame_list = []
|
||||
for idx in sampled_frame_idx_list:
|
||||
sampled_frame = Image.fromarray(vr[idx].asnumpy())
|
||||
sampled_frame_list.append(sampled_frame)
|
||||
|
||||
return sampled_frame_list
|
||||
|
||||
|
||||
def stream_output(output_stream):
|
||||
for outputs in output_stream:
|
||||
output_text = outputs["text"]
|
||||
output_text = output_text.strip().split(" ")
|
||||
# print(f"output_text: {output_text}.")
|
||||
return " ".join(output_text)
|
||||
|
||||
|
||||
def skip(*args, **kwargs):
|
||||
pass
|
||||
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="Recaption videos with VILA1.5.")
|
||||
parser.add_argument(
|
||||
"--video_metadata_path",
|
||||
type=str,
|
||||
default=None,
|
||||
help="The path to the video dataset metadata (csv/jsonl).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--video_path_column",
|
||||
type=str,
|
||||
default="video_path",
|
||||
help="The column contains the video path (an absolute path or a relative path w.r.t the video_folder).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--caption_column",
|
||||
type=str,
|
||||
default="caption",
|
||||
help="The column contains the caption.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--video_folder", type=str, default="", help="The video folder."
|
||||
)
|
||||
parser.add_argument("--input_prompt", type=str, default="<video>\\n Elaborate on the visual and narrative elements of the video in detail.")
|
||||
parser.add_argument(
|
||||
"--model_type", type=str, default="LLaMa", help="type of the model"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--model_path", type=str, default="Efficient-Large-Model/Llama-3-VILA1.5-8b-AWQ"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--quant_path",
|
||||
type=str,
|
||||
default=None,
|
||||
)
|
||||
parser.add_argument(
|
||||
"--precision", type=str, default="W4A16", help="compute precision"
|
||||
)
|
||||
parser.add_argument("--num_sampled_frames", type=int, default=8)
|
||||
parser.add_argument(
|
||||
"--saved_path",
|
||||
type=str,
|
||||
required=True,
|
||||
help="The save path to the output results (csv/jsonl).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--saved_freq",
|
||||
type=int,
|
||||
default=100,
|
||||
help="The frequency to save the output results.",
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--basic_metadata_path", type=str, default=None, help="The path to the basic metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_resolution", type=float, default=0, help="The resolution threshold.")
|
||||
parser.add_argument("--min_duration", type=float, default=-1, help="The minimum duration.")
|
||||
parser.add_argument("--max_duration", type=float, default=-1, help="The maximum duration.")
|
||||
parser.add_argument(
|
||||
"--asethetic_score_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_asethetic_score", type=float, default=4.0, help="The asethetic score threshold.")
|
||||
parser.add_argument(
|
||||
"--asethetic_score_siglip_metadata_path", type=str, default=None, help="The path to the video quality metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_asethetic_score_siglip", type=float, default=4.0, help="The asethetic score (SigLIP) threshold.")
|
||||
parser.add_argument(
|
||||
"--text_score_metadata_path", type=str, default=None, help="The path to the video text score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_text_score", type=float, default=0.02, help="The text threshold.")
|
||||
parser.add_argument(
|
||||
"--motion_score_metadata_path", type=str, default=None, help="The path to the video motion score metadata (csv/jsonl)."
|
||||
)
|
||||
parser.add_argument("--min_motion_score", type=float, default=2, help="The motion threshold.")
|
||||
|
||||
args = parser.parse_args()
|
||||
return args
|
||||
|
||||
|
||||
def main(args):
|
||||
if args.video_metadata_path.endswith(".csv"):
|
||||
video_metadata_df = pd.read_csv(args.video_metadata_path)
|
||||
elif args.video_metadata_path.endswith(".jsonl"):
|
||||
video_metadata_df = pd.read_json(args.video_metadata_path, lines=True)
|
||||
else:
|
||||
raise ValueError("The video_metadata_path must end with .csv or .jsonl.")
|
||||
video_path_list = video_metadata_df[args.video_path_column].tolist()
|
||||
video_path_list = [os.path.basename(video_path) for video_path in video_path_list]
|
||||
|
||||
if not (args.saved_path.endswith(".csv") or args.saved_path.endswith(".jsonl")):
|
||||
raise ValueError("The saved_path must end with .csv or .jsonl.")
|
||||
|
||||
if os.path.exists(args.saved_path):
|
||||
if args.saved_path.endswith(".csv"):
|
||||
saved_metadata_df = pd.read_csv(args.saved_path)
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
saved_metadata_df = pd.read_json(args.saved_path, lines=True)
|
||||
saved_video_path_list = saved_metadata_df[args.video_path_column].tolist()
|
||||
video_path_list = list(set(video_path_list).difference(set(saved_video_path_list)))
|
||||
logger.info(
|
||||
f"Resume from {args.saved_path}: {len(saved_video_path_list)} processed and {len(video_path_list)} to be processed."
|
||||
)
|
||||
|
||||
video_path_list = filter(
|
||||
video_path_list,
|
||||
basic_metadata_path=args.basic_metadata_path,
|
||||
min_resolution=args.min_resolution,
|
||||
min_duration=args.min_duration,
|
||||
max_duration=args.max_duration,
|
||||
asethetic_score_metadata_path=args.asethetic_score_metadata_path,
|
||||
min_asethetic_score=args.min_asethetic_score,
|
||||
asethetic_score_siglip_metadata_path=args.asethetic_score_siglip_metadata_path,
|
||||
min_asethetic_score_siglip=args.min_asethetic_score_siglip,
|
||||
text_score_metadata_path=args.text_score_metadata_path,
|
||||
min_text_score=args.min_text_score,
|
||||
motion_score_metadata_path=args.motion_score_metadata_path,
|
||||
min_motion_score=args.min_motion_score,
|
||||
)
|
||||
video_path_list = [os.path.join(args.video_folder, video_path) for video_path in video_path_list]
|
||||
# Sorting to guarantee the same result for each process.
|
||||
video_path_list = natsorted(video_path_list)
|
||||
|
||||
state = PartialState()
|
||||
|
||||
# Accelerate model initialization
|
||||
setattr(torch.nn.Linear, "reset_parameters", lambda self: None)
|
||||
setattr(torch.nn.LayerNorm, "reset_parameters", lambda self: None)
|
||||
torch.nn.init.kaiming_uniform_ = skip
|
||||
torch.nn.init.kaiming_normal_ = skip
|
||||
torch.nn.init.uniform_ = skip
|
||||
torch.nn.init.normal_ = skip
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(os.path.join(args.model_path, "llm"), use_fast=False)
|
||||
tinychat.utils.constants.LLAVA_DEFAULT_IMAGE_PATCH_TOKEN_IDX = (
|
||||
tokenizer.convert_tokens_to_ids(
|
||||
[tinychat.utils.constants.LLAVA_DEFAULT_IMAGE_PATCH_TOKEN]
|
||||
)[0]
|
||||
)
|
||||
config = AutoConfig.from_pretrained(args.model_path, trust_remote_code=True)
|
||||
model = VilaLlamaForCausalLM(config).half()
|
||||
tinychat.utils.constants.LLAVA_DEFAULT_IMAGE_PATCH_TOKEN_IDX = (
|
||||
tokenizer.convert_tokens_to_ids(
|
||||
[tinychat.utils.constants.LLAVA_DEFAULT_IMAGE_PATCH_TOKEN]
|
||||
)[0]
|
||||
)
|
||||
vision_tower = model.get_vision_tower()
|
||||
# if not vision_tower.is_loaded:
|
||||
# vision_tower.load_model()
|
||||
image_processor = vision_tower.image_processor
|
||||
# vision_tower = vision_tower.half()
|
||||
|
||||
if args.precision == "W16A16":
|
||||
pbar = tqdm(range(1))
|
||||
pbar.set_description("Loading checkpoint shards")
|
||||
for i in pbar:
|
||||
model.llm = load_checkpoint_and_dispatch(
|
||||
model.llm,
|
||||
os.path.join(args.model_path, "llm"),
|
||||
no_split_module_classes=[
|
||||
"OPTDecoderLayer",
|
||||
"LlamaDecoderLayer",
|
||||
"BloomBlock",
|
||||
"MPTBlock",
|
||||
"DecoderLayer",
|
||||
"CLIPEncoderLayer",
|
||||
],
|
||||
).to(state.device)
|
||||
model = model.to(state.device)
|
||||
|
||||
elif args.precision == "W4A16":
|
||||
from tinychat.utils.load_quant import load_awq_model
|
||||
# Auto load quant_path from the 3b/8b/13b/40b model.
|
||||
if args.quant_path is None:
|
||||
if "VILA1.5-3b-s2-AWQ" in args.model_path:
|
||||
args.quant_path = os.path.join(args.model_path, "llm/vila-1.5-3b-s2-w4-g128-awq-v2.pt")
|
||||
elif "VILA1.5-3b-AWQ" in args.model_path:
|
||||
args.quant_path = os.path.join(args.model_path, "llm/vila-1.5-3b-w4-g128-awq-v2.pt")
|
||||
elif "Llama-3-VILA1.5-8b-AWQ" in args.model_path:
|
||||
args.quant_path = os.path.join(args.model_path, "llm/llama-3-vila1.5-8b-w4-g128-awq-v2.pt")
|
||||
elif "VILA1.5-13b-AWQ" in args.model_path:
|
||||
args.quant_path = os.path.join(args.model_path, "llm/vila-1.5-13b-w4-g128-awq-v2.pt")
|
||||
elif "VILA1.5-40b-AWQ" in args.model_path:
|
||||
args.quant_path = os.path.join(args.model_path, "llm/vila-1.5-40b-w4-g128-awq-v2.pt")
|
||||
model.llm = load_awq_model(model.llm, args.quant_path, 4, 128, state.device)
|
||||
from tinychat.modules import (
|
||||
make_fused_mlp,
|
||||
make_fused_vision_attn,
|
||||
make_quant_attn,
|
||||
make_quant_norm,
|
||||
)
|
||||
|
||||
make_quant_attn(model.llm, state.device)
|
||||
make_quant_norm(model.llm)
|
||||
# make_fused_mlp(model)
|
||||
# make_fused_vision_attn(model,state.device)
|
||||
model = model.to(state.device)
|
||||
|
||||
else:
|
||||
raise NotImplementedError(f"Precision {args.precision} is not supported.")
|
||||
|
||||
device_warmup(state.device)
|
||||
tune_llava_patch_embedding(vision_tower, device=state.device)
|
||||
|
||||
stream_generator = LlavaStreamGenerator
|
||||
|
||||
model_prompter = get_prompter(
|
||||
args.model_type, args.model_path, False, False
|
||||
)
|
||||
stop_token_ids = get_stop_token_ids(args.model_type, args.model_path)
|
||||
|
||||
model.eval()
|
||||
|
||||
index = len(video_path_list) - len(video_path_list) % state.num_processes
|
||||
# Avoid the NCCL timeout in the final gather operation.
|
||||
logger.info(f"Drop {len(video_path_list) % state.num_processes} videos to ensure each process handles the same number of videos.")
|
||||
video_path_list = video_path_list[:index]
|
||||
logger.info(f"{len(video_path_list)} videos are to be processed.")
|
||||
|
||||
result_dict = {args.video_path_column: [], args.caption_column: []}
|
||||
with state.split_between_processes(video_path_list) as splitted_video_path_list:
|
||||
# TODO: Use VideoDataset.
|
||||
for i, video_path in enumerate(tqdm(splitted_video_path_list)):
|
||||
try:
|
||||
image_list = extract_uniform_frames(video_path, args.num_sampled_frames)
|
||||
image_num = len(image_list)
|
||||
# Similar operation in model_worker.py
|
||||
image_tensor = process_images(image_list, image_processor, model.config)
|
||||
if type(image_tensor) is list:
|
||||
image_tensor = [
|
||||
image.to(state.device, dtype=torch.float16) for image in image_tensor
|
||||
]
|
||||
else:
|
||||
image_tensor = image_tensor.to(state.device, dtype=torch.float16)
|
||||
|
||||
input_prompt = args.input_prompt
|
||||
# Insert image here
|
||||
image_token = get_image_token(model, args.model_path)
|
||||
image_token_holder = tinychat.utils.constants.LLAVA_DEFAULT_IM_TOKEN_PLACE_HOLDER
|
||||
im_token_count = input_prompt.count(image_token_holder)
|
||||
if im_token_count == 0:
|
||||
model_prompter.insert_prompt(image_token * image_num + input_prompt)
|
||||
else:
|
||||
assert im_token_count == image_num
|
||||
input_prompt = input_prompt.replace(image_token_holder, image_token)
|
||||
model_prompter.insert_prompt(input_prompt)
|
||||
output_stream = stream_generator(
|
||||
model,
|
||||
tokenizer,
|
||||
model_prompter.model_input,
|
||||
gen_params,
|
||||
device=state.device,
|
||||
stop_token_ids=stop_token_ids,
|
||||
image_tensor=image_tensor,
|
||||
)
|
||||
outputs = stream_output(output_stream)
|
||||
if len(outputs) != 0:
|
||||
result_dict[args.video_path_column].append(Path(video_path).name)
|
||||
result_dict[args.caption_column].append(outputs)
|
||||
|
||||
except Exception as e:
|
||||
logger.warning(f"VILA with {video_path} failed. Error is {e}.")
|
||||
|
||||
if i != 0 and i % args.saved_freq == 0:
|
||||
state.wait_for_everyone()
|
||||
gathered_result_dict = {k: gather_object(v) for k, v in result_dict.items()}
|
||||
if state.is_main_process and len(gathered_result_dict[args.video_path_column]) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_dict)
|
||||
if args.saved_path.endswith(".csv"):
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
for k in result_dict.keys():
|
||||
result_dict[k] = []
|
||||
|
||||
state.wait_for_everyone()
|
||||
gathered_result_dict = {k: gather_object(v) for k, v in result_dict.items()}
|
||||
if state.is_main_process and len(gathered_result_dict[args.video_path_column]) != 0:
|
||||
result_df = pd.DataFrame(gathered_result_dict)
|
||||
if args.saved_path.endswith(".csv"):
|
||||
header = False if os.path.exists(args.saved_path) else True
|
||||
result_df.to_csv(args.saved_path, header=header, index=False, mode="a")
|
||||
elif args.saved_path.endswith(".jsonl"):
|
||||
result_df.to_json(args.saved_path, orient="records", lines=True, mode="a", force_ascii=False)
|
||||
logger.info(f"Save result to {args.saved_path}.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
args = parse_args()
|
||||
main(args)
|
||||
@@ -1,62 +1,65 @@
|
||||
# ComfyUI VideoX-Fun
|
||||
Easily use VideoX-Fun inside ComfyUI!
|
||||
# ComfyUI CogVideoX-Fun
|
||||
Easily use CogVideoX-Fun inside ComfyUI!
|
||||
|
||||
- [Installation](#1-installation)
|
||||
- [Node types](#node-types)
|
||||
- [Example workflows](#example-workflows)
|
||||
|
||||
## Installation
|
||||
### 1. ComfyUI Installation
|
||||
## 1. Installation
|
||||
|
||||
#### Option 1: Install via ComfyUI Manager
|
||||

|
||||
### Option 1: Install via ComfyUI Manager
|
||||
TBD
|
||||
|
||||
#### Option 2: Install manually
|
||||
The VideoX-Fun repository needs to be placed at `ComfyUI/custom_nodes/VideoX-Fun/`.
|
||||
### Option 2: Install manually
|
||||
The CogVideoX-Fun repository needs to be placed at `ComfyUI/custom_nodes/CogVideoX-Fun/`.
|
||||
|
||||
```
|
||||
cd ComfyUI/custom_nodes/
|
||||
|
||||
# Git clone the cogvideox_fun itself
|
||||
git clone https://github.com/aigc-apps/VideoX-Fun.git
|
||||
git clone https://github.com/aigc-apps/CogVideoX-Fun.git
|
||||
|
||||
# Git clone the video outout node
|
||||
git clone https://github.com/Kosinkadink/ComfyUI-VideoHelperSuite.git
|
||||
|
||||
# Git clone the KJ Nodes
|
||||
git clone https://github.com/kijai/ComfyUI-KJNodes.git
|
||||
|
||||
cd VideoX-Fun/
|
||||
cd CogVideoX-Fun/
|
||||
python install.py
|
||||
```
|
||||
|
||||
### 2. Download models
|
||||
#### i、Full loading
|
||||
Download full model into `ComfyUI/models/Fun_Models/`.
|
||||
### 2. Download models into `ComfyUI/models/CogVideoX-Fun/`
|
||||
|
||||
#### ii、Chunked loading
|
||||
Put the transformer model weights to the `ComfyUI/models/diffusion_models/`.
|
||||
Put the text encoer model weights to the `ComfyUI/models/text_encoders/`.
|
||||
Put the clip vision model weights to the `ComfyUI/models/clip_vision/`.
|
||||
Put the vae model weights to the `ComfyUI/models/vae/`.
|
||||
Put the tokenizer files to the `ComfyUI/models/Fun_Models/` (For example: `ComfyUI/models/Fun_Models/umt5-xxl`).
|
||||
| Name | Storage Space | Url | Hugging Face | Description |
|
||||
|--|--|--|--|--|
|
||||
| CogVideoX-Fun-2b-InP.tar.gz | Before extraction:9.69 GB \/ After extraction: 13.0 GB | [Download](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/Diffusion_Transformer/CogVideoX-Fun-2b-InP.tar.gz) | [🤗Link](https://huggingface.co/alibaba-pai/CogVideoX-Fun-2b-InP)| Our official graph-generated video model is capable of predicting videos at multiple resolutions (512, 768, 1024, 1280) and has been trained on 144 frames at a rate of 24 frames per second. |
|
||||
|
||||
### 3. (Optional) Download preprocess weights into `ComfyUI/custom_nodes/Fun_Models/Third_Party/`.
|
||||
Except for the fun models' weights, if you want to use the control preprocess nodes, you can download the preprocess weights to `ComfyUI/custom_nodes/Fun_Models/Third_Party/`.
|
||||
## Node types
|
||||
- **LoadCogVideoX_Fun_Model**
|
||||
- Loads the CogVideoX-Fun model
|
||||
- **TextBox**
|
||||
- Write the prompt for CogVideoX-Fun model
|
||||
- **CogVideoX_Fun_I2VSampler**
|
||||
- CogVideoX-Fun Sampler for Image to Video
|
||||
- **CogVideoX_Fun_T2VSampler**
|
||||
- CogVideoX-Fun Sampler for Text to Video
|
||||
- **CogVideoX_Fun_V2VSampler**
|
||||
- CogVideoX-Fun Sampler for Video to Video
|
||||
|
||||
```
|
||||
remote_onnx_det = "https://huggingface.co/yzd-v/DWPose/resolve/main/yolox_l.onnx"
|
||||
remote_onnx_pose = "https://huggingface.co/yzd-v/DWPose/resolve/main/dw-ll_ucoco_384.onnx"
|
||||
remote_zoe= "https://huggingface.co/lllyasviel/Annotators/resolve/main/ZoeD_M12_N.pt"
|
||||
```
|
||||
## Example workflows
|
||||
|
||||
## Support models
|
||||
### Video to video generation
|
||||
Our ui is shown as follow, this is the [download link](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_v2v.json) of the json:
|
||||

|
||||
|
||||
- [CogVideox-Fun](cogvideox_fun/README.md)
|
||||
- [Qwen-Image](qwenimage/README.md)
|
||||
- [Wan2.1](wan2_1/README.md)
|
||||
- [Wan2.2](wan2_2/README.md)
|
||||
- [Wan2.1-Fun](wan2_1_fun/README.md)
|
||||
- [Wan2.2-Fun](wan2_2_fun/README.md)
|
||||
- [Wan2.2-VACE-Fun](wan2_2_fun/README.md)
|
||||
- [Z-Image](z_image/README.md)
|
||||
You can run the demo using following video:
|
||||
[demo video](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/play_guitar.mp4)
|
||||
|
||||
### Image to video generation
|
||||
Our ui is shown as follow, this is the [download link](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_i2v.json) of the json:
|
||||

|
||||
|
||||
You can run the demo using following photo:
|
||||

|
||||
|
||||
### Text to video generation
|
||||
Our ui is shown as follow, this is the [download link](https://pai-aigc-photog.oss-cn-hangzhou.aliyuncs.com/cogvideox_fun/asset/v1/cogvideoxfunv1_workflow_t2v.json) of the json:
|
||||

|
||||
@@ -1,46 +0,0 @@
|
||||
# This folder is modified from the https://github.com/Mikubill/sd-webui-controlnet
|
||||
# Openpose
|
||||
# Original from CMU https://github.com/CMU-Perceptual-Computing-Lab/openpose
|
||||
# 2nd Edited by https://github.com/Hzzone/pytorch-openpose
|
||||
# 3rd Edited by ControlNet
|
||||
# 4th Edited by ControlNet (added face and correct hands)
|
||||
|
||||
import os
|
||||
|
||||
os.environ["KMP_DUPLICATE_LIB_OK"]="TRUE"
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from . import util
|
||||
from .wholebody import Wholebody
|
||||
|
||||
|
||||
def draw_pose(poses, H, W):
|
||||
canvas = np.zeros(shape=(H, W, 3), dtype=np.uint8)
|
||||
|
||||
for pose in poses:
|
||||
canvas = util.draw_bodypose(canvas, pose.body.keypoints)
|
||||
|
||||
canvas = util.draw_handpose(canvas, pose.left_hand)
|
||||
canvas = util.draw_handpose(canvas, pose.right_hand)
|
||||
|
||||
canvas = util.draw_facepose(canvas, pose.face)
|
||||
return canvas
|
||||
|
||||
|
||||
class DWposeDetector:
|
||||
def __init__(self, onnx_det, onnx_pose):
|
||||
self.pose_estimation = Wholebody(onnx_det, onnx_pose)
|
||||
|
||||
def __call__(self, oriImg):
|
||||
oriImg = oriImg.copy()
|
||||
H, W, C = oriImg.shape
|
||||
with torch.no_grad():
|
||||
keypoints_info = self.pose_estimation(oriImg)
|
||||
return draw_pose(
|
||||
Wholebody.format_result(keypoints_info),
|
||||
H,
|
||||
W,
|
||||
)
|
||||
|
||||