Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
cec2bc5ff6 |
@@ -1,94 +0,0 @@
|
||||
# Agent Infrastructure — Status Dashboard
|
||||
|
||||
Developer-maintained overview of all agent components and their maturity.
|
||||
Use this to understand what exists, how complete it is, and how much to trust it.
|
||||
|
||||
_Last synced: 2026-03-02_
|
||||
|
||||
> To resync this dashboard, use the workflow: `.agents/workflows/sync-dashboard.md`
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Category | Total | ✅ Ready | 🟡 Draft | 🔴 Stub | Trust |
|
||||
|----------|-------|---------|---------|---------|-------|
|
||||
| Skills | 8 | 0 | 8 | 0 | Low — newly created, untested |
|
||||
| Workflows (SOPs) | 4 | 0 | 4 | 0 | Low — newly created, untested |
|
||||
| Memory files | 4 | 1 | 3 | 0 | Medium — codebase_map is solid |
|
||||
| Lessons | 0 | — | — | — | N/A — empty |
|
||||
| Exploration logs | 0 | — | — | — | N/A — empty |
|
||||
|
||||
---
|
||||
|
||||
## Skills (`.agents/skills/`)
|
||||
|
||||
| Skill | File | Status | Trust | Tested | Notes |
|
||||
|-------|------|--------|-------|--------|-------|
|
||||
| Launch Experiment | `launch-experiment.md` | 🟡 Draft | Low | ❌ | Needs dry-run validation |
|
||||
| Monitor Experiment | `monitor-experiment.md` | 🟡 Draft | Low | ❌ | Requires W&B API access to test |
|
||||
| Summarize Run | `summarize-run.md` | 🟡 Draft | Low | ❌ | Pattern from existing test infra |
|
||||
| Log Experiment | `log-experiment.md` | 🟡 Draft | Low | ❌ | Journal formatting only |
|
||||
| Evaluate Video Quality | `evaluate-video-quality.md` | 🟡 Draft | Low | ❌ | SSIM section most mature |
|
||||
| Index Related Work | `index-related-work.md` | 🟡 Draft | Low | ❌ | Schema defined, no entries yet |
|
||||
| Search Related Work | `search-related-work.md` | 🟡 Draft | Low | ❌ | Depends on indexed entries |
|
||||
| Skill Template | `SKILL_TEMPLATE.md` | ✅ Ready | High | ✅ | Meta-template, stable |
|
||||
|
||||
### Trust Level Definitions
|
||||
- **High**: Tested in production, validated against real experiments
|
||||
- **Medium**: Logic is sound, partially tested or based on existing patterns
|
||||
- **Low**: Newly written, not yet validated
|
||||
- **None**: Placeholder only
|
||||
|
||||
---
|
||||
|
||||
## Workflows / SOPs (`.agents/workflows/`)
|
||||
|
||||
| Workflow | File | Status | Trust | Tested | Notes |
|
||||
|----------|------|--------|-------|--------|-------|
|
||||
| Experiment Lifecycle | `experiment-lifecycle.md` | 🟡 Draft | Low | ❌ | End-to-end flow, untested |
|
||||
| Evaluation Development | `evaluation-development.md` | 🟡 Draft | Low | ❌ | Metric dev process |
|
||||
| Experiment Journaling | `experiment-journaling.md` | 🟡 Draft | Low | ❌ | Journaling cadence |
|
||||
| Lesson Capture | `lesson-capture.md` | 🟡 Draft | Low | ❌ | Post-experiment reflection |
|
||||
| Sync Dashboard | `sync-dashboard.md` | 🟡 Draft | Low | ❌ | This dashboard's updater |
|
||||
|
||||
---
|
||||
|
||||
## Memory (`.agents/memory/`)
|
||||
|
||||
| File | Status | Trust | Notes |
|
||||
|------|--------|-------|-------|
|
||||
| `codebase_map.md` | ✅ Ready | High | Synthesized from full repo research |
|
||||
| `experiment_journal.md` | 🟡 Draft | Medium | Schema defined, no entries yet |
|
||||
| `evaluation_registry.md` | 🟡 Draft | Medium | SSIM/loss metrics documented |
|
||||
| `related_work/README.md` | 🟡 Draft | Medium | Schema defined, no entries yet |
|
||||
|
||||
---
|
||||
|
||||
## Lessons (`.agents/lessons/`)
|
||||
|
||||
| File | Category | Severity | Notes |
|
||||
|------|----------|----------|-------|
|
||||
|
||||
_No lessons captured yet._
|
||||
|
||||
---
|
||||
|
||||
## Exploration Logs (`.agents/exploration/`)
|
||||
|
||||
| File | Status | Topic | Notes |
|
||||
|------|--------|-------|-------|
|
||||
|
||||
_No exploration logs yet._
|
||||
|
||||
---
|
||||
|
||||
## What to Do Next
|
||||
|
||||
1. **Validate skills**: Run a minimal training experiment using the
|
||||
`experiment-lifecycle` SOP to test `launch-experiment` → `monitor-experiment`
|
||||
→ `summarize-run` end-to-end.
|
||||
2. **Index first related work**: Use `index-related-work` to add at least one
|
||||
paper (e.g., the Self-Forcing paper used in the codebase).
|
||||
3. **Capture first lesson**: After the validation run, capture any findings.
|
||||
4. **Promote to Ready**: As each skill/SOP is tested, update its status here.
|
||||
@@ -1,46 +0,0 @@
|
||||
# Exploration Logs
|
||||
|
||||
This directory holds draft procedures and investigation notes for tasks that
|
||||
don't yet have a standardized skill or SOP. Each exploration should follow this
|
||||
template.
|
||||
|
||||
## When to Create an Exploration Log
|
||||
|
||||
- You are working on a task with no existing skill or workflow.
|
||||
- You are experimenting with a new metric, training technique, or tool.
|
||||
- You want to document findings before they are promoted to a standard.
|
||||
|
||||
## File Naming
|
||||
|
||||
`<topic-slug>.md` — e.g., `fvd-metric-investigation.md`
|
||||
|
||||
## Template
|
||||
|
||||
```markdown
|
||||
# Exploration Log: <Topic>
|
||||
|
||||
## Status: draft | under_review | promoted | abandoned
|
||||
|
||||
## Context
|
||||
<Why this exploration is needed — link to experiment or task if applicable.>
|
||||
|
||||
## Progress
|
||||
- [ ] Step 1: ...
|
||||
- [ ] Step 2: ...
|
||||
|
||||
## Findings
|
||||
<What you have learned so far.>
|
||||
|
||||
## Mistakes / Dead Ends
|
||||
<What didn't work and why — these become lessons.>
|
||||
|
||||
## Proposed Standardization
|
||||
<If this works, describe the skill/SOP/workflow to create.>
|
||||
```
|
||||
|
||||
## Lifecycle
|
||||
|
||||
1. **Create** during exploration mode.
|
||||
2. **Update** as you make progress.
|
||||
3. **Promote**: If findings are solid, create a skill in `.agents/skills/` or an SOP in `.agents/workflows/`.
|
||||
4. **Archive mistakes**: Move failures into `.agents/lessons/`.
|
||||
@@ -1,48 +0,0 @@
|
||||
# Lessons Learned Database
|
||||
|
||||
This directory stores documented mistakes, unexpected behaviors, and their fixes.
|
||||
Each lesson is a permanent record that helps agents and humans avoid repeating
|
||||
past errors.
|
||||
|
||||
## When to Create a Lesson
|
||||
|
||||
- An experiment failed for a non-obvious reason.
|
||||
- A configuration or hyperparameter choice led to wasted compute.
|
||||
- A porting, data, or infrastructure issue was discovered and resolved.
|
||||
- A workaround was needed for a known framework/library bug.
|
||||
|
||||
## File Naming
|
||||
|
||||
`<YYYY-MM-DD>_<short-slug>.md` — e.g., `2026-03-02_lr-too-high-for-lora.md`
|
||||
|
||||
## Template
|
||||
|
||||
```markdown
|
||||
---
|
||||
date: <ISO-8601>
|
||||
experiment: <reference to experiment_journal.md entry, if applicable>
|
||||
category: hyperparameter | data | infrastructure | evaluation | porting | other
|
||||
severity: critical | important | minor
|
||||
---
|
||||
|
||||
# <Short Descriptive Title>
|
||||
|
||||
## What Happened
|
||||
<Description of the problem and its symptoms.>
|
||||
|
||||
## Root Cause
|
||||
<Analysis of why it happened.>
|
||||
|
||||
## Fix / Workaround
|
||||
<What resolved the issue.>
|
||||
|
||||
## Prevention
|
||||
<How to avoid this in the future — updated skills, SOPs, or checks.>
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
- Before starting a task, **search this directory** for relevant lessons.
|
||||
- After completing or failing a task, **check if a new lesson should be created**.
|
||||
- Periodically review lessons for **patterns** — recurring themes may warrant
|
||||
a new skill, SOP, or codebase fix.
|
||||
@@ -1,129 +0,0 @@
|
||||
# FastVideo-WorldModel — Codebase Map
|
||||
|
||||
High-level structural index for agent orientation. Updated 2026-03-08.
|
||||
|
||||
## Repository Layout
|
||||
|
||||
```
|
||||
FastVideo-WorldModel/
|
||||
├── fastvideo/ # Core Python package
|
||||
│ ├── models/ # Model implementations
|
||||
│ │ ├── dits/ # DiT transformers (wanvideo, ltx2, ...)
|
||||
│ │ ├── vaes/ # VAE models
|
||||
│ │ ├── encoders/ # Text/image encoders (T5, CLIP)
|
||||
│ │ ├── schedulers/ # Noise schedulers
|
||||
│ │ ├── upsamplers/ # Super-resolution models
|
||||
│ │ ├── audio/ # Audio models
|
||||
│ │ └── loader/ # Component loaders for HF repos
|
||||
│ ├── configs/ # Configuration system
|
||||
│ │ ├── models/ # Arch configs + param_names_mapping
|
||||
│ │ ├── pipelines/ # Pipeline wiring
|
||||
│ │ └── sample/ # Default sampling parameters
|
||||
│ ├── pipelines/ # End-to-end pipelines
|
||||
│ │ ├── basic/ # Per-model pipelines (wan/, ltx2/, ...)
|
||||
│ │ └── stages/ # Reusable pipeline stages
|
||||
│ ├── train/ # Refactored training framework (YAML-driven, preferred)
|
||||
│ │ ├── trainer.py # Main training loop coordinator
|
||||
│ │ ├── entrypoint/ # Training entrypoint (train.py) + checkpoint conversion
|
||||
│ │ ├── methods/ # Training algorithms (FineTune, DFSFT, DMD2, SelfForcing)
|
||||
│ │ │ ├── base.py # TrainingMethod ABC
|
||||
│ │ │ ├── fine_tuning/ # FineTuneMethod, DiffusionForcingSFTMethod
|
||||
│ │ │ └── distribution_matching/ # DMD2Method, SelfForcingMethod
|
||||
│ │ ├── models/ # Per-role model wrappers (ModelBase, CausalModelBase)
|
||||
│ │ │ └── wan/ # WanModel, WanCausalModel
|
||||
│ │ ├── callbacks/ # Composable hooks (grad_clip, ema, validation)
|
||||
│ │ └── utils/ # Config, builder, checkpoint, optimizer, tracking
|
||||
│ ├── training/ # Legacy training infrastructure (being phased out)
|
||||
│ │ ├── trackers.py # W&B tracker (BaseTracker → WandbTracker)
|
||||
│ │ ├── training_utils.py # Checkpointing, grad clipping, state dicts
|
||||
│ │ ├── training_pipeline.py # Base training pipeline
|
||||
│ │ ├── wan_training_pipeline.py # Wan T2V training
|
||||
│ │ ├── wan_i2v_training_pipeline.py # Wan I2V training
|
||||
│ │ ├── distillation_pipeline.py # Distillation base
|
||||
│ │ ├── wan_distillation_pipeline.py # Wan distillation
|
||||
│ │ ├── self_forcing_distillation_pipeline.py # Self-forcing distill
|
||||
│ │ ├── ltx2_training_pipeline.py # LTX-2 training
|
||||
│ │ └── matrixgame_training_pipeline.py # MatrixGame training
|
||||
│ ├── attention/ # Attention backends
|
||||
│ ├── distributed/ # Sequence/tensor parallel utilities
|
||||
│ ├── layers/ # Tensor-parallel layers
|
||||
│ ├── tests/ # Package-level tests
|
||||
│ │ ├── training/ # Training regression tests (W&B summary comparison)
|
||||
│ │ ├── ssim/ # SSIM visual regression tests
|
||||
│ │ ├── encoders/ # Encoder parity tests
|
||||
│ │ └── modal/ # Modal CI test runner
|
||||
│ └── registry.py # Unified config registry
|
||||
├── fastvideo-kernel/ # CUDA/custom kernels (separate build: ./build.sh)
|
||||
├── scripts/ # Utility scripts
|
||||
│ ├── distill/ # Distillation launch scripts
|
||||
│ ├── inference/ # Inference scripts
|
||||
│ ├── checkpoint_conversion/ # Weight conversion tools
|
||||
│ ├── finetune/ # Finetune scripts
|
||||
│ └── preprocess/ # Data preprocessing
|
||||
├── examples/ # Ready-to-run examples
|
||||
│ ├── training/ # Training examples (finetune/, consistency_finetune/)
|
||||
│ ├── distill/ # Distillation examples
|
||||
│ ├── inference/ # Inference examples
|
||||
│ └── dataset/ # Dataset examples
|
||||
├── docs/ # MkDocs documentation source
|
||||
│ ├── design/overview.md # Architecture overview
|
||||
│ ├── training/ # Training guides
|
||||
│ └── contributing/ # Contributor guides + coding_agents.md
|
||||
├── tests/ # Top-level tests (local_tests/)
|
||||
├── AGENTS.md # Agent coding guidelines
|
||||
└── .agents/ # Agent infrastructure (you are here)
|
||||
```
|
||||
|
||||
## Key Training Entrypoints
|
||||
|
||||
### New framework (`fastvideo/train/`) — preferred
|
||||
|
||||
| Method | Config Example | Launch Pattern |
|
||||
|--------|---------------|----------------|
|
||||
| FineTune (Wan) | `examples/train/finetune_wan2.1_t2v_1.3B_vsa_*.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
|
||||
| DFSFT (Wan causal) | `examples/train/dfsft_wan_causal_t2v_1.3B.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
|
||||
| DMD2 distillation | `examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
|
||||
| Self-Forcing | `examples/train/self_forcing_wan_causal_t2v_1.3B.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
|
||||
|
||||
### Legacy pipelines (`fastvideo/training/`) — being phased out
|
||||
|
||||
| Pipeline | Entrypoint | Launch Pattern |
|
||||
|----------|-----------|----------------|
|
||||
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| Wan distillation (DMD) | `fastvideo/training/wan_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| Self-forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
| MatrixGame | `fastvideo/training/matrixgame_training_pipeline.py` | `torchrun --nproc_per_node N` |
|
||||
|
||||
## W&B Integration
|
||||
|
||||
- **Tracker classes**: `fastvideo/training/trackers.py`
|
||||
- `WandbTracker` — logs metrics, videos, timing
|
||||
- `SequentialTracker` — fan-out to multiple trackers
|
||||
- `DummyTracker` — no-op for offline/test
|
||||
- **Run summary location**: `<output_dir>/tracker/wandb/latest-run/files/wandb-summary.json`
|
||||
- **Reference summaries**: `fastvideo/tests/training/*/` (e.g., `a40_reference_wandb_summary.json`)
|
||||
- **Environment**: `WANDB_API_KEY`, `WANDB_BASE_URL`, `WANDB_MODE`
|
||||
|
||||
## Critical Environment Variables
|
||||
|
||||
| Variable | Purpose |
|
||||
|----------|---------|
|
||||
| `WANDB_API_KEY` | W&B authentication |
|
||||
| `WANDB_MODE` | `online` / `offline` |
|
||||
| `FASTVIDEO_ATTENTION_BACKEND` | `FLASH_ATTN` / `TORCH_SDPA` |
|
||||
| `TOKENIZERS_PARALLELISM` | Set `false` to avoid fork warnings |
|
||||
| `HF_HOME` | HuggingFace cache directory |
|
||||
|
||||
## Build & Test Commands
|
||||
|
||||
```bash
|
||||
uv pip install -e .[dev] # Editable install
|
||||
pre-commit run --all-files # Lint/format/spell
|
||||
pytest tests/ # Top-level tests
|
||||
pytest fastvideo/tests/ -v # Package tests
|
||||
pytest fastvideo/tests/training/Vanilla -srP # Training loss regression
|
||||
pytest fastvideo/tests/ssim/ -vs # SSIM visual regression
|
||||
cd fastvideo-kernel && ./build.sh # Build kernels
|
||||
```
|
||||
@@ -1,327 +0,0 @@
|
||||
# Evaluation Metrics Registry
|
||||
|
||||
Living catalog of all evaluation metrics for FastVideo-WorldModel video quality
|
||||
assessment. Each metric includes a detailed explanation, implementation status,
|
||||
usage instructions, and interpretation guide.
|
||||
|
||||
_Last updated: 2026-03-02_
|
||||
|
||||
---
|
||||
|
||||
## Metric Summary
|
||||
|
||||
| Metric | Category | Status | Location | Trust |
|
||||
|--------|----------|--------|----------|-------|
|
||||
| **FVD** | Distribution | ✅ Implemented | `benchmarks/fvd/` | High |
|
||||
| **SSIM** | Reference | ✅ Implemented | `fastvideo/tests/ssim/` | High |
|
||||
| **LPIPS** | Perceptual | ✅ Implemented | `scripts/lora_extraction/` | Medium |
|
||||
| **Loss trajectory** | Training signal | ✅ Implemented | W&B `train_loss` | Medium |
|
||||
| **Grad norm stability** | Training signal | ✅ Implemented | W&B `grad_norm` | Medium |
|
||||
| **GameWorld Score** | Multi-dim benchmark | 🟡 External | Matrix-Game repo | Low |
|
||||
| **Human preference** | Gold standard | 🔴 Manual | N/A | Highest |
|
||||
|
||||
---
|
||||
|
||||
## Implemented Metrics
|
||||
|
||||
### FVD — Fréchet Video Distance
|
||||
|
||||
**Category**: Distribution-level quality metric
|
||||
**Status**: ✅ Fully implemented in `benchmarks/fvd/`
|
||||
**Trust**: High — standard protocol, I3D feature extractor
|
||||
|
||||
#### What It Measures
|
||||
FVD measures the distance between the **distribution** of generated videos and
|
||||
a distribution of real/reference videos. It works by:
|
||||
1. Extracting spatiotemporal features from both real and generated video sets
|
||||
using a pretrained **I3D** (Inflated 3D ConvNet) model.
|
||||
2. Modeling each set of features as a multivariate Gaussian (mean + covariance).
|
||||
3. Computing the **Fréchet distance** between the two Gaussians.
|
||||
|
||||
Lower FVD = generated videos are more statistically similar to real videos.
|
||||
|
||||
#### Why It Matters
|
||||
- FVD is the **de facto standard** for benchmarking video generation models.
|
||||
- It captures both **visual quality** (are individual frames realistic?) and
|
||||
**temporal coherence** (do frames flow naturally?).
|
||||
- Matrix-Game 2.0, Open-Sora, and most video generation papers report FVD.
|
||||
|
||||
#### Limitations
|
||||
- Requires a **large sample set** (standard protocol uses 2048 videos) to
|
||||
produce stable statistics. Small sample sizes yield noisy results.
|
||||
- Measures **distributional similarity**, not per-video quality. A model could
|
||||
have low FVD by generating a diverse set of "roughly okay" videos.
|
||||
- The I3D model was trained on Kinetics-400 (human actions). It may be less
|
||||
sensitive to domain-specific artifacts in non-human-action videos (e.g.,
|
||||
driving, game environments).
|
||||
- Does not directly measure text-video alignment or action controllability.
|
||||
|
||||
#### How to Use
|
||||
|
||||
```python
|
||||
# Programmatic
|
||||
from benchmarks.fvd import compute_fvd_with_config, FVDConfig
|
||||
|
||||
config = FVDConfig.fvd2048_16f() # Standard: 2048 videos, 16 frames
|
||||
results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
|
||||
print(f"FVD: {results['fvd']:.2f}")
|
||||
```
|
||||
|
||||
```bash
|
||||
# CLI
|
||||
python -m benchmarks.fvd.cli \
|
||||
--real-path data/real/ \
|
||||
--gen-path outputs/gen/ \
|
||||
--protocol fvd2048_16f
|
||||
```
|
||||
|
||||
**Preset protocols**:
|
||||
| Protocol | Videos | Frames | Use Case |
|
||||
|----------|--------|--------|----------|
|
||||
| `fvd2048_16f` | 2048 | 16 | Standard benchmark (papers) |
|
||||
| `fvd2048_128f` | 2048 | 128 | Long video evaluation |
|
||||
| `quick_test` | 100 | 16 | Fast dev iteration |
|
||||
|
||||
**Feature extractors**: `i3d` (default, standard), `clip`, `videomae`
|
||||
|
||||
#### Interpretation
|
||||
| FVD Range | Interpretation |
|
||||
|-----------|---------------|
|
||||
| < 100 | Excellent — near-real quality |
|
||||
| 100–300 | Good — competitive with SOTA |
|
||||
| 300–600 | Fair — noticeable gap from real |
|
||||
| > 600 | Poor — significant quality issues |
|
||||
|
||||
> FVD values are dataset-dependent. Always compare against baselines evaluated
|
||||
> on the same real video distribution.
|
||||
|
||||
---
|
||||
|
||||
### SSIM — Structural Similarity Index
|
||||
|
||||
**Category**: Per-frame reference comparison
|
||||
**Status**: ✅ Implemented in `fastvideo/tests/ssim/`
|
||||
**Trust**: High — used in CI regression tests
|
||||
|
||||
#### What It Measures
|
||||
SSIM compares two images (or video frames) based on three components:
|
||||
1. **Luminance**: brightness similarity
|
||||
2. **Contrast**: dynamic range similarity
|
||||
3. **Structure**: spatial pattern similarity
|
||||
|
||||
The final score is a value in [0, 1] where 1.0 = identical.
|
||||
|
||||
#### Why It Matters
|
||||
- Used as a **regression guard** in CI: ensures model updates don't degrade
|
||||
visual output below a threshold.
|
||||
- More perceptually meaningful than raw pixel MSE.
|
||||
- Fast to compute — suitable for automated testing.
|
||||
|
||||
#### Limitations
|
||||
- Requires a **pixel-aligned reference** video. Cannot compare videos with
|
||||
different seeds, prompts, or angles.
|
||||
- Operates **per-frame** — does not capture temporal coherence.
|
||||
- Insensitive to some perceptual artifacts (color shifts, high-frequency noise).
|
||||
|
||||
#### How to Use
|
||||
|
||||
```bash
|
||||
pytest fastvideo/tests/ssim/ -vs
|
||||
```
|
||||
|
||||
#### Interpretation
|
||||
| SSIM Range | Quality |
|
||||
|------------|---------|
|
||||
| > 0.90 | Excellent — very close to reference |
|
||||
| 0.80–0.90 | Good — acceptable for most uses |
|
||||
| 0.70–0.80 | Fair — noticeable differences |
|
||||
| < 0.70 | Poor — significant divergence |
|
||||
|
||||
---
|
||||
|
||||
### LPIPS — Learned Perceptual Image Patch Similarity
|
||||
|
||||
**Category**: Per-frame perceptual distance
|
||||
**Status**: ✅ Implemented in `scripts/lora_extraction/lora_inference_comparison.py`
|
||||
**Trust**: Medium — available but only used for LoRA comparison currently
|
||||
|
||||
#### What It Measures
|
||||
LPIPS uses a pretrained neural network (AlexNet by default) to extract
|
||||
deep features from two images and computes the distance between them in
|
||||
feature space. Unlike SSIM, LPIPS correlates much more strongly with
|
||||
**human perceptual judgments**.
|
||||
|
||||
Lower LPIPS = more perceptually similar.
|
||||
|
||||
#### Why It Matters
|
||||
- Best available automated proxy for **human visual judgments** at the frame
|
||||
level.
|
||||
- Captures semantic and structural differences that SSIM misses (e.g., texture
|
||||
changes, minor recoloring).
|
||||
- Used for validating LoRA merge quality.
|
||||
|
||||
#### Limitations
|
||||
- Per-frame metric — no temporal awareness.
|
||||
- Requires reference video (paired comparison only).
|
||||
- Slightly slower than SSIM due to neural network forward pass.
|
||||
|
||||
#### How to Use
|
||||
|
||||
```bash
|
||||
python scripts/lora_extraction/lora_inference_comparison.py \
|
||||
--base merged_model \
|
||||
--ft path/to/finetuned \
|
||||
--adapter NONE \
|
||||
--output-dir results \
|
||||
--prompt "A cat" \
|
||||
--compute-lpips
|
||||
```
|
||||
|
||||
#### Interpretation
|
||||
| LPIPS Range | Quality |
|
||||
|-------------|---------|
|
||||
| < 0.10 | Excellent — nearly indistinguishable |
|
||||
| 0.10–0.20 | Good — minor perceptual differences |
|
||||
| 0.20–0.40 | Fair — noticeable differences |
|
||||
| > 0.40 | Poor — clearly different |
|
||||
|
||||
---
|
||||
|
||||
### Loss Trajectory
|
||||
|
||||
**Category**: Training signal proxy
|
||||
**Status**: ✅ Active (from W&B `train_loss`)
|
||||
**Trust**: Medium — proxy, not direct quality measure
|
||||
|
||||
#### What It Measures
|
||||
Tracks the training loss over time. A healthy training run shows:
|
||||
- **Decreasing loss** over the first hundreds of steps.
|
||||
- **Stable gradient norms** (no wild spikes).
|
||||
- **Consistent step times** (no infrastructure issues).
|
||||
|
||||
#### Why It Matters
|
||||
- Cheapest evaluation signal — available in real-time from W&B.
|
||||
- Critical for the **30-minute quality check** workflow.
|
||||
- At later training stages (when loss becomes meaningful), trajectory shape
|
||||
can predict final model quality.
|
||||
|
||||
#### Context: How This Evolves
|
||||
The team's experience shows evaluation signals change during a project:
|
||||
- **Early stage**: Loss may be flat or meaningless → focus on SSIM & visual
|
||||
inspection instead.
|
||||
- **Mid stage**: Loss starts decreasing → trajectory shape becomes useful.
|
||||
- **Late stage**: Loss is meaningful → can compare trajectories across runs.
|
||||
|
||||
This dynamic is a key insight from the team's workflow: don't over-rely on
|
||||
loss early; don't ignore it late.
|
||||
|
||||
---
|
||||
|
||||
### Grad Norm Stability
|
||||
|
||||
**Category**: Training health diagnostic
|
||||
**Status**: ✅ Active (from W&B `grad_norm`)
|
||||
**Trust**: Medium — diagnostic, not quality metric
|
||||
|
||||
#### What It Measures
|
||||
The magnitude of gradients during training. Stable grad norms indicate
|
||||
healthy optimization. Spikes or NaN values indicate training instability.
|
||||
|
||||
#### Alert Thresholds
|
||||
| Condition | Meaning |
|
||||
|-----------|---------|
|
||||
| Stable ~0.3–0.5 | Normal training |
|
||||
| Single spike > 3× average | Possible bad batch, monitor |
|
||||
| NaN or Inf | 🔴 Training has diverged — stop run |
|
||||
| Increasing trend | Learning rate may be too high |
|
||||
|
||||
---
|
||||
|
||||
## External Benchmarks
|
||||
|
||||
### GameWorld Score Benchmark (Matrix-Game)
|
||||
|
||||
**Category**: Multi-dimensional evaluation framework for interactive world models
|
||||
**Status**: 🟡 External — not implemented in-repo
|
||||
**Source**: [Matrix-Game 1.0 benchmark](https://github.com/SkyworkAI/Matrix-Game), used in [Matrix-Game 2.0 paper](https://arxiv.org/abs/2508.13009)
|
||||
|
||||
#### What It Measures
|
||||
A comprehensive benchmark examining **four critical capabilities**:
|
||||
|
||||
| Dimension | What It Evaluates | Example Signals |
|
||||
|-----------|-------------------|-----------------|
|
||||
| **Visual quality** | Frame-level realism, absence of artifacts | Color fidelity, sharpness, coherence |
|
||||
| **Temporal quality** | Smoothness across frames, motion consistency | Jitter, flickering, temporal aliasing |
|
||||
| **Action controllability** | Response to input actions (keyboard/mouse) | Action delay, correctness, smoothness |
|
||||
| **Physical rule understanding** | Adherence to physics (gravity, collision) | Object persistence, plausible motion |
|
||||
|
||||
#### Context from Matrix-Game 2.0
|
||||
- Evaluation uses **597-frame composite action sequences** over 32 Minecraft
|
||||
scenes and 16 wild scenes.
|
||||
- Action controllability assessment is **Minecraft-specific** — cannot be
|
||||
directly applied to wild/general scenes.
|
||||
- The paper notes that models that "collapse" to static frames can
|
||||
paradoxically score higher on consistency metrics — beware of this confound.
|
||||
|
||||
#### Relevance to FastVideo
|
||||
- Matrix-Game 2.0 is built on SkyReels-V2/Wan2.1 architecture — **same model
|
||||
family as FastVideo**.
|
||||
- Their distillation uses DMD-based Self-Forcing — **same technique** as our
|
||||
`self_forcing_distillation_pipeline.py`.
|
||||
- GameWorld Score dimensions are a useful framework for thinking about world
|
||||
model quality even outside gaming contexts.
|
||||
|
||||
---
|
||||
|
||||
## Human Preference Evaluation
|
||||
|
||||
**Category**: Gold-standard quality assessment
|
||||
**Status**: 🔴 Manual process — no automated implementation
|
||||
**Priority**: **Highest** — this is the most important evaluation signal
|
||||
**Trust**: Highest — but expensive
|
||||
|
||||
### What It Measures
|
||||
Human evaluators compare generated videos and rate them on dimensions like:
|
||||
- Overall quality and realism
|
||||
- Temporal coherence and smoothness
|
||||
- Prompt adherence / action correctness
|
||||
- Absence of artifacts
|
||||
|
||||
#### Why It's the Most Important Metric
|
||||
All automated metrics are **proxies** for human judgment. They can be gamed
|
||||
or may miss artifacts that humans easily notice. Human preference is the
|
||||
ultimate ground truth for video generation quality.
|
||||
|
||||
#### Cost & Practicality
|
||||
| Approach | Cost | Scale | When to Use |
|
||||
|----------|------|-------|-------------|
|
||||
| Internal team review | Low | ~10–50 videos | Every major checkpoint |
|
||||
| Crowdsource (MTurk, Scale) | Medium | 100+ videos | Pre-release validation |
|
||||
| A/B preference test | Medium | Pairs | Comparing two model versions |
|
||||
|
||||
#### Recommended Protocol
|
||||
1. Sample 10–20 videos from the model at a checkpoint.
|
||||
2. Include diverse prompts (easy + hard, short + long).
|
||||
3. Have 2–3 evaluators score each video 1–5 on: quality, coherence, fidelity.
|
||||
4. Record scores in the experiment journal.
|
||||
|
||||
---
|
||||
|
||||
## Metrics NOT Used
|
||||
|
||||
| Metric | Reason |
|
||||
|--------|--------|
|
||||
| ~~CLIP-Score~~ | Not used by the team. Measures text-image alignment using CLIP embeddings, but not well-suited for video temporal quality. |
|
||||
| Inception Score (IS) | Less informative than FVD for video; primarily an image metric. |
|
||||
| PSNR | Pixel-level metric; less perceptually meaningful than SSIM/LPIPS. |
|
||||
|
||||
---
|
||||
|
||||
## Adding a New Metric
|
||||
|
||||
Follow the SOP: `.agents/workflows/evaluation-development.md`
|
||||
|
||||
1. Prototype in `.agents/exploration/`
|
||||
2. Validate on known-good and known-bad samples
|
||||
3. Add to this registry
|
||||
4. Update the `evaluate-video-quality` skill
|
||||
@@ -1,21 +0,0 @@
|
||||
# Experiment Journal
|
||||
|
||||
Living log of all experiments. Each entry captures what was tried, the result,
|
||||
and any insights. Newest entries go at the top.
|
||||
|
||||
_No experiments logged yet. Use the `log-experiment` skill to add entries._
|
||||
|
||||
<!-- TEMPLATE — copy and fill for each new experiment:
|
||||
|
||||
## [YYYY-MM-DD] Experiment: <name>
|
||||
- **Hypothesis**: <what you expected to learn>
|
||||
- **Config**: model=..., lr=..., sp_size=..., gpus=..., script=...
|
||||
- **W&B run**: <run_id or URL>
|
||||
- **Duration**: <total wall time>
|
||||
- **Key metrics**: loss=..., step_time=..., grad_norm=...
|
||||
- **Checkpoint**: <path>
|
||||
- **Insight**: <what was learned>
|
||||
- **Status**: running | completed | failed | abandoned
|
||||
- **Related lessons**: `.agents/lessons/<filename>.md`
|
||||
|
||||
-->
|
||||
@@ -1,4 +0,0 @@
|
||||
{"name": "codebase-map", "description": "High-level structural index of the FastVideo-WorldModel repository", "path": "codebase-map/README.md", "status": "ready", "trust": "high"}
|
||||
{"name": "evaluation-registry", "description": "Catalog of all evaluation metrics with detailed explanations, implementation status, and usage guides", "path": "evaluation-registry/README.md", "status": "draft", "trust": "medium"}
|
||||
{"name": "experiment-journal", "description": "Living log of all experiments with hypotheses, configs, metrics, and insights", "path": "experiment-journal/README.md", "status": "draft", "trust": "medium"}
|
||||
{"name": "related-work", "description": "Index of related papers, repos, and blog posts with structured comparisons to FastVideo", "path": "related-work/README.md", "status": "draft", "trust": "low"}
|
||||
@@ -1,34 +0,0 @@
|
||||
# Related Work Index
|
||||
|
||||
Each file in this directory is a structured summary of a related paper, repo,
|
||||
or blog post relevant to FastVideo-WorldModel training.
|
||||
|
||||
## File Format
|
||||
|
||||
Each file is named `<slug>.md` and follows this structure:
|
||||
|
||||
```markdown
|
||||
---
|
||||
title: <paper/repo title>
|
||||
source: <URL or citation>
|
||||
type: paper | repo | blog
|
||||
date_indexed: <ISO-8601>
|
||||
tags: [world-model, distillation, evaluation, reward-shaping, ...]
|
||||
---
|
||||
|
||||
## Summary
|
||||
<1-2 paragraph summary of the work.>
|
||||
|
||||
## Key Differences from FastVideo
|
||||
- <Bullet points comparing their approach to ours.>
|
||||
|
||||
## Actionable Insights
|
||||
- <What we could adopt or adapt.>
|
||||
```
|
||||
|
||||
## How to Add New Entries
|
||||
|
||||
Use the `index-related-work` skill, or manually create a file following the
|
||||
template above.
|
||||
|
||||
_No related work indexed yet._
|
||||
@@ -1,76 +0,0 @@
|
||||
# Agent Onboarding — FastVideo-WorldModel
|
||||
|
||||
Welcome, agent. This is the **master onboarding** guide. Follow the steps below,
|
||||
then check if a **domain-specific onboarding** exists for your task.
|
||||
|
||||
## Domain-Specific Onboarding
|
||||
|
||||
If your task falls into one of these areas, read the specialized guide **after**
|
||||
completing the general steps below:
|
||||
|
||||
| Domain | Guide | When to Use |
|
||||
|--------|-------|-------------|
|
||||
| **WorldModel Training** | `worldmodel-training/README.md` | Training, finetuning, distillation, experiment management |
|
||||
|
||||
---
|
||||
|
||||
## Step 1: Understand the Codebase
|
||||
|
||||
Read these files to build your context:
|
||||
|
||||
| Priority | File | What you learn |
|
||||
|----------|------|----------------|
|
||||
| 1 | `AGENTS.md` | Coding guidelines, build/test commands, PR conventions |
|
||||
| 2 | `docs/design/overview.md` | Architecture: models, pipelines, configs, registry |
|
||||
| 3 | `fastvideo/train/` | Refactored training framework (YAML-driven, modular methods/models/callbacks) |
|
||||
| 4 | `docs/training/overview.md` | Training data flow and preprocessing |
|
||||
| 5 | `docs/training/finetune.md` | Training arguments, parallelism, LoRA, validation |
|
||||
| 6 | `docs/contributing/coding_agents.md` | How to add model pipelines with agent assistance |
|
||||
|
||||
## Step 2: Discover Available Resources
|
||||
|
||||
Read these two index files to see what skills and memory modules exist:
|
||||
|
||||
- **`.agents/skills/index.jsonl`** — catalog of all agent skills (name + description)
|
||||
- **`.agents/memory/index.jsonl`** — catalog of all memory modules (name + description)
|
||||
|
||||
Each entry has a `path` field pointing to the full content. Only load the
|
||||
full README.md for modules relevant to your current task.
|
||||
|
||||
## Step 3: Check for Existing Skills & SOPs
|
||||
|
||||
Before writing new code or procedures:
|
||||
|
||||
1. **Skills**: Read `.agents/skills/index.jsonl` — find a matching skill by description.
|
||||
2. **Workflows/SOPs**: Browse `.agents/workflows/` — step-by-step procedures for common tasks.
|
||||
3. **Lessons**: Browse `.agents/lessons/` — known pitfalls and their fixes.
|
||||
|
||||
If a skill or SOP exists for your task, **use it**. If not, you are in **exploration mode** — see Step 4.
|
||||
|
||||
## Step 4: Exploration Mode
|
||||
|
||||
If no existing skill/SOP covers your task:
|
||||
|
||||
1. Document your progress in `.agents/exploration/<topic>.md` using the template in `.agents/exploration/README.md`.
|
||||
2. At the end of your session, reflect:
|
||||
- **What worked** → propose a new skill or SOP in the exploration log.
|
||||
- **What failed** → create a lesson in `.agents/lessons/`.
|
||||
3. Flag the exploration log for human review.
|
||||
|
||||
## Quick Reference
|
||||
|
||||
```
|
||||
.agents/
|
||||
├── ONBOARDING.md ← you are here
|
||||
├── STATUS.md ← dashboard: completeness & trust of all components
|
||||
├── skills/ ← reusable agent skills
|
||||
├── workflows/ ← SOPs and procedures
|
||||
├── memory/ ← persistent context (folder per topic + index.jsonl)
|
||||
│ ├── index.jsonl
|
||||
│ ├── codebase-map/
|
||||
│ ├── experiment-journal/
|
||||
│ ├── evaluation-registry/
|
||||
│ └── related-work/
|
||||
├── lessons/ ← mistakes and fixes
|
||||
└── exploration/ ← draft procedures
|
||||
```
|
||||
@@ -1,302 +0,0 @@
|
||||
# WorldModel Training — Agent Onboarding
|
||||
|
||||
Specialized onboarding for agents working on FastVideo-WorldModel training,
|
||||
distillation, and evaluation. Read the master onboarding (`.agents/onboarding/README.md`)
|
||||
first, then come here.
|
||||
|
||||
---
|
||||
|
||||
## Domain Context
|
||||
|
||||
FastVideo-WorldModel trains **interactive world models** — video generation systems
|
||||
that respond to user actions (keyboard/mouse) in real-time. The architecture is
|
||||
based on **Wan2.1** (SkyReels-V2) DiT models with causal attention for
|
||||
auto-regressive streaming generation.
|
||||
|
||||
**Key techniques you will work with:**
|
||||
- Full finetuning and LoRA on Wan / LTX-2 / MatrixGame models
|
||||
- DMD-based distillation (few-step generation)
|
||||
- Self-Forcing distillation (causal streaming)
|
||||
- Diffusion-Forcing SFT (DFSFT) for causal models
|
||||
- VSA (Variable Sparsity Acceleration) for efficient training
|
||||
|
||||
---
|
||||
|
||||
## Training Code: Two Generations
|
||||
|
||||
### New modular framework: `fastvideo/train/` (preferred)
|
||||
|
||||
The refactored training code uses a **YAML-only config-driven** architecture
|
||||
with composable methods, per-role models, and a callback system. All new
|
||||
training work should use this framework.
|
||||
|
||||
### Legacy pipelines: `fastvideo/training/` (deprecated)
|
||||
|
||||
The old monolithic pipeline classes (`WanTrainingPipeline`,
|
||||
`DistillationPipeline`, etc.) still exist but are being phased out. The new
|
||||
framework imports select utilities from `fastvideo/training/` for backward
|
||||
compatibility (EMA, gradient clipping, checkpoint wrappers).
|
||||
|
||||
---
|
||||
|
||||
## Essential Reading (Training-Specific)
|
||||
|
||||
Read these **in order** before touching any training code:
|
||||
|
||||
| # | File | What You Learn |
|
||||
|---|------|----------------|
|
||||
| 1 | `docs/training/overview.md` | Training data flow: raw video → text embeddings + video latents → training |
|
||||
| 2 | `docs/training/finetune.md` | Training arguments, parallelism (SP/TP), LoRA, validation settings |
|
||||
| 3 | `docs/training/data_preprocess.md` | How to preprocess datasets into the expected format |
|
||||
| 4 | `docs/design/overview.md` | Architecture: models, pipelines, configs, registry |
|
||||
|
||||
---
|
||||
|
||||
## New Training Framework (`fastvideo/train/`)
|
||||
|
||||
### Architecture Overview
|
||||
|
||||
```
|
||||
fastvideo/train/
|
||||
├── __init__.py → exports Trainer
|
||||
├── trainer.py → main training loop coordinator
|
||||
├── entrypoint/
|
||||
│ ├── train.py → YAML-only training entrypoint
|
||||
│ └── dcp_to_diffusers.py → checkpoint conversion utility
|
||||
├── methods/ → training algorithms (TrainingMethod ABC)
|
||||
│ ├── base.py → TrainingMethod base class
|
||||
│ ├── fine_tuning/
|
||||
│ │ ├── finetune.py → FineTuneMethod (supervised finetuning)
|
||||
│ │ └── dfsft.py → DiffusionForcingSFTMethod (causal)
|
||||
│ ├── distribution_matching/
|
||||
│ │ ├── dmd2.py → DMD2Method (distribution matching distill)
|
||||
│ │ └── self_forcing.py → SelfForcingMethod (causal streaming)
|
||||
│ ├── knowledge_distillation/ → (stub, not yet implemented)
|
||||
│ └── consistency_model/ → (stub, not yet implemented)
|
||||
├── models/ → per-role model instances
|
||||
│ ├── base.py → ModelBase & CausalModelBase (ABC)
|
||||
│ └── wan/
|
||||
│ ├── wan.py → WanModel (non-causal)
|
||||
│ └── wan_causal.py → WanCausalModel (causal streaming)
|
||||
├── callbacks/ → training hooks & monitoring
|
||||
│ ├── callback.py → Callback base class + CallbackDict
|
||||
│ ├── grad_clip.py → GradNormClipCallback
|
||||
│ ├── ema.py → EMACallback (shadow weights)
|
||||
│ └── validation.py → ValidationCallback (sampling + eval)
|
||||
└── utils/ → configuration, building, checkpointing
|
||||
├── builder.py → build_from_config() (config → runtime)
|
||||
├── checkpoint.py → CheckpointManager (DCP-based)
|
||||
├── config.py → load_run_config() (YAML → RunConfig)
|
||||
├── training_config.py → TypedConfig dataclasses
|
||||
├── optimizer.py → build_optimizer_and_scheduler()
|
||||
├── instantiate.py → resolve_target() + instantiate()
|
||||
├── tracking.py → build_tracker() (W&B, etc.)
|
||||
├── dataloader.py → dataloader utilities
|
||||
├── module_state.py → apply_trainable()
|
||||
└── moduleloader.py → load_module_from_path()
|
||||
```
|
||||
|
||||
### Key Concepts
|
||||
|
||||
**TrainingMethod** (`methods/base.py`): Abstract base class for all training
|
||||
algorithms. Owns role models (student, teacher, critic), manages checkpoint
|
||||
state, and defines the training step interface.
|
||||
|
||||
**ModelBase** (`models/base.py`): Per-role model wrapper. Each role (student,
|
||||
teacher, critic) gets its own `ModelBase` instance owning a `transformer` and
|
||||
`noise_scheduler`. `CausalModelBase` extends this for streaming models.
|
||||
|
||||
**Callback system** (`callbacks/`): Composable hooks for gradient clipping,
|
||||
EMA, validation, etc. Configured via YAML, dispatched by `CallbackDict`.
|
||||
|
||||
**Config system** (`utils/config.py`, `utils/training_config.py`): YAML files
|
||||
are parsed into typed `RunConfig` dataclass trees. Models and methods use
|
||||
`_target_` fields for instantiation (similar to Hydra).
|
||||
|
||||
### Training Flow
|
||||
|
||||
```
|
||||
run_training_from_config(config_path)
|
||||
→ load_run_config() # YAML → RunConfig
|
||||
→ init_distributed() # TP/SP setup
|
||||
→ build_from_config() # instantiate models, method, dataloader
|
||||
→ Trainer.run() # main loop:
|
||||
├─ callbacks.on_train_start()
|
||||
├─ checkpoint_manager.maybe_resume()
|
||||
├─ for step in range(max_steps):
|
||||
│ ├─ method.single_train_step(batch)
|
||||
│ ├─ method.backward()
|
||||
│ ├─ callbacks.on_before_optimizer_step()
|
||||
│ ├─ method.optimizers_schedulers_step()
|
||||
│ ├─ tracker.log(metrics, step)
|
||||
│ ├─ callbacks.on_training_step_end()
|
||||
│ └─ checkpoint_manager.maybe_save(step)
|
||||
├─ callbacks.on_train_end()
|
||||
└─ checkpoint_manager.save_final()
|
||||
```
|
||||
|
||||
### Training Methods
|
||||
|
||||
| Method | Class | Use Case |
|
||||
|--------|-------|----------|
|
||||
| **FineTune** | `FineTuneMethod` | Single-role supervised finetuning |
|
||||
| **DFSFT** | `DiffusionForcingSFTMethod` | Diffusion-forcing SFT with inhomogeneous timesteps |
|
||||
| **DMD2** | `DMD2Method` | Multi-role distribution matching distillation (student + teacher + critic) |
|
||||
| **Self-Forcing** | `SelfForcingMethod` | Extends DMD2 for causal student rollouts |
|
||||
|
||||
### Launching Training (New Framework)
|
||||
|
||||
Training is launched via `torchrun` with a single YAML config:
|
||||
|
||||
```bash
|
||||
torchrun --nproc_per_node <N_GPUS> \
|
||||
-m fastvideo.train.entrypoint.train \
|
||||
--config examples/train/<config>.yaml
|
||||
```
|
||||
|
||||
### Example YAML Configs
|
||||
|
||||
| Config | Method | Description |
|
||||
|--------|--------|-------------|
|
||||
| `examples/train/finetune_wan2.1_t2v_1.3B_vsa_phase3.4_0.9sparsity.yaml` | FineTune | Wan 1.3B finetuning with VSA sparsity |
|
||||
| `examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml` | DMD2 | Wan 1.3B distillation (student + teacher + critic) |
|
||||
| `examples/train/dfsft_wan_causal_t2v_1.3B.yaml` | DFSFT | Causal Wan 1.3B diffusion-forcing SFT |
|
||||
| `examples/train/self_forcing_wan_causal_t2v_1.3B.yaml` | Self-Forcing | Causal streaming distillation |
|
||||
|
||||
### Checkpointing (New Framework)
|
||||
|
||||
**CheckpointManager** (`utils/checkpoint.py`) saves via `torch.distributed.checkpoint`:
|
||||
|
||||
```
|
||||
output_dir/
|
||||
└─ checkpoint-{step}/
|
||||
├─ dcp/ # DCP state dict
|
||||
├─ config.json # resolved training config
|
||||
└─ .fastvideo_metadata.json
|
||||
```
|
||||
|
||||
Checkpoint state includes: role model weights, per-role optimizers/schedulers,
|
||||
CUDA RNG state, and callback state (e.g., EMA shadow weights).
|
||||
|
||||
### Config Structure
|
||||
|
||||
A YAML config defines the full training pipeline:
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
model_path: ...
|
||||
trainable: true
|
||||
teacher: # optional, for distillation
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
model_path: ...
|
||||
trainable: false
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.fine_tuning.FineTuneMethod
|
||||
# method-specific params...
|
||||
|
||||
training:
|
||||
distributed: { num_gpus: 8, tp_size: 1, sp_size: 8 }
|
||||
data: { data_path: ..., batch_size: 1 }
|
||||
optimizer: { lr: 1e-5, lr_scheduler: constant_with_warmup }
|
||||
loop: { max_train_steps: 1000 }
|
||||
checkpoint: { output_dir: ./outputs }
|
||||
tracker: { trackers: [wandb], project_name: ... }
|
||||
|
||||
callbacks:
|
||||
grad_clip:
|
||||
_target_: fastvideo.train.callbacks.GradNormClipCallback
|
||||
max_grad_norm: 1.0
|
||||
validation:
|
||||
_target_: fastvideo.train.callbacks.ValidationCallback
|
||||
validation_steps: 100
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Legacy Training Pipelines (`fastvideo/training/`)
|
||||
|
||||
> **Note:** Use the new `fastvideo/train/` framework for new work. This section
|
||||
> is retained for reference on existing pipelines not yet migrated.
|
||||
|
||||
| Pipeline | Entrypoint | Use Case |
|
||||
|----------|-----------|----------|
|
||||
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | Standard text-to-video finetune / LoRA |
|
||||
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | Image-to-video (first frame conditioned) |
|
||||
| MatrixGame finetune | `fastvideo/training/matrixgame_training_pipeline.py` | Action-conditioned world model |
|
||||
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | LTX-2 architecture finetuning |
|
||||
| Wan DMD distillation | `fastvideo/training/wan_distillation_pipeline.py` | Few-step distillation via DMD |
|
||||
| Self-Forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | Causal streaming distillation |
|
||||
|
||||
---
|
||||
|
||||
## Key Infrastructure
|
||||
|
||||
### W&B Integration
|
||||
- **Tracker**: `fastvideo/training/trackers.py` — `WandbTracker` class
|
||||
- **New framework tracker**: `fastvideo/train/utils/tracking.py` — `build_tracker()`
|
||||
- **Env vars**: `WANDB_API_KEY`, `WANDB_BASE_URL`, `WANDB_MODE`
|
||||
|
||||
### Parallelism
|
||||
- **SP** (Sequence Parallel): splits video frames across GPUs — `sp_size: N`
|
||||
- **TP** (Tensor Parallel): splits model layers across GPUs — `tp_size: N`
|
||||
- Typical configs: SP=2–8, TP=1–2
|
||||
|
||||
---
|
||||
|
||||
## Evaluation (for training runs)
|
||||
|
||||
Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
|
||||
|
||||
**Quick summary for training agents:**
|
||||
| Metric | When to Use | Trust |
|
||||
|--------|-------------|-------|
|
||||
| **Loss trajectory** | Every run, real-time from W&B | Medium |
|
||||
| **SSIM** | When comparing against reference outputs | High |
|
||||
| **FVD** | For benchmarking model quality (`benchmarks/fvd/`) | High |
|
||||
| **LPIPS** | LoRA merge validation | Medium |
|
||||
| **Human preference** | Major checkpoints | Highest |
|
||||
|
||||
---
|
||||
|
||||
## Common Workflows
|
||||
|
||||
| Task | Skill / SOP |
|
||||
|------|-------------|
|
||||
| Launch a training run | `.agents/skills/launch-experiment/SKILL.md` |
|
||||
| Monitor a running experiment | `.agents/skills/monitor-experiment/SKILL.md` |
|
||||
| Summarize final results | `.agents/skills/summarize-run/SKILL.md` |
|
||||
| Full experiment lifecycle | `.agents/workflows/experiment-lifecycle.md` |
|
||||
| Capture lessons from failures | `.agents/workflows/lesson-capture.md` |
|
||||
|
||||
---
|
||||
|
||||
## World Model–Specific Concepts
|
||||
|
||||
### Action Injection (MatrixGame)
|
||||
The MatrixGame pipeline adds **action modules** to each DiT block, enabling
|
||||
frame-level mouse/keyboard input conditioning. The action sequence is injected
|
||||
per-frame alongside the latent video tokens.
|
||||
|
||||
### Causal Architecture
|
||||
For streaming generation, the model uses **causal attention** (each frame only
|
||||
attends to previous frames). This enables auto-regressive chunk-by-chunk
|
||||
generation — critical for real-time interactive world models.
|
||||
|
||||
### Self-Forcing Distillation
|
||||
A **data-free** distillation method where the student model is trained to
|
||||
generate coherent video sequences by being forced to use its own previous
|
||||
outputs (rather than ground-truth) as context. This produces models robust to
|
||||
their own error accumulation during long auto-regressive generation.
|
||||
|
||||
### DMD Distillation (Distribution Matching Distillation)
|
||||
Reduces inference steps from ~50 to 3–4 by training a student model to match
|
||||
the output distribution of the teacher model. Uses a critic network to estimate
|
||||
distribution divergence.
|
||||
|
||||
### Diffusion-Forcing SFT (DFSFT)
|
||||
Supervised finetuning with **inhomogeneous timesteps** across chunks — each
|
||||
chunk in a causal sequence can have a different noise level, training the model
|
||||
to handle mixed-fidelity contexts.
|
||||
@@ -1,57 +0,0 @@
|
||||
---
|
||||
name: <skill-name>
|
||||
description: <one-line description — Codex uses this for implicit invocation matching>
|
||||
---
|
||||
|
||||
# <Skill Name>
|
||||
|
||||
## Purpose
|
||||
<Why this skill exists and when to use it.>
|
||||
|
||||
## Prerequisites
|
||||
- <What must be true before using this skill>
|
||||
|
||||
## Inputs
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `param1` | Yes | ... |
|
||||
|
||||
## Steps
|
||||
|
||||
1. **Step 1 title**
|
||||
- Detail...
|
||||
|
||||
2. **Step 2 title**
|
||||
- Detail...
|
||||
|
||||
## Outputs
|
||||
- <What this skill produces>
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
<Example invocation or prompt snippet>
|
||||
```
|
||||
|
||||
## References
|
||||
- <Links to relevant files in the codebase>
|
||||
|
||||
---
|
||||
|
||||
## Folder Structure
|
||||
|
||||
Each skill lives in its own directory under `.agents/skills/`:
|
||||
|
||||
```
|
||||
.agents/skills/<skill-name>/
|
||||
├── SKILL.md # Required: instructions + metadata (this file)
|
||||
├── scripts/ # Optional: executable helper scripts
|
||||
├── references/ # Optional: documentation, papers
|
||||
└── assets/ # Optional: templates, resources
|
||||
```
|
||||
|
||||
After creating a new skill, add an entry to `.agents/skills/index.jsonl`:
|
||||
|
||||
```json
|
||||
{"name": "<skill-name>", "description": "<description>", "path": "<skill-name>/SKILL.md", "status": "draft", "trust": "low"}
|
||||
```
|
||||
@@ -1,128 +0,0 @@
|
||||
---
|
||||
name: evaluate-video-quality
|
||||
description: Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)
|
||||
---
|
||||
|
||||
# Evaluate Video Quality
|
||||
|
||||
## Purpose
|
||||
Assess the quality of videos generated by a training run. Combines multiple
|
||||
signals to give a holistic quality assessment. This skill is **evolving** —
|
||||
new metrics will be added as they are developed.
|
||||
|
||||
## Prerequisites
|
||||
- Generated videos available locally or via W&B artifacts.
|
||||
- For SSIM: reference videos from official implementations.
|
||||
- For caption consistency: LLM access (optional, stub for now).
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `video_paths` | Yes | List of paths to generated videos |
|
||||
| `reference_paths` | No | Paths to reference videos (for SSIM) |
|
||||
| `prompts` | No | Prompts used to generate videos (for caption check) |
|
||||
| `loss_summary` | No | Path to W&B summary JSON (for loss trajectory) |
|
||||
| `metrics` | No | Which metrics to run (default: all available) |
|
||||
|
||||
## Available Metrics
|
||||
|
||||
Check `.agents/memory/evaluation-registry/README.md` for the current catalog.
|
||||
|
||||
### SSIM (Active)
|
||||
|
||||
Leverages the existing infrastructure in `fastvideo/tests/ssim/`.
|
||||
|
||||
```bash
|
||||
pytest fastvideo/tests/ssim/ -vs --video-path <generated> --reference-path <reference>
|
||||
```
|
||||
|
||||
Or use the SSIM utility directly:
|
||||
|
||||
```python
|
||||
from fastvideo.tests.ssim.ssim_utils import compute_ssim
|
||||
score = compute_ssim(generated_video, reference_video)
|
||||
# score > 0.85 is typically "acceptable"
|
||||
```
|
||||
|
||||
**Interpretation**:
|
||||
| SSIM Range | Quality |
|
||||
|------------|---------|
|
||||
| > 0.90 | Excellent — very close to reference |
|
||||
| 0.80–0.90 | Good — acceptable for most uses |
|
||||
| 0.70–0.80 | Fair — noticeable differences |
|
||||
| < 0.70 | Poor — significant quality issues |
|
||||
|
||||
### Loss Trajectory (Active)
|
||||
|
||||
Analyze the loss curve shape from W&B summary:
|
||||
|
||||
```python
|
||||
import json
|
||||
with open(loss_summary_path) as f:
|
||||
summary = json.load(f)
|
||||
|
||||
final_loss = summary["train_loss"]
|
||||
runtime = summary["_runtime"]
|
||||
steps = summary["_step"]
|
||||
```
|
||||
|
||||
**Early-stage heuristics** (first 500 steps):
|
||||
- Loss should be decreasing (even slightly).
|
||||
- Grad norm should be stable (no wild oscillations).
|
||||
- If loss is flat or increasing, flag for review.
|
||||
|
||||
### Caption Consistency (Draft — Not Yet Calibrated)
|
||||
|
||||
Use an LLM to evaluate whether the video content matches the input prompt.
|
||||
|
||||
```
|
||||
Prompt: "A golden retriever playing in the snow"
|
||||
Video: <path>
|
||||
|
||||
Score the video on:
|
||||
1. Object presence (is there a golden retriever?)
|
||||
2. Action accuracy (is it playing?)
|
||||
3. Environment match (is there snow?)
|
||||
4. Overall coherence (does it look natural?)
|
||||
|
||||
Each 1-5, total /20.
|
||||
```
|
||||
|
||||
> ⚠️ This metric is in **draft** status. Results should not be treated as
|
||||
> ground truth until calibrated against human judgments.
|
||||
|
||||
## Steps
|
||||
|
||||
1. **Identify available metrics** — Check `.agents/memory/evaluation-registry/README.md`.
|
||||
2. **Run each metric** — Collect scores.
|
||||
3. **Aggregate** — Produce a combined quality report.
|
||||
4. **Log** — Update the experiment journal with quality results.
|
||||
|
||||
## Outputs
|
||||
|
||||
```markdown
|
||||
## Video Quality Report: <experiment_name>
|
||||
|
||||
| Metric | Score | Threshold | Status |
|
||||
|--------|-------|-----------|--------|
|
||||
| SSIM (avg) | 0.87 | > 0.80 | ✅ Pass |
|
||||
| Loss trajectory | decreasing | decreasing | ✅ Pass |
|
||||
| Caption consistency | 16/20 | > 14/20 | ✅ Pass |
|
||||
|
||||
### Per-Video Scores
|
||||
| Video | SSIM | Caption |
|
||||
|-------|------|---------|
|
||||
| video_001.mp4 | 0.89 | 17/20 |
|
||||
| video_002.mp4 | 0.85 | 15/20 |
|
||||
```
|
||||
|
||||
## References
|
||||
- `fastvideo/tests/ssim/` — SSIM test infrastructure
|
||||
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — loss comparison
|
||||
- `.agents/memory/evaluation-registry/README.md` — metric catalog
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version with SSIM, loss trajectory, caption consistency stub |
|
||||
@@ -1,94 +0,0 @@
|
||||
---
|
||||
name: index-related-work
|
||||
description: Ingest a paper or repository into the related work index
|
||||
---
|
||||
|
||||
# Index Related Work
|
||||
|
||||
## Purpose
|
||||
Create a structured summary of a related paper, repository, or blog post and
|
||||
add it to `.agents/memory/related-work/` for future reference. This builds the
|
||||
agent's knowledge base for making informed decisions about training, evaluation,
|
||||
and architecture choices.
|
||||
|
||||
## Prerequisites
|
||||
- Access to the paper/repo (URL, PDF, or local clone).
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `source` | Yes | URL, citation, or local path |
|
||||
| `type` | Yes | `paper`, `repo`, or `blog` |
|
||||
| `tags` | No | List of tags (default: inferred from content) |
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Extract key information
|
||||
|
||||
For **papers**: Read abstract, method section, experimental setup, and results.
|
||||
For **repos**: Read README, key source files, and training scripts.
|
||||
For **blogs**: Read the full post.
|
||||
|
||||
Focus on:
|
||||
- What problem does it solve?
|
||||
- What architecture/technique is used?
|
||||
- How does it relate to FastVideo's approach?
|
||||
|
||||
### 2. Create the index entry
|
||||
|
||||
Write to `.agents/memory/related-work/<slug>.md`:
|
||||
|
||||
```markdown
|
||||
---
|
||||
title: <title>
|
||||
source: <URL or citation>
|
||||
type: paper | repo | blog
|
||||
date_indexed: <ISO-8601>
|
||||
tags: [world-model, distillation, evaluation, ...]
|
||||
---
|
||||
|
||||
## Summary
|
||||
<1-2 paragraph summary.>
|
||||
|
||||
## Key Differences from FastVideo
|
||||
- <comparison points>
|
||||
|
||||
## Actionable Insights
|
||||
- <what we could adopt or adapt>
|
||||
```
|
||||
|
||||
### 3. Update the catalog
|
||||
|
||||
If `.agents/memory/related-work/_catalog.md` exists, append the new entry.
|
||||
If not, create it:
|
||||
|
||||
```markdown
|
||||
# Related Work Catalog
|
||||
|
||||
| Slug | Title | Type | Tags | Date |
|
||||
|------|-------|------|------|------|
|
||||
| <slug> | <title> | <type> | <tags> | <date> |
|
||||
```
|
||||
|
||||
## Outputs
|
||||
- New file in `.agents/memory/related-work/<slug>.md`.
|
||||
- Updated catalog.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Index the Self-Forcing paper:
|
||||
|
||||
source: https://arxiv.org/abs/2406.xxxxx
|
||||
type: paper
|
||||
tags: [world-model, self-forcing, distillation]
|
||||
```
|
||||
|
||||
## References
|
||||
- `.agents/memory/related-work/README.md` — schema documentation
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,7 +0,0 @@
|
||||
{"name": "launch-experiment", "description": "Generate and execute a training launch command for FastVideo models", "path": "launch-experiment/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "monitor-experiment", "description": "Poll a running W&B training run for progress and emit structured alerts", "path": "monitor-experiment/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "summarize-run", "description": "Extract a W&B run summary into a structured experiment report", "path": "summarize-run/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "log-experiment", "description": "Append or update an experiment entry in the experiment journal", "path": "log-experiment/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "evaluate-video-quality", "description": "Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)", "path": "evaluate-video-quality/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "index-related-work", "description": "Ingest a paper or repository into the related work index", "path": "index-related-work/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "search-related-work", "description": "Query the related work index for relevant papers, repos, or comparisons", "path": "search-related-work/SKILL.md", "status": "draft", "trust": "low"}
|
||||
@@ -1,127 +0,0 @@
|
||||
---
|
||||
name: launch-experiment
|
||||
description: Generate and execute a training launch command for FastVideo models
|
||||
---
|
||||
|
||||
# Launch Experiment
|
||||
|
||||
## Purpose
|
||||
Construct a fully-specified `torchrun` training command for a FastVideo model
|
||||
given a target pipeline, dataset, and hyperparameter overrides. This skill
|
||||
automates the boilerplate of setting environment variables, picking the right
|
||||
entrypoint, and applying defaults from the closest example script.
|
||||
|
||||
## Prerequisites
|
||||
- The repo is cloned and `fastvideo` is installed (`uv pip install -e .[dev]`).
|
||||
- Dataset is preprocessed (see `docs/training/data_preprocess.md`).
|
||||
- `WANDB_API_KEY` is set in the environment (or `WANDB_MODE=offline` for local).
|
||||
- GPU resources are available (multi-GPU requires NCCL).
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `pipeline` | Yes | Training pipeline type: `finetune`, `distill-dmd`, `self-forcing`, `lora`, `consistency` |
|
||||
| `model` | Yes | Model family: `wan-t2v-1.3B`, `wan-i2v-14B`, `ltx2`, `matrixgame` |
|
||||
| `data_path` | Yes | Path to preprocessed dataset (parquet) |
|
||||
| `num_gpus` | Yes | Number of GPUs |
|
||||
| `overrides` | No | Dict of hyperparameter overrides (any CLI arg) |
|
||||
| `output_dir` | No | Output directory (default: `outputs/<model>_<pipeline>`) |
|
||||
| `run_name` | No | W&B run name (default: auto-generated) |
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Identify the training entrypoint
|
||||
|
||||
| Pipeline | Entrypoint |
|
||||
|----------|-----------|
|
||||
| `finetune` (Wan T2V) | `fastvideo/training/wan_training_pipeline.py` |
|
||||
| `finetune` (Wan I2V) | `fastvideo/training/wan_i2v_training_pipeline.py` |
|
||||
| `finetune` (LTX-2) | `fastvideo/training/ltx2_training_pipeline.py` |
|
||||
| `finetune` (MatrixGame) | `fastvideo/training/matrixgame_training_pipeline.py` |
|
||||
| `distill-dmd` | `fastvideo/training/wan_distillation_pipeline.py` |
|
||||
| `self-forcing` | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` |
|
||||
|
||||
### 2. Resolve default hyperparameters
|
||||
|
||||
Find the closest example script in `examples/training/` for the model:
|
||||
|
||||
| Model | Example Script Directory |
|
||||
|-------|-------------------------|
|
||||
| `wan-t2v-1.3B` | `examples/training/finetune/wan_t2v_1.3B/crush_smol/` |
|
||||
| `wan-i2v-14B` | `examples/training/finetune/wan_i2v_14B_480p/crush_smol/` |
|
||||
| `ltx2` | `examples/training/finetune/ltx2/` |
|
||||
| `matrixgame` | `examples/training/finetune/MatrixGame2.0/` |
|
||||
| `distill-dmd` | `scripts/distill/v1_distill_dmd_wan.sh` |
|
||||
|
||||
Read the script to extract default values for:
|
||||
- `--learning_rate`, `--train_batch_size`, `--sp_size`, `--tp_size`
|
||||
- `--num_latent_t`, `--num_height`, `--num_width`, `--num_frames`
|
||||
- `--gradient_accumulation_steps`, `--max_train_steps`
|
||||
- `--mixed_precision`, `--weight_decay`, `--max_grad_norm`
|
||||
- `--validation_steps`, `--validation_sampling_steps`
|
||||
|
||||
### 3. Set environment variables
|
||||
|
||||
```bash
|
||||
export WANDB_API_KEY="${WANDB_API_KEY}"
|
||||
export WANDB_BASE_URL="https://api.wandb.ai"
|
||||
export FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN
|
||||
export TOKENIZERS_PARALLELISM=false
|
||||
export TRITON_CACHE_DIR=/tmp/triton_cache
|
||||
```
|
||||
|
||||
### 4. Construct the torchrun command
|
||||
|
||||
```bash
|
||||
torchrun --nnodes 1 --nproc_per_node <num_gpus> \
|
||||
<entrypoint> \
|
||||
--pretrained_model_name_or_path <model_hf_id> \
|
||||
--data_path "<data_path>" \
|
||||
--output_dir "<output_dir>" \
|
||||
--wandb_run_name "<run_name>" \
|
||||
--tracker_project_name "<project_name>" \
|
||||
--log_validation \
|
||||
<...all hyperparameters...>
|
||||
```
|
||||
|
||||
### 5. Log to experiment journal
|
||||
|
||||
After launching, append an entry to `.agents/memory/experiment-journal/README.md`:
|
||||
|
||||
```markdown
|
||||
## [YYYY-MM-DD] Experiment: <run_name>
|
||||
- **Hypothesis**: <user-provided or auto-generated>
|
||||
- **Config**: model=<model>, lr=<lr>, sp_size=<sp>, gpus=<n>, script=<entrypoint>
|
||||
- **W&B run**: <pending — will be updated by monitor skill>
|
||||
- **Status**: running
|
||||
```
|
||||
|
||||
## Outputs
|
||||
- A ready-to-execute shell command.
|
||||
- An experiment journal entry.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Launch a Wan T2V 1.3B finetune on 4 GPUs with lr=5e-5 and max_train_steps=1000:
|
||||
|
||||
pipeline: finetune
|
||||
model: wan-t2v-1.3B
|
||||
data_path: data/crush_smol_preprocessed/
|
||||
num_gpus: 4
|
||||
overrides:
|
||||
learning_rate: 5e-5
|
||||
max_train_steps: 1000
|
||||
```
|
||||
|
||||
## References
|
||||
- `examples/training/finetune/wan_t2v_1.3B/crush_smol/finetune_t2v.sh`
|
||||
- `scripts/distill/v1_distill_dmd_wan.sh`
|
||||
- `docs/training/finetune.md` (training arguments table)
|
||||
- `fastvideo/training/trackers.py` (tracker initialization)
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,87 +0,0 @@
|
||||
---
|
||||
name: log-experiment
|
||||
description: Append or update an experiment entry in the experiment journal
|
||||
---
|
||||
|
||||
# Log Experiment
|
||||
|
||||
## Purpose
|
||||
Create or update an entry in `.agents/memory/experiment-journal/README.md` to maintain
|
||||
a living record of all experiments and their outcomes.
|
||||
|
||||
## Prerequisites
|
||||
- `.agents/memory/experiment-journal/README.md` exists.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `name` | Yes | Experiment name / identifier |
|
||||
| `hypothesis` | No | What you expected to learn |
|
||||
| `config` | Yes | Key config: model, lr, sp_size, gpus, script |
|
||||
| `wandb_run` | No | W&B run ID or URL |
|
||||
| `duration` | No | Total wall time |
|
||||
| `metrics` | No | Key metrics dict (loss, step_time, grad_norm) |
|
||||
| `checkpoint` | No | Path to checkpoint |
|
||||
| `insight` | No | What was learned |
|
||||
| `status` | Yes | `running`, `completed`, `failed`, `abandoned` |
|
||||
| `lessons` | No | Paths to related lesson files |
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Check for existing entry
|
||||
|
||||
Search `.agents/memory/experiment-journal/README.md` for an entry with the same name.
|
||||
If found, update it instead of creating a duplicate.
|
||||
|
||||
### 2. Format the entry
|
||||
|
||||
```markdown
|
||||
## [YYYY-MM-DD] Experiment: <name>
|
||||
- **Hypothesis**: <hypothesis or "N/A">
|
||||
- **Config**: model=<model>, lr=<lr>, sp_size=<sp>, gpus=<n>, script=<script>
|
||||
- **W&B run**: <wandb_run or "pending">
|
||||
- **Duration**: <duration or "in progress">
|
||||
- **Key metrics**: loss=<loss>, step_time=<step_time>, grad_norm=<grad_norm>
|
||||
- **Checkpoint**: <checkpoint or "N/A">
|
||||
- **Insight**: <insight or "pending">
|
||||
- **Status**: <status>
|
||||
- **Related lessons**: <lessons or "none">
|
||||
```
|
||||
|
||||
### 3. Insert at the top of the journal
|
||||
|
||||
New entries go at the top of the file (after the header), so the most recent
|
||||
experiments are always visible first.
|
||||
|
||||
### 4. Warn on duplicates
|
||||
|
||||
If a similar experiment name exists with `status: completed`, warn that this
|
||||
may be a repeat. If it's `status: running`, assume this is an update.
|
||||
|
||||
## Outputs
|
||||
- Updated `.agents/memory/experiment-journal/README.md`.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Log a completed experiment:
|
||||
|
||||
name: wan-t2v-finetune-lr5e5-sp4
|
||||
config: model=wan-t2v-1.3B, lr=5e-5, sp_size=4, gpus=4
|
||||
wandb_run: fastvideo/training/run_abc123
|
||||
duration: 2h 15m
|
||||
metrics: {loss: 0.065, step_time: 2.3, grad_norm: 0.35}
|
||||
checkpoint: outputs/wan_finetune/checkpoint-1000
|
||||
insight: LR 5e-5 converges 30% faster than 1e-5 with no quality loss
|
||||
status: completed
|
||||
```
|
||||
|
||||
## References
|
||||
- `.agents/memory/experiment-journal/README.md` — journal file
|
||||
- `.agents/workflows/experiment-lifecycle.md` — when to log
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,134 +0,0 @@
|
||||
---
|
||||
name: monitor-experiment
|
||||
description: Poll a running W&B training run for progress and emit structured alerts
|
||||
---
|
||||
|
||||
# Monitor Experiment
|
||||
|
||||
## Purpose
|
||||
Continuously (or on-demand) check a running experiment's W&B metrics and emit
|
||||
alerts for anomalies. Supports the "30-minute quality check" paradigm: after
|
||||
the first 30 minutes of a long training run, produce a checkpoint quality
|
||||
report before committing more resources.
|
||||
|
||||
## Prerequisites
|
||||
- `WANDB_API_KEY` is set in the environment.
|
||||
- The experiment is actively logging to W&B (not in `WANDB_MODE=offline`).
|
||||
- For offline mode: read from local `wandb-summary.json` instead.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `run_id` | Yes* | W&B run ID (e.g., `entity/project/run_id`) |
|
||||
| `output_dir` | Yes* | Local output directory (for offline mode fallback) |
|
||||
| `poll_interval` | No | Seconds between polls (default: 60) |
|
||||
| `alert_on` | No | List of alert conditions to enable (default: all) |
|
||||
|
||||
\* One of `run_id` or `output_dir` is required.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Connect to the run
|
||||
|
||||
**Online mode** (preferred):
|
||||
|
||||
```python
|
||||
import wandb
|
||||
api = wandb.Api()
|
||||
run = api.run("<run_id>")
|
||||
```
|
||||
|
||||
**Offline fallback**:
|
||||
|
||||
```python
|
||||
import json
|
||||
summary_path = f"{output_dir}/tracker/wandb/latest-run/files/wandb-summary.json"
|
||||
with open(summary_path) as f:
|
||||
summary = json.load(f)
|
||||
```
|
||||
|
||||
### 2. Track key metrics
|
||||
|
||||
| Metric | W&B Key | Description |
|
||||
|--------|---------|-------------|
|
||||
| Training loss | `train_loss` | Primary training loss |
|
||||
| Gradient norm | `grad_norm` | Gradient magnitude |
|
||||
| Step time | `step_time` | Wall-clock seconds per step |
|
||||
| Learning rate | `learning_rate` | Current LR |
|
||||
| Avg step time | `avg_step_time` | Running average step time |
|
||||
| Validation videos | `validation_videos_*` | Generated validation samples |
|
||||
|
||||
### 3. Evaluate alert conditions
|
||||
|
||||
| Alert | Condition | Severity |
|
||||
|-------|-----------|----------|
|
||||
| **Loss spike** | `current_loss > 3 × rolling_avg_loss` | 🔴 Critical |
|
||||
| **NaN/Inf gradient** | `grad_norm` is NaN or Inf | 🔴 Critical |
|
||||
| **Step time regression** | `step_time > 2 × baseline_step_time` | 🟡 Warning |
|
||||
| **No progress** | No new W&B logs for > 10 minutes | 🟡 Warning |
|
||||
| **Loss plateau** | Loss change < 1% over last 100 steps | 🟢 Info |
|
||||
|
||||
### 4. Emit structured status
|
||||
|
||||
Output format (agent-consumable):
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "...",
|
||||
"step": 500,
|
||||
"metrics": {
|
||||
"train_loss": 0.078,
|
||||
"grad_norm": 0.41,
|
||||
"step_time": 2.5,
|
||||
"learning_rate": 1e-6
|
||||
},
|
||||
"alerts": [
|
||||
{"type": "loss_spike", "severity": "critical", "message": "Loss jumped to 0.45 (avg: 0.08)"}
|
||||
],
|
||||
"status": "running"
|
||||
}
|
||||
```
|
||||
|
||||
### 5. 30-Minute Quality Check
|
||||
|
||||
After the first 30 minutes of wall-clock time:
|
||||
1. Summarize the loss curve shape (decreasing? at what rate?).
|
||||
2. Check if validation videos have been generated.
|
||||
3. Report step count, loss at start vs. current, and estimated time to completion.
|
||||
4. Produce a go/no-go recommendation.
|
||||
|
||||
```markdown
|
||||
## 30-Minute Check: <run_name>
|
||||
- **Steps completed**: 150
|
||||
- **Loss**: 0.12 → 0.08 (↓ 33%)
|
||||
- **Grad norm**: stable at ~0.4
|
||||
- **Step time**: 2.5s/step (consistent)
|
||||
- **Validation videos**: 5 generated at step 100
|
||||
- **Recommendation**: ✅ Continue — loss is decreasing normally
|
||||
```
|
||||
|
||||
## Outputs
|
||||
- Structured JSON status updates.
|
||||
- Alert messages for anomalous conditions.
|
||||
- 30-minute checkpoint quality report.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Monitor W&B run "fastvideo/Wan_distillation/abc123":
|
||||
|
||||
run_id: fastvideo/Wan_distillation/abc123
|
||||
poll_interval: 120
|
||||
alert_on: [loss_spike, nan_gradient, step_time_regression]
|
||||
```
|
||||
|
||||
## References
|
||||
- `fastvideo/training/trackers.py` — `WandbTracker` implementation
|
||||
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — how summaries are compared
|
||||
- `fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json` — reference summary format
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,82 +0,0 @@
|
||||
---
|
||||
name: search-related-work
|
||||
description: Query the related work index for relevant papers, repos, or comparisons
|
||||
---
|
||||
|
||||
# Search Related Work
|
||||
|
||||
## Purpose
|
||||
Search through `.agents/memory/related-work/` to find indexed papers, repos,
|
||||
or blog posts relevant to a query. Use this when you need to understand how
|
||||
other work compares to FastVideo's approach, or when looking for techniques
|
||||
to adopt.
|
||||
|
||||
## Prerequisites
|
||||
- The related work index has entries (`.agents/memory/related-work/*.md`).
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `query` | Yes | Natural language query |
|
||||
| `tags` | No | Filter by tags (e.g., `[distillation, evaluation]`) |
|
||||
| `type` | No | Filter by type (`paper`, `repo`, `blog`) |
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Search the index
|
||||
|
||||
Use grep-based search through `.agents/memory/related-work/`:
|
||||
|
||||
```bash
|
||||
# Search by content
|
||||
grep -rl "<query>" .agents/memory/related-work/
|
||||
|
||||
# Search by tags (in frontmatter)
|
||||
grep -l "tags:.*<tag>" .agents/memory/related-work/*.md
|
||||
```
|
||||
|
||||
### 2. Rank results
|
||||
|
||||
For each matching file:
|
||||
1. Read the file.
|
||||
2. Score relevance to the query based on:
|
||||
- Title match
|
||||
- Tag match
|
||||
- Content match (summary, differences, insights)
|
||||
3. Return top results.
|
||||
|
||||
### 3. Format output
|
||||
|
||||
```markdown
|
||||
## Related Work Search: "<query>"
|
||||
|
||||
### 1. <Title> (relevance: high)
|
||||
- **Source**: <URL>
|
||||
- **Tags**: <tags>
|
||||
- **Key insight**: <most relevant excerpt>
|
||||
- **File**: `.agents/memory/related-work/<slug>.md`
|
||||
|
||||
### 2. <Title> (relevance: medium)
|
||||
...
|
||||
```
|
||||
|
||||
## Outputs
|
||||
- Ranked list of relevant related work entries with excerpts.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Search for work related to video quality evaluation metrics:
|
||||
|
||||
query: "video generation quality evaluation metrics"
|
||||
tags: [evaluation]
|
||||
```
|
||||
|
||||
## References
|
||||
- `.agents/memory/related-work/README.md` — index schema
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,137 +0,0 @@
|
||||
---
|
||||
name: summarize-run
|
||||
description: Extract a W&B run summary into a structured experiment report
|
||||
---
|
||||
|
||||
# Summarize Run
|
||||
|
||||
## Purpose
|
||||
After a training run completes (or at any checkpoint), extract key metrics from
|
||||
the W&B run summary and produce a structured markdown report. Supports both
|
||||
online (W&B API) and offline (local `wandb-summary.json`) modes.
|
||||
|
||||
## Prerequisites
|
||||
- Run has completed or reached a checkpoint with a saved summary.
|
||||
- For online: `WANDB_API_KEY` set in environment.
|
||||
- For offline: access to `<output_dir>/tracker/wandb/latest-run/files/wandb-summary.json`.
|
||||
|
||||
## Inputs
|
||||
|
||||
| Parameter | Required | Description |
|
||||
|-----------|----------|-------------|
|
||||
| `run_id` | Yes* | W&B run ID for online access |
|
||||
| `output_dir` | Yes* | Local output dir for offline access |
|
||||
| `reference_run` | No | Path to reference `wandb-summary.json` for comparison |
|
||||
| `experiment_name` | No | Name for the journal entry (default: from W&B) |
|
||||
|
||||
\* One of `run_id` or `output_dir` is required.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Load run summary
|
||||
|
||||
**Online**:
|
||||
|
||||
```python
|
||||
import wandb
|
||||
api = wandb.Api()
|
||||
run = api.run("<run_id>")
|
||||
summary = dict(run.summary)
|
||||
config = dict(run.config)
|
||||
```
|
||||
|
||||
**Offline** (existing codebase pattern from `fastvideo/tests/training/`):
|
||||
|
||||
```python
|
||||
import json
|
||||
summary_path = f"{output_dir}/tracker/wandb/latest-run/files/wandb-summary.json"
|
||||
with open(summary_path) as f:
|
||||
summary = json.load(f)
|
||||
```
|
||||
|
||||
### 2. Extract key fields
|
||||
|
||||
| Field | Source | Description |
|
||||
|-------|--------|-------------|
|
||||
| `train_loss` | `summary["train_loss"]` | Final training loss |
|
||||
| `avg_step_time` | `summary["avg_step_time"]` | Average seconds per step |
|
||||
| `step_time` | `summary["step_time"]` | Last step time |
|
||||
| `grad_norm` | `summary["grad_norm"]` | Final gradient norm |
|
||||
| `learning_rate` | `summary["learning_rate"]` | Final LR |
|
||||
| `_step` | `summary["_step"]` | Total steps completed |
|
||||
| `_runtime` | `summary["_runtime"]` | Total wall-clock seconds |
|
||||
| `validation_videos_*` | `summary[key]` | Validation video artifacts |
|
||||
|
||||
### 3. Compare against reference (optional)
|
||||
|
||||
Follow the pattern in `fastvideo/tests/training/Vanilla/test_training_loss.py`:
|
||||
|
||||
```python
|
||||
# Fields to compare
|
||||
compare_fields = ["train_loss", "grad_norm", "avg_step_time"]
|
||||
tolerance = 0.05 # 5% relative tolerance
|
||||
|
||||
for field in compare_fields:
|
||||
ref_val = reference_summary[field]
|
||||
cur_val = summary[field]
|
||||
diff_pct = abs(cur_val - ref_val) / abs(ref_val) * 100
|
||||
status = "✅" if diff_pct < tolerance * 100 else "⚠️"
|
||||
print(f"{status} {field}: {cur_val:.4f} (ref: {ref_val:.4f}, diff: {diff_pct:.1f}%)")
|
||||
```
|
||||
|
||||
### 4. Generate report
|
||||
|
||||
```markdown
|
||||
# Run Summary: <experiment_name>
|
||||
|
||||
| Metric | Value | Reference | Diff |
|
||||
|--------|-------|-----------|------|
|
||||
| Train Loss | 0.0788 | 0.0800 | -1.5% ✅ |
|
||||
| Avg Step Time | 2.81s | 2.80s | +0.4% ✅ |
|
||||
| Grad Norm | 0.408 | 0.410 | -0.5% ✅ |
|
||||
| Total Steps | 500 | — | — |
|
||||
| Wall Time | 23m 30s | — | — |
|
||||
|
||||
## Configuration
|
||||
- Model: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
- Learning Rate: 1e-6
|
||||
- Batch Size: 1
|
||||
- GPUs: 8 × (SP=1, TP=1)
|
||||
- Mixed Precision: bf16
|
||||
|
||||
## Validation Videos
|
||||
<list of validation video paths if available>
|
||||
|
||||
## Notes
|
||||
<any observations or anomalies>
|
||||
```
|
||||
|
||||
### 5. Update experiment journal
|
||||
|
||||
Append or update the experiment's entry in `.agents/memory/experiment-journal/README.md`
|
||||
with the final metrics and status.
|
||||
|
||||
## Outputs
|
||||
- Structured markdown report.
|
||||
- Updated experiment journal entry.
|
||||
|
||||
## Example Usage
|
||||
|
||||
```
|
||||
Summarize the run in output directory "outputs/wan_finetune":
|
||||
|
||||
output_dir: outputs/wan_finetune
|
||||
reference_run: fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json
|
||||
experiment_name: wan-t2v-finetune-lr1e6
|
||||
```
|
||||
|
||||
## References
|
||||
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — reference comparison pattern
|
||||
- `fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json` — example summary
|
||||
- `fastvideo/tests/training/lora/test_lora_training.py` — LoRA summary comparison
|
||||
- `fastvideo/training/trackers.py` — tracker summary generation
|
||||
|
||||
## Changelog
|
||||
| Date | Change |
|
||||
|------|--------|
|
||||
| 2026-03-02 | Initial version |
|
||||
@@ -1,54 +0,0 @@
|
||||
---
|
||||
description: How to develop, validate, and register a new evaluation metric
|
||||
---
|
||||
|
||||
# Evaluation Development SOP
|
||||
|
||||
Standard procedure for adding new video quality evaluation metrics to the
|
||||
FastVideo agent toolkit.
|
||||
|
||||
## When to Use
|
||||
|
||||
- You need a metric that doesn't exist in `.agents/memory/evaluation-registry/README.md`.
|
||||
- An existing metric needs significant changes to its methodology.
|
||||
- You're exploring a new evaluation approach.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Research
|
||||
|
||||
- Search `.agents/memory/related-work/` for existing evaluation approaches.
|
||||
- Check the `evaluation_registry.md` for current metrics and their limitations.
|
||||
- Review literature: FVD, CLIP-Score, human preference, etc.
|
||||
|
||||
### 2. Prototype
|
||||
|
||||
- Write a standalone script in `.agents/exploration/<metric-name>.md`.
|
||||
- Keep it simple: one script, minimal dependencies.
|
||||
- Test on a few known-good and known-bad video samples.
|
||||
|
||||
### 3. Validate
|
||||
|
||||
- **Known-good test**: Metric should score high on reference-quality videos.
|
||||
- **Known-bad test**: Metric should score low on degraded/unrelated videos.
|
||||
- **Sensitivity test**: Small quality differences should produce meaningful
|
||||
score differences.
|
||||
- Document thresholds and their justification.
|
||||
|
||||
### 4. Register
|
||||
|
||||
Update `.agents/memory/evaluation-registry/README.md`:
|
||||
- Add the metric with status `Active`.
|
||||
- Document location, thresholds, and trust level.
|
||||
|
||||
### 5. Integrate
|
||||
|
||||
Update `.agents/skills/evaluate-video-quality.md`:
|
||||
- Add the new metric as a section.
|
||||
- Include code examples and interpretation guide.
|
||||
|
||||
### 6. Document
|
||||
|
||||
- Move the exploration log content into the skill.
|
||||
- Clean up the exploration file or mark it as `promoted`.
|
||||
- If anything went wrong during development, create a lesson.
|
||||
@@ -1,47 +0,0 @@
|
||||
---
|
||||
description: When and how to log experiments in the experiment journal
|
||||
---
|
||||
|
||||
# Experiment Journaling SOP
|
||||
|
||||
Ensures every experiment is properly recorded with context and outcomes.
|
||||
|
||||
## When to Log
|
||||
|
||||
**Always.** Every experiment — even quick tests — should be journaled.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Before Launch — Create Draft Entry
|
||||
|
||||
Use the `log-experiment` skill with `status: running`:
|
||||
- Include hypothesis and config.
|
||||
- Leave metrics, duration, and insight blank.
|
||||
|
||||
### 2. After 30-Minute Check — Update with Initial Metrics
|
||||
|
||||
Update the entry with:
|
||||
- Current loss and its trajectory direction.
|
||||
- Step time.
|
||||
- Number of validation videos generated.
|
||||
- Preliminary go/no-go assessment.
|
||||
|
||||
### 3. On Completion — Fill Final Entry
|
||||
|
||||
Update the entry with `status: completed`:
|
||||
- Final loss, grad norm, avg step time.
|
||||
- Total duration and steps.
|
||||
- Checkpoint path.
|
||||
- Key insight.
|
||||
|
||||
### 4. On Failure — Document Failure Mode
|
||||
|
||||
Update the entry with `status: failed`:
|
||||
- What went wrong (OOM, NaN, crash, etc.).
|
||||
- At what step the failure occurred.
|
||||
- Create a lesson in `.agents/lessons/` for non-trivial failures.
|
||||
|
||||
### 5. Cross-Reference
|
||||
|
||||
- Link related lessons: `**Related lessons**: .agents/lessons/<filename>.md`
|
||||
- Link related experiments: if this is a follow-up, reference the prior entry.
|
||||
@@ -1,87 +0,0 @@
|
||||
---
|
||||
description: End-to-end experiment lifecycle from hypothesis to lessons learned
|
||||
---
|
||||
|
||||
# Experiment Lifecycle SOP
|
||||
|
||||
Standard operating procedure for running ML training experiments on
|
||||
FastVideo-WorldModel. Every experiment should follow this flow.
|
||||
|
||||
## Overview
|
||||
|
||||
```
|
||||
Plan → Launch → Monitor → Summarize → Journal → Reflect
|
||||
```
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Plan the Experiment
|
||||
|
||||
Before launching:
|
||||
- [ ] Define a clear **hypothesis** (what you expect to learn).
|
||||
- [ ] Select the **model** and **pipeline** type (finetune, distill, lora, etc.).
|
||||
- [ ] Prepare the **dataset** (preprocessed into parquet format).
|
||||
- [ ] Review existing experiments in `.agents/memory/experiment-journal/README.md` for related work.
|
||||
- [ ] Check `.agents/lessons/` for known pitfalls with this configuration.
|
||||
- [ ] Document the plan in the experiment journal as a draft entry.
|
||||
|
||||
### 2. Launch the Experiment
|
||||
|
||||
Use the `launch-experiment` skill:
|
||||
- Provide: pipeline, model, data_path, num_gpus, and any hyperparameter overrides.
|
||||
- The skill generates the `torchrun` command and creates a journal entry.
|
||||
- Verify the command looks correct before executing.
|
||||
|
||||
Reference: `.agents/skills/launch-experiment.md`
|
||||
|
||||
### 3. Monitor the Experiment
|
||||
|
||||
Use the `monitor-experiment` skill:
|
||||
- Provide the W&B run ID (or output_dir for offline).
|
||||
- Monitor alerts: loss spikes, NaN gradients, step time regressions.
|
||||
- At the **30-minute mark**: perform the quality check.
|
||||
- Is loss decreasing?
|
||||
- Are validation videos reasonable?
|
||||
- Is step time consistent?
|
||||
- **Decision point**: Continue or abort based on the 30-min check.
|
||||
|
||||
Reference: `.agents/skills/monitor-experiment.md`
|
||||
|
||||
### 4. Summarize the Run
|
||||
|
||||
After completion (or at any checkpoint), use the `summarize-run` skill:
|
||||
- Extract final metrics from W&B summary.
|
||||
- Compare against reference runs if available.
|
||||
- Generate a structured report.
|
||||
|
||||
Reference: `.agents/skills/summarize-run.md`
|
||||
|
||||
### 5. Update the Experiment Journal
|
||||
|
||||
Use the `log-experiment` skill to update the journal entry:
|
||||
- Fill in final metrics, duration, checkpoint paths.
|
||||
- Record the key insight learned.
|
||||
- Set status to `completed`, `failed`, or `abandoned`.
|
||||
|
||||
Reference: `.agents/skills/log-experiment.md`
|
||||
|
||||
### 6. Reflect and Capture Lessons
|
||||
|
||||
After every experiment:
|
||||
- **What went right?** → Note in the journal insight field.
|
||||
- **What went wrong?** → Create a lesson in `.agents/lessons/`:
|
||||
- Use the template in `.agents/lessons/README.md`.
|
||||
- Cross-reference the experiment journal entry.
|
||||
- **What was surprising?** → Consider creating an exploration log if this
|
||||
warrants further investigation.
|
||||
|
||||
Reference: `.agents/workflows/lesson-capture.md`
|
||||
|
||||
## Validation Criteria
|
||||
|
||||
This SOP is validated when an agent can:
|
||||
1. Follow steps 1–6 end-to-end for a minimal training run
|
||||
(e.g., `examples/training/finetune/wan_t2v_1.3B/crush_smol/finetune_t2v.sh`
|
||||
with `--max_train_steps 5`).
|
||||
2. Produce a complete experiment journal entry.
|
||||
3. Generate a run summary report.
|
||||
@@ -1,71 +0,0 @@
|
||||
---
|
||||
description: Post-experiment reflection to capture lessons learned
|
||||
---
|
||||
|
||||
# Lesson Capture SOP
|
||||
|
||||
Systematic procedure for turning experiment outcomes into persistent knowledge.
|
||||
|
||||
## When to Use
|
||||
|
||||
After **every** completed or failed experiment. Even successful experiments
|
||||
can yield lessons (e.g., "LR 5e-5 works better than 1e-5 for LoRA").
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Review the Experiment
|
||||
|
||||
Read the experiment journal entry. Ask:
|
||||
- Did anything go wrong?
|
||||
- Was anything surprising?
|
||||
- Did anything take longer than expected?
|
||||
- Was a workaround needed?
|
||||
|
||||
### 2. Decide: Lesson or Not?
|
||||
|
||||
| Situation | Action |
|
||||
|-----------|--------|
|
||||
| Something broke | Create a lesson (category: `infrastructure` or `data`) |
|
||||
| Hyperparameter choice mattered | Create a lesson (category: `hyperparameter`) |
|
||||
| Porting issue found | Create a lesson (category: `porting`) |
|
||||
| Evaluation metric was misleading | Create a lesson (category: `evaluation`) |
|
||||
| Everything went smoothly | No lesson needed, but note in the journal insight |
|
||||
|
||||
### 3. Create the Lesson File
|
||||
|
||||
In `.agents/lessons/`, create `<YYYY-MM-DD>_<short-slug>.md`:
|
||||
|
||||
```markdown
|
||||
---
|
||||
date: <ISO-8601>
|
||||
experiment: <journal entry reference>
|
||||
category: hyperparameter | data | infrastructure | evaluation | porting
|
||||
severity: critical | important | minor
|
||||
---
|
||||
|
||||
# <Short Descriptive Title>
|
||||
|
||||
## What Happened
|
||||
<description>
|
||||
|
||||
## Root Cause
|
||||
<analysis>
|
||||
|
||||
## Fix / Workaround
|
||||
<resolution>
|
||||
|
||||
## Prevention
|
||||
<how to avoid in future>
|
||||
```
|
||||
|
||||
### 4. Cross-Reference
|
||||
|
||||
- Update the experiment journal entry with a link to the lesson file.
|
||||
- If a similar lesson already exists, add a reference or update it.
|
||||
|
||||
### 5. Periodic Pattern Review
|
||||
|
||||
Every ~10 lessons, scan for patterns:
|
||||
- Multiple lessons in the same category → consider a new skill or SOP.
|
||||
- Repeated mistakes → strengthen the relevant SOP with a checklist item.
|
||||
- Infrastructure issues → propose a codebase fix.
|
||||
@@ -1,67 +0,0 @@
|
||||
---
|
||||
description: Synchronize the STATUS.md dashboard by scanning .agents/ directories
|
||||
---
|
||||
|
||||
# Sync Dashboard
|
||||
|
||||
Updates `.agents/STATUS.md` by scanning the skills, workflows, memory, lessons,
|
||||
and exploration directories to reflect what actually exists on disk.
|
||||
|
||||
## When to Use
|
||||
|
||||
- After adding, removing, or renaming any file in `.agents/`.
|
||||
- Periodically (e.g., at end of each conversation session).
|
||||
- When the dashboard feels out of date.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Scan directories
|
||||
|
||||
List all files in each directory:
|
||||
|
||||
```bash
|
||||
echo "=== Skills ==="
|
||||
ls -1 .agents/skills/*.md 2>/dev/null | grep -v SKILL_TEMPLATE
|
||||
|
||||
echo "=== Workflows ==="
|
||||
ls -1 .agents/workflows/*.md 2>/dev/null
|
||||
|
||||
echo "=== Memory ==="
|
||||
ls -1 .agents/memory/*.md 2>/dev/null
|
||||
ls -1 .agents/memory/related-work/*.md 2>/dev/null | grep -v README
|
||||
|
||||
echo "=== Lessons ==="
|
||||
ls -1 .agents/lessons/*.md 2>/dev/null | grep -v README
|
||||
|
||||
echo "=== Exploration ==="
|
||||
ls -1 .agents/exploration/*.md 2>/dev/null | grep -v README
|
||||
```
|
||||
|
||||
### 2. Compare with STATUS.md
|
||||
|
||||
For each file found:
|
||||
- If it's in STATUS.md → leave it (preserve status/trust/tested fields).
|
||||
- If it's NOT in STATUS.md → add it with status `🔴 Stub`, trust `None`, tested `❌`.
|
||||
|
||||
For each entry in STATUS.md:
|
||||
- If the file no longer exists → mark it as `❌ Removed` or delete the row.
|
||||
|
||||
### 3. Update counts
|
||||
|
||||
Recalculate the summary table at the top:
|
||||
- Count files per category.
|
||||
- Count by status (Ready, Draft, Stub).
|
||||
|
||||
### 4. Update timestamp
|
||||
|
||||
Set `_Last synced: <current date>_` at the top of STATUS.md.
|
||||
|
||||
### 5. Review
|
||||
|
||||
Read through the updated STATUS.md for accuracy. Flag anything that looks wrong.
|
||||
|
||||
## Notes
|
||||
|
||||
- Do NOT change trust levels during sync — those are set manually after testing.
|
||||
- Do NOT change status during sync — status changes require actual validation.
|
||||
- This workflow only handles structural sync (file existence), not content review.
|
||||
@@ -1,46 +0,0 @@
|
||||
{
|
||||
"benchmark_id": "wan-t2v-1.3b-2gpu",
|
||||
"description": "Wan2.1 T2V 1.3B inference performance",
|
||||
"model": {
|
||||
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
"model_short_name": "Wan2.1-T2V-1.3B"
|
||||
},
|
||||
"init_kwargs": {
|
||||
"num_gpus": 2,
|
||||
"flow_shift": 7.0,
|
||||
"sp_size": 2,
|
||||
"tp_size": 1,
|
||||
"vae_sp": true,
|
||||
"vae_tiling": true,
|
||||
"text_encoder_precisions": ["fp32"]
|
||||
},
|
||||
"generation_kwargs": {
|
||||
"height": 480,
|
||||
"width": 832,
|
||||
"num_frames": 45,
|
||||
"num_inference_steps": 4,
|
||||
"guidance_scale": 3,
|
||||
"embedded_cfg_scale": 6,
|
||||
"seed": 1024,
|
||||
"fps": 24,
|
||||
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
|
||||
},
|
||||
"test_prompts": [
|
||||
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
|
||||
],
|
||||
"run_config": {
|
||||
"num_warmup_runs": 1,
|
||||
"num_measurement_runs": 3,
|
||||
"required_gpus": 2
|
||||
},
|
||||
"thresholds": {
|
||||
"L40S": {
|
||||
"max_generation_time_s": 34.0,
|
||||
"max_peak_memory_mb": 11000.0
|
||||
},
|
||||
"default": {
|
||||
"max_generation_time_s": 120.0,
|
||||
"max_peak_memory_mb": 30000.0
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -61,7 +61,7 @@ steps:
|
||||
- "pyproject.toml"
|
||||
- "docker/Dockerfile.python3.12"
|
||||
config:
|
||||
command: "timeout 90m .buildkite/scripts/pr_test.sh"
|
||||
command: "timeout 60m .buildkite/scripts/pr_test.sh"
|
||||
label: "SSIM Tests"
|
||||
env:
|
||||
- TEST_TYPE=ssim
|
||||
@@ -139,6 +139,18 @@ steps:
|
||||
- TEST_TYPE=training_vsa
|
||||
agents:
|
||||
queue: "default"
|
||||
- path:
|
||||
- "fastvideo/**"
|
||||
- "fastvideo-kernel/**"
|
||||
- "pyproject.toml"
|
||||
- "docker/Dockerfile.python3.12"
|
||||
config:
|
||||
command: "timeout 15m .buildkite/scripts/pr_test.sh"
|
||||
label: "Inference Tests STA"
|
||||
env:
|
||||
- TEST_TYPE=inference_sta
|
||||
agents:
|
||||
queue: "default"
|
||||
- path:
|
||||
- "fastvideo-kernel/**"
|
||||
- "pyproject.toml"
|
||||
@@ -171,37 +183,6 @@ steps:
|
||||
- TEST_TYPE=unit_test
|
||||
agents:
|
||||
queue: "default"
|
||||
- path:
|
||||
- "fastvideo/models/dits/**"
|
||||
- "fastvideo/pipelines/**"
|
||||
- "fastvideo/attention/**"
|
||||
- "fastvideo/layers/**"
|
||||
- "fastvideo/worker/**"
|
||||
- "fastvideo/entrypoints/**"
|
||||
- "fastvideo/tests/performance/**"
|
||||
- ".buildkite/performance-benchmarks/**"
|
||||
- "pyproject.toml"
|
||||
- "docker/Dockerfile.python3.12"
|
||||
config:
|
||||
command: "timeout 30m .buildkite/scripts/pr_test.sh"
|
||||
label: "Performance Tests"
|
||||
env:
|
||||
- TEST_TYPE=performance
|
||||
agents:
|
||||
queue: "default"
|
||||
- path:
|
||||
- "fastvideo/entrypoints/openai/**"
|
||||
- "fastvideo/entrypoints/cli/serve.py"
|
||||
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
|
||||
- "pyproject.toml"
|
||||
- "docker/Dockerfile.python3.12"
|
||||
config:
|
||||
command: "timeout 30m .buildkite/scripts/pr_test.sh"
|
||||
label: "API Server Tests"
|
||||
env:
|
||||
- TEST_TYPE=api_server
|
||||
agents:
|
||||
queue: "default"
|
||||
# - path:
|
||||
# - "scripts/lora_extraction/**"
|
||||
# - "pyproject.toml"
|
||||
|
||||
@@ -51,7 +51,6 @@ else
|
||||
fi
|
||||
|
||||
MODAL_TEST_FILE="fastvideo/tests/modal/pr_test.py"
|
||||
MODAL_SSIM_TEST_FILE="fastvideo/tests/modal/ssim_test.py"
|
||||
|
||||
if [ -z "${TEST_TYPE:-}" ]; then
|
||||
log "Error: TEST_TYPE environment variable is not set"
|
||||
@@ -76,7 +75,7 @@ case "$TEST_TYPE" in
|
||||
;;
|
||||
"ssim")
|
||||
log "Running SSIM tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_SSIM_TEST_FILE::run_ssim_tests"
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_ssim_tests"
|
||||
;;
|
||||
"training")
|
||||
log "Running training tests..."
|
||||
@@ -90,6 +89,10 @@ case "$TEST_TYPE" in
|
||||
log "Running training VSA tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV WANDB_API_KEY=$WANDB_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_training_tests_VSA"
|
||||
;;
|
||||
"inference_sta")
|
||||
log "Running inference STA tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_inference_tests_STA"
|
||||
;;
|
||||
"kernel_tests")
|
||||
log "Running kernel tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_kernel_tests"
|
||||
@@ -119,14 +122,6 @@ case "$TEST_TYPE" in
|
||||
log "Running LoRA extraction tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_lora_extraction_tests"
|
||||
;;
|
||||
"performance")
|
||||
log "Running performance tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_performance_tests"
|
||||
;;
|
||||
"api_server")
|
||||
log "Running API server integration tests..."
|
||||
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_api_server_tests"
|
||||
;;
|
||||
*)
|
||||
log "Error: Unknown test type: $TEST_TYPE"
|
||||
exit 1
|
||||
|
||||
@@ -1,43 +0,0 @@
|
||||
## Purpose
|
||||
|
||||
<!-- What does this PR do? Link the related issue if applicable. -->
|
||||
|
||||
Fixes #
|
||||
|
||||
## Changes
|
||||
|
||||
<!-- Describe your changes concisely. What approach did you take? -->
|
||||
|
||||
-
|
||||
|
||||
## Test Plan
|
||||
|
||||
<!-- How did you verify your changes? Paste exact commands and output. -->
|
||||
|
||||
```bash
|
||||
# Commands you ran
|
||||
```
|
||||
|
||||
## Test Results
|
||||
|
||||
<!-- Paste test output, before/after comparisons, or SSIM scores for model changes. -->
|
||||
|
||||
<details>
|
||||
<summary>Test output</summary>
|
||||
|
||||
```
|
||||
# Paste output here
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] I ran `pre-commit run --all-files` and fixed all issues
|
||||
- [ ] I added or updated tests for my changes
|
||||
- [ ] I updated documentation if needed
|
||||
- [ ] I considered GPU memory impact of my changes
|
||||
|
||||
**For model/pipeline changes, also check:**
|
||||
- [ ] I verified SSIM regression tests pass
|
||||
- [ ] I updated the support matrix if adding a new model
|
||||
@@ -45,12 +45,6 @@ jobs:
|
||||
- name: Setup Pages
|
||||
uses: actions/configure-pages@v4
|
||||
|
||||
- name: Generate docs examples
|
||||
run: python docs/generate_examples.py
|
||||
|
||||
- name: Check docs links
|
||||
run: python scripts/check_docs_links.py
|
||||
|
||||
- name: Build documentation
|
||||
run: mkdocs build
|
||||
|
||||
@@ -69,4 +63,4 @@ jobs:
|
||||
steps:
|
||||
- name: Deploy to GitHub Pages
|
||||
id: deployment
|
||||
uses: actions/deploy-pages@v4
|
||||
uses: actions/deploy-pages@v4
|
||||
@@ -64,10 +64,7 @@ jobs:
|
||||
# - torch-version: '2.7.1'
|
||||
# cuda-version: '12.8.0'
|
||||
# torch-cuda-short: 'cu128'
|
||||
# - torch-version: '2.9.1'
|
||||
# cuda-version: '12.8.0'
|
||||
# torch-cuda-short: 'cu128'
|
||||
- torch-version: '2.10.0'
|
||||
- torch-version: '2.9.1'
|
||||
cuda-version: '12.8.0'
|
||||
torch-cuda-short: 'cu128'
|
||||
|
||||
@@ -159,23 +156,8 @@ jobs:
|
||||
|
||||
# Fix the wheel to be manylinux compliant
|
||||
pip install auditwheel
|
||||
# Point auditwheel at torch libs, but do not vendor them into the wheel.
|
||||
TORCH_LIB_DIR=$(python - <<'PY'
|
||||
import os
|
||||
import torch
|
||||
|
||||
print(os.path.join(os.path.dirname(torch.__file__), "lib"))
|
||||
PY
|
||||
)
|
||||
export LD_LIBRARY_PATH="${TORCH_LIB_DIR}:${LD_LIBRARY_PATH}"
|
||||
# Target manylinux_2_35 (Ubuntu 22.04 native)
|
||||
auditwheel repair dist/*.whl --plat manylinux_2_35_x86_64 -w fixed_dist \
|
||||
--exclude libtorch_cuda.so \
|
||||
--exclude libtorch_cpu.so \
|
||||
--exclude libtorch.so \
|
||||
--exclude libc10.so \
|
||||
--exclude libc10_cuda.so \
|
||||
--exclude libtorch_python.so
|
||||
auditwheel repair dist/*.whl --plat manylinux_2_35_x86_64 -w fixed_dist
|
||||
# Move fixed wheels back to dist for upload consistency
|
||||
rm dist/*.whl
|
||||
mv fixed_dist/*.whl dist/
|
||||
|
||||
@@ -47,6 +47,16 @@ on:
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
run_inference_test_STA:
|
||||
description: "Run inference-test-STA"
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
run_precision_test_STA:
|
||||
description: "Run precision-test-STA"
|
||||
required: false
|
||||
default: false
|
||||
type: boolean
|
||||
run_precision_test_VSA:
|
||||
description: "Run precision-test-VSA"
|
||||
required: false
|
||||
@@ -80,6 +90,8 @@ jobs:
|
||||
transformer-test: ${{ steps.filter.outputs.transformer-test }}
|
||||
training-test: ${{ steps.filter.outputs.training-test }}
|
||||
training-test-VSA: ${{ steps.filter.outputs.training-test-VSA }}
|
||||
inference-test-STA: ${{ steps.filter.outputs.inference-test-STA }}
|
||||
precision-test-STA: ${{ steps.filter.outputs.precision-test-STA }}
|
||||
precision-test-VSA: ${{ steps.filter.outputs.precision-test-VSA }}
|
||||
unit-test: ${{ steps.filter.outputs.unit-test }}
|
||||
steps:
|
||||
@@ -94,6 +106,12 @@ jobs:
|
||||
- 'docker/Dockerfile.python3.10'
|
||||
- 'docker/Dockerfile.python3.11'
|
||||
- 'docker/Dockerfile.python3.12'
|
||||
sta-kernel-paths: &sta-kernel-paths
|
||||
- 'csrc/attn/sliding_tile_attn/**'
|
||||
- 'csrc/attn/sliding_tile_attn/tk/**'
|
||||
- 'csrc/attn/sliding_tile_attn/setup.py'
|
||||
- 'csrc/attn/sliding_tile_attn/config_sta.py'
|
||||
- 'csrc/attn/sliding_tile_attn/st_attn.cpp'
|
||||
vsa-kernel-paths: &vsa-kernel-paths
|
||||
- 'csrc/attn/video_sparse_attn/**'
|
||||
- 'csrc/attn/video_sparse_attn/tk/**'
|
||||
@@ -130,6 +148,13 @@ jobs:
|
||||
- 'fastvideo/**'
|
||||
- *common-paths
|
||||
- *vsa-kernel-paths
|
||||
inference-test-STA:
|
||||
- 'fastvideo/**'
|
||||
- *common-paths
|
||||
- *sta-kernel-paths
|
||||
precision-test-STA:
|
||||
- *common-paths
|
||||
- *sta-kernel-paths
|
||||
precision-test-VSA:
|
||||
- *common-paths
|
||||
- *vsa-kernel-paths
|
||||
@@ -257,6 +282,44 @@ jobs:
|
||||
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
|
||||
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
|
||||
|
||||
inference-test-STA:
|
||||
needs: change-filter
|
||||
if: >-
|
||||
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.inference-test-STA == 'true') ||
|
||||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_inference_test_STA == 'true')
|
||||
uses: ./.github/workflows/runpod-test.yml
|
||||
with:
|
||||
job_id: "inference-test-STA"
|
||||
gpu_type: "NVIDIA H100 NVL"
|
||||
gpu_count: 2
|
||||
volume_size: 100
|
||||
disk_size: 100
|
||||
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
|
||||
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/inference/STA -srP"
|
||||
timeout_minutes: 30
|
||||
secrets:
|
||||
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
|
||||
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
|
||||
|
||||
precision-test-STA:
|
||||
needs: change-filter
|
||||
if: >-
|
||||
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.precision-test-STA == 'true') ||
|
||||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_precision_test_STA == 'true')
|
||||
uses: ./.github/workflows/runpod-test.yml
|
||||
with:
|
||||
job_id: "precision-test-STA"
|
||||
gpu_type: "NVIDIA H100 NVL"
|
||||
gpu_count: 1
|
||||
volume_size: 100
|
||||
disk_size: 100
|
||||
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
|
||||
test_command: "uv pip install -e .[test] && python csrc/attn/tests/test_sta.py"
|
||||
timeout_minutes: 30
|
||||
secrets:
|
||||
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
|
||||
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
|
||||
|
||||
precision-test-VSA:
|
||||
needs: change-filter
|
||||
if: >-
|
||||
@@ -315,7 +378,7 @@ jobs:
|
||||
|
||||
runpod-cleanup:
|
||||
# Add other jobs to this list as you create them
|
||||
needs: [encoder-test, vae-test, transformer-test, ssim-test, training-test, training-test-VSA, precision-test-VSA]
|
||||
needs: [encoder-test, vae-test, transformer-test, ssim-test, training-test, training-test-VSA, inference-test-STA, precision-test-STA, precision-test-VSA]
|
||||
if: ${{ always() && ((github.event_name != 'workflow_dispatch' && github.event.pull_request.draft == false) || github.event_name == 'workflow_dispatch') }}
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
@@ -332,7 +395,7 @@ jobs:
|
||||
|
||||
- name: Cleanup all RunPod instances
|
||||
env:
|
||||
JOB_IDS: '["encoder-test", "vae-test", "transformer-test", "ssim-test-py3.10", "ssim-test-py3.11", "ssim-test-py3.12", "training-test", "training-test-VSA", "precision-test-VSA"]'
|
||||
JOB_IDS: '["encoder-test", "vae-test", "transformer-test", "ssim-test-py3.10", "ssim-test-py3.11", "ssim-test-py3.12", "training-test", "training-test-VSA", "inference-test-STA", "precision-test-STA", "precision-test-VSA"]'
|
||||
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
|
||||
GITHUB_RUN_ID: ${{ github.run_id }}
|
||||
run: python .github/scripts/runpod_cleanup.py
|
||||
|
||||
@@ -0,0 +1,249 @@
|
||||
name: Publish Sliding Tile Attention Kernel to PyPI on Version Change
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- "csrc/attn/sliding_tile_attn/setup.py"
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
check-version-change:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
version-changed: ${{ steps.check-version.outputs.changed }}
|
||||
new-version: ${{ steps.check-version.outputs.new-version }}
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 2
|
||||
|
||||
- name: Check if version changed
|
||||
id: check-version
|
||||
run: |
|
||||
cd csrc/attn/sliding_tile_attn
|
||||
# Get current commit's version
|
||||
NEW_VERSION=$(grep -oP 'VERSION\s*=\s*"\K[^"]+' setup.py)
|
||||
echo "New version: $NEW_VERSION"
|
||||
|
||||
# Get previous version from git history
|
||||
OLD_VERSION=$(git show HEAD~1:./setup.py | grep -oP 'VERSION\s*=\s*"\K[^"]+' || echo "0.0.0")
|
||||
echo "Old version: $OLD_VERSION"
|
||||
|
||||
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
|
||||
echo "Version changed from $OLD_VERSION to $NEW_VERSION"
|
||||
echo "changed=true" >> $GITHUB_OUTPUT
|
||||
echo "new-version=$NEW_VERSION" >> $GITHUB_OUTPUT
|
||||
else
|
||||
echo "Version did not change"
|
||||
echo "changed=false" >> $GITHUB_OUTPUT
|
||||
fi
|
||||
|
||||
build_wheels:
|
||||
name: Build Wheel
|
||||
needs: check-version-change
|
||||
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
|
||||
runs-on: ${{ matrix.os }}
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
# Using ubuntu-20.04 instead of 22.04 for more compatibility (glibc). Ideally we'd use the
|
||||
# manylinux docker image, but I haven't figured out how to install CUDA on manylinux.
|
||||
os: [ubuntu-22.04]
|
||||
python-version: ['3.10', '3.11', '3.12', '3.13']
|
||||
torch-version: ['2.5.1', '2.6.0']
|
||||
cuda-version: ['12.4.1', '12.5.1', '12.6.3']
|
||||
|
||||
steps:
|
||||
- name: Free up disk space
|
||||
run: |
|
||||
echo "Initial disk space:"
|
||||
df -h
|
||||
|
||||
# Remove large directories
|
||||
sudo rm -rf /usr/share/dotnet
|
||||
sudo rm -rf /usr/local/lib/android
|
||||
sudo rm -rf /opt/ghc
|
||||
sudo rm -rf /usr/local/share/boost
|
||||
sudo rm -rf /usr/share/swift
|
||||
sudo rm -rf /usr/local/lib/node_modules
|
||||
sudo rm -rf /usr/local/share/powershell
|
||||
sudo rm -rf /usr/share/rust
|
||||
sudo rm -rf /usr/local/.ghcup
|
||||
|
||||
# Remove cached files
|
||||
sudo rm -rf /var/lib/apt/lists/*
|
||||
sudo rm -rf /var/cache/apt/archives/*
|
||||
|
||||
echo "Disk space after cleanup:"
|
||||
df -h
|
||||
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Set up Python
|
||||
uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: ${{ matrix.python-version }}
|
||||
|
||||
- name: Install CUDA ${{ matrix.cuda-version }}
|
||||
uses: Jimver/cuda-toolkit@v0.2.21
|
||||
id: cuda-toolkit
|
||||
with:
|
||||
cuda: ${{ matrix.cuda-version }}
|
||||
linux-local-args: '["--toolkit"]'
|
||||
method: 'network'
|
||||
|
||||
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
|
||||
run: |
|
||||
sudo apt update
|
||||
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
|
||||
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
|
||||
|
||||
# Allow Git to Access Safe Directory
|
||||
git config --global --add safe.directory /__w/FastVideo/FastVideo
|
||||
|
||||
# Set CUDA environment variables
|
||||
export CUDA_HOME=/usr/local/cuda-${{ matrix.cuda-version }}
|
||||
export PATH=${CUDA_HOME}/bin:${PATH}
|
||||
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
|
||||
|
||||
# Verify installation
|
||||
gcc --version
|
||||
g++ --version
|
||||
clang-11 --version
|
||||
nvcc --version
|
||||
|
||||
- name: Install PyTorch ${{ matrix.torch-version }}+cu${{ matrix.cuda-version }}
|
||||
run: |
|
||||
pip install --upgrade pip
|
||||
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
|
||||
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
|
||||
pip install typing-extensions==4.12.2
|
||||
# We want to figure out the CUDA version to download pytorch
|
||||
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
|
||||
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
|
||||
export TORCH_CUDA_VERSION=124
|
||||
pip install --no-cache-dir torch==${{ matrix.torch-version }} --index-url https://download.pytorch.org/whl/cu${TORCH_CUDA_VERSION}
|
||||
nvcc --version
|
||||
python --version
|
||||
python -c "import torch; print('PyTorch:', torch.__version__)"
|
||||
python -c "import torch; print('CUDA:', torch.version.cuda)"
|
||||
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
|
||||
|
||||
- name: Build wheel
|
||||
run: |
|
||||
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
|
||||
|
||||
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
|
||||
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
|
||||
# However this still fails so I'm using a newer version of setuptools
|
||||
pip install setuptools
|
||||
pip install ninja packaging wheel
|
||||
|
||||
cd csrc/attn/sliding_tile_attn # Move into the correct folder
|
||||
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
|
||||
python setup.py bdist_wheel --dist-dir=dist
|
||||
|
||||
- name: Rename wheel file
|
||||
run: |
|
||||
cd csrc/attn/sliding_tile_attn
|
||||
|
||||
CUDA_SHORT_VERSION=$(echo ${{ matrix.cuda-version }} | cut -d. -f1,2 | sed 's/\.//g')
|
||||
TORCH_SHORT_VERSION=$(echo ${{ matrix.torch-version }} | cut -d. -f1,2)
|
||||
# Get the correct version format
|
||||
tmpname=cu${CUDA_SHORT_VERSION}torch${TORCH_SHORT_VERSION}
|
||||
wheel_name=$(ls dist/*whl | xargs -n 1 basename | sed "s/-/+$tmpname-/2")
|
||||
# Rename with version information
|
||||
ls dist/*whl |xargs -I {} mv {} dist/${wheel_name}
|
||||
echo "wheel_name=${wheel_name}" >> $GITHUB_ENV
|
||||
|
||||
- name: Upload wheel artifact
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: ${{ env.wheel_name }}
|
||||
path: csrc/attn/sliding_tile_attn/dist/*.whl
|
||||
retention-days: 90
|
||||
|
||||
publish_package:
|
||||
name: Publish package
|
||||
needs: [build_wheels, check-version-change]
|
||||
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
|
||||
runs-on: ubuntu-22.04
|
||||
permissions:
|
||||
id-token: write # Needed for OIDC Trusted Publishing
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: '3.10'
|
||||
|
||||
- name: Install CUDA 12.4.1
|
||||
uses: Jimver/cuda-toolkit@v0.2.21
|
||||
id: cuda-toolkit
|
||||
with:
|
||||
cuda: 12.4.1
|
||||
linux-local-args: '["--toolkit"]'
|
||||
method: 'network'
|
||||
sub-packages: '["nvcc"]'
|
||||
|
||||
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
|
||||
run: |
|
||||
sudo apt update
|
||||
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
|
||||
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
|
||||
|
||||
# Allow Git to Access Safe Directory
|
||||
git config --global --add safe.directory /__w/FastVideo/FastVideo
|
||||
|
||||
# Set CUDA environment variables
|
||||
export CUDA_HOME=/usr/local/cuda-12.4.1
|
||||
export PATH=${CUDA_HOME}/bin:${PATH}
|
||||
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
|
||||
|
||||
# Verify installation
|
||||
gcc --version
|
||||
g++ --version
|
||||
clang-11 --version
|
||||
nvcc --version
|
||||
|
||||
- name: Install PyTorch 2.5.1+cu12.4.1
|
||||
run: |
|
||||
pip install --upgrade pip
|
||||
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
|
||||
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
|
||||
pip install typing-extensions==4.12.2
|
||||
# We want to figure out the CUDA version to download pytorch
|
||||
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
|
||||
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
|
||||
export TORCH_CUDA_VERSION=124
|
||||
pip install --no-cache-dir torch==2.5.1 --index-url https://download.pytorch.org/whl/cu${TORCH_CUDA_VERSION}
|
||||
nvcc --version
|
||||
python --version
|
||||
python -c "import torch; print('PyTorch:', torch.__version__)"
|
||||
python -c "import torch; print('CUDA:', torch.version.cuda)"
|
||||
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
|
||||
|
||||
- name: Build source distribution
|
||||
run: |
|
||||
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
|
||||
|
||||
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
|
||||
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
|
||||
# However this still fails so I'm using a newer version of setuptools
|
||||
pip install setuptools
|
||||
pip install ninja packaging wheel
|
||||
|
||||
cd csrc/attn/sliding_tile_attn # Move into the correct folder
|
||||
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
|
||||
python setup.py sdist --dist-dir=dist
|
||||
|
||||
- name: Publish release distributions to PyPI
|
||||
uses: pypa/gh-action-pypi-publish@release/v1
|
||||
with:
|
||||
packages-dir: csrc/attn/sliding_tile_attn/dist/
|
||||
@@ -18,7 +18,6 @@ venv/
|
||||
.venv/
|
||||
runs/
|
||||
samples/
|
||||
Miniconda3-latest-Linux-x86_64.sh
|
||||
*validation/
|
||||
data/
|
||||
outputs/
|
||||
@@ -34,11 +33,6 @@ env
|
||||
*.log
|
||||
weights/
|
||||
|
||||
# SSIM test outputs
|
||||
fastvideo/tests/ssim/generated_videos/
|
||||
**/.cache/**
|
||||
|
||||
|
||||
# Distribution / packaging
|
||||
build/
|
||||
dist/
|
||||
@@ -75,15 +69,6 @@ docs/distillation/examples/
|
||||
!docs/assets/images/**/*.png
|
||||
!comfyui/assets/**/*.png
|
||||
!comfyui/assets/**/*.gif
|
||||
!assets/images/**/*.png
|
||||
!assets/images/**/*.jpg
|
||||
!assets/images/**/*.jpeg
|
||||
!assets/images/**/*.gif
|
||||
!assets/videos/**/*.mp4
|
||||
|
||||
dmd_t2v_output/
|
||||
preprocess_output_text/
|
||||
|
||||
.claude/
|
||||
.codex/
|
||||
openspec/
|
||||
|
||||
@@ -10,7 +10,7 @@ exclude: |
|
||||
demo/.*|
|
||||
predict\.py|
|
||||
scripts/.*|
|
||||
assets/prompts/.*|
|
||||
prompts/.*|
|
||||
fastvideo/data_preprocess/.*|
|
||||
fastvideo/dataset/.*|
|
||||
fastvideo/models/.*|
|
||||
@@ -60,7 +60,7 @@ repos:
|
||||
hooks:
|
||||
- id: mypy
|
||||
args: [--python-version, '3.10', --follow-imports, "skip", "--disable-error-code", "union-attr", "--disable-error-code", "override" ]
|
||||
additional_dependencies: [types-aiofiles, types-cachetools, types-setuptools, types-PyYAML, types-requests]
|
||||
additional_dependencies: [types-cachetools, types-setuptools, types-PyYAML, types-requests]
|
||||
- repo: local
|
||||
hooks:
|
||||
- id: check-filenames
|
||||
@@ -68,7 +68,7 @@ repos:
|
||||
entry: bash
|
||||
args:
|
||||
- -c
|
||||
- 'git ls-files | grep -v "^\"*fastvideo/tests/ssim/" | grep -v "^\"*fastvideo/tests/inference/lora/L40S_reference_videos/" | grep " " && echo "Filenames should not contain spaces!" && exit 1 || exit 0'
|
||||
- 'git ls-files | grep -v "^fastvideo/tests/ssim/" | grep -v "^fastvideo/tests/inference/lora/L40S_reference_videos/" | grep " " && echo "Filenames should not contain spaces!" && exit 1 || exit 0'
|
||||
language: system
|
||||
always_run: true
|
||||
pass_filenames: false
|
||||
|
||||
@@ -1,56 +0,0 @@
|
||||
# Repository Guidelines
|
||||
|
||||
## Project Structure & Module Organization
|
||||
- Core Python package: `fastvideo/` (models, pipelines, training, distributed runtime, CLI entrypoints).
|
||||
- CUDA/custom kernels: `fastvideo-kernel/` (separate build/test flow).
|
||||
- Tests:
|
||||
- `fastvideo/tests/` for package-level tests (dataset, encoders, inference, training, SSIM, workflow).
|
||||
- `tests/local_tests/` for additional local/component checks.
|
||||
- Docs and guides: `docs/` (MkDocs source), with contributor docs in `docs/contributing/`.
|
||||
- Runnable examples and scripts: `examples/` and `scripts/`.
|
||||
- Static assets: `assets/` (including `assets/images/`, `assets/videos/`, and `assets/prompts/`) and `comfyui/assets/`.
|
||||
|
||||
## Build, Test, and Development Commands
|
||||
- `uv pip install -e .[dev]`: editable install with lint/test extras.
|
||||
- `pre-commit install --hook-type pre-commit --hook-type commit-msg`: enable local hooks.
|
||||
- `pre-commit run --all-files`: run formatter/lint/type/spelling checks.
|
||||
- `pytest tests/`: run top-level test suite.
|
||||
- `pytest fastvideo/tests/ -v`: run package tests.
|
||||
- `pytest fastvideo/tests/ssim/ -vs`: run SSIM regression tests (GPU-heavy).
|
||||
- `cd fastvideo-kernel && ./build.sh`: build kernel extensions.
|
||||
|
||||
## Coding Style & Naming Conventions
|
||||
- Python 3.10+; 4-space indentation; keep code and imports readable and explicit.
|
||||
- Style tools are configured in `pyproject.toml` and `.pre-commit-config.yaml`:
|
||||
- `yapf` (format), `ruff` (lint, auto-fix), `mypy` (typing), `codespell`.
|
||||
- Target line length is 80.
|
||||
- Naming: `snake_case` for functions/files, `PascalCase` for classes, `UPPER_SNAKE_CASE` for constants.
|
||||
|
||||
## Testing Guidelines
|
||||
- Use `pytest` and place tests near relevant domains (e.g., `fastvideo/tests/encoders/`).
|
||||
- Prefer descriptive names like `test_<feature>_<expected_behavior>.py`.
|
||||
- For new pipelines/backends, include at least one regression-oriented test; add SSIM coverage when output quality must be preserved.
|
||||
- Document GPU assumptions in tests that require specific hardware.
|
||||
|
||||
## Commit & Pull Request Guidelines
|
||||
- Follow existing commit style: short subject with optional tag prefix, e.g. `[bugfix]: ...`, `[feat]: ...`, `[misc]: ...`, and include PR reference like `(#1234)` when applicable.
|
||||
- Keep commits focused by concern (feature, refactor, fix).
|
||||
- PRs should include:
|
||||
- clear problem/solution summary,
|
||||
- test evidence (`pytest`/SSIM outputs or rationale if skipped),
|
||||
- linked issue/PR context,
|
||||
- screenshots or sample outputs for UI/demo/docs changes.
|
||||
|
||||
## Agent Infrastructure
|
||||
|
||||
This repository is agent-friendly. Before doing any work, read:
|
||||
|
||||
1. `.agents/onboarding/README.md` — full onboarding guide with step-by-step instructions.
|
||||
2. `.agents/memory/codebase-map/README.md` — structural index of the entire repository.
|
||||
3. `.agents/skills/` — available agent skills (check if one exists before writing code).
|
||||
4. `.agents/workflows/` — SOPs for common procedures (experiment lifecycle, evaluation, etc.).
|
||||
5. `.agents/lessons/` — known pitfalls and their documented fixes.
|
||||
|
||||
If you are exploring a new procedure that has no existing SOP, document your
|
||||
progress in `.agents/exploration/` and flag it for review at the end of your
|
||||
session.
|
||||
@@ -3,32 +3,33 @@
|
||||
</div>
|
||||
|
||||
<p align="center">
|
||||
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
|
||||
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://ibb.co/sv3MMKyv" target="_blank"> <b> WeChat </b> </a> |
|
||||
</p>
|
||||
|
||||
**FastVideo is a unified post-training and inference framework for accelerated video generation.**
|
||||
|
||||
## NEWS
|
||||
- ```2025/11/19```: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py)
|
||||
- ```2025/08/04```: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
|
||||
|
||||
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py)
|
||||
- `2025/08/04`: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
|
||||
<details>
|
||||
<summary>More</summary>
|
||||
|
||||
### More News
|
||||
- ```2025/06/14```: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389)
|
||||
- ```2025/04/24```: [FastVideo V1](https://hao-ai-lab.github.io/blogs/fastvideo/) is released!
|
||||
- ```2025/02/18```: Release the inference code for [Sliding Tile Attention](https://hao-ai-lab.github.io/blogs/sta/).
|
||||
|
||||
- `2025/06/14`: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389)
|
||||
- `2025/04/24`: [FastVideo V1](https://hao-ai-lab.github.io/blogs/fastvideo/) is released!
|
||||
- `2025/02/18`: Release the inference code for [Sliding Tile Attention](https://hao-ai-lab.github.io/blogs/sta/).
|
||||
</details>
|
||||
|
||||
## Key Features
|
||||
|
||||
FastVideo has the following features:
|
||||
|
||||
- End-to-end post-training support for bidirectional and autoregressive models:
|
||||
- Support full finetuning and LoRA finetuning for state-of-the-art open video DiTs
|
||||
- Data preprocessing pipeline for video, image, and text data
|
||||
- Distribution Matching Distillation (DMD2) stepwise distillation.
|
||||
- Sparse attention with [Video Sparse Attention](https://arxiv.org/pdf/2505.13389)
|
||||
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
|
||||
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achineve >50x denoising speedup
|
||||
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
|
||||
- Causal distillation through Self-Forcing
|
||||
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
|
||||
@@ -40,39 +41,36 @@ FastVideo has the following features:
|
||||
- Diverse hardware and OS support
|
||||
- Support H100, A100, 4090
|
||||
- Support Linux, Windows, MacOS
|
||||
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
|
||||
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/hardware_support/) for full list of supported hardware and OS.
|
||||
|
||||
## Getting Started
|
||||
|
||||
We recommend using [uv](https://docs.astral.sh/uv/) to create a clean environment. If you previously used Conda, switching to uv generally gives faster and more stable installs.
|
||||
We recommend using an environment manager such as `Conda` to create a clean environment:
|
||||
|
||||
```bash
|
||||
# Create and activate a new uv environment
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
# Create and activate a new conda environment
|
||||
conda create -n fastvideo python=3.12
|
||||
conda activate fastvideo
|
||||
|
||||
# Install FastVideo
|
||||
uv pip install fastvideo
|
||||
pip install fastvideo
|
||||
```
|
||||
|
||||
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
|
||||
|
||||
## Sparse Distillation
|
||||
|
||||
For our sparse distillation techniques, please see our [distillation docs](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/) and check out our [blog](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
|
||||
|
||||
See below for recipes and datasets:
|
||||
|
||||
| Model | Sparse Distillation | Dataset |
|
||||
| ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
||||
| [FastWan2.1-T2V-1.3B](https://huggingface.co/FastVideo/FastWan2.1-T2V-1.3B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.1-T2V/Wan-Syn-Data-480P) | [FastVideo Synthetic Wan2.1 480P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x448x832_600k) |
|
||||
| [FastWan2.2-TI2V-5B](https://huggingface.co/FastVideo/FastWan2.2-TI2V-5B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.2-TI2V-5B-Diffusers/Data-free) | [FastVideo Synthetic Wan2.2 720P](https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k) |
|
||||
| Model | Sparse Distillation | Dataset |
|
||||
|:-------------------------------------------------------------------------------------------: |:---------------------------------------------------------------------------------------------------------------: |:--------------------------------------------------------------------------------------------------------: |
|
||||
| [FastWan2.1-T2V-1.3B](https://huggingface.co/FastVideo/FastWan2.1-T2V-1.3B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.1-T2V/Wan-Syn-Data-480P) | [FastVideo Synthetic Wan2.1 480P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x448x832_600k) |
|
||||
| [FastWan2.1-T2V-14B-Preview](https://huggingface.co/FastVideo/FastWan2.1-T2V-14B-Diffusers) | Coming soon! | [FastVideo Synthetic Wan2.1 720P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x768x1280_250k) |
|
||||
| [FastWan2.2-TI2V-5B](https://huggingface.co/FastVideo/FastWan2.2-TI2V-5B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.2-TI2V-5B-Diffusers/Data-free) | [FastVideo Synthetic Wan2.2 720P](https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k) |
|
||||
|
||||
## Inference
|
||||
|
||||
### Generating Your First Video
|
||||
|
||||
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/attention/vsa/#installation). Create a file called `example.py` with the following code:
|
||||
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/video_sparse_attention/installation/). Create a file called `example.py` with the following code:
|
||||
|
||||
```python
|
||||
import os
|
||||
@@ -93,6 +91,7 @@ def main():
|
||||
# Generate the video
|
||||
video = generator.generate_video(
|
||||
prompt,
|
||||
return_frames=True, # Also return frames from this call (defaults to False)
|
||||
output_path="my_videos/", # Controls where videos are saved
|
||||
save_video=True
|
||||
)
|
||||
@@ -109,37 +108,55 @@ python example.py
|
||||
|
||||
For a more detailed guide, please see our [inference quick start](https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/).
|
||||
|
||||
## More Guides
|
||||
### Other docs:
|
||||
|
||||
- [Design Overview](https://hao-ai-lab.github.io/FastVideo/design/overview/)
|
||||
- [Contribution Guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/)
|
||||
|
||||
## Distillation and Finetuning
|
||||
- [Distillation Guide](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/)
|
||||
- [Contribution Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
|
||||
<!-- - [Finetuning Guide](https://hao-ai-lab.github.io/FastVideo/training/finetune.html) -->
|
||||
|
||||
## Awesome work using FastVideo or our research projects
|
||||
|
||||
- [SGLang](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen): SGLang's diffusion inference functionality is based on a fork of FastVideo on Sept. 24, 2025.
|
||||
- [DanceGRPO](https://github.com/XueZeyue/DanceGRPO): A unified framework to adapt Group Relative Policy Optimization (GRPO) to visual generation paradigms. Code based on FastVideo.
|
||||
- [SRPO](https://github.com/Tencent-Hunyuan/SRPO): A method to directly align the full diffusion trajectory with fine-grained human preference. Code based on FastVideo.
|
||||
- [DCM](https://github.com/Vchitect/DCM): Dual-expert consistency model for efficient and high-quality video generation. Code based on FastVideo.
|
||||
- [HY-WorldPlay](https://github.com/Tencent-Hunyuan/HY-WorldPlay): An action-conditioned world model model trained using FastVideo framework.
|
||||
- [Hunyuan Video 1.5](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5): A leading lightweight video generation model, where they proposed SSTA based on Sliding Tile Attention.
|
||||
- [Kandinsky-5.0](https://github.com/kandinskylab/kandinsky-5): A family of diffusion models for video & image generation, where their NABLA attention includes a Sliding Tile Attention branch.
|
||||
- [LongCat Video](https://github.com/meituan-longcat/LongCat-Video): A foundational video generation model with 13.6B parameters with block-sparse attention similar to Video Sparse Attention.
|
||||
- [SGLang](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen): SGLang's diffusion inference functionality is based on a fork of FastVideo on Sept. 24, 2025. [](https://github.com/sgl-project/sglang)
|
||||
|
||||
- [DanceGRPO](https://github.com/XueZeyue/DanceGRPO): A unified framework to adapt Group Relative Policy Optimization (GRPO) to visual generation paradigms. Code based on FastVideo. [](https://github.com/XueZeyue/DanceGRPO)
|
||||
- [SRPO](https://github.com/Tencent-Hunyuan/SRPO): A method to directly align the full diffusion trajectory with fine-grained human preference. Code based on FastVideo. [](https://github.com/Tencent-Hunyuan/SRPO)
|
||||
- [DCM](https://github.com/Vchitect/DCM): Dual-expert consistency model for efficient and high-quality video generation. Code based on FastVideo. [](https://github.com/Vchitect/DCM)
|
||||
- [Hunyuan Video 1.5](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5): A leading lightweight video generation model, where they proposed SSTA based on Sliding Tile Attention. [](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5)
|
||||
- [Kandinsky-5.0](https://github.com/kandinskylab/kandinsky-5): A family of diffusion models for video & image generation, where their NABLA attention includes a Sliding Tile Attention branch. [](https://github.com/kandinskylab/kandinsky-5)
|
||||
- [LongCat Video](https://github.com/meituan-longcat/LongCat-Video): A foundational video generation model with 13.6B parameters with block-sparse attention similar to Video Sparse Attention. [](https://github.com/meituan-longcat/LongCat-Video)
|
||||
|
||||
## 🤝 Contributing
|
||||
|
||||
We welcome all contributions. Please check out our guide [here](https://hao-ai-lab.github.io/FastVideo/contributing/overview/).
|
||||
See details in [development roadmap](https://github.com/hao-ai-lab/FastVideo/issues/899).
|
||||
|
||||
## Acknowledgement
|
||||
We learned and reused code from the following projects:
|
||||
- [Wan-Video](https://github.com/Wan-Video)
|
||||
- [ThunderKittens](https://github.com/HazyResearch/ThunderKittens)
|
||||
- [Triton](https://github.com/triton-lang/triton)
|
||||
- [DMD2](https://github.com/tianweiy/DMD2)
|
||||
- [diffusers](https://github.com/huggingface/diffusers)
|
||||
- [xDiT](https://github.com/xdit-project/xDiT)
|
||||
- [vLLM](https://github.com/vllm-project/vllm)
|
||||
- [SGLang](https://github.com/sgl-project/sglang)
|
||||
|
||||
We learned the design and reused code from the following projects: [Wan-Video](https://github.com/Wan-Video), [ThunderKittens](https://github.com/HazyResearch/ThunderKittens), [DMD2](https://github.com/tianweiy/DMD2), [diffusers](https://github.com/huggingface/diffusers), [xDiT](https://github.com/xdit-project/xDiT), [vLLM](https://github.com/vllm-project/vllm), [SGLang](https://github.com/sgl-project/sglang). We thank [MBZUAI](https://ifm.mbzuai.ac.ae/), [Anyscale](https://www.anyscale.com/), and [GMI Cloud](https://www.gmicloud.ai/) for their support throughout this project.
|
||||
We thank [MBZUAI](https://ifm.mbzuai.ac.ae/), [Anyscale](https://www.anyscale.com/), and [GMI Cloud](https://www.gmicloud.ai/) for their support throughout this project.
|
||||
|
||||
## Citation
|
||||
|
||||
If you find FastVideo useful, please consider citing our research work:
|
||||
If you find FastVideo useful, please considering citing our work:
|
||||
|
||||
```bibtex
|
||||
@software{fastvideo2024,
|
||||
title = {FastVideo: A Unified Framework for Accelerated Video Generation},
|
||||
author = {The FastVideo Team},
|
||||
url = {https://github.com/hao-ai-lab/FastVideo},
|
||||
month = apr,
|
||||
year = {2024},
|
||||
}
|
||||
|
||||
@article{zhang2025vsa,
|
||||
title={Vsa: Faster video diffusion with trainable sparse attention},
|
||||
author={Zhang, Peiyuan and Chen, Yongqi and Huang, Haofeng and Lin, Will and Liu, Zhengzhong and Stoica, Ion and Xing, Eric and Zhang, Hao},
|
||||
|
||||
|
Before Width: | Height: | Size: 1.2 MiB |
@@ -1,7 +0,0 @@
|
||||
# FastVideo/assets/videos
|
||||
|
||||
This folder is used to store **video assets for examples**, primarily **input videos** consumed by scripts under `FastVideo/examples/`.
|
||||
|
||||
- **Typical contents**: short input clips for demos (e.g., video2world / image2video examples).
|
||||
- **Non-critical**: these assets are for convenience and are not required to use the FastVideo library.
|
||||
- **Large files**: avoid committing large videos to git; prefer shared storage or download-on-demand.
|
||||
@@ -9,12 +9,11 @@ import datetime
|
||||
import locale
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
# This script outputs relevant system environment info.
|
||||
# Run it with: python collect_env.py
|
||||
# Requires Python 3.10+ (matches fastvideo); uses shutil.which (Python 3.3+).
|
||||
# Unlike the rest of the PyTorch this file must be python2 compliant.
|
||||
# This script outputs relevant system environment info
|
||||
# Run it with `python collect_env.py` or `python -m torch.utils.collect_env`
|
||||
from collections import namedtuple
|
||||
|
||||
from fastvideo.envs import environment_variables
|
||||
@@ -496,7 +495,8 @@ def get_pip_packages(run_lambda, patterns=None):
|
||||
|
||||
if pip_available:
|
||||
cmd = [sys.executable, '-mpip', 'list', '--format=freeze']
|
||||
elif shutil.which("uv") is not None:
|
||||
elif os.environ.get("UV") is not None:
|
||||
print("uv is set")
|
||||
cmd = ["uv", "pip", "list", "--format=freeze"]
|
||||
else:
|
||||
raise RuntimeError(
|
||||
|
||||
@@ -443,6 +443,7 @@
|
||||
1025,
|
||||
"fixed",
|
||||
24,
|
||||
-99999,
|
||||
-99999
|
||||
],
|
||||
"auto_widget_states": {
|
||||
@@ -490,6 +491,11 @@
|
||||
"isAuto": true,
|
||||
"value": -99999,
|
||||
"cachedValue": "X://insert/path/here.mp4"
|
||||
},
|
||||
"enable_teacache": {
|
||||
"isAuto": true,
|
||||
"value": -99999,
|
||||
"cachedValue": true
|
||||
}
|
||||
}
|
||||
},
|
||||
@@ -636,4 +642,4 @@
|
||||
"VHS_KeepIntermediate": true
|
||||
},
|
||||
"version": 0.4
|
||||
}
|
||||
}
|
||||
@@ -353,7 +353,8 @@
|
||||
1024,
|
||||
"fixed",
|
||||
24,
|
||||
"X://insert/path/here.mp4"
|
||||
"X://insert/path/here.mp4",
|
||||
true
|
||||
],
|
||||
"auto_widget_states": {
|
||||
"height": {
|
||||
@@ -400,6 +401,11 @@
|
||||
"isAuto": true,
|
||||
"value": "X://insert/path/here.mp4",
|
||||
"cachedValue": "X://insert/path/here.mp4"
|
||||
},
|
||||
"enable_teacache": {
|
||||
"isAuto": true,
|
||||
"value": true,
|
||||
"cachedValue": true
|
||||
}
|
||||
}
|
||||
},
|
||||
@@ -688,4 +694,4 @@
|
||||
"VHS_KeepIntermediate": true
|
||||
},
|
||||
"version": 0.4
|
||||
}
|
||||
}
|
||||
@@ -31,6 +31,9 @@ class InferenceArgs:
|
||||
"image_path": ("STRING", {
|
||||
"default": "X://insert/path/here.mp4"
|
||||
}),
|
||||
"enable_teacache": ([True, False], {
|
||||
"default": False
|
||||
}),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -54,6 +57,7 @@ class InferenceArgs:
|
||||
seed,
|
||||
fps,
|
||||
image_path,
|
||||
enable_teacache,
|
||||
):
|
||||
raw_args = {
|
||||
"height": height,
|
||||
@@ -65,6 +69,7 @@ class InferenceArgs:
|
||||
"seed": seed,
|
||||
"fps": fps,
|
||||
"image_path": image_path,
|
||||
"enable_teacache": enable_teacache,
|
||||
}
|
||||
|
||||
# Filter out keys where value is -99999, handling different types properly
|
||||
|
||||
@@ -552,7 +552,7 @@ app.registerExtension({
|
||||
]
|
||||
const floatWidgetNames = ["embedded_cfg_scale", "guidance_scale"]
|
||||
const comboWidgetNames = ["vae_tiling", "vae_precision", "vae_sp", "text_encoder_precision", "precision",
|
||||
"load_encoder", "load_decoder", "use_tiling", "use_temporal_tiling", "use_parallel_tiling", "dit_cpu_offload"
|
||||
"load_encoder", "load_decoder", "use_tiling", "use_temporal_tiling", "use_parallel_tiling", "dit_cpu_offload", "enable_teacache"
|
||||
]
|
||||
const stringWidgetNames = ["prefix", "quant_config", "lora_config", "image_path"]
|
||||
|
||||
|
||||
@@ -0,0 +1,195 @@
|
||||
import argparse
|
||||
import os
|
||||
import tempfile
|
||||
|
||||
import gradio as gr
|
||||
import torch
|
||||
from diffusers import FlowMatchEulerDiscreteScheduler
|
||||
from diffusers.utils import export_to_video
|
||||
|
||||
from fastvideo.distill.solver import PCMFMScheduler
|
||||
from fastvideo.models.mochi_hf.modeling_mochi import MochiTransformer3DModel
|
||||
from fastvideo.models.mochi_hf.pipeline_mochi import MochiPipeline
|
||||
|
||||
|
||||
def init_args():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--prompts", nargs="+", default=[])
|
||||
parser.add_argument("--num_frames", type=int, default=25)
|
||||
parser.add_argument("--height", type=int, default=480)
|
||||
parser.add_argument("--width", type=int, default=848)
|
||||
parser.add_argument("--num_inference_steps", type=int, default=8)
|
||||
parser.add_argument("--guidance_scale", type=float, default=4.5)
|
||||
parser.add_argument("--model_path", type=str, default="data/mochi")
|
||||
parser.add_argument("--seed", type=int, default=12345)
|
||||
parser.add_argument("--transformer_path", type=str, default=None)
|
||||
parser.add_argument("--scheduler_type", type=str, default="pcm_linear_quadratic")
|
||||
parser.add_argument("--lora_checkpoint_dir", type=str, default=None)
|
||||
parser.add_argument("--shift", type=float, default=8.0)
|
||||
parser.add_argument("--num_euler_timesteps", type=int, default=50)
|
||||
parser.add_argument("--linear_threshold", type=float, default=0.1)
|
||||
parser.add_argument("--linear_range", type=float, default=0.75)
|
||||
parser.add_argument("--cpu_offload", action="store_true")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def load_model(args):
|
||||
if args.scheduler_type == "euler":
|
||||
scheduler = FlowMatchEulerDiscreteScheduler()
|
||||
else:
|
||||
linear_quadratic = True if "linear_quadratic" in args.scheduler_type else False
|
||||
scheduler = PCMFMScheduler(
|
||||
1000,
|
||||
args.shift,
|
||||
args.num_euler_timesteps,
|
||||
linear_quadratic,
|
||||
args.linear_threshold,
|
||||
args.linear_range,
|
||||
)
|
||||
|
||||
if args.transformer_path:
|
||||
transformer = MochiTransformer3DModel.from_pretrained(args.transformer_path)
|
||||
else:
|
||||
transformer = MochiTransformer3DModel.from_pretrained(args.model_path, subfolder="transformer/")
|
||||
|
||||
pipe = MochiPipeline.from_pretrained(args.model_path, transformer=transformer, scheduler=scheduler)
|
||||
pipe.enable_vae_tiling()
|
||||
# pipe.to(device)
|
||||
# if args.cpu_offload:
|
||||
pipe.enable_sequential_cpu_offload()
|
||||
return pipe
|
||||
|
||||
|
||||
def generate_video(
|
||||
prompt,
|
||||
negative_prompt,
|
||||
use_negative_prompt,
|
||||
seed,
|
||||
guidance_scale,
|
||||
num_frames,
|
||||
height,
|
||||
width,
|
||||
num_inference_steps,
|
||||
randomize_seed=False,
|
||||
):
|
||||
if randomize_seed:
|
||||
seed = torch.randint(0, 1000000, (1, )).item()
|
||||
|
||||
generator = torch.Generator(device="cuda").manual_seed(seed)
|
||||
|
||||
if not use_negative_prompt:
|
||||
negative_prompt = None
|
||||
|
||||
with torch.autocast("cuda", dtype=torch.bfloat16):
|
||||
output = pipe(
|
||||
prompt=[prompt],
|
||||
negative_prompt=negative_prompt,
|
||||
height=height,
|
||||
width=width,
|
||||
num_frames=num_frames,
|
||||
num_inference_steps=num_inference_steps,
|
||||
guidance_scale=guidance_scale,
|
||||
generator=generator,
|
||||
).frames[0]
|
||||
|
||||
output_path = os.path.join(tempfile.mkdtemp(), "output.mp4")
|
||||
export_to_video(output, output_path, fps=30)
|
||||
return output_path, seed
|
||||
|
||||
|
||||
examples = [
|
||||
"A hand enters the frame, pulling a sheet of plastic wrap over three balls of dough placed on a wooden surface. The plastic wrap is stretched to cover the dough more securely. The hand adjusts the wrap, ensuring that it is tight and smooth over the dough. The scene focuses on the hand’s movements as it secures the edges of the plastic wrap. No new objects appear, and the camera remains stationary, focusing on the action of covering the dough.",
|
||||
"A vintage train snakes through the mountains, its plume of white steam rising dramatically against the jagged peaks. The cars glint in the late afternoon sun, their deep crimson and gold accents lending a touch of elegance. The tracks carve a precarious path along the cliffside, revealing glimpses of a roaring river far below. Inside, passengers peer out the large windows, their faces lit with awe as the landscape unfolds.",
|
||||
"A crowded rooftop bar buzzes with energy, the city skyline twinkling like a field of stars in the background. Strings of fairy lights hang above, casting a warm, golden glow over the scene. Groups of people gather around high tables, their laughter blending with the soft rhythm of live jazz. The aroma of freshly mixed cocktails and charred appetizers wafts through the air, mingling with the cool night breeze.",
|
||||
]
|
||||
|
||||
args = init_args()
|
||||
pipe = load_model(args)
|
||||
print("load model successfully")
|
||||
with gr.Blocks() as demo:
|
||||
gr.Markdown("# Fastvideo Mochi Video Generation Demo")
|
||||
|
||||
with gr.Group():
|
||||
with gr.Row():
|
||||
prompt = gr.Text(
|
||||
label="Prompt",
|
||||
show_label=False,
|
||||
max_lines=1,
|
||||
placeholder="Enter your prompt",
|
||||
container=False,
|
||||
)
|
||||
run_button = gr.Button("Run", scale=0)
|
||||
result = gr.Video(label="Result", show_label=False)
|
||||
|
||||
with gr.Accordion("Advanced options", open=False):
|
||||
with gr.Group():
|
||||
with gr.Row():
|
||||
height = gr.Slider(
|
||||
label="Height",
|
||||
minimum=256,
|
||||
maximum=1024,
|
||||
step=32,
|
||||
value=args.height,
|
||||
)
|
||||
width = gr.Slider(label="Width", minimum=256, maximum=1024, step=32, value=args.width)
|
||||
|
||||
with gr.Row():
|
||||
num_frames = gr.Slider(
|
||||
label="Number of Frames",
|
||||
minimum=21,
|
||||
maximum=163,
|
||||
value=args.num_frames,
|
||||
)
|
||||
guidance_scale = gr.Slider(
|
||||
label="Guidance Scale",
|
||||
minimum=1,
|
||||
maximum=12,
|
||||
value=args.guidance_scale,
|
||||
)
|
||||
num_inference_steps = gr.Slider(
|
||||
label="Inference Steps",
|
||||
minimum=4,
|
||||
maximum=100,
|
||||
value=args.num_inference_steps,
|
||||
)
|
||||
|
||||
with gr.Row():
|
||||
use_negative_prompt = gr.Checkbox(label="Use negative prompt", value=False)
|
||||
negative_prompt = gr.Text(
|
||||
label="Negative prompt",
|
||||
max_lines=1,
|
||||
placeholder="Enter a negative prompt",
|
||||
visible=False,
|
||||
)
|
||||
|
||||
seed = gr.Slider(label="Seed", minimum=0, maximum=1000000, step=1, value=args.seed)
|
||||
randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
|
||||
seed_output = gr.Number(label="Used Seed")
|
||||
|
||||
gr.Examples(examples=examples, inputs=prompt)
|
||||
|
||||
use_negative_prompt.change(
|
||||
fn=lambda x: gr.update(visible=x),
|
||||
inputs=use_negative_prompt,
|
||||
outputs=negative_prompt,
|
||||
)
|
||||
|
||||
run_button.click(
|
||||
fn=generate_video,
|
||||
inputs=[
|
||||
prompt,
|
||||
negative_prompt,
|
||||
use_negative_prompt,
|
||||
seed,
|
||||
guidance_scale,
|
||||
num_frames,
|
||||
height,
|
||||
width,
|
||||
num_inference_steps,
|
||||
randomize_seed,
|
||||
],
|
||||
outputs=[result, seed_output],
|
||||
)
|
||||
|
||||
if __name__ == "__main__":
|
||||
demo.queue(max_size=20).launch(server_name="0.0.0.0", server_port=7860)
|
||||
@@ -0,0 +1,15 @@
|
||||
Fast-Hunyuan comparison with original Hunyuan, achieving an 8X diffusion speed boost with the FastVideo framework.
|
||||
|
||||
https://github.com/user-attachments/assets/064ac1d2-11ed-4a0c-955b-4d412a96ef30
|
||||
|
||||
Fast-Mochi comparison with original Mochi, achieving an 8X diffusion speed boost with the FastVideo framework.
|
||||
|
||||
https://github.com/user-attachments/assets/5fbc4596-56d6-43aa-98e0-da472cf8e26c
|
||||
|
||||
Comparison between OpenAI Sora, original Hunyuan and FastHunyuan
|
||||
|
||||
https://github.com/user-attachments/assets/d323b712-3f68-42b2-952b-94f6a49c4836
|
||||
|
||||
Comparison between original FastHunyuan, LLM-INT8 quantized FastHunyuan and NF4 quantized FastHunyuan
|
||||
|
||||
https://github.com/user-attachments/assets/cf89efb5-5f68-4949-a085-f41c1ef26c94
|
||||
@@ -43,7 +43,7 @@ RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
uv pip install --no-cache-dir --upgrade pip && \
|
||||
uv pip install --no-cache-dir .[dev] && \
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.7.16/flash_attn-2.8.3+cu128torch2.10-cp310-cp310-linux_x86_64.whl
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.5.4/flash_attn-2.8.3%2Bcu128torch2.9-cp310-cp310-linux_x86_64.whl
|
||||
|
||||
COPY . .
|
||||
|
||||
|
||||
@@ -43,7 +43,7 @@ RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
uv pip install --no-cache-dir --upgrade pip && \
|
||||
uv pip install --no-cache-dir .[dev] && \
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.7.16/flash_attn-2.8.3+cu128torch2.10-cp311-cp311-linux_x86_64.whl
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.5.4/flash_attn-2.8.3%2Bcu128torch2.9-cp311-cp311-linux_x86_64.whl
|
||||
|
||||
COPY . .
|
||||
|
||||
|
||||
@@ -43,7 +43,7 @@ RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
uv pip install --no-cache-dir --upgrade pip && \
|
||||
uv pip install --no-cache-dir .[dev] && \
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.7.16/flash_attn-2.8.3+cu128torch2.10-cp312-cp312-linux_x86_64.whl
|
||||
uv pip install --no-cache-dir https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.5.4/flash_attn-2.8.3%2Bcu128torch2.9-cp312-cp312-linux_x86_64.whl
|
||||
|
||||
COPY . .
|
||||
|
||||
|
||||
@@ -121,15 +121,6 @@ This page contains the complete API reference for the FastVideo library.
|
||||
show_root_toc_entry: true
|
||||
heading_level: 4
|
||||
|
||||
## fastvideo.registry
|
||||
|
||||
::: fastvideo.registry
|
||||
options:
|
||||
show_source: true
|
||||
show_root_heading: true
|
||||
show_root_toc_entry: true
|
||||
heading_level: 3
|
||||
|
||||
## fastvideo.pipelines
|
||||
|
||||
::: fastvideo.pipelines
|
||||
|
||||
|
Before Width: | Height: | Size: 211 KiB |
|
Before Width: | Height: | Size: 117 KiB |
|
Before Width: | Height: | Size: 461 KiB After Width: | Height: | Size: 18 KiB |
|
Before Width: | Height: | Size: 210 KiB |
|
Before Width: | Height: | Size: 437 KiB After Width: | Height: | Size: 27 KiB |
@@ -54,8 +54,7 @@ To use this:
|
||||
1. **Set Context**: In your pipeline or generation loop, use the `set_forward_context` context manager.
|
||||
2. **Access Context**: Inside your attention backend, use `get_forward_context()`.
|
||||
|
||||
See [`docs/attention/sta/index.md`](../sta/index.md) for a legacy STA example
|
||||
of passing complex configuration (window sizes) through `ForwardContext`.
|
||||
See [`docs/attention/sta/index.md`](../sta/index.md) (Sliding Tile Attention) for an example of how complex configuration (window sizes) is passed this way.
|
||||
|
||||
## 3. Adding Compiled Kernels (C++/CUDA)
|
||||
|
||||
|
||||
@@ -5,16 +5,11 @@ FastVideo provides highly optimized custom attention kernels to accelerate video
|
||||
## Supported Kernels
|
||||
|
||||
* **[Video Sparse Attention (VSA)](vsa/index.md)**: Sparse attention mechanism selecting top-k blocks.
|
||||
* **[Sliding Tile Attention (STA)](sta/index.md)**: STA kernel support is kept in
|
||||
`fastvideo-kernel`; full FastVideo STA pipeline workflow is archived in
|
||||
`sta_do_not_delete`.
|
||||
* **Backend development guide**: See the developer guide at
|
||||
[Attention Backend Development](../contributing/attention_backend.md).
|
||||
* **[Sliding Tile Attention (STA)](sta/index.md)**: Optimized attention for window-based video generation.
|
||||
|
||||
## General Build Instructions
|
||||
|
||||
These instructions apply to building the `fastvideo-kernel` package from
|
||||
source, which includes both STA and VSA kernels.
|
||||
These instructions apply to building the `fastvideo-kernel` package from source, which includes both STA and VSA kernels.
|
||||
|
||||
### Prerequisites
|
||||
|
||||
|
||||
@@ -1,84 +1,27 @@
|
||||
# Sliding Tile Attention (STA)
|
||||
|
||||
STA inference integration is archived from `main`.
|
||||
Optimized attention for window-based video generation (e.g., HunyuanVideo).
|
||||
|
||||
The full STA pipeline code (including mask search and STA inference wiring in
|
||||
`fastvideo/`) is preserved in:
|
||||
## Installation
|
||||
|
||||
- https://github.com/hao-ai-lab/FastVideo/tree/sta_do_not_delete
|
||||
STA is included in the `fastvideo-kernel` package. See the [main Attention page](../index.md) for build instructions.
|
||||
|
||||
In this branch, STA kernels in `fastvideo-kernel` are still kept.
|
||||
## Usage
|
||||
|
||||
## Why STA is not in `main`
|
||||
```python
|
||||
from fastvideo_kernel import sliding_tile_attention
|
||||
|
||||
We do not keep STA pipeline integration in `main` because we believe Video
|
||||
Sparse Attention (VSA) is strictly better than STA for the actively maintained
|
||||
FastVideo inference path.
|
||||
# q, k, v: [batch_size, num_heads, seq_length, head_dim]
|
||||
# window_size: List of (t, h, w) tiles. Tile size is (6, 8, 8).
|
||||
# text_length: Number of text tokens (0-256)
|
||||
|
||||
## What to checkout for STA workflows
|
||||
|
||||
To run the full STA workflow, switch to the archived branch:
|
||||
|
||||
```bash
|
||||
git fetch origin
|
||||
git checkout sta_do_not_delete
|
||||
out = sliding_tile_attention(
|
||||
q, k, v,
|
||||
window_size=[(3, 3, 3)], # Example window
|
||||
text_length=256
|
||||
)
|
||||
```
|
||||
|
||||
## Mask Search (archive branch)
|
||||
|
||||
The reference script is:
|
||||
|
||||
- `examples/inference/sta_mask_search/inference_wan_sta.sh`
|
||||
|
||||
It runs two stages:
|
||||
|
||||
1. `STA_searching` (full search), output at
|
||||
`inference_results/sta/mask_search_full`
|
||||
2. `STA_tuning` (sparse tuning), output at
|
||||
`inference_results/sta/mask_search_sparse`
|
||||
|
||||
Run:
|
||||
|
||||
```bash
|
||||
export FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN
|
||||
export FASTVIDEO_ATTENTION_CONFIG=assets/mask_strategy_wan.json
|
||||
|
||||
bash examples/inference/sta_mask_search/inference_wan_sta.sh
|
||||
```
|
||||
|
||||
## STA Inference (archive branch)
|
||||
|
||||
With a selected mask strategy, run inference with:
|
||||
|
||||
```bash
|
||||
export FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN
|
||||
export FASTVIDEO_ATTENTION_CONFIG=assets/mask_strategy_wan.json
|
||||
|
||||
fastvideo generate \
|
||||
--model-path Wan-AI/Wan2.1-T2V-14B-Diffusers \
|
||||
--num-gpus 2 \
|
||||
--tp-size 2 \
|
||||
--sp-size 2 \
|
||||
--height 768 \
|
||||
--width 1280 \
|
||||
--num-frames 69 \
|
||||
--num-inference-steps 50 \
|
||||
--prompt "A cinematic wildlife shot of a lion walking in golden grasslands." \
|
||||
--output-path outputs_video/STA/
|
||||
```
|
||||
|
||||
Python usage on the archive branch can also set `STA_mode` in
|
||||
`VideoGenerator.from_pretrained(...)`:
|
||||
|
||||
- `STA_searching`
|
||||
- `STA_tuning`
|
||||
- `STA_inference`
|
||||
|
||||
## Kernel-level API (current branch)
|
||||
|
||||
STA kernels remain available from `fastvideo-kernel`. See
|
||||
[Attention overview](../index.md) for build instructions.
|
||||
|
||||
## Citation
|
||||
|
||||
If you use Sliding Tile Attention in your research, please cite:
|
||||
|
||||
@@ -13,13 +13,11 @@ from fastvideo_kernel import video_sparse_attn
|
||||
|
||||
# q, k, v: [batch_size, num_heads, seq_len, head_dim]
|
||||
# variable_block_sizes: Number of valid tokens per block
|
||||
# q_variable_block_sizes: Number of valid tokens per q block (can differ from KV for q/k of different lengths)
|
||||
# topk: Number of blocks to attend
|
||||
|
||||
output = video_sparse_attn(
|
||||
q, k, v,
|
||||
block_sizes,
|
||||
block_sizes,
|
||||
variable_block_sizes=block_sizes,
|
||||
topk=32
|
||||
)
|
||||
```
|
||||
|
||||
@@ -1,176 +0,0 @@
|
||||
# Attention Backend Development
|
||||
|
||||
This guide is for contributors adding a new attention backend (kernel or
|
||||
implementation) to FastVideo. If you just want to use existing kernels or build
|
||||
fastvideo-kernel, see [Attention overview](../attention/index.md).
|
||||
|
||||
## When you need this guide
|
||||
|
||||
Use this guide when you are:
|
||||
|
||||
- Adding a new attention kernel or algorithm.
|
||||
- Wiring an existing kernel into FastVideo's attention selection.
|
||||
- Extending attention support to a new platform.
|
||||
|
||||
## 0) Choose a backend name and scope
|
||||
|
||||
Pick a backend name in `UPPER_SNAKE_CASE` and decide where it should run.
|
||||
Example: `MY_NEW_ATTN`.
|
||||
|
||||
You will use this name in:
|
||||
|
||||
- `AttentionBackendEnum` (global list of backends).
|
||||
- `get_name()` in your backend class (match the enum name).
|
||||
- Platform selectors (CUDA/ROCm/MPS/NPU) to return your backend class.
|
||||
|
||||
## 1) Add enum + platform selection
|
||||
|
||||
1) Add your backend to `fastvideo/platforms/interface.py`:
|
||||
|
||||
```python
|
||||
class AttentionBackendEnum(enum.Enum):
|
||||
...
|
||||
MY_NEW_ATTN = enum.auto()
|
||||
```
|
||||
|
||||
1) Register it in platform selection (example: CUDA). Update
|
||||
`fastvideo/platforms/cuda.py` inside `get_attn_backend_cls`:
|
||||
|
||||
```python
|
||||
elif selected_backend == AttentionBackendEnum.MY_NEW_ATTN:
|
||||
try:
|
||||
from fastvideo.attention.backends.my_new_attn import MyNewAttnBackend
|
||||
return "fastvideo.attention.backends.my_new_attn.MyNewAttnBackend"
|
||||
except ImportError as e:
|
||||
logger.error("Failed to import MY_NEW_ATTN backend: %s", str(e))
|
||||
raise
|
||||
```
|
||||
|
||||
If you want support on other platforms, add a similar branch in
|
||||
`fastvideo/platforms/rocm.py`, `fastvideo/platforms/mps.py`, or `fastvideo/platforms/npu.py`.
|
||||
|
||||
## 2) Implement the backend
|
||||
|
||||
Create `fastvideo/attention/backends/my_new_attn.py` and implement the required
|
||||
classes.
|
||||
|
||||
Minimal skeleton (no custom metadata):
|
||||
|
||||
```python
|
||||
import torch
|
||||
from dataclasses import dataclass
|
||||
from fastvideo.attention.backends.abstract import (
|
||||
AttentionBackend,
|
||||
AttentionImpl,
|
||||
AttentionMetadata,
|
||||
AttentionMetadataBuilder,
|
||||
)
|
||||
|
||||
class MyNewAttnBackend(AttentionBackend):
|
||||
@staticmethod
|
||||
def get_name() -> str:
|
||||
return "MY_NEW_ATTN"
|
||||
|
||||
@staticmethod
|
||||
def get_impl_cls() -> type["MyNewAttnImpl"]:
|
||||
return MyNewAttnImpl
|
||||
|
||||
@staticmethod
|
||||
def get_metadata_cls() -> type["AttentionMetadata"]:
|
||||
return AttentionMetadata
|
||||
|
||||
@staticmethod
|
||||
def get_builder_cls() -> type["AttentionMetadataBuilder"]:
|
||||
return MyNewAttnMetadataBuilder
|
||||
|
||||
|
||||
@dataclass
|
||||
class MyNewAttnMetadata(AttentionMetadata):
|
||||
current_timestep: int
|
||||
|
||||
|
||||
class MyNewAttnMetadataBuilder(AttentionMetadataBuilder):
|
||||
def __init__(self) -> None:
|
||||
pass
|
||||
|
||||
def prepare(self) -> None:
|
||||
pass
|
||||
|
||||
def build(self, current_timestep: int, **kwargs):
|
||||
return MyNewAttnMetadata(current_timestep=current_timestep)
|
||||
|
||||
|
||||
class MyNewAttnImpl(AttentionImpl):
|
||||
def __init__(
|
||||
self,
|
||||
num_heads: int,
|
||||
head_size: int,
|
||||
softmax_scale: float,
|
||||
causal: bool = False,
|
||||
num_kv_heads: int | None = None,
|
||||
prefix: str = "",
|
||||
**extra_impl_args,
|
||||
) -> None:
|
||||
self.softmax_scale = softmax_scale
|
||||
self.causal = causal
|
||||
|
||||
def forward(
|
||||
self,
|
||||
query: torch.Tensor,
|
||||
key: torch.Tensor,
|
||||
value: torch.Tensor,
|
||||
attn_metadata: MyNewAttnMetadata,
|
||||
) -> torch.Tensor:
|
||||
# Implement attention
|
||||
return torch.nn.functional.scaled_dot_product_attention(
|
||||
query.transpose(1, 2),
|
||||
key.transpose(1, 2),
|
||||
value.transpose(1, 2),
|
||||
is_causal=self.causal,
|
||||
scale=self.softmax_scale,
|
||||
).transpose(1, 2)
|
||||
```
|
||||
|
||||
Optional:
|
||||
|
||||
- Implement `preprocess_qkv` / `postprocess_output` if your kernel needs tiling
|
||||
or reshaping.
|
||||
- Use `fastvideo.forward_context.get_forward_context()` if you need dynamic
|
||||
per-step data (e.g., window sizes).
|
||||
- Set `accept_output_buffer = True` if your backend writes into a provided
|
||||
output buffer.
|
||||
|
||||
## 3) Wire into attention layers
|
||||
|
||||
Backends are used by `LocalAttention` and `DistributedAttention`. These layers
|
||||
accept a `supported_attention_backends` tuple. If your backend should be
|
||||
eligible, update the call sites that construct these layers (search for
|
||||
`supported_attention_backends=`).
|
||||
|
||||
## 4) Add compiled kernels (optional)
|
||||
|
||||
If you have a custom CUDA kernel:
|
||||
|
||||
1) Add sources in `fastvideo-kernel/csrc/attention/`.
|
||||
2) Register bindings in `fastvideo-kernel/csrc/common_extension.cpp`.
|
||||
3) Add to `fastvideo-kernel/CMakeLists.txt` (and any feature flags).
|
||||
4) Expose in `fastvideo-kernel/python/fastvideo_kernel/ops.py`.
|
||||
5) Export in `fastvideo-kernel/python/fastvideo_kernel/__init__.py`.
|
||||
|
||||
Keep a Python/Triton fallback so the backend runs even when the kernel is not
|
||||
available.
|
||||
|
||||
## 5) Testing and debugging
|
||||
|
||||
- Add a small parity test or microbenchmark comparing to SDPA.
|
||||
- Force your backend with the env var:
|
||||
`FASTVIDEO_ATTENTION_BACKEND=MY_NEW_ATTN`.
|
||||
- Check logs from `fastvideo/attention/selector.py` to confirm selection.
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] Added enum entry in `fastvideo/platforms/interface.py`.
|
||||
- [ ] Implemented backend in `fastvideo/attention/backends/`.
|
||||
- [ ] Registered selection in platform(s).
|
||||
- [ ] Updated layer call sites to include the backend where appropriate.
|
||||
- [ ] Added tests and documentation.
|
||||
@@ -1,480 +0,0 @@
|
||||
# FastVideo + Coding Agents
|
||||
|
||||
Coding agents are now strong at navigating large codebases and iterating fast
|
||||
with parity tests and examples. This guide shows how to use them to add new
|
||||
model pipelines and ship PRs in a production-grade video diffusion framework.
|
||||
|
||||
FastVideo is a great project to contribute to, with production-grade
|
||||
infrastructure, active collaborations (including NVIDIA), and a pipeline design
|
||||
and inference architecture that has been forked by [SGLang’s
|
||||
multimodal generation stack](https://github.com/sgl-project/sglang/tree/main/python/sglang/multimodal_gen).
|
||||
|
||||
Goal: run the new pipeline with a minimal script like
|
||||
`examples/inference/basic/basic.py`. In production, FastVideo can download
|
||||
models automatically via `HF_HOME`; for development, use local directories so
|
||||
agents can run scripts and tests deterministically. We standardize local paths
|
||||
as:
|
||||
|
||||
- `official_weights/<model_name>/` for official checkpoints
|
||||
- `converted_weights/<model_name>/` if conversion is required
|
||||
|
||||
## Tips when prompting the agent
|
||||
|
||||
When prompting the agent, include:
|
||||
|
||||
- This guide and the [FastVideo design overview](../design/overview.md).
|
||||
- Exact file paths to edit.
|
||||
- A closest reference example file in FastVideo.
|
||||
- Expected behavior and acceptance criteria.
|
||||
- Repro steps (command, inputs, logs).
|
||||
- Constraints (performance, memory, compatibility).
|
||||
- Local paths (e.g., `official_weights/<model_name>/` or
|
||||
`converted_weights/<model_name>/`) for parity tests.
|
||||
|
||||
## FastVideo structure at a glance
|
||||
|
||||
Before diving in, scan these references:
|
||||
|
||||
- [Contributing overview](overview.md) for environment/setup context.
|
||||
- [FastVideo design overview](../design/overview.md) for pipeline architecture, configs, and HF layout.
|
||||
|
||||
FastVideo maps a Diffusers-style repo into a pipeline like:
|
||||
|
||||
- `fastvideo/models/*`: model implementations (DiT, VAE, encoders, upsamplers).
|
||||
- `fastvideo/configs/models/*`: arch configs and `param_names_mapping` for
|
||||
weight name translation.
|
||||
- `fastvideo/configs/pipelines/*`: pipeline wiring (component classes + names).
|
||||
- `fastvideo/configs/sample/*`: default runtime sampling parameters.
|
||||
- `fastvideo/pipelines/basic/*`: end-to-end pipeline logic built from stages.
|
||||
- `model_index.json`: the HF repo entrypoint that maps component names to
|
||||
classes and weight files.
|
||||
- Component loading happens in `VideoGenerator.from_pretrained`, which reads
|
||||
`model_index.json`, resolves configs, and loads weights.
|
||||
|
||||
Minimal usage example (based on `examples/inference/basic/basic.py`):
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.sample import SamplingParam
|
||||
|
||||
model_id = "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" # or official_weights/<model_name>/
|
||||
generator = VideoGenerator.from_pretrained(model_id, num_gpus=1)
|
||||
|
||||
sampling = SamplingParam.from_pretrained(model_id)
|
||||
sampling.num_frames = 45
|
||||
video = generator.generate_video(
|
||||
"A vibrant city street at sunset.",
|
||||
sampling_param=sampling,
|
||||
output_path="video_samples",
|
||||
save_video=True,
|
||||
)
|
||||
```
|
||||
|
||||
## Some questions to ask yourself before starting
|
||||
|
||||
Answering these upfront clarifies the work and speeds up implementation.
|
||||
|
||||
### Is the model already supported by SGLang's multimodal generation stack?
|
||||
If yes, you can port many components from SGLang. It is a FastVideo fork, so
|
||||
interfaces line up, but you still need to swap layers/modules to match
|
||||
FastVideo's architecture and attention stack.
|
||||
|
||||
If not, implement the model directly in FastVideo.
|
||||
|
||||
### Is there an official implementation of the model you are adding?
|
||||
|
||||
If yes, use it as the numerical reference. For example, LTX‑2 has an official
|
||||
implementation here: https://github.com/Lightricks/LTX-2. Prefer official code
|
||||
even if Diffusers also has one.
|
||||
|
||||
### Is there a HuggingFace repo for the model you are adding? Is it in Diffusers format?
|
||||
|
||||
If yes, load it directly in FastVideo after setting tensor mapping rules in the
|
||||
config. Otherwise, convert the weights to Diffusers format. See [Weights and
|
||||
Diffusers format](../design/overview.md#weights-and-diffusers-format) for details.
|
||||
|
||||
### Do I have official weights + local paths ready?
|
||||
|
||||
Standardize local paths as:
|
||||
|
||||
- `official_weights/<model_name>/` for official checkpoints
|
||||
- `converted_weights/<model_name>/` if conversion is required (can be created later)
|
||||
|
||||
### What pipeline components are required for the model you are adding?
|
||||
|
||||
Usually you need a transformer (DiT), VAE, text encoder, and tokenizer. Some
|
||||
models add extra components.
|
||||
|
||||
### What tasks does the model support?
|
||||
|
||||
Usually a video diffusion model supports text‑to‑video (T2V),
|
||||
image‑to‑video (I2V), and video‑to‑video (V2V). Some add extra tasks (two‑stage
|
||||
generation, keyframe interpolation), which require extra components.
|
||||
|
||||
It's usually easiest to start with a T2V pipeline and add the other tasks later.
|
||||
|
||||
You can refer to the [Pipeline system](../design/overview.md#pipeline-system)
|
||||
section for more details.
|
||||
|
||||
### Am I able to generate videos with the official implementation?
|
||||
|
||||
These videos and prompts are your reference. Once the FastVideo pipeline works,
|
||||
compare outputs to the official implementation. Due to seeding and other
|
||||
factors, outputs may not match exactly, but they should be comparable.
|
||||
|
||||
## Workflow: adding a full pipeline
|
||||
|
||||
This is an example workflow for adding a full model pipeline (model + configs +
|
||||
examples + tests). This guide is in active development; feedback is welcome.
|
||||
|
||||
!!! note
|
||||
If you get stuck, refer to existing models/pipelines in FastVideo or ask in Slack.
|
||||
|
||||
### 0) Fetch official model's code and weights
|
||||
|
||||
Purpose:
|
||||
|
||||
- Keep official checkpoints and source code local so conversion, parity tests,
|
||||
and reference runs are reproducible.
|
||||
- Clone the official repo so you can use it as a numerical reference.
|
||||
|
||||
Action:
|
||||
|
||||
- Download official weights into `official_weights/<model_name>/`
|
||||
(Diffusers format or not).
|
||||
- Clone the official repo under the project root (e.g., `FastVideo/LTX-2/`).
|
||||
- If a Diffusers-format HF repo already exists, you can skip manual weight
|
||||
handling and download it directly with
|
||||
`scripts/huggingface/download_hf.py`.
|
||||
|
||||
!!! note
|
||||
This step is best done manually because large downloads can time out.
|
||||
Example:
|
||||
```bash
|
||||
python scripts/huggingface/download_hf.py \
|
||||
--repo_id Wan-AI/Wan2.1-T2V-1.3B-Diffusers \
|
||||
--local_dir official_weights/Wan2.1-T2V-1.3B-Diffusers \
|
||||
--repo_type model
|
||||
```
|
||||
|
||||
### 1) Implement the model + config mapping
|
||||
|
||||
Purpose:
|
||||
|
||||
- Model weights are a dictionary of named tensors (`state_dict`). If the names
|
||||
don’t line up with FastVideo’s module names, weights won’t load correctly (or
|
||||
will silently load into the wrong layer).
|
||||
- Official checkpoints often use different prefixes or module layouts than
|
||||
FastVideo, so we translate names via the mapping (during load or conversion).
|
||||
- Mapping aligns three things:
|
||||
1. the official implementation’s module names,
|
||||
2. the checkpoint `state_dict` keys,
|
||||
3. FastVideo’s model classes and layer naming conventions.
|
||||
- If names don’t align, weights won’t load; implement the FastVideo model and
|
||||
define mapping rules first.
|
||||
|
||||
Action:
|
||||
|
||||
- Implement the FastVideo model + config mapping.
|
||||
- Add/extend the model in `fastvideo/models/...` and config in
|
||||
`fastvideo/configs/models/...` (including `param_names_mapping`).
|
||||
- Reuse existing FastVideo layers/modules where possible.
|
||||
- Use FastVideo’s attention layers:
|
||||
- `DistributedAttention` only for full‑sequence self‑attention in the DiT.
|
||||
- `LocalAttention` for cross‑attention and other attention layers.
|
||||
- See the “Configuration System” and “Weights and Diffusers format” sections
|
||||
in `docs/design/overview.md` for how these pieces connect.
|
||||
- If you are using an agent, ask it to implement the model, config mapping,
|
||||
and a parity test together so you can validate numerics immediately.
|
||||
|
||||
!!! note
|
||||
After the first component is aligned and parity‑tested, open a **DRAFT PR**
|
||||
on FastVideo so the rest of the pipeline work can build on top of it.
|
||||
|
||||
!!! note
|
||||
If a Diffusers-format HF repo already exists and loads correctly, you can
|
||||
skip conversion entirely (no conversion script needed) and just download it
|
||||
with `scripts/huggingface/download_hf.py`. Otherwise, you may need a
|
||||
conversion script + a `converted_weights/<model>/` staging directory.
|
||||
|
||||
Example (key renaming via arch config mapping, Wan2.1‑style):
|
||||
|
||||
```python
|
||||
# Official model (simplified) in the upstream repo.
|
||||
class OfficialWanTransformer(torch.nn.Module):
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
self.patch_embedding = torch.nn.Conv3d(16, 1536, kernel_size=2, padding=0)
|
||||
|
||||
def forward(self, x):
|
||||
return self.patch_embedding(x)
|
||||
|
||||
# FastVideo model (simplified) in fastvideo/models/dits/wanvideo.py
|
||||
class PatchEmbed(torch.nn.Module):
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
self.proj = torch.nn.Conv3d(16, 1536, kernel_size=2, padding=0)
|
||||
|
||||
def forward(self, x):
|
||||
return self.proj(x)
|
||||
|
||||
class WanTransformer3DModel(torch.nn.Module):
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
self.patch_embedding = PatchEmbed()
|
||||
|
||||
def forward(self, x):
|
||||
return self.patch_embedding(x)
|
||||
|
||||
# Mapping defined in a config (simplified; see the real mapping in
|
||||
# fastvideo/configs/models/dits/wanvideo.py)
|
||||
param_names_mapping = {
|
||||
r"^patch_embedding\.(.*)$": r"patch_embedding.proj.\1",
|
||||
r"^blocks\.(\d+)\.attn1\.to_q\.(.*)$": r"blocks.\1.to_q.\2",
|
||||
}
|
||||
|
||||
def apply_regex_map(state_dict, mapping):
|
||||
# Pseudocode: apply regex substitutions in order
|
||||
...
|
||||
|
||||
# Official checkpoint keys (example)
|
||||
official = {
|
||||
"patch_embedding.weight": ...,
|
||||
"blocks.0.attn1.to_q.weight": ...,
|
||||
}
|
||||
|
||||
# Apply mapping so keys match FastVideo modules
|
||||
converted = apply_regex_map(official, param_names_mapping)
|
||||
|
||||
```
|
||||
|
||||
Optional helper (print a few checkpoint keys quickly):
|
||||
|
||||
```bash
|
||||
python - <<'PY'
|
||||
import safetensors.torch as st
|
||||
keys = list(st.load_file("official_weights/<model>/transformer/diffusion_pytorch_model.safetensors").keys())
|
||||
print(keys[:20])
|
||||
PY
|
||||
```
|
||||
|
||||
Example agent prompt (task request):
|
||||
|
||||
```
|
||||
Please add the Wan2.1 T2V 1.3B Diffusers pipeline to FastVideo:
|
||||
- Add a FastVideo native Wan2.1 DiT implementation + config mapping.
|
||||
- Make sure to use the existing FastVideo layers and attention modules where possible.
|
||||
- Add a parity test that loads the official model alongside the FastVideo model and compares outputs numerically with fixed seeds and inputs.
|
||||
|
||||
Paths:
|
||||
- Official repo: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
- Local download: official_weights/Wan2.1-T2V-1.3B-Diffusers
|
||||
Mapping steps:
|
||||
- Load the official DiT weights from
|
||||
official_weights/Wan2.1-T2V-1.3B-Diffusers/transformer/diffusion_pytorch_model.safetensors.
|
||||
- Instantiate the FastVideo DiT (`WanTransformer3DModel`) and compare
|
||||
its `state_dict().keys()` to the official keys.
|
||||
- Update `param_names_mapping` in
|
||||
fastvideo/configs/models/dits/wanvideo.py to resolve missing/unexpected keys.
|
||||
- Use `load_state_dict(strict=False)` during iteration to surface mismatches.
|
||||
```
|
||||
|
||||
External examples of the same pattern:
|
||||
- SGLang uses prefix-based routing in its weight loader to map checkpoint keys
|
||||
into internal submodules (e.g., stripping a top-level prefix before delegating).
|
||||
- vLLM includes model-specific renamers for certain checkpoints that adjust
|
||||
key prefixes so weights match its internal naming.
|
||||
|
||||
### 2) Test numerical alignment with the official implementation
|
||||
|
||||
Purpose:
|
||||
|
||||
- Verify that the FastVideo component is numerically aligned with the official
|
||||
implementation.
|
||||
|
||||
Action:
|
||||
|
||||
- Add or reuse a numerical parity test that loads the official model and the
|
||||
FastVideo model and compares outputs.
|
||||
- See examples in `tests/local_tests/` (e.g., `tests/local_tests/upsamplers/`)
|
||||
and the commands in `tests/local_tests/README.md`.
|
||||
- If there are discrepancies, add opt‑in logging to both models and compare
|
||||
activation summaries (layer output sums, per‑stage logs).
|
||||
- First align the loaded weights (validate `param_names_mapping`).
|
||||
- Then align forward outputs using fixed seeds and inputs.
|
||||
- Start with `atol=1e-4, rtol=1e-4` in `assert_close`.
|
||||
- Keep dtype consistent (bf16 if available; otherwise fp32).
|
||||
- If attention parity is unstable, align backends (e.g.,
|
||||
`FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA`).
|
||||
|
||||
### 3) Repeat the process for each component
|
||||
|
||||
If the model requires additional components, repeat Steps 1–2 for each one.
|
||||
For example, implement the VAE in `fastvideo/models/vaes/` and its config in
|
||||
`fastvideo/configs/models/vaes/`, then add parity coverage for it.
|
||||
|
||||
### 4) Add a pipeline config + sample defaults
|
||||
|
||||
Purpose:
|
||||
|
||||
- `fastvideo/configs/pipelines/` describes pipeline wiring and model module
|
||||
names.
|
||||
- `fastvideo/configs/sample/` defines default runtime parameters.
|
||||
|
||||
Action:
|
||||
|
||||
- Add a new pipeline config + sampling params.
|
||||
- Register them in `fastvideo/registry.py` using explicit
|
||||
`register_configs(...)` blocks (this file is the single source of truth now).
|
||||
|
||||
### 5) Wire pipeline stages
|
||||
|
||||
Purpose:
|
||||
|
||||
- `fastvideo/pipelines/basic/<pipeline>/` contains the actual pipeline logic.
|
||||
- `fastvideo/pipelines/stages/` holds reusable, testable stages.
|
||||
|
||||
Action:
|
||||
|
||||
- Build the pipeline using stages; keep new stages isolated and documented.
|
||||
- Prefer opt‑in flags for expensive or optional steps.
|
||||
|
||||
### 6) Add pipeline‑level tests
|
||||
|
||||
Purpose:
|
||||
|
||||
- Ensure the end‑to‑end pipeline works and stays aligned as pieces evolve.
|
||||
|
||||
Action:
|
||||
|
||||
- Add a pipeline parity test under `tests/local_tests/pipelines/`.
|
||||
- See the [Testing Guide](testing.md) for test conventions.
|
||||
|
||||
### 7) Add user‑facing examples
|
||||
|
||||
Purpose:
|
||||
|
||||
- `examples/inference/basic/` is the entry point for simple, runnable scripts.
|
||||
|
||||
Action:
|
||||
|
||||
- Provide a minimal “hello world” example plus advanced variations.
|
||||
- Use fixed seeds and stable prompts.
|
||||
- Run the example locally to confirm end‑to‑end behavior.
|
||||
|
||||
### 8) Add SSIM tests for CI checks
|
||||
|
||||
Purpose:
|
||||
|
||||
- Ensure visual similarity stays within expected bounds for regression testing.
|
||||
- SSIM tests act as a higher‑level guardrail beyond unit/parity tests.
|
||||
|
||||
Action:
|
||||
|
||||
- Add SSIM tests under `fastvideo/tests/ssim/` and include reference videos
|
||||
(see the structure in the Testing Guide).
|
||||
- Use stable prompts/seeds and document any GPU‑specific requirements.
|
||||
- Follow the [Testing Guide](testing.md) for reference video placement and
|
||||
execution details.
|
||||
|
||||
### 9) Document it
|
||||
|
||||
Purpose:
|
||||
|
||||
- `docs/` is where users find the new pipeline usage and limitations.
|
||||
|
||||
Action:
|
||||
|
||||
- Add a short doc page or update an existing one.
|
||||
- Mention any caveats (memory, speed, constraints).
|
||||
|
||||
## Common pitfalls when porting models
|
||||
|
||||
- **Attention backend mismatch**: parity can fail if the official model uses a
|
||||
different attention backend (e.g., SDPA vs custom). Align backends before
|
||||
debugging deeper issues.
|
||||
- **Patchifier shape mistakes**: wrong patchification or reshape lengths can
|
||||
silently corrupt outputs. Validate patch shapes early.
|
||||
- **Mask handling**: attention masks must match the official behavior (padding,
|
||||
causal masks, and broadcast shapes).
|
||||
- **Scheduler / sigma schedule mismatch**: even small differences in schedules
|
||||
or timestep shapes can cause noticeable drift.
|
||||
|
||||
## Diffusers vs manual conversion
|
||||
|
||||
If a model already ships in Diffusers format (with a proper `model_index.json`),
|
||||
prefer downloading it directly and loading it via FastVideo. In that case:
|
||||
|
||||
- You usually **do not need** a conversion script.
|
||||
- You still need a correct `param_names_mapping` if the internal module names
|
||||
differ from FastVideo’s implementation.
|
||||
|
||||
If the model does **not** have a Diffusers-format repo:
|
||||
|
||||
- You will need a conversion script to rewrite `state_dict` keys into FastVideo
|
||||
naming and stage the result (e.g., under `converted_weights/<model>/`).
|
||||
- You may still use the official repo for reference parity and debugging.
|
||||
|
||||
In both cases, parity testing is required to validate correctness.
|
||||
|
||||
If you want to publish a Diffusers‑style repo after conversion, use
|
||||
`scripts/checkpoint_conversion/create_hf_repo.py` to assemble a HuggingFace‑ready
|
||||
directory before uploading.
|
||||
|
||||
## FAQ
|
||||
|
||||
**Q: Why do we implement the FastVideo model before conversion?**
|
||||
A: You can’t define the key‑mapping rules until the FastVideo module names are
|
||||
known. The implementation determines the target `state_dict` schema.
|
||||
|
||||
**Q: Do we always need a conversion script?**
|
||||
A: No. If a Diffusers‑format repo exists and loads correctly, download it and
|
||||
skip conversion.
|
||||
|
||||
**Q: How do I figure out `param_names_mapping` quickly?**
|
||||
A: Load the official weights, instantiate the FastVideo model, and diff
|
||||
`state_dict().keys()` on both sides. Add regex rules until missing/unexpected
|
||||
keys are resolved. Agents can help you with this.
|
||||
|
||||
**Q: What if parity fails even after mapping?**
|
||||
A: Align attention backends, sigma schedules, and timestep shapes first. Then
|
||||
add opt‑in activation logging to locate the first divergent layer.
|
||||
|
||||
## Case study: LTX‑2 port (from PLAN.md)
|
||||
|
||||
The LTX‑2 port in `PLAN.md` shows the real sequence of steps and backtracking
|
||||
that happened during integration. Use it as a reference for how parity work
|
||||
actually unfolds:
|
||||
|
||||
- Ported components first (transformer, VAE, audio, text encoder).
|
||||
- Added parity tests per component; used SDPA for reference parity.
|
||||
- Added debug logging to compare per‑block activations and isolate divergence.
|
||||
- Fixed cross‑attention reshape and patch grid bounds issues after logging.
|
||||
- Aligned sigma schedule and masking behavior to match the official pipeline.
|
||||
|
||||
Recommendation:
|
||||
|
||||
- Keep raw step‑by‑step logs in your own local `PLAN.md` for large ports.
|
||||
|
||||
## Worked example: Wan2.1 T2V 1.3B pipeline
|
||||
|
||||
The Wan2.1 T2V 1.3B Diffusers pipeline is a good “standard” example for
|
||||
FastVideo integration.
|
||||
|
||||
1. Verify model config + mapping.
|
||||
- DiT mapping: `fastvideo/configs/models/dits/wanvideo.py`
|
||||
- VAE: `fastvideo/models/vaes/wanvae.py`
|
||||
- Text encoder: `fastvideo/models/encoders/t5.py`
|
||||
|
||||
2. Parity test the core components.
|
||||
- Example tests: `fastvideo/tests/transformers/test_wanvideo.py`,
|
||||
`fastvideo/tests/vaes/test_wan_vae.py`,
|
||||
`fastvideo/tests/encoders/test_t5_encoder.py`
|
||||
|
||||
3. Pipeline wiring.
|
||||
- Pipeline: `fastvideo/pipelines/basic/wan/wan_pipeline.py`
|
||||
- Pipeline config: `fastvideo/configs/pipelines/wan.py`
|
||||
- Sampling defaults: `fastvideo/configs/sample/wan.py`
|
||||
|
||||
4. Minimal example.
|
||||
- Script: `examples/inference/basic/basic.py`
|
||||
@@ -1,55 +1,16 @@
|
||||
|
||||
# 📦 Developing FastVideo on RunPod
|
||||
|
||||
You can easily use the FastVideo Pod Template on [RunPod](https://www.runpod.io) for development or experimentation.
|
||||
You can easily use the FastVideo Docker image as a custom container on [RunPod](https://www.runpod.io) for development or experimentation.
|
||||
|
||||
## Creating a new pod
|
||||
|
||||
- Make sure you are using the correct RunPod account.
|
||||

|
||||
Choose a GPU that supports CUDA 12.8
|
||||
|
||||
Pick 1 or 2 L40S GPU(s)
|
||||
|
||||
- Use "Additional Filters" to select CUDA 12.8.
|
||||

|
||||
|
||||
- Click "Deploy" and Pick a single A40 or RTX 4090 GPU.
|
||||

|
||||
|
||||
- Select the "FastVideo" or "fastvideo-dev" Pod Template.
|
||||

|
||||
|
||||
- Set the Pod name to "`<name>-<FastVideo>-<date>`".
|
||||
|
||||
- Finally, once the pod is deployed (will take a few minutes as the image is being pulled), you can SSH into it using "SSH exposed over TCP". You'll need to use the matching private ssh key you provided.
|
||||

|
||||
|
||||
## Working with the pod
|
||||
|
||||
After SSH'ing into your pod, you'll find the correct `uv` environment already activated and you should be in /FastVideo/ directory. Make sure to use /FastVideo/ for all your work.
|
||||
|
||||
To pull in the latest changes from the GitHub repo:
|
||||
|
||||
```bash
|
||||
cd /FastVideo
|
||||
git pull
|
||||
```
|
||||
|
||||
Run your development workflows as usual:
|
||||
|
||||
```bash
|
||||
# Run linters
|
||||
pre-commit run --all-files
|
||||
|
||||
# Run tests
|
||||
pytest tests/
|
||||
```
|
||||
|
||||
Make sure to push your changes back to the GitHub repo as nothing will be saved to the pod when it is terminated.
|
||||
|
||||
After you are done with your work, you can terminate the pod by clicking the "Terminate" and "Delete" buttons. Remember if the pod is not completely deleted, Runpod will keep charging you for it.
|
||||
|
||||
## Extra Information:
|
||||
If you need to customize the pod template this section has some useful information. For the most part you can leave the defaults of the FastVideo Pod Template.
|
||||
|
||||
When creating your pod template, use this image:
|
||||
|
||||
```
|
||||
@@ -63,3 +24,30 @@ bash -c "apt update;DEBIAN_FRONTEND=noninteractive apt-get install openssh-serve
|
||||
```
|
||||
|
||||

|
||||
|
||||
After deploying, the pod will take a few minutes to pull the image and start the SSH service.
|
||||
|
||||

|
||||
|
||||
## Working with the pod
|
||||
|
||||
After SSH'ing into your pod, you'll find the `fastvideo-dev` Conda environment already activated.
|
||||
|
||||
To pull in the latest changes from the GitHub repo:
|
||||
|
||||
```bash
|
||||
cd /FastVideo
|
||||
git pull
|
||||
```
|
||||
|
||||
`If you have a persistent volume and want to keep your code changes, you can move /FastVideo to /workspace/FastVideo, or simply clone the repository there.`
|
||||
|
||||
Run your development workflows as usual:
|
||||
|
||||
```bash
|
||||
# Run linters
|
||||
pre-commit run --all-files
|
||||
|
||||
# Run tests
|
||||
pytest tests/
|
||||
```
|
||||
|
||||
@@ -1,85 +1,76 @@
|
||||
|
||||
# 🛠️ Contributing to FastVideo
|
||||
|
||||
Thank you for your interest in contributing to FastVideo. We want the process
|
||||
to be smooth and beginner‑friendly, whether you are adding a new pipeline,
|
||||
improving performance, or fixing a bug.
|
||||
Thank you for your interest in contributing to FastVideo. We want to make the process as smooth for you as possible and this is a guide to help get you started!
|
||||
|
||||
## Quick prerequisites
|
||||
Our community is open to everyone and welcomes any contributions no matter how large or small.
|
||||
|
||||
- **OS**: Linux is the primary development target (WSL can work).
|
||||
- **GPU**: NVIDIA GPU recommended for inference and training workflows.
|
||||
- **CUDA**: Use a recent CUDA 12.x toolchain (see the installation guide for
|
||||
the current recommendation).
|
||||
# Developer Environment:
|
||||
Do make sure you have CUDA 12.4 installed and supported. FastVideo currently only supports Linux and CUDA GPUs, but we hope to support other platforms in the future.
|
||||
|
||||
For a full install checklist, see `docs/getting_started/installation/gpu.md`.
|
||||
We recommend using a fresh Python 3.10 Conda environment to develop FastVideo:
|
||||
|
||||
## Local development (UV + editable install)
|
||||
Install Miniconda:
|
||||
|
||||
If you previously used Conda for local setup, switch to uv for a faster and more stable development environment.
|
||||
|
||||
Install `uv`:
|
||||
|
||||
```bash
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
# or
|
||||
wget -qO- https://astral.sh/uv/install.sh | sh
|
||||
```
|
||||
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
|
||||
bash Miniconda3-latest-Linux-x86_64.sh
|
||||
source ~/.bashrc
|
||||
```
|
||||
|
||||
Create and activate a uv environment (recommended):
|
||||
Create and activate a Conda environment for FastVideo:
|
||||
|
||||
```bash
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
Conda alternative (supported):
|
||||
|
||||
```bash
|
||||
conda create -n fastvideo python=3.12 -y
|
||||
conda activate fastvideo
|
||||
```
|
||||
|
||||
Clone the repo:
|
||||
Install `uv` (optional, but recommended):
|
||||
|
||||
```bash
|
||||
git clone https://github.com/hao-ai-lab/FastVideo.git && cd FastVideo
|
||||
From instructions on [uv](https://astral.sh/uv/):
|
||||
|
||||
```
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
# or
|
||||
wget -qO- https://astral.sh/uv/install.sh | sh
|
||||
```
|
||||
|
||||
Install FastVideo in editable mode and set up hooks:
|
||||
Clone the FastVideo repository and go to the FastVideo directory:
|
||||
|
||||
```
|
||||
git clone https://github.com/hao-ai-lab/FastVideo.git && cd FastVideo
|
||||
|
||||
```
|
||||
|
||||
Now you can install FastVideo and setup git hooks for running linting. By using `pre-commit`, the linters will run and have to pass before you'll be able to make a commit.
|
||||
|
||||
```bash
|
||||
uv pip install -e .[dev]
|
||||
|
||||
# Optional: FlashAttention (builds native kernels)
|
||||
uv pip install flash-attn --no-build-isolation -v
|
||||
# Can also install flash-attn (optional)
|
||||
uv pip install flash-attn --no-build-isolation
|
||||
|
||||
# Linting, formatting, static typing
|
||||
# Linting, formatting and static type checking
|
||||
pre-commit install --hook-type pre-commit --hook-type commit-msg
|
||||
|
||||
# You can manually run pre-commit with
|
||||
pre-commit run --all-files
|
||||
|
||||
# Unit tests
|
||||
pytest tests/
|
||||
```
|
||||
|
||||
If you are on a Hopper GPU, installing FlashAttention 3 can improve
|
||||
performance (see `docs/inference/optimizations.md`).
|
||||
If you are on a Hopper GPU, you should also install [FA3](https://github.com/Dao-AILab/flash-attention) for much better performance:
|
||||
|
||||
## Docker development (optional)
|
||||
```
|
||||
git clone https://github.com/Dao-AILab/flash-attention.git && cd flash-attention/hopper
|
||||
|
||||
If you prefer a containerized environment, use the dev image documented in
|
||||
`docs/contributing/developer_env/docker.md`.
|
||||
# make sure you have ninja installed
|
||||
uv pip install ninja
|
||||
|
||||
python setup.py install
|
||||
```
|
||||
|
||||
## Testing
|
||||
|
||||
See the [Testing Guide](testing.md) for how to add and run tests in FastVideo.
|
||||
|
||||
## Attention backend development
|
||||
|
||||
If you are adding a new attention kernel or backend, follow
|
||||
[Attention Backend Development](attention_backend.md).
|
||||
|
||||
## Contributing with coding agents
|
||||
|
||||
For a step‑by‑step workflow on adding pipelines or components with coding
|
||||
agents, see `docs/contributing/coding_agents.md`.
|
||||
Please refer to the [Testing Guide](testing.md) for more information on how to add and run tests in FastVideo.
|
||||
|
||||
@@ -8,7 +8,7 @@ This guide explains how to add and run tests in FastVideo. The testing suite is
|
||||
* **Component Tests**: Located in `fastvideo/tests/encoders`, `fastvideo/tests/transformers`, and `fastvideo/tests/vaes`. These verify the loading and basic functionality of model components.
|
||||
* **SSIM Tests**: Located in `fastvideo/tests/ssim`. These are regression tests that compare generated videos against reference videos using the Structural Similarity Index Measure (SSIM) to detect quality degradation.
|
||||
* **Training Tests**: Located in `fastvideo/tests/training`. These validate training loops, loss calculations, and specific training techniques like LoRA, Distillation, and VSA.
|
||||
* **Inference Tests**: Located in `fastvideo/tests/inference`. These test specialized inference pipelines and optimizations (e.g., VSA, V-MoBA).
|
||||
* **Inference Tests**: Located in `fastvideo/tests/inference`. These test specialized inference pipelines and optimizations (e.g., STA, V-MoBA).
|
||||
|
||||
For now, we will focus on **SSIM Tests**.
|
||||
|
||||
|
||||
@@ -1,179 +1,424 @@
|
||||
# FastVideo Architecture Overview
|
||||
# 🔍 FastVideo Overview
|
||||
|
||||
This document summarizes how FastVideo is structured and how a Diffusers-style
|
||||
model repo maps into a runnable pipeline. It is intended for contributors who
|
||||
need the high-level layout and key entrypoints, not every internal detail.
|
||||
This document outlines FastVideo's architecture for developers interested in framework internals or contributions. It serves as an onboarding guide for new contributors by providing an overview of the most important directories and files within the `fastvideo/` codebase.
|
||||
|
||||
## FastVideo structure at a glance
|
||||
## Table of Contents - Directory Structure and Files
|
||||
|
||||
FastVideo maps a Diffusers-style repo into a pipeline like this:
|
||||
- [`fastvideo/pipelines/`](#pipeline-system) - Core diffusion pipeline components
|
||||
- [`fastvideo/models/`](#model-components) - Model implementations
|
||||
- [`dits/`](#transformer-models) - Transformer-based diffusion models
|
||||
- [`vaes/`](#vae-variational-auto-encoder) - Variational autoencoders
|
||||
- [`encoders/`](#text-and-image-encoders) - Text and image encoders
|
||||
- [`schedulers/`](#schedulers) - Diffusion schedulers
|
||||
- [`fastvideo/attention/`](#optimized-attention) - Optimized attention implementations
|
||||
- [`fastvideo/distributed/`](#distributed-processing) - Distributed computing utilities
|
||||
- [`fastvideo/layers/`](#tensor-parallelism) - Custom neural network layers
|
||||
- [`fastvideo/platforms/`](#platforms) - Hardware platform abstractions
|
||||
- [`fastvideo/worker/`](#executor-and-worker-system) - Multi-GPU process management
|
||||
- [`fastvideo/fastvideo_args.py`](#fastvideoargs) - Argument handling
|
||||
- [`fastvideo/forward_context.py`](#forward-context-management) - Forward pass context management
|
||||
- `fastvideo/utils.py` - Utility functions
|
||||
- [`fastvideo/logger.py`](#logger) - Logging infrastructure
|
||||
|
||||
- `fastvideo/models/*`: model implementations (DiT, VAE, encoders, upsamplers).
|
||||
- `fastvideo/configs/models/*`: arch configs and `param_names_mapping` for
|
||||
weight name translation.
|
||||
- `fastvideo/configs/pipelines/*`: pipeline wiring (component classes + names).
|
||||
- `fastvideo/configs/sample/*`: default runtime sampling parameters.
|
||||
- `fastvideo/pipelines/basic/*`: end-to-end pipelines.
|
||||
- `fastvideo/pipelines/stages/*`: reusable pipeline stages.
|
||||
- `fastvideo/models/loader/*`: component loaders for Diffusers-style repos.
|
||||
- `model_index.json`: HF repo entrypoint mapping component names to classes.
|
||||
## Core Architecture
|
||||
|
||||
Flow:
|
||||
`model_index.json` -> component loaders -> model modules -> pipeline stages ->
|
||||
sampling params.
|
||||
FastVideo separates model components from execution logic with these principles:
|
||||
|
||||
Minimal usage (from `examples/inference/basic/basic.py`):
|
||||
- **Component Isolation**: Models (encoders, VAEs, transformers) are isolated from execution (pipelines, stages, distributed processing)
|
||||
- **Modular Design**: Components can be independently replaced
|
||||
- **Distributed Execution**: Supports various parallelism strategies (Tensor, Sequence)
|
||||
- **Custom Attention Backends**: Components can support and use different Attention implementations
|
||||
- **Pipeline Abstraction**: Consistent interface across diffusion models
|
||||
|
||||
## FastVideoArgs
|
||||
|
||||
The `FastVideoArgs` class in `fastvideo/fastvideo_args.py` serves as the central configuration system for FastVideo. It contains all parameters needed to control model loading, inference configuration, performance optimization settings, and more.
|
||||
|
||||
Key features include:
|
||||
|
||||
- **Command-line Interface**: Automatic conversion between CLI arguments and dataclass fields
|
||||
- **Configuration Groups**: Organized by functional areas (model loading, video params, optimization settings)
|
||||
- **Context Management**: Global access to current settings via `get_current_fastvideo_args()`
|
||||
- **Parameter Validation**: Ensures valid combinations of settings
|
||||
|
||||
Common configuration areas:
|
||||
|
||||
- **Model paths and loading options**: `model_path`, `trust_remote_code`, `revision`
|
||||
- **Distributed execution settings**: `num_gpus`, `tp_size`, `sp_size`
|
||||
- **Video generation parameters**: `height`, `width`, `num_frames`, `num_inference_steps`
|
||||
- **Precision settings**: Control computation precision for different components
|
||||
|
||||
Example usage:
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.sample import SamplingParam
|
||||
# Load arguments from command line
|
||||
fastvideo_args = prepare_fastvideo_args(sys.argv[1:])
|
||||
|
||||
model_id = "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" # or official_weights/<model_name>/
|
||||
generator = VideoGenerator.from_pretrained(model_id, num_gpus=1)
|
||||
# Access parameters
|
||||
model = load_model(fastvideo_args.model_path)
|
||||
|
||||
sampling = SamplingParam.from_pretrained(model_id)
|
||||
sampling.num_frames = 45
|
||||
video = generator.generate_video(
|
||||
"A vibrant city street at sunset.",
|
||||
sampling_param=sampling,
|
||||
output_path="video_samples",
|
||||
save_video=True,
|
||||
# Set as global context
|
||||
with set_current_fastvideo_args(fastvideo_args):
|
||||
# Code that requires access to these arguments
|
||||
result = generate_video()
|
||||
```
|
||||
|
||||
## Pipeline System
|
||||
|
||||
### `ComposedPipelineBase`
|
||||
|
||||
This foundational class provides:
|
||||
|
||||
- **Model Loading**: Automatically loads components from HuggingFace-Diffusers-compatible model directories
|
||||
- **Stage Management**: Creates and orchestrates processing stages
|
||||
- **Data Flow Coordination**: Ensures proper state flow between stages
|
||||
|
||||
```python
|
||||
class MyCustomPipeline(ComposedPipelineBase):
|
||||
_required_config_modules = [
|
||||
"text_encoder", "tokenizer", "vae", "transformer", "scheduler"
|
||||
]
|
||||
|
||||
def initialize_pipeline(self, fastvideo_args: FastVideoArgs):
|
||||
# Pipeline-specific initialization
|
||||
pass
|
||||
|
||||
def create_pipeline_stages(self, fastvideo_args: FastVideoArgs):
|
||||
self.add_stage("input_validation_stage", InputValidationStage())
|
||||
self.add_stage("text_encoding_stage", CLIPTextEncodingStage(
|
||||
text_encoder=self.get_module("text_encoder"),
|
||||
tokenizer=self.get_module("tokenizer")
|
||||
))
|
||||
# Additional stages...
|
||||
```
|
||||
|
||||
### Pipeline Stages
|
||||
|
||||
Each stage handles a specific diffusion process component:
|
||||
|
||||
- **Input Validation**: Parameter verification
|
||||
- **Text Encoding**: CLIP, LLaMA, or T5-based encoding
|
||||
- **Image Encoding**: Image input processing
|
||||
- **Timestep & Latent Preparation**: Setup for diffusion
|
||||
- **Denoising**: Core diffusion loop
|
||||
- **Decoding**: Latent-to-pixel conversion
|
||||
|
||||
Each stage implements a standard interface:
|
||||
|
||||
```python
|
||||
def forward(self, batch: ForwardBatch, fastvideo_args: FastVideoArgs) -> ForwardBatch:
|
||||
# Process batch and update state
|
||||
return batch
|
||||
```
|
||||
|
||||

|
||||
|
||||
### ForwardBatch
|
||||
|
||||
Defined in `fastvideo/pipelines/pipeline_batch_info.py`, `ForwardBatch` encapsulates the data payload passed between pipeline stages. It typically holds:
|
||||
|
||||
- **Input Data**: Prompts, images, generation parameters
|
||||
- **Intermediate State**: Embeddings, latents, timesteps, accumulated during stage execution
|
||||
- **Output Storage**: Generated results and metadata
|
||||
- **Configuration**: Sampling parameters, precision settings
|
||||
|
||||
This structure facilitates clear state transitions between stages.
|
||||
|
||||
## Model Components
|
||||
|
||||
The `fastvideo/models/` directory contains implementations of the core neural network models used in video diffusion:
|
||||
|
||||
### Transformer Models
|
||||
|
||||
Transformer networks perform the actual denoising during diffusion:
|
||||
|
||||
- **Location**: `fastvideo/models/dits/`
|
||||
- **Examples**:
|
||||
- `WanTransformer3DModel`
|
||||
- `HunyuanVideoTransformer3DModel`
|
||||
|
||||
Features include:
|
||||
|
||||
- Text/image conditioning
|
||||
- Standardized interface for model-specific optimizations
|
||||
|
||||
```python
|
||||
def forward(
|
||||
self,
|
||||
latents, # [B, T, C, H, W]
|
||||
encoder_hidden_states, # Text embeddings
|
||||
timestep, # Current diffusion timestep
|
||||
encoder_hidden_states_image=None, # Optional image embeddings
|
||||
**kwargs
|
||||
):
|
||||
# Perform denoising computation
|
||||
return noise_pred # Predicted noise residual
|
||||
```
|
||||
|
||||
### VAE (Variational Auto-Encoder)
|
||||
|
||||
VAEs handle conversion between pixel space and latent space:
|
||||
|
||||
- **Location**: `fastvideo/models/vaes/`
|
||||
- **Examples**:
|
||||
- `AutoencoderKLWan`
|
||||
- `AutoencoderKLHunyuanVideo`
|
||||
|
||||
These models compress image/video data to a more efficient latent representation (typically 4x-8x smaller in each dimension).
|
||||
|
||||
FastVideo's VAE implementations include:
|
||||
|
||||
- Efficient video batch processing
|
||||
- Memory optimization
|
||||
- Optional tiling for large frames
|
||||
- Distributed weight support
|
||||
|
||||
### Text and Image Encoders
|
||||
|
||||
Encoders process conditioning inputs into embeddings:
|
||||
|
||||
- **Location**: `fastvideo/models/encoders/`
|
||||
- **Text Encoders**:
|
||||
- `CLIPTextModel`
|
||||
- `LlamaModel`
|
||||
- `UMT5EncoderModel`
|
||||
- **Image Encoders**:
|
||||
- `CLIPVisionModel`
|
||||
|
||||
FastVideo implements optimizations such as:
|
||||
|
||||
- Vocab parallelism for distributed processing
|
||||
- Caching for common prompts
|
||||
- Precision-tuned computation
|
||||
|
||||
### Schedulers
|
||||
|
||||
Schedulers manage the diffusion sampling process:
|
||||
|
||||
- **Location**: `fastvideo/models/schedulers/`
|
||||
- **Examples**:
|
||||
- `UniPCMultistepScheduler`
|
||||
- `FlowMatchEulerDiscreteScheduler`
|
||||
|
||||
These components control:
|
||||
|
||||
- Diffusion timestep sequences
|
||||
- Noise prediction to latent update conversions
|
||||
- Quality/speed trade-offs
|
||||
|
||||
```python
|
||||
def step(
|
||||
self,
|
||||
model_output: torch.Tensor,
|
||||
timestep: torch.LongTensor,
|
||||
sample: torch.Tensor,
|
||||
**kwargs
|
||||
) -> torch.Tensor:
|
||||
# Process model output and update latents
|
||||
# Return updated latents
|
||||
return prev_sample
|
||||
```
|
||||
|
||||
This diagram shows how models are discovered, validated, and loaded across entrypoints, executors, pipelines, and model loaders.
|
||||
|
||||

|
||||
|
||||
## Optimized Attention
|
||||
|
||||
The `fastvideo/attention/` directory contains optimized attention implementations crucial for efficient video diffusion:
|
||||
|
||||
### Attention Backends
|
||||
|
||||
Multiple implementations with automatic selection:
|
||||
|
||||
- **FLASH_ATTN**: Optimized for supporting hardware
|
||||
- **TORCH_SDPA**: Built-in PyTorch scaled dot-product attention
|
||||
- **SLIDING_TILE_ATTN**: For very long sequences
|
||||
|
||||
```python
|
||||
# Configure available attention backends for this layer
|
||||
self.attn = LocalAttention(
|
||||
num_heads=num_heads,
|
||||
head_size=head_dim,
|
||||
causal=False,
|
||||
supported_attention_backends=(_Backend.FLASH_ATTN, _Backend.TORCH_SDPA)
|
||||
)
|
||||
|
||||
# Override via environment variable
|
||||
# export FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN
|
||||
```
|
||||
|
||||

|
||||
|
||||
### Attention Patterns
|
||||
|
||||
Supports various patterns with memory optimization techniques:
|
||||
|
||||
- **Cross/Self/Temporal/Global-Local Attention**
|
||||
- Chunking, progressive computation, optimized masking
|
||||
|
||||
## Distributed Processing
|
||||
|
||||
The `fastvideo/distributed/` directory contains implementations for distributed model execution:
|
||||
|
||||
### Tensor Parallelism
|
||||
|
||||
Tensor parallelism splits model weights across devices:
|
||||
|
||||
- **Implementation**: Through `RowParallelLinear` and `ColumnParallelLinear` layers
|
||||
- **Use cases**: Will be used by encoder models as their sequence lengths are shorter and enables efficient sharding.
|
||||
|
||||
```python
|
||||
# Tensor-parallel layers in a transformer block
|
||||
from fastvideo.layers.linear import ColumnParallelLinear, RowParallelLinear
|
||||
|
||||
# Split along output dimension
|
||||
self.qkv_proj = ColumnParallelLinear(
|
||||
input_size=hidden_size,
|
||||
output_size=3 * hidden_size,
|
||||
bias=True,
|
||||
gather_output=False
|
||||
)
|
||||
|
||||
# Split along input dimension
|
||||
self.out_proj = RowParallelLinear(
|
||||
input_size=hidden_size,
|
||||
output_size=hidden_size,
|
||||
bias=True,
|
||||
input_is_parallel=True
|
||||
)
|
||||
```
|
||||
|
||||
## Configuration system
|
||||
### Sequence Parallelism
|
||||
|
||||
FastVideo uses typed configs to keep model definitions, pipeline wiring, and
|
||||
runtime parameters consistent:
|
||||
Sequence parallelism splits sequences across devices:
|
||||
|
||||
- `fastvideo/configs/models/`: architecture definitions, layer shapes, and
|
||||
`param_names_mapping` rules for key renaming.
|
||||
- `fastvideo/configs/pipelines/`: pipeline wiring and required components.
|
||||
- `fastvideo/configs/sample/`: default sampling parameters (steps, frames,
|
||||
guidance scale, resolution, fps).
|
||||
- `fastvideo/registry.py`: unified registry for pipeline config + sampling
|
||||
defaults and model metadata resolution, defined via explicit
|
||||
`register_configs(...)` blocks (no separate dict registries).
|
||||
- **Implementation**: Through `DistributedAttention` and sequence splitting
|
||||
- **Use cases**: Long video sequences or high-resolution processing. Used by DiT models.
|
||||
|
||||
`FastVideoArgs` (in `fastvideo/fastvideo_args.py`) provides runtime settings and
|
||||
is passed into pipeline construction and stages.
|
||||
```python
|
||||
# Distributed attention for long sequences
|
||||
from fastvideo.attention import DistributedAttention
|
||||
|
||||
## Weights and Diffusers format
|
||||
|
||||
FastVideo follows the HuggingFace Diffusers repo layout. This keeps loaders
|
||||
compatible with HF repos and makes it easy to add new components.
|
||||
|
||||
Typical Diffusers repo:
|
||||
|
||||
```
|
||||
<model-repo>/
|
||||
model_index.json
|
||||
scheduler/
|
||||
scheduler_config.json
|
||||
transformer/ # or unet/ for image models
|
||||
config.json
|
||||
diffusion_pytorch_model.safetensors
|
||||
vae/
|
||||
config.json
|
||||
diffusion_pytorch_model.safetensors
|
||||
text_encoder/
|
||||
config.json
|
||||
model.safetensors
|
||||
tokenizer/
|
||||
tokenizer_config.json
|
||||
tokenizer.json
|
||||
self.attn = DistributedAttention(
|
||||
num_heads=num_heads,
|
||||
head_size=head_dim,
|
||||
causal=False,
|
||||
supported_attention_backends=(_Backend.SLIDING_TILE_ATTN, _Backend.FLASH_ATTN)
|
||||
)
|
||||
```
|
||||
|
||||
Key points:
|
||||
### Communication Primitives
|
||||
|
||||
- `model_index.json` is the root map that tells FastVideo which components to
|
||||
load and which classes implement them.
|
||||
- Each component lives in its own folder with a `config.json` and weights.
|
||||
- Weights are usually in `diffusion_pytorch_model.safetensors`.
|
||||
Efficient distributed operations via AllGather, AllReduce, and synchronization mechanisms.
|
||||
|
||||
Note on tensor names:
|
||||
Efficient communication primitives minimize distributed overhead:
|
||||
|
||||
Official checkpoints often use different `state_dict` names than FastVideo's
|
||||
module layout. We translate tensor names via the DiT arch config mapping
|
||||
(`param_names_mapping` under `fastvideo/configs/models/dits/`). This is similar
|
||||
in spirit to name-translation layers used in systems like vLLM and SGLang.
|
||||
- **Sequence-Parallel AllGather**: Collects sequence chunks
|
||||
- **Tensor-Parallel AllReduce**: Combines partial results
|
||||
- **Distributed Synchronization**: Coordinates execution
|
||||
|
||||
Example HF repo (Wan 2.1 T2V 1.3B Diffusers):
|
||||
## Forward Context Management
|
||||
|
||||
```
|
||||
https://huggingface.co/Wan-AI/Wan2.1-T2V-1.3B-Diffusers/tree/main
|
||||
### ForwardContext
|
||||
|
||||
Defined in `fastvideo/forward_context.py`, `ForwardContext` manages execution-specific state *within* a forward pass, particularly for low-level optimizations. It is accessed via `get_forward_context()`.
|
||||
|
||||
- **Attention Metadata**: Configuration for optimized attention kernels (`attn_metadata`)
|
||||
- **Profiling Data**: Potential hooks for performance metrics collection
|
||||
|
||||
This context-based approach enables:
|
||||
|
||||
- Dynamic optimization based on execution state (e.g., attention backend selection)
|
||||
- Step-specific customizations within model components
|
||||
|
||||
Usage example:
|
||||
|
||||
```python
|
||||
with set_forward_context(current_timestep, attn_metadata, fastvideo_args):
|
||||
# During this forward pass, components can access context
|
||||
# through get_forward_context()
|
||||
output = model(inputs)
|
||||
```
|
||||
|
||||
Example `model_index.json` from that repo:
|
||||
## Executor and Worker System
|
||||
|
||||
```json
|
||||
{
|
||||
"_class_name": "WanPipeline",
|
||||
"_diffusers_version": "0.33.0.dev0",
|
||||
"scheduler": [
|
||||
"diffusers",
|
||||
"UniPCMultistepScheduler"
|
||||
],
|
||||
"text_encoder": [
|
||||
"transformers",
|
||||
"UMT5EncoderModel"
|
||||
],
|
||||
"tokenizer": [
|
||||
"transformers",
|
||||
"T5TokenizerFast"
|
||||
],
|
||||
"transformer": [
|
||||
"diffusers",
|
||||
"WanTransformer3DModel"
|
||||
],
|
||||
"vae": [
|
||||
"diffusers",
|
||||
"AutoencoderKLWan"
|
||||
]
|
||||
}
|
||||
The `fastvideo/worker/` directory contains the distributed execution framework:
|
||||
|
||||
### Executor Abstraction
|
||||
|
||||
FastVideo implements a flexible execution model for distributed processing:
|
||||
|
||||
- **Executor Base Class**: An abstract base class defining the interface for all executors
|
||||
- **MultiProcExecutor**: Primary implementation that spawns and manages worker processes
|
||||
- **GPU Workers**: Handle actual model execution on individual GPUs
|
||||
|
||||
The MultiProcExecutor implementation:
|
||||
|
||||
1. Spawns worker processes for each GPU
|
||||
2. Establishes communication channels via pipes
|
||||
3. Coordinates distributed operations across workers
|
||||
4. Handles graceful startup and shutdown of the process group
|
||||
|
||||
Each GPU worker:
|
||||
|
||||
1. Initializes the distributed environment
|
||||
2. Builds the pipeline for the specified model
|
||||
3. Executes requested operations on its assigned GPU
|
||||
4. Manages local resources and communicates results back to the executor
|
||||
|
||||
This design allows FastVideo to efficiently utilize multiple GPUs while providing a simple, unified interface for model execution.
|
||||
|
||||
## Platforms
|
||||
|
||||
The `fastvideo/platforms/` directory provides hardware platform abstractions that enable FastVideo to run efficiently on different hardware configurations:
|
||||
|
||||
### Platform Abstraction
|
||||
|
||||
FastVideo's platform abstraction layer enables:
|
||||
|
||||
- **Hardware Detection**: Automatic detection of available hardware
|
||||
- **Backend Selection**: Appropriate selection of compute kernels
|
||||
- **Memory Management**: Efficient utilization of hardware-specific memory features
|
||||
|
||||
The primary components include:
|
||||
|
||||
- **Platform Interface**: Defines the common API for all platform implementations
|
||||
- **CUDA Platform**: Optimized implementation for NVIDIA GPUs
|
||||
- **Backend Enum**: Used throughout the codebase for feature selection
|
||||
|
||||
Usage example:
|
||||
|
||||
```python
|
||||
from fastvideo.platforms import current_platform, _Backend
|
||||
|
||||
# Check hardware capabilities
|
||||
if current_platform.supports_backend(_Backend.FLASH_ATTN):
|
||||
# Use FlashAttention implementation
|
||||
else:
|
||||
# Fall back to standard implementation
|
||||
```
|
||||
|
||||
How this maps to FastVideo:
|
||||
The platform system is designed to be extensible for future hardware targets.
|
||||
|
||||
- `WanPipeline` -> `fastvideo/pipelines/basic/wan/wan_pipeline.py`
|
||||
- `WanTransformer3DModel` -> `fastvideo/models/dits/wanvideo.py`
|
||||
- `AutoencoderKLWan` -> `fastvideo/models/vaes/wanvae.py`
|
||||
- `UMT5EncoderModel` -> `fastvideo/models/encoders/t5.py`
|
||||
- `T5TokenizerFast` -> loaded via HF in `fastvideo/models/loader/`
|
||||
- `UniPCMultistepScheduler` -> loaded via Diffusers scheduler utilities
|
||||
- Pipeline defaults -> `fastvideo/configs/pipelines/wan.py`
|
||||
- Sampling defaults -> `fastvideo/configs/sample/wan.py`
|
||||
## Logger
|
||||
|
||||
## Pipeline system
|
||||
See [PR](https://github.com/hao-ai-lab/FastVideo/pull/356)
|
||||
|
||||
- `fastvideo/pipelines/basic/*` contains end-to-end pipelines for each model
|
||||
family.
|
||||
- `fastvideo/pipelines/stages/*` contains reusable, testable stages.
|
||||
- Pipelines subclass `ComposedPipelineBase` and declare required components via
|
||||
`_required_config_modules`.
|
||||
- `ForwardBatch` (in `fastvideo/pipelines/pipeline_batch_info.py`) carries
|
||||
prompts, latents, timesteps, and intermediate state across stages.
|
||||
*TODO*: (help wanted) Add an environment variable that disables process-aware logging.
|
||||
|
||||
## Model components
|
||||
## Contributing to FastVideo
|
||||
|
||||
- DiT models: `fastvideo/models/dits/`
|
||||
- VAEs: `fastvideo/models/vaes/`
|
||||
- Text/image encoders: `fastvideo/models/encoders/`
|
||||
- Schedulers: `fastvideo/models/schedulers/`
|
||||
- Upsamplers: `fastvideo/models/upsamplers/`
|
||||
- Optional audio models: `fastvideo/models/audio/`
|
||||
If you're a new contributor, here are some common areas to explore:
|
||||
|
||||
## Attention and distributed execution
|
||||
1. **Adding a new model**: Implement new model types in the appropriate subdirectory of `fastvideo/models/`
|
||||
2. **Optimizing performance**: Look at attention implementations or memory management
|
||||
3. **Adding a new pipeline**: Create a new pipeline subclass in `fastvideo/pipelines/`
|
||||
4. **Hardware support**: Extend the `platforms` module for new hardware targets
|
||||
|
||||
- Attention backends live in `fastvideo/attention/` and can be selected via
|
||||
`FASTVIDEO_ATTENTION_BACKEND`.
|
||||
- `LocalAttention` is used for cross-attention and most attention layers.
|
||||
- `DistributedAttention` is used for full-sequence self-attention in the DiT.
|
||||
- Tensor-parallel layers live in `fastvideo/layers/`.
|
||||
- Sequence/tensor parallel utilities live in `fastvideo/distributed/`.
|
||||
When adding code, follow these practices:
|
||||
|
||||
## Related docs
|
||||
|
||||
- [Contributing overview](../contributing/overview.md)
|
||||
- [Coding agents workflow](../contributing/coding_agents.md)
|
||||
- [Testing guide](../contributing/testing.md)
|
||||
- Use type hints for better code readability
|
||||
- Add appropriate docstrings
|
||||
- Maintain the separation between model components and execution logic
|
||||
- Follow existing patterns for distributed processing
|
||||
|
||||
@@ -1,369 +0,0 @@
|
||||
# Training Architecture
|
||||
|
||||
!!! warning "Work in Progress"
|
||||
This training architecture (`fastvideo/train/`) is under active development
|
||||
and is replacing the older `fastvideo/training/` module. APIs, config
|
||||
formats, and supported methods may change. See the
|
||||
[Current Status](#current-status) section for what is implemented so far.
|
||||
|
||||
FastVideo's training framework (`fastvideo/train/`) is built around a
|
||||
**pluggable, YAML-driven architecture** that cleanly separates **models**,
|
||||
**training methods**, and **infrastructure** into independent, composable
|
||||
layers. A single YAML config file is all that is needed to train any supported
|
||||
model with any supported algorithm — no code changes required to mix and match.
|
||||
|
||||
---
|
||||
|
||||
## Motivation
|
||||
|
||||
Training video diffusion models involves a tangle of concerns: model loading,
|
||||
noise scheduling, distillation algorithms, distributed strategies,
|
||||
checkpointing, and validation. Existing training scripts tend to hard-wire
|
||||
these together, making it painful to:
|
||||
|
||||
1. **Try a new distillation algorithm** on an existing model (requires forking
|
||||
the training loop).
|
||||
2. **Add a new model** to an existing algorithm (requires re-implementing
|
||||
boilerplate).
|
||||
3. **Switch distributed strategies** (FSDP, TP, SP) without touching algorithm
|
||||
code.
|
||||
4. **Resume, checkpoint, and validate** uniformly across all combinations.
|
||||
|
||||
The training framework solves this by making each axis of variation an
|
||||
independent plugin.
|
||||
|
||||
---
|
||||
|
||||
## Architecture Overview
|
||||
|
||||
```
|
||||
YAML Config
|
||||
|
|
||||
v
|
||||
+------------------+ +---------------------+ +------------------+
|
||||
| Models Layer | | Methods Layer | | Infrastructure |
|
||||
| (per-role) | | (algorithm) | | Layer |
|
||||
| | | | | |
|
||||
| - ModelBase |<----| - TrainingMethod |---->| - Trainer |
|
||||
| - CausalModelBase| | - single_train_step| | - Callbacks |
|
||||
| | | - backward | | - Checkpoint |
|
||||
| Roles: | | - optimizers | | - Tracker (W&B) |
|
||||
| student | | | | - Dataloader |
|
||||
| teacher | | Algorithms: | | |
|
||||
| critic | | DMD2, SelfForcing, | | Distributed: |
|
||||
| | | SFT, DFSFT | | HSDP, TP, SP |
|
||||
+------------------+ +---------------------+ +------------------+
|
||||
```
|
||||
|
||||
### Three Layers
|
||||
|
||||
| Layer | Responsibility | Extension point |
|
||||
|-------|---------------|-----------------|
|
||||
| **Models** (`fastvideo/train/models/`) | Load transformer + scheduler, define `predict_noise`, `predict_x0`, `add_noise`, `backward`. Each training role (student/teacher/critic) is an independent instance. | Subclass `ModelBase` (or `CausalModelBase` for streaming). |
|
||||
| **Methods** (`fastvideo/train/methods/`) | Implement the training algorithm: own role models, define `single_train_step` + `backward`, manage optimizers/schedulers. | Subclass `TrainingMethod`. |
|
||||
| **Infrastructure** (`fastvideo/train/trainer.py`, `utils/`, `callbacks/`) | Training loop, gradient accumulation, distributed setup, checkpointing (DCP), W&B tracking, validation, EMA, grad clipping. | Add callbacks; everything else is shared. |
|
||||
|
||||
---
|
||||
|
||||
## YAML-Driven Configuration
|
||||
|
||||
Everything is configured declaratively. The `_target_` field selects the Python
|
||||
class to instantiate:
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
teacher:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: false
|
||||
critic:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.distribution_matching.dmd2.DMD2Method
|
||||
rollout_mode: simulate
|
||||
dmd_denoising_steps: [1000, 850, 700, 550, 350, 275, 200, 125]
|
||||
generator_update_interval: 5
|
||||
real_score_guidance_scale: 3.5
|
||||
# ...
|
||||
|
||||
training:
|
||||
distributed: { num_gpus: 8, sp_size: 1, tp_size: 1 }
|
||||
data: { data_path: ..., num_latent_t: 20, num_frames: 77 }
|
||||
optimizer: { learning_rate: 2.0e-6, betas: [0.0, 0.999] }
|
||||
loop: { max_train_steps: 4000 }
|
||||
checkpoint: { output_dir: outputs/my_run }
|
||||
|
||||
callbacks:
|
||||
grad_clip: { max_grad_norm: 1.0 }
|
||||
validation: { pipeline_target: ..., every_steps: 100 }
|
||||
```
|
||||
|
||||
To switch from DMD2 to SFT, change the `method._target_` and remove the
|
||||
teacher/critic — no code changes needed.
|
||||
|
||||
---
|
||||
|
||||
## Model Abstraction
|
||||
|
||||
### `ModelBase` — Standard (Bidirectional) Models
|
||||
|
||||
Every role gets its own `ModelBase` instance owning a `transformer` and
|
||||
`noise_scheduler`. The base class defines:
|
||||
|
||||
- **`prepare_batch()`** — Convert raw dataloader output into forward-ready
|
||||
`TrainingBatch`.
|
||||
- **`add_noise()`** — Apply forward-process noise at a given timestep.
|
||||
- **`predict_noise()` / `predict_x0()`** — Run the transformer and return
|
||||
predictions.
|
||||
- **`backward()`** — Backward pass that restores forward context (attention
|
||||
metadata, timesteps).
|
||||
- **`init_preprocessors()`** — Lazy-load VAE, build dataloader (called only on
|
||||
the student).
|
||||
|
||||
### `CausalModelBase` — Streaming / Causal Models
|
||||
|
||||
Extends `ModelBase` with streaming inference primitives for causal video
|
||||
generation:
|
||||
|
||||
```python
|
||||
class CausalModelBase(ModelBase):
|
||||
def clear_caches(self, *, cache_tag: str = "pos") -> None: ...
|
||||
def predict_noise_streaming(
|
||||
self, ..., cache_tag, store_kv, cur_start_frame
|
||||
) -> Tensor | None: ...
|
||||
def predict_x0_streaming(
|
||||
self, ..., cache_tag, store_kv, cur_start_frame
|
||||
) -> Tensor | None: ...
|
||||
```
|
||||
|
||||
KV caches are **internal** to the model instance, keyed by `cache_tag`. The
|
||||
method controls when to store (`store_kv=True`) vs. read-only
|
||||
(`store_kv=False`), enabling block-by-block causal rollout during training.
|
||||
|
||||
---
|
||||
|
||||
## Training Methods
|
||||
|
||||
### DMD2 (Distribution Matching Distillation)
|
||||
|
||||
**Roles:** student (trainable) + teacher (frozen) + critic (trainable)
|
||||
|
||||
The student learns to generate clean video in few steps by matching the
|
||||
teacher's score function, with a critic network providing a learned fake-score
|
||||
baseline.
|
||||
|
||||
- **Rollout modes:**
|
||||
- `simulate` — Student starts from pure noise and iteratively denoises
|
||||
through the full step schedule.
|
||||
- `data_latent` — Student denoises from a single randomly-noised data
|
||||
sample.
|
||||
- **Losses:** Generator loss (DMD gradient) + critic flow-matching loss, with
|
||||
alternating updates (`generator_update_interval`).
|
||||
|
||||
### Self-Forcing (Causal DMD)
|
||||
|
||||
**Roles:** student (causal, trainable) + teacher (frozen) + critic (trainable)
|
||||
|
||||
Extends DMD2 for **streaming/causal video generation**. The key idea: during
|
||||
training, the student processes video in temporal chunks, using its own
|
||||
previously-denoised outputs as context for future chunks — simulating online
|
||||
autoregressive rollout.
|
||||
|
||||
- Video is split into blocks of `chunk_size` latent frames.
|
||||
- Each block is denoised through the student's step schedule; a random
|
||||
early-exit step is sampled per block.
|
||||
- After denoising a block, its output is fed back (with optional
|
||||
`context_noise`) as KV cache context for subsequent blocks via
|
||||
`predict_noise_streaming(store_kv=True)`.
|
||||
- Supports SDE and ODE sampling during rollout.
|
||||
- Selective gradient control: `enable_gradient_in_rollout`,
|
||||
`start_gradient_frame`.
|
||||
|
||||
### Supervised Fine-Tuning (SFT)
|
||||
|
||||
**Roles:** student only
|
||||
|
||||
Standard flow-matching loss between predicted and ground-truth noise/x0.
|
||||
|
||||
### Diffusion-Forcing SFT (DFSFT)
|
||||
|
||||
**Roles:** student only
|
||||
|
||||
SFT with **inhomogeneous (per-chunk) timesteps** — each temporal chunk in a
|
||||
video gets a different noise level. This trains the model to handle mixed-noise
|
||||
inputs, which is a prerequisite for causal/streaming inference where earlier
|
||||
frames are cleaner than later ones.
|
||||
|
||||
---
|
||||
|
||||
## Training Loop
|
||||
|
||||
The `Trainer` runs a standard loop with pluggable method and callbacks:
|
||||
|
||||
```
|
||||
for step in range(start_step, max_steps):
|
||||
for accum_iter in range(grad_accum_steps):
|
||||
batch <- dataloader
|
||||
loss_map, outputs, metrics <- method.single_train_step(batch, step)
|
||||
method.backward(loss_map, outputs)
|
||||
|
||||
callbacks.on_before_optimizer_step() # grad clipping
|
||||
method.optimizers_schedulers_step()
|
||||
method.optimizers_zero_grad()
|
||||
callbacks.on_training_step_end() # logging
|
||||
checkpoint_manager.maybe_save(step)
|
||||
callbacks.on_validation_begin() # periodic inference
|
||||
```
|
||||
|
||||
### Callbacks
|
||||
|
||||
- **GradNormClipCallback** — Per-module gradient norm logging + global
|
||||
clipping.
|
||||
- **ValidationCallback** — Periodic inference sampling with configurable
|
||||
pipeline, sampling steps, and guidance scale.
|
||||
- **EMACallback** — Exponential moving average of student weights.
|
||||
|
||||
### Checkpointing
|
||||
|
||||
- DCP (Distributed Checkpoint) format, compatible with FSDP/HSDP.
|
||||
- Saves: model weights, optimizer states, scheduler states, RNG states (per
|
||||
role).
|
||||
- Full resume support: auto-restores step counter and all RNG states.
|
||||
|
||||
---
|
||||
|
||||
## Getting Started
|
||||
|
||||
```bash
|
||||
# Install
|
||||
uv pip install -e .[dev]
|
||||
|
||||
# Run DMD2 distillation on Wan 2.1
|
||||
torchrun --nproc_per_node=8 -m fastvideo.train.entrypoint.train \
|
||||
--config examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml
|
||||
|
||||
# Run SFT fine-tuning
|
||||
torchrun --nproc_per_node=8 -m fastvideo.train.entrypoint.train \
|
||||
--config examples/train/finetune_wan2.1_t2v_1.3B_vsa_phase3.4_0.9sparsity.yaml
|
||||
```
|
||||
|
||||
Example configs are in `examples/train/`.
|
||||
|
||||
---
|
||||
|
||||
## File Structure
|
||||
|
||||
```
|
||||
fastvideo/train/
|
||||
trainer.py # Training loop
|
||||
models/
|
||||
base.py # ModelBase, CausalModelBase ABCs
|
||||
wan/wan.py # Wan 2.1 T2V model plugin
|
||||
wangame/wangame.py # WanGame 2.1 I2V model plugin
|
||||
wangame/wangame_causal.py # WanGame causal (streaming) plugin
|
||||
methods/
|
||||
base.py # TrainingMethod ABC
|
||||
distribution_matching/
|
||||
dmd2.py # DMD2 distillation
|
||||
self_forcing.py # Self-Forcing (causal DMD)
|
||||
fine_tuning/
|
||||
finetune.py # Supervised fine-tuning
|
||||
dfsft.py # Diffusion-forcing SFT
|
||||
callbacks/
|
||||
grad_clip.py # Gradient clipping + norm logging
|
||||
validation.py # Periodic inference validation
|
||||
ema.py # EMA weight averaging
|
||||
entrypoint/
|
||||
train.py # CLI entrypoint (torchrun)
|
||||
utils/
|
||||
config.py # YAML parser -> RunConfig
|
||||
builder.py # build_from_config: model/method instantiation
|
||||
training_config.py # TrainingConfig dataclass
|
||||
dataloader.py # Dataset/dataloader construction
|
||||
optimizer.py # Optimizer/scheduler construction
|
||||
checkpoint.py # DCP save/resume
|
||||
tracking.py # W&B tracker
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Current Status
|
||||
|
||||
| Component | Status |
|
||||
|-----------|--------|
|
||||
| Core framework (trainer, config, callbacks) | Implemented and tested |
|
||||
| `WanModel` (Wan 2.1 T2V) | Implemented and tested |
|
||||
| `WanGameModel` (WanGame 2.1 I2V) | Implemented and tested |
|
||||
| `WanGameCausalModel` (streaming) | Implemented and tested |
|
||||
| `WanCausalModel` (Wan T2V causal) | In progress |
|
||||
| DMD2 method | Implemented and tested |
|
||||
| Self-Forcing method | Implemented and tested |
|
||||
| SFT method | Implemented and tested |
|
||||
| DFSFT method | Implemented and tested |
|
||||
| DCP checkpointing + resume | Implemented and tested |
|
||||
| EMA callback | Implemented |
|
||||
| Validation callback | Implemented and tested |
|
||||
| Causal DMD inference pipeline | Implemented |
|
||||
|
||||
---
|
||||
|
||||
## Open Questions
|
||||
|
||||
We welcome community feedback on the following topics:
|
||||
|
||||
### Model Plugin API
|
||||
|
||||
The current `ModelBase` interface requires implementing 6 methods. Is this the
|
||||
right granularity?
|
||||
|
||||
- Should `prepare_batch` be split into separate concerns (noise sampling,
|
||||
timestep sampling, attention metadata)?
|
||||
- Should `backward` be lifted out of the model and into the method/trainer?
|
||||
|
||||
### Causal Streaming Interface
|
||||
|
||||
`CausalModelBase` adds `predict_noise_streaming` / `predict_x0_streaming` with
|
||||
cache management. Alternatives considered:
|
||||
|
||||
- **(a) Current:** Cache is internal to the model, keyed by `cache_tag`.
|
||||
Simple but couples cache lifecycle to model.
|
||||
- **(b) External cache:** Method owns the cache dict, passes it into predict
|
||||
calls. More explicit but verbose.
|
||||
- **(c) Context manager:** `with model.streaming_context(tag) as ctx: ...` —
|
||||
cleaner lifecycle but harder to compose.
|
||||
|
||||
### Method Composition
|
||||
|
||||
Currently, methods are monolithic classes. Should we support composing methods
|
||||
(e.g., DFSFT pre-training followed by Self-Forcing distillation) within a
|
||||
single config? Or is sequential training with checkpoint handoff sufficient?
|
||||
|
||||
### New Models and Methods
|
||||
|
||||
What models and training methods should we prioritize next?
|
||||
|
||||
- **Models:** HunyuanVideo, CogVideoX, other Wan variants?
|
||||
- **Methods:** Consistency models, progressive distillation, reward-based
|
||||
fine-tuning?
|
||||
|
||||
### Distributed Strategy
|
||||
|
||||
Currently supports HSDP (hybrid sharded data parallel) + TP + SP. Are there
|
||||
scenarios where the current distributed setup is insufficient? Should we add
|
||||
pipeline parallelism for very large models?
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [Self-Forcing paper](https://arxiv.org/abs/2406.05477) — Chen et al., 2024.
|
||||
- [DMD2 paper](https://arxiv.org/abs/2405.14867) — Yin et al., 2024.
|
||||
- [Diffusion Forcing paper](https://arxiv.org/abs/2407.01392) — Chen et al.,
|
||||
2024.
|
||||
@@ -6,7 +6,7 @@ All documented examples are autogenerated using [generate_examples.py](https://g
|
||||
|
||||
## Examples
|
||||
|
||||
- [Examples Distillation Index](../distillation/examples/examples_distillation_index.md)
|
||||
- [Examples Training Index](../training/examples/examples_training_index.md)
|
||||
- [Examples Inference Index](../inference/examples/examples_inference_index.md)
|
||||
- [Examples Distillation Index](distillation/examples/examples_distillation_index.md)
|
||||
- [Examples Training Index](training/examples/examples_training_index.md)
|
||||
- [Examples Inference Index](inference/examples/examples_inference_index.md)
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
# adapted from vllm: https://github.com/vllm-project/vllm/blob/v0.7.3/docs/source/generate_examples.py
|
||||
|
||||
import itertools
|
||||
import os
|
||||
import re
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
@@ -127,20 +126,8 @@ class Example:
|
||||
Raises:
|
||||
IndexError: If no Markdown files are found in the directory.
|
||||
""" # noqa: E501
|
||||
if self.path.is_file():
|
||||
return self.path
|
||||
|
||||
markdown_files = sorted(self.path.glob("*.md"))
|
||||
if not markdown_files:
|
||||
raise IndexError(f"No Markdown files found in {self.path}")
|
||||
|
||||
readme_files = [
|
||||
f for f in markdown_files if f.name.lower() == "readme.md"
|
||||
]
|
||||
if readme_files:
|
||||
return readme_files[0]
|
||||
|
||||
return markdown_files[0]
|
||||
return self.path if self.path.is_file() else list(
|
||||
self.path.glob("*.md")).pop()
|
||||
|
||||
def determine_other_files(self) -> list[Path]:
|
||||
"""
|
||||
@@ -534,11 +521,11 @@ def generate_examples(generate_main_index: bool = False) -> None:
|
||||
# Add to main index if it exists
|
||||
if generate_main_index and examples_index:
|
||||
main_index_dir = examples_index.path.parent
|
||||
rel_path = os.path.relpath(category_index.path,
|
||||
start=main_index_dir)
|
||||
rel_path = category_index.path.relative_to(
|
||||
main_index_dir.parent)
|
||||
examples_index.documents.insert(
|
||||
0,
|
||||
str(rel_path).replace("\\", "/").replace(".md", ""))
|
||||
str(rel_path).replace(".md", ""))
|
||||
|
||||
# Write the category index file
|
||||
with open(category_index.path, "w+") as f:
|
||||
|
||||
@@ -8,23 +8,11 @@ FastVideo supports the following hardware platforms:
|
||||
|
||||
## Quick Installation
|
||||
|
||||
### Using uv (recommended)
|
||||
|
||||
Use uv as the default environment manager for faster and more stable installs.
|
||||
|
||||
```bash
|
||||
# Create and activate a new uv environment
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
|
||||
uv pip install fastvideo
|
||||
```
|
||||
|
||||
### Using Conda (alternative)
|
||||
### Using pip
|
||||
|
||||
```bash
|
||||
# Create and activate a new conda environment
|
||||
conda create -n fastvideo python=3.12 -y
|
||||
conda create -n fastvideo python=3.12
|
||||
conda activate fastvideo
|
||||
|
||||
pip install fastvideo
|
||||
@@ -35,17 +23,13 @@ pip install fastvideo
|
||||
```bash
|
||||
git clone https://github.com/hao-ai-lab/FastVideo.git
|
||||
cd FastVideo
|
||||
uv pip install -e .
|
||||
|
||||
# optional: install flash-attn
|
||||
uv pip install flash-attn --no-build-isolation -v
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Alternative with Conda environment:
|
||||
Also optionally install flash-attn:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
pip install flash-attn --no-build-isolation -v
|
||||
pip install flash-attn --no-build-isolation
|
||||
```
|
||||
|
||||
## Hardware Requirements
|
||||
|
||||
@@ -12,20 +12,8 @@ Instructions to install FastVideo for NVIDIA CUDA GPUs.
|
||||
## Set up using Python
|
||||
### Create a new Python environment
|
||||
|
||||
#### uv
|
||||
Recommended default: use [uv](https://docs.astral.sh/uv/) for faster and more stable environment setup.
|
||||
|
||||
Please follow the [documentation](https://docs.astral.sh/uv/#getting-started) to install `uv`. After installing `uv`, create a new environment using:
|
||||
|
||||
```console
|
||||
# (Recommended) Create a new uv environment. Use `--seed` to install `pip` and `setuptools`.
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
#### Conda (alternative)
|
||||
You can also create a Python environment using [Conda](https://docs.conda.io/projects/conda/en/stable/user-guide/getting-started.html).
|
||||
|
||||
#### Conda
|
||||
You can create a new python environment using [Conda](https://docs.conda.io/projects/conda/en/stable/user-guide/getting-started.html)
|
||||
##### 1. Install Miniconda (if not already installed)
|
||||
|
||||
```bash
|
||||
@@ -37,35 +25,34 @@ source ~/.bashrc
|
||||
##### 2. Create and activate a Conda environment for FastVideo
|
||||
|
||||
```bash
|
||||
# Create and activate a Conda environment
|
||||
# (Recommended) Create a new conda environment.
|
||||
conda create -n fastvideo python=3.12 -y
|
||||
conda activate fastvideo
|
||||
```
|
||||
|
||||
#### uv
|
||||
|
||||
Or you can create a new Python environment using [uv](https://docs.astral.sh/uv/), a very fast Python environment manager. Please follow the [documentation](https://docs.astral.sh/uv/#getting-started) to install `uv`. After installing `uv`, you can create a new Python environment using the following command:
|
||||
|
||||
```console
|
||||
# (Recommended) Create a new uv environment. Use `--seed` to install `pip` and `setuptools` in the environment.
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
### Installation
|
||||
|
||||
#### With uv (recommended)
|
||||
|
||||
```bash
|
||||
uv pip install fastvideo
|
||||
```
|
||||
|
||||
Also optionally install FlashAttention:
|
||||
|
||||
```bash
|
||||
uv pip install flash-attn --no-build-isolation -v
|
||||
```
|
||||
|
||||
#### With Conda environment (alternative)
|
||||
|
||||
```bash
|
||||
pip install fastvideo
|
||||
|
||||
# or if you are using uv
|
||||
uv pip install fastvideo
|
||||
```
|
||||
|
||||
Also optionally install FlashAttention:
|
||||
Also optionally install flash-attn:
|
||||
|
||||
```bash
|
||||
pip install flash-attn --no-build-isolation -v
|
||||
pip install flash-attn --no-build-isolation
|
||||
```
|
||||
|
||||
### Installation from Source
|
||||
@@ -80,14 +67,11 @@ git clone https://github.com/hao-ai-lab/FastVideo.git && cd FastVideo
|
||||
|
||||
Basic installation:
|
||||
|
||||
```bash
|
||||
uv pip install -e .
|
||||
```
|
||||
|
||||
Alternative with Conda environment:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
|
||||
# or if you are using uv
|
||||
uv pip install -e .
|
||||
```
|
||||
|
||||
### Optional Dependencies
|
||||
@@ -95,13 +79,7 @@ pip install -e .
|
||||
#### Flash Attention
|
||||
|
||||
```bash
|
||||
uv pip install flash-attn --no-build-isolation -v
|
||||
```
|
||||
|
||||
Alternative with Conda environment:
|
||||
|
||||
```bash
|
||||
pip install flash-attn --no-build-isolation -v
|
||||
pip install flash-attn --no-build-isolation
|
||||
```
|
||||
|
||||
## Set up using Docker
|
||||
|
||||
@@ -11,20 +11,9 @@ Instructions to install FastVideo for Apple Silicon.
|
||||
|
||||
### Create a new Python environment
|
||||
|
||||
#### uv
|
||||
Recommended default: use [uv](https://docs.astral.sh/uv/) for faster and more stable environment setup.
|
||||
#### Conda
|
||||
|
||||
Please follow the [documentation](https://docs.astral.sh/uv/#getting-started) to install `uv`. After installing `uv`, create a new environment using:
|
||||
|
||||
```console
|
||||
# (Recommended) Create a new uv environment. Use `--seed` to install `pip` and `setuptools`.
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
#### Conda (alternative)
|
||||
|
||||
You can also create a Python environment using [Conda](https://docs.conda.io/projects/conda/en/stable/user-guide/getting-started.html).
|
||||
You can create a new python environment using [Conda](https://docs.conda.io/projects/conda/en/stable/user-guide/getting-started.html)
|
||||
|
||||
##### 1. Install Miniconda (if not already installed)
|
||||
|
||||
@@ -37,10 +26,21 @@ source ~/.zshrc
|
||||
##### 2. Create and activate a Conda environment for FastVideo
|
||||
|
||||
```bash
|
||||
# (Recommended) Create a new conda environment.
|
||||
conda create -n fastvideo python=3.12.4 -y
|
||||
conda activate fastvideo
|
||||
```
|
||||
|
||||
#### uv
|
||||
|
||||
Or you can create a new Python environment using [uv](https://docs.astral.sh/uv/), a very fast Python environment manager. Please follow the [documentation](https://docs.astral.sh/uv/#getting-started) to install `uv`. After installing `uv`, you can create a new Python environment using the following command:
|
||||
|
||||
```console
|
||||
# (Recommended) Create a new uv environment. Use `--seed` to install `pip` and `setuptools` in the environment.
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
### Dependencies
|
||||
|
||||
```
|
||||
@@ -49,16 +49,11 @@ brew install ffmpeg
|
||||
|
||||
### Installation
|
||||
|
||||
#### With uv (recommended)
|
||||
|
||||
```bash
|
||||
uv pip install fastvideo
|
||||
```
|
||||
|
||||
#### With Conda environment (alternative)
|
||||
|
||||
```bash
|
||||
pip install fastvideo
|
||||
|
||||
# or if you are using uv
|
||||
uv pip install fastvideo
|
||||
```
|
||||
|
||||
### Installation from Source
|
||||
@@ -73,14 +68,11 @@ git clone https://github.com/hao-ai-lab/FastVideo.git && cd FastVideo
|
||||
|
||||
Basic installation:
|
||||
|
||||
```bash
|
||||
uv pip install -e .
|
||||
```
|
||||
|
||||
Alternative with Conda environment:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
|
||||
# or if you are using uv
|
||||
uv pip install -e .
|
||||
```
|
||||
|
||||
## Development Environment Setup
|
||||
|
||||
@@ -7,18 +7,18 @@ Get up and running with FastVideo in minutes!
|
||||
First, install FastVideo:
|
||||
|
||||
```bash
|
||||
# If you previously used Conda, use uv instead for a faster, more stable setup
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
# Create and activate a new conda environment
|
||||
conda create -n fastvideo python=3.12
|
||||
conda activate fastvideo
|
||||
|
||||
# Install FastVideo
|
||||
uv pip install fastvideo
|
||||
pip install fastvideo
|
||||
```
|
||||
|
||||
Also optionally install flash-attn:
|
||||
|
||||
```bash
|
||||
uv pip install flash-attn --no-build-isolation -v
|
||||
pip install flash-attn --no-build-isolation
|
||||
```
|
||||
|
||||
## Basic Usage
|
||||
@@ -41,6 +41,7 @@ def main():
|
||||
# Generate the video
|
||||
video = generator.generate_video(
|
||||
prompt,
|
||||
return_frames=True, # Also return frames from this call (defaults to False)
|
||||
output_path="my_videos/", # Controls where videos are saved
|
||||
save_video=True
|
||||
)
|
||||
@@ -78,6 +79,5 @@ if __name__ == '__main__':
|
||||
|
||||
- [Installation Guide](installation.md) - Detailed installation instructions
|
||||
- [Configuration](../inference/configuration.md) - Learn about configuration options
|
||||
- [Examples](../inference/examples/examples_inference_index.md) - Explore more
|
||||
examples
|
||||
- [Examples](../inference/examples/) - Explore more examples
|
||||
- [Optimizations](../inference/optimizations.md) - Performance optimization tips
|
||||
|
||||
@@ -26,7 +26,8 @@ FastVideo is an inference and post-training framework for diffusion models. It f
|
||||
FastVideo has the following features:
|
||||
|
||||
- State-of-the-art performance optimizations for inference
|
||||
- [Sliding Tile Attention](attention/sta/index.md)
|
||||
- [Sliding Tile Attention](https://arxiv.org/pdf/2502.04507)
|
||||
- [TeaCache](https://arxiv.org/pdf/2411.19108)
|
||||
- [Sage Attention](https://arxiv.org/abs/2410.02367)
|
||||
- E2E post-training support
|
||||
- Data preprocessing pipeline for video data
|
||||
@@ -44,7 +45,7 @@ Use the navigation menu on the left to explore different sections:
|
||||
- **Inference**: Learn how to use FastVideo for video generation
|
||||
- **Training**: Data preprocessing and fine-tuning workflows
|
||||
- **Distillation**: Post-training optimization techniques
|
||||
- **Sliding Tile Attention**: Legacy workflow docs and kernel notes
|
||||
- **Sliding Tile Attention**: Advanced attention mechanisms
|
||||
- **Video Sparse Attention**: Efficient attention for video models
|
||||
- **Design**: Framework architecture and design principles
|
||||
- **Developer Guide**: Contributing and development setup
|
||||
|
||||
@@ -101,17 +101,14 @@ Replace standard attention with FastVideo's optimized attention:
|
||||
```python
|
||||
# Local attention patterns
|
||||
from fastvideo.attention import LocalAttention
|
||||
from fastvideo.platforms.interface import AttentionBackendEnum
|
||||
from fastvideo.attention.backends.abstract import _Backend
|
||||
self.attn = LocalAttention(
|
||||
num_heads=num_heads,
|
||||
head_size=head_dim,
|
||||
dropout_rate=0.0,
|
||||
softmax_scale=None,
|
||||
causal=False,
|
||||
supported_attention_backends=(
|
||||
AttentionBackendEnum.FLASH_ATTN,
|
||||
AttentionBackendEnum.TORCH_SDPA,
|
||||
)
|
||||
supported_attention_backends=(_Backend.FLASH_ATTN, _Backend.TORCH_SDPA)
|
||||
)
|
||||
|
||||
# Distributed attention for long sequences
|
||||
@@ -122,60 +119,29 @@ self.attn = DistributedAttention(
|
||||
dropout_rate=0.0,
|
||||
softmax_scale=None,
|
||||
causal=False,
|
||||
supported_attention_backends=(
|
||||
AttentionBackendEnum.VIDEO_SPARSE_ATTN,
|
||||
AttentionBackendEnum.FLASH_ATTN,
|
||||
AttentionBackendEnum.TORCH_SDPA,
|
||||
)
|
||||
supported_attention_backends=(_Backend.SLIDING_TILE_ATTN, _Backend.FLASH_ATTN, _Backend.TORCH_SDPA)
|
||||
)
|
||||
```
|
||||
|
||||
#### Define supported backend selection
|
||||
|
||||
```python
|
||||
_supported_attention_backends = (
|
||||
AttentionBackendEnum.FLASH_ATTN,
|
||||
AttentionBackendEnum.TORCH_SDPA,
|
||||
)
|
||||
_supported_attention_backends = (_Backend.FLASH_ATTN, _Backend.TORCH_SDPA)
|
||||
```
|
||||
|
||||
### Registering Models
|
||||
|
||||
Register implemented modules for auto‑discovery by adding `EntryClass` in each
|
||||
model module (the registry scans for it):
|
||||
Register implemented modules in the model registry:
|
||||
|
||||
```python
|
||||
# In fastvideo/models/dits/your_module.py
|
||||
class YourTransformerModel(...):
|
||||
...
|
||||
# In fastvideo/models/registry.py
|
||||
_TEXT_TO_VIDEO_DIT_MODELS = {
|
||||
"YourTransformerModel": ("dits", "yourmodule", "YourTransformerClass"),
|
||||
}
|
||||
|
||||
# Entry point for model registry
|
||||
EntryClass = YourTransformerModel
|
||||
```
|
||||
|
||||
```python
|
||||
# In fastvideo/models/vaes/your_vae.py
|
||||
class YourVAEModel(...):
|
||||
...
|
||||
|
||||
# Entry point for model registry
|
||||
EntryClass = YourVAEModel
|
||||
```
|
||||
|
||||
Register pipeline config + sampling defaults in the unified registry:
|
||||
|
||||
```python
|
||||
# In fastvideo/registry.py
|
||||
register_configs(
|
||||
sampling_param_cls=YourSamplingParam,
|
||||
pipeline_config_cls=YourPipelineConfig,
|
||||
hf_model_paths=[
|
||||
"org/your-model-id",
|
||||
],
|
||||
model_detectors=[
|
||||
lambda path: "your-model" in path.lower(),
|
||||
],
|
||||
)
|
||||
_VAE_MODELS = {
|
||||
"YourVAEModel": ("vaes", "yourvae", "YourVAEClass"),
|
||||
}
|
||||
```
|
||||
|
||||
## Step 2: Directory Structure
|
||||
|
||||
@@ -1,460 +0,0 @@
|
||||
# Inference Architecture
|
||||
|
||||
This section documents the FastVideo inference pipeline: how models are
|
||||
discovered, configs resolved, components loaded, and stages composed to
|
||||
generate video. Training-specific code paths (FSDP, gradient checkpointing)
|
||||
are out of scope.
|
||||
|
||||
## Registries
|
||||
|
||||
FastVideo uses three registries that work together to resolve a
|
||||
user-provided `model_path` into a runnable pipeline.
|
||||
|
||||
### Model Registry (`fastvideo/models/registry.py`)
|
||||
|
||||
Maps HuggingFace architecture class names (e.g. `"WanTransformer3DModel"`)
|
||||
to FastVideo model classes. Two discovery mechanisms:
|
||||
|
||||
1. **Hardcoded dicts** — `_TEXT_TO_VIDEO_DIT_MODELS`, `_VAE_MODELS`,
|
||||
`_SCHEDULERS`, `_TEXT_ENCODER_MODELS`, `_IMAGE_ENCODER_MODELS`,
|
||||
`_UPSAMPLERS`, `_AUDIO_MODELS`. Each entry is
|
||||
`{hf_class_name: (component_name, module_name, class_name)}`.
|
||||
2. **AST-based discovery** — `_discover_and_register_models()` walks
|
||||
`fastvideo/models/` and parses each `.py` file's AST looking for an
|
||||
`EntryClass` variable assignment. Discovered models take priority over
|
||||
hardcoded entries. For example,
|
||||
`fastvideo/models/dits/wanvideo.py` exports
|
||||
`EntryClass = WanTransformer3DModel`.
|
||||
|
||||
Both feed into a unified `_FAST_VIDEO_MODELS` dict, which populates the
|
||||
singleton `ModelRegistry` — an instance of `_ModelRegistry`. Components are
|
||||
wrapped in `_LazyRegisteredModel` for deferred import.
|
||||
|
||||
**Key API:** `ModelRegistry.resolve_model_cls(architectures)` iterates
|
||||
candidate architecture strings and returns the first matching
|
||||
`(model_cls, arch)` tuple. Called by `TransformerLoader`, `VAELoader`, etc.
|
||||
|
||||
### Config Registry (`fastvideo/registry.py`)
|
||||
|
||||
Maps model paths and names to `(PipelineConfig, SamplingParam)` class
|
||||
pairs. Registration happens at module load via `_register_configs()`, which
|
||||
calls `register_configs()` for each model family:
|
||||
|
||||
```python
|
||||
register_configs(
|
||||
sampling_param_cls=WanT2V_1_3B_SamplingParam,
|
||||
pipeline_config_cls=WanT2V480PConfig,
|
||||
hf_model_paths=["Wan-AI/Wan2.1-T2V-1.3B-Diffusers"],
|
||||
model_detectors=[lambda path: "wanpipeline" in path.lower()],
|
||||
)
|
||||
```
|
||||
|
||||
Each call populates three data structures:
|
||||
- `_CONFIG_REGISTRY: dict[str, ConfigInfo]` — auto-incrementing ID to
|
||||
`ConfigInfo(sampling_param_cls, pipeline_config_cls)`.
|
||||
- `_MODEL_HF_PATH_TO_NAME: dict[str, str]` — HF path to registry ID.
|
||||
- `_MODEL_NAME_DETECTORS: list[tuple[str, Callable]]` — lambda detectors.
|
||||
|
||||
**Resolution priority** (`_get_config_info()`):
|
||||
1. Exact HF path match in `_MODEL_HF_PATH_TO_NAME`.
|
||||
2. Partial match on short model name (last path segment, case-insensitive).
|
||||
3. Detector-based match — runs each detector against the lowercased path
|
||||
and the `_class_name` from `model_index.json`.
|
||||
4. `RuntimeError` if no match.
|
||||
|
||||
**Top-level resolver:** `get_model_info(model_path, pipeline_type,
|
||||
workload_type)` combines config resolution with pipeline resolution to
|
||||
return a `ModelInfo(pipeline_cls, sampling_param_cls, pipeline_config_cls)`.
|
||||
|
||||
### Pipeline Registry (`fastvideo/pipelines/pipeline_registry.py`)
|
||||
|
||||
Discovers pipeline classes by scanning Python packages under
|
||||
`fastvideo/pipelines/{basic,preprocess,training}/`.
|
||||
|
||||
`import_pipeline_classes()` iterates architecture subdirectories
|
||||
(e.g. `wan/`, `hunyuan/`), imports each module, and collects those
|
||||
exporting an `EntryClass` attribute. Supports single class or list.
|
||||
Returns `{pipeline_type_str: {pipeline_class_name: pipeline_cls}}`.
|
||||
|
||||
`_PipelineRegistry.resolve_pipeline_cls(pipeline_name, pipeline_type,
|
||||
workload_type)` looks up the pipeline class by the `_class_name` field
|
||||
from `model_index.json`.
|
||||
|
||||
## Config Mechanism
|
||||
|
||||
### Config Hierarchy
|
||||
|
||||
```
|
||||
PipelineConfig (fastvideo/configs/pipelines/base.py)
|
||||
├── WanT2V480PConfig (fastvideo/configs/pipelines/wan.py)
|
||||
│ ├── WanT2V720PConfig
|
||||
│ └── WanI2V480PConfig
|
||||
├── HunyuanConfig (fastvideo/configs/pipelines/hunyuan.py)
|
||||
├── LTX2T2VConfig (fastvideo/configs/pipelines/ltx2.py)
|
||||
├── CosmosConfig (fastvideo/configs/pipelines/cosmos.py)
|
||||
└── ... (15+ model families)
|
||||
```
|
||||
|
||||
`PipelineConfig` holds:
|
||||
- Video generation params: `embedded_cfg_scale`, `flow_shift`,
|
||||
`disable_autocast`, `is_causal`.
|
||||
- Nested model configs: `dit_config: DiTConfig`, `vae_config: VAEConfig`,
|
||||
`text_encoder_configs: tuple[EncoderConfig, ...]`.
|
||||
- Precision settings: `dit_precision`, `vae_precision`,
|
||||
`text_encoder_precisions`.
|
||||
|
||||
Model-specific subclasses override defaults. For example,
|
||||
`WanT2V480PConfig` sets `flow_shift=3.0` and uses `WanVideoConfig` as
|
||||
its DiT config.
|
||||
|
||||
### ModelConfig / ArchConfig (`fastvideo/configs/models/base.py`)
|
||||
|
||||
`ModelConfig` wraps an `ArchConfig` using `__getattr__` proxy — attribute
|
||||
access falls through to `arch_config` transparently. `ArchConfig` holds
|
||||
architecture fields from `config.json` (hidden_size, num_attention_heads,
|
||||
etc.) and is immutable after initialization. `update_model_arch()` writes
|
||||
to `ArchConfig`; `update_model_config()` writes to `ModelConfig` fields.
|
||||
|
||||
Concrete hierarchy: `DiTConfig` → `DiTArchConfig`, `VAEConfig` →
|
||||
`VAEArchConfig`, `EncoderConfig` → `EncoderArchConfig`.
|
||||
|
||||
### Config Construction
|
||||
|
||||
- `PipelineConfig.from_pretrained(model_path)` — resolves config class
|
||||
via `get_pipeline_config_cls_from_name()`, instantiates with defaults.
|
||||
- `PipelineConfig.from_kwargs(kwargs)` — resolves class, optionally loads
|
||||
JSON via `load_from_json()`, then applies CLI overrides via
|
||||
`update_config_from_dict()`.
|
||||
- `dump_to_json()` / `load_from_json()` — JSON persistence. Callable
|
||||
fields and `arch_config` are excluded from dumps.
|
||||
|
||||
### SamplingParam (`fastvideo/configs/sample/`)
|
||||
|
||||
Generation parameters separate from pipeline config. Each model family
|
||||
provides defaults:
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
class WanT2V_1_3B_SamplingParam(SamplingParam):
|
||||
height: int = 480
|
||||
width: int = 832
|
||||
num_frames: int = 81
|
||||
guidance_scale: float = 3.0
|
||||
num_inference_steps: int = 50
|
||||
```
|
||||
|
||||
## Component Loading
|
||||
|
||||
### ComponentLoader (`fastvideo/models/loader/component_loader.py`)
|
||||
|
||||
Abstract base with a `load(model_path, fastvideo_args)` method.
|
||||
`ComponentLoader.for_module_type(module_type, library)` is a factory
|
||||
that dispatches to specialized loaders via a `module_loaders` dict:
|
||||
|
||||
| Module type | Loader class | Library |
|
||||
|---|---|---|
|
||||
| `scheduler` | `SchedulerLoader` | diffusers |
|
||||
| `transformer`, `transformer_2`, `transformer_3` | `TransformerLoader` | diffusers |
|
||||
| `vae` | `VAELoader` | diffusers |
|
||||
| `text_encoder`, `text_encoder_2`, `text_encoder_3` | `TextEncoderLoader` | transformers |
|
||||
| `tokenizer`, `tokenizer_2`, `tokenizer_3` | `TokenizerLoader` | transformers |
|
||||
| `image_encoder` | `ImageEncoderLoader` | transformers |
|
||||
| `image_processor`, `feature_extractor` | `ImageProcessorLoader` | transformers |
|
||||
| `audio_vae`, `audio_decoder` | `AudioDecoderLoader` | diffusers |
|
||||
| `vocoder` | `VocoderLoader` | diffusers |
|
||||
| `upsampler`, `upsampler_2` | `UpsamplerLoader` | diffusers |
|
||||
|
||||
`TransformerLoader` reads `config.json` from the component directory,
|
||||
resolves the class via `ModelRegistry.resolve_model_cls()`, instantiates
|
||||
the model, and loads safetensors weights. CPU offload and layerwise
|
||||
offload are applied based on `FastVideoArgs`.
|
||||
|
||||
Unknown module types fall back to `GenericComponentLoader`.
|
||||
|
||||
### model_index.json
|
||||
|
||||
Diffusers-format JSON at the model root. Keys are module names; values
|
||||
are `[library, class_name]` tuples:
|
||||
|
||||
```json
|
||||
{
|
||||
"_class_name": "WanPipeline",
|
||||
"_diffusers_version": "0.24.0",
|
||||
"transformer": ["diffusers", "WanTransformer3DModel"],
|
||||
"vae": ["diffusers", "AutoencoderKLWan"],
|
||||
"text_encoder": ["transformers", "UMT5EncoderModel"],
|
||||
"tokenizer": ["transformers", "AutoTokenizer"],
|
||||
"scheduler": ["diffusers", "FlowMatchEulerDiscreteScheduler"]
|
||||
}
|
||||
```
|
||||
|
||||
`ComposedPipelineBase.load_modules()` reads this file via
|
||||
`_load_config()`, strips metadata keys (`_class_name`,
|
||||
`_diffusers_version`, `_name_or_path`), detects MoE pipelines via
|
||||
`boundary_ratio`, then loads each module listed in the pipeline's
|
||||
`required_config_modules`. Modules not in `required_config_modules` are
|
||||
skipped. `_extra_config_module_map` allows aliasing (e.g. mapping
|
||||
`"transformer_2"` to an alternate directory name).
|
||||
|
||||
`PipelineComponentLoader.load_module()` orchestrates per-component
|
||||
loading by calling `ComponentLoader.for_module_type()` then `.load()`.
|
||||
|
||||
## Stage Design
|
||||
|
||||
### PipelineStage (`fastvideo/pipelines/stages/base.py`)
|
||||
|
||||
Abstract base class using the Template Method pattern:
|
||||
|
||||
- `__call__(batch, fastvideo_args)` — orchestrates verification, timing,
|
||||
and error handling. Not overridden by subclasses.
|
||||
- `forward(batch, fastvideo_args) -> ForwardBatch` — abstract, contains
|
||||
the stage logic.
|
||||
- `verify_input()` / `verify_output()` — optional hooks returning
|
||||
`VerificationResult`. Default: no checks.
|
||||
|
||||
When `fastvideo_args.enable_stage_verification` is `True`, `__call__`
|
||||
runs input verification before `forward()` and output verification after.
|
||||
When `envs.FASTVIDEO_STAGE_LOGGING` is set, execution time is measured
|
||||
with `torch.cuda.synchronize()` and logged.
|
||||
|
||||
### ForwardBatch (`fastvideo/pipelines/pipeline_batch_info.py`)
|
||||
|
||||
Dataclass carrying all pipeline state between stages. Key field groups:
|
||||
|
||||
- **Inputs**: `prompt`, `negative_prompt`, `image_path`, `pil_image`,
|
||||
`video_path`.
|
||||
- **Embeddings**: `prompt_embeds: list[Tensor]`,
|
||||
`negative_prompt_embeds`, `prompt_attention_mask`, `image_embeds`.
|
||||
- **Latents**: `latents`, `image_latent`, `noise_pred`,
|
||||
`lq_latents`.
|
||||
- **Dimensions**: `height`, `width`, `num_frames`, `height_latents`,
|
||||
`width_latents`.
|
||||
- **Scheduler**: `timesteps`, `num_inference_steps`, `guidance_scale`,
|
||||
`sigmas`.
|
||||
- **Task-specific**: `mouse_cond`/`keyboard_cond` (MatrixGame), `pose`
|
||||
(HYWorld), `camera_states` (GameCraft), `c2ws_plucker_emb`
|
||||
(LingBotWorld).
|
||||
- **Output**: `output: Tensor | None`.
|
||||
- **Logging**: `logging_info: PipelineLoggingInfo`.
|
||||
|
||||
`__post_init__` enables CFG when `guidance_scale > 1.0` or LTX2 text
|
||||
CFG scales differ from 1.0.
|
||||
|
||||
### Stage Catalog
|
||||
|
||||
Standard stages (typical execution order):
|
||||
|
||||
| Stage | File | Purpose |
|
||||
|---|---|---|
|
||||
| `InputValidationStage` | `stages/input_validation.py` | Validates input dimensions and types |
|
||||
| `TextEncodingStage` | `stages/text_encoding.py` | Encodes prompts via text encoders |
|
||||
| `ImageEncodingStage` | `stages/image_encoding.py` | Encodes input images (I2V pipelines) |
|
||||
| `ConditioningStage` | `stages/conditioning.py` | Prepares conditioning embeddings |
|
||||
| `TimestepPreparationStage` | `stages/timestep_preparation.py` | Sets up scheduler timesteps |
|
||||
| `LatentPreparationStage` | `stages/latent_preparation.py` | Initializes noise latents |
|
||||
| `DenoisingStage` | `stages/denoising.py` | Main diffusion denoising loop |
|
||||
| `DecodingStage` | `stages/decoding.py` | Decodes latents to video via VAE |
|
||||
|
||||
Specialized variants: `CausalDenoisingStage`, `LTX2DenoisingStage`,
|
||||
`LongCatDenoisingStage`, `GameCraftDenoisingStage`,
|
||||
`HYWorldDenoisingStage`, `MatrixGameDenoisingStage`,
|
||||
`SRDenoisingStage`, `LTX2AudioDecodingStage`, `SD35ConditioningStage`,
|
||||
`LTX2TextEncodingStage`, `LTX2LatentPreparationStage`.
|
||||
|
||||
### Verification System (`fastvideo/pipelines/stages/validators.py`)
|
||||
|
||||
`StageValidators` (aliased as `V`) provides static validators:
|
||||
`not_none`, `positive_int`, `is_tensor`, `tensor_with_dims`,
|
||||
`positive_int_divisible(divisor)`, etc.
|
||||
|
||||
`VerificationResult` collects check results:
|
||||
|
||||
```python
|
||||
result = VerificationResult()
|
||||
result.add_check("height", batch.height, V.positive_int_divisible(8))
|
||||
result.add_check("width", batch.width, V.positive_int_divisible(8))
|
||||
```
|
||||
|
||||
`is_valid()` returns whether all checks passed. `get_failure_summary()`
|
||||
provides detailed error messages. Failed verification raises
|
||||
`StageVerificationError`.
|
||||
|
||||
## Pipeline Architecture
|
||||
|
||||
### ComposedPipelineBase (`fastvideo/pipelines/composed_pipeline_base.py`)
|
||||
|
||||
Abstract base for all inference pipelines. Lifecycle:
|
||||
|
||||
1. **`__init__(model_path, fastvideo_args)`** — initializes distributed
|
||||
environment via `maybe_init_distributed_environment_and_model_parallel
|
||||
(tp_size, sp_size)`, then calls `load_modules()` to populate
|
||||
`self.modules`.
|
||||
2. **`post_init()`** — calls `initialize_pipeline()` (model-specific
|
||||
setup), `create_pipeline_stages()` (abstract — subclasses wire stages),
|
||||
optionally applies `torch.compile` to transformers, and calls
|
||||
`warmup_sequence_parallel_communication()`.
|
||||
3. **`forward(batch, fastvideo_args)`** — iterates `self.stages` calling
|
||||
each stage in order. Decorated with `@torch.no_grad()`.
|
||||
|
||||
Key class attributes:
|
||||
- `_required_config_modules: list[str]` — module names to load from
|
||||
`model_index.json`.
|
||||
- `_extra_config_module_map: dict[str, str]` — aliases for module dirs.
|
||||
- `is_video_pipeline: bool` — whether this produces video output.
|
||||
|
||||
Key methods:
|
||||
- `add_stage(name, stage)` — appends to `_stages` list and
|
||||
`_stage_name_mapping` dict, also sets attribute on `self`.
|
||||
- `get_module(name, default)` — retrieves a loaded module.
|
||||
- `from_pretrained(model_path, **kwargs)` — class method constructing
|
||||
`FastVideoArgs` and calling `cls(...)` then `post_init()`.
|
||||
|
||||
### LoRAPipeline (`fastvideo/pipelines/lora_pipeline.py`)
|
||||
|
||||
Extends `ComposedPipelineBase` with LoRA adapter support. Sits in the
|
||||
MRO between the concrete pipeline and `ComposedPipelineBase`:
|
||||
|
||||
```python
|
||||
class WanPipeline(LoRAPipeline, ComposedPipelineBase):
|
||||
...
|
||||
```
|
||||
|
||||
Key functionality:
|
||||
- `convert_to_lora_layers()` — scans transformer blocks, replaces target
|
||||
linear layers (default: q/k/v/o projections) with LoRA equivalents via
|
||||
`get_lora_layer()`.
|
||||
- `set_lora_adapter(path)` — loads safetensors containing `lora_A`,
|
||||
`lora_B`, `lora_alpha` and maps weights to internal layers.
|
||||
- `merge_lora_weights()` / `unmerge_lora_weights()` — activates or
|
||||
deactivates LoRA in the forward pass.
|
||||
- `LoRAModelLayers` — groups LoRA layers by transformer block for
|
||||
efficient layerwise offload.
|
||||
|
||||
### Distributed Inference
|
||||
|
||||
`maybe_init_distributed_environment_and_model_parallel(tp_size, sp_size)`
|
||||
in `fastvideo/distributed/` initializes `torch.distributed` and creates
|
||||
tensor-parallel (TP) and sequence-parallel (SP) process groups.
|
||||
|
||||
Key APIs: `get_tp_rank()`, `get_tp_world_size()`, `get_sp_rank()`,
|
||||
`get_sp_world_size()`, `get_world_rank()`, `get_world_size()`.
|
||||
|
||||
`warmup_sequence_parallel_communication()` pre-warms NCCL communicators
|
||||
to avoid slow first forward passes.
|
||||
|
||||
Usage: `torchrun --nproc-per-node=N -m fastvideo.entrypoints.cli.main
|
||||
generate --model-path ... --tp-size N --sp-size M`.
|
||||
|
||||
### torch.compile Integration
|
||||
|
||||
When `fastvideo_args.enable_torch_compile` is `True`,
|
||||
`_maybe_compile_pipeline_module()` checks for a `_compile_conditions`
|
||||
attribute on the module. If present, only matching submodules are
|
||||
compiled. Otherwise, the entire module is compiled. FSDP-wrapped
|
||||
modules are skipped.
|
||||
|
||||
### Entry Points
|
||||
|
||||
**Python API** (`fastvideo/entrypoints/video_generator.py`):
|
||||
|
||||
```python
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_path="Wan-AI/Wan2.1-T2V-14B-Diffusers",
|
||||
num_gpus=1, tp_size=1, sp_size=1,
|
||||
)
|
||||
result = generator.generate_video(
|
||||
prompt="A cat dancing",
|
||||
height=720, width=1280, num_frames=81,
|
||||
)
|
||||
```
|
||||
|
||||
**CLI** (`fastvideo/entrypoints/cli/`):
|
||||
|
||||
```bash
|
||||
fastvideo generate \
|
||||
--model-path "Wan-AI/Wan2.1-T2V-14B-Diffusers" \
|
||||
--prompt "A cat dancing" \
|
||||
--num-gpus 1
|
||||
```
|
||||
|
||||
**FastVideoArgs** (`fastvideo/fastvideo_args.py`): Central args dataclass.
|
||||
Key fields: `model_path`, `mode` (`ExecutionMode`), `workload_type`
|
||||
(`WorkloadType`), `pipeline_config` (`PipelineConfig`), `num_gpus`,
|
||||
`tp_size`, `sp_size`, `lora_path`, `dit_cpu_offload`,
|
||||
`dit_layerwise_offload`, `enable_torch_compile`,
|
||||
`enable_stage_verification`.
|
||||
|
||||
Constructed via `FastVideoArgs.from_kwargs(**kwargs)` which resolves the
|
||||
`PipelineConfig` from the registry, applies JSON config if provided, and
|
||||
merges CLI overrides.
|
||||
|
||||
## End-to-End Inference Flow
|
||||
|
||||
```
|
||||
User: VideoGenerator.from_pretrained(model_path, **kwargs)
|
||||
│
|
||||
├─ FastVideoArgs.from_kwargs() → PipelineConfig resolved via registry
|
||||
├─ get_model_info() → ModelInfo(pipeline_cls, sampling_param_cls, ...)
|
||||
│ ├─ model_index.json read → _class_name extracted
|
||||
│ ├─ pipeline_registry resolves pipeline_cls from _class_name
|
||||
│ └─ config_registry resolves config classes from model_path
|
||||
│
|
||||
├─ pipeline_cls.__init__(model_path, fastvideo_args)
|
||||
│ ├─ maybe_init_distributed(tp_size, sp_size)
|
||||
│ └─ load_modules() → reads model_index.json, loads each component
|
||||
│ ├─ ComponentLoader.for_module_type() → specialized loader
|
||||
│ └─ loader.load() → model class resolved, weights loaded
|
||||
│
|
||||
└─ pipeline.post_init()
|
||||
├─ initialize_pipeline() → model-specific setup
|
||||
├─ create_pipeline_stages() → stages wired with modules
|
||||
├─ torch.compile (if enabled)
|
||||
└─ warmup_sequence_parallel_communication()
|
||||
|
||||
User: generator.generate_video(prompt, ...)
|
||||
│
|
||||
├─ ForwardBatch constructed from SamplingParam + user args
|
||||
└─ pipeline.forward(batch, fastvideo_args)
|
||||
├─ InputValidationStage → validates dims
|
||||
├─ TextEncodingStage → prompt → embeddings
|
||||
├─ ConditioningStage → prepares conditioning
|
||||
├─ TimestepPreparationStage → scheduler timesteps
|
||||
├─ LatentPreparationStage → random noise
|
||||
├─ DenoisingStage → iterative denoising loop
|
||||
└─ DecodingStage → latents → video frames
|
||||
```
|
||||
|
||||
## Adding a New Model Family — Checklist
|
||||
|
||||
1. **Pipeline config** — Create a `PipelineConfig` subclass in
|
||||
`fastvideo/configs/pipelines/<model>.py`. Set DiT/VAE/encoder configs,
|
||||
flow_shift, precision defaults.
|
||||
|
||||
2. **Sampling param** — Create a `SamplingParam` subclass in
|
||||
`fastvideo/configs/sample/<model>.py`. Set default height, width,
|
||||
num_frames, guidance_scale, num_inference_steps.
|
||||
|
||||
3. **Register configs** — In `fastvideo/registry.py`, add a
|
||||
`register_configs()` call inside `_register_configs()` with
|
||||
`hf_model_paths` and/or `model_detectors`.
|
||||
|
||||
4. **Pipeline class** — Create a subclass of `ComposedPipelineBase` (or
|
||||
`LoRAPipeline` + `ComposedPipelineBase`) in
|
||||
`fastvideo/pipelines/basic/<model>/<model>_pipeline.py`.
|
||||
- Set `_required_config_modules` listing needed components.
|
||||
- Implement `create_pipeline_stages()` wiring stages via `add_stage()`.
|
||||
- Optionally override `initialize_pipeline()` for custom setup.
|
||||
- Export `EntryClass = YourPipeline` at module level.
|
||||
|
||||
5. **Model classes** (if custom) — Add DiT/VAE implementations in
|
||||
`fastvideo/models/dits/` or `fastvideo/models/vaes/` with
|
||||
`EntryClass = YourModel`. The AST discovery will register them
|
||||
automatically.
|
||||
|
||||
6. **Custom stages** (if needed) — Subclass `PipelineStage` in
|
||||
`fastvideo/pipelines/stages/`, implement `forward()`, optionally
|
||||
implement `verify_input()`/`verify_output()`.
|
||||
|
||||
7. **Verify** — Run `fastvideo generate --model-path <path> --prompt
|
||||
"test" --num-inference-steps 2` to confirm the pipeline loads and
|
||||
generates output.
|
||||
@@ -1,85 +1,104 @@
|
||||
# FastVideo CLI Inference
|
||||
|
||||
The FastVideo CLI exposes the same core inference controls as the Python API.
|
||||
The FastVideo CLI provides a quick way to access the FastVideo inference pipeline for video generation. For more advanced usage,
|
||||
see the Python interface [here](examples/basic.md).
|
||||
|
||||
## Basic Usage
|
||||
|
||||
Use either:
|
||||
|
||||
1. `--model-path` + `--prompt`
|
||||
2. `--model-path` + `--prompt-txt` (batch prompts, one line per prompt)
|
||||
3. `--config` (JSON/YAML)
|
||||
The basic command to generate a video is:
|
||||
|
||||
```bash
|
||||
fastvideo generate --model-path Wan-AI/Wan2.1-T2V-1.3B-Diffusers \
|
||||
--prompt "A cat playing with a ball of yarn"
|
||||
fastvideo generate --model-path {MODEL_PATH} --prompt {PROMPT}
|
||||
```
|
||||
|
||||
```bash
|
||||
fastvideo generate --model-path Wan-AI/Wan2.1-T2V-1.3B-Diffusers \
|
||||
--prompt-txt prompts.txt
|
||||
```
|
||||
### Required Parameters
|
||||
|
||||
You cannot provide both `--prompt` and `--prompt-txt` in the same run.
|
||||
- `--model-path {MODEL_PATH}`: Path to the model or model ID
|
||||
- `--prompt {PROMPT}`: Text description for the video you want to generate
|
||||
|
||||
## View All Arguments
|
||||
## Common Arguments
|
||||
|
||||
To see all the options, you can use the `--help` flag:
|
||||
|
||||
```bash
|
||||
fastvideo generate --help
|
||||
```
|
||||
|
||||
Arguments come from:
|
||||
### Hardware Configuration
|
||||
|
||||
- FastVideo runtime args (`FastVideoArgs`)
|
||||
- Sampling args (`SamplingParam`)
|
||||
- Pipeline config args (`PipelineConfig`)
|
||||
- `--num-gpus {NUM_GPUS}`: Number of GPUs to use
|
||||
- `--tp-size {TP_SIZE}`: Tensor parallelism size (only for the encoder, should not be larger than 1 if text encoder offload is enabled, as layerwise offload + prefetch is faster)
|
||||
- `--sp-size {SP_SIZE}`: Sequence parallelism size (Typically should match the number of GPUs)
|
||||
|
||||
## Common Arguments
|
||||
#### Video Configuration
|
||||
|
||||
### Parallelism
|
||||
- `--height {HEIGHT}`: Height of the generated video
|
||||
- `--width {WIDTH}`: Width of the generated video
|
||||
- `--num-frames {NUM_FRAMES}`: Number of frames to generate
|
||||
- `--fps {FPS}`: Frames per second for the saved video
|
||||
|
||||
- `--num-gpus`
|
||||
- `--sp-size`
|
||||
- `--tp-size`
|
||||
#### Generation Parameters
|
||||
|
||||
### Sampling
|
||||
- `--num-inference-steps {STEPS}`: Number of denoising steps
|
||||
- `--negative-prompt {PROMPT}`: Negative prompt to guide generation away from certain concepts
|
||||
- `--seed {SEED}`: Random seed for reproducible generation
|
||||
|
||||
- `--num-frames`
|
||||
- `--height` / `--width`
|
||||
- `--num-inference-steps`
|
||||
- `--guidance-scale`
|
||||
- `--seed`
|
||||
- `--negative-prompt`
|
||||
#### Output Options
|
||||
|
||||
### Output
|
||||
- `--output-path {PATH}`: Directory to save the generated video
|
||||
- `--save-video`: Whether to save the video to disk
|
||||
- `--return-frames`: Whether to return the raw frames
|
||||
|
||||
- `--output-path`
|
||||
- `--save-video` / `--no-save-video`
|
||||
- `--return-frames`
|
||||
## Using Configuration Files
|
||||
|
||||
### Offloading and Performance
|
||||
|
||||
- `--dit-layerwise-offload`
|
||||
- `--use-fsdp-inference`
|
||||
- `--text-encoder-cpu-offload`
|
||||
- `--image-encoder-cpu-offload`
|
||||
- `--vae-cpu-offload`
|
||||
- `--enable-torch-compile`
|
||||
- `--torch-compile-kwargs`
|
||||
|
||||
## Using Config Files
|
||||
Instead of specifying all parameters on the command line, you can use a configuration file:
|
||||
|
||||
```bash
|
||||
fastvideo generate --config config.yaml
|
||||
fastvideo generate --config {CONFIG_FILE_PATH}
|
||||
```
|
||||
|
||||
Config files can be JSON or YAML. CLI flags override config-file values.
|
||||
The config file should be in JSON or YAML format with the same parameter names as the CLI options. Command-line arguments will take precedence over settings in the configuration file, allowing you to override specific values while keeping the rest from the config file.
|
||||
|
||||
Example `config.yaml`:
|
||||
Example configuration file (config.json):
|
||||
|
||||
```json
|
||||
{
|
||||
"model_path": "FastVideo/FastHunyuan-diffusers",
|
||||
"prompt": "A beautiful woman in a red dress walking down a street",
|
||||
"output_path": "outputs/",
|
||||
"num_gpus": 2,
|
||||
"sp_size": 2,
|
||||
"tp_size": 1,
|
||||
"num_frames": 45,
|
||||
"height": 720,
|
||||
"width": 1280,
|
||||
"num_inference_steps": 6,
|
||||
"seed": 1024,
|
||||
"fps": 24,
|
||||
"precision": "bf16",
|
||||
"vae_precision": "fp16",
|
||||
"vae_tiling": true,
|
||||
"vae_sp": true,
|
||||
"vae_config": {
|
||||
"load_encoder": false,
|
||||
"load_decoder": true,
|
||||
"tile_sample_min_height": 256,
|
||||
"tile_sample_min_width": 256
|
||||
},
|
||||
"text_encoder_precisions": [
|
||||
"fp16",
|
||||
"fp16"
|
||||
],
|
||||
"mask_strategy_file_path": null,
|
||||
"enable_torch_compile": false
|
||||
}
|
||||
```
|
||||
|
||||
Or using YAML format (config.yaml):
|
||||
|
||||
```yaml
|
||||
model_path: "FastVideo/FastHunyuan-diffusers"
|
||||
prompt: "A capybara lounging in a hammock"
|
||||
prompt: "A beautiful woman in a red dress walking down a street"
|
||||
output_path: "outputs/"
|
||||
num_gpus: 2
|
||||
sp_size: 2
|
||||
@@ -89,34 +108,44 @@ height: 720
|
||||
width: 1280
|
||||
num_inference_steps: 6
|
||||
seed: 1024
|
||||
dit_precision: "bf16"
|
||||
fps: 24
|
||||
precision: "bf16"
|
||||
vae_precision: "fp16"
|
||||
vae_tiling: true
|
||||
vae_sp: true
|
||||
vae_config:
|
||||
load_encoder: false
|
||||
load_decoder: true
|
||||
tile_sample_min_height: 256
|
||||
tile_sample_min_width: 256
|
||||
text_encoder_precisions:
|
||||
- "fp16"
|
||||
- "fp16"
|
||||
mask_strategy_file_path: null
|
||||
enable_torch_compile: false
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- Use `dit_precision` / `vae_precision` (not `precision`).
|
||||
- Nested config objects are supported, for example `vae_config` and
|
||||
`dit_config`.
|
||||
|
||||
## Examples
|
||||
|
||||
Simple generation:
|
||||
Generating a simple video:
|
||||
|
||||
```bash
|
||||
fastvideo generate \
|
||||
--model-path FastVideo/FastHunyuan-diffusers \
|
||||
--prompt "A cat playing with a ball of yarn" \
|
||||
--num-frames 45 --height 720 --width 1280 \
|
||||
--num-inference-steps 6 --seed 1024 \
|
||||
--output-path outputs/
|
||||
fastvideo generate --model-path FastVideo/FastHunyuan-diffusers --prompt "A cat playing with a ball of yarn" --num-frames 45 --height 720 --width 1280 --num-inference-steps 6 --seed 1024 --output-path outputs/
|
||||
```
|
||||
|
||||
Config + CLI override:
|
||||
Using a negative prompt to avoid certain elements:
|
||||
|
||||
```bash
|
||||
fastvideo generate --config config.yaml --prompt "A panda skiing at sunset"
|
||||
fastvideo generate --model-path FastVideo/FastHunyuan-diffusers --prompt "A beautiful forest landscape" --negative-prompt "people, buildings, roads"
|
||||
```
|
||||
|
||||
Combining command line arguments and a configuration file:
|
||||
|
||||
```bash
|
||||
fastvideo generate --config config.json --prompt "A capybara lounging in a hammock"
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- If you encounter CUDA out-of-memory errors, try reducing the video dimensions or number of frames, or the number of inference steps.
|
||||
- For reproducible results, set the same seed value between runs.
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
|
||||
# Configuration
|
||||
|
||||
## Multi-GPU Setup
|
||||
@@ -17,8 +18,7 @@ generator = VideoGenerator.from_pretrained(
|
||||
- `PipelineConfig`: Initialization time parameters
|
||||
- `SamplingParam`: Generation time parameters
|
||||
|
||||
You can customize generation behavior using `PipelineConfig` and
|
||||
`SamplingParam`:
|
||||
You can customize various parameters when generating videos using the `PipelineConfig` and `SamplingParam` class:
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator, SamplingParam, PipelineConfig
|
||||
@@ -27,12 +27,12 @@ def main():
|
||||
model_name = "Wan-AI/Wan2.1-T2V-1.3B-Diffusers"
|
||||
config = PipelineConfig.from_pretrained(model_name)
|
||||
config.vae_precision = "fp16"
|
||||
config.dit_cpu_offload = True
|
||||
|
||||
# Create the generator
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_name,
|
||||
num_gpus=1,
|
||||
dit_layerwise_offload=True, # FastVideoArgs option
|
||||
pipeline_config=config
|
||||
)
|
||||
|
||||
@@ -61,46 +61,17 @@ def main():
|
||||
prompt,
|
||||
sampling_param=sampling_param,
|
||||
output_path="my_videos/", # Controls where videos are saved
|
||||
return_frames=True, # Also return frames from this call (defaults to False)
|
||||
save_video=True
|
||||
)
|
||||
|
||||
# If return_frames=True, frames are available in video["frames"]
|
||||
print(f"Generated {len(video['frames'])} frames")
|
||||
# If return_frames=True, video contains the generated frames as a NumPy array
|
||||
print(f"Generated {len(video)} frames")
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
```
|
||||
|
||||
## JSON/YAML Config Files (CLI)
|
||||
|
||||
The CLI supports `--config` with JSON or YAML. Command-line arguments override
|
||||
config file values.
|
||||
By default, `fastvideo generate` uses `return_frames=false` unless you set
|
||||
`--return-frames` (or `return_frames: true` in config).
|
||||
|
||||
```bash
|
||||
fastvideo generate --config config.yaml
|
||||
```
|
||||
|
||||
Use CLI argument names as keys (underscore or hyphen is accepted). Example:
|
||||
|
||||
```yaml
|
||||
model_path: "FastVideo/FastHunyuan-diffusers"
|
||||
prompt: "A capybara relaxing in a hammock"
|
||||
num_gpus: 2
|
||||
sp_size: 2
|
||||
num_frames: 45
|
||||
height: 720
|
||||
width: 1280
|
||||
num_inference_steps: 6
|
||||
seed: 1024
|
||||
dit_precision: "bf16"
|
||||
vae_precision: "fp16"
|
||||
vae_tiling: true
|
||||
vae_sp: true
|
||||
enable_torch_compile: false
|
||||
```
|
||||
|
||||
## Performance Optimization
|
||||
|
||||
For configuring optimizations, please see our [optimizations guide](optimizations.md)
|
||||
|
||||
@@ -11,15 +11,15 @@ This page contains step-by-step instructions to get you quickly started with vid
|
||||
|
||||
## Installation
|
||||
|
||||
If you previously used Conda, we recommend using [uv](https://docs.astral.sh/uv/) instead for a faster and more stable environment setup:
|
||||
We recommend using an environment manager such as `Conda` to create a clean environment:
|
||||
|
||||
```bash
|
||||
# Create and activate a new uv environment
|
||||
uv venv --python 3.12 --seed
|
||||
source .venv/bin/activate
|
||||
# Create and activate a new conda environment
|
||||
conda create -n fastvideo python=3.12
|
||||
conda activate fastvideo
|
||||
|
||||
# Install FastVideo
|
||||
uv pip install fastvideo
|
||||
pip install fastvideo
|
||||
```
|
||||
|
||||
For advanced installation options, see the [Installation Guide](../getting_started/installation.md).
|
||||
@@ -44,6 +44,7 @@ def main():
|
||||
# Generate the video
|
||||
video = generator.generate_video(
|
||||
prompt,
|
||||
return_frames=True, # Also return frames from this call (defaults to False)
|
||||
output_path="my_videos/", # Controls where videos are saved
|
||||
save_video=True
|
||||
)
|
||||
@@ -60,8 +61,7 @@ python example.py
|
||||
|
||||
The generated video will be saved in the current directory under `my_videos/`
|
||||
|
||||
More inference scripts and recipes can be found in `examples/inference/` and
|
||||
`scripts/inference/`.
|
||||
More inference example scripts can be found in `scripts/inference/`
|
||||
|
||||
## Available Models
|
||||
|
||||
@@ -103,10 +103,11 @@ Common issues and their solutions:
|
||||
If you encounter CUDA out of memory errors:
|
||||
|
||||
- Reduce `num_frames` or video resolution
|
||||
- Enable FastVideo offloading options such as `dit_layerwise_offload=True`
|
||||
(single GPU) or `use_fsdp_inference=True` (multi-GPU)
|
||||
- Enable memory optimization with `enable_model_cpu_offload`
|
||||
- Try a smaller model or use distilled versions
|
||||
- Use `num_gpus` > 1 if multiple GPUs are available
|
||||
- Try enabling FSDP inference with `use_fsdp_inference=True` (may slow down generation)
|
||||
- Try enabling DiT layerwise offload with `dit_layerwise_offload=True` (now only a few models support this, but may introduce less overhead than FSDP)
|
||||
|
||||
### Slow Generation
|
||||
|
||||
|
||||
@@ -1,140 +0,0 @@
|
||||
# Offloading
|
||||
|
||||
This page describes how to use offloading techniques for inference to reduce GPU memory usage while maintaining acceptable performance.
|
||||
|
||||
## Default Behavior
|
||||
|
||||
```python
|
||||
dit_cpu_offload: bool = True
|
||||
use_fsdp_inference: bool = False
|
||||
dit_layerwise_offload: bool = True
|
||||
text_encoder_cpu_offload: bool = True
|
||||
image_encoder_cpu_offload: bool = True
|
||||
vae_cpu_offload: bool = True
|
||||
pin_cpu_memory: bool = True
|
||||
```
|
||||
|
||||
## Behavior Explanation
|
||||
|
||||
!!! note
|
||||
For CLI usage, replace underscores (`_`) with hyphens (`-`).
|
||||
|
||||
### `use_fsdp_inference`
|
||||
|
||||
Enables [FSDP](https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html) for inference. The model weights are sharded across multiple GPUs to reduce memory usage per GPU, and weights are broadcast to all GPUs layer by layer during inference.
|
||||
|
||||
#### Performance Impact
|
||||
|
||||
FSDP inference introduces negligible performance overhead due to weight prefetching. Performance overhead may be visible when GPU interconnect is slow (e.g., multiple consumer-level GPUs connected by slow PCIe without GPU P2P support).
|
||||
|
||||
#### Usage Recommendation
|
||||
|
||||
We recommend enabling this option when multiple GPUs are available.
|
||||
|
||||
### `dit_cpu_offload`
|
||||
|
||||
Enables CPU offloading for FSDP inference. When enabled, the model weights are offloaded to CPU memory, and the weight of each layer is moved to GPU memory only when that layer is being computed.
|
||||
|
||||
#### Performance Impact
|
||||
|
||||
The PyTorch FSDP implementation does not overlap computation and data transfer perfectly for inference, so enabling this option will harm performance.
|
||||
|
||||
#### Usage Recommendation
|
||||
|
||||
This option only takes effect when FSDP is enabled. For single GPU usage, we recommend using `dit_layerwise_offload` instead.
|
||||
|
||||
### `dit_layerwise_offload`
|
||||
|
||||
This option is similar to `dit_cpu_offload`, but with two key differences:
|
||||
|
||||
1. It overlaps computation and PCIe data transfer
|
||||
2. It only works for single GPU inference
|
||||
|
||||
#### Performance Impact
|
||||
|
||||
This option introduces negligible performance overhead.
|
||||
|
||||
#### Usage Recommendation
|
||||
|
||||
We recommend enabling this option for single GPU usage. This option is not compatible with FSDP.
|
||||
|
||||
### `text_encoder_cpu_offload`
|
||||
|
||||
When enabled, the text encoder model weights are offloaded to CPU memory, and text encoding is computed on CPU.
|
||||
|
||||
#### Performance Impact
|
||||
|
||||
This option significantly slows down text encoding computation, but text encoding is usually not the bottleneck.
|
||||
|
||||
#### Usage Recommendation
|
||||
|
||||
We recommend enabling this option only when OOM happens.
|
||||
|
||||
### `image_encoder_cpu_offload` and `vae_cpu_offload`
|
||||
|
||||
When enabled, the weights are stored in CPU memory and moved to GPU memory when the corresponding module is being computed. After computation, the weights are moved back to CPU memory.
|
||||
|
||||
#### Performance Impact
|
||||
|
||||
These options introduce performance overhead due to PCIe data transfer.
|
||||
|
||||
#### Usage Recommendation
|
||||
|
||||
We recommend enabling these options when OOM happens.
|
||||
|
||||
## General Recommendations
|
||||
|
||||
### Single GPU Inference
|
||||
|
||||
We recommend enabling `dit_layerwise_offload`. If OOM happens, also enable `image_encoder_cpu_offload` and `vae_cpu_offload`. If OOM still happens, consider enabling `text_encoder_cpu_offload`.
|
||||
|
||||
### Multi-GPU Inference
|
||||
|
||||
We recommend enabling `use_fsdp_inference` and disabling both `dit_layerwise_offload` and `dit_cpu_offload`. If OOM happens, consider enabling `text_encoder_cpu_offload`, `image_encoder_cpu_offload`, and `vae_cpu_offload`. If OOM still happens, consider enabling `dit_cpu_offload`.
|
||||
|
||||
## Examples
|
||||
|
||||
### Single GPU with Layerwise Offloading
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
num_gpus=1,
|
||||
# Recommended for single GPU
|
||||
dit_layerwise_offload=True,
|
||||
# Enable if OOM happens
|
||||
vae_cpu_offload=True,
|
||||
image_encoder_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
# Speeds up CPU-GPU transfer
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
prompt = "A curious raccoon peers through a vibrant field of yellow sunflowers."
|
||||
video = generator.generate_video(prompt, output_path="output/", save_video=True)
|
||||
```
|
||||
|
||||
### Multi-GPU with FSDP
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
num_gpus=2,
|
||||
# Recommended for multi-GPU
|
||||
use_fsdp_inference=True,
|
||||
dit_layerwise_offload=False,
|
||||
dit_cpu_offload=False,
|
||||
# Enable if OOM happens
|
||||
vae_cpu_offload=True,
|
||||
image_encoder_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
prompt = "A majestic lion strides across the golden savanna."
|
||||
video = generator.generate_video(prompt, output_path="output/", save_video=True)
|
||||
```
|
||||
@@ -8,24 +8,23 @@ This page describes the various options for speeding up generation times in Fast
|
||||
- Optimized Attention Backends
|
||||
|
||||
- [Flash Attention](#flash-attention)
|
||||
- [Sliding Tile Attention (Archived)](#sliding-tile-attention-archived)
|
||||
- [Sliding Tile Attention](#sliding-tile-attention)
|
||||
- [Sage Attention](#sage-attention)
|
||||
- [Sage Attention 3](#sage-attention-3)
|
||||
|
||||
- Caching Techniques
|
||||
- [TeaCache](#teacache)
|
||||
|
||||
## Attention Backends
|
||||
|
||||
### Available Backends
|
||||
|
||||
- Torch SDPA: `FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA`
|
||||
- Flash Attention 2 and 3: `FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN`
|
||||
- Sliding Tile Attention: `FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN`
|
||||
- Video Sparse Attention: `FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN`
|
||||
- Sage Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN`
|
||||
- Sage Attention 3: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN_THREE`
|
||||
- Video MoBA Attention: `FASTVIDEO_ATTENTION_BACKEND=VMOBA_ATTN`
|
||||
- Sparse Linear Attention: `FASTVIDEO_ATTENTION_BACKEND=SLA_ATTN`
|
||||
- SageSLA Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_SLA_ATTN`
|
||||
- Sliding Tile Attention (archived branch only):
|
||||
`FASTVIDEO_ATTENTION_BACKEND=SLIDING_TILE_ATTN`
|
||||
|
||||
### Configuring Backends
|
||||
|
||||
@@ -36,7 +35,7 @@ There are two ways to configure the attention backend in FastVideo.
|
||||
In python, set the `FASTVIDEO_ATTENTION_BACKEND` environment variable before instantiating `VideoGenerator` like this:
|
||||
|
||||
```python
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "VIDEO_SPARSE_ATTN"
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "SLIDING_TILE_ATTN"
|
||||
```
|
||||
|
||||
#### 2. In CLI
|
||||
@@ -67,27 +66,26 @@ pip install ninja
|
||||
python setup.py install
|
||||
```
|
||||
|
||||
### Sliding Tile Attention (Archived)
|
||||
### Sliding Tile Attention
|
||||
|
||||
**`SLIDING_TILE_ATTN`**
|
||||
|
||||
The full STA integration in `fastvideo/` is archived from `main` and preserved
|
||||
at:
|
||||
```bash
|
||||
pip install st_attn==0.0.4
|
||||
```
|
||||
|
||||
- https://github.com/hao-ai-lab/FastVideo/tree/sta_do_not_delete
|
||||
|
||||
We keep STA off `main` because we believe VSA is strictly better than STA for
|
||||
the actively maintained FastVideo path.
|
||||
|
||||
Kernel code in `fastvideo-kernel` is still retained. For mask search and STA
|
||||
inference workflow, see [STA docs](../attention/sta/index.md).
|
||||
Please see [this page](../attention/sta/index.md) for more installation instructions.
|
||||
|
||||
### Video Sparse Attention
|
||||
|
||||
**`VIDEO_SPARSE_ATTN`**
|
||||
|
||||
Video Sparse Attention is provided by `fastvideo-kernel`.
|
||||
See [VSA docs](../attention/vsa/index.md) for installation details.
|
||||
```bash
|
||||
git submodule update --init --recursive
|
||||
python setup_vsa.py install
|
||||
```
|
||||
|
||||
Please see [this page](../attention/vsa/index.md) for more installation instructions.
|
||||
|
||||
### Sage Attention
|
||||
|
||||
@@ -105,7 +103,7 @@ python setup.py install # or pip install -e .
|
||||
|
||||
**`SAGE_ATTN_THREE`**
|
||||
|
||||
[SageAttention 3](https://github.com/thu-ml/SageAttention/tree/main/sageattention3_blackwell) is an advanced attention mechanism that leverages FP4 quantization and Blackwell GPU Tensor Cores for significant performance improvements.
|
||||
[SageAttention 3](https://huggingface.co/jt-zhang/SageAttention3) is an advanced attention mechanism that leverages FP4 quantization and Blackwell GPU Tensor Cores for significant performance improvements.
|
||||
|
||||
#### Hardware Requirements
|
||||
|
||||
@@ -115,32 +113,69 @@ python setup.py install # or pip install -e .
|
||||
|
||||
Note that Sage Attention 3 requires `python>=3.13`, `torch>=2.8.0`, `CUDA >=12.8`. If you are using `uv` and using `torch==2.8.0` make sure that `sentencepiece==0.2.1` in the pyproject.toml file.
|
||||
|
||||
To use Sage Attention 3 in FastVideo, follow the `README.md` in the linked repository to install the package from source.
|
||||
To use Sage Attention 3 in FastVideo, first get access to the SageAttention3 code, then move `sageattn/` and `setup.py` to the directory `fastvideo/attention/backends`, then install from using:
|
||||
|
||||
### V-MoBA / SLA / SageSLA
|
||||
```bash
|
||||
python setup.py install
|
||||
```
|
||||
|
||||
These backends are model-specific and require the corresponding kernels and
|
||||
dependencies. Use the support matrix and model examples to confirm compatibility
|
||||
before enabling them.
|
||||
## Teacache
|
||||
|
||||
TeaCache is an optimization technique supported in FastVideo that can significantly speed up video generation by skipping redundant calculations across diffusion steps. This guide explains how to enable and configure TeaCache for optimal performance in FastVideo.
|
||||
|
||||
### What is TeaCache?
|
||||
|
||||
See the official [TeaCache](https://github.com/ali-vilab/TeaCache) repo and their paper for more details.
|
||||
|
||||
### How to Enable TeaCache
|
||||
|
||||
Enabling TeaCache is straightforward - simply add the `enable_teacache=True` parameter to your `generate_video()` call:
|
||||
|
||||
```python
|
||||
# ... previous code
|
||||
generator.generate_video(
|
||||
prompt="Your prompt here",
|
||||
sampling_param=params,
|
||||
enable_teacache=True
|
||||
)
|
||||
# more code ...
|
||||
```
|
||||
|
||||
### Complete Example
|
||||
|
||||
At the bottom is a complete example of using TeaCache for faster video generation. You can run it using the following command:
|
||||
|
||||
```bash
|
||||
python examples/inference/optimizations/teacache_example.py
|
||||
```
|
||||
|
||||
### Advanced Configuration
|
||||
|
||||
While TeaCache works well with default settings, you can fine-tune its behavior by adjusting the threshold value:
|
||||
|
||||
1. Lower threshold values (e.g., 0.1) will result in more skipped calculations and faster generation with slightly more potential for quality degradation
|
||||
2. Higher threshold values (e.g., 0.15-0.23) will skip fewer calculations but maintain quality closer to the original
|
||||
|
||||
Note that the optimal threshold depends on your specific model and content.
|
||||
|
||||
## Benchmarking different optimizations
|
||||
|
||||
To benchmark backend performance, generate the same prompt with the same seed and compare end-to-end generation times:
|
||||
To benchmark the performance improvement, try generating the same video with and without TeaCache enabled and compare the generation times:
|
||||
|
||||
```python
|
||||
import os
|
||||
import time
|
||||
# Without TeaCache
|
||||
start_time = time.perf_counter()
|
||||
generator.generate_video(prompt="Your prompt", enable_teacache=False)
|
||||
standard_time = time.perf_counter() - start_time
|
||||
|
||||
for backend in ["TORCH_SDPA", "FLASH_ATTN", "SAGE_ATTN"]:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = backend
|
||||
generator = VideoGenerator.from_pretrained("your-model-id")
|
||||
start_time = time.perf_counter()
|
||||
generator.generate_video(
|
||||
prompt="Your prompt",
|
||||
seed=1024,
|
||||
)
|
||||
elapsed = time.perf_counter() - start_time
|
||||
print(f"{backend}: {elapsed:.2f}s")
|
||||
# With TeaCache
|
||||
start_time = time.perf_counter()
|
||||
generator.generate_video(prompt="Your prompt", enable_teacache=True)
|
||||
teacache_time = time.perf_counter() - start_time
|
||||
|
||||
print(f"Standard generation: {standard_time:.2f} seconds")
|
||||
print(f"TeaCache generation: {teacache_time:.2f} seconds")
|
||||
print(f"Speedup: {standard_time/teacache_time:.2f}x")
|
||||
```
|
||||
|
||||
Note: reinstantiate `VideoGenerator` after changing `FASTVIDEO_ATTENTION_BACKEND`.
|
||||
Note: If you want to benchmark different attention backends, you'll need to reinstantiate `VideoGenerator`.
|
||||
|
||||
@@ -1,17 +1,6 @@
|
||||
# Compatibility Matrix
|
||||
|
||||
This page summarizes common model + optimization combinations.
|
||||
|
||||
For the canonical, code-level list of model IDs recognized by
|
||||
`VideoGenerator.from_pretrained(...)`, see the registrations in
|
||||
`fastvideo/registry.py` (`register_configs(...)` entries).
|
||||
|
||||
!!! note
|
||||
The full STA integration in `fastvideo/` is archived from `main` and kept
|
||||
in `sta_do_not_delete`:
|
||||
https://github.com/hao-ai-lab/FastVideo/tree/sta_do_not_delete
|
||||
We do this because we believe VSA is strictly better than STA for the
|
||||
actively maintained `main` inference path.
|
||||
The table below shows every supported model and optimizations supported for them.
|
||||
|
||||
The symbols used have the following meanings:
|
||||
|
||||
@@ -21,9 +10,7 @@ The symbols used have the following meanings:
|
||||
|
||||
## Models x Optimization
|
||||
|
||||
The `HuggingFace Model ID` can be passed directly to
|
||||
`from_pretrained()`. FastVideo then uses model-specific default settings for
|
||||
pipeline initialization and sampling.
|
||||
The `HuggingFace Model ID` can be directly pass to `from_pretrained()` methods and FastVideo will use the optimal default parameters when initializing and generating videos.
|
||||
|
||||
<style>
|
||||
/* Target tables in this section */
|
||||
@@ -53,7 +40,7 @@ pipeline initialization and sampling.
|
||||
}
|
||||
</style>
|
||||
|
||||
| Model Name | HuggingFace Model ID | Resolutions | TeaCache | Sliding Tile Attn (Legacy Branch) | Sage Attn | VSA | BSA |
|
||||
| Model Name | HuggingFace Model ID | Resolutions | TeaCache | Sliding Tile Attn | Sage Attn | VSA | BSA |
|
||||
|------------|---------------------|-------------|----------|-------------------|-----------|-----|-----|
|
||||
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
@@ -66,9 +53,9 @@ pipeline initialization and sampling.
|
||||
| Wan2.1 T2V 14B | `Wan-AI/Wan2.1-T2V-14B-Diffusers` | 480P, 720P | ✅ | ✅* | ✅ | ⭕ | ⭕ |
|
||||
| Wan2.1 I2V 480P | `Wan-AI/Wan2.1-I2V-14B-480P-Diffusers` | 480P | ✅ | ✅* | ✅ | ⭕ | ⭕ |
|
||||
| Wan2.1 I2V 720P | `Wan-AI/Wan2.1-I2V-14B-720P-Diffusers` | 720P | ✅ | ✅ | ✅ | ⭕ | ⭕ |
|
||||
| StepVideo T2V | `FastVideo/stepvideo-t2v-diffusers` | 768px768px204f<br>544px992px204f<br>544px992px136f | ❌ | ❌ | ✅ | ⭕ | ⭕ |
|
||||
| TurboWan2.1 T2V 1.3B | `loayrashid/TurboWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| TurboWan2.1 T2V 14B | `loayrashid/TurboWan2.1-T2V-14B-Diffusers` | 480P, 720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| TurboWan2.2 I2V A14B | `loayrashid/TurboWan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| LongCat T2V 13.6B | See note** | 480P<br>720P | ❌ | ❌ | ❌ | ⭕ | ✅ |
|
||||
| Matrix Game 2.0 Base | `FastVideo/Matrix-Game-2.0-Base-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Matrix Game 2.0 GTA | `FastVideo/Matrix-Game-2.0-GTA-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
@@ -76,21 +63,13 @@ pipeline initialization and sampling.
|
||||
|
||||
**Note**: Wan2.2 TI2V 5B has some quality issues when performing I2V generation. We are working on fixing this issue.
|
||||
|
||||
`Sliding Tile Attn (Legacy Branch)` entries refer to the archived
|
||||
`sta_do_not_delete` branch workflow, not active `main` inference wiring.
|
||||
|
||||
## Canonical Supported IDs
|
||||
|
||||
The authoritative source for model-ID recognition is
|
||||
`fastvideo/registry.py`. If a model ID is registered there, FastVideo can
|
||||
resolve default pipeline and sampling configuration for it.
|
||||
|
||||
## Special requirements
|
||||
|
||||
### StepVideo T2V
|
||||
- The self-attention in text-encoder (step_llm) only supports CUDA capabilities sm_80 sm_86 and sm_90
|
||||
|
||||
### Sliding Tile Attention
|
||||
- Full STA pipeline usage is on the archived branch:
|
||||
https://github.com/hao-ai-lab/FastVideo/tree/sta_do_not_delete
|
||||
- STA currently requires Hopper GPUs (H100s).
|
||||
- Currently only Hopper GPUs (H100s) are supported.
|
||||
|
||||
### TurboWan2.1 (TurboDiffusion)
|
||||
- Uses TurboDiffusionPipeline with RCM scheduler for 1-4 step generation
|
||||
|
||||
@@ -1,595 +0,0 @@
|
||||
# Training Infrastructure
|
||||
|
||||
FastVideo's training infrastructure (`fastvideo/train/`) is a YAML-driven
|
||||
framework for training and distilling video diffusion models. A single config
|
||||
file controls everything — models, algorithms, distributed strategy,
|
||||
checkpointing, and validation — with no code changes needed to mix and match.
|
||||
|
||||
!!! note "Relationship to legacy training"
|
||||
This system replaces the older script-based training in `fastvideo/training/`.
|
||||
The legacy scripts still work for basic fine-tuning, but new development
|
||||
should use the config-driven system documented here.
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Launch with the helper script
|
||||
|
||||
```bash
|
||||
bash examples/train/run.sh examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml
|
||||
```
|
||||
|
||||
The script auto-detects available GPUs and sets up `torchrun`. Override with
|
||||
environment variables:
|
||||
|
||||
```bash
|
||||
NUM_GPUS=4 NNODES=2 NODE_RANK=0 \
|
||||
MASTER_ADDR=10.0.0.1 MASTER_PORT=29501 \
|
||||
bash examples/train/run.sh my_config.yaml
|
||||
```
|
||||
|
||||
### Launch directly with torchrun
|
||||
|
||||
```bash
|
||||
torchrun --nproc_per_node=8 \
|
||||
fastvideo/train/entrypoint/train.py \
|
||||
--config examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml
|
||||
```
|
||||
|
||||
### CLI flags
|
||||
|
||||
| Flag | Description |
|
||||
|------|-------------|
|
||||
| `--config` | Path to YAML config file (required) |
|
||||
| `--resume-from-checkpoint` | Path to a DCP checkpoint directory to resume from |
|
||||
| `--override-output-dir` | Override `training.checkpoint.output_dir` |
|
||||
| `--dry-run` | Validate config and exit without training |
|
||||
|
||||
---
|
||||
|
||||
## Config Format
|
||||
|
||||
Every run is defined by a single YAML file with five top-level sections.
|
||||
See `examples/train/example.yaml` for a fully-commented reference.
|
||||
|
||||
### `models` — Role-based model instances
|
||||
|
||||
Each entry defines a model role. The `_target_` field specifies the Python class
|
||||
to instantiate:
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
teacher:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: false
|
||||
disable_custom_init_weights: true
|
||||
```
|
||||
|
||||
Common model parameters:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
|-----------|---------|-------------|
|
||||
| `_target_` | *(required)* | Python class path for the model |
|
||||
| `init_from` | *(required)* | HuggingFace repo ID or local checkpoint path |
|
||||
| `trainable` | `true` | Whether the model's parameters require gradients |
|
||||
| `disable_custom_init_weights` | `false` | Skip custom weight initialization (use for teacher/critic) |
|
||||
| `flow_shift` | `3.0` | Timestep shifting factor |
|
||||
| `enable_gradient_checkpointing_type` | `null` | Gradient checkpointing (`"full"` or `null`) |
|
||||
|
||||
Which roles are needed depends on the training method:
|
||||
|
||||
| Method | Required roles |
|
||||
|--------|---------------|
|
||||
| Fine-tune (SFT) | `student` |
|
||||
| Diffusion-Forcing SFT | `student` |
|
||||
| DMD2 | `student`, `teacher`, `critic` |
|
||||
| Self-Forcing | `student` (causal), `teacher`, `critic` |
|
||||
|
||||
### `method` — Training algorithm
|
||||
|
||||
Selects and configures the training algorithm:
|
||||
|
||||
```yaml
|
||||
method:
|
||||
_target_: fastvideo.train.methods.distribution_matching.dmd2.DMD2Method
|
||||
rollout_mode: simulate
|
||||
dmd_denoising_steps: [1000, 750, 500, 250]
|
||||
generator_update_interval: 5
|
||||
```
|
||||
|
||||
To switch algorithms, change `_target_` and adjust the method-specific keys.
|
||||
See [Training Methods](#training-methods) for details on each algorithm.
|
||||
|
||||
### `training` — Typed infrastructure config
|
||||
|
||||
This section maps to typed dataclasses with defaults and validation:
|
||||
|
||||
```yaml
|
||||
training:
|
||||
distributed:
|
||||
num_gpus: 8
|
||||
sp_size: 1 # sequence parallelism
|
||||
tp_size: 1 # tensor parallelism
|
||||
hsdp_replicate_dim: 1 # HSDP replication dimension
|
||||
hsdp_shard_dim: 8 # HSDP sharding dimension
|
||||
|
||||
data:
|
||||
data_path: data/my_dataset
|
||||
train_batch_size: 1
|
||||
dataloader_num_workers: 4
|
||||
training_cfg_rate: 0.1 # classifier-free guidance dropout rate
|
||||
seed: 1000
|
||||
num_latent_t: 20
|
||||
num_height: 448
|
||||
num_width: 832
|
||||
num_frames: 77
|
||||
|
||||
optimizer:
|
||||
learning_rate: 2.0e-6
|
||||
betas: [0.9, 0.999]
|
||||
weight_decay: 0.01
|
||||
lr_scheduler: constant # constant, linear, cosine, polynomial
|
||||
lr_warmup_steps: 0
|
||||
|
||||
loop:
|
||||
max_train_steps: 4000
|
||||
gradient_accumulation_steps: 1
|
||||
|
||||
checkpoint:
|
||||
output_dir: outputs/my_run
|
||||
training_state_checkpointing_steps: 1000 # 0 = disabled
|
||||
checkpoints_total_limit: 3 # 0 = keep all
|
||||
|
||||
tracker:
|
||||
project_name: my_project
|
||||
run_name: my_run
|
||||
|
||||
model:
|
||||
weighting_scheme: uniform # uniform, logit_normal, mode
|
||||
precondition_outputs: false
|
||||
enable_gradient_checkpointing_type: full
|
||||
|
||||
vsa:
|
||||
sparsity: 0.0 # 0.0 = disabled
|
||||
decay_rate: 0.0
|
||||
decay_interval_steps: 0
|
||||
```
|
||||
|
||||
### `callbacks` — Pluggable hooks
|
||||
|
||||
Callbacks run at specific points in the training loop (before/after optimizer
|
||||
steps, at validation time, etc.):
|
||||
|
||||
```yaml
|
||||
callbacks:
|
||||
grad_clip:
|
||||
max_grad_norm: 1.0
|
||||
|
||||
ema:
|
||||
_target_: fastvideo.train.callbacks.ema.EMACallback
|
||||
decay: 0.9999
|
||||
start_iter: 0
|
||||
|
||||
validation:
|
||||
_target_: fastvideo.train.callbacks.validation.ValidationCallback
|
||||
pipeline_target: fastvideo.pipelines.basic.wan.wan_pipeline.WanPipeline
|
||||
dataset_file: path/to/validation.json
|
||||
every_steps: 100
|
||||
sampling_steps: [4]
|
||||
guidance_scale: 5.0
|
||||
```
|
||||
|
||||
See [Callbacks](#callbacks) for details on each callback.
|
||||
|
||||
### `pipeline` — Inference pipeline overrides
|
||||
|
||||
Optional overrides for the inference pipeline used during validation:
|
||||
|
||||
```yaml
|
||||
pipeline:
|
||||
flow_shift: 8
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Training Methods
|
||||
|
||||
### Supervised Fine-Tuning (SFT)
|
||||
|
||||
Standard flow-matching loss. The simplest method — train the student to predict
|
||||
noise (or clean x0) from noised data samples.
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.fine_tuning.finetune.FineTuneMethod
|
||||
attn_kind: dense # "dense" or "vsa"
|
||||
```
|
||||
|
||||
| Parameter | Default | Description |
|
||||
|-----------|---------|-------------|
|
||||
| `attn_kind` | `"dense"` | Attention mode: `"dense"` (standard) or `"vsa"` (sparse) |
|
||||
|
||||
### Diffusion-Forcing SFT (DFSFT)
|
||||
|
||||
SFT with **per-chunk inhomogeneous timesteps** — each temporal chunk of the
|
||||
video gets a different noise level. This is a prerequisite for training causal /
|
||||
streaming models that must handle mixed-noise inputs.
|
||||
|
||||
```yaml
|
||||
method:
|
||||
_target_: fastvideo.train.methods.fine_tuning.dfsft.DiffusionForcingSFTMethod
|
||||
chunk_size: 3
|
||||
min_timestep_ratio: 0.0
|
||||
max_timestep_ratio: 1.0
|
||||
attn_kind: dense
|
||||
```
|
||||
|
||||
| Parameter | Default | Description |
|
||||
|-----------|---------|-------------|
|
||||
| `chunk_size` | `3` | Latent frames per temporal chunk |
|
||||
| `min_timestep_ratio` | `0.0` | Lower bound of timestep sampling range |
|
||||
| `max_timestep_ratio` | `1.0` | Upper bound of timestep sampling range |
|
||||
| `attn_kind` | `"dense"` | `"dense"` or `"vsa"` |
|
||||
|
||||
### DMD2 (Distribution Matching Distillation)
|
||||
|
||||
Distill a many-step teacher into a few-step student. The student learns to match
|
||||
the teacher's score function, guided by a trainable critic network.
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
teacher:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: false
|
||||
disable_custom_init_weights: true
|
||||
critic:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
disable_custom_init_weights: true
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.distribution_matching.dmd2.DMD2Method
|
||||
rollout_mode: simulate
|
||||
dmd_denoising_steps: [1000, 750, 500, 250]
|
||||
generator_update_interval: 5
|
||||
real_score_guidance_scale: 4.5
|
||||
|
||||
fake_score_learning_rate: 8.0e-6
|
||||
fake_score_betas: [0.0, 0.999]
|
||||
fake_score_lr_scheduler: constant
|
||||
```
|
||||
|
||||
| Parameter | Default | Description |
|
||||
|-----------|---------|-------------|
|
||||
| `rollout_mode` | *(required)* | `"simulate"` (pure noise) or `"data_latent"` (from data) |
|
||||
| `dmd_denoising_steps` | *(required)* | Timestep schedule for student rollout |
|
||||
| `generator_update_interval` | `1` | Update student every N critic steps |
|
||||
| `real_score_guidance_scale` | `1.0` | CFG scale for teacher predictions |
|
||||
| `fake_score_learning_rate` | *(required)* | Critic optimizer learning rate |
|
||||
| `fake_score_betas` | *(required)* | Critic optimizer Adam betas |
|
||||
| `fake_score_lr_scheduler` | *(required)* | Critic LR scheduler type |
|
||||
|
||||
### Self-Forcing (Causal DMD)
|
||||
|
||||
Extends DMD2 for **streaming / causal video generation**. The student processes
|
||||
video in temporal chunks, feeding its own denoised outputs as context for future
|
||||
chunks — simulating autoregressive rollout during training.
|
||||
|
||||
Requires a causal model class (e.g., `WanCausalModel`) for the student:
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.wan_causal.WanCausalModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.distribution_matching.self_forcing.SelfForcingMethod
|
||||
rollout_mode: simulate
|
||||
dmd_denoising_steps: [1000, 750, 500, 250]
|
||||
student_sample_type: sde
|
||||
context_noise: 0.0
|
||||
enable_gradient_in_rollout: true
|
||||
start_gradient_frame: 0
|
||||
```
|
||||
|
||||
Self-Forcing inherits all DMD2 parameters, plus:
|
||||
|
||||
| Parameter | Default | Description |
|
||||
|-----------|---------|-------------|
|
||||
| `student_sample_type` | `"sde"` | `"sde"` or `"ode"` for intermediate steps |
|
||||
| `same_step_across_blocks` | `false` | Use same exit timestep for all blocks |
|
||||
| `last_step_only` | `false` | Always exit at the final denoising step |
|
||||
| `context_noise` | `0.0` | Noise added to context frames (0 = clean) |
|
||||
| `enable_gradient_in_rollout` | `true` | Enable backprop through rollout |
|
||||
| `start_gradient_frame` | `0` | Frame index where gradients begin |
|
||||
|
||||
---
|
||||
|
||||
## Callbacks
|
||||
|
||||
Callbacks are pluggable hooks that run at specific points in the training loop.
|
||||
Configure them under the `callbacks` section.
|
||||
|
||||
### GradNormClipCallback
|
||||
|
||||
Clips gradient norms before the optimizer step. Optionally logs per-module
|
||||
gradient norms to the tracker.
|
||||
|
||||
```yaml
|
||||
callbacks:
|
||||
grad_clip:
|
||||
max_grad_norm: 1.0 # 0.0 = disabled
|
||||
log_grad_norms: false
|
||||
```
|
||||
|
||||
### EMACallback
|
||||
|
||||
Maintains an exponential moving average of the student's weights. The EMA
|
||||
weights are automatically swapped in during validation.
|
||||
|
||||
```yaml
|
||||
callbacks:
|
||||
ema:
|
||||
_target_: fastvideo.train.callbacks.ema.EMACallback
|
||||
decay: 0.9999
|
||||
start_iter: 0 # delay EMA updates until this iteration
|
||||
```
|
||||
|
||||
The EMA callback owns its own state and checkpoints independently — EMA weights
|
||||
are saved and restored automatically on resume.
|
||||
|
||||
### ValidationCallback
|
||||
|
||||
Runs inference with the trained model at regular intervals, saving generated
|
||||
videos and logging them to the tracker (W&B).
|
||||
|
||||
```yaml
|
||||
callbacks:
|
||||
validation:
|
||||
_target_: fastvideo.train.callbacks.validation.ValidationCallback
|
||||
pipeline_target: fastvideo.pipelines.basic.wan.wan_pipeline.WanPipeline
|
||||
dataset_file: path/to/validation.json
|
||||
every_steps: 100
|
||||
sampling_steps: [4]
|
||||
sampling_timesteps: [1000, 750, 500, 250] # explicit timestep list
|
||||
guidance_scale: 5.0
|
||||
rollout_mode: parallel # "parallel" or "streaming"
|
||||
```
|
||||
|
||||
The validation dataset is a JSON file containing a list of prompt strings.
|
||||
If EMA is enabled, validation automatically uses the EMA weights.
|
||||
|
||||
---
|
||||
|
||||
## Checkpointing and Resume
|
||||
|
||||
### Checkpoint format
|
||||
|
||||
Checkpoints use PyTorch Distributed Checkpoint (DCP) format, compatible with
|
||||
FSDP/HSDP sharding. Each checkpoint saves:
|
||||
|
||||
- Model weights (all roles)
|
||||
- Optimizer states (all roles)
|
||||
- LR scheduler states
|
||||
- RNG states (for exact reproducibility)
|
||||
- EMA shadow weights (if enabled)
|
||||
- Training step counter
|
||||
|
||||
Checkpoints are saved to `<output_dir>/checkpoint-<step>/`.
|
||||
|
||||
### Saving checkpoints
|
||||
|
||||
```yaml
|
||||
training:
|
||||
checkpoint:
|
||||
output_dir: outputs/my_run
|
||||
training_state_checkpointing_steps: 1000 # save every N steps (0 = off)
|
||||
checkpoints_total_limit: 3 # rolling window (0 = keep all)
|
||||
```
|
||||
|
||||
### Resuming training
|
||||
|
||||
Use `--resume-from-checkpoint` to resume from a specific checkpoint:
|
||||
|
||||
```bash
|
||||
# Via the helper script
|
||||
bash examples/train/run.sh my_config.yaml --resume outputs/my_run/checkpoint-2000
|
||||
|
||||
# Via torchrun directly
|
||||
torchrun --nproc_per_node=8 \
|
||||
fastvideo/train/entrypoint/train.py \
|
||||
--config my_config.yaml \
|
||||
--resume-from-checkpoint outputs/my_run/checkpoint-2000
|
||||
```
|
||||
|
||||
Or set it in the YAML:
|
||||
|
||||
```yaml
|
||||
training:
|
||||
checkpoint:
|
||||
resume_from_checkpoint: outputs/my_run/checkpoint-2000
|
||||
```
|
||||
|
||||
### Reproducibility
|
||||
|
||||
The training entrypoint enables deterministic mode automatically:
|
||||
|
||||
- `torch.backends.cudnn.benchmark = False`
|
||||
- `torch.backends.cudnn.deterministic = True`
|
||||
- `torch.use_deterministic_algorithms(True)`
|
||||
|
||||
A shared CUDA RNG generator is seeded from `training.data.seed` and threaded
|
||||
through all random operations (noise sampling, timestep sampling, etc.).
|
||||
Ranks within the same sequence-parallel group share a seed, ensuring identical
|
||||
noise across SP shards.
|
||||
|
||||
---
|
||||
|
||||
## Distributed Training
|
||||
|
||||
The framework supports HSDP (Hybrid Sharded Data Parallel), Tensor Parallelism
|
||||
(TP), and Sequence Parallelism (SP):
|
||||
|
||||
```yaml
|
||||
training:
|
||||
distributed:
|
||||
num_gpus: 8
|
||||
sp_size: 1 # sequence parallelism group size
|
||||
tp_size: 1 # tensor parallelism group size
|
||||
hsdp_replicate_dim: 1 # number of HSDP replicas
|
||||
hsdp_shard_dim: 8 # number of HSDP shards
|
||||
```
|
||||
|
||||
**HSDP** shards model parameters across `hsdp_shard_dim` GPUs and replicates
|
||||
across `hsdp_replicate_dim` groups. The product
|
||||
`hsdp_replicate_dim * hsdp_shard_dim` should equal `num_gpus`.
|
||||
|
||||
**Sequence parallelism** splits the sequence (video frames) across `sp_size`
|
||||
GPUs within each data-parallel group. Useful for long videos that don't fit on a
|
||||
single GPU.
|
||||
|
||||
---
|
||||
|
||||
## VSA (Variable Sparse Attention)
|
||||
|
||||
VSA progressively increases attention sparsity during training, reducing compute
|
||||
while maintaining quality:
|
||||
|
||||
```yaml
|
||||
training:
|
||||
vsa:
|
||||
sparsity: 0.9 # target sparsity level
|
||||
decay_rate: 0.03 # sparsity increment per decay interval
|
||||
decay_interval_steps: 1 # steps between sparsity increases
|
||||
```
|
||||
|
||||
The effective sparsity at step `t` is
|
||||
`min(sparsity, decay_rate * (t // decay_interval_steps))`.
|
||||
|
||||
---
|
||||
|
||||
## Extending the Framework
|
||||
|
||||
### Adding a new model
|
||||
|
||||
1. Create a new module under `fastvideo/train/models/` (e.g.,
|
||||
`fastvideo/train/models/mymodel/mymodel.py`).
|
||||
2. Subclass `ModelBase` (or `CausalModelBase` for streaming models).
|
||||
3. Implement the required methods:
|
||||
- `prepare_batch()` — convert raw dataloader output to `TrainingBatch`
|
||||
- `add_noise()` — forward-process noise addition
|
||||
- `predict_noise()` — run the transformer forward pass
|
||||
- `backward()` — backward pass with forward context restoration
|
||||
4. Reference it in your YAML config:
|
||||
|
||||
```yaml
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.mymodel.mymodel.MyModel
|
||||
init_from: my-org/my-model
|
||||
trainable: true
|
||||
```
|
||||
|
||||
### Adding a new training method
|
||||
|
||||
1. Create a new module under `fastvideo/train/methods/`.
|
||||
2. Subclass `TrainingMethod`.
|
||||
3. Implement the required methods:
|
||||
- `single_train_step()` — one forward pass returning losses, outputs, metrics
|
||||
- `get_optimizers()` — return optimizer list
|
||||
- `get_lr_schedulers()` — return scheduler list
|
||||
4. Reference it in your config:
|
||||
|
||||
```yaml
|
||||
method:
|
||||
_target_: fastvideo.train.methods.my_method.MyMethod
|
||||
my_param: 42
|
||||
```
|
||||
|
||||
Method-specific parameters are accessible via `self.method_config` (a plain
|
||||
dict).
|
||||
|
||||
### Adding a new callback
|
||||
|
||||
1. Create a new module under `fastvideo/train/callbacks/`.
|
||||
2. Subclass `Callback`.
|
||||
3. Override the hooks you need: `on_train_start`, `on_training_step_end`,
|
||||
`on_before_optimizer_step`, etc.
|
||||
4. Optionally implement `state_dict()` / `load_state_dict()` for checkpoint
|
||||
persistence.
|
||||
5. Add it to your config:
|
||||
|
||||
```yaml
|
||||
callbacks:
|
||||
my_callback:
|
||||
_target_: fastvideo.train.callbacks.my_callback.MyCallback
|
||||
my_param: 42
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## File Structure
|
||||
|
||||
```
|
||||
fastvideo/train/
|
||||
entrypoint/
|
||||
train.py # CLI entrypoint (torchrun)
|
||||
trainer.py # Training loop orchestrator
|
||||
models/
|
||||
base.py # ModelBase, CausalModelBase ABCs
|
||||
wan/
|
||||
wan.py # Wan 2.1 T2V model
|
||||
wan_causal.py # Wan causal (streaming) model
|
||||
methods/
|
||||
base.py # TrainingMethod ABC
|
||||
distribution_matching/
|
||||
dmd2.py # DMD2 distillation
|
||||
self_forcing.py # Self-Forcing (causal DMD)
|
||||
fine_tuning/
|
||||
finetune.py # Supervised fine-tuning
|
||||
dfsft.py # Diffusion-forcing SFT
|
||||
callbacks/
|
||||
callback.py # Callback ABC and CallbackDict
|
||||
grad_clip.py # Gradient clipping + norm logging
|
||||
ema.py # EMA weight averaging
|
||||
validation.py # Periodic inference validation
|
||||
utils/
|
||||
config.py # YAML parser -> RunConfig
|
||||
training_config.py # Typed config dataclasses
|
||||
builder.py # Model/method instantiation
|
||||
optimizer.py # Optimizer/scheduler construction
|
||||
checkpoint.py # DCP save/resume
|
||||
dataloader.py # Dataset/dataloader construction
|
||||
tracking.py # W&B tracker
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Related Docs
|
||||
|
||||
- [Training Architecture](../design/training_architecture.md) — design
|
||||
rationale, model/method abstractions, and open questions.
|
||||
- [Training Overview](overview.md) — data requirements and preprocessing.
|
||||
- [Data Preprocessing](data_preprocess.md) — how to prepare datasets.
|
||||
- [Config Reference](../../examples/train/configs/example.yaml) — fully-commented
|
||||
YAML config with all fields and defaults.
|
||||
@@ -1,77 +0,0 @@
|
||||
# Debugging
|
||||
|
||||
This page collects practical debugging steps for FastVideo inference issues.
|
||||
|
||||
## Collect Environment Info
|
||||
|
||||
From the repository root, run:
|
||||
|
||||
```bash
|
||||
python collect_env.py
|
||||
```
|
||||
|
||||
Attach the output when filing a GitHub issue.
|
||||
|
||||
## Increase Logging
|
||||
|
||||
FastVideo logging level is controlled by environment variables:
|
||||
|
||||
```bash
|
||||
FASTVIDEO_LOGGING_LEVEL=DEBUG \
|
||||
FASTVIDEO_STAGE_LOGGING=1 \
|
||||
python your_script.py
|
||||
```
|
||||
|
||||
Useful variables:
|
||||
|
||||
- `FASTVIDEO_LOGGING_LEVEL`: `DEBUG`, `INFO`, `WARNING`, `ERROR`
|
||||
- `FASTVIDEO_STAGE_LOGGING`: print per-stage timings during pipeline execution
|
||||
- `FASTVIDEO_ATTENTION_BACKEND`: force an attention backend (for example
|
||||
`TORCH_SDPA` or `FLASH_ATTN`)
|
||||
|
||||
## Common Failure Modes
|
||||
|
||||
### Out-of-memory
|
||||
|
||||
Try, in order:
|
||||
|
||||
1. Reduce `height`, `width`, `num_frames`, or `num_inference_steps`.
|
||||
2. Enable offloading flags such as `dit_layerwise_offload` (single GPU) or
|
||||
`use_fsdp_inference` (multi-GPU).
|
||||
3. Enable `vae_cpu_offload`, `image_encoder_cpu_offload`, and
|
||||
`text_encoder_cpu_offload`.
|
||||
|
||||
See [Inference Offloading](../inference/offloading.md) for recommended
|
||||
combinations.
|
||||
|
||||
### Attention backend import errors
|
||||
|
||||
If forcing a backend fails, verify optional dependencies are installed:
|
||||
|
||||
- `FLASH_ATTN`: `flash-attn`
|
||||
- `VIDEO_SPARSE_ATTN`: `fastvideo-kernel`
|
||||
- `SLIDING_TILE_ATTN`: STA legacy workflow in
|
||||
`sta_do_not_delete` + `fastvideo-kernel`
|
||||
- `SAGE_ATTN` / `SAGE_ATTN_THREE`: SageAttention packages
|
||||
|
||||
As a fallback, use:
|
||||
|
||||
```bash
|
||||
export FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA
|
||||
```
|
||||
|
||||
### Configuration parsing errors
|
||||
|
||||
When using `--config`, keep keys aligned with CLI argument names (underscores or
|
||||
hyphens are both accepted). For nested config values, use nested objects
|
||||
(`vae_config`, `dit_config`) rather than dotted keys.
|
||||
|
||||
## Issue Template
|
||||
|
||||
When opening an issue, include:
|
||||
|
||||
- exact command or Python snippet,
|
||||
- model ID/path,
|
||||
- full traceback,
|
||||
- `collect_env.py` output,
|
||||
- whether the problem reproduces with `FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA`.
|
||||
@@ -1,50 +0,0 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.sample import SamplingParam
|
||||
|
||||
|
||||
def main():
|
||||
# Point this to your local diffusers model dir (or replace with a HF model ID).
|
||||
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_path,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # set True if GPU is out of memory
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
sampling_param = SamplingParam.from_pretrained(model_path)
|
||||
|
||||
# image2world example from official repo
|
||||
image_path = "assets/images/bus_terminal.jpg"
|
||||
|
||||
prompt = (
|
||||
"A nighttime city bus terminal gradually shifts from stillness to subtle movement. "
|
||||
"At first, multiple double-decker buses are parked under the glow of overhead lights, "
|
||||
"with a central bus labeled '87D' facing forward and stationary. "
|
||||
"As the video progresses, the bus in the middle moves ahead slowly, its headlights brightening the surrounding area "
|
||||
"and casting reflections onto adjacent vehicles. "
|
||||
"The motion creates space in the lineup, signaling activity within the otherwise quiet station. "
|
||||
"It then comes to a smooth stop, resuming its position in line. "
|
||||
"Overhead signage in Chinese characters remains illuminated, enhancing the vibrant, urban night scene."
|
||||
)
|
||||
|
||||
generator.generate_video(
|
||||
prompt,
|
||||
sampling_param=sampling_param,
|
||||
image_path=str(image_path),
|
||||
num_cond_frames=1,
|
||||
output_path="outputs_video/cosmos2_5_i2w.mp4",
|
||||
save_video=True,
|
||||
)
|
||||
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
||||
@@ -1,51 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.sample import SamplingParam
|
||||
|
||||
|
||||
def main():
|
||||
# Point this to your local diffusers model dir (or replace with a HF model ID).
|
||||
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_path,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # set True if GPU is out of memory
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
# Load default sampling parameters (negative_prompt, resolution, steps, etc.)
|
||||
sampling_param = SamplingParam.from_pretrained(model_path)
|
||||
|
||||
prompt = (
|
||||
"A high-definition video captures the precision of robotic welding in an industrial setting. "
|
||||
"The first frame showcases a robotic arm, equipped with a welding torch, positioned over a large metal structure. "
|
||||
"The welding process is in full swing, with bright sparks and intense light illuminating the scene, "
|
||||
"creating a vivid display of blue and white hues. "
|
||||
"A significant amount of smoke billows around the welding area, partially obscuring the view but emphasizing the heat and activity. "
|
||||
"The background reveals parts of the workshop environment, including a ventilation system and various pieces of machinery, "
|
||||
"indicating a busy and functional industrial workspace. "
|
||||
"As the video progresses, the robotic arm maintains its steady position, continuing the welding process and moving to its left. "
|
||||
"The welding torch consistently emits sparks and light, and the smoke continues to rise, diffusing slightly as it moves upward. "
|
||||
"The metal surface beneath the torch shows ongoing signs of heating and melting. "
|
||||
"The scene retains its industrial ambiance, with the welding sparks and smoke dominating the visual field, "
|
||||
"underscoring the ongoing nature of the welding operation."
|
||||
)
|
||||
|
||||
generator.generate_video(
|
||||
prompt,
|
||||
sampling_param=sampling_param,
|
||||
output_path="outputs_video/cosmos2_5_t2w.mp4",
|
||||
save_video=True,
|
||||
)
|
||||
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
||||
|
||||
|
||||
@@ -1,53 +0,0 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.sample import SamplingParam
|
||||
|
||||
|
||||
def main():
|
||||
# Point this to your local diffusers model dir (or replace with a HF model ID).
|
||||
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_path,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # set True if GPU is out of memory
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
sampling_param = SamplingParam.from_pretrained(model_path)
|
||||
|
||||
# video2world example from official repo
|
||||
video_path = "assets/videos/robot_pouring.mp4"
|
||||
|
||||
prompt = (
|
||||
"A robotic arm, primarily white with black joints and cables, is shown in a clean, modern indoor setting with a white tabletop. "
|
||||
"The arm, equipped with a gripper holding a small, light green pitcher, is positioned above a clear glass containing a reddish-brown liquid and a spoon. "
|
||||
"The robotic arm is in the process of pouring a transparent liquid into the glass. "
|
||||
"To the left of the pitcher, there is an opened jar with a similar reddish-brown substance visible through its transparent body. "
|
||||
"In the background, a vase with white flowers and a brown couch are partially visible, adding to the contemporary ambiance. "
|
||||
"The lighting is bright, casting soft shadows on the table. "
|
||||
"The robotic arm's movements are smooth and controlled, demonstrating precision in its task. "
|
||||
"As the video progresses, the robotic arm completes the pour, leaving the glass half-filled with the reddish-brown liquid. "
|
||||
"The jar remains untouched throughout the sequence, and the spoon inside the glass remains stationary. "
|
||||
"The other robotic arm on the right side also stays stationary throughout the video. "
|
||||
"The final frame captures the robotic arm with the pitcher finishing the pour, with the glass now filled to a higher level, while the pitcher is slightly tilted but still held securely by the gripper."
|
||||
)
|
||||
|
||||
generator.generate_video(
|
||||
prompt,
|
||||
sampling_param=sampling_param,
|
||||
video_path=str(video_path),
|
||||
num_cond_frames=1,
|
||||
output_path="outputs_video/cosmos2_5_v2w.mp4",
|
||||
save_video=True,
|
||||
)
|
||||
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
||||
@@ -1,119 +0,0 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""
|
||||
Basic inference script for HunyuanGameCraft video generation.
|
||||
|
||||
HunyuanGameCraft generates game-like videos with camera/action control.
|
||||
It takes an optional image input and generates video with camera motion
|
||||
based on simple action commands (forward, left, right, backward, rotations).
|
||||
|
||||
Available actions:
|
||||
- forward (w): Move camera forward
|
||||
- backward (s): Move camera backward
|
||||
- left (a): Move camera left (strafe)
|
||||
- right (d): Move camera right (strafe)
|
||||
- left_rot: Rotate camera left (pan)
|
||||
- right_rot: Rotate camera right (pan)
|
||||
- up_rot: Rotate camera up (tilt)
|
||||
- down_rot: Rotate camera down (tilt)
|
||||
|
||||
T2V vs I2V:
|
||||
- Default: I2V (uses a default reference image). Set GAMECRAFT_I2V_IMAGE to a
|
||||
URL or path to use a different image.
|
||||
- T2V only (no reference image): run with GAMECRAFT_I2V_IMAGE= (empty).
|
||||
"""
|
||||
import os
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.models.camera import create_camera_trajectory
|
||||
|
||||
# Model configuration (use GAMECRAFT_MODEL_PATH for local weights)
|
||||
MODEL_PATH = os.environ.get("GAMECRAFT_MODEL_PATH", "FastVideo/HunyuanGameCraft-Diffusers")
|
||||
|
||||
# Default prompts for demo
|
||||
DEFAULT_PROMPTS = {
|
||||
"village": "A charming medieval village with cobblestone streets, thatched-roof houses, and vibrant flower gardens under a bright blue sky.",
|
||||
"temple": "A majestic ancient temple stands under a clear blue sky, its grandeur highlighted by towering Doric columns and intricate architectural details.",
|
||||
"forest": "A lush green forest with tall trees, dappled sunlight filtering through the leaves, and a winding dirt path.",
|
||||
"beach": "A tropical beach with crystal clear turquoise water, white sand, and palm trees swaying in the breeze.",
|
||||
}
|
||||
|
||||
# I2V: default reference image (URL). Can override with a local path.
|
||||
DEFAULT_I2V_IMAGE_URL = (
|
||||
"https://huggingface.co/datasets/huggingface/documentation-images/"
|
||||
"resolve/main/diffusers/astronaut.jpg"
|
||||
)
|
||||
DEFAULT_I2V_PROMPT = (
|
||||
"An astronaut hatching from an egg, on the surface of the moon, "
|
||||
"the darkness and depth of space realised in the background."
|
||||
)
|
||||
|
||||
OUTPUT_PATH = "video_samples_gamecraft"
|
||||
|
||||
|
||||
def main():
|
||||
# Initialize generator
|
||||
# FastVideo will automatically download weights from HuggingFace
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
MODEL_PATH,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=True,
|
||||
dit_cpu_offload=True,
|
||||
vae_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
# Video parameters
|
||||
height = 704
|
||||
width = 1280
|
||||
num_frames = 33
|
||||
action = "forward"
|
||||
action_speed = 0.2
|
||||
|
||||
# Create camera trajectory (Plücker coordinates)
|
||||
camera_states = create_camera_trajectory(
|
||||
action=action,
|
||||
height=height,
|
||||
width=width,
|
||||
num_frames=num_frames,
|
||||
action_speed=action_speed,
|
||||
dtype=torch.bfloat16,
|
||||
)
|
||||
print(f"Camera states shape: {camera_states.shape}")
|
||||
|
||||
# I2V vs T2V: unset GAMECRAFT_I2V_IMAGE -> I2V (default image). Set to "" -> T2V.
|
||||
env_image = os.environ.get("GAMECRAFT_I2V_IMAGE")
|
||||
if env_image is None:
|
||||
image_path = DEFAULT_I2V_IMAGE_URL # default: I2V
|
||||
elif env_image.strip() == "":
|
||||
image_path = None # T2V
|
||||
else:
|
||||
image_path = env_image.strip() # I2V with given URL/path
|
||||
|
||||
is_i2v = image_path is not None
|
||||
prompt = DEFAULT_I2V_PROMPT if is_i2v else DEFAULT_PROMPTS["temple"]
|
||||
print(f"Mode: {'I2V' if is_i2v else 'T2V'}, prompt: {prompt[:60]}...")
|
||||
|
||||
gen_kw = dict(
|
||||
prompt=prompt,
|
||||
negative_prompt="",
|
||||
camera_states=camera_states,
|
||||
height=height,
|
||||
width=width,
|
||||
num_frames=num_frames,
|
||||
num_inference_steps=50,
|
||||
guidance_scale=6.0,
|
||||
seed=42,
|
||||
fps=24,
|
||||
output_path=OUTPUT_PATH,
|
||||
save_video=True,
|
||||
)
|
||||
if is_i2v:
|
||||
gen_kw["image_path"] = image_path
|
||||
generator.generate_video(**gen_kw)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -39,4 +39,4 @@ def main():
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
main()
|
||||
|
||||
@@ -1,43 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
import json
|
||||
# from fastvideo.configs.sample import SamplingParam
|
||||
|
||||
OUTPUT_PATH = "video_samples_hy15_1080p"
|
||||
def main():
|
||||
# FastVideo will automatically use the optimal default arguments for the
|
||||
# model.
|
||||
# If a local path is provided, FastVideo will make a best effort
|
||||
# attempt to identify the optimal arguments.
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"weizhou03/HunyuanVideo-1.5-Diffusers-1080p-2SR", # 480p -> 720p -> 1080p
|
||||
# or "weizhou03/HunyuanVideo-1.5-Diffusers-1080p" # 720p -> 1080p
|
||||
# FastVideo will automatically handle distributed setup
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # set to True if GPU is out of memory
|
||||
dit_cpu_offload=True,
|
||||
vae_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True, # set to false if low CPU RAM or hit obscure "CUDA error: Invalid argument"
|
||||
# image_encoder_cpu_offload=False,
|
||||
)
|
||||
|
||||
prompt = (
|
||||
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
|
||||
"wide with interest. The playful yet serene atmosphere is complemented by soft "
|
||||
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
|
||||
)
|
||||
|
||||
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, save_video=True, negative_prompt="")
|
||||
|
||||
prompt2 = (
|
||||
"A majestic lion strides across the golden savanna, its powerful frame "
|
||||
"glistening under the warm afternoon sun. The tall grass ripples gently in "
|
||||
"the breeze, enhancing the lion's commanding presence. The tone is vibrant, "
|
||||
"embodying the raw energy of the wild. Low angle, steady tracking shot, "
|
||||
"cinematic.")
|
||||
|
||||
video2 = generator.generate_video(prompt2, output_path=OUTPUT_PATH, save_video=True, negative_prompt="")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,66 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.models.dits.hyworld.resolution_utils import get_resolution_from_image
|
||||
|
||||
# Default prompt from HY-WorldPlay run.sh
|
||||
DEFAULT_PROMPT = 'A paved pathway leads towards a stone arch bridge spanning a calm body of water. Lush green trees and foliage line the path and the far bank of the water. A traditional-style pavilion with a tiered, reddish-brown roof sits on the far shore. The water reflects the surrounding greenery and the sky. The scene is bathed in soft, natural light, creating a tranquil and serene atmosphere. The pathway is composed of large, rectangular stones, and the bridge is constructed of light gray stone. The overall composition emphasizes the peaceful and harmonious nature of the landscape.'
|
||||
DEFAULT_IMAGE = 'https://raw.githubusercontent.com/Tencent-Hunyuan/HY-WorldPlay/main/assets/img/test.png'
|
||||
|
||||
OUTPUT_PATH = "video_samples_hyworld"
|
||||
def main():
|
||||
import argparse
|
||||
|
||||
# pose: (a, w, s, d) - (15, 31)
|
||||
# num_frames: (61, 125)
|
||||
parser = argparse.ArgumentParser(description="HYWorld video generation with FastVideo")
|
||||
parser.add_argument("--prompt", type=str, default=DEFAULT_PROMPT, help="Text prompt for video generation")
|
||||
parser.add_argument("--image", type=str, default=DEFAULT_IMAGE, help="Path or URL to input image")
|
||||
parser.add_argument("--pose", type=str, default='w-31', help="Pose string (e.g., 'a-31', 'w-31', 's-31', 'd-31')")
|
||||
parser.add_argument("--output_path", type=str, default=OUTPUT_PATH, help="Output video path")
|
||||
parser.add_argument("--num-frames", type=int, default=125, help="Number of frames")
|
||||
parser.add_argument("--seed", type=int, default=1, help="Random seed")
|
||||
parser.add_argument("--resolution", type=str, default="480p", help="Only support 480p for now")
|
||||
args = parser.parse_args()
|
||||
|
||||
# Automatically determine resolution from input image
|
||||
HEIGHT, WIDTH = get_resolution_from_image(args.image, args.resolution)
|
||||
print(f"Image: {args.image}")
|
||||
print(f"Pose: {args.pose}")
|
||||
print(f"Resolution: {HEIGHT}x{WIDTH} (from {args.resolution} buckets)")
|
||||
print(f"Num frames: {args.num_frames}")
|
||||
print(f"Output path: {args.output_path}")
|
||||
|
||||
# Initialize generator
|
||||
print("\nInitializing VideoGenerator for HYWorld...")
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"FastVideo/HY-WorldPlay-Bidirectional-Diffusers",
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=True,
|
||||
dit_cpu_offload=True,
|
||||
vae_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
image_encoder_cpu_offload=True,
|
||||
)
|
||||
|
||||
# Generate video
|
||||
# The pose string is automatically converted to camera matrices by the pipeline
|
||||
print("\nGenerating video...")
|
||||
generator.generate_video(
|
||||
prompt=args.prompt,
|
||||
image_path=args.image,
|
||||
pose=args.pose, # Camera trajectory control
|
||||
output_path=args.output_path,
|
||||
save_video=True,
|
||||
negative_prompt="",
|
||||
num_frames=args.num_frames,
|
||||
fps=24,
|
||||
height=HEIGHT,
|
||||
width=WIDTH,
|
||||
seed=args.seed,
|
||||
)
|
||||
|
||||
print(f"\nVideo saved to: {args.output_path}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,49 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.models.dits.lingbotworld.cam_utils import prepare_camera_embedding
|
||||
|
||||
# from fastvideo.configs.sample import SamplingParam
|
||||
OUTPUT_PATH = "video_samples_lingbotworld"
|
||||
def main():
|
||||
# FastVideo will automatically use the optimal default arguments for the
|
||||
# model.
|
||||
# If a local path is provided, FastVideo will make a best effort
|
||||
# attempt to identify the optimal arguments.
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"FastVideo/LingBot-World-Base-Cam-Diffusers",
|
||||
# FastVideo will automatically handle distributed setup
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False, # set to True if GPU is out of memory
|
||||
dit_cpu_offload=True, # DiT need to be offloaded for MoE
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
# Set pin_cpu_memory to false if CPU RAM is limited and there're no frequent CPU-GPU transfer
|
||||
pin_cpu_memory=True,
|
||||
# image_encoder_cpu_offload=False,
|
||||
)
|
||||
|
||||
num_frames = 81
|
||||
prompt = "The video presents a soaring journey through a fantasy jungle. The wind whips past the rider's blue hands gripping the reins, causing the leather straps to vibrate. The ancient gothic castle approaches steadily, its stone details becoming clearer against the backdrop of floating islands and distant waterfalls."
|
||||
image_path = "https://raw.githubusercontent.com/Robbyant/lingbot-world/main/examples/00/image.jpg"
|
||||
action_path = "examples/inference/basic/lingbotworld_examples/00"
|
||||
c2ws_plucker_emb, num_frames = prepare_camera_embedding(
|
||||
action_path=action_path,
|
||||
num_frames=num_frames,
|
||||
height=480,
|
||||
width=832,
|
||||
spatial_scale=8,
|
||||
)
|
||||
|
||||
generator.generate_video(
|
||||
prompt,
|
||||
image_path=image_path,
|
||||
output_path=OUTPUT_PATH,
|
||||
save_video=True,
|
||||
num_frames=num_frames,
|
||||
height=480,
|
||||
width=832,
|
||||
c2ws_plucker_emb=c2ws_plucker_emb,
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,39 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
|
||||
PROMPT = (
|
||||
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
|
||||
"of a woman and a man in their 30s, facing each other with serious "
|
||||
"expressions. The woman, emotional and dramatic, says softly, \"That's "
|
||||
"it... Dad's lost it. And we've lost Dad.\" The man exhales, slightly "
|
||||
"annoyed: \"Stop being so dramatic, Jess.\" A beat. He glances aside, "
|
||||
"then mutters defensively, \"He's just having fun.\" The camera slowly "
|
||||
"pans right, revealing the grandfather in the garden wearing enormous "
|
||||
"butterfly wings, waving his arms in the air like he's trying to take "
|
||||
"off. He shouts, \"Wheeeew!\" as he flaps his wings with full commitment. "
|
||||
"The woman covers her face, on the verge of tears. The tone is deadpan, "
|
||||
"absurd, and quietly tragic."
|
||||
)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
# Uses FastVideo default sampling settings for LTX2 base.
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Davids048/LTX2-Base-Diffusers",
|
||||
num_gpus=1,
|
||||
)
|
||||
|
||||
output_path = "outputs_video/ltx2_basic/output_ltx2_base_t2v_1088_1920_1.1.mp4"
|
||||
generator.generate_video(
|
||||
prompt=PROMPT,
|
||||
output_path=output_path,
|
||||
save_video=True,
|
||||
num_frames=121,
|
||||
height=1088,
|
||||
width=1920,
|
||||
)
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,35 +0,0 @@
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
PROMPT = (
|
||||
"A warm sunny backyard. The camera starts in a tight cinematic close-up "
|
||||
"of a woman and a man in their 30s, facing each other with serious "
|
||||
"expressions. The woman, emotional and dramatic, says softly, \"That's "
|
||||
"it... Dad's lost it. And we've lost Dad.\" The man exhales, slightly "
|
||||
"annoyed: \"Stop being so dramatic, Jess.\" A beat. He glances aside, "
|
||||
"then mutters defensively, \"He's just having fun.\" The camera slowly "
|
||||
"pans right, revealing the grandfather in the garden wearing enormous "
|
||||
"butterfly wings, waving his arms in the air like he's trying to take "
|
||||
"off. He shouts, \"Wheeeew!\" as he flaps his wings with full commitment. "
|
||||
"The woman covers her face, on the verge of tears. The tone is deadpan, "
|
||||
"absurd, and quietly tragic."
|
||||
)
|
||||
import os
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
|
||||
|
||||
def main() -> None:
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"FastVideo/LTX2-Distilled-Diffusers",
|
||||
num_gpus=4,
|
||||
)
|
||||
|
||||
output_path = "outputs_video/ltx2_basic/output_ltx2_distilled_t2v.mp4"
|
||||
generator.generate_video(
|
||||
prompt=PROMPT,
|
||||
output_path=output_path,
|
||||
save_video=True,
|
||||
)
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,5 +1,6 @@
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.models.dits.matrixgame.utils import create_action_presets
|
||||
from fastvideo.configs.pipelines.wan import MatrixGameI2V480PConfig
|
||||
from fastvideo.models.dits.matrix_game.utils import create_action_presets
|
||||
|
||||
import torch
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
from fastvideo.entrypoints.streaming_generator import StreamingVideoGenerator
|
||||
from fastvideo.models.dits.matrixgame.utils import get_current_action_async, expand_action_to_frames
|
||||
from fastvideo.models.dits.matrix_game.utils import get_current_action_async, expand_action_to_frames
|
||||
|
||||
import torch
|
||||
import asyncio
|
||||
|
||||
@@ -1,137 +0,0 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import os
|
||||
import re
|
||||
from typing import List
|
||||
|
||||
|
||||
DEFAULT_PROMPTS = [
|
||||
"a photo of a cat",
|
||||
"a cinematic photo of a red panda wearing a tiny backpack, standing on a rainy neon-lit street at night, shallow depth of field, sharp focus, 35mm, bokeh",
|
||||
]
|
||||
|
||||
|
||||
def _safe_filename(text: str, max_len: int = 100) -> str:
|
||||
"""
|
||||
Make a stable, filesystem-friendly filename base.
|
||||
VideoGenerator uses prompt[:100].strip() internally, so we mirror that,
|
||||
but also remove path separators and other problematic characters.
|
||||
"""
|
||||
s = text[:max_len].strip()
|
||||
s = s.replace(os.sep, "_")
|
||||
if os.altsep:
|
||||
s = s.replace(os.altsep, "_")
|
||||
s = re.sub(r"\s+", " ", s)
|
||||
s = re.sub(r"[^A-Za-z0-9 .,_-]", "_", s)
|
||||
s = s.strip(" .")
|
||||
return s or "prompt"
|
||||
|
||||
|
||||
def _remove_existing_outputs(out_dir: str, filename_base: str) -> None:
|
||||
"""
|
||||
Ensure deterministic naming by deleting any existing outputs that would
|
||||
cause VideoGenerator to append suffixes like _1, _2, etc.
|
||||
"""
|
||||
if not os.path.isdir(out_dir):
|
||||
return
|
||||
|
||||
pattern = re.compile(rf"^{re.escape(filename_base)}(_\d+)?\.(mp4|png)$")
|
||||
for fn in os.listdir(out_dir):
|
||||
if pattern.match(fn):
|
||||
try:
|
||||
os.remove(os.path.join(out_dir, fn))
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
p = argparse.ArgumentParser(description="Run SD3.5 Medium text-to-image with FastVideo VideoGenerator.")
|
||||
p.add_argument("--model-path", default="stabilityai/stable-diffusion-3.5-medium", help="Path to local diffusers-format SD3.5 weights directory.")
|
||||
p.add_argument(
|
||||
"--out-dir",
|
||||
"--outdir",
|
||||
default="outputs/sd35/samples",
|
||||
help="Output directory for generated mp4 files.",
|
||||
)
|
||||
p.add_argument(
|
||||
"--prompt",
|
||||
action="append",
|
||||
default=None,
|
||||
help="Prompt text. Repeat --prompt multiple times to generate multiple samples.",
|
||||
)
|
||||
p.add_argument("--negative", default="lowres, blurry, jpeg artifacts, watermark, text", help="Negative prompt.")
|
||||
p.add_argument(
|
||||
"--backend",
|
||||
default=None,
|
||||
help="Set FASTVIDEO_ATTENTION_BACKEND (e.g. TORCH_SDPA). If omitted, respects the existing env var.",
|
||||
)
|
||||
p.add_argument("--seed", type=int, default=42, help="Base seed. Each prompt uses seed + prompt_idx.")
|
||||
p.add_argument("--height", type=int, default=768, help="Output height.")
|
||||
p.add_argument("--width", type=int, default=768, help="Output width.")
|
||||
p.add_argument("--steps", type=int, default=28, help="Number of inference steps.")
|
||||
p.add_argument("--guidance", type=float, default=6.0, help="Guidance scale.")
|
||||
p.add_argument("--num-gpus", type=int, default=1, help="Number of GPUs to use.")
|
||||
return p.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
prompts: List[str] = args.prompt if args.prompt else DEFAULT_PROMPTS
|
||||
|
||||
if args.backend:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
os.makedirs(args.out_dir, exist_ok=True)
|
||||
|
||||
init_kwargs = {
|
||||
"num_gpus": args.num_gpus,
|
||||
"workload_type": "t2i",
|
||||
"sp_size": 1,
|
||||
"tp_size": 1,
|
||||
"dit_cpu_offload": False,
|
||||
"dit_layerwise_offload": False,
|
||||
"text_encoder_cpu_offload": False,
|
||||
"vae_cpu_offload": False,
|
||||
"image_encoder_cpu_offload": False,
|
||||
"pin_cpu_memory": False,
|
||||
"use_fsdp_inference": False,
|
||||
}
|
||||
|
||||
generator = VideoGenerator.from_pretrained(model_path=args.model_path, **init_kwargs)
|
||||
try:
|
||||
for i, prompt in enumerate(prompts):
|
||||
seed = args.seed + i
|
||||
|
||||
filename_base = f"sd35_{i:02d}_seed{seed}_{_safe_filename(prompt, max_len=80)}"
|
||||
_remove_existing_outputs(args.out_dir, filename_base)
|
||||
|
||||
output_path = os.path.join(args.out_dir, f"{filename_base}.png")
|
||||
print(f"[sd35] prompt_idx={i} seed={seed} output_path={output_path}")
|
||||
|
||||
generation_kwargs = {
|
||||
"output_path": output_path,
|
||||
"height": args.height,
|
||||
"width": args.width,
|
||||
"num_frames": 1,
|
||||
"fps": 1,
|
||||
"num_inference_steps": args.steps,
|
||||
"guidance_scale": args.guidance,
|
||||
"seed": seed,
|
||||
"negative_prompt": args.negative,
|
||||
"save_video": True,
|
||||
}
|
||||
|
||||
generator.generate_video(prompt, **generation_kwargs)
|
||||
|
||||
print(f"[sd35] done. outputs written to: {args.out_dir}")
|
||||
finally:
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||