Compare commits

..
Author SHA1 Message Date
SolitaryThinker d1a1b52f78 [docs]: add release skill for future version bumps 2026-06-04 14:29:35 -07:00
537 changed files with 537 additions and 60571 deletions
-207
View File
@@ -1,207 +0,0 @@
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
Inferix, and FlashDrive as techniques integratable into the graph runtime.
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
corrections in; it is the corrected, authoritative synthesis.
---
## 1. Executive summary
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
is the exact MoT serving property the omni cards in this repo already express.
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
flywheel driving the *same* serving Loop).
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
already has the seam M\* only gestures at.**
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
### Verified P0 correctness findings (spot-checked by hand)
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
disaggregated runners execute *different node sets*. ✅ confirmed.
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
---
## 2. Dimension table (corrected)
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|---|---|---|---|---|---|---|---|
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
---
## 3. P0/P1 deep-dives (sequenced)
```
PR-1 (P0) min-components ──┐
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
PR-3 (P1) output_determinism (independent)
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
```
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
already exist and are tested).
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
admission per (class,label).
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
---
## 4. What v2 already does ≥ M\* — do NOT regress
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
must pass it.
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
interceptor) — M\* describes extensibility; v2 has the mechanism.
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
self-contained runtime.
---
## 5. Dropped / merged / deferred (and why)
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
(per-step capture ships in 14 cards, not just wan21).
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
---
## 6. The first integration test, if/when multi-GPU placement work starts
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
placement work — it is the one place this repo already has real multi-rank composite-model execution.
---
## Anchor files for P0/P1
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
`v2/program/workflow.py` (registry-driven delivery).
-331
View File
@@ -1,331 +0,0 @@
# Performance Dashboard Memory
Date: 2026-06-16
Branch: `ci/dashboard`
## Purpose
This branch adds a local live dashboard for FastVideo performance benchmark
history. It is intended for maintainer/operator use: inspect latest benchmark
status, compare current values with recent baseline context, and view trends
from the Hugging Face performance-tracking dataset.
The dashboard is a FastAPI + React app. It is separate from the existing
Svelte `ui/` app.
## Main Files Added Or Changed
Backend:
- `fastvideo/performance_dashboard/__init__.py`
- `fastvideo/performance_dashboard/__main__.py`
- `fastvideo/performance_dashboard/api.py`
- `fastvideo/performance_dashboard/metrics.py`
- `fastvideo/performance_dashboard/service.py`
Frontend:
- `performance_dashboard/frontend/package.json`
- `performance_dashboard/frontend/package-lock.json`
- `performance_dashboard/frontend/tsconfig.json`
- `performance_dashboard/frontend/vite.config.ts`
- `performance_dashboard/frontend/index.html`
- `performance_dashboard/frontend/scripts/build.mjs`
- `performance_dashboard/frontend/src/main.tsx`
- `performance_dashboard/frontend/src/api.ts`
- `performance_dashboard/frontend/src/App.tsx`
- `performance_dashboard/frontend/src/styles.css`
Docs/tests:
- `performance_dashboard/README.md`
- `docs/contributing/performance_benchmarks.md`
- `fastvideo/tests/performance/test_dashboard_service.py`
- `fastvideo/tests/performance/test_dashboard_api.py`
Shared HF utility change:
- `fastvideo/tests/performance/hf_store.py`
## Data Source
The source of truth remains the Hugging Face dataset repo used by existing
performance CI:
```text
HF_REPO_ID=FastVideo/performance-tracking
```
The dataset stores normalized JSON records emitted by
`fastvideo/tests/performance/compare_baseline.py`. The current v1 normalized
schema includes:
- `model_id`
- `timestamp`
- `commit_sha`
- `gpu_type`
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
- `success`
Records are grouped by `(model_id, gpu_type)` for v1 dashboard behavior.
## Local Cache
The backend syncs the HF dataset to a local cache directory:
```text
PERFORMANCE_TRACKING_ROOT=/tmp/fastvideo-perf-dashboard
```
If `PERFORMANCE_TRACKING_ROOT` is not set, the dashboard defaults to:
```text
/tmp/fastvideo-perf-dashboard
```
The sync is performed through the existing helper:
```python
fastvideo.tests.performance.hf_store.sync_from_hf(...)
```
The dashboard then loads JSON files from the local cache through:
```python
fastvideo.tests.performance.hf_store.load_records(...)
```
## Authentication
Originally `hf_store.py` only read `HF_API_KEY`. This caused local dashboard
runs to fail when users had standard Hugging Face token variables set.
`hf_store.py` now resolves tokens from the first available variable in:
```text
HF_API_KEY
HUGGINGFACE_HUB_TOKEN
HF_TOKEN
```
For local use:
```bash
export HF_TOKEN=hf_...
```
If the HF repo is private or gated, the token must have dataset read access.
## Backend API
The FastAPI app is created by:
```python
fastvideo.performance_dashboard.api:create_app
```
The module-level app is:
```python
fastvideo.performance_dashboard.api:app
```
Endpoints:
- `GET /api/performance/health`
- `POST /api/performance/refresh`
- `GET /api/performance/records?days=90`
- `GET /api/performance/summary?days=90`
- `GET /api/performance/trends?days=90`
`POST /api/performance/refresh` forces a fresh HF sync.
## Status Semantics
Important: the dashboard intentionally separates stored CI status from
recomputed context.
Stored status:
- Comes directly from the latest JSON record's `success` field.
- This is what the dashboard displays as `Stored Status`.
- This is the primary latest status.
Recomputed status:
- Calculated locally from cached records for explanatory context.
- Uses the latest record's metric values compared to the median of the latest
five previous successful records in the same `(model_id, gpu_type)` group.
- Displayed separately as `Recomputed`.
- Does not override the stored JSON `success` status.
This distinction was added after observing that recomputing pass/fail from the
local cache can disagree with the status originally uploaded by CI.
## Time Window Behavior
The default dashboard time window is 90 days.
The selected `days` value affects:
- trend charts
- record browsing/filtering
The selected `days` value does not affect:
- latest stored status
- latest summary baseline context
Reason: latest status should not change when users widen or narrow the trend
window. The API keeps `days` on `/summary` only for shared frontend filter
state, but summary loading uses all cached records.
This fixed a bug where changing from roughly 35 days to 42 days could change
the latest status from pass to fail because older records entered the local
baseline window.
## Metric Logic
Dashboard metric definitions live in:
```text
fastvideo/performance_dashboard/metrics.py
```
Tracked metrics:
- `latency` lower is better
- `throughput` higher is better
- `memory` lower is better
- `text_encoder_time_s` lower is better
- `dit_time_s` lower is better
- `vae_decode_time_s` lower is better
Baseline context uses the median of up to five previous successful records for
the same `(model_id, gpu_type)`.
## Frontend Behavior
The React app:
- fetches `/api/performance/summary`
- fetches `/api/performance/trends`
- displays summary cards
- displays latest rows by model/GPU
- displays native SVG trend charts
- has model/GPU/day filters
- includes a refresh button
- auto-refreshes every five minutes
The UI is implemented without a charting library. Trend charts are native SVG
in `performance_dashboard/frontend/src/App.tsx`.
The production frontend build uses `esbuild` through
`performance_dashboard/frontend/scripts/build.mjs`. Vite is still used for the
dev server and `/api` proxy.
Why esbuild for production build:
- Vite/Rollup hit a local macOS native optional dependency code-signing issue
in this environment.
- Direct esbuild worked reliably and is sufficient for this small dashboard.
## Static Serving
After frontend build, the FastAPI server serves:
- static JS/CSS from `performance_dashboard/frontend/dist/assets`
- `performance_dashboard/frontend/dist/index.html` for the dashboard page
This allows a single local port to serve both the API and UI.
## Local Run Workflow
Build frontend:
```bash
cd performance_dashboard/frontend
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm install
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
```
Run dashboard:
```bash
export HF_TOKEN=hf_...
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
```
Open locally:
```text
http://127.0.0.1:8000
```
## ngrok Workflow
`python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000`
starts the actual local dashboard server.
`ngrok http 8000` does not start the dashboard. It exposes the already-running
local server through a temporary public URL.
Typical flow:
```bash
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
ngrok http 8000
```
Use the HTTPS URL printed by ngrok to view the dashboard remotely.
## Verification Commands
Backend tests:
```bash
conda run -n fastvideo python -m pytest \
fastvideo/tests/performance/test_dashboard_service.py \
fastvideo/tests/performance/test_dashboard_api.py \
-q
```
Expected after latest changes:
```text
8 passed
```
Frontend build:
```bash
cd performance_dashboard/frontend
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
```
Expected:
```text
tsc && node scripts/build.mjs
```
with exit code 0.
## Known Notes
- `performance_dashboard/frontend/node_modules/` and
`performance_dashboard/frontend/dist/` are ignored by git.
- `npm install` reported two high-severity audit findings in dependency tree.
`npm audit fix --force` was not run because it can introduce breaking
dependency upgrades.
- Existing `fastvideo` package imports may emit platform warnings such as NPU
or macOS torch distributed messages. These are not dashboard-specific errors.
-32
View File
@@ -1,32 +0,0 @@
---
name: add-reward-model
description: Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.
---
# Add Reward Model
Use for reward models consumed by RL methods.
## Placement
- Put reusable reward code under `fastvideo/train/methods/rl/rewards/`.
- Expose public builders from `fastvideo/train/methods/rl/rewards/__init__.py`.
- Keep method-specific aggregation or advantage logic out of reward classes.
## Media Inputs
- Reward callables receive decoded media tensors.
- Accept single-frame tensors as `[B, C, H, W]` and multi-frame tensors as `[B, C, T, H, W]` when practical.
- Frame selection is reward-specific. Frame scorers such as PickScore and CLIPScore should explicitly select frame `0`; temporal rewards should inspect whichever frames they need.
- Return one scalar reward per prompt/sample.
## Attribution
- If code is ported or closely adapted from another repo, add a short comment or docstring naming the source file/function.
- Preserve SPDX headers used by FastVideo files.
## Tests
- Unit-test tensor layout handling without loading large reward checkpoints.
- Allow fake scorer injection for multi-reward tests.
- Test weighted reward aggregation and metric keys.
-38
View File
@@ -1,38 +0,0 @@
---
name: add-rl-method
description: Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.
---
# Add RL Method
Use for new RL methods in the modular `fastvideo/train` stack.
## Required Shape
- Add the method under `fastvideo/train/methods/rl/`.
- Subclass `TrainingMethod`.
- Keep model-family logic in `ModelBase` wrappers.
- Decode generated latents through `ModelBase.decode_latents`; add that hook to the new model wrapper instead of decoding inside the RL method.
- Use `fastvideo/train/methods/rl/common/sampling.py` for generation unless the method has a documented reason to avoid sampling.
- Use `fastvideo/train/methods/rl/common/prompt_sampling.py` for reusable grouped prompt sampling patterns such as DiffusionNFT K-repeat.
- Use `fastvideo/train/methods/rl/rewards/` for reward models.
## Optimization
- Return `manages_optimization() == True` only when the method must own a nonstandard outer/inner loop.
- If using managed optimization, implement `managed_train_step(data_stream, iteration)`.
- Existing trainer callbacks, checkpointing, tracking, and validation should still work.
## Config
- Put method knobs under `method`.
- Put sampler knobs under `method.sampling`.
- Do not put scheduler or trajectory policy into model configs.
- Do not split a diffusers-style scheduler from its built-in `step()` solver in YAML; use `trajectory` only for higher-level ODE vs re-noise behavior.
- Avoid fixed timestep lists in examples unless reproducing a known baseline; prefer scheduler-generated defaults.
## Tests
- Add fake-model tests for sampler/method behavior.
- Add config parse tests for the public YAML.
- Confirm existing train methods stay on the default Trainer path.
+1 -3
View File
@@ -8,6 +8,4 @@
{"name": "decompose-pipeline-pr", "description": "Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body. Worked example: PR #1280 daVinci-MagiHuman (9.8k LOC) decomposed into 10 stacked PRs.", "path": "decompose-pipeline-pr/SKILL.md", "status": "tested", "trust": "medium"}
{"name": "reseed-performance-baseline", "description": "Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline must be advanced by replicating one reviewed shifted source result into three success=true records, or five records when explicitly requested", "path": "reseed-performance-baseline/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-model", "description": "Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.", "path": "add-model/SKILL.md", "status": "draft", "trust": "low"}
{"name": "rlhf-training-abstractions", "description": "Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.", "path": "rlhf-training-abstractions/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-rl-method", "description": "Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.", "path": "add-rl-method/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-reward-model", "description": "Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.", "path": "add-reward-model/SKILL.md", "status": "draft", "trust": "low"}
{"name": "release", "description": "Cut a new FastVideo release. Bumps the version across the three authoritative files (fastvideo/version.py, pyproject.toml, pyproject_other.toml), opens a [chore]: release PR, and documents the post-merge tag + GitHub release ritual. Triggers on requests like \"release X.Y.Z\", \"cut a release\", \"bump version to X.Y.Z\", \"publish to PyPI\".", "path": "release/SKILL.md", "status": "draft", "trust": "low"}
+248
View File
@@ -0,0 +1,248 @@
---
name: release
description: Cut a new FastVideo release. Bumps the version across the three authoritative files (fastvideo/version.py, pyproject.toml, pyproject_other.toml), opens a [chore]: release PR, and documents the post-merge tag + GitHub release ritual. Triggers on requests like "release X.Y.Z", "cut a release", "bump version to X.Y.Z", "publish to PyPI".
---
# FastVideo release skill
End-to-end recipe for cutting a FastVideo release. The PyPI publish is automatic — pushing a `pyproject.toml` version change to `main` triggers `.github/workflows/publish-fastvideo.yml`. Your job is to land the version bump cleanly and follow up with a git tag + GitHub Release for the changelog.
## Inputs
- `${NEW}` — the new version (e.g. `0.2.0`). Required.
- `${OLD}` — the current version. Auto-detect with: `grep -oP '__version__ = "\K[^"]+' fastvideo/version.py` from the repo root.
## When to use
Trigger phrases: "release X.Y.Z", "cut a release", "bump version to X.Y.Z", "publish to PyPI", "tag a release".
## Files to update (3 — the authoritative list)
These are the ONLY files that carry the version as a Python/package declaration:
| File | Line | Change |
|---|---|---|
| `fastvideo/version.py` | 1 | `__version__ = "${OLD}"` → `__version__ = "${NEW}"` |
| `pyproject.toml` | 7 | `version = "${OLD}"` → `version = "${NEW}"` |
| `pyproject_other.toml` | 7 | `version = "${OLD}"` → `version = "${NEW}"` |
`fastvideo/__init__.py` re-exports `__version__` from `fastvideo.version`, so no edit needed there.
## Files NOT to touch
- `apps/dreamverse/pyproject.toml` — declares `"fastvideo>=X.Y.Z"` as a floor. A new release usually still satisfies the floor; bumping it is a separate policy call (does dreamverse strictly require the new version?). Leave alone unless explicitly asked.
- `.agents/memory/**/*.md` — historical notes; the version strings in there are snapshots, not declarations.
- `examples/`, `docs/` — version mentions are illustrative; not authoritative.
- `uv.lock` — main does NOT track a `uv.lock`. Do NOT run `uv lock` as part of a release.
## Workflow
### 1. Verify clean state
```bash
# From the primary FastVideo jj workspace
jj git fetch
OLD=$(grep -oP '__version__ = "\K[^"]+' fastvideo/version.py)
echo "current: $OLD → target: $NEW"
```
Confirm `$NEW > $OLD` follows semver. Check prior tags for the pattern:
```bash
gh release list --repo hao-ai-lab/FastVideo --limit 5
```
### 2. Create a dedicated jj workspace + bookmark
```bash
WS=/home/william5lin/FastVideo_release_${NEW//./_}
jj workspace add --name release-${NEW//./-} "$WS"
cd "$WS"
jj new main@origin -m "[chore]: release v${NEW}"
jj bookmark create chore/release-${NEW} -r @
```
### 3. Apply the 3-file bump
Use the `edit` tool or `sed -i` with exact context. Example with sed:
```bash
sed -i "s/__version__ = \"${OLD}\"/__version__ = \"${NEW}\"/" fastvideo/version.py
sed -i "0,/version = \"${OLD}\"/s//version = \"${NEW}\"/" pyproject.toml
sed -i "0,/version = \"${OLD}\"/s//version = \"${NEW}\"/" pyproject_other.toml
```
(The `0,/.../s//.../` form replaces only the FIRST match in each `pyproject*.toml`, since `${OLD}` might appear elsewhere as a constraint.)
### 4. Verify
```bash
jj diff --name-only -r @ # MUST be exactly 3 files
jj diff --stat -r @ # MUST be +3 / -3
grep -nE "${OLD//./\\.}" fastvideo/version.py pyproject.toml pyproject_other.toml
# expect NO matches in the three files
```
### 5. Lint
```bash
pre-commit run --files fastvideo/version.py pyproject.toml pyproject_other.toml
```
Must pass. Never `--no-verify`.
### 6. Describe + push
```bash
jj describe -m "[chore]: release v${NEW}
Bumps FastVideo from ${OLD} to ${NEW}.
Files updated:
fastvideo/version.py
pyproject.toml
pyproject_other.toml
Note: pushing this to main triggers .github/workflows/publish-fastvideo.yml,
which detects the pyproject.toml version change and publishes to PyPI.
Tag v${NEW} + GitHub release notes follow merge."
jj git push --bookmark chore/release-${NEW}
```
### 7. Open PR
```bash
gh pr create \
--repo hao-ai-lab/FastVideo \
--base main \
--head chore/release-${NEW} \
--title "[chore]: release v${NEW}" \
--body-file - <<EOF
## Summary
Bumps FastVideo from \`${OLD}\` to \`${NEW}\`.
## Files updated (3)
- \`fastvideo/version.py\`
- \`pyproject.toml\`
- \`pyproject_other.toml\`
## Out of scope
\`apps/dreamverse/pyproject.toml\` floor (\`fastvideo>=${OLD}\`) — \`${NEW}\` satisfies it; bumping is a separate policy call.
## After merge
\`.github/workflows/publish-fastvideo.yml\` auto-publishes to PyPI on push-to-main when \`pyproject.toml\` changes.
Manual follow-up:
- Tag the merge commit: \`git tag v${NEW} <merge-sha> && git push origin v${NEW}\`
- Create GitHub Release \`v${NEW}\` matching the prior \`Release X.Y.Z\` pattern.
EOF
```
### 8. Post-merge ritual (do AFTER the PR merges)
1. **Tag the merge commit**:
```bash
git fetch origin
MERGE_SHA=$(gh pr view <PR-NUMBER> --repo hao-ai-lab/FastVideo --json mergeCommit --jq .mergeCommit.oid)
git tag v${NEW} ${MERGE_SHA}
git push origin v${NEW}
```
2. **Confirm PyPI publish workflow ran**:
```bash
gh run list --repo hao-ai-lab/FastVideo --workflow publish-fastvideo.yml --limit 3
```
3. **Create the GitHub Release**:
```bash
gh release create v${NEW} \
--repo hao-ai-lab/FastVideo \
--title "Release ${NEW}" \
--notes "<changelog highlights — what shipped since v${OLD}>" \
--target main
```
Use `gh release view v${OLD}` to mirror tone/structure from the prior release.
4. **Cleanup**: after merge + tag + release land, tear down the workspace:
```bash
jj workspace forget release-${NEW//./-}
rm -rf "$WS"
jj bookmark delete chore/release-${NEW}
```
## Verification gates (must all pass before pushing)
- `jj diff --name-only -r @` returns exactly 3 files
- `jj diff --stat -r @` shows `+3 / -3`
- `grep -E "${OLD//./\\.}" fastvideo/version.py pyproject.toml pyproject_other.toml` returns no matches
- `pre-commit run --files <the-three>` passes
- No `uv.lock` in the change
- No source-code files touched
## Conventions (enforced)
- Commit subject: `[chore]: release v${NEW}` (under 72 chars).
- NEVER add AI co-author trailers (`Co-Authored-By: Claude`, "Generated with…", etc.).
- NEVER `--no-verify`.
- NEVER `uv lock` as part of a release — main doesn't track the lockfile.
- Tag format: `vX.Y.Z` (with leading `v`), matching prior releases.
## Why three files?
`pyproject.toml` and `pyproject_other.toml` are two co-existing project metadata files (the project ships both — the latter is a slimmer variant without dreamverse/job-runner extras). Both carry an authoritative `version = "X.Y.Z"` field and must stay in lock-step. `fastvideo/version.py` is the runtime source of truth re-exported by `fastvideo/__init__.py`.
## Publish workflow contract
`.github/workflows/publish-fastvideo.yml` triggers on `push` to `main` when `pyproject.toml` changes. It compares the new `version` field to the previous commit's `version` field and, if different, builds + publishes to PyPI. The version bump in `pyproject_other.toml` does NOT trigger the workflow (only `pyproject.toml` is in the `paths:` filter), but keeping the two in sync prevents installer surprises for users of the alternate metadata file.
## PyPI publish failure modes
The publish workflow ran on the merge commit but the PyPI upload can still fail at the OIDC trusted-publishing exchange. Always verify the workflow succeeded — do not assume "merge implies published":
```bash
gh run list --repo hao-ai-lab/FastVideo --workflow publish-fastvideo.yml --limit 3
```
Look for the run on the release merge commit. If it shows `failure`, dump the failed log:
```bash
gh run view <run-id> --repo hao-ai-lab/FastVideo --log-failed | tail -80
```
### Known failure: `invalid-publisher` (Trusted Publisher claim mismatch)
The most common failure surfaces as:
```
Trusted publishing exchange failure:
* `invalid-publisher`: valid token, but no corresponding publisher
(Publisher with matching claims was not found)
* environment: MISSING
```
This means the PyPI Trusted Publisher registered for the project expects an `environment` claim (e.g. `pypi`) that the workflow job does not set. Two recovery paths:
**A. Fix the trusted publisher + re-run the workflow** (cleaner long-term):
1. On `pypi.org/manage/project/fastvideo/settings/publishing/`, either remove the `Environment name` field from the registered publisher, OR add `environment: pypi` (matching the existing PyPI config) to the `build-publish-main` job in `.github/workflows/publish-fastvideo.yml`.
2. Re-run the failed workflow:
```bash
gh run rerun <run-id> --repo hao-ai-lab/FastVideo --failed
```
**B. Manual one-shot publish** (faster, no infra change):
```bash
git checkout <merge-sha> # the v${NEW} merge commit on main
uv build # builds sdist + wheel into dist/
uv publish --token <PYPI_TOKEN> # or: twine upload dist/*
```
PyPI is **immutable per version** — if any artifact for `${NEW}` got uploaded (sdist or wheel), you cannot re-upload it. Check before retrying:
```bash
curl -s https://pypi.org/pypi/fastvideo/${NEW}/json | python3 -c "import sys,json; d=json.load(sys.stdin); print('on pypi:', list(d['urls'][0].keys()) if d.get('urls') else 'NOT_PUBLISHED')"
```
If `NOT_PUBLISHED`, either recovery path works. If anything is already up, you have to cut a `${NEW}.postN` patch release instead.
### Tag and GitHub Release are independent
The `git tag v${NEW}` and `gh release create v${NEW}` steps are **independent of PyPI publish success**. If you created the tag + release before noticing the publish failure, that's fine — keep them; just complete the PyPI publish via path A or B above. Do NOT delete and re-create the tag, because doing so will cause confusion in dependents that pin to the tag.
@@ -1,41 +0,0 @@
---
name: rlhf-training-abstractions
description: Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.
---
# RLHF Training Abstractions
Use this skill before editing RLHF-style training code in `fastvideo/train`.
## Boundaries
- RL methods live under `fastvideo/train/methods/rl/` and own algorithm logic: reward collection, advantage computation, policy loss, KL/reference terms, and optimizer cadence.
- Rewards live under `fastvideo/train/methods/rl/rewards/` and must be reusable across RL methods.
- RL methods pass decoded media to rewards; each reward decides whether to use the first frame, sampled frames, or the full video.
- Sampling lives under `fastvideo/train/methods/rl/common/` and must use `ModelBase` primitives plus scheduler math, not model-family inference pipelines.
- Model wrappers under `fastvideo/train/models/` own model-specific forward details.
- Model wrappers also own model-specific latent decoding via `ModelBase.decode_latents`; RL methods should not reach into VAE normalization internals.
- Shared RL helpers such as K-repeat prompt sampling belong under `fastvideo/train/methods/rl/common/` when they are reusable across RL methods.
## Anti-Patterns
- Do not bind RL methods to inference pipeline classes such as `WanDMDPipeline`.
- Do not hardcode timestep lists in a method when the scheduler can generate them.
- Do not put reward-model code inside one RL method.
- Do not make existing non-RL methods use method-managed optimization unless explicitly requested.
## Sampling Policy
- Prefer YAML-configured `method.sampling` with `scheduler`, `trajectory`, `num_steps`, `timesteps`, and `sigmas`.
- Treat diffusers-style scheduler classes as owning both the timestep schedule and their `step()` update rule; avoid a separate `solver` field unless a new sampler truly implements solver math outside the scheduler object.
- Missing `timesteps` means “ask the scheduler”; explicit `timesteps` or `sigmas` are overrides.
- ODE-style trajectories should not re-noise between denoising steps.
- SDE/re-noise behavior must be explicit in config.
## Validation
- Run focused local tests for sampler config and Trainer opt-in behavior.
- Verify existing train methods still report `manages_optimization() == False`.
- Keep fixed-prompt validation helpers in `fastvideo/train/methods/rl/common/validation.py` so new RL methods can reuse sharding and captions.
- Test distributed prompt grouping helpers separately from heavyweight model loading.
- Run `pre-commit run --files <changed paths>`; respect configured excludes.
-2
View File
@@ -34,8 +34,6 @@ env
*.log
weights/
logs/
official_weights/
converted_weights/
# SSIM test outputs
fastvideo/tests/ssim/generated_videos/
-56
View File
@@ -1,56 +0,0 @@
# FastVideo — Design Philosophy
One page on *why* FastVideo is built the way it is. The full architecture, the as-built status, and the
forward roadmap live in **[`v2/README.md`](v2/README.md)** — this is the philosophy beneath it.
---
**A deployable model is a post-training artifact.** Unlike an LLM — where inference optimizes frozen weights
after the fact — a *usable* video/omni model is *created* by training: step distillation for latency, QAT for
precision, distillation + self-forcing for causal/world models. So every inference capability is a
**(recipe, runtime) pair**: the weights and the loop that produced-and-assumes them are one versioned object.
This is the source of the moat — whoever owns *both* sides of the pair owns the optimization frontier — and it
is why training and serving cannot be two systems.
**The work is loops, not `forward()`.** Denoise timesteps, AR decode, chunked rollout, VAE tiles, encoder
chunks, audio tokens, reward batches, optimizer steps, media chunks — video and omni inference is iteration. A
runtime that collapses everything to a single `forward` can't schedule, batch, cancel, stream, reserve memory
for, or capture the behavior of what actually runs. So loops are first-class, and they are **driven**: the
model describes the next step it needs, the runtime decides when and with whom it runs, the model folds the
result back. The model keeps content-adaptive control flow; the runtime keeps admission, batching, streaming,
and behavior capture. Per-request state lives in typed `LoopState`, never in module globals — so interleaving
requests through one model instance cannot smear state, by construction.
**The model is the center; everything else is a view over it.** A typed `ModelCard` owns components, loops,
the recipe, and the parity contract. Programs compose a card's loops into a task; Workflows compose cards into
pipelines; the scheduler runs the *steps* of all loops as `WorkUnit`s under one currency (predicted GPU-time,
because a bidirectional denoise step and an AR token are ~1000× apart and incommensurable in counts);
deployment places and routes; products stream artifacts. None of them define model semantics — they reference
the Model Plane. One resident instance can run many loop types on shared weights, which is what makes omni/MoT
native rather than a DAG that doubles weights.
**Correctness is a typed contract, not a hope.** Caches are correct by *key* — if a field can change output
semantics it is in the key, so reuse is partitioned, never blindly flushed. Parity between the train-forward
and the serve-forward is *measured* on a declared ladder (component → loop → behavioral → distribution →
artifact-quality), never assumed. And the non-negotiable gate is **interleave bit-parity**: N requests
interleaved at step granularity must be bit-identical to running them serially — the test the whole
loop-inversion bet lives or dies on.
**One substrate for inference, training, and RL.** The rollout forward *is* the serve forward plus capture —
same loop, same caches, same batcher, same numerics — so every serving optimization is automatically a rollout
optimization, and there is one numerics surface the ladder measures rather than a correction layer papering
over it. The engine doubles as the RL rollout engine under a strict rule: `training` consumes the engine; the
**engine never imports `training`**.
**Borrow aggressively; copy nothing as the core.** vLLM/SGLang scheduling, vLLM-Omni/SGLang-Omni omni serving,
Dynamo fleet orchestration, diffusers components, xDiT parallelism, TorchTitan mesh discipline,
verl-omni/miles RL lessons, ComfyUI workflows, Dreamverse/LiveKit sessions — each contributes a take, none is
the center. Deployment orchestration (Dynamo) sits *above* the engine, never inside it. Extensions are
versioned hook points, never monkeypatching. New frontier capabilities arrive as a card, a method, a loop, a
workflow, or a controller — **not a rewrite**.
> A model card is a (recipe, runtime) pair with a parity obligation. The model owns loop semantics; the runtime
> owns loop lifecycle. One resident instance runs many loops; one scheduler runs their steps in one currency.
> Caches are correct by key; parity is correct by test; the interleave gate is non-negotiable. Training records
> behavior on the same loops it serves. Deployment places and routes; products stream artifacts; neither defines
> the model.
+2 -4
View File
@@ -55,14 +55,12 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
# Install FastVideo Unified Kernel
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
./build.sh
EXPOSE 22
+2 -4
View File
@@ -55,14 +55,12 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
# Install FastVideo Unified Kernel
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
./build.sh
EXPOSE 22
+2 -4
View File
@@ -55,13 +55,11 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
# Install FastVideo Unified Kernel
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
./build.sh
EXPOSE 22
+2 -4
View File
@@ -55,14 +55,12 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
# Install FastVideo Unified Kernel
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
./build.sh
EXPOSE 22
@@ -41,13 +41,6 @@ when you want local Markdown/normalized-result artifacts from the comparator.
`fastvideo/tests/performance/results/`; remove stale result files if you only
want to compare the latest local run.
## Local live dashboard
For an app-style local dashboard backed by the same HF performance-tracking
records, see `performance_dashboard/README.md`. The dashboard provides a
FastAPI API plus a React UI and can be exposed with `ngrok` after building the
frontend.
## Architecture
```
@@ -108,7 +108,6 @@ surfaces:
vae_sp: generator.pipeline.preset_overrides.vae_sp
dmd_denoising_steps: generator.pipeline.preset_overrides.dmd_denoising_steps
ti2v_task: generator.pipeline.preset_overrides.ti2v_task
lucy_edit_task: generator.pipeline.preset_overrides.lucy_edit_task
boundary_ratio: generator.pipeline.preset_overrides.boundary_ratio
compatibility_only:
model_path: "Redundant with generator.model_path."
@@ -126,18 +125,9 @@ surfaces:
text_encoder_configs: "Legacy internal component config object."
preprocess_text_funcs: "Internal text preprocessing hooks."
postprocess_text_funcs: "Internal text postprocessing hooks."
scheduler_step_in_fp32: "Runtime scheduler precision toggle; not part of the public typed inference API."
pipeline_config_extensions:
preset_owned:
flux2_text_encoder_type:
sources:
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
text_encoder_out_layers:
sources:
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
conditioning_strategy:
sources:
- fastvideo.configs.pipelines.cosmos.CosmosConfig
@@ -501,8 +491,6 @@ surfaces:
inpaint_mask: request.extensions.stable_audio.inpaint_mask
internal_only:
data_type: "Derived from the request shape and not a public input."
latents: "Pre-generated diffusion latents supplied by parity/debug harnesses; not a public input."
max_sequence_length: "Model-specific text-encoder sequence cap; not part of the public typed inference API."
sampling_param_extensions: {}
-40
View File
@@ -24,7 +24,6 @@ This page describes the various options for speeding up generation times in Fast
- Video Sparse Attention: `FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN`
- Sage Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN`
- Sage Attention 3: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN_THREE`
- Attn-QAT inference (modified SageAttention3 FP4, sm_120/RTX 5090): `FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER`
- Video MoBA Attention: `FASTVIDEO_ATTENTION_BACKEND=VMOBA_ATTN`
- Sparse Linear Attention: `FASTVIDEO_ATTENTION_BACKEND=SLA_ATTN`
- SageSLA Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_SLA_ATTN`
@@ -123,45 +122,6 @@ gen.generate_video(prompt="A raccoon in sunflowers", save_video=True)
- Per-call cosine similarity vs BF16: ~0.99 (slight quantization error accumulates over denoising steps)
- Only supports `headdim >= 128`
### NVFP4 + Attn-QAT (modified SageAttention3, Blackwell sm_120)
**`ATTN_QAT_INFER`** with **`transformer_quant=nvfp4_qat`**
Runs the DiT fully in 4-bit: NVFP4 linear layers (activations quantized on the
fly) plus the modified SageAttention3 FP4 attention backend. This is the
inference half of the Quantization-Aware Distillation (QAD) recipe and the path
used for the RTX 5090 release.
The `attn_qat_infer` kernel hard-gates on **sm_120 (consumer Blackwell / RTX
5090)**; on other GPUs the backend logs a notice and falls back to Flash
Attention. See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040).
Enable both halves — attention via the env var, linear via `transformer_quant`:
```python
import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "ATTN_QAT_INFER"
from fastvideo import VideoGenerator
from fastvideo.layers.quantization import get_quantization_config
gen = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
# Wan-2.1 uses the nvfp4_qat config (NVFP4 is LTX2-specific). Pass an
# instance — the bare string is not resolved on the from_pretrained path.
transformer_quant=get_quantization_config("nvfp4_qat")(),
use_fsdp_inference=False, # FSDP shards invalidate the FP4 tensor pointers
)
gen.generate(request={"prompt": "A raccoon in sunflowers", "output": {"save_video": True}})
```
Or run the example script:
```bash
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py --bf16 # baseline
```
### Sliding Tile Attention (Archived)
**`SLIDING_TILE_ATTN`**
-4
View File
@@ -58,7 +58,6 @@ pipeline initialization and sampling.
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| Wan2.2 TI2V 5B | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | 720P | ⭕ | ⭕ | ✅ | ⭕ | ⭕ |
| Lucy Edit Dev 5B*** | `decart-ai/Lucy-Edit-Dev` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Wan2.2 T2V A14B | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
| Wan2.2 I2V A14B | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
| HunyuanVideo | `hunyuanvideo-community/HunyuanVideo` | 720px1280p<br>544px960p | ❌ | ✅ | ✅ | ⭕ | ⭕ |
@@ -79,9 +78,6 @@ pipeline initialization and sampling.
**Note**: Wan2.2 TI2V 5B has some quality issues when performing I2V generation. We are working on fixing this issue.
***Lucy Edit Dev uses a non-commercial model license. FastVideo support is
focused on inference integration for video editing workflows.
`Sliding Tile Attn (Legacy Branch)` entries refer to the archived
`sta_do_not_delete` branch workflow, not active `main` inference wiring.
@@ -1,94 +0,0 @@
# v2 porting status — fastvideo models → the v2 (recipe, runtime) substrate
Goal: every model in fastvideo's registry resolves through the **v2 `VideoGenerator`** / `Engine`
(typed `fastvideo.api` configs + the real torch backend) to a recipe that can construct and run it.
**Scope: ALL fastvideo models (achieved).** v2 now resolves **63/64** of fastvideo's registered HF ids
by exact id (PRIMARY), plus the architecture fallback for local/unregistered checkpoints. The single
remaining id — `FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers` — is **environment-blocked**: its VSA
(Sparse-Linear Attention) kernels require `nvcc` (not built in this bring-up). It arch-resolves to the
base Wan card but needs the VSA kernel build to run faithfully.
Dispatch is **architecture-driven** (`v2/registry.py`): exact HF id → short-name → architecture
inference from the checkpoint (pipeline / transformer / VAE class names + `z_dim`, `transformer_2`,
`spatial_upsampler`). Adding a model is one `_BUCKET_C` row (HF ids → builders + transformer class).
## The porting mechanism — self-contained recipe packages
Every net-new arch is a **self-contained recipe package** (`v2/recipes/<arch>/` = `card.py` `loop.py`
`program.py` [+ `sampler.py`] + an optional `v2/platform/backends/torch_<arch>.py` adapter). The card
declares its torch adapter via **`ComponentSpec.adapter="module:Class"`** (the `_explicit_adapter` seam in
`torch_backend.py`) instead of editing the shared `_make_dit`/`_make_vae`/`_make_text_encoder` dispatch —
so a port adds **only new files**, never touching shared code, and parallel ports never conflict. New
samplers/loops live in-package. Registration is one row in `v2/registry.py:_BUCKET_C`.
## Working today (GPU-verified, real video/audio) — committed on `v2`
| Official example(s) | Model | v2 card |
|---|---|---|
| `basic.py`, `basic_mps.py`, `basic_ray.py` | `Wan-AI/Wan2.1-T2V-1.3B-Diffusers` | wan21 |
| `basic_self_forcing_causal.py` | `wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers` | wan_causal |
| `basic_ltx2_distilled.py` | `FastVideo/LTX2-Distilled-Diffusers` (2-stage + spatial upsampler) | ltx2 |
| `basic_wan2_2_ti2v.py` | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | wan2.2-ti2v |
| `basic_wan2_2.py` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` (MoE + CPU expert offload) | wan2.2-a14b |
| `basic_ltx2.py` | `Davids048/LTX2-Base-Diffusers` | ltx2 base |
| `basic_ltx2_3_distilled.py` | `FastVideo/LTX-2.3-Distilled-Diffusers` (joint T2VS, video+audio) | ltx2.3-distilled |
Plus the **Wan2.1 i2v cluster** (Fun-1.3B-InP GPU-verified; I2V-14B-480P/720P + Wan2.2-I2V-A14B MoE reuse
the i2v card) — CLIP image-encoder + first-frame `[mask|cond]` → 36ch DiT.
## GPU bring-up results (real weights on H100 NVL, single-GPU, TORCH_SDPA)
**20 models generate real video/audio on GPU** — the 7 above + **13 of the newly-ported** archs, each run
end-to-end through the real `VideoGenerator` (resolve → stamp → CUDA load → generate). The rest are blocked
by a **fastvideo-shared-code / missing-kernel / HF-access** wall, NOT a v2 recipe bug (the v2 recipes are
faithful — e.g. cosmos25's DiT+VAE produced finite output; only its Qwen2.5-VL encoder hit a library
incompat). All ports also resolve + run end-to-end on the CPU toy backend (`test_bucket_c_ports.py`).
| GPU status | Models |
|---|---|
| ✅ **Verified** (real GPU output) | stable_audio (audio), matrixgame2, matrixgame3, gen3c, wan_fun_control, lucy_edit, hunyuangamecraft, hunyuan_video, hunyuan_video15, longcat (13.58B), sfwan22 (2×14B MoE, expert offload), lingbotworld (2×14B, offload), fastwan (TI2V-5B-FullAttn DMD) |
| 🚫 fastvideo/env-blocked | **cosmos25** (DiT+VAE ran; Qwen2.5-VL encoder → transformers 5.12.1 incompat in fastvideo); **kandinsky5** (fastvideo registry registers a bare `PipelineConfig`); **hyworld** (fastvideo DiT hardcodes `flash_attn`, not built); **turbowan** 1.3B/i2v + **fastwan** VSA-variants (SLA/VSA sparse-attn params + Triton kernels need nvcc) |
| 🚫 access-blocked (HF-gated) | cosmos2, flux2, sd35 (no HF token in this env) |
To unblock the env-blocked: build `fastvideo-kernel` (SLA/VSA Triton, needs nvcc); pin a fastvideo-compatible
`transformers` for the Qwen2.5-VL encoder; add a Kandinsky5 `PipelineConfig` + an SDPA fallback in the
hyworld DiT (all fastvideo-side / environment, not v2 recipe work).
## Newly ported (recipe details)
Each resolves through the registry AND runs end-to-end on the CPU toy backend via the public `Engine`
path (the `v2/tests/test_bucket_c_ports.py` regression guard), emitting the correct modality artifact.
**15 net-new architectures** (each a new `TorchComponent` adapter + recipe):
- **cosmos2** (Cosmos-Predict2-2B-Video2World) — EDM-Karras denoiser; new `CosmosDenoiseLoop` +
`build_karras_sigmas` (the reference port). **cosmos25** (Cosmos-Predict2.5 2B/14B) — flow-match,
per-frame plain-sigma timestep, Reason1/Qwen2.5-VL encoder. **gen3c** (GEN3C) — EDM + 82ch pose-buffer.
- **hunyuan_video** (+FastHunyuan) — reuses WanDenoiseLoop, dual LLaMA+CLIP encoders, Hunyuan VAE.
**hunyuan_video15** (480p/720p). **hunyuangamecraft**, **hyworld** — interactive (camera/action).
- **longcat** (T2V/I2V/VC). **kandinsky5** (5.0 T2V Lite).
- **sd35** (MMDiT, image, triple-encoder). **flux2** (dev/klein, MMDiT image). **stable_audio** (audio).
- **lingbotworld** (camera/Plucker), **matrixgame2**, **matrixgame3** — interactive world models.
**5 Wan-family variants** (reuse the Wan/Causal arch, new in-package sampler/loop/conditioning):
- **turbowan** — rCM few-step (faithful RCMScheduler port), 1.3B/14B T2V + I2V-A14B MoE.
- **lucy_edit** — v2v editor (video-VAE-encode node → 96ch DiT input). **wan_fun_control** — control input.
- **sfwan22** — Self-Forcing Wan2.2-A14B causal + MoE (i2v + t2v). **fastwan** — DMD 3-step (TI2V-5B-FullAttn
loadable; VSA-trained variants + non-strict `to_gate_compress` load are BRINGUP).
BRINGUP scope per port (documented in each package): GPU load/run; for interactive/world-model archs the
action/camera/memory conditioning needs a request-API extension (the t2v/degenerate path is what
CPU-verifies); video2world/i2v frame-replace conditioning is threaded but inert without conditioning inputs.
## Environment
v2 bring-up runs **single-GPU, resident, on the `TORCH_SDPA` backend** (no fastvideo-kernel / VSA / FP4).
The box has been rescheduled across hosts/arches/python versions mid-session; rebuild the venv for the
current arch when that happens: `uv venv --python 3.12 .venv`; comment out `fastvideo-kernel` in
`pyproject.toml`; `uv pip install -e ".[dev]"`. Source `/home/scratch.willlin_ent/.bringup_env`
(`HF_HOME=./.cache` on scratch, `FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA`). v2 CPU mini: 240 passed, 2 skipped.
## How to add a model to the v2 substrate
1. `v2/recipes/<arch>/` — card (declare adapters via `ComponentSpec.adapter`; per-model `SamplingDefaults`),
loop (reuse `WanDenoiseLoop`/`chunk_rollout` or a new in-package loop+sampler), program.
2. `v2/platform/backends/torch_<arch>.py` — a `TorchComponent` subclass (only the forward semantics) if the
arch is genuinely new; reuse `WanDiT`/`LTX2DiT`/`WanVAE`/`T5Encoder` via `load_id` when it isn't.
3. One row in `v2/registry.py:_BUCKET_C` (HF ids → builders; `transformer_cls` for the arch fallback, or
`""` for explicit-id-only capability variants of an existing arch).
4. CPU-verify: it resolves + runs on the toy backend (auto-covered by `test_bucket_c_ports.py`). Then GPU
bring-up (`stamp_*_checkpoints` → real weights) per BRINGUP notes.
-124
View File
@@ -1,124 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""Run full Flux2 text-to-image generation through FastVideo.
User story:
"I have a local or HF Diffusers-format full Flux2 checkpoint and want a
minimal text-to-image generation command that uses embedded guidance."
"""
import argparse
import os
from pathlib import Path
from fastvideo import VideoGenerator
from fastvideo.api import (
ComponentConfig,
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
ParallelismConfig,
PipelineSelection,
SamplingConfig,
)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run full Flux2 text-to-image generation.")
parser.add_argument(
"--model-path",
default="black-forest-labs/FLUX.2-dev",
help="HF id or local diffusers-format full Flux2 weights directory.",
)
parser.add_argument(
"--output",
default="outputs/flux2/flux2.png",
help="Output PNG path.",
)
parser.add_argument(
"--prompt",
default="a photo of a banana on a wooden table, studio lighting",
help="Text prompt.",
)
parser.add_argument("--height", type=int, default=1024)
parser.add_argument("--width", type=int, default=1024)
parser.add_argument("--steps", type=int, default=50)
parser.add_argument("--guidance-scale", type=float, default=4.0)
parser.add_argument("--max-sequence-length", type=int, default=None)
parser.add_argument("--seed", type=int, default=0)
parser.add_argument("--num-gpus", type=int, default=1)
parser.add_argument("--tp-size", type=int, default=None)
parser.add_argument("--sp-size", type=int, default=None)
parser.add_argument(
"--backend",
default=None,
help="Set FASTVIDEO_ATTENTION_BACKEND, for example TORCH_SDPA.",
)
return parser.parse_args()
def main() -> None:
args = parse_args()
if args.backend:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
output = Path(args.output)
output.parent.mkdir(parents=True, exist_ok=True)
tp_size = args.tp_size if args.tp_size is not None else (
args.num_gpus if args.num_gpus > 1 else 1
)
sp_size = args.sp_size if args.sp_size is not None else (
1 if args.num_gpus > 1 else args.num_gpus
)
generator_config = GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=args.num_gpus,
parallelism=ParallelismConfig(tp_size=tp_size, sp_size=sp_size),
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(
workload_type="t2i",
components=ComponentConfig(override_pipeline_cls_name="Flux2Pipeline"),
),
)
generator = VideoGenerator.from_config(generator_config)
try:
sampling = SamplingConfig(
height=args.height,
width=args.width,
num_frames=1,
fps=1,
num_inference_steps=args.steps,
guidance_scale=args.guidance_scale,
seed=args.seed,
)
extensions = {}
if args.max_sequence_length is not None:
extensions["max_sequence_length"] = args.max_sequence_length
request = GenerationRequest(
prompt=args.prompt,
sampling=sampling,
output=OutputConfig(
output_path=str(output),
save_video=True,
return_frames=False,
),
extensions=extensions,
)
generator.generate(request)
finally:
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,98 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""Run Flux2 Klein text-to-image generation through FastVideo.
User story:
"I need a short local smoke for the Flux2 Klein checkpoint before wiring it
into an image workflow. Use the model's distilled four-step defaults and
write a single PNG so I can compare the output against the reference."
"""
import argparse
import os
from fastvideo import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
PipelineSelection,
SamplingConfig,
)
DEFAULT_PROMPT = "a brushed steel espresso machine on a marble counter, morning window light"
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run Flux2 Klein text-to-image generation.")
parser.add_argument(
"--model-path",
default="black-forest-labs/FLUX.2-klein-4B",
help="HF id or local diffusers-format Flux2 Klein weights directory.",
)
parser.add_argument(
"--output-path",
default="outputs/flux2/flux2_klein.png",
help="PNG output path or output directory.",
)
parser.add_argument("--prompt", default=DEFAULT_PROMPT, help="Prompt text.")
parser.add_argument("--seed", type=int, default=0, help="Generation seed.")
parser.add_argument("--height", type=int, default=1024, help="Output image height.")
parser.add_argument("--width", type=int, default=1024, help="Output image width.")
parser.add_argument("--steps", type=int, default=4, help="Number of denoising steps.")
parser.add_argument("--num-gpus", type=int, default=1, help="Number of GPUs to use.")
parser.add_argument(
"--backend",
default=None,
help="Set FASTVIDEO_ATTENTION_BACKEND, for example TORCH_SDPA.",
)
return parser.parse_args()
def main() -> None:
args = parse_args()
if args.backend:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
generator_config = GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=args.num_gpus,
use_fsdp_inference=False,
offload=OffloadConfig(
dit=False,
vae=True,
text_encoder=True,
pin_cpu_memory=False,
),
),
pipeline=PipelineSelection(workload_type="t2i"),
)
generator = VideoGenerator.from_config(generator_config)
try:
request = GenerationRequest(
prompt=args.prompt,
sampling=SamplingConfig(
height=args.height,
width=args.width,
num_frames=1,
fps=1,
num_inference_steps=args.steps,
guidance_scale=1.0,
seed=args.seed,
),
output=OutputConfig(
output_path=args.output_path,
save_video=True,
),
)
generator.generate(request)
finally:
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,289 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2.3 distilled image-to-video with torch.compile + timing breakdown.
This example runs the LTX-2.3 distilled student model on a single GPU with
torch.compile fully enabled, then prints a per-stage timing breakdown so the
user can see where wall-time goes. It is meant as the canonical entry point
for trying out the LTX-2.3 i2v path on `hao-ai-lab/FastVideo:main`.
Quick start
-----------
export LTX23_I2V_IMAGE=/path/to/your/portrait_or_product.jpg
# optional overrides:
# export LTX23_I2V_PROMPT="a fashion model walks toward camera..."
# export LTX23_OUTPUT_DIR=outputs_video/ltx2_3_distilled_i2v
python examples/inference/basic/basic_ltx2_3_distilled_i2v.py
What the script does
--------------------
1. Loads FastVideo/LTX-2.3-Distilled-Diffusers (8 denoise + 3 refine steps,
CFG=1, no refine LoRA — the distilled production recipe).
2. Compiles the DiT, text encoder, and VAE (fullgraph, Inductor default
mode — autotune adds ~7 min cold-compile here with no measurable
e2e gain).
3. Runs 2 warmup calls (untimed) + 2 measured calls. Two warmups are kept
as a safety net — the first call pays cold compile + first-shape guard
work, and a second warmup ensures any residual recompiles settle before
we measure.
4. Prints a per-stage breakdown and an average over the measured runs.
Hardware notes
--------------
- Single-GPU example; for multi-GPU sequence-parallel see the gradio demo
under `examples/inference/gradio/local/gradio_local_demo_ltx2_3/`.
- First-time compile takes ~30-40 min on GB200 (~20 min on H100; cached
in `$TORCHINDUCTOR_CACHE_DIR` afterwards). Subsequent invocations only
pay the one-time process load + a few seconds of dynamo trace.
- On GB200 / Blackwell, run with `env -u LD_LIBRARY_PATH ...` to avoid a
system-cuBLAS / torch-cuBLAS mismatch that fails every GEMM. The
`_inductor.shape_padding = False` line below also avoids a pad_mm
landmine on the same generation of cards.
"""
from __future__ import annotations
import os
import time
from collections import OrderedDict
from pathlib import Path
import torch._inductor.config as _inductor
from fastvideo import VideoGenerator
from fastvideo.configs.pipelines.base import PipelineConfig
from fastvideo.utils import maybe_download_model
# Env knobs (set BEFORE importing fastvideo where possible — but
# FASTVIDEO_ATTENTION_BACKEND is fine here because the worker reads it
# on generator construction).
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "FLASH_ATTN")
os.environ.setdefault("FASTVIDEO_STAGE_LOGGING", "1")
# Inductor knobs. The first one (shape_padding=False) is mandatory on
# Blackwell to avoid a cuBLAS INVALID_VALUE crash inside pad_mm during
# the refine path. The rest are autotune-friendliness flags.
_inductor.shape_padding = False
_inductor.conv_1x1_as_mm = True
_inductor.coordinate_descent_tuning = True
_inductor.coordinate_descent_check_all_directions = True
_inductor.epilogue_fusion = False
MODEL_ID = os.path.expandvars(
os.path.expanduser(
os.getenv("LTX23_MODEL_PATH", "FastVideo/LTX-2.3-Distilled-Diffusers")
)
)
OUTPUT_DIR = Path(
os.getenv("LTX23_OUTPUT_DIR", "outputs_video/ltx2_3_distilled_i2v")
)
I2V_IMAGE = os.getenv("LTX23_I2V_IMAGE", "")
DEFAULT_PROMPT = (
"A fashion model takes a slow step forward and shifts her weight, "
"the soft fabric of her clothing swaying and rippling with the "
"motion, her hair shifting gently, soft even studio lighting on a "
"clean light background, elegant slow-motion runway feel."
)
PROMPT = os.getenv("LTX23_I2V_PROMPT", DEFAULT_PROMPT)
# Per-stage timing helpers --------------------------------------------------
def _print_stage_breakdown(result: dict, label: str) -> float | None:
"""Print stage execution times and return the sum, or None if missing."""
logging_info = result.get("logging_info")
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
print(f" [{label}] stage breakdown unavailable")
return None
print(f" [{label}] stage breakdown:")
total = 0.0
for name, metrics in stages.items():
exec_s = float(metrics.get("execution_time", 0.0))
total += exec_s
print(f" - {name}: {exec_s:.3f}s")
print(f" - stage_sum: {total:.3f}s")
return total
def _collect_stage_times(
result: dict,
stage_times: dict[str, list[float]],
stage_order: OrderedDict[str, None],
) -> None:
logging_info = result.get("logging_info")
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
return
for name, metrics in stages.items():
stage_order.setdefault(name, None)
stage_times.setdefault(name, []).append(
float(metrics.get("execution_time", 0.0))
)
def _resolve_refine_upsampler(model_root: str) -> Path:
"""LTX-2.3 distilled snapshots ship a `spatial_upscaler/` subdir."""
for name in ("spatial_upscaler", "spatial_upsampler"):
cand = Path(model_root) / name
if (cand / "config.json").is_file():
return cand
raise FileNotFoundError(
f"No refine upsampler directory under {model_root}. "
f"Expected `{model_root}/spatial_upscaler/config.json`."
)
# Main ---------------------------------------------------------------------
def main() -> None:
if not I2V_IMAGE:
raise SystemExit(
"LTX23_I2V_IMAGE is required for i2v. Example:\n"
" export LTX23_I2V_IMAGE=/path/to/portrait_or_product.jpg\n"
" python examples/inference/basic/basic_ltx2_3_distilled_i2v.py"
)
if not Path(I2V_IMAGE).is_file():
raise SystemExit(f"LTX23_I2V_IMAGE not found: {I2V_IMAGE}")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
model_root = maybe_download_model(MODEL_ID)
refine_upsampler_path = _resolve_refine_upsampler(model_root)
print(f"Model: {model_root}")
print(f"Refine upsampler: {refine_upsampler_path}")
print(f"i2v image: {I2V_IMAGE}")
print(f"Output dir: {OUTPUT_DIR.resolve()}")
# mode="default" — Inductor's default schedule matches max-autotune on
# this pipeline (denoise/refine/decode all within ~5 ms, n=2) while
# saving ~7 min of cold compile on a single GB200.
torch_compile_kwargs = {
"backend": "inductor",
"fullgraph": True,
"mode": "default",
"dynamic": False,
}
# Loading the pipeline config *with model_path* binds model-specific
# tuning (notably VAE precision/decoder defaults) into the config. Without
# this, the generic pipeline config gives a substantially slower VAE
# decode stage. `basic_ltx2_distilled_fast_profile.py` uses the same
# pattern.
pipeline_config = PipelineConfig.from_pretrained(model_root)
pipeline_config.dit_config.quant_config = None
generator = VideoGenerator.from_pretrained(
model_root,
num_gpus=1,
# LTX-2.3 distilled uses the two-stage refine pipeline; the refine
# LoRA is intentionally empty for the distilled student.
ltx2_refine_enabled=True,
ltx2_refine_upsampler_path=str(refine_upsampler_path),
ltx2_refine_lora_path="",
ltx2_refine_num_inference_steps=3,
ltx2_refine_guidance_scale=1.0,
ltx2_refine_add_noise=True,
pipeline_config=pipeline_config,
enable_torch_compile=True,
enable_torch_compile_text_encoder=True,
# Compile the VAE codec submodules (encoder / decoder) too. The
# `LTX2CausalVideoAutoencoder` declares `_compile_conditions` so
# `_compile_with_conditions` targets just those submodules and
# leaves the surrounding tiling control flow eager — needed for
# fullgraph + dynamic=False to succeed. VAE eager decode is
# ~1.0s; compiling it brings the stage to ~0.3s.
enable_torch_compile_vae=True,
torch_compile_kwargs=torch_compile_kwargs,
torch_compile_kwargs_vae=torch_compile_kwargs,
# Keep everything resident — no CPU offload for serving-style runs.
dit_cpu_offload=False,
text_encoder_cpu_offload=False,
vae_cpu_offload=False,
ltx2_vae_tiling=False,
)
common_kwargs = dict(
prompt=PROMPT,
negative_prompt="", # distilled is CFG-free; no negative needed
guidance_scale=1.0, # CFG=1 for distilled
height=1280, width=832, # portrait runway aspect
num_frames=121, fps=24, # ~5s clip
num_inference_steps=8, # distilled denoise steps
# i2v: anchor the input image at frame 0 with full strength.
# `ltx2_image_crf=0.0` skips an extra JPEG re-encode of an already
# JPEG conditioning image.
ltx2_images=[(I2V_IMAGE, 0, 1.0)],
ltx2_image_crf=0.0,
save_video=True,
)
warmup_runs = 2
measured_runs = 2
warmup_secs: list[float] = []
measured_secs: list[float] = []
stage_times: dict[str, list[float]] = {}
stage_order: OrderedDict[str, None] = OrderedDict()
try:
# Warmup: untimed (but we still wall-clock them so the first compile
# cost is visible to the reader).
for w in range(warmup_runs):
t0 = time.perf_counter()
print(f"\n[warmup {w + 1}/{warmup_runs}] compiling + generating…")
generator.generate_video(
output_path=str(OUTPUT_DIR / f"_warmup_{w + 1}.mp4"),
seed=7,
**common_kwargs,
)
dt = time.perf_counter() - t0
warmup_secs.append(dt)
print(f"[warmup {w + 1}/{warmup_runs}] wall={dt:.1f}s")
# Cleanup warmup artifacts so the user only sees measured outputs.
for w in range(warmup_runs):
(OUTPUT_DIR / f"_warmup_{w + 1}.mp4").unlink(missing_ok=True)
# Measured.
for m in range(measured_runs):
out_path = OUTPUT_DIR / f"output_ltx2_3_distilled_i2v_run_{m + 1}.mp4"
print(f"\n[measured {m + 1}/{measured_runs}] generating: {out_path}")
t0 = time.perf_counter()
result = generator.generate_video(
output_path=str(out_path),
seed=2002 + m,
**common_kwargs,
)
wall = time.perf_counter() - t0
e2e = (
result.get("e2e_latency")
if isinstance(result, dict) else None
) or wall
measured_secs.append(e2e)
print(f"[measured {m + 1}/{measured_runs}] e2e={e2e:.2f}s wall={wall:.2f}s")
if isinstance(result, dict):
_print_stage_breakdown(result, f"measured {m + 1}")
_collect_stage_times(result, stage_times, stage_order)
# Summary.
print("\n=== summary ===")
print(f"warmup wall-times: {[round(x, 1) for x in warmup_secs]}")
if measured_secs:
avg = sum(measured_secs) / len(measured_secs)
print(
f"measured e2e (n={len(measured_secs)}): "
f"{[round(x, 2) for x in measured_secs]} -> avg {avg:.2f}s"
)
if stage_times:
print(f"average stage times over {measured_runs} measured runs:")
avg_total = 0.0
for name in stage_order:
vals = stage_times.get(name) or []
if not vals:
continue
avg_v = sum(vals) / len(vals)
avg_total += avg_v
print(f" - {name}: {avg_v:.3f}s")
print(f" - stage_sum_avg: {avg_total:.3f}s")
finally:
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,350 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""LTX-2.3 distilled image-to-video — typed API (``from_config`` / ``generate``).
Identical generation behavior to ``basic_ltx2_3_distilled_i2v.py``, but
expressed through the newer typed surface (``GeneratorConfig`` /
``GenerationRequest``) instead of the ``from_pretrained(**legacy_kwargs)``
bridge. The typed API is now the preferred entry point — the legacy
example still works but emits a ``DeprecationWarning`` for the LTX-2.3
specific knobs.
Quick start
-----------
export LTX23_I2V_IMAGE=/path/to/your/portrait_or_product.jpg
# optional overrides:
# export LTX23_I2V_PROMPT="a fashion model walks toward camera..."
# export LTX23_OUTPUT_DIR=outputs_video/ltx2_3_distilled_i2v_typed
python examples/inference/basic/basic_ltx2_3_distilled_i2v_typed.py
What the script does
--------------------
1. Loads FastVideo/LTX-2.3-Distilled-Diffusers (8 denoise + 3 refine
steps, CFG=1, no refine LoRA — the distilled production recipe).
2. Compiles the DiT, text encoder, and VAE (fullgraph, Inductor default
mode — autotune adds ~7 min cold-compile here with no measurable
e2e gain).
3. Runs 2 warmup calls (untimed) + 2 measured calls. Two warmups are
kept as a safety net — the first call pays cold compile + first-shape
guard work, and a second warmup ensures any residual recompiles
settle before we measure.
4. Prints a per-stage breakdown and an average over the measured runs.
Hardware notes
--------------
- Single-GPU example; for multi-GPU sequence-parallel see the gradio
demo under ``examples/inference/gradio/local/gradio_local_demo_ltx2_3/``.
- First-time compile takes ~30-40 min on GB200 (~20 min on H100;
cached in ``$TORCHINDUCTOR_CACHE_DIR`` afterwards). Subsequent
invocations only pay the one-time process load + a few seconds of
dynamo trace.
- On GB200 / Blackwell, run with ``env -u LD_LIBRARY_PATH ...`` to
avoid a system-cuBLAS / torch-cuBLAS mismatch that fails every GEMM.
The ``_inductor.shape_padding = False`` line below also avoids a
``pad_mm`` landmine on the same generation of cards.
Typed-API mapping (legacy kwarg ↔ typed field)
----------------------------------------------
- ``num_gpus`` ↔ ``engine.num_gpus``
- ``enable_torch_compile`` ↔ ``engine.compile.enabled``
- ``enable_torch_compile_text_encoder`` ↔ ``engine.compile.text_encoder_enabled``
- ``enable_torch_compile_vae`` ↔ ``engine.compile.vae_enabled``
- ``torch_compile_kwargs`` ↔ ``engine.compile.backend/fullgraph/mode/dynamic``
- ``torch_compile_kwargs_vae`` ↔ empty ``compile.vae_kwargs`` (inherits master)
- ``dit_cpu_offload`` ↔ ``engine.offload.dit``
- ``text_encoder_cpu_offload`` ↔ ``engine.offload.text_encoder``
- ``vae_cpu_offload`` ↔ ``engine.offload.vae``
- ``ltx2_vae_tiling`` ↔ ``pipeline.vae_tiling``
- ``ltx2_refine_enabled`` ↔ ``pipeline.preset_overrides["refine"]["enabled"]``
- ``ltx2_refine_upsampler_path`` ↔ ``pipeline.components.upsampler_weights``
- ``ltx2_refine_lora_path`` ↔ ``pipeline.components.lora_path``
- ``ltx2_refine_num_inference_steps`` ↔ ``pipeline.preset_overrides["refine"]["num_inference_steps"]``
- ``ltx2_refine_guidance_scale`` ↔ ``pipeline.preset_overrides["refine"]["guidance_scale"]``
- ``ltx2_refine_add_noise`` ↔ ``pipeline.preset_overrides["refine"]["add_noise"]``
- ``pipeline_config=PipelineConfig.from_pretrained(model_root)`` ↔ (no-op — ``PipelineConfig.from_kwargs`` already resolves the model-specific class from ``model_path``)
- ``pipeline_config.dit_config.quant_config = None`` ↔ leave ``engine.quantization`` unset
- ``ltx2_images`` / ``ltx2_image_crf`` ↔ ``request.extensions`` (LTX-2 specific, no
first-class typed field yet)
"""
from __future__ import annotations
import os
import time
from collections import OrderedDict
from pathlib import Path
import torch._inductor.config as _inductor
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig,
ComponentConfig,
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
PipelineSelection,
SamplingConfig,
)
from fastvideo.utils import maybe_download_model
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "FLASH_ATTN")
os.environ.setdefault("FASTVIDEO_STAGE_LOGGING", "1")
# Inductor knobs. ``shape_padding=False`` is mandatory on Blackwell to
# avoid a cuBLAS INVALID_VALUE crash inside pad_mm during the refine
# path. The rest are autotune-friendliness flags.
_inductor.shape_padding = False
_inductor.conv_1x1_as_mm = True
_inductor.coordinate_descent_tuning = True
_inductor.coordinate_descent_check_all_directions = True
_inductor.epilogue_fusion = False
MODEL_ID = os.path.expandvars(
os.path.expanduser(
os.getenv("LTX23_MODEL_PATH", "FastVideo/LTX-2.3-Distilled-Diffusers")
)
)
OUTPUT_DIR = Path(
os.getenv(
"LTX23_OUTPUT_DIR", "outputs_video/ltx2_3_distilled_i2v_typed"
)
)
I2V_IMAGE = os.getenv("LTX23_I2V_IMAGE", "")
DEFAULT_PROMPT = (
"A fashion model takes a slow step forward and shifts her weight, "
"the soft fabric of her clothing swaying and rippling with the "
"motion, her hair shifting gently, soft even studio lighting on a "
"clean light background, elegant slow-motion runway feel."
)
PROMPT = os.getenv("LTX23_I2V_PROMPT", DEFAULT_PROMPT)
def _print_stage_breakdown(result, label: str) -> float | None:
logging_info = getattr(result, "logging_info", None)
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
print(f" [{label}] stage breakdown unavailable")
return None
print(f" [{label}] stage breakdown:")
total = 0.0
for name, metrics in stages.items():
exec_s = float(metrics.get("execution_time", 0.0))
total += exec_s
print(f" - {name}: {exec_s:.3f}s")
print(f" - stage_sum: {total:.3f}s")
return total
def _collect_stage_times(
result,
stage_times: dict[str, list[float]],
stage_order: OrderedDict[str, None],
) -> None:
logging_info = getattr(result, "logging_info", None)
stages = getattr(logging_info, "stages", None) if logging_info else None
if not stages:
return
for name, metrics in stages.items():
stage_order.setdefault(name, None)
stage_times.setdefault(name, []).append(
float(metrics.get("execution_time", 0.0))
)
def _resolve_refine_upsampler(model_root: str) -> Path:
for name in ("spatial_upscaler", "spatial_upsampler"):
cand = Path(model_root) / name
if (cand / "config.json").is_file():
return cand
raise FileNotFoundError(
f"No refine upsampler directory under {model_root}. "
f"Expected `{model_root}/spatial_upscaler/config.json`."
)
def main() -> None:
if not I2V_IMAGE:
raise SystemExit(
"LTX23_I2V_IMAGE is required for i2v. Example:\n"
" export LTX23_I2V_IMAGE=/path/to/portrait_or_product.jpg\n"
" python examples/inference/basic/"
"basic_ltx2_3_distilled_i2v_typed.py"
)
if not Path(I2V_IMAGE).is_file():
raise SystemExit(f"LTX23_I2V_IMAGE not found: {I2V_IMAGE}")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
model_root = maybe_download_model(MODEL_ID)
refine_upsampler_path = _resolve_refine_upsampler(model_root)
print(f"Model: {model_root}")
print(f"Refine upsampler: {refine_upsampler_path}")
print(f"i2v image: {I2V_IMAGE}")
print(f"Output dir: {OUTPUT_DIR.resolve()}")
# mode="default" — Inductor's default schedule matches max-autotune on
# this pipeline (denoise/refine/decode all within ~5 ms, n=2) while
# saving ~7 min of cold compile on a single GB200.
generator_config = GeneratorConfig(
model_path=model_root,
engine=EngineConfig(
num_gpus=1,
# Keep DiT / text encoder / VAE resident on GPU — no CPU offload
# for serving-style runs. ``image_encoder`` and
# ``pin_cpu_memory`` are left at their schema defaults
# (matches the legacy example, which only set these three).
offload=OffloadConfig(
dit=False,
text_encoder=False,
vae=False,
),
compile=CompileConfig(
enabled=True,
text_encoder_enabled=True,
# ``vae_enabled`` triggers ``_compile_with_conditions`` on
# ``LTX2CausalVideoAutoencoder``, which compiles just the
# encoder/decoder submodules and leaves the surrounding
# tiling control flow eager (required for ``fullgraph``).
# Empty ``vae_kwargs`` → inherits the master kwargs below.
vae_enabled=True,
backend="inductor",
fullgraph=True,
mode="default",
dynamic=False,
),
),
pipeline=PipelineSelection(
# ``PipelineConfig.from_kwargs`` resolves the model-specific
# pipeline-config class from ``model_path`` automatically, so we
# don't need to set ``components.pipeline_config_path`` — the
# model-specific VAE precision / decoder defaults are picked up
# the same way the legacy example's
# ``PipelineConfig.from_pretrained(model_root)`` did them.
components=ComponentConfig(
upsampler_weights=str(refine_upsampler_path),
# Distilled has no refine LoRA — omit ``lora_path``.
),
vae_tiling=False,
preset_overrides={
"refine": {
"enabled": True,
"num_inference_steps": 3,
"guidance_scale": 1.0,
"add_noise": True,
},
},
),
)
generator = VideoGenerator.from_config(generator_config)
def build_request(out_path: Path, seed: int) -> GenerationRequest:
return GenerationRequest(
prompt=PROMPT,
# distilled is CFG-free; no negative prompt
negative_prompt="",
sampling=SamplingConfig(
num_videos_per_prompt=1,
seed=seed,
height=1280,
width=832,
num_frames=121,
fps=24,
num_inference_steps=8,
guidance_scale=1.0,
),
output=OutputConfig(
output_path=str(out_path),
save_video=True,
return_frames=False,
),
# LTX-2.3 i2v fields don't have first-class typed slots yet;
# extensions is the documented bridge. ``ltx2_image_crf=0.0``
# skips an extra JPEG re-encode of an already JPEG image.
extensions={
"ltx2_images": [(I2V_IMAGE, 0, 1.0)],
"ltx2_image_crf": 0.0,
},
)
warmup_runs = 2
measured_runs = 2
warmup_secs: list[float] = []
measured_secs: list[float] = []
stage_times: dict[str, list[float]] = {}
stage_order: OrderedDict[str, None] = OrderedDict()
try:
for w in range(warmup_runs):
print(f"\n[warmup {w + 1}/{warmup_runs}] compiling + generating…")
t0 = time.perf_counter()
generator.generate(
build_request(
OUTPUT_DIR / f"_warmup_{w + 1}.mp4", seed=7
)
)
dt = time.perf_counter() - t0
warmup_secs.append(dt)
print(f"[warmup {w + 1}/{warmup_runs}] wall={dt:.1f}s")
for w in range(warmup_runs):
(OUTPUT_DIR / f"_warmup_{w + 1}.mp4").unlink(missing_ok=True)
for m in range(measured_runs):
out_path = (
OUTPUT_DIR
/ f"output_ltx2_3_distilled_i2v_typed_run_{m + 1}.mp4"
)
print(
f"\n[measured {m + 1}/{measured_runs}] generating: {out_path}"
)
t0 = time.perf_counter()
result = generator.generate(
build_request(out_path, seed=2002 + m)
)
wall = time.perf_counter() - t0
# ``e2e_latency`` is currently surfaced via ``result.extra``;
# ``GenerationResult`` exposes ``generation_time`` as a
# first-class field but the LTX-2 pipeline only fills the
# legacy ``e2e_latency`` key. Prefer the explicit one, fall
# back to wall-clock.
e2e = (
result.extra.get("e2e_latency")
if hasattr(result, "extra") else None
) or wall
measured_secs.append(e2e)
print(
f"[measured {m + 1}/{measured_runs}] "
f"e2e={e2e:.2f}s wall={wall:.2f}s"
)
_print_stage_breakdown(result, f"measured {m + 1}")
_collect_stage_times(result, stage_times, stage_order)
print("\n=== summary ===")
print(
f"warmup wall-times: "
f"{[round(x, 1) for x in warmup_secs]}"
)
if measured_secs:
avg = sum(measured_secs) / len(measured_secs)
print(
f"measured e2e (n={len(measured_secs)}): "
f"{[round(x, 2) for x in measured_secs]} -> avg {avg:.2f}s"
)
if stage_times:
print(f"average stage times over {measured_runs} measured runs:")
avg_total = 0.0
for name in stage_order:
vals = stage_times.get(name) or []
if not vals:
continue
avg_v = sum(vals) / len(vals)
avg_total += avg_v
print(f" - {name}: {avg_v:.3f}s")
print(f" - stage_sum_avg: {avg_total:.3f}s")
finally:
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,38 +0,0 @@
from fastvideo import VideoGenerator
OUTPUT_PATH = "video_samples_lucy_edit"
def main():
generator = VideoGenerator.from_pretrained(
"decart-ai/Lucy-Edit-Dev",
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=True,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
)
prompt = ("Change the apron and blouse to a classic clown costume: satin "
"polka-dot jumpsuit in bright primary colors, ruffled white collar, "
"oversized pom-pom buttons, white gloves, oversized red shoes, red "
"foam nose; soft window light from left, eye-level medium shot.")
video_path = "https://d2drjpuinn46lb.cloudfront.net/painter_original_edit.mp4"
generator.generate_video(
prompt,
negative_prompt="",
video_path=video_path,
output_path=OUTPUT_PATH,
save_video=True,
height=480,
width=832,
num_frames=81,
fps=24,
guidance_scale=5.0,
)
if __name__ == "__main__":
main()
-36
View File
@@ -1,36 +0,0 @@
"""v2 port of basic.py — Wan2.1-T2V-1.3B through the v2 VideoGenerator.
Same convenience API as upstream (from_pretrained + generate_video); only delta is importing
VideoGenerator from v2. v2 bring-up: single-GPU, resident, SDPA; modest res/frames for a quick run.
"""
from v2 import VideoGenerator
OUTPUT_PATH = "v2_video_samples"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
pin_cpu_memory=False,
)
common = dict(output_path=OUTPUT_PATH, save_video=True,
num_frames=25, height=480, width=832, num_inference_steps=30, guidance_scale=5.0)
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with "
"interest. The playful yet serene atmosphere is complemented by soft natural light "
"filtering through the petals. Mid-shot, warm and cheerful tones.")
video = generator.generate_video(prompt, output_video_name="wan21_raccoon", **common)
prompt2 = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
"the warm afternoon sun. Low angle, steady tracking shot, cinematic.")
video2 = generator.generate_video(prompt2, output_video_name="wan21_lion", **common)
print(f"Outputs: {video.video_path} , {video2.video_path}")
if __name__ == "__main__":
main()
-29
View File
@@ -1,29 +0,0 @@
"""v2 port of basic_ltx2.py — LTX-2 base (single-stage) through the v2 VideoGenerator.
Same convenience API as upstream; only delta is importing VideoGenerator from v2. LTX-2 base is the
single-stage (non-distilled) model: the v2 single-stage card (build_ltx2_base_card) runs a request-driven
many-step flow-match at FULL latent res (no distilled base/refine split, no spatial upsampler), reusing
the LTX-2 DiT/VAE/Gemma adapters. The SAME single-stage card also serves LTX-2.3-Distilled (which is also
single-stage) — just pass fewer num_inference_steps for the few-step distilled schedule.
NOTE: modest res/frames here — upstream defaults to 1088x1920x121, which on an 18.88B base is very slow;
raise them for full quality. v2 bring-up: single-GPU, resident, SDPA.
"""
from v2 import VideoGenerator
PROMPT = ("A warm sunny backyard, cinematic close-up of two people talking; the camera slowly pans right "
"to reveal a grandfather in the garden wearing enormous butterfly wings, flapping his arms like "
"he is trying to take off. Deadpan, absurd, quietly tragic.")
def main() -> None:
generator = VideoGenerator.from_pretrained("Davids048/LTX2-Base-Diffusers", num_gpus=1)
video = generator.generate_video(
prompt=PROMPT, output_path="v2_video_samples_ltx2_base", output_video_name="ltx2_base_backyard",
save_video=True, num_frames=25, height=512, width=768, num_inference_steps=30)
print(f"Output: {video.video_path}")
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,36 +0,0 @@
"""v2 port of basic_ltx2_3_distilled.py — LTX-2.3 Distilled (single-stage, joint A/V) through the v2
VideoGenerator.
Unlike LTX-2.0 distilled (two-stage, video-only), LTX-2.3 is a single-stage *audio+video* model. The
shared registry (v2/registry.py) maps ``FastVideo/LTX-2.3-Distilled-Diffusers`` to its OWN card,
``build_ltx2_3_card`` — distinct from the LTX-2 base/2-stage cards — which wires the 2.3-specific path:
* SEPARATE video + audio text connectors (the Gemma encoder projects the prompt to two embeddings,
2048-dim for audio, 4096-dim for video) plus gated attention;
* a JOINT DiT forward where video and audio latents cross-attend in a single denoise per step;
* a video VAE decode + an AudioDecoder→Vocoder decode → video frames AND a stereo waveform @24kHz.
Because the model advertises TEXT_TO_VIDEO_SOUND, the VideoGenerator issues a T2VS request by default,
so ``generate_video`` returns BOTH modalities: the mp4 plus a sibling ``.wav`` (and ``result.audio`` /
``result.audio_sample_rate`` in memory). Being distilled, it wants FEW steps (8). GPU-verified on the
rebuilt x86 stack: video (3,33,256,384) + stereo audio (2×61920 @ 24kHz).
"""
from v2 import VideoGenerator
PROMPT = "ocean waves crashing on rocks at sunset, seagulls calling in the distance, cinematic, highly detailed"
def main() -> None:
generator = VideoGenerator.from_pretrained("FastVideo/LTX-2.3-Distilled-Diffusers", num_gpus=1)
# audio=None auto-enables sound for this A/V model (pass audio=False to force video-only).
result = generator.generate_video(
prompt=PROMPT, output_path="v2_video_samples_ltx2_3", output_video_name="ltx2_3_ocean",
save_video=True, num_frames=33, height=512, width=768, num_inference_steps=8, seed=1)
print(f"Video: {result.video_path}")
audio_path = result.extra.get("audio_path")
if audio_path:
print(f"Audio: {audio_path} ({result.audio_sample_rate} Hz)")
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,87 +0,0 @@
"""v2 typed-API inference example — mirrors ``basic_dmd_new_api.py`` but drives the **v2
(recipe, runtime) substrate + real torch backend** for the three models brought up on GPU
(Wan2.1, SF-causal Wan, LTX-2).
The ONLY delta from the upstream example is importing ``VideoGenerator`` from ``v2`` instead of
``fastvideo`` — the typed config classes are the SAME ``fastvideo.api`` dataclasses.
Run (on a GPU box, with the v2 venv active):
python examples/inference/basic/v2_basic_new_api.py
Notes vs upstream: the v2 bring-up runs single-GPU, resident, on the TORCH_SDPA backend (no
fastvideo-kernel / VSA), so resolutions/steps are modest here for a quick runnable demo. LTX-2 loads
an 18.88B DiT (slow first load).
"""
import os
import time
from v2 import VideoGenerator
from fastvideo.api import (
EngineConfig,
GenerationRequest,
GeneratorConfig,
OffloadConfig,
OutputConfig,
SamplingConfig,
)
OUTPUT_PATH = "v2_video_samples"
MODELS = [
{
"family": "wan21",
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"prompt": "a red panda surfing on ocean waves at sunset, cinematic, highly detailed",
"sampling": SamplingConfig(num_frames=25, height=480, width=832,
num_inference_steps=30, guidance_scale=5.0, seed=1, fps=16),
},
{
"family": "wan_causal",
"model_path": "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers",
"prompt": "a cat walking through a sunlit garden, cinematic",
"sampling": SamplingConfig(num_frames=25, height=480, width=832,
num_inference_steps=4, guidance_scale=5.0, seed=1, fps=16),
},
{
"family": "ltx2",
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"prompt": "surfers riding ocean waves at sunset, cinematic, highly detailed",
"sampling": SamplingConfig(num_frames=9, height=512, width=768,
num_inference_steps=8, guidance_scale=1.0, seed=1, fps=16),
},
]
def run_one(m: dict) -> None:
generator_config = GeneratorConfig(
model_path=m["model_path"],
engine=EngineConfig(
num_gpus=1,
use_fsdp_inference=False,
offload=OffloadConfig(text_encoder=False, dit=False, vae=False, pin_cpu_memory=False),
),
)
load_start = time.perf_counter()
generator = VideoGenerator.from_config(generator_config)
load_time = time.perf_counter() - load_start
request = GenerationRequest(
prompt=m["prompt"],
sampling=m["sampling"],
output=OutputConfig(output_path=OUTPUT_PATH, output_video_name=f"v2_{m['family']}",
save_video=True, return_frames=False),
)
gen_start = time.perf_counter()
result = generator.generate(request)
gen_time = time.perf_counter() - gen_start
print(f"[{m['family']:10s}] load={load_time:6.1f}s gen={gen_time:6.1f}s -> {result.video_path}")
def main() -> None:
for m in MODELS:
run_one(m)
if __name__ == "__main__":
main()
@@ -1,30 +0,0 @@
"""v2 port of basic_self_forcing_causal.py — SF-causal Wan2.1 (CausalWanTransformer3DModel) through
the v2 VideoGenerator (chunk_rollout loop).
Same convenience API as upstream; only delta is importing VideoGenerator from v2. NOTE: the v2 causal
loop runs per-chunk few-step (not the upstream kv-cache streaming + SF schedule), so output is coherent
but lower-fidelity (a documented gap). num_frames is set by the card's chunk schedule; height/width
drive the latent geometry.
"""
from v2 import VideoGenerator
from fastvideo.api.sampling_param import SamplingParam
OUTPUT_PATH = "v2_video_samples_causal"
def main() -> None:
model_name = "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_name, num_gpus=1, text_encoder_cpu_offload=False, dit_cpu_offload=False)
sampling_param = SamplingParam.from_pretrained(model_name)
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with "
"interest. The playful yet serene atmosphere is complemented by soft natural light "
"filtering through the petals. Mid-shot, warm and cheerful tones.")
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="causal_raccoon",
save_video=True, sampling_param=sampling_param, height=480, width=832)
print(f"Output: {video.video_path}")
if __name__ == "__main__":
main()
@@ -1,40 +0,0 @@
"""v2 port of basic_wan2_2.py — Wan2.2-T2V-A14B (MoE) through the v2 VideoGenerator.
Same convenience API as upstream; only delta is importing VideoGenerator from v2. A14B is a 2-expert
MoE: WanTransformer3DModel x2 (in_ch=16, Wan2.1 geometry) with a boundary-timestep switch
(boundary_ratio 0.875) — ported via build_wan22_a14b_card (BoundaryTimestepRouting: transformer =
high-noise expert, transformer_2 = low-noise), reusing the Wan adapters for both experts.
NOTE: upstream runs A14B with num_gpus=2 + dit_cpu_offload=True ("DiT need to be offloaded for MoE").
The v2 bring-up is single-GPU + resident (no offload), so the two 14B experts (~56GB bf16) + UMT5 are
near an 80GB GPU's limit — this example uses reduced res/frames to fit. If it OOMs, the A14B card is
still correct; it just needs the (not-yet-ported) MoE DiT CPU offload. See V2_PORTING_STATUS.md.
"""
from v2 import VideoGenerator
OUTPUT_PATH = "v2_video_samples_wan2_2_14B_t2v"
def main() -> None:
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.2-T2V-A14B-Diffusers",
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
pin_cpu_memory=False,
)
prompt = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
"the warm afternoon sun. The tall grass ripples gently in the breeze. Low angle, steady "
"tracking shot, cinematic.")
# Reduced res/frames so the two resident 14B experts fit a single 80GB GPU (upstream: 720x1280x81).
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="wan22_a14b_lion",
save_video=True, num_frames=17, height=480, width=832,
num_inference_steps=20, guidance_scale=5.0)
print(f"Output: {video.video_path}")
if __name__ == "__main__":
main()
@@ -1,38 +0,0 @@
"""v2 port of basic_wan2_2_ti2v.py — Wan2.2-TI2V-5B (T2V mode) through the v2 VideoGenerator.
Same convenience API as upstream (from_pretrained + generate_video); only delta is importing
VideoGenerator from v2. Wan2.2-TI2V-5B reuses the Wan adapter classes (WanTransformer3DModel /
AutoencoderKLWan / UMT5) with the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal).
NOTE: upstream also runs I2V (image_path=...). The v2 program here is T2V-only (image conditioning is
not yet ported), so this mirrors the upstream *T2V* branch (prompt2). Modest res/frames for a quick run.
"""
from v2 import VideoGenerator
OUTPUT_PATH = "v2_video_samples_wan2_2_5B_ti2v"
def main() -> None:
model_name = "Wan-AI/Wan2.2-TI2V-5B-Diffusers"
generator = VideoGenerator.from_pretrained(
model_name,
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
pin_cpu_memory=False,
)
# T2V mode (the v2 program is text-to-video; upstream's image_path I2V branch is not ported yet).
prompt = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
"the warm afternoon sun. The tall grass ripples gently in the breeze, enhancing the lion's "
"commanding presence. Low angle, steady tracking shot, cinematic.")
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="wan22_ti2v_lion",
save_video=True, num_frames=25, height=448, width=768,
num_inference_steps=20, guidance_scale=5.0)
print(f"Output: {video.video_path}")
if __name__ == "__main__":
main()
@@ -1,91 +0,0 @@
"""Run ``judge.third_person_separation`` (needs ``.[eval-judge]`` + a Gemini key)
over each baseline and print the candidate's win-rate table — from a ``--manifest``
of pairs, or by pairing ``--candidate-dir`` against each ``--reference`` dir by
filename stem.
"""
from __future__ import annotations
import argparse
import json
from collections import defaultdict
from pathlib import Path
from fastvideo.eval import create_evaluator
METRIC = "judge.third_person_separation"
VIDEO_EXTS = {".mp4", ".avi", ".mov", ".mkv", ".webm"}
IMAGE_EXTS = {".png", ".jpg", ".jpeg", ".webp"}
def _by_stem(directory: Path, exts: set[str]) -> dict[str, Path]:
"""Map filename stem -> path for files with the given extensions."""
return {p.stem: p for p in sorted(directory.iterdir()) if p.suffix.lower() in exts}
def main() -> None:
p = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
p.add_argument("--candidate-dir", type=Path, default=None,
help="Directory of candidate clips (directory mode).")
p.add_argument("--reference", action="append", default=[], metavar="NAME=DIR",
help="Baseline directory, repeatable: 'name=dir' or bare 'dir'.")
p.add_argument("--image-dir", type=Path, default=None,
help="Optional first-frame images, matched to clips by stem.")
p.add_argument("--prompts-json", type=Path, default=None,
help="Optional {stem: control-signal text} JSON.")
p.add_argument("--actions-json", type=Path, default=None,
help="Optional {stem: action-label} JSON for the per-action breakdown.")
p.add_argument("--manifest", type=Path, default=None,
help="JSON list of {baseline, video_path, reference_path, ...} rows.")
p.add_argument("--output", type=Path, default=None)
args = p.parse_args()
# Group path-only samples per baseline: {baseline: [sample dict, ...]}.
by_baseline: dict[str, list[dict]] = defaultdict(list)
if args.manifest is not None:
for row in json.loads(args.manifest.read_text()):
by_baseline[row.get("baseline", "baseline")].append(
{k: v for k, v in row.items() if k != "baseline"})
elif args.candidate_dir is not None and args.reference:
cands = _by_stem(args.candidate_dir, VIDEO_EXTS)
images = _by_stem(args.image_dir, IMAGE_EXTS) if args.image_dir else {}
prompts = json.loads(args.prompts_json.read_text()) if args.prompts_json else {}
actions = json.loads(args.actions_json.read_text()) if args.actions_json else {}
for spec in args.reference:
name, sep, ref_dir = spec.partition("=")
if not sep:
name, ref_dir = Path(spec).name, spec
refs = _by_stem(Path(ref_dir), VIDEO_EXTS)
for stem in sorted(cands.keys() & refs.keys()):
sample = {"video_path": str(cands[stem]), "reference_path": str(refs[stem])}
if stem in images:
sample["image_path"] = str(images[stem])
if stem in prompts:
sample["text_prompt"] = prompts[stem]
if stem in actions:
sample["action"] = actions[stem]
by_baseline[name].append(sample)
else:
p.error("provide either --manifest, or --candidate-dir with at least one --reference")
ev = create_evaluator(metrics=[METRIC], device="cpu")
print("\n| Baseline | Candidate win-rate (excl. ties) | W / L / T | n |")
print("|---|---|---|---|")
rows = {}
for baseline, samples in by_baseline.items():
res = ev.evaluate(samples=samples).corpus[METRIC]
rows[baseline] = res
d = res.details
if res.score is None:
print(f"| {baseline} | — | — | 0 |")
else:
print(f"| {baseline} | {100 * res.score:.1f}% | {d['wins']}/{d['losses']}/{d['ties']} | {d['n']} |")
if args.output is not None:
payload = {b: {"score": r.score, "details": r.details} for b, r in rows.items()}
args.output.write_text(json.dumps(payload, indent=2))
print(f"\nWrote {args.output}")
if __name__ == "__main__":
main()
@@ -1,95 +0,0 @@
"""NVFP4 + Attn-QAT (modified SageAttention3) inference on Blackwell.
Runs Wan2.1-T2V-1.3B fully in 4-bit: NVFP4 linear layers (activations
quantized on the fly) together with the modified SageAttention3 FP4 attention
backend (``ATTN_QAT_INFER``). This is the inference half of the
Quantization-Aware Distillation (QAD) recipe.
Requirements:
- RTX 5090 / consumer Blackwell (sm_120a). The attn_qat_infer kernel hard
gates on sm_120; on other GPUs it falls back to Flash Attention.
- The attn_qat_infer kernel built into fastvideo-kernel (see #1455) and
flashinfer for the NVFP4 linear matmuls.
Usage:
python nvfp4_qat_wan2_1_1_3b.py # NVFP4 linear + Attn-QAT attn
python nvfp4_qat_wan2_1_1_3b.py --bf16 # BF16 baseline
"""
import argparse
import os
import time
OUTPUT_PATH = "video_samples"
def main():
parser = argparse.ArgumentParser(description="NVFP4 + Attn-QAT video generation")
parser.add_argument("--bf16", action="store_true",
help="BF16 baseline (no NVFP4 linear, default attention)")
parser.add_argument("--model", default="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
help="Model path or HuggingFace ID")
parser.add_argument("--quant-method", default="nvfp4_qat", choices=["nvfp4_qat", "NVFP4"],
help="Linear quantization config. Wan-2.1 uses nvfp4_qat (matches its "
"to_q/k/v/out + ffn layers); NVFP4 is LTX2-specific and will NOT "
"quantize Wan.")
parser.add_argument("--compile", action="store_true", help="Enable torch.compile for the DiT")
parser.add_argument("--num_gpus", type=int, default=1)
parser.add_argument("--infer_steps", type=int, default=50)
args = parser.parse_args()
# The attention backend is selected via env var before the engine starts.
# ATTN_QAT_INFER -> AttnQatInferBackend (modified SageAttention3 FP4).
if not args.bf16:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "ATTN_QAT_INFER"
# Import after the env var so the platform picks up the selection.
from fastvideo import VideoGenerator
from fastvideo.layers.quantization import get_quantization_config
mode = "bf16" if args.bf16 else args.quant_method
if args.compile:
mode += "_compile"
print(f"Mode: {mode.upper()}")
# transformer_quant needs a QuantizationConfig *instance* — the bare string
# is not resolved on the from_pretrained kwarg path.
extra = {} if args.bf16 else {"transformer_quant": get_quantization_config(args.quant_method)()}
generator = VideoGenerator.from_pretrained(
args.model,
num_gpus=args.num_gpus,
use_fsdp_inference=args.bf16,
dit_cpu_offload=False,
vae_cpu_offload=True,
text_encoder_cpu_offload=True,
enable_torch_compile=args.compile,
**extra,
)
prompt = (
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
"wide with interest. The playful yet serene atmosphere is complemented by soft "
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
)
n_warmup = 2 if args.compile else 1
for _ in range(n_warmup):
generator.generate(request={"prompt": prompt, "sampling": {"num_inference_steps": 2},
"output": {"save_video": False}})
os.makedirs(OUTPUT_PATH, exist_ok=True)
start = time.time()
generator.generate(request={
"prompt": prompt,
"sampling": {"num_inference_steps": args.infer_steps},
"output": {"save_video": True, "output_path": os.path.join(OUTPUT_PATH, f"raccoon_{mode}.mp4")},
})
elapsed = time.time() - start
print(f"[{mode.upper()}] {args.infer_steps} steps in {elapsed:.2f}s "
f"({args.infer_steps / elapsed:.2f} it/s)")
generator.shutdown()
if __name__ == "__main__":
main()
@@ -1,52 +0,0 @@
{
"data": [
{
"caption": "gold tip pyramid in the night, extremely detailed , rain, stars"
},
{
"caption": "a fruit stacking in the shape of a dog, stock image, shutterstock"
},
{
"caption": "Landscape, By Lee madgwick, by Luis Royo, by Louise nevelson"
},
{
"caption": "A colorful poster that says \"philo is a weird\""
},
{
"caption": "Danish male with blue eyes, realistic, viking"
},
{
"caption": "futuristic, cityscape, flying cars, neon lights, towering skyscrapers, glowing purple sky."
},
{
"caption": "a crow with cameras for eyes, sitting on a mans shoulder, anime, studio ghibli, fantasy, fairytale, sketch, digital art, watercolor, dnd, rustic, professional photograph, medieval, hd, 4k"
},
{
"caption": "a background image mixing the matrix and AI"
},
{
"caption": "Golden sunset, a bright orange and yellow sky is visible, lit up by the setting sun, the horizon is a mix of bright colors and deep shadows"
},
{
"caption": "Grim reaper playing an electric guitar"
},
{
"caption": "an epic view of a demonic Rose-ringed parakeet cyborg inside an ironmaiden robot,wearing a noble robe,large view,a surrealist painting, aralan bean and Philippe Druillet,hiromu arakawa,volumetric lighting,detailed shadows"
},
{
"caption": "Ben Shapiro as the cover of ministry's filth pig album, but covered in milk"
},
{
"caption": "king charles spaniel with , ethereal, midjourney style lighting and shadows, insanely detailed, 8k, photorealistic"
},
{
"caption": "A website for a party resort service"
},
{
"caption": "full shot of a steampunk horse"
},
{
"caption": "60s psycedelic spiritual jazz album art"
}
]
}
-1
View File
@@ -87,7 +87,6 @@ training:
# --- training.data [TYPED] -> DataConfig ---
data:
data_path: data/my_dataset # default: ""
preprocessed_data_type: t2v # default: "t2v" ("text_only" for simulate-only DMD text prompts)
train_batch_size: 1 # default: 1
dataloader_num_workers: 4 # default: 0
training_cfg_rate: 0.1 # default: 0.0
@@ -1,116 +0,0 @@
# DiffusionNFT multi-reward single-frame RL: Wan 2.1 T2V 1.3B on text-only PickScore prompts.
#
# Single-frame RL is represented as a one-latent-frame Wan run:
# num_latent_t: 1
# num_frames: 1
#
# The method trains the full transformer (no LoRA) and keeps an old-policy
# transformer plus a frozen reference transformer, matching the non-LoRA
# DiffusionNFT loss path.
models:
student:
_target_: fastvideo.train.models.wan.WanModel
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
trainable: true
old:
_target_: fastvideo.train.models.wan.WanModel
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
trainable: false
disable_custom_init_weights: true
reference:
_target_: fastvideo.train.models.wan.WanModel
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
trainable: false
disable_custom_init_weights: true
method:
_target_: fastvideo.train.methods.rl.diffusion_nft.DiffusionNFTMethod
reward_fn:
pickscore: 1.0
clipscore: 1.0
sampling:
num_steps: 25
scheduler: flow_match_euler
trajectory: ode
flow_shift: inherit
validation:
every_steps: 10
num_steps: 40
num_prompts: 16
batch_size: 16
log_samples: true
seed: 42
# Null reuses training.data.data_path. Override this with a held-out
# preprocessed parquet path when one is available.
data_path:
# DiffusionNFT sd3_multi_reward on 4 GPUs resolves to per-GPU sample batch
# size 6, 48 sample batches per outer epoch, and grad accumulation 48.
sample_train_batch_size: 6
train_batch_size: 6
num_batches_per_epoch: 48
num_video_per_prompt: 24
num_inner_epochs: 1
timestep_fraction: 0.99
beta: 0.1
kl_beta: 0.0001
decay_type: 1
adv_mode: all
adv_clip_max: 5
max_grad_norm: 1.0
ema:
enabled: true
decay: 0.9
update_after_step: 0
validation: true
terminal_progress: true
training:
distributed:
num_gpus: 4
sp_size: 1
tp_size: 1
hsdp_replicate_dim: 1
hsdp_shard_dim: 4
data:
data_path: data/pickscore_text_only_preprocessed
preprocessed_data_type: text_only
dataloader_num_workers: 0
train_batch_size: 1
training_cfg_rate: 0.0
seed: 42
num_latent_t: 1
num_height: 448
num_width: 832
num_frames: 1
optimizer:
learning_rate: 3.0e-5
betas: [0.9, 0.999]
weight_decay: 0.0001
lr_scheduler: constant
lr_warmup_steps: 0
loop:
max_train_steps: 100000
gradient_accumulation_steps: 48
checkpoint:
output_dir: outputs/wan2.1_diffusion_nft_pick_clip
training_state_checkpointing_steps: 30
checkpoints_total_limit: 3
tracker:
project_name: diffusion_nft_wan
run_name: wan2.1_diffusion_nft_pick_clip
model:
enable_gradient_checkpointing_type: full
pipeline:
flow_shift: 8
@@ -1,119 +0,0 @@
# MixKit training data (QAD 5090 recipe)
The QAD 5090 models are distilled from Wan2.1-T2V-1.3B on a MixKit subset at
**480×832, 77 frames, 16 fps**. FastVideo training consumes **Parquet** shards of
precomputed VAE latents + text embeddings (no text encoder / VAE needed at train
time).
## Option A — download the preprocessed data (recommended)
The encoded dataset is published on the Hugging Face Hub, ready to train:
```bash
# from the repo root
bash examples/training/finetune/wan_t2v_1.3B/mixkit/download_mixkit_data.sh
```
This pulls [`weizhou03/HD-Mixkit-Finetune-Wan`](https://huggingface.co/datasets/weizhou03/HD-Mixkit-Finetune-Wan)
into `data/HD-Mixkit-Finetune-Wan/`:
```
data/HD-Mixkit-Finetune-Wan/
├── combined_parquet_dataset/ # training shards -> point --data_path here
│ └── worker_0/data_chunk_*.parquet
└── validation_parquet_dataset/ # validation shards
└── worker_0/data_chunk_0.parquet
```
Each Parquet row holds the VAE latent bytes + text-embedding bytes (plus
shape/dtype metadata), matching FastVideo's standard preprocessing output.
## Option B — build the Parquet from raw videos
If you want to reproduce the encoding from your own MixKit videos, arrange them as
a `merged` dataset (videos + a captions JSON), then run FastVideo's standard
preprocessing to VAE-encode and text-embed them into Parquet:
```bash
GPU_NUM=2
torchrun --nproc_per_node=$GPU_NUM \
-m fastvideo.pipelines.preprocess.v1_preprocessing_new \
--model_path "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" \
--mode preprocess \
--workload_type t2v \
--preprocess.video_loader_type torchvision \
--preprocess.dataset_type merged \
--preprocess.dataset_path "data/mixkit_raw/" \
--preprocess.dataset_output_dir "data/HD-Mixkit-Finetune-Wan/" \
--preprocess.max_height 480 \
--preprocess.max_width 832 \
--preprocess.num_frames 77 \
--preprocess.train_fps 16 \
--preprocess.samples_per_file 8
```
The raw videos are full-HD MixKit clips (≈1080p/30fps); preprocessing resizes to
480×832, resamples to 16 fps, and extracts 77 frames per clip. See
[`docs/training/data_preprocess.md`](../../../../../docs/training/data_preprocess.md)
for the full parameter reference.
## Train (QAT finetune)
With the data in place, run the quantization-aware finetune. The 4-bit attention
path is **config-driven** — selected purely by an env var, no monkey-patching:
```bash
bash examples/training/finetune/wan_t2v_1.3B/mixkit/finetune_qat.sh
# or point at your own parquet dir / GPU count:
NUM_GPUS=4 bash .../mixkit/finetune_qat.sh data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/
```
`FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN` routes attention through the
fake-quantized Triton kernel (straight-through estimator), so the DiT learns to
absorb FP4 attention error. This kernel is Triton, so it runs on both `sm_100`
(B200/GB200) and `sm_120` (RTX 5090).
## Train stage 2 (QAT DMD distillation to 3 steps)
Distill the QAT-finetuned generator down to **3 sampling steps**. Only the
generator is quantized (Attn-QAT); the teacher (`real_score`) and critic
(`fake_score`) stay full precision. This is enforced in the loader
(`component_loader.py`, via the `_loading_teacher_critic_model` flag), so the
same global `ATTN_QAT_TRAIN` env reaches **only** the generator — no per-model
flags or monkey-patching.
```bash
# generator init = the stage-1 finetune checkpoint
bash examples/training/finetune/wan_t2v_1.3B/mixkit/distill_dmd_qat.sh \
data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/ \
checkpoints/wan_t2v_qat_finetune/checkpoint-2000/transformer/diffusion_pytorch_model.safetensors
```
DMD runs a double loop (critic every step, generator every
`generator_update_interval`), and validation samples the distilled student at
3 steps — the final 4-bit-attention model.
## Inference (NVFP4 4-bit linear)
For Wan-2.1, enable the FP4 linear layers with the **`nvfp4_qat`** quantization
config (it matches Wan's `to_q/k/v/out` + `ffn` layers; the plain `NVFP4` config
is LTX2-specific and will not quantize Wan):
```python
from fastvideo import VideoGenerator
from fastvideo.layers.quantization import get_quantization_config
gen = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers", num_gpus=1,
transformer_quant=get_quantization_config("nvfp4_qat")(), # a config instance, not the string
use_fsdp_inference=False,
)
gen.generate(request={"prompt": "...", "output": {"save_video": True}})
```
The loader converts the tagged linear weights to FP4 at load time
(`_maybe_convert_model_to_nvfp4`). Combine with
`FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER` on an RTX 5090 (`sm_120`) for the
full 4-bit path; on other GPUs the attention falls back to Flash while the FP4
linear layers still run. `flashinfer` (and a host C++ compiler for its FP4
kernel JIT) are required.
@@ -1,51 +0,0 @@
#!/bin/bash
# QAD recipe stage 2 — quantization-aware DMD distillation of Wan2.1-T2V-1.3B
# down to 3 sampling steps, with the GENERATOR in fake-quant Attn-QAT and the
# teacher (real_score) + critic (fake_score) at full precision.
#
# Generator-only QAT is config-driven: FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN
# is applied to the generator only, because the loader masks it (and the
# nvfp4_qat quant) for the teacher/critic via the `_loading_teacher_critic_model`
# flag (see fastvideo/models/loader/component_loader.py). No monkey-patching.
#
# Init the generator from the stage-1 finetune checkpoint (finetune_qat.sh).
# Data: run download_mixkit_data.sh first.
#
# Verified end-to-end on Blackwell (GB200/sm_100): generator loads with
# ATTN_QAT_TRAIN while teacher/critic load full-precision; the DMD double loop
# runs (generator updates every generator_update_interval steps, critic every
# step), 3-step validation generates videos, checkpoint saved.
set -euo pipefail
export FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN # generator-only (loader-gated)
export WANDB_MODE=${WANDB_MODE:-online}
export TOKENIZERS_PARALLELISM=false
BASE="Wan-AI/Wan2.1-T2V-1.3B-Diffusers"
DATA_DIR=${1:-"data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/"}
# Generator init weights = the stage-1 QAT-finetune checkpoint.
INIT_WEIGHTS=${2:-"checkpoints/wan_t2v_qat_finetune/checkpoint-2000/transformer/diffusion_pytorch_model.safetensors"}
VALIDATION_FILE="$(dirname "$0")/../crush_smol/validation.json"
NUM_GPUS=${NUM_GPUS:-4}
torchrun --nnodes 1 --nproc_per_node "${NUM_GPUS}" \
fastvideo/training/wan_distillation_pipeline.py \
--num_gpus "${NUM_GPUS}" --sp_size 1 --tp_size 1 \
--hsdp_replicate_dim "${NUM_GPUS}" --hsdp_shard_dim 1 \
--model_path "${BASE}" --pretrained_model_name_or_path "${BASE}" \
--real_score_model_path "${BASE}" --fake_score_model_path "${BASE}" \
--init_weights_from_safetensors "${INIT_WEIGHTS}" \
--data_path "${DATA_DIR}" --dataloader_num_workers 4 \
--max_train_steps 2000 --train_batch_size 1 --train_sp_batch_size 1 \
--gradient_accumulation_steps 1 \
--num_latent_t 20 --num_height 480 --num_width 832 --num_frames 77 \
--enable_gradient_checkpointing_type full \
--log_validation --validation_dataset_file "${VALIDATION_FILE}" \
--validation_steps 200 --validation_sampling_steps 3 --validation_guidance_scale 6.0 \
--learning_rate 2e-6 --mixed_precision bf16 --weight_decay 0.01 --max_grad_norm 1.0 \
--weight_only_checkpointing_steps 500 --training_state_checkpointing_steps 500 \
--tracker_project_name wan_t2v_distill_dmd_qat \
--output_dir checkpoints/wan_t2v_distill_dmd_qat \
--inference_mode False --dit_precision fp32 --ema_start_step 0 --training_cfg_rate 0.0 \
--generator_update_interval 5 --real_score_guidance_scale 2.0 \
--dmd_denoising_steps '1000,757,522' --min_timestep_ratio 0.02 --max_timestep_ratio 0.98
@@ -1,22 +0,0 @@
#!/bin/bash
# Download the preprocessed MixKit finetune dataset used for the QAD 5090 recipe.
#
# This is the MixKit subset already VAE-encoded (Wan2.1-T2V-1.3B) and text-embedded
# into Parquet shards, so it can be fed straight to training with no further
# preprocessing. To build the Parquet from raw videos yourself, see README.md.
#
# Usage (run from the repo root):
# bash examples/training/finetune/wan_t2v_1.3B/mixkit/download_mixkit_data.sh [DATA_ROOT]
set -euo pipefail
DATA_ROOT=${1:-data/HD-Mixkit-Finetune-Wan}
python scripts/huggingface/download_hf.py \
--repo_id "weizhou03/HD-Mixkit-Finetune-Wan" \
--local_dir "${DATA_ROOT}" \
--repo_type "dataset"
echo "Done."
echo " Train data: ${DATA_ROOT}/combined_parquet_dataset"
echo " Validation data: ${DATA_ROOT}/validation_parquet_dataset"
echo "Point your training script's data path at the combined_parquet_dataset directory."
@@ -1,44 +0,0 @@
#!/bin/bash
# QAD recipe — quantization-aware finetune of Wan2.1-T2V-1.3B with fake-quant
# (Attn-QAT) attention.
#
# The 4-bit attention path is selected purely by env var (config-driven, no
# monkey-patching): FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN routes attention
# through the fake-quantized Triton kernel (straight-through estimator), so the
# DiT learns to absorb FP4 attention error instead of fighting it.
#
# Data: run download_mixkit_data.sh first (preprocessed Parquet).
#
# Verified end-to-end on Blackwell (GB200/sm_100): the ATTN_QAT_TRAIN backend is
# selected (not a fallback), forward+backward run, loss/grad are healthy, and
# validation generates videos. The kernel is Triton so it runs on sm_100 and
# sm_120 alike (the FP4 inference kernel, by contrast, is sm_120-only).
set -euo pipefail
export FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN # <-- enables Attn-QAT training
export WANDB_MODE=${WANDB_MODE:-online}
export TOKENIZERS_PARALLELISM=false
MODEL_PATH="Wan-AI/Wan2.1-T2V-1.3B-Diffusers"
DATA_DIR=${1:-"data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/"}
VALIDATION_FILE="$(dirname "$0")/../crush_smol/validation.json"
NUM_GPUS=${NUM_GPUS:-4}
torchrun --nnodes 1 --nproc_per_node "${NUM_GPUS}" \
fastvideo/training/wan_training_pipeline.py \
--num_gpus "${NUM_GPUS}" --sp_size "${NUM_GPUS}" --tp_size 1 \
--hsdp_replicate_dim 1 --hsdp_shard_dim "${NUM_GPUS}" \
--model_path "${MODEL_PATH}" --pretrained_model_name_or_path "${MODEL_PATH}" \
--data_path "${DATA_DIR}" --dataloader_num_workers 1 \
--max_train_steps 2000 --train_batch_size 1 --train_sp_batch_size 1 \
--gradient_accumulation_steps 1 \
--num_latent_t 20 --num_height 480 --num_width 832 --num_frames 77 \
--enable_gradient_checkpointing_type full \
--log_validation --validation_dataset_file "${VALIDATION_FILE}" \
--validation_steps 200 --validation_sampling_steps 50 --validation_guidance_scale 3.0 \
--learning_rate 5e-5 --mixed_precision bf16 --weight_decay 1e-4 --max_grad_norm 1.0 \
--weight_only_checkpointing_steps 500 --training_state_checkpointing_steps 500 \
--tracker_project_name wan_t2v_qat_finetune --output_dir checkpoints/wan_t2v_qat_finetune \
--inference_mode False --training_cfg_rate 0.1 --not_apply_cfg_solver \
--dit_precision fp32 --num_euler_timesteps 50 --ema_start_step 0 \
--multi_phased_distill_schedule "4000-1"
-135
View File
@@ -50,21 +50,6 @@ include_directories(
set(FASTVIDEO_KERNEL_BUILD_TK "AUTO" CACHE STRING "Build ThunderKittens kernels: AUTO/ON/OFF")
set_property(CACHE FASTVIDEO_KERNEL_BUILD_TK PROPERTY STRINGS AUTO ON OFF)
set(_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT "AUTO")
if(DEFINED FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3 AND NOT DEFINED CACHE{FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER})
set(_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT "${FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3}")
endif()
set(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER "${_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT}" CACHE STRING
"Build attn_qat_infer Blackwell inference kernels: AUTO/ON/OFF")
set_property(CACHE FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER PROPERTY STRINGS AUTO ON OFF)
if(DEFINED FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3)
message(DEPRECATION
"FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3 is deprecated. "
"Use FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER instead.")
endif()
# Prefer environment variable (used by CI) if CMake var is not explicitly set.
if(NOT DEFINED TORCH_CUDA_ARCH_LIST AND DEFINED ENV{TORCH_CUDA_ARCH_LIST})
set(TORCH_CUDA_ARCH_LIST "$ENV{TORCH_CUDA_ARCH_LIST}")
@@ -72,7 +57,6 @@ endif()
message(STATUS "TORCH_CUDA_ARCH_LIST (cmake/env): ${TORCH_CUDA_ARCH_LIST}")
message(STATUS "FASTVIDEO_KERNEL_BUILD_TK: ${FASTVIDEO_KERNEL_BUILD_TK}")
message(STATUS "FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER: ${FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER}")
set(ENABLE_TK_KERNELS OFF)
if(FASTVIDEO_KERNEL_BUILD_TK STREQUAL "ON")
@@ -107,54 +91,6 @@ else()
message(STATUS "ThunderKittens kernels: DISABLED (will use Triton fallbacks at runtime)")
endif()
set(ENABLE_ATTN_QAT_INFER OFF)
if(GPU_BACKEND STREQUAL "ROCM")
message(STATUS "attn_qat_infer kernels: DISABLED (ROCm build)")
else()
set(_WANTS_ATTN_QAT_INFER OFF)
if(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER STREQUAL "ON")
set(_WANTS_ATTN_QAT_INFER ON)
elseif(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER STREQUAL "AUTO")
if(TORCH_CUDA_ARCH_LIST)
string(REGEX MATCH
"(^|[; ,])((12\\.0a)|(120a)|(sm_120a))([; ,]|$)"
_HAS_120A "${TORCH_CUDA_ARCH_LIST}")
if(_HAS_120A)
set(_WANTS_ATTN_QAT_INFER ON)
endif()
else()
execute_process(
COMMAND "${Python_EXECUTABLE}" -c
"import torch; print('1' if (torch.cuda.is_available() and torch.version.cuda and torch.cuda.get_device_capability()[0] >= 12) else '0')"
OUTPUT_VARIABLE _LOCAL_HAS_BLACKWELL
OUTPUT_STRIP_TRAILING_WHITESPACE
ERROR_QUIET
)
if(_LOCAL_HAS_BLACKWELL STREQUAL "1")
set(_WANTS_ATTN_QAT_INFER ON)
endif()
endif()
endif()
if(_WANTS_ATTN_QAT_INFER)
if(CUDAToolkit_VERSION VERSION_LESS 12.8)
message(WARNING
"attn_qat_infer kernels require CUDA Toolkit 12.8+. "
"Skipping because CUDAToolkit_VERSION=${CUDAToolkit_VERSION}.")
else()
set(ENABLE_ATTN_QAT_INFER ON)
endif()
endif()
if(ENABLE_ATTN_QAT_INFER)
message(STATUS "attn_qat_infer kernels: ENABLED")
else()
message(STATUS
"attn_qat_infer kernels: DISABLED "
"(requires CUDA 12.8+ and Blackwell sm_120a)")
endif()
endif()
# Always try to build the extension if CUDA is available, but conditionally add sources/flags
set(BUILD_CXX_KERNELS ON)
@@ -247,74 +183,3 @@ if(BUILD_CXX_KERNELS)
install(TARGETS fastvideo_kernel_ops LIBRARY DESTINATION fastvideo_kernel/_C)
endif()
if(ENABLE_ATTN_QAT_INFER)
set(ATTN_QAT_INFER_DIR ${CMAKE_SOURCE_DIR}/attn_qat_infer)
set(ATTN_QAT_INFER_INCLUDE_DIRS
${ATTN_QAT_INFER_DIR}
${CMAKE_SOURCE_DIR}/include/cutlass/include
${CMAKE_SOURCE_DIR}/include/cutlass/tools/util/include
${TORCH_INCLUDE_DIRS}
)
set(ATTN_QAT_INFER_CUDA_FLAGS
"-O3"
"-std=c++17"
"-U__CUDA_NO_HALF_OPERATORS__"
"-U__CUDA_NO_HALF_CONVERSIONS__"
"-U__CUDA_NO_BFLOAT16_OPERATORS__"
"-U__CUDA_NO_BFLOAT16_CONVERSIONS__"
"-U__CUDA_NO_BFLOAT162_OPERATORS__"
"-U__CUDA_NO_BFLOAT162_CONVERSIONS__"
"--expt-relaxed-constexpr"
"--expt-extended-lambda"
"--use_fast_math"
"--ptxas-options=--verbose,--warn-on-local-memory-usage"
"-lineinfo"
"-DCUTLASS_DEBUG_TRACE_LEVEL=0"
"-DNDEBUG"
"-DQBLKSIZE=128"
"-DKBLKSIZE=128"
"-DCTA256"
"-DDQINRMEM"
)
Python_add_library(fp4attn_cuda MODULE WITH_SOABI
attn_qat_infer/blackwell/api.cu
)
target_include_directories(fp4attn_cuda PRIVATE ${ATTN_QAT_INFER_INCLUDE_DIRS})
target_compile_definitions(fp4attn_cuda PRIVATE TORCH_EXTENSION_NAME=fp4attn_cuda)
target_compile_options(fp4attn_cuda PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:-O3 -std=c++17>
$<$<COMPILE_LANGUAGE:CUDA>:${ATTN_QAT_INFER_CUDA_FLAGS}>
)
set_target_properties(fp4attn_cuda PROPERTIES
CUDA_ARCHITECTURES "120a"
CXX_STANDARD 17
CUDA_STANDARD 17
)
target_link_libraries(fp4attn_cuda PRIVATE ${TORCH_LIBRARIES} CUDA::cudart CUDA::cuda_driver)
Python_add_library(fp4quant_cuda MODULE WITH_SOABI
attn_qat_infer/quantization/fp4_quantization_4d.cu
)
target_include_directories(fp4quant_cuda PRIVATE ${ATTN_QAT_INFER_INCLUDE_DIRS})
target_compile_definitions(fp4quant_cuda PRIVATE TORCH_EXTENSION_NAME=fp4quant_cuda)
target_compile_options(fp4quant_cuda PRIVATE
$<$<COMPILE_LANGUAGE:CXX>:-O3 -std=c++17>
$<$<COMPILE_LANGUAGE:CUDA>:${ATTN_QAT_INFER_CUDA_FLAGS}>
)
set_target_properties(fp4quant_cuda PROPERTIES
CUDA_ARCHITECTURES "120a"
CXX_STANDARD 17
CUDA_STANDARD 17
)
target_link_libraries(fp4quant_cuda PRIVATE ${TORCH_LIBRARIES} CUDA::cudart CUDA::cuda_driver)
if(TORCH_PYTHON_LIBRARY_PATH)
target_link_libraries(fp4attn_cuda PRIVATE "${TORCH_PYTHON_LIBRARY_PATH}")
target_link_libraries(fp4quant_cuda PRIVATE "${TORCH_PYTHON_LIBRARY_PATH}")
endif()
install(TARGETS fp4attn_cuda LIBRARY DESTINATION .)
install(TARGETS fp4quant_cuda LIBRARY DESTINATION .)
endif()
-1
View File
@@ -2,6 +2,5 @@ include LICENSE
include README.md
include pyproject.toml
recursive-include python/fastvideo_kernel *.py
recursive-include attn_qat_infer *.py *.cu *.cuh *.cpp *.h
recursive-include csrc *.cu *.cuh *.cpp *.h
recursive-include include/tk *.cu *.cuh *.cpp *.h *.src
@@ -1,16 +0,0 @@
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
from .api import sageattn_blackwell
-189
View File
@@ -1,189 +0,0 @@
# Modified from the original SageATtention3 code
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
import torch
import triton
import triton.language as tl
import torch.nn.functional as F
from typing import Tuple
from torch.nn.functional import scaled_dot_product_attention as sdpa
import fp4attn_cuda
import fp4quant_cuda
# Centralized block size configuration for sageattn_blackwell kernels
# These should match the values in fastvideo/attention/backends/sageattn/blackwell/block_config.h
BLOCK_M = 128 # Block size for M dimension (query sequence length)
BLOCK_N = 128 # Block size for N dimension (key/value sequence length)
@triton.jit
def group_mean_kernel(
q_ptr,
q_out_ptr,
qm_out_ptr,
B, H, L, D: tl.constexpr,
stride_qb, stride_qh, stride_ql, stride_qd,
stride_qmb, stride_qmh, stride_qml, stride_qmd,
GROUP_SIZE: tl.constexpr
):
pid_b = tl.program_id(0)
pid_h = tl.program_id(1)
pid_group = tl.program_id(2)
group_start = pid_group * GROUP_SIZE
offsets = group_start + tl.arange(0, GROUP_SIZE)
q_offsets = pid_b * stride_qb + pid_h * stride_qh + offsets[:, None] * stride_ql + tl.arange(0, D)[None, :] * stride_qd
q_group = tl.load(q_ptr + q_offsets)
qm_group = tl.sum(q_group, axis=0) / GROUP_SIZE
q_group = q_group - qm_group
tl.store(q_out_ptr + q_offsets, q_group)
qm_offset = pid_b * stride_qmb + pid_h * stride_qmh + pid_group * stride_qml + tl.arange(0, D) * stride_qmd
tl.store(qm_out_ptr + qm_offset, qm_group)
def triton_group_mean(q: torch.Tensor):
B, H, L, D = q.shape
GROUP_SIZE = BLOCK_M
num_groups = L // GROUP_SIZE
q_out = torch.empty_like(q) # [B, H, L, D]
qm = torch.empty(B, H, num_groups, D, device=q.device, dtype=q.dtype)
grid = (B, H, num_groups)
group_mean_kernel[grid](
q, q_out, qm,
B, H, L, D,
q.stride(0), q.stride(1), q.stride(2), q.stride(3),
qm.stride(0), qm.stride(1), qm.stride(2), qm.stride(3),
GROUP_SIZE=GROUP_SIZE
)
return q_out, qm
def preprocess_qkv(q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, per_block_mean: bool = True, enable_smoothing_q: bool = False, enable_smoothing_k: bool = False):
def pad_to_block_size(x):
L = x.size(2)
pad_len = (BLOCK_M - L % BLOCK_M) % BLOCK_M
if pad_len == 0:
return x.contiguous()
return F.pad(x, (0, 0, 0, pad_len), value=0).contiguous()
if enable_smoothing_k:
k -= k.mean(dim=-2, keepdim=True)
q, k, v = map(lambda x: pad_to_block_size(x), [q, k, v])
if per_block_mean and enable_smoothing_q:
q, qm = triton_group_mean(q)
elif enable_smoothing_q:
qm = q.mean(dim=-2, keepdim=True)
q = q - qm
if enable_smoothing_q:
delta_s = torch.matmul(qm, k.transpose(-2, -1)).to(torch.float32).contiguous()
else: # used to disable q smoothing
B, H, L, D = q.shape
delta_s = torch.zeros((B, H, L // BLOCK_M, k.shape[2]), device=q.device, dtype=torch.float32)
return q, k, v, delta_s
def scale_and_quant_fp4(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
assert x.ndim == 4
B, H, N, D = x.shape
packed_fp4 = torch.empty((B, H, N, D // 2), device=x.device, dtype=torch.uint8)
fp8_scale = torch.empty((B, H, N, D // 16), device=x.device, dtype=torch.float8_e4m3fn)
fp4quant_cuda.scaled_fp4_quant(x, packed_fp4, fp8_scale, 1)
return packed_fp4, fp8_scale
def scale_and_quant_fp4_permute(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
assert x.ndim == 4
B, H, N, D = x.shape
packed_fp4 = torch.empty((B, H, N, D // 2), device=x.device, dtype=torch.uint8)
fp8_scale = torch.empty((B, H, N, D // 16), device=x.device, dtype=torch.float8_e4m3fn)
fp4quant_cuda.scaled_fp4_quant_permute(x, packed_fp4, fp8_scale, 1)
return packed_fp4, fp8_scale
def scale_and_quant_fp4_transpose(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
assert x.ndim == 4
B, H, N, D = x.shape
packed_fp4 = torch.empty((B, H, D, N // 2), device=x.device, dtype=torch.uint8)
fp8_scale = torch.empty((B, H, D, N // 16), device=x.device, dtype=torch.float8_e4m3fn)
fp4quant_cuda.scaled_fp4_quant_trans(x, packed_fp4, fp8_scale, 1)
return packed_fp4, fp8_scale
def blockscaled_fp4_attn(qlist: Tuple,
klist: Tuple,
vlist: Tuple,
delta_s: torch.Tensor,
KL: int,
is_causal: bool = False,
per_block_mean: bool = True,
is_bf16: bool = True,
single_level_p_quant: bool = False,
sm_scale: float | None = None
):
softmax_scale = sm_scale if sm_scale is not None else (qlist[0].shape[-1] * 2) ** (-0.5)
return fp4attn_cuda.fwd(qlist[0], klist[0], vlist[0], qlist[1], klist[1], vlist[1], delta_s, KL, None, softmax_scale, is_causal, per_block_mean, is_bf16, single_level_p_quant)
def sageattn_blackwell(q, k, v, attn_mask = None, is_causal = False, per_block_mean = True, single_level_p_quant = True, sm_scale: float | None = None, **kwargs):
"""
SageAttention3 Blackwell kernel for FP4 attention.
Args:
q: Query tensor [B, H, L, D]
k: Key tensor [B, H, L, D]
v: Value tensor [B, H, L, D]
attn_mask: Attention mask (not used)
is_causal: Whether to use causal masking
per_block_mean: Whether to use per-block mean for Q smoothing
single_level_p_quant: If True, use single-level quantization: s_P2, P̂_2 = φ(P̃) directly
(standard per-block FP4 quantization like V, no s_P1).
If False (default), use two-level quantization:
s_P1 = rowmax(P̃)/(448×6), then s_P2, P̂_2 = φ(P̃/s_P1).
sm_scale: Softmax scale to pass through to the CUDA kernel. If None,
defaults to the kernel's 1/sqrt(D) scale.
**kwargs: Additional arguments (ignored)
Returns:
Output tensor [B, H, L, D]
"""
if q.size(-1) >= 256:
print(f"Unsupported Headdim {q.size(-1)}")
return sdpa(q, k, v, is_causal = is_causal)
QL = q.size(2)
KL = k.size(2)
is_bf16 = q.dtype == torch.bfloat16
q, k, v, delta_s = preprocess_qkv(q, k, v, per_block_mean)
qlist_from_cuda = scale_and_quant_fp4(q)
klist_from_cuda = scale_and_quant_fp4_permute(k)
vlist_from_cuda = scale_and_quant_fp4_transpose(v)
o_fp4 = blockscaled_fp4_attn(
qlist_from_cuda,
klist_from_cuda,
vlist_from_cuda,
delta_s,
KL,
is_causal,
per_block_mean,
is_bf16,
single_level_p_quant,
sm_scale
)[0][:, :, :QL, :].contiguous()
return o_fp4
@@ -1 +0,0 @@
__version__ = "3.0.0.b1"
@@ -1,347 +0,0 @@
// Modified from the original SageAttention3 code
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
// Include these 2 headers instead of torch/extension.h since we don't need all of the torch headers.
#include <torch/python.h>
#include <torch/nn/functional.h>
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>
#include <cutlass/numeric_types.h>
#include "params.h"
#include "launch.h"
#include "static_switch.h"
#include "block_config.h"
#define CHECK_DEVICE(x) TORCH_CHECK(x.is_cuda(), #x " must be on CUDA")
#define CHECK_SHAPE(x, ...) TORCH_CHECK(x.sizes() == torch::IntArrayRef({__VA_ARGS__}), #x " must have shape (" #__VA_ARGS__ ")")
#define CHECK_CONTIGUOUS(x) TORCH_CHECK(x.is_contiguous(), #x " must be contiguous")
void set_params_fprop(Flash_fwd_params &params,
// sizes
const size_t b,
const size_t seqlen_q,
const size_t seqlen_k,
const size_t unpadded_seqlen_k,
const size_t seqlen_q_rounded,
const size_t seqlen_k_rounded,
const size_t h,
const size_t h_k,
const size_t d,
const size_t d_rounded,
// device pointers
const at::Tensor q,
const at::Tensor k,
const at::Tensor v,
const at::Tensor delta_s,
at::Tensor out,
const at::Tensor sfq,
const at::Tensor sfk,
const at::Tensor sfv,
void *cu_seqlens_q_d,
void *cu_seqlens_k_d,
void *seqused_k,
void *p_d,
void *softmax_lse_d,
float p_dropout,
float softmax_scale,
int window_size_left,
int window_size_right,
bool per_block_mean,
bool is_bf16,
bool single_level_p_quant=false,
bool seqlenq_ngroups_swapped=false) {
// Reset the parameters
params = {};
// Set the pointers and strides.
params.q_ptr = q.data_ptr();
params.k_ptr = k.data_ptr();
params.v_ptr = v.data_ptr();
params.delta_s_ptr = delta_s.data_ptr();
params.sfq_ptr = sfq.data_ptr();
params.sfk_ptr = sfk.data_ptr();
params.sfv_ptr = sfv.data_ptr();
// All stride are in elements, not bytes.
params.q_row_stride = q.stride(-2) * 2;
params.k_row_stride = k.stride(-2) * 2;
params.v_row_stride = v.stride(-2) * 2;;
params.q_head_stride = q.stride(-3) * 2;
params.k_head_stride = k.stride(-3) * 2;
params.v_head_stride = v.stride(-3) * 2; // for packed q k v
params.ds_row_stride = delta_s.stride(-2);
params.ds_head_stride = delta_s.stride(-3);
params.sfq_row_stride = sfq.stride(-2);
params.sfk_row_stride = sfk.stride(-2);
params.sfv_row_stride = sfv.stride(-2);
params.sfq_head_stride = sfq.stride(-3);
params.sfk_head_stride = sfk.stride(-3);
params.sfv_head_stride = sfv.stride(-3);
params.o_ptr = out.data_ptr();
params.o_row_stride = out.stride(-2);
params.o_head_stride = out.stride(-3);
if (cu_seqlens_q_d == nullptr) {
params.q_batch_stride = q.stride(0) * 2;
params.k_batch_stride = k.stride(0) * 2;
params.v_batch_stride = v.stride(0) * 2;
params.ds_batch_stride = delta_s.stride(0);
params.sfq_batch_stride = sfq.stride(0);
params.sfk_batch_stride = sfk.stride(0);
params.sfv_batch_stride = sfv.stride(0);
params.o_batch_stride = out.stride(0);
if (seqlenq_ngroups_swapped) {
params.q_batch_stride *= seqlen_q;
params.o_batch_stride *= seqlen_q;
}
}
params.cu_seqlens_q = static_cast<int *>(cu_seqlens_q_d);
params.cu_seqlens_k = static_cast<int *>(cu_seqlens_k_d);
params.seqused_k = static_cast<int *>(seqused_k);
// P = softmax(QK^T)
params.p_ptr = p_d;
// Softmax sum
params.softmax_lse_ptr = softmax_lse_d;
// Set the dimensions.
params.b = b;
params.h = h;
params.h_k = h_k;
params.h_h_k_ratio = h / h_k;
params.seqlen_q = seqlen_q;
params.seqlen_k = seqlen_k;
params.unpadded_seqlen_k = unpadded_seqlen_k;
params.seqlen_q_rounded = seqlen_q_rounded;
params.seqlen_k_rounded = seqlen_k_rounded;
params.d = d;
params.d_rounded = d_rounded;
params.head_divmod = cutlass::FastDivmod(int(h));
// Set the different scale values.
params.scale_softmax = softmax_scale;
params.scale_softmax_log2 = softmax_scale * M_LOG2E;
__half scale_softmax_log2_half = __float2half(params.scale_softmax_log2);
__half2 scale_softmax_log2_half2 = __half2(scale_softmax_log2_half, scale_softmax_log2_half);
params.scale_softmax_log2_half2 = reinterpret_cast<uint32_t&>(scale_softmax_log2_half2);
// Set this to probability of keeping an element to simplify things.
params.p_dropout = 1.f - p_dropout;
// Convert p from float to int so we don't have to convert the random uint to float to compare.
// [Minor] We want to round down since when we do the comparison we use <= instead of <
// params.p_dropout_in_uint = uint32_t(std::floor(params.p_dropout * 4294967295.0));
// params.p_dropout_in_uint16_t = uint16_t(std::floor(params.p_dropout * 65535.0));
params.p_dropout_in_uint8_t = uint8_t(std::floor(params.p_dropout * 255.0));
params.rp_dropout = 1.f / params.p_dropout;
params.scale_softmax_rp_dropout = params.rp_dropout * params.scale_softmax;
TORCH_CHECK(p_dropout < 1.f);
#ifdef FLASHATTENTION_DISABLE_DROPOUT
TORCH_CHECK(p_dropout == 0.0f, "This flash attention build does not support dropout.");
#endif
// Causal is the special case where window_size_right == 0 and window_size_left < 0.
// Local is the more general case where window_size_right >= 0 or window_size_left >= 0.
params.is_causal = window_size_left < 0 && window_size_right == 0;
params.per_block_mean = per_block_mean;
if (per_block_mean) {
params.seqlen_s = seqlen_q;
} else {
params.seqlen_s = flash::BLOCK_M; // size of BLOCK_M
}
if (window_size_left < 0 && window_size_right >= 0) { window_size_left = seqlen_k; }
if (window_size_left >= 0 && window_size_right < 0) { window_size_right = seqlen_k; }
params.window_size_left = window_size_left;
params.window_size_right = window_size_right;
#ifdef FLASHATTENTION_DISABLE_LOCAL
TORCH_CHECK(params.is_causal || (window_size_left < 0 && window_size_right < 0),
"This flash attention build does not support local attention.");
#endif
params.is_seqlens_k_cumulative = true;
params.is_bf16 = is_bf16;
params.single_level_p_quant = single_level_p_quant;
#ifdef FLASHATTENTION_DISABLE_UNEVEN_K
TORCH_CHECK(d == d_rounded, "This flash attention build does not support headdim not being a multiple of 32.");
#endif
}
template<bool IsBF16>
void run_mha_fwd_dispatch_dtype(Flash_fwd_params &params, cudaStream_t stream) {
using OType = std::conditional_t<IsBF16, cutlass::bfloat16_t, cutlass::half_t>;
if (params.d == 64) {
run_mha_fwd_<cutlass::nv_float4_t<cutlass::float_e2m1_t>, 64, OType>(params, stream);
} else if (params.d == 128) {
run_mha_fwd_<cutlass::nv_float4_t<cutlass::float_e2m1_t>, 128, OType>(params, stream);
}
}
void run_mha_fwd(Flash_fwd_params &params, cudaStream_t stream, bool force_split_kernel = false) {
BOOL_SWITCH(params.is_bf16, IsBF16, ([&] {
run_mha_fwd_dispatch_dtype<IsBF16>(params, stream);
}));
}
std::vector<at::Tensor>
mha_fwd(at::Tensor &q, // batch_size x seqlen_q x num_heads x (head_size // 2)
const at::Tensor &k, // batch_size x seqlen_k x num_heads_k x (head_size // 2)
const at::Tensor &v, // batch_size x seqlen_k x num_heads_k x (head_size // 2)
const at::Tensor &sfq,
const at::Tensor &sfk,
const at::Tensor &sfv,
const at::Tensor &delta_s,
int unpadded_k,
c10::optional<at::Tensor> &out_, // batch_size x seqlen_q x num_heads x head_size
const float softmax_scale,
bool is_causal,
bool per_block_mean,
bool is_bf16,
bool single_level_p_quant=false // If true, use only per-row scale s_P2 (no per-block s_P1)
) {
auto dprops = at::cuda::getCurrentDeviceProperties();
bool is_blackwell_or_newer = dprops->major >= 12;
TORCH_CHECK(is_blackwell_or_newer, "only supports Blackwell GPUs or newer.");
auto q_dtype = q.dtype();
auto sfq_dtype = sfq.dtype();
TORCH_CHECK(q_dtype == torch::kUInt8, "q dtype must be uint8");
TORCH_CHECK(k.dtype() == q_dtype, "query and key must have the same dtype");
TORCH_CHECK(v.dtype() == q_dtype, "query and value must have the same dtype");
CHECK_DEVICE(q); CHECK_DEVICE(k); CHECK_DEVICE(v);
TORCH_CHECK(sfq_dtype == torch::kFloat8_e4m3fn, "q dtype must be uint8");
TORCH_CHECK(sfk.dtype() == sfq_dtype, "query and key must have the same dtype");
TORCH_CHECK(sfv.dtype() == sfq_dtype, "query and value must have the same dtype");
CHECK_DEVICE(sfq); CHECK_DEVICE(sfk); CHECK_DEVICE(sfv);
TORCH_CHECK(q.stride(-1) == 1, "Input tensor must have contiguous last dimension");
TORCH_CHECK(k.stride(-1) == 1, "Input tensor must have contiguous last dimension");
TORCH_CHECK(v.stride(-1) == 1, "Input tensor must have contiguous last dimension");
TORCH_CHECK(delta_s.stride(-1) == 1, "Input tensor must have contiguous last dimension");
TORCH_CHECK(q.is_contiguous(), "Input tensor must be contiguous");
TORCH_CHECK(k.is_contiguous(), "Input tensor must be contiguous");
TORCH_CHECK(v.is_contiguous(), "Input tensor must be contiguous");
const auto sizes = q.sizes();
auto opts = q.options();
const int batch_size = sizes[0];
int seqlen_q = sizes[2];
int num_heads = sizes[1];
const int head_size_og = sizes[3];
const int unpacked_head_size = head_size_og * 2;
const int seqlen_k = k.size(2);
const int num_heads_k = k.size(1);
TORCH_CHECK(batch_size > 0, "batch size must be postive");
TORCH_CHECK(unpacked_head_size <= 256, "FlashAttention forward only supports head dimension at most 256");
TORCH_CHECK(num_heads % num_heads_k == 0, "Number of heads in key/value must divide number of heads in query");
TORCH_CHECK(num_heads == num_heads_k, "We do not support MQA/GQA yet");
TORCH_CHECK(unpacked_head_size == 64 || unpacked_head_size == 128 || unpacked_head_size == 256, "Only support head size 64, 128, and 256 for now");
CHECK_SHAPE(q, batch_size, num_heads, seqlen_q, head_size_og);
CHECK_SHAPE(k, batch_size, num_heads_k, seqlen_k, head_size_og);
CHECK_SHAPE(v, batch_size, num_heads_k, unpacked_head_size, seqlen_k/2);
// CHECK_SHAPE(delta_s, batch_size, num_heads, seqlen_q / 128, seqlen_k);
// CHECK_SHAPE(sfq, batch_size, seqlen_q, num_heads, unpacked_head_size);
// CHECK_SHAPE(sfk, batch_size, seqlen_k, num_heads_k, unpacked_head_size);
// CHECK_SHAPE(sfv, batch_size, unpacked_head_size, num_heads_k, seqlen_k);
TORCH_CHECK(unpacked_head_size % 8 == 0, "head_size must be a multiple of 8");
auto dtype = is_bf16 ? at::ScalarType::BFloat16 : at::ScalarType::Half;
at::Tensor out = torch::empty({batch_size, num_heads, seqlen_q, unpacked_head_size}, opts.dtype(dtype));
auto round_multiple = [](int x, int m) { return (x + m - 1) / m * m; };
// const int head_size = round_multiple(head_size_og, 8);
// const int head_size_rounded = round_multiple(head_size, 32);
const int seqlen_q_rounded = round_multiple(seqlen_q, flash::BLOCK_M);
const int seqlen_k_rounded = round_multiple(seqlen_k, flash::BLOCK_N);
// Otherwise the kernel will be launched from cuda:0 device
// Cast to char to avoid compiler warning about narrowing
at::cuda::CUDAGuard device_guard{(char)q.get_device()};
auto softmax_lse = torch::empty({batch_size, num_heads, seqlen_q}, opts.dtype(at::kFloat));
at::Tensor p;
Flash_fwd_params params;
set_params_fprop(params,
batch_size,
seqlen_q, seqlen_k, unpadded_k,
seqlen_q_rounded, seqlen_k_rounded,
num_heads, num_heads_k,
unpacked_head_size, unpacked_head_size,
q, k, v, delta_s, out,
sfq, sfk, sfv,
/*cu_seqlens_q_d=*/nullptr,
/*cu_seqlens_k_d=*/nullptr,
/*seqused_k=*/nullptr,
nullptr,
softmax_lse.data_ptr(),
/*p_dropout=*/0.f,
softmax_scale,
/*window_size_left=*/-1,
/*window_size_right=*/is_causal ? 0 : -1,
per_block_mean,
is_bf16,
single_level_p_quant
);
// StaticPersistentTileScheduler does not use tile_count_semaphore; avoid a
// stack-local tensor whose data pointer would dangle after mha_fwd returns
// while the async kernel may still be running.
params.tile_count_semaphore = nullptr;
if (seqlen_k > 0) {
auto stream = at::cuda::getCurrentCUDAStream().stream();
run_mha_fwd(params, stream);
} else {
// If seqlen_k == 0, then we have an empty tensor. We need to set the output to 0.
out.zero_();
softmax_lse.fill_(std::numeric_limits<float>::infinity());
}
// at::Tensor out_padded = out;
// if (head_size_og % 8 != 0) {
// out = out.index({"...", torch::indexing::Slice(torch::indexing::None, head_size_og)});
// if (out_.has_value()) { out_.value().copy_(out); }
// }
// return {out, q_padded, k_padded, v_padded, out_padded, softmax_lse, p};
// cudaDeviceSynchronize();
// auto err = cudaGetLastError();
// printf("%s\n", cudaGetErrorString(err));
return {out, softmax_lse};
}
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
m.doc() = "FlashAttention";
m.def("fwd", &mha_fwd, "Forward pass");
}
@@ -1,28 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
// Centralized block size configuration for sageattn_blackwell kernels
// Block sizes for M and N dimensions
namespace flash {
// Block size for M dimension (query sequence length)
static constexpr int BLOCK_M = 128;
// Block size for N dimension (key/value sequence length)
static constexpr int BLOCK_N = 128;
}
@@ -1,60 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* This code is based on code from FlashAttention3, https://github.com/Dao-AILab/flash-attention
* Copyright (c) 2024, Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao.
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
namespace flash {
////////////////////////////////////////////////////////////////////////////////////////////////////
template<bool Varlen=true>
struct BlockInfo {
template<typename Params>
__device__ BlockInfo(const Params &params, const int bidb)
: sum_s_q(!Varlen || params.cu_seqlens_q == nullptr ? -1 : params.cu_seqlens_q[bidb])
, sum_s_k(!Varlen || params.cu_seqlens_k == nullptr || !params.is_seqlens_k_cumulative ? -1 : params.cu_seqlens_k[bidb])
, actual_seqlen_q(!Varlen || params.cu_seqlens_q == nullptr ? params.seqlen_q : params.cu_seqlens_q[bidb + 1] - sum_s_q)
// If is_seqlens_k_cumulative, then seqlen_k is cu_seqlens_k[bidb + 1] - cu_seqlens_k[bidb].
// Otherwise it's cu_seqlens_k[bidb], i.e., we use cu_seqlens_k to store the sequence lengths of K.
, seqlen_k_cache(!Varlen || params.cu_seqlens_k == nullptr ? params.seqlen_k : (params.is_seqlens_k_cumulative ? params.cu_seqlens_k[bidb + 1] - sum_s_k : params.cu_seqlens_k[bidb]))
, actual_seqlen_k(params.seqused_k ? params.seqused_k[bidb] : seqlen_k_cache + (params.knew_ptr == nullptr ? 0 : params.seqlen_knew))
{
}
template <typename index_t>
__forceinline__ __device__ index_t q_offset(const index_t batch_stride, const index_t row_stride, const int bidb) const {
return sum_s_q == -1 ? bidb * batch_stride : uint32_t(sum_s_q) * row_stride;
}
template <typename index_t>
__forceinline__ __device__ index_t k_offset(const index_t batch_stride, const index_t row_stride, const int bidb) const {
return sum_s_k == -1 ? bidb * batch_stride : uint32_t(sum_s_k) * row_stride;
}
const int sum_s_q;
const int sum_s_k;
const int actual_seqlen_q;
// We have to have seqlen_k_cache declared before actual_seqlen_k, otherwise actual_seqlen_k is set to 0.
const int seqlen_k_cache;
const int actual_seqlen_k;
};
////////////////////////////////////////////////////////////////////////////////////////////////////
} // namespace flash
@@ -1,149 +0,0 @@
/***************************************************************************************************
* Copyright (c) 2023 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
* SPDX-License-Identifier: BSD-3-Clause
*
* Redistribution and use in source and binary forms, with or without
* modification, are permitted provided that the following conditions are met:
*
* 1. Redistributions of source code must retain the above copyright notice, this
* list of conditions and the following disclaimer.
*
* 2. Redistributions in binary form must reproduce the above copyright notice,
* this list of conditions and the following disclaimer in the documentation
* and/or other materials provided with the distribution.
*
* 3. Neither the name of the copyright holder nor the names of its
* contributors may be used to endorse or promote products derived from
* this software without specific prior written permission.
*
* THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
* AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
* DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
* SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
* CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
* OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
* OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
*
**************************************************************************************************/
/*! \file
\brief Blocked Scale configs specific for SM100 BlockScaled MMA
*/
#pragma once
#include "cutlass/layout/matrix.h"
#include "cute/int_tuple.hpp"
#include "cute/atom/mma_traits_sm100.hpp"
namespace flash {
/////////////////////////////////////////////////////////////////////////////////////////////////
using namespace cute;
template<int SFVecSize, UMMA::Major major = UMMA::Major::K>
struct BlockScaledBasicChunk {
using Blk_MN = _64;
using Blk_SF = _4;
using SfAtom = Layout< Shape< Shape<_16,_4>, Shape<Int<SFVecSize>, _4>>,
Stride<Stride<_16,_4>, Stride< _0, _1>>>;
};
template<int SFVecSize_>
struct BlockScaledConfig {
// We are creating the SFA and SFB tensors' layouts in the collective since they always have the same layout.
// k-major order
static constexpr int SFVecSize = SFVecSize_;
static constexpr int MMA_NSF = 4; // SFVecSize, MMA_NSF
using BlkScaledChunk = BlockScaledBasicChunk<SFVecSize>;
using Blk_MN = _64;
using Blk_SF = _4;
using mnBasicBlockShape = Shape<_16,_4>;
using mnBasicBlockStride = Stride<_16,_4>;
using kBasicBlockShape = Shape<Int<SFVecSize>, Int<MMA_NSF>>; // SFVecSize, MMA_NSF
using kBasicBlockStride = Stride<_0, _1>;
using SfAtom = Layout< Shape< mnBasicBlockShape, kBasicBlockShape>,
Stride<mnBasicBlockStride, kBasicBlockStride>>;
using LayoutSF = decltype(blocked_product(SfAtom{},
make_layout(
make_shape(int32_t(0), int32_t(0), int32_t(0), int32_t(0)),
make_stride(int32_t(0), _1{}, int32_t(0), int32_t(0)))));
// A single indivisible block will hold 4 scale factors of 64 rows/columns (A/B matrix).
// 4 is chosen to make consecutive 32bits of data to have scale factors for only a single row (col). 32bits corresponds to the TMEM word size
using Blk_Elems = decltype(Blk_MN{} * Blk_SF{});
using sSF_strideMN = decltype(prepend(Blk_Elems{}, mnBasicBlockStride{}));
// The following function is provided for user fill dynamic problem size to the layout_SFA.
template < class ProblemShape>
CUTE_HOST_DEVICE
static constexpr auto
tile_atom_to_shape_SFQKV(ProblemShape problem_shape) {
auto [Seqlen, Dim, HeadNum, Batch] = problem_shape;
return tile_to_shape(SfAtom{}, make_shape(Seqlen, Dim, HeadNum, Batch), Step<_2,_1,_3,_4>{});
}
// The following function is provided for user fill dynamic problem size to the layout_SFB.
template <class ProblemShape>
CUTE_HOST_DEVICE
static constexpr auto
tile_atom_to_shape_SFVt(ProblemShape problem_shape) {
auto [Dim, Seqlen, HeadNum, Batch] = problem_shape;
return tile_to_shape(SfAtom{}, make_shape(Dim, Seqlen, HeadNum, Batch), Step<_2,_1,_3,_4>{});
}
template<class TiledMma, class TileShape_MNK>
CUTE_HOST_DEVICE
static constexpr auto
deduce_smem_layoutSFQ(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
using sSFQ_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
using sSFQ_shapeM = decltype(prepend(size<0>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
using sSFQ_strideM = sSF_strideMN;
using sSFQ_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<0>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
using sSFQ_shape = decltype(make_shape(sSFQ_shapeM{}, sSFQ_shapeK{}));
using sSFQ_stride = decltype(make_stride(sSFQ_strideM{}, sSFQ_strideK{}));
using SmemLayoutAtomSFQ = decltype(make_layout(sSFQ_shape{}, sSFQ_stride{}));
return SmemLayoutAtomSFQ{};
}
template<class TiledMma, class TileShape_MNK>
CUTE_HOST_DEVICE
static constexpr auto
deduce_smem_layoutSFKV(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
using sSFK_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
using sSFK_shapeN = decltype(prepend(size<1>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
using sSFK_strideN = sSF_strideMN;
using sSFK_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<1>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
using sSFK_shape = decltype(make_shape(sSFK_shapeN{}, sSFK_shapeK{}));
using sSFK_stride = decltype(make_stride(sSFK_strideN{}, sSFK_strideK{}));
using SmemLayoutAtomSFK = decltype(make_layout(sSFK_shape{}, sSFK_stride{}));
return SmemLayoutAtomSFK{};
}
template<class TiledMma, class TileShape_MNK>
CUTE_HOST_DEVICE
static constexpr auto
deduce_smem_layoutSFVt(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
using sSFVt_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
using sSFVt_shapeN = decltype(prepend(size<1>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
using sSFVt_strideN = sSF_strideMN;
using sSFVt_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<1>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
using sSFVt_shape = decltype(make_shape(sSFVt_shapeN{}, sSFVt_shapeK{}));
using sSFVt_stride = decltype(make_stride(sSFVt_strideN{}, sSFVt_strideK{}));
using SmemLayoutAtomSFVt = decltype(make_layout(sSFVt_shape{}, sSFVt_stride{}));
return SmemLayoutAtomSFVt{};
}
};
} // namespace flash
@@ -1,327 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include "cute/arch/mma_sm120.hpp"
#include "cute/atom/mma_traits_sm120.hpp"
#include "cute/atom/mma_atom.hpp"
#include "cutlass/cutlass.h"
#include "cutlass/float8.h"
#include "cutlass/float_subbyte.h"
namespace cute::SM120::BLOCKSCALED {
using cutlass::float_e2m1_t;
using cutlass::float_ue4m3_t;
// MMA.SF 16x32x64 TN E2M1 x E2M1 with SF E4M3
struct SM120_16x32x64_TN_VS_NVFP4 {
using DRegisters = float[16];
using ARegisters = uint32_t[4];
using BRegisters = uint32_t[8];
using CRegisters = float[16];
static constexpr int SFBits = 32;
using RegTypeSF = cute::uint_bit_t<SFBits>;
using SFARegisters = RegTypeSF[1];
using SFBRegisters = RegTypeSF[1];
CUTE_HOST_DEVICE static void
fma(float & d0 , float & d1 , float & d2 , float & d3 ,
float & d4 , float & d5 , float & d6 , float & d7 ,
float & d8 , float & d9 , float & d10, float & d11,
float & d12, float & d13, float & d14, float & d15,
uint32_t const& a0 , uint32_t const& a1 , uint32_t const& a2 , uint32_t const& a3 ,
uint32_t const& b0 , uint32_t const& b1 , uint32_t const& b2 , uint32_t const& b3 ,
uint32_t const& b4 , uint32_t const& b5 , uint32_t const& b6 , uint32_t const& b7 ,
float const & c0 , float const & c1 , float const & c2 , float const & c3 ,
float const & c4 , float const & c5 , float const & c6 , float const & c7 ,
float const & c8 , float const & c9 , float const & c10 , float const & c11,
float const & c12, float const & c13, float const & c14, float const & c15,
RegTypeSF const& sfa0,
RegTypeSF const& sfb0)
{
static constexpr uint16_t tidA = 0;
static constexpr uint16_t bidA = 0;
static constexpr uint16_t bidB = 0;
static constexpr uint16_t tidB0 = 0;
static constexpr uint16_t tidB1 = 1;
static constexpr uint16_t tidB2 = 2;
static constexpr uint16_t tidB3 = 3;
#if defined(CUTE_ARCH_MXF4NVF4_4X_UE4M3_MMA_ENABLED)
asm volatile(
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
"{%0, %1, %2, %3},"
"{%4, %5, %6, %7},"
"{%8, %9},"
"{%10, %11, %12, %13},"
"{%14},"
"{%15, %16},"
"{%17},"
"{%18, %19};\n"
: "=f"(d0), "=f"(d1), "=f"(d8), "=f"(d9)
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
"r"(b0), "r"(b1),
"f"(c0), "f"(c1), "f"(c8), "f"(c9),
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB0));
asm volatile(
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
"{%0, %1, %2, %3},"
"{%4, %5, %6, %7},"
"{%8, %9},"
"{%10, %11, %12, %13},"
"{%14},"
"{%15, %16},"
"{%17},"
"{%18, %19};\n"
: "=f"(d2), "=f"(d3), "=f"(d10), "=f"(d11)
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
"r"(b2), "r"(b3),
"f"(c2), "f"(c3), "f"(c10), "f"(c11),
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB1));
asm volatile(
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
"{%0, %1, %2, %3},"
"{%4, %5, %6, %7},"
"{%8, %9},"
"{%10, %11, %12, %13},"
"{%14},"
"{%15, %16},"
"{%17},"
"{%18, %19};\n"
: "=f"(d4), "=f"(d5), "=f"(d12), "=f"(d13)
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
"r"(b4), "r"(b5),
"f"(c4), "f"(c5), "f"(c12), "f"(c13),
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB2));
asm volatile(
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
"{%0, %1, %2, %3},"
"{%4, %5, %6, %7},"
"{%8, %9},"
"{%10, %11, %12, %13},"
"{%14},"
"{%15, %16},"
"{%17},"
"{%18, %19};\n"
: "=f"(d6), "=f"(d7), "=f"(d14), "=f"(d15)
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
"r"(b6), "r"(b7),
"f"(c6), "f"(c7), "f"(c14), "f"(c15),
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB3));
#else
CUTE_INVALID_CONTROL_PATH("Attempting to use SM120::BLOCKSCALED::SM120_16x8x64_TN_VS without CUTE_ARCH_MXF4NVF4_4X_UE4M3_MMA_ENABLED");
#endif
}
};
} // namespace cute::SM120::BLOCKSCALED
namespace cute {
// MMA NVFP4 16x32x64 TN
template <>
struct MMA_Traits<SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4>
{
// The MMA accepts 4-bit inputs regardless of the types for A and B
using ValTypeA = uint4_t;
using ValTypeB = uint4_t;
using ValTypeD = float;
using ValTypeC = float;
using ValTypeSF = cutlass::float_ue4m3_t;
constexpr static int SFVecSize = 16;
using Shape_MNK = Shape<_16,_32,_64>;
using ThrID = Layout<_32>;
// (T32,V32) -> (M16,K64)
using ALayout = Layout<Shape <Shape < _4,_8>,Shape < _8,_2, _2>>,
Stride<Stride<_128,_1>,Stride<_16,_8,_512>>>;
// (T32,V64) -> (N32,K64)
using BLayout = Layout<Shape <Shape < _4,_8>,Shape <_8, _2, _4>>,
Stride<Stride<_256,_1>,Stride<_32,_1024, _8>>>;
// (T32,V64) -> (M16,K64)
using SFALayout = Layout<Shape <Shape <_2,_2,_8>,_64>,
Stride<Stride<_8,_0,_1>,_16>>;
// (T32,V64) -> (N32,K64)
using SFBLayout = Layout<Shape <Shape <_4,_8>,_64>,
Stride<Stride<_8,_1>, _32>>;
// (T32,V16) -> (M16,N32)
using CLayout = Layout<Shape <Shape < _4,_8>,Shape < Shape<_2, _4>,_2>>,
Stride<Stride<_32,_1>,Stride<Stride<_16, _128>,_8>>>;
};
template <class SFATensor, class Atom, class TiledThr, class TiledPerm>
CUTE_HOST_DEVICE constexpr
auto
thrfrg_SFA(SFATensor&& sfatensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
{
CUTE_STATIC_ASSERT_V(rank(sfatensor) >= Int<2>{});
using AtomShape_MNK = typename Atom::Shape_MNK;
using AtomLayoutSFA_TV = typename Atom::Traits::SFALayout;
auto permutation_mnk = TiledPerm{};
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// Reorder the tensor for the TiledAtom
auto t_tile = make_tile(get<0>(permutation_mnk),
get<2>(permutation_mnk));
auto t_tensor = logical_divide(sfatensor, t_tile); // (PermM,PermK)
// Tile the tensor for the Atom
auto a_tile = make_tile(make_layout(size<0>(AtomShape_MNK{})),
make_layout(size<2>(AtomShape_MNK{})));
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomM,AtomK),(RestM,RestK))
// Transform the Atom mode from (M,K) to (Thr,Val)
auto tv_tensor = a_tensor.compose(AtomLayoutSFA_TV{},_); // ((ThrV,FrgV),(RestM,RestK))
// Tile the tensor for the Thread
auto thr_tile = make_tile(_,
make_tile(make_layout(size<1>(thr_layout_vmnk)),
make_layout(size<3>(thr_layout_vmnk))));
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrM,ThrK)),(FrgV,(RestM,RestK)))
return thr_tensor;
}
template <class SFBTensor, class Atom, class TiledThr, class TiledPerm>
CUTE_HOST_DEVICE constexpr
auto
thrfrg_SFB(SFBTensor&& sfbtensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
{
CUTE_STATIC_ASSERT_V(rank(sfbtensor) >= Int<2>{});
using AtomShape_MNK = typename Atom::Shape_MNK;
using AtomLayoutSFB_TV = typename Atom::Traits::SFBLayout;
auto permutation_mnk = TiledPerm{};
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// Reorder the tensor for the TiledAtom
auto t_tile = make_tile(get<1>(permutation_mnk),
get<2>(permutation_mnk));
auto t_tensor = logical_divide(sfbtensor, t_tile); // (PermN,PermK)
// Tile the tensor for the Atom
auto a_tile = make_tile(make_layout(size<1>(AtomShape_MNK{})),
make_layout(size<2>(AtomShape_MNK{})));
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomN,AtomK),(RestN,RestK))
// Transform the Atom mode from (M,K) to (Thr,Val)
auto tv_tensor = a_tensor.compose(AtomLayoutSFB_TV{},_); // ((ThrV,FrgV),(RestN,RestK))
// Tile the tensor for the Thread
auto thr_tile = make_tile(_,
make_tile(make_layout(size<2>(thr_layout_vmnk)),
make_layout(size<3>(thr_layout_vmnk))));
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrN,ThrK)),(FrgV,(RestN,RestK)))
return thr_tensor;
}
template <class SFATensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_SFA(SFATensor&& sfatensor, ThrMma& thread_mma) {
auto thr_tensor = make_tensor(static_cast<SFATensor&&>(sfatensor).data(), thrfrg_SFA(sfatensor.layout(),thread_mma));
auto thr_vmnk = thread_mma.thr_vmnk_;
auto thr_vmk = make_coord(get<0>(thr_vmnk), make_coord(get<1>(thr_vmnk), get<3>(thr_vmnk)));
return thr_tensor(thr_vmk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
}
template <class SFATensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_fragment_SFA(SFATensor&& sfatensor, ThrMma& thread_mma) {
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
return make_fragment_like<ValTypeSF>(partition_SFA(sfatensor, thread_mma));
}
template <class SFBTensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma) {
auto thr_tensor = make_tensor(static_cast<SFBTensor&&>(sfbtensor).data(), thrfrg_SFB(sfbtensor.layout(),thread_mma));
auto thr_vmnk = thread_mma.thr_vmnk_;
auto thr_vnk = make_coord(get<0>(thr_vmnk), make_coord(get<2>(thr_vmnk), get<3>(thr_vmnk)));
return thr_tensor(thr_vnk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
}
template <class SFBTensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_fragment_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma) {
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
return make_fragment_like<ValTypeSF>(partition_SFB(sfbtensor, thread_mma));
}
template<class TiledMma>
CUTE_HOST_DEVICE constexpr
auto
get_layoutSFA_TV(TiledMma& mma)
{
// (M,K) -> (M,K)
auto tile_shape_mnk = tile_shape(mma);
auto ref_A = make_layout(make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk)));
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
auto atile = make_tile(_,
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
make_stride( Int<1>{} , Int<0>{} )),
_));
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
// (thr_idx,val) -> (M,K)
return thrfrg_SFA(ref_A, mma).compose(atile, _).compose(thridx_2_thrid, _);
}
template<class TiledMma>
CUTE_HOST_DEVICE constexpr
auto
get_layoutSFB_TV(TiledMma& mma)
{
// (N,K) -> (N,K)
auto tile_shape_mnk = tile_shape(mma);
auto ref_B = make_layout(make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk)));
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
auto btile = make_tile(_,
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
make_stride( Int<0>{} , Int<1>{} )),
_));
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
// (thr_idx,val) -> (M,K)
return thrfrg_SFB(ref_B, mma).compose(btile, _).compose(thridx_2_thrid, _);
}
} // namespace cute
@@ -1,222 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <cutlass/cutlass.h>
#include "cute/tensor.hpp"
#include "cutlass/gemm/collective/collective_builder.hpp"
#include "named_barrier.h"
#include "utils.h"
namespace flash {
using namespace cute;
template <typename Ktraits>
struct CollectiveEpilogueFwd{
using Element = typename Ktraits::ElementOut;
static constexpr int kBlockM = Ktraits::kBlockM;
static constexpr int kBlockN = Ktraits::kBlockN;
static constexpr int kHeadDim = Ktraits::kHeadDim;
using TileShape_MNK = Shape<Int<kBlockM>, Int<kBlockN>, Int<kHeadDim>>;
static constexpr int kNWarps = Ktraits::kNWarps;
static constexpr int kNThreads = kNWarps * cutlass::NumThreadsPerWarp;
static constexpr int NumMmaThreads = kNThreads - cutlass::NumThreadsPerWarpGroup;
using GmemTiledCopyOTMA = cute::SM90_TMA_STORE;
// These are for storing the output tensor without TMA (e.g., for setting output to zero)
static constexpr int kGmemElemsPerLoad = sizeof(cute::uint128_t) / sizeof(Element);
static_assert(kHeadDim % kGmemElemsPerLoad == 0, "kHeadDim must be a multiple of kGmemElemsPerLoad");
static constexpr int kGmemThreadsPerRow = kHeadDim / kGmemElemsPerLoad;
static_assert(NumMmaThreads % kGmemThreadsPerRow == 0, "NumMmaThreads must be a multiple of kGmemThreadsPerRow");
using GmemLayoutAtom = Layout<Shape <Int<NumMmaThreads / kGmemThreadsPerRow>, Int<kGmemThreadsPerRow>>,
Stride<Int<kGmemThreadsPerRow>, _1>>;
using GmemTiledCopyO = decltype(
make_tiled_copy(Copy_Atom<DefaultCopy, Element>{},
GmemLayoutAtom{},
Layout<Shape<_1, Int<kGmemElemsPerLoad>>>{})); // Val layout, 8 or 16 vals per store
using SmemLayoutO = typename Ktraits::SmemLayoutO;
using SmemCopyAtomO = Copy_Atom<SM90_U32x2_STSM_N, Element>;
using SharedStorage = cute::array_aligned<Element, cute::cosize_v<SmemLayoutO>>;
using ShapeO = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen_q, d, head, batch)
using StrideO = cute::Stride<int64_t, _1, int64_t, int64_t>;
using StrideLSE = cute::Stride<_1, int64_t, int64_t>; // (seqlen_q, head, batch)
using TMA_O = decltype(make_tma_copy(
GmemTiledCopyOTMA{},
make_tensor(make_gmem_ptr(static_cast<Element*>(nullptr)), repeat_like(StrideO{}, int32_t(0)), StrideO{}),
SmemLayoutO{},
select<0, 2>(TileShape_MNK{}),
_1{})); // no mcast for O
// Host side kernel arguments
struct Arguments {
Element* ptr_O;
ShapeO const shape_O;
StrideO const stride_O;
float* ptr_LSE;
StrideLSE const stride_LSE;
};
// Device side kernel params
struct Params {
Element* ptr_O;
ShapeO const shape_O;
StrideO const stride_O;
float* ptr_LSE;
StrideLSE const stride_LSE;
TMA_O tma_store_O;
};
static Params
to_underlying_arguments(Arguments const& args) {
Tensor mO = make_tensor(make_gmem_ptr(args.ptr_O), args.shape_O, args.stride_O);
TMA_O tma_store_O = make_tma_copy(
GmemTiledCopyOTMA{},
mO,
SmemLayoutO{},
select<0, 2>(TileShape_MNK{}),
_1{}); // no mcast for O
return {args.ptr_O, args.shape_O, args.stride_O, args.ptr_LSE, args.stride_LSE, tma_store_O};
}
/// Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
CUTLASS_DEVICE
static void prefetch_tma_descriptors(Params const& epilogue_params) {
cute::prefetch_tma_descriptor(epilogue_params.tma_store_O.get_tma_descriptor());
}
template <typename SharedStorage, typename FrgTensorO, typename TiledMma>
CUTLASS_DEVICE void
mma_store(
SharedStorage& shared_storage,
TiledMma tiled_mma,
FrgTensorO const& tOrO,
int thread_idx
){
Tensor sO = cute::as_position_independent_swizzle_tensor(make_tensor(make_smem_ptr(shared_storage.smem_o.begin()), SmemLayoutO{}));
auto smem_tiled_copy_O = make_tiled_copy_C(SmemCopyAtomO{}, tiled_mma);
auto smem_thr_copy_O = smem_tiled_copy_O.get_thread_slice(thread_idx);
constexpr int numel = decltype(size(tOrO))::value;
cutlass::NumericArrayConverter<Element, float, numel> convert_op;
// HACK: this requires tensor to be "contiguous"
auto frag = convert_op(*reinterpret_cast<const cutlass::Array<float, numel> *>(tOrO.data()));
auto tOrO_out = make_tensor(make_rmem_ptr<Element>(&frag), tOrO.layout());
Tensor taccOrO = smem_thr_copy_O.retile_S(tOrO_out); // ((Atom,AtomNum), MMA_M, MMA_N)
Tensor taccOsO = smem_thr_copy_O.partition_D(sO); // ((Atom,AtomNum),PIPE_M,PIPE_N)
cute::copy(smem_tiled_copy_O, taccOrO, taccOsO);
cutlass::arch::fence_view_async_shared(); // ensure smem writes are visible to TMA
}
template<typename SharedStorage, typename Params, typename WorkTileInfo, typename SchedulerParams>
CUTLASS_DEVICE void
tma_store(
SharedStorage& shared_storage,
Params const& epilogue_params,
WorkTileInfo work_tile_info,
SchedulerParams const& scheduler_params,
int thread_idx
) {
auto [m_block, bidh, bidb] = work_tile_info.get_block_coord(scheduler_params);
Tensor sO = cute::as_position_independent_swizzle_tensor(make_tensor(make_smem_ptr(shared_storage.smem_o.begin()), SmemLayoutO{}));
Tensor mO = epilogue_params.tma_store_O.get_tma_tensor(epilogue_params.shape_O);
Tensor gO = local_tile(mO(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
auto block_tma_O = epilogue_params.tma_store_O.get_slice(_0{});
Tensor tOgO = block_tma_O.partition_D(gO); // (TMA, TMA_M, TMA_K)
Tensor tOsO = block_tma_O.partition_S(sO); // (TMA, TMA_M, TMA_K)
// auto shape_LSE = select<0, 2, 3>(epilogue_params.shape_O);
// Tensor mLSE = make_tensor(make_gmem_ptr(epilogue_params.ptr_LSE), shape_LSE, epilogue_params.stride_LSE);
// Tensor gLSE = local_tile(mLSE(_, bidh, bidb), Shape<Int<kBlockM>>{}, make_coord(m_block));
// Tensor caccO = cute::make_identity_tensor(select<0, 2>(TileShape_MNK{}));
// auto thread_mma = tiled_mma.get_thread_slice(thread_idx);
// Tensor taccOcO = thread_mma.partition_C(caccO); // (MMA,MMA_M,MMA_K)
// static_assert(decltype(size<0, 0>(taccOcO))::value == 2);
// static_assert(decltype(size<0, 1>(taccOcO))::value == 2);
// // // // taccOcO has shape ((2, 2, V), MMA_M, MMA_K), we only take only the row indices.
// Tensor taccOcO_row = taccOcO(make_coord(_0{}, _), _, _0{});
// CUTE_STATIC_ASSERT_V(size(lse) == size(taccOcO_row)); // MMA_M
// if (get<1>(taccOcO_row(_0{})) == 0) {
// #pragma unroll
// for (int mi = 0; mi < size(lse); ++mi) {
// const int row = get<0>(taccOcO_row(mi));
// if (row < get<0>(shape_LSE) - m_block * kBlockM) { gLSE(row) = lse(mi); }
// }
// }
// if (cutlass::canonical_warp_idx_sync() == kNWarps - 1) {
// cutlass::arch::NamedBarrier::sync(NumMmaThreads + cutlass::NumThreadsPerWarp,
// static_cast<uint32_t>(FP4NamedBarriers::EpilogueBarrier));
// int const lane_predicate = cute::elect_one_sync();
// if (lane_predicate) {
// cute::copy(epilogue_params.tma_store_O, tOsO, tOgO);
// tma_store_arrive();
// }
// }
cute::copy(epilogue_params.tma_store_O, tOsO, tOgO);
tma_store_arrive();
}
CUTLASS_DEVICE void
store_tail() {
tma_store_wait<0>();
}
// Write 0 to output and -inf to LSE
CUTLASS_DEVICE void
store_zero(
Params const& epilogue_params,
int thread_idx,
cute::tuple<int32_t, int32_t, int32_t> const& block_coord
) {
auto [m_block, bidh, bidb] = block_coord;
Tensor mO = make_tensor(make_gmem_ptr(epilogue_params.ptr_O), epilogue_params.shape_O, epilogue_params.stride_O);
Tensor gO = local_tile(mO(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
auto shape_LSE = select<0, 2, 3>(epilogue_params.shape_O);
Tensor mLSE = make_tensor(make_gmem_ptr(epilogue_params.ptr_LSE), shape_LSE, epilogue_params.stride_LSE);
Tensor gLSE = local_tile(mLSE(_, bidh, bidb), Shape<Int<kBlockM>>{}, make_coord(m_block));
GmemTiledCopyO gmem_tiled_copy_O;
auto gmem_thr_copy_O = gmem_tiled_copy_O.get_thread_slice(thread_idx);
Tensor tOgO = gmem_thr_copy_O.partition_D(gO);
Tensor tOrO = make_fragment_like(tOgO);
clear(tOrO);
// Construct identity layout for sO
Tensor cO = cute::make_identity_tensor(select<0, 2>(TileShape_MNK{})); // (BLK_M,BLK_K) -> (blk_m,blk_k)
// Repeat the partitioning with identity layouts
Tensor tOcO = gmem_thr_copy_O.partition_D(cO);
Tensor tOpO = make_tensor<bool>(make_shape(size<2>(tOgO)));
#pragma unroll
for (int k = 0; k < size(tOpO); ++k) { tOpO(k) = get<1>(tOcO(_0{}, _0{}, k)) < get<1>(epilogue_params.shape_O); }
// Clear_OOB_K must be false since we don't want to write zeros to gmem
flash::copy</*Is_even_MN=*/false, /*Is_even_K=*/false, /*Clear_OOB_MN=*/false, /*Clear_OOB_K=*/false>(
gmem_tiled_copy_O, tOrO, tOgO, tOcO, tOpO, get<0>(epilogue_params.shape_O) - m_block * kBlockM
);
static_assert(kBlockM <= NumMmaThreads);
if (thread_idx < get<0>(shape_LSE) - m_block * kBlockM) { gLSE(thread_idx) = INFINITY; }
}
};
} // namespace flash
@@ -1,202 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include "cute/algorithm/copy.hpp"
#include "cute/atom/mma_atom.hpp"
#include "cutlass/gemm/collective/collective_builder.hpp"
#include "cute/tensor.hpp"
#include "cutlass/cutlass.h"
#include "cutlass/layout/layout.h"
#include "cutlass/numeric_types.h"
#include "cutlass/pipeline/pipeline.hpp"
#include "blockscaled_layout.h"
#include "cute_extension.h"
#include "named_barrier.h"
using namespace cute;
template <
int kStages,
int EpiStages,
typename Element,
typename ElementSF,
typename OutputType,
typename SmemLayoutQ,
typename SmemLayoutK,
typename SmemLayoutV,
typename SmemLayoutDS,
typename SmemLayoutO,
typename SmemLayoutSFQ,
typename SmemLayoutSFK,
typename SmemLayoutSFV
>
struct SharedStorageQKVOwithSF : cute::aligned_struct<128, _0>{
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutQ>> smem_q;
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutK>> smem_k;
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFQ>> smem_SFQ;
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFK>> smem_SFK;
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFV>> smem_SFV;
alignas(1024) cute::ArrayEngine<float, cute::cosize_v<SmemLayoutDS>> smem_ds;
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutV>> smem_v;
alignas(1024) cute::ArrayEngine<OutputType, cute::cosize_v<SmemLayoutO>> smem_o;
struct {
alignas(16) typename cutlass::PipelineTmaAsync<1>::SharedStorage pipeline_q;
alignas(16) typename cutlass::PipelineTmaAsync<kStages>::SharedStorage pipeline_k;
alignas(16) typename cutlass::PipelineTmaAsync<kStages>::SharedStorage pipeline_v;
alignas(16) typename flash::OrderedSequenceBarrierVarGroupSize<EpiStages, 2>::SharedStorage barrier_o;
int tile_count_semaphore;
};
};
template <
int kHeadDim_,
int kBlockM_,
int kBlockN_,
int kStages_,
int kClusterM_,
bool BlockMean_,
typename ElementPairType_ = cutlass::nv_float4_t<cutlass::float_e2m1_t>,
typename ElementOut_ = cutlass::bfloat16_t
>
struct Flash_fwd_kernel_traits {
static constexpr int kBlockM = kBlockM_;
static constexpr int kBlockN = kBlockN_;
static constexpr int kHeadDim = kHeadDim_;
static constexpr bool BlockMean = BlockMean_;
static constexpr bool SmoothQ = true;
static_assert(kHeadDim % 32 == 0);
static_assert(kBlockM == 64 || kBlockM == 128);
static constexpr int kNWarps = kBlockM == 128 ? 12 : 8;
static constexpr int kNThreads = kNWarps * cutlass::NumThreadsPerWarp;
static constexpr int kClusterM = kClusterM_;
static constexpr int kStages = kStages_;
static constexpr int EpiStages = 1;
static constexpr int NumSFQK = kHeadDim / 16;
static constexpr int NumSFPV = kBlockN / 16;
using ElementSF = cutlass::float_ue4m3_t;
using Element = cutlass::float_e2m1_t;
using ElementAccum = float;
using ElementOut = ElementOut_;
using index_t = int64_t;
static constexpr auto SFVectorSize = 16;
using TileShape_MNK = Shape<Int<kBlockM>, Int<kBlockN>, Int<kHeadDim>>;
using ClusterShape_MNK = Shape<_1, _1, _1>;
using PermTileM = decltype(cute::min(size<0>(TileShape_MNK{}), _128{}));
using PermTileN = _32;
using PermTileK = Int<kHeadDim>;
using ElementQMma = decltype(cutlass::gemm::collective::detail::sm1xx_kernel_input_element_to_mma_input_element<Element>());
using ElementKMma = decltype(cutlass::gemm::collective::detail::sm1xx_kernel_input_element_to_mma_input_element<Element>());
using AtomLayoutMNK = std::conditional_t<kBlockM == 128,
Layout<Shape<_8, _1, _1>>,
Layout<Shape<_4, _1, _1>>
>;
using TiledMmaQK = decltype(cute::make_tiled_mma(
cute::SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4{},
AtomLayoutMNK{},
Tile<PermTileM, PermTileN, PermTileK>{}
));
using TiledMmaPV = decltype(cute::make_tiled_mma(
cute::SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4{},
AtomLayoutMNK{},
Tile<PermTileM, _32, PermTileK>{}
));
static constexpr int MMA_NSF = size<2>(typename TiledMmaQK::AtomShape_MNK{}) / SFVectorSize;
using GmemTiledCopy = SM90_TMA_LOAD;
using GmemTiledCopySF = SM90_TMA_LOAD;
using SmemLayoutAtomQ = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
using SmemLayoutAtomK = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
using SmemLayoutAtomV = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
using SmemLayoutAtomVt = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<1>(TileShape_MNK{}))>());
using SmemLayoutQ = decltype(tile_to_shape(SmemLayoutAtomQ{}, select<0, 2>(TileShape_MNK{})));
using SmemLayoutK =
decltype(tile_to_shape(SmemLayoutAtomK{},
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{}), Int<kStages>{})));
using SmemLayoutV =
decltype(tile_to_shape(SmemLayoutAtomV{},
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{}), Int<kStages>{})));
using SmemLayoutVt =
decltype(tile_to_shape(SmemLayoutAtomVt{},
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{}), Int<kStages>{})));
using SmemLayoutAtomDS = Layout<Shape<Int<kBlockM>, Int<kBlockN>>, Stride<_0, _1>>;
using SmemLayoutDS =
decltype(tile_to_shape(SmemLayoutAtomDS{},
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{}), Int<kStages>{})));
using SmemCopyAtomQ = Copy_Atom<SM75_U32x4_LDSM_N, Element>;
using SmemCopyAtomKV = Copy_Atom<SM75_U32x4_LDSM_N, Element>;
using SmemCopyAtomSF = Copy_Atom<UniversalCopy<ElementSF>, ElementSF>;
using SmemCopyAtomDS = Copy_Atom<UniversalCopy<float>, float>;
using BlkScaledConfig = flash::BlockScaledConfig<SFVectorSize>;
using LayoutSF = typename BlkScaledConfig::LayoutSF;
using SfAtom = typename BlkScaledConfig::SfAtom;
using SmemLayoutAtomSFQ = decltype(BlkScaledConfig::deduce_smem_layoutSFQ(TiledMmaQK{}, TileShape_MNK{}));
using SmemLayoutAtomSFK = decltype(BlkScaledConfig::deduce_smem_layoutSFKV(TiledMmaQK{}, TileShape_MNK{}));
using SmemLayoutAtomSFV = decltype(BlkScaledConfig::deduce_smem_layoutSFKV(TiledMmaPV{}, TileShape_MNK{}));
using SmemLayoutAtomSFVt = decltype(BlkScaledConfig::deduce_smem_layoutSFVt(TiledMmaPV{}, Shape<Int<kBlockM>, Int<kHeadDim>, Int<kBlockN>>{}));
using LayoutSFP = decltype(
make_layout(
make_shape(make_shape(_16{}, _4{}), _1{}, Int<kBlockN / 64>{}),
make_stride(make_stride(_0{}, _1{}), _0{}, _4{})
)
);
using LayoutP = decltype(
make_layout(
make_shape(make_shape(_8{}, _2{}, _2{}), _1{}, Int<kBlockN / 64>{}),
make_stride(make_stride(_1{}, _8{}, _16{}), _0{}, _32{})
)
);
using SmemLayoutSFQ = decltype(make_layout(
shape(SmemLayoutAtomSFQ{}),
stride(SmemLayoutAtomSFQ{})
));
using SmemLayoutSFK = decltype(make_layout(
append(shape(SmemLayoutAtomSFK{}), Int<kStages>{}),
append(stride(SmemLayoutAtomSFK{}), size(filter_zeros(SmemLayoutAtomSFK{})))
));
using SmemLayoutSFV = decltype(make_layout(
append(shape(SmemLayoutAtomSFV{}), Int<kStages>{}),
append(stride(SmemLayoutAtomSFV{}), size(filter_zeros(SmemLayoutAtomSFV{})))
));
using SmemLayoutSFVt = decltype(make_layout(
append(shape(SmemLayoutAtomSFVt{}), Int<kStages>{}),
append(stride(SmemLayoutAtomSFVt{}), size(filter_zeros(SmemLayoutAtomSFVt{})))
));
using SmemLayoutAtomO = decltype(cutlass::gemm::collective::detail::ss_smem_selector<GMMA::Major::K, ElementOut,
decltype(cute::get<0>(TileShape_MNK{})), decltype(cute::get<2>(TileShape_MNK{}))>());
using SmemLayoutO = decltype(tile_to_shape(SmemLayoutAtomO{}, select<0, 2>(TileShape_MNK{}), Step<_1, _2>{}));
using SharedStorage = SharedStorageQKVOwithSF<kStages, EpiStages, Element, ElementSF, ElementOut,
SmemLayoutQ, SmemLayoutK, SmemLayoutV, SmemLayoutDS,
SmemLayoutO, SmemLayoutSFQ, SmemLayoutSFK, SmemLayoutSFVt>;
using MainloopPipeline = typename cutlass::PipelineTmaAsync<kStages>;
using PipelineState = typename cutlass::PipelineState<kStages>;
using MainloopPipelineQ = cutlass::PipelineTmaAsync<1>;
using PipelineParamsQ = typename MainloopPipelineQ::Params;
using PipelineStateQ = typename cutlass::PipelineState<1>;
using EpilogueBarrier = typename flash::OrderedSequenceBarrierVarGroupSize<EpiStages, 2>;
};
@@ -1,204 +0,0 @@
// Modified from the original SageAttention3 code
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include "cute/tensor.hpp"
#include <cutlass/cutlass.h>
#include <cutlass/arch/reg_reconfig.h>
#include <cutlass/array.h>
#include <cutlass/numeric_types.h>
#include <cutlass/numeric_conversion.h>
#include "cutlass/pipeline/pipeline.hpp"
#include "params.h"
#include "utils.h"
#include "tile_scheduler.h"
#include "mainloop_tma_ws.h"
#include "epilogue_tma_ws.h"
#include "named_barrier.h"
#include "softmax_fused.h"
namespace flash {
using namespace cute;
template <typename Ktraits, bool Is_causal, typename TileScheduler>
__global__ void __launch_bounds__(Ktraits::kNWarps * cutlass::NumThreadsPerWarp, 1)
compute_attn_ws(CUTE_GRID_CONSTANT Flash_fwd_params const params,
CUTE_GRID_CONSTANT typename CollectiveMainloopFwd<Ktraits, Is_causal>::Params const mainloop_params,
CUTE_GRID_CONSTANT typename CollectiveEpilogueFwd<Ktraits>::Params const epilogue_params,
CUTE_GRID_CONSTANT typename TileScheduler::Params const scheduler_params
) {
using Element = typename Ktraits::Element;
using ElementAccum = typename Ktraits::ElementAccum;
using SoftType = ElementAccum;
using TileShape_MNK = typename Ktraits::TileShape_MNK;
using ClusterShape = typename Ktraits::ClusterShape_MNK;
static constexpr int NumMmaThreads = size(typename Ktraits::TiledMmaQK{});
static constexpr int NumCopyThreads = cutlass::NumThreadsPerWarpGroup;
static constexpr int kBlockM = Ktraits::kBlockM;
using CollectiveMainloop = CollectiveMainloopFwd<Ktraits, Is_causal>;
using CollectiveEpilogue = CollectiveEpilogueFwd<Ktraits>;
using MainloopPipeline = typename Ktraits::MainloopPipeline;
using PipelineParams = typename MainloopPipeline::Params;
using PipelineState = typename MainloopPipeline::PipelineState;
using MainloopPipelineQ = typename Ktraits::MainloopPipelineQ;
using PipelineParamsQ = typename Ktraits::PipelineParamsQ;
using PipelineStateQ = typename Ktraits::PipelineStateQ;
using EpilogueBarrier = typename Ktraits::EpilogueBarrier;
enum class WarpGroupRole {
Producer = 0,
Consumer0 = 1,
Consumer1 = 2
};
enum class ProducerWarpRole {
Mainloop = 0,
Epilogue = 1,
Warp2 = 2,
Warp3 = 3
};
extern __shared__ char shared_memory[];
auto &shared_storage = *reinterpret_cast<typename Ktraits::SharedStorage*>(shared_memory);
int const lane_predicate = cute::elect_one_sync();
int const warp_idx = cutlass::canonical_warp_idx_sync();
int warp_group_idx = cutlass::canonical_warp_group_idx();
int const warp_group_thread_idx = threadIdx.x % cutlass::NumThreadsPerWarpGroup;
int warp_idx_in_warp_group = warp_idx % cutlass::NumWarpsPerWarpGroup;
auto warp_group_role = WarpGroupRole(warp_group_idx);
auto producer_warp_role = ProducerWarpRole(warp_idx_in_warp_group);
// Issue Tma Descriptor Prefetch from a single thread
if (warp_idx == 0 && lane_predicate) {
CollectiveMainloop::prefetch_tma_descriptors(mainloop_params);
CollectiveEpilogue::prefetch_tma_descriptors(epilogue_params);
}
// Obtain warp index
PipelineParams pipeline_params_v;
pipeline_params_v.transaction_bytes = CollectiveMainloop::TmaTransactionBytesV;
pipeline_params_v.role = warp_group_role == WarpGroupRole::Producer
? MainloopPipeline::ThreadCategory::Producer
: MainloopPipeline::ThreadCategory::Consumer;
pipeline_params_v.is_leader = warp_group_thread_idx == 0;
pipeline_params_v.num_consumers = NumMmaThreads;
PipelineParams pipeline_params_k;
pipeline_params_k.transaction_bytes = CollectiveMainloop::TmaTransactionBytesK;
pipeline_params_k.role = warp_group_role == WarpGroupRole::Producer
? MainloopPipeline::ThreadCategory::Producer
: MainloopPipeline::ThreadCategory::Consumer;
pipeline_params_k.is_leader = warp_group_thread_idx == 0;
pipeline_params_k.num_consumers = NumMmaThreads;
PipelineParamsQ pipeline_params_q;
pipeline_params_q.transaction_bytes = CollectiveMainloop::TmaTransactionBytesQ;
pipeline_params_q.role = warp_group_role == WarpGroupRole::Producer
? MainloopPipelineQ::ThreadCategory::Producer
: MainloopPipelineQ::ThreadCategory::Consumer;
pipeline_params_q.is_leader = warp_group_thread_idx == 0;
pipeline_params_q.num_consumers = NumMmaThreads;
// We're counting on pipeline_k to call cutlass::arch::fence_barrier_init();
MainloopPipelineQ pipeline_q(shared_storage.pipeline_q, pipeline_params_q, ClusterShape{});
MainloopPipeline pipeline_k(shared_storage.pipeline_k, pipeline_params_k, ClusterShape{});
MainloopPipeline pipeline_v(shared_storage.pipeline_v, pipeline_params_v, ClusterShape{});
uint32_t epilogue_barrier_group_size_list[2] = {cutlass::NumThreadsPerWarp, NumMmaThreads};
typename EpilogueBarrier::Params params_epilogue_barrier;
params_epilogue_barrier.group_id = (warp_group_role == WarpGroupRole::Producer);
params_epilogue_barrier.group_size_list = epilogue_barrier_group_size_list;
EpilogueBarrier barrier_o(shared_storage.barrier_o, params_epilogue_barrier);
CollectiveMainloop collective_mainloop;
CollectiveEpilogue collective_epilogue;
__syncthreads();
if (warp_group_role == WarpGroupRole::Producer) {
cutlass::arch::warpgroup_reg_dealloc<24>();
TileScheduler scheduler;
if (producer_warp_role == ProducerWarpRole::Mainloop) { // Load Q, K, V
PipelineStateQ smem_pipe_write_q = cutlass::make_producer_start_state<MainloopPipelineQ>();
PipelineState smem_pipe_write_k = cutlass::make_producer_start_state<MainloopPipeline>();
PipelineState smem_pipe_write_v = cutlass::make_producer_start_state<MainloopPipeline>();
int work_idx = 0;
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
int tile_count_semaphore = 0;
collective_mainloop.load(mainloop_params, scheduler_params,
pipeline_q, pipeline_k, pipeline_v,
smem_pipe_write_q, smem_pipe_write_k, smem_pipe_write_v,
shared_storage, work_tile_info, work_idx, tile_count_semaphore);
}
collective_mainloop.load_tail(pipeline_q, pipeline_k, pipeline_v,
smem_pipe_write_q, smem_pipe_write_k, smem_pipe_write_v);
} else if (producer_warp_role == ProducerWarpRole::Epilogue) {
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
barrier_o.wait();
collective_epilogue.tma_store(shared_storage, epilogue_params, work_tile_info, scheduler_params, threadIdx.x);
collective_epilogue.store_tail();
barrier_o.arrive();
}
}
} else if (warp_group_role == WarpGroupRole::Consumer0 || warp_group_role == WarpGroupRole::Consumer1) {
cutlass::arch::warpgroup_reg_alloc<232>();
typename Ktraits::TiledMmaPV tiled_mma_pv;
TileScheduler scheduler{};
PipelineState smem_pipe_read_k, smem_pipe_read_v;
PipelineStateQ smem_pipe_read_q;
int work_idx = 0;
CUTLASS_PRAGMA_NO_UNROLL
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
// Attention output (GEMM-II) accumulator.
Tensor tOrO = partition_fragment_C(tiled_mma_pv, select<0, 2>(TileShape_MNK{}));
// flash::Softmax<2 * (2 * kBlockM / NumMmaThreads)> softmax;
// Pass single_level_p_quant flag to control P quantization mode
flash::SoftmaxFused<2 * (2 * kBlockM / NumMmaThreads)> softmax_fused(params.single_level_p_quant);
auto block_coord = work_tile_info.get_block_coord(scheduler_params);
auto [m_block, bidh, bidb] = block_coord;
int n_block_max = collective_mainloop.get_n_block_max(mainloop_params, m_block);
if (Is_causal && n_block_max <= 0) { // We exit early and write 0 to gO and -inf to gLSE.
collective_epilogue.store_zero(epilogue_params, threadIdx.x - NumCopyThreads, block_coord);
continue;
}
collective_mainloop.mma(mainloop_params, pipeline_q, pipeline_k, pipeline_v, smem_pipe_read_q, smem_pipe_read_k, smem_pipe_read_v,
tOrO, softmax_fused, n_block_max, threadIdx.x - NumCopyThreads, work_idx, m_block, shared_storage);
barrier_o.wait();
collective_epilogue.mma_store(shared_storage, tiled_mma_pv, tOrO, threadIdx.x - NumCopyThreads);
barrier_o.arrive();
++work_idx;
}
}
}
} // namespace flash
@@ -1,114 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <ATen/cuda/CUDAContext.h>
#include "cute/tensor.hpp"
#include "cutlass/cluster_launch.hpp"
#include "static_switch.h"
#include "params.h"
#include "tile_scheduler.h"
#include "kernel_ws.h"
#include "kernel_traits.h"
#include "block_config.h"
template<typename Kernel_traits, bool Is_causal>
void run_flash_fwd(Flash_fwd_params &params, cudaStream_t stream) {
using Element = typename Kernel_traits::Element;
using ElementSF = typename Kernel_traits::ElementSF;
using ElementOut = typename Kernel_traits::ElementOut;
using TileShape_MNK = typename Kernel_traits::TileShape_MNK;
using ClusterShape = typename Kernel_traits::ClusterShape_MNK;
using CollectiveMainloop = flash::CollectiveMainloopFwd<Kernel_traits, Is_causal>;
using CollectiveEpilogue = flash::CollectiveEpilogueFwd<Kernel_traits>;
// using Scheduler = flash::SingleTileScheduler;
using Scheduler = flash::StaticPersistentTileScheduler;
typename CollectiveMainloop::Params mainloop_params =
CollectiveMainloop::to_underlying_arguments({
static_cast<Element const*>(params.q_ptr),
{params.seqlen_q, params.d, params.h, params.b}, // shape_Q
{params.q_row_stride, _1{}, params.q_head_stride, params.q_batch_stride}, // stride_Q
static_cast<Element const*>(params.k_ptr),
{params.seqlen_k, params.d, params.h_k, params.b}, // shape_K
{params.k_row_stride, _1{}, params.k_head_stride, params.k_batch_stride}, // stride_K
{params.unpadded_seqlen_k, params.d, params.h_k, params.b}, // shape_K
static_cast<Element const*>(params.v_ptr),
{params.d, params.seqlen_k, params.h_k, params.b}, // shape_Vt
{params.v_row_stride, _1{}, params.v_head_stride, params.v_batch_stride}, // stride_Vt
static_cast<ElementSF const*>(params.sfq_ptr),
{params.seqlen_q, params.d, params.h, params.b}, // shape_SFQ
static_cast<ElementSF const*>(params.sfk_ptr),
{params.seqlen_k, params.d, params.h_k, params.b}, // shape_SFK
static_cast<ElementSF const*>(params.sfv_ptr),
{params.d, params.seqlen_k, params.h_k, params.b}, // shape_SFVt
static_cast<float const*>(params.delta_s_ptr),
{params.seqlen_s, params.seqlen_k, params.h_k, params.b},
{params.ds_row_stride, _1{}, params.ds_head_stride, params.ds_batch_stride},
params.scale_softmax_log2
});
typename CollectiveEpilogue::Params epilogue_params =
CollectiveEpilogue::to_underlying_arguments({
static_cast<ElementOut*>(params.o_ptr),
{params.seqlen_q, params.d, params.h, params.b}, // shape_O
{params.o_row_stride, _1{}, params.o_head_stride, params.o_batch_stride}, // stride_O
static_cast<float*>(params.softmax_lse_ptr),
{_1{}, params.seqlen_q, params.h * params.seqlen_q}, // stride_LSE
});
int num_blocks_m = cutlass::ceil_div(params.seqlen_q, Kernel_traits::kBlockM);
num_blocks_m = cutlass::ceil_div(num_blocks_m, size<0>(ClusterShape{})) * size<0>(ClusterShape{});
typename Scheduler::Arguments scheduler_args = {num_blocks_m, params.h, params.b};
typename Scheduler::Params scheduler_params = Scheduler::to_underlying_arguments(scheduler_args);
// Get the ptr to kernel function.
void *kernel;
kernel = (void *)flash::compute_attn_ws<Kernel_traits, Is_causal, Scheduler>;
int smem_size = sizeof(typename Kernel_traits::SharedStorage);
if (smem_size >= 48 * 1024) {
C10_CUDA_CHECK(cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, smem_size));
}
static constexpr int ctaSize = Kernel_traits::kNWarps * 32;
params.m_block_divmod = cutlass::FastDivmod(num_blocks_m);
params.total_blocks = num_blocks_m * params.h * params.b;
dim3 grid_dims = Scheduler::get_grid_dim(scheduler_args, 170);
dim3 block_dims(ctaSize);
dim3 cluster_dims(size<0>(ClusterShape{}), size<1>(ClusterShape{}), size<2>(ClusterShape{}));
cutlass::ClusterLaunchParams launch_params{grid_dims, block_dims, cluster_dims, smem_size, stream};
cutlass::launch_kernel_on_cluster(launch_params, kernel, params, mainloop_params, epilogue_params, scheduler_params);
C10_CUDA_KERNEL_LAUNCH_CHECK();
}
template<typename T, int Headdim, typename O = cutlass::bfloat16_t>
void run_mha_fwd_(Flash_fwd_params &params, cudaStream_t stream) {
BOOL_SWITCH(params.is_causal, Is_causal, [&] {
BOOL_SWITCH(params.per_block_mean, per_block, [&] {
if constexpr (Headdim == 64 || Headdim == 128) {
run_flash_fwd<
Flash_fwd_kernel_traits<Headdim, flash::BLOCK_M, flash::BLOCK_N, 3, 1, per_block, T, O>,
Is_causal
>(params, stream);
} else {
static_assert(Headdim == 64 || Headdim == 128, "Unsupported Headdim");
}
});
});
}
@@ -1,920 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <cutlass/cutlass.h>
#include <cutlass/array.h>
#include <cutlass/numeric_types.h>
#include <cutlass/numeric_conversion.h>
#include "cutlass/pipeline/pipeline.hpp"
#include "cute/tensor.hpp"
#include "cutlass/gemm/collective/collective_builder.hpp"
#include "utils.h"
#include "named_barrier.h"
namespace flash {
using namespace cute;
template <typename Ktraits, bool Is_causal>
struct CollectiveMainloopFwd {
using Element = typename Ktraits::Element;
using ElementSF = typename Ktraits::ElementSF;
// using TMAElement = Element;
// using TMAElementSF = typename Ktraits::ElementSF;
using TileShape_MNK = typename Ktraits::TileShape_MNK;
using ClusterShape = typename Ktraits::ClusterShape_MNK;
static constexpr int kStages = Ktraits::kStages;
static constexpr int kHeadDim = Ktraits::kHeadDim;
static constexpr int BlockMean = Ktraits::BlockMean;
using GmemTiledCopy = typename Ktraits::GmemTiledCopy;
using SmemLayoutQ = typename Ktraits::SmemLayoutQ;
using SmemLayoutK = typename Ktraits::SmemLayoutK;
using SmemLayoutV = typename Ktraits::SmemLayoutV;
using SmemLayoutVt = typename Ktraits::SmemLayoutVt;
using SmemLayoutDS = typename Ktraits::SmemLayoutDS;
using SmemLayoutAtomDS = typename Ktraits::SmemLayoutAtomDS;
using LayoutDS = decltype(
blocked_product(
SmemLayoutAtomDS{},
make_layout(
make_shape(int32_t(0), int32_t(0), int32_t(0), int32_t(0)),
make_stride(int32_t(0), _1{}, int32_t(0), int32_t(0)))
)
);
using ShapeQKV = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen, d, head, batch)
using StrideQKV = cute::Stride<int64_t, _1, int64_t, int64_t>;
using ShapeSF = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen, d // 16, head, batch)
using LayoutSF = typename Ktraits::LayoutSF;
using LayoutP = typename Ktraits::LayoutP;
using LayoutSFP = typename Ktraits::LayoutSFP;
using SfAtom = typename Ktraits::SfAtom;
using TMA_Q = decltype(make_tma_copy(
GmemTiledCopy{},
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
SmemLayoutQ{},
select<0, 2>(TileShape_MNK{}),
_1{}));
using TMA_KV = decltype(make_tma_copy(
GmemTiledCopy{},
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
take<0, 2>(SmemLayoutK{}),
select<1, 2>(TileShape_MNK{}),
_1{}));
using TMA_Vt = decltype(make_tma_copy(
GmemTiledCopy{},
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
take<0, 2>(SmemLayoutVt{}),
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{}));
using TMA_DS = decltype(make_tma_copy(
GmemTiledCopy{},
make_tensor(make_gmem_ptr(static_cast<float const*>(nullptr)), LayoutDS{}),
take<0, 2>(SmemLayoutDS{}),
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{}));
using BlkScaledConfig = typename Ktraits::BlkScaledConfig;
using GmemTiledCopySF = typename Ktraits::GmemTiledCopySF;
using SmemLayoutSFQ = typename Ktraits::SmemLayoutSFQ;
using SmemLayoutSFK = typename Ktraits::SmemLayoutSFK;
using SmemLayoutSFV = typename Ktraits::SmemLayoutSFV;
using SmemLayoutSFVt = typename Ktraits::SmemLayoutSFVt;
using TMA_SFQ = decltype(make_tma_copy<uint16_t>(
GmemTiledCopySF{},
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
SmemLayoutSFQ{},
make_shape(shape<0>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
_1{})); // No programmatic multicast
using TMA_SFKV = decltype(make_tma_copy<uint16_t>(
GmemTiledCopySF{},
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
SmemLayoutSFK{}(_,_,cute::Int<0>{}),
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
_1{}));
using TMA_SFVt = decltype(make_tma_copy<uint16_t>(
GmemTiledCopySF{},
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
SmemLayoutSFVt{}(_,_,cute::Int<0>{}),
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{}));
using SmemCopyAtomQ = typename Ktraits::SmemCopyAtomQ;
using SmemCopyAtomKV = typename Ktraits::SmemCopyAtomKV;
using SmemCopyAtomSF = typename Ktraits::SmemCopyAtomSF;
using TiledMmaQK = typename Ktraits::TiledMmaQK;
using TiledMmaPV = typename Ktraits::TiledMmaPV;
static constexpr int NumMmaThreads = size(TiledMmaQK{});
using MainloopPipeline = typename Ktraits::MainloopPipeline;
using PipelineParams = typename MainloopPipeline::Params;
using PipelineState = typename MainloopPipeline::PipelineState;
using MainloopPipelineQ = typename Ktraits::MainloopPipelineQ;
using PipelineParamsQ = typename Ktraits::PipelineParamsQ;
using PipelineStateQ = typename Ktraits::PipelineStateQ;
using EpilogueBarrier = typename Ktraits::EpilogueBarrier;
// Set the bytes transferred in this TMA transaction (may involve multiple issues)
static constexpr uint32_t TmaTransactionBytesQ = static_cast<uint32_t>(
cutlass::bits_to_bytes(cosize((SmemLayoutSFQ{})) * cute::sizeof_bits_v<ElementSF>) +
cutlass::bits_to_bytes(size((SmemLayoutQ{})) * sizeof_bits<Element>::value));
static constexpr uint32_t TmaTransactionBytesK = static_cast<uint32_t>(
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutSFK{})) * cute::sizeof_bits_v<ElementSF>) +
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutDS{})) * cute::sizeof_bits_v<float>) +
cutlass::bits_to_bytes(size(take<0,2>(SmemLayoutK{})) * sizeof_bits<Element>::value));
static constexpr uint32_t TmaTransactionBytesV = static_cast<uint32_t>(
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutSFVt{})) * cute::sizeof_bits_v<ElementSF>) +
cutlass::bits_to_bytes(size(take<0,2>(SmemLayoutVt{})) * sizeof_bits<Element>::value));
// Host side kernel arguments
struct Arguments {
Element const* ptr_Q;
ShapeQKV const shape_Q;
StrideQKV const stride_Q;
Element const* ptr_K;
ShapeQKV const shape_K;
StrideQKV const stride_K;
ShapeQKV const unpadded_shape_K;
Element const* ptr_Vt;
ShapeQKV const shape_Vt;
StrideQKV const stride_Vt;
ElementSF const* ptr_SFQ{nullptr};
ShapeSF const shape_SFQ{};
ElementSF const* ptr_SFK{nullptr};
ShapeSF const shape_SFK{};
ElementSF const* ptr_SFVt{nullptr};
ShapeSF const shape_SFVt{};
float const* ptr_ds;
ShapeQKV const shape_ds;
StrideQKV const stride_ds;
float const softmax_scale_log2;
};
// Device side kernel params
struct Params {
ShapeQKV const shape_Q;
LayoutSF const layout_SFQ;
ShapeQKV const shape_K;
ShapeQKV const unpadded_shape_K;
LayoutSF const layout_SFK;
ShapeQKV const shape_Vt;
LayoutSF const layout_SFVt;
LayoutDS const layout_DS;
TMA_Q tma_load_Q;
TMA_SFQ tma_load_SFQ;
TMA_KV tma_load_K;
TMA_SFKV tma_load_SFK;
TMA_Vt tma_load_Vt;
TMA_SFVt tma_load_SFVt;
TMA_DS tma_load_DS;
float const softmax_scale_log2;
};
static Params
to_underlying_arguments(Arguments const& args) {
Tensor mQ = make_tensor(make_gmem_ptr(args.ptr_Q), args.shape_Q, args.stride_Q);
TMA_Q tma_load_Q = make_tma_copy(
GmemTiledCopy{},
mQ,
SmemLayoutQ{},
select<0, 2>(TileShape_MNK{}),
_1{}); // no mcast for Q
Tensor mK = make_tensor(make_gmem_ptr(args.ptr_K), args.shape_K, args.stride_K);
TMA_KV tma_load_K = make_tma_copy(
GmemTiledCopy{},
mK,
SmemLayoutK{}(_, _, _0{}),
select<1, 2>(TileShape_MNK{}),
_1{}); // mcast along M mode for this N load, if any
Tensor mVt = make_tensor(make_gmem_ptr(args.ptr_Vt), args.shape_Vt, args.stride_Vt);
TMA_Vt tma_load_Vt = make_tma_copy(
GmemTiledCopy{},
mVt,
SmemLayoutVt{}(_, _, _0{}),
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{}); // mcast along M mode for this N load, if any
auto [Seqlen_Q, Seqlen_K, HeadNum, Batch] = args.shape_ds;
LayoutDS layout_ds = tile_to_shape(SmemLayoutAtomDS{}, make_shape(Seqlen_Q, Seqlen_K, HeadNum, Batch), Step<_2,_1,_3,_4>{});
Tensor mDS = make_tensor(make_gmem_ptr(args.ptr_ds), layout_ds);
TMA_DS tma_load_ds = make_tma_copy (
GmemTiledCopy{},
mDS,
SmemLayoutDS{}(_, _, _0{}),
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{});
LayoutSF layout_sfq = BlkScaledConfig::tile_atom_to_shape_SFQKV(args.shape_SFQ);
Tensor mSFQ = make_tensor(make_gmem_ptr(args.ptr_SFQ), layout_sfq);
TMA_SFQ tma_load_sfq = make_tma_copy<uint16_t>(
GmemTiledCopySF{},
mSFQ,
SmemLayoutSFQ{},
make_shape(shape<0>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
_1{});
LayoutSF layout_sfk = BlkScaledConfig::tile_atom_to_shape_SFQKV(args.shape_SFK);
Tensor mSFK = make_tensor(make_gmem_ptr(args.ptr_SFK), layout_sfk);
TMA_SFKV tma_load_sfk = make_tma_copy<uint16_t>(
GmemTiledCopySF{},
mSFK,
SmemLayoutSFK{}(_, _, _0{}),
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
_1{});
LayoutSF layout_sfvt = BlkScaledConfig::tile_atom_to_shape_SFVt(args.shape_SFVt);
Tensor mSFVt = make_tensor(make_gmem_ptr(args.ptr_SFVt), layout_sfvt);
TMA_SFVt tma_load_sfvt = make_tma_copy<uint16_t>(
GmemTiledCopySF{},
mSFVt,
SmemLayoutSFVt{}(_, _, _0{}),
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
_1{});
return {args.shape_Q, layout_sfq,
args.shape_K, args.unpadded_shape_K, layout_sfk,
args.shape_Vt, layout_sfvt,
layout_ds,
tma_load_Q, tma_load_sfq,
tma_load_K, tma_load_sfk,
tma_load_Vt, tma_load_sfvt,
tma_load_ds,
args.softmax_scale_log2};
}
/// Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
CUTLASS_DEVICE
static void prefetch_tma_descriptors(Params const& mainloop_params) {
cute::prefetch_tma_descriptor(mainloop_params.tma_load_Q.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_K.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_Vt.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFQ.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFK.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFVt.get_tma_descriptor());
cute::prefetch_tma_descriptor(mainloop_params.tma_load_DS.get_tma_descriptor());
}
CUTLASS_DEVICE
int get_n_block_max(Params const& mainloop_params, int m_block) {
static constexpr int kBlockM = get<0>(TileShape_MNK{});
static constexpr int kBlockN = get<1>(TileShape_MNK{});
int const seqlen_q = get<0>(mainloop_params.shape_Q);
int const seqlen_k = get<0>(mainloop_params.shape_K);
int n_block_max = cute::ceil_div(seqlen_k, kBlockN);
if constexpr (Is_causal) {
n_block_max = std::min(n_block_max,
cute::ceil_div((m_block + 1) * kBlockM + seqlen_k - seqlen_q, kBlockN));
}
return n_block_max;
}
template <class SFATensor, class Atom, class TiledThr, class TiledPerm>
CUTE_HOST_DEVICE constexpr
auto
thrfrg_SFA(SFATensor&& sfatensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
{
CUTE_STATIC_ASSERT_V(rank(sfatensor) >= Int<2>{});
using AtomShape_MNK = typename Atom::Shape_MNK;
using AtomLayoutSFA_TV = typename Atom::Traits::SFALayout;
auto permutation_mnk = TiledPerm{};
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// Reorder the tensor for the TiledAtom
auto t_tile = make_tile(get<0>(permutation_mnk),
get<2>(permutation_mnk));
auto t_tensor = logical_divide(sfatensor, t_tile); // (PermM,PermK)
// Tile the tensor for the Atom
auto a_tile = make_tile(make_layout(size<0>(AtomShape_MNK{})),
make_layout(size<2>(AtomShape_MNK{})));
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomM,AtomK),(RestM,RestK))
// Transform the Atom mode from (M,K) to (Thr,Val)
auto tv_tensor = a_tensor.compose(AtomLayoutSFA_TV{},_); // ((ThrV,FrgV),(RestM,RestK))
// Tile the tensor for the Thread
auto thr_tile = make_tile(_,
make_tile(make_layout(size<1>(thr_layout_vmnk)),
make_layout(size<3>(thr_layout_vmnk))));
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrM,ThrK)),(FrgV,(RestM,RestK)))
return thr_tensor;
}
template <class SFBTensor, class Atom, class TiledThr, class TiledPerm>
CUTE_HOST_DEVICE constexpr
auto
thrfrg_SFB(SFBTensor&& sfbtensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
{
CUTE_STATIC_ASSERT_V(rank(sfbtensor) >= Int<2>{});
using AtomShape_MNK = typename Atom::Shape_MNK;
using AtomLayoutSFB_TV = typename Atom::Traits::SFBLayout;
auto permutation_mnk = TiledPerm{};
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// Reorder the tensor for the TiledAtom
auto t_tile = make_tile(get<1>(permutation_mnk),
get<2>(permutation_mnk));
auto t_tensor = logical_divide(sfbtensor, t_tile); // (PermN,PermK)
// Tile the tensor for the Atom
auto a_tile = make_tile(make_layout(size<1>(AtomShape_MNK{})),
make_layout(size<2>(AtomShape_MNK{})));
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomN,AtomK),(RestN,RestK))
// Transform the Atom mode from (M,K) to (Thr,Val)
auto tv_tensor = a_tensor.compose(AtomLayoutSFB_TV{},_); // ((ThrV,FrgV),(RestN,RestK))
// Tile the tensor for the Thread
auto thr_tile = make_tile(_,
make_tile(make_layout(size<2>(thr_layout_vmnk)),
make_layout(size<3>(thr_layout_vmnk))));
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrN,ThrK)),(FrgV,(RestN,RestK)))
return thr_tensor;
}
template <class SFATensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_fragment_SFA(SFATensor&& sfatensor, ThrMma& thread_mma)
{
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
auto thr_tensor = make_tensor(static_cast<SFATensor&&>(sfatensor).data(), thrfrg_SFA(sfatensor.layout(),thread_mma));
auto thr_vmnk = thread_mma.thr_vmnk_;
auto thr_vmk = make_coord(get<0>(thr_vmnk), make_coord(get<1>(thr_vmnk), get<3>(thr_vmnk)));
auto partition_SFA = thr_tensor(thr_vmk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
return make_fragment_like<ValTypeSF>(partition_SFA);
}
template <class SFBTensor, class ThrMma>
CUTE_HOST_DEVICE constexpr
auto
partition_fragment_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma)
{
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
auto thr_tensor = make_tensor(static_cast<SFBTensor&&>(sfbtensor).data(), thrfrg_SFB(sfbtensor.layout(),thread_mma));
auto thr_vmnk = thread_mma.thr_vmnk_;
auto thr_vnk = make_coord(get<0>(thr_vmnk), make_coord(get<2>(thr_vmnk), get<3>(thr_vmnk)));
auto partition_SFB = thr_tensor(thr_vnk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
return make_fragment_like<ValTypeSF>(partition_SFB);
}
template<class TiledMma>
CUTE_HOST_DEVICE constexpr
auto
get_layoutSFA_TV(TiledMma& mma)
{
// (M,K) -> (M,K)
auto tile_shape_mnk = tile_shape(mma);
auto ref_A = make_layout(make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk)));
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
auto atile = make_tile(_,
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
make_stride( Int<1>{} , Int<0>{} )),
_));
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
// (thr_idx,val) -> (M,K)
return thrfrg_SFA(ref_A, mma).compose(atile, _).compose(thridx_2_thrid, _);
}
template<class TiledMma>
CUTE_HOST_DEVICE constexpr
auto
get_layoutSFB_TV(TiledMma& mma)
{
// (N,K) -> (N,K)
auto tile_shape_mnk = tile_shape(mma);
auto ref_B = make_layout(make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk)));
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
auto btile = make_tile(_,
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
make_stride( Int<0>{} , Int<1>{} )),
_));
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
// (thr_idx,val) -> (M,K)
return thrfrg_SFB(ref_B, mma).compose(btile, _).compose(thridx_2_thrid, _);
}
template <typename SchedulerParams, typename SharedStorage, typename WorkTileInfo>
CUTLASS_DEVICE void
load(Params const& mainloop_params,
SchedulerParams const& scheduler_params,
MainloopPipelineQ pipeline_q,
MainloopPipeline pipeline_k,
MainloopPipeline pipeline_v,
PipelineStateQ& smem_pipe_write_q,
PipelineState& smem_pipe_write_k,
PipelineState& smem_pipe_write_v,
SharedStorage &shared_storage,
WorkTileInfo work_tile_info,
int& work_idx,
int& tile_count_semaphore
) {
static constexpr int kBlockM = get<0>(TileShape_MNK{});
static constexpr int kBlockN = get<1>(TileShape_MNK{});
auto [m_block, bidh, bidb] = work_tile_info.get_block_coord(scheduler_params);
int n_block_max = get_n_block_max(mainloop_params, m_block);
Tensor sQ = make_tensor(make_smem_ptr(shared_storage.smem_q.begin()), SmemLayoutQ{});
Tensor sK = make_tensor(make_smem_ptr(shared_storage.smem_k.begin()), SmemLayoutK{});
Tensor sVt = make_tensor(make_smem_ptr(shared_storage.smem_v.begin()), SmemLayoutVt{});
Tensor sSFQ = make_tensor(make_smem_ptr(shared_storage.smem_SFQ.begin()), SmemLayoutSFQ{});
Tensor sSFK = make_tensor(make_smem_ptr(shared_storage.smem_SFK.begin()), SmemLayoutSFK{});
Tensor sSFVt = make_tensor(make_smem_ptr(shared_storage.smem_SFV.begin()), SmemLayoutSFVt{});
Tensor sDS = make_tensor(make_smem_ptr(shared_storage.smem_ds.begin()), SmemLayoutDS{});
Tensor mQ = mainloop_params.tma_load_Q.get_tma_tensor(mainloop_params.shape_Q);
Tensor mK = mainloop_params.tma_load_K.get_tma_tensor(mainloop_params.shape_K);
Tensor mVt = mainloop_params.tma_load_Vt.get_tma_tensor(mainloop_params.shape_Vt);
Tensor mDS = mainloop_params.tma_load_DS.get_tma_tensor(shape(mainloop_params.layout_DS));
Tensor mSFQ = mainloop_params.tma_load_SFQ.get_tma_tensor(shape(mainloop_params.layout_SFQ));
Tensor mSFK = mainloop_params.tma_load_SFK.get_tma_tensor(shape(mainloop_params.layout_SFK));
Tensor mSFVt = mainloop_params.tma_load_SFVt.get_tma_tensor(shape(mainloop_params.layout_SFVt));
uint32_t block_rank_in_cluster = cute::block_rank_in_cluster();
constexpr uint32_t cluster_shape_x = get<0>(ClusterShape());
uint2 cluster_local_block_id = {block_rank_in_cluster % cluster_shape_x, block_rank_in_cluster / cluster_shape_x};
Tensor gQ = local_tile(mQ(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
Tensor gK = local_tile(mK(_, _, bidh, bidb), select<1, 2>(TileShape_MNK{}), make_coord(_, _0{})); // (N, K, _)
Tensor gVt = local_tile(mVt(_, _, bidh, bidb), make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})), make_coord(_0{}, _)); // (N, K, _)
Tensor gDS = [&] {
if constexpr (BlockMean) {
return local_tile(mDS(_, _, bidh, bidb), select<0, 1>(TileShape_MNK{}), make_coord(m_block, _));
} else {
return local_tile(mDS(_, _, bidh, bidb), select<0, 1>(TileShape_MNK{}), make_coord(_0{}, _));
}
}();
Tensor gSFQ = local_tile(mSFQ(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{}));
Tensor gSFK = local_tile(mSFK(_, _, bidh, bidb), select<1, 2>(TileShape_MNK{}), make_coord(_, _0{}));
Tensor gSFVt = local_tile(mSFVt(_, _, bidh, bidb), make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})), make_coord(_0{}, _));
auto block_tma_q = mainloop_params.tma_load_Q.get_slice(_0{});
Tensor tQgQ = block_tma_q.partition_S(gQ);
Tensor tQsQ = block_tma_q.partition_D(sQ);
auto block_tma_sfq = mainloop_params.tma_load_SFQ.get_slice(_0{});
Tensor tQgSFQ = block_tma_sfq.partition_S(gSFQ);
Tensor tQsSFQ = block_tma_sfq.partition_D(sSFQ);
auto block_tma_k = mainloop_params.tma_load_K.get_slice(cluster_local_block_id.x);
Tensor tKgK = group_modes<0, 3>(block_tma_k.partition_S(gK));
Tensor tKsK = group_modes<0, 3>(block_tma_k.partition_D(sK));
auto block_tma_sfk = mainloop_params.tma_load_SFK.get_slice(cluster_local_block_id.x);
Tensor tKgSFK = group_modes<0, 3>(block_tma_sfk.partition_S(gSFK));
Tensor tKsSFK = group_modes<0, 3>(block_tma_sfk.partition_D(sSFK));
auto block_tma_vt = mainloop_params.tma_load_Vt.get_slice(cluster_local_block_id.x);
Tensor tVgVt = group_modes<0, 3>(block_tma_vt.partition_S(gVt));
Tensor tVsVt = group_modes<0, 3>(block_tma_vt.partition_D(sVt));
auto block_tma_sfvt = mainloop_params.tma_load_SFVt.get_slice(cluster_local_block_id.x);
Tensor tVgSFVt = group_modes<0, 3>(block_tma_sfvt.partition_S(gSFVt));
Tensor tVsSFVt = group_modes<0, 3>(block_tma_sfvt.partition_D(sSFVt));
auto block_tma_ds = mainloop_params.tma_load_DS.get_slice(cluster_local_block_id.x);
Tensor tDSgDS = group_modes<0, 3>(block_tma_ds.partition_S(gDS));
Tensor tDSsDS = group_modes<0, 3>(block_tma_ds.partition_D(sDS));
uint16_t mcast_mask_kv = 0;
int n_block = n_block_max - 1;
int lane_predicate = cute::elect_one_sync();
if (lane_predicate) {
pipeline_q.producer_acquire(smem_pipe_write_q);
copy(mainloop_params.tma_load_Q.with(*pipeline_q.producer_get_barrier(smem_pipe_write_q), 0), tQgQ, tQsQ);
copy(mainloop_params.tma_load_SFQ.with(*pipeline_q.producer_get_barrier(smem_pipe_write_q), 0), tQgSFQ, tQsSFQ);
++smem_pipe_write_q;
pipeline_k.producer_acquire(smem_pipe_write_k);
copy(mainloop_params.tma_load_K.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tKgK(_, n_block), tKsK(_, smem_pipe_write_k.index()));
copy(mainloop_params.tma_load_SFK.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tKgSFK(_, n_block), tKsSFK(_, smem_pipe_write_k.index()));
copy(mainloop_params.tma_load_DS.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tDSgDS(_, n_block), tDSsDS(_, smem_pipe_write_k.index()));
++smem_pipe_write_k;
pipeline_v.producer_acquire(smem_pipe_write_v);
copy(mainloop_params.tma_load_Vt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
tVgVt(_, n_block), tVsVt(_, smem_pipe_write_v.index()));
copy(mainloop_params.tma_load_SFVt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
tVgSFVt(_, n_block), tVsSFVt(_, smem_pipe_write_v.index()));
++smem_pipe_write_v;
}
n_block--;
if (lane_predicate) {
// CUTLASS_PRAGMA_NO_UNROLL
#pragma unroll 2
for (; n_block >= 0; --n_block) {
pipeline_k.producer_acquire(smem_pipe_write_k);
copy(mainloop_params.tma_load_K.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tKgK(_, n_block), tKsK(_, smem_pipe_write_k.index()));
copy(mainloop_params.tma_load_SFK.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tKgSFK(_, n_block), tKsSFK(_, smem_pipe_write_k.index()));
copy(mainloop_params.tma_load_DS.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
tDSgDS(_, n_block), tDSsDS(_, smem_pipe_write_k.index()));
++smem_pipe_write_k;
pipeline_v.producer_acquire(smem_pipe_write_v);
copy(mainloop_params.tma_load_Vt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
tVgVt(_, n_block), tVsVt(_, smem_pipe_write_v.index()));
copy(mainloop_params.tma_load_SFVt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
tVgSFVt(_, n_block), tVsSFVt(_, smem_pipe_write_v.index()));
++smem_pipe_write_v;
}
}
++work_idx;
}
/// Perform a Producer Epilogue to prevent early exit of blocks in a Cluster
CUTLASS_DEVICE void
load_tail(MainloopPipelineQ pipeline_q,
MainloopPipeline pipeline_k,
MainloopPipeline pipeline_v,
PipelineStateQ& smem_pipe_write_q,
PipelineState& smem_pipe_write_k,
PipelineState& smem_pipe_write_v) {
int lane_predicate = cute::elect_one_sync();
// Issue the epilogue waits
if (lane_predicate) {
pipeline_q.producer_tail(smem_pipe_write_q);
pipeline_k.producer_tail(smem_pipe_write_k);
pipeline_v.producer_tail(smem_pipe_write_v);
}
}
template <typename SharedStorage, typename FrgTensorO, typename SoftmaxFused>
CUTLASS_DEVICE void
mma(Params const& mainloop_params,
MainloopPipelineQ pipeline_q,
MainloopPipeline pipeline_k,
MainloopPipeline pipeline_v,
PipelineStateQ& smem_pipe_read_q,
PipelineState& smem_pipe_read_k,
PipelineState& smem_pipe_read_v,
FrgTensorO& tOrO_store,
SoftmaxFused& softmax_fused,
int n_block_count,
int thread_idx,
int work_idx,
int m_block,
SharedStorage& shared_storage
) {
static_assert(is_rmem<FrgTensorO>::value, "O tensor must be rmem resident.");
static constexpr int kBlockM = get<0>(TileShape_MNK{});
static constexpr int kBlockN = get<1>(TileShape_MNK{});
static constexpr int kBlockK = get<2>(TileShape_MNK{});
Tensor sQ = make_tensor(make_smem_ptr(shared_storage.smem_q.begin()), SmemLayoutQ{});
Tensor sK = make_tensor(make_smem_ptr(shared_storage.smem_k.begin()), SmemLayoutK{});
Tensor sVt = make_tensor(make_smem_ptr(shared_storage.smem_v.begin()), SmemLayoutVt{});
Tensor sDS = make_tensor(make_smem_ptr(shared_storage.smem_ds.begin()), SmemLayoutDS{});
Tensor sSFQ = make_tensor(make_smem_ptr(shared_storage.smem_SFQ.begin()), SmemLayoutSFQ{});
Tensor sSFK = make_tensor(make_smem_ptr(shared_storage.smem_SFK.begin()), SmemLayoutSFK{});
Tensor sSFVt = make_tensor(make_smem_ptr(shared_storage.smem_SFV.begin()), SmemLayoutSFVt{});
Tensor cQ = make_identity_tensor(make_shape(size<0>(sQ), size<1>(sQ)));
Tensor cKV = make_identity_tensor(make_shape(size<0>(sK), size<1>(sK)));
TiledMmaQK tiled_mma_qk;
TiledMmaPV tiled_mma_pv;
auto thread_mma_qk = tiled_mma_qk.get_thread_slice(thread_idx);
auto thread_mma_pv = tiled_mma_pv.get_thread_slice(thread_idx);
Tensor tSrQ = thread_mma_qk.partition_fragment_A(sQ);
Tensor tSrK = thread_mma_qk.partition_fragment_B(sK(_,_,Int<0>{}));
Tensor tOrVt = thread_mma_pv.partition_fragment_B(sVt(_,_,Int<0>{}));
Tensor tOrP = make_tensor_like<Element>(LayoutP{});
Tensor tSrSFQ = partition_fragment_SFA(sSFQ, thread_mma_qk);
Tensor tSrSFK = partition_fragment_SFB(sSFK(_,_,Int<0>{}), thread_mma_qk);
Tensor tOrSFVt = partition_fragment_SFB(sSFVt(_,_,Int<0>{}), thread_mma_pv);
Tensor tOrSFP = make_tensor<ElementSF>(LayoutSFP{});
Tensor tOrSFP_flt = filter_zeros(tOrSFP);
Tensor tSrDS = make_tensor<float>(make_shape(_8{}, _4{}), make_stride(_1{}, _8{}));
// copy qk and sf from smem to rmem
auto smem_tiled_copy_Q = make_tiled_copy_A(SmemCopyAtomQ{}, tiled_mma_qk);
auto smem_thr_copy_Q = smem_tiled_copy_Q.get_thread_slice(thread_idx);
Tensor tSsQ = smem_thr_copy_Q.partition_S(as_position_independent_swizzle_tensor(sQ));
Tensor tSrQ_copy_view = smem_thr_copy_Q.retile_D(tSrQ);
auto smem_tiled_copy_K = make_tiled_copy_B(SmemCopyAtomKV{}, tiled_mma_qk);
auto smem_thr_copy_K = smem_tiled_copy_K.get_thread_slice(thread_idx);
Tensor tSsK = smem_thr_copy_K.partition_S(as_position_independent_swizzle_tensor(sK));
Tensor tSrK_copy_view = smem_thr_copy_K.retile_D(tSrK);
auto smem_tiled_copy_V = make_tiled_copy_B(SmemCopyAtomKV{}, tiled_mma_pv);
auto smem_thr_copy_V = smem_tiled_copy_V.get_thread_slice(thread_idx);
Tensor tOsVt = smem_thr_copy_V.partition_S(as_position_independent_swizzle_tensor(sVt));
Tensor tOrVt_copy_view = smem_thr_copy_V.retile_D(tOrVt);
auto tile_shape_mnk = tile_shape(tiled_mma_qk);
auto smem_tiled_copy_SFQ = make_tiled_copy_impl(SmemCopyAtomSF{},
get_layoutSFA_TV(tiled_mma_qk),
make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk))
);
auto smem_thr_copy_SFQ = smem_tiled_copy_SFQ.get_thread_slice(thread_idx);
Tensor tSsSFQ = smem_thr_copy_SFQ.partition_S(as_position_independent_swizzle_tensor(sSFQ));
Tensor tSrSFQ_copy_view = smem_thr_copy_SFQ.retile_D(tSrSFQ);
auto smem_tiled_copy_SFK = make_tiled_copy_impl(SmemCopyAtomSF{},
get_layoutSFB_TV(tiled_mma_qk),
make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk))
);
auto smem_thr_copy_SFK = smem_tiled_copy_SFK.get_thread_slice(thread_idx);
Tensor tSsSFK = smem_thr_copy_SFK.partition_S(as_position_independent_swizzle_tensor(sSFK));
Tensor tSrSFK_copy_view = smem_thr_copy_SFK.retile_D(tSrSFK);
auto smem_tiled_copy_SFV = make_tiled_copy_impl(SmemCopyAtomSF{},
get_layoutSFB_TV(tiled_mma_pv),
make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk))
);
auto smem_thr_copy_SFV = smem_tiled_copy_SFV.get_thread_slice(thread_idx);
Tensor tOsSFVt = smem_thr_copy_SFV.partition_S(as_position_independent_swizzle_tensor(sSFVt));
Tensor tOrSFVt_copy_view = smem_thr_copy_SFV.retile_D(tOrSFVt);
auto consumer_wait = [](auto& pipeline, auto& smem_pipe_read) {
auto barrier_token = pipeline.consumer_try_wait(smem_pipe_read);
pipeline.consumer_wait(smem_pipe_read, barrier_token);
};
int const seqlen_q = get<0>(mainloop_params.shape_Q);
int const seqlen_k = get<0>(mainloop_params.shape_K);
int const unpadded_seqlen_k = get<0>(mainloop_params.unpadded_shape_K);
int n_block = n_block_count - 1;
auto copy_k_block = [&](auto block_id) {
auto tSsK_stage = tSsK(_, _, _, smem_pipe_read_k.index());
auto tSsSFK_stage = tSsSFK(_, _, _, smem_pipe_read_k.index());
copy(smem_tiled_copy_K, tSsK_stage(_, _, block_id), tSrK_copy_view(_, _, block_id));
copy(smem_tiled_copy_SFK, tSsSFK_stage(_, _, block_id), tSrSFK_copy_view(_, _, block_id));
};
auto copy_v_block = [&](auto block_id) {
auto tOsVt_stage = tOsVt(_, _, _, smem_pipe_read_v.index());
auto tOsSFVt_stage = tOsSFVt(_, _, _, smem_pipe_read_v.index());
copy(smem_tiled_copy_V, tOsVt_stage(_, _, block_id), tOrVt_copy_view(_, _, block_id));
copy(smem_tiled_copy_SFV, tOsSFVt_stage(_, _, block_id), tOrSFVt_copy_view(_, _, block_id));
};
// auto gemm_qk = [&](auto block_id) {
// cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, block_id), tSrSFQ(_, _, block_id)), make_zip_tensor(tSrK(_, _, block_id), tSrSFK(_, _, block_id)), tSrS);
// };
// auto gemm_pv = [&](auto block_id) {
// cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, block_id), tOrSFP(_, _, block_id)), make_zip_tensor(tOrVt(_, _, block_id), tOrSFVt(_, _, block_id)), tOrO);
// };
auto add_delta_s = [&](auto& acc) {
// The MMA atom composites 4 sub-MMA m16n8k64 covering N=0-7, 8-15, 16-23, 24-31.
// Each float4 register group spans two sub-MMAs, so N positions are scattered
// (e.g., {2t, 2t+1, 8+2t, 9+2t}), not consecutive.
float const* ds_ptr = reinterpret_cast<float const*>(
&sDS(_0{}, _0{}, smem_pipe_read_k.index()));
auto acc_float4 = recast<float4>(acc);
int tid = threadIdx.x % 4;
for (int i = 0; i < 4; i++) {
int base_n = i * 32 + tid * 2;
float4 delta_s_0 = make_float4(
ds_ptr[base_n], ds_ptr[base_n + 1],
ds_ptr[base_n + 8], ds_ptr[base_n + 9]);
float4 delta_s_1 = make_float4(
ds_ptr[base_n + 16], ds_ptr[base_n + 17],
ds_ptr[base_n + 24], ds_ptr[base_n + 25]);
acc_float4(make_coord(make_coord(_0{}, _0{}), _0{}), _0{}, i) = delta_s_0;
acc_float4(make_coord(make_coord(_0{}, _0{}), _1{}), _0{}, i) = delta_s_0;
acc_float4(make_coord(make_coord(_0{}, _1{}), _0{}), _0{}, i) = delta_s_1;
acc_float4(make_coord(make_coord(_0{}, _1{}), _1{}), _0{}, i) = delta_s_1;
}
};
consumer_wait(pipeline_q, smem_pipe_read_q);
copy(smem_tiled_copy_Q, tSsQ, tSrQ_copy_view);
copy(smem_tiled_copy_SFQ, tSsSFQ, tSrSFQ_copy_view);
pipeline_q.consumer_release(smem_pipe_read_q);
++smem_pipe_read_q;
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
Tensor AbsMaxP = make_tensor_like<float>(
make_layout(shape(group<1, 4>(flatten(tSrS_converion_view.layout()(make_coord(_0{}, _), _, _)))))
);
consumer_wait(pipeline_k, smem_pipe_read_k);
copy_k_block(_0{});
add_delta_s(tSrS);
CUTLASS_PRAGMA_UNROLL
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
if (k_block < size<2>(tSrQ) - 1) {
copy_k_block(k_block + 1);
} else {
pipeline_k.consumer_release(smem_pipe_read_k);
++smem_pipe_read_k;
}
}
auto col_limit_causal = [&](int row, int n_block) {
return row + 1 + seqlen_k - n_block * kBlockN - seqlen_q + m_block * kBlockM;
};
{
Tensor cS = cute::make_identity_tensor(select<0, 1>(TileShape_MNK{}));
Tensor tScS = thread_mma_qk.partition_C(cS);
CUTLASS_PRAGMA_UNROLL
for (int i = 0; i < size(tSrS); ++i) {
if constexpr (!Is_causal) { // Just masking based on col
if (int(get<1>(tScS(i))) >= int(unpadded_seqlen_k - n_block * kBlockN)) { tSrS(i) = -INFINITY; }
} else {
if (int(get<1>(tScS(i))) >= std::min(seqlen_k - n_block * kBlockN,
col_limit_causal(int(get<0>(tScS(i))), n_block))) {
tSrS(i) = -INFINITY;
}
}
}
}
auto quantize = [&](auto mma_k, auto acc_conversion_view) {
Tensor AbsMaxP_stagek = AbsMaxP(_, make_coord(_, _, mma_k));
Tensor acc_conversion_stagek = acc_conversion_view(_, _, mma_k);
Tensor SFP = make_tensor_like<cutlass::float_ue4m3_t>(AbsMaxP_stagek.layout());
Tensor SFP_uint32_view = recast<uint32_t>(SFP);
CUTLASS_PRAGMA_UNROLL
for (int i = 0; i < size(AbsMaxP_stagek); i += 4) {
uint32_t& tmp = SFP_uint32_view(i / 4);
flash::packed_float_to_ue4m3(
AbsMaxP_stagek(i),
AbsMaxP_stagek(i + 1),
AbsMaxP_stagek(i + 2),
AbsMaxP_stagek(i + 3),
tmp
);
}
int const quad_id = threadIdx.x & 3;
uint32_t MASK = (0xFF00FF) << ((quad_id & 1) * 8);
Tensor tOrSFP_uint32_view = recast<uint32_t>(tOrSFP(_, _, mma_k));
Tensor tOrP_uint32_view = recast<uint32_t>(tOrP(_, _, mma_k));
CUTLASS_PRAGMA_UNROLL
for (int mma_m = 0; mma_m < size<1>(tOrP); ++mma_m) {
CUTLASS_PRAGMA_UNROLL
for (int i = 0; i < 4; ++i) {
flash::packed_float_to_e2m1(
acc_conversion_stagek(make_coord(_0{}, i), mma_m),
acc_conversion_stagek(make_coord(_1{}, i), mma_m),
acc_conversion_stagek(make_coord(_2{}, i), mma_m),
acc_conversion_stagek(make_coord(_3{}, i), mma_m),
acc_conversion_stagek(make_coord(_4{}, i), mma_m),
acc_conversion_stagek(make_coord(_5{}, i), mma_m),
acc_conversion_stagek(make_coord(_6{}, i), mma_m),
acc_conversion_stagek(make_coord(_7{}, i), mma_m),
tOrP_uint32_view(i, mma_m)
);
}
uint32_t local_sfp = SFP_uint32_view(_0{}, _0{}, mma_m);
uint32_t peer_sfp = __shfl_xor_sync(int32_t(-1), local_sfp, 2);
if ((quad_id & 1) == 0) {
uint32_t sfp = (local_sfp & MASK) | ((peer_sfp & MASK) << 8);
tOrSFP_uint32_view(_0{}, mma_m) = sfp;
} else {
uint32_t sfp = (peer_sfp & MASK) | ((local_sfp & MASK) >> 8);
tOrSFP_uint32_view(_0{}, mma_m) = sfp;
}
}
};
softmax_fused.template online_softmax_with_quant</*Is_first=*/true>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
consumer_wait(pipeline_v, smem_pipe_read_v);
copy_v_block(_0{});
quantize(_0{}, tSrS_converion_view);
CUTLASS_PRAGMA_UNROLL
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO_store);
if (v_block < size<2>(tOrP) - 1) {
copy_v_block(v_block + 1);
quantize(v_block + 1, tSrS_converion_view);
} else {
pipeline_v.consumer_release(smem_pipe_read_v);
++smem_pipe_read_v;
}
}
n_block--;
constexpr int n_masking_steps = !Is_causal ? 1 : cute::ceil_div(kBlockM, kBlockN) + 1;
// // Only go through these if Is_causal, since n_masking_steps = 1 when !Is_causal
CUTLASS_PRAGMA_UNROLL
for (int masking_step = 0; masking_step < n_masking_steps - 1 && n_block >= 0; ++masking_step, --n_block) {
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
consumer_wait(pipeline_k, smem_pipe_read_k);
copy_k_block(_0{});
add_delta_s(tSrS);
CUTLASS_PRAGMA_UNROLL
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
if (k_block < size<2>(tSrQ) - 1) {
copy_k_block(k_block + 1);
}
}
pipeline_k.consumer_release(smem_pipe_read_k); // release K
++smem_pipe_read_k;
Tensor cS = cute::make_identity_tensor(select<0, 1>(TileShape_MNK{}));
Tensor tScS = thread_mma_qk.partition_C(cS);
#pragma unroll
for (int i = 0; i < size(tSrS); ++i) {
if (int(get<1>(tScS(i))) >= col_limit_causal(int(get<0>(tScS(i))), n_block)) {
tSrS(i) = -INFINITY;
}
}
softmax_fused.template online_softmax_with_quant</*Is_first=*/false>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
Tensor tOrO = make_fragment_like(tOrO_store);
consumer_wait(pipeline_v, smem_pipe_read_v);
copy_v_block(_0{});
quantize(_0{}, tSrS_converion_view);
CUTLASS_PRAGMA_UNROLL
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO);
if (v_block < size<2>(tOrP) - 1) {
copy_v_block(v_block + 1);
quantize(v_block + 1, tSrS_converion_view);
}
}
pipeline_v.consumer_release(smem_pipe_read_v);
++smem_pipe_read_v;
if (masking_step > 0) { softmax_fused.rescale_o(tOrO_store, tOrO); }
}
#pragma unroll 1
for (; n_block >= 0; --n_block) {
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
consumer_wait(pipeline_k, smem_pipe_read_k);
copy_k_block(_0{});
add_delta_s(tSrS);
CUTLASS_PRAGMA_UNROLL
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
if (k_block < size<2>(tSrQ) - 1) {
copy_k_block(k_block + 1);
} else {
pipeline_k.consumer_release(smem_pipe_read_k);
++smem_pipe_read_k;
}
}
softmax_fused.template online_softmax_with_quant</*Is_first=*/false>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
Tensor tOrO = make_fragment_like(tOrO_store);
consumer_wait(pipeline_v, smem_pipe_read_v);
copy_v_block(_0{});
quantize(_0{}, tSrS_converion_view);
CUTLASS_PRAGMA_UNROLL
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO);
if (v_block < size<2>(tOrP) - 1) {
copy_v_block(v_block + 1);
quantize(v_block + 1, tSrS_converion_view);
} else {
pipeline_v.consumer_release(smem_pipe_read_v);
++smem_pipe_read_v;
}
}
softmax_fused.rescale_o(tOrO_store, tOrO);
}
softmax_fused.finalize(tOrO_store);
return;
}
};
} // namespace flash
@@ -1,119 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include "cutlass/arch/barrier.h"
#include "cutlass/pipeline/sm90_pipeline.hpp"
namespace flash {
enum class FP4NamedBarriers {
QueryEmpty = 1,
WarpSpecializedConsumer = 2,
WarpSpecializedPingPongConsumer1 = 3,
WarpSpecializedPingPongConsumer2 = 4,
ProducerEnd = 5,
ConsumerEnd = 6,
EpilogueBarrier = 7
};
template<int SequenceDepth, int SequenceLength>
struct OrderedSequenceBarrierVarGroupSizeSharedStorage {
using Barrier = cutlass::arch::ClusterBarrier;
Barrier barrier_[SequenceDepth][SequenceLength];
};
template<int SequenceDepth_, int SequenceLength_>
class OrderedSequenceBarrierVarGroupSize {
public:
static constexpr int SequenceDepth = SequenceDepth_;
static constexpr int SequenceLength = SequenceLength_;
using Barrier = cutlass::arch::ClusterBarrier;
using SharedStorage = flash::OrderedSequenceBarrierVarGroupSizeSharedStorage<SequenceDepth, SequenceLength>;
struct Params {
uint32_t group_id;
uint32_t* group_size_list;
};
private :
// In future this Params object can be replaced easily with a CG object
Params params_;
Barrier *barrier_ptr_;
cutlass::PipelineState<SequenceDepth> stage_;
static constexpr int Depth = SequenceDepth;
static constexpr int Length = SequenceLength;
public:
OrderedSequenceBarrierVarGroupSize() = delete;
OrderedSequenceBarrierVarGroupSize(const OrderedSequenceBarrierVarGroupSize&) = delete;
OrderedSequenceBarrierVarGroupSize(OrderedSequenceBarrierVarGroupSize&&) = delete;
OrderedSequenceBarrierVarGroupSize& operator=(const OrderedSequenceBarrierVarGroupSize&) = delete;
OrderedSequenceBarrierVarGroupSize& operator=(OrderedSequenceBarrierVarGroupSize&&) = delete;
~OrderedSequenceBarrierVarGroupSize() = default;
CUTLASS_DEVICE
OrderedSequenceBarrierVarGroupSize(SharedStorage& storage, Params const& params) :
params_(params),
barrier_ptr_(&storage.barrier_[0][0]),
// Group 0 - starts with an opposite phase
stage_({0, params.group_id == 0, 0}) {
int warp_idx = cutlass::canonical_warp_idx_sync();
int lane_predicate = cute::elect_one_sync();
// Barrier FULL, EMPTY init
// Init is done only by the one elected thread of the block
if (warp_idx == 0 && lane_predicate) {
for (int d = 0; d < Depth; ++d) {
for (int l = 0; l < Length; ++l) {
barrier_ptr_[d * Length + l].init(*(params.group_size_list + l));
}
}
}
cutlass::arch::fence_barrier_init();
}
// Wait on a stage to be unlocked
CUTLASS_DEVICE
void wait() {
get_barrier_for_current_stage(params_.group_id).wait(stage_.phase());
}
// Signal completion of Stage and move to the next stage
// (group_id) signals to (group_id+1)
CUTLASS_DEVICE
void arrive() {
int signalling_id = (params_.group_id + 1) % Length;
get_barrier_for_current_stage(signalling_id).arrive();
++stage_;
}
CUTLASS_DEVICE
void advance() {
++stage_;
}
private:
CUTLASS_DEVICE
Barrier& get_barrier_for_current_stage(int group_id) {
return barrier_ptr_[stage_.index() * Length + group_id];
}
};
} // flash
@@ -1,180 +0,0 @@
// Modified from the original SageAttention3 code
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <cuda.h>
#include <vector>
#ifdef OLD_GENERATOR_PATH
#include <ATen/CUDAGeneratorImpl.h>
#else
#include <ATen/cuda/CUDAGeneratorImpl.h>
#endif
#include <ATen/cuda/CUDAGraphsUtils.cuh> // For at::cuda::philox::unpack
#include "cutlass/fast_math.h" // For cutlass::FastDivmod
////////////////////////////////////////////////////////////////////////////////////////////////////
struct Qkv_params {
using index_t = int64_t;
// The QKV matrices.
void *__restrict__ q_ptr;
void *__restrict__ k_ptr;
void *__restrict__ v_ptr;
void *__restrict__ delta_s_ptr;
// The QKV scale factor matrices.
void *__restrict__ sfq_ptr;
void *__restrict__ sfk_ptr;
void *__restrict__ sfv_ptr;
// The stride between rows of the Q, K and V matrices.
index_t q_batch_stride;
index_t k_batch_stride;
index_t v_batch_stride;
index_t q_row_stride;
index_t k_row_stride;
index_t v_row_stride;
index_t q_head_stride;
index_t k_head_stride;
index_t v_head_stride;
index_t ds_batch_stride;
index_t ds_row_stride;
index_t ds_head_stride;
// The stride of the Q, K and V scale factor matrices.
index_t sfq_batch_stride;
index_t sfk_batch_stride;
index_t sfv_batch_stride;
index_t sfq_row_stride;
index_t sfk_row_stride;
index_t sfv_row_stride;
index_t sfq_head_stride;
index_t sfk_head_stride;
index_t sfv_head_stride;
// The number of heads.
int h, h_k;
// In the case of multi-query and grouped-query attention (MQA/GQA), nheads_k could be
// different from nheads (query).
int h_h_k_ratio; // precompute h / h_k,
};
////////////////////////////////////////////////////////////////////////////////////////////////////
struct Flash_fwd_params : public Qkv_params {
// The O matrix (output).
void * __restrict__ o_ptr;
void * __restrict__ oaccum_ptr;
void * __restrict__ s_ptr;
// The stride between rows of O.
index_t o_batch_stride;
index_t o_row_stride;
index_t o_head_stride;
// The pointer to the P matrix.
void * __restrict__ p_ptr;
// The pointer to the softmax sum.
void * __restrict__ softmax_lse_ptr;
void * __restrict__ softmax_lseaccum_ptr;
// The dimensions.
int b, seqlen_q, seqlen_k, seqlen_knew, d, seqlen_q_rounded, seqlen_k_rounded, d_rounded, rotary_dim, unpadded_seqlen_k;
cutlass::FastDivmod head_divmod, m_block_divmod;
int total_blocks;
int seqlen_s;
// The scaling factors for the kernel.
float scale_softmax;
float scale_softmax_log2;
uint32_t scale_softmax_log2_half2;
// array of length b+1 holding starting offset of each sequence.
int * __restrict__ cu_seqlens_q;
int * __restrict__ cu_seqlens_k;
// If provided, the actual length of each k sequence.
int * __restrict__ seqused_k;
int *__restrict__ blockmask;
// The K_new and V_new matrices.
void * __restrict__ knew_ptr;
void * __restrict__ vnew_ptr;
// The stride between rows of the Q, K and V matrices.
index_t knew_batch_stride;
index_t vnew_batch_stride;
index_t knew_row_stride;
index_t vnew_row_stride;
index_t knew_head_stride;
index_t vnew_head_stride;
// The cos and sin matrices for rotary embedding.
void * __restrict__ rotary_cos_ptr;
void * __restrict__ rotary_sin_ptr;
// The indices to index into the KV cache.
int * __restrict__ cache_batch_idx;
// Paged KV cache
int * __restrict__ block_table;
index_t block_table_batch_stride;
int page_block_size;
// The dropout probability (probability of keeping an activation).
float p_dropout;
// uint32_t p_dropout_in_uint;
// uint16_t p_dropout_in_uint16_t;
uint8_t p_dropout_in_uint8_t;
// Scale factor of 1 / (1 - p_dropout).
float rp_dropout;
float scale_softmax_rp_dropout;
// Local window size
int window_size_left, window_size_right;
// Random state.
at::PhiloxCudaState philox_args;
// Pointer to the RNG seed (idx 0) and offset (idx 1).
uint64_t * rng_state;
bool is_bf16;
bool is_e4m3;
bool is_causal;
bool per_block_mean;
bool single_level_p_quant; // If true, use single-level 1x16 block scale quantization for P (like V), instead of two-level quantization
// If is_seqlens_k_cumulative, then seqlen_k is cu_seqlens_k[bidb + 1] - cu_seqlens_k[bidb].
// Otherwise it's cu_seqlens_k[bidb], i.e., we use cu_seqlens_k to store the sequence lengths of K.
bool is_seqlens_k_cumulative;
bool is_rotary_interleaved;
int num_splits; // For split-KV version
void * __restrict__ alibi_slopes_ptr;
index_t alibi_slopes_batch_stride;
int * __restrict__ tile_count_semaphore;
};
////////////////////////////////////////////////////////////////////////////////////////////////////
@@ -1,190 +0,0 @@
// Modified from the original SageAttention3 code
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <cmath>
#include "cute/tensor.hpp"
#include "cutlass/numeric_types.h"
#include "utils.h"
namespace flash {
using namespace cute;
template <int Rows>
struct SoftmaxFused{
using TensorT = decltype(make_fragment_like<float>(Shape<Int<Rows>>{}));
TensorT row_sum, row_max, scores_scale;
static constexpr float fp8_scalexfp4_scale = 1.f / (448 * 6);
static constexpr float fp8_scalexfp4_scale_log2 = -11.392317422778762f; //log2f(fp8_scalexfp4_scale)
static constexpr float fp4_scale_log2 = -2.584962500721156f; // log2f(fp4_scale)
static constexpr int RowReductionThr = 4;
// If true, use single-level quantization: s_P2, P̂_2 = φ(P̃) directly (standard per-block FP4 quantization like V)
// If false (default), use two-level quantization: s_P1 = rowmax(P̃)/(448×6), then s_P2, P̂_2 = φ(P̃/s_P1)
bool single_level_p_quant;
CUTLASS_DEVICE SoftmaxFused(bool single_level = false) : single_level_p_quant(single_level) {};
template<bool FirstTile, bool InfCheck = false, typename TensorAcc, typename TensorMax>
CUTLASS_DEVICE auto online_softmax_with_quant(
TensorAcc& acc,
TensorMax& AbsMaxP,
const float softmax_scale_log2
) {
Tensor acc_reduction_view = make_tensor(acc.data(), flash::convert_to_reduction_layout(acc.layout()));
Tensor acc_conversion_view = make_tensor(acc.data(), flash::convert_to_conversion_layout(acc.layout()));
Tensor acc_conversion_flatten = group_modes<1, 5>(group_modes<0, 2>(flatten(acc_conversion_view)));
if constexpr (FirstTile) {
fill(row_max, -INFINITY);
clear(row_sum);
fill(scores_scale, 1.f);
CUTLASS_PRAGMA_UNROLL
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1, 1>(acc_reduction_view); ni++) {
CUTLASS_PRAGMA_UNROLL
for (int ei = 0; ei < size<1, 0>(acc_reduction_view); ei++) {
AbsMaxP(mi, ni) = fmaxf(AbsMaxP(mi, ni), acc_reduction_view(mi, make_coord(ei, ni)));
}
float max_recv = __shfl_xor_sync(int32_t(-1), AbsMaxP(mi, ni), 1); // exchange max with neighbour thread of 8 elements
AbsMaxP(mi, ni) = fmaxf(AbsMaxP(mi, ni), max_recv);
row_max(mi) = fmaxf(row_max(mi), AbsMaxP(mi, ni));
}
float max_recv = __shfl_xor_sync(int32_t(-1), row_max(mi), 2); // exchange max in a quad in a row
row_max(mi) = fmaxf(row_max(mi), max_recv);
// Two-level P quantization (default): s_P1 = rowmax(P̃)/(448×6), then s_P2,P̂_2 = φ(P̃/s_P1)
// - Pre-scales P to [0, 448×6] range before φ, output scaled by s_P1
// Single-level P quantization: s_P2, P̂_2 = φ(P̃) directly (like V quantization)
// - No s_P1, just standard per-block FP4 quantization φ
const float s_P1_offset = single_level_p_quant ? 0.f : fp8_scalexfp4_scale_log2;
const float max_scaled = InfCheck
? (row_max(mi) == -INFINITY ? 0.f : (row_max(mi) * softmax_scale_log2 + s_P1_offset))
: (row_max(mi) * softmax_scale_log2 + s_P1_offset);
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
acc_reduction_view(mi, ni) = flash::ptx_exp2(acc_reduction_view(mi, ni) * softmax_scale_log2 - max_scaled);
}
// s_P2 = max(P_block)/6 — per-block scale factor from φ function (same formula for both modes)
// The difference is in max_scaled: two-level includes 448×6 pre-scaling, single-level doesn't
CUTLASS_PRAGMA_UNROLL
for (int sfi = 0; sfi < size<1>(AbsMaxP); sfi++) {
AbsMaxP(mi, sfi) = flash::ptx_exp2(AbsMaxP(mi, sfi) * softmax_scale_log2 - max_scaled + fp4_scale_log2);
}
}
CUTLASS_PRAGMA_UNROLL
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
row_sum(mi) += acc_reduction_view(mi, ni);
}
}
}
else {
Tensor scores_max_prev = make_fragment_like(row_max);
cute::copy(row_max, scores_max_prev);
CUTLASS_PRAGMA_UNROLL
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1, 1>(acc_reduction_view); ni++) {
float local_max = -INFINITY;
CUTLASS_PRAGMA_UNROLL
for (int ei = 0; ei < size<1, 0>(acc_reduction_view); ei++) {
local_max = fmaxf(local_max, acc_reduction_view(mi, make_coord(ei, ni)));
}
float max_recv = __shfl_xor_sync(int32_t(-1), local_max, 1); // exchange max with neighbour thread of 8 elements
AbsMaxP(mi, ni) = fmaxf(local_max, max_recv);
row_max(mi) = fmaxf(row_max(mi), AbsMaxP(mi, ni));
}
float max_recv = __shfl_xor_sync(int32_t(-1), row_max(mi), 2); // exchange max in a quad in a row
row_max(mi) = fmaxf(row_max(mi), max_recv);
float scores_max_cur = !InfCheck
? row_max(mi)
: (row_max(mi) == -INFINITY ? 0.0f : row_max(mi));
scores_scale(mi) = flash::ptx_exp2((scores_max_prev(mi) - scores_max_cur) * softmax_scale_log2);
// Two-level P quantization (default): s_P1 = rowmax(P̃)/(448×6), then s_P2,P̂_2 = φ(P̃/s_P1)
// Single-level P quantization: s_P2, P̂_2 = φ(P̃) directly (like V quantization)
const float s_P1_offset = single_level_p_quant ? 0.f : fp8_scalexfp4_scale_log2;
const float max_scaled = InfCheck
? (row_max(mi) == -INFINITY ? 0.f : (row_max(mi) * softmax_scale_log2 + s_P1_offset))
: (row_max(mi) * softmax_scale_log2 + s_P1_offset);
row_sum(mi) = row_sum(mi) * scores_scale(mi);
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
acc_reduction_view(mi, ni) = flash::ptx_exp2(acc_reduction_view(mi, ni) * softmax_scale_log2 - max_scaled);
row_sum(mi) += acc_reduction_view(mi, ni);
}
// s_P2 = max(P_block)/6 — per-block scale factor from φ function
CUTLASS_PRAGMA_UNROLL
for (int sfi = 0; sfi < size<1>(AbsMaxP); sfi++) {
AbsMaxP(mi, sfi) = flash::ptx_exp2(AbsMaxP(mi, sfi) * softmax_scale_log2 - max_scaled + fp4_scale_log2);
}
// scores_scale(mi) = max_scaled;
}
}
CUTLASS_PRAGMA_UNROLL
for (int i = 0; i < size(AbsMaxP); ++i) {
CUTLASS_PRAGMA_UNROLL
for (int j = 0; j < size<0>(acc_conversion_flatten); ++j)
acc_conversion_flatten(j, i) /= AbsMaxP(i);
}
}
template<typename TensorAcc>
CUTLASS_DEVICE void finalize(TensorAcc& o_store) {
Tensor o_store_reduction_view = make_tensor(o_store.data(), flash::convert_to_reduction_layout(o_store.layout()));
CUTLASS_PRAGMA_UNROLL
for (int mi = 0; mi < size(row_max); ++mi) {
CUTLASS_PRAGMA_UNROLL
for (int i = 1; i < RowReductionThr; i <<= 1) {
float sum_recv = __shfl_xor_sync(int32_t(-1), row_sum(mi), i);
row_sum(mi) += sum_recv;
}
float sum = row_sum(mi);
float inv_sum = (sum == 0.f || sum != sum) ? 0.f : 1 / sum;
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1>(o_store_reduction_view); ++ni) {
o_store_reduction_view(mi, ni) *= inv_sum;
}
}
}
template<typename TensorAcc>
CUTLASS_DEVICE void rescale_o(TensorAcc& o_store, TensorAcc const& o_tmp) {
Tensor o_store_reduction_view = make_tensor(o_store.data(), flash::convert_to_reduction_layout(o_store.layout()));
Tensor o_tmp_reduction_view = make_tensor(o_tmp.data(), flash::convert_to_reduction_layout(o_tmp.layout()));
CUTLASS_PRAGMA_UNROLL
for (int mi = 0; mi < size(row_max); ++mi) {
CUTLASS_PRAGMA_UNROLL
for (int ni = 0; ni < size<1>(o_store_reduction_view); ++ni) {
o_store_reduction_view(mi, ni) = o_store_reduction_view(mi, ni) * scores_scale(mi) + o_tmp_reduction_view(mi, ni);
}
}
}
};
} // namespace flash
@@ -1,83 +0,0 @@
// Inspired by
// https://github.com/NVIDIA/DALI/blob/main/include/dali/core/static_switch.h
// and https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/Dispatch.h
#pragma once
/// @param COND - a boolean expression to switch by
/// @param CONST_NAME - a name given for the constexpr bool variable.
/// @param ... - code to execute for true and false
///
/// Usage:
/// ```
/// BOOL_SWITCH(flag, BoolConst, [&] {
/// some_function<BoolConst>(...);
/// });
/// ```
//
#define BOOL_SWITCH(COND, CONST_NAME, ...) \
[&] { \
if (COND) { \
constexpr static bool CONST_NAME = true; \
return __VA_ARGS__(); \
} else { \
constexpr static bool CONST_NAME = false; \
return __VA_ARGS__(); \
} \
}()
#define PREC_SWITCH(PRECTYPE, ...) \
[&] { \
if (PRECTYPE == 1) { \
using kPrecType = cutlass::half_t; \
constexpr static bool kSoftFp16 = false; \
constexpr static bool kHybrid = false; \
return __VA_ARGS__(); \
} else if (PRECTYPE == 2) { \
using kPrecType = cutlass::float_e4m3_t; \
constexpr static bool kSoftFp16 = false; \
constexpr static bool kHybrid = false; \
return __VA_ARGS__(); \
} else if (PRECTYPE == 3) { \
using kPrecType = cutlass::float_e4m3_t; \
constexpr static bool kSoftFp16 = false; \
constexpr static bool kHybrid = true; \
return __VA_ARGS__(); \
} else if (PRECTYPE == 4) { \
using kPrecType = cutlass::float_e4m3_t; \
constexpr static bool kSoftFp16 = true; \
constexpr static bool kHybrid = false; \
return __VA_ARGS__(); \
} \
}()
#define HEADDIM_SWITCH(HEADDIM, ...) \
[&] { \
if (HEADDIM == 64) { \
constexpr static int kHeadSize = 64; \
return __VA_ARGS__(); \
} else if (HEADDIM == 128) { \
constexpr static int kHeadSize = 128; \
return __VA_ARGS__(); \
} else if (HEADDIM == 256) { \
constexpr static int kHeadSize = 256; \
return __VA_ARGS__(); \
} \
}()
#define SEQLEN_SWITCH(USE_VAR_SEQ_LEN, SEQ_LEN_OUT_OF_BOUND_CHECK, ...) \
[&] { \
if (!USE_VAR_SEQ_LEN) { \
if (SEQ_LEN_OUT_OF_BOUND_CHECK) { \
using kSeqLenTraitsType = FixedSeqLenTraits<true>; \
return __VA_ARGS__(); \
} else { \
using kSeqLenTraitsType = FixedSeqLenTraits<false>; \
return __VA_ARGS__(); \
} \
} else { \
using kSeqLenTraitsType = VarSeqLenTraits; \
return __VA_ARGS__(); \
} \
}()
@@ -1,304 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* This code is based on code from FlashAttention3, https://github.com/Dao-AILab/flash-attention
* Copyright (c) 2024, Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao.
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include "cutlass/fast_math.h"
namespace flash {
///////////////////////////////////////////////////////////////////////////////
class StaticPersistentTileSchedulerOld {
//
// Data members
//
private:
int current_work_linear_idx_;
cutlass::FastDivmod const &m_block_divmod, &head_divmod;
int const total_blocks;
public:
struct WorkTileInfo {
int M_idx = 0;
int H_idx = 0;
int B_idx = 0;
bool is_valid_tile = false;
CUTLASS_HOST_DEVICE
bool
is_valid() const {
return is_valid_tile;
}
CUTLASS_HOST_DEVICE
static WorkTileInfo
invalid_work_tile() {
return {-1, -1, -1, false};
}
};
public:
CUTLASS_DEVICE explicit StaticPersistentTileSchedulerOld(cutlass::FastDivmod const &m_block_divmod_,
cutlass::FastDivmod const &head_divmod_,
int const total_blocks_) :
m_block_divmod(m_block_divmod_), head_divmod(head_divmod_), total_blocks(total_blocks_) {
// MSVC requires protecting use of CUDA-specific nonstandard syntax,
// like blockIdx and gridDim, with __CUDA_ARCH__.
#if defined(__CUDA_ARCH__)
// current_work_linear_idx_ = blockIdx.x + blockIdx.y * gridDim.x + blockIdx.z * gridDim.x * gridDim.y;
current_work_linear_idx_ = blockIdx.x;
#else
CUTLASS_ASSERT(false && "This line should never be reached");
#endif
}
CUTLASS_DEVICE
WorkTileInfo
get_current_work() const {
return get_current_work_for_linear_idx(current_work_linear_idx_);
}
CUTLASS_DEVICE
WorkTileInfo
get_current_work_for_linear_idx(int linear_idx) const {
if (linear_idx >= total_blocks) {
return WorkTileInfo::invalid_work_tile();
}
// Map worker's linear index into the CTA tiled problem shape to the corresponding MHB indices
int M_idx, H_idx, B_idx;
int quotient = m_block_divmod.divmod(M_idx, linear_idx);
B_idx = head_divmod.divmod(H_idx, quotient);
return {M_idx, H_idx, B_idx, true};
}
CUTLASS_DEVICE
void
// advance_to_next_work(int advance_count = 1) {
advance_to_next_work() {
// current_work_linear_idx_ += int(gridDim.x * gridDim.y * gridDim.z);
current_work_linear_idx_ += int(gridDim.x);
}
CUTLASS_DEVICE
WorkTileInfo
fetch_next_work() {
WorkTileInfo new_work_tile_info;
advance_to_next_work();
new_work_tile_info = get_current_work();
return new_work_tile_info;
}
};
///////////////////////////////////////////////////////////////////////////////
class SingleTileScheduler {
public:
// Host side kernel arguments
struct Arguments {
int const num_blocks_m, num_head, num_batch;
int const* tile_count_semaphore = nullptr;
};
// Device side kernel params
struct Params {};
static Params
to_underlying_arguments(Arguments const& args) {
return {};
}
static dim3
get_grid_dim(Arguments const& args, int num_sm) {
return {uint32_t(args.num_blocks_m), uint32_t(args.num_head), uint32_t(args.num_batch)};
}
struct WorkTileInfo {
int M_idx = 0;
int H_idx = 0;
int B_idx = 0;
bool is_valid_tile = false;
CUTLASS_DEVICE
bool
is_valid(Params const& params) const {
return is_valid_tile;
}
CUTLASS_DEVICE
cute::tuple<int32_t, int32_t, int32_t>
get_block_coord(Params const& params) const {
return {M_idx, H_idx, B_idx};
}
CUTLASS_DEVICE
WorkTileInfo
get_next_work(Params const& params) const {
return {-1, -1, -1, false};
}
};
CUTLASS_DEVICE
WorkTileInfo
get_initial_work() const {
return {int(blockIdx.x), int(blockIdx.y), int(blockIdx.z), true};
}
CUTLASS_DEVICE
WorkTileInfo
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
return {-1, -1, -1, false};
}
};
///////////////////////////////////////////////////////////////////////////////
class StaticPersistentTileScheduler {
public:
// Host side kernel arguments
struct Arguments {
int const num_blocks_m, num_head, num_batch;
int const* tile_count_semaphore = nullptr;
};
// Device side kernel params
struct Params {
int total_blocks;
cutlass::FastDivmod m_block_divmod, head_divmod;
};
static Params
to_underlying_arguments(Arguments const& args) {
return {args.num_blocks_m * args.num_head * args.num_batch,
cutlass::FastDivmod(args.num_blocks_m), cutlass::FastDivmod(args.num_head)};
}
static dim3
get_grid_dim(Arguments const& args, int num_sm) {
return {uint32_t(num_sm)};
}
struct WorkTileInfo {
int tile_idx;
CUTLASS_DEVICE
bool
is_valid(Params const& params) const {
return tile_idx < params.total_blocks;
}
CUTLASS_DEVICE
cute::tuple<int32_t, int32_t, int32_t>
get_block_coord(Params const& params) const {
int m_block, bidh, bidb;
bidb = params.head_divmod.divmod(bidh, params.m_block_divmod.divmod(m_block, tile_idx));
return {m_block, bidh, bidb};
}
};
CUTLASS_DEVICE
WorkTileInfo
get_initial_work() const {
return {int(blockIdx.x)};
}
CUTLASS_DEVICE
WorkTileInfo
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
return {current_work.tile_idx + int(gridDim.x)};
}
};
class DynamicPersistentTileScheduler {
public:
// Host side kernel arguments
struct Arguments {
int const num_blocks_m, num_head, num_batch;
int const* tile_count_semaphore;
};
// Device side kernel params
struct Params {
int const total_blocks;
cutlass::FastDivmod const m_block_divmod, head_divmod;
int const* tile_count_semaphore;
};
static Params
to_underlying_arguments(Arguments const& args) {
return {args.num_blocks_m * args.num_head * args.num_batch,
cutlass::FastDivmod(args.num_blocks_m), cutlass::FastDivmod(args.num_head),
args.tile_count_semaphore};
}
static dim3
get_grid_dim(Arguments const& args, int num_sm) {
return {uint32_t(num_sm)};
}
using WorkTileInfo = StaticPersistentTileScheduler::WorkTileInfo;
// struct WorkTileInfo {
// int tile_idx;
// CUTLASS_DEVICE
// bool
// is_valid(Params const& params) const {
// return tile_idx < params.total_blocks;
// }
// CUTLASS_DEVICE
// cute::tuple<int32_t, int32_t, int32_t>
// get_block_coord(Params const& params) const {
// int m_block, bidh, bidb;
// bidb = params.head_divmod.divmod(bidh, params.m_block_divmod.divmod(m_block, tile_idx));
// return {m_block, bidh, bidb};
// }
// };
CUTLASS_DEVICE
WorkTileInfo
get_initial_work() const {
return {int(blockIdx.x)};
}
CUTLASS_DEVICE
WorkTileInfo
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
return {current_work.tile_idx + int(gridDim.x)};
}
};
} // flash
@@ -1,408 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#pragma once
#include <assert.h>
#include <stdint.h>
#include <stdlib.h>
#include <cuda_fp16.h>
#if defined(__CUDA_ARCH__) && __CUDA_ARCH__ >= 800
#include <cuda_bf16.h>
#endif
#include <cute/tensor.hpp>
#include <cutlass/array.h>
#include <cutlass/cutlass.h>
#include <cutlass/numeric_conversion.h>
#include <cutlass/numeric_types.h>
namespace flash {
using namespace cute;
////////////////////////////////////////////////////////////////////////////////////////////////////
template<typename T>
struct MaxOp {
__device__ __forceinline__ T operator()(T const & x, T const & y) { return x > y ? x : y; }
};
template <>
struct MaxOp<float> {
// This is slightly faster
__device__ __forceinline__ float operator()(float const &x, float const &y) { return max(x, y); }
};
////////////////////////////////////////////////////////////////////////////////////////////////////
template<typename T>
struct SumOp {
__device__ __forceinline__ T operator()(T const & x, T const & y) { return x + y; }
};
////////////////////////////////////////////////////////////////////////////////////////////////////
template<int THREADS>
struct Allreduce {
static_assert(THREADS == 32 || THREADS == 16 || THREADS == 8 || THREADS == 4);
template<typename T, typename Operator>
static __device__ __forceinline__ T run(T x, Operator &op) {
constexpr int OFFSET = THREADS / 2;
x = op(x, __shfl_xor_sync(uint32_t(-1), x, OFFSET));
return Allreduce<OFFSET>::run(x, op);
}
};
////////////////////////////////////////////////////////////////////////////////////////////////////
template<>
struct Allreduce<2> {
template<typename T, typename Operator>
static __device__ __forceinline__ T run(T x, Operator &op) {
x = op(x, __shfl_xor_sync(uint32_t(-1), x, 1));
return x;
}
};
////////////////////////////////////////////////////////////////////////////////////////////////////
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
__device__ __forceinline__ void thread_reduce_(Tensor<Engine0, Layout0> const &tensor, Tensor<Engine1, Layout1> &summary, Operator &op) {
static_assert(Layout0::rank == 2, "Only support 2D Tensor");
static_assert(Layout1::rank == 1, "Only support 1D Tensor");
CUTE_STATIC_ASSERT_V(size<0>(summary) == size<0>(tensor));
#pragma unroll
for (int mi = 0; mi < size<0>(tensor); mi++) {
summary(mi) = zero_init ? tensor(mi, 0) : op(summary(mi), tensor(mi, 0));
#pragma unroll
for (int ni = 1; ni < size<1>(tensor); ni++) {
summary(mi) = op(summary(mi), tensor(mi, ni));
}
}
}
template<typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
__device__ __forceinline__ void quad_allreduce_(Tensor<Engine0, Layout0> &dst, Tensor<Engine1, Layout1> &src, Operator &op) {
CUTE_STATIC_ASSERT_V(size(dst) == size(src));
#pragma unroll
for (int i = 0; i < size(dst); i++){
dst(i) = Allreduce<4>::run(src(i), op);
}
}
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
__device__ __forceinline__ void reduce_(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &summary, Operator &op) {
thread_reduce_<zero_init>(tensor, summary, op);
quad_allreduce_(summary, summary, op);
}
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
__device__ __forceinline__ void reduce_max(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &max){
MaxOp<float> max_op;
reduce_<zero_init>(tensor, max, max_op);
}
template<bool zero_init=true, bool warp_reduce=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
__device__ __forceinline__ void reduce_sum(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &sum){
SumOp<float> sum_op;
thread_reduce_<zero_init>(tensor, sum, sum_op);
if constexpr (warp_reduce) { quad_allreduce_(sum, sum, sum_op); }
}
__forceinline__ __device__ __half2 half_exp(__half2 x) {
uint32_t tmp_out, tmp_in;
tmp_in = reinterpret_cast<uint32_t&>(x);
asm ("ex2.approx.f16x2 %0, %1;\n"
: "=r"(tmp_out)
: "r"(tmp_in));
__half2 out = reinterpret_cast<__half2&>(tmp_out);
return out;
}
// Apply the exp to all the elements.
template <bool zero_init=false, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
__forceinline__ __device__ void max_scale_exp2_sum(Tensor<Engine0, Layout0> &tensor, Tensor<Engine1, Layout1> &max, Tensor<Engine1, Layout1> &sum, const float scale) {
static_assert(Layout0::rank == 2, "Only support 2D Tensor"); static_assert(Layout1::rank == 1, "Only support 1D Tensor"); CUTE_STATIC_ASSERT_V(size<0>(max) == size<0>(tensor));
#pragma unroll
for (int mi = 0; mi < size<0>(tensor); ++mi) {
MaxOp<float> max_op;
max(mi) = zero_init ? tensor(mi, 0) : max_op(max(mi), tensor(mi, 0));
#pragma unroll
for (int ni = 1; ni < size<1>(tensor); ni++) {
max(mi) = max_op(max(mi), tensor(mi, ni));
}
max(mi) = Allreduce<4>::run(max(mi), max_op);
// If max is -inf, then all elements must have been -inf (possibly due to masking).
// We don't want (-inf - (-inf)) since that would give NaN.
const float max_scaled = max(mi) == -INFINITY ? 0.f : max(mi) * scale;
sum(mi) = 0;
#pragma unroll
for (int ni = 0; ni < size<1>(tensor); ++ni) {
// Instead of computing exp(x - max), we compute exp2(x * log_2(e) -
// max * log_2(e)) This allows the compiler to use the ffma
// instruction instead of fadd and fmul separately.
tensor(mi, ni) = exp2f(tensor(mi, ni) * scale - max_scaled);
sum(mi) += tensor(mi, ni);
}
}
}
// Apply the exp to all the elements.
template <bool Scale_max=true, bool Check_inf=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
__forceinline__ __device__ void scale_apply_exp2(Tensor<Engine0, Layout0> &tensor, Tensor<Engine1, Layout1> const &max, const float scale) {
static_assert(Layout0::rank == 2, "Only support 2D Tensor");
static_assert(Layout1::rank == 1, "Only support 1D Tensor");
CUTE_STATIC_ASSERT_V(size<0>(max) == size<0>(tensor));
#pragma unroll
for (int mi = 0; mi < size<0>(tensor); ++mi) {
// If max is -inf, then all elements must have been -inf (possibly due to masking).
// We don't want (-inf - (-inf)) since that would give NaN.
// If we don't have float around M_LOG2E the multiplication is done in fp64.
const float max_scaled = Check_inf
? (max(mi) == -INFINITY ? 0.f : (max(mi) * (Scale_max ? scale : float(M_LOG2E))))
: (max(mi) * (Scale_max ? scale : float(M_LOG2E)));
#pragma unroll
for (int ni = 0; ni < size<1>(tensor); ++ni) {
// Instead of computing exp(x - max), we compute exp2(x * log_2(e) -
// max * log_2(e)) This allows the compiler to use the ffma
// instruction instead of fadd and fmul separately.
tensor(mi, ni) = exp2f(tensor(mi, ni) * scale - max_scaled);
}
}
}
////////////////////////////////////////////////////////////////////////////////////////////////////
__forceinline__ __device__ float ptx_exp2(float x) {
float y;
asm volatile("ex2.approx.ftz.f32 %0, %1;" : "=f"(y) : "f"(x));
return y;
}
CUTLASS_DEVICE void
packed_float_to_ue4m3(
float const &f0, float const &f1, float const &f2, float const &f3,
uint32_t &out
) {
asm volatile( \
"{\n" \
".reg .b16 lo;\n" \
".reg .b16 hi;\n" \
"cvt.rn.satfinite.e4m3x2.f32 lo, %2, %1;\n" \
"cvt.rn.satfinite.e4m3x2.f32 hi, %4, %3;\n" \
"mov.b32 %0, {lo, hi};\n" \
"}" \
: "=r"(out) : "f"(f0), "f"(f1), "f"(f2), "f"(f3));
}
CUTLASS_DEVICE void
packed_float_to_e2m1(
float const &f0, float const &f1, float const &f2, float const& f3,
float const &f4, float const &f5, float const &f6, float const& f7,
uint32_t &out
) {
asm volatile( \
"{\n" \
".reg .b8 byte0;\n" \
".reg .b8 byte1;\n" \
".reg .b8 byte2;\n" \
".reg .b8 byte3;\n" \
"cvt.rn.satfinite.e2m1x2.f32 byte0, %2, %1;\n" \
"cvt.rn.satfinite.e2m1x2.f32 byte1, %4, %3;\n" \
"cvt.rn.satfinite.e2m1x2.f32 byte2, %6, %5;\n" \
"cvt.rn.satfinite.e2m1x2.f32 byte3, %8, %7;\n" \
"mov.b32 %0, {byte0, byte1, byte2, byte3};\n" \
"}" \
: "=r"(out) : "f"(f0), "f"(f1), "f"(f2), "f"(f3),
"f"(f4), "f"(f5), "f"(f6), "f"(f7));
}
CUTLASS_DEVICE void
add(float2 & c,
float2 const& a,
float2 const& b)
{
asm volatile("add.f32x2 %0, %1, %2;\n"
: "=l"(reinterpret_cast<uint64_t &>(c))
: "l"(reinterpret_cast<uint64_t const&>(a)),
"l"(reinterpret_cast<uint64_t const&>(b)));
}
CUTLASS_DEVICE void
add_inplace(float2 &a,
float2 const& b)
{
asm volatile("add.f32x2 %0, %0, %1;\n"
: "+l"(reinterpret_cast<uint64_t &>(a)) // a: input/output
: "l"(reinterpret_cast<uint64_t const&>(b)) // b: input
);
}
CUTLASS_DEVICE void
sub(float2 & c,
float2 const& a,
float2 const& b)
{
asm volatile("sub.f32x2 %0, %1, %2;\n"
: "=l"(reinterpret_cast<uint64_t &>(c))
: "l"(reinterpret_cast<uint64_t const&>(a)),
"l"(reinterpret_cast<uint64_t const&>(b)));
}
CUTLASS_DEVICE void
sub_inplace(float2 &a,
float2 const& b)
{
asm volatile("sub.f32x2 %0, %0, %1;\n"
: "+l"(reinterpret_cast<uint64_t &>(a)) // a: input/output
: "l"(reinterpret_cast<uint64_t const&>(b)) // b: input
);
}
CUTLASS_DEVICE void
mul(float2 & c,
float2 const& a,
float2 const& b)
{
asm volatile("mul.f32x2 %0, %1, %2;\n"
: "=l"(reinterpret_cast<uint64_t &>(c))
: "l"(reinterpret_cast<uint64_t const&>(a)),
"l"(reinterpret_cast<uint64_t const&>(b)));
}
CUTLASS_DEVICE void
fma(float2 & d,
float2 const& a,
float2 const& b,
float2 const& c)
{
asm volatile("fma.rn.f32x2 %0, %1, %2, %3;\n"
: "=l"(reinterpret_cast<uint64_t &>(d))
: "l"(reinterpret_cast<uint64_t const&>(a)),
"l"(reinterpret_cast<uint64_t const&>(b)),
"l"(reinterpret_cast<uint64_t const&>(c)));
}
CUTLASS_DEVICE void
fma_inplace(float2 &a,
float2 const& b,
float2 const& c)
{
asm volatile("fma.rn.f32x2 %0, %0, %1, %2;\n"
: "+l"(reinterpret_cast<uint64_t &>(a))
: "l"(reinterpret_cast<uint64_t const&>(b)),
"l"(reinterpret_cast<uint64_t const&>(c)));
}
////////////////////////////////////////////////////////////////////////////////////////////////////
template <
class Layout
>
CUTLASS_DEVICE constexpr
auto convert_to_reduction_layout(Layout mma_layout) {
static_assert(rank(mma_layout) == 3, "Mma Layout should be (MmaAtom, MmaM, MmaN)");
static_assert(rank(get<0>(shape(mma_layout))) == 2, "MmaAtom should be (AtomN, AtomM)");
return make_layout(
make_layout(get<0,1>(mma_layout), get<1>(mma_layout)),
make_layout(get<0,0>(mma_layout), get<2>(mma_layout))
);
}
template <
class Tensor
>
CUTLASS_DEVICE constexpr
auto convert_to_reduction_tensor(Tensor mma_tensor) {
return make_tensor(mma_tensor.data(), convert_to_reduction_layout(mma_tensor.layout()));
}
template <
class Layout
>
CUTLASS_DEVICE constexpr
auto convert_to_conversion_layout(Layout mma_layout) {
static_assert(rank(mma_layout) == 3, "Mma Layout should be (MmaAtom, MmaM, MmaN)");
static_assert(rank(get<0>(shape(mma_layout))) == 2, "MmaAtom should be (AtomN, AtomM)");
constexpr int MmaAtomN = size<0, 0>(mma_layout);
constexpr int MmaAtomM = size<0, 1>(mma_layout);
constexpr int MmaM = size<1>(mma_layout);
constexpr int MmaN = size<2>(mma_layout);
static_assert(MmaAtomN == 8, "MmaAtomN should be 8.");
static_assert(MmaAtomM == 2, "MmaAtomM should be 2.");
static_assert(MmaN % 2 == 0, "MmaN should be multiple of 2.");
auto mma_n_division = zipped_divide(
layout<2>(mma_layout), make_tile(_2{})
);
return make_layout(
make_layout(layout<0,0>(mma_layout), make_layout(layout<0,1>(mma_layout), layout<0>(mma_n_division))),
layout<1>(mma_layout), layout<1>(mma_n_division)
);
}
template <
class Tensor
>
CUTLASS_DEVICE constexpr
auto convert_to_conversion_tensor(Tensor mma_tensor) {
return make_tensor(mma_tensor.data(), convert_to_conversion_layout(mma_tensor.layout()));
}
////////////////////////////////////////////////////////////////////////////////////////////////////
template <bool Is_even_MN=true, bool Is_even_K=true, bool Clear_OOB_MN=false, bool Clear_OOB_K=true,
typename TiledCopy, typename Engine0, typename Layout0, typename Engine1, typename Layout1,
typename Engine2, typename Layout2, typename Engine3, typename Layout3>
CUTLASS_DEVICE void copy(TiledCopy tiled_copy, Tensor<Engine0, Layout0> const &S,
Tensor<Engine1, Layout1> &D, Tensor<Engine2, Layout2> const &identity_MN,
Tensor<Engine3, Layout3> const &predicate_K, const int max_MN=0) {
CUTE_STATIC_ASSERT_V(rank(S) == Int<3>{});
CUTE_STATIC_ASSERT_V(rank(D) == Int<3>{});
CUTE_STATIC_ASSERT_V(size<0>(S) == size<0>(D)); // MMA
CUTE_STATIC_ASSERT_V(size<1>(S) == size<1>(D)); // MMA_M
CUTE_STATIC_ASSERT_V(size<2>(S) == size<2>(D)); // MMA_K
// There's no case where !Clear_OOB_K && Clear_OOB_MN
static_assert(!(Clear_OOB_MN && !Clear_OOB_K));
#pragma unroll
for (int m = 0; m < size<1>(S); ++m) {
if (Is_even_MN || get<0>(identity_MN(0, m, 0)) < max_MN) {
#pragma unroll
for (int k = 0; k < size<2>(S); ++k) {
if (Is_even_K || predicate_K(k)) {
cute::copy(tiled_copy, S(_, m, k), D(_, m, k));
} else if (Clear_OOB_K) {
cute::clear(D(_, m, k));
}
}
} else if (Clear_OOB_MN) {
cute::clear(D(_, m, _));
}
}
}
} // namespace flash
@@ -1 +0,0 @@
__version__ = "3.0.0.b1"
@@ -1,90 +0,0 @@
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
import torch
import fp4quant
from triton.tools.mxfp import MXFP4Tensor
from bench_utils import bench_kineto
b = 1
h = 32
n = 16384
d = 128
def test():
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
o = torch.empty((b, h, n, d // 2), device="cuda", dtype=torch.uint8)
o_s = torch.empty((b, h, n, d // 16), device="cuda", dtype=torch.float8_e4m3fn)
fp4quant.scaled_fp4_quant_permute(q, o, o_s, 1)
test()
t = bench_kineto(test, "scaled_fp4_quant_kernel", suppress_kineto_output=True)
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
throughput = IO / t * 1e-9
print(f"Throughput: {throughput:.2f} GB/s")
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
B, H, M, N = x.shape
x = x.view(B, H, M, N // 16, 16)
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
if all_ones:
scales = torch.ones_like(scales)
x_scaled = x / scales
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
permuted_fp8_scale = None
if permuted:
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
b = 2
h = 4
n = 251
n_padded = (n + 127) // 128 * 128
d = 128
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
o = torch.empty((b, h, n, d // 2), dtype=torch.uint8, device='cuda')
o_s = torch.empty((b, h, n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
k_permute = [0, 1, 8, 9, 16, 17, 24, 25, 2, 3, 10, 11, 18, 19, 26, 27, 4, 5, 12, 13, 20, 21, 28, 29, 6, 7, 14, 15, 22, 23, 30, 31]
o_permuted = torch.empty((b, h, n_padded, d // 2), dtype=torch.uint8, device='cuda')
o_s_permuted = torch.empty((b, h, n_padded, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
fp4quant.scaled_fp4_quant_permute(q, o_permuted, o_s_permuted, 1)
# padding
if n % 128 != 0:
o_permuted_gt = torch.cat([o, torch.zeros((b, h, n_padded - n, d // 2), dtype=torch.uint8, device='cuda')], dim=2)
o_s_permuted_gt = torch.cat([o_s, torch.zeros((b, h, n_padded - n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')], dim=2)
else:
o_permuted_gt = o
o_s_permuted_gt = o_s
# use scale_and_fp4_tensor + torch permutation to get the ground truth
o_permuted_gt = o_permuted_gt.reshape(b, h, n_padded // 32, 32, d // 2)[:, :, :, k_permute, :].reshape(b, h, n_padded, d // 2)
o_s_permuted_gt = o_s_permuted_gt.reshape(b, h, n_padded // 32, 32, d // 16)[:, :, :, k_permute, :].reshape(b, h, n_padded, d // 16)
assert((o_permuted - o_permuted_gt).abs().max() == 0)
assert((o_s_permuted.float() - o_s_permuted_gt.float()).abs().max() == 0)
print("All tests passed!")
@@ -1,86 +0,0 @@
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
import torch
import fp4quant
from triton.tools.mxfp import MXFP4Tensor
from bench_utils import bench_kineto
b = 1
h = 32
n = 16384
d = 128
def test():
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
o = torch.empty((b, h, n, d // 2), device="cuda", dtype=torch.uint8)
o_s = torch.empty((b, h, n, d // 16), device="cuda", dtype=torch.float8_e4m3fn)
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
test()
t = bench_kineto(test, "scaled_fp4_quant_kernel", suppress_kineto_output=True)
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
throughput = IO / t * 1e-9
print(f"Throughput: {throughput:.2f} GB/s")
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
B, H, M, N = x.shape
x = x.view(B, H, M, N // 16, 16)
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
if all_ones:
scales = torch.ones_like(scales)
x_scaled = x / scales
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
permuted_fp8_scale = None
if permuted:
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
b = 2
h = 4
n = 251
d = 128
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
o = torch.empty((b, h, n, d // 2), dtype=torch.uint8, device='cuda')
o_s = torch.empty((b, h, n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale = scale_and_fp4_tensor(q, packed_dim=3)
assert((fp8_scale.float() - o_s.float()).abs().max() == 0)
o_binary = [
(int(bin_str[:4], 2), int(bin_str[4:], 2))
for bin_str in [format(x.item(), '08b') for x in o.view(-1)]
]
o_binary_gt = [
(int(bin_str[:4], 2), int(bin_str[4:], 2))
for bin_str in [format(x.item(), '08b') for x in packed_fp4.view(-1)]
]
for i in range(len(o_binary)):
# check contiguous 4 bits. Difference should be at most one
assert(abs(o_binary[i][0] - o_binary_gt[i][0]) <= 1)
assert(abs(o_binary[i][1] - o_binary_gt[i][1]) <= 1)
print("All tests passed!")
@@ -1,86 +0,0 @@
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
import torch
import fp4quant
from triton.tools.mxfp import MXFP4Tensor
from bench_utils import bench_kineto
b = 1
h = 32
n = 16384
d = 128
def test():
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
o = torch.empty((b, h, d, n // 2), device="cuda", dtype=torch.uint8)
o_s = torch.empty((b, h, d, n // 16), device="cuda", dtype=torch.float8_e4m3fn)
fp4quant.scaled_fp4_quant_trans(q, o, o_s, 1)
test()
t = bench_kineto(test, "scaled_fp4_quant_trans_kernel", suppress_kineto_output=True)
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
throughput = IO / t * 1e-9
print(f"Throughput: {throughput:.2f} GB/s")
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
B, H, M, N = x.shape
x = x.view(B, H, M, N // 16, 16)
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
if all_ones:
scales = torch.ones_like(scales)
x_scaled = x / scales
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
permuted_fp8_scale = None
if permuted:
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
b = 2
h = 4
n = 491
n_padded = (n + 127) // 128 * 128
d = 128
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
o = torch.empty((b, h, d, n_padded // 2), dtype=torch.uint8, device='cuda')
o_s = torch.empty((b, h, d, n_padded // 16), dtype=torch.float8_e4m3fn, device='cuda')
fp4quant.scaled_fp4_quant_trans(q, o, o_s, 1)
if n % 128 != 0:
q_padded = torch.cat([q, torch.zeros((b, h, n_padded - n, d), dtype=torch.float16, device='cuda')], dim=2)
else:
q_padded = q
# use torch transpose + scaled_fp4_quant to get the ground truth
q_padded = q_padded.transpose(2, 3).reshape(b, h, n_padded, d).contiguous()
o_gt = torch.empty((b, h, n_padded, d // 2), dtype=torch.uint8, device='cuda')
o_s_gt = torch.empty((b, h, n_padded, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
fp4quant.scaled_fp4_quant(q_padded, o_gt, o_s_gt, 1)
o_gt = o_gt.reshape(b, h, d, n_padded // 2).contiguous()
o_s_gt = o_s_gt.reshape(b, h, d, n_padded // 16).contiguous()
assert((o_s_gt.float() - o_s.float()).abs().max() == 0)
assert((o_gt - o).abs().max() == 0)
print("All tests passed!")
@@ -1,169 +0,0 @@
"""
Copyright (c) 2025 by SageAttention team.
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
"""
import os
import sys
import torch
import torch.distributed as dist
def bench(fn, num_warmups: int = 5, num_tests: int = 10,
high_precision: bool = False):
# Flush L2 cache with 256 MB data
torch.cuda.synchronize()
cache = torch.empty(int(256e6 // 4), dtype=torch.int, device='cuda')
cache.zero_()
# Warmup
for _ in range(num_warmups):
fn()
# Add a large kernel to eliminate the CPU launch overhead
if high_precision:
x = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
y = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
x @ y
# Testing
start_event = torch.cuda.Event(enable_timing=True)
end_event = torch.cuda.Event(enable_timing=True)
start_event.record()
for i in range(num_tests):
fn()
end_event.record()
torch.cuda.synchronize()
return start_event.elapsed_time(end_event) / num_tests
class empty_suppress:
def __enter__(self):
return self
def __exit__(self, *_):
pass
class suppress_stdout_stderr:
def __enter__(self):
self.outnull_file = open(os.devnull, 'w')
self.errnull_file = open(os.devnull, 'w')
self.old_stdout_fileno_undup = sys.stdout.fileno()
self.old_stderr_fileno_undup = sys.stderr.fileno()
self.old_stdout_fileno = os.dup(sys.stdout.fileno())
self.old_stderr_fileno = os.dup(sys.stderr.fileno())
self.old_stdout = sys.stdout
self.old_stderr = sys.stderr
os.dup2(self.outnull_file.fileno(), self.old_stdout_fileno_undup)
os.dup2(self.errnull_file.fileno(), self.old_stderr_fileno_undup)
sys.stdout = self.outnull_file
sys.stderr = self.errnull_file
return self
def __exit__(self, *_):
sys.stdout = self.old_stdout
sys.stderr = self.old_stderr
os.dup2(self.old_stdout_fileno, self.old_stdout_fileno_undup)
os.dup2(self.old_stderr_fileno, self.old_stderr_fileno_undup)
os.close(self.old_stdout_fileno)
os.close(self.old_stderr_fileno)
self.outnull_file.close()
self.errnull_file.close()
def bench_kineto(fn, kernel_names, num_tests: int = 30, suppress_kineto_output: bool = False,
trace_path: str = None, barrier_comm_profiling: bool = False, flush_l2: bool = False):
# Conflict with Nsight Systems
using_nsys = os.environ.get('DG_NSYS_PROFILING', False)
# For some auto-tuning kernels with prints
fn()
# Profile
suppress = suppress_stdout_stderr if suppress_kineto_output and not using_nsys else empty_suppress
with suppress():
schedule = torch.profiler.schedule(wait=0, warmup=1, active=1, repeat=1) if not using_nsys else None
profiler = torch.profiler.profile(activities=[torch.profiler.ProfilerActivity.CUDA], schedule=schedule) if not using_nsys else empty_suppress()
with profiler:
for i in range(2):
# NOTES: use a large kernel and a barrier to eliminate the unbalanced CPU launch overhead
if barrier_comm_profiling:
lhs = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
rhs = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
lhs @ rhs
dist.all_reduce(torch.ones(1, dtype=torch.float, device='cuda'))
for _ in range(num_tests):
if flush_l2:
torch.empty(int(256e6 // 4), dtype=torch.int, device='cuda').zero_()
fn()
if not using_nsys:
profiler.step()
# Return 1 if using Nsight Systems
if using_nsys:
return 1
# Parse the profiling table
assert isinstance(kernel_names, str) or isinstance(kernel_names, tuple)
is_tupled = isinstance(kernel_names, tuple)
prof_lines = profiler.key_averages().table(sort_by='cuda_time_total', max_name_column_width=100).split('\n')
kernel_names = (kernel_names, ) if isinstance(kernel_names, str) else kernel_names
assert all([isinstance(name, str) for name in kernel_names])
for name in kernel_names:
assert sum([name in line for line in prof_lines]) == 1, f'Errors of the kernel {name} in the profiling table'
# Save chrome traces
if trace_path is not None:
profiler.export_chrome_trace(trace_path)
# Return average kernel times
units = {'ms': 1e3, 'us': 1e6}
kernel_times = []
for name in kernel_names:
for line in prof_lines:
if name in line:
time_str = line.split()[-2]
for unit, scale in units.items():
if unit in time_str:
kernel_times.append(float(time_str.replace(unit, '')) / scale)
break
break
return tuple(kernel_times) if is_tupled else kernel_times[0]
def calc_diff(x, y):
x, y = x.double(), y.double()
denominator = (x * x + y * y).sum()
sim = 2 * (x * y).sum() / denominator
return 1 - sim
def count_bytes(tensors):
total = 0
for t in tensors:
if isinstance(t, tuple):
total += count_bytes(t)
else:
total += t.numel() * t.element_size()
return total
@@ -1,52 +0,0 @@
#pragma once
#include <stdio.h>
#if defined(__HIPCC__)
#define HOST_DEVICE_INLINE __host__ __device__
#define DEVICE_INLINE __device__
#define HOST_INLINE __host__
#elif defined(__CUDACC__) || defined(_NVHPC_CUDA)
#define HOST_DEVICE_INLINE __host__ __device__ __forceinline__
#define DEVICE_INLINE __device__ __forceinline__
#define HOST_INLINE __host__ __forceinline__
#else
#define HOST_DEVICE_INLINE inline
#define DEVICE_INLINE inline
#define HOST_INLINE inline
#endif
#define CUDA_CHECK(cmd) \
do { \
cudaError_t e = cmd; \
if (e != cudaSuccess) { \
printf("Failed: Cuda error %s:%d '%s'\n", __FILE__, __LINE__, \
cudaGetErrorString(e)); \
exit(EXIT_FAILURE); \
} \
} while (0)
int64_t get_device_attribute(int64_t attribute, int64_t device_id) {
static int value = [=]() {
int device = static_cast<int>(device_id);
if (device < 0) {
CUDA_CHECK(cudaGetDevice(&device));
}
int value;
CUDA_CHECK(cudaDeviceGetAttribute(
&value, static_cast<cudaDeviceAttr>(attribute), device));
return static_cast<int>(value);
}();
return value;
}
namespace cuda_utils {
template <typename T>
HOST_DEVICE_INLINE constexpr std::enable_if_t<std::is_integral_v<T>, T>
ceil_div(T a, T b) {
return (a + b - 1) / b;
}
}; // namespace cuda_utils
@@ -1,644 +0,0 @@
/*
* Copyright (c) 2025 by SageAttention team.
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include <torch/all.h>
#include <torch/python.h>
#include <torch/nn/functional.h>
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>
#include <cuda_runtime_api.h>
#include <cuda_runtime.h>
#include <ATen/cuda/CUDAContext.h>
#include <c10/cuda/CUDAGuard.h>
#include <cuda_fp8.h>
#include "cuda_utils.h"
#include "../blackwell/block_config.h"
#define DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(pytorch_dtype, c_type, ...) \
if (pytorch_dtype == at::ScalarType::Half) { \
using c_type = half; \
__VA_ARGS__ \
} else if (pytorch_dtype == at::ScalarType::BFloat16) { \
using c_type = nv_bfloat16; \
__VA_ARGS__ \
} else { \
std::ostringstream oss; \
oss << __PRETTY_FUNCTION__ << " failed to dispatch data type " << pytorch_dtype; \
TORCH_CHECK(false, oss.str()); \
}
#define DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, ...) \
if (head_dim == 64) { \
constexpr int HEAD_DIM = 64; \
__VA_ARGS__ \
} else if (head_dim == 128) { \
constexpr int HEAD_DIM = 128; \
__VA_ARGS__ \
} else { \
std::ostringstream err_msg; \
err_msg << "Unsupported head dim: " << int(head_dim); \
throw std::invalid_argument(err_msg.str()); \
}
#define CHECK_CUDA(x) \
TORCH_CHECK(x.is_cuda(), "Tensor " #x " must be on CUDA")
#define CHECK_DTYPE(x, true_dtype) \
TORCH_CHECK(x.dtype() == true_dtype, \
"Tensor " #x " must have dtype (" #true_dtype ")")
#define CHECK_DIMS(x, true_dim) \
TORCH_CHECK(x.dim() == true_dim, \
"Tensor " #x " must have dimension number (" #true_dim ")")
#define CHECK_SHAPE(x, ...) \
TORCH_CHECK(x.sizes() == torch::IntArrayRef({__VA_ARGS__}), \
"Tensor " #x " must have shape (" #__VA_ARGS__ ")")
#define CHECK_CONTIGUOUS(x) \
TORCH_CHECK(x.is_contiguous(), "Tensor " #x " must be contiguous")
#define CHECK_LASTDIM_CONTIGUOUS(x) \
TORCH_CHECK(x.stride(-1) == 1, \
"Tensor " #x " must be contiguous at the last dimension")
constexpr int CVT_FP4_ELTS_PER_THREAD = 16;
// Convert 4 float2 values into 8 e2m1 values (represented as one uint32_t).
inline __device__ uint32_t fp32_vec_to_e2m1(float2 *array) {
#if defined(__CUDA_ARCH__) && (__CUDA_ARCH__ >= 1000)
uint32_t val;
asm volatile(
"{\n"
".reg .b8 byte0;\n"
".reg .b8 byte1;\n"
".reg .b8 byte2;\n"
".reg .b8 byte3;\n"
"cvt.rn.satfinite.e2m1x2.f32 byte0, %2, %1;\n"
"cvt.rn.satfinite.e2m1x2.f32 byte1, %4, %3;\n"
"cvt.rn.satfinite.e2m1x2.f32 byte2, %6, %5;\n"
"cvt.rn.satfinite.e2m1x2.f32 byte3, %8, %7;\n"
"mov.b32 %0, {byte0, byte1, byte2, byte3};\n"
"}"
: "=r"(val)
: "f"(array[0].x), "f"(array[0].y), "f"(array[1].x), "f"(array[1].y),
"f"(array[2].x), "f"(array[2].y), "f"(array[3].x), "f"(array[3].y));
return val;
#else
return 0;
#endif
}
// Get type2 from type or vice versa (applied to half and bfloat16)
template <typename T>
struct TypeConverter {
using Type = half2;
}; // keep for generality
template <>
struct TypeConverter<half2> {
using Type = half;
};
template <>
struct TypeConverter<half> {
using Type = half2;
};
template <>
struct TypeConverter<__nv_bfloat162> {
using Type = __nv_bfloat16;
};
template <>
struct TypeConverter<__nv_bfloat16> {
using Type = __nv_bfloat162;
};
// Define a 32 bytes packed data type.
template <class Type>
struct PackedVec {
typename TypeConverter<Type>::Type elts[8];
};
template <uint32_t head_dim, uint32_t BLOCK_SIZE, bool permute, typename T>
__global__ void scaled_fp4_quant_kernel(
const T* input, uint8_t* output, uint8_t* output_sf,
int batch_size, int num_heads, int num_tokens,
int stride_bz_input, int stride_h_input, int stride_seq_input,
int stride_bz_output, int stride_h_output, int stride_seq_output,
int stride_bz_output_sf, int stride_h_output_sf, int stride_seq_output_sf) {
static_assert(std::is_same<T, half>::value || std::is_same<T, nv_bfloat16>::value, "Only half and bfloat16 input are supported");
using PackedVec = PackedVec<T>;
const int batch_id = blockIdx.y;
const int head_id = blockIdx.z;
const int token_block_id = blockIdx.x;
static_assert(CVT_FP4_ELTS_PER_THREAD == 8 || CVT_FP4_ELTS_PER_THREAD == 16,
"CVT_FP4_ELTS_PER_THREAD must be 8 or 16");
static_assert(sizeof(PackedVec) == sizeof(T) * CVT_FP4_ELTS_PER_THREAD,
"Vec size is not matched.");
constexpr uint32_t NUM_THREADS_PER_TOKEN = head_dim / CVT_FP4_ELTS_PER_THREAD;
// load input
const int token_id = token_block_id * BLOCK_SIZE + threadIdx.x / NUM_THREADS_PER_TOKEN;
int load_token_id;
if constexpr (!permute) {
load_token_id = token_id;
} else {
int local_token_id = threadIdx.x / NUM_THREADS_PER_TOKEN;
int local_token_id_residue = local_token_id % 32;
// [0, 1, 8, 9, 16, 17, 24, 25, 2, 3, 10, 11, 18, 19, 26, 27, 4, 5, 12, 13, 20, 21, 28, 29, 6, 7, 14, 15, 22, 23, 30, 31]
load_token_id = token_block_id * BLOCK_SIZE + (local_token_id / 32) * 32 +
(local_token_id_residue / 8) * 2 +
((local_token_id_residue % 8) / 2) * 8 +
(local_token_id_residue % 8) % 2;
}
PackedVec in_vec;
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
reinterpret_cast<uint32_t&>(in_vec.elts[i]) = 0;
}
if (load_token_id < num_tokens) {
in_vec = reinterpret_cast<PackedVec const*>(input +
batch_id * stride_bz_input + // batch dim
head_id * stride_h_input + // head dim
load_token_id * stride_seq_input + // seq dim
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD)[0]; // feature dim
}
// calculate max of every consecutive 16 elements
auto localMax = __habs2(in_vec.elts[0]);
#pragma unroll
for (int i = 1; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) { // local max
localMax = __hmax2(localMax, __habs2(in_vec.elts[i]));
}
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) { // shuffle across two threads
localMax = __hmax2(__shfl_xor_sync(0xffffffff, localMax, 1, 32), localMax);
}
float vecMax = float(__hmax(localMax.x, localMax.y));
// scaling factor
float SFValue = vecMax / 6.0f;
uint8_t SFValueFP8;
reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8) = __nv_fp8_e4m3(SFValue);
SFValue = float(reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8));
float SFValueInv = (SFValue == 0.0f) ? 0.0f : 1.0f / SFValue;
// convert input to float2 and apply scale
float2 fp2Vals[CVT_FP4_ELTS_PER_THREAD / 2];
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
if constexpr (std::is_same<T, half>::value) {
fp2Vals[i] = __half22float2(in_vec.elts[i]);
} else {
fp2Vals[i] = __bfloat1622float2(in_vec.elts[i]);
}
fp2Vals[i].x = fp2Vals[i].x * SFValueInv;
fp2Vals[i].y = fp2Vals[i].y * SFValueInv;
}
// convert to e2m1
uint32_t e2m1Vals[CVT_FP4_ELTS_PER_THREAD / 8];
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 8; i++) {
e2m1Vals[i] = fp32_vec_to_e2m1(fp2Vals + i * 4);
}
// Skip out-of-range tokens: never write past the sequence when num_tokens
// is not a multiple of BLOCK_SIZE (out-of-bounds write fix).
if (token_id >= num_tokens) return;
// save
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) {
reinterpret_cast<uint32_t*>(output +
batch_id * stride_bz_output +
head_id * stride_h_output +
token_id * stride_seq_output +
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD / 2)[0] = e2m1Vals[0];
} else {
reinterpret_cast<uint64_t*>(output +
batch_id * stride_bz_output +
head_id * stride_h_output +
token_id * stride_seq_output +
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD / 2)[0] = reinterpret_cast<uint64_t*>(e2m1Vals)[0];
}
uint8_t* output_sf_save_base = output_sf + batch_id * stride_bz_output_sf + head_id * stride_h_output_sf + (token_id / 64) * 64 * stride_seq_output_sf;
uint32_t token_id_local = token_id % 64;
if constexpr (CVT_FP4_ELTS_PER_THREAD == 16) {
uint32_t col_id_local = threadIdx.x % NUM_THREADS_PER_TOKEN;
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
(token_id_local / 16) * 4 + (token_id_local % 16) * 16;
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
} else {
if (threadIdx.x % 2 == 0) {
uint32_t col_id_local = (threadIdx.x % NUM_THREADS_PER_TOKEN) / 2;
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
(token_id_local / 16) * 4 + (token_id_local % 16) * 16;
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
}
}
}
template <uint32_t head_dim, uint32_t BLOCK_SIZE, typename T>
__global__ void scaled_fp4_quant_trans_kernel(
const T* input, uint8_t* output, uint8_t* output_sf,
int batch_size, int num_heads, int num_tokens,
int stride_bz_input, int stride_h_input, int stride_seq_input,
int stride_bz_output, int stride_h_output, int stride_d_output,
int stride_bz_output_sf, int stride_h_output_sf, int stride_d_output_sf) {
static_assert(std::is_same<T, half>::value || std::is_same<T, nv_bfloat16>::value, "Only half and bfloat16 input are supported");
using PackedVec = PackedVec<T>;
const int batch_id = blockIdx.y;
const int head_id = blockIdx.z;
const int token_block_id = blockIdx.x;
static_assert(CVT_FP4_ELTS_PER_THREAD == 8 || CVT_FP4_ELTS_PER_THREAD == 16,
"CVT_FP4_ELTS_PER_THREAD must be 8 or 16");
static_assert(sizeof(PackedVec) == sizeof(T) * CVT_FP4_ELTS_PER_THREAD,
"Vec size is not matched.");
constexpr uint32_t NUM_THREADS_PER_TOKEN = head_dim / CVT_FP4_ELTS_PER_THREAD;
constexpr uint32_t NUM_THREADS_PER_SEQ = BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD;
// load input
const int token_id = token_block_id * BLOCK_SIZE + threadIdx.x / NUM_THREADS_PER_TOKEN;
// Permute V rows within each 32-element block so the PV MMA K-indexed
// access reads the correct CLayout N-indexed values (Edenzzzz causal fix).
const int k_intra = token_id & 31;
const int load_token_id = (token_id & ~31)
| ((k_intra & 6) << 2) | ((k_intra & 24) >> 2) | (k_intra & 1);
PackedVec in_vec;
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
reinterpret_cast<uint32_t&>(in_vec.elts[i]) = 0;
}
if (load_token_id < num_tokens) {
in_vec = reinterpret_cast<PackedVec const*>(input +
batch_id * stride_bz_input + // batch dim
head_id * stride_h_input + // head dim
load_token_id * stride_seq_input + // seq dim (permuted)
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD)[0]; // feature dim
}
// transpose
__shared__ T shared_input[BLOCK_SIZE * head_dim];
reinterpret_cast<PackedVec*>(shared_input)[threadIdx.x] = in_vec;
__syncthreads();
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
in_vec.elts[i].x = shared_input[(threadIdx.x / NUM_THREADS_PER_SEQ) + ((threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD + 2 * i) * head_dim];
in_vec.elts[i].y = shared_input[(threadIdx.x / NUM_THREADS_PER_SEQ) + ((threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD + 2 * i + 1) * head_dim];
}
// calculate max of every consecutive 16 elements
auto localMax = __habs2(in_vec.elts[0]);
#pragma unroll
for (int i = 1; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) { // local max
localMax = __hmax2(localMax, __habs2(in_vec.elts[i]));
}
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) { // shuffle across two threads
localMax = __hmax2(__shfl_xor_sync(0xffffffff, localMax, 1, 32), localMax);
}
float vecMax = float(__hmax(localMax.x, localMax.y));
// scaling factor
float SFValue = vecMax / 6.0f;
uint8_t SFValueFP8;
reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8) = __nv_fp8_e4m3(SFValue);
SFValue = float(reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8));
float SFValueInv = (SFValue == 0.0f) ? 0.0f : 1.0f / SFValue;
// convert input to float2 and apply scale
float2 fp2Vals[CVT_FP4_ELTS_PER_THREAD / 2];
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
if constexpr (std::is_same<T, half>::value) {
fp2Vals[i] = __half22float2(in_vec.elts[i]);
} else {
fp2Vals[i] = __bfloat1622float2(in_vec.elts[i]);
}
fp2Vals[i].x = fp2Vals[i].x * SFValueInv;
fp2Vals[i].y = fp2Vals[i].y * SFValueInv;
}
// convert to e2m1
uint32_t e2m1Vals[CVT_FP4_ELTS_PER_THREAD / 8];
#pragma unroll
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 8; i++) {
e2m1Vals[i] = fp32_vec_to_e2m1(fp2Vals + i * 4);
}
// Skip out-of-range tokens: never write past the sequence when num_tokens
// is not a multiple of BLOCK_SIZE (out-of-bounds write fix).
const int write_token_id = token_block_id * BLOCK_SIZE +
(threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD;
if (write_token_id >= num_tokens) return;
// save
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) {
reinterpret_cast<uint32_t*>(output +
batch_id * stride_bz_output +
head_id * stride_h_output +
(threadIdx.x / NUM_THREADS_PER_SEQ) * stride_d_output +
(token_block_id * BLOCK_SIZE + (threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD) / 2)[0] = e2m1Vals[0];
} else {
reinterpret_cast<uint64_t*>(output +
batch_id * stride_bz_output +
head_id * stride_h_output +
(threadIdx.x / NUM_THREADS_PER_SEQ) * stride_d_output +
(token_block_id * BLOCK_SIZE + (threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD) / 2)[0] = reinterpret_cast<uint64_t*>(e2m1Vals)[0];
}
uint8_t *output_sf_save_base = output_sf +
batch_id * stride_bz_output_sf +
head_id * stride_h_output_sf +
(threadIdx.x / NUM_THREADS_PER_SEQ / 64) * 64 * stride_d_output_sf;
uint32_t row_id_local = (threadIdx.x / NUM_THREADS_PER_SEQ) % 64;
if constexpr (CVT_FP4_ELTS_PER_THREAD == 16) {
uint32_t col_id_local = token_block_id * BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD + threadIdx.x % NUM_THREADS_PER_SEQ;
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
(row_id_local / 16) * 4 + (row_id_local % 16) * 16;
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
} else {
if (threadIdx.x % 2 == 0) {
uint32_t col_id_local = token_block_id * BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD + (threadIdx.x % NUM_THREADS_PER_SEQ) / 2;
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
(row_id_local / 16) * 4 + (row_id_local % 16) * 16;
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
}
}
}
void scaled_fp4_quant(torch::Tensor const& input,
torch::Tensor const& output,
torch::Tensor const& output_sf,
int tensor_layout) {
constexpr int BLOCK_SIZE = flash::BLOCK_M;
CHECK_CUDA(input);
CHECK_CUDA(output);
CHECK_CUDA(output_sf);
CHECK_LASTDIM_CONTIGUOUS(input);
CHECK_LASTDIM_CONTIGUOUS(output);
CHECK_LASTDIM_CONTIGUOUS(output_sf);
CHECK_DTYPE(output, at::ScalarType::Byte);
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
CHECK_DIMS(input, 4);
CHECK_DIMS(output, 4);
CHECK_DIMS(output_sf, 4);
const int batch_size = input.size(0);
const int head_dim = input.size(3);
const int stride_bz_input = input.stride(0);
const int stride_bz_output = output.stride(0);
const int stride_bz_output_sf = output_sf.stride(0);
int num_tokens, num_heads;
int stride_seq_input, stride_seq_output, stride_seq_output_sf;
int stride_h_input, stride_h_output, stride_h_output_sf;
if (tensor_layout == 0) {
num_tokens = input.size(1);
num_heads = input.size(2);
stride_seq_input = input.stride(1);
stride_seq_output = output.stride(1);
stride_seq_output_sf = output_sf.stride(1);
stride_h_input = input.stride(2);
stride_h_output = output.stride(2);
stride_h_output_sf = output_sf.stride(2);
CHECK_SHAPE(output, batch_size, num_tokens, num_heads, head_dim / 2);
CHECK_SHAPE(output_sf, batch_size, num_tokens, num_heads, head_dim / 16);
} else {
num_tokens = input.size(2);
num_heads = input.size(1);
stride_seq_input = input.stride(2);
stride_seq_output = output.stride(2);
stride_seq_output_sf = output_sf.stride(2);
stride_h_input = input.stride(1);
stride_h_output = output.stride(1);
stride_h_output_sf = output_sf.stride(1);
CHECK_SHAPE(output, batch_size, num_heads, num_tokens, head_dim / 2);
CHECK_SHAPE(output_sf, batch_size, num_heads, num_tokens, head_dim / 16);
}
auto input_dtype = input.scalar_type();
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
scaled_fp4_quant_kernel<HEAD_DIM, BLOCK_SIZE, false, c_type>
<<<grid, block, 0, stream>>>(
reinterpret_cast<c_type*>(input.data_ptr()),
reinterpret_cast<uint8_t*>(output.data_ptr()),
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
batch_size, num_heads, num_tokens,
stride_bz_input, stride_h_input, stride_seq_input,
stride_bz_output, stride_h_output, stride_seq_output,
stride_bz_output_sf, stride_h_output_sf, stride_seq_output_sf);
});
});
}
void scaled_fp4_quant_permute(torch::Tensor const& input,
torch::Tensor const& output,
torch::Tensor const& output_sf,
int tensor_layout) {
constexpr int BLOCK_SIZE = flash::BLOCK_M;
CHECK_CUDA(input);
CHECK_CUDA(output);
CHECK_CUDA(output_sf);
CHECK_LASTDIM_CONTIGUOUS(input);
CHECK_LASTDIM_CONTIGUOUS(output);
CHECK_LASTDIM_CONTIGUOUS(output_sf);
CHECK_DTYPE(output, at::ScalarType::Byte);
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
CHECK_DIMS(input, 4);
CHECK_DIMS(output, 4);
CHECK_DIMS(output_sf, 4);
const int batch_size = input.size(0);
const int head_dim = input.size(3);
const int stride_bz_input = input.stride(0);
const int stride_bz_output = output.stride(0);
const int stride_bz_output_sf = output_sf.stride(0);
int num_tokens, num_heads;
int stride_seq_input, stride_seq_output, stride_seq_output_sf;
int stride_h_input, stride_h_output, stride_h_output_sf;
if (tensor_layout == 0) {
num_tokens = input.size(1);
num_heads = input.size(2);
stride_seq_input = input.stride(1);
stride_seq_output = output.stride(1);
stride_seq_output_sf = output_sf.stride(1);
stride_h_input = input.stride(2);
stride_h_output = output.stride(2);
stride_h_output_sf = output_sf.stride(2);
CHECK_SHAPE(output, batch_size, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, num_heads, head_dim / 2);
CHECK_SHAPE(output_sf, batch_size, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, num_heads, head_dim / 16);
} else {
num_tokens = input.size(2);
num_heads = input.size(1);
stride_seq_input = input.stride(2);
stride_seq_output = output.stride(2);
stride_seq_output_sf = output_sf.stride(2);
stride_h_input = input.stride(1);
stride_h_output = output.stride(1);
stride_h_output_sf = output_sf.stride(1);
CHECK_SHAPE(output, batch_size, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, head_dim / 2);
CHECK_SHAPE(output_sf, batch_size, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, head_dim / 16);
}
auto input_dtype = input.scalar_type();
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
constexpr int BLOCK_SIZE = flash::BLOCK_M;
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
scaled_fp4_quant_kernel<HEAD_DIM, BLOCK_SIZE, true, c_type>
<<<grid, block, 0, stream>>>(
reinterpret_cast<c_type*>(input.data_ptr()),
reinterpret_cast<uint8_t*>(output.data_ptr()),
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
batch_size, num_heads, num_tokens,
stride_bz_input, stride_h_input, stride_seq_input,
stride_bz_output, stride_h_output, stride_seq_output,
stride_bz_output_sf, stride_h_output_sf, stride_seq_output_sf);
});
});
}
void scaled_fp4_quant_trans(torch::Tensor const& input,
torch::Tensor const& output,
torch::Tensor const& output_sf,
int tensor_layout) {
constexpr int BLOCK_SIZE = flash::BLOCK_M;
CHECK_CUDA(input);
CHECK_CUDA(output);
CHECK_CUDA(output_sf);
CHECK_LASTDIM_CONTIGUOUS(input);
CHECK_LASTDIM_CONTIGUOUS(output);
CHECK_LASTDIM_CONTIGUOUS(output_sf);
CHECK_DTYPE(output, at::ScalarType::Byte);
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
CHECK_DIMS(input, 4);
CHECK_DIMS(output, 4);
CHECK_DIMS(output_sf, 4);
const int batch_size = input.size(0);
const int head_dim = input.size(3);
const int stride_bz_input = input.stride(0);
const int stride_bz_output = output.stride(0);
const int stride_bz_output_sf = output_sf.stride(0);
int num_tokens, num_heads;
int stride_seq_input;
int stride_d_output, stride_d_output_sf;
int stride_h_input, stride_h_output, stride_h_output_sf;
if (tensor_layout == 0) {
num_tokens = input.size(1);
num_heads = input.size(2);
stride_seq_input = input.stride(1);
stride_d_output = output.stride(1);
stride_d_output_sf = output_sf.stride(1);
stride_h_input = input.stride(2);
stride_h_output = output.stride(2);
stride_h_output_sf = output_sf.stride(2);
CHECK_SHAPE(output, batch_size, head_dim, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 2);
CHECK_SHAPE(output_sf, batch_size, head_dim, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 16);
} else {
num_tokens = input.size(2);
num_heads = input.size(1);
stride_seq_input = input.stride(2);
stride_d_output = output.stride(2);
stride_d_output_sf = output_sf.stride(2);
stride_h_input = input.stride(1);
stride_h_output = output.stride(1);
stride_h_output_sf = output_sf.stride(1);
CHECK_SHAPE(output, batch_size, num_heads, head_dim, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 2);
CHECK_SHAPE(output_sf, batch_size, num_heads, head_dim, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 16);
}
auto input_dtype = input.scalar_type();
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
scaled_fp4_quant_trans_kernel<HEAD_DIM, BLOCK_SIZE, c_type>
<<<grid, block, 0, stream>>>(
reinterpret_cast<c_type*>(input.data_ptr()),
reinterpret_cast<uint8_t*>(output.data_ptr()),
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
batch_size, num_heads, num_tokens,
stride_bz_input, stride_h_input, stride_seq_input,
stride_bz_output, stride_h_output, stride_d_output,
stride_bz_output_sf, stride_h_output_sf, stride_d_output_sf);
});
});
}
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
m.def("scaled_fp4_quant", &scaled_fp4_quant);
m.def("scaled_fp4_quant_permute", &scaled_fp4_quant_permute);
m.def("scaled_fp4_quant_trans", &scaled_fp4_quant_trans);
}
+9 -19
View File
@@ -78,26 +78,16 @@ print(f'{mj}.{mn}')"
}
if [ "${GPU_BACKEND}" = "CUDA" ]; then
# Compute capability drives the arch/TK defaults below. Prefer an explicit
# TORCH_CUDA_ARCH_LIST (works on GPU-less build machines such as CI/Docker);
# only probe a live GPU via torch when no arch was provided.
if [ -n "${TORCH_CUDA_ARCH_LIST:-}" ]; then
echo "Using TORCH_CUDA_ARCH_LIST=${TORCH_CUDA_ARCH_LIST} (skipping torch GPU probe)"
first_arch="${TORCH_CUDA_ARCH_LIST%%[;, ]*}" # first entry, e.g. 9.0a
first_arch="${first_arch%[af]}" # strip trailing a/f suffix
cc_major="${first_arch%%.*}"
cc_minor="${first_arch##*.}"
else
detected_cc="$(detect_with_torch)" || {
echo "ERROR: torch-based CUDA arch detection failed and TORCH_CUDA_ARCH_LIST is unset." >&2
echo " Set TORCH_CUDA_ARCH_LIST (e.g. 9.0a) for GPU-less builds, or build where CUDA is available." >&2
exit 1
}
cc_major="${detected_cc%%.*}"
cc_minor="${detected_cc##*.}"
echo "Detected compute capability via torch: ${detected_cc}"
fi
detected_cc="$(detect_with_torch)" || {
echo "ERROR: torch-based CUDA arch detection failed in uv environment." >&2
echo " Ensure torch is installed and CUDA is available in the uv-selected Python." >&2
exit 1
}
cc_major="${detected_cc%%.*}"
cc_minor="${detected_cc##*.}"
cmake_arch="${cc_major}${cc_minor}"
echo "Detected compute capability via torch: ${detected_cc} (sm_${cmake_arch})"
# Respect explicit overrides.
if [ -z "${TORCH_CUDA_ARCH_LIST:-}" ]; then
+1 -1
View File
@@ -32,4 +32,4 @@ dependencies = [
[tool.scikit-build]
cmake.build-type = "Release"
minimum-version = "build-system.requires"
wheel.packages = ["python/fastvideo_kernel", "attn_qat_infer"]
wheel.packages = ["python/fastvideo_kernel"]
@@ -25,17 +25,12 @@ from fastvideo_kernel.turbodiffusion_ops import (
int8_quant,
)
from fastvideo_kernel.block_sparse_attn_varlen import (
block_sparse_attn_varlen,
)
__all__ = [
"sliding_tile_attention",
"video_sparse_attn",
"video_sparse_attn_bshd",
"block_sparse_attn",
"block_sparse_attn_from_indices",
"block_sparse_attn_varlen",
"moba_attn_varlen",
"process_moba_input",
"process_moba_output",
@@ -1,207 +0,0 @@
"""Variable-length block-sparse attention via sequence packing.
Packs multiple variable-length sequences into a single [1, H, T_total, D]
tensor and delegates to the existing block_sparse_attn_from_indices kernel
in a single launch. No kernel modifications required.
"""
from __future__ import annotations
from typing import Sequence
import torch
from .block_sparse_attn import block_sparse_attn_from_indices
BLOCK_SIZE = 64
def _scatter_to_padded(
src: torch.Tensor,
block_sizes: torch.Tensor,
block_size: int,
dst: torch.Tensor,
dst_offset: int,
src_start: int,
src_end: int,
) -> None:
"""Copy tokens from a flat source into block-aligned positions in dst.
Each block occupies exactly `block_size` slots in dst. The first
`block_sizes[b]` slots of block *b* receive real tokens; the remainder
stays zero (padding the kernel expects).
src: [total_tokens, H, D]
dst: [1, H, total_padded, D]
block_sizes: [num_blocks] int32, actual token count per block.
"""
src_pos = src_start
dst_pos = dst_offset
sizes = block_sizes.cpu().tolist()
for actual in sizes:
actual = min(actual, src_end - src_pos)
if actual > 0:
dst[:, :, dst_pos:dst_pos + actual, :] = (
src[src_pos:src_pos + actual].transpose(0, 1).unsqueeze(0)
)
src_pos += actual
dst_pos += block_size
def _gather_from_padded(
src: torch.Tensor,
block_sizes: torch.Tensor,
block_size: int,
dst: torch.Tensor,
src_offset: int,
dst_start: int,
dst_end: int,
) -> None:
"""Inverse of _scatter_to_padded: extract real tokens from padded blocks.
src: [1, H, total_padded, D]
dst: [total_tokens, H, D]
"""
src_pos = src_offset
dst_pos = dst_start
sizes = block_sizes.cpu().tolist()
for actual in sizes:
actual = min(actual, dst_end - dst_pos)
if actual > 0:
dst[dst_pos:dst_pos + actual] = (
src[0, :, src_pos:src_pos + actual, :].transpose(0, 1)
)
dst_pos += actual
src_pos += block_size
def block_sparse_attn_varlen(
q: torch.Tensor,
k: torch.Tensor,
v: torch.Tensor,
cu_seqlens_q: torch.Tensor,
cu_seqlens_kv: torch.Tensor,
q2k_idx_list: Sequence[torch.Tensor],
q2k_num_list: Sequence[torch.Tensor],
variable_block_sizes_list: Sequence[torch.Tensor],
q_variable_block_sizes_list: Sequence[torch.Tensor] | None = None,
block_size: int = BLOCK_SIZE,
) -> torch.Tensor:
"""Block-sparse attention over packed variable-length sequences.
Args:
q: [total_q_tokens, H, D] packed query tensor.
k: [total_kv_tokens, H, D] packed key tensor.
v: [total_kv_tokens, H, D] packed value tensor.
cu_seqlens_q: [N+1] int32, cumulative Q token offsets.
cu_seqlens_kv: [N+1] int32, cumulative KV token offsets.
q2k_idx_list: Per-sequence q2k_idx tensors, each [1, H, Nq_i, Mk].
q2k_num_list: Per-sequence q2k_num tensors, each [1, H, Nq_i].
variable_block_sizes_list: Per-sequence KV block sizes, each [Nkv_i].
q_variable_block_sizes_list: Per-sequence Q block sizes, each [Nq_i].
If None, each Q block is assumed to be exactly `block_size` tokens.
block_size: Attention block size (default 64).
Returns:
out: [total_q_tokens, H, D] packed output tensor.
"""
device = q.device
dtype = q.dtype
num_heads = q.shape[1]
head_dim = q.shape[2]
num_seqs = cu_seqlens_q.shape[0] - 1
cu_q = cu_seqlens_q.cpu().tolist()
cu_kv = cu_seqlens_kv.cpu().tolist()
padded_q_lens = []
padded_kv_lens = []
q_block_offsets = [0]
kv_block_offsets = [0]
q_vbs_resolved = []
for i in range(num_seqs):
n_q_blocks = q2k_num_list[i].shape[-1]
n_kv_blocks = variable_block_sizes_list[i].numel()
padded_q_lens.append(n_q_blocks * block_size)
padded_kv_lens.append(n_kv_blocks * block_size)
q_block_offsets.append(q_block_offsets[-1] + n_q_blocks)
kv_block_offsets.append(kv_block_offsets[-1] + n_kv_blocks)
if q_variable_block_sizes_list is not None:
q_vbs_resolved.append(q_variable_block_sizes_list[i])
else:
q_vbs_resolved.append(
torch.full((n_q_blocks,), block_size, dtype=torch.int32)
)
total_padded_q = sum(padded_q_lens)
total_padded_kv = sum(padded_kv_lens)
q_packed = torch.zeros(1, num_heads, total_padded_q, head_dim, device=device, dtype=dtype)
k_packed = torch.zeros(1, num_heads, total_padded_kv, head_dim, device=device, dtype=dtype)
v_packed = torch.zeros(1, num_heads, total_padded_kv, head_dim, device=device, dtype=dtype)
q_offset = 0
kv_offset = 0
for i in range(num_seqs):
_scatter_to_padded(
q, q_vbs_resolved[i], block_size,
q_packed, q_offset, cu_q[i], cu_q[i + 1],
)
_scatter_to_padded(
k, variable_block_sizes_list[i], block_size,
k_packed, kv_offset, cu_kv[i], cu_kv[i + 1],
)
_scatter_to_padded(
v, variable_block_sizes_list[i], block_size,
v_packed, kv_offset, cu_kv[i], cu_kv[i + 1],
)
q_offset += padded_q_lens[i]
kv_offset += padded_kv_lens[i]
total_q_blocks = q_block_offsets[-1]
max_kv_per_q = max(t.shape[-1] for t in q2k_idx_list)
global_q2k_idx = torch.zeros(
1, num_heads, total_q_blocks, max_kv_per_q,
dtype=torch.int32, device=device,
)
global_q2k_num = torch.zeros(
1, num_heads, total_q_blocks,
dtype=torch.int32, device=device,
)
global_vbs_parts = []
for i in range(num_seqs):
qb_start = q_block_offsets[i]
qb_end = q_block_offsets[i + 1]
n_q_blocks = qb_end - qb_start
kv_offset_blocks = kv_block_offsets[i]
idx = q2k_idx_list[i]
num = q2k_num_list[i]
vbs = variable_block_sizes_list[i]
mk = idx.shape[-1]
global_q2k_idx[:, :, qb_start:qb_end, :mk] = idx[:, :, :n_q_blocks, :] + kv_offset_blocks
global_q2k_num[:, :, qb_start:qb_end] = num[:, :, :n_q_blocks]
global_vbs_parts.append(vbs)
global_vbs = torch.cat(global_vbs_parts, dim=0).to(torch.int32).contiguous()
out_packed, _ = block_sparse_attn_from_indices(
q_packed, k_packed, v_packed,
global_q2k_idx, global_q2k_num, global_vbs,
)
out = torch.zeros(cu_q[-1], num_heads, head_dim, device=device, dtype=dtype)
q_offset = 0
for i in range(num_seqs):
_gather_from_padded(
out_packed, q_vbs_resolved[i], block_size,
out, q_offset, cu_q[i], cu_q[i + 1],
)
q_offset += padded_q_lens[i]
return out
File diff suppressed because it is too large Load Diff
@@ -1,55 +0,0 @@
"""Compatibility shim for the legacy non-QAT Triton attention import path.
Historically callers imported
``fastvideo_kernel.triton_kernels.fused_attention`` directly. The shared
implementation now lives in ``attn_qat_train.py`` and is parameterized by the
``IS_QAT`` flag. This module preserves the original public API for tests and
downstream users while always dispatching to the non-QAT configuration.
"""
from __future__ import annotations
import torch
from .attn_qat_train import attention as _attention
def attention(
q: torch.Tensor,
k: torch.Tensor,
v: torch.Tensor,
causal: bool,
sm_scale: float,
warp_specialize: bool = True,
) -> torch.Tensor:
"""Run the shared Triton attention kernel in non-QAT mode."""
use_qat_qkv_backward = True
smooth_k = False
is_qat = False
two_level_quant_p = False
fake_quant_p = False
use_high_prec_o = False
smooth_q = False
use_global_sf_p = False
use_global_sf_qkv = False
return _attention(
q,
k,
v,
causal,
sm_scale,
use_qat_qkv_backward,
smooth_k,
warp_specialize,
is_qat,
two_level_quant_p,
fake_quant_p,
use_high_prec_o,
smooth_q,
use_global_sf_p,
use_global_sf_qkv,
)
__all__ = ["attention"]
@@ -1,237 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# Adapted from https://github.com/triton-lang/triton/blob/main/python/triton_kernels/triton_kernels/numerics_details/mxfp_details/_upcast_from_mxfp.py
# and https://github.com/triton-lang/triton/blob/main/python/triton_kernels/triton_kernels/numerics_details/mxfp_details/_downcast_to_mxfp.py
import triton
import triton.language as tl
from triton.language.target_info import cuda_capability_geq
MXFP_BLOCK_SIZE = tl.constexpr(16)
@triton.jit
def _compute_quant_and_scale(
src_tensor,
valid_src_mask,
mx_tensor_dtype: tl.constexpr = tl.uint8,
use_global_sf=True,
two_level_quant_P=False,
):
BLOCK_SIZE_OUT_DIM: tl.constexpr = src_tensor.shape[0]
BLOCK_SIZE_QUANT_DIM: tl.constexpr = src_tensor.shape[1]
BLOCK_SIZE_QUANT_MX_SCALE: tl.constexpr = src_tensor.shape[1] // MXFP_BLOCK_SIZE
is_fp4: tl.constexpr = mx_tensor_dtype == tl.uint8
tl.static_assert(
is_fp4
or mx_tensor_dtype == tl.float8e4nv
or mx_tensor_dtype == tl.float8e5,
"mx_tensor_dtype must be uint8, float8e4nv, or float8e5",
)
# Explicit cast to fp32 since most ops are not supported on bfloat16. We avoid needless conversions to and from bf16
f32_tensor = src_tensor.to(tl.float32)
abs_tensor = tl.abs(f32_tensor)
abs_tensor = tl.where(valid_src_mask, abs_tensor, -1.0) # Don't consider padding tensors in scale computation
if two_level_quant_P:
# row max from SageAttn3 paper
global_max_val = tl.max(f32_tensor, axis=1, keep_dims=True) # (BLOCK_SIZE_OUT_DIM, 1)
global_max_val = tl.maximum(global_max_val, 1e-8)
s_enc = ((6 * 448) / global_max_val).reshape([BLOCK_SIZE_OUT_DIM, 1, 1])
s_dec = (1 / s_enc)
abs_tensor = tl.reshape(abs_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
if use_global_sf and not two_level_quant_P:
global_max_val = tl.max(abs_tensor)
# Avoid division by zero: if all values are padding (max is 0), use a default scale
global_max_val = tl.maximum(global_max_val, 1e-8)
s_enc = (6 * 448) / global_max_val
s_dec = (1 / s_enc)
elif not two_level_quant_P and not use_global_sf:
s_dec = 1.0
s_enc = 1.0
max_val = tl.max(abs_tensor, axis=2, keep_dims=True) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1) # per block maxima
s_dec_b = max_val / 6 # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
s_dec_b_e4m3 = (s_dec_b * s_enc).to(tl.float8e4nv) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
s_enc_b = 1 / (s_dec_b_e4m3.to(tl.float32) * s_dec) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
f32_tensor = tl.reshape(f32_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
quant_tensor = f32_tensor * s_enc_b
# Reshape the tensors after scaling
quant_tensor = quant_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM])
# Set the invalid portions of the tensor to 0. This will ensure that any padding tensors are 0 in the mx format.
quant_tensor = tl.where(valid_src_mask, quant_tensor, 0.0)
dequant_scale = s_dec_b_e4m3.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE])
if is_fp4 and cuda_capability_geq(10, 0):
# Convert scaled values to two f32 lanes and use PTX cvt to e2m1x2 with two f32 operands.
pairs = tl.reshape(quant_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM // 2, 2])
lo_f, hi_f = tl.split(pairs)
lo_f32 = lo_f.to(tl.float32)
hi_f32 = hi_f.to(tl.float32)
# Inline PTX: cvt.rn.satfinite.e2m1x2.f32 takes two f32 sources and produces one .b8 packed e2m1x2.
out_tensor = tl.inline_asm_elementwise(
"""
{
.reg .b8 r;
cvt.rn.satfinite.e2m1x2.f32 r, $1, $2;
mov.b32 $0, {r, r, r, r};
}
""",
constraints="=r,f,f",
args=[hi_f32, lo_f32],
dtype=tl.uint8,
is_pure=True,
pack=1,
)
elif is_fp4:
quant_tensor = quant_tensor.to(tl.uint32, bitcast=True)
signs = quant_tensor & 0x80000000
exponents = (quant_tensor >> 23) & 0xFF
mantissas_orig = (quant_tensor & 0x7FFFFF)
# For RTNE: 0.25 < x < 0.75 maps to 0.5 (denormal); exactly 0.25 maps to 0.0
E8_BIAS = 127
E2_BIAS = 1
# Move implicit bit 1 at the beginning to mantissa for denormals
is_subnormal = exponents < E8_BIAS
adjusted_exponents = tl.core.sub(E8_BIAS, exponents + 1, sanitize_overflow=False)
mantissas_pre = (0x400000 | (mantissas_orig >> 1))
mantissas = tl.where(is_subnormal, mantissas_pre >> adjusted_exponents, mantissas_orig)
# For normal numbers, we change the bias from 127 to 1, and for subnormals, we keep exponent as 0.
exponents = tl.maximum(exponents, E8_BIAS - E2_BIAS) - (E8_BIAS - E2_BIAS)
# Combine sign, exponent, and mantissa, while saturating
# Round to nearest, ties to even (RTNE): use guard/sticky and LSB to decide increment
m2bits = mantissas >> 21
lsb_keep = (m2bits >> 1) & 0x1
guard = m2bits & 0x1
IS_SRC_FP32: tl.constexpr = src_tensor.dtype == tl.float32
if IS_SRC_FP32:
bit0_dropped = (mantissas_orig & 0x1) != 0
mask = (1 << tl.minimum(adjusted_exponents, 31)) - 1
dropped_post = (mantissas_pre & mask) != 0
sticky = is_subnormal & (bit0_dropped | dropped_post)
sticky |= ((mantissas & 0x1FFFFF) != 0).to(tl.uint32)
else:
sticky = ((mantissas & 0x1FFFFF) != 0).to(tl.uint32)
round_inc = guard & (sticky | lsb_keep)
e2m1_tmp = tl.minimum((((exponents << 2) | m2bits) + round_inc) >> 1, 0x7)
e2m1_value = ((signs >> 28) | e2m1_tmp).to(tl.uint8)
e2m1_value = tl.reshape(e2m1_value, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM // 2, 2])
evens, odds = tl.split(e2m1_value)
out_tensor = evens | (odds << 4)
else:
out_tensor = quant_tensor.to(mx_tensor_dtype)
return out_tensor, dequant_scale, s_dec
@triton.jit
def _compute_dequant(
mx_tensor,
scale,
s_dec,
BLOCK_SIZE_OUT_DIM: tl.constexpr,
BLOCK_SIZE_QUANT_DIM: tl.constexpr,
dst_dtype: tl.constexpr,
):
tl.static_assert(BLOCK_SIZE_QUANT_DIM % MXFP_BLOCK_SIZE == 0, f"Block size along quantization block must be a multiple of {MXFP_BLOCK_SIZE=}")
# uint8 signifies two fp4 e2m1 values packed into a single byte
mx_tensor_dtype: tl.constexpr = mx_tensor.dtype
tl.static_assert(dst_dtype == tl.float16 or dst_dtype == tl.bfloat16 or dst_dtype == tl.float32)
tl.static_assert(
mx_tensor_dtype == tl.uint8
or ((mx_tensor_dtype == tl.float8e4nv or mx_tensor_dtype == tl.float8e5) or mx_tensor_dtype == dst_dtype),
"mx_tensor_ptr must be uint8 or float8 or dst_dtype")
tl.static_assert(scale.dtype == tl.float8e4nv, "scale must be float8e4nv")
# Determine if we are dealing with fp8 types.
is_fp4: tl.constexpr = mx_tensor_dtype == tl.uint8
BLOCK_SIZE_QUANT_MX_SCALE: tl.constexpr = BLOCK_SIZE_QUANT_DIM // MXFP_BLOCK_SIZE
# Upcast the scale to the destination type.
if dst_dtype == tl.bfloat16:
dst_scale = scale.to(tl.bfloat16)
else:
dst_scale = scale.to(tl.float32)
if dst_dtype == tl.float16:
dst_scale = dst_scale.to(tl.float16)
# Now upcast the tensor.
intermediate_dtype: tl.constexpr = tl.bfloat16 if dst_dtype == tl.float32 else dst_dtype
if cuda_capability_geq(10, 0):
assert is_fp4
packed_u32 = tl.inline_asm_elementwise(
asm="""
{
.reg .b8 in_8;
.reg .f16x2 out;
cvt.u8.u32 in_8, $1;
cvt.rn.f16x2.e2m1x2 out, in_8;
mov.b32 $0, out;
}
""",
constraints="=r,r",
args=[mx_tensor], # tl.uint8 passed in as a 32-bit reg with value in low 8 bits
dtype=tl.uint32,
is_pure=True,
pack=1,
)
lo_u16 = (packed_u32 & 0xFFFF).to(tl.uint16)
hi_u16 = (packed_u32 >> 16).to(tl.uint16)
lo_f16 = lo_u16.to(tl.float16, bitcast=True)
hi_f16 = hi_u16.to(tl.float16, bitcast=True)
if intermediate_dtype == tl.float16:
x0, x1 = lo_f16, hi_f16
else:
x0 = lo_f16.to(intermediate_dtype)
x1 = hi_f16.to(intermediate_dtype)
dst_tensor = tl.interleave(x0, x1)
else:
assert is_fp4
dst_bias: tl.constexpr = 127 if intermediate_dtype == tl.bfloat16 else 15 # exponent bias
dst_0p5: tl.constexpr = 16128 if intermediate_dtype == tl.bfloat16 else 0x3800
dst_m_bits: tl.constexpr = 7 if intermediate_dtype == tl.bfloat16 else 10 # mantissa bits
# e2m1
em0 = mx_tensor & 0x07
em1 = mx_tensor & 0x70
x0 = (em0.to(tl.uint16) << (dst_m_bits - 1)) | ((mx_tensor & 0x08).to(tl.uint16) << 12)
x1 = (em1.to(tl.uint16) << (dst_m_bits - 5)) | ((mx_tensor & 0x80).to(tl.uint16) << 8)
# Three cases:
# 1) x is normal and non-zero: Correct bias
x0 = tl.where((em0 & 0x06) != 0, x0 + ((dst_bias - 1) << dst_m_bits), x0)
x1 = tl.where((em1 & 0x60) != 0, x1 + ((dst_bias - 1) << dst_m_bits), x1)
# 2) x is subnormal (x == 0bs001 where s is the sign): Map to +-0.5 in the dst type
x0 = tl.where(em0 == 0x01, dst_0p5 | (x0 & 0x8000), x0)
x1 = tl.where(em1 == 0x10, dst_0p5 | (x1 & 0x8000), x1)
# 3) x is zero, do nothing
dst_tensor = tl.interleave(x0, x1).to(intermediate_dtype, bitcast=True)
dst_tensor = dst_tensor.to(dst_dtype)
# Reshape for proper broadcasting: the scale was stored with a 16‐sized “inner” grouping.
dst_tensor = dst_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
dst_scale = dst_scale.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1])
scale = scale.reshape(dst_scale.shape)
out_tensor = dst_tensor * dst_scale * s_dec # NVFP4 has the additional global scale factor
if dst_dtype == tl.float32:
max_fin = 3.4028234663852886e+38
elif dst_dtype == tl.bfloat16:
max_fin = 3.3895313892515355e+38
else:
tl.static_assert(dst_dtype == tl.float16)
max_fin = 65504
out_tensor = tl.clamp(out_tensor, min=-max_fin, max=max_fin)
out_tensor = out_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM])
out_tensor = out_tensor.to(dst_dtype)
return out_tensor
@@ -1,80 +0,0 @@
import triton
import triton.language as tl
from .nvfp4_utils import _compute_quant_and_scale, _compute_dequant
@triton.jit
def fake_quantize(src_tensor, valid_src_mask, BLOCK_SIZE_OUT_DIM: tl.constexpr,
BLOCK_SIZE_QUANT_DIM: tl.constexpr,
dst_dtype: tl.constexpr,
mx_tensor_dtype: tl.constexpr = tl.uint8,
use_global_sf: tl.constexpr = True,
two_level_quant_P: tl.constexpr = False):
high_prec_src_tensor = src_tensor
src_tensor, src_scale, src_s_dec = _compute_quant_and_scale(src_tensor=src_tensor,
valid_src_mask=valid_src_mask,
mx_tensor_dtype=mx_tensor_dtype,
use_global_sf=use_global_sf,
two_level_quant_P=two_level_quant_P)
src_tensor = _compute_dequant(mx_tensor=src_tensor,
scale=src_scale,
s_dec=src_s_dec,
BLOCK_SIZE_OUT_DIM=BLOCK_SIZE_OUT_DIM,
BLOCK_SIZE_QUANT_DIM=BLOCK_SIZE_QUANT_DIM,
dst_dtype=dst_dtype)
return src_tensor, high_prec_src_tensor.to(src_tensor.dtype)
@triton.jit
def fake_quantize_q(Q, fake_Q, stride_z_q, stride_h_q,
stride_tok_q, stride_d_q,
fake_stride_z_q, fake_stride_h_q,
fake_stride_tok_q, fake_stride_d_q,
H, N_CTX_Q,
BLOCK_M: tl.constexpr,
HEAD_DIM: tl.constexpr,
use_global_sf: tl.constexpr = True):
bhid = tl.program_id(1)
adj_q = (stride_h_q * (bhid % H) + stride_z_q * (bhid // H))
fake_adj_q = (fake_stride_h_q * (bhid % H) + fake_stride_z_q * (bhid // H))
Q += adj_q
fake_Q += fake_adj_q
pid = tl.program_id(0)
start_m = pid * BLOCK_M
offs_m = start_m + tl.arange(0, BLOCK_M)
offs_k = tl.arange(0, HEAD_DIM)
q_valid = offs_m < N_CTX_Q
q = tl.load(Q + offs_m[:, None] * stride_tok_q + offs_k[None, :] * stride_d_q, mask=q_valid[:, None], other=0.0)
q, _ = fake_quantize(src_tensor=q, valid_src_mask=q_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_M, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=q.dtype, use_global_sf=use_global_sf)
tl.store(fake_Q + offs_m[:, None] * fake_stride_tok_q + offs_k[None, :] * fake_stride_d_q, q, mask=q_valid[:, None])
@triton.jit
def fake_quantize_kv(K, V, fake_K, fake_V, stride_z_kv, stride_h_kv,
stride_tok_kv, stride_d_kv,
fake_stride_z_kv, fake_stride_h_kv,
fake_stride_tok_kv, fake_stride_d_kv,
H, N_CTX_KV,
BLOCK_N: tl.constexpr,
HEAD_DIM: tl.constexpr,
use_global_sf: tl.constexpr = True):
bhid = tl.program_id(1)
adj_kv = (stride_h_kv * (bhid % H) + stride_z_kv * (bhid // H))
fake_adj_kv = (fake_stride_h_kv * (bhid % H) + fake_stride_z_kv * (bhid // H))
K += adj_kv
V += adj_kv
fake_K += fake_adj_kv
fake_V += fake_adj_kv
pid = tl.program_id(0)
start_n = pid * BLOCK_N
offs_n = start_n + tl.arange(0, BLOCK_N)
offs_k = tl.arange(0, HEAD_DIM)
kv_valid = offs_n < N_CTX_KV
k_block = tl.load(K + offs_n[:, None] * stride_tok_kv + offs_k[None, :] * stride_d_kv, mask=kv_valid[:, None], other=0.0)
v_block = tl.load(V + offs_n[:, None] * stride_tok_kv + offs_k[None, :] * stride_d_kv, mask=kv_valid[:, None], other=0.0)
k, _ = fake_quantize(src_tensor=k_block, valid_src_mask=kv_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_N, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=k_block.dtype, use_global_sf=use_global_sf)
v, _ = fake_quantize(src_tensor=v_block, valid_src_mask=kv_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_N, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=v_block.dtype, use_global_sf=use_global_sf)
tl.store(fake_K + offs_n[:, None] * fake_stride_tok_kv + offs_k[None, :] * fake_stride_d_kv, k, mask=kv_valid[:, None])
tl.store(fake_V + offs_n[:, None] * fake_stride_tok_kv + offs_k[None, :] * fake_stride_d_kv, v, mask=kv_valid[:, None])
File diff suppressed because it is too large Load Diff
-434
View File
@@ -1,434 +0,0 @@
"""Correctness tests for variable-length block-sparse attention.
Reference: per-sequence calls to block_sparse_attn_from_indices.
Test: single-launch via block_sparse_attn_varlen.
Tests cover both forward and backward (gradient) correctness.
"""
import torch
import pytest
from .test_vsa import (
BLOCK_M,
generate_variable_block_sizes,
get_non_pad_index,
vsa_pad,
generate_tensor,
)
from .utils import generate_block_sparse_mask_for_function
from fastvideo_kernel.block_sparse_attn import (
block_sparse_attn_from_indices,
_map_to_index,
)
from fastvideo_kernel.block_sparse_attn_varlen import block_sparse_attn_varlen
def _reference_per_sequence(
q_list, k_list, v_list,
block_masks, vbs_list,
non_pad_q_list, non_pad_kv_list,
q_nblocks_list, kv_nblocks_list,
):
"""Run per-sequence block_sparse_attn and concat outputs."""
outs = []
for i in range(len(q_list)):
q_pad = vsa_pad(q_list[i], non_pad_q_list[i], q_nblocks_list[i], BLOCK_M)
k_pad = vsa_pad(k_list[i], non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
v_pad = vsa_pad(v_list[i], non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
q2k_idx, q2k_num = _map_to_index(block_masks[i].unsqueeze(0))
o_pad, _ = block_sparse_attn_from_indices(
q_pad, k_pad, v_pad, q2k_idx, q2k_num, vbs_list[i],
)
o = o_pad[:, :, non_pad_q_list[i], :]
outs.append(o.squeeze(0).transpose(0, 1))
return torch.cat(outs, dim=0)
def _run_varlen_test(
seq_configs: list,
h: int = 8,
d: int = 64,
topk: int = 2,
atol: float = 0.05,
rtol: float = 0.02,
):
"""Core test: compare varlen vs per-sequence reference.
seq_configs: list of (num_q_blocks, num_kv_blocks) per sequence.
"""
device = "cuda"
num_seqs = len(seq_configs)
q_list = []
k_list = []
v_list = []
block_masks = []
vbs_list = []
q_vbs_list = []
non_pad_q_list = []
non_pad_kv_list = []
q_nblocks_list = []
kv_nblocks_list = []
q2k_idx_list = []
q2k_num_list = []
q_vbs_for_varlen = []
cu_q = [0]
cu_kv = [0]
for nq, nkv in seq_configs:
vbs_kv = generate_variable_block_sizes(nkv, device=device)
vbs_q = generate_variable_block_sizes(nq, device=device)
sq = int(vbs_q.sum().item())
skv = int(vbs_kv.sum().item())
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
npq = get_non_pad_index(vbs_q, nq, BLOCK_M)
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
q_list.append(q)
k_list.append(k)
v_list.append(v)
block_masks.append(mask)
vbs_list.append(vbs_kv)
q_vbs_list.append(vbs_q)
non_pad_q_list.append(npq)
non_pad_kv_list.append(npkv)
q_nblocks_list.append(nq)
kv_nblocks_list.append(nkv)
q2k_idx_list.append(q2k_idx)
q2k_num_list.append(q2k_num)
q_vbs_for_varlen.append(vbs_q)
cu_q.append(cu_q[-1] + sq)
cu_kv.append(cu_kv[-1] + skv)
ref_out = _reference_per_sequence(
q_list, k_list, v_list,
block_masks, vbs_list,
non_pad_q_list, non_pad_kv_list,
q_nblocks_list, kv_nblocks_list,
)
q_packed = torch.cat(
[qi.squeeze(0).transpose(0, 1) for qi in q_list], dim=0,
)
k_packed = torch.cat(
[ki.squeeze(0).transpose(0, 1) for ki in k_list], dim=0,
)
v_packed = torch.cat(
[vi.squeeze(0).transpose(0, 1) for vi in v_list], dim=0,
)
cu_seqlens_q = torch.tensor(cu_q, dtype=torch.int32, device=device)
cu_seqlens_kv = torch.tensor(cu_kv, dtype=torch.int32, device=device)
varlen_out = block_sparse_attn_varlen(
q_packed, k_packed, v_packed,
cu_seqlens_q, cu_seqlens_kv,
q2k_idx_list, q2k_num_list,
vbs_list,
q_variable_block_sizes_list=q_vbs_for_varlen,
)
max_abs = (ref_out - varlen_out).abs().max().item()
mean_abs = ref_out.abs().mean().item()
max_rel = max_abs / (mean_abs + 1e-8)
print(f" seqs={[c for c in seq_configs]}, max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
assert max_rel < rtol, f"max relative error {max_rel:.4e} exceeds threshold {rtol}"
@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required")
class TestVSAVarlen:
def test_equal_length(self):
"""Two sequences with same number of blocks."""
_run_varlen_test([(4, 4), (4, 4)], h=8, d=64)
def test_different_lengths(self):
"""Three sequences with different block counts."""
_run_varlen_test([(2, 3), (5, 4), (3, 6)], h=8, d=64)
def test_single_sequence(self):
"""Degenerate case: single sequence should match non-varlen path."""
_run_varlen_test([(8, 8)], h=8, d=64)
def test_many_heads(self):
"""More heads to stress the packing logic."""
_run_varlen_test([(3, 4), (5, 3)], h=16, d=128)
def test_many_sequences(self):
"""Stress test: 8 sequences with varying block counts."""
configs = [(i + 2, i + 3) for i in range(8)]
_run_varlen_test(configs, h=8, d=64)
def test_topk_equals_num_blocks(self):
"""Edge: topk covers all KV blocks (dense attention)."""
_run_varlen_test([(3, 3), (4, 4)], h=8, d=64, topk=8)
def test_single_block_per_sequence(self):
"""Minimal: each sequence has exactly 1 Q block and 1 KV block."""
_run_varlen_test([(1, 1), (1, 1), (1, 1)], h=8, d=64, topk=1)
def test_asymmetric_q_kv(self):
"""Q and KV have very different block counts."""
_run_varlen_test([(1, 8), (8, 1)], h=8, d=64, topk=1)
def test_without_q_vbs(self):
"""Test the default path where q_variable_block_sizes_list is None.
Uses full block_size=64 for Q blocks so the None path is valid.
"""
device = "cuda"
h, d, topk = 4, 64, 2
nq, nkv = 3, 4
vbs_kv = generate_variable_block_sizes(nkv, device=device)
sq = nq * BLOCK_M
skv = int(vbs_kv.sum().item())
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
k_pad = vsa_pad(k, npkv, nkv, BLOCK_M)
v_pad = vsa_pad(v, npkv, nkv, BLOCK_M)
ref_out, _ = block_sparse_attn_from_indices(q, k_pad, v_pad, q2k_idx, q2k_num, vbs_kv)
ref_flat = ref_out.squeeze(0).transpose(0, 1)
q_flat = q.squeeze(0).transpose(0, 1)
k_flat = k.squeeze(0).transpose(0, 1)
v_flat = v.squeeze(0).transpose(0, 1)
cu_q = torch.tensor([0, sq], dtype=torch.int32, device=device)
cu_kv = torch.tensor([0, skv], dtype=torch.int32, device=device)
varlen_out = block_sparse_attn_varlen(
q_flat, k_flat, v_flat,
cu_q, cu_kv,
[q2k_idx], [q2k_num], [vbs_kv],
)
max_abs = (ref_flat - varlen_out).abs().max().item()
mean_abs = ref_flat.abs().mean().item()
max_rel = max_abs / (mean_abs + 1e-8)
print(f" without_q_vbs: max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
assert max_rel < 0.02, f"max relative error {max_rel:.4e} exceeds threshold"
def _run_varlen_backward_test(
seq_configs: list,
h: int = 8,
d: int = 64,
topk: int = 2,
grad_rtol: float = 0.05,
):
"""Backward correctness: compare dQ/dK/dV from varlen vs per-sequence reference.
Both paths use the same underlying block_sparse_attn_from_indices kernel
(which has registered autograd). The varlen wrapper's scatter/gather must
correctly propagate gradients through PyTorch's in-place slice assignment.
"""
device = "cuda"
num_seqs = len(seq_configs)
q_list = []
k_list = []
v_list = []
block_masks = []
vbs_list = []
q_vbs_list = []
non_pad_q_list = []
non_pad_kv_list = []
q_nblocks_list = []
kv_nblocks_list = []
q2k_idx_list = []
q2k_num_list = []
cu_q = [0]
cu_kv = [0]
for nq, nkv in seq_configs:
vbs_kv = generate_variable_block_sizes(nkv, device=device)
vbs_q = generate_variable_block_sizes(nq, device=device)
sq = int(vbs_q.sum().item())
skv = int(vbs_kv.sum().item())
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
npq = get_non_pad_index(vbs_q, nq, BLOCK_M)
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
q_list.append(q)
k_list.append(k)
v_list.append(v)
block_masks.append(mask)
vbs_list.append(vbs_kv)
q_vbs_list.append(vbs_q)
non_pad_q_list.append(npq)
non_pad_kv_list.append(npkv)
q_nblocks_list.append(nq)
kv_nblocks_list.append(nkv)
q2k_idx_list.append(q2k_idx)
q2k_num_list.append(q2k_num)
cu_q.append(cu_q[-1] + sq)
cu_kv.append(cu_kv[-1] + skv)
# --- Reference: per-sequence backward ---
ref_q_grads = []
ref_k_grads = []
ref_v_grads = []
ref_outs = []
for i in range(num_seqs):
qi = q_list[i].detach().requires_grad_(True)
ki = k_list[i].detach().requires_grad_(True)
vi = v_list[i].detach().requires_grad_(True)
q_pad = vsa_pad(qi, non_pad_q_list[i], q_nblocks_list[i], BLOCK_M)
k_pad = vsa_pad(ki, non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
v_pad = vsa_pad(vi, non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
q2k_idx, q2k_num = _map_to_index(block_masks[i].unsqueeze(0))
o_pad, _ = block_sparse_attn_from_indices(
q_pad, k_pad, v_pad, q2k_idx, q2k_num, vbs_list[i],
)
o = o_pad[:, :, non_pad_q_list[i], :]
o_flat = o.squeeze(0).transpose(0, 1)
ref_outs.append(o_flat)
dO = torch.ones_like(o_flat)
o_flat.backward(dO)
ref_q_grads.append(qi.grad.squeeze(0).transpose(0, 1))
ref_k_grads.append(ki.grad.squeeze(0).transpose(0, 1))
ref_v_grads.append(vi.grad.squeeze(0).transpose(0, 1))
ref_dq = torch.cat(ref_q_grads, dim=0)
ref_dk = torch.cat(ref_k_grads, dim=0)
ref_dv = torch.cat(ref_v_grads, dim=0)
# --- Varlen backward ---
q_packed = torch.cat(
[qi.squeeze(0).transpose(0, 1) for qi in q_list], dim=0,
).detach().requires_grad_(True)
k_packed = torch.cat(
[ki.squeeze(0).transpose(0, 1) for ki in k_list], dim=0,
).detach().requires_grad_(True)
v_packed = torch.cat(
[vi.squeeze(0).transpose(0, 1) for vi in v_list], dim=0,
).detach().requires_grad_(True)
cu_seqlens_q = torch.tensor(cu_q, dtype=torch.int32, device=device)
cu_seqlens_kv = torch.tensor(cu_kv, dtype=torch.int32, device=device)
varlen_out = block_sparse_attn_varlen(
q_packed, k_packed, v_packed,
cu_seqlens_q, cu_seqlens_kv,
q2k_idx_list, q2k_num_list,
vbs_list,
q_variable_block_sizes_list=q_vbs_list,
)
dO = torch.ones_like(varlen_out)
varlen_out.backward(dO)
varlen_dq = q_packed.grad
varlen_dk = k_packed.grad
varlen_dv = v_packed.grad
for name, ref, actual in [
("dQ", ref_dq, varlen_dq),
("dK", ref_dk, varlen_dk),
("dV", ref_dv, varlen_dv),
]:
assert actual is not None, f"{name}: gradient is None (autograd chain broken)"
max_abs = (ref - actual).abs().max().item()
mean_abs = ref.abs().mean().item()
max_rel = max_abs / (mean_abs + 1e-8)
print(f" {name}: max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
assert max_rel < grad_rtol, (
f"{name}: max relative error {max_rel:.4e} exceeds threshold {grad_rtol}"
)
@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required")
class TestVSAVarlenBackward:
def test_backward_equal_length(self):
"""Backward: two sequences with same number of blocks."""
_run_varlen_backward_test([(4, 4), (4, 4)], h=8, d=64)
def test_backward_different_lengths(self):
"""Backward: three sequences with different block counts."""
_run_varlen_backward_test([(2, 3), (5, 4), (3, 6)], h=8, d=64)
def test_backward_single_sequence(self):
"""Backward: single sequence should match non-varlen gradient path."""
_run_varlen_backward_test([(8, 8)], h=8, d=64)
def test_backward_many_heads(self):
"""Backward: more heads to stress gradient routing."""
_run_varlen_backward_test([(3, 4), (5, 3)], h=16, d=128)
def test_backward_asymmetric_q_kv(self):
"""Backward: Q and KV have very different block counts."""
_run_varlen_backward_test([(1, 8), (8, 1)], h=8, d=64, topk=1)
def test_backward_grad_nonzero(self):
"""Smoke test: gradients are non-zero (autograd chain is connected)."""
device = "cuda"
h, d, topk = 4, 64, 2
nq, nkv = 3, 4
vbs_kv = generate_variable_block_sizes(nkv, device=device)
vbs_q = generate_variable_block_sizes(nq, device=device)
sq = int(vbs_q.sum().item())
skv = int(vbs_kv.sum().item())
q = torch.randn(sq, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
k = torch.randn(skv, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
v = torch.randn(skv, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
cu_q = torch.tensor([0, sq], dtype=torch.int32, device=device)
cu_kv = torch.tensor([0, skv], dtype=torch.int32, device=device)
out = block_sparse_attn_varlen(
q, k, v,
cu_q, cu_kv,
[q2k_idx], [q2k_num], [vbs_kv],
q_variable_block_sizes_list=[vbs_q],
)
loss = out.sum()
loss.backward()
assert q.grad is not None, "q.grad is None"
assert k.grad is not None, "k.grad is None"
assert v.grad is not None, "v.grad is None"
assert q.grad.abs().sum().item() > 0, "q.grad is all zeros"
assert k.grad.abs().sum().item() > 0, "k.grad is all zeros"
assert v.grad.abs().sum().item() > 0, "v.grad is all zeros"
if __name__ == "__main__":
pytest.main([__file__, "-v", "-s"])
-5
View File
@@ -29,10 +29,6 @@ class SamplingParam:
# Video inputs
video_path: str | None = None
# Optional pre-generated diffusion latents. Used by parity/debug harnesses
# and advanced callers that need deterministic latent reuse.
latents: Any | None = None
# Action control inputs (Matrix-Game)
mouse_cond: Any | None = None # Shape: (B, T, 2)
keyboard_cond: Any | None = None # Shape: (B, T, K)
@@ -68,7 +64,6 @@ class SamplingParam:
# Text inputs
prompt: str | list[str] | None = None
negative_prompt: str = "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
max_sequence_length: int | None = None
prompt_path: str | None = None
output_path: str = "outputs/"
output_video_name: str | None = None
@@ -49,10 +49,6 @@ def _get_attn_qat_train_attention() -> Callable[..., torch.Tensor] | None:
return _attn_qat_train_attention
def is_attn_qat_train_available() -> bool:
return _get_attn_qat_train_attention() is not None
def attn_qat_train(q_BLHD: torch.Tensor,
k_BLHD: torch.Tensor,
v_BLHD: torch.Tensor,
@@ -150,11 +150,14 @@ class VideoSparseAttentionMetadata(AttentionMetadata):
# in postprocess_output(). Avoids materializing the intermediate
# ``[B, len(non_pad_index), H, D]`` tensor on every layer.
untile_combined_index: torch.LongTensor
# Per-step shared padded buffer used by tile(). Inference can reuse this
# across VSA layers, but training disables it so activation checkpointing
# can release the large tiled QKVG scratch tensor after each attention call.
# Per-step shared padded buffer used by tile(). Lazily populated on
# the first layer's call and reused by every subsequent VSA layer in
# the same denoising step. Scoping to metadata (not class/instance)
# makes the reuse thread-safe across concurrent requests and keeps
# the "pad positions are zero" invariant trivially true (the buffer
# is freshly zeroed alongside ``non_pad_index`` so the index set
# cannot drift between calls).
tile_buf: torch.Tensor | None = None
cache_tile_buf: bool = True
class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
@@ -172,7 +175,6 @@ class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
patch_size: tuple[int, int, int],
VSA_sparsity: float,
device: torch.device,
cache_tile_buf: bool = True,
**kwargs: dict[str, Any],
) -> VideoSparseAttentionMetadata:
patch_size = patch_size
@@ -199,8 +201,7 @@ class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
reverse_tile_partition_indices=reverse_tile_partition_indices,
variable_block_sizes=variable_block_sizes,
non_pad_index=non_pad_index,
untile_combined_index=untile_combined_index,
cache_tile_buf=cache_tile_buf)
untile_combined_index=untile_combined_index)
class VideoSparseAttentionImpl(AttentionImpl):
@@ -236,11 +237,6 @@ class VideoSparseAttentionImpl(AttentionImpl):
w_padded_size = num_tiles[2] * VSA_TILE_SIZE[2]
target_shape = (x.shape[0], t_padded_size * h_padded_size * w_padded_size, x.shape[-2], x.shape[-1])
if not attn_metadata.cache_tile_buf:
buf = torch.zeros(target_shape, device=x.device, dtype=x.dtype)
buf[:, attn_metadata.non_pad_index] = x[:, attn_metadata.tile_partition_indices]
return buf
# Reuse the per-step buffer stashed on metadata (lazily allocated
# on the first VSA layer's call within a denoising step). Pad
# positions are zero from the initial torch.zeros and never
+1 -2
View File
@@ -1,6 +1,5 @@
from fastvideo.configs.models.dits.cosmos import CosmosVideoConfig
from fastvideo.configs.models.dits.cosmos2_5 import Cosmos25VideoConfig
from fastvideo.configs.models.dits.flux_2 import Flux2Config
from fastvideo.configs.models.dits.hunyuangamecraft import HunyuanGameCraftConfig
from fastvideo.configs.models.dits.hunyuanvideo import HunyuanVideoConfig
from fastvideo.configs.models.dits.hunyuanvideo15 import HunyuanVideo15Config
@@ -15,5 +14,5 @@ from fastvideo.configs.models.dits.kandinsky5 import Kandinsky5VideoConfig
__all__ = [
"HunyuanVideoConfig", "HunyuanVideo15Config", "HunyuanGameCraftConfig", "WanVideoConfig", "CosmosVideoConfig",
"Cosmos25VideoConfig", "LongCatVideoConfig", "LTX2VideoConfig", "HYWorldConfig", "Kandinsky5VideoConfig",
"MagiHumanVideoConfig", "StableAudioConfig", "Flux2Config"
"MagiHumanVideoConfig", "StableAudioConfig"
]
+1 -8
View File
@@ -14,19 +14,12 @@ class DiTArchConfig(ArchConfig):
param_names_mapping: dict = field(default_factory=dict)
reverse_param_names_mapping: dict = field(default_factory=dict)
lora_param_names_mapping: dict = field(default_factory=dict)
# When True, the denoising stage casts text/prompt embeddings to the DiT's
# working dtype before the diffusion loop. Flux2 requires this (BFL casts ctx
# to bf16 before denoising); models with fp32 text encoders (Wan, Hunyuan15,
# SD3.5) leave it False to preserve full-precision embeddings.
cast_prompt_embeds_to_dit_dtype: bool = False
_supported_attention_backends: tuple[AttentionBackendEnum,
...] = (AttentionBackendEnum.SAGE_ATTN, AttentionBackendEnum.FLASH_ATTN,
AttentionBackendEnum.TORCH_SDPA,
AttentionBackendEnum.VIDEO_SPARSE_ATTN,
AttentionBackendEnum.VMOBA_ATTN, AttentionBackendEnum.SAGE_ATTN_THREE,
AttentionBackendEnum.ATTN_QAT_INFER,
AttentionBackendEnum.ATTN_QAT_TRAIN, AttentionBackendEnum.SLA_ATTN,
AttentionBackendEnum.SAGE_SLA_ATTN)
AttentionBackendEnum.SLA_ATTN, AttentionBackendEnum.SAGE_SLA_ATTN)
hidden_size: int = 0
num_attention_heads: int = 0
-77
View File
@@ -1,77 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# Copied and adapted from: https://github.com/sglang-ai/sglang
from dataclasses import dataclass, field
from fastvideo.configs.models.dits.base import DiTArchConfig, DiTConfig
from fastvideo.logger import init_logger
logger = init_logger(__name__)
@dataclass
class Flux2ArchConfig(DiTArchConfig):
"""Architecture configuration for Flux2 transformer model."""
cast_prompt_embeds_to_dit_dtype: bool = True
# Flux2-specific architecture parameters
patch_size: int = 1
in_channels: int = 64
out_channels: int | None = None
num_layers: int = 19 # Number of double-stream transformer blocks
num_single_layers: int = 38 # Number of single-stream transformer blocks
attention_head_dim: int = 128
num_attention_heads: int = 24
joint_attention_dim: int = 4096 # Dimension for text encoder output
timestep_guidance_channels: int = 256 # Dimension for timestep embedding
mlp_ratio: float = 3.0
axes_dims_rope: tuple[int, ...] = (32, 32, 32, 32) # RoPE dimensions per axis (match diffusers Flux2)
rope_theta: int = 2000 # Base frequency for RoPE (match diffusers Flux2)
eps: float = 1e-6
guidance_embeds: bool = True # Whether to use guidance embeddings
# When True, compute SwiGLU in fp32 inside ``ff_context`` only (bf16 noise mitigation).
ff_context_swiglu_fp32: bool = False
# Parameter name mapping for loading HuggingFace checkpoints
param_names_mapping: dict = field(default_factory=lambda: {
r"transformer\.(\w*)\.(.*)$": r"\1.\2",
})
def __post_init__(self) -> None:
super().__post_init__()
self.out_channels = self.out_channels or self.in_channels
self.hidden_size = self.num_attention_heads * self.attention_head_dim
self.num_channels_latents = self.out_channels
def update_from_weight_keys(self, all_keys: set[str]) -> None:
"""Infer num_layers and num_single_layers from checkpoint weight keys so the model is built with the same number of blocks as the weights."""
if not all_keys:
return
num_layers = 0
num_single_layers = 0
for k in all_keys:
if "single_transformer_blocks." not in k and "transformer_blocks." in k:
parts = k.split("transformer_blocks.")[-1].split(".")
if parts[0].isdigit():
num_layers = max(num_layers, int(parts[0]) + 1)
if "single_transformer_blocks." in k:
parts = k.split("single_transformer_blocks.")[-1].split(".")
if parts[0].isdigit():
num_single_layers = max(num_single_layers, int(parts[0]) + 1)
if num_layers > 0:
self.num_layers = num_layers
logger.info("Inferred num_layers=%s from checkpoint keys", num_layers)
if num_single_layers > 0:
self.num_single_layers = num_single_layers
logger.info("Inferred num_single_layers=%s from checkpoint keys", num_single_layers)
if num_layers > 0 or num_single_layers > 0:
self.__post_init__()
@dataclass
class Flux2Config(DiTConfig):
"""Configuration for Flux2 transformer model."""
arch_config: DiTArchConfig = field(default_factory=Flux2ArchConfig)
prefix: str = "Flux"
@@ -7,8 +7,6 @@ from fastvideo.configs.models.encoders.qwen2_5 import Qwen2_5_VLConfig
from fastvideo.configs.models.encoders.siglip import SiglipVisionConfig
from fastvideo.configs.models.encoders.reason1 import Reason1ArchConfig, Reason1Config
from fastvideo.configs.models.encoders.gemma import LTX2GemmaConfig
from fastvideo.configs.models.encoders.mistral3 import Mistral3TextConfig
from fastvideo.configs.models.encoders.qwen3 import Qwen3TextConfig
from fastvideo.configs.models.encoders.stable_audio_conditioner import (StableAudioConditionerArchConfig,
StableAudioConditionerConfig)
from fastvideo.configs.models.encoders.t5gemma import T5GemmaEncoderConfig
@@ -17,5 +15,5 @@ __all__ = [
"EncoderConfig", "TextEncoderConfig", "ImageEncoderConfig", "BaseEncoderOutput", "CLIPTextConfig",
"CLIPVisionConfig", "WAN2_1ControlCLIPVisionConfig", "LlamaConfig", "T5Config", "T5LargeConfig", "Qwen2_5_VLConfig",
"Reason1ArchConfig", "Reason1Config", "LTX2GemmaConfig", "SiglipVisionConfig", "StableAudioConditionerArchConfig",
"StableAudioConditionerConfig", "T5GemmaEncoderConfig", "Qwen3TextConfig", "Mistral3TextConfig"
"StableAudioConditionerConfig", "T5GemmaEncoderConfig"
]
@@ -36,11 +36,6 @@ class TextEncoderArchConfig(EncoderArchConfig):
default_factory=list) # mapping from huggingface weight names to custom names
tokenizer_kwargs: dict[str, Any] = field(default_factory=dict)
_fsdp_shard_conditions: list = field(default_factory=lambda: [])
# When True, the tokenizer loader prefers AutoProcessor over AutoTokenizer
# for encoders whose tokenizer dir ships a processor_config.json (e.g. Flux2
# full's Mistral3 multimodal processor). Default False keeps every existing
# encoder on the historical AutoTokenizer path.
require_processor: bool = False
def __post_init__(self) -> None:
self.tokenizer_kwargs = {
@@ -1,38 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""Mistral3 text encoder configuration for full Flux2."""
from dataclasses import dataclass, field
from fastvideo.configs.models.encoders.base import (
TextEncoderArchConfig,
TextEncoderConfig,
)
@dataclass
class Mistral3TextArchConfig(TextEncoderArchConfig):
"""Architecture config for the Mistral3 text encoder used by full Flux2."""
architectures: list[str] = field(default_factory=lambda: ["Mistral3ForConditionalGeneration"])
hidden_size: int = 5120
num_hidden_layers: int = 40
text_len: int = 512
output_hidden_states: bool = True
# Mistral3 (full Flux2) ships a multimodal processor; load via AutoProcessor.
require_processor: bool = True
def __post_init__(self) -> None:
self.tokenizer_kwargs = {
"padding": "max_length",
"truncation": True,
"max_length": self.text_len,
"return_tensors": "pt",
}
@dataclass
class Mistral3TextConfig(TextEncoderConfig):
"""Top-level config for the Mistral3 full Flux2 text encoder."""
arch_config: TextEncoderArchConfig = field(default_factory=Mistral3TextArchConfig)
prefix: str = "mistral3"
is_chat_model: bool = True
@@ -1,82 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# Ported from SGLang: python/sglang/multimodal_gen/configs/models/encoders/qwen3.py
"""Qwen3 text encoder configuration for FastVideo diffusion models (e.g. Flux2 Klein)."""
from dataclasses import dataclass, field
from typing import Any
from fastvideo.configs.models.encoders.base import (
TextEncoderArchConfig,
TextEncoderConfig,
)
def _is_transformer_layer(n: str, m: Any) -> bool:
return "layers" in n and str.isdigit(n.split(".")[-1])
def _is_embeddings(n: str, m: Any) -> bool:
return n.endswith("embed_tokens")
def _is_final_norm(n: str, m: Any) -> bool:
return n.endswith("norm")
@dataclass
class Qwen3TextArchConfig(TextEncoderArchConfig):
"""Architecture config for Qwen3 text encoder.
Qwen3 is similar to LLaMA but with QK-Norm (RMSNorm on Q and K before attention).
Used by Flux2 Klein.
"""
vocab_size: int = 151936
hidden_size: int = 2560
intermediate_size: int = 9728
num_hidden_layers: int = 36
num_attention_heads: int = 32
num_key_value_heads: int = 8
hidden_act: str = "silu"
max_position_embeddings: int = 40960
initializer_range: float = 0.02
rms_norm_eps: float = 1e-6
use_cache: bool = True
pad_token_id: int = 151643
bos_token_id: int = 151643
eos_token_id: int = 151645
tie_word_embeddings: bool = True
rope_theta: float = 1000000.0
rope_scaling: dict | None = None
attention_bias: bool = False
attention_dropout: float = 0.0
mlp_bias: bool = False
head_dim: int = 128
text_len: int = 512
output_hidden_states: bool = True # Klein needs hidden states from layers 9, 18, 27
stacked_params_mapping: list[tuple[str, str, str | int]] = field(default_factory=lambda: [
(".qkv_proj", ".q_proj", "q"),
(".qkv_proj", ".k_proj", "k"),
(".qkv_proj", ".v_proj", "v"),
(".gate_up_proj", ".gate_proj", 0),
(".gate_up_proj", ".up_proj", 1),
])
_fsdp_shard_conditions: list = field(
default_factory=lambda: [_is_transformer_layer, _is_embeddings, _is_final_norm])
def __post_init__(self) -> None:
self.tokenizer_kwargs = {
"padding": "max_length",
"truncation": True,
"max_length": self.text_len,
"return_tensors": "pt",
}
@dataclass
class Qwen3TextConfig(TextEncoderConfig):
"""Top-level config for Qwen3 text encoder."""
arch_config: TextEncoderArchConfig = field(default_factory=Qwen3TextArchConfig)
prefix: str = "qwen3"
is_chat_model: bool = True
@@ -6,7 +6,6 @@ from fastvideo.configs.models.vaes.hunyuanvae import HunyuanVAEConfig
from fastvideo.configs.models.vaes.hunyuan15vae import Hunyuan15VAEConfig
from fastvideo.configs.models.vaes.ltx2vae import LTX2VAEConfig
from fastvideo.configs.models.vaes.oobleck import OobleckVAEArchConfig, OobleckVAEConfig
from fastvideo.configs.models.vaes.flux2vae import Flux2VAEConfig
from fastvideo.configs.models.vaes.wanvae import WanVAEConfig
__all__ = [
@@ -20,5 +19,4 @@ __all__ = [
"LTX2VAEConfig",
"OobleckVAEArchConfig",
"OobleckVAEConfig",
"Flux2VAEConfig",
]
-58
View File
@@ -1,58 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# Copied and adapted from: https://github.com/sglang-ai/sglang
from dataclasses import dataclass, field
from fastvideo.configs.models.vaes.base import VAEArchConfig, VAEConfig
@dataclass
class Flux2VAEArchConfig(VAEArchConfig):
"""Architecture configuration for Flux2 VAE model."""
# Flux2 VAE-specific architecture parameters
in_channels: int = 3
out_channels: int = 3
down_block_types: tuple[str, ...] = (
"DownEncoderBlock2D",
"DownEncoderBlock2D",
"DownEncoderBlock2D",
"AttnDownEncoderBlock2D",
)
up_block_types: tuple[str, ...] = (
"AttnUpDecoderBlock2D",
"UpDecoderBlock2D",
"UpDecoderBlock2D",
"UpDecoderBlock2D",
)
block_out_channels: tuple[int, ...] = (128, 256, 512, 512)
layers_per_block: int = 2
act_fn: str = "silu"
latent_channels: int = 16
norm_num_groups: int = 32
sample_size: int = 512
force_upcast: bool = False
use_quant_conv: bool = True
use_post_quant_conv: bool = True
mid_block_add_attention: bool = True
batch_norm_eps: float = 1e-5
batch_norm_momentum: float = 0.1
patch_size: tuple[int, int] = (1, 1)
# Latent scaling for decode: avoid division-by-zero; match Flux/Flux2 convention (e.g. 0.13025)
scaling_factor: float = 0.13025
# Spatial compression (for images, this is typically 8)
spatial_compression_ratio: int = 8
temporal_compression_ratio: int = 1 # Images don't have temporal dimension
@dataclass
class Flux2VAEConfig(VAEConfig):
"""Configuration for Flux2 VAE model."""
arch_config: Flux2VAEArchConfig = field(default_factory=Flux2VAEArchConfig)
# Flux2 is an image model, so disable temporal tiling
use_tiling: bool = False
use_temporal_tiling: bool = False
use_parallel_tiling: bool = False
+4 -4
View File
@@ -9,12 +9,12 @@ from fastvideo.configs.pipelines.matrixgame2 import MatrixGame2I2V480PConfig
from fastvideo.configs.pipelines.matrixgame3 import MatrixGame3I2V720PConfig
from fastvideo.pipelines.basic.ltx2.pipeline_configs import LTX2T2VConfig
from fastvideo.registry import get_pipeline_config_cls_from_name
from fastvideo.configs.pipelines.wan import (LucyEditDevConfig, SelfForcingWanT2V480PConfig, WanI2V480PConfig,
WanI2V720PConfig, WanT2V480PConfig, WanT2V720PConfig)
from fastvideo.configs.pipelines.wan import (SelfForcingWanT2V480PConfig, WanI2V480PConfig, WanI2V720PConfig,
WanT2V480PConfig, WanT2V720PConfig)
__all__ = [
"HunyuanConfig", "FastHunyuanConfig", "HunyuanGameCraftPipelineConfig", "PipelineConfig", "Hunyuan15T2V480PConfig",
"Hunyuan15T2V720PConfig", "WanT2V480PConfig", "WanI2V480PConfig", "WanT2V720PConfig", "WanI2V720PConfig",
"SelfForcingWanT2V480PConfig", "LucyEditDevConfig", "CosmosConfig", "Cosmos25Config", "LTX2T2VConfig",
"HYWorldConfig", "MatrixGame2I2V480PConfig", "MatrixGame3I2V720PConfig", "get_pipeline_config_cls_from_name"
"SelfForcingWanT2V480PConfig", "CosmosConfig", "Cosmos25Config", "LTX2T2VConfig", "HYWorldConfig",
"MatrixGame2I2V480PConfig", "MatrixGame3I2V720PConfig", "get_pipeline_config_cls_from_name"
]
+1 -7
View File
@@ -35,11 +35,6 @@ class PipelineConfig:
flow_shift: float | None = None
flow_shift_sr: float | None = None
disable_autocast: bool = False
# When True, the scheduler's Euler update runs in fp32 outside the autocast
# block (Diffusers-style; avoids BF16 drift over multiple steps). Flux2 sets
# this True for reference parity; other models keep the legacy in-autocast
# behavior to preserve existing SSIM references.
scheduler_step_in_fp32: bool = False
is_causal: bool = False
# Model configuration
@@ -69,9 +64,8 @@ class PipelineConfig:
# DMD parameters
dmd_denoising_steps: list[int] | None = field(default=None)
# Wan2.2 task modifiers
# Wan2.2 TI2V parameters
ti2v_task: bool = False
lucy_edit_task: bool = False
boundary_ratio: float | None = None
# Compilation
-86
View File
@@ -1,86 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# Copied and adapted from: https://github.com/sglang-ai/sglang
from collections.abc import Callable
from dataclasses import dataclass, field
import torch
from fastvideo.configs.models import DiTConfig, EncoderConfig, VAEConfig
from fastvideo.configs.models.dits.flux_2 import Flux2Config
from fastvideo.configs.models.encoders import BaseEncoderOutput
from fastvideo.configs.models.encoders.base import EncoderArchConfig
from fastvideo.configs.models.encoders.mistral3 import Mistral3TextConfig
from fastvideo.configs.models.encoders.qwen3 import Qwen3TextConfig
from fastvideo.configs.models.vaes.flux2vae import Flux2VAEConfig
from fastvideo.configs.pipelines.base import PipelineConfig, preprocess_text
@dataclass
class Flux2PipelineConfig(PipelineConfig):
"""Configuration for Flux2 image generation pipeline."""
# Flux2-specific parameters
embedded_cfg_scale: float | None = 4.0
scheduler_step_in_fp32: bool = True
flux2_text_encoder_type: str = "mistral3"
text_encoder_out_layers: tuple[int, ...] = (10, 20, 30)
# DiT configuration
dit_config: DiTConfig = field(default_factory=Flux2Config)
dit_precision: str = "bf16"
# VAE configuration
vae_config: VAEConfig = field(default_factory=Flux2VAEConfig)
vae_precision: str = "fp32"
vae_tiling: bool = False # Flux2 is image model, disable tiling by default
vae_sp: bool = False
# Text encoder configuration (full Flux2 uses Mistral3)
text_encoder_configs: tuple[EncoderConfig, ...] = field(default_factory=lambda: (Mistral3TextConfig(), ))
text_encoder_precisions: tuple[str, ...] = field(default_factory=lambda: ("bf16", ))
# Default postprocess function (can be overridden)
@staticmethod
def default_postprocess_text(outputs: BaseEncoderOutput) -> torch.Tensor:
"""Default text postprocessing for Flux2."""
return outputs.last_hidden_state
postprocess_text_funcs: tuple[Callable[[BaseEncoderOutput], torch.Tensor],
...] = field(default_factory=lambda: (Flux2PipelineConfig.default_postprocess_text, ))
def flux2_klein_postprocess_text(outputs: BaseEncoderOutput) -> torch.Tensor:
"""Klein postprocess: hidden states from layers 9, 18, 27 (Qwen3)."""
hidden_states_layers: list[int] = [9, 18, 27]
if outputs.hidden_states is None:
raise ValueError("Flux2 Klein requires output_hidden_states=True from text encoder")
out = torch.stack([outputs.hidden_states[k] for k in hidden_states_layers], dim=1)
batch_size, num_channels, seq_len, hidden_dim = out.shape
prompt_embeds = out.permute(0, 2, 1, 3).reshape(batch_size, seq_len, num_channels * hidden_dim)
return prompt_embeds
@dataclass
class Flux2KleinEncoderArchConfig(EncoderArchConfig):
"""Encoder arch config for Flux2 Klein (Qwen3); needs hidden states for layers 9, 18, 27."""
output_hidden_states: bool = True
@dataclass
class Flux2KleinTextEncoderConfig(EncoderConfig):
"""Text encoder config for Flux2 Klein (Qwen3)."""
arch_config: EncoderArchConfig = field(default_factory=Flux2KleinEncoderArchConfig)
@dataclass
class Flux2KleinPipelineConfig(Flux2PipelineConfig):
"""Configuration for Flux2 Klein (distilled, 4-step, no guidance)."""
embedded_cfg_scale: float | None = None # Klein distilled: no guidance embedding (matches Diffusers)
scheduler_step_in_fp32: bool = True
flux2_text_encoder_type: str = "qwen3"
text_encoder_out_layers: tuple[int, ...] = (9, 18, 27)
text_encoder_configs: tuple[EncoderConfig, ...] = field(default_factory=lambda: (Qwen3TextConfig(), ))
text_encoder_precisions: tuple[str, ...] = field(default_factory=lambda: ("bf16", ))
preprocess_text_funcs: tuple[Callable[[str], str], ...] = field(default_factory=lambda: (preprocess_text, ))
postprocess_text_funcs: tuple[Callable[[BaseEncoderOutput], torch.Tensor],
...] = field(default_factory=lambda: (flux2_klein_postprocess_text, ))
-138
View File
@@ -6,11 +6,9 @@ import torch
from fastvideo.configs.models import DiTConfig, EncoderConfig, VAEConfig
from fastvideo.configs.models.dits import WanVideoConfig
from fastvideo.configs.models.dits.wanvideo import WanVideoArchConfig
from fastvideo.configs.models.encoders import (BaseEncoderOutput, CLIPVisionConfig, T5Config,
WAN2_1ControlCLIPVisionConfig)
from fastvideo.configs.models.vaes import WanVAEConfig
from fastvideo.configs.models.vaes.wanvae import WanVAEArchConfig
from fastvideo.configs.pipelines.base import PipelineConfig
@@ -122,142 +120,6 @@ class Wan2_2_TI2V_5B_Config(WanT2V480PConfig):
expand_timesteps: bool = True
def __post_init__(self) -> None:
assert not (self.ti2v_task and self.lucy_edit_task)
self.vae_config.load_encoder = True
self.vae_config.load_decoder = True
self.dit_config.expand_timesteps = self.expand_timesteps
@dataclass
class LucyEditDevConfig(Wan2_2_TI2V_5B_Config):
"""Configuration for Decart Lucy Edit Dev video editing."""
dit_config: DiTConfig = field(default_factory=lambda: WanVideoConfig(arch_config=WanVideoArchConfig(
num_attention_heads=24,
in_channels=96,
out_channels=48,
ffn_dim=14336,
num_layers=30,
)))
vae_config: VAEConfig = field(default_factory=lambda: WanVAEConfig(arch_config=WanVAEArchConfig(
base_dim=160,
decoder_base_dim=256,
z_dim=48,
in_channels=12,
out_channels=12,
scale_factor_spatial=16,
patch_size=2,
is_residual=True,
clip_output=False,
latents_mean=(
-0.2289,
-0.0052,
-0.1323,
-0.2339,
-0.2799,
0.0174,
0.1838,
0.1557,
-0.1382,
0.0542,
0.2813,
0.0891,
0.1570,
-0.0098,
0.0375,
-0.1825,
-0.2246,
-0.1207,
-0.0698,
0.5109,
0.2665,
-0.2108,
-0.2158,
0.2502,
-0.2055,
-0.0322,
0.1109,
0.1567,
-0.0729,
0.0899,
-0.2799,
-0.1230,
-0.0313,
-0.1649,
0.0117,
0.0723,
-0.2839,
-0.2083,
-0.0520,
0.3748,
0.0152,
0.1957,
0.1433,
-0.2944,
0.3573,
-0.0548,
-0.1681,
-0.0667,
),
latents_std=(
0.4765,
1.0364,
0.4514,
1.1677,
0.5313,
0.4990,
0.4818,
0.5013,
0.8158,
1.0344,
0.5894,
1.0901,
0.6885,
0.6165,
0.8454,
0.4978,
0.5759,
0.3523,
0.7135,
0.6804,
0.5833,
1.4146,
0.8986,
0.5659,
0.7069,
0.5338,
0.4889,
0.4917,
0.4069,
0.4999,
0.6866,
0.4093,
0.5709,
0.6065,
0.6415,
0.4944,
0.5726,
1.2042,
0.5458,
1.6887,
0.3971,
1.0600,
0.3943,
0.5537,
0.5444,
0.4089,
0.7468,
0.7744,
),
)))
ti2v_task: bool = False
lucy_edit_task: bool = True
def __post_init__(self) -> None:
assert not (self.ti2v_task and self.lucy_edit_task)
# Lucy uses Wan2.2's enhanced 48-channel VAE latents. Denoising
# concatenates noise + video latents, matching the 96-channel
# transformer input declared above.
self.vae_config.load_encoder = True
self.vae_config.load_decoder = True
self.dit_config.expand_timesteps = self.expand_timesteps
-7
View File
@@ -517,7 +517,6 @@ class VideoGenerator:
if _ek in kwargs:
extra_overrides[_ek] = kwargs.pop(_ek)
prompt_embeds = kwargs.pop("prompt_embeds", None)
sampling_param.update(kwargs)
kwargs["_extra_overrides"] = extra_overrides
@@ -568,8 +567,6 @@ class VideoGenerator:
raise ValueError("Either prompt or prompt_txt must be provided")
output_path = self._prepare_output_path(sampling_param.output_path, prompt)
kwargs["output_path"] = output_path
if prompt_embeds is not None:
kwargs["prompt_embeds"] = prompt_embeds
return self._generate_single_video(
prompt=prompt,
sampling_param=sampling_param,
@@ -672,7 +669,6 @@ class VideoGenerator:
prompt = prompt.strip()
sampling_param = deepcopy(sampling_param)
output_path = kwargs["output_path"]
prompt_embeds = kwargs.get("prompt_embeds")
sampling_param.prompt = prompt
# Process negative prompt
if sampling_param.negative_prompt is not None:
@@ -719,9 +715,6 @@ class VideoGenerator:
n_tokens=n_tokens,
VSA_sparsity=fastvideo_args.VSA_sparsity,
)
# Allow precomputed prompt_embeds (e.g. from diffusers) to skip text encoding
if prompt_embeds is not None:
batch.prompt_embeds = (list(prompt_embeds) if isinstance(prompt_embeds, list | tuple) else [prompt_embeds])
extra_overrides = kwargs.pop("_extra_overrides", {})
for _ek, _ev in extra_overrides.items():
+2 -38
View File
@@ -2,9 +2,8 @@
In-process evaluation suite for video generations. Includes pixel
metrics (SSIM, PSNR, LPIPS), Fréchet Video Distance (FVD), optical-flow
comparisons, the full VBench suite, Physics-IQ, audio metrics, an
absolute VLM scorer (`videoscore2`), and a pairwise VLM judge
(`judge.third_person_separation`) — all behind a single registry-driven API.
comparisons, the full VBench suite, Physics-IQ, audio metrics, and a
VLM scorer behind a single registry-driven API.
## Install
@@ -160,7 +159,6 @@ fastvideo/
│ ├── audio/ # clap_score, audiobox_aesthetics, kl_divergence,
│ │ # frechet_distance, wer, desync, imagebind_score
│ ├── videoscore2/ # VideoScore-2 (Qwen2.5-VL)
│ ├── judge/ # pairwise VLM judges (third_person_separation)
│ ├── physics_iq/ # PhysicsIQ + sub-metrics
│ └── vbench/ # adapter: sys.path bootstrap + shims
│ ├── __init__.py
@@ -315,40 +313,6 @@ to control read/write behavior. The example script
`examples/inference/eval/eval_fvd.py` demonstrates the full
two-directory workflow.
## `judge.third_person_separation` — pairwise VLM judge
A **preference** metric (a judge, not an absolute score), and the suite's first
remote-API one. For each pair the judge (Gemini) sees the shared first frame and
two rollouts — a candidate and a reference model under the same control signal —
and picks the one that better separates the third-person CHARACTER (foreground)
from the BACKGROUND. The corpus score is the candidate's win-rate, excluding
ties. Set-vs-set, motion-first; it reads native mp4s, so samples carry path
strings, not decoded tensors.
```bash
uv pip install -e .[eval-judge] # opt-in: needs network + an API key
export GEMINI_API_KEY=... # or GOOGLE_API_KEY, or ~/.gemini_token
```
```python
from fastvideo.eval import create_evaluator
ev = create_evaluator(metrics=["judge.third_person_separation"], device="cpu")
result = ev.evaluate(samples=[
{"video_path": "cand/000.mp4", "reference_path": "base/000.mp4",
"image_path": "frames/000.png", "text_prompt": "W: moves forward", "action": "W"},
# ... more pairs ...
]).corpus["judge.third_person_separation"]
result.score # candidate win-rate excl. ties; result.details has the breakdown
```
Only `video_path`/`reference_path` are required; `image_path`/`text_prompt`/
`action` are optional. Verdicts are cached under `${FASTVIDEO_EVAL_CACHE}/eval/judge/`.
The judge separates best when the control yields genuine parallax (e.g.
translation); rigid whole-frame motion (e.g. pure camera rotation) is harder. To
sweep several baselines into a table, see
`examples/inference/eval/eval_third_person_separation.py`.
## Out of scope (follow-up PRs)
- **MIND** metrics. Depend on a separate `vipe` upstream submodule.
@@ -1,306 +0,0 @@
"""Pairwise VLM judge (Gemini): of two rollouts under the same control, which
better separates the third-person character (foreground) from the background;
the score is the candidate's win-rate over a reference.
"""
from __future__ import annotations
import hashlib
import json
import os
import time
from pathlib import Path
from typing import Any
from fastvideo.eval.metrics.base import BaseMetric
from fastvideo.eval.models import get_cache_dir
from fastvideo.eval.registry import register
from fastvideo.eval.types import MetricResult, Video
# Part of the on-disk cache key; bump to invalidate cached verdicts.
RUBRIC_ID = "v1"
DEFAULT_MODEL = "gemini-2.5-pro"
DEFAULT_K = 3
SYSTEM_PROMPT = ("You are a strict comparative evaluator of third-person video-game "
"rollouts. Two videos were generated by two different models from the "
"SAME first frame and the SAME control signal. You judge which video "
"better demonstrates that the model SEPARATES the third-person CHARACTER "
"(foreground) from the BACKGROUND SCENE — i.e. the control animates the "
"character as an INDEPENDENT AGENT with its own trajectory while the "
"background moves with the camera. You do NOT reward whichever video "
"merely looks cleaner, sharper, or higher-res.")
RUBRIC = """\
You are watching TWO generated video rollouts (Video 1, Video 2) played in full,
from the same first frame under the same control signal. Pick the better
third-person world-model rollout. Judge THREE things together, in this order:
(C) MOTION / ACTION EXECUTION FIRST. The rollout must actually CARRY OUT the
control signal with substantial motion (the scene/character clearly moves as
commanded). A clip that is near-static, barely drifts, or only twitches has
NOT demonstrated controllable separation — it FAILS, no matter how clean it
looks. CRITICAL: do NOT reward a clip for looking "smoother" or "more
stable" when that smoothness is really just the ABSENCE OF MOTION. Less
motion is NOT better. If one clip executes the action with clear motion and
the other is comparatively static, the MOVING one wins (unless it fails B).
(A) TEMPORAL COHERENCE — among clips that actually move, penalize GENUINE
corruption: flicker/strobing, texture boiling/crawling, geometry swimming,
the character or scene morphing/warping into mush, colors pulsing, or
progressive degradation into noise. Do NOT confuse LEGITIMATE large motion
(camera sweeping, character running, scene flowing past) with instability —
fast correct motion is GOOD, not a defect. Only true frame-to-frame
INCOHERENCE counts against a clip.
(B) FOREGROUND/BACKGROUND SEPARATION — the character stays a distinct, coherent
entity with its own trajectory while the background responds to the control;
it does not dissolve/smear into the bg, and the whole frame does not slide
as one rigid sheet.
Decision: among clips that genuinely execute the motion (C), pick the one that
is both temporally coherent (A) and shows cleaner separation (B). A static or
barely-moving clip loses to a moving one. Real motion is not instability. Do not
reward resolution or placidity. "tie" only if truly equivalent on all three."""
USER_TASK = ("Output a JSON object with four fields:\n"
" - video_1_analysis: FIRST, how much does Video 1 actually move — does it "
"execute the control with clear motion, or is it near-static / barely "
"drifting? THEN: among its motion, is there GENUINE corruption (flicker, "
"boiling, morphing into mush) as opposed to legitimate fast motion? THEN: "
"is the character a distinct coherent entity vs rigid-slide / dissolve?\n"
" - video_2_analysis: the same three checks for Video 2.\n"
" - comparison: apply (C) motion-first, then (A) genuine-coherence, then "
"(B) separation. A near-static clip loses to a moving one; legitimate large "
"motion is NOT a defect.\n"
" - winner: \"video_1\", \"video_2\", or \"tie\".")
RESPONSE_SCHEMA = {
"type": "object",
"properties": {
"video_1_analysis": {
"type": "string"
},
"video_2_analysis": {
"type": "string"
},
"comparison": {
"type": "string"
},
"winner": {
"type": "string",
"enum": ["video_1", "video_2", "tie"]
},
},
"required": ["video_1_analysis", "video_2_analysis", "comparison", "winner"],
"propertyOrdering": ["video_1_analysis", "video_2_analysis", "comparison", "winner"],
}
def _path_of(sample: dict, key: str) -> str | None:
"""Resolve a native-file path from a string key or a Video wrapper."""
p = sample.get(f"{key}_path")
if isinstance(p, str):
return p
v = sample.get(key)
if isinstance(v, Video) and isinstance(v.source, str):
return v.source
if isinstance(v, str):
return v
return None
def _resolve_api_key() -> str:
for env in ("GEMINI_API_KEY", "GOOGLE_API_KEY"):
key = os.environ.get(env)
if key:
return key.strip()
token = Path("~/.gemini_token").expanduser()
if token.is_file():
return token.read_text().strip()
raise ValueError("judge.third_person_separation needs a Gemini API key. Set "
"GEMINI_API_KEY (or GOOGLE_API_KEY), or write it to ~/.gemini_token.")
@register("judge.third_person_separation")
class ThirdPersonSeparationMetric(BaseMetric):
"""Pairwise VLM judge of third-person fg/bg separation; corpus win-rate."""
name = "judge.third_person_separation"
requires_reference = True
higher_is_better = True
needs_gpu = False
is_set_metric = True
dependencies = ["google.genai"]
def __init__(self, model: str = DEFAULT_MODEL, k: int = DEFAULT_K) -> None:
super().__init__()
self.model = model
self.k = k
self._client: Any = None
self._files: dict[str, Any] = {} # path -> uploaded Gemini file handle
self._records: list[dict] = [] # one per accumulated pair
# --- model / client -----------------------------------------------------
def setup(self) -> None:
if self._client is not None:
return
from google import genai
self._client = genai.Client(api_key=_resolve_api_key())
# --- set-vs-set protocol ------------------------------------------------
def reset(self) -> None:
self._records = []
self._files = {}
def accumulate(self, sample: dict) -> None:
cand = _path_of(sample, "video")
base = _path_of(sample, "reference")
if cand is None or base is None:
return # nothing to compare
image = sample.get("image_path")
action_text = sample.get("text_prompt") or ""
action = sample.get("action")
rec = self._cached(cand, base, action_text)
if rec is None:
if self._client is None:
self.setup()
rec = self._judge_pair(cand, base, image, action_text)
self._write_cache(cand, base, action_text, rec)
rec = {**rec, "action": action}
self._records.append(rec)
def finalize(self) -> MetricResult:
recs = [r for r in self._records if r.get("verdict")]
if not recs:
return MetricResult(name=self.name, score=None, details={"skipped": "no pairs judged"})
wins = sum(r["verdict"] == "candidate" for r in recs)
losses = sum(r["verdict"] == "baseline" for r in recs)
ties = sum(r["verdict"] == "tie" for r in recs)
decided = wins + losses
score = wins / decided if decided else None
# Per-action win-rate, grouped by the raw label (no assumed control scheme).
per_action: dict[str, dict] = {}
labels: set[str] = {str(r["action"]) for r in recs if r.get("action")}
for action in sorted(labels):
gr = [r for r in recs if r.get("action") == action]
gw = sum(r["verdict"] == "candidate" for r in gr)
gl = sum(r["verdict"] == "baseline" for r in gr)
per_action[action] = {
"n": len(gr),
"wins": gw,
"losses": gl,
"ties": len(gr) - gw - gl,
"win_rate": (gw / (gw + gl)) if (gw + gl) else None,
}
return MetricResult(name=self.name,
score=score,
details={
"wins": wins,
"losses": losses,
"ties": ties,
"n": len(recs),
"win_rate_excl_ties": score,
"per_action": per_action,
})
def merge_from(self, other: BaseMetric) -> None:
assert isinstance(other, ThirdPersonSeparationMetric)
self._records.extend(other._records)
# --- judging ------------------------------------------------------------
def _judge_pair(self, cand: str, base: str, image: str | None, action_text: str) -> dict:
"""k counterbalanced comparisons → aggregated per-pair verdict."""
# Seed the A/B alternation from the pair itself so it is reproducible and
# independent of evaluation order or which subset is being run.
seed = int(hashlib.sha1(f"{cand}|{base}".encode()).hexdigest(), 16)
mapped: list[str] = []
for i in range(self.k):
cand_first = (seed + i) % 2 == 0
v1, v2 = (cand, base) if cand_first else (base, cand)
winner = self._one_call(image, v1, v2, action_text)
if winner == "tie":
mapped.append("tie")
elif (winner == "video_1") == cand_first:
mapped.append("candidate")
else:
mapped.append("baseline")
cand_w = mapped.count("candidate")
base_w = mapped.count("baseline")
verdict = ("candidate" if cand_w > base_w else "baseline" if base_w > cand_w else "tie")
return {
"verdict": verdict,
"candidate_wins": cand_w,
"baseline_wins": base_w,
"ties": mapped.count("tie"),
"k": self.k,
"rubric_id": RUBRIC_ID
}
def _one_call(self, image: str | None, vid1: str, vid2: str, action_text: str) -> str:
from google.genai import types
contents: list[Any] = [action_text or "Compare these two rollouts."]
if image is not None:
contents += ["\nFirst frame (input condition, shared by BOTH "
"videos):", self._upload(image)]
contents += [
"\nVideo 1 (model A's full rollout — watch it in motion):",
self._upload(vid1),
"\nVideo 2 (model B's full rollout — watch it in motion):",
self._upload(vid2),
"\n" + RUBRIC + "\n\n" + USER_TASK,
]
for attempt in range(6):
try:
resp = self._client.models.generate_content(model=self.model,
contents=contents,
config=types.GenerateContentConfig(
system_instruction=SYSTEM_PROMPT,
response_mime_type="application/json",
response_schema=RESPONSE_SCHEMA,
temperature=0.4))
return json.loads(resp.text).get("winner", "tie")
except Exception as exc: # noqa: BLE001 - transient API errors
if attempt == 5:
print(f"[judge] giving up after 6 attempts ({exc}); scoring this call a tie")
break
is_429 = "429" in str(exc) or "RESOURCE_EXHAUSTED" in str(exc)
time.sleep(40 if is_429 else 2**attempt)
return "tie"
def _upload(self, path: str) -> Any:
f = self._files.get(path)
if f is not None:
return f
f = self._client.files.upload(file=path)
while f.state.name != "ACTIVE":
time.sleep(1)
f = self._client.files.get(name=f.name)
if f.state.name == "FAILED":
raise RuntimeError(f"Gemini upload failed for {path}")
self._files[path] = f
return f
# --- per-pair cache -----------------------------------------------------
def _cache_path(self, cand: str, base: str, action_text: str) -> Path:
# k is intentionally NOT in the key so a larger-k run can reuse an
# existing verdict with enough samples (see ``_cached``).
h = hashlib.sha1(f"{RUBRIC_ID}|{self.model}|{cand}|{base}|{action_text}".encode()).hexdigest()[:16]
return get_cache_dir() / "judge" / "third_person_separation" / f"{h}.json"
def _cached(self, cand: str, base: str, action_text: str) -> dict | None:
cp = self._cache_path(cand, base, action_text)
if not cp.is_file():
return None
try:
rec = json.loads(cp.read_text())
except Exception:
return None
return rec if rec.get("k", 0) >= self.k and "verdict" in rec else None
def _write_cache(self, cand: str, base: str, action_text: str, rec: dict) -> None:
cp = self._cache_path(cand, base, action_text)
cp.parent.mkdir(parents=True, exist_ok=True)
cp.write_text(json.dumps(rec, indent=2))
-2
View File
@@ -85,8 +85,6 @@ def _extra_for(metric_name: str) -> str:
return "eval-physics-iq"
if metric_name.startswith("audio."):
return "eval-audio"
if metric_name.startswith("judge."):
return "eval-judge"
return "eval"
-46
View File
@@ -219,15 +219,6 @@ class ReplicatedLinear(LinearBase):
(e.g. model.layers.0.qkv_proj)
"""
# Opt-in instrumentation: when ``enable_shape_tracking`` is set to True,
# ``forward`` records every unique ``(input_shape, output_shape)`` pair
# observed across all ``ReplicatedLinear`` instances, along with the
# subclass name that produced it. Used by upcoming QAT-aware backends
# to discover which GEMM shapes need quantized kernels. Defaults to
# False; default forward path is bit-identical to pre-slice behavior.
enable_shape_tracking = False
_shape_to_layer_types: dict[tuple[torch.Size, torch.Size], set[str]] = {}
def __init__(
self,
input_size: int,
@@ -294,8 +285,6 @@ class ReplicatedLinear(LinearBase):
bias = self.bias if not self.skip_bias_add else None
assert self.quant_method is not None
output = self.quant_method.apply(self, x, bias)
if self.enable_shape_tracking:
self._track_shape(x.shape, output.shape)
output_bias = self.bias if self.skip_bias_add else None
return output, output_bias
@@ -305,41 +294,6 @@ class ReplicatedLinear(LinearBase):
s += f", bias={self.bias is not None}"
return s
@classmethod
def get_shape_mapping(cls) -> dict:
"""Get the mapping from (input_shape, output_shape) to layer types."""
return cls._shape_to_layer_types.copy()
@classmethod
def reset_shape_tracking(cls) -> None:
"""Clear tracked shapes and layer type mappings."""
cls._shape_to_layer_types.clear()
def _track_shape(self, input_shape: torch.Size, output_shape: torch.Size) -> None:
shape_key = (input_shape, output_shape)
if shape_key not in self._shape_to_layer_types:
self._shape_to_layer_types[shape_key] = set()
logger.debug("Layer: %s | input shape: %s --> output shape: %s, Quant Method: %s", self.prefix, input_shape,
output_shape, self.quant_method.__class__.__name__)
self._shape_to_layer_types[shape_key].add(self.__class__.__name__)
@classmethod
def print_shape_summary(cls) -> None:
"""Log a summary of all unique shapes and their layer types."""
if not cls._shape_to_layer_types:
logger.info("No shapes have been processed yet.")
return
lines = [
"=== Matrix Multiplication Shape Summary ===",
f"Total unique shapes: {len(cls._shape_to_layer_types)}",
]
for i, (shape_key, layer_types) in enumerate(cls._shape_to_layer_types.items(), 1):
input_shape, output_shape = shape_key
lines.append(f"{i}. Input: {input_shape} → Output: {output_shape}")
lines.append(f" Layer types: {', '.join(sorted(layer_types))}")
logger.info("\n".join(lines))
class ColumnParallelLinear(LinearBase):
"""Linear layer with column parallelism.

Some files were not shown because too many files have changed in this diff Show More