Compare commits
101
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4a33540c08 | ||
|
|
ba83e9bef7 | ||
|
|
21e59268d8 | ||
|
|
139577639f | ||
|
|
db866dfb04 | ||
|
|
53d63afa76 | ||
|
|
685702e330 | ||
|
|
9e8d515842 | ||
|
|
1302a27470 | ||
|
|
e1a2c9ede8 | ||
|
|
8ee7270368 | ||
|
|
26c46b901a | ||
|
|
fcc5a26b46 | ||
|
|
6450f9fa26 | ||
|
|
4d079c65ed | ||
|
|
55fcf27759 | ||
|
|
3149d0ec89 | ||
|
|
1f4b729c4a | ||
|
|
ac29750b55 | ||
|
|
50fd379f3a | ||
|
|
fa9d8040e1 | ||
|
|
9794be7202 | ||
|
|
e4b14e124d | ||
|
|
60ca94796c | ||
|
|
060bdd85b6 | ||
|
|
0a18b6a031 | ||
|
|
2e5bad1d6c | ||
|
|
705f580542 | ||
|
|
cbe67337c3 | ||
|
|
113478fd74 | ||
|
|
9b2007c058 | ||
|
|
1542b7a03d | ||
|
|
e4b7261b66 | ||
|
|
b39b68c6f5 | ||
|
|
5f698f03f5 | ||
|
|
799a9b69a2 | ||
|
|
27b320afd8 | ||
|
|
c74e9cb4fe | ||
|
|
0a2c985b56 | ||
|
|
8b343db604 | ||
|
|
24b76d372c | ||
|
|
f7256ab197 | ||
|
|
9de0e2bcd1 | ||
|
|
5f33f80d6c | ||
|
|
e7d4aaeef4 | ||
|
|
2a085a5c39 | ||
|
|
261d3007a9 | ||
|
|
b2da3b2b4d | ||
|
|
7ec4d01afc | ||
|
|
8b0e2bbb4b | ||
|
|
7fd58f26ea | ||
|
|
15b6278e9f | ||
|
|
cd0fbbd3a0 | ||
|
|
02151584cd | ||
|
|
51bb5e349a | ||
|
|
1bdbf1d5b3 | ||
|
|
de295403c7 | ||
|
|
34834eb4f4 | ||
|
|
bd36d69680 | ||
|
|
94b84b03b9 | ||
|
|
a16ce1a07f | ||
|
|
37544c68a5 | ||
|
|
88e4a4355c | ||
|
|
f1dead1361 | ||
|
|
cfc2ad7487 | ||
|
|
97fedaf432 | ||
|
|
3ec511419f | ||
|
|
145082cf74 | ||
|
|
71caa58ba9 | ||
|
|
8240970ea6 | ||
|
|
44d10b5c79 | ||
|
|
9d874aa609 | ||
|
|
8e828680c6 | ||
|
|
735ee6aa56 | ||
|
|
068a532b31 | ||
|
|
87f98c9b8b | ||
|
|
e60601df7f | ||
|
|
eed9c4bfbf | ||
|
|
1dee77f4a4 | ||
|
|
b80148819c | ||
|
|
c3b971488e | ||
|
|
88e753f281 | ||
|
|
77832059cc | ||
|
|
633d393568 | ||
|
|
5854aec2ce | ||
|
|
30e45c2411 | ||
|
|
2a4fe697a6 | ||
|
|
921db7479d | ||
|
|
7f539424cb | ||
|
|
19a838f54f | ||
|
|
d922ab2cbc | ||
|
|
9ea77d37f3 | ||
|
|
2e35b0c6bd | ||
|
|
1c627a3f98 | ||
|
|
a931efe33a | ||
|
|
041e5e9029 | ||
|
|
efcc245c2e | ||
|
|
922e7e0813 | ||
|
|
c62a8514b0 | ||
|
|
3eb8081801 | ||
|
|
3505d09564 |
@@ -0,0 +1,207 @@
|
||||
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
|
||||
|
||||
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
|
||||
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
|
||||
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
|
||||
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
|
||||
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
|
||||
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
|
||||
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
|
||||
Inferix, and FlashDrive as techniques integratable into the graph runtime.
|
||||
|
||||
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
|
||||
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
|
||||
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
|
||||
corrections in; it is the corrected, authoritative synthesis.
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive summary
|
||||
|
||||
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
|
||||
|
||||
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
|
||||
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
|
||||
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
|
||||
is the exact MoT serving property the omni cards in this repo already express.
|
||||
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
|
||||
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
|
||||
flywheel driving the *same* serving Loop).
|
||||
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
|
||||
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
|
||||
already has the seam M\* only gestures at.**
|
||||
|
||||
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
|
||||
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
|
||||
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
|
||||
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
|
||||
|
||||
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
|
||||
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
|
||||
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
|
||||
|
||||
### Verified P0 correctness findings (spot-checked by hand)
|
||||
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
|
||||
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
|
||||
disaggregated runners execute *different node sets*. ✅ confirmed.
|
||||
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
|
||||
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
|
||||
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
|
||||
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
|
||||
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
|
||||
|
||||
---
|
||||
|
||||
## 2. Dimension table (corrected)
|
||||
|
||||
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
|
||||
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
|
||||
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
|
||||
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
|
||||
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
|
||||
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
|
||||
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
|
||||
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
|
||||
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
|
||||
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
|
||||
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
|
||||
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
|
||||
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
|
||||
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
|
||||
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
|
||||
|
||||
---
|
||||
|
||||
## 3. P0/P1 deep-dives (sequenced)
|
||||
|
||||
```
|
||||
PR-1 (P0) min-components ──┐
|
||||
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
|
||||
PR-3 (P1) output_determinism (independent)
|
||||
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
|
||||
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
|
||||
```
|
||||
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
|
||||
|
||||
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
|
||||
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
|
||||
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
|
||||
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
|
||||
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
|
||||
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
|
||||
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
|
||||
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
|
||||
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
|
||||
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
|
||||
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
|
||||
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
|
||||
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
|
||||
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
|
||||
|
||||
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
|
||||
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
|
||||
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
|
||||
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
|
||||
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
|
||||
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
|
||||
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
|
||||
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
|
||||
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
|
||||
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
|
||||
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
|
||||
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
|
||||
|
||||
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
|
||||
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
|
||||
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
|
||||
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
|
||||
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
|
||||
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
|
||||
|
||||
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
|
||||
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
|
||||
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
|
||||
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
|
||||
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
|
||||
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
|
||||
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
|
||||
already exist and are tested).
|
||||
|
||||
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
|
||||
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
|
||||
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
|
||||
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
|
||||
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
|
||||
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
|
||||
admission per (class,label).
|
||||
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
|
||||
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
|
||||
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
|
||||
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
|
||||
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
|
||||
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
|
||||
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
|
||||
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
|
||||
|
||||
---
|
||||
|
||||
## 4. What v2 already does ≥ M\* — do NOT regress
|
||||
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
|
||||
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
|
||||
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
|
||||
must pass it.
|
||||
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
|
||||
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
|
||||
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
|
||||
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
|
||||
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
|
||||
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
|
||||
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
|
||||
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
|
||||
interceptor) — M\* describes extensibility; v2 has the mechanism.
|
||||
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
|
||||
self-contained runtime.
|
||||
|
||||
---
|
||||
|
||||
## 5. Dropped / merged / deferred (and why)
|
||||
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
|
||||
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
|
||||
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
|
||||
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
|
||||
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
|
||||
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
|
||||
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
|
||||
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
|
||||
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
|
||||
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
|
||||
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
|
||||
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
|
||||
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
|
||||
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
|
||||
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
|
||||
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
|
||||
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
|
||||
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
|
||||
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
|
||||
(per-step capture ships in 14 cards, not just wan21).
|
||||
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
|
||||
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
|
||||
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
|
||||
|
||||
---
|
||||
|
||||
## 6. The first integration test, if/when multi-GPU placement work starts
|
||||
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
|
||||
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
|
||||
placement work — it is the one place this repo already has real multi-rank composite-model execution.
|
||||
|
||||
---
|
||||
|
||||
## Anchor files for P0/P1
|
||||
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
|
||||
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
|
||||
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
|
||||
`v2/program/workflow.py` (registry-driven delivery).
|
||||
@@ -0,0 +1,331 @@
|
||||
# Performance Dashboard Memory
|
||||
|
||||
Date: 2026-06-16
|
||||
Branch: `ci/dashboard`
|
||||
|
||||
## Purpose
|
||||
|
||||
This branch adds a local live dashboard for FastVideo performance benchmark
|
||||
history. It is intended for maintainer/operator use: inspect latest benchmark
|
||||
status, compare current values with recent baseline context, and view trends
|
||||
from the Hugging Face performance-tracking dataset.
|
||||
|
||||
The dashboard is a FastAPI + React app. It is separate from the existing
|
||||
Svelte `ui/` app.
|
||||
|
||||
## Main Files Added Or Changed
|
||||
|
||||
Backend:
|
||||
|
||||
- `fastvideo/performance_dashboard/__init__.py`
|
||||
- `fastvideo/performance_dashboard/__main__.py`
|
||||
- `fastvideo/performance_dashboard/api.py`
|
||||
- `fastvideo/performance_dashboard/metrics.py`
|
||||
- `fastvideo/performance_dashboard/service.py`
|
||||
|
||||
Frontend:
|
||||
|
||||
- `performance_dashboard/frontend/package.json`
|
||||
- `performance_dashboard/frontend/package-lock.json`
|
||||
- `performance_dashboard/frontend/tsconfig.json`
|
||||
- `performance_dashboard/frontend/vite.config.ts`
|
||||
- `performance_dashboard/frontend/index.html`
|
||||
- `performance_dashboard/frontend/scripts/build.mjs`
|
||||
- `performance_dashboard/frontend/src/main.tsx`
|
||||
- `performance_dashboard/frontend/src/api.ts`
|
||||
- `performance_dashboard/frontend/src/App.tsx`
|
||||
- `performance_dashboard/frontend/src/styles.css`
|
||||
|
||||
Docs/tests:
|
||||
|
||||
- `performance_dashboard/README.md`
|
||||
- `docs/contributing/performance_benchmarks.md`
|
||||
- `fastvideo/tests/performance/test_dashboard_service.py`
|
||||
- `fastvideo/tests/performance/test_dashboard_api.py`
|
||||
|
||||
Shared HF utility change:
|
||||
|
||||
- `fastvideo/tests/performance/hf_store.py`
|
||||
|
||||
## Data Source
|
||||
|
||||
The source of truth remains the Hugging Face dataset repo used by existing
|
||||
performance CI:
|
||||
|
||||
```text
|
||||
HF_REPO_ID=FastVideo/performance-tracking
|
||||
```
|
||||
|
||||
The dataset stores normalized JSON records emitted by
|
||||
`fastvideo/tests/performance/compare_baseline.py`. The current v1 normalized
|
||||
schema includes:
|
||||
|
||||
- `model_id`
|
||||
- `timestamp`
|
||||
- `commit_sha`
|
||||
- `gpu_type`
|
||||
- `latency`
|
||||
- `throughput`
|
||||
- `memory`
|
||||
- `text_encoder_time_s`
|
||||
- `dit_time_s`
|
||||
- `vae_decode_time_s`
|
||||
- `success`
|
||||
|
||||
Records are grouped by `(model_id, gpu_type)` for v1 dashboard behavior.
|
||||
|
||||
## Local Cache
|
||||
|
||||
The backend syncs the HF dataset to a local cache directory:
|
||||
|
||||
```text
|
||||
PERFORMANCE_TRACKING_ROOT=/tmp/fastvideo-perf-dashboard
|
||||
```
|
||||
|
||||
If `PERFORMANCE_TRACKING_ROOT` is not set, the dashboard defaults to:
|
||||
|
||||
```text
|
||||
/tmp/fastvideo-perf-dashboard
|
||||
```
|
||||
|
||||
The sync is performed through the existing helper:
|
||||
|
||||
```python
|
||||
fastvideo.tests.performance.hf_store.sync_from_hf(...)
|
||||
```
|
||||
|
||||
The dashboard then loads JSON files from the local cache through:
|
||||
|
||||
```python
|
||||
fastvideo.tests.performance.hf_store.load_records(...)
|
||||
```
|
||||
|
||||
## Authentication
|
||||
|
||||
Originally `hf_store.py` only read `HF_API_KEY`. This caused local dashboard
|
||||
runs to fail when users had standard Hugging Face token variables set.
|
||||
|
||||
`hf_store.py` now resolves tokens from the first available variable in:
|
||||
|
||||
```text
|
||||
HF_API_KEY
|
||||
HUGGINGFACE_HUB_TOKEN
|
||||
HF_TOKEN
|
||||
```
|
||||
|
||||
For local use:
|
||||
|
||||
```bash
|
||||
export HF_TOKEN=hf_...
|
||||
```
|
||||
|
||||
If the HF repo is private or gated, the token must have dataset read access.
|
||||
|
||||
## Backend API
|
||||
|
||||
The FastAPI app is created by:
|
||||
|
||||
```python
|
||||
fastvideo.performance_dashboard.api:create_app
|
||||
```
|
||||
|
||||
The module-level app is:
|
||||
|
||||
```python
|
||||
fastvideo.performance_dashboard.api:app
|
||||
```
|
||||
|
||||
Endpoints:
|
||||
|
||||
- `GET /api/performance/health`
|
||||
- `POST /api/performance/refresh`
|
||||
- `GET /api/performance/records?days=90`
|
||||
- `GET /api/performance/summary?days=90`
|
||||
- `GET /api/performance/trends?days=90`
|
||||
|
||||
`POST /api/performance/refresh` forces a fresh HF sync.
|
||||
|
||||
## Status Semantics
|
||||
|
||||
Important: the dashboard intentionally separates stored CI status from
|
||||
recomputed context.
|
||||
|
||||
Stored status:
|
||||
|
||||
- Comes directly from the latest JSON record's `success` field.
|
||||
- This is what the dashboard displays as `Stored Status`.
|
||||
- This is the primary latest status.
|
||||
|
||||
Recomputed status:
|
||||
|
||||
- Calculated locally from cached records for explanatory context.
|
||||
- Uses the latest record's metric values compared to the median of the latest
|
||||
five previous successful records in the same `(model_id, gpu_type)` group.
|
||||
- Displayed separately as `Recomputed`.
|
||||
- Does not override the stored JSON `success` status.
|
||||
|
||||
This distinction was added after observing that recomputing pass/fail from the
|
||||
local cache can disagree with the status originally uploaded by CI.
|
||||
|
||||
## Time Window Behavior
|
||||
|
||||
The default dashboard time window is 90 days.
|
||||
|
||||
The selected `days` value affects:
|
||||
|
||||
- trend charts
|
||||
- record browsing/filtering
|
||||
|
||||
The selected `days` value does not affect:
|
||||
|
||||
- latest stored status
|
||||
- latest summary baseline context
|
||||
|
||||
Reason: latest status should not change when users widen or narrow the trend
|
||||
window. The API keeps `days` on `/summary` only for shared frontend filter
|
||||
state, but summary loading uses all cached records.
|
||||
|
||||
This fixed a bug where changing from roughly 35 days to 42 days could change
|
||||
the latest status from pass to fail because older records entered the local
|
||||
baseline window.
|
||||
|
||||
## Metric Logic
|
||||
|
||||
Dashboard metric definitions live in:
|
||||
|
||||
```text
|
||||
fastvideo/performance_dashboard/metrics.py
|
||||
```
|
||||
|
||||
Tracked metrics:
|
||||
|
||||
- `latency` lower is better
|
||||
- `throughput` higher is better
|
||||
- `memory` lower is better
|
||||
- `text_encoder_time_s` lower is better
|
||||
- `dit_time_s` lower is better
|
||||
- `vae_decode_time_s` lower is better
|
||||
|
||||
Baseline context uses the median of up to five previous successful records for
|
||||
the same `(model_id, gpu_type)`.
|
||||
|
||||
## Frontend Behavior
|
||||
|
||||
The React app:
|
||||
|
||||
- fetches `/api/performance/summary`
|
||||
- fetches `/api/performance/trends`
|
||||
- displays summary cards
|
||||
- displays latest rows by model/GPU
|
||||
- displays native SVG trend charts
|
||||
- has model/GPU/day filters
|
||||
- includes a refresh button
|
||||
- auto-refreshes every five minutes
|
||||
|
||||
The UI is implemented without a charting library. Trend charts are native SVG
|
||||
in `performance_dashboard/frontend/src/App.tsx`.
|
||||
|
||||
The production frontend build uses `esbuild` through
|
||||
`performance_dashboard/frontend/scripts/build.mjs`. Vite is still used for the
|
||||
dev server and `/api` proxy.
|
||||
|
||||
Why esbuild for production build:
|
||||
|
||||
- Vite/Rollup hit a local macOS native optional dependency code-signing issue
|
||||
in this environment.
|
||||
- Direct esbuild worked reliably and is sufficient for this small dashboard.
|
||||
|
||||
## Static Serving
|
||||
|
||||
After frontend build, the FastAPI server serves:
|
||||
|
||||
- static JS/CSS from `performance_dashboard/frontend/dist/assets`
|
||||
- `performance_dashboard/frontend/dist/index.html` for the dashboard page
|
||||
|
||||
This allows a single local port to serve both the API and UI.
|
||||
|
||||
## Local Run Workflow
|
||||
|
||||
Build frontend:
|
||||
|
||||
```bash
|
||||
cd performance_dashboard/frontend
|
||||
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
|
||||
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm install
|
||||
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
|
||||
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
|
||||
```
|
||||
|
||||
Run dashboard:
|
||||
|
||||
```bash
|
||||
export HF_TOKEN=hf_...
|
||||
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
|
||||
```
|
||||
|
||||
Open locally:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:8000
|
||||
```
|
||||
|
||||
## ngrok Workflow
|
||||
|
||||
`python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000`
|
||||
starts the actual local dashboard server.
|
||||
|
||||
`ngrok http 8000` does not start the dashboard. It exposes the already-running
|
||||
local server through a temporary public URL.
|
||||
|
||||
Typical flow:
|
||||
|
||||
```bash
|
||||
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
|
||||
ngrok http 8000
|
||||
```
|
||||
|
||||
Use the HTTPS URL printed by ngrok to view the dashboard remotely.
|
||||
|
||||
## Verification Commands
|
||||
|
||||
Backend tests:
|
||||
|
||||
```bash
|
||||
conda run -n fastvideo python -m pytest \
|
||||
fastvideo/tests/performance/test_dashboard_service.py \
|
||||
fastvideo/tests/performance/test_dashboard_api.py \
|
||||
-q
|
||||
```
|
||||
|
||||
Expected after latest changes:
|
||||
|
||||
```text
|
||||
8 passed
|
||||
```
|
||||
|
||||
Frontend build:
|
||||
|
||||
```bash
|
||||
cd performance_dashboard/frontend
|
||||
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
|
||||
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
|
||||
```
|
||||
|
||||
Expected:
|
||||
|
||||
```text
|
||||
tsc && node scripts/build.mjs
|
||||
```
|
||||
|
||||
with exit code 0.
|
||||
|
||||
## Known Notes
|
||||
|
||||
- `performance_dashboard/frontend/node_modules/` and
|
||||
`performance_dashboard/frontend/dist/` are ignored by git.
|
||||
- `npm install` reported two high-severity audit findings in dependency tree.
|
||||
`npm audit fix --force` was not run because it can introduce breaking
|
||||
dependency upgrades.
|
||||
- Existing `fastvideo` package imports may emit platform warnings such as NPU
|
||||
or macOS torch distributed messages. These are not dashboard-specific errors.
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
name: add-reward-model
|
||||
description: Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.
|
||||
---
|
||||
|
||||
# Add Reward Model
|
||||
|
||||
Use for reward models consumed by RL methods.
|
||||
|
||||
## Placement
|
||||
|
||||
- Put reusable reward code under `fastvideo/train/methods/rl/rewards/`.
|
||||
- Expose public builders from `fastvideo/train/methods/rl/rewards/__init__.py`.
|
||||
- Keep method-specific aggregation or advantage logic out of reward classes.
|
||||
|
||||
## Media Inputs
|
||||
|
||||
- Reward callables receive decoded media tensors.
|
||||
- Accept single-frame tensors as `[B, C, H, W]` and multi-frame tensors as `[B, C, T, H, W]` when practical.
|
||||
- Frame selection is reward-specific. Frame scorers such as PickScore and CLIPScore should explicitly select frame `0`; temporal rewards should inspect whichever frames they need.
|
||||
- Return one scalar reward per prompt/sample.
|
||||
|
||||
## Attribution
|
||||
|
||||
- If code is ported or closely adapted from another repo, add a short comment or docstring naming the source file/function.
|
||||
- Preserve SPDX headers used by FastVideo files.
|
||||
|
||||
## Tests
|
||||
|
||||
- Unit-test tensor layout handling without loading large reward checkpoints.
|
||||
- Allow fake scorer injection for multi-reward tests.
|
||||
- Test weighted reward aggregation and metric keys.
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: add-rl-method
|
||||
description: Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.
|
||||
---
|
||||
|
||||
# Add RL Method
|
||||
|
||||
Use for new RL methods in the modular `fastvideo/train` stack.
|
||||
|
||||
## Required Shape
|
||||
|
||||
- Add the method under `fastvideo/train/methods/rl/`.
|
||||
- Subclass `TrainingMethod`.
|
||||
- Keep model-family logic in `ModelBase` wrappers.
|
||||
- Decode generated latents through `ModelBase.decode_latents`; add that hook to the new model wrapper instead of decoding inside the RL method.
|
||||
- Use `fastvideo/train/methods/rl/common/sampling.py` for generation unless the method has a documented reason to avoid sampling.
|
||||
- Use `fastvideo/train/methods/rl/common/prompt_sampling.py` for reusable grouped prompt sampling patterns such as DiffusionNFT K-repeat.
|
||||
- Use `fastvideo/train/methods/rl/rewards/` for reward models.
|
||||
|
||||
## Optimization
|
||||
|
||||
- Return `manages_optimization() == True` only when the method must own a nonstandard outer/inner loop.
|
||||
- If using managed optimization, implement `managed_train_step(data_stream, iteration)`.
|
||||
- Existing trainer callbacks, checkpointing, tracking, and validation should still work.
|
||||
|
||||
## Config
|
||||
|
||||
- Put method knobs under `method`.
|
||||
- Put sampler knobs under `method.sampling`.
|
||||
- Do not put scheduler or trajectory policy into model configs.
|
||||
- Do not split a diffusers-style scheduler from its built-in `step()` solver in YAML; use `trajectory` only for higher-level ODE vs re-noise behavior.
|
||||
- Avoid fixed timestep lists in examples unless reproducing a known baseline; prefer scheduler-generated defaults.
|
||||
|
||||
## Tests
|
||||
|
||||
- Add fake-model tests for sampler/method behavior.
|
||||
- Add config parse tests for the public YAML.
|
||||
- Confirm existing train methods stay on the default Trainer path.
|
||||
@@ -8,3 +8,6 @@
|
||||
{"name": "decompose-pipeline-pr", "description": "Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body. Worked example: PR #1280 daVinci-MagiHuman (9.8k LOC) decomposed into 10 stacked PRs.", "path": "decompose-pipeline-pr/SKILL.md", "status": "tested", "trust": "medium"}
|
||||
{"name": "reseed-performance-baseline", "description": "Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline must be advanced by replicating one reviewed shifted source result into three success=true records, or five records when explicitly requested", "path": "reseed-performance-baseline/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-model", "description": "Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.", "path": "add-model/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "rlhf-training-abstractions", "description": "Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.", "path": "rlhf-training-abstractions/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-rl-method", "description": "Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.", "path": "add-rl-method/SKILL.md", "status": "draft", "trust": "low"}
|
||||
{"name": "add-reward-model", "description": "Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.", "path": "add-reward-model/SKILL.md", "status": "draft", "trust": "low"}
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
name: rlhf-training-abstractions
|
||||
description: Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.
|
||||
---
|
||||
|
||||
# RLHF Training Abstractions
|
||||
|
||||
Use this skill before editing RLHF-style training code in `fastvideo/train`.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- RL methods live under `fastvideo/train/methods/rl/` and own algorithm logic: reward collection, advantage computation, policy loss, KL/reference terms, and optimizer cadence.
|
||||
- Rewards live under `fastvideo/train/methods/rl/rewards/` and must be reusable across RL methods.
|
||||
- RL methods pass decoded media to rewards; each reward decides whether to use the first frame, sampled frames, or the full video.
|
||||
- Sampling lives under `fastvideo/train/methods/rl/common/` and must use `ModelBase` primitives plus scheduler math, not model-family inference pipelines.
|
||||
- Model wrappers under `fastvideo/train/models/` own model-specific forward details.
|
||||
- Model wrappers also own model-specific latent decoding via `ModelBase.decode_latents`; RL methods should not reach into VAE normalization internals.
|
||||
- Shared RL helpers such as K-repeat prompt sampling belong under `fastvideo/train/methods/rl/common/` when they are reusable across RL methods.
|
||||
|
||||
## Anti-Patterns
|
||||
|
||||
- Do not bind RL methods to inference pipeline classes such as `WanDMDPipeline`.
|
||||
- Do not hardcode timestep lists in a method when the scheduler can generate them.
|
||||
- Do not put reward-model code inside one RL method.
|
||||
- Do not make existing non-RL methods use method-managed optimization unless explicitly requested.
|
||||
|
||||
## Sampling Policy
|
||||
|
||||
- Prefer YAML-configured `method.sampling` with `scheduler`, `trajectory`, `num_steps`, `timesteps`, and `sigmas`.
|
||||
- Treat diffusers-style scheduler classes as owning both the timestep schedule and their `step()` update rule; avoid a separate `solver` field unless a new sampler truly implements solver math outside the scheduler object.
|
||||
- Missing `timesteps` means “ask the scheduler”; explicit `timesteps` or `sigmas` are overrides.
|
||||
- ODE-style trajectories should not re-noise between denoising steps.
|
||||
- SDE/re-noise behavior must be explicit in config.
|
||||
|
||||
## Validation
|
||||
|
||||
- Run focused local tests for sampler config and Trainer opt-in behavior.
|
||||
- Verify existing train methods still report `manages_optimization() == False`.
|
||||
- Keep fixed-prompt validation helpers in `fastvideo/train/methods/rl/common/validation.py` so new RL methods can reuse sharding and captions.
|
||||
- Test distributed prompt grouping helpers separately from heavyweight model loading.
|
||||
- Run `pre-commit run --files <changed paths>`; respect configured excludes.
|
||||
@@ -34,6 +34,8 @@ env
|
||||
*.log
|
||||
weights/
|
||||
logs/
|
||||
official_weights/
|
||||
converted_weights/
|
||||
|
||||
# SSIM test outputs
|
||||
fastvideo/tests/ssim/generated_videos/
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
# FastVideo — Design Philosophy
|
||||
|
||||
One page on *why* FastVideo is built the way it is. The full architecture, the as-built status, and the
|
||||
forward roadmap live in **[`v2/README.md`](v2/README.md)** — this is the philosophy beneath it.
|
||||
|
||||
---
|
||||
|
||||
**A deployable model is a post-training artifact.** Unlike an LLM — where inference optimizes frozen weights
|
||||
after the fact — a *usable* video/omni model is *created* by training: step distillation for latency, QAT for
|
||||
precision, distillation + self-forcing for causal/world models. So every inference capability is a
|
||||
**(recipe, runtime) pair**: the weights and the loop that produced-and-assumes them are one versioned object.
|
||||
This is the source of the moat — whoever owns *both* sides of the pair owns the optimization frontier — and it
|
||||
is why training and serving cannot be two systems.
|
||||
|
||||
**The work is loops, not `forward()`.** Denoise timesteps, AR decode, chunked rollout, VAE tiles, encoder
|
||||
chunks, audio tokens, reward batches, optimizer steps, media chunks — video and omni inference is iteration. A
|
||||
runtime that collapses everything to a single `forward` can't schedule, batch, cancel, stream, reserve memory
|
||||
for, or capture the behavior of what actually runs. So loops are first-class, and they are **driven**: the
|
||||
model describes the next step it needs, the runtime decides when and with whom it runs, the model folds the
|
||||
result back. The model keeps content-adaptive control flow; the runtime keeps admission, batching, streaming,
|
||||
and behavior capture. Per-request state lives in typed `LoopState`, never in module globals — so interleaving
|
||||
requests through one model instance cannot smear state, by construction.
|
||||
|
||||
**The model is the center; everything else is a view over it.** A typed `ModelCard` owns components, loops,
|
||||
the recipe, and the parity contract. Programs compose a card's loops into a task; Workflows compose cards into
|
||||
pipelines; the scheduler runs the *steps* of all loops as `WorkUnit`s under one currency (predicted GPU-time,
|
||||
because a bidirectional denoise step and an AR token are ~1000× apart and incommensurable in counts);
|
||||
deployment places and routes; products stream artifacts. None of them define model semantics — they reference
|
||||
the Model Plane. One resident instance can run many loop types on shared weights, which is what makes omni/MoT
|
||||
native rather than a DAG that doubles weights.
|
||||
|
||||
**Correctness is a typed contract, not a hope.** Caches are correct by *key* — if a field can change output
|
||||
semantics it is in the key, so reuse is partitioned, never blindly flushed. Parity between the train-forward
|
||||
and the serve-forward is *measured* on a declared ladder (component → loop → behavioral → distribution →
|
||||
artifact-quality), never assumed. And the non-negotiable gate is **interleave bit-parity**: N requests
|
||||
interleaved at step granularity must be bit-identical to running them serially — the test the whole
|
||||
loop-inversion bet lives or dies on.
|
||||
|
||||
**One substrate for inference, training, and RL.** The rollout forward *is* the serve forward plus capture —
|
||||
same loop, same caches, same batcher, same numerics — so every serving optimization is automatically a rollout
|
||||
optimization, and there is one numerics surface the ladder measures rather than a correction layer papering
|
||||
over it. The engine doubles as the RL rollout engine under a strict rule: `training` consumes the engine; the
|
||||
**engine never imports `training`**.
|
||||
|
||||
**Borrow aggressively; copy nothing as the core.** vLLM/SGLang scheduling, vLLM-Omni/SGLang-Omni omni serving,
|
||||
Dynamo fleet orchestration, diffusers components, xDiT parallelism, TorchTitan mesh discipline,
|
||||
verl-omni/miles RL lessons, ComfyUI workflows, Dreamverse/LiveKit sessions — each contributes a take, none is
|
||||
the center. Deployment orchestration (Dynamo) sits *above* the engine, never inside it. Extensions are
|
||||
versioned hook points, never monkeypatching. New frontier capabilities arrive as a card, a method, a loop, a
|
||||
workflow, or a controller — **not a rewrite**.
|
||||
|
||||
> A model card is a (recipe, runtime) pair with a parity obligation. The model owns loop semantics; the runtime
|
||||
> owns loop lifecycle. One resident instance runs many loops; one scheduler runs their steps in one currency.
|
||||
> Caches are correct by key; parity is correct by test; the interleave gate is non-negotiable. Training records
|
||||
> behavior on the same loops it serves. Deployment places and routes; products stream artifacts; neither defines
|
||||
> the model.
|
||||
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,11 +55,13 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
|
||||
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
|
||||
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
|
||||
|
||||
# Install FastVideo Unified Kernel
|
||||
# Install FastVideo Unified Kernel.
|
||||
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
|
||||
# of probing a live device for the arch (matches the released kernel wheel).
|
||||
RUN source $HOME/.local/bin/env && \
|
||||
source /opt/venv/bin/activate && \
|
||||
cd fastvideo-kernel && \
|
||||
git submodule update --init --recursive && \
|
||||
./build.sh
|
||||
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
|
||||
|
||||
|
||||
EXPOSE 22
|
||||
|
||||
@@ -41,6 +41,13 @@ when you want local Markdown/normalized-result artifacts from the comparator.
|
||||
`fastvideo/tests/performance/results/`; remove stale result files if you only
|
||||
want to compare the latest local run.
|
||||
|
||||
## Local live dashboard
|
||||
|
||||
For an app-style local dashboard backed by the same HF performance-tracking
|
||||
records, see `performance_dashboard/README.md`. The dashboard provides a
|
||||
FastAPI API plus a React UI and can be exposed with `ngrok` after building the
|
||||
frontend.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
|
||||
@@ -108,6 +108,7 @@ surfaces:
|
||||
vae_sp: generator.pipeline.preset_overrides.vae_sp
|
||||
dmd_denoising_steps: generator.pipeline.preset_overrides.dmd_denoising_steps
|
||||
ti2v_task: generator.pipeline.preset_overrides.ti2v_task
|
||||
lucy_edit_task: generator.pipeline.preset_overrides.lucy_edit_task
|
||||
boundary_ratio: generator.pipeline.preset_overrides.boundary_ratio
|
||||
compatibility_only:
|
||||
model_path: "Redundant with generator.model_path."
|
||||
@@ -125,9 +126,18 @@ surfaces:
|
||||
text_encoder_configs: "Legacy internal component config object."
|
||||
preprocess_text_funcs: "Internal text preprocessing hooks."
|
||||
postprocess_text_funcs: "Internal text postprocessing hooks."
|
||||
scheduler_step_in_fp32: "Runtime scheduler precision toggle; not part of the public typed inference API."
|
||||
|
||||
pipeline_config_extensions:
|
||||
preset_owned:
|
||||
flux2_text_encoder_type:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
|
||||
text_encoder_out_layers:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
|
||||
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
|
||||
conditioning_strategy:
|
||||
sources:
|
||||
- fastvideo.configs.pipelines.cosmos.CosmosConfig
|
||||
@@ -491,6 +501,8 @@ surfaces:
|
||||
inpaint_mask: request.extensions.stable_audio.inpaint_mask
|
||||
internal_only:
|
||||
data_type: "Derived from the request shape and not a public input."
|
||||
latents: "Pre-generated diffusion latents supplied by parity/debug harnesses; not a public input."
|
||||
max_sequence_length: "Model-specific text-encoder sequence cap; not part of the public typed inference API."
|
||||
|
||||
sampling_param_extensions: {}
|
||||
|
||||
|
||||
@@ -24,6 +24,7 @@ This page describes the various options for speeding up generation times in Fast
|
||||
- Video Sparse Attention: `FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN`
|
||||
- Sage Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN`
|
||||
- Sage Attention 3: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN_THREE`
|
||||
- Attn-QAT inference (modified SageAttention3 FP4, sm_120/RTX 5090): `FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER`
|
||||
- Video MoBA Attention: `FASTVIDEO_ATTENTION_BACKEND=VMOBA_ATTN`
|
||||
- Sparse Linear Attention: `FASTVIDEO_ATTENTION_BACKEND=SLA_ATTN`
|
||||
- SageSLA Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_SLA_ATTN`
|
||||
@@ -122,6 +123,45 @@ gen.generate_video(prompt="A raccoon in sunflowers", save_video=True)
|
||||
- Per-call cosine similarity vs BF16: ~0.99 (slight quantization error accumulates over denoising steps)
|
||||
- Only supports `headdim >= 128`
|
||||
|
||||
### NVFP4 + Attn-QAT (modified SageAttention3, Blackwell sm_120)
|
||||
|
||||
**`ATTN_QAT_INFER`** with **`transformer_quant=nvfp4_qat`**
|
||||
|
||||
Runs the DiT fully in 4-bit: NVFP4 linear layers (activations quantized on the
|
||||
fly) plus the modified SageAttention3 FP4 attention backend. This is the
|
||||
inference half of the Quantization-Aware Distillation (QAD) recipe and the path
|
||||
used for the RTX 5090 release.
|
||||
|
||||
The `attn_qat_infer` kernel hard-gates on **sm_120 (consumer Blackwell / RTX
|
||||
5090)**; on other GPUs the backend logs a notice and falls back to Flash
|
||||
Attention. See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040).
|
||||
|
||||
Enable both halves — attention via the env var, linear via `transformer_quant`:
|
||||
|
||||
```python
|
||||
import os
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "ATTN_QAT_INFER"
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.layers.quantization import get_quantization_config
|
||||
gen = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
num_gpus=1,
|
||||
# Wan-2.1 uses the nvfp4_qat config (NVFP4 is LTX2-specific). Pass an
|
||||
# instance — the bare string is not resolved on the from_pretrained path.
|
||||
transformer_quant=get_quantization_config("nvfp4_qat")(),
|
||||
use_fsdp_inference=False, # FSDP shards invalidate the FP4 tensor pointers
|
||||
)
|
||||
gen.generate(request={"prompt": "A raccoon in sunflowers", "output": {"save_video": True}})
|
||||
```
|
||||
|
||||
Or run the example script:
|
||||
|
||||
```bash
|
||||
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py
|
||||
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py --bf16 # baseline
|
||||
```
|
||||
|
||||
### Sliding Tile Attention (Archived)
|
||||
|
||||
**`SLIDING_TILE_ATTN`**
|
||||
|
||||
@@ -58,6 +58,7 @@ pipeline initialization and sampling.
|
||||
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
|
||||
| Wan2.2 TI2V 5B | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | 720P | ⭕ | ⭕ | ✅ | ⭕ | ⭕ |
|
||||
| Lucy Edit Dev 5B*** | `decart-ai/Lucy-Edit-Dev` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
|
||||
| Wan2.2 T2V A14B | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
|
||||
| Wan2.2 I2V A14B | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
|
||||
| HunyuanVideo | `hunyuanvideo-community/HunyuanVideo` | 720px1280p<br>544px960p | ❌ | ✅ | ✅ | ⭕ | ⭕ |
|
||||
@@ -78,6 +79,9 @@ pipeline initialization and sampling.
|
||||
|
||||
**Note**: Wan2.2 TI2V 5B has some quality issues when performing I2V generation. We are working on fixing this issue.
|
||||
|
||||
***Lucy Edit Dev uses a non-commercial model license. FastVideo support is
|
||||
focused on inference integration for video editing workflows.
|
||||
|
||||
`Sliding Tile Attn (Legacy Branch)` entries refer to the archived
|
||||
`sta_do_not_delete` branch workflow, not active `main` inference wiring.
|
||||
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
# v2 porting status — fastvideo models → the v2 (recipe, runtime) substrate
|
||||
|
||||
Goal: every model in fastvideo's registry resolves through the **v2 `VideoGenerator`** / `Engine`
|
||||
(typed `fastvideo.api` configs + the real torch backend) to a recipe that can construct and run it.
|
||||
|
||||
**Scope: ALL fastvideo models (achieved).** v2 now resolves **63/64** of fastvideo's registered HF ids
|
||||
by exact id (PRIMARY), plus the architecture fallback for local/unregistered checkpoints. The single
|
||||
remaining id — `FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers` — is **environment-blocked**: its VSA
|
||||
(Sparse-Linear Attention) kernels require `nvcc` (not built in this bring-up). It arch-resolves to the
|
||||
base Wan card but needs the VSA kernel build to run faithfully.
|
||||
|
||||
Dispatch is **architecture-driven** (`v2/registry.py`): exact HF id → short-name → architecture
|
||||
inference from the checkpoint (pipeline / transformer / VAE class names + `z_dim`, `transformer_2`,
|
||||
`spatial_upsampler`). Adding a model is one `_BUCKET_C` row (HF ids → builders + transformer class).
|
||||
|
||||
## The porting mechanism — self-contained recipe packages
|
||||
Every net-new arch is a **self-contained recipe package** (`v2/recipes/<arch>/` = `card.py` `loop.py`
|
||||
`program.py` [+ `sampler.py`] + an optional `v2/platform/backends/torch_<arch>.py` adapter). The card
|
||||
declares its torch adapter via **`ComponentSpec.adapter="module:Class"`** (the `_explicit_adapter` seam in
|
||||
`torch_backend.py`) instead of editing the shared `_make_dit`/`_make_vae`/`_make_text_encoder` dispatch —
|
||||
so a port adds **only new files**, never touching shared code, and parallel ports never conflict. New
|
||||
samplers/loops live in-package. Registration is one row in `v2/registry.py:_BUCKET_C`.
|
||||
|
||||
## Working today (GPU-verified, real video/audio) — committed on `v2`
|
||||
| Official example(s) | Model | v2 card |
|
||||
|---|---|---|
|
||||
| `basic.py`, `basic_mps.py`, `basic_ray.py` | `Wan-AI/Wan2.1-T2V-1.3B-Diffusers` | wan21 |
|
||||
| `basic_self_forcing_causal.py` | `wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers` | wan_causal |
|
||||
| `basic_ltx2_distilled.py` | `FastVideo/LTX2-Distilled-Diffusers` (2-stage + spatial upsampler) | ltx2 |
|
||||
| `basic_wan2_2_ti2v.py` | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | wan2.2-ti2v |
|
||||
| `basic_wan2_2.py` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` (MoE + CPU expert offload) | wan2.2-a14b |
|
||||
| `basic_ltx2.py` | `Davids048/LTX2-Base-Diffusers` | ltx2 base |
|
||||
| `basic_ltx2_3_distilled.py` | `FastVideo/LTX-2.3-Distilled-Diffusers` (joint T2VS, video+audio) | ltx2.3-distilled |
|
||||
|
||||
Plus the **Wan2.1 i2v cluster** (Fun-1.3B-InP GPU-verified; I2V-14B-480P/720P + Wan2.2-I2V-A14B MoE reuse
|
||||
the i2v card) — CLIP image-encoder + first-frame `[mask|cond]` → 36ch DiT.
|
||||
|
||||
## GPU bring-up results (real weights on H100 NVL, single-GPU, TORCH_SDPA)
|
||||
**20 models generate real video/audio on GPU** — the 7 above + **13 of the newly-ported** archs, each run
|
||||
end-to-end through the real `VideoGenerator` (resolve → stamp → CUDA load → generate). The rest are blocked
|
||||
by a **fastvideo-shared-code / missing-kernel / HF-access** wall, NOT a v2 recipe bug (the v2 recipes are
|
||||
faithful — e.g. cosmos25's DiT+VAE produced finite output; only its Qwen2.5-VL encoder hit a library
|
||||
incompat). All ports also resolve + run end-to-end on the CPU toy backend (`test_bucket_c_ports.py`).
|
||||
|
||||
| GPU status | Models |
|
||||
|---|---|
|
||||
| ✅ **Verified** (real GPU output) | stable_audio (audio), matrixgame2, matrixgame3, gen3c, wan_fun_control, lucy_edit, hunyuangamecraft, hunyuan_video, hunyuan_video15, longcat (13.58B), sfwan22 (2×14B MoE, expert offload), lingbotworld (2×14B, offload), fastwan (TI2V-5B-FullAttn DMD) |
|
||||
| 🚫 fastvideo/env-blocked | **cosmos25** (DiT+VAE ran; Qwen2.5-VL encoder → transformers 5.12.1 incompat in fastvideo); **kandinsky5** (fastvideo registry registers a bare `PipelineConfig`); **hyworld** (fastvideo DiT hardcodes `flash_attn`, not built); **turbowan** 1.3B/i2v + **fastwan** VSA-variants (SLA/VSA sparse-attn params + Triton kernels need nvcc) |
|
||||
| 🚫 access-blocked (HF-gated) | cosmos2, flux2, sd35 (no HF token in this env) |
|
||||
|
||||
To unblock the env-blocked: build `fastvideo-kernel` (SLA/VSA Triton, needs nvcc); pin a fastvideo-compatible
|
||||
`transformers` for the Qwen2.5-VL encoder; add a Kandinsky5 `PipelineConfig` + an SDPA fallback in the
|
||||
hyworld DiT (all fastvideo-side / environment, not v2 recipe work).
|
||||
|
||||
## Newly ported (recipe details)
|
||||
Each resolves through the registry AND runs end-to-end on the CPU toy backend via the public `Engine`
|
||||
path (the `v2/tests/test_bucket_c_ports.py` regression guard), emitting the correct modality artifact.
|
||||
|
||||
**15 net-new architectures** (each a new `TorchComponent` adapter + recipe):
|
||||
- **cosmos2** (Cosmos-Predict2-2B-Video2World) — EDM-Karras denoiser; new `CosmosDenoiseLoop` +
|
||||
`build_karras_sigmas` (the reference port). **cosmos25** (Cosmos-Predict2.5 2B/14B) — flow-match,
|
||||
per-frame plain-sigma timestep, Reason1/Qwen2.5-VL encoder. **gen3c** (GEN3C) — EDM + 82ch pose-buffer.
|
||||
- **hunyuan_video** (+FastHunyuan) — reuses WanDenoiseLoop, dual LLaMA+CLIP encoders, Hunyuan VAE.
|
||||
**hunyuan_video15** (480p/720p). **hunyuangamecraft**, **hyworld** — interactive (camera/action).
|
||||
- **longcat** (T2V/I2V/VC). **kandinsky5** (5.0 T2V Lite).
|
||||
- **sd35** (MMDiT, image, triple-encoder). **flux2** (dev/klein, MMDiT image). **stable_audio** (audio).
|
||||
- **lingbotworld** (camera/Plucker), **matrixgame2**, **matrixgame3** — interactive world models.
|
||||
|
||||
**5 Wan-family variants** (reuse the Wan/Causal arch, new in-package sampler/loop/conditioning):
|
||||
- **turbowan** — rCM few-step (faithful RCMScheduler port), 1.3B/14B T2V + I2V-A14B MoE.
|
||||
- **lucy_edit** — v2v editor (video-VAE-encode node → 96ch DiT input). **wan_fun_control** — control input.
|
||||
- **sfwan22** — Self-Forcing Wan2.2-A14B causal + MoE (i2v + t2v). **fastwan** — DMD 3-step (TI2V-5B-FullAttn
|
||||
loadable; VSA-trained variants + non-strict `to_gate_compress` load are BRINGUP).
|
||||
|
||||
BRINGUP scope per port (documented in each package): GPU load/run; for interactive/world-model archs the
|
||||
action/camera/memory conditioning needs a request-API extension (the t2v/degenerate path is what
|
||||
CPU-verifies); video2world/i2v frame-replace conditioning is threaded but inert without conditioning inputs.
|
||||
|
||||
## Environment
|
||||
v2 bring-up runs **single-GPU, resident, on the `TORCH_SDPA` backend** (no fastvideo-kernel / VSA / FP4).
|
||||
The box has been rescheduled across hosts/arches/python versions mid-session; rebuild the venv for the
|
||||
current arch when that happens: `uv venv --python 3.12 .venv`; comment out `fastvideo-kernel` in
|
||||
`pyproject.toml`; `uv pip install -e ".[dev]"`. Source `/home/scratch.willlin_ent/.bringup_env`
|
||||
(`HF_HOME=./.cache` on scratch, `FASTVIDEO_ATTENTION_BACKEND=TORCH_SDPA`). v2 CPU mini: 240 passed, 2 skipped.
|
||||
|
||||
## How to add a model to the v2 substrate
|
||||
1. `v2/recipes/<arch>/` — card (declare adapters via `ComponentSpec.adapter`; per-model `SamplingDefaults`),
|
||||
loop (reuse `WanDenoiseLoop`/`chunk_rollout` or a new in-package loop+sampler), program.
|
||||
2. `v2/platform/backends/torch_<arch>.py` — a `TorchComponent` subclass (only the forward semantics) if the
|
||||
arch is genuinely new; reuse `WanDiT`/`LTX2DiT`/`WanVAE`/`T5Encoder` via `load_id` when it isn't.
|
||||
3. One row in `v2/registry.py:_BUCKET_C` (HF ids → builders; `transformer_cls` for the arch fallback, or
|
||||
`""` for explicit-id-only capability variants of an existing arch).
|
||||
4. CPU-verify: it resolves + runs on the toy backend (auto-covered by `test_bucket_c_ports.py`). Then GPU
|
||||
bring-up (`stamp_*_checkpoints` → real weights) per BRINGUP notes.
|
||||
@@ -0,0 +1,124 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""Run full Flux2 text-to-image generation through FastVideo.
|
||||
|
||||
User story:
|
||||
"I have a local or HF Diffusers-format full Flux2 checkpoint and want a
|
||||
minimal text-to-image generation command that uses embedded guidance."
|
||||
"""
|
||||
import argparse
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
ComponentConfig,
|
||||
EngineConfig,
|
||||
GenerationRequest,
|
||||
GeneratorConfig,
|
||||
OffloadConfig,
|
||||
OutputConfig,
|
||||
ParallelismConfig,
|
||||
PipelineSelection,
|
||||
SamplingConfig,
|
||||
)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run full Flux2 text-to-image generation.")
|
||||
parser.add_argument(
|
||||
"--model-path",
|
||||
default="black-forest-labs/FLUX.2-dev",
|
||||
help="HF id or local diffusers-format full Flux2 weights directory.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output",
|
||||
default="outputs/flux2/flux2.png",
|
||||
help="Output PNG path.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--prompt",
|
||||
default="a photo of a banana on a wooden table, studio lighting",
|
||||
help="Text prompt.",
|
||||
)
|
||||
parser.add_argument("--height", type=int, default=1024)
|
||||
parser.add_argument("--width", type=int, default=1024)
|
||||
parser.add_argument("--steps", type=int, default=50)
|
||||
parser.add_argument("--guidance-scale", type=float, default=4.0)
|
||||
parser.add_argument("--max-sequence-length", type=int, default=None)
|
||||
parser.add_argument("--seed", type=int, default=0)
|
||||
parser.add_argument("--num-gpus", type=int, default=1)
|
||||
parser.add_argument("--tp-size", type=int, default=None)
|
||||
parser.add_argument("--sp-size", type=int, default=None)
|
||||
parser.add_argument(
|
||||
"--backend",
|
||||
default=None,
|
||||
help="Set FASTVIDEO_ATTENTION_BACKEND, for example TORCH_SDPA.",
|
||||
)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
if args.backend:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
|
||||
|
||||
output = Path(args.output)
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
tp_size = args.tp_size if args.tp_size is not None else (
|
||||
args.num_gpus if args.num_gpus > 1 else 1
|
||||
)
|
||||
sp_size = args.sp_size if args.sp_size is not None else (
|
||||
1 if args.num_gpus > 1 else args.num_gpus
|
||||
)
|
||||
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=args.model_path,
|
||||
engine=EngineConfig(
|
||||
num_gpus=args.num_gpus,
|
||||
parallelism=ParallelismConfig(tp_size=tp_size, sp_size=sp_size),
|
||||
use_fsdp_inference=False,
|
||||
offload=OffloadConfig(
|
||||
dit=False,
|
||||
vae=True,
|
||||
text_encoder=True,
|
||||
pin_cpu_memory=False,
|
||||
),
|
||||
),
|
||||
pipeline=PipelineSelection(
|
||||
workload_type="t2i",
|
||||
components=ComponentConfig(override_pipeline_cls_name="Flux2Pipeline"),
|
||||
),
|
||||
)
|
||||
|
||||
generator = VideoGenerator.from_config(generator_config)
|
||||
try:
|
||||
sampling = SamplingConfig(
|
||||
height=args.height,
|
||||
width=args.width,
|
||||
num_frames=1,
|
||||
fps=1,
|
||||
num_inference_steps=args.steps,
|
||||
guidance_scale=args.guidance_scale,
|
||||
seed=args.seed,
|
||||
)
|
||||
extensions = {}
|
||||
if args.max_sequence_length is not None:
|
||||
extensions["max_sequence_length"] = args.max_sequence_length
|
||||
|
||||
request = GenerationRequest(
|
||||
prompt=args.prompt,
|
||||
sampling=sampling,
|
||||
output=OutputConfig(
|
||||
output_path=str(output),
|
||||
save_video=True,
|
||||
return_frames=False,
|
||||
),
|
||||
extensions=extensions,
|
||||
)
|
||||
generator.generate(request)
|
||||
finally:
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,98 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""Run Flux2 Klein text-to-image generation through FastVideo.
|
||||
|
||||
User story:
|
||||
"I need a short local smoke for the Flux2 Klein checkpoint before wiring it
|
||||
into an image workflow. Use the model's distilled four-step defaults and
|
||||
write a single PNG so I can compare the output against the reference."
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import os
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
EngineConfig,
|
||||
GenerationRequest,
|
||||
GeneratorConfig,
|
||||
OffloadConfig,
|
||||
OutputConfig,
|
||||
PipelineSelection,
|
||||
SamplingConfig,
|
||||
)
|
||||
|
||||
|
||||
DEFAULT_PROMPT = "a brushed steel espresso machine on a marble counter, morning window light"
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run Flux2 Klein text-to-image generation.")
|
||||
parser.add_argument(
|
||||
"--model-path",
|
||||
default="black-forest-labs/FLUX.2-klein-4B",
|
||||
help="HF id or local diffusers-format Flux2 Klein weights directory.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output-path",
|
||||
default="outputs/flux2/flux2_klein.png",
|
||||
help="PNG output path or output directory.",
|
||||
)
|
||||
parser.add_argument("--prompt", default=DEFAULT_PROMPT, help="Prompt text.")
|
||||
parser.add_argument("--seed", type=int, default=0, help="Generation seed.")
|
||||
parser.add_argument("--height", type=int, default=1024, help="Output image height.")
|
||||
parser.add_argument("--width", type=int, default=1024, help="Output image width.")
|
||||
parser.add_argument("--steps", type=int, default=4, help="Number of denoising steps.")
|
||||
parser.add_argument("--num-gpus", type=int, default=1, help="Number of GPUs to use.")
|
||||
parser.add_argument(
|
||||
"--backend",
|
||||
default=None,
|
||||
help="Set FASTVIDEO_ATTENTION_BACKEND, for example TORCH_SDPA.",
|
||||
)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
args = parse_args()
|
||||
if args.backend:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = args.backend
|
||||
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=args.model_path,
|
||||
engine=EngineConfig(
|
||||
num_gpus=args.num_gpus,
|
||||
use_fsdp_inference=False,
|
||||
offload=OffloadConfig(
|
||||
dit=False,
|
||||
vae=True,
|
||||
text_encoder=True,
|
||||
pin_cpu_memory=False,
|
||||
),
|
||||
),
|
||||
pipeline=PipelineSelection(workload_type="t2i"),
|
||||
)
|
||||
|
||||
generator = VideoGenerator.from_config(generator_config)
|
||||
try:
|
||||
request = GenerationRequest(
|
||||
prompt=args.prompt,
|
||||
sampling=SamplingConfig(
|
||||
height=args.height,
|
||||
width=args.width,
|
||||
num_frames=1,
|
||||
fps=1,
|
||||
num_inference_steps=args.steps,
|
||||
guidance_scale=1.0,
|
||||
seed=args.seed,
|
||||
),
|
||||
output=OutputConfig(
|
||||
output_path=args.output_path,
|
||||
save_video=True,
|
||||
),
|
||||
)
|
||||
generator.generate(request)
|
||||
finally:
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,289 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""LTX-2.3 distilled image-to-video with torch.compile + timing breakdown.
|
||||
|
||||
This example runs the LTX-2.3 distilled student model on a single GPU with
|
||||
torch.compile fully enabled, then prints a per-stage timing breakdown so the
|
||||
user can see where wall-time goes. It is meant as the canonical entry point
|
||||
for trying out the LTX-2.3 i2v path on `hao-ai-lab/FastVideo:main`.
|
||||
|
||||
Quick start
|
||||
-----------
|
||||
export LTX23_I2V_IMAGE=/path/to/your/portrait_or_product.jpg
|
||||
# optional overrides:
|
||||
# export LTX23_I2V_PROMPT="a fashion model walks toward camera..."
|
||||
# export LTX23_OUTPUT_DIR=outputs_video/ltx2_3_distilled_i2v
|
||||
python examples/inference/basic/basic_ltx2_3_distilled_i2v.py
|
||||
|
||||
What the script does
|
||||
--------------------
|
||||
1. Loads FastVideo/LTX-2.3-Distilled-Diffusers (8 denoise + 3 refine steps,
|
||||
CFG=1, no refine LoRA — the distilled production recipe).
|
||||
2. Compiles the DiT, text encoder, and VAE (fullgraph, Inductor default
|
||||
mode — autotune adds ~7 min cold-compile here with no measurable
|
||||
e2e gain).
|
||||
3. Runs 2 warmup calls (untimed) + 2 measured calls. Two warmups are kept
|
||||
as a safety net — the first call pays cold compile + first-shape guard
|
||||
work, and a second warmup ensures any residual recompiles settle before
|
||||
we measure.
|
||||
4. Prints a per-stage breakdown and an average over the measured runs.
|
||||
|
||||
Hardware notes
|
||||
--------------
|
||||
- Single-GPU example; for multi-GPU sequence-parallel see the gradio demo
|
||||
under `examples/inference/gradio/local/gradio_local_demo_ltx2_3/`.
|
||||
- First-time compile takes ~30-40 min on GB200 (~20 min on H100; cached
|
||||
in `$TORCHINDUCTOR_CACHE_DIR` afterwards). Subsequent invocations only
|
||||
pay the one-time process load + a few seconds of dynamo trace.
|
||||
- On GB200 / Blackwell, run with `env -u LD_LIBRARY_PATH ...` to avoid a
|
||||
system-cuBLAS / torch-cuBLAS mismatch that fails every GEMM. The
|
||||
`_inductor.shape_padding = False` line below also avoids a pad_mm
|
||||
landmine on the same generation of cards.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import time
|
||||
from collections import OrderedDict
|
||||
from pathlib import Path
|
||||
|
||||
import torch._inductor.config as _inductor
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.configs.pipelines.base import PipelineConfig
|
||||
from fastvideo.utils import maybe_download_model
|
||||
|
||||
# Env knobs (set BEFORE importing fastvideo where possible — but
|
||||
# FASTVIDEO_ATTENTION_BACKEND is fine here because the worker reads it
|
||||
# on generator construction).
|
||||
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "FLASH_ATTN")
|
||||
os.environ.setdefault("FASTVIDEO_STAGE_LOGGING", "1")
|
||||
|
||||
# Inductor knobs. The first one (shape_padding=False) is mandatory on
|
||||
# Blackwell to avoid a cuBLAS INVALID_VALUE crash inside pad_mm during
|
||||
# the refine path. The rest are autotune-friendliness flags.
|
||||
_inductor.shape_padding = False
|
||||
_inductor.conv_1x1_as_mm = True
|
||||
_inductor.coordinate_descent_tuning = True
|
||||
_inductor.coordinate_descent_check_all_directions = True
|
||||
_inductor.epilogue_fusion = False
|
||||
|
||||
MODEL_ID = os.path.expandvars(
|
||||
os.path.expanduser(
|
||||
os.getenv("LTX23_MODEL_PATH", "FastVideo/LTX-2.3-Distilled-Diffusers")
|
||||
)
|
||||
)
|
||||
OUTPUT_DIR = Path(
|
||||
os.getenv("LTX23_OUTPUT_DIR", "outputs_video/ltx2_3_distilled_i2v")
|
||||
)
|
||||
I2V_IMAGE = os.getenv("LTX23_I2V_IMAGE", "")
|
||||
DEFAULT_PROMPT = (
|
||||
"A fashion model takes a slow step forward and shifts her weight, "
|
||||
"the soft fabric of her clothing swaying and rippling with the "
|
||||
"motion, her hair shifting gently, soft even studio lighting on a "
|
||||
"clean light background, elegant slow-motion runway feel."
|
||||
)
|
||||
PROMPT = os.getenv("LTX23_I2V_PROMPT", DEFAULT_PROMPT)
|
||||
|
||||
# Per-stage timing helpers --------------------------------------------------
|
||||
|
||||
def _print_stage_breakdown(result: dict, label: str) -> float | None:
|
||||
"""Print stage execution times and return the sum, or None if missing."""
|
||||
logging_info = result.get("logging_info")
|
||||
stages = getattr(logging_info, "stages", None) if logging_info else None
|
||||
if not stages:
|
||||
print(f" [{label}] stage breakdown unavailable")
|
||||
return None
|
||||
print(f" [{label}] stage breakdown:")
|
||||
total = 0.0
|
||||
for name, metrics in stages.items():
|
||||
exec_s = float(metrics.get("execution_time", 0.0))
|
||||
total += exec_s
|
||||
print(f" - {name}: {exec_s:.3f}s")
|
||||
print(f" - stage_sum: {total:.3f}s")
|
||||
return total
|
||||
|
||||
|
||||
def _collect_stage_times(
|
||||
result: dict,
|
||||
stage_times: dict[str, list[float]],
|
||||
stage_order: OrderedDict[str, None],
|
||||
) -> None:
|
||||
logging_info = result.get("logging_info")
|
||||
stages = getattr(logging_info, "stages", None) if logging_info else None
|
||||
if not stages:
|
||||
return
|
||||
for name, metrics in stages.items():
|
||||
stage_order.setdefault(name, None)
|
||||
stage_times.setdefault(name, []).append(
|
||||
float(metrics.get("execution_time", 0.0))
|
||||
)
|
||||
|
||||
|
||||
def _resolve_refine_upsampler(model_root: str) -> Path:
|
||||
"""LTX-2.3 distilled snapshots ship a `spatial_upscaler/` subdir."""
|
||||
for name in ("spatial_upscaler", "spatial_upsampler"):
|
||||
cand = Path(model_root) / name
|
||||
if (cand / "config.json").is_file():
|
||||
return cand
|
||||
raise FileNotFoundError(
|
||||
f"No refine upsampler directory under {model_root}. "
|
||||
f"Expected `{model_root}/spatial_upscaler/config.json`."
|
||||
)
|
||||
|
||||
|
||||
# Main ---------------------------------------------------------------------
|
||||
|
||||
def main() -> None:
|
||||
if not I2V_IMAGE:
|
||||
raise SystemExit(
|
||||
"LTX23_I2V_IMAGE is required for i2v. Example:\n"
|
||||
" export LTX23_I2V_IMAGE=/path/to/portrait_or_product.jpg\n"
|
||||
" python examples/inference/basic/basic_ltx2_3_distilled_i2v.py"
|
||||
)
|
||||
if not Path(I2V_IMAGE).is_file():
|
||||
raise SystemExit(f"LTX23_I2V_IMAGE not found: {I2V_IMAGE}")
|
||||
|
||||
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||
model_root = maybe_download_model(MODEL_ID)
|
||||
refine_upsampler_path = _resolve_refine_upsampler(model_root)
|
||||
print(f"Model: {model_root}")
|
||||
print(f"Refine upsampler: {refine_upsampler_path}")
|
||||
print(f"i2v image: {I2V_IMAGE}")
|
||||
print(f"Output dir: {OUTPUT_DIR.resolve()}")
|
||||
|
||||
# mode="default" — Inductor's default schedule matches max-autotune on
|
||||
# this pipeline (denoise/refine/decode all within ~5 ms, n=2) while
|
||||
# saving ~7 min of cold compile on a single GB200.
|
||||
torch_compile_kwargs = {
|
||||
"backend": "inductor",
|
||||
"fullgraph": True,
|
||||
"mode": "default",
|
||||
"dynamic": False,
|
||||
}
|
||||
|
||||
# Loading the pipeline config *with model_path* binds model-specific
|
||||
# tuning (notably VAE precision/decoder defaults) into the config. Without
|
||||
# this, the generic pipeline config gives a substantially slower VAE
|
||||
# decode stage. `basic_ltx2_distilled_fast_profile.py` uses the same
|
||||
# pattern.
|
||||
pipeline_config = PipelineConfig.from_pretrained(model_root)
|
||||
pipeline_config.dit_config.quant_config = None
|
||||
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_root,
|
||||
num_gpus=1,
|
||||
# LTX-2.3 distilled uses the two-stage refine pipeline; the refine
|
||||
# LoRA is intentionally empty for the distilled student.
|
||||
ltx2_refine_enabled=True,
|
||||
ltx2_refine_upsampler_path=str(refine_upsampler_path),
|
||||
ltx2_refine_lora_path="",
|
||||
ltx2_refine_num_inference_steps=3,
|
||||
ltx2_refine_guidance_scale=1.0,
|
||||
ltx2_refine_add_noise=True,
|
||||
pipeline_config=pipeline_config,
|
||||
enable_torch_compile=True,
|
||||
enable_torch_compile_text_encoder=True,
|
||||
# Compile the VAE codec submodules (encoder / decoder) too. The
|
||||
# `LTX2CausalVideoAutoencoder` declares `_compile_conditions` so
|
||||
# `_compile_with_conditions` targets just those submodules and
|
||||
# leaves the surrounding tiling control flow eager — needed for
|
||||
# fullgraph + dynamic=False to succeed. VAE eager decode is
|
||||
# ~1.0s; compiling it brings the stage to ~0.3s.
|
||||
enable_torch_compile_vae=True,
|
||||
torch_compile_kwargs=torch_compile_kwargs,
|
||||
torch_compile_kwargs_vae=torch_compile_kwargs,
|
||||
# Keep everything resident — no CPU offload for serving-style runs.
|
||||
dit_cpu_offload=False,
|
||||
text_encoder_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
ltx2_vae_tiling=False,
|
||||
)
|
||||
|
||||
common_kwargs = dict(
|
||||
prompt=PROMPT,
|
||||
negative_prompt="", # distilled is CFG-free; no negative needed
|
||||
guidance_scale=1.0, # CFG=1 for distilled
|
||||
height=1280, width=832, # portrait runway aspect
|
||||
num_frames=121, fps=24, # ~5s clip
|
||||
num_inference_steps=8, # distilled denoise steps
|
||||
# i2v: anchor the input image at frame 0 with full strength.
|
||||
# `ltx2_image_crf=0.0` skips an extra JPEG re-encode of an already
|
||||
# JPEG conditioning image.
|
||||
ltx2_images=[(I2V_IMAGE, 0, 1.0)],
|
||||
ltx2_image_crf=0.0,
|
||||
save_video=True,
|
||||
)
|
||||
|
||||
warmup_runs = 2
|
||||
measured_runs = 2
|
||||
warmup_secs: list[float] = []
|
||||
measured_secs: list[float] = []
|
||||
stage_times: dict[str, list[float]] = {}
|
||||
stage_order: OrderedDict[str, None] = OrderedDict()
|
||||
|
||||
try:
|
||||
# Warmup: untimed (but we still wall-clock them so the first compile
|
||||
# cost is visible to the reader).
|
||||
for w in range(warmup_runs):
|
||||
t0 = time.perf_counter()
|
||||
print(f"\n[warmup {w + 1}/{warmup_runs}] compiling + generating…")
|
||||
generator.generate_video(
|
||||
output_path=str(OUTPUT_DIR / f"_warmup_{w + 1}.mp4"),
|
||||
seed=7,
|
||||
**common_kwargs,
|
||||
)
|
||||
dt = time.perf_counter() - t0
|
||||
warmup_secs.append(dt)
|
||||
print(f"[warmup {w + 1}/{warmup_runs}] wall={dt:.1f}s")
|
||||
|
||||
# Cleanup warmup artifacts so the user only sees measured outputs.
|
||||
for w in range(warmup_runs):
|
||||
(OUTPUT_DIR / f"_warmup_{w + 1}.mp4").unlink(missing_ok=True)
|
||||
|
||||
# Measured.
|
||||
for m in range(measured_runs):
|
||||
out_path = OUTPUT_DIR / f"output_ltx2_3_distilled_i2v_run_{m + 1}.mp4"
|
||||
print(f"\n[measured {m + 1}/{measured_runs}] generating: {out_path}")
|
||||
t0 = time.perf_counter()
|
||||
result = generator.generate_video(
|
||||
output_path=str(out_path),
|
||||
seed=2002 + m,
|
||||
**common_kwargs,
|
||||
)
|
||||
wall = time.perf_counter() - t0
|
||||
e2e = (
|
||||
result.get("e2e_latency")
|
||||
if isinstance(result, dict) else None
|
||||
) or wall
|
||||
measured_secs.append(e2e)
|
||||
print(f"[measured {m + 1}/{measured_runs}] e2e={e2e:.2f}s wall={wall:.2f}s")
|
||||
if isinstance(result, dict):
|
||||
_print_stage_breakdown(result, f"measured {m + 1}")
|
||||
_collect_stage_times(result, stage_times, stage_order)
|
||||
|
||||
# Summary.
|
||||
print("\n=== summary ===")
|
||||
print(f"warmup wall-times: {[round(x, 1) for x in warmup_secs]}")
|
||||
if measured_secs:
|
||||
avg = sum(measured_secs) / len(measured_secs)
|
||||
print(
|
||||
f"measured e2e (n={len(measured_secs)}): "
|
||||
f"{[round(x, 2) for x in measured_secs]} -> avg {avg:.2f}s"
|
||||
)
|
||||
if stage_times:
|
||||
print(f"average stage times over {measured_runs} measured runs:")
|
||||
avg_total = 0.0
|
||||
for name in stage_order:
|
||||
vals = stage_times.get(name) or []
|
||||
if not vals:
|
||||
continue
|
||||
avg_v = sum(vals) / len(vals)
|
||||
avg_total += avg_v
|
||||
print(f" - {name}: {avg_v:.3f}s")
|
||||
print(f" - stage_sum_avg: {avg_total:.3f}s")
|
||||
finally:
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,350 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""LTX-2.3 distilled image-to-video — typed API (``from_config`` / ``generate``).
|
||||
|
||||
Identical generation behavior to ``basic_ltx2_3_distilled_i2v.py``, but
|
||||
expressed through the newer typed surface (``GeneratorConfig`` /
|
||||
``GenerationRequest``) instead of the ``from_pretrained(**legacy_kwargs)``
|
||||
bridge. The typed API is now the preferred entry point — the legacy
|
||||
example still works but emits a ``DeprecationWarning`` for the LTX-2.3
|
||||
specific knobs.
|
||||
|
||||
Quick start
|
||||
-----------
|
||||
export LTX23_I2V_IMAGE=/path/to/your/portrait_or_product.jpg
|
||||
# optional overrides:
|
||||
# export LTX23_I2V_PROMPT="a fashion model walks toward camera..."
|
||||
# export LTX23_OUTPUT_DIR=outputs_video/ltx2_3_distilled_i2v_typed
|
||||
python examples/inference/basic/basic_ltx2_3_distilled_i2v_typed.py
|
||||
|
||||
What the script does
|
||||
--------------------
|
||||
1. Loads FastVideo/LTX-2.3-Distilled-Diffusers (8 denoise + 3 refine
|
||||
steps, CFG=1, no refine LoRA — the distilled production recipe).
|
||||
2. Compiles the DiT, text encoder, and VAE (fullgraph, Inductor default
|
||||
mode — autotune adds ~7 min cold-compile here with no measurable
|
||||
e2e gain).
|
||||
3. Runs 2 warmup calls (untimed) + 2 measured calls. Two warmups are
|
||||
kept as a safety net — the first call pays cold compile + first-shape
|
||||
guard work, and a second warmup ensures any residual recompiles
|
||||
settle before we measure.
|
||||
4. Prints a per-stage breakdown and an average over the measured runs.
|
||||
|
||||
Hardware notes
|
||||
--------------
|
||||
- Single-GPU example; for multi-GPU sequence-parallel see the gradio
|
||||
demo under ``examples/inference/gradio/local/gradio_local_demo_ltx2_3/``.
|
||||
- First-time compile takes ~30-40 min on GB200 (~20 min on H100;
|
||||
cached in ``$TORCHINDUCTOR_CACHE_DIR`` afterwards). Subsequent
|
||||
invocations only pay the one-time process load + a few seconds of
|
||||
dynamo trace.
|
||||
- On GB200 / Blackwell, run with ``env -u LD_LIBRARY_PATH ...`` to
|
||||
avoid a system-cuBLAS / torch-cuBLAS mismatch that fails every GEMM.
|
||||
The ``_inductor.shape_padding = False`` line below also avoids a
|
||||
``pad_mm`` landmine on the same generation of cards.
|
||||
|
||||
Typed-API mapping (legacy kwarg ↔ typed field)
|
||||
----------------------------------------------
|
||||
- ``num_gpus`` ↔ ``engine.num_gpus``
|
||||
- ``enable_torch_compile`` ↔ ``engine.compile.enabled``
|
||||
- ``enable_torch_compile_text_encoder`` ↔ ``engine.compile.text_encoder_enabled``
|
||||
- ``enable_torch_compile_vae`` ↔ ``engine.compile.vae_enabled``
|
||||
- ``torch_compile_kwargs`` ↔ ``engine.compile.backend/fullgraph/mode/dynamic``
|
||||
- ``torch_compile_kwargs_vae`` ↔ empty ``compile.vae_kwargs`` (inherits master)
|
||||
- ``dit_cpu_offload`` ↔ ``engine.offload.dit``
|
||||
- ``text_encoder_cpu_offload`` ↔ ``engine.offload.text_encoder``
|
||||
- ``vae_cpu_offload`` ↔ ``engine.offload.vae``
|
||||
- ``ltx2_vae_tiling`` ↔ ``pipeline.vae_tiling``
|
||||
- ``ltx2_refine_enabled`` ↔ ``pipeline.preset_overrides["refine"]["enabled"]``
|
||||
- ``ltx2_refine_upsampler_path`` ↔ ``pipeline.components.upsampler_weights``
|
||||
- ``ltx2_refine_lora_path`` ↔ ``pipeline.components.lora_path``
|
||||
- ``ltx2_refine_num_inference_steps`` ↔ ``pipeline.preset_overrides["refine"]["num_inference_steps"]``
|
||||
- ``ltx2_refine_guidance_scale`` ↔ ``pipeline.preset_overrides["refine"]["guidance_scale"]``
|
||||
- ``ltx2_refine_add_noise`` ↔ ``pipeline.preset_overrides["refine"]["add_noise"]``
|
||||
- ``pipeline_config=PipelineConfig.from_pretrained(model_root)`` ↔ (no-op — ``PipelineConfig.from_kwargs`` already resolves the model-specific class from ``model_path``)
|
||||
- ``pipeline_config.dit_config.quant_config = None`` ↔ leave ``engine.quantization`` unset
|
||||
- ``ltx2_images`` / ``ltx2_image_crf`` ↔ ``request.extensions`` (LTX-2 specific, no
|
||||
first-class typed field yet)
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import time
|
||||
from collections import OrderedDict
|
||||
from pathlib import Path
|
||||
|
||||
import torch._inductor.config as _inductor
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
CompileConfig,
|
||||
ComponentConfig,
|
||||
EngineConfig,
|
||||
GenerationRequest,
|
||||
GeneratorConfig,
|
||||
OffloadConfig,
|
||||
OutputConfig,
|
||||
PipelineSelection,
|
||||
SamplingConfig,
|
||||
)
|
||||
from fastvideo.utils import maybe_download_model
|
||||
|
||||
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "FLASH_ATTN")
|
||||
os.environ.setdefault("FASTVIDEO_STAGE_LOGGING", "1")
|
||||
|
||||
# Inductor knobs. ``shape_padding=False`` is mandatory on Blackwell to
|
||||
# avoid a cuBLAS INVALID_VALUE crash inside pad_mm during the refine
|
||||
# path. The rest are autotune-friendliness flags.
|
||||
_inductor.shape_padding = False
|
||||
_inductor.conv_1x1_as_mm = True
|
||||
_inductor.coordinate_descent_tuning = True
|
||||
_inductor.coordinate_descent_check_all_directions = True
|
||||
_inductor.epilogue_fusion = False
|
||||
|
||||
MODEL_ID = os.path.expandvars(
|
||||
os.path.expanduser(
|
||||
os.getenv("LTX23_MODEL_PATH", "FastVideo/LTX-2.3-Distilled-Diffusers")
|
||||
)
|
||||
)
|
||||
OUTPUT_DIR = Path(
|
||||
os.getenv(
|
||||
"LTX23_OUTPUT_DIR", "outputs_video/ltx2_3_distilled_i2v_typed"
|
||||
)
|
||||
)
|
||||
I2V_IMAGE = os.getenv("LTX23_I2V_IMAGE", "")
|
||||
DEFAULT_PROMPT = (
|
||||
"A fashion model takes a slow step forward and shifts her weight, "
|
||||
"the soft fabric of her clothing swaying and rippling with the "
|
||||
"motion, her hair shifting gently, soft even studio lighting on a "
|
||||
"clean light background, elegant slow-motion runway feel."
|
||||
)
|
||||
PROMPT = os.getenv("LTX23_I2V_PROMPT", DEFAULT_PROMPT)
|
||||
|
||||
|
||||
def _print_stage_breakdown(result, label: str) -> float | None:
|
||||
logging_info = getattr(result, "logging_info", None)
|
||||
stages = getattr(logging_info, "stages", None) if logging_info else None
|
||||
if not stages:
|
||||
print(f" [{label}] stage breakdown unavailable")
|
||||
return None
|
||||
print(f" [{label}] stage breakdown:")
|
||||
total = 0.0
|
||||
for name, metrics in stages.items():
|
||||
exec_s = float(metrics.get("execution_time", 0.0))
|
||||
total += exec_s
|
||||
print(f" - {name}: {exec_s:.3f}s")
|
||||
print(f" - stage_sum: {total:.3f}s")
|
||||
return total
|
||||
|
||||
|
||||
def _collect_stage_times(
|
||||
result,
|
||||
stage_times: dict[str, list[float]],
|
||||
stage_order: OrderedDict[str, None],
|
||||
) -> None:
|
||||
logging_info = getattr(result, "logging_info", None)
|
||||
stages = getattr(logging_info, "stages", None) if logging_info else None
|
||||
if not stages:
|
||||
return
|
||||
for name, metrics in stages.items():
|
||||
stage_order.setdefault(name, None)
|
||||
stage_times.setdefault(name, []).append(
|
||||
float(metrics.get("execution_time", 0.0))
|
||||
)
|
||||
|
||||
|
||||
def _resolve_refine_upsampler(model_root: str) -> Path:
|
||||
for name in ("spatial_upscaler", "spatial_upsampler"):
|
||||
cand = Path(model_root) / name
|
||||
if (cand / "config.json").is_file():
|
||||
return cand
|
||||
raise FileNotFoundError(
|
||||
f"No refine upsampler directory under {model_root}. "
|
||||
f"Expected `{model_root}/spatial_upscaler/config.json`."
|
||||
)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
if not I2V_IMAGE:
|
||||
raise SystemExit(
|
||||
"LTX23_I2V_IMAGE is required for i2v. Example:\n"
|
||||
" export LTX23_I2V_IMAGE=/path/to/portrait_or_product.jpg\n"
|
||||
" python examples/inference/basic/"
|
||||
"basic_ltx2_3_distilled_i2v_typed.py"
|
||||
)
|
||||
if not Path(I2V_IMAGE).is_file():
|
||||
raise SystemExit(f"LTX23_I2V_IMAGE not found: {I2V_IMAGE}")
|
||||
|
||||
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||
model_root = maybe_download_model(MODEL_ID)
|
||||
refine_upsampler_path = _resolve_refine_upsampler(model_root)
|
||||
print(f"Model: {model_root}")
|
||||
print(f"Refine upsampler: {refine_upsampler_path}")
|
||||
print(f"i2v image: {I2V_IMAGE}")
|
||||
print(f"Output dir: {OUTPUT_DIR.resolve()}")
|
||||
|
||||
# mode="default" — Inductor's default schedule matches max-autotune on
|
||||
# this pipeline (denoise/refine/decode all within ~5 ms, n=2) while
|
||||
# saving ~7 min of cold compile on a single GB200.
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=model_root,
|
||||
engine=EngineConfig(
|
||||
num_gpus=1,
|
||||
# Keep DiT / text encoder / VAE resident on GPU — no CPU offload
|
||||
# for serving-style runs. ``image_encoder`` and
|
||||
# ``pin_cpu_memory`` are left at their schema defaults
|
||||
# (matches the legacy example, which only set these three).
|
||||
offload=OffloadConfig(
|
||||
dit=False,
|
||||
text_encoder=False,
|
||||
vae=False,
|
||||
),
|
||||
compile=CompileConfig(
|
||||
enabled=True,
|
||||
text_encoder_enabled=True,
|
||||
# ``vae_enabled`` triggers ``_compile_with_conditions`` on
|
||||
# ``LTX2CausalVideoAutoencoder``, which compiles just the
|
||||
# encoder/decoder submodules and leaves the surrounding
|
||||
# tiling control flow eager (required for ``fullgraph``).
|
||||
# Empty ``vae_kwargs`` → inherits the master kwargs below.
|
||||
vae_enabled=True,
|
||||
backend="inductor",
|
||||
fullgraph=True,
|
||||
mode="default",
|
||||
dynamic=False,
|
||||
),
|
||||
),
|
||||
pipeline=PipelineSelection(
|
||||
# ``PipelineConfig.from_kwargs`` resolves the model-specific
|
||||
# pipeline-config class from ``model_path`` automatically, so we
|
||||
# don't need to set ``components.pipeline_config_path`` — the
|
||||
# model-specific VAE precision / decoder defaults are picked up
|
||||
# the same way the legacy example's
|
||||
# ``PipelineConfig.from_pretrained(model_root)`` did them.
|
||||
components=ComponentConfig(
|
||||
upsampler_weights=str(refine_upsampler_path),
|
||||
# Distilled has no refine LoRA — omit ``lora_path``.
|
||||
),
|
||||
vae_tiling=False,
|
||||
preset_overrides={
|
||||
"refine": {
|
||||
"enabled": True,
|
||||
"num_inference_steps": 3,
|
||||
"guidance_scale": 1.0,
|
||||
"add_noise": True,
|
||||
},
|
||||
},
|
||||
),
|
||||
)
|
||||
|
||||
generator = VideoGenerator.from_config(generator_config)
|
||||
|
||||
def build_request(out_path: Path, seed: int) -> GenerationRequest:
|
||||
return GenerationRequest(
|
||||
prompt=PROMPT,
|
||||
# distilled is CFG-free; no negative prompt
|
||||
negative_prompt="",
|
||||
sampling=SamplingConfig(
|
||||
num_videos_per_prompt=1,
|
||||
seed=seed,
|
||||
height=1280,
|
||||
width=832,
|
||||
num_frames=121,
|
||||
fps=24,
|
||||
num_inference_steps=8,
|
||||
guidance_scale=1.0,
|
||||
),
|
||||
output=OutputConfig(
|
||||
output_path=str(out_path),
|
||||
save_video=True,
|
||||
return_frames=False,
|
||||
),
|
||||
# LTX-2.3 i2v fields don't have first-class typed slots yet;
|
||||
# extensions is the documented bridge. ``ltx2_image_crf=0.0``
|
||||
# skips an extra JPEG re-encode of an already JPEG image.
|
||||
extensions={
|
||||
"ltx2_images": [(I2V_IMAGE, 0, 1.0)],
|
||||
"ltx2_image_crf": 0.0,
|
||||
},
|
||||
)
|
||||
|
||||
warmup_runs = 2
|
||||
measured_runs = 2
|
||||
warmup_secs: list[float] = []
|
||||
measured_secs: list[float] = []
|
||||
stage_times: dict[str, list[float]] = {}
|
||||
stage_order: OrderedDict[str, None] = OrderedDict()
|
||||
|
||||
try:
|
||||
for w in range(warmup_runs):
|
||||
print(f"\n[warmup {w + 1}/{warmup_runs}] compiling + generating…")
|
||||
t0 = time.perf_counter()
|
||||
generator.generate(
|
||||
build_request(
|
||||
OUTPUT_DIR / f"_warmup_{w + 1}.mp4", seed=7
|
||||
)
|
||||
)
|
||||
dt = time.perf_counter() - t0
|
||||
warmup_secs.append(dt)
|
||||
print(f"[warmup {w + 1}/{warmup_runs}] wall={dt:.1f}s")
|
||||
|
||||
for w in range(warmup_runs):
|
||||
(OUTPUT_DIR / f"_warmup_{w + 1}.mp4").unlink(missing_ok=True)
|
||||
|
||||
for m in range(measured_runs):
|
||||
out_path = (
|
||||
OUTPUT_DIR
|
||||
/ f"output_ltx2_3_distilled_i2v_typed_run_{m + 1}.mp4"
|
||||
)
|
||||
print(
|
||||
f"\n[measured {m + 1}/{measured_runs}] generating: {out_path}"
|
||||
)
|
||||
t0 = time.perf_counter()
|
||||
result = generator.generate(
|
||||
build_request(out_path, seed=2002 + m)
|
||||
)
|
||||
wall = time.perf_counter() - t0
|
||||
# ``e2e_latency`` is currently surfaced via ``result.extra``;
|
||||
# ``GenerationResult`` exposes ``generation_time`` as a
|
||||
# first-class field but the LTX-2 pipeline only fills the
|
||||
# legacy ``e2e_latency`` key. Prefer the explicit one, fall
|
||||
# back to wall-clock.
|
||||
e2e = (
|
||||
result.extra.get("e2e_latency")
|
||||
if hasattr(result, "extra") else None
|
||||
) or wall
|
||||
measured_secs.append(e2e)
|
||||
print(
|
||||
f"[measured {m + 1}/{measured_runs}] "
|
||||
f"e2e={e2e:.2f}s wall={wall:.2f}s"
|
||||
)
|
||||
_print_stage_breakdown(result, f"measured {m + 1}")
|
||||
_collect_stage_times(result, stage_times, stage_order)
|
||||
|
||||
print("\n=== summary ===")
|
||||
print(
|
||||
f"warmup wall-times: "
|
||||
f"{[round(x, 1) for x in warmup_secs]}"
|
||||
)
|
||||
if measured_secs:
|
||||
avg = sum(measured_secs) / len(measured_secs)
|
||||
print(
|
||||
f"measured e2e (n={len(measured_secs)}): "
|
||||
f"{[round(x, 2) for x in measured_secs]} -> avg {avg:.2f}s"
|
||||
)
|
||||
if stage_times:
|
||||
print(f"average stage times over {measured_runs} measured runs:")
|
||||
avg_total = 0.0
|
||||
for name in stage_order:
|
||||
vals = stage_times.get(name) or []
|
||||
if not vals:
|
||||
continue
|
||||
avg_v = sum(vals) / len(vals)
|
||||
avg_total += avg_v
|
||||
print(f" - {name}: {avg_v:.3f}s")
|
||||
print(f" - stage_sum_avg: {avg_total:.3f}s")
|
||||
finally:
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,38 @@
|
||||
from fastvideo import VideoGenerator
|
||||
|
||||
OUTPUT_PATH = "video_samples_lucy_edit"
|
||||
|
||||
|
||||
def main():
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"decart-ai/Lucy-Edit-Dev",
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
dit_cpu_offload=True,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=True,
|
||||
pin_cpu_memory=True,
|
||||
)
|
||||
|
||||
prompt = ("Change the apron and blouse to a classic clown costume: satin "
|
||||
"polka-dot jumpsuit in bright primary colors, ruffled white collar, "
|
||||
"oversized pom-pom buttons, white gloves, oversized red shoes, red "
|
||||
"foam nose; soft window light from left, eye-level medium shot.")
|
||||
video_path = "https://d2drjpuinn46lb.cloudfront.net/painter_original_edit.mp4"
|
||||
|
||||
generator.generate_video(
|
||||
prompt,
|
||||
negative_prompt="",
|
||||
video_path=video_path,
|
||||
output_path=OUTPUT_PATH,
|
||||
save_video=True,
|
||||
height=480,
|
||||
width=832,
|
||||
num_frames=81,
|
||||
fps=24,
|
||||
guidance_scale=5.0,
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,36 @@
|
||||
"""v2 port of basic.py — Wan2.1-T2V-1.3B through the v2 VideoGenerator.
|
||||
|
||||
Same convenience API as upstream (from_pretrained + generate_video); only delta is importing
|
||||
VideoGenerator from v2. v2 bring-up: single-GPU, resident, SDPA; modest res/frames for a quick run.
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
|
||||
OUTPUT_PATH = "v2_video_samples"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=False,
|
||||
pin_cpu_memory=False,
|
||||
)
|
||||
common = dict(output_path=OUTPUT_PATH, save_video=True,
|
||||
num_frames=25, height=480, width=832, num_inference_steps=30, guidance_scale=5.0)
|
||||
|
||||
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with "
|
||||
"interest. The playful yet serene atmosphere is complemented by soft natural light "
|
||||
"filtering through the petals. Mid-shot, warm and cheerful tones.")
|
||||
video = generator.generate_video(prompt, output_video_name="wan21_raccoon", **common)
|
||||
|
||||
prompt2 = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
|
||||
"the warm afternoon sun. Low angle, steady tracking shot, cinematic.")
|
||||
video2 = generator.generate_video(prompt2, output_video_name="wan21_lion", **common)
|
||||
print(f"Outputs: {video.video_path} , {video2.video_path}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,29 @@
|
||||
"""v2 port of basic_ltx2.py — LTX-2 base (single-stage) through the v2 VideoGenerator.
|
||||
|
||||
Same convenience API as upstream; only delta is importing VideoGenerator from v2. LTX-2 base is the
|
||||
single-stage (non-distilled) model: the v2 single-stage card (build_ltx2_base_card) runs a request-driven
|
||||
many-step flow-match at FULL latent res (no distilled base/refine split, no spatial upsampler), reusing
|
||||
the LTX-2 DiT/VAE/Gemma adapters. The SAME single-stage card also serves LTX-2.3-Distilled (which is also
|
||||
single-stage) — just pass fewer num_inference_steps for the few-step distilled schedule.
|
||||
|
||||
NOTE: modest res/frames here — upstream defaults to 1088x1920x121, which on an 18.88B base is very slow;
|
||||
raise them for full quality. v2 bring-up: single-GPU, resident, SDPA.
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
|
||||
PROMPT = ("A warm sunny backyard, cinematic close-up of two people talking; the camera slowly pans right "
|
||||
"to reveal a grandfather in the garden wearing enormous butterfly wings, flapping his arms like "
|
||||
"he is trying to take off. Deadpan, absurd, quietly tragic.")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
generator = VideoGenerator.from_pretrained("Davids048/LTX2-Base-Diffusers", num_gpus=1)
|
||||
video = generator.generate_video(
|
||||
prompt=PROMPT, output_path="v2_video_samples_ltx2_base", output_video_name="ltx2_base_backyard",
|
||||
save_video=True, num_frames=25, height=512, width=768, num_inference_steps=30)
|
||||
print(f"Output: {video.video_path}")
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,36 @@
|
||||
"""v2 port of basic_ltx2_3_distilled.py — LTX-2.3 Distilled (single-stage, joint A/V) through the v2
|
||||
VideoGenerator.
|
||||
|
||||
Unlike LTX-2.0 distilled (two-stage, video-only), LTX-2.3 is a single-stage *audio+video* model. The
|
||||
shared registry (v2/registry.py) maps ``FastVideo/LTX-2.3-Distilled-Diffusers`` to its OWN card,
|
||||
``build_ltx2_3_card`` — distinct from the LTX-2 base/2-stage cards — which wires the 2.3-specific path:
|
||||
* SEPARATE video + audio text connectors (the Gemma encoder projects the prompt to two embeddings,
|
||||
2048-dim for audio, 4096-dim for video) plus gated attention;
|
||||
* a JOINT DiT forward where video and audio latents cross-attend in a single denoise per step;
|
||||
* a video VAE decode + an AudioDecoder→Vocoder decode → video frames AND a stereo waveform @24kHz.
|
||||
|
||||
Because the model advertises TEXT_TO_VIDEO_SOUND, the VideoGenerator issues a T2VS request by default,
|
||||
so ``generate_video`` returns BOTH modalities: the mp4 plus a sibling ``.wav`` (and ``result.audio`` /
|
||||
``result.audio_sample_rate`` in memory). Being distilled, it wants FEW steps (8). GPU-verified on the
|
||||
rebuilt x86 stack: video (3,33,256,384) + stereo audio (2×61920 @ 24kHz).
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
|
||||
PROMPT = "ocean waves crashing on rocks at sunset, seagulls calling in the distance, cinematic, highly detailed"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
generator = VideoGenerator.from_pretrained("FastVideo/LTX-2.3-Distilled-Diffusers", num_gpus=1)
|
||||
# audio=None auto-enables sound for this A/V model (pass audio=False to force video-only).
|
||||
result = generator.generate_video(
|
||||
prompt=PROMPT, output_path="v2_video_samples_ltx2_3", output_video_name="ltx2_3_ocean",
|
||||
save_video=True, num_frames=33, height=512, width=768, num_inference_steps=8, seed=1)
|
||||
print(f"Video: {result.video_path}")
|
||||
audio_path = result.extra.get("audio_path")
|
||||
if audio_path:
|
||||
print(f"Audio: {audio_path} ({result.audio_sample_rate} Hz)")
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,87 @@
|
||||
"""v2 typed-API inference example — mirrors ``basic_dmd_new_api.py`` but drives the **v2
|
||||
(recipe, runtime) substrate + real torch backend** for the three models brought up on GPU
|
||||
(Wan2.1, SF-causal Wan, LTX-2).
|
||||
|
||||
The ONLY delta from the upstream example is importing ``VideoGenerator`` from ``v2`` instead of
|
||||
``fastvideo`` — the typed config classes are the SAME ``fastvideo.api`` dataclasses.
|
||||
|
||||
Run (on a GPU box, with the v2 venv active):
|
||||
python examples/inference/basic/v2_basic_new_api.py
|
||||
|
||||
Notes vs upstream: the v2 bring-up runs single-GPU, resident, on the TORCH_SDPA backend (no
|
||||
fastvideo-kernel / VSA), so resolutions/steps are modest here for a quick runnable demo. LTX-2 loads
|
||||
an 18.88B DiT (slow first load).
|
||||
"""
|
||||
import os
|
||||
import time
|
||||
|
||||
from v2 import VideoGenerator
|
||||
from fastvideo.api import (
|
||||
EngineConfig,
|
||||
GenerationRequest,
|
||||
GeneratorConfig,
|
||||
OffloadConfig,
|
||||
OutputConfig,
|
||||
SamplingConfig,
|
||||
)
|
||||
|
||||
OUTPUT_PATH = "v2_video_samples"
|
||||
|
||||
MODELS = [
|
||||
{
|
||||
"family": "wan21",
|
||||
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
"prompt": "a red panda surfing on ocean waves at sunset, cinematic, highly detailed",
|
||||
"sampling": SamplingConfig(num_frames=25, height=480, width=832,
|
||||
num_inference_steps=30, guidance_scale=5.0, seed=1, fps=16),
|
||||
},
|
||||
{
|
||||
"family": "wan_causal",
|
||||
"model_path": "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers",
|
||||
"prompt": "a cat walking through a sunlit garden, cinematic",
|
||||
"sampling": SamplingConfig(num_frames=25, height=480, width=832,
|
||||
num_inference_steps=4, guidance_scale=5.0, seed=1, fps=16),
|
||||
},
|
||||
{
|
||||
"family": "ltx2",
|
||||
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
|
||||
"prompt": "surfers riding ocean waves at sunset, cinematic, highly detailed",
|
||||
"sampling": SamplingConfig(num_frames=9, height=512, width=768,
|
||||
num_inference_steps=8, guidance_scale=1.0, seed=1, fps=16),
|
||||
},
|
||||
]
|
||||
|
||||
|
||||
def run_one(m: dict) -> None:
|
||||
generator_config = GeneratorConfig(
|
||||
model_path=m["model_path"],
|
||||
engine=EngineConfig(
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
offload=OffloadConfig(text_encoder=False, dit=False, vae=False, pin_cpu_memory=False),
|
||||
),
|
||||
)
|
||||
load_start = time.perf_counter()
|
||||
generator = VideoGenerator.from_config(generator_config)
|
||||
load_time = time.perf_counter() - load_start
|
||||
|
||||
request = GenerationRequest(
|
||||
prompt=m["prompt"],
|
||||
sampling=m["sampling"],
|
||||
output=OutputConfig(output_path=OUTPUT_PATH, output_video_name=f"v2_{m['family']}",
|
||||
save_video=True, return_frames=False),
|
||||
)
|
||||
gen_start = time.perf_counter()
|
||||
result = generator.generate(request)
|
||||
gen_time = time.perf_counter() - gen_start
|
||||
|
||||
print(f"[{m['family']:10s}] load={load_time:6.1f}s gen={gen_time:6.1f}s -> {result.video_path}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
for m in MODELS:
|
||||
run_one(m)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,30 @@
|
||||
"""v2 port of basic_self_forcing_causal.py — SF-causal Wan2.1 (CausalWanTransformer3DModel) through
|
||||
the v2 VideoGenerator (chunk_rollout loop).
|
||||
|
||||
Same convenience API as upstream; only delta is importing VideoGenerator from v2. NOTE: the v2 causal
|
||||
loop runs per-chunk few-step (not the upstream kv-cache streaming + SF schedule), so output is coherent
|
||||
but lower-fidelity (a documented gap). num_frames is set by the card's chunk schedule; height/width
|
||||
drive the latent geometry.
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
from fastvideo.api.sampling_param import SamplingParam
|
||||
|
||||
OUTPUT_PATH = "v2_video_samples_causal"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
model_name = "wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers"
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_name, num_gpus=1, text_encoder_cpu_offload=False, dit_cpu_offload=False)
|
||||
sampling_param = SamplingParam.from_pretrained(model_name)
|
||||
|
||||
prompt = ("A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with "
|
||||
"interest. The playful yet serene atmosphere is complemented by soft natural light "
|
||||
"filtering through the petals. Mid-shot, warm and cheerful tones.")
|
||||
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="causal_raccoon",
|
||||
save_video=True, sampling_param=sampling_param, height=480, width=832)
|
||||
print(f"Output: {video.video_path}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,40 @@
|
||||
"""v2 port of basic_wan2_2.py — Wan2.2-T2V-A14B (MoE) through the v2 VideoGenerator.
|
||||
|
||||
Same convenience API as upstream; only delta is importing VideoGenerator from v2. A14B is a 2-expert
|
||||
MoE: WanTransformer3DModel x2 (in_ch=16, Wan2.1 geometry) with a boundary-timestep switch
|
||||
(boundary_ratio 0.875) — ported via build_wan22_a14b_card (BoundaryTimestepRouting: transformer =
|
||||
high-noise expert, transformer_2 = low-noise), reusing the Wan adapters for both experts.
|
||||
|
||||
NOTE: upstream runs A14B with num_gpus=2 + dit_cpu_offload=True ("DiT need to be offloaded for MoE").
|
||||
The v2 bring-up is single-GPU + resident (no offload), so the two 14B experts (~56GB bf16) + UMT5 are
|
||||
near an 80GB GPU's limit — this example uses reduced res/frames to fit. If it OOMs, the A14B card is
|
||||
still correct; it just needs the (not-yet-ported) MoE DiT CPU offload. See V2_PORTING_STATUS.md.
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
|
||||
OUTPUT_PATH = "v2_video_samples_wan2_2_14B_t2v"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.2-T2V-A14B-Diffusers",
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=False,
|
||||
pin_cpu_memory=False,
|
||||
)
|
||||
|
||||
prompt = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
|
||||
"the warm afternoon sun. The tall grass ripples gently in the breeze. Low angle, steady "
|
||||
"tracking shot, cinematic.")
|
||||
# Reduced res/frames so the two resident 14B experts fit a single 80GB GPU (upstream: 720x1280x81).
|
||||
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="wan22_a14b_lion",
|
||||
save_video=True, num_frames=17, height=480, width=832,
|
||||
num_inference_steps=20, guidance_scale=5.0)
|
||||
print(f"Output: {video.video_path}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,38 @@
|
||||
"""v2 port of basic_wan2_2_ti2v.py — Wan2.2-TI2V-5B (T2V mode) through the v2 VideoGenerator.
|
||||
|
||||
Same convenience API as upstream (from_pretrained + generate_video); only delta is importing
|
||||
VideoGenerator from v2. Wan2.2-TI2V-5B reuses the Wan adapter classes (WanTransformer3DModel /
|
||||
AutoencoderKLWan / UMT5) with the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal).
|
||||
|
||||
NOTE: upstream also runs I2V (image_path=...). The v2 program here is T2V-only (image conditioning is
|
||||
not yet ported), so this mirrors the upstream *T2V* branch (prompt2). Modest res/frames for a quick run.
|
||||
"""
|
||||
from v2 import VideoGenerator
|
||||
|
||||
OUTPUT_PATH = "v2_video_samples_wan2_2_5B_ti2v"
|
||||
|
||||
|
||||
def main() -> None:
|
||||
model_name = "Wan-AI/Wan2.2-TI2V-5B-Diffusers"
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
model_name,
|
||||
num_gpus=1,
|
||||
use_fsdp_inference=False,
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=False,
|
||||
text_encoder_cpu_offload=False,
|
||||
pin_cpu_memory=False,
|
||||
)
|
||||
|
||||
# T2V mode (the v2 program is text-to-video; upstream's image_path I2V branch is not ported yet).
|
||||
prompt = ("A majestic lion strides across the golden savanna, its powerful frame glistening under "
|
||||
"the warm afternoon sun. The tall grass ripples gently in the breeze, enhancing the lion's "
|
||||
"commanding presence. Low angle, steady tracking shot, cinematic.")
|
||||
video = generator.generate_video(prompt, output_path=OUTPUT_PATH, output_video_name="wan22_ti2v_lion",
|
||||
save_video=True, num_frames=25, height=448, width=768,
|
||||
num_inference_steps=20, guidance_scale=5.0)
|
||||
print(f"Output: {video.video_path}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,91 @@
|
||||
"""Run ``judge.third_person_separation`` (needs ``.[eval-judge]`` + a Gemini key)
|
||||
over each baseline and print the candidate's win-rate table — from a ``--manifest``
|
||||
of pairs, or by pairing ``--candidate-dir`` against each ``--reference`` dir by
|
||||
filename stem.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from collections import defaultdict
|
||||
from pathlib import Path
|
||||
|
||||
from fastvideo.eval import create_evaluator
|
||||
|
||||
METRIC = "judge.third_person_separation"
|
||||
VIDEO_EXTS = {".mp4", ".avi", ".mov", ".mkv", ".webm"}
|
||||
IMAGE_EXTS = {".png", ".jpg", ".jpeg", ".webp"}
|
||||
|
||||
|
||||
def _by_stem(directory: Path, exts: set[str]) -> dict[str, Path]:
|
||||
"""Map filename stem -> path for files with the given extensions."""
|
||||
return {p.stem: p for p in sorted(directory.iterdir()) if p.suffix.lower() in exts}
|
||||
|
||||
|
||||
def main() -> None:
|
||||
p = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
p.add_argument("--candidate-dir", type=Path, default=None,
|
||||
help="Directory of candidate clips (directory mode).")
|
||||
p.add_argument("--reference", action="append", default=[], metavar="NAME=DIR",
|
||||
help="Baseline directory, repeatable: 'name=dir' or bare 'dir'.")
|
||||
p.add_argument("--image-dir", type=Path, default=None,
|
||||
help="Optional first-frame images, matched to clips by stem.")
|
||||
p.add_argument("--prompts-json", type=Path, default=None,
|
||||
help="Optional {stem: control-signal text} JSON.")
|
||||
p.add_argument("--actions-json", type=Path, default=None,
|
||||
help="Optional {stem: action-label} JSON for the per-action breakdown.")
|
||||
p.add_argument("--manifest", type=Path, default=None,
|
||||
help="JSON list of {baseline, video_path, reference_path, ...} rows.")
|
||||
p.add_argument("--output", type=Path, default=None)
|
||||
args = p.parse_args()
|
||||
|
||||
# Group path-only samples per baseline: {baseline: [sample dict, ...]}.
|
||||
by_baseline: dict[str, list[dict]] = defaultdict(list)
|
||||
if args.manifest is not None:
|
||||
for row in json.loads(args.manifest.read_text()):
|
||||
by_baseline[row.get("baseline", "baseline")].append(
|
||||
{k: v for k, v in row.items() if k != "baseline"})
|
||||
elif args.candidate_dir is not None and args.reference:
|
||||
cands = _by_stem(args.candidate_dir, VIDEO_EXTS)
|
||||
images = _by_stem(args.image_dir, IMAGE_EXTS) if args.image_dir else {}
|
||||
prompts = json.loads(args.prompts_json.read_text()) if args.prompts_json else {}
|
||||
actions = json.loads(args.actions_json.read_text()) if args.actions_json else {}
|
||||
for spec in args.reference:
|
||||
name, sep, ref_dir = spec.partition("=")
|
||||
if not sep:
|
||||
name, ref_dir = Path(spec).name, spec
|
||||
refs = _by_stem(Path(ref_dir), VIDEO_EXTS)
|
||||
for stem in sorted(cands.keys() & refs.keys()):
|
||||
sample = {"video_path": str(cands[stem]), "reference_path": str(refs[stem])}
|
||||
if stem in images:
|
||||
sample["image_path"] = str(images[stem])
|
||||
if stem in prompts:
|
||||
sample["text_prompt"] = prompts[stem]
|
||||
if stem in actions:
|
||||
sample["action"] = actions[stem]
|
||||
by_baseline[name].append(sample)
|
||||
else:
|
||||
p.error("provide either --manifest, or --candidate-dir with at least one --reference")
|
||||
|
||||
ev = create_evaluator(metrics=[METRIC], device="cpu")
|
||||
print("\n| Baseline | Candidate win-rate (excl. ties) | W / L / T | n |")
|
||||
print("|---|---|---|---|")
|
||||
rows = {}
|
||||
for baseline, samples in by_baseline.items():
|
||||
res = ev.evaluate(samples=samples).corpus[METRIC]
|
||||
rows[baseline] = res
|
||||
d = res.details
|
||||
if res.score is None:
|
||||
print(f"| {baseline} | — | — | 0 |")
|
||||
else:
|
||||
print(f"| {baseline} | {100 * res.score:.1f}% | {d['wins']}/{d['losses']}/{d['ties']} | {d['n']} |")
|
||||
|
||||
if args.output is not None:
|
||||
payload = {b: {"score": r.score, "details": r.details} for b, r in rows.items()}
|
||||
args.output.write_text(json.dumps(payload, indent=2))
|
||||
print(f"\nWrote {args.output}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,95 @@
|
||||
"""NVFP4 + Attn-QAT (modified SageAttention3) inference on Blackwell.
|
||||
|
||||
Runs Wan2.1-T2V-1.3B fully in 4-bit: NVFP4 linear layers (activations
|
||||
quantized on the fly) together with the modified SageAttention3 FP4 attention
|
||||
backend (``ATTN_QAT_INFER``). This is the inference half of the
|
||||
Quantization-Aware Distillation (QAD) recipe.
|
||||
|
||||
Requirements:
|
||||
- RTX 5090 / consumer Blackwell (sm_120a). The attn_qat_infer kernel hard
|
||||
gates on sm_120; on other GPUs it falls back to Flash Attention.
|
||||
- The attn_qat_infer kernel built into fastvideo-kernel (see #1455) and
|
||||
flashinfer for the NVFP4 linear matmuls.
|
||||
|
||||
Usage:
|
||||
python nvfp4_qat_wan2_1_1_3b.py # NVFP4 linear + Attn-QAT attn
|
||||
python nvfp4_qat_wan2_1_1_3b.py --bf16 # BF16 baseline
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import os
|
||||
import time
|
||||
|
||||
OUTPUT_PATH = "video_samples"
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description="NVFP4 + Attn-QAT video generation")
|
||||
parser.add_argument("--bf16", action="store_true",
|
||||
help="BF16 baseline (no NVFP4 linear, default attention)")
|
||||
parser.add_argument("--model", default="Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
|
||||
help="Model path or HuggingFace ID")
|
||||
parser.add_argument("--quant-method", default="nvfp4_qat", choices=["nvfp4_qat", "NVFP4"],
|
||||
help="Linear quantization config. Wan-2.1 uses nvfp4_qat (matches its "
|
||||
"to_q/k/v/out + ffn layers); NVFP4 is LTX2-specific and will NOT "
|
||||
"quantize Wan.")
|
||||
parser.add_argument("--compile", action="store_true", help="Enable torch.compile for the DiT")
|
||||
parser.add_argument("--num_gpus", type=int, default=1)
|
||||
parser.add_argument("--infer_steps", type=int, default=50)
|
||||
args = parser.parse_args()
|
||||
|
||||
# The attention backend is selected via env var before the engine starts.
|
||||
# ATTN_QAT_INFER -> AttnQatInferBackend (modified SageAttention3 FP4).
|
||||
if not args.bf16:
|
||||
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "ATTN_QAT_INFER"
|
||||
|
||||
# Import after the env var so the platform picks up the selection.
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.layers.quantization import get_quantization_config
|
||||
|
||||
mode = "bf16" if args.bf16 else args.quant_method
|
||||
if args.compile:
|
||||
mode += "_compile"
|
||||
print(f"Mode: {mode.upper()}")
|
||||
|
||||
# transformer_quant needs a QuantizationConfig *instance* — the bare string
|
||||
# is not resolved on the from_pretrained kwarg path.
|
||||
extra = {} if args.bf16 else {"transformer_quant": get_quantization_config(args.quant_method)()}
|
||||
generator = VideoGenerator.from_pretrained(
|
||||
args.model,
|
||||
num_gpus=args.num_gpus,
|
||||
use_fsdp_inference=args.bf16,
|
||||
dit_cpu_offload=False,
|
||||
vae_cpu_offload=True,
|
||||
text_encoder_cpu_offload=True,
|
||||
enable_torch_compile=args.compile,
|
||||
**extra,
|
||||
)
|
||||
|
||||
prompt = (
|
||||
"A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes "
|
||||
"wide with interest. The playful yet serene atmosphere is complemented by soft "
|
||||
"natural light filtering through the petals. Mid-shot, warm and cheerful tones."
|
||||
)
|
||||
|
||||
n_warmup = 2 if args.compile else 1
|
||||
for _ in range(n_warmup):
|
||||
generator.generate(request={"prompt": prompt, "sampling": {"num_inference_steps": 2},
|
||||
"output": {"save_video": False}})
|
||||
|
||||
os.makedirs(OUTPUT_PATH, exist_ok=True)
|
||||
start = time.time()
|
||||
generator.generate(request={
|
||||
"prompt": prompt,
|
||||
"sampling": {"num_inference_steps": args.infer_steps},
|
||||
"output": {"save_video": True, "output_path": os.path.join(OUTPUT_PATH, f"raccoon_{mode}.mp4")},
|
||||
})
|
||||
elapsed = time.time() - start
|
||||
print(f"[{mode.upper()}] {args.infer_steps} steps in {elapsed:.2f}s "
|
||||
f"({args.infer_steps / elapsed:.2f} it/s)")
|
||||
|
||||
generator.shutdown()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"data": [
|
||||
{
|
||||
"caption": "gold tip pyramid in the night, extremely detailed , rain, stars"
|
||||
},
|
||||
{
|
||||
"caption": "a fruit stacking in the shape of a dog, stock image, shutterstock"
|
||||
},
|
||||
{
|
||||
"caption": "Landscape, By Lee madgwick, by Luis Royo, by Louise nevelson"
|
||||
},
|
||||
{
|
||||
"caption": "A colorful poster that says \"philo is a weird\""
|
||||
},
|
||||
{
|
||||
"caption": "Danish male with blue eyes, realistic, viking"
|
||||
},
|
||||
{
|
||||
"caption": "futuristic, cityscape, flying cars, neon lights, towering skyscrapers, glowing purple sky."
|
||||
},
|
||||
{
|
||||
"caption": "a crow with cameras for eyes, sitting on a mans shoulder, anime, studio ghibli, fantasy, fairytale, sketch, digital art, watercolor, dnd, rustic, professional photograph, medieval, hd, 4k"
|
||||
},
|
||||
{
|
||||
"caption": "a background image mixing the matrix and AI"
|
||||
},
|
||||
{
|
||||
"caption": "Golden sunset, a bright orange and yellow sky is visible, lit up by the setting sun, the horizon is a mix of bright colors and deep shadows"
|
||||
},
|
||||
{
|
||||
"caption": "Grim reaper playing an electric guitar"
|
||||
},
|
||||
{
|
||||
"caption": "an epic view of a demonic Rose-ringed parakeet cyborg inside an ironmaiden robot,wearing a noble robe,large view,a surrealist painting, aralan bean and Philippe Druillet,hiromu arakawa,volumetric lighting,detailed shadows"
|
||||
},
|
||||
{
|
||||
"caption": "Ben Shapiro as the cover of ministry's filth pig album, but covered in milk"
|
||||
},
|
||||
{
|
||||
"caption": "king charles spaniel with , ethereal, midjourney style lighting and shadows, insanely detailed, 8k, photorealistic"
|
||||
},
|
||||
{
|
||||
"caption": "A website for a party resort service"
|
||||
},
|
||||
{
|
||||
"caption": "full shot of a steampunk horse"
|
||||
},
|
||||
{
|
||||
"caption": "60s psycedelic spiritual jazz album art"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -87,6 +87,7 @@ training:
|
||||
# --- training.data [TYPED] -> DataConfig ---
|
||||
data:
|
||||
data_path: data/my_dataset # default: ""
|
||||
preprocessed_data_type: t2v # default: "t2v" ("text_only" for simulate-only DMD text prompts)
|
||||
train_batch_size: 1 # default: 1
|
||||
dataloader_num_workers: 4 # default: 0
|
||||
training_cfg_rate: 0.1 # default: 0.0
|
||||
|
||||
@@ -0,0 +1,116 @@
|
||||
# DiffusionNFT multi-reward single-frame RL: Wan 2.1 T2V 1.3B on text-only PickScore prompts.
|
||||
#
|
||||
# Single-frame RL is represented as a one-latent-frame Wan run:
|
||||
# num_latent_t: 1
|
||||
# num_frames: 1
|
||||
#
|
||||
# The method trains the full transformer (no LoRA) and keeps an old-policy
|
||||
# transformer plus a frozen reference transformer, matching the non-LoRA
|
||||
# DiffusionNFT loss path.
|
||||
|
||||
models:
|
||||
student:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: true
|
||||
old:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: false
|
||||
disable_custom_init_weights: true
|
||||
reference:
|
||||
_target_: fastvideo.train.models.wan.WanModel
|
||||
init_from: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
|
||||
trainable: false
|
||||
disable_custom_init_weights: true
|
||||
|
||||
method:
|
||||
_target_: fastvideo.train.methods.rl.diffusion_nft.DiffusionNFTMethod
|
||||
reward_fn:
|
||||
pickscore: 1.0
|
||||
clipscore: 1.0
|
||||
|
||||
sampling:
|
||||
num_steps: 25
|
||||
scheduler: flow_match_euler
|
||||
trajectory: ode
|
||||
flow_shift: inherit
|
||||
|
||||
validation:
|
||||
every_steps: 10
|
||||
num_steps: 40
|
||||
num_prompts: 16
|
||||
batch_size: 16
|
||||
log_samples: true
|
||||
seed: 42
|
||||
# Null reuses training.data.data_path. Override this with a held-out
|
||||
# preprocessed parquet path when one is available.
|
||||
data_path:
|
||||
|
||||
# DiffusionNFT sd3_multi_reward on 4 GPUs resolves to per-GPU sample batch
|
||||
# size 6, 48 sample batches per outer epoch, and grad accumulation 48.
|
||||
sample_train_batch_size: 6
|
||||
train_batch_size: 6
|
||||
num_batches_per_epoch: 48
|
||||
num_video_per_prompt: 24
|
||||
num_inner_epochs: 1
|
||||
timestep_fraction: 0.99
|
||||
|
||||
beta: 0.1
|
||||
kl_beta: 0.0001
|
||||
decay_type: 1
|
||||
adv_mode: all
|
||||
adv_clip_max: 5
|
||||
max_grad_norm: 1.0
|
||||
ema:
|
||||
enabled: true
|
||||
decay: 0.9
|
||||
update_after_step: 0
|
||||
validation: true
|
||||
terminal_progress: true
|
||||
|
||||
training:
|
||||
distributed:
|
||||
num_gpus: 4
|
||||
sp_size: 1
|
||||
tp_size: 1
|
||||
hsdp_replicate_dim: 1
|
||||
hsdp_shard_dim: 4
|
||||
|
||||
data:
|
||||
data_path: data/pickscore_text_only_preprocessed
|
||||
preprocessed_data_type: text_only
|
||||
dataloader_num_workers: 0
|
||||
train_batch_size: 1
|
||||
training_cfg_rate: 0.0
|
||||
seed: 42
|
||||
num_latent_t: 1
|
||||
num_height: 448
|
||||
num_width: 832
|
||||
num_frames: 1
|
||||
|
||||
optimizer:
|
||||
learning_rate: 3.0e-5
|
||||
betas: [0.9, 0.999]
|
||||
weight_decay: 0.0001
|
||||
lr_scheduler: constant
|
||||
lr_warmup_steps: 0
|
||||
|
||||
loop:
|
||||
max_train_steps: 100000
|
||||
gradient_accumulation_steps: 48
|
||||
|
||||
checkpoint:
|
||||
output_dir: outputs/wan2.1_diffusion_nft_pick_clip
|
||||
training_state_checkpointing_steps: 30
|
||||
checkpoints_total_limit: 3
|
||||
|
||||
tracker:
|
||||
project_name: diffusion_nft_wan
|
||||
run_name: wan2.1_diffusion_nft_pick_clip
|
||||
|
||||
model:
|
||||
enable_gradient_checkpointing_type: full
|
||||
|
||||
pipeline:
|
||||
flow_shift: 8
|
||||
@@ -0,0 +1,119 @@
|
||||
# MixKit training data (QAD 5090 recipe)
|
||||
|
||||
The QAD 5090 models are distilled from Wan2.1-T2V-1.3B on a MixKit subset at
|
||||
**480×832, 77 frames, 16 fps**. FastVideo training consumes **Parquet** shards of
|
||||
precomputed VAE latents + text embeddings (no text encoder / VAE needed at train
|
||||
time).
|
||||
|
||||
## Option A — download the preprocessed data (recommended)
|
||||
|
||||
The encoded dataset is published on the Hugging Face Hub, ready to train:
|
||||
|
||||
```bash
|
||||
# from the repo root
|
||||
bash examples/training/finetune/wan_t2v_1.3B/mixkit/download_mixkit_data.sh
|
||||
```
|
||||
|
||||
This pulls [`weizhou03/HD-Mixkit-Finetune-Wan`](https://huggingface.co/datasets/weizhou03/HD-Mixkit-Finetune-Wan)
|
||||
into `data/HD-Mixkit-Finetune-Wan/`:
|
||||
|
||||
```
|
||||
data/HD-Mixkit-Finetune-Wan/
|
||||
├── combined_parquet_dataset/ # training shards -> point --data_path here
|
||||
│ └── worker_0/data_chunk_*.parquet
|
||||
└── validation_parquet_dataset/ # validation shards
|
||||
└── worker_0/data_chunk_0.parquet
|
||||
```
|
||||
|
||||
Each Parquet row holds the VAE latent bytes + text-embedding bytes (plus
|
||||
shape/dtype metadata), matching FastVideo's standard preprocessing output.
|
||||
|
||||
## Option B — build the Parquet from raw videos
|
||||
|
||||
If you want to reproduce the encoding from your own MixKit videos, arrange them as
|
||||
a `merged` dataset (videos + a captions JSON), then run FastVideo's standard
|
||||
preprocessing to VAE-encode and text-embed them into Parquet:
|
||||
|
||||
```bash
|
||||
GPU_NUM=2
|
||||
torchrun --nproc_per_node=$GPU_NUM \
|
||||
-m fastvideo.pipelines.preprocess.v1_preprocessing_new \
|
||||
--model_path "Wan-AI/Wan2.1-T2V-1.3B-Diffusers" \
|
||||
--mode preprocess \
|
||||
--workload_type t2v \
|
||||
--preprocess.video_loader_type torchvision \
|
||||
--preprocess.dataset_type merged \
|
||||
--preprocess.dataset_path "data/mixkit_raw/" \
|
||||
--preprocess.dataset_output_dir "data/HD-Mixkit-Finetune-Wan/" \
|
||||
--preprocess.max_height 480 \
|
||||
--preprocess.max_width 832 \
|
||||
--preprocess.num_frames 77 \
|
||||
--preprocess.train_fps 16 \
|
||||
--preprocess.samples_per_file 8
|
||||
```
|
||||
|
||||
The raw videos are full-HD MixKit clips (≈1080p/30fps); preprocessing resizes to
|
||||
480×832, resamples to 16 fps, and extracts 77 frames per clip. See
|
||||
[`docs/training/data_preprocess.md`](../../../../../docs/training/data_preprocess.md)
|
||||
for the full parameter reference.
|
||||
|
||||
## Train (QAT finetune)
|
||||
|
||||
With the data in place, run the quantization-aware finetune. The 4-bit attention
|
||||
path is **config-driven** — selected purely by an env var, no monkey-patching:
|
||||
|
||||
```bash
|
||||
bash examples/training/finetune/wan_t2v_1.3B/mixkit/finetune_qat.sh
|
||||
# or point at your own parquet dir / GPU count:
|
||||
NUM_GPUS=4 bash .../mixkit/finetune_qat.sh data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/
|
||||
```
|
||||
|
||||
`FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN` routes attention through the
|
||||
fake-quantized Triton kernel (straight-through estimator), so the DiT learns to
|
||||
absorb FP4 attention error. This kernel is Triton, so it runs on both `sm_100`
|
||||
(B200/GB200) and `sm_120` (RTX 5090).
|
||||
|
||||
## Train stage 2 (QAT DMD distillation to 3 steps)
|
||||
|
||||
Distill the QAT-finetuned generator down to **3 sampling steps**. Only the
|
||||
generator is quantized (Attn-QAT); the teacher (`real_score`) and critic
|
||||
(`fake_score`) stay full precision. This is enforced in the loader
|
||||
(`component_loader.py`, via the `_loading_teacher_critic_model` flag), so the
|
||||
same global `ATTN_QAT_TRAIN` env reaches **only** the generator — no per-model
|
||||
flags or monkey-patching.
|
||||
|
||||
```bash
|
||||
# generator init = the stage-1 finetune checkpoint
|
||||
bash examples/training/finetune/wan_t2v_1.3B/mixkit/distill_dmd_qat.sh \
|
||||
data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/ \
|
||||
checkpoints/wan_t2v_qat_finetune/checkpoint-2000/transformer/diffusion_pytorch_model.safetensors
|
||||
```
|
||||
|
||||
DMD runs a double loop (critic every step, generator every
|
||||
`generator_update_interval`), and validation samples the distilled student at
|
||||
3 steps — the final 4-bit-attention model.
|
||||
|
||||
## Inference (NVFP4 4-bit linear)
|
||||
|
||||
For Wan-2.1, enable the FP4 linear layers with the **`nvfp4_qat`** quantization
|
||||
config (it matches Wan's `to_q/k/v/out` + `ffn` layers; the plain `NVFP4` config
|
||||
is LTX2-specific and will not quantize Wan):
|
||||
|
||||
```python
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.layers.quantization import get_quantization_config
|
||||
|
||||
gen = VideoGenerator.from_pretrained(
|
||||
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers", num_gpus=1,
|
||||
transformer_quant=get_quantization_config("nvfp4_qat")(), # a config instance, not the string
|
||||
use_fsdp_inference=False,
|
||||
)
|
||||
gen.generate(request={"prompt": "...", "output": {"save_video": True}})
|
||||
```
|
||||
|
||||
The loader converts the tagged linear weights to FP4 at load time
|
||||
(`_maybe_convert_model_to_nvfp4`). Combine with
|
||||
`FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER` on an RTX 5090 (`sm_120`) for the
|
||||
full 4-bit path; on other GPUs the attention falls back to Flash while the FP4
|
||||
linear layers still run. `flashinfer` (and a host C++ compiler for its FP4
|
||||
kernel JIT) are required.
|
||||
@@ -0,0 +1,51 @@
|
||||
#!/bin/bash
|
||||
# QAD recipe stage 2 — quantization-aware DMD distillation of Wan2.1-T2V-1.3B
|
||||
# down to 3 sampling steps, with the GENERATOR in fake-quant Attn-QAT and the
|
||||
# teacher (real_score) + critic (fake_score) at full precision.
|
||||
#
|
||||
# Generator-only QAT is config-driven: FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN
|
||||
# is applied to the generator only, because the loader masks it (and the
|
||||
# nvfp4_qat quant) for the teacher/critic via the `_loading_teacher_critic_model`
|
||||
# flag (see fastvideo/models/loader/component_loader.py). No monkey-patching.
|
||||
#
|
||||
# Init the generator from the stage-1 finetune checkpoint (finetune_qat.sh).
|
||||
# Data: run download_mixkit_data.sh first.
|
||||
#
|
||||
# Verified end-to-end on Blackwell (GB200/sm_100): generator loads with
|
||||
# ATTN_QAT_TRAIN while teacher/critic load full-precision; the DMD double loop
|
||||
# runs (generator updates every generator_update_interval steps, critic every
|
||||
# step), 3-step validation generates videos, checkpoint saved.
|
||||
set -euo pipefail
|
||||
|
||||
export FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN # generator-only (loader-gated)
|
||||
export WANDB_MODE=${WANDB_MODE:-online}
|
||||
export TOKENIZERS_PARALLELISM=false
|
||||
|
||||
BASE="Wan-AI/Wan2.1-T2V-1.3B-Diffusers"
|
||||
DATA_DIR=${1:-"data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/"}
|
||||
# Generator init weights = the stage-1 QAT-finetune checkpoint.
|
||||
INIT_WEIGHTS=${2:-"checkpoints/wan_t2v_qat_finetune/checkpoint-2000/transformer/diffusion_pytorch_model.safetensors"}
|
||||
VALIDATION_FILE="$(dirname "$0")/../crush_smol/validation.json"
|
||||
NUM_GPUS=${NUM_GPUS:-4}
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node "${NUM_GPUS}" \
|
||||
fastvideo/training/wan_distillation_pipeline.py \
|
||||
--num_gpus "${NUM_GPUS}" --sp_size 1 --tp_size 1 \
|
||||
--hsdp_replicate_dim "${NUM_GPUS}" --hsdp_shard_dim 1 \
|
||||
--model_path "${BASE}" --pretrained_model_name_or_path "${BASE}" \
|
||||
--real_score_model_path "${BASE}" --fake_score_model_path "${BASE}" \
|
||||
--init_weights_from_safetensors "${INIT_WEIGHTS}" \
|
||||
--data_path "${DATA_DIR}" --dataloader_num_workers 4 \
|
||||
--max_train_steps 2000 --train_batch_size 1 --train_sp_batch_size 1 \
|
||||
--gradient_accumulation_steps 1 \
|
||||
--num_latent_t 20 --num_height 480 --num_width 832 --num_frames 77 \
|
||||
--enable_gradient_checkpointing_type full \
|
||||
--log_validation --validation_dataset_file "${VALIDATION_FILE}" \
|
||||
--validation_steps 200 --validation_sampling_steps 3 --validation_guidance_scale 6.0 \
|
||||
--learning_rate 2e-6 --mixed_precision bf16 --weight_decay 0.01 --max_grad_norm 1.0 \
|
||||
--weight_only_checkpointing_steps 500 --training_state_checkpointing_steps 500 \
|
||||
--tracker_project_name wan_t2v_distill_dmd_qat \
|
||||
--output_dir checkpoints/wan_t2v_distill_dmd_qat \
|
||||
--inference_mode False --dit_precision fp32 --ema_start_step 0 --training_cfg_rate 0.0 \
|
||||
--generator_update_interval 5 --real_score_guidance_scale 2.0 \
|
||||
--dmd_denoising_steps '1000,757,522' --min_timestep_ratio 0.02 --max_timestep_ratio 0.98
|
||||
@@ -0,0 +1,22 @@
|
||||
#!/bin/bash
|
||||
# Download the preprocessed MixKit finetune dataset used for the QAD 5090 recipe.
|
||||
#
|
||||
# This is the MixKit subset already VAE-encoded (Wan2.1-T2V-1.3B) and text-embedded
|
||||
# into Parquet shards, so it can be fed straight to training with no further
|
||||
# preprocessing. To build the Parquet from raw videos yourself, see README.md.
|
||||
#
|
||||
# Usage (run from the repo root):
|
||||
# bash examples/training/finetune/wan_t2v_1.3B/mixkit/download_mixkit_data.sh [DATA_ROOT]
|
||||
set -euo pipefail
|
||||
|
||||
DATA_ROOT=${1:-data/HD-Mixkit-Finetune-Wan}
|
||||
|
||||
python scripts/huggingface/download_hf.py \
|
||||
--repo_id "weizhou03/HD-Mixkit-Finetune-Wan" \
|
||||
--local_dir "${DATA_ROOT}" \
|
||||
--repo_type "dataset"
|
||||
|
||||
echo "Done."
|
||||
echo " Train data: ${DATA_ROOT}/combined_parquet_dataset"
|
||||
echo " Validation data: ${DATA_ROOT}/validation_parquet_dataset"
|
||||
echo "Point your training script's data path at the combined_parquet_dataset directory."
|
||||
@@ -0,0 +1,44 @@
|
||||
#!/bin/bash
|
||||
# QAD recipe — quantization-aware finetune of Wan2.1-T2V-1.3B with fake-quant
|
||||
# (Attn-QAT) attention.
|
||||
#
|
||||
# The 4-bit attention path is selected purely by env var (config-driven, no
|
||||
# monkey-patching): FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN routes attention
|
||||
# through the fake-quantized Triton kernel (straight-through estimator), so the
|
||||
# DiT learns to absorb FP4 attention error instead of fighting it.
|
||||
#
|
||||
# Data: run download_mixkit_data.sh first (preprocessed Parquet).
|
||||
#
|
||||
# Verified end-to-end on Blackwell (GB200/sm_100): the ATTN_QAT_TRAIN backend is
|
||||
# selected (not a fallback), forward+backward run, loss/grad are healthy, and
|
||||
# validation generates videos. The kernel is Triton so it runs on sm_100 and
|
||||
# sm_120 alike (the FP4 inference kernel, by contrast, is sm_120-only).
|
||||
set -euo pipefail
|
||||
|
||||
export FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_TRAIN # <-- enables Attn-QAT training
|
||||
export WANDB_MODE=${WANDB_MODE:-online}
|
||||
export TOKENIZERS_PARALLELISM=false
|
||||
|
||||
MODEL_PATH="Wan-AI/Wan2.1-T2V-1.3B-Diffusers"
|
||||
DATA_DIR=${1:-"data/HD-Mixkit-Finetune-Wan/combined_parquet_dataset/"}
|
||||
VALIDATION_FILE="$(dirname "$0")/../crush_smol/validation.json"
|
||||
NUM_GPUS=${NUM_GPUS:-4}
|
||||
|
||||
torchrun --nnodes 1 --nproc_per_node "${NUM_GPUS}" \
|
||||
fastvideo/training/wan_training_pipeline.py \
|
||||
--num_gpus "${NUM_GPUS}" --sp_size "${NUM_GPUS}" --tp_size 1 \
|
||||
--hsdp_replicate_dim 1 --hsdp_shard_dim "${NUM_GPUS}" \
|
||||
--model_path "${MODEL_PATH}" --pretrained_model_name_or_path "${MODEL_PATH}" \
|
||||
--data_path "${DATA_DIR}" --dataloader_num_workers 1 \
|
||||
--max_train_steps 2000 --train_batch_size 1 --train_sp_batch_size 1 \
|
||||
--gradient_accumulation_steps 1 \
|
||||
--num_latent_t 20 --num_height 480 --num_width 832 --num_frames 77 \
|
||||
--enable_gradient_checkpointing_type full \
|
||||
--log_validation --validation_dataset_file "${VALIDATION_FILE}" \
|
||||
--validation_steps 200 --validation_sampling_steps 50 --validation_guidance_scale 3.0 \
|
||||
--learning_rate 5e-5 --mixed_precision bf16 --weight_decay 1e-4 --max_grad_norm 1.0 \
|
||||
--weight_only_checkpointing_steps 500 --training_state_checkpointing_steps 500 \
|
||||
--tracker_project_name wan_t2v_qat_finetune --output_dir checkpoints/wan_t2v_qat_finetune \
|
||||
--inference_mode False --training_cfg_rate 0.1 --not_apply_cfg_solver \
|
||||
--dit_precision fp32 --num_euler_timesteps 50 --ema_start_step 0 \
|
||||
--multi_phased_distill_schedule "4000-1"
|
||||
@@ -50,6 +50,21 @@ include_directories(
|
||||
set(FASTVIDEO_KERNEL_BUILD_TK "AUTO" CACHE STRING "Build ThunderKittens kernels: AUTO/ON/OFF")
|
||||
set_property(CACHE FASTVIDEO_KERNEL_BUILD_TK PROPERTY STRINGS AUTO ON OFF)
|
||||
|
||||
set(_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT "AUTO")
|
||||
if(DEFINED FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3 AND NOT DEFINED CACHE{FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER})
|
||||
set(_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT "${FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3}")
|
||||
endif()
|
||||
|
||||
set(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER "${_FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER_DEFAULT}" CACHE STRING
|
||||
"Build attn_qat_infer Blackwell inference kernels: AUTO/ON/OFF")
|
||||
set_property(CACHE FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER PROPERTY STRINGS AUTO ON OFF)
|
||||
|
||||
if(DEFINED FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3)
|
||||
message(DEPRECATION
|
||||
"FASTVIDEO_KERNEL_BUILD_MODIFIED_SAGE3 is deprecated. "
|
||||
"Use FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER instead.")
|
||||
endif()
|
||||
|
||||
# Prefer environment variable (used by CI) if CMake var is not explicitly set.
|
||||
if(NOT DEFINED TORCH_CUDA_ARCH_LIST AND DEFINED ENV{TORCH_CUDA_ARCH_LIST})
|
||||
set(TORCH_CUDA_ARCH_LIST "$ENV{TORCH_CUDA_ARCH_LIST}")
|
||||
@@ -57,6 +72,7 @@ endif()
|
||||
|
||||
message(STATUS "TORCH_CUDA_ARCH_LIST (cmake/env): ${TORCH_CUDA_ARCH_LIST}")
|
||||
message(STATUS "FASTVIDEO_KERNEL_BUILD_TK: ${FASTVIDEO_KERNEL_BUILD_TK}")
|
||||
message(STATUS "FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER: ${FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER}")
|
||||
|
||||
set(ENABLE_TK_KERNELS OFF)
|
||||
if(FASTVIDEO_KERNEL_BUILD_TK STREQUAL "ON")
|
||||
@@ -91,6 +107,54 @@ else()
|
||||
message(STATUS "ThunderKittens kernels: DISABLED (will use Triton fallbacks at runtime)")
|
||||
endif()
|
||||
|
||||
set(ENABLE_ATTN_QAT_INFER OFF)
|
||||
if(GPU_BACKEND STREQUAL "ROCM")
|
||||
message(STATUS "attn_qat_infer kernels: DISABLED (ROCm build)")
|
||||
else()
|
||||
set(_WANTS_ATTN_QAT_INFER OFF)
|
||||
if(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER STREQUAL "ON")
|
||||
set(_WANTS_ATTN_QAT_INFER ON)
|
||||
elseif(FASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER STREQUAL "AUTO")
|
||||
if(TORCH_CUDA_ARCH_LIST)
|
||||
string(REGEX MATCH
|
||||
"(^|[; ,])((12\\.0a)|(120a)|(sm_120a))([; ,]|$)"
|
||||
_HAS_120A "${TORCH_CUDA_ARCH_LIST}")
|
||||
if(_HAS_120A)
|
||||
set(_WANTS_ATTN_QAT_INFER ON)
|
||||
endif()
|
||||
else()
|
||||
execute_process(
|
||||
COMMAND "${Python_EXECUTABLE}" -c
|
||||
"import torch; print('1' if (torch.cuda.is_available() and torch.version.cuda and torch.cuda.get_device_capability()[0] >= 12) else '0')"
|
||||
OUTPUT_VARIABLE _LOCAL_HAS_BLACKWELL
|
||||
OUTPUT_STRIP_TRAILING_WHITESPACE
|
||||
ERROR_QUIET
|
||||
)
|
||||
if(_LOCAL_HAS_BLACKWELL STREQUAL "1")
|
||||
set(_WANTS_ATTN_QAT_INFER ON)
|
||||
endif()
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(_WANTS_ATTN_QAT_INFER)
|
||||
if(CUDAToolkit_VERSION VERSION_LESS 12.8)
|
||||
message(WARNING
|
||||
"attn_qat_infer kernels require CUDA Toolkit 12.8+. "
|
||||
"Skipping because CUDAToolkit_VERSION=${CUDAToolkit_VERSION}.")
|
||||
else()
|
||||
set(ENABLE_ATTN_QAT_INFER ON)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(ENABLE_ATTN_QAT_INFER)
|
||||
message(STATUS "attn_qat_infer kernels: ENABLED")
|
||||
else()
|
||||
message(STATUS
|
||||
"attn_qat_infer kernels: DISABLED "
|
||||
"(requires CUDA 12.8+ and Blackwell sm_120a)")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
# Always try to build the extension if CUDA is available, but conditionally add sources/flags
|
||||
set(BUILD_CXX_KERNELS ON)
|
||||
|
||||
@@ -183,3 +247,74 @@ if(BUILD_CXX_KERNELS)
|
||||
install(TARGETS fastvideo_kernel_ops LIBRARY DESTINATION fastvideo_kernel/_C)
|
||||
endif()
|
||||
|
||||
if(ENABLE_ATTN_QAT_INFER)
|
||||
set(ATTN_QAT_INFER_DIR ${CMAKE_SOURCE_DIR}/attn_qat_infer)
|
||||
set(ATTN_QAT_INFER_INCLUDE_DIRS
|
||||
${ATTN_QAT_INFER_DIR}
|
||||
${CMAKE_SOURCE_DIR}/include/cutlass/include
|
||||
${CMAKE_SOURCE_DIR}/include/cutlass/tools/util/include
|
||||
${TORCH_INCLUDE_DIRS}
|
||||
)
|
||||
set(ATTN_QAT_INFER_CUDA_FLAGS
|
||||
"-O3"
|
||||
"-std=c++17"
|
||||
"-U__CUDA_NO_HALF_OPERATORS__"
|
||||
"-U__CUDA_NO_HALF_CONVERSIONS__"
|
||||
"-U__CUDA_NO_BFLOAT16_OPERATORS__"
|
||||
"-U__CUDA_NO_BFLOAT16_CONVERSIONS__"
|
||||
"-U__CUDA_NO_BFLOAT162_OPERATORS__"
|
||||
"-U__CUDA_NO_BFLOAT162_CONVERSIONS__"
|
||||
"--expt-relaxed-constexpr"
|
||||
"--expt-extended-lambda"
|
||||
"--use_fast_math"
|
||||
"--ptxas-options=--verbose,--warn-on-local-memory-usage"
|
||||
"-lineinfo"
|
||||
"-DCUTLASS_DEBUG_TRACE_LEVEL=0"
|
||||
"-DNDEBUG"
|
||||
"-DQBLKSIZE=128"
|
||||
"-DKBLKSIZE=128"
|
||||
"-DCTA256"
|
||||
"-DDQINRMEM"
|
||||
)
|
||||
|
||||
Python_add_library(fp4attn_cuda MODULE WITH_SOABI
|
||||
attn_qat_infer/blackwell/api.cu
|
||||
)
|
||||
target_include_directories(fp4attn_cuda PRIVATE ${ATTN_QAT_INFER_INCLUDE_DIRS})
|
||||
target_compile_definitions(fp4attn_cuda PRIVATE TORCH_EXTENSION_NAME=fp4attn_cuda)
|
||||
target_compile_options(fp4attn_cuda PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:-O3 -std=c++17>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:${ATTN_QAT_INFER_CUDA_FLAGS}>
|
||||
)
|
||||
set_target_properties(fp4attn_cuda PROPERTIES
|
||||
CUDA_ARCHITECTURES "120a"
|
||||
CXX_STANDARD 17
|
||||
CUDA_STANDARD 17
|
||||
)
|
||||
target_link_libraries(fp4attn_cuda PRIVATE ${TORCH_LIBRARIES} CUDA::cudart CUDA::cuda_driver)
|
||||
|
||||
Python_add_library(fp4quant_cuda MODULE WITH_SOABI
|
||||
attn_qat_infer/quantization/fp4_quantization_4d.cu
|
||||
)
|
||||
target_include_directories(fp4quant_cuda PRIVATE ${ATTN_QAT_INFER_INCLUDE_DIRS})
|
||||
target_compile_definitions(fp4quant_cuda PRIVATE TORCH_EXTENSION_NAME=fp4quant_cuda)
|
||||
target_compile_options(fp4quant_cuda PRIVATE
|
||||
$<$<COMPILE_LANGUAGE:CXX>:-O3 -std=c++17>
|
||||
$<$<COMPILE_LANGUAGE:CUDA>:${ATTN_QAT_INFER_CUDA_FLAGS}>
|
||||
)
|
||||
set_target_properties(fp4quant_cuda PROPERTIES
|
||||
CUDA_ARCHITECTURES "120a"
|
||||
CXX_STANDARD 17
|
||||
CUDA_STANDARD 17
|
||||
)
|
||||
target_link_libraries(fp4quant_cuda PRIVATE ${TORCH_LIBRARIES} CUDA::cudart CUDA::cuda_driver)
|
||||
|
||||
if(TORCH_PYTHON_LIBRARY_PATH)
|
||||
target_link_libraries(fp4attn_cuda PRIVATE "${TORCH_PYTHON_LIBRARY_PATH}")
|
||||
target_link_libraries(fp4quant_cuda PRIVATE "${TORCH_PYTHON_LIBRARY_PATH}")
|
||||
endif()
|
||||
|
||||
install(TARGETS fp4attn_cuda LIBRARY DESTINATION .)
|
||||
install(TARGETS fp4quant_cuda LIBRARY DESTINATION .)
|
||||
endif()
|
||||
|
||||
|
||||
@@ -2,5 +2,6 @@ include LICENSE
|
||||
include README.md
|
||||
include pyproject.toml
|
||||
recursive-include python/fastvideo_kernel *.py
|
||||
recursive-include attn_qat_infer *.py *.cu *.cuh *.cpp *.h
|
||||
recursive-include csrc *.cu *.cuh *.cpp *.h
|
||||
recursive-include include/tk *.cu *.cuh *.cpp *.h *.src
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
from .api import sageattn_blackwell
|
||||
@@ -0,0 +1,189 @@
|
||||
# Modified from the original SageATtention3 code
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
import torch
|
||||
import triton
|
||||
import triton.language as tl
|
||||
import torch.nn.functional as F
|
||||
from typing import Tuple
|
||||
from torch.nn.functional import scaled_dot_product_attention as sdpa
|
||||
import fp4attn_cuda
|
||||
import fp4quant_cuda
|
||||
|
||||
# Centralized block size configuration for sageattn_blackwell kernels
|
||||
# These should match the values in fastvideo/attention/backends/sageattn/blackwell/block_config.h
|
||||
BLOCK_M = 128 # Block size for M dimension (query sequence length)
|
||||
BLOCK_N = 128 # Block size for N dimension (key/value sequence length)
|
||||
|
||||
|
||||
@triton.jit
|
||||
def group_mean_kernel(
|
||||
q_ptr,
|
||||
q_out_ptr,
|
||||
qm_out_ptr,
|
||||
B, H, L, D: tl.constexpr,
|
||||
stride_qb, stride_qh, stride_ql, stride_qd,
|
||||
stride_qmb, stride_qmh, stride_qml, stride_qmd,
|
||||
GROUP_SIZE: tl.constexpr
|
||||
):
|
||||
pid_b = tl.program_id(0)
|
||||
pid_h = tl.program_id(1)
|
||||
pid_group = tl.program_id(2)
|
||||
|
||||
group_start = pid_group * GROUP_SIZE
|
||||
offsets = group_start + tl.arange(0, GROUP_SIZE)
|
||||
|
||||
q_offsets = pid_b * stride_qb + pid_h * stride_qh + offsets[:, None] * stride_ql + tl.arange(0, D)[None, :] * stride_qd
|
||||
q_group = tl.load(q_ptr + q_offsets)
|
||||
|
||||
qm_group = tl.sum(q_group, axis=0) / GROUP_SIZE
|
||||
|
||||
q_group = q_group - qm_group
|
||||
tl.store(q_out_ptr + q_offsets, q_group)
|
||||
|
||||
qm_offset = pid_b * stride_qmb + pid_h * stride_qmh + pid_group * stride_qml + tl.arange(0, D) * stride_qmd
|
||||
tl.store(qm_out_ptr + qm_offset, qm_group)
|
||||
|
||||
|
||||
def triton_group_mean(q: torch.Tensor):
|
||||
B, H, L, D = q.shape
|
||||
GROUP_SIZE = BLOCK_M
|
||||
num_groups = L // GROUP_SIZE
|
||||
|
||||
q_out = torch.empty_like(q) # [B, H, L, D]
|
||||
qm = torch.empty(B, H, num_groups, D, device=q.device, dtype=q.dtype)
|
||||
|
||||
grid = (B, H, num_groups)
|
||||
|
||||
group_mean_kernel[grid](
|
||||
q, q_out, qm,
|
||||
B, H, L, D,
|
||||
q.stride(0), q.stride(1), q.stride(2), q.stride(3),
|
||||
qm.stride(0), qm.stride(1), qm.stride(2), qm.stride(3),
|
||||
GROUP_SIZE=GROUP_SIZE
|
||||
)
|
||||
return q_out, qm
|
||||
|
||||
|
||||
def preprocess_qkv(q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, per_block_mean: bool = True, enable_smoothing_q: bool = False, enable_smoothing_k: bool = False):
|
||||
|
||||
def pad_to_block_size(x):
|
||||
L = x.size(2)
|
||||
pad_len = (BLOCK_M - L % BLOCK_M) % BLOCK_M
|
||||
if pad_len == 0:
|
||||
return x.contiguous()
|
||||
return F.pad(x, (0, 0, 0, pad_len), value=0).contiguous()
|
||||
|
||||
if enable_smoothing_k:
|
||||
k -= k.mean(dim=-2, keepdim=True)
|
||||
q, k, v = map(lambda x: pad_to_block_size(x), [q, k, v])
|
||||
if per_block_mean and enable_smoothing_q:
|
||||
q, qm = triton_group_mean(q)
|
||||
elif enable_smoothing_q:
|
||||
qm = q.mean(dim=-2, keepdim=True)
|
||||
q = q - qm
|
||||
if enable_smoothing_q:
|
||||
delta_s = torch.matmul(qm, k.transpose(-2, -1)).to(torch.float32).contiguous()
|
||||
else: # used to disable q smoothing
|
||||
B, H, L, D = q.shape
|
||||
delta_s = torch.zeros((B, H, L // BLOCK_M, k.shape[2]), device=q.device, dtype=torch.float32)
|
||||
|
||||
return q, k, v, delta_s
|
||||
|
||||
def scale_and_quant_fp4(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
assert x.ndim == 4
|
||||
B, H, N, D = x.shape
|
||||
packed_fp4 = torch.empty((B, H, N, D // 2), device=x.device, dtype=torch.uint8)
|
||||
fp8_scale = torch.empty((B, H, N, D // 16), device=x.device, dtype=torch.float8_e4m3fn)
|
||||
fp4quant_cuda.scaled_fp4_quant(x, packed_fp4, fp8_scale, 1)
|
||||
return packed_fp4, fp8_scale
|
||||
|
||||
def scale_and_quant_fp4_permute(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
assert x.ndim == 4
|
||||
B, H, N, D = x.shape
|
||||
packed_fp4 = torch.empty((B, H, N, D // 2), device=x.device, dtype=torch.uint8)
|
||||
fp8_scale = torch.empty((B, H, N, D // 16), device=x.device, dtype=torch.float8_e4m3fn)
|
||||
fp4quant_cuda.scaled_fp4_quant_permute(x, packed_fp4, fp8_scale, 1)
|
||||
return packed_fp4, fp8_scale
|
||||
|
||||
def scale_and_quant_fp4_transpose(x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
assert x.ndim == 4
|
||||
B, H, N, D = x.shape
|
||||
packed_fp4 = torch.empty((B, H, D, N // 2), device=x.device, dtype=torch.uint8)
|
||||
fp8_scale = torch.empty((B, H, D, N // 16), device=x.device, dtype=torch.float8_e4m3fn)
|
||||
fp4quant_cuda.scaled_fp4_quant_trans(x, packed_fp4, fp8_scale, 1)
|
||||
return packed_fp4, fp8_scale
|
||||
|
||||
def blockscaled_fp4_attn(qlist: Tuple,
|
||||
klist: Tuple,
|
||||
vlist: Tuple,
|
||||
delta_s: torch.Tensor,
|
||||
KL: int,
|
||||
is_causal: bool = False,
|
||||
per_block_mean: bool = True,
|
||||
is_bf16: bool = True,
|
||||
single_level_p_quant: bool = False,
|
||||
sm_scale: float | None = None
|
||||
):
|
||||
softmax_scale = sm_scale if sm_scale is not None else (qlist[0].shape[-1] * 2) ** (-0.5)
|
||||
return fp4attn_cuda.fwd(qlist[0], klist[0], vlist[0], qlist[1], klist[1], vlist[1], delta_s, KL, None, softmax_scale, is_causal, per_block_mean, is_bf16, single_level_p_quant)
|
||||
|
||||
|
||||
def sageattn_blackwell(q, k, v, attn_mask = None, is_causal = False, per_block_mean = True, single_level_p_quant = True, sm_scale: float | None = None, **kwargs):
|
||||
"""
|
||||
SageAttention3 Blackwell kernel for FP4 attention.
|
||||
|
||||
Args:
|
||||
q: Query tensor [B, H, L, D]
|
||||
k: Key tensor [B, H, L, D]
|
||||
v: Value tensor [B, H, L, D]
|
||||
attn_mask: Attention mask (not used)
|
||||
is_causal: Whether to use causal masking
|
||||
per_block_mean: Whether to use per-block mean for Q smoothing
|
||||
single_level_p_quant: If True, use single-level quantization: s_P2, P̂_2 = φ(P̃) directly
|
||||
(standard per-block FP4 quantization like V, no s_P1).
|
||||
If False (default), use two-level quantization:
|
||||
s_P1 = rowmax(P̃)/(448×6), then s_P2, P̂_2 = φ(P̃/s_P1).
|
||||
sm_scale: Softmax scale to pass through to the CUDA kernel. If None,
|
||||
defaults to the kernel's 1/sqrt(D) scale.
|
||||
**kwargs: Additional arguments (ignored)
|
||||
|
||||
Returns:
|
||||
Output tensor [B, H, L, D]
|
||||
"""
|
||||
if q.size(-1) >= 256:
|
||||
print(f"Unsupported Headdim {q.size(-1)}")
|
||||
return sdpa(q, k, v, is_causal = is_causal)
|
||||
QL = q.size(2)
|
||||
KL = k.size(2)
|
||||
is_bf16 = q.dtype == torch.bfloat16
|
||||
q, k, v, delta_s = preprocess_qkv(q, k, v, per_block_mean)
|
||||
qlist_from_cuda = scale_and_quant_fp4(q)
|
||||
klist_from_cuda = scale_and_quant_fp4_permute(k)
|
||||
vlist_from_cuda = scale_and_quant_fp4_transpose(v)
|
||||
o_fp4 = blockscaled_fp4_attn(
|
||||
qlist_from_cuda,
|
||||
klist_from_cuda,
|
||||
vlist_from_cuda,
|
||||
delta_s,
|
||||
KL,
|
||||
is_causal,
|
||||
per_block_mean,
|
||||
is_bf16,
|
||||
single_level_p_quant,
|
||||
sm_scale
|
||||
)[0][:, :, :QL, :].contiguous()
|
||||
return o_fp4
|
||||
@@ -0,0 +1 @@
|
||||
__version__ = "3.0.0.b1"
|
||||
@@ -0,0 +1,347 @@
|
||||
// Modified from the original SageAttention3 code
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
// Include these 2 headers instead of torch/extension.h since we don't need all of the torch headers.
|
||||
#include <torch/python.h>
|
||||
#include <torch/nn/functional.h>
|
||||
#include <ATen/cuda/CUDAContext.h>
|
||||
#include <c10/cuda/CUDAGuard.h>
|
||||
|
||||
#include <cutlass/numeric_types.h>
|
||||
|
||||
#include "params.h"
|
||||
#include "launch.h"
|
||||
#include "static_switch.h"
|
||||
#include "block_config.h"
|
||||
|
||||
#define CHECK_DEVICE(x) TORCH_CHECK(x.is_cuda(), #x " must be on CUDA")
|
||||
#define CHECK_SHAPE(x, ...) TORCH_CHECK(x.sizes() == torch::IntArrayRef({__VA_ARGS__}), #x " must have shape (" #__VA_ARGS__ ")")
|
||||
#define CHECK_CONTIGUOUS(x) TORCH_CHECK(x.is_contiguous(), #x " must be contiguous")
|
||||
|
||||
|
||||
void set_params_fprop(Flash_fwd_params ¶ms,
|
||||
// sizes
|
||||
const size_t b,
|
||||
const size_t seqlen_q,
|
||||
const size_t seqlen_k,
|
||||
const size_t unpadded_seqlen_k,
|
||||
const size_t seqlen_q_rounded,
|
||||
const size_t seqlen_k_rounded,
|
||||
const size_t h,
|
||||
const size_t h_k,
|
||||
const size_t d,
|
||||
const size_t d_rounded,
|
||||
// device pointers
|
||||
const at::Tensor q,
|
||||
const at::Tensor k,
|
||||
const at::Tensor v,
|
||||
const at::Tensor delta_s,
|
||||
at::Tensor out,
|
||||
const at::Tensor sfq,
|
||||
const at::Tensor sfk,
|
||||
const at::Tensor sfv,
|
||||
void *cu_seqlens_q_d,
|
||||
void *cu_seqlens_k_d,
|
||||
void *seqused_k,
|
||||
void *p_d,
|
||||
void *softmax_lse_d,
|
||||
float p_dropout,
|
||||
float softmax_scale,
|
||||
int window_size_left,
|
||||
int window_size_right,
|
||||
bool per_block_mean,
|
||||
bool is_bf16,
|
||||
bool single_level_p_quant=false,
|
||||
bool seqlenq_ngroups_swapped=false) {
|
||||
|
||||
// Reset the parameters
|
||||
params = {};
|
||||
// Set the pointers and strides.
|
||||
params.q_ptr = q.data_ptr();
|
||||
params.k_ptr = k.data_ptr();
|
||||
params.v_ptr = v.data_ptr();
|
||||
params.delta_s_ptr = delta_s.data_ptr();
|
||||
params.sfq_ptr = sfq.data_ptr();
|
||||
params.sfk_ptr = sfk.data_ptr();
|
||||
params.sfv_ptr = sfv.data_ptr();
|
||||
|
||||
// All stride are in elements, not bytes.
|
||||
params.q_row_stride = q.stride(-2) * 2;
|
||||
params.k_row_stride = k.stride(-2) * 2;
|
||||
params.v_row_stride = v.stride(-2) * 2;;
|
||||
params.q_head_stride = q.stride(-3) * 2;
|
||||
params.k_head_stride = k.stride(-3) * 2;
|
||||
params.v_head_stride = v.stride(-3) * 2; // for packed q k v
|
||||
|
||||
params.ds_row_stride = delta_s.stride(-2);
|
||||
params.ds_head_stride = delta_s.stride(-3);
|
||||
|
||||
params.sfq_row_stride = sfq.stride(-2);
|
||||
params.sfk_row_stride = sfk.stride(-2);
|
||||
params.sfv_row_stride = sfv.stride(-2);
|
||||
params.sfq_head_stride = sfq.stride(-3);
|
||||
params.sfk_head_stride = sfk.stride(-3);
|
||||
params.sfv_head_stride = sfv.stride(-3);
|
||||
params.o_ptr = out.data_ptr();
|
||||
params.o_row_stride = out.stride(-2);
|
||||
params.o_head_stride = out.stride(-3);
|
||||
|
||||
if (cu_seqlens_q_d == nullptr) {
|
||||
params.q_batch_stride = q.stride(0) * 2;
|
||||
params.k_batch_stride = k.stride(0) * 2;
|
||||
params.v_batch_stride = v.stride(0) * 2;
|
||||
params.ds_batch_stride = delta_s.stride(0);
|
||||
params.sfq_batch_stride = sfq.stride(0);
|
||||
params.sfk_batch_stride = sfk.stride(0);
|
||||
params.sfv_batch_stride = sfv.stride(0);
|
||||
params.o_batch_stride = out.stride(0);
|
||||
if (seqlenq_ngroups_swapped) {
|
||||
params.q_batch_stride *= seqlen_q;
|
||||
params.o_batch_stride *= seqlen_q;
|
||||
}
|
||||
}
|
||||
|
||||
params.cu_seqlens_q = static_cast<int *>(cu_seqlens_q_d);
|
||||
params.cu_seqlens_k = static_cast<int *>(cu_seqlens_k_d);
|
||||
params.seqused_k = static_cast<int *>(seqused_k);
|
||||
|
||||
// P = softmax(QK^T)
|
||||
params.p_ptr = p_d;
|
||||
|
||||
// Softmax sum
|
||||
params.softmax_lse_ptr = softmax_lse_d;
|
||||
|
||||
// Set the dimensions.
|
||||
params.b = b;
|
||||
params.h = h;
|
||||
params.h_k = h_k;
|
||||
params.h_h_k_ratio = h / h_k;
|
||||
params.seqlen_q = seqlen_q;
|
||||
params.seqlen_k = seqlen_k;
|
||||
params.unpadded_seqlen_k = unpadded_seqlen_k;
|
||||
params.seqlen_q_rounded = seqlen_q_rounded;
|
||||
params.seqlen_k_rounded = seqlen_k_rounded;
|
||||
params.d = d;
|
||||
params.d_rounded = d_rounded;
|
||||
|
||||
params.head_divmod = cutlass::FastDivmod(int(h));
|
||||
|
||||
// Set the different scale values.
|
||||
params.scale_softmax = softmax_scale;
|
||||
params.scale_softmax_log2 = softmax_scale * M_LOG2E;
|
||||
__half scale_softmax_log2_half = __float2half(params.scale_softmax_log2);
|
||||
__half2 scale_softmax_log2_half2 = __half2(scale_softmax_log2_half, scale_softmax_log2_half);
|
||||
params.scale_softmax_log2_half2 = reinterpret_cast<uint32_t&>(scale_softmax_log2_half2);
|
||||
|
||||
// Set this to probability of keeping an element to simplify things.
|
||||
params.p_dropout = 1.f - p_dropout;
|
||||
// Convert p from float to int so we don't have to convert the random uint to float to compare.
|
||||
// [Minor] We want to round down since when we do the comparison we use <= instead of <
|
||||
// params.p_dropout_in_uint = uint32_t(std::floor(params.p_dropout * 4294967295.0));
|
||||
// params.p_dropout_in_uint16_t = uint16_t(std::floor(params.p_dropout * 65535.0));
|
||||
params.p_dropout_in_uint8_t = uint8_t(std::floor(params.p_dropout * 255.0));
|
||||
params.rp_dropout = 1.f / params.p_dropout;
|
||||
params.scale_softmax_rp_dropout = params.rp_dropout * params.scale_softmax;
|
||||
TORCH_CHECK(p_dropout < 1.f);
|
||||
#ifdef FLASHATTENTION_DISABLE_DROPOUT
|
||||
TORCH_CHECK(p_dropout == 0.0f, "This flash attention build does not support dropout.");
|
||||
#endif
|
||||
|
||||
// Causal is the special case where window_size_right == 0 and window_size_left < 0.
|
||||
// Local is the more general case where window_size_right >= 0 or window_size_left >= 0.
|
||||
params.is_causal = window_size_left < 0 && window_size_right == 0;
|
||||
params.per_block_mean = per_block_mean;
|
||||
if (per_block_mean) {
|
||||
params.seqlen_s = seqlen_q;
|
||||
} else {
|
||||
params.seqlen_s = flash::BLOCK_M; // size of BLOCK_M
|
||||
}
|
||||
if (window_size_left < 0 && window_size_right >= 0) { window_size_left = seqlen_k; }
|
||||
if (window_size_left >= 0 && window_size_right < 0) { window_size_right = seqlen_k; }
|
||||
params.window_size_left = window_size_left;
|
||||
params.window_size_right = window_size_right;
|
||||
|
||||
#ifdef FLASHATTENTION_DISABLE_LOCAL
|
||||
TORCH_CHECK(params.is_causal || (window_size_left < 0 && window_size_right < 0),
|
||||
"This flash attention build does not support local attention.");
|
||||
#endif
|
||||
|
||||
params.is_seqlens_k_cumulative = true;
|
||||
params.is_bf16 = is_bf16;
|
||||
params.single_level_p_quant = single_level_p_quant;
|
||||
#ifdef FLASHATTENTION_DISABLE_UNEVEN_K
|
||||
TORCH_CHECK(d == d_rounded, "This flash attention build does not support headdim not being a multiple of 32.");
|
||||
#endif
|
||||
}
|
||||
|
||||
template<bool IsBF16>
|
||||
void run_mha_fwd_dispatch_dtype(Flash_fwd_params ¶ms, cudaStream_t stream) {
|
||||
using OType = std::conditional_t<IsBF16, cutlass::bfloat16_t, cutlass::half_t>;
|
||||
if (params.d == 64) {
|
||||
run_mha_fwd_<cutlass::nv_float4_t<cutlass::float_e2m1_t>, 64, OType>(params, stream);
|
||||
} else if (params.d == 128) {
|
||||
run_mha_fwd_<cutlass::nv_float4_t<cutlass::float_e2m1_t>, 128, OType>(params, stream);
|
||||
}
|
||||
}
|
||||
|
||||
void run_mha_fwd(Flash_fwd_params ¶ms, cudaStream_t stream, bool force_split_kernel = false) {
|
||||
BOOL_SWITCH(params.is_bf16, IsBF16, ([&] {
|
||||
run_mha_fwd_dispatch_dtype<IsBF16>(params, stream);
|
||||
}));
|
||||
}
|
||||
|
||||
std::vector<at::Tensor>
|
||||
mha_fwd(at::Tensor &q, // batch_size x seqlen_q x num_heads x (head_size // 2)
|
||||
const at::Tensor &k, // batch_size x seqlen_k x num_heads_k x (head_size // 2)
|
||||
const at::Tensor &v, // batch_size x seqlen_k x num_heads_k x (head_size // 2)
|
||||
const at::Tensor &sfq,
|
||||
const at::Tensor &sfk,
|
||||
const at::Tensor &sfv,
|
||||
const at::Tensor &delta_s,
|
||||
int unpadded_k,
|
||||
c10::optional<at::Tensor> &out_, // batch_size x seqlen_q x num_heads x head_size
|
||||
const float softmax_scale,
|
||||
bool is_causal,
|
||||
bool per_block_mean,
|
||||
bool is_bf16,
|
||||
bool single_level_p_quant=false // If true, use only per-row scale s_P2 (no per-block s_P1)
|
||||
) {
|
||||
|
||||
auto dprops = at::cuda::getCurrentDeviceProperties();
|
||||
bool is_blackwell_or_newer = dprops->major >= 12;
|
||||
TORCH_CHECK(is_blackwell_or_newer, "only supports Blackwell GPUs or newer.");
|
||||
|
||||
auto q_dtype = q.dtype();
|
||||
auto sfq_dtype = sfq.dtype();
|
||||
TORCH_CHECK(q_dtype == torch::kUInt8, "q dtype must be uint8");
|
||||
TORCH_CHECK(k.dtype() == q_dtype, "query and key must have the same dtype");
|
||||
TORCH_CHECK(v.dtype() == q_dtype, "query and value must have the same dtype");
|
||||
CHECK_DEVICE(q); CHECK_DEVICE(k); CHECK_DEVICE(v);
|
||||
|
||||
TORCH_CHECK(sfq_dtype == torch::kFloat8_e4m3fn, "q dtype must be uint8");
|
||||
TORCH_CHECK(sfk.dtype() == sfq_dtype, "query and key must have the same dtype");
|
||||
TORCH_CHECK(sfv.dtype() == sfq_dtype, "query and value must have the same dtype");
|
||||
CHECK_DEVICE(sfq); CHECK_DEVICE(sfk); CHECK_DEVICE(sfv);
|
||||
|
||||
TORCH_CHECK(q.stride(-1) == 1, "Input tensor must have contiguous last dimension");
|
||||
TORCH_CHECK(k.stride(-1) == 1, "Input tensor must have contiguous last dimension");
|
||||
TORCH_CHECK(v.stride(-1) == 1, "Input tensor must have contiguous last dimension");
|
||||
TORCH_CHECK(delta_s.stride(-1) == 1, "Input tensor must have contiguous last dimension");
|
||||
|
||||
TORCH_CHECK(q.is_contiguous(), "Input tensor must be contiguous");
|
||||
TORCH_CHECK(k.is_contiguous(), "Input tensor must be contiguous");
|
||||
TORCH_CHECK(v.is_contiguous(), "Input tensor must be contiguous");
|
||||
|
||||
const auto sizes = q.sizes();
|
||||
auto opts = q.options();
|
||||
const int batch_size = sizes[0];
|
||||
int seqlen_q = sizes[2];
|
||||
int num_heads = sizes[1];
|
||||
const int head_size_og = sizes[3];
|
||||
const int unpacked_head_size = head_size_og * 2;
|
||||
const int seqlen_k = k.size(2);
|
||||
const int num_heads_k = k.size(1);
|
||||
|
||||
TORCH_CHECK(batch_size > 0, "batch size must be postive");
|
||||
TORCH_CHECK(unpacked_head_size <= 256, "FlashAttention forward only supports head dimension at most 256");
|
||||
TORCH_CHECK(num_heads % num_heads_k == 0, "Number of heads in key/value must divide number of heads in query");
|
||||
TORCH_CHECK(num_heads == num_heads_k, "We do not support MQA/GQA yet");
|
||||
|
||||
TORCH_CHECK(unpacked_head_size == 64 || unpacked_head_size == 128 || unpacked_head_size == 256, "Only support head size 64, 128, and 256 for now");
|
||||
|
||||
CHECK_SHAPE(q, batch_size, num_heads, seqlen_q, head_size_og);
|
||||
CHECK_SHAPE(k, batch_size, num_heads_k, seqlen_k, head_size_og);
|
||||
CHECK_SHAPE(v, batch_size, num_heads_k, unpacked_head_size, seqlen_k/2);
|
||||
// CHECK_SHAPE(delta_s, batch_size, num_heads, seqlen_q / 128, seqlen_k);
|
||||
// CHECK_SHAPE(sfq, batch_size, seqlen_q, num_heads, unpacked_head_size);
|
||||
// CHECK_SHAPE(sfk, batch_size, seqlen_k, num_heads_k, unpacked_head_size);
|
||||
// CHECK_SHAPE(sfv, batch_size, unpacked_head_size, num_heads_k, seqlen_k);
|
||||
TORCH_CHECK(unpacked_head_size % 8 == 0, "head_size must be a multiple of 8");
|
||||
|
||||
auto dtype = is_bf16 ? at::ScalarType::BFloat16 : at::ScalarType::Half;
|
||||
at::Tensor out = torch::empty({batch_size, num_heads, seqlen_q, unpacked_head_size}, opts.dtype(dtype));
|
||||
|
||||
auto round_multiple = [](int x, int m) { return (x + m - 1) / m * m; };
|
||||
// const int head_size = round_multiple(head_size_og, 8);
|
||||
// const int head_size_rounded = round_multiple(head_size, 32);
|
||||
const int seqlen_q_rounded = round_multiple(seqlen_q, flash::BLOCK_M);
|
||||
const int seqlen_k_rounded = round_multiple(seqlen_k, flash::BLOCK_N);
|
||||
|
||||
// Otherwise the kernel will be launched from cuda:0 device
|
||||
// Cast to char to avoid compiler warning about narrowing
|
||||
at::cuda::CUDAGuard device_guard{(char)q.get_device()};
|
||||
|
||||
|
||||
|
||||
auto softmax_lse = torch::empty({batch_size, num_heads, seqlen_q}, opts.dtype(at::kFloat));
|
||||
at::Tensor p;
|
||||
|
||||
Flash_fwd_params params;
|
||||
set_params_fprop(params,
|
||||
batch_size,
|
||||
seqlen_q, seqlen_k, unpadded_k,
|
||||
seqlen_q_rounded, seqlen_k_rounded,
|
||||
num_heads, num_heads_k,
|
||||
unpacked_head_size, unpacked_head_size,
|
||||
q, k, v, delta_s, out,
|
||||
sfq, sfk, sfv,
|
||||
/*cu_seqlens_q_d=*/nullptr,
|
||||
/*cu_seqlens_k_d=*/nullptr,
|
||||
/*seqused_k=*/nullptr,
|
||||
nullptr,
|
||||
softmax_lse.data_ptr(),
|
||||
/*p_dropout=*/0.f,
|
||||
softmax_scale,
|
||||
/*window_size_left=*/-1,
|
||||
/*window_size_right=*/is_causal ? 0 : -1,
|
||||
per_block_mean,
|
||||
is_bf16,
|
||||
single_level_p_quant
|
||||
);
|
||||
// StaticPersistentTileScheduler does not use tile_count_semaphore; avoid a
|
||||
// stack-local tensor whose data pointer would dangle after mha_fwd returns
|
||||
// while the async kernel may still be running.
|
||||
params.tile_count_semaphore = nullptr;
|
||||
|
||||
if (seqlen_k > 0) {
|
||||
auto stream = at::cuda::getCurrentCUDAStream().stream();
|
||||
run_mha_fwd(params, stream);
|
||||
} else {
|
||||
// If seqlen_k == 0, then we have an empty tensor. We need to set the output to 0.
|
||||
out.zero_();
|
||||
softmax_lse.fill_(std::numeric_limits<float>::infinity());
|
||||
}
|
||||
|
||||
// at::Tensor out_padded = out;
|
||||
// if (head_size_og % 8 != 0) {
|
||||
// out = out.index({"...", torch::indexing::Slice(torch::indexing::None, head_size_og)});
|
||||
// if (out_.has_value()) { out_.value().copy_(out); }
|
||||
// }
|
||||
|
||||
// return {out, q_padded, k_padded, v_padded, out_padded, softmax_lse, p};
|
||||
// cudaDeviceSynchronize();
|
||||
// auto err = cudaGetLastError();
|
||||
// printf("%s\n", cudaGetErrorString(err));
|
||||
return {out, softmax_lse};
|
||||
}
|
||||
|
||||
|
||||
|
||||
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
|
||||
m.doc() = "FlashAttention";
|
||||
m.def("fwd", &mha_fwd, "Forward pass");
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
// Centralized block size configuration for sageattn_blackwell kernels
|
||||
// Block sizes for M and N dimensions
|
||||
namespace flash {
|
||||
// Block size for M dimension (query sequence length)
|
||||
static constexpr int BLOCK_M = 128;
|
||||
|
||||
// Block size for N dimension (key/value sequence length)
|
||||
static constexpr int BLOCK_N = 128;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,60 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* This code is based on code from FlashAttention3, https://github.com/Dao-AILab/flash-attention
|
||||
* Copyright (c) 2024, Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao.
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
namespace flash {
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<bool Varlen=true>
|
||||
struct BlockInfo {
|
||||
|
||||
template<typename Params>
|
||||
__device__ BlockInfo(const Params ¶ms, const int bidb)
|
||||
: sum_s_q(!Varlen || params.cu_seqlens_q == nullptr ? -1 : params.cu_seqlens_q[bidb])
|
||||
, sum_s_k(!Varlen || params.cu_seqlens_k == nullptr || !params.is_seqlens_k_cumulative ? -1 : params.cu_seqlens_k[bidb])
|
||||
, actual_seqlen_q(!Varlen || params.cu_seqlens_q == nullptr ? params.seqlen_q : params.cu_seqlens_q[bidb + 1] - sum_s_q)
|
||||
// If is_seqlens_k_cumulative, then seqlen_k is cu_seqlens_k[bidb + 1] - cu_seqlens_k[bidb].
|
||||
// Otherwise it's cu_seqlens_k[bidb], i.e., we use cu_seqlens_k to store the sequence lengths of K.
|
||||
, seqlen_k_cache(!Varlen || params.cu_seqlens_k == nullptr ? params.seqlen_k : (params.is_seqlens_k_cumulative ? params.cu_seqlens_k[bidb + 1] - sum_s_k : params.cu_seqlens_k[bidb]))
|
||||
, actual_seqlen_k(params.seqused_k ? params.seqused_k[bidb] : seqlen_k_cache + (params.knew_ptr == nullptr ? 0 : params.seqlen_knew))
|
||||
{
|
||||
}
|
||||
|
||||
template <typename index_t>
|
||||
__forceinline__ __device__ index_t q_offset(const index_t batch_stride, const index_t row_stride, const int bidb) const {
|
||||
return sum_s_q == -1 ? bidb * batch_stride : uint32_t(sum_s_q) * row_stride;
|
||||
}
|
||||
|
||||
template <typename index_t>
|
||||
__forceinline__ __device__ index_t k_offset(const index_t batch_stride, const index_t row_stride, const int bidb) const {
|
||||
return sum_s_k == -1 ? bidb * batch_stride : uint32_t(sum_s_k) * row_stride;
|
||||
}
|
||||
|
||||
const int sum_s_q;
|
||||
const int sum_s_k;
|
||||
const int actual_seqlen_q;
|
||||
// We have to have seqlen_k_cache declared before actual_seqlen_k, otherwise actual_seqlen_k is set to 0.
|
||||
const int seqlen_k_cache;
|
||||
const int actual_seqlen_k;
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
} // namespace flash
|
||||
@@ -0,0 +1,149 @@
|
||||
/***************************************************************************************************
|
||||
* Copyright (c) 2023 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
* SPDX-License-Identifier: BSD-3-Clause
|
||||
*
|
||||
* Redistribution and use in source and binary forms, with or without
|
||||
* modification, are permitted provided that the following conditions are met:
|
||||
*
|
||||
* 1. Redistributions of source code must retain the above copyright notice, this
|
||||
* list of conditions and the following disclaimer.
|
||||
*
|
||||
* 2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
* this list of conditions and the following disclaimer in the documentation
|
||||
* and/or other materials provided with the distribution.
|
||||
*
|
||||
* 3. Neither the name of the copyright holder nor the names of its
|
||||
* contributors may be used to endorse or promote products derived from
|
||||
* this software without specific prior written permission.
|
||||
*
|
||||
* THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
* AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
* IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
* DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
* FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
* DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
* SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
* CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
* OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
* OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
*
|
||||
**************************************************************************************************/
|
||||
|
||||
/*! \file
|
||||
\brief Blocked Scale configs specific for SM100 BlockScaled MMA
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "cutlass/layout/matrix.h"
|
||||
|
||||
#include "cute/int_tuple.hpp"
|
||||
#include "cute/atom/mma_traits_sm100.hpp"
|
||||
|
||||
namespace flash {
|
||||
|
||||
/////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
using namespace cute;
|
||||
|
||||
template<int SFVecSize, UMMA::Major major = UMMA::Major::K>
|
||||
struct BlockScaledBasicChunk {
|
||||
|
||||
using Blk_MN = _64;
|
||||
using Blk_SF = _4;
|
||||
|
||||
using SfAtom = Layout< Shape< Shape<_16,_4>, Shape<Int<SFVecSize>, _4>>,
|
||||
Stride<Stride<_16,_4>, Stride< _0, _1>>>;
|
||||
};
|
||||
|
||||
template<int SFVecSize_>
|
||||
struct BlockScaledConfig {
|
||||
// We are creating the SFA and SFB tensors' layouts in the collective since they always have the same layout.
|
||||
// k-major order
|
||||
static constexpr int SFVecSize = SFVecSize_;
|
||||
static constexpr int MMA_NSF = 4; // SFVecSize, MMA_NSF
|
||||
using BlkScaledChunk = BlockScaledBasicChunk<SFVecSize>;
|
||||
using Blk_MN = _64;
|
||||
using Blk_SF = _4;
|
||||
using mnBasicBlockShape = Shape<_16,_4>;
|
||||
using mnBasicBlockStride = Stride<_16,_4>;
|
||||
using kBasicBlockShape = Shape<Int<SFVecSize>, Int<MMA_NSF>>; // SFVecSize, MMA_NSF
|
||||
using kBasicBlockStride = Stride<_0, _1>;
|
||||
using SfAtom = Layout< Shape< mnBasicBlockShape, kBasicBlockShape>,
|
||||
Stride<mnBasicBlockStride, kBasicBlockStride>>;
|
||||
|
||||
using LayoutSF = decltype(blocked_product(SfAtom{},
|
||||
make_layout(
|
||||
make_shape(int32_t(0), int32_t(0), int32_t(0), int32_t(0)),
|
||||
make_stride(int32_t(0), _1{}, int32_t(0), int32_t(0)))));
|
||||
// A single indivisible block will hold 4 scale factors of 64 rows/columns (A/B matrix).
|
||||
// 4 is chosen to make consecutive 32bits of data to have scale factors for only a single row (col). 32bits corresponds to the TMEM word size
|
||||
using Blk_Elems = decltype(Blk_MN{} * Blk_SF{});
|
||||
using sSF_strideMN = decltype(prepend(Blk_Elems{}, mnBasicBlockStride{}));
|
||||
|
||||
|
||||
// The following function is provided for user fill dynamic problem size to the layout_SFA.
|
||||
template < class ProblemShape>
|
||||
CUTE_HOST_DEVICE
|
||||
static constexpr auto
|
||||
tile_atom_to_shape_SFQKV(ProblemShape problem_shape) {
|
||||
auto [Seqlen, Dim, HeadNum, Batch] = problem_shape;
|
||||
return tile_to_shape(SfAtom{}, make_shape(Seqlen, Dim, HeadNum, Batch), Step<_2,_1,_3,_4>{});
|
||||
}
|
||||
|
||||
// The following function is provided for user fill dynamic problem size to the layout_SFB.
|
||||
template <class ProblemShape>
|
||||
CUTE_HOST_DEVICE
|
||||
static constexpr auto
|
||||
tile_atom_to_shape_SFVt(ProblemShape problem_shape) {
|
||||
auto [Dim, Seqlen, HeadNum, Batch] = problem_shape;
|
||||
return tile_to_shape(SfAtom{}, make_shape(Dim, Seqlen, HeadNum, Batch), Step<_2,_1,_3,_4>{});
|
||||
}
|
||||
|
||||
template<class TiledMma, class TileShape_MNK>
|
||||
CUTE_HOST_DEVICE
|
||||
static constexpr auto
|
||||
deduce_smem_layoutSFQ(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
|
||||
|
||||
using sSFQ_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
|
||||
using sSFQ_shapeM = decltype(prepend(size<0>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
|
||||
using sSFQ_strideM = sSF_strideMN;
|
||||
using sSFQ_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<0>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
|
||||
using sSFQ_shape = decltype(make_shape(sSFQ_shapeM{}, sSFQ_shapeK{}));
|
||||
using sSFQ_stride = decltype(make_stride(sSFQ_strideM{}, sSFQ_strideK{}));
|
||||
using SmemLayoutAtomSFQ = decltype(make_layout(sSFQ_shape{}, sSFQ_stride{}));
|
||||
return SmemLayoutAtomSFQ{};
|
||||
}
|
||||
|
||||
template<class TiledMma, class TileShape_MNK>
|
||||
CUTE_HOST_DEVICE
|
||||
static constexpr auto
|
||||
deduce_smem_layoutSFKV(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
|
||||
|
||||
using sSFK_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
|
||||
using sSFK_shapeN = decltype(prepend(size<1>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
|
||||
using sSFK_strideN = sSF_strideMN;
|
||||
using sSFK_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<1>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
|
||||
using sSFK_shape = decltype(make_shape(sSFK_shapeN{}, sSFK_shapeK{}));
|
||||
using sSFK_stride = decltype(make_stride(sSFK_strideN{}, sSFK_strideK{}));
|
||||
using SmemLayoutAtomSFK = decltype(make_layout(sSFK_shape{}, sSFK_stride{}));
|
||||
return SmemLayoutAtomSFK{};
|
||||
}
|
||||
|
||||
template<class TiledMma, class TileShape_MNK>
|
||||
CUTE_HOST_DEVICE
|
||||
static constexpr auto
|
||||
deduce_smem_layoutSFVt(TiledMma tiled_mma, TileShape_MNK tileshape_mnk) {
|
||||
|
||||
using sSFVt_shapeK = decltype(prepend(make_shape(Blk_SF{}/Int<MMA_NSF>{}, size<2>(TileShape_MNK{}) / Int<SFVecSize>{} / Blk_SF{}), kBasicBlockShape{}));
|
||||
using sSFVt_shapeN = decltype(prepend(size<1>(TileShape_MNK{}) / Blk_MN{}, mnBasicBlockShape{}));
|
||||
using sSFVt_strideN = sSF_strideMN;
|
||||
using sSFVt_strideK = decltype(prepend(make_stride(Int<MMA_NSF>{}, size<1>(TileShape_MNK{}) / Blk_MN{} * Blk_Elems{}), kBasicBlockStride{}));
|
||||
using sSFVt_shape = decltype(make_shape(sSFVt_shapeN{}, sSFVt_shapeK{}));
|
||||
using sSFVt_stride = decltype(make_stride(sSFVt_strideN{}, sSFVt_strideK{}));
|
||||
using SmemLayoutAtomSFVt = decltype(make_layout(sSFVt_shape{}, sSFVt_stride{}));
|
||||
return SmemLayoutAtomSFVt{};
|
||||
}
|
||||
};
|
||||
|
||||
|
||||
} // namespace flash
|
||||
@@ -0,0 +1,327 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
#include "cute/arch/mma_sm120.hpp"
|
||||
#include "cute/atom/mma_traits_sm120.hpp"
|
||||
#include "cute/atom/mma_atom.hpp"
|
||||
#include "cutlass/cutlass.h"
|
||||
#include "cutlass/float8.h"
|
||||
#include "cutlass/float_subbyte.h"
|
||||
|
||||
namespace cute::SM120::BLOCKSCALED {
|
||||
|
||||
using cutlass::float_e2m1_t;
|
||||
using cutlass::float_ue4m3_t;
|
||||
|
||||
// MMA.SF 16x32x64 TN E2M1 x E2M1 with SF E4M3
|
||||
struct SM120_16x32x64_TN_VS_NVFP4 {
|
||||
using DRegisters = float[16];
|
||||
using ARegisters = uint32_t[4];
|
||||
using BRegisters = uint32_t[8];
|
||||
using CRegisters = float[16];
|
||||
|
||||
static constexpr int SFBits = 32;
|
||||
using RegTypeSF = cute::uint_bit_t<SFBits>;
|
||||
|
||||
using SFARegisters = RegTypeSF[1];
|
||||
using SFBRegisters = RegTypeSF[1];
|
||||
|
||||
CUTE_HOST_DEVICE static void
|
||||
fma(float & d0 , float & d1 , float & d2 , float & d3 ,
|
||||
float & d4 , float & d5 , float & d6 , float & d7 ,
|
||||
float & d8 , float & d9 , float & d10, float & d11,
|
||||
float & d12, float & d13, float & d14, float & d15,
|
||||
uint32_t const& a0 , uint32_t const& a1 , uint32_t const& a2 , uint32_t const& a3 ,
|
||||
uint32_t const& b0 , uint32_t const& b1 , uint32_t const& b2 , uint32_t const& b3 ,
|
||||
uint32_t const& b4 , uint32_t const& b5 , uint32_t const& b6 , uint32_t const& b7 ,
|
||||
float const & c0 , float const & c1 , float const & c2 , float const & c3 ,
|
||||
float const & c4 , float const & c5 , float const & c6 , float const & c7 ,
|
||||
float const & c8 , float const & c9 , float const & c10 , float const & c11,
|
||||
float const & c12, float const & c13, float const & c14, float const & c15,
|
||||
RegTypeSF const& sfa0,
|
||||
RegTypeSF const& sfb0)
|
||||
{
|
||||
static constexpr uint16_t tidA = 0;
|
||||
static constexpr uint16_t bidA = 0;
|
||||
static constexpr uint16_t bidB = 0;
|
||||
static constexpr uint16_t tidB0 = 0;
|
||||
static constexpr uint16_t tidB1 = 1;
|
||||
static constexpr uint16_t tidB2 = 2;
|
||||
static constexpr uint16_t tidB3 = 3;
|
||||
|
||||
#if defined(CUTE_ARCH_MXF4NVF4_4X_UE4M3_MMA_ENABLED)
|
||||
asm volatile(
|
||||
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
|
||||
"{%0, %1, %2, %3},"
|
||||
"{%4, %5, %6, %7},"
|
||||
"{%8, %9},"
|
||||
"{%10, %11, %12, %13},"
|
||||
"{%14},"
|
||||
"{%15, %16},"
|
||||
"{%17},"
|
||||
"{%18, %19};\n"
|
||||
: "=f"(d0), "=f"(d1), "=f"(d8), "=f"(d9)
|
||||
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
|
||||
"r"(b0), "r"(b1),
|
||||
"f"(c0), "f"(c1), "f"(c8), "f"(c9),
|
||||
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
|
||||
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB0));
|
||||
|
||||
asm volatile(
|
||||
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
|
||||
"{%0, %1, %2, %3},"
|
||||
"{%4, %5, %6, %7},"
|
||||
"{%8, %9},"
|
||||
"{%10, %11, %12, %13},"
|
||||
"{%14},"
|
||||
"{%15, %16},"
|
||||
"{%17},"
|
||||
"{%18, %19};\n"
|
||||
: "=f"(d2), "=f"(d3), "=f"(d10), "=f"(d11)
|
||||
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
|
||||
"r"(b2), "r"(b3),
|
||||
"f"(c2), "f"(c3), "f"(c10), "f"(c11),
|
||||
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
|
||||
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB1));
|
||||
|
||||
asm volatile(
|
||||
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
|
||||
"{%0, %1, %2, %3},"
|
||||
"{%4, %5, %6, %7},"
|
||||
"{%8, %9},"
|
||||
"{%10, %11, %12, %13},"
|
||||
"{%14},"
|
||||
"{%15, %16},"
|
||||
"{%17},"
|
||||
"{%18, %19};\n"
|
||||
: "=f"(d4), "=f"(d5), "=f"(d12), "=f"(d13)
|
||||
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
|
||||
"r"(b4), "r"(b5),
|
||||
"f"(c4), "f"(c5), "f"(c12), "f"(c13),
|
||||
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
|
||||
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB2));
|
||||
|
||||
asm volatile(
|
||||
"mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::4X.m16n8k64.row.col.f32.e2m1.e2m1.f32.ue4m3 "
|
||||
"{%0, %1, %2, %3},"
|
||||
"{%4, %5, %6, %7},"
|
||||
"{%8, %9},"
|
||||
"{%10, %11, %12, %13},"
|
||||
"{%14},"
|
||||
"{%15, %16},"
|
||||
"{%17},"
|
||||
"{%18, %19};\n"
|
||||
: "=f"(d6), "=f"(d7), "=f"(d14), "=f"(d15)
|
||||
: "r"(a0), "r"(a1), "r"(a2), "r"(a3),
|
||||
"r"(b6), "r"(b7),
|
||||
"f"(c6), "f"(c7), "f"(c14), "f"(c15),
|
||||
"r"(uint32_t(sfa0)) , "h"(bidA), "h"(tidA),
|
||||
"r"(uint32_t(sfb0)) , "h"(bidB), "h"(tidB3));
|
||||
#else
|
||||
CUTE_INVALID_CONTROL_PATH("Attempting to use SM120::BLOCKSCALED::SM120_16x8x64_TN_VS without CUTE_ARCH_MXF4NVF4_4X_UE4M3_MMA_ENABLED");
|
||||
#endif
|
||||
|
||||
}
|
||||
};
|
||||
|
||||
} // namespace cute::SM120::BLOCKSCALED
|
||||
|
||||
namespace cute {
|
||||
|
||||
// MMA NVFP4 16x32x64 TN
|
||||
template <>
|
||||
struct MMA_Traits<SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4>
|
||||
{
|
||||
// The MMA accepts 4-bit inputs regardless of the types for A and B
|
||||
using ValTypeA = uint4_t;
|
||||
using ValTypeB = uint4_t;
|
||||
|
||||
using ValTypeD = float;
|
||||
using ValTypeC = float;
|
||||
|
||||
using ValTypeSF = cutlass::float_ue4m3_t;
|
||||
constexpr static int SFVecSize = 16;
|
||||
|
||||
using Shape_MNK = Shape<_16,_32,_64>;
|
||||
using ThrID = Layout<_32>;
|
||||
|
||||
// (T32,V32) -> (M16,K64)
|
||||
using ALayout = Layout<Shape <Shape < _4,_8>,Shape < _8,_2, _2>>,
|
||||
Stride<Stride<_128,_1>,Stride<_16,_8,_512>>>;
|
||||
// (T32,V64) -> (N32,K64)
|
||||
using BLayout = Layout<Shape <Shape < _4,_8>,Shape <_8, _2, _4>>,
|
||||
Stride<Stride<_256,_1>,Stride<_32,_1024, _8>>>;
|
||||
// (T32,V64) -> (M16,K64)
|
||||
using SFALayout = Layout<Shape <Shape <_2,_2,_8>,_64>,
|
||||
Stride<Stride<_8,_0,_1>,_16>>;
|
||||
// (T32,V64) -> (N32,K64)
|
||||
using SFBLayout = Layout<Shape <Shape <_4,_8>,_64>,
|
||||
Stride<Stride<_8,_1>, _32>>;
|
||||
// (T32,V16) -> (M16,N32)
|
||||
using CLayout = Layout<Shape <Shape < _4,_8>,Shape < Shape<_2, _4>,_2>>,
|
||||
Stride<Stride<_32,_1>,Stride<Stride<_16, _128>,_8>>>;
|
||||
};
|
||||
|
||||
|
||||
template <class SFATensor, class Atom, class TiledThr, class TiledPerm>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
thrfrg_SFA(SFATensor&& sfatensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
|
||||
{
|
||||
CUTE_STATIC_ASSERT_V(rank(sfatensor) >= Int<2>{});
|
||||
|
||||
using AtomShape_MNK = typename Atom::Shape_MNK;
|
||||
using AtomLayoutSFA_TV = typename Atom::Traits::SFALayout;
|
||||
|
||||
auto permutation_mnk = TiledPerm{};
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// Reorder the tensor for the TiledAtom
|
||||
auto t_tile = make_tile(get<0>(permutation_mnk),
|
||||
get<2>(permutation_mnk));
|
||||
auto t_tensor = logical_divide(sfatensor, t_tile); // (PermM,PermK)
|
||||
|
||||
// Tile the tensor for the Atom
|
||||
auto a_tile = make_tile(make_layout(size<0>(AtomShape_MNK{})),
|
||||
make_layout(size<2>(AtomShape_MNK{})));
|
||||
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomM,AtomK),(RestM,RestK))
|
||||
|
||||
// Transform the Atom mode from (M,K) to (Thr,Val)
|
||||
auto tv_tensor = a_tensor.compose(AtomLayoutSFA_TV{},_); // ((ThrV,FrgV),(RestM,RestK))
|
||||
|
||||
// Tile the tensor for the Thread
|
||||
auto thr_tile = make_tile(_,
|
||||
make_tile(make_layout(size<1>(thr_layout_vmnk)),
|
||||
make_layout(size<3>(thr_layout_vmnk))));
|
||||
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrM,ThrK)),(FrgV,(RestM,RestK)))
|
||||
|
||||
return thr_tensor;
|
||||
}
|
||||
|
||||
template <class SFBTensor, class Atom, class TiledThr, class TiledPerm>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
thrfrg_SFB(SFBTensor&& sfbtensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
|
||||
{
|
||||
CUTE_STATIC_ASSERT_V(rank(sfbtensor) >= Int<2>{});
|
||||
|
||||
using AtomShape_MNK = typename Atom::Shape_MNK;
|
||||
using AtomLayoutSFB_TV = typename Atom::Traits::SFBLayout;
|
||||
|
||||
auto permutation_mnk = TiledPerm{};
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// Reorder the tensor for the TiledAtom
|
||||
auto t_tile = make_tile(get<1>(permutation_mnk),
|
||||
get<2>(permutation_mnk));
|
||||
auto t_tensor = logical_divide(sfbtensor, t_tile); // (PermN,PermK)
|
||||
|
||||
// Tile the tensor for the Atom
|
||||
auto a_tile = make_tile(make_layout(size<1>(AtomShape_MNK{})),
|
||||
make_layout(size<2>(AtomShape_MNK{})));
|
||||
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomN,AtomK),(RestN,RestK))
|
||||
|
||||
// Transform the Atom mode from (M,K) to (Thr,Val)
|
||||
auto tv_tensor = a_tensor.compose(AtomLayoutSFB_TV{},_); // ((ThrV,FrgV),(RestN,RestK))
|
||||
|
||||
// Tile the tensor for the Thread
|
||||
auto thr_tile = make_tile(_,
|
||||
make_tile(make_layout(size<2>(thr_layout_vmnk)),
|
||||
make_layout(size<3>(thr_layout_vmnk))));
|
||||
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrN,ThrK)),(FrgV,(RestN,RestK)))
|
||||
return thr_tensor;
|
||||
}
|
||||
|
||||
template <class SFATensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_SFA(SFATensor&& sfatensor, ThrMma& thread_mma) {
|
||||
auto thr_tensor = make_tensor(static_cast<SFATensor&&>(sfatensor).data(), thrfrg_SFA(sfatensor.layout(),thread_mma));
|
||||
auto thr_vmnk = thread_mma.thr_vmnk_;
|
||||
auto thr_vmk = make_coord(get<0>(thr_vmnk), make_coord(get<1>(thr_vmnk), get<3>(thr_vmnk)));
|
||||
return thr_tensor(thr_vmk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
|
||||
}
|
||||
|
||||
template <class SFATensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_fragment_SFA(SFATensor&& sfatensor, ThrMma& thread_mma) {
|
||||
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
|
||||
return make_fragment_like<ValTypeSF>(partition_SFA(sfatensor, thread_mma));
|
||||
}
|
||||
|
||||
template <class SFBTensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma) {
|
||||
auto thr_tensor = make_tensor(static_cast<SFBTensor&&>(sfbtensor).data(), thrfrg_SFB(sfbtensor.layout(),thread_mma));
|
||||
auto thr_vmnk = thread_mma.thr_vmnk_;
|
||||
auto thr_vnk = make_coord(get<0>(thr_vmnk), make_coord(get<2>(thr_vmnk), get<3>(thr_vmnk)));
|
||||
return thr_tensor(thr_vnk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
|
||||
}
|
||||
|
||||
template <class SFBTensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_fragment_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma) {
|
||||
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
|
||||
return make_fragment_like<ValTypeSF>(partition_SFB(sfbtensor, thread_mma));
|
||||
}
|
||||
|
||||
template<class TiledMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
get_layoutSFA_TV(TiledMma& mma)
|
||||
{
|
||||
// (M,K) -> (M,K)
|
||||
auto tile_shape_mnk = tile_shape(mma);
|
||||
auto ref_A = make_layout(make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk)));
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
|
||||
auto atile = make_tile(_,
|
||||
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
|
||||
make_stride( Int<1>{} , Int<0>{} )),
|
||||
_));
|
||||
|
||||
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
|
||||
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
|
||||
// (thr_idx,val) -> (M,K)
|
||||
return thrfrg_SFA(ref_A, mma).compose(atile, _).compose(thridx_2_thrid, _);
|
||||
}
|
||||
|
||||
template<class TiledMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
get_layoutSFB_TV(TiledMma& mma)
|
||||
{
|
||||
// (N,K) -> (N,K)
|
||||
auto tile_shape_mnk = tile_shape(mma);
|
||||
auto ref_B = make_layout(make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk)));
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
|
||||
auto btile = make_tile(_,
|
||||
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
|
||||
make_stride( Int<0>{} , Int<1>{} )),
|
||||
_));
|
||||
|
||||
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
|
||||
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
|
||||
// (thr_idx,val) -> (M,K)
|
||||
return thrfrg_SFB(ref_B, mma).compose(btile, _).compose(thridx_2_thrid, _);
|
||||
}
|
||||
|
||||
} // namespace cute
|
||||
@@ -0,0 +1,222 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <cutlass/cutlass.h>
|
||||
#include "cute/tensor.hpp"
|
||||
|
||||
#include "cutlass/gemm/collective/collective_builder.hpp"
|
||||
#include "named_barrier.h"
|
||||
#include "utils.h"
|
||||
|
||||
namespace flash {
|
||||
|
||||
using namespace cute;
|
||||
|
||||
template <typename Ktraits>
|
||||
struct CollectiveEpilogueFwd{
|
||||
|
||||
using Element = typename Ktraits::ElementOut;
|
||||
static constexpr int kBlockM = Ktraits::kBlockM;
|
||||
static constexpr int kBlockN = Ktraits::kBlockN;
|
||||
static constexpr int kHeadDim = Ktraits::kHeadDim;
|
||||
using TileShape_MNK = Shape<Int<kBlockM>, Int<kBlockN>, Int<kHeadDim>>;
|
||||
static constexpr int kNWarps = Ktraits::kNWarps;
|
||||
static constexpr int kNThreads = kNWarps * cutlass::NumThreadsPerWarp;
|
||||
static constexpr int NumMmaThreads = kNThreads - cutlass::NumThreadsPerWarpGroup;
|
||||
|
||||
using GmemTiledCopyOTMA = cute::SM90_TMA_STORE;
|
||||
|
||||
// These are for storing the output tensor without TMA (e.g., for setting output to zero)
|
||||
static constexpr int kGmemElemsPerLoad = sizeof(cute::uint128_t) / sizeof(Element);
|
||||
static_assert(kHeadDim % kGmemElemsPerLoad == 0, "kHeadDim must be a multiple of kGmemElemsPerLoad");
|
||||
static constexpr int kGmemThreadsPerRow = kHeadDim / kGmemElemsPerLoad;
|
||||
static_assert(NumMmaThreads % kGmemThreadsPerRow == 0, "NumMmaThreads must be a multiple of kGmemThreadsPerRow");
|
||||
using GmemLayoutAtom = Layout<Shape <Int<NumMmaThreads / kGmemThreadsPerRow>, Int<kGmemThreadsPerRow>>,
|
||||
Stride<Int<kGmemThreadsPerRow>, _1>>;
|
||||
using GmemTiledCopyO = decltype(
|
||||
make_tiled_copy(Copy_Atom<DefaultCopy, Element>{},
|
||||
GmemLayoutAtom{},
|
||||
Layout<Shape<_1, Int<kGmemElemsPerLoad>>>{})); // Val layout, 8 or 16 vals per store
|
||||
|
||||
using SmemLayoutO = typename Ktraits::SmemLayoutO;
|
||||
|
||||
using SmemCopyAtomO = Copy_Atom<SM90_U32x2_STSM_N, Element>;
|
||||
using SharedStorage = cute::array_aligned<Element, cute::cosize_v<SmemLayoutO>>;
|
||||
|
||||
using ShapeO = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen_q, d, head, batch)
|
||||
using StrideO = cute::Stride<int64_t, _1, int64_t, int64_t>;
|
||||
using StrideLSE = cute::Stride<_1, int64_t, int64_t>; // (seqlen_q, head, batch)
|
||||
|
||||
using TMA_O = decltype(make_tma_copy(
|
||||
GmemTiledCopyOTMA{},
|
||||
make_tensor(make_gmem_ptr(static_cast<Element*>(nullptr)), repeat_like(StrideO{}, int32_t(0)), StrideO{}),
|
||||
SmemLayoutO{},
|
||||
select<0, 2>(TileShape_MNK{}),
|
||||
_1{})); // no mcast for O
|
||||
|
||||
// Host side kernel arguments
|
||||
struct Arguments {
|
||||
Element* ptr_O;
|
||||
ShapeO const shape_O;
|
||||
StrideO const stride_O;
|
||||
float* ptr_LSE;
|
||||
StrideLSE const stride_LSE;
|
||||
};
|
||||
|
||||
// Device side kernel params
|
||||
struct Params {
|
||||
Element* ptr_O;
|
||||
ShapeO const shape_O;
|
||||
StrideO const stride_O;
|
||||
float* ptr_LSE;
|
||||
StrideLSE const stride_LSE;
|
||||
TMA_O tma_store_O;
|
||||
};
|
||||
|
||||
static Params
|
||||
to_underlying_arguments(Arguments const& args) {
|
||||
Tensor mO = make_tensor(make_gmem_ptr(args.ptr_O), args.shape_O, args.stride_O);
|
||||
TMA_O tma_store_O = make_tma_copy(
|
||||
GmemTiledCopyOTMA{},
|
||||
mO,
|
||||
SmemLayoutO{},
|
||||
select<0, 2>(TileShape_MNK{}),
|
||||
_1{}); // no mcast for O
|
||||
return {args.ptr_O, args.shape_O, args.stride_O, args.ptr_LSE, args.stride_LSE, tma_store_O};
|
||||
}
|
||||
|
||||
/// Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
|
||||
CUTLASS_DEVICE
|
||||
static void prefetch_tma_descriptors(Params const& epilogue_params) {
|
||||
cute::prefetch_tma_descriptor(epilogue_params.tma_store_O.get_tma_descriptor());
|
||||
}
|
||||
|
||||
template <typename SharedStorage, typename FrgTensorO, typename TiledMma>
|
||||
CUTLASS_DEVICE void
|
||||
mma_store(
|
||||
SharedStorage& shared_storage,
|
||||
TiledMma tiled_mma,
|
||||
FrgTensorO const& tOrO,
|
||||
int thread_idx
|
||||
){
|
||||
Tensor sO = cute::as_position_independent_swizzle_tensor(make_tensor(make_smem_ptr(shared_storage.smem_o.begin()), SmemLayoutO{}));
|
||||
auto smem_tiled_copy_O = make_tiled_copy_C(SmemCopyAtomO{}, tiled_mma);
|
||||
auto smem_thr_copy_O = smem_tiled_copy_O.get_thread_slice(thread_idx);
|
||||
constexpr int numel = decltype(size(tOrO))::value;
|
||||
cutlass::NumericArrayConverter<Element, float, numel> convert_op;
|
||||
// HACK: this requires tensor to be "contiguous"
|
||||
auto frag = convert_op(*reinterpret_cast<const cutlass::Array<float, numel> *>(tOrO.data()));
|
||||
auto tOrO_out = make_tensor(make_rmem_ptr<Element>(&frag), tOrO.layout());
|
||||
Tensor taccOrO = smem_thr_copy_O.retile_S(tOrO_out); // ((Atom,AtomNum), MMA_M, MMA_N)
|
||||
Tensor taccOsO = smem_thr_copy_O.partition_D(sO); // ((Atom,AtomNum),PIPE_M,PIPE_N)
|
||||
cute::copy(smem_tiled_copy_O, taccOrO, taccOsO);
|
||||
cutlass::arch::fence_view_async_shared(); // ensure smem writes are visible to TMA
|
||||
}
|
||||
|
||||
template<typename SharedStorage, typename Params, typename WorkTileInfo, typename SchedulerParams>
|
||||
CUTLASS_DEVICE void
|
||||
tma_store(
|
||||
SharedStorage& shared_storage,
|
||||
Params const& epilogue_params,
|
||||
WorkTileInfo work_tile_info,
|
||||
SchedulerParams const& scheduler_params,
|
||||
int thread_idx
|
||||
) {
|
||||
auto [m_block, bidh, bidb] = work_tile_info.get_block_coord(scheduler_params);
|
||||
Tensor sO = cute::as_position_independent_swizzle_tensor(make_tensor(make_smem_ptr(shared_storage.smem_o.begin()), SmemLayoutO{}));
|
||||
Tensor mO = epilogue_params.tma_store_O.get_tma_tensor(epilogue_params.shape_O);
|
||||
Tensor gO = local_tile(mO(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
|
||||
auto block_tma_O = epilogue_params.tma_store_O.get_slice(_0{});
|
||||
Tensor tOgO = block_tma_O.partition_D(gO); // (TMA, TMA_M, TMA_K)
|
||||
Tensor tOsO = block_tma_O.partition_S(sO); // (TMA, TMA_M, TMA_K)
|
||||
|
||||
// auto shape_LSE = select<0, 2, 3>(epilogue_params.shape_O);
|
||||
// Tensor mLSE = make_tensor(make_gmem_ptr(epilogue_params.ptr_LSE), shape_LSE, epilogue_params.stride_LSE);
|
||||
// Tensor gLSE = local_tile(mLSE(_, bidh, bidb), Shape<Int<kBlockM>>{}, make_coord(m_block));
|
||||
|
||||
// Tensor caccO = cute::make_identity_tensor(select<0, 2>(TileShape_MNK{}));
|
||||
// auto thread_mma = tiled_mma.get_thread_slice(thread_idx);
|
||||
// Tensor taccOcO = thread_mma.partition_C(caccO); // (MMA,MMA_M,MMA_K)
|
||||
// static_assert(decltype(size<0, 0>(taccOcO))::value == 2);
|
||||
// static_assert(decltype(size<0, 1>(taccOcO))::value == 2);
|
||||
// // // // taccOcO has shape ((2, 2, V), MMA_M, MMA_K), we only take only the row indices.
|
||||
// Tensor taccOcO_row = taccOcO(make_coord(_0{}, _), _, _0{});
|
||||
// CUTE_STATIC_ASSERT_V(size(lse) == size(taccOcO_row)); // MMA_M
|
||||
// if (get<1>(taccOcO_row(_0{})) == 0) {
|
||||
// #pragma unroll
|
||||
// for (int mi = 0; mi < size(lse); ++mi) {
|
||||
// const int row = get<0>(taccOcO_row(mi));
|
||||
// if (row < get<0>(shape_LSE) - m_block * kBlockM) { gLSE(row) = lse(mi); }
|
||||
// }
|
||||
// }
|
||||
|
||||
// if (cutlass::canonical_warp_idx_sync() == kNWarps - 1) {
|
||||
// cutlass::arch::NamedBarrier::sync(NumMmaThreads + cutlass::NumThreadsPerWarp,
|
||||
// static_cast<uint32_t>(FP4NamedBarriers::EpilogueBarrier));
|
||||
// int const lane_predicate = cute::elect_one_sync();
|
||||
// if (lane_predicate) {
|
||||
// cute::copy(epilogue_params.tma_store_O, tOsO, tOgO);
|
||||
// tma_store_arrive();
|
||||
// }
|
||||
// }
|
||||
cute::copy(epilogue_params.tma_store_O, tOsO, tOgO);
|
||||
tma_store_arrive();
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
store_tail() {
|
||||
tma_store_wait<0>();
|
||||
}
|
||||
|
||||
// Write 0 to output and -inf to LSE
|
||||
CUTLASS_DEVICE void
|
||||
store_zero(
|
||||
Params const& epilogue_params,
|
||||
int thread_idx,
|
||||
cute::tuple<int32_t, int32_t, int32_t> const& block_coord
|
||||
) {
|
||||
auto [m_block, bidh, bidb] = block_coord;
|
||||
Tensor mO = make_tensor(make_gmem_ptr(epilogue_params.ptr_O), epilogue_params.shape_O, epilogue_params.stride_O);
|
||||
Tensor gO = local_tile(mO(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
|
||||
auto shape_LSE = select<0, 2, 3>(epilogue_params.shape_O);
|
||||
Tensor mLSE = make_tensor(make_gmem_ptr(epilogue_params.ptr_LSE), shape_LSE, epilogue_params.stride_LSE);
|
||||
Tensor gLSE = local_tile(mLSE(_, bidh, bidb), Shape<Int<kBlockM>>{}, make_coord(m_block));
|
||||
|
||||
GmemTiledCopyO gmem_tiled_copy_O;
|
||||
auto gmem_thr_copy_O = gmem_tiled_copy_O.get_thread_slice(thread_idx);
|
||||
Tensor tOgO = gmem_thr_copy_O.partition_D(gO);
|
||||
Tensor tOrO = make_fragment_like(tOgO);
|
||||
clear(tOrO);
|
||||
// Construct identity layout for sO
|
||||
Tensor cO = cute::make_identity_tensor(select<0, 2>(TileShape_MNK{})); // (BLK_M,BLK_K) -> (blk_m,blk_k)
|
||||
// Repeat the partitioning with identity layouts
|
||||
Tensor tOcO = gmem_thr_copy_O.partition_D(cO);
|
||||
Tensor tOpO = make_tensor<bool>(make_shape(size<2>(tOgO)));
|
||||
#pragma unroll
|
||||
for (int k = 0; k < size(tOpO); ++k) { tOpO(k) = get<1>(tOcO(_0{}, _0{}, k)) < get<1>(epilogue_params.shape_O); }
|
||||
// Clear_OOB_K must be false since we don't want to write zeros to gmem
|
||||
flash::copy</*Is_even_MN=*/false, /*Is_even_K=*/false, /*Clear_OOB_MN=*/false, /*Clear_OOB_K=*/false>(
|
||||
gmem_tiled_copy_O, tOrO, tOgO, tOcO, tOpO, get<0>(epilogue_params.shape_O) - m_block * kBlockM
|
||||
);
|
||||
static_assert(kBlockM <= NumMmaThreads);
|
||||
if (thread_idx < get<0>(shape_LSE) - m_block * kBlockM) { gLSE(thread_idx) = INFINITY; }
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
} // namespace flash
|
||||
@@ -0,0 +1,202 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "cute/algorithm/copy.hpp"
|
||||
#include "cute/atom/mma_atom.hpp"
|
||||
#include "cutlass/gemm/collective/collective_builder.hpp"
|
||||
#include "cute/tensor.hpp"
|
||||
#include "cutlass/cutlass.h"
|
||||
#include "cutlass/layout/layout.h"
|
||||
#include "cutlass/numeric_types.h"
|
||||
#include "cutlass/pipeline/pipeline.hpp"
|
||||
|
||||
#include "blockscaled_layout.h"
|
||||
#include "cute_extension.h"
|
||||
#include "named_barrier.h"
|
||||
using namespace cute;
|
||||
|
||||
template <
|
||||
int kStages,
|
||||
int EpiStages,
|
||||
typename Element,
|
||||
typename ElementSF,
|
||||
typename OutputType,
|
||||
typename SmemLayoutQ,
|
||||
typename SmemLayoutK,
|
||||
typename SmemLayoutV,
|
||||
typename SmemLayoutDS,
|
||||
typename SmemLayoutO,
|
||||
typename SmemLayoutSFQ,
|
||||
typename SmemLayoutSFK,
|
||||
typename SmemLayoutSFV
|
||||
>
|
||||
struct SharedStorageQKVOwithSF : cute::aligned_struct<128, _0>{
|
||||
|
||||
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutQ>> smem_q;
|
||||
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutK>> smem_k;
|
||||
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFQ>> smem_SFQ;
|
||||
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFK>> smem_SFK;
|
||||
cute::ArrayEngine<ElementSF, cute::cosize_v<SmemLayoutSFV>> smem_SFV;
|
||||
alignas(1024) cute::ArrayEngine<float, cute::cosize_v<SmemLayoutDS>> smem_ds;
|
||||
alignas(1024) cute::ArrayEngine<Element, cute::cosize_v<SmemLayoutV>> smem_v;
|
||||
alignas(1024) cute::ArrayEngine<OutputType, cute::cosize_v<SmemLayoutO>> smem_o;
|
||||
|
||||
struct {
|
||||
alignas(16) typename cutlass::PipelineTmaAsync<1>::SharedStorage pipeline_q;
|
||||
alignas(16) typename cutlass::PipelineTmaAsync<kStages>::SharedStorage pipeline_k;
|
||||
alignas(16) typename cutlass::PipelineTmaAsync<kStages>::SharedStorage pipeline_v;
|
||||
alignas(16) typename flash::OrderedSequenceBarrierVarGroupSize<EpiStages, 2>::SharedStorage barrier_o;
|
||||
int tile_count_semaphore;
|
||||
};
|
||||
};
|
||||
|
||||
template <
|
||||
int kHeadDim_,
|
||||
int kBlockM_,
|
||||
int kBlockN_,
|
||||
int kStages_,
|
||||
int kClusterM_,
|
||||
bool BlockMean_,
|
||||
typename ElementPairType_ = cutlass::nv_float4_t<cutlass::float_e2m1_t>,
|
||||
typename ElementOut_ = cutlass::bfloat16_t
|
||||
>
|
||||
struct Flash_fwd_kernel_traits {
|
||||
static constexpr int kBlockM = kBlockM_;
|
||||
static constexpr int kBlockN = kBlockN_;
|
||||
static constexpr int kHeadDim = kHeadDim_;
|
||||
static constexpr bool BlockMean = BlockMean_;
|
||||
static constexpr bool SmoothQ = true;
|
||||
static_assert(kHeadDim % 32 == 0);
|
||||
static_assert(kBlockM == 64 || kBlockM == 128);
|
||||
static constexpr int kNWarps = kBlockM == 128 ? 12 : 8;
|
||||
static constexpr int kNThreads = kNWarps * cutlass::NumThreadsPerWarp;
|
||||
static constexpr int kClusterM = kClusterM_;
|
||||
static constexpr int kStages = kStages_;
|
||||
static constexpr int EpiStages = 1;
|
||||
static constexpr int NumSFQK = kHeadDim / 16;
|
||||
static constexpr int NumSFPV = kBlockN / 16;
|
||||
using ElementSF = cutlass::float_ue4m3_t;
|
||||
using Element = cutlass::float_e2m1_t;
|
||||
using ElementAccum = float;
|
||||
using ElementOut = ElementOut_;
|
||||
using index_t = int64_t;
|
||||
static constexpr auto SFVectorSize = 16;
|
||||
using TileShape_MNK = Shape<Int<kBlockM>, Int<kBlockN>, Int<kHeadDim>>;
|
||||
using ClusterShape_MNK = Shape<_1, _1, _1>;
|
||||
using PermTileM = decltype(cute::min(size<0>(TileShape_MNK{}), _128{}));
|
||||
using PermTileN = _32;
|
||||
using PermTileK = Int<kHeadDim>;
|
||||
|
||||
using ElementQMma = decltype(cutlass::gemm::collective::detail::sm1xx_kernel_input_element_to_mma_input_element<Element>());
|
||||
using ElementKMma = decltype(cutlass::gemm::collective::detail::sm1xx_kernel_input_element_to_mma_input_element<Element>());
|
||||
|
||||
using AtomLayoutMNK = std::conditional_t<kBlockM == 128,
|
||||
Layout<Shape<_8, _1, _1>>,
|
||||
Layout<Shape<_4, _1, _1>>
|
||||
>;
|
||||
using TiledMmaQK = decltype(cute::make_tiled_mma(
|
||||
cute::SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4{},
|
||||
AtomLayoutMNK{},
|
||||
Tile<PermTileM, PermTileN, PermTileK>{}
|
||||
));
|
||||
|
||||
using TiledMmaPV = decltype(cute::make_tiled_mma(
|
||||
cute::SM120::BLOCKSCALED::SM120_16x32x64_TN_VS_NVFP4{},
|
||||
AtomLayoutMNK{},
|
||||
Tile<PermTileM, _32, PermTileK>{}
|
||||
));
|
||||
|
||||
static constexpr int MMA_NSF = size<2>(typename TiledMmaQK::AtomShape_MNK{}) / SFVectorSize;
|
||||
|
||||
using GmemTiledCopy = SM90_TMA_LOAD;
|
||||
using GmemTiledCopySF = SM90_TMA_LOAD;
|
||||
|
||||
using SmemLayoutAtomQ = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
|
||||
using SmemLayoutAtomK = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
|
||||
using SmemLayoutAtomV = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<2>(TileShape_MNK{}))>());
|
||||
using SmemLayoutAtomVt = decltype(cutlass::gemm::collective::detail::sm120_rr_smem_selector<Element, decltype(size<1>(TileShape_MNK{}))>());
|
||||
using SmemLayoutQ = decltype(tile_to_shape(SmemLayoutAtomQ{}, select<0, 2>(TileShape_MNK{})));
|
||||
using SmemLayoutK =
|
||||
decltype(tile_to_shape(SmemLayoutAtomK{},
|
||||
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{}), Int<kStages>{})));
|
||||
using SmemLayoutV =
|
||||
decltype(tile_to_shape(SmemLayoutAtomV{},
|
||||
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{}), Int<kStages>{})));
|
||||
using SmemLayoutVt =
|
||||
decltype(tile_to_shape(SmemLayoutAtomVt{},
|
||||
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{}), Int<kStages>{})));
|
||||
using SmemLayoutAtomDS = Layout<Shape<Int<kBlockM>, Int<kBlockN>>, Stride<_0, _1>>;
|
||||
using SmemLayoutDS =
|
||||
decltype(tile_to_shape(SmemLayoutAtomDS{},
|
||||
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{}), Int<kStages>{})));
|
||||
|
||||
using SmemCopyAtomQ = Copy_Atom<SM75_U32x4_LDSM_N, Element>;
|
||||
using SmemCopyAtomKV = Copy_Atom<SM75_U32x4_LDSM_N, Element>;
|
||||
using SmemCopyAtomSF = Copy_Atom<UniversalCopy<ElementSF>, ElementSF>;
|
||||
using SmemCopyAtomDS = Copy_Atom<UniversalCopy<float>, float>;
|
||||
|
||||
using BlkScaledConfig = flash::BlockScaledConfig<SFVectorSize>;
|
||||
using LayoutSF = typename BlkScaledConfig::LayoutSF;
|
||||
using SfAtom = typename BlkScaledConfig::SfAtom;
|
||||
using SmemLayoutAtomSFQ = decltype(BlkScaledConfig::deduce_smem_layoutSFQ(TiledMmaQK{}, TileShape_MNK{}));
|
||||
using SmemLayoutAtomSFK = decltype(BlkScaledConfig::deduce_smem_layoutSFKV(TiledMmaQK{}, TileShape_MNK{}));
|
||||
using SmemLayoutAtomSFV = decltype(BlkScaledConfig::deduce_smem_layoutSFKV(TiledMmaPV{}, TileShape_MNK{}));
|
||||
using SmemLayoutAtomSFVt = decltype(BlkScaledConfig::deduce_smem_layoutSFVt(TiledMmaPV{}, Shape<Int<kBlockM>, Int<kHeadDim>, Int<kBlockN>>{}));
|
||||
using LayoutSFP = decltype(
|
||||
make_layout(
|
||||
make_shape(make_shape(_16{}, _4{}), _1{}, Int<kBlockN / 64>{}),
|
||||
make_stride(make_stride(_0{}, _1{}), _0{}, _4{})
|
||||
)
|
||||
);
|
||||
using LayoutP = decltype(
|
||||
make_layout(
|
||||
make_shape(make_shape(_8{}, _2{}, _2{}), _1{}, Int<kBlockN / 64>{}),
|
||||
make_stride(make_stride(_1{}, _8{}, _16{}), _0{}, _32{})
|
||||
)
|
||||
);
|
||||
using SmemLayoutSFQ = decltype(make_layout(
|
||||
shape(SmemLayoutAtomSFQ{}),
|
||||
stride(SmemLayoutAtomSFQ{})
|
||||
));
|
||||
using SmemLayoutSFK = decltype(make_layout(
|
||||
append(shape(SmemLayoutAtomSFK{}), Int<kStages>{}),
|
||||
append(stride(SmemLayoutAtomSFK{}), size(filter_zeros(SmemLayoutAtomSFK{})))
|
||||
));
|
||||
using SmemLayoutSFV = decltype(make_layout(
|
||||
append(shape(SmemLayoutAtomSFV{}), Int<kStages>{}),
|
||||
append(stride(SmemLayoutAtomSFV{}), size(filter_zeros(SmemLayoutAtomSFV{})))
|
||||
));
|
||||
using SmemLayoutSFVt = decltype(make_layout(
|
||||
append(shape(SmemLayoutAtomSFVt{}), Int<kStages>{}),
|
||||
append(stride(SmemLayoutAtomSFVt{}), size(filter_zeros(SmemLayoutAtomSFVt{})))
|
||||
));
|
||||
|
||||
using SmemLayoutAtomO = decltype(cutlass::gemm::collective::detail::ss_smem_selector<GMMA::Major::K, ElementOut,
|
||||
decltype(cute::get<0>(TileShape_MNK{})), decltype(cute::get<2>(TileShape_MNK{}))>());
|
||||
using SmemLayoutO = decltype(tile_to_shape(SmemLayoutAtomO{}, select<0, 2>(TileShape_MNK{}), Step<_1, _2>{}));
|
||||
using SharedStorage = SharedStorageQKVOwithSF<kStages, EpiStages, Element, ElementSF, ElementOut,
|
||||
SmemLayoutQ, SmemLayoutK, SmemLayoutV, SmemLayoutDS,
|
||||
SmemLayoutO, SmemLayoutSFQ, SmemLayoutSFK, SmemLayoutSFVt>;
|
||||
using MainloopPipeline = typename cutlass::PipelineTmaAsync<kStages>;
|
||||
using PipelineState = typename cutlass::PipelineState<kStages>;
|
||||
using MainloopPipelineQ = cutlass::PipelineTmaAsync<1>;
|
||||
using PipelineParamsQ = typename MainloopPipelineQ::Params;
|
||||
using PipelineStateQ = typename cutlass::PipelineState<1>;
|
||||
using EpilogueBarrier = typename flash::OrderedSequenceBarrierVarGroupSize<EpiStages, 2>;
|
||||
};
|
||||
|
||||
@@ -0,0 +1,204 @@
|
||||
// Modified from the original SageAttention3 code
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "cute/tensor.hpp"
|
||||
|
||||
#include <cutlass/cutlass.h>
|
||||
#include <cutlass/arch/reg_reconfig.h>
|
||||
#include <cutlass/array.h>
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/numeric_conversion.h>
|
||||
#include "cutlass/pipeline/pipeline.hpp"
|
||||
|
||||
#include "params.h"
|
||||
#include "utils.h"
|
||||
#include "tile_scheduler.h"
|
||||
#include "mainloop_tma_ws.h"
|
||||
#include "epilogue_tma_ws.h"
|
||||
#include "named_barrier.h"
|
||||
#include "softmax_fused.h"
|
||||
|
||||
namespace flash {
|
||||
|
||||
using namespace cute;
|
||||
|
||||
template <typename Ktraits, bool Is_causal, typename TileScheduler>
|
||||
__global__ void __launch_bounds__(Ktraits::kNWarps * cutlass::NumThreadsPerWarp, 1)
|
||||
compute_attn_ws(CUTE_GRID_CONSTANT Flash_fwd_params const params,
|
||||
CUTE_GRID_CONSTANT typename CollectiveMainloopFwd<Ktraits, Is_causal>::Params const mainloop_params,
|
||||
CUTE_GRID_CONSTANT typename CollectiveEpilogueFwd<Ktraits>::Params const epilogue_params,
|
||||
CUTE_GRID_CONSTANT typename TileScheduler::Params const scheduler_params
|
||||
) {
|
||||
|
||||
using Element = typename Ktraits::Element;
|
||||
using ElementAccum = typename Ktraits::ElementAccum;
|
||||
using SoftType = ElementAccum;
|
||||
using TileShape_MNK = typename Ktraits::TileShape_MNK;
|
||||
using ClusterShape = typename Ktraits::ClusterShape_MNK;
|
||||
|
||||
static constexpr int NumMmaThreads = size(typename Ktraits::TiledMmaQK{});
|
||||
static constexpr int NumCopyThreads = cutlass::NumThreadsPerWarpGroup;
|
||||
static constexpr int kBlockM = Ktraits::kBlockM;
|
||||
|
||||
using CollectiveMainloop = CollectiveMainloopFwd<Ktraits, Is_causal>;
|
||||
using CollectiveEpilogue = CollectiveEpilogueFwd<Ktraits>;
|
||||
|
||||
using MainloopPipeline = typename Ktraits::MainloopPipeline;
|
||||
using PipelineParams = typename MainloopPipeline::Params;
|
||||
using PipelineState = typename MainloopPipeline::PipelineState;
|
||||
using MainloopPipelineQ = typename Ktraits::MainloopPipelineQ;
|
||||
using PipelineParamsQ = typename Ktraits::PipelineParamsQ;
|
||||
using PipelineStateQ = typename Ktraits::PipelineStateQ;
|
||||
using EpilogueBarrier = typename Ktraits::EpilogueBarrier;
|
||||
|
||||
|
||||
enum class WarpGroupRole {
|
||||
Producer = 0,
|
||||
Consumer0 = 1,
|
||||
Consumer1 = 2
|
||||
};
|
||||
enum class ProducerWarpRole {
|
||||
Mainloop = 0,
|
||||
Epilogue = 1,
|
||||
Warp2 = 2,
|
||||
Warp3 = 3
|
||||
};
|
||||
|
||||
extern __shared__ char shared_memory[];
|
||||
auto &shared_storage = *reinterpret_cast<typename Ktraits::SharedStorage*>(shared_memory);
|
||||
|
||||
int const lane_predicate = cute::elect_one_sync();
|
||||
int const warp_idx = cutlass::canonical_warp_idx_sync();
|
||||
int warp_group_idx = cutlass::canonical_warp_group_idx();
|
||||
int const warp_group_thread_idx = threadIdx.x % cutlass::NumThreadsPerWarpGroup;
|
||||
int warp_idx_in_warp_group = warp_idx % cutlass::NumWarpsPerWarpGroup;
|
||||
auto warp_group_role = WarpGroupRole(warp_group_idx);
|
||||
auto producer_warp_role = ProducerWarpRole(warp_idx_in_warp_group);
|
||||
|
||||
// Issue Tma Descriptor Prefetch from a single thread
|
||||
if (warp_idx == 0 && lane_predicate) {
|
||||
CollectiveMainloop::prefetch_tma_descriptors(mainloop_params);
|
||||
CollectiveEpilogue::prefetch_tma_descriptors(epilogue_params);
|
||||
}
|
||||
|
||||
// Obtain warp index
|
||||
|
||||
PipelineParams pipeline_params_v;
|
||||
pipeline_params_v.transaction_bytes = CollectiveMainloop::TmaTransactionBytesV;
|
||||
pipeline_params_v.role = warp_group_role == WarpGroupRole::Producer
|
||||
? MainloopPipeline::ThreadCategory::Producer
|
||||
: MainloopPipeline::ThreadCategory::Consumer;
|
||||
pipeline_params_v.is_leader = warp_group_thread_idx == 0;
|
||||
pipeline_params_v.num_consumers = NumMmaThreads;
|
||||
|
||||
PipelineParams pipeline_params_k;
|
||||
pipeline_params_k.transaction_bytes = CollectiveMainloop::TmaTransactionBytesK;
|
||||
pipeline_params_k.role = warp_group_role == WarpGroupRole::Producer
|
||||
? MainloopPipeline::ThreadCategory::Producer
|
||||
: MainloopPipeline::ThreadCategory::Consumer;
|
||||
pipeline_params_k.is_leader = warp_group_thread_idx == 0;
|
||||
pipeline_params_k.num_consumers = NumMmaThreads;
|
||||
|
||||
PipelineParamsQ pipeline_params_q;
|
||||
pipeline_params_q.transaction_bytes = CollectiveMainloop::TmaTransactionBytesQ;
|
||||
pipeline_params_q.role = warp_group_role == WarpGroupRole::Producer
|
||||
? MainloopPipelineQ::ThreadCategory::Producer
|
||||
: MainloopPipelineQ::ThreadCategory::Consumer;
|
||||
pipeline_params_q.is_leader = warp_group_thread_idx == 0;
|
||||
pipeline_params_q.num_consumers = NumMmaThreads;
|
||||
|
||||
// We're counting on pipeline_k to call cutlass::arch::fence_barrier_init();
|
||||
MainloopPipelineQ pipeline_q(shared_storage.pipeline_q, pipeline_params_q, ClusterShape{});
|
||||
MainloopPipeline pipeline_k(shared_storage.pipeline_k, pipeline_params_k, ClusterShape{});
|
||||
MainloopPipeline pipeline_v(shared_storage.pipeline_v, pipeline_params_v, ClusterShape{});
|
||||
|
||||
uint32_t epilogue_barrier_group_size_list[2] = {cutlass::NumThreadsPerWarp, NumMmaThreads};
|
||||
typename EpilogueBarrier::Params params_epilogue_barrier;
|
||||
params_epilogue_barrier.group_id = (warp_group_role == WarpGroupRole::Producer);
|
||||
params_epilogue_barrier.group_size_list = epilogue_barrier_group_size_list;
|
||||
EpilogueBarrier barrier_o(shared_storage.barrier_o, params_epilogue_barrier);
|
||||
|
||||
CollectiveMainloop collective_mainloop;
|
||||
CollectiveEpilogue collective_epilogue;
|
||||
__syncthreads();
|
||||
|
||||
if (warp_group_role == WarpGroupRole::Producer) {
|
||||
cutlass::arch::warpgroup_reg_dealloc<24>();
|
||||
TileScheduler scheduler;
|
||||
|
||||
if (producer_warp_role == ProducerWarpRole::Mainloop) { // Load Q, K, V
|
||||
PipelineStateQ smem_pipe_write_q = cutlass::make_producer_start_state<MainloopPipelineQ>();
|
||||
PipelineState smem_pipe_write_k = cutlass::make_producer_start_state<MainloopPipeline>();
|
||||
PipelineState smem_pipe_write_v = cutlass::make_producer_start_state<MainloopPipeline>();
|
||||
|
||||
int work_idx = 0;
|
||||
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
|
||||
int tile_count_semaphore = 0;
|
||||
collective_mainloop.load(mainloop_params, scheduler_params,
|
||||
pipeline_q, pipeline_k, pipeline_v,
|
||||
smem_pipe_write_q, smem_pipe_write_k, smem_pipe_write_v,
|
||||
shared_storage, work_tile_info, work_idx, tile_count_semaphore);
|
||||
}
|
||||
collective_mainloop.load_tail(pipeline_q, pipeline_k, pipeline_v,
|
||||
smem_pipe_write_q, smem_pipe_write_k, smem_pipe_write_v);
|
||||
} else if (producer_warp_role == ProducerWarpRole::Epilogue) {
|
||||
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
|
||||
barrier_o.wait();
|
||||
collective_epilogue.tma_store(shared_storage, epilogue_params, work_tile_info, scheduler_params, threadIdx.x);
|
||||
collective_epilogue.store_tail();
|
||||
barrier_o.arrive();
|
||||
}
|
||||
|
||||
}
|
||||
} else if (warp_group_role == WarpGroupRole::Consumer0 || warp_group_role == WarpGroupRole::Consumer1) {
|
||||
cutlass::arch::warpgroup_reg_alloc<232>();
|
||||
typename Ktraits::TiledMmaPV tiled_mma_pv;
|
||||
TileScheduler scheduler{};
|
||||
PipelineState smem_pipe_read_k, smem_pipe_read_v;
|
||||
PipelineStateQ smem_pipe_read_q;
|
||||
|
||||
int work_idx = 0;
|
||||
|
||||
CUTLASS_PRAGMA_NO_UNROLL
|
||||
for (auto work_tile_info = scheduler.get_initial_work(); work_tile_info.is_valid(scheduler_params); work_tile_info = scheduler.get_next_work(scheduler_params, work_tile_info)) {
|
||||
// Attention output (GEMM-II) accumulator.
|
||||
Tensor tOrO = partition_fragment_C(tiled_mma_pv, select<0, 2>(TileShape_MNK{}));
|
||||
// flash::Softmax<2 * (2 * kBlockM / NumMmaThreads)> softmax;
|
||||
// Pass single_level_p_quant flag to control P quantization mode
|
||||
flash::SoftmaxFused<2 * (2 * kBlockM / NumMmaThreads)> softmax_fused(params.single_level_p_quant);
|
||||
auto block_coord = work_tile_info.get_block_coord(scheduler_params);
|
||||
auto [m_block, bidh, bidb] = block_coord;
|
||||
|
||||
int n_block_max = collective_mainloop.get_n_block_max(mainloop_params, m_block);
|
||||
if (Is_causal && n_block_max <= 0) { // We exit early and write 0 to gO and -inf to gLSE.
|
||||
collective_epilogue.store_zero(epilogue_params, threadIdx.x - NumCopyThreads, block_coord);
|
||||
continue;
|
||||
}
|
||||
|
||||
collective_mainloop.mma(mainloop_params, pipeline_q, pipeline_k, pipeline_v, smem_pipe_read_q, smem_pipe_read_k, smem_pipe_read_v,
|
||||
tOrO, softmax_fused, n_block_max, threadIdx.x - NumCopyThreads, work_idx, m_block, shared_storage);
|
||||
barrier_o.wait();
|
||||
collective_epilogue.mma_store(shared_storage, tiled_mma_pv, tOrO, threadIdx.x - NumCopyThreads);
|
||||
barrier_o.arrive();
|
||||
++work_idx;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace flash
|
||||
@@ -0,0 +1,114 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <ATen/cuda/CUDAContext.h>
|
||||
|
||||
#include "cute/tensor.hpp"
|
||||
|
||||
#include "cutlass/cluster_launch.hpp"
|
||||
|
||||
#include "static_switch.h"
|
||||
#include "params.h"
|
||||
#include "tile_scheduler.h"
|
||||
#include "kernel_ws.h"
|
||||
#include "kernel_traits.h"
|
||||
#include "block_config.h"
|
||||
|
||||
|
||||
template<typename Kernel_traits, bool Is_causal>
|
||||
void run_flash_fwd(Flash_fwd_params ¶ms, cudaStream_t stream) {
|
||||
using Element = typename Kernel_traits::Element;
|
||||
using ElementSF = typename Kernel_traits::ElementSF;
|
||||
using ElementOut = typename Kernel_traits::ElementOut;
|
||||
using TileShape_MNK = typename Kernel_traits::TileShape_MNK;
|
||||
using ClusterShape = typename Kernel_traits::ClusterShape_MNK;
|
||||
using CollectiveMainloop = flash::CollectiveMainloopFwd<Kernel_traits, Is_causal>;
|
||||
using CollectiveEpilogue = flash::CollectiveEpilogueFwd<Kernel_traits>;
|
||||
// using Scheduler = flash::SingleTileScheduler;
|
||||
using Scheduler = flash::StaticPersistentTileScheduler;
|
||||
typename CollectiveMainloop::Params mainloop_params =
|
||||
CollectiveMainloop::to_underlying_arguments({
|
||||
static_cast<Element const*>(params.q_ptr),
|
||||
{params.seqlen_q, params.d, params.h, params.b}, // shape_Q
|
||||
{params.q_row_stride, _1{}, params.q_head_stride, params.q_batch_stride}, // stride_Q
|
||||
static_cast<Element const*>(params.k_ptr),
|
||||
{params.seqlen_k, params.d, params.h_k, params.b}, // shape_K
|
||||
{params.k_row_stride, _1{}, params.k_head_stride, params.k_batch_stride}, // stride_K
|
||||
{params.unpadded_seqlen_k, params.d, params.h_k, params.b}, // shape_K
|
||||
static_cast<Element const*>(params.v_ptr),
|
||||
{params.d, params.seqlen_k, params.h_k, params.b}, // shape_Vt
|
||||
{params.v_row_stride, _1{}, params.v_head_stride, params.v_batch_stride}, // stride_Vt
|
||||
static_cast<ElementSF const*>(params.sfq_ptr),
|
||||
{params.seqlen_q, params.d, params.h, params.b}, // shape_SFQ
|
||||
static_cast<ElementSF const*>(params.sfk_ptr),
|
||||
{params.seqlen_k, params.d, params.h_k, params.b}, // shape_SFK
|
||||
static_cast<ElementSF const*>(params.sfv_ptr),
|
||||
{params.d, params.seqlen_k, params.h_k, params.b}, // shape_SFVt
|
||||
static_cast<float const*>(params.delta_s_ptr),
|
||||
{params.seqlen_s, params.seqlen_k, params.h_k, params.b},
|
||||
{params.ds_row_stride, _1{}, params.ds_head_stride, params.ds_batch_stride},
|
||||
params.scale_softmax_log2
|
||||
});
|
||||
typename CollectiveEpilogue::Params epilogue_params =
|
||||
CollectiveEpilogue::to_underlying_arguments({
|
||||
static_cast<ElementOut*>(params.o_ptr),
|
||||
{params.seqlen_q, params.d, params.h, params.b}, // shape_O
|
||||
{params.o_row_stride, _1{}, params.o_head_stride, params.o_batch_stride}, // stride_O
|
||||
static_cast<float*>(params.softmax_lse_ptr),
|
||||
{_1{}, params.seqlen_q, params.h * params.seqlen_q}, // stride_LSE
|
||||
});
|
||||
|
||||
int num_blocks_m = cutlass::ceil_div(params.seqlen_q, Kernel_traits::kBlockM);
|
||||
num_blocks_m = cutlass::ceil_div(num_blocks_m, size<0>(ClusterShape{})) * size<0>(ClusterShape{});
|
||||
typename Scheduler::Arguments scheduler_args = {num_blocks_m, params.h, params.b};
|
||||
typename Scheduler::Params scheduler_params = Scheduler::to_underlying_arguments(scheduler_args);
|
||||
// Get the ptr to kernel function.
|
||||
void *kernel;
|
||||
kernel = (void *)flash::compute_attn_ws<Kernel_traits, Is_causal, Scheduler>;
|
||||
int smem_size = sizeof(typename Kernel_traits::SharedStorage);
|
||||
if (smem_size >= 48 * 1024) {
|
||||
C10_CUDA_CHECK(cudaFuncSetAttribute(kernel, cudaFuncAttributeMaxDynamicSharedMemorySize, smem_size));
|
||||
}
|
||||
static constexpr int ctaSize = Kernel_traits::kNWarps * 32;
|
||||
params.m_block_divmod = cutlass::FastDivmod(num_blocks_m);
|
||||
params.total_blocks = num_blocks_m * params.h * params.b;
|
||||
dim3 grid_dims = Scheduler::get_grid_dim(scheduler_args, 170);
|
||||
dim3 block_dims(ctaSize);
|
||||
dim3 cluster_dims(size<0>(ClusterShape{}), size<1>(ClusterShape{}), size<2>(ClusterShape{}));
|
||||
cutlass::ClusterLaunchParams launch_params{grid_dims, block_dims, cluster_dims, smem_size, stream};
|
||||
cutlass::launch_kernel_on_cluster(launch_params, kernel, params, mainloop_params, epilogue_params, scheduler_params);
|
||||
|
||||
C10_CUDA_KERNEL_LAUNCH_CHECK();
|
||||
}
|
||||
|
||||
|
||||
template<typename T, int Headdim, typename O = cutlass::bfloat16_t>
|
||||
void run_mha_fwd_(Flash_fwd_params ¶ms, cudaStream_t stream) {
|
||||
BOOL_SWITCH(params.is_causal, Is_causal, [&] {
|
||||
BOOL_SWITCH(params.per_block_mean, per_block, [&] {
|
||||
if constexpr (Headdim == 64 || Headdim == 128) {
|
||||
run_flash_fwd<
|
||||
Flash_fwd_kernel_traits<Headdim, flash::BLOCK_M, flash::BLOCK_N, 3, 1, per_block, T, O>,
|
||||
Is_causal
|
||||
>(params, stream);
|
||||
} else {
|
||||
static_assert(Headdim == 64 || Headdim == 128, "Unsupported Headdim");
|
||||
}
|
||||
});
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,920 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <cutlass/cutlass.h>
|
||||
#include <cutlass/array.h>
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/numeric_conversion.h>
|
||||
#include "cutlass/pipeline/pipeline.hpp"
|
||||
|
||||
#include "cute/tensor.hpp"
|
||||
|
||||
#include "cutlass/gemm/collective/collective_builder.hpp"
|
||||
|
||||
#include "utils.h"
|
||||
#include "named_barrier.h"
|
||||
namespace flash {
|
||||
|
||||
using namespace cute;
|
||||
|
||||
template <typename Ktraits, bool Is_causal>
|
||||
struct CollectiveMainloopFwd {
|
||||
|
||||
using Element = typename Ktraits::Element;
|
||||
using ElementSF = typename Ktraits::ElementSF;
|
||||
// using TMAElement = Element;
|
||||
// using TMAElementSF = typename Ktraits::ElementSF;
|
||||
using TileShape_MNK = typename Ktraits::TileShape_MNK;
|
||||
using ClusterShape = typename Ktraits::ClusterShape_MNK;
|
||||
|
||||
static constexpr int kStages = Ktraits::kStages;
|
||||
static constexpr int kHeadDim = Ktraits::kHeadDim;
|
||||
static constexpr int BlockMean = Ktraits::BlockMean;
|
||||
using GmemTiledCopy = typename Ktraits::GmemTiledCopy;
|
||||
using SmemLayoutQ = typename Ktraits::SmemLayoutQ;
|
||||
using SmemLayoutK = typename Ktraits::SmemLayoutK;
|
||||
using SmemLayoutV = typename Ktraits::SmemLayoutV;
|
||||
using SmemLayoutVt = typename Ktraits::SmemLayoutVt;
|
||||
using SmemLayoutDS = typename Ktraits::SmemLayoutDS;
|
||||
using SmemLayoutAtomDS = typename Ktraits::SmemLayoutAtomDS;
|
||||
using LayoutDS = decltype(
|
||||
blocked_product(
|
||||
SmemLayoutAtomDS{},
|
||||
make_layout(
|
||||
make_shape(int32_t(0), int32_t(0), int32_t(0), int32_t(0)),
|
||||
make_stride(int32_t(0), _1{}, int32_t(0), int32_t(0)))
|
||||
)
|
||||
);
|
||||
using ShapeQKV = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen, d, head, batch)
|
||||
using StrideQKV = cute::Stride<int64_t, _1, int64_t, int64_t>;
|
||||
using ShapeSF = cute::Shape<int32_t, int32_t, int32_t, int32_t>; // (seqlen, d // 16, head, batch)
|
||||
using LayoutSF = typename Ktraits::LayoutSF;
|
||||
using LayoutP = typename Ktraits::LayoutP;
|
||||
using LayoutSFP = typename Ktraits::LayoutSFP;
|
||||
using SfAtom = typename Ktraits::SfAtom;
|
||||
using TMA_Q = decltype(make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
|
||||
SmemLayoutQ{},
|
||||
select<0, 2>(TileShape_MNK{}),
|
||||
_1{}));
|
||||
|
||||
using TMA_KV = decltype(make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
|
||||
take<0, 2>(SmemLayoutK{}),
|
||||
select<1, 2>(TileShape_MNK{}),
|
||||
_1{}));
|
||||
|
||||
using TMA_Vt = decltype(make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
make_tensor(make_gmem_ptr(static_cast<Element const*>(nullptr)), repeat_like(StrideQKV{}, int32_t(0)), StrideQKV{}),
|
||||
take<0, 2>(SmemLayoutVt{}),
|
||||
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{}));
|
||||
|
||||
using TMA_DS = decltype(make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
make_tensor(make_gmem_ptr(static_cast<float const*>(nullptr)), LayoutDS{}),
|
||||
take<0, 2>(SmemLayoutDS{}),
|
||||
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{}));
|
||||
|
||||
using BlkScaledConfig = typename Ktraits::BlkScaledConfig;
|
||||
using GmemTiledCopySF = typename Ktraits::GmemTiledCopySF;
|
||||
using SmemLayoutSFQ = typename Ktraits::SmemLayoutSFQ;
|
||||
using SmemLayoutSFK = typename Ktraits::SmemLayoutSFK;
|
||||
using SmemLayoutSFV = typename Ktraits::SmemLayoutSFV;
|
||||
using SmemLayoutSFVt = typename Ktraits::SmemLayoutSFVt;
|
||||
|
||||
using TMA_SFQ = decltype(make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
|
||||
SmemLayoutSFQ{},
|
||||
make_shape(shape<0>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
|
||||
_1{})); // No programmatic multicast
|
||||
|
||||
|
||||
using TMA_SFKV = decltype(make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
|
||||
SmemLayoutSFK{}(_,_,cute::Int<0>{}),
|
||||
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
|
||||
_1{}));
|
||||
|
||||
using TMA_SFVt = decltype(make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
make_tensor(static_cast<ElementSF const*>(nullptr), LayoutSF{}),
|
||||
SmemLayoutSFVt{}(_,_,cute::Int<0>{}),
|
||||
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{}));
|
||||
|
||||
using SmemCopyAtomQ = typename Ktraits::SmemCopyAtomQ;
|
||||
using SmemCopyAtomKV = typename Ktraits::SmemCopyAtomKV;
|
||||
using SmemCopyAtomSF = typename Ktraits::SmemCopyAtomSF;
|
||||
using TiledMmaQK = typename Ktraits::TiledMmaQK;
|
||||
using TiledMmaPV = typename Ktraits::TiledMmaPV;
|
||||
static constexpr int NumMmaThreads = size(TiledMmaQK{});
|
||||
using MainloopPipeline = typename Ktraits::MainloopPipeline;
|
||||
using PipelineParams = typename MainloopPipeline::Params;
|
||||
using PipelineState = typename MainloopPipeline::PipelineState;
|
||||
using MainloopPipelineQ = typename Ktraits::MainloopPipelineQ;
|
||||
using PipelineParamsQ = typename Ktraits::PipelineParamsQ;
|
||||
using PipelineStateQ = typename Ktraits::PipelineStateQ;
|
||||
using EpilogueBarrier = typename Ktraits::EpilogueBarrier;
|
||||
|
||||
// Set the bytes transferred in this TMA transaction (may involve multiple issues)
|
||||
static constexpr uint32_t TmaTransactionBytesQ = static_cast<uint32_t>(
|
||||
cutlass::bits_to_bytes(cosize((SmemLayoutSFQ{})) * cute::sizeof_bits_v<ElementSF>) +
|
||||
cutlass::bits_to_bytes(size((SmemLayoutQ{})) * sizeof_bits<Element>::value));
|
||||
|
||||
static constexpr uint32_t TmaTransactionBytesK = static_cast<uint32_t>(
|
||||
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutSFK{})) * cute::sizeof_bits_v<ElementSF>) +
|
||||
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutDS{})) * cute::sizeof_bits_v<float>) +
|
||||
cutlass::bits_to_bytes(size(take<0,2>(SmemLayoutK{})) * sizeof_bits<Element>::value));
|
||||
|
||||
static constexpr uint32_t TmaTransactionBytesV = static_cast<uint32_t>(
|
||||
cutlass::bits_to_bytes(cosize(take<0,2>(SmemLayoutSFVt{})) * cute::sizeof_bits_v<ElementSF>) +
|
||||
cutlass::bits_to_bytes(size(take<0,2>(SmemLayoutVt{})) * sizeof_bits<Element>::value));
|
||||
|
||||
// Host side kernel arguments
|
||||
struct Arguments {
|
||||
Element const* ptr_Q;
|
||||
ShapeQKV const shape_Q;
|
||||
StrideQKV const stride_Q;
|
||||
Element const* ptr_K;
|
||||
ShapeQKV const shape_K;
|
||||
StrideQKV const stride_K;
|
||||
ShapeQKV const unpadded_shape_K;
|
||||
Element const* ptr_Vt;
|
||||
ShapeQKV const shape_Vt;
|
||||
StrideQKV const stride_Vt;
|
||||
ElementSF const* ptr_SFQ{nullptr};
|
||||
ShapeSF const shape_SFQ{};
|
||||
ElementSF const* ptr_SFK{nullptr};
|
||||
ShapeSF const shape_SFK{};
|
||||
ElementSF const* ptr_SFVt{nullptr};
|
||||
ShapeSF const shape_SFVt{};
|
||||
float const* ptr_ds;
|
||||
ShapeQKV const shape_ds;
|
||||
StrideQKV const stride_ds;
|
||||
float const softmax_scale_log2;
|
||||
};
|
||||
|
||||
// Device side kernel params
|
||||
struct Params {
|
||||
ShapeQKV const shape_Q;
|
||||
LayoutSF const layout_SFQ;
|
||||
ShapeQKV const shape_K;
|
||||
ShapeQKV const unpadded_shape_K;
|
||||
LayoutSF const layout_SFK;
|
||||
ShapeQKV const shape_Vt;
|
||||
LayoutSF const layout_SFVt;
|
||||
LayoutDS const layout_DS;
|
||||
TMA_Q tma_load_Q;
|
||||
TMA_SFQ tma_load_SFQ;
|
||||
TMA_KV tma_load_K;
|
||||
TMA_SFKV tma_load_SFK;
|
||||
TMA_Vt tma_load_Vt;
|
||||
TMA_SFVt tma_load_SFVt;
|
||||
TMA_DS tma_load_DS;
|
||||
float const softmax_scale_log2;
|
||||
};
|
||||
|
||||
|
||||
static Params
|
||||
to_underlying_arguments(Arguments const& args) {
|
||||
Tensor mQ = make_tensor(make_gmem_ptr(args.ptr_Q), args.shape_Q, args.stride_Q);
|
||||
TMA_Q tma_load_Q = make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
mQ,
|
||||
SmemLayoutQ{},
|
||||
select<0, 2>(TileShape_MNK{}),
|
||||
_1{}); // no mcast for Q
|
||||
Tensor mK = make_tensor(make_gmem_ptr(args.ptr_K), args.shape_K, args.stride_K);
|
||||
TMA_KV tma_load_K = make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
mK,
|
||||
SmemLayoutK{}(_, _, _0{}),
|
||||
select<1, 2>(TileShape_MNK{}),
|
||||
_1{}); // mcast along M mode for this N load, if any
|
||||
Tensor mVt = make_tensor(make_gmem_ptr(args.ptr_Vt), args.shape_Vt, args.stride_Vt);
|
||||
TMA_Vt tma_load_Vt = make_tma_copy(
|
||||
GmemTiledCopy{},
|
||||
mVt,
|
||||
SmemLayoutVt{}(_, _, _0{}),
|
||||
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{}); // mcast along M mode for this N load, if any
|
||||
auto [Seqlen_Q, Seqlen_K, HeadNum, Batch] = args.shape_ds;
|
||||
LayoutDS layout_ds = tile_to_shape(SmemLayoutAtomDS{}, make_shape(Seqlen_Q, Seqlen_K, HeadNum, Batch), Step<_2,_1,_3,_4>{});
|
||||
Tensor mDS = make_tensor(make_gmem_ptr(args.ptr_ds), layout_ds);
|
||||
TMA_DS tma_load_ds = make_tma_copy (
|
||||
GmemTiledCopy{},
|
||||
mDS,
|
||||
SmemLayoutDS{}(_, _, _0{}),
|
||||
make_shape(shape<0>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{});
|
||||
LayoutSF layout_sfq = BlkScaledConfig::tile_atom_to_shape_SFQKV(args.shape_SFQ);
|
||||
Tensor mSFQ = make_tensor(make_gmem_ptr(args.ptr_SFQ), layout_sfq);
|
||||
TMA_SFQ tma_load_sfq = make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
mSFQ,
|
||||
SmemLayoutSFQ{},
|
||||
make_shape(shape<0>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
|
||||
_1{});
|
||||
LayoutSF layout_sfk = BlkScaledConfig::tile_atom_to_shape_SFQKV(args.shape_SFK);
|
||||
Tensor mSFK = make_tensor(make_gmem_ptr(args.ptr_SFK), layout_sfk);
|
||||
TMA_SFKV tma_load_sfk = make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
mSFK,
|
||||
SmemLayoutSFK{}(_, _, _0{}),
|
||||
make_shape(shape<1>(TileShape_MNK{}), shape<2>(TileShape_MNK{})),
|
||||
_1{});
|
||||
LayoutSF layout_sfvt = BlkScaledConfig::tile_atom_to_shape_SFVt(args.shape_SFVt);
|
||||
Tensor mSFVt = make_tensor(make_gmem_ptr(args.ptr_SFVt), layout_sfvt);
|
||||
TMA_SFVt tma_load_sfvt = make_tma_copy<uint16_t>(
|
||||
GmemTiledCopySF{},
|
||||
mSFVt,
|
||||
SmemLayoutSFVt{}(_, _, _0{}),
|
||||
make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})),
|
||||
_1{});
|
||||
return {args.shape_Q, layout_sfq,
|
||||
args.shape_K, args.unpadded_shape_K, layout_sfk,
|
||||
args.shape_Vt, layout_sfvt,
|
||||
layout_ds,
|
||||
tma_load_Q, tma_load_sfq,
|
||||
tma_load_K, tma_load_sfk,
|
||||
tma_load_Vt, tma_load_sfvt,
|
||||
tma_load_ds,
|
||||
args.softmax_scale_log2};
|
||||
}
|
||||
|
||||
/// Issue Tma Descriptor Prefetch -- ideally from a single thread for best performance
|
||||
CUTLASS_DEVICE
|
||||
static void prefetch_tma_descriptors(Params const& mainloop_params) {
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_Q.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_K.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_Vt.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFQ.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFK.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_SFVt.get_tma_descriptor());
|
||||
cute::prefetch_tma_descriptor(mainloop_params.tma_load_DS.get_tma_descriptor());
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
int get_n_block_max(Params const& mainloop_params, int m_block) {
|
||||
static constexpr int kBlockM = get<0>(TileShape_MNK{});
|
||||
static constexpr int kBlockN = get<1>(TileShape_MNK{});
|
||||
int const seqlen_q = get<0>(mainloop_params.shape_Q);
|
||||
int const seqlen_k = get<0>(mainloop_params.shape_K);
|
||||
int n_block_max = cute::ceil_div(seqlen_k, kBlockN);
|
||||
if constexpr (Is_causal) {
|
||||
n_block_max = std::min(n_block_max,
|
||||
cute::ceil_div((m_block + 1) * kBlockM + seqlen_k - seqlen_q, kBlockN));
|
||||
}
|
||||
return n_block_max;
|
||||
}
|
||||
|
||||
template <class SFATensor, class Atom, class TiledThr, class TiledPerm>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
thrfrg_SFA(SFATensor&& sfatensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
|
||||
{
|
||||
CUTE_STATIC_ASSERT_V(rank(sfatensor) >= Int<2>{});
|
||||
|
||||
using AtomShape_MNK = typename Atom::Shape_MNK;
|
||||
using AtomLayoutSFA_TV = typename Atom::Traits::SFALayout;
|
||||
|
||||
auto permutation_mnk = TiledPerm{};
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// Reorder the tensor for the TiledAtom
|
||||
auto t_tile = make_tile(get<0>(permutation_mnk),
|
||||
get<2>(permutation_mnk));
|
||||
auto t_tensor = logical_divide(sfatensor, t_tile); // (PermM,PermK)
|
||||
|
||||
// Tile the tensor for the Atom
|
||||
auto a_tile = make_tile(make_layout(size<0>(AtomShape_MNK{})),
|
||||
make_layout(size<2>(AtomShape_MNK{})));
|
||||
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomM,AtomK),(RestM,RestK))
|
||||
|
||||
// Transform the Atom mode from (M,K) to (Thr,Val)
|
||||
auto tv_tensor = a_tensor.compose(AtomLayoutSFA_TV{},_); // ((ThrV,FrgV),(RestM,RestK))
|
||||
|
||||
// Tile the tensor for the Thread
|
||||
auto thr_tile = make_tile(_,
|
||||
make_tile(make_layout(size<1>(thr_layout_vmnk)),
|
||||
make_layout(size<3>(thr_layout_vmnk))));
|
||||
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrM,ThrK)),(FrgV,(RestM,RestK)))
|
||||
|
||||
return thr_tensor;
|
||||
}
|
||||
|
||||
template <class SFBTensor, class Atom, class TiledThr, class TiledPerm>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
thrfrg_SFB(SFBTensor&& sfbtensor, TiledMMA<Atom, TiledThr, TiledPerm>& mma)
|
||||
{
|
||||
CUTE_STATIC_ASSERT_V(rank(sfbtensor) >= Int<2>{});
|
||||
|
||||
using AtomShape_MNK = typename Atom::Shape_MNK;
|
||||
using AtomLayoutSFB_TV = typename Atom::Traits::SFBLayout;
|
||||
|
||||
auto permutation_mnk = TiledPerm{};
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// Reorder the tensor for the TiledAtom
|
||||
auto t_tile = make_tile(get<1>(permutation_mnk),
|
||||
get<2>(permutation_mnk));
|
||||
auto t_tensor = logical_divide(sfbtensor, t_tile); // (PermN,PermK)
|
||||
|
||||
// Tile the tensor for the Atom
|
||||
auto a_tile = make_tile(make_layout(size<1>(AtomShape_MNK{})),
|
||||
make_layout(size<2>(AtomShape_MNK{})));
|
||||
auto a_tensor = zipped_divide(t_tensor, a_tile); // ((AtomN,AtomK),(RestN,RestK))
|
||||
|
||||
// Transform the Atom mode from (M,K) to (Thr,Val)
|
||||
auto tv_tensor = a_tensor.compose(AtomLayoutSFB_TV{},_); // ((ThrV,FrgV),(RestN,RestK))
|
||||
|
||||
// Tile the tensor for the Thread
|
||||
auto thr_tile = make_tile(_,
|
||||
make_tile(make_layout(size<2>(thr_layout_vmnk)),
|
||||
make_layout(size<3>(thr_layout_vmnk))));
|
||||
auto thr_tensor = zipped_divide(tv_tensor, thr_tile); // ((ThrV,(ThrN,ThrK)),(FrgV,(RestN,RestK)))
|
||||
return thr_tensor;
|
||||
}
|
||||
|
||||
template <class SFATensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_fragment_SFA(SFATensor&& sfatensor, ThrMma& thread_mma)
|
||||
{
|
||||
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
|
||||
auto thr_tensor = make_tensor(static_cast<SFATensor&&>(sfatensor).data(), thrfrg_SFA(sfatensor.layout(),thread_mma));
|
||||
auto thr_vmnk = thread_mma.thr_vmnk_;
|
||||
auto thr_vmk = make_coord(get<0>(thr_vmnk), make_coord(get<1>(thr_vmnk), get<3>(thr_vmnk)));
|
||||
auto partition_SFA = thr_tensor(thr_vmk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
|
||||
return make_fragment_like<ValTypeSF>(partition_SFA);
|
||||
}
|
||||
|
||||
template <class SFBTensor, class ThrMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
partition_fragment_SFB(SFBTensor&& sfbtensor, ThrMma& thread_mma)
|
||||
{
|
||||
using ValTypeSF = typename ThrMma::Atom::Traits::ValTypeSF;
|
||||
auto thr_tensor = make_tensor(static_cast<SFBTensor&&>(sfbtensor).data(), thrfrg_SFB(sfbtensor.layout(),thread_mma));
|
||||
auto thr_vmnk = thread_mma.thr_vmnk_;
|
||||
auto thr_vnk = make_coord(get<0>(thr_vmnk), make_coord(get<2>(thr_vmnk), get<3>(thr_vmnk)));
|
||||
auto partition_SFB = thr_tensor(thr_vnk, make_coord(_, repeat<rank<1,1>(thr_tensor)>(_)));
|
||||
return make_fragment_like<ValTypeSF>(partition_SFB);
|
||||
}
|
||||
|
||||
template<class TiledMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
get_layoutSFA_TV(TiledMma& mma)
|
||||
{
|
||||
// (M,K) -> (M,K)
|
||||
auto tile_shape_mnk = tile_shape(mma);
|
||||
auto ref_A = make_layout(make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk)));
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
|
||||
auto atile = make_tile(_,
|
||||
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
|
||||
make_stride( Int<1>{} , Int<0>{} )),
|
||||
_));
|
||||
|
||||
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
|
||||
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
|
||||
// (thr_idx,val) -> (M,K)
|
||||
return thrfrg_SFA(ref_A, mma).compose(atile, _).compose(thridx_2_thrid, _);
|
||||
}
|
||||
|
||||
template<class TiledMma>
|
||||
CUTE_HOST_DEVICE constexpr
|
||||
auto
|
||||
get_layoutSFB_TV(TiledMma& mma)
|
||||
{
|
||||
// (N,K) -> (N,K)
|
||||
auto tile_shape_mnk = tile_shape(mma);
|
||||
auto ref_B = make_layout(make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk)));
|
||||
auto thr_layout_vmnk = mma.get_thr_layout_vmnk();
|
||||
|
||||
// (ThrV,(ThrM,ThrK)) -> (ThrV,(ThrM,ThrN,ThrK))
|
||||
auto btile = make_tile(_,
|
||||
make_tile(make_layout(make_shape (size<1>(thr_layout_vmnk), size<2>(thr_layout_vmnk)),
|
||||
make_stride( Int<0>{} , Int<1>{} )),
|
||||
_));
|
||||
|
||||
// thr_idx -> (ThrV,ThrM,ThrN,ThrK)
|
||||
auto thridx_2_thrid = right_inverse(thr_layout_vmnk);
|
||||
// (thr_idx,val) -> (M,K)
|
||||
return thrfrg_SFB(ref_B, mma).compose(btile, _).compose(thridx_2_thrid, _);
|
||||
}
|
||||
|
||||
template <typename SchedulerParams, typename SharedStorage, typename WorkTileInfo>
|
||||
CUTLASS_DEVICE void
|
||||
load(Params const& mainloop_params,
|
||||
SchedulerParams const& scheduler_params,
|
||||
MainloopPipelineQ pipeline_q,
|
||||
MainloopPipeline pipeline_k,
|
||||
MainloopPipeline pipeline_v,
|
||||
PipelineStateQ& smem_pipe_write_q,
|
||||
PipelineState& smem_pipe_write_k,
|
||||
PipelineState& smem_pipe_write_v,
|
||||
SharedStorage &shared_storage,
|
||||
WorkTileInfo work_tile_info,
|
||||
int& work_idx,
|
||||
int& tile_count_semaphore
|
||||
) {
|
||||
|
||||
static constexpr int kBlockM = get<0>(TileShape_MNK{});
|
||||
static constexpr int kBlockN = get<1>(TileShape_MNK{});
|
||||
|
||||
auto [m_block, bidh, bidb] = work_tile_info.get_block_coord(scheduler_params);
|
||||
|
||||
int n_block_max = get_n_block_max(mainloop_params, m_block);
|
||||
|
||||
Tensor sQ = make_tensor(make_smem_ptr(shared_storage.smem_q.begin()), SmemLayoutQ{});
|
||||
Tensor sK = make_tensor(make_smem_ptr(shared_storage.smem_k.begin()), SmemLayoutK{});
|
||||
Tensor sVt = make_tensor(make_smem_ptr(shared_storage.smem_v.begin()), SmemLayoutVt{});
|
||||
Tensor sSFQ = make_tensor(make_smem_ptr(shared_storage.smem_SFQ.begin()), SmemLayoutSFQ{});
|
||||
Tensor sSFK = make_tensor(make_smem_ptr(shared_storage.smem_SFK.begin()), SmemLayoutSFK{});
|
||||
Tensor sSFVt = make_tensor(make_smem_ptr(shared_storage.smem_SFV.begin()), SmemLayoutSFVt{});
|
||||
Tensor sDS = make_tensor(make_smem_ptr(shared_storage.smem_ds.begin()), SmemLayoutDS{});
|
||||
|
||||
Tensor mQ = mainloop_params.tma_load_Q.get_tma_tensor(mainloop_params.shape_Q);
|
||||
Tensor mK = mainloop_params.tma_load_K.get_tma_tensor(mainloop_params.shape_K);
|
||||
Tensor mVt = mainloop_params.tma_load_Vt.get_tma_tensor(mainloop_params.shape_Vt);
|
||||
Tensor mDS = mainloop_params.tma_load_DS.get_tma_tensor(shape(mainloop_params.layout_DS));
|
||||
Tensor mSFQ = mainloop_params.tma_load_SFQ.get_tma_tensor(shape(mainloop_params.layout_SFQ));
|
||||
Tensor mSFK = mainloop_params.tma_load_SFK.get_tma_tensor(shape(mainloop_params.layout_SFK));
|
||||
Tensor mSFVt = mainloop_params.tma_load_SFVt.get_tma_tensor(shape(mainloop_params.layout_SFVt));
|
||||
uint32_t block_rank_in_cluster = cute::block_rank_in_cluster();
|
||||
constexpr uint32_t cluster_shape_x = get<0>(ClusterShape());
|
||||
uint2 cluster_local_block_id = {block_rank_in_cluster % cluster_shape_x, block_rank_in_cluster / cluster_shape_x};
|
||||
Tensor gQ = local_tile(mQ(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{})); // (M, K)
|
||||
Tensor gK = local_tile(mK(_, _, bidh, bidb), select<1, 2>(TileShape_MNK{}), make_coord(_, _0{})); // (N, K, _)
|
||||
Tensor gVt = local_tile(mVt(_, _, bidh, bidb), make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})), make_coord(_0{}, _)); // (N, K, _)
|
||||
Tensor gDS = [&] {
|
||||
if constexpr (BlockMean) {
|
||||
return local_tile(mDS(_, _, bidh, bidb), select<0, 1>(TileShape_MNK{}), make_coord(m_block, _));
|
||||
} else {
|
||||
return local_tile(mDS(_, _, bidh, bidb), select<0, 1>(TileShape_MNK{}), make_coord(_0{}, _));
|
||||
}
|
||||
}();
|
||||
Tensor gSFQ = local_tile(mSFQ(_, _, bidh, bidb), select<0, 2>(TileShape_MNK{}), make_coord(m_block, _0{}));
|
||||
Tensor gSFK = local_tile(mSFK(_, _, bidh, bidb), select<1, 2>(TileShape_MNK{}), make_coord(_, _0{}));
|
||||
Tensor gSFVt = local_tile(mSFVt(_, _, bidh, bidb), make_shape(shape<2>(TileShape_MNK{}), shape<1>(TileShape_MNK{})), make_coord(_0{}, _));
|
||||
auto block_tma_q = mainloop_params.tma_load_Q.get_slice(_0{});
|
||||
Tensor tQgQ = block_tma_q.partition_S(gQ);
|
||||
Tensor tQsQ = block_tma_q.partition_D(sQ);
|
||||
auto block_tma_sfq = mainloop_params.tma_load_SFQ.get_slice(_0{});
|
||||
Tensor tQgSFQ = block_tma_sfq.partition_S(gSFQ);
|
||||
Tensor tQsSFQ = block_tma_sfq.partition_D(sSFQ);
|
||||
auto block_tma_k = mainloop_params.tma_load_K.get_slice(cluster_local_block_id.x);
|
||||
Tensor tKgK = group_modes<0, 3>(block_tma_k.partition_S(gK));
|
||||
Tensor tKsK = group_modes<0, 3>(block_tma_k.partition_D(sK));
|
||||
auto block_tma_sfk = mainloop_params.tma_load_SFK.get_slice(cluster_local_block_id.x);
|
||||
Tensor tKgSFK = group_modes<0, 3>(block_tma_sfk.partition_S(gSFK));
|
||||
Tensor tKsSFK = group_modes<0, 3>(block_tma_sfk.partition_D(sSFK));
|
||||
auto block_tma_vt = mainloop_params.tma_load_Vt.get_slice(cluster_local_block_id.x);
|
||||
Tensor tVgVt = group_modes<0, 3>(block_tma_vt.partition_S(gVt));
|
||||
Tensor tVsVt = group_modes<0, 3>(block_tma_vt.partition_D(sVt));
|
||||
auto block_tma_sfvt = mainloop_params.tma_load_SFVt.get_slice(cluster_local_block_id.x);
|
||||
Tensor tVgSFVt = group_modes<0, 3>(block_tma_sfvt.partition_S(gSFVt));
|
||||
Tensor tVsSFVt = group_modes<0, 3>(block_tma_sfvt.partition_D(sSFVt));
|
||||
auto block_tma_ds = mainloop_params.tma_load_DS.get_slice(cluster_local_block_id.x);
|
||||
Tensor tDSgDS = group_modes<0, 3>(block_tma_ds.partition_S(gDS));
|
||||
Tensor tDSsDS = group_modes<0, 3>(block_tma_ds.partition_D(sDS));
|
||||
uint16_t mcast_mask_kv = 0;
|
||||
|
||||
int n_block = n_block_max - 1;
|
||||
int lane_predicate = cute::elect_one_sync();
|
||||
if (lane_predicate) {
|
||||
pipeline_q.producer_acquire(smem_pipe_write_q);
|
||||
copy(mainloop_params.tma_load_Q.with(*pipeline_q.producer_get_barrier(smem_pipe_write_q), 0), tQgQ, tQsQ);
|
||||
copy(mainloop_params.tma_load_SFQ.with(*pipeline_q.producer_get_barrier(smem_pipe_write_q), 0), tQgSFQ, tQsSFQ);
|
||||
++smem_pipe_write_q;
|
||||
pipeline_k.producer_acquire(smem_pipe_write_k);
|
||||
copy(mainloop_params.tma_load_K.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tKgK(_, n_block), tKsK(_, smem_pipe_write_k.index()));
|
||||
copy(mainloop_params.tma_load_SFK.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tKgSFK(_, n_block), tKsSFK(_, smem_pipe_write_k.index()));
|
||||
copy(mainloop_params.tma_load_DS.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tDSgDS(_, n_block), tDSsDS(_, smem_pipe_write_k.index()));
|
||||
++smem_pipe_write_k;
|
||||
pipeline_v.producer_acquire(smem_pipe_write_v);
|
||||
copy(mainloop_params.tma_load_Vt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
|
||||
tVgVt(_, n_block), tVsVt(_, smem_pipe_write_v.index()));
|
||||
copy(mainloop_params.tma_load_SFVt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
|
||||
tVgSFVt(_, n_block), tVsSFVt(_, smem_pipe_write_v.index()));
|
||||
++smem_pipe_write_v;
|
||||
}
|
||||
|
||||
n_block--;
|
||||
if (lane_predicate) {
|
||||
// CUTLASS_PRAGMA_NO_UNROLL
|
||||
#pragma unroll 2
|
||||
for (; n_block >= 0; --n_block) {
|
||||
pipeline_k.producer_acquire(smem_pipe_write_k);
|
||||
copy(mainloop_params.tma_load_K.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tKgK(_, n_block), tKsK(_, smem_pipe_write_k.index()));
|
||||
copy(mainloop_params.tma_load_SFK.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tKgSFK(_, n_block), tKsSFK(_, smem_pipe_write_k.index()));
|
||||
copy(mainloop_params.tma_load_DS.with(*pipeline_k.producer_get_barrier(smem_pipe_write_k), mcast_mask_kv),
|
||||
tDSgDS(_, n_block), tDSsDS(_, smem_pipe_write_k.index()));
|
||||
++smem_pipe_write_k;
|
||||
pipeline_v.producer_acquire(smem_pipe_write_v);
|
||||
copy(mainloop_params.tma_load_Vt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
|
||||
tVgVt(_, n_block), tVsVt(_, smem_pipe_write_v.index()));
|
||||
copy(mainloop_params.tma_load_SFVt.with(*pipeline_v.producer_get_barrier(smem_pipe_write_v), mcast_mask_kv),
|
||||
tVgSFVt(_, n_block), tVsSFVt(_, smem_pipe_write_v.index()));
|
||||
++smem_pipe_write_v;
|
||||
}
|
||||
}
|
||||
++work_idx;
|
||||
}
|
||||
|
||||
/// Perform a Producer Epilogue to prevent early exit of blocks in a Cluster
|
||||
CUTLASS_DEVICE void
|
||||
load_tail(MainloopPipelineQ pipeline_q,
|
||||
MainloopPipeline pipeline_k,
|
||||
MainloopPipeline pipeline_v,
|
||||
PipelineStateQ& smem_pipe_write_q,
|
||||
PipelineState& smem_pipe_write_k,
|
||||
PipelineState& smem_pipe_write_v) {
|
||||
int lane_predicate = cute::elect_one_sync();
|
||||
// Issue the epilogue waits
|
||||
if (lane_predicate) {
|
||||
pipeline_q.producer_tail(smem_pipe_write_q);
|
||||
pipeline_k.producer_tail(smem_pipe_write_k);
|
||||
pipeline_v.producer_tail(smem_pipe_write_v);
|
||||
}
|
||||
}
|
||||
|
||||
template <typename SharedStorage, typename FrgTensorO, typename SoftmaxFused>
|
||||
CUTLASS_DEVICE void
|
||||
mma(Params const& mainloop_params,
|
||||
MainloopPipelineQ pipeline_q,
|
||||
MainloopPipeline pipeline_k,
|
||||
MainloopPipeline pipeline_v,
|
||||
PipelineStateQ& smem_pipe_read_q,
|
||||
PipelineState& smem_pipe_read_k,
|
||||
PipelineState& smem_pipe_read_v,
|
||||
FrgTensorO& tOrO_store,
|
||||
SoftmaxFused& softmax_fused,
|
||||
int n_block_count,
|
||||
int thread_idx,
|
||||
int work_idx,
|
||||
int m_block,
|
||||
SharedStorage& shared_storage
|
||||
) {
|
||||
|
||||
static_assert(is_rmem<FrgTensorO>::value, "O tensor must be rmem resident.");
|
||||
|
||||
static constexpr int kBlockM = get<0>(TileShape_MNK{});
|
||||
static constexpr int kBlockN = get<1>(TileShape_MNK{});
|
||||
static constexpr int kBlockK = get<2>(TileShape_MNK{});
|
||||
Tensor sQ = make_tensor(make_smem_ptr(shared_storage.smem_q.begin()), SmemLayoutQ{});
|
||||
Tensor sK = make_tensor(make_smem_ptr(shared_storage.smem_k.begin()), SmemLayoutK{});
|
||||
Tensor sVt = make_tensor(make_smem_ptr(shared_storage.smem_v.begin()), SmemLayoutVt{});
|
||||
Tensor sDS = make_tensor(make_smem_ptr(shared_storage.smem_ds.begin()), SmemLayoutDS{});
|
||||
Tensor sSFQ = make_tensor(make_smem_ptr(shared_storage.smem_SFQ.begin()), SmemLayoutSFQ{});
|
||||
Tensor sSFK = make_tensor(make_smem_ptr(shared_storage.smem_SFK.begin()), SmemLayoutSFK{});
|
||||
Tensor sSFVt = make_tensor(make_smem_ptr(shared_storage.smem_SFV.begin()), SmemLayoutSFVt{});
|
||||
|
||||
Tensor cQ = make_identity_tensor(make_shape(size<0>(sQ), size<1>(sQ)));
|
||||
Tensor cKV = make_identity_tensor(make_shape(size<0>(sK), size<1>(sK)));
|
||||
TiledMmaQK tiled_mma_qk;
|
||||
TiledMmaPV tiled_mma_pv;
|
||||
auto thread_mma_qk = tiled_mma_qk.get_thread_slice(thread_idx);
|
||||
auto thread_mma_pv = tiled_mma_pv.get_thread_slice(thread_idx);
|
||||
|
||||
Tensor tSrQ = thread_mma_qk.partition_fragment_A(sQ);
|
||||
Tensor tSrK = thread_mma_qk.partition_fragment_B(sK(_,_,Int<0>{}));
|
||||
Tensor tOrVt = thread_mma_pv.partition_fragment_B(sVt(_,_,Int<0>{}));
|
||||
Tensor tOrP = make_tensor_like<Element>(LayoutP{});
|
||||
Tensor tSrSFQ = partition_fragment_SFA(sSFQ, thread_mma_qk);
|
||||
Tensor tSrSFK = partition_fragment_SFB(sSFK(_,_,Int<0>{}), thread_mma_qk);
|
||||
Tensor tOrSFVt = partition_fragment_SFB(sSFVt(_,_,Int<0>{}), thread_mma_pv);
|
||||
Tensor tOrSFP = make_tensor<ElementSF>(LayoutSFP{});
|
||||
Tensor tOrSFP_flt = filter_zeros(tOrSFP);
|
||||
Tensor tSrDS = make_tensor<float>(make_shape(_8{}, _4{}), make_stride(_1{}, _8{}));
|
||||
// copy qk and sf from smem to rmem
|
||||
auto smem_tiled_copy_Q = make_tiled_copy_A(SmemCopyAtomQ{}, tiled_mma_qk);
|
||||
auto smem_thr_copy_Q = smem_tiled_copy_Q.get_thread_slice(thread_idx);
|
||||
Tensor tSsQ = smem_thr_copy_Q.partition_S(as_position_independent_swizzle_tensor(sQ));
|
||||
Tensor tSrQ_copy_view = smem_thr_copy_Q.retile_D(tSrQ);
|
||||
|
||||
auto smem_tiled_copy_K = make_tiled_copy_B(SmemCopyAtomKV{}, tiled_mma_qk);
|
||||
auto smem_thr_copy_K = smem_tiled_copy_K.get_thread_slice(thread_idx);
|
||||
Tensor tSsK = smem_thr_copy_K.partition_S(as_position_independent_swizzle_tensor(sK));
|
||||
Tensor tSrK_copy_view = smem_thr_copy_K.retile_D(tSrK);
|
||||
|
||||
auto smem_tiled_copy_V = make_tiled_copy_B(SmemCopyAtomKV{}, tiled_mma_pv);
|
||||
auto smem_thr_copy_V = smem_tiled_copy_V.get_thread_slice(thread_idx);
|
||||
Tensor tOsVt = smem_thr_copy_V.partition_S(as_position_independent_swizzle_tensor(sVt));
|
||||
Tensor tOrVt_copy_view = smem_thr_copy_V.retile_D(tOrVt);
|
||||
|
||||
auto tile_shape_mnk = tile_shape(tiled_mma_qk);
|
||||
auto smem_tiled_copy_SFQ = make_tiled_copy_impl(SmemCopyAtomSF{},
|
||||
get_layoutSFA_TV(tiled_mma_qk),
|
||||
make_shape(size<0>(tile_shape_mnk), size<2>(tile_shape_mnk))
|
||||
);
|
||||
auto smem_thr_copy_SFQ = smem_tiled_copy_SFQ.get_thread_slice(thread_idx);
|
||||
Tensor tSsSFQ = smem_thr_copy_SFQ.partition_S(as_position_independent_swizzle_tensor(sSFQ));
|
||||
Tensor tSrSFQ_copy_view = smem_thr_copy_SFQ.retile_D(tSrSFQ);
|
||||
|
||||
auto smem_tiled_copy_SFK = make_tiled_copy_impl(SmemCopyAtomSF{},
|
||||
get_layoutSFB_TV(tiled_mma_qk),
|
||||
make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk))
|
||||
);
|
||||
auto smem_thr_copy_SFK = smem_tiled_copy_SFK.get_thread_slice(thread_idx);
|
||||
Tensor tSsSFK = smem_thr_copy_SFK.partition_S(as_position_independent_swizzle_tensor(sSFK));
|
||||
Tensor tSrSFK_copy_view = smem_thr_copy_SFK.retile_D(tSrSFK);
|
||||
|
||||
auto smem_tiled_copy_SFV = make_tiled_copy_impl(SmemCopyAtomSF{},
|
||||
get_layoutSFB_TV(tiled_mma_pv),
|
||||
make_shape(size<1>(tile_shape_mnk), size<2>(tile_shape_mnk))
|
||||
);
|
||||
auto smem_thr_copy_SFV = smem_tiled_copy_SFV.get_thread_slice(thread_idx);
|
||||
Tensor tOsSFVt = smem_thr_copy_SFV.partition_S(as_position_independent_swizzle_tensor(sSFVt));
|
||||
Tensor tOrSFVt_copy_view = smem_thr_copy_SFV.retile_D(tOrSFVt);
|
||||
|
||||
auto consumer_wait = [](auto& pipeline, auto& smem_pipe_read) {
|
||||
auto barrier_token = pipeline.consumer_try_wait(smem_pipe_read);
|
||||
pipeline.consumer_wait(smem_pipe_read, barrier_token);
|
||||
};
|
||||
|
||||
int const seqlen_q = get<0>(mainloop_params.shape_Q);
|
||||
int const seqlen_k = get<0>(mainloop_params.shape_K);
|
||||
int const unpadded_seqlen_k = get<0>(mainloop_params.unpadded_shape_K);
|
||||
int n_block = n_block_count - 1;
|
||||
|
||||
auto copy_k_block = [&](auto block_id) {
|
||||
auto tSsK_stage = tSsK(_, _, _, smem_pipe_read_k.index());
|
||||
auto tSsSFK_stage = tSsSFK(_, _, _, smem_pipe_read_k.index());
|
||||
copy(smem_tiled_copy_K, tSsK_stage(_, _, block_id), tSrK_copy_view(_, _, block_id));
|
||||
copy(smem_tiled_copy_SFK, tSsSFK_stage(_, _, block_id), tSrSFK_copy_view(_, _, block_id));
|
||||
};
|
||||
|
||||
auto copy_v_block = [&](auto block_id) {
|
||||
auto tOsVt_stage = tOsVt(_, _, _, smem_pipe_read_v.index());
|
||||
auto tOsSFVt_stage = tOsSFVt(_, _, _, smem_pipe_read_v.index());
|
||||
copy(smem_tiled_copy_V, tOsVt_stage(_, _, block_id), tOrVt_copy_view(_, _, block_id));
|
||||
copy(smem_tiled_copy_SFV, tOsSFVt_stage(_, _, block_id), tOrSFVt_copy_view(_, _, block_id));
|
||||
};
|
||||
// auto gemm_qk = [&](auto block_id) {
|
||||
// cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, block_id), tSrSFQ(_, _, block_id)), make_zip_tensor(tSrK(_, _, block_id), tSrSFK(_, _, block_id)), tSrS);
|
||||
// };
|
||||
// auto gemm_pv = [&](auto block_id) {
|
||||
// cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, block_id), tOrSFP(_, _, block_id)), make_zip_tensor(tOrVt(_, _, block_id), tOrSFVt(_, _, block_id)), tOrO);
|
||||
// };
|
||||
auto add_delta_s = [&](auto& acc) {
|
||||
// The MMA atom composites 4 sub-MMA m16n8k64 covering N=0-7, 8-15, 16-23, 24-31.
|
||||
// Each float4 register group spans two sub-MMAs, so N positions are scattered
|
||||
// (e.g., {2t, 2t+1, 8+2t, 9+2t}), not consecutive.
|
||||
float const* ds_ptr = reinterpret_cast<float const*>(
|
||||
&sDS(_0{}, _0{}, smem_pipe_read_k.index()));
|
||||
auto acc_float4 = recast<float4>(acc);
|
||||
int tid = threadIdx.x % 4;
|
||||
for (int i = 0; i < 4; i++) {
|
||||
int base_n = i * 32 + tid * 2;
|
||||
float4 delta_s_0 = make_float4(
|
||||
ds_ptr[base_n], ds_ptr[base_n + 1],
|
||||
ds_ptr[base_n + 8], ds_ptr[base_n + 9]);
|
||||
float4 delta_s_1 = make_float4(
|
||||
ds_ptr[base_n + 16], ds_ptr[base_n + 17],
|
||||
ds_ptr[base_n + 24], ds_ptr[base_n + 25]);
|
||||
acc_float4(make_coord(make_coord(_0{}, _0{}), _0{}), _0{}, i) = delta_s_0;
|
||||
acc_float4(make_coord(make_coord(_0{}, _0{}), _1{}), _0{}, i) = delta_s_0;
|
||||
acc_float4(make_coord(make_coord(_0{}, _1{}), _0{}), _0{}, i) = delta_s_1;
|
||||
acc_float4(make_coord(make_coord(_0{}, _1{}), _1{}), _0{}, i) = delta_s_1;
|
||||
}
|
||||
};
|
||||
consumer_wait(pipeline_q, smem_pipe_read_q);
|
||||
copy(smem_tiled_copy_Q, tSsQ, tSrQ_copy_view);
|
||||
copy(smem_tiled_copy_SFQ, tSsSFQ, tSrSFQ_copy_view);
|
||||
pipeline_q.consumer_release(smem_pipe_read_q);
|
||||
++smem_pipe_read_q;
|
||||
|
||||
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
|
||||
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
|
||||
Tensor AbsMaxP = make_tensor_like<float>(
|
||||
make_layout(shape(group<1, 4>(flatten(tSrS_converion_view.layout()(make_coord(_0{}, _), _, _)))))
|
||||
);
|
||||
consumer_wait(pipeline_k, smem_pipe_read_k);
|
||||
copy_k_block(_0{});
|
||||
add_delta_s(tSrS);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
|
||||
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
|
||||
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
|
||||
if (k_block < size<2>(tSrQ) - 1) {
|
||||
copy_k_block(k_block + 1);
|
||||
} else {
|
||||
pipeline_k.consumer_release(smem_pipe_read_k);
|
||||
++smem_pipe_read_k;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
auto col_limit_causal = [&](int row, int n_block) {
|
||||
return row + 1 + seqlen_k - n_block * kBlockN - seqlen_q + m_block * kBlockM;
|
||||
};
|
||||
{
|
||||
Tensor cS = cute::make_identity_tensor(select<0, 1>(TileShape_MNK{}));
|
||||
Tensor tScS = thread_mma_qk.partition_C(cS);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int i = 0; i < size(tSrS); ++i) {
|
||||
if constexpr (!Is_causal) { // Just masking based on col
|
||||
if (int(get<1>(tScS(i))) >= int(unpadded_seqlen_k - n_block * kBlockN)) { tSrS(i) = -INFINITY; }
|
||||
} else {
|
||||
if (int(get<1>(tScS(i))) >= std::min(seqlen_k - n_block * kBlockN,
|
||||
col_limit_causal(int(get<0>(tScS(i))), n_block))) {
|
||||
tSrS(i) = -INFINITY;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
auto quantize = [&](auto mma_k, auto acc_conversion_view) {
|
||||
Tensor AbsMaxP_stagek = AbsMaxP(_, make_coord(_, _, mma_k));
|
||||
Tensor acc_conversion_stagek = acc_conversion_view(_, _, mma_k);
|
||||
Tensor SFP = make_tensor_like<cutlass::float_ue4m3_t>(AbsMaxP_stagek.layout());
|
||||
Tensor SFP_uint32_view = recast<uint32_t>(SFP);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int i = 0; i < size(AbsMaxP_stagek); i += 4) {
|
||||
uint32_t& tmp = SFP_uint32_view(i / 4);
|
||||
flash::packed_float_to_ue4m3(
|
||||
AbsMaxP_stagek(i),
|
||||
AbsMaxP_stagek(i + 1),
|
||||
AbsMaxP_stagek(i + 2),
|
||||
AbsMaxP_stagek(i + 3),
|
||||
tmp
|
||||
);
|
||||
}
|
||||
int const quad_id = threadIdx.x & 3;
|
||||
uint32_t MASK = (0xFF00FF) << ((quad_id & 1) * 8);
|
||||
Tensor tOrSFP_uint32_view = recast<uint32_t>(tOrSFP(_, _, mma_k));
|
||||
Tensor tOrP_uint32_view = recast<uint32_t>(tOrP(_, _, mma_k));
|
||||
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mma_m = 0; mma_m < size<1>(tOrP); ++mma_m) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int i = 0; i < 4; ++i) {
|
||||
flash::packed_float_to_e2m1(
|
||||
acc_conversion_stagek(make_coord(_0{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_1{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_2{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_3{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_4{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_5{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_6{}, i), mma_m),
|
||||
acc_conversion_stagek(make_coord(_7{}, i), mma_m),
|
||||
tOrP_uint32_view(i, mma_m)
|
||||
);
|
||||
}
|
||||
uint32_t local_sfp = SFP_uint32_view(_0{}, _0{}, mma_m);
|
||||
uint32_t peer_sfp = __shfl_xor_sync(int32_t(-1), local_sfp, 2);
|
||||
if ((quad_id & 1) == 0) {
|
||||
uint32_t sfp = (local_sfp & MASK) | ((peer_sfp & MASK) << 8);
|
||||
tOrSFP_uint32_view(_0{}, mma_m) = sfp;
|
||||
} else {
|
||||
uint32_t sfp = (peer_sfp & MASK) | ((local_sfp & MASK) >> 8);
|
||||
tOrSFP_uint32_view(_0{}, mma_m) = sfp;
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
softmax_fused.template online_softmax_with_quant</*Is_first=*/true>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
|
||||
|
||||
consumer_wait(pipeline_v, smem_pipe_read_v);
|
||||
copy_v_block(_0{});
|
||||
quantize(_0{}, tSrS_converion_view);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
|
||||
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
|
||||
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO_store);
|
||||
if (v_block < size<2>(tOrP) - 1) {
|
||||
copy_v_block(v_block + 1);
|
||||
quantize(v_block + 1, tSrS_converion_view);
|
||||
} else {
|
||||
pipeline_v.consumer_release(smem_pipe_read_v);
|
||||
++smem_pipe_read_v;
|
||||
}
|
||||
}
|
||||
|
||||
n_block--;
|
||||
constexpr int n_masking_steps = !Is_causal ? 1 : cute::ceil_div(kBlockM, kBlockN) + 1;
|
||||
// // Only go through these if Is_causal, since n_masking_steps = 1 when !Is_causal
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int masking_step = 0; masking_step < n_masking_steps - 1 && n_block >= 0; ++masking_step, --n_block) {
|
||||
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
|
||||
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
|
||||
consumer_wait(pipeline_k, smem_pipe_read_k);
|
||||
copy_k_block(_0{});
|
||||
add_delta_s(tSrS);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
|
||||
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
|
||||
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
|
||||
if (k_block < size<2>(tSrQ) - 1) {
|
||||
copy_k_block(k_block + 1);
|
||||
}
|
||||
}
|
||||
pipeline_k.consumer_release(smem_pipe_read_k); // release K
|
||||
++smem_pipe_read_k;
|
||||
Tensor cS = cute::make_identity_tensor(select<0, 1>(TileShape_MNK{}));
|
||||
Tensor tScS = thread_mma_qk.partition_C(cS);
|
||||
#pragma unroll
|
||||
for (int i = 0; i < size(tSrS); ++i) {
|
||||
if (int(get<1>(tScS(i))) >= col_limit_causal(int(get<0>(tScS(i))), n_block)) {
|
||||
tSrS(i) = -INFINITY;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
softmax_fused.template online_softmax_with_quant</*Is_first=*/false>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
|
||||
Tensor tOrO = make_fragment_like(tOrO_store);
|
||||
consumer_wait(pipeline_v, smem_pipe_read_v);
|
||||
copy_v_block(_0{});
|
||||
quantize(_0{}, tSrS_converion_view);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
|
||||
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
|
||||
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO);
|
||||
if (v_block < size<2>(tOrP) - 1) {
|
||||
copy_v_block(v_block + 1);
|
||||
quantize(v_block + 1, tSrS_converion_view);
|
||||
}
|
||||
}
|
||||
pipeline_v.consumer_release(smem_pipe_read_v);
|
||||
++smem_pipe_read_v;
|
||||
if (masking_step > 0) { softmax_fused.rescale_o(tOrO_store, tOrO); }
|
||||
}
|
||||
|
||||
#pragma unroll 1
|
||||
for (; n_block >= 0; --n_block) {
|
||||
Tensor tSrS = partition_fragment_C(tiled_mma_qk, select<0, 1>(TileShape_MNK{}));
|
||||
Tensor tSrS_converion_view = make_tensor(tSrS.data(), flash::convert_to_conversion_layout(tSrS.layout()));
|
||||
consumer_wait(pipeline_k, smem_pipe_read_k);
|
||||
copy_k_block(_0{});
|
||||
add_delta_s(tSrS);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int k_block = 0; k_block < size<2>(tSrQ); ++k_block) {
|
||||
cute::gemm(tiled_mma_qk, make_zip_tensor(tSrQ(_, _, k_block), tSrSFQ(_, _, k_block)),
|
||||
make_zip_tensor(tSrK(_, _, k_block), tSrSFK(_, _, k_block)), tSrS);
|
||||
if (k_block < size<2>(tSrQ) - 1) {
|
||||
copy_k_block(k_block + 1);
|
||||
} else {
|
||||
pipeline_k.consumer_release(smem_pipe_read_k);
|
||||
++smem_pipe_read_k;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
softmax_fused.template online_softmax_with_quant</*Is_first=*/false>(tSrS, AbsMaxP, mainloop_params.softmax_scale_log2);
|
||||
Tensor tOrO = make_fragment_like(tOrO_store);
|
||||
consumer_wait(pipeline_v, smem_pipe_read_v);
|
||||
copy_v_block(_0{});
|
||||
quantize(_0{}, tSrS_converion_view);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int v_block = 0; v_block < size<2>(tOrP); ++v_block) {
|
||||
cute::gemm(tiled_mma_pv, make_zip_tensor(tOrP(_, _, v_block), tOrSFP(_, _, v_block)),
|
||||
make_zip_tensor(tOrVt(_, _, v_block), tOrSFVt(_, _, v_block)), tOrO);
|
||||
if (v_block < size<2>(tOrP) - 1) {
|
||||
copy_v_block(v_block + 1);
|
||||
quantize(v_block + 1, tSrS_converion_view);
|
||||
} else {
|
||||
pipeline_v.consumer_release(smem_pipe_read_v);
|
||||
++smem_pipe_read_v;
|
||||
}
|
||||
}
|
||||
softmax_fused.rescale_o(tOrO_store, tOrO);
|
||||
}
|
||||
softmax_fused.finalize(tOrO_store);
|
||||
return;
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
} // namespace flash
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
#pragma once
|
||||
|
||||
#include "cutlass/arch/barrier.h"
|
||||
#include "cutlass/pipeline/sm90_pipeline.hpp"
|
||||
|
||||
namespace flash {
|
||||
|
||||
enum class FP4NamedBarriers {
|
||||
QueryEmpty = 1,
|
||||
WarpSpecializedConsumer = 2,
|
||||
WarpSpecializedPingPongConsumer1 = 3,
|
||||
WarpSpecializedPingPongConsumer2 = 4,
|
||||
ProducerEnd = 5,
|
||||
ConsumerEnd = 6,
|
||||
EpilogueBarrier = 7
|
||||
};
|
||||
|
||||
template<int SequenceDepth, int SequenceLength>
|
||||
struct OrderedSequenceBarrierVarGroupSizeSharedStorage {
|
||||
using Barrier = cutlass::arch::ClusterBarrier;
|
||||
Barrier barrier_[SequenceDepth][SequenceLength];
|
||||
};
|
||||
|
||||
template<int SequenceDepth_, int SequenceLength_>
|
||||
class OrderedSequenceBarrierVarGroupSize {
|
||||
public:
|
||||
static constexpr int SequenceDepth = SequenceDepth_;
|
||||
static constexpr int SequenceLength = SequenceLength_;
|
||||
using Barrier = cutlass::arch::ClusterBarrier;
|
||||
using SharedStorage = flash::OrderedSequenceBarrierVarGroupSizeSharedStorage<SequenceDepth, SequenceLength>;
|
||||
|
||||
|
||||
struct Params {
|
||||
uint32_t group_id;
|
||||
uint32_t* group_size_list;
|
||||
};
|
||||
|
||||
private :
|
||||
// In future this Params object can be replaced easily with a CG object
|
||||
Params params_;
|
||||
Barrier *barrier_ptr_;
|
||||
cutlass::PipelineState<SequenceDepth> stage_;
|
||||
|
||||
static constexpr int Depth = SequenceDepth;
|
||||
static constexpr int Length = SequenceLength;
|
||||
|
||||
public:
|
||||
OrderedSequenceBarrierVarGroupSize() = delete;
|
||||
OrderedSequenceBarrierVarGroupSize(const OrderedSequenceBarrierVarGroupSize&) = delete;
|
||||
OrderedSequenceBarrierVarGroupSize(OrderedSequenceBarrierVarGroupSize&&) = delete;
|
||||
OrderedSequenceBarrierVarGroupSize& operator=(const OrderedSequenceBarrierVarGroupSize&) = delete;
|
||||
OrderedSequenceBarrierVarGroupSize& operator=(OrderedSequenceBarrierVarGroupSize&&) = delete;
|
||||
~OrderedSequenceBarrierVarGroupSize() = default;
|
||||
|
||||
CUTLASS_DEVICE
|
||||
OrderedSequenceBarrierVarGroupSize(SharedStorage& storage, Params const& params) :
|
||||
params_(params),
|
||||
barrier_ptr_(&storage.barrier_[0][0]),
|
||||
// Group 0 - starts with an opposite phase
|
||||
stage_({0, params.group_id == 0, 0}) {
|
||||
int warp_idx = cutlass::canonical_warp_idx_sync();
|
||||
int lane_predicate = cute::elect_one_sync();
|
||||
|
||||
// Barrier FULL, EMPTY init
|
||||
// Init is done only by the one elected thread of the block
|
||||
if (warp_idx == 0 && lane_predicate) {
|
||||
for (int d = 0; d < Depth; ++d) {
|
||||
for (int l = 0; l < Length; ++l) {
|
||||
barrier_ptr_[d * Length + l].init(*(params.group_size_list + l));
|
||||
}
|
||||
}
|
||||
}
|
||||
cutlass::arch::fence_barrier_init();
|
||||
}
|
||||
|
||||
// Wait on a stage to be unlocked
|
||||
CUTLASS_DEVICE
|
||||
void wait() {
|
||||
get_barrier_for_current_stage(params_.group_id).wait(stage_.phase());
|
||||
}
|
||||
|
||||
// Signal completion of Stage and move to the next stage
|
||||
// (group_id) signals to (group_id+1)
|
||||
CUTLASS_DEVICE
|
||||
void arrive() {
|
||||
int signalling_id = (params_.group_id + 1) % Length;
|
||||
get_barrier_for_current_stage(signalling_id).arrive();
|
||||
++stage_;
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
void advance() {
|
||||
++stage_;
|
||||
}
|
||||
|
||||
private:
|
||||
|
||||
CUTLASS_DEVICE
|
||||
Barrier& get_barrier_for_current_stage(int group_id) {
|
||||
return barrier_ptr_[stage_.index() * Length + group_id];
|
||||
}
|
||||
};
|
||||
|
||||
} // flash
|
||||
@@ -0,0 +1,180 @@
|
||||
// Modified from the original SageAttention3 code
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <cuda.h>
|
||||
#include <vector>
|
||||
|
||||
#ifdef OLD_GENERATOR_PATH
|
||||
#include <ATen/CUDAGeneratorImpl.h>
|
||||
#else
|
||||
#include <ATen/cuda/CUDAGeneratorImpl.h>
|
||||
#endif
|
||||
|
||||
#include <ATen/cuda/CUDAGraphsUtils.cuh> // For at::cuda::philox::unpack
|
||||
|
||||
#include "cutlass/fast_math.h" // For cutlass::FastDivmod
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
struct Qkv_params {
|
||||
using index_t = int64_t;
|
||||
// The QKV matrices.
|
||||
void *__restrict__ q_ptr;
|
||||
void *__restrict__ k_ptr;
|
||||
void *__restrict__ v_ptr;
|
||||
void *__restrict__ delta_s_ptr;
|
||||
// The QKV scale factor matrices.
|
||||
void *__restrict__ sfq_ptr;
|
||||
void *__restrict__ sfk_ptr;
|
||||
void *__restrict__ sfv_ptr;
|
||||
// The stride between rows of the Q, K and V matrices.
|
||||
index_t q_batch_stride;
|
||||
index_t k_batch_stride;
|
||||
index_t v_batch_stride;
|
||||
index_t q_row_stride;
|
||||
index_t k_row_stride;
|
||||
index_t v_row_stride;
|
||||
index_t q_head_stride;
|
||||
index_t k_head_stride;
|
||||
index_t v_head_stride;
|
||||
index_t ds_batch_stride;
|
||||
index_t ds_row_stride;
|
||||
index_t ds_head_stride;
|
||||
// The stride of the Q, K and V scale factor matrices.
|
||||
index_t sfq_batch_stride;
|
||||
index_t sfk_batch_stride;
|
||||
index_t sfv_batch_stride;
|
||||
index_t sfq_row_stride;
|
||||
index_t sfk_row_stride;
|
||||
index_t sfv_row_stride;
|
||||
index_t sfq_head_stride;
|
||||
index_t sfk_head_stride;
|
||||
index_t sfv_head_stride;
|
||||
|
||||
// The number of heads.
|
||||
int h, h_k;
|
||||
// In the case of multi-query and grouped-query attention (MQA/GQA), nheads_k could be
|
||||
// different from nheads (query).
|
||||
int h_h_k_ratio; // precompute h / h_k,
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
struct Flash_fwd_params : public Qkv_params {
|
||||
|
||||
// The O matrix (output).
|
||||
void * __restrict__ o_ptr;
|
||||
void * __restrict__ oaccum_ptr;
|
||||
void * __restrict__ s_ptr;
|
||||
|
||||
// The stride between rows of O.
|
||||
index_t o_batch_stride;
|
||||
index_t o_row_stride;
|
||||
index_t o_head_stride;
|
||||
|
||||
// The pointer to the P matrix.
|
||||
void * __restrict__ p_ptr;
|
||||
|
||||
// The pointer to the softmax sum.
|
||||
void * __restrict__ softmax_lse_ptr;
|
||||
void * __restrict__ softmax_lseaccum_ptr;
|
||||
|
||||
// The dimensions.
|
||||
int b, seqlen_q, seqlen_k, seqlen_knew, d, seqlen_q_rounded, seqlen_k_rounded, d_rounded, rotary_dim, unpadded_seqlen_k;
|
||||
cutlass::FastDivmod head_divmod, m_block_divmod;
|
||||
int total_blocks;
|
||||
int seqlen_s;
|
||||
|
||||
// The scaling factors for the kernel.
|
||||
float scale_softmax;
|
||||
float scale_softmax_log2;
|
||||
uint32_t scale_softmax_log2_half2;
|
||||
|
||||
// array of length b+1 holding starting offset of each sequence.
|
||||
int * __restrict__ cu_seqlens_q;
|
||||
int * __restrict__ cu_seqlens_k;
|
||||
|
||||
// If provided, the actual length of each k sequence.
|
||||
int * __restrict__ seqused_k;
|
||||
|
||||
int *__restrict__ blockmask;
|
||||
|
||||
// The K_new and V_new matrices.
|
||||
void * __restrict__ knew_ptr;
|
||||
void * __restrict__ vnew_ptr;
|
||||
|
||||
// The stride between rows of the Q, K and V matrices.
|
||||
index_t knew_batch_stride;
|
||||
index_t vnew_batch_stride;
|
||||
index_t knew_row_stride;
|
||||
index_t vnew_row_stride;
|
||||
index_t knew_head_stride;
|
||||
index_t vnew_head_stride;
|
||||
|
||||
// The cos and sin matrices for rotary embedding.
|
||||
void * __restrict__ rotary_cos_ptr;
|
||||
void * __restrict__ rotary_sin_ptr;
|
||||
|
||||
// The indices to index into the KV cache.
|
||||
int * __restrict__ cache_batch_idx;
|
||||
|
||||
// Paged KV cache
|
||||
int * __restrict__ block_table;
|
||||
index_t block_table_batch_stride;
|
||||
int page_block_size;
|
||||
|
||||
// The dropout probability (probability of keeping an activation).
|
||||
float p_dropout;
|
||||
// uint32_t p_dropout_in_uint;
|
||||
// uint16_t p_dropout_in_uint16_t;
|
||||
uint8_t p_dropout_in_uint8_t;
|
||||
|
||||
// Scale factor of 1 / (1 - p_dropout).
|
||||
float rp_dropout;
|
||||
float scale_softmax_rp_dropout;
|
||||
|
||||
// Local window size
|
||||
int window_size_left, window_size_right;
|
||||
|
||||
// Random state.
|
||||
at::PhiloxCudaState philox_args;
|
||||
|
||||
// Pointer to the RNG seed (idx 0) and offset (idx 1).
|
||||
uint64_t * rng_state;
|
||||
|
||||
bool is_bf16;
|
||||
bool is_e4m3;
|
||||
bool is_causal;
|
||||
bool per_block_mean;
|
||||
bool single_level_p_quant; // If true, use single-level 1x16 block scale quantization for P (like V), instead of two-level quantization
|
||||
// If is_seqlens_k_cumulative, then seqlen_k is cu_seqlens_k[bidb + 1] - cu_seqlens_k[bidb].
|
||||
// Otherwise it's cu_seqlens_k[bidb], i.e., we use cu_seqlens_k to store the sequence lengths of K.
|
||||
bool is_seqlens_k_cumulative;
|
||||
|
||||
bool is_rotary_interleaved;
|
||||
|
||||
int num_splits; // For split-KV version
|
||||
|
||||
void * __restrict__ alibi_slopes_ptr;
|
||||
index_t alibi_slopes_batch_stride;
|
||||
|
||||
int * __restrict__ tile_count_semaphore;
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
@@ -0,0 +1,190 @@
|
||||
// Modified from the original SageAttention3 code
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
#pragma once
|
||||
|
||||
#include <cmath>
|
||||
#include "cute/tensor.hpp"
|
||||
#include "cutlass/numeric_types.h"
|
||||
#include "utils.h"
|
||||
|
||||
namespace flash {
|
||||
|
||||
using namespace cute;
|
||||
|
||||
template <int Rows>
|
||||
struct SoftmaxFused{
|
||||
|
||||
using TensorT = decltype(make_fragment_like<float>(Shape<Int<Rows>>{}));
|
||||
TensorT row_sum, row_max, scores_scale;
|
||||
static constexpr float fp8_scalexfp4_scale = 1.f / (448 * 6);
|
||||
static constexpr float fp8_scalexfp4_scale_log2 = -11.392317422778762f; //log2f(fp8_scalexfp4_scale)
|
||||
static constexpr float fp4_scale_log2 = -2.584962500721156f; // log2f(fp4_scale)
|
||||
static constexpr int RowReductionThr = 4;
|
||||
|
||||
// If true, use single-level quantization: s_P2, P̂_2 = φ(P̃) directly (standard per-block FP4 quantization like V)
|
||||
// If false (default), use two-level quantization: s_P1 = rowmax(P̃)/(448×6), then s_P2, P̂_2 = φ(P̃/s_P1)
|
||||
bool single_level_p_quant;
|
||||
|
||||
CUTLASS_DEVICE SoftmaxFused(bool single_level = false) : single_level_p_quant(single_level) {};
|
||||
|
||||
template<bool FirstTile, bool InfCheck = false, typename TensorAcc, typename TensorMax>
|
||||
CUTLASS_DEVICE auto online_softmax_with_quant(
|
||||
TensorAcc& acc,
|
||||
TensorMax& AbsMaxP,
|
||||
const float softmax_scale_log2
|
||||
) {
|
||||
Tensor acc_reduction_view = make_tensor(acc.data(), flash::convert_to_reduction_layout(acc.layout()));
|
||||
Tensor acc_conversion_view = make_tensor(acc.data(), flash::convert_to_conversion_layout(acc.layout()));
|
||||
Tensor acc_conversion_flatten = group_modes<1, 5>(group_modes<0, 2>(flatten(acc_conversion_view)));
|
||||
|
||||
if constexpr (FirstTile) {
|
||||
fill(row_max, -INFINITY);
|
||||
clear(row_sum);
|
||||
fill(scores_scale, 1.f);
|
||||
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1, 1>(acc_reduction_view); ni++) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ei = 0; ei < size<1, 0>(acc_reduction_view); ei++) {
|
||||
AbsMaxP(mi, ni) = fmaxf(AbsMaxP(mi, ni), acc_reduction_view(mi, make_coord(ei, ni)));
|
||||
}
|
||||
float max_recv = __shfl_xor_sync(int32_t(-1), AbsMaxP(mi, ni), 1); // exchange max with neighbour thread of 8 elements
|
||||
AbsMaxP(mi, ni) = fmaxf(AbsMaxP(mi, ni), max_recv);
|
||||
row_max(mi) = fmaxf(row_max(mi), AbsMaxP(mi, ni));
|
||||
}
|
||||
|
||||
float max_recv = __shfl_xor_sync(int32_t(-1), row_max(mi), 2); // exchange max in a quad in a row
|
||||
row_max(mi) = fmaxf(row_max(mi), max_recv);
|
||||
|
||||
// Two-level P quantization (default): s_P1 = rowmax(P̃)/(448×6), then s_P2,P̂_2 = φ(P̃/s_P1)
|
||||
// - Pre-scales P to [0, 448×6] range before φ, output scaled by s_P1
|
||||
// Single-level P quantization: s_P2, P̂_2 = φ(P̃) directly (like V quantization)
|
||||
// - No s_P1, just standard per-block FP4 quantization φ
|
||||
const float s_P1_offset = single_level_p_quant ? 0.f : fp8_scalexfp4_scale_log2;
|
||||
const float max_scaled = InfCheck
|
||||
? (row_max(mi) == -INFINITY ? 0.f : (row_max(mi) * softmax_scale_log2 + s_P1_offset))
|
||||
: (row_max(mi) * softmax_scale_log2 + s_P1_offset);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
|
||||
acc_reduction_view(mi, ni) = flash::ptx_exp2(acc_reduction_view(mi, ni) * softmax_scale_log2 - max_scaled);
|
||||
}
|
||||
// s_P2 = max(P_block)/6 — per-block scale factor from φ function (same formula for both modes)
|
||||
// The difference is in max_scaled: two-level includes 448×6 pre-scaling, single-level doesn't
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int sfi = 0; sfi < size<1>(AbsMaxP); sfi++) {
|
||||
AbsMaxP(mi, sfi) = flash::ptx_exp2(AbsMaxP(mi, sfi) * softmax_scale_log2 - max_scaled + fp4_scale_log2);
|
||||
}
|
||||
}
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
|
||||
row_sum(mi) += acc_reduction_view(mi, ni);
|
||||
}
|
||||
}
|
||||
}
|
||||
else {
|
||||
Tensor scores_max_prev = make_fragment_like(row_max);
|
||||
cute::copy(row_max, scores_max_prev);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mi = 0; mi < size<0>(acc_reduction_view); mi++) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1, 1>(acc_reduction_view); ni++) {
|
||||
float local_max = -INFINITY;
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ei = 0; ei < size<1, 0>(acc_reduction_view); ei++) {
|
||||
local_max = fmaxf(local_max, acc_reduction_view(mi, make_coord(ei, ni)));
|
||||
}
|
||||
float max_recv = __shfl_xor_sync(int32_t(-1), local_max, 1); // exchange max with neighbour thread of 8 elements
|
||||
AbsMaxP(mi, ni) = fmaxf(local_max, max_recv);
|
||||
row_max(mi) = fmaxf(row_max(mi), AbsMaxP(mi, ni));
|
||||
}
|
||||
|
||||
float max_recv = __shfl_xor_sync(int32_t(-1), row_max(mi), 2); // exchange max in a quad in a row
|
||||
row_max(mi) = fmaxf(row_max(mi), max_recv);
|
||||
|
||||
float scores_max_cur = !InfCheck
|
||||
? row_max(mi)
|
||||
: (row_max(mi) == -INFINITY ? 0.0f : row_max(mi));
|
||||
scores_scale(mi) = flash::ptx_exp2((scores_max_prev(mi) - scores_max_cur) * softmax_scale_log2);
|
||||
|
||||
// Two-level P quantization (default): s_P1 = rowmax(P̃)/(448×6), then s_P2,P̂_2 = φ(P̃/s_P1)
|
||||
// Single-level P quantization: s_P2, P̂_2 = φ(P̃) directly (like V quantization)
|
||||
const float s_P1_offset = single_level_p_quant ? 0.f : fp8_scalexfp4_scale_log2;
|
||||
const float max_scaled = InfCheck
|
||||
? (row_max(mi) == -INFINITY ? 0.f : (row_max(mi) * softmax_scale_log2 + s_P1_offset))
|
||||
: (row_max(mi) * softmax_scale_log2 + s_P1_offset);
|
||||
row_sum(mi) = row_sum(mi) * scores_scale(mi);
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1>(acc_reduction_view); ni++) {
|
||||
acc_reduction_view(mi, ni) = flash::ptx_exp2(acc_reduction_view(mi, ni) * softmax_scale_log2 - max_scaled);
|
||||
row_sum(mi) += acc_reduction_view(mi, ni);
|
||||
}
|
||||
// s_P2 = max(P_block)/6 — per-block scale factor from φ function
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int sfi = 0; sfi < size<1>(AbsMaxP); sfi++) {
|
||||
AbsMaxP(mi, sfi) = flash::ptx_exp2(AbsMaxP(mi, sfi) * softmax_scale_log2 - max_scaled + fp4_scale_log2);
|
||||
}
|
||||
// scores_scale(mi) = max_scaled;
|
||||
}
|
||||
}
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int i = 0; i < size(AbsMaxP); ++i) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int j = 0; j < size<0>(acc_conversion_flatten); ++j)
|
||||
acc_conversion_flatten(j, i) /= AbsMaxP(i);
|
||||
}
|
||||
}
|
||||
|
||||
template<typename TensorAcc>
|
||||
CUTLASS_DEVICE void finalize(TensorAcc& o_store) {
|
||||
Tensor o_store_reduction_view = make_tensor(o_store.data(), flash::convert_to_reduction_layout(o_store.layout()));
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mi = 0; mi < size(row_max); ++mi) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int i = 1; i < RowReductionThr; i <<= 1) {
|
||||
float sum_recv = __shfl_xor_sync(int32_t(-1), row_sum(mi), i);
|
||||
row_sum(mi) += sum_recv;
|
||||
}
|
||||
float sum = row_sum(mi);
|
||||
float inv_sum = (sum == 0.f || sum != sum) ? 0.f : 1 / sum;
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1>(o_store_reduction_view); ++ni) {
|
||||
o_store_reduction_view(mi, ni) *= inv_sum;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
template<typename TensorAcc>
|
||||
CUTLASS_DEVICE void rescale_o(TensorAcc& o_store, TensorAcc const& o_tmp) {
|
||||
Tensor o_store_reduction_view = make_tensor(o_store.data(), flash::convert_to_reduction_layout(o_store.layout()));
|
||||
Tensor o_tmp_reduction_view = make_tensor(o_tmp.data(), flash::convert_to_reduction_layout(o_tmp.layout()));
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int mi = 0; mi < size(row_max); ++mi) {
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int ni = 0; ni < size<1>(o_store_reduction_view); ++ni) {
|
||||
o_store_reduction_view(mi, ni) = o_store_reduction_view(mi, ni) * scores_scale(mi) + o_tmp_reduction_view(mi, ni);
|
||||
}
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
|
||||
};
|
||||
} // namespace flash
|
||||
@@ -0,0 +1,83 @@
|
||||
// Inspired by
|
||||
// https://github.com/NVIDIA/DALI/blob/main/include/dali/core/static_switch.h
|
||||
// and https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/Dispatch.h
|
||||
|
||||
#pragma once
|
||||
|
||||
/// @param COND - a boolean expression to switch by
|
||||
/// @param CONST_NAME - a name given for the constexpr bool variable.
|
||||
/// @param ... - code to execute for true and false
|
||||
///
|
||||
/// Usage:
|
||||
/// ```
|
||||
/// BOOL_SWITCH(flag, BoolConst, [&] {
|
||||
/// some_function<BoolConst>(...);
|
||||
/// });
|
||||
/// ```
|
||||
//
|
||||
|
||||
#define BOOL_SWITCH(COND, CONST_NAME, ...) \
|
||||
[&] { \
|
||||
if (COND) { \
|
||||
constexpr static bool CONST_NAME = true; \
|
||||
return __VA_ARGS__(); \
|
||||
} else { \
|
||||
constexpr static bool CONST_NAME = false; \
|
||||
return __VA_ARGS__(); \
|
||||
} \
|
||||
}()
|
||||
|
||||
#define PREC_SWITCH(PRECTYPE, ...) \
|
||||
[&] { \
|
||||
if (PRECTYPE == 1) { \
|
||||
using kPrecType = cutlass::half_t; \
|
||||
constexpr static bool kSoftFp16 = false; \
|
||||
constexpr static bool kHybrid = false; \
|
||||
return __VA_ARGS__(); \
|
||||
} else if (PRECTYPE == 2) { \
|
||||
using kPrecType = cutlass::float_e4m3_t; \
|
||||
constexpr static bool kSoftFp16 = false; \
|
||||
constexpr static bool kHybrid = false; \
|
||||
return __VA_ARGS__(); \
|
||||
} else if (PRECTYPE == 3) { \
|
||||
using kPrecType = cutlass::float_e4m3_t; \
|
||||
constexpr static bool kSoftFp16 = false; \
|
||||
constexpr static bool kHybrid = true; \
|
||||
return __VA_ARGS__(); \
|
||||
} else if (PRECTYPE == 4) { \
|
||||
using kPrecType = cutlass::float_e4m3_t; \
|
||||
constexpr static bool kSoftFp16 = true; \
|
||||
constexpr static bool kHybrid = false; \
|
||||
return __VA_ARGS__(); \
|
||||
} \
|
||||
}()
|
||||
|
||||
#define HEADDIM_SWITCH(HEADDIM, ...) \
|
||||
[&] { \
|
||||
if (HEADDIM == 64) { \
|
||||
constexpr static int kHeadSize = 64; \
|
||||
return __VA_ARGS__(); \
|
||||
} else if (HEADDIM == 128) { \
|
||||
constexpr static int kHeadSize = 128; \
|
||||
return __VA_ARGS__(); \
|
||||
} else if (HEADDIM == 256) { \
|
||||
constexpr static int kHeadSize = 256; \
|
||||
return __VA_ARGS__(); \
|
||||
} \
|
||||
}()
|
||||
|
||||
#define SEQLEN_SWITCH(USE_VAR_SEQ_LEN, SEQ_LEN_OUT_OF_BOUND_CHECK, ...) \
|
||||
[&] { \
|
||||
if (!USE_VAR_SEQ_LEN) { \
|
||||
if (SEQ_LEN_OUT_OF_BOUND_CHECK) { \
|
||||
using kSeqLenTraitsType = FixedSeqLenTraits<true>; \
|
||||
return __VA_ARGS__(); \
|
||||
} else { \
|
||||
using kSeqLenTraitsType = FixedSeqLenTraits<false>; \
|
||||
return __VA_ARGS__(); \
|
||||
} \
|
||||
} else { \
|
||||
using kSeqLenTraitsType = VarSeqLenTraits; \
|
||||
return __VA_ARGS__(); \
|
||||
} \
|
||||
}()
|
||||
@@ -0,0 +1,304 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* This code is based on code from FlashAttention3, https://github.com/Dao-AILab/flash-attention
|
||||
* Copyright (c) 2024, Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao.
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "cutlass/fast_math.h"
|
||||
|
||||
namespace flash {
|
||||
|
||||
///////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
class StaticPersistentTileSchedulerOld {
|
||||
//
|
||||
// Data members
|
||||
//
|
||||
|
||||
private:
|
||||
int current_work_linear_idx_;
|
||||
cutlass::FastDivmod const &m_block_divmod, &head_divmod;
|
||||
int const total_blocks;
|
||||
|
||||
public:
|
||||
struct WorkTileInfo {
|
||||
int M_idx = 0;
|
||||
int H_idx = 0;
|
||||
int B_idx = 0;
|
||||
bool is_valid_tile = false;
|
||||
|
||||
CUTLASS_HOST_DEVICE
|
||||
bool
|
||||
is_valid() const {
|
||||
return is_valid_tile;
|
||||
}
|
||||
|
||||
CUTLASS_HOST_DEVICE
|
||||
static WorkTileInfo
|
||||
invalid_work_tile() {
|
||||
return {-1, -1, -1, false};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
public:
|
||||
|
||||
CUTLASS_DEVICE explicit StaticPersistentTileSchedulerOld(cutlass::FastDivmod const &m_block_divmod_,
|
||||
cutlass::FastDivmod const &head_divmod_,
|
||||
int const total_blocks_) :
|
||||
m_block_divmod(m_block_divmod_), head_divmod(head_divmod_), total_blocks(total_blocks_) {
|
||||
|
||||
// MSVC requires protecting use of CUDA-specific nonstandard syntax,
|
||||
// like blockIdx and gridDim, with __CUDA_ARCH__.
|
||||
#if defined(__CUDA_ARCH__)
|
||||
// current_work_linear_idx_ = blockIdx.x + blockIdx.y * gridDim.x + blockIdx.z * gridDim.x * gridDim.y;
|
||||
current_work_linear_idx_ = blockIdx.x;
|
||||
#else
|
||||
CUTLASS_ASSERT(false && "This line should never be reached");
|
||||
#endif
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_current_work() const {
|
||||
return get_current_work_for_linear_idx(current_work_linear_idx_);
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_current_work_for_linear_idx(int linear_idx) const {
|
||||
if (linear_idx >= total_blocks) {
|
||||
return WorkTileInfo::invalid_work_tile();
|
||||
}
|
||||
|
||||
// Map worker's linear index into the CTA tiled problem shape to the corresponding MHB indices
|
||||
int M_idx, H_idx, B_idx;
|
||||
int quotient = m_block_divmod.divmod(M_idx, linear_idx);
|
||||
B_idx = head_divmod.divmod(H_idx, quotient);
|
||||
return {M_idx, H_idx, B_idx, true};
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
void
|
||||
// advance_to_next_work(int advance_count = 1) {
|
||||
advance_to_next_work() {
|
||||
// current_work_linear_idx_ += int(gridDim.x * gridDim.y * gridDim.z);
|
||||
current_work_linear_idx_ += int(gridDim.x);
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
fetch_next_work() {
|
||||
WorkTileInfo new_work_tile_info;
|
||||
advance_to_next_work();
|
||||
new_work_tile_info = get_current_work();
|
||||
return new_work_tile_info;
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
///////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
class SingleTileScheduler {
|
||||
|
||||
public:
|
||||
|
||||
// Host side kernel arguments
|
||||
struct Arguments {
|
||||
int const num_blocks_m, num_head, num_batch;
|
||||
int const* tile_count_semaphore = nullptr;
|
||||
};
|
||||
|
||||
// Device side kernel params
|
||||
struct Params {};
|
||||
|
||||
static Params
|
||||
to_underlying_arguments(Arguments const& args) {
|
||||
return {};
|
||||
}
|
||||
|
||||
static dim3
|
||||
get_grid_dim(Arguments const& args, int num_sm) {
|
||||
return {uint32_t(args.num_blocks_m), uint32_t(args.num_head), uint32_t(args.num_batch)};
|
||||
}
|
||||
|
||||
struct WorkTileInfo {
|
||||
int M_idx = 0;
|
||||
int H_idx = 0;
|
||||
int B_idx = 0;
|
||||
bool is_valid_tile = false;
|
||||
|
||||
CUTLASS_DEVICE
|
||||
bool
|
||||
is_valid(Params const& params) const {
|
||||
return is_valid_tile;
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
cute::tuple<int32_t, int32_t, int32_t>
|
||||
get_block_coord(Params const& params) const {
|
||||
return {M_idx, H_idx, B_idx};
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_next_work(Params const& params) const {
|
||||
return {-1, -1, -1, false};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_initial_work() const {
|
||||
return {int(blockIdx.x), int(blockIdx.y), int(blockIdx.z), true};
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
|
||||
return {-1, -1, -1, false};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
///////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
class StaticPersistentTileScheduler {
|
||||
|
||||
public:
|
||||
|
||||
// Host side kernel arguments
|
||||
struct Arguments {
|
||||
int const num_blocks_m, num_head, num_batch;
|
||||
int const* tile_count_semaphore = nullptr;
|
||||
};
|
||||
|
||||
// Device side kernel params
|
||||
struct Params {
|
||||
int total_blocks;
|
||||
cutlass::FastDivmod m_block_divmod, head_divmod;
|
||||
};
|
||||
|
||||
static Params
|
||||
to_underlying_arguments(Arguments const& args) {
|
||||
return {args.num_blocks_m * args.num_head * args.num_batch,
|
||||
cutlass::FastDivmod(args.num_blocks_m), cutlass::FastDivmod(args.num_head)};
|
||||
}
|
||||
|
||||
static dim3
|
||||
get_grid_dim(Arguments const& args, int num_sm) {
|
||||
return {uint32_t(num_sm)};
|
||||
}
|
||||
|
||||
struct WorkTileInfo {
|
||||
int tile_idx;
|
||||
|
||||
CUTLASS_DEVICE
|
||||
bool
|
||||
is_valid(Params const& params) const {
|
||||
return tile_idx < params.total_blocks;
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
cute::tuple<int32_t, int32_t, int32_t>
|
||||
get_block_coord(Params const& params) const {
|
||||
int m_block, bidh, bidb;
|
||||
bidb = params.head_divmod.divmod(bidh, params.m_block_divmod.divmod(m_block, tile_idx));
|
||||
return {m_block, bidh, bidb};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_initial_work() const {
|
||||
return {int(blockIdx.x)};
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
|
||||
return {current_work.tile_idx + int(gridDim.x)};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
class DynamicPersistentTileScheduler {
|
||||
|
||||
public:
|
||||
|
||||
// Host side kernel arguments
|
||||
struct Arguments {
|
||||
int const num_blocks_m, num_head, num_batch;
|
||||
int const* tile_count_semaphore;
|
||||
};
|
||||
|
||||
// Device side kernel params
|
||||
struct Params {
|
||||
int const total_blocks;
|
||||
cutlass::FastDivmod const m_block_divmod, head_divmod;
|
||||
int const* tile_count_semaphore;
|
||||
};
|
||||
|
||||
static Params
|
||||
to_underlying_arguments(Arguments const& args) {
|
||||
return {args.num_blocks_m * args.num_head * args.num_batch,
|
||||
cutlass::FastDivmod(args.num_blocks_m), cutlass::FastDivmod(args.num_head),
|
||||
args.tile_count_semaphore};
|
||||
}
|
||||
|
||||
static dim3
|
||||
get_grid_dim(Arguments const& args, int num_sm) {
|
||||
return {uint32_t(num_sm)};
|
||||
}
|
||||
|
||||
using WorkTileInfo = StaticPersistentTileScheduler::WorkTileInfo;
|
||||
// struct WorkTileInfo {
|
||||
// int tile_idx;
|
||||
|
||||
// CUTLASS_DEVICE
|
||||
// bool
|
||||
// is_valid(Params const& params) const {
|
||||
// return tile_idx < params.total_blocks;
|
||||
// }
|
||||
|
||||
// CUTLASS_DEVICE
|
||||
// cute::tuple<int32_t, int32_t, int32_t>
|
||||
// get_block_coord(Params const& params) const {
|
||||
// int m_block, bidh, bidb;
|
||||
// bidb = params.head_divmod.divmod(bidh, params.m_block_divmod.divmod(m_block, tile_idx));
|
||||
// return {m_block, bidh, bidb};
|
||||
// }
|
||||
|
||||
// };
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_initial_work() const {
|
||||
return {int(blockIdx.x)};
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE
|
||||
WorkTileInfo
|
||||
get_next_work(Params const& params, WorkTileInfo const& current_work) const {
|
||||
return {current_work.tile_idx + int(gridDim.x)};
|
||||
}
|
||||
|
||||
};
|
||||
|
||||
} // flash
|
||||
@@ -0,0 +1,408 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#pragma once
|
||||
|
||||
#include <assert.h>
|
||||
#include <stdint.h>
|
||||
#include <stdlib.h>
|
||||
|
||||
#include <cuda_fp16.h>
|
||||
|
||||
#if defined(__CUDA_ARCH__) && __CUDA_ARCH__ >= 800
|
||||
#include <cuda_bf16.h>
|
||||
#endif
|
||||
|
||||
#include <cute/tensor.hpp>
|
||||
|
||||
#include <cutlass/array.h>
|
||||
#include <cutlass/cutlass.h>
|
||||
#include <cutlass/numeric_conversion.h>
|
||||
#include <cutlass/numeric_types.h>
|
||||
|
||||
namespace flash {
|
||||
|
||||
using namespace cute;
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<typename T>
|
||||
struct MaxOp {
|
||||
__device__ __forceinline__ T operator()(T const & x, T const & y) { return x > y ? x : y; }
|
||||
};
|
||||
|
||||
template <>
|
||||
struct MaxOp<float> {
|
||||
// This is slightly faster
|
||||
__device__ __forceinline__ float operator()(float const &x, float const &y) { return max(x, y); }
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<typename T>
|
||||
struct SumOp {
|
||||
__device__ __forceinline__ T operator()(T const & x, T const & y) { return x + y; }
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<int THREADS>
|
||||
struct Allreduce {
|
||||
static_assert(THREADS == 32 || THREADS == 16 || THREADS == 8 || THREADS == 4);
|
||||
template<typename T, typename Operator>
|
||||
static __device__ __forceinline__ T run(T x, Operator &op) {
|
||||
constexpr int OFFSET = THREADS / 2;
|
||||
x = op(x, __shfl_xor_sync(uint32_t(-1), x, OFFSET));
|
||||
return Allreduce<OFFSET>::run(x, op);
|
||||
}
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<>
|
||||
struct Allreduce<2> {
|
||||
template<typename T, typename Operator>
|
||||
static __device__ __forceinline__ T run(T x, Operator &op) {
|
||||
x = op(x, __shfl_xor_sync(uint32_t(-1), x, 1));
|
||||
return x;
|
||||
}
|
||||
};
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
|
||||
__device__ __forceinline__ void thread_reduce_(Tensor<Engine0, Layout0> const &tensor, Tensor<Engine1, Layout1> &summary, Operator &op) {
|
||||
static_assert(Layout0::rank == 2, "Only support 2D Tensor");
|
||||
static_assert(Layout1::rank == 1, "Only support 1D Tensor");
|
||||
CUTE_STATIC_ASSERT_V(size<0>(summary) == size<0>(tensor));
|
||||
#pragma unroll
|
||||
for (int mi = 0; mi < size<0>(tensor); mi++) {
|
||||
summary(mi) = zero_init ? tensor(mi, 0) : op(summary(mi), tensor(mi, 0));
|
||||
#pragma unroll
|
||||
for (int ni = 1; ni < size<1>(tensor); ni++) {
|
||||
summary(mi) = op(summary(mi), tensor(mi, ni));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
template<typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
|
||||
__device__ __forceinline__ void quad_allreduce_(Tensor<Engine0, Layout0> &dst, Tensor<Engine1, Layout1> &src, Operator &op) {
|
||||
CUTE_STATIC_ASSERT_V(size(dst) == size(src));
|
||||
#pragma unroll
|
||||
for (int i = 0; i < size(dst); i++){
|
||||
dst(i) = Allreduce<4>::run(src(i), op);
|
||||
}
|
||||
}
|
||||
|
||||
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1, typename Operator>
|
||||
__device__ __forceinline__ void reduce_(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &summary, Operator &op) {
|
||||
thread_reduce_<zero_init>(tensor, summary, op);
|
||||
quad_allreduce_(summary, summary, op);
|
||||
}
|
||||
|
||||
template<bool zero_init=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
|
||||
__device__ __forceinline__ void reduce_max(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &max){
|
||||
MaxOp<float> max_op;
|
||||
reduce_<zero_init>(tensor, max, max_op);
|
||||
}
|
||||
|
||||
template<bool zero_init=true, bool warp_reduce=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
|
||||
__device__ __forceinline__ void reduce_sum(Tensor<Engine0, Layout0> const& tensor, Tensor<Engine1, Layout1> &sum){
|
||||
SumOp<float> sum_op;
|
||||
thread_reduce_<zero_init>(tensor, sum, sum_op);
|
||||
if constexpr (warp_reduce) { quad_allreduce_(sum, sum, sum_op); }
|
||||
}
|
||||
|
||||
__forceinline__ __device__ __half2 half_exp(__half2 x) {
|
||||
uint32_t tmp_out, tmp_in;
|
||||
tmp_in = reinterpret_cast<uint32_t&>(x);
|
||||
asm ("ex2.approx.f16x2 %0, %1;\n"
|
||||
: "=r"(tmp_out)
|
||||
: "r"(tmp_in));
|
||||
__half2 out = reinterpret_cast<__half2&>(tmp_out);
|
||||
return out;
|
||||
}
|
||||
|
||||
// Apply the exp to all the elements.
|
||||
template <bool zero_init=false, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
|
||||
__forceinline__ __device__ void max_scale_exp2_sum(Tensor<Engine0, Layout0> &tensor, Tensor<Engine1, Layout1> &max, Tensor<Engine1, Layout1> &sum, const float scale) {
|
||||
static_assert(Layout0::rank == 2, "Only support 2D Tensor"); static_assert(Layout1::rank == 1, "Only support 1D Tensor"); CUTE_STATIC_ASSERT_V(size<0>(max) == size<0>(tensor));
|
||||
#pragma unroll
|
||||
for (int mi = 0; mi < size<0>(tensor); ++mi) {
|
||||
MaxOp<float> max_op;
|
||||
max(mi) = zero_init ? tensor(mi, 0) : max_op(max(mi), tensor(mi, 0));
|
||||
#pragma unroll
|
||||
for (int ni = 1; ni < size<1>(tensor); ni++) {
|
||||
max(mi) = max_op(max(mi), tensor(mi, ni));
|
||||
}
|
||||
max(mi) = Allreduce<4>::run(max(mi), max_op);
|
||||
// If max is -inf, then all elements must have been -inf (possibly due to masking).
|
||||
// We don't want (-inf - (-inf)) since that would give NaN.
|
||||
const float max_scaled = max(mi) == -INFINITY ? 0.f : max(mi) * scale;
|
||||
sum(mi) = 0;
|
||||
#pragma unroll
|
||||
for (int ni = 0; ni < size<1>(tensor); ++ni) {
|
||||
// Instead of computing exp(x - max), we compute exp2(x * log_2(e) -
|
||||
// max * log_2(e)) This allows the compiler to use the ffma
|
||||
// instruction instead of fadd and fmul separately.
|
||||
tensor(mi, ni) = exp2f(tensor(mi, ni) * scale - max_scaled);
|
||||
sum(mi) += tensor(mi, ni);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Apply the exp to all the elements.
|
||||
template <bool Scale_max=true, bool Check_inf=true, typename Engine0, typename Layout0, typename Engine1, typename Layout1>
|
||||
__forceinline__ __device__ void scale_apply_exp2(Tensor<Engine0, Layout0> &tensor, Tensor<Engine1, Layout1> const &max, const float scale) {
|
||||
static_assert(Layout0::rank == 2, "Only support 2D Tensor");
|
||||
static_assert(Layout1::rank == 1, "Only support 1D Tensor");
|
||||
CUTE_STATIC_ASSERT_V(size<0>(max) == size<0>(tensor));
|
||||
#pragma unroll
|
||||
for (int mi = 0; mi < size<0>(tensor); ++mi) {
|
||||
// If max is -inf, then all elements must have been -inf (possibly due to masking).
|
||||
// We don't want (-inf - (-inf)) since that would give NaN.
|
||||
// If we don't have float around M_LOG2E the multiplication is done in fp64.
|
||||
const float max_scaled = Check_inf
|
||||
? (max(mi) == -INFINITY ? 0.f : (max(mi) * (Scale_max ? scale : float(M_LOG2E))))
|
||||
: (max(mi) * (Scale_max ? scale : float(M_LOG2E)));
|
||||
#pragma unroll
|
||||
for (int ni = 0; ni < size<1>(tensor); ++ni) {
|
||||
// Instead of computing exp(x - max), we compute exp2(x * log_2(e) -
|
||||
// max * log_2(e)) This allows the compiler to use the ffma
|
||||
// instruction instead of fadd and fmul separately.
|
||||
tensor(mi, ni) = exp2f(tensor(mi, ni) * scale - max_scaled);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
__forceinline__ __device__ float ptx_exp2(float x) {
|
||||
float y;
|
||||
asm volatile("ex2.approx.ftz.f32 %0, %1;" : "=f"(y) : "f"(x));
|
||||
return y;
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
packed_float_to_ue4m3(
|
||||
float const &f0, float const &f1, float const &f2, float const &f3,
|
||||
uint32_t &out
|
||||
) {
|
||||
asm volatile( \
|
||||
"{\n" \
|
||||
".reg .b16 lo;\n" \
|
||||
".reg .b16 hi;\n" \
|
||||
"cvt.rn.satfinite.e4m3x2.f32 lo, %2, %1;\n" \
|
||||
"cvt.rn.satfinite.e4m3x2.f32 hi, %4, %3;\n" \
|
||||
"mov.b32 %0, {lo, hi};\n" \
|
||||
"}" \
|
||||
: "=r"(out) : "f"(f0), "f"(f1), "f"(f2), "f"(f3));
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
packed_float_to_e2m1(
|
||||
float const &f0, float const &f1, float const &f2, float const& f3,
|
||||
float const &f4, float const &f5, float const &f6, float const& f7,
|
||||
uint32_t &out
|
||||
) {
|
||||
|
||||
asm volatile( \
|
||||
"{\n" \
|
||||
".reg .b8 byte0;\n" \
|
||||
".reg .b8 byte1;\n" \
|
||||
".reg .b8 byte2;\n" \
|
||||
".reg .b8 byte3;\n" \
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte0, %2, %1;\n" \
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte1, %4, %3;\n" \
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte2, %6, %5;\n" \
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte3, %8, %7;\n" \
|
||||
"mov.b32 %0, {byte0, byte1, byte2, byte3};\n" \
|
||||
"}" \
|
||||
: "=r"(out) : "f"(f0), "f"(f1), "f"(f2), "f"(f3),
|
||||
"f"(f4), "f"(f5), "f"(f6), "f"(f7));
|
||||
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
add(float2 & c,
|
||||
float2 const& a,
|
||||
float2 const& b)
|
||||
{
|
||||
asm volatile("add.f32x2 %0, %1, %2;\n"
|
||||
: "=l"(reinterpret_cast<uint64_t &>(c))
|
||||
: "l"(reinterpret_cast<uint64_t const&>(a)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(b)));
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
add_inplace(float2 &a,
|
||||
float2 const& b)
|
||||
{
|
||||
asm volatile("add.f32x2 %0, %0, %1;\n"
|
||||
: "+l"(reinterpret_cast<uint64_t &>(a)) // a: input/output
|
||||
: "l"(reinterpret_cast<uint64_t const&>(b)) // b: input
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
sub(float2 & c,
|
||||
float2 const& a,
|
||||
float2 const& b)
|
||||
{
|
||||
asm volatile("sub.f32x2 %0, %1, %2;\n"
|
||||
: "=l"(reinterpret_cast<uint64_t &>(c))
|
||||
: "l"(reinterpret_cast<uint64_t const&>(a)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(b)));
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
sub_inplace(float2 &a,
|
||||
float2 const& b)
|
||||
{
|
||||
asm volatile("sub.f32x2 %0, %0, %1;\n"
|
||||
: "+l"(reinterpret_cast<uint64_t &>(a)) // a: input/output
|
||||
: "l"(reinterpret_cast<uint64_t const&>(b)) // b: input
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
mul(float2 & c,
|
||||
float2 const& a,
|
||||
float2 const& b)
|
||||
{
|
||||
asm volatile("mul.f32x2 %0, %1, %2;\n"
|
||||
: "=l"(reinterpret_cast<uint64_t &>(c))
|
||||
: "l"(reinterpret_cast<uint64_t const&>(a)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(b)));
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
fma(float2 & d,
|
||||
float2 const& a,
|
||||
float2 const& b,
|
||||
float2 const& c)
|
||||
{
|
||||
asm volatile("fma.rn.f32x2 %0, %1, %2, %3;\n"
|
||||
: "=l"(reinterpret_cast<uint64_t &>(d))
|
||||
: "l"(reinterpret_cast<uint64_t const&>(a)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(b)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(c)));
|
||||
}
|
||||
|
||||
CUTLASS_DEVICE void
|
||||
fma_inplace(float2 &a,
|
||||
float2 const& b,
|
||||
float2 const& c)
|
||||
{
|
||||
asm volatile("fma.rn.f32x2 %0, %0, %1, %2;\n"
|
||||
: "+l"(reinterpret_cast<uint64_t &>(a))
|
||||
: "l"(reinterpret_cast<uint64_t const&>(b)),
|
||||
"l"(reinterpret_cast<uint64_t const&>(c)));
|
||||
}
|
||||
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template <
|
||||
class Layout
|
||||
>
|
||||
CUTLASS_DEVICE constexpr
|
||||
auto convert_to_reduction_layout(Layout mma_layout) {
|
||||
static_assert(rank(mma_layout) == 3, "Mma Layout should be (MmaAtom, MmaM, MmaN)");
|
||||
static_assert(rank(get<0>(shape(mma_layout))) == 2, "MmaAtom should be (AtomN, AtomM)");
|
||||
|
||||
return make_layout(
|
||||
make_layout(get<0,1>(mma_layout), get<1>(mma_layout)),
|
||||
make_layout(get<0,0>(mma_layout), get<2>(mma_layout))
|
||||
);
|
||||
}
|
||||
|
||||
template <
|
||||
class Tensor
|
||||
>
|
||||
CUTLASS_DEVICE constexpr
|
||||
auto convert_to_reduction_tensor(Tensor mma_tensor) {
|
||||
return make_tensor(mma_tensor.data(), convert_to_reduction_layout(mma_tensor.layout()));
|
||||
}
|
||||
|
||||
|
||||
template <
|
||||
class Layout
|
||||
>
|
||||
CUTLASS_DEVICE constexpr
|
||||
auto convert_to_conversion_layout(Layout mma_layout) {
|
||||
static_assert(rank(mma_layout) == 3, "Mma Layout should be (MmaAtom, MmaM, MmaN)");
|
||||
static_assert(rank(get<0>(shape(mma_layout))) == 2, "MmaAtom should be (AtomN, AtomM)");
|
||||
|
||||
constexpr int MmaAtomN = size<0, 0>(mma_layout);
|
||||
constexpr int MmaAtomM = size<0, 1>(mma_layout);
|
||||
constexpr int MmaM = size<1>(mma_layout);
|
||||
constexpr int MmaN = size<2>(mma_layout);
|
||||
|
||||
static_assert(MmaAtomN == 8, "MmaAtomN should be 8.");
|
||||
static_assert(MmaAtomM == 2, "MmaAtomM should be 2.");
|
||||
static_assert(MmaN % 2 == 0, "MmaN should be multiple of 2.");
|
||||
|
||||
auto mma_n_division = zipped_divide(
|
||||
layout<2>(mma_layout), make_tile(_2{})
|
||||
);
|
||||
return make_layout(
|
||||
make_layout(layout<0,0>(mma_layout), make_layout(layout<0,1>(mma_layout), layout<0>(mma_n_division))),
|
||||
layout<1>(mma_layout), layout<1>(mma_n_division)
|
||||
);
|
||||
}
|
||||
|
||||
template <
|
||||
class Tensor
|
||||
>
|
||||
CUTLASS_DEVICE constexpr
|
||||
auto convert_to_conversion_tensor(Tensor mma_tensor) {
|
||||
return make_tensor(mma_tensor.data(), convert_to_conversion_layout(mma_tensor.layout()));
|
||||
}
|
||||
////////////////////////////////////////////////////////////////////////////////////////////////////
|
||||
|
||||
template <bool Is_even_MN=true, bool Is_even_K=true, bool Clear_OOB_MN=false, bool Clear_OOB_K=true,
|
||||
typename TiledCopy, typename Engine0, typename Layout0, typename Engine1, typename Layout1,
|
||||
typename Engine2, typename Layout2, typename Engine3, typename Layout3>
|
||||
CUTLASS_DEVICE void copy(TiledCopy tiled_copy, Tensor<Engine0, Layout0> const &S,
|
||||
Tensor<Engine1, Layout1> &D, Tensor<Engine2, Layout2> const &identity_MN,
|
||||
Tensor<Engine3, Layout3> const &predicate_K, const int max_MN=0) {
|
||||
CUTE_STATIC_ASSERT_V(rank(S) == Int<3>{});
|
||||
CUTE_STATIC_ASSERT_V(rank(D) == Int<3>{});
|
||||
CUTE_STATIC_ASSERT_V(size<0>(S) == size<0>(D)); // MMA
|
||||
CUTE_STATIC_ASSERT_V(size<1>(S) == size<1>(D)); // MMA_M
|
||||
CUTE_STATIC_ASSERT_V(size<2>(S) == size<2>(D)); // MMA_K
|
||||
// There's no case where !Clear_OOB_K && Clear_OOB_MN
|
||||
static_assert(!(Clear_OOB_MN && !Clear_OOB_K));
|
||||
#pragma unroll
|
||||
for (int m = 0; m < size<1>(S); ++m) {
|
||||
if (Is_even_MN || get<0>(identity_MN(0, m, 0)) < max_MN) {
|
||||
#pragma unroll
|
||||
for (int k = 0; k < size<2>(S); ++k) {
|
||||
if (Is_even_K || predicate_K(k)) {
|
||||
cute::copy(tiled_copy, S(_, m, k), D(_, m, k));
|
||||
} else if (Clear_OOB_K) {
|
||||
cute::clear(D(_, m, k));
|
||||
}
|
||||
}
|
||||
} else if (Clear_OOB_MN) {
|
||||
cute::clear(D(_, m, _));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace flash
|
||||
@@ -0,0 +1 @@
|
||||
__version__ = "3.0.0.b1"
|
||||
@@ -0,0 +1,90 @@
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
import torch
|
||||
import fp4quant
|
||||
from triton.tools.mxfp import MXFP4Tensor
|
||||
|
||||
from bench_utils import bench_kineto
|
||||
b = 1
|
||||
h = 32
|
||||
n = 16384
|
||||
d = 128
|
||||
|
||||
def test():
|
||||
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
|
||||
o = torch.empty((b, h, n, d // 2), device="cuda", dtype=torch.uint8)
|
||||
o_s = torch.empty((b, h, n, d // 16), device="cuda", dtype=torch.float8_e4m3fn)
|
||||
fp4quant.scaled_fp4_quant_permute(q, o, o_s, 1)
|
||||
|
||||
test()
|
||||
|
||||
t = bench_kineto(test, "scaled_fp4_quant_kernel", suppress_kineto_output=True)
|
||||
|
||||
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
|
||||
throughput = IO / t * 1e-9
|
||||
|
||||
print(f"Throughput: {throughput:.2f} GB/s")
|
||||
|
||||
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
|
||||
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
|
||||
B, H, M, N = x.shape
|
||||
x = x.view(B, H, M, N // 16, 16)
|
||||
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
|
||||
if all_ones:
|
||||
scales = torch.ones_like(scales)
|
||||
x_scaled = x / scales
|
||||
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
|
||||
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
|
||||
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
|
||||
permuted_fp8_scale = None
|
||||
if permuted:
|
||||
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
|
||||
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
|
||||
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
|
||||
|
||||
b = 2
|
||||
h = 4
|
||||
n = 251
|
||||
n_padded = (n + 127) // 128 * 128
|
||||
d = 128
|
||||
|
||||
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
|
||||
o = torch.empty((b, h, n, d // 2), dtype=torch.uint8, device='cuda')
|
||||
o_s = torch.empty((b, h, n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
|
||||
|
||||
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
|
||||
|
||||
k_permute = [0, 1, 8, 9, 16, 17, 24, 25, 2, 3, 10, 11, 18, 19, 26, 27, 4, 5, 12, 13, 20, 21, 28, 29, 6, 7, 14, 15, 22, 23, 30, 31]
|
||||
o_permuted = torch.empty((b, h, n_padded, d // 2), dtype=torch.uint8, device='cuda')
|
||||
o_s_permuted = torch.empty((b, h, n_padded, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
|
||||
fp4quant.scaled_fp4_quant_permute(q, o_permuted, o_s_permuted, 1)
|
||||
|
||||
# padding
|
||||
if n % 128 != 0:
|
||||
o_permuted_gt = torch.cat([o, torch.zeros((b, h, n_padded - n, d // 2), dtype=torch.uint8, device='cuda')], dim=2)
|
||||
o_s_permuted_gt = torch.cat([o_s, torch.zeros((b, h, n_padded - n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')], dim=2)
|
||||
else:
|
||||
o_permuted_gt = o
|
||||
o_s_permuted_gt = o_s
|
||||
|
||||
# use scale_and_fp4_tensor + torch permutation to get the ground truth
|
||||
o_permuted_gt = o_permuted_gt.reshape(b, h, n_padded // 32, 32, d // 2)[:, :, :, k_permute, :].reshape(b, h, n_padded, d // 2)
|
||||
o_s_permuted_gt = o_s_permuted_gt.reshape(b, h, n_padded // 32, 32, d // 16)[:, :, :, k_permute, :].reshape(b, h, n_padded, d // 16)
|
||||
|
||||
assert((o_permuted - o_permuted_gt).abs().max() == 0)
|
||||
assert((o_s_permuted.float() - o_s_permuted_gt.float()).abs().max() == 0)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,86 @@
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
import torch
|
||||
import fp4quant
|
||||
from triton.tools.mxfp import MXFP4Tensor
|
||||
|
||||
from bench_utils import bench_kineto
|
||||
b = 1
|
||||
h = 32
|
||||
n = 16384
|
||||
d = 128
|
||||
|
||||
def test():
|
||||
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
|
||||
o = torch.empty((b, h, n, d // 2), device="cuda", dtype=torch.uint8)
|
||||
o_s = torch.empty((b, h, n, d // 16), device="cuda", dtype=torch.float8_e4m3fn)
|
||||
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
|
||||
|
||||
test()
|
||||
|
||||
t = bench_kineto(test, "scaled_fp4_quant_kernel", suppress_kineto_output=True)
|
||||
|
||||
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
|
||||
throughput = IO / t * 1e-9
|
||||
|
||||
print(f"Throughput: {throughput:.2f} GB/s")
|
||||
|
||||
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
|
||||
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
|
||||
B, H, M, N = x.shape
|
||||
x = x.view(B, H, M, N // 16, 16)
|
||||
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
|
||||
if all_ones:
|
||||
scales = torch.ones_like(scales)
|
||||
x_scaled = x / scales
|
||||
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
|
||||
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
|
||||
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
|
||||
permuted_fp8_scale = None
|
||||
if permuted:
|
||||
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
|
||||
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
|
||||
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
|
||||
|
||||
b = 2
|
||||
h = 4
|
||||
n = 251
|
||||
d = 128
|
||||
|
||||
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
|
||||
o = torch.empty((b, h, n, d // 2), dtype=torch.uint8, device='cuda')
|
||||
o_s = torch.empty((b, h, n, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
|
||||
|
||||
fp4quant.scaled_fp4_quant(q, o, o_s, 1)
|
||||
|
||||
fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale = scale_and_fp4_tensor(q, packed_dim=3)
|
||||
|
||||
assert((fp8_scale.float() - o_s.float()).abs().max() == 0)
|
||||
|
||||
o_binary = [
|
||||
(int(bin_str[:4], 2), int(bin_str[4:], 2))
|
||||
for bin_str in [format(x.item(), '08b') for x in o.view(-1)]
|
||||
]
|
||||
o_binary_gt = [
|
||||
(int(bin_str[:4], 2), int(bin_str[4:], 2))
|
||||
for bin_str in [format(x.item(), '08b') for x in packed_fp4.view(-1)]
|
||||
]
|
||||
for i in range(len(o_binary)):
|
||||
# check contiguous 4 bits. Difference should be at most one
|
||||
assert(abs(o_binary[i][0] - o_binary_gt[i][0]) <= 1)
|
||||
assert(abs(o_binary[i][1] - o_binary_gt[i][1]) <= 1)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,86 @@
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
import torch
|
||||
import fp4quant
|
||||
from triton.tools.mxfp import MXFP4Tensor
|
||||
|
||||
from bench_utils import bench_kineto
|
||||
b = 1
|
||||
h = 32
|
||||
n = 16384
|
||||
d = 128
|
||||
|
||||
def test():
|
||||
q = torch.randn((b, h, n, d), device="cuda", dtype=torch.float16)
|
||||
o = torch.empty((b, h, d, n // 2), device="cuda", dtype=torch.uint8)
|
||||
o_s = torch.empty((b, h, d, n // 16), device="cuda", dtype=torch.float8_e4m3fn)
|
||||
fp4quant.scaled_fp4_quant_trans(q, o, o_s, 1)
|
||||
|
||||
test()
|
||||
|
||||
t = bench_kineto(test, "scaled_fp4_quant_trans_kernel", suppress_kineto_output=True)
|
||||
|
||||
IO = b * h * n * d * 2 + b * h * n * d * 0.5 + b * h * n * d // 16 * 1
|
||||
throughput = IO / t * 1e-9
|
||||
|
||||
print(f"Throughput: {throughput:.2f} GB/s")
|
||||
|
||||
def scale_and_fp4_tensor(x: torch.Tensor, packed_dim: int = 3, all_ones: bool = False, permuted: bool = False):
|
||||
assert x.is_contiguous() and x.ndim == 4 and x.shape[-1] % 16 == 0
|
||||
B, H, M, N = x.shape
|
||||
x = x.view(B, H, M, N // 16, 16)
|
||||
scales = (x.abs().amax(dim=-1, keepdim=True) / 6).to(torch.float32)
|
||||
if all_ones:
|
||||
scales = torch.ones_like(scales)
|
||||
x_scaled = x / scales
|
||||
packed_fp4 = MXFP4Tensor(x_scaled.flatten(start_dim=-2)).to_packed_tensor(dim=packed_dim)
|
||||
dequant_x = (MXFP4Tensor(x_scaled).to(torch.float32) * scales.to(torch.float8_e4m3fn).to(torch.float32)).flatten(start_dim=-2)
|
||||
fp8_scale = scales.flatten(start_dim=-2).to(torch.float8_e4m3fn)
|
||||
permuted_fp8_scale = None
|
||||
if permuted:
|
||||
scales = scales.view(B, H // 64, 4, 16, M, N // 16).permute(0, 1, 3, 2, 4, 5).reshape(B, H, M, N // 16)
|
||||
permuted_fp8_scale = scales.view(B, H // 64, 64, M, N // 64, 4).permute(0, 1, 4, 3, 2, 5).reshape(B, H, M, N // 16).to(torch.float8_e4m3fn)
|
||||
return fp8_scale, packed_fp4, dequant_x, permuted_fp8_scale
|
||||
|
||||
b = 2
|
||||
h = 4
|
||||
n = 491
|
||||
n_padded = (n + 127) // 128 * 128
|
||||
d = 128
|
||||
|
||||
q = torch.randn(b, h, n, d, dtype=torch.float16, device='cuda')
|
||||
o = torch.empty((b, h, d, n_padded // 2), dtype=torch.uint8, device='cuda')
|
||||
o_s = torch.empty((b, h, d, n_padded // 16), dtype=torch.float8_e4m3fn, device='cuda')
|
||||
|
||||
fp4quant.scaled_fp4_quant_trans(q, o, o_s, 1)
|
||||
|
||||
if n % 128 != 0:
|
||||
q_padded = torch.cat([q, torch.zeros((b, h, n_padded - n, d), dtype=torch.float16, device='cuda')], dim=2)
|
||||
else:
|
||||
q_padded = q
|
||||
|
||||
# use torch transpose + scaled_fp4_quant to get the ground truth
|
||||
q_padded = q_padded.transpose(2, 3).reshape(b, h, n_padded, d).contiguous()
|
||||
o_gt = torch.empty((b, h, n_padded, d // 2), dtype=torch.uint8, device='cuda')
|
||||
o_s_gt = torch.empty((b, h, n_padded, d // 16), dtype=torch.float8_e4m3fn, device='cuda')
|
||||
fp4quant.scaled_fp4_quant(q_padded, o_gt, o_s_gt, 1)
|
||||
o_gt = o_gt.reshape(b, h, d, n_padded // 2).contiguous()
|
||||
o_s_gt = o_s_gt.reshape(b, h, d, n_padded // 16).contiguous()
|
||||
|
||||
assert((o_s_gt.float() - o_s.float()).abs().max() == 0)
|
||||
assert((o_gt - o).abs().max() == 0)
|
||||
|
||||
print("All tests passed!")
|
||||
@@ -0,0 +1,169 @@
|
||||
"""
|
||||
Copyright (c) 2025 by SageAttention team.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
import torch
|
||||
import torch.distributed as dist
|
||||
|
||||
|
||||
def bench(fn, num_warmups: int = 5, num_tests: int = 10,
|
||||
high_precision: bool = False):
|
||||
# Flush L2 cache with 256 MB data
|
||||
torch.cuda.synchronize()
|
||||
cache = torch.empty(int(256e6 // 4), dtype=torch.int, device='cuda')
|
||||
cache.zero_()
|
||||
|
||||
# Warmup
|
||||
for _ in range(num_warmups):
|
||||
fn()
|
||||
|
||||
# Add a large kernel to eliminate the CPU launch overhead
|
||||
if high_precision:
|
||||
x = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
|
||||
y = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
|
||||
x @ y
|
||||
|
||||
# Testing
|
||||
start_event = torch.cuda.Event(enable_timing=True)
|
||||
end_event = torch.cuda.Event(enable_timing=True)
|
||||
start_event.record()
|
||||
for i in range(num_tests):
|
||||
fn()
|
||||
end_event.record()
|
||||
torch.cuda.synchronize()
|
||||
|
||||
return start_event.elapsed_time(end_event) / num_tests
|
||||
|
||||
|
||||
class empty_suppress:
|
||||
def __enter__(self):
|
||||
return self
|
||||
|
||||
def __exit__(self, *_):
|
||||
pass
|
||||
|
||||
|
||||
class suppress_stdout_stderr:
|
||||
def __enter__(self):
|
||||
self.outnull_file = open(os.devnull, 'w')
|
||||
self.errnull_file = open(os.devnull, 'w')
|
||||
|
||||
self.old_stdout_fileno_undup = sys.stdout.fileno()
|
||||
self.old_stderr_fileno_undup = sys.stderr.fileno()
|
||||
|
||||
self.old_stdout_fileno = os.dup(sys.stdout.fileno())
|
||||
self.old_stderr_fileno = os.dup(sys.stderr.fileno())
|
||||
|
||||
self.old_stdout = sys.stdout
|
||||
self.old_stderr = sys.stderr
|
||||
|
||||
os.dup2(self.outnull_file.fileno(), self.old_stdout_fileno_undup)
|
||||
os.dup2(self.errnull_file.fileno(), self.old_stderr_fileno_undup)
|
||||
|
||||
sys.stdout = self.outnull_file
|
||||
sys.stderr = self.errnull_file
|
||||
return self
|
||||
|
||||
def __exit__(self, *_):
|
||||
sys.stdout = self.old_stdout
|
||||
sys.stderr = self.old_stderr
|
||||
|
||||
os.dup2(self.old_stdout_fileno, self.old_stdout_fileno_undup)
|
||||
os.dup2(self.old_stderr_fileno, self.old_stderr_fileno_undup)
|
||||
|
||||
os.close(self.old_stdout_fileno)
|
||||
os.close(self.old_stderr_fileno)
|
||||
|
||||
self.outnull_file.close()
|
||||
self.errnull_file.close()
|
||||
|
||||
|
||||
def bench_kineto(fn, kernel_names, num_tests: int = 30, suppress_kineto_output: bool = False,
|
||||
trace_path: str = None, barrier_comm_profiling: bool = False, flush_l2: bool = False):
|
||||
# Conflict with Nsight Systems
|
||||
using_nsys = os.environ.get('DG_NSYS_PROFILING', False)
|
||||
|
||||
# For some auto-tuning kernels with prints
|
||||
fn()
|
||||
|
||||
# Profile
|
||||
suppress = suppress_stdout_stderr if suppress_kineto_output and not using_nsys else empty_suppress
|
||||
with suppress():
|
||||
schedule = torch.profiler.schedule(wait=0, warmup=1, active=1, repeat=1) if not using_nsys else None
|
||||
profiler = torch.profiler.profile(activities=[torch.profiler.ProfilerActivity.CUDA], schedule=schedule) if not using_nsys else empty_suppress()
|
||||
with profiler:
|
||||
for i in range(2):
|
||||
# NOTES: use a large kernel and a barrier to eliminate the unbalanced CPU launch overhead
|
||||
if barrier_comm_profiling:
|
||||
lhs = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
|
||||
rhs = torch.randn((8192, 8192), dtype=torch.float, device='cuda')
|
||||
lhs @ rhs
|
||||
dist.all_reduce(torch.ones(1, dtype=torch.float, device='cuda'))
|
||||
for _ in range(num_tests):
|
||||
if flush_l2:
|
||||
torch.empty(int(256e6 // 4), dtype=torch.int, device='cuda').zero_()
|
||||
fn()
|
||||
|
||||
if not using_nsys:
|
||||
profiler.step()
|
||||
|
||||
# Return 1 if using Nsight Systems
|
||||
if using_nsys:
|
||||
return 1
|
||||
|
||||
# Parse the profiling table
|
||||
assert isinstance(kernel_names, str) or isinstance(kernel_names, tuple)
|
||||
is_tupled = isinstance(kernel_names, tuple)
|
||||
prof_lines = profiler.key_averages().table(sort_by='cuda_time_total', max_name_column_width=100).split('\n')
|
||||
kernel_names = (kernel_names, ) if isinstance(kernel_names, str) else kernel_names
|
||||
assert all([isinstance(name, str) for name in kernel_names])
|
||||
for name in kernel_names:
|
||||
assert sum([name in line for line in prof_lines]) == 1, f'Errors of the kernel {name} in the profiling table'
|
||||
|
||||
# Save chrome traces
|
||||
if trace_path is not None:
|
||||
profiler.export_chrome_trace(trace_path)
|
||||
|
||||
# Return average kernel times
|
||||
units = {'ms': 1e3, 'us': 1e6}
|
||||
kernel_times = []
|
||||
for name in kernel_names:
|
||||
for line in prof_lines:
|
||||
if name in line:
|
||||
time_str = line.split()[-2]
|
||||
for unit, scale in units.items():
|
||||
if unit in time_str:
|
||||
kernel_times.append(float(time_str.replace(unit, '')) / scale)
|
||||
break
|
||||
break
|
||||
return tuple(kernel_times) if is_tupled else kernel_times[0]
|
||||
|
||||
|
||||
def calc_diff(x, y):
|
||||
x, y = x.double(), y.double()
|
||||
denominator = (x * x + y * y).sum()
|
||||
sim = 2 * (x * y).sum() / denominator
|
||||
return 1 - sim
|
||||
|
||||
|
||||
def count_bytes(tensors):
|
||||
total = 0
|
||||
for t in tensors:
|
||||
if isinstance(t, tuple):
|
||||
total += count_bytes(t)
|
||||
else:
|
||||
total += t.numel() * t.element_size()
|
||||
return total
|
||||
@@ -0,0 +1,52 @@
|
||||
#pragma once
|
||||
|
||||
#include <stdio.h>
|
||||
|
||||
#if defined(__HIPCC__)
|
||||
#define HOST_DEVICE_INLINE __host__ __device__
|
||||
#define DEVICE_INLINE __device__
|
||||
#define HOST_INLINE __host__
|
||||
#elif defined(__CUDACC__) || defined(_NVHPC_CUDA)
|
||||
#define HOST_DEVICE_INLINE __host__ __device__ __forceinline__
|
||||
#define DEVICE_INLINE __device__ __forceinline__
|
||||
#define HOST_INLINE __host__ __forceinline__
|
||||
#else
|
||||
#define HOST_DEVICE_INLINE inline
|
||||
#define DEVICE_INLINE inline
|
||||
#define HOST_INLINE inline
|
||||
#endif
|
||||
|
||||
#define CUDA_CHECK(cmd) \
|
||||
do { \
|
||||
cudaError_t e = cmd; \
|
||||
if (e != cudaSuccess) { \
|
||||
printf("Failed: Cuda error %s:%d '%s'\n", __FILE__, __LINE__, \
|
||||
cudaGetErrorString(e)); \
|
||||
exit(EXIT_FAILURE); \
|
||||
} \
|
||||
} while (0)
|
||||
|
||||
int64_t get_device_attribute(int64_t attribute, int64_t device_id) {
|
||||
static int value = [=]() {
|
||||
int device = static_cast<int>(device_id);
|
||||
if (device < 0) {
|
||||
CUDA_CHECK(cudaGetDevice(&device));
|
||||
}
|
||||
int value;
|
||||
CUDA_CHECK(cudaDeviceGetAttribute(
|
||||
&value, static_cast<cudaDeviceAttr>(attribute), device));
|
||||
return static_cast<int>(value);
|
||||
}();
|
||||
|
||||
return value;
|
||||
}
|
||||
|
||||
namespace cuda_utils {
|
||||
|
||||
template <typename T>
|
||||
HOST_DEVICE_INLINE constexpr std::enable_if_t<std::is_integral_v<T>, T>
|
||||
ceil_div(T a, T b) {
|
||||
return (a + b - 1) / b;
|
||||
}
|
||||
|
||||
}; // namespace cuda_utils
|
||||
@@ -0,0 +1,644 @@
|
||||
/*
|
||||
* Copyright (c) 2025 by SageAttention team.
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
#include <torch/all.h>
|
||||
#include <torch/python.h>
|
||||
#include <torch/nn/functional.h>
|
||||
#include <ATen/cuda/CUDAContext.h>
|
||||
#include <c10/cuda/CUDAGuard.h>
|
||||
#include <cuda_runtime_api.h>
|
||||
#include <cuda_runtime.h>
|
||||
|
||||
#include <ATen/cuda/CUDAContext.h>
|
||||
#include <c10/cuda/CUDAGuard.h>
|
||||
|
||||
#include <cuda_fp8.h>
|
||||
|
||||
#include "cuda_utils.h"
|
||||
#include "../blackwell/block_config.h"
|
||||
|
||||
#define DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(pytorch_dtype, c_type, ...) \
|
||||
if (pytorch_dtype == at::ScalarType::Half) { \
|
||||
using c_type = half; \
|
||||
__VA_ARGS__ \
|
||||
} else if (pytorch_dtype == at::ScalarType::BFloat16) { \
|
||||
using c_type = nv_bfloat16; \
|
||||
__VA_ARGS__ \
|
||||
} else { \
|
||||
std::ostringstream oss; \
|
||||
oss << __PRETTY_FUNCTION__ << " failed to dispatch data type " << pytorch_dtype; \
|
||||
TORCH_CHECK(false, oss.str()); \
|
||||
}
|
||||
|
||||
#define DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, ...) \
|
||||
if (head_dim == 64) { \
|
||||
constexpr int HEAD_DIM = 64; \
|
||||
__VA_ARGS__ \
|
||||
} else if (head_dim == 128) { \
|
||||
constexpr int HEAD_DIM = 128; \
|
||||
__VA_ARGS__ \
|
||||
} else { \
|
||||
std::ostringstream err_msg; \
|
||||
err_msg << "Unsupported head dim: " << int(head_dim); \
|
||||
throw std::invalid_argument(err_msg.str()); \
|
||||
}
|
||||
|
||||
#define CHECK_CUDA(x) \
|
||||
TORCH_CHECK(x.is_cuda(), "Tensor " #x " must be on CUDA")
|
||||
#define CHECK_DTYPE(x, true_dtype) \
|
||||
TORCH_CHECK(x.dtype() == true_dtype, \
|
||||
"Tensor " #x " must have dtype (" #true_dtype ")")
|
||||
#define CHECK_DIMS(x, true_dim) \
|
||||
TORCH_CHECK(x.dim() == true_dim, \
|
||||
"Tensor " #x " must have dimension number (" #true_dim ")")
|
||||
#define CHECK_SHAPE(x, ...) \
|
||||
TORCH_CHECK(x.sizes() == torch::IntArrayRef({__VA_ARGS__}), \
|
||||
"Tensor " #x " must have shape (" #__VA_ARGS__ ")")
|
||||
#define CHECK_CONTIGUOUS(x) \
|
||||
TORCH_CHECK(x.is_contiguous(), "Tensor " #x " must be contiguous")
|
||||
#define CHECK_LASTDIM_CONTIGUOUS(x) \
|
||||
TORCH_CHECK(x.stride(-1) == 1, \
|
||||
"Tensor " #x " must be contiguous at the last dimension")
|
||||
|
||||
constexpr int CVT_FP4_ELTS_PER_THREAD = 16;
|
||||
|
||||
// Convert 4 float2 values into 8 e2m1 values (represented as one uint32_t).
|
||||
inline __device__ uint32_t fp32_vec_to_e2m1(float2 *array) {
|
||||
#if defined(__CUDA_ARCH__) && (__CUDA_ARCH__ >= 1000)
|
||||
uint32_t val;
|
||||
asm volatile(
|
||||
"{\n"
|
||||
".reg .b8 byte0;\n"
|
||||
".reg .b8 byte1;\n"
|
||||
".reg .b8 byte2;\n"
|
||||
".reg .b8 byte3;\n"
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte0, %2, %1;\n"
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte1, %4, %3;\n"
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte2, %6, %5;\n"
|
||||
"cvt.rn.satfinite.e2m1x2.f32 byte3, %8, %7;\n"
|
||||
"mov.b32 %0, {byte0, byte1, byte2, byte3};\n"
|
||||
"}"
|
||||
: "=r"(val)
|
||||
: "f"(array[0].x), "f"(array[0].y), "f"(array[1].x), "f"(array[1].y),
|
||||
"f"(array[2].x), "f"(array[2].y), "f"(array[3].x), "f"(array[3].y));
|
||||
return val;
|
||||
#else
|
||||
return 0;
|
||||
#endif
|
||||
}
|
||||
|
||||
// Get type2 from type or vice versa (applied to half and bfloat16)
|
||||
template <typename T>
|
||||
struct TypeConverter {
|
||||
using Type = half2;
|
||||
}; // keep for generality
|
||||
|
||||
template <>
|
||||
struct TypeConverter<half2> {
|
||||
using Type = half;
|
||||
};
|
||||
|
||||
template <>
|
||||
struct TypeConverter<half> {
|
||||
using Type = half2;
|
||||
};
|
||||
|
||||
template <>
|
||||
struct TypeConverter<__nv_bfloat162> {
|
||||
using Type = __nv_bfloat16;
|
||||
};
|
||||
|
||||
template <>
|
||||
struct TypeConverter<__nv_bfloat16> {
|
||||
using Type = __nv_bfloat162;
|
||||
};
|
||||
|
||||
// Define a 32 bytes packed data type.
|
||||
template <class Type>
|
||||
struct PackedVec {
|
||||
typename TypeConverter<Type>::Type elts[8];
|
||||
};
|
||||
|
||||
template <uint32_t head_dim, uint32_t BLOCK_SIZE, bool permute, typename T>
|
||||
__global__ void scaled_fp4_quant_kernel(
|
||||
const T* input, uint8_t* output, uint8_t* output_sf,
|
||||
int batch_size, int num_heads, int num_tokens,
|
||||
int stride_bz_input, int stride_h_input, int stride_seq_input,
|
||||
int stride_bz_output, int stride_h_output, int stride_seq_output,
|
||||
int stride_bz_output_sf, int stride_h_output_sf, int stride_seq_output_sf) {
|
||||
static_assert(std::is_same<T, half>::value || std::is_same<T, nv_bfloat16>::value, "Only half and bfloat16 input are supported");
|
||||
using PackedVec = PackedVec<T>;
|
||||
|
||||
const int batch_id = blockIdx.y;
|
||||
const int head_id = blockIdx.z;
|
||||
const int token_block_id = blockIdx.x;
|
||||
|
||||
static_assert(CVT_FP4_ELTS_PER_THREAD == 8 || CVT_FP4_ELTS_PER_THREAD == 16,
|
||||
"CVT_FP4_ELTS_PER_THREAD must be 8 or 16");
|
||||
static_assert(sizeof(PackedVec) == sizeof(T) * CVT_FP4_ELTS_PER_THREAD,
|
||||
"Vec size is not matched.");
|
||||
|
||||
constexpr uint32_t NUM_THREADS_PER_TOKEN = head_dim / CVT_FP4_ELTS_PER_THREAD;
|
||||
|
||||
// load input
|
||||
const int token_id = token_block_id * BLOCK_SIZE + threadIdx.x / NUM_THREADS_PER_TOKEN;
|
||||
|
||||
int load_token_id;
|
||||
if constexpr (!permute) {
|
||||
load_token_id = token_id;
|
||||
} else {
|
||||
int local_token_id = threadIdx.x / NUM_THREADS_PER_TOKEN;
|
||||
int local_token_id_residue = local_token_id % 32;
|
||||
// [0, 1, 8, 9, 16, 17, 24, 25, 2, 3, 10, 11, 18, 19, 26, 27, 4, 5, 12, 13, 20, 21, 28, 29, 6, 7, 14, 15, 22, 23, 30, 31]
|
||||
load_token_id = token_block_id * BLOCK_SIZE + (local_token_id / 32) * 32 +
|
||||
(local_token_id_residue / 8) * 2 +
|
||||
((local_token_id_residue % 8) / 2) * 8 +
|
||||
(local_token_id_residue % 8) % 2;
|
||||
}
|
||||
|
||||
PackedVec in_vec;
|
||||
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
|
||||
reinterpret_cast<uint32_t&>(in_vec.elts[i]) = 0;
|
||||
}
|
||||
|
||||
if (load_token_id < num_tokens) {
|
||||
in_vec = reinterpret_cast<PackedVec const*>(input +
|
||||
batch_id * stride_bz_input + // batch dim
|
||||
head_id * stride_h_input + // head dim
|
||||
load_token_id * stride_seq_input + // seq dim
|
||||
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD)[0]; // feature dim
|
||||
}
|
||||
|
||||
// calculate max of every consecutive 16 elements
|
||||
auto localMax = __habs2(in_vec.elts[0]);
|
||||
#pragma unroll
|
||||
for (int i = 1; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) { // local max
|
||||
localMax = __hmax2(localMax, __habs2(in_vec.elts[i]));
|
||||
}
|
||||
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) { // shuffle across two threads
|
||||
localMax = __hmax2(__shfl_xor_sync(0xffffffff, localMax, 1, 32), localMax);
|
||||
}
|
||||
|
||||
float vecMax = float(__hmax(localMax.x, localMax.y));
|
||||
|
||||
// scaling factor
|
||||
float SFValue = vecMax / 6.0f;
|
||||
uint8_t SFValueFP8;
|
||||
reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8) = __nv_fp8_e4m3(SFValue);
|
||||
SFValue = float(reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8));
|
||||
|
||||
float SFValueInv = (SFValue == 0.0f) ? 0.0f : 1.0f / SFValue;
|
||||
|
||||
// convert input to float2 and apply scale
|
||||
float2 fp2Vals[CVT_FP4_ELTS_PER_THREAD / 2];
|
||||
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
|
||||
if constexpr (std::is_same<T, half>::value) {
|
||||
fp2Vals[i] = __half22float2(in_vec.elts[i]);
|
||||
} else {
|
||||
fp2Vals[i] = __bfloat1622float2(in_vec.elts[i]);
|
||||
}
|
||||
fp2Vals[i].x = fp2Vals[i].x * SFValueInv;
|
||||
fp2Vals[i].y = fp2Vals[i].y * SFValueInv;
|
||||
}
|
||||
|
||||
// convert to e2m1
|
||||
uint32_t e2m1Vals[CVT_FP4_ELTS_PER_THREAD / 8];
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 8; i++) {
|
||||
e2m1Vals[i] = fp32_vec_to_e2m1(fp2Vals + i * 4);
|
||||
}
|
||||
|
||||
// Skip out-of-range tokens: never write past the sequence when num_tokens
|
||||
// is not a multiple of BLOCK_SIZE (out-of-bounds write fix).
|
||||
if (token_id >= num_tokens) return;
|
||||
|
||||
// save
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) {
|
||||
reinterpret_cast<uint32_t*>(output +
|
||||
batch_id * stride_bz_output +
|
||||
head_id * stride_h_output +
|
||||
token_id * stride_seq_output +
|
||||
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD / 2)[0] = e2m1Vals[0];
|
||||
} else {
|
||||
reinterpret_cast<uint64_t*>(output +
|
||||
batch_id * stride_bz_output +
|
||||
head_id * stride_h_output +
|
||||
token_id * stride_seq_output +
|
||||
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD / 2)[0] = reinterpret_cast<uint64_t*>(e2m1Vals)[0];
|
||||
}
|
||||
|
||||
uint8_t* output_sf_save_base = output_sf + batch_id * stride_bz_output_sf + head_id * stride_h_output_sf + (token_id / 64) * 64 * stride_seq_output_sf;
|
||||
uint32_t token_id_local = token_id % 64;
|
||||
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 16) {
|
||||
uint32_t col_id_local = threadIdx.x % NUM_THREADS_PER_TOKEN;
|
||||
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
|
||||
(token_id_local / 16) * 4 + (token_id_local % 16) * 16;
|
||||
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
|
||||
} else {
|
||||
if (threadIdx.x % 2 == 0) {
|
||||
uint32_t col_id_local = (threadIdx.x % NUM_THREADS_PER_TOKEN) / 2;
|
||||
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
|
||||
(token_id_local / 16) * 4 + (token_id_local % 16) * 16;
|
||||
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
template <uint32_t head_dim, uint32_t BLOCK_SIZE, typename T>
|
||||
__global__ void scaled_fp4_quant_trans_kernel(
|
||||
const T* input, uint8_t* output, uint8_t* output_sf,
|
||||
int batch_size, int num_heads, int num_tokens,
|
||||
int stride_bz_input, int stride_h_input, int stride_seq_input,
|
||||
int stride_bz_output, int stride_h_output, int stride_d_output,
|
||||
int stride_bz_output_sf, int stride_h_output_sf, int stride_d_output_sf) {
|
||||
static_assert(std::is_same<T, half>::value || std::is_same<T, nv_bfloat16>::value, "Only half and bfloat16 input are supported");
|
||||
using PackedVec = PackedVec<T>;
|
||||
|
||||
const int batch_id = blockIdx.y;
|
||||
const int head_id = blockIdx.z;
|
||||
const int token_block_id = blockIdx.x;
|
||||
|
||||
static_assert(CVT_FP4_ELTS_PER_THREAD == 8 || CVT_FP4_ELTS_PER_THREAD == 16,
|
||||
"CVT_FP4_ELTS_PER_THREAD must be 8 or 16");
|
||||
static_assert(sizeof(PackedVec) == sizeof(T) * CVT_FP4_ELTS_PER_THREAD,
|
||||
"Vec size is not matched.");
|
||||
|
||||
constexpr uint32_t NUM_THREADS_PER_TOKEN = head_dim / CVT_FP4_ELTS_PER_THREAD;
|
||||
constexpr uint32_t NUM_THREADS_PER_SEQ = BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD;
|
||||
|
||||
// load input
|
||||
const int token_id = token_block_id * BLOCK_SIZE + threadIdx.x / NUM_THREADS_PER_TOKEN;
|
||||
// Permute V rows within each 32-element block so the PV MMA K-indexed
|
||||
// access reads the correct CLayout N-indexed values (Edenzzzz causal fix).
|
||||
const int k_intra = token_id & 31;
|
||||
const int load_token_id = (token_id & ~31)
|
||||
| ((k_intra & 6) << 2) | ((k_intra & 24) >> 2) | (k_intra & 1);
|
||||
|
||||
PackedVec in_vec;
|
||||
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
|
||||
reinterpret_cast<uint32_t&>(in_vec.elts[i]) = 0;
|
||||
}
|
||||
|
||||
if (load_token_id < num_tokens) {
|
||||
in_vec = reinterpret_cast<PackedVec const*>(input +
|
||||
batch_id * stride_bz_input + // batch dim
|
||||
head_id * stride_h_input + // head dim
|
||||
load_token_id * stride_seq_input + // seq dim (permuted)
|
||||
(threadIdx.x % NUM_THREADS_PER_TOKEN) * CVT_FP4_ELTS_PER_THREAD)[0]; // feature dim
|
||||
}
|
||||
|
||||
// transpose
|
||||
__shared__ T shared_input[BLOCK_SIZE * head_dim];
|
||||
reinterpret_cast<PackedVec*>(shared_input)[threadIdx.x] = in_vec;
|
||||
__syncthreads();
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
|
||||
in_vec.elts[i].x = shared_input[(threadIdx.x / NUM_THREADS_PER_SEQ) + ((threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD + 2 * i) * head_dim];
|
||||
in_vec.elts[i].y = shared_input[(threadIdx.x / NUM_THREADS_PER_SEQ) + ((threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD + 2 * i + 1) * head_dim];
|
||||
}
|
||||
|
||||
// calculate max of every consecutive 16 elements
|
||||
auto localMax = __habs2(in_vec.elts[0]);
|
||||
#pragma unroll
|
||||
for (int i = 1; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) { // local max
|
||||
localMax = __hmax2(localMax, __habs2(in_vec.elts[i]));
|
||||
}
|
||||
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) { // shuffle across two threads
|
||||
localMax = __hmax2(__shfl_xor_sync(0xffffffff, localMax, 1, 32), localMax);
|
||||
}
|
||||
|
||||
float vecMax = float(__hmax(localMax.x, localMax.y));
|
||||
|
||||
// scaling factor
|
||||
float SFValue = vecMax / 6.0f;
|
||||
uint8_t SFValueFP8;
|
||||
reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8) = __nv_fp8_e4m3(SFValue);
|
||||
SFValue = float(reinterpret_cast<__nv_fp8_e4m3&>(SFValueFP8));
|
||||
|
||||
float SFValueInv = (SFValue == 0.0f) ? 0.0f : 1.0f / SFValue;
|
||||
|
||||
// convert input to float2 and apply scale
|
||||
float2 fp2Vals[CVT_FP4_ELTS_PER_THREAD / 2];
|
||||
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 2; i++) {
|
||||
if constexpr (std::is_same<T, half>::value) {
|
||||
fp2Vals[i] = __half22float2(in_vec.elts[i]);
|
||||
} else {
|
||||
fp2Vals[i] = __bfloat1622float2(in_vec.elts[i]);
|
||||
}
|
||||
fp2Vals[i].x = fp2Vals[i].x * SFValueInv;
|
||||
fp2Vals[i].y = fp2Vals[i].y * SFValueInv;
|
||||
}
|
||||
|
||||
// convert to e2m1
|
||||
uint32_t e2m1Vals[CVT_FP4_ELTS_PER_THREAD / 8];
|
||||
#pragma unroll
|
||||
for (int i = 0; i < CVT_FP4_ELTS_PER_THREAD / 8; i++) {
|
||||
e2m1Vals[i] = fp32_vec_to_e2m1(fp2Vals + i * 4);
|
||||
}
|
||||
|
||||
// Skip out-of-range tokens: never write past the sequence when num_tokens
|
||||
// is not a multiple of BLOCK_SIZE (out-of-bounds write fix).
|
||||
const int write_token_id = token_block_id * BLOCK_SIZE +
|
||||
(threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD;
|
||||
if (write_token_id >= num_tokens) return;
|
||||
|
||||
// save
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 8) {
|
||||
reinterpret_cast<uint32_t*>(output +
|
||||
batch_id * stride_bz_output +
|
||||
head_id * stride_h_output +
|
||||
(threadIdx.x / NUM_THREADS_PER_SEQ) * stride_d_output +
|
||||
(token_block_id * BLOCK_SIZE + (threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD) / 2)[0] = e2m1Vals[0];
|
||||
} else {
|
||||
reinterpret_cast<uint64_t*>(output +
|
||||
batch_id * stride_bz_output +
|
||||
head_id * stride_h_output +
|
||||
(threadIdx.x / NUM_THREADS_PER_SEQ) * stride_d_output +
|
||||
(token_block_id * BLOCK_SIZE + (threadIdx.x % NUM_THREADS_PER_SEQ) * CVT_FP4_ELTS_PER_THREAD) / 2)[0] = reinterpret_cast<uint64_t*>(e2m1Vals)[0];
|
||||
}
|
||||
|
||||
uint8_t *output_sf_save_base = output_sf +
|
||||
batch_id * stride_bz_output_sf +
|
||||
head_id * stride_h_output_sf +
|
||||
(threadIdx.x / NUM_THREADS_PER_SEQ / 64) * 64 * stride_d_output_sf;
|
||||
uint32_t row_id_local = (threadIdx.x / NUM_THREADS_PER_SEQ) % 64;
|
||||
|
||||
if constexpr (CVT_FP4_ELTS_PER_THREAD == 16) {
|
||||
uint32_t col_id_local = token_block_id * BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD + threadIdx.x % NUM_THREADS_PER_SEQ;
|
||||
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
|
||||
(row_id_local / 16) * 4 + (row_id_local % 16) * 16;
|
||||
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
|
||||
} else {
|
||||
if (threadIdx.x % 2 == 0) {
|
||||
uint32_t col_id_local = token_block_id * BLOCK_SIZE / CVT_FP4_ELTS_PER_THREAD + (threadIdx.x % NUM_THREADS_PER_SEQ) / 2;
|
||||
uint32_t offset_local = (col_id_local / 4) * 256 + (col_id_local % 4) +
|
||||
(row_id_local / 16) * 4 + (row_id_local % 16) * 16;
|
||||
reinterpret_cast<uint8_t*>(output_sf_save_base + offset_local)[0] = SFValueFP8;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void scaled_fp4_quant(torch::Tensor const& input,
|
||||
torch::Tensor const& output,
|
||||
torch::Tensor const& output_sf,
|
||||
int tensor_layout) {
|
||||
constexpr int BLOCK_SIZE = flash::BLOCK_M;
|
||||
|
||||
CHECK_CUDA(input);
|
||||
CHECK_CUDA(output);
|
||||
CHECK_CUDA(output_sf);
|
||||
|
||||
CHECK_LASTDIM_CONTIGUOUS(input);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output_sf);
|
||||
|
||||
CHECK_DTYPE(output, at::ScalarType::Byte);
|
||||
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
|
||||
|
||||
CHECK_DIMS(input, 4);
|
||||
CHECK_DIMS(output, 4);
|
||||
CHECK_DIMS(output_sf, 4);
|
||||
|
||||
const int batch_size = input.size(0);
|
||||
const int head_dim = input.size(3);
|
||||
|
||||
const int stride_bz_input = input.stride(0);
|
||||
const int stride_bz_output = output.stride(0);
|
||||
const int stride_bz_output_sf = output_sf.stride(0);
|
||||
|
||||
int num_tokens, num_heads;
|
||||
int stride_seq_input, stride_seq_output, stride_seq_output_sf;
|
||||
int stride_h_input, stride_h_output, stride_h_output_sf;
|
||||
if (tensor_layout == 0) {
|
||||
num_tokens = input.size(1);
|
||||
num_heads = input.size(2);
|
||||
stride_seq_input = input.stride(1);
|
||||
stride_seq_output = output.stride(1);
|
||||
stride_seq_output_sf = output_sf.stride(1);
|
||||
stride_h_input = input.stride(2);
|
||||
stride_h_output = output.stride(2);
|
||||
stride_h_output_sf = output_sf.stride(2);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, num_tokens, num_heads, head_dim / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, num_tokens, num_heads, head_dim / 16);
|
||||
} else {
|
||||
num_tokens = input.size(2);
|
||||
num_heads = input.size(1);
|
||||
stride_seq_input = input.stride(2);
|
||||
stride_seq_output = output.stride(2);
|
||||
stride_seq_output_sf = output_sf.stride(2);
|
||||
stride_h_input = input.stride(1);
|
||||
stride_h_output = output.stride(1);
|
||||
stride_h_output_sf = output_sf.stride(1);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, num_heads, num_tokens, head_dim / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, num_heads, num_tokens, head_dim / 16);
|
||||
}
|
||||
|
||||
auto input_dtype = input.scalar_type();
|
||||
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
|
||||
|
||||
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
|
||||
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
|
||||
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
|
||||
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
|
||||
|
||||
scaled_fp4_quant_kernel<HEAD_DIM, BLOCK_SIZE, false, c_type>
|
||||
<<<grid, block, 0, stream>>>(
|
||||
reinterpret_cast<c_type*>(input.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
|
||||
batch_size, num_heads, num_tokens,
|
||||
stride_bz_input, stride_h_input, stride_seq_input,
|
||||
stride_bz_output, stride_h_output, stride_seq_output,
|
||||
stride_bz_output_sf, stride_h_output_sf, stride_seq_output_sf);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
void scaled_fp4_quant_permute(torch::Tensor const& input,
|
||||
torch::Tensor const& output,
|
||||
torch::Tensor const& output_sf,
|
||||
int tensor_layout) {
|
||||
constexpr int BLOCK_SIZE = flash::BLOCK_M;
|
||||
|
||||
CHECK_CUDA(input);
|
||||
CHECK_CUDA(output);
|
||||
CHECK_CUDA(output_sf);
|
||||
|
||||
CHECK_LASTDIM_CONTIGUOUS(input);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output_sf);
|
||||
|
||||
CHECK_DTYPE(output, at::ScalarType::Byte);
|
||||
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
|
||||
|
||||
CHECK_DIMS(input, 4);
|
||||
CHECK_DIMS(output, 4);
|
||||
CHECK_DIMS(output_sf, 4);
|
||||
|
||||
const int batch_size = input.size(0);
|
||||
const int head_dim = input.size(3);
|
||||
|
||||
const int stride_bz_input = input.stride(0);
|
||||
const int stride_bz_output = output.stride(0);
|
||||
const int stride_bz_output_sf = output_sf.stride(0);
|
||||
|
||||
int num_tokens, num_heads;
|
||||
int stride_seq_input, stride_seq_output, stride_seq_output_sf;
|
||||
int stride_h_input, stride_h_output, stride_h_output_sf;
|
||||
if (tensor_layout == 0) {
|
||||
num_tokens = input.size(1);
|
||||
num_heads = input.size(2);
|
||||
stride_seq_input = input.stride(1);
|
||||
stride_seq_output = output.stride(1);
|
||||
stride_seq_output_sf = output_sf.stride(1);
|
||||
stride_h_input = input.stride(2);
|
||||
stride_h_output = output.stride(2);
|
||||
stride_h_output_sf = output_sf.stride(2);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, num_heads, head_dim / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, num_heads, head_dim / 16);
|
||||
} else {
|
||||
num_tokens = input.size(2);
|
||||
num_heads = input.size(1);
|
||||
stride_seq_input = input.stride(2);
|
||||
stride_seq_output = output.stride(2);
|
||||
stride_seq_output_sf = output_sf.stride(2);
|
||||
stride_h_input = input.stride(1);
|
||||
stride_h_output = output.stride(1);
|
||||
stride_h_output_sf = output_sf.stride(1);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, head_dim / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE, head_dim / 16);
|
||||
}
|
||||
|
||||
auto input_dtype = input.scalar_type();
|
||||
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
|
||||
|
||||
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
|
||||
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
|
||||
constexpr int BLOCK_SIZE = flash::BLOCK_M;
|
||||
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
|
||||
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
|
||||
|
||||
scaled_fp4_quant_kernel<HEAD_DIM, BLOCK_SIZE, true, c_type>
|
||||
<<<grid, block, 0, stream>>>(
|
||||
reinterpret_cast<c_type*>(input.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
|
||||
batch_size, num_heads, num_tokens,
|
||||
stride_bz_input, stride_h_input, stride_seq_input,
|
||||
stride_bz_output, stride_h_output, stride_seq_output,
|
||||
stride_bz_output_sf, stride_h_output_sf, stride_seq_output_sf);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
void scaled_fp4_quant_trans(torch::Tensor const& input,
|
||||
torch::Tensor const& output,
|
||||
torch::Tensor const& output_sf,
|
||||
int tensor_layout) {
|
||||
constexpr int BLOCK_SIZE = flash::BLOCK_M;
|
||||
|
||||
CHECK_CUDA(input);
|
||||
CHECK_CUDA(output);
|
||||
CHECK_CUDA(output_sf);
|
||||
|
||||
CHECK_LASTDIM_CONTIGUOUS(input);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output);
|
||||
CHECK_LASTDIM_CONTIGUOUS(output_sf);
|
||||
|
||||
CHECK_DTYPE(output, at::ScalarType::Byte);
|
||||
CHECK_DTYPE(output_sf, at::ScalarType::Float8_e4m3fn);
|
||||
|
||||
CHECK_DIMS(input, 4);
|
||||
CHECK_DIMS(output, 4);
|
||||
CHECK_DIMS(output_sf, 4);
|
||||
|
||||
const int batch_size = input.size(0);
|
||||
const int head_dim = input.size(3);
|
||||
|
||||
const int stride_bz_input = input.stride(0);
|
||||
const int stride_bz_output = output.stride(0);
|
||||
const int stride_bz_output_sf = output_sf.stride(0);
|
||||
|
||||
int num_tokens, num_heads;
|
||||
int stride_seq_input;
|
||||
int stride_d_output, stride_d_output_sf;
|
||||
int stride_h_input, stride_h_output, stride_h_output_sf;
|
||||
if (tensor_layout == 0) {
|
||||
num_tokens = input.size(1);
|
||||
num_heads = input.size(2);
|
||||
stride_seq_input = input.stride(1);
|
||||
stride_d_output = output.stride(1);
|
||||
stride_d_output_sf = output_sf.stride(1);
|
||||
stride_h_input = input.stride(2);
|
||||
stride_h_output = output.stride(2);
|
||||
stride_h_output_sf = output_sf.stride(2);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, head_dim, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, head_dim, num_heads, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 16);
|
||||
} else {
|
||||
num_tokens = input.size(2);
|
||||
num_heads = input.size(1);
|
||||
stride_seq_input = input.stride(2);
|
||||
stride_d_output = output.stride(2);
|
||||
stride_d_output_sf = output_sf.stride(2);
|
||||
stride_h_input = input.stride(1);
|
||||
stride_h_output = output.stride(1);
|
||||
stride_h_output_sf = output_sf.stride(1);
|
||||
|
||||
CHECK_SHAPE(output, batch_size, num_heads, head_dim, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 2);
|
||||
CHECK_SHAPE(output_sf, batch_size, num_heads, head_dim, ((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE) * BLOCK_SIZE / 16);
|
||||
}
|
||||
|
||||
auto input_dtype = input.scalar_type();
|
||||
auto stream = at::cuda::getCurrentCUDAStream(input.get_device());
|
||||
|
||||
DISPATCH_PYTORCH_DTYPE_TO_CTYPE_FP16(input_dtype, c_type, {
|
||||
DISPATCH_HEAD_DIM(head_dim, HEAD_DIM, {
|
||||
dim3 block(BLOCK_SIZE * HEAD_DIM / CVT_FP4_ELTS_PER_THREAD, 1, 1);
|
||||
dim3 grid((num_tokens + BLOCK_SIZE - 1) / BLOCK_SIZE, batch_size, num_heads);
|
||||
|
||||
scaled_fp4_quant_trans_kernel<HEAD_DIM, BLOCK_SIZE, c_type>
|
||||
<<<grid, block, 0, stream>>>(
|
||||
reinterpret_cast<c_type*>(input.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output.data_ptr()),
|
||||
reinterpret_cast<uint8_t*>(output_sf.data_ptr()),
|
||||
batch_size, num_heads, num_tokens,
|
||||
stride_bz_input, stride_h_input, stride_seq_input,
|
||||
stride_bz_output, stride_h_output, stride_d_output,
|
||||
stride_bz_output_sf, stride_h_output_sf, stride_d_output_sf);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
|
||||
m.def("scaled_fp4_quant", &scaled_fp4_quant);
|
||||
m.def("scaled_fp4_quant_permute", &scaled_fp4_quant_permute);
|
||||
m.def("scaled_fp4_quant_trans", &scaled_fp4_quant_trans);
|
||||
}
|
||||
@@ -78,16 +78,26 @@ print(f'{mj}.{mn}')"
|
||||
}
|
||||
|
||||
if [ "${GPU_BACKEND}" = "CUDA" ]; then
|
||||
detected_cc="$(detect_with_torch)" || {
|
||||
echo "ERROR: torch-based CUDA arch detection failed in uv environment." >&2
|
||||
echo " Ensure torch is installed and CUDA is available in the uv-selected Python." >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
cc_major="${detected_cc%%.*}"
|
||||
cc_minor="${detected_cc##*.}"
|
||||
# Compute capability drives the arch/TK defaults below. Prefer an explicit
|
||||
# TORCH_CUDA_ARCH_LIST (works on GPU-less build machines such as CI/Docker);
|
||||
# only probe a live GPU via torch when no arch was provided.
|
||||
if [ -n "${TORCH_CUDA_ARCH_LIST:-}" ]; then
|
||||
echo "Using TORCH_CUDA_ARCH_LIST=${TORCH_CUDA_ARCH_LIST} (skipping torch GPU probe)"
|
||||
first_arch="${TORCH_CUDA_ARCH_LIST%%[;, ]*}" # first entry, e.g. 9.0a
|
||||
first_arch="${first_arch%[af]}" # strip trailing a/f suffix
|
||||
cc_major="${first_arch%%.*}"
|
||||
cc_minor="${first_arch##*.}"
|
||||
else
|
||||
detected_cc="$(detect_with_torch)" || {
|
||||
echo "ERROR: torch-based CUDA arch detection failed and TORCH_CUDA_ARCH_LIST is unset." >&2
|
||||
echo " Set TORCH_CUDA_ARCH_LIST (e.g. 9.0a) for GPU-less builds, or build where CUDA is available." >&2
|
||||
exit 1
|
||||
}
|
||||
cc_major="${detected_cc%%.*}"
|
||||
cc_minor="${detected_cc##*.}"
|
||||
echo "Detected compute capability via torch: ${detected_cc}"
|
||||
fi
|
||||
cmake_arch="${cc_major}${cc_minor}"
|
||||
echo "Detected compute capability via torch: ${detected_cc} (sm_${cmake_arch})"
|
||||
|
||||
# Respect explicit overrides.
|
||||
if [ -z "${TORCH_CUDA_ARCH_LIST:-}" ]; then
|
||||
|
||||
@@ -32,4 +32,4 @@ dependencies = [
|
||||
[tool.scikit-build]
|
||||
cmake.build-type = "Release"
|
||||
minimum-version = "build-system.requires"
|
||||
wheel.packages = ["python/fastvideo_kernel"]
|
||||
wheel.packages = ["python/fastvideo_kernel", "attn_qat_infer"]
|
||||
|
||||
@@ -25,12 +25,17 @@ from fastvideo_kernel.turbodiffusion_ops import (
|
||||
int8_quant,
|
||||
)
|
||||
|
||||
from fastvideo_kernel.block_sparse_attn_varlen import (
|
||||
block_sparse_attn_varlen,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"sliding_tile_attention",
|
||||
"video_sparse_attn",
|
||||
"video_sparse_attn_bshd",
|
||||
"block_sparse_attn",
|
||||
"block_sparse_attn_from_indices",
|
||||
"block_sparse_attn_varlen",
|
||||
"moba_attn_varlen",
|
||||
"process_moba_input",
|
||||
"process_moba_output",
|
||||
|
||||
@@ -0,0 +1,207 @@
|
||||
"""Variable-length block-sparse attention via sequence packing.
|
||||
|
||||
Packs multiple variable-length sequences into a single [1, H, T_total, D]
|
||||
tensor and delegates to the existing block_sparse_attn_from_indices kernel
|
||||
in a single launch. No kernel modifications required.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Sequence
|
||||
|
||||
import torch
|
||||
|
||||
from .block_sparse_attn import block_sparse_attn_from_indices
|
||||
|
||||
BLOCK_SIZE = 64
|
||||
|
||||
|
||||
def _scatter_to_padded(
|
||||
src: torch.Tensor,
|
||||
block_sizes: torch.Tensor,
|
||||
block_size: int,
|
||||
dst: torch.Tensor,
|
||||
dst_offset: int,
|
||||
src_start: int,
|
||||
src_end: int,
|
||||
) -> None:
|
||||
"""Copy tokens from a flat source into block-aligned positions in dst.
|
||||
|
||||
Each block occupies exactly `block_size` slots in dst. The first
|
||||
`block_sizes[b]` slots of block *b* receive real tokens; the remainder
|
||||
stays zero (padding the kernel expects).
|
||||
|
||||
src: [total_tokens, H, D]
|
||||
dst: [1, H, total_padded, D]
|
||||
block_sizes: [num_blocks] int32, actual token count per block.
|
||||
"""
|
||||
src_pos = src_start
|
||||
dst_pos = dst_offset
|
||||
sizes = block_sizes.cpu().tolist()
|
||||
for actual in sizes:
|
||||
actual = min(actual, src_end - src_pos)
|
||||
if actual > 0:
|
||||
dst[:, :, dst_pos:dst_pos + actual, :] = (
|
||||
src[src_pos:src_pos + actual].transpose(0, 1).unsqueeze(0)
|
||||
)
|
||||
src_pos += actual
|
||||
dst_pos += block_size
|
||||
|
||||
|
||||
def _gather_from_padded(
|
||||
src: torch.Tensor,
|
||||
block_sizes: torch.Tensor,
|
||||
block_size: int,
|
||||
dst: torch.Tensor,
|
||||
src_offset: int,
|
||||
dst_start: int,
|
||||
dst_end: int,
|
||||
) -> None:
|
||||
"""Inverse of _scatter_to_padded: extract real tokens from padded blocks.
|
||||
|
||||
src: [1, H, total_padded, D]
|
||||
dst: [total_tokens, H, D]
|
||||
"""
|
||||
src_pos = src_offset
|
||||
dst_pos = dst_start
|
||||
sizes = block_sizes.cpu().tolist()
|
||||
for actual in sizes:
|
||||
actual = min(actual, dst_end - dst_pos)
|
||||
if actual > 0:
|
||||
dst[dst_pos:dst_pos + actual] = (
|
||||
src[0, :, src_pos:src_pos + actual, :].transpose(0, 1)
|
||||
)
|
||||
dst_pos += actual
|
||||
src_pos += block_size
|
||||
|
||||
|
||||
def block_sparse_attn_varlen(
|
||||
q: torch.Tensor,
|
||||
k: torch.Tensor,
|
||||
v: torch.Tensor,
|
||||
cu_seqlens_q: torch.Tensor,
|
||||
cu_seqlens_kv: torch.Tensor,
|
||||
q2k_idx_list: Sequence[torch.Tensor],
|
||||
q2k_num_list: Sequence[torch.Tensor],
|
||||
variable_block_sizes_list: Sequence[torch.Tensor],
|
||||
q_variable_block_sizes_list: Sequence[torch.Tensor] | None = None,
|
||||
block_size: int = BLOCK_SIZE,
|
||||
) -> torch.Tensor:
|
||||
"""Block-sparse attention over packed variable-length sequences.
|
||||
|
||||
Args:
|
||||
q: [total_q_tokens, H, D] packed query tensor.
|
||||
k: [total_kv_tokens, H, D] packed key tensor.
|
||||
v: [total_kv_tokens, H, D] packed value tensor.
|
||||
cu_seqlens_q: [N+1] int32, cumulative Q token offsets.
|
||||
cu_seqlens_kv: [N+1] int32, cumulative KV token offsets.
|
||||
q2k_idx_list: Per-sequence q2k_idx tensors, each [1, H, Nq_i, Mk].
|
||||
q2k_num_list: Per-sequence q2k_num tensors, each [1, H, Nq_i].
|
||||
variable_block_sizes_list: Per-sequence KV block sizes, each [Nkv_i].
|
||||
q_variable_block_sizes_list: Per-sequence Q block sizes, each [Nq_i].
|
||||
If None, each Q block is assumed to be exactly `block_size` tokens.
|
||||
block_size: Attention block size (default 64).
|
||||
|
||||
Returns:
|
||||
out: [total_q_tokens, H, D] packed output tensor.
|
||||
"""
|
||||
device = q.device
|
||||
dtype = q.dtype
|
||||
num_heads = q.shape[1]
|
||||
head_dim = q.shape[2]
|
||||
num_seqs = cu_seqlens_q.shape[0] - 1
|
||||
|
||||
cu_q = cu_seqlens_q.cpu().tolist()
|
||||
cu_kv = cu_seqlens_kv.cpu().tolist()
|
||||
|
||||
padded_q_lens = []
|
||||
padded_kv_lens = []
|
||||
q_block_offsets = [0]
|
||||
kv_block_offsets = [0]
|
||||
q_vbs_resolved = []
|
||||
|
||||
for i in range(num_seqs):
|
||||
n_q_blocks = q2k_num_list[i].shape[-1]
|
||||
n_kv_blocks = variable_block_sizes_list[i].numel()
|
||||
padded_q_lens.append(n_q_blocks * block_size)
|
||||
padded_kv_lens.append(n_kv_blocks * block_size)
|
||||
q_block_offsets.append(q_block_offsets[-1] + n_q_blocks)
|
||||
kv_block_offsets.append(kv_block_offsets[-1] + n_kv_blocks)
|
||||
|
||||
if q_variable_block_sizes_list is not None:
|
||||
q_vbs_resolved.append(q_variable_block_sizes_list[i])
|
||||
else:
|
||||
q_vbs_resolved.append(
|
||||
torch.full((n_q_blocks,), block_size, dtype=torch.int32)
|
||||
)
|
||||
|
||||
total_padded_q = sum(padded_q_lens)
|
||||
total_padded_kv = sum(padded_kv_lens)
|
||||
|
||||
q_packed = torch.zeros(1, num_heads, total_padded_q, head_dim, device=device, dtype=dtype)
|
||||
k_packed = torch.zeros(1, num_heads, total_padded_kv, head_dim, device=device, dtype=dtype)
|
||||
v_packed = torch.zeros(1, num_heads, total_padded_kv, head_dim, device=device, dtype=dtype)
|
||||
|
||||
q_offset = 0
|
||||
kv_offset = 0
|
||||
for i in range(num_seqs):
|
||||
_scatter_to_padded(
|
||||
q, q_vbs_resolved[i], block_size,
|
||||
q_packed, q_offset, cu_q[i], cu_q[i + 1],
|
||||
)
|
||||
_scatter_to_padded(
|
||||
k, variable_block_sizes_list[i], block_size,
|
||||
k_packed, kv_offset, cu_kv[i], cu_kv[i + 1],
|
||||
)
|
||||
_scatter_to_padded(
|
||||
v, variable_block_sizes_list[i], block_size,
|
||||
v_packed, kv_offset, cu_kv[i], cu_kv[i + 1],
|
||||
)
|
||||
q_offset += padded_q_lens[i]
|
||||
kv_offset += padded_kv_lens[i]
|
||||
|
||||
total_q_blocks = q_block_offsets[-1]
|
||||
max_kv_per_q = max(t.shape[-1] for t in q2k_idx_list)
|
||||
|
||||
global_q2k_idx = torch.zeros(
|
||||
1, num_heads, total_q_blocks, max_kv_per_q,
|
||||
dtype=torch.int32, device=device,
|
||||
)
|
||||
global_q2k_num = torch.zeros(
|
||||
1, num_heads, total_q_blocks,
|
||||
dtype=torch.int32, device=device,
|
||||
)
|
||||
global_vbs_parts = []
|
||||
|
||||
for i in range(num_seqs):
|
||||
qb_start = q_block_offsets[i]
|
||||
qb_end = q_block_offsets[i + 1]
|
||||
n_q_blocks = qb_end - qb_start
|
||||
kv_offset_blocks = kv_block_offsets[i]
|
||||
|
||||
idx = q2k_idx_list[i]
|
||||
num = q2k_num_list[i]
|
||||
vbs = variable_block_sizes_list[i]
|
||||
|
||||
mk = idx.shape[-1]
|
||||
global_q2k_idx[:, :, qb_start:qb_end, :mk] = idx[:, :, :n_q_blocks, :] + kv_offset_blocks
|
||||
global_q2k_num[:, :, qb_start:qb_end] = num[:, :, :n_q_blocks]
|
||||
global_vbs_parts.append(vbs)
|
||||
|
||||
global_vbs = torch.cat(global_vbs_parts, dim=0).to(torch.int32).contiguous()
|
||||
|
||||
out_packed, _ = block_sparse_attn_from_indices(
|
||||
q_packed, k_packed, v_packed,
|
||||
global_q2k_idx, global_q2k_num, global_vbs,
|
||||
)
|
||||
|
||||
out = torch.zeros(cu_q[-1], num_heads, head_dim, device=device, dtype=dtype)
|
||||
q_offset = 0
|
||||
for i in range(num_seqs):
|
||||
_gather_from_padded(
|
||||
out_packed, q_vbs_resolved[i], block_size,
|
||||
out, q_offset, cu_q[i], cu_q[i + 1],
|
||||
)
|
||||
q_offset += padded_q_lens[i]
|
||||
|
||||
return out
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,55 @@
|
||||
"""Compatibility shim for the legacy non-QAT Triton attention import path.
|
||||
|
||||
Historically callers imported
|
||||
``fastvideo_kernel.triton_kernels.fused_attention`` directly. The shared
|
||||
implementation now lives in ``attn_qat_train.py`` and is parameterized by the
|
||||
``IS_QAT`` flag. This module preserves the original public API for tests and
|
||||
downstream users while always dispatching to the non-QAT configuration.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import torch
|
||||
|
||||
from .attn_qat_train import attention as _attention
|
||||
|
||||
|
||||
def attention(
|
||||
q: torch.Tensor,
|
||||
k: torch.Tensor,
|
||||
v: torch.Tensor,
|
||||
causal: bool,
|
||||
sm_scale: float,
|
||||
warp_specialize: bool = True,
|
||||
) -> torch.Tensor:
|
||||
"""Run the shared Triton attention kernel in non-QAT mode."""
|
||||
use_qat_qkv_backward = True
|
||||
smooth_k = False
|
||||
is_qat = False
|
||||
two_level_quant_p = False
|
||||
fake_quant_p = False
|
||||
use_high_prec_o = False
|
||||
smooth_q = False
|
||||
use_global_sf_p = False
|
||||
use_global_sf_qkv = False
|
||||
|
||||
return _attention(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
causal,
|
||||
sm_scale,
|
||||
use_qat_qkv_backward,
|
||||
smooth_k,
|
||||
warp_specialize,
|
||||
is_qat,
|
||||
two_level_quant_p,
|
||||
fake_quant_p,
|
||||
use_high_prec_o,
|
||||
smooth_q,
|
||||
use_global_sf_p,
|
||||
use_global_sf_qkv,
|
||||
)
|
||||
|
||||
|
||||
__all__ = ["attention"]
|
||||
@@ -0,0 +1,237 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
# Adapted from https://github.com/triton-lang/triton/blob/main/python/triton_kernels/triton_kernels/numerics_details/mxfp_details/_upcast_from_mxfp.py
|
||||
# and https://github.com/triton-lang/triton/blob/main/python/triton_kernels/triton_kernels/numerics_details/mxfp_details/_downcast_to_mxfp.py
|
||||
|
||||
import triton
|
||||
import triton.language as tl
|
||||
from triton.language.target_info import cuda_capability_geq
|
||||
|
||||
MXFP_BLOCK_SIZE = tl.constexpr(16)
|
||||
|
||||
@triton.jit
|
||||
def _compute_quant_and_scale(
|
||||
src_tensor,
|
||||
valid_src_mask,
|
||||
mx_tensor_dtype: tl.constexpr = tl.uint8,
|
||||
use_global_sf=True,
|
||||
two_level_quant_P=False,
|
||||
):
|
||||
BLOCK_SIZE_OUT_DIM: tl.constexpr = src_tensor.shape[0]
|
||||
BLOCK_SIZE_QUANT_DIM: tl.constexpr = src_tensor.shape[1]
|
||||
BLOCK_SIZE_QUANT_MX_SCALE: tl.constexpr = src_tensor.shape[1] // MXFP_BLOCK_SIZE
|
||||
is_fp4: tl.constexpr = mx_tensor_dtype == tl.uint8
|
||||
|
||||
tl.static_assert(
|
||||
is_fp4
|
||||
or mx_tensor_dtype == tl.float8e4nv
|
||||
or mx_tensor_dtype == tl.float8e5,
|
||||
"mx_tensor_dtype must be uint8, float8e4nv, or float8e5",
|
||||
)
|
||||
|
||||
# Explicit cast to fp32 since most ops are not supported on bfloat16. We avoid needless conversions to and from bf16
|
||||
f32_tensor = src_tensor.to(tl.float32)
|
||||
abs_tensor = tl.abs(f32_tensor)
|
||||
abs_tensor = tl.where(valid_src_mask, abs_tensor, -1.0) # Don't consider padding tensors in scale computation
|
||||
|
||||
if two_level_quant_P:
|
||||
# row max from SageAttn3 paper
|
||||
global_max_val = tl.max(f32_tensor, axis=1, keep_dims=True) # (BLOCK_SIZE_OUT_DIM, 1)
|
||||
global_max_val = tl.maximum(global_max_val, 1e-8)
|
||||
s_enc = ((6 * 448) / global_max_val).reshape([BLOCK_SIZE_OUT_DIM, 1, 1])
|
||||
s_dec = (1 / s_enc)
|
||||
|
||||
abs_tensor = tl.reshape(abs_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
|
||||
|
||||
if use_global_sf and not two_level_quant_P:
|
||||
global_max_val = tl.max(abs_tensor)
|
||||
# Avoid division by zero: if all values are padding (max is 0), use a default scale
|
||||
global_max_val = tl.maximum(global_max_val, 1e-8)
|
||||
s_enc = (6 * 448) / global_max_val
|
||||
s_dec = (1 / s_enc)
|
||||
elif not two_level_quant_P and not use_global_sf:
|
||||
s_dec = 1.0
|
||||
s_enc = 1.0
|
||||
|
||||
max_val = tl.max(abs_tensor, axis=2, keep_dims=True) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1) # per block maxima
|
||||
s_dec_b = max_val / 6 # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
|
||||
s_dec_b_e4m3 = (s_dec_b * s_enc).to(tl.float8e4nv) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
|
||||
s_enc_b = 1 / (s_dec_b_e4m3.to(tl.float32) * s_dec) # (BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1)
|
||||
|
||||
f32_tensor = tl.reshape(f32_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
|
||||
quant_tensor = f32_tensor * s_enc_b
|
||||
|
||||
# Reshape the tensors after scaling
|
||||
quant_tensor = quant_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM])
|
||||
# Set the invalid portions of the tensor to 0. This will ensure that any padding tensors are 0 in the mx format.
|
||||
quant_tensor = tl.where(valid_src_mask, quant_tensor, 0.0)
|
||||
dequant_scale = s_dec_b_e4m3.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE])
|
||||
|
||||
if is_fp4 and cuda_capability_geq(10, 0):
|
||||
# Convert scaled values to two f32 lanes and use PTX cvt to e2m1x2 with two f32 operands.
|
||||
pairs = tl.reshape(quant_tensor, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM // 2, 2])
|
||||
lo_f, hi_f = tl.split(pairs)
|
||||
lo_f32 = lo_f.to(tl.float32)
|
||||
hi_f32 = hi_f.to(tl.float32)
|
||||
|
||||
# Inline PTX: cvt.rn.satfinite.e2m1x2.f32 takes two f32 sources and produces one .b8 packed e2m1x2.
|
||||
out_tensor = tl.inline_asm_elementwise(
|
||||
"""
|
||||
{
|
||||
.reg .b8 r;
|
||||
cvt.rn.satfinite.e2m1x2.f32 r, $1, $2;
|
||||
mov.b32 $0, {r, r, r, r};
|
||||
}
|
||||
""",
|
||||
constraints="=r,f,f",
|
||||
args=[hi_f32, lo_f32],
|
||||
dtype=tl.uint8,
|
||||
is_pure=True,
|
||||
pack=1,
|
||||
)
|
||||
elif is_fp4:
|
||||
quant_tensor = quant_tensor.to(tl.uint32, bitcast=True)
|
||||
signs = quant_tensor & 0x80000000
|
||||
exponents = (quant_tensor >> 23) & 0xFF
|
||||
mantissas_orig = (quant_tensor & 0x7FFFFF)
|
||||
|
||||
# For RTNE: 0.25 < x < 0.75 maps to 0.5 (denormal); exactly 0.25 maps to 0.0
|
||||
E8_BIAS = 127
|
||||
E2_BIAS = 1
|
||||
# Move implicit bit 1 at the beginning to mantissa for denormals
|
||||
is_subnormal = exponents < E8_BIAS
|
||||
adjusted_exponents = tl.core.sub(E8_BIAS, exponents + 1, sanitize_overflow=False)
|
||||
mantissas_pre = (0x400000 | (mantissas_orig >> 1))
|
||||
mantissas = tl.where(is_subnormal, mantissas_pre >> adjusted_exponents, mantissas_orig)
|
||||
|
||||
# For normal numbers, we change the bias from 127 to 1, and for subnormals, we keep exponent as 0.
|
||||
exponents = tl.maximum(exponents, E8_BIAS - E2_BIAS) - (E8_BIAS - E2_BIAS)
|
||||
|
||||
# Combine sign, exponent, and mantissa, while saturating
|
||||
# Round to nearest, ties to even (RTNE): use guard/sticky and LSB to decide increment
|
||||
m2bits = mantissas >> 21
|
||||
lsb_keep = (m2bits >> 1) & 0x1
|
||||
guard = m2bits & 0x1
|
||||
IS_SRC_FP32: tl.constexpr = src_tensor.dtype == tl.float32
|
||||
if IS_SRC_FP32:
|
||||
bit0_dropped = (mantissas_orig & 0x1) != 0
|
||||
mask = (1 << tl.minimum(adjusted_exponents, 31)) - 1
|
||||
dropped_post = (mantissas_pre & mask) != 0
|
||||
sticky = is_subnormal & (bit0_dropped | dropped_post)
|
||||
sticky |= ((mantissas & 0x1FFFFF) != 0).to(tl.uint32)
|
||||
else:
|
||||
sticky = ((mantissas & 0x1FFFFF) != 0).to(tl.uint32)
|
||||
round_inc = guard & (sticky | lsb_keep)
|
||||
e2m1_tmp = tl.minimum((((exponents << 2) | m2bits) + round_inc) >> 1, 0x7)
|
||||
e2m1_value = ((signs >> 28) | e2m1_tmp).to(tl.uint8)
|
||||
|
||||
e2m1_value = tl.reshape(e2m1_value, [BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM // 2, 2])
|
||||
evens, odds = tl.split(e2m1_value)
|
||||
out_tensor = evens | (odds << 4)
|
||||
else:
|
||||
out_tensor = quant_tensor.to(mx_tensor_dtype)
|
||||
|
||||
return out_tensor, dequant_scale, s_dec
|
||||
|
||||
@triton.jit
|
||||
def _compute_dequant(
|
||||
mx_tensor,
|
||||
scale,
|
||||
s_dec,
|
||||
BLOCK_SIZE_OUT_DIM: tl.constexpr,
|
||||
BLOCK_SIZE_QUANT_DIM: tl.constexpr,
|
||||
dst_dtype: tl.constexpr,
|
||||
):
|
||||
tl.static_assert(BLOCK_SIZE_QUANT_DIM % MXFP_BLOCK_SIZE == 0, f"Block size along quantization block must be a multiple of {MXFP_BLOCK_SIZE=}")
|
||||
# uint8 signifies two fp4 e2m1 values packed into a single byte
|
||||
mx_tensor_dtype: tl.constexpr = mx_tensor.dtype
|
||||
tl.static_assert(dst_dtype == tl.float16 or dst_dtype == tl.bfloat16 or dst_dtype == tl.float32)
|
||||
tl.static_assert(
|
||||
mx_tensor_dtype == tl.uint8
|
||||
or ((mx_tensor_dtype == tl.float8e4nv or mx_tensor_dtype == tl.float8e5) or mx_tensor_dtype == dst_dtype),
|
||||
"mx_tensor_ptr must be uint8 or float8 or dst_dtype")
|
||||
tl.static_assert(scale.dtype == tl.float8e4nv, "scale must be float8e4nv")
|
||||
|
||||
# Determine if we are dealing with fp8 types.
|
||||
is_fp4: tl.constexpr = mx_tensor_dtype == tl.uint8
|
||||
BLOCK_SIZE_QUANT_MX_SCALE: tl.constexpr = BLOCK_SIZE_QUANT_DIM // MXFP_BLOCK_SIZE
|
||||
|
||||
# Upcast the scale to the destination type.
|
||||
if dst_dtype == tl.bfloat16:
|
||||
dst_scale = scale.to(tl.bfloat16)
|
||||
else:
|
||||
dst_scale = scale.to(tl.float32)
|
||||
if dst_dtype == tl.float16:
|
||||
dst_scale = dst_scale.to(tl.float16)
|
||||
|
||||
# Now upcast the tensor.
|
||||
intermediate_dtype: tl.constexpr = tl.bfloat16 if dst_dtype == tl.float32 else dst_dtype
|
||||
if cuda_capability_geq(10, 0):
|
||||
assert is_fp4
|
||||
packed_u32 = tl.inline_asm_elementwise(
|
||||
asm="""
|
||||
{
|
||||
.reg .b8 in_8;
|
||||
.reg .f16x2 out;
|
||||
cvt.u8.u32 in_8, $1;
|
||||
cvt.rn.f16x2.e2m1x2 out, in_8;
|
||||
mov.b32 $0, out;
|
||||
}
|
||||
""",
|
||||
constraints="=r,r",
|
||||
args=[mx_tensor], # tl.uint8 passed in as a 32-bit reg with value in low 8 bits
|
||||
dtype=tl.uint32,
|
||||
is_pure=True,
|
||||
pack=1,
|
||||
)
|
||||
lo_u16 = (packed_u32 & 0xFFFF).to(tl.uint16)
|
||||
hi_u16 = (packed_u32 >> 16).to(tl.uint16)
|
||||
lo_f16 = lo_u16.to(tl.float16, bitcast=True)
|
||||
hi_f16 = hi_u16.to(tl.float16, bitcast=True)
|
||||
|
||||
if intermediate_dtype == tl.float16:
|
||||
x0, x1 = lo_f16, hi_f16
|
||||
else:
|
||||
x0 = lo_f16.to(intermediate_dtype)
|
||||
x1 = hi_f16.to(intermediate_dtype)
|
||||
|
||||
dst_tensor = tl.interleave(x0, x1)
|
||||
|
||||
else:
|
||||
assert is_fp4
|
||||
dst_bias: tl.constexpr = 127 if intermediate_dtype == tl.bfloat16 else 15 # exponent bias
|
||||
dst_0p5: tl.constexpr = 16128 if intermediate_dtype == tl.bfloat16 else 0x3800
|
||||
dst_m_bits: tl.constexpr = 7 if intermediate_dtype == tl.bfloat16 else 10 # mantissa bits
|
||||
# e2m1
|
||||
em0 = mx_tensor & 0x07
|
||||
em1 = mx_tensor & 0x70
|
||||
x0 = (em0.to(tl.uint16) << (dst_m_bits - 1)) | ((mx_tensor & 0x08).to(tl.uint16) << 12)
|
||||
x1 = (em1.to(tl.uint16) << (dst_m_bits - 5)) | ((mx_tensor & 0x80).to(tl.uint16) << 8)
|
||||
# Three cases:
|
||||
# 1) x is normal and non-zero: Correct bias
|
||||
x0 = tl.where((em0 & 0x06) != 0, x0 + ((dst_bias - 1) << dst_m_bits), x0)
|
||||
x1 = tl.where((em1 & 0x60) != 0, x1 + ((dst_bias - 1) << dst_m_bits), x1)
|
||||
# 2) x is subnormal (x == 0bs001 where s is the sign): Map to +-0.5 in the dst type
|
||||
x0 = tl.where(em0 == 0x01, dst_0p5 | (x0 & 0x8000), x0)
|
||||
x1 = tl.where(em1 == 0x10, dst_0p5 | (x1 & 0x8000), x1)
|
||||
# 3) x is zero, do nothing
|
||||
dst_tensor = tl.interleave(x0, x1).to(intermediate_dtype, bitcast=True)
|
||||
|
||||
dst_tensor = dst_tensor.to(dst_dtype)
|
||||
|
||||
# Reshape for proper broadcasting: the scale was stored with a 16‐sized “inner” grouping.
|
||||
dst_tensor = dst_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, MXFP_BLOCK_SIZE])
|
||||
dst_scale = dst_scale.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_MX_SCALE, 1])
|
||||
scale = scale.reshape(dst_scale.shape)
|
||||
|
||||
out_tensor = dst_tensor * dst_scale * s_dec # NVFP4 has the additional global scale factor
|
||||
if dst_dtype == tl.float32:
|
||||
max_fin = 3.4028234663852886e+38
|
||||
elif dst_dtype == tl.bfloat16:
|
||||
max_fin = 3.3895313892515355e+38
|
||||
else:
|
||||
tl.static_assert(dst_dtype == tl.float16)
|
||||
max_fin = 65504
|
||||
out_tensor = tl.clamp(out_tensor, min=-max_fin, max=max_fin)
|
||||
out_tensor = out_tensor.reshape([BLOCK_SIZE_OUT_DIM, BLOCK_SIZE_QUANT_DIM])
|
||||
out_tensor = out_tensor.to(dst_dtype)
|
||||
return out_tensor
|
||||
@@ -0,0 +1,80 @@
|
||||
import triton
|
||||
import triton.language as tl
|
||||
|
||||
from .nvfp4_utils import _compute_quant_and_scale, _compute_dequant
|
||||
|
||||
@triton.jit
|
||||
def fake_quantize(src_tensor, valid_src_mask, BLOCK_SIZE_OUT_DIM: tl.constexpr,
|
||||
BLOCK_SIZE_QUANT_DIM: tl.constexpr,
|
||||
dst_dtype: tl.constexpr,
|
||||
mx_tensor_dtype: tl.constexpr = tl.uint8,
|
||||
use_global_sf: tl.constexpr = True,
|
||||
two_level_quant_P: tl.constexpr = False):
|
||||
high_prec_src_tensor = src_tensor
|
||||
src_tensor, src_scale, src_s_dec = _compute_quant_and_scale(src_tensor=src_tensor,
|
||||
valid_src_mask=valid_src_mask,
|
||||
mx_tensor_dtype=mx_tensor_dtype,
|
||||
use_global_sf=use_global_sf,
|
||||
two_level_quant_P=two_level_quant_P)
|
||||
src_tensor = _compute_dequant(mx_tensor=src_tensor,
|
||||
scale=src_scale,
|
||||
s_dec=src_s_dec,
|
||||
BLOCK_SIZE_OUT_DIM=BLOCK_SIZE_OUT_DIM,
|
||||
BLOCK_SIZE_QUANT_DIM=BLOCK_SIZE_QUANT_DIM,
|
||||
dst_dtype=dst_dtype)
|
||||
return src_tensor, high_prec_src_tensor.to(src_tensor.dtype)
|
||||
|
||||
@triton.jit
|
||||
def fake_quantize_q(Q, fake_Q, stride_z_q, stride_h_q,
|
||||
stride_tok_q, stride_d_q,
|
||||
fake_stride_z_q, fake_stride_h_q,
|
||||
fake_stride_tok_q, fake_stride_d_q,
|
||||
H, N_CTX_Q,
|
||||
BLOCK_M: tl.constexpr,
|
||||
HEAD_DIM: tl.constexpr,
|
||||
use_global_sf: tl.constexpr = True):
|
||||
bhid = tl.program_id(1)
|
||||
adj_q = (stride_h_q * (bhid % H) + stride_z_q * (bhid // H))
|
||||
fake_adj_q = (fake_stride_h_q * (bhid % H) + fake_stride_z_q * (bhid // H))
|
||||
Q += adj_q
|
||||
fake_Q += fake_adj_q
|
||||
|
||||
pid = tl.program_id(0)
|
||||
start_m = pid * BLOCK_M
|
||||
offs_m = start_m + tl.arange(0, BLOCK_M)
|
||||
offs_k = tl.arange(0, HEAD_DIM)
|
||||
|
||||
q_valid = offs_m < N_CTX_Q
|
||||
q = tl.load(Q + offs_m[:, None] * stride_tok_q + offs_k[None, :] * stride_d_q, mask=q_valid[:, None], other=0.0)
|
||||
q, _ = fake_quantize(src_tensor=q, valid_src_mask=q_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_M, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=q.dtype, use_global_sf=use_global_sf)
|
||||
tl.store(fake_Q + offs_m[:, None] * fake_stride_tok_q + offs_k[None, :] * fake_stride_d_q, q, mask=q_valid[:, None])
|
||||
|
||||
@triton.jit
|
||||
def fake_quantize_kv(K, V, fake_K, fake_V, stride_z_kv, stride_h_kv,
|
||||
stride_tok_kv, stride_d_kv,
|
||||
fake_stride_z_kv, fake_stride_h_kv,
|
||||
fake_stride_tok_kv, fake_stride_d_kv,
|
||||
H, N_CTX_KV,
|
||||
BLOCK_N: tl.constexpr,
|
||||
HEAD_DIM: tl.constexpr,
|
||||
use_global_sf: tl.constexpr = True):
|
||||
bhid = tl.program_id(1)
|
||||
adj_kv = (stride_h_kv * (bhid % H) + stride_z_kv * (bhid // H))
|
||||
fake_adj_kv = (fake_stride_h_kv * (bhid % H) + fake_stride_z_kv * (bhid // H))
|
||||
K += adj_kv
|
||||
V += adj_kv
|
||||
fake_K += fake_adj_kv
|
||||
fake_V += fake_adj_kv
|
||||
|
||||
pid = tl.program_id(0)
|
||||
start_n = pid * BLOCK_N
|
||||
offs_n = start_n + tl.arange(0, BLOCK_N)
|
||||
offs_k = tl.arange(0, HEAD_DIM)
|
||||
|
||||
kv_valid = offs_n < N_CTX_KV
|
||||
k_block = tl.load(K + offs_n[:, None] * stride_tok_kv + offs_k[None, :] * stride_d_kv, mask=kv_valid[:, None], other=0.0)
|
||||
v_block = tl.load(V + offs_n[:, None] * stride_tok_kv + offs_k[None, :] * stride_d_kv, mask=kv_valid[:, None], other=0.0)
|
||||
k, _ = fake_quantize(src_tensor=k_block, valid_src_mask=kv_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_N, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=k_block.dtype, use_global_sf=use_global_sf)
|
||||
v, _ = fake_quantize(src_tensor=v_block, valid_src_mask=kv_valid[:, None], BLOCK_SIZE_OUT_DIM=BLOCK_N, BLOCK_SIZE_QUANT_DIM=HEAD_DIM, dst_dtype=v_block.dtype, use_global_sf=use_global_sf)
|
||||
tl.store(fake_K + offs_n[:, None] * fake_stride_tok_kv + offs_k[None, :] * fake_stride_d_kv, k, mask=kv_valid[:, None])
|
||||
tl.store(fake_V + offs_n[:, None] * fake_stride_tok_kv + offs_k[None, :] * fake_stride_d_kv, v, mask=kv_valid[:, None])
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,434 @@
|
||||
"""Correctness tests for variable-length block-sparse attention.
|
||||
|
||||
Reference: per-sequence calls to block_sparse_attn_from_indices.
|
||||
Test: single-launch via block_sparse_attn_varlen.
|
||||
Tests cover both forward and backward (gradient) correctness.
|
||||
"""
|
||||
|
||||
import torch
|
||||
import pytest
|
||||
|
||||
from .test_vsa import (
|
||||
BLOCK_M,
|
||||
generate_variable_block_sizes,
|
||||
get_non_pad_index,
|
||||
vsa_pad,
|
||||
generate_tensor,
|
||||
)
|
||||
from .utils import generate_block_sparse_mask_for_function
|
||||
from fastvideo_kernel.block_sparse_attn import (
|
||||
block_sparse_attn_from_indices,
|
||||
_map_to_index,
|
||||
)
|
||||
from fastvideo_kernel.block_sparse_attn_varlen import block_sparse_attn_varlen
|
||||
|
||||
|
||||
def _reference_per_sequence(
|
||||
q_list, k_list, v_list,
|
||||
block_masks, vbs_list,
|
||||
non_pad_q_list, non_pad_kv_list,
|
||||
q_nblocks_list, kv_nblocks_list,
|
||||
):
|
||||
"""Run per-sequence block_sparse_attn and concat outputs."""
|
||||
outs = []
|
||||
for i in range(len(q_list)):
|
||||
q_pad = vsa_pad(q_list[i], non_pad_q_list[i], q_nblocks_list[i], BLOCK_M)
|
||||
k_pad = vsa_pad(k_list[i], non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
|
||||
v_pad = vsa_pad(v_list[i], non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
|
||||
|
||||
q2k_idx, q2k_num = _map_to_index(block_masks[i].unsqueeze(0))
|
||||
o_pad, _ = block_sparse_attn_from_indices(
|
||||
q_pad, k_pad, v_pad, q2k_idx, q2k_num, vbs_list[i],
|
||||
)
|
||||
o = o_pad[:, :, non_pad_q_list[i], :]
|
||||
outs.append(o.squeeze(0).transpose(0, 1))
|
||||
return torch.cat(outs, dim=0)
|
||||
|
||||
|
||||
def _run_varlen_test(
|
||||
seq_configs: list,
|
||||
h: int = 8,
|
||||
d: int = 64,
|
||||
topk: int = 2,
|
||||
atol: float = 0.05,
|
||||
rtol: float = 0.02,
|
||||
):
|
||||
"""Core test: compare varlen vs per-sequence reference.
|
||||
|
||||
seq_configs: list of (num_q_blocks, num_kv_blocks) per sequence.
|
||||
"""
|
||||
device = "cuda"
|
||||
num_seqs = len(seq_configs)
|
||||
|
||||
q_list = []
|
||||
k_list = []
|
||||
v_list = []
|
||||
block_masks = []
|
||||
vbs_list = []
|
||||
q_vbs_list = []
|
||||
non_pad_q_list = []
|
||||
non_pad_kv_list = []
|
||||
q_nblocks_list = []
|
||||
kv_nblocks_list = []
|
||||
q2k_idx_list = []
|
||||
q2k_num_list = []
|
||||
q_vbs_for_varlen = []
|
||||
|
||||
cu_q = [0]
|
||||
cu_kv = [0]
|
||||
|
||||
for nq, nkv in seq_configs:
|
||||
vbs_kv = generate_variable_block_sizes(nkv, device=device)
|
||||
vbs_q = generate_variable_block_sizes(nq, device=device)
|
||||
sq = int(vbs_q.sum().item())
|
||||
skv = int(vbs_kv.sum().item())
|
||||
|
||||
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
|
||||
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
|
||||
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
|
||||
npq = get_non_pad_index(vbs_q, nq, BLOCK_M)
|
||||
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
|
||||
|
||||
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
|
||||
|
||||
q_list.append(q)
|
||||
k_list.append(k)
|
||||
v_list.append(v)
|
||||
block_masks.append(mask)
|
||||
vbs_list.append(vbs_kv)
|
||||
q_vbs_list.append(vbs_q)
|
||||
non_pad_q_list.append(npq)
|
||||
non_pad_kv_list.append(npkv)
|
||||
q_nblocks_list.append(nq)
|
||||
kv_nblocks_list.append(nkv)
|
||||
q2k_idx_list.append(q2k_idx)
|
||||
q2k_num_list.append(q2k_num)
|
||||
q_vbs_for_varlen.append(vbs_q)
|
||||
|
||||
cu_q.append(cu_q[-1] + sq)
|
||||
cu_kv.append(cu_kv[-1] + skv)
|
||||
|
||||
ref_out = _reference_per_sequence(
|
||||
q_list, k_list, v_list,
|
||||
block_masks, vbs_list,
|
||||
non_pad_q_list, non_pad_kv_list,
|
||||
q_nblocks_list, kv_nblocks_list,
|
||||
)
|
||||
|
||||
q_packed = torch.cat(
|
||||
[qi.squeeze(0).transpose(0, 1) for qi in q_list], dim=0,
|
||||
)
|
||||
k_packed = torch.cat(
|
||||
[ki.squeeze(0).transpose(0, 1) for ki in k_list], dim=0,
|
||||
)
|
||||
v_packed = torch.cat(
|
||||
[vi.squeeze(0).transpose(0, 1) for vi in v_list], dim=0,
|
||||
)
|
||||
|
||||
cu_seqlens_q = torch.tensor(cu_q, dtype=torch.int32, device=device)
|
||||
cu_seqlens_kv = torch.tensor(cu_kv, dtype=torch.int32, device=device)
|
||||
|
||||
varlen_out = block_sparse_attn_varlen(
|
||||
q_packed, k_packed, v_packed,
|
||||
cu_seqlens_q, cu_seqlens_kv,
|
||||
q2k_idx_list, q2k_num_list,
|
||||
vbs_list,
|
||||
q_variable_block_sizes_list=q_vbs_for_varlen,
|
||||
)
|
||||
|
||||
max_abs = (ref_out - varlen_out).abs().max().item()
|
||||
mean_abs = ref_out.abs().mean().item()
|
||||
max_rel = max_abs / (mean_abs + 1e-8)
|
||||
|
||||
print(f" seqs={[c for c in seq_configs]}, max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
|
||||
assert max_rel < rtol, f"max relative error {max_rel:.4e} exceeds threshold {rtol}"
|
||||
|
||||
|
||||
@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required")
|
||||
class TestVSAVarlen:
|
||||
|
||||
def test_equal_length(self):
|
||||
"""Two sequences with same number of blocks."""
|
||||
_run_varlen_test([(4, 4), (4, 4)], h=8, d=64)
|
||||
|
||||
def test_different_lengths(self):
|
||||
"""Three sequences with different block counts."""
|
||||
_run_varlen_test([(2, 3), (5, 4), (3, 6)], h=8, d=64)
|
||||
|
||||
def test_single_sequence(self):
|
||||
"""Degenerate case: single sequence should match non-varlen path."""
|
||||
_run_varlen_test([(8, 8)], h=8, d=64)
|
||||
|
||||
def test_many_heads(self):
|
||||
"""More heads to stress the packing logic."""
|
||||
_run_varlen_test([(3, 4), (5, 3)], h=16, d=128)
|
||||
|
||||
def test_many_sequences(self):
|
||||
"""Stress test: 8 sequences with varying block counts."""
|
||||
configs = [(i + 2, i + 3) for i in range(8)]
|
||||
_run_varlen_test(configs, h=8, d=64)
|
||||
|
||||
def test_topk_equals_num_blocks(self):
|
||||
"""Edge: topk covers all KV blocks (dense attention)."""
|
||||
_run_varlen_test([(3, 3), (4, 4)], h=8, d=64, topk=8)
|
||||
|
||||
def test_single_block_per_sequence(self):
|
||||
"""Minimal: each sequence has exactly 1 Q block and 1 KV block."""
|
||||
_run_varlen_test([(1, 1), (1, 1), (1, 1)], h=8, d=64, topk=1)
|
||||
|
||||
def test_asymmetric_q_kv(self):
|
||||
"""Q and KV have very different block counts."""
|
||||
_run_varlen_test([(1, 8), (8, 1)], h=8, d=64, topk=1)
|
||||
|
||||
def test_without_q_vbs(self):
|
||||
"""Test the default path where q_variable_block_sizes_list is None.
|
||||
|
||||
Uses full block_size=64 for Q blocks so the None path is valid.
|
||||
"""
|
||||
device = "cuda"
|
||||
h, d, topk = 4, 64, 2
|
||||
nq, nkv = 3, 4
|
||||
|
||||
vbs_kv = generate_variable_block_sizes(nkv, device=device)
|
||||
sq = nq * BLOCK_M
|
||||
skv = int(vbs_kv.sum().item())
|
||||
|
||||
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
|
||||
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
|
||||
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
|
||||
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
|
||||
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
|
||||
|
||||
k_pad = vsa_pad(k, npkv, nkv, BLOCK_M)
|
||||
v_pad = vsa_pad(v, npkv, nkv, BLOCK_M)
|
||||
ref_out, _ = block_sparse_attn_from_indices(q, k_pad, v_pad, q2k_idx, q2k_num, vbs_kv)
|
||||
ref_flat = ref_out.squeeze(0).transpose(0, 1)
|
||||
|
||||
q_flat = q.squeeze(0).transpose(0, 1)
|
||||
k_flat = k.squeeze(0).transpose(0, 1)
|
||||
v_flat = v.squeeze(0).transpose(0, 1)
|
||||
cu_q = torch.tensor([0, sq], dtype=torch.int32, device=device)
|
||||
cu_kv = torch.tensor([0, skv], dtype=torch.int32, device=device)
|
||||
|
||||
varlen_out = block_sparse_attn_varlen(
|
||||
q_flat, k_flat, v_flat,
|
||||
cu_q, cu_kv,
|
||||
[q2k_idx], [q2k_num], [vbs_kv],
|
||||
)
|
||||
|
||||
max_abs = (ref_flat - varlen_out).abs().max().item()
|
||||
mean_abs = ref_flat.abs().mean().item()
|
||||
max_rel = max_abs / (mean_abs + 1e-8)
|
||||
print(f" without_q_vbs: max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
|
||||
assert max_rel < 0.02, f"max relative error {max_rel:.4e} exceeds threshold"
|
||||
|
||||
|
||||
def _run_varlen_backward_test(
|
||||
seq_configs: list,
|
||||
h: int = 8,
|
||||
d: int = 64,
|
||||
topk: int = 2,
|
||||
grad_rtol: float = 0.05,
|
||||
):
|
||||
"""Backward correctness: compare dQ/dK/dV from varlen vs per-sequence reference.
|
||||
|
||||
Both paths use the same underlying block_sparse_attn_from_indices kernel
|
||||
(which has registered autograd). The varlen wrapper's scatter/gather must
|
||||
correctly propagate gradients through PyTorch's in-place slice assignment.
|
||||
"""
|
||||
device = "cuda"
|
||||
num_seqs = len(seq_configs)
|
||||
|
||||
q_list = []
|
||||
k_list = []
|
||||
v_list = []
|
||||
block_masks = []
|
||||
vbs_list = []
|
||||
q_vbs_list = []
|
||||
non_pad_q_list = []
|
||||
non_pad_kv_list = []
|
||||
q_nblocks_list = []
|
||||
kv_nblocks_list = []
|
||||
q2k_idx_list = []
|
||||
q2k_num_list = []
|
||||
|
||||
cu_q = [0]
|
||||
cu_kv = [0]
|
||||
|
||||
for nq, nkv in seq_configs:
|
||||
vbs_kv = generate_variable_block_sizes(nkv, device=device)
|
||||
vbs_q = generate_variable_block_sizes(nq, device=device)
|
||||
sq = int(vbs_q.sum().item())
|
||||
skv = int(vbs_kv.sum().item())
|
||||
|
||||
q = generate_tensor((1, h, sq, d), torch.bfloat16, device)
|
||||
k = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
v = generate_tensor((1, h, skv, d), torch.bfloat16, device)
|
||||
|
||||
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
|
||||
npq = get_non_pad_index(vbs_q, nq, BLOCK_M)
|
||||
npkv = get_non_pad_index(vbs_kv, nkv, BLOCK_M)
|
||||
|
||||
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
|
||||
|
||||
q_list.append(q)
|
||||
k_list.append(k)
|
||||
v_list.append(v)
|
||||
block_masks.append(mask)
|
||||
vbs_list.append(vbs_kv)
|
||||
q_vbs_list.append(vbs_q)
|
||||
non_pad_q_list.append(npq)
|
||||
non_pad_kv_list.append(npkv)
|
||||
q_nblocks_list.append(nq)
|
||||
kv_nblocks_list.append(nkv)
|
||||
q2k_idx_list.append(q2k_idx)
|
||||
q2k_num_list.append(q2k_num)
|
||||
|
||||
cu_q.append(cu_q[-1] + sq)
|
||||
cu_kv.append(cu_kv[-1] + skv)
|
||||
|
||||
# --- Reference: per-sequence backward ---
|
||||
ref_q_grads = []
|
||||
ref_k_grads = []
|
||||
ref_v_grads = []
|
||||
ref_outs = []
|
||||
for i in range(num_seqs):
|
||||
qi = q_list[i].detach().requires_grad_(True)
|
||||
ki = k_list[i].detach().requires_grad_(True)
|
||||
vi = v_list[i].detach().requires_grad_(True)
|
||||
|
||||
q_pad = vsa_pad(qi, non_pad_q_list[i], q_nblocks_list[i], BLOCK_M)
|
||||
k_pad = vsa_pad(ki, non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
|
||||
v_pad = vsa_pad(vi, non_pad_kv_list[i], kv_nblocks_list[i], BLOCK_M)
|
||||
|
||||
q2k_idx, q2k_num = _map_to_index(block_masks[i].unsqueeze(0))
|
||||
o_pad, _ = block_sparse_attn_from_indices(
|
||||
q_pad, k_pad, v_pad, q2k_idx, q2k_num, vbs_list[i],
|
||||
)
|
||||
o = o_pad[:, :, non_pad_q_list[i], :]
|
||||
o_flat = o.squeeze(0).transpose(0, 1)
|
||||
ref_outs.append(o_flat)
|
||||
|
||||
dO = torch.ones_like(o_flat)
|
||||
o_flat.backward(dO)
|
||||
|
||||
ref_q_grads.append(qi.grad.squeeze(0).transpose(0, 1))
|
||||
ref_k_grads.append(ki.grad.squeeze(0).transpose(0, 1))
|
||||
ref_v_grads.append(vi.grad.squeeze(0).transpose(0, 1))
|
||||
|
||||
ref_dq = torch.cat(ref_q_grads, dim=0)
|
||||
ref_dk = torch.cat(ref_k_grads, dim=0)
|
||||
ref_dv = torch.cat(ref_v_grads, dim=0)
|
||||
|
||||
# --- Varlen backward ---
|
||||
q_packed = torch.cat(
|
||||
[qi.squeeze(0).transpose(0, 1) for qi in q_list], dim=0,
|
||||
).detach().requires_grad_(True)
|
||||
k_packed = torch.cat(
|
||||
[ki.squeeze(0).transpose(0, 1) for ki in k_list], dim=0,
|
||||
).detach().requires_grad_(True)
|
||||
v_packed = torch.cat(
|
||||
[vi.squeeze(0).transpose(0, 1) for vi in v_list], dim=0,
|
||||
).detach().requires_grad_(True)
|
||||
|
||||
cu_seqlens_q = torch.tensor(cu_q, dtype=torch.int32, device=device)
|
||||
cu_seqlens_kv = torch.tensor(cu_kv, dtype=torch.int32, device=device)
|
||||
|
||||
varlen_out = block_sparse_attn_varlen(
|
||||
q_packed, k_packed, v_packed,
|
||||
cu_seqlens_q, cu_seqlens_kv,
|
||||
q2k_idx_list, q2k_num_list,
|
||||
vbs_list,
|
||||
q_variable_block_sizes_list=q_vbs_list,
|
||||
)
|
||||
|
||||
dO = torch.ones_like(varlen_out)
|
||||
varlen_out.backward(dO)
|
||||
|
||||
varlen_dq = q_packed.grad
|
||||
varlen_dk = k_packed.grad
|
||||
varlen_dv = v_packed.grad
|
||||
|
||||
for name, ref, actual in [
|
||||
("dQ", ref_dq, varlen_dq),
|
||||
("dK", ref_dk, varlen_dk),
|
||||
("dV", ref_dv, varlen_dv),
|
||||
]:
|
||||
assert actual is not None, f"{name}: gradient is None (autograd chain broken)"
|
||||
max_abs = (ref - actual).abs().max().item()
|
||||
mean_abs = ref.abs().mean().item()
|
||||
max_rel = max_abs / (mean_abs + 1e-8)
|
||||
print(f" {name}: max_abs={max_abs:.4e}, max_rel={max_rel:.4e}")
|
||||
assert max_rel < grad_rtol, (
|
||||
f"{name}: max relative error {max_rel:.4e} exceeds threshold {grad_rtol}"
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required")
|
||||
class TestVSAVarlenBackward:
|
||||
|
||||
def test_backward_equal_length(self):
|
||||
"""Backward: two sequences with same number of blocks."""
|
||||
_run_varlen_backward_test([(4, 4), (4, 4)], h=8, d=64)
|
||||
|
||||
def test_backward_different_lengths(self):
|
||||
"""Backward: three sequences with different block counts."""
|
||||
_run_varlen_backward_test([(2, 3), (5, 4), (3, 6)], h=8, d=64)
|
||||
|
||||
def test_backward_single_sequence(self):
|
||||
"""Backward: single sequence should match non-varlen gradient path."""
|
||||
_run_varlen_backward_test([(8, 8)], h=8, d=64)
|
||||
|
||||
def test_backward_many_heads(self):
|
||||
"""Backward: more heads to stress gradient routing."""
|
||||
_run_varlen_backward_test([(3, 4), (5, 3)], h=16, d=128)
|
||||
|
||||
def test_backward_asymmetric_q_kv(self):
|
||||
"""Backward: Q and KV have very different block counts."""
|
||||
_run_varlen_backward_test([(1, 8), (8, 1)], h=8, d=64, topk=1)
|
||||
|
||||
def test_backward_grad_nonzero(self):
|
||||
"""Smoke test: gradients are non-zero (autograd chain is connected)."""
|
||||
device = "cuda"
|
||||
h, d, topk = 4, 64, 2
|
||||
nq, nkv = 3, 4
|
||||
|
||||
vbs_kv = generate_variable_block_sizes(nkv, device=device)
|
||||
vbs_q = generate_variable_block_sizes(nq, device=device)
|
||||
sq = int(vbs_q.sum().item())
|
||||
skv = int(vbs_kv.sum().item())
|
||||
|
||||
q = torch.randn(sq, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
|
||||
k = torch.randn(skv, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
|
||||
v = torch.randn(skv, h, d, device=device, dtype=torch.bfloat16, requires_grad=True)
|
||||
|
||||
mask = generate_block_sparse_mask_for_function(h, nq, nkv, topk, device)
|
||||
q2k_idx, q2k_num = _map_to_index(mask.unsqueeze(0))
|
||||
|
||||
cu_q = torch.tensor([0, sq], dtype=torch.int32, device=device)
|
||||
cu_kv = torch.tensor([0, skv], dtype=torch.int32, device=device)
|
||||
|
||||
out = block_sparse_attn_varlen(
|
||||
q, k, v,
|
||||
cu_q, cu_kv,
|
||||
[q2k_idx], [q2k_num], [vbs_kv],
|
||||
q_variable_block_sizes_list=[vbs_q],
|
||||
)
|
||||
|
||||
loss = out.sum()
|
||||
loss.backward()
|
||||
|
||||
assert q.grad is not None, "q.grad is None"
|
||||
assert k.grad is not None, "k.grad is None"
|
||||
assert v.grad is not None, "v.grad is None"
|
||||
assert q.grad.abs().sum().item() > 0, "q.grad is all zeros"
|
||||
assert k.grad.abs().sum().item() > 0, "k.grad is all zeros"
|
||||
assert v.grad.abs().sum().item() > 0, "v.grad is all zeros"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
pytest.main([__file__, "-v", "-s"])
|
||||
@@ -29,6 +29,10 @@ class SamplingParam:
|
||||
# Video inputs
|
||||
video_path: str | None = None
|
||||
|
||||
# Optional pre-generated diffusion latents. Used by parity/debug harnesses
|
||||
# and advanced callers that need deterministic latent reuse.
|
||||
latents: Any | None = None
|
||||
|
||||
# Action control inputs (Matrix-Game)
|
||||
mouse_cond: Any | None = None # Shape: (B, T, 2)
|
||||
keyboard_cond: Any | None = None # Shape: (B, T, K)
|
||||
@@ -64,6 +68,7 @@ class SamplingParam:
|
||||
# Text inputs
|
||||
prompt: str | list[str] | None = None
|
||||
negative_prompt: str = "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
|
||||
max_sequence_length: int | None = None
|
||||
prompt_path: str | None = None
|
||||
output_path: str = "outputs/"
|
||||
output_video_name: str | None = None
|
||||
|
||||
@@ -49,6 +49,10 @@ def _get_attn_qat_train_attention() -> Callable[..., torch.Tensor] | None:
|
||||
return _attn_qat_train_attention
|
||||
|
||||
|
||||
def is_attn_qat_train_available() -> bool:
|
||||
return _get_attn_qat_train_attention() is not None
|
||||
|
||||
|
||||
def attn_qat_train(q_BLHD: torch.Tensor,
|
||||
k_BLHD: torch.Tensor,
|
||||
v_BLHD: torch.Tensor,
|
||||
|
||||
@@ -150,14 +150,11 @@ class VideoSparseAttentionMetadata(AttentionMetadata):
|
||||
# in postprocess_output(). Avoids materializing the intermediate
|
||||
# ``[B, len(non_pad_index), H, D]`` tensor on every layer.
|
||||
untile_combined_index: torch.LongTensor
|
||||
# Per-step shared padded buffer used by tile(). Lazily populated on
|
||||
# the first layer's call and reused by every subsequent VSA layer in
|
||||
# the same denoising step. Scoping to metadata (not class/instance)
|
||||
# makes the reuse thread-safe across concurrent requests and keeps
|
||||
# the "pad positions are zero" invariant trivially true (the buffer
|
||||
# is freshly zeroed alongside ``non_pad_index`` so the index set
|
||||
# cannot drift between calls).
|
||||
# Per-step shared padded buffer used by tile(). Inference can reuse this
|
||||
# across VSA layers, but training disables it so activation checkpointing
|
||||
# can release the large tiled QKVG scratch tensor after each attention call.
|
||||
tile_buf: torch.Tensor | None = None
|
||||
cache_tile_buf: bool = True
|
||||
|
||||
|
||||
class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
|
||||
@@ -175,6 +172,7 @@ class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
|
||||
patch_size: tuple[int, int, int],
|
||||
VSA_sparsity: float,
|
||||
device: torch.device,
|
||||
cache_tile_buf: bool = True,
|
||||
**kwargs: dict[str, Any],
|
||||
) -> VideoSparseAttentionMetadata:
|
||||
patch_size = patch_size
|
||||
@@ -201,7 +199,8 @@ class VideoSparseAttentionMetadataBuilder(AttentionMetadataBuilder):
|
||||
reverse_tile_partition_indices=reverse_tile_partition_indices,
|
||||
variable_block_sizes=variable_block_sizes,
|
||||
non_pad_index=non_pad_index,
|
||||
untile_combined_index=untile_combined_index)
|
||||
untile_combined_index=untile_combined_index,
|
||||
cache_tile_buf=cache_tile_buf)
|
||||
|
||||
|
||||
class VideoSparseAttentionImpl(AttentionImpl):
|
||||
@@ -237,6 +236,11 @@ class VideoSparseAttentionImpl(AttentionImpl):
|
||||
w_padded_size = num_tiles[2] * VSA_TILE_SIZE[2]
|
||||
target_shape = (x.shape[0], t_padded_size * h_padded_size * w_padded_size, x.shape[-2], x.shape[-1])
|
||||
|
||||
if not attn_metadata.cache_tile_buf:
|
||||
buf = torch.zeros(target_shape, device=x.device, dtype=x.dtype)
|
||||
buf[:, attn_metadata.non_pad_index] = x[:, attn_metadata.tile_partition_indices]
|
||||
return buf
|
||||
|
||||
# Reuse the per-step buffer stashed on metadata (lazily allocated
|
||||
# on the first VSA layer's call within a denoising step). Pad
|
||||
# positions are zero from the initial torch.zeros and never
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
from fastvideo.configs.models.dits.cosmos import CosmosVideoConfig
|
||||
from fastvideo.configs.models.dits.cosmos2_5 import Cosmos25VideoConfig
|
||||
from fastvideo.configs.models.dits.flux_2 import Flux2Config
|
||||
from fastvideo.configs.models.dits.hunyuangamecraft import HunyuanGameCraftConfig
|
||||
from fastvideo.configs.models.dits.hunyuanvideo import HunyuanVideoConfig
|
||||
from fastvideo.configs.models.dits.hunyuanvideo15 import HunyuanVideo15Config
|
||||
@@ -14,5 +15,5 @@ from fastvideo.configs.models.dits.kandinsky5 import Kandinsky5VideoConfig
|
||||
__all__ = [
|
||||
"HunyuanVideoConfig", "HunyuanVideo15Config", "HunyuanGameCraftConfig", "WanVideoConfig", "CosmosVideoConfig",
|
||||
"Cosmos25VideoConfig", "LongCatVideoConfig", "LTX2VideoConfig", "HYWorldConfig", "Kandinsky5VideoConfig",
|
||||
"MagiHumanVideoConfig", "StableAudioConfig"
|
||||
"MagiHumanVideoConfig", "StableAudioConfig", "Flux2Config"
|
||||
]
|
||||
|
||||
@@ -14,12 +14,19 @@ class DiTArchConfig(ArchConfig):
|
||||
param_names_mapping: dict = field(default_factory=dict)
|
||||
reverse_param_names_mapping: dict = field(default_factory=dict)
|
||||
lora_param_names_mapping: dict = field(default_factory=dict)
|
||||
# When True, the denoising stage casts text/prompt embeddings to the DiT's
|
||||
# working dtype before the diffusion loop. Flux2 requires this (BFL casts ctx
|
||||
# to bf16 before denoising); models with fp32 text encoders (Wan, Hunyuan15,
|
||||
# SD3.5) leave it False to preserve full-precision embeddings.
|
||||
cast_prompt_embeds_to_dit_dtype: bool = False
|
||||
_supported_attention_backends: tuple[AttentionBackendEnum,
|
||||
...] = (AttentionBackendEnum.SAGE_ATTN, AttentionBackendEnum.FLASH_ATTN,
|
||||
AttentionBackendEnum.TORCH_SDPA,
|
||||
AttentionBackendEnum.VIDEO_SPARSE_ATTN,
|
||||
AttentionBackendEnum.VMOBA_ATTN, AttentionBackendEnum.SAGE_ATTN_THREE,
|
||||
AttentionBackendEnum.SLA_ATTN, AttentionBackendEnum.SAGE_SLA_ATTN)
|
||||
AttentionBackendEnum.ATTN_QAT_INFER,
|
||||
AttentionBackendEnum.ATTN_QAT_TRAIN, AttentionBackendEnum.SLA_ATTN,
|
||||
AttentionBackendEnum.SAGE_SLA_ATTN)
|
||||
|
||||
hidden_size: int = 0
|
||||
num_attention_heads: int = 0
|
||||
|
||||
@@ -0,0 +1,77 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
# Copied and adapted from: https://github.com/sglang-ai/sglang
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from fastvideo.configs.models.dits.base import DiTArchConfig, DiTConfig
|
||||
from fastvideo.logger import init_logger
|
||||
|
||||
logger = init_logger(__name__)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2ArchConfig(DiTArchConfig):
|
||||
"""Architecture configuration for Flux2 transformer model."""
|
||||
|
||||
cast_prompt_embeds_to_dit_dtype: bool = True
|
||||
|
||||
# Flux2-specific architecture parameters
|
||||
patch_size: int = 1
|
||||
in_channels: int = 64
|
||||
out_channels: int | None = None
|
||||
num_layers: int = 19 # Number of double-stream transformer blocks
|
||||
num_single_layers: int = 38 # Number of single-stream transformer blocks
|
||||
attention_head_dim: int = 128
|
||||
num_attention_heads: int = 24
|
||||
joint_attention_dim: int = 4096 # Dimension for text encoder output
|
||||
timestep_guidance_channels: int = 256 # Dimension for timestep embedding
|
||||
mlp_ratio: float = 3.0
|
||||
axes_dims_rope: tuple[int, ...] = (32, 32, 32, 32) # RoPE dimensions per axis (match diffusers Flux2)
|
||||
rope_theta: int = 2000 # Base frequency for RoPE (match diffusers Flux2)
|
||||
eps: float = 1e-6
|
||||
guidance_embeds: bool = True # Whether to use guidance embeddings
|
||||
# When True, compute SwiGLU in fp32 inside ``ff_context`` only (bf16 noise mitigation).
|
||||
ff_context_swiglu_fp32: bool = False
|
||||
|
||||
# Parameter name mapping for loading HuggingFace checkpoints
|
||||
param_names_mapping: dict = field(default_factory=lambda: {
|
||||
r"transformer\.(\w*)\.(.*)$": r"\1.\2",
|
||||
})
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
super().__post_init__()
|
||||
self.out_channels = self.out_channels or self.in_channels
|
||||
self.hidden_size = self.num_attention_heads * self.attention_head_dim
|
||||
self.num_channels_latents = self.out_channels
|
||||
|
||||
def update_from_weight_keys(self, all_keys: set[str]) -> None:
|
||||
"""Infer num_layers and num_single_layers from checkpoint weight keys so the model is built with the same number of blocks as the weights."""
|
||||
if not all_keys:
|
||||
return
|
||||
num_layers = 0
|
||||
num_single_layers = 0
|
||||
for k in all_keys:
|
||||
if "single_transformer_blocks." not in k and "transformer_blocks." in k:
|
||||
parts = k.split("transformer_blocks.")[-1].split(".")
|
||||
if parts[0].isdigit():
|
||||
num_layers = max(num_layers, int(parts[0]) + 1)
|
||||
if "single_transformer_blocks." in k:
|
||||
parts = k.split("single_transformer_blocks.")[-1].split(".")
|
||||
if parts[0].isdigit():
|
||||
num_single_layers = max(num_single_layers, int(parts[0]) + 1)
|
||||
if num_layers > 0:
|
||||
self.num_layers = num_layers
|
||||
logger.info("Inferred num_layers=%s from checkpoint keys", num_layers)
|
||||
if num_single_layers > 0:
|
||||
self.num_single_layers = num_single_layers
|
||||
logger.info("Inferred num_single_layers=%s from checkpoint keys", num_single_layers)
|
||||
if num_layers > 0 or num_single_layers > 0:
|
||||
self.__post_init__()
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2Config(DiTConfig):
|
||||
"""Configuration for Flux2 transformer model."""
|
||||
|
||||
arch_config: DiTArchConfig = field(default_factory=Flux2ArchConfig)
|
||||
|
||||
prefix: str = "Flux"
|
||||
@@ -7,6 +7,8 @@ from fastvideo.configs.models.encoders.qwen2_5 import Qwen2_5_VLConfig
|
||||
from fastvideo.configs.models.encoders.siglip import SiglipVisionConfig
|
||||
from fastvideo.configs.models.encoders.reason1 import Reason1ArchConfig, Reason1Config
|
||||
from fastvideo.configs.models.encoders.gemma import LTX2GemmaConfig
|
||||
from fastvideo.configs.models.encoders.mistral3 import Mistral3TextConfig
|
||||
from fastvideo.configs.models.encoders.qwen3 import Qwen3TextConfig
|
||||
from fastvideo.configs.models.encoders.stable_audio_conditioner import (StableAudioConditionerArchConfig,
|
||||
StableAudioConditionerConfig)
|
||||
from fastvideo.configs.models.encoders.t5gemma import T5GemmaEncoderConfig
|
||||
@@ -15,5 +17,5 @@ __all__ = [
|
||||
"EncoderConfig", "TextEncoderConfig", "ImageEncoderConfig", "BaseEncoderOutput", "CLIPTextConfig",
|
||||
"CLIPVisionConfig", "WAN2_1ControlCLIPVisionConfig", "LlamaConfig", "T5Config", "T5LargeConfig", "Qwen2_5_VLConfig",
|
||||
"Reason1ArchConfig", "Reason1Config", "LTX2GemmaConfig", "SiglipVisionConfig", "StableAudioConditionerArchConfig",
|
||||
"StableAudioConditionerConfig", "T5GemmaEncoderConfig"
|
||||
"StableAudioConditionerConfig", "T5GemmaEncoderConfig", "Qwen3TextConfig", "Mistral3TextConfig"
|
||||
]
|
||||
|
||||
@@ -36,6 +36,11 @@ class TextEncoderArchConfig(EncoderArchConfig):
|
||||
default_factory=list) # mapping from huggingface weight names to custom names
|
||||
tokenizer_kwargs: dict[str, Any] = field(default_factory=dict)
|
||||
_fsdp_shard_conditions: list = field(default_factory=lambda: [])
|
||||
# When True, the tokenizer loader prefers AutoProcessor over AutoTokenizer
|
||||
# for encoders whose tokenizer dir ships a processor_config.json (e.g. Flux2
|
||||
# full's Mistral3 multimodal processor). Default False keeps every existing
|
||||
# encoder on the historical AutoTokenizer path.
|
||||
require_processor: bool = False
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
self.tokenizer_kwargs = {
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""Mistral3 text encoder configuration for full Flux2."""
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from fastvideo.configs.models.encoders.base import (
|
||||
TextEncoderArchConfig,
|
||||
TextEncoderConfig,
|
||||
)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Mistral3TextArchConfig(TextEncoderArchConfig):
|
||||
"""Architecture config for the Mistral3 text encoder used by full Flux2."""
|
||||
|
||||
architectures: list[str] = field(default_factory=lambda: ["Mistral3ForConditionalGeneration"])
|
||||
hidden_size: int = 5120
|
||||
num_hidden_layers: int = 40
|
||||
text_len: int = 512
|
||||
output_hidden_states: bool = True
|
||||
# Mistral3 (full Flux2) ships a multimodal processor; load via AutoProcessor.
|
||||
require_processor: bool = True
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
self.tokenizer_kwargs = {
|
||||
"padding": "max_length",
|
||||
"truncation": True,
|
||||
"max_length": self.text_len,
|
||||
"return_tensors": "pt",
|
||||
}
|
||||
|
||||
|
||||
@dataclass
|
||||
class Mistral3TextConfig(TextEncoderConfig):
|
||||
"""Top-level config for the Mistral3 full Flux2 text encoder."""
|
||||
|
||||
arch_config: TextEncoderArchConfig = field(default_factory=Mistral3TextArchConfig)
|
||||
prefix: str = "mistral3"
|
||||
is_chat_model: bool = True
|
||||
@@ -0,0 +1,82 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
# Ported from SGLang: python/sglang/multimodal_gen/configs/models/encoders/qwen3.py
|
||||
"""Qwen3 text encoder configuration for FastVideo diffusion models (e.g. Flux2 Klein)."""
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.configs.models.encoders.base import (
|
||||
TextEncoderArchConfig,
|
||||
TextEncoderConfig,
|
||||
)
|
||||
|
||||
|
||||
def _is_transformer_layer(n: str, m: Any) -> bool:
|
||||
return "layers" in n and str.isdigit(n.split(".")[-1])
|
||||
|
||||
|
||||
def _is_embeddings(n: str, m: Any) -> bool:
|
||||
return n.endswith("embed_tokens")
|
||||
|
||||
|
||||
def _is_final_norm(n: str, m: Any) -> bool:
|
||||
return n.endswith("norm")
|
||||
|
||||
|
||||
@dataclass
|
||||
class Qwen3TextArchConfig(TextEncoderArchConfig):
|
||||
"""Architecture config for Qwen3 text encoder.
|
||||
|
||||
Qwen3 is similar to LLaMA but with QK-Norm (RMSNorm on Q and K before attention).
|
||||
Used by Flux2 Klein.
|
||||
"""
|
||||
|
||||
vocab_size: int = 151936
|
||||
hidden_size: int = 2560
|
||||
intermediate_size: int = 9728
|
||||
num_hidden_layers: int = 36
|
||||
num_attention_heads: int = 32
|
||||
num_key_value_heads: int = 8
|
||||
hidden_act: str = "silu"
|
||||
max_position_embeddings: int = 40960
|
||||
initializer_range: float = 0.02
|
||||
rms_norm_eps: float = 1e-6
|
||||
use_cache: bool = True
|
||||
pad_token_id: int = 151643
|
||||
bos_token_id: int = 151643
|
||||
eos_token_id: int = 151645
|
||||
tie_word_embeddings: bool = True
|
||||
rope_theta: float = 1000000.0
|
||||
rope_scaling: dict | None = None
|
||||
attention_bias: bool = False
|
||||
attention_dropout: float = 0.0
|
||||
mlp_bias: bool = False
|
||||
head_dim: int = 128
|
||||
text_len: int = 512
|
||||
output_hidden_states: bool = True # Klein needs hidden states from layers 9, 18, 27
|
||||
|
||||
stacked_params_mapping: list[tuple[str, str, str | int]] = field(default_factory=lambda: [
|
||||
(".qkv_proj", ".q_proj", "q"),
|
||||
(".qkv_proj", ".k_proj", "k"),
|
||||
(".qkv_proj", ".v_proj", "v"),
|
||||
(".gate_up_proj", ".gate_proj", 0),
|
||||
(".gate_up_proj", ".up_proj", 1),
|
||||
])
|
||||
_fsdp_shard_conditions: list = field(
|
||||
default_factory=lambda: [_is_transformer_layer, _is_embeddings, _is_final_norm])
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
self.tokenizer_kwargs = {
|
||||
"padding": "max_length",
|
||||
"truncation": True,
|
||||
"max_length": self.text_len,
|
||||
"return_tensors": "pt",
|
||||
}
|
||||
|
||||
|
||||
@dataclass
|
||||
class Qwen3TextConfig(TextEncoderConfig):
|
||||
"""Top-level config for Qwen3 text encoder."""
|
||||
|
||||
arch_config: TextEncoderArchConfig = field(default_factory=Qwen3TextArchConfig)
|
||||
prefix: str = "qwen3"
|
||||
is_chat_model: bool = True
|
||||
@@ -6,6 +6,7 @@ from fastvideo.configs.models.vaes.hunyuanvae import HunyuanVAEConfig
|
||||
from fastvideo.configs.models.vaes.hunyuan15vae import Hunyuan15VAEConfig
|
||||
from fastvideo.configs.models.vaes.ltx2vae import LTX2VAEConfig
|
||||
from fastvideo.configs.models.vaes.oobleck import OobleckVAEArchConfig, OobleckVAEConfig
|
||||
from fastvideo.configs.models.vaes.flux2vae import Flux2VAEConfig
|
||||
from fastvideo.configs.models.vaes.wanvae import WanVAEConfig
|
||||
|
||||
__all__ = [
|
||||
@@ -19,4 +20,5 @@ __all__ = [
|
||||
"LTX2VAEConfig",
|
||||
"OobleckVAEArchConfig",
|
||||
"OobleckVAEConfig",
|
||||
"Flux2VAEConfig",
|
||||
]
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
# Copied and adapted from: https://github.com/sglang-ai/sglang
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from fastvideo.configs.models.vaes.base import VAEArchConfig, VAEConfig
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2VAEArchConfig(VAEArchConfig):
|
||||
"""Architecture configuration for Flux2 VAE model."""
|
||||
|
||||
# Flux2 VAE-specific architecture parameters
|
||||
in_channels: int = 3
|
||||
out_channels: int = 3
|
||||
down_block_types: tuple[str, ...] = (
|
||||
"DownEncoderBlock2D",
|
||||
"DownEncoderBlock2D",
|
||||
"DownEncoderBlock2D",
|
||||
"AttnDownEncoderBlock2D",
|
||||
)
|
||||
up_block_types: tuple[str, ...] = (
|
||||
"AttnUpDecoderBlock2D",
|
||||
"UpDecoderBlock2D",
|
||||
"UpDecoderBlock2D",
|
||||
"UpDecoderBlock2D",
|
||||
)
|
||||
block_out_channels: tuple[int, ...] = (128, 256, 512, 512)
|
||||
layers_per_block: int = 2
|
||||
act_fn: str = "silu"
|
||||
latent_channels: int = 16
|
||||
norm_num_groups: int = 32
|
||||
sample_size: int = 512
|
||||
force_upcast: bool = False
|
||||
use_quant_conv: bool = True
|
||||
use_post_quant_conv: bool = True
|
||||
mid_block_add_attention: bool = True
|
||||
batch_norm_eps: float = 1e-5
|
||||
batch_norm_momentum: float = 0.1
|
||||
patch_size: tuple[int, int] = (1, 1)
|
||||
|
||||
# Latent scaling for decode: avoid division-by-zero; match Flux/Flux2 convention (e.g. 0.13025)
|
||||
scaling_factor: float = 0.13025
|
||||
|
||||
# Spatial compression (for images, this is typically 8)
|
||||
spatial_compression_ratio: int = 8
|
||||
temporal_compression_ratio: int = 1 # Images don't have temporal dimension
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2VAEConfig(VAEConfig):
|
||||
"""Configuration for Flux2 VAE model."""
|
||||
|
||||
arch_config: Flux2VAEArchConfig = field(default_factory=Flux2VAEArchConfig)
|
||||
|
||||
# Flux2 is an image model, so disable temporal tiling
|
||||
use_tiling: bool = False
|
||||
use_temporal_tiling: bool = False
|
||||
use_parallel_tiling: bool = False
|
||||
@@ -9,12 +9,12 @@ from fastvideo.configs.pipelines.matrixgame2 import MatrixGame2I2V480PConfig
|
||||
from fastvideo.configs.pipelines.matrixgame3 import MatrixGame3I2V720PConfig
|
||||
from fastvideo.pipelines.basic.ltx2.pipeline_configs import LTX2T2VConfig
|
||||
from fastvideo.registry import get_pipeline_config_cls_from_name
|
||||
from fastvideo.configs.pipelines.wan import (SelfForcingWanT2V480PConfig, WanI2V480PConfig, WanI2V720PConfig,
|
||||
WanT2V480PConfig, WanT2V720PConfig)
|
||||
from fastvideo.configs.pipelines.wan import (LucyEditDevConfig, SelfForcingWanT2V480PConfig, WanI2V480PConfig,
|
||||
WanI2V720PConfig, WanT2V480PConfig, WanT2V720PConfig)
|
||||
|
||||
__all__ = [
|
||||
"HunyuanConfig", "FastHunyuanConfig", "HunyuanGameCraftPipelineConfig", "PipelineConfig", "Hunyuan15T2V480PConfig",
|
||||
"Hunyuan15T2V720PConfig", "WanT2V480PConfig", "WanI2V480PConfig", "WanT2V720PConfig", "WanI2V720PConfig",
|
||||
"SelfForcingWanT2V480PConfig", "CosmosConfig", "Cosmos25Config", "LTX2T2VConfig", "HYWorldConfig",
|
||||
"MatrixGame2I2V480PConfig", "MatrixGame3I2V720PConfig", "get_pipeline_config_cls_from_name"
|
||||
"SelfForcingWanT2V480PConfig", "LucyEditDevConfig", "CosmosConfig", "Cosmos25Config", "LTX2T2VConfig",
|
||||
"HYWorldConfig", "MatrixGame2I2V480PConfig", "MatrixGame3I2V720PConfig", "get_pipeline_config_cls_from_name"
|
||||
]
|
||||
|
||||
@@ -35,6 +35,11 @@ class PipelineConfig:
|
||||
flow_shift: float | None = None
|
||||
flow_shift_sr: float | None = None
|
||||
disable_autocast: bool = False
|
||||
# When True, the scheduler's Euler update runs in fp32 outside the autocast
|
||||
# block (Diffusers-style; avoids BF16 drift over multiple steps). Flux2 sets
|
||||
# this True for reference parity; other models keep the legacy in-autocast
|
||||
# behavior to preserve existing SSIM references.
|
||||
scheduler_step_in_fp32: bool = False
|
||||
is_causal: bool = False
|
||||
|
||||
# Model configuration
|
||||
@@ -64,8 +69,9 @@ class PipelineConfig:
|
||||
# DMD parameters
|
||||
dmd_denoising_steps: list[int] | None = field(default=None)
|
||||
|
||||
# Wan2.2 TI2V parameters
|
||||
# Wan2.2 task modifiers
|
||||
ti2v_task: bool = False
|
||||
lucy_edit_task: bool = False
|
||||
boundary_ratio: float | None = None
|
||||
|
||||
# Compilation
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
# Copied and adapted from: https://github.com/sglang-ai/sglang
|
||||
from collections.abc import Callable
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
import torch
|
||||
|
||||
from fastvideo.configs.models import DiTConfig, EncoderConfig, VAEConfig
|
||||
from fastvideo.configs.models.dits.flux_2 import Flux2Config
|
||||
from fastvideo.configs.models.encoders import BaseEncoderOutput
|
||||
from fastvideo.configs.models.encoders.base import EncoderArchConfig
|
||||
from fastvideo.configs.models.encoders.mistral3 import Mistral3TextConfig
|
||||
from fastvideo.configs.models.encoders.qwen3 import Qwen3TextConfig
|
||||
from fastvideo.configs.models.vaes.flux2vae import Flux2VAEConfig
|
||||
from fastvideo.configs.pipelines.base import PipelineConfig, preprocess_text
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2PipelineConfig(PipelineConfig):
|
||||
"""Configuration for Flux2 image generation pipeline."""
|
||||
|
||||
# Flux2-specific parameters
|
||||
embedded_cfg_scale: float | None = 4.0
|
||||
scheduler_step_in_fp32: bool = True
|
||||
flux2_text_encoder_type: str = "mistral3"
|
||||
text_encoder_out_layers: tuple[int, ...] = (10, 20, 30)
|
||||
|
||||
# DiT configuration
|
||||
dit_config: DiTConfig = field(default_factory=Flux2Config)
|
||||
dit_precision: str = "bf16"
|
||||
|
||||
# VAE configuration
|
||||
vae_config: VAEConfig = field(default_factory=Flux2VAEConfig)
|
||||
vae_precision: str = "fp32"
|
||||
vae_tiling: bool = False # Flux2 is image model, disable tiling by default
|
||||
vae_sp: bool = False
|
||||
|
||||
# Text encoder configuration (full Flux2 uses Mistral3)
|
||||
text_encoder_configs: tuple[EncoderConfig, ...] = field(default_factory=lambda: (Mistral3TextConfig(), ))
|
||||
text_encoder_precisions: tuple[str, ...] = field(default_factory=lambda: ("bf16", ))
|
||||
|
||||
# Default postprocess function (can be overridden)
|
||||
@staticmethod
|
||||
def default_postprocess_text(outputs: BaseEncoderOutput) -> torch.Tensor:
|
||||
"""Default text postprocessing for Flux2."""
|
||||
return outputs.last_hidden_state
|
||||
|
||||
postprocess_text_funcs: tuple[Callable[[BaseEncoderOutput], torch.Tensor],
|
||||
...] = field(default_factory=lambda: (Flux2PipelineConfig.default_postprocess_text, ))
|
||||
|
||||
|
||||
def flux2_klein_postprocess_text(outputs: BaseEncoderOutput) -> torch.Tensor:
|
||||
"""Klein postprocess: hidden states from layers 9, 18, 27 (Qwen3)."""
|
||||
hidden_states_layers: list[int] = [9, 18, 27]
|
||||
if outputs.hidden_states is None:
|
||||
raise ValueError("Flux2 Klein requires output_hidden_states=True from text encoder")
|
||||
out = torch.stack([outputs.hidden_states[k] for k in hidden_states_layers], dim=1)
|
||||
batch_size, num_channels, seq_len, hidden_dim = out.shape
|
||||
prompt_embeds = out.permute(0, 2, 1, 3).reshape(batch_size, seq_len, num_channels * hidden_dim)
|
||||
return prompt_embeds
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2KleinEncoderArchConfig(EncoderArchConfig):
|
||||
"""Encoder arch config for Flux2 Klein (Qwen3); needs hidden states for layers 9, 18, 27."""
|
||||
output_hidden_states: bool = True
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2KleinTextEncoderConfig(EncoderConfig):
|
||||
"""Text encoder config for Flux2 Klein (Qwen3)."""
|
||||
arch_config: EncoderArchConfig = field(default_factory=Flux2KleinEncoderArchConfig)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Flux2KleinPipelineConfig(Flux2PipelineConfig):
|
||||
"""Configuration for Flux2 Klein (distilled, 4-step, no guidance)."""
|
||||
embedded_cfg_scale: float | None = None # Klein distilled: no guidance embedding (matches Diffusers)
|
||||
scheduler_step_in_fp32: bool = True
|
||||
flux2_text_encoder_type: str = "qwen3"
|
||||
text_encoder_out_layers: tuple[int, ...] = (9, 18, 27)
|
||||
text_encoder_configs: tuple[EncoderConfig, ...] = field(default_factory=lambda: (Qwen3TextConfig(), ))
|
||||
text_encoder_precisions: tuple[str, ...] = field(default_factory=lambda: ("bf16", ))
|
||||
preprocess_text_funcs: tuple[Callable[[str], str], ...] = field(default_factory=lambda: (preprocess_text, ))
|
||||
postprocess_text_funcs: tuple[Callable[[BaseEncoderOutput], torch.Tensor],
|
||||
...] = field(default_factory=lambda: (flux2_klein_postprocess_text, ))
|
||||
@@ -6,9 +6,11 @@ import torch
|
||||
|
||||
from fastvideo.configs.models import DiTConfig, EncoderConfig, VAEConfig
|
||||
from fastvideo.configs.models.dits import WanVideoConfig
|
||||
from fastvideo.configs.models.dits.wanvideo import WanVideoArchConfig
|
||||
from fastvideo.configs.models.encoders import (BaseEncoderOutput, CLIPVisionConfig, T5Config,
|
||||
WAN2_1ControlCLIPVisionConfig)
|
||||
from fastvideo.configs.models.vaes import WanVAEConfig
|
||||
from fastvideo.configs.models.vaes.wanvae import WanVAEArchConfig
|
||||
from fastvideo.configs.pipelines.base import PipelineConfig
|
||||
|
||||
|
||||
@@ -120,6 +122,142 @@ class Wan2_2_TI2V_5B_Config(WanT2V480PConfig):
|
||||
expand_timesteps: bool = True
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
assert not (self.ti2v_task and self.lucy_edit_task)
|
||||
self.vae_config.load_encoder = True
|
||||
self.vae_config.load_decoder = True
|
||||
self.dit_config.expand_timesteps = self.expand_timesteps
|
||||
|
||||
|
||||
@dataclass
|
||||
class LucyEditDevConfig(Wan2_2_TI2V_5B_Config):
|
||||
"""Configuration for Decart Lucy Edit Dev video editing."""
|
||||
|
||||
dit_config: DiTConfig = field(default_factory=lambda: WanVideoConfig(arch_config=WanVideoArchConfig(
|
||||
num_attention_heads=24,
|
||||
in_channels=96,
|
||||
out_channels=48,
|
||||
ffn_dim=14336,
|
||||
num_layers=30,
|
||||
)))
|
||||
vae_config: VAEConfig = field(default_factory=lambda: WanVAEConfig(arch_config=WanVAEArchConfig(
|
||||
base_dim=160,
|
||||
decoder_base_dim=256,
|
||||
z_dim=48,
|
||||
in_channels=12,
|
||||
out_channels=12,
|
||||
scale_factor_spatial=16,
|
||||
patch_size=2,
|
||||
is_residual=True,
|
||||
clip_output=False,
|
||||
latents_mean=(
|
||||
-0.2289,
|
||||
-0.0052,
|
||||
-0.1323,
|
||||
-0.2339,
|
||||
-0.2799,
|
||||
0.0174,
|
||||
0.1838,
|
||||
0.1557,
|
||||
-0.1382,
|
||||
0.0542,
|
||||
0.2813,
|
||||
0.0891,
|
||||
0.1570,
|
||||
-0.0098,
|
||||
0.0375,
|
||||
-0.1825,
|
||||
-0.2246,
|
||||
-0.1207,
|
||||
-0.0698,
|
||||
0.5109,
|
||||
0.2665,
|
||||
-0.2108,
|
||||
-0.2158,
|
||||
0.2502,
|
||||
-0.2055,
|
||||
-0.0322,
|
||||
0.1109,
|
||||
0.1567,
|
||||
-0.0729,
|
||||
0.0899,
|
||||
-0.2799,
|
||||
-0.1230,
|
||||
-0.0313,
|
||||
-0.1649,
|
||||
0.0117,
|
||||
0.0723,
|
||||
-0.2839,
|
||||
-0.2083,
|
||||
-0.0520,
|
||||
0.3748,
|
||||
0.0152,
|
||||
0.1957,
|
||||
0.1433,
|
||||
-0.2944,
|
||||
0.3573,
|
||||
-0.0548,
|
||||
-0.1681,
|
||||
-0.0667,
|
||||
),
|
||||
latents_std=(
|
||||
0.4765,
|
||||
1.0364,
|
||||
0.4514,
|
||||
1.1677,
|
||||
0.5313,
|
||||
0.4990,
|
||||
0.4818,
|
||||
0.5013,
|
||||
0.8158,
|
||||
1.0344,
|
||||
0.5894,
|
||||
1.0901,
|
||||
0.6885,
|
||||
0.6165,
|
||||
0.8454,
|
||||
0.4978,
|
||||
0.5759,
|
||||
0.3523,
|
||||
0.7135,
|
||||
0.6804,
|
||||
0.5833,
|
||||
1.4146,
|
||||
0.8986,
|
||||
0.5659,
|
||||
0.7069,
|
||||
0.5338,
|
||||
0.4889,
|
||||
0.4917,
|
||||
0.4069,
|
||||
0.4999,
|
||||
0.6866,
|
||||
0.4093,
|
||||
0.5709,
|
||||
0.6065,
|
||||
0.6415,
|
||||
0.4944,
|
||||
0.5726,
|
||||
1.2042,
|
||||
0.5458,
|
||||
1.6887,
|
||||
0.3971,
|
||||
1.0600,
|
||||
0.3943,
|
||||
0.5537,
|
||||
0.5444,
|
||||
0.4089,
|
||||
0.7468,
|
||||
0.7744,
|
||||
),
|
||||
)))
|
||||
ti2v_task: bool = False
|
||||
lucy_edit_task: bool = True
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
assert not (self.ti2v_task and self.lucy_edit_task)
|
||||
# Lucy uses Wan2.2's enhanced 48-channel VAE latents. Denoising
|
||||
# concatenates noise + video latents, matching the 96-channel
|
||||
# transformer input declared above.
|
||||
self.vae_config.load_encoder = True
|
||||
self.vae_config.load_decoder = True
|
||||
self.dit_config.expand_timesteps = self.expand_timesteps
|
||||
|
||||
@@ -517,6 +517,7 @@ class VideoGenerator:
|
||||
if _ek in kwargs:
|
||||
extra_overrides[_ek] = kwargs.pop(_ek)
|
||||
|
||||
prompt_embeds = kwargs.pop("prompt_embeds", None)
|
||||
sampling_param.update(kwargs)
|
||||
kwargs["_extra_overrides"] = extra_overrides
|
||||
|
||||
@@ -567,6 +568,8 @@ class VideoGenerator:
|
||||
raise ValueError("Either prompt or prompt_txt must be provided")
|
||||
output_path = self._prepare_output_path(sampling_param.output_path, prompt)
|
||||
kwargs["output_path"] = output_path
|
||||
if prompt_embeds is not None:
|
||||
kwargs["prompt_embeds"] = prompt_embeds
|
||||
return self._generate_single_video(
|
||||
prompt=prompt,
|
||||
sampling_param=sampling_param,
|
||||
@@ -669,6 +672,7 @@ class VideoGenerator:
|
||||
prompt = prompt.strip()
|
||||
sampling_param = deepcopy(sampling_param)
|
||||
output_path = kwargs["output_path"]
|
||||
prompt_embeds = kwargs.get("prompt_embeds")
|
||||
sampling_param.prompt = prompt
|
||||
# Process negative prompt
|
||||
if sampling_param.negative_prompt is not None:
|
||||
@@ -715,6 +719,9 @@ class VideoGenerator:
|
||||
n_tokens=n_tokens,
|
||||
VSA_sparsity=fastvideo_args.VSA_sparsity,
|
||||
)
|
||||
# Allow precomputed prompt_embeds (e.g. from diffusers) to skip text encoding
|
||||
if prompt_embeds is not None:
|
||||
batch.prompt_embeds = (list(prompt_embeds) if isinstance(prompt_embeds, list | tuple) else [prompt_embeds])
|
||||
|
||||
extra_overrides = kwargs.pop("_extra_overrides", {})
|
||||
for _ek, _ev in extra_overrides.items():
|
||||
|
||||
@@ -2,8 +2,9 @@
|
||||
|
||||
In-process evaluation suite for video generations. Includes pixel
|
||||
metrics (SSIM, PSNR, LPIPS), Fréchet Video Distance (FVD), optical-flow
|
||||
comparisons, the full VBench suite, Physics-IQ, audio metrics, and a
|
||||
VLM scorer behind a single registry-driven API.
|
||||
comparisons, the full VBench suite, Physics-IQ, audio metrics, an
|
||||
absolute VLM scorer (`videoscore2`), and a pairwise VLM judge
|
||||
(`judge.third_person_separation`) — all behind a single registry-driven API.
|
||||
|
||||
## Install
|
||||
|
||||
@@ -159,6 +160,7 @@ fastvideo/
|
||||
│ ├── audio/ # clap_score, audiobox_aesthetics, kl_divergence,
|
||||
│ │ # frechet_distance, wer, desync, imagebind_score
|
||||
│ ├── videoscore2/ # VideoScore-2 (Qwen2.5-VL)
|
||||
│ ├── judge/ # pairwise VLM judges (third_person_separation)
|
||||
│ ├── physics_iq/ # PhysicsIQ + sub-metrics
|
||||
│ └── vbench/ # adapter: sys.path bootstrap + shims
|
||||
│ ├── __init__.py
|
||||
@@ -313,6 +315,40 @@ to control read/write behavior. The example script
|
||||
`examples/inference/eval/eval_fvd.py` demonstrates the full
|
||||
two-directory workflow.
|
||||
|
||||
## `judge.third_person_separation` — pairwise VLM judge
|
||||
|
||||
A **preference** metric (a judge, not an absolute score), and the suite's first
|
||||
remote-API one. For each pair the judge (Gemini) sees the shared first frame and
|
||||
two rollouts — a candidate and a reference model under the same control signal —
|
||||
and picks the one that better separates the third-person CHARACTER (foreground)
|
||||
from the BACKGROUND. The corpus score is the candidate's win-rate, excluding
|
||||
ties. Set-vs-set, motion-first; it reads native mp4s, so samples carry path
|
||||
strings, not decoded tensors.
|
||||
|
||||
```bash
|
||||
uv pip install -e .[eval-judge] # opt-in: needs network + an API key
|
||||
export GEMINI_API_KEY=... # or GOOGLE_API_KEY, or ~/.gemini_token
|
||||
```
|
||||
|
||||
```python
|
||||
from fastvideo.eval import create_evaluator
|
||||
|
||||
ev = create_evaluator(metrics=["judge.third_person_separation"], device="cpu")
|
||||
result = ev.evaluate(samples=[
|
||||
{"video_path": "cand/000.mp4", "reference_path": "base/000.mp4",
|
||||
"image_path": "frames/000.png", "text_prompt": "W: moves forward", "action": "W"},
|
||||
# ... more pairs ...
|
||||
]).corpus["judge.third_person_separation"]
|
||||
result.score # candidate win-rate excl. ties; result.details has the breakdown
|
||||
```
|
||||
|
||||
Only `video_path`/`reference_path` are required; `image_path`/`text_prompt`/
|
||||
`action` are optional. Verdicts are cached under `${FASTVIDEO_EVAL_CACHE}/eval/judge/`.
|
||||
The judge separates best when the control yields genuine parallax (e.g.
|
||||
translation); rigid whole-frame motion (e.g. pure camera rotation) is harder. To
|
||||
sweep several baselines into a table, see
|
||||
`examples/inference/eval/eval_third_person_separation.py`.
|
||||
|
||||
## Out of scope (follow-up PRs)
|
||||
|
||||
- **MIND** metrics. Depend on a separate `vipe` upstream submodule.
|
||||
|
||||
@@ -0,0 +1,306 @@
|
||||
"""Pairwise VLM judge (Gemini): of two rollouts under the same control, which
|
||||
better separates the third-person character (foreground) from the background;
|
||||
the score is the candidate's win-rate over a reference.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from fastvideo.eval.metrics.base import BaseMetric
|
||||
from fastvideo.eval.models import get_cache_dir
|
||||
from fastvideo.eval.registry import register
|
||||
from fastvideo.eval.types import MetricResult, Video
|
||||
|
||||
# Part of the on-disk cache key; bump to invalidate cached verdicts.
|
||||
RUBRIC_ID = "v1"
|
||||
DEFAULT_MODEL = "gemini-2.5-pro"
|
||||
DEFAULT_K = 3
|
||||
|
||||
SYSTEM_PROMPT = ("You are a strict comparative evaluator of third-person video-game "
|
||||
"rollouts. Two videos were generated by two different models from the "
|
||||
"SAME first frame and the SAME control signal. You judge which video "
|
||||
"better demonstrates that the model SEPARATES the third-person CHARACTER "
|
||||
"(foreground) from the BACKGROUND SCENE — i.e. the control animates the "
|
||||
"character as an INDEPENDENT AGENT with its own trajectory while the "
|
||||
"background moves with the camera. You do NOT reward whichever video "
|
||||
"merely looks cleaner, sharper, or higher-res.")
|
||||
|
||||
RUBRIC = """\
|
||||
You are watching TWO generated video rollouts (Video 1, Video 2) played in full,
|
||||
from the same first frame under the same control signal. Pick the better
|
||||
third-person world-model rollout. Judge THREE things together, in this order:
|
||||
|
||||
(C) MOTION / ACTION EXECUTION FIRST. The rollout must actually CARRY OUT the
|
||||
control signal with substantial motion (the scene/character clearly moves as
|
||||
commanded). A clip that is near-static, barely drifts, or only twitches has
|
||||
NOT demonstrated controllable separation — it FAILS, no matter how clean it
|
||||
looks. CRITICAL: do NOT reward a clip for looking "smoother" or "more
|
||||
stable" when that smoothness is really just the ABSENCE OF MOTION. Less
|
||||
motion is NOT better. If one clip executes the action with clear motion and
|
||||
the other is comparatively static, the MOVING one wins (unless it fails B).
|
||||
|
||||
(A) TEMPORAL COHERENCE — among clips that actually move, penalize GENUINE
|
||||
corruption: flicker/strobing, texture boiling/crawling, geometry swimming,
|
||||
the character or scene morphing/warping into mush, colors pulsing, or
|
||||
progressive degradation into noise. Do NOT confuse LEGITIMATE large motion
|
||||
(camera sweeping, character running, scene flowing past) with instability —
|
||||
fast correct motion is GOOD, not a defect. Only true frame-to-frame
|
||||
INCOHERENCE counts against a clip.
|
||||
|
||||
(B) FOREGROUND/BACKGROUND SEPARATION — the character stays a distinct, coherent
|
||||
entity with its own trajectory while the background responds to the control;
|
||||
it does not dissolve/smear into the bg, and the whole frame does not slide
|
||||
as one rigid sheet.
|
||||
|
||||
Decision: among clips that genuinely execute the motion (C), pick the one that
|
||||
is both temporally coherent (A) and shows cleaner separation (B). A static or
|
||||
barely-moving clip loses to a moving one. Real motion is not instability. Do not
|
||||
reward resolution or placidity. "tie" only if truly equivalent on all three."""
|
||||
|
||||
USER_TASK = ("Output a JSON object with four fields:\n"
|
||||
" - video_1_analysis: FIRST, how much does Video 1 actually move — does it "
|
||||
"execute the control with clear motion, or is it near-static / barely "
|
||||
"drifting? THEN: among its motion, is there GENUINE corruption (flicker, "
|
||||
"boiling, morphing into mush) as opposed to legitimate fast motion? THEN: "
|
||||
"is the character a distinct coherent entity vs rigid-slide / dissolve?\n"
|
||||
" - video_2_analysis: the same three checks for Video 2.\n"
|
||||
" - comparison: apply (C) motion-first, then (A) genuine-coherence, then "
|
||||
"(B) separation. A near-static clip loses to a moving one; legitimate large "
|
||||
"motion is NOT a defect.\n"
|
||||
" - winner: \"video_1\", \"video_2\", or \"tie\".")
|
||||
|
||||
RESPONSE_SCHEMA = {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"video_1_analysis": {
|
||||
"type": "string"
|
||||
},
|
||||
"video_2_analysis": {
|
||||
"type": "string"
|
||||
},
|
||||
"comparison": {
|
||||
"type": "string"
|
||||
},
|
||||
"winner": {
|
||||
"type": "string",
|
||||
"enum": ["video_1", "video_2", "tie"]
|
||||
},
|
||||
},
|
||||
"required": ["video_1_analysis", "video_2_analysis", "comparison", "winner"],
|
||||
"propertyOrdering": ["video_1_analysis", "video_2_analysis", "comparison", "winner"],
|
||||
}
|
||||
|
||||
|
||||
def _path_of(sample: dict, key: str) -> str | None:
|
||||
"""Resolve a native-file path from a string key or a Video wrapper."""
|
||||
p = sample.get(f"{key}_path")
|
||||
if isinstance(p, str):
|
||||
return p
|
||||
v = sample.get(key)
|
||||
if isinstance(v, Video) and isinstance(v.source, str):
|
||||
return v.source
|
||||
if isinstance(v, str):
|
||||
return v
|
||||
return None
|
||||
|
||||
|
||||
def _resolve_api_key() -> str:
|
||||
for env in ("GEMINI_API_KEY", "GOOGLE_API_KEY"):
|
||||
key = os.environ.get(env)
|
||||
if key:
|
||||
return key.strip()
|
||||
token = Path("~/.gemini_token").expanduser()
|
||||
if token.is_file():
|
||||
return token.read_text().strip()
|
||||
raise ValueError("judge.third_person_separation needs a Gemini API key. Set "
|
||||
"GEMINI_API_KEY (or GOOGLE_API_KEY), or write it to ~/.gemini_token.")
|
||||
|
||||
|
||||
@register("judge.third_person_separation")
|
||||
class ThirdPersonSeparationMetric(BaseMetric):
|
||||
"""Pairwise VLM judge of third-person fg/bg separation; corpus win-rate."""
|
||||
|
||||
name = "judge.third_person_separation"
|
||||
requires_reference = True
|
||||
higher_is_better = True
|
||||
needs_gpu = False
|
||||
is_set_metric = True
|
||||
dependencies = ["google.genai"]
|
||||
|
||||
def __init__(self, model: str = DEFAULT_MODEL, k: int = DEFAULT_K) -> None:
|
||||
super().__init__()
|
||||
self.model = model
|
||||
self.k = k
|
||||
self._client: Any = None
|
||||
self._files: dict[str, Any] = {} # path -> uploaded Gemini file handle
|
||||
self._records: list[dict] = [] # one per accumulated pair
|
||||
|
||||
# --- model / client -----------------------------------------------------
|
||||
def setup(self) -> None:
|
||||
if self._client is not None:
|
||||
return
|
||||
from google import genai
|
||||
self._client = genai.Client(api_key=_resolve_api_key())
|
||||
|
||||
# --- set-vs-set protocol ------------------------------------------------
|
||||
def reset(self) -> None:
|
||||
self._records = []
|
||||
self._files = {}
|
||||
|
||||
def accumulate(self, sample: dict) -> None:
|
||||
cand = _path_of(sample, "video")
|
||||
base = _path_of(sample, "reference")
|
||||
if cand is None or base is None:
|
||||
return # nothing to compare
|
||||
|
||||
image = sample.get("image_path")
|
||||
action_text = sample.get("text_prompt") or ""
|
||||
action = sample.get("action")
|
||||
|
||||
rec = self._cached(cand, base, action_text)
|
||||
if rec is None:
|
||||
if self._client is None:
|
||||
self.setup()
|
||||
rec = self._judge_pair(cand, base, image, action_text)
|
||||
self._write_cache(cand, base, action_text, rec)
|
||||
rec = {**rec, "action": action}
|
||||
self._records.append(rec)
|
||||
|
||||
def finalize(self) -> MetricResult:
|
||||
recs = [r for r in self._records if r.get("verdict")]
|
||||
if not recs:
|
||||
return MetricResult(name=self.name, score=None, details={"skipped": "no pairs judged"})
|
||||
wins = sum(r["verdict"] == "candidate" for r in recs)
|
||||
losses = sum(r["verdict"] == "baseline" for r in recs)
|
||||
ties = sum(r["verdict"] == "tie" for r in recs)
|
||||
decided = wins + losses
|
||||
score = wins / decided if decided else None
|
||||
|
||||
# Per-action win-rate, grouped by the raw label (no assumed control scheme).
|
||||
per_action: dict[str, dict] = {}
|
||||
labels: set[str] = {str(r["action"]) for r in recs if r.get("action")}
|
||||
for action in sorted(labels):
|
||||
gr = [r for r in recs if r.get("action") == action]
|
||||
gw = sum(r["verdict"] == "candidate" for r in gr)
|
||||
gl = sum(r["verdict"] == "baseline" for r in gr)
|
||||
per_action[action] = {
|
||||
"n": len(gr),
|
||||
"wins": gw,
|
||||
"losses": gl,
|
||||
"ties": len(gr) - gw - gl,
|
||||
"win_rate": (gw / (gw + gl)) if (gw + gl) else None,
|
||||
}
|
||||
return MetricResult(name=self.name,
|
||||
score=score,
|
||||
details={
|
||||
"wins": wins,
|
||||
"losses": losses,
|
||||
"ties": ties,
|
||||
"n": len(recs),
|
||||
"win_rate_excl_ties": score,
|
||||
"per_action": per_action,
|
||||
})
|
||||
|
||||
def merge_from(self, other: BaseMetric) -> None:
|
||||
assert isinstance(other, ThirdPersonSeparationMetric)
|
||||
self._records.extend(other._records)
|
||||
|
||||
# --- judging ------------------------------------------------------------
|
||||
def _judge_pair(self, cand: str, base: str, image: str | None, action_text: str) -> dict:
|
||||
"""k counterbalanced comparisons → aggregated per-pair verdict."""
|
||||
# Seed the A/B alternation from the pair itself so it is reproducible and
|
||||
# independent of evaluation order or which subset is being run.
|
||||
seed = int(hashlib.sha1(f"{cand}|{base}".encode()).hexdigest(), 16)
|
||||
mapped: list[str] = []
|
||||
for i in range(self.k):
|
||||
cand_first = (seed + i) % 2 == 0
|
||||
v1, v2 = (cand, base) if cand_first else (base, cand)
|
||||
winner = self._one_call(image, v1, v2, action_text)
|
||||
if winner == "tie":
|
||||
mapped.append("tie")
|
||||
elif (winner == "video_1") == cand_first:
|
||||
mapped.append("candidate")
|
||||
else:
|
||||
mapped.append("baseline")
|
||||
cand_w = mapped.count("candidate")
|
||||
base_w = mapped.count("baseline")
|
||||
verdict = ("candidate" if cand_w > base_w else "baseline" if base_w > cand_w else "tie")
|
||||
return {
|
||||
"verdict": verdict,
|
||||
"candidate_wins": cand_w,
|
||||
"baseline_wins": base_w,
|
||||
"ties": mapped.count("tie"),
|
||||
"k": self.k,
|
||||
"rubric_id": RUBRIC_ID
|
||||
}
|
||||
|
||||
def _one_call(self, image: str | None, vid1: str, vid2: str, action_text: str) -> str:
|
||||
from google.genai import types
|
||||
contents: list[Any] = [action_text or "Compare these two rollouts."]
|
||||
if image is not None:
|
||||
contents += ["\nFirst frame (input condition, shared by BOTH "
|
||||
"videos):", self._upload(image)]
|
||||
contents += [
|
||||
"\nVideo 1 (model A's full rollout — watch it in motion):",
|
||||
self._upload(vid1),
|
||||
"\nVideo 2 (model B's full rollout — watch it in motion):",
|
||||
self._upload(vid2),
|
||||
"\n" + RUBRIC + "\n\n" + USER_TASK,
|
||||
]
|
||||
for attempt in range(6):
|
||||
try:
|
||||
resp = self._client.models.generate_content(model=self.model,
|
||||
contents=contents,
|
||||
config=types.GenerateContentConfig(
|
||||
system_instruction=SYSTEM_PROMPT,
|
||||
response_mime_type="application/json",
|
||||
response_schema=RESPONSE_SCHEMA,
|
||||
temperature=0.4))
|
||||
return json.loads(resp.text).get("winner", "tie")
|
||||
except Exception as exc: # noqa: BLE001 - transient API errors
|
||||
if attempt == 5:
|
||||
print(f"[judge] giving up after 6 attempts ({exc}); scoring this call a tie")
|
||||
break
|
||||
is_429 = "429" in str(exc) or "RESOURCE_EXHAUSTED" in str(exc)
|
||||
time.sleep(40 if is_429 else 2**attempt)
|
||||
return "tie"
|
||||
|
||||
def _upload(self, path: str) -> Any:
|
||||
f = self._files.get(path)
|
||||
if f is not None:
|
||||
return f
|
||||
f = self._client.files.upload(file=path)
|
||||
while f.state.name != "ACTIVE":
|
||||
time.sleep(1)
|
||||
f = self._client.files.get(name=f.name)
|
||||
if f.state.name == "FAILED":
|
||||
raise RuntimeError(f"Gemini upload failed for {path}")
|
||||
self._files[path] = f
|
||||
return f
|
||||
|
||||
# --- per-pair cache -----------------------------------------------------
|
||||
def _cache_path(self, cand: str, base: str, action_text: str) -> Path:
|
||||
# k is intentionally NOT in the key so a larger-k run can reuse an
|
||||
# existing verdict with enough samples (see ``_cached``).
|
||||
h = hashlib.sha1(f"{RUBRIC_ID}|{self.model}|{cand}|{base}|{action_text}".encode()).hexdigest()[:16]
|
||||
return get_cache_dir() / "judge" / "third_person_separation" / f"{h}.json"
|
||||
|
||||
def _cached(self, cand: str, base: str, action_text: str) -> dict | None:
|
||||
cp = self._cache_path(cand, base, action_text)
|
||||
if not cp.is_file():
|
||||
return None
|
||||
try:
|
||||
rec = json.loads(cp.read_text())
|
||||
except Exception:
|
||||
return None
|
||||
return rec if rec.get("k", 0) >= self.k and "verdict" in rec else None
|
||||
|
||||
def _write_cache(self, cand: str, base: str, action_text: str, rec: dict) -> None:
|
||||
cp = self._cache_path(cand, base, action_text)
|
||||
cp.parent.mkdir(parents=True, exist_ok=True)
|
||||
cp.write_text(json.dumps(rec, indent=2))
|
||||
@@ -85,6 +85,8 @@ def _extra_for(metric_name: str) -> str:
|
||||
return "eval-physics-iq"
|
||||
if metric_name.startswith("audio."):
|
||||
return "eval-audio"
|
||||
if metric_name.startswith("judge."):
|
||||
return "eval-judge"
|
||||
return "eval"
|
||||
|
||||
|
||||
|
||||
@@ -219,6 +219,15 @@ class ReplicatedLinear(LinearBase):
|
||||
(e.g. model.layers.0.qkv_proj)
|
||||
"""
|
||||
|
||||
# Opt-in instrumentation: when ``enable_shape_tracking`` is set to True,
|
||||
# ``forward`` records every unique ``(input_shape, output_shape)`` pair
|
||||
# observed across all ``ReplicatedLinear`` instances, along with the
|
||||
# subclass name that produced it. Used by upcoming QAT-aware backends
|
||||
# to discover which GEMM shapes need quantized kernels. Defaults to
|
||||
# False; default forward path is bit-identical to pre-slice behavior.
|
||||
enable_shape_tracking = False
|
||||
_shape_to_layer_types: dict[tuple[torch.Size, torch.Size], set[str]] = {}
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
input_size: int,
|
||||
@@ -285,6 +294,8 @@ class ReplicatedLinear(LinearBase):
|
||||
bias = self.bias if not self.skip_bias_add else None
|
||||
assert self.quant_method is not None
|
||||
output = self.quant_method.apply(self, x, bias)
|
||||
if self.enable_shape_tracking:
|
||||
self._track_shape(x.shape, output.shape)
|
||||
output_bias = self.bias if self.skip_bias_add else None
|
||||
return output, output_bias
|
||||
|
||||
@@ -294,6 +305,41 @@ class ReplicatedLinear(LinearBase):
|
||||
s += f", bias={self.bias is not None}"
|
||||
return s
|
||||
|
||||
@classmethod
|
||||
def get_shape_mapping(cls) -> dict:
|
||||
"""Get the mapping from (input_shape, output_shape) to layer types."""
|
||||
return cls._shape_to_layer_types.copy()
|
||||
|
||||
@classmethod
|
||||
def reset_shape_tracking(cls) -> None:
|
||||
"""Clear tracked shapes and layer type mappings."""
|
||||
cls._shape_to_layer_types.clear()
|
||||
|
||||
def _track_shape(self, input_shape: torch.Size, output_shape: torch.Size) -> None:
|
||||
shape_key = (input_shape, output_shape)
|
||||
if shape_key not in self._shape_to_layer_types:
|
||||
self._shape_to_layer_types[shape_key] = set()
|
||||
logger.debug("Layer: %s | input shape: %s --> output shape: %s, Quant Method: %s", self.prefix, input_shape,
|
||||
output_shape, self.quant_method.__class__.__name__)
|
||||
self._shape_to_layer_types[shape_key].add(self.__class__.__name__)
|
||||
|
||||
@classmethod
|
||||
def print_shape_summary(cls) -> None:
|
||||
"""Log a summary of all unique shapes and their layer types."""
|
||||
if not cls._shape_to_layer_types:
|
||||
logger.info("No shapes have been processed yet.")
|
||||
return
|
||||
|
||||
lines = [
|
||||
"=== Matrix Multiplication Shape Summary ===",
|
||||
f"Total unique shapes: {len(cls._shape_to_layer_types)}",
|
||||
]
|
||||
for i, (shape_key, layer_types) in enumerate(cls._shape_to_layer_types.items(), 1):
|
||||
input_shape, output_shape = shape_key
|
||||
lines.append(f"{i}. Input: {input_shape} → Output: {output_shape}")
|
||||
lines.append(f" Layer types: {', '.join(sorted(layer_types))}")
|
||||
logger.info("\n".join(lines))
|
||||
|
||||
|
||||
class ColumnParallelLinear(LinearBase):
|
||||
"""Linear layer with column parallelism.
|
||||
|
||||
+12
-2
@@ -5,6 +5,7 @@ import torch.nn as nn
|
||||
|
||||
from fastvideo.layers.activation import get_act_fn
|
||||
from fastvideo.layers.linear import ReplicatedLinear
|
||||
from fastvideo.layers.quantization import QuantizationConfig
|
||||
|
||||
|
||||
class MLP(nn.Module):
|
||||
@@ -21,18 +22,27 @@ class MLP(nn.Module):
|
||||
act_type: str = "gelu_pytorch_tanh",
|
||||
dtype: torch.dtype | None = None,
|
||||
prefix: str = "",
|
||||
quant_config: QuantizationConfig | None = None,
|
||||
):
|
||||
super().__init__()
|
||||
self.fc_in = ReplicatedLinear(
|
||||
input_dim,
|
||||
mlp_hidden_dim, # For activation func like SiLU that need 2x width
|
||||
bias=bias,
|
||||
params_dtype=dtype)
|
||||
params_dtype=dtype,
|
||||
quant_config=quant_config,
|
||||
prefix=f"{prefix}.fc_in",
|
||||
)
|
||||
|
||||
self.act = get_act_fn(act_type)
|
||||
if output_dim is None:
|
||||
output_dim = input_dim
|
||||
self.fc_out = ReplicatedLinear(mlp_hidden_dim, output_dim, bias=bias, params_dtype=dtype)
|
||||
self.fc_out = ReplicatedLinear(mlp_hidden_dim,
|
||||
output_dim,
|
||||
bias=bias,
|
||||
params_dtype=dtype,
|
||||
quant_config=quant_config,
|
||||
prefix=f"{prefix}.fc_out")
|
||||
|
||||
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
||||
x, _ = self.fc_in(x)
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user