Compare commits

...
Author SHA1 Message Date
Davids048andClaude Fable 5.1 1ea7278c00 [refactor]: group model-family settings under pipeline.model
pipeline.model is a tagged union: a mapping with exactly one key, which
names the model family (ltx2, minimax_h3, longcat, or generic) and selects
that family's options class. The former sibling sections pipeline.ltx2,
pipeline.minimax_h3, and pipeline.longcat, and the DiT/VAE architecture
override dicts pipeline.dit and pipeline.vae, live inside the family block.
The parser implements tagged unions as a general mechanism (a dataclass
with a TAGS table); the family is readable as pipeline.model.family.

Resolution validates the tag against the family that the registry chose
for model_path (validate_model_family raises and names the registry class
and the right block) and fills the family's empty block when the input has
none (fill_model_family), so readers never see None. Family-only fills and
derivations (the MiniMax-H3 parallel-VAE environment variables, the LTX-2
tile-size derivation and refine preset copy, the LongCat BSA mirrors) run
only for their family.

Against the 3f6893a0 goldens every PipelineConfig is identical and no
other field differs beyond the allowed categories; family-only settings
are no longer decided on models of another family. LingBot-Video reads
its refiner switch from pipeline.preset_overrides.refine.enabled and loads
the refiner by default when the checkpoint has one, as at 3f6893a0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CbSuTKA9oX4YsECdcBPAsX
2026-10-05 06:04:17 +00:00
Davids048andClaude Fable 5.1 a5aa64ab6b [refactor]: decide every config value before the config freezes, and drop the from_pretrained keywords
VideoGenerator.from_pretrained(model_path, config) takes the same nested
settings as from_config; the 24 flat keywords, from_pretrained_kwargs_to_config,
FROM_PRETRAINED_KWARGS, and every flat_field in the schema are deleted.
nvfp4_fa4 is the typed field engine.attention.nvfp4_fa4, applied by a
resolution step. 65 call sites and the docs use the nested form; the three
kwargs golden cases are config cases with byte-identical results.

Every value is decided during resolution; no runtime code calls
with_override. The device offload policy (unified memory, MPS, layerwise
conflicts, lazy module load) runs as resolution steps in the main process
(fastvideo/api/device_policy.py), and the worker runs on the config it
receives. The LTX-2 refine defaults from model_index.json and the MiniMax-H3
schedule from fastvideo_inference.json are resolution steps
(fastvideo/api/checkpoint_defaults.py) that read local checkpoints or download
only those files; the H3 pipeline validates the schedule against the loaded
schedulers but no longer writes it. The preprocessing entry point passes the
downloaded local path as run state instead of a config override. Teacher and
critic models load with override_transformer_cls_name as a load argument.
The test isolation turns the device policy off, as it blocks downloads, so
that the goldens do not depend on the machine.

Against the 3f6893a0 goldens, every field still matches except boundary_ratio
and ltx2_vae_tiling (DESIGN section 9) and the fields that no longer exist.
The launcher and trainer YAML verifications report 0 unexplained differences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 07:14:31 +00:00
Davids048andClaude Fable 5.1 0bcd15d5e2 [refactor]: delete FastVideoArgs, TrainingArgs, the flattening, and argparse
ResolvedGeneratorConfig is the only runtime config. It answers only its
typed fields, pipeline_config, mode, training_mode, and inference_mode;
the transitional flat-name fallback is gone.

Deleted: fastvideo/fastvideo_args.py (FastVideoArgs, TrainingArgs, their
argparse, prepare/set/get_current_fastvideo_args), the flat-name fallback,
generator_config_to_fastvideo_args and the other flattening helpers,
PipelineConfig/model-config/PreprocessConfig/SamplingParam add_cli_args,
configs/utils.py, and the tests of the deleted code.

Materialization builds the model's PipelineConfig directly from the typed
values through an explicit table of typed path to PipelineConfig mirror
attribute (PIPELINE_CONFIG_MIRRORS); PipelineConfig.from_source replaces
from_kwargs. flat_field metadata stays only on the fields behind the 24
from_pretrained keywords; a pipeline.experimental key that names a mirror
attribute, and pipeline.preset / preset_version / components.vae_weights,
are rejected during resolution. The training root keeps the DiT on the
device through a resolution step instead of a from_pretrained override.

The goldens snapshot the end-state objects (resolved config, PipelineConfig,
decisions, request); the cli category is gone. Against the 3f6893a0
goldens, every field matches except disable_autocast, boundary_ratio,
ltx2_vae_tiling (DESIGN section 9), and the fields that no longer exist.
The schema parity inventory classifies the typed GeneratorConfig paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-04 04:45:02 +00:00
Davids048andClaude Opus 5.5 0b0c692012 [refactor]: read typed config paths in the loaders
- The component loaders read engine.offload.*, engine.precision.*,
  engine.parallelism.hsdp_*, engine.use_fsdp_inference, engine.compile.*,
  engine.attention.*, pipeline.components.*, and pipeline.flow_shift (the
  scheduler shift) from the resolved config, and compile settings through
  torch_compile_kwargs(). They no longer write model_paths (the pipeline's
  ComponentState records paths) or apply the unified-memory policy to the
  config; the text-encoder loader only queries the policy for its target
  device.
- PipelineComponentLoader.load_module and TransformerLoader.load take
  loading_teacher_critic_model instead of reading an attribute that
  training code set on the config.
- Loader, encoder, transformer, VAE, and device-policy tests build resolved
  configs; the env-var docs name typed paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:46:52 +00:00
Davids048andClaude Opus 5.5 6c2a250fd0 [refactor]: read typed config paths in the pipeline base, worker, and entry points
- ComposedPipelineBase owns a ComponentState: it records each component's
  path after loading and attaches the state to every stage it registers.
  from_pretrained(model_path, *, resolved_config, ...) takes a resolved
  config instead of building FastVideoArgs from keywords or argparse; the
  training_args alias and TrainingArgs checks are gone, and the training
  DiT offload is a recorded override. Compile settings come from
  engine.compile and torch_compile_kwargs(); LoRA settings from
  pipeline.components.lora_* and training.lora.*.
- The worker and executors read engine.num_gpus, engine.parallelism.*,
  engine.execution_backend, and pipeline.output_type; the Ray executor owns
  its placement group and runtime env.
- VideoGenerator builds from a resolved config only (_from_resolved_config);
  the prompt_txt fallback is gone (requests carry inputs.prompt_path). The
  OpenAI server state holds the resolved config (get_resolved_config()).
- Tests use resolved configs; the SSIM and performance helpers map their
  settings to dotted typed paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:35:19 +00:00
Davids048andClaude Opus 5.5 0607bd78c7 [refactor]: run the legacy training stack on a TrainingRunConfig
- The 12 fastvideo/training/*_pipeline.py entry points load a
  TrainingRunConfig from --config <yaml> plus dotted overrides (mode
  finetuning or distillation) instead of TrainingArgs argparse flags.
- The training pipelines read self.resolved_config.training.<section>.<field>
  and typed paths; the training_args alias, the inference_mode toggles,
  and the writes to the training config are gone. Teacher and critic
  models load with a with_override config and the loading_teacher_critic_model
  argument; validation pipelines are built from their own resolved config
  (resolve_validation_config), and the tracker receives to_dict().
- Each of the 37 training launchers has a YAML file next to it with the
  values its flags held; shell variables stay as dotted overrides. Against
  the 3f6893a0 TrainingArgs snapshots every value and PipelineConfig
  matches, except mode/inference_mode (the stage readers saw training mode
  before and after), the dropped flags without a reader, lora_alpha (now
  derived from the rank by a step), and the vae_tiling and compile-kwargs
  defaults.
- Validation stages read the validation pipeline's config: modules that
  validation loads itself use the inference offload defaults, and the
  bf16-autocast VAE decode of two distillation launchers becomes fp32.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:28:44 +00:00
Davids048andClaude Opus 5.5 c4bb13a650 [refactor]: decide the v1 preprocessing VAE precision during resolution
The fp32 VAE that every v1 preprocessing task except text_only needs is a
resolution step (derive_video_preprocess_vae_precision) instead of an
override that the entry point applied after materialization, so the
PipelineConfig carries it too. Sections of a resolved config pickle on
their own, so a DataLoader worker can receive one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:25:18 +00:00
Davids048andClaude Opus 5.5 d051603677 [refactor]: run preprocessing on a PreprocessRunConfig
- v1_preprocess.py and v1_preprocessing_new.py load a PreprocessRunConfig
  from --config <yaml> plus dotted overrides; their argparse flags are
  gone. The preprocessing pipelines, workflows, datasets, and stages read
  resolved_config.preprocess.* and typed paths instead of
  preprocess_config and an argparse namespace.
- The overfit preprocessors build resolved configs instead of FastVideoArgs.
- Each preprocessing launcher has a YAML file next to it with the values
  its flags held; shell variables stay as dotted overrides. Against the
  3f6893a0 snapshots of the 17 launchers, every value and PipelineConfig
  matches, except mode (preprocess instead of inference), ltx2_vae_tiling,
  and text_only's vae_config.load_encoder. Five launchers that the old
  parser rejected now carry their intended values; the never-defined
  --model_type and --validation_dataset_file flags are dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:24:03 +00:00
Davids048andClaude Opus 5.5 00a8e5bcb6 [refactor]: build resolved configs in the modular trainer
- moduleloader builds the component-loading config and the inference
  config as resolved configs (_build_training_resolved_config and
  build_inference_resolved_config) instead of TrainingArgs objects mutated
  after construction; the transformer override, weights path, and
  teacher/critic flag are config values or load_module arguments.
  keep_checkpoint_component_config keeps the trainer's PipelineConfig view
  of the checkpoint-filled VAE and text-encoder configs.
- The validation callback builds its pipeline with
  from_pretrained(path, resolved_config=...) and forwards with a resolved
  config built per run; it no longer replaces or mutates the trainer's
  pipeline_config, and generator_overrides takes nested GeneratorConfig
  fields instead of flat names.
- Against the 3f6893a0 snapshots of the 67 trainer YAML files, every
  PipelineConfig and value matches except ltx2_vae_tiling (the model's
  vae_tiling) and the unread pretrained_model_name_or_path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:19:13 +00:00
Davids048andClaude Opus 5.5 4117696a4f [refactor]: describe and use the resolved config in docs, examples, and local tests
- Docs, AGENTS.md files, and skills describe ResolvedGeneratorConfig,
  resolve_inference_config, typed YAML, and --config plus dotted
  overrides for training and preprocessing instead of FastVideoArgs and
  flat flags.
- The local tests build resolved configs (or small fakes that expose typed
  paths) instead of FastVideoArgs; the LoRA extraction scripts build
  pipelines with from_pretrained(model_path, resolved_config=...).
- The mixkit README's NVFP4 snippet uses from_config with
  engine.quantization.transformer_quant (the from_pretrained keyword it
  used is rejected).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:16:10 +00:00
Davids048andClaude Opus 5.5 23cb8c0c2b [refactor]: read typed config paths in the shared stages
- PipelineStage gains component_state; the denoising, SR denoising, and
  decoding stages reload released components through it instead of the
  model_paths and model_loaded fields on the config.
- The shared stages read engine.disable_autocast, engine.offload.*,
  engine.precision.*, engine.attention.{vsa_sparsity, moba_config},
  pipeline.{vae_tiling, flow_shift, output_type}, and
  engine.enable_stage_verification; embedded_cfg_scale_for_batch falls
  back to pipeline.embedded_cfg_scale.
- The stage tests build resolved configs (make_resolved_config) instead of
  FastVideoArgs fakes; the sequential-load CLI test uses dotted overrides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:12:37 +00:00
Davids048andClaude Opus 5.5 ca7965ed28 [refactor]: read typed config paths in the remaining model pipelines
The Cosmos, DreamX-World, Flux, GLM-Image, HunyuanVideo, HyWorld,
Kandinsky5, LingBot, LongCat, Matrix-Game, SD3.5, TurboDiffusion, and
Z-Image pipelines read flow_shift, precisions, vae_tiling,
dmd_denoising_steps, output_type, offload, autocast, and the LongCat BSA
settings from typed paths. LingBot's refine switch reads
pipeline.ltx2.refine.enabled, so the typed field also turns its refiner off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:09:11 +00:00
Davids048andClaude Opus 5.5 61dffbe904 [refactor]: read typed config paths in the LTX-2, MiniMax-H3, Wan, MMAudio, MagiHuman, and Stable Audio pipelines
- LTX-2: the refine settings are read from pipeline.ltx2.refine.* and
  pipeline.components.upsampler_weights. The checkpoint's model_index.json
  defaults fill the fields that are still None, and the remaining switches
  get their defaults, in one recorded override that rebinds the pipeline's
  config. An explicit input value now always wins over the checkpoint
  default; the flat sentinels (the generic refine_* names and the class
  default of the step count) are gone.
- MiniMax-H3: the sequential load, decode backend, TAEH3, parallel VAE,
  offload, and checkpoint schedule settings are read from typed paths.
- Wan, MMAudio, MagiHuman, Stable Audio: flow_shift, dmd_denoising_steps,
  output_type, workload_type, offload, and autocast settings are read from
  typed paths. MagiHuman no longer reads the never-set log_level_progress
  name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:06:25 +00:00
Davids048andClaude Opus 5.5 7ec6251133 [refactor]: read typed config paths in the model-specific stages
The Flux, GameCraft, Gen3C, HyWorld, Kandinsky5, LongCat, Matrix-Game,
and SD3.5 stages read engine.precision.*, engine.disable_autocast,
engine.offload.vae, pipeline.vae_tiling, pipeline.dmd_denoising_steps,
pipeline.flow_shift, and engine.attention.vsa_sparsity from the resolved
config, and component residency from the pipeline's ComponentState.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 06:01:11 +00:00
Davids048andClaude Opus 5.5 b323cb787c [refactor]: add the shared runtime helpers for reader migration
- ComponentState (fastvideo/pipelines/component_state.py) holds each
  pipeline's component paths and residency, which runtime code kept on
  FastVideoArgs as model_paths and model_loaded.
- torch_compile_kwargs(resolved_config) derives the torch.compile keyword
  arguments from engine.compile.
- thaw(value) returns a mutable copy of a value read from a resolved
  config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 05:48:05 +00:00
Davids048andClaude Opus 5.5 e263249175 [refactor]: name the runtime config parameter resolved_config
Every parameter, attribute, and keyword argument that carries the runtime
config is named resolved_config instead of fastvideo_args, including the
RPC kwargs keys between the executors and the workers. The module path
fastvideo.fastvideo_args is unchanged. The rename is mechanical (a
token-level rewrite) plus yapf line wrapping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 05:45:37 +00:00
Davids048andClaude Opus 5.5 6c776a219b [refactor]: run the inference path on the resolved config
ResolvedGeneratorConfig becomes the runtime config of every mode. Resolution
builds the model's PipelineConfig once, freezes it, and attaches it as
resolved_config.pipeline_config; it also exposes mode, training_mode, and
inference_mode. VideoGenerator, fastvideo serve, the OpenAI server state,
the executors, the worker, and ComposedPipelineBase receive the resolved
config instead of a FastVideoArgs, and generator_config_to_fastvideo_args
is no longer on the inference path.

- Schema: ExecutionMode and WorkloadType move to fastvideo/api/schema.py;
  typed homes for mode, output_type, boundary_ratio, master_port,
  moba_config, and the LTX-2 refine settings; Enum parsing.
- Resolution: steps for the refine preset overrides, flat experimental
  keys, the V-MoBA config, the validation that check_fastvideo_args did,
  the num_gpus bump, the model's vae_tiling default, and a final
  fill_runtime_defaults step for typed fields no earlier step decides.
- fastvideo/api/training_schema.py: TrainingRunConfig and
  PreprocessRunConfig roots, their resolvers, and a --config plus dotted
  override loader.
- fastvideo/api/device_policy.py: the device offload policy as functions
  that return an overridden config; the worker keeps the result. The LTX-2
  and MiniMax-H3 checkpoint defaults rebind the pipeline's config.
- fastvideo/api/flat_name_fallback.py answers the flat FastVideoArgs names
  on the resolved config until the readers move to typed paths.

Golden changes: the decisions of the steps above are recorded, and
FastVideoArgs.ltx2_vae_tiling holds the model's vae_tiling (it held None;
no runtime code reads it). Against the 3f6893a0 goldens every
FastVideoArgs field, PipelineConfig, and SamplingParam matches, except
disable_autocast and boundary_ratio, which FastVideoArgs held at their
defaults because flattening routed them only to PipelineConfig.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 05:35:14 +00:00
Davids048andClaude Opus 5.5 3f6893a098 [refactor]: accept only the convenience keywords in from_pretrained
VideoGenerator.from_pretrained accepts the 24 common engine and offload
keywords in FROM_PRETRAINED_KWARGS. Any other keyword raises TypeError
with the config path to use with VideoGenerator.from_config. The
conversion of the removed flat keywords (LTX-2 refine names, component
config prefixes, pipeline_config objects, empty-string paths) is removed
from from_pretrained_kwargs_to_config.

The golden kwargs cases that used removed keywords move to a config
category of typed configs; their golden files are byte-identical. The
DreamVerse contract test and the GPU pool forward-translation tests
checked only the removed keyword mapping; the GPU pool tests keep the
typed-config flattening checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:23:41 +00:00
Davids048andClaude Opus 5.5 d728a0dc5f [refactor]: remove VideoGenerator.generate_video
VideoGenerator.generate(request) is the one inference entry point. The
generate_video keyword call, its mapping from flat keywords to a request
(legacy_generate_call_to_request and its helpers), and the tests of that
mapping are removed. Comments and docs that named generate_video describe
generate and request extensions instead.

scripts/ltx2_sr_alignment.py keeps its generate_video call, because its
reference run imports the separate FastVideo-internal package.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:17:49 +00:00
Davids048andClaude Opus 5.5 d6ba5ced94 [refactor]: move apps and docs to generate(request) and from_config
DreamVerse, FastVideo Studio, the ComfyUI node, the README, and the
inference docs call VideoGenerator.generate with nested requests and
VideoGenerator.from_config with typed configs, and read GenerationResult
attributes instead of result dict keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:15:04 +00:00
Davids048andClaude Opus 5.5 9139c0411b [refactor]: move tests and scripts to generate(request) and from_config
The SSIM, LoRA, performance, nightly, and local tests, two scripts, and the
pipeline parity skill template call VideoGenerator.generate with nested
requests and VideoGenerator.from_config with typed configs. The SSIM and
performance helpers convert their flat keyword tables into a typed config
and a request after the overrides merge. Text encoder precisions are passed
as lists, which the typed config parser requires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:14:42 +00:00
Davids048andClaude Opus 5.5 f571621ae7 [refactor]: move examples to generate(request) and from_config
Every example calls VideoGenerator.generate with a nested request instead
of generate_video keywords, and passes settings outside the from_pretrained
convenience keywords through VideoGenerator.from_config at their typed
paths. Each request and generator config equals what the keyword call built.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:14:36 +00:00
Davids048andClaude Opus 5.5 96b7c0d223 [refactor]: build OpenAI image requests as typed requests
The image routes call VideoGenerator.generate with a GenerationRequest
instead of passing flat generate_video keywords. The flat projection
_build_generation_kwargs in video_api.py had no runtime caller, so it is
removed, and its tests exercise build_generation_request instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 02:03:51 +00:00
Davids048andClaude Opus 5.5 74a52c027b [refactor]: remove the flat-flag entry point of the OpenAI API server
python -m fastvideo.entrypoints.openai.api_server parsed the hand-written
FastVideoArgs argparse flags, an inference input path next to the typed config
that fastvideo serve reads. It is removed; start the server with
fastvideo serve --config SERVE_CONFIG and dotted overrides such as
--server.port 8000. The GPU integration test starts the server that way, with
the same model, GPU count, and DiT offload.

FastVideoArgs.add_cli_args stays for the legacy training and preprocessing
scripts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 01:54:34 +00:00
Davids048andClaude Opus 5.5 2cd2018137 [refactor]: build OpenAI requests as nested typed requests
build_generation_request flattened the serve config's default_request with
explicit_request_updates, merged the request body into the flat dict, and then
rebuilt a nested GenerationRequest through legacy_generate_call_to_request,
the generate_video keyword mapping. It now builds the nested request directly:

- default_request contributes its explicit fields with their sections, through
  the new compat.explicit_request_raw;
- each value collected from the OpenAI request goes to a fixed typed path from
  _REQUEST_PATHS, and any other collected value goes to extensions.

The resulting SamplingParam is unchanged. A request-body field still wins over
the same-named field in default_request.stage_overrides, because
request_to_sampling_param applies stage overrides to sampling fields of the same
name. The only difference in the typed request: an operator stage override
that the body does not set now stays under stage_overrides instead of moving
into sampling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-03 01:52:27 +00:00
Davids048andClaude Opus 5.5 6627366805 [docs]: show how to look up where a config value came from
docs/inference/configuration.md describes resolved_config.provenance(path)
on a VideoGenerator and resolved_request.provenance(path) on a
GenerationResult, with examples checked against the Wan2.1 T2V 1.3B preset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 20:27:14 +00:00
Davids048andClaude Opus 5.5 7aa1f78e10 [refactor]: carry the resolved config through VideoGenerator and deprecate raw FastVideoArgs
VideoGenerator.__init__ takes resolved_config from the FastVideoArgs it
receives, so a generator that fastvideo serve builds through
from_fastvideo_args exposes the resolved config too, not only one built
through from_config.

from_fastvideo_args(...) with a FastVideoArgs that was not built from a typed
config now emits a DeprecationWarning that points to from_config. This is the
only public path that hands VideoGenerator a hand-built FastVideoArgs, and it
has to go before FastVideoArgs can be removed. The deprecated
fastvideo.entrypoints.openai.api_server flag entry point reaches it, as
before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 20:25:33 +00:00
Davids048andClaude Opus 5.5 38096b94a3 [refactor]: resolve request settings against the model's preset defaults
resolve_generation_request in fastvideo/api/resolution.py runs resolution
steps over the settings sections of a GenerationRequest (sampling, runtime,
output, stage_overrides) and freezes them as a ResolvedRequest. The explicit
paths are the ones the request tracked while it was parsed or built. Prompts,
inputs, state, plan, and extensions are request data and are left out, so
large tensors are not copied.

fastvideo/api/request_resolution.py adds resolve_request(request, model_path).
Its step, fill_sampling_defaults[preset <name>], sets each setting that the
request did not write to the model's preset default
(SamplingParam.from_pretrained). This is the same source that
request_to_sampling_param starts from, and a test checks that the two agree
field by field.

VideoGenerator attaches the ResolvedRequest to each GenerationResult as
resolved_request, so the source of every setting can be looked up. The golden
request snapshots record request_resolution_decisions; nothing else changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 20:23:05 +00:00
Davids048andClaude Opus 5.5 4d05e9c020 [bugfix]: apply per-request embedded_cfg_scale in the worker
A request's embedded_cfg_scale (request extension or legacy generate_video
keyword) was applied to a deep copy of FastVideoArgs.pipeline_config in
VideoGenerator. The worker runs the pipeline with its own FastVideoArgs
(GpuWorker.execute_forward passes self.fastvideo_args), so the copy never
reached the denoising stages and the request value was ignored.

embedded_cfg_scale is now a per-request sampling value:

- the request field request.sampling.embedded_cfg_scale (the extension key
  still works);
- SamplingParam.embedded_cfg_scale and ForwardBatch.embedded_cfg_scale;
- embedded_cfg_scale_for_batch(batch, fastvideo_args), which the denoising,
  SR denoising, Wan DMD, and Flux2 text-encoding stages read. It returns the
  batch value, else pipeline_config.embedded_cfg_scale, so training and
  requests without the field are unchanged.

VideoGenerator no longer copies or changes FastVideoArgs per request, and
request_to_pipeline_overrides is removed. The golden snapshots record the new
SamplingParam field, and the request with extensions.embedded_cfg_scale now
carries 3.0 on SamplingParam instead of on a pipeline override.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 20:17:03 +00:00
Davids048andClaude Opus 5.5 62780cf53b [refactor]: freeze FastVideoArgs built from a resolved config
generator_config_to_fastvideo_args passes the ResolvedGeneratorConfig to
FastVideoArgs through the resolved_config init-only argument. Such an
object keeps the config as resolved_config, skips the environment folds that
resolution has already decided, and becomes read-only after construction.
Its PipelineConfig becomes read-only too:

- Assigning a configuration field raises AttributeError and names
  FastVideoArgs.override.
- Runtime state stays writable: model_paths, model_loaded, and the Ray
  placement handles. So do nested component configs, which loaders fill
  with checkpoint data.
- A FastVideoArgs built directly, as the training and argparse entry points
  do, is unchanged and stays writable.

FastVideoArgs.override(source, values) is the one way to change fields after
construction. It records (source, values) in override_log and records the
change on resolved_config for fields that have a typed path. The writers that
changed configuration in place now go through it:

- device_policy:* for the unified-memory, lazy-load, MPS, and layerwise
  offload decisions that each worker makes after binding its device;
- checkpoint:model_index.json for the LTX-2 refine defaults;
- checkpoint:fastvideo_inference.json for the FastH3 DMD schedule;
- request for the per-request embedded_cfg_scale copy.

Behavior change: an explicit false for engine.compile.regional or
pipeline.minimax_h3.vae_parallel_decode/encode now wins over
FASTVIDEO_INFERENCE_TORCH_COMPILE and FASTVIDEO_VAE_PARALLEL_*. The resolution
steps fill these fields only while they are unset, and FastVideoArgs no
longer folds the environment a second time.

A test run that froze every FastVideoArgs logged no other non-test writer in
the stages, pipelines, loader, worker, hooks, entrypoints, and platforms
suites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 20:08:29 +00:00
Davids048andClaude Opus 5.5 40ff9af3fb [refactor]: derive VAE tiling from LTX-2 tile sizes during resolution
FastVideoArgs._apply_ltx2_vae_overrides turns on PipelineConfig.vae_tiling
when an LTX-2 VAE tile size is set and ltx2_vae_tiling is unset. The
resolution step derive_vae_tiling_from_ltx2_tile_sizes makes the same
decision on pipeline.vae_tiling, so the resolved config holds it with its
source.

PipelineConfig is unchanged. In the golden case kwargs/ltx2_vae_tile_sizes,
the carrier field FastVideoArgs.ltx2_vae_tiling is now True instead of
None, because the decision now reaches FastVideoArgs as input.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:24:08 +00:00
Davids048andClaude Opus 5.5 93d921fe33 [misc]: add a golden case for the LTX-2 VAE tile-size keywords
kwargs/ltx2_vae_tile_sizes records that setting an LTX-2 VAE tile size turns
on VAE tiling and reaches the VAE config, before that rule moves into
resolution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:22:23 +00:00
Davids048andClaude Opus 5.5 d26aa6b2d7 [refactor]: fill unset typed fields from the model's PipelineConfig during resolution
The resolution step fill_pipeline_config_defaults[<PipelineConfig class>]
fills each typed field that is still None with the value that the model's
PipelineConfig holds for the field's flat name. That PipelineConfig is the
one flattening starts from: the registry class for model_path, plus a
pipeline config file or object when one is given. The covered fields are the
engine.precision fields, pipeline.vae_sp, flow_shift, embedded_cfg_scale,
dmd_denoising_steps, and the pipeline.longcat BSA settings.

The resolved config therefore holds the effective value of these fields,
and provenance tells a model default apart from user input. The step runs
after the environment steps and before the derived values, so the precedence
is user input, then environment variables, then model defaults.
inference_resolution_steps(config) builds the ordered step list.

The values that the step fills are the values that PipelineConfig already
had, so FastVideoArgs and PipelineConfig are unchanged; the golden snapshots
add the decisions. Tests that flatten unregistered model paths skip this
step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:21:25 +00:00
Davids048andClaude Opus 5.5 b9c397dd1a [bugfix]: reject unsupported FASTVIDEO_ATTENTION_BACKEND names
An unsupported FASTVIDEO_ATTENTION_BACKEND name, such as a typo, was ignored
and the run fell back to automatic backend selection without telling the
user. An explicit attention_backend argument with the same name raised.

get_env_variable_attn_backend now raises ValueError for an unsupported name,
using the same parse as an explicit request (coerce_attn_backend). Every
reader goes through it: the fill_attention_backend_from_env resolution
step, FastVideoArgs.__post_init__, and the attention selector's
environment fallbacks.

The golden case environment/attention_backend_unknown_name now records the
error, and the registry description and env_vars.md table say that an
unsupported name raises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:09:50 +00:00
Davids048andClaude Opus 5.5 867f960bf0 [bugfix]: commit the environment-variable golden snapshots
The golden snapshots of the environment-variable cases lived under
config_snapshot_goldens/env/, which the repository .gitignore entry "env"
ignores, so they were never committed and the five cases would fail in CI
with a missing golden file. The case category is renamed environment, and its
golden files, generated from the previous commit, are added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:04:40 +00:00
Davids048andClaude Opus 5.5 1d7e7e2e35 [refactor]: fold inference env vars and parallel placeholders during resolution
The first resolution steps take over startup rewrites that
FastVideoArgs.__post_init__ and check_fastvideo_args make in place:

- fill_attention_backend_from_env: FASTVIDEO_ATTENTION_BACKEND fills
  engine.attention.backend while it is unset.
- fill_regional_compile_from_env: FASTVIDEO_INFERENCE_TORCH_COMPILE turns
  engine.compile.regional on.
- fill_vae_parallel_from_env: the FASTVIDEO_VAE_PARALLEL_* variables set
  pipeline.minimax_h3, and the decode strategy defaults to gather.
- derive_parallel_sizes: the -1 placeholders of tp_size, sp_size, and
  hsdp_shard_dim become 1, num_gpus, and num_gpus.

Each step repeats the current FastVideoArgs rule exactly, so FastVideoArgs is
unchanged: the golden snapshots differ only in their resolution decisions.
The FastVideoArgs code stays for FastVideoArgs built directly by the
training and argparse entry points; for a resolved config it finds the
values already decided.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 19:01:09 +00:00
Davids048andClaude Opus 5.5 0d93b16fdd [refactor]: resolve inference configs before flattening them
generator_config_to_fastvideo_args now runs resolve_inference_config, in
the new fastvideo/api/inference_resolution.py, before it flattens the
config. The resolution goes through the ordered, provenance-recording driver
in fastvideo/api/resolution.py. VideoGenerator.from_config keeps the frozen
result as resolved_config, so the source of every startup value can be
inspected.

A mapping input is the raw user input, and every leaf in it counts as
explicit. A GeneratorConfig object does not record which fields the caller
wrote. written_fields therefore treats the fields that differ from the
schema defaults as explicit, and parsing its output gives back an equal
config.

INFERENCE_RESOLUTION_STEPS is empty in this commit, so FastVideoArgs is
unchanged. The golden snapshots of the YAML and from_pretrained cases add
the list of resolution decisions, which is empty here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 18:58:12 +00:00
Davids048andClaude Opus 5.5 7ba7dbbbd5 [misc]: deprecate the flat-flag entry point of the OpenAI API server
python -m fastvideo.entrypoints.openai.api_server parses the hand-written
FastVideoArgs argparse flags, a second input system next to the typed config
that fastvideo serve reads. It keeps working unchanged during the transition
and logs a warning that points to fastvideo serve --config with dotted
overrides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 18:30:31 +00:00
Davids048andClaude Opus 5.5 c7c2d77fae [bugfix]: apply the LTX-2 refine LoRA to the refine transformer
The from_pretrained keyword ltx2_refine_lora_path set
pipeline.components.lora_path, which reaches FastVideoArgs.lora_path and loads
as the main LoRA on the stage-1 transformer. The refine stage reads
FastVideoArgs.ltx2_refine_lora_path, which stayed unset. The conversion also
turned an empty string into None. An empty string is how callers such as the
Dreamverse GPU pool disable the refine LoRA, and None makes LTX2Pipeline load
the checkpoint default instead: fastvideo_refine_lora_path in model_index.json
is FastVideo/LTX2-Distilled-LoRA for the LTX-2 Diffusers checkpoints.

The refine LoRA gets its own field, pipeline.ltx2.refine.lora_path, whose
flat name is ltx2_refine_lora_path. None keeps the checkpoint default, and an
empty string disables the refine LoRA. The Dreamverse contract test, the
Dynamo contract example, and the parity inventory point at the field.

Golden changes: kwargs/ltx2_refine_lora_path now records the path on
ltx2_refine_lora_path instead of lora_path, and the added case
kwargs/ltx2_refine_lora_disabled pins the empty string.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 18:29:37 +00:00
Davids048andClaude Opus 5.5 991fa391ff [refactor]: name the model-family config partitions *Options
The schema class MiniMaxH3Config had the same name as the MiniMax-H3 DiT
config class in fastvideo/configs/models/dits/minimax_h3.py. The family
partitions of PipelineSelection are renamed LTX2Options, LTX2RefineOptions,
MiniMaxH3Options, and LongCatOptions, so their names cannot be confused with
the model and pipeline config classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 17:49:25 +00:00
Davids048andClaude Opus 5.5 ecb15bf9e0 [refactor]: move repository configs and examples off pipeline.experimental
Every key that repository YAML configs, Python examples, the Dreamverse
backend, and tests passed through pipeline.experimental has a typed
GeneratorConfig field, so they use those fields instead:
engine.attention, engine.compile.regional, pipeline.flow_shift,
pipeline.embedded_cfg_scale, pipeline.dmd_denoising_steps,
pipeline.components.lora_nickname, pipeline.minimax_h3, and
pipeline.longcat. No repository config outside the tests uses
pipeline.experimental any more.

The flat FastVideoArgs keywords that each migrated YAML file and each
basic_fasth3.py option set produces match the previous ones. The golden
snapshots of the two Hunyuan scripts now record embedded_cfg_scale and
flow_shift as floats (6.0 instead of 6), because the typed fields are
floats. The guidance tensor is built as float32 either way. The golden case
untyped_experimental_keys is renamed attention_precision_and_flow_shift with
identical content, and keys_without_typed_fields covers keywords that still
go to pipeline.experimental.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 17:37:22 +00:00
Davids048andClaude Opus 5.5 b24396561a [refactor]: add component overrides and model-family config partitions
Parameters that belong to one model family, or to one component config,
could be set only through argparse flags or pipeline.experimental. They gain
typed GeneratorConfig fields, each with the flat name that FastVideoArgs or
PipelineConfig uses:

- pipeline.dit and pipeline.vae: dicts of DiT and VAE config field
  overrides, passed as dit_config.<key> and vae_config.<key>
- pipeline.ltx2: VAE tile sizes, initial video and audio latent paths, noise
  order, distilled sigmas, and refine.{transformer,noise,audio_noise}_path
- pipeline.minimax_h3: sequential load, video decode backend, TAEH3
  checkpoint and chunk size, and the sequence-parallel VAE switches
- pipeline.longcat: the block sparse attention (BSA) settings

legacy_from_pretrained_to_config routes dit_config.* and vae_config.*
keywords into the new dicts. Every new scalar field defaults to None and is
not passed to FastVideoArgs while unset. Behavior is unchanged: the golden
config snapshots match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 17:28:11 +00:00
Davids048andClaude Opus 5.5 0b80dd52db [refactor]: add typed fields for attention, precision, and pipeline knobs
These inference parameters could be set only through argparse flags or the
untyped pipeline.experimental dict. Each one gains a typed GeneratorConfig
field whose flat name is its FastVideoArgs or PipelineConfig field, so YAML,
dotted CLI overrides, and from_pretrained keywords all reach it:

- engine.attention: backend, vsa_sparsity, vsa_tile_size, moba_config_path
- engine.precision: dit, vae, vae_decode, image_encoder, text_encoders
- engine.compile.regional (inference_torch_compile)
- pipeline: vae_sp, flow_shift, embedded_cfg_scale, dmd_denoising_steps
- pipeline.components.lora_target_modules

Every field defaults to None, which keeps the runtime or model default; a
None field is not passed to FastVideoArgs. pipeline.experimental still
accepts the flat names and still wins over the typed fields. The parity
inventory lists the typed paths, and the docs show the typed forms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 17:24:19 +00:00
Davids048andClaude Opus 5.5 04878f0584 [refactor]: declare flat keyword names on GeneratorConfig fields
Each GeneratorConfig field that has a name in the flat keyword API
(VideoGenerator.from_pretrained kwargs and FastVideoArgs fields) declares it
through flat_field(...) metadata. legacy_from_pretrained_to_config and
generator_config_to_fastvideo_args read those names instead of keeping two
hand-written mappings, so adding a typed field with a flat name needs no
compat change.

The keywords that need a conversion keep explicit branches: the
torch_compile_kwargs split, the transformer_quant instance, the LTX-2 refine
preset keys, empty-string refine paths, and a non-string pipeline_config.
_SPECIALLY_MAPPED_PATHS lists every field without a flat name, and
test_generator_config_flat_names.py fails when a field has neither.

Behavior is unchanged: the 150 golden config snapshots match. The flat
keyword dict no longer passes None for revision, dist_timeout, and
lazy_module_load, and always passes the default LoRA nickname and strength and
empty per-component compile dicts; each equals the FastVideoArgs default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 17:17:45 +00:00
Davids048andClaude Opus 5.5 cd1ffc09c9 [refactor]: add golden snapshots of resolved inference configuration
Record the final FastVideoArgs, PipelineConfig, and SamplingParam for every
registered model path, every repository YAML config with a generator section,
and a set of from_pretrained kwargs, argparse, environment-variable, and
GenerationRequest cases. Each case runs with registered FastVideo environment
variables unset and model-index downloads blocked, so the snapshot depends only
on the source tree. Later configuration refactors compare against these files
to show that behavior is unchanged, or to make an intended change visible as a
JSON diff.

Regenerate with: python -m fastvideo.tests.api.config_snapshot

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 10:03:56 +00:00
Davids048andClaude Opus 5.5 db4a60c5db [refactor]: add ordered config resolution with provenance
Add fastvideo/api/resolution.py. resolve_generator_config parses a raw
GeneratorConfig mapping and runs an ordered list of steps over it. Each
step receives a read-only ResolutionView and returns {dotted_path: value};
it never assigns fields. The driver records every decision as
(source, values), where source is the step's qualified name, and the last
decision for a path wins. Explicit paths are taken from the keys of the
raw mapping, so "the user set this" is never inferred from defaults.

The result is a frozen ResolvedGeneratorConfig with the same nesting as
GeneratorConfig, per-path provenance, and with_override(source, values)
as the only way to change it after resolution; every override is logged.

The module is not wired into FastVideoArgs, compat.py, or any entry
point yet, so runtime behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012M91wnVFmPEvJ9h39r5BH7
2026-10-02 09:55:15 +00:00
Kyle Hu 9491c8638a [bugfix]: keep the full audio track when saving through the ffmpeg pipe (#1910) 2026-10-01 22:22:17 -07:00
Ishan 6809a751fb [bugfix] Fixed the MiniMax-H3 pin-fallback decode test setup (#1909) 2026-10-01 22:17:46 -07:00
Kyle Huandaryan5v 8322b01815 [bugfix]: restore the NVFP4 module after test_nvfp4_config re-imports it (#1908)
Co-authored-by: aryan5v <email.aryan1005@gmail.com>
2026-10-01 22:17:32 -07:00
Junda Su 9edc8adf5f [misc] Move test-suite environment access into the registry and allowlists (#1898) 2026-09-30 00:15:39 -07:00
e1f3904799 [kernel] sm_100a CUDA backward for VSA block-sparse attention, 128-token blocks (#1866)
Co-authored-by: Pengcheng Li <pengchengl@ptyche0203.ptyche.clusters.nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-29 16:57:22 -07:00
Junda Su 02f1ce11ae [misc] Move environment var reads into env registry and rename unprefixed variables (#1897) 2026-09-29 14:36:16 -07:00
Aryan KumarandAryan Kumar 7f03e03dc6 [docs]: FastH3 V2 cookbook serving and MLX install guide (#1884)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-29 13:03:18 -07:00
Junda Su e3b88bb12a [feat] Add a typed environment-variable registry, policy doc, and contract test (#1896) 2026-09-29 12:50:45 -07:00
Junda Su cb66acd400 [bugfix] Fix environment-variable bugs in LTX-2 debug logging, HF token lookup, and attention backend reads (#1895) 2026-09-29 11:48:31 -07:00
li-lizhe 442e2d2e18 [bugfix] fix(metrics): set SceneMetric device_map for any non-CPU device (#1817) 2026-09-28 10:24:00 -07:00
Max LI dd35763ad6 [bugfix] Fix MLX prompt enhancement sampler compatibility (#1891) 2026-09-27 14:53:25 -07:00
Keith e90be598e5 [perf] MiniMax H3: return uint8 frames from the decode worker (#1828) 2026-09-24 18:08:56 -07:00
Leleand武垚乐 ba5e81083c [bugfix] Preserve video frames when decoded audio is shorter (#1857)
Signed-off-by: 武垚乐 <wuyaole@mininglamp.com>
Co-authored-by: 武垚乐 <wuyaole@mininglamp.com>
2026-09-24 17:20:37 -07:00
YZJF 76ce9c7fd6 [bugfix] Reject mismatched model names on image generation routes (#1881) 2026-09-24 17:19:29 -07:00
Junda Su 08d99c089e [perf] Reuse cached VAE offload for H3 conditioning (#1882) 2026-09-24 17:18:39 -07:00
Jerry-XY 20751a21aa [bugfix] Keep the VSA-H3 tile buffer out of the autograd graph (#1858) 2026-09-24 17:16:49 -07:00
alanhuangyooandSolitaryThinker 9dd2a837f4 [bugfix] encoders: keep Qwen2.5-VL importable on transformers 5 (#1791)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-24 17:13:23 -07:00
alanhuangyoo 93aab45ac2 [bugfix] tests: include ltx2_3_base in local LTX-2 preset set (#1750) 2026-09-24 17:11:31 -07:00
Yogya Mehrotra 017ce6602d [bugfix] Fix dead NVFP4 QAT plumbing: bench imports, int8 error string, smooth_q crash (#1723) 2026-09-24 17:09:33 -07:00
Yogya Mehrotra a575055eec [perf] benchmark_attn_qat_train: device peak-TFLOPS table instead of silent RTX 5090 default (#1724) 2026-09-24 17:08:03 -07:00
Kyle Hu 81f3fec7fd [bugfix]: workers killed by a signal now log the reason (#1725) 2026-09-24 17:06:55 -07:00
Raghav K d265a454bf [bugfix]: Fall back when large pinned output allocation fails (#1759) 2026-09-24 17:05:59 -07:00
c100c66578 [bugfix]: isolate cancelled streaming requests from GPU result readers (#1848)
Co-authored-by: Gxj230958 <222823329+Gxj230958@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-24 17:04:59 -07:00
Aryan Kumar 8760eb7a06 [feat] Add FastH3 8-Step V2 MLX inference (#1863)
Keep the four-step uniform AdaLN cache, and use explicit DMD rungs only when the snapshot declares its own scheduler shifts.
2026-09-23 12:42:03 -07:00
Aryan Kumar 361f919c88 [feat]: add a narrow FastH3 8-step MLX ladder
Keep the four-step uniform AdaLN cache, and use explicit DMD rungs only when the snapshot declares its own scheduler shifts.
2026-09-23 11:06:55 -07:00
Aryan Kumar 8b5377aab2 Merge remote-tracking branch 'origin/main' into aryan/pr-1863-fasth3-8step 2026-09-23 10:52:20 -07:00
Yixing Wangandleo d995516da0 [bugfix] Restore page interaction after dismissing Create Job or successfully creating a job (#1869)
Co-authored-by: leo <yixingwang@YIXINGs-MacBook-Pro.local>
2026-09-21 12:03:54 -07:00
Keith f47ad3f5b7 [bugfix] fastvideo-kernel: install the Python + Triton package when the HIP toolchain cannot be configured (#1871) 2026-09-21 11:59:13 -07:00
Hexu ZhaoandClaude Opus 5 10bdf5e076 [bugfix] frozen offload: no host copy for an already resident module
`load` took the host copies before checking which tensors were missing from
the device, so with `vae_cpu_offload=False` — where the loader builds the VAEs
on CUDA and `unload` is never called — it pulled about 11 GB per rank back over
PCIe once and kept a pinned mirror of it for the life of the process, for a
module that never leaves the device.

It now asks what is absent first and returns early when nothing is. The
offloaded path is unchanged: weights still round-trip and come back bit
identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:46:13 -07:00
Hexu ZhaoandClaude Opus 5 3e26db40b0 [perf] H3 decode: pin the frozen VAEs' host copies
With the device-to-host copy gone, the remaining host-to-device copy is the
stage's cost, and from pageable memory it is staged through a bounce buffer at
roughly 2.6 GB/s instead of full PCIe speed. The host copies are made once, so
pinning them costs nothing per request.

It does cost the module's size in non-pageable host memory (10.4 GB for the
H3 video VAE, 0.6 GB for the audio VAE, per rank) for as long as the pipeline
is alive. It follows --pin-cpu-memory, which is on by default; --no-pin-cpu-
memory keeps the copies pageable and only loses the bandwidth. Note that this
flag did not already reach the VAE: its other use is FSDP's own CPUOffloadPolicy
and the VAE does not go through FSDP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ATw7g6cxq9N2LtnMratsru
2026-09-20 22:46:13 -07:00
Hexu ZhaoandClaude Opus 5 c73dd0ab55 [perf] H3 decode: offload the frozen VAEs without the device-to-host copy
The decode stages move the 10.4 GB fp32 video VAE (and the 0.6 GB audio VAE)
with module.to(device) / module.to("cpu") on every request. The weights are
frozen -- the whole stage runs under no_grad -- so the copy back to the host
returns bytes the host already has.

Keep one host copy per tensor, made the first time the module is loaded, and
on unload point .data back at it instead of copying. One H2D per request, no
D2H at all; the module still lives on the host between requests. Where the old
path allocated a fresh host tensor per cycle and freed the previous one, this
reuses one buffer, so peak host memory can only go down.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ATw7g6cxq9N2LtnMratsru
2026-09-20 22:46:13 -07:00
430e52154e [bugfix]: Restore exact Qwen3-VL vision interpolation (#1737)
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-18 05:00:05 -07:00
vanch 384eee8aef feat(mlx): add FastH3-8-Step-V2 DMD schedule and thermal-safe MLX inference
- Support official 8-step DMD schedule from fastvideo_inference.json in MLX pipeline
- Implement MiniMaxH3SchedulerState.from_dmd_steps for exact sigma ladder calculation
- Add auto-detection of fastvideo_inference.json in checkpoint converter AdaLN precomputation
- Add inter_step_cooldown_s and per-step MLX graph/cache cleanup to prevent thermal throttling
- Support configurable VAE tile sizes to guarantee 256px tiled decode without grid artifacts
- Skip non-existent shard files gracefully in Qwen3-VL conditioner
2026-09-17 16:12:02 +08:00
KeithandSolitaryThinker c4824c7764 [bugfix] FP8: gate the ROCm _scaled_mm path on CDNA4 (gfx950) instead of the capability tuple (#1859)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-16 19:58:55 -07:00
Raghav K 9b0e57fe4b [ci] Make Dreamverse provider race test deterministic (#1729) 2026-09-15 14:13:56 -07:00
William Lin 0100218594 [feat] Support the FastH3 8-Step V2 checkpoint: checkpoint-defined shifts, explicit DMD schedule, new example (#1852) 2026-09-15 14:13:45 -07:00
li-lizhe 39718cd54d [bugfix] fix(cosmos): make AdaLayerNorm autocast device-agnostic (#1818) 2026-09-15 07:44:51 -07:00
IshanandSolitaryThinker 37d06a832f [feat] Add fastvideo serve configs for Wan CUDA models (#1801)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 18:23:35 -07:00
sudhirpol522 8839ba8d4d [bugfix] Add OpenAI-compatible image generation endpoint (#1840) 2026-09-14 17:55:56 -07:00
sudhirpol522andSolitaryThinker 61b91220c0 [bugfix] Validate image response format before generation (#1841)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 17:01:40 -07:00
Lele 316f3876c2 [bugfix]: allow MiniMax H3 frame padding at the 15-second limit
Accept the causal-VAE-aligned 362-frame bucket (15.083 s) for 15-second H3 requests.

- Hoist MINIMAX_H3_MIN/MAX_ALIGNED_FRAMES next to align_num_frames and reuse them in the CUDA stage, the MLX runtime, and the LoRA example so the bound is consistent across entry points.
- State the accepted frame range in the error messages.
- Cover seconds="15" -> 362 at OpenAI admission and the MLX resolve_geometry bound.
- Update the Spark/openai cookbook frame caps from 345 to 362.

Known gaps: no GPU-lane coverage for the 362 bucket (SSIM runs 124 frames; the golden gate is a fixed-geometry fingerprint), and explicit num_frames=360 still hits the pre-existing grid gate.
2026-09-14 16:54:41 -07:00
Yaegaki1Erika 614b59543c [bugfix]: respect serialized tensor dtype in parquet dataloader (#1843) 2026-09-14 16:34:25 -07:00
William Lin 1c14afd559 [docs]: refresh AGENTS.md maps and add Wan SP/I2V and CI test notes (#1845) 2026-09-14 16:06:53 -07:00
Yaegaki1Erika 9a3c45779c [bugfix]: fix Wan I2V DMD conditioning under sequence parallelism (#1844)
The pipeline pre-sharded only the image conditioning along the temporal axis while the noise input stayed full-length, so concatenation could not match for sp_world_size > 1. WanTransformer3DModel shards the flattened token sequence after patch embedding, so pass full-length mask+latent conditioning and let the transformer shard.

Also reads temporal_compression_ratio from the VAE config, simplifies the conditioning mask, concatenates in the transformer's (bs, c, t, h, w) layout, and adds unit coverage.

Co-authored-by: Yaegaki1Erika <70182590+Yaegaki1Erika@users.noreply.github.com>
2026-09-14 15:58:13 -07:00
Kyle Hu bfc9c01797 [feat]: convert the MiniMax H3 text encoder to NVFP4 (#1838) 2026-09-12 18:48:07 -07:00
Kyle Hu 3a3ad3d209 [feat]: NVFP4 text encoder for MiniMax H3 (#1837) 2026-09-12 18:31:59 -07:00
lpc0220andClaude Fable 5.1 aef4e9b3b1 [kernel] VSA kernel: one sm_100a / sm_103a image per listed arch; un-gate backward on sm_103a (#1833)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:49:36 -07:00
a943220c11 [bugfix]: drop dead h3_sequential_load from Spark FastH3 presets (#1831)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-08 12:38:06 -07:00
William Lin 556ac7088e [refactor] Simplify Wan sampling and tests (#1825) 2026-09-07 15:57:19 -07:00
Junda Su e7456f1b75 Add H3 support into Dreamverse (#1800) 2026-09-07 13:09:44 -07:00
lpc0220 c993d7393e [kernel] sm_100a CUDA backward for VSA block-sparse attention (blk64) (#1819) 2026-09-06 21:00:18 -07:00
William Lin 7f83164233 [refactor] Move Wan VAE into the Wan package (#1824) 2026-09-05 18:43:55 -07:00
Junda Su 4e52f47d1e [feat] add MXFP8 support on H3 (#1796) 2026-09-05 17:08:58 -07:00
William Lin e19913f6e9 [refactor] Group Wan transformer and config (#1823) 2026-09-05 17:05:32 -07:00
William Lin 2413a57651 Disable old SSIM models (#1820) 2026-09-04 21:20:48 -07:00
SYLAR 7bb76b5ec9 [feat] Add native SM103a VSA support (#1812)
Signed-off-by: lishunyang12 <lishunyang12@163.com>
2026-09-03 15:00:28 -07:00
Aryan KumarandAryan Kumar 0bd19a976b [docs]: add one-Spark FastH3 cookbook runtime with a device-count row (#1811)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-02 10:05:59 -07:00
William Lin 40b93784d2 [docs]: add cookbook link to README (#1810) 2026-09-01 12:42:19 -07:00
Aryan Kumar 33d3478bad [docs] Announce local FastH3 support (#1809) 2026-09-01 12:28:04 -07:00
3d8ac9d14b [feat]: collapse cookbook recipe pages into an accordion layout, add … (#1805)
Co-authored-by: Vaish, Ishan <isvaish@UCSD.EDU>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 01:51:51 -07:00
aaef49bfc6 [feat]: run FastH3 across two DGX Sparks with Ray sequence parallel (#1803)
Co-authored-by: Kyle <shh075@ucsd.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-01 01:37:23 -07:00
Aryan KumarandAryan Kumar cf6a00b9be [feat]: add opt-in CUDA TAEH3 preview decode for FastH3 (#1795)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 23:34:56 -07:00
Shahrad ZomorrodiandShahrad Zomorrodi 1ae39562dd [bugfix] Write generated documentation as UTF-8 (#1797)
Co-authored-by: Shahrad Zomorrodi <264690209+shahradzomorrodi@users.noreply.github.com>
2026-08-31 22:45:03 -07:00
Aryan KumarandAryan Kumar 26064193e2 [feat] Add an H3 server cookbook and prompt playground (#1798)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 22:14:33 -07:00
William Lin 8446fc003e [docs]: add FastH3 Preview v1 news links (#1804) 2026-08-31 22:12:02 -07:00
Aryan KumarandAryan Kumar a28f2bab4b [feat] Add an optional MLX TAEH3 preview decoder (#1794)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 05:08:06 -07:00
Aryan KumarandAryan Kumar f82d8be4bf [perf] Sequential MiniMax H3 start with GPU-direct DiT load (#1793)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:20:21 -07:00
Aryan KumarandAryan Kumar 8e1775183e [perf] Speed up exact MiniMax H3 MLX inference (#1792)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:14:55 -07:00
Aryan KumarandAryan Kumar 620bc36dc4 [docs]: Cookbook catalog improvements (#1790)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 17:50:57 -07:00
Suhaan Khurana 29ff16ec96 [feat] Add MiniMax H3 MLX spatial fast mode (#1789) 2026-08-30 17:47:02 -07:00
Aryan KumarandAryan Kumar 8f9d76a80d [perf]: dispatch wide-M affine H3 MLX linears through dequant plus dense GEMM (#1788)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 14:42:01 -07:00
a4d9a75e2c [perf] Add MiniMax H3 MLX VSA and SIMD attention (#1776)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-30 13:44:54 -07:00
KyleNeverGivesUp b2db0c0a13 [ci]: seed stable GB10 grad-norm references (#1756) 2026-08-30 03:09:18 -07:00
Kevin Lin 6aa7d8a278 [misc] FastVideo Studio UI Additions (H3 Ref2V support) (#1783) 2026-08-30 03:06:47 -07:00
Aryan KumarandAryan Kumar ccc9014430 [docs] Add model-family inference cookbook (#1787)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 02:26:10 -07:00
William Lin a159b63c67 [bugfix] Harden OpenAI serving after post-merge review (#1782) 2026-08-28 22:29:24 -07:00
ac48bb3cd1 [feat] Add MiniMax H3 MLX T2VA inference (#1770)
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
2026-08-28 12:54:33 -07:00
William Lin 3987b9ddcd [feat] Align multimodal OpenAI serving APIs (#1781) 2026-08-28 10:09:58 -07:00
William Lin c7da2f5d60 [chore]: release v0.2.1 (#1778) 2026-08-28 02:03:18 -07:00
William Lin 39ae1decc0 [misc] pin fastvideo-kernel to exact 0.3.5 (#1777) 2026-08-28 02:02:27 -07:00
William Lin 1aed667377 [chore] release fastvideo-kernel 0.3.5 (#1775) 2026-08-27 23:18:59 -07:00
William Lin c1612ff397 [bugfix]: pin fastvideo-kernel to Torch 2.12.0 (#1774) 2026-08-27 23:13:54 -07:00
William Linandshaoxiongduan a534ba20a0 [feat] Add MiniMax H3 LoRA inference and preview launchers (#1771)
Co-authored-by: shaoxiongduan <shaoxiongduan@gmail.com>
2026-08-27 14:43:39 -07:00
KyleNeverGivesUp e9bbaca07d [perf] Disable every offload path on unified memory, unblocking MiniMax H3 generation on one GB10 (#1715) 2026-08-26 15:32:56 -07:00
KyleNeverGivesUp 9bfa585448 [perf]: stop holding the whole checkpoint during DiT load, unblocking MiniMax H3 on one GB10 (#1714) 2026-08-26 15:23:51 -07:00
Raghav K b2062556a9 [perf] VSA Triton: widen the autotune num_stages range (the optimum was outside it) (#1706) 2026-08-26 12:30:56 -07:00
KyleNeverGivesUp c9c5585758 [perf]: MiniMax H3 on GB10 - skip text encoder CPU offload on unified memory (5m49s to 30ms) (#1710) 2026-08-26 12:03:14 -07:00
William Lin 9212f4f218 [ci] make GPU validation change-aware (#1747) 2026-08-25 21:26:25 -07:00
Aryan KumarandAryan Kumar 6388db815b [bugfix] FastMetal-QAD MLX support: refuse CUDA QAD trees, use packed mlx_dit config, stream loads (#1736) (#1758)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-25 14:51:40 -07:00
lpc0220 7a4285189f [kernel] Route block-sparse VSA to the sm_100a forward behind FASTVIDEO_VSA_SM100A (opt-in) (#1754) 2026-08-24 15:26:56 -07:00
William Lin a837fe841a [docs] Update FastH3 README (#1749) 2026-08-23 05:05:58 -07:00
William Lin f9e3680f11 [perf] Align FastH3 optimized inference profile (#1748) 2026-08-23 02:01:03 -07:00
William Lin 98f761ec45 [bugfix] validation: inherit the trained denoising ladder (#1738) 2026-08-22 23:09:26 -07:00
Shao Duan c041318f2c [perf] Add fused NVLink all-to-all for Ulysses (#1740) 2026-08-22 23:06:24 -07:00
William Lin 604e0205a4 [perf] Keep odd MiniMax-H3 VSA tiles on sm100a (#1745) 2026-08-22 18:29:30 -07:00
William Lin 13213395b4 [perf] Parallelize MiniMax-H3 VAE over sequence ranks (#1744) 2026-08-22 18:05:31 -07:00
William Lin 46afee5998 [bugfix] Classify MiniMax-H3 inference controls in schema inventory (#1743) 2026-08-22 17:17:02 -07:00
William Lin c488fa1211 [perf] Add opt-in packed-varlen FA4 for MiniMax-H3 (#1742) 2026-08-22 17:16:47 -07:00
William Lin d3cff517cd [perf] Add opt-in regional fullgraph compile for DiT inference (#1741) 2026-08-22 12:14:39 -07:00
Junda Su 2f3d407406 [perf] Optimize MiniMax H3 VAE decoding (#1734) 2026-08-21 14:57:32 -07:00
Kaiqin Kong bcffa4026e [perf] Optimize MiniMax-H3 text encoder memory (#1732) 2026-08-21 14:57:06 -07:00
William Lin 6d6a10be7a [feat] FastVideo-Minimax-FastH3-Preview few-step example + 64-token-tile VSA-H3 inference path (#1731) 2026-08-21 12:40:09 -05:00
Kaiqin Kong 73dd105f3d [perf] Add opt-in MiniMax-H3 Sol-Engine fusions (#1735) 2026-08-21 12:39:28 -05:00
William LinandClaude Fable 5 56d4a6074f [bugfix] fastvideo-kernel: fix Triton block-sparse backward logit scaling (bf16 K pre-scaling) (#1730)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 04:50:06 -05:00
Shao Duan c4ad4227c0 [misc] MiniMax-H3: move the AdaLN converter into scripts/checkpoint_conversion (#1712) 2026-08-21 02:15:00 -05:00
KyleNeverGivesUpandSolitaryThinker a63ccce73d [docs]: add a maintained inference cookbook (#1290)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-21 01:53:44 -05:00
KyleNeverGivesUp 0462e1b0e7 [perf]: MiniMax H3 - build the Qwen3-VL encoder only as far as it is read (-13.7 GB) (#1711) 2026-08-20 23:19:25 -05:00
lpc0220 907f2100ec [kernel] sm_100a CUDA block-sparse VSA forward (Blackwell), 64- and 128-token blocks (#1719) 2026-08-20 23:16:25 -05:00
Kaiqin Kong e0a3db5651 [perf] Reduce MiniMax-H3 VAE peak memory (#1703) 2026-08-20 23:13:23 -05:00
Raghav K fca45bc8e1 [perf] Quantize frames to uint8 on-device before the post-decode D->H copy (#1362) 2026-08-20 21:55:55 -05:00
Aryan Kumar 86d639c848 [docs] Announce FastMetal-QAD (#1721) 2026-08-19 13:45:49 -07:00
00338aa9ca [perf] Add FA4 CuTe backward support for VSA-256 (#1639)
Co-authored-by: Hyunsung Lee <hyunsungl@sizigistudios.com>
Co-authored-by: alexzms <3036648523@qq.com>
2026-08-19 11:46:49 -07:00
Aryan KumarandAryan Kumar 8537dcd6de [feat]: Apple Silicon MLX runtime — INT8 Wan2.1 and Wan2.2 inference (#1638)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-18 15:08:00 -07:00
William Lin 8208536cd1 [bugfix] profiler region system: record + export actually work, usability roll-up (#1691) 2026-08-09 20:35:53 -07:00
Kai 0653f8f3af [new-model] Add V2A: native MMAudio inference pipeline (#1622) 2026-08-09 17:29:27 -07:00
William Lin e0d702decb [feat] VSA for MiniMax H3: packed mixed-modality sparse attention (#1695) 2026-08-09 13:10:51 -07:00
Shao Duan 541ef014ee [perf] MiniMax-H3: rank-reduced AdaLN pruned model option (-39% params, -23 GiB VRAM) (#1699) 2026-08-09 12:31:57 -07:00
William Lin ffc1a7a58b [refactor] H3 pipeline cleanup: shared helpers, dead machinery, loop-invariant hoists (#1698) 2026-08-09 04:51:31 -07:00
William Lin 9028953625 [misc] yapf pass under CI's interpreter (3.12) + pin hook language_version (#1702) 2026-08-08 22:24:17 -07:00
KyleNeverGivesUpandClaude Opus 5 6eb95693a1 [misc]: re-run yapf on main so pre-commit passes again (#1700)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:17:52 -07:00
Junda Su c3567eb468 [feat] add Minimax H3 sft pipeline (#1688) 2026-08-07 16:12:04 -07:00
William Lin 15568f27db [perf]: H3 torch.compile + CUDA graphs (1.2-1.3x) with denoising step marking (#1689) 2026-08-06 16:23:33 -07:00
Kaiqin Kong 126a75ad63 [misc] Support partial Hugging Face model downloads (#1684) 2026-08-06 16:09:46 -07:00
Raghav KandSolitaryThinker a2bfc7cdb2 [docs] DGX Spark (GB10) performance & tuning guide + reproduction examples (#1631)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-06 14:36:38 -07:00
William Lin b963a24612 [bugfix]: hard-fail when ATTN_QAT_INFER is selected but the kernel is unusable (#1690) 2026-08-06 13:12:27 -07:00
William Lin fb7be2fe2c [ci]: add golden-gate lane — single-layer bitwise DiT fingerprints for all SSIM-covered families (#1682) 2026-08-05 12:47:03 -07:00
William Lin ab00392664 [bugfix]: add GB200 to the inline SSIM device tables #1676 missed (#1681) 2026-08-05 10:33:46 -07:00
Shao Duan 9f1e7c19d2 [bugfix] Wan I2V: CLIP image conditioning silently dropped when passed as a tensor during training (#1673) 2026-08-05 01:35:09 -07:00
Kaiqin Kong e8b0e4c61e [feat] Add MiniMax H3 (#1674) 2026-08-04 13:54:39 -07:00
Haochen Jiang 9145ffdc46 [bugfix]: keep _resolved_attention_backend out of the positional config signature (#1678) 2026-08-03 17:09:15 -07:00
William Lin c3d07c870b [bugfix]: give GB200 its own SSIM reference folder instead of B200's (#1676) 2026-08-03 15:14:25 -07:00
William Lin e8812bef0b [docs]: batched docs cleanup (landing page, links, requirements, nav) (#1644) 2026-08-02 18:03:26 -07:00
William Lin b9be2449dc [refactor]: delete the dead global attention-backend override (#1672) 2026-08-02 18:00:42 -07:00
Adhvay IyerandSolitaryThinker bc7a804618 [bugfix]: harden FastVideo Studio UI reliability and accessibility (#1659)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-02 15:26:18 -07:00
William Lin 7b094c945b [refactor]: resolve attention backend once per component at load time (#1657) 2026-08-02 15:07:42 -07:00
William Lin eeb3e8a597 [bugfix]: left-align Gemma connector tokens per batch row (#1664) 2026-08-02 15:01:34 -07:00
Suhaan Khurana 05406c5d1b [misc]: consolidate dataset download scripts under examples/datasets/ (#1667) 2026-08-02 14:21:24 -07:00
KyleNeverGivesUpandClaude Opus 5 99d04a7f98 [bugfix] keep loader-populated text encoder configs in the validation pipeline (#1669)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:10:16 -07:00
William Lin 1b2b2a0161 [bugfix]: copy text encoder outputs out of CUDAGraph static buffers (#1650) 2026-07-27 21:52:53 -07:00
William Lin 98d65835b5 [misc]: refresh stale sm_120-only validation notes in QAT recipes (#1655) 2026-07-27 19:35:29 -07:00
William Lin 422585d08f [misc]: add LTX-2.3 fine-tuning example recipes (#1651) 2026-07-27 18:42:44 -07:00
ryanM154 e59a1ce16a [bugfix] Report actual package version in fastvideo --version (#1652) 2026-07-27 16:32:01 -07:00
William Lin d71acc0eb5 [bugfix]: fix stale imports in LTX-2.3 gradio local demo (#1640) 2026-07-27 16:31:07 -07:00
William Lin af2934dd6b [feat]: FA4-FP4 ATTN_QAT_INFER on sm_100/sm_103 + NVFP4 weight purge (#1647) 2026-07-27 15:46:52 -07:00
William Lin 1801512818 [docs]: cover all registered models in the support matrix (#1641) 2026-07-27 11:13:53 -07:00
William Lin 5ae05b032e [misc]: add LTX-2 fine-tuning example recipes (#1645) 2026-07-27 11:11:49 -07:00
Mac Lee 7a592ff09a [ci]: cache FastVideo kernel builds in Modal (#1562) 2026-07-26 04:21:09 -07:00
Lev NovitskiyandClaude Sonnet 5 8b23984c79 [feat] Add Kandinsky5 QAD training pipeline: data preprocessing, QAT finetune, QAT-aware DMD distillation (#1601)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 02:59:21 -07:00
1787 changed files with 224046 additions and 46083 deletions
@@ -10,9 +10,7 @@ from pathlib import Path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Clone a reference repo for FastVideo parity tests."
)
parser = argparse.ArgumentParser(description="Clone a reference repo for FastVideo parity tests.")
parser.add_argument("repo_url", help="Official reference repository URL")
parser.add_argument("target_dir", help="Directory to clone into")
parser.add_argument("--branch", help="Branch or tag to clone")
@@ -62,9 +60,7 @@ def gitignore_entry_for(target: Path) -> str:
try:
relative = resolved.relative_to(root)
except ValueError as exc:
raise ValueError(
"--update-gitignore requires target_dir to be under the current directory"
) from exc
raise ValueError("--update-gitignore requires target_dir to be under the current directory") from exc
text = relative.as_posix().rstrip("/")
return "/" + text + "/"
@@ -8,14 +8,12 @@ import os
import sys
from pathlib import Path
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Download a HF model snapshot or selected files into a local directory."
)
description="Download a HF model snapshot or selected files into a local directory.")
parser.add_argument("repo_id", help="HF repo id, for example Org/Model")
parser.add_argument("local_dir", help="Destination directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
@@ -10,7 +10,6 @@ import sys
from pathlib import Path
from typing import Any
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
RAW_WEIGHT_SUFFIXES = (".safetensors", ".pt", ".pth", ".ckpt", ".bin")
KNOWN_COMPONENTS = {
@@ -34,8 +33,7 @@ KNOWN_COMPONENTS = {
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown."
)
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown.")
parser.add_argument("source", help="HF repo id or local weights directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
parser.add_argument("--revision", help="HF revision to inspect")
@@ -94,14 +92,12 @@ def load_remote_files(
) -> list[str]:
from huggingface_hub import list_repo_files
return sorted(
list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
)
)
return sorted(list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
))
def load_remote_model_index(
@@ -215,24 +211,24 @@ def build_result(args: argparse.Namespace) -> dict[str, Any]:
"components_seen": components,
"file_count": len(files),
"file_scan_truncated": truncated,
"files_sample": files[: args.sample_limit],
"files_sample": files[:args.sample_limit],
}
def print_human(result: dict[str, Any]) -> None:
for key in (
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
):
value = result.get(key)
if value is not None:
@@ -18,7 +18,6 @@ import pytest
import torch
from torch.testing import assert_close
os.environ.setdefault("MASTER_ADDR", "localhost")
os.environ.setdefault("MASTER_PORT", "29519")
os.environ.setdefault("DISABLE_SP", "1")
@@ -35,15 +34,10 @@ FASTVIDEO_CONFIG_CLASS = "<FastVideoConfig>" # TODO.
FASTVIDEO_MODEL_MODULE = "fastvideo.models.<bucket>.<module>" # TODO.
FASTVIDEO_MODEL_CLASS = "<FastVideoModel>" # TODO.
OFFICIAL_REF_DIR = Path(
os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>")
)
LOCAL_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY)
)
CONVERTED_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR", REPO_ROOT / "converted_weights" / FAMILY)
)
OFFICIAL_REF_DIR = Path(os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>"))
LOCAL_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY))
CONVERTED_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR",
REPO_ROOT / "converted_weights" / FAMILY))
def _resolve_hf_token() -> str | None:
@@ -99,18 +93,14 @@ def _load_official_model(device: torch.device, dtype: torch.dtype) -> torch.nn.M
model = OfficialClass() # TODO: pass official config kwargs.
state_dict = {} # TODO: load official state dict from LOCAL_WEIGHTS_DIR.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
assert not missing and not unexpected, (f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.Module:
"""Load the FastVideo component with the same tensor content."""
if not CONVERTED_WEIGHTS_DIR.exists() and not LOCAL_WEIGHTS_DIR.exists():
pytest.skip(
f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}"
)
pytest.skip(f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}")
# TODO: replace with the bucket-specific FastVideo config/class/loader.
# DiT examples:
@@ -127,8 +117,7 @@ def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.
state_dict = {} # TODO: load converted or directly mapped state dict.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
@@ -187,11 +176,9 @@ def test_component_parity():
assert official_out.shape == fastvideo_out.shape
diff = (official_out - fastvideo_out).abs()
print(
f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}"
)
print(f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}")
# TODO: pick tolerance by scope:
# - single block / same kernel: 1e-4
@@ -27,7 +27,6 @@ try:
except ImportError: # pragma: no cover - optional local conversion dependency
snapshot_download = None
# TODO: fill with authoritative component prefixes for monolithic checkpoints.
# Example: {"model.model.": "transformer", "pretransform.model.": "vae"}
COMPONENT_PREFIXES: dict[str, str] = {}
@@ -47,10 +46,7 @@ SKIP_PATTERNS: tuple[str, ...] = ()
def _hf_token() -> str | None:
return (
os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
or os.environ.get("HF_API_KEY")
)
return (os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN") or os.environ.get("HF_API_KEY"))
def resolve_src(src: str, revision: str | None) -> Path:
@@ -95,11 +91,10 @@ def apply_mapping(key: str) -> str | None:
return key
def split_monolithic(
state: dict[str, torch.Tensor],
) -> dict[str, OrderedDict[str, torch.Tensor]]:
def split_monolithic(state: dict[str, torch.Tensor], ) -> dict[str, OrderedDict[str, torch.Tensor]]:
components: dict[str, OrderedDict[str, torch.Tensor]] = {
name: OrderedDict() for name in set(COMPONENT_PREFIXES.values())
name: OrderedDict()
for name in set(COMPONENT_PREFIXES.values())
}
intentionally_skipped: list[str] = []
unowned: list[str] = []
@@ -117,10 +112,8 @@ def split_monolithic(
unowned.append(key)
if unowned:
sample = ", ".join(unowned[:10])
raise ValueError(
f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}"
)
raise ValueError(f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}")
if intentionally_skipped:
print(f"Intentionally skipped {len(intentionally_skipped)} keys")
return {name: weights for name, weights in components.items() if weights}
@@ -143,8 +136,12 @@ def build_component_configs(_src_dir: Path) -> dict[str, dict[str, Any]]:
# TODO: emit config content accepted by FastVideo loaders. Most components use
# config.json; schedulers use scheduler_config.json.
return {
"transformer": {"_class_name": "<FastVideoTransformerClass>"},
"vae": {"_class_name": "<FastVideoVAEClass>"},
"transformer": {
"_class_name": "<FastVideoTransformerClass>"
},
"vae": {
"_class_name": "<FastVideoVAEClass>"
},
}
@@ -177,19 +174,13 @@ def build_model_index(
}
if revision:
index["_fastvideo_converted_revision"] = revision
return {
key: value
for key, value in index.items()
if key.startswith("_") or key in available_components
}
return {key: value for key, value in index.items() if key.startswith("_") or key in available_components}
def validate_component_configs(configs: dict[str, dict[str, Any]]) -> None:
# TODO: instantiate each FastVideo config and call update_model_arch(...) or
# update_model_config(...) with this JSON so unknown emitted keys fail here.
placeholder_configs = [
name for name, config in configs.items() if "<" in json.dumps(config)
]
placeholder_configs = [name for name, config in configs.items() if "<" in json.dumps(config)]
if placeholder_configs:
raise ValueError(f"Replace config placeholders for: {placeholder_configs}")
@@ -201,9 +192,7 @@ def verify_conversion(
del dst_dir, components
# TODO: load each emitted stateful component through its production loader and
# assert strict load, or document exact allowed missing/unexpected keys.
raise NotImplementedError(
"Implement production config validation and strict-load checks"
)
raise NotImplementedError("Implement production config validation and strict-load checks")
def write_component(
@@ -216,9 +205,7 @@ def write_component(
if component_dir.exists() and any(component_dir.iterdir()):
shutil.rmtree(component_dir)
component_dir.mkdir(parents=True, exist_ok=True)
save_file(
dict(state), str(component_dir / "diffusion_pytorch_model.safetensors")
)
save_file(dict(state), str(component_dir / "diffusion_pytorch_model.safetensors"))
if config is not None:
config_path = component_dir / config_filename(name)
with config_path.open("w", encoding="utf-8") as f:
@@ -261,9 +248,7 @@ def convert(
if layout in {"monolithic", "raw_official"}:
# TODO: replace model.safetensors with the official monolithic file name.
components = split_monolithic(
load_checkpoint(default_monolithic_checkpoint(src_path))
)
components = split_monolithic(load_checkpoint(default_monolithic_checkpoint(src_path)))
elif layout in {"separate_components", "mixed"}:
if not src_path.is_dir():
raise ValueError(f"{layout} layout requires a source directory: {src_path}")
@@ -271,9 +256,7 @@ def convert(
else:
raise ValueError(f"Unsupported template layout: {layout}")
copied = (
copy_passthrough(src_path, dst_dir) if src_path.is_dir() else []
)
copied = (copy_passthrough(src_path, dst_dir) if src_path.is_dir() else [])
configs = build_component_configs(src_path if src_path.is_dir() else src_path.parent)
validate_component_configs(configs)
for name, state in components.items():
@@ -289,9 +272,7 @@ def convert(
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--src", required=True, help="HF repo id, local dir, or checkpoint path"
)
parser.add_argument("--src", required=True, help="HF repo id, local dir, or checkpoint path")
parser.add_argument("--revision", help="HF branch, tag, or commit for repo sources")
parser.add_argument(
"--dst",
@@ -24,8 +24,8 @@ from typing import Any
import torch
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
DRILL_LAYER_ENV: str = "<FAMILY>_DEBUG_DRILL_LAYER"
HYPOTHESIS_ENV: str = "<FAMILY>_DEBUG_PATCH_<HYPOTHESIS>"
REL_THRESHOLD: float = 0.005 # 0.5% abs_mean drift flags a block as divergent
@@ -94,6 +94,7 @@ def _attach_block_hooks(
handles: list[Any] = []
def _hook(name: str):
def fn(_module, _inputs, outputs):
t = outputs[0] if isinstance(outputs, tuple) else outputs
if not torch.is_tensor(t):
@@ -101,6 +102,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(name, t)})
if tensors is not None:
tensors[name] = t.detach().float().cpu()
return fn
def _pre_hook(name: str):
@@ -114,6 +116,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(key, t)})
if tensors is not None:
tensors[key] = t.detach().float().cpu()
return fn
# TODO: adapt attribute paths to your model. Remove adapter block if absent.
@@ -131,43 +134,21 @@ def _attach_block_hooks(
# magi-human uses: attention, mlp.pre_norm, mlp.up_gate_proj,
# mlp.down_proj (pre+post), mlp, attn_post_norm, mlp_post_norm.
if hasattr(layer, "attention"):
handles.append(
layer.attention.register_forward_hook(_hook(f"{tag}.attention"))
)
handles.append(layer.attention.register_forward_hook(_hook(f"{tag}.attention")))
if hasattr(layer, "mlp"):
mlp = layer.mlp
if hasattr(mlp, "pre_norm"):
handles.append(
mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm"))
)
handles.append(mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm")))
if hasattr(mlp, "up_gate_proj"):
handles.append(
mlp.up_gate_proj.register_forward_hook(
_hook(f"{tag}.mlp.up_gate_proj")
)
)
handles.append(mlp.up_gate_proj.register_forward_hook(_hook(f"{tag}.mlp.up_gate_proj")))
if hasattr(mlp, "down_proj"):
handles.append(
mlp.down_proj.register_forward_pre_hook(
_pre_hook(f"{tag}.mlp.down_proj")
)
)
handles.append(
mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj"))
)
handles.append(mlp.down_proj.register_forward_pre_hook(_pre_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.register_forward_hook(_hook(f"{tag}.mlp")))
if hasattr(layer, "attn_post_norm"):
handles.append(
layer.attn_post_norm.register_forward_hook(
_hook(f"{tag}.attn_post_norm")
)
)
handles.append(layer.attn_post_norm.register_forward_hook(_hook(f"{tag}.attn_post_norm")))
if hasattr(layer, "mlp_post_norm"):
handles.append(
layer.mlp_post_norm.register_forward_hook(
_hook(f"{tag}.mlp_post_norm")
)
)
handles.append(layer.mlp_post_norm.register_forward_hook(_hook(f"{tag}.mlp_post_norm")))
return handles
@@ -193,11 +174,9 @@ def _write_log(entries: list[dict], path: Path) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
with open(path, "w") as f:
for e in entries:
f.write(
f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n"
)
f.write(f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n")
def _sort_key(name: str, drill_layer: int) -> tuple:
@@ -205,9 +184,14 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
return (0, "")
if name.startswith(f"L{drill_layer:02d}."):
sub_order = {
"attention": 0, "attn_post_norm": 1, "mlp.pre_norm": 2,
"mlp.up_gate_proj": 3, "mlp.down_proj<in>": 4,
"mlp.down_proj": 5, "mlp": 6, "mlp_post_norm": 7,
"attention": 0,
"attn_post_norm": 1,
"mlp.pre_norm": 2,
"mlp.up_gate_proj": 3,
"mlp.down_proj<in>": 4,
"mlp.down_proj": 5,
"mlp": 6,
"mlp_post_norm": 7,
}.get(name.split(".", 1)[1], 9)
return (1, f"block[{drill_layer:02d}]", sub_order)
if name.startswith("block["):
@@ -216,10 +200,8 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
hdr = (
f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}"
)
hdr = (f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}")
print(f"\n{hdr}\n{'-' * len(hdr)}")
first_div: int | None = None
for name in sorted(by_name.keys(), key=lambda n: _sort_key(n, drill_layer)):
@@ -235,11 +217,9 @@ def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
flag = " <<< DIVERGE"
if first_div is None:
first_div = int(name[len("block["):-1])
print(
f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}"
)
print(f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}")
return first_div
@@ -255,10 +235,8 @@ def _print_elementwise(up_t: dict[str, torch.Tensor], fv_t: dict[str, torch.Tens
continue
diff = (a - b).abs()
rel = (diff.mean().item() / max(a.abs().mean().item(), 1e-9)) * 100
print(
f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%"
)
print(f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%")
def main() -> None:
@@ -43,12 +43,10 @@ def _add_official_to_path() -> Path:
def _log_tensor_stats(label: str, tensor: torch.Tensor) -> None:
value = tensor.detach().float()
print(
f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}"
)
print(f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}")
def _extract_tensor(output: Any, key: str) -> torch.Tensor:
@@ -73,10 +71,8 @@ def _run_official_pipeline(
device: torch.device,
) -> Any:
del official_path, params, device
pytest.skip(
"TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target."
)
pytest.skip("TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target.")
def _run_fastvideo_pipeline(model_path: Path, params: dict[str, Any]) -> Any:
@@ -84,26 +80,36 @@ def _run_fastvideo_pipeline(model_path: Path, params: dict[str, Any]) -> Any:
generator = VideoGenerator.from_pretrained(
str(model_path),
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
{
"engine": {
"num_gpus": 1,
"use_fsdp_inference": False,
"offload": {
"dit": False,
"vae": False,
"text_encoder": False,
},
},
},
)
try:
return generator.generate_video(
prompt=params["prompt"],
negative_prompt=params.get("negative_prompt"),
output_path=f"outputs_{_MODEL_FAMILY}/pipeline_parity",
save_video=False,
height=params.get("height"),
width=params.get("width"),
num_frames=params.get("num_frames"),
fps=params.get("fps"),
num_inference_steps=params["num_inference_steps"],
guidance_scale=params.get("guidance_scale"),
seed=params["seed"],
)
return generator.generate({
"prompt": params["prompt"],
"negative_prompt": params.get("negative_prompt"),
"sampling": {
"height": params.get("height"),
"width": params.get("width"),
"num_frames": params.get("num_frames"),
"fps": params.get("fps"),
"num_inference_steps": params["num_inference_steps"],
"guidance_scale": params.get("guidance_scale"),
"seed": params["seed"],
},
"output": {
"output_path": f"outputs_{_MODEL_FAMILY}/pipeline_parity",
"save_video": False,
},
})
finally:
generator.shutdown()
@@ -146,8 +152,6 @@ def test_todo_model_family_pipeline_official_parity() -> None:
assert official_tensor.shape == fastvideo_tensor.shape
diff = (official_tensor - fastvideo_tensor).abs()
print(
f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}"
)
print(f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}")
assert_close(fastvideo_tensor, official_tensor, atol=1e-2, rtol=1e-2)
+99
View File
@@ -0,0 +1,99 @@
---
name: ci-runner
description: Work on FastVideo's Slurm-only, change-aware GPU CI lanes, static Buildkite graph, trusted ci-runner policy, lane scripts, and GB200 validation.
---
# Slinky Slurm CI lanes
FastVideo's `ci-runner` Buildkite queue is the control plane for all active
GPU CI. A private host-owned dispatcher leases GPUs from the Slinky Slurm tray
and runs the immutable PR SHA inside an isolated Enroot container. Buildkite
pipeline upload and Slurm submission occur on the login plane; every test
payload executes on Slurm compute.
The files under `fastvideo/tests/modal/` and `.buildkite/scripts/pr_test.sh`
are dormant rollback code. Never add an active Buildkite or slash-command
route to them. `pr_test.sh` must continue to reject Buildkite invocations.
The private operator bundle is deliberately outside this repository because
it contains site paths and credentials. See
`docs/contributing/ci_architecture.md`; this skill covers the repository half
and the coordination contract with that bundle.
## Invariants
- `.buildkite/pipeline.yml` contains exactly one static step for every active
GPU lane. Each step pins a unique key and label, a 90-minute timeout, the
trusted `/opt/fastvideo-ci-runner/run-ci` command (`run-unit` is the one
compatibility wrapper), step-level internal `TEST_TYPE`, and
`queue: "ci-runner"`.
- Active CI contains no `pr_test.sh` command, Modal invocation, default queue,
Buildkite plugin, `soft_fail`, or job-controlled artifact glob.
- The six Fastcheck lanes use `:microscope:` labels. Full-Suite-only lanes use
`:test_tube:` or `:bar_chart:` so direct reruns update the right aggregate.
- SSIM and vanilla training request all four GPUs. Keep both in the
`fastvideo/slinky/whole-tray` Buildkite concurrency group with a limit of one
so the second job does not consume an agent or command timeout while waiting
for the same tray.
- `/test full` schedules all twenty lanes. `/merge`, `ready`, and new pushes to
ready PRs use the trusted base-branch planner in
`.github/scripts/plan_merge_ci.py`: automatic Fastcheck remains the universal
six-lane baseline, and the merge build adds only path-relevant integration
lanes. Unknown source/build paths fail closed to all fourteen additive lanes.
The trusted uploader still normalizes and validates the complete static graph
before Buildkite evaluates its plan conditions.
- Focused merge builds may pass allowlisted golden-gate and SSIM test basenames.
The private host validates the lane plan and basenames before staging them,
and the in-container scripts validate them again. Direct `/test ssim`,
explicit `/test full`, and the weekly main-branch schedule run the complete
SSIM matrix.
- The trusted uploader serves exactly three entry pipelines:
`pr-fastcheck` for automatic PR builds, `ci` for slash-command/ready-label
API builds, and `fastvideo-performance-lane` for the weekly schedule. Keep
incoming GitHub webhook processing disabled on `ci` so it cannot duplicate
`pr-fastcheck` on every PR update.
- Test payloads live in `.buildkite/scripts/unit_test.sh` or executable
`.buildkite/scripts/lanes/<lane>.sh`. Backend policy (GPU count, extras,
secrets, kernel build, artifacts) stays in the agent-owned lane table.
- Tests must preserve an inherited `MASTER_PORT`. Packed containers share the
tray network namespace, so the private runner assigns a distinct port range
per GPU lease and the SSIM scheduler assigns task offsets within its range.
- The ARM64 runner image includes the pinned FA4 CuTe overlay validated on
GB200. Keep SSIM at `FASTVIDEO_FA4=1` because its references were seeded with
FA4; keep lanes with FA2 baselines at `FASTVIDEO_FA4=0`. A runner image change
must revalidate both the FA4 import and an actual GB200 forward kernel.
- `fastvideo/tests/ssim/ci_runner.py` is the active four-GPU SSIM scheduler.
New SSIM files are discovered through `REQUIRED_GPUS` and
`*_MODEL_TO_PARAMS`; do not wire them through the dormant Modal scheduler.
- The host policy fail-closes unknown tuples. A repository-side lane change is
inert until the operator updates the private lane table and uploader policy
in the same rollout.
## Adding or changing a lane
1. Read the closest `AGENTS.md` and the domain-specific testing guide.
2. Add or update the executable lane payload under `.buildkite/scripts/`.
Keep it deterministic and free of host-specific paths or credential fetches.
3. Add the static pipeline step and canonical `/test <name>` mapping. Keep the
`<name>-ci` alias only when compatibility requires it.
4. Add its source/test path ownership to `.github/scripts/plan_merge_ci.py`.
Prefer the narrowest correctness-preserving lane set; leave unknown paths
fail-closed. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
`test_merge_ci_plan.py`, and focused CPU-only scheduler/policy tests.
5. Coordinate the private lane row: GPU count (1-4), wall time, script, scope
pairs, step key, command, HF cache/token, tracking mode, extras, attention
backend policy, kernel policy, and artifact relay. Active training lanes
keep W&B offline and do not stage a W&B credential.
6. Update the trusted pipeline-uploader schema. A mismatch must reject the
pipeline rather than silently skip a lane.
7. Run `pre-commit run --files <changed paths>`, the planner's representative
diff matrix, contract tests, private driver tests, and a real GB200 canary.
Multi-GPU, hardware-reference, training, performance, and SSIM changes need
their own target-hardware evidence.
## Rollback
Rollback the Slurm routing/configuration change or pause the `ci-runner` queue.
Do not silently reactivate Modal. A manual Modal experiment requires the
explicit local opt-in documented in `ci_architecture.md`; returning it to
production CI needs a separate reviewed decision.
@@ -84,7 +84,9 @@ fastvideo/configs/models/dits/__init__.py
fastvideo/configs/models/encoders/__init__.py
fastvideo/configs/models/vaes/__init__.py
fastvideo/envs.py
fastvideo/fastvideo_args.py
fastvideo/api/schema.py
fastvideo/api/resolution.py
fastvideo/api/inference_resolution.py
fastvideo/distributed/**
fastvideo/layers/**
fastvideo/attention/**
@@ -0,0 +1,80 @@
---
name: env-var-conventions
description: Add, read, rename, or remove an environment variable in FastVideo, or change the environment-variable policy. Use before touching fastvideo/envs.py, os.environ, os.getenv, or monkeypatch.setenv in fastvideo/, and when fastvideo/tests/contract/test_env_policy.py fails.
---
# Environment Variable Conventions
## Purpose
FastVideo registers its environment variables as typed fields in
`fastvideo/envs.py`. The policy that governs them is
`docs/contributing/env_vars.md`, and the contract test
`fastvideo/tests/contract/test_env_policy.py` enforces the policy in the unit
CI lane. This skill routes an environment-variable change through that policy.
The policy doc is the single source of the rules; read it instead of relying
on a summary here.
## Prerequisites
- Read `docs/contributing/env_vars.md` in full.
- Decide whether the setting belongs in an environment variable or an argument
(rule 5 in the policy doc). Settings that users change per deployment are
arguments; add them as typed config fields in `fastvideo/api/schema.py`
instead.
## Inputs
| Parameter | Required | Description |
| ---------- | -------- | -------------------------------------------------------------- |
| `change` | Yes | Add, read, rename, or remove a variable, or change the policy. |
| `variable` | Yes | The variable name, with the `FASTVIDEO_` prefix. |
## Steps
1. **Declare or edit the variable in `fastvideo/envs.py`.**
- Pick the field type and category that the policy doc lists.
- Write a description that states what the variable does and its units.
- To rename, keep the old name in `deprecated_names`. To remove, add the
name to `DEPRECATED_VARIABLES`. Update the uses in `examples/`,
`scripts/`, `docs/`, `apps/`, and the tests.
2. **Read the variable with `envs.NAME.get()` inside a function.**
- In tests, change the value with `envs.NAME.override(value)`, and a variable
outside the registry with `envs.override_external(name, value)`; the
`env_overrides` fixture keeps either until the end of the test.
- Name a variable that only tests read `FASTVIDEO_TEST_*`.
- Do not call `os.environ`, `os.getenv`, or `monkeypatch.setenv` for a
FastVideo variable.
- To set a variable that another tool reads, call `envs.set_external`,
`envs.setdefault_external`, or `envs.unset_external`.
3. **Regenerate the table in the policy doc.**
- Run `python fastvideo/tests/contract/test_env_policy.py`.
4. **Run the contract test.**
- Run `pytest fastvideo/tests/contract/test_env_policy.py`.
- When the test reports a fixed known violation, delete or lower its entry
in `KNOWN_VIOLATIONS`. Never add an entry to `KNOWN_VIOLATIONS`.
5. **When the policy itself changes, update the policy doc and the contract
test in the same pull request.**
- The rules in `docs/contributing/env_vars.md`, the checks and allowlist in
`fastvideo/tests/contract/test_env_policy.py`, and this skill must agree.
## Outputs
- A registry entry in `fastvideo/envs.py` and call sites that use
`envs.NAME.get()`.
- A regenerated table in `docs/contributing/env_vars.md`.
- A passing `fastvideo/tests/contract/test_env_policy.py`.
## Example Usage
```
Add a FASTVIDEO_DEBUG_MY_STAGE switch that logs MyStage inputs.
```
## References
- `docs/contributing/env_vars.md`: the policy, the field types, and the
violation kinds that the contract test reports.
- `fastvideo/envs.py`: the registry.
- `fastvideo/tests/contract/test_env_policy.py`: the contract test,
`EXTERNAL_ALLOWLIST`, and `KNOWN_VIOLATIONS`.
@@ -1,6 +1,6 @@
---
name: reseed-ssim-references
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted model subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
@@ -13,7 +13,7 @@ on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
3. Regenerates through the manual legacy Modal L40S maintenance path.
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
@@ -51,12 +51,13 @@ harder to recover from than failing closed.
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Modal GPU: **L40S**. This is a manual reference-maintenance target, not the
active Slurm CI compute path; changing the SKU also changes the historical
`L40S_reference_videos` contract.
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
`FASTVIDEO_SSIM_REFERENCE_HF_REPO`).
`FASTVIDEO_TEST_SSIM_REFERENCE_HF_REPO`).
- Device folder: `L40S_reference_videos`.
## Prerequisites
+4 -2
View File
@@ -35,7 +35,8 @@ The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an artefact without crashing (the missing-reference assertion at
the end is expected). The skill does not re-test locally; it goes straight
to Modal L40S (which is what CI uses).
to the manual legacy Modal L40S reference-maintenance target. Active CI runs
on the Slinky Slurm cluster and only consumes the resulting references.
## When to use
@@ -61,7 +62,8 @@ Prompt the user for it if they didn't supply it.
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Modal maintenance GPU: **L40S** (hardcoded in
`fastvideo/tests/modal/ssim_test.py`; this is not the active CI compute path).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
seeded by this skill.
+461 -517
View File
@@ -1,7 +1,8 @@
env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
# Buildkite only launches Modal; remote jobs initialize their own submodules.
# Slurm workers clone the immutable commit and initialize submodules inside
# their isolated container. The Buildkite login-plane checkout is a no-op.
BUILDKITE_GIT_SUBMODULES: false
notify:
@@ -10,528 +11,471 @@ notify:
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
- github_commit_status:
context: "full-suite-passed"
if: build.env("TEST_SCOPE") == "full"
if: build.env("TEST_SCOPE") == "full" || build.env("TEST_SCOPE") == "merge"
- github_commit_status:
context: "direct-test-completed"
if: build.env("TEST_SCOPE") == "direct"
- github_commit_status:
context: "scheduled-ssim-passed"
if: build.env("TEST_SCOPE") == "scheduled"
# This is the complete active GPU CI surface. Every command is a trusted host
# dispatcher, and every test payload executes inside the Slinky Slurm tray.
# fastvideo/tests/modal remains available only for an explicit manual rollback;
# no active pipeline or slash-command route invokes it.
steps:
# ============================================================
# Direct test: triggered by /test <name> slash command.
# Labels match fastcheck/full-suite counterparts so the GitHub
# check status overwrites the original failed check.
# Only ONE step executes per build (gated by TEST_TYPE).
# ============================================================
- label: ":microscope: Encoder Tests"
key: "encoder"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,encoder,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "encoder" || build.env("TEST_TYPE") == "encoder_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "encoder_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Fastcheck-scope direct tests ---
- label: ":microscope: Encoder Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "encoder"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "vae"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "transformer"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "kernel_tests"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
key: "vae"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,vae,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "vae" || build.env("TEST_TYPE") == "vae_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "vae_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Full-suite-scope direct tests ---
- label: ":bar_chart: SSIM Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "ssim"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Inference Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Extraction Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "lora_extraction"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Distillation DMD Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "distillation_dmd"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Self-Forcing Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "self_forcing"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests VSA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_vsa"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Inference Tests VMoBA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_vmoba"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Performance Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "performance"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: API Server Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "api_server"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Train Framework Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "train_framework"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Eval Metrics Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "eval"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
key: "transformer"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,transformer,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "transformer" || build.env("TEST_TYPE") == "transformer_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "transformer_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Fastcheck: Runs on every PR (~10-15 min parallel)
# Core component validation: encoders, VAEs, transformers,
# CUDA kernels, and unit tests.
# ============================================================
- label: "Trigger Fastcheck"
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/models/encoders/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/encoders/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: Encoder Tests"
env:
- TEST_TYPE=encoder
agents:
queue: "default"
- path:
- "fastvideo/models/vaes/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/vaes/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: VAE Tests"
env:
- TEST_TYPE=vae
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/layers/**"
- "fastvideo/attention/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Transformer Tests"
env:
- TEST_TYPE=transformer
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Kernel Tests"
env:
- TEST_TYPE=kernel_tests
agents:
queue: "default"
- path:
- "fastvideo/**"
- ".buildkite/**"
- ".github/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Unit Tests"
env:
- TEST_TYPE=unit_test
agents:
queue: "default"
- path:
- "apps/dreamverse/**"
- "pyproject.toml"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":microscope: DreamVerse App Tests"
env:
- TEST_TYPE=dreamverse_app
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
key: "kernel-tests"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,kernel-tests,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "kernel_tests" || build.env("TEST_TYPE") == "kernel_tests_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "kernel_tests_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Full Suite: Runs when TEST_SCOPE=full
# Triggered by adding the 'ready' label (via ci-trigger-full-suite.yml)
# or on-demand via /test full slash command.
# Includes integration tests, SSIM regression, training pipelines,
# and performance benchmarks.
# ============================================================
- label: "Trigger Full Suite"
if: build.env("TEST_SCOPE") == "full"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/**/*.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":bar_chart: SSIM Tests"
env:
- TEST_TYPE=ssim
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/tests/lora/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/pipelines/**"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Inference Tests"
env:
- TEST_TYPE=inference_lora
agents:
queue: "default"
- path:
- "scripts/lora_extraction/**"
- "fastvideo/tests/lora_extraction/**"
- "fastvideo/models/loader/**"
- "fastvideo/training/training_utils.py"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Extraction Tests"
env:
- TEST_TYPE=lora_extraction
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
agents:
queue: "default"
- path:
- "fastvideo/training/*distillation_pipeline.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
agents:
queue: "default"
- path:
- "fastvideo/training/*self_forcing_distillation_pipeline.py"
- "fastvideo/tests/training/self-forcing/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Self-Forcing Tests"
env:
- TEST_TYPE=self_forcing
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Training Tests"
env:
- TEST_TYPE=training_lora
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests VSA"
env:
- TEST_TYPE=training_vsa
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "fastvideo/attention/backends/vmoba.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Inference Tests VMoBA"
env:
- TEST_TYPE=inference_vmoba
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/pipelines/**"
- "fastvideo/attention/**"
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/performance/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Performance Tests"
env:
- TEST_TYPE=performance
agents:
queue: "default"
- path:
- "fastvideo/entrypoints/openai/**"
- "fastvideo/entrypoints/cli/serve.py"
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: API Server Tests"
env:
- TEST_TYPE=api_server
agents:
queue: "default"
- path:
- "fastvideo/train/**"
- "fastvideo/tests/train/models/**"
- "fastvideo/tests/train/fixtures/**"
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Train Framework Tests"
env:
- TEST_TYPE=train_framework
agents:
queue: "default"
- path:
- "fastvideo/eval/**"
- "fastvideo/tests/eval/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Eval Metrics Tests"
env:
- TEST_TYPE=eval
agents:
queue: "default"
- label: ":microscope: Unit Tests"
key: "unit"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,unit,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "unit_test" || build.env("TEST_TYPE") == "unit_test_ci"))
command: "/opt/fastvideo-ci-runner/run-unit"
timeout_in_minutes: 90
env:
TEST_TYPE: "unit_test_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: DreamVerse App Tests"
key: "dreamverse"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,dreamverse,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "dreamverse_app" || build.env("TEST_TYPE") == "dreamverse_app_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "dreamverse_app_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Golden-Gate Tests"
key: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,golden-gate,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "golden_gate" || build.env("TEST_TYPE") == "golden_gate_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "golden_gate_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":bar_chart: SSIM Tests"
key: "ssim"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
build.env("TEST_SCOPE") == "scheduled" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,ssim,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "ssim" || build.env("TEST_TYPE") == "ssim_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "ssim_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Inference Tests"
key: "lora-inference"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-inference,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_lora" || build.env("TEST_TYPE") == "inference_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Extraction Tests"
key: "lora-extraction"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-extraction,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "lora_extraction" || build.env("TEST_TYPE") == "lora_extraction_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "lora_extraction_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests"
key: "training"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training" || build.env("TEST_TYPE") == "training_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "training_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Distillation DMD Tests"
key: "distillation"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,distillation,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "distillation_dmd" || build.env("TEST_TYPE") == "distillation_dmd_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "distillation_dmd_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Self-Forcing Tests"
key: "self-forcing"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,self-forcing,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "self_forcing" || build.env("TEST_TYPE") == "self_forcing_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "self_forcing_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Training Tests"
key: "lora-training"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_lora" || build.env("TEST_TYPE") == "training_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests VSA"
key: "training-vsa"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training-vsa,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_vsa" || build.env("TEST_TYPE") == "training_vsa_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_vsa_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Inference Tests VMoBA"
key: "inference-vmoba"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,inference-vmoba,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_vmoba" || build.env("TEST_TYPE") == "inference_vmoba_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_vmoba_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Performance Tests"
key: "performance"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,performance,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "performance" || build.env("TEST_TYPE") == "performance_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "performance_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: API Server Tests"
key: "api-server"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,api-server,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "api_server" || build.env("TEST_TYPE") == "api_server_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "api_server_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Train Framework Tests"
key: "train-framework"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,train-framework,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "train_framework" || build.env("TEST_TYPE") == "train_framework_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "train_framework_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Eval Metrics Tests"
key: "eval"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,eval,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "eval" || build.env("TEST_TYPE") == "eval_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "eval_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the OpenAI-compatible API lane.
set -euo pipefail
exec pytest ./fastvideo/tests/entrypoints/test_openai_api_integration.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the distillation-DMD lane.
set -euo pipefail
exec pytest ./fastvideo/tests/training/distill/test_distill_dmd.py -vs
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# DreamVerse needs a GPU for import-time device resolution, but it does not
# build or exercise fastvideo-kernel. A checksummed Node archive is installed
# in the disposable Slurm container because the shared CI image is
# Python/CUDA focused.
set -euo pipefail
node_version=v22.23.2
case $(uname -m) in
aarch64 | arm64)
node_arch=arm64
node_archive_sha256=013b59cfd2819703a6f4a14ab891fc46fc2a4e3f5bcd92de3fb4929b43e35b30
;;
x86_64 | amd64)
node_arch=x64
node_archive_sha256=b294a556e639d64338823920e5866c21c02741742d2e1529ee1a225c1ec9252a
;;
*)
echo "Unsupported architecture for DreamVerse Node runtime: $(uname -m)" >&2
exit 2
;;
esac
node_archive="node-${node_version}-linux-${node_arch}.tar.gz"
node_runtime_root=$(mktemp -d -t fastvideo-node.XXXXXX)
node_archive_path="${node_runtime_root}/${node_archive}"
node_install_dir="${node_runtime_root}/${node_archive%.tar.gz}"
curl --proto '=https' --tlsv1.2 --retry 5 --retry-all-errors \
--location --fail --silent --show-error \
"https://nodejs.org/dist/${node_version}/${node_archive}" \
--output "$node_archive_path"
printf '%s %s\n' "$node_archive_sha256" "$node_archive_path" | sha256sum --check --status
tar -xzf "$node_archive_path" -C "$node_runtime_root"
export PATH="${node_install_dir}/bin:${PATH}"
node --version
npm --version
export PYTHONPATH="$(pwd)/apps/dreamverse${PYTHONPATH:+:$PYTHONPATH}"
pytest apps/dreamverse/dreamverse/tests -q
cd apps/dreamverse/web
npm ci
npm run typecheck
npm test
machine_arch=$(uname -m)
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
npx playwright install --with-deps chromium firefox
else
npx playwright install --with-deps chromium webkit firefox
fi
master_port=${MASTER_PORT:-7959}
BACKEND_PORT=${BACKEND_PORT:-$((master_port + 50))}
python -m uvicorn dreamverse.mock_server:app --host 127.0.0.1 --port "$BACKEND_PORT" &
mock_server_pid=$!
cleanup() {
kill "$mock_server_pid" 2>/dev/null || true
wait "$mock_server_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
for _ in {1..30}; do
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz" && break
sleep 1
done
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz"
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
# Playwright WebKit traps before opening a page on Linux ARM64, and its
# bundled Chromium lacks the H.264/AAC codecs used by the fMP4 assertions.
# Firefox covers every flow, including streaming. Chromium and its mobile
# profile still cover all codec-independent UI behavior on GB200.
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- --project=firefox
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=mobile-chromium \
--grep-invert='streams, plays, and surfaces a downloadable clip|starts a new project and switches back to the prior session|saved projects persist across a page reload'
else
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=webkit \
--project=firefox \
--project=mobile-safari \
--project=mobile-chromium
fi
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the encoder lane.
set -euo pipefail
exec pytest ./fastvideo/tests/encoders -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the evaluation lane.
set -euo pipefail
exec pytest ./fastvideo/tests/eval -vs
+35
View File
@@ -0,0 +1,35 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the golden-gate lane. Environment (HF_HOME
# and authentication) is the runner's responsibility.
set -euo pipefail
golden_root=./fastvideo/tests/golden_gate
selected=${FASTVIDEO_GOLDEN_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_GOLDEN_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" = all ]; then
exec pytest "$golden_root" -xvs
fi
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_GOLDEN_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a golden_files <<< "$selected"
golden_paths=()
for golden_file in "${golden_files[@]}"; do
golden_path="$golden_root/$golden_file"
[ -f "$golden_path" ] || {
echo "Selected golden test does not exist: $golden_file" >&2
exit 2
}
golden_paths+=("$golden_path")
done
exec pytest "${golden_paths[@]}" -xvs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-inference lane.
set -euo pipefail
exec pytest ./fastvideo/tests/inference/lora/test_lora_inference_similarity.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VMoBA-inference lane.
set -euo pipefail
exec python fastvideo/tests/inference/vmoba/test_vmoba_inference.py
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the custom-kernel lane.
set -euo pipefail
exec pytest fastvideo-kernel/tests/ -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-extraction lane.
set -euo pipefail
exec pytest ./fastvideo/tests/lora_extraction/test_lora_extraction.py -vs
+52
View File
@@ -0,0 +1,52 @@
#!/usr/bin/env bash
# Canonical Slurm performance lane. Reports are written outside the checkout
# so the trusted host driver can upload them after untrusted code exits.
set -uo pipefail
export PERFORMANCE_TRACKING_ROOT=/tmp/perf-tracking
export PERF_REPORTS_DIR=/workspace/artifacts/performance
mkdir -p "$PERF_REPORTS_DIR"
if [[ ${BUILDKITE_PULL_REQUEST:-false} =~ ^[1-9][0-9]*$ ]]; then
export PERF_RUN_SOURCE=pr
export PERF_UPLOAD_POLICY=pass
elif [ "${BUILDKITE_BRANCH:-}" = main ] \
&& { [ "${BUILDKITE_SOURCE:-}" = schedule ] || [ "${TEST_SCOPE:-}" = full ]; }; then
export PERF_RUN_SOURCE=scheduled_main
export PERF_UPLOAD_POLICY=always
elif [ "${TEST_SCOPE:-}" = direct ]; then
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=pass
else
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=never
fi
nvidia-smi \
--query-gpu=index,timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu \
--format=csv -l 10 > "$PERF_REPORTS_DIR/gpu_telemetry.csv" 2>/dev/null &
telemetry_pid=$!
cleanup() {
kill "$telemetry_pid" 2>/dev/null || true
wait "$telemetry_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
pytest ./fastvideo/tests/performance -vs
pytest_rc=$?
compare_rc=0
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
PERF_PYTEST_RC=$pytest_rc python ./fastvideo/tests/performance/compare_baseline.py
compare_rc=$?
fi
python ./fastvideo/tests/performance/dashboard.py || true
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
final_rc=$pytest_rc
if [ "$final_rc" -eq 0 ]; then
final_rc=$compare_rc
fi
exit "$final_rc"
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the self-forcing lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/self-forcing/test_self_forcing.py -vs
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# Canonical four-GPU SSIM lane for the Slinky Slurm worker.
set -euo pipefail
args=()
if [ "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-0}" = 1 ]; then
args+=(--bootstrap-mode)
fi
selected=${FASTVIDEO_SSIM_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_SSIM_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" != all ]; then
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_SSIM_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a ssim_files <<< "$selected"
for ssim_file in "${ssim_files[@]}"; do
args+=(--test-file "$ssim_file")
done
fi
# MoGe's utils3d dependency builds glcontext from source on ARM64. The current
# runner image predates the baked-in X11 headers below, so keep this guarded
# bootstrap until every deployed image digest contains libx11-dev.
if [ ! -f /usr/include/X11/Xlib.h ]; then
apt-get -o Acquire::Retries=5 update
apt-get -o Acquire::Retries=5 install -y --no-install-recommends libx11-dev
rm -rf /var/lib/apt/lists/*
fi
uv pip install git+https://github.com/microsoft/MoGe.git
uv pip install k_diffusion einops_exts alias_free_torch torchsde
exec python fastvideo/tests/ssim/ci_runner.py "${args[@]}"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the modular training-framework lane.
set -euo pipefail
exec pytest ./fastvideo/tests/train/models ./fastvideo/tests/train/methods -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy vanilla-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/Vanilla -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy LoRA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/lora/test_lora_training.py -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy VSA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/VSA -srP
+9
View File
@@ -0,0 +1,9 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the transformer lane.
set -euo pipefail
# The existing block reference records an absent FASTVIDEO_FA4 (FA2). Keep
# that reference identity; the component lane also selects FA2 explicitly.
env -u FASTVIDEO_FA4 pytest ./fastvideo/tests/golden_gate/test_wan_t2v.py -xvs
pytest ./fastvideo/tests/golden_gate/test_wan_causal.py -xvs
exec pytest ./fastvideo/tests/transformers -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VAE lane.
set -euo pipefail
pytest ./fastvideo/tests/golden_gate/test_wan_vae.py -xvs
exec pytest ./fastvideo/tests/vaes -vs
+17
View File
@@ -1,6 +1,19 @@
#!/bin/bash
set -uo pipefail
# DORMANT ROLLBACK ONLY. Active CI is Slurm-only and pipeline.yml never calls
# this launcher. Refuse every Buildkite invocation even if a stale step or
# operator typo reaches this file; local rollback experiments require an
# explicit opt-in.
if [ -n "${BUILDKITE:-}" ]; then
echo "Legacy Modal CI is disabled; use the Slinky Slurm runner." >&2
exit 2
fi
if [ "${FASTVIDEO_ENABLE_LEGACY_MODAL_CI:-0}" != 1 ]; then
echo "Legacy Modal CI is dormant. Set FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1 only for a manual rollback test." >&2
exit 2
fi
log() {
echo "[$(date '+%Y-%m-%d %H:%M:%S')] $1"
}
@@ -187,6 +200,10 @@ case "$TEST_TYPE" in
log "Running transformer tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_transformer_tests"
;;
"golden_gate")
log "Running golden-gate tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_golden_gate_tests"
;;
"ssim")
log "Running SSIM tests..."
SSIM_BOOTSTRAP_ARGS=$(ssim_bootstrap_args)
+26
View File
@@ -0,0 +1,26 @@
#!/usr/bin/env bash
set -euo pipefail
exec pytest \
./fastvideo/tests/api/ \
./fastvideo/tests/contract/ \
./fastvideo/tests/dataset/ \
./fastvideo/tests/workflow/ \
./fastvideo/tests/entrypoints/ \
./fastvideo/tests/loader/ \
./fastvideo/tests/pipelines/ \
./fastvideo/tests/platforms/ \
./fastvideo/tests/train/ \
./fastvideo/tests/stages/ \
./fastvideo/tests/ops/ \
./fastvideo/tests/worker/ \
./fastvideo/tests/training/test_trackers.py \
./fastvideo/tests/attention/test_sdpa_metadata_mask_contract.py \
./fastvideo/tests/attention/test_vsa_h3_tile_grad_safety.py \
./fastvideo/tests/modal/test_kernel_build_cache.py \
./fastvideo/tests/modal/test_pr_test.py \
./fastvideo/tests/modal/test_ssim_test.py \
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
--ignore=./fastvideo/tests/train/models \
--ignore=./fastvideo/tests/train/methods \
-vs
+2 -2
View File
@@ -8,10 +8,10 @@ PR TITLE: Must start with a type tag, e.g.:
MERGE WORKFLOW:
1. Ensure pre-commit passes and you have at least 1 approval
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
3. A path-aware merge gate runs only relevant integration tests → auto-merge on success
ON-DEMAND TESTING (write access required):
/test full — Full Test Suite /test ssim — SSIM regression
/test full — Explicit all-lane run /test ssim — Full SSIM regression
/test training — Training pipeline /test encoder — Encoder tests
/test transformer — Transformer tests /test vae — VAE tests
/test kernel — CUDA kernel tests /test unit — Unit tests
+10 -10
View File
@@ -1,14 +1,14 @@
#!/usr/bin/env bash
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
# Gate the path-aware Buildkite merge plan on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
# already doomed.
# 'ready' label cannot burn path-selected GPU lanes on a head that a cheap
# check has already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no full suite; the next push re-arms via the 'synchronize' trigger)
# no merge gate; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
@@ -29,7 +29,7 @@ set -euo pipefail
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the full suite
# Workflow-level `name:` values that must be green before the merge gate
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
@@ -56,7 +56,7 @@ recheck_ready_label() {
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
"NOT triggering the Buildkite merge gate. Re-add the label to re-arm."
exit 1
fi
else
@@ -84,7 +84,7 @@ while true; do
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
"NOT triggering the Buildkite merge gate. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
@@ -97,7 +97,7 @@ while true; do
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — full suite may proceed."
echo "All watched cheap checks are green — merge gate may proceed."
exit 0
fi
case "$missing" in
@@ -119,14 +119,14 @@ while true; do
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the merge gate WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the merge gate anyway."
exit 0
fi
sleep "$POLL_SECS"
+581
View File
@@ -0,0 +1,581 @@
#!/usr/bin/env python3
"""Select the additive GPU integration lanes needed by a PR diff.
Fastcheck is the universal six-lane baseline and is intentionally not repeated
here. This planner selects only the more expensive merge-gate lanes. Unknown
source/build paths fail closed to the complete integration set, while explicit
documentation and repository-metadata paths require no additional GPU work.
"""
from __future__ import annotations
import argparse
import fnmatch
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import TextIO
MERGE_LANES = (
"golden-gate",
"ssim",
"lora-inference",
"lora-extraction",
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
"inference-vmoba",
"performance",
"api-server",
"train-framework",
"eval",
)
LANE_SCRIPT_TO_KEY = {
"api_server.sh": "api-server",
"distillation_dmd.sh": "distillation",
"eval.sh": "eval",
"golden_gate.sh": "golden-gate",
"inference_lora.sh": "lora-inference",
"inference_vmoba.sh": "inference-vmoba",
"lora_extraction.sh": "lora-extraction",
"performance.sh": "performance",
"self_forcing.sh": "self-forcing",
"ssim.sh": "ssim",
"train_framework.sh": "train-framework",
"training.sh": "training",
"training_lora.sh": "lora-training",
"training_vsa.sh": "training-vsa",
}
FASTCHECK_LANE_SCRIPTS = {
"dreamverse.sh",
"encoder.sh",
"kernel_tests.sh",
"transformer.sh",
"vae.sh",
}
LEGACY_TRAINING_LANES = (
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
)
ALL_TRAINING_LANES = (*LEGACY_TRAINING_LANES, "train-framework")
SSIM_SMOKE_TESTS = (
"test_flux_t2i_similarity.py",
"test_wan_t2v_similarity.py",
)
SAFE_PATTERNS = (
"*.md",
"*.rst",
".agents/**",
".claude/**",
".codex/**",
".github/ISSUE_TEMPLATE/**",
".github/PULL_REQUEST_TEMPLATE.md",
".github/dependabot.yml",
".github/mergify.yml",
".github/scripts/**",
".github/workflows/**",
".buildkite/scripts/pre_commit.sh",
".git-blame-ignore-revs",
".gitattributes",
".gitignore",
".pre-commit-config.yaml",
"AGENTS.md",
"CITATION.cff",
"CODE_OF_CONDUCT.md",
"CONTRIBUTING.md",
"LICENSE",
"NOTICE",
"__init__.py",
"collect_env.py",
"SECURITY.md",
"assets/**",
"comfyui/**",
"docs/**",
"examples/**",
"mkdocs.yml",
"requirements-mkdocs.in",
"requirements-mkdocs.txt",
"scripts/**",
"tests/__init__.py",
"tests/local_tests/**",
)
ALL_IMPACT_PATTERNS = (
".buildkite/pipeline.yml",
"docker/**",
"pyproject.toml",
"requirements*.txt",
"setup.cfg",
"setup.py",
"uv.lock",
)
@dataclass(frozen=True)
class FamilyCoverage:
pattern: re.Pattern[str]
golden_tests: tuple[str, ...]
ssim_tests: tuple[str, ...]
FAMILY_COVERAGE = (
FamilyCoverage(
re.compile(r"(^|[/_.-])dreamx(_world)?([/_.-]|$)"),
("test_dreamx.py", ),
("test_dreamx_world_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux[_-]?2([/_.-]|$)"),
("test_flux2_klein.py", ),
("test_flux2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux(?![_-]?2)([/_.-]|$)"),
("test_flux.py", ),
("test_flux_t2i_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])(hunyuan)?gamecraft([/_.-]|$)"),
("test_gamecraft.py", ),
("test_gamecraft_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])gen3c([/_.-]|$)"),
("test_gen3c.py", ),
("test_gen3c_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])glm[_-]?image([/_.-]|$)"),
("test_glm_image.py", ),
("test_glm_image_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
("test_kandinsky5.py", ),
("test_kandinsky5_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])lingbot([a-z0-9_-]*)([/_.-]|$)"),
("test_lingbot.py", ),
("test_lingbot_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])longcat([/_.-]|$)"),
("test_longcat.py", ),
("test_longcat_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])ltx[_-]?2([/_.-]|$)"),
("test_ltx2.py", ),
("test_ltx2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?2([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?3([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])minimax[_-]?h3([/_.-]|$)"),
("test_minimax_h3_t2v.py", ),
("test_minimax_h3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])sd[_-]?3([._-]?5)?([/_.-]|$)"),
("test_sd35.py", ),
("test_sd35_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])stable[_-]?audio([/_.-]|$)"),
("test_stable_audio.py", ),
("test_stable_audio_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])turbo(diffusion)?([/_.-]|$)"),
(),
("test_turbodiffusion_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)"),
("test_wan_t2v.py", "test_wan_vae.py", "test_wan_causal.py", "test_wan_denoising.py"),
(
"test_causal_similarity.py",
"test_wan_i2v_similarity.py",
"test_wan_t2v_similarity.py",
),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])z[_-]?image([/_.-]|$)"),
("test_zimage.py", ),
("test_zimage_similarity.py", ),
),
)
@dataclass
class MergePlan:
lanes: set[str] = field(default_factory=set)
golden_tests: set[str] = field(default_factory=set)
ssim_tests: set[str] = field(default_factory=set)
golden_all: bool = False
ssim_all: bool = False
reasons: list[str] = field(default_factory=list)
def add_lanes(self, *lanes: str, reason: str) -> None:
unknown = set(lanes) - set(MERGE_LANES)
if unknown:
raise ValueError(f"Unknown merge lanes: {sorted(unknown)}")
self.lanes.update(lanes)
self.reasons.append(reason)
def add_golden(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("golden-gate", reason=reason)
self.golden_tests.update(tests)
def add_ssim(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("ssim", reason=reason)
self.ssim_tests.update(tests)
def require_all(self, reason: str) -> None:
self.lanes.update(MERGE_LANES)
self.golden_all = True
self.ssim_all = True
self.reasons.append(reason)
def ordered_lanes(self) -> tuple[str, ...]:
return tuple(lane for lane in MERGE_LANES if lane in self.lanes)
def encoded_lanes(self) -> str:
lanes = self.ordered_lanes()
return "," + ",".join(lanes or ("none", )) + ","
def encoded_golden_tests(self) -> str:
if "golden-gate" not in self.lanes:
return "none"
if self.golden_all or not self.golden_tests:
return "all"
return ",".join(sorted(self.golden_tests))
def encoded_ssim_tests(self) -> str:
if "ssim" not in self.lanes:
return "none"
if self.ssim_all or not self.ssim_tests:
return "all"
return ",".join(sorted(self.ssim_tests))
def _matches_any(path: str, patterns: tuple[str, ...]) -> bool:
return any(fnmatch.fnmatchcase(path, pattern) for pattern in patterns)
def _family_coverage(path: str) -> tuple[set[str], set[str]]:
normalized = path.lower()
golden: set[str] = set()
ssim: set[str] = set()
for family in FAMILY_COVERAGE:
if family.pattern.search(normalized):
golden.update(family.golden_tests)
ssim.update(family.ssim_tests)
# Select the component actually touched, including compatibility paths.
# Family configs/pipeline wiring can affect all four Wan gates.
if re.search(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)", normalized):
if (normalized.endswith(("/wan/vae.py", "/wan/vae_config.py", "/vaes/wanvae.py"))
or normalized.endswith("/wan/stages/conditioning.py")):
golden = {"test_wan_vae.py"}
elif normalized.endswith(("/wan/causal_transformer.py", "/dits/causal_wanvideo.py",
"/wan/stages/causal_denoising.py")):
golden = {"test_wan_causal.py"}
elif (normalized == "fastvideo/models/dits/wanvideo.py"
or normalized.endswith(("/wan/transformer.py", "/wan/stages/denoising.py", "/wan/stages/dmd.py"))):
golden = {"test_wan_t2v.py", "test_wan_denoising.py"}
return golden, ssim
def _select_output_coverage(plan: MergePlan, path: str) -> None:
golden, ssim = _family_coverage(path)
if golden:
plan.add_golden(tuple(sorted(golden)), reason=f"model-family golden coverage: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared output golden coverage: {path}")
if ssim:
plan.add_ssim(tuple(sorted(ssim)), reason=f"model-family SSIM coverage: {path}")
else:
plan.add_ssim(SSIM_SMOKE_TESTS, reason=f"shared output SSIM smoke coverage: {path}")
def classify_paths(paths: list[str]) -> MergePlan:
plan = MergePlan()
normalized_paths: list[str] = []
for raw_path in paths:
path = raw_path.strip()
while path.startswith("./"):
path = path[2:]
if path:
normalized_paths.append(path)
normalized_paths = sorted(set(normalized_paths))
if not normalized_paths:
plan.require_all("changed-file list was empty; failing closed")
return plan
for path in normalized_paths:
if path == "__FASTVIDEO_CI_PLAN_ALL__":
plan.require_all("changed-file API failed; failing closed")
continue
if path in {"requirements-mkdocs.in", "requirements-mkdocs.txt"}:
plan.reasons.append(f"documentation dependencies need no GPU integration: {path}")
continue
if _matches_any(path, ALL_IMPACT_PATTERNS):
plan.require_all(f"cross-cutting build/runtime surface: {path}")
continue
lane_script_prefix = ".buildkite/scripts/lanes/"
if path.startswith(lane_script_prefix):
script_name = Path(path).name
lane = LANE_SCRIPT_TO_KEY.get(script_name)
if lane is None:
if script_name in FASTCHECK_LANE_SCRIPTS:
plan.reasons.append(f"covered by automatic Fastcheck lane: {path}")
else:
plan.require_all(f"unknown lane script: {path}")
elif lane == "golden-gate":
plan.golden_all = True
plan.add_lanes(lane, reason=f"golden lane implementation: {path}")
elif lane == "ssim":
plan.ssim_all = True
plan.add_lanes(lane, reason=f"SSIM lane implementation: {path}")
else:
plan.add_lanes(lane, reason=f"lane implementation: {path}")
continue
if path.startswith("fastvideo/tests/golden_gate/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_golden((name, ), reason=f"changed golden test: {path}")
elif name in {"AGENTS.md", "README.md"}:
plan.reasons.append(f"golden documentation only: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared golden harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/ssim/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_ssim((name, ), reason=f"changed SSIM test: {path}")
elif path.endswith((".py", ".json", ".pt", ".png", ".mp4")):
plan.ssim_all = True
plan.add_lanes("ssim", reason=f"shared SSIM harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/performance/") or path.startswith(".buildkite/performance-benchmarks/"):
plan.add_lanes("performance", reason=f"performance coverage: {path}")
continue
if path.startswith(("fastvideo/performance/", "fastvideo/performance_dashboard/",
"apps/performance_dashboard/")):
plan.add_lanes("performance", reason=f"performance implementation: {path}")
continue
if path.startswith("fastvideo/benchmarks/"):
if "/mlx_" in path or Path(path).name.startswith("mlx_"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
else:
plan.add_lanes("performance", reason=f"benchmark implementation: {path}")
continue
if path.startswith("fastvideo/tests/eval/") or path.startswith("fastvideo/eval/"):
plan.add_lanes("eval", reason=f"evaluation coverage: {path}")
continue
if path.startswith("fastvideo/third_party/eval/"):
plan.add_lanes("eval", reason=f"vendored evaluation implementation: {path}")
continue
if path.startswith("fastvideo/tests/lora_extraction/") or path.startswith("scripts/lora_extraction/"):
plan.add_lanes("lora-extraction", reason=f"LoRA extraction coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/lora/"):
plan.add_lanes("lora-inference", reason=f"LoRA inference coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/vmoba/"):
plan.add_lanes("inference-vmoba", reason=f"VMoBA inference coverage: {path}")
continue
if path.startswith(("fastvideo/dataset/", "fastvideo/workflow/", "fastvideo/pipelines/preprocess/",
"fastvideo/pipelines/training/")):
plan.add_lanes(*ALL_TRAINING_LANES, reason=f"shared data/training input surface: {path}")
continue
if path.startswith("fastvideo/tests/train/") or path.startswith("fastvideo/train/"):
plan.add_lanes("train-framework", reason=f"modular training coverage: {path}")
continue
if path.startswith("fastvideo/tests/training/"):
lowered = path.lower()
if "/vanilla/" in lowered:
plan.add_lanes("training", reason=f"vanilla training coverage: {path}")
elif "/distill/" in lowered:
plan.add_lanes("distillation", reason=f"distillation coverage: {path}")
elif "/self-forcing/" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing coverage: {path}")
elif "/lora/" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training coverage: {path}")
elif "/vsa/" in lowered:
plan.add_lanes("training-vsa", reason=f"VSA training coverage: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training coverage: {path}")
continue
if path.startswith("fastvideo/training/"):
lowered = path.lower()
if "self_forcing" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing implementation: {path}")
elif "distill" in lowered:
plan.add_lanes("distillation", reason=f"distillation implementation: {path}")
elif "lora" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training implementation: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training implementation: {path}")
continue
lowered = path.lower()
if "vmoba" in lowered and path.startswith(("fastvideo/", ".buildkite/")):
plan.add_lanes("inference-vmoba", reason=f"VMoBA implementation: {path}")
plan.add_golden(("test_wan_t2v.py", ), reason=f"VMoBA end-to-end coverage: {path}")
continue
if "lora" in lowered and path.startswith("fastvideo/"):
plan.add_lanes(
"lora-inference",
"lora-extraction",
"lora-training",
reason=f"shared LoRA implementation: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/entrypoints/") or path.startswith("fastvideo/api/"):
plan.add_lanes("api-server", reason=f"API/entrypoint integration: {path}")
if "openai" not in lowered and "/cli/" not in lowered:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/worker/"):
plan.add_lanes("api-server", reason=f"worker/API integration: {path}")
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/distributed/"):
plan.add_lanes(
"training",
"train-framework",
reason=f"distributed runtime integration: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/hooks/", "fastvideo/platforms/", "fastvideo/third_party/")):
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/models/", "fastvideo/pipelines/", "fastvideo/configs/",
"fastvideo/layers/", "fastvideo/attention/")):
_select_output_coverage(plan, path)
continue
if path in {
"fastvideo/forward_context.py",
"fastvideo/image_processor.py",
"fastvideo/registry.py",
"fastvideo/utils.py",
}:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/mlx_runtime/"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
continue
if path.startswith("fastvideo/logging_utils/") or path in {
"fastvideo/__init__.py",
"fastvideo/envs.py",
"fastvideo/logger.py",
"fastvideo/profiler.py",
"fastvideo/version.py",
}:
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path.startswith(("fastvideo-kernel/", "csrc/")):
plan.add_golden(("test_wan_t2v.py", ), reason=f"kernel integration smoke: {path}")
plan.add_ssim(("test_wan_t2v_similarity.py", ), reason=f"kernel numerical smoke: {path}")
continue
if path.startswith("apps/dreamverse/"):
# DreamVerse is already one of the six automatic Fastcheck lanes.
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
continue
if path.startswith("fastvideo/tests/"):
# The automatic unit/component Fastcheck lanes own the remaining
# package tests. Domain-specific expensive test roots were handled
# above.
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path in {".buildkite/scripts/unit_test.sh", ".buildkite/scripts/pr_test.sh"}:
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if _matches_any(path, SAFE_PATTERNS):
plan.reasons.append(f"no additional GPU integration needed: {path}")
continue
plan.require_all(f"unclassified path; failing closed: {path}")
return plan
def _write_github_output(output: TextIO, plan: MergePlan) -> None:
output.write(f"merge_test_plan={plan.encoded_lanes()}\n")
output.write(f"merge_golden_tests={plan.encoded_golden_tests()}\n")
output.write(f"merge_ssim_tests={plan.encoded_ssim_tests()}\n")
output.write(f"merge_plan_label={','.join(plan.ordered_lanes()) or 'none'}\n")
def _write_summary(output: TextIO, plan: MergePlan) -> None:
output.write("## Change-aware merge test plan\n\n")
output.write("| Selection | Value |\n|---|---|\n")
output.write(f"| Additional Slurm lanes | `{','.join(plan.ordered_lanes()) or 'none'}` |\n")
output.write(f"| Golden tests | `{plan.encoded_golden_tests()}` |\n")
output.write(f"| SSIM tests | `{plan.encoded_ssim_tests()}` |\n\n")
output.write("Fastcheck remains the universal six-lane baseline.\n")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--paths-file", type=Path, required=True)
parser.add_argument("--github-output", type=Path)
parser.add_argument("--summary-file", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
paths = args.paths_file.read_text(encoding="utf-8").splitlines()
plan = classify_paths(paths)
print(f"MERGE_TEST_PLAN={plan.encoded_lanes()}")
print(f"MERGE_GOLDEN_TESTS={plan.encoded_golden_tests()}")
print(f"MERGE_SSIM_TESTS={plan.encoded_ssim_tests()}")
for reason in plan.reasons:
print(f"- {reason}")
if args.github_output:
with args.github_output.open("a", encoding="utf-8") as output:
_write_github_output(output, plan)
if args.summary_file:
with args.summary_file.open("a", encoding="utf-8") as output:
_write_summary(output, plan)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -1
View File
@@ -53,7 +53,7 @@ PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
OTHER='{"name": "Trigger Merge Gate", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
@@ -190,6 +190,7 @@ jobs:
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "Digest: ${{ steps.build-push.outputs.digest }}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
+39 -25
View File
@@ -26,29 +26,48 @@ jobs:
per_page: 100,
});
const bkStatuses = data.statuses.filter(
s => s.context.startsWith('buildkite/ci/')
);
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
// Buildkite derives the GitHub context prefix from the label emoji.
// Keep hard Full Suite lanes in test-tube/bar-chart namespaces and
// Fastcheck lanes in microscope so targeted reruns cannot clear the
// wrong aggregate status. Automatic PR jobs use pr-fastcheck while
// slash-command and Full Suite jobs use ci; normalize the suffix
// and keep the newest status for each logical lane.
const FASTCHECK_PREFIXES = [
'buildkite/pr-fastcheck/microscope-',
'buildkite/ci/microscope-',
];
const FULL_SUITE_PREFIXES = [
'buildkite/ci/test-tube-',
'buildkite/ci/bar-chart-',
];
const fastcheck = bkStatuses.filter(
s => s.context.startsWith(FASTCHECK_PREFIX)
);
const fullSuite = bkStatuses.filter(
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
);
function newestByLane(prefixes) {
const statuses = new Map();
for (const status of data.statuses) {
const prefix = prefixes.find(p => status.context.startsWith(p));
if (!prefix) continue;
const lane = status.context.slice(prefix.length);
const previous = statuses.get(lane);
if (!previous || Date.parse(status.updated_at) > Date.parse(previous.updated_at)) {
statuses.set(lane, status);
}
}
return statuses;
}
if (
fastcheck.length > 0
&& fastcheck.every(s => s.state === 'success')
) {
const fastcheck = newestByLane(FASTCHECK_PREFIXES);
const fullSuiteOnly = newestByLane(FULL_SUITE_PREFIXES);
const fastcheckPassed =
fastcheck.size === 6
&& [...fastcheck.values()].every(s => s.state === 'success');
const fullSuitePassed =
fastcheckPassed
&& fullSuiteOnly.size === 14
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
if (fastcheckPassed) {
core.info(
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -56,17 +75,13 @@ jobs:
sha,
state: 'success',
context: 'fastcheck-passed',
description:
`All ${fastcheck.length} fastcheck tests passed`,
description: `All ${fastcheck.size} fastcheck tests passed`,
});
}
if (
fullSuite.length > 0
&& fullSuite.every(s => s.state === 'success')
) {
if (fullSuitePassed) {
core.info(
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
'All 20 full suite tests passed — updating full-suite-passed'
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -74,7 +89,6 @@ jobs:
sha,
state: 'success',
context: 'full-suite-passed',
description:
`All ${fullSuite.length} full suite tests passed`,
description: 'All 20 full suite tests passed',
});
}
+169
View File
@@ -0,0 +1,169 @@
name: macOS MLX Smoke
on:
pull_request:
branches: [main]
paths:
- ".github/workflows/ci-macos-mlx.yml"
- "fastvideo/mlx_runtime/**"
- "fastvideo/tests/mlx/**"
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
- "fastvideo/tests/platforms/test_cpu_sdpa.py"
- "fastvideo/platforms/cpu.py"
- "fastvideo/platforms/mps.py"
- "fastvideo/platforms/__init__.py"
- "fastvideo/__init__.py"
- "examples/inference/basic/mlx_*.py"
- "fastvideo/benchmarks/mlx_*.py"
- "pyproject.toml"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: macos-mlx-${{ github.ref }}
cancel-in-progress: true
jobs:
mlx-smoke:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: macos-15
timeout-minutes: 25
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru mlx \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Show Apple runtime
run: |
python - <<'PY'
import platform
import mlx.core as mx
import torch
print("machine:", platform.machine())
print("processor:", platform.processor())
print("mlx default device:", mx.default_device())
memory_size = mx.metal.device_info().get("memory_size") if mx.metal.is_available() else "metal unavailable"
print("mlx memory_size:", memory_size)
print("torch:", torch.__version__)
print("torch mps available:", torch.backends.mps.is_available())
PY
- name: Run MLX smoke tests
run: |
python -m pytest \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
# graph (the parity tests were designed to be backend-agnostic), while the
# macOS job above stays the source of truth for Metal behavior.
mlx-smoke-linux-cpu:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: ubuntu-latest
timeout-minutes: 20
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies (CPU backend)
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru "mlx[cpu]" \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Run MLX smoke tests (CPU backend)
run: |
python -m pytest \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
+44
View File
@@ -0,0 +1,44 @@
name: Scheduled Full SSIM
on:
schedule:
- cron: "0 5 * * 0"
workflow_dispatch:
permissions:
contents: read
jobs:
trigger:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- name: Trigger weekly full SSIM on Slinky Slurm
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
SOURCE_SHA: ${{ github.sha }}
SOURCE_BRANCH: ${{ github.event.repository.default_branch }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
set -euo pipefail
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$SOURCE_SHA" \
--arg branch "$SOURCE_BRANCH" \
'{
commit: $commit,
branch: $branch,
message: "Weekly full SSIM on Slinky Slurm",
ignore_pipeline_branch_filters: true,
env: {
TEST_SCOPE: "scheduled",
FULL_SUITE: "false",
TEST_TYPE: "ssim",
PR_NUMBER: "false",
PR_TITLE: "Scheduled full SSIM"
}
}')"
+13 -45
View File
@@ -33,7 +33,6 @@ jobs:
core.setOutput('has_write', String(hasWrite));
- name: Add ready label and react
id: label
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
@@ -48,47 +47,6 @@ jobs:
comment_id: context.payload.comment.id,
content: 'rocket',
});
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
core.setOutput('pr_title', pr.title);
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
PR_TITLE: ${{ steps.label.outputs.pr_title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
parse-command:
if: >-
@@ -129,7 +87,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval unit-ci kernel-ci dreamverse-ci ssim-ci golden-gate-ci encoder-ci vae-ci transformer-ci lora-inference-ci lora-training-ci lora-extraction-ci training-ci distillation-ci self-forcing-ci vsa-ci vmoba-ci performance-ci api-ci train-framework-ci eval-ci full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -137,8 +95,18 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [training]=training
[kernel]=kernel_tests [unit]=unit_test [unit-ci]=unit_test_ci
[kernel-ci]=kernel_tests_ci [dreamverse-ci]=dreamverse_app_ci
[ssim-ci]=ssim_ci [vmoba-ci]=inference_vmoba_ci
[golden-gate-ci]=golden_gate_ci [training-ci]=training_ci
[encoder-ci]=encoder_ci [vae-ci]=vae_ci [transformer-ci]=transformer_ci
[lora-inference-ci]=inference_lora_ci [lora-training-ci]=training_lora_ci
[lora-extraction-ci]=lora_extraction_ci [distillation-ci]=distillation_dmd_ci
[self-forcing-ci]=self_forcing_ci [vsa-ci]=training_vsa_ci
[performance-ci]=performance_ci [api-ci]=api_server_ci
[train-framework-ci]=train_framework_ci [eval-ci]=eval_ci
[dreamverse]=dreamverse_app
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
[distillation]=distillation_dmd [self-forcing]=self_forcing
+66 -13
View File
@@ -1,4 +1,4 @@
name: Trigger Full Suite
name: Trigger Merge Gate
on:
pull_request_target:
@@ -10,7 +10,7 @@ permissions:
actions: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
group: merge-gate-${{ github.event.pull_request.number }}
cancel-in-progress: false
jobs:
@@ -34,29 +34,72 @@ jobs:
});
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
core.setOutput('changed_files', String(pr.changed_files));
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
- name: Cancel previous Buildkite builds
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
# Find running builds for this branch with TEST_SCOPE=full and cancel them
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
# Match both branch and PR number: forks can reuse the same branch name.
builds=$(curl -sS --get -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state=running,scheduled" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds" \
| jq -r --arg pr_number "$PR_NUMBER" \
'.[] | select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number)) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
# Checks out the BASE branch (default for pull_request_target), so PR
# authors cannot tamper with the gate script.
- name: Checkout gate script
# Check out the immutable BASE SHA: pull_request_target must never run a
# planner or gate script from the untrusted PR head.
- name: Checkout trusted merge planner
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
ref: ${{ github.event.pull_request.base.sha }}
persist-credentials: false
- name: Collect changed paths
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ github.event.pull_request.number }}
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
run: |
set -euo pipefail
changed_json="$RUNNER_TEMP/merge-changed-files.json"
changed_paths="$RUNNER_TEMP/merge-changed-paths.txt"
if gh api --paginate --slurp \
"repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/files?per_page=100" \
> "$changed_json"; then
observed=$(jq '[.[][] | .filename] | unique | length' "$changed_json")
if [ "$observed" = "$EXPECTED_CHANGED_FILES" ]; then
jq -r '.[][] | .filename, (.previous_filename // empty)' "$changed_json" \
| sort -u > "$changed_paths"
else
echo "::warning::Changed-file API returned $observed of $EXPECTED_CHANGED_FILES paths; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
else
echo "::warning::Changed-file API failed; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
- name: Select minimal merge tests
id: plan
if: steps.check.outputs.has_ready == 'true'
run: |
python3 .github/scripts/plan_merge_ci.py \
--paths-file "$RUNNER_TEMP/merge-changed-paths.txt" \
--github-output "$GITHUB_OUTPUT" \
--summary-file "$GITHUB_STEP_SUMMARY"
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
@@ -66,7 +109,7 @@ jobs:
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite Full Suite
- name: Trigger Buildkite merge gate
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
@@ -76,6 +119,10 @@ jobs:
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
MERGE_GOLDEN_TESTS: ${{ steps.plan.outputs.merge_golden_tests }}
MERGE_SSIM_TESTS: ${{ steps.plan.outputs.merge_ssim_tests }}
MERGE_PLAN_LABEL: ${{ steps.plan.outputs.merge_plan_label }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
@@ -84,8 +131,11 @@ jobs:
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg message "Merge gate [${MERGE_PLAN_LABEL}] for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--arg merge_test_plan "$MERGE_TEST_PLAN" \
--arg merge_golden_tests "$MERGE_GOLDEN_TESTS" \
--arg merge_ssim_tests "$MERGE_SSIM_TESTS" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -95,8 +145,11 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
TEST_SCOPE: "merge",
FULL_SUITE: "true",
MERGE_TEST_PLAN: $merge_test_plan,
MERGE_GOLDEN_TESTS: $merge_golden_tests,
MERGE_SSIM_TESTS: $merge_ssim_tests,
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
+4 -4
View File
@@ -38,17 +38,17 @@ jobs:
**How our CI works:**
PRs run a two-tier CI system:
PRs run a three-tier CI system:
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
2. **Fastcheck** — six core GPU lanes run automatically via Buildkite (~10-15 min).
3. **Merge gate** — a reviewer adds `ready`; changed paths select only the relevant integration, training, golden, or SSIM coverage.
**Before your PR is reviewed:**
- [ ] `pre-commit run --all-files` passes locally
- [ ] You've added or updated tests for your changes
- [ ] The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
**Useful links:**
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
+41 -7
View File
@@ -13,16 +13,28 @@ on:
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when their Dockerfile changes on main. The CUDA
# matrix is the only lane that builds from docker/Dockerfile, so a path-scoped
# push trigger is a sufficient change detector on its own -- no separate
# detect-changes/paths-filter job is needed now that there is a single
# in-scope Dockerfile. Dreamverse (apps/dreamverse/docker/Dockerfile) and the
# rocm Dockerfile stay manual-dispatch only.
build_ci_runner_image:
description: 'Build the ARM64 CUDA 13 CI runner image (sm_100)'
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
# Dreamverse (apps/dreamverse/docker/Dockerfile) and the ROCm Dockerfile stay
# manual-dispatch only.
push:
branches: [main]
paths:
- '.dockerignore'
- '.github/workflows/_template-build-image.yml'
- '.github/workflows/infra-build-image.yml'
- '.gitmodules'
- 'docker/Dockerfile'
- 'docker/uv-excludes'
- 'fastvideo-kernel/**'
- 'fastvideo/tests/modal/kernel_build_cache.py'
- 'pyproject.toml'
permissions:
@@ -50,7 +62,7 @@ jobs:
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
# Runs on a manual dispatch when build_cuda_matrix is set, or automatically
# on a push that changed docker/Dockerfile (inputs are null on push). The
# on an in-scope main push (inputs are null on push). The
# repository guard keeps fork syncs from auto-building; manual dispatch
# still works in forks.
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_cuda_matrix == 'true' }}
@@ -191,6 +203,28 @@ jobs:
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# The CI runner is ARM64 like DGX Spark, but targets sm_100 rather than sm_121.
# Publish a single-architecture variant so the self-hosted CI runner can reuse
# the exact prebuilt kernel instead of compiling it in every job.
build-ci-runner-image:
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_ci_runner_image == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12-cuda13.0.0-sm100
runner: ubuntu-24.04-arm
architecture: arm64
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=13.0.0
UV_TORCH_BACKEND=cu130
TORCH_CUDA_ARCH_LIST=10.0
CMAKE_BUILD_PARALLEL_LEVEL=1
FLASH_ATTN_WHEEL_TAG=cu130torch2.12
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
+2
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
@@ -16,6 +17,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
+15 -8
View File
@@ -62,8 +62,9 @@ jobs:
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
# x86_64 builds the full cu126 + cu130 set. cu130 ships the
# data-center Blackwell sm_100a/sm_103a VSA and consumer sm_120a FP4
# kernels.
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
@@ -124,7 +125,7 @@ jobs:
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo apt install -y git gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
@@ -168,17 +169,18 @@ jobs:
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
# + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="9.0a;10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
@@ -194,7 +196,11 @@ jobs:
python -m build --wheel --outdir dist
# Fix the wheel to be manylinux compliant
uv pip install --system auditwheel
# Ubuntu 22.04 ships patchelf 0.14.3, while current auditwheel
# requires at least 0.14.5. Use the stable PyPI binary on both
# x86_64 and aarch64 release runners.
uv pip install --system auditwheel patchelf==0.17.2.4
patchelf --version
# Point auditwheel at torch libs, but do not vendor them into the wheel.
TORCH_LIB_DIR=$(python - <<'PY'
import os
@@ -211,7 +217,8 @@ jobs:
--exclude libtorch.so \
--exclude libc10.so \
--exclude libc10_cuda.so \
--exclude libtorch_python.so
--exclude libtorch_python.so \
--exclude libnccl.so.2
# Move fixed wheels back to dist for upload consistency
rm dist/*.whl
mv fixed_dist/*.whl dist/
+8 -1
View File
@@ -6,7 +6,7 @@ results/
wandb/
*.ipynb
*.jpg
!examples/dataset/lingbotworld2/image.jpg
!examples/datasets/lingbotworld2/image.jpg
*.safetensors
*.mp4
*.png
@@ -23,6 +23,7 @@ Miniconda3-latest-Linux-x86_64.sh
*validation/
data/
outputs/
outputs_audio/
outputs_video
checkpoints/
sbatch.sh
@@ -54,6 +55,8 @@ eggs/
# MkDocs documentation
site/
docs/assets/cookbook-serving.json
examples/serving/clients/node_modules/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
@@ -75,6 +78,7 @@ docs/distillation/examples/
*.pkl
# Reference videos (negations must come after the catch-all on line below)
!fastvideo/tests/nightly/reference_video_*.mp4
# Static images
!docs/assets/images/**/*.png
@@ -131,6 +135,9 @@ fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
.kernel_bench/
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
+2 -1
View File
@@ -9,7 +9,7 @@ exclude: |
tests/.*|
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
fastvideo/models/(?!wan/(config|vae_config|pipeline_config|definition|__init__)\.py$).*|
^apps/dreamverse/web/.*|
examples/.*|
\.agents/.*|
@@ -22,6 +22,7 @@ repos:
hooks:
- id: yapf
args: [--in-place, --verbose]
language_version: python3.12
additional_dependencies: [toml] # TODO: Remove when yapf is upgraded
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.11.12
+4
View File
@@ -66,14 +66,18 @@ Local guidance lives next to the code. Read the in-scope file before editing:
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
| `fastvideo/configs/AGENTS.md` | Arch + pipeline config dataclasses, `param_names_mapping` |
| `fastvideo/models/AGENTS.md` | DiT / VAE / encoder / scheduler / loader layout (pre-commit excluded) |
| `fastvideo/models/wan/AGENTS.md` | Wan family-local transformers, VAE, configs, and the SP sharding invariant |
| `fastvideo/layers/AGENTS.md` | Tensor-parallel linear/attention layer rules for ports |
| `fastvideo/attention/AGENTS.md` | Backend registry + env-var override |
| `fastvideo/pipelines/AGENTS.md` | Stage ABC, `basic/<model>/`, `preprocess/`, presets |
| `fastvideo/pipelines/basic/wan/AGENTS.md` | Wan sampling stages, first-frame conditioning, DMD/causal boundaries |
| `fastvideo/pipelines/basic/magi_human/AGENTS.md` | MagiHuman umbrella repo, lazy-loaded components, packing invariants |
| `fastvideo/training/AGENTS.md` | Legacy monolithic pipelines (frozen for existing models) |
| `fastvideo/train/AGENTS.md` | New modular trainer (methods × models × callbacks, YAML) |
| `fastvideo/tests/AGENTS.md` | Test taxonomy, conftest, pre-commit-excluded path |
| `fastvideo/tests/ssim/AGENTS.md` | GPU SSIM regression authoring + reference video sync |
| `scripts/checkpoint_conversion/AGENTS.md` | Adding a converter for a new HF/official checkpoint |
| `apps/dreamverse/AGENTS.md` | DreamVerse app structure and conventions |
## Critical: Two Training Stacks Coexist
+21 -10
View File
@@ -3,12 +3,16 @@
</div>
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://haoailab.com/FastVideo/cookbook/"><b>Cookbook</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
</p>
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/09/15`: Release [FastH3 8-Step V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2), an eight-forward data-free DMD2 checkpoint distilled from MiniMax-H3 with 80% Video Sparse Attention. Run it with `examples/inference/basic/basic_fasth3_8step.py` or the [FastH3 8-Step V2 recipe](https://haoailab.com/FastVideo/cookbook/minimax-h3/).
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac. Follow the [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
@@ -33,7 +37,7 @@ FastVideo has the following features:
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for the supported training workflows, and the [support matrix](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for supported models.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
@@ -60,7 +64,12 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
```
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
[MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
> **On an Apple Silicon Mac?** Install with `uv pip install -e '.[mlx]'` from
> a clone, then pick a recipe in the
> [cookbook](https://haoailab.com/FastVideo/cookbook/). See the
> [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
@@ -78,7 +87,7 @@ Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virt
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
- Apple Silicon, macOS -> docs/getting_started/installation/mlx.md
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
4. Verify the result:
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
@@ -126,18 +135,20 @@ def main():
# Create a video generator with a pre-trained model
generator = VideoGenerator.from_pretrained(
"FastVideo/FastWan2.1-T2V-1.3B-Diffusers",
num_gpus=1, # Adjust based on your hardware
{"engine": {"num_gpus": 1}}, # Adjust based on your hardware
)
# Define a prompt for your video
prompt = "A curious raccoon peers through a vibrant field of yellow sunflowers, its eyes wide with interest."
# Generate the video
video = generator.generate_video(
prompt,
output_path="my_videos/", # Controls where videos are saved
save_video=True
)
video = generator.generate({
"prompt": prompt,
"output": {
"output_path": "my_videos/", # Controls where videos are saved
"save_video": True,
},
})
if __name__ == '__main__':
main()
+23 -2
View File
@@ -97,13 +97,33 @@ dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
### Run Dreamverse with FastH3
Select the VSA data-free FastH3 Preview profile when you start the backend:
```bash
DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
The `fast-h3` profile uses four visible GPUs by default. It loads the `MiniMaxAI/MiniMax-H3` base checkpoint and the
`vsa-datafree/adapter_model.safetensors` adapter from
`FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA`. Each request generates a 124-frame, 768×1344 video with
synchronized audio and five sigma-grid points. Dreamverse uses the last frame of each segment as first-frame
conditioning for the following segment.
Set `CUDA_VISIBLE_DEVICES` when you need to choose the four physical GPUs:
```bash
CUDA_VISIBLE_DEVICES=0,1,2,3 DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
> (the default), the backend compiles the segment 1 and segment 2 inference
> paths before it reports ready — this can take **tens of minutes on a cold
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
> warmup finishes. For a faster, uncompiled startup while testing, set
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
> warmup finishes. To defer compilation until the first generated request while
> testing, set `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
## Frontend Setup
@@ -219,6 +239,7 @@ selection, and mock-server behavior:
pytest apps/dreamverse/dreamverse/tests/test_config.py \
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
apps/dreamverse/dreamverse/tests/test_minimax_h3_generation.py \
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
```
+13 -2
View File
@@ -42,7 +42,7 @@ Near-term OSS note:
- `apps/dreamverse/dreamverse/main.py`: websocket endpoint, request handling,
session state machine, rewrite orchestration, REST routes, and stream relay
- `apps/dreamverse/dreamverse/gpu_pool.py`: GPU worker processes, warmup, model
loading, and `generate_video()` calls through FastVideo
loading, and `generate()` calls through FastVideo
- `apps/dreamverse/dreamverse/prompt_enhancer.py`: prompt enhancement, rollout
rewrite execution, provider selection, and timeout/fallback behavior
- `apps/dreamverse/dreamverse/rewrite_prompt_payload.py`: canonical rewrite request payload
@@ -139,7 +139,18 @@ session.
- startup warmup
- user join/leave commands
- `USER_STEP` execution for each segment
- continuation state between segments
- generation-command routing and stream-result delivery
Model generation has a separate ownership boundary inside each GPU process:
- `apps/dreamverse/dreamverse/generation_worker.py` selects the backend that the active model profile declares and owns
the backend lifecycle.
- `apps/dreamverse/dreamverse/ltx2_generation.py` owns LTX-2 generator configuration, video and audio continuation, and
runtime LoRA application.
- `apps/dreamverse/dreamverse/minimax_h3_generation.py` owns the VSA data-free FastH3 adapter, FastH3 generator and
request configuration, and last-frame continuation through MiniMax H3 first-frame conditioning.
- `apps/dreamverse/dreamverse/generation_contracts.py` defines the decoded media and stream-trimming result that both
model backends return to `apps/dreamverse/dreamverse/gpu_pool.py`.
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
@@ -1,6 +1,6 @@
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
Mirrors how ``apps/dreamverse/dreamverse/video_generation.py`` constructs
Mirrors how ``apps/dreamverse/dreamverse/ltx2_generation.py`` constructs
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
@@ -52,7 +52,8 @@ import torch # noqa: E402
from fastvideo import VideoGenerator # noqa: E402
from fastvideo.api import ( # noqa: E402
ComponentConfig, CompileConfig, EngineConfig, GeneratorConfig, OffloadConfig, PipelineSelection, QuantizationConfig,
ComponentConfig, CompileConfig, EngineConfig, GenerationResult, GeneratorConfig, OffloadConfig, PipelineSelection,
QuantizationConfig,
)
DEFAULT_PROMPT = ("A cinematic drone shot over coastal cliffs at sunrise, golden "
@@ -128,9 +129,9 @@ def _build_generator_config(model_path: str, enable_compile: bool, num_gpus: int
)
def _extract_stage_times(result: dict) -> OrderedDict[str, float]:
def _extract_stage_times(result: GenerationResult) -> OrderedDict[str, float]:
out: OrderedDict[str, float] = OrderedDict()
info = result.get("logging_info") if isinstance(result, dict) else None
info = result.logging_info if isinstance(result, GenerationResult) else None
if info is None:
return out
stages = getattr(info, "stages", None)
@@ -162,19 +163,25 @@ def _do_one_run(generator: VideoGenerator, prompt: str, *, height: int, width: i
_reset_peak_gpu()
t0 = time.perf_counter()
try:
result = generator.generate_video(
prompt=prompt,
negative_prompt="",
save_video=False,
height=height,
width=width,
num_frames=num_frames,
fps=24,
num_inference_steps=num_inference_steps,
guidance_scale=1.0,
seed=seed,
ltx2_image_crf=0.0,
)
result = generator.generate({
"prompt": prompt,
"negative_prompt": "",
"sampling": {
"height": height,
"width": width,
"num_frames": num_frames,
"fps": 24,
"num_inference_steps": num_inference_steps,
"guidance_scale": 1.0,
"seed": seed,
},
"output": {
"save_video": False
},
"extensions": {
"ltx2_image_crf": 0.0
},
})
if torch.cuda.is_available():
torch.cuda.synchronize()
except Exception as exc:
+21 -2
View File
@@ -1,5 +1,6 @@
import os
from pathlib import Path
from typing import cast
_REPO_ROOT = Path(__file__).resolve().parents[1]
_SERVER_ROOT = Path(__file__).resolve().parent
@@ -55,16 +56,34 @@ FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
MODEL_REGISTRY = {
"fast-ltx2": {
"name": "FastLTX2",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
},
"fast-ltx23": {
"name": "FastLTX23",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
},
"fast-h3": {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
@@ -171,7 +190,7 @@ def _optional_env(*names: str) -> str | None:
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", cast(int, MODEL_CONFIG["default_sp_size"])))
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
if DREAMVERSE_MODEL_PATH:
@@ -213,7 +232,7 @@ def _resolve_lora_spec(spec: str) -> str | None:
if not spec:
return None
if spec.lower() == "omninft":
return MODEL_CONFIG.get("lora_repo")
return cast(str | None, MODEL_CONFIG.get("lora_repo"))
if spec.lower() in AVAILABLE_LORAS:
return AVAILABLE_LORAS[spec.lower()]["repo"]
return spec
@@ -0,0 +1,46 @@
"""Shared contract between DreamVerse generation backends and GPU workers."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Protocol
@dataclass
class StepResult:
"""Decoded media and stream-trimming metadata for one DreamVerse segment."""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict[str, float]
head_trim_frames: int
head_trim_audio_frames: int
class GenerationBackend(Protocol):
"""Model-owned generation operations used by one GPU worker process."""
def initialize(self, model_config: dict | None = None) -> None:
...
def shutdown(self) -> None:
...
def clear_conditioning(self) -> None:
...
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> StepResult:
...
def warmup(self, prompt: str) -> dict[str, float]:
...
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
...
@@ -0,0 +1,96 @@
"""Select and own one model-specific generation backend per GPU process."""
from __future__ import annotations
from dreamverse.config import MODEL_CONFIG
from dreamverse.generation_contracts import GenerationBackend, StepResult
def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBackend:
"""Construct the backend that owns the selected model family's behavior."""
if backend_name == "ltx2":
from dreamverse.ltx2_generation import LTX2GenerationBackend
return LTX2GenerationBackend(gpu_id)
if backend_name == "minimax_h3":
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
return MiniMaxH3GenerationBackend(gpu_id)
raise ValueError(f"Unsupported DreamVerse generation backend: {backend_name!r}")
class VideoGenerationWorker:
"""Delegate GPU lifecycle and generation calls to the active model backend."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.model_config: dict = dict(MODEL_CONFIG)
self.backend_name: str | None = None
self.backend: GenerationBackend | None = None
def initialize(self, model_config: dict | None = None) -> None:
"""Load the requested model through its generation backend.
Model selection belongs here so the GPU process and streaming layers
use one stable media contract without importing model-specific code.
"""
requested_model_config = dict(model_config) if model_config is not None else dict(self.model_config)
backend_name = requested_model_config.get("generation_backend")
if not isinstance(backend_name, str) or not backend_name:
raise ValueError("DreamVerse model configuration requires `generation_backend`.")
candidate_backend = self.backend
if candidate_backend is None or self.backend_name != backend_name:
if candidate_backend is not None:
candidate_backend.shutdown()
candidate_backend = _create_generation_backend(backend_name, self.gpu_id)
try:
candidate_backend.initialize(requested_model_config)
except Exception:
try:
candidate_backend.shutdown()
except Exception as shutdown_error:
print(f"[GPU {self.gpu_id}] Backend cleanup after initialization failure: {shutdown_error}")
self.backend = None
self.backend_name = None
raise
self.model_config = requested_model_config
self.backend = candidate_backend
self.backend_name = backend_name
def _require_backend(self) -> GenerationBackend:
"""Return the initialized backend or fail before processing a command."""
if self.backend is None:
raise RuntimeError("Generation backend is not initialized.")
return self.backend
def shutdown(self) -> None:
"""Release model resources owned by the selected backend."""
if self.backend is not None:
self.backend.shutdown()
def clear_conditioning(self) -> None:
self._require_backend().clear_conditioning()
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> StepResult:
"""Generate one segment through the selected model backend."""
return self._require_backend().generate_step(
prompt,
segment_idx,
image_path,
reset_conditioning,
)
def warmup(self, prompt: str) -> dict[str, float]:
return self._require_backend().warmup(prompt)
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
return self._require_backend().apply_lora_stack(stack)
+17 -10
View File
@@ -12,7 +12,7 @@ from enum import Enum
from multiprocessing import Process, Queue
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
DREAMVERSE_SP_SIZE,
MODEL_REGISTRY,
STARTUP_WARMUP_ENABLED,
@@ -54,7 +54,7 @@ from dreamverse.worker_ipc import (
def _parse_requested_gpu_limit() -> int | None:
raw_value = os.getenv("FASTVIDEO_GPU_COUNT", "").strip().lower()
if not raw_value:
return 1
return DREAMVERSE_SP_SIZE
if raw_value == "all":
return None
try:
@@ -164,12 +164,12 @@ def gpu_worker_process(
os.environ["CUDA_VISIBLE_DEVICES"] = cuda_device
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
from dreamverse.video_generation import VideoGenerationWorker
from dreamverse.generation_worker import VideoGenerationWorker
worker = VideoGenerationWorker(gpu_id)
def event_loop(first_cmd: Command = None):
"""Blocking event loop for LTX2; dispatches user commands."""
"""Block on generation commands after the model is initialized."""
print(f"[GPU {gpu_id}] Entering event loop")
def handle_command(cmd: Command):
@@ -435,7 +435,7 @@ class GPUSlot:
self._response_reader_task: asyncio.Task | None = None
self._active: bool = False
self._reader_lock: asyncio.Lock | None = None
self.current_model_id: str = DEFAULT_MODEL_ID
self.current_model_id: str | None = ACTIVE_MODEL_ID
self.shared_stream_buffer = None
self.shared_stream_buffer_size = SHARED_STREAM_BUFFER_BYTES
@@ -690,7 +690,7 @@ class GPUSlot:
async def join_user(self, user_id: str, model_id: str = None) -> JoinAck:
"""Add a user to this GPU."""
if model_id is None:
model_id = DEFAULT_MODEL_ID
model_id = ACTIVE_MODEL_ID
# Reload model if a different one is requested
if model_id != self.current_model_id and model_id in MODEL_REGISTRY:
@@ -705,16 +705,23 @@ class GPUSlot:
self.connected_users.clear()
model_config = MODEL_REGISTRY[model_id]
reload_response = await self._send_command(Command(CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
try:
reload_response = await self._send_command(Command(
CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
except Exception:
self.current_model_id = None
raise
match reload_response:
case ReloadAck():
pass
case WorkerError(message=msg):
self.current_model_id = None
raise RuntimeError(f"Model reload failed: {msg}")
case _:
self.current_model_id = None
raise RuntimeError(f"Unexpected reload response: "
f"{type(reload_response).__name__}")
@@ -1,9 +1,9 @@
"""LTX2 model lifecycle and continuation conditioning.
"""LTX-2 model lifecycle and continuation conditioning.
Runs inside a GPU worker subprocess. Owns the model, the audio
encoder, and the per-session continuation state carried across
segments. Callers must set ``os.environ["CUDA_VISIBLE_DEVICES"]``
before constructing ``VideoGenerationWorker`` — all ``fastvideo.*``
before constructing ``LTX2GenerationBackend`` — all ``fastvideo.*``
imports are deferred to method bodies so nothing touches CUDA at
module import time.
"""
@@ -14,9 +14,7 @@ import gc
import os
import re
import time
from dataclasses import dataclass
from typing import Any
from typing import TYPE_CHECKING
import numpy as np
import torch
@@ -35,6 +33,10 @@ from dreamverse.config import (
DREAMVERSE_LORA_STACK,
_resolve_lora_spec,
)
from dreamverse.generation_contracts import StepResult
if TYPE_CHECKING:
from fastvideo.api import GenerationResult
# Multi-frame decoded continuation defaults from
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
@@ -80,22 +82,6 @@ def _reset_lora_registry(worker) -> dict:
return {"status": "lora_registry_reset"}
@dataclass
class StepResult:
"""Output of one generation step.
``head_trim_frames`` / ``head_trim_audio_frames`` are derived here
so downstream AV streaming never needs to import conditioning
constants.
"""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict
head_trim_frames: int
head_trim_audio_frames: int
class ContinuationState:
"""Per-session video + audio conditioning carried across segments."""
@@ -113,8 +99,8 @@ class ContinuationState:
self.video_images = None
self.audio_latents = None
def apply_video(self, request_kwargs: dict, segment_idx: int) -> None:
"""Seed next-segment kwargs with the cached tail frames."""
def apply_video(self, request: dict, segment_idx: int) -> None:
"""Seed the next-segment request with the cached tail frames."""
if segment_idx <= 1 or not self.video_images:
return
from PIL import Image
@@ -128,21 +114,21 @@ class ContinuationState:
arr = np.clip(arr, 0, 255).astype(np.uint8)
noisy.append(Image.fromarray(arr))
cond_images = noisy
request_kwargs["ltx2_video_conditions"] = [(
request["extensions"]["ltx2_video_conditions"] = [(
cond_images,
LTX2_VIDEO_CONDITIONING_FRAME_IDX,
LTX2_VIDEO_CONDITIONING_STRENGTH,
)]
request_kwargs["ltx2_images"] = None
request_kwargs["image_path"] = None
request["extensions"]["ltx2_images"] = None
request["inputs"]["image_path"] = None
def apply_audio(
self,
request_kwargs: dict,
request: dict,
segment_idx: int,
audio_lps: float,
) -> None:
"""Seed next-segment kwargs with clean audio latents + denoise mask.
"""Seed the next-segment request with clean audio latents + denoise mask.
When audio conditioning is longer than video, extend audio
generation and shift video RoPE forward so the audio prefix
@@ -159,9 +145,9 @@ class ContinuationState:
audio_extra = max(0, AUDIO_CONDITIONING_NUM_FRAMES - LTX2_VIDEO_CONDITIONING_NUM_FRAMES)
if audio_extra > 0:
audio_num_frames = NUM_FRAMES + audio_extra
request_kwargs["audio_num_frames"] = (audio_num_frames)
request["extensions"]["audio_num_frames"] = (audio_num_frames)
prefix_sec = float(audio_extra) / 24.0
request_kwargs["video_position_offset_sec"] = prefix_sec
request["extensions"]["video_position_offset_sec"] = prefix_sec
new_duration = float(NUM_FRAMES + audio_extra) / 24.0
total_T = max(
@@ -179,8 +165,8 @@ class ContinuationState:
mask = torch.ones((B, 1, total_T, 1), dtype=torch.float32)
mask[:, :, :audio_cond_T, :] = (1.0 - AUDIO_CONDITIONING_STRENGTH)
request_kwargs["ltx2_audio_clean_latent"] = clean
request_kwargs["ltx2_audio_denoise_mask"] = mask
request["extensions"]["ltx2_audio_clean_latent"] = clean
request["extensions"]["ltx2_audio_denoise_mask"] = mask
def save_video(self, frames: list) -> None:
"""Snapshot trailing N frames as PIL images for next-segment conditioning."""
@@ -202,7 +188,7 @@ class ContinuationState:
self.audio_latents = latents.detach().clone().cpu()
class VideoGenerationWorker:
class LTX2GenerationBackend:
"""Single-GPU LTX2 generator with continuation state.
Caller must set ``os.environ["CUDA_VISIBLE_DEVICES"]`` before
@@ -324,7 +310,7 @@ class VideoGenerationWorker:
),
)
self.generator = VideoGenerator.from_pretrained(config=generator_config)
self.generator = VideoGenerator.from_config(generator_config)
print(f"[GPU {self.gpu_id}] After model load: {self._gpu_mem()}")
lora_stack = DREAMVERSE_LORA_STACK or ([(DREAMVERSE_LORA_PATH,
@@ -421,7 +407,7 @@ class VideoGenerationWorker:
return
loader = ComponentLoader.for_module_type("audio_encoder", "diffusers")
enc = loader.load(audio_vae_path, self.generator.fastvideo_args)
enc = loader.load(audio_vae_path, self.generator.resolved_config)
target = getattr(enc, "model", enc)
proc = AudioProcessor(
@@ -478,51 +464,59 @@ class VideoGenerationWorker:
prompt = self._inject_style_trigger(prompt)
request_kwargs = dict(
prompt=prompt,
negative_prompt="",
save_video=False,
height=FRAME_HEIGHT,
width=FRAME_WIDTH,
num_frames=NUM_FRAMES,
fps=24,
num_inference_steps=NUM_INFERENCE_STEPS,
guidance_scale=1.0,
seed=10,
ltx2_image_crf=0.0,
image_path=image_path if segment_idx == 1 else None,
return_continuation_state=False,
)
request = {
"prompt": prompt,
"negative_prompt": "",
"inputs": {
"image_path": image_path if segment_idx == 1 else None
},
"sampling": {
"height": FRAME_HEIGHT,
"width": FRAME_WIDTH,
"num_frames": NUM_FRAMES,
"fps": 24,
"num_inference_steps": NUM_INFERENCE_STEPS,
"guidance_scale": 1.0,
"seed": 10,
},
"output": {
"save_video": False
},
"extensions": {
"ltx2_image_crf": 0.0,
"return_continuation_state": False,
},
}
if reset_conditioning:
self.continuation.clear()
audio_lps = (DEFAULT_LTX2_AUDIO_SAMPLE_RATE / DEFAULT_LTX2_AUDIO_HOP_LENGTH / DEFAULT_LTX2_AUDIO_DOWNSAMPLE)
# Phase 1: seed kwargs with prior-segment conditioning.
self.continuation.apply_video(request_kwargs, segment_idx)
self.continuation.apply_audio(request_kwargs, segment_idx, audio_lps)
# Phase 1: seed the request with prior-segment conditioning.
self.continuation.apply_video(request, segment_idx)
self.continuation.apply_audio(request, segment_idx, audio_lps)
# Phase 2: generate.
t0 = time.perf_counter()
result = self.generator.generate_video(**request_kwargs)
result = self.generator.generate(request)
torch.cuda.synchronize()
timings["generation_ms"] = (time.perf_counter() - t0) * 1000
if not isinstance(result, dict):
raise RuntimeError("Expected dictionary output from generate_video.")
frames = result.get("frames")
if isinstance(result, list):
raise RuntimeError("Expected a single GenerationResult from generate.")
frames = result.frames
if not isinstance(frames, list) or len(frames) == 0:
raise RuntimeError("Generation did not return frames.")
audio = result.get("audio")
audio_sample_rate = result.get("audio_sample_rate")
audio = result.audio
audio_sample_rate = result.audio_sample_rate
if audio is not None and audio_sample_rate is None:
# LTX2 audio decoding stage uses 24kHz output by default.
audio_sample_rate = 24000
print(f"[GPU {self.gpu_id}] audio_sample_rate missing from result; "
f"defaulting to {audio_sample_rate}Hz")
timings["generation_time_ms"] = result.get("generation_time", 0.0) * 1000
timings["generation_time_ms"] = (result.generation_time or 0.0) * 1000
# Phase 3: snapshot continuation state for the next segment.
t_save_start = time.perf_counter()
@@ -560,7 +554,7 @@ class VideoGenerationWorker:
self,
audio: object,
audio_sample_rate: int | None,
result: dict,
result: "GenerationResult",
segment_idx: int,
) -> torch.Tensor | None:
"""Pick which tensor to cache for next-segment audio conditioning."""
@@ -575,7 +569,7 @@ class VideoGenerationWorker:
f"for segment {segment_idx + 1}")
return re_encoded
return None
audio_latents = result.get("ltx2_audio_latents")
audio_latents = result.extra.get("ltx2_audio_latents")
if audio_latents is not None:
print(f"[GPU {self.gpu_id}] Cached audio latents "
f"shape={tuple(audio_latents.shape)} "
@@ -0,0 +1,294 @@
"""FastH3 model lifecycle and first-frame continuation for DreamVerse."""
from __future__ import annotations
import gc
import os
import time
from typing import TYPE_CHECKING, Any
import numpy as np
import torch
from dreamverse.config import DREAMVERSE_SP_SIZE
from dreamverse.generation_contracts import StepResult
if TYPE_CHECKING:
from PIL.Image import Image
def _required_config_str(model_config: dict, field_name: str) -> str:
"""Read one required non-empty string from a DreamVerse model profile."""
value = model_config.get(field_name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"FastH3 model configuration requires `{field_name}`.")
return value.strip()
class MiniMaxH3GenerationBackend:
"""Run the VSA data-free FastH3 adapter and retain one continuation frame."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
reserved_gib = torch.cuda.memory_reserved() / 1024**3
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
@staticmethod
def _configure_environment(attention_backend: str) -> None:
"""Apply the fixed boot-time switches from the FastH3 reference recipe."""
os.environ.update({
"FASTVIDEO_ATTENTION_BACKEND": attention_backend,
"FASTVIDEO_FA4": "1",
"FASTVIDEO_MINIMAX_H3_FUSIONS": "all",
"FASTVIDEO_VSA_SM100A": "0",
})
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
def initialize(self, model_config: dict | None = None) -> None:
"""Download the fixed Preview adapter and load the FastH3 generator.
The model profile owns the base checkpoint, adapter file, attention
backend, and generation geometry. The backend translates that profile
into FastVideo's typed generator configuration.
"""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("FastH3 initialization requires a model configuration.")
if self.generator is not None:
self.generator.shutdown()
self.generator = None
gc.collect()
torch.cuda.empty_cache()
self.clear_conditioning()
model_path = _required_config_str(self.model_config, "model_path")
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
from huggingface_hub import hf_hub_download
from fastvideo import VideoGenerator
from fastvideo.api import (
AttentionConfig,
CompileConfig,
ComponentConfig,
EngineConfig,
GeneratorConfig,
MiniMaxH3Options,
OffloadConfig,
ParallelismConfig,
PipelineSelection,
)
adapter_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
use_vsa = attention_backend == "VIDEO_SPARSE_ATTN_H3"
generator_config = GeneratorConfig(
model_path=model_path,
pipeline=PipelineSelection(
components=ComponentConfig(lora_path=adapter_path, lora_strength=1.0),
model=MiniMaxH3Options(vae_parallel_decode=True, vae_parallel_decode_strategy="gather"),
),
engine=EngineConfig(
num_gpus=DREAMVERSE_SP_SIZE,
parallelism=ParallelismConfig(tp_size=1, sp_size=DREAMVERSE_SP_SIZE),
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
text_encoder=True,
image_encoder=True,
vae=True,
pin_cpu_memory=True,
),
compile=CompileConfig(enabled=False, vae_enabled=True, regional=attention_backend == "FLASH_ATTN"),
attention=AttentionConfig(
backend=attention_backend,
vsa_sparsity=0.9 if use_vsa else None,
vsa_tile_size=64 if use_vsa else None,
),
use_fsdp_inference=False,
),
)
print(f"[GPU {self.gpu_id}] Loading FastH3 model: {model_path}")
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
self.generator = VideoGenerator.from_config(generator_config)
print(f"[GPU {self.gpu_id}] FastH3 loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
"""Release the FastVideo generator and cached continuation image."""
self.clear_conditioning()
if self.generator is not None:
self.generator.shutdown()
self.generator = None
def clear_conditioning(self) -> None:
"""Release the first-frame image retained for the next segment."""
if self.continuation_image is not None:
self.continuation_image.close()
self.continuation_image = None
@staticmethod
def _load_rgb_image(image_path: str) -> Image:
"""Load an image into an independent RGB buffer with no open file handle."""
from PIL import Image
with Image.open(image_path) as image:
return image.convert("RGB").copy()
def _select_conditioning_image(
self,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> tuple[Image | None, bool]:
"""Select the initial upload or retained last frame for one segment."""
if reset_conditioning:
self.clear_conditioning()
if segment_idx > 1 and self.continuation_image is not None:
return self.continuation_image.copy(), True
if segment_idx > 1 and not reset_conditioning:
raise RuntimeError(f"FastH3 segment {segment_idx} requires a retained continuation frame.")
if segment_idx == 1 and image_path:
return self._load_rgb_image(image_path), False
return None, False
def _build_request(self, prompt: str, conditioning_image: Image | None):
"""Build the typed FastVideo request owned by the FastH3 profile."""
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
return GenerationRequest(
prompt=prompt,
negative_prompt="",
inputs=InputConfig(pil_image=conditioning_image),
sampling=SamplingConfig(
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
num_frames=int(self.model_config["num_frames"]),
fps=24,
num_inference_steps=int(self.model_config["num_inference_steps"]),
guidance_scale=1.0,
batch_cfg=False,
seed=int(self.model_config["seed"]),
),
output=OutputConfig(save_video=False, return_frames=True),
)
def _save_continuation_frame(self, frames: list) -> None:
"""Retain the last decoded frame as first-frame conditioning."""
from PIL import Image
self.clear_conditioning()
self.continuation_image = Image.fromarray(np.ascontiguousarray(frames[-1])).convert("RGB")
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> StepResult:
"""Generate one synchronized FastH3 segment and retain its last frame.
Later segments use MiniMax H3's first-frame-to-video path. The first
conditioned frame and its matching audio duration are trimmed before
streaming so adjacent segments do not duplicate media.
"""
if self.generator is None:
raise RuntimeError("FastH3 generator is not initialized.")
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
request = self._build_request(prompt, conditioning_image)
started = time.perf_counter()
try:
result = self.generator.generate(request)
finally:
if conditioning_image is not None:
conditioning_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
if isinstance(result, list):
raise RuntimeError("FastH3 returned multiple results for one DreamVerse segment.")
frames = result.frames
if not isinstance(frames, list) or not frames:
raise RuntimeError("FastH3 generation did not return decoded frames.")
audio = result.audio
audio_sample_rate = result.audio_sample_rate
if audio is not None and audio_sample_rate is None:
raise RuntimeError("FastH3 returned audio without an audio sample rate.")
save_started = time.perf_counter()
self._save_continuation_frame(frames)
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
timings = {
"generation_ms": generation_ms,
"generation_time_ms": float(result.generation_time or 0.0) * 1000.0,
"save_conditioning_ms": save_conditioning_ms,
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
}
trim_frames = 1 if uses_continuation else 0
print(f"[GPU {self.gpu_id}] FastH3 segment {segment_idx}: "
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
f"save_conditioning={save_conditioning_ms:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
return StepResult(
frames=frames,
audio=audio,
audio_sample_rate=audio_sample_rate,
timings=timings,
head_trim_frames=trim_frames,
head_trim_audio_frames=trim_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
"""Compile the FastH3 text and first-frame paths before readiness."""
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] FastH3 startup warmup starting "
"(synthetic segments: text-to-video, first-frame-to-video)")
started = time.perf_counter()
text_result = self.generate_step(
warmup_prompt,
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
first_frame_result = self.generate_step(
warmup_prompt,
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
total_ms = (time.perf_counter() - started) * 1000.0
self.clear_conditioning()
text_ms = float(text_result.timings.get("e2e_latency_ms", 0.0))
first_frame_ms = float(first_frame_result.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] FastH3 startup warmup complete: "
f"text_to_video={text_ms:.0f}ms, "
f"first_frame_to_video={first_frame_ms:.0f}ms, "
f"total={total_ms:.0f}ms")
return {
"warmup_text_to_video_ms": text_ms,
"warmup_first_frame_to_video_ms": first_frame_ms,
"warmup_total_ms": total_ms,
}
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
"""Reject runtime LoRA mutation because FastH3 uses one startup adapter."""
del stack
raise RuntimeError("FastH3 uses its fixed startup adapter and does not support runtime LoRA changes.")
@@ -30,7 +30,7 @@ from dreamverse.session_init_image import cleanup_session_init_image, persist_se
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
GENERATION_SEGMENT_CAP,
PROMPT_AUTO_SLEEP_MS,
PROMPT_AUTO_TIMEOUT_MS,
@@ -264,7 +264,7 @@ class SessionController:
timeout_task = asyncio.create_task(session_timeout())
# Join the engine on this GPU.
await slot.join_user(client_id, model_id=DEFAULT_MODEL_ID)
await slot.join_user(client_id, model_id=ACTIVE_MODEL_ID)
# Notify client they're connected to a GPU.
await ws_send_json({
@@ -3,7 +3,6 @@ from __future__ import annotations
import sys
from pathlib import Path
TESTS_DIR = Path(__file__).resolve().parent
DREAMVERSE_PACKAGE_DIR = TESTS_DIR.parent
DREAMVERSE_APP_DIR = DREAMVERSE_PACKAGE_DIR.parent
+46 -22
View File
@@ -2,14 +2,14 @@ from __future__ import annotations
import importlib.util
from pathlib import Path
from types import ModuleType
import pytest
SERVER_DIR = Path(__file__).resolve().parents[1]
def _load_config_module():
def _load_config_module() -> ModuleType:
spec = importlib.util.spec_from_file_location(
"server_config_test_module",
SERVER_DIR / "config.py",
@@ -21,7 +21,7 @@ def _load_config_module():
return module
def _set_required_prompt_keys(monkeypatch):
def _set_required_prompt_keys(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
@@ -53,9 +53,7 @@ def test_config_defaults_to_cerebras_with_parallel_groq_fallback_stage(monkeypat
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -86,9 +84,7 @@ def test_config_ignores_legacy_groq_primary_override(monkeypatch):
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -106,24 +102,17 @@ def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_
assert module.DEVTOOLS_ENABLED is True
assert module.FRONTEND_ROOT.as_posix().endswith("apps/dreamverse/web")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/next_segment_system_prompt.md"
)
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith("dreamverse/prompts.local/next_segment_system_prompt.md")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/next_segment_system_prompt.md"
)
"dreamverse/prompts/next_segment_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/rewrite_user_system_prompt.md"
)
"dreamverse/prompts.local/rewrite_user_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/rewrite_user_system_prompt.md"
)
"dreamverse/prompts/rewrite_user_system_prompt.md")
assert module.CURATED_PRESETS_FILE_PATH.endswith(
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json")
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json")
assert module.FRONTEND_STATIC_DIR_CANDIDATES[:2] == (
str(module.FRONTEND_ROOT / "out"),
str(module.FRONTEND_ROOT / "dist"),
@@ -162,3 +151,38 @@ def test_config_rejects_invalid_prompt_provider(monkeypatch):
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_PROMPT_PROVIDER"):
_load_config_module()
def test_config_registers_vsa_datafree_fasth3_profile(monkeypatch):
"""The FastH3 registry entry owns the complete fixed Preview recipe."""
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.MODEL_REGISTRY["fast-h3"] == {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
def test_config_uses_fasth3_sequence_parallel_default(monkeypatch):
"""Selecting FastH3 defaults DreamVerse to its four-GPU topology."""
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "fast-h3")
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
module = _load_config_module()
assert module.ACTIVE_MODEL_ID == "fast-h3"
assert module.MODEL_CONFIG["generation_backend"] == "minimax_h3"
assert module.DREAMVERSE_SP_SIZE == 4
@@ -9,6 +9,7 @@ from fastapi.testclient import TestClient
import fastvideo.entrypoints.streaming as streaming_entrypoints
import pytest
def _install_stack03_import_stubs(monkeypatch):
"""Keep entrypoint tests focused while later-stack runtime modules are absent."""
if not hasattr(streaming_entrypoints, "build_health_router"):
@@ -17,6 +18,7 @@ def _install_stack03_import_stubs(monkeypatch):
gpu_pool_stub = types.ModuleType("dreamverse.gpu_pool")
class GPUPool:
def __init__(self, _gpu_ids):
pass
@@ -49,6 +51,7 @@ def _install_stack03_import_stubs(monkeypatch):
controller_stub = types.ModuleType("dreamverse.session.controller")
class SessionController:
def __init__(self, **_kwargs):
pass
@@ -76,13 +79,11 @@ def _run_cli(module, monkeypatch, argv: list[str]) -> list[dict[str, object]]:
uvicorn_stub = types.ModuleType("uvicorn")
def run(app, host: str, port: int) -> None:
calls.append(
{
"app": app,
"host": host,
"port": port,
}
)
calls.append({
"app": app,
"host": host,
"port": port,
})
uvicorn_stub.run = run
monkeypatch.setitem(sys.modules, "uvicorn", uvicorn_stub)
@@ -99,13 +100,11 @@ def test_server_cli_defaults_to_local_web_port(monkeypatch):
server_main = _import_server_main(monkeypatch)
calls = _run_cli(server_main, monkeypatch, ["dreamverse-server"])
assert calls == [
{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_server_cli_allows_explicit_host_and_port(monkeypatch):
@@ -116,13 +115,11 @@ def test_server_cli_allows_explicit_host_and_port(monkeypatch):
["dreamverse-server", "--host", "127.0.0.1", "--port", "8123"],
)
assert calls == [
{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}
]
assert calls == [{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}]
def test_server_does_not_expose_backend_source_as_static_assets(monkeypatch):
@@ -142,13 +139,11 @@ def test_mock_server_cli_defaults_to_local_web_port(monkeypatch):
["dreamverse-mock-server"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_mock_server_cli_updates_latency(monkeypatch):
@@ -161,13 +156,11 @@ def test_mock_server_cli_updates_latency(monkeypatch):
["dreamverse-mock-server", "--latency", "321", "--port", "8111"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}]
assert mock_server.LATENCY_MS == 321
finally:
mock_server.LATENCY_MS = old_latency_ms
@@ -0,0 +1,9 @@
from dreamverse.generation_worker import _create_generation_backend
from dreamverse.ltx2_generation import LTX2GenerationBackend
def test_create_generation_backend_ltx2_module_import():
backend = _create_generation_backend("ltx2", gpu_id=3)
assert isinstance(backend, LTX2GenerationBackend)
assert backend.gpu_id == 3
@@ -7,7 +7,6 @@ from types import SimpleNamespace
import pytest
import dreamverse.gpu_pool as gpu_pool
@@ -64,6 +63,14 @@ def test_get_available_gpus_defaults_to_first_visible_device(monkeypatch):
assert gpu_pool.get_available_gpus() == [3]
def test_get_available_gpus_defaults_to_active_model_sequence_parallel_size(monkeypatch):
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "0,1,2,3,4")
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
monkeypatch.setattr(gpu_pool, "DREAMVERSE_SP_SIZE", 4)
assert gpu_pool.get_available_gpus() == [0, 1, 2, 3]
def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "zero")
@@ -72,6 +79,23 @@ def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
gpu_pool.get_available_gpus()
def test_join_user_failed_reload_marks_model_uninitialized(monkeypatch):
"""A failed model reload forces the next join to reload a model."""
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.current_model_id = "fast-ltx2"
async def fake_send_command(command, timeout):
del command, timeout
return gpu_pool.WorkerError(user_id="__reload__", message="load failed")
monkeypatch.setattr(slot, "_send_command", fake_send_command)
with pytest.raises(RuntimeError, match="Model reload failed"):
asyncio.run(slot.join_user("client-id", model_id="fast-h3"))
assert slot.current_model_id is None
def test_send_command_raises_on_worker_death():
"""A worker that consumes a command and exits without replying must
surface as RuntimeError via sentinel detection, not after the long
@@ -85,9 +109,7 @@ def test_send_command_raises_on_worker_death():
cmd_q = ctx.Queue()
resp_q = ctx.Queue()
proc = ctx.Process(
target=_child_consume_and_exit, args=(cmd_q, resp_q)
)
proc = ctx.Process(target=_child_consume_and_exit, args=(cmd_q, resp_q))
proc.start()
# Wait for the spawn child to fully boot. Allow generous time —
@@ -95,9 +117,9 @@ def test_send_command_raises_on_worker_death():
ready = resp_q.get(timeout=30.0)
assert ready == "READY"
async def runner():
async def runner() -> None:
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.process = proc
slot.process = proc # type: ignore[assignment]
slot.command_queue = cmd_q
slot.response_queue = resp_q
@@ -7,21 +7,20 @@ ALLOWED_PREFIXES = (
"fastvideo.entrypoints.video_generator",
"fastvideo.configs",
)
ALLOWED_EXACT = ("fastvideo",)
ALLOWED_EXACT = ("fastvideo", )
FORBIDDEN_PREFIXES = (
"fastvideo.pipelines",
"fastvideo.models",
"fastvideo.layers",
"fastvideo.worker",
"fastvideo.fastvideo_args",
)
ALLOWED_INTERNAL_IMPORTS = {
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.audio.ltx2_audio_processing",
),
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.loader.component_loader",
),
}
@@ -38,19 +37,13 @@ def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
except SyntaxError as task_exc:
raise AssertionError(f"Failed to parse {path}") from task_exc
for node in ast.walk(tree):
names = (
[a.name for a in node.names] if isinstance(node, ast.Import)
else [node.module] if isinstance(node, ast.ImportFrom) and node.module
else []
)
names = ([a.name for a in node.names] if isinstance(node, ast.Import) else
[node.module] if isinstance(node, ast.ImportFrom) and node.module else [])
for name in names:
if not name:
continue
rel_path = str(path.relative_to(root))
if (
name.startswith(FORBIDDEN_PREFIXES)
and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS
):
if (name.startswith(FORBIDDEN_PREFIXES) and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS):
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
assert bad == [], f"Forbidden internal imports: {bad}"
@@ -0,0 +1,246 @@
from __future__ import annotations
import os
from types import SimpleNamespace
from typing import Any
import numpy as np
import pytest
import dreamverse.generation_worker as generation_worker
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
FASTH3_MODEL_CONFIG = {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
class _RecordingGenerator:
"""Record typed requests and return small synchronized media fixtures."""
def __init__(self) -> None:
self.requests: list[Any] = []
self.conditioning_pixels: list[np.ndarray | None] = []
def generate(self, request):
"""Capture the request and return two tiny video frames with audio."""
self.requests.append(request)
conditioning_image = request.inputs.pil_image
self.conditioning_pixels.append(
None if conditioning_image is None else np.asarray(conditioning_image).copy())
frames = [
np.full((2, 3, 3), 10, dtype=np.uint8),
np.full((2, 3, 3), 20, dtype=np.uint8),
]
return SimpleNamespace(
frames=frames,
audio=np.zeros((2, 16), dtype=np.float32),
audio_sample_rate=44100,
generation_time=0.25,
)
def test_initialize_builds_vsa_datafree_fasth3_generator(monkeypatch):
"""Initialization translates the DreamVerse profile into typed FastVideo config."""
from fastvideo import VideoGenerator
captured = {}
fake_generator = SimpleNamespace(shutdown=lambda: None)
def fake_from_config(config):
captured["config"] = config
return fake_generator
def fake_download(**kwargs):
captured["download"] = kwargs
return f"/models/{kwargs['filename']}"
monkeypatch.setattr("huggingface_hub.hf_hub_download", fake_download)
monkeypatch.setattr(VideoGenerator, "from_config", fake_from_config)
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
monkeypatch.setenv("FASTVIDEO_FA4", "0")
monkeypatch.setenv("FASTVIDEO_MINIMAX_H3_FUSIONS", "0")
monkeypatch.setenv("FASTVIDEO_VSA_SM100A", "1")
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
backend = MiniMaxH3GenerationBackend(gpu_id=0)
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
backend.initialize(FASTH3_MODEL_CONFIG)
config = captured["config"]
assert captured["download"] == {
"repo_id": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"filename": "vsa-datafree/adapter_model.safetensors",
}
assert config.model_path == "MiniMaxAI/MiniMax-H3"
assert config.pipeline.components.lora_path.endswith("vsa-datafree/adapter_model.safetensors")
assert config.pipeline.components.lora_strength == 1.0
assert config.pipeline.experimental == {}
assert config.engine.attention.backend == "VIDEO_SPARSE_ATTN_H3"
assert config.engine.attention.vsa_sparsity == 0.9
assert config.engine.attention.vsa_tile_size == 64
assert config.engine.compile.regional is False
assert config.pipeline.model.vae_parallel_decode is True
assert config.pipeline.model.vae_parallel_decode_strategy == "gather"
assert config.engine.num_gpus == 4
assert config.engine.parallelism.tp_size == 1
assert config.engine.parallelism.sp_size == 4
assert config.engine.offload.dit is False
assert config.engine.offload.dit_layerwise is False
assert config.engine.offload.text_encoder is True
assert config.engine.offload.vae is True
assert config.engine.compile.vae_enabled is True
assert config.engine.use_fsdp_inference is False
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "VIDEO_SPARSE_ATTN_H3"
assert os.environ["FASTVIDEO_FA4"] == "1"
assert os.environ["FASTVIDEO_MINIMAX_H3_FUSIONS"] == "all"
assert os.environ["FASTVIDEO_VSA_SM100A"] == "0"
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
def test_initialize_selects_declared_generation_backend(monkeypatch):
"""The GPU worker constructs the backend that the active model profile declares."""
from unittest.mock import Mock
selected_backend = Mock()
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: selected_backend,
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend_name == "minimax_h3"
assert worker.backend is selected_backend
selected_backend.initialize.assert_called_once_with(FASTH3_MODEL_CONFIG)
def test_initialize_failure_clears_backend_ownership(monkeypatch):
"""A failed family change leaves the GPU worker explicitly uninitialized."""
ltx_backend = SimpleNamespace(initialize=lambda config: None, shutdown=lambda: None)
def fail_initialize(config):
del config
raise RuntimeError("load failed")
fasth3_backend = SimpleNamespace(
initialize=fail_initialize,
shutdown=lambda: None,
)
backends = {
"ltx2": ltx_backend,
"minimax_h3": fasth3_backend,
}
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: backends[backend_name],
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize({"generation_backend": "ltx2"})
with pytest.raises(RuntimeError, match="load failed"):
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend is None
assert worker.backend_name is None
assert worker.model_config == {"generation_backend": "ltx2"}
def test_generate_step_uses_last_frame_for_continuation(monkeypatch):
"""A later segment receives the prior segment's last decoded frame."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
first_result = backend.generate_step(
"first prompt",
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
second_result = backend.generate_step(
"second prompt",
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
first_request = backend.generator.requests[0]
assert first_request.inputs.pil_image is None
assert first_request.negative_prompt == ""
assert first_request.sampling.height == 768
assert first_request.sampling.width == 1344
assert first_request.sampling.num_frames == 124
assert first_request.sampling.num_inference_steps == 5
assert first_request.sampling.fps == 24
assert first_request.sampling.guidance_scale == 1.0
assert first_request.sampling.batch_cfg is False
assert first_request.sampling.seed == 1000
assert first_request.output.save_video is False
assert first_request.output.return_frames is True
assert backend.generator.conditioning_pixels[1].tolist() == np.full((2, 3, 3), 20).tolist()
assert first_result.head_trim_frames == 0
assert first_result.head_trim_audio_frames == 0
assert second_result.head_trim_frames == 1
assert second_result.head_trim_audio_frames == 1
assert second_result.audio_sample_rate == 44100
def test_generate_step_reset_uses_text_to_video_path(monkeypatch):
"""Resetting continuation produces an unconditioned text-to-video request."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
backend.generate_step("first prompt", 1, None, True)
reset_result = backend.generate_step("reset prompt", 2, None, True)
assert backend.generator.requests[-1].inputs.pil_image is None
assert reset_result.head_trim_frames == 0
assert reset_result.head_trim_audio_frames == 0
def test_generate_step_missing_continuation_frame(monkeypatch):
"""A later segment fails when no reset or retained frame defines its input."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
backend.generate_step("later prompt", 2, None, False)
assert backend.generator.requests == []
def test_warmup_exercises_text_and_first_frame_paths(monkeypatch):
"""Warmup covers both request shapes used by a DreamVerse session."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
timings = backend.warmup("warmup prompt")
assert backend.generator.conditioning_pixels[0] is None
assert backend.generator.conditioning_pixels[1] is not None
assert backend.continuation_image is None
assert "warmup_text_to_video_ms" in timings
assert "warmup_first_frame_to_video_ms" in timings
@@ -6,7 +6,6 @@ import os
from fastapi import WebSocketDisconnect
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -14,6 +13,7 @@ import dreamverse.mock_server as mock_server
class _FakeWebSocket:
def __init__(self, messages: list[tuple[float, dict[str, object]]]):
self._messages = messages
self._index = 0
@@ -49,34 +49,34 @@ def test_mock_server_matches_current_single5s_protocol():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -92,24 +92,14 @@ def test_mock_server_matches_current_single5s_protocol():
assert message_types.count("ltx2_stream_complete") == 2
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
gpu_assigned_event = next(
payload for payload in ws.sent_json if payload["type"] == "gpu_assigned"
)
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
gpu_assigned_event = next(payload for payload in ws.sent_json if payload["type"] == "gpu_assigned")
assert gpu_assigned_event["session_timeout"] == mock_server.SESSION_TIMEOUT_SECONDS
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "selected prompt"
assert segment_start_events[1]["prompt"] == "custom prompt"
step_complete_events = [
payload
for payload in ws.sent_json
if payload["type"] == "step_complete"
]
step_complete_events = [payload for payload in ws.sent_json if payload["type"] == "step_complete"]
assert len(step_complete_events) == 2
assert step_complete_events[0]["latency_ms"] == {
"total": 121.0,
@@ -134,29 +124,29 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
mock_server.LATENCY_MS = 1
mock_server.GENERATION_SEGMENT_CAP = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -166,11 +156,7 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
assert "generation_cap_reached" not in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "segment one"
assert segment_start_events[1]["prompt"] == "segment one [start a new rollout]"
@@ -187,54 +173,40 @@ def test_mock_server_rewrite_during_active_segment_restarts_from_first_rewritten
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 100
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment one [restart from rewrite]",
]
assert all(
payload["prompt"] != "segment two"
for payload in segment_start_events[1:]
)
reset_events = [
payload
for payload in ws.sent_json
if payload.get("type") == "seed_prompts_reset_applied"
]
assert any(
payload.get("reason") == "rewrite_during_generation"
for payload in reset_events
)
assert all(payload["prompt"] != "segment two" for payload in segment_start_events[1:])
reset_events = [payload for payload in ws.sent_json if payload.get("type") == "seed_prompts_reset_applied"]
assert any(payload.get("reason") == "rewrite_during_generation" for payload in reset_events)
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -247,24 +219,24 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -276,15 +248,9 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
assert "ltx2_stream_start" in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert segment_start_events
assert segment_start_events[0]["prompt"] == (
"A moonbase corridor thriller with flooding [segment 1]"
)
assert segment_start_events[0]["prompt"] == ("A moonbase corridor thriller with flooding [segment 1]")
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -297,35 +263,37 @@ def test_mock_server_can_start_new_project_without_reconnecting():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 40
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {"type": "end_project_keep_session"}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {
"type": "end_project_keep_session"
}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -336,16 +304,11 @@ def test_mock_server_can_start_new_project_without_reconnecting():
project_idle_index = message_types.index("project_idle")
stream_start_indexes = [
index for index, message_type in enumerate(message_types)
if message_type == "ltx2_stream_start"
index for index, message_type in enumerate(message_types) if message_type == "ltx2_stream_start"
]
assert stream_start_indexes[0] < project_idle_index < stream_start_indexes[1]
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment two",
@@ -6,7 +6,6 @@ import os
import re
import time
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -22,6 +21,7 @@ from dreamverse.prompt_enhancer import (
class _FakeResponse:
def __init__(self, payload: dict):
self._payload = payload
@@ -30,6 +30,7 @@ class _FakeResponse:
class _FakeSyncCompletions:
def __init__(self, payload: dict):
self._payload = payload
@@ -38,6 +39,7 @@ class _FakeSyncCompletions:
class _FakeSyncClient:
def __init__(self, payload: dict):
self.chat = type(
"_FakeChat",
@@ -47,6 +49,7 @@ class _FakeSyncClient:
class _DelayedSyncCompletions:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self._payload = payload
self._delay_s = delay_s
@@ -61,29 +64,26 @@ class _DelayedSyncCompletions:
class _DelayedSyncClient:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self.chat = type(
"_FakeChat",
(),
{
"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)
},
{"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)},
)()
def _chat_payload_with_content(content: str) -> dict:
return {
"choices": [
{
"message": {
"content": content,
}
"choices": [{
"message": {
"content": content,
}
]
}]
}
@@ -172,6 +172,7 @@ def _build_staged_enhancer(
class _FakeOpenAIClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -182,6 +183,7 @@ class _FakeOpenAIClient:
class _FakeCerebrasClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -192,16 +194,12 @@ class _FakeCerebrasClient:
def test_parse_json_response_accepts_fenced_json_with_prose():
parsed = _parse_json_response(
"Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks."
)
parsed = _parse_json_response("Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks.")
assert parsed == {"segment_prompts": ["A", "B"]}
def test_parse_json_response_extracts_first_embedded_object():
parsed = _parse_json_response(
"Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)"
)
parsed = _parse_json_response("Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)")
assert parsed == {"segment_prompts": ["A", "B"]}
@@ -268,16 +266,12 @@ def test_build_client_supports_groq_provider(monkeypatch):
def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -286,15 +280,12 @@ def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"rewritten_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"rewritten_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -303,19 +294,14 @@ def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
def test_rewrite_prompt_sequence_accepts_segment_dicts_without_top_level_rollout_metadata():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"segments":[{"prompt":"A"},{"text":"B"}]}'
)
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segments":[{"prompt":"A"},{"text":"B"}]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
preset_id="preset_a",
preset_label="Preset A",
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -329,14 +315,12 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
"The user is asking for a cinematic rewrite.\n\n"
"1. A dog bounds across the moon's dusty surface, kicking up silver regolith as it chases a rabbit beneath the black sky.\n"
"2. The rabbit darts around a crater rim while the dog lunges after it, Earth glowing blue in the distance.\n"
)
)
))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -347,29 +331,27 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
]
def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
def test_enhance_prompt_uses_groq_when_it_returns_first():
enhancer = _build_staged_enhancer(
cerebras_payload=_chat_payload_with_content('{"prompt":"Cerebras prompt"}'),
groq_payload=_chat_payload_with_content('{"prompt":"Groq prompt"}'),
cerebras_delay_s=0.01,
cerebras_delay_s=0.08,
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
assert result.provider == "cerebras"
assert result.provider == "groq"
assert result.model == "gpt-test"
assert result.prompt == "Cerebras prompt"
assert result.prompt == "Groq prompt"
assert enhancer.get_provider_success_counts() == {
"cerebras": 1,
"groq": 0,
"cerebras": 0,
"groq": 1,
}
@@ -381,12 +363,10 @@ def test_enhance_prompt_uses_groq_when_cerebras_fails():
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -408,13 +388,11 @@ def test_enhance_prompt_can_use_groq_when_cerebras_times_out():
enhancer.http_timeout_ms = 50
enhancer.default_timeout_ms = 50
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
))
assert result.fallback_used is False
assert result.error is None
@@ -434,12 +412,10 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
groq_delay_s=0.08,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -453,15 +429,12 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
enhancer = _build_test_enhancer(
_chat_payload_with_content("I cannot comply with JSON right now.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("I cannot comply with JSON right now."))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in (result.error or "")
assert result.raw_response_text == "I cannot comply with JSON right now."
@@ -473,9 +446,7 @@ def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
)
)
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'))
captured = {
"body": None,
"timeout_seconds": None,
@@ -486,8 +457,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
captured["timeout_seconds"] = timeout_seconds
return (
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}',
)
@@ -502,8 +472,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -512,12 +481,12 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
}
assert captured["body"]["messages"][1]["role"] == "user"
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "edit_existing_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "make it cinematic",
"mode":
"edit_existing_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"make it cinematic",
"current_rollout": {
"id": "preset_a",
"label": "Preset A",
@@ -528,11 +497,8 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
def test_rewrite_prompt_sequence_supports_new_rollout_mode():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
)
)
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'))
captured = {
"body": None,
}
@@ -541,10 +507,8 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
),
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'),
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}',
)
@@ -560,30 +524,29 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.prompts == ["A", "B", "C", "D", "E", "F"]
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "new_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count": 6,
"rollout_id_hint": "custom_editable",
"rollout_label_hint": "Custom rollout",
"mode":
"new_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count":
6,
"rollout_id_hint":
"custom_editable",
"rollout_label_hint":
"Custom rollout",
}
def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared system prompt"
captured = {
"body": None,
@@ -593,9 +556,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
),
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'),
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}',
)
@@ -609,8 +570,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
rewrite_instruction="make it cinematic",
rewrite_model="gpt-test",
system_prompt_override="session specific system prompt",
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -621,10 +581,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
@@ -635,24 +592,17 @@ def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_prefers_override():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt(
"session specific system prompt"
)
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt("session specific system prompt")
assert resolved == "session specific system prompt"
def test_generate_auto_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Auto next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Auto next"}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
enhancer.rewrite_default_model = "gpt-test"
@@ -680,8 +630,7 @@ def test_generate_auto_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Auto next"
@@ -690,9 +639,7 @@ def test_generate_auto_prompt_uses_selected_model():
def test_enhance_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Enhanced next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Enhanced next"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -722,8 +669,7 @@ def test_enhance_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Enhanced next"
@@ -732,9 +678,7 @@ def test_enhance_prompt_uses_selected_model():
def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"prompt":"Extended single clip"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"prompt":"Extended single clip"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -764,14 +708,12 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
enhancer._request_content = _fake_request_content # type: ignore[attr-defined]
result = asyncio.run(
enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Extended single clip"
@@ -784,17 +726,15 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
"single 5-second LTX-2.3 video clip. Respond with "
'valid JSON only as {"prompt": "..."}.' # noqa: E501
),
"user_prompt": "short 5s idea",
"user_prompt":
"short 5s idea",
}
def test_enhance_prompt_single_clip_rejects_plain_text_response():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
"Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."
)
)
_chat_payload_with_content("Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."))
enhancer.auto_system_prompt = "auto system prompt"
result = asyncio.run(
@@ -803,17 +743,14 @@ def test_enhance_prompt_single_clip_rejects_plain_text_response():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in result.error
assert result.prompt == ""
def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"segment_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segment_prompts":["A","B"]}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -824,17 +761,14 @@ def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "Missing prompt string." in (result.error or "")
def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A cinematic continuation with slow dolly movement.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A cinematic continuation with slow dolly movement."))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -846,17 +780,14 @@ def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A calm, grounded continuation with subtle motion.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A calm, grounded continuation with subtle motion."))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -867,34 +798,30 @@ def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
enhancer = _build_test_enhancer(
{
"choices": [
{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}
],
"usage": {"completion_tokens": 0},
}
)
enhancer = _build_test_enhancer({
"choices": [{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}],
"usage": {
"completion_tokens": 0
},
})
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No rewrite segment prompts found in assistant response." in (result.error or "")
assert isinstance(result.raw_response_text, str)
@@ -903,10 +830,7 @@ def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
def test_get_rewrite_model_config_returns_fixed_defaults():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_default_model = "gpt-oss-120b"
enhancer.rewrite_model_options = ["gpt-oss-120b"]
@@ -918,10 +842,7 @@ def test_get_rewrite_model_config_returns_fixed_defaults():
def test_get_prompt_config_includes_auto_extension_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.enhance_system_prompt_path = "/tmp/next.md"
enhancer.auto_system_prompt_path = "/tmp/auto.md"
enhancer.rewrite_all_system_prompt_path = "/tmp/rewrite.md"
@@ -948,19 +869,14 @@ def test_get_prompt_config_includes_auto_extension_prompt():
def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
rewrite_fallback_path = tmp_path / "rewrite_window_system_prompt.md"
rewrite_fallback_path.write_text("rewrite prompt\n", encoding="utf-8")
next_path = tmp_path / "next.md"
next_path.write_text("next prompt\n", encoding="utf-8")
auto_path = tmp_path / "auto.md"
auto_path.write_text("auto prompt\n", encoding="utf-8")
enhancer.rewrite_all_system_prompt_path = str(
tmp_path / "prompts.local" / "rewrite_window_system_prompt.md"
)
enhancer.rewrite_all_system_prompt_path = str(tmp_path / "prompts.local" / "rewrite_window_system_prompt.md")
enhancer.rewrite_all_system_prompt_fallback_path = str(rewrite_fallback_path)
enhancer.enhance_system_prompt_path = str(next_path)
enhancer.auto_system_prompt_path = str(auto_path)
@@ -973,14 +889,9 @@ def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
assert config["rewrite_window_system_prompt_path"] == str(rewrite_fallback_path)
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(
tmp_path,
):
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(tmp_path, ):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1007,10 +918,7 @@ def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_emp
def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1021,9 +929,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer.auto_system_prompt_path = str(auto_path)
enhancer.rewrite_all_system_prompt_path = str(rewrite_path)
config = enhancer.save_prompt_config(
auto_extension_system_prompt="auto updated",
)
config = enhancer.save_prompt_config(auto_extension_system_prompt="auto updated", )
assert auto_path.read_text(encoding="utf-8").strip() == "auto updated"
assert config["auto_extension_system_prompt"] == "auto updated"
@@ -1031,10 +937,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1052,9 +955,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.rewrite_user_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_user_system_prompt="rewrite user updated",
)
config = enhancer.save_prompt_config(rewrite_user_system_prompt="rewrite user updated", )
assert rewrite_user_path.read_text(encoding="utf-8").strip() == "rewrite user updated"
assert config["rewrite_user_system_prompt"] == "rewrite user updated"
@@ -1062,10 +963,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1081,9 +979,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer.rewrite_default_model = "gpt-test"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
config = enhancer.save_prompt_config(
rewrite_model="gpt-alt",
)
config = enhancer.save_prompt_config(rewrite_model="gpt-alt", )
assert enhancer.rewrite_default_model == "gpt-alt"
assert config["rewrite_model"] == "gpt-alt"
@@ -1092,10 +988,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1109,9 +1002,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_temperature=1.3,
)
config = enhancer.save_prompt_config(rewrite_temperature=1.3, )
assert enhancer.rewrite_default_temperature == 1.3
assert config["rewrite_temperature"] == 1.3
@@ -1119,10 +1010,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1136,13 +1024,9 @@ def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_pat
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.save_prompt_config(
rewrite_window_system_prompt="rewrite updated",
)
enhancer.save_prompt_config(rewrite_window_system_prompt="rewrite updated", )
backup_paths = sorted(
tmp_path.glob("rewrite_window_system_prompt.*.bak.md")
)
backup_paths = sorted(tmp_path.glob("rewrite_window_system_prompt.*.bak.md"))
assert rewrite_path.read_text(encoding="utf-8").strip() == "rewrite updated"
assert len(backup_paths) == 1
@@ -27,13 +27,8 @@ try:
except ModuleNotFoundError:
websockets = None # type: ignore[assignment]
DEFAULT_PRESET_FILE = (
Path(__file__).resolve().parents[2]
/ "web"
/ "prompts"
/ "selected_ltx2_continuation_story_presets.json"
)
DEFAULT_PRESET_FILE = (Path(__file__).resolve().parents[2] / "web" / "prompts" /
"selected_ltx2_continuation_story_presets.json")
def utc_now_iso() -> str:
@@ -65,10 +60,7 @@ def safe_percentile(values: list[float], percentile: float) -> float | None:
if lower == upper:
return sorted_values[lower]
fraction = rank - lower
return (
sorted_values[lower]
+ (sorted_values[upper] - sorted_values[lower]) * fraction
)
return (sorted_values[lower] + (sorted_values[upper] - sorted_values[lower]) * fraction)
def summarize_series(values: list[float]) -> dict[str, float | int | None]:
@@ -145,24 +137,16 @@ def load_curated_prompts(
selected_id = str(selected.get("id", "")).strip() or "unknown_preset"
raw_prompts = selected.get("segment_prompts", [])
if not isinstance(raw_prompts, list):
raise ValueError(
f"Preset {selected_id} has invalid segment_prompts (must be list)."
)
raise ValueError(f"Preset {selected_id} has invalid segment_prompts (must be list).")
prompts = [
str(prompt).strip()
for prompt in raw_prompts
if isinstance(prompt, str) and str(prompt).strip()
]
prompts = [str(prompt).strip() for prompt in raw_prompts if isinstance(prompt, str) and str(prompt).strip()]
if not prompts:
raise ValueError(f"Preset {selected_id} has no non-empty prompts.")
limited = prompts[:curated_limit]
if not limited:
raise ValueError(
f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}."
)
raise ValueError(f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}.")
return selected_id, limited, len(prompts)
@@ -224,11 +208,11 @@ async def run_single_session(
try:
async with websockets.connect(
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
) as ws:
connect_finish_monotonic = time.monotonic()
session_data["connect_finish_ts_utc"] = utc_now_iso()
@@ -249,9 +233,7 @@ async def run_single_session(
timeout_remaining = session_timeout_s - elapsed_s
if timeout_remaining <= 0:
session_data["status"] = "timeout"
session_data["error"] = (
f"Session timed out after {session_timeout_s:.1f}s."
)
session_data["error"] = (f"Session timed out after {session_timeout_s:.1f}s.")
break
recv_start_epoch = time.time()
@@ -265,9 +247,7 @@ async def run_single_session(
)
except asyncio.TimeoutError:
session_data["status"] = "timeout"
session_data["error"] = (
"Timed out waiting for websocket message."
)
session_data["error"] = ("Timed out waiting for websocket message.")
break
except Exception as exc:
session_data["status"] = "failed"
@@ -288,20 +268,16 @@ async def run_single_session(
chunk_gap_ms: float | None = None
if last_chunk_finish_monotonic is not None:
chunk_gap_ms = (
recv_finish_monotonic - last_chunk_finish_monotonic
) * 1000.0
chunk_gap_ms = (recv_finish_monotonic - last_chunk_finish_monotonic) * 1000.0
session_data["chunks"].append(
{
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
}
)
session_data["chunks"].append({
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
})
last_chunk_finish_monotonic = recv_finish_monotonic
last_chunk_finish_epoch = recv_finish_epoch
session_data["last_chunk_finish_ts_utc"] = recv_finish_iso
@@ -321,9 +297,7 @@ async def run_single_session(
if msg_type == "gpu_assigned":
session_data["gpu_assigned_ts_utc"] = recv_finish_iso
if connect_finish_monotonic is not None:
session_data["queue_wait_ms"] = (
recv_finish_monotonic - connect_finish_monotonic
) * 1000.0
session_data["queue_wait_ms"] = (recv_finish_monotonic - connect_finish_monotonic) * 1000.0
elif msg_type == "ltx2_stream_start":
if initial_total_segments is None:
parsed_total = parse_int(data.get("total_segments"))
@@ -338,20 +312,13 @@ async def run_single_session(
session_data["media_segments_completed"] += 1
if first_media_segment_complete_epoch is None:
first_media_segment_complete_epoch = recv_finish_epoch
session_data[
"first_media_segment_complete_ts_utc"
] = recv_finish_iso
session_data["first_media_segment_complete_ts_utc"] = recv_finish_iso
elif msg_type == "ltx2_segment_complete":
session_data["segments_completed"] += 1
seg_idx = parse_int(data.get("segment_idx"))
if (
initial_total_segments is not None
and seg_idx is not None
and seg_idx >= initial_total_segments
):
session_data[
"target_segment_complete_ts_utc"
] = recv_finish_iso
if (initial_total_segments is not None and seg_idx is not None
and seg_idx >= initial_total_segments):
session_data["target_segment_complete_ts_utc"] = recv_finish_iso
await asyncio.sleep(post_complete_wait_s)
session_data["leave_sent_ts_utc"] = utc_now_iso()
try:
@@ -362,15 +329,11 @@ async def run_single_session(
break
elif msg_type == "session_timeout":
session_data["status"] = "timeout"
session_data["error"] = str(
data.get("message") or "Backend session timeout"
)
session_data["error"] = str(data.get("message") or "Backend session timeout")
break
elif msg_type == "error":
session_data["status"] = "failed"
session_data["error"] = str(
data.get("message") or "Backend error message"
)
session_data["error"] = str(data.get("message") or "Backend error message")
break
if session_data["status"] == "failed" and session_data["error"] is None:
@@ -379,29 +342,18 @@ async def run_single_session(
session_data["status"] = "failed"
session_data["error"] = f"WebSocket connect/run failed: {exc}"
if (
first_chunk_finish_epoch is not None
and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0
):
if (first_chunk_finish_epoch is not None and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0):
duration_s = last_chunk_finish_epoch - first_chunk_finish_epoch
if duration_s > 0:
session_data["session_goodput_mbps"] = (
session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0
)
session_data["session_goodput_mbps"] = (session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0)
if (
first_chunk_finish_epoch is not None
and first_media_segment_complete_epoch is not None
):
session_data["first_chunk_before_first_media_complete"] = (
first_chunk_finish_epoch < first_media_segment_complete_epoch
)
if (first_chunk_finish_epoch is not None and first_media_segment_complete_epoch is not None):
session_data["first_chunk_before_first_media_complete"] = (first_chunk_finish_epoch
< first_media_segment_complete_epoch)
session_data["close_ts_utc"] = utc_now_iso()
session_data["duration_ms"] = (
time.monotonic() - session_start_monotonic
) * 1000.0
session_data["duration_ms"] = (time.monotonic() - session_start_monotonic) * 1000.0
return session_data
@@ -412,14 +364,11 @@ async def run_worker_sessions(
config: dict[str, Any],
) -> list[dict[str, Any]]:
tasks = [
asyncio.create_task(
run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)
)
for idx in range(session_count)
asyncio.create_task(run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)) for idx in range(session_count)
]
if not tasks:
return []
@@ -437,29 +386,23 @@ def worker_entry(
try:
ready_queue.put({"worker_id": worker_id, "status": "ready"})
start_event.wait()
sessions = asyncio.run(
run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
)
)
result_queue.put(
{
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
}
)
sessions = asyncio.run(run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
))
result_queue.put({
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
})
except Exception as exc:
result_queue.put(
{
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
}
)
result_queue.put({
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
})
def build_summary(
@@ -517,33 +460,22 @@ def build_summary(
if len(all_chunk_finish_epochs) >= 2 and total_chunk_bytes > 0:
duration_s = max(all_chunk_finish_epochs) - min(all_chunk_finish_epochs)
if duration_s > 0:
global_goodput_mbps = (
total_chunk_bytes * 8.0 / duration_s / 1_000_000.0
)
global_goodput_mbps = (total_chunk_bytes * 8.0 / duration_s / 1_000_000.0)
bucket_throughputs_mbps = [
(bytes_count * 8.0) / 1_000_000.0
for _, bytes_count in sorted(bucket_bytes.items())
]
bucket_throughputs_mbps = [(bytes_count * 8.0) / 1_000_000.0 for _, bytes_count in sorted(bucket_bytes.items())]
bucket_stats = summarize_series(bucket_throughputs_mbps)
chunk_gap_threshold_breaches = [
value for value in chunk_gaps if value >= chunk_gap_threshold_ms
]
chunk_gap_threshold_breaches = [value for value in chunk_gaps if value >= chunk_gap_threshold_ms]
non_success = len(sessions) - status_counts.get("success", 0)
fail_reasons: list[str] = []
if non_success > 0:
fail_reasons.append(
f"{non_success} session(s) did not complete successfully."
)
fail_reasons.append(f"{non_success} session(s) did not complete successfully.")
if not chunk_gaps:
fail_reasons.append("No chunk gap data collected.")
if chunk_gap_threshold_breaches:
fail_reasons.append(
f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms."
)
fail_reasons.append(f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms.")
passed = len(fail_reasons) == 0
progressive_ratio = None
@@ -554,20 +486,18 @@ def build_summary(
"passed": passed,
"fail_reasons": fail_reasons,
"sessions": {
"total": len(sessions),
"success": status_counts.get("success", 0),
"failed": status_counts.get("failed", 0),
"timeout": status_counts.get("timeout", 0),
"protocol_error": status_counts.get("protocol_error", 0),
"other": (
len(sessions)
- (
status_counts.get("success", 0)
+ status_counts.get("failed", 0)
+ status_counts.get("timeout", 0)
+ status_counts.get("protocol_error", 0)
)
),
"total":
len(sessions),
"success":
status_counts.get("success", 0),
"failed":
status_counts.get("failed", 0),
"timeout":
status_counts.get("timeout", 0),
"protocol_error":
status_counts.get("protocol_error", 0),
"other": (len(sessions) - (status_counts.get("success", 0) + status_counts.get("failed", 0) +
status_counts.get("timeout", 0) + status_counts.get("protocol_error", 0))),
},
"chunk_gap_ms": {
**chunk_gap_stats,
@@ -606,51 +536,39 @@ def print_summary(
bucket_bw = bandwidth["bucketed_1s"]
print("=== LTX2 Realtime Stress Test Summary ===")
print(
"Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}"
)
print(
"Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}"
)
print(
"Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}"
)
print(
"Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}"
)
print("Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}")
print("Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}")
print("Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}")
print("Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}")
ratio = progressive["ratio"]
ratio_text = "n/a" if ratio is None else f"{ratio * 100:.2f}%"
print(
"Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})"
)
print(
"Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}"
)
print("Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})")
print("Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}")
print(f"VERDICT: {'PASS' if summary['passed'] else 'FAIL'}")
if summary["fail_reasons"]:
print("Fail reasons:")
@@ -670,10 +588,8 @@ def distribute_sessions(total_clients: int, process_count: int) -> list[int]:
def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if websockets is None:
raise RuntimeError(
"Missing dependency: websockets. Install it before running this "
"stress test."
)
raise RuntimeError("Missing dependency: websockets. Install it before running this "
"stress test.")
preset_file = Path(args.preset_file).expanduser().resolve()
selected_preset_id, curated_prompts, total_prompt_count = load_curated_prompts(
@@ -735,13 +651,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
start_event.set()
result_deadline = (
time.monotonic()
+ args.connect_timeout_s
+ args.session_timeout_s
+ args.post_complete_wait_s
+ 180.0
)
result_deadline = (time.monotonic() + args.connect_timeout_s + args.session_timeout_s +
args.post_complete_wait_s + 180.0)
worker_results: list[dict[str, Any]] = []
while len(worker_results) < len(processes):
timeout_s = max(0.1, result_deadline - time.monotonic())
@@ -765,24 +676,20 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if result.get("status") == "ok":
sessions.extend(result.get("sessions", []))
else:
worker_errors.append(
{
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
}
)
worker_errors.append({
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
})
received_workers = {result.get("worker_id") for result in worker_results}
expected_workers = set(range(len(processes)))
missing_workers = sorted(expected_workers - received_workers)
for worker_id in missing_workers:
worker_errors.append(
{
"worker_id": worker_id,
"error": "No worker result received.",
}
)
worker_errors.append({
"worker_id": worker_id,
"error": "No worker result received.",
})
run_end_epoch = time.time()
run_end_iso = iso_from_epoch(run_end_epoch)
@@ -795,9 +702,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if worker_errors:
summary["passed"] = False
summary["fail_reasons"] = list(summary["fail_reasons"]) + [
f"{len(worker_errors)} worker error(s) occurred."
]
summary["fail_reasons"] = list(
summary["fail_reasons"]) + [f"{len(worker_errors)} worker error(s) occurred."]
output_payload = {
"run_info": {
@@ -833,9 +739,7 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Multiprocess realtime stress test for LTX2 streaming.",
)
parser = argparse.ArgumentParser(description="Multiprocess realtime stress test for LTX2 streaming.", )
parser.add_argument(
"-u",
"--url",
@@ -47,13 +47,11 @@ def test_persist_session_init_image_returns_none_when_missing_data():
def test_persist_session_init_image_rejects_unsupported_mime():
with pytest.raises(ValueError, match="PNG, JPEG, or WebP"):
persist_session_init_image(
{
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
}
)
persist_session_init_image({
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
})
def test_persist_session_init_image_rejects_large_payload(monkeypatch):
@@ -66,10 +64,8 @@ def test_persist_session_init_image_rejects_large_payload(monkeypatch):
monkeypatch.setattr(base64, "b64decode", fake_b64decode)
with pytest.raises(ValueError, match="15 MB or smaller"):
persist_session_init_image(
{
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
}
)
persist_session_init_image({
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
})
File diff suppressed because it is too large Load Diff
+8 -8
View File
@@ -70,7 +70,7 @@
<mxCell id="dispatcher" value="command dispatcher&#xa;&#xa;gpu_worker_process() branches on&#xa;CommandType; asserts payload type&#xa;&#xa;INIT / WARMUP / RELOAD_MODEL&#xa;USER_JOIN / USER_STEP / USER_LEAVE&#xa;SHUTDOWN" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe6cc;strokeColor=#d79b00;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="120" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;video_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;ltx2_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="460" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="stream_av" value="stream_fmp4()&#xa;av_streaming.py:121&#xa;&#xa;trims overlap, pipes to ffmpeg,&#xa;publishes StreamInit / StreamChunk /&#xa;StreamComplete via callback" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#b1d8d7;strokeColor=#23445d;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
@@ -79,13 +79,13 @@
<mxCell id="Ot8BU52QTIb4EhyRSe7I-2" value="" style="edgeStyle=none;html=1;" parent="1" source="generator" target="Ot8BU52QTIb4EhyRSe7I-1" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;video_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;ltx2_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="460" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="ffmpeg" value="ffmpeg subprocess&#xa;&#xa;libx264 / *_nvenc&#xa;fragmented mp4" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="800" y="1300" width="260" height="100" as="geometry"/>
</mxCell>
<mxCell id="caches" value="ContinuationState&#xa;video_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="caches" value="ContinuationState&#xa;ltx2_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="120" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="e_cp" value="acquire" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="client" target="pool" edge="1">
@@ -207,7 +207,7 @@
</Array>
</mxGeometry>
</mxCell>
<mxCell id="e_dsg" value="generator.generate_video()" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#9673a6;endArrow=classic;fontSize=10;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="do_step" target="generator" edge="1">
<mxCell id="e_dsg" value="generator.generate()" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#9673a6;endArrow=classic;fontSize=10;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="do_step" target="generator" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="e_dscache" value="read / write" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#d6b656;endArrow=classic;startArrow=classic;fontSize=10;exitX=0;exitY=0.8;exitDx=0;exitDy=0;entryX=1;entryY=0.2;entryDx=0;entryDy=0;" parent="1" source="do_step" target="caches" edge="1">
@@ -250,7 +250,7 @@
<mxPoint x="690" y="880"/>
</Array>
</mxCell>
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender video_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender ltx2_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="39" y="-200" width="270" height="380" as="geometry"/>
</mxCell>
<mxCell id="Ot8BU52QTIb4EhyRSe7I-1" value="FastVideo video_generator" style="whiteSpace=wrap;html=1;fontSize=11;fillColor=#e1d5e7;strokeColor=#9673a6;rounded=1;" parent="1" vertex="1">
@@ -389,10 +389,10 @@
<mxCell id="cw2" value="from fastvideo.entrypoints.video_generator import VideoGenerator&#xa;from fastvideo.models.dits.ltx2 import DEFAULT_LTX2_AUDIO_*&#xa;&#xa;** Dreamverse reaches into fastvideo internals here **" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="695" width="550" height="60" as="geometry"/>
</mxCell>
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (video_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (ltx2_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="765" width="550" height="95" as="geometry"/>
</mxCell>
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (video_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (ltx2_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="870" width="550" height="55" as="geometry"/>
</mxCell>
<mxCell id="cw5" value="enter main worker loop → waits for JOIN_USER / USER_STEP / LEAVE" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
@@ -534,7 +534,7 @@
<mxPoint x="1040" y="1610" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (video_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (ltx2_generation.py:380)&#xa; → generator.generate()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="955" y="1640" width="180" height="70" as="geometry"/>
</mxCell>
<mxCell id="dm11" value="10b. resp_q.put(MediaInit / MediaChunk / MediaComplete / StepComplete)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 85 KiB

After

Width:  |  Height:  |  Size: 85 KiB

+9 -15
View File
@@ -8,12 +8,10 @@ import modal
IMAGE = os.environ.get("DREAMVERSE_IMAGE")
if not IMAGE:
raise RuntimeError(
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag."
)
raise RuntimeError("DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag.")
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
# ENTRYPOINT (``docker/docker_entrypoint.sh``). That entrypoint normally
@@ -65,14 +63,10 @@ def serve():
# ``or ""`` collapses ``None`` (unset) into an empty string, ``.strip()``
# collapses whitespace-only values (e.g. ``" "``) — both should be
# treated as missing.
missing = [
k for k in _REQUIRED_SECRET_KEYS
if not (os.environ.get(k) or "").strip()
]
missing = [k for k in _REQUIRED_SECRET_KEYS if not (os.environ.get(k) or "").strip()]
if missing:
raise RuntimeError(
"dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
raise RuntimeError("dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
subprocess.Popen(["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"])
@@ -64,8 +64,8 @@ generator:
# internal: pipeline_config.dit_config.quant_config = FP4Config()
# set in gpu_pool.py:280 (via the legacy in-place mutation). The
# public typed surface resolves "NVFP4" to NVFP4Config() and pins
# it on dit_config in FastVideoArgs.__post_init__. Comment this
# block out on hosts without flashinfer / NVFP4 hardware.
# it on dit_config when resolution materializes the PipelineConfig.
# Comment this block out on hosts without flashinfer / NVFP4 hardware.
quantization:
transformer_quant: NVFP4
@@ -67,7 +67,7 @@ test.describe('preset prompt generation', () => {
// on a B200 plus encode/transfer time. The "Continuation flipped
// to Generating + Leave button rendered" pair above is the proof
// the integration works: FE → /readyz → /curated-presets → WS
// /ws → BE → GPU pool → VideoGenerator.generate_video, all green.
// /ws → BE → GPU pool → VideoGenerator.generate, all green.
const video = page.locator('video').first();
await expect(video).toHaveCount(1);
});
+20 -4
View File
@@ -9,6 +9,7 @@ from __future__ import annotations
import contextlib
import logging
import json
import sqlite3
import threading
from pathlib import Path
@@ -46,8 +47,13 @@ DEFAULT_SETTINGS: dict[str, Any] = {
def _sqlite_row_get(row: sqlite3.Row, key: str, default: Any) -> Any:
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10)."""
return row[key] if key in row else default # noqa: SIM401
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10).
NOTE: `key in row` tests Row *values*, not column names, so the membership
check has to go through .keys() -- otherwise every lookup falls back to the
default and jobs restored from the database lose their stored fields.
"""
return row[key] if key in row.keys() else default # noqa: SIM401, SIM118
def _get_db_path(data_dir: Path) -> Path:
@@ -83,6 +89,9 @@ def _migrate_db(conn: sqlite3.Connection) -> None:
_add_column_if_missing(conn, "jobs", "fps", "INTEGER", "24")
_add_column_if_missing(conn, "jobs", "workload_type", "TEXT", "'t2v'")
_add_column_if_missing(conn, "jobs", "image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "name", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "last_image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "references_json", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "job_type", "TEXT", "'inference'")
_add_column_if_missing(conn, "jobs", "data_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "max_train_steps", "INTEGER", "1000")
@@ -242,7 +251,8 @@ class Database:
self._execute(
"""
INSERT INTO jobs (
id, model_id, prompt, workload_type, image_path, job_type, status,
id, model_id, name, prompt, workload_type, image_path,
last_image_path, references_json, job_type, status,
created_at, started_at, finished_at, error, output_path, log_file_path,
num_inference_steps, num_frames, height, width, guidance_scale,
guidance_rescale, fps, seed, num_gpus, dit_cpu_offload,
@@ -254,14 +264,17 @@ class Database:
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
real_score_guidance_scale,
generator_update_interval, real_score_model_path, fake_score_model_path
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
job["id"],
job["model_id"],
job.get("name", ""),
job["prompt"],
job.get("workload_type", "t2v"),
job.get("image_path", ""),
job.get("last_image_path", ""),
json.dumps(job.get("references") or []),
job.get("job_type", "inference"),
job["status"],
job["created_at"],
@@ -540,9 +553,12 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
result = {
"id": row["id"],
"model_id": row["model_id"],
"name": _sqlite_row_get(row, "name", "") or "",
"prompt": row["prompt"],
"workload_type": _sqlite_row_get(row, "workload_type", "t2v"),
"image_path": _sqlite_row_get(row, "image_path", "") or "",
"last_image_path": _sqlite_row_get(row, "last_image_path", "") or "",
"references": _sqlite_row_get(row, "references_json", "") or "",
"job_type": _sqlite_row_get(row, "job_type", "inference"),
"status": row["status"],
"created_at": row["created_at"],
@@ -0,0 +1,63 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
test.describe('create job interactions', () => {
skipWithoutMock();
for (const jobType of ['inference', 'finetuning', 'distillation']) {
test(`${jobType} remains interactive after repeated dialog dismissals`, async ({ page }) => {
await page.goto(`/${jobType}`);
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
const dialog = page.getByRole('dialog');
// Exercise both dismissal paths and reopen without reloading the page.
for (const closeWithEscape of [false, true]) {
await trigger.click();
await page.getByRole('menuitem').first().click();
await expect(dialog).toBeVisible();
if (closeWithEscape) {
await page.keyboard.press('Escape');
} else {
await dialog.getByRole('button', { name: 'Close', exact: true }).click();
}
await expect(dialog).toBeHidden();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
await expect(trigger).toBeFocused();
}
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
await expect(page).toHaveURL(/\/datasets$/);
});
}
test('preserves keyboard menu dismissal and dialog focus trapping', async ({ page }) => {
await page.goto('/inference');
const trigger = page.getByRole('button', { name: 'Create Job', exact: true });
await trigger.focus();
await page.keyboard.press('Enter');
const firstItem = page.getByRole('menuitem').first();
await expect(firstItem).toBeFocused();
await page.keyboard.press('Escape');
await expect(page.getByRole('menu')).toBeHidden();
await expect(trigger).toBeFocused();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
await page.keyboard.press('Enter');
await expect(firstItem).toBeFocused();
await page.keyboard.press('Enter');
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
// Shift+Tab from the first field wraps to Close, then Tab wraps back.
await page.keyboard.press('Shift+Tab');
await expect(dialog.getByRole('button', { name: 'Close', exact: true })).toBeFocused();
await page.keyboard.press('Tab');
await expect(dialog.getByLabel('Name (optional)')).toBeFocused();
await page.keyboard.press('Escape');
await expect(dialog).toBeHidden();
await expect(trigger).toBeFocused();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
});
});
+21 -6
View File
@@ -1,6 +1,6 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
import { API_BASE, skipWithoutMock } from './helpers';
/**
* Create-job flow: open the Create Job modal on /inference, fill the prompt
@@ -10,13 +10,13 @@ import { skipWithoutMock } from './helpers';
test.describe('create inference job', () => {
skipWithoutMock();
test('creates a T2V job and shows it in the queue', async ({ page }) => {
test('creates a T2V job and starts it without refreshing', async ({ page, request }) => {
await request.put(`${API_BASE}/settings`, { data: { autoStartJob: false } });
await page.goto('/inference');
// The "Create Job" button reveals a workload menu on hover; wait for the
// T2V item to become visible before clicking so the CSS hover transition
// can't race the click.
await page.getByRole('button', { name: /create job/i }).hover();
// The trigger opens a real menu on click, so this path works for touch,
// mouse, and keyboard users.
await page.getByRole('button', { name: /create job/i }).click();
const t2vItem = page.getByRole('menuitem', { name: /T2V/i });
await expect(t2vItem).toBeVisible();
await t2vItem.click();
@@ -39,5 +39,20 @@ test.describe('create inference job', () => {
// Modal closes and the queue refreshes with the newly created job.
await expect(dialog).toBeHidden();
await expect(page.getByText(prompt)).toBeVisible();
await expect(page.locator('body')).toHaveCSS('pointer-events', 'auto');
const card = page.getByRole('article').filter({ hasText: prompt });
await expect(card.getByText('pending', { exact: true })).toBeVisible();
const started = page.waitForResponse((response) =>
response.url().startsWith(`${API_BASE}/jobs/`) &&
response.url().endsWith('/start') &&
response.request().method() === 'POST',
);
await card.getByRole('button', { name: 'Start', exact: true }).click();
expect((await started).ok()).toBe(true);
await expect(card.getByText('running', { exact: true })).toBeVisible();
await page.getByRole('link', { name: 'Datasets', exact: true }).click();
await expect(page).toHaveURL(/\/datasets$/);
});
});
+10 -7
View File
@@ -4,7 +4,7 @@ import { API_BASE, skipWithoutMock } from './helpers';
/**
* Gallery page: the seeded completed inference job surfaces as a media tile
* (an <article> wrapping a <video>) captioned with its prompt.
* with playback controls or an explicit media-error fallback.
*/
test.describe('gallery', () => {
skipWithoutMock();
@@ -30,12 +30,15 @@ test.describe('gallery', () => {
page.getByRole('heading', { level: 1, name: 'Gallery' }),
).toBeVisible();
// The completed job renders as an <article> containing a <video> tile.
const tile = page
.locator('article')
.filter({ has: page.locator('video') });
await expect(tile.first()).toBeVisible();
const tile = page.locator('article').filter({ hasText: completed!.prompt });
await expect(tile).toBeVisible();
await expect(
tile.locator('video').or(tile.getByText('Preview unavailable')),
).toBeVisible();
await expect(page.getByText(completed!.prompt)).toBeVisible();
const video = tile.locator('video');
if (await video.isVisible()) {
await expect(video).toHaveAttribute('controls', '');
}
});
});
+68
View File
@@ -42,6 +42,74 @@ test.describe('app shell', () => {
await expect(
page.getByRole('heading', { level: 1, name: section.title }),
).toBeVisible();
await expect(page.getByRole('main')).toHaveCount(1);
}
});
test('keeps navigation and content usable at responsive breakpoints', async ({
page,
}) => {
for (const width of [320, 375, 414, 768]) {
await page.setViewportSize({ width, height: 800 });
await page.goto('/inference');
const main = page.getByRole('main');
await expect(main).toBeVisible();
await expect(
page.getByRole('button', { name: /Create Job/i }),
).toBeVisible();
const initialBox = await main.boundingBox();
expect(initialBox?.x).toBe(width < 768 ? 0 : 220);
expect(initialBox?.width).toBe(width < 768 ? width : width - 220);
const navigation = page.getByRole('navigation', {
name: 'Primary navigation',
});
if (width < 768) {
await expect(
page.getByRole('button', { name: 'Open navigation' }),
).toBeVisible();
await page.getByRole('button', { name: 'Open navigation' }).click();
}
await expect(navigation).toBeVisible();
await navigation.getByRole('link', { name: 'Datasets' }).click();
await expect(page).toHaveURL(/\/datasets$/);
expect(
await page.evaluate(
() => document.documentElement.scrollWidth <= window.innerWidth,
),
).toBe(true);
}
});
test('uses full-width detail drawers on mobile', async ({ page }) => {
await page.setViewportSize({ width: 320, height: 800 });
await page.goto('/inference');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const jobDrawer = page.getByRole('dialog', { name: 'Job details' });
await expect(jobDrawer).toBeVisible();
expect(await jobDrawer.boundingBox()).toMatchObject({ x: 0, width: 320 });
await jobDrawer.getByRole('button', { name: 'Close' }).click();
await page.goto('/datasets');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const datasetDrawer = page.getByRole('dialog', {
name: /dataset details$/,
});
await expect(datasetDrawer).toBeVisible();
expect(await datasetDrawer.boundingBox()).toMatchObject({
x: 0,
width: 320,
});
});
});
+284 -83
View File
@@ -10,7 +10,9 @@ from __future__ import annotations
import atexit
import collections
import contextlib
import copy
import enum
import json
import logging
import logging.handlers
import multiprocessing as mp
@@ -123,10 +125,13 @@ class LogBufferHandler(logging.Handler):
class Job:
id: str
model_id: str
prompt: str
name: str = ""
prompt: str = ""
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] = field(default_factory=list)
status: JobStatus = JobStatus.PENDING
created_at: float = field(default_factory=time.time)
started_at: float | None = None
@@ -145,6 +150,7 @@ class Job:
negative_prompt: str = ""
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
@@ -180,10 +186,13 @@ class Job:
return {
"id": self.id,
"model_id": self.model_id,
"name": self.name,
"prompt": self.prompt,
"workload_type": self.workload_type,
"job_type": self.job_type,
"image_path": self.image_path,
"last_image_path": self.last_image_path,
"references": self.references,
"status": self.status.value,
"created_at": self.created_at,
"started_at": self.started_at,
@@ -202,6 +211,7 @@ class Job:
"negative_prompt": self.negative_prompt,
"num_gpus": self.num_gpus,
"dit_cpu_offload": self.dit_cpu_offload,
"dit_layerwise_offload": self.dit_layerwise_offload,
"text_encoder_cpu_offload": self.text_encoder_cpu_offload,
"vae_cpu_offload": self.vae_cpu_offload,
"image_encoder_cpu_offload": self.image_encoder_cpu_offload,
@@ -232,6 +242,75 @@ class Job:
}
MINIMAX_H3_REF2VA_PIPELINE = "MiniMaxH3Ref2VAModularPipeline"
def _build_h3_references(raw: list[dict[str, Any]]) -> list[Any]:
"""Turn the API's reference dicts into MiniMaxH3Reference objects.
Imported lazily so the API server starts without pulling in fastvideo.
"""
from fastvideo.pipelines.basic.minimax_h3 import MiniMaxH3Reference
built = []
for i, ref in enumerate(raw):
source = (ref or {}).get("source")
if not source:
raise ValueError(f"reference {i} has no source")
if not os.path.isfile(source):
raise ValueError(f"reference {i} source not found: {source}")
kwargs: dict[str, Any] = {
"source": source,
"media_type": (ref.get("media_type") or "image"),
}
for opt in ("soundtrack", "fps", "sample_rate"):
if ref.get(opt) not in (None, ""):
kwargs[opt] = ref[opt]
built.append(MiniMaxH3Reference(**kwargs))
return built
JOB_LOG_FILENAME = "out.log"
def _job_log_path(output_dir: str, job_id: str) -> str:
"""Each job's log lives beside its outputs: <output_dir>/<job_id>/out.log."""
return os.path.join(output_dir, job_id, JOB_LOG_FILENAME)
def _decode_references(value: Any) -> list[dict[str, Any]]:
"""Reference lists round-trip through the DB as JSON text."""
if not value:
return []
if isinstance(value, list):
return list(value)
try:
decoded = json.loads(value)
except (TypeError, ValueError):
logger.warning("Could not decode stored references: %r", value)
return []
return list(decoded) if isinstance(decoded, list) else []
def _generator_is_alive(generator: Any) -> bool:
"""True if the generator's worker processes are all still running.
A cached VideoGenerator holds a MultiprocExecutor whose workers are separate
processes; nothing notices when they exit. Probing `proc.is_alive()` is what
the executor itself uses during shutdown. Anything unexpected in the object
graph is treated as alive so a probe failure can never wedge the cache.
"""
executor = getattr(generator, "executor", None)
workers = getattr(executor, "workers", None)
if not workers:
return True
try:
return all(w.proc.is_alive() for w in workers)
except Exception:
logger.debug("Worker liveness probe failed", exc_info=True)
return True
class JobRunner:
"""Manages video generation jobs, their execution, and generator caching."""
@@ -276,7 +355,7 @@ class JobRunner:
"""Populate job's log buffer from its log file if it exists."""
path = job.log_file_path
if not path:
path = os.path.join(self.log_dir, f"{job.id}.log")
path = _job_log_path(self.output_dir, job.id)
if not os.path.isfile(path):
return
try:
@@ -314,10 +393,13 @@ class JobRunner:
job = Job(
id=row["id"],
model_id=row["model_id"],
name=row.get("name", "") or "",
prompt=row["prompt"],
workload_type=row.get("workload_type", "t2v"),
job_type=row.get("job_type", "inference"),
image_path=row.get("image_path", "") or "",
last_image_path=row.get("last_image_path", "") or "",
references=_decode_references(row.get("references")),
data_path=row.get("data_path", "") or "",
max_train_steps=row.get("max_train_steps", 1000),
train_batch_size=row.get("train_batch_size", 1),
@@ -350,6 +432,7 @@ class JobRunner:
negative_prompt=row.get("negative_prompt", "") or "",
num_gpus=row.get("num_gpus", 1),
dit_cpu_offload=row.get("dit_cpu_offload", False),
dit_layerwise_offload=row.get("dit_layerwise_offload", False),
text_encoder_cpu_offload=row.get("text_encoder_cpu_offload", False),
vae_cpu_offload=row.get("vae_cpu_offload", False),
image_encoder_cpu_offload=row.get("image_encoder_cpu_offload", False),
@@ -394,9 +477,12 @@ class JobRunner:
job_id: str,
model_id: str,
prompt: str,
name: str = "",
workload_type: str = "t2v",
job_type: str = "inference",
image_path: str = "",
last_image_path: str = "",
references: list[dict[str, Any]] | None = None,
data_path: str = "",
max_train_steps: int = 1000,
train_batch_size: int = 1,
@@ -422,6 +508,7 @@ class JobRunner:
num_gpus: int = 1,
negative_prompt: str = "",
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -435,10 +522,13 @@ class JobRunner:
job = Job(
id=job_id,
model_id=model_id,
name=(name or "").strip(),
prompt=prompt.strip(),
workload_type=workload_type or "t2v",
job_type=job_type or "inference",
image_path=image_path or "",
last_image_path=last_image_path or "",
references=list(references or []),
data_path=data_path or "",
max_train_steps=max_train_steps,
train_batch_size=train_batch_size,
@@ -464,6 +554,7 @@ class JobRunner:
negative_prompt=negative_prompt or "",
num_gpus=num_gpus,
dit_cpu_offload=dit_cpu_offload,
dit_layerwise_offload=dit_layerwise_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
image_encoder_cpu_offload=image_encoder_cpu_offload,
@@ -521,6 +612,84 @@ class JobRunner:
logger.info("Deleted job %s", job.id)
return True
CONFIG_FIELDS: tuple[str, ...] = (
"model_id",
"name",
"prompt",
"workload_type",
"job_type",
"image_path",
"last_image_path",
"references",
"negative_prompt",
"num_inference_steps",
"num_frames",
"height",
"width",
"guidance_scale",
"guidance_rescale",
"fps",
"seed",
"num_gpus",
"dit_cpu_offload",
"dit_layerwise_offload",
"text_encoder_cpu_offload",
"vae_cpu_offload",
"image_encoder_cpu_offload",
"use_fsdp_inference",
"enable_torch_compile",
"vsa_sparsity",
"tp_size",
"sp_size",
"data_path",
"max_train_steps",
"train_batch_size",
"learning_rate",
"num_latent_t",
"validation_dataset_file",
"lora_rank",
"dmd_use_vsa",
"dmd_vsa_sparsity",
"dmd_denoising_steps",
"real_score_guidance_scale",
"generator_update_interval",
"real_score_model_path",
"fake_score_model_path",
)
def duplicate_job(self, job_id: str, new_job_id: str) -> Job:
"""Create a new pending job with an existing job's configuration.
Runtime state (status, timings, logs, outputs) is not carried over.
"""
with self._jobs_lock:
source = self._jobs.get(job_id)
if source is None:
raise ValueError(f"Job {job_id} not found")
config = {f: copy.deepcopy(getattr(source, f)) for f in self.CONFIG_FIELDS}
return self.create_job(job_id=new_job_id, **config)
#: Editable exactly when startable: the same set start_job() accepts.
EDITABLE_STATUSES = (JobStatus.PENDING, JobStatus.FAILED, JobStatus.STOPPED)
def update_job_config(self, job_id: str, updates: dict[str, Any]) -> Job:
"""Edit the configuration of a job that has not produced a result."""
with self._jobs_lock:
job = self._jobs.get(job_id)
if job is None:
raise ValueError(f"Job {job_id} not found")
if job.status not in self.EDITABLE_STATUSES:
allowed = ", ".join(s.value for s in self.EDITABLE_STATUSES)
raise ValueError(f"Job is {job.status.value}; only {allowed} jobs can be edited. "
"Duplicate it instead.")
unknown = set(updates) - set(self.CONFIG_FIELDS)
if unknown:
raise ValueError(f"Not editable: {', '.join(sorted(unknown))}")
for field_name, value in updates.items():
setattr(job, field_name, value)
self._save_job(job)
return job
def start_job(self, job_id: str) -> Job:
"""Start (or restart) a pending / stopped / failed job.
@@ -623,6 +792,8 @@ class JobRunner:
workload_type: str,
num_gpus: int,
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
override_pipeline_cls_name: str | None = None,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -638,6 +809,10 @@ class JobRunner:
workload_type,
num_gpus,
dit_cpu_offload,
dit_layerwise_offload,
# Ref2VA loads different DiT weights (transformer_ref), so the
# override must key the cache or a t2v/i2v generator gets reused.
override_pipeline_cls_name,
text_encoder_cpu_offload,
vae_cpu_offload,
image_encoder_cpu_offload,
@@ -650,8 +825,21 @@ class JobRunner:
# Generators are cached by model_id and configuration parameters
with self._generators_lock:
if cache_key in self._generators:
return self._generators[cache_key]
cached = self._generators.get(cache_key)
if cached is not None:
if _generator_is_alive(cached):
return cached
# Workers can exit while a generator sits idle in the cache;
# reusing it fails every later job with the same config.
logger.warning(
"Cached generator for %s has dead workers; reloading.",
model_id,
)
self._generators.pop(cache_key, None)
try:
cached.shutdown()
except Exception:
logger.debug("Shutdown of the dead generator failed", exc_info=True)
# Import lazily so starting the server is fast even without a GPU.
from fastvideo import VideoGenerator
@@ -674,18 +862,40 @@ class JobRunner:
sp_size,
)
gen = VideoGenerator.from_pretrained(
model_id,
workload_type=workload_type,
dit_cpu_offload=dit_cpu_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
image_encoder_cpu_offload=image_encoder_cpu_offload,
use_fsdp_inference=use_fsdp_inference,
enable_torch_compile=enable_torch_compile,
VSA_sparsity=vsa_sparsity,
tp_size=tp_size,
sp_size=sp_size,
gen = VideoGenerator.from_config(
{
"model_path": model_id,
"engine": {
"num_gpus": num_gpus,
"parallelism": {
"tp_size": tp_size,
"sp_size": sp_size,
},
"offload": {
"dit": dit_cpu_offload,
"dit_layerwise": dit_layerwise_offload,
"text_encoder": text_encoder_cpu_offload,
"image_encoder": image_encoder_cpu_offload,
"vae": vae_cpu_offload,
},
"compile": {
"enabled": enable_torch_compile
},
"attention": {
"vsa_sparsity": vsa_sparsity
},
"use_fsdp_inference": use_fsdp_inference,
},
"pipeline": {
"workload_type":
workload_type,
**({
"components": {
"override_pipeline_cls_name": override_pipeline_cls_name
}
} if override_pipeline_cls_name else {}),
},
},
log_queue=log_queue,
)
@@ -706,10 +916,9 @@ class JobRunner:
def _run_training_job(self, job: Job):
"""Run a finetuning, distillation, or LoRA job via subprocess."""
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
job_output_dir = os.path.join(self.output_dir, job.id)
os.makedirs(job_output_dir, exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
if not job.data_path or not os.path.isdir(job.data_path):
job.status = JobStatus.FAILED
@@ -827,8 +1036,8 @@ class JobRunner:
def _run_inference_job(self, job: Job):
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
os.makedirs(os.path.join(self.output_dir, job.id), exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
# Add file handler to persist logs
file_handler = logging.FileHandler(job.log_file_path, mode='w', encoding='utf-8')
@@ -875,77 +1084,69 @@ class JobRunner:
buf.phase = "loading model"
logger.info("Loading model...")
# Run generator creation in a background thread so we
# can poll _stop_event while the (potentially slow)
# model download / load is in progress.
_gen_result: list[Any] = []
_gen_error: list[BaseException] = []
# The generator MUST be created on this thread: building it spawns
# the executor's worker processes, and they are torn down if the
# creating thread exits. Running it in a helper thread (to poll
# _stop_event during load) made every collective_rpc fail with
# ConnectionResetError.
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning("Job %s stopped before model loading", job.id)
buf.phase = "stopped"
return
def _load_generator() -> None:
try:
gen = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
_gen_result.append(gen)
except BaseException as exc:
_gen_error.append(exc)
loader = threading.Thread(
target=_load_generator,
daemon=True,
generator = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
dit_layerwise_offload=job.dit_layerwise_offload,
override_pipeline_cls_name=(MINIMAX_H3_REF2VA_PIPELINE if job.references else None),
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
loader.start()
while loader.is_alive():
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning(
"Job %s stopped during model loading",
job.id,
)
buf.phase = "stopped"
return
loader.join(timeout=0.5)
if _gen_error:
raise _gen_error[0]
generator = _gen_result[0]
buf.phase = "generating"
logger.info("Starting generation for job %s (model=%s)", job.id, job.model_id)
gen_kwargs: dict[str, Any] = {
# Without a name FastVideo derives the filename from the prompt.
safe_name = re.sub(r'[\\/:*?"<>|]+', "", job.name).strip().strip(".")
output_target = (os.path.join(job_output_dir, f"{safe_name[:80]}.mp4") if safe_name else job_output_dir)
request: dict[str, Any] = {
"prompt": job.prompt,
"output_path": job_output_dir,
"save_video": True,
"num_inference_steps": job.num_inference_steps,
"num_frames": job.num_frames,
"height": job.height,
"width": job.width,
"guidance_scale": job.guidance_scale,
"guidance_rescale": job.guidance_rescale,
"fps": job.fps,
"seed": job.seed,
"negative_prompt": job.negative_prompt or "",
"log_queue": log_queue,
"sampling": {
"num_inference_steps": job.num_inference_steps,
"num_frames": job.num_frames,
"height": job.height,
"width": job.width,
"guidance_scale": job.guidance_scale,
"guidance_rescale": job.guidance_rescale,
"fps": job.fps,
"seed": job.seed,
},
"output": {
"output_path": output_target,
"save_video": True,
},
}
if job.image_path:
gen_kwargs["image_path"] = job.image_path
generator.generate_video(**gen_kwargs)
request.setdefault("inputs", {})["image_path"] = job.image_path
if job.references:
request.setdefault("inputs", {})["references"] = _build_h3_references(job.references)
if job.last_image_path:
# _prepare_fl2va requires a PIL image, not a path.
from PIL import Image as _PILImage
request.setdefault("inputs", {})["last_image"] = _PILImage.open(job.last_image_path)
generator.generate(request, log_queue=log_queue)
buf.phase = "saving"
logger.info("Generation completed, searching for output file...")
@@ -977,7 +1178,7 @@ class JobRunner:
except Exception as exception:
error_msg = str(exception)
logger.error("Critical error in job thread: %s", error_msg)
logger.exception("Critical error in job thread: %s", error_msg)
job.status = JobStatus.FAILED
job.error = f"Critical error ({type(exception).__name__}): {error_msg}"
job.finished_at = time.time()
@@ -1,15 +1,20 @@
# SPDX-License-Identifier: Apache-2.0
"""Request model for creating a job."""
from typing import Any
from pydantic import BaseModel
class CreateJobRequest(BaseModel):
model_id: str
name: str = ""
prompt: str
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] | None = None
data_path: str = ""
max_train_steps: int = 1000
train_batch_size: int = 1
@@ -28,6 +33,7 @@ class CreateJobRequest(BaseModel):
seed: int = 1024
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
+1052 -225
View File
File diff suppressed because it is too large Load Diff
+1 -9
View File
@@ -17,19 +17,11 @@
"start:all": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\""
},
"dependencies": {
"@radix-ui/react-dialog": "^1.1.0",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.2.0",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-switch": "^1.1.0",
"@radix-ui/react-tabs": "^1.1.0",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"lucide-react": "^0.577.0",
"next": "15.5.18",
"radix-ui": "^1.6.7",
"react": "^19.1.0",
"react-dom": "^19.1.0",
"sonner": "^2.0.7",
+92 -2
View File
@@ -18,6 +18,7 @@ import argparse
import contextlib
import logging
import os
import re
import shutil
import signal
import time
@@ -116,7 +117,30 @@ def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _available_models
def _safe_upload_name(filename: str | None, ext: str) -> str:
"""A filesystem-safe version of the client's filename, keeping it readable.
Uploads live under a per-file uuid directory, so the basename does not have
to be unique -- only safe. Keeping the original name means the path stays
self-describing wherever it travels: the database, job logs, and payloads
copied back out to the API.
"""
stem = os.path.basename(filename or "").rsplit(".", 1)[0]
stem = re.sub(r"[^A-Za-z0-9._-]+", "_", stem).strip("._-")
return f"{stem[:80] or 'upload'}{ext}"
def _upload_destination(ext: str, filename: str | None) -> str:
"""<upload_dir>/<uuid4>/<safe original name><ext>"""
directory = os.path.join(upload_dir, uuid.uuid4().hex)
os.makedirs(directory, exist_ok=True)
return os.path.join(directory, _safe_upload_name(filename, ext))
ALLOWED_IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".webp", ".bmp"}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".mov", ".mkv", ".webm", ".avi"}
ALLOWED_AUDIO_EXTENSIONS = {".wav", ".mp3", ".flac", ".m4a", ".ogg"}
ALLOWED_MEDIA_EXTENSIONS = (ALLOWED_IMAGE_EXTENSIONS | ALLOWED_VIDEO_EXTENSIONS | ALLOWED_AUDIO_EXTENSIONS)
@app.post("/api/upload-image")
@@ -136,8 +160,7 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
f"{', '.join(ALLOWED_IMAGE_EXTENSIONS)}"),
)
os.makedirs(upload_dir, exist_ok=True)
unique_name = f"{uuid.uuid4().hex}{ext}"
dest_path = os.path.join(upload_dir, unique_name)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
@@ -150,6 +173,47 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
return {"path": os.path.abspath(dest_path)}
@app.post("/api/upload-media")
async def upload_media(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
"""Upload an image, video or audio file for Ref2VA references.
Returns the absolute path plus the media_type MiniMax-H3 expects, so the
caller does not have to re-derive it from the extension.
"""
global upload_dir # noqa: PLW0603
if not upload_dir:
raise HTTPException(
status_code=503,
detail="Upload directory not configured",
)
ext = Path(file.filename or "").suffix.lower()
if ext not in ALLOWED_MEDIA_EXTENSIONS:
raise HTTPException(
status_code=400,
detail=(f"Invalid file type. Allowed: "
f"{', '.join(sorted(ALLOWED_MEDIA_EXTENSIONS))}"),
)
if ext in ALLOWED_VIDEO_EXTENSIONS:
media_type = "video"
elif ext in ALLOWED_AUDIO_EXTENSIONS:
media_type = "audio"
else:
media_type = "image"
os.makedirs(upload_dir, exist_ok=True)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
f.write(contents)
except OSError as e:
raise HTTPException(
status_code=500,
detail=f"Failed to save upload: {e}",
) from e
return {"path": os.path.abspath(dest_path), "media_type": media_type}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
@@ -282,10 +346,13 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
job = job_runner.create_job(
job_id=str(uuid.uuid4()),
model_id=req.model_id,
name=req.name or "",
prompt=req.prompt,
workload_type=req.workload_type or "t2v",
job_type=job_type,
image_path=req.image_path or "",
last_image_path=req.last_image_path or "",
references=req.references or [],
data_path=data_path,
max_train_steps=req.max_train_steps,
train_batch_size=req.train_batch_size,
@@ -304,6 +371,7 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
seed=req.seed,
num_gpus=req.num_gpus,
dit_cpu_offload=req.dit_cpu_offload,
dit_layerwise_offload=req.dit_layerwise_offload,
text_encoder_cpu_offload=req.text_encoder_cpu_offload,
vae_cpu_offload=req.vae_cpu_offload,
image_encoder_cpu_offload=req.image_encoder_cpu_offload,
@@ -336,6 +404,28 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
return job.to_dict()
@app.post("/api/jobs/{job_id}/duplicate", status_code=201)
def duplicate_job(job_id: str) -> dict[str, Any]:
"""Create a new pending job with the same configuration as an existing one."""
try:
job = job_runner.duplicate_job(job_id, str(uuid.uuid4()))
except ValueError as e:
raise HTTPException(status_code=404, detail=str(e)) from e
return job.to_dict()
@app.patch("/api/jobs/{job_id}")
def update_job(job_id: str, updates: dict[str, Any]) -> dict[str, Any]:
"""Edit a pending job's configuration. Started jobs cannot be edited."""
try:
job = job_runner.update_job_config(job_id, updates)
except ValueError as e:
detail = str(e)
status = 404 if "not found" in detail else 400
raise HTTPException(status_code=status, detail=detail) from e
return job.to_dict()
@app.post("/api/jobs/{job_id}/start")
def start_job(job_id: str) -> dict[str, Any]:
"""Start (or restart) a pending / stopped / failed job."""
@@ -0,0 +1,64 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { getDatasets, type Dataset } from '@/lib/api';
import DatasetsPage from './page';
vi.mock('@/lib/api', () => ({
getDatasets: vi.fn(),
}));
vi.mock('@/components/datasets/AddDatasetButton', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/CreateDatasetModal', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/DatasetCard', () => ({
default: ({ dataset }: { dataset: Dataset }) => <div>{dataset.name}</div>,
}));
function renderPage() {
return render(
<HeaderActionsProvider>
<DatasetsPage />
</HeaderActionsProvider>,
);
}
describe('DatasetsPage', () => {
it('shows loading content before the initial request settles', async () => {
let resolveDatasets: (datasets: Dataset[]) => void = () => {};
vi.mocked(getDatasets).mockReturnValue(
new Promise<Dataset[]>((resolve) => {
resolveDatasets = resolve;
}),
);
renderPage();
expect(screen.getByLabelText('Loading datasets')).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
act(() => resolveDatasets([]));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
it('shows API failures separately from an empty list and retries', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getDatasets).mockRejectedValueOnce(new Error('network down'));
renderPage();
expect(
await screen.findByText(/Could not load datasets from the Studio API/),
).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
vi.mocked(getDatasets).mockResolvedValueOnce([]);
fireEvent.click(screen.getByRole('button', { name: 'Try Again' }));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
});
+73 -22
View File
@@ -1,12 +1,14 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import AddDatasetButton from '@/components/datasets/AddDatasetButton';
import CreateDatasetModal from '@/components/datasets/CreateDatasetModal';
import DatasetCard from '@/components/datasets/DatasetCard';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import { Card } from '@/components/ui/card';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import { getDatasets } from '@/lib/api';
import type { Dataset } from '@/lib/api';
@@ -21,18 +23,28 @@ import {
export default function DatasetsPage() {
const [datasets, setDatasets] = React.useState<Dataset[]>([]);
const [isInitialLoading, setIsInitialLoading] = React.useState(true);
const [error, setError] = React.useState<string | null>(null);
const { open } = useStore(createDatasetModalStore);
const fetchSequence = React.useRef(0);
const fetchDatasets = React.useCallback(async () => {
const sequence = ++fetchSequence.current;
try {
setDatasets(await getDatasets());
setError(null);
const next = await getDatasets();
if (sequence === fetchSequence.current) {
setDatasets(next);
setError(null);
}
} catch (err) {
console.error('Failed to fetch datasets:', err);
// Distinguish an API outage from a genuinely empty list, so the user
// isn't told they have no datasets when the server is unreachable.
setError(err instanceof Error ? err.message : 'Failed to load datasets');
if (sequence === fetchSequence.current) {
setError(
'Could not load datasets from the Studio API. Check the server and try again.',
);
}
} finally {
if (sequence === fetchSequence.current) setIsInitialLoading(false);
}
}, []);
@@ -50,28 +62,67 @@ export default function DatasetsPage() {
<HeaderActions>
<AddDatasetButton />
</HeaderActions>
<main className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<Card className="p-6">
<div>
{error ? (
<p className="py-8 text-center text-destructive">{error}</p>
) : datasets.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No datasets yet.
</p>
) : (
datasets.map((ds) => (
<DatasetCard
key={ds.id}
dataset={ds}
onUpdated={fetchDatasets}
onSelect={() => handleSelectDataset(ds)}
<div aria-busy={isInitialLoading}>
{isInitialLoading ? (
<div
aria-label="Loading datasets"
className="flex flex-col gap-3 py-2"
>
{[0, 1, 2].map((item) => (
<div
key={item}
className="h-24 animate-pulse rounded-lg border border-border bg-muted/50"
/>
))}
</div>
) : error && datasets.length === 0 ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle
className="size-6 text-destructive"
aria-hidden
/>
))
<p className="max-w-md text-sm text-muted-foreground">
{error}
</p>
<Button type="button" variant="outline" onClick={fetchDatasets}>
Try Again
</Button>
</div>
) : (
<>
{error && (
<p
role="status"
className="mb-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm text-foreground"
>
Dataset updates are temporarily unavailable. Showing the
most recent results.
</p>
)}
{datasets.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No datasets yet.
</p>
) : (
datasets.map((ds) => (
<DatasetCard
key={ds.id}
dataset={ds}
onUpdated={fetchDatasets}
onSelect={() => handleSelectDataset(ds)}
/>
))
)}
</>
)}
</div>
</Card>
</main>
</div>
<CreateDatasetModal
isOpen={open}
onClose={() => setCreateDatasetModalOpen(false)}
@@ -1,4 +1,4 @@
import { render, screen } from '@testing-library/react';
import { fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import GalleryPage from './page';
@@ -43,6 +43,22 @@ describe('GalleryPage', () => {
expect(getJobsList).toHaveBeenCalledWith('inference');
});
it('provides video controls and a visible fallback when media fails', async () => {
vi.mocked(getJobsList).mockResolvedValue([makeJob()]);
renderGallery();
const video = await screen.findByLabelText(
'Generated video: a cat surfing a wave',
);
expect(video).toHaveAttribute('controls');
fireEvent.error(video);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(
screen.getByText('The generated file could not be loaded.'),
).toBeInTheDocument();
});
it('shows the empty state when no completed videos exist', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ status: 'running', output_path: null }),
+70 -21
View File
@@ -1,8 +1,9 @@
'use client';
import { Loader2 } from 'lucide-react';
import { AlertTriangle, ImageOff, Loader2 } from 'lucide-react';
import { useEffect, useState } from 'react';
import { Button } from '@/components/ui/button';
import { Card } from '@/components/ui/card';
import { getJobVideoUrl, getJobsList } from '@/lib/api';
import type { Job } from '@/lib/types';
@@ -11,11 +12,61 @@ function isImage(job: Job): boolean {
return job.output_path?.toLowerCase().endsWith('.png') ?? false;
}
function GalleryMedia({ job }: { job: Job }) {
const [failed, setFailed] = useState(false);
if (failed) {
return (
<div
role="status"
className="flex h-full flex-col items-center justify-center gap-2 px-4 text-center text-muted-foreground"
>
<ImageOff className="size-7" aria-hidden />
<span className="text-sm font-medium">Preview unavailable</span>
<span className="text-xs">
The generated file could not be loaded.
</span>
</div>
);
}
if (isImage(job)) {
return (
// eslint-disable-next-line @next/next/no-img-element
<img
src={getJobVideoUrl(job.id)}
alt={job.prompt}
className="block h-full w-full object-contain"
loading="lazy"
onError={() => setFailed(true)}
/>
);
}
return (
<video
src={getJobVideoUrl(job.id)}
aria-label={
job.prompt ? `Generated video: ${job.prompt}` : 'Generated video'
}
className="block h-full w-full object-contain"
controls
muted
loop
playsInline
preload="metadata"
onError={() => setFailed(true)}
/>
);
}
export default function GalleryPage() {
const [jobs, setJobs] = useState<Job[]>([]);
const [isLoading, setIsLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const [reloadKey, setReloadKey] = useState(0);
useEffect(() => {
let cancelled = false;
async function load() {
@@ -40,7 +91,13 @@ export default function GalleryPage() {
return () => {
cancelled = true;
};
}, []);
}, [reloadKey]);
function retry() {
setError(null);
setIsLoading(true);
setReloadKey((k) => k + 1);
}
const galleryJobs = jobs.filter(
(j) =>
@@ -64,7 +121,16 @@ export default function GalleryPage() {
<span>Loading gallery…</span>
</div>
) : error ? (
<p className="py-8 text-destructive">{error}</p>
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="max-w-md text-sm text-muted-foreground">{error}</p>
<Button type="button" variant="outline" onClick={retry}>
Try Again
</Button>
</div>
) : galleryJobs.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No completed videos yet
@@ -77,24 +143,7 @@ export default function GalleryPage() {
className="flex flex-col overflow-hidden rounded-lg border border-border bg-background"
>
<div className="relative aspect-video overflow-hidden bg-muted">
{isImage(job) ? (
// eslint-disable-next-line @next/next/no-img-element
<img
src={getJobVideoUrl(job.id)}
alt={job.prompt}
className="block h-full w-full object-contain"
loading="lazy"
/>
) : (
<video
src={getJobVideoUrl(job.id)}
className="block h-full w-full object-contain"
muted
loop
playsInline
preload="metadata"
/>
)}
<GalleryMedia job={job} />
</div>
<p
className="line-clamp-3 border-t border-border px-4 py-3 text-sm text-muted-foreground"
+19 -2
View File
@@ -41,7 +41,7 @@
--border: #e2e8f0;
--input: #cbd5e1;
--ring: #94a3b8;
--ring: #1d4ed8;
--radius: 0.5rem;
}
@@ -77,7 +77,7 @@
--border: #334155;
--input: #334155;
--ring: #cbd5e1;
--ring: #7dd3fc;
}
@theme inline {
@@ -125,6 +125,7 @@
html,
body {
min-height: 100%;
overflow-x: clip;
}
html {
@@ -163,6 +164,22 @@ a {
color: inherit;
}
:where(
a,
button,
input,
textarea,
select,
summary,
[role="button"],
[role="menuitem"],
[role="slider"],
[tabindex]
):focus-visible {
outline: 3px solid var(--ring) !important;
outline-offset: 2px !important;
}
summary {
list-style: none;
}
@@ -0,0 +1,55 @@
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { describe, expect, it } from 'vitest';
const css = readFileSync(join(process.cwd(), 'src/app/globals.css'), 'utf8');
function token(block: string, name: string): string {
const match = block.match(new RegExp(`--${name}:\\s*(#[0-9a-fA-F]{6})`));
if (!match) throw new Error(`Missing --${name} token`);
return match[1];
}
function luminance(hex: string): number {
const channels = hex
.slice(1)
.match(/.{2}/g)!
.map((channel) => parseInt(channel, 16) / 255)
.map((channel) =>
channel <= 0.04045
? channel / 12.92
: ((channel + 0.055) / 1.055) ** 2.4,
);
return (
0.2126 * channels[0] + 0.7152 * channels[1] + 0.0722 * channels[2]
);
}
function contrast(first: string, second: string): number {
const firstLuminance = luminance(first);
const secondLuminance = luminance(second);
return (
(Math.max(firstLuminance, secondLuminance) + 0.05) /
(Math.min(firstLuminance, secondLuminance) + 0.05)
);
}
describe('global focus styles', () => {
it('keeps focus tokens above 3:1 against both page themes', () => {
const light = css.match(/:root\s*{([\s\S]*?)\n}/)?.[1] ?? '';
const dark = css.match(/\.dark\s*{([\s\S]*?)\n}/)?.[1] ?? '';
expect(contrast(token(light, 'ring'), token(light, 'background'))).toBeGreaterThanOrEqual(
3,
);
expect(contrast(token(dark, 'ring'), token(dark, 'background'))).toBeGreaterThanOrEqual(
3,
);
});
it('applies a non-animated three-pixel outline to focus-visible controls', () => {
expect(css).toContain('):focus-visible {');
expect(css).toContain('outline: 3px solid var(--ring) !important;');
expect(css).toContain('outline-offset: 2px !important;');
});
});
+2 -2
View File
@@ -4,8 +4,8 @@ import GpuGrid from '@/components/system/GpuGrid';
export default function GpusPage() {
return (
<main className="mx-auto flex w-full max-w-[1100px] flex-col gap-6 px-4 pb-12 pt-6">
<div className="mx-auto flex w-full max-w-[1100px] flex-col gap-6 px-4 pb-12 pt-6">
<GpuGrid />
</main>
</div>
);
}
@@ -52,6 +52,20 @@ describe('Settings page', () => {
expect(updateOption).toHaveBeenCalledWith('numFrames', expect.any(Number));
});
it('gives every slider an accessible name', () => {
renderPage();
const sliders = screen.getAllByRole('slider');
expect(sliders).toHaveLength(11);
for (const slider of sliders) {
expect(slider).toHaveAccessibleName();
}
expect(screen.getByRole('slider', { name: 'Frames' })).toBeInTheDocument();
expect(
screen.getByRole('slider', { name: 'Guidance Scale' }),
).toBeInTheDocument();
});
it('calls resetToDefaults when Reset to Defaults is clicked', () => {
renderPage();
fireEvent.click(screen.getByRole('button', { name: 'Reset to Defaults' }));
@@ -25,6 +25,16 @@ beforeEach(() => {
});
describe('DatasetCard', () => {
it('keeps selection and delete buttons as semantic siblings', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
const selectButton = screen.getByRole('button', { pressed: false });
const deleteButton = screen.getByRole('button', { name: 'Delete' });
expect(selectButton).toHaveTextContent('My Dataset');
expect(selectButton).not.toContainElement(deleteButton);
});
it('renders the name, file count and human-readable size', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
expect(screen.getByText('My Dataset')).toBeInTheDocument();
@@ -71,7 +81,7 @@ describe('DatasetCard', () => {
expect(onSelect).not.toHaveBeenCalled();
});
it('selects on keyboard activation of the card body but not of the Delete button', () => {
it('keeps the selection and delete actions separate', () => {
const onSelect = vi.fn();
render(
<DatasetCard dataset={dataset} onUpdated={() => {}} onSelect={onSelect} />,
@@ -83,8 +93,12 @@ describe('DatasetCard', () => {
});
expect(onSelect).not.toHaveBeenCalled();
// Activating the card body itself does select.
fireEvent.keyDown(screen.getByText('My Dataset'), { key: 'Enter' });
// Activating the dedicated selection button selects the dataset.
fireEvent.click(
screen.getByRole('button', {
name: /My Dataset.*3 files.*2.0 KB/,
}),
);
expect(onSelect).toHaveBeenCalledTimes(1);
});
@@ -55,43 +55,33 @@ export default function DatasetCard({
}
}
function handleKeyDown(e: React.KeyboardEvent) {
if ((e.target as HTMLElement).closest('button')) return;
if (e.key === 'Enter' || e.key === ' ') {
e.preventDefault();
onSelect();
}
}
return (
<div
<article
className={cn(
'mb-3 flex cursor-pointer flex-col gap-[0.6rem] rounded-lg border border-border bg-background px-[1.15rem] py-4',
'mb-3 flex items-start gap-3 rounded-lg border border-border bg-background px-[1.15rem] py-4',
isSelected && 'border-accent-blue bg-accent-blue/5',
)}
onClick={(e) => {
if ((e.target as HTMLElement).closest('button')) return;
onSelect();
}}
onKeyDown={handleKeyDown}
role="button"
tabIndex={0}
>
<div className="flex flex-wrap items-center justify-between gap-2">
<button
type="button"
aria-pressed={isSelected}
onClick={onSelect}
className="flex min-w-0 flex-1 cursor-pointer flex-col gap-[0.6rem] rounded-md text-left"
>
<span className="text-[0.95rem] font-semibold">{dataset.name}</span>
<Button
type="button"
variant="destructive"
size="sm"
onClick={handleDelete}
disabled={isLoading}
>
Delete
</Button>
</div>
<div className="text-sm text-muted-foreground">
{fileCount} {fileCount === 1 ? 'file' : 'files'} · {sizeLabel}
</div>
</div>
<span className="text-sm text-muted-foreground">
{fileCount} {fileCount === 1 ? 'file' : 'files'} · {sizeLabel}
</span>
</button>
<Button
type="button"
variant="destructive"
size="sm"
onClick={handleDelete}
disabled={isLoading}
>
Delete
</Button>
</article>
);
}
@@ -6,6 +6,9 @@ import * as api from '@/lib/api';
import type { Dataset } from '@/lib/api';
vi.mock('@/lib/api');
vi.mock('sonner', () => ({
toast: { error: vi.fn() },
}));
const mockedApi = vi.mocked(api);
@@ -27,6 +30,27 @@ beforeEach(() => {
});
describe('DatasetSidebar', () => {
it('fills the mobile viewport without reserving main-content width', async () => {
const onWidthChange = vi.fn();
render(
<DatasetSidebar
dataset={dataset}
isMobile
onClose={() => {}}
onWidthChange={onWidthChange}
/>,
);
const drawer = screen.getByRole('dialog', {
name: 'My Dataset dataset details',
});
expect(drawer).toHaveStyle({ width: '100%', maxWidth: 'none' });
expect(drawer).toHaveAttribute('aria-modal', 'true');
expect(drawer).toHaveFocus();
expect(onWidthChange).toHaveBeenCalledWith(0);
});
it('lists dataset files after loading', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
@@ -43,6 +67,16 @@ describe('DatasetSidebar', () => {
expect(mockedApi.getDatasetMediaUrl).toHaveBeenCalledWith('ds-1', 'b.mp4');
});
it('shows a fallback when a dataset preview cannot load', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const preview = await screen.findByLabelText('Preview of a.mp4');
fireEvent.error(preview);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(screen.queryByLabelText('Preview of a.mp4')).not.toBeInTheDocument();
});
it('debounces caption save by 500ms', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
@@ -99,6 +133,39 @@ describe('DatasetSidebar', () => {
}
});
it('shows a failed save and lets the user retry it', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
mockedApi.updateDatasetCaption
.mockRejectedValueOnce(new Error('network down'))
.mockResolvedValueOnce(undefined);
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'needs retry' } });
await act(async () => {
await vi.advanceTimersByTimeAsync(500);
});
expect(screen.getByText(/Not saved/)).toBeInTheDocument();
fireEvent.click(screen.getByRole('button', { name: 'Retry' }));
await act(async () => {
await Promise.resolve();
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(2);
expect(mockedApi.updateDatasetCaption).toHaveBeenLastCalledWith(
'ds-1',
'a.mp4',
'needs retry',
);
expect(screen.getByText('Saved')).toBeInTheDocument();
} finally {
vi.useRealTimers();
}
});
it('debounces per file: editing another caption does not cancel a pending save', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
await screen.findByDisplayValue('cap a');

Some files were not shown because too many files have changed in this diff Show More