VideoGenerator.from_pretrained(model_path, config) takes the same nested
settings as from_config; the 24 flat keywords, from_pretrained_kwargs_to_config,
FROM_PRETRAINED_KWARGS, and every flat_field in the schema are deleted.
nvfp4_fa4 is the typed field engine.attention.nvfp4_fa4, applied by a
resolution step. 65 call sites and the docs use the nested form; the three
kwargs golden cases are config cases with byte-identical results.
Every value is decided during resolution; no runtime code calls
with_override. The device offload policy (unified memory, MPS, layerwise
conflicts, lazy module load) runs as resolution steps in the main process
(fastvideo/api/device_policy.py), and the worker runs on the config it
receives. The LTX-2 refine defaults from model_index.json and the MiniMax-H3
schedule from fastvideo_inference.json are resolution steps
(fastvideo/api/checkpoint_defaults.py) that read local checkpoints or download
only those files; the H3 pipeline validates the schedule against the loaded
schedulers but no longer writes it. The preprocessing entry point passes the
downloaded local path as run state instead of a config override. Teacher and
critic models load with override_transformer_cls_name as a load argument.
The test isolation turns the device policy off, as it blocks downloads, so
that the goldens do not depend on the machine.
Against the 3f6893a0 goldens, every field still matches except boundary_ratio
and ltx2_vae_tiling (DESIGN section 9) and the fields that no longer exist.
The launcher and trainer YAML verifications report 0 unexplained differences.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
50 lines
1.8 KiB
Python
50 lines
1.8 KiB
Python
# SPDX-License-Identifier: Apache-2.0
|
|
from fastvideo import VideoGenerator
|
|
|
|
|
|
def main():
|
|
# Point this to your local diffusers model dir (or replace with a HF model ID).
|
|
model_path = "KyleShao/Cosmos-Predict2.5-2B-Diffusers"
|
|
|
|
generator = VideoGenerator.from_pretrained(
|
|
model_path,
|
|
{
|
|
"engine": {
|
|
"num_gpus": 1,
|
|
"use_fsdp_inference": False, # set True if GPU is out of memory
|
|
"offload": {
|
|
"dit": False,
|
|
"vae": False,
|
|
"text_encoder": True,
|
|
"pin_cpu_memory": True,
|
|
},
|
|
},
|
|
},
|
|
)
|
|
|
|
# image2world example from official repo
|
|
image_path = "assets/images/bus_terminal.jpg"
|
|
|
|
prompt = (
|
|
"A nighttime city bus terminal gradually shifts from stillness to subtle movement. "
|
|
"At first, multiple double-decker buses are parked under the glow of overhead lights, "
|
|
"with a central bus labeled '87D' facing forward and stationary. "
|
|
"As the video progresses, the bus in the middle moves ahead slowly, its headlights brightening the surrounding area "
|
|
"and casting reflections onto adjacent vehicles. "
|
|
"The motion creates space in the lineup, signaling activity within the otherwise quiet station. "
|
|
"It then comes to a smooth stop, resuming its position in line. "
|
|
"Overhead signage in Chinese characters remains illuminated, enhancing the vibrant, urban night scene.")
|
|
|
|
generator.generate({
|
|
"prompt": prompt,
|
|
"inputs": {"image_path": str(image_path)},
|
|
"output": {"output_path": "outputs_video/cosmos2_5_i2w.mp4", "save_video": True},
|
|
"extensions": {"num_cond_frames": 1},
|
|
})
|
|
|
|
generator.shutdown()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|