4 Commits
Author SHA1 Message Date
WildAi 567fb03e30 feat: align VibeVoice loading with core DynamicVRAM, fix RAM ghost (v2.12.0)
Loading filled host RAM 20->40GB while streaming from SSD and never released
it; inference then re-served the whole model per AR step. Teardown against
core's aimdo/DynamicVRAM machinery, measured on the real 9.47GB 7B fp8 and
5.16GB 1.5B bf16 checkpoints, found four loader defects plus one that was
never ours to fix.

Defects fixed (each with before/after numbers):

* Quant pass 1 pre-read each dequant-at-load layer's scale with
  safetensors.safe_open + get_tensor. On Windows one such call commits ~1x
  file size as private, untouched memory, pinned for the lifetime of the
  returned tensor: +9,050MB on the 7B fp8, invisible to both the working-set
  counter and the storage census. Now maps the file once through core's
  aimdo arm and clones only the tiny scales. Whole load: 1.51x -> 0.13x file
  size; retained after free 11.3GB -> 96MB.

* Quant families were excluded from core's dynamic patcher by an inherited
  "quant streams natively" stop condition. Because a legacy patcher has
  nowhere to page from, those routes cloned every tensor into host RAM and
  fully H2D'd it. select_patcher_class now follows core's availability rule
  for every family, with gguf_block still excluded for a measured reason
  (GGUFTensor.from_reader_tensor clones the reader's view).

* replace_linears_for_quant built its resident modules outside the meta
  context: 8.08GB of never-written host allocation per 7B fp8 load.

* Dormant weight_function/bias_function double application: core already
  applies them inside cast_bias_weight (ops.py:431-438).

Exonerated with numbers, not assumed: the dense read path is core's own
(13.5MB private for a 5.16GB file), and core's file->VRAM paging read is
cache-clean (+0.13GB per 2GB). Only host-side view reads reproduce the
reported 1:1 RAM at 0.55GB/s signature.

Inference: reproduced headlessly that when the tree does not fit in the
VRAM that is actually free, every forward re-reads every weight from the
checkpoint file (3072/3072 file reads, 21ms -> 125ms per step). Our wrappers
are protocol-correct; residency is stable at 100% with headroom. [vvpull]
now reports resident vs reread with bytes and free VRAM so one live line
settles it, behind VIBEVOICE_VBAR_OBSERVER=0.

Not changed after being tried and reverted: deriving fast_disk from the
checkpoint. It measured as a no-op on the dev host but converted RAM-speed
pinned re-reads into disk-speed reads live, which made loading dramatically
slower. Reverted in full; see the report for the probe-design lesson.

Adds host-RAM instrumentation ([vvrss] with machine-level start/end, [vvcensus]
storage census, tts-generate bracket), four standalone probes, and ~4000
lines of tests pinning the invariants above.
2026-09-30 18:24:52 +03:00
WildAi d2390eb476 fix: correct realtime VibeVoice output on transformers 5.3
Garbled, parameter-insensitive speech came from a randomized EOS head:
5.3 re-runs _initialize_weights over acoustic_connector and
tts_eos_classifier because the vendored _init_weights override had no
_is_hf_initialized guard. Guard it; only checkpoint-absent weights are
initialized now.

- MockCacheLayer exposes both the 4.x and 5.x cache APIs, so the
  prefilled voice prompt is visible to 5.3's mask builder.
- _ensure_cache_has_layers covers the container: offload/prefetch,
  batch ops, crop, and a copyable lazy prefetch stream.
- max_new_tokens is a combined text+speech budget, not latents.
- cfg_scale floor of 1.5 for the realtime family.
- voice presets resolve against every registered TTS root.
- sage excluded from the realtime path: it ignores the attention mask.
- bound transformers to >=5.3.0,<5.4, the measured line.
2026-09-26 23:30:37 +03:00
WildAi 10e70d8e25 perf: fp8-resident weights + streaming safetensors load (v2.8.0)
Loading VibeVoice-*-fp8_e4m3.safetensors materialized the full state
dict and dequantized all ~380 fp8 layers to bf16 in RAM: 18->43 GB
spike, Pin-error flood, ~4.5 GB partial offload on 16 GB cards.

- FP8Linear keeps fp8 weight + scalar fp32 scale resident in VRAM
  (~1 byte/weight); forward dequantizes per-tensor via comfy-kitchen
  into the activation dtype. Missing kitchen backend falls back to
  dequant-at-load.
- Quantized safetensors load streams per-tensor (never a full
  in-memory state dict); dense files keep the batch path.
- Non-Linear fp8 targets (real exports quantize embed_tokens) demote
  to dequant-at-load; ConvRot non-Linear targets keep the hard fail.
- Diffusion head resolves the timestep-mlp input dtype via the
  module's compute_dtype, never the fp8 storage dtype — casting
  activations to weight.dtype fed fp8 into the dequant kernel and
  crashed generate() (NoCapableBackendError).

GPU-gated on 1.5B + 7B fp8: full VRAM fit, no partial offload, no
pin flood, audio generates.
2026-08-28 01:24:26 +03:00
WildAi 29488e4a1f feat: unload-on-change eviction + quant-resident runtime for GGUF and quantized safetensors (v2.4.0 -> v2.5.0)
Model management (v2.4.0):
- single-active-per-family registry releases the previous model fully
  (RAM, VRAM, ComfyUI current_loaded_models) before a new one loads
- file-identity cache keys (basename+mtime+size+attention+q4+dtype)
  prevent cross-file collisions; ASR request keys share the consumer
  namespace so identical re-runs never spuriously evict

Quant-resident runtime (v2.5.0):
- GGUF weights stay raw-block resident end-to-end: uint8 parameters,
  per-matmul dequant kernels for Q8_0/Q4_K/Q5_K/Q6_K pinned bitwise to
  the gguf-py oracle; F32/F16/BF16 pass through at native dtype via
  zero-copy views; load-time RAM spike (~2x float size) eliminated
- quantized safetensors via *.comfy_quant metadata: rotated ConvRot
  INT8 residents through comfy-kitchen, plain rowwise int8 / fp8
  e4m3+e5m2 / int8_blockwise dequant-at-load (per-row, scalar, and
  per-gs-block scale layouts); unsupported formats hard-fail with
  actionable errors; dense gate rejects unplanned quant storages;
  rotated non-Linear targets (embeddings) fail with re-export guidance
- dtype casts filter quant-resident storage (_quant_resident markers,
  fp32 weight_scale protection); SageAttention wrapper resolves
  activation dtype per module kind, fixing uint8-weight crash

Tests: 937 passed / 5 pre-existing failures / 4 skipped
2026-08-26 12:17:24 +03:00