21 Commits
Author SHA1 Message Date
WildAi f71eef465a refactor(log): rename console prefix to [VibeVoice TTS]
Partial rename completed across 24 source files (73 sites). Tests pinned
the old literal; test_logging_idiom.py now derives from PREFIX.
2026-10-03 16:03:14 +03:00
WildAi 8389ab4ce6 refactor: log through ComfyUI's own root-logger idiom, re-audit every level
The package installed its own handler, Formatter and level on a
module-level logger, then logged through that logger. ComfyUI core already
installs app.logger.ColoredFormatter on the ROOT logger, so a bare
logging.<level>() call whose message starts with "[ComfyUI-VibeVoice] " is
tagged and coloured for free. This removes the duplicate machinery rather
than extending it -- no vv_logging module, Logger subclass, LoggerAdapter,
vendored ANSI table, new handler, setLevel or propagate.

Three hazards that motivated deleting the setup block rather than moving it:

- propagate=False severed pytest's caplog; the tests only passed because
  __init__.py's `if "pytest" in sys.modules` guard skipped the block.
  Bare root calls propagate by default.
- logger.setLevel(INFO) on the package logger pinned every descendant, so
  ComfyUI's --verbose never applied to this node. Deleting the block fixes it.
- The level must stay a literal at the call site; no env-var switch was
  added. modules/diagnostics.py gates diagnostic CONTENT and is untouched.

Every call site (235 across 34 files) is now a bare root call with the
prefix. The vendored src/vibevoice tree used transformers.utils.logging,
which has no module-level info/warning/debug; those files now import stdlib
logging, so a new call site there would fail loudly instead of silently.

Level audit against ERROR=failure / WARNING=degradation / INFO=user-facing
progress / DEBUG=internals. Deleted as noise: the resample notice (44.1 kHz
reference audio is the normal case) and the two "Successfully loaded
external VibeVoice" confirmations, which duplicated the patcher's load line.
Promoted to WARNING: a mixed-naming GGUF, which is resolved by heuristic
majority vote and silently aliases the rest. Demoted to DEBUG: the four
SageAttention kernel-selection lines, model discovery, shard counts,
per-retry download attempts, attention-mode confirmation and the
save_pretrained notice. Kept at INFO: generation complete, transcription
results, model downloads and load starts.

Judgement calls recorded in docs/2026-10-01-vv-logging-cleanup-design.md.
The two I am least sure of: the ungated memory-census profile line was
demoted rather than gated (adding a gate would not be presentation-only),
and patcher.py's "Loading VibeVoice models for..." was kept at INFO against
the ask's example list because it is the only load-start line the package has.

48 new tests in tests/test_logging_idiom.py plus tests/test_audio_utils.py:
byte-exact ColoredFormatter rendering, no-leftover-machinery, AST prefix and
two-direction level policy, and the 44.1 kHz resample regression (the call
still fires with (44100, 24000); nothing is logged). 37 caplog.at_level
pins that named a module logger were stripped -- they only lowered a named
logger and left root at WARNING, so the record was discarded before capture.

Gate: 1924 passed, 30 skipped, 0 failed (1876 before this change). Twelve
deliberate mutations -- re-deleting a message, re-leveling, re-adding
setLevel and propagate=False, stripping a prefix, moving the resample out
of its branch -- were each caught by at least one test.

Not run: ComfyUI was never launched and no checkpoint was loaded.
2026-10-01 23:08:04 +03:00
WildAi a5f8103c66 feat: load every weight route file->VRAM via core's aimdo DMA (v2.13.0)
Loading problem closed for every model type. Live, 2026-09-30 (RTX 4070 Ti
SUPER 16GB), 1.5B bf16 and 7B fp8 both load straight into VRAM with
+0.69 GB / +1.45 GB of machine RAM and no sustained SSD traffic.

What was wrong
--------------
The per-tensor placement used `Tensor.to(cuda)` on a memory-mapped view. That
is a HOST-side read: the copy engine faults every page in through the CPU, at
0.55 GB/s and +2.18 GB of machine RAM per 2 GB (report 2026-09-30 section F6).
Every route paid it, on every model type.

The 7B fp8 load had looked SSD-free only because pass 1's `safe_open` was
committing the whole file as private memory first, so those host-side copies
were served from RAM. The 19.7 -> 28 -> 20 GB spike was the receipt for that.

What changed
------------
* base_loader.place_tensor_on_device(): one placement helper for every route.
  With an aimdo mapping (main.py sets aimdo_enabled at startup) it uses core's
  own `read_tensor_file_slice_into` to DMA the file byte range straight into a
  preallocated CUDA tensor -- the same primitive core already uses to page
  weights into VRAM, measured at 2.6 GB/s with the page cache left clean.
  Falls back to `.to()` when core declines or when the aimdo native library
  raises; a one-time `[vvload]` log line confirms the DMA is live.

* One streaming assign for all routes. `_stream_apply_dense` gained
  `target_device`, and a new `_stream_apply_dense_safetensors` is the dense
  twin of `_stream_apply_safetensors`. The TTS and ASR dense branches no longer
  build a whole CPU state dict and call `model.to(cuda)`; GGUF and the internal
  official-model loader joined the same path.

* `read_safetensors_tensors_by_name()`: pass-1 dequant scales are read from the
  byte ranges the safetensors header already records, instead of `safe_open`.
  That drops a genuine ~1x-file private commit.

* select_patcher_class() always returns the standard ModelPatcher, so core can
  full-load instead of trapping weights in VBAR for an autoregressive model.

Tests
-----
5 new tests for DMA routing and both fallbacks; load-path tests updated to the
new read seam. Targeted run: 373 passed / 34 failed, all 34 pre-existing and
asserting the ModelPatcherDynamic/vbar architecture this change removes
(26 in test_dynamic_patcher_selection.py). Full suite not run, per user rule.
2026-09-30 23:02:25 +03:00
WildAi 567fb03e30 feat: align VibeVoice loading with core DynamicVRAM, fix RAM ghost (v2.12.0)
Loading filled host RAM 20->40GB while streaming from SSD and never released
it; inference then re-served the whole model per AR step. Teardown against
core's aimdo/DynamicVRAM machinery, measured on the real 9.47GB 7B fp8 and
5.16GB 1.5B bf16 checkpoints, found four loader defects plus one that was
never ours to fix.

Defects fixed (each with before/after numbers):

* Quant pass 1 pre-read each dequant-at-load layer's scale with
  safetensors.safe_open + get_tensor. On Windows one such call commits ~1x
  file size as private, untouched memory, pinned for the lifetime of the
  returned tensor: +9,050MB on the 7B fp8, invisible to both the working-set
  counter and the storage census. Now maps the file once through core's
  aimdo arm and clones only the tiny scales. Whole load: 1.51x -> 0.13x file
  size; retained after free 11.3GB -> 96MB.

* Quant families were excluded from core's dynamic patcher by an inherited
  "quant streams natively" stop condition. Because a legacy patcher has
  nowhere to page from, those routes cloned every tensor into host RAM and
  fully H2D'd it. select_patcher_class now follows core's availability rule
  for every family, with gguf_block still excluded for a measured reason
  (GGUFTensor.from_reader_tensor clones the reader's view).

* replace_linears_for_quant built its resident modules outside the meta
  context: 8.08GB of never-written host allocation per 7B fp8 load.

* Dormant weight_function/bias_function double application: core already
  applies them inside cast_bias_weight (ops.py:431-438).

Exonerated with numbers, not assumed: the dense read path is core's own
(13.5MB private for a 5.16GB file), and core's file->VRAM paging read is
cache-clean (+0.13GB per 2GB). Only host-side view reads reproduce the
reported 1:1 RAM at 0.55GB/s signature.

Inference: reproduced headlessly that when the tree does not fit in the
VRAM that is actually free, every forward re-reads every weight from the
checkpoint file (3072/3072 file reads, 21ms -> 125ms per step). Our wrappers
are protocol-correct; residency is stable at 100% with headroom. [vvpull]
now reports resident vs reread with bytes and free VRAM so one live line
settles it, behind VIBEVOICE_VBAR_OBSERVER=0.

Not changed after being tried and reverted: deriving fast_disk from the
checkpoint. It measured as a no-op on the dev host but converted RAM-speed
pinned re-reads into disk-speed reads live, which made loading dramatically
slower. Reverted in full; see the report for the probe-design lesson.

Adds host-RAM instrumentation ([vvrss] with machine-level start/end, [vvcensus]
storage census, tts-generate bracket), four standalone probes, and ~4000
lines of tests pinning the invariants above.
2026-09-30 18:24:52 +03:00
WildAi cfa57c8738 feat: streaming quant load, single-rounding dequant, release hygiene (v2.11.0)
Loading
- GGUF install is now two passes: a metadata pass that decides each
  tensor's disposition, then an install pass that assigns residents in
  place and streams dense tensors. The whole checkpoint is no longer
  buffered in a dict alongside the model being built.
- dequantize_reader_tensor takes a target dtype, so dequant-at-load
  writes straight into the destination parameter and the fp32
  intermediate is never allocated.
- A bundle whose heavy fields were released is rebuilt from its recorded
  source_path instead of failing the consumer.
- Host memory is released after install.

Numerics
- Q8_0 dequant computes in fp32 so the result is rounded once, at the
  final cast. The activation-dtype path was reverted: it rounded twice
  and moved stored weights.
- Removed a redundant weight-sized copy from the dequant kernel.
  Bitwise-identical, ~1.16x.
- Precision gates are bitwise rather than tolerance-based.

Docs and tests
- README condensed; changelog moved to CHANGELOG.md.
- Third-party project references removed from source comments.
- Tests no longer assert README prose; the e2e smoke contract follows
  the developer script to its new location and skips when absent.
- Version guard reads CHANGELOG.md.
2026-09-28 18:25:19 +03:00
WildAi d2390eb476 fix: correct realtime VibeVoice output on transformers 5.3
Garbled, parameter-insensitive speech came from a randomized EOS head:
5.3 re-runs _initialize_weights over acoustic_connector and
tts_eos_classifier because the vendored _init_weights override had no
_is_hf_initialized guard. Guard it; only checkpoint-absent weights are
initialized now.

- MockCacheLayer exposes both the 4.x and 5.x cache APIs, so the
  prefilled voice prompt is visible to 5.3's mask builder.
- _ensure_cache_has_layers covers the container: offload/prefetch,
  batch ops, crop, and a copyable lazy prefetch stream.
- max_new_tokens is a combined text+speech budget, not latents.
- cfg_scale floor of 1.5 for the realtime family.
- voice presets resolve against every registered TTS root.
- sage excluded from the realtime path: it ignores the attention mask.
- bound transformers to >=5.3.0,<5.4, the measured line.
2026-09-26 23:30:37 +03:00
WildAi 8404100f31 perf: stream sharded/quant loads, console progress, dep cleanup
- streaming per-tensor assign (tensor.clone) replaces merged-dict
  load_state_dict_sharded — mmap views were keeping all shard files
  resident after partial offload (+15GB ghost RSS, Pin-error flood);
  private RAM copies pin reliably and release shard mappings early
- console progress bar during TTS/ASR generation (progress_utils)
- GGUF read-only mmap no longer warns (copy before from_numpy)
- demote internal load-narration logs to DEBUG (success-only INFO)
- drop dead deps (einops, tokenizers, s3tokenizer, conformer),
  add accelerate to pyproject (4-bit path needs it), requirements
  trimmed to non-ComfyUI-bundled packages; parity guard tests
2026-09-01 21:39:58 +03:00
WildAi 10e70d8e25 perf: fp8-resident weights + streaming safetensors load (v2.8.0)
Loading VibeVoice-*-fp8_e4m3.safetensors materialized the full state
dict and dequantized all ~380 fp8 layers to bf16 in RAM: 18->43 GB
spike, Pin-error flood, ~4.5 GB partial offload on 16 GB cards.

- FP8Linear keeps fp8 weight + scalar fp32 scale resident in VRAM
  (~1 byte/weight); forward dequantizes per-tensor via comfy-kitchen
  into the activation dtype. Missing kitchen backend falls back to
  dequant-at-load.
- Quantized safetensors load streams per-tensor (never a full
  in-memory state dict); dense files keep the batch path.
- Non-Linear fp8 targets (real exports quantize embed_tokens) demote
  to dequant-at-load; ConvRot non-Linear targets keep the hard fail.
- Diffusion head resolves the timestep-mlp input dtype via the
  module's compute_dtype, never the fp8 storage dtype — casting
  activations to weight.dtype fed fp8 into the dequant kernel and
  crashed generate() (NoCapableBackendError).

GPU-gated on 1.5B + 7B fp8: full VRAM fit, no partial offload, no
pin flood, audio generates.
2026-08-28 01:24:26 +03:00
WildAi 2fc53b85a3 fix: version-safe config dtype access — silence torch_dtype deprecation (v2.7.1)
transformers 5.x deprecated PretrainedConfig.torch_dtype; both getter and setter log a warning on every load. Add set_config_dtype/get_config_dtype helpers (configuration_vibevoice.py + modules/dtype_utils.py) that use config.dtype on v5 and torch_dtype on 4.x, and route the loader/external_loader writes plus the 5 vendored modeling reader blocks through them. +17 regression tests incl. no-warning assertions on the real load path.
2026-08-27 21:49:17 +03:00
WildAi 08df29df25 feat: auto-detect config_name, drop VibeVoice-Large (v2.7.0)
Config auto-detection:
- config_detect: read the checkpoint embedding shape as an architecture
  fingerprint (header-only for safetensors, reuses the open GGUF reader);
  7B=[152064,3584], 1.5B=[151936,1536]; orientation-agnostic for
  shape-reversed GGUF files
- Auto-detect is the new default config_name; resolves the family before
  any heavy load, or fails fast with an actionable error (.bin/.pt and
  unknown families cannot be fingerprinted)
- an explicit config_name that contradicts the weights self-corrects to
  the detected family with one WARNING (reconcile_config)
- loader: friendly shape pre-check in _apply_state_dict names the
  offending tensors and hints at config_name instead of torch's raw
  size-mismatch RuntimeError

Dropdown dedup:
- VibeVoice-Large removed from config_name options (duplicate of 7B);
  kept as a legacy alias so saved workflows still load (normalize at node
  + loader entry; validate_inputs(**kwargs) override makes core skip its
  combo-membership check)
- node resolves Auto-detect before computing the cache identity so the
  request key matches the consumer's bundle-derived key (no per-run churn)

Console noise:
- demote ~30 internal INFO logs to DEBUG across loaders/patcher/registry
- drop two stray tie_weights prints; tied lm_head.weight no longer warned
  as missing (expected under tie_word_embeddings)

Tests: 1060 passed / 5 pre-existing failures / 4 skipped
2026-08-27 20:33:19 +03:00
WildAi 2aaef8228b feat: native lowvram streaming for oversized models (v2.6.0)
- comfy_stream: convert the transformers tree to comfy-cast streaming modules; direct-param containers (Block1D gammas) self-relocate to the compute device on attribute access
- GGUF/ConvRot quant residents stream natively, dtype-preserving
- patcher: drop force-placement workaround, pure core delegation
- vendored generate: derive stage devices from activations, not parameter residency (lies under partial offload)
- generation: place inputs on the runtime compute device
- tests: lowvram protocol, streaming parity, residents, e2e
2026-08-27 17:52:17 +03:00
WildAi 964e164414 fix: lowvram partial-load completion, VibeVoice-7B rename follow-through, copy-free packaged tokenizer
- patcher: complete the H2D transfer for modules ComfyUI's lowvram
  arbiter leaves on CPU during partial loads (plain transformers trees
  have no comfy.ops streaming hooks; crashed at forward with
  'Input type CUDABFloat16Type and weight type CPUBFloat16Type' on
  models larger than free VRAM, e.g. 7B)
- finish Large->7B model_info rename: tokenizer repo selection now
  matches '7B' (previously fell back to the 1.5B tokenizer), 7B added
  to external-loader config options (shares Large packaged defaults),
  stale test assertions updated, tests for the removed external-model
  example workflow deleted
- loader: packaged tokenizer.json loads directly from the node folder
  instead of being copied into user model directories; HF download
  stays as last resort; model dirs stay untouched
2026-08-26 13:24:55 +03:00
WildAi 29488e4a1f feat: unload-on-change eviction + quant-resident runtime for GGUF and quantized safetensors (v2.4.0 -> v2.5.0)
Model management (v2.4.0):
- single-active-per-family registry releases the previous model fully
  (RAM, VRAM, ComfyUI current_loaded_models) before a new one loads
- file-identity cache keys (basename+mtime+size+attention+q4+dtype)
  prevent cross-file collisions; ASR request keys share the consumer
  namespace so identical re-runs never spuriously evict

Quant-resident runtime (v2.5.0):
- GGUF weights stay raw-block resident end-to-end: uint8 parameters,
  per-matmul dequant kernels for Q8_0/Q4_K/Q5_K/Q6_K pinned bitwise to
  the gguf-py oracle; F32/F16/BF16 pass through at native dtype via
  zero-copy views; load-time RAM spike (~2x float size) eliminated
- quantized safetensors via *.comfy_quant metadata: rotated ConvRot
  INT8 residents through comfy-kitchen, plain rowwise int8 / fp8
  e4m3+e5m2 / int8_blockwise dequant-at-load (per-row, scalar, and
  per-gs-block scale layouts); unsupported formats hard-fail with
  actionable errors; dense gate rejects unplanned quant storages;
  rotated non-Linear targets (embeddings) fail with re-export guidance
- dtype casts filter quant-resident storage (_quant_resident markers,
  fp32 weight_scale protection); SageAttention wrapper resolves
  activation dtype per module kind, fixing uint8-weight crash

Tests: 937 passed / 5 pre-existing failures / 4 skipped
2026-08-26 12:17:24 +03:00
WildAi 84840718ea fix: recompute RoPE inv_freq after meta materialization — gibberish output regression (v2.3.2) 2026-08-20 12:18:45 +03:00
WildAi 95757d22cc fix: restore sentinel buffers (nan) after meta materialization — silent output regression (v2.3.1) 2026-08-20 02:06:51 +03:00
WildAi 5c03866a17 fix: materialize meta buffers after assign load (NotImplementedError on unpatch) 2026-08-19 23:28:23 +03:00
WildAi 7139ca22c6 perf: meta-init + assign loading, non-destructive offload contract (v2.3.0) 2026-08-19 18:56:35 +03:00
WildAi dc64c17c00 fix: cpu-first loading + neg-branch rope fix
- DF-001..006: loader builds the model entirely on CPU (state dict, load_state_dict, dtype cast, 4-bit); patcher owns the single H2D transfer after ComfyUI load_models_gpu arbitration. Removes the disk->VRAM->RAM->VRAM round-trip; peak load VRAM ~2x -> ~1x model size.

- BUG-011: negative CFG forward passed full-length position_ids with single-token inputs_embeds; transformers 5.x RoPE broadcast expanded q/k (not v) -> KV cache key/value desync -> SDPA crash [12,3] vs [12,2]. Now current-only position_ids.

- Tests: +36 device-flow, +4 neg-position-ids (red/green verified), +scheduler/e2e/mm-surface suites. Full suite 461p/5 pre-existing f/4s.

- Version 2.0.2; untrack .workbuddy-ai, ignore tmp/
2026-08-15 16:46:56 +03:00
WildAi 851647983b refactor init 2026-08-14 20:26:31 +03:00
WildAi 68ccb138b1 fix tokenizer.json issue, fix num_hidden_layers 2025-09-25 13:15:07 +03:00
WildAi ef1b65b9fa major refactoring 2025-09-10 12:06:26 +03:00