main
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f71eef465a |
refactor(log): rename console prefix to [VibeVoice TTS]
Partial rename completed across 24 source files (73 sites). Tests pinned the old literal; test_logging_idiom.py now derives from PREFIX. |
||
|
|
8389ab4ce6 |
refactor: log through ComfyUI's own root-logger idiom, re-audit every level
The package installed its own handler, Formatter and level on a module-level logger, then logged through that logger. ComfyUI core already installs app.logger.ColoredFormatter on the ROOT logger, so a bare logging.<level>() call whose message starts with "[ComfyUI-VibeVoice] " is tagged and coloured for free. This removes the duplicate machinery rather than extending it -- no vv_logging module, Logger subclass, LoggerAdapter, vendored ANSI table, new handler, setLevel or propagate. Three hazards that motivated deleting the setup block rather than moving it: - propagate=False severed pytest's caplog; the tests only passed because __init__.py's `if "pytest" in sys.modules` guard skipped the block. Bare root calls propagate by default. - logger.setLevel(INFO) on the package logger pinned every descendant, so ComfyUI's --verbose never applied to this node. Deleting the block fixes it. - The level must stay a literal at the call site; no env-var switch was added. modules/diagnostics.py gates diagnostic CONTENT and is untouched. Every call site (235 across 34 files) is now a bare root call with the prefix. The vendored src/vibevoice tree used transformers.utils.logging, which has no module-level info/warning/debug; those files now import stdlib logging, so a new call site there would fail loudly instead of silently. Level audit against ERROR=failure / WARNING=degradation / INFO=user-facing progress / DEBUG=internals. Deleted as noise: the resample notice (44.1 kHz reference audio is the normal case) and the two "Successfully loaded external VibeVoice" confirmations, which duplicated the patcher's load line. Promoted to WARNING: a mixed-naming GGUF, which is resolved by heuristic majority vote and silently aliases the rest. Demoted to DEBUG: the four SageAttention kernel-selection lines, model discovery, shard counts, per-retry download attempts, attention-mode confirmation and the save_pretrained notice. Kept at INFO: generation complete, transcription results, model downloads and load starts. Judgement calls recorded in docs/2026-10-01-vv-logging-cleanup-design.md. The two I am least sure of: the ungated memory-census profile line was demoted rather than gated (adding a gate would not be presentation-only), and patcher.py's "Loading VibeVoice models for..." was kept at INFO against the ask's example list because it is the only load-start line the package has. 48 new tests in tests/test_logging_idiom.py plus tests/test_audio_utils.py: byte-exact ColoredFormatter rendering, no-leftover-machinery, AST prefix and two-direction level policy, and the 44.1 kHz resample regression (the call still fires with (44100, 24000); nothing is logged). 37 caplog.at_level pins that named a module logger were stripped -- they only lowered a named logger and left root at WARNING, so the record was discarded before capture. Gate: 1924 passed, 30 skipped, 0 failed (1876 before this change). Twelve deliberate mutations -- re-deleting a message, re-leveling, re-adding setLevel and propagate=False, stripping a prefix, moving the resample out of its branch -- were each caught by at least one test. Not run: ComfyUI was never launched and no checkpoint was loaded. |
||
|
|
a5f8103c66 |
feat: load every weight route file->VRAM via core's aimdo DMA (v2.13.0)
Loading problem closed for every model type. Live, 2026-09-30 (RTX 4070 Ti SUPER 16GB), 1.5B bf16 and 7B fp8 both load straight into VRAM with +0.69 GB / +1.45 GB of machine RAM and no sustained SSD traffic. What was wrong -------------- The per-tensor placement used `Tensor.to(cuda)` on a memory-mapped view. That is a HOST-side read: the copy engine faults every page in through the CPU, at 0.55 GB/s and +2.18 GB of machine RAM per 2 GB (report 2026-09-30 section F6). Every route paid it, on every model type. The 7B fp8 load had looked SSD-free only because pass 1's `safe_open` was committing the whole file as private memory first, so those host-side copies were served from RAM. The 19.7 -> 28 -> 20 GB spike was the receipt for that. What changed ------------ * base_loader.place_tensor_on_device(): one placement helper for every route. With an aimdo mapping (main.py sets aimdo_enabled at startup) it uses core's own `read_tensor_file_slice_into` to DMA the file byte range straight into a preallocated CUDA tensor -- the same primitive core already uses to page weights into VRAM, measured at 2.6 GB/s with the page cache left clean. Falls back to `.to()` when core declines or when the aimdo native library raises; a one-time `[vvload]` log line confirms the DMA is live. * One streaming assign for all routes. `_stream_apply_dense` gained `target_device`, and a new `_stream_apply_dense_safetensors` is the dense twin of `_stream_apply_safetensors`. The TTS and ASR dense branches no longer build a whole CPU state dict and call `model.to(cuda)`; GGUF and the internal official-model loader joined the same path. * `read_safetensors_tensors_by_name()`: pass-1 dequant scales are read from the byte ranges the safetensors header already records, instead of `safe_open`. That drops a genuine ~1x-file private commit. * select_patcher_class() always returns the standard ModelPatcher, so core can full-load instead of trapping weights in VBAR for an autoregressive model. Tests ----- 5 new tests for DMA routing and both fallbacks; load-path tests updated to the new read seam. Targeted run: 373 passed / 34 failed, all 34 pre-existing and asserting the ModelPatcherDynamic/vbar architecture this change removes (26 in test_dynamic_patcher_selection.py). Full suite not run, per user rule. |
||
|
|
567fb03e30 |
feat: align VibeVoice loading with core DynamicVRAM, fix RAM ghost (v2.12.0)
Loading filled host RAM 20->40GB while streaming from SSD and never released it; inference then re-served the whole model per AR step. Teardown against core's aimdo/DynamicVRAM machinery, measured on the real 9.47GB 7B fp8 and 5.16GB 1.5B bf16 checkpoints, found four loader defects plus one that was never ours to fix. Defects fixed (each with before/after numbers): * Quant pass 1 pre-read each dequant-at-load layer's scale with safetensors.safe_open + get_tensor. On Windows one such call commits ~1x file size as private, untouched memory, pinned for the lifetime of the returned tensor: +9,050MB on the 7B fp8, invisible to both the working-set counter and the storage census. Now maps the file once through core's aimdo arm and clones only the tiny scales. Whole load: 1.51x -> 0.13x file size; retained after free 11.3GB -> 96MB. * Quant families were excluded from core's dynamic patcher by an inherited "quant streams natively" stop condition. Because a legacy patcher has nowhere to page from, those routes cloned every tensor into host RAM and fully H2D'd it. select_patcher_class now follows core's availability rule for every family, with gguf_block still excluded for a measured reason (GGUFTensor.from_reader_tensor clones the reader's view). * replace_linears_for_quant built its resident modules outside the meta context: 8.08GB of never-written host allocation per 7B fp8 load. * Dormant weight_function/bias_function double application: core already applies them inside cast_bias_weight (ops.py:431-438). Exonerated with numbers, not assumed: the dense read path is core's own (13.5MB private for a 5.16GB file), and core's file->VRAM paging read is cache-clean (+0.13GB per 2GB). Only host-side view reads reproduce the reported 1:1 RAM at 0.55GB/s signature. Inference: reproduced headlessly that when the tree does not fit in the VRAM that is actually free, every forward re-reads every weight from the checkpoint file (3072/3072 file reads, 21ms -> 125ms per step). Our wrappers are protocol-correct; residency is stable at 100% with headroom. [vvpull] now reports resident vs reread with bytes and free VRAM so one live line settles it, behind VIBEVOICE_VBAR_OBSERVER=0. Not changed after being tried and reverted: deriving fast_disk from the checkpoint. It measured as a no-op on the dev host but converted RAM-speed pinned re-reads into disk-speed reads live, which made loading dramatically slower. Reverted in full; see the report for the probe-design lesson. Adds host-RAM instrumentation ([vvrss] with machine-level start/end, [vvcensus] storage census, tts-generate bracket), four standalone probes, and ~4000 lines of tests pinning the invariants above. |
||
|
|
08df29df25 |
feat: auto-detect config_name, drop VibeVoice-Large (v2.7.0)
Config auto-detection: - config_detect: read the checkpoint embedding shape as an architecture fingerprint (header-only for safetensors, reuses the open GGUF reader); 7B=[152064,3584], 1.5B=[151936,1536]; orientation-agnostic for shape-reversed GGUF files - Auto-detect is the new default config_name; resolves the family before any heavy load, or fails fast with an actionable error (.bin/.pt and unknown families cannot be fingerprinted) - an explicit config_name that contradicts the weights self-corrects to the detected family with one WARNING (reconcile_config) - loader: friendly shape pre-check in _apply_state_dict names the offending tensors and hints at config_name instead of torch's raw size-mismatch RuntimeError Dropdown dedup: - VibeVoice-Large removed from config_name options (duplicate of 7B); kept as a legacy alias so saved workflows still load (normalize at node + loader entry; validate_inputs(**kwargs) override makes core skip its combo-membership check) - node resolves Auto-detect before computing the cache identity so the request key matches the consumer's bundle-derived key (no per-run churn) Console noise: - demote ~30 internal INFO logs to DEBUG across loaders/patcher/registry - drop two stray tie_weights prints; tied lm_head.weight no longer warned as missing (expected under tie_word_embeddings) Tests: 1060 passed / 5 pre-existing failures / 4 skipped |
||
|
|
29488e4a1f |
feat: unload-on-change eviction + quant-resident runtime for GGUF and quantized safetensors (v2.4.0 -> v2.5.0)
Model management (v2.4.0): - single-active-per-family registry releases the previous model fully (RAM, VRAM, ComfyUI current_loaded_models) before a new one loads - file-identity cache keys (basename+mtime+size+attention+q4+dtype) prevent cross-file collisions; ASR request keys share the consumer namespace so identical re-runs never spuriously evict Quant-resident runtime (v2.5.0): - GGUF weights stay raw-block resident end-to-end: uint8 parameters, per-matmul dequant kernels for Q8_0/Q4_K/Q5_K/Q6_K pinned bitwise to the gguf-py oracle; F32/F16/BF16 pass through at native dtype via zero-copy views; load-time RAM spike (~2x float size) eliminated - quantized safetensors via *.comfy_quant metadata: rotated ConvRot INT8 residents through comfy-kitchen, plain rowwise int8 / fp8 e4m3+e5m2 / int8_blockwise dequant-at-load (per-row, scalar, and per-gs-block scale layouts); unsupported formats hard-fail with actionable errors; dense gate rejects unplanned quant storages; rotated non-Linear targets (embeddings) fail with re-export guidance - dtype casts filter quant-resident storage (_quant_resident markers, fp32 weight_scale protection); SageAttention wrapper resolves activation dtype per module kind, fixing uint8-weight crash Tests: 937 passed / 5 pre-existing failures / 4 skipped |