4 Commits
Author SHA1 Message Date
WildAi cfa57c8738 feat: streaming quant load, single-rounding dequant, release hygiene (v2.11.0)
Loading
- GGUF install is now two passes: a metadata pass that decides each
  tensor's disposition, then an install pass that assigns residents in
  place and streams dense tensors. The whole checkpoint is no longer
  buffered in a dict alongside the model being built.
- dequantize_reader_tensor takes a target dtype, so dequant-at-load
  writes straight into the destination parameter and the fp32
  intermediate is never allocated.
- A bundle whose heavy fields were released is rebuilt from its recorded
  source_path instead of failing the consumer.
- Host memory is released after install.

Numerics
- Q8_0 dequant computes in fp32 so the result is rounded once, at the
  final cast. The activation-dtype path was reverted: it rounded twice
  and moved stored weights.
- Removed a redundant weight-sized copy from the dequant kernel.
  Bitwise-identical, ~1.16x.
- Precision gates are bitwise rather than tolerance-based.

Docs and tests
- README condensed; changelog moved to CHANGELOG.md.
- Third-party project references removed from source comments.
- Tests no longer assert README prose; the e2e smoke contract follows
  the developer script to its new location and skips when absent.
- Version guard reads CHANGELOG.md.
2026-09-28 18:25:19 +03:00
WildAi d718db15bf feat: drop the dedicated realtime node, move to WMNodes categories (v2.10.0)
VibeVoice TTS already routes realtime models correctly, so nodes/realtime_node.py
and the VibeVoiceRealtime ID are removed along with their shim tests and the legacy
workflow fixture. That node only existed in the unreleased 2.x line; origin/main is
still 1.5.0, so no published workflow referenced it.

- VibeVoice TTS and VibeVoice External Loader -> WMNodes/sound/tts
- VibeVoice ASR -> WMNodes/sound/asr (node IDs unchanged, only menu location moves)
- Tests no longer read local-only files: the transformers range guard now checks a
  measured-version literal, and the GPU-audit fixtures skip when the local
  diagnostic script is absent. A clone without that tree went from 9 failures and
  18 collection errors to matching a full checkout exactly.
- audio_acceptance.py withdrawn: it imported an untracked module and raised
  ModuleNotFoundError on first use, and nothing in the node code imports it.
- The realtime node's removal is announced in 2.10.0 rather than backdated into
  2.9.0, which recorded the shim that 2.9.0 actually shipped.

Plan: .dev/plans/2026-09-27-hide-internal-dev-process-from-repo.md (2026-09-27)
2026-09-27 01:45:18 +03:00
WildAi 08df29df25 feat: auto-detect config_name, drop VibeVoice-Large (v2.7.0)
Config auto-detection:
- config_detect: read the checkpoint embedding shape as an architecture
  fingerprint (header-only for safetensors, reuses the open GGUF reader);
  7B=[152064,3584], 1.5B=[151936,1536]; orientation-agnostic for
  shape-reversed GGUF files
- Auto-detect is the new default config_name; resolves the family before
  any heavy load, or fails fast with an actionable error (.bin/.pt and
  unknown families cannot be fingerprinted)
- an explicit config_name that contradicts the weights self-corrects to
  the detected family with one WARNING (reconcile_config)
- loader: friendly shape pre-check in _apply_state_dict names the
  offending tensors and hints at config_name instead of torch's raw
  size-mismatch RuntimeError

Dropdown dedup:
- VibeVoice-Large removed from config_name options (duplicate of 7B);
  kept as a legacy alias so saved workflows still load (normalize at node
  + loader entry; validate_inputs(**kwargs) override makes core skip its
  combo-membership check)
- node resolves Auto-detect before computing the cache identity so the
  request key matches the consumer's bundle-derived key (no per-run churn)

Console noise:
- demote ~30 internal INFO logs to DEBUG across loaders/patcher/registry
- drop two stray tie_weights prints; tied lm_head.weight no longer warned
  as missing (expected under tie_word_embeddings)

Tests: 1060 passed / 5 pre-existing failures / 4 skipped
2026-08-27 20:33:19 +03:00
WildAi e20b4fd9d8 feat: external model input via VibeVoiceLoadExternalModel node (v2.2.0)
Add a dedicated loader node that loads VibeVoice checkpoints from user-provided .safetensors/.pt files in models/diffusion_models, since ComfyUI's stock Load Diffusion Model cannot detect VibeVoice architecture. The loader emits a VIBEVOICE_MODEL custom type bundle (model/processor/config/state_dict) consumed by an optional external_model input on the TTS, Realtime, and ASR nodes.

- modules/custom_types.py: VibeVoiceModel = io.Custom("VIBEVOICE_MODEL")

- modules/external_loader.py: sidecar config/preprocessor/tokenizer resolution + load_external_vibevoice_model() (TTS/streaming) and load_external_vibevoice_asr_model() (ASR); CPU-first load, in-memory state-dict injection, dtype cast, optional 4-bit quant (TTS only), SageAttention

- nodes/external_loader_node.py: VibeVoiceLoadExternalModel node

- modules/generation.py: ExternalVibeVoiceModelHandler + load_vibevoice_from_external()

- modules/asr_generation.py: ExternalVibeVoiceASRModelHandler + load_asr_from_external()

- nodes/tts_node.py, realtime_node.py, asr_node.py: optional external_model input with kind guards (streaming/ASR/TTS mismatch rejection)

- README: Loading External Models section + v2.2.0 changelog

- example_workflows/VibeVoice_external_model_example.json

- docs/plans/2026-08-16-external-model-input.md (status: COMPLETED)

Tests: +122 new/extended (test_custom_types, test_external_loader, test_external_loader_node, plus extensions to node_schema, generation, asr_generation, realtime_node, asr_node, patcher_behavioral, integration, docs_consistency, workflow, imports, extension). Full suite: 687 passed, 5 pre-existing failures, 4 skipped.
2026-08-16 23:59:04 +03:00