8 Commits
Author SHA1 Message Date
WildAi cfa57c8738 feat: streaming quant load, single-rounding dequant, release hygiene (v2.11.0)
Loading
- GGUF install is now two passes: a metadata pass that decides each
  tensor's disposition, then an install pass that assigns residents in
  place and streams dense tensors. The whole checkpoint is no longer
  buffered in a dict alongside the model being built.
- dequantize_reader_tensor takes a target dtype, so dequant-at-load
  writes straight into the destination parameter and the fp32
  intermediate is never allocated.
- A bundle whose heavy fields were released is rebuilt from its recorded
  source_path instead of failing the consumer.
- Host memory is released after install.

Numerics
- Q8_0 dequant computes in fp32 so the result is rounded once, at the
  final cast. The activation-dtype path was reverted: it rounded twice
  and moved stored weights.
- Removed a redundant weight-sized copy from the dequant kernel.
  Bitwise-identical, ~1.16x.
- Precision gates are bitwise rather than tolerance-based.

Docs and tests
- README condensed; changelog moved to CHANGELOG.md.
- Third-party project references removed from source comments.
- Tests no longer assert README prose; the e2e smoke contract follows
  the developer script to its new location and skips when absent.
- Version guard reads CHANGELOG.md.
2026-09-28 18:25:19 +03:00
WildAi 29488e4a1f feat: unload-on-change eviction + quant-resident runtime for GGUF and quantized safetensors (v2.4.0 -> v2.5.0)
Model management (v2.4.0):
- single-active-per-family registry releases the previous model fully
  (RAM, VRAM, ComfyUI current_loaded_models) before a new one loads
- file-identity cache keys (basename+mtime+size+attention+q4+dtype)
  prevent cross-file collisions; ASR request keys share the consumer
  namespace so identical re-runs never spuriously evict

Quant-resident runtime (v2.5.0):
- GGUF weights stay raw-block resident end-to-end: uint8 parameters,
  per-matmul dequant kernels for Q8_0/Q4_K/Q5_K/Q6_K pinned bitwise to
  the gguf-py oracle; F32/F16/BF16 pass through at native dtype via
  zero-copy views; load-time RAM spike (~2x float size) eliminated
- quantized safetensors via *.comfy_quant metadata: rotated ConvRot
  INT8 residents through comfy-kitchen, plain rowwise int8 / fp8
  e4m3+e5m2 / int8_blockwise dequant-at-load (per-row, scalar, and
  per-gs-block scale layouts); unsupported formats hard-fail with
  actionable errors; dense gate rejects unplanned quant storages;
  rotated non-Linear targets (embeddings) fail with re-export guidance
- dtype casts filter quant-resident storage (_quant_resident markers,
  fp32 weight_scale protection); SageAttention wrapper resolves
  activation dtype per module kind, fixing uint8-weight crash

Tests: 937 passed / 5 pre-existing failures / 4 skipped
2026-08-26 12:17:24 +03:00
WildAi e20b4fd9d8 feat: external model input via VibeVoiceLoadExternalModel node (v2.2.0)
Add a dedicated loader node that loads VibeVoice checkpoints from user-provided .safetensors/.pt files in models/diffusion_models, since ComfyUI's stock Load Diffusion Model cannot detect VibeVoice architecture. The loader emits a VIBEVOICE_MODEL custom type bundle (model/processor/config/state_dict) consumed by an optional external_model input on the TTS, Realtime, and ASR nodes.

- modules/custom_types.py: VibeVoiceModel = io.Custom("VIBEVOICE_MODEL")

- modules/external_loader.py: sidecar config/preprocessor/tokenizer resolution + load_external_vibevoice_model() (TTS/streaming) and load_external_vibevoice_asr_model() (ASR); CPU-first load, in-memory state-dict injection, dtype cast, optional 4-bit quant (TTS only), SageAttention

- nodes/external_loader_node.py: VibeVoiceLoadExternalModel node

- modules/generation.py: ExternalVibeVoiceModelHandler + load_vibevoice_from_external()

- modules/asr_generation.py: ExternalVibeVoiceASRModelHandler + load_asr_from_external()

- nodes/tts_node.py, realtime_node.py, asr_node.py: optional external_model input with kind guards (streaming/ASR/TTS mismatch rejection)

- README: Loading External Models section + v2.2.0 changelog

- example_workflows/VibeVoice_external_model_example.json

- docs/plans/2026-08-16-external-model-input.md (status: COMPLETED)

Tests: +122 new/extended (test_custom_types, test_external_loader, test_external_loader_node, plus extensions to node_schema, generation, asr_generation, realtime_node, asr_node, patcher_behavioral, integration, docs_consistency, workflow, imports, extension). Full suite: 687 passed, 5 pre-existing failures, 4 skipped.
2026-08-16 23:59:04 +03:00
WildAi 62912003f4 voice bleeding fix, audio quality, input speakers tags, zero-shot voices 2025-09-24 17:42:30 +03:00
WildAi d5d17c87bc SageAttention support, fixes 2025-09-03 11:42:43 +03:00
WildAi 3345eadab8 model path update, fixes 2025-09-01 11:57:35 +03:00
WildAi 8e1fbb8e2d small fixes 2025-08-28 15:35:04 +03:00
WildAi 13b70cf0cc init examples 2025-08-27 15:53:34 +03:00