Loading
- GGUF install is now two passes: a metadata pass that decides each
tensor's disposition, then an install pass that assigns residents in
place and streams dense tensors. The whole checkpoint is no longer
buffered in a dict alongside the model being built.
- dequantize_reader_tensor takes a target dtype, so dequant-at-load
writes straight into the destination parameter and the fp32
intermediate is never allocated.
- A bundle whose heavy fields were released is rebuilt from its recorded
source_path instead of failing the consumer.
- Host memory is released after install.
Numerics
- Q8_0 dequant computes in fp32 so the result is rounded once, at the
final cast. The activation-dtype path was reverted: it rounded twice
and moved stored weights.
- Removed a redundant weight-sized copy from the dequant kernel.
Bitwise-identical, ~1.16x.
- Precision gates are bitwise rather than tolerance-based.
Docs and tests
- README condensed; changelog moved to CHANGELOG.md.
- Third-party project references removed from source comments.
- Tests no longer assert README prose; the e2e smoke contract follows
the developer script to its new location and skips when absent.
- Version guard reads CHANGELOG.md.
VibeVoice TTS already routes realtime models correctly, so nodes/realtime_node.py
and the VibeVoiceRealtime ID are removed along with their shim tests and the legacy
workflow fixture. That node only existed in the unreleased 2.x line; origin/main is
still 1.5.0, so no published workflow referenced it.
- VibeVoice TTS and VibeVoice External Loader -> WMNodes/sound/tts
- VibeVoice ASR -> WMNodes/sound/asr (node IDs unchanged, only menu location moves)
- Tests no longer read local-only files: the transformers range guard now checks a
measured-version literal, and the GPU-audit fixtures skip when the local
diagnostic script is absent. A clone without that tree went from 9 failures and
18 collection errors to matching a full checkout exactly.
- audio_acceptance.py withdrawn: it imported an untracked module and raised
ModuleNotFoundError on first use, and nothing in the node code imports it.
- The realtime node's removal is announced in 2.10.0 rather than backdated into
2.9.0, which recorded the shim that 2.9.0 actually shipped.
Plan: .dev/plans/2026-09-27-hide-internal-dev-process-from-repo.md (2026-09-27)
Config auto-detection:
- config_detect: read the checkpoint embedding shape as an architecture
fingerprint (header-only for safetensors, reuses the open GGUF reader);
7B=[152064,3584], 1.5B=[151936,1536]; orientation-agnostic for
shape-reversed GGUF files
- Auto-detect is the new default config_name; resolves the family before
any heavy load, or fails fast with an actionable error (.bin/.pt and
unknown families cannot be fingerprinted)
- an explicit config_name that contradicts the weights self-corrects to
the detected family with one WARNING (reconcile_config)
- loader: friendly shape pre-check in _apply_state_dict names the
offending tensors and hints at config_name instead of torch's raw
size-mismatch RuntimeError
Dropdown dedup:
- VibeVoice-Large removed from config_name options (duplicate of 7B);
kept as a legacy alias so saved workflows still load (normalize at node
+ loader entry; validate_inputs(**kwargs) override makes core skip its
combo-membership check)
- node resolves Auto-detect before computing the cache identity so the
request key matches the consumer's bundle-derived key (no per-run churn)
Console noise:
- demote ~30 internal INFO logs to DEBUG across loaders/patcher/registry
- drop two stray tie_weights prints; tied lm_head.weight no longer warned
as missing (expected under tie_word_embeddings)
Tests: 1060 passed / 5 pre-existing failures / 4 skipped