Loading
- GGUF install is now two passes: a metadata pass that decides each
tensor's disposition, then an install pass that assigns residents in
place and streams dense tensors. The whole checkpoint is no longer
buffered in a dict alongside the model being built.
- dequantize_reader_tensor takes a target dtype, so dequant-at-load
writes straight into the destination parameter and the fp32
intermediate is never allocated.
- A bundle whose heavy fields were released is rebuilt from its recorded
source_path instead of failing the consumer.
- Host memory is released after install.
Numerics
- Q8_0 dequant computes in fp32 so the result is rounded once, at the
final cast. The activation-dtype path was reverted: it rounded twice
and moved stored weights.
- Removed a redundant weight-sized copy from the dequant kernel.
Bitwise-identical, ~1.16x.
- Precision gates are bitwise rather than tolerance-based.
Docs and tests
- README condensed; changelog moved to CHANGELOG.md.
- Third-party project references removed from source comments.
- Tests no longer assert README prose; the e2e smoke contract follows
the developer script to its new location and skips when absent.
- Version guard reads CHANGELOG.md.