Multiple improvements and fixes across GGUF loading and model patching:
- dequant.py: Fix index dtype for IQ4 dequant gathers by casting indices to int64 to avoid dtype issues.
- loader.py:
- Extend TXT_ARCH_LIST with gemma3.
- Fix get_field to extract scalar values via .item().
- Add get_gguf_metadata to collect simple GGUF metadata (string/int/float/bool).
- Change gguf_sd_loader to return (state_dict, extra) where extra includes arch_str and metadata.
- Dequantize 1D BF16 tensors to float32 to avoid incorrect quantization for bias/1D params.
- Add GEMMA3_SD_MAP and gemma3_norm_corrections to reverse a Gemma3-specific norm offset (apply -1.0 correction and dequantize if needed).
- Add gguf_gemma3_tokenizer_loader to reconstruct a SentencePiece tokenizer from GGUF metadata.
- Update gguf_clip_loader to handle gemma3: map keys, apply norm corrections, and produce tokenizer data.
- Ensure gguf_mmproj_loader and other callers unpack the new gguf_sd_loader return value.
- nodes.py:
- Improve GGUFModelPatcher mmap handling: track named modules to unmap, add pin_weight_to_device to safely move modules when releasing mmap, clear tracking after release.
- Pass GGUF metadata into comfy.sd.load_diffusion_model_state_dict when supported, and add error checks when loading fails.
These changes add Gemma3 model/tokenizer support, fix dtype and BF16 edge cases, and improve low-memory mmap/unmap handling for safer weight pinning and loading.