larsupb/main carried the sampler rework, released and published to the Comfy
Registry as 2.3.0 on 2026-06-22. This branch had independently claimed 2.3.0
for the delta-space merge work, so that release is renumbered 2.4.0 and the
published 2.3.0 entry is kept intact below it in the changelog.
- README: both sides added a `## Changelog` at the same spot. Resolved by
keeping the sampler entry as 2.3.0 and moving the merge work to 2.4.0.
- pyproject.toml / __init__.py -> 2.4.0.
- tests/conftest.py: mock comfy.sample and latent_preview, which the incoming
sampler code imports. Without them collection failed for 6 test files --
`comfy` is a MagicMock, not a package, so every submodule needs an explicit
sys.modules entry.
Suite verified at 238 passed, both in-tree and from a copy outside the ComfyUI
tree where `import comfy` raises ModuleNotFoundError.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`pytest tests/` previously crashed during collection and 7 of 17 test files
were dead: 61 tests were reachable, all via ad-hoc standalone scripts. Now a
bare `pytest` collects everything and passes 238 tests with ComfyUI absent
(verified by running the suite from outside the ComfyUI tree, where
`import comfy` raises ModuleNotFoundError).
Import structure:
- Drop tests/__init__.py. With it, pytest walks up to the project root's
__init__.py -- the ComfyUI node entry point -- and imports ComfyUI before
any test runs.
- Import project code as `src.<module>` instead of putting src/ on sys.path
and importing bare `merge.algorithms` / `validation` / `types`. Modules in
src/ use package-relative imports (`from ..types import ...`) that cannot
resolve when loaded top-level, and `types` collided with the stdlib module.
Same change for the mock.patch targets in test_algorithms.
- Consolidate conftest.py in tests/, mocking comfy, folder_paths,
comfy_extras and nodes. It stays in tests/ rather than the project root
because pytest imports a root-level conftest as part of the root package,
executing the ComfyUI entry point.
- Guard the script-style runners behind `if __name__ == "__main__":` so they
no longer sys.exit() during collection. Those files still run standalone.
- Drop run_pytest.py: a mocking wrapper made redundant by conftest, unused
and pointing at an unresolvable default path.
Bugs the dead tests were hiding:
- validators: the INCOMPATIBLE_DIMENSIONS check sat after the `continue` that
skips the reference tensor, so a lone LoRA with mismatched up/down ranks
passed validation unchecked. It is a per-LoRA check and now runs for every
entry.
- decomposition: __init__ exported a QRDecomposer that exists nowhere, so
`import src.decomposition` raised ImportError. Export and tests removed.
Stale expectations corrected:
- return_statistics is a constructor argument, not a decompose() kwarg.
- The zero-matrix rank guard only applies under dynamic rank selection; the
test now exercises that path, plus a new case pinning fixed-rank behavior.
- `reconstruction_error < 0.5` for a rank-10 truncation of a random 100x50
Gaussian is unreachable -- the optimum is 0.7557 and the decomposer hits
0.7568. Assert near-optimality instead, and add a genuinely low-rank case
that reconstructs to 0.003.
- sym/asym distributions differ only by float32 rounding (~5e-7), below the
default atol of 1e-8.
RUN_TESTS.md is rewritten against the real setup: correct interpreter path,
the two test-file styles, the import rules for adding tests, and a per-file
coverage table. It no longer documents test_gradient_analyzer_integration.py,
which is not in the repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Extends the delta-space merge path beyond the GTA family and hardens the
large-layer code paths that OOM'd on 8 GB cards.
- Merger node: route slerp/nuslerp/karcher/nearswap and SCE through the
delta-space path (reconstruct each LoRA's full delta, merge, refactor),
instead of merging up/down factors separately and injecting meaningless
up_i @ down_j cross-terms. Serialize these on CUDA like GTA.
- Merger node: `offload_models` toggle (default on) evicts resident
DIT/CLIP/VAE from VRAM before a CUDA merge; ComfyUI reloads them lazily.
- gta: add memory-frugal `karcher_delta_merge` and `sce_delta_merge` that
consume the delta list in place instead of stacking [N, out, in]; the
stock mergekit paths make ~3N full-tensor copies and OOM large layers.
- gta: pick magnitude / magnitude_outliers / SCE-select thresholds from a
GPU histogram above 16M elements, avoiding the int64 argsort + sort
workspace (the OOM behind TIES/Breadcrumbs) without the ~300x host
kthvalue penalty. Small tensors keep exact top-k for mergekit parity.
- gta: cap della chunks by element budget as well as rows, and build
per-row ranks with a single int32 `scatter_` instead of a second int64
argsort -- wide layers (e.g. KREA2 mlp [16384, 6144]) no longer need a
~1 GiB contiguous allocation.
- lora_save: sanitize tensors before saving; refactoring produces
transposed/sliced views that safetensors refuses to serialize.
- Tests for interp fidelity/integration, save sanitization, VRAM offload
and the new sparsify paths, plus RUN_TESTS.md and a .gitignore.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
average_weights (formerly normalize) ON computes a weighted AVERAGE
(÷Σweights), which silently halves stacked LoRAs -- the original "normalize
kills my LoRA" complaint. Native ComfyUI stacks LoRAs additively
(W + Σ strengthᵢ·Δᵢ), so OFF (additive sum) is the least-surprising default.
Flipped the widget + signature defaults True->False on Task Arithmetic,
TIES, DARE, DELLA (Breadcrumbs was already False). Also fixes the latent
`average_weights: bool = 0.5` default on Task Arithmetic. Tooltip updated to
say OFF matches ComfyUI stacking; turn ON only to blend/interpolate.
Safe for saved workflows: widget values serialize positionally, so existing
nodes keep their stored value; only newly-added nodes get the new default.
New guard test asserts the default is False on all GTA method nodes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
In this delta-space pipeline Linear is identical to Task Arithmetic at
default settings (both: no trim, no sign consensus, weighted combine of
deltas). Linear only exposed fewer knobs, so it added menu clutter with
no distinct behavior. Removed the LinearMergeMethod node class and its
three registrations; updated the widget-name guard test.
The "linear" GTA mode itself is kept — it is still referenced by the
dispatcher, algorithms registry, mergekit_utils and the experimental
checkpoint merger. Only the user-facing node is gone.
Note: saved workflows using PM Linear will load it as an undefined node;
replace with PM Task Arithmetic (identical output at defaults).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The test imports lora_mergekit_merge, which imports comfy.*; a script only
puts its own dir on sys.path, so `import comfy` failed unless run from a
cwd that happened to expose it. Add the ComfyUI root to sys.path explicitly
so the test runs from any working directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_della_magprune materialized several full-layer temporaries (two int64
argsorts + a stack of fp32 buffers), peaking ~2.1 GB for one FLUX
[21504,3072] delta -- ~11x the delta. On an 8 GB card with a resident
diffusion model this OOMs (della, unlike ties which does one argsort,
sits just over the edge).
della's ranking is per-row (argsort(dim=1)), so rows are independent and
block-wise processing is exact -- only the bernoulli draw order changes,
and della is stochastic pruning anyway. Tensors with rows <= 4096 keep
the single-shot whole-tensor path so they stay bit-for-bit identical to
mergekit (the parity unit test relies on that); larger layers stream in
row blocks with the mask applied in place and a global l1/l2/linf rescale.
Measured: single FLUX delta peak 2114 MB -> 566 MB; della+ties merge that
OOM'd on a genuinely loaded GPU (5.9 GB resident, ~1.3 GB free) now fits.
Density and l1-norm preservation verified; 4 new characterization tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
gta_merge stacked all N deltas into [N,out,in] plus several full copies,
peaking at ~3.5 GB for a single 21504x3072 FLUX layer (~13x the delta). It now
streams over a list: one accumulator for the elected sign, then numerator/
divisor accumulators, freeing each input delta as consumed. Peak ~1.7 GB/key,
independent of N. Preserves mergekit parity incl. the density<1 sparsified
sign vote (3-state torch.sign so sparsified-out zeros are excluded). gta_merge
now consumes its deltas list (documented); behavior test snapshots first.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>