larsupb/main carried the sampler rework, released and published to the Comfy
Registry as 2.3.0 on 2026-06-22. This branch had independently claimed 2.3.0
for the delta-space merge work, so that release is renumbered 2.4.0 and the
published 2.3.0 entry is kept intact below it in the changelog.
- README: both sides added a `## Changelog` at the same spot. Resolved by
keeping the sampler entry as 2.3.0 and moving the merge work to 2.4.0.
- pyproject.toml / __init__.py -> 2.4.0.
- tests/conftest.py: mock comfy.sample and latent_preview, which the incoming
sampler code imports. Without them collection failed for 6 test files --
`comfy` is a MagicMock, not a package, so every submodule needs an explicit
sys.modules entry.
Suite verified at 238 passed, both in-tree and from a copy outside the ComfyUI
tree where `import comfy` raises ModuleNotFoundError.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Bumps from 2.2.5, which was released on 2026-06-22; the 51 commits since are
unreleased. Minor rather than patch: a removed node, three renamed widgets and
a flipped default all break saved workflows.
- pyproject.toml -> 2.3.0
- __init__.py version_code -> [2, 3, 0]. It still read [2, 2, 4] and so had
been printing a stale version on node load since the 2.2.5 release.
- README: add a Changelog section covering the delta-space interpolation work,
the VRAM/OOM fixes, the breaking UX changes and the test-suite overhaul.
Correct the delta-space note, which was labelled v2.2.5 although that work
landed after the 2.2.5 release.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`pytest tests/` previously crashed during collection and 7 of 17 test files
were dead: 61 tests were reachable, all via ad-hoc standalone scripts. Now a
bare `pytest` collects everything and passes 238 tests with ComfyUI absent
(verified by running the suite from outside the ComfyUI tree, where
`import comfy` raises ModuleNotFoundError).
Import structure:
- Drop tests/__init__.py. With it, pytest walks up to the project root's
__init__.py -- the ComfyUI node entry point -- and imports ComfyUI before
any test runs.
- Import project code as `src.<module>` instead of putting src/ on sys.path
and importing bare `merge.algorithms` / `validation` / `types`. Modules in
src/ use package-relative imports (`from ..types import ...`) that cannot
resolve when loaded top-level, and `types` collided with the stdlib module.
Same change for the mock.patch targets in test_algorithms.
- Consolidate conftest.py in tests/, mocking comfy, folder_paths,
comfy_extras and nodes. It stays in tests/ rather than the project root
because pytest imports a root-level conftest as part of the root package,
executing the ComfyUI entry point.
- Guard the script-style runners behind `if __name__ == "__main__":` so they
no longer sys.exit() during collection. Those files still run standalone.
- Drop run_pytest.py: a mocking wrapper made redundant by conftest, unused
and pointing at an unresolvable default path.
Bugs the dead tests were hiding:
- validators: the INCOMPATIBLE_DIMENSIONS check sat after the `continue` that
skips the reference tensor, so a lone LoRA with mismatched up/down ranks
passed validation unchecked. It is a per-LoRA check and now runs for every
entry.
- decomposition: __init__ exported a QRDecomposer that exists nowhere, so
`import src.decomposition` raised ImportError. Export and tests removed.
Stale expectations corrected:
- return_statistics is a constructor argument, not a decompose() kwarg.
- The zero-matrix rank guard only applies under dynamic rank selection; the
test now exercises that path, plus a new case pinning fixed-rank behavior.
- `reconstruction_error < 0.5` for a rank-10 truncation of a random 100x50
Gaussian is unreachable -- the optimum is 0.7557 and the decomposer hits
0.7568. Assert near-optimality instead, and add a genuinely low-rank case
that reconstructs to 0.003.
- sym/asym distributions differ only by float32 rounding (~5e-7), below the
default atol of 1e-8.
RUN_TESTS.md is rewritten against the real setup: correct interpreter path,
the two test-file styles, the import rules for adding tests, and a per-file
coverage table. It no longer documents test_gradient_analyzer_integration.py,
which is not in the repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Extends the delta-space merge path beyond the GTA family and hardens the
large-layer code paths that OOM'd on 8 GB cards.
- Merger node: route slerp/nuslerp/karcher/nearswap and SCE through the
delta-space path (reconstruct each LoRA's full delta, merge, refactor),
instead of merging up/down factors separately and injecting meaningless
up_i @ down_j cross-terms. Serialize these on CUDA like GTA.
- Merger node: `offload_models` toggle (default on) evicts resident
DIT/CLIP/VAE from VRAM before a CUDA merge; ComfyUI reloads them lazily.
- gta: add memory-frugal `karcher_delta_merge` and `sce_delta_merge` that
consume the delta list in place instead of stacking [N, out, in]; the
stock mergekit paths make ~3N full-tensor copies and OOM large layers.
- gta: pick magnitude / magnitude_outliers / SCE-select thresholds from a
GPU histogram above 16M elements, avoiding the int64 argsort + sort
workspace (the OOM behind TIES/Breadcrumbs) without the ~300x host
kthvalue penalty. Small tensors keep exact top-k for mergekit parity.
- gta: cap della chunks by element budget as well as rows, and build
per-row ranks with a single int32 `scatter_` instead of a second int64
argsort -- wide layers (e.g. KREA2 mlp [16384, 6144]) no longer need a
~1 GiB contiguous allocation.
- lora_save: sanitize tensors before saving; refactoring produces
transposed/sliced views that safetensors refuses to serialize.
- Tests for interp fidelity/integration, save sanitization, VRAM offload
and the new sparsify paths, plus RUN_TESTS.md and a .gitignore.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
average_weights (formerly normalize) ON computes a weighted AVERAGE
(÷Σweights), which silently halves stacked LoRAs -- the original "normalize
kills my LoRA" complaint. Native ComfyUI stacks LoRAs additively
(W + Σ strengthᵢ·Δᵢ), so OFF (additive sum) is the least-surprising default.
Flipped the widget + signature defaults True->False on Task Arithmetic,
TIES, DARE, DELLA (Breadcrumbs was already False). Also fixes the latent
`average_weights: bool = 0.5` default on Task Arithmetic. Tooltip updated to
say OFF matches ComfyUI stacking; turn ON only to blend/interpolate.
Safe for saved workflows: widget values serialize positionally, so existing
nodes keep their stored value; only newly-added nodes get the new default.
New guard test asserts the default is False on all GTA method nodes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
In this delta-space pipeline Linear is identical to Task Arithmetic at
default settings (both: no trim, no sign consensus, weighted combine of
deltas). Linear only exposed fewer knobs, so it added menu clutter with
no distinct behavior. Removed the LinearMergeMethod node class and its
three registrations; updated the widget-name guard test.
The "linear" GTA mode itself is kept — it is still referenced by the
dispatcher, algorithms registry, mergekit_utils and the experimental
checkpoint merger. Only the user-facing node is gone.
Note: saved workflows using PM Linear will load it as an undefined node;
replace with PM Task Arithmetic (identical output at defaults).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The test imports lora_mergekit_merge, which imports comfy.*; a script only
puts its own dir on sys.path, so `import comfy` failed unless run from a
cwd that happened to expose it. Add the ComfyUI root to sys.path explicitly
so the test runs from any working directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
TDD plan for ties->sign_consensus, normalize->average_weights,
lambda_->output_scale. Naming-only; internal keys unchanged. Occurrence
counts for each replace_all verified against the source.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rename ties -> sign_consensus, normalize -> average_weights,
lambda_ -> output_scale. Naming only; internal keys unchanged, so
downstream plumbing is untouched. Positional widgets_values makes the
renames safe for saved graph workflows.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Display-only; no widget keys renamed, so saved workflows keep their
values (only visible titles/tooltips change).
- Drop the misleading "(Mergekit)" suffix from the GTA-family nodes
(Merger, Della, TIES, DARE, Linear, Task Arithmetic, Breadcrumbs) --
their UNet path runs our own gta.py now. Slerp/NuSlerp/NearSwap/SCE/
KArcher keep the suffix (still genuinely mergekit).
- Fix dead display keys: "PM TIES"/"PM DARE" never matched the class
keys "PM Ties"/"PM Dare", so those titles never applied.
- normalize tooltip now spells out the footgun: ON = weighted average
(strengths as ratios, two 1.0 LoRAs -> ~50%), OFF = additive sum.
- ties tooltip explains the sign-consensus switch (vs the _linear variant).
- svd_rank tooltip leads with "-1 = auto".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_della_magprune materialized several full-layer temporaries (two int64
argsorts + a stack of fp32 buffers), peaking ~2.1 GB for one FLUX
[21504,3072] delta -- ~11x the delta. On an 8 GB card with a resident
diffusion model this OOMs (della, unlike ties which does one argsort,
sits just over the edge).
della's ranking is per-row (argsort(dim=1)), so rows are independent and
block-wise processing is exact -- only the bernoulli draw order changes,
and della is stochastic pruning anyway. Tensors with rows <= 4096 keep
the single-shot whole-tensor path so they stay bit-for-bit identical to
mergekit (the parity unit test relies on that); larger layers stream in
row blocks with the mask applied in place and a global l1/l2/linf rescale.
Measured: single FLUX delta peak 2114 MB -> 566 MB; della+ties merge that
OOM'd on a genuinely loaded GPU (5.9 GB resident, ~1.3 GB free) now fits.
Density and l1-norm preservation verified; 4 new characterization tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
process_key materializes full dense deltas for the GTA family; the 8-way
thread pool ran 8 concurrently, multiplying peak VRAM ~8x and exhausting an
8 GB card with a resident FLUX model -- OOM, or a hard segfault under the
quantized-model CUDA context (surfacing at the first CUDA op in the branch,
the u@d matmul). Use 1 worker for GTA methods on CUDA (GPU-serial ~11s for
256 keys); the light non-GTA factored path keeps the 8-way pool.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
gta_merge stacked all N deltas into [N,out,in] plus several full copies,
peaking at ~3.5 GB for a single 21504x3072 FLUX layer (~13x the delta). It now
streams over a list: one accumulator for the elected sign, then numerator/
divisor accumulators, freeing each input delta as consumed. Peak ~1.7 GB/key,
independent of N. Preserves mergekit parity incl. the density<1 sparsified
sign vote (3-state torch.sign so sparsified-out zeros are excluded). gta_merge
now consumes its deltas list (documented); behavior test snapshots first.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds a refactor_method choice (energy_rSVD default, rSVD) to the PM Ties /
mergekit merge node, controlling how the merged dense delta is decomposed
back into a LoRA. Threaded through lora_mergekit -> merge -> process_key ->
merged_delta_to_lora, carried on merge_context, and reused by the parameter
sweep sampler. Full SVD is intentionally not selectable (the OOM/crash path).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The delta-space GTA path refactored the merged dense delta with a full
torch.linalg.svd. On a real (model-resident) run across the 8-worker merge
pool that is the fragile, heavy path: ~190 MB workspace and ~2.9 s per
3072x3072 layer via cuSOLVER, which OOMs / segfaults under VRAM pressure.
Replace it with torch.svd_lowrank (randomized SVD) computing only the top-r
components with energy-based dynamic rank: ~6 MB and ~4 ms for the same layer
(29x less memory, ~700x faster), and no cuSOLVER full-SVD driver. Also drops
the fragile utility.py import-stub machinery (svd_lowrank is pure torch).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Apply sqrt-magnitude weights to each up/down factor so the reconstructed
delta scales linearly by strength rather than strength**2. Sign is carried
on the up factor only. Lambda scaling moved to the reconstructed delta
(up factor only) so it also scales linearly, not quadratically. CLIP
layers get normalized weighted average with overall strength re-applied
once to the up factor.