113 Commits
Author SHA1 Message Date
larsupbandClaude Opus 5 dad8a4a5ee Merge larsupb/main: samplers (2.3.0) + delta-space merge work (2.4.0)
larsupb/main carried the sampler rework, released and published to the Comfy
Registry as 2.3.0 on 2026-06-22. This branch had independently claimed 2.3.0
for the delta-space merge work, so that release is renumbered 2.4.0 and the
published 2.3.0 entry is kept intact below it in the changelog.

- README: both sides added a `## Changelog` at the same spot. Resolved by
  keeping the sampler entry as 2.3.0 and moving the merge work to 2.4.0.
- pyproject.toml / __init__.py -> 2.4.0.
- tests/conftest.py: mock comfy.sample and latent_preview, which the incoming
  sampler code imports. Without them collection failed for 6 test files --
  `comfy` is a MagicMock, not a package, so every submodule needs an explicit
  sys.modules entry.

Suite verified at 238 passed, both in-tree and from a copy outside the ComfyUI
tree where `import comfy` raises ModuleNotFoundError.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 11:54:20 +02:00
larsupbandClaude Opus 5 fe92e95f65 release: 2.3.0 with changelog
Bumps from 2.2.5, which was released on 2026-06-22; the 51 commits since are
unreleased. Minor rather than patch: a removed node, three renamed widgets and
a flipped default all break saved workflows.

- pyproject.toml -> 2.3.0
- __init__.py version_code -> [2, 3, 0]. It still read [2, 2, 4] and so had
  been printing a stale version on node load since the 2.2.5 release.
- README: add a Changelog section covering the delta-space interpolation work,
  the VRAM/OOM fixes, the breaking UX changes and the test-suite overhaul.
  Correct the delta-space note, which was labelled v2.2.5 although that work
  landed after the 2.2.5 release.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 11:48:08 +02:00
larsupbandClaude Opus 5 7973203f38 test: run the whole suite without a ComfyUI installation
`pytest tests/` previously crashed during collection and 7 of 17 test files
were dead: 61 tests were reachable, all via ad-hoc standalone scripts. Now a
bare `pytest` collects everything and passes 238 tests with ComfyUI absent
(verified by running the suite from outside the ComfyUI tree, where
`import comfy` raises ModuleNotFoundError).

Import structure:
- Drop tests/__init__.py. With it, pytest walks up to the project root's
  __init__.py -- the ComfyUI node entry point -- and imports ComfyUI before
  any test runs.
- Import project code as `src.<module>` instead of putting src/ on sys.path
  and importing bare `merge.algorithms` / `validation` / `types`. Modules in
  src/ use package-relative imports (`from ..types import ...`) that cannot
  resolve when loaded top-level, and `types` collided with the stdlib module.
  Same change for the mock.patch targets in test_algorithms.
- Consolidate conftest.py in tests/, mocking comfy, folder_paths,
  comfy_extras and nodes. It stays in tests/ rather than the project root
  because pytest imports a root-level conftest as part of the root package,
  executing the ComfyUI entry point.
- Guard the script-style runners behind `if __name__ == "__main__":` so they
  no longer sys.exit() during collection. Those files still run standalone.
- Drop run_pytest.py: a mocking wrapper made redundant by conftest, unused
  and pointing at an unresolvable default path.

Bugs the dead tests were hiding:
- validators: the INCOMPATIBLE_DIMENSIONS check sat after the `continue` that
  skips the reference tensor, so a lone LoRA with mismatched up/down ranks
  passed validation unchecked. It is a per-LoRA check and now runs for every
  entry.
- decomposition: __init__ exported a QRDecomposer that exists nowhere, so
  `import src.decomposition` raised ImportError. Export and tests removed.

Stale expectations corrected:
- return_statistics is a constructor argument, not a decompose() kwarg.
- The zero-matrix rank guard only applies under dynamic rank selection; the
  test now exercises that path, plus a new case pinning fixed-rank behavior.
- `reconstruction_error < 0.5` for a rank-10 truncation of a random 100x50
  Gaussian is unreachable -- the optimum is 0.7557 and the decomposer hits
  0.7568. Assert near-optimality instead, and add a genuinely low-rank case
  that reconstructs to 0.003.
- sym/asym distributions differ only by float32 rounding (~5e-7), below the
  default atol of 1e-8.

RUN_TESTS.md is rewritten against the real setup: correct interpreter path,
the two test-file styles, the import rules for adding tests, and a per-file
coverage table. It no longer documents test_gradient_analyzer_integration.py,
which is not in the repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 11:41:31 +02:00
larsupbandClaude Opus 5 1ca216dcbb feat(merge): delta-space interpolation, VRAM offload, and OOM-safe kernels
Extends the delta-space merge path beyond the GTA family and hardens the
large-layer code paths that OOM'd on 8 GB cards.

- Merger node: route slerp/nuslerp/karcher/nearswap and SCE through the
  delta-space path (reconstruct each LoRA's full delta, merge, refactor),
  instead of merging up/down factors separately and injecting meaningless
  up_i @ down_j cross-terms. Serialize these on CUDA like GTA.
- Merger node: `offload_models` toggle (default on) evicts resident
  DIT/CLIP/VAE from VRAM before a CUDA merge; ComfyUI reloads them lazily.
- gta: add memory-frugal `karcher_delta_merge` and `sce_delta_merge` that
  consume the delta list in place instead of stacking [N, out, in]; the
  stock mergekit paths make ~3N full-tensor copies and OOM large layers.
- gta: pick magnitude / magnitude_outliers / SCE-select thresholds from a
  GPU histogram above 16M elements, avoiding the int64 argsort + sort
  workspace (the OOM behind TIES/Breadcrumbs) without the ~300x host
  kthvalue penalty. Small tensors keep exact top-k for mergekit parity.
- gta: cap della chunks by element budget as well as rows, and build
  per-row ranks with a single int32 `scatter_` instead of a second int64
  argsort -- wide layers (e.g. KREA2 mlp [16384, 6144]) no longer need a
  ~1 GiB contiguous allocation.
- lora_save: sanitize tensors before saving; refactoring produces
  transposed/sliced views that safetensors refuses to serialize.
- Tests for interp fidelity/integration, save sanitization, VRAM offload
  and the new sparsify paths, plus RUN_TESTS.md and a .gitignore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 11:13:13 +02:00
larsupb d6b5ad87e6 feat(nodes): average_weights toggle on slerp/nuslerp/karcher/nearswap 2026-07-21 00:57:42 +02:00
larsupb 48150ddd2e feat(merge): interp_delta_merge helper for delta-space slerp/nuslerp/karcher/nearswap 2026-07-21 00:54:33 +02:00
larsupbandClaude Opus 4.8 fbaf1e9ae7 docs(plan): delta-space + average_weights for slerp/nuslerp/karcher/nearswap
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 00:48:06 +02:00
larsupbandClaude Opus 4.8 ca17e1c2fc docs(spec): delta-space + average_weights for slerp/nuslerp/karcher/nearswap
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 00:41:58 +02:00
larsupbandClaude Opus 4.8 fd0578ffe5 docs(spec): VRAM offload before merge on PM LoRA Merger
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 15:04:16 +02:00
larsupbandClaude Opus 4.8 7a1d7851cd docs(readme): reflect renames, Linear removal, default flip, refactor_method
- Node title: "PM LoRA Merger (Mergekit)" -> "PM LoRA Merger".
- Merger params: _lambda -> output_scale; document refactor_method
  (energy_rSVD/rSVD), spectral_norm_scale, merge_clip; note device-aware
  worker count (GTA on CUDA = single worker).
- Merge methods: remove PM Linear; document the shared GTA controls
  (sign_consensus, average_weights default OFF, rescale_norm) with a
  migration note; fix DELLA (was malformed) and Breadcrumbs (bogus
  tie_method) and add per-method density/epsilon/gamma defaults.
- Task Arithmetic: note it matches ComfyUI's native additive stacking
  when average_weights is off.
- Parameter Sweep + Features: normalize -> average_weights; add a
  native delta-space GTA feature bullet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 18:54:09 +02:00
larsupbandClaude Opus 4.8 36df70a2b1 feat(nodes): default average_weights to OFF (additive, like ComfyUI)
average_weights (formerly normalize) ON computes a weighted AVERAGE
(÷Σweights), which silently halves stacked LoRAs -- the original "normalize
kills my LoRA" complaint. Native ComfyUI stacks LoRAs additively
(W + Σ strengthᵢ·Δᵢ), so OFF (additive sum) is the least-surprising default.

Flipped the widget + signature defaults True->False on Task Arithmetic,
TIES, DARE, DELLA (Breadcrumbs was already False). Also fixes the latent
`average_weights: bool = 0.5` default on Task Arithmetic. Tooltip updated to
say OFF matches ComfyUI stacking; turn ON only to blend/interpolate.

Safe for saved workflows: widget values serialize positionally, so existing
nodes keep their stored value; only newly-added nodes get the new default.
New guard test asserts the default is False on all GTA method nodes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 18:47:52 +02:00
larsupbandClaude Opus 4.8 bec8c3ec3d ux(nodes): remove redundant PM Linear node
In this delta-space pipeline Linear is identical to Task Arithmetic at
default settings (both: no trim, no sign consensus, weighted combine of
deltas). Linear only exposed fewer knobs, so it added menu clutter with
no distinct behavior. Removed the LinearMergeMethod node class and its
three registrations; updated the widget-name guard test.

The "linear" GTA mode itself is kept — it is still referenced by the
dispatcher, algorithms registry, mergekit_utils and the experimental
checkpoint merger. Only the user-facing node is gone.

Note: saved workflows using PM Linear will load it as an undefined node;
replace with PM Task Arithmetic (identical output at defaults).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 18:17:04 +02:00
larsupbandClaude Opus 4.8 bbc74069ab test(nodes): make widget-name test cwd-robust
The test imports lora_mergekit_merge, which imports comfy.*; a script only
puts its own dir on sys.path, so `import comfy` failed unless run from a
cwd that happened to expose it. Add the ComfyUI root to sys.path explicitly
so the test runs from any working directory.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:33:14 +02:00
larsupbandClaude Opus 4.8 1123befb2e ux(nodes): rename lambda_->output_scale widget on the Merger node
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:31:03 +02:00
larsupb 5b3d59c43a ux(nodes): rename ties->sign_consensus, normalize->average_weights widgets 2026-07-19 15:29:24 +02:00
larsupb 5b9c38d30f test(nodes): guard clearer widget names (ties/normalize/lambda_ renames) 2026-07-19 15:27:28 +02:00
larsupbandClaude Opus 4.8 999baae4a0 docs(plan): clearer merge-node widget names
TDD plan for ties->sign_consensus, normalize->average_weights,
lambda_->output_scale. Naming-only; internal keys unchanged. Occurrence
counts for each replace_all verified against the source.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 14:40:29 +02:00
larsupbandClaude Opus 4.8 b38b3f873a docs(spec): clearer widget names for merge nodes
Rename ties -> sign_consensus, normalize -> average_weights,
lambda_ -> output_scale. Naming only; internal keys unchanged, so
downstream plumbing is untouched. Positional widgets_values makes the
renames safe for saved graph workflows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 12:07:58 +02:00
larsupbandClaude Opus 4.8 ab00821e43 ux(nodes): clearer titles and tooltips for merge params
Display-only; no widget keys renamed, so saved workflows keep their
values (only visible titles/tooltips change).

- Drop the misleading "(Mergekit)" suffix from the GTA-family nodes
  (Merger, Della, TIES, DARE, Linear, Task Arithmetic, Breadcrumbs) --
  their UNet path runs our own gta.py now. Slerp/NuSlerp/NearSwap/SCE/
  KArcher keep the suffix (still genuinely mergekit).
- Fix dead display keys: "PM TIES"/"PM DARE" never matched the class
  keys "PM Ties"/"PM Dare", so those titles never applied.
- normalize tooltip now spells out the footgun: ON = weighted average
  (strengths as ratios, two 1.0 LoRAs -> ~50%), OFF = additive sum.
- ties tooltip explains the sign-consensus switch (vs the _linear variant).
- svd_rank tooltip leads with "-1 = auto".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 11:49:49 +02:00
larsupbandClaude Opus 4.8 13b66a06c6 perf(gta): chunk della sparsify over rows to bound VRAM
_della_magprune materialized several full-layer temporaries (two int64
argsorts + a stack of fp32 buffers), peaking ~2.1 GB for one FLUX
[21504,3072] delta -- ~11x the delta. On an 8 GB card with a resident
diffusion model this OOMs (della, unlike ties which does one argsort,
sits just over the edge).

della's ranking is per-row (argsort(dim=1)), so rows are independent and
block-wise processing is exact -- only the bernoulli draw order changes,
and della is stochastic pruning anyway. Tensors with rows <= 4096 keep
the single-shot whole-tensor path so they stay bit-for-bit identical to
mergekit (the parity unit test relies on that); larger layers stream in
row blocks with the mask applied in place and a global l1/l2/linf rescale.

Measured: single FLUX delta peak 2114 MB -> 566 MB; della+ties merge that
OOM'd on a genuinely loaded GPU (5.9 GB resident, ~1.3 GB free) now fits.
Density and l1-norm preservation verified; 4 new characterization tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 11:49:38 +02:00
larsupbandClaude Opus 4.8 abeec6129a fix: serialize GTA merge on CUDA to bound peak VRAM (fixes FLUX OOM/segfault)
process_key materializes full dense deltas for the GTA family; the 8-way
thread pool ran 8 concurrently, multiplying peak VRAM ~8x and exhausting an
8 GB card with a resident FLUX model -- OOM, or a hard segfault under the
quantized-model CUDA context (surfacing at the first CUDA op in the branch,
the u@d matmul). Use 1 worker for GTA methods on CUDA (GPU-serial ~11s for
256 keys); the light non-GTA factored path keeps the 8-way pool.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 02:31:16 +02:00
larsupbandClaude Opus 4.8 030ab02fe7 perf(gta): stream the merge instead of stacking (cut peak VRAM ~2x)
gta_merge stacked all N deltas into [N,out,in] plus several full copies,
peaking at ~3.5 GB for a single 21504x3072 FLUX layer (~13x the delta). It now
streams over a list: one accumulator for the elected sign, then numerator/
divisor accumulators, freeing each input delta as consumed. Peak ~1.7 GB/key,
independent of N. Preserves mergekit parity incl. the density<1 sparsified
sign vote (3-state torch.sign so sparsified-out zeros are excluded). gta_merge
now consumes its deltas list (documented); behavior test snapshots first.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 02:31:16 +02:00
larsupbandClaude Opus 4.8 036934f911 docs: note rSVD refactor deviation + refactor_method widget
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 02:14:10 +02:00
larsupbandClaude Opus 4.8 df041f81d9 feat(gta): expose refactor_method widget on the merge node
Adds a refactor_method choice (energy_rSVD default, rSVD) to the PM Ties /
mergekit merge node, controlling how the merged dense delta is decomposed
back into a LoRA. Threaded through lora_mergekit -> merge -> process_key ->
merged_delta_to_lora, carried on merge_context, and reused by the parameter
sweep sampler. Full SVD is intentionally not selectable (the OOM/crash path).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 02:12:55 +02:00
larsupbandClaude Opus 4.8 7c52c13e81 fix(gta): randomized SVD for merged-delta refactor (fixes OOM/segfault)
The delta-space GTA path refactored the merged dense delta with a full
torch.linalg.svd. On a real (model-resident) run across the 8-worker merge
pool that is the fragile, heavy path: ~190 MB workspace and ~2.9 s per
3072x3072 layer via cuSOLVER, which OOMs / segfaults under VRAM pressure.

Replace it with torch.svd_lowrank (randomized SVD) computing only the top-r
components with energy-based dynamic rank: ~6 MB and ~4 ms for the same layer
(29x less memory, ~700x faster), and no cuSOLVER full-SVD driver. Also drops
the fragile utility.py import-stub machinery (svd_lowrank is pure torch).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 02:12:55 +02:00
larsupb 7424fe0274 tests(blocks): rewrite TestApplySelection, add BlockSelection dict/resolve tests 2026-07-19 01:55:52 +02:00
larsupb 9ab7115695 refactor(LoraDecompose): use resolve_block_selection for index-based block selection 2026-07-19 01:54:54 +02:00
larsupb 271009d741 types: add BlockSelectionConfig alias 2026-07-19 01:54:28 +02:00
larsupb 40d060ee71 refactor(BlockSelector): remove lora_stack input, chain index-based configs 2026-07-19 01:54:14 +02:00
larsupb 8b76f73401 feat(blocks): add build_block_selection_dict and resolve_block_selection 2026-07-19 01:53:41 +02:00
larsupb 68069046b2 docs: add implementation plan for block selector lora_stack removal 2026-07-19 01:46:36 +02:00
larsupb 479099fb8c docs: add spec for block selector lora_stack removal 2026-07-19 01:44:26 +02:00
larsupb bfdc5fe5dc docs: mark delta-space GTA merge implemented 2026-07-19 01:33:16 +02:00
larsupb d0080d4a37 test(gta): expose gta module in gta_helpers for behavior tests 2026-07-19 01:31:37 +02:00
larsupb 4e44e5b67f test(gta): end-to-end behavioral tests encoding the normalize fix 2026-07-19 01:31:35 +02:00
larsupb 8e2a40d440 feat: route GTA merge family through delta-space path 2026-07-19 01:30:33 +02:00
larsupb 3c15e6d72d feat(gta): merged-delta -> LoRA refactor with energy-dynamic rank 2026-07-19 01:28:57 +02:00
larsupb 7456bee6a3 feat(gta): resolve per-mode rescale_norm defaults 2026-07-19 01:24:43 +02:00
larsupb 08ea66cae6 feat(gta): gta_merge entry point with mode config, parity vs mergekit 2026-07-19 01:23:39 +02:00
larsupb 2926caa600 feat(gta): sign election + disjoint merge with per-element normalize 2026-07-19 01:21:46 +02:00
larsupb e85ec67057 feat(gta): own sparsify primitives matching mergekit 2026-07-19 01:20:52 +02:00
larsupb 2f86588fad test: add standalone loader harness for gta module 2026-07-19 01:19:39 +02:00
larsupb 35426a39f8 fix(mergekit): strength linearization and correct lambda scaling
Apply sqrt-magnitude weights to each up/down factor so the reconstructed
delta scales linearly by strength rather than strength**2. Sign is carried
on the up factor only. Lambda scaling moved to the reconstructed delta
(up factor only) so it also scales linearly, not quadratically. CLIP
layers get normalized weighted average with overall strength re-applied
once to the up factor.
2026-07-19 01:16:35 +02:00
larsupb aafda5c809 feat: PM Block Selector + model-specific block nodes (KREA2, FLUX.2-Klein)
Add per-block, per-LoRA weighting to the LoRA PowerMerge pipeline:
- PM Block Selector: bind a BlockDefinition to one LoRA by stack index;
  chain outputs to weight multiple LoRAs, feed into PM LoRA Stack Decompose.
- PM KREA 2 Blocks: model-specific block definition for KREA2 LoRAs
  (diffusion_model.blocks.N + txtfusion + txtmlp pathways).
- PM FLUX.2.Klein Blocks: block definition for FLUX.2-Klein LoRAs
  (double_blocks + single_blocks streams).

Implementation:
- src/blocks.py: pure logic (key normalization, weight-string parsing,
  category/pathway matching, per-LoRA weight computation, up-factor scaling).
  Already existed, fully unit-tested in tests/test_blocks.py.
- src/nodes_block_selector.py: thin ComfyUI node wrappers around blocks.py.
- src/lora_decompose.py: optional BlockSelection input threads per-block
  weights through the decompose path; scales the up factor (linear delta
  scaling), weight 0 drops the LoRA from that key; cache-aware.
- tests/conftest.py: fix pytest_ignore_collect to exclude __init__.py and
  src/ from collection, enabling clean test isolation.

3 new nodes registered under LoRA PowerMerge.
2026-07-19 01:14:03 +02:00
larsupbandClaude Opus 4.8 acdf80b5da docs: implementation plan for delta-space GTA merge
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 00:40:35 +02:00
larsupb 205824a2b3 feat(blocks): KREA2 and FLUX.2-Klein definition builders 2026-07-19 00:39:10 +02:00
larsupb bd2b1b6390 feat(blocks): apply per-block weights to up factor (linear delta scaling) 2026-07-19 00:37:38 +02:00
larsupb ffc272b312 feat(blocks): per-LoRA weight computation, selection merge, index selection 2026-07-19 00:36:45 +02:00
larsupb 84c12e352d feat(blocks): category construction and per-key weight lookup 2026-07-19 00:35:53 +02:00
larsupb db9f94fb5d feat(blocks): key normalization and weight-string parsing 2026-07-19 00:34:19 +02:00