10 KiB
Blend Modes
There are so many blend modes now that it's probably pretty overwhelming. Also a lot of them are junk/don't work well but can't be removed without breaking existing workflows that might use them. So here is some incomplete, low-effort documentation!
Meta Modes
Every blend function takes at least three parameters: a, b and t (the ratio to blend). lerp(a, b, 0.25) would mean a * 0.75 + b * 0.25.
revWHATEVER- Just flips the inputs, so if we were doing something likelerp(a, b, 0.25)(a * 0.75 + b * 0.5) the blend would be applied likelerp(b, a, 0.25).normWHATEVER- Tries to scale the input to -1...1 and then uses a simple LERP to find a range for the output. Pretty much garbage.
Some modes also have suffixes. Not precisely a meta mode but this is probably the logical place to cover it.
_d1,_d2, etc - Indicates the mode will operate on that dimension. Dimension 1 in this case, which is typically channels in most latents._copysign_a- The blended result copies the sign from theaparameter._avoidsign_a- The blended result avoids the sign of theaparameter. In other words, negate A and then copy the sign from it to the result.- Digits like
_025- Usually indicates a multiplier.025would stand for0.25. _base_a- Mostly for CFG type blend modes. This means ifais cond andbis uncond, a ratio of 0 will give you cond. Normal CFG islerp(uncond, cond, ratio)so if ratio is 0 you get pure uncond.
Custom Blend Parameters
Nodes taking a blend mode use a selection but you can define modes as a string (whitespace is ignored so I recommend a multiline string widget) and force the connection to the mode parameter with the BlehCast node.
Custom definitions use this syntax: mode_name:param1=val1:param2=val2. Integer and float values don't require any special handling. Boolean parameters use true and false. Nullable parameters can use none. An empty list is (). Lists are comma-separated and a singleton list is specified like 1,. String literals (like mode names) should start with the caret, like ^whatever otherwise they will be interpreted as a blend mode. Finally, some blend modes are wrappers for other blend modes. If it exists, a single trailing underscore will be stripped from the parameters when they are passed to the nested blend function. Not very convenient to use, but it allows specifying parameters where the names may clash. I.E. some_mode:blend_mode=blah:blend_mode=other would let some_mode use the blend_mode parameter and then pass the second to the blah blend mode handler.
This is clunky/inconvenient and pretty limited but it does allow specifying custom parameters in most cases.
All blend modes support some common parameters:
rev- boolean. Examplelerp:rev=true. Flips the inputs.scale_multiplier- float. Rescales the blend ratio. Example:lerp:scale_multiplier=0.5.lerp(a, b, 1.0)would result inlerp(a, b, 0.5).invert_scale- float. Adjusts the blend ratio by doingscale_value - ratio. Example:lerp:invert_scale=1.0.lerp(a, b, 0.4)would result inlerp(a, b, 0.6)(1 - 0.4 == 0.6).fork_rng- boolean. Forks the random number generator state when calling the blend mode. Can be useful for probalistic blend modes likeproblerpwhich would perturb the RNG and affect stuff like noise used for ancestral sampling, changing your seed even if the ratio is tiny. Example:problerp:fork_rng=true
Realistically, the unique parameters for custom blend modes will probably never get documented. Unfortunately, you will need to read the source in latent_utils.py to find out what the options are.
Generally Useful
lerpand friends. Linear interpolation, the most common blend mode. CFG is also just LERP.slerp- Can sometimes be better than LERP for blending latents.inject- Simple addition.inject(a, b, 0.3)is justa + b * 0.3. Useful for adding stuff like a CFG diff.
Garbage/Redundant
bislerp_wrong- This is just LERP with the useless normalize.hslerp(anything starting with HSLERP).colorize- Literally just LERP.colordodge,difference,exclusion,glow,hardlight,linearlight,overlay,pinlight,reflect,screen,vividlight- Photoshop filter type modes. They are designed for images and assume certain value ranges for the input so they are essentually useless for blending latents.linear_dodge- Same asinject. This is just scaled addition.cosinesimilarity- I actually like using this but it is roughly just a worse SLERP.
Experimental/Exotic Modes
Note: Many of these are vibe-mathed, so the description explains what I was attempting to do and what I believe the mode does. I can't guarantee it is doing what it purports to because I don't always fully understand the math involved.
ortho- Orthogonal addition. You may want to specify the dimension parameters to control how the normalization happens. Example:ortho:start_dim=2:end_dim=3- For 4D latents, this would normalize over the height/width dimensions. The ortho blend function has many parameters, you will need to look at the source to see them.ortho_rescaled_lerpish- Orthogonal blending means the parallel component gets thrown away, in other words you may end up adding less of something than you expected. This mode tries to adjust the result to target something like the result of a LERP.ortho_dyn_lerp- Similar to the previous, except it calculates how muchbgot scaled down and LERPs to compensate.ortho_cfgandortho_cfg_base_a- Does ortho addition of the CFG diff.contrastive_ortho_cfg(and_base_a) - Mostly useful for CFG. Let's say we're in base A mode andais cond andbis uncond. The mode takes aa_ortho_scale(what is unique to cond) andb_ortho_scale(what is unique to uncond) parameters.b_ortho(AKA what is unique to uncond) gets _subtracted. Socontrastive_ortho_cfg:a_ortho_scale=0.0:b_ortho_scale=1.0would mean only subtract what is unique to uncond but don't scale up what is unique to cond. The reverse is also possible, only enhance what is unique to cond but don't subtract uncond's unique features. Compare this with normal CFG:cond + (cond - uncond) * scaleor in other wordscond + cond * scale - uncond * scale. Ifcondis 1 and uncond is 0 then the result would effectively becond + cond * scale- we scale up cond since there isn't a value on the uncond side to cancel it out.- Modes starting with
slice- Slices along a dimension. The ratio is the size of the dimension multiplied by the ratio. Example:slice_d1- slices dimension 1 (second dimension in zero-based dimension indexing).slice_d1(a, b, 0.5)would use the first 50% of channels fromaand the remaining ones fromb. Or maybe it's the other way around, I forget! These modes also have a_flipvariant which would make it so thearesult comes first or vice versa. The blend function supports various parameters from smoothing the result and controling the offset so if you wanted something like just the middle 25% of a dimension that is possible with custom parameters. - Modes starting with
wavelet.wavelet_b_hi_100_lo_0means 100% of the high frequency components ofband none of the low frequency.wavelet_b_hi_0_lo_100is the reverse. Can be interesting as a CFG function. The blend function supports many parameters such as setting the wavelet type and ratios. This requires wavelet support from thepytorch_waveletspackage which is unfortunately broken with recent Python versions and hasn't been updated in years. You can use my pull with a fix (or the repo it links to): https://github.com/fbcotter/pytorch_wavelets/pull/66 pct_limited_025- Limits the result to a maximum change (relative toaby default). This is a wrapper for other blend modes which defaults to LERP. Example:pct_limited_025:diff_limit=0.25:blend_mode=injectLet's say we doblend(0.5, 100.0, 1.0)Without limiting this is0.5 + 100.0 * 1.0which results in something like a 2,000% change. If we limit to adjustingaby 25% at most we get a limit of0.125(0.5 * 0.25) so the output is0.5 + 0.125 == 0.625. This blend function has various other parameters, for example to allow soft clamping, prevent sign flipping (which not limiting the actual value), calculating percentage change over a dimension rather than elementwise, etc.distro_aligned- By default matches the distribution ofbtoa.distro_aligned_resultaligns the output from the blend to the distribution ofa. The blend function supports many options for controlling which gets aligned to what else, how it occurs, etc. You will need to look at the source.gaussian_alignedandgaussian_aligned_result. The latter only aligns the result to the Gaussian distribution, the former aligns bothaandbbefore blending and then also aligns the result. SDXL latents (and probably most latent formats?) should be in the Gaussian distribution (zero mean, std 1) so using this as a CFG function actually works pretty well.rms_interpolation_lerpsign- By default, this is roughlylerp(a**2, b**2, ratio)**0.5and then copies the sign fromlerp(a, b, ratio). There are many options to control what blend mode is used, the power, whether the input gets converted to absolute values and how the sign is restored (sincewhatever**2will always be positive).magnitude_interpolation_lerpsignis just a preset for this with power 1 and using absolute inputs.moment_aligned- Similar todistro_aligned, maybe just a worse version of it. Attempts to align the std and mean (of the input) to a reference. It does not align the result, though you could probably do that through nesting.pythagorean_lerp- Like LERP but attempts to preserve the variance of the inputs. Let's say ratio is 0.4, so the multiplier forawould be0.6andbwould be0.4. We'd calculate a variance division(0.6**2 + 0.4 ** 2)**0.5(roughly0.7211) and scale the weights (0.6 / 0.7211 == 0.832,0.4 / 0.7211 == 0.554). If the ratio was0.1you'd get something likea * 0.993 + b * 0.11.