Cleanups, documentation updates
This commit is contained in:
@@ -9,86 +9,347 @@ Current status: In flux, not suitable for general use.
|
||||
|
||||
*Note*: You will basically always have to tweak settings like `s_noise` to get a good result. If the generation looks smooth/undetailed increase `s_noise` somewhere. If it looks crunchy, super high contrast, etc then try reducing noise.
|
||||
|
||||
## Usage
|
||||
|
||||
First, a note on the basic structure:
|
||||
|
||||

|
||||
|
||||
The sampler node connects to a group node. You can chain group nodes, however only one can match per step. Checking for a match starts at the group furthest from the sampler node. So if you have `Group1 -> Group2 -> Group3 -> Sampler`, the groups will be tried starting from `Group1`. Currently matching groups is only time based.
|
||||
|
||||
You will then connect a substeps node to the group. These can also be chained and like groups, execution starts with the node furthest from the group. I.E.: `Substeps1 -> Substeps2 -> Substeps3 -> Group` will start with `Substeps1`.
|
||||
|
||||
Most of these nodes have a text parameter input (YAML format - JSON is also valid YAML so you can use that instead if you prefer) and a parameter input. The parameter input can be used to specify stuff like custom noise types.
|
||||
|
||||
## Nodes
|
||||
|
||||
### ComposableSampler
|
||||
### `OCS Sampler`
|
||||
|
||||
**Possible Parameters**
|
||||
The main sampler node, with an output suitable for connecting to a `SamplerCustom`. This node has builtin support for Restart sampling, if you are
|
||||
using Restart don't use the `RestartSampler` node.
|
||||
|
||||
* `avgmerge_stretch`(`0.4`): Used for `average` and `sample` merge types. See below.
|
||||
* `model_call_cache`(unset): Caches the result of model calls. For example, Bogacki is 3 model calls per step. The first one usually depends on the merge strategy: `average` for example shares the first model evaluation between substeps, but subsequent model calls (i.e. Bogacki 2nd and 3rd model evaluations) still occur. When the model call cache is active, it's possible to cache those evalutions and avoid a model call for the remaining substeps. If you set `model_call_cache` to `1` then the result of that second call will be cached and if you're running two Bogacki substeps then the second one will use the cached version. Massively accelerates inference (especially when using the `average` merge strategy) but is likely very unsound and inaccurate. Does not apply to the sampler call for the `sample` merge strategy.
|
||||
* `model_call_cache_threshold`(`1`): Disables caching model call results with a call index below the threshold value (starting at 0). For example, if set to `2` and using a sampler like Bogacki that calls the model two extra times, the first will never be cached. The default value of `1` disabling caching for the first model call per substep. I generally would not recommend setting it to `0`, especially with `average` or `sample` merge strategies.
|
||||
* `model_call_cache_max_use`(`1000000`): The number of times cache items can be re-used. The default is effectively no limit. Where would this be useful? Let's say you're using the `average` merge strategy and a multi step sampler that calls the model at least one more time with 50 substeps. If you set the value to `25`, the model cache result will be updated around substep 25 which _may_ produce better results than reusing the result 50 times.
|
||||
You can connect a chain of `OCS Group` nodes to it and it will choose one per step (based on conditions like time).
|
||||
|
||||
#### Input Parameters
|
||||
|
||||
* `restart_custom_noise`: Value type: `SONAR_CUSTOM_NOISE`. Allows specifying a custom noise type when used with Restart sampling.
|
||||
|
||||
#### Text Parameters
|
||||
|
||||
Show in YAML with default values.
|
||||
|
||||
```yaml
|
||||
# Noise scale. May not do anything currently.
|
||||
s_noise: 1.0
|
||||
|
||||
# ETA (basically ancestralness). May not do anything currently.
|
||||
eta: 1.0
|
||||
|
||||
# Reversible ETA (used for reversible samplers). May not do anything currently.
|
||||
reta: 1.0
|
||||
|
||||
# Scales the noise added when used with Restart sampling.
|
||||
restart_s_noise: 1.0
|
||||
|
||||
|
||||
# The noise block allows defining global noise sampling parameters.
|
||||
noise:
|
||||
# You can disable this to allow GPU noise generation. I believe it only makes a difference for Brownian.
|
||||
cpu_noise: true
|
||||
|
||||
# ComfyUI has a bug where if you disable add_noise in the sampler, no seed gets set. If you
|
||||
# are manually noising a sample and have add_noise turned off then you should enable this if
|
||||
# you want reproducible generations.
|
||||
set_seed: false
|
||||
|
||||
# Global scale scale for generated noise
|
||||
scale: 1.0
|
||||
|
||||
# Whether the generated noise should be normalized before use. Generally a good idea to leave enabled.
|
||||
normalize_noise: true
|
||||
|
||||
# Dimensions to normalize over (when normalization is enabled). Negative values mean starting
|
||||
# from the end (i.e. -1 means the last dimension, -2 means the penultimate dimension).
|
||||
# Latents generally have these dimensions: batch, channels, height, width
|
||||
# The default of [-3, -2, -1] normalizes noise over the batch. You can try something like
|
||||
# [-2, -1] to normalize over the batch and channels. See: https://pytorch.org/docs/stable/generated/torch.std.html
|
||||
normalize_dims: [-3, -2, -1]
|
||||
|
||||
# When caching, the batch size for chunks of noise to generate in advance. Generating a batch of noise
|
||||
# can be more efficient than generating on demand when using a high number of substeps (>10) per step.
|
||||
batch_size: 32
|
||||
|
||||
# Whether to cache noise.
|
||||
caching: true
|
||||
|
||||
# Interval (in full steps) to reset the cache. Brownian noise takes time into account so
|
||||
# if using Brownian you will generally want to reset each step.
|
||||
cache_reset_interval: 1
|
||||
|
||||
|
||||
# Model calls can be cached. This is very experimental: I don't recommend using it
|
||||
# unless you know what you're doing.
|
||||
model_call_cache:
|
||||
|
||||
# The cache size.
|
||||
size: 0
|
||||
|
||||
# Threshold for model call caching. For example if you have size=3 and threshold=1
|
||||
# then model calls 1 through 3 will be cached, but model call 0 will not be (the first one).
|
||||
# Additional explanation: Some samplers call the model multiple times per step. For example,
|
||||
# Bogacki uses three model calls: 0, 1, 2
|
||||
threshold: 1
|
||||
|
||||
# Maximum use count for cache items.
|
||||
max_use: 1000000
|
||||
```
|
||||
|
||||
Any parameters you don't specify will use the defaults. For example if your text parameter block is:
|
||||
|
||||
```yaml
|
||||
noise:
|
||||
cpu_noise: false
|
||||
```
|
||||
|
||||
Then the rest of the parameters will use the defaults shown above.
|
||||
|
||||
### `OCS Group`
|
||||
|
||||
Defines a group of substeps.
|
||||
|
||||
Since it's kind of confusing even for me, a little more explanation: The model call cache caches results for model call indexes between `model_call_cache_threshold` and `model_call_cache - 1`. If you set `model_call_cache_threshold` to `0` and `model_call_cache` to `1` then only the first model call will be cached. If you set `model_call_cache_threshold` to `1` and `model_call_cache` to `2` then call 0 will not be cached, call 1 will be cached, call 2 will be cached, call 3 will not be cached, and so on.
|
||||
|
||||
#### Merging
|
||||
|
||||
When running multiple substeps per step, the results will combined based on the merge strategy. Possible strategies (in order of least weird to most weird):
|
||||
|
||||
* `simple`: Doesn't merge anything: only runs a single substep per step.
|
||||
* `divide`: Creates a linear schedule between the current sigma and the next and runs the substeps in sequence. The model is called at least once per substep.
|
||||
* `normal`: The model is called at least once per substep (and possibly additional times for higher order samplers). The result of each substep is noised and the next substep uses that result. Then all the results are averaged.
|
||||
* `normal`: The model is called at least once per substep (and possibly additional times for higher order samplers). Each substep shares the first model call result. The results are averaged together. *Note*: Since the first model call is shared and the initial input is the same for each substep, there is no point in running multiple identical substeps. Also note: This merge strategy doesn't work well with non-ancestral samplers (i.e. dpmpp_2m or any sampler with `eta: 0`).
|
||||
<!--
|
||||
* `average`: The model is called once at the beginning of the step and substeps share that result (but it may be called additional times for higher order samplers). This means substeps for samplers like reversible Euler, Heun 1s, DPM++ 2m SDE are essentially free. May be theoretically very unsound and inaccurate, requires manual tweaking of settings like `s_noise`. Supports the parameter `avgmerge_stretch`(`0.4`) which basically rolls back the current sigma and adds some noise (otherwise running a substep is deterministic and there would be no point to running a sampler like Euler more than once).
|
||||
* `sample`: Like `average` (and uses `avgmerge_stretch`) but instead of simply using the average, it does a sampler step toward that instead. You can plug in any substep sampler to the `merge_sampler_opt` input (if unconnected and the merge method is `sample` then Euler will be used). *Note*: Substeps in the attached sampler will be ignored.
|
||||
* `sample_uncached`: Similar to `sample`, however it calls the model per substep instead of caching the result and sharing it. Aside from sampling toward the result, it works more like the `normal` merge strategy. Theoretically it should be better because it's taking less shortcuts but results seem worse.
|
||||
|
||||
When using `average` and `sample` merge strategies and with model call caching enabled you can get away with setting substeps super high. Running something like 100 substeps is actually quite practical and seems to work well.
|
||||
-->
|
||||
|
||||
### ComposableStepSampler
|
||||
#### Node Parameters
|
||||
|
||||
This node has a text input for YAML (or JSON) advanced parameters.
|
||||
* `merge_method`: One of `simple`, `divide`, `normal` <!--, `average`, `sample`, `sample_uncached`. -->
|
||||
* `time_mode`(`step`): One of `step`, `step_pct`, `sigma`. Time matching mode. Matching based on steps generally will be simplest. Matches are inclusive and steps start at 0 (so step 0 is the first step). `step_pct` is the percentage of total steps (1.0=100%, 0.5=50%, etc).
|
||||
* `time_start`(`0`): Match start time.
|
||||
* `time_end`(`999`): Match end time.
|
||||
|
||||
For example, you could enter something like this in the field:
|
||||
Example:
|
||||
|
||||

|
||||
|
||||
The left side group matches steps 0, 1, 2. The right side group matches all steps. This setup will use whatever substeps are connected to the first group for the first three steps and the second group will handle the rest.
|
||||
|
||||
|
||||
#### Input Parameters
|
||||
|
||||
Currently unused for groups.
|
||||
|
||||
<!--
|
||||
* `merge_sampler`: Value type: `OCS_SUBSTEPS`. Only used when `merge_method` is `sample` or `sample_uncached`. Allows defining the sampler used for merging substeps.
|
||||
-->
|
||||
|
||||
|
||||
#### Text Parameters
|
||||
|
||||
Show in YAML with default values.
|
||||
|
||||
```yaml
|
||||
reta: 1.1
|
||||
leap: 3
|
||||
dyn_deta_mode: "deta"
|
||||
# Noise scale. May not do anything currently.
|
||||
s_noise: 1.0
|
||||
|
||||
# ETA (basically ancestralness). May not do anything currently.
|
||||
eta: 1.0
|
||||
|
||||
# Reversible ETA (used for reversible samplers). May not do anything currently.
|
||||
reta: 1.0
|
||||
|
||||
# Currently unused.
|
||||
avgmerge_stretch: 0.4
|
||||
```
|
||||
|
||||
**Possible Parameters**
|
||||
### `OCS Substeps`
|
||||
|
||||
#### General
|
||||
#### Step Methods (Samplers)
|
||||
|
||||
* `eta`(`1.0`): Will override `eta` in the node if set.
|
||||
* `dyn_eta_start`(`unset`) and `dyn_eta_end`(`unset`): No effect unless both values are set. Will interpolate between start and end based on the percentage of sampling. *Note*: This is a factor applied to ETA, not a flat value.
|
||||
* `s_noise`(`1.0`): Will override `s_noise` in the node if set.
|
||||
* `solver_type`(`midpoint`): Applies to DPM++ 2m SDE. May be one of `midpoint` or `heun` (`midpoint` is generally recommended).
|
||||
In alphabetical order.
|
||||
|
||||
#### Reversible
|
||||
* `bogacki`: CFG++ enabled.
|
||||
* `deis`: May not work work ETA > 0. See parameters: `history_limit`
|
||||
* `dpmpp_2m_sde`: See parameters: `history_limit`
|
||||
* `dpmpp_2m`: `eta` and `s_noise` parameters are ignored. See parameters: `history_limit`
|
||||
* `dpmpp_2s`
|
||||
* `dpmpp_3m_sde`: See parameters: `history_limit`
|
||||
* `dpmpp_sde`
|
||||
* `euler_cycle`: CFG++ enabled. See parameters: `cycle_pct`
|
||||
* `euler_dancing`: Pretty broken currently, will probably require increased `s_noise` values. See parameters: `deta`, `leap`, `deta_mode`
|
||||
* `euler`: CFG++ enabled.
|
||||
* `heunpp`: May not work work ETA > 0. See parameters: `max_order`
|
||||
* `ipndm_v`: May not work work ETA > 0. See parameters: `history_limit`
|
||||
* `ipndm`: May not work work ETA > 0. See parameters: `history_limit`
|
||||
* `res`
|
||||
* `reversible_bogacki`: CFG++ enabled.
|
||||
* `reversible_heun`: CFG++ enabled.
|
||||
* `reversible_heun_1s`: CFG++ enabled. See parameters: `history_limit`
|
||||
* `rk4`: CFG++ enabled.
|
||||
* `tde`: Uses the [torchdiffeq](https://github.com/rtqichen/torchdiffeq) ODE backend. See `ode_*` parameters below. CFG++ enabled.
|
||||
* `tode`: Uses the [torchode]((https://github.com/martenlienen/torchode)) ODE backend. See `ode_*` parameters below. CFG++ enabled.
|
||||
* `trapezoidal`: CFG++ enabled.
|
||||
* `trapezoidal_cycle`: CFG++ enabled. See parameters: `cycle_pct`
|
||||
* `ttm_jvp`: TTM is a weird sampler. If you're using model caching you must make sure the entries TTM uses are populated first (by having it run before any other samplers that call the model multiple times). It may also not work with some other model patches and upscale methods. See parameters: `alternate_phi_2_calc`
|
||||
|
||||
* `reta`(`1.0`): Reverse ETA.
|
||||
* `dyn_reta_start`(`unset`) and `dyn_reta_end`(`unset`): No effect unless both values are set. Will interpolate between start and end based on the percentage of sampling. *Note*: This is a factor applied to RETA, not a flat value.
|
||||
**ODE Solvers** (`tde`, `tode`):
|
||||
|
||||
#### Dancing
|
||||
You will need to have the relevant Python package installed in your venv to use these. TDE cannot handle batches and
|
||||
each batch item will be evaluated separately. Using `tode` may be faster for batch sizes over 1.
|
||||
|
||||
* `leap`(`2`): Distance to try to leap forward. If you set `leap` to `1` you just get plain old Euler ancestral.
|
||||
* `deta`(`1.0`): ETA used for dance steps.
|
||||
* `dyn_deta_start`(`unset`) and `dyn_deta_end`(`unset`): No effect unless both values are set. Will interpolate between start and end based on the percentage of sampling. *Note*: This is a factor applied to DETA, not a flat value.
|
||||
* `dyn_deta_mode`(`lerp`): May be one of:
|
||||
* `deta`: Scales `deta` based on the value from `dyn_deta_start/end`.
|
||||
* `lerp`: Does the dance step according to `deta` and then LERPs the non-dance sample result with the dance sample result based on the scale calculated from `dyn_deta_start/end` (which is `1.0` if they are unset). For example, if the dance scale is `0.5` you will get 50% normal sampling, 50% dancing sampling.
|
||||
* `lerp_alt`: Similar to `lerp` except it LERPs with the leap result instead of a normal Euler ancestral result.
|
||||
`ode_solver` types for TDE: adaptive: `dopri8`, `dopri5`, `bosh3`, `fehlberg2`, `adaptive_heun`, fixed step: `euler`, `midpoint`, `rk4`, `explicit_adams`, `implicit_adams`
|
||||
|
||||
#### RES
|
||||
`ode_solver` types for TODE: adaptive only: `dopri5`, `tsit5`, `heun`. I haven't much luck with anything other than `dopri5`.
|
||||
|
||||
* `res_simple_phi`(`false`): Uses a faster but possibly less accurate method for calculating phi. What does phi do? I haven't the foggiest!
|
||||
* `res_c2`(`0.5`): Solver partial step size, the default of `0.5` appears to use the midpoint. Setting it to a lower value might possibly be more accurate but slower?
|
||||
Note that adaptive solvers may be _very_ slow. Think along the lines of 20-100 model calls per substep (or in other words, the equivalent for running that many `euler` steps). Tolerances only apply to adaptive solvers.
|
||||
|
||||
#### TTM JVP
|
||||
**Cycle Samplers** (`euler_cycle`, `trapezoidal_cycle`)
|
||||
|
||||
`alterate_phi_2_calc`(`true`): Supposedly works better than disabled when ETA isn't 0. I didn't notice a difference.
|
||||
Basically a different approach to ancestral sampling. First a crash course on how sampling works:
|
||||
|
||||
Each step has an expected noise level, with the first step generally being pure noise and the end of the last step aiming to end with no noise remaining. Let's say the image on the current step is called `x`, calling the model with `x` gives us a prediction of what the image looks like with all noise removed (`denoised`), however the model is not capable of just removing all the noise in a single step: its prediction will be imprecise. `x - denoised` leaves us with just the noise (we subtract the prediction which theoretically has no noise from the noisy sample). This is a very simplified, but the idea is basically to add the noise back into `denoised`, but scaled so that it matches the amount of noise expected on the _next_ step. `denoised + noise * expected_noise_at_next_step`.
|
||||
|
||||
When doing ancestral sampling, we actually _overshoot_ expected noise for the next step and add less than that amount back to `denoised`. Then we generate some of our own noise and add it, scaled so that the result matches `expected_noise_at_next_step`. `eta` controls how the scale of the overshoot.
|
||||
|
||||
The difference with cycle is that instead of adding `noise * expected_noise_at_next_step`, we instead first add `noise * (expected_noise_at_next_step * (1.0 - cycle_pct))` and then we generate noise and scale it to `cycle_pct` and add it too. Just for example, suppose `cycle_pct` is `0.2`: we'll add 80% of the expected noise at the next step (`1.0 - 0.2 == 0.8`) and then generate the remaining 20% and add it in to meet the expected amount. I don't recommend setting `cycle_pct` to values over `0.5`, especially if using "weird" noise types.
|
||||
|
||||
#### Node Parameters
|
||||
|
||||
* `substeps`(`1`): Number of substeps. Generally involves a model call per substep, so for example setting this to 4 would approximately quadruple sampling time.
|
||||
* `step_method`(`euler`): Method used for sampling the substeps. May include a parenthesized number (i.e. `rk4 (3)`) which denotes the number of _extra_ model calls required per sample. At least one is always required. So `euler` requires 1 in total, `rk4` requires 4 in total. RK4 is about 4 times slower than `euler`.
|
||||
|
||||
#### Input Parameters
|
||||
|
||||
* `custom_noise`: Value type: `SONAR_CUSTOM_NOISE`. Allows specifying a custom noise type for samplers that generate noise (most of them).
|
||||
|
||||
#### Text Parameters
|
||||
|
||||
Show in YAML with default values.
|
||||
|
||||
```yaml
|
||||
# Scale for added noise.
|
||||
s_noise: 1.0
|
||||
|
||||
# ETA (basically ancestralness).
|
||||
eta: 1.0
|
||||
# No effect unless both start and end are set. Will scale the eta value based on the
|
||||
# percentage of sampling. In other words, eta*dyn_eta_start at the beginning,
|
||||
# eta*dyn_eta_end at the end.
|
||||
dyn_eta_start: null
|
||||
dyn_eta_end: null
|
||||
|
||||
# CFG++ scale (see https://cfgpp-diffusion.github.io/)
|
||||
# Setting this to 1.0 is the equivalent of enabling it. Can also be set
|
||||
# to a negative value (I don't recommend going lower than -0.5).
|
||||
cfgpp_scale: 0
|
||||
|
||||
### Reversible Settings ###
|
||||
|
||||
# Reversible ETA (used for reversible samplers).
|
||||
reta: 1.0
|
||||
# Scale of the reversible correction. Can also be set to a negative value.
|
||||
reversible_scale: 1.0
|
||||
# No effect unless both start and end are set. Will scale the reta value based on the
|
||||
# percentage of sampling. In other words, reta*dyn_reta_start at the beginning,
|
||||
# reta*dyn_reta_end at the end.
|
||||
dyn_reta_start: null
|
||||
dyn_reta_end: null
|
||||
|
||||
### ODE Sampler Settings ###
|
||||
|
||||
# Solver type.
|
||||
ode_solver: dopri5
|
||||
# Relative tolerance (log 10)
|
||||
ode_rtol: -1.5
|
||||
# Absolute tolerance (log 10)
|
||||
ode_atol: -3.5
|
||||
# Max model calls allowed to compute the solution. If the limit is exceeded, it is an error.
|
||||
ode_max_nfe: 1000
|
||||
# Hack that seems to help results. Set to 0 to disable.
|
||||
ode_fixup_hack: 0.025
|
||||
|
||||
## torchdiffeq (tde) specific parameters ##
|
||||
# Used to split the step into sections. Useful for fixed step methods.
|
||||
ode_split: 1
|
||||
|
||||
## torchode (tode) specific parameters ##
|
||||
# Initial step size (as a percentage).
|
||||
ode_initial_step: 0.25
|
||||
|
||||
# Coefficients for the step size PID controller.
|
||||
# See https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller
|
||||
# These values seem okay with dopri5.
|
||||
ode_ctl_pcoeff: 0.3
|
||||
ode_ctl_icoeff: 0.9
|
||||
ode_ctl_dcoeff: 0.2
|
||||
|
||||
# Controls whether to compile the solver. May or may not work,
|
||||
# also may or may not be a speed increase as the compiled solver is
|
||||
# not cached between substeps.
|
||||
ode_compile: false
|
||||
|
||||
### Other Sampler Specific Parameters ###
|
||||
|
||||
# Used for some samplers that use history from previous steps.
|
||||
# List of samplers and default value below:
|
||||
# dpmpp_2m: 1
|
||||
# dpmpp_2m_sde: 1
|
||||
# dpmpp_3m_sde: 2
|
||||
# reversible_heun_1s: 1
|
||||
# ipndm: 3
|
||||
# ipndm_v: 3
|
||||
# deis: 2 (max 3)
|
||||
history_limit: 999 # Varies based on sampler.
|
||||
|
||||
# Used for some samplers with variable order. List of samplers and default value below:
|
||||
# heunpp2: 3
|
||||
max_order: 999 # Varies based on sampler.
|
||||
|
||||
# Used for dpmpp_2m. One of midpoint, heun
|
||||
solver_type: "midpoint"
|
||||
|
||||
# Coefficients mode for DEIS. One of tab or rhoab.
|
||||
deis_mode: "tab"
|
||||
|
||||
# Used for samplers with cycle in the name. Controls how much noise is cycled per step.
|
||||
cycle_pct: 0.25
|
||||
|
||||
# Used for ttm_jvp. Supposed works better when ETA > 0
|
||||
alternate_phi_2_calc: true
|
||||
|
||||
# Parameters for dancing samplers:
|
||||
# Number of steps to leap ahead.
|
||||
leap: 2
|
||||
# ETA for dance steps
|
||||
deta: 1.0
|
||||
# dyn_deta works the same as dyn_eta/reta. See above.
|
||||
dyn_deta_start: null
|
||||
dyn_deta_end: null
|
||||
# One of lerp, lerp_alt, deta
|
||||
dyn_deta_mode: "lerp"
|
||||
```
|
||||
|
||||
**Note**: TTM is a weird sampler. If you're using model caching you must make sure the entries TTM uses are populated first (by having before any other samplers that call the model multiple times). It may also not work with some other model patches and upscale methods.
|
||||
|
||||
## Credits
|
||||
|
||||
I can move code around but sampling math and creating samplers is far beyond my ability. I didn't write any of the original samplers:
|
||||
|
||||
* Euler, DPMPP SDE, DPMPP 2S, DPM++ 2m, 2m SDE and 3m SDE samplers based on ComfyUI's implementation.
|
||||
* Euler, Heun++2, DPMPP SDE, DPMPP 2S, DPM++ 2m, 2m SDE and 3m SDE samplers based on ComfyUI's implementation.
|
||||
* Reversible Heun, Reversible Heun 1s, RES, Trapezoidal, Bogacki, Reversible Bogacki, RK4 and Euler Dancing samplers based on implementation from https://github.com/Clybius/ComfyUI-Extra-Samplers
|
||||
* TTM JVP sampler based on implementation written by Katherine Crowson (but yoinked from the Extra-Samplers repo mentioned above).
|
||||
* IPNDM and IPNDM_V adapted from https://github.com/zju-pi/diff-sampler/blob/main/diff-solvers-main/solvers.py (I used the Comfy version as a reference).
|
||||
* IPNDM, IPNDM_V and DEIS adapted from https://github.com/zju-pi/diff-sampler/blob/main/diff-solvers-main/solvers.py (I used the Comfy version as a reference).
|
||||
* Normal substep merge strategy based on implementation from https://github.com/Clybius/ComfyUI-Extra-Samplers
|
||||
|
||||
This repo wouldn't be possible without building on the work of others. Thanks!
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 48 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 31 KiB |
+20
-187
@@ -93,8 +93,9 @@ class NormalMergeSubstepsSampler(MergeSubstepsSampler):
|
||||
noise = z_avg.clone()
|
||||
noise_total = 0.0
|
||||
substep = 0
|
||||
curr_x = x
|
||||
orig_std = None
|
||||
pbar = tqdm.tqdm(total=self.substeps, initial=1, disable=ss.disable_status)
|
||||
ss.hist.push(ss.model(x, ss.sigma))
|
||||
ss.callback()
|
||||
for ssampler in self.samplers:
|
||||
custom_noise = ssampler.options.get(
|
||||
"custom_noise", self.options.get("custom_noise")
|
||||
@@ -103,77 +104,18 @@ class NormalMergeSubstepsSampler(MergeSubstepsSampler):
|
||||
custom_noise, ssampler.max_noise_samples(), ss.sigma, ss.sigma_next
|
||||
)
|
||||
ssampler.noise_sampler = noise_sampler
|
||||
# STEP 1: 20.6 -> 15.9 || up=10.1, down=12.3
|
||||
for subidx in range(ssampler.substeps):
|
||||
print(f" SUBSTEP {substep + 1}: {ssampler.name}")
|
||||
ss.denoised = ss.model(curr_x, ss.sigma)
|
||||
hmm = curr_x - ss.denoised
|
||||
hmmm = hmm.mean(dim=(-3, -2, -1), keepdim=True)
|
||||
hmms = hmm.std(dim=(-3, -2, -1), keepdim=True)
|
||||
if substep == 0:
|
||||
orig_std = hmms
|
||||
print(">>>", hmmm, hmms, "|", orig_std)
|
||||
sr = self.simple_substep(curr_x, ssampler)
|
||||
# x = sr.x
|
||||
pbar.set_description(f"{ssampler.name}: {substep + 1}/{substeps}")
|
||||
sr = self.simple_substep(x, ssampler)
|
||||
z_avg += renoise_weight * sr.x
|
||||
noise_strength = sr.noise_scale
|
||||
if sr.noise_scale != 0 and ss.sigma_next != 0:
|
||||
noise_total += renoise_weight * sr.noise_scale
|
||||
noise += renoise_weight * sr.get_noise()
|
||||
substep += 1
|
||||
if ss.sigma_next == 0:
|
||||
continue
|
||||
noise_curr = sr.get_noise(scaled=False)
|
||||
# rnoise = ssampler.s_noise * ((ss.sigma**2 - ss.sigma_next**2) ** 0.5)
|
||||
# rnoise = ss.sigma
|
||||
# rnoise = ssampler.s_noise * (
|
||||
# noise_strength + ((ss.sigma**2 - ss.sigma_next**2) ** 0.4)
|
||||
# )
|
||||
rnoise = ssampler.s_noise * orig_std
|
||||
if substep < substeps:
|
||||
# print(
|
||||
# "NOISE?",
|
||||
# sr.noise_scale,
|
||||
# ((ss.sigma**2 - ss.sigma_next**2) ** 0.5),
|
||||
# )
|
||||
# x = x + ss.noise.scale_noise(noise_curr, rnoise)
|
||||
# noise_strength = torch.lerp(noise_strength, rnoise, 0.75)
|
||||
noise_curr = ss.noise.scale_noise(
|
||||
noise_curr, orig_std, normalized=True
|
||||
)
|
||||
# curr_x = ss.denoised + ss.noise.scale_noise(noise_curr, rnoise)
|
||||
curr_x = ss.denoised + noise_curr
|
||||
pbar.update(1)
|
||||
|
||||
noise_strength = rnoise
|
||||
# noise_total += rnoise * renoise_weight
|
||||
# noise_total += (
|
||||
# (noise_strength + rnoise) * 0.5 * renoise_weight
|
||||
# if substep < substeps
|
||||
# else noise_strength * renoise_weight
|
||||
# )
|
||||
|
||||
print(
|
||||
"NOISE?", noise_strength, rnoise, "::", ss.sigma_next - ss.sigma_up
|
||||
)
|
||||
noise_total += (
|
||||
noise_strength - (ss.sigma_next - ss.sigma_up)
|
||||
) * renoise_weight
|
||||
# noise_total += torch.lerp(noise_strength, rnoise, 0.75) * renoise_weight
|
||||
# noise_total += ((noise_strength.item() + rnoise) * 0.5) * renoise_weight
|
||||
# noise = noise + ss.noise.scale_noise(noise_curr, noise_strength)
|
||||
# noise = noise + noise_curr * noise_strength * renoise_weight
|
||||
# noise = ss.noise.scale_noise(
|
||||
# noise + sr.noise_scale * noise_curr, noise_total, normalized=True
|
||||
# )
|
||||
print(
|
||||
"NOISE TOT",
|
||||
noise_total,
|
||||
"--",
|
||||
ss.sigma_up,
|
||||
)
|
||||
|
||||
noise_sampler = ss.noise.make_caching_noise_sampler(
|
||||
self.options.get("custom_noise"), 1, ss.sigma, ss.sigma_next
|
||||
)
|
||||
noise = ss.noise.scale_noise(
|
||||
noise_sampler(),
|
||||
noise,
|
||||
noise_total * self.options.get("s_noise", 1.0),
|
||||
normalized=True,
|
||||
)
|
||||
@@ -181,118 +123,6 @@ class NormalMergeSubstepsSampler(MergeSubstepsSampler):
|
||||
x, z_avg, noise=None if noise_total == 0 else noise, denoised=ss.denoised
|
||||
)
|
||||
|
||||
# def step(self, x):
|
||||
# ss = self.ss
|
||||
# substeps = self.substeps
|
||||
# renoise_weight = 1.0 / substeps
|
||||
# z_avg = torch.zeros_like(x)
|
||||
# noise = z_avg.clone()
|
||||
# noise_total = 0.0
|
||||
# substep = 0
|
||||
# curr_x = x
|
||||
# for ssampler in self.samplers:
|
||||
# custom_noise = ssampler.options.get(
|
||||
# "custom_noise", self.options.get("custom_noise")
|
||||
# )
|
||||
# noise_sampler = ss.noise.make_caching_noise_sampler(
|
||||
# custom_noise, ssampler.max_noise_samples(), ss.sigma, ss.sigma_next
|
||||
# )
|
||||
# ssampler.noise_sampler = noise_sampler
|
||||
# for subidx in range(ssampler.substeps):
|
||||
# print(f" SUBSTEP {substep + 1}: {ssampler.name}")
|
||||
# ss.denoised = ss.model(x, ss.sigma)
|
||||
# sr = self.simple_substep(x, ssampler)
|
||||
# x = sr.x
|
||||
# z_avg += renoise_weight * x
|
||||
# noise_strength = sr.noise_scale
|
||||
# substep += 1
|
||||
# if ss.sigma_next == 0 or noise_strength == 0:
|
||||
# continue
|
||||
# noise_curr = sr.get_noise(scaled=False)
|
||||
# rnoise = ssampler.s_noise * ((ss.sigma**2 - ss.sigma_next**2) ** 0.5)
|
||||
# if substep < substeps:
|
||||
# # print(
|
||||
# # "NOISE?",
|
||||
# # sr.noise_scale,
|
||||
# # ((ss.sigma**2 - ss.sigma_next**2) ** 0.5),
|
||||
# # )
|
||||
# x = x + ss.noise.scale_noise(noise_curr, rnoise)
|
||||
# # noise_strength = torch.lerp(noise_strength, rnoise, 0.75)
|
||||
# noise_strength = rnoise
|
||||
# # noise_total += rnoise * renoise_weight
|
||||
# # noise_total += (
|
||||
# # (noise_strength + rnoise) * 0.5 * renoise_weight
|
||||
# # if substep < substeps
|
||||
# # else noise_strength * renoise_weight
|
||||
# # )
|
||||
|
||||
# print("NOISE?", noise_strength, rnoise)
|
||||
# noise_total += noise_strength * renoise_weight
|
||||
# # noise_total += torch.lerp(noise_strength, rnoise, 0.75) * renoise_weight
|
||||
# # noise_total += ((noise_strength.item() + rnoise) * 0.5) * renoise_weight
|
||||
# # noise = noise + ss.noise.scale_noise(noise_curr, noise_strength)
|
||||
# # noise = noise + noise_curr * noise_strength * renoise_weight
|
||||
# # noise = ss.noise.scale_noise(
|
||||
# # noise + sr.noise_scale * noise_curr, noise_total, normalized=True
|
||||
# # )
|
||||
# print(
|
||||
# "NOISE TOT",
|
||||
# noise_total,
|
||||
# "--",
|
||||
# ss.sigma_up,
|
||||
# )
|
||||
|
||||
# noise_sampler = ss.noise.make_caching_noise_sampler(
|
||||
# self.options.get("custom_noise"), 1, ss.sigma, ss.sigma_next
|
||||
# )
|
||||
# noise = ss.noise.scale_noise(
|
||||
# noise_sampler(),
|
||||
# noise_total * self.options.get("s_noise", 1.0),
|
||||
# normalized=True,
|
||||
# )
|
||||
# return self.merge_steps(
|
||||
# x, z_avg, noise=None if noise_total == 0 else noise, denoised=ss.denoised
|
||||
# )
|
||||
|
||||
# def step(self, x):
|
||||
# ss = self.ss
|
||||
# substeps = self.substeps
|
||||
# renoise_weight = 1.0 / substeps
|
||||
# z_avg = torch.zeros_like(x)
|
||||
# noise = torch.zeros_like(x)
|
||||
# noise_total = 0.0
|
||||
# substep = 0
|
||||
# for ssampler in self.samplers:
|
||||
# custom_noise = ssampler.options.get(
|
||||
# "custom_noise", self.options.get("custom_noise")
|
||||
# )
|
||||
# noise_sampler = ss.noise.make_caching_noise_sampler(
|
||||
# custom_noise, ssampler.max_noise_samples(), ss.sigma, ss.sigma_next
|
||||
# )
|
||||
# ssampler.noise_sampler = noise_sampler
|
||||
# for subidx in range(ssampler.substeps):
|
||||
# print(f" SUBSTEP {substep + 1}: {ssampler.name}")
|
||||
# ss.denoised = ss.model(x, ss.sigma)
|
||||
# sr = self.simple_substep(x, ssampler)
|
||||
# x = sr.x
|
||||
# z_avg += renoise_weight * x
|
||||
# noise_strength = sr.noise_scale
|
||||
# substep += 1
|
||||
# if ss.sigma_next == 0 or noise_strength == 0:
|
||||
# continue
|
||||
# noise_curr = sr.get_noise()
|
||||
# if substep < substeps or True:
|
||||
# x = x + noise_curr
|
||||
# noise_total += noise_strength.item() * renoise_weight
|
||||
# noise += noise_curr
|
||||
# return self.merge_steps(
|
||||
# x,
|
||||
# z_avg,
|
||||
# noise=None
|
||||
# if not noise_total
|
||||
# else ss.noise.scale_noise(noise, noise_total * ss.s_noise, normalized=True),
|
||||
# )
|
||||
|
||||
|
||||
class AverageMergeSubstepsSampler(NormalMergeSubstepsSampler):
|
||||
name = "average"
|
||||
@@ -349,9 +179,11 @@ class AverageMergeSubstepsSampler(NormalMergeSubstepsSampler):
|
||||
noise_strength = sr.noise_scale
|
||||
if ss.sigma_next == 0 or noise_strength == 0:
|
||||
continue
|
||||
noise_curr = sr.get_noise()
|
||||
noise_total += noise_strength.item() * renoise_weight
|
||||
noise += noise_curr
|
||||
if noise_strength != 0 and ss.sigma_next != 0:
|
||||
noise_curr = sr.get_noise()
|
||||
noise_total += noise_strength.item() * renoise_weight
|
||||
noise += noise_curr
|
||||
substep += 1
|
||||
substep += ssampler.substeps
|
||||
return self.merge_steps(
|
||||
x,
|
||||
@@ -515,7 +347,7 @@ class DivideMergeSubstepsSampler(MergeSubstepsSampler):
|
||||
f"substep({ssampler.name}): {subss.sigma.item():.03} -> {subss.sigma_next.item():.03}"
|
||||
)
|
||||
subss.hist.push(subss.model(x, subss.sigma))
|
||||
if substep == self.substeps - 1:
|
||||
if substep == 0:
|
||||
subss.callback()
|
||||
sr = self.simple_substep(x, ssampler, ss=subss)
|
||||
x = sr.x
|
||||
@@ -528,10 +360,11 @@ class DivideMergeSubstepsSampler(MergeSubstepsSampler):
|
||||
|
||||
|
||||
MERGE_SUBSTEPS_CLASSES = {
|
||||
"default (simple)": SimpleSubstepsSampler,
|
||||
"normal": NormalMergeSubstepsSampler,
|
||||
"divide": DivideMergeSubstepsSampler,
|
||||
"average": AverageMergeSubstepsSampler,
|
||||
"sample": SampleMergeSubstepsSampler,
|
||||
"sample_uncached": SampleUncachedMergeSubstepsSampler,
|
||||
# "average": AverageMergeSubstepsSampler,
|
||||
# "sample": SampleMergeSubstepsSampler,
|
||||
# "sample_uncached": SampleUncachedMergeSubstepsSampler,
|
||||
"simple": SimpleSubstepsSampler,
|
||||
}
|
||||
|
||||
+20
-26
@@ -11,7 +11,7 @@ from comfy.k_diffusion.sampling import (
|
||||
)
|
||||
|
||||
from .res_support import _de_second_order
|
||||
from .utils import find_first_unsorted, fallback
|
||||
from .utils import find_first_unsorted
|
||||
|
||||
HAVE_TDE = HAVE_TODE = False
|
||||
|
||||
@@ -73,15 +73,8 @@ class CFGPPStepMixin:
|
||||
def init_cfgpp(self, /, cfgpp_scale=0.0):
|
||||
self.cfgpp_scale = 0.0 if not self.allow_cfgpp else cfgpp_scale
|
||||
|
||||
def to_d(self, mr, /, x=None, sigma=None, denoised=None, denoised_uncond=None):
|
||||
x = fallback(x, mr.x)
|
||||
sigma = fallback(sigma, mr.sigma)
|
||||
denoised = fallback(denoised, mr.denoised)
|
||||
scale = self.cfgpp_scale
|
||||
if scale == 0:
|
||||
return to_d(x, sigma, denoised)
|
||||
denoised_uncond = fallback(denoised_uncond, mr.denoised_uncond)
|
||||
return to_d(x - denoised * scale + denoised_uncond * scale, sigma, denoised)
|
||||
def to_d(self, mr, **kwargs):
|
||||
return mr.to_d(**kwargs)
|
||||
|
||||
|
||||
class SingleStepSampler(CFGPPStepMixin):
|
||||
@@ -230,7 +223,7 @@ class EulerStep(SingleStepSampler):
|
||||
|
||||
|
||||
class CycleSingleStepSampler(SingleStepSampler):
|
||||
def __init__(self, *, cycle_pct=1.0, **kwargs):
|
||||
def __init__(self, *, cycle_pct=0.25, **kwargs):
|
||||
super().__init__(**kwargs)
|
||||
self.cycle_pct = cycle_pct
|
||||
|
||||
@@ -1085,7 +1078,7 @@ class HeunPP2Step(SingleStepSampler):
|
||||
|
||||
def __init__(self, *args, max_order=3, **kwargs):
|
||||
super().__init__(*args, **kwargs)
|
||||
self.max_order = max_order
|
||||
self.max_order = max(1, min(self.model_calls + 1, max_order))
|
||||
|
||||
def step(self, x, ss):
|
||||
steps_remain = max(0, len(ss.sigmas) - (ss.idx + 2))
|
||||
@@ -1364,29 +1357,30 @@ class TODEStep(SingleStepSampler, MinSigmaStepMixin):
|
||||
|
||||
|
||||
STEP_SAMPLERS = {
|
||||
"bogacki": BogackiStep,
|
||||
"default (euler)": EulerStep,
|
||||
"bogacki (2)": BogackiStep,
|
||||
"deis": DEISStep,
|
||||
"dpmpp_2m_sde": DPMPP2MSDEStep,
|
||||
"dpmpp_2m": DPMPP2MStep,
|
||||
"dpmpp_2s": DPMPP2SStep,
|
||||
"dpmpp_3m_sde": DPMPP3MSDEStep,
|
||||
"dpmpp_sde": DPMPPSDEStep,
|
||||
"dpmpp_sde (1)": DPMPPSDEStep,
|
||||
"euler_cycle": EulerCycleStep,
|
||||
"euler_dancing": EulerDancingStep,
|
||||
"euler": EulerStep,
|
||||
"heunpp": HeunPP2Step,
|
||||
"ipndm": IPNDMStep,
|
||||
"heunpp (1-2)": HeunPP2Step,
|
||||
"ipndm_v": IPNDMVStep,
|
||||
"deis": DEISStep,
|
||||
"res": RESStep,
|
||||
"reversible_bogacki": ReversibleBogackiStep,
|
||||
"ipndm": IPNDMStep,
|
||||
"res (1)": RESStep,
|
||||
"reversible_bogacki (2)": ReversibleBogackiStep,
|
||||
"reversible_heun (1)": ReversibleHeunStep,
|
||||
"reversible_heun_1s": ReversibleHeun1SStep,
|
||||
"reversible_heun": ReversibleHeunStep,
|
||||
"rk4": RK4Step,
|
||||
"tde": TDEStep,
|
||||
"tode": TODEStep,
|
||||
"trapezoidal_cycle": TrapezoidalCycleStep,
|
||||
"trapezoidal": TrapezoidalStep,
|
||||
"ttm_jvp": TTMJVPStep,
|
||||
"rk4 (3)": RK4Step,
|
||||
"tde (variable)": TDEStep,
|
||||
"tode (variable)": TODEStep,
|
||||
"trapezoidal (1)": TrapezoidalStep,
|
||||
"trapezoidal_cycle (1)": TrapezoidalCycleStep,
|
||||
"ttm_jvp (1)": TTMJVPStep,
|
||||
}
|
||||
|
||||
__all__ = (
|
||||
|
||||
+19
-35
@@ -113,34 +113,11 @@ class History:
|
||||
def reset(self):
|
||||
self.history = []
|
||||
|
||||
|
||||
class History_:
|
||||
def __init__(self, x, size):
|
||||
self.history = torch.zeros(size, *x.shape, device=x.device, dtype=x.dtype)
|
||||
self.size = size
|
||||
self.pos = 0
|
||||
self.last = None
|
||||
|
||||
def __len__(self):
|
||||
return min(self.pos, self.size)
|
||||
|
||||
def __getitem__(self, k):
|
||||
idx = (self.pos + k if k < 0 else self.pos + -self.size + k) % self.size
|
||||
print(f"\nFETCH {k}: pos={self.pos}, size={self.size}, at={idx}")
|
||||
return self.history[idx]
|
||||
|
||||
def push(self, val):
|
||||
self.last = self.pos % self.size
|
||||
print(
|
||||
f"\nPUSH {self.pos % self.size}: pos={self.pos}, size={self.size}, at={self.last}"
|
||||
)
|
||||
self.history[self.last] = val
|
||||
self.pos += 1
|
||||
|
||||
def reset(self):
|
||||
print("HRESET")
|
||||
self.pos = 0
|
||||
self.last = None
|
||||
def clone(self):
|
||||
obj = self.__new__(self.__class__)
|
||||
obj.__init__(self.size)
|
||||
obj.history = self.history.copy()
|
||||
return obj
|
||||
|
||||
|
||||
class NoiseSamplerCache:
|
||||
@@ -155,7 +132,7 @@ class NoiseSamplerCache:
|
||||
cpu_noise=True,
|
||||
batch_size=32,
|
||||
caching=True,
|
||||
cache_reset_interval=9999,
|
||||
cache_reset_interval=1,
|
||||
set_seed=False,
|
||||
scale=1.0,
|
||||
normalize_dims=(-3, -2, -1),
|
||||
@@ -286,13 +263,20 @@ class ModelResult:
|
||||
self.denoised = denoised
|
||||
self.denoised_uncond = denoised_uncond
|
||||
|
||||
def to_d(
|
||||
self, /, x=None, sigma=None, denoised=None, denoised_uncond=None, cfgpp_scale=0
|
||||
):
|
||||
x = fallback(x, self.x)
|
||||
sigma = fallback(sigma, self.sigma)
|
||||
denoised = fallback(denoised, self.denoised)
|
||||
denoised_uncond = fallback(denoised_uncond, self.denoised_uncond)
|
||||
if cfgpp_scale != 0:
|
||||
x = x - denoised * cfgpp_scale + denoised_uncond * cfgpp_scale
|
||||
return to_d(x, sigma, denoised)
|
||||
|
||||
@property
|
||||
def d(self, /, x=None, sigma=None, denoised=None):
|
||||
return to_d(
|
||||
fallback(x, self.x),
|
||||
fallback(sigma, self.sigma),
|
||||
fallback(denoised, self.denoised),
|
||||
)
|
||||
def d(self):
|
||||
return self.to_d()
|
||||
|
||||
|
||||
class ModelCallCache:
|
||||
|
||||
Reference in New Issue
Block a user