Several nodes only exposed a 'cuda'/'cpu' device dropdown, forcing users on
any other ComfyUI-supported accelerator (Ascend NPU, XPU, MPS) to run on CPU
even when their device is available.
Add a shared DEVICE_LIST_OPTIONS constant and get_device() helper in
imagefunc.py that resolve 'auto' to ComfyUI's default device, and use them in
the VITMatte-based matting nodes and the VQA model loader so non-CUDA devices
can be selected from the UI.
Several model-loading and processing paths pick their device with
`"cuda" if torch.cuda.is_available() else "cpu"`, and move tensors with
`.cuda()` guarded by the same check:
- `imagefunc.load_RMBG_model` / `RMBG`
- `imagefunc.get_masked_VITMatte_image` (vitmatte)
- `imagefunc.UformGen2QwenChat`
- `imagefunc.clear_memory`
- `blendmodes.rgb_to_hsv_via_torch` / `hsv_to_rgb_via_torch`
On any non-CUDA accelerator that ComfyUI itself supports (e.g. Ascend NPU
via torch_npu, which ComfyUI's own `comfy.model_management.get_torch_device`
returns a `torch.device("npu")` for), these paths silently fall back to the
CPU or, for `.cuda()`, move the input to a CUDA device that does not exist.
This repo already imports `comfy.model_management` (e.g. in `purge_vram.py`),
so use its device selection for consistency with the rest of ComfyUI:
- `load_RMBG_model`, `UformGen2QwenChat`, and the blendmode helpers now use
`comfy.model_management.get_torch_device()`.
- `RMBG` moves the input to the device of the (lru-cached) model itself
instead of probing CUDA, so it also follows device changes correctly.
- the vitmatte path uses the ComfyUI device unless the caller explicitly
asked for CPU; a hardcoded "cuda" request without CUDA available now logs
and uses the ComfyUI default.
- cache clearing also calls `torch.accelerator.empty_cache()` when a
non-CUDA accelerator is present (`torch.accelerator` is the device-agnostic
PyTorch API).
Verified on Ascend 910B (torch 2.14.0a0 + torch_npu 2.14.0) with
ComfyUI-master's `comfy.model_management` implementation:
`get_torch_device()` returns `torch.device("npu", 0)`, a conv2d forward and
an RMBG-style model/input flow run on `npu:0`, and
`torch.accelerator.empty_cache()` is callable. On CUDA machines the selected
device is unchanged (`cuda`).