Files
Limbicnation-ComfyUI-Prompt…/tests/unit
limbicnation f0aaf1314a Free Ollama VRAM after node execution
Pass keep_alive="0s" on every Ollama generate call so models are
evicted from GPU VRAM immediately instead of lingering 5 minutes,
which caused CUDA OOM when downstream diffusion models loaded.

Each node now runs async VRAM cleanup in a try/finally block. The
cleanup uses unload=True so it also evicts a model loaded via the
subprocess fallback path (which carries the default keep_alive).

Add unload_model(), release_vram(), cleanup() and cleanup_async()
to OllamaClient, plus a per-node unload toggle on PromptRefiner.
2026-06-23 04:21:47 +02:00
..