Pass keep_alive="0s" on every Ollama generate call so models are
evicted from GPU VRAM immediately instead of lingering 5 minutes,
which caused CUDA OOM when downstream diffusion models loaded.
Each node now runs async VRAM cleanup in a try/finally block. The
cleanup uses unload=True so it also evicts a model loaded via the
subprocess fallback path (which carries the default keep_alive).
Add unload_model(), release_vram(), cleanup() and cleanup_async()
to OllamaClient, plus a per-node unload toggle on PromptRefiner.
Add PromptDualStreamRefinerNode that produces a positive and negative
prompt pair in a single pass via Ollama, intended for the shipped Q8
GGUF of qwen2-5-7b-dual-stream-prompt-lora. Reuses OllamaClient for
streaming, timeouts, progress, and llama-runner crash handling, and
parses Positive/Negative output defensively across label variants.
Includes config/Modelfile.dualstream, unit tests, node registration,
and a corrected implementation plan replacing the invalid local
transformers/PEFT approach.
- Fix _weighted_average deduplication bug (now uses weight-ratio emphasis markers)
- Replace bare except Exception with specific exception handling in OllamaClient
- Propagate API contract mismatches (TypeError/AttributeError) as RuntimeError
- Replace print() with logging.getLogger(__name__) in all new nodes
- Expose top_p in PromptRefinerNode INPUT_TYPES for consistency
- Update PR-8-REVIEW.md with fix log and merge recommendation