fix: replace blocking ollama.generate() with streaming to eliminate 120s timeout errors

Switch to streamed generation with per-chunk timeout enforcement and
ComfyUI ProgressBar integration. Add user-configurable timeout slider
(30-600s), cold-start detection via ollama.ps(), and graceful fallback
to subprocess. Bump version to 1.1.6.
This commit is contained in:
limbicnation
2026-02-16 06:01:47 +01:00
parent b171550164
commit 6e2688f5ba
-2
View File
@@ -226,8 +226,6 @@ Format the response as a single, detailed sci-fi prompt.""",
Fetch available Ollama models with caching.
Prioritizes LoRA-enhanced models (containing 'lora', 'limbicnation', 'fine').
"""
import time
# Cache for 60 seconds
if cls._cached_models and (time.time() - cls._cache_time) < 60:
return cls._cached_models