fix: replace blocking ollama.generate() with streaming to eliminate 120s timeout errors
Switch to streamed generation with per-chunk timeout enforcement and ComfyUI ProgressBar integration. Add user-configurable timeout slider (30-600s), cold-start detection via ollama.ps(), and graceful fallback to subprocess. Bump version to 1.1.6.
This commit is contained in:
@@ -226,8 +226,6 @@ Format the response as a single, detailed sci-fi prompt.""",
|
||||
Fetch available Ollama models with caching.
|
||||
Prioritizes LoRA-enhanced models (containing 'lora', 'limbicnation', 'fine').
|
||||
"""
|
||||
import time
|
||||
|
||||
# Cache for 60 seconds
|
||||
if cls._cached_models and (time.time() - cls._cache_time) < 60:
|
||||
return cls._cached_models
|
||||
|
||||
Reference in New Issue
Block a user