20 Commits
Author SHA1 Message Date
Sebastian Monroy 552252b84c worker consumer normalizes incoming workflows based on operating system 2026-01-15 21:17:35 +00:00
Sebastian Monroy 266fe598c6 tweaks to config.json5 and .env.example 2026-01-07 13:00:46 +00:00
Sebastian Monroy 527e8dc4ab 663 implement comfyui memory guardian to prevent oom errors (#12)
* feat(config): add Memory Hygiene config scaffolding and defaults for nilor-nodes

- introduce MemoryHygieneConfig and wire into NilorNodesConfig
- parse NILOR_MEMORY_HYGIENE_* from JSON5 and apply env overrides
- add validation for thresholds, policy, cooldowns, retries, and durations
- extend config.json5 with sane Memory Guardian defaults
- update .env.example

* feat(client): add ComfyUI capability detection and supports_hygiene cache

- add one-time probe for /system_stats and /free, cached per session
- expose supports_hygiene() that logs a single warning when unsupported
- use short timeouts and no retries; mark false only on 404/405
- leave transient failures retryable by keeping capability as unknown

* feat(memory): add MemoryHygiene module with typed skeleton and API

- introduce RemediationAction and RemediationResult dataclass
- add MemoryHygiene class with DI for client/config/logger
- implement check_and_remediate skeleton with enablement/capability checks
- add safe stats helper; defer policy/remediation to later commits
- export public symbols via all

* feat(memory): add metrics collector with usage pct and vram_total

- extend SystemStats with vram_total and parse from /system_stats
- add DerivedStats and collect_metrics() computing vram/ram used pct (0–100)
- integrate metrics collection in check_and_remediate skeleton
- safe math with clamping and None handling for incomplete stats

* feat(memory): implement policy engine thresholds, cooldown, and action selection

- add cooldown handling and respect it in check_and_remediate
- detect pressure via percent or absolute MB thresholds for vram/ram
- normalize policy and choose staged initial action (auto => free)
- return actionable RemediationResult with reason; no remediation yet
- helper functions for conversions and comparisons

* feat(memory): add remediation loop with retries, time caps, and cooldown

- implement remediate_cycle with free/unload flags and staged auto escalation
- respect MAX_RETRIES, SLEEP_BETWEEN_ATTEMPTS_SECONDS, MAX_CYCLE_DURATION_SECONDS
- set cooldown after cycle; return after-stats, attempts, action, and outcome reason
- integrate cycle into check_and_remediate; keep helpers in module

* feat(worker): wire MemoryHygiene into idle and post-completion paths

- initialize MemoryHygiene with comfy client and config in consume_jobs
- run hygiene before polling when idle and after prompt finalize with 1s debounce
- guard remediation by setting is_busy to block new intake; reset after
- keep websocket listener unaffected; schedule post-completion hygiene as background task

* feat(worker): throttle hygiene checks and add cadence tracking

- throttle idle hygiene by cfg.hygiene.idle_poll_seconds using monotonic clock
- add last_hygiene_check_ts to avoid overly frequent checks
- keep is_busy gating and post-completion debounce execution

* chore(memory): add structured logs for decisions and remediation

- log disabled/unsupported/cooldown/no-pressure branches
- log start/end of remediation cycles with before/after VRAM/RAM stats
- log each /free invocation flags; guard logging to avoid exceptions

* fix(memory): harden hygiene with session disable and single-warning on unsupported

- add _capability_disabled to short-circuit future runs after unsupported endpoints
- emit a single warning then quietly skip further cycles for the session
- preserve existing retry/backoff/cooldown and safe exception handling

* .env.example update

* improved get_system_stats and added more logging

* improve logging formatting for comfyui_client and memory_hygiene

* feat(nilor-nodes): add startup hygiene summary log; convert hygiene lambda to class method

- add WorkerConsumer._run_memory_hygiene() async method; remove late-bound lambda injection
- delegate to guarded helper to respect busy gate and optional debounce
- emit startup hygiene summary with effective thresholds from _CFG.hygiene
  (enabled, idle_poll_s, vram/ram pct caps, min_free_mb, policy, retries, cooldown,
  sleep_between, max_cycle)
- keep existing call sites in consume_loop() and _finalize_prompt() using the new method
- no functional changes to remediation logic; new log improves observability at boot

* got rid of redundant .env loading in worker_consumer

* chore(config): simplify global config caching

- keep process-wide _CONFIG singleton for shared configuration instance
- no functional behavior change to config loading paths

* fixed valueerror logging in media_stream
2025-10-20 13:56:38 +01:00
Sebastian Monroy 5668e09a79 removed legacy code that only existed for controlling rollout, idc bout dat 2025-10-16 20:51:03 +01:00
Sebastian Monroy a94c1a9d60 update .env.example and config.json5 2025-10-16 20:39:33 +01:00
Sebastian Monroy 675fce6a16 update logger to use NILOR_LOG_LEVEL instead of LOG_LEVEL 2025-10-16 20:39:22 +01:00
Sebastian Monroy c1f5eb352c update .env.example 2025-10-16 20:00:28 +01:00
Sebastian Monroy b6bf6c1b6c update .env.example 2025-10-16 18:25:19 +01:00
Sebastian Monroy 848899298c update the worker consumer to poll jobs_to_process-comfyui and include job_type data 2025-10-14 14:57:19 +01:00
Sebastian Monroy c541bc18dd commented out .env.example params that fall-back to same value already by default 2025-10-08 15:32:24 +01:00
Sebastian Monroy 7beda857c8 added SQS_POLL_WAIT_TIME to .env 2025-10-08 15:31:42 +01:00
Sebastian Monroy e8c0d451d9 added logger class 2025-10-08 15:30:09 +01:00
Sebastian Monroy 5af42118fe add SQS_ENABLED flag to .env to toggle SQS functionality related to worker_consumer.py 2025-09-25 11:02:48 +01:00
Sebastian Monroy 04ac2b655a set up websocket connection with ComfyUI so that it can report whether it has started a ComfyUI job, and then set job status to "running" via the queue 2025-09-08 14:33:55 +01:00
Sebastian Monroy cc9056e11c implement status update publishing to SQS
-   Adds configuration for the new `job_status_updates` queue in the worker consumer.
-   After successfully submitting a job to the local ComfyUI instance, the worker publishes a message with `{"status": "running"}`.
-   Includes a critical check for `client_id` in the workflow data to ensure a job ID is present before publishing.
-   Logs a warning and skips the update if `client_id` is missing, preventing silent failures.
-   Errors during the status update publication are logged but do not interrupt the primary job, maintaining system resilience.
-   Update .env.example
2025-09-08 12:16:57 +01:00
Sebastian Monroy e5b165605a update README and .env.example 2025-08-11 13:52:32 +01:00
Sebastian Monroy 2387d0e0ad update worker_consumer.py to work with new SQS requirements and update .env.example 2025-08-08 16:54:09 +01:00
Sebastian Monroy 67c5159cfc decouple ComfyUI workers from Brain API by introducing a second SQS queue to mediate job completion reporting, update requirements.txt and .env.example 2025-08-08 16:07:03 +01:00
Sebastian Monroy 8f355e360c 3.3: successful end-to-end test of client initiating comfyui job, brain api creating the job, worker consuming the job, and comfui running the job and outputting to minio 2025-08-04 13:32:35 +01:00
Sebastian Monroy 3171063c50 first implementation of image_stream_input node 2025-07-17 14:02:14 +01:00