Replace torch.stack(frames) with a pre-allocated tensor filled in-place.
torch.stack keeps the full frame list alive while allocating a second
equally-sized tensor, doubling peak RAM. The new approach keeps peak at
~1× the output tensor size, allowing 482-frame 4K video to fit within
machine RAM limits.
Made-with: Cursor
* feat(config): add Memory Hygiene config scaffolding and defaults for nilor-nodes
- introduce MemoryHygieneConfig and wire into NilorNodesConfig
- parse NILOR_MEMORY_HYGIENE_* from JSON5 and apply env overrides
- add validation for thresholds, policy, cooldowns, retries, and durations
- extend config.json5 with sane Memory Guardian defaults
- update .env.example
* feat(client): add ComfyUI capability detection and supports_hygiene cache
- add one-time probe for /system_stats and /free, cached per session
- expose supports_hygiene() that logs a single warning when unsupported
- use short timeouts and no retries; mark false only on 404/405
- leave transient failures retryable by keeping capability as unknown
* feat(memory): add MemoryHygiene module with typed skeleton and API
- introduce RemediationAction and RemediationResult dataclass
- add MemoryHygiene class with DI for client/config/logger
- implement check_and_remediate skeleton with enablement/capability checks
- add safe stats helper; defer policy/remediation to later commits
- export public symbols via all
* feat(memory): add metrics collector with usage pct and vram_total
- extend SystemStats with vram_total and parse from /system_stats
- add DerivedStats and collect_metrics() computing vram/ram used pct (0–100)
- integrate metrics collection in check_and_remediate skeleton
- safe math with clamping and None handling for incomplete stats
* feat(memory): implement policy engine thresholds, cooldown, and action selection
- add cooldown handling and respect it in check_and_remediate
- detect pressure via percent or absolute MB thresholds for vram/ram
- normalize policy and choose staged initial action (auto => free)
- return actionable RemediationResult with reason; no remediation yet
- helper functions for conversions and comparisons
* feat(memory): add remediation loop with retries, time caps, and cooldown
- implement remediate_cycle with free/unload flags and staged auto escalation
- respect MAX_RETRIES, SLEEP_BETWEEN_ATTEMPTS_SECONDS, MAX_CYCLE_DURATION_SECONDS
- set cooldown after cycle; return after-stats, attempts, action, and outcome reason
- integrate cycle into check_and_remediate; keep helpers in module
* feat(worker): wire MemoryHygiene into idle and post-completion paths
- initialize MemoryHygiene with comfy client and config in consume_jobs
- run hygiene before polling when idle and after prompt finalize with 1s debounce
- guard remediation by setting is_busy to block new intake; reset after
- keep websocket listener unaffected; schedule post-completion hygiene as background task
* feat(worker): throttle hygiene checks and add cadence tracking
- throttle idle hygiene by cfg.hygiene.idle_poll_seconds using monotonic clock
- add last_hygiene_check_ts to avoid overly frequent checks
- keep is_busy gating and post-completion debounce execution
* chore(memory): add structured logs for decisions and remediation
- log disabled/unsupported/cooldown/no-pressure branches
- log start/end of remediation cycles with before/after VRAM/RAM stats
- log each /free invocation flags; guard logging to avoid exceptions
* fix(memory): harden hygiene with session disable and single-warning on unsupported
- add _capability_disabled to short-circuit future runs after unsupported endpoints
- emit a single warning then quietly skip further cycles for the session
- preserve existing retry/backoff/cooldown and safe exception handling
* .env.example update
* improved get_system_stats and added more logging
* improve logging formatting for comfyui_client and memory_hygiene
* feat(nilor-nodes): add startup hygiene summary log; convert hygiene lambda to class method
- add WorkerConsumer._run_memory_hygiene() async method; remove late-bound lambda injection
- delegate to guarded helper to respect busy gate and optional debounce
- emit startup hygiene summary with effective thresholds from _CFG.hygiene
(enabled, idle_poll_s, vram/ram pct caps, min_free_mb, policy, retries, cooldown,
sleep_between, max_cycle)
- keep existing call sites in consume_loop() and _finalize_prompt() using the new method
- no functional changes to remediation logic; new log improves observability at boot
* got rid of redundant .env loading in worker_consumer
* chore(config): simplify global config caching
- keep process-wide _CONFIG singleton for shared configuration instance
- no functional behavior change to config loading paths
* fixed valueerror logging in media_stream
- load shared NilorNodesConfig in media_stream and use cfg.worker for SQS client
- replace endpoint/credentials/region env reads with typed config fields
- keep logger LOG_LEVEL and package SQS_ENABLED gate unchanged by design
- copy content_id, venue, canvas, scene from job payload into status updates
- use running_status for first progress; fail_status on execution errors (fallbacks preserved)
- manage per-content context lifecycles
The purpose of this code is to notify the backend that a specific output file has been successfully generated and uploaded. The backend (ComfyUIContentHandler) needs to know which output file this message corresponds to.
The original code sent the entire final_outputs_dict. This would work, but it's inefficient and sends redundant information. If a workflow has five MediaStreamOutput nodes, each one would send a completion message containing the information for all five outputs. The backend would receive five identical messages.
The new code is more precise. It filters the dictionary to include only the key-value pair for the output it just handled. This is a much cleaner and more correct approach. It ensures that each completion message is atomic and only contains the information relevant to the event that triggered it.
Removes the MASK output from the MediaStreamInput node to simplify its API and align with the capabilities of the Brain API server.
- The `RETURN_TYPES` is now just `("IMAGE",)`.
- All internal processing methods (`_process_image`, `_process_video`, `_process_image_batch`) have been updated to no longer extract or generate mask data.
- This change simplifies the node's logic and removes an unused feature, improving maintainability.
feat:
- Add image_batch format support to MediaStreamInput node INPUT_TYPES
- Implement two-phase download: fetch manifest first, then download individual assets
- Add _process_image_batch method for converting multiple images to tensor batches
- Sort assets by sequence number from manifest to maintain proper ordering
- Add comprehensive error handling for network failures during asset downloads
- Preserve backward compatibility for existing single-file image and video workflows
- Create proper tensor concatenation along batch dimension for ComfyUI processing
- Handle varying image formats and alpha channels within batches consistently
- Add detailed logging for manifest processing and batch creation debugging
Completes Phase 4 of multi-image support plan enabling end-to-end batch processing from brain_rnd manifest generation to ComfyUI tensor consumption.
This commit aligns the nilor-nodes with the project's new unified, name-based I/O system, as specified in the workflow override fix plan. This change establishes a stable, human-readable API contract for all workflows, replacing the previous fragile node-ID-based system.
Key Changes:
- **`MediaStreamInput` & `NilorUserInput`**: Added a static, non-overridable `input_name` string widget. Workflow authors now assign a logical name to each input, which is used by the Brain API to inject data.
- **`MediaStreamOutput`**: Added a static `output_name` widget. This provides a stable key for the Brain API to identify and retrieve specific outputs.
- **`MediaStreamOutput` (Logic)**: Corrected the completion logic to properly parse the full dictionary of named outputs it receives from the Brain API, ensuring it sends the correct, complete payload upon job completion.
These changes are a critical part of the larger refactor to improve the security, scalability, and maintainability of the ComfyUI integration.