1.0 KiB
1.0 KiB
FastVLM-7B ComfyUI Node
A custom ComfyUI node for Apple’s FastVLM-7B vision-language model.
This node lets you pass an image + instruction and returns a generated text response.
✨ Features
- 🔹 Uses apple/FastVLM-7B from Hugging Face.
- 🔹 Accepts ComfyUI images (
B,H,W,Cfloat tensors). - 🔹 Converts to PIL image internally for model input.
- 🔹 Text input (
instruction) with multiline support. - 🔹 Output is a STRING containing the model’s answer.
- 🔹 Automatically checks
ComfyUI/models/LLM/FastVLM:- Creates folder if missing.
- Downloads model into it on first run.
- 🔹 Works on GPU (fp16) or CPU (fp32) depending on your system.
support us on Patreon Check our TBG Enhanced upscaler and Refiner for ComfyUI at https://github.com/Ltamann/ComfyUI-TBG-ETUR