ComfyUI-DataSet
Data research, preparation, and manipulation nodes for model trainers and artists.
Drag & drop image into your workspace for node layout
Installation
Using comfy-cli (https://github.com/yoland68/comfy-cli)
comfy node registry-install ComfyUI-DataSet- https://registry.comfy.org/publishers/daxcay/nodes/comfyui-dataset
Manual Method
- Go to your Comfyui > Custom Nodes folder path > Run CMD
- Copy and Paste this command git clone
https://github.com/daxcay/ComfyUI-DataSet.git - Then go inside ComfyUI-DataSet with cmd or open new.
- type
pip install -r requirements.txtto install the dependencies
Automatic Method with Comfy Manager
-
Inside ComfyUI > Click Manager Button on Side.
-
Click
Custom Nodes Managerand Search forDataSetand Install this node: -
Restart ComfyUI and it should be good to go
Extra Nodes Needed
- ComfyUI-JDCN (https://github.com/daxcay/ComfyUI-JDCN)
You can find DataSet under this category:
DataSet_Visualizer
The DataSet_Visualizer node is designed to visualize dataset captions. It generates graphs offering various perspectives on token analysis. The word cloud represents token frequency with different sized fonts. The network graph illustrates the relationships between tokens. The frequency graph provides an exact metric of how often each token appears in your captions.
Inputs
- TextFileContents(STRING, required): the contents of the text file to be processed.
- Seperator(['comma', 'colon', 'space', 'pipe'], required): the delimiter used to separate tags in the text file.
- WordCloudTop(INT, min: 1, max: 9999, required): the number of top tokens to be plotted in WordCloud.
- NetworkGraphTop(INT, min: 1, max: 9999, required): the number of top tokens having the highest interconnections within the captions.
- FrequencyGraphTop(INT, min: 1, max: 9999, required): the number of top tags with highest frequency from highest to lowest.
Outputs
- GraphsPaths(STRING, list): the file paths of the generated visualizations. It includes paths for: WordCloud image, NetworkGraph image, FrequencyTable image
- GraphsImages(IMAGE, list): the generated images for the visualizations which can be used with PreviewImage and SaveImage node.
DataSet_CopyFiles
The DataSet_CopyFiles node provides a method to copy files from a source folder to a destination folder using different modes: BlindCopy and CopyByDestinationFiles.
Inputs
- source_folder (STRING, default: "directory path", required): source folder path to the files.
- destination_folder (STRING, default: "directory path", required): destination folder path for the files copied.
- copy_mode (['BlindCopy', 'CopyByDestinationFiles'], required):
BlindCopy: copies all files from source to the destination folder.CopyByDestinationFiles: copies files from source folder to the destination only if there is a matching file (based on the base name) already present in the destination.
DataSet_TriggerWords
The DataSet_TriggerWords node is designed to identify and extract trigger words or phrases from text file contents. Trigger words are identified based on the presence of digits within the words.
Inputs
- TextFileContents (
STRING, required): the contents of the text file(s) to be processed. - search(
['trigger_word_only', 'trigger_word_phrase'], required): the mode of searching for trigger words:'trigger_word_only': extracts individual trigger words containing digits.'trigger_word_phrase': extracts entire phrases up to the next comma if any word in the phrase contains a digit.
Outputs
- Words: (
STRING, list) - the extracted trigger words or phrases from the text file(s).
DataSet_TextFilesLoadFromList
The DataSet_TextFilesLoadFromList node is designed to load and read contents from a list of text file paths. It extracts file names, file names without extensions, file paths, and file contents.
Inputs
- TextFilePathsList(
STRING, required): a list of file paths to the text files to be loaded. Only paths ending with.txtwill be processed.
Outputs
- TextFileNames(
STRING, list): the names of the text files. - TextFileNamesWithoutExtension(
STRING, list): the names of the text files without their extensions. - TextFilePaths(
STRING, list): the file paths of the text files. - TextFileContents(
STRING, list): the contents of the text files.
DataSet_TextFilesLoad
The DataSet_TextFilesLoad node is designed to load and read contents from text files within a specified directory. It extracts file names, file names without extensions, file paths, and file contents.
Inputs
- directory(
STRING, required): the directory path where the text files are located. The path should be specified as a string.
Outputs
- TextFileNames(
STRING, list): the names of the text files in the directory. - TextFileNamesWithoutExtension(
STRING, list): the names of the text files without their extensions. - TextFilePaths(
STRING, list): the file paths of the text files in the directory. - TextFileContents(
STRING, list): the contents of the text files in the directory.
DataSet_TextFilesSave
The DataSet_TextFilesSave node is designed to save text file contents to a specified directory with various saving modes. It supports overwriting, merging, creating new files, and merging before saving new files.
Inputs
- TextFileNames(
STRING, required): the names of the text files to be saved. - TextFileContents(
STRING, required): the contents of the text files to be saved. - destination(
STRING, required): the directory path where the text files will be saved. - save_mode(['Overwrite', 'Merge', 'SaveNew', 'MergeAndSaveNew'], required): the mode of saving the files:
Overwrite: overwrites existing files with the same name.Merge: appends content to existing files with the same name.SaveNew: saves new files with a unique name if a file with the same name already exists.MergeAndSaveNew: merges content with existing files and then saves as a new file with a unique name if a file with the same name already exists.
DataSet_FindAndReplace
The DataSet_FindAndReplace node facilitates finding and replacing specific text patterns within text file contents.
Inputs
- TextFileContents(
STRING, required): the contents of the text file(s) where the search and replace operation will be performed. - SearchFor(
STRING, default: "concept", required): the text pattern to search for within theTextFileContents. Supports multiline input. - ReplaceWith(
STRING, default: "concept", required): the replacement text for theSearchForpattern. Supports multiline input.
Outputs
- TextFileContents(
STRING, list): the modified contents of the text file(s) after performing the find and replace operation.
DataSet_PathSelector
The DataSet_PathSelector node is designed to search for files with specific extensions in one directory and then select files with matching names (excluding extensions) from another directory.
Inputs
- search_in_directory(
STRING, required): the directory to search for files. - search_for_extensions(
STRING, required): the extensions of files to search for, separated by commas (e.g.,.txt, .csv). - select_from_directory(
STRING, required): the directory to select matching files from. - select_extensions(
STRING, required): the extensions of files to select, separated by commas (e.g.,.txt, .csv).
Outputs
- SelectedNamesWithExtension(
STRING, list): the names of the selected files with their extensions. - SelectedNamesWithoutExtension(
STRING, list): the names of the selected files without their extensions. - SelectedPaths(
STRING, list): the full paths of the selected files.
DataSet_ConceptManager
The DataSet_ConceptManager node is designed to manage concepts within text file contents. It allows adding or removing specified concepts at defined positions.
Inputs
- TextFileContents(
STRING, required): the contents of the text file(s) to be processed. - Mode(
STRING, required): the mode of operation:'add'to add concepts or'remove'to remove concepts. - Concepts(
STRING, required): the concepts to add or remove, formatted as text-position pairs (e.g.,"concept1 0, concept2 2"for adding,"concept1, concept2"for removing).
Outputs
- TextFileContents(
STRING, list): the modified contents of the text file(s) after adding or removing concepts.
DataSet_OpenAIChat
The DataSet_OpenAIChat node integrates with the OpenAI API to generate responses based on given prompts using various GPT models.
Inputs
- model(STRING, required): the OpenAI model to use for generating responses. Options include
"gpt-4","gpt-4-32k","gpt-3.5-turbo", and others. - api_url(STRING, default:
"https://api.openai.com/v1"): the base URL of the OpenAI API. - api_key(STRING, required): the API key required for authentication with the OpenAI API.
- prompt(STRING, default: ""): the prompt to start the conversation or generate responses.
- token_length(INT, default: 1024): the maximum number of tokens (words) in the generated response.
Outputs
- STRING: the generated response from the OpenAI model based on the provided prompt.
DataSet_LoadImage
The DataSet_LoadImage node provides functionality to load and process images from a specified directory using Pillow and numpy.
Inputs
- image (STRING, required): the name of the image file to load from the input directory.
Outputs
- IMAGE: the loaded image.
- MASK: the mask associated with the image.
- STRING: the name of the image file.
- STRING: the name of the image file without extension.
- STRING: the full path of the image file.
- STRING: the directory path of the image file.
DataSet_SaveImage
The DataSet_SaveImage node facilitates batch saving of images to a specified directory with optional PNG metadata using Pillow and numpy.
Inputs
- Images(IMAGE, required): list of images to save.
- ImageFilePrefix(STRING, default: "Image"): prefix for the saved image filenames.
- destination(STRING): directory path where images will be saved.
DataSet_OpenAIChatImage
The DataSet_OpenAIChatImage node integrates image input with OpenAI's chat API for generating text-based responses.
Inputs
- image(IMAGE, required): image to be processed.
- image_detail(STRING, default: "high": detail level of the image ("low" or "high").
- prompt(STRING, default: ""): text prompt for the AI model.
- model(STRING, default: "gpt-4o"): OpenAI model to use ("gpt-4o", "gpt-4", etc.).
- api_url(STRING, default: "https://api.openai.com/v1"): OpenAI API endpoint URL.
- api_key(STRING): OpenAI API key for authentication.
- token_length(INT, default: 1024): maximum token length for the generated response.
Outputs
- STRING: Text-based response generated by the AI model.
DataSet_OpenAIChatImageBatch
The DataSet_OpenAIChatImageBatch class extends the functionality of DataSet_OpenAIChatImage to process batches of images with OpenAI's chat API for generating text-based responses.
Inputs
- images(IMAGE, required): list of images to be processed.
- image_detail(STRING, default: "high"): detail level of the images ("low" or "high").
- prompt(STRING, default: ""): text prompt for the AI model.
- model(STRING, default: "gpt-4o"): OpenAI model to use ("gpt-4o", "gpt-4", etc.).
- api_url(STRING, default: "https://api.openai.com/v1"): OpenAI API endpoint URL.
- api_key(STRING): OpenAI API key for authentication.
- token_length(INT, default: 1024): maximum token length for the generated response.
Outputs
- STRING: List of text-based responses generated by the AI model for each input image.
Credits ❤️
Daxton Caylor - ComfyUI Node Developer
Contact
- Twitter: @daxcay27
- Email - daxtoncaylor@gmail.com
- Discord - daxtoncaylor
- DiscordServer: https://discord.gg/Z44Zjpurjp
Support
- Buy me a coffee: https://buymeacoffee.com/daxtoncaylor
- Support me on paypal: https://paypal.me/daxtoncaylor