Merge pull request #2 from whatbirdisthat/imageneering

Imageneering
This commit is contained in:
whatbirdisthat
2023-10-12 19:34:17 +11:00
committed by GitHub
7 changed files with 134 additions and 22 deletions
+34 -13
View File
@@ -1,13 +1,18 @@
# cyberdolphin
The dolphin is wiring up local models and / or APIs.
## Installation
Git clone this repo into the `custom_nodes` folder. If necessary, check the pip requirements.
Git clone this repo into the `custom_nodes` folder.
If necessary, check the pip requirements. It will be necessary.
## Examples
There are workflows in the [examples folder](./examples)
![img.png](examples/img.png)
---
## Nodes
The nodes all share a config file at `settings.yaml`. Provided with the repo is the
@@ -15,6 +20,7 @@ The nodes all share a config file at `settings.yaml`. Provided with the repo is
for editing. The `settings.yaml` file is ignored by git.
### OpenAI GPT Node
**REQUIRES** STRING `user_prompt`
The text is the user portion of the gpt prompt.
@@ -28,6 +34,7 @@ Runs the prompt gpt-3.5-turbo (or a user-selected alternative) with the text.
**PRODUCES** STRING.
### OpenAI Compatible Node
**REQUIRES** STRING `text`
The text is embedded in the user prompt.
Generates an "engineered" prompt from template.
@@ -35,50 +42,65 @@ The user text is embedded in the engineered prompt.
Calls for completion of the prompt to the user-defined URL.
**PRODUCES** STRING
### OpenAI DALL·E Node
**REQUIRES** STRING `text`
Calls OpenAI DALL·E with the text.
**PRODUCES** STRING
---
### Pip requirements
This collection has some extra requirements that are not present in the ComfyUI distribution.
Things like openai, gradio-client and technologist tools.
### Experimental
This is an experimental collection of nodes. This project needs validation on MacOS, Windows and Linux.
So far, it works on my machine which is a Linux distribution.
## Contributions
Looking for participants, happy to work on PRs!
**Guidelines for the Dolphin:**
**Guidelines for the Dolphin:**
* Keep it small - PRs should be quick and easy.
* Large things must be compositions of smaller things.
* Dependencies should be external - i.e. loaded by a node
* For example:
* _the Llava loader node passes the Llava model to the recogniser node which uses the Llava model to emit a list of objects_
* _and not, the "Llava node does everything"_
* _the Llava loader node passes the Llava model to the recogniser node which uses the Llava model to emit a list of
objects_
* _and not, the "Llava node does everything"_
**Keep it small**
In the spirit of "Keep it small", I'm trying to make sure my big ideas for the dolphin
stay within the realm of LLMs -
In the spirit of "Keep it small", I'm trying to make sure my big ideas for the dolphin
stay within the realm of LLMs -
Here are some big ideas that didn't make it into the roadmap for CyberDolphin:
#### Big Ideas I have for future things that are not the dolphin:
**Cam Nodes**
* **Webcam Node** for phone/laptop
* **Cam Node** for hdmi type input devices
* **Live Stream Node** to capture vision from a _Thing of the Internet_
* **Webcam Node** for phone/laptop
* **Cam Node** for HDMI type input devices
* **Live Stream Node** to capture vision from a _Thing of the Internet_
**Speech to Text**
* **Microphone node** Captures spoken instructions into **audio node**
* Instructions are transcribed using
* **OpenAI-Whisper node** or
* _TTS model_ loaded by the **TTS node**
* **OpenAI-Whisper node** or
* _TTS model_ loaded by the **TTS node**
**The Simple Storybook Production Kit**
Where "LLM-node" is short for "LLM powered node":
```text
LLM-node dreams up the story type
LLM-node dreams up the character names, their badge
@@ -97,7 +119,6 @@ LLM-node generates prompt for page illustration
LLM-node generates page text
```
## License
GPL 3.
+3
View File
@@ -2,12 +2,14 @@ from .cyberdolphin_gradio import CyberDolphinGradioApi
from .cyberdolphin_openai_advanced import CyberdolphinOpenAIAdvanced
from .cyberdolphin_openai_simple import CyberdolphinOpenAISimple
from .cyberdolphin_openai_compatible import CyberdolphinOpenAICompatible
from .cyberdolphin_imageneering import CyberDolphinImageneering
NODE_CLASS_MAPPINGS = {
"🐬 Gradio ChatInterface": CyberDolphinGradioApi,
"🐬 OpenAI Simple": CyberdolphinOpenAISimple,
"🐬 OpenAI Advanced": CyberdolphinOpenAIAdvanced,
"🐬 OpenAI Compatible": CyberdolphinOpenAICompatible,
"🐬 OpenAI DALL·E": CyberDolphinImageneering,
}
NODE_DISPLAY_NAME_MAPPINGS = {
@@ -15,4 +17,5 @@ NODE_DISPLAY_NAME_MAPPINGS = {
"CyberDolphin GPT-3.5 (Simple)": "🐬 CyberDolphin GPT-3.5 (Simple)",
"CyberDolphin OpenAI (Advanced)": "🐬 CyberDolphin OpenAI (Advanced)",
"CyberDolphin OpenAI Compatible": "🐬 CyberDolphin OpenAI Compatible",
"CyberDolphin OpenAI DALL·E": "🐬 CyberDolphin OpenAI DALL·E",
}
+62
View File
@@ -0,0 +1,62 @@
import hashlib
from PIL import ImageOps
import torch
import numpy as np
import folder_paths
from .openai_client import OpenAiClient
class CyberDolphinImageneering:
@classmethod
def INPUT_TYPES(s):
return {"required": {
"prompt": ('STRING', {'default': 'a white siamese cat'}),
"size": (["256x256", "512x512", "1024x1024"], {'default': "1024x1024"}),
}}
CATEGORY = "🐬 CyberDolphin"
RETURN_TYPES = ("IMAGE", "MASK")
FUNCTION = "load_image"
def load_image(self, prompt: str, size: str):
"""
Loads an image from api.openai.com.
https://platform.openai.com/docs/api-reference/images/create
Args:
prompt: A text description of the desired image. The maximum length is 1000 characters.
size: Must be one of: 256x256, 512x512, 1024x1024
Returns:
"""
i = OpenAiClient.image_create(prompt=prompt, size=size)
# copy/pasted from class LoadImage:
i = ImageOps.exif_transpose(i)
image = i.convert("RGB")
image = np.array(image).astype(np.float32) / 255.0
image = torch.from_numpy(image)[None,]
if 'A' in i.getbands():
mask = np.array(i.getchannel('A')).astype(np.float32) / 255.0
mask = 1. - torch.from_numpy(mask)
else:
mask = torch.zeros((64, 64), dtype=torch.float32, device="cpu")
return image, mask.unsqueeze(0)
@classmethod
def IS_CHANGED(s, image):
image_path = folder_paths.get_annotated_filepath(image)
m = hashlib.sha256()
with open(image_path, 'rb') as f:
m.update(f.read())
return m.digest().hex()
# @classmethod
# def VALIDATE_INPUTS(s, image):
# # if not folder_paths.exists_annotated_filepath(image):
# # return "Invalid image file: {}".format(image)
#
# return True
+1 -4
View File
@@ -14,10 +14,7 @@ class CyberdolphinOpenAIAdvanced:
the_settings = load_settings()
gpt_prompt = the_settings['prompt_templates']['gpt-3.5-turbo']
example_system_prompt = gpt_prompt['system']
example_user_prompt = f"\
{gpt_prompt['prefix']}\
{the_settings['example_user_prompt']}\
{gpt_prompt['suffix']}"
example_user_prompt = f"{gpt_prompt['prefix']}{the_settings['example_user_prompt']}{gpt_prompt['suffix']}"
return {
"required": {
+11 -5
View File
@@ -1,11 +1,17 @@
# Examples
## Cyber Dolphin
![img.png](img.png)
[dolphin-openai workflow](./dolphin-openai.workflow.json)
---
![Comfy-PNG](./Gpt-3.5-turbo-gen_00005_.png)
_(comfyui png depends on a stable api)
---
Get image prompt from OpenAI or a compatible API. By default `gpt-3.5-turbo` is used.
![img.png](img.png)
---
Compare the results of DALL·E and CLIP to see which one is better at generating images from text.
![img_1.png](img_1.png)
---
Binary file not shown.

After

Width:  |  Height:  |  Size: 638 KiB

+23
View File
@@ -1,3 +1,4 @@
import PIL.Image
import openai
from custom_nodes.cyberdolphin.settings import api_settings
@@ -18,8 +19,30 @@ def validation(temperature: float, top_p: float = None) -> list[str]:
return errors_list
def convert_bson_to_image(image_bson: str) -> PIL.Image.Image:
import base64
import io
from PIL import Image
image_bytes = base64.b64decode(image_bson)
image = Image.open(io.BytesIO(image_bytes))
return image
class OpenAiClient:
@staticmethod
def image_create(prompt: str, size: str = "1024x1024", api='openai') -> PIL.Image.Image:
openai.api_base, openai.api_key, openai.organization = api_settings(api)
response = openai.Image.create(
n=1,
size=size,
prompt=prompt,
response_format="b64_json"
)
image_bson = response['data'][0]['b64_json']
i = convert_bson_to_image(image_bson)
return i
@staticmethod
def complete(key: str, model: str, temperature: float, top_p: float, system_content: str, user_content: str):
errors = validation(temperature, top_p)