First Commit

This commit is contained in:
Sinphaltimus
2025-05-09 07:42:44 -04:00
commit 680e004bf0
10 changed files with 369 additions and 0 deletions
+9
View File
@@ -0,0 +1,9 @@
MIT License
Copyright (c) 2025 Reverend Pope of the Poconos Doktor Sinphaltimus Exmortus of the First Ever Digital Church Of Mind Slack, Destroyer of Chairs, Splitter of Aircraft, Raiser of Packs, Pixel Pushing, Sound Dabbling, AI Fondling Enabler of AiRTwerx, Yeti Bellowing, Ambassador of Slack.
Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS," WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING FROM, OUT OF, OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
+55
View File
@@ -0,0 +1,55 @@
# Function to validate user-input paths
function Get-ValidPath($prompt) {
do {
$path = Read-Host $prompt
if (Test-Path $path) {
return $path
} else {
Write-Host "❌ Invalid path. Please enter a valid directory." -ForegroundColor Red
}
} while ($true)
}
# Function to filter directories that contain files
function Get-FullPaths($source) {
Get-ChildItem -Path $source -Recurse | Where-Object { $_.PSIsContainer -eq $false } | ForEach-Object { $_.FullName }
}
# Get valid model directory path from user
$source = Get-ValidPath "Please enter the location of your models directory (eg. c:\ComfyUI_windows_portable\ComfyUI\models):"
# Get valid save destination from user
$destination = Get-ValidPath "Please enter the location where you would like to save your model_list.txt file (eg. c:\temp):"
# Generate output file path
$mlistPath = "$destination\model_list.txt"
# Check if model_list.txt already exists
if (Test-Path $mlistPath) {
do {
Write-Host "`n⚠️ model_list.txt already exists at $mlistPath"
$choice = Read-Host "Do you want to overwrite it? (Y/N)"
switch ($choice.ToUpper()) {
"Y" {
Write-Host "✏️ Overwriting existing file..." -ForegroundColor Yellow
Remove-Item $mlistPath -Force
break
}
"N" {
$destination = Get-ValidPath "Please enter a new save location:"
$mlistPath = "$destination\model_list.txt"
break
}
default {
Write-Host "❌ Invalid choice! Please enter Y or N." -ForegroundColor Red
}
}
} while ($choice -notmatch "^[YN]$")
}
# Generate the model list excluding empty directories
Write-Host "`n📂 Scanning directory: $source"
Get-FullPaths $source | Out-File $mlistPath
Write-Host "`n✅ Model list saved successfully at: $mlistPath" -ForegroundColor Green
+2
View File
@@ -0,0 +1,2 @@
This is a simple powershell script to generate a model_list file for you to copy paste from.
When I get better with Python, I plan to have better functionality built into the nodes themselves.
+57
View File
@@ -0,0 +1,57 @@
Readme.md
Workflow Guide: Extracting Model Metadata
This workflow begins with running Model_Lister_with_paths.ps1, which lists all available model files along with their paths. Use this output to copy-paste the file paths into each node above for metadata extraction.
You can preload requirements if you like.
.\ComfyUI_windows_portable\python_embeded\python.exe -m pip install -r requirements.txt
(For Windows ComfyUI Portable as an example)
Simply copy from the mocel_list.txt file and paste it into one or all nodes above. Connect the String Outputs from anyone of the three nodes to the string input connector of the Display String Node.
Click RUN and wait. As you progess to Enahnced and Advanced nodes, the data extraction times will increase. Also, for large models, expect log extraction times. Please be patient and let the workflow finish.
1️⃣ Model Metadata Reader
Purpose:
Extract basic metadata from models in various formats, including Safetensors, Checkpoints (.ckpt, .pth, .pt, .bin).
Provides a structured metadata report for supported model formats.
Detects the model format automatically and applies the correct extraction method.
How It Works: ✔ Reads metadata from Safetensors models using safetensors.safe_open(). ✔ Extracts available keys from Torch-based models (ckpt, .bin, .pth). ✔ Returns structured metadata when available, otherwise reports unsupported formats. ✔ Logs errors in case extraction fails.
Use this node for a quick overview of model metadata without deep metadata parsing.
2️⃣ Enhanced Model Metadata Reader
Purpose:
Extract deep metadata from models, including structured attributes and raw text parsing.
Focuses heavily on ONNX models, using direct binary parsing to retrieve metadata without relying on the ONNX Python package.
How It Works: ✔ Reads ONNX files as raw binary, searching for readable metadata like author, description, version, etc. ✔ Extracts ASCII-readable strings directly from the binary file if structured metadata isn't available. ✔ Provides warnings when metadata is missing but still displays raw extracted text. ✔ Enhanced logging for debugging failed extractions and unsupported formats.
This node is ideal for ONNX models, offering both metadata and raw text extraction for deeper insights.
3️⃣ Advanced Model Data Extractor
Purpose:
Extract structured metadata and raw text together from various model formats.
Supports Safetensors, Checkpoints (.ckpt, .pth, .bin, .gguf, .onnx).
How It Works: ✔ Extracts metadata for Safetensors using direct access to model properties. ✔ Retrieves Torch model metadata such as available keys. ✔ Attempts raw text extraction from the binary file using character encoding detection (chardet). ✔ Limits raw text output for readability while keeping detailed extraction logs.
This node provides both metadata and raw text from models, making it the most comprehensive extraction tool in the workflow.
🚀 Final Notes
Run Model_Lister_with_paths.ps1 first, then copy a model path into each node.
Use ModelMetadataReader for quick metadata lookup.
Use EnhancedModelMetadataReader for deep metadata parsing, especially for ONNX models.
Use AdvancedModelDataExtractor for full metadata + raw text extraction.
File diff suppressed because one or more lines are too long
+15
View File
@@ -0,0 +1,15 @@
from .modelmetadatareader import ModelMetadataReader
from .modeldataextractor import ModelDataExtractor
from .enhancedmodelmetadatareader import EnhancedModelMetadataReader # ✅ New Node Added
NODE_CLASS_MAPPINGS = {
"ModelMetadataReader": ModelMetadataReader,
"ModelDataExtractor": ModelDataExtractor,
"EnhancedModelMetadataReader": EnhancedModelMetadataReader, # ✅ Third Node Included
}
NODE_DISPLAY_NAME_MAPPINGS = {
"ModelMetadataReader": "Model Metadata Reader",
"ModelDataExtractor": "Advanced Model Data Extractor",
"EnhancedModelMetadataReader": "Enhanced Model Metadata Reader", # ✅ Custom UI Label
}
+67
View File
@@ -0,0 +1,67 @@
import os
import re
import json
import datetime
class EnhancedModelMetadataReader:
CATEGORY = "Model Tools"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"model_path": ("STRING", {"default": "Enter full model path here", "trigger": True}),
}
}
RETURN_TYPES = ("STRING",)
FUNCTION = "extract_data"
def extract_data(self, model_path):
"""Extracts metadata and raw text from an ONNX file with enhanced logging."""
log_entries = []
log_entries.append(f"📌 Starting Metadata Extraction: {self.get_timestamp()}")
log_entries.append(f"🔎 Checking model path: {model_path}")
if not os.path.exists(model_path):
log_entries.append("❌ Error: Model not found.")
return ("\n".join(log_entries),)
# ✅ Extract Metadata and Raw Text from ONNX
metadata, raw_text = self.extract_onnx_metadata(model_path)
log_entries.append("📂 Metadata extraction method: Direct Binary Parsing")
log_entries.append(f"✅ Extraction Complete: {self.get_timestamp()}")
return ("\n".join(log_entries) + "\n\n" + json.dumps(metadata, indent=4) + "\n\n🔍 Extracted Raw Text:\n" + raw_text[:2000],)
def extract_onnx_metadata(self, file_path):
"""Extracts readable metadata and raw text from an ONNX file using direct binary parsing."""
try:
with open(file_path, "rb") as f:
data = f.read()
# Extract human-readable text sections
extracted_text = re.findall(rb'[ -~]{4,}', data) # Captures ASCII-readable characters
decoded_text = [text.decode("utf-8", errors="ignore") for text in extracted_text]
# Look for potential metadata-related fields
metadata_keys = ["author", "description", "license", "version", "model_name"]
metadata_found = {key: value for value in decoded_text if any(key in value.lower() for key in metadata_keys)}
# Combine raw extracted text
raw_text_output = "\n".join(decoded_text)
return metadata_found if metadata_found else {"warning": "No structured metadata found."}, raw_text_output
except Exception as e:
return {"error": f"Failed to extract metadata: {str(e)}"}, ""
def get_timestamp(self):
"""Returns formatted timestamp."""
return datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
NODE_CLASS_MAPPINGS = {
"EnhancedModelMetadataReader": EnhancedModelMetadataReader
}
+85
View File
@@ -0,0 +1,85 @@
import os
import json
import torch
import safetensors
import chardet
import datetime
class ModelDataExtractor:
CATEGORY = "Model Tools"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"model_path": ("STRING", {"default": "Enter full model path here", "trigger": True}),
}
}
RETURN_TYPES = ("STRING",)
FUNCTION = "extract_data"
def extract_data(self, model_path):
"""Extracts metadata and raw text with enhanced logging."""
log_entries = []
log_entries.append(f"📌 Starting Data Extraction: {self.get_timestamp()}")
log_entries.append(f"🔎 Checking model path: {model_path}")
if not os.path.exists(model_path):
log_entries.append("❌ Error: Model not found.")
return ("\n".join(log_entries),)
extracted_data = {}
# ✅ Extract structured metadata based on model format
if model_path.endswith(".safetensors"):
extracted_data["structured_metadata"] = self.read_safetensors_metadata(model_path)
log_entries.append("📂 Metadata extraction method: Safetensors")
elif model_path.endswith((".ckpt", ".pth", ".pt", ".bin", ".gguf", ".onnx")):
extracted_data["structured_metadata"] = self.read_torch_metadata(model_path)
log_entries.append("📂 Metadata extraction method: Checkpoint/Torch")
# ✅ Always attempt raw text extraction
extracted_data["raw_text"] = self.extract_raw_text(model_path)
log_entries.append("📂 Attempting raw text extraction.")
log_entries.append(f"✅ Extraction Complete: {self.get_timestamp()}")
return ("\n".join(log_entries) + "\n\n" + json.dumps(extracted_data, indent=4),)
def read_safetensors_metadata(self, model_path):
"""Reads metadata from Safetensors models."""
try:
with safetensors.safe_open(model_path, framework="pt") as f:
return f.metadata()
except Exception as e:
return {"error": f"Safetensors extraction failed: {str(e)}"}
def read_torch_metadata(self, model_path):
"""Reads metadata from Torch-based models."""
try:
model_data = torch.load(model_path, map_location="cpu")
return {"metadata_keys": list(model_data.keys())}
except Exception as e:
return {"error": f"Torch model extraction failed: {str(e)}"}
def extract_raw_text(self, model_path):
"""Attempts raw text extraction from model binaries."""
try:
with open(model_path, "rb") as f:
data = f.read()
encoding = chardet.detect(data)["encoding"]
text_data = data.decode(encoding, errors="ignore") if encoding else "Encoding not detected"
return text_data[:2000] # ✅ Increased limit for better visibility
except Exception as e:
return f"Error extracting raw text: {str(e)}"
def get_timestamp(self):
"""Returns formatted timestamp."""
return datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
NODE_CLASS_MAPPINGS = {
"ModelDataExtractor": ModelDataExtractor
}
+71
View File
@@ -0,0 +1,71 @@
import os
import json
import torch
import safetensors
import datetime
class ModelMetadataReader:
CATEGORY = "Model Tools"
@classmethod
def INPUT_TYPES(cls):
return {
"required": {
"model_path": ("STRING", {"default": "Enter full model path here", "trigger": True}),
}
}
RETURN_TYPES = ("STRING",)
FUNCTION = "get_metadata"
def get_metadata(self, model_path):
"""Retrieves metadata from a model file with enhanced logging."""
log_entries = []
log_entries.append(f"📌 Starting Metadata Extraction: {self.get_timestamp()}")
log_entries.append(f"🔎 Checking model path: {model_path}")
if not os.path.exists(model_path):
log_entries.append("❌ Error: Model not found.")
return ("\n".join(log_entries),)
metadata = {}
# ✅ Extract metadata based on file type
if model_path.endswith(".safetensors"):
metadata = self.read_safetensors_metadata(model_path)
log_entries.append("📂 Metadata extraction method: Safetensors")
elif model_path.endswith((".ckpt", ".pth", ".pt", ".bin")):
metadata = self.read_torch_metadata(model_path)
log_entries.append("📂 Metadata extraction method: Checkpoint/Torch")
else:
metadata = {"error": "Unsupported model format"}
log_entries.append("⚠ Unsupported model format detected.")
log_entries.append(f"✅ Extraction Complete: {self.get_timestamp()}")
return ("\n".join(log_entries) + "\n\n" + json.dumps(metadata, indent=4),)
def read_safetensors_metadata(self, model_path):
"""Reads metadata from Safetensors models."""
try:
with safetensors.safe_open(model_path, framework="pt") as f:
return f.metadata()
except Exception as e:
return {"error": f"Safetensors extraction failed: {str(e)}"}
def read_torch_metadata(self, model_path):
"""Reads metadata from Torch-based models."""
try:
model_data = torch.load(model_path, map_location="cpu")
return {"metadata_keys": list(model_data.keys())}
except Exception as e:
return {"error": f"Torch model extraction failed: {str(e)}"}
def get_timestamp(self):
"""Returns formatted timestamp."""
return datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
NODE_CLASS_MAPPINGS = {
"ModelMetadataReader": ModelMetadataReader
}
+7
View File
@@ -0,0 +1,7 @@
os
json
torch
safetensors
re
datetime
chardet