63 changed files with 17492 additions and 0 deletions
+7
View File
@@ -152,6 +152,13 @@ dmypy.json
# Cython debug symbols
cython_debug/
# Sites needs this small source plugin; it is not a generated build output.
!benchmarks/site/build/
!benchmarks/site/build/sites-vite-plugin.ts
# Rebuildable TensorRT engine archives are too large for Git.
benchmarks/results/*.ep
# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
+10
View File
@@ -7,6 +7,16 @@ build. It removes startup installers and global accelerator cache flushes,
adds real image/video batches and live token streaming, and uses ComfyUI model
residency and offloading.
## VLM Speed Lab
Performance work is tracked as reproducible, quality-gated iterations in the
[VLM Speed Lab](benchmarks/README.md). The first target is the default
`Qwen/Qwen3-VL-2B-Instruct`: Transformers baseline, visual-work reduction,
Flash Attention 2, compiled execution, SGLang/FlashInfer, and TensorRT-LLM.
Every promoted speedup must attach raw outputs and remain inside the declared
quality tolerance on the same checkpoint, media, prompts, and decode settings.
Planned GPU results stay visibly unreported until a run artifact exists.
## Modern model coverage
The **Modern VLM** node provides one stable interface with a deliberately
+135
View File
@@ -0,0 +1,135 @@
# Qwen3-VL cold-start research
This note separates weight loading from warm inference. The current promoted
runtime remains SGLang 0.5.10 native + Triton multimodal attention + compiled
decode at 190.5 ms end to end. FlashPack does not make a resident model decode
faster; it targets the much larger cold-start path.
## Local profile
Host: RTX 3090 24 GB, WSL2 ext4, one 4,255,140,312-byte
`Qwen/Qwen3-VL-2B-Instruct` safetensors checkpoint. Each cold sample ran in a
fresh process after `POSIX_FADV_DONTNEED` was applied only to the measured file.
Conversion to FlashPack was excluded. Three tensors spanning the packed file
were checked bit-for-bit against safetensors and all passed.
| Loader | Reader staging | Cold seconds | Cold p50 / p95 | Effective p50 | Warm p50 | Result |
| --- | --- | ---: | ---: | ---: | ---: | --- |
| safetensors | library default | 58.55, 59.31, 60.18 | 59.31 / 60.18 s | 0.574 Gbit/s | 1.058 s | control |
| safetensors fast GPU | library default | 62.31, 58.51, 60.63 | 60.63 / 62.31 s | 0.561 Gbit/s | 1.137 s | slower |
| FlashPack direct I/O | 4 readers x 2 buffers x 32 MiB = 256 MiB | 44.78, 38.82, 43.74 | **43.74 / 44.78 s** | **0.779 Gbit/s** | not applicable to direct I/O | **26.3% faster** |
| FlashPack direct I/O | 8 readers x 2 buffers x 16 MiB = 256 MiB | 45.74 (probe) | — | 0.744 Gbit/s | — | no improvement |
| FlashPack buffered legacy | bounded internal buffer | 87.24 (probe) | — | 0.390 Gbit/s | 1.00 s | cold regression |
The upstream FlashPack default at the audited `a923a6c` revision attempted
16 readers x 2 buffers x 64 MiB, a 2 GiB pinned staging pool, and failed with a
CUDA pinned-allocation out-of-memory error on this host. The local profiler
therefore defaults to the measured 256 MiB configuration. Production code
must budget pinned memory from available host and GPU pressure rather than
assuming that the upstream default is safe.
These are local storage results, not fal `/data` results. The approximately
56x gap between cold safetensors (59.31 s) and warm safetensors (1.06 s) shows
that this WSL profile is storage-bound. fal documents up to 25 Gbit/s for
FlashPack on its infrastructure, but that number must not be presented as this
model's measured startup speed until the same profiler runs inside the target
fal machine.
## What FlashPack and ComfyUI contribute
FlashPack flattens a state dictionary into large dtype-grouped blocks, reads
chunks in parallel, overlaps host reads with CUDA copies, and creates parameter
views without a second GPU allocation. fal's persistent `/data` cache makes the
packed file reusable across runners and deployments.
Current ComfyUI adds a complementary set of mechanisms:
- read-only safetensors memory maps annotated with exact file offsets;
- direct file-slice-to-device reads where AIMDO is available;
- bounded host buffers and asynchronous device copies otherwise;
- pressure-aware pinned-memory registration and eviction;
- model deduplication, residency, partial unload, and reuse;
- module-ahead prefetch with stream synchronization;
- two asynchronous offload streams by default on supported NVIDIA systems.
ComfyUI's dynamic-VRAM path is primarily a memory-capacity and model-switching
feature. For a 2B checkpoint that fits comfortably on a 24 GB GPU, eagerly
loading the complete pack once and retaining the SGLang process minimizes first
request latency. Lazy layer materialization should be an explicit low-VRAM or
multi-model mode, not the fast default.
## Proposed combined loader: FlashSlice
1. Convert the pinned checkpoint revision to one FlashPack file during image
build or a one-time `/data` preparation job. Store its index, checksum,
dtype, model revision, FlashPack revision, Torch version, and CUDA version.
2. Instantiate the model with empty/meta parameters and map each parameter to
the packed file's offset, borrowing ComfyUI's `TensorFileSlice` abstraction.
3. For the latency path, eagerly stream the entire pack through a bounded pool.
Start with a 256 MiB budget, four read workers, two buffers per worker, and
two CUDA copy streams; autotune against the target machine and checkpoint.
4. Pipeline file read, host staging, H2D copy, parameter binding, and runtime
initialization. Never allocate a second full GPU state dictionary.
5. Keep the initialized SGLang engine resident and reuse it for every ComfyUI
execution. Do not reconstruct the engine per graph run.
6. For low-VRAM or rapid model switching, retain the file-offset map and enable
ComfyUI-style layer-ahead prefetch, bounded pinning, and pressure-aware
eviction. Record this as a distinct runtime because its first-request shape
differs from the eager path.
```text
/data packed checkpoint
|
v
bounded parallel reads --> pinned ring --> 2 CUDA streams --> empty parameters
| |
+------ file offsets for optional lazy/prefetch mode -----+
|
v
resident SGLang engine
```
## End-to-end startup ladder
Every deployment benchmark should emit timestamps for these phases. A single
"cold start" duration is not actionable.
| Mark | Phase | Optimization |
| --- | --- | --- |
| T0 | request accepted | client region, upload size, connection reuse |
| T1 | runner allocated | fal `min_concurrency`, `keep_alive`, capacity |
| T2 | imports complete | small image, pinned dependencies, lazy imports |
| T3 | checkpoint available | persistent `/data`, checksum hit, no download |
| T4 | model skeleton ready | empty/meta initialization |
| T5 | weights resident | bounded FlashPack/FlashSlice pipeline |
| T6 | kernels ready | synchronized Inductor cache, GPU/version key |
| T7 | serving ready | in-process engine or explicit readiness barrier |
| T8 | first token | preprocessed fixed shape, CUDA graph/compile cache |
| T9 | final token | existing SGLang steady-state benchmark |
Recommended production sequence:
1. Measure a true zero-runner fal cold start and a `/data`-cached cold start.
2. Add the bounded packed loader; accept it only with exact tensor and output
gates.
3. Persist the compiled Inductor cache and warm the real 448-edge, batch-one
image/decode shape during setup.
4. Reuse the model process. For latency-critical traffic, compare
`min_concurrency=1` against cost; for sporadic traffic, start with a longer
`keep_alive` such as 300 seconds and measure the hit rate.
5. Stream output so perceived latency follows TTFT, resize media before upload,
and avoid base64 copies when a region-local URL is available.
## Primary sources
- [fal FlashPack optimization](https://fal.ai/docs/documentation/serverless/optimizations/flashpack)
- [fal cold-start phases](https://fal.ai/docs/documentation/serverless/optimizations/optimize-cold-starts)
- [fal compiled-cache synchronization](https://fal.ai/docs/documentation/serverless/optimizations/optimize-startup-with-compiled-caches)
- [fal cold-start scaling controls](https://fal.ai/docs/documentation/serverless/optimizations/cold-start-scaling)
- [fal parallel file loading](https://fal.ai/docs/documentation/serverless/optimizations/parallel-file-loading)
- [FlashPack source](https://github.com/fal-ai/flashpack)
- [ComfyUI tensor loading and mmap metadata](https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/utils.py)
- [ComfyUI model residency and loading](https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/model_management.py)
- [ComfyUI file-slice-to-device pipeline](https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/memory_management.py)
- [ComfyUI module prefetch](https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/model_prefetch.py)
- [ComfyUI bounded pinned memory](https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/pinned_memory.py)
+196
View File
@@ -0,0 +1,196 @@
# VLM Speed Lab
This directory turns performance work into a sequence of reproducible,
quality-gated experiments. The first target is the repository default:
`Qwen/Qwen3-VL-2B-Instruct`.
## Rule zero
A result is a speedup only when it uses the same checkpoint revision, media,
prompts, seed, precision policy, and decoding settings as its baseline, and its
task-quality score remains inside the declared tolerance. A faster result that
misses the quality gate is recorded as a regression.
## Iteration order
1. Transformers BF16 + SDPA baseline.
2. Existing adaptive sampling and pixel-budget nodes.
3. Flash Attention 2.
4. `torch.compile` / CUDA graph experiments.
5. SGLang with its declared attention backend (including FlashInfer where
selected by the runtime).
6. TensorRT component engines where the model is exportable; TensorRT-LLM only
where the upstream runtime supports the complete architecture.
Change one performance variable at a time. Run single-request latency first,
then concurrency sweeps. Never mix cold-start and steady-state samples.
## Reproduce the first RTX 3090 matrix in WSL
The committed `qwen3-vl-2b-matrix-tf5-rubric.json` artifact was generated on
Ubuntu 22.04 under WSL2 with an RTX 3090, PyTorch 2.8.0+cu128, and Transformers
5.12.1. Model files, the virtual environment, media, and results all lived on
the WSL ext4 disk rather than a `/mnt/c` or `/mnt/d` mount.
```bash
HF_ENABLE_PARALLEL_LOADING=true \
HF_PARALLEL_LOADING_WORKERS=8 \
./.venv-bench/bin/python benchmarks/qwen3_vl_matrix.py \
--image benchmarks/media/qwen-demo.jpeg \
--runs 10 \
--max-new-tokens 96 \
--output benchmarks/results/qwen3-vl-2b-matrix-tf5-rubric.json
```
Ten measured runs follow two warmups for dynamic-cache variants and six for
the compiled static-cache variant. The one-time compilation sample remains in
`warmup_samples`; it is never mixed into steady-state percentiles.
| Iteration | Input | TTFT p50 | E2E p50 | Output tok/s | Peak VRAM | Quality |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| 00 SDPA + dynamic | 2048x1365 | 700.3 ms | 1395.7 ms | 42.3 | 4.55 GiB | rubric pass |
| 01a SDPA + dynamic | 672x448 | 112.8 ms | 844.0 ms | 42.2 | 4.04 GiB | rubric pass |
| 01b SDPA + dynamic | 448x299 | 88.2 ms | 774.1 ms | 43.7 | 4.00 GiB | rubric pass |
| 02 SDPA + static compiled | 448x299 | 76.6 ms | 290.1 ms | 139.4 | 4.02 GiB | rubric + exact-output pass vs 01b |
| 03a FA2 + dynamic | 448x299 | 106.6 ms | 1033.3 ms | 32.3 | 4.00 GiB | exact pass; performance regression |
| 03b FA2 + static compiled | 448x299 | 265.0 ms | 3028.0 ms | 34.4 | 4.02 GiB | **fail; corrupted repetitive output** |
| 04 SDPA + static + scoped TF32 | 448x299 | 74.3 ms | 274.5 ms | 149.2 | 4.03 GiB | rubric + exact-output pass vs 01b |
Iteration 04 is 9.43x faster to first token, 5.08x faster end to end, and
3.53x higher output throughput than iteration 00. Resizing preserves the task
rubric but is not byte-identical to source-resolution output; the artifact
records both facts. The cache/compiler change is byte-identical to iteration
01b, as is scoped TF32. Flash Attention 2 is retained as negative evidence:
its dynamic-cache run was correct but slower, while its static-cache pairing
failed the exact-output gate. These are single-image, batch-one latency
results—not yet a general VLM quality claim.
Parallel safetensor loading reduced warm-filesystem model/processor setup from
88.351 seconds to 6.858 seconds. Treat this as a warm-cache startup result;
network download time is outside the measurement.
The separate [cold-start study](COLD_START_RESEARCH.md) profiles the same
checkpoint from disk to GPU and combines a bounded FlashPack reader with
ComfyUI's file-slice, pinned-memory, residency, and prefetch ideas. On the local
WSL host, the validated bounded FlashPack configuration reduced cold weight
loading from 59.31 seconds to 43.74 seconds p50 (26.3%). This is explicitly a
local storage result; fal `/data` remains to be measured independently.
## SGLang and FlashInfer matrix
The same 448x299 image, prompt, greedy decode, 96-token cap, RTX 3090, three
warmups, and ten measured requests were used for the serving-runtime matrix.
SGLang 0.5.10.post1 ran with PyTorch 2.9.1+cu128, Transformers 5.3.0, and
FlashInfer 0.6.7.post3. The concept gate requires the woman, golden retriever,
beach, and high-five action; inflection aliases such as `high-fiving` are
accepted within that action concept.
| Iteration | Runtime change | TTFT p50 / p95 | E2E p50 / p95 | Output tok/s | Quality |
| --- | --- | ---: | ---: | ---: | --- |
| 05a SGLang 0.5.9 native | FlashInfer + SDPA vision | 38.3 / 42.9 ms | 43.3 / 48.0 ms | 393.9 | **fail; output was only a code fence** |
| 05b SGLang 0.5.10 Transformers backend | Version + model implementation | 75.6 / 79.2 ms | 254.2 / 257.5 ms | 173.5 | pass; exact vs 01b |
| 05c SGLang 0.5.10 native | Native model implementation | 35.2 / 38.3 ms | 240.6 / 243.6 ms | 194.7 | concept pass |
| 05d Triton multimodal attention | SDPA vision -> Triton vision | 35.5 / 37.9 ms | 193.6 / 195.3 ms | 196.4 | pass; exact vs 01b |
| 05e compiled decode | `torch.compile`, max batch 4 | 37.5 / 41.0 ms | 190.5 / 194.7 ms | 202.6 | pass; exact vs 01b |
Iteration 05e is 7.33x faster end to end and delivers 4.79x higher output
throughput than iteration 00. Iteration 05c retains the best TTFT at 19.88x
faster than iteration 00, while 05e trades 2.2 ms of TTFT for the best E2E and
decode throughput. The one-request 0.5.10 cold probe took 17.6 seconds because
of one-time compilation and is kept separate from steady-state percentiles.
The 0.5.9 result demonstrates why latency cannot be promoted without output
evidence: its apparently extraordinary timing came from terminating after two
invalid tokens. The 0.5.10 release fixed the native vision path for this case.
The current 0.5.15.post1 release was also installed and audited, but its CUDA
13 / PyTorch 2.11 build cannot initialize CUDA on the machine's NVIDIA 560.94
driver, so it is recorded as incompatible rather than benchmarked.
## TensorRT vision engine
Current TensorRT-LLM does not list Qwen3-VL as a supported multimodal serving
architecture, so iteration 06 does not mislabel its PyTorch backend as a
TensorRT engine. Instead, Torch-TensorRT 2.9.0 and TensorRT 10.13.3 compile the
fixed-shape Qwen3-VL vision tower into one real BF16 engine on the RTX 3090.
The graph has zero PyTorch fallback partitions.
```bash
./.venv-tensorrt/bin/python benchmarks/qwen3_vl_tensorrt.py \
--image benchmarks/media/qwen-demo.jpeg \
--longest-edge 448 \
--warmups 3 \
--runs 10 \
--generation-warmups 1 \
--generation-runs 3 \
--output benchmarks/results/qwen3-vl-2b-tensorrt-vision-full.json
```
| Path | Vision p50 | TTFT p50 / p95 | E2E p50 / p95 | Output tok/s | Quality |
| --- | ---: | ---: | ---: | ---: | --- |
| Torch 2.9 eager control | 2385.3 ms | 2452.9 / 2464.3 ms | 2660.6 / 2674.7 ms | 143.3 | 3/3 identical |
| TensorRT vision + unchanged decoder | 9.1 ms | 61.4 / 62.4 ms | 273.4 / 274.0 ms | 142.1 | exact output vs eager |
Engine construction took 98.070 seconds and is reported separately from
inference. TensorRT produced the same 31-token sentence in every full-model
sample. Its isolated 262.8x vision speedup is real relative to the Torch 2.9
eager control but is not the cross-stack headline: the established Torch 2.8
Transformers path already runs end to end in 274.5 ms, and SGLang iteration
05e remains the overall winner at 190.5 ms. The useful result is a verified
9.1 ms vision engine and a new 61.4 ms Transformers TTFT.
Iteration 07 serializes that engine and injects its packed pooler plus three
deep-stack tensors into SGLang's native decoder. The static bridge only accepts
the compiled `(1, 18, 28)` grid; other image shapes fall back to SGLang's
unchanged vision path.
| Path | TTFT p50 / p95 | E2E p50 / p95 | Output tok/s | Semantic gate | Exact gate |
| --- | ---: | ---: | ---: | --- | --- |
| 05e SGLang control | 37.5 / 41.0 ms | **190.5 / 194.7 ms** | **202.6** | pass | pass; 31 tokens |
| 07 TensorRT + SGLang | **34.9 / 37.8 ms** | 250.7 / 366.1 ms | 176.1 | pass | **fail; 40 tokens** |
The bridge reduced TTFT by 7.0%, but numerical differences in the Transformers
vision engine changed greedy decoding to a longer, semantically correct
caption. That makes iteration 07 a measured regression rather than a promoted
speedup. The next experiment is to compile SGLang-native vision weights and
preserve the exact 31-token output.
## Run the OpenAI-compatible benchmark
SGLang and TensorRT-LLM both expose OpenAI-compatible chat endpoints. Start
one server, copy `suite.example.json`, point its cases to local benchmark media,
and run:
```bash
python benchmarks/vlm_bench.py \
--suite benchmarks/suite.local.json \
--base-url http://127.0.0.1:8000/v1 \
--backend sglang \
--label qwen3-vl-2b-sglang \
--warmups 3 \
--runs 30
```
The runner writes one immutable JSON artifact under `benchmarks/results/`.
It records raw model output, per-request latency and time-to-first-token,
aggregate percentiles, quality scores, media hashes, server identity, and the
local Git commit. Do not hand-edit result artifacts.
## Required suite fields
Each case declares a task and an evaluator:
- `keywords`: case-insensitive keyword recall for captions.
- `concepts`: required semantic concepts, each with one or more accepted aliases.
- `exact`: normalized exact match for OCR and constrained answers.
- `number`: extracts the first integer for counting tasks.
Detection, segmentation, and tracking evaluators will be added after the first
text-output baseline is frozen. Their artifacts will use the same run envelope
and add box, mask, or track data rather than creating a separate leaderboard.
## Result review
The comparison site lives in `benchmarks/site`. It shows regressions alongside
winners and never substitutes estimates for missing GPU runs.
The existing `11.38×` figure is explicitly labeled as frame-by-pixel input-work
reduction, not end-to-end model acceleration.
+1
View File
@@ -0,0 +1 @@
"""Reproducible performance benchmarks for ComfyUI VLM Nodes."""
+11
View File
@@ -0,0 +1,11 @@
# Benchmark media
`qwen-demo.jpeg` is the public demonstration image linked by the Qwen-VL
project and used only as a reproducible benchmark input.
- Source: <https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg>
- Dimensions: 2048x1365
- SHA-256: `9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2`
The image is not presented as repository-owned content. Keep its provenance
with any redistributed benchmark artifact.
Binary file not shown.

After

Width:  |  Height:  |  Size: 485 KiB

+227
View File
@@ -0,0 +1,227 @@
"""Profile cold and warm disk-to-GPU loading for Qwen3-VL weights.
FlashPack conversion is deliberately outside the timed path. Each measured
run uses a new Python process so CUDA allocator state cannot leak across runs.
"""
from __future__ import annotations
import argparse
import gc
import json
import os
import statistics
import subprocess
import sys
import time
from pathlib import Path
from typing import Any
def drop_file_cache(path: Path) -> None:
"""Ask Linux to evict this file's pages without dropping global caches."""
if not hasattr(os, "posix_fadvise"):
return
descriptor = os.open(path, os.O_RDONLY)
try:
os.posix_fadvise(descriptor, 0, 0, os.POSIX_FADV_DONTNEED)
finally:
os.close(descriptor)
def load_once(method: str, path: Path) -> dict[str, Any]:
import torch
torch.cuda.empty_cache()
torch.cuda.reset_peak_memory_stats()
started = time.perf_counter()
if method.startswith("safetensors"):
if method == "safetensors_fast_gpu":
os.environ["SAFETENSORS_FAST_GPU"] = "1"
else:
os.environ.pop("SAFETENSORS_FAST_GPU", None)
from safetensors.torch import load_file
loaded = load_file(str(path), device="cuda")
tensor_count = len(loaded)
elif method == "flashpack":
# FlashPack main currently defaults to 16 readers, two 64 MiB pinned
# buffers per reader (2 GiB total). That failed on the RTX 3090 WSL
# test host. Keep the benchmark's default bounded and let callers
# override every value explicitly when tuning another machine.
os.environ.setdefault("FLASHPACK_READ_THREADS", "4")
os.environ.setdefault("FLASHPACK_READ_CHUNK_BYTES", str(32 * 1024 * 1024))
os.environ.setdefault("FLASHPACK_CACHE_PINNED", "0")
from flashpack.deserialization import read_flashpack_file
loaded, metadata = read_flashpack_file(path=str(path), device="cuda")
tensor_count = len(metadata["index"])
else:
raise ValueError(f"Unknown method: {method}")
torch.cuda.synchronize()
elapsed = time.perf_counter() - started
peak_bytes = torch.cuda.max_memory_allocated()
del loaded
gc.collect()
torch.cuda.empty_cache()
return {
"seconds": elapsed,
"tensor_count": tensor_count,
"peak_gpu_bytes": peak_bytes,
}
def worker(method: str, path: Path, runs: int, cold_only: bool) -> None:
samples = []
for _ in range(runs):
drop_file_cache(path)
cold = load_once(method, path)
warm = None if cold_only else load_once(method, path)
samples.append({"method": method, "cold": cold, "warm": warm})
print(json.dumps(samples[0] if runs == 1 else {"samples": samples}))
def prepare(safetensors_path: Path, flashpack_path: Path) -> dict[str, Any]:
import torch
from flashpack import is_flashpack_file, pack_to_file
from flashpack.deserialization import (
iterate_from_flash_tensor,
read_flashpack_file,
)
from safetensors import safe_open
from safetensors.torch import load_file
conversion_seconds = 0.0
if not flashpack_path.exists() or not is_flashpack_file(str(flashpack_path)):
flashpack_path.parent.mkdir(parents=True, exist_ok=True)
started = time.perf_counter()
state_dict = load_file(str(safetensors_path), device="cpu")
pack_to_file(
state_dict,
str(flashpack_path),
target_dtype=None,
silent=False,
)
conversion_seconds = time.perf_counter() - started
del state_dict
gc.collect()
storage, metadata = read_flashpack_file(str(flashpack_path), device="cpu")
packed_tensors = dict(iterate_from_flash_tensor(storage, metadata))
names = list(packed_tensors)
sample_names = [names[0], names[len(names) // 2], names[-1]]
exact = {}
with safe_open(str(safetensors_path), framework="pt", device="cpu") as source:
for name in sample_names:
exact[name] = bool(torch.equal(source.get_tensor(name), packed_tensors[name]))
del packed_tensors, storage
gc.collect()
drop_file_cache(safetensors_path)
drop_file_cache(flashpack_path)
return {
"conversion_seconds": conversion_seconds,
"flashpack_bytes": flashpack_path.stat().st_size,
"tensor_count": len(metadata["index"]),
"sample_exact": exact,
}
def percentile(values: list[float], fraction: float) -> float:
ordered = sorted(values)
index = min(len(ordered) - 1, int(round((len(ordered) - 1) * fraction)))
return ordered[index]
def summarize(samples: list[dict[str, Any]], file_bytes: int) -> dict[str, Any]:
summary = {}
for cache_state in ("cold", "warm"):
seconds = [sample[cache_state]["seconds"] for sample in samples]
median = statistics.median(seconds)
summary[cache_state] = {
"seconds": seconds,
"p50_seconds": median,
"p95_seconds": percentile(seconds, 0.95),
"p50_throughput_gbps": file_bytes * 8 / median / 1e9,
"peak_gpu_bytes": max(
sample[cache_state]["peak_gpu_bytes"] for sample in samples
),
}
return summary
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--safetensors", type=Path, required=True)
parser.add_argument("--flashpack", type=Path, required=True)
parser.add_argument("--runs", type=int, default=3)
parser.add_argument("--output", type=Path)
parser.add_argument("--worker", choices=("safetensors", "safetensors_fast_gpu", "flashpack"))
parser.add_argument("--worker-runs", type=int, default=1)
parser.add_argument("--cold-only", action="store_true")
args = parser.parse_args()
if args.worker:
worker(
args.worker,
args.flashpack if args.worker == "flashpack" else args.safetensors,
args.worker_runs,
args.cold_only,
)
return
preparation = prepare(args.safetensors, args.flashpack)
methods = ("safetensors", "safetensors_fast_gpu", "flashpack")
result: dict[str, Any] = {
"schema_version": 1,
"checkpoint": "Qwen/Qwen3-VL-2B-Instruct",
"safetensors_bytes": args.safetensors.stat().st_size,
"flashpack_reader": {
"threads": int(os.environ.get("FLASHPACK_READ_THREADS", "4")),
"chunk_bytes": int(
os.environ.get("FLASHPACK_READ_CHUNK_BYTES", str(32 * 1024 * 1024))
),
"cache_pinned": os.environ.get("FLASHPACK_CACHE_PINNED", "0"),
},
"preparation": preparation,
"methods": {},
}
for method in methods:
samples = []
for _ in range(args.runs):
completed = subprocess.run(
[
sys.executable,
str(Path(__file__).resolve()),
"--safetensors",
str(args.safetensors),
"--flashpack",
str(args.flashpack),
"--worker",
method,
],
check=True,
capture_output=True,
text=True,
)
samples.append(json.loads(completed.stdout.strip().splitlines()[-1]))
file_bytes = (
preparation["flashpack_bytes"]
if method == "flashpack"
else args.safetensors.stat().st_size
)
result["methods"][method] = summarize(samples, file_bytes)
if args.output:
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(
json.dumps(result, indent=2) + "\n", encoding="utf-8"
)
payload = json.dumps(result, indent=2)
print(payload)
if args.output:
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(payload + "\n", encoding="utf-8")
if __name__ == "__main__":
main()
+213
View File
@@ -0,0 +1,213 @@
"""Run the core Qwen3-VL optimization matrix in one loaded-model process."""
from __future__ import annotations
import argparse
import json
import platform
import time
from datetime import UTC, datetime
from pathlib import Path
import torch
from PIL import Image
from qwen3_vl_transformers import aggregate, resize_to_longest_edge, run_sample
from transformers import AutoModelForImageTextToText, AutoProcessor
VARIANTS = (
{
"id": "00",
"label": "BF16 SDPA / dynamic cache / source resolution",
"longest_edge": None,
"cache": "dynamic",
"warmups": 2,
},
{
"id": "01a",
"label": "BF16 SDPA / dynamic cache / 672px edge",
"longest_edge": 672,
"cache": "dynamic",
"warmups": 2,
},
{
"id": "01b",
"label": "BF16 SDPA / dynamic cache / 448px edge",
"longest_edge": 448,
"cache": "dynamic",
"warmups": 2,
},
{
"id": "02",
"label": "BF16 SDPA / static compiled cache / 448px edge",
"longest_edge": 448,
"cache": "static",
"warmups": 6,
"exact_reference": "01b",
},
)
DEFAULT_CONCEPT_GROUPS = (
("woman", "person"),
("golden retriever", "dog"),
("beach", "sand"),
("high-five", "high five"),
)
def evaluate_concepts(output: str, groups: tuple[tuple[str, ...], ...]) -> dict:
normalized = output.casefold()
matched = [next((term for term in group if term in normalized), None) for group in groups]
return {
"passed": all(matched),
"matched": matched,
"required": [list(group) for group in groups],
}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--image", type=Path, required=True)
parser.add_argument("--model", default="Qwen/Qwen3-VL-2B-Instruct")
parser.add_argument("--prompt", default="Describe this image precisely in one sentence.")
parser.add_argument("--runs", type=int, default=10)
parser.add_argument("--max-new-tokens", type=int, default=96)
parser.add_argument(
"--output",
type=Path,
default=Path("benchmarks/results/qwen3-vl-2b-matrix-tf5.json"),
)
args = parser.parse_args()
source_image = Image.open(args.image).convert("RGB")
load_started = time.perf_counter()
processor = AutoProcessor.from_pretrained(args.model)
model = AutoModelForImageTextToText.from_pretrained(
args.model,
dtype=torch.bfloat16,
attn_implementation="sdpa",
device_map="cuda",
).eval()
torch.cuda.synchronize()
load_seconds = time.perf_counter() - load_started
results = []
output_hashes: dict[str, str] = {}
baseline_hash: str | None = None
baseline_summary = None
for variant in VARIANTS:
image = resize_to_longest_edge(source_image, variant["longest_edge"])
warmup_samples = []
measured_samples = []
total = int(variant["warmups"]) + args.runs
for index in range(total):
sample = run_sample(
model,
processor,
image,
args.prompt,
max_new_tokens=args.max_new_tokens,
cache_implementation=str(variant["cache"]),
min_pixels=None,
max_pixels=None,
disable_compile=False,
)
target = warmup_samples if index < int(variant["warmups"]) else measured_samples
target.append(sample)
print(
f"{variant['id']} {index + 1}/{total} "
f"ttft={sample['ttft_ms']:.1f}ms "
f"e2e={sample['e2e_ms']:.1f}ms "
f"tok/s={sample['output_tokens_per_second']}",
flush=True,
)
summary = aggregate(measured_samples)
if baseline_hash is None:
baseline_hash = measured_samples[0]["output_sha256"]
baseline_summary = summary
output_hashes[str(variant["id"])] = measured_samples[0]["output_sha256"]
rubric_results = [
evaluate_concepts(sample["output"], DEFAULT_CONCEPT_GROUPS)
for sample in measured_samples
]
exact_reference = variant.get("exact_reference")
exact_hash = (
output_hashes[str(exact_reference)] if exact_reference is not None else None
)
exact_passed = (
all(sample["output_sha256"] == exact_hash for sample in measured_samples)
if exact_hash is not None
else None
)
rubric_passed = all(result["passed"] for result in rubric_results)
speedup = {
"ttft": round(
baseline_summary["ttft_ms"]["p50"] / summary["ttft_ms"]["p50"], 3
),
"e2e": round(
baseline_summary["e2e_ms"]["p50"] / summary["e2e_ms"]["p50"], 3
),
"throughput": round(
summary["output_tokens_per_second_mean"]
/ baseline_summary["output_tokens_per_second_mean"],
3,
),
}
results.append(
{
**variant,
"processed_width": image.width,
"processed_height": image.height,
"quality_gate": {
"method": "required visual concepts"
+ (
f" plus byte-identical output against variant {exact_reference}"
if exact_reference is not None
else ""
),
"passed": rubric_passed and exact_passed is not False,
"concepts": rubric_results[0],
"exact_output_reference": exact_reference,
"exact_output_passed": exact_passed,
"exact_output_vs_baseline": all(
sample["output_sha256"] == baseline_hash
for sample in measured_samples
),
},
"speedup_vs_baseline": speedup,
"summary": summary,
"warmup_samples": warmup_samples,
"samples": measured_samples,
}
)
artifact = {
"schema": "comfyui-vlm/optimization-matrix",
"version": 1,
"created_at": datetime.now(UTC).isoformat(),
"model": args.model,
"media": {
"path": str(args.image.resolve()),
"source_width": source_image.width,
"source_height": source_image.height,
},
"prompt": args.prompt,
"model_load_seconds": round(load_seconds, 3),
"environment": {
"platform": platform.platform(),
"python": platform.python_version(),
"torch": torch.__version__,
"cuda": torch.version.cuda,
"gpu": torch.cuda.get_device_name(),
"transformers": __import__("transformers").__version__,
},
"runs_per_variant": args.runs,
"variants": results,
}
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(json.dumps(artifact, indent=2) + "\n", encoding="utf-8")
print(json.dumps([{v["id"]: v["summary"]} for v in results], indent=2))
print(args.output)
if __name__ == "__main__":
main()
@@ -0,0 +1,113 @@
"""Inject a serialized TensorRT Qwen3-VL vision engine into SGLang.
The bridge is deliberately static-shape and quality-safe. Requests matching the
compiled 448px benchmark grid use TensorRT; every other shape takes SGLang's
unchanged native vision path.
"""
from __future__ import annotations
import argparse
import logging
import os
import time
from pathlib import Path
from typing import Any
import torch
LOGGER = logging.getLogger("sglang.tensorrt_bridge")
ENGINE_ENV = "QWEN3_VL_TRT_ENGINE"
EXPECTED_GRID = ((1, 18, 28),)
def _load_engine(path: Path) -> torch.nn.Module:
# Importing Torch-TensorRT registers the serialized engine operators used by
# the ExportedProgram.
import torch_tensorrt # noqa: F401
started = time.perf_counter()
engine = torch.export.load(path).module().cuda()
LOGGER.info(
"Loaded Qwen3-VL TensorRT vision engine path=%s elapsed=%.3fs",
path,
time.perf_counter() - started,
)
return engine
def _grid_tuple(grid: torch.Tensor) -> tuple[tuple[int, ...], ...]:
return tuple(tuple(int(value) for value in row) for row in grid.cpu().tolist())
def install_bridge() -> bool:
engine_value = os.environ.get(ENGINE_ENV)
if not engine_value:
return False
engine_path = Path(engine_value).expanduser().resolve()
if not engine_path.is_file():
raise FileNotFoundError(f"TensorRT vision engine not found: {engine_path}")
from sglang.srt.models.qwen3_vl import Qwen3VLForConditionalGeneration
if getattr(Qwen3VLForConditionalGeneration, "_trt_bridge_installed", False):
return True
native_get_image_feature = Qwen3VLForConditionalGeneration.get_image_feature
def get_image_feature(self: Any, items: list[Any]) -> torch.Tensor:
image_grid_thw = torch.concat(
[item.image_grid_thw for item in items], dim=0
)
if _grid_tuple(image_grid_thw) != EXPECTED_GRID:
self._trt_bridge_fallbacks = getattr(self, "_trt_bridge_fallbacks", 0) + 1
return native_get_image_feature(self, items)
engine = getattr(self, "_trt_vision_engine", None)
if engine is None:
engine = _load_engine(engine_path)
self._trt_vision_engine = engine
pixel_values = torch.cat([item.feature for item in items], dim=0).to(
device="cuda", dtype=torch.bfloat16
)
outputs = engine(pixel_values.contiguous())
# Output 0 is the unmerged vision state. SGLang consumes the merged
# language embedding followed by all three packed deep-stack features.
packed = torch.cat(tuple(outputs[1:]), dim=-1)
if packed.shape != (126, 8192):
raise RuntimeError(
f"Unexpected TensorRT packed vision shape: {tuple(packed.shape)}"
)
self._trt_bridge_hits = getattr(self, "_trt_bridge_hits", 0) + 1
return packed
Qwen3VLForConditionalGeneration.get_image_feature = get_image_feature
Qwen3VLForConditionalGeneration._trt_bridge_installed = True
LOGGER.info(
"Installed static Qwen3-VL TensorRT/SGLang bridge engine=%s grid=%s",
engine_path,
EXPECTED_GRID,
)
return True
install_bridge()
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--smoke-test", type=Path)
args = parser.parse_args()
if args.smoke_test is None:
return
engine = _load_engine(args.smoke_test.resolve())
sample = torch.zeros((504, 1536), device="cuda", dtype=torch.bfloat16)
with torch.inference_mode():
outputs = engine(sample)
torch.cuda.synchronize()
print([list(output.shape) for output in outputs])
if __name__ == "__main__":
main()
+389
View File
@@ -0,0 +1,389 @@
"""Probe and benchmark a real TensorRT vision path for Qwen3-VL.
The experiment deliberately compiles only the vision tower. It reports
TensorRT graph coverage, numerical drift, isolated vision latency, and (when
conversion succeeds) can be extended to the unchanged language decoder.
"""
from __future__ import annotations
import argparse
import json
import platform
import statistics
import time
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
import torch
import torch_tensorrt
from PIL import Image
from qwen3_vl_transformers import (
aggregate,
prepare_inputs,
resize_to_longest_edge,
run_sample,
)
from transformers import AutoModelForImageTextToText, AutoProcessor
from transformers.models.qwen3_vl.modeling_qwen3_vl import (
BaseModelOutputWithDeepstackFeatures,
get_vision_bilinear_indices_and_weights,
get_vision_cu_seqlens,
get_vision_position_ids,
)
class StaticVisionTensorOutputs(torch.nn.Module):
"""Tensor-only vision tower with fixed-shape positional metadata.
Transformers derives this metadata from ``grid_thw`` using Python integer
conversions. Hoisting it is both export-safe and valid for our explicitly
static benchmark shape.
"""
def __init__(self, visual: torch.nn.Module, grid_thw: torch.Tensor) -> None:
super().__init__()
self.visual = visual
indices, weights = get_vision_bilinear_indices_and_weights(
grid_thw,
num_grid_per_side=visual.num_grid_per_side,
spatial_merge_size=visual.config.spatial_merge_size,
kwargs={},
)
position_ids = get_vision_position_ids(
grid_thw, visual.spatial_merge_size, kwargs={}
)
cu_seqlens = get_vision_cu_seqlens(grid_thw, kwargs={})
self.register_buffer("bilinear_indices", indices)
self.register_buffer("bilinear_weights", weights)
self.register_buffer("position_ids", position_ids)
self.register_buffer("cu_seqlens", cu_seqlens)
def forward(self, pixel_values: torch.Tensor) -> tuple[torch.Tensor, ...]:
hidden_states = self.visual.patch_embed(pixel_values)
pos_embeds = (
self.visual.pos_embed(self.bilinear_indices)
* self.bilinear_weights[:, :, None]
).sum(0)
hidden_states = hidden_states + pos_embeds.to(hidden_states.dtype)
rotary_pos_emb = self.visual.rotary_pos_emb(self.position_ids)
seq_len, _ = hidden_states.size()
hidden_states = hidden_states.reshape(seq_len, -1)
rotary_pos_emb = rotary_pos_emb.reshape(seq_len, -1)
embedding = torch.cat((rotary_pos_emb, rotary_pos_emb), dim=-1)
position_embeddings = (embedding.cos(), embedding.sin())
deepstack_features = []
for layer_num, block in enumerate(self.visual.blocks):
hidden_states = block(
hidden_states,
cu_seqlens=self.cu_seqlens,
position_embeddings=position_embeddings,
)
if layer_num in self.visual.deepstack_visual_indexes:
merger_index = self.visual.deepstack_visual_indexes.index(layer_num)
deepstack_features.append(
self.visual.deepstack_merger_list[merger_index](hidden_states)
)
return (
hidden_states,
self.visual.merger(hidden_states),
*deepstack_features,
)
class CompiledVisionAdapter(torch.nn.Module):
"""Restore the Transformers vision API around a compiled tensor graph."""
def __init__(
self,
compiled: torch.nn.Module,
*,
dtype: torch.dtype,
spatial_merge_size: int,
) -> None:
super().__init__()
self.compiled = compiled
self._output_dtype = dtype
self.spatial_merge_size = spatial_merge_size
@property
def dtype(self) -> torch.dtype:
return self._output_dtype
def forward(
self,
pixel_values: torch.Tensor,
grid_thw: torch.Tensor | None = None,
return_dict: bool = True,
**_: Any,
) -> BaseModelOutputWithDeepstackFeatures | tuple[torch.Tensor, ...]:
del grid_thw
outputs = self.compiled(pixel_values)
if not return_dict:
return outputs
return BaseModelOutputWithDeepstackFeatures(
last_hidden_state=outputs[0],
pooler_output=outputs[1],
deepstack_features=list(outputs[2:]),
)
def timed_samples(
module: torch.nn.Module,
pixel_values: torch.Tensor,
*,
warmups: int,
runs: int,
) -> tuple[tuple[torch.Tensor, ...], list[float]]:
output: tuple[torch.Tensor, ...] | None = None
samples: list[float] = []
with torch.inference_mode():
for index in range(warmups + runs):
torch.cuda.synchronize()
started = time.perf_counter()
output = module(pixel_values)
torch.cuda.synchronize()
elapsed_ms = (time.perf_counter() - started) * 1000
if index >= warmups:
samples.append(elapsed_ms)
assert output is not None
return output, samples
def tensor_errors(
eager: tuple[torch.Tensor, ...], compiled: tuple[torch.Tensor, ...]
) -> list[dict[str, Any]]:
errors = []
for index, (reference, candidate) in enumerate(zip(eager, compiled, strict=True)):
difference = (reference.float() - candidate.float()).abs()
errors.append(
{
"output_index": index,
"shape": list(reference.shape),
"max_absolute_error": float(difference.max()),
"mean_absolute_error": float(difference.mean()),
"cosine_similarity": float(
torch.nn.functional.cosine_similarity(
reference.float().flatten(),
candidate.float().flatten(),
dim=0,
)
),
}
)
return errors
def graph_coverage(module: torch.nn.Module) -> dict[str, Any]:
graph = getattr(module, "graph", None)
if graph is None:
return {"available": False}
nodes = list(graph.nodes)
call_modules = [node for node in nodes if node.op == "call_module"]
targets = [str(node.target) for node in call_modules]
engine_targets = [target for target in targets if "run_on_acc" in target]
fallback_targets = [target for target in targets if "run_on_gpu" in target]
return {
"available": True,
"graph_nodes": len(nodes),
"call_modules": targets,
"tensorrt_engine_partitions": len(engine_targets),
"pytorch_fallback_partitions": len(fallback_targets),
}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--image", type=Path, required=True)
parser.add_argument("--model", default="Qwen/Qwen3-VL-2B-Instruct")
parser.add_argument(
"--prompt", default="Describe this image precisely in one sentence."
)
parser.add_argument("--longest-edge", type=int, default=448)
parser.add_argument("--warmups", type=int, default=5)
parser.add_argument("--runs", type=int, default=20)
parser.add_argument("--generation-warmups", type=int, default=1)
parser.add_argument("--generation-runs", type=int, default=3)
parser.add_argument("--max-new-tokens", type=int, default=96)
parser.add_argument("--min-block-size", type=int, default=5)
parser.add_argument("--optimization-level", type=int, default=3)
parser.add_argument("--require-full-compilation", action="store_true")
parser.add_argument(
"--save-engine",
type=Path,
help="Serialize the compiled vision graph as a portable ExportedProgram.",
)
parser.add_argument(
"--output",
type=Path,
default=Path("benchmarks/results/qwen3-vl-2b-tensorrt-vision.json"),
)
args = parser.parse_args()
if not torch.cuda.is_available():
raise RuntimeError("CUDA is required")
image = resize_to_longest_edge(
Image.open(args.image).convert("RGB"), args.longest_edge
)
load_started = time.perf_counter()
processor = AutoProcessor.from_pretrained(args.model)
model = AutoModelForImageTextToText.from_pretrained(
args.model,
dtype=torch.bfloat16,
attn_implementation="sdpa",
device_map="cuda",
).eval()
torch.cuda.synchronize()
load_seconds = time.perf_counter() - load_started
inputs = prepare_inputs(
processor,
image,
args.prompt,
min_pixels=None,
max_pixels=None,
)
pixel_values = inputs["pixel_values"].to("cuda", dtype=torch.bfloat16)
grid_thw = inputs["image_grid_thw"].to("cuda")
visual = StaticVisionTensorOutputs(model.model.visual, grid_thw).eval()
eager_output, eager_ms = timed_samples(
visual,
pixel_values,
warmups=args.warmups,
runs=args.runs,
)
eager_generation = []
for index in range(args.generation_warmups + args.generation_runs):
sample = run_sample(
model,
processor,
image,
args.prompt,
max_new_tokens=args.max_new_tokens,
cache_implementation="static",
min_pixels=None,
max_pixels=None,
disable_compile=False,
)
if index >= args.generation_warmups:
eager_generation.append(sample)
compile_started = time.perf_counter()
compiled = torch_tensorrt.compile(
visual,
ir="dynamo",
arg_inputs=(pixel_values,),
enabled_precisions={torch.bfloat16},
min_block_size=args.min_block_size,
optimization_level=args.optimization_level,
require_full_compilation=args.require_full_compilation,
pass_through_build_failures=True,
enable_experimental_decompositions=True,
cache_built_engines=True,
reuse_cached_engines=True,
engine_cache_dir="benchmarks/results/tensorrt-engine-cache",
)
torch.cuda.synchronize()
compile_seconds = time.perf_counter() - compile_started
compiled_output, compiled_ms = timed_samples(
compiled,
pixel_values,
warmups=args.warmups,
runs=args.runs,
)
if args.save_engine is not None:
args.save_engine.parent.mkdir(parents=True, exist_ok=True)
torch_tensorrt.save(
compiled,
str(args.save_engine),
output_format="exported_program",
pickle_protocol=4,
)
original_visual = model.model.visual
model.model.visual = CompiledVisionAdapter(
compiled,
dtype=original_visual.dtype,
spatial_merge_size=original_visual.spatial_merge_size,
)
tensorrt_generation = []
for index in range(args.generation_warmups + args.generation_runs):
sample = run_sample(
model,
processor,
image,
args.prompt,
max_new_tokens=args.max_new_tokens,
cache_implementation="static",
min_pixels=None,
max_pixels=None,
disable_compile=False,
)
if index >= args.generation_warmups:
tensorrt_generation.append(sample)
eager_median = statistics.median(eager_ms)
compiled_median = statistics.median(compiled_ms)
artifact = {
"schema": "comfyui-vlm/tensorrt-vision-probe",
"version": 1,
"created_at": datetime.now(UTC).isoformat(),
"model": args.model,
"media": {
"path": str(args.image.resolve()),
"processed_size": list(image.size),
"pixel_values_shape": list(pixel_values.shape),
"image_grid_thw": grid_thw.cpu().tolist(),
},
"environment": {
"platform": platform.platform(),
"python": platform.python_version(),
"torch": torch.__version__,
"cuda": torch.version.cuda,
"torch_tensorrt": torch_tensorrt.__version__,
"tensorrt": __import__("tensorrt").__version__,
"transformers": __import__("transformers").__version__,
"gpu": torch.cuda.get_device_name(),
},
"configuration": {
"precision": "bfloat16",
"min_block_size": args.min_block_size,
"optimization_level": args.optimization_level,
"require_full_compilation": args.require_full_compilation,
"warmups": args.warmups,
"runs": args.runs,
"generation_warmups": args.generation_warmups,
"generation_runs": args.generation_runs,
},
"model_load_seconds": round(load_seconds, 3),
"compile_seconds": round(compile_seconds, 3),
"coverage": graph_coverage(compiled),
"fidelity": tensor_errors(eager_output, compiled_output),
"latency_ms": {
"eager_samples": [round(value, 3) for value in eager_ms],
"tensorrt_samples": [round(value, 3) for value in compiled_ms],
"eager_median": round(eager_median, 3),
"tensorrt_median": round(compiled_median, 3),
"speedup": round(eager_median / compiled_median, 3),
},
"generation": {
"eager": aggregate(eager_generation),
"tensorrt": aggregate(tensorrt_generation),
"exact_output_match": all(
sample["output_sha256"] == eager_generation[0]["output_sha256"]
for sample in tensorrt_generation
),
"eager_output": eager_generation[0]["output"],
"tensorrt_output": tensorrt_generation[0]["output"],
"eager_samples": eager_generation,
"tensorrt_samples": tensorrt_generation,
},
}
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(json.dumps(artifact, indent=2) + "\n", encoding="utf-8")
print(json.dumps(artifact, indent=2), flush=True)
print(args.output, flush=True)
if __name__ == "__main__":
main()
+351
View File
@@ -0,0 +1,351 @@
"""Direct Qwen3-VL Transformers benchmark with quality-preserving artifacts.
This runner measures the same local model path used by Modern VLM without
requiring a running ComfyUI server. It records preprocessing, user-visible
time-to-first-text, end-to-end latency, decode throughput, peak VRAM, and the
complete output for exact cross-iteration comparisons.
"""
from __future__ import annotations
import argparse
import hashlib
import json
import math
import os
import platform
import statistics
import subprocess
import threading
import time
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
import torch
from PIL import Image
from transformers import (
AutoModelForImageTextToText,
AutoProcessor,
TextIteratorStreamer,
)
def percentile(values: list[float], quantile: float) -> float:
ordered = sorted(values)
position = (len(ordered) - 1) * quantile
lower = math.floor(position)
upper = math.ceil(position)
if lower == upper:
return ordered[lower]
return ordered[lower] * (upper - position) + ordered[upper] * (position - lower)
def git_value(*args: str) -> str | None:
try:
return subprocess.check_output(
["git", *args], text=True, stderr=subprocess.DEVNULL
).strip()
except (OSError, subprocess.CalledProcessError):
return None
def sha256_file(path: Path) -> str:
digest = hashlib.sha256()
with path.open("rb") as handle:
for chunk in iter(lambda: handle.read(1024 * 1024), b""):
digest.update(chunk)
return digest.hexdigest()
def resize_to_longest_edge(image: Image.Image, longest_edge: int | None) -> Image.Image:
if longest_edge is None or max(image.size) <= longest_edge:
return image
scale = longest_edge / max(image.size)
size = (
max(1, round(image.width * scale)),
max(1, round(image.height * scale)),
)
return image.resize(size, Image.Resampling.BOX)
def prepare_inputs(
processor: Any,
image: Image.Image,
prompt: str,
*,
min_pixels: int | None,
max_pixels: int | None,
) -> dict[str, torch.Tensor]:
image_part: dict[str, Any] = {"type": "image", "image": image}
if min_pixels is not None:
image_part["min_pixels"] = min_pixels
if max_pixels is not None:
image_part["max_pixels"] = max_pixels
messages = [
{
"role": "user",
"content": [image_part, {"type": "text", "text": prompt}],
}
]
return processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
)
def run_sample(
model: Any,
processor: Any,
image: Image.Image,
prompt: str,
*,
max_new_tokens: int,
cache_implementation: str,
min_pixels: int | None,
max_pixels: int | None,
disable_compile: bool,
) -> dict[str, Any]:
torch.cuda.reset_peak_memory_stats()
torch.cuda.synchronize()
started = time.perf_counter()
inputs = prepare_inputs(
processor,
image,
prompt,
min_pixels=min_pixels,
max_pixels=max_pixels,
)
prepared_at = time.perf_counter()
inputs = {name: value.to(model.device) for name, value in inputs.items()}
input_length = int(inputs["input_ids"].shape[-1])
streamer = TextIteratorStreamer(
processor.tokenizer,
skip_prompt=True,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)
generated: list[torch.Tensor] = []
errors: list[BaseException] = []
def generate() -> None:
try:
with torch.inference_mode():
generated.append(
model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False,
cache_implementation=cache_implementation,
disable_compile=disable_compile,
streamer=streamer,
)
)
except BaseException as exc:
errors.append(exc)
streamer.end()
first_text_at: float | None = None
chunks: list[str] = []
worker = threading.Thread(target=generate, daemon=True)
worker.start()
for chunk in streamer:
if chunk and first_text_at is None:
first_text_at = time.perf_counter()
chunks.append(chunk)
worker.join()
if errors:
raise errors[0]
torch.cuda.synchronize()
finished = time.perf_counter()
output_ids = generated[0][:, input_length:]
output_tokens = int(output_ids.shape[-1])
output = processor.batch_decode(
output_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0].strip()
ttft_seconds = (first_text_at or finished) - started
decode_seconds = max(0.0, finished - (first_text_at or finished))
return {
"preprocess_ms": round((prepared_at - started) * 1000, 3),
"ttft_ms": round(ttft_seconds * 1000, 3),
"e2e_ms": round((finished - started) * 1000, 3),
"output_tokens": output_tokens,
"output_tokens_per_second": (
round(max(0, output_tokens - 1) / decode_seconds, 3)
if output_tokens > 1 and decode_seconds > 0
else None
),
"peak_vram_gib": round(torch.cuda.max_memory_allocated() / 1024**3, 3),
"input_tokens": input_length,
"vision_tokens": int(inputs.get("pixel_values", torch.empty(0)).shape[0]),
"output": output,
"output_sha256": hashlib.sha256(output.encode("utf-8")).hexdigest(),
}
def aggregate(samples: list[dict[str, Any]]) -> dict[str, Any]:
def metric(name: str) -> list[float]:
return [float(sample[name]) for sample in samples]
rates = [
float(sample["output_tokens_per_second"])
for sample in samples
if sample["output_tokens_per_second"] is not None
]
return {
"preprocess_ms_mean": round(statistics.fmean(metric("preprocess_ms")), 3),
"ttft_ms": {
"p50": round(percentile(metric("ttft_ms"), 0.50), 3),
"p95": round(percentile(metric("ttft_ms"), 0.95), 3),
},
"e2e_ms": {
"p50": round(percentile(metric("e2e_ms"), 0.50), 3),
"p95": round(percentile(metric("e2e_ms"), 0.95), 3),
},
"output_tokens_per_second_mean": round(statistics.fmean(rates), 3),
"peak_vram_gib": round(max(metric("peak_vram_gib")), 3),
"output_tokens_mean": round(statistics.fmean(metric("output_tokens")), 3),
"outputs_identical": len({sample["output_sha256"] for sample in samples}) == 1,
}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--image", type=Path, required=True)
parser.add_argument("--prompt", default="Describe this image precisely in one sentence.")
parser.add_argument("--model", default="Qwen/Qwen3-VL-2B-Instruct")
parser.add_argument("--label", required=True)
parser.add_argument(
"--attention",
choices=("sdpa", "flash_attention_2", "eager"),
default="sdpa",
)
parser.add_argument("--cache", choices=("dynamic", "static"), default="dynamic")
parser.add_argument("--disable-compile", action="store_true")
parser.add_argument("--min-pixels", type=int)
parser.add_argument("--max-pixels", type=int)
parser.add_argument("--longest-edge", type=int)
parser.add_argument("--max-new-tokens", type=int, default=96)
parser.add_argument("--warmups", type=int, default=2)
parser.add_argument("--runs", type=int, default=5)
parser.add_argument("--expected-output-sha256")
parser.add_argument(
"--float32-matmul-precision",
choices=("highest", "high", "medium"),
default="highest",
)
parser.add_argument("--output-dir", type=Path, default=Path("benchmarks/results"))
args = parser.parse_args()
if not torch.cuda.is_available():
raise RuntimeError("This benchmark requires a CUDA GPU.")
if args.runs < 1 or args.warmups < 0:
parser.error("--runs must be positive and --warmups non-negative")
torch.set_float32_matmul_precision(args.float32_matmul_precision)
image_path = args.image.resolve()
source_image = Image.open(image_path).convert("RGB")
image = resize_to_longest_edge(source_image, args.longest_edge)
load_started = time.perf_counter()
processor = AutoProcessor.from_pretrained(args.model)
model = AutoModelForImageTextToText.from_pretrained(
args.model,
dtype=torch.bfloat16,
attn_implementation=args.attention,
device_map="cuda",
).eval()
torch.cuda.synchronize()
load_seconds = time.perf_counter() - load_started
samples = []
for index in range(args.warmups + args.runs):
sample = run_sample(
model,
processor,
image,
args.prompt,
max_new_tokens=args.max_new_tokens,
cache_implementation=args.cache,
min_pixels=args.min_pixels,
max_pixels=args.max_pixels,
disable_compile=args.disable_compile,
)
print(
f"{index + 1}/{args.warmups + args.runs} "
f"ttft={sample['ttft_ms']:.1f}ms "
f"e2e={sample['e2e_ms']:.1f}ms "
f"tok/s={sample['output_tokens_per_second']}"
)
if index >= args.warmups:
samples.append(sample)
artifact = {
"schema": "comfyui-vlm/transformers-benchmark",
"version": 1,
"created_at": datetime.now(UTC).isoformat(),
"label": args.label,
"model": args.model,
"git_commit": git_value("rev-parse", "HEAD"),
"git_dirty": bool(git_value("status", "--porcelain")),
"media": {
"path": os.fspath(image_path),
"sha256": sha256_file(image_path),
"source_width": source_image.width,
"source_height": source_image.height,
"processed_width": image.width,
"processed_height": image.height,
},
"environment": {
"platform": platform.platform(),
"python": platform.python_version(),
"torch": torch.__version__,
"cuda": torch.version.cuda,
"gpu": torch.cuda.get_device_name(),
"transformers": __import__("transformers").__version__,
"flash_attn": (
__import__("flash_attn").__version__
if args.attention == "flash_attention_2"
else None
),
},
"settings": {
"attention": args.attention,
"cache": args.cache,
"disable_compile": args.disable_compile,
"min_pixels": args.min_pixels,
"max_pixels": args.max_pixels,
"longest_edge": args.longest_edge,
"max_new_tokens": args.max_new_tokens,
"warmups": args.warmups,
"runs": args.runs,
"float32_matmul_precision": args.float32_matmul_precision,
},
"model_load_seconds": round(load_seconds, 3),
"quality_gate": {
"method": "byte-identical output SHA-256",
"reference_sha256": args.expected_output_sha256,
"passed": (
all(
sample["output_sha256"] == args.expected_output_sha256
for sample in samples
)
if args.expected_output_sha256
else None
),
},
"summary": aggregate(samples),
"samples": samples,
}
args.output_dir.mkdir(parents=True, exist_ok=True)
output_path = args.output_dir / f"{args.label}.json"
output_path.write_text(json.dumps(artifact, indent=2) + "\n", encoding="utf-8")
print(json.dumps(artifact["summary"], indent=2))
print(output_path)
if __name__ == "__main__":
main()
+4
View File
@@ -0,0 +1,4 @@
*.log
qwen3-vl-2b-sdpa-*.json
!qwen3-vl-2b-sdpa-static-edge448-tf32.json
qwen3-vl-2b-matrix-tf5.json
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-07T23:33:29.188812+00:00",
"label": "qwen3-vl-2b-sglang-flashinfer-448",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "d9584e1e35c4373ef99ac07d7c4a866852363c99",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 43.271,
"p95": 47.958,
"p99": 50.233
},
"ttft_ms": {
"p50": 38.276,
"p95": 42.913,
"p99": 45.215
},
"output_tokens_per_second_mean": 393.882,
"quality_mean": 0.0
},
"quality_gate": {
"threshold": 1.0,
"passed": false
},
"samples": [
{
"output": "```",
"latency_ms": 43.967,
"ttft_ms": 38.911,
"completion_tokens": 2,
"output_tokens_per_second": 395.577,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 44.482,
"ttft_ms": 39.396,
"completion_tokens": 2,
"output_tokens_per_second": 393.213,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 44.431,
"ttft_ms": 39.33,
"completion_tokens": 2,
"output_tokens_per_second": 392.065,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 50.802,
"ttft_ms": 45.791,
"completion_tokens": 2,
"output_tokens_per_second": 399.098,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 41.64,
"ttft_ms": 36.399,
"completion_tokens": 2,
"output_tokens_per_second": 381.592,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 42.047,
"ttft_ms": 37.063,
"completion_tokens": 2,
"output_tokens_per_second": 401.332,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 42.536,
"ttft_ms": 37.641,
"completion_tokens": 2,
"output_tokens_per_second": 408.589,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 43.94,
"ttft_ms": 38.958,
"completion_tokens": 2,
"output_tokens_per_second": 401.461,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 42.602,
"ttft_ms": 37.429,
"completion_tokens": 2,
"output_tokens_per_second": 386.593,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
},
{
"output": "```",
"latency_ms": 42.116,
"ttft_ms": 36.844,
"completion_tokens": 2,
"output_tokens_per_second": 379.305,
"usage": {
"prompt_tokens": 144,
"total_tokens": 146,
"completion_tokens": 2,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 0.0
}
]
}
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-07T23:51:14.661282+00:00",
"label": "qwen3-vl-2b-sglang-0510-transformers-flashinfer-edge448",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "d9584e1e35c4373ef99ac07d7c4a866852363c99",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 254.247,
"p95": 257.515,
"p99": 258.626
},
"ttft_ms": {
"p50": 75.627,
"p95": 79.154,
"p99": 80.361
},
"output_tokens_per_second_mean": 173.546,
"quality_mean": 1.0
},
"quality_gate": {
"threshold": 1.0,
"passed": true
},
"samples": [
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 254.057,
"ttft_ms": 74.753,
"completion_tokens": 31,
"output_tokens_per_second": 172.891,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 258.904,
"ttft_ms": 80.663,
"completion_tokens": 31,
"output_tokens_per_second": 173.922,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 254.891,
"ttft_ms": 76.593,
"completion_tokens": 31,
"output_tokens_per_second": 173.866,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 252.898,
"ttft_ms": 74.213,
"completion_tokens": 31,
"output_tokens_per_second": 173.489,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 253.15,
"ttft_ms": 74.329,
"completion_tokens": 31,
"output_tokens_per_second": 173.358,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 254.437,
"ttft_ms": 75.904,
"completion_tokens": 31,
"output_tokens_per_second": 173.638,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 253.923,
"ttft_ms": 75.351,
"completion_tokens": 31,
"output_tokens_per_second": 173.599,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 255.317,
"ttft_ms": 76.96,
"completion_tokens": 31,
"output_tokens_per_second": 173.808,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 253.695,
"ttft_ms": 74.735,
"completion_tokens": 31,
"output_tokens_per_second": 173.222,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 255.817,
"ttft_ms": 77.31,
"completion_tokens": 31,
"output_tokens_per_second": 173.662,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
}
]
}
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-07T23:54:42.051771+00:00",
"label": "qwen3-vl-2b-sglang-0510-native-flashinfer-edge448-concepts",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "d9584e1e35c4373ef99ac07d7c4a866852363c99",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 240.643,
"p95": 243.55,
"p99": 243.925
},
"ttft_ms": {
"p50": 35.228,
"p95": 38.264,
"p99": 38.455
},
"output_tokens_per_second_mean": 194.675,
"quality_mean": 1.0
},
"quality_gate": {
"threshold": 1.0,
"passed": true
},
"samples": [
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.312,
"ttft_ms": 34.498,
"completion_tokens": 40,
"output_tokens_per_second": 194.35,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.597,
"ttft_ms": 35.33,
"completion_tokens": 40,
"output_tokens_per_second": 194.868,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.689,
"ttft_ms": 35.043,
"completion_tokens": 40,
"output_tokens_per_second": 194.509,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 241.979,
"ttft_ms": 36.655,
"completion_tokens": 40,
"output_tokens_per_second": 194.815,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 241.043,
"ttft_ms": 35.436,
"completion_tokens": 40,
"output_tokens_per_second": 194.546,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.182,
"ttft_ms": 34.754,
"completion_tokens": 40,
"output_tokens_per_second": 194.716,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.477,
"ttft_ms": 35.127,
"completion_tokens": 40,
"output_tokens_per_second": 194.789,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 240.458,
"ttft_ms": 34.706,
"completion_tokens": 40,
"output_tokens_per_second": 194.409,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 244.019,
"ttft_ms": 38.503,
"completion_tokens": 40,
"output_tokens_per_second": 194.631,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 242.976,
"ttft_ms": 37.973,
"completion_tokens": 40,
"output_tokens_per_second": 195.119,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
}
]
}
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-07T23:56:26.751727+00:00",
"label": "qwen3-vl-2b-sglang-0510-native-triton-mm-edge448",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "d9584e1e35c4373ef99ac07d7c4a866852363c99",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 193.603,
"p95": 195.324,
"p99": 195.594
},
"ttft_ms": {
"p50": 35.484,
"p95": 37.948,
"p99": 38.154
},
"output_tokens_per_second_mean": 196.424,
"quality_mean": 1.0
},
"quality_gate": {
"threshold": 1.0,
"passed": true
},
"samples": [
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 194.306,
"ttft_ms": 37.462,
"completion_tokens": 31,
"output_tokens_per_second": 197.649,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.679,
"ttft_ms": 36.034,
"completion_tokens": 31,
"output_tokens_per_second": 196.645,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.82,
"ttft_ms": 35.359,
"completion_tokens": 31,
"output_tokens_per_second": 195.632,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 195.662,
"ttft_ms": 38.205,
"completion_tokens": 31,
"output_tokens_per_second": 196.879,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.234,
"ttft_ms": 35.429,
"completion_tokens": 31,
"output_tokens_per_second": 196.444,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.526,
"ttft_ms": 35.539,
"completion_tokens": 31,
"output_tokens_per_second": 196.219,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 192.795,
"ttft_ms": 34.703,
"completion_tokens": 31,
"output_tokens_per_second": 196.088,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 192.793,
"ttft_ms": 34.319,
"completion_tokens": 31,
"output_tokens_per_second": 195.616,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 194.911,
"ttft_ms": 37.633,
"completion_tokens": 31,
"output_tokens_per_second": 197.103,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.316,
"ttft_ms": 35.129,
"completion_tokens": 31,
"output_tokens_per_second": 195.97,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
}
]
}
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-07T23:59:00.162586+00:00",
"label": "qwen3-vl-2b-sglang-0510-native-triton-mm-compile-edge448",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "d9584e1e35c4373ef99ac07d7c4a866852363c99",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 190.452,
"p95": 194.69,
"p99": 195.155
},
"ttft_ms": {
"p50": 37.46,
"p95": 40.955,
"p99": 41.035
},
"output_tokens_per_second_mean": 202.626,
"quality_mean": 1.0
},
"quality_gate": {
"threshold": 1.0,
"passed": true
},
"samples": [
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 189.65,
"ttft_ms": 37.316,
"completion_tokens": 31,
"output_tokens_per_second": 203.5,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 191.164,
"ttft_ms": 37.604,
"completion_tokens": 31,
"output_tokens_per_second": 201.876,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 193.981,
"ttft_ms": 41.055,
"completion_tokens": 31,
"output_tokens_per_second": 202.712,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 195.271,
"ttft_ms": 40.832,
"completion_tokens": 31,
"output_tokens_per_second": 200.727,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 188.249,
"ttft_ms": 35.967,
"completion_tokens": 31,
"output_tokens_per_second": 203.569,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 191.006,
"ttft_ms": 36.928,
"completion_tokens": 31,
"output_tokens_per_second": 201.197,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 189.897,
"ttft_ms": 38.167,
"completion_tokens": 31,
"output_tokens_per_second": 204.311,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 191.132,
"ttft_ms": 37.715,
"completion_tokens": 31,
"output_tokens_per_second": 202.064,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 188.008,
"ttft_ms": 35.408,
"completion_tokens": 31,
"output_tokens_per_second": 203.146,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"latency_ms": 188.646,
"ttft_ms": 36.057,
"completion_tokens": 31,
"output_tokens_per_second": 203.161,
"usage": {
"prompt_tokens": 144,
"total_tokens": 175,
"completion_tokens": 31,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
}
]
}
@@ -0,0 +1,304 @@
{
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": "2026-08-08T07:29:10.320582+00:00",
"label": "qwen3-vl-2b-sglang-tensorrt-bridge-edge448",
"backend": "sglang",
"suite": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "7bfc87bc04789febd24142ed27facde5e5fd54bf",
"git_dirty": true,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"server_base_url": "http://127.0.0.1:30000/v1"
},
"settings": {
"warmups": 3,
"runs": 10,
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0
},
"summary": {
"requests": 10,
"latency_ms": {
"p50": 250.673,
"p95": 366.093,
"p99": 439.235
},
"ttft_ms": {
"p50": 34.85,
"p95": 37.752,
"p99": 38.232
},
"output_tokens_per_second_mean": 176.12,
"quality_mean": 1.0
},
"quality_gate": {
"threshold": 1.0,
"passed": true
},
"samples": [
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 254.239,
"ttft_ms": 38.352,
"completion_tokens": 40,
"output_tokens_per_second": 185.282,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 0,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 254.348,
"ttft_ms": 35.315,
"completion_tokens": 40,
"output_tokens_per_second": 182.621,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 1,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 253.791,
"ttft_ms": 37.018,
"completion_tokens": 40,
"output_tokens_per_second": 184.525,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 2,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 249.899,
"ttft_ms": 34.987,
"completion_tokens": 40,
"output_tokens_per_second": 186.122,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 3,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 249.934,
"ttft_ms": 34.695,
"completion_tokens": 40,
"output_tokens_per_second": 185.84,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 4,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 249.348,
"ttft_ms": 34.744,
"completion_tokens": 40,
"output_tokens_per_second": 186.39,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 5,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 250.471,
"ttft_ms": 34.284,
"completion_tokens": 40,
"output_tokens_per_second": 185.025,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 6,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 250.24,
"ttft_ms": 34.789,
"completion_tokens": 40,
"output_tokens_per_second": 185.657,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 7,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 457.52,
"ttft_ms": 34.346,
"completion_tokens": 40,
"output_tokens_per_second": 94.524,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 8,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
},
{
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, playfully high-fiving each other as the golden hour light bathes the scene in warm, soft light.",
"latency_ms": 250.874,
"ttft_ms": 34.911,
"completion_tokens": 40,
"output_tokens_per_second": 185.218,
"usage": {
"prompt_tokens": 144,
"total_tokens": 184,
"completion_tokens": 40,
"prompt_tokens_details": null,
"reasoning_tokens": 0
},
"sample": 9,
"case_id": "caption-qwen-demo-001",
"task": "caption",
"media_sha256": "188eb59f12f8da458d6cb77ce19fb471519a94cd43f6701364593cc5843dce23",
"media": {
"source_sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"quality": 1.0
}
]
}
+10
View File
@@ -0,0 +1,10 @@
# Result artifacts
Committed JSON files are immutable raw benchmark evidence. Each artifact
contains environment identity, input dimensions and token counts, warmups,
every measured sample, full model output, quality-gate details, percentiles,
VRAM, and speedups.
Console logs and scratch experiments are ignored. Promote a result by rerunning
the benchmark with its final runner and committing the resulting JSON rather
than editing an artifact by hand.
@@ -0,0 +1,120 @@
{
"schema": "comfyui-vlm/transformers-benchmark",
"version": 1,
"created_at": "2026-08-07T23:01:55.323925+00:00",
"label": "qwen3-vl-2b-fa2-dynamic-edge448",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "858a4a4998e68bd99b26a12831e861be49904564",
"git_dirty": true,
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.8.0+cu128",
"cuda": "12.8",
"gpu": "NVIDIA GeForce RTX 3090",
"transformers": "5.12.1",
"flash_attn": "2.8.3"
},
"settings": {
"attention": "flash_attention_2",
"cache": "dynamic",
"disable_compile": false,
"min_pixels": null,
"max_pixels": null,
"longest_edge": 448,
"max_new_tokens": 96,
"warmups": 2,
"runs": 5
},
"model_load_seconds": 6.845,
"quality_gate": {
"method": "byte-identical output SHA-256",
"reference_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1",
"passed": true
},
"summary": {
"preprocess_ms_mean": 4.445,
"ttft_ms": {
"p50": 106.553,
"p95": 114.012
},
"e2e_ms": {
"p50": 1033.269,
"p95": 1054.373
},
"output_tokens_per_second_mean": 32.297,
"peak_vram_gib": 4.002,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"samples": [
{
"preprocess_ms": 4.993,
"ttft_ms": 106.308,
"e2e_ms": 1029.213,
"output_tokens": 31,
"output_tokens_per_second": 32.506,
"peak_vram_gib": 4.002,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.881,
"ttft_ms": 114.994,
"e2e_ms": 1059.265,
"output_tokens": 31,
"output_tokens_per_second": 31.771,
"peak_vram_gib": 4.002,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.982,
"ttft_ms": 110.083,
"e2e_ms": 1034.803,
"output_tokens": 31,
"output_tokens_per_second": 32.442,
"peak_vram_gib": 4.002,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.823,
"ttft_ms": 106.553,
"e2e_ms": 1033.269,
"output_tokens": 31,
"output_tokens_per_second": 32.372,
"peak_vram_gib": 4.002,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.547,
"ttft_ms": 103.872,
"e2e_ms": 1030.035,
"output_tokens": 31,
"output_tokens_per_second": 32.392,
"peak_vram_gib": 4.002,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
}
@@ -0,0 +1,180 @@
{
"schema": "comfyui-vlm/transformers-benchmark",
"version": 1,
"created_at": "2026-08-07T23:01:20.543267+00:00",
"label": "qwen3-vl-2b-fa2-static-edge448",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "858a4a4998e68bd99b26a12831e861be49904564",
"git_dirty": true,
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.8.0+cu128",
"cuda": "12.8",
"gpu": "NVIDIA GeForce RTX 3090",
"transformers": "5.12.1",
"flash_attn": "2.8.3"
},
"settings": {
"attention": "flash_attention_2",
"cache": "static",
"disable_compile": false,
"min_pixels": null,
"max_pixels": null,
"longest_edge": 448,
"max_new_tokens": 96,
"warmups": 6,
"runs": 10
},
"model_load_seconds": 84.208,
"quality_gate": {
"method": "byte-identical output SHA-256",
"reference_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1",
"passed": false
},
"summary": {
"preprocess_ms_mean": 4.404,
"ttft_ms": {
"p50": 265.004,
"p95": 273.819
},
"e2e_ms": {
"p50": 3027.985,
"p95": 3042.934
},
"output_tokens_per_second_mean": 34.427,
"peak_vram_gib": 4.023,
"output_tokens_mean": 96.0,
"outputs_identical": true
},
"samples": [
{
"preprocess_ms": 5.119,
"ttft_ms": 262.459,
"e2e_ms": 3016.629,
"output_tokens": 96,
"output_tokens_per_second": 34.493,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 4.519,
"ttft_ms": 265.289,
"e2e_ms": 3014.475,
"output_tokens": 96,
"output_tokens_per_second": 34.556,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 4.943,
"ttft_ms": 273.965,
"e2e_ms": 3044.849,
"output_tokens": 96,
"output_tokens_per_second": 34.285,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 3.913,
"ttft_ms": 262.719,
"e2e_ms": 3039.894,
"output_tokens": 96,
"output_tokens_per_second": 34.207,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 4.66,
"ttft_ms": 261.946,
"e2e_ms": 3040.593,
"output_tokens": 96,
"output_tokens_per_second": 34.189,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 3.929,
"ttft_ms": 273.64,
"e2e_ms": 2997.449,
"output_tokens": 96,
"output_tokens_per_second": 34.878,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 3.926,
"ttft_ms": 264.987,
"e2e_ms": 3028.232,
"output_tokens": 96,
"output_tokens_per_second": 34.38,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 4.683,
"ttft_ms": 267.937,
"e2e_ms": 3027.738,
"output_tokens": 96,
"output_tokens_per_second": 34.423,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 4.571,
"ttft_ms": 262.483,
"e2e_ms": 3016.256,
"output_tokens": 96,
"output_tokens_per_second": 34.498,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
},
{
"preprocess_ms": 3.78,
"ttft_ms": 265.021,
"e2e_ms": 3029.975,
"output_tokens": 96,
"output_tokens_per_second": 34.359,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "s:V:col, 1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1:1",
"output_sha256": "ab0d23104f87597bcea8948a30df54a967214df93df48121fb73729585a29026"
}
]
}
@@ -0,0 +1,100 @@
{
"schema_version": 1,
"checkpoint": "Qwen/Qwen3-VL-2B-Instruct",
"safetensors_bytes": 4255140312,
"environment": {
"gpu": "NVIDIA GeForce RTX 3090 24GB",
"filesystem": "WSL2 ext4",
"flashpack_revision": "a923a6c",
"cache_eviction": "POSIX_FADV_DONTNEED on the measured file only"
},
"flashpack_reader": {
"direct_io": true,
"threads": 4,
"buffers_per_thread": 2,
"chunk_bytes": 33554432,
"pinned_staging_bytes": 268435456,
"cache_pinned": false
},
"preparation": {
"conversion_seconds": 0.0,
"flashpack_bytes": 4255133230,
"tensor_count": 625,
"sample_exact": {
"model.language_model.embed_tokens.weight": true,
"model.visual.blocks.17.mlp.linear_fc1.bias": true,
"model.language_model.layers.9.self_attn.q_norm.weight": true
}
},
"methods": {
"safetensors": {
"cold": {
"seconds": [
58.55027909799992,
59.30713190999995,
60.17983154800004
],
"p50_seconds": 59.30713190999995,
"p95_seconds": 60.17983154800004,
"p50_throughput_gbps": 0.5739802516105188,
"peak_gpu_bytes": 4256651264
},
"warm": {
"seconds": [
1.033367304999956,
1.0584651120000217,
1.1987151169998924
],
"p50_seconds": 1.0584651120000217,
"p95_seconds": 1.1987151169998924,
"p50_throughput_gbps": 32.16083563838739,
"peak_gpu_bytes": 4256651264
}
},
"safetensors_fast_gpu": {
"cold": {
"seconds": [
62.305181905999916,
58.50693893500011,
60.62746760899995
],
"p50_seconds": 60.62746760899995,
"p95_seconds": 62.305181905999916,
"p50_throughput_gbps": 0.5614801976480164,
"peak_gpu_bytes": 4256651264
},
"warm": {
"seconds": [
1.137218976999975,
1.0898004420000689,
1.1920738910000637
],
"p50_seconds": 1.137218976999975,
"p95_seconds": 1.1920738910000637,
"p50_throughput_gbps": 29.933656740236362,
"peak_gpu_bytes": 4256651264
}
},
"flashpack_bounded_direct_io": {
"cold": {
"seconds": [
44.775754307999705,
38.81784143599998,
43.73773216100017
],
"p50_seconds": 43.73773216100017,
"p95_seconds": 44.775754307999705,
"p50_throughput_gbps": 0.7782997461938267,
"peak_gpu_bytes": 4255121408
},
"warm": null,
"sample_exact": true
}
},
"probes": {
"flashpack_8x16mib_cold_seconds": 45.7384941790001,
"flashpack_buffered_cold_seconds": 87.244453406,
"flashpack_buffered_warm_seconds": 0.997375596,
"upstream_default": "failed: 2 GiB pinned staging allocation exceeded this host's available pinned-memory budget"
}
}
@@ -0,0 +1,917 @@
{
"schema": "comfyui-vlm/optimization-matrix",
"version": 1,
"created_at": "2026-08-07T22:40:42.770088+00:00",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"source_width": 2048,
"source_height": 1365
},
"prompt": "Describe this image precisely in one sentence.",
"model_load_seconds": 6.858,
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.8.0+cu128",
"cuda": "12.8",
"gpu": "NVIDIA GeForce RTX 3090",
"transformers": "5.12.1"
},
"runs_per_variant": 10,
"variants": [
{
"id": "00",
"label": "BF16 SDPA / dynamic cache / source resolution",
"longest_edge": null,
"cache": "dynamic",
"warmups": 2,
"processed_width": 2048,
"processed_height": 1365,
"quality_gate": {
"method": "required visual concepts",
"passed": true,
"concepts": {
"passed": true,
"matched": [
"woman",
"golden retriever",
"beach",
"high-five"
],
"required": [
[
"woman",
"person"
],
[
"golden retriever",
"dog"
],
[
"beach",
"sand"
],
[
"high-five",
"high five"
]
]
},
"exact_output_reference": null,
"exact_output_passed": null,
"exact_output_vs_baseline": true
},
"speedup_vs_baseline": {
"ttft": 1.0,
"e2e": 1.0,
"throughput": 1.0
},
"summary": {
"preprocess_ms_mean": 74.368,
"ttft_ms": {
"p50": 700.296,
"p95": 723.148
},
"e2e_ms": {
"p50": 1395.726,
"p95": 1487.708
},
"output_tokens_per_second_mean": 42.32,
"peak_vram_gib": 4.548,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"warmup_samples": [
{
"preprocess_ms": 119.344,
"ttft_ms": 1411.57,
"e2e_ms": 2117.774,
"output_tokens": 31,
"output_tokens_per_second": 42.481,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 74.884,
"ttft_ms": 705.019,
"e2e_ms": 1410.964,
"output_tokens": 31,
"output_tokens_per_second": 42.496,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
}
],
"samples": [
{
"preprocess_ms": 72.204,
"ttft_ms": 701.128,
"e2e_ms": 1389.643,
"output_tokens": 31,
"output_tokens_per_second": 43.572,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 72.348,
"ttft_ms": 704.423,
"e2e_ms": 1402.979,
"output_tokens": 31,
"output_tokens_per_second": 42.946,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 72.379,
"ttft_ms": 699.464,
"e2e_ms": 1390.29,
"output_tokens": 31,
"output_tokens_per_second": 43.426,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 72.929,
"ttft_ms": 698.669,
"e2e_ms": 1401.161,
"output_tokens": 31,
"output_tokens_per_second": 42.705,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 76.526,
"ttft_ms": 731.448,
"e2e_ms": 1537.253,
"output_tokens": 31,
"output_tokens_per_second": 37.23,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 76.624,
"ttft_ms": 713.003,
"e2e_ms": 1422.268,
"output_tokens": 31,
"output_tokens_per_second": 42.297,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 75.459,
"ttft_ms": 666.823,
"e2e_ms": 1383.653,
"output_tokens": 31,
"output_tokens_per_second": 41.851,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 76.597,
"ttft_ms": 691.338,
"e2e_ms": 1375.888,
"output_tokens": 31,
"output_tokens_per_second": 43.824,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 75.082,
"ttft_ms": 711.574,
"e2e_ms": 1427.153,
"output_tokens": 31,
"output_tokens_per_second": 41.924,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
},
{
"preprocess_ms": 73.533,
"ttft_ms": 673.702,
"e2e_ms": 1364.577,
"output_tokens": 31,
"output_tokens_per_second": 43.423,
"peak_vram_gib": 4.548,
"input_tokens": 2770,
"vision_tokens": 11008,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully giving a high-five to her.",
"output_sha256": "7f7b9f725ce3dd1bd0e815bfcff7bf187fa2abe0d53fa7bcd7bfd612db7085da"
}
]
},
{
"id": "01a",
"label": "BF16 SDPA / dynamic cache / 672px edge",
"longest_edge": 672,
"cache": "dynamic",
"warmups": 2,
"processed_width": 672,
"processed_height": 448,
"quality_gate": {
"method": "required visual concepts",
"passed": true,
"concepts": {
"passed": true,
"matched": [
"woman",
"golden retriever",
"beach",
"high-five"
],
"required": [
[
"woman",
"person"
],
[
"golden retriever",
"dog"
],
[
"beach",
"sand"
],
[
"high-five",
"high five"
]
]
},
"exact_output_reference": null,
"exact_output_passed": null,
"exact_output_vs_baseline": false
},
"speedup_vs_baseline": {
"ttft": 6.21,
"e2e": 1.654,
"throughput": 0.998
},
"summary": {
"preprocess_ms_mean": 7.266,
"ttft_ms": {
"p50": 112.763,
"p95": 122.823
},
"e2e_ms": {
"p50": 843.966,
"p95": 881.66
},
"output_tokens_per_second_mean": 42.225,
"peak_vram_gib": 4.036,
"output_tokens_mean": 32.0,
"outputs_identical": true
},
"warmup_samples": [
{
"preprocess_ms": 6.031,
"ttft_ms": 122.476,
"e2e_ms": 844.441,
"output_tokens": 32,
"output_tokens_per_second": 42.938,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.525,
"ttft_ms": 113.414,
"e2e_ms": 867.517,
"output_tokens": 32,
"output_tokens_per_second": 41.108,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
}
],
"samples": [
{
"preprocess_ms": 8.086,
"ttft_ms": 123.813,
"e2e_ms": 872.379,
"output_tokens": 32,
"output_tokens_per_second": 41.412,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.525,
"ttft_ms": 114.325,
"e2e_ms": 831.252,
"output_tokens": 32,
"output_tokens_per_second": 43.24,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.52,
"ttft_ms": 110.647,
"e2e_ms": 889.253,
"output_tokens": 32,
"output_tokens_per_second": 39.815,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 6.243,
"ttft_ms": 112.581,
"e2e_ms": 835.841,
"output_tokens": 32,
"output_tokens_per_second": 42.862,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 6.796,
"ttft_ms": 113.04,
"e2e_ms": 840.013,
"output_tokens": 32,
"output_tokens_per_second": 42.643,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.75,
"ttft_ms": 110.994,
"e2e_ms": 858.836,
"output_tokens": 32,
"output_tokens_per_second": 41.453,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.667,
"ttft_ms": 112.917,
"e2e_ms": 847.365,
"output_tokens": 32,
"output_tokens_per_second": 42.209,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 6.15,
"ttft_ms": 110.293,
"e2e_ms": 825.91,
"output_tokens": 32,
"output_tokens_per_second": 43.319,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.394,
"ttft_ms": 121.612,
"e2e_ms": 846.696,
"output_tokens": 32,
"output_tokens_per_second": 42.754,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
},
{
"preprocess_ms": 7.532,
"ttft_ms": 112.61,
"e2e_ms": 841.236,
"output_tokens": 32,
"output_tokens_per_second": 42.546,
"peak_vram_gib": 4.036,
"input_tokens": 312,
"vision_tokens": 1176,
"output": "A woman and her golden retriever share a joyful moment on a sandy beach at sunset, with the dog playfully reaching out to give a high-five.",
"output_sha256": "96d246f1d52b0845af007cc02f126776b655eda16d09a07f4ac1bfb4be8404f0"
}
]
},
{
"id": "01b",
"label": "BF16 SDPA / dynamic cache / 448px edge",
"longest_edge": 448,
"cache": "dynamic",
"warmups": 2,
"processed_width": 448,
"processed_height": 299,
"quality_gate": {
"method": "required visual concepts",
"passed": true,
"concepts": {
"passed": true,
"matched": [
"woman",
"golden retriever",
"beach",
"high-five"
],
"required": [
[
"woman",
"person"
],
[
"golden retriever",
"dog"
],
[
"beach",
"sand"
],
[
"high-five",
"high five"
]
]
},
"exact_output_reference": null,
"exact_output_passed": null,
"exact_output_vs_baseline": false
},
"speedup_vs_baseline": {
"ttft": 7.942,
"e2e": 1.803,
"throughput": 1.033
},
"summary": {
"preprocess_ms_mean": 4.583,
"ttft_ms": {
"p50": 88.172,
"p95": 90.227
},
"e2e_ms": {
"p50": 774.062,
"p95": 780.506
},
"output_tokens_per_second_mean": 43.707,
"peak_vram_gib": 4.001,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"warmup_samples": [
{
"preprocess_ms": 3.98,
"ttft_ms": 88.636,
"e2e_ms": 768.672,
"output_tokens": 31,
"output_tokens_per_second": 44.115,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.299,
"ttft_ms": 88.223,
"e2e_ms": 773.967,
"output_tokens": 31,
"output_tokens_per_second": 43.748,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
],
"samples": [
{
"preprocess_ms": 4.915,
"ttft_ms": 89.425,
"e2e_ms": 778.012,
"output_tokens": 31,
"output_tokens_per_second": 43.567,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.91,
"ttft_ms": 90.767,
"e2e_ms": 773.458,
"output_tokens": 31,
"output_tokens_per_second": 43.944,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.475,
"ttft_ms": 87.596,
"e2e_ms": 780.921,
"output_tokens": 31,
"output_tokens_per_second": 43.27,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 5.156,
"ttft_ms": 89.35,
"e2e_ms": 772.663,
"output_tokens": 31,
"output_tokens_per_second": 43.904,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.618,
"ttft_ms": 87.633,
"e2e_ms": 773.253,
"output_tokens": 31,
"output_tokens_per_second": 43.756,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.432,
"ttft_ms": 88.058,
"e2e_ms": 779.999,
"output_tokens": 31,
"output_tokens_per_second": 43.356,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.676,
"ttft_ms": 89.566,
"e2e_ms": 768.915,
"output_tokens": 31,
"output_tokens_per_second": 44.16,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.846,
"ttft_ms": 87.046,
"e2e_ms": 778.431,
"output_tokens": 31,
"output_tokens_per_second": 43.391,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.991,
"ttft_ms": 88.103,
"e2e_ms": 774.665,
"output_tokens": 31,
"output_tokens_per_second": 43.696,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.815,
"ttft_ms": 88.242,
"e2e_ms": 769.596,
"output_tokens": 31,
"output_tokens_per_second": 44.03,
"peak_vram_gib": 4.001,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
},
{
"id": "02",
"label": "BF16 SDPA / static compiled cache / 448px edge",
"longest_edge": 448,
"cache": "static",
"warmups": 6,
"exact_reference": "01b",
"processed_width": 448,
"processed_height": 299,
"quality_gate": {
"method": "required visual concepts plus byte-identical output against variant 01b",
"passed": true,
"concepts": {
"passed": true,
"matched": [
"woman",
"golden retriever",
"beach",
"high-five"
],
"required": [
[
"woman",
"person"
],
[
"golden retriever",
"dog"
],
[
"beach",
"sand"
],
[
"high-five",
"high five"
]
]
},
"exact_output_reference": "01b",
"exact_output_passed": true,
"exact_output_vs_baseline": false
},
"speedup_vs_baseline": {
"ttft": 9.142,
"e2e": 4.811,
"throughput": 3.295
},
"summary": {
"preprocess_ms_mean": 4.24,
"ttft_ms": {
"p50": 76.6,
"p95": 77.349
},
"e2e_ms": {
"p50": 290.124,
"p95": 299.671
},
"output_tokens_per_second_mean": 139.435,
"peak_vram_gib": 4.023,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"warmup_samples": [
{
"preprocess_ms": 4.676,
"ttft_ms": 5838.639,
"e2e_ms": 6436.632,
"output_tokens": 31,
"output_tokens_per_second": 50.168,
"peak_vram_gib": 4.011,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.922,
"ttft_ms": 79.28,
"e2e_ms": 294.627,
"output_tokens": 31,
"output_tokens_per_second": 139.31,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.893,
"ttft_ms": 76.419,
"e2e_ms": 292.685,
"output_tokens": 31,
"output_tokens_per_second": 138.718,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.068,
"ttft_ms": 76.42,
"e2e_ms": 287.574,
"output_tokens": 31,
"output_tokens_per_second": 142.076,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.804,
"ttft_ms": 73.609,
"e2e_ms": 289.638,
"output_tokens": 31,
"output_tokens_per_second": 138.871,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.888,
"ttft_ms": 77.168,
"e2e_ms": 292.967,
"output_tokens": 31,
"output_tokens_per_second": 139.018,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
],
"samples": [
{
"preprocess_ms": 4.199,
"ttft_ms": 73.284,
"e2e_ms": 289.897,
"output_tokens": 31,
"output_tokens_per_second": 138.496,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.649,
"ttft_ms": 76.752,
"e2e_ms": 288.657,
"output_tokens": 31,
"output_tokens_per_second": 141.573,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.599,
"ttft_ms": 76.872,
"e2e_ms": 288.993,
"output_tokens": 31,
"output_tokens_per_second": 141.429,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.903,
"ttft_ms": 76.448,
"e2e_ms": 287.537,
"output_tokens": 31,
"output_tokens_per_second": 142.121,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.784,
"ttft_ms": 76.333,
"e2e_ms": 303.535,
"output_tokens": 31,
"output_tokens_per_second": 132.041,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.625,
"ttft_ms": 77.362,
"e2e_ms": 294.949,
"output_tokens": 31,
"output_tokens_per_second": 137.876,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.001,
"ttft_ms": 75.874,
"e2e_ms": 290.573,
"output_tokens": 31,
"output_tokens_per_second": 139.731,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.878,
"ttft_ms": 75.599,
"e2e_ms": 287.892,
"output_tokens": 31,
"output_tokens_per_second": 141.315,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.702,
"ttft_ms": 77.109,
"e2e_ms": 290.35,
"output_tokens": 31,
"output_tokens_per_second": 140.686,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 5.057,
"ttft_ms": 77.332,
"e2e_ms": 293.032,
"output_tokens": 31,
"output_tokens_per_second": 139.082,
"peak_vram_gib": 4.023,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
}
]
}
@@ -0,0 +1,181 @@
{
"schema": "comfyui-vlm/transformers-benchmark",
"version": 1,
"created_at": "2026-08-07T23:06:56.398807+00:00",
"label": "qwen3-vl-2b-sdpa-static-edge448-tf32",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"git_commit": "858a4a4998e68bd99b26a12831e861be49904564",
"git_dirty": true,
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"sha256": "9eeaa87013b4e800930e8a411b58ff9e2fd5383906b1a022f4a712720af34cc2",
"source_width": 2048,
"source_height": 1365,
"processed_width": 448,
"processed_height": 299
},
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.8.0+cu128",
"cuda": "12.8",
"gpu": "NVIDIA GeForce RTX 3090",
"transformers": "5.12.1",
"flash_attn": null
},
"settings": {
"attention": "sdpa",
"cache": "static",
"disable_compile": false,
"min_pixels": null,
"max_pixels": null,
"longest_edge": 448,
"max_new_tokens": 96,
"warmups": 6,
"runs": 10,
"float32_matmul_precision": "high"
},
"model_load_seconds": 83.958,
"quality_gate": {
"method": "byte-identical output SHA-256",
"reference_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1",
"passed": true
},
"summary": {
"preprocess_ms_mean": 4.357,
"ttft_ms": {
"p50": 74.256,
"p95": 79.529
},
"e2e_ms": {
"p50": 274.547,
"p95": 289.619
},
"output_tokens_per_second_mean": 149.167,
"peak_vram_gib": 4.025,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"samples": [
{
"preprocess_ms": 5.029,
"ttft_ms": 79.329,
"e2e_ms": 278.01,
"output_tokens": 31,
"output_tokens_per_second": 150.996,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.893,
"ttft_ms": 79.693,
"e2e_ms": 276.153,
"output_tokens": 31,
"output_tokens_per_second": 152.703,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.754,
"ttft_ms": 71.308,
"e2e_ms": 269.219,
"output_tokens": 31,
"output_tokens_per_second": 151.584,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.838,
"ttft_ms": 74.473,
"e2e_ms": 272.941,
"output_tokens": 31,
"output_tokens_per_second": 151.158,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.331,
"ttft_ms": 74.039,
"e2e_ms": 281.754,
"output_tokens": 31,
"output_tokens_per_second": 144.429,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.426,
"ttft_ms": 73.694,
"e2e_ms": 272.792,
"output_tokens": 31,
"output_tokens_per_second": 150.68,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 3.908,
"ttft_ms": 72.445,
"e2e_ms": 269.177,
"output_tokens": 31,
"output_tokens_per_second": 152.492,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.856,
"ttft_ms": 74.703,
"e2e_ms": 276.675,
"output_tokens": 31,
"output_tokens_per_second": 148.536,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.905,
"ttft_ms": 78.645,
"e2e_ms": 296.054,
"output_tokens": 31,
"output_tokens_per_second": 137.989,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.63,
"ttft_ms": 73.816,
"e2e_ms": 272.358,
"output_tokens": 31,
"output_tokens_per_second": 151.101,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
}
@@ -0,0 +1,247 @@
{
"schema": "comfyui-vlm/tensorrt-vision-probe",
"version": 1,
"created_at": "2026-08-08T00:49:06.960727+00:00",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"processed_size": [
448,
299
],
"pixel_values_shape": [
504,
1536
],
"image_grid_thw": [
[
1,
18,
28
]
]
},
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.9.0+cu128",
"cuda": "12.8",
"torch_tensorrt": "2.9.0+cu128",
"tensorrt": "10.13.3.9.post1",
"transformers": "5.12.1",
"gpu": "NVIDIA GeForce RTX 3090"
},
"configuration": {
"precision": "bfloat16",
"min_block_size": 5,
"optimization_level": 3,
"require_full_compilation": false,
"warmups": 3,
"runs": 10,
"generation_warmups": 1,
"generation_runs": 3
},
"model_load_seconds": 78.171,
"compile_seconds": 98.07,
"coverage": {
"available": true,
"graph_nodes": 8,
"call_modules": [
"_run_on_acc_0"
],
"tensorrt_engine_partitions": 1,
"pytorch_fallback_partitions": 0
},
"fidelity": [
{
"output_index": 0,
"shape": [
504,
1024
],
"max_absolute_error": 1380.0,
"mean_absolute_error": 0.563834547996521,
"cosine_similarity": 0.9949914216995239
},
{
"output_index": 1,
"shape": [
126,
2048
],
"max_absolute_error": 2.375,
"mean_absolute_error": 0.025903113186359406,
"cosine_similarity": 0.9966588020324707
},
{
"output_index": 2,
"shape": [
126,
2048
],
"max_absolute_error": 0.09375,
"mean_absolute_error": 0.004035853315144777,
"cosine_similarity": 0.9999096393585205
},
{
"output_index": 3,
"shape": [
126,
2048
],
"max_absolute_error": 0.978515625,
"mean_absolute_error": 0.009971227496862411,
"cosine_similarity": 0.9982407093048096
},
{
"output_index": 4,
"shape": [
126,
2048
],
"max_absolute_error": 3.34375,
"mean_absolute_error": 0.018642043694853783,
"cosine_similarity": 0.998470664024353
}
],
"latency_ms": {
"eager_samples": [
2391.378,
2382.472,
2386.845,
2408.021,
2382.034,
2381.989,
2431.839,
2419.309,
2382.089,
2383.834
],
"tensorrt_samples": [
9.526,
8.409,
8.817,
9.337,
8.556,
9.596,
8.549,
9.892,
8.632,
9.572
],
"eager_median": 2385.339,
"tensorrt_median": 9.077,
"speedup": 262.779
},
"generation": {
"eager": {
"preprocess_ms_mean": 4.546,
"ttft_ms": {
"p50": 2452.88,
"p95": 2464.33
},
"e2e_ms": {
"p50": 2660.594,
"p95": 2674.669
},
"output_tokens_per_second_mean": 143.283,
"peak_vram_gib": 4.028,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"tensorrt": {
"preprocess_ms_mean": 6.582,
"ttft_ms": {
"p50": 61.388,
"p95": 62.364
},
"e2e_ms": {
"p50": 273.449,
"p95": 274.016
},
"output_tokens_per_second_mean": 142.09,
"peak_vram_gib": 4.024,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"exact_output_match": true,
"eager_output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"tensorrt_output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"eager_samples": [
{
"preprocess_ms": 4.702,
"ttft_ms": 2465.602,
"e2e_ms": 2676.233,
"output_tokens": 31,
"output_tokens_per_second": 142.429,
"peak_vram_gib": 4.028,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.443,
"ttft_ms": 2452.88,
"e2e_ms": 2660.594,
"output_tokens": 31,
"output_tokens_per_second": 144.429,
"peak_vram_gib": 4.028,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 4.493,
"ttft_ms": 2446.741,
"e2e_ms": 2656.543,
"output_tokens": 31,
"output_tokens_per_second": 142.992,
"peak_vram_gib": 4.028,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
],
"tensorrt_samples": [
{
"preprocess_ms": 6.038,
"ttft_ms": 60.673,
"e2e_ms": 270.425,
"output_tokens": 31,
"output_tokens_per_second": 143.026,
"peak_vram_gib": 4.024,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 6.937,
"ttft_ms": 62.472,
"e2e_ms": 273.449,
"output_tokens": 31,
"output_tokens_per_second": 142.196,
"peak_vram_gib": 4.024,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
},
{
"preprocess_ms": 6.772,
"ttft_ms": 61.388,
"e2e_ms": 274.079,
"output_tokens": 31,
"output_tokens_per_second": 141.049,
"peak_vram_gib": 4.024,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
}
}
@@ -0,0 +1,189 @@
{
"schema": "comfyui-vlm/tensorrt-vision-probe",
"version": 1,
"created_at": "2026-08-08T07:02:36.436372+00:00",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"media": {
"path": "/home/gokaygokay/ComfyUI_VLM_nodes/benchmarks/media/qwen-demo.jpeg",
"processed_size": [
448,
299
],
"pixel_values_shape": [
504,
1536
],
"image_grid_thw": [
[
1,
18,
28
]
]
},
"environment": {
"platform": "Linux-6.18.33.2-microsoft-standard-WSL2-x86_64-with-glibc2.35",
"python": "3.11.14",
"torch": "2.9.0+cu128",
"cuda": "12.8",
"torch_tensorrt": "2.9.0+cu128",
"tensorrt": "10.13.3.9.post1",
"transformers": "5.12.1",
"gpu": "NVIDIA GeForce RTX 3090"
},
"configuration": {
"precision": "bfloat16",
"min_block_size": 5,
"optimization_level": 3,
"require_full_compilation": true,
"warmups": 2,
"runs": 5,
"generation_warmups": 0,
"generation_runs": 1
},
"model_load_seconds": 80.059,
"compile_seconds": 96.692,
"coverage": {
"available": true,
"graph_nodes": 8,
"call_modules": [
"_run_on_acc_0"
],
"tensorrt_engine_partitions": 1,
"pytorch_fallback_partitions": 0
},
"fidelity": [
{
"output_index": 0,
"shape": [
504,
1024
],
"max_absolute_error": 1380.0,
"mean_absolute_error": 0.563834547996521,
"cosine_similarity": 0.9949914216995239
},
{
"output_index": 1,
"shape": [
126,
2048
],
"max_absolute_error": 2.375,
"mean_absolute_error": 0.025903113186359406,
"cosine_similarity": 0.9966588020324707
},
{
"output_index": 2,
"shape": [
126,
2048
],
"max_absolute_error": 0.09375,
"mean_absolute_error": 0.004035853315144777,
"cosine_similarity": 0.9999096393585205
},
{
"output_index": 3,
"shape": [
126,
2048
],
"max_absolute_error": 0.978515625,
"mean_absolute_error": 0.009971227496862411,
"cosine_similarity": 0.9982407093048096
},
{
"output_index": 4,
"shape": [
126,
2048
],
"max_absolute_error": 3.34375,
"mean_absolute_error": 0.018642043694853783,
"cosine_similarity": 0.998470664024353
}
],
"latency_ms": {
"eager_samples": [
2371.396,
2388.794,
2374.26,
2374.196,
2362.955
],
"tensorrt_samples": [
8.405,
9.066,
8.388,
9.738,
8.271
],
"eager_median": 2374.196,
"tensorrt_median": 8.405,
"speedup": 282.491
},
"generation": {
"eager": {
"preprocess_ms_mean": 6.771,
"ttft_ms": {
"p50": 42960.862,
"p95": 42960.862
},
"e2e_ms": {
"p50": 43196.275,
"p95": 43196.275
},
"output_tokens_per_second_mean": 127.436,
"peak_vram_gib": 4.025,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"tensorrt": {
"preprocess_ms_mean": 8.081,
"ttft_ms": {
"p50": 325.253,
"p95": 325.253
},
"e2e_ms": {
"p50": 552.777,
"p95": 552.777
},
"output_tokens_per_second_mean": 131.854,
"peak_vram_gib": 4.022,
"output_tokens_mean": 31.0,
"outputs_identical": true
},
"exact_output_match": true,
"eager_output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"tensorrt_output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"eager_samples": [
{
"preprocess_ms": 6.771,
"ttft_ms": 42960.862,
"e2e_ms": 43196.275,
"output_tokens": 31,
"output_tokens_per_second": 127.436,
"peak_vram_gib": 4.025,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
],
"tensorrt_samples": [
{
"preprocess_ms": 8.081,
"ttft_ms": 325.253,
"e2e_ms": 552.777,
"output_tokens": 31,
"output_tokens_per_second": 131.854,
"peak_vram_gib": 4.022,
"input_tokens": 144,
"vision_tokens": 504,
"output": "A woman and her golden retriever share a joyful moment on a sunlit beach, with the dog playfully reaching out to give a high-five.",
"output_sha256": "7b1c4202212d7ecbb5c90cb79ac7d6395cf487e97fdd7980f0c43687104184c1"
}
]
}
}
@@ -0,0 +1,5 @@
"""Process bootstrap for the Qwen3-VL TensorRT/SGLang benchmark."""
from qwen3_vl_sglang_tensorrt_bridge import install_bridge
install_bridge()
+40
View File
@@ -0,0 +1,40 @@
# See https://help.github.com/articles/ignoring-files/ for more about ignoring files.
# dependencies
/node_modules
/.pnp
.pnp.*
.yarn/*
!.yarn/patches
!.yarn/plugins
!.yarn/releases
!.yarn/versions
# testing
/coverage
# next.js
/.next/
/.vinext/
/out/
# misc
.DS_Store
*.pem
# debug
npm-debug.log*
yarn-debug.log*
yarn-error.log*
.pnpm-debug.log*
# env files (can opt-in for committing if needed)
.env*
# vercel
.vercel
/dist/
/.wrangler/
/outputs/
/work/
+5
View File
@@ -0,0 +1,5 @@
{
"project_id": "appgprj_6a7654db28d48191858ee067884b6d34",
"d1": null,
"r2": null
}
+100
View File
@@ -0,0 +1,100 @@
# vinext-starter
A clean full-stack starter running on
[vinext](https://github.com/cloudflare/vinext), with optional Cloudflare D1 and
Drizzle support.
## Prerequisites
- Node.js `>=22.13.0`
## Quick Start
```bash
npm install
npm run dev
npm run build
```
This starter does not use `wrangler.jsonc`.
## Included Shape
- edit site code under `app/`
- `.openai/hosting.json` declares optional Sites D1 and R2 bindings
- `vite.config.ts` simulates declared bindings for local development
- `db/schema.ts` starts intentionally empty
- `examples/d1/` contains an optional D1 example surface
- `drizzle.config.ts` supports local migration generation when needed
## Workspace Auth Headers
Signed-in visitors receive both `oai-authenticated-user-id` and `oai-authenticated-user-email`. Private Sites require every visitor to sign in; public Sites may also have anonymous visitors, for whom neither header is present.
The user ID is stable for the same user on the same Site and different across Sites. Email and name are intended for display or contact purposes.
SIWC-authenticated workspace sites may also receive
`oai-authenticated-user-full-name` when the user's SIWC profile has a non-empty
`name` claim. The full-name value is percent-encoded UTF-8 and is accompanied by
`oai-authenticated-user-full-name-encoding: percent-encoded-utf-8`.
Treat the full name as optional and fall back to email when it is absent:
```tsx
import { headers } from "next/headers";
export default async function Home() {
const requestHeaders = await headers();
const userId = requestHeaders.get("oai-authenticated-user-id");
const email = requestHeaders.get("oai-authenticated-user-email");
const encodedFullName = requestHeaders.get("oai-authenticated-user-full-name");
const fullName =
encodedFullName &&
requestHeaders.get("oai-authenticated-user-full-name-encoding") ===
"percent-encoded-utf-8"
? decodeURIComponent(encodedFullName)
: null;
const displayName = fullName ?? email;
// ...
}
```
## Optional Dispatch-Owned ChatGPT Sign-In
Import the ready-to-use helpers from `app/chatgpt-auth.ts` when the site needs
optional or required ChatGPT sign-in:
- Use `getChatGPTUser()` for optional signed-in UI.
- Use `requireChatGPTUser(returnTo)` for server-rendered pages that should send
anonymous visitors through Sign in with ChatGPT.
- Use `chatGPTSignInPath(returnTo)` and `chatGPTSignOutPath(returnTo)` for
browser links or actions.
- Pass a same-origin relative `returnTo` path for the destination after sign-in
or sign-out. The helper validates and safely encodes it.
- Mark protected pages with `export const dynamic = "force-dynamic"` because
they depend on per-request identity headers.
Dispatch owns `/signin-with-chatgpt`, `/signout-with-chatgpt`, `/callback`, the
OAuth cookies, and identity header injection. Do not implement app routes for
those reserved paths. Routes that do not import and call the helper remain
anonymous-compatible.
SIWC establishes identity only; it does not prove workspace membership. Use the
Sites hosting platform's access policy controls for workspace-wide restrictions,
or enforce explicit server-side membership or allowlist checks.
Use SIWC for account pages, user-specific dashboards, saved records, and write
actions tied to the current ChatGPT user. Leave public content anonymous.
## Useful Commands
- `npm run dev`: start local development
- `npm run build`: verify the vinext build output
- `npm test`: build the starter and verify its rendered loading skeleton
- `npm run db:generate`: generate Drizzle migrations after schema changes
## Learn More
- [vinext Documentation](https://github.com/cloudflare/vinext)
- [Drizzle D1 Guide](https://orm.drizzle.team/docs/get-started/d1-new)
+90
View File
@@ -0,0 +1,90 @@
import { headers } from "next/headers";
import { redirect } from "next/navigation";
export type ChatGPTUser = {
userId: string;
displayName: string;
email: string;
fullName: string | null;
};
const USER_ID_HEADER = "oai-authenticated-user-id";
const USER_EMAIL_HEADER = "oai-authenticated-user-email";
const USER_FULL_NAME_HEADER = "oai-authenticated-user-full-name";
const USER_FULL_NAME_ENCODING_HEADER =
"oai-authenticated-user-full-name-encoding";
const PERCENT_ENCODED_UTF8 = "percent-encoded-utf-8";
const SIGN_IN_PATH = "/signin-with-chatgpt";
const SIGN_OUT_PATH = "/signout-with-chatgpt";
const CALLBACK_PATH = "/callback";
export async function getChatGPTUser(): Promise<ChatGPTUser | null> {
const requestHeaders = await headers();
const userId = requestHeaders.get(USER_ID_HEADER);
const email = requestHeaders.get(USER_EMAIL_HEADER);
if (!userId || !email) return null;
const encodedFullName = requestHeaders.get(USER_FULL_NAME_HEADER);
const fullName =
encodedFullName &&
requestHeaders.get(USER_FULL_NAME_ENCODING_HEADER) === PERCENT_ENCODED_UTF8
? safeDecodeURIComponent(encodedFullName)
: null;
return {
userId,
displayName: fullName ?? email,
email,
fullName,
};
}
export async function requireChatGPTUser(
returnTo: string,
): Promise<ChatGPTUser> {
const user = await getChatGPTUser();
if (user) return user;
redirect(chatGPTSignInPath(returnTo));
}
export function chatGPTSignInPath(returnTo: string): string {
const safeReturnTo = safeRelativeReturnPath(returnTo);
return `${SIGN_IN_PATH}?return_to=${encodeURIComponent(safeReturnTo)}`;
}
export function chatGPTSignOutPath(returnTo = "/"): string {
const safeReturnTo = safeRelativeReturnPath(returnTo);
return `${SIGN_OUT_PATH}?return_to=${encodeURIComponent(safeReturnTo)}`;
}
function safeRelativeReturnPath(value: string): string {
if (!value.startsWith("/") || value.startsWith("//")) return "/";
let url: URL;
try {
url = new URL(value, "https://app.local");
} catch {
return "/";
}
if (url.origin !== "https://app.local") return "/";
if (isReservedAuthPath(url.pathname)) return "/";
return `${url.pathname}${url.search}${url.hash}`;
}
function isReservedAuthPath(pathname: string): boolean {
return (
pathname === SIGN_IN_PATH ||
pathname === SIGN_OUT_PATH ||
pathname === CALLBACK_PATH
);
}
function safeDecodeURIComponent(value: string): string | null {
try {
return decodeURIComponent(value);
} catch {
return null;
}
}
+157
View File
@@ -0,0 +1,157 @@
@import "tailwindcss";
:root { --ink:#11130f; --paper:#f3f0e7; --lime:#c8ff36; --orange:#ff6b35; --muted:#706f68; --line:#d3d0c5; }
* { box-sizing:border-box; }
html { scroll-behavior:smooth; }
body { margin:0; background:var(--paper); color:var(--ink); font-family:var(--font-geist-sans), Arial, sans-serif; }
a { color:inherit; text-decoration:none; }
.topbar { height:76px; padding:0 4.5vw; display:flex; align-items:center; justify-content:space-between; border-bottom:1px solid var(--line); position:sticky; top:0; z-index:10; background:rgba(243,240,231,.92); backdrop-filter:blur(16px); }
.brand { display:flex; align-items:center; gap:11px; font-weight:760; letter-spacing:-.02em; }
.brand-mark { display:grid; place-items:center; width:34px; height:34px; background:var(--ink); color:var(--lime); font:700 11px var(--font-geist-mono); transform:rotate(-3deg); }
nav { display:flex; gap:34px; color:#55564f; font-size:13px; }
nav a:hover { color:var(--ink); }
.repo-link { font:650 12px var(--font-geist-mono); border-bottom:1px solid var(--ink); padding-bottom:3px; }
.hero { min-height:690px; padding:72px 6vw 70px; display:grid; grid-template-columns:1.18fr .82fr; gap:7vw; align-items:center; overflow:hidden; background-image:linear-gradient(rgba(17,19,15,.035) 1px,transparent 1px),linear-gradient(90deg,rgba(17,19,15,.035) 1px,transparent 1px); background-size:42px 42px; }
.eyebrow { font:700 11px/1.2 var(--font-geist-mono); letter-spacing:.12em; text-transform:uppercase; display:flex; align-items:center; gap:9px; }
.eyebrow.light { color:var(--lime); }
.live-dot { width:8px; height:8px; border-radius:99px; background:var(--orange); box-shadow:0 0 0 5px rgba(255,107,53,.15); }
h1 { font-size:clamp(60px,7.2vw,116px); line-height:.84; letter-spacing:-.075em; margin:31px 0 30px; font-weight:770; }
h1 em { font-family:Georgia,serif; font-weight:400; color:var(--orange); }
.lede { font-size:18px; line-height:1.55; max-width:620px; color:#484a44; }
.hero-actions { margin-top:38px; display:flex; align-items:center; gap:24px; }
.primary-button { display:inline-flex; gap:24px; align-items:center; background:var(--ink); color:white; padding:18px 21px; font-weight:650; font-size:14px; }
.primary-button span { color:var(--lime); font-size:20px; }
.artifact-note { font:600 10px var(--font-geist-mono); color:var(--muted); text-transform:uppercase; letter-spacing:.08em; }
.hero-metric { position:relative; border:1px solid var(--ink); padding:24px 25px 0; background:#e9e6dc; box-shadow:13px 13px 0 var(--ink); transform:rotate(1deg); }
.metric-topline { display:flex; justify-content:space-between; font:600 10px var(--font-geist-mono); text-transform:uppercase; letter-spacing:.08em; }
.verified { color:#497400; }
.big-number { font-size:clamp(95px,12vw,184px); font-weight:800; line-height:.9; letter-spacing:-.085em; margin:25px 0 0; }
.big-number span { color:var(--orange); font-size:.45em; vertical-align:top; position:relative; top:18px; }
.metric-label { font-size:22px; font-weight:680; letter-spacing:-.03em; margin-bottom:35px; }
.work-bars { display:grid; gap:12px; padding:20px 0; border-top:1px solid var(--line); }
.work-row { display:grid; grid-template-columns:52px 1fr 52px; gap:12px; align-items:center; font:600 10px var(--font-geist-mono); }
.work-row b { text-align:right; }
.bar { height:11px; background:var(--ink); display:block; }
.bar.after { width:9%; background:var(--orange); min-width:9px; }
.hero-metric>p { font:500 10px/1.5 var(--font-geist-mono); color:var(--muted); }
.honesty-strip { margin:20px -25px 0; padding:12px 25px; background:var(--lime); font:700 9px var(--font-geist-mono); text-transform:uppercase; letter-spacing:.07em; }
.manifesto-band { background:var(--ink); color:white; min-height:74px; display:flex; align-items:center; justify-content:space-around; gap:24px; padding:16px 4vw; font:650 10px var(--font-geist-mono); text-transform:uppercase; letter-spacing:.08em; }
.manifesto-band span::first-letter { color:var(--lime); }
.section { padding:110px 6vw; }
.section-heading { display:grid; grid-template-columns:1fr minmax(280px,440px); align-items:end; gap:40px; margin-bottom:58px; }
h2 { font-size:clamp(43px,5vw,75px); line-height:.97; letter-spacing:-.058em; margin:18px 0 0; }
.section-heading>p { color:var(--muted); line-height:1.6; font-size:14px; margin:0; }
.run-context { display:grid; grid-template-columns:repeat(5,minmax(0,1fr)); border:1px solid var(--ink); border-bottom:0; background:#e8e5db; }
.run-context span { min-width:0; padding:12px 14px; border-right:1px solid var(--line); font:500 9px/1.45 var(--font-geist-mono); color:var(--muted); text-transform:uppercase; }
.run-context span:last-child { border-right:0; }
.run-context b { color:var(--ink); margin-right:7px; }
.comparison-table-wrap { overflow-x:auto; border:1px solid var(--ink); }
.comparison-table { width:100%; min-width:1420px; border-collapse:collapse; table-layout:fixed; font-size:11px; }
.comparison-table th,.comparison-table td { padding:14px 12px; text-align:left; border-right:1px solid var(--line); border-bottom:1px solid var(--line); vertical-align:middle; }
.comparison-table th:last-child,.comparison-table td:last-child { border-right:0; }
.comparison-table tbody tr:last-child>* { border-bottom:0; }
.comparison-table thead { background:var(--ink); color:white; }
.comparison-table thead th { padding-top:11px; padding-bottom:11px; font:650 9px/1.25 var(--font-geist-mono); text-transform:uppercase; letter-spacing:.06em; color:#deddd7; border-color:#3c3d38; }
.comparison-table thead span { color:#858780; font-size:8px; }
.comparison-table th:nth-child(1) { width:42px; }
.comparison-table th:nth-child(2) { width:180px; }
.comparison-table th:nth-child(3) { width:160px; }
.comparison-table th:nth-child(4) { width:105px; }
.comparison-table th:nth-child(5),.comparison-table th:nth-child(6) { width:92px; }
.comparison-table th:nth-child(7),.comparison-table th:nth-child(8) { width:72px; }
.comparison-table th:nth-child(9) { width:150px; }
.comparison-table th:nth-child(10),.comparison-table th:nth-child(11) { width:100px; }
.comparison-table th:nth-child(12) { width:82px; }
.comparison-table tbody tr:hover { background:#eae7de; }
.comparison-table tbody tr.measured { background:rgba(200,255,54,.11); }
.comparison-table tbody tr.measured:hover { background:rgba(200,255,54,.2); }
.comparison-table tbody tr.regression { background:rgba(255,180,53,.09); }
.comparison-table tbody tr.rejected { background:rgba(255,107,53,.1); }
.row-id { color:var(--muted); font:600 10px var(--font-geist-mono); }
.variant-cell strong { display:block; font-size:12px; letter-spacing:-.015em; }
.variant-cell span { display:block; margin-top:4px; color:var(--muted); font:500 8px var(--font-geist-mono); text-transform:uppercase; }
.input-cell strong { display:block; font:650 10px var(--font-geist-mono); }
.input-cell span { display:block; margin-top:4px; color:var(--muted); font:500 8px var(--font-geist-mono); }
.metric-cell { font:650 11px var(--font-geist-mono); font-variant-numeric:tabular-nums; }
.metric-cell strong { display:block; font:inherit; }
.metric-cell span { display:block; margin-top:3px; color:var(--muted); font-size:8px; }
.muted-cell { color:var(--muted); }
.quality-ok { color:#557900; font-weight:700; }
.speedup-cell { color:#456700; font:700 9px/1.45 var(--font-geist-mono); }
.regression-cell { color:#a5421a; font:700 9px/1.45 var(--font-geist-mono); }
.fidelity-cell { font-weight:750; color:#456700; }
.quality-fail { color:#bd351f; font-weight:750; }
.status { height:23px; padding:5px 9px; border:1px solid; border-radius:99px; font:700 8px var(--font-geist-mono); letter-spacing:.08em; text-transform:uppercase; }
.status-measured { background:var(--lime); border-color:var(--lime); }
.status-regression { color:#8a4710; background:#ffe0a3; border-color:#f2bd58; }
.status-rejected { color:#9e2d1c; background:#ffd3c4; border-color:#ff9e7b; }
.status-ready,.status-queued { color:#ad451c; background:#ffe1d5; border-color:#ffc1a9; }
.status-planned { color:var(--muted); border-color:var(--line); }
.table-notes { display:flex; gap:26px; padding:13px 2px 0; color:var(--muted); font:500 9px var(--font-geist-mono); }
.table-notes b { color:var(--ink); margin-right:5px; }
.quality-section { background:var(--ink); color:white; padding:115px 6vw; display:grid; grid-template-columns:.8fr 1.2fr; gap:8vw; }
.quality-intro>p { color:#aaa9a3; max-width:480px; line-height:1.6; margin-top:24px; }
.gate-formula { display:flex; flex-direction:column; gap:9px; margin-top:45px; border-left:2px solid var(--lime); padding:4px 0 4px 18px; }
.gate-formula span { color:#8c8d87; font:600 9px var(--font-geist-mono); text-transform:uppercase; }
.gate-formula code { color:var(--lime); font-size:13px; }
.quality-table { border-top:1px solid #555650; }
.quality-row { display:grid; grid-template-columns:1fr 1.25fr 1fr; gap:15px; padding:23px 10px; border-bottom:1px solid #3b3c37; font-size:12px; }
.quality-row.header { color:#777973; font:600 9px var(--font-geist-mono); text-transform:uppercase; }
.quality-row span { color:#aaa9a3; }
.quality-row b { color:var(--lime); font:600 10px var(--font-geist-mono); }
.quality-proof { background:var(--lime); color:var(--ink); margin-top:24px; padding:22px; display:flex; gap:17px; }
.proof-icon { display:grid; place-items:center; flex:0 0 36px; height:36px; border-radius:99px; background:var(--ink); color:var(--lime); }
.quality-proof strong { font-size:14px; }
.quality-proof p { margin:5px 0 0; font-size:11px; line-height:1.5; }
.protocol-heading { align-items:center; }
.commit-chip { justify-self:end; border:1px solid var(--line); padding:11px 15px; font:500 10px var(--font-geist-mono); }
.commit-chip code { color:var(--orange); }
.protocol-grid { display:grid; grid-template-columns:repeat(4,1fr); border-top:1px solid var(--ink); border-bottom:1px solid var(--ink); }
.protocol-grid article { min-height:230px; padding:25px; border-right:1px solid var(--line); }
.protocol-grid article:last-child { border-right:0; }
.protocol-grid article>span { color:var(--orange); font:700 11px var(--font-geist-mono); }
.protocol-grid h3 { margin:48px 0 10px; font-size:23px; letter-spacing:-.04em; }
.protocol-grid p { color:var(--muted); font-size:12px; line-height:1.6; }
.metric-strip { margin-top:60px; display:grid; grid-template-columns:repeat(5,1fr); background:#e7e4da; }
.metric-strip>div { padding:21px; border-right:1px solid var(--paper); }
.metric-strip small { display:block; color:var(--muted); font:600 9px var(--font-geist-mono); text-transform:uppercase; margin-bottom:8px; }
.metric-strip strong { font-size:13px; }
.next-run { background:var(--orange); padding:80px 6vw; display:grid; grid-template-columns:1.2fr .8fr; gap:8vw; align-items:end; }
.next-run .eyebrow { color:var(--ink); }
.next-run-copy { line-height:1.6; font-size:14px; }
.next-run-copy a { font:700 10px var(--font-geist-mono); text-transform:uppercase; border-bottom:1px solid; padding-bottom:4px; }
footer { min-height:110px; padding:25px 4.5vw; background:var(--ink); color:#aaa9a3; display:flex; justify-content:space-between; align-items:center; font-size:10px; }
footer .brand { color:white; }
@media (max-width:900px) {
nav { display:none; }
.hero { grid-template-columns:1fr; padding-top:60px; }
.hero-metric { max-width:600px; }
.section-heading,.quality-section,.next-run { grid-template-columns:1fr; }
.protocol-grid { grid-template-columns:1fr 1fr; }
.protocol-grid article:nth-child(2) { border-right:0; }
.metric-strip { grid-template-columns:1fr 1fr; }
.run-context { grid-template-columns:1fr 1fr; }
.run-context span:nth-child(2) { border-right:0; }
.run-context span:nth-child(-n+2) { border-bottom:1px solid var(--line); }
}
@media (max-width:560px) {
.topbar { padding:0 20px; }
.repo-link { font-size:9px; }
.hero,.section,.quality-section,.next-run { padding-left:22px; padding-right:22px; }
h1 { font-size:58px; }
.hero-actions { align-items:flex-start; flex-direction:column; }
.manifesto-band { justify-content:flex-start; overflow:auto; }
.manifesto-band span { white-space:nowrap; }
.section-heading { grid-template-columns:1fr; }
.run-context { grid-template-columns:1fr; }
.run-context span { border-right:0; border-bottom:1px solid var(--line); }
.run-context span:nth-child(3) { border-bottom:1px solid var(--line); }
.table-notes { flex-direction:column; gap:7px; }
.quality-row { grid-template-columns:.75fr 1.2fr; }
.quality-row>*:last-child { grid-column:2; }
.protocol-grid,.metric-strip { grid-template-columns:1fr; }
.protocol-grid article { border-right:0; border-bottom:1px solid var(--line); }
footer { align-items:flex-start; gap:20px; flex-direction:column; }
}
@media (prefers-reduced-motion:reduce) { html { scroll-behavior:auto; } * { transition:none!important; } }
+37
View File
@@ -0,0 +1,37 @@
import type { Metadata } from "next";
import { Geist, Geist_Mono } from "next/font/google";
import { headers } from "next/headers";
import "./globals.css";
const geistSans = Geist({ variable: "--font-geist-sans", subsets: ["latin"] });
const geistMono = Geist_Mono({ variable: "--font-geist-mono", subsets: ["latin"] });
export async function generateMetadata(): Promise<Metadata> {
const requestHeaders = await headers();
const host = requestHeaders.get("x-forwarded-host") ?? requestHeaders.get("host") ?? "localhost:3000";
const protocol = requestHeaders.get("x-forwarded-proto") ?? (host.startsWith("localhost") ? "http" : "https");
const origin = `${protocol}://${host}`;
return {
metadataBase: new URL(origin),
title: { default: "VLM Speed Lab", template: "%s · VLM Speed Lab" },
description: "Measured VLM speedups with reproducible quality evidence.",
icons: { icon: "/favicon.svg", shortcut: "/favicon.svg" },
openGraph: {
title: "VLM Speed Lab",
description: "Make it faster. Prove it stayed good.",
type: "website",
url: origin,
images: [{ url: `${origin}/og.png`, width: 1200, height: 630, alt: "VLM Speed Lab — measured, not marketed" }],
},
twitter: {
card: "summary_large_image",
title: "VLM Speed Lab",
description: "Make it faster. Prove it stayed good.",
images: [`${origin}/og.png`],
},
};
}
export default function RootLayout({ children }: Readonly<{ children: React.ReactNode }>) {
return <html lang="en"><body className={`${geistSans.variable} ${geistMono.variable}`}>{children}</body></html>;
}
+407
View File
@@ -0,0 +1,407 @@
import type { Metadata } from "next";
export const metadata: Metadata = {
title: "VLM Speed Lab — Qwen3-VL 2B",
description:
"Reproducible VLM performance iterations with latency, throughput, memory, and quality evidence.",
};
const iterations = [
{
id: "00",
name: "Source-resolution control",
stack: "BF16 · SDPA · dynamic",
input: "2048×1365",
tokens: "11,008 vision · 2,770 input",
status: "measured",
change: "Frozen control",
ttft: ["700.3", "723.1"],
e2e: ["1395.7", "1487.7"],
throughput: "42.3",
vram: "4.55",
speedup: "1.00× / 1.00× / 1.00×",
quality: "4/4 concepts",
exact: "10/10",
},
{
id: "01a",
name: "Medium visual budget",
stack: "BF16 · SDPA · dynamic",
input: "672×448",
tokens: "1,176 vision · 312 input",
status: "measured",
change: "Resize only",
ttft: ["112.8", "122.8"],
e2e: ["844.0", "881.7"],
throughput: "42.2",
vram: "4.04",
speedup: "6.21× / 1.65× / 1.00×",
quality: "4/4 concepts",
exact: "Semantic",
},
{
id: "01b",
name: "Aggressive visual budget",
stack: "BF16 · SDPA · dynamic",
input: "448×299",
tokens: "504 vision · 144 input",
status: "measured",
change: "Resize only",
ttft: ["88.2", "90.2"],
e2e: ["774.1", "780.5"],
throughput: "43.7",
vram: "4.00",
speedup: "7.94× / 1.80× / 1.03×",
quality: "4/4 concepts",
exact: "Semantic",
},
{
id: "02",
name: "Compiled execution",
stack: "BF16 · SDPA · static cache",
input: "448×299",
tokens: "504 vision · 144 input",
status: "measured",
change: "Cache + compile",
ttft: ["76.6", "77.3"],
e2e: ["290.1", "299.7"],
throughput: "139.4",
vram: "4.02",
speedup: "9.14× / 4.81× / 3.30×",
quality: "4/4 concepts",
exact: "Exact vs 01b",
},
{
id: "03a",
name: "Flash Attention 2 isolated",
stack: "BF16 · FA2 · dynamic",
input: "448×299",
tokens: "504 vision · 144 input",
status: "regression",
change: "Attention kernel",
ttft: ["106.6", "114.0"],
e2e: ["1033.3", "1054.4"],
throughput: "32.3",
vram: "4.00",
speedup: "6.57× / 1.35× / 0.76×",
quality: "PASS",
exact: "Exact vs 01b",
},
{
id: "03b",
name: "FA2 + compiled decode",
stack: "BF16 · FA2 · static cache",
input: "448×299",
tokens: "hit 96-token cap",
status: "rejected",
change: "Cache + compile",
ttft: ["265.0", "273.8"],
e2e: ["3028.0", "3042.9"],
throughput: "34.4",
vram: "4.02",
speedup: "2.64× / 0.46× / 0.81×",
quality: "FAIL",
exact: "Corrupt repeat",
},
{
id: "04",
name: "Scoped TF32",
stack: "BF16 · SDPA · static · TF32",
input: "448×299",
tokens: "504 vision · 144 input",
status: "measured",
change: "FP32 matmul policy",
ttft: ["74.3", "79.5"],
e2e: ["274.5", "289.6"],
throughput: "149.2",
vram: "4.03",
speedup: "9.43× / 5.08× / 3.53×",
quality: "4/4 concepts",
exact: "Exact vs 01b",
},
{
id: "05a",
name: "SGLang 0.5.9 native",
stack: "FlashInfer · SDPA vision",
input: "448×299",
tokens: "2 output tokens",
status: "rejected",
change: "Serving runtime",
ttft: ["38.3", "42.9"],
e2e: ["43.3", "48.0"],
throughput: "393.9",
vram: null,
speedup: "Invalid — gate failed",
quality: "0/4 concepts",
exact: "Output was ```",
},
{
id: "05b",
name: "SGLang 0.5.10 TF backend",
stack: "FlashInfer · Transformers VLM",
input: "448×299",
tokens: "144 input · 31 output",
status: "measured",
change: "Version + model impl",
ttft: ["75.6", "79.2"],
e2e: ["254.2", "257.5"],
throughput: "173.5",
vram: null,
speedup: "9.26× / 5.49× / 4.10×",
quality: "4/4 concepts",
exact: "Exact vs 01b",
},
{
id: "05c",
name: "SGLang 0.5.10 native",
stack: "FlashInfer · SDPA vision",
input: "448×299",
tokens: "144 input · 40 output",
status: "measured",
change: "Native model impl",
ttft: ["35.2", "38.3"],
e2e: ["240.6", "243.6"],
throughput: "194.7",
vram: null,
speedup: "19.88× / 5.80× / 4.60×",
quality: "4/4 concepts",
exact: "Semantic",
},
{
id: "05d",
name: "Triton vision attention",
stack: "FlashInfer · Triton vision",
input: "448×299",
tokens: "144 input · 31 output",
status: "measured",
change: "Vision attention only",
ttft: ["35.5", "37.9"],
e2e: ["193.6", "195.3"],
throughput: "196.4",
vram: null,
speedup: "19.74× / 7.21× / 4.64×",
quality: "4/4 concepts",
exact: "Exact vs 01b",
},
{
id: "05e",
name: "Compiled SGLang decode",
stack: "FlashInfer · Triton · compile",
input: "448×299",
tokens: "144 input · 31 output",
status: "measured",
change: "Torch compile only",
ttft: ["37.5", "41.0"],
e2e: ["190.5", "194.7"],
throughput: "202.6",
vram: null,
speedup: "18.69× / 7.33× / 4.79×",
quality: "4/4 concepts",
exact: "Exact vs 01b",
},
{
id: "06",
name: "TensorRT vision engine",
stack: "TRT 10.13 · BF16 · static",
input: "448×299",
tokens: "504 vision · 144 input · 31 output",
status: "measured",
change: "Vision tower only",
ttft: ["61.4", "62.4"],
e2e: ["273.4", "274.0"],
throughput: "142.1",
vram: "4.02",
speedup: "11.41× / 5.10× / 3.36×",
quality: "4/4 concepts",
exact: "Exact vs Torch 2.9",
},
{
id: "07",
name: "TensorRT + SGLang bridge",
stack: "TRT vision · SGLang decode",
input: "448×299",
tokens: "144 input · 40 output",
status: "regression",
change: "Runtime composition",
ttft: ["34.9", "37.8"],
e2e: ["250.7", "366.1"],
throughput: "176.1",
vram: null,
speedup: "20.09× / 5.57× / 4.16×",
quality: "4/4 concepts",
exact: "Semantic · exact fail",
},
];
const qualityTasks = [
["Caption facts", "concept groups + aliases", "4 / 4 in every run"],
["Resize fidelity", "task rubric", "pass · wording changed"],
["Compiler fidelity", "SHA-256 output", "exact vs iteration 01b"],
["Repeatability", "within variant", "10 / 10 identical"],
];
export default function Home() {
return (
<main>
<header className="topbar">
<a className="brand" href="#top" aria-label="VLM Speed Lab home">
<span className="brand-mark">VL</span>
<span>VLM Speed Lab</span>
</a>
<nav aria-label="Primary navigation">
<a href="#iterations">Iterations</a>
<a href="#quality">Quality gate</a>
<a href="#protocol">Protocol</a>
</nav>
<a className="repo-link" href="https://github.com/gokayfem/ComfyUI_VLM_nodes">
View repository ↗
</a>
</header>
<section className="hero" id="top">
<div className="hero-copy">
<div className="eyebrow"><span className="live-dot" /> Experiment 001 · Qwen3-VL 2B Instruct</div>
<h1>Make it faster.<br /><em>Prove</em> it stayed good.</h1>
<p className="lede">
One model. One frozen test set. One change per iteration. Every speed claim ships with its output, configuration, and quality score.
</p>
<div className="hero-actions">
<a className="primary-button" href="#iterations">Explore the iterations <span>↓</span></a>
<span className="artifact-note">No synthetic leaderboard numbers</span>
</div>
</div>
<div className="hero-metric" aria-label="Measured end-to-end speedup">
<div className="metric-topline"><span>Measured now</span><span className="verified">● VERIFIED</span></div>
<div className="big-number">7.33<span>×</span></div>
<div className="metric-label">faster end to end</div>
<div className="work-bars" aria-hidden="true">
<div className="work-row"><span>Before</span><i className="bar before" /><b>1395.7</b></div>
<div className="work-row"><span>After</span><i className="bar after" /><b>190.5</b></div>
</div>
<p>Milliseconds p50 · 10 measured runs · output throughput 42.3 → 202.6 tok/s</p>
<div className="honesty-strip">RTX 3090 · batch 1 · task rubric passed · raw samples attached</div>
</div>
</section>
<section className="manifesto-band" aria-label="Benchmark principles">
<span>01 / Same checkpoint</span>
<span>02 / Same media</span>
<span>03 / Same decode</span>
<span>04 / Quality gated</span>
<span>05 / Raw artifacts</span>
</section>
<section className="section iterations-section" id="iterations">
<div className="section-heading">
<div>
<div className="eyebrow">THE OPTIMIZATION LOG</div>
<h2>Every millisecond has a paper trail.</h2>
</div>
<p>Primary numbers are p50; the smaller number is p95. Every row keeps input work, memory, speedup, and quality evidence in view.</p>
</div>
<div className="run-context" aria-label="Benchmark run context">
<span><b>Model</b> Qwen3-VL 2B Instruct</span>
<span><b>Mode</b> Single request</span>
<span><b>Sample</b> 10 measured / variant</span>
<span><b>Warmup</b> 2–6 local / 3 server</span>
<span><b>Runtime</b> Torch 2.8/2.9 · SGLang 0.5.10 · TRT 10.13</span>
</div>
<div className="comparison-table-wrap">
<table className="comparison-table">
<thead>
<tr>
<th scope="col">#</th>
<th scope="col">Variant</th>
<th scope="col">Input work</th>
<th scope="col">One change</th>
<th scope="col">TTFT<br /><span>p50 / p95 ms</span></th>
<th scope="col">E2E<br /><span>p50 / p95 ms</span></th>
<th scope="col">Output<br /><span>tok/s</span></th>
<th scope="col">Peak<br /><span>VRAM GiB</span></th>
<th scope="col">Speedup<br /><span>TTFT / E2E / tok/s</span></th>
<th scope="col">Quality</th>
<th scope="col">Output fidelity</th>
<th scope="col">Status</th>
</tr>
</thead>
<tbody>
{iterations.map((item) => (
<tr className={item.status} key={item.id}>
<td className="row-id">{item.id}</td>
<th scope="row" className="variant-cell"><strong>{item.name}</strong><span>{item.stack}</span></th>
<td className="input-cell"><strong>{item.input}</strong><span>{item.tokens}</span></td>
<td>{item.change}</td>
<td className="metric-cell">{item.ttft ? <><strong>{item.ttft[0]}</strong><span>{item.ttft[1]}</span></> : "—"}</td>
<td className="metric-cell">{item.e2e ? <><strong>{item.e2e[0]}</strong><span>{item.e2e[1]}</span></> : "—"}</td>
<td className="metric-cell">{item.throughput ?? "—"}</td>
<td className="metric-cell">{item.vram ?? "—"}</td>
<td className={item.status === "measured" ? "speedup-cell" : item.status === "planned" ? "muted-cell" : "regression-cell"}>{item.speedup}</td>
<td className={item.status === "rejected" ? "quality-fail" : item.status === "planned" ? "muted-cell" : "quality-ok"}>{item.quality}</td>
<td className={item.status === "rejected" ? "quality-fail" : item.status === "planned" ? "muted-cell" : "fidelity-cell"}>{item.exact}</td>
<td><span className={`status status-${item.status}`}>{item.status}</span></td>
</tr>
))}
</tbody>
</table>
</div>
<div className="table-notes">
<span><b>—</b> Not measured; never estimated</span>
<span><b>Semantic</b> Required facts pass; wording changed</span>
<span><b>Exact</b> SHA-256-identical generated text</span>
<span><b>VRAM —</b> Server peak not yet instrumented</span>
<span><b>Load</b> 88.351s → 6.858s warm cache</span>
<span><b>TRT</b> 1 engine · 0 fallback · 98.070s compile</span>
</div>
</section>
<section className="quality-section" id="quality">
<div className="quality-intro">
<div className="eyebrow light">QUALITY IS A HARD CONSTRAINT</div>
<h2>Fast and wrong<br />doesn’t ship.</h2>
<p>A speedup is promoted only after it clears its declared task gate. Semantic preservation and exact bytes are reported separately.</p>
<div className="gate-formula"><span>promotion rule</span><code>speed ↑ &amp;&amp; quality ≥ tolerance</code></div>
</div>
<div className="quality-table" role="table" aria-label="Quality thresholds">
<div className="quality-row header" role="row"><span>Capability</span><span>Primary score</span><span>Pass threshold</span></div>
{qualityTasks.map(([task, metric, threshold]) => (
<div className="quality-row" role="row" key={task}><strong>{task}</strong><span>{metric}</span><b>{threshold}</b></div>
))}
<div className="quality-proof">
<span className="proof-icon">✓</span>
<div><strong>Outputs stay attached</strong><p>Prompts, model text, boxes, masks, tracks, timing traces, and environment metadata live beside each result.</p></div>
</div>
</div>
</section>
<section className="section protocol-section" id="protocol">
<div className="section-heading protocol-heading">
<div><div className="eyebrow">REPRODUCIBLE BY DEFAULT</div><h2>The benchmark contract.</h2></div>
<div className="commit-chip">artifact <code>TF5 · RTX3090 · B1</code></div>
</div>
<div className="protocol-grid">
<article><span>1</span><h3>Freeze</h3><p>Checkpoint revision, media hashes, prompts, seed, precision, and generation parameters.</p></article>
<article><span>2</span><h3>Warm</h3><p>Cold start is recorded once. Warmups are declared and excluded from steady-state percentiles.</p></article>
<article><span>3</span><h3>Measure</h3><p>TTFT, inter-token latency, output tokens/sec, end-to-end time, peak VRAM, and concurrency.</p></article>
<article><span>4</span><h3>Gate</h3><p>Compare outputs to the baseline and ground truth. Publish pass, regression, or inconclusive.</p></article>
</div>
<div className="metric-strip">
<div><small>Latency</small><strong>p50 / p95 / p99</strong></div>
<div><small>Throughput</small><strong>output tok/s</strong></div>
<div><small>Responsiveness</small><strong>TTFT + ITL</strong></div>
<div><small>Efficiency</small><strong>GB VRAM / request</strong></div>
<div><small>Quality</small><strong>task-specific score</strong></div>
</div>
</section>
<section className="next-run">
<div><span className="eyebrow light">NEXT ON THE RIG</span><h2>Recover exact output.</h2></div>
<div className="next-run-copy"><p>The bridge cut TTFT to 34.9 ms, but changed the exact caption and generated 40 tokens, raising end-to-end latency to 250.7 ms. Next: compile SGLang-native vision weights so TensorRT preserves the 31-token output.</p><a href="https://github.com/gokayfem/ComfyUI_VLM_nodes/tree/codex/vlm-benchmark-lab/benchmarks">Open benchmark kit ↗</a></div>
</section>
<footer><div className="brand"><span className="brand-mark">VL</span><span>VLM Speed Lab</span></div><p>Built in public. Measured, not marketed.</p><span>ComfyUI VLM Nodes · 2026</span></footer>
</main>
);
}
@@ -0,0 +1,45 @@
import { access, cp, mkdir, rm } from "node:fs/promises";
import { resolve } from "node:path";
import type { Plugin } from "vite";
async function exists(path: string): Promise<boolean> {
try {
await access(path);
return true;
} catch (error) {
if ((error as NodeJS.ErrnoException).code === "ENOENT") {
return false;
}
throw error;
}
}
// Packages Sites metadata and migrations after Vite finishes compiling.
export function sites(): Plugin {
let root = process.cwd();
return {
name: "sites",
apply: "build",
configResolved(config) {
root = config.root;
},
async closeBundle() {
const outputDirectory = resolve(root, "dist", ".openai");
const hostingConfig = resolve(root, ".openai", "hosting.json");
const drizzleSource = resolve(root, "drizzle");
await rm(outputDirectory, { recursive: true, force: true });
await mkdir(outputDirectory, { recursive: true });
if (await exists(hostingConfig)) {
await cp(hostingConfig, resolve(outputDirectory, "hosting.json"));
}
if (await exists(drizzleSource)) {
await cp(drizzleSource, resolve(outputDirectory, "drizzle"), {
recursive: true,
});
}
},
};
}
+13
View File
@@ -0,0 +1,13 @@
import { env } from "cloudflare:workers";
import { drizzle } from "drizzle-orm/d1";
import * as schema from "./schema";
export function getDb() {
if (!env.DB) {
throw new Error(
"Cloudflare D1 binding `DB` is unavailable. Set the `d1` field in .openai/hosting.json to `DB` or let your control plane inject the real binding values before using the database."
);
}
return drizzle(env.DB, { schema });
}
+4
View File
@@ -0,0 +1,4 @@
// Intentionally empty by default.
// Add Drizzle tables here when the site actually needs a database.
// See examples/d1/db/schema.ts for an opt-in example.
export {};
+7
View File
@@ -0,0 +1,7 @@
import { defineConfig } from "drizzle-kit";
export default defineConfig({
out: "./drizzle",
schema: "./db/schema.ts",
dialect: "sqlite",
});
@@ -0,0 +1,5 @@
{
"version": "7",
"dialect": "sqlite",
"entries": []
}
+41
View File
@@ -0,0 +1,41 @@
import { defineConfig, globalIgnores } from "eslint/config";
import eslint from "@eslint/js";
import next from "@next/eslint-plugin-next";
import jsxA11y from "eslint-plugin-jsx-a11y";
import react from "eslint-plugin-react";
import reactHooks from "eslint-plugin-react-hooks";
import globals from "globals";
import tseslint from "typescript-eslint";
const eslintConfig = defineConfig([
globalIgnores([
".next/**",
"dist/**",
"out/**",
"build/**",
"next-env.d.ts",
]),
eslint.configs.recommended,
...tseslint.configs.recommended,
react.configs.flat.recommended,
react.configs.flat["jsx-runtime"],
reactHooks.configs.flat["recommended-latest"],
jsxA11y.flatConfigs.recommended,
next.configs["core-web-vitals"],
{
languageOptions: {
globals: {
...globals.browser,
...globals.node,
...globals.serviceworker,
},
},
settings: {
react: {
version: "detect",
},
},
},
]);
export default eslintConfig;
@@ -0,0 +1,58 @@
import { desc } from "drizzle-orm";
import { getDb } from "../../../../../db";
import { notes } from "../../../db/schema";
function toRouteErrorMessage(error: unknown) {
const message = error instanceof Error ? error.message : "Unexpected error";
const detail =
error instanceof Error && error.cause instanceof Error ? error.cause.message : "";
const combined = `${message}\n${detail}`;
if (combined.includes("no such table") || combined.includes('from "notes"')) {
return "The notes table is unavailable. Generate the migration locally with `npm run db:generate`, then deploy so the platform can apply the generated SQL to the real D1 database.";
}
return message;
}
export async function GET() {
try {
const db = getDb();
const rows = await db
.select()
.from(notes)
.orderBy(desc(notes.createdAt), desc(notes.id))
.limit(20);
return Response.json({ notes: rows });
} catch (error) {
return Response.json(
{ error: toRouteErrorMessage(error) },
{ status: 500 }
);
}
}
export async function POST(request: Request) {
try {
const payload = (await request.json()) as {
title?: string;
content?: string;
};
const title = payload.title?.trim() ?? "";
const content = payload.content?.trim() ?? "";
if (!title) {
return Response.json({ error: "title is required" }, { status: 400 });
}
const db = getDb();
const [note] = await db.insert(notes).values({ title, content }).returning();
return Response.json({ note }, { status: 201 });
} catch (error) {
return Response.json(
{ error: toRouteErrorMessage(error) },
{ status: 500 }
);
}
}
+9
View File
@@ -0,0 +1,9 @@
import { sql } from "drizzle-orm";
import { integer, sqliteTable, text } from "drizzle-orm/sqlite-core";
export const notes = sqliteTable("notes", {
id: integer("id").primaryKey({ autoIncrement: true }),
title: text("title").notNull(),
content: text("content").notNull().default(""),
createdAt: text("created_at").notNull().default(sql`CURRENT_TIMESTAMP`),
});
+5
View File
@@ -0,0 +1,5 @@
import "vinext/types";
import "./.next/types/routes.d.ts";
// NOTE: This file should not be edited
// see https://nextjs.org/docs/app/api-reference/config/typescript for more information.
+7
View File
@@ -0,0 +1,7 @@
import type { NextConfig } from "next";
const nextConfig: NextConfig = {
/* config options here */
};
export default nextConfig;
+10269
View File
File diff suppressed because it is too large Load Diff
+46
View File
@@ -0,0 +1,46 @@
{
"name": "site-creator-vinext-starter",
"version": "0.1.0",
"private": true,
"engines": {
"node": ">=22.13.0"
},
"scripts": {
"dev": "WRANGLER_LOG_PATH=.wrangler/wrangler.log vinext dev",
"build": "WRANGLER_LOG_PATH=.wrangler/wrangler.log vinext build",
"start": "WRANGLER_LOG_PATH=.wrangler/wrangler.log vinext start",
"test": "npm run build && node --test tests/rendered-html.test.mjs",
"lint": "eslint . --ignore-pattern dist --ignore-pattern .next",
"db:generate": "drizzle-kit generate"
},
"dependencies": {
"drizzle-orm": "0.45.2",
"react": "19.2.6",
"react-dom": "19.2.6"
},
"devDependencies": {
"@cloudflare/vite-plugin": "1.37.1",
"@eslint/js": "9.39.4",
"@next/eslint-plugin-next": "16.2.6",
"@tailwindcss/postcss": "4.2.1",
"@types/node": "22.19.19",
"@types/react": "19.2.14",
"@types/react-dom": "19.2.3",
"@vitejs/plugin-react": "6.0.2",
"@vitejs/plugin-rsc": "0.5.26",
"drizzle-kit": "0.31.10",
"eslint": "9.39.4",
"eslint-plugin-jsx-a11y": "6.10.2",
"eslint-plugin-react": "7.37.5",
"eslint-plugin-react-hooks": "7.1.1",
"globals": "16.4.0",
"react-server-dom-webpack": "19.2.6",
"tailwindcss": "4.2.1",
"typescript": "5.9.3",
"typescript-eslint": "8.59.3",
"vinext": "1.0.0-beta.2",
"vite": "8.0.13",
"wrangler": "4.92.0"
},
"type": "module"
}
+7
View File
@@ -0,0 +1,7 @@
const config = {
plugins: {
"@tailwindcss/postcss": {},
},
};
export default config;
+6
View File
@@ -0,0 +1,6 @@
<svg width="24" height="24" viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M22 19.2727C22 20.779 20.779 22 19.2727 22H14.7273C13.221 22 12 20.779 12 19.2727V12H19.2727C20.779 12 22 13.221 22 14.7273V19.2727Z" fill="#68C4FF"/>
<path d="M20 2C21.1046 2 22 2.89543 22 4V7C22 8.10457 21.1046 9 20 9H17C15.8954 9 15 8.10457 15 7V4C15 2.89543 15.8954 2 17 2H20Z" fill="#0C79D8"/>
<path d="M7 15C8.10457 15 9 15.8954 9 17V20C9 21.1046 8.10457 22 7 22H4C2.89543 22 2 21.1046 2 20V17C2 15.8954 2.89543 15 4 15H7Z" fill="#0C79D8"/>
<path d="M12 12H4.72727C3.22104 12 2 10.779 2 9.27273V4.72727C2 3.22104 3.22104 2 4.72727 2H9.27273C10.779 2 12 3.22104 12 4.72727V12Z" fill="#2E9EFF"/>
</svg>

After

Width:  |  Height:  |  Size: 712 B

+1
View File
@@ -0,0 +1 @@
<svg fill="none" viewBox="0 0 16 16" xmlns="http://www.w3.org/2000/svg"><path d="M14.5 13.5V5.41a1 1 0 0 0-.3-.7L9.8.29A1 1 0 0 0 9.08 0H1.5v13.5A2.5 2.5 0 0 0 4 16h8a2.5 2.5 0 0 0 2.5-2.5m-1.5 0v-7H8v-5H3v12a1 1 0 0 0 1 1h8a1 1 0 0 0 1-1M9.5 5V2.12L12.38 5zM5.13 5h-.62v1.25h2.12V5zm-.62 3h7.12v1.25H4.5zm.62 3h-.62v1.25h7.12V11z" clip-rule="evenodd" fill="#666" fill-rule="evenodd"/></svg>

After

Width:  |  Height:  |  Size: 392 B

+1
View File
@@ -0,0 +1 @@
<svg fill="none" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 16 16"><g clip-path="url(#a)"><path fill-rule="evenodd" clip-rule="evenodd" d="M10.27 14.1a6.5 6.5 0 0 0 3.67-3.45q-1.24.21-2.7.34-.31 1.83-.97 3.1M8 16A8 8 0 1 0 8 0a8 8 0 0 0 0 16m.48-1.52a7 7 0 0 1-.96 0H7.5a4 4 0 0 1-.84-1.32q-.38-.89-.63-2.08a40 40 0 0 0 3.92 0q-.25 1.2-.63 2.08a4 4 0 0 1-.84 1.31zm2.94-4.76q1.66-.15 2.95-.43a7 7 0 0 0 0-2.58q-1.3-.27-2.95-.43a18 18 0 0 1 0 3.44m-1.27-3.54a17 17 0 0 1 0 3.64 39 39 0 0 1-4.3 0 17 17 0 0 1 0-3.64 39 39 0 0 1 4.3 0m1.1-1.17q1.45.13 2.69.34a6.5 6.5 0 0 0-3.67-3.44q.65 1.26.98 3.1M8.48 1.5l.01.02q.41.37.84 1.31.38.89.63 2.08a40 40 0 0 0-3.92 0q.25-1.2.63-2.08a4 4 0 0 1 .85-1.32 7 7 0 0 1 .96 0m-2.75.4a6.5 6.5 0 0 0-3.67 3.44 29 29 0 0 1 2.7-.34q.31-1.83.97-3.1M4.58 6.28q-1.66.16-2.95.43a7 7 0 0 0 0 2.58q1.3.27 2.95.43a18 18 0 0 1 0-3.44m.17 4.71q-1.45-.12-2.69-.34a6.5 6.5 0 0 0 3.67 3.44q-.65-1.27-.98-3.1" fill="#666"/></g><defs><clipPath id="a"><path fill="#fff" d="M0 0h16v16H0z"/></clipPath></defs></svg>

After

Width:  |  Height:  |  Size: 1.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.0 MiB

+1
View File
@@ -0,0 +1 @@
<svg fill="none" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 16 16"><path fill-rule="evenodd" clip-rule="evenodd" d="M1.5 2.5h13v10a1 1 0 0 1-1 1h-11a1 1 0 0 1-1-1zM0 1h16v11.5a2.5 2.5 0 0 1-2.5 2.5h-11A2.5 2.5 0 0 1 0 12.5zm3.75 4.5a.75.75 0 1 0 0-1.5.75.75 0 0 0 0 1.5M7 4.75a.75.75 0 1 1-1.5 0 .75.75 0 0 1 1.5 0m1.75.75a.75.75 0 1 0 0-1.5.75.75 0 0 0 0 1.5" fill="#666"/></svg>

After

Width:  |  Height:  |  Size: 386 B

@@ -0,0 +1,59 @@
import assert from "node:assert/strict";
import { readFile } from "node:fs/promises";
import test from "node:test";
async function render() {
const workerUrl = new URL("../dist/server/index.js", import.meta.url);
workerUrl.searchParams.set("test", `${process.pid}-${Date.now()}`);
const { default: worker } = await import(workerUrl.href);
return worker.fetch(
new Request("http://localhost/", {
headers: { accept: "text/html" },
}),
{ ASSETS: { fetch: async () => new Response("Not found", { status: 404 }) } },
{ waitUntil() {}, passThroughOnException() {} },
);
}
test("server-renders the measured optimization matrix", async () => {
const response = await render();
assert.equal(response.status, 200);
assert.match(response.headers.get("content-type") ?? "", /^text\/html\b/i);
const html = await response.text();
assert.match(html, /VLM Speed Lab/);
assert.match(html, /7\.33/);
assert.match(html, /202\.6/);
assert.match(html, /Source-resolution control/);
assert.match(html, /Compiled execution/);
assert.match(html, /Exact vs 01b/);
assert.match(html, /Corrupt repeat/);
assert.match(html, /SGLang 0\.5\.9 native/);
assert.match(html, /Triton vision attention/);
assert.match(html, /TensorRT \+ SGLang bridge/);
assert.match(html, /Semantic · exact fail/);
assert.match(html, /Invalid — gate failed/);
assert.match(html, /88\.351s → 6\.858s/);
assert.doesNotMatch(html, /GPU run pending|end-to-end run pending/);
assert.doesNotMatch(html, /codex-preview|react-loading-skeleton/);
});
test("keeps measured regressions visually honest", async () => {
const [page, css] = await Promise.all([
readFile(new URL("../app/page.tsx", import.meta.url), "utf8"),
readFile(new URL("../app/globals.css", import.meta.url), "utf8"),
]);
assert.match(page, /p50 \/ p95 ms/);
assert.match(page, /Not measured/);
assert.match(page, /status: "regression"/);
assert.match(page, /status: "rejected"/);
assert.match(page, /ttft: \["34\.9", "37\.8"\]/);
assert.match(page, /4\/4 concepts/);
assert.match(page, /Semantic preservation and exact bytes/);
assert.match(css, /\.comparison-table-wrap \{ overflow-x:auto/);
assert.match(css, /\.status-measured/);
assert.match(css, /\.status-rejected/);
assert.match(css, /\.status-planned/);
});
+29
View File
@@ -0,0 +1,29 @@
{
"compilerOptions": {
"target": "ES2017",
"lib": ["dom", "dom.iterable", "esnext"],
"allowJs": true,
"skipLibCheck": true,
"strict": true,
"noEmit": true,
"esModuleInterop": true,
"module": "esnext",
"moduleResolution": "bundler",
"resolveJsonModule": true,
"isolatedModules": true,
"jsx": "react-jsx",
"incremental": true,
"paths": {
"@/*": ["./*"]
}
},
"include": [
"next-env.d.ts",
"**/*.ts",
"**/*.tsx",
".next/types/**/*.ts",
".next/dev/types/**/*.ts",
"**/*.mts"
],
"exclude": ["node_modules"]
}
+59
View File
@@ -0,0 +1,59 @@
import vinext from "vinext";
import { defineConfig } from "vite";
import hostingConfig from "./.openai/hosting.json";
import { sites } from "./build/sites-vite-plugin";
const SITE_CREATOR_PLACEHOLDER_DATABASE_ID =
"00000000-0000-4000-8000-000000000000";
const { d1, r2 } = hostingConfig;
// macOS Seatbelt blocks FSEvents, so Codex previews need polling for HMR.
const isCodexSeatbeltSandbox = process.env.CODEX_SANDBOX === "seatbelt";
const localBindingConfig = {
main: "./worker/index.ts",
compatibility_flags: ["nodejs_compat"],
d1_databases: d1
? [
{
binding: d1,
database_name: "site-creator-d1",
database_id: SITE_CREATOR_PLACEHOLDER_DATABASE_ID,
},
]
: [],
r2_buckets: r2
? [
{
binding: r2,
bucket_name: "site-creator-r2",
},
]
: [],
};
export default defineConfig(async () => {
// Keep Wrangler and Miniflare state project-local. These are non-secret tool
// settings; application environment belongs in ignored `.env*` files.
process.env.WRANGLER_WRITE_LOGS ??= "false";
process.env.WRANGLER_LOG_PATH ??= ".wrangler/logs";
process.env.MINIFLARE_REGISTRY_PATH ??= ".wrangler/registry";
// Wrangler snapshots its log path while the Cloudflare plugin is imported.
const { cloudflare } = await import("@cloudflare/vite-plugin");
return {
server: isCodexSeatbeltSandbox
? { watch: { useFsEvents: false, usePolling: true } }
: undefined,
plugins: [
vinext(),
sites(),
cloudflare({
viteEnvironment: { name: "rsc", childEnvironments: ["ssr"] },
config: localBindingConfig,
}),
],
};
});
+47
View File
@@ -0,0 +1,47 @@
/** Cloudflare Worker entry point for the vinext-starter template. */
import { handleImageOptimization, DEFAULT_DEVICE_SIZES, DEFAULT_IMAGE_SIZES } from "vinext/server/image-optimization";
import handler from "vinext/server/app-router-entry";
interface Env {
ASSETS: Fetcher;
DB: D1Database;
IMAGES: {
input(stream: ReadableStream): {
transform(options: Record<string, unknown>): {
output(options: { format: string; quality: number }): Promise<{ response(): Response }>;
};
};
};
}
interface ExecutionContext {
waitUntil(promise: Promise<unknown>): void;
passThroughOnException(): void;
}
// Image security config. SVG sources with .svg extension auto-skip the
// optimization endpoint on the client side (served directly, no proxy).
// To route SVGs through the optimizer (with security headers), set
// dangerouslyAllowSVG: true in next.config.js and uncomment below:
// const imageConfig: ImageConfig = { dangerouslyAllowSVG: true };
const worker = {
async fetch(request: Request, env: Env, ctx: ExecutionContext): Promise<Response> {
const url = new URL(request.url);
if (url.pathname === "/_vinext/image") {
const allowedWidths = [...DEFAULT_DEVICE_SIZES, ...DEFAULT_IMAGE_SIZES];
return handleImageOptimization(request, {
fetchAsset: (path) => env.ASSETS.fetch(new Request(new URL(path, request.url))),
transformImage: async (body, { width, format, quality }) => {
const result = await env.IMAGES.input(body).transform(width > 0 ? { width } : {}).output({ format, quality });
return result.response();
},
}, allowedWidths);
}
return handler.fetch(request, env, ctx);
},
};
export default worker;
+34
View File
@@ -0,0 +1,34 @@
{
"name": "qwen3-vl-2b-control-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"max_tokens": 128,
"temperature": 0.0,
"quality_tolerance": 0.98,
"cases": [
{
"id": "caption-001",
"task": "caption",
"image": "media/caption-001.jpg",
"prompt": "Describe the image in one precise sentence.",
"evaluator": "keywords",
"expected": ["replace", "with", "ground-truth", "keywords"]
},
{
"id": "ocr-001",
"task": "ocr",
"image": "media/ocr-001.png",
"prompt": "Return only the text visible in the image.",
"evaluator": "exact",
"expected": "REPLACE WITH GROUND TRUTH"
},
{
"id": "count-001",
"task": "count",
"image": "media/count-001.png",
"prompt": "How many red objects are visible? Return only the integer.",
"evaluator": "number",
"expected": 0
}
]
}
+23
View File
@@ -0,0 +1,23 @@
{
"name": "qwen3-vl-2b-demo-448-v1",
"model": "Qwen/Qwen3-VL-2B-Instruct",
"max_tokens": 96,
"temperature": 0.0,
"quality_tolerance": 1.0,
"cases": [
{
"id": "caption-qwen-demo-001",
"task": "caption",
"image": "media/qwen-demo.jpeg",
"longest_edge": 448,
"prompt": "Describe this image precisely in one sentence.",
"evaluator": "concepts",
"expected": [
["woman"],
["golden retriever"],
["beach"],
["high-five", "high-fiving", "high five", "high fiving"]
]
}
]
}
+310
View File
@@ -0,0 +1,310 @@
"""Quality-gated benchmark for OpenAI-compatible VLM servers.
The runner intentionally depends only on packages already required by this
repository. It is suitable for SGLang and TensorRT-LLM chat endpoints and keeps
the raw evidence required to audit every aggregate number.
"""
from __future__ import annotations
import argparse
import base64
import hashlib
import io
import json
import math
import mimetypes
import platform
import re
import statistics
import subprocess
import time
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
import httpx
from PIL import Image
def normalize_text(value: str) -> str:
return " ".join(re.sub(r"[^\w\s]", " ", value.casefold()).split())
def score_output(output: str, evaluator: str, expected: Any) -> float:
normalized = normalize_text(output)
if evaluator == "exact":
return float(normalized == normalize_text(str(expected)))
if evaluator == "keywords":
terms = [normalize_text(str(term)) for term in expected]
terms = [term for term in terms if term]
return sum(term in normalized for term in terms) / len(terms) if terms else 0.0
if evaluator == "concepts":
concepts = []
for concept in expected:
aliases = concept if isinstance(concept, list) else [concept]
aliases = [normalize_text(str(alias)) for alias in aliases]
aliases = [alias for alias in aliases if alias]
if aliases:
concepts.append(aliases)
return (
sum(any(alias in normalized for alias in aliases) for aliases in concepts)
/ len(concepts)
if concepts
else 0.0
)
if evaluator == "number":
match = re.search(r"-?\d+", output.replace(",", ""))
return float(match is not None and int(match.group()) == int(expected))
raise ValueError(f"Unsupported evaluator: {evaluator!r}")
def percentile(values: list[float], quantile: float) -> float:
if not values:
return math.nan
ordered = sorted(values)
position = (len(ordered) - 1) * quantile
lower = math.floor(position)
upper = math.ceil(position)
if lower == upper:
return ordered[lower]
return ordered[lower] * (upper - position) + ordered[upper] * (position - lower)
def file_data_url(path: Path, longest_edge: int | None) -> tuple[str, str, dict]:
content = path.read_bytes()
source_digest = hashlib.sha256(content).hexdigest()
image = Image.open(io.BytesIO(content)).convert("RGB")
source_size = image.size
if longest_edge is not None and max(image.size) > longest_edge:
scale = longest_edge / max(image.size)
image = image.resize(
(round(image.width * scale), round(image.height * scale)),
Image.Resampling.BOX,
)
buffer = io.BytesIO()
image.save(buffer, format="PNG")
content = buffer.getvalue()
mime = "image/png"
else:
mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
encoded = base64.b64encode(content).decode("ascii")
return (
f"data:{mime};base64,{encoded}",
hashlib.sha256(content).hexdigest(),
{
"source_sha256": source_digest,
"source_width": source_size[0],
"source_height": source_size[1],
"processed_width": image.width,
"processed_height": image.height,
},
)
def git_value(*args: str) -> str | None:
try:
return subprocess.check_output(
["git", *args], text=True, stderr=subprocess.DEVNULL
).strip()
except (OSError, subprocess.CalledProcessError):
return None
def parse_sse_line(line: str) -> dict[str, Any] | None:
if not line.startswith("data:"):
return None
payload = line[5:].strip()
if not payload or payload == "[DONE]":
return None
return json.loads(payload)
def run_request(
client: httpx.Client,
*,
base_url: str,
model: str,
prompt: str,
image_url: str,
max_tokens: int,
temperature: float,
) -> dict[str, Any]:
payload = {
"model": model,
"messages": [
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": image_url}},
{"type": "text", "text": prompt},
],
}
],
"max_tokens": max_tokens,
"temperature": temperature,
"stream": True,
"stream_options": {"include_usage": True},
}
started = time.perf_counter()
first_content_at: float | None = None
pieces: list[str] = []
usage: dict[str, Any] = {}
with client.stream(
"POST", f"{base_url.rstrip('/')}/chat/completions", json=payload
) as response:
response.raise_for_status()
for line in response.iter_lines():
event = parse_sse_line(line)
if event is None:
continue
usage = event.get("usage") or usage
for choice in event.get("choices", []):
content = (choice.get("delta") or {}).get("content")
if content:
if first_content_at is None:
first_content_at = time.perf_counter()
pieces.append(content)
finished = time.perf_counter()
output = "".join(pieces)
completion_tokens = usage.get("completion_tokens")
decode_seconds = finished - (first_content_at or finished)
return {
"output": output,
"latency_ms": round((finished - started) * 1000, 3),
"ttft_ms": round(((first_content_at or finished) - started) * 1000, 3),
"completion_tokens": completion_tokens,
"output_tokens_per_second": (
round(completion_tokens / decode_seconds, 3)
if completion_tokens and decode_seconds > 0
else None
),
"usage": usage,
}
def aggregate(samples: list[dict[str, Any]]) -> dict[str, Any]:
latencies = [float(sample["latency_ms"]) for sample in samples]
ttfts = [float(sample["ttft_ms"]) for sample in samples]
rates = [
float(sample["output_tokens_per_second"])
for sample in samples
if sample.get("output_tokens_per_second") is not None
]
return {
"requests": len(samples),
"latency_ms": {
"p50": round(percentile(latencies, 0.50), 3),
"p95": round(percentile(latencies, 0.95), 3),
"p99": round(percentile(latencies, 0.99), 3),
},
"ttft_ms": {
"p50": round(percentile(ttfts, 0.50), 3),
"p95": round(percentile(ttfts, 0.95), 3),
"p99": round(percentile(ttfts, 0.99), 3),
},
"output_tokens_per_second_mean": round(statistics.fmean(rates), 3) if rates else None,
"quality_mean": round(statistics.fmean(sample["quality"] for sample in samples), 6),
}
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--suite", type=Path, required=True)
parser.add_argument("--base-url", required=True)
parser.add_argument("--backend", choices=("sglang", "tensorrt-llm", "other"), required=True)
parser.add_argument("--label", required=True)
parser.add_argument("--warmups", type=int, default=3)
parser.add_argument("--runs", type=int, default=30)
parser.add_argument("--timeout", type=float, default=180.0)
parser.add_argument("--output-dir", type=Path, default=Path("benchmarks/results"))
args = parser.parse_args()
if args.warmups < 0 or args.runs < 1:
parser.error("--warmups must be non-negative and --runs must be positive")
suite_path = args.suite.resolve()
suite = json.loads(suite_path.read_text(encoding="utf-8"))
cases = suite.get("cases") or []
if not cases:
raise ValueError("The suite must contain at least one case.")
prepared = []
for case in cases:
media_path = (suite_path.parent / case["image"]).resolve()
if not media_path.is_file():
raise FileNotFoundError(f"Missing benchmark media: {media_path}")
data_url, digest, media = file_data_url(
media_path,
int(case["longest_edge"]) if case.get("longest_edge") else None,
)
prepared.append((case, data_url, digest, media))
samples: list[dict[str, Any]] = []
with httpx.Client(timeout=args.timeout) as client:
for index in range(args.warmups + args.runs):
case, data_url, digest, media = prepared[index % len(prepared)]
result = run_request(
client,
base_url=args.base_url,
model=suite["model"],
prompt=case["prompt"],
image_url=data_url,
max_tokens=int(suite.get("max_tokens", 128)),
temperature=float(suite.get("temperature", 0.0)),
)
if index < args.warmups:
continue
result.update(
{
"sample": index - args.warmups,
"case_id": case["id"],
"task": case["task"],
"media_sha256": digest,
"media": media,
"quality": score_output(
result["output"], case["evaluator"], case["expected"]
),
}
)
samples.append(result)
summary = aggregate(samples)
tolerance = float(suite.get("quality_tolerance", 0.98))
artifact = {
"schema": "comfyui-vlm/benchmark-run",
"version": 1,
"created_at": datetime.now(UTC).isoformat(),
"label": args.label,
"backend": args.backend,
"suite": suite["name"],
"model": suite["model"],
"git_commit": git_value("rev-parse", "HEAD"),
"git_dirty": bool(git_value("status", "--porcelain")),
"environment": {
"platform": platform.platform(),
"python": platform.python_version(),
"server_base_url": args.base_url,
},
"settings": {
"warmups": args.warmups,
"runs": args.runs,
"max_tokens": suite.get("max_tokens", 128),
"temperature": suite.get("temperature", 0.0),
"quality_tolerance": tolerance,
},
"summary": summary,
"quality_gate": {
"threshold": tolerance,
"passed": summary["quality_mean"] >= tolerance,
},
"samples": samples,
}
args.output_dir.mkdir(parents=True, exist_ok=True)
timestamp = datetime.now(UTC).strftime("%Y%m%dT%H%M%SZ")
output = args.output_dir / f"{timestamp}-{args.label}.json"
output.write_text(json.dumps(artifact, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
print(output)
if __name__ == "__main__":
main()
+80
View File
@@ -7,8 +7,10 @@ small and large VLM families while keeping downloads and VRAM allocation lazy.
from __future__ import annotations
import os
import threading
from collections.abc import Callable
from contextlib import contextmanager
from dataclasses import dataclass
from typing import Any
@@ -217,6 +219,39 @@ MEMORY_MODES = (
"CPU",
)
ATTENTION_MODES = ("Auto (SDPA)", "Flash Attention 2", "Eager")
GENERATION_CACHE_MODES = (
"Dynamic (compatible)",
"Static compiled (fastest repeated shape)",
)
MATMUL_PRECISION_MODES = (
"Highest (strict)",
"High / TF32 (fast on NVIDIA)",
)
_MATMUL_PRECISION_LOCK = threading.RLock()
def _enable_parallel_weight_loading() -> None:
"""Use Transformers' threaded safetensor loader unless the user opted out."""
os.environ.setdefault("HF_ENABLE_PARALLEL_LOADING", "true")
os.environ.setdefault(
"HF_PARALLEL_LOADING_WORKERS",
str(min(8, os.cpu_count() or 1)),
)
@contextmanager
def _float32_matmul_precision(mode: str):
if mode not in MATMUL_PRECISION_MODES:
raise ValueError(f"Unknown float32 matmul precision mode {mode!r}.")
requested = "high" if mode.startswith("High / TF32") else "highest"
with _MATMUL_PRECISION_LOCK:
previous = torch.get_float32_matmul_precision()
torch.set_float32_matmul_precision(requested)
try:
yield
finally:
torch.set_float32_matmul_precision(previous)
def _progress_text_sender(node_id: str | None) -> Callable[[str], None] | None:
@@ -354,6 +389,10 @@ class ModernVLMPredictor:
elif memory_mode == "CPU":
kwargs["dtype"] = torch.float32
# This only affects checkpoint deserialization. It leaves inference,
# precision, placement, and model outputs unchanged, and respects any
# explicit environment settings supplied by the user.
_enable_parallel_weight_loading()
try:
model = _model_class(transformers).from_pretrained(
model_path, **kwargs
@@ -464,6 +503,8 @@ class ModernVLMPredictor:
video_frames=None,
fps: float = 1.0,
enable_thinking: bool = False,
generation_cache: str = "Dynamic (compatible)",
matmul_precision: str = "Highest (strict)",
stream_callback: Callable[[str], None] | None = None,
video_selection: VideoFrameSelection | None = None,
) -> str:
@@ -581,6 +622,17 @@ class ModernVLMPredictor:
"max_new_tokens": int(max_new_tokens),
"do_sample": float(temperature) > 0,
}
if generation_cache not in GENERATION_CACHE_MODES:
raise ValueError(
f"Unknown generation cache mode {generation_cache!r}."
)
if generation_cache.startswith("Static compiled"):
# A fixed-size cache lets maintained Transformers releases
# compile the token-decoding stage. The first several calls
# pay compilation cost; repeated identical shapes then reuse
# the optimized graph. Keep dynamic cache as the compatibility
# default for one-shot and frequently changing workloads.
generation["cache_implementation"] = "static"
if generation["do_sample"]:
generation.update(
temperature=float(temperature), top_p=float(top_p)
@@ -600,6 +652,7 @@ class ModernVLMPredictor:
def generate_in_background() -> None:
try:
with (
_float32_matmul_precision(matmul_precision),
torch.inference_mode(),
inference_context(device, self.dtype),
):
@@ -645,6 +698,7 @@ class ModernVLMPredictor:
results.append(decoded)
else:
with (
_float32_matmul_precision(matmul_precision),
torch.inference_mode(),
inference_context(device, self.dtype),
):
@@ -713,6 +767,28 @@ class ModernVLM(CachedModelNode):
ATTENTION_MODES,
{"default": "Auto (SDPA)"},
),
"generation_cache": (
GENERATION_CACHE_MODES,
{
"default": "Dynamic (compatible)",
"tooltip": (
"Static compiled is fastest after several warmups "
"when image and output shapes repeat. Its first "
"run can be much slower while kernels compile."
),
},
),
"matmul_precision": (
MATMUL_PRECISION_MODES,
{
"default": "Highest (strict)",
"tooltip": (
"High / TF32 can accelerate Ampere-or-newer NVIDIA "
"GPUs. It is scoped to this generation and restored "
"afterward. Validate output quality for each model."
),
},
),
"enable_thinking": ("BOOLEAN", {"default": False}),
"unload_after": ("BOOLEAN", {"default": False}),
"stream_output": (
@@ -757,6 +833,8 @@ class ModernVLM(CachedModelNode):
video_selection=None,
fps=1.0,
attention_mode="Auto (SDPA)",
generation_cache="Dynamic (compatible)",
matmul_precision="Highest (strict)",
enable_thinking=False,
unload_after=False,
stream_output=True,
@@ -792,6 +870,8 @@ class ModernVLM(CachedModelNode):
fps=fps,
video_selection=video_selection,
enable_thinking=enable_thinking,
generation_cache=generation_cache,
matmul_precision=matmul_precision,
stream_callback=stream_callback,
),
)
+40
View File
@@ -0,0 +1,40 @@
from benchmarks.vlm_bench import aggregate, percentile, score_output
def test_task_specific_quality_scores_are_deterministic():
assert score_output("A red bird on a branch.", "keywords", ["red", "bird", "branch"]) == 1.0
assert (
score_output(
"A woman and dog are high-fiving on a beach.",
"concepts",
[["woman"], ["dog", "retriever"], ["beach"], ["high-five", "high-fiving"]],
)
== 1.0
)
assert (
score_output(
"A woman and dog on a beach.",
"concepts",
[["woman"], ["dog"], ["beach"], ["high-fiving"]],
)
== 0.75
)
assert score_output("Hello, WORLD!", "exact", "hello world") == 1.0
assert score_output("There are 12 objects.", "number", 12) == 1.0
assert score_output("There are 11 objects.", "number", 12) == 0.0
def test_percentile_interpolates_small_samples():
assert percentile([10.0, 20.0], 0.5) == 15.0
def test_aggregate_keeps_latency_ttft_throughput_and_quality_separate():
samples = [
{"latency_ms": 100.0, "ttft_ms": 40.0, "output_tokens_per_second": 20.0, "quality": 1.0},
{"latency_ms": 200.0, "ttft_ms": 60.0, "output_tokens_per_second": 30.0, "quality": 0.5},
]
result = aggregate(samples)
assert result["latency_ms"]["p50"] == 150.0
assert result["ttft_ms"]["p50"] == 50.0
assert result["output_tokens_per_second_mean"] == 25.0
assert result["quality_mean"] == 0.75
+20
View File
@@ -69,6 +69,20 @@ def test_source_has_no_runtime_installer_or_direct_cuda_cache():
assert "pip install" not in source
def test_parallel_weight_loading_respects_user_environment(monkeypatch):
monkeypatch.delenv("HF_ENABLE_PARALLEL_LOADING", raising=False)
monkeypatch.delenv("HF_PARALLEL_LOADING_WORKERS", raising=False)
modern_vlm._enable_parallel_weight_loading()
assert modern_vlm.os.environ["HF_ENABLE_PARALLEL_LOADING"] == "true"
assert int(modern_vlm.os.environ["HF_PARALLEL_LOADING_WORKERS"]) in range(1, 9)
monkeypatch.setenv("HF_ENABLE_PARALLEL_LOADING", "false")
monkeypatch.setenv("HF_PARALLEL_LOADING_WORKERS", "2")
modern_vlm._enable_parallel_weight_loading()
assert modern_vlm.os.environ["HF_ENABLE_PARALLEL_LOADING"] == "false"
assert modern_vlm.os.environ["HF_PARALLEL_LOADING_WORKERS"] == "2"
def test_portable_device_dtype_and_backend_contracts(monkeypatch):
assert torch_dtype("float16", torch.device("cpu")) == torch.float32
assert torch_dtype("float16", torch.device("mps")) == torch.float16
@@ -500,6 +514,8 @@ def test_modern_vlm_streams_cumulative_text_without_changing_final_output(
class FakeModel:
def generate(self, **kwargs):
assert isinstance(kwargs["streamer"], FakeStreamer)
assert kwargs["cache_implementation"] == "static"
assert torch.get_float32_matmul_precision() == "high"
return torch.tensor([[10, 11, 12]], dtype=torch.long)
class FakeProcessor:
@@ -528,6 +544,7 @@ def test_modern_vlm_streams_cumulative_text_without_changing_final_output(
lambda *_args: nullcontext(),
)
partials = []
original_precision = torch.get_float32_matmul_precision()
result = predictor.generate(
torch.zeros((1, 8, 8, 3), dtype=torch.float32),
"Describe it.",
@@ -535,11 +552,14 @@ def test_modern_vlm_streams_cumulative_text_without_changing_final_output(
16,
0.0,
0.9,
generation_cache="Static compiled (fastest repeated shape)",
matmul_precision="High / TF32 (fast on NVIDIA)",
stream_callback=partials.append,
)
assert result == "Hello from the VLM."
assert partials == ["Hello", "Hello from", "Hello from the VLM."]
assert torch.get_float32_matmul_precision() == original_precision
def test_view_text_frontend_rehydrates_and_uses_native_progress_channel():