Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
39b015713f | ||
|
|
da9cbe3da9 | ||
|
|
04bfddb7fa |
@@ -232,13 +232,26 @@ by the harness and is not config-declarable.)
|
||||
|
||||
Recipe fingerprinting, hardware/software profile IDs, exact-identity
|
||||
comparison, and dashboard cohort grouping land with this change: v2 records
|
||||
compare only within their identity cohort, and a record that opens a NEW
|
||||
cohort is marked `baseline_status: "initialized_new_cohort"` (regression
|
||||
gating starts once that cohort accumulates history). Legacy v1 configs still
|
||||
run and are normalized for reporting, but their records skip rolling-baseline
|
||||
comparison entirely (`baseline_status: "skipped_missing_identity"`, never
|
||||
baseline eligible); only static thresholds gate them. Metric-specific
|
||||
threshold policies and promoted baselines remain separate follow-ups.
|
||||
compare only within their identity cohort and are never compared against v1
|
||||
records (or vice versa). Legacy records missing any comparison identity field
|
||||
still run and are normalized for reporting, but they skip rolling-baseline
|
||||
comparison entirely (`baseline_status: "skipped_missing_identity"`, comparator
|
||||
verdict `PASS`, never baseline eligible); only static thresholds gate them.
|
||||
Each comparison reports an explicit `comparator_status` on the normalized
|
||||
record and in the Markdown summary:
|
||||
|
||||
| Status | Meaning | CI |
|
||||
|---|---|---|
|
||||
| `PASS` | Comparable baseline found with no gated regression, or comparison was skipped for a record without the v2 identity block. | passes |
|
||||
| `REGRESSION` | A gated metric regressed past both its percent and absolute floors. | fails |
|
||||
| `CALIBRATION_NEEDED` | No comparable baseline exists. Gating is inactive and the record does **not** seed a baseline — seed new cohorts explicitly via the reseed workflow. The record also carries `baseline_status: "initialized_new_cohort"`. | passes |
|
||||
| `RECIPE_MISMATCH` | Baseline history exists for the same variant/hardware/software cohort but under a different `recipe_fingerprint`. Represent recipe changes as a new `variant_id`, or reseed. | fails |
|
||||
| `INFRA_ERROR` | The comparison itself failed (for example, baseline history could not be loaded). | fails |
|
||||
| `HOST_BELOW_PROFILE` | The harness's ~2s host CPU probe (`host_probe.py`) scored below the healthy-host floor (`PERF_HOST_CPU_MIN_SCORE`, default 0.75): the shared runner was packed, so the measurements are not comparable. Gating is skipped and the record never seeds a baseline. | passes |
|
||||
| `QUALITY_BLOCKED` | Reserved for the promoted-baseline workflow; never emitted by the comparator. | n/a |
|
||||
|
||||
Metric-specific threshold policies and promoted baselines remain separate
|
||||
follow-ups.
|
||||
|
||||
### Raw record (`results/perf_*.json`)
|
||||
|
||||
@@ -367,6 +380,8 @@ result, used as the rolling-baseline source of truth.
|
||||
"build_id": "<buildkite-build-id>",
|
||||
"job_id": "<buildkite-job-id>",
|
||||
"quality_metadata": { "quality_status": "canonical" },
|
||||
"baseline_status": "compared",
|
||||
"comparator_status": "PASS",
|
||||
"success": true
|
||||
}
|
||||
```
|
||||
@@ -487,10 +502,10 @@ When the rolling-baseline phase runs, it emits:
|
||||
2. The pytest test auto-discovers all configs — no test code needed. CI
|
||||
picks it up on the next `/test performance` run.
|
||||
|
||||
3. The first persisted main-branch run with no HF history initializes the
|
||||
baseline (passes automatically). Subsequent runs compare against it. Local
|
||||
and pull-request runs with no HF history also pass, but they do not seed the
|
||||
shared baseline.
|
||||
3. The first run with no HF history reports `CALIBRATION_NEEDED`: it passes,
|
||||
but it does not seed the shared baseline. Seed the cohort explicitly with
|
||||
the `reseed-performance-baseline` skill; subsequent scheduled-main runs
|
||||
then compare against it and keep the rolling baseline advancing.
|
||||
|
||||
4. If the benchmark targets a GPU not currently in `thresholds`, either add
|
||||
that GPU as a key or rely on the `default` block. Note that `default` is
|
||||
@@ -510,8 +525,21 @@ When the rolling-baseline phase runs, it emits:
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**"No baseline for ... Initializing"** — first run for this comparison cohort.
|
||||
Run will pass and (if persisting) seed the first record.
|
||||
**`CALIBRATION_NEEDED: no comparable baseline`** — first run for this cohort.
|
||||
The run passes but does not seed a baseline; seed it explicitly with the
|
||||
`reseed-performance-baseline` skill.
|
||||
|
||||
**`RECIPE_MISMATCH`** — the benchmark recipe changed without a new
|
||||
`variant_id`. Either bump `variant_id` to open a new cohort, or reseed the
|
||||
baseline if the existing variant should adopt the new recipe.
|
||||
|
||||
**`HOST_BELOW_PROFILE: host below profile — measurements not comparable`** —
|
||||
the shared Modal host was packed during the run (host CPU contention; GPU
|
||||
telemetry typically still looks healthy). Not a property of the PR: the gate
|
||||
is skipped and CI passes. Retry the lane if you need a comparable measurement.
|
||||
If every host suddenly scores low after an image/python/torch bump, the probe
|
||||
reference is stale — recalibrate `PERF_HOST_CPU_ST_REF_KOPS` (see
|
||||
`host_probe.py`, which documents the calibration procedure and sampled data).
|
||||
|
||||
**Persistent failure right after a torch / kernel / image upgrade** —
|
||||
genuine regression *or* baseline drift. Compare the failing normalized record
|
||||
|
||||
@@ -89,6 +89,34 @@ COMPARISON_IDENTITY_KEYS = (
|
||||
"hardware_profile_id",
|
||||
"software_profile_id",
|
||||
)
|
||||
# Comparator verdicts (issue #1532), stored as record["comparator_status"].
|
||||
STATUS_PASS = "PASS"
|
||||
STATUS_REGRESSION = "REGRESSION"
|
||||
STATUS_CALIBRATION_NEEDED = "CALIBRATION_NEEDED"
|
||||
STATUS_RECIPE_MISMATCH = "RECIPE_MISMATCH"
|
||||
STATUS_INFRA_ERROR = "INFRA_ERROR"
|
||||
STATUS_HOST_BELOW_PROFILE = "HOST_BELOW_PROFILE"
|
||||
# Reserved for the promoted-baseline workflow; never emitted by this comparator.
|
||||
STATUS_QUALITY_BLOCKED = "QUALITY_BLOCKED"
|
||||
|
||||
# Floor for host_cpu_score (host_probe.py: ~1.0 on a healthy perf-lane host).
|
||||
# Calibration basis (see host_probe.py for the sampled data): healthy lane
|
||||
# hosts scored 0.97-1.17; the packed-host signature (review r30; TE 3.595s vs
|
||||
# the 2.03s healthy cluster, DiT +58-80%) inflates CPU-bound wall time by
|
||||
# 1.58-1.77x, i.e. a packed host projects to 0.54-0.73. 0.75 sits above every
|
||||
# packed projection and below every healthy sample.
|
||||
DEFAULT_HOST_CPU_MIN_SCORE = 0.75
|
||||
|
||||
|
||||
def _host_cpu_min_score() -> float:
|
||||
raw = os.environ.get("PERF_HOST_CPU_MIN_SCORE", "").strip()
|
||||
if raw:
|
||||
try:
|
||||
return float(raw)
|
||||
except ValueError:
|
||||
print(f"Invalid PERF_HOST_CPU_MIN_SCORE={raw!r}; "
|
||||
f"using {DEFAULT_HOST_CPU_MIN_SCORE}")
|
||||
return DEFAULT_HOST_CPU_MIN_SCORE
|
||||
|
||||
|
||||
def _should_persist_tracking() -> bool:
|
||||
@@ -260,6 +288,13 @@ def normalize_performance_result(result: dict[str, Any]) -> dict[str, Any]:
|
||||
"success": True,
|
||||
**_record_metadata(_detect_run_source(), result),
|
||||
}
|
||||
# Host CPU probe fields (host_probe.py); absent on records measured
|
||||
# before the probe existed or when the probe failed.
|
||||
for key in ("host_cpu_score", "host_cpu_single_thread_kops",
|
||||
"host_cpu_multi_thread_gflops"):
|
||||
value = safe_float(result.get(key))
|
||||
if value is not None:
|
||||
record[key] = value
|
||||
record.update(_identity_metadata(result))
|
||||
return record
|
||||
|
||||
@@ -343,6 +378,110 @@ def _check_regressions(
|
||||
return failures
|
||||
|
||||
|
||||
def _recipe_mismatch_fingerprints(
|
||||
record: dict[str, Any],
|
||||
identity_filters: dict[str, str],
|
||||
) -> list[str]:
|
||||
"""Return baseline recipe fingerprints for the same variant cohort.
|
||||
|
||||
Non-empty means the same (workload, variant, version, hardware, software)
|
||||
cohort has baseline history under a DIFFERENT recipe fingerprint: the
|
||||
recipe changed without being represented as a new variant.
|
||||
"""
|
||||
if not identity_filters:
|
||||
return []
|
||||
variant_filters = {key: value for key, value in identity_filters.items() if key != "recipe_fingerprint"}
|
||||
same_variant = load_records_for_model(
|
||||
TRACKING_ROOT,
|
||||
record["model_id"],
|
||||
record["gpu_type"],
|
||||
**variant_filters,
|
||||
successful_only=True,
|
||||
baseline_eligible_only=True,
|
||||
)
|
||||
fingerprints = {str(r.get("recipe_fingerprint")) for r in same_variant}
|
||||
return sorted(fingerprints - {identity_filters["recipe_fingerprint"]})
|
||||
|
||||
|
||||
def _compare_record(
|
||||
record: dict[str, Any],
|
||||
metric_policies: tuple[MetricPolicy, ...],
|
||||
) -> tuple[list[str], list[dict[str, Any]]]:
|
||||
"""Compare one normalized record against its exact-identity cohort.
|
||||
|
||||
Sets ``record["comparator_status"]`` (and ``record["baseline_status"]``
|
||||
for cohort bookkeeping) and returns ``(failures, baseline_records)``.
|
||||
"""
|
||||
model = record.get("model_id", "unknown")
|
||||
|
||||
# Host-profile guard (r30 option a): a measurement taken on a packed host
|
||||
# is not comparable to anything, so skip gating loudly instead of failing
|
||||
# the PR. The record keeps its non-PASS status, so it can never advance
|
||||
# the rolling baseline. Records without a score gate normally (fail-open).
|
||||
host_cpu_score = safe_float(record.get("host_cpu_score"))
|
||||
min_score = _host_cpu_min_score()
|
||||
if host_cpu_score is not None and host_cpu_score < min_score:
|
||||
record["comparator_status"] = STATUS_HOST_BELOW_PROFILE
|
||||
print("=" * 72)
|
||||
print(f"HOST_BELOW_PROFILE: host below profile — measurements not "
|
||||
f"comparable (host_cpu_score={host_cpu_score:.3f} < floor "
|
||||
f"{min_score:.2f}) for {model}.")
|
||||
print("Regression gating is SKIPPED for this record and it will NOT "
|
||||
"seed a baseline. This is host CPU contention on the shared "
|
||||
"runner, not a property of the PR; retry the lane to land on a "
|
||||
"healthy host.")
|
||||
print("=" * 72)
|
||||
return [], []
|
||||
|
||||
try:
|
||||
identity_filters = _comparison_identity_filters(record)
|
||||
baseline_records = load_records_for_model(
|
||||
TRACKING_ROOT,
|
||||
record["model_id"],
|
||||
record["gpu_type"],
|
||||
**identity_filters,
|
||||
last_n=5,
|
||||
successful_only=True,
|
||||
baseline_eligible_only=True,
|
||||
)
|
||||
if baseline_records:
|
||||
record["baseline_status"] = "compared"
|
||||
failures = _check_regressions(record, baseline_records, metric_policies)
|
||||
record["comparator_status"] = STATUS_REGRESSION if failures else STATUS_PASS
|
||||
return failures, baseline_records
|
||||
|
||||
mismatched = _recipe_mismatch_fingerprints(record, identity_filters)
|
||||
except Exception as exc:
|
||||
record["comparator_status"] = STATUS_INFRA_ERROR
|
||||
return [f"{model} baseline comparison hit an infra error: {exc}"], []
|
||||
|
||||
if mismatched:
|
||||
record["comparator_status"] = STATUS_RECIPE_MISMATCH
|
||||
return [
|
||||
f"{model} recipe fingerprint {record.get('recipe_fingerprint')} does not "
|
||||
f"match baseline fingerprint(s) {', '.join(mismatched)} for variant "
|
||||
f"{record.get('variant_id')}. Represent recipe changes as a new "
|
||||
"variant_id, or reseed the baseline."
|
||||
], []
|
||||
|
||||
# No comparable baseline anywhere: a brand-new cohort. Make that loud and
|
||||
# machine-readable instead of an indistinguishable pass, so a cohort shift
|
||||
# (intended or accidental, e.g. an identity-field change) never silently
|
||||
# blinds the comparison — and never silently seeds a passing baseline.
|
||||
record["baseline_status"] = "initialized_new_cohort"
|
||||
record["comparator_status"] = STATUS_CALIBRATION_NEEDED
|
||||
print("=" * 72)
|
||||
print(f"CALIBRATION_NEEDED: no comparable baseline for {model} on "
|
||||
f"{record.get('gpu_type', 'unknown')}"
|
||||
f"{_format_identity_filters(identity_filters)}")
|
||||
print("Regression gating is INACTIVE for this record and it will NOT "
|
||||
"seed a baseline. Seed the cohort explicitly via the reseed "
|
||||
"workflow. If this cohort shift is unexpected, check the identity "
|
||||
"fields above.")
|
||||
print("=" * 72)
|
||||
return [], []
|
||||
|
||||
|
||||
def _compact_value(value: float | None, precision: int = 3) -> str:
|
||||
if value is None:
|
||||
return "n/a"
|
||||
@@ -394,6 +533,7 @@ def _build_summary_row(
|
||||
return {
|
||||
"model_id": record["model_id"],
|
||||
"gpu_type": record["gpu_type"],
|
||||
"comparator_status": record.get("comparator_status", STATUS_PASS),
|
||||
"baseline_n": len(baseline_records),
|
||||
"metrics": metric_values,
|
||||
"worst_regression_pct": worst_regression_pct,
|
||||
@@ -435,7 +575,10 @@ def _build_markdown_summary(
|
||||
else "none"
|
||||
)
|
||||
failing_metrics = ", ".join(row["failing_metrics"]) if row["failing_metrics"] else "none"
|
||||
status = "FAIL" if row["failed"] else "PASS"
|
||||
status = row["comparator_status"]
|
||||
if row["failed"] and status == STATUS_PASS:
|
||||
# Baseline comparison passed but the fixed-threshold phase failed.
|
||||
status = "FAIL"
|
||||
|
||||
lines.append(f"| {row['model_id']} | {row['gpu_type']} | "
|
||||
f"{row['baseline_n']} | "
|
||||
@@ -501,46 +644,26 @@ def main() -> int:
|
||||
if identity_filters is None:
|
||||
# Records without the full v2 identity block skip rolling-baseline
|
||||
# comparison entirely: only the static thresholds gate them and
|
||||
# they never become baseline eligible.
|
||||
# they never become baseline eligible. Nothing was compared and
|
||||
# nothing failed, so the comparator verdict is PASS;
|
||||
# baseline_status keeps the skip machine-readable.
|
||||
record["baseline_status"] = "skipped_missing_identity"
|
||||
record["comparator_status"] = STATUS_PASS
|
||||
failures: list[str] = []
|
||||
else:
|
||||
baseline_records = load_records_for_model(
|
||||
TRACKING_ROOT,
|
||||
record["model_id"],
|
||||
record["gpu_type"],
|
||||
**identity_filters,
|
||||
last_n=5,
|
||||
successful_only=True,
|
||||
baseline_eligible_only=True,
|
||||
)
|
||||
|
||||
if not baseline_records:
|
||||
# A brand-new cohort has NO regression gating until history
|
||||
# accumulates — make that loud and machine-readable instead of
|
||||
# an indistinguishable pass, so a cohort shift (intended or
|
||||
# accidental, e.g. an identity-field change) never silently
|
||||
# blinds the comparison.
|
||||
record["baseline_status"] = "initialized_new_cohort"
|
||||
print("=" * 72)
|
||||
print(f"WARNING: NO BASELINE — initializing a NEW cohort for "
|
||||
f"{record['model_id']} on {record['gpu_type']}"
|
||||
f"{_format_identity_filters(identity_filters)}")
|
||||
print("Regression gating is INACTIVE for this cohort until "
|
||||
"baseline history accumulates. If this cohort shift is "
|
||||
"unexpected, check the identity fields above.")
|
||||
print("=" * 72)
|
||||
failures = []
|
||||
else:
|
||||
record["baseline_status"] = "compared"
|
||||
failures = _check_regressions(record, baseline_records, metric_policies)
|
||||
failures, baseline_records = _compare_record(record, metric_policies)
|
||||
if static_threshold_failed:
|
||||
failures.append(f"{record['model_id']} fixed-threshold phase failed "
|
||||
f"(PERF_PYTEST_RC={os.environ.get('PERF_PYTEST_RC')})")
|
||||
|
||||
record["success"] = not failures
|
||||
# Only compared-and-PASS records may advance the rolling baseline: a
|
||||
# CALIBRATION_NEEDED record must never silently seed a new cohort, and
|
||||
# records without the v2 identity block never become eligible.
|
||||
record["baseline_eligible"] = (
|
||||
identity_filters is not None and _is_baseline_eligible(record["run_source"], record["success"])
|
||||
identity_filters is not None
|
||||
and record["comparator_status"] == STATUS_PASS
|
||||
and _is_baseline_eligible(record["run_source"], record["success"])
|
||||
)
|
||||
all_failures.extend(failures)
|
||||
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""Host CPU health probe for the performance lane.
|
||||
|
||||
The perf lane runs on shared Modal hosts. A packed host slows CPU-bound
|
||||
pipeline work by ~1.6-1.8x (observed: text encoder pinned at 3.595s vs the
|
||||
2.03s healthy cluster, DiT +58-80%) while GPU telemetry stays healthy, which
|
||||
produces false REGRESSION verdicts against PRs. CPU reservations are a floor,
|
||||
not isolation, so the comparator instead refuses to gate on measurements
|
||||
taken on a degraded host (review r30, option a).
|
||||
|
||||
Two fixed-work, wall-clock-timed sub-probes (~1.5s total), run once per
|
||||
pytest process before the first benchmark:
|
||||
|
||||
- single-thread: a pure-Python arithmetic loop. Captures per-core slowdown
|
||||
(CPU steal, scheduling delay, frequency/cache pressure). This is the axis
|
||||
``host_cpu_score`` gates on: it directly mirrors the packed-host signature
|
||||
(single-stream stage wall time), and healthy lane hosts cluster tightly.
|
||||
- multi-thread: torch float32 matmuls pinned to the lane's CPU reservation
|
||||
(cpu=8 in fastvideo/tests/modal/pr_test.py). Recorded as informational
|
||||
telemetry only: healthy-pool throughput spans 582-1010 GFLOP/s across host
|
||||
SKUs (1.74x), wider than the packed-host signal, so it cannot gate.
|
||||
|
||||
Calibration basis (2026-07-07, fastvideo-dev:latest image /opt/venv
|
||||
python 3.10.19 + torch 2.9.1, five Modal L40S:2 cpu=8 memory=32GiB
|
||||
containers — the exact perf-lane reservation): four hosts (cpu_count=21)
|
||||
scored 19.0-20.1k kops/s single-thread, one faster host (cpu_count=24)
|
||||
scored 22.9k. Reference = 19700 (common-SKU median), so healthy samples
|
||||
score 0.97-1.17. The packed-host signature (r30) inflates CPU-bound wall
|
||||
time 1.58-1.77x, i.e. throughput drops to 0.56-0.63x: 0.54-0.61 on the
|
||||
common SKU, at most ~0.73 on the fast SKU. The comparator's default floor
|
||||
of 0.75 (PERF_HOST_CPU_MIN_SCORE) sits above every packed projection and
|
||||
below every healthy sample. Recalibrate the reference after an image /
|
||||
python / torch bump via PERF_HOST_CPU_ST_REF_KOPS (a stale reference can
|
||||
only skip the gate loudly, never fail a PR).
|
||||
"""
|
||||
|
||||
import os
|
||||
import time
|
||||
|
||||
_DEFAULT_ST_REF_KOPS = 19700.0
|
||||
|
||||
_PROBE_THREADS = 8 # matches the perf lane's cpu=8.0 Modal reservation
|
||||
_ST_ITERS = 20_000_000 # ~1.0s single-threaded on a healthy lane host
|
||||
_MT_N = 2048
|
||||
_MT_REPS = 6
|
||||
|
||||
|
||||
def _single_thread_kops(iters: int = _ST_ITERS) -> float:
|
||||
"""Fixed pure-Python workload; returns kilo-iterations per second."""
|
||||
start = time.perf_counter()
|
||||
acc = 0
|
||||
for i in range(iters):
|
||||
acc += i * i
|
||||
elapsed = time.perf_counter() - start
|
||||
return iters / elapsed / 1e3
|
||||
|
||||
|
||||
def _multi_thread_gflops(n: int = _MT_N, reps: int = _MT_REPS) -> float:
|
||||
"""Fixed torch float32 matmul workload on the lane's CPU reservation."""
|
||||
import torch
|
||||
old_threads = torch.get_num_threads()
|
||||
torch.set_num_threads(min(_PROBE_THREADS, os.cpu_count() or _PROBE_THREADS))
|
||||
try:
|
||||
a = torch.rand(n, n, dtype=torch.float32)
|
||||
b = torch.rand(n, n, dtype=torch.float32)
|
||||
a @ b # warmup: thread-pool spin-up and first-touch page faults
|
||||
start = time.perf_counter()
|
||||
for _ in range(reps):
|
||||
a @ b
|
||||
elapsed = time.perf_counter() - start
|
||||
finally:
|
||||
torch.set_num_threads(old_threads)
|
||||
return reps * 2 * n**3 / elapsed / 1e9
|
||||
|
||||
|
||||
def measure_host_cpu_profile() -> dict[str, float]:
|
||||
"""Run both sub-probes and return the normalized host CPU profile.
|
||||
|
||||
``host_cpu_score`` is ~1.0 on a healthy perf-lane host (single-thread
|
||||
axis); the raw sub-probe throughputs are kept alongside for auditing
|
||||
and recalibration.
|
||||
"""
|
||||
st_ref = float(os.environ.get("PERF_HOST_CPU_ST_REF_KOPS", _DEFAULT_ST_REF_KOPS))
|
||||
st = _single_thread_kops()
|
||||
mt = _multi_thread_gflops()
|
||||
return {
|
||||
"host_cpu_score": round(st / st_ref, 3),
|
||||
"host_cpu_single_thread_kops": round(st, 1),
|
||||
"host_cpu_multi_thread_gflops": round(mt, 1),
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
# Recalibration helper: run this on a perf-lane container (Modal L40S:2,
|
||||
# cpu=8) a few times, take healthy-cluster medians, update the refs.
|
||||
print(measure_host_cpu_profile())
|
||||
@@ -421,3 +421,247 @@ def test_informational_metric_remains_visible_without_failing():
|
||||
assert row["metrics"]["throughput"]["regressed"] is False
|
||||
assert row["threshold_exceeded_metrics"] == ["throughput"]
|
||||
assert row["failing_metrics"] == []
|
||||
|
||||
|
||||
def _v2_raw_result(**overrides):
|
||||
raw = _raw_result()
|
||||
raw.update({
|
||||
"workload_id": "wan-t2v",
|
||||
"variant_id": "1.3b-sp2",
|
||||
"benchmark_version": 2,
|
||||
"recipe_fingerprint": "recipe-1",
|
||||
"hardware_profile_id": "hw-1",
|
||||
"software_profile_id": "sw-1",
|
||||
})
|
||||
raw.update(overrides)
|
||||
return raw
|
||||
|
||||
|
||||
def _v2_baseline_record(**overrides):
|
||||
record = {
|
||||
"gpu_type": "NVIDIA L40S",
|
||||
"workload_id": "wan-t2v",
|
||||
"variant_id": "1.3b-sp2",
|
||||
"benchmark_version": 2,
|
||||
"recipe_fingerprint": "recipe-1",
|
||||
"hardware_profile_id": "hw-1",
|
||||
"software_profile_id": "sw-1",
|
||||
"latency": 10.0,
|
||||
"throughput": 4.5,
|
||||
"memory": 10000.0,
|
||||
"timestamp": "2026-06-15T00:00:00+00:00",
|
||||
"success": True,
|
||||
"run_source": "scheduled_main",
|
||||
"baseline_eligible": True,
|
||||
}
|
||||
record.update(overrides)
|
||||
return record
|
||||
|
||||
|
||||
def _run_compare(monkeypatch, tmp_path, raw_result, baseline_records):
|
||||
"""Run compare_baseline.main() against a local tracking root.
|
||||
|
||||
Returns (exit_code, normalized_record, markdown_report).
|
||||
"""
|
||||
results_dir = tmp_path / "results"
|
||||
results_dir.mkdir()
|
||||
(results_dir / "perf_current.json").write_text(json.dumps(raw_result))
|
||||
|
||||
model_dir = tmp_path / "tracking" / raw_result["benchmark_id"]
|
||||
model_dir.mkdir(parents=True)
|
||||
for index, record in enumerate(baseline_records):
|
||||
(model_dir / f"rec{index}.json").write_text(json.dumps(record))
|
||||
|
||||
reports_dir = tmp_path / "reports"
|
||||
monkeypatch.setattr(compare_baseline, "RESULTS_DIR", str(results_dir))
|
||||
monkeypatch.setattr(compare_baseline, "TRACKING_ROOT", str(tmp_path / "tracking"))
|
||||
monkeypatch.setattr(compare_baseline, "PERF_REPORTS_DIR", str(reports_dir))
|
||||
monkeypatch.setattr(compare_baseline, "UPLOAD_POLICY", "never")
|
||||
monkeypatch.setattr(compare_baseline, "sync_from_hf", lambda *args, **kwargs: None)
|
||||
monkeypatch.setenv("PERF_RUN_SOURCE", "scheduled_main")
|
||||
monkeypatch.delenv("PERF_PYTEST_RC", raising=False)
|
||||
monkeypatch.delenv("GITHUB_STEP_SUMMARY", raising=False)
|
||||
|
||||
exit_code = compare_baseline.main()
|
||||
|
||||
normalized_paths = list((reports_dir / "results").glob("normalized_perf_*.json"))
|
||||
assert len(normalized_paths) == 1
|
||||
markdown = "\n".join(path.read_text() for path in reports_dir.glob("perf_*.md"))
|
||||
return exit_code, json.loads(normalized_paths[0].read_text()), markdown
|
||||
|
||||
|
||||
def test_main_pass_with_comparable_baseline(monkeypatch, tmp_path):
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(), [_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "PASS"
|
||||
assert record["baseline_status"] == "compared"
|
||||
assert record["success"] is True
|
||||
# Scheduled-main PASS records keep advancing the rolling baseline.
|
||||
assert record["baseline_eligible"] is True
|
||||
assert "| PASS |" in markdown
|
||||
|
||||
|
||||
def test_main_regression_on_gated_metric_fails_ci(monkeypatch, tmp_path):
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(avg_generation_time_s=20.0),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 1
|
||||
assert record["comparator_status"] == "REGRESSION"
|
||||
assert record["success"] is False
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| REGRESSION |" in markdown
|
||||
|
||||
|
||||
def test_main_missing_baseline_is_calibration_needed_and_does_not_seed(monkeypatch, tmp_path):
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(), [])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "CALIBRATION_NEEDED"
|
||||
assert record["baseline_status"] == "initialized_new_cohort"
|
||||
assert record["success"] is True
|
||||
# Visible, but never silently seeds a passing baseline.
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| CALIBRATION_NEEDED |" in markdown
|
||||
|
||||
|
||||
def test_main_recipe_change_without_new_variant_is_recipe_mismatch(monkeypatch, tmp_path):
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(recipe_fingerprint="recipe-2"),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 1
|
||||
assert record["comparator_status"] == "RECIPE_MISMATCH"
|
||||
assert record["success"] is False
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| RECIPE_MISMATCH |" in markdown
|
||||
|
||||
|
||||
def test_main_recipe_change_as_new_variant_is_calibration_needed(monkeypatch, tmp_path):
|
||||
exit_code, record, _ = _run_compare(
|
||||
monkeypatch, tmp_path,
|
||||
_v2_raw_result(variant_id="1.3b-sp2-r2", recipe_fingerprint="recipe-2"),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "CALIBRATION_NEEDED"
|
||||
|
||||
|
||||
def test_main_v2_record_never_compares_against_v1_baselines(monkeypatch, tmp_path):
|
||||
legacy_v1_baseline = {
|
||||
"gpu_type": "NVIDIA L40S",
|
||||
# Would be a >100% latency regression if the comparator (wrongly)
|
||||
# matched the v2 record against v1 history.
|
||||
"latency": 1.0,
|
||||
"timestamp": "2026-06-15T00:00:00+00:00",
|
||||
"success": True,
|
||||
}
|
||||
|
||||
exit_code, record, _ = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(), [legacy_v1_baseline])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "CALIBRATION_NEEDED"
|
||||
|
||||
|
||||
def test_main_legacy_v1_record_skips_comparison_with_pass_verdict(monkeypatch, tmp_path):
|
||||
legacy_v1_baseline = {
|
||||
"gpu_type": "NVIDIA L40S",
|
||||
# Would be a >100% latency regression if the comparator (wrongly)
|
||||
# compared the identity-less v1 record against this history.
|
||||
"latency": 1.0,
|
||||
"timestamp": "2026-06-15T00:00:00+00:00",
|
||||
"success": True,
|
||||
}
|
||||
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _raw_result(), [legacy_v1_baseline])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "PASS"
|
||||
assert record["baseline_status"] == "skipped_missing_identity"
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| PASS |" in markdown
|
||||
|
||||
|
||||
def test_main_partial_identity_skips_comparison(monkeypatch, tmp_path):
|
||||
raw = _v2_raw_result()
|
||||
del raw["software_profile_id"]
|
||||
|
||||
exit_code, record, _ = _run_compare(monkeypatch, tmp_path, raw, [])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "PASS"
|
||||
assert record["baseline_status"] == "skipped_missing_identity"
|
||||
assert record["success"] is True
|
||||
assert record["baseline_eligible"] is False
|
||||
|
||||
|
||||
def test_main_baseline_load_failure_is_infra_error_and_fails_ci(monkeypatch, tmp_path):
|
||||
def _boom(*_args, **_kwargs):
|
||||
raise OSError("hf store unavailable")
|
||||
|
||||
monkeypatch.setattr(compare_baseline, "load_records_for_model", _boom)
|
||||
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(), [])
|
||||
|
||||
assert exit_code == 1
|
||||
assert record["comparator_status"] == "INFRA_ERROR"
|
||||
assert record["success"] is False
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| INFRA_ERROR |" in markdown
|
||||
|
||||
|
||||
def test_main_host_below_profile_skips_gate_and_never_seeds_baseline(
|
||||
monkeypatch, tmp_path, capsys):
|
||||
# A 2x latency "regression" measured on a packed host must not fail CI.
|
||||
exit_code, record, markdown = _run_compare(
|
||||
monkeypatch, tmp_path,
|
||||
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.55),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "HOST_BELOW_PROFILE"
|
||||
assert record["success"] is True
|
||||
assert record["baseline_eligible"] is False
|
||||
assert "| HOST_BELOW_PROFILE |" in markdown
|
||||
assert ("host below profile — measurements not comparable"
|
||||
in capsys.readouterr().out)
|
||||
|
||||
|
||||
def test_main_healthy_host_score_gates_normally(monkeypatch, tmp_path):
|
||||
exit_code, record, _ = _run_compare(
|
||||
monkeypatch, tmp_path,
|
||||
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.97),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 1
|
||||
assert record["comparator_status"] == "REGRESSION"
|
||||
|
||||
|
||||
def test_main_healthy_host_pass_keeps_baseline_eligibility_and_score(
|
||||
monkeypatch, tmp_path):
|
||||
exit_code, record, _ = _run_compare(
|
||||
monkeypatch, tmp_path, _v2_raw_result(host_cpu_score=1.02),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 0
|
||||
assert record["comparator_status"] == "PASS"
|
||||
assert record["baseline_eligible"] is True
|
||||
assert record["host_cpu_score"] == 1.02
|
||||
|
||||
|
||||
def test_host_cpu_min_score_env_override(monkeypatch, tmp_path):
|
||||
monkeypatch.setenv("PERF_HOST_CPU_MIN_SCORE", "0.5")
|
||||
|
||||
exit_code, record, _ = _run_compare(
|
||||
monkeypatch, tmp_path,
|
||||
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.55),
|
||||
[_v2_baseline_record()])
|
||||
|
||||
assert exit_code == 1
|
||||
assert record["comparator_status"] == "REGRESSION"
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""CPU-runnable smoke test for the host CPU health probe (~2s)."""
|
||||
|
||||
import pytest
|
||||
|
||||
from fastvideo.tests.performance import host_probe
|
||||
|
||||
|
||||
def test_probe_reports_positive_normalized_profile(monkeypatch):
|
||||
monkeypatch.setenv("PERF_HOST_CPU_ST_REF_KOPS", "1000")
|
||||
|
||||
profile = host_probe.measure_host_cpu_profile()
|
||||
|
||||
assert profile["host_cpu_single_thread_kops"] > 0
|
||||
assert profile["host_cpu_multi_thread_gflops"] > 0
|
||||
# The gated score is the single-thread axis normalized by the reference.
|
||||
assert profile["host_cpu_score"] == pytest.approx(
|
||||
profile["host_cpu_single_thread_kops"] / 1000.0, rel=0.05)
|
||||
@@ -19,6 +19,7 @@ import pytest
|
||||
|
||||
from fastvideo import VideoGenerator
|
||||
from fastvideo.logger import init_logger
|
||||
from fastvideo.tests.performance.host_probe import measure_host_cpu_profile
|
||||
from fastvideo.tests.performance.identity import (
|
||||
benchmark_identity_from_config,
|
||||
build_recipe_from_benchmark_config,
|
||||
@@ -153,6 +154,24 @@ _BENCHMARK_CONFIGS = _discover_benchmarks()
|
||||
|
||||
# -- Helpers ----------------------------------------------------------------
|
||||
|
||||
# Probed once per pytest process (before the first benchmark's GPU work) and
|
||||
# attached to every result: the comparator skips gating when the host is
|
||||
# below the healthy-host CPU profile. Fail-open: a probe error just omits the
|
||||
# fields, and the comparator gates normally.
|
||||
_HOST_CPU_PROFILE: dict[str, float] | None = None
|
||||
|
||||
|
||||
def _host_cpu_profile() -> dict[str, float]:
|
||||
global _HOST_CPU_PROFILE
|
||||
if _HOST_CPU_PROFILE is None:
|
||||
try:
|
||||
_HOST_CPU_PROFILE = measure_host_cpu_profile()
|
||||
logger.info("Host CPU probe: %s", _HOST_CPU_PROFILE)
|
||||
except Exception as e:
|
||||
logger.warning("Host CPU probe failed (%s); records will carry no host_cpu_score", e)
|
||||
_HOST_CPU_PROFILE = {}
|
||||
return _HOST_CPU_PROFILE
|
||||
|
||||
|
||||
def _get_thresholds(cfg):
|
||||
"""Return thresholds dict for the current GPU from config."""
|
||||
@@ -477,6 +496,7 @@ def _run_benchmark(cfg):
|
||||
|
||||
num_warmup, num_measure = _validate_run_counts(run_config, cfg["benchmark_id"])
|
||||
thresholds = _get_thresholds(cfg)
|
||||
host_cpu_profile = _host_cpu_profile()
|
||||
|
||||
# Remap JSON keys to VideoGenerator kwargs
|
||||
text_enc_prec = init_kwargs.pop("text_encoder_precisions", None)
|
||||
@@ -537,6 +557,9 @@ def _run_benchmark(cfg):
|
||||
runtime_identity=runtime_identity,
|
||||
device_name=device_name,
|
||||
)
|
||||
# Attach the host CPU probe so the comparator can skip gating on
|
||||
# below-profile (packed) hosts. Fail-open: an empty probe adds nothing.
|
||||
results.update(host_cpu_profile)
|
||||
|
||||
logger.info(
|
||||
"Performance results: avg_time=%.2fs, "
|
||||
|
||||
@@ -14,6 +14,7 @@ from fastvideo.training.wan_training_pipeline import main
|
||||
from fastvideo.fastvideo_args import FastVideoArgs, TrainingArgs
|
||||
from fastvideo.utils import FlexibleArgumentParser
|
||||
from fastvideo.training.wan_training_pipeline import WanTrainingPipeline
|
||||
from fastvideo.tests.performance.host_probe import measure_host_cpu_profile
|
||||
|
||||
wandb_name = "test_training_loss"
|
||||
a40_reference_wandb_summary_file = "fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json"
|
||||
@@ -142,6 +143,31 @@ def test_distributed_training():
|
||||
'train_loss': 0.0025
|
||||
}
|
||||
|
||||
# Host-profile guard (mirrors compare_baseline.py's HOST_BELOW_PROFILE):
|
||||
# a packed shared host inflates wall time (observed ~8x) with bit-clean
|
||||
# loss, so the step-time gates are skipped loudly on a below-profile
|
||||
# host. Loss and grad-norm stay asserted — they are host-independent.
|
||||
# Fail-open: a probe error keeps the step-time gates active.
|
||||
try:
|
||||
min_score = float(os.environ.get("PERF_HOST_CPU_MIN_SCORE", "") or 0.75)
|
||||
except ValueError:
|
||||
min_score = 0.75
|
||||
try:
|
||||
host_cpu_score = measure_host_cpu_profile()["host_cpu_score"]
|
||||
except Exception as e:
|
||||
host_cpu_score = None
|
||||
print(f"Host CPU probe failed ({e}); step-time gates stay active")
|
||||
if host_cpu_score is not None and host_cpu_score < min_score:
|
||||
print("=" * 72)
|
||||
print(f"HOST_BELOW_PROFILE: host_cpu_score={host_cpu_score:.3f} < "
|
||||
f"floor {min_score:.2f} — skipping the step-time gates "
|
||||
"(avg_step_time, step_time). This is host CPU contention on "
|
||||
"the shared runner, not a property of the PR. Loss values are "
|
||||
"still asserted.")
|
||||
print("=" * 72)
|
||||
del fields_and_thresholds['avg_step_time']
|
||||
del fields_and_thresholds['step_time']
|
||||
|
||||
failures = []
|
||||
for field, threshold in fields_and_thresholds.items():
|
||||
ref_value = reference_wandb_summary[field]
|
||||
|
||||
Reference in New Issue
Block a user