Compare commits

...
Author SHA1 Message Date
SolitaryThinker 39b015713f [ci] training-loss lane: skip step-time gate on below-profile hosts 2026-07-13 16:25:26 -07:00
SolitaryThinker da9cbe3da9 [ci]: skip perf gating on packed hosts via host CPU profile guard
The wan-t2v perf lane intermittently runs ~1.7x slow on packed Modal
hosts (TE pinned at 3.595s vs the 2.03s healthy cluster, DiT +58-80%,
GPU clocks at full boost): host CPU contention that cpu=8/16
reservations cannot isolate. 9 of the last ~11 full suites redded only
on this lane. A perf gate must never fail a PR on a measurement taken
on a degraded host (review r30, option a).

The harness now runs a ~1.5s host CPU probe (single-thread pure-Python
loop, gated axis; 8-thread torch matmul, informational) once per pytest
process and stores host_cpu_score on every record. The comparator emits
HOST_BELOW_PROFILE when the score is under the calibrated floor
(PERF_HOST_CPU_MIN_SCORE, default 0.75): gating is skipped loudly, CI
passes, and the record can never seed a baseline. Records without a
score gate normally.

Calibrated against five fresh perf-lane containers (L40S:2 cpu=8):
healthy hosts score 0.97-1.17, packed-host projections 0.54-0.73;
sampled data and recalibration procedure documented in host_probe.py.

Stacked on maint/perf-comparator-statuses (f46c9e61).
2026-07-13 16:25:26 -07:00
SolitaryThinker 04bfddb7fa [ci]: add exact-identity perf comparator statuses
Comparator verdicts per issue #1532: PASS / REGRESSION /
CALIBRATION_NEEDED / RECIPE_MISMATCH / INFRA_ERROR (QUALITY_BLOCKED
reserved for the promoted-baseline workflow). REGRESSION and
RECIPE_MISMATCH fail CI; CALIBRATION_NEEDED is loud but no longer
auto-seeds a passing baseline.
2026-07-13 16:24:52 -07:00
7 changed files with 603 additions and 45 deletions
+41 -13
View File
@@ -232,13 +232,26 @@ by the harness and is not config-declarable.)
Recipe fingerprinting, hardware/software profile IDs, exact-identity
comparison, and dashboard cohort grouping land with this change: v2 records
compare only within their identity cohort, and a record that opens a NEW
cohort is marked `baseline_status: "initialized_new_cohort"` (regression
gating starts once that cohort accumulates history). Legacy v1 configs still
run and are normalized for reporting, but their records skip rolling-baseline
comparison entirely (`baseline_status: "skipped_missing_identity"`, never
baseline eligible); only static thresholds gate them. Metric-specific
threshold policies and promoted baselines remain separate follow-ups.
compare only within their identity cohort and are never compared against v1
records (or vice versa). Legacy records missing any comparison identity field
still run and are normalized for reporting, but they skip rolling-baseline
comparison entirely (`baseline_status: "skipped_missing_identity"`, comparator
verdict `PASS`, never baseline eligible); only static thresholds gate them.
Each comparison reports an explicit `comparator_status` on the normalized
record and in the Markdown summary:
| Status | Meaning | CI |
|---|---|---|
| `PASS` | Comparable baseline found with no gated regression, or comparison was skipped for a record without the v2 identity block. | passes |
| `REGRESSION` | A gated metric regressed past both its percent and absolute floors. | fails |
| `CALIBRATION_NEEDED` | No comparable baseline exists. Gating is inactive and the record does **not** seed a baseline — seed new cohorts explicitly via the reseed workflow. The record also carries `baseline_status: "initialized_new_cohort"`. | passes |
| `RECIPE_MISMATCH` | Baseline history exists for the same variant/hardware/software cohort but under a different `recipe_fingerprint`. Represent recipe changes as a new `variant_id`, or reseed. | fails |
| `INFRA_ERROR` | The comparison itself failed (for example, baseline history could not be loaded). | fails |
| `HOST_BELOW_PROFILE` | The harness's ~2s host CPU probe (`host_probe.py`) scored below the healthy-host floor (`PERF_HOST_CPU_MIN_SCORE`, default 0.75): the shared runner was packed, so the measurements are not comparable. Gating is skipped and the record never seeds a baseline. | passes |
| `QUALITY_BLOCKED` | Reserved for the promoted-baseline workflow; never emitted by the comparator. | n/a |
Metric-specific threshold policies and promoted baselines remain separate
follow-ups.
### Raw record (`results/perf_*.json`)
@@ -367,6 +380,8 @@ result, used as the rolling-baseline source of truth.
"build_id": "<buildkite-build-id>",
"job_id": "<buildkite-job-id>",
"quality_metadata": { "quality_status": "canonical" },
"baseline_status": "compared",
"comparator_status": "PASS",
"success": true
}
```
@@ -487,10 +502,10 @@ When the rolling-baseline phase runs, it emits:
2. The pytest test auto-discovers all configs — no test code needed. CI
picks it up on the next `/test performance` run.
3. The first persisted main-branch run with no HF history initializes the
baseline (passes automatically). Subsequent runs compare against it. Local
and pull-request runs with no HF history also pass, but they do not seed the
shared baseline.
3. The first run with no HF history reports `CALIBRATION_NEEDED`: it passes,
but it does not seed the shared baseline. Seed the cohort explicitly with
the `reseed-performance-baseline` skill; subsequent scheduled-main runs
then compare against it and keep the rolling baseline advancing.
4. If the benchmark targets a GPU not currently in `thresholds`, either add
that GPU as a key or rely on the `default` block. Note that `default` is
@@ -510,8 +525,21 @@ When the rolling-baseline phase runs, it emits:
## Troubleshooting
**"No baseline for ... Initializing"** — first run for this comparison cohort.
Run will pass and (if persisting) seed the first record.
**`CALIBRATION_NEEDED: no comparable baseline`** — first run for this cohort.
The run passes but does not seed a baseline; seed it explicitly with the
`reseed-performance-baseline` skill.
**`RECIPE_MISMATCH`** — the benchmark recipe changed without a new
`variant_id`. Either bump `variant_id` to open a new cohort, or reseed the
baseline if the existing variant should adopt the new recipe.
**`HOST_BELOW_PROFILE: host below profile — measurements not comparable`** —
the shared Modal host was packed during the run (host CPU contention; GPU
telemetry typically still looks healthy). Not a property of the PR: the gate
is skipped and CI passes. Retry the lane if you need a comparable measurement.
If every host suddenly scores low after an image/python/torch bump, the probe
reference is stale — recalibrate `PERF_HOST_CPU_ST_REF_KOPS` (see
`host_probe.py`, which documents the calibration procedure and sampled data).
**Persistent failure right after a torch / kernel / image upgrade** —
genuine regression *or* baseline drift. Compare the failing normalized record
+155 -32
View File
@@ -89,6 +89,34 @@ COMPARISON_IDENTITY_KEYS = (
"hardware_profile_id",
"software_profile_id",
)
# Comparator verdicts (issue #1532), stored as record["comparator_status"].
STATUS_PASS = "PASS"
STATUS_REGRESSION = "REGRESSION"
STATUS_CALIBRATION_NEEDED = "CALIBRATION_NEEDED"
STATUS_RECIPE_MISMATCH = "RECIPE_MISMATCH"
STATUS_INFRA_ERROR = "INFRA_ERROR"
STATUS_HOST_BELOW_PROFILE = "HOST_BELOW_PROFILE"
# Reserved for the promoted-baseline workflow; never emitted by this comparator.
STATUS_QUALITY_BLOCKED = "QUALITY_BLOCKED"
# Floor for host_cpu_score (host_probe.py: ~1.0 on a healthy perf-lane host).
# Calibration basis (see host_probe.py for the sampled data): healthy lane
# hosts scored 0.97-1.17; the packed-host signature (review r30; TE 3.595s vs
# the 2.03s healthy cluster, DiT +58-80%) inflates CPU-bound wall time by
# 1.58-1.77x, i.e. a packed host projects to 0.54-0.73. 0.75 sits above every
# packed projection and below every healthy sample.
DEFAULT_HOST_CPU_MIN_SCORE = 0.75
def _host_cpu_min_score() -> float:
raw = os.environ.get("PERF_HOST_CPU_MIN_SCORE", "").strip()
if raw:
try:
return float(raw)
except ValueError:
print(f"Invalid PERF_HOST_CPU_MIN_SCORE={raw!r}; "
f"using {DEFAULT_HOST_CPU_MIN_SCORE}")
return DEFAULT_HOST_CPU_MIN_SCORE
def _should_persist_tracking() -> bool:
@@ -260,6 +288,13 @@ def normalize_performance_result(result: dict[str, Any]) -> dict[str, Any]:
"success": True,
**_record_metadata(_detect_run_source(), result),
}
# Host CPU probe fields (host_probe.py); absent on records measured
# before the probe existed or when the probe failed.
for key in ("host_cpu_score", "host_cpu_single_thread_kops",
"host_cpu_multi_thread_gflops"):
value = safe_float(result.get(key))
if value is not None:
record[key] = value
record.update(_identity_metadata(result))
return record
@@ -343,6 +378,110 @@ def _check_regressions(
return failures
def _recipe_mismatch_fingerprints(
record: dict[str, Any],
identity_filters: dict[str, str],
) -> list[str]:
"""Return baseline recipe fingerprints for the same variant cohort.
Non-empty means the same (workload, variant, version, hardware, software)
cohort has baseline history under a DIFFERENT recipe fingerprint: the
recipe changed without being represented as a new variant.
"""
if not identity_filters:
return []
variant_filters = {key: value for key, value in identity_filters.items() if key != "recipe_fingerprint"}
same_variant = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**variant_filters,
successful_only=True,
baseline_eligible_only=True,
)
fingerprints = {str(r.get("recipe_fingerprint")) for r in same_variant}
return sorted(fingerprints - {identity_filters["recipe_fingerprint"]})
def _compare_record(
record: dict[str, Any],
metric_policies: tuple[MetricPolicy, ...],
) -> tuple[list[str], list[dict[str, Any]]]:
"""Compare one normalized record against its exact-identity cohort.
Sets ``record["comparator_status"]`` (and ``record["baseline_status"]``
for cohort bookkeeping) and returns ``(failures, baseline_records)``.
"""
model = record.get("model_id", "unknown")
# Host-profile guard (r30 option a): a measurement taken on a packed host
# is not comparable to anything, so skip gating loudly instead of failing
# the PR. The record keeps its non-PASS status, so it can never advance
# the rolling baseline. Records without a score gate normally (fail-open).
host_cpu_score = safe_float(record.get("host_cpu_score"))
min_score = _host_cpu_min_score()
if host_cpu_score is not None and host_cpu_score < min_score:
record["comparator_status"] = STATUS_HOST_BELOW_PROFILE
print("=" * 72)
print(f"HOST_BELOW_PROFILE: host below profile — measurements not "
f"comparable (host_cpu_score={host_cpu_score:.3f} < floor "
f"{min_score:.2f}) for {model}.")
print("Regression gating is SKIPPED for this record and it will NOT "
"seed a baseline. This is host CPU contention on the shared "
"runner, not a property of the PR; retry the lane to land on a "
"healthy host.")
print("=" * 72)
return [], []
try:
identity_filters = _comparison_identity_filters(record)
baseline_records = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**identity_filters,
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
if baseline_records:
record["baseline_status"] = "compared"
failures = _check_regressions(record, baseline_records, metric_policies)
record["comparator_status"] = STATUS_REGRESSION if failures else STATUS_PASS
return failures, baseline_records
mismatched = _recipe_mismatch_fingerprints(record, identity_filters)
except Exception as exc:
record["comparator_status"] = STATUS_INFRA_ERROR
return [f"{model} baseline comparison hit an infra error: {exc}"], []
if mismatched:
record["comparator_status"] = STATUS_RECIPE_MISMATCH
return [
f"{model} recipe fingerprint {record.get('recipe_fingerprint')} does not "
f"match baseline fingerprint(s) {', '.join(mismatched)} for variant "
f"{record.get('variant_id')}. Represent recipe changes as a new "
"variant_id, or reseed the baseline."
], []
# No comparable baseline anywhere: a brand-new cohort. Make that loud and
# machine-readable instead of an indistinguishable pass, so a cohort shift
# (intended or accidental, e.g. an identity-field change) never silently
# blinds the comparison — and never silently seeds a passing baseline.
record["baseline_status"] = "initialized_new_cohort"
record["comparator_status"] = STATUS_CALIBRATION_NEEDED
print("=" * 72)
print(f"CALIBRATION_NEEDED: no comparable baseline for {model} on "
f"{record.get('gpu_type', 'unknown')}"
f"{_format_identity_filters(identity_filters)}")
print("Regression gating is INACTIVE for this record and it will NOT "
"seed a baseline. Seed the cohort explicitly via the reseed "
"workflow. If this cohort shift is unexpected, check the identity "
"fields above.")
print("=" * 72)
return [], []
def _compact_value(value: float | None, precision: int = 3) -> str:
if value is None:
return "n/a"
@@ -394,6 +533,7 @@ def _build_summary_row(
return {
"model_id": record["model_id"],
"gpu_type": record["gpu_type"],
"comparator_status": record.get("comparator_status", STATUS_PASS),
"baseline_n": len(baseline_records),
"metrics": metric_values,
"worst_regression_pct": worst_regression_pct,
@@ -435,7 +575,10 @@ def _build_markdown_summary(
else "none"
)
failing_metrics = ", ".join(row["failing_metrics"]) if row["failing_metrics"] else "none"
status = "FAIL" if row["failed"] else "PASS"
status = row["comparator_status"]
if row["failed"] and status == STATUS_PASS:
# Baseline comparison passed but the fixed-threshold phase failed.
status = "FAIL"
lines.append(f"| {row['model_id']} | {row['gpu_type']} | "
f"{row['baseline_n']} | "
@@ -501,46 +644,26 @@ def main() -> int:
if identity_filters is None:
# Records without the full v2 identity block skip rolling-baseline
# comparison entirely: only the static thresholds gate them and
# they never become baseline eligible.
# they never become baseline eligible. Nothing was compared and
# nothing failed, so the comparator verdict is PASS;
# baseline_status keeps the skip machine-readable.
record["baseline_status"] = "skipped_missing_identity"
record["comparator_status"] = STATUS_PASS
failures: list[str] = []
else:
baseline_records = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**identity_filters,
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
if not baseline_records:
# A brand-new cohort has NO regression gating until history
# accumulates — make that loud and machine-readable instead of
# an indistinguishable pass, so a cohort shift (intended or
# accidental, e.g. an identity-field change) never silently
# blinds the comparison.
record["baseline_status"] = "initialized_new_cohort"
print("=" * 72)
print(f"WARNING: NO BASELINE — initializing a NEW cohort for "
f"{record['model_id']} on {record['gpu_type']}"
f"{_format_identity_filters(identity_filters)}")
print("Regression gating is INACTIVE for this cohort until "
"baseline history accumulates. If this cohort shift is "
"unexpected, check the identity fields above.")
print("=" * 72)
failures = []
else:
record["baseline_status"] = "compared"
failures = _check_regressions(record, baseline_records, metric_policies)
failures, baseline_records = _compare_record(record, metric_policies)
if static_threshold_failed:
failures.append(f"{record['model_id']} fixed-threshold phase failed "
f"(PERF_PYTEST_RC={os.environ.get('PERF_PYTEST_RC')})")
record["success"] = not failures
# Only compared-and-PASS records may advance the rolling baseline: a
# CALIBRATION_NEEDED record must never silently seed a new cohort, and
# records without the v2 identity block never become eligible.
record["baseline_eligible"] = (
identity_filters is not None and _is_baseline_eligible(record["run_source"], record["success"])
identity_filters is not None
and record["comparator_status"] == STATUS_PASS
and _is_baseline_eligible(record["run_source"], record["success"])
)
all_failures.extend(failures)
+96
View File
@@ -0,0 +1,96 @@
# SPDX-License-Identifier: Apache-2.0
"""Host CPU health probe for the performance lane.
The perf lane runs on shared Modal hosts. A packed host slows CPU-bound
pipeline work by ~1.6-1.8x (observed: text encoder pinned at 3.595s vs the
2.03s healthy cluster, DiT +58-80%) while GPU telemetry stays healthy, which
produces false REGRESSION verdicts against PRs. CPU reservations are a floor,
not isolation, so the comparator instead refuses to gate on measurements
taken on a degraded host (review r30, option a).
Two fixed-work, wall-clock-timed sub-probes (~1.5s total), run once per
pytest process before the first benchmark:
- single-thread: a pure-Python arithmetic loop. Captures per-core slowdown
(CPU steal, scheduling delay, frequency/cache pressure). This is the axis
``host_cpu_score`` gates on: it directly mirrors the packed-host signature
(single-stream stage wall time), and healthy lane hosts cluster tightly.
- multi-thread: torch float32 matmuls pinned to the lane's CPU reservation
(cpu=8 in fastvideo/tests/modal/pr_test.py). Recorded as informational
telemetry only: healthy-pool throughput spans 582-1010 GFLOP/s across host
SKUs (1.74x), wider than the packed-host signal, so it cannot gate.
Calibration basis (2026-07-07, fastvideo-dev:latest image /opt/venv
python 3.10.19 + torch 2.9.1, five Modal L40S:2 cpu=8 memory=32GiB
containers — the exact perf-lane reservation): four hosts (cpu_count=21)
scored 19.0-20.1k kops/s single-thread, one faster host (cpu_count=24)
scored 22.9k. Reference = 19700 (common-SKU median), so healthy samples
score 0.97-1.17. The packed-host signature (r30) inflates CPU-bound wall
time 1.58-1.77x, i.e. throughput drops to 0.56-0.63x: 0.54-0.61 on the
common SKU, at most ~0.73 on the fast SKU. The comparator's default floor
of 0.75 (PERF_HOST_CPU_MIN_SCORE) sits above every packed projection and
below every healthy sample. Recalibrate the reference after an image /
python / torch bump via PERF_HOST_CPU_ST_REF_KOPS (a stale reference can
only skip the gate loudly, never fail a PR).
"""
import os
import time
_DEFAULT_ST_REF_KOPS = 19700.0
_PROBE_THREADS = 8 # matches the perf lane's cpu=8.0 Modal reservation
_ST_ITERS = 20_000_000 # ~1.0s single-threaded on a healthy lane host
_MT_N = 2048
_MT_REPS = 6
def _single_thread_kops(iters: int = _ST_ITERS) -> float:
"""Fixed pure-Python workload; returns kilo-iterations per second."""
start = time.perf_counter()
acc = 0
for i in range(iters):
acc += i * i
elapsed = time.perf_counter() - start
return iters / elapsed / 1e3
def _multi_thread_gflops(n: int = _MT_N, reps: int = _MT_REPS) -> float:
"""Fixed torch float32 matmul workload on the lane's CPU reservation."""
import torch
old_threads = torch.get_num_threads()
torch.set_num_threads(min(_PROBE_THREADS, os.cpu_count() or _PROBE_THREADS))
try:
a = torch.rand(n, n, dtype=torch.float32)
b = torch.rand(n, n, dtype=torch.float32)
a @ b # warmup: thread-pool spin-up and first-touch page faults
start = time.perf_counter()
for _ in range(reps):
a @ b
elapsed = time.perf_counter() - start
finally:
torch.set_num_threads(old_threads)
return reps * 2 * n**3 / elapsed / 1e9
def measure_host_cpu_profile() -> dict[str, float]:
"""Run both sub-probes and return the normalized host CPU profile.
``host_cpu_score`` is ~1.0 on a healthy perf-lane host (single-thread
axis); the raw sub-probe throughputs are kept alongside for auditing
and recalibration.
"""
st_ref = float(os.environ.get("PERF_HOST_CPU_ST_REF_KOPS", _DEFAULT_ST_REF_KOPS))
st = _single_thread_kops()
mt = _multi_thread_gflops()
return {
"host_cpu_score": round(st / st_ref, 3),
"host_cpu_single_thread_kops": round(st, 1),
"host_cpu_multi_thread_gflops": round(mt, 1),
}
if __name__ == "__main__":
# Recalibration helper: run this on a perf-lane container (Modal L40S:2,
# cpu=8) a few times, take healthy-cluster medians, update the refs.
print(measure_host_cpu_profile())
@@ -421,3 +421,247 @@ def test_informational_metric_remains_visible_without_failing():
assert row["metrics"]["throughput"]["regressed"] is False
assert row["threshold_exceeded_metrics"] == ["throughput"]
assert row["failing_metrics"] == []
def _v2_raw_result(**overrides):
raw = _raw_result()
raw.update({
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"recipe_fingerprint": "recipe-1",
"hardware_profile_id": "hw-1",
"software_profile_id": "sw-1",
})
raw.update(overrides)
return raw
def _v2_baseline_record(**overrides):
record = {
"gpu_type": "NVIDIA L40S",
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"recipe_fingerprint": "recipe-1",
"hardware_profile_id": "hw-1",
"software_profile_id": "sw-1",
"latency": 10.0,
"throughput": 4.5,
"memory": 10000.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
"run_source": "scheduled_main",
"baseline_eligible": True,
}
record.update(overrides)
return record
def _run_compare(monkeypatch, tmp_path, raw_result, baseline_records):
"""Run compare_baseline.main() against a local tracking root.
Returns (exit_code, normalized_record, markdown_report).
"""
results_dir = tmp_path / "results"
results_dir.mkdir()
(results_dir / "perf_current.json").write_text(json.dumps(raw_result))
model_dir = tmp_path / "tracking" / raw_result["benchmark_id"]
model_dir.mkdir(parents=True)
for index, record in enumerate(baseline_records):
(model_dir / f"rec{index}.json").write_text(json.dumps(record))
reports_dir = tmp_path / "reports"
monkeypatch.setattr(compare_baseline, "RESULTS_DIR", str(results_dir))
monkeypatch.setattr(compare_baseline, "TRACKING_ROOT", str(tmp_path / "tracking"))
monkeypatch.setattr(compare_baseline, "PERF_REPORTS_DIR", str(reports_dir))
monkeypatch.setattr(compare_baseline, "UPLOAD_POLICY", "never")
monkeypatch.setattr(compare_baseline, "sync_from_hf", lambda *args, **kwargs: None)
monkeypatch.setenv("PERF_RUN_SOURCE", "scheduled_main")
monkeypatch.delenv("PERF_PYTEST_RC", raising=False)
monkeypatch.delenv("GITHUB_STEP_SUMMARY", raising=False)
exit_code = compare_baseline.main()
normalized_paths = list((reports_dir / "results").glob("normalized_perf_*.json"))
assert len(normalized_paths) == 1
markdown = "\n".join(path.read_text() for path in reports_dir.glob("perf_*.md"))
return exit_code, json.loads(normalized_paths[0].read_text()), markdown
def test_main_pass_with_comparable_baseline(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "compared"
assert record["success"] is True
# Scheduled-main PASS records keep advancing the rolling baseline.
assert record["baseline_eligible"] is True
assert "| PASS |" in markdown
def test_main_regression_on_gated_metric_fails_ci(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(avg_generation_time_s=20.0),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "REGRESSION"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| REGRESSION |" in markdown
def test_main_missing_baseline_is_calibration_needed_and_does_not_seed(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
assert record["baseline_status"] == "initialized_new_cohort"
assert record["success"] is True
# Visible, but never silently seeds a passing baseline.
assert record["baseline_eligible"] is False
assert "| CALIBRATION_NEEDED |" in markdown
def test_main_recipe_change_without_new_variant_is_recipe_mismatch(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(recipe_fingerprint="recipe-2"),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "RECIPE_MISMATCH"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| RECIPE_MISMATCH |" in markdown
def test_main_recipe_change_as_new_variant_is_calibration_needed(monkeypatch, tmp_path):
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path,
_v2_raw_result(variant_id="1.3b-sp2-r2", recipe_fingerprint="recipe-2"),
[_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
def test_main_v2_record_never_compares_against_v1_baselines(monkeypatch, tmp_path):
legacy_v1_baseline = {
"gpu_type": "NVIDIA L40S",
# Would be a >100% latency regression if the comparator (wrongly)
# matched the v2 record against v1 history.
"latency": 1.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
}
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [legacy_v1_baseline])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
def test_main_legacy_v1_record_skips_comparison_with_pass_verdict(monkeypatch, tmp_path):
legacy_v1_baseline = {
"gpu_type": "NVIDIA L40S",
# Would be a >100% latency regression if the comparator (wrongly)
# compared the identity-less v1 record against this history.
"latency": 1.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
}
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _raw_result(), [legacy_v1_baseline])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "skipped_missing_identity"
assert record["baseline_eligible"] is False
assert "| PASS |" in markdown
def test_main_partial_identity_skips_comparison(monkeypatch, tmp_path):
raw = _v2_raw_result()
del raw["software_profile_id"]
exit_code, record, _ = _run_compare(monkeypatch, tmp_path, raw, [])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "skipped_missing_identity"
assert record["success"] is True
assert record["baseline_eligible"] is False
def test_main_baseline_load_failure_is_infra_error_and_fails_ci(monkeypatch, tmp_path):
def _boom(*_args, **_kwargs):
raise OSError("hf store unavailable")
monkeypatch.setattr(compare_baseline, "load_records_for_model", _boom)
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [])
assert exit_code == 1
assert record["comparator_status"] == "INFRA_ERROR"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| INFRA_ERROR |" in markdown
def test_main_host_below_profile_skips_gate_and_never_seeds_baseline(
monkeypatch, tmp_path, capsys):
# A 2x latency "regression" measured on a packed host must not fail CI.
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path,
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.55),
[_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "HOST_BELOW_PROFILE"
assert record["success"] is True
assert record["baseline_eligible"] is False
assert "| HOST_BELOW_PROFILE |" in markdown
assert ("host below profile — measurements not comparable"
in capsys.readouterr().out)
def test_main_healthy_host_score_gates_normally(monkeypatch, tmp_path):
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path,
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.97),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "REGRESSION"
def test_main_healthy_host_pass_keeps_baseline_eligibility_and_score(
monkeypatch, tmp_path):
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(host_cpu_score=1.02),
[_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_eligible"] is True
assert record["host_cpu_score"] == 1.02
def test_host_cpu_min_score_env_override(monkeypatch, tmp_path):
monkeypatch.setenv("PERF_HOST_CPU_MIN_SCORE", "0.5")
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path,
_v2_raw_result(avg_generation_time_s=20.0, host_cpu_score=0.55),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "REGRESSION"
@@ -0,0 +1,18 @@
# SPDX-License-Identifier: Apache-2.0
"""CPU-runnable smoke test for the host CPU health probe (~2s)."""
import pytest
from fastvideo.tests.performance import host_probe
def test_probe_reports_positive_normalized_profile(monkeypatch):
monkeypatch.setenv("PERF_HOST_CPU_ST_REF_KOPS", "1000")
profile = host_probe.measure_host_cpu_profile()
assert profile["host_cpu_single_thread_kops"] > 0
assert profile["host_cpu_multi_thread_gflops"] > 0
# The gated score is the single-thread axis normalized by the reference.
assert profile["host_cpu_score"] == pytest.approx(
profile["host_cpu_single_thread_kops"] / 1000.0, rel=0.05)
@@ -19,6 +19,7 @@ import pytest
from fastvideo import VideoGenerator
from fastvideo.logger import init_logger
from fastvideo.tests.performance.host_probe import measure_host_cpu_profile
from fastvideo.tests.performance.identity import (
benchmark_identity_from_config,
build_recipe_from_benchmark_config,
@@ -153,6 +154,24 @@ _BENCHMARK_CONFIGS = _discover_benchmarks()
# -- Helpers ----------------------------------------------------------------
# Probed once per pytest process (before the first benchmark's GPU work) and
# attached to every result: the comparator skips gating when the host is
# below the healthy-host CPU profile. Fail-open: a probe error just omits the
# fields, and the comparator gates normally.
_HOST_CPU_PROFILE: dict[str, float] | None = None
def _host_cpu_profile() -> dict[str, float]:
global _HOST_CPU_PROFILE
if _HOST_CPU_PROFILE is None:
try:
_HOST_CPU_PROFILE = measure_host_cpu_profile()
logger.info("Host CPU probe: %s", _HOST_CPU_PROFILE)
except Exception as e:
logger.warning("Host CPU probe failed (%s); records will carry no host_cpu_score", e)
_HOST_CPU_PROFILE = {}
return _HOST_CPU_PROFILE
def _get_thresholds(cfg):
"""Return thresholds dict for the current GPU from config."""
@@ -477,6 +496,7 @@ def _run_benchmark(cfg):
num_warmup, num_measure = _validate_run_counts(run_config, cfg["benchmark_id"])
thresholds = _get_thresholds(cfg)
host_cpu_profile = _host_cpu_profile()
# Remap JSON keys to VideoGenerator kwargs
text_enc_prec = init_kwargs.pop("text_encoder_precisions", None)
@@ -537,6 +557,9 @@ def _run_benchmark(cfg):
runtime_identity=runtime_identity,
device_name=device_name,
)
# Attach the host CPU probe so the comparator can skip gating on
# below-profile (packed) hosts. Fail-open: an empty probe adds nothing.
results.update(host_cpu_profile)
logger.info(
"Performance results: avg_time=%.2fs, "
@@ -14,6 +14,7 @@ from fastvideo.training.wan_training_pipeline import main
from fastvideo.fastvideo_args import FastVideoArgs, TrainingArgs
from fastvideo.utils import FlexibleArgumentParser
from fastvideo.training.wan_training_pipeline import WanTrainingPipeline
from fastvideo.tests.performance.host_probe import measure_host_cpu_profile
wandb_name = "test_training_loss"
a40_reference_wandb_summary_file = "fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json"
@@ -142,6 +143,31 @@ def test_distributed_training():
'train_loss': 0.0025
}
# Host-profile guard (mirrors compare_baseline.py's HOST_BELOW_PROFILE):
# a packed shared host inflates wall time (observed ~8x) with bit-clean
# loss, so the step-time gates are skipped loudly on a below-profile
# host. Loss and grad-norm stay asserted — they are host-independent.
# Fail-open: a probe error keeps the step-time gates active.
try:
min_score = float(os.environ.get("PERF_HOST_CPU_MIN_SCORE", "") or 0.75)
except ValueError:
min_score = 0.75
try:
host_cpu_score = measure_host_cpu_profile()["host_cpu_score"]
except Exception as e:
host_cpu_score = None
print(f"Host CPU probe failed ({e}); step-time gates stay active")
if host_cpu_score is not None and host_cpu_score < min_score:
print("=" * 72)
print(f"HOST_BELOW_PROFILE: host_cpu_score={host_cpu_score:.3f} < "
f"floor {min_score:.2f} — skipping the step-time gates "
"(avg_step_time, step_time). This is host CPU contention on "
"the shared runner, not a property of the PR. Loss values are "
"still asserted.")
print("=" * 72)
del fields_and_thresholds['avg_step_time']
del fields_and_thresholds['step_time']
failures = []
for field, threshold in fields_and_thresholds.items():
ref_value = reference_wandb_summary[field]