[ci]: add exact-identity perf comparator statuses

Comparator verdicts per issue #1532: PASS / REGRESSION /
CALIBRATION_NEEDED / RECIPE_MISMATCH / INFRA_ERROR (QUALITY_BLOCKED
reserved for the promoted-baseline workflow). REGRESSION and
RECIPE_MISMATCH fail CI; CALIBRATION_NEEDED is loud but no longer
auto-seeds a passing baseline.
This commit is contained in:
SolitaryThinker
2026-07-13 16:24:52 -07:00
parent c1abc42782
commit 04bfddb7fa
3 changed files with 333 additions and 45 deletions
+32 -13
View File
@@ -232,13 +232,25 @@ by the harness and is not config-declarable.)
Recipe fingerprinting, hardware/software profile IDs, exact-identity
comparison, and dashboard cohort grouping land with this change: v2 records
compare only within their identity cohort, and a record that opens a NEW
cohort is marked `baseline_status: "initialized_new_cohort"` (regression
gating starts once that cohort accumulates history). Legacy v1 configs still
run and are normalized for reporting, but their records skip rolling-baseline
comparison entirely (`baseline_status: "skipped_missing_identity"`, never
baseline eligible); only static thresholds gate them. Metric-specific
threshold policies and promoted baselines remain separate follow-ups.
compare only within their identity cohort and are never compared against v1
records (or vice versa). Legacy records missing any comparison identity field
still run and are normalized for reporting, but they skip rolling-baseline
comparison entirely (`baseline_status: "skipped_missing_identity"`, comparator
verdict `PASS`, never baseline eligible); only static thresholds gate them.
Each comparison reports an explicit `comparator_status` on the normalized
record and in the Markdown summary:
| Status | Meaning | CI |
|---|---|---|
| `PASS` | Comparable baseline found with no gated regression, or comparison was skipped for a record without the v2 identity block. | passes |
| `REGRESSION` | A gated metric regressed past both its percent and absolute floors. | fails |
| `CALIBRATION_NEEDED` | No comparable baseline exists. Gating is inactive and the record does **not** seed a baseline — seed new cohorts explicitly via the reseed workflow. The record also carries `baseline_status: "initialized_new_cohort"`. | passes |
| `RECIPE_MISMATCH` | Baseline history exists for the same variant/hardware/software cohort but under a different `recipe_fingerprint`. Represent recipe changes as a new `variant_id`, or reseed. | fails |
| `INFRA_ERROR` | The comparison itself failed (for example, baseline history could not be loaded). | fails |
| `QUALITY_BLOCKED` | Reserved for the promoted-baseline workflow; never emitted by the comparator. | n/a |
Metric-specific threshold policies and promoted baselines remain separate
follow-ups.
### Raw record (`results/perf_*.json`)
@@ -367,6 +379,8 @@ result, used as the rolling-baseline source of truth.
"build_id": "<buildkite-build-id>",
"job_id": "<buildkite-job-id>",
"quality_metadata": { "quality_status": "canonical" },
"baseline_status": "compared",
"comparator_status": "PASS",
"success": true
}
```
@@ -487,10 +501,10 @@ When the rolling-baseline phase runs, it emits:
2. The pytest test auto-discovers all configs — no test code needed. CI
picks it up on the next `/test performance` run.
3. The first persisted main-branch run with no HF history initializes the
baseline (passes automatically). Subsequent runs compare against it. Local
and pull-request runs with no HF history also pass, but they do not seed the
shared baseline.
3. The first run with no HF history reports `CALIBRATION_NEEDED`: it passes,
but it does not seed the shared baseline. Seed the cohort explicitly with
the `reseed-performance-baseline` skill; subsequent scheduled-main runs
then compare against it and keep the rolling baseline advancing.
4. If the benchmark targets a GPU not currently in `thresholds`, either add
that GPU as a key or rely on the `default` block. Note that `default` is
@@ -510,8 +524,13 @@ When the rolling-baseline phase runs, it emits:
## Troubleshooting
**"No baseline for ... Initializing"** — first run for this comparison cohort.
Run will pass and (if persisting) seed the first record.
**`CALIBRATION_NEEDED: no comparable baseline`** — first run for this cohort.
The run passes but does not seed a baseline; seed it explicitly with the
`reseed-performance-baseline` skill.
**`RECIPE_MISMATCH`** — the benchmark recipe changed without a new
`variant_id`. Either bump `variant_id` to open a new cohort, or reseed the
baseline if the existing variant should adopt the new recipe.
**Persistent failure right after a torch / kernel / image upgrade** —
genuine regression *or* baseline drift. Compare the failing normalized record
+108 -32
View File
@@ -89,6 +89,14 @@ COMPARISON_IDENTITY_KEYS = (
"hardware_profile_id",
"software_profile_id",
)
# Comparator verdicts (issue #1532), stored as record["comparator_status"].
STATUS_PASS = "PASS"
STATUS_REGRESSION = "REGRESSION"
STATUS_CALIBRATION_NEEDED = "CALIBRATION_NEEDED"
STATUS_RECIPE_MISMATCH = "RECIPE_MISMATCH"
STATUS_INFRA_ERROR = "INFRA_ERROR"
# Reserved for the promoted-baseline workflow; never emitted by this comparator.
STATUS_QUALITY_BLOCKED = "QUALITY_BLOCKED"
def _should_persist_tracking() -> bool:
@@ -343,6 +351,90 @@ def _check_regressions(
return failures
def _recipe_mismatch_fingerprints(
record: dict[str, Any],
identity_filters: dict[str, str],
) -> list[str]:
"""Return baseline recipe fingerprints for the same variant cohort.
Non-empty means the same (workload, variant, version, hardware, software)
cohort has baseline history under a DIFFERENT recipe fingerprint: the
recipe changed without being represented as a new variant.
"""
if not identity_filters:
return []
variant_filters = {key: value for key, value in identity_filters.items() if key != "recipe_fingerprint"}
same_variant = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**variant_filters,
successful_only=True,
baseline_eligible_only=True,
)
fingerprints = {str(r.get("recipe_fingerprint")) for r in same_variant}
return sorted(fingerprints - {identity_filters["recipe_fingerprint"]})
def _compare_record(
record: dict[str, Any],
metric_policies: tuple[MetricPolicy, ...],
) -> tuple[list[str], list[dict[str, Any]]]:
"""Compare one normalized record against its exact-identity cohort.
Sets ``record["comparator_status"]`` (and ``record["baseline_status"]``
for cohort bookkeeping) and returns ``(failures, baseline_records)``.
"""
model = record.get("model_id", "unknown")
try:
identity_filters = _comparison_identity_filters(record)
baseline_records = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**identity_filters,
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
if baseline_records:
record["baseline_status"] = "compared"
failures = _check_regressions(record, baseline_records, metric_policies)
record["comparator_status"] = STATUS_REGRESSION if failures else STATUS_PASS
return failures, baseline_records
mismatched = _recipe_mismatch_fingerprints(record, identity_filters)
except Exception as exc:
record["comparator_status"] = STATUS_INFRA_ERROR
return [f"{model} baseline comparison hit an infra error: {exc}"], []
if mismatched:
record["comparator_status"] = STATUS_RECIPE_MISMATCH
return [
f"{model} recipe fingerprint {record.get('recipe_fingerprint')} does not "
f"match baseline fingerprint(s) {', '.join(mismatched)} for variant "
f"{record.get('variant_id')}. Represent recipe changes as a new "
"variant_id, or reseed the baseline."
], []
# No comparable baseline anywhere: a brand-new cohort. Make that loud and
# machine-readable instead of an indistinguishable pass, so a cohort shift
# (intended or accidental, e.g. an identity-field change) never silently
# blinds the comparison — and never silently seeds a passing baseline.
record["baseline_status"] = "initialized_new_cohort"
record["comparator_status"] = STATUS_CALIBRATION_NEEDED
print("=" * 72)
print(f"CALIBRATION_NEEDED: no comparable baseline for {model} on "
f"{record.get('gpu_type', 'unknown')}"
f"{_format_identity_filters(identity_filters)}")
print("Regression gating is INACTIVE for this record and it will NOT "
"seed a baseline. Seed the cohort explicitly via the reseed "
"workflow. If this cohort shift is unexpected, check the identity "
"fields above.")
print("=" * 72)
return [], []
def _compact_value(value: float | None, precision: int = 3) -> str:
if value is None:
return "n/a"
@@ -394,6 +486,7 @@ def _build_summary_row(
return {
"model_id": record["model_id"],
"gpu_type": record["gpu_type"],
"comparator_status": record.get("comparator_status", STATUS_PASS),
"baseline_n": len(baseline_records),
"metrics": metric_values,
"worst_regression_pct": worst_regression_pct,
@@ -435,7 +528,10 @@ def _build_markdown_summary(
else "none"
)
failing_metrics = ", ".join(row["failing_metrics"]) if row["failing_metrics"] else "none"
status = "FAIL" if row["failed"] else "PASS"
status = row["comparator_status"]
if row["failed"] and status == STATUS_PASS:
# Baseline comparison passed but the fixed-threshold phase failed.
status = "FAIL"
lines.append(f"| {row['model_id']} | {row['gpu_type']} | "
f"{row['baseline_n']} | "
@@ -501,46 +597,26 @@ def main() -> int:
if identity_filters is None:
# Records without the full v2 identity block skip rolling-baseline
# comparison entirely: only the static thresholds gate them and
# they never become baseline eligible.
# they never become baseline eligible. Nothing was compared and
# nothing failed, so the comparator verdict is PASS;
# baseline_status keeps the skip machine-readable.
record["baseline_status"] = "skipped_missing_identity"
record["comparator_status"] = STATUS_PASS
failures: list[str] = []
else:
baseline_records = load_records_for_model(
TRACKING_ROOT,
record["model_id"],
record["gpu_type"],
**identity_filters,
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
if not baseline_records:
# A brand-new cohort has NO regression gating until history
# accumulates — make that loud and machine-readable instead of
# an indistinguishable pass, so a cohort shift (intended or
# accidental, e.g. an identity-field change) never silently
# blinds the comparison.
record["baseline_status"] = "initialized_new_cohort"
print("=" * 72)
print(f"WARNING: NO BASELINE — initializing a NEW cohort for "
f"{record['model_id']} on {record['gpu_type']}"
f"{_format_identity_filters(identity_filters)}")
print("Regression gating is INACTIVE for this cohort until "
"baseline history accumulates. If this cohort shift is "
"unexpected, check the identity fields above.")
print("=" * 72)
failures = []
else:
record["baseline_status"] = "compared"
failures = _check_regressions(record, baseline_records, metric_policies)
failures, baseline_records = _compare_record(record, metric_policies)
if static_threshold_failed:
failures.append(f"{record['model_id']} fixed-threshold phase failed "
f"(PERF_PYTEST_RC={os.environ.get('PERF_PYTEST_RC')})")
record["success"] = not failures
# Only compared-and-PASS records may advance the rolling baseline: a
# CALIBRATION_NEEDED record must never silently seed a new cohort, and
# records without the v2 identity block never become eligible.
record["baseline_eligible"] = (
identity_filters is not None and _is_baseline_eligible(record["run_source"], record["success"])
identity_filters is not None
and record["comparator_status"] == STATUS_PASS
and _is_baseline_eligible(record["run_source"], record["success"])
)
all_failures.extend(failures)
@@ -421,3 +421,196 @@ def test_informational_metric_remains_visible_without_failing():
assert row["metrics"]["throughput"]["regressed"] is False
assert row["threshold_exceeded_metrics"] == ["throughput"]
assert row["failing_metrics"] == []
def _v2_raw_result(**overrides):
raw = _raw_result()
raw.update({
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"recipe_fingerprint": "recipe-1",
"hardware_profile_id": "hw-1",
"software_profile_id": "sw-1",
})
raw.update(overrides)
return raw
def _v2_baseline_record(**overrides):
record = {
"gpu_type": "NVIDIA L40S",
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"recipe_fingerprint": "recipe-1",
"hardware_profile_id": "hw-1",
"software_profile_id": "sw-1",
"latency": 10.0,
"throughput": 4.5,
"memory": 10000.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
"run_source": "scheduled_main",
"baseline_eligible": True,
}
record.update(overrides)
return record
def _run_compare(monkeypatch, tmp_path, raw_result, baseline_records):
"""Run compare_baseline.main() against a local tracking root.
Returns (exit_code, normalized_record, markdown_report).
"""
results_dir = tmp_path / "results"
results_dir.mkdir()
(results_dir / "perf_current.json").write_text(json.dumps(raw_result))
model_dir = tmp_path / "tracking" / raw_result["benchmark_id"]
model_dir.mkdir(parents=True)
for index, record in enumerate(baseline_records):
(model_dir / f"rec{index}.json").write_text(json.dumps(record))
reports_dir = tmp_path / "reports"
monkeypatch.setattr(compare_baseline, "RESULTS_DIR", str(results_dir))
monkeypatch.setattr(compare_baseline, "TRACKING_ROOT", str(tmp_path / "tracking"))
monkeypatch.setattr(compare_baseline, "PERF_REPORTS_DIR", str(reports_dir))
monkeypatch.setattr(compare_baseline, "UPLOAD_POLICY", "never")
monkeypatch.setattr(compare_baseline, "sync_from_hf", lambda *args, **kwargs: None)
monkeypatch.setenv("PERF_RUN_SOURCE", "scheduled_main")
monkeypatch.delenv("PERF_PYTEST_RC", raising=False)
monkeypatch.delenv("GITHUB_STEP_SUMMARY", raising=False)
exit_code = compare_baseline.main()
normalized_paths = list((reports_dir / "results").glob("normalized_perf_*.json"))
assert len(normalized_paths) == 1
markdown = "\n".join(path.read_text() for path in reports_dir.glob("perf_*.md"))
return exit_code, json.loads(normalized_paths[0].read_text()), markdown
def test_main_pass_with_comparable_baseline(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "compared"
assert record["success"] is True
# Scheduled-main PASS records keep advancing the rolling baseline.
assert record["baseline_eligible"] is True
assert "| PASS |" in markdown
def test_main_regression_on_gated_metric_fails_ci(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(avg_generation_time_s=20.0),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "REGRESSION"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| REGRESSION |" in markdown
def test_main_missing_baseline_is_calibration_needed_and_does_not_seed(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
assert record["baseline_status"] == "initialized_new_cohort"
assert record["success"] is True
# Visible, but never silently seeds a passing baseline.
assert record["baseline_eligible"] is False
assert "| CALIBRATION_NEEDED |" in markdown
def test_main_recipe_change_without_new_variant_is_recipe_mismatch(monkeypatch, tmp_path):
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(recipe_fingerprint="recipe-2"),
[_v2_baseline_record()])
assert exit_code == 1
assert record["comparator_status"] == "RECIPE_MISMATCH"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| RECIPE_MISMATCH |" in markdown
def test_main_recipe_change_as_new_variant_is_calibration_needed(monkeypatch, tmp_path):
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path,
_v2_raw_result(variant_id="1.3b-sp2-r2", recipe_fingerprint="recipe-2"),
[_v2_baseline_record()])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
def test_main_v2_record_never_compares_against_v1_baselines(monkeypatch, tmp_path):
legacy_v1_baseline = {
"gpu_type": "NVIDIA L40S",
# Would be a >100% latency regression if the comparator (wrongly)
# matched the v2 record against v1 history.
"latency": 1.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
}
exit_code, record, _ = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [legacy_v1_baseline])
assert exit_code == 0
assert record["comparator_status"] == "CALIBRATION_NEEDED"
def test_main_legacy_v1_record_skips_comparison_with_pass_verdict(monkeypatch, tmp_path):
legacy_v1_baseline = {
"gpu_type": "NVIDIA L40S",
# Would be a >100% latency regression if the comparator (wrongly)
# compared the identity-less v1 record against this history.
"latency": 1.0,
"timestamp": "2026-06-15T00:00:00+00:00",
"success": True,
}
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _raw_result(), [legacy_v1_baseline])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "skipped_missing_identity"
assert record["baseline_eligible"] is False
assert "| PASS |" in markdown
def test_main_partial_identity_skips_comparison(monkeypatch, tmp_path):
raw = _v2_raw_result()
del raw["software_profile_id"]
exit_code, record, _ = _run_compare(monkeypatch, tmp_path, raw, [])
assert exit_code == 0
assert record["comparator_status"] == "PASS"
assert record["baseline_status"] == "skipped_missing_identity"
assert record["success"] is True
assert record["baseline_eligible"] is False
def test_main_baseline_load_failure_is_infra_error_and_fails_ci(monkeypatch, tmp_path):
def _boom(*_args, **_kwargs):
raise OSError("hf store unavailable")
monkeypatch.setattr(compare_baseline, "load_records_for_model", _boom)
exit_code, record, markdown = _run_compare(
monkeypatch, tmp_path, _v2_raw_result(), [])
assert exit_code == 1
assert record["comparator_status"] == "INFRA_ERROR"
assert record["success"] is False
assert record["baseline_eligible"] is False
assert "| INFRA_ERROR |" in markdown