# Scoring rubric

## Silver-release rubric

`100-talk-v1` is ranked only on disagreement and operations. It does not apply
the human-gold gates below. A system is eligible for an aggregate row only when:

- all submitted audio checksums match the frozen cohort;
- case coverage and every failure/retry are reported;
- substitutions, insertions, and deletions are retained separately;
- word-timestamp coverage is reported before drift among aligned words;
- speaker-count and cluster-mapped attribution are explicitly disagreements
  with anonymous recording-local Scribe clusters;
- actual billed cost is separated from rate-card estimates, and self-hosted
  estimates name hardware, runtime, and pricing assumptions;
- macro, micro, per-stratum, and worst-decile results are available; and
- the UI calls Scribe a silver reference and never labels its self-score as
  evidence of correctness.

Provider completion and artifact integrity are hard gates. Silver disagreement
is descriptive and has no pass threshold because the reference can be wrong.
The later human-gold gates are the only quality pass/fail thresholds.

Passing a hard gate is required before weighted ranking.

## Weighted score

| Category | Weight |
| --- | ---: |
| Word error rate | 25% |
| Technical named-entity quality | 15% |
| Diarization, speaker count, and identity | 25% |
| Timestamp coverage, drift, and seek correctness | 15% |
| Summary and chapter factuality | 10% |
| Latency, cost, and operations | 10% |

## Standalone-talk gates

- Normalized WER ≤ 5%.
- Technical named-entity error rate ≤ 2%.
- DER ≤ 5% and JER ≤ 10%.
- Exact speaker count.
- Exact source-token timing coverage ≥ 99.5%.
- Word midpoint absolute-error p95 ≤ 300 ms.
- At least 99% of deep links land inside the cited utterance or within 500 ms before it.
- No unsupported summary claims.

## Panel and interview gates

- Normalized WER ≤ 8%.
- Technical named-entity error rate ≤ 3%.
- DER ≤ 10% and JER ≤ 15%.
- Exact speaker count for the scored recording.
- Exact source-token timing coverage ≥ 99%.
- Word midpoint absolute-error p95 ≤ 500 ms.
- Overlap miss rate ≤ 20%.
- No unsupported summary claims.

## Operational gates

- Every public transcript can be regenerated from a source checksum, run manifest, and immutable normalized artifact.
- Corrections are overlays with reviewer, reason, timestamp, and previous-artifact hash; raw outputs are not overwritten.
- Provider aliases without immutable revisions are labeled as such.
- Paid acquisition and GPU inference are optional/manual; checked receipts and regression scoring remain offline and credential-free.
- Pricing is snapshot-dated and excluded from reproducibility claims.

## Current case

The initial `CEvIs9y1uog` case cannot pass or fail these quality gates because human transcript, entity, RTTM/UEM, and word-boundary gold do not exist. WER, technical entity error rate, DER, JER, and absolute timing error must remain `null` until that annotation is complete.
