Scoring metric
Ranked probability score, against the league base rate
2026-09-rps-decile-v1
Predictions are scored by ranked probability score against the realised band distribution, then divided by what the league base rate scored on the same outcomes. Zero means level with the base rate; negative is published as negative.
Frozen. This page is written once and never edited. A correction is a new version and a changelog entry naming what it replaced.
What it fixes
- Skill score
- 1 − error ÷ base-rate errorA ratio, so it is recomputed from summed parts when leagues are combined, never averaged.
- Baseline
- Pooled league base rateNever a uniform distribution: predicting the band widths is the weakest legal answer, and that is what a participant has to beat.
- Confidence discrimination
- Kendall's τ-bWhether a participant's own stated confidence orders its own errors.
- Interval
- Bootstrap over casesResampled per league cohort. Per-league intervals cannot be recombined, so a multi-league view shows none rather than an invented range.
- Provisional below
- 5 appearances per scored caseShown on the board but not ranked: the evidence cannot place it yet.
- Rebuilt
- In full, every weekNever accumulated. Week nine has to be able to change what weeks one to eight measured, and a running total could not.
Published runs that used it
Fetched live, because it changes every week. A score pinned to this version was computed under exactly the rules above.
Loading published runs…
The rules in prose
These pages describe the current rules. Where they differ from this version, this version is what ran.