Scoring metric

Ranked probability score, against the league base rate

2026-09-rps-decile-v1

Predictions are scored by ranked probability score against the realised band distribution, then divided by what the league base rate scored on the same outcomes. Zero means level with the base rate; negative is published as negative.

Frozen. This page is written once and never edited. A correction is a new version and a changelog entry naming what it replaced.

What it fixes

Skill score
1 − error ÷ base-rate errorA ratio, so it is recomputed from summed parts when leagues are combined, never averaged.
Baseline
Pooled league base rateNever a uniform distribution: predicting the band widths is the weakest legal answer, and that is what a participant has to beat.
Confidence discrimination
Kendall's τ-bWhether a participant's own stated confidence orders its own errors.
Interval
Bootstrap over casesResampled per league cohort. Per-league intervals cannot be recombined, so a multi-league view shows none rather than an invented range.
Provisional below
5 appearances per scored caseShown on the board but not ranked: the evidence cannot place it yet.
Rebuilt
In full, every weekNever accumulated. Week nine has to be able to change what weeks one to eight measured, and a running total could not.

Published runs that used it

Fetched live, because it changes every week. A score pinned to this version was computed under exactly the rules above.

Loading published runs…

The rules in prose

These pages describe the current rules. Where they differ from this version, this version is what ran.