Glossary

The terms a board uses, defined once. Several of them are ordinary words doing specific work, and reading them loosely is the fastest way to misread a score.

Case

One transfer. A player, a move, and a destination — the unit a participant is asked about and the unit a score is built from.

A case is not a player: the same player transferred twice is two cases. It is not a match either; one case accumulates many appearances over a season, and all of them together produce the one realised distribution that case is scored against.

Cases are weighted equally in a participant's score, regardless of how many appearances each accumulated. An ever-present signing does not outweigh the rest of the cohort.

Appearance

One player in one match. Every appearance counts, with no minutes floor: playing time is already expressed in the percentile a short appearance lands on, and a substitute's twenty minutes produces a low percentile that is the correct observation rather than noise.

Values are raw per-match totals, never per-90 rates.

Cohort

Every case for one competition and season. A cohort's counts are properties of the benchmark, identical for every participant, and are published once rather than per row:

  • Candidate — cases in the registry.
  • Void — the player transferred again inside the observation window, so he can no longer be observed at the club he was predicted for. Voided for everyone.
  • Unobserved — signed, resolved, and has not played. A football outcome, not a pipeline failure.
  • Scorable — has recorded at least one appearance. This is the denominator for coverage.

Coverage

Scored cases over scorable cases, per participant.

Both a refusal and a case never answered reduce it. Coverage exists so that declining is visible and costed without a refusal being scored as a wrong answer — a participant answering 40 of 110 cases very well should not read as a better model than one answering all 110 nearly as well.

Refusal

A participant declining a case, with HTTP 422. All-or-nothing: every metric is answered, or the case is refused.

A refusal is not scored as a maximally wrong prediction. A model that knows a case is outside its scope is behaving correctly, and scoring it as wrong would make good judgement indistinguishable from bad modelling. It costs coverage instead.

Realised distribution

What actually happened, as a distribution over the same bands the prediction used. Each appearance is banded, and the histogram of those bands across every appearance to date is what a prediction is scored against.

It is rebuilt from full history every week, not accumulated. So a score is always the whole season to date, and adding a week cannot corrupt what earlier weeks measured.

Ranked probability score

RPS — the distance between the predicted distribution and the realised one, comparing their cumulative distributions band by band. Bounded in [0, 1]; lower is better.

Two properties matter. It is order-aware, so being one band out costs less than being five bands out. And it is proper, meaning a participant minimises its expected score by reporting what it actually believes — there is no shading that improves the expected result.

Skill score

1 − participant RPS ÷ baseline RPS, where the baseline is the league base rate for that role and metric.

  • Zero — level with the league base rate. Knowing nothing about the individual player.
  • Above zero — the participant knew something about this specific player the league average did not.
  • Below zero — worse than knowing nothing. Published as negative, never clamped: a confidently wrong model belongs below a shrug.

The baseline is reachable by anyone. A band's width is its share of peer appearances, so predicting the widths of the grid you were sent is the base rate, which is what makes zero a fair floor rather than an insider's number.

Interval

The 95% range around a published skill score, from a bootstrap over cases.

Early in a season the interval is the story. Ten cases with three appearances each cannot separate a good model from a lucky one, and a board that showed only point estimates would invite a conclusion the data does not support.

Cases are resampled, never appearances: appearances within a case share a player and are not independent evidence, so one signing is one unit of evidence however often he plays.

Confidence discrimination

Whether a participant's stated confidence tracked its realised error — a rank correlation between confidence and -RPS across its scored cases.

Published as its own ranked column, never folded into the skill score. A participant that sends no confidence shows no data, never zero, and its accuracy ranking is untouched.

A constant confidence, high or low, carries no rank information and scores no data. Asserting certainty gains nothing.

Provisional

A row shown but not ranked, because it sits below the minimum number of observed appearances.

Provisional is not a judgement about the participant. It says the evidence is too thin to place the row, which early in a season is true of everyone.

Version pins

Every published snapshot records the case registry, entity crosswalk, reference pool, scoring metric and contract versions it was computed under.

That is what makes a score citable weeks later: the rules that produced it are named, frozen, and readable, rather than being whatever the current documentation happens to say.