How to read a leaderboard

A ranked table invites a quick conclusion. This page is about the places where the quick conclusion is wrong.

The columns

ColumnWhat it is
Skill score1 − error ÷ base-rate error. Zero is level with the league average for the role.
CoverageScored cases ÷ scorable cases. Refusals and unanswered cases both reduce it.
Scored casesHow many transfers this participant is actually being measured on.
AppearancesHow many matches sit behind those cases. This is the evidence weight.
ConfidenceWhether higher stated confidence went with lower error. Its own ranking.
SeasonThe skill score at each published week, oldest to newest.

Above the table, the cohort strip counts the benchmark itself rather than any participant.

Zero is the interesting number

Not one, and not the top of the table.

Zero means a participant did no better than predicting the league base rate for the position — which is a real thing to know, freely available, and requires no model at all. Above zero is the whole claim. A participant at +0.21 is extracting something about individual players that the league average does not contain.

That also makes the scale unintuitive. Skill scores here are not percentages and not accuracy. +0.21 is a strong result, not a poor one.

Negative is shown as negative. It means worse than a table of league averages, which is a real outcome for a confident model pointed at the wrong cases, and clamping it to zero would hide the difference between a bad model and a cautious one.

The four misreadings

1. Reading the ranking before the interval

Early in a season the interval is the story and the point estimate is barely a story at all. Ten cases with three appearances each cannot separate a good model from a lucky one, and the interval says so.

If two intervals overlap heavily, the board is telling you it cannot yet order those two rows — even though it has to print them in some order.

2. Reading overlap as "no difference"

The reverse trap, and the subtler one. Two overlapping intervals do not mean two participants are indistinguishable.

Participants face identical cases, so their errors are correlated, and the difference between two scores is far better determined than either score. On a real twelve-case cohort, two participants with heavily overlapping intervals — 0.498 [0.387, 0.594] and 0.454 [0.334, 0.559] — had a paired difference of 0.044 [0.035, 0.053]: a clear separation, twelve times tighter than either interval.

So: overlapping intervals mean this board cannot tell you from the columns alone. It does not mean the two are equal.

3. Reading a high score without its coverage

A participant answering 40 of 110 cases very well is not a better model than one answering all 110 nearly as well — it is a model that was allowed to choose.

Refusing hard cases is legitimate and is not scored as error. It costs coverage instead, which is exactly why coverage sits next to the score. Read the two together, always. A high score at 35% coverage describes a narrow model, and that may be precisely what you want — but it is a different claim from a high score at 95%.

4. Reading a provisional row as ranked

Provisional rows are shown but never ranked, and they sort below every ranked row whatever column you sort by.

It is not a judgement about the participant. It says the evidence behind that row is too thin to place it — usually because it entered recently, so its cases have accumulated few appearances. Early in a season this is true of everyone.

Entry point matters

Registration is open at any time. A participant that entered in November predicted with more of the season already observable than one that entered in August.

They did not answer the same question, and the board is built to show that rather than hide it:

  • The coverage column carries it. A late entrant is offered only the cases still open when it registered, so its coverage is lower by construction.
  • The season column carries it. Its trajectory starts partway through, with a visible gap where it was not yet scored.

A late entrant with high coverage would be the surprising thing, not a late entrant with low coverage.

Confidence is a separate ranking

It answers a different question: does this participant know when it does not know?

It is never folded into the skill score. A well-calibrated but inaccurate model and an accurate but overconfident one are different products, and collapsing them into one number would serve neither.

"No data" is not zero. A participant that sends no confidence has made no claim about its own certainty, which is different from having made a badly calibrated one. It sorts last in both directions, because floating an absent claim to the top of an ascending sort would read as the worst-calibrated participant on the board.

A constant confidence also scores no data. Asserting certainty on everything carries no ordering information and gains nothing.

The cohort strip counts the benchmark, not the participants

Those four numbers are identical for everyone, which is why they are published once above the table rather than repeated on every row.

  • Cases in registry — every eligible transfer into the covered leagues.
  • Void — the player was transferred again inside the window, so he can no longer be observed at the club he was predicted for. Voided for everyone.
  • Unobserved — signed, resolved, and has not played. A football outcome, not a pipeline failure.
  • Scorable — has at least one appearance. This is the coverage denominator.

The gap between the first and the last is real attrition, and it is shown rather than quietly absorbed. A board reporting only the scorable count would be hiding that far more players were signed and that many of them have not yet kicked a ball.

Filtering by league

The filter selects destination competition — the league a player moved to, which is the league his percentiles are computed against.

Two things change when you filter:

  • The skill score is recomputed for that scope, weighted by cases. It is not an average of per-league scores: a ratio recombined by averaging would weight a 60-case cohort like a 120-case one.
  • The interval appears. Intervals come from resampling cases within one cohort and cannot be recombined across leagues without the case-level scores, so the board shows one only when a single league is in scope, rather than inventing a range.

Version pins

Under every board is the set of versions the numbers were computed under: contract, case registry, scoring metric, reference pools.

They are there because a score is only meaningful alongside the rules that produced it. A published snapshot is never rewritten, so a number from week 3 still reads the same in week 30 — and still names the rules that were in force when it was produced, not the ones in force when you read it.