Scoring

Two boards, scored separately against different baselines. Both are in draft and neither is scoring yet.

Board one: signal beyond demographics

What is compared

Players are compared in pairs, and a pair is formed only when all of these hold:

  • the same broad role
  • the same competition
  • ages differing by at most two years
  • both have a value in force before the window, and appearances after it

Within such a pair, age, position and league are fixed. The more valuable player is predicted to perform better, and the outcome is what he actually did.

Why the null is exactly zero

Because the stratification holds the demographic terms fixed, a value carrying no quality information orders players at chance. Zero means no signal beyond demographics — and there is no baseline series to argue about, because the pairing itself is the baseline.

Why two years, and not age bands

Bands are the obvious choice and are worse. Four bands would pair a 22-year-old with a 25-year-old while refusing 21-with-22 across a boundary, and yield fewer comparisons than a window that never does either. The window is the primary; a stricter one-year window is published as a sensitivity.

The outcome is two claims, not one

OutcomeWhat it measures
QualityPercentile across a player's appearances, on the metric mask for his role
AvailabilityMinutes played as a share of his team's available minutes

A valuation predicts both, and they can disagree: a player who is excellent when fit and never fit is a bad valuation and a good quality score. Blending them would need a weighting nobody can defend, so both are published.

A player with no appearances has no quality outcome, and those are disproportionately the ones a valuation got most wrong. The availability board catches them, and the share of a participant's players absent from the quality board is published beside its score.

Board two: fee prediction

The lead-time curve

A value taken a week before a transfer knows about the deal. One taken a year before cannot. So the board does not score one moment — it scores four, and publishes the curve:

Lead timeA score here means
1 monthThe value is tracking a deal already being reported
3 monthsIt anticipates the window
6 monthsIt is pricing the player rather than the transaction
12 monthsIt is a genuine forecast

The headline is the 6-month score. A participant whose skill collapses between three months and six is reacting to news; one whose skill holds across the curve is valuing the player.

The baseline is the value it replaced

A revaluation is scored against the participant's own previous value for that player. Where there is no history — a newly tracked player — a demographic stand-in is used instead.

This asks a harder question than a table of median prices would, and a fairer one:

Did revaluing this player add anything to what you already said about him?

No participant can call that unfair, because the thing to beat is itself. It does mean a participant that never revalues scores exactly zero by construction, so the score against a median-price baseline is published beside it, and the two answer different questions.

What is excluded, and counted

  • Undisclosed fees — there is no outcome to score.
  • Free transfers — a fee of zero is a statement about an expiring contract, not about a player's worth. They are outside the evaluation entirely.
  • Add-ons and contingent money — the scored figure is the base fee, so a score does not depend on how a deal was structured.

Each exclusion is published with its count.

Where the fee comes from

A fee is not taken from a source. It is established by agreement between sources, under a published criterion, and a transfer whose sources disagree is unresolved and unscored rather than settled by preference.

A participant may be one of those sources. The safeguard is the criterion rather than exclusion: every participant's score is published a second time on fees established without that participant as a source, so a criterion that is too weak is visible on the row instead of being argued about.

The statistic

Both boards use Somers' D, with Kendall's tau-b published beside it, as everywhere else in the Arena. It is a rank measure, so it is unaffected by the scale a participant reports on and cannot be dominated by a handful of very large transfers.

Intervals come from a bootstrap that resamples players — or transfers on the fee board — never pairs, since one player appears in many pairs and they are not independent observations of him.

Two participants are compared by a paired bootstrap on their difference, not by checking whether their published intervals overlap. Facing identical players, their errors move together, so the uncertainty in the difference is much smaller than in either score.

Reading a null result

A score of zero and a score that cannot be measured are different findings, and the interval is what separates them.

  • A tight interval around zero. The evidence was there, and the signal is not.
  • A wide interval around zero. Not yet measurable — which is what the board says, rather than the stronger claim a reader might hear.

Noise in football performance attenuates a real signal toward zero and never past it: it can hide a result, but it cannot manufacture one.