Game Value — Player board
Does your rating of a player tell us more about what happens while he is on the pitch than his own on-pitch goal difference already does?
That is the whole claim a player rating makes: that it separates a player's contribution from the crude fact of what happened around him. It is what this board measures, and nothing else.
A player rating is the most widely shipped football analytics product there is — live-score apps, media outlets and scouting platforms publish one number per player per match to millions of people every week — and there is almost no public evidence about which of them is any good. Any player-rating product can enter this board, with or without a game value model behind it.
This benchmark is in draft and is not scoring yet. The protocol is published before any score exists, so that it cannot be changed to suit a result.
Why not whole matches
The obvious test compares a team's results with a player and without him. Across whole matches that compares different matches — different opponents, different venues, and a squad that changed for the same reasons the player is missing.
Three segments of the same match
Each match is cut into three segments, and each is scored as one comparison between the two sides:
| Segment | Window |
|---|---|
| 1 | kick-off → half-time, with first-half stoppage time |
| 2 | half-time → 70' |
| 3 | 70' → full time, with second-half stoppage time |
A player withdrawn on 68 minutes gives a with-him and a without-him observation inside the same match, against the same opponent, on the same afternoon. Everything that makes two matches hard to compare is held fixed.
The boundaries sit where the lineup actually changes. Measured on a full season, substitutions fall 42.8% between half-time and 70' and 54.8% after it, with a median entry at the 71st minute, and only about one in a hundred arrives before the 25th minute — so a cut there would only separate the same eleven players from themselves. Between one segment and the next, the players on the pitch change in 69.3% of cases. That is what gives the comparison something to compare.
Who counts as on the pitch
A player is on the pitch for a segment when he is present for more than half of it. Occupancy is taken from timed match events — the starting eleven, every substitution and every sending-off — and is exact: eleven players on the pitch, less any sent off, at every minute of all but one of 3,504 team-matches in a season.
A side reduced to ten is not a data fault. It is scored with ten, because discarding matches with a red card would discard the matches where who is on the pitch matters most.
What is scored
The signal
For each segment, the average rating of the players on the pitch, each rating taken from that player's appearances before this match. The signal is the difference between the two sides.
That is a hybrid of two times, and deliberately so: who is on the pitch comes from this segment, what they are worth comes from before this match. Price the players on this match's ratings and the test becomes a restatement of what just happened; take the eleven from before the match and the substitutions — the whole source of the comparison — disappear.
The outcome
The goal difference during that segment. It is sparse: a side scores 0.46 goals in an average segment, and nearly half of all segments end level. The statistic handles ties on both sides, so none is discarded — but it is why this board's intervals are wider per season than the team board's.
The baseline: naive plus-minus
Each player's own on-pitch goal difference per segment so far, averaged over the players on the pitch exactly as the ratings are. It is the demanding choice, for three reasons.
- It neutralises the outcome mimic. A rating that quietly rewards whoever scored tracks plus-minus almost by construction, so it scores near zero however convincing it looks.
- It absorbs team strength. Good players play for good teams, so a signal built from the eleven on the pitch mostly says which team is which. Plus-minus carries exactly that, and subtracting it leaves only what the rating adds about the player.
- It is the honest competitor. Plus-minus is what the lineup sheet and the scoreline give for free. A rating that does not beat it is not selling information.
A lineup-blind baseline — the team's record alone — is published as a floor beside the score, because a rating that cannot beat even that has no player signal at all.
Players the baseline cannot see. A debutant has no plus-minus, while a rating may have a real opinion of him. That is a legitimate advantage — a rating that can judge a player plus-minus cannot is genuinely better — but it could also come from squad churn rather than judgement. So the share of on-pitch players without enough plus-minus history is published beside the score, and the skill score is published a second time restricted to segments where every player has it.
Comparing like with like
Substitutions are decided by the scoreline: attackers on when chasing, defenders on when protecting a lead. Left alone, being on the pitch late in a match would measure when a player tends to come on rather than what he contributes.
So segments are compared only with segments in the same situation — the same segment of the match, and the same score at its start, from two or more down to two or more up. A player introduced at 70' with his side two down is measured against other segments that also began two down.
Two views, both published
| View | Compares | What it holds fixed |
|---|---|---|
| Across fixtures | All segments in the same situation | The score state, exactly; team strength through the baseline |
| Within the fixture | Only a team's own segments in the same match | Team, opponent, venue and afternoon, exactly — but not the score, which is what changes between segments |
Neither view controls everything, and they control complementary things. A rating that beats plus-minus under both is evidence; one that beats it under only one is a lead, and the board lets a reader tell the difference rather than choosing on their behalf.
Reading the score
The statistic is Somers' D, as on the team board: of the pairs of segments where one genuinely went better than the other, the net share the ratings put in the right order. With this many level segments, about 29% of comparable pairs are tied on the outcome, so the older Kendall's τ-b could never exceed about 0.84 here. Somers' D has a ceiling of 1 on every board, so an improvement of +0.02 means the same thing here as on the team board.
The scale does not matter. Ratings arrive out of 10, out of 100, or as z-scores; the score depends only on order, so nothing is rescaled.
Every score carries an interval, and the intervals here will be wider than the team board's: half the outcomes are level, and one match carries three linked segments. The board says so beside the score rather than implying a precision it does not have.
Taking part
A participant sends one rating per player per appearance — not one per segment. Requiring per-segment ratings would exclude exactly the products this board exists to measure. A game value vendor sends its own player ratings rather than having the Arena derive them from its action values: how value is shared out among players is a modelling choice, and the vendor is judged on it.
A rating counts only toward segments of later matches, never the match it rates.
Still to be decided
This board's protocol is in draft, and these are open:
- whether prior ratings are weighted by minutes played, or a simple average of appearances;
- how much plus-minus history a player needs before the baseline uses it;
- how many of the players on the pitch must be rated for a segment to count;
- whether goalkeepers are included in the average;
- whether the board leads with the across-fixtures view or the within-fixture one.
What this does not claim
It does not measure a rating against what a match report said, or against who was named man of the match — neither has an answerable form. It measures one thing: whether a rating tells us more about what happens while a player is on the pitch than his own on-pitch goal difference already does.
It is also independent of the other two boards. A participant may enter any of them, and none checks a submission against another.