Game Value — Team board
Does the value you assign to what teams did carry information about results that the results themselves do not already carry?
Game value models — expected threat, VAEP, on-ball and possession-value models, and the in-house variants clubs and broadcasters run — put a number on what each part of a match was worth. They disagree with each other constantly, and there is no public evidence about which of them is right.
Game Value has three boards, scored separately: team, player and element. This page is the team board; the player and element boards have their own.
The team board is open. The fixture list is published, values are accepted, and the scoring run is weekly. The protocol was published before any score existed, so that it could not be changed to suit a result — and no participant appears on the board until it has been scored. Onboarding is where to start.
Why agreeing with the scoreline is not evidence
The obvious test of a game value model is whether it agrees with what happened: the home side won 3–0, so did the model give it more value?
That test is worthless, because its perfect score is trivial. A model that ignores football and returns team value = goals scored agrees with every scoreline, in every match, forever.
Nor is the scoreline a weak rival to beat. Goal difference per match played is one of the strongest cheap predictors in football — it is what decides league tables — so a copy of the scoreline would look competitive on almost any test built around results.
What is measured instead: the increment
The claim every game value vendor makes is not "my numbers agree with the result". It is "my numbers see something the result does not" — that a side was better than its 1–0 suggested, that a team creating this much will start winning. The team board measures exactly that claim:
A model is scored on how much better it predicts results than the scoreline already does.
A model that copies the scoreline therefore scores exactly zero — not "is flagged", not "is detectable", but zero by construction, because what it copied is what is being subtracted.
The forecast
For each fixture, each side's net value per match played so far — its value minus the value its opponents generated against it — is compared across the two teams. The side with the higher rate is the model's pick, and the size of the gap is how strongly it picks.
- Per match played, never a running total, so a team with a game in hand is not penalised.
- Only matches known before kick-off. A fixture is forecast from what was known before it, and never from itself.
- Value conceded is never asked for. In a match between A and B, the value against A is B's value.
The baseline
The identical forecast, built from goals instead of values: each side's goal difference per match played so far. It stands for everything the results already say. The skill score is the model's forecast quality minus the baseline's, over the same fixtures.
Three horizons
Each fixture is forecast three times: from every earlier match of both teams, then with the most recent withheld, then with the last two withheld. The three are scored separately and averaged.
That does two things. It stops a single freak weekend dominating, and the drop from the first horizon to the third is published as a diagnostic — a model whose score falls fast is measuring current form, one whose score holds is measuring strength, and both are legitimate products.
Draws count
A quarter of all matches are draws, and none is discarded. The statistic handles ties on both sides, so a draw is an observation like any other.
Reading the score
The statistic is Somers' D, a ranking measure over pairs of fixtures: of the pairs where one match genuinely went better than the other, the net share the model put in the right order. Its ceiling is 1, and it is published beside Kendall's τ-b for readers who know that one.
Scale does not matter. Values arrive on completely different scales — expected-threat sums, action values, goal equivalents — and the score depends only on order, so nothing is normalised and no participant can move its score by rescaling.
Every score carries an interval. Each skill score is published with a 95% interval, and two participants are compared on the same resampled fixtures rather than by checking whether their intervals overlap — which would call genuinely different models a tie.
Expect small numbers. A real improvement over the scoreline is likely to be a few hundredths, and a first season can only distinguish so much. The board says so beside the score rather than implying a precision it does not have.
Published beside the score, never ranked
| Diagnostic | What it shows |
|---|---|
| Scoreline agreement | Whether a model's value gap matches each result — high here proves reconstruction, not insight |
| Draw centring | Whether value gaps in drawn matches centre on zero |
| Calibration to goals | Whether a season's team values track goals scored |
| Horizon decay | How much the score falls from the first horizon to the third: form or strength |
Scope and timing
- 2026/27, the top five European leagues, with one score across all five.
- A fixture is scored once both teams have three matches behind them, so each league's first scored fixtures fall on its fourth matchday.
- Values for a round are due at 23:59 Paris time the day after it ends — Tuesday for a weekend round — and scores are updated every Saturday morning, once the previous round's results are final.
Taking part
A participant integrates the Arena's team identifiers once, then its own scheduled job delivers one number per team per match after every round. Values can be corrected until their deadline and freeze at it; every delivery gets a receipt, and a participant is reminded before any deadline where something is still missing.
A participant can join mid-season and is scored from the fourth round after it joins. Values for matches played before it joined are not accepted: they could be priced knowing every result since.
Onboarding is the page to read before writing any code: every step from registering to appearing on the board, each with a way to tell that it worked. The delivery contract is the reference behind it.
What this does not claim
It is not a match-prediction contest. A great many things decide a football match that no possession model should be trying to capture, and absolute accuracy is beside the point. What is measured is the increment — and increments can be compared even when nobody's absolute number is impressive.
Nor does the team board say anything about the other two. A participant may enter any board, and no board checks a submission against another: how a model's actions add up to a team's match is its own choice.