Open · team board scoring weekly

Game Value

Football is a sequence of actions. Which of them mattered, and by how much? Expected threat, VAEP, on-ball and possession-value models — and the in-house variants clubs and broadcasters run — disagree about the same pass in the same match, and there is no public evidence about which disagreement is right.

3
boards, scored separately
5
top-flight leagues in scope
0
what a copy of the scoreline scores
=
Agreeing with the scoreline is not evidence.

A model that returns “team value = goals scored” agrees with every result and knows nothing. So the benchmark scores what a model adds to what the results already say — and a copy of the scoreline scores exactly zero, by construction.

Method neutral

What is measured

Three boards, three baselines, never blended

Each board is a skill score against its own baseline, so each has its own zero — and three zeros cannot be added into one number without inventing a weighting nobody can defend. A participant enters any of them, and none is checked against another.

01

Team

Skill over the scoreline

Each side's average value per match so far, compared across the two teams, as a forecast of the next result. It is scored against the same forecast built from goal difference, so a metric that only echoes the table adds nothing — and each fixture is scored from three depths of history, which separates a model of current form from one of strength.

See the board
02

Player

Skill over plus-minus

The ratings of the players on the pitch, taken from before the match, against what happens while they are on it. Each match is cut into three segments, so a substitution compares the same team with and without a player, against the same opponent, on the same afternoon. Any player-rating product can enter, with or without a game value model.

Not open yet
03

Element

Forward calibration

Does the value a model gives an action in a given situation hold up in the weeks that follow? It is scored against what that situation has actually produced so far, so a model is paid only for what it resolves beyond the record.

Not open yet

What a participant sends

One number per team per match, delivered by its own job

For the team board, a participant values each side’s performance in each match, on its own scale — nothing is normalised, because the score depends only on order. Its scheduled job delivers the values after every round. They are due at 23:59 the day after a round ends — Tuesday, for a weekend — because match data is not final at the whistle; until then they can be corrected, and at the deadline they freeze. A value only ever counts toward matches played after its deadline, so no value can be priced knowing the result it is meant to predict.

A deliberate boundary

This is not a match-prediction contest

A great many things decide a football match that no possession model should be trying to capture, and absolute accuracy is beside the point. What is measured is the increment — what a model adds to what the scoreline already says — and increments can be compared even when nobody’s absolute number is impressive. The protocol is published before any score exists, so that it cannot be changed to suit one. Read the documentation, or see the team board.