Game Value — Element board
Does your value function assign value in a way that generalises to football it was not looking at when it assigned it?
This is the board where a game value model makes its actual claim — this pass was worth 0.04 goals, this tackle 0.06 — and the only one of the three that needs a possession-value model to enter.
This benchmark is in draft and is not scoring yet. The protocol is published before any score exists, so that it cannot be changed to suit a result.
The question that cannot be asked
Not "was this pass really worth 0.04?" That has no observable answer. The claim is about what would have happened otherwise, and what would have happened otherwise is never seen — not unmeasured, but unobservable.
Why a match's own outcome is off limits
Element values are delivered after the match has been played. Any score computed against that same match's outcomes is hindsight, and hindsight is trivially gamed: give a high value to whatever happened in the fifteen seconds before each goal, a low one to everything else, and score perfectly while modelling nothing.
So the values a participant sends for a match are never scored against that match. What is scored is a summary of them, formed on matches already delivered and tested on the matches that follow.
What the board can see, and what it cannot
The Arena receives a value function's outputs, never the function itself, so it cannot run a participant's model on new football — which is what testing a model out of sample normally means. That sets a precise limit, and it is stated before anything is built on it:
A value function can be tested only at the resolution of the situation it declares. Whatever a model knows that it cannot express in a declared situation — the exact coordinates, the shape of the defensive block, the six passes that came before — is real, may well be what makes it good, and is not measurable here.
This is a limit of the benchmark, not of the models, and it is why the element board is scored more modestly than the other two. A participant does control how fine its declared situations are, and that turns out to be a genuine strategic choice.
What a participant sends
For every match, the value of every element, and the situation each one sits in:
| Field | Meaning |
|---|---|
| match, team, player | who produced it |
| second | when, from kick-off |
| element type | what kind of thing it is, in the participant's own words |
| value | the value assigned, on the participant's own scale |
| situation | drawn from the Arena's closed vocabulary |
No ontology is forced. One model values actions, another states, another possessions; the element type carries whichever it is, unchanged. What has to be shared is the situation, because that is what makes two participants comparable.
At something like 1,500 elements a match, a season runs to millions, so elements are delivered as a file rather than in a request.
A closed vocabulary of situations
Each element is labelled with a situation from a fixed vocabulary the Arena publishes. A participant cannot invent labels, and that is a safeguard rather than tidiness: a free-text label is a channel for hindsight. A participant could label elements pass that led to a goal, and the score would reward the label rather than the model. A closed vocabulary makes every label a statement about the state of the match, which is checkable, rather than about what happened next, which is not.
The vocabulary proposed for v0:
| Dimension | Values |
|---|---|
| Pitch zone | 12 zones — thirds by channels, attacking direction normalised |
| Action family | pass, carry, dribble, shot, cross, tackle, interception, clearance, recovery, error |
| Game state | leading by 2+, leading by 1, level, trailing by 1, trailing by 2+ |
| Phase | open play, transition, set piece, counter |
That is 2,400 situations: fine enough to be interesting, populated enough to estimate. The size is still being settled.
The score: forward calibration
Each scoring week splits the season in two — the weeks so far, and the weeks still to come.
For every situation, three numbers:
| Built from | Is | |
|---|---|---|
| The participant's value | its average value for the situation, over the weeks so far | what the model says the situation is worth |
| The record | how often the situation was followed by a goal for or against, over the same weeks | what the situation was observed to be worth |
| What followed | the same frequency, over the weeks that come after | what it turned out to be worth |
The score asks one thing: does the model order situations by their future worth better than their own past record does? Skill is the model's ordering against what followed, minus the record's ordering against the same. Zero means the model knew nothing the record did not.
An element's outcome is simply whether the team scored, or conceded, within a short window after it. The window's length is one of the decisions still open.
The situation is the scored unit, not the element
Every element in a situation gets the same fitted value, so elements within one are indistinguishable anyway. Scoring situations makes that explicit, and it matters twice over. Almost no single element is followed by a goal, so an element-level comparison would be nearly all ties; averaged over a situation, the outcome becomes a rate that is almost never tied. And a season's elements would mean hundreds of billions of comparisons, where a few thousand situations mean a few million.
Why a copy of the outcomes scores zero
A participant labelling by hindsight — value 1 if a goal followed, 0 otherwise — produces situation averages that are the record. Its ordering coincides with the baseline's and its skill is zero, however perfectly it labelled the past. As on the other two boards, the thing it copied is the thing being subtracted.
That is also why the baseline is the record and not zero: a situation's past frequency is a perfectly respectable predictor of its future one, and scoring against zero would pay the copy handsomely.
What a model is actually paid for
A model wins only where it estimates a situation's worth better than that situation's own record:
- Rare situations, where the record is an average over a handful of elements and a model that borrows strength from similar states is genuinely better.
- Situations whose record is distorted by a short run — a phase of the season, a streak of unusual finishing — which a model built on structure rather than frequency will not chase.
This is a real but modest claim, and the board does not oversell it. A participant declaring coarse situations will find the record matches it almost exactly and will score near zero. Declaring finer ones opens a real gap, and requires actually beating a sparse estimate. That trade-off is the closest this board comes to measuring the resolution it cannot see directly.
A participant declaring near-unique situations — one per element, in the limit — finds that almost none recur in the weeks that follow. It is published as having described a great deal and predicted almost nothing; the design needs no rule against it.
Reading the score
The statistic is Somers' D, as on the other two boards. Because the outcome here is a rate rather than a count, ties are rare and the older Kendall's τ-b would reach almost as high — D is kept so that a reader comparing boards compares like with like.
A second view compares situations only within the same game state. It asks whether a model orders zones and actions correctly inside a match situation, rather than being carried by knowing that trailing sides attack more, and it is published beside the main score.
Intervals resample matches, never situations and never elements. Situations share matches, and two situations from the same match share its goals, so treating them as independent would print an interval narrow enough to separate participants that cannot actually be told apart.
Published beside the score, never ranked
Repeatability. A participant's value profile across situations, compared between the first and second halves of a season. It measures stability rather than truth, so it is never the score — but a hindsight copy fails it badly, because which situations happen to precede goals is noise, and noise does not repeat. It is set against the record's own stability.
The disagreement map. Where do providers most disagree? That needs no ground truth, so it is published from the first week: the elements two participants rank furthest apart. Elements are matched across participants by match, team, player, time and action family, and the share that could not be matched is published beside the map — two providers who cannot be aligned at all is itself among the most interesting findings about them.
Still to be decided
This board's protocol is in draft, and these are open:
- the length of the window after an element in which a goal counts — the score's sensitivity to it will be published either way;
- how many elements a situation needs before it is scored;
- whether "the weeks so far" means the whole season to date or a trailing window;
- whether 2,400 situations is the right size for the vocabulary;
- whether a participant entering a competition must describe every match in it;
- whether this board ranks at all in v0, or publishes its measurements without a leaderboard — given the limit above, a ranking may claim more than the method supports.
What this does not claim
It does not say whether any single value is right; that question has no observable answer. It measures whether a value function generalises to football it was not looking at, at the resolution of the situations it declares, and says plainly that finer resolution is out of its reach.
It is also independent of the other two boards. Element values are not required to add up to a participant's team values or player ratings — three blocked shots in three seconds before a goal are three actions but one situation, and how a model aggregates is its own choice.