Evaluation protocol
The rules that hold across every benchmark. A task decides what is measured; this page decides how any measurement here is conducted, published, and allowed to change.
Capabilities, not companies
There is no universal benchmark, because there is no universal task. League normalization is not player similarity. Player similarity is not recruitment. Recruitment is not opponent analysis. Each deserves its own methodology, and a single vendor score would average away the only thing a club actually wants to know.
So a participant does not receive a score. It receives one per benchmark, and a provider strong at projecting transfers and weak at normalising across leagues reads as exactly that.
Six principles
Task-first. Every benchmark is designed around one question with a checkable answer, rather than a shared scoring template stretched to fit.
Use the strongest available signal. Where ground truth exists — what a player actually did — it is used. Where it does not, the benchmark says so rather than substituting a proxy and calling it truth.
Neutrality. Every participant is evaluated under exactly the same protocol, on the same cases, with the same request payload. The Arena's own models take part on those terms and have no privileged access to reference data.
Transparency. Every benchmark states what is measured, why, how the score is computed, and where the data comes from. Every published result is reproducible from stored artifacts.
Versioning. Protocols, datasets and metrics carry versions. Every published score names the versions it was computed under, so historical results stay readable while the methodology improves.
Continuous evolution. Football changes and so does analytics. A frozen protocol would be accurate for one season and misleading afterwards.
Participation rules
Registration
An endpoint is validated live, before it is stored. It must answer a fixed sample case in the contract's shape; values are never judged. Registering first and validating later would put a row on the board that cannot be scored.
Registration is open at any point in a season. See rolling entry.
Freezing
Everything a scored round depends on is pinned before the round begins:
- The contract version pins the request schema, the response rules, the role taxonomy, the metric vocabulary and the band grids together. A change to any of them is a new contract version.
- The reference pool is frozen before the season and never recomputed during it. A pool that moved would change the percentile of an appearance already scored, silently rewriting a published leaderboard.
- The case registry is versioned, so the cohort a score was computed on is nameable later.
Locking
A prediction is locked when it is stored, and locking must precede the first observation of that case.
- A participant cannot resubmit a case, in any round.
model_versionis recorded at lock time and published with the score, so a board row is always traceable to the model that produced it and a later model cannot claim an earlier result.- Requests carry no provider identifiers and no ground-truth fields, so a case cannot be keyed against a scraped copy of the answer.
Rolling entry
A benchmark that only accepted entrants in August would be a benchmark almost nobody could enter.
A participant registering mid-season is offered the cases whose outcomes are still open when it registers. Its predictions are made conditioned on everything observable at its lock date, and that date is recorded.
What arriving late costs is coverage, not accuracy. Two participants that entered at different points did not answer the same question, and the board is built to show that rather than hide it.
Refusal
A participant may decline a case. Refusal is all-or-nothing: every metric is answered, or the case is refused.
A refused case is not scored as a maximally wrong prediction. A model that knows a case is outside its scope is behaving correctly, and scoring it as wrong would make good judgement indistinguishable from bad modelling. It reduces coverage instead, which is published beside the score.
Exclusion
Malformed responses, timeouts and transport failures are contract failures, not refusals. Repeated failures within a round can exclude a participant from that round. A broken endpoint is left retryable rather than recorded as a decline, so a fixed endpoint can rejoin.
Publication policy
Scores are appended, never rewritten. Each weekly run is computed afresh from full history, and the result is a new immutable snapshot. A value published in week 3 still reads the same in week 30.
That is what allows a board to show a participant's trajectory across a season rather than only its current value, and what makes a score citable weeks after it was produced.
Every snapshot names its versions. Case registry, entity crosswalk, reference pool, scoring metric and contract. Click a score and you can reach the rules that produced it — not the current documentation, those rules, as of that week.
Uncertainty is published with the number. Every score carries an interval, because early in a season the interval is the story and a point estimate alone would invite a conclusion the data does not support.
Nothing is private. There is no vendor-only view. A participant sees its coverage and its refusals on the same pages everyone else does. A benchmark whose credibility rests on there being no privileged view cannot have one.
Predictions are published individually, not only in aggregate. A participant's answer to a single case — the distribution it submitted, the error it scored, how that error moved as the player played, and how much football had already been played when it locked — is published under its name alongside the same figures for every other participant on that case.
This follows from what a benchmark is. A score nobody can inspect is a claim, not a measurement: a reader who cannot see the individual predictions has to take the aggregate on trust, which is the position the Arena exists to make unnecessary. It is also the only way a participant can be defended against a bad-looking number, because the case, the evidence and the base rate are all visible next to it.
Two things are published beside every individual call so that it cannot be read as more than it is. How many appearances it has been scored on, because a case with two matches behind it is close to noise and no single case is evidence about a model. And how much was already known when the prediction was locked, because a participant that entered in November answered an easier question than one that answered blind in August, and putting the two side by side without saying so would misrepresent both.
A cohort below its publication floor is marked, not hidden. A board measuring recruitment in a league has to cover the league rather than three busy clubs, so a competition-season is published only once it yields at least one scorable case per club on average — and says so when it does not.
What a benchmark must document
Before a benchmark accepts registrations it publishes: the question, the request and response contract, the scoring formula with worked examples, the data it is built on with its gaps stated and dated, and the versions in force.
If something is not decided, the documentation says it is open rather than implying it is settled.