How it works
Four forecasts, each measured against the floor it has to clear. This page is the short version; the handbook is the long one.
- Match outcomeWho wins, and the scoreline
- Season shapeTitle, relegation, final table
- Knockout tiesWho advances, tie by tie
- TrophiesWho lifts it, simulated
Every number against its floor
Lower is better on both scales, so the short bar is the better forecaster. The gap is the claim; the level is not.
A match · three outcomes
- 0.5930This modelBrier
- 0.6667A blind one-in-three guess
43,433 matches · walk-forward
A knockout tie · two outcomes
- 0.2500Coin flip
- 0.2381Higher-rated side advances
- 0.2175This modelBrier
2,141 ties · trained only on earlier seasons
A trophy · log loss on the team that actually won it
- 1.9686This model
- 2.1453An unfitted Elo simulation
- 2.5606Uniform over the field
85 tournaments
Does 70% mean 70%
The claim this project cares about most. A calibrated band has its two bars the same length.
- 50–60%817Said55.1%Happened55.7%
- 60–70%731Said64.7%Happened64.8%
- 70–80%417Said74.3%Happened74.3%
- 80–90%175Said83.9%Happened86.3%
On match forecasts the same check gives a calibration error of 0.0099 over 43,433 matches.
Why calibration, not accuracyTwo records, never added together
One is retrospective and large. The other is what this site published before kickoff, scored afterwards.
What this is not
Not a betting product. The bookmaker's price is a yardstick here, not a target: measured across every bucket where this model disagrees with the closing line, backing it loses money — and loses more the more confident it is. That measurement is published rather than buried.
Documentation
Tutorials, the concepts in full, and a reference for the API, the artifacts and the commands.