Skip to main content

How it works

Four forecasts, each measured against the floor it has to clear. This page is the short version; the handbook is the long one.

Every number against its floor

Lower is better on both scales, so the short bar is the better forecaster. The gap is the claim; the level is not.

A match · three outcomes

  • This modelBrier
    0.5930
  • A blind one-in-three guess
    0.6667

43,433 matches · walk-forward

A knockout tie · two outcomes

  • Coin flip
    0.2500
  • Higher-rated side advances
    0.2381
  • This modelBrier
    0.2175

2,141 ties · trained only on earlier seasons

A trophy · log loss on the team that actually won it

  • This model
    1.9686
  • An unfitted Elo simulation
    2.1453
  • Uniform over the field
    2.5606

85 tournaments

What each floor is

Does 70% mean 70%

The claim this project cares about most. A calibrated band has its two bars the same length.

  • 50–60%817
    Said
    55.1%
    Happened
    55.7%
  • 60–70%731
    Said
    64.7%
    Happened
    64.8%
  • 70–80%417
    Said
    74.3%
    Happened
    74.3%
  • 80–90%175
    Said
    83.9%
    Happened
    86.3%

On match forecasts the same check gives a calibration error of 0.0099 over 43,433 matches.

Why calibration, not accuracy

Two records, never added together

One is retrospective and large. The other is what this site published before kickoff, scored afterwards.

Backtest
43,433
matches, refit as the corpus advanced
Published live
328
forecasts scored after the result

What this is not

Not a betting product. The bookmaker's price is a yardstick here, not a target: measured across every bucket where this model disagrees with the closing line, backing it loses money — and loses more the more confident it is. That measurement is published rather than buried.

Documentation

Tutorials, the concepts in full, and a reference for the API, the artifacts and the commands.