Calibration over accuracy.
The Council architecture is the best-calibrated AI shape we tested. Read the finding two ways, in two languages, with the open dataset and citation.
Read the paper→The eight analysts
Final scoreboard
Ranked by calibration error (ECE), lower is better. Calibration, not raw accuracy, is what this benchmark measures. Click any column to re-sort.
| # | Analyst | Type |
|---|
A note on The Human. Mo is a reference line, not a competitor. As the study's operator he could see the AI predictions before locking his own picks, so any calibration edge is confounded by information access, not an independent human-versus-machine result. The finding is about the machines: the Council is the best-calibrated architecture tested.
- Brier — The real score. It measures whether the confidence is honest: being confident and wrong is punished hard, being cautious and wrong barely costs anything. Lower is better. (Wikipedia ↗)
- ECE (Expected Calibration Error) — The trust check. When an analyst says 70 percent, does it actually win about 70 percent of the time? It catches bluffing. Lower is better. (Wikipedia ↗)
- Picks / skipped — Skipping a match is allowed. An analyst that sits one out instead of guessing is not penalized, so picks and skips sit side by side.
- Hit rate — Plain percentage of correct calls. Fun to watch, but it is not the verdict here. A bold guesser can have a high hit rate and still be poorly calibrated.
- Why wait for the curve? — A calibration curve needs about 24 scored matches before it means anything. That is one full group matchday. Showing it sooner would be noise pretending to be signal.
How it works
An hour before kickoff, every AI gets the same match dossier and locks its call: winner, score, confidence. No revisions, ever.
Picks lock. In that same final hour we record what the betting market believes, the hardest benchmark in sports. The human enters his pick on instinct alone, any time before kickoff.
Then we score everyone. Not just right or wrong: does 70 percent confident actually win 70 percent of the time? The misses go up next to the wins.
- ●Rules locked and published before the tournament ↗
- ●Scored by three independent AIs plus a human judge
- ●AI models frozen for the whole tournament, no mid-game upgrades
- ●Full dataset open on OSF
Calibration over accuracy.
When an analyst says 70 percent, does that pick win 70 percent of the time? That is what this chart answers (statisticians call it a reliability curve). Updated after every completed match. It appears after the first scored matchday.
Calibration →Solo, Pipeline, Council. The same three AI systems we put to work inside real businesses.
An AI that says 70 percent is making you a promise. Calibration checks whether it keeps it, and that decides what you can delegate. Football is just our test track, with a scoreboard nobody can argue with.
