The analysts
The eight analysts
The three AI teams
Three AI teams, three benchmarks, one set of matches and one scoring rule. Each team also has a twin with one ingredient removed, which tells us exactly what that ingredient is worth.
Solo reads everything we have on a match: the form, the lineups, the rivalries, the weather, even notes from its own past predictions. One model, one call, one prediction. No second opinion, no chain of specialists. Just discipline and synthesis. The risk is over-confidence on heavy favorites. The strength is speed and a clear single voice.
- 1.Claude Opus 4.7single call · the full match dossier
Pipeline works like a newsroom. A statistician reads the numbers, an operations reader catches what the numbers miss, an editor writes the final call. It keeps notes on its own misses and rereads them before every match. And when its confidence drops under 60 percent, it hands the ball to Council, on the record.
- 1.Claude Opus 4.7statistician · quantitative read
- 2.Claude Opus 4.7operations reader · what the numbers miss
- 3.Claude Opus 4.7voice editor · final synthesis
Council convenes three different model families for each match. They argue, sometimes sharply. A synthesizer weighs the three views by how far apart they sit and delivers one prediction, with the disagreement itself logged as a risk factor. Best at catching what the other setups miss. Most expensive to run.
- 1.Claude Opus 4.7council member · the structural reader
- 2.GPT-5.4council member · the contrarian
- 3.Gemini 3.5 Flashcouncil member · the historian
- 4.Claude Opus 4.7synthesizer · final call
The three benchmarks
The benchmarks don't compete. They are what the AI teams get measured against.
A pure math rating built from four years of results, like a chess ranking, plus a home-field bump. No AI involved.
The betting odds at kickoff, turned into percentages. The AIs never see them. The hardest score to beat.
The human participant. Mo predicts every match on instinct alone, before kickoff, without looking at any data. The open question: does human instinct hold up against machines and math over 104 matches?