inocta.iomatch
Published study · World Cup 2026

The study: calibration over accuracy.

A pre-registered live benchmark of three AI reasoning architectures, run across all 104 matches of the 2026 FIFA World Cup. The finding, the open dataset, and four ways to read it.

What we found

The Council architecture recorded the lowest expected calibration error of any machine tested, ECE 0.100 over 104 scored matches, ahead of the betting market at 0.104, the Pipeline at 0.112, and the Solo at 0.114 (the human reference aside at 0.074). Among the three AI reasoning architectures, the Council is the best calibrated, and it is the shape to reach for when a business needs an AI-supplied probability it can trust.

The Council architecture, three AIs that debate with an anonymous fourth deciding, is the best-calibrated shape we tested. Of every machine on the board, it posts the lowest calibration error. When an operator asks which shape of AI to reach for when the number has to be trustworthy, that is the answer.

The result is directional, not a photo finish erased by noise. The bootstrap interval around the leaders overlaps, which we disclose as a limit a second tournament would sharpen. It does not turn the recommendation into a null. You do not refuse to call the faster car because two lap times sit within measurement error.

The three architectures

Three ways to build an AI forecaster, compared on the same 104 matches. Same question, three shapes of reasoning.

Solo
The Soloist

One AI, one pass, no helpers. It gets the full match dossier in a single call.

Pipeline
The Assembly Line

A chain of specialists, each step building on the last. It keeps notes on its own misses and reads them before the next match.

Council
The Council

Three different AIs talk it out, sometimes disagreeing sharply. A fourth makes the final call.

Meet all eight analysts →

The evidence

The study compares three AI reasoning architectures on calibration over all 104 matches. Lower ECE means confidence you can trust. The Council leads.

Calibration error (ECE), the three architectures

1Council0.100
2Pipeline0.112
3Solo0.114

Lower is better. The Council posts the lowest calibration error of the three architectures, and the lowest of every machine tested (the human reference aside).

Reliability curve

Each point is a confidence band: what the architecture said (horizontal) against what actually happened (vertical). The diagonal is perfect calibration. Closer to the line is better.

See all eight analysts on the calibration page →

Why this matters to a business

On football, the AI forecasts land near the betting market because football already has a massive liquid market solving this. A company has no such market. There is no closing line for a Q3 pipeline. So the calibrated AI forecast is not competing with a reference, it IS the reference a business otherwise lacks. That makes the Council more valuable to an operator, not less.

Four ways to read it

Two reads, two languages. The business read is the decision-grade brief. The scientific read is the full peer-review-grade paper.

Business read

The decision-grade brief. What the finding means for choosing an AI shape, in operator language.

EnglishViewDownload PDF
FrançaisViewDownload PDF

Scientific read

The full paper. Methodology, pre-registration, scoring, reliability curves, and every result.

EnglishViewDownload PDF
FrançaisViewDownload PDF

Cite this work

Pre-registered on OSF before the first scored match. Open dataset, CC-BY 4.0.

Open dataset (OSF)
https://osf.io/r5sbj/
Pre-registration (OSF)
https://osf.io/ke35t/
APA
Kahlain, M. (2026). Calibration over Accuracy: A Pre-Registered Live Benchmark of Three AI Reasoning Architectures on the 2026 FIFA World Cup. OSF. https://doi.org/10.17605/OSF.IO/R5SBJ
BibTeX
@misc{kahlain2026calibration,
  author = {Kahlain, Mohamed},
  title = {Calibration over Accuracy: A Pre-Registered Live Benchmark of Three AI Reasoning Architectures on the 2026 FIFA World Cup},
  year = {2026},
  publisher = {OSF},
  doi = {10.17605/OSF.IO/R5SBJ},
  url = {https://doi.org/10.17605/OSF.IO/R5SBJ}
}

DOI · 10.17605/OSF.IO/R5SBJ

Explore the evidence

The full benchmark is preserved as a public archive. The scoreboard, the reliability curves, and every match prediction stay online.