What we found
The Council architecture recorded the lowest expected calibration error of any machine tested, ECE 0.100 over 104 scored matches, ahead of the betting market at 0.104, the Pipeline at 0.112, and the Solo at 0.114 (the human reference aside at 0.074). Among the three AI reasoning architectures, the Council is the best calibrated, and it is the shape to reach for when a business needs an AI-supplied probability it can trust.
The Council architecture, three AIs that debate with an anonymous fourth deciding, is the best-calibrated shape we tested. Of every machine on the board, it posts the lowest calibration error. When an operator asks which shape of AI to reach for when the number has to be trustworthy, that is the answer.
The result is directional, not a photo finish erased by noise. The bootstrap interval around the leaders overlaps, which we disclose as a limit a second tournament would sharpen. It does not turn the recommendation into a null. You do not refuse to call the faster car because two lap times sit within measurement error.
The three architectures
Three ways to build an AI forecaster, compared on the same 104 matches. Same question, three shapes of reasoning.
One AI, one pass, no helpers. It gets the full match dossier in a single call.
A chain of specialists, each step building on the last. It keeps notes on its own misses and reads them before the next match.
Three different AIs talk it out, sometimes disagreeing sharply. A fourth makes the final call.
The evidence
The study compares three AI reasoning architectures on calibration over all 104 matches. Lower ECE means confidence you can trust. The Council leads.
Calibration error (ECE), the three architectures
Lower is better. The Council posts the lowest calibration error of the three architectures, and the lowest of every machine tested (the human reference aside).
Reliability curve
Each point is a confidence band: what the architecture said (horizontal) against what actually happened (vertical). The diagonal is perfect calibration. Closer to the line is better.
Why this matters to a business
On football, the AI forecasts land near the betting market because football already has a massive liquid market solving this. A company has no such market. There is no closing line for a Q3 pipeline. So the calibrated AI forecast is not competing with a reference, it IS the reference a business otherwise lacks. That makes the Council more valuable to an operator, not less.
Four ways to read it
Two reads, two languages. The business read is the decision-grade brief. The scientific read is the full peer-review-grade paper.
Business read
The decision-grade brief. What the finding means for choosing an AI shape, in operator language.
Scientific read
The full paper. Methodology, pre-registration, scoring, reliability curves, and every result.
Cite this work
Pre-registered on OSF before the first scored match. Open dataset, CC-BY 4.0.
Kahlain, M. (2026). Calibration over Accuracy: A Pre-Registered Live Benchmark of Three AI Reasoning Architectures on the 2026 FIFA World Cup. OSF. https://doi.org/10.17605/OSF.IO/R5SBJ
@misc{kahlain2026calibration,
author = {Kahlain, Mohamed},
title = {Calibration over Accuracy: A Pre-Registered Live Benchmark of Three AI Reasoning Architectures on the 2026 FIFA World Cup},
year = {2026},
publisher = {OSF},
doi = {10.17605/OSF.IO/R5SBJ},
url = {https://doi.org/10.17605/OSF.IO/R5SBJ}
}DOI · 10.17605/OSF.IO/R5SBJ
Explore the evidence
The full benchmark is preserved as a public archive. The scoreboard, the reliability curves, and every match prediction stay online.
