It has been an exciting (and in many ways unpredictable!) World Cup. As the dust settles, we can see how Quintessa’s N-Estimates algorithm has performed.
Throughout the FIFA World Cup 2026, Quintessa’s N-Estimates
algorithm was busy predicting the results. Here, we analyse the algorithm’s performance, and compare it with independent predictions produced by the BBC using their expert pundit Chris Sutton (Figure 1) and by asking Microsoft Copilot to Predict the results of the World Cup [round] ties.
Previously, each of Quintessa's predictions was accompanied by a plot displaying all the possible scoreline probabilities with the most likely outcome marked by a cross. These plots have been updated with a green circle representing the actual result. The final plots for all the matches can be found at the end of this article.
Group Winners
Before the tournament, we presented N-Estimates’ predicted winners for each of the twelve groups. Nine of these predictions were correct. Of the others, Turkey finished bottom of their group after failing to score (despite a large number of chances) in their first two matches; Portugal’s draws against DR Congo and Colombia saw them finish second in the group behind the South Americans; and Norway also finished second in their group after falling to a heavy defeat against France while resting top goal scorer Erling Haaland.
Algorithm Success Rates
When assessing the performance of the algorithm during the competition, we have considered five performance metrics:
- the percentage of correct winners or draws predicted;
- the percentage of correct goal difference predictions;
- the percentage of correct exact scoreline predictions;
- the goal difference discrepancy; and
- the goal discrepancy.
The first three of these measure the success rate of the algorithm and the last two measure how close the algorithm was in cases where the goal difference or scoreline prediction was incorrect. The goal difference discrepancy is the cumulative difference between the predicted and observed goal differences. This distinguishes between the two teams (so, for example, if the prediction is 2-0 and the actual result is 0-2, the goal difference discrepancy is 4). The goal discrepancy is the cumulative difference between predicted and observed goals scored by each team. For example, if the prediction is 2-1 and the actual result is 1-2, the goal discrepancy is 2 (as each team was one goal away from the prediction).
Throughout this analysis, the predictions are compared with the scores at full time (including extra time if applicable; matches that were decided by a penalty shoot-out are counted as draws).
Figure 2 compares the algorithm’s performance over the whole tournament against a benchmark in each of the first three metrics. These benchmarks represent the expected success rate if predictions were made at random. They were calculated by taking all matches in the dataset used to inform the N-Estimates algorithm, after the FIFA World Cup in Qatar in 2022, and finding the frequency of each kind of result (wins and draws for the correct winner/draw
benchmark, each goal difference for the correct goal difference
benchmark, and each scoreline for the correct scoreline
benchmark). As home advantage is irrelevant for the World Cup (except for matches involving the host nations), the frequencies of equivalent results for home and away teams (e.g. 1-0 and 0-1) were averaged. These frequencies were normalised to give the probability of selecting that outcome at random for a given match. These probabilities were then multiplied by the number of matches that occurred with that outcome during the tournament to give the expected number of correct predictions, and rounded to the nearest integer.
The algorithm performed much better than chance. The correct winner/draw was predicted in 66 matches (63% of the 104 matches in the tournament), nearly twice what would be expected by chance. The correct goal difference was predicted in 31 matches (30%), just over double the expectation by chance, and the exact scoreline was correctly predicted in 19 matches (18%), nearly four times the expectation by chance.
N-Estimates versus AI and Expert Judgement
Overall, N-Estimates performed best on three of the five metrics considered, and was only narrowly beaten by Copilot on the other two.
The N-Estimates algorithm was more successful than Chris on all three success rate metrics, and was more successful than Copilot except for the correct winner/draw
metric, where Copilot won by two matches (68 matches with the correct prediction compared with 66 for N-Estimates).
Table 1 presents the goal difference discrepancy and goal discrepancy values for the N-Estimates algorithm, Chris Sutton and Copilot. This shows that, overall, N-Estimates was comfortably closer to the number of goals scored by each team than Copilot or Chris. After a closely-fought competition, Copilot was closer than N-Estimates on goal difference predictions by just two goals (remarkably close, over 104 matches).
| Goal Difference Discrepancy | Goal Discrepancy | |
|---|---|---|
| N-Estimates | 126 | 180 |
| Copilot | 124 | 190 |
| Chris Sutton | 143 | 211 |
Capturing the Uncertainty
There is a large degree of inherent variability in the outcomes of football matches. For each match, our predictions included calculated probabilities for various scorelines. We can use these to calculate the percentile within the distribution of predicted results that contains the actual match result, for a given match. Plotting the cumulative distribution of these percentiles allows an assessment of how well (or otherwise) the variability of the predictions agrees with that of the real-life results. Such a plot is shown in Figure 3. This plot excludes the 19 matches in which the observed result was predicted as the most likely scoreline; this is because those matches sit at the 100th percentile by definition, and therefore can obscure the interpretation of the plot.
If the distribution of predictions exactly matched the distribution of observed match results, the cumulative distribution would follow the black dashed line in the figure. In practice, we expect the distribution to lie within the 3σ envelope denoted by the dotted lines.
The blue line shows the cumulative distribution of the N-Estimates predictions (excluding perfect predictions). There is good agreement between the blue and black lines across most of the range, indicating that the inherent uncertainty in match results has been well-captured. However, there was one imperfection in the uncertainty modelling, as shown by the larger deviation above the 70th percentile. This arises because there were more results than expected between the 70th and 75th percentiles (and correspondingly fewer above the 90th percentile). When each match is predicted, there will be several possible outcomes surrounding the most likely outcome with similar probabilities, so even when the prediction and observed outcome only differ by a single goal, that prediction might only have the 4th or 5th highest probability. During this tournament, the 4th-most-likely outcomes were observed more often than the 2nd or 3rd-most-likely, indicating a small opportunity for improvement.
Prediction Plots with Observed Results
Group Stage
Figure 4: Predictions and actual scores for every group stage match. For each plot, the circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Last 32
Figure 5: Predictions and actual scores for every last-32 match. For each plot, the circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Last 16
Figure 6: Predictions and actual scores for every last-16 match. For each plot, the circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Quarter Finals
Figure 7: Predictions and actual scores for every quarter finals match. For each plot, the circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Semi Finals
Figure 8: Predictions and actual scores for both semi final matches. For each plot, the circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Third-Place Playoff/Bronze Final
Figure 9: Predictions and actual scores for the third-place playoff match. The circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Final
Figure 10: Predictions and actual scores for the final. The circles represent possible final scores, with the number of goals scored by each team plotted on the axes. Each circle has been colour coded to indicate the probability of that result occurring, with the most likely outcome marked with a black cross, and the actual outcome marked with a green ring. The dashed black line indicates a goal difference of zero.
Quintessa is not affiliated in any way with FIFA, UEFA, the BBC or Microsoft. Its application of the N-Estimates algorithm to the Fifa World Cup 2026 is an independent and non-commercial endeavour.