Is It Really a Draw-Heavy World Cup?
How to test questions with data
Day twelve, and the tournament has a habit it cannot shake: the draw. It feels like every other match ends level, and today brings a few more candidates: Norway-Senegal, France-Iraq, Argentina-Austria, and Jordan-Algeria. So it’s a good day to ask a question that sounds simple and is not: have there really been a lot of draws, or does it just feel that way? But first, a quick recap.
Previously, at the World Cup
Day eleven split cleanly between favorites taking care of business and favorites adding to the draw count. Spain steadied themselves against Saudi Arabia, and Egypt beat New Zealand 3-1. In between, Belgium and Iran played out a 0-0 draw, and Uruguay were held 2-2 by Cape Verde.
The real story was on the forecasting board, where our two DSWC models finished at opposite ends of the field. Classic, the stubborn results-only model, posted the worst score of anyone. Pro, our market-aware competition model, posted the best. The scoreboard at the bottom updated accordingly.
Question 12: How do you tell a trend from a fluke?
So far, World Cup 2026 feels like it’s running hot on draws. Through the first 39 group-stage matches, 13 have ended even, a rate of 33%. That sits above the 25% the group stage usually runs at.
The twist is that the expanded 48-team field might have pointed the other way. More debutants, more mismatches, more chances for the strongest teams to separate. Instead, the newcomers keep digging in. So the question is not whether the tournament feels unusual. It does. The question is whether it is unusual enough to treat as evidence, or whether 39 matches can produce this much weirdness on their own. That is what the rest of this post is about, and the honest preview is: not yet.
The historical baseline
Start with history. This chart shows the share of group-stage matches that ended level in every World Cup since the 32-team format arrived in 1998. Tournament to tournament, the draw rate sits right around a quarter: 1998 ran hot at 33 percent, 2014 and 2018 came in low at 19 percent, and pooled across all seven it is 24.7 percent. That is relatively stable and gives us a clean expectation to measure against.
It also delivers the first cold splash. This year’s 33 percent is high, but it is not unprecedented. The 1998 tournament hit exactly the same mark. Before we run a single calculation, the long view already whispers that a third of group games drawn is something this format has done before.
Turning a feeling into a test
Now let’s do something more precise. Assume, as a starting point, that 2026 is a perfectly ordinary World Cup and draws are landing at the usual group-stage rate of about 25 percent. We’ll call that the null hypothesis. It is the boring explanation, the one that says nothing special is happening and what we see is normal variation.
Here’s the really important point. Statistical tests do not try to prove the exciting story. They ask whether the boring one can be ruled out.
If draws truly arrive a quarter of the time, then in 39 matches we would expect about 10. We have seen 13. So the question turns mechanical: if the real rate were 25 percent, how often would luck alone hand us 13 or more draws in 39 games?
With luck alone
Imagine we could replay the first eleven days of this World Cup over and over. Same teams, same schedule, same underlying tendency to draw. The only thing that changes from replay to replay is luck: which tight games happen to end level and which tip to a winner. The chart is what that giant pile of replays looks like. The x-axis represents all the draw totals we could have seen through 39 matches: 0 draws, 1 draw, 2 draws, and so on. The height of each bar shows how likely each total would be if this year’s draw rate were really in line with past tournaments.
If the true draw rate is the historical 25 percent, the replays cluster around 10 draws in 39 matches. That is what “normal” predicts. But the cluster is wide. Ten draws is the heart of it, eleven is common, twelve would not raise an eyebrow. The red bars are 13 or more, where this tournament actually sits, and added together they come to about 14 percent of all the replays.
That 14 percent is the p-value: the share of normal tournament starts that would, by luck alone, serve up 13 or more draws in 39 games. By convention, we usually start calling a result statistically significant once that share drops below 5 percent, when the boring explanation has been squeezed down to a one-in-twenty fluke. Fourteen percent is not nothing, but it is not enough. The draw surge is real in the standings; it is not yet strong evidence that this tournament is different.
Here is the same idea from the other direction. To actually rule out the null, to push luck below that 5 percent bar, we would have needed to see not 13 draws but 15, a draw rate of 38 percent rather than 33 (14 draws still lands at a p-value of about 8 percent). So the distance between “what happened” and “what would have convinced us” is two matches. Two coin-flip games falling the other way is the whole difference between a headline and a shrug.
So yes, 13 draws in 39 matches is high. But it is not impossibly high. Even in a completely normal World Cup you would see a start this draw-heavy, or heavier, about one time in seven. That is the lesson worth carrying out of this: unusual is not the same as proven.
Up next
Tomorrow we’ll stay with the theme of noisy evidence, but move from draws to goals. Goals are the thing every match is built around, and also one of the most stubbornly awkward things to model. A team can dominate for ninety minutes and score once. Another can create almost nothing and finish twice.
So we’ll ask what goals actually tell us once the tournament is underway. Are high-scoring wins real information? Do expected goals, shots, goal difference, and finishing luck help us update faster than the market?
Today’s scorecard and forecasts
Yesterday was a draw day again, and the models felt it. DSWC Pro had the best day on the board, helped by staying closer to the market on Belgium-Iran and Uruguay-Cape Verde while avoiding Classic’s bigger swings. Classic nearly landed its Belgium skepticism, but paid for being too confident on Uruguay and too cool on Egypt. Through 39 matches, Dimers still leads at 0.570, with the market, Kalshi, Opta, and PELE close behind. Classic trails at 0.610, still better than the 0.667 know-nothing line, but clearly behind the pack.
Today’s slate is a useful test of what kind of model you want. The board mostly agrees that France should beat Iraq and Algeria should beat Jordan. The interesting games are Argentina-Austria and Norway-Senegal. Classic is the most aggressive on Argentina, pushing the defending champions to 76 percent where the market sits at 62. Pro pulls that back toward the field at 63. Norway-Senegal is the cleanest draw candidate: everyone sees a close match, with Norway only a narrow favorite and the draw sitting near 30 percent. If today is going to keep the draw story alive, that is the place to look.







https://epiphanym3.substack.com/p/the-architecture-of-desolation