Skip to content

Honest Stats · World Cup 2026 · A/B Testing

Norway just beat Brazil. Penalty shootouts are coin flips. The World Cup is a terrible A/B test, and your store might be running one.

Favorites win shootouts 47.4% of the time, a famous 60% finding died at n = 7,116, and a champion gets crowned on seven matches. What the wildest World Cup in memory teaches about sample size, peeking, and calling winners your data has not earned.

On July 5, in East Rutherford, Norway beat Brazil 2 to 1 in the round of 16. Erling Haaland scored twice, Neymar answered with a stoppage-time penalty and then retired from international football, and a country that had not even qualified for a World Cup since 1998 knocked out the five-time champions (Sky Sports; Al Jazeera, July 5–6, 2026). On Saturday, July 11, Norway plays England in Miami Gardens, the first World Cup quarterfinal in Norwegian history.

It is a wonderful story. It is also, if you squint at it the way we look at experiment dashboards all day, a live demonstration of the single most expensive mistake in A/B testing: trusting a verdict the sample cannot support.

The World Cup decides the biggest question in football on some of the smallest samples in sport. One-off knockout matches, shootouts that are statistically near coin flips, a champion crowned on seven games. Merchants do the same thing every week: call a winner on 200 sessions, trust a four-test winning streak, stop a test the first day the dashboard turns green. This tournament is a masterclass in why that fails, with better footage.

What does Norway's run actually prove?

TL;DR The big sample (8-for-8 in qualifying, 37–5 on goals) proves Norway is excellent. The single match against Brazil proves very little by itself. Knowing which evidence to trust is the whole game, in football and in testing.

Look at Norway through three different sample sizes and you get three different levels of certainty. That gradient, from noise to signal, is exactly the gradient every A/B test lives on.

Sample one: qualifying, 8 matches. Norway won all eight, scored 37, conceded 5, and beat Italy twice; Haaland scored 16 goals, the most in UEFA qualifying (Yahoo Sports). Eight matches is still football-small, but a perfect record with a +32 goal difference is a loud, consistent signal. This is the evidence that Norway belongs.

Sample two: the group, 3 matches. Beat Iraq 4–1, edged Senegal 3–2, then lost 1–4 to France (UEFA.com). Three matches produced a 2–1 record and two very different faces. If you had to judge Norway on the group alone, you would be guessing.

Sample three: the knockouts, n = 1 per round. A late Haaland winner against Côte d'Ivoire (2–1, June 30, FIFA), then Brazil (2–1, July 5). Glorious, and statistically almost mute. Brazil out-ranked Norway on every long-run measure; on the night, one match got decided by two moments. That is not a complaint. It is what single observations do.

Haaland now has 7 tournament goals, more than Messi, Mbappé and Ronaldo combined at this World Cup (ESPN). The tournament around him has been carnage for favorites: Uruguay and Türkiye went out in the group stage, Germany and the Netherlands lost round-of-32 shootouts, Brazil fell in the round of 16 (SI; ESPN). Meanwhile Cape Verde, population about 525,000, drew with Spain, drew with Uruguay, and took Argentina to extra time (SI; Al Jazeera). A record 215 goals in the 72 group matches, about 3.0 per game, the highest rate since 1958 (FIFA), and the widest World Cup ever is behaving like what it is: a small-sample festival where variance eats reputations.

Why penalty shootouts are coin flips with medals attached

TL;DR Ten pressure kicks decide a tournament. Stronger teams win shootouts 47.4% of the time, first kickers at World Cups 48.6%. If a metric ignores who is better, it is measuring noise.

This World Cup has already had three shootouts: Paraguay over Germany (4–3), Morocco over the Netherlands (3–2), Egypt over Australia (4–2), all in the round of 32 (FIFA; Sky Sports; Al Jazeera). Germany had never lost a World Cup shootout, 4-for-4 across half a century, until June 29 (Opta Analyst). Here is the uncomfortable arithmetic: a fair coin goes 4-for-4 one time in sixteen. Germany's perfect record never needed a penalty-taking gene. It needed sixteen national teams trying, and someone was going to be the lucky one.

The wider numbers say the same thing louder. Across World Cup history through 2022 there have been 35 shootouts, and the team kicking first won 17 of them: 48.6%, a coin flip (Opta Analyst, June 2026). Conversion drops from 79.1% for in-game penalties to 69.4% in shootouts, so pressure is real, but it does not sort teams by quality. The cleanest test came in 2026, in the Journal of Sports Economics: across UEFA club shootouts from 2000 to 2025, teams that were more than 100 Elo points stronger, clear favorites by any definition, won just 47.4% of their 154 shootouts (Csátó & Petróczy). Read that again: the better team wins the shootout less than half the time in that sample. A shootout is roughly ten kicks. Ten trials cannot separate a great team from a good one, so they do not.

Same mistake, different scoreboard
At the World CupOn your storeThe honest read
A penalty shootout (~10 kicks)Calling a winner on a few hundred sessionsA coin flip wearing a scoreboard. Favorites win shootouts 47.4% of the time.
Germany: 4-for-4, then a loss“Our last four tests all won, we've got the touch”4-for-4 happens 1 in 16 by pure chance. Streaks are cheap at small n.
A 3-match group stageA one-week testEnough to catch a disaster, nowhere near enough to crown a winner.
8 qualifiers, 37–5 goalsA test run to its precomputed sample sizeNow the verdict means something. This is the evidence Norway is good.
Watching the bracket match by matchChecking the dashboard daily, stopping at the first greenPeeking. Every look is another chance for noise to cross the line.

Even the scientists got fooled by a small sample

TL;DR The famous “first kicker wins 60%” finding shrank to 53.3% at 540 shootouts and to nothing at 7,116. Fifteen years of argument, settled only by sample size. Your dashboard runs the same movie on fast-forward.

If small samples only fooled fans, this would be a football story. They fooled the American Economic Review.

In 2010, Apesteguia and Palacios-Huerta published a celebrated paper showing the team kicking first in a shootout won 60.5% of the time, across the 129 shootouts they studied. The finding traveled everywhere: coin-toss winners were advised to always kick first, pundits repeated it every tournament. Then other researchers widened the lens. Kocher, Lenz and Sutter examined 540 shootouts and found 53.3%, not statistically distinguishable from a coin flip. Later replications kept agreeing with the bigger number: no order effect in 1,759 shootouts (Vollmer et al., 2024), and no order effect in 7,116 shootouts, the largest sample ever assembled (Pipke, 2025; the arc is summarized in Csátó & Petróczy, 2026).

Nobody cheated. The original data really did show 60.5%. That is precisely the trap: 129 trials of a noisy binary event will happily hand you a large, publishable, wrong-at-scale effect, and it took the field fifteen years and fifty times the data to walk it back. When your test hits 95% significance on day three and your gut says ship it, remember that professional economists, with referees and replication incentives, rode a 129-observation effect for over a decade.

The “first-kicker advantage” vs. sample size

Bars show the reported first-kicker win rate above the 50% coin-flip line. Sources: Apesteguia & Palacios-Huerta (AER 2010); Kocher, Lenz & Sutter (540 shootouts, not significant); Pipke 2025 via Csátó & Petróczy (2026).

The best team often does not win, and that is measurable

TL;DR Football's favorite wins roughly 50% of single matches (basketball's wins 65%+). Hungary 1954 and Brazil 1982 are what variance does to the best team in the world. Seasons reveal skill; single matches roll dice.

Anderson and Sally, in The Numbers Game, put a number on football's randomness from 8,000+ matches: the sport is roughly half skill, half luck, and the pre-match favorite wins only about 50% of the time, against more than 65% in basketball. Football is the major sport where a single result carries the least information. The World Cup then compounds it by making every decisive match a single result.

History keeps receipts. Hungary arrived at the 1954 final unbeaten in over 30 matches, having beaten West Germany 8–3 in the group stage two weeks earlier, and lost the final 3–2 (FIFA; the “Miracle of Bern”). The 1982 Brazil of Zico, Sócrates and Falcão, still routinely called the greatest team never to win it, went out 3–2 to Italy on a Paolo Rossi hat-trick (CNN; ESPN). Neither team stopped being the best team in the world that day. The format simply asked one match to answer a question one match cannot answer.

Leagues answer it. Over 38 matches, the table sorts skill from luck well enough that the same handful of clubs keep finishing on top. Nobody crowns a Premier League champion after three fixtures. Your store should hold itself to league standards, and most dashboards quietly run knockout football.

Your store runs a World Cup every week

TL;DR The typical Shopify store converts around 1.4%. At that rate, small session counts make your test results as random as a shootout, and daily dashboard-watching makes it worse.

Translate the football into sessions. The typical Shopify store converts about 1.4% of visitors (median; our sourced benchmarks). At 1.4%, a hundred sessions on a variant means one or two orders. Variant B beating variant A “3 orders to 1” is not a result. It is a penalty shootout: a handful of binary events deciding a question they cannot carry.

The streak trap follows the same math. Four tests in a row where the new variant “won” feels like a hot hand, the way Germany's 4-for-4 shootout record felt like a national trait. If each call was actually a coin flip, a 4-streak arrives 1 time in 16. Run tests the underpowered way and you will manufacture streaks, ship “winners” that do nothing, and eventually watch your conversion rate fail to move even though the dashboard says you have been winning all year.

And peeking is bracket-watching. Check a running test daily and stop the moment it crosses 95%, and your real false-positive rate climbs to a multiple of the 5% you think you accepted, because every look is another kick at the target. The mechanics, and the always-valid sequential methods that fix it, are in the significance section of our A/B testing guide. The one-line version: decide the stopping rule before kickoff, then do not renegotiate with the scoreboard mid-match.

The sample size your test actually needs

TL;DR Visitors per variant ≈ 16 / (baseline × lift²). At a 2% baseline and a 10% target lift, that is roughly 78,000 visitors per variant. If you cannot afford the sample, change the plan, not the standard.

The honest approximation fits in one line: visitors per variant ≈ 16 / (baseline conversion rate × relative lift²). It is the same power math used across the industry, worked through with examples in our guide's traffic-and-time section.

Plug in a common case: a 2% baseline and a 10% relative lift you would love to detect. The formula asks for roughly 78,000 visitors per variant, over 150,000 in total (the sample-size trap, worked through here). A store doing 3,000 sessions a month would need years. That is not a reason to fudge the math. It is the math telling you to test differently: bigger, bolder changes (a 30% lift needs about a ninth of the traffic of a 10% lift, since the lift term is squared), template-level tests pooled across a collection instead of one page, or the apply-and-measure methods in the low-traffic playbook.

Time matters alongside volume. Run complete weeks, because weekday and weekend shoppers are different populations, and let at least two business cycles through before judging. Norway earned trust across eight qualifiers spread over months, home and away. A test that has only seen your store's version of a home Tuesday has not met your customers yet.

Testing while the world watches football

TL;DR A mega-event shifts who is on your store and why. Keep tests running, but do not conclude during the spike; extend past July 19 and confirm on normal weeks.

There is one directly practical World Cup question for merchants: what does a month-long global event do to a running experiment? It changes your traffic mix. Different countries, different hours, more phones on couches, more distracted half-watching browsers. A test spanning the tournament measures a blended audience that stops existing after the final.

The honest protocol is short. First, do not launch a decisive, close-call test mid-event if your niche is visibly swept up in it; start it after, or accept that the read extends past the event. Second, if a test is already running, keep it running: pausing and resuming creates its own problems with returning visitors and assignment. Third, extend the window past July 19 and check whether the verdict holds on post-event weeks before shipping. Fourth, write the event into your test notes, so a strange fortnight in the data has its explanation attached when you reread it in October. This is ordinary seasonality discipline; the World Cup is just seasonality with better television.

The honest-test checklist

Everything above, compressed to what you do differently starting today.

  • Compute the sample size before starting: 16 / (baseline × lift²) per variant. If you cannot afford it, redesign the test, not the threshold.
  • Decide the stopping rule in advance and hold it. No stopping at the first green day; that is how coin flips get shipped.
  • Run complete weeks, minimum two business cycles. Weekday and weekend visitors are different teams.
  • Judge on revenue per visitor, not conversion rate alone, so one whale or one discount-hunter wave cannot fake a win (why RPV).
  • Distrust streaks. Four straight winners is 1-in-16 coin-flip territory, not proof of a golden touch.
  • During big events, collect but do not conclude. Extend past the spike, confirm on normal traffic.
  • When a result looks miraculous, remember the 60.5% paper: the sample was real, the effect was not. Wait for your 540, not your 129.

One disclosure, since this whole post argues for honesty: StorePilot, the product we are building at EVDEV, is an AI CRO agent for Shopify that exists because of exactly these failure modes. It enforces the minimum-sample gate and the pre-committed stopping rule in software, judges tests on revenue per visitor, and never lets a model call a winner (deterministic math decides; the AI only writes the words). It is pre-launch, and none of the statistics above depend on it. The math works the same whether or not you ever use our tool.

Norway might beat England on Saturday. If they do, it will be one more single match, and the qualifying campaign will still be the real evidence they belong. Enjoy the football precisely because anything can happen in one game. Then go home and make sure your revenue decisions run on eight qualifiers, not one shootout.

Questions merchants keep asking

Why is the World Cup a terrible A/B test?

Because it decides enormous questions on tiny samples. A knockout match is one observation, a penalty shootout is roughly ten kicks, and football is so noisy that the favorite wins only about half the time (Anderson & Sally, The Numbers Game). An A/B test that calls a winner after a few days or a few hundred sessions has exactly the same structure: a confident verdict resting on a sample that cannot support it.

Are penalty shootouts really a coin flip?

Very close. In a 2026 Journal of Sports Economics study of UEFA club shootouts (2000 to 2025), teams that were more than 100 Elo points stronger won only 47.4% of their 154 shootouts (Csátó & Petróczy). At World Cups, the team kicking first has won 17 of 35 shootouts, 48.6% (Opta Analyst). Being the better team buys you almost nothing across ten pressure kicks.

Was Norway beating Brazil a fluke?

One match cannot tell you either way, and that is the point. The strong evidence that Norway is genuinely good is the big sample: 8 wins in 8 qualifiers, 37 goals scored, 5 conceded, including two wins over Italy (Yahoo Sports). The 2 to 1 over Brazil on July 5 is a single observation. Real quality and match-level randomness are both true at once; only the large sample separates them.

What was the first-kicker advantage, and why did it vanish?

A 2010 American Economic Review paper found the team kicking first won 60.5% of 129 shootouts. Independent replications on bigger samples kept shrinking it: 53.3% across 540 shootouts (not statistically significant), then no effect at all in 7,116 shootouts, the largest sample ever assembled (Pipke 2025, summarized in Csátó & Petróczy 2026). It took fifteen years and fifty times the data to settle. Small samples produce confident conclusions that do not survive scale.

What sample size does an A/B test need?

A useful approximation is visitors per variant ≈ 16 / (baseline conversion rate × relative lift²). At a 2% baseline, detecting a 10% relative lift takes roughly 78,000 visitors per variant. The full math and worked examples are in our A/B testing guide. If that number looks impossible for your store, that is a real answer: pick bigger swings or a different method, covered in the low-traffic playbook.

How long should I run an A/B test?

Until you hit the sample size you computed before starting, over complete weeks (weekday and weekend traffic behave differently), and typically two business cycles at minimum. Decide the stopping rule in advance and hold to it. Stopping on the first day the dashboard shows significance is how noise gets crowned.

What is peeking in A/B testing?

Checking a running test repeatedly and stopping the moment it looks significant. Each look is another chance for random noise to cross the threshold, so the real false-positive rate climbs far above the 5% you think you accepted. It is the dashboard version of judging a team by whichever ten minutes of the match you happened to watch. Fixed-horizon tests need a fixed stop; sequential methods (always-valid p-values) exist precisely to make continuous looking honest.

Should I pause my A/B tests during the World Cup or other big events?

Pausing is usually unnecessary, but do not conclude during one. A global event shifts who is on your store and why (new countries, new devices, distracted browsing). A test that spans the spike measures a blended audience that will not exist in August. Practical rule: keep collecting, extend the test past the event (the final is July 19), and check that the verdict holds on the normal-traffic weeks before shipping the winner.

My variant won after three days. Is it real?

Three days of data is a group stage, not a season, and usually not even that. Germany had won every World Cup shootout it ever contested, 4 of 4, until Paraguay beat them on June 29, 2026. A 4-for-4 streak happens 1 time in 16 by pure coin-flipping. Unless three days already contains your precomputed sample size (rare outside very high-traffic stores), wait. The pattern that survives four more weeks is the one that pays.

Do upsets mean skill does not matter?

No. Skill shows up exactly where samples get big. Norway's 8-for-8 qualifying run is skill you can trust; a league table after 38 matches is skill you can trust. Anderson & Sally estimate a single football match is roughly half skill, half luck, which is why the favorite wins only about 50% of the time, against more than 65% in basketball. Single matches lie; seasons tell the truth. Sessions lie; adequate samples tell the truth.

Founding-merchant offer

You already paid for the traffic. Let's convert more of it.

Join the StorePilot AI waitlist and lock in the founding-merchant offer.

+ add your store URL (optional)

Free for your first 3 months · No spam, just launch news. Unsubscribe anytime.

  • 24/7 support from a real human, not a bot
  • Hands-on setup for your store and catalog
  • A CRO expert reviews your first A/B tests by hand