Skip to content
← CRO for Shopify

Run A/B tests you can actually trust

Most 'winners' are noise called too early. Honest testing is the whole point.

Misha Gavura, Senior CRO Specialist Reviewed by Misha Gavura, Senior CRO Specialist · EVDEV Top Rated Plus Last updated

In short

  • Only 20% of 28,304 real experiments ever hit 95% significance, so most 'winners' were called too early.
  • Roughly 1 in 7 tests produces a real winner. If a tool says you're winning every time, it's lying.
  • Every time you peek at a running test and call it, you raise your odds of a false positive. Set the finish line before you start.

A "winner" you saw on day one isn't a winner. It's a sample size that hasn't argued back yet. In an analysis of 28,304 experiments run by Convert customers, only 20% ever crossed the 95% significance threshold, which means four out of five tests people called were guesses wearing a percentage sign. The whole job of an honest tool is to tell you when you don't know yet.

What's the problem?

You've been burned by tools that declared a 'winner' early, you shipped it, and conversion didn't actually improve. You can't trust results you can't trust.

Why does this happen?

  • Tools call winners before reaching significance.
  • Results aren't segmented, so a mobile loss hides inside a desktop win.
  • There's pressure to show a quick result over a true one.
  • Early data swings hard and then regresses. The first 50 visitors split unevenly by pure chance, so a variant can show +20% on Monday and -3% by Friday. The lift you saw early wasn't real signal; it was the spread you a…
  • Stopping a test the moment it 'looks significant' inflates your false-positive rate. Peeking at a running test and calling it whenever the line crosses 95% is the most common way to manufacture fake winners. Every extr…
  • Most variants genuinely don't beat the control, and that's the math, not your failure. VWO's own data puts the rate of meaningful winners at roughly 1 in 7 tests. A tool that declares a winner on most of your tests is l…
  • Low-traffic stores get strung along the longest. If you're doing a few hundred checkouts a week, a real 5% lift needs weeks of data to confirm, so a tool that promises a fast answer is either ignoring significance or t…

What does the research show?

Independent research

Figures below are from independent studies, not StorePilot data. They're why this problem is worth testing on your own store.

  • In an analysis of 28,304 experiments run by Convert customers, only 20% ever reached the 95% statistical-significance threshold, so most stores never gather enough traffic to call a clear winner.

    Convert
  • Only about 1 in 7 A/B tests (~14%) produces a meaningful winning variation that actually lifts conversions, and most variations never beat the original at all.

    VWO
  • Better checkout design alone can raise the average large ecommerce site's conversion rate by roughly 35%, so there's real upside worth testing honestly for.

    Baymard Institute, E-Commerce Checkout Usability research
  • Monitoring an A/B test continuously and stopping at the first significant result drives the false positive rate to 26.1% against a nominal 5% threshold, more than five times the intended risk.

    Evan Miller, 'How Not To Run An A/B Test'
  • To hold a true 5% error rate you must require 2.9% reported significance if you look once more, 1.4% across 5 looks, and 1.0% across 10 looks.

    Evan Miller, 'How Not To Run An A/B Test'

How does StorePilot AI fix it?

  • StorePilot enforces minimum-traffic and significance thresholds and never declares early winners.
  • It reports honestly, including 'Variant B +12% but not enough data yet', and segments by device and visitor type.
  • Every result carries a recommended decision (publish B / keep A / split-ship per device) with one clear action.

How do you fix it, step by step?

  1. Set the sample size before you launch

    Decide up front how many visitors (or conversions) per variant the test needs to detect a realistic lift, based on your current conversion rate and traffic. If you can't reach that number in a few weeks, the test is too ambitious for your traffic; pick a bigger swing or a higher-traffic page.

  2. Pick one primary metric and commit to it

    Choose revenue per visitor or conversion rate as the single metric that decides the test, before you see any data. Tracking ten metrics and celebrating whichever one turns green is how you find a 'win' in pure noise.

  3. Stop peeking. Let it run to the pre-set finish

    Don't call the test the moment it crosses 95% on day two. Run it to the sample size and the minimum duration you set (at least one full business cycle, usually 1-2 weeks), so weekday/weekend and payday traffic are all represented.

  4. Read the result by segment before you publish

    Split the outcome by device at minimum. A +8% desktop win can hide a mobile loss, and since most of your traffic is mobile, the blended number can point you exactly the wrong way.

  5. Accept 'no difference' as a valid, useful outcome

    If the variants finish statistically tied, that's a real answer: this change doesn't matter, keep the simpler version and move on. Forcing a winner out of a flat result is how you ship changes that quietly do nothing.

  6. Ship the winner, then verify it held

    After publishing the winning variant, watch the live conversion rate for a couple of weeks to confirm the lift shows up in reality. If it evaporates, the original 'win' was noise and you've just learned to trust the number less, not more.

  7. Name your stopping rule before the test starts

    Sample size per variant, significance threshold, minimum duration in full weeks, and the single metric that decides. Record all four before launch. The rule is not there to be clever; it is there so that the version of you looking at an encouraging chart on day four cannot quietly overrule the version of you who designed the test.

  8. If you must look early, use a method that expects it

    Sequential testing sets boundaries that account for repeated looks, so checking is legitimate rather than corrosive. What you cannot do is run a fixed-horizon test, peek daily, and stop when it goes green. Looking 20 times turns a 5% false-positive rate into roughly 24% (Armitage, McPherson and Rowe, 'Repeated significance tests on accumulating data', JRSS Series A, 1969), so on a three-week test peeked at daily, about one in four of your shipped winners is noise.

  9. Cap the number of variants, or correct for them

    Five variants at a 5% threshold carry about a 22.6% chance that one looks like a winner by chance. Either test two arms at a time, or apply a correction and accept that you now need substantially more traffic. Running a five-way test on thin traffic and keeping the top performer is the most reliable way to ship noise.

  10. Run through at least one full business cycle

    Weekend and weekday shoppers differ, payday cycles are real, and returning customers react to novelty before they react to design. A test that starts Tuesday and ends Friday has measured a slice of your audience under unusual conditions. Full weeks, always, even when the numbers look decided.

  11. Decide in advance what an inconclusive result means

    It means you keep the control. Write that down before launch, because only about 20% of experiments ever reach 95% significance and inconclusive is therefore the most likely ending. Deciding the rule while the data is still unseen is what stops a null result from being renegotiated into a win.

An illustrative example

Demo data
What StorePilot detects
A variant looks +12% after a day, but the sample is far too small to trust.
The fix it builds & tests
StorePilot holds the call until significance, then recommends a clear decision.
The projected outcome
Example: 'Variant B, 94% confidence, +8.4% revenue/visitor, recommend publish.' (Illustrative wording of an honest result.)

Key takeaways

  • Only 20% of 28,304 real experiments ever hit 95% significance, so most 'winners' were called too early.
  • Roughly 1 in 7 tests produces a real winner. If a tool says you're winning every time, it's lying.
  • Every time you peek at a running test and call it, you raise your odds of a false positive. Set the finish line before you start.
  • A flat result is a real answer. Don't manufacture a winner out of noise.
Founding-merchant offer
$129/mo Free for your first 3 months, plus the playbook and 3 fixes now

Be one of the first stores to get this fixed automatically.

StorePilot is in private testing right now on the best-performing stores from the 180+ Shopify builds EVDEV has shipped. It runs on Misha Gavura's CRO playbook — the same judgment behind those 180+ builds — and works autonomously: it watches, finds the friction, builds the fix, tests it honestly, and brings you the winner. There is nothing like it on the Shopify App Store.

Leave your email by October 15 and three things happen, in this order.

  1. Right now The Shopify CRO Playbook 10 pages of field notes from 180+ builds and an audit of 5,858 live stores. Instant download.
  2. Within days 3 CRO fixes on your store, by hand Misha looks at your actual store and does three optimizations himself. No app needed, no charge.
  3. At launch 3 months of StorePilot, free The full app on your store the day it is ready for you, with founding pricing locked for life.
  • 180+ Shopify stores shipped
  • In live testing on real client stores
  • Works autonomously, 24/7
-- days
-- hrs
-- min
-- sec

Founding spots close October 15, 2026. Your 3 free months start when your store gets access.

  • The CRO playbook, right now. 10 pages on where Shopify stores leak money and the rules that catch it. Downloads the moment you enter your email.
  • 3 CRO fixes by hand, before the app. Misha Gavura looks at your store and does three optimizations himself while we finish the build.
  • Your first 3 months free. From the day your store gets access. Everything unlocked, and Shopify handles billing, so no card details here.
  • Done-for-you setup. We install and configure StorePilot for your store and catalog.
  • Expert-reviewed first tests. Misha Gavura — whose CRO playbook the agent runs on — checks your first A/B tests by hand before they ship.
  • A real human, not a chatbot. Direct support from the team, every message read by a person.

Sign up now, keep these forever

  • Founding price, locked for life. When paid plans turn on, you keep a permanent founding rate that never goes up.
  • Every new feature, included. Founding members are grandfathered into everything we ship next, at no extra cost.
  • Founding-member priority support. A direct line to the team for as long as you run StorePilot.

Real people, not a black box

Misha Gavura, Senior CRO · EVDEV

Misha Gavura

Senior CRO · EVDEV

Top Rated Plus · Upwork

“I set StorePilot up on your store myself and review your first A/B tests by hand: the setup, the stats, the call, before anything ships. Founding merchants get me directly.”

Never miss a revenue leak

We ping you the moment there's a new opportunity worth testing, with the projected dollars. No dashboard to babysit.

Get the playbook, claim your founding spot

The PDF downloads straight away. The 3 hand-done fixes follow within days, while we finish the app.

Free, no card. The playbook downloads straight away. No spam, unsubscribe anytime.

  • No credit card
  • Nothing to install yet
  • Fully reversible
  • Cancel anytime

Founding deal for the first stores to install.

Frequently asked questions

Why won't StorePilot just give me a fast answer?

Because a fast wrong answer costs you real money. StorePilot optimizes for decisions you can trust, with the timeline shown up front.

How long should I run a Shopify A/B test before trusting it?

Run it until it reaches the sample size you calculated AND covers at least one full week (usually one to two weeks minimum) so weekend, weekday, and payday traffic are all in the data. Duration alone isn't enough; a low-traffic store may need several weeks to gather enough conversions to call anything.

What's a good statistical significance level for store experiments?

95% confidence is the standard bar, meaning there's roughly a 5% chance the result is a fluke. Below that you don't have a result, you have a hint, and per Convert's data, only about 20% of real experiments ever even reach 95%.

Can I run an A/B test with low traffic?

You can, but you can only detect large effects. A few hundred sessions a week means small lifts will never reach significance in a reasonable window, so test bold, high-impact changes instead of button colors, or expect to run for many weeks.

Why did my A/B test winner stop working after I shipped it?

Almost always because it was never a real winner. It was called before reaching significance, so the 'lift' was random variation that regressed to the mean once it went live. This is exactly why honest testing holds the call until the numbers settle and then verifies the result post-launch.

What is peeking in A/B testing and why is it such a problem?

Peeking is checking results repeatedly and stopping as soon as they look significant. It is a problem because a fixed-horizon significance test assumes you look once, at the end. Evan Miller's analysis shows that monitoring continuously and stopping at the first significant reading pushes the false positive rate to 26.1% against a nominal 5%. You are not accepting a 1-in-20 chance of being wrong; you are accepting better than 1 in 4.

How many variants can I test at once?

Two, unless you are correcting for it and have the traffic to spare. Each additional arm is another chance for noise to look like a winner: at five variants and a 5% threshold there is roughly a 22.6% probability that at least one appears to win by chance alone. Multivariate tests are not wrong in principle, they are just expensive in traffic, and most Shopify stores do not have it to spend.

What is a novelty effect and how do I avoid being fooled by it?

It is the early bump you get because returning shoppers notice something changed, not because the change is better. It fades. The defence is duration: run through at least one full business cycle, and if you have the traffic, compare new visitors against returning ones. If the lift lives entirely in returning visitors and decays week over week, you measured surprise.