Shopify CRO agency vs software: what each one actually does
Not a pricing argument. A capability one: which parts of CRO need a human, which parts a tool does better, and how to tell which you actually need.
In short
- CRO is five jobs. Agencies are better at research, strategy, and bespoke build. Software is better at continuous measurement and honest test analysis.
- Break-even lift equals monthly fee divided by monthly revenue. At 30,000 a month a 3,000 dollar retainer needs 10%; at 500,000 it needs 0.6%.
- Judge on cost per test attempt, not cost per month. With roughly 1 in 7 tests winning and only 20% reaching significance, attempts are what you are buying.
A monthly retainer bills whether or not a test wins.
Illustrative. Real lift is measured on your traffic first.
The useful version of this comparison is not a pricing table. It is a list of the jobs CRO actually involves, with an honest note on which ones need a human and which ones a tool does better. There are five: qualitative research, strategy and brand judgement, design and build, continuous measurement, and honest test analysis. Agencies are strong at the first three. Software is strong at the last two. Almost every bad buying decision in this category comes from paying agency rates for the last two, or expecting a tool to do the first two.
What's the problem?
You've decided your store needs CRO. Now you have to choose between hiring an agency and buying software, and every page you read is written by one or the other. What's missing is the boring version: a list of the jobs CRO involves, and an honest note on which of them a firm does better and which a tool does better.
Why does this happen?
- 'CRO' names a bundle of five different jobs, and agencies and tools are strong at opposite ends of that bundle.
- Agencies lead with research depth and bespoke strategy, which is real, and quietly bill for the repetitive measurement work too.
- Tools lead with automation and price, and tend not to mention that qualitative research and brand judgement are not things software does well.
- Comparisons get written as pricing tables, which hides the actual question: how much of this work is judgement, and how much is measurement?
- Agencies win on judgement, and that is not a soft advantage. Deciding that your problem is a positioning failure rather than a button colour, running customer interviews, and rebuilding how a brand explains itself are jobs that need taste and experience. No tool does this, and any tool claiming to is selling you a template.
- Software wins on repetition and patience. Watching every session, re-ranking leaks as traffic shifts, and refusing to call a winner before the numbers support it are jobs that reward doing the same thing thousands of times without getting bored. A human on a monthly reporting cadence is structurally worse at this, not because they are worse practitioners but because they are billing by the hour.
- The base rate makes the split matter more than it looks. VWO puts roughly 1 in 7 A/B tests at a winner, and across 28,304 experiments Convert found only 20% ever reached 95% significance. Since most attempts do not pay, the cost per attempt drives the whole economics, and that is precisely the axis where automation and human hours diverge most.
- Store size flips the answer, and the flip is arithmetic rather than opinion. Divide the monthly fee by your monthly revenue to get the lift the engagement must produce to break even. Against Clutch's listed agency minimums of 1,000 to 10,000 dollars and up, a store doing 30,000 a month is often looking at a double-digit required lift, while a store doing 500,000 is looking at low single digits. Same agency, same quality, opposite decision.
What does the research show?
Independent researchFigures below are from independent studies, not StorePilot data. They're why this problem is worth testing on your own store.
-
Clutch's public directory of conversion optimization agencies lists hourly rates from 50 to 199 dollars and stated minimum project sizes from 1,000 dollars to 10,000 dollars and above.
Clutch, Conversion Optimization Agencies directory ↗ -
Only about 1 in 7 (roughly 14%) of A/B tests produces a winning variation, so cost per attempt, not cost per month, is the number that decides the economics.
VWO ↗ -
Across 28,304 experiments run by Convert customers, only 20% reached the 95% statistical-significance threshold.
Convert ↗ -
Baymard's benchmark of 335 leading ecommerce sites finds the average site needs 32 unique checkout improvements, against a documented 70.19% average cart abandonment rate.
Baymard Institute, Checkout Usability research ↗
How does StorePilot AI fix it?
- StorePilot is deliberately built for the measurement half of the bundle: continuous behavior analysis, leak ranking, variant generation, and honest significance testing.
- It doesn't pretend to replace qualitative research, brand strategy, or a redesign. Those stay human jobs and the app says so.
- Because the measurement layer runs continuously, an agency using it spends its hours on the judgement half instead of on data pulls and reporting.
- For a store that can't justify a retainer at all, it covers the jobs that would otherwise simply not happen.
How do you fix it, step by step?
-
Name the job you are actually short of
Write down which of the five jobs your store is missing right now. If you cannot explain why a shopper should pick you, that is a strategy gap and no tool will close it. If you know what is broken but nothing ever gets measured properly, that is a measurement gap and an agency retainer is an expensive way to fill it.
-
Run the break-even calculation on every option
Monthly fee divided by monthly revenue gives the lift each option must produce to cover itself. Do it for the agency quote and for the software subscription on the same sheet. The gap between those two numbers is usually larger than any difference in capability the sales calls will discuss.
-
Convert the fee into cost per test attempt
Ask how many tests per month the engagement includes, then divide. Apply the roughly 1-in-7 win rate to get a rough cost per winner. Do the same for the tool. This is where automation and hourly work separate hardest, and it is the comparison neither side volunteers.
-
Check who is allowed to say 'we cannot call this'
Only 20% of experiments reach significance, so any honest partner will regularly report inconclusive results. Ask directly what happens in that case. A partner whose incentive is to show monthly wins will find wins, and that is the failure mode that costs you more than the fee.
-
Decide whether you need both, and in what order
For many stores the honest answer is a fixed-scope strategy engagement once, then software running the measurement loop continuously afterwards. That sequence buys the judgement you cannot automate and stops paying hourly rates for the repetition you can.
-
Re-run the decision when your revenue changes materially
Break-even lift moves with your top line. A retainer that was indefensible at 40,000 a month is straightforward at 400,000, and a tool that was the obvious call at the bottom may become the smaller half of your programme at the top. Put the calculation back on the table annually.
An illustrative example
Demo data- What StorePilot detects
- A store needs a full brand and messaging rethink, plus ongoing testing of buy-flow friction.
- The fix it builds & tests
- Those are two different jobs. The messaging rethink is agency work with qualitative research behind it. The buy-flow testing is a loop that software runs cheaper and more often.
- The projected outcome
- Example: the honest answer is often 'both, for different things,' and knowing which half you're buying is what stops you overpaying for one of them. (Illustrative of the decision, not a recommendation for your store.)
Key takeaways
- CRO is five jobs. Agencies are better at research, strategy, and bespoke build. Software is better at continuous measurement and honest test analysis.
- Break-even lift equals monthly fee divided by monthly revenue. At 30,000 a month a 3,000 dollar retainer needs 10%; at 500,000 it needs 0.6%.
- Judge on cost per test attempt, not cost per month. With roughly 1 in 7 tests winning and only 20% reaching significance, attempts are what you are buying.
- The common right answer is both, in sequence: buy judgement once, automate the measurement loop continuously. Paying agency rates for repetition is the expensive mistake.