Shopify SimGym · AI shoppers · 2026
Shopify SimGym Review: How It Works and How Accurate It Is
Last verified: September 3, 2026
Shopify SimGym is a free-to-install Shopify app, still in AI Research Preview, that sends AI shoppers through your live theme and a draft theme and reports which one produced more add-to-carts. A single theme analysis costs 1 credit and a comparison costs 2; credits cannot be bought outright, Shopify may grant free ones, and otherwise you pay per run on your Shopify bill (Help Center, read September 3, 2026). On Shopify's own 50-shop evaluation it called the direction of the add-to-cart change correctly 77% of the time, with a confidence interval of 66% to 87%, and it has never been validated against orders (arXiv:2605.19219v1).
What Shopify SimGym is, who can install it, what a simulation costs, how to run one from the admin, what it cannot test, and how accurate its AI shoppers are according to the two papers Shopify published on it.
Shopify SimGym is the first-party app that lets you test a storefront change on AI shoppers before any real customer sees it. This page covers what it is, who can install it, what a simulation costs, how to run one, what it cannot test, and then the part most write-ups skip: how accurate it actually is, using the two papers Shopify published on it.
Every product fact below was read from Shopify's Help Center, App Store listing, changelog, or engineering blog on September 3, 2026. Every accuracy figure comes from a Shopify research paper, cited by table. Where we do arithmetic of our own, we say so.
What SimGym is
SimGym is a Shopify app that runs AI shoppers through your online store in real browsers and reports what they did, chiefly whether they added a product to cart. It can analyze one theme on its own or compare your live theme against a draft theme, and for a comparison it names a winner based on which theme produced the higher add-to-cart rate. Shopify calls it an AI Research Preview, which means it is public, unfinished, and priced at Shopify's discretion.
The Help Center's definition is one sentence: "a Shopify app in AI Research Preview that uses AI-powered shoppers to simulate buyer behavior on your online store." The App Store tagline is "Simulate buyer behavior with AI shoppers. Gain the confidence of high traffic brands before launch."
Shopify announced it in the Winter '26 Edition on December 10, 2025, describing models trained on "real commerce sessions across our platform" that "approximate what shoppers might do and feel as they land on your storefront." The listing went live on December 8, 2025. On March 11, 2026 the changelog opened it "to all eligible merchants with no waitlist required."
Under the hood, per Shopify's engineering post of February 27, 2026, the platform runs "up to 2,000 concurrent Chromium sessions" through Browserbase and can spawn "400,000 shopping sessions a day." These are real browsers loading your real theme. That explains both why the results are plausible and why the tool can show up in your analytics.
Who can use it: eligibility and plan requirements
TL;DR Four conditions from the Help Center. No Shopify plan requirement is listed.
The Help Center lists these requirements, quoted directly:
- "Your store uses a Liquid storefront. Hydrogen and headless storefronts aren't supported."
- "Your store has Shopify Network Intelligence (SNI) activated."
- "Your store access isn't set to private mode."
- "Your store has at least one product listing."
For a theme comparison there is a fifth: "Theme comparisons require your active theme and at least one other theme from your Draft themes."
There is no plan gate on the page. Basic, Shopify, Advanced and Plus are not mentioned at all, which is different from Rollouts, where Shopify's changelog states "Rollouts is available for merchants on Basic plans and above."
Note the private-mode condition: you cannot simulate a store before it is public, so in practice SimGym tests a draft theme on a store that is already live.
The Shopify Network Intelligence condition is the one to read twice
Shopify Network Intelligence is switched on under Settings > Customer privacy in the admin. Shopify's requirements page describes it as Shopify using "your customer data together with customer data from other merchants to generate insights that power advanced features, known as Enhanced Services."
Turning it on creates obligations. Per that page you must "Include a link to Shopify's Consumer Privacy Policy in your privacy policy, and post a link to your privacy policy prominently in your store," explain that visitor data "will be shared with Shopify and other third parties that may be located in other countries," offer an opt-out in certain US states, and in the EEA, UK and Switzerland "Obtain consent from customers to see ads based on their activity in your store, with other merchants, and with Shopify."
So for an EU merchant the real cost of a first run is a credit plus a privacy policy review. Neither the listing nor the SimGym Help Center page says this.
What it costs
TL;DR Free to install. One credit per single-theme analysis, two per comparison. Credits cannot be bought. No credits left means you pay per run on your Shopify bill, at a price the app shows you before you start.
The App Store listing says "Free to install. Additional charges may apply" and, under the pricing block, "Charges per simulation run. More details available in the app." So "free" is accurate for installing and inaccurate for using.
The Help Center is more specific. Quoting it in order:
- "SimGym uses a pay-per-use pricing model, and you're charged only for a simulation that completes successfully."
- "A single theme analysis costs 1 credit."
- "A theme comparison costs 2 credits."
- "Credits can't be purchased."
- "During the research preview, you might be granted free credits."
- "If you don't have enough credits, then you can still run a simulation and pay for it on your Shopify bill."
- "SimGym displays the cost of a simulation, in credits or as an amount, before you start it."
What Shopify does not publish anywhere we could find is the dollar price of a paid run. The engineering post says "the cost per merchant run is in single digits and falling," but that is Shopify's compute cost, not your price. One agency write-up (Blend Commerce, July 23, 2026) reports $10 USD per simulation once free credits are used. We could not confirm that against a Shopify page, so treat it as a reader report and check the figure the app shows you on the start screen.
How to install and run a simulation
TL;DR Install from the App Store, open Apps > SimGym, click Create simulation, pick a type, name it, choose a theme or two, optionally a focus area, and start. It queues and runs on its own.
Before you install
- Check your theme is a Liquid theme (the normal Online Store theme editor). If you are on Hydrogen or headless, stop here.
- Go to Settings > Customer privacy and confirm Shopify Network Intelligence is activated. Read the obligations above first if you sell into the EU or UK.
- Make sure the store is not in private mode and has at least one product.
- If you want a comparison, save the change you want to test as a draft theme under Online Store > Themes. Duplicate your live theme, make the change, and leave it unpublished.
Install and run
- Install Shopify SimGym from the App Store. The changelog's instruction is "Install the Shopify SimGym app from the Shopify App Store to get started."
- In the admin, go to Apps > SimGym.
- Click Create simulation.
- Choose a simulation type. The Help Center defines the two: "Theme comparison: Compare your live theme against another installed theme to analyze how theme changes affect shopper behavior," and "Single theme analysis: Analyze any live or draft theme on its own to get feedback on what's working and where to improve."
- Enter a simulation name.
- Select the theme (or, for a comparison, your live theme plus a draft).
- Optionally pick a focus area. The Help Center: "select a Focus area to direct AI shoppers to a specific part of your store: Home page, Products, Collections, Cart, or Search."
- Check the cost shown, then click Start simulation. "Your simulation is added to the queue and will begin shortly."
Shopify's docs do not say how long a run takes. Digismoothie's March 14, 2026 test reports about 30 minutes for theirs, and the research paper's best configuration averages 5.3 minutes of agent time per shop, so expect somewhere in that range plus queue time.
What comes back
For a comparison, "The winning theme is the one with the highest add-to-cart rate." For a single-theme analysis, "Results include a summary of the shopping experience, strengths, areas for improvement, and prioritized recommendations ordered by conversion impact." The changelog adds that reports cover "navigation, product discovery, add-to-cart flows, and performance indicators" and that the tool "recommends which theme performs better."
Two things the report does not include, and they matter later in this article: a confidence interval on the add-to-cart gap, and any statement of which AI model or how many agents ran your simulation.
What it can and cannot simulate
TL;DR It simulates browsing, product discovery and add-to-cart on a whole theme. It does not simulate checkout, purchases, pricing, discounts, or anything after the cart, and it tests themes, not isolated elements.
| Can simulate | Cannot simulate |
|---|---|
| A draft theme against your live theme | Two draft themes against each other (comparisons require your active theme) |
| A single live or draft theme on its own | A single section or element in isolation; the unit is the whole theme |
| Navigation, product discovery, add-to-cart behaviour | Checkout, payment, purchase completion; the papers validate add-to-cart only |
| Focus on Home page, Products, Collections, Cart or Search | Pricing, discounts, shipping rules, checkout configuration |
| A public Liquid storefront with at least one product | Hydrogen, headless, or a store in private mode |
| Shoppers modelled on your store's own traffic | Shoppers you define yourself; there is no persona editor |
The "whole theme" point trips people up. Digismoothie's test removed quantity bundles from a product page and got back feedback on parts of the theme they had not touched. That is by design. To isolate one change, make the draft differ from the live theme in exactly that one way.
The "no persona editor" row is the one that decides whether SimGym works for your store, and we come back to it in the accuracy section.
How SimGym pairs with Rollouts
TL;DR SimGym predicts. Rollouts measures. Shopify shipped them together as a pair and that is exactly how to use them: simulate to pick a candidate, then run the real test.
Shopify's Winter '26 Edition grouped them in one breath: "more confident launches with Rollouts and SimGym." Rollouts is the native feature that, per the March 31, 2026 changelog, lets you "Run an A/B test to compare your changes against your current theme with real traffic, review the impact on conversions, and select the winning version." It lives under Markets > Rollouts and is "available for merchants on Basic plans and above."
The two tools work on the same object, a draft theme against your live theme, which makes the handoff clean:
- Build two or three draft themes with different approaches to the same problem.
- Run a SimGym comparison on each. Read the direction of the add-to-cart result, not the size of it.
- Take the draft that won, or the one that did not lose, into a Rollouts A/B test with real traffic.
- Judge the Rollouts result on revenue per visitor and check significance yourself, because Rollouts shows side-by-side performance and leaves the call to you.
Our Shopify Rollouts guide covers what it can and cannot test, the plan gates, and where it stops short of an honest result. The rest of this page is about why step 2 needs the word "direction" in it.
How accurate is SimGym? What Shopify's own papers say
TL;DR 77% directional accuracy on add-to-cart, with a 95% confidence interval of 66% to 87%, and a 0.55 correlation with the size of the change, interval 0.32 to 0.72. That is the best published configuration on 50 shops. Nothing was validated against orders.
Shopify's researchers published two papers on SimGym. Paper A (arXiv:2602.01443v1, February 1, 2026) validated on 20 shops in 12 countries. Paper B (arXiv:2605.19219v1, May 19, 2026) grew that to 50 shops in 16 countries, added vision to the agents, and published confidence intervals. Both use 600 agents per shop.
The validation design is better than most vendor case studies. Shopify found real stores that had switched themes, measured what real customers did before and after, filtered out stores with promotions, seasonality, pricing or assortment changes, checked that filtering "using double machine learning to confirm consistent treatment effect estimates," and then asked whether the agents predicted the same direction of change.
The headline table
| Configuration | Direction right | 95% CI | Correlation with size | 95% CI | Runtime/shop |
|---|---|---|---|---|---|
| Gemini 3 Flash (with vision) | 77.0% | 66.0 to 87.0 | 0.55 | 0.32 to 0.72 | 5.3 min |
| Gemini 3 Flash (text only) | 70.0% | 58.0 to 82.0 | 0.49 | 0.25 to 0.68 | 4.5 min |
| GPT-OSS (text only) | 59.0% | 47.0 to 71.0 | 0.41 | 0.15 to 0.62 | 12.8 min |
Read the top row two ways. Direction: three times in four the agents said "this theme wins" and real customers agreed. That is a real signal. Magnitude: a correlation of 0.55 means the simulator explains about 30% of the variance in how big the real change was, and the interval runs from 0.32 to 0.72, so the same evidence is consistent with anywhere from about 10% to 52%. If SimGym tells you a draft theme lifted add-to-cart by 14%, the honest reading is "the draft was better." The 14% is not a forecast.
Now read the bottom row. The GPT-OSS interval runs from 47% to 71%. Its lower bound is below a coin flip.
Which model runs your simulation?
That bottom row matters because Shopify's engineering post of February 27, 2026 documents the production stack at that time as "48 dedicated NVIDIA B200 GPUs" running gpt-oss-120b, reading "a representation tree of the current page" rather than screenshots. That is the text-only GPT-OSS configuration, which Paper B scores at 59%.
We cannot say that is what runs today. Three months separate the engineering post from the paper, and Shopify may have moved production to the vision configuration. What we can say is that Shopify has not published which model or how many agents serve a merchant's run, and the report does not tell you. The gap between the top and bottom rows is the gap between a useful screen and a coin flip with a written report attached. It is a fair question to put to Shopify support.
The three numbers people quote, and which one to use
You will see 69%, 73% and 77% cited as SimGym's accuracy. 69% with 0.64 correlation is Paper A's headline on 20 shops (Tables 2 and 5). 73% with 0.65 is from Paper A's Section 4.1, a bootstrap measuring how consistent two runs are with each other, not with humans. 77% with 0.55 is Paper B's headline on 50 shops. Quote 77%, and note that magnitude correlation fell from 0.64 to 0.55 as the sample grew, the usual sign of a small-sample estimate being corrected.
It was validated on add-to-cart, not sales
Paper B: "We use A2C rate as the primary outcome, as it directly reflects purchase intent." Neither paper reports orders, revenue, or revenue per visitor. The agents never complete a checkout.
Add-to-cart is a proximate metric. Jakub Linowski's GoodUI analysis of November 30, 2023 puts the correlation between adds-to-cart and orders at R = 0.4983 across 44 experiments. So add-to-cart explains about a quarter of the variance in whether orders move.
Chain the two links and the implied correlation between a SimGym prediction and your actual orders is roughly 0.55 × 0.50, or about 0.27. That is our arithmetic, not a published finding, and it assumes the simulator carries information about orders only through add-to-cart, which is reasonable because the agents never buy. The two coefficients also come from different samples. Treat 0.27 as an order-of-magnitude ceiling. Nobody has measured SimGym against orders, and that is the point.
What makes it work, and what breaks it
Paper B strips components out one at a time. The results are stark.
| Configuration | Direction right | Correlation with size |
|---|---|---|
| Full persona, with memory | 77.0% | 0.55 |
| Shopping intent only | 44.0% | 0.02 |
| Product only | 51.0% | -0.14 |
| No session memory | 42.0% | 0.00 |
Two components carry all of the signal: personas grounded in your store's own clickstream, and session memory. Remove either and the correlation drops to zero. Paper A measured the memory effect from the other side: with memory, agents reached their goal 90% of the time and timed out 9.59% of the time; without it, 45% and 54.93%. An agent stuck in a navigation loop never reaches the part of your theme you changed.
The persona finding is the one that predicts merchant complaints. Paper A tested a "generic persona" built from similar stores instead of yours and found it produced the highest rate of agents that could not find a product, 43.56% of diverging agents against 36.36% for the properly grounded version. The paper's own line: "the issue is not how agents search but who is searching." Personas come from your clickstream and there is no way to edit them, so if your store is unusual, the grounding misfires and the paper's numbers say you slide from 0.55 toward 0.02.
How many AI shoppers is enough?
Both papers use 600 agents per shop. Paper B shows why: directional alignment rises from 67% at 50 agents to 73% at 300, with little gain after, and the spread between the 10th and 90th percentile narrows from 0.17 at 50 agents to 0.05 at 600. Sample size governs a synthetic test exactly as it governs a real one. Shopify does not publish how many agents a merchant credit buys, so if a paid run uses far fewer than 600, the published accuracy does not describe your run.
What merchants report
TL;DR 2.8 out of 5 across 37 App Store reviews as of September 3, 2026. The complaints map almost exactly onto the failure modes Shopify measured in its own papers.
The listing's structured data reads a 2.8 rating on 37 reviews today. Reviews skew negative on any research-preview product, so the score itself is not the interesting part. The pattern is.
Agents shop for things you do not sell. A US merchant, June 15, 2026: "searched for things we do not even sell." A Canadian merchant, May 20, 2026: "Unable to get accurate results as the AI Shoppers are based on a location our website doesnt support." An Australian university store, April 13, 2026: "without being able to customise Customer Persona's there is only a small benefit." That is Paper A's Product Not Found category showing up in the wild, and the reviewers have correctly identified that they cannot fix it.
Confident false positives. The same June 15 reviewer: it "claimed our Checkout button is broken, when it is clearly not." A Japanese merchant, June 30, 2026, reports footer links flagged as not switching pages when they were transitioning correctly. A German merchant, July 19, 2026, gives the sharpest version: the tool appears to test against a stale snapshot, re-reporting problems already fixed, and cannot distinguish real site bugs from its own test-session artefacts.
It is not a real test. A Cyprus merchant, January 24, 2026, did the comparison this article is about: "we have noticed very inaccurate results when comparing to our current active AB tests." Digismoothie's own test, March 14, 2026, removed quantity bundles from a product page and got back a verdict that the original theme was better, with recommendations they describe as rarely making sense. Their conclusion, and ours, is that it works as a validation pass on a full redesign rather than as a conversion-optimization tool.
Operational gaps. Reviewers ask for export ("MAKE A FULL REPORT EXPORT!", UAE, June 27, 2026), reports in the store's language rather than English (France, June 18, 2026), and clearer credits (US, June 27, 2026).
The positive reviews are consistent too. They describe finding friction and getting ideas, not deciding winners. The satisfied users are already using it the way the evidence supports.
The analytics side effect
TL;DR Real browsers on your real storefront can fire your real tracking. One merchant reports SimGym creating anomalies in Google Analytics. Shopify's documentation does not address it.
A UK merchant, March 19, 2026: "Creates an anomaly in your traffic in google analytics, including conversion tracking." The architecture predicts this. Shopify's engineering post describes real Chromium sessions driving live storefronts, and a real browser loading your real theme executes your tags unless something stops it. We have not been able to verify what filtering, if any, Shopify applies, and the SimGym Help Center page is silent on it.
Synthetic cart events arrive in a burst and never become orders, so your conversion rate dips at exactly the stage you were investigating, and any live experiment running at the same time absorbs sessions that never buy.
Hygiene when running simulations
- Write down the start and end time of every run. Without it you cannot clean anything later.
- Never simulate while a live A/B test is collecting data. Finish the Rollouts test first.
- Annotate the window in GA4 so the next person reviewing the period does not panic.
- Exclude simulation windows from any baseline you use to size a future test.
- Check your own numbers. Run one simulation on a quiet day and compare that day's sessions and add-to-cart events with the surrounding week. Ten minutes settles it for your store.
Verdict
TL;DR Use SimGym where being wrong is cheap and being fast is valuable: finding friction on a draft and eliminating weak candidates. Then confirm the survivor with real traffic.
SimGym is a screening tool, and screening tools are judged on a different standard from decision tools. A screen is allowed to be wrong sometimes as long as it is cheap and fails in the right direction. At 77% directional accuracy with a 66% floor, it is a poor instrument for promoting a candidate and a decent one for discarding several, because wrongly discarding a variant you never had the traffic to test costs you little, while wrongly shipping a bad one costs you real revenue.
So: build more draft themes than you can afford to test. Simulate them. Read only the direction. Drop the losers. Put the survivor into Rollouts, judge it on revenue per visitor, and run the result through a significance check before you publish. Log every simulation call and its real outcome; after a dozen pairs you will know whether SimGym is systematically wrong on your catalogue, which the persona ablation and the reviews both suggest is possible.
The 30-second decision
- Redesigning a theme and want friction found before launch? Worth a credit. This is what it is good at.
- Choosing between several redesign candidates? Use it to eliminate, then test the survivor for real.
- Have enough traffic to run a real A/B test? Run the real test. A merchant who compared both reached that conclusion in January.
- Need to justify a change to a client or board? Not on simulated data. The confidence interval will not survive a competent question.
- Running a live experiment right now? Do not simulate until it finishes.
- Headless, private mode, or an unusual catalogue? Ineligible, ineligible, and likely to misfire.
What would change our verdict
Four things, and three are disclosure. Validate against orders, not just add-to-cart. Show the confidence interval in the app; Shopify already computes it with a 10,000-resample bootstrap. Disclose the model and agent count per run. Let merchants correct the personas. Only the first needs new research.
Where we sit on this
We build StorePilot, an AI CRO agent for Shopify, so we have an interest here, and it is in development rather than on the App Store, so nothing above is a suggestion to install something of ours instead. We think SimGym is the right idea. Our disagreement is about presentation: Shopify's researchers published their intervals and failure modes, and the product wrapped that in "try new ideas without risk" and one sentence of caveat. The uncertainty belongs on the screen where the decision gets made. Hold us to the same standard when we ship.
If you take one thing from this page: SimGym is a cheap way to be wrong less often about which idea to test next. It is not a way to find out whether you were right. Only one of those jobs can be done without real customers.
Sources
All pages fetched and read on September 3, 2026.
- Shopify Help Center: SimGym. Requirements, simulation types, admin path, credits, caveat.
- Shopify App Store: SimGym listing and reviews. Pricing block, launch date, aggregate rating 2.8 on 37 reviews, review quotes.
- Shopify changelog, 11 Mar 2026. SimGym open to all eligible merchants.
- Shopify changelog, 31 Mar 2026. Rollouts: real-traffic A/B tests, Basic plans and above, Markets > Rollouts.
- Shopify Winter '26 Edition, 10 Dec 2025.
- Shopify Engineering, "2,000 robots walk into a shop," Javier Moreno, 27 Feb 2026.
- Shopify Help Center: Requirements when Shopify Network Intelligence is enabled.
- arXiv:2605.19219v1, "SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents," 19 May 2026. Tables 2, 3, 4; Figure 4.
- arXiv:2602.01443v1, "SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce," 1 Feb 2026. Tables 2, 5, 6; Section 4.1.
- GoodUI, Jakub Linowski, 30 Nov 2023. Adds-to-cart vs orders, R = 0.4983, n = 44.
- Digismoothie, SimGym review, 14 Mar 2026. Hands-on test, run time.
- Blend Commerce, 23 Jul 2026. Reader report of $10 USD per simulation; not confirmed against a Shopify page.
Frequently asked questions
What is Shopify SimGym?
SimGym is a first-party Shopify app that runs AI shoppers through your storefront and reports how they behaved. Shopify's Help Center defines it as "a Shopify app in AI Research Preview that uses AI-powered shoppers to simulate buyer behavior on your online store." It launched on the App Store on December 8, 2025 and has been open to all eligible merchants without a waitlist since March 11, 2026.
Is Shopify SimGym free?
Free to install, then pay per simulation. The App Store listing reads "Free to install. Additional charges may apply" with "Charges per simulation run." The Help Center says a single theme analysis costs 1 credit and a theme comparison costs 2 credits, that credits cannot be purchased, that free credits may be granted during the research preview, and that with no credits left you can still run a simulation and pay for it on your Shopify bill. The app shows the cost before you start.
Which Shopify plans can use SimGym?
The Help Center lists no plan requirement. It lists four conditions instead: a Liquid storefront (Hydrogen and headless are not supported), Shopify Network Intelligence activated, store access not set to private mode, and at least one product listing. If you meet those, you can install it from the App Store.
How accurate is SimGym?
On Shopify's own best published result, SimGym called the direction of the real add-to-cart change correctly on 77% of 50 test shops, with a 95% confidence interval of 66% to 87% (arXiv:2605.19219v1, Table 2). Its correlation with the size of the change was 0.55, with an interval of 0.32 to 0.72. Direction is predicted reasonably well. Magnitude is predicted weakly.
How is SimGym different from a real A/B test?
A real A/B test splits live visitors between two versions and measures what they actually did. SimGym sends AI agents through both versions and predicts what visitors would do. The prediction was right about direction three times in four in Shopify's validation, and it was only ever checked against add-to-cart, never orders. Shopify's own Rollouts feature runs the real test with real traffic.
Can SimGym replace A/B testing?
No, and Shopify does not claim it can. The Help Center's caveat is that "results might differ from actual buyer behavior." Use SimGym to eliminate weak ideas before they consume real traffic, then confirm the survivor with a real test and a real significance check.
How many AI shoppers does a SimGym simulation run?
Shopify's research papers use 600 agents per shop, chosen because accuracy stops improving much past about 300 agents. Shopify does not publish how many agents a merchant simulation runs, and the Help Center does not state a number.
What does research preview mean for SimGym?
It means the product is public but unfinished and Shopify has not committed to a general-availability date. The App Store listing still reads "Currently in AI Research Preview" as of September 3, 2026, nine months after launch. Practically, expect free credits at Shopify's discretion, features that change, and a support experience that reviewers describe as thin.
Does SimGym affect Google Analytics?
One merchant reports that it does. A March 19, 2026 App Store review from the United Kingdom says SimGym "Creates an anomaly in your traffic in google analytics, including conversion tracking." The agents run in real browsers against your live storefront, so this is plausible by design. Shopify's SimGym documentation does not address it. Note the start and end time of every run and never simulate while a real A/B test is collecting data.
What can SimGym not simulate?
Anything past the cart. The only validated outcome in Shopify's papers is add-to-cart rate, and the agents do not complete purchases. The Help Center also limits comparisons to your active theme plus a draft theme, so it cannot test a price change, a discount, or a checkout configuration. It cannot run on Hydrogen or headless stores, and it cannot run on a store in private mode, which rules out pre-launch testing.