Five rules to remember
- Set the end date before the test starts, and decide at that date.
- Pick one primary metric before the test starts, and do not change it later.
- Use two or three guardrail metrics to catch side effects, not to create a second winner rule.
- Check the traffic split before you decide, even when Confidence is high.
- Roll out only when the low end of the uplift range is still worth the work.
Set the end date before you start
Choose the end date when you create or schedule the test, then make the decision at that date. If you roll out the first time Confidence crosses 95%, you will call some tests a winner that are not. Give the test at least one full business cycle, so it covers both weekday and weekend behavior. That is usually 1 to 2 weeks. A high-traffic store may collect enough orders sooner, but still run at least one full week. A low-traffic store often needs longer. Two situations change the plan:- The test is harming your store, such as a large drop in orders or profit. Stop it. An early stop protects your store, but it does not give you a result you can act on.
- The test’s Health is not all green. Pause the test and clear the problem, or contact support if you cannot resolve it. Once it is fixed, duplicate the test and start collecting data fresh, so earlier dirty data does not pollute the result.
Pick one primary metric
The primary metric is the one number that decides whether the test worked. You choose it when you create the test, and the result summary reports Confidence for it.
Do not change the primary metric after the test starts. Every extra metric you judge is one more chance for a random result to cross 95%, so switching raises the risk of a false positive.
For metric definitions, see Analytics metrics.
Add two or three guardrails
Guardrails are the checks that stop you from rolling out a harmful change. They do not need to beat Control. They need to stay healthy enough for the rollout to make business sense.- Primary metric Conversion Rate: guard with Revenue per Visitor or Profit per Visitor, so you can see whether the extra purchases are lower-value or lower-margin orders.
- Primary metric Average Order Value: guard with Conversion Rate and Profit per Visitor. A higher threshold or a bundle can raise order value while fewer visitors buy.
- Running an offer test: guard the discount metrics. A lift bought entirely with discount is a transfer, not a gain.
Check the data before you decide
There is no single visitor count that works for every test. How much data you need depends on your normal Conversion Rate and on how small a change you want to detect. A change smaller than about 20% usually needs several hundred Orders in each test group.
If the split mismatch does not go away, the result is not valid. File a support ticket so we can find the cause, then run the test again.
File a support ticket
In the app, go to Help > Support > Open ticket. Include your store and the test, and we’ll look into it.
Turn the result into a decision
Check Orders and the traffic split first, then use this matrix.
Rollout cost includes setup work, testing, customer support risk, operational changes, and margin or brand tradeoffs. For a small gain, do not roll out a change that needs theme edits, checkout testing, pricing policy changes, new support templates, or extra fulfillment work. If the result is flat, test a bigger change next.
A conclusive loss is a real answer too. Keeping Control on the evidence is a decision, and knowing what a change costs you is worth as much as knowing what it earns.
Common mistakes
- Ignoring guardrails. A test group can win the primary metric and still hurt profit, order value, or purchase likelihood.
- Reading the estimated monthly revenue impact as a promise. It is an estimate from the measured difference, and its range shows how much it could vary.
- Deciding before Health is all green. When a Health check is not passing, the numbers under it may not be measuring what you think. Fix the check first, then read the result.