1. Start with one question
Write down what you are changing and why. For example: “Will a shorter hero explanation help more visitors reach the Buy Box?” Use your current page as the control and create a variant with the proposed change. Starting with two variants keeps the comparison easier to manage. For a full-page redesign, treat the result as a comparison between two page approaches—not proof that a particular headline caused the difference.2. Decide what success means
Choose the business result you care about, such as conversion rate or revenue per session. Use supporting metrics to understand the result: more button clicks are not necessarily better if fewer shoppers complete checkout. See Reading results and the Metrics glossary for the measures shown in Jurni.3. Check the experience before launch
Test both variants on mobile and desktop. Verify the products, variants, purchase options, discounts, and final checkout. Enter through the experiment’s campaign link so you are testing the intended routing, not just opening a page directly. If you are comparing against an existing store page, follow Testing against your website.4. Keep the comparison fair
Keep the campaign audience and offer consistent where possible. Avoid editing the pages repeatedly during a test; otherwise, the results describe several versions rather than the comparison you planned. If you must fix a broken page or checkout, fix it. Record when the issue happened and reconsider which results are usable. Do not leave a broken purchase flow running simply to preserve a test schedule.5. Review the result in context
Look at the amount of traffic and the number of completed purchases behind the percentages. A large percentage difference based on very few orders can change quickly. Review a representative period for your campaign rather than reacting to one unusually strong day.There is no universal rule that 500 sessions, three weeks, or a particular daily ad budget makes a result conclusive. A higher metric by itself is not proof of statistical significance. Use only the confidence or significance information actually available in your results, and do not assume a fixed confidence threshold applies to every test.
