> For the complete documentation index, see [llms.txt](https://docs.shoplift.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.shoplift.ai/analyze/reports/statistical-significance.md).

# Understanding Test Outcomes

While a test runs, Shoplift continuously evaluates test results using Bayesian statistical methods. Throughout the life of a test, we provide clear indicators to help you understand how your test is performing and when it’s safe to make decisions based on the results.

### The question your report answers

A/B testing has traditionally been built around statistical significance — a strict bar (usually 95% confidence) borrowed from large-scale academic studies. Most real ecommerce tests never reach it in a reasonable window, so a lot of genuine signal ends up filed away as "inconclusive."

Shoplift focuses on the decision instead. Your report is built around two questions: how likely is this variant to win, and is that read stable enough to act on? Everything below supports those two questions.

### Probability to Win

**Probability to Win** is the main metric on your report. It's the chance your variant beats the original.

The number means exactly what it says. A Probability to Win of 92 means the variant comes out ahead in about 92 of every 100 likely outcomes. The higher the number, the more confident Shoplift is that the variant is the better performer. Either side can win: if your original is stronger, its Probability to Win climbs instead.

As more visitors enter the test, the number settles and becomes more reliable.

<figure><img src="/files/zOn8QPs7gWIDcqRvbn9a" alt=""><figcaption></figcaption></figure>

### Lift

**Lift** is how big the win is: the difference in performance between your original and your variant. Probability to Win tells you *which* variant is ahead. Lift tells you *how much* it's worth.

Lift is shown with a range rather than a single point. The range reflects how much uncertainty is still in the estimate. Early in a test the range is wide. As data accumulates, it narrows and the estimate gets more precise.

Read the two together. A high Probability to Win on a tiny Lift may not be worth the effort to ship. A large Lift is only worth acting on once Probability to Win is high enough to trust.

<figure><img src="/files/wZM806Hmb7B391OT8S7p" alt=""><figcaption></figcaption></figure>

### The test lifecycle: stages and what to do

Every test moves through a set of clear stages. Each one is shown as a color-coded card on your report and in your test list, so you can tell at a glance where a test stands.

#### Collecting data

<figure><img src="/files/8YPR7xPQCSZZIcdh6D2i" alt=""><figcaption></figcaption></figure>

**What you see:** Your test is live, but Probability to Win and Lift aren't shown in a headline position yet. This covers roughly the first day.

**What it means:** There isn't enough data to say anything reliable. Early numbers swing hard from hour to hour, and showing them prominently would only anchor you to noise.

**What to do:** Let it run. Ending a test at this stage tells you nothing.

#### Keep running

<figure><img src="/files/f2upRdYmt5M7tzM0nCFs" alt=""><figcaption></figcaption></figure>

**What you see:** Probability to Win and Lift now appear, but are still not shown in a headline position yet because they haven't stabilized. This starts around day 2, once Shoplift has completed its first daily read of the data.

**What it means:** You can see the direction the test is pointing, but it hasn't held long enough to trust. A test at this stage can show a high number — even 95% — on a single day. Until that level holds for multiple days, it isn't stable.

**What to do:** Keep the test running, and don't act on the number yet, however good it looks. Shoplift will tell you when a result can be acted on.

#### Leaning

<figure><img src="/files/yQY0Cd4ch0U4aRGHOjax" alt=""><figcaption></figcaption></figure>

**What you see:** A blue **Leaning** card. Probability to Win has stabilized and the data indicates a real directional lean.

**What it means:** A directional signal has stabilized. One variant is reliably ahead. This is not a final call, but the lean is consistent enough to be useful.

**What to do:** This is enough to act on when the stakes are low or you want to move quickly. For higher-stakes changes, let the test keep running toward a confident outcome.

#### Test Complete - Winner

<figure><img src="/files/0VwzVJRiBrCVLYwbJWsC" alt=""><figcaption></figcaption></figure>

**What you see:** A green **Test Complete** card showing the winning variant and its Probability to Win. There are two confidence levels:

* **High Confidence** — a confident outcome. You can ship it, or keep running for even higher confidence.
* **Clear Winner** — a clear outcome. You can act on these test results.

**What it means:** The result has held steady long enough to trust. This is a call you can act on.

**What to do:** Ship the winner. If a test reaches **High Confidence** and the decision is high-stakes, you can let it keep running toward **Clear Winner** before committing.

#### Test Complete — No Consistent Difference Detected

<figure><img src="/files/qfsgo4MUBnfWUnCCp9Y9" alt=""><figcaption></figcaption></figure>

**What you see:** A green **Test Complete** card. This appears when a test has run for two full weeks, but no meaningful signal has appeared during its duration (a consistent lean or a winner).

**What it means:** The test ran its course and neither variant made a meaningful difference. This is a real, conclusive result, not a failure. Learning that a change doesn't move the needle saves you from shipping something that wouldn't have helped, or allows you to ship something based on preference, knowing there is no significant difference in performance.

**What to do:** Move on with confidence. Roll what you learned into your next test.

#### Quick reference

| Stage                           | Color | What it means                                                               |
| ------------------------------- | ----- | --------------------------------------------------------------------------- |
| Collecting data                 | White | Too early to show a reliable read                                           |
| Keep running                    | White | Results are stabilizing, but still noisy                                    |
| Leaning                         | Blue  | A real, directional signal has emerged — act on it when stakes are low      |
| Test Complete (High Confidence) | Green | A high confidence outcome you can act on for most tests                     |
| Test Complete (Clear Winner)    | Green | The strongest outcome, indicating a significant difference between variants |
| Test Complete — No Change       | Green | Conclusive: no meaningful difference between variants                       |

### Why Shoplift waits for a stable read

Shoplift updates Probability to Win as new visitors arrive. Early on, that number can move a lot: a variant sitting at 90% today might read very differently tomorrow, simply because so few visitors have been counted.

That's why an outcome isn't called the first time a number crosses a threshold. Probability to Win has to hold at or above a level for multiple days before Shoplift assigns a stage. That stability hold is the guardrail that keeps you from acting on a lucky moment instead of a settled result.

### FAQ

**Probability to Win already looks high — why can't I act on it yet?**

If the stage still says Collecting data or Keep running, the number hasn't held long enough to trust. A single high reading early in a test often comes from a small number of visitors and can swing the next day. Once Probability to Win stays above a threshold for multiple days, the test moves into Leaning or Test Complete, and that's your signal to act.

**Is "No Change" a failed test?**

No. "No Change" is a conclusive result: your two variants performed about the same. That's a genuine learning. It saves you from shipping a change that wouldn't have moved your numbers, and it frees your traffic for the next test.

**What do the numbers around Lift mean?**

Lift is shown as a range because there's always some uncertainty in the estimate. The range is where your true lift most likely falls. A tighter range means more certainty about the size of the effect. A wider range means more uncertainty. The range narrows as your test collects more data.
