A lot of advice about onsite A/B testing comes from gut feeling: change your button color, add a countdown, never show a popup on the first visit. We wanted to see how that advice holds up against real data, so we looked at every controlled onsite experiment run on Wisepops over the past two and a half years.
Every experiment is a real test on live traffic, counted in full between January 2024 and June 2026. This report covers how businesses set their tests up, how often a test lands a clear answer, how long it takes, and how much the winners gain. Read it as a benchmark for your own program.
Most A/B tests never crown a winner. And that's the point. You don't test to be right every time; you test so the winners, which lift results by a median of 58%, actually show up.
When a test wins, it usually wins big
Onsite A/B testing is accelerating
The pace of onsite A/B testing has climbed sharply. In the first half of 2026, businesses ran experiments at roughly twice the monthly rate of the previous two years.
A third of every experiment in this report ran in just the last six months. More businesses are testing, and the ones that test are running experiments more often.
How businesses set their A/B tests up
Onsite experiments lean to the two-version test far more than to multivariate designs, which pit three or more versions against each other at once.
In 83% of experiments, the design is a straight A versus B, one version against another. The rest go multivariate with three or more versions, and a small number run far more than that at once. A two-version test isolates a single change and reaches a verdict faster than splitting traffic across many variations, which is part of why it stays the default.
A control-group win looks smaller because the bar is higher
Every test needs something to compare against, and that baseline comes in two forms that answer different questions.
The common one pits a new version of a campaign, with a different headline, offer, or trigger, against the original. The other holds back a control group: a slice of visitors who see no campaign at all, which isolates the campaign's real effect, for instance whether a discount popup actually lifted revenue. Each answers a different question, and each lands on a different median uplift (how much the winning version beat its baseline), shown below.
The control group sets a higher bar: proving a campaign moves a metric like revenue at all is a tougher test than showing one version edges out another. That is why its median sits closer to +20%, and the sample is small enough to read as directional.
The two are far from evenly used. Almost every test compares one version against another; only a small fraction hold back a true control group, as the breakdown below shows.
How long does an A/B test take?
A/B tests tend to conclude quickly, though some run for months.
About a third of tests conclude within a week, and roughly half within two weeks. How long a test takes mostly comes down to traffic: a busy page gathers enough visitors to call a result in days, while a quieter one needs longer to reach the same confidence. Around one in six run for 60 days or more, usually on low-traffic pages or where the effect is small.
What happens when a test concludes
A concluded test can do one of three things: keep the version already live, end inconclusive, or declare a clear, significant winner.
Keeping the live version means the tested change did not beat what was already running, which is still a useful result. Inconclusive means the data could not separate the versions, usually too little traffic or too small a difference, and a clear winner beat the baseline by a statistically significant margin.
Clear winners are the minority, and that is expected: a test that rules an idea out is doing its job just as much as one that names a new winner. Among the experiments that did produce a clear winner, 83% reached statistical significance (p < 0.05), and 79% across H1 2026.
How much do winning tests gain?
The gain is how much a winning version beat the baseline on the chosen metric, and it is usually a lot more than marginal.
We report medians rather than averages: a handful of tests posted gains above 20,000%, which would make the mean meaningless. The typical winner lands between +32% and +195%, and single-digit wins, though they happen, are the uncommon case.
One precision point: that +58% comes almost entirely from version-versus-version tests, so it measures the value of optimizing one version against another rather than the lift over showing nothing at all. For that, look back at the control group in section 04, where the median sits closer to +20%. Both numbers are real; they answer different questions.
Which metric gains the most when a test wins
The three metrics businesses optimize for gain at very different rates.
Click-through is the most-tested metric, and it gains the most when a test wins. Revenue moves nearly as much on a much smaller sample, while conversion is the hardest to shift but the most reliably significant when it does move.
What the winning variations have in common
Across every experiment that produced a significant winner, the same patterns recur in the winning variations, holding across formats, industries, and languages.
They are directional rather than firm rules. Context still decides, so they describe what has tended to work rather than what always will.
Lead with an image over text
Image- and hero-led creative tends to beat text-only layouts, with two exceptions the data insists on: a static image often beats an autoplay video, and an oversized hero sometimes loses to a compact one.
Appears in ~21% of winnersBreak a big ask into steps
Two-step teasers, yes/no micro-commitments, and short multi-step forms tend to beat the single-step ask. Lowering the friction of that first click does more than asking for everything up front.
Strong among the top winnersA real offer beats a clever mechanic
An explicit offer is the most common ingredient in winning variations. A guaranteed fixed discount repeatedly beat a gamified spin-the-wheel, though gamification still beats a plain form.
~28% mention an offerFull-screen and centered for capture
For email and lead capture, centered overlays and full-screen takeovers tend to beat bottom-corner slide-ins and compact banners. The effect is strongest on mobile, where a lot of onsite engagement happens.
Strongest on mobileTrigger on behaviour over a timer
Early and behaviour-based triggers tend to beat arbitrary time delays. An immediate first-page trigger beat a 20-second delay; an "after three pageviews" trigger beat a fixed 5-second timer.
~17% mention timingWhat you change matters more than the format
The surface barely moves the result: popups and bars both land near a median of +60%. The size of a win comes from what you change rather than the format you change it in.
Holds across formatsHow to run your tests
- Choose one primary metric before you launch Pick the single number you're optimizing, whether click-through, revenue, or conversion, and size what counts as a win to it. Click-through swings most (median +63%); conversion moves least (+24%).
- Start with a straight A/B There’s a reason 83% of experiments are two-way: more variations split your traffic across more arms and push the verdict further out. Add a third or fourth only once you have the volume.
- Run for two weeks, and avoid early calls Half of tests need two weeks to conclude, and about one in six runs 60 days or more. Reading results in the first few days is the most common way noise gets mistaken for a winner.
- Use a control group when the question is value Comparing two versions tells you which design wins. Holding back a group that sees no campaign tells you whether it's worth running at all, a median of about +20% here. Reach for it when you need to prove value.
- Test in volume, and expect plenty of draws Only about one in nine concluded tests produces a clear winner, and that's normal. The inconclusive runs are the price of the ones that land, and the winners gain a median of +58%.
How we measured the data
Population. All CAMPAIGN-scope onsite experiments run on Wisepops from January 2024 to June 2026, counted in full rather than sampled. Everything is aggregated and anonymized.
Definitions. A test is concluded once it has finished running. A winner is a concluded test with a declared winning version, and significant means the result cleared a p-value below 0.05. The gain is how much the winning version beat the baseline on the primary metric, reported as a median, the midpoint of all results rather than the average, so a few extreme outliers do not distort it. Decisiveness (how often a test produces a winner) and payoff (how large those winners are) are measured separately and kept apart.
Two kinds of control. The baseline is either another version of the campaign, which measures optimization, or a control group, a held-back slice of traffic that sees no campaign, which measures incremental value. The two are never mixed. About 247 experiments used a control group, including 69 three-way "A vs B vs control" tests, so the +58% headline is overwhelmingly version-versus-version, while the +20% figure compares against a control group.
Formats. The all-time median across every format is +58%. Popups (n=473) and bars (n=19) each sit close to +60% on their own, and a few less common formats with very small samples (CTA, one-click, embeds, recommendations) carry lower medians that pull the pooled figure down. The two figures are consistent.
Statistics. Significance uses Welch's t-test (ratio) and z-test (ratio). Uplift is reported as medians with interquartile ranges; means are omitted because extreme outliers (gains above 20,000%) make them meaningless. Revenue and conversion samples are smaller (tens of tests) than click-through (hundreds), so read those medians as directional.