Your AI conversion tool ships a winner every week. Here's how to tell which ones are real.

by , Founder & Growth Lead

Open the dashboard on a Friday and the AI conversion tool has good news again. Another test called. Another green number next to it — +9%, +14%, converts better. Over a quarter the list gets long: a dozen winners, most of them shipped, each one logged as a lift. Then finance pulls the actual site-wide conversion rate for the same three months, and it's flat. Not down — flat. The wins are all on the board, and none of them are in the total.

Nobody in the room is lying. The tool really did run those tests, and each one really did cross the line it was told to watch for. But a pile of reported wins and a conversion rate that never moved can't both be the truth, and the gap between them is where a forecast quietly goes wrong. If you run or are evaluating an AI conversion tool — anything that "tests variations," "optimizes 24/7," or ships you a winner on a schedule — this piece is about why more tests can grow your win-pile without moving your business, and the one check that tells you which of your winners are real.

The running argument on this blog: the lift comes from research, not auto-tuning, and once a tool has run, the number it hands back rarely means what the slide says it means. That last piece was about a single test — how one reported lift gets contaminated when the tool moves traffic mid-run. This one is about what happens across many tests, once speed is the whole pitch.

Why more tests make more phantom winners

Speed is what you're buying. The sell is that an AI-native conversion tool runs more experiments than your team ever could — while you sleep, across every page, forever. That part is true. The problem is what more tests do to the odds.

Run one A/B test on two genuinely identical pages and, by pure chance, one of them will usually edge ahead. Run twenty such tests and you'll get a fistful of "winners" that were nothing but noise. This is the plain version of what statisticians call the winner's curse: when you pick the variant that looks best out of a big batch, the one that looks best is also the one most likely to have gotten lucky — and its lift shrinks, or vanishes, the moment you look again. More shots on goal doesn't just find more real wins. It manufactures more fake ones, and the two look identical on the dashboard.

Test velocity has a second, quieter multiplier. Most tools watch a test live and call it the instant it looks significant — the practice A/B-testing people call peeking, checking the scoreboard over and over and stopping the moment it crosses the line. It feels efficient. It isn't. As Evan Miller put it in the reference every experimentation team eventually reads, "Repeated significance testing always increases the rate of false positives, that is, you'll think many insignificant results are significant." A tool built to call tests fast is a tool built to peek constantly — so the faster it ships you winners, the larger the share of them that were never there.

Put the two together and the machine's headline feature works against its headline number. The thing that makes it valuable — volume, speed, always-on — is the same thing inflating the count of wins you can't bank.

The one check: do the wins reconcile?

Here's the test that cuts through all of it, and it takes one afternoon a quarter. Add up the lifts the tool reported over a window — every winner it shipped, every percent it claimed. Then pull what a holdout actually earned over the same window: a slice of traffic the tool never touched, left on the original the whole time. Compare the two totals.

If the tool's optimization is real, the holdout group should be visibly worse off than everyone else — that gap is the true, banked lift. If the summed wins are mostly phantom, the two totals collapse together: a long list of reported percentages on one side, a conversion rate that barely moved on the other. Eppo's experimentation team named this exactly. Summing your experiment wins, they write, "often paints an overly optimistic picture of true impact" — and once you measure against a proper holdout, "the gains revealed by the holdouts are far less dramatic than the naive summation of individual 'wins'."

That's the whole check — not "was this one test valid," but do the wins, added up, reconcile with the number finance can see? One snag catches exactly this audience: plenty of always-on tools don't hold any traffic out by default. They optimize across everyone and leave you no untouched group to measure against. If yours is one of them, standing up a permanent holdout is step zero — and it's the most useful thing you can ask the vendor to switch on.

The market sells you the win, not the count

None of this is hidden. It's just not on the page you're shown. Of the seven live AI-CRO vendor pages we read while writing this, six lead with an uplift — "boost conversions," a case-study percentage, a testimonial number — and not one of them reconciles that win against a holdout or a site-wide total anywhere on the page. Unbounce's page states the pattern about as plainly as it gets: turn on its optimizer, it says, and marketers get "30% more sales and signups (on average)."

Read that as the honest sentence it is. It's an average lift, headlined as the payoff — with no test count behind it and no aggregate next to it. That isn't a lie, and it isn't Unbounce's job to talk a buyer out of the number; every serious tool in the category markets the same way, because the win is what sells and the reconciliation is the buyer's homework. But it means the surface you evaluate a tool on is engineered to show you the wins and stay quiet on how many tries produced them — which is precisely the information you'd need to know whether they're real.

Two moves the reconciliation doesn't cover

The quarterly reconciliation tells you whether your program as a whole is real. Two smaller moves protect you between reconciliations — and neither one is the reconciliation.

Before you bank a single winner, re-run it frozen. When one specific result is about to go into a forecast or roll out everywhere, put it back at a fixed even split for a defined window with the tool's optimizing switched off, and see if it holds. This is the frozen-split discipline the Vol 6 piece laid out for reading a contaminated single test — here it does a second job: it's the cheapest way to catch a lucky winner before it becomes a number someone defends in a meeting.

Ask the vendor one question: how many tests did this winner come out of? A win pulled from a batch of two is worth taking seriously. The same win pulled from a batch of forty is mostly the winner's curse wearing a green badge. The good tools can answer in a sentence and will show you the denominator. If the honest answer is "we don't surface that," you've learned something more useful than the lift: you've learned the tool is built to declare winners, not to help you audit them.

Optimizing more isn't finding more truth

The tool that runs a hundred tests a quarter is doing real work, and at enough volume it's the right machine for the job — that's the case the earlier piece on auto-tuning made about high traffic, and nothing here walks it back. The boundary is just this: a machine tuned to produce wins fast is not the same machine as one tuned to tell you which wins are true, and the first one, run hard, will hand you more false positives, not fewer. Speed is a feature of the optimizer. It's a liability for the scoreboard.

So next quarter, before the win-pile goes anywhere near a forecast, run the reconciliation — and on any winner big enough to bet on, freeze it and ask how many tests it came from. If the wins reconcile, bank them. If they don't, you don't have a conversion problem the tool solved. You have a counting problem the tool created. Run it on your own numbers and tell us what the holdout said.

More articles

How to get the most out of Meta Advantage+ and Google Performance Max: build the outer loop

Meta Advantage+ and Google Performance Max run the buying now — and the same machine ships to every advertiser. How to feed it what the platforms themselves ask for, and build the outer loop: the record of why winners won that compounds for you alone.

Read more

AI can fix most of your CRO problems. The two it can’t are probably your real bottleneck.

Every CRO program stalls on the same handful of problems. AI solves some, only speeds up others, and can’t touch two — the traffic floor and where your learnings live. Sort yours before you buy a tool.

Read more

Tell us about your project

Our offices

  • Cascais
    Rua do Cabo 6
    2755-669 Cascais, Portugal
  • Rio de Janeiro
    Honório de Barros 12
    22250-120, Rio de Janeiro, Brazil