AI, usefully · 4 min read ·

The numbers were right. The recommendation wasn’t.

You upload a spreadsheet to an AI assistant and ask which version of a signup page performed better.

It calculates conversion rates, produces a chart and recommends Version B.

You check the arithmetic. Everything matches.

Should you trust the recommendation?

There is one more thing to check: whether the comparison was fair in the first place.

A result that looks convincing

Imagine these fictional results from two signup pages:

Fictional signup-page results
Visitor typeVersion AVersion B
Returning visitors81 signups / 90 visitors — 90%234 / 270 — 86.7%
New visitors192 signups / 240 visitors — 80%55 / 70 — 78.6%
Overall273 / 330 — 82.7%289 / 340 — 85%

Version B has the higher overall conversion rate. An assistant might recommend rolling it out.

But look at the two visitor groups separately.

Version A performed better among returning visitors. It also performed better among new visitors.

How can B win overall while losing in both groups?

B received far more returning visitors—the group that was more likely to sign up under either version. Its overall result benefited from the audience it received.

The calculation is correct. The conclusion that B is the better page does not follow from it.

The total hides who was counted

An overall conversion rate combines two things: how each group behaved and how much of each group appeared in the data.

Change the audience mix, and the total can change even when the experience gets worse within every group.

This reversal is an example of Simpson’s paradox: a pattern in combined data can differ from the patterns within its groups.

It is easy to imagine this happening outside a spreadsheet exercise. One page receives more traffic from existing customers. Another receives more first-time visitors from a new campaign. One week includes a sale; the next does not.

A comparison that looks like a product result may partly reflect a distribution decision.

Before recommending a change, I would want to know how visitors reached each version. Were they randomly assigned? Did traffic sources differ? Were both versions available during the same period?

Those questions determine what the numbers can support.

Ask AI to investigate its recommendation

“Check your answer” is a reasonable request, but it leaves the assistant to decide what checking means. It may simply recalculate the same totals.

Give it a more specific job:

Review this comparison before recommending a winner.

Calculate the overall conversion rate and the rate within each visitor group. Show the numerator and denominator.

Compare the audience mix between versions. Explain whether that mix could account for the overall difference.

State what we know about assignment to each version. If that information is missing, do not describe the difference as a causal effect.

Separate the observed result from the recommendation.

Now the review has something concrete to examine.

For our example, a useful answer would acknowledge B’s higher overall rate while pointing out A’s higher rate within both groups. It would also flag the unequal audience mix and missing information about assignment.

That is a more defensible starting point than declaring a winner.

Hold the audience mix constant

We can make the comparison easier to understand by asking a hypothetical question:

What would each version’s conversion rate be if both received an audience split equally between returning and new visitors?

For A, the average of the two group rates is 85%.

For B, it is approximately 82.6%.

This is called standardization: applying the same group weights to both versions so the mix does not drive the comparison.

It reveals why the overall ranking changed. It does not prove A caused better conversion. The equal split is an illustrative choice, and other differences between the groups may remain.

For a business decision, the relevant weights might reflect the audience expected next month. Those weights should be chosen deliberately and explained.

More segmentation is not always better

Once you notice this problem, it is tempting to split everything by device, country, channel, tenure and dozens of other attributes.

That can create tiny groups and a long list of apparent winners. Some differences will arise through ordinary variation.

Start with groups that have a plausible relationship to the outcome and existed before exposure to the change. In this example, returning status matters because returning visitors may already know the product.

Also distinguish a descriptive check from a reliable experiment. These fictional counts illustrate a reversal; they are not enough to settle a rollout decision without understanding assignment, uncertainty and other outcomes that matter.

AI can help identify those missing pieces. It cannot recover experimental conditions that were never recorded.

Try this

Paste the table above into an assistant and ask only:

Which version should we launch?

Then use the review prompt.

Compare the two responses. Did the assistant notice that A won within both groups? Did it ask how visitors were assigned? Did its recommendation change when the audience mix became visible?

Finally, verify the rates yourself with a calculator or spreadsheet.

Checking the arithmetic is necessary. Checking what the arithmetic means is where the decision becomes trustworthy.