THE DECISION LABWhat is the next decision worth getting right?Explore both briefs

AI & Measurement · Published

The A/B Test Illusion

THE ARGUMENT IN BRIEF

An A/B label cannot rescue an experiment with weak randomisation, contamination or an unsuitable comparison. Anil asks marketers to examine platform tests before treating their reported lift as causal evidence.

Rethinking Experimentation in Digital Advertising

A/B testing has long been the marketing industry’s gold standard for measuring advertising effectiveness. But recent research reveals that many A/B tests conducted via platforms like Google and Meta don’t meet the core principles of experimental design.

Instead of offering clean, randomized comparisons, these platforms frequently introduce algorithmic interference, selective audience delivery, and predictive targeting — all of which compromise test validity.

This paper outlines:

  • Why many industry-standard A/B tests are flawed.

  • How platforms present prediction as causation.

  • The real-world consequences of trusting skewed tests.

  • What marketers must do to rebuild measurement integrity.

📍 Introduction: The Sacred Cow of Marketing

Marketers have long treated A/B testing as sacred — reliable, scientific, and neutral. But as machine learning enters the test environment, even experienced professionals are being misled.

When the same platform controlling ad delivery also controls the experiment design, biases creep in. And when algorithms pick who sees an ad — based on likelihood to engage — the test is no longer a fair assessment of ad impact.

This is not just semantics. It’s a systemic problem.

🚩 The Myth of True Randomization

Two landmark papers — Braun & Schwartz (2024) and Boegershausen, Cornil, Yi & Hardisty (2024) — highlight key failures in so-called A/B tests on Google and Meta:

  • ❌ No real random assignment. Algorithms optimize delivery based on predicted responsiveness.

  • ❌ Skewed groups. The treatment group is often more likely to convert before the ad is even shown.

  • ❌ Confounding variables. Audience composition differs in ways unrelated to the ad.

This breaks the very foundation of causal testing.

“The results are confounded. It’s not an experiment. It’s a biased observation.” — Boegershausen et al.

🌻 Case Example: The Gardener Test (Braun & Schwartz)

A landscape gardening business runs a campaign on Meta. The ad is mostly shown to users who already like gardens, crafts, or DIY.

The control group, meanwhile, is more randomly composed.

Results: The ad appears to “work.” Reality: The targeting algorithm just picked users already primed to convert.

This isn't an ad impact test. It's a selection bias showcase.

📉 What Platforms Promote (and Why)

Major platforms regularly release case studies citing 60–80% ROI or conversion uplift. But here’s what’s happening behind the scenes:

  • Meta optimizes delivery mid-flight using user prediction scores.

  • Google suggests ROI metrics as evidence of ad performance — without a real counterfactual.

  • Twitter compares users who saw ads to those who didn’t — but these are apples and oranges demographically and behaviorally.

“These platforms blur the line between prediction and causation — and marketers pay the price.”

🧪 Prediction ≠ Causation

This is the industry’s biggest measurement confusion.

  • ✅ Prediction answers: “Who is most likely to act?”

  • ✅ Causation answers: “Did the ad make them act?”

Most platform A/B tests focus on the former while claiming to prove the latter. That’s a dangerous conflation.

💥 Why It Matters

When tests are flawed, consequences follow:

  • 📉 Marketing budgets are misallocated.

  • 🎯 Creative strategies are built on false signals.

  • 💸 Brands spend more on platforms than justified.

  • ⚠️ Incrementality becomes a mirage.

And we continue believing in ad success that never really existed.

🔍 The Industry Needs Better Questions

Before trusting any test, ask:

  • Was the audience truly randomized?

  • Were both groups treated identically, except for the ad?

  • Was algorithmic optimization turned off during the test?

  • Is this result causal or just correlational?

If you can’t answer confidently — the test isn’t valid.

✅ A Better Way Forward

1. Use Ghost Ads & Real Holdouts Randomly assign users, but only show ads to one group — no algorithmic cherry-picking.

2. Educate Teams in Causal Thinking Train marketers, analysts, and CMOs to understand what constitutes real experimental rigor.

3. Demand Transparency from Platforms Platforms must reveal how their A/B tools work — and whether machine learning interferes.

4. Shift to Causal Measurement Tools Use incrementality platforms, causal lift experiments, and counterfactual modeling.

5. Push for Industry Standards Let’s stop treating biased observational setups as “experiments.” Measurement integrity must be non-negotiable.

The A/B test is not broken. But our faith in platform-driven tests must be re-examined.

It’s time we moved beyond pretty dashboards and back to real science. Because without true measurement, there is no optimization. Only illusion.

Let’s end the A/B illusion. And bring clarity back to marketing.


References

  1. Boegershausen, J., Cornil, Y., Yi, S., & Hardisty, D. On the Persistent Mischaracterization of Google and Facebook A/B Tests 👉 Full text

  2. Braun & Schwartz (2024) Selection Bias in Online Ad Effectiveness 👉 Summary + Diagram

  3. Langhe, B. & Puntoni, S. (2021) Does Personalized Advertising Work as Claimed? (Harvard Business Review) 👉 HBR Article


Anil Pandit

EVP- Publicis Media

India


*Disclaimer: This post is for informational purposes only and does not endorse or disapprove of any specific tools, platforms, or technologies. The views and opinions expressed in this article are those of the author and do not reflect the official policy or position of the company where he is employed.

FROM THE AUTHOR’S ARCHIVE

Original text from Anil Pandit’s article export. Claims and references reflect the time of writing.

View on LinkedIn ↗

What does this mean for your business?

Bring the question into a focused advisory conversation.

Work with Anil ↗
WhatsApp Anil