attribution vs incrementality vs mmm

Attribution vs Incrementality vs MMM: When Each One Lies to You

Every marketer eventually has the same conversation with their CFO. The dashboard says a channel is delivering 5x ROAS. The CFO asks, reasonably, why company revenue hasn’t grown 5x alongside the ad spend. There’s no good answer, because the dashboard was never measuring what either of you assumed it was measuring.

This isn’t a data quality problem you can fix with better tracking. It’s a structural feature of how attribution works. And the uncomfortable follow-up is that incrementality testing and marketing mix modeling, the two methods marketers reach for as the “more honest” alternative, have their own blind spots that get far less airtime.

This article isn’t going to tell you incrementality is right and attribution is wrong, or that MMM is the grown-up answer everyone should be running. It’s going to walk through exactly where each of these three methods lies to you, with real numbers, so you know which question you’re actually answering when you look at any one of them.

The Case That Should Change How You Think About Attribution

In 2012, eBay ran an experiment that should be required reading for anyone who has ever looked at a ROAS number and made a budget decision. Economists Thomas Blake, Chris Nosko, and Steven Tadelis, working inside eBay’s own research team, shut off paid search ads on branded keywords, searches that already included the word “eBay”, across a sample of U.S. markets, while leaving them running elsewhere as a control.

If attribution models were telling the truth, traffic should have collapsed. eBay’s own platform reporting had been crediting those branded search ads with a huge share of conversions. Instead, when the ads went dark, 99.5% of the clicks eBay had been paying for were simply retained through organic search results instead. Customers who typed “eBay” into Google were going to click on eBay’s own listing whether or not eBay paid for the privilege of being there.

In a peer-reviewed field experiment published in Econometrica (2015), eBay found that shutting off branded search advertising retained 99.5% of the traffic that paid attribution had been crediting to those ads, demonstrating that attributed performance and actual causal impact can diverge almost completely for high-intent, branded queries.

The paper’s broader finding was just as important: for non-branded keywords, ads genuinely worked for new and infrequent customers, but delivered negative average returns once you accounted for frequent buyers who would have converted anyway. The study has since been replicated with different magnitudes at other companies, one 2017 study on Edmunds.com found a smaller but still substantial substitution effect, confirming the pattern generalizes even if the exact size varies by brand and category.

This is the case worth anchoring on, not because it’s the most dramatic number available (plenty of vendor blogs cite flashier figures), but because it’s peer-reviewed, independently replicated, and comes from inside a company with every incentive to want its own advertising to look effective. If eBay’s branded search ROAS was fiction, yours probably has some fiction in it too.

What Attribution Actually Measures (and Why That’s Not “What Worked”)

Attribution answers a narrow, specific question: which tracked touchpoint was present when a conversion happened. It does not, and structurally cannot, answer whether that touchpoint caused the conversion.

This distinction sounds academic until you sit with what it means in practice. A platform like Meta or Google Ads reports conversions based on its own tracking, its own attribution window, and its own model, last-click, data-driven, or otherwise. Every one of those platforms has a direct incentive to credit itself generously. This isn’t a conspiracy, it’s just structural: the entity measuring the ad’s effectiveness is the same entity selling you the ad.

Where attribution genuinely earns its keep: fast, tactical, day-to-day decisions. Creative testing, campaign-level optimisation, catching a broken landing page, spotting which ad set is burning budget with zero engagement. For these questions, attribution’s speed matters more than its causal precision. You don’t need a randomised controlled trial to know a creative with a 0.3% CTR is underperforming one with 2.1%.

Where attribution consistently misleads: retargeting and branded search specifically, the two channel types that reach people already deep in a buying decision. Industry benchmark data compiled across multiple holdout studies consistently shows retargeting campaigns reporting platform ROAS in the 3-4x range collapsing to closer to 1-2x, sometimes below breakeven, once measured against a true holdout group, because the audience being retargeted was disproportionately likely to convert regardless of the ad.

The mistake most teams make isn’t using attribution. It’s treating attribution’s number as a ledger of truth rather than a fast, biased signal that needs occasional correction.

Where Incrementality Testing Lies to You

Incrementality testing is the corrective most marketers reach for once they’ve internalised the eBay lesson, and it deserves the reputation it’s earned. A properly randomised holdout test, where a portion of your audience is deliberately withheld from seeing your ads and you compare outcomes between the two groups, produces something attribution structurally cannot: a causal estimate of what your advertising actually caused, not what merely happened near it.

But “properly randomised” is doing enormous work in that sentence, and this is where incrementality testing quietly lies to teams that don’t respect its requirements.

Sample size is the first lie. Most incrementality tests need a meaningfully large holdout, often 200,000 or more users per group for the statistical power to detect a real effect, and weeks of runtime to capture delayed conversions. Run a holdout test on a small account with modest traffic, and the “result” you get back is frequently statistical noise wearing the costume of an answer. A test that comes back saying “no significant lift” from an underpowered sample doesn’t mean the channel isn’t working. It might just mean the test couldn’t see far enough to tell.

Contamination is the second lie. Geo-holdout tests assume the holdout market is genuinely isolated from the treatment market. In practice, people travel, click through VPNs, see the same brand on a different channel entirely, or get exposed via word of mouth from someone in the treatment group. Every one of these contaminates the “clean” comparison the test depends on, and it rarely shows up as an obvious anomaly. It just quietly biases the lift estimate toward zero.

Access and cost used to be the third lie, though this has genuinely improved. Historically, a properly powered incrementality test carried a real cost floor, vendor estimates have put it as high as $100,000 for a rigorous test in years past. That floor has come down substantially as platforms like Google and Meta have expanded native, lower-budget testing tools, but smaller advertisers with limited traffic still frequently lack the volume to power a trustworthy test at all, regardless of what it costs.

Incrementality testing produces a genuinely causal estimate, but only when the test is adequately powered and the holdout group is truly uncontaminated, requirements that smaller advertisers frequently cannot meet, which means an underpowered “incrementality result” can be just as misleading as the attribution number it was meant to correct.

None of this makes incrementality testing less valuable. It makes it a tool with real prerequisites, not a magic corrective you can bolt onto any account and trust blindly.

What MMM Gets Right That the Other Two Can’t See

Marketing mix modeling takes a completely different approach: instead of tracking individual users or running live experiments, it uses statistical regression on aggregate data, spend, sales, seasonality, pricing, promotions, across time, to estimate each channel’s marginal contribution to revenue.

This structure gives MMM two real advantages the other two methods don’t have. It sees every channel, including offline media, TV, radio, out-of-home, that user-level tracking never touches at all. And it’s inherently privacy-resilient: MMM doesn’t depend on cookies, device IDs, or individual identity, which makes it the one measurement method that isn’t degrading as browsers and platforms restrict tracking further. 46.9% of US marketers report planning to increase MMM investment in 2026, a meaningful shift, and open-source tooling like Google’s Meridian and Meta’s Robyn has lowered the technical barrier to entry considerably from where it stood even two years ago.

Here’s where MMM lies, and it’s a subtler lie than attribution’s. MMM is still, underneath the statistical sophistication, a regression on correlated spend data. It can tell you which channels moved alongside revenue. It cannot, on its own, prove those channels caused that revenue movement, particularly when multiple channels scale their budgets together, which is exactly what most marketing teams do. Two channels that both increased spend the same quarter revenue grew will show correlated contribution in an MMM model regardless of which one, if either, actually drove the growth.

Marketing mix modeling estimates each channel’s marginal contribution to revenue from aggregate historical data, but remains a correlational method under strong modeling assumptions, which is why the rigorous position is that MMM is the strongest strategic allocation instrument available while incrementality testing remains the only genuinely causal method among the three.

The fix, and this is the part most MMM implementations skip, is calibration. Feeding real experimental results, from holdout tests, into the MMM’s assumptions meaningfully improves its accuracy. An MMM that’s never been checked against a single real experiment is a model built entirely on faith in its own assumptions holding true.

The Real Complication: AI-Mediated Discovery

There’s a structural shift that started in 2026 that complicated all three methods simultaneously, and it’s worth naming directly: AI-mediated discovery.

When someone asks ChatGPT, Gemini, or an AI Overview to recommend a product and then clicks through directly to the brand’s site, that visit typically registers in analytics as direct traffic. The AI system compressed discovery, comparison, and recommendation into a single interaction, and the influence layer that led to the visit disappears from every attribution model built for a world of trackable click paths. Some platforms have started adding referral parameters when users click directly from within an AI chat, which helps, but it only captures a fraction of the actual influence, since much of the decision-making now happens inside the conversation itself, invisibly, before any link is ever clicked.

This affects attribution most obviously, since it’s built entirely on tracked touchpoints. But it complicates incrementality testing too: a holdout test measures the effect of withholding your own paid media, not the effect of an AI system independently recommending or failing to recommend your brand in response to someone else’s query. And it complicates MMM by introducing a new demand-generation channel, organic AI visibility, that most models weren’t built to isolate as a distinct variable at all.

AI-mediated discovery compresses the traditional multi-step customer journey into a single interaction inside a chat interface, and the resulting traffic typically registers as direct or organic in analytics, meaning brand influence exercised entirely inside an AI conversation is currently invisible to attribution, incrementality testing, and MMM alike, not just to one of the three.

The honest answer here is that measurement for AI-mediated discovery is still genuinely unsettled as a discipline. Teams are experimenting with tracking branded search volume as a proxy for AI-driven awareness, monitoring referral traffic specifically from AI platforms, and building qualitative tracking of brand mentions inside LLM responses. None of these is a mature, validated methodology yet. Treat any framework claiming to have “solved” AI-attribution measurement with real skepticism.

How to Actually Combine All Three

The practitioner-level answer isn’t to pick a favourite. It’s to use each method for the specific question it can honestly answer, and let the other two calibrate its blind spots.

Use attribution for tactical, daily decisions. Which creative is underperforming. Which ad set needs a fresh audience. Where a landing page is leaking conversions. Speed matters more than causal precision here, and attribution’s bias is tolerable because the decisions are small and reversible.

Use incrementality testing to validate the channels attribution flatters most. Retargeting and branded search are the highest-priority candidates for a holdout test precisely because they’re the channels most likely to be over-credited. If a channel is consuming a large share of budget and has never been incrementality tested, that’s a real risk sitting in your media plan, not a hypothetical one.

Use MMM for strategic, quarterly-or-longer budget allocation, especially across channels that attribution can’t see cleanly at all, offline, brand, upper-funnel. Feed your incrementality test results back into the model as calibration data rather than treating MMM as a standalone oracle.

Watch for the specific warning sign of disagreement between methods. If your MMM suggests a channel is a strong performer but you’ve never validated that with an experiment, that’s exactly the candidate for a holdout test. If your attribution dashboard shows retargeting crushing it but you haven’t checked incrementality, that’s the other priority candidate. Disagreement between methods isn’t a data problem to smooth over. It’s the signal telling you where to spend your next testing budget.

A Field Test You Can Run This Quarter

If you’ve never run a real incrementality test, start narrow rather than comprehensive. Pick your single highest-spend channel that has never been validated, retargeting is usually the strongest first candidate given how consistently it’s over-credited across published studies. Use your ad platform’s native conversion lift tool (both Meta and Google offer these, and the access threshold has dropped considerably from where it sat a few years ago) rather than building a custom geo-holdout from scratch for your first attempt. Run it for a minimum of two to four weeks to capture delayed conversions, and be honest with yourself about whether your traffic volume is actually large enough to power a real result, if the platform’s own tool flags the test as underpowered, believe it.

Whatever the result, feed it back into how you talk about that channel’s attributed ROAS internally. The gap between the two numbers is the actual size of the problem you’ve been budgeting around.

The Honest Version

None of these three methods gives you the truth on its own. Attribution is fast and structurally biased toward whichever platform is measuring itself. Incrementality testing is genuinely causal but only within the narrow bounds of what it was properly powered to detect. MMM sees the whole picture but only through a correlational lens that needs real experiments to stay honest.

The teams making good budget decisions today aren’t the ones who picked a favourite method. They’re the ones who know exactly which question each method can honestly answer, and who’ve built the habit of checking one against another before a number gets treated as ground truth in a budget meeting.

If measurement, incrementality design, and budget allocation are areas you want to get genuinely sharp on, not just conceptually but with real frameworks you can run on your own accounts, YUP’s Performance Marketing course covers exactly this, from attribution mechanics through to running your first properly powered holdout test.

FAQs

What’s the difference between attribution and incrementality?

Attribution assigns credit to tracked touchpoints present near a conversion, without proving those touchpoints caused it. Incrementality testing uses a randomised holdout group to measure what a campaign actually caused by comparing outcomes between an exposed group and an identical group that wasn’t shown the ads. Attribution answers “what happened,” incrementality answers “what would have happened anyway.”

Is my ROAS real?

Partially, and the gap depends heavily on the channel. Branded search and retargeting are the channels most consistently over-credited by attribution, since they disproportionately reach people already likely to convert. Prospecting and upper-funnel channels tend to be closer to their attributed numbers, sometimes even under-credited, since attribution often misses the delayed, assisted conversions those channels generate.

What is marketing mix modeling used for?

MMM estimates each marketing channel’s contribution to revenue using aggregate historical data on spend, sales, and external factors like seasonality and pricing. It’s used primarily for strategic, longer-horizon budget allocation decisions, and it’s the only one of the three methods that can meaningfully account for offline channels and remains resilient to the privacy changes eroding user-level tracking.

Why did eBay’s branded search ads show no effect when turned off?

Because the customers clicking on eBay’s branded search ads were, overwhelmingly, people who had already decided to visit eBay and would have found it through organic search results regardless. The peer-reviewed 2012 field experiment found 99.5% of that “paid” traffic was simply retained through free organic listings once the ads were removed, showing the attributed value of those ads had been almost entirely fictional.

How big does a holdout group need to be for a valid incrementality test?

It varies by expected effect size and baseline conversion rate, but many properly powered tests require 200,000 or more users per group, with a minimum runtime of two to four weeks to capture delayed conversions. Smaller advertisers frequently lack the traffic volume to power a trustworthy test at all, which is worth checking before you run one and trust the result.

Can I trust MMM without ever running an experiment?

Not fully. MMM is a correlational method under the hood, and without real experimental data to calibrate its assumptions, it’s vulnerable to confusing channels that scaled spend together for channels that actually drove the resulting revenue. Feeding holdout test results back into the model as calibration data materially improves its reliability.

How does AI search change marketing attribution?

When AI systems like ChatGPT or Google’s AI Overviews recommend a brand and a user clicks through directly, that visit typically registers as direct or organic traffic in analytics, with the influence exercised inside the AI conversation invisible to standard tracking. This affects attribution, incrementality testing, and MMM simultaneously, and measurement approaches for it are still an unsettled, actively developing area as of today.

Which channels should I incrementality test first?

Prioritise the channels that consume the largest share of budget and have never been validated, retargeting and branded search are consistently the highest-priority candidates across published research, since both disproportionately reach people already likely to convert regardless of the ad.

Do I need a data science team to run any of these three methods?

Not necessarily. Attribution comes built into every ad platform by default. Basic incrementality testing is now accessible through native conversion lift tools on Meta and Google without custom infrastructure. MMM has become more accessible through open-source tools like Google Meridian and Meta Robyn, though a meaningfully rigorous MMM implementation still benefits from someone with statistical modeling experience.