• Research Paper
  • Current
  • Advanced
  • 15 min read

What Your ROAS Report Actually Measures

Every ad platform grades its own homework, on its own clock, using conversions it partly estimated. The field experiments that checked the answer found the reported number can be several times the real one.

Marketing · Advertising

Key takeaways

  • Observational attribution does not recover the experimental answer. Across 663 large-scale Facebook experiments, the median randomized lower-funnel lift was 5%; double/debiased machine learning put the same effect at 24% and stratified propensity matching at 64%.
  • Every channel grades its own homework on its own clock, so claimed conversions routinely sum past your actual order count. The excess is arithmetic, not fraud, and it stays invisible until somebody totals it.
  • A meaningful share of what a dashboard reports was never observed. Google states that modeled conversions “use data that doesn't identify individual users to estimate conversions that Google is unable to observe directly” (consent, cookie and cross-device gaps).
  • The rule-based multi-touch models are gone: Google Ads says first click, linear, time decay and position-based attribution “are no longer supported,” leaving last click and data-driven.

The number on the dashboard is not a return

Return on ad spend, as reported by an advertising platform, is a specific and much narrower quantity than the phrase suggests. It is the revenue that platform claims credit for, under its own attribution rule, within its own lookback window, divided by what you paid it. Three separate substitutions are buried in that sentence, and each one moves the number in the same direction: up.

The first substitution is correlation for causation. The platform can see that a person saw or clicked an ad and later bought. It cannot see what that person would have done if the ad had not been shown, because it never withheld the ad from a comparable person and looked. Everything a standard report contains is the first half of a comparison whose second half was never run.

The second is estimated for observed. Google's own documentation is direct about this: modeled conversions “use data that doesn't identify individual users to estimate conversions that Google is unable to observe directly,” filling three specific gaps: browsers that block cookies, users who did not consent to advertising cookies, and journeys that start on one device and finish on another. Modeling those gaps is a defensible engineering response to a real problem. It also means a portion of the conversions in your report is a prediction produced by a model whose parameters you cannot inspect, trained on patterns from users who did consent.

The third is claimed for unique. No platform reports on the condition that no other platform also reports. Each applies its own window to the same buyer, and windows overlap.

Illustrative only: invented round numbers, chosen to make the arithmetic visible. A store does 5,000 online orders in a month at an $80 average order value, so $400,000 of revenue. Meta's dashboard claims 1,875 conversions. Google Ads claims 1,510. The affiliate network claims 620. The email platform claims 1,240. That totals 5,245 claimed conversions against 5,000 orders that actually happened, or 105% of the month's business, allocated entirely to paid and owned channels, leaving less than nothing for organic search, direct traffic, the retail partner, and everyone who bought because a friend recommended it. Nobody lied. Four systems each answered a question about their own contribution and none of them was asked to reconcile with the others.

The practical lever here costs nothing: add the column up. Sum every channel's claimed conversions for a full month and divide by orders actually shipped. That ratio is your double-count factor, and it is the first honest number most teams have ever put on their reporting. The risk of not doing it is specific and expensive. Budget gets allocated in proportion to claimed credit, so the channels best at claiming credit get the most money, which is not the same property as producing the most sales.

Attribution allocates credit; it does not measure cause

An attribution model is a rule for dividing one conversion among several touchpoints. Last click gives everything to the final interaction. First click gives everything to the first. Linear splits it evenly, time decay weights the recent, position-based front-loads and back-loads. Data-driven models fit weights from historical patterns in the account.

Google has quietly narrowed the menu. Its attribution documentation now states that “the first click, linear, time decay, and position-based attribution models are no longer supported by Google,” with conversion actions that used them upgraded to data-driven attribution and last click remaining as the alternative. That is a meaningful admission from the largest seller of advertising in the world: the rule-based multi-touch models, which an industry spent a decade arguing about, were not worth maintaining.

But the retirement of five models does not fix the flaw the remaining two share. Every attribution model, including the data-driven one, starts by assuming the conversion belongs to the ads and then argues about proportions. All of them allocate exactly 100% of the sale across the touchpoints they can see. None of them entertains the possibility that the correct allocation was 0%: the customer was already going to buy, and the ad was a toll booth on a road they were driving anyway.

The distinction is worth stating in one line, because it is the whole discipline. Attribution answers who was nearby when the sale happened. Incrementality answers what would not have happened otherwise. They are different questions, they have different answers, and only the second one has a dollar value attached, because only the second one tells you what happens if you stop spending.

This matters most exactly where reported performance looks best. The channels that report the highest ROAS are structurally the ones positioned closest to an already-decided purchase: retargeting people who put the item in a cart, branded search sitting on your own company name, the affiliate coupon site the customer visited after choosing to buy. High reported ROAS is a signal of good positioning within the purchase journey. It is not, on its own, evidence of causation, and the correlation between the two is weaker than any dashboard implies.

What the experiments found when someone checked

This is not a theoretical concern, and the reason we know is that the largest advertising platform in the world ran the comparison at scale and published it twice.

The first study is Gordon, Zettelmeyer, Bhargava and Chapsky, “A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook,” Marketing Science 38(2), 193–225 (2019). It used 15 US advertising experiments covering 500 million user-experiment observations and 1.6 billion ad impressions, and compared the randomized results with what conventional observational methods would have concluded from the same data. The finding, in the authors' own framing, is that the observational methods often fail to produce the same effects as the randomized experiments even after conditioning on extensive demographic and behavioural variables.

The second study is larger and more uncomfortable. Gordon, Moakler and Zettelmeyer, “Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement,” Marketing Science 42(4), 768–793 (2023), analysed 663 large-scale Facebook experiments with access to over 5,000 user-level features. That is richer data, as the authors note, than most advertisers or their measurement partners can access. They tested two serious modern methods: double/debiased machine learning and stratified propensity score matching.

The numbers are the point. Median randomized lift was 29% for upper-funnel outcomes, 18% for middle-funnel and 5% for lower-funnel. The same effects estimated by double/debiased machine learning came out at 83%, 58% and 24%. By stratified propensity score matching, 173%, 176% and 64%. At the lower funnel (purchases, the outcome most advertisers actually care about), the machine-learning estimate was roughly five times the experimental truth and the matching estimate roughly thirteen times it. The authors' conclusion is that despite having access to large-scale experiments and rich user-level data, they were unable to reliably estimate an ad campaign's causal effect.

Sit with the implication rather than the numbers. These researchers had the platform's own logs, thousands of user-level features, the experimental ground truth to check themselves against, and no commercial incentive to flatter any channel. They still could not get there without the experiment. Whatever your attribution vendor is doing, it is working with less than that.

The honest reading is not that advertising does not work: the randomized lifts were real and positive at every funnel stage. It is that the direction of the error is consistent. Observational methods overstated, at every level, in both studies. That is what selection does: platforms optimise delivery toward people most likely to convert, which is the optimiser working exactly as designed, and it is precisely the behaviour that makes the exposed group incomparable to the unexposed one.

From a reported 3.0 to a real 0.96

The gap between reported and incremental is not a rounding error you can adjust for with a haircut. It routinely flips the sign on a channel's contribution. Here is the arithmetic end to end.

Illustrative only: the same invented store as before, 5,000 orders a month at an $80 average order value, $400,000 of revenue. It spends $50,000 a month on one paid channel. The platform reports 1,875 conversions, which at $80 each is $150,000 of attributed revenue, so a reported ROAS of 3.0. Contribution margin after product cost, payment processing, packing, shipping and returns is 45%, which makes break-even ROAS 1 ÷ 0.45 = 2.22. Reported 3.0 comfortably clears 2.22, so the channel looks profitable: $150,000 × 0.45 = $67,500 of contribution against $50,000 of spend, an apparent gain of $17,500 a month.

Now run a geo holdout. Split the country into 20 comparable markets, turn the channel off in 10 of them for four weeks and leave it on in the other 10. In the four weeks before the test each half did $200,000, together making up the store's $400,000 monthly total. During the test, the on-half did $222,000 and the off-half did $198,000. The on-half moved by +$22,000; the off-half moved by −$2,000, which is the seasonal and market drift you would have mistaken for advertising. Difference-in-differences gives $22,000 − (−$2,000) = $24,000 of incremental revenue, produced by the $25,000 of spend that ran in the on-half. Incremental ROAS is $24,000 ÷ $25,000 = 0.96.

Scale that back to the full country: about $48,000 of incremental revenue for $50,000 of spend. At 45% contribution margin that is $21,600 of contribution against $50,000 of cost, a real loss of $28,400 a month against an apparent gain of $17,500. The reported figure overstated the truth by roughly 3.1 times, and more importantly it sat on the profitable side of a break-even threshold that the real number never came close to.

The method matters more than the specific outcome, so be clear about what it does. Geo experiments, the design formalised in Vaver and Koehler's 2011 Google paper, randomise non-overlapping geographic regions into treatment and control and use location targeting to deliver the condition. They work when individual-level randomisation is unavailable, which after the privacy changes is most of the time. Their weaknesses are equally specific: you need enough comparable regions, you need a clean pre-period, and spillover across region borders biases the estimate toward zero.

The risk in this section is over-reading a single test. One four-week holdout in one channel in one season is a data point, not a policy. A channel that reads 0.96 in November may read differently in March; a brand campaign judged over four weeks will look worse than it is, because the effects it produces have not arrived yet. Run the test, act on it, and then run it again.

Why the signal will not come back

A reasonable person reading the above might conclude the answer is better tracking. It is worth being clear that this is not available, and that the direction of travel is away from it.

Apple's App Tracking Transparency requires an app to ask permission before tracking a user across other companies' apps and websites, and if the user declines, the advertising identifier is returned as all zeros. Not degraded, not noisy — zeros. Meta's fiscal 2025 Form 10-K names the consequence plainly, stating that its advertising revenue has been negatively impacted by marketer reaction to targeting and measurement challenges associated with iOS changes beginning in 2021. The same filing says regulatory developments and browser and operating-system changes have limited its ability to target and measure the effectiveness of ads on its platform.

On the web the picture is not resolution but stalemate. On 22 April 2025, Anthony Chavez, VP of Privacy Sandbox, announced that Google would “maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies.” Third-party cookies did not go away and the replacement did not arrive. What that means operationally is that the measurement substrate is now permanently heterogeneous: some browsers block, some allow, some users consent and some do not, some journeys cross devices and some do not. Every gap gets filled by a model rather than closed.

The context makes the stakes concrete. IAB and PwC's Internet Advertising Revenue Report puts US digital advertising at nearly $300 billion of revenue in 2025, up 13.9% year over year. That is a very large market whose primary performance metric is, at every level, partly modelled and structurally biased upward.

The strategic consequence is a reordering of what to invest in. Better pixels, longer lookback windows, and server-side conversion APIs improve the completeness of a measure that was answering the wrong question in the first place. Deliberate holdouts, and first-party data you own outright, answer the right one. The teams that will be measuring anything trustworthy in three years are the ones building experiment capability now, not the ones optimising their tag manager.

The risk of ignoring this is not that a report is slightly wrong. It is directional and compounding: budget flows to whichever channel reports best, reported performance is systematically inflated where causation is weakest, and the money therefore drifts steadily toward the parts of the funnel closest to purchases that were going to happen regardless, until growth stops and nobody can find the reason in a dashboard that still says 3.0.

Designing a test that can actually decide something

A holdout only settles an argument if it was capable of detecting the effect before you ran it. Most are not, and the failure mode is silent: an underpowered test returns no significant difference, the team reads that as proof the channel does nothing, and a working channel gets cut on the strength of a measurement that never had a chance.

Start with power, because it is arithmetic and it is unforgiving. Suppose your baseline conversion rate is 2% and the smallest lift worth acting on is a relative 10%, moving 2.0% to 2.2%. A standard two-proportion power calculation at 95% confidence and 80% power needs roughly 81,000 users per arm, about 161,000 in total, to detect that difference reliably. If your test can only reach 20,000 people, the honest thing to do is not run it. Change the design: test a bigger effect, test at the geo level where the unit of observation is a market's whole revenue rather than one person's binary purchase, or accept that this channel will be judged on judgement rather than evidence.

Then pick the design that fits the constraint. A geo holdout is the general-purpose tool: it needs no user-level identity, survives the privacy changes intact, and measures total business outcome rather than tracked conversions. A platform-run conversion lift study randomises at the user level inside one platform, which is cleaner statistically but leaves the platform holding the ruler. A media mix model reads history at the aggregate level and is genuinely useful for allocation across channels and for capturing long-lag effects, but it is a correlational model fitted to observational data, the exact category of method the Facebook studies found unable to recover experimental truth. Treat an MMM as a hypothesis generator that experiments then adjudicate, never as the adjudicator.

Build the cadence rather than the one-off. A sensible rhythm is one holdout per quarter on the largest single line of spend, rotating channels, with each result written down alongside what the platform reported for the same period. Two years of that produces something no vendor can sell you: a channel-by-channel record of your own overstatement factor, which lets you discount tomorrow's report with a number you measured instead of a feeling.

And hold two honest limits. First, you cannot experiment on everything. Testing has an opportunity cost measured in the revenue you deliberately gave up in the control group, and on small lines that cost exceeds the value of the answer. Measure the big lines properly and accept a wide range on the rest. Second, a null result is not a zero. It means the effect, if any, was smaller than your test could see. Say it that way in the readout, every time, because the alternative is a team that gradually concludes advertising does not work when what actually happened is that nobody built a test large enough to notice it working.

Put it to work

Stop reading platform-reported ROAS as a return. Add up every channel's claimed conversions this month and divide by your actual order count. The excess is your double-count. Then pick one channel and run a geo holdout: half your markets dark for four weeks, difference-in-differences against the pre-period. Judge the channel on that number against your break-even ROAS, and re-run it quarterly.

Sources & references

Linked entries open the named source directly. Entries without a link say exactly what kind of reference they are — and how to check them yourself.

Educational note: This briefing is general business education, not financial, legal, tax, or investment advice. Figures and rules change and vary by situation — verify current specifics with primary sources and qualified professionals before acting.