ROAS answers a credit question
Return on ad spend is usually calculated as revenue attributed to advertising divided by advertising cost. The equation looks causal because “return” suggests that the spend produced the revenue. In practice, the numerator depends on an attribution system: windows, identity, view-through rules, click rules, channel coverage, consent, modelling, and platform-specific definitions.
That makes ROAS useful for some operational comparisons, but it does not reveal the missing counterfactual: what would these customers have done if the advertising had not run?
Incrementality attempts to estimate the difference caused by the intervention. It asks a different question, under assumptions that also need scrutiny. A lift test can be underpowered, contaminated, poorly randomised, or unrepresentative. Experimental language does not guarantee a credible estimate.
The responsible sequence is to name the decision first, then use the lightest evidence capable of supporting it.
Attribution and causality are not rivals
Attribution helps teams organise observed journeys, operate campaigns, reconcile systems, and form hypotheses. Incrementality helps estimate whether an intervention changed an outcome relative to a comparison.
A campaign can have strong attributed ROAS and low incremental effect if many credited customers would have bought anyway. It can have modest attributed ROAS and useful incremental effect if tracking misses cross-device or delayed behaviour. Both measures can be wrong when revenue, cost, identity, or assignment data is poor.
Do not “correct” platform ROAS with an arbitrary incrementality factor borrowed from another company. The causal effect depends on brand demand, audience, channel, creative, offer, period, competitor activity, and test design.
State the claim:
- Tracking claim: events and cost are recorded and reconciled within known coverage.
- Attribution claim: observed outcomes receive credit under a documented rule.
- Triangulation claim: several non-experimental signals point in a consistent direction.
- Experimental claim: a defined comparison estimated lift under stated assumptions.
- Generalisation claim: repeated evidence suggests a planning range beyond one test.
Each level can improve a decision when labelled.
The Advertising Evidence Ladder
The ladder has five rungs: tracking, attribution, triangulation, holdout, and repeated learning. Climb only as far as the decision needs and the organisation can support.
Tracking
Reconcile spend, delivery, site events, orders, cancellations, refunds, contribution, and customer status. Document timezone, currency, tax, attribution windows, view-through treatment, consent loss, cross-device limits, and how duplicate events are removed.
Inspect data by campaign and outcome rather than trusting a dashboard total. Confirm that a purchase event represents a valid order and that later refunds can be connected. Include agency, creative, discount, fulfilment, and platform costs appropriate to the decision, not only media.
Tracking answers whether the observable system is coherent enough for analysis. It does not determine causality.
Attribution
Choose models for an operational reason. Last click may help manage lower-funnel capture but ignore earlier exposure. Platform attribution can help the platform optimise within its own signals but overlap with other systems. Multi-touch models distribute credit according to assumptions that are often hard to validate.
Compare models to identify dependence on the rule. If a campaign looks strong only under a long view-through window, that is a question for investigation. Do not average model outputs as though several assumptions create one truth.
Use attributed ROAS with guardrails: contribution, new-customer quality, repeat, returns, brand demand, support load, and capacity. A campaign can improve attributed revenue while harming a more important outcome.
Triangulation
Combine imperfect signals with different failure modes: branded and direct demand, geographic patterns, search interest, customer self-report, promotion codes, CRM source, time series, and channel interruption. Look for convergence and contradiction.
Triangulation is not a hidden algorithm. Write the observations and alternative explanations. A sales increase after launch may reflect seasonality, distribution expansion, price, publicity, competitor stock, or organic demand as well as advertising.
Use triangulation to decide whether a controlled comparison is worth the cost and where it should focus.
Holdout
A holdout creates a comparison group that does not receive the intervention, while the exposed group does. Random assignment makes the groups comparable on average when implemented correctly.
Google documents Conversion Lift as a product for eligible advertisers, with particular setup and measurement constraints. Its documentation is useful for understanding the platform method, not a guarantee of eligibility, statistical power, independence, or unbiased implementation.
User-level tests can be difficult when identity and consent are incomplete, exposure crosses channels, or platform rules restrict design. Geo tests assign regions rather than individuals. Google’s geography-based lift documentation describes one provider implementation. Geography introduces assumptions about comparability, spillover, local media, inventory, and external shocks.
Academic and industry methods offer deeper guidance. Google Research’s paper on randomised paired-geo experiments develops a trimmed-match estimator intended to improve robustness under its design. It does not make every small geo test valid; enough suitable geographies, variation, data, and statistical skill remain necessary.
Repeated learning
One experiment estimates one treatment in one context. Campaign effects can change with saturation, creative, audience, season, competitors, price, and brand maturity.
Store the full result: hypothesis, assignment, pre-period, intervention, outcome, cost, exclusions, estimate, interval, diagnostics, deviations, and decision. Keep null and negative results. Repeat material questions and update planning ranges.
Avoid using a lift result forever or applying it to channels and audiences not tested. Generalisation is another inference that needs evidence.
Define the decision before the design
Different decisions need different evidence.
“Is tracking reliable enough to operate tomorrow?” may need reconciliation and test transactions. “Which of two creatives should the platform serve?” may use a platform experiment on an appropriate proximal outcome. “Should we increase annual channel investment by 40%?” may justify a stronger incrementality design and downstream economics. “Did advertising build future demand?” may require repeated, longer-horizon evidence beyond immediate conversion lift.
Write:
- the action that will change;
- the smallest effect that would change it;
- the primary outcome and economic definition;
- the decision window;
- material risks and guardrails;
- feasible unit of assignment;
- what conclusion remains unavailable.
If the test result cannot alter spend, targeting, creative, or strategy, do not run it for theatre.
Test feasibility honestly
Power is not just a software output. The test needs enough units, outcome volume, expected signal, and stable assignment. Geo designs with a few highly different regions can produce fragile estimates. User tests can lose power through noncompliance or contamination. Long tests can collide with seasonality and business changes.
Run an analysis on historical data before launch. Inspect outcome variance, geographic fit, pre-period stability, sample availability, minimum detectable effect, and how business-as-usual campaigns interact. Predefine exclusions and avoid changing them after seeing the answer.
Protect the customer. Do not withhold essential service, transactional communication, accessibility, or safety. A holdout from promotional advertising is different from withholding an entitled support message.
The joint IAB and IAB Europe incrementality guidelines provide useful industry vocabulary and reporting considerations, especially in commerce media. They are industry guidance, not proof that a particular experiment satisfies causal assumptions.
Interpret uncertainty as part of the result
A lift estimate without an uncertainty interval is incomplete. An interval that includes zero does not prove “no effect”; the study may be unable to distinguish a meaningful effect from noise. A statistically distinguishable effect may be too small to justify cost.
Report absolute incremental outcome, relative lift, incremental contribution, incremental cost, and a return measure with intervals where the method permits. Show assumptions and diagnostics. Explain whether the result applies to the tested spend level; marginal effect can change as budget scales.
Avoid selecting the most favourable outcome from many metrics. Predefine the primary outcome. Treat secondary and subgroup findings as exploratory unless the design supports them.
Include implementation deviations: campaign leakage into control geographies, tracking outage, price change, stock constraint, competitor event, or mid-test creative change. A clean slide should not hide a messy intervention.
What a smaller team can do
Small organisations often lack the volume for sophisticated platform or geo lift tests. The answer is not to fabricate precision. Improve the ladder.
First, reconcile orders, refunds, margin, customer status, and spend. Second, shorten unreasonable attribution windows and compare rules. Third, track promotion or landing mechanisms without assuming exclusivity. Fourth, record major campaign changes and external events. Fifth, use bounded pauses or staggered launches only where operationally sensible, interpreting them as weaker quasi-experiments. Sixth, pool learning across repeated comparable periods cautiously.
Focus on large decisions. It may be impossible to estimate the incremental effect of every weekly campaign, but feasible to test whether a major prospecting programme adds new-customer contribution during a stable period.
Use qualitative evidence to improve mechanism, not to calculate lift. Customer interviews can reveal how advertising entered consideration and which proof mattered. They cannot provide a precise population effect from a small sample.
Build action rules before seeing the result
For a material test, agree:
- Scale if the conservative estimate clears the contribution threshold and guardrails hold.
- Maintain and retest if the estimate is promising but uncertainty remains decision-relevant.
- Revise the mechanism if exposure occurs but intermediate behaviour does not change.
- Stop or redirect if the upper plausible effect cannot justify cost.
- Treat a compromised test as inconclusive rather than rescuing it with favourable slices.
These rules reduce pressure to reinterpret the result around a committed budget. They also make the evidence operational.
Measurement is organisational, not only statistical
Marketing, finance, commerce, analytics, privacy, sales, and operations may use different revenue, customer, and cost definitions. Resolve those differences before arguing about estimators.
Assign ownership for event quality, experiment design, campaign implementation, economic calculation, and decision. Keep a measurement register with tested question, design, status, result, limitation, and next review.
The IAB’s State of Data 2026 offers directional industry context on measurement and data practice. As an industry report, it should inform questions rather than serve as causal evidence for a campaign.
Ask what the advertising changed
ROAS remains useful when its definition and job are clear. It can monitor attributed commercial efficiency, compare operational views, and flag changes. It becomes dangerous when the ratio is presented as proof of causality.
Incrementality is not a superior dashboard number. It is an inference from a designed comparison, with uncertainty and limits. Use it when the decision warrants the cost and the design can support the conclusion.
Name the decision. Climb the evidence ladder. Preserve the counterfactual. Report uncertainty. Repeat the learning. The objective is not to make advertising look accountable; it is to understand what the advertising actually changed.
Sources and further reading
- About Conversion Lift
Google Ads Help · Technical documentation
- About Conversion Lift based on geography
Google Ads Help · Technical documentation
- Robust causal inference for incremental return on ad spend with randomized paired geo experiments
Google Research · Peer-reviewed research
- Guidelines for Incremental Measurement in Commerce Media
IAB and IAB Europe · Industry evidence
- State of Data 2026
IAB · Industry evidence
