Geo Lift Testing: 7 Limitations to Understand Before Trusting the Result

Geo lift testing has become an increasingly popular approach to marketing incrementality testing. The appeal is obvious. Instead of relying on attribution to assign credit to channels, a geo lift test attempts to estimate what happened because marketing activity changed relative to what would likely have happened without that change.

That can provide valuable causal evidence. The problem begins when the result from a relatively small part of the market is converted into a national incremental ROAS and used to make a much larger budget decision.

A geo lift test can be statistically robust in the markets tested and still produce a misleading national ROI.

The central issue is heterogeneity. Markets differ in brand strength, customer mix, competition, distribution, price sensitivity, media saturation, underlying demand and many other factors that influence how they respond to additional advertising.

This creates a question that deserves much more attention in geo incrementality testing: transportability.

It is one thing to identify an effect in the geographies included in an experiment. It is another to establish that the same effect can be expected across the rest of the market.

For businesses using geo lift testing to allocate millions in marketing budget, that distinction matters.

What does a geo lift test actually measure?

A geo lift test changes marketing activity in selected geographic markets and compares their subsequent performance with an estimate of what would have happened without the intervention.

Different methodologies construct this counterfactual in different ways. Some approaches match test and control regions. Others stratify geographies before treatment. Synthetic control methods construct a weighted combination of untreated markets that attempts to reproduce the historical behaviour of the treated geography.

The objective is to answer a causal question: did changing marketing activity create incremental revenue, conversions or another business outcome in the markets tested?

If the experiment is well designed, geo lift testing can answer that question effectively.

The commercial question is usually larger, however. A CMO does not simply want to know whether another £100,000 generated incremental revenue in a handful of cities. They want to know what will happen if they deploy another £5 million nationally.

Moving from the first question to the second requires assumptions about how representative the experiment is of the wider market.

That is where many geo lift analyses become much less certain than the headline ROI suggests.

1. A test covering 20% of the market is not automatically a national read

Suppose a company tests paid social across 20% of its national market and estimates an incremental ROAS of 3.2.

It is tempting to use 3.2 as the expected return on additional national investment. Doing so assumes that the remaining 80% of the market will respond to additional spend in approximately the same way.

There may be little reason to believe that.

A geography with strong brand awareness, high category demand and relatively low competition may respond very differently from one where awareness is weak and competitors dominate. A region where the channel is already saturated may produce little incremental response to additional spend, while an underexposed market may still have significant available reach.

The same issue applies to distribution, customer demographics, price elasticity, offline sales and historical media investment.

These are not small details around the edge of the experiment. They can determine the marginal return generated by additional advertising.

The geo lift result may therefore be valid for the markets tested without being a reliable estimate of what will happen nationally.

That is the essence of the transportability problem.

2. Stratifying on revenue does not necessarily create representative groups

One of the most important questions in a geo experiment is how markets are selected and stratified.

Some approaches rely heavily on the KPI being measured, such as historical revenue, orders or conversions. This can produce treatment and control groups that look well balanced before the experiment.

The problem is that similar outcomes do not imply similar underlying markets.

Consider two regions that both generate £10 million of annual revenue. One might have high brand awareness, extensive retail distribution, limited competition and relatively low price sensitivity. The other could have weak brand awareness, predominantly online sales, aggressive competitors and high dependence on promotional activity.

Their historical revenue may look almost identical while the mechanisms producing that revenue are completely different.

If marketing spend increases by 50%, there is no reason to assume that those two markets will respond similarly.

This is why matching or stratifying on the KPI alone can create a false sense of representativeness.

A stronger design should consider the underlying factors likely to influence treatment response. Depending on the business, these might include brand strength, category demand, customer demographics, competitive intensity, distribution, price elasticity, previous media investment, existing reach and frequency, and the split between online and offline sales.

Matching the result is not the same as matching the mechanism that created it.

That distinction matters because the experiment is not simply trying to predict revenue. It is trying to understand what happens to revenue when marketing activity changes.

3. Synthetic control can match historical behaviour without matching response to spend

Synthetic control is a powerful method for constructing a counterfactual. By combining untreated geographies in different proportions, it can sometimes reproduce the historical behaviour of a treated market extremely closely.

A strong pre-period fit is useful evidence that the synthetic control can approximate what would have happened without treatment.

It does not prove that the geographies involved share the same relationship between marketing spend and revenue.

Two markets can have highly similar historical sales patterns for very different reasons. One may be driven heavily by existing brand demand, while another depends more strongly on paid media. One may have substantial untapped reach, while another is close to saturation.

Those differences may be relatively invisible while both markets continue operating normally. They become important as soon as spend changes.

For marketing measurement, the relevant question is therefore not only whether the synthetic control reproduces historical revenue. It is whether the factors determining the response to marketing are sufficiently comparable for the counterfactual to remain meaningful under treatment.

A beautiful pre-period fit cannot establish that by itself.

This does not make synthetic control a bad methodology. It means that the quality of a synthetic control should not be judged only by how closely two revenue curves align before the experiment.

4. The planned budget change may not be the treatment that was actually delivered

Another important limitation appears between experimental design and media delivery.

Suppose every treated geography receives a planned 50% increase in budget. On paper, the intervention looks consistent.

In practice, advertising platforms do not allocate budget to respect an experimental design. Their bidding systems allocate budget to meet an optimisation objective.

An additional 50% of budget may therefore create a substantial increase in reach in one geography, mostly higher frequency in another, a different audience composition elsewhere, or relatively little additional delivery in a saturated market.

Auction conditions, audience availability, expected conversion rates, competition and existing saturation all influence what the platform actually does with the additional money.

The planned treatment may be identical across geographies. The treatment actually delivered is not.

This distinction matters because the experiment may appear to be estimating the effect of one intervention when it is actually averaging several different interventions.

A proper review of a geo lift test should therefore examine more than the budget change. It should investigate what happened to reach, frequency, impressions, CPMs, audience composition and delivery within each treated market.

Budget is an input. It is not necessarily the treatment.

5. A single incremental ROAS can hide substantial local variation

Aggregating results into one incremental ROAS makes them easy to communicate, but it can also hide information that matters for investment decisions.

Imagine ten treated geographies generating an average incremental ROAS of 3.0.

One possibility is that all ten markets produced returns close to 3.0. Another is that two markets generated exceptional results while the remaining eight produced very little incremental revenue.

Both experiments can produce the same headline average. They do not support the same commercial conclusion.

Before using the average to inform a national budget decision, it is worth understanding the distribution underneath it.

Which markets created the lift? How variable were the effects? Did unusually strong regions carry the result? Were those regions less saturated or experiencing stronger underlying demand? Did the advertising platform deliver a materially different treatment in those markets?

The aggregate result may be statistically correct while still concealing the variables that determine whether the result will survive national deployment.

This is another reason why transportability should be considered explicitly rather than assumed after the experiment.

6. Market selection can materially influence the reported result

Market selection also creates a particular risk in channel-sponsored geo lift testing.

A platform-sponsored test is not automatically unreliable. The issue is that the geographies selected for an experiment can materially affect the incremental ROI subsequently observed.

Testing markets with unusually strong brand demand, favourable auction conditions, low competition or significant unused reach can produce a strong result. That result may be entirely genuine within those regions.

The problem arises when the resulting incremental ROAS is extrapolated to markets with very different conditions.

This is why a channel helping to select markets, implement the experiment and interpret the result should create additional scrutiny around test design and national extrapolation.

The objective is not to assume bad intent. It is to understand the incentives involved and independently evaluate whether the selected markets provide a reasonable representation of the investment decision the business ultimately wants to make.

Identifiability inside the experiment does not automatically solve that problem.

7. Repeating geo tests does not eliminate the transportability problem

Some measurement approaches recommend testing major channels several times per year.

In principle, repeated experimentation is valuable. In practice, it quickly creates operational constraints if every experiment requires a meaningful share of the market to produce an actionable result.

Consider a company testing paid search, paid social, video, CTV, display and other significant channels. If each experiment requires 20% or 25% of the market, and each channel is tested several times per year, clean experimental space becomes difficult to maintain.

Tests can overlap. Other campaigns contaminate the treatment. Seasonality changes. Competitors react. Brand conditions evolve. Marketing teams may also be reluctant to repeatedly suppress or increase spend in commercially important markets purely to preserve experimental design.

More tests can provide more information, but frequency does not automatically create representativeness.

Five tests on poorly representative geographies do not necessarily provide a better national answer than one carefully designed experiment.

The important question remains the same: what justifies transporting the measured effect from the markets tested to the markets where the budget will ultimately be deployed?

Geo lift testing should be part of a measurement system, not the entire system

None of these limitations mean that geo lift testing is ineffective.

The same principle applies to every marketing measurement method.

Attribution provides useful information about observable customer interactions, but observed journeys should not be confused with causality.

Marketing mix modelling can provide valuable insight into the contribution of larger media channels, but statistical signal becomes harder to isolate as channel spend becomes small relative to the overall variation in the business.

Geo incrementality testing provides stronger experimental evidence, but it introduces questions around market selection, heterogeneous treatment effects, media delivery and transportability.

Every method has a blind spot.

The problem comes when a useful measurement technique is presented as a complete measurement system.

The goal of marketing measurement should not be to defend a particular model. It should be to combine evidence in a way that improves investment decisions.

Depending on the question, that may involve geo lift testing, other forms of incrementality testing, marketing mix modelling, attribution, brand tracking, Share of Search, customer research, pricing analysis and broader demand signals.

The method should follow the decision that needs to be made.

Seven questions to ask before trusting a geo lift result

Before using a geo incrementality test to make a major budget decision, I would want clear answers to seven questions.

1. What proportion of the actual market was tested?
The smaller the tested share, the more important the assumptions required to extrapolate the result nationally.

2. How were the geographies selected and stratified?
Historical revenue may be useful, but it should not be the only consideration if the underlying drivers of that revenue differ substantially.

3. Do the tested markets represent the markets where the budget will ultimately be deployed?
Consider brand strength, customer mix, competition, distribution, demand, saturation and other factors likely to affect response.

4. What treatment was actually delivered?
A 50% change in budget does not necessarily translate into a comparable change in reach, frequency or audience exposure.

5. How heterogeneous were the individual market responses?
Understand the distribution behind the average incremental ROAS.

6. How sensitive is the result to market selection?
If another reasonable set of geographies would produce a substantially different answer, that uncertainty should be reflected in the investment decision.

7. What other evidence supports the conclusion?
A geo lift result should contribute to the marketing measurement framework, not replace it.

A statistically robust geo lift can still lead to the wrong commercial decision

The most dangerous measurement result is often not an obviously bad one. It is a precise number built on an assumption that nobody thought to question.

Geo lift testing can provide valuable evidence of marketing incrementality. A strong synthetic control, statistically significant lift and narrow confidence interval can all strengthen confidence that an effect occurred in the markets tested.

They do not automatically prove that the same incremental ROAS should be applied nationally.

Before turning a local geo lift into a national investment decision, businesses need to understand whether the test markets represent the wider market, whether the stratification captures the factors that determine treatment response, whether the intervention was actually delivered consistently, and how much local variation exists underneath the aggregate result.

Identifiability tells you whether you measured an effect in the markets tested.

Transportability tells you whether that effect is useful anywhere else.

For marketing budget allocation, you need both.

Growth Dynamics helps brands design and independently review geo lift testing, incrementality testing and broader marketing measurement frameworks. If you are planning a geo lift test, or already have an incremental ROAS that is about to influence a significant budget decision, we can review the test design, stratification, transportability assumptions and commercial implications before you act on the result.

Previous
Previous

Incremental ROAS vs Attributed ROAS: How to Find Wasted Marketing Spend

Next
Next

You can only sell to people willing to buy from you