How to Validate a Geo Lift Test Before Moving National Budget

This is the validation checklist for a geo lift result that is already on the table.

The broader geo lift article explains the main limitations of the method: market selection, transportability, synthetic control fit, treatment heterogeneity and national extrapolation. This article starts later in the process. It assumes a test has been run, a vendor has delivered a lift number and the business is deciding whether that number deserves to move budget.

At that point, the measurement work is unfinished. The team needs to know whether the treatment was actually delivered, whether the counterfactual was credible, whether the calendar created the gap, whether a few markets carried the average and whether the tested markets represent the country the budget will be projected onto.

The practical question is whether this result can support the decision ahead.

Sometimes the answer is yes. Sometimes the test should be rerun. Sometimes the result is directionally useful but too fragile to move a large amount of money.

That is the purpose of geo lift test validation.

Validate the treatment before reading the result

The first check is whether the treatment in the test was the treatment you think you ran.

This sounds obvious, but it is one of the easiest places for a geo test to break. Marketing teams often read the revenue result first. If revenue moved, the test worked. If revenue did not move, the channel failed. That skips the most important operational question: did the treated and control markets actually receive different media exposure?

In one review, a social channel looked ready to cut. The holdout appeared clean. The channel had been switched off in test markets, total revenue barely moved and the spend no longer looked defensible.

Then we checked the traffic.

The supposedly held-out campaigns were still driving about half their original traffic into the test markets. The channel was never truly off. The flat revenue result could mean the channel was not incremental, or it could mean the test had failed because the market was still being partially treated. The result was unusable in either direction.

Delivery validation comes before effect validation.

For each market, check whether spend, impressions, clicks, sessions and meaningful downstream behaviour changed in the way the design required. If the control markets were supposed to go quiet, confirm they went quiet in analytics. If the treatment markets were supposed to receive a spend increase, confirm that increase actually happened. If the platform reallocated budget inside the treatment arm, understand which markets received more pressure and which received less.

The platform does not care about your test design. It has an optimisation objective. It can move delivery towards markets where response is easiest to find, even when the experiment needs balanced pressure across the cell.

If the bidder changed the treatment, the lift number is no longer measuring the planned intervention.

Add a spend-balance gate

Most geo test validation focuses on outcome similarity before the test. Parallel trends, holdout R-squared, beta drift and rolling beta stability are all useful checks. They tell you whether the treated and control markets had a credible historical relationship in the outcome.

They do not tell you whether the treatment markets were clean.

A test can pass those outcome checks while still being biased by the way spend is distributed across geographies. Imagine the treatment cell contains markets where spend share is much higher than revenue share because of fake traffic, bot farms or last-click attribution. Those markets can be structurally leaky before the test starts.

If the rest of the split is built to make the revenue trends match, the design can look statistically acceptable. The missing check is spend balance. The outcome history may line up while the underlying media economics do not.

That creates a dangerous pattern in channel-sponsored tests.

The test passes the validation gates the vendor reports. The lift reads well. The result gets projected nationally. But the design may have placed the most favourable or most distorted markets in treatment, then used outcome matching to make the rest of the experiment look balanced.

Revenue matching is not enough.

Before accepting the split, inspect the relationship between spend, revenue and traffic by geo. Look for markets where spend share, revenue share and traffic quality are out of line. Compare media mix, channel saturation, branded demand, offline exposure and baseline conversion patterns. If the treatment group starts with unusually favourable economics, the test can measure a real local gap while still giving you the wrong national answer.

A good validation pack should show why the selected markets are credible from a spend and response perspective, not only why their historical revenue trend matched the control.

Test the counterfactual, not only the lift

The counterfactual is the heart of the test. If the test says revenue was higher than it would have been, the entire claim depends on the "would have been" line.

That line deserves more scrutiny than it usually gets.

I once built a counterfactual for markets where no test had run. Business as usual continued. There was no campaign lift to detect. Then I ran a normal geo lift read anyway, using synthetic control logic to build the expected revenue line and compare it with actual revenue.

Every significant effect found in that exercise was false by construction.

The results were uncomfortable. A naive formula interval called 44% of the 30-day windows significant and 69% of the 90-day windows significant. CausalImpact's own interval was better, but still called 29% and 15% significant. A placebo test, where the treated markets were swapped and the same logic was rerun, stayed around 5% or below.

The same data. The same counterfactual. A very different answer depending on how uncertainty was computed.

This matters because a confidence interval can be tight for the wrong reason. If the synthetic control matched the treated markets closely before the test, a formula may assume that any later gap is unlikely to be noise. But markets stop tracking each other for reasons that have nothing to do with the campaign.

A coastal market and an inland market may move together through winter. Summer arrives, tourism changes one market more than the other and the counterfactual breaks. The model sees the gap. The interval says it is meaningful. The campaign may have had nothing to do with it.

Placebos and backtests help expose that problem. They ask whether the same method finds effects when it should not. If the method produces frequent "wins" in periods or markets where no intervention occurred, the clean lift number deserves much less confidence.

Before trusting the lift, ask how the interval was computed and how often the method generates false positives on the same business.

If the validation cannot answer that, the result is not ready to move money.

Treat the calendar as a competing explanation

Short natural experiments are especially vulnerable to the calendar.

A team I spoke to had a major channel stop delivering for three days because of a payment issue. Conversions fell, and the implied effect looked much larger than the company's MMM estimate for that channel.

The first interpretation was that the MMM was under-crediting the channel.

Maybe it was. But the pause ran Friday to Sunday for a B2C audience that buys heavily around weekends and payday. Those days already carried unusual demand. Move the same outage to a quiet Tuesday and the apparent channel effect could have looked completely different.

The outage may still contain useful evidence, but the result needs to survive a calendar check before becoming a budget rule.

Geo tests face the same issue. Day-of-week, payday cycles, holidays, local events, product drops, promotional calendars and weather can all move revenue independently of media. The shorter the test, the more those patterns can dominate the result.

Difference-in-differences can help because the control group absorbs some shared movement. It does not remove every local calendar problem. If the treated markets have a different timing pattern from the controls, or if the intervention happens to overlap a locally unusual period, the lift can still be distorted.

The validation pack should show the test window against relevant business calendars. It should explain why the pre-period and read period were chosen before the result was known. It should also show whether the result survives reasonable alternative windows that were defined as robustness checks rather than selected after the fact.

A read window chosen after seeing the data is a way to pick the most flattering version of the experiment.

Open the average

National budget decisions often receive one aggregate lift number. That number can hide a very uneven local reality.

Geographies are not interchangeable. Each market has its own demand, brand strength, competition, pricing sensitivity, channel mix, offline exposure, store footprint, customer mix and sales pattern. Two markets can have similar historical revenue trends while responding very differently to the same spend change.

If the test reports a positive average, inspect the market-level response.

Did most treated markets move in the same direction, or did two large markets carry the result? Did the strongest lift appear in markets with unusually strong brand demand? Did the weakest markets have different media saturation, offline availability, competitor pressure or traffic quality? Were there outliers that deserve to be removed, or are those outliers part of the reality you will face nationally?

The average matters, but it is not enough.

A few high-response markets can pull the national estimate upward and make a weak rollout look attractive. A few saturated markets can pull it downward and hide a treatment that works in the parts of the country where the next dollar would actually go.

The result should be read as a distribution before it becomes an allocation recommendation.

At minimum, I want to see market-level spend movement, outcome movement, pre-period fit, traffic movement, brand-demand movement and any local anomalies that could explain the result. If the treated markets disagree sharply with each other, the validation should say that clearly.

The decision may still be to invest. But the investment should reflect where the effect appeared, not only the blended average.

Check transportability before national projection

Identifiability and transportability are different questions.

Identifiability asks whether the test measured an effect in the markets tested. Transportability asks whether that effect can be applied somewhere else.

A geo lift test can be locally valid and nationally misleading.

This is the part many teams underweight. They test 10% or 20% of a country, estimate an incremental ROI and then apply that ROI nationally. The projection looks like arithmetic, but it contains a major assumption: the tested markets respond to spend in a way that represents the rest of the market.

That assumption often does too much work.

Market response can vary because brand strength varies. It can vary because the media mix is different. It can vary because online and offline sales behave differently. It can vary because some geographies already have higher saturation, different price sensitivity, different store coverage, different competitors or different traffic quality.

Synthetic control can create a strong counterfactual by matching historical revenue. That does not prove the treated and control markets have the same relationship between spend and revenue. Two regions can track together historically while having very different response mechanics.

Matching only on the KPI is insufficient.

A better validation reads the features of the regions, not just the outcome line. Compare average ROI by market, media mix, share of online and offline sales, brand strength, traffic quality, customer mix and baseline spend intensity. Where possible, look at spend and revenue per household by geo, because it helps reveal whether the treatment markets resemble the wider market economically rather than only statistically.

Transportability cannot be solved perfectly. The useful work is making the assumption visible and deciding whether it is acceptable for the size of the budget move.

Larger, more representative treatment designs can reduce the problem. They do not remove it. A test covering a tiny, unusually favourable slice of the country should not be allowed to dictate national budget with the same confidence as a broader design that has been checked across the features that shape response.

Local impact is fact. National projection is an assumption.

Good validation tells you how large that assumption is.

Read the channels around the holdout

A holdout does more than tell you whether revenue moved when one channel changed.

It also shows how the rest of the system reacted.

When a channel is turned off, organic demand may rise, paid search may absorb conversions, direct traffic may change, branded search may move or another channel may pick up sales that were going to happen anyway. Those movements are not noise around the test. They are part of the learning.

This matters because a flat revenue result has several possible meanings.

The channel may have been non-incremental. It may have been doing discovery work that other channels closed. It may have been cannibalising demand that transferred cleanly when the spend stopped. It may also have been held out badly, with traffic still leaking into the market.

The top-line lift number will not separate those explanations by itself.

Before cutting or scaling a channel, inspect cross-channel movement in the treated and control markets. Look at organic search, branded search, direct traffic, paid search, retargeting, email, store demand and meaningful mid-funnel actions where relevant.

If a channel disappears and another channel immediately captures the same demand, that can indicate cannibalisation. If a channel disappears and later-stage demand weakens downstream, that can indicate contribution that the final-click dashboard had been assigning elsewhere.

The budget decision should use that information.

Incrementality testing should show how demand moves through the system when the treatment changes, not only whether one channel passed or failed.

Clean the traffic before using it as validation

Traffic checks are useful in geo tests, but only if the traffic is worth reading.

Pulling raw GA4 sessions into a test platform can create false comfort. A campaign can produce large volumes of low-quality or bot-driven traffic that never had a meaningful chance of turning into revenue. Programmatic and display activity can also create noisy geographies, suspicious device mixes and inflated engagement signals.

Traffic still belongs in the validation pack, but it should be cleaned before it becomes evidence.

When validating a geo test, I want to know whether the traffic movement matches the treatment and whether the traffic has enough quality to matter. Sessions alone are weak. Time on site, page depth, important page views, add-to-cart, quote starts, store searches, pricing views and other meaningful actions can be more useful, depending on the business.

The same logic applies to halo effects. If a treatment is expected to create demand, the evidence may appear first in branded search, direct visits, organic behaviour or high-intent actions rather than immediate revenue. That is especially relevant for longer consideration journeys.

Those signals cannot prove causal revenue lift. They help validate whether the customer system moved in a way that is consistent with the test story.

If the lift number says the channel worked but traffic quality deteriorated, meaningful intent did not move and the result comes from a small set of strange geographies, the test needs more scrutiny. If the revenue result is inconclusive but demand and intent signals moved coherently across the treated markets, the correct decision may be more nuanced than "failed test".

Revenue remains the decision target. Clean supporting signals help explain whether the result is credible.

The decision is approve, reject or rerun

Geo lift validation should end with a decision, not a collection of charts.

I would not approve a national budget move from a geo lift test until these questions have been answered:

  • - Was the planned treatment actually delivered in the treated markets?

  • - Did the holdout or control markets remain materially untreated?

  • - Were spend, traffic and revenue balanced well enough before the test?

  • - Did the counterfactual survive placebo or backtest validation?

  • - Was the confidence interval computed in a way that controls false positives on this business?

  • - Did the test window avoid obvious calendar distortion?

  • - Was the read window pre-committed?

  • - Did the market-level results support the aggregate lift?

  • - Were the tested markets representative enough for national projection?

  • - Did cross-channel movement support or contradict the interpretation?

  • - Were traffic and mid-funnel signals cleaned before being used as evidence?

If those checks hold, the result can support an allocation decision. The exact decision still depends on effect size, uncertainty, marginal spend level and the commercial risk of being wrong.

If some checks fail, the answer is not automatically that the channel failed. It may mean the test failed.

That distinction matters. Cutting a channel because a contaminated holdout showed no lift is as dangerous as scaling one because a narrow interval found a false positive. In both cases the company is moving capital on a number that does not deserve that much authority.

The cleanest geo lift result is the one whose assumptions have been inspected and whose decision still survives.

That is the standard before a local result becomes a national budget move.

Previous
Previous

Mid-Funnel Marketing Measurement: How to Understand Demand Without an MMM

Next
Next

Retargeting Incrementality: How Much of Your Retargeting Budget Is Actually Creating Sales?