Marketing Measurement for Long Purchase Cycles: Why More Complex Models Do Not Solve the Problem

Marketing measurement becomes considerably harder when the purchase happens weeks or months after the marketing activity that may have influenced it. This is common in high-ticket categories, businesses with substantial offline sales, and products where customers spend a long time researching and comparing before they make a decision.

The natural response is often to make the measurement model more sophisticated. Extend the lag structure in an MMM. Introduce more flexible priors. Add time-varying parameters. Track more customer touchpoints. Run longer incrementality tests.

These approaches can all be useful. But they do not solve the fundamental problem created by a long consideration journey.

The ability to model a long-term marketing effect is not evidence that the data can identify that effect correctly.

An econometric model can easily be designed to allow an advertising campaign to influence revenue for three months, six months or longer. It can estimate decay rates, carryover, saturation and interactions with other variables. It can also fit historical revenue extremely well.

None of that proves that the resulting decomposition is correct.

As the gap between marketing activity and purchase increases, more things can influence the final outcome. Brand preference changes. Competitors alter prices. Distribution moves. Promotions begin and end. Other marketing channels change. Customers encounter the brand offline. Economic conditions shift. Category demand rises or falls.

The modelling problem therefore becomes an identification problem. There may be several plausible explanations for the revenue observed, and a more flexible model may simply become capable of representing more of them.

That is not the same as creating more evidence.

Long purchase cycles create an identification problem

Consider a business selling a product that customers typically buy five minutes after clicking an advert. Connecting marketing activity with the commercial outcome is not necessarily easy, but the time between intervention and outcome is short.

Now consider a £5,000 purchase with a three-month consideration cycle. A customer may become aware of the brand in January, visit the website several times, search the category, compare competitors, visit a store, discuss the purchase with their partner and finally buy in April.

The second business does not simply need a longer attribution window. It has a fundamentally harder measurement problem.

Revenue in April can potentially be related to advertising from January, February, March or April. It can also be related to underlying brand demand, competitor changes, distribution, pricing, economic conditions and dozens of other factors.

The challenge is not creating a model capable of associating January advertising with April revenue. That is straightforward.

The challenge is establishing what evidence allows us to distinguish that explanation from the alternatives.

This distinction matters because marketing models are usually used for capital allocation. A company does not merely want to explain why revenue moved historically. It wants to know whether another £5 million invested in a particular channel will create sufficient incremental return.

That decision requires more than a plausible historical story.

Modelling a long-term effect does not prove the effect exists

Marketing mix models commonly use carryover or adstock functions to represent the possibility that advertising affects outcomes after the initial exposure. This is entirely reasonable. Advertising effects do not necessarily disappear the week the money is spent.

The problem appears when the existence of the modelled lag is confused with evidence that the lag has been identified.

Suppose an MMM concludes that television continues influencing revenue for 26 weeks. Before using that result to move budget, I want to understand what information in the historical data distinguishes a 26-week effect from an 8-week effect or a 12-week effect.

This becomes particularly difficult when several variables move together.

Television investment increases. Paid social investment rises at the same time. Brand search subsequently increases. Revenue rises later. The company also expands distribution and a major competitor changes its pricing.

The model now has to separate the contribution of multiple correlated variables while estimating lag, saturation, baseline demand and other business effects.

A sufficiently flexible model will usually produce an answer.

The question is whether the data contained enough independent information to establish that this particular answer was the correct one.

If several different specifications fit historical revenue reasonably well but produce materially different channel contributions, the problem has not been solved by adding modelling flexibility. The uncertainty has simply moved inside the model.

Prediction and causal decomposition are not the same problem

This is why I do not think an MMM should be validated simply because it predicts holdout revenue well.

Prediction is useful. If a model consistently fails to predict outcomes outside its training period, that is an obvious warning sign.

But successful prediction does not prove that the decomposition underneath the prediction is correct.

Imagine three models all forecast an eight-week holdout with similar accuracy. The first concludes that television generated £20 million of incremental revenue. The second estimates £8 million. The third attributes only £3 million to television and places substantially more of the growth into baseline brand demand.

All three cannot simultaneously provide the correct causal decomposition simply because they predicted total revenue well.

This matters because the business decision is based on the decomposition. If the recommendation is to move millions into television because the estimated response curve says it is underinvested, we need evidence that supports that interpretation rather than merely evidence that the model can forecast total sales.

A model can therefore be useful for prediction while remaining uncertain about the contribution of correlated marketing variables.

Long purchase cycles amplify this problem because there is more temporal distance over which competing explanations can operate.

Tracking more of the customer journey does not solve it either

Attribution approaches the same problem from the opposite direction. Instead of using aggregate historical relationships, it attempts to observe more of the customer journey.

For a long-consideration business this can feel attractive. If we can track the original visit, the paid social click, the email interaction, the branded search and the final purchase, perhaps we can reconstruct what caused the sale.

What we have actually reconstructed is the part of the journey we were able to observe.

A customer may already have known the brand for years. They may have received a recommendation from a friend, seen offline advertising, visited a store or simply considered the brand more trustworthy than its competitors. None of those influences necessarily appear inside an attribution system.

Even when a digital interaction is observable, its position in the journey does not establish its causal contribution. A branded paid-search click immediately before a £5,000 purchase may be extremely easy to attribute while contributing very little incremental demand.

Tracking more touchpoints can improve operational understanding, but it does not turn an observed customer journey into a causal one.

Long purchase cycles also make revenue too slow for many decisions

There is another practical issue. Even if final revenue could be measured perfectly, it may arrive too slowly to manage marketing effectively.

If the average customer takes twelve weeks to purchase, a campaign launched today may not generate enough mature revenue for useful analysis until the following quarter. Marketing teams cannot wait three months for every decision.

Platforms fill that gap with faster metrics such as clicks, sessions, view-through conversions and platform-attributed revenue. The problem is that these metrics often describe activity inside the advertising system rather than whether meaningful customer demand is progressing.

This is where observed mid-funnel behaviour becomes particularly valuable.

A customer may not buy today, but they can still show signs of increasing commercial intent. Depending on the business, that might include repeated product exploration, visiting pricing information, using a store locator, adding a product to a basket or returning to commercially important parts of the website.

These signals do not establish causality. Their value is that they are observable and traceable.

We can see what happened, where the traffic came from, how the customer behaved afterwards and whether the pattern exists consistently across campaigns. That can tell us a great deal about where meaningful demand is appearing while final revenue is still developing.

Start with evidence you can actually inspect

This is why I do not believe a long consideration journey should automatically push a business towards heavier modelling.

Before attempting to infer a three-month causal chain, there is usually a great deal of evidence available directly.

Is branded demand growing relative to competitors? Is baseline revenue changing? Are more customers reaching commercially important parts of the journey? Which campaigns are associated with those behaviours? Which campaigns generate substantial traffic but almost no meaningful progression?

Observed evidence cannot answer every question, but it has an important advantage: the path from behaviour to conclusion is relatively easy to inspect.

If a campaign appears to generate high-intent activity, we can examine the underlying sessions. We can see whether the result is widespread or concentrated in a strange placement. We can examine whether traffic quality changed at the same time and whether the pattern persists.

Compare that with receiving a model output saying that a channel contributed £4.2 million. Understanding why the model produced that number may require interrogating the priors, transformations, lag structure, saturation assumptions, controls and correlations between variables.

Sometimes that additional assumption load is justified. Sometimes the decision simply did not require it.

Qualitative evidence answers a different part of the journey

Long purchase cycles also make qualitative customer evidence more valuable because much of the purchase decision happens outside observable marketing systems.

Behavioural data can tell us what somebody did. It is much weaker at explaining why they ultimately chose one brand over another.

For that, I want to hear from the customer.

Post-purchase surveys, non-buyer research, sales-call transcripts and reviews can expose motivations that never appear in attribution data. A buyer may explain that they had known the brand for years, trusted its reputation, received a recommendation or believed it offered better quality than a competitor.

This evidence should not be confused with causal measurement either. A customer saying that brand reputation influenced them does not quantify the incremental revenue generated by a television campaign.

It answers a different question: what factors does the buyer report as important to their decision?

That evidence can then inform what we measure elsewhere, which hypotheses deserve testing and whether the causal story produced by a model is consistent with what customers are actually telling us.

Media execution needs to be understood before the channel is judged

The same principle applies inside the media channel.

A channel-level result is an average of whatever happened inside that channel. It says very little about the quality of the intervention itself.

Creative matters. Placement matters. Audience selection matters. Campaign configuration matters. The bidding system matters.

Two companies can each spend £10 million on paid social while delivering radically different marketing interventions.

One might put most of the budget behind highly effective creative shown to relevant audiences in valuable environments. Another might distribute weak creative through poor placements while its bidding system concentrates spend in areas that are easy to convert but unlikely to create additional demand.

A channel coefficient averages all of that together.

Before concluding that the channel deserves more or less money, it is worth understanding what actually happened inside it.

This becomes relevant to causal measurement as well. Creative, placements, audiences and bidding behaviour are part of the treatment. If they change materially across markets or over time, then the treatment itself is changing.

Incrementality gives stronger causal evidence, but within defined conditions

When the allocation decision genuinely requires causal evidence, incrementality testing becomes extremely valuable.

But reducing incrementality to the question “would the revenue have happened anyway?” misses much of what needs to be understood.

The more useful questions are: what changed because of the intervention, how large was the effect, what treatment was actually delivered, how consistent was the response, what uncertainty sits around the estimate, under what conditions did the effect occur, and how far can we transport that result beyond the experiment?

Long purchase cycles make those questions harder.

If an experiment increases advertising for six weeks but customers normally take twelve weeks to purchase, stopping the analysis when the treatment ends may miss revenue generated by customers whose journeys began during the experiment.

Extending the observation period allows more outcomes to mature, but it also gives more time for competitor activity, pricing, promotions, other campaigns and normal business variation to affect the result.

The experiment therefore needs an explicit view of timing rather than simply a longer window.

The same applies to transportability. A causal effect identified in a set of geographies is evidence about those markets under the conditions tested. Applying the resulting ROI nationally requires another set of assumptions about whether the remaining markets will respond similarly.

Designed evidence is powerful because the counterfactual is created deliberately. It is not assumption-free truth.

Why our answer is not simply a better MMM

It would be easy to respond to all of these problems by saying that companies need a more advanced econometric model.

I do not think that is the right conclusion.

If the historical data cannot distinguish between several plausible causal explanations, giving the model more freedom does not create identification. In some cases it simply allows a more sophisticated model to produce a more sophisticated story.

The better answer is to build independent evidence around the model.

Understand whether preference and demand are changing. Understand what buyers say influenced their decision. Observe where meaningful customer intent is appearing. Understand whether media execution is actually effective. Run designed causal tests when the decision justifies them.

Then build the MMM.

By the time econometrics is used, it no longer has to construct the entire marketing story from two years of correlated weekly data. There are external pieces of evidence capable of informing its assumptions and, more importantly, challenging its conclusions.

That is a fundamentally different role for the model.

Observed first, designed second, modelled last

This is why Growth Dynamics approaches measurement in layers.

We begin with growth quality: whether preference is increasing and whether that preference is being monetised.

We then use qualitative buyer evidence to understand what customers themselves report as influencing their choice.

Next comes observed traffic quality and mid-funnel behaviour, which helps identify where meaningful intent is appearing while revenue is still developing.

We also examine execution economics across creative, placements and campaign mechanics. A poor channel result may be the consequence of poor execution rather than evidence that the channel itself has no value.

When a decision requires causal evidence, incrementality testing asks what changed because of a deliberate intervention, how large the effect was, under what conditions it occurred and how confidently that result can support the wider allocation decision.

MMM comes after those layers, not before them.

Its role is to help decompose the broader marketing system and support strategic budget decisions, informed and challenged by the evidence underneath it.

Putting econometrics last does not make it less important. It makes the model more useful because we have more information available to interrogate what it claims.

How I would assess a claimed long-term marketing effect

When a model concludes that advertising produced a substantial long-term effect, the first thing I want to understand is what variation in the data actually identifies that duration.

Allowing a 26-week carryover is not evidence for a 26-week effect.

I also want to know how sensitive the conclusion is to modelling choices. If reasonable changes to priors, lag structures, controls or modelling windows materially change the channel contribution, that uncertainty matters to the capital allocation decision.

Then I want to compare the decomposition with independent evidence. Does observed demand support the story? Does customer research? Do experiments? Does the implied response curve make sense given what happened inside the channel?

Finally, the amount of scrutiny should depend on the decision being made. A model output used to generate a hypothesis can tolerate more uncertainty than an output being used to reallocate £20 million.

The purpose of validation is not to prove that the model is perfect.

It is to understand how far we can trust the answer.

Long consideration journeys require more evidence, not more modelling confidence

The difficulty of measuring a long purchase cycle does not come from a shortage of sophisticated methodologies.

The difficulty comes from trying to identify small marketing effects across long periods in which many other things are changing.

An econometric model can represent those effects mathematically. Attribution can record more of the observable journey. Experiments can create stronger causal evidence under controlled conditions.

None of those methods reconstructs the complete customer decision perfectly.

The solution is therefore not to choose one methodology and ask it to explain everything.

It is to build a chain of evidence, beginning with what can be observed directly and increasing the assumption load only when the question genuinely requires it.

That produces something more useful than a complicated model that everybody agrees not to question.

It produces an allocation decision whose evidence, assumptions and uncertainty can actually be understood.

Growth Dynamics works with B2C brands where long consideration journeys, high-ticket purchases and offline sales make simplistic marketing measurement particularly unreliable. We combine observable growth signals, qualitative buyer evidence, mid-funnel behaviour, execution economics, designed causal evidence and MMM to determine where marketing capital should be invested, held or removed.

The objective is not to build the cleverest marketing model.

It is to make an allocation decision you can defend.

Next
Next

Incremental ROAS vs Attributed ROAS: How to Find Wasted Marketing Spend