Marketing Mix Modeling Limitations: 7 Checks Before You Trust an MMM

Marketing mix modeling is one of the most useful tools available for strategic marketing allocation, but I think its limitations are often discussed in the wrong order.

Most lists begin with data quality, multicollinearity, priors and model fit. Those things matter, but they sit underneath a more fundamental problem: signal-to-noise.

An MMM is trying to identify relatively small marketing effects inside a much larger business outcome that is changing for many other reasons. The easier that marketing effect is to see in the historical data, the more useful the model can become. Large channels that move materially over time and create effects reasonably close to the period in which the investment occurs provide much more information than a small channel whose effect is spread gradually across several months.

This is one reason MMM can be extremely useful while also creating a false impression that every number in the final waterfall has been measured with similar reliability.

They have not.

Google's current Meridian documentation makes this problem unusually explicit. It notes that national MMMs often contain many variables relative to the amount of available data, that insufficient media variation makes estimation harder, and that low-spend channels are particularly likely to contain too little signal for the data to meaningfully update the model. Google even recommends combining or removing low-spend channels in some circumstances because the information simply may not be there.

That is the limitation I would understand before all the others.

A model can always produce an answer. The important question is how much of that answer came from information that genuinely existed in the data, and how much came from the decisions required to make the model produce one.

1. Is the marketing signal large enough to separate from normal business noise?

Imagine a company generating hundreds of millions in revenue while spending a relatively small amount on one media channel. Even if that channel genuinely produces a positive commercial effect, the resulting movement in revenue can be tiny relative to ordinary variation caused by seasonality, price, competitors, promotions, distribution, economic conditions and everything else happening to the business.

The problem gets harder when the effect is spread over time.

I demonstrated this separately using an incrementality simulation because it makes the signal-to-noise problem easy to see. We created two scenarios containing exactly the same $5,000 revenue impact. In the first, the effect appeared largely within seven days. In the second, the same impact was distributed across 60 days. A Difference-in-Difference analysis recovered roughly $4,900 in the short scenario but could not detect the 60-day effect at conventional significance levels, even though we knew the effect existed because we had added it ourselves.

That was an experimental demonstration rather than an MMM, but the statistical problem is closely related. Spreading an effect over time reduces its amplitude relative to everything else moving the outcome. The signal has not necessarily disappeared economically. It has become harder to distinguish statistically.

This naturally creates an asymmetry in what marketing measurement can identify well. Large interventions with reasonably immediate effects tend to leave stronger fingerprints in the data. Smaller investments and long-term effects leave weaker ones.

That matters because those weaker effects can include precisely the marketing activity a business most wants to understand: smaller emerging channels, long-consideration activity and investments intended to build future demand rather than generate an immediate transaction.

In my own content I have described MMM as a small-data problem for exactly this reason. Two or three years of weekly history gives roughly 100 to 150 observations, yet businesses often want the model to separate numerous media channels, controls, seasonal patterns, lags and response curves from those observations. Google's Meridian documentation now gives a similar example: with 12 media channels, six controls and eight knots, two years of weekly data provides only around four observations per parameter before even accounting for adstock and Hill parameters, which it describes as too little for reliable estimation under that strict calculation.

This is why model sophistication has a boundary. Better methods can extract information more efficiently and impose more credible structure, but they cannot manufacture a strong marketing signal where the historical business never generated one.

2. For long-term effects, how much is the data saying and how much is the modeller saying?

This is where I think the most important hidden limitation of MMM appears.

An econometric model can represent long-term advertising effects perfectly well. Adstock, carryover functions, time-varying parameters and Bayesian priors can all be used to allow marketing investment today to influence outcomes weeks or months later.

The fact that a model is capable of representing a six-month advertising effect does not establish that the data identified a six-month effect.

Suppose television increased in January and the model estimates that part of its commercial impact persisted for 26 weeks. During those 26 weeks, paid social changed, search demand moved, competitors altered their activity, pricing changed, the creative rotated, distribution shifted and the underlying brand continued developing.

The statistical question becomes much harder than whether January television spend correlates with later revenue. We need enough variation in the data to distinguish the specific lag structure from the many alternative explanations capable of producing a similar historical outcome.

When the signal is weak, modelling decisions become increasingly important.

The modeller decides which lag structures are possible, what priors are reasonable, which controls enter the model, how the baseline behaves, how saturation is represented and how much flexibility exists over time. None of these choices is inherently illegitimate. A model needs structure.

The problem is forgetting that the structure helped create the answer.

This is why I would be much more comfortable saying that an MMM has modelled a long-term effect than saying that it has proven one.

For strong short-term signals, reasonable modelling choices may produce broadly similar conclusions because the data dominates the decision. For weaker long-term effects, several plausible specifications can potentially explain the same history while assigning materially different value to marketing.

At that point, the modeller's decisions can speak more loudly than the data itself.

Google's own Bayesian documentation makes the underlying mechanism clear. The posterior estimate combines the information contained in the data with the prior information supplied to the model. When the data is weak, the prior understandably has more influence.

There is nothing wrong with that statistically.

The problem is commercial interpretation. A precise-looking ROI derived largely from assumptions should not be presented to leadership in the same way as an effect that the historical data strongly identifies.

3. What does a two-year channel coefficient actually represent when the advertising kept changing?

MMM normally operates at a relatively aggregated channel level. Google itself describes Meridian as a macro tool focused on channel-level analysis rather than recommending campaign-level MMM.

That aggregation is useful, but it creates another important limitation.

What exactly is “Meta” over two years?

The creative changed. Campaign objectives changed. Audiences changed. Placements changed. Bidding systems changed. Media costs changed. The platform itself changed. Some campaigns may have been excellent and others poor.

The coefficient produced by the MMM is therefore not measuring some permanent causal property belonging to Meta.

It is describing the historical portfolio of Meta activity that happened to exist during the modelling period.

Creative makes this particularly important.

System1 and Effie's current Creative Dividend work combines evidence from 1,265 campaigns and more than 100,000 tested advertisements. Their findings emphasise the joint role of creative quality and media support, with those factors together accounting for 60.1% of reported campaign business results on average in their analysis. System1 also cites Paul Dyson's profitability analysis, which places creative quality behind only brand size and therefore as the largest profitability multiplier that marketers can directly control.

Whether one accepts every estimate in that literature is secondary to the measurement problem it exposes.

Creative is not decoration around a stable channel treatment.

Creative is part of the treatment.

If a company spent £20 million on Meta over the previous two years using hundreds of different advertisements, the MMM coefficient describes the average economics of that changing creative portfolio combined with the media, audiences, placements and bidding system that distributed it.

Next year's treatment may be very different.

This is already explicit in the Growth Dynamics methodology: a channel coefficient averaged over two years of rotating creative is a property of a portfolio that no longer exists.

That does not make the coefficient useless. It limits what we should infer from it.

If the model says historical Meta investment produced an ROI of 2.4, that can be useful evidence for strategic allocation. I would be much more cautious about interpreting it as evidence that another £5 million deployed behind entirely different creative next year will inherit the same economics.

The more advertising effectiveness depends on the quality of the execution inside the channel, the weaker the idea of a channel coefficient as a permanent characteristic becomes.

4. Is the model measuring marketing, or demand that caused the marketing?

The next limitation is confounding.

Marketing investment does not normally vary randomly. Companies change spend because something else is happening.

Budgets increase during important commercial periods. Search spend rises when more people search. Promotions trigger more advertising. New products receive larger launches. Strong sales periods can encourage teams to increase investment further.

This makes it possible for marketing activity and revenue to move together even when part of the relationship originates somewhere else.

Paid search is the clearest example.

When underlying demand increases, more people search. More queries create more opportunities to serve paid-search advertisements, which can increase clicks and spend. Revenue also increases because more customers wanted the product in the first place.

Google's own Meridian guidance explicitly identifies query volume as a potential confounder when modelling paid search because failing to control for it can overestimate search's causal effect.

The same conceptual problem exists more broadly around brand demand. Someone searching for a company by name already possesses some level of awareness or preference. The paid advertisement can still have incremental value, but the query itself was created upstream.

This is why a coefficient should never be interpreted without understanding the system that generated the input variable.

Organic search creates another awkward boundary. In the Growth Dynamics approach, improvements to organic search and specific paid-search account changes remain outside the MMM decision layer because a weekly model over roughly two years is too coarse for many of those operational questions, while organic demand is intertwined with the baseline and paid search carries demand created elsewhere.

The fact that a variable can be included in a model does not mean that the model is the best instrument for every decision involving that variable.

5. How much of the answer depends on assumptions that reasonable people could change?

Every MMM requires modelling decisions, and this becomes much more important as the signal weakens.

The modeller has to decide which controls matter, how seasonality behaves, what prior distributions should be used, how long advertising effects may persist, which variables can vary over time and what saturation functions are plausible.

With strong data, reasonable choices may converge towards a similar commercial conclusion.

With weak data, the specification can become part of the result.

That is why I think sensitivity to modelling decisions is often more important than the confidence interval people focus on at the end.

An ROI credible interval describes uncertainty conditional on a particular model specification. It does not automatically contain the uncertainty created by choosing a different but equally defensible specification.

What happens if we alter the television lag?

What happens if we change the baseline structure?

What happens when a competitor variable is added?

How dependent is the channel contribution on a specific prior?

If several reasonable versions of the model fit the business similarly but recommend moving money in different directions, that specification uncertainty matters more to me than whether one version has a narrow posterior interval.

This is also why Bayesian priors need to be visible.

Meridian explicitly treats prior-posterior movement as a model-health check because when the posterior remains close to the prior, the data may simply not contain enough information to change the model's initial belief. Google specifically notes that sparse, noisy, low-variation and low-spend channels are susceptible to this problem.

A prior can be excellent. If it came from a strong relevant experiment, it may be considerably better than forcing a noisy observational history to generate an unconstrained answer.

But an experimentally informed assumption is still different from the historical MMM independently identifying the effect.

Understanding that difference is part of using the model honestly.

6. Are the response curves based on behaviour the business has actually experienced?

Response curves are one of the reasons MMM becomes so attractive for capital allocation. A historical decomposition says what the model believes happened. A response curve appears to say what happens next if spending changes.

That requires another modelling leap.

Suppose a company has historically spent between £900,000 and £1.1 million per week on a channel. The MMM can still draw a response curve extending to £2 million and calculate a marginal ROI at that level.

But the business has supplied very little direct evidence about what happens at £2 million.

The curve outside the historical range is increasingly determined by the functional form and parameter estimates rather than observed variation.

That does not mean the optimiser should never recommend a level of spend the business has not tried before. If it could not, it would be considerably less useful.

It means the recommendation should carry more uncertainty as it moves away from the region where the business actually generated data.

The same principle applies to saturation. A model estimating that a channel is close to saturation is much more persuasive if the historical record contains meaningful increases and decreases around that point than if spend remained almost constant throughout the modelling period.

Response curves also inherit the earlier problem around creative. If the historical curve was estimated using one portfolio of advertising, moving dramatically up the curve with different creative, inventory and bidding conditions may produce a different response.

The model is estimating the economics of the historical treatment.

The future business is deciding whether to deploy a new one.

7. Does the MMM survive evidence that was created outside the MMM?

The final check is the one that determines where MMM sits in the Growth Dynamics measurement system.

I do not want an MMM to be the only witness in its own trial.

If the model says a major channel has an ROI of 6 while a well-designed incrementality experiment repeatedly finds something closer to 2 under relevant conditions, I want to understand the disagreement.

The experiment is not automatically right. It has a treatment, a population, a time period and limits around transportability.

But it provides evidence that was not produced by fitting another version of the same historical revenue series.

The same applies to observed and qualitative evidence.

If an MMM says brand activity contributes almost nothing while Share of Search, buyer research, consideration and baseline revenue are moving strongly, that discrepancy deserves investigation.

If search appears to be driving the business while customers overwhelmingly report knowing the brand before they entered search, I want to understand whether the model is assigning demand creation to the channel that captured it.

If a channel's historical coefficient remains strong but the company's creative quality has deteriorated materially, I would be cautious about using that coefficient to project future returns.

These lower layers of evidence do not exist to compete with MMM.

They constrain it.

That is why Growth Dynamics places MMM last. The model consumes buyer evidence, observed demand, mid-funnel behaviour, knowledge of creative and media execution, and designed causal evidence where it exists. The offer itself describes an MMM bought first as a model forced to infer too much from correlated historical data, while an MMM built after the earlier evidence can become the strategic backbone.

This is also why backtesting matters. The model should survive different historical regimes rather than being trusted because one training window produced an attractive decomposition. If contribution shares and baseline estimates move dramatically when the window changes, that instability belongs in the allocation decision rather than being hidden in the modelling appendix.

The biggest limitation of MMM is not the mathematics

None of this is an argument that marketing mix modeling does not work.

It is an argument for being much more precise about where it works well.

MMM is particularly useful when the business has large media investments, meaningful variation in those investments, enough historical data and a strategic allocation question at the level the model can realistically support.

In those conditions, it can provide an enormously useful whole-mix view. It can quantify uncertainty, estimate diminishing returns, incorporate experiments, model online and offline channels together and give leadership a framework for discussing large capital movements.

The problem begins when that strength is extended indiscriminately to every piece of marketing.

The same model may contain a large, variable performance channel whose commercial response appears quickly and a small brand channel whose impact accumulates gradually over months. Both can receive an ROI number in the output.

That does not mean the two numbers were identified with equal strength.

In practice, the first effect leaves much more signal for the model to work with. The second requires the model to distinguish a smaller, slower effect from a much larger amount of business variation. As that signal weakens, assumptions about priors, lag, baseline, controls and functional form play a larger role.

This is the part of MMM that I think buyers should understand before purchasing one.

The model gets more opinionated precisely where the data gets quieter.

That does not make the opinion useless. A well-informed model can combine previous experiments, sensible priors, customer behaviour and domain knowledge into a far better allocation framework than raw attribution.

But we should stop pretending that the sophistication of the model has removed the uncertainty underneath it.

What I would want to know before moving budget

Before acting on an MMM, I want to understand the strength of signal behind the recommendation rather than simply the ROI estimate itself.

For each major channel, how much spend and variation existed? Was the effect large enough relative to normal revenue noise for the data to learn anything meaningful? Did the posterior actually move away from the prior?

For long-term effects, what variation identifies the lag? Do alternative reasonable decay assumptions materially change the result? If they do, I want that uncertainty exposed.

I also want to know what the channel variable represents operationally. How much did creative, placement, audience and platform execution change during the modelling period? Is the model describing something sufficiently similar to the advertising the company intends to run next year?

Then I want the response curves compared with the historical spend range, the model rebuilt over different periods and the major conclusions compared with independent experiments and the observed evidence elsewhere in the business.

The point is not to make MMM pass an impossible standard.

The point is to understand what decision the evidence can actually support.

A directional estimate for a small channel may still be useful. It should not carry the same confidence as a strongly identified large channel.

A long-term brand estimate may help inform the wider story. It should not suddenly become factual because the adstock function returned a median.

And a two-year historical coefficient can inform next year's allocation without pretending that a completely different creative portfolio will behave identically.

MMM should be the strategic backbone, not the whole measurement system

The most useful role for MMM is not to replace every other form of marketing measurement.

It is to bring the wider system together once the evidence underneath it exists.

Observable demand tells us whether preference is moving. Buyer evidence helps explain why customers choose the company. Mid-funnel behaviour shows where demand is appearing before revenue arrives. Creative and placement analysis tells us what treatment the media budget is actually delivering. Incrementality creates stronger causal evidence around important interventions.

MMM then helps connect those lenses across the whole portfolio.

That sequence matters because the weakest part of an MMM is often exactly where marketers most want certainty: small effects, long-term effects and activity whose execution changes constantly.

Those are the areas where I want more evidence, not simply more sophisticated modelling.

Growth Dynamics builds and independently reviews MMM for B2C brands making major allocation decisions. The objective is not to produce an ROI for every row in the media plan. It is to establish which effects the data genuinely identifies, where assumptions are doing more of the work, what the model cannot reasonably answer and where the wider evidence suggests capital should move.

A model that admits where the signal is weak is considerably more useful than one that gives every channel a confident number.

Previous
Previous

Should You Bid on Your Own Brand? How to Measure Brand Search Incrementality

Next
Next

How to Audit Paid Media Spend for Waste Before You Cut the Channel