Marketing Mix Modeling Assumptions: How Much Are We Really Letting the Data Speak?

There is a paradox at the heart of Marketing Mix Modeling that I cannot get past.

We want the model to tell us what caused the outcome, but before it gives us an answer, we tell it quite a lot about how we think causality works.

We decide how advertising effects decay, what saturation should look like, what belongs in the baseline, which variables should be controlled for and what range of channel effects we consider plausible. Bayesian models add priors to that list.

Those choices define the causal world the model is allowed to consider. Once the decomposition comes out, it is easy to talk about the result as though the data itself told us what caused revenue.

That is the part I struggle with.

Assumptions are part of any causal model. The real question is how much those assumptions shape an answer that is later presented as data-driven.

The setup can materially change the answer

A channel contribution is not sitting inside the historical data waiting to be extracted.

Adstock determines how later revenue can be connected with earlier advertising. Saturation determines which response curves are possible as spend increases. The baseline determines how much revenue marketing can potentially explain. Controls change which competing explanations are separated from media. Priors constrain the range of effects the model considers plausible.

Change those choices and the decomposition can change with them.

I have seen the same underlying observations produce very different channel contributions depending on the MMM setup. The spend and revenue data stayed the same. The structure around that data changed.

That matters when an ROI or contribution number is presented as though it were simply a measurement of what happened. The number reflects both the historical observations and the system we used to interpret them.

I want to know how much the observations actually determined the answer.

Signal-to-noise determines how much room the assumptions have

A large part of the problem comes down to signal-to-noise.

MMM is trying to identify marketing effects inside revenue that is also moving because of price, promotions, seasonality, distribution, competitors, economic conditions and normal business volatility.

Large, variable channels with effects relatively close to revenue create a stronger signal. When investment changes materially and there is a clear commercial response, the observations have more information with which to constrain the model.

Smaller channels are harder. They can create real incremental value while producing movements that are tiny compared with normal changes in revenue. There is less signal for the model to separate from everything else happening in the business.

Long-term effects have the same problem. A meaningful economic effect spread across several months creates a weaker signal in each period. Over those months, many other things also change.

An MMM can represent a six-month advertising effect by giving the model a structure that allows the effect to persist for six months. That does not tell us how strongly the historical data identified that duration.

When the signal is weak, decisions around adstock, baseline, controls and priors have more influence over how much later revenue gets attributed back to earlier advertising. The model has been given a way to believe in a long effect, while the observations may have limited power to determine how much of that effect really occurred.

Being able to model a six-month effect does not prove that the data identified one.

Variation matters for the same reason. A channel that barely changes gives the model little information about what happens when investment changes. Two channels that move together are difficult to separate. Adding more weeks of the same pattern does not necessarily create more identifying information.

The model can still return a contribution for every channel. Those contributions are not necessarily equally well identified.

As the marketing signal becomes weaker relative to the noise, the modelling choices have more room to shape the causal story.

The modeller already decides what to believe

An MMM can produce several statistically plausible explanations of the same history. The modeller still has to decide which one is credible.

That decision uses knowledge that does not come from the model itself. The modeller knows what happened in the media plan, which channels changed, what previous tests showed, how the business normally behaves and which results look commercially plausible.

That knowledge can improve the model. It also means the final decomposition reflects more than the observations.

We ask the MMM to tell us what caused revenue, then rely on the modeller to decide which causal stories should be believed.

When the signal in the data is strong, the observations can constrain those choices more tightly. When the signal is weak, judgement around priors, response curves, carryover, controls and baseline can have much more influence over the contribution that eventually gets reported.

This is why the modeller matters so much. The model does not remove judgement from the process. It gives that judgement a formal structure.

The important part is being clear about where the answer came from.

Automation does not make those choices disappear

MMM is becoming much easier to automate. That can make the process faster, cheaper and more consistent.

I have seen the same underlying data produce very different channel contributions in Meridian and PyMC-Marketing. The observations did not change. The modelling system around them did.

Different frameworks and setups can make different choices around priors, adstock, saturation, baseline structure, controls, parameterisation and regularisation. Those choices can materially change the decomposition.

Automation moves more of those decisions into the framework itself. Someone still chose which response curves are available, how carryover can work, what default priors look like and how the baseline is constructed.

The signal-to-noise problem also remains. A small marketing effect buried inside normal revenue volatility does not become easier to identify because the model can be fitted in a few minutes. Two channels that always move together do not suddenly contain more independent information because we can test more specifications.

More flexible models can represent more possible explanations of the same history. The observations still need enough information to distinguish between them.

Automation can make MMM easier to build. It cannot manufacture identifying information that was never present in the data.

MMM needs evidence around it

I still think MMM is extremely useful. It gives us a way to look across a large marketing portfolio, estimate response curves and support allocation decisions across channels that cannot all be tested independently at the same time.

I do not think it should sit above the rest of the measurement system as the final authority on what caused growth.

Our framework starts with evidence that asks less of the model.

1. Market behaviour

We first look at what we can observe about demand.

Share of Search tells us whether the brand is gaining or losing relative search demand against competitors. Branded organic search shows whether more people are actively looking for the company. Baseline revenue helps us understand how much demand the business is monetising without paid channels claiming the final interaction.

These measures do not establish channel causality. They show us movements that actually happened in the market.

2. Buyer evidence

We then look at what buyers say.

Post-purchase surveys, non-buyer research, sales conversations and reviews help us understand what drove consideration, choice and rejection.

Tracking can show how someone reached the business. Buyer evidence helps explain why the business was considered in the first place.

3. Buyer behaviour

The next layer looks at what buyers do.

Clean traffic and meaningful commercial actions can show where demand is moving before final revenue matures. Pricing-page visits, quote starts, store searches, product configuration, add-to-cart behaviour and other high-intent actions can all provide useful evidence depending on the business.

This becomes especially valuable when the purchase cycle is long and the final sale arrives well after the marketing activity.

4. Execution economics

We also need to understand what actually happened inside the media investment.

A historical channel label can hide very different creatives, placements, campaign objectives, audiences and bidding systems. A two-year Meta coefficient compresses all of that into one much simpler representation.

Knowing what actually ran helps us understand what the historical coefficient represents and whether it says anything useful about the next dollar.

5. Incrementality

Some allocation decisions need stronger causal evidence.

A deliberate intervention can create variation that did not exist in the historical data. We change the treatment and compare what happened with a counterfactual.

That evidence can answer a specific allocation question directly. It can also challenge the causal story coming from the MMM.

6. MMM

MMM comes after those layers.

By then we know much more about the business the model is trying to explain. We have observed demand, listened to buyers, measured meaningful behaviour, understood what happened inside the channels and created causal evidence for some of the most important decisions.

There is less left for the MMM to infer from aggregate spend and revenue alone.

There is also more evidence available when the model produces a result that looks wrong. A large modelled brand effect with no movement in observable demand deserves scrutiny. A response curve that conflicts with an experiment deserves scrutiny. A stable historical coefficient covering radically different periods of creative and execution deserves scrutiny.

The MMM becomes one source of evidence inside the system rather than the judge of all the others.

Start with what you can observe

My discomfort with MMM comes back to a simple question: how much are we really letting the data speak?

The answer is not to remove every assumption from the model. We need structure to interpret observational data.

The practical response is to reduce how much we ask one model to infer.

Observe customer demand where you can. Listen to buyers. Measure meaningful behaviour. Understand what happened inside the media investment. Create causal information through experiments when the decision justifies it.

Then model what remains.

Observed evidence first. Designed causal evidence next. Modelled evidence last.

As we move through those layers, we ask the measurement method to make more inference. The evidence from the earlier layers gives us something independent against which to judge the later ones.

MMM asks the data a causal question after we have already defined much of the causal world in which the answer can exist.

When the marketing signal is strong, the observations can constrain those decisions. As the signal gets weaker relative to the noise, the modelling choices have more influence over the decomposition.

We should know which situation we are in before moving the budget.

MMM should inform that decision. It should never be the only evidence behind it.

Previous
Previous

How to Measure Brand Marketing Effectiveness Before Revenue Shows Up

Next
Next

The Double Jeopardy of Performance Marketing: Why Smaller Brands May Pay More to Compete