Marketing Measurement Data Layer: What You Need Before MMM or Incrementality

Most marketing measurement problems start before the model.

The spend is in a planning sheet. Revenue is in finance. Site behaviour is in analytics. CRM activity is in a separate system. Offline sales are somewhere else. Product taxonomy changes by team. IDs are missing, duplicated or inconsistent.

Then the business asks for MMM, incrementality, attribution or an AI measurement system. The model becomes the visible object, but the real problem is the data layer underneath it.

Disconnected spend, revenue, behaviour and identity layers create false precision downstream. The model can still produce a number. The number will reflect the gaps in the foundation.

What The Data Layer Has To Do

A marketing measurement data layer does not need to be complicated at the start. It needs to make the core commercial journey visible enough for decisions.

At minimum, the business should be able to connect campaign spend, user traffic, meaningful site actions, CRM signals where relevant, product-level revenue and geography or market context. For offline or sales-assisted journeys, the layer also needs a path from marketing activity to store, sales or CRM outcomes.

The goal is not perfect tracking. Perfect tracking does not exist, especially in long consideration journeys. The goal is to create a reliable operating base that lets the team ask better questions.

Can we see spend by campaign and time? Can we see revenue by product, channel, geography and customer type? Can we see which site actions are meaningful? Can we tie CRM events to accounts, leads or customer IDs where the business structure allows it? Can we separate clean traffic from suspicious traffic? Can we compare markets without mixing incompatible definitions?

If the answer is no, build the data foundation before the model.

The Minimum Inputs

Start with spend data at campaign level. Campaign-level spend is usually the lowest useful granularity for strategic decisions. It is detailed enough to diagnose allocation while avoiding a false sense of precision from every ad or placement.

Add commercial outcomes. Revenue is useful, but gross profit is often more useful because it accounts for margin. Product-level outcomes matter because a campaign can look average overall while creating valuable demand for one product group and weak demand for another.

Add website and app behaviour. Unbounced traffic, product views, pricing-page views, store-locator visits, add-to-carts, chat requests, demo starts and other meaningful actions help bridge the gap between exposure and final sale.

Add CRM and sales activity when the journey is sales-assisted. Emails, calls, sales chats, trials, opportunities, pipeline stages, lost reasons and competitor mentions can explain demand quality in a way click paths cannot.

This is not only a data completeness issue. It is also a trust issue. If an attribution report claims marketing created a deal while the sales notes show that the buyer came through previous experience, a referral, an event or a sales-led conversation, the measurement layer will lose credibility with the people closest to the customer. Sales evidence should not be treated as a footnote to the click path.

Add search data with the right split. Organic traffic should be separated into branded and non-branded queries where possible, and paid search should be connected to Search Console data. Otherwise the data layer can treat branded organic traffic as SEO-created demand, or treat paid search as the origin of demand that another channel created earlier.

Add identity and context. Customer ID, account ID, email where permitted, geography, product, time and campaign taxonomy are the joins that make the data useful.

Taxonomy Is Measurement Infrastructure

Many teams treat taxonomy as admin work. It is measurement infrastructure.

If campaign names, product categories, regions and funnel stages are inconsistent, the analysis will be inconsistent too. A model cannot know that two teams used different names for the same product, or that one market includes stores while another does not, unless that logic is defined.

A clean taxonomy lets the team split the funnel. Spend can be connected to unbounced traffic, product views, carts, sales activity and revenue. Once the funnel is split, correlations between neighbouring steps are often more useful than trying to jump straight from spend to final revenue.

This is especially important for high-ticket B2C and B2B. The longer the consideration journey, the weaker the direct relationship between media spend this week and revenue this week. A split funnel gives the team earlier and more stable reads.

Clean Traffic Before Valuing It

Traffic quality belongs in the data layer because bad traffic can corrupt every downstream method.

Pulling raw site traffic straight into a test or model is risky. Bot traffic, low-quality placements, strange geography, weak engagement and odd device behaviour can make a campaign look active while adding little commercial value. If those sessions are counted as demand, the model learns the wrong thing.

The data layer should flag obvious traffic problems before the measurement method sees them. Useful checks include engagement rate, time on site, page depth, product-page behaviour, add-to-cart rates, geography quality, device patterns and halo effects across channels.

Cleaning traffic first gives it the right commercial meaning.

Why MMM And Incrementality Need The Same Foundation

MMM needs stable historical spend, outcome and context data. Incrementality needs clean test inputs, clear treatment and control definitions, reliable outcome measurement and enough visibility to catch leakage.

Both methods break when the data layer is weak. MMM can fit noise from disconnected or misclassified inputs. Incrementality can misread a market split if treatment and control differ in media delivery, traffic quality, product mix or data capture.

The same foundation also makes the methods work together. Mid-funnel evidence can explain why a campaign moved product interest before revenue. Incrementality can test whether a high-ROAS channel actually caused growth. MMM can use cleaner, better-structured inputs to support broader allocation.

Without the data layer, the business debates model outputs. With the data layer, the business can inspect the assumptions behind the output.

What AI Changes And What It Does Not

AI can help write code, inspect assumptions, document validation logic, translate technical results and accelerate modelling work. It cannot recover a customer ID that was never captured. It cannot join systems that do not share a key. It cannot make a bad taxonomy consistent after the fact without business rules.

This is the practical AI boundary. The cost of analytical work can fall, but the value of clean data plumbing rises.

The companies that benefit most from AI measurement will not be the ones with the flashiest interface. They will be the ones whose spend, revenue, behaviour and customer context are already organised well enough for AI to reason on top of them.

The Decision Rule

Before buying MMM, running a major incrementality test or building an AI measurement layer, ask for one table or view that ties the commercial system together.

It should show spend, traffic quality, meaningful behaviour, CRM signals where relevant, product-level revenue, gross profit, geography and the IDs or rules that connect them. It does not need to be perfect. It needs to be good enough that the team understands what is visible, what is missing and what assumptions any model will have to make.

If that view does not exist, build the data layer first. The measurement method comes after the business can see what it is measuring.

Previous
Previous

Retargeting Incrementality: How Much of Your Retargeting Budget Is Actually Creating Sales?

Next
Next

How to Measure a Price Change Without an A/B Test