Extending the test window doesn't make incrementality fit to measure long-term effects and I'll show you why

incrementality-testing-long-term-effects
I built two scenarios with the same impact: $5,000 in revenue.

In the first, the impact shows up in the first 7 days.
In the second, it shows up gradually over 60 days.

Using a Difference-in-Difference test:

In scenario 1, we detected $4,900 of the $5,000.
In scenario 2, we detected nothing. The p-value was 0.26 which means "inconclusive".

Even though we know the lift was there: we added it ourselves.

𝐓𝐡𝐞 𝐩𝐫𝐨𝐛𝐥𝐞𝐦 𝐢𝐬 𝐧𝐨𝐢𝐬𝐞. And it hits in two ways.

𝟏. 𝐓𝐢𝐦𝐞 𝐝𝐢𝐥𝐮𝐭𝐢𝐨𝐧. The longer the buying cycle, the more the signal spreads across days, the harder it is to separate from the noise.

𝟐. 𝐑𝐞𝐥𝐚𝐭𝐢𝐯𝐞 𝐬𝐢𝐳𝐞. Smaller channels sit under the noise created by the bigger ones. A channel at 10% of your mix has to drive ridiculously high ROI just to be visible.

Brand initiatives usually get both. Long payback and a smaller share of spend. So they get penalised twice, even when they're working.

So what works?

For short-term effects, lift tests are great. Use them.

For long-term effects, lift tests alone aren't enough. We layer them with:

- Qualitative data (surveys, self-attribution). People remember and tell you about effects long after a revenue-target lift test is capable to trace the signal.
- Website traffic and behavioural signals.
- Share of search.


NB: Still better than attribution though...

Related reading

Previous
Previous

Finance Already Ran This Experiment

Next
Next

The Financial Limits of Performance Marketing