Runs the checks that decide whether a marketing mix model can work at all. None of them are about the model; all of them are about whether the data contains the information the model will be asked to find. Run this first.
Usage
diagnose_media(
data,
media,
spend = NULL,
impressions = NULL,
by = NULL,
vif_threshold = 5,
cv_threshold = 0.15,
decay = NULL
)Arguments
- data
A data frame containing the media columns.
- media
Character vector naming the media columns.
- spend, impressions
Optional character vectors of matching spend and impression columns, in the same order, for the cost-per-mille consistency check.
- by
Optional character vector of grouping columns. Every diagnostic is computed within each group and the worst case reported: the highest variance inflation factor, the strongest absolute correlation, the lowest coefficient of variation, the highest share of dark periods. Pooling geographies that differ mainly in scale would manufacture correlation that exists in no single series, so
byis worth supplying whenever the rows are a panel.- vif_threshold
Variance inflation factor above which a channel is flagged.
- cv_threshold
Coefficient of variation below which a channel is flagged as insufficiently varying.
- decay
Optional geometric decay for computing collinearity on adstocked media: a single number for every channel, or a named vector with one per channel (unnamed channels are left raw). Adstock smooths each series, and smoothed series are usually more correlated than the raw spend, so this is the collinearity the model will actually face. Rows must be in time order within each group. Variation and flighting are always reported on the raw spend.
Value
An object of class mm_diagnosis: a named list with elements
variationPer channel:
meanandsdaveraged across groups, the lowest groupcvanddistinctcount, the highest groupzero_share, and the flagslow_variation,heavily_flightedandusable.collinearityPer channel: the worst group
vif, themost_correlated_withpartner and itscorrelation, andidentified–TRUE,FALSE, orNAwhere the factor could not be computed at all.NAis notTRUE.correlationsPairwise correlation matrix. With
bysupplied, each cell is the group value with the largest magnitude, sign preserved.cpmImplied cost per mille for each row of
data(rowis the row number), with the series median and an outlier flag, whenspendandimpressionsare supplied; withby, the median and outlier rule are computed within each series and agroupcolumn is added.NULLotherwise.flagsCharacter vector of the problems found, in the order they should be dealt with. Empty when nothing was found.
n_obs,n_groups,grouped,adstockedRows examined, series examined, whether
bywas supplied, and whether collinearity was measured on adstocked media.
Details
Four things sink marketing mix projects, and all four are visible before a model is fitted.
Channel collinearity is the biggest practical problem in the field. When two channels move together – because they were planned together, which is usually the case – the model cannot tell their effects apart. It will still produce coefficients, and those coefficients will flip sign on small changes to the specification or the sample. A variance inflation factor above 5 deserves attention and above 10 means the split between those channels is not identified, however tight the overall fit looks.
Insufficient variation is the quieter version of the same problem. A channel spending nearly the same amount every week carries almost no information about what different amounts would do. No adstock or saturation transform can recover an effect that the data never varied enough to reveal.
Zero inflation and flighting matter because a channel that is dark most of the time has far less effective sample than its row count suggests, and because carryover across dark periods is where transform bugs hide.
Spend and impression inconsistency – a wildly varying implied cost per mille – almost always means a data-join error upstream rather than a real change in media pricing. It is worth catching before it becomes a modelling puzzle.
Examples
data(mm_weekly)
channels <- c("tv", "video", "search", "social", "display")
# Always pass `by` on panel data: pooling geographies that differ in scale
# invents correlation that is not in any single series.
d <- diagnose_media(mm_weekly, media = channels, by = "geo")
d
#>
#> ── Media diagnostics
#> 468 observations across 3 series, 5 channels
#> ✔ No structural problems found.
#>
#> Inspect $variation, $collinearity, $correlations.
d$variation
#> channel mean sd cv distinct zero_share low_variation
#> 1 tv 1700.5021 881.6773 0.5096251 131 0.1474359 FALSE
#> 2 video 336.2778 308.5805 0.8938791 84 0.4166667 FALSE
#> 3 search 487.1432 115.5527 0.2286063 118 0.0000000 FALSE
#> 4 social 251.1923 174.8592 0.6612980 101 0.2756410 FALSE
#> 5 display 188.0641 149.3027 0.7584458 88 0.3333333 FALSE
#> heavily_flighted usable
#> 1 FALSE TRUE
#> 2 FALSE TRUE
#> 3 FALSE TRUE
#> 4 FALSE TRUE
#> 5 FALSE TRUE
d$collinearity
#> channel vif most_correlated_with correlation identified
#> 1 video 1.098660 social 0.2246800 TRUE
#> 2 social 1.093837 video 0.2246800 TRUE
#> 3 tv 1.066918 video 0.1589877 TRUE
#> 4 search 1.052593 video -0.2063315 TRUE
#> 5 display 1.045907 social -0.1879395 TRUE
# Collinearity after adstocking is what the model actually sees
diagnose_media(mm_weekly, media = channels, by = "geo",
decay = attr(mm_weekly, "truth")$decay)$collinearity
#> channel vif most_correlated_with correlation identified
#> 1 tv 1.122759 video 0.3245009 TRUE
#> 2 video 1.122127 tv 0.3245009 TRUE
#> 3 social 1.111494 display 0.2368121 TRUE
#> 4 display 1.100370 social 0.2368121 TRUE
#> 5 search 1.028554 tv 0.1534401 TRUE
# A deliberately collinear pair is caught
fake <- mm_weekly[mm_weekly$geo == "north", ]
fake$twin <- fake$tv * 1.02 + 5
diagnose_media(fake, media = c("tv", "twin", "search"))$collinearity
#> channel vif most_correlated_with correlation identified
#> 1 tv Inf twin 1.00000000 FALSE
#> 2 twin Inf tv 1.00000000 FALSE
#> 3 search 1.000823 twin 0.02868335 TRUE