mediamix: media transforms and attribution path construction
Source:R/mediamix-package.R
mediamix-package.RdFeature engineering for marketing mix modelling and multi-touch attribution. mediamix turns raw marketing data into model-ready features: media spend into carryover- and saturation-adjusted regressors, and raw event logs into attributed customer journeys.
Details
It is a preprocessing and reporting package. It does not fit models –
lm(), glmnet and brms do that better than a marketing package would.
It computes order-1 Markov removal effects natively; for higher-order chains
as_channel_paths() hands your journeys to ChannelAttribution.
Full documentation, an end-to-end walkthrough and the methods behind each function are at https://elkronos.github.io/mediamix/.
Where to start
Five vignettes, in the order most people need them:
vignette("mediamix")Getting started: raw weekly spend through diagnostics, transforms and a model to contributions and ROI.
vignette("journeys")Building journeys from event logs, and what a naive
group_by()andpaste()gets wrong.vignette("carryover")Choosing decay and lag by cross-validation against your KPI rather than by eye.
vignette("tidymodels")Tuning carryover, saturation and model penalty jointly in one search.
vignette("spine")Why carryover and credit are the same idea.
The media half
- Carryover
adstock_geometric()for the standard geometric kernel,adstock_weibull()andadstock_delayed()when the response peaks after the spend,adstock_filter()for a kernel of your own.adstock_weights()andadstock_weights_weibull()expose the kernels themselves, andadstock_state()carries a filter across a boundary.- Saturation
saturate_hill(),saturate_exponential(),saturate_michaelis_menten(),saturate_power(), and thesaturate()dispatcher.- Composing them
media_transform()applies carryover then saturation, in that order, and will not do it backwards quietly.- Choosing parameters
decay_from_half_life(),half_life()andeffective_window()translate between half-lives and decay coefficients.tune_carryover()andtune_carryover_joint()select them by cross-validation against an actual KPI;adstock_steady_state()removes start-up bias.
The attribution half
- Journeys
build_paths()turns an event log into journeys, handling the seven things that go wrong on the way.- Credit
credit_linear(),credit_first(),credit_last(),credit_position(),credit_time_decay()andcredit_custom(), and the data-drivenmarkov_removal().attribute()runs several at once andattribution_spread()reports how much the answer depends on which you picked.- Diagnostics
path_summary(),path_lengths(),channel_positions(),assisted_conversions(),top_paths()andconversion_lag(), orpath_diagnostics()for all of them.- Interop
as_channel_paths()exports to ChannelAttribution's format.
Reporting, diagnostics and tidymodels
diagnose_media() checks whether the data can support a model at all –
run it first. contributions(), roi(), marginal_roi(), mroi(),
response_curve() and spend_for() turn a fitted model into a
deliverable, and block_bootstrap() puts intervals on it. Each result has
a plot() method; mm_palette() exposes their colours. step_adstock() and
step_saturation() are recipes steps that carry filter state across
the train/test boundary, with carryover_decay() and friends as their
dials parameters.
Data
mm_weekly is a synthetic weekly panel generated from known parameters, so a workflow can be checked against the truth it is trying to recover. mm_events is a synthetic touchpoint log containing, on purpose, every case that breaks a naive journey pipeline.
Dependencies
The core transforms need only base R, stats, and cli for error messages.
data.table is used by the attribution half, where event logs are large.
recipes, dials and the rest of tidymodels are optional and
registered conditionally, so the transforms work equally well from a Stan or
brms workflow that will never touch them.
Author
Maintainer: Justin Chase jchase.msu@gmail.com
Examples
# The media half: spend in, model-ready regressor out
spend <- c(0, 500, 800, 300, 0, 0, 1200, 400)
media_transform(
spend,
adstock = list(decay = decay_from_half_life(2)),
saturation = list(half_max = 300)
)
#> [1] 0.0000000 0.3280272 0.5296832 0.5213606 0.4350985 0.3525949 0.6088682
#> [8] 0.5985976
# The attribution half: event log in, journeys out
data(mm_events)
paths <- build_paths(mm_events, id = "customer_id", channel = "channel",
timestamp = "timestamp", conversion = "conversion")
path_summary(paths)
#> events_in events_out journeys converting_journeys mean_length median_length
#> 1 20799 17306 6137 2741 2.819945 3
#> max_length direct_events conversion_events conversions_observed
#> 1 9 1597 2741 2741
#> conversions_unattributable
#> 1 0