Thin wrapper over cv_spatial that refits a ranger
forest on each training fold. Unlike the out-of-bag error, this holds out
whole spatial blocks, so neighbours of a held-out point are not sitting in
the training set.
Usage
cv_rf(
data_sf,
response_var,
predictor_vars,
folds = NULL,
k = 5,
seed = 123,
parallel = FALSE,
block_size = NULL,
auto_range = FALSE,
boundary = NULL,
pointize = "auto",
metrics = NULL,
...
)Arguments
- data_sf
An sf object.
- response_var
Response column name.
- predictor_vars
Predictor column names.
- folds
Optional fold definitions, in any of three shapes: a
make_folds()return value; a bare list oflist(train =, test =)pairs of..row_idvalues; or a vector of fold labels, one per row, which becomes leave-that-label-out splits. The label vector is how folds built by another package are used here, sinceblockCV::cv_spatial()'s$folds_listholds two unnamed vectors per fold and is refused by name. That function returns a label vector as$folds_ids. Train and test must be disjoint: a fold that trains on its own test rows is not a cross-validation split and is refused with an error. IDs naming no row in the prepared data are dropped with a logged count.- k
Number of folds when
foldsisNULL. Default 5.- seed
RNG seed. Default 123. It seeds fold construction and, through a per-fold draw, each fold's forest, so two different seeds give different results even on identical
folds. To grow every fold's forest from one fixed ranger seed instead, callcv_spatialwith your ownfit_fnwrappingfit_rf_model(seed = ).- parallel
Passed to
cv_spatial. DefaultFALSE. Under forked workers each fold's forest runs single-threaded unlessnum_threadsis passed explicitly, soparallel = 4means four threads in total, where it would otherwise mean four times the session'smc.cores.- block_size, auto_range, boundary
Passed to
cv_spatial.- pointize
How non-POINT geometry is reduced to a point before fitting; passed to
cv_spatial. Default"auto".- metrics
Optional scoring function of your own, passed to
cv_spatial; see Your own metrics there.- ...
Passed to
fit_rf_modelon every fold:num_trees,mtry,importance,include_coordsand so on.data_sf,response_var,predictor_varsand.already_preppedare set by this function and must not be passed here (every fold would fail with "matched by multiple actual arguments").seedis this function's own argument and never reachesfit_rf_model()through here; seeseedabove for growing every fold's forest from one fixed seed.
Value
The cv_spatial result.
Percentage errors on responses with zeros
MAPE divides by the observed value and SMAPE by
\(|y| + |\hat{y}|\), so neither is defined where its denominator is zero.
Neither returns Inf or NaN. Both are averaged over the rows
whose denominator is non-zero, and are NA when no row qualifies.
Non-zero is judged at the scale of the data: a denominator no larger
than 100 machine epsilons times the largest one counts as zero, so the
rule does not depend on the units of the response.
The n_MAPE and n_SMAPE columns record how many rows that was;
the n column counts finite observation/prediction pairs. Read a
percentage error next to its count: when n_MAPE < n, MAPE is
an average over a subset of the data, whatever its value.
This bites on any response taking exact zeros: counts, rainfall,
abundance, claim amounts. On a zero-inflated response with 62 zeros out of
120, MAPE is an average over the 58 non-zero rows, which
n_MAPE = 58 now says. SMAPE fails differently and more
subtly: it drops the rows where observation and prediction are both near
zero (which on a well-fitted zero-inflated model are the rows it got
right), so it averages the harder rows only and reads worse than
the fit deserves; n_SMAPE shows how many rows it kept, and the
count is only a label, not a repair.
RMSE, MAE and \(R^2\) use every finite row and are
unaffected; prefer them whenever the response can be zero. For a Bayesian
fit, cv_bayes() additionally reports CRPS and interval
coverage, which are proper scoring rules and have no such failure mode.
Examples
if (requireNamespace("ranger", quietly = TRUE)) {
library(sf)
set.seed(1)
n <- 200
dat <- st_as_sf(
data.frame(x = 5e5 + runif(n, 0, 1000), y = 5e6 + runif(n, 0, 1000),
a = rnorm(n)),
coords = c("x", "y"), crs = 32632
)
dat$z <- 2 * dat$a + rnorm(n, 0, 0.3)
cv_rf(dat, "z", "a", k = 4)$overall
}
#> cv_rf(): no folds supplied -- using spatial block k-fold CV (k=4).
#> RMSE MAE MAPE SMAPE R2 Adj_R2 n_pred n_MAPE n_SMAPE
#> 1 0.417048 0.3264843 83.10643 42.3203 0.9587525 NA 200 200 200