A blocked cross-validation with blocks smaller than the autocorrelation
range leaks: every held-out point has a near-identical neighbour in the
training set, and the score is optimistic in proportion. The single
number a cv_*() call returns cannot show this. This runs the same
cross-validation at a ladder of block sizes, and by default once with
random folds as the fully leaky reference. It returns the metric at each,
with the estimated autocorrelation range alongside so that the curve can be
read against it: it rises as the blocks pass the range and then plateaus,
and the height of the rise is how much the random-fold number overstated
the model.
Usage
cv_block_size_sweep(
data_sf,
response_var,
predictor_vars,
fit_fn,
block_sizes = NULL,
n_sizes = 6L,
k = 5L,
metric = "RMSE",
include_random = TRUE,
max_fits = 60L,
sac = NULL,
seed = 123L,
quiet = FALSE,
...
)Arguments
- data_sf
An sf object with the response and predictors.
- response_var, predictor_vars
Column names.
- fit_fn
A
function(train_sf)returning aspatial_fit, as forcv_spatial(); see the example for wrapping a built-in backend.- block_sizes
Optional numeric vector of block edge lengths to sweep, in the CRS units the folds are built in (plain numbers; a
unitsobject is refused). DefaultNULL: the ladder described above. A size at which the grid would hold fewer thankblocks is not run, with a warning naming it and the largest size that still giveskblocks; if no size is left, the call is an error.- n_sizes
Number of sizes in the default ladder. Default 6.
- k
Folds per cross-validation. Default 5.
- metric
Which column of
overallto read. Default"RMSE".- include_random
Also cross-validate with random folds, as the leaky reference. Default
TRUE.- max_fits
The fit budget; see above. Default 60.
- sac
Optional
sac_rangefromestimate_sac_range()to mark on the curve, or a single number in the units of the CRS the folds are built in. Anestimate_sac_range()result records the CRS it was fitted in; when that is not the sweep's, the range is converted to the sweep's units, with a warning. DefaultNULL: estimated here from the response, detrended onpredictor_vars, when gstat is installed.- seed
Seed for the fold construction at every size.
- quiet
Suppress the progress messages. Default
FALSE.- ...
Passed to
cv_spatial()at every size (predict_args,p,parallel,metrics, ...).block_size,folds,k,seedandauto_rangeare set here and cannot be passed.
Value
A data.frame of class "block_size_sweep" with one row per
cross-validation: block_size (NA for the random
reference), method, blocks_used, k (the folds
actually built), n_folds_succeeded, value (the pooled
metric), fold_min, fold_max and fold_sd (its spread
across folds). Attributes: metric, sac_range (the
effective range, or NA), crs, k (the folds
requested, which print() and plot() report),
n_fits, response_var, and results, the full
cv_spatial() result at every size.
plot() draws it.
The fit budget
Each block size is a full cross-validation, so the cost is
length(block_sizes) * k fits, plus k for the random
reference. max_fits caps that (default 60: room for up to eleven
sizes at k = 5 plus the reference; the default six-size ladder at
k = 5 needs at most 35). A sweep that would run past the cap
refuses to start, naming the number of fits it would have needed. Raise
max_fits deliberately; the RF example below takes seconds, a Bayesian
fit_fn takes minutes per fit.
The ladder
When block_sizes is NULL, n_sizes values are
log-spaced from a twenty-fifth of the shorter side of the data's extent
(of the longer side when the points lie on one line parallel to an axis)
to the largest size, at most half that side, at which the grid still
holds k blocks. Half the side gives a grid two blocks across,
enough for k up to 4 and, for larger k, on an extent long
enough in the other direction; on a roughly square extent at the default
k = 5 the top is about a third of the side. Any size at which the
grid would hold more than the 1,000,000 make_folds() will build is
dropped. The count is of grid cells: on clustered data a grid
of k or more cells can have fewer than k that hold points,
and make_folds() then lowers k at that size, which the
k column shows. On an extent much longer than it is wide every
rung can fall below the autocorrelation range while longer blocks would
still fit k times along the longer side; the sweep warns when
that happens, and block_sizes is then the way to reach past the
range. Sizes are in the units of the CRS the folds are built in
(make_folds()'s params$crs, metres for geographic input),
and the returned table records that CRS. Each is the minimum block
edge handed to make_folds(), which fits a whole number of cells
across the extent, so the cells are somewhat longer than the size on the
axis (up to twice as long).
Examples
if (requireNamespace("ranger", quietly = TRUE) &&
requireNamespace("gstat", quietly = TRUE)) {
library(sf)
set.seed(1)
n <- 200
x <- 5e5 + runif(n, 0, 1000); y <- 5e6 + runif(n, 0, 1000)
d <- as.matrix(dist(cbind(x, y)))
field <- as.numeric(t(chol(exp(-d / 100) + diag(1e-6, n))) %*% rnorm(n))
dat <- st_as_sf(data.frame(x = x, y = y, a = rnorm(n)), coords = c("x", "y"),
crs = 32632)
dat$z <- field + 0.5 * dat$a + rnorm(n, 0, 0.2)
rf_fn <- function(train_sf)
fit_rf_model(train_sf, "z", "a", include_coords = TRUE, num_trees = 100, seed = 1)
sw <- cv_block_size_sweep(dat, "z", "a", fit_fn = rf_fn, k = 4, n_sizes = 4,
quiet = TRUE)
print(sw) # the curve as a table, with the random-fold reference
if (requireNamespace("ggplot2", quietly = TRUE)) plot(sw)
}
#> Cross-validation RMSE against block size (20 fits, k = 4, sizes in EPSG:32632)
#> estimated autocorrelation range: 452.5
#> block_size method blocks_used k n_folds_succeeded value fold_min
#> random random_kfold NA 4 4 0.9639 0.8346
#> 38.73 block_kfold 172 4 4 0.9438 0.7825
#> 89.89 block_kfold 87 4 4 1.0380 0.8706
#> 208.60 block_kfold 16 4 4 1.1289 1.0262
#> 484.10 block_kfold 4 4 4 1.2438 0.9707
#> fold_max
#> 1.146
#> 1.079
#> 1.117
#> 1.233
#> 1.540