Skip to contents

A blocked cross-validation with blocks smaller than the autocorrelation range leaks: every held-out point has a near-identical neighbour in the training set, and the score is optimistic in proportion. The single number a cv_*() call returns cannot show this. This runs the same cross-validation at a ladder of block sizes, and by default once with random folds as the fully leaky reference. It returns the metric at each, with the estimated autocorrelation range alongside so that the curve can be read against it: it rises as the blocks pass the range and then plateaus, and the height of the rise is how much the random-fold number overstated the model.

Usage

cv_block_size_sweep(
  data_sf,
  response_var,
  predictor_vars,
  fit_fn,
  block_sizes = NULL,
  n_sizes = 6L,
  k = 5L,
  metric = "RMSE",
  include_random = TRUE,
  max_fits = 60L,
  sac = NULL,
  seed = 123L,
  quiet = FALSE,
  ...
)

Arguments

data_sf

An sf object with the response and predictors.

response_var, predictor_vars

Column names.

fit_fn

A function(train_sf) returning a spatial_fit, as for cv_spatial(); see the example for wrapping a built-in backend.

block_sizes

Optional numeric vector of block edge lengths to sweep, in the CRS units the folds are built in (plain numbers; a units object is refused). Default NULL: the ladder described above. A size at which the grid would hold fewer than k blocks is not run, with a warning naming it and the largest size that still gives k blocks; if no size is left, the call is an error.

n_sizes

Number of sizes in the default ladder. Default 6.

k

Folds per cross-validation. Default 5.

metric

Which column of overall to read. Default "RMSE".

include_random

Also cross-validate with random folds, as the leaky reference. Default TRUE.

max_fits

The fit budget; see above. Default 60.

sac

Optional sac_range from estimate_sac_range() to mark on the curve, or a single number in the units of the CRS the folds are built in. An estimate_sac_range() result records the CRS it was fitted in; when that is not the sweep's, the range is converted to the sweep's units, with a warning. Default NULL: estimated here from the response, detrended on predictor_vars, when gstat is installed.

seed

Seed for the fold construction at every size.

quiet

Suppress the progress messages. Default FALSE.

...

Passed to cv_spatial() at every size (predict_args, p, parallel, metrics, ...). block_size, folds, k, seed and auto_range are set here and cannot be passed.

Value

A data.frame of class "block_size_sweep" with one row per cross-validation: block_size (NA for the random reference), method, blocks_used, k (the folds actually built), n_folds_succeeded, value (the pooled metric), fold_min, fold_max and fold_sd (its spread across folds). Attributes: metric, sac_range (the effective range, or NA), crs, k (the folds requested, which print() and plot() report), n_fits, response_var, and results, the full cv_spatial() result at every size. plot() draws it.

The fit budget

Each block size is a full cross-validation, so the cost is length(block_sizes) * k fits, plus k for the random reference. max_fits caps that (default 60: room for up to eleven sizes at k = 5 plus the reference; the default six-size ladder at k = 5 needs at most 35). A sweep that would run past the cap refuses to start, naming the number of fits it would have needed. Raise max_fits deliberately; the RF example below takes seconds, a Bayesian fit_fn takes minutes per fit.

The ladder

When block_sizes is NULL, n_sizes values are log-spaced from a twenty-fifth of the shorter side of the data's extent (of the longer side when the points lie on one line parallel to an axis) to the largest size, at most half that side, at which the grid still holds k blocks. Half the side gives a grid two blocks across, enough for k up to 4 and, for larger k, on an extent long enough in the other direction; on a roughly square extent at the default k = 5 the top is about a third of the side. Any size at which the grid would hold more than the 1,000,000 make_folds() will build is dropped. The count is of grid cells: on clustered data a grid of k or more cells can have fewer than k that hold points, and make_folds() then lowers k at that size, which the k column shows. On an extent much longer than it is wide every rung can fall below the autocorrelation range while longer blocks would still fit k times along the longer side; the sweep warns when that happens, and block_sizes is then the way to reach past the range. Sizes are in the units of the CRS the folds are built in (make_folds()'s params$crs, metres for geographic input), and the returned table records that CRS. Each is the minimum block edge handed to make_folds(), which fits a whole number of cells across the extent, so the cells are somewhat longer than the size on the axis (up to twice as long).

Examples

if (requireNamespace("ranger", quietly = TRUE) &&
    requireNamespace("gstat", quietly = TRUE)) {
  library(sf)
  set.seed(1)
  n <- 200
  x <- 5e5 + runif(n, 0, 1000); y <- 5e6 + runif(n, 0, 1000)
  d <- as.matrix(dist(cbind(x, y)))
  field <- as.numeric(t(chol(exp(-d / 100) + diag(1e-6, n))) %*% rnorm(n))
  dat <- st_as_sf(data.frame(x = x, y = y, a = rnorm(n)), coords = c("x", "y"),
                  crs = 32632)
  dat$z <- field + 0.5 * dat$a + rnorm(n, 0, 0.2)
  rf_fn <- function(train_sf)
    fit_rf_model(train_sf, "z", "a", include_coords = TRUE, num_trees = 100, seed = 1)
  sw <- cv_block_size_sweep(dat, "z", "a", fit_fn = rf_fn, k = 4, n_sizes = 4,
                            quiet = TRUE)
  print(sw)               # the curve as a table, with the random-fold reference
  if (requireNamespace("ggplot2", quietly = TRUE)) plot(sw)
}
#> Cross-validation RMSE against block size (20 fits, k = 4, sizes in EPSG:32632)
#>   estimated autocorrelation range: 452.5
#>  block_size       method blocks_used k n_folds_succeeded  value fold_min
#>      random random_kfold          NA 4                 4 0.9639   0.8346
#>       38.73  block_kfold         172 4                 4 0.9438   0.7825
#>       89.89  block_kfold          87 4                 4 1.0380   0.8706
#>      208.60  block_kfold          16 4                 4 1.1289   1.0262
#>      484.10  block_kfold           4 4                 4 1.2438   0.9707
#>  fold_max
#>     1.146
#>     1.079
#>     1.117
#>     1.233
#>     1.540