Skip to contents

Blocked cross-validation exists to put distance between a test point and the training points that could tell you its value. Whether it succeeded is a measurement, and until now the package only offered a proxy for it: the block size compared against the estimated autocorrelation range, reported by make_folds() as a warning. That comparison is about the design. This is about the result: for every held-out point, the distance to its nearest training point, summarised per fold.

Usage

fold_separation(folds, data_sf, sac = NULL)

Arguments

folds

A make_folds() result, or its $folds element (a list of train/test splits).

data_sf

The layer the folds were built on. Row identifiers are matched through ..row_id when the layer carries one, and by row position otherwise, which is what make_folds() and every cv_*() do. As in cv_*(), a make_folds() result whose recorded rows sit at other locations here (folds built on another layer, such as the points before assign_features_to_polygons() dropped some) is refused; the location check is skipped when one of the two layers is POINT and the other is not.

sac

Optional: an estimate_sac_range() result or a single number. Defaults to the range the folds carry, if any. Supplying one adds the within_range column and the closing verdict. A range that records its CRS (the folds' own, or an estimate_sac_range() result) is compared with distances measured in that CRS, whatever CRS data_sf is in. A bare number is taken to be in the units the distances are otherwise measured in: those of data_sf if it is projected, and for geographic (lon/lat) input metres, in the CRS ensure_projected() chooses (as for make_folds()'s block_size), not degrees. A units object is refused.

Value

A data.frame of class fold_separation, one row per fold: fold (the fold's number: for the $folds of a cv_*() result, the fold_id its fold_metrics use, which differs from the list position once a fold has been dropped), n_train, n_test, n_blocks (NA for a scheme with no blocks), min_dist and median_dist (distance from a held-out point to its nearest training point, in the units of the CRS the crs attribute names), and within_range (the share of held-out points closer to training data than sac; NA without one). Attributes: method, sac_range, crs (the CRS the distances were measured in) and n_unknown_ids.

Details

The two can disagree, and the direction is not obvious. Blocks wider than the range still leak wherever a test point sits near a block edge with training data just across it, which is most of the points in a fine block grid; conversely a fold whose blocks are narrower than the range can still separate well if the points inside them are clustered. The share of held-out points closer to training data than the correlation range is the number that settles it, and it is the last column here.

Nothing is estimated: the distances come from the geometry, and the range, when one is shown, is the one the folds already carry (make_folds( auto_range = TRUE) records it) or the one you pass as sac.

See also

make_folds() for the fold schemes and the block sizing this measures the outcome of; cv_block_size_sweep() for choosing a block size by cross-validated error instead.

Other cross-validation: area_of_applicability(), cv_bayes(), cv_block_size_sweep(), cv_gwr(), cv_rf(), cv_spatial(), estimate_sac_range(), gwr_model_selection(), make_folds(), sac_nugget(), select_features_forward()

Examples

library(sf)
set.seed(1)
n <- 200
pts <- st_as_sf(
  data.frame(x = 5e5 + runif(n, 0, 1000), y = 5e6 + runif(n, 0, 1000)),
  coords = c("x", "y"), crs = 32632
)

# Random folds put a training point almost on top of every held-out one.
random  <- make_folds(pts, k = 4, method = "random_kfold", seed = 1)
print(fold_separation(random, pts, sac = 200))
#> Fold separation: random_kfold, 4 fold(s), 200 held-out point(s) (EPSG:32632)
#>   autocorrelation range: 200 (in metres, like the distances)
#> 
#>  fold n_train n_test min_dist median_dist within_range
#>     1     150     50    4.696       36.95         100%
#>     2     150     50    3.835       42.14         100%
#>     3     150     50    3.835       38.52         100%
#>     4     150     50    4.696       45.79         100%
#> 
#>   100% of held-out points sit closer to a training point than the
#>   correlation range (200), and the closest is 3.83 away. Most of the
#>   hold-out is inside the range of its own training data, so this score is
#>   optimistic: use blocked or buffered folds.

# Blocked folds hold out whole neighbourhoods, so the distances grow.
blocked <- make_folds(pts, k = 4, method = "block_kfold",
                      block_size = 250, seed = 1)
print(fold_separation(blocked, pts, sac = 200))
#> Fold separation: block_kfold, 4 fold(s), 200 held-out point(s) (EPSG:32632)
#>   autocorrelation range: 200 (in metres, like the distances)
#> 
#>  fold n_train n_test n_blocks min_dist median_dist within_range
#>     1     152     48        2    23.02       111.9         100%
#>     2     143     57        3    29.04       169.5          67%
#>     3     152     48        2    21.03       118.3          96%
#>     4     153     47        2    21.03       115.9          85%
#> 
#>   86% of held-out points sit closer to a training point than the
#>   correlation range (200), and the closest is 21 away. Most of the
#>   hold-out is inside the range of its own training data, so this score is
#>   optimistic: widen the blocks.