How far the held-out points actually sit from the training data
Source:R/fold-separation.R
fold_separation.RdBlocked cross-validation exists to put distance between a test point and
the training points that could tell you its value. Whether it succeeded is
a measurement, and until now the package only offered a proxy for it: the
block size compared against the estimated autocorrelation range, reported
by make_folds() as a warning. That comparison is about the
design. This is about the result: for every held-out point,
the distance to its nearest training point, summarised per fold.
Arguments
- folds
A
make_folds()result, or its$foldselement (a list oftrain/testsplits).- data_sf
The layer the folds were built on. Row identifiers are matched through
..row_idwhen the layer carries one, and by row position otherwise, which is whatmake_folds()and everycv_*()do. As incv_*(), amake_folds()result whose recorded rows sit at other locations here (folds built on another layer, such as the points beforeassign_features_to_polygons()dropped some) is refused; the location check is skipped when one of the two layers is POINT and the other is not.- sac
Optional: an
estimate_sac_range()result or a single number. Defaults to the range the folds carry, if any. Supplying one adds thewithin_rangecolumn and the closing verdict. A range that records its CRS (the folds' own, or anestimate_sac_range()result) is compared with distances measured in that CRS, whatever CRSdata_sfis in. A bare number is taken to be in the units the distances are otherwise measured in: those ofdata_sfif it is projected, and for geographic (lon/lat) input metres, in the CRSensure_projected()chooses (as formake_folds()'sblock_size), not degrees. Aunitsobject is refused.
Value
A data.frame of class fold_separation, one row per fold:
fold (the fold's number: for the $folds of a
cv_*() result, the fold_id its fold_metrics use,
which differs from the list position once a fold has been dropped),
n_train, n_test, n_blocks (NA
for a scheme with no blocks), min_dist and median_dist
(distance from a held-out point to its nearest training point, in the
units of the CRS the crs attribute names), and
within_range (the share of held-out points closer to
training data than sac; NA without one). Attributes:
method, sac_range, crs (the CRS the distances were
measured in) and n_unknown_ids.
Details
The two can disagree, and the direction is not obvious. Blocks wider than the range still leak wherever a test point sits near a block edge with training data just across it, which is most of the points in a fine block grid; conversely a fold whose blocks are narrower than the range can still separate well if the points inside them are clustered. The share of held-out points closer to training data than the correlation range is the number that settles it, and it is the last column here.
Nothing is estimated: the distances come from the geometry, and the range,
when one is shown, is the one the folds already carry (make_folds(
auto_range = TRUE) records it) or the one you pass as sac.
See also
make_folds() for the fold schemes and the block
sizing this measures the outcome of; cv_block_size_sweep()
for choosing a block size by cross-validated error instead.
Other cross-validation:
area_of_applicability(),
cv_bayes(),
cv_block_size_sweep(),
cv_gwr(),
cv_rf(),
cv_spatial(),
estimate_sac_range(),
gwr_model_selection(),
make_folds(),
sac_nugget(),
select_features_forward()
Examples
library(sf)
set.seed(1)
n <- 200
pts <- st_as_sf(
data.frame(x = 5e5 + runif(n, 0, 1000), y = 5e6 + runif(n, 0, 1000)),
coords = c("x", "y"), crs = 32632
)
# Random folds put a training point almost on top of every held-out one.
random <- make_folds(pts, k = 4, method = "random_kfold", seed = 1)
print(fold_separation(random, pts, sac = 200))
#> Fold separation: random_kfold, 4 fold(s), 200 held-out point(s) (EPSG:32632)
#> autocorrelation range: 200 (in metres, like the distances)
#>
#> fold n_train n_test min_dist median_dist within_range
#> 1 150 50 4.696 36.95 100%
#> 2 150 50 3.835 42.14 100%
#> 3 150 50 3.835 38.52 100%
#> 4 150 50 4.696 45.79 100%
#>
#> 100% of held-out points sit closer to a training point than the
#> correlation range (200), and the closest is 3.83 away. Most of the
#> hold-out is inside the range of its own training data, so this score is
#> optimistic: use blocked or buffered folds.
# Blocked folds hold out whole neighbourhoods, so the distances grow.
blocked <- make_folds(pts, k = 4, method = "block_kfold",
block_size = 250, seed = 1)
print(fold_separation(blocked, pts, sac = 200))
#> Fold separation: block_kfold, 4 fold(s), 200 held-out point(s) (EPSG:32632)
#> autocorrelation range: 200 (in metres, like the distances)
#>
#> fold n_train n_test n_blocks min_dist median_dist within_range
#> 1 152 48 2 23.02 111.9 100%
#> 2 143 57 3 29.04 169.5 67%
#> 3 152 48 2 21.03 118.3 96%
#> 4 153 47 2 21.03 115.9 85%
#>
#> 86% of held-out points sit closer to a training point than the
#> correlation range (200), and the closest is 21 away. Most of the
#> hold-out is inside the range of its own training data, so this score is
#> optimistic: widen the blocks.