Greedy forward feature selection with spatially blocked inner folds
Source:R/feature-selection.R
select_features_forward.RdSelects predictors by repeatedly adding whichever candidate most improves a
cross-validated score, stopping when no candidate improves it by more than
tol.
Arguments
- train_sf
Training data (
sf).- response_var
Character(1).
- candidate_vars
Character vector of predictors to choose among.
- fit_fn
A function
(train_sf, predictor_vars)returning aspatial_fit. The signature takes two arguments because selection has to refit with different predictor sets.- k
Inner fold count. Default 5.
- method
Inner fold method. Default
"block_kfold".- block_size
Passed to
make_folds(); inherit the outer block size so inner and outer blocks are on the same spatial scale.- metric
Score to optimise:
"RMSE","MAE"(minimised) or"R2"(maximised). Default"RMSE".- tol
Minimum improvement required to accept a variable. Default 0, meaning any improvement is accepted. The first variable is judged against the null (intercept-only) model, so
tolbites from step 1, but only when that null model can be scored. Backends that refuse a zero-lengthpredictor_vars(fit_rf_modelandfit_gwr_modelboth do) have no null score, and there the first variable is accepted unconditionally.- max_vars
Optional cap on how many predictors to select.
- max_fits
Abort if the sweep would exceed this many model fits. Default 5000.
- seed
RNG seed. It governs both the inner fold construction and the cross-validation itself: it is forwarded to
cv_spatial(seed = ), which draws one RNG stream per fold from it, so it also seeds the learner inside every fold. A stochasticfit_fnis therefore reproducible from this one value.- quiet
Logical; suppress this function's progress
message()s. It does not silence R warnings, nor the package's console log echo (seespatialkit_quietfor that). DefaultFALSE.- auto_range
Logical. If
TRUEandmethodis"block_kfold", the autocorrelation range is estimated from the response (detrended on the candidates) and used as the minimum block size of the inner folds, exactly as inmake_folds();block_sizestill applies as a floor. DefaultFALSE, which keeps the geometric blocks. Inner blocks smaller than the range let a candidate be selected for spatial proximity to the response rather than for predicting it, which is the same failure therandom_kfoldcaution above exists to prevent, so the leakage warningmake_folds()raises applies here with more force than usual.- select_on
"all"(default) runs the sweep on every row oftrain_sf."split"runs it on one spatially blocked half, then fits the selected set on that half and scores it on the other:score_holdoutis then the selected model'smetricon rows whose response the sweep never read (the sweep's ownscoreis not; see "The score is not a performance estimate"). It is one estimate from one region: the halves share a border with no buffer, so rows near it are still correlated with the selection half. Both halves come back in$split. See the "Post-selection inference" section ofdetermine_optimal_levelsfor the trade: coverage for half the sample.
Value
A list of class "feature_selection" (so that
plot.feature_selection() draws the selection path) with
selected (the chosen predictors, in the order they were added),
score, score_holdout, history, params and
split.
score is the winning set's cross-validated metric at the
final step: the selection-internal optimum, optimistically
biased because it was chosen as the best of many (see the section above),
and NA when nothing was selected. history is a data.frame
with step, variable, score and n_pred,
holding every candidate evaluated at every step; when the null model
could be scored it also carries a step = 0 row named
"<none>" giving that baseline, so the first variable's gain can be
read off directly. Every set is scored on the same rows: those the null
model's cross-validation predicted or, when there is no null model,
those any step-1 set predicted; params$n_scored counts them (a
warning says so when that is fewer than all). n_pred is how many
rows the set's cross-validation predicted. A set that left some of the
scored rows unpredicted, because a fold failed for it, has score
NA, with a warning naming it: scored on the rows it did predict
it would be compared on fewer, usually easier, rows than its rivals. A
factor with a level found in one spatial block only is the usual case,
and cannot be selected.
score_holdout is NA unless select_on = "split", and
then the selected set's metric when fitted on the selection half
and predicted on the estimation half (\(R^2\) against the selection
half's mean, the out-of-sample convention); NA when nothing was
selected or the prediction failed. split is NULL or a
list with selection and estimation, integer row positions
in train_sf as passed; rows the completeness filter above dropped
are in neither.
params records metric, method, k,
tol, seed, auto_range, select_on,
n_candidates, estimated_fits and n_scored.
Details
The inner folds must be spatial, and that is the entire point.
Nested selection is only worth doing if the inner loop is blocked the same
way the outer one is. Random inner folds inside blocked outer folds select
variables that look predictive only because nearby points leak between
train and test. The outer loop then reports honest-looking numbers for a
dishonestly chosen feature set, which is worse than not selecting at all,
because the dishonesty is now hidden behind a defensible-looking validation.
method therefore defaults to "block_kfold" and logs a loud
caution if set to "random_kfold" (a deliberate choice, so it is not
raised as an R warning).
Call this inside the fit_fn you pass to cv_spatial().
.cv_fit_one_fold() calls fit_fn(train_sf) on the training
slice only, so anything done inside it is automatically nested and
leak-free; no extra plumbing is needed. The cost grows fast: a sweep over
p candidates costs roughly p^2 / 2 * k model fits, and nesting
that inside n outer leave-one-out folds multiplies it by n.
max_fits guards against that.
The score is not a performance estimate
$score is the cross-validated metric of the winning set at
the final step: the best of every candidate set the sweep scored. That
is the number the selection optimised, and a number optimised over many
candidates is optimistically biased by construction: Cawley and Talbot
(2010) show the bias can exceed the genuine differences between the models
being compared. Quote it as the selection criterion, not as the
performance of the selected model. An honest performance estimate needs
data the selection never saw. Two ways to get one: run this function
inside the fit_fn of cv_spatial(), so the outer
folds score a model whose predictors were chosen on the inner ones alone;
or pass select_on = "split", which selects on one spatial half of
train_sf and returns the selected set's score on the other as
score_holdout.
References
Cawley, G. C. and Talbot, N. L. C. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11, 2079–2107. https://jmlr.org/papers/v11/cawley10a.html
See also
Other cross-validation:
area_of_applicability(),
cv_bayes(),
cv_block_size_sweep(),
cv_gwr(),
cv_rf(),
cv_spatial(),
estimate_sac_range(),
fold_separation(),
gwr_model_selection(),
make_folds(),
sac_nugget()
Examples
if (requireNamespace("GWmodel", quietly = TRUE) &&
requireNamespace("sp", quietly = TRUE)) {
library(sf)
set.seed(1)
n <- 120
pts <- st_as_sf(
data.frame(x = 5e5 + runif(n, 0, 1000), y = 5e6 + runif(n, 0, 1000),
a = rnorm(n), b = rnorm(n), noise = rnorm(n)),
coords = c("x", "y"), crs = 32632
)
pts$resp <- 3 * pts$a + 2 * pts$b + rnorm(n, 0, 0.3)
fit_fn <- function(tr, vars) fit_gwr_model(tr, "resp", vars, bandwidth = 30)
sel <- select_features_forward(pts, "resp", c("a", "b", "noise"), fit_fn,
k = 3, quiet = TRUE)
sel$selected
}
#> [1] "a" "b"