Answers "which subset of these columns does AIC keep in a linear model?"
Uses MASS::stepAIC() to perform forward, backward, or both-direction
stepwise selection on a linear regression of target against all other
columns of data. For direction = "forward" and "both" a proper null
model and scope are set up so that forward moves are possible.
Usage
fs_stepwise(
data,
target,
direction = c("both", "backward", "forward"),
verbose = FALSE,
...
)Arguments
- data
A data.frame (or data.table) with at least two columns: the numeric target and at least one candidate predictor. Every column other than
targetis offered to the search.- target
Single string naming the numeric target column of
data. Unquoted symbols and column indices are not accepted. Non-syntactic names are backticked into the formula, so they work.- direction
Direction of the search:
"both"(the default),"backward", or"forward", matched withmatch.arg().- verbose
Logical. If
TRUE, emits progress messages and enables thestepAIC()trace output on the console. DefaultFALSE.- ...
Additional arguments passed to
MASS::stepAIC(), for examplek(usek = log(nrow(data))for BIC),steps, orscope(which then replaces the default scope described above).traceis the one exception: it is controlled byverbose, and a user-suppliedtraceis dropped with a warning. See Details for what else...quietly absorbs.
Value
An object of class fs_result with:
- selected
Character vector of the selected predictor terms (excluding the intercept).
- scores
Named numeric vector of absolute t statistics from the final model's coefficient table, excluding the intercept. Scores are per coefficient, so a selected factor term appears here under its dummy names (e.g.
grpb,grpcfor the termgrp), not under the term label reported inselected. Caveat: these statistics (and the p-values indetails$coefficients) are computed after selection on the same data, so they are optimistically biased and are not valid for inference.- method
"stepwise".
- task
"regression";
fs_stepwise()fits linear models only.- model
The
lmselected bystepAIC(). It carries a copy of the complete-case data in a private environment (see Details).- details
A list with
final_model(the samelm),coefficients(thesummary()coefficient matrix, same caveat asscores),selected_terms(the same term labels asselected),direction(the search direction actually used),n_features(the number of candidate predictors offered to the search, i.e.ncol(data) - 1) anddropped_na_rows(how many rows were removed for missing values,0Lwhen none were).- call
The matched call.
Details
This is the cheapest wrapper method in featR and the easiest to read: what
comes back is an ordinary lm that the usual generics work on. The price is
that the search is greedy, so it can walk past the AIC-best subset, and
unstable, so small changes in the data can change the retained set. On top
of that, the statistics it reports about its own answer are not valid; see
the post-selection inference caveat below.
Requires the suggested package 'MASS'. The function is
linear-regression-only: the target must be numeric, and every other column
of data is offered to the search as a candidate predictor. There is no
seed argument, because stepAIC() is deterministic: repeating a call on
the same data returns the same model, and the RNG is never touched.
Post-selection inference caveat. The reported scores (absolute
t statistics) and the p-values in details$coefficients are computed on
the same data that drove the search. They are optimistically biased – the
selective-inference problem – and must not be used for formal inference,
only as a rough ordering of the retained terms.
The fitted models reach their data through a private environment attached to
the model formula, so predict(), summary(), anova(), add1()/drop1()
and similar generics keep working on the returned model; nothing is assigned
to the global environment and nothing is written to disk. That environment
holds a copy of the complete-case data and is kept alive by the returned
model, so the result carries the data with it and is correspondingly large.
To refit the returned model with update() from another environment, pass
the data explicitly, e.g. update(model, . ~ . - x, data = my_data). Plain
update(model) does not work: update() re-evaluates the stored
call in the caller's environment, where the private data object is not
visible, so it fails with an object-not-found error.
Non-syntactic column names are backticked into the formula and reported
back without the backticks, so selected matches names(data).
... is forwarded to MASS::stepAIC() verbatim, which means it also
absorbs arguments this function does not have. The removed seed and
return_models are passed through and ignored rather than rejected, so a
call written against the older API still runs and quietly does nothing with
them; check your argument names against the list below.
Rows containing missing values in any column are dropped (with a warning)
before the search, because stepAIC() cannot compare models fitted on
differing row sets.
Examples
# \donttest{
if (requireNamespace("MASS", quietly = TRUE)) {
res <- fs_stepwise(mtcars, target = "mpg", direction = "both")
print(res$selected)
# |t| of the retained terms: a rough ordering, not valid inference
print(res$scores)
# the returned model is an ordinary lm, but update() needs 'data'
refit <- stats::update(res$model, . ~ . - wt, data = mtcars)
print(stats::formula(refit))
}
#> [1] "wt" "qsec" "am"
#> wt qsec am
#> 5.506882 4.246676 2.080819
#> mpg ~ qsec + am
#> <environment: 0x55b388988e80>
# }