featR has fourteen selection functions and two dimensionality-reduction functions. This page helps you pick one. It gives a decision guide, the trade-offs between the method families, and a reference table of what each function needs and accepts.
Start with the question
Different methods answer different questions. Pick the question first:
| Question | Function |
|---|---|
| Which columns are constant, near-constant, or mostly missing? | fs_unsupervised() |
| Which columns duplicate each other? | fs_correlation() |
| Which numeric columns are individually associated with the outcome? | fs_supervised() |
| Which categorical columns are associated with a categorical outcome? |
fs_chi(), fs_infogain()
|
| How much does each column reduce uncertainty about the outcome, whatever its type? | fs_infogain() |
| Which predictors survive in a sparse linear model? |
fs_lasso(), fs_elastic()
|
| Which predictors carry any information, including redundant ones? | fs_boruta() |
| How does each predictor rank in a random forest, and how good is that forest on held-out rows? | fs_randomforest() |
| Which predictors does a model with automatic non-linearities and interactions use? | fs_mars() |
| What is the smallest subset a particular model needs? |
fs_recursivefeature(), fs_svm()
|
| Which terms does AIC keep in a linear regression? | fs_stepwise() |
| Which predictor subset has the best out-of-sample predictive fit under a Bayesian model? | fs_bayes() |
| Can I summarize these columns in a few directions? |
fs_pca(), fs_svd()
|
Filters, embedded methods, and wrappers
Filters (fs_unsupervised(),
fs_supervised(), fs_chi(),
fs_infogain(), fs_correlation()) score each
feature without fitting a predictive model. They are fast and scale to
thousands of columns. The cost is that they are
univariate:
- Two copies of the same signal both score highly.
- A feature that matters only in combination with another scores low.
Use filters for a first cut, then refine with something else.
Embedded methods (fs_lasso(),
fs_elastic(), fs_randomforest(),
fs_mars()) fit one model and read the selection off its
structure: non-zero coefficients, importance, or retained terms. They
account for the other predictors, but the answer belongs to that model
family.
Wrappers (fs_recursivefeature(),
fs_svm(), fs_boruta(),
fs_stepwise(), fs_bayes()) refit a model over
many candidate subsets. They are the most expensive and the most
tailored to one model. They are also the most prone to overfitting the
selection itself, so validate their output on held-out data.
A sensible default pipeline
For a new tabular problem:
-
Clean. Drop constant or mostly missing columns with
fs_unsupervised(). This uses no outcome, so it is safe to run before any split. - Split. Hold out a test set now, before anything looks at the outcome.
-
Deduplicate. Prune near-duplicate columns with
fs_correlation()on the training rows. -
Select. Run the method that matches your final
model on the training rows:
fs_lasso()for a linear model,fs_randomforest()orfs_boruta()for trees,fs_svm()for an SVM. - Validate. Evaluate the final model on the test rows.
The validation and leakage article shows this pipeline end to end.
Reference: what each function accepts
Every selection function takes data first. The functions
with an outcome take target, the name of
the outcome column, second. The housekeeping arguments differ between
functions, so check the table before you write a loop over several
methods:
| Function | target |
Outcome types | seed |
Parallel | Engine packages (Suggests) |
|---|---|---|---|---|---|
fs_unsupervised() |
— | none | — | — | none |
fs_correlation() |
— | none | yes |
parallel, n_cores (point-biserial
only) |
polycor (polychoric), foreach + doParallel |
fs_supervised() |
yes | numeric, factor | — | — | none |
fs_chi() |
yes | categorical | yes |
parallel, n_cores
|
furrr + future when parallel |
fs_infogain() |
yes | any (discretized) | — | — | none |
fs_lasso() |
yes | numeric | yes |
parallel, n_cores
|
glmnet, Matrix |
fs_elastic() |
yes | numeric, factor | yes | n_cores |
caret, glmnet, Matrix |
fs_randomforest() |
yes | set by task
|
yes | n_cores |
randomForest, caret, pROC (AUC) |
fs_mars() |
yes | numeric, factor | yes | n_cores |
caret, earth, pROC/PRROC (AUC) |
fs_recursivefeature() |
yes | numeric, factor | yes | parallel |
caret, randomForest (default functions), e1071 |
fs_svm() |
yes | set by task
|
yes | n_cores |
caret, kernlab, e1071, randomForest (rf_rfe) |
fs_boruta() |
yes | numeric, factor | yes | — | Boruta |
fs_stepwise() |
yes | numeric | — (deterministic) | — | MASS |
fs_bayes() |
yes | any brms family | yes (subset sampling) |
parallel_combinations, n_cores
|
brms, loo, a Stan toolchain |
fs_pca() |
— | none | — | — | ggplot2 (plot), bigstatsr (large data) |
fs_svd() |
— (x) |
none | — | — | RSpectra (approximate solver) |
Every function takes verbose. All of them are sequential
by default. fs_chi(), fs_correlation(), and
fs_lasso() default to n_cores = 2, but that
value is used only when you also set parallel = TRUE.
If an engine package is missing, the function stops and prints the
exact install.packages() call you need. Nothing fails deep
inside a model fit.