Skip to contents

featR has fourteen selection functions and two dimensionality-reduction functions. This page helps you pick one. It gives a decision guide, the trade-offs between the method families, and a reference table of what each function needs and accepts.

Start with the question

Different methods answer different questions. Pick the question first:

Question Function
Which columns are constant, near-constant, or mostly missing? fs_unsupervised()
Which columns duplicate each other? fs_correlation()
Which numeric columns are individually associated with the outcome? fs_supervised()
Which categorical columns are associated with a categorical outcome? fs_chi(), fs_infogain()
How much does each column reduce uncertainty about the outcome, whatever its type? fs_infogain()
Which predictors survive in a sparse linear model? fs_lasso(), fs_elastic()
Which predictors carry any information, including redundant ones? fs_boruta()
How does each predictor rank in a random forest, and how good is that forest on held-out rows? fs_randomforest()
Which predictors does a model with automatic non-linearities and interactions use? fs_mars()
What is the smallest subset a particular model needs? fs_recursivefeature(), fs_svm()
Which terms does AIC keep in a linear regression? fs_stepwise()
Which predictor subset has the best out-of-sample predictive fit under a Bayesian model? fs_bayes()
Can I summarize these columns in a few directions? fs_pca(), fs_svd()

Filters, embedded methods, and wrappers

Filters (fs_unsupervised(), fs_supervised(), fs_chi(), fs_infogain(), fs_correlation()) score each feature without fitting a predictive model. They are fast and scale to thousands of columns. The cost is that they are univariate:

  • Two copies of the same signal both score highly.
  • A feature that matters only in combination with another scores low.

Use filters for a first cut, then refine with something else.

Embedded methods (fs_lasso(), fs_elastic(), fs_randomforest(), fs_mars()) fit one model and read the selection off its structure: non-zero coefficients, importance, or retained terms. They account for the other predictors, but the answer belongs to that model family.

Wrappers (fs_recursivefeature(), fs_svm(), fs_boruta(), fs_stepwise(), fs_bayes()) refit a model over many candidate subsets. They are the most expensive and the most tailored to one model. They are also the most prone to overfitting the selection itself, so validate their output on held-out data.

A sensible default pipeline

For a new tabular problem:

  1. Clean. Drop constant or mostly missing columns with fs_unsupervised(). This uses no outcome, so it is safe to run before any split.
  2. Split. Hold out a test set now, before anything looks at the outcome.
  3. Deduplicate. Prune near-duplicate columns with fs_correlation() on the training rows.
  4. Select. Run the method that matches your final model on the training rows: fs_lasso() for a linear model, fs_randomforest() or fs_boruta() for trees, fs_svm() for an SVM.
  5. Validate. Evaluate the final model on the test rows.

The validation and leakage article shows this pipeline end to end.

Reference: what each function accepts

Every selection function takes data first. The functions with an outcome take target, the name of the outcome column, second. The housekeeping arguments differ between functions, so check the table before you write a loop over several methods:

Function target Outcome types seed Parallel Engine packages (Suggests)
fs_unsupervised() — none — — none
fs_correlation() — none yes parallel, n_cores (point-biserial only) polycor (polychoric), foreach + doParallel
fs_supervised() yes numeric, factor — — none
fs_chi() yes categorical yes parallel, n_cores furrr + future when parallel
fs_infogain() yes any (discretized) — — none
fs_lasso() yes numeric yes parallel, n_cores glmnet, Matrix
fs_elastic() yes numeric, factor yes n_cores caret, glmnet, Matrix
fs_randomforest() yes set by task yes n_cores randomForest, caret, pROC (AUC)
fs_mars() yes numeric, factor yes n_cores caret, earth, pROC/PRROC (AUC)
fs_recursivefeature() yes numeric, factor yes parallel caret, randomForest (default functions), e1071
fs_svm() yes set by task yes n_cores caret, kernlab, e1071, randomForest (rf_rfe)
fs_boruta() yes numeric, factor yes — Boruta
fs_stepwise() yes numeric — (deterministic) — MASS
fs_bayes() yes any brms family yes (subset sampling) parallel_combinations, n_cores brms, loo, a Stan toolchain
fs_pca() — none — — ggplot2 (plot), bigstatsr (large data)
fs_svd() — (x) none — — RSpectra (approximate solver)

Every function takes verbose. All of them are sequential by default. fs_chi(), fs_correlation(), and fs_lasso() default to n_cores = 2, but that value is used only when you also set parallel = TRUE.

If an engine package is missing, the function stops and prints the exact install.packages() call you need. Nothing fails deep inside a model fit.