Fits a lasso (or elastic-net) model with glmnet::cv.glmnet() and reports
which predictors survive at lambda.min (default) or lambda.1se. Numeric
outcomes only (gaussian family).
Arguments
- data
A data.frame or data.table holding the target and the candidate predictors. A numeric matrix with column names is also accepted; non-numeric matrices are not (supply a data.frame so factors and characters can be expanded by
model.matrix()).- target
Single string naming the outcome column in
data. It must be numeric and free of NA/NaN/Inf.- alpha
Numeric in (0, 1]; default 1 (lasso). Use values in (0, 1) for elastic-net.
alpha = 0(pure ridge) is rejected: it never sets a coefficient to exactly zero, so it cannot select.- nfolds
Integer >= 3; default 5. Number of cross-validation folds passed to
glmnet::cv.glmnet(), which rejects anything smaller than 3. Whencustom_foldsis supplied those fold IDs define the folds instead, andnfoldsonly bounds which IDs are valid (socustom_foldsmust also define at least 3 folds).- standardize
Logical; default TRUE. Passed to
glmnet::cv.glmnet()and also controls the scalescoresis ranked on (see Details).- custom_folds
Optional integer vector of fold IDs (one per row of
data, covering 1..nfoldswith no empty folds); default NULL.- impute
How to handle missing predictor values:
"none"(default) errors and names the offending columns,"mean"imputes column means computed on the full data and warns about the leakage.- return_model
Logical; keep the fitted cv.glmnet object in the result (and pass
keep = TRUEtocv.glmnet()); default FALSE.- seed
Optional whole number for reproducibility, applied locally and restored on exit; default NULL (never seeds by default).
- verbose
Logical; default FALSE. When TRUE, reports the parallel backend status and how many design-matrix columns were selected.
- parallel
Logical; default FALSE. When TRUE, cross-validation runs on
n_coresworkers (requires the 'doParallel' and 'foreach' packages); ifn_coresresolves to one worker it stays sequential and says so underverbose = TRUE.- n_cores
Integer >= 1; number of workers used only when
parallel = TRUE. Default 2. Values above the detected core count are capped.- lambda
Which cross-validated penalty to select at:
"min"(default,lambda.min) or"1se"(lambda.1se, sparser). See Details.
Value
An fs_result object with:
- selected
Design-matrix columns whose raw coefficient at the chosen lambda (
lambda.minby default) is non-zero, ordered by decreasingscoresmagnitude. (Selection reads the raw coefficients, so a constant column that glmnet nonetheless gave a non-zero coefficient is reported even though its standardized score is 0.)- scores
data.frame with columns Variable, Coefficient and AbsCoefficient, one row per design-matrix column (including those shrunk to zero), ordered by decreasing AbsCoefficient, on the standardized scale when
standardize = TRUE(see Details).- method
"lasso", regardless ofalpha.- task
"regression"(gaussian family only).- model
The fitted cv.glmnet object when
return_model = TRUE, else NULL.- details
List of
lambda_min(lambda minimizing CV error),lambda_1se(largest lambda within 1 SE of the minimum),lambda_used("lambda.min"or"lambda.1se", the one selection and all coefficients are read at),coefficients(the same three-column table asscoresbut always on the raw coefficient scale) andn_features(number of design-matrix columns considered).- call
The matched call.
Details
Use this when the question is "which predictors keep a non-zero coefficient
under a cross-validated L1 penalty (or, for alpha < 1, an elastic-net
penalty)?". By default selection happens at lambda.min, the penalty that
minimizes cross-validated error. lambda.min is tuned for prediction and
is known to over-select (it tends to keep noise variables alongside the
true ones); lambda = "1se" selects at lambda.1se, the largest penalty
whose CV error is within one standard error of the minimum, which gives a
sparser, more conservative set. Both values are reported in details
whichever one is used. The main caveat is that lasso tends to
keep one member of a group of strongly correlated predictors and zero out
the rest, so an absent feature is not evidence that it is unrelated to the
outcome.
The design matrix is built internally from every column of data except
target, via stats::model.frame(na.action = stats::na.pass) followed by
stats::model.matrix(): factors and characters are expanded to dummies and
rows carrying NAs are preserved rather than silently dropped. Non-syntactic
column names are reported as the caller wrote them (without the backticks
model.matrix() adds). The design matrix must have at least two columns,
because glmnet cannot fit a single-column x. The matrix
carries no intercept column of its own (glmnet fits its own intercept,
which is dropped from the reported coefficients), so selected, scores
and details$coefficients name design-matrix columns – for a factor
predictor, its expanded dummy columns rather than the original column.
scores ranks features on the STANDARDIZED coefficient scale when
standardize = TRUE: each coefficient is multiplied by the standard
deviation of its design-matrix column, which makes the ranking independent
of the units the predictors happen to be measured in. glmnet standardizes
internally for fitting but reports coefficients back on the input scale, so
those raw coefficients are kept in details$coefficients. With
standardize = FALSE, scores is the raw table and the two are identical.
Raw coefficient magnitude is a scale-dependent notion of importance: a
predictor measured in small units earns a large coefficient for the same
effect. That applies to details$coefficients always, and to scores when
standardize = FALSE, so compare those numbers across predictors only when
the predictors share a scale.
Missing predictor values are an error under the default
impute = "none", because imputing before cross-validation lets the folds
see each other. impute = "mean" fills them with column means computed on
the whole data set – not on the training part of each fold – and warns
that this leaks. Columns that are entirely NA are an error either way.
Examples
# \donttest{
if (requireNamespace("glmnet", quietly = TRUE) &&
requireNamespace("Matrix", quietly = TRUE)) {
n <- 100
df <- data.frame(
x1 = rnorm(n),
x2 = rnorm(n),
cat = sample(letters[1:3], n, TRUE)
)
df$y <- 2 * df$x1 - 3 * df$x2 + rnorm(n)
res <- fs_lasso(df, "y", seed = 123)
selected(res)
head(res$scores)
}
#> Variable Coefficient AbsCoefficient
#> 1 x2 -2.87232033 2.87232033
#> 2 x1 2.16440132 2.16440132
#> 3 catc -0.09940822 0.09940822
#> 4 cata 0.08988327 0.08988327
#> 5 catb 0.00000000 0.00000000
# }