Performs supervised, univariate, filter-based feature selection: every
column of data except target is scored against the target, and
features are selected or dropped by comparing that score to threshold.
Arguments
- data
A data.frame, data.table, or matrix containing
targetand the candidate feature columns. Every column other thantargetmust be numeric. The input is copied, never modified in place.- target
Character. Name of the target column in
data. It is removed from the candidate features and used as the scoring target. The target column itself may be numeric, factor, character, or logical; any other type is an error.- method
One of
"auto"(default),"correlation","anova".- threshold
Non-negative, finite numeric scalar threshold applied to the feature scores (not to the target directly). Default 0. Note that the two methods put scores on different scales:
[0, 1]for"correlation"and[0, Inf]for"anova".- direction
One of
"above"(default),"below"; compares scores tothreshold.- action
One of
"keep"(default),"remove"; determines whether features meeting the condition are retained or dropped.- include_equal
Logical; if TRUE, comparisons are inclusive (greater/less than or equal) instead of strict. Default FALSE.
- na_rm
Logical; if TRUE (the default), rows with NA in a feature or in the target are dropped when computing that feature's score. If FALSE, an NA anywhere in a feature or the target makes that feature's score undefined.
- output
One of
"result"(default),"matrix","dt","data.frame","mask","indices","names","list". Throughout, "candidate features" means the columns ofdataother thantarget, in their original order."result"(default): anfs_resultobject (see Value)."matrix": numeric matrix of the selected features."dt": data.table of the selected features."data.frame": data.frame of the selected features."mask": logical vector, one element per candidate feature, named after the candidate features; TRUE marks a selected column."indices": integer vector of the selected column indices, indexing the candidate features (that is,datawithouttarget), named after the selected columns."names": character vector of the selected column names."list": list with componentsfiltered(matrix),mask,indices,names,scores(all as above), andmeta, a list recordingmethod_arg(the method as requested),method_used(the method after"auto"was resolved),threshold,direction,action,include_equal,na_rm,n_input_cols, andn_kept_cols.
- verbose
Logical; emit progress messages. Default FALSE.
Value
With the default output = "result", an object of class
fs_result with elements:
selected: character vector of selected feature names.scores: named numeric vector of per-feature scores, in column order,NAwhere the score is undefined.method:"supervised_"followed by the resolved scoring method, for example"supervised_correlation".task:"regression"for a numeric target,"classification"for a factor target,NAotherwise.model:NULL.details: a list holding, in this order,mask(the logical keep-mask over the candidate features),indices(the selected column indices, named after the selected columns),filtered(the selected columns as a data.table),threshold,direction,action, andn_features(the number of candidate features, that isncol(data) - 1).call: the matched call.
Any other output returns that shape instead, exactly as documented
above. When no feature meets the selection criteria, a warning is issued
and the tabular shapes come back empty: "matrix" and
"data.frame" have zero columns and keep the input row count, while
the "dt" shape (and details$filtered) is only guaranteed to have
zero columns – data.table represents a zero-column table as having zero
rows, so its row count is not preserved.
Details
This is one of the cheapest supervised screens in featR: no predictive
model is fitted, each feature is scored on its own, and the cost is linear
in the number of columns, so it scales to wide data. The flip side is that
scoring is strictly univariate: it cannot see interactions between
features, and it will happily keep a whole group of near-duplicate columns
that all correlate with the target. Use fs_correlation() to prune that
redundancy afterwards, or a wrapper such as fs_recursivefeature() or
fs_svm() when the joint contribution is what matters.
Supported methods:
"correlation": Absolute Pearson correlation (numeric target), so scores lie in[0, 1]. A constant feature (or target) has no defined correlation and scoresNA; the check is relative to the column's magnitude, so features measured on a very small scale are scored normally."anova": One-way ANOVA F-statistic (categorical target), so scores lie in[0, Inf]: a feature that separates the classes perfectly (no spread within any class) scoresInf, and a constant feature scoresNA."auto": Chooses"correlation"for a numeric target and"anova"for a categorical one. Because the two score scales differ, a message reports the resolved method whenverbose = TRUE.
Features whose score is undefined (NA) are never selected, under
both action = "keep" and action = "remove"; a warning
reports how many such features were excluded.
Columns are subset by integer index, never by name, so duplicated column names cannot select the wrong columns.
Examples
df <- data.frame(
strong = c(1, 2, 3, 4),
mirror = c(4, 3, 2, 1),
weak = c(1, 0, 1, 0),
y = c(1, 2, 3, 4)
)
# Default: an fs_result
res <- fs_supervised(df, target = "y", method = "correlation",
threshold = 0.5)
res$selected
#> [1] "strong" "mirror"
res$scores
#> strong mirror weak
#> 1.0000000 1.0000000 0.4472136
res$details$filtered
#> strong mirror
#> <num> <num>
#> 1: 1 4
#> 2: 2 3
#> 3: 3 2
#> 4: 4 1
# The classic shapes are still available
fs_supervised(df, target = "y", method = "correlation", threshold = 0.5,
output = "names")
#> [1] "strong" "mirror"
# ANOVA against a factor target
df_fac <- data.frame(
wide = c(1, 2, 10, 11),
mild = c(1, 3, 2, 4),
grp = factor(c("a", "a", "b", "b"))
)
fs_supervised(df_fac, target = "grp", method = "anova", threshold = 1,
output = "matrix")
#> wide
#> [1,] 1
#> [2,] 2
#> [3,] 10
#> [4,] 11