Performs unsupervised, univariate, filter-based feature selection by scoring each feature using a chosen unsupervised criterion and selecting or dropping features based on a threshold on that score.
Usage
fs_unsupervised(
data,
method = c("variance", "mad", "iqr", "range", "missing_prop", "n_unique"),
threshold = 0,
direction = c("above", "below"),
action = c("keep", "remove"),
include_equal = FALSE,
na_rm = TRUE,
output = c("result", "matrix", "dt", "data.frame", "mask", "indices", "names", "list"),
verbose = FALSE
)Arguments
- data
A data.frame, data.table, or matrix; all columns must be numeric. Every column is a candidate feature. The input is copied, never modified in place.
- method
One of
"variance"(default),"mad","iqr","range","missing_prop","n_unique".- threshold
Non-negative, finite numeric scalar threshold applied to the feature scores. Default 0. The scales differ by method (a variance is in squared units,
"missing_prop"is in[0, 1],"n_unique"is a count), so a threshold is not portable between methods.- direction
One of
"above"(default),"below"; compares scores tothreshold.- action
One of
"keep"(default),"remove"; determines whether features meeting the condition are retained or dropped.- include_equal
Logical; if TRUE, comparisons are inclusive (greater/less than or equal) instead of strict. Default FALSE.
- na_rm
Logical; if TRUE (the default), remove NAs when computing scores. Has no effect on
"missing_prop"and"n_unique", which always account for NAs the same way.- output
One of
"result"(default),"matrix","dt","data.frame","mask","indices","names","list"."result"(default): anfs_resultobject (see Value)."matrix": numeric matrix of the selected features."dt": data.table of the selected features."data.frame": data.frame of the selected features."mask": logical vector of lengthncol(data), named after the columns ofdata; TRUE marks a selected column."indices": integer vector of the selected column indices, named after the selected columns."names": character vector of the selected column names."list": list with componentsfiltered(matrix),mask,indices,names,scores(all as above), andmeta, a list recordingmethod,threshold,direction,action,include_equal,na_rm,n_input_cols, andn_kept_cols.
- verbose
Logical; emit progress messages. Default FALSE.
Value
With the default output = "result", an object of class
fs_result with elements:
selected: character vector of selected feature names.scores: named numeric vector of per-feature scores, in column order,NAwhere the score is undefined.method:"unsupervised_"followed by the scoring method, for example"unsupervised_variance".task:NA_character_(unsupervised selection has no task).model:NULL.details: a list holding, in this order,mask(the logical keep-mask over the candidate features),indices(the selected column indices, named after the selected columns),filtered(the selected columns as a data.table),threshold,direction,action, andn_features(the number of candidate features, that isncol(data)).call: the matched call.
Any other output returns that shape instead, exactly as documented
above. When no feature meets the selection criteria, a warning is issued
and the tabular shapes come back empty: "matrix" and
"data.frame" have zero columns and keep the input row count, while
the "dt" shape (and details$filtered) is only guaranteed to have
zero columns – data.table represents a zero-column table as having zero
rows, so its row count is not preserved.
Details
No target is involved, so this is the right tool for the first cleaning
pass – dropping constant or near-constant columns, columns that are mostly
missing, or columns with too few distinct values – and it is safe to run
before a train/test split, since nothing about the outcome informs it. It
says nothing about whether a feature is useful: a high-variance
column can be pure noise, and a low-variance one can be the best predictor
you have. Use fs_supervised() or a model-based method for that judgment.
Supported methods:
"variance": Sample variance (denominator n - 1)."mad": Median absolute deviation, computed withstats::mad()'s default consistency constant 1.4826 (normal-consistent). Scores are therefore on the standard-deviation (sigma) estimate scale, not the raw median-absolute-deviation scale, and thresholds should be chosen accordingly."iqr": Interquartile range."range": Max - Min."missing_prop": Proportion of missing values, in[0, 1]. Note that with the defaultthreshold = 0,direction = "above", andaction = "keep", this KEEPS the features with the most missing values, which is almost never the intent; a warning suggestsaction = "remove"ordirection = "below". The warning fires only when all three ofthreshold,direction, andactionare left at their defaults, so passing any one of them explicitly opts out of the advisory."n_unique": Number of unique non-NA values.
na_rm affects only the methods that summarize the observed values
("variance", "mad", "iqr", "range");
"missing_prop" and "n_unique" always look at the whole
column and treat NA as NA.
Features whose score is undefined (NA) are never selected, under
both action = "keep" and action = "remove"; a warning
reports how many such features were excluded.
Columns are subset by integer index, never by name, so duplicated column names cannot select the wrong columns.
Examples
df <- data.frame(
spread = c(1, 2, 3, 4, 100),
flat = c(2, 2, 2, 2, 2),
gappy = c(1, NA, 3, NA, 5)
)
# Default: an fs_result
res <- fs_unsupervised(df, method = "variance", threshold = 0.5)
res$selected
#> [1] "spread" "gappy"
res$scores
#> spread flat gappy
#> 1902.5 0.0 4.0
res$details$filtered
#> spread gappy
#> <num> <num>
#> 1: 1 1
#> 2: 2 NA
#> 3: 3 3
#> 4: 4 NA
#> 5: 100 5
# The classic shapes are still available
fs_unsupervised(df, method = "variance", threshold = 0.5,
output = "matrix")
#> spread gappy
#> [1,] 1 1
#> [2,] 2 NA
#> [3,] 3 3
#> [4,] 4 NA
#> [5,] 100 5
# Remove features with missing proportion >= 0.2
fs_unsupervised(df, method = "missing_prop", threshold = 0.2,
direction = "above", action = "remove",
include_equal = TRUE, output = "names")
#> [1] "spread" "flat"